AIニュース 2026-08-21
自動生成: 2026-08-21 10:40 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Introducing AI FuturesOpenAI
Introducing AI Futures, a new OpenAI blog exploring how transformativ…
-
Googleのオープンモデル「Gemma」、累計10億ダウンロード超 GitHubに公式ディレクトリ公開ITmedia AI+
Googleは、オープンモデル「Gemma」ファミリーの累計ダウンロード数が10億回を突破したと発表した。派生モデルは10万種を超え、公式…
-
macOS版ChatGPT、Appleの「メッセージ」と連携 会話検索や下書き、送信に対応ITmedia AI+
OpenAIは、macOS版ChatGPT向けにAppleの「メッセージ」アプリと連携するプラグインを公開した。CodexやChatGPT…
-
「チャピる」「ギュられる」って何? 今年流行った「就活用語」にAI関連ワード マイナビ調査ITmedia AI+
マイナビは、2027年卒の就職活動で流行した用語を発表した。米OpenAIのチャットAI「ChatGPT」を示す「チャッピー」に加え、「チ…
-
「孫さんはOpenAIだが、僕はAnthropic」 SBI北尾会長が語る「AI投資5億円→増収27億円」の勝算ITmedia AI+
SBIホールディングスが生成AI「Claude」を開発する米Anthropicとの全社提携を発表した。当面の最優先戦略にAI化を掲げ、社外…
-
Slack、AIとチームで協働する「Slack Code」を発表 ClaudeやDevinを専用チャネルで操作ITmedia AI+
Slackは、AIコーディングエージェントと協働するための新機能「Slack Code」を発表した。メンションで専用の「コードチャネル」が…
-
A third of web pages published since ChatGPT’s launch show signs of AI authorship, study findsTechCrunch AI
ChatGPT and other AI models are now authoring and editing much of the…
トピック別件数
- LLM/生成AI 129件
- 研究/論文 106件
- エージェント 74件
- 画像/動画生成 35件
- その他 15件
- ビジネス/資金調達 15件
- ロボティクス 15件
- ハードウェア/半導体 11件
- 規制/政策 1件
日本語メディア19件
ITmedia AI+ (日本語)
Googleのオープンモデル「Gemma」、累計10億ダウンロード超 GitHubに公式ディレクトリ公開
Googleは、オープンモデル「Gemma」ファミリーの累計ダウンロード数が10億回を突破したと発表した。派生モデルは10万種を超え、公式リポジトリ「Awesome Gemma」をGitHubで公開。宇宙空間での衛星データ解析や新規のがん治療経路の発見など、多様な活用事例を紹介…
macOS版ChatGPT、Appleの「メッセージ」と連携 会話検索や下書き、送信に対応
OpenAIは、macOS版ChatGPT向けにAppleの「メッセージ」アプリと連携するプラグインを公開した。CodexやChatGPT Workのチャット上で過去の会話検索や下書き作成、送信が可能になる。誤送信防止のため都度承認フローを備える。Appleシリコン搭載Mac向…
「Fable禁止」で仕事が止まったあの日々を振り返る 日本企業が取るべき「脱・単一モデル」戦略
米政府の輸出管理で「Claude Fable 5」の提供が突如停止し、業務が止まった実体験を基に、単一AIモデル依存の地政学リスクを指摘。オンプレミス化やマルチモデル統合基盤など、日本企業が取るべき分散戦略を整理する。
データをつなぎ、AI活用へ――オートデスクが示す設計/製造DXの未来像
オートデスクは「Design & Make Summit Japan 2026」を東京都内で開催した。本稿では、米Autodeskのビック・ベダンサム氏による基調講演から、AI時代の製造業に求められるデータ基盤と設計/製造の変革について紹介する。
業務標準化の手間を9割減 三菱UFJ銀行は生成AIに「業務知識」をどう教えた?
海外事務の標準化を進めている三菱UFJ銀行。熟練従業員に頼ってきた業務プロセスの精査を生成AIに置き換える過程で直面したのは、AI特有の誤情報や回答のバラつきだった。同行はこれをどう克服したのか。
Snowflakeが過去最高業績 CEOが「他社との差別化は容易になった」と語るワケ
競合が同じ方向に走り出した今こそ、差別化はむしろ容易になっている――。SnowflakeのCEOが年次カンファレンスでこう言い切った根拠はどこにあるのか。AIエージェント時代のSnowflakeの戦略に迫る。
「チャピる」「ギュられる」って何? 今年流行った「就活用語」にAI関連ワード マイナビ調査
マイナビは、2027年卒の就職活動で流行した用語を発表した。米OpenAIのチャットAI「ChatGPT」を示す「チャッピー」に加え、「チャピる」「ギュられる」といったAIに関する新語が登場した。
ChatGPTに「おすすめの○○は?」 実は答えが決まっているらしい:893rd Lap
ChatGPTに「おすすめの○○は?」と聞けば、いくつかのブランドや商品を教えてくれる。では、その候補はどうやって選ばれているのだろうか。どうやらChatGPTは、検索を始める前から「この分野ならこれ」と、ある程度の候補を持っているらしい。
「たった14人」の挑戦から7兆円の逆転劇へ ラピダス小池社長の「TSMCとは戦わない」2ナノ半導体の勝算
世界の半導体市場を台湾TSMCが席巻する中、7兆円規模の国家プロジェクトとして最先端「2ナノ」の量産化に挑むのがラピダスだ。同社はTSMCとの規模の勝負を避け、設計から前後工程を一棟で完結させる「RUMS」による多品種生産で勝負する。「たった14人」の同志でスタートした原点から…
「孫さんはOpenAIだが、僕はAnthropic」 SBI北尾会長が語る「AI投資5億円→増収27億円」の勝算
SBIホールディングスが生成AI「Claude」を開発する米Anthropicとの全社提携を発表した。当面の最優先戦略にAI化を掲げ、社外から専門人材を起用。SBI証券では顧客対応のAIエージェント開発に5億円を投資し、口座再活性化などを通じて年27億円の増収を見込む。「孫正義…
「Gemini Notebook」で利用者10倍 シニア社員をAIヘビーユーザーにした首都高の考え
「生成AIを何に使えばよいかわからない」という理由により、生成AIの活用が停滞してしまう企業は多い。安全を最優先する故に慎重な組織風土であった首都高速道路でも同様の課題を抱えていた。しかし同社では「Google Gemini」を起点としたある工夫により、劇的に活用状況を改善した…
Slack、AIとチームで協働する「Slack Code」を発表 ClaudeやDevinを専用チャネルで操作
Slackは、AIコーディングエージェントと協働するための新機能「Slack Code」を発表した。メンションで専用の「コードチャネル」が自動生成され、計画やコード差分、プレビューを確認しながら指示できる。ClaudeやDevinなど複数社のエージェントに対応し、人間の承認を経…
カルビーが挑むジャガイモ収量の限界――自社開発AIでサプライチェーン最適化
「ポテトチップス」や「じゃがりこ」など、カルビーの主力商品に欠かせないばれいしょには、収穫量の限界がある。後手の意思決定から脱却すべく、同社はAIを活用した全社最適シミュレーター「C-BOSS」を自社開発。いかにして現場定着の壁を越え、データに基づく攻めのサプライチェーンを構築…
GoogleはAI競争に負けたのか 「最強のAI」ではなく「AIの“電力網”」を選ぶ賭け
GoogleからAI研究の中心人物が相次いで去った。「Geminiは終わった」という見方に対し、「最先端ではなく、AIを社会全体に行き渡らせる“電力網”で勝つ賭けだ」という別の解釈もある。電気の歴史になぞらえながら整理する。
「ロボットのChatGPTモーメントが近づいている」 中国UnitreeのCEO、世界ロボット大会で言及
Unitreeのワン・シンシンCEOは世界ロボット大会で、ロボットの「ChatGPTモーメント」が近づいていると発言した。
日本精工が「国産人型ロボ」開発を後押し スタートアップのアトムと協力、アクチュエータの検証など
人型ロボットを開発するスタートアップのアトムは、日本精工(NSK)と国産人型ロボットの開発・実装に向け、戦略的パートナーシップに関する基本合意書を締結したと発表した。
エイベックス松浦会長、noteの“バズり記事”を「ほぼAI」で作成 「僕の60年分のデータを入れた」
エイベックスの松浦勝人会長は、8月13日から「note」に投稿している記事について、ほとんどAIで作成していたと自身のXアカウント(@maxmatsuuratwit)で明かした。
「AIで生産性向上」日本の従業員は57%、世界平均は81% 仕事の満足度でも大差──アクセンチュア調査
アクセンチュアの世界20カ国調査で、AIによる生産性向上を実感する日本の従業員は57%と、世界平均の81%を大きく下回った。仕事の満足度や成果の実感でも世界との差が開き、人材と組織の変革が課題となっている。
Stripe、AIモデルゲートウェイのOpenRouter買収 400以上のAIモデルを束ねる中立基盤は維持
Stripeは、AIモデルのゲートウェイを手掛けるOpenRouterを買収することで合意したと発表した。報道による買収額は約75億ドル。OpenRouterは単一APIで400超のモデル切り替えを可能にする。Stripeはトークンコスト最適化を強化し、OpenRouterは買…
海外メディア13件
TechCrunch AI (英語)
AI data startup Micro1 reaches $500M gross run rate amid AI training boom
Surging demand for AI training data is driving rapid growth for the startup and its rivals.
ChatGPT can now send texts for you with new Apple Messages plug-in
Ever wanted someone else to do your texting for you? ChatGPT is being offered up as an automated text scribe via a new Apple Messages integ…
OK, can we actually cool data centers with our pee?
Jason Kelce joked that people should cool data centers with their pee, rather than potable water -- but his suggestion is not completely lu…
Google gives publishers a new way to fight AI-driven traffic losses
Google is giving publishers a new button that lets readers make them a preferred source across Search, Discover, and Google News, potential…
Runlayer, Rippling drop lawsuits — but the brouhaha is still a cautionary tale for founders
Runlayer and Rippling have dropped their lawsuits. No money was paid. Rippling celebrated by releasing a competing product.
Linkdaze’s smart calendar is built to run a household, not just track a schedule
Linkdaze's smart digital calendar stands out for not putting its features behind a paywall, including an AI meal planner tool.
Grok keeps sending gibberish responses to users
Affected users told TechCrunch they were using Grok Lite, and noticed the issues as early as Wednesday morning.
A third of web pages published since ChatGPT’s launch show signs of AI authorship, study finds
ChatGPT and other AI models are now authoring and editing much of the new web.
Ramp launches its own AI model router, called Router
Ramp has launched its own AI model routing service, dubbed Router, that lets users and companies use and switch between various large langu…
Meta brings Pocket, an app that lets you vibe-code and share games, to US users
Meta is bringing Pocket, its experimental AI-powered app for creating and sharing interactive games, to users across the U.S. after quietly…
Inertia Enterprises finds a way to make its fusion fuel fast
Fusion power startup Inertia Enterprises reduced the fuel filling process from a week to just a few hours. It's one of 10 hurdles the compa…
Meta AI’s new Mac app wants you to talk to your apps
The company said that the dictation feature works across all apps, just like other tools such as Wispr Flow, Superwhisper, and Monologue.
Binance now lets AI agents trade, but keeping them in check is largely up to users
Binance's Agent OS works with tools such as ChatGPT, Claude Code, and Cursor.
公式ブログ1件
OpenAI (英語)
Introducing AI Futures
Introducing AI Futures, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual free…
論文267件
arXiv cs.AI (英語)
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions
This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and sh…
役職: 遷移の複雑さによるゲーム世界のプロファイリング
ゲーム ワールド モデリング (GWM) と強化学習 (RL) は、しばしば混同されます。研究論文では、宣言されたインターフェイス (有限履歴を持つピクセル/トークン/潜在) における基礎となる遷移予測問題がどれほど難しいかを定量化することがほとんどないためです。我々は、遷移複雑性プロファイル (TCP) を提案します。これは、環境 (またはゲームプレイ データセット) によって引き起こされる遷移カーネルを、(i) 固有の 1 ステップ分岐、(ii) 観測可能な場合の相互作用に起因する不確実性と対戦相手の影響、および (iii) 標準化されたプローブ曲線を介した時間的/空間的依存性の範囲によって特徴付ける、小規模で再現可能なメトリクスのセットです。 TCP は、明示的なリファレンス分布、プロトコルの確率論、およびバージョン管理された測定バジェット (サンプリング/リサンプリングおよび固定プローブ計算) とともにレポートされ、ベンチマーク間で比較可能な数値を可能にします。私たちは、一般的なゲーム ファミリと最新の「ニューラル ゲーム エンジン」ドメインがこの状況にどのように組み込まれるかを概説し、TCP が標準ベンチマーク メタデータとなり、GWM および RL 論文で必須の統計になることを求めます。
原文 (English)
Position: Profiling Game Worlds by Transition Complexity
Game world modeling (GWM) and reinforcement learning (RL) are often confounded because research papers rarely quantify how difficult the underlying transition prediction problem is at the declared interface (pixels/tokens/latents with finite history). We propose the Transition Complexity Profile (TCP): a small, reproducible set of metrics that characterizes an environment's (or gameplay dataset's) induced transition kernel by (i) intrinsic one-step branching, (ii) interaction-induced uncertainty and opponent influence when observable, and (iii) temporal/spatial dependency span via standardized probe curves. TCP is reported with an explicit reference distribution, protocol stochasticity, and a versioned measurement budget (sampling/resampling and fixed probe compute), enabling comparable numbers across benchmarks. We outline how common game families and modern "neural game engine" domains populate this landscape and call for TCP to become standard benchmark metadata and a required statistic in GWM and RL papers.
メンタルヘルスにおける大規模言語モデル: 応用、革新、倫理的課題の系統的レビュー
我々は、ソーシャルメディア分析、臨床会話エージェント、治療支援ツール、プロンプトエンジニアリング、マルチモーダル学習、倫理的配慮など、健康における大規模言語モデル(LLM)の応用に関するレビューを紹介します。私たちは、ソーシャルメディア投稿、電子医療記録、マルチモーダル入力などの多様なデータソースを利用した学際的な研究の結果を統合して、うつ病の早期発見、自殺リスク評価、個別化された治療サポート、心理教育コンテンツの生成を可能にします。私たちのレビューでは、解釈可能性と臨床関連性を強化する LLM モデルとアノテーション戦略の進歩に焦点を当てていますが、ドメイン適応のための迅速なエンジニアリングの重要な役割も強調しています。また、メンタルヘルスの診断とモニタリングを改善するために、テキスト、音声、センサーデータを統合する新たなマルチモーダル融合技術についても説明します。最後に、私たちは現在進行中の倫理的、社会技術的、規制上の課題に取り組み、現実世界のメンタルヘルスケアにおける LLM の安全、公平、責任ある導入を確保するためのフレームワークを提唱します。
原文 (English)
Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges
We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational agents, therapy support tools, prompt engineering, multimodal learning, and ethical considerations. We integrate findings from interdisciplinary studies utilizing diverse data sources such as social media posts, electronic medical records, and multimodal inputs to enable early detection of depression, suicide risk assessment, personalized therapy support, and psychoeducational content generation. Our review highlights advancements in LLM models and annotation strategies that enhance interpretability and clinical relevance, while we also emphasize the critical role of prompt engineering for domain adaptation. We also discuss emerging multimodal fusion techniques integrating text, speech, and sensor data for improved mental health diagnosis and monitoring. Finally, we address ongoing ethical, sociotechnical, and regulatory challenges, and advocate frameworks to ensure safe, equitable, and accountable deployment of LLMs in real-world mental health care.
立場: 行動システムには行動テストが必要
人工エージェント システムは、動的環境と相互作用し、目標を追求し、時間の経過とともに適応することによって、行動システムとして動作することがますます増えています。しかし、現在の評価方法は主にパフォーマンスの結果に焦点を当てており、パフォーマンスの結果を生み出す根本的な行動プロセスには焦点を当てていません。この論文は、AI エージェントは他の行動システムと同様に、その動作の系統的な観察、摂動、解釈を通じて評価される必要があると主張しています。私たちはこの立場を動機づけるために行動科学から得た教訓を活用し、厳密な行動テストの開発に焦点を当てた研究課題を提案します。これらには、アクションシーケンスから意思決定戦略を復元する方法、行動の違いを分離する環境を構築する方法、マルチエージェントシステムにおける新たなダイナミクスを調査する方法などが含まれます。これらの方向性を総合すると、AI の動作科学を開発するためのロードマップが提供されます。
原文 (English)
Position: Behavioral Systems Require Behavioral Tests
Artificial agentic systems increasingly operate as behavioral systems by interacting with dynamic environments, pursuing goals, and adapting over time. Yet, current evaluation methods largely focus on performance outcomes, not the underlying behavioral processes that produce them. This paper argues that AI agents must be evaluated like other behavioral systems: through systematic observation, perturbation, and interpretation of their actions. We draw on lessons from the behavioral sciences to motivate this position, and propose a research agenda focused on developing rigorous behavioral tests. These include methods for recovering decision strategies from action sequences, constructing environments that isolate behavioral differences, and probing emergent dynamics in multi-agent systems. Taken together, these directions offer a roadmap for developing a science of AI behavior.
立場: 現在のモデル カードはオープンウェイト ファウンデーション モデルの下流ガバナンスには不十分です
オープンウェイト基盤モデル (OWFM) の成長により、AI コミュニティは効果的な下流ガバナンスのための戦略を再評価するようになりました。モデル カードはモデル リポジトリの透明性アーティファクトとして広く採用されていますが、既存のフレームワークでは、OWFM によってもたらされる明確な安全性の課題について、下流の開発者やユーザーに適切に通知できないことがよくあります。この意見書は、Hugging Face でホストされている 500 のモデル カードを分析し、OWFM の効果的なガバナンスには、(i) モデル カード、(ii) 許容使用ポリシー (AUP)、および (iii) ライセンスの 3 つの補完的なコンポーネントを統合する多層アプローチが必要であると主張しています。この主張を動機付けるために、安全性が重要な情報を含むモデルカードの分析を通じて、モデルの伝統、アライメントの来歴、経験的に観察された動作など、既存の規制アプローチによって残された安全性のギャップを特定します。さらに、標準のオープンソース ライセンス (OSL) は OWFM にはあまり適しておらず、AUP の強制力が弱まる可能性があると主張します。これらの観察に基づいて、モデル カード、AUP、ライセンスを統合された安全成果物に進化させ、情報、規範、および法的側面を一貫して統合する、より包括的なガバナンス フレームワークを可能にするための方向性を概説します。
原文 (English)
Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models
The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate strategies for effective downstream governance. Although model cards have been widely adopted as transparency artifacts in model repositories, existing frameworks often fail to adequately inform downstream developers and users about the distinct safety challenges posed by OWFMs. This position paper analyzes 500 model cards hosted on Hugging Face and argues that effective governance of OWFMs requires a multi-layered approach integrating three complementary components: (i) model cards, (ii) acceptable use policies (AUPs), and (iii) licenses. To motivate this claim, we identify a safety gap left by existing regulatory approaches, including model heritage, alignment provenance, and empirically observed behaviors, through an analysis of model cards with safety-critical information. We further argue that standard open-source licenses (OSLs) are not well suited for OWFMs and may weaken the enforceability of AUPs. Building on these observations, we outline directions for evolving model cards, AUPs, and licenses into integrated safety artifacts to enable a more comprehensive governance framework that coherently integrates informational, normative, and legal dimensions.
飛行ログベースのドローンプロペラ健康状態モニタリングのための変形型人工年齢スコア意思決定支援プロトタイプ
ドローンのプロペラの故障は、その影響が単一の診断信号として現れるのではなく、複数の飛行ログ チャネルに分散される場合、安全性と信頼性のリスクを引き起こす可能性があります。この論文では、飛行ログに基づくドローン プロペラの健全性モニタリングのための、メタモーフィック人工年齢スコア (AAS) 意思決定支援プロトタイプを提案します。このフレームワークは、2024 年の DronePropA 公開データセットから選択された過去の実際の飛行ログを使用して、生の MATLAB 行列から 6 つの健康関連指標 (軌道追跡エラー、姿勢の不安定性、推力コマンドの負担、モーター コマンドの不均衡、ESC コマンドの不安定性、およびバッテリー レベルのストレス) を計算します。これらの指標は健全なベースラインに対して正規化され、候補者のスコアリング ポリシー、メタモーフィックな適切性関係、および冗長性が調整された AAS 定式化を通じて評価されます。この文脈では、AAS は実年齢の尺度としてではなく、構造的な政策の適切性と負担の尺度として使用されます。制御された遡及的評価は、同じ速度プロファイルと軌道の下で、1 つの健全なベースラインと 3 つの欠陥のあるプロペラ ケースを使用して実行されました。健康な症例は定期的なモニタリングに割り当てられました。重大度 1 のケースは ESC コマンドの不安定性が大半を占め、メンテナンス レビューに割り当てられました。重大度 2 のケースではモーター コマンドと ESC コマンドの負担が最大に達しましたが、重大度 3 のケースでは軌道追跡エラーが最大に達しました。どちらも強制検査を引き起こした。この結果は、プロペラ故障の影響がさまざまな運用チャネルを通じて現れる可能性があることを示しており、飛行後のメンテナンスの優先順位付けと自律システムの監視のための複数の指標による意思決定支援層の必要性を裏付けています。
原文 (English)
A Metamorphic Artificial Age Score Decision-Support Prototype for Flight-Log-Based Drone Propeller Health Monitoring
Drone propeller faults can create safety and reliability risks when their effects are distributed across multiple flight-log channels rather than appearing as a single diagnostic signal. This paper proposes a Metamorphic Artificial Age Score (AAS) decision-support prototype for flight-log-based drone propeller health monitoring. Using selected historical real flight logs from the 2024 DronePropA public dataset, the framework computes six health-related indicators from raw MATLAB matrices: trajectory tracking error, attitude instability, thrust-command burden, motor-command imbalance, ESC-command instability, and battery-level stress. These indicators are normalized relative to a healthy baseline and evaluated through candidate scoring policies, metamorphic adequacy relations, and a redundancy-adjusted AAS formulation. In this context, AAS is used as a structural policy-adequacy and burden measure rather than as a chronological age measure. A controlled retrospective evaluation was performed using one healthy baseline and three defective propeller cases under the same speed profile and trajectory. The healthy case was assigned to routine monitoring. The Severity 1 case was dominated by ESC-command instability and assigned to maintenance review. The Severity 2 case reached maximum motor-command and ESC-command burden, while the Severity 3 case reached maximum trajectory tracking error; both triggered mandatory inspection. The results show that propeller fault effects may appear through different operational channels, supporting the need for a multi-indicator decision-support layer for post-flight maintenance prioritization and autonomous-system oversight.
立場: マルチエージェント システムでは同時実行制御を優先する必要があります
LLM ベースのマルチエージェント システム (MAS) は、スケーラブルなコラボレーションを約束しますが、エージェントを追加すると信頼性が低下することがよくあります。この意見書では、MAS 障害の多くは基本的に同時実行制御の問題であると主張しています。エージェントは共有状態の読み取りと書き込みを同時に行い、LLM 推論ウィンドウが長いため、古い読み取り、更新の喪失、一貫性のない結果のリスクが増幅されます。一般に調整または通信の障害に起因すると考えられる障害モードは、従来の同時実行異常に直接マッピングできます。私たちは、MAS フレームワークが明示的な同時実行制御メカニズム (競合検出、分離保証、共有リソースへの構造化アクセス) を通じてこれらの障害に対処する必要があると主張します。同時実行制御は、後付けの考えではなく、最優先の設計上の考慮事項である必要があります。
原文 (English)
Position: Multi-Agent Systems Should Prioritize Concurrency Control
LLM-based multi-agent systems (MAS) promise scalable collaboration, yet adding agents often reduces reliability. This position paper argues that many MAS failures are fundamentally concurrency control problems: agents concurrently read and write shared state, and long LLM inference windows amplify the risk of stale reads, lost updates, and inconsistent outcomes. Failure modes commonly attributed to coordination or communication breakdowns can be mapped directly onto classical concurrency anomalies. We contend that MAS frameworks should address these failures through explicit concurrency control mechanisms: conflict detection, isolation guarantees, and structured access to shared resources. Concurrency control should be a first-class design concern, not an afterthought.
FinSkillBench: 投資管理のための AI エージェントとドメイン スキルの評価
投資管理は一か八かの分野であり、エージェント AI システムはもっともらしいテキストを生成する以上のことを行う必要があります。特定の時点のデータを取得し、正しい計算入力を組み立て、特殊なメソッドを呼び出し、監査可能な構造化された出力を生成する必要があります。 FinSkillBench は、言語モデル エージェントが財務ドメインのスキルを効果的に使用して投資管理タスクを解決できるかどうかを測定するように設計された評価スイートです。このベンチマークは、ポートフォリオ構築、リスク管理、ファンダメンタルズ分析の 3 つのドメインにまたがり、2,603 のタスク エピソードを含む 12 のサブタスクが含まれています。各エピソードでは、ポイントインタイムの入力、隠されたグラウンド トゥルース、およびタスク固有の検証ツールが提供されます。スキルなし、手順ドキュメントと実行可能コンポーネントで構成される厳選されたスキル パッケージ、エージェントがエピソード内で独自の手順を記述して再利用する自己生成スキルの 3 つの条件を比較します。 9 つのモデルと大規模な評価にわたって、厳選されたスキルにより一貫してパフォーマンスが向上し、平均スコアが 0.366 から 0.528 に上昇し、ポートフォリオ構築とリスク管理において最大の向上が見られました。対照的に、自己生成スキルは、計算コストが高いにもかかわらず、ほとんどメリットがありません。別のエージェント フレームワーク (Hermes Agent、8 モデル、合計 5,280 エピソード) を使用した独立した評価により、3 つのドメインすべてにわたる指向性パターンが再現され、スキル効果の大きさはサブタスクとハーネスによって異なります。これらの結果は、投資管理エージェントにおいては、信頼できる手順スキルへのアクセスがモデルの選択と同じくらい重要である一方、スキルの単純な自己生成は効果がないことが多いことを示しています。さらなる研究をサポートするために、ベンチマーク、評価ツール、厳選されたスキル パッケージ、および完全な軌跡をリリースします。
原文 (English)
FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management
Investment management is a high-stakes domain in which agentic AI systems must do more than generate plausible text. They must retrieve point-in-time data, assemble correct computational inputs, invoke specialized methods, and produce auditable structured outputs. We introduce FinSkillBench, an evaluation suite designed to measure whether language model agents can effectively use financial domain skills to solve investment management tasks. The benchmark spans three domains, portfolio construction, risk management, and fundamental analysis, and includes 12 subtasks with 2,603 task episodes. Each episode provides point-in-time inputs, hidden ground truth, and a task-specific verifier.We compare three conditions: no skill, curated skill packages consisting of procedural documents and executable components, and self-generated skills in which the agent writes and reuses its own procedures within an episode. Across 9 models and a large-scale evaluation, curated skills consistently improve performance, raising mean scores from 0.366 to 0.528, with the largest gains in portfolio construction and risk management. In contrast, self-generated skills provide little benefit despite higher computational cost. An independent evaluation using a separate agent framework (Hermes Agent, 8 models, 5,280 episodes total) reproduces the directional pattern across all three domains, with the magnitude of skill effects varying by subtask and harness. These results showthat in investment management agents, access to reliable procedural skills can be as important as model choice, while naive self-generation of skills is often ineffective. We release the benchmark, evaluation tools, curated skill packages, and full trajectories to support further research.
動的グラフ変換としての自己進化エージェント: 調査と新たな視点
大規模言語モデル (LLM) ベースのエージェントは、インタラクション全体にわたって持続し、記憶を維持し、ツールを使用し、スキルを習得し、ワークフローを改良し、他のエージェントと調整する自己進化型システムになりつつあります。これらの機能により、エージェントの状態が構造的かつ動的になります。エンティティ、関係、属性、依存関係、実行構造は、新しい証拠、フィードバック、環境条件によって変化します。既存のグラフエージェント調査は通常、グラフを進化する基盤としてではなくエージェント機能のサポート構造として扱いますが、自己進化エージェント調査はエージェントレベルのメカニズムに焦点を当てており、グラフトポロジの進化についてはほとんど議論しません。したがって、進化するエージェントの状態と動的グラフ トポロジの間の結合は、まだ調査されていません。この調査は、\textit{動的なグラフ変換としてのエージェントの進化} を構成することによって、これら 2 つの研究ラインを結び付けます。エージェントの状態を動的グラフとしてモデル化します。そこでは、記憶、ツール、スキル、ワークフロー、エージェント間の関係が、スキーマに制約された書き換えによって更新される型付きノード、エッジ、およびサブグラフとして表現されます。この定式化に基づいて、自己進化エージェントのための既存の動的グラフベースの手法を、ノード/フィーチャー進化、エッジ/トポロジー進化、サブグラフ活性化、およびクロスコンポーネント共進化の 4 つの分類に整理します。この分類に基づいて、自己進化するエージェントのための再利用可能なインフラストラクチャとして動的グラフ学習を提案し、9 つの動的グラフ学習サブフィールドをエージェント進化機能にマッピングし、それらの適応と考えられる障害モードについて説明します。最後に、エンドタスクの評価を補完する、動的グラフの観点から見た 5 種類のグラフ認識評価およびガバナンス プロトコルについて説明します。目標は、自己進化するエージェントを設計および管理するためのコンパクトな構造レンズを提供することです。
原文 (English)
Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective
Large language model (LLM)-based agents are increasingly becoming self-evolving systems that persist across interactions, maintain memories, use tools, acquire skills, refine workflows, and coordinate with other agents. These capabilities make agent states structural and dynamic: entities, relations, attributes, dependencies, and execution structures change with new evidence, feedback, and environmental conditions. Existing graph-agent surveys typically treat graphs as support structures for agent functions rather than as evolving substrates, while self-evolving-agent surveys focus on agent-level mechanisms and rarely discuss graph topology evolution. Thus, the coupling between evolving agent state and dynamic graph topology remains underexplored. This survey connects these two research lines by framing \textit{agent evolution as dynamic graph transformation}. We model agent state as a dynamic graph, where memories, tools, skills, workflows, and inter-agent relations are represented as typed nodes, edges, and subgraphs updated through schema-constrained rewrites. Based on this formulation, we organize existing dynamic-graph-based methods for self-evolving agents into four taxonomies: node/feature evolution, edge/topology evolution, subgraph activation, and cross-component co-evolution. Building on this taxonomy, we propose dynamic graph learning as reusable infrastructure for self-evolving agents and map nine dynamic-graph-learning subfields to agent-evolution capabilities, discussing their adaptations and possible failure modes. Finally, we discuss five types of graph-aware evaluation and governance protocols from a dynamic-graph perspective, which complement end-task evaluation. The goal is to provide a compact structural lens for designing and governing self-evolving agents.
エージェントティック AI の出現: 進化、背景、動作原理、アプリケーション、導入要因、および将来の研究の方向性に関するレビュー
エージェントティック AI は、人工知能の分野で新たな洞察と進歩を獲得しており、さまざまなドメインにわたる急速な変革を可能にする大きな可能性を促進しています。この急速な進歩とさまざまなドメインに革命を起こす可能性は、テクノロジーをより深く理解し、しっかりと把握する必要性を提唱しています。さらに、エージェントティック AI における最先端の研究の方向性に関する調査は、改善と応用の潜在的な範囲を包括的に評価するために実施する必要があります。したがって、これらの目的に対処するために、包括的なレビューは、研究者や実践者にエージェントティック AI の現状と将来の研究範囲に関する貴重な洞察を提供することができます。したがって、この研究では、さまざまな領域にわたるエージェントティック AI における最近出版された学術的貢献を検討し、エージェントティック AI の基礎と動作原理について議論し、人工知能におけるエージェンシーの歴史的および理論的進化を追跡します。システムを調査し、Agentic AI のアーキテクチャ、動作原理、および機能を調査および議論し、さまざまなドメインにわたる Agentic AI の実際のアプリケーションを調査し、研究結果を分析し、現在の課題を特定し、潜在的な将来の研究の方向性について議論し、提案されたシステム品質次元の助けを借りて、Agentic AI を使用および採用する利害関係者の意図に関する包括的なフレームワークを提案します。したがって、この体系的なレビューは、研究者と実践者に Agentic AI、その現在の開発と応用についての包括的な理解を提供し、主要な研究ギャップを強調します。今後の研究の方向性を概説します。
原文 (English)
Emergence of Agentic AI: A Review on Evolution, Background, Working Principles, Applications, Adoption Factors, and Future Research Directions
Agentic AI is gaining new insights and advancements in the field of Artificial Intelligence, fostering significant potential to enable rapid transformation across various domains.This rapid advancement and the potential to revolutionize various domains advocate the need for a deeper understanding and firm grasp of the technology. Moreover, an investigation into state of the art research directions in agentic AI needs to be conducted to comprehensively assess the potential scope for improvement and application.Therefore, to address these objectives, a comprehensive review can provide researchers and practitioners with valuable insights into the current state and future research scopes of agentic AI.Hence, this work considers the recently published scholarly contributions in agentic AI across various domains and discusses the fundamentals and working principles of Agentic AI, traces the historical and theoretical evolution of agency in artificial systems, explores and discusses Agentic AIs architecture, working principles, and functionalities, explores real-world applications of Agentic AI across various domains, analyzes the research findings, identifies current challenges, and discuss potential future research directions, and proposes a comprehensive framework of stakeholders intention to use and adopt Agentic AI with the help of proposed system quality dimensions.Therefore, this systematic review provides researchers and practitioners with a comprehensive understanding of Agentic AI, its current developments and applications, highlights key research gaps, and outlines future research directions.
解決は描画ではありません: オリンピック幾何学における図式的推論のベンチマーク
GPT やクロードなどの基礎モデルは、現在ではオリンピック レベルの数学を驚くべき熟練度で解決しており、幾何学の問題解決が数学的推論の標準的な代用となっています。しかし、幾何学問題を解くことと、それが依存する図形を描くことは、同じスキルではありません。多くの場合、進歩は適切な補助構造と出現を備えた忠実な図にかかっており、答えに至る道筋を推論するモデルが同様に答えを生み出すことができるかどうかは不明です。 MathVista や MathVerse など、モデルが正しい答えに到達するかどうかを測定するベンチマークのコレクションは増え続けていますが、私たちの知る限りでは、図自体を構築する明確な能力を分離するものはなく、この能力は測定されないままになっています。私たちは、このギャップをターゲットにするオープンソースのベンチマークを導入します。297 の問題のハード サブセットを含む 954 の自己完結型のオリンピック幾何学問題で、それぞれがその解法と、レンダリング可能な漸近線コードで人間が作成した忠実度の高い図と、図式的推論と呼ばれる一連のテキスト、コード、画像、VLM、および制約ベースのメトリクスと組み合わせられています。現在の基礎モデルを評価すると、解決と描画の間に顕著なギャップがあることが明らかになります。それらの図は著しく忠実度が低く、平均コンパイル成功率はわずか 36.14\% です。強力な数学的推論は、正確な幾何学図を構築する能力を意味しないことがわかりました。私たちのベンチマークとデータセットには、https://huggingface.co/datasets/max98765/hard_geometry_problems_with_diagrams からアクセスできます。
原文 (English)
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry
Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem solving has become a standard proxy for their mathematical reasoning. Yet solving a geometry problem and drawing the figure it depends on are not the same skill: progress often hinges on a faithful diagram with the right auxiliary constructions and incidences, and it is unclear that a model which reasons its way to the answer can also produce one. A growing collection of benchmarks, including MathVista, and MathVerse, measures whether models reach the correct answer, but to our knowledge, none isolate the distinct ability to construct the diagram itself, leaving this capability unmeasured. We introduce an open-source benchmark that targets this gap: 954 self-contained olympiad geometry problems, with a 297-problem hard subset, each paired with its solution and a human-authored, high-fidelity diagram in renderable Asymptote code, together with a suite of text-, code-, image-, VLM-, and constraint-based metrics for what we term diagrammatic reasoning. Evaluating current foundation models reveals a pronounced gap between solving and drawing: their diagrams are markedly less faithful, with an average compile success rate of only 36.14\%. Strong mathematical reasoning, we find, does not imply the ability to construct accurate geometric diagrams. Our benchmark and dataset can be accessed at https://huggingface.co/datasets/max98765/hard_geometry_problems_with_diagrams.
立場: AI のリーダーボードがグローバル サウスの恩恵を受けている: インドの事例研究
この意見書では、AI リーダーボードには独立したガバナンス、利益相反ポリシー、指標進化のメカニズムが欠如しているため、構造的にグローバル・サウスにサービスを提供するには不向きであると主張しています。障壁となるのはデータが欠落しているわけではありません。高品質の地域ベンチマークはすでに存在します: インドの IndicSUPERB、MILU、LAHAJA。アフリカ向けIrokoBench。アラビア語の「アルガーファ」。障壁となるのは制度設計です。グローバル リーダーボードにはこれらのベンチマークは含まれておらず、それを強制するガバナンス メカニズムもありません。商業的圧力により、グローバル ノースの有料顧客が影響を受ける場合、リーダーボードの失敗が修正されます。グローバル・サウスには同等の影響力がありません。ガバナンスがなければ、ヒンディー語、スワヒリ語、またはアラビア語話者に影響を与える障害は、文書化されているものの対処されていないギャップとして無期限に残ります。インドをケーススタディとして使用し(人口 14 億人、22 の予定言語、高品質のベンチマークがあるが、信頼できる集計は存在しない)、正式なガバナンスと開示ベースの紛争管理を一貫して好むことを示す 58 人の AI 実践者との協議から得られた結果を報告します。解決策は、より多くのデータではなく、より良い組織、つまり最初から独立したガバナンスを備えた地域のリーダーボードです。
原文 (English)
Position: AI Leaderboards Are Underserving the Global South: A Case Study from India
This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent governance, conflict-of-interest policies, and mechanisms for metric evolution. The barrier is not missing data; high-quality regional benchmarks already exist: IndicSUPERB, MILU, and LAHAJA for India; IrokoBench for Africa; AlGhafa for Arabic. The barrier is institutional design. Global leaderboards do not include these benchmarks, and no governance mechanism compels them to do so. Commercial pressure corrects leaderboard failures when paying customers in the Global North are affected. The Global South lacks equivalent leverage. Without governance, failures affecting Hindi, Swahili, or Arabic speakers persist indefinitely as documented but unaddressed gaps. Using India as a case study (1.4 billion people, 22 scheduled languages, high-quality benchmarks, but no trusted aggregation), we report findings from a consultation with 58 AI practitioners showing consistent preference for formal governance and disclosure-based conflict management. The solution is not more data but better institutions: regional leaderboards with independent governance from the start.
Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs
Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-Engli…
溶存ガス分析を使用した電力変圧器の故障診断のための IEEE キーガス法による最適化されたファジー ロジック アプローチ
電力システムの安定性を維持するには、信頼性の高い変圧器の故障診断が不可欠です。溶存ガス分析 (DGA) で広く利用されているアプローチである IEEE キー ガス メソッド (KGM) は、あいまいなデータに対処し、高い診断精度を確保するには限界があります。この研究では、ファジー ロジックと IEEE Key Gas Method (FL-KGM) を組み合わせた強化されたモデルを紹介します。このモデルでは、洗練されたメンバーシップ関数、最適化されたファジー ルール セット、および診断の不一致を排除するための CO と CO2 の新しい分離が導入されています。多次元ガス比分析と適応型分類フレームワークを活用することで、FL-KGM は優れた障害の特定と分類を実現します。実世界のデータセットを利用した実験による検証では、FL-KGM が最大 98.6% の精度を達成し、KGM や他の FL ベースのアプローチを大幅に上回ることが実証されました。これらの発見は、現代の電力システムにおける変圧器監視の進歩、インテリジェントな故障検出の実現、予知保全戦略の強化における FL-KGM の可能性を明らかにしています。
原文 (English)
Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas Analysis
Reliable transformer fault diagnosis is essential for maintaining power system stability. The IEEE Key Gas Method (KGM), a widely utilized approach in Dissolved Gas Analysis (DGA), exhibits limitations in addressing ambiguous data and ensuring high diagnostic accuracy. This study presents An enhanced model combining Fuzzy Logic with the IEEE Key Gas Method (FL-KGM) that introduces refined membership functions, optimized fuzzy rule sets, and a novel separation of CO and CO2 to eliminate diagnostic inconsistencies. By leveraging multidimensional gas ratio analysis and an adaptive classification framework, FL-KGM delivers superior fault identification and classification. Experimental validation utilizing real-world datasets demonstrates that FL-KGM achieves up to 98.6% accuracy, significantly outperforming KGM and other FL-based approaches. These findings elucidate the potential of FL-KGM in advancing transformer monitoring, enabling intelligent fault detection, and enhancing predictive maintenance strategies in modern power systems.
AI による農村部の医薬品の安全性の向上: スコーピングレビュー
はじめに: 投薬過誤 (ME) は世界の医療システムに対する重大な脅威であり、患者に損害を与える一因となっています。地方の医療に人工知能 (AI) を導入すると、患者の安全性が向上します。その目的は、地方の医療現場における患者の安全性の向上と投薬ミスの削減における AI テクノロジーの応用と有効性を調査することです。方法: EBSCohost、Emcare (Ovid)、MEDLINE、ProQuest Consumer Health Database を含む複数のデータベースにわたって、2012 年から 2025 年にわたる体系的な文献検索を通じて範囲レビューが実施されました。 9 つの異なる国からの 12 の主要研究が調査されました。データはテーマ別に分析され、投薬プロセス全体にわたる AI 介入に関する洞察が得られました。結果: AI テクノロジーは、処方と調剤から投与と投与後のモニタリングに至るまで、薬剤管理のあらゆる段階に統合されています。 4 つの重要なテーマが明らかになりました。(1) 利用されているさまざまな種類の AI (臨床意思決定支援システム、機械学習、自然言語処理、スマート ポンプなど)。 (2) 影響を受ける投薬プロセスの段階。 (3) これらのテクノロジーがエラーを最小限に抑え、ワークフローの安全性を高めるのにどれだけ効果的であるか。 (4) インフラストラクチャ、スタッフトレーニング、システム統合、警戒疲労などの地方特有の課題。いくつかの研究では、機械学習ベースの監視によりインシデント検出が向上し、処方と転写のエラーが 34% ~ 80% も減少することが実証されています。障壁には、ガバナンスの枠組みの欠如、財政的制限、臨床医の抵抗などがあり、依然として大きな障害となっています。結論: 地方の医療において、AI テクノロジーは医薬品の安全性を高める大きな可能性を秘めています。データ駆動型のモニタリングを可能にし、プロセスを自動化し、臨床上の意思決定支援を提供します。
原文 (English)
Improving Rural Medication Safety with AI: A Scoping Review
Introduction: Medication errors (MEs) represent a significant threat to global healthcare systems, contributing to patient harm. Introducing artificial intelligence (AI) in rural healthcare enhances patient safety. The aim is to explore the applications and effectiveness of AI technologies in enhancing patient safety and reducing medication errors in rural health settings. Methods: A scoping review was conducted through a systematic literature search spanning 2012 to 2025 across multiple databases, including EBSCohost, Emcare (Ovid), MEDLINE, and the ProQuest Consumer Health Database. Twelve primary studies from nine different nations were examined. Data were analysed thematically to obtain insights on AI interventions across the medication process. Results: AI technologies have been integrated into every stage of medication management, right from prescribing and dispensing to administration and post-administration monitoring. Four key themes came to light: (1) the various types of AI being utilised (like Clinical Decision Support Systems, Machine Learning, Natural Language Processing, and smart pumps); (2) the phases of the medication process that are affected; (3) how effective these technologies are in minimising errors and boosting workflow safety; and (4) rural-specific challenges including infrastructure, staff training, system integration, and alert fatigue. Several studies have demonstrated that machine learning-based surveillance improves incident detection and reduces prescribing and transcription errors by an impressive 34% to 80%. Barriers included lack of governance frameworks, financial limitations, and clinician resistance, which still present major obstacles. Conclusion: In rural healthcare, AI technologies hold great potential for enhancing pharmaceutical safety. They can allow data-driven monitoring, automate processes, and offer clinical decision assistance.
FraudBench: 適応型詐欺に対する銀行代理店のストレス テスト ポリシーに基づいたポリシー
会話型エージェントは、ツールを通じてエンド ユーザーの代わりに機能すると同時に、発信者が対話だけでアクセスできる顧客データベースや内部ポリシー文書へのアクセスを保持します。銀行取引は最も明確なケースです。質問に答える同じエージェントが、連絡先の詳細を変更したり、PIN をリセットしたり、送金したりすることもできるため、通常の顧客サービスは、承認、不正行為の検出、およびポリシーの遵守と切り離すことができません。既存の金融詐欺ベンチマークは静的なトランザクションまたはメッセージを分類しており、一般的なエージェントの安全性ベンチマークは即時注入または一般的な有害な使用を対象としています。なし。発信者が会話を通じて ID、認可、信頼を操作するときに、ポリシーに基づいた銀行代理店が安全に動作するかどうかをテストします。 $\tau^2$-bench デュアルコントロール フレームワークと $\tau$-Knowledge Banking 環境に基づいて構築された実行可能ベンチマークである FraudBench を紹介します。エージェントとシミュレートされた呼び出し元は両方とも、共有された変更可能なアカウント状態に対してツールを通じて動作し、エージェントは呼び出し元に選択されたツールへのアクセスを許可する場合があります。この環境は、エージェントが取得する必要がある 698 個の文書からなる内部ポリシー コーパスを公開します。 FraudBench には、150 の作成された敵対的シナリオが含まれています。凍結された公開セットの 107 件 (10 件の詐欺メカニズムにわたる 90 件と 17 件の連鎖適応型攻撃) が報告されたすべての実行に使用され、さらに 43 件の連鎖攻撃が阻止されました。安全性は履歴に依存します。単一制御タスクは 1 つを除くすべての前提条件を満たしますが、適応型攻撃では、以前のプローブ、許可、試行の失敗により、後のローカルで有効なリクエストが安全でなくなります。各シナリオには、観察可能な証拠、禁止された行為、安全な処置、および介入ポイントの注釈が付けられます。 107 の段階的タスクに関する 4 人のエージェントの予備的な単回試用評価では、攻撃セキュリティは 49\% ~ 65\% であり、モデル間で最も一般的な弱点はマネーミュールとファーストパーティ詐欺でした。
原文 (English)
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud
Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a caller can reach through dialogue alone. Banking is the clearest case: the same agent that answers a question can also change contact details, reset a PIN, or move money, so ordinary customer service is inseparable from authorization, fraud detection, and policy compliance. Existing financial-fraud benchmarks classify static transactions or messages, and general agent-safety benchmarks target prompt injection or generic harmful use; none test whether a policy-grounded banking agent safely acts when a caller manipulates identity, authorization, and trust over a conversation. We introduce FraudBench, an executable benchmark built on the $\tau^2$-bench dual-control framework and the $\tau$-Knowledge banking environment. Both the agent and the simulated caller act through tools over shared, mutable account state, and the agent may grant the caller access to selected tools; the environment exposes a 698-document internal policy corpus that the agent must retrieve from. FraudBench contains 150 authored adversarial scenarios; a frozen public set of 107 (90 across ten fraud mechanisms plus 17 chained adaptive attacks) is used for all reported runs, with 43 further chained attacks held out. Safety is history-dependent: single-control tasks satisfy every precondition but one, and adaptive attacks make a later, locally valid request unsafe because of an earlier probe, admission, or failed attempt. Each scenario is annotated with observable evidence, prohibited actions, safe dispositions, and intervention points. A preliminary single-trial evaluation of four agents on the 107 graded tasks yields attack-security between 49\% and 65\%, with money-mule and first-party fraud the most common cross-model weaknesses.
低リソース言語におけるヘイトスピーチ検出のための LLM の効率的な適応: ローマ字ウルドゥー語の比較研究
注釈付きデータが存在しないこと、言語構造が非公式であること、標準化された文法が存在しないことなどにより、低リソース言語 (LRL) でヘイトスピーチを検出することは困難です。このような課題の好例はローマ字ウルドゥー語です。これはソーシャルメディア上で南アジア人に広く使用されており、文脈上一貫したスペルが欠けているものの、バリエーションが豊富です。このペーパーの目的は、ウルドゥー語ローマ字でのヘイトスピーチ検出 (HSD) のための大規模言語モデル (LLM) の包括的な評価を実施し、低ランク適応 (LoRA) と呼ばれるパラメータ効率微調整 (PEFT) 手法を使用してこれらのモデルを微調整することです。ゼロショット推論を評価するために、Mistral、LLaMA、Falcon、多言語 BERT などのさまざまなトランスフォーマー モデルで PEFT に対してベンチマークを実行しました。実験は、72,000 を超える注釈付きコメントを含む PURUTT (有害なコメントと音訳のための並列ウルドゥー語およびローマ字ウルドゥー語コーパス) データセットで行われます。結果は、ゼロ ショット モデルのパフォーマンスは中程度 (F1 = 0.56) ですが、モデルのトレーニング可能なパラメーターのごく一部を更新すると、分類パフォーマンスが大幅に向上する (F1 > 0.93) ことを示唆しています。私たちの結果は、PEFT が優れた計算効率とともに卓越したパフォーマンスを提供し、低リソースの言語処理タスクに非常に適していることを示しています。
原文 (English)
Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu
It is challenging to detect hate speech in Low Resource Languages (LRLs) because of the absence of annotated data, the informality of its language structure, and the lack of standardized grammar. A good example of such a challenge is Roman Urdu which is broadly used by South Asians on social media and has a high variation while lacking contextually consistent spellings. The objective of this paper is to conduct a comprehensive assessment of Large Language Models (LLMs) for Hate Speech Detection (HSD) in Roman Urdu script and fine-tune these models using the Parameter-Efficient Fine-Tuning (PEFT) method called Low-Rank Adaptation (LoRA). To evaluate zero-shot inference, we benchmarked it against PEFT on different transformer models, including Mistral, LLaMA, Falcon, and multilingual BERT. Experiments are conducted on the PURUTT (Parallel Urdu and Roman Urdu Corpus for Toxic Comments and Transliteration) dataset with over 72,000 annotated comments. The results suggest that zero shot models perform moderately (F1 = 0.56), but updating a small fraction of the model trainable parameters improves the classification performance significantly (F1 > 0.93). Our results have shown that PEFT delivers outstanding performance alongside excellent computational efficiency, making it highly suitable for low-resource language processing tasks.
RDFdL: RDF と差動ダイナミック ロジックの統合
RDF でモデル化されたナレッジ グラフは、静的な知識を記述するのに強力ですが、物理システム (微分方程式で記述されるシステムなど) の動的な動作を捕捉したり推論したりすることはできません。これは、AI 駆動のサイバー物理システムにとって重大なギャップです。これを解決するために、RDF と差分ダイナミック ロジック (dL) を統合して、物理システムの静的な知識と連続的なダイナミクスの両方を表現および推論するフレームワークである RDFdL を提案します。動的部分では、RDF と SHACL の状態空間の微分方程式と範囲を構文的に表現し、dL への変換を使用してセマンティクスを提供します。一次ロジックの共有基盤を介して RDF と dL をリンクすると、独自の統合が実現します。動的ロジック ドメインの安全性と到達可能性のプロパティの検証結果が、RDF データに対する SPARQL クエリの含意として利用可能になります。オントロジー駆動の RDF 推論のための Apache Jena と、dL の定理証明者である KeYmaera X を使用してパイプラインを実装し、製造におけるその適用性をスケッチします。
原文 (English)
RDFdL: Integrating RDF with Differential Dynamic Logic
Knowledge graphs modeled in RDF are powerful for describing static knowledge, but they cannot capture or reason about the dynamic behavior of physical systems, e.g., systems described by differential equations, which is a critical gap for AI-driven cyber-physical systems. To solve this, we propose RDFdL, a framework that integrates RDF with Differential Dynamic Logic (dL) to represent and reason about both static knowledge and the continuous dynamics of physical systems. For the dynamic part, we syntactically represent differential equations and ranges in the state space in RDF and SHACL and provide semantics using a translation to dL. Linking RDF and dL through their shared foundation in first-order logic achieves a unique integration: verification results for safety and reachability properties in the dynamic logic domain become available as entailment to SPARQL queries over RDF data. We implement the pipeline using Apache Jena for ontology-driven RDF reasoning and KeYmaera X, the theorem prover for dL, and sketch its applicability in manufacturing.
Adversarial Review: Structured Disagreement for Grounded Agentic Code Review
Early multi-agent LLM systems often used role-separated teams, yet scaling agent count yields diminishing returns on repository-level codin…
ループ言語モデルにより構成ツールの呼び出しが改善される
ループ言語モデルは、推論ベンチマークで有望な結果を示していますが、エージェント ツールの使用の可能性はほとんど解明されていません。私たちはこの問題を、モデルが複数の API 呼び出しを調整し、中間状態を維持し、ツールの相互作用全体で依存関係を保持する必要がある構成的なツール呼び出し設定で研究します。 API-Bank、BFCL、NESTful でネイティブおよび後付けのループ言語モデルを評価し、一致する教師あり微調整レシピと推論時のさまざまな反復深度に基づいてトレーニングされたループ モデルと非ループ モデルを比較します。制御された実験では、反復計算は一般に、構成および依存関係を意識したツールの使用に利益をもたらしますが、分離された API 呼び出しではより小規模でモデルに依存する利益が得られます。通常、複数ステップのツールを使用する場合の精度は、繰り返しの深さが増すにつれて向上します。ただし、適応推論は、必要な場合にのみ追加の計算を割り当てることで、より有利な計算パフォーマンスのトレードオフを実現します。私たちの結果は、ループ言語モデルが、構成ツール使用ワークフローの信頼性の高い計画、調整、実行を必要とするエージェント システムにとって有望なアーキテクチャであることを示唆しています。
原文 (English)
Looped Language Models Improve Compositional Tool Calling
Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use remains largely unexplored. We study this question in compositional tool-calling settings, where models must coordinate multiple API calls, maintain intermediate state, and preserve dependencies across tool interactions. We evaluate native and retrofitted looped language models on API-Bank, BFCL, and NESTful, comparing looped and non-looped models trained under matched supervised fine-tuning recipes and varying recurrent depth at inference time. In controlled experiments, recurrent computation generally benefits compositional and dependency-aware tool use, while providing smaller and more model-dependent gains on isolated API invocation. Accuracy on multi-step tool use generally increases with recurrent depth; adaptive inference, however, achieves a more favorable compute-performance trade-off by allocating additional computation only when needed. Our results suggest that looped language models are a promising architecture for agentic systems that require reliable planning, coordination, and execution of compositional tool use workflows.
任意の格子におけるジャッカード距離の三角形不等式について
この論文では、格子と実際の評価に対する Jaccard 距離の一般化に関する新しい理論的結果を示します。評価が厳密に正、単調、モジュラーである場合、ジャッカード距離は任意の格子上の三角不等式を満たし、分布性に大きく依存していた以前の結果を効果的に一般化することを示します。相対的に補完された分布格子 (ブール代数に見られる大域境界の要件を安全に削除します) に移り、評価が正、単調、超モジュール、$\log$-サブモジュールである限り、三角不等式が成り立つことを証明します。さらに、部分モジュール評価のための対称差分 Jaccard 定式化を部分的に補完された分布格子に適用します。必要な条件に移ると、標準の一般化された Jaccard 距離が有効なメトリックとして機能するには、スーパーモジュール性が厳格な要件であることが証明されます。最後に、これらの構造的制約を緩和することの実際的な価値を、量子情報理論、形式概念分析、機械学習などの計算分野にマッピングし、未解決の数学的問題を簡単に見て終わります。
原文 (English)
On the Triangle Inequality for the Jaccard Distance in Arbitrary Lattices
This paper presents new theoretical results on generalizing the Jaccard distance for lattices and real valuations. We demonstrate that when the valuation is strictly positive, monotone, and modular, the Jaccard distance satisfies the triangle inequality on arbitrary lattices, effectively generalizing earlier results that depended heavily on distributivity. Moving to relatively complemented distributive lattices (which safely drop the requirement for the global bounds found in Boolean algebras), we prove the triangle inequality holds as long as the valuation is positive, monotone, supermodular, and $\log$-submodular. Additionally, we adapt the symmetric-difference Jaccard formulation for submodular valuations to sectionally complemented distributive lattices. Shifting to necessary conditions, we prove that supermodularity is a strict requirement for the standard generalized Jaccard distance to operate as a valid metric. Finally, we map the practical value of relaxing these structural constraints to computational fields like quantum information theory, formal concept analysis, and machine learning, closing with a brief look at open mathematical problems.
GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks
Genomic analysis on viruses such as SARS-CoV-2 variants: Beta, Gamma, Delta, and Omicron is heavily dominated by classical bioinformatics m…
Redakto - LLM 用のシークレット タブ
大規模言語モデル (LLM) は、日常のアプリケーションでますます使用されています。一般に、LLM または人工知能 (AI) のコンテキストにおける大きな課題は、LLM を使用する際のプライバシーを確保することです。これは、LLM に入力されるテキストから個人を特定できる情報 (PII) が削除されることを意味します。これらの課題は、EU の新しい法律によってさらに緊急性が高まっています。 EU 諸国におけるプライバシー上の懸念に関する LLM の使用に関する不確実性は、イノベーションの速度と研究からアプリケーションへの移行の大きな妨げとなる可能性があります。ここでは、\textbf{Redakto} を紹介します。これは、LLM またはその他の下流のテキスト処理にテキストを供給する前にテキストを匿名化するために使用できるツールです。当社は、PII の編集だけでなく、仮名化に使用される場合にも最先端の機能を提供します。これらの機能は、エンドユーザーが Redakto Web アプリケーションを通じて、また開発者や研究者が REST API とモデル コンテキスト プロトコル (MCP) フックを通じて簡単に使用できるように公開されています。この実装は完全にオープン ソースであり、必要なコンピューティング リソースは控えめで、ローカル ハードウェアに簡単に導入できます。これまでの研究とは対照的に、匿名化されたテキストの品質をより適切に評価するために、私たちはプライバシーと編集されたテキストの有用性の両方に関して、法的および医療分野のテキストデータに対して広範な実証的評価を実施しました。私たちの経験的結果は、さまざまな編集戦略で匿名化されたテキストが元のテキストと同等の有用性スコアを達成することを示しており、調査したタスクに大きな悪影響を与えることなく、Redakto による匿名化を LLM タスクに使用できることを示唆しています。
原文 (English)
Redakto - The Incognito Tab for LLMs
Large Language Models (LLMs) are being increasingly used in everyday applications. A major challenge in the context of LLMs or Artificial Intelligence (AI) in general is to ensure privacy when using them, meaning that personally identifiable information (PII) is removed from any text that enters an LLM. These challenges have become more urgent with novel EU legislation. Uncertainty around LLM usage with respect to privacy concerns in EU countries can be a major blocker for the speed of innovation and transfer from research to applications. Here we present \textbf{Redakto}, a tool that can be used for anonymizing text prior to feeding it to an LLM or other downstream text processing. We provide state-of-the-art functionalities for both redaction of PII but also when used for pseudonymization. These functionalities are exposed such that they can easily be used by end-users, through the Redakto web application, and by developers and researchers, via REST APIs and model context protocol (MCP) hooks. The implementation is fully open source, requires modest compute resources, and can be readily deployed on local hardware. In contrast to prior work and in order to better assess the quality of the anonymized texts, we conduct extensive empirical evaluations on textual data from legal and medical domain with respect to both privacy and utility of the redacted texts. Our empirical results demonstrate that the texts anonymized with different redaction strategies achieve utility scores on par with the original texts, suggesting that anonymization with Redakto can be used for LLM tasks without substantial negative impact for the tasks we explored.
設計上キャッシュ可能ですか?エッジのメモリ帯域幅の壁に対する局所性を高めるための専門家混合ルーターのトレーニング: システム測定調査で事前に登録された否定的な結果
単一の 8 GB GPU で 235B パラメータの Mixture-of-Experts (MoE) モデルを提供する場合、コンピューティングではなくメモリ帯域幅がボトルネックとなります。デコードでは、各トークンのアクティブなエキスパートを保持している層からストリーミングする必要があり、コンシューマ ハードウェアでは、ほとんどのエキスパートは RAM よりもはるかに遅い SSD に配置されます。この帯域幅の壁を Qwen3-235B (Q4_K_M、134 GB) で定量化します。測定されたデコードは 0.44 tok/s ウォームであり、トークンあたりのバイト数 / 帯域幅モデルと一致しますが、1 つのディスク スイープを代わりに償却する必要があるバッチ スキームは、ページング スラッシングによりバッチ 32 で崩壊します。ゼロサージェリー ルーター テレメトリ ツールである llama-moe-trace を構築し、Qwen3-30B でルーティングを測定します。隣接トークン エキスパートの再利用の確率は 2.0 倍、トラフィックの 95% でエキスパートの 52.5% が使用され、エキスパートの 13.4% の LRU キャッシュがリクエストの 66% に対応します。次に、キャッシュ可能性がトレーニング可能かどうかを尋ねます。キャッシュミス削減とパープレキシティに関する共同基準に基づいて、補助局所性とドメインルーター損失を伴う 1 億 3,700 万 MoE 言語モデルのトレーニングを事前登録します。このメカニズムは機能しますが (ミスは 60% まで減少し、静的ピンのヒット率は 99%)、すべての構成が事前に登録された <=1% の複雑性ゲートを通過できません。ミスの削減と品質は密接に関係しています。同時実行 StickyMoE は、単一ドメインの 2500 万未満のモデルでほぼ無料と同じ損失を報告します。マルチドメイン 1 億 3,700 万では、税金が現実のものであることがわかります。私たちの貢献は、この事前登録された、より厳格な基準のマルチドメイン評価とエッジ サービング測定です。 3 億 4,000 万段は、税が規模に応じて縮小しない (わずかに増加する) ことを示しています。さらに、訓練された局所性を備えた訓練不要のキャッシュ対応再ルーティング スタックを示します。両方のサイズで、両方のサイズで複雑度 3.4% 以下でミスが最大 80% 削減され、どちらか単独よりもはるかに安価です。ただし、ドメイン プライムのプリフェッチは役に立ちません。すべてのコード、トレース、事前登録が公開されています。
原文 (English)
Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study
Serving a 235B-parameter Mixture-of-Experts (MoE) model on a single 8 GB GPU is bottlenecked not by compute but by memory bandwidth: decode must stream each token's active experts from whichever tier holds them, and on consumer hardware most experts sit on an SSD far slower than RAM. We quantify this bandwidth wall on Qwen3-235B (Q4_K_M, 134 GB): measured decode is 0.44 tok/s warm, matching a bytes-per-token / bandwidth model, while a batching scheme that should amortize one disk sweep instead collapses at batch 32 from paging thrash. We build llama-moe-trace, a zero-surgery router-telemetry tool, and measure routing on Qwen3-30B: adjacent-token expert reuse is 2.0x chance, 95% of traffic uses 52.5% of experts, and an LRU cache of 13.4% of experts serves 66% of requests. We then ask whether cacheability is trainable: we pre-register training of 137M MoE language models with auxiliary locality and domain router losses, under joint criteria on cache-miss reduction and perplexity. The mechanism works (misses down up to 60%; a 99% static-pin hit rate) but every configuration fails the pre-registered <=1% perplexity gate -- miss reduction and quality are tightly coupled. Concurrent StickyMoE reports the same loss as near-free on single-domain sub-25M models; on multi-domain 137M we find the tax real. Our contribution is this pre-registered, stricter-criterion, multi-domain evaluation plus edge-serving measurements. A 340M rung shows the tax does not shrink with scale (it rises slightly). We further show training-free cache-aware rerouting stacks with trained locality -- together ~80% miss reduction at <=3.4% perplexity at both sizes, far cheaper than either alone -- while domain-primed prefetching does not help. All code, traces, and the pre-registration are released.
リスクの高い公共部門アプリケーションにおけるオープン モデルを使用した構造化情報抽出の評価
非構造化ドキュメントから構造化情報を抽出することは、あらゆる分野におけるデジタル変革の重要な要素となります。商用アプリケーションでは独自のソリューションが主流を占めていますが、急速に成長しているオープンソースの光学式文字認識 (OCR) エンジン、ラージ言語モデル (LLM)、およびビジョン言語モデル (VLM) のエコシステムは、アクセス可能な代替ソリューションを提供しています。ただし、現実的な複数ステップの抽出パイプラインに関する体系的な評価は依然として不足しています。このような抽出ツールを責任を持って使用するには、特にこれらのソリューションが EU AI 法で高リスクとして分類されている公共部門のアプリケーションの主要コンポーネントとなるため、現実的なタスクに関する包括的な評価が必要です。このギャップに対処するために、高リスクとして分類された現実世界の複雑な文書処理タスクにおけるオープンソース システムのエンドツーエンドのパフォーマンスを評価する包括的なベンチマークを紹介します。それは、国際学習プログラムへの学生の申請です。当社は、最先端の OCR エンジン、LLM、VLM を使用して包括的な実証評価を実施します。私たちの結果から、VLM は一般的に OCR+LLM パイプラインよりも優れたパフォーマンスを発揮しますが、最先端のオープンソース モデルでさえ、ゼロショット設定でそのようなタスクを確実に処理するのに苦労していることが明らかになりました。 35 構成のうち 4 つだけが 0.5 を超える F1 スコアを達成し、最高の OCR+LLM パイプラインが最高の VLM パフォーマンスと一致しましたが、ほとんどの OCR+LLM の組み合わせのパフォーマンスは大幅に悪かったです。すべての構成の約 75\% のスコアが 0.25 未満でした。モデルの規模はパフォーマンスに影響しますが、その関係は非線形です。大幅に大きなモデルが比例してより良い結果を保証するわけではありません。入力品質、特に OCR 出力の構造保持は、下流モデルの機能とは独立した重要な要素として浮上します。
原文 (English)
Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
The extraction of structured information from unstructured documents represents a critical component of digital transformations in all sectors. While proprietary solutions dominate commercial applications, a rapidly growing ecosystem of open-source Optical Character Recognition (OCR) engines, Large Language Models (LLMs), and Vision-Language Models (VLMs) offers accessible alternatives. However, systematic evaluations on realistic, multi-step extraction pipelines remain scarce. Responsible usage of such extraction tools require comprehensive evaluations on realistic tasks, especially as these solutions will be key components of applications in the public sector that the EU AI act categorizes as high risk. To address this gap we present a comprehensive benchmark assessing the end-to-end performance of open-source systems on a complex real-world document processing task classified as high risk: Student applications for an international study program. We conduct a comprehensive empirical evaluation with state-of-the-art OCR engines, LLMs and VLMs. Our results reveal that while VLMs generally outperform OCR+LLM pipelines, even state-of-the-art open-source models struggle to handle such tasks reliably in zero-shot settings. Only 4 of 35 configurations achieved F1 scores above 0.5, with the best OCR+LLM pipeline matching top VLM performance, though most OCR+LLM combinations performed substantially worse. Roughly 75\% of all configurations scored below 0.25. Model scale influences performance, yet the relationship is non-linear: substantially larger models do not guarantee proportionally better results. Input quality, particularly the structural preservation of OCR output, emerges as a critical factor independent of downstream model capability.
The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has becom…
SESSE: スケッチ、展開、並べ替え、要約、評価 -- 構造化分解による LLM-as-Judge 評価
審査員としての LLM 評価は、応答品質評価を単一の全体的な A/B 優先順位の選択に減らし、どの品質次元が優先順位を決定したかを分離したり、モデルのエラーを真のラベルの曖昧さから区別したりするメカニズムを提供しません。私たちは、全体的な判断を、裁判官自身の誤り事例から直接抽出された構造化されたサブ質問に分解する、トレーニング不要のフレームワークである SESSE (スケッチ、展開、並べ替え、要約、評価) を提案します。オラクルの応答、タスク固有のルーブリック、または微調整は必要ありません。 RewardBench (n=1,000) では、SESSE は思考連鎖ベースラインとほぼ同等の成績を達成し、完全にトレーニングを受けていないにもかかわらず、微調整されたスペシャリストである RISE-Judge-32B (92.7%) と競合します。基準ごとの投票証拠は、ラベルのあいまいさを診断し、単一の総合的な出力トークンからは利用できない失敗モードを判断するための解釈可能な監査証跡を提供します。
原文 (English)
SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition
LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice, providing no mechanism to isolate which quality dimensions drove the preference or distinguish model errors from genuine label ambiguity. We propose SESSE (Sketch, Expand, Sort, Summarize, Evaluate), a training-free framework that decomposes holistic judgment into structured sub-questions mined directly from the judge's own error cases; requiring no oracle responses, task-specific rubrics, or fine-tuning. On RewardBench (n=1,000), SESSE achieves near-parity with the chain-of-thought baseline and is competitive with RISE-Judge-32B (92.7%), a fine-tuned specialist, while remaining fully training-free. Per-criterion vote evidence provides an interpretable audit trail for diagnosing label ambiguity and judge failure modes unavailable from a single holistic output token.
ComponentBench: コンピュータ使用エージェントのコンポーネント レベルの障害を診断する
コンピュータ使用エージェントの現在の評価は、長期的なワークフロー ベンチマークとアトミックな GUI 基礎テストに分かれています。これにより、インスツルメントが不十分な中間層が残ります。つまり、診断するには十分短く、最新のインターフェイスの負担を把握するには十分に豊富な、現実的なコンポーネント中心のインタラクション (ボタン セットの切り替えなど) です。最新の Web UI 上のコンピューター使用エージェントをコンポーネント レベルで評価するためのベンチマークおよび診断パイプラインである ComponentBench を紹介します。 ComponentBench は、広く使用されているコンポーネント ライブラリ全体でプログラム的に検証された 2,910 のタスクとしてインスタンス化された 97 の正規 UI コンポーネントのライブラリに依存しないオントロジーを中心に構成されており、タスクの成功とインタラクション効率の両方の評価を可能にするクリーンな人間の参照軌跡と組み合わせられています。タスク収集を超えて、実装後に実現された構造的困難を監査し、タスクとコンポーネント ファミリ全体で構造化された障害分析を合成するためのスケーラブルなパイプラインを導入します。 4 つの観察およびアクション スペースにわたって 7 つのモデル (GPT-5.4、Gemini 3 Flash、GPT-5.4 mini、GPT-5 mini、Gemini 3.1 Flash-Lite、Qwen3-VL-235B、および UI-TARS-1.5-7B) を評価したところ、これらの設計の選択がパフォーマンスに重大な影響を与えることがわかりました。単一の共有ハーネス内で、観察スペースとアクション スペースのみを変更すると、同じモデルのタスクの成功率が 30% 以上シフトします。GPT-5 mini は、アクセシビリティ ツリー観察による 83.1% から、座標のみのピクセル制御による 48.9% に低下しました。さらに、最速の構成でも人間の基準と一致する場合の 3.7 倍の時間がかかり、人間にとっては些細な空間操作が現在のエージェントの課題となり続けています。
原文 (English)
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an under-instrumented middle layer: realistic component-centered interactions (e.g., toggle a button set) that are short enough to diagnose and rich enough to capture the burdens of modern interfaces. We present ComponentBench, a benchmark and diagnostic pipeline for component-level evaluation of computer-use agents on modern web UIs. ComponentBench is organized around a library-agnostic ontology of 97 canonical UI components instantiated as 2,910 programmatically verified tasks across widely used component libraries, paired with cleaned human reference trajectories that enable evaluation of both task success and interaction efficiency. Beyond task collection, we introduce a scalable pipeline for auditing realized structural difficulty after implementation and synthesizing structured failure analyses across tasks and component families. Evaluating seven models -- GPT-5.4, Gemini 3 Flash, GPT-5.4 mini, GPT-5 mini, Gemini 3.1 Flash-Lite, Qwen3-VL-235B, and UI-TARS-1.5-7B -- across four observation and action spaces, we show that these design choices critically impact performance. Within a single shared harness, changing only the observation and action space shifts task success by more than 30% for the same model: GPT-5 mini falls from 83.1% with accessibility-tree observations to 48.9% with coordinate-only Pixel control. Moreover, even the fastest configuration takes 3.7x as long as the matched human reference, and spatial manipulations that are trivial for humans continue to challenge current agents.
監督としてのガバナンス記録: 構造化ワークフロー修復のための検証者が選択した自己トレーニング
機械検証可能なワークフローは、タスクの契約、モデルの試行、検証者の決定、受け入れられた出力、およびターゲットの生成元をリンクするガバナンス レコードを生成します。これらのレコードが有界モデルを監視し、臨時または高価な機能を信頼性の高いワンショット実行に統合できるかどうかをテストします。新しい、構造的に切り離された PlanBench 再計画ケースでは、Qwen3-14B の考え方により、独立して作成された VAL 検証者によって認められた 24 の計画が生成されました。これらの計画は、神託のターゲットや強力な教師なしで、同じチェックポイントを非思考実行用に訓練しました。 80 の未開封のケースでは、VAL で受け入れられた計画は 1 から 57 に増加し、56 のペアゲインとゼロの回帰が発生しました。アダプターはすべてのケースでスキーマが有効であり、思考の平均レイテンシーの約 1/56 を使用しました。別のペアの界面硬化ゲートは通過しませんでした。マッチしたアブレーションにより、ターゲットの選択を変更しながら、ソース ケース、52 の候補プール、24 のターゲット数、モデル、レシピ、シードが固定されました。 160 件の新しいケースで、ベース、スキーマ選択、モデル自己選択、および VAL 選択の実行は、1、55、69、および 102 の承認済みプランに達しました。 VAL はペア純 +33 (p=0.0000019647) で自己選択を上回り、両方の難易度階層で増加しました。したがって、独立したセマンティック選択は、この帯域内で一致する代替案に比べて負荷がかかります。補完的なファイのより強力な教師アームにより、ベースのファイ-4 の承認されたプランが 2 から 51 に、スキーマ有効な出力が 35 から 80 に増加しました。初期の合成実験により、教育可能性、累積学習、構築の堅牢性、および停止境界が確立されます。その結果は、恣意的な計画、企業の正当性、無制限の自己改善ではなく、制限された機械チェック可能な機能に対する検証者が選択した監視をサポートします。
原文 (English)
Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair
Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and target origin. We test whether these records can supervise bounded models, consolidating occasional or expensive capability into reliable one-shot execution. On fresh, structure-disjoint PlanBench replanning cases, Qwen3-14B thinking generated 24 plans admitted by the independently authored VAL verifier. Those plans trained the same checkpoint for non-thinking execution, without oracle targets or a stronger teacher. On 80 unopened cases, VAL-accepted plans increased from 1 to 57, with 56 paired gains and zero regressions; thinking reached 30. The adapter was schema-valid on all cases and used approximately 1/56 of thinking's mean latency. The separate paired interface-cure gate did not pass. A matched ablation fixed the source cases, 52-candidate pool, 24-target count, model, recipe, and seed while changing target selection. On 160 new cases, base, schema-selected, model-self-selected, and VAL-selected execution reached 1, 55, 69, and 102 accepted plans. VAL exceeded self-selection by paired net +33 (p=0.0000019647), with gains in both difficulty strata. Independent semantic selection is therefore load-bearing relative to matched alternatives within this band. A complementary Phi stronger-teacher arm raised base Phi-4 from 2 to 51 accepted plans and from 35 to 80 schema-valid outputs. Earlier synthetic experiments establish teachability, cumulative learning, construction robustness, and stopping boundaries. The results support verifier-selected supervision for bounded, machine-checkable capabilities, not arbitrary planning, enterprise validity, or unrestricted self-improvement.
部分信用ギャップの測定: ベトナムの 2025 年凸型マーキング制度に関する厳格なベンチマーク
人間の試験で言語モデルを評価する場合、ベンチマークは通常、各回答の正誤をスコアリングし、全体的な精度を報告します。このアプローチは、部分的な知識が比例的に単位を取得する価値があることを前提としていますが、試験で非加法的な採点方式が使用されている場合、この前提は当てはまりません。 2025 年のベトナムの全国高等学校卒業試験の改革は、この代替のコストを示しています。試験のパート II では、受験者は質問ごとに 4 つの正誤記述を評価します。採点は凸型です。正しいステートメントの数により、0、0.10、0.25、0.50、または 1.00 ポイントが獲得されます。 3 つのステートメントを正しく識別すると、標準の精度指標で与えられる 0.75 ポイントではなく、0.50 ポイントが支払われます。パート II は試験の 10.00 点のうち 4.00 点を占めるため、レポートの正確性は、州が明示的に罰則を課している部分的な知識に報酬を与えることでスコアを吊り上げます。 THPT-Ladder は、11 科目にわたる 21 の公式試験からの 632 項目のベンチマークであり、同省が学生を採点するのとまったく同じように採点されます。同省は100万人を超える候補者のマークを公開しているため、モデルを人間のコホートに直接配置することができる。 8 つのモデルにわたって、公式ルーブリックの支払いは比例単位よりもパート II の質問ごとに 0.020 ~ 0.159 ポイント低くなります。この不足により、モデルの見かけの能力が変化します。 2025 年の歴史試験の Qwen3.5-27B では、0.042 ポイントの不足により、受験者 481,293 人中 90 パーセンタイルから 77 パーセンタイルに順位が下がりました。モデルの精度はこのペナルティを予測しません。クロード ソネット 5 の精度レベルでは、誤差のさまざまな分布により、質問ごとに 0.869 から 0.932 ポイントの範囲でスコアが得られます。公式マークは、正しいステートメントがどのようにグループ化されるかによって決まります。つまり、標準ベンチマークは、教育機関が認定しない能力を報告します。
原文 (English)
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme
When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part II accounts for 4.00 of the exam's 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students. The ministry publishes the marks of over a million candidates, allowing us to place models directly into the human cohort. Across eight models, the official rubric pays 0.020 to 0.159 points less per Part II question than proportional credit. This shortfall changes a model's apparent competence. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile among 481,293 candidates. A model's accuracy does not predict this penalty. At Claude Sonnet 5's accuracy level, different distributions of errors yield scores varying from 0.869 to 0.932 points per question. Official marks depend on how correct statements are grouped, meaning standard benchmarks report a competence the institution would not certify.
ギザギザのフロンティア: セマンティクスを保持する変換に対するコード エージェントの堅牢性の評価
AI コード エージェントは、実際のソフトウェアの問題を解決するために導入されることが増えていますが、表面的なコードのバリエーションにおける AI コード エージェントの信頼性については、依然として十分に理解されていません。周囲のコードベースが意味的に同等の形式に書き換えられたときに、リポジトリ レベルの問題を修復するコーディング エージェントの信頼性が維持されるかどうかを評価します。制御フローの書き換え、デッドコードの挿入、識別子の名前変更など、一般的なセマンティクス保持変換 (SPT) を適用して、摂動されたバリアントを生成するランダム バリアント サンプラーを導入します。 SWE ベンチ Verified および SWE ベンチ Pro から抽出されたインスタンスにわたる 4 つのフロンティア モデル (Claude Opus 4.5、Kimi K2.5、MiniMax M2.5、および Qwen 3.6-27B) のいずれかによってそれぞれサポートされる 2 つのエージェント スキャフォールド (ミニ SWE エージェントと OpenCode) を評価します。インスタンスごとに、エージェントは摂動のないバリアントと摂動のあるバリアントに対して複数回実行され、摂動の影響を固有の確率性から分離するペアの解決率推定値が得られます。ほとんどの構成でわずかな低下が見られます。モデル、足場、データセットの 16 構成のうち 6 構成で統計的に有意な低下が見られ、最も影響を受けた構成では平均最大 6.7 パーセント ポイントの解決率低下が見られます。重要なのは、堅牢性による単一のモデルのランキングがスキャフォールド全体で保持されるということはありません。Qwen は、SWE ベンチ検証済みのミニ SWE エージェントの下では最も堅牢なモデルの 1 つですが、OpenCode の下では最も脆弱です。これにより、ギザギザの堅牢性のフロンティアが明らかになります。より単純な足場 (ミニ SWE エージェント) は摂動に対してより堅牢です。私たちの結果は、最上位のフロンティア モデルであっても、その影響は均一ではないものの、セマンティクスを保持する摂動の影響を受けやすいことを示しており、現実世界の多様なコードベースにおける AI コード エージェントの展開の信頼性についての懸念が生じています。
原文 (English)
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations
AI code agents are increasingly deployed to resolve real software issues, yet their reliability under superficial code variations remains poorly understood. We evaluate whether coding agents that repair repository-level issues remain reliable when the surrounding codebase is rewritten into a semantically equivalent form. We introduce a random variant sampler that applies common semantics-preserving transformations (SPTs) - spanning control-flow rewrites, dead-code injection, and identifier renaming - to produce perturbed variants. We evaluate two agentic scaffolds (mini-SWE agent and OpenCode) each backed by one of four frontier models (Claude Opus 4.5, Kimi K2.5, MiniMax M2.5, and Qwen 3.6-27B) across instances drawn from SWE-bench Verified and SWE-bench Pro. For each instance, the agent is run multiple times on the unperturbed and perturbed variants, yielding paired resolve-rate estimates that isolate the perturbation effect from intrinsic stochasticity. We find small degradation in most configurations: up to 6.7 percentage points mean resolve-rate drop in the most affected configurations with statistically significant degradations in 6 of 16 configurations of model, scaffold, and dataset. Crucially, no single model ranking by robustness holds across scaffolds - Qwen is among the most robust under mini-SWE agent on SWE-bench Verified yet the most brittle under OpenCode - revealing a jagged robustness frontier. The simpler scaffold (mini-SWE agent) is more robust to perturbation. Our results demonstrate that even top frontier models are susceptible to semantics-preserving perturbations although the effect is not uniform, raising concerns about the deployment reliability of AI code agents in diverse real-world codebases.
クリーンな信号では不十分な場合: 安全なウェアラブル ストレス分類のための構造的曖昧性の検出
ウェアラブルストレス分類器は、平均的には優れたパフォーマンスを達成できますが、特定の個人に対しては完全に失敗します。 WESAD では、ランダム フォレストは平均精度 93.0% に達しますが、被験者 14 の F1 = 0 が得られます。被験者 14 の相互信号結合はストレスの開始近くで弱まります。私たちはこれを構造的曖昧性と呼んでいます。個別にもっともらしい生理学的チャネルは、人の非ストレス基準によって十分に裏付けられていない信号間パターンを形成します。我々は、被験者固有の結合発散を定量化し、下流の分類器を再トレーニングすることなく分類、延期、または棄権するために各ウィンドウをルーティングする、軽量で透明な推論前モニターである Individual Conformal Coupling Monitor (ICCM) を導入します。 WESAD (N = 15) と Stress-Predict (N = 35) 全体で、曖昧さと精度の間のフルコホートのピアソン関連は負です (r = -0.607、p = 0.016; r = -0.412、p = 0.014)。ロバストネス分析により、この結果は緩和されます。順位相関は有意ではなく、被験者 14 が削除されると WESAD の関連性は消失します。 ICCM は、偽陽性数を 29 から 27、94 から 92 に変更しますが、どちらの変化も重要ではありません。対象者 14 の 21 のストレス ウィンドウのうち 3 つを保留しますが、ストレスの失敗は修復されません。これらの結果は、ICCM を独立した安全保証ではなく、裏付けのない生理機能や個人の失敗を示す解釈可能なシグナルとして位置づけています。
原文 (English)
When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification
Wearable stress classifiers can achieve strong average performance while failing completely for a particular individual. On WESAD, a Random Forest reaches 93.0% mean accuracy yet yields F1 = 0 for Subject 14, whose cross-signal coupling weakens near stress onset. We call this structural ambiguity: individually plausible physiological channels form an inter-signal pattern that is poorly supported by the person's non-stress reference. We introduce the Individual Conformal Coupling Monitor (ICCM), a lightweight and transparent pre-inference monitor that quantifies subject-specific coupling divergence and routes each window to classify, defer, or abstain without retraining the downstream classifier. Across WESAD (N = 15) and Stress-Predict (N = 35), full-cohort Pearson associations between ambiguity and accuracy are negative (r = -0.607, p = 0.016; r = -0.412, p = 0.014). Robustness analyses temper this finding: rank correlations are not significant, and the WESAD association disappears when Subject 14 is removed. ICCM changes false-positive counts from 29 to 27 and 94 to 92, although neither paired change is significant. It withholds 3 of Subject 14's 21 stress windows but does not repair the missed-stress failure. These results position ICCM as an interpretable signal of unsupported physiology and individual failure, rather than a stand-alone safety guarantee.
形式的抽象化によるリソース制約のある言語モデルにおける自然言語の組み合わせ最適化の精度の向上
組み合わせスケジューリングは言語モデルにとって大きな課題であり、複雑な制約を満たしながら、指数関数的に大きな検索空間内で実現可能なソリューションを特定する必要があります。この課題は、リソースに制約のある設定で特に顕著です。この設定では、大規模な言語モデルは現実的ではなく、選択が小規模なモデルに限定されるため、自然言語から直接スケジュールする場合に実現可能性を維持できないことがよくあります。これらの制限に対処するために、低レベルのモデリングと検索を決定論的コンパイラーと外部ソルバーに委任しながら、自然言語スケジューリング問題をタスク、リソース、制約、および目的のソルバーに合わせたコンパクトな表現に変換するニューロシンボリック フレームワークである SDDL を導入します。 SDDL は、スケジューリング問題の 300 インスタンスのマルチファミリー サブセットにおいて、テストされたリソース制約のあるすべてのモデルについて個別に検証された実現可能性を向上させます。 2 つの最も強力な SDDL 構成は 55.3% と 28.3% に達し、直接生成ベースラインの 23.7% と 1.3%、ソルバーコード ベースラインの 21.7% と 7.0% よりも上昇しており、実現可能なスケジュール間の最適性の中央値ギャップは 0.0% です。 SDDL では、解決策やソルバー コードを生成するのではなく問題構造を表現することで、より小さなモデルが、実質的に大規模なフロンティア モデルを含む、最も強力に評価された直接コードおよびソルバー コードの構成に近づくことができます。
原文 (English)
Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions
Combinatorial scheduling poses a significant challenge for language models, requiring them to identify feasible solutions within exponentially large search spaces while satisfying complex constraints. This challenge is especially pronounced in resource-constrained settings, where larger language models are impractical and selection is limited to smaller models which often fail to preserve feasibility when scheduling directly from natural language. To address these limitations, we introduce SDDL, a neuro-symbolic framework that translates natural-language scheduling problems into compact, solver-aligned representations of tasks, resources, constraints, and objectives, while delegating low-level modeling and search to a deterministic compiler and external solver. On a 300-instance, multi-family subset of scheduling problems, SDDL improves independently verified feasibility for every resource-constrained model tested. The two strongest SDDL configurations reach 55.3% and 28.3%, up from direct-generation baselines of 23.7% and 1.3% and solver-code baselines of 21.7% and 7.0%, with a 0.0% median optimality gap among feasible schedules. By expressing problem structure rather than generating solutions or solver code, SDDL enables smaller models to approach the strongest evaluated direct- and solver-code configurations, including substantially larger frontier models.
FM-Bench: 競合エージェントとの長期的な管理のためのベンチマーク
言語モデル エージェントは、制限されたタスクを確実に実行するようになりました。行動が累積的な結果をもたらし、環境が彼らの選択に反応する場合、彼らが長期にわたって効果的な意思決定を維持できるかどうかは、ほとんど測定されていないままです。 FM-Bench (フットボールマネジメントベンチマーク) はこれを測定します。 LLM エージェントは、26 のツールとおよそ 340 ~ 400 の意思決定ストップを通じて、ゲーム内 20 年間フットボール クラブを運営します。すべてのライバルと同じ予算でチームをドラフトし、選手をトレードし、契約交渉をし、施設と若手に投資し、ラインナップを設定し、それを解雇できる理事会に回答する一方で、決定論的なエンジンが毎年、LLMの裁判官や人間の評価者なしで最終的なスコアを1つに蓄積していきます。ソロ トラックでは、凍結したスクリプト化された世界に対して 15 のフロンティア モデルのそれぞれが再生され、アリーナでは同じモデルとスクリプト化されたアンカーが 1 つの共有された 20 年間の世界に配置されます。私たちの知る限り、この規模での直接評価は初めてです。スコアの背後にある 6 つの行動能力を測定します。 3 つのシード全体で、15 モデルすべてがすべてのレベルを達成していますが、ほとんどのモデルでブラインド スクリプトベースラインは消滅しており、claude-fable-5 は平均スコアとアリーナでソロ ボードのトップに立っており、それでもタイトルは 10 モデル間で交代します。規模、価格、ベンダーのいずれも注文を予測しません。順序は地平線の後半でしか決まらず、最高のファーストプレイ人間がモデルボードの最下位にのみ着地します。モデルを区別するのは、計算ではなく管理動作です。より高いスコアのモデルは、終わり近くでペイオフの遅い投資を減らし、現金を遊ばせるのではなく投資し続け、期限のかなり前に更新を開始しますが、トークンの支出は何も予測しません。何百もの拒否された入札から市場の隠れた価格を学習するモデルはなく、自己管理メモリは、増大するだけのアーカイブかシーズンごとに書き換えられる計画という 2 つの相反するモードで失敗します。コードは https://github.com/Analogy-AI/fm-bench で入手できます。
原文 (English)
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget as every rival, trades players, negotiates contracts, invests in facilities and youth, sets lineups, and answers to a board that can fire it, while a deterministic engine accumulates every year into one final score with no LLM judge or human rater. The solo track plays each of 15 frontier models against a frozen scripted world, and the Arena places the same models plus a scripted anchor in one shared 20-year world; to our knowledge, the first head-to-head evaluation at this scale. We measure six behavioral capabilities behind the score. Across three seeds, all 15 models complete every horizon while the blind scripted baselines die out in most of theirs, and claude-fable-5 tops the solo board on mean score and the Arena, where the title nonetheless rotates among ten models. Neither scale, price, nor vendor predicts the order; the order settles only late in the horizon, and the best first-play human lands only at the bottom of the model board. What separates the models is managerial behavior rather than computation. Higher-scoring models reduce slow-payoff investment near the end, keep cash invested rather than idle, and open renewals well before the deadline, while token spend predicts nothing. No model learns the market's hidden prices from hundreds of rejected bids, and self-managed memory fails in two opposite modes: an archive that only grows or a plan rewritten every season. Code is available at https://github.com/Analogy-AI/fm-bench.
UMER: ユニバーサルマルチモーダル検索のためのペア認識判別推論による埋め込みとランキングの統合
ユニバーサル マルチモーダル検索は、効率的なコーパス スケールのマッチングときめ細かい意味論的推論の両方を必要とする、命令を意識した多様な検索タスクをサポートすることを目的としています。最近の MLLM ベースの埋め込み手法は通常、隠れ状態から表現を導き出しますが、思考連鎖 (CoT) 推論は、中間の意味論的証拠を表現空間にエンコードすることによる強化を埋め込むための有望な戦略として浮上しています。ただし、既存の CoT 手法は通常、クエリと候補に対して項目ごとの推論を単独で使用するため、肯定的なものと意味的に紛らわしい厳密な否定的なものを区別する明確な証拠は提供されません。さらに、対照的な埋め込みは全体的な類似性を捉えますが、回答の検証、カテゴリの判断、または詳細な推論を必要とするメタタスクに苦労します。この論文では、ユニバーサルマルチモーダル検索のための統合マルチモーダル埋め込みおよびランキングフレームワークである UMER を提案します。 UMER は、項目ごとの反映をペア認識判別推論に置き換えます。ペア認識判別推論は、クエリと候補のペアを比較して、指示に関連した一致および不一致の証拠を特定します。 UMER は、効率的なグローバル マッチングのための対照的な埋め込みと、単一の MLLM 内で明示的なペアごとの関連性の判断のための識別ランキングを共同学習します。相補的な相互蒸留戦略により、埋め込み関数とランキング関数の間で信頼性の高いペアごとの優先順位がさらに転送されます。 MMEB-V2 ベンチマークでは、UMER は、予算に応じて調整可能な推論をサポートしながら、同等の実験設定の下で最先端のパフォーマンスを達成します。
原文 (English)
UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval
Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching and fine-grained semantic reasoning. Recent MLLM-based embedding methods typically derive representations from hidden states, while Chain-of-Thought (CoT) reasoning is emerging as a promising strategy for embedding enhancement by encoding intermediate semantic evidence into the representation space. However, existing CoT methods typically use item-wise reasoning over queries and candidates in isolation, providing no explicit evidence to distinguish a positive from a semantically confusable hard negative. Moreover, contrastive embeddings capture global similarity but struggle with meta-tasks requiring answer verification, category judgment or fine-grained reasoning. In this paper, we propose UMER, a Unified Multimodal Embedding and Ranking framework for universal multimodal retrieval. UMER replaces item-wise reflection with Pair-Aware Discriminative Reasoning, which compares query--candidate pairs to identify instruction-relevant matching and discrepancy evidence. UMER jointly learns contrastive embeddings for efficient global matching and discriminative ranking for explicit pairwise relevance judgment within a single MLLM. A complementary mutual distillation strategy further transfers reliable pairwise preferences between the embedding and ranking functions. On the MMEB-V2 benchmark, UMER achieves state-of-the-art performance under comparable experimental settings while supporting budget-adjustable inference.
どのネガティブなことが重要ですか?テキスト エンコーダーに尋ねる: 高密度キャプション取得のための適応型類似マージン
密なキャプションの取得は、最近、セグメンテーション、エッジ マップ、LLM フィルター処理されたキャプション、およびクロスモーダル モジュールをコントラストの微調整に導入することによって改善されました。ただし、これらの手法は主に同じ InfoNCE 目標を継承しており、その最適化は、強力な事前トレーニング済み初期化の下では時期尚早に飽和する可能性があります。つまり、密なキャプションでは、最初のエポック内のバッチの 80% で損失が 10^{-3} を下回りますが、その勾配は測定の 47% で fp32 で正確にゼロに達します。この動作は、高密度キャプション ベンチマークの多数のほぼ重複したキャプションと密接に関連していることがわかりました。つまり、簡単な大部分がすでに分離された後、いくつかの非常に類似したネガが未解決のままになっています。解決策として、HN-CLIP を導入します。HN-CLIP は、テキスト エンコーダー独自のテキスト間ジオメトリを使用して、ネガティブごとの適応類似マージンを構築します。具体的には、分離されたキャプション類似性行列がネガティブ ロジットに追加され、ネガティブのマイニング、合成、またはリサンプリングを行わずに、より類似したキャプションにより大きなマージンが割り当てられます。結果として得られる目標には、トレーニング中に 1 つのキャプション類似度行列とマスクされたロジット加算のみが必要であり、補助データ、追加パラメーター、オフライン前処理、または推論時間のオーバーヘッドはありません。 4 つの高密度キャプション検索ベンチマークに関する広範な実験により、HN-CLIP は最強の競合他社よりも +2.4 ~ +4.3 R@1 向上し、同時に GOAL よりも 2.4 倍、StructXLIP よりも 5.4 倍速くトレーニングできることがわかりました。さらに、提案された目標は、ドメイン内ベンチマークでテストされた 6 つの微調整フレームワークすべてを改善し、わずか 20% のトレーニング データで最強のフルデータ ベースラインに達します。
原文 (English)
Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules into contrastive fine-tuning. However, these methods largely inherit the same InfoNCE objective, whose optimization can prematurely saturate under a strong pre-trained initialization: on dense captions, the loss falls below 10^{-3} on 80% of batches within the first epoch, while its gradient reaches exact zero in fp32 in 47% of measurements. We find that this behavior is closely related to the large number of near-duplicate captions in dense-caption benchmarks, where a few highly similar negatives remain unresolved after the easy majority has already been separated. As a remedy, we introduce HN-CLIP, which uses the text encoder's own text-text geometry to construct per-negative adaptive similarity margins. Specifically, a detached caption-similarity matrix is added to the negative logits, assigning larger margins to more similar captions without mining, synthesizing, or resampling negatives. The resulting objective requires only one caption-similarity matrix and a masked logit addition during training, with no auxiliary data, additional parameters, offline preprocessing, or inference-time overhead. Extensive experiments on four dense-caption retrieval benchmarks show that HN-CLIP improves over the strongest competitors by +2.4--+4.3 R@1 while training 2.4x faster than GOAL and 5.4x faster than StructXLIP. Moreover, the proposed objective improves all six tested fine-tuning frameworks on the in-domain benchmarks and reaches the strongest full-data baseline with only 20% of the training data.
ペアごとのランキングがオフライン説明選択のシングルアクション RL を上回る: 実践的なレッスン
LLM 上に構築された産業用の説明可能な推奨システムには、相当なサービス コストがかかります。リクエストごとに LLM 生成がトリガーされ、レイテンシは数百ミリ秒で、コストはトラフィックに比例して増加します。生成と選択を分離します。説明は凍結された候補プール (6 つのプロンプト スタイル、2 つのコモディティ LLM) として事前に生成され、CPU 常駐の小さなセレクターがリクエスト時に 1 つを選択します。スタックは GPU を必要とせず、100 ミリ秒未満で戻ります。私たちの主なベンチマークは、2,958 ペアの XRec Google Local サブセットで、6 つのオフライン プール セレクター (LambdaRank、PPO、GRPO、DPO、教師と生徒の蒸留) と 3 つの KG パス セレクター (ランダム ウォーク、エッジの素の列挙、MMR で再ランク付けされたパス) を評価します。 Claude-Sonnet-4.5 参照を使用した 300 ペアの MovieLens-1M 分割は、この設定に対する公開ベンチマークが存在しないため、内部データセット チェックとして機能します。すべてのバリアントは、XRec および G-Refer と同じ BERTScore-F1 プロトコルを使用し、5 つのシード間で平均化されます。 LambdaRank は Google Local で F1 = 0.500 に達し、G-Refer と XRec の両方を超え、MovieLens-1M チェックでは F1 = 0.329 に達しました。シード分散が 0.003 F1 未満の場合、順序付けは信頼できます。ペアごとのランク付け学習は、ロールアウトごとに 1 つのラベル付き候補のみを使用し、K-1 ラベルを未使用のままにするシングルアクション RL (PPO、GRPO、DPO) よりも優れたパフォーマンスを発揮します。 KG パス ファミリは異なる目標をターゲットとしています。3 つのバリアントすべてが、Google Local では USR = 1.000、MovieLens-1M では 0.997 ~ 1.000 に達します。これは、リクエストごとのパス グラウンディングによりクエリごとに一意の出力が生成され、キャッシュされた LLM 出力に影響を与えるテンプレートの折りたたみエラーが回避されるためです。 Claude 3 Haiku と Claude Haiku 4.5 を比較したジェネレーター プールの調査では、セレクターのランキングを維持しながら、F1 の小さなシフト (0.001 ~ 0.006) が示されています。セレクターとジェネレーターは独立して評価できますが、絶対的な F1 はジェネレーターに依存します。エンドツーエンドの構築コストは、汎用ハードウェアで 15 ドル近くかかります。
原文 (English)
Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson
Industrial explainable-recommendation systems built on LLMs incur a substantial serving cost: each request triggers an LLM generation, with latency in the hundreds of milliseconds and cost that scales linearly with traffic. We separate generation from selection: explanations are produced ahead of time as a frozen candidate pool (six prompt styles, two commodity LLMs), and a small CPU-resident selector picks one at request time. The stack needs no GPU and returns in under 100 ms. Our primary benchmark is a 2,958-pair XRec Google Local subset, evaluating six offline-pool selectors (LambdaRank, PPO, GRPO, DPO, teacher-student distillation) and three KG-path selectors (random walks, edge-disjoint enumeration, MMR-reranked paths). A 300-pair MovieLens-1M split with Claude-Sonnet-4.5 references serves as an internal cross-dataset check, since no public benchmark exists for this setting. All variants use the same BERTScore-F1 protocol as XRec and G-Refer, averaged across five seeds. LambdaRank reaches F1 = 0.500 on Google Local, exceeding both G-Refer and XRec, and F1 = 0.329 on the MovieLens-1M check. With seed variance below 0.003 F1, the ordering is reliable: pairwise learning-to-rank outperforms single-action RL (PPO, GRPO, DPO), which use only one labelled candidate per rollout, leaving K-1 labels unused. The KG-path family targets a different objective: all three variants reach USR = 1.000 on Google Local and 0.997-1.000 on MovieLens-1M, since per-request path grounding yields a unique output per query, avoiding template-collapse failures affecting cached-LLM outputs. A generator-pool study comparing Claude 3 Haiku and Claude Haiku 4.5 shows small F1 shifts (0.001-0.006) while preserving selector ranking: selector and generator can be evaluated independently, though absolute F1 depends on the generator. End-to-end build cost is near $15 on commodity hardware.
FinRCA-Bench: 金融 AI システムのベンチマーク証拠の検索と推論
金融業務をサポートするために大規模な言語モデルがますます使用されていますが、その見かけの推論パフォーマンスは、適切な証拠を受け取るかどうかに依存する可能性があります。財務調整では、診断に必要な証拠が請求書、注文書、承認、割り当て、支払い、元帳入力、銀行活動に渡って分散され、テキストの類似性ではなく取引関係によって関連付けられます。したがって、エンドツーエンドの精度は、証拠へのアクセスと推論の品質を混同する可能性があります。 FinRCA-Bench は、14 の運用テーブルにわたる 2,250 件の買掛金と銀行の調整ケースの決定論的合成ベンチマークであり、これには 15 の因果関係カテゴリにわたる 1,500 件の注入失敗と、750 件の正当またはハードネガティブなケースが含まれます。根本原因ラベルとレコードレベルの証拠コントラクトはモデルから隠蔽されるため、回答の正しさとは無関係に検索を評価できます。ルール/SQL、古典的な機械学習、高密度セマンティック検索、決定論的リレーショナル拡張、および永続的なトランザクション関係に限定された型付きトラバーサルである型付き来歴グラフ検索 (TPGR) を比較します。ルール/SQL は 84.97% の正確な精度に達し、従来の ML は 95.44% に達します。推論モデル、プロンプト、および生成の設定を固定し、取得のみを変更すると、マクロの必須レコード再現率が 0.83% から 77.70% に増加し、正確な 16 クラスの精度が 2.05% から 72.44% に増加します。構造的検索の失敗は、十分な検索が行われた場合の推論の失敗よりも 95 対 15 多くなります。検索が不完全であるにもかかわらず、254 件の正しい予測が発生し、厳密に返された証拠契約の精度はわずか 5.72% です。 FinRCA-Bench では、検索アーキテクチャが観察された AI システムのパフォーマンスを強力に形成し、正しい根本原因ラベルは監査可能な診断の弱い代用手段となります。
原文 (English)
FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems
Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether they receive the right evidence. In financial reconciliation, the evidence needed for diagnosis is distributed across invoices, purchase orders, approvals, allocations, payments, ledger entries, and bank activity, linked by transactional relationships rather than textual similarity. End-to-end accuracy can therefore conflate evidence access with reasoning quality. We introduce FinRCA-Bench, a deterministic synthetic benchmark of 2,250 accounts-payable-to-bank reconciliation cases spanning 14 operational tables, including 1,500 injected failures across 15 causal categories and 750 legitimate or hard-negative cases. Root-cause labels and record-level evidence contracts are hidden from the model, allowing retrieval to be evaluated independently of answer correctness. We compare Rules/SQL, classical machine learning, dense semantic retrieval, deterministic relational expansion, and Typed Provenance Graph Retrieval (TPGR), a typed traversal restricted to persisted transaction relationships. Rules/SQL reaches 84.97% held-out exact accuracy and classical ML reaches 95.44%. Holding the reasoning model, prompt, and generation settings fixed while changing only retrieval increases macro required-record recall from 0.83% to 77.70% and exact 16-class accuracy from 2.05% to 72.44%. Structural retrieval failures outnumber reasoning failures with sufficient retrieval by 95 to 15; 254 correct predictions occur despite incomplete retrieval, and strict returned-evidence contract accuracy is only 5.72%. On FinRCA-Bench, retrieval architecture strongly shapes observed AI-system performance, and a correct root-cause label is a weak proxy for an auditable diagnosis.
検索と CRM の橋渡し: 顧客の再エンゲージメントのための AI 製品調査エージェントの生産化
最新の e コマース プラットフォームは、検索、レコメンデーション、パーソナライゼーション、CRM システムを独立して運用していることが多く、積極的な顧客の再エンゲージメントの機会が制限されています。これは、ユーザーが購入前に外部調査のためにプラットフォームを離れる可能性がある、最高のスマートフォンや最新の 5G 携帯電話などの探索目的の場合に特に困難です。 AI を活用した製品調査エージェントを通じて検索と CRM ワークフローの橋渡しをする、スケーラブルな運用環境に導入されたフレームワークを紹介します。このシステムは、探索的な購入意図とエンゲージメントの低いユーザーを識別し、行動シグナル、外部知識、企業カタログ データを使用して根拠のあるマルチエージェント製品調査を実施し、WhatsApp を通じてパーソナライズされた推奨事項を提供します。モバイル製品検出のための約 15,000 件の WhatsApp 通知を含む 23 日間の実稼働環境でフレームワークを評価しました。このキャンペーンは、メッセージの転送と共有による二次的なエンゲージメントの証拠により、従来の WhatsApp レコメンデーション キャンペーンと比較して CTR の大幅な向上を達成しました。この導入により、下流での購入と GMV への影響も生じ、プロアクティブな顧客の再エンゲージメントとエンドツーエンドのカスタマー ジャーニーの最適化に対する AI 製品リサーチ エージェントの実際的な有効性が実証されました。
原文 (English)
Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement
Modern e-commerce platforms often operate search, recommendation, personalization, and CRM systems independently, limiting opportunities for proactive customer re-engagement. This is particularly challenging for exploratory intents such as best smartphones or latest 5G phones, where users may leave the platform for external research before purchasing. We present a scalable, production-deployed framework that bridges search and CRM workflows through AI-powered Product Research Agents. The system identifies users with exploratory purchase intent and low engagement, conducts grounded multi-agent product research using behavioral signals, external knowledge, and enterprise catalog data, and delivers personalized recommendations through WhatsApp. We evaluate the framework in a 23-day production deployment involving approximately 15K WhatsApp notifications for mobile product discovery. The campaign achieved substantial CTR improvements over traditional WhatsApp recommendation campaigns, with evidence of secondary engagement through message forwarding and sharing. The deployment also generated downstream purchases and GMV impact, demonstrating the practical effectiveness of AI Product Research Agents for proactive customer re-engagement and end-to-end customer journey optimization.
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each t…
軽量マルチモーダル モデルは LLM 推論のパフォーマンスを推定できますか?コンピューティングに最適なドキュメント推論に関する研究
推論推論の予算を LLM に一律に割り当てるとコストがかかり、考えすぎるペナルティが発生する傾向があります。特に、視覚的なレイアウトによって複雑さが増す文書タスクの場合はそうです。これに対処するために、3 つのドキュメント タスクにわたるモデル、予算、パフォーマンスのトレードオフを明示的に監視する初のマルチモーダル ベンチマークである BudgetDoc を導入します。 BudgetDoc を使用して、DRB (Document-Reasoning Balancer) をトレーニングします。 1B パラメーターの飛行前推定器 (SigLIP-2 + Qwen3-0.6B) は、予算レベル全体で順序モデルのパフォーマンスを予測し、0.753 の重み付け F1 を達成します。 5 つのフロンティア モデルと 3 つのデータセットに推論予算を動的に割り当てる場合、DRB はコストを大幅に削減しながら、15 構成中 9 構成で常に最大予算のベースラインと比較して F1 スコアと同等または向上します。最後に、予備評価により、DRB がクロスモデル選択に一般化できる可能性が実証されました。
原文 (English)
Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference
Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking penalties; especially in document tasks where visual layouts drive complexity. To address this, we introduce BudgetDoc, the first multimodal benchmark providing explicit supervision for model-budget-performance trade-offs across three document tasks. Using BudgetDoc, we train DRB (Document-Reasoning Balancer), an approx. 1B-parameter pre-flight estimator (SigLIP-2 + Qwen3-0.6B) that predicts ordinal model performance across budget levels, achieving a 0.753 weighted F1. When dynamically allocating reasoning budgets across five frontier models and three datasets, DRB matches or improves F1 scores compared to always-maximum-budget baselines in 9 of 15 configurations while drastically reducing cost. Finally, preliminary evaluations demonstrate DRB's potential to generalize to cross-model selection.
CTIFoundry: サイバー脅威インテリジェンスのためのエージェントネイティブのコーパス足場
サイバー脅威インテリジェンス (CTI) は、人間のアナリストではなく、クエリ時に複数ステップの調査を構成する LLM エージェントによって利用されることが増えています。この移行のハーネス側 (計画ループ、ツール プロトコル、コンテキスト管理) は急速に成熟しましたが、コーパス側は成熟していません。脅威レポートと脆弱性データベースは、埋め込みインデックスの背後にある不透明なチャンクとして、検索拡張生成用にまだパッケージ化されています。私たちは、モデルの能力ではなく、この基盤がエージェントによる CTI 調査のボトルネックであると主張し、エージェントネイティブのコーパス足場である CTIFoundry を提案します。ビルド時に、CTIFoundry は CTI コーパスの潜在構造を具体化します。つまり、4 つの権威ある知識ベース (CVE、CWE、CAPEC、ATT&CK) にわたる決定論的なオントロジー グラフであり、公式の相互参照は型付きのトラバース可能なエッジになります。スパンベースのレポート層。その正規のエイリアス解決されたクロスベンダーエンティティが来歴を保持するチャンクにインデックスを付けます。ハイブリッド密集 + 語彙検索サーフェス。クエリ時に、この構造は、標準のオープンソース エージェント ハーネスに搭載された 7 つの型指定されたツールと 3 つの手続き型スキルを通じて公開されます。公開されている CTIConnect ベンチマークでは、アクション サーフェスのみを交換すると、4 モデル、2 プロバイダー パネル全体で、同じように利用されたエージェントの F1 が +0.19 ~ +0.28 上昇します。CTIFoundry の小型モデルは、フラット サブストレートのフラッグシップモデルを上回っています。また、どちらのクロード モデルでも、ツール呼び出しの約半分でスキャフォールド エージェントの精度が高いため、検索の労力によって得られる利益はありません。アブレーションはそれを特徴づけます。型指定された構造がより大きなシェアを占め、手続き型スキルは構造を規律に変換し、スキルは存在する構造にのみ結合するため、この 2 つは超相加的に構成されます。
原文 (English)
CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence
Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at query time. The harness side of this shift has matured rapidly (planning loops, tool protocols, context management), but the corpus side has not: threat reports and vulnerability databases are still packaged for retrieval-augmented generation, as opaque chunks behind an embedding index. We argue that this substrate, not model capability, is the bottleneck on agentic CTI investigation, and present CTIFoundry, an agent-native corpus scaffold. At build time, CTIFoundry materializes the latent structure of a CTI corpus: a deterministic ontology graph over four authoritative knowledge bases (CVE, CWE, CAPEC, ATT&CK) whose official cross-references become typed, traversable edges; a span-grounded report layer whose canonical, alias-resolved cross-vendor entities index provenance-carrying chunks; and hybrid dense+lexical retrieval surfaces. At query time this structure is exposed through seven typed tools and three procedural skills mounted on a stock open-source agent harness. On the public CTIConnect benchmark, swapping only the action surface lifts the identically-harnessed agent by +0.19 to +0.28 overall F1 across a four-model, two-provider panel: a small model on CTIFoundry surpasses a flagship on the flat substrate, and the gain is not bought with search effort, since on both Claude models the scaffolded agent is more accurate at roughly half the tool calls. An ablation attributes it: typed structure carries the larger share, procedural skills convert structure into discipline, and the two compose super-additively, because skills bind only to structure that exists.
大規模言語モデルの不確定性下での優先推論
大規模な言語モデルが意思決定エージェントに進化するにつれ、好みを推論する能力が連携、調整、集合知の基礎となります。しかし、標準的なベンチマークとは異なり、現実世界の好みの推論は本質的に不確定です。情報が不完全である可能性があり、有効なソリューションが存在しない可能性があります。私たちは、正しさだけではなく、不確定性が AI 推論の中心的な課題であると主張します。我々は、この課題を 2 つの軸、(i) 不完全、部分的、または表現的な選好から生じる認識論的不決定性、および (ii) 標準的な社会的選択概念の下での解決策が存在しないことから生じる構造的不決定性、に沿って形式化します。タスクの階層全体にわたって、最先端の言語モデルは決定されたインスタンスと未決定のインスタンスを体系的に区別できず、検証設定においても推論が誤って調整されていることを示します。
原文 (English)
Preference Reasoning under Indeterminacy in Large Language Models
As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. Yet, unlike standard benchmarks, real-world preference reasoning is inherently indeterminate: information may be incomplete, and valid solutions may not exist. We argue that indeterminacy, rather than correctness alone, is a central challenge for AI reasoning. We formalize this challenge along two axes, (i) epistemic indeterminacy, arising from incomplete, partial, or expressive preferences, and (ii) structural indeterminacy, arising from the non-existence of solutions under standard social choice concepts. Across a hierarchy of tasks, we show that state-of-the-art language models systematically fail to distinguish between determined and undetermined instances, exhibiting miscalibrated reasoning even in verification settings.
透明センサー診断パイプライン検索の候補運命の説明
産業用センサーの診断は前処理、表現、分類パイプラインに依存しているため、自動化されたパイプライン検索は手動設計コストの削減に役立ちます。ただし、既存の自動機械/深層学習 (AutoML/AutoDL) レポートは通常、適合したトライアル、スコア、および勝者のみを保持し、無効、プルーニング、スキップ、キャッシュされた、または適合しない生成された候補を省略します。この省略により、信号の制約、予算の使用、および未評価の法的代替案をチェックするレビュー担当者の能力が制限されます。これに対処するために、診断検索トレースのための候補者レベルの監査フレームワークである候補者運命アカウンティングを提案します。観察された各候補者を監査可能な証拠として記録します。ハッシュは繰り返しの観察をマージし、合法性チェックは無効な候補者にフラグを立て、割り当ての根拠は予算決定を説明し、非公開の運命台帳は各候補者に 1 つの最終運命を割り当てます。 3 つの方位診断データセットの実験では、このフレームワークが無効な候補を検出し、適合トライアルのみのレポートで省略された 30 ~ 41 の候補を特定し、クローズド運命記録により、競争力のある診断パフォーマンスを維持しながら完全な候補の説明を検証できることが示されました。コードは https://github.com/XXIE999/candidate-fate-accounting で入手できます。
原文 (English)
Candidate-Fate Accounting for Transparent Sensor Diagnostic Pipeline Search
Industrial sensor diagnostics relies on preprocessing, representation, and classification pipelines, making automated pipeline search useful for reducing manual design cost. However, existing automated machine/deep learning (AutoML/AutoDL) reports typically retain only fitted trials, scores, and winners, omitting generated candidates that are invalid, pruned, skipped, cached, or unfitted. This omission limits reviewers' ability to check signal constraints, budget use, and unevaluated legal alternatives. To address this, we propose candidate-fate accounting, a candidate-level audit framework for diagnostic search traces. It records each observed candidate as auditable evidence: hashes merge repeated observations, legality checks flag invalid candidates, allocation rationales explain budget decisions, and a closed fate ledger assigns one terminal fate to each candidate. Experiments on three bearing-diagnostic datasets show that the framework detects invalid candidates and identifies 30--41 candidates omitted by fitted-trial-only reports, with closed fate records verifying complete candidate accounting while maintaining competitive diagnostic performance. The code is available at https://github.com/XXIE999/candidate-fate-accounting.
Sanyu Studio: 美術史的物語構築のためのマルチエージェント システム
生成型 AI が芸術解釈を標準化する可能性があるという懸念の中で、この論文では、LLM ベースのインタラクションが複数の芸術史的物語の構築をサポートできるかどうかを検討します。 Sanyu Studio は、321 枚の Sanyu の油絵を事実、解釈、構成、記憶フィルタリングのメカニズムを備えたエージェントとしてモデル化するマルチエージェント対話システムです。芸術大学の参加者8名による7日間のワークショップに基づくこの研究では、ユーザーのプロンプト、証拠の整理、認知傾向が、デジタルSanyuの多様でありながら一貫したバージョンを形成していることを示している。この研究結果は、歴史的証拠が限られている状況下で、AI が人間の主体性を増幅し、一般の観客に美術史解釈へのインタラクティブな入り口を提供できることを示唆しています。
原文 (English)
Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction
Amid concerns that generative AI may standardize art interpretation, this paper examines whether LLM-based interaction can support plural art-historical narrative construction. We present Sanyu Studio, a multi-agent dialogue system that models 321 Sanyu oil paintings as agents with fact, interpretation, organization, and memory-filtering mechanisms. Based on a seven-day workshop with eight art-university participants, the study shows that user prompts, evidence organization, and cognitive tendencies shaped divergent yet coherent versions of digital Sanyu. The findings suggest that, under conditions of limited historical evidence, AI can amplify human agency and offer public audiences an interactive entry point into art-historical interpretation.
RTPO: エージェントティック RL トレーニングを安定させるためのリバース ターン ポリシーの最適化
強化学習 (RL) を使用してマルチターン エージェント ワークフローをトレーニングすると、大規模な言語モデルで複雑な推論を実行し、外部ツールを使用し、シングル ターン設定を超えた反復検索を実行できるようになります。しかし、マルチターン RL トレーニングは非常に不安定なままであり、ターン数が増加するにつれてパフォーマンスが大幅に低下することがよくあります。理論分析を通じて、ロールアウト トレーニング コンテキストの不一致、最終報酬がまばらな場合の弱いターンレベルのクレジット割り当て、および異なるポリシー バージョンで短距離と長距離の軌道が最適化された場合の非同期ポリシー ドリフトの 3 つの密接に結合した不安定性の原因を特定します。我々は、これらの問題が平坦軌道の最適化に共通の構造的起源を共有していることを示し、統一されたリバースターン定式化を通じてそれらの問題に対処します。私たちは、マルチターン ロールアウトを疎なリバース ツリーとして編成し、ターンレベルのポリシー更新を時間的な逆順で実行し、各決定をその下流の継続と整合させる、リバース ターン ポリシー最適化 (RTPO) を提案します。 RTPO により、因果的に一貫したターンレベルのクレジット割り当てとポリシー上の継続が可能になり、非同期ドリフトを制御できます。我々は、RTPO が提案されたターンレベル定式化の下でコンテキストの不一致と非同期ドリフトを排除し、信用バイアスを低減し、再帰的最適性に収束することを示す理論的保証を提供します。マルチターン エージェント RL ベンチマークの実験では、RTPO が軌道レベルとターン レベルのベースラインをそれぞれ 21.50% と 10.76% 改善することが示されており、ツールを使用するエージェントのより安定したトレーニングをサポートする可能性が強調されています。
原文 (English)
RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training
Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use external tools, and conduct iterative search beyond single-turn settings. Yet multi-turn RL training remains highly unstable, often causing severe performance degradation as the number of turns increases. Through theoretical analysis, we identify three tightly coupled sources of instability: rollout-training context mismatch, weak turn-level credit assignment under sparse terminal rewards, and asynchronous policy drift when short and long trajectories are optimized under different policy versions. We show that these issues share a common structural origin in flattened trajectory optimization and address them through a unified reverse-turn formulation. We propose Reverse-Turn Policy Optimization (RTPO), which organizes multi-turn rollouts as sparse reverse trees and performs turn-level policy updates in temporal reverse order, aligning each decision with its downstream continuation. RTPO enables causally consistent turn-level credit assignment and on-policy continuation to control asynchronous drift. We provide theoretical guarantees showing that RTPO eliminates context mismatch and asynchronous drift under the proposed turn-level formulation, reduces credit bias, and converges to recursive optimality. Experiments on multi-turn agentic RL benchmarks show that RTPO improves upon trajectory- and turn-level baselines by 21.50% and 10.76%, respectively, highlighting its potential to support more stable training for tool-using agents.
正確さではなく、能力: スキル最適化におけるリファレンスフリーのジャッジゲートの診断
テキスト空間スキルの最適化は、自然言語スキル文書を進化させることによって凍結されたエージェントを適応させ、検証ゲートを通じて各候補者を受け入れます。既存のゲートは検証可能な報酬に依存しており、これらの方法は自動検証機能を備えたタスクに限定されています。ベリファイアを LLM 判定ゲートに置き換えると、その制限は解除されますが、そのようなゲートが使用可能な信号を伝送するかどうかはテストされていません。私たちは事前に質問をします。審査員をループに入れる前に、そのスコアが正解と不正解を区別できるかどうかを知ることができるでしょうか?私たちは、参照のない裁判官を潜在的な解決者として形式化します。その評決は、それ自体が結論するであろうものとの一致に基づいているため、裁判官の評価能力は解決能力によって制限されます。このモデルは、裁判官の能力 $c$ と解答空間サイズ $k$ における閉形式識別限界 (ROC-AUC)、必要条件 $c > 1/k$、および限界 AUC が項目の難易度によって混同されるが、質問内推定量は混同されないという結果をもたらします。非介入プローブは、決定を変更することなく、実際の最適化実行の審査員スコアを記録します。私たちは、能力が床の近くにあり、その上で使用できる場合に、偶然に識別可能性を見出します。裁判官のベンチマークの正確性が、重要な能力を誇張していること。そして、閉ループ研究では、画面がどの種類のゲート エラーが発生するかを予測します。その結果、ジャッジゲート用の安価な導入前診断が実現しました。
原文 (English)
Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization
Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document, accepting each candidate through a validation gate. Existing gates rely on verifiable rewards, confining these methods to tasks with an automatic verifier. Replacing the verifier with an LLM-judge gate would lift that restriction, but whether such a gate carries usable signal is untested. We ask a prior question: can we tell, before placing a judge in the loop, whether its scores separate correct from incorrect answers at all? We formalize a reference-free judge as a latent solver -- its verdict rests on agreement with whatever it would itself conclude, so its capacity to evaluate is bounded by its capacity to solve. The model yields a closed-form bound on discriminability (ROC-AUC) in the judge's competence $c$ and answer-space size $k$, a necessary condition $c > 1/k$, and the result that the marginal AUC is confounded by item difficulty while a within-question estimator is not. A non-intervening probe records judge scores on genuine optimization runs without altering any decision. We find discriminability at chance where competence sits near the floor and usable above it; that a judge's benchmark accuracy overstates the competence that matters; and, in a closed-loop study, that the screen predicts which kind of gating error occurs. The result is a cheap pre-deployment diagnostic for judge gates.
自動化されたエンタープライズ分析と洞察生成のためのマルチエージェント プラットフォーム
この論文では、会話型ビジネス インテリジェンスのために CrewAI [1] に基づいて構築されたマルチエージェント フレームワークを提案します。 5 つの専門 AI エージェントが順次パイプラインで動作し、自然言語クエリを処理し、データを取得および分析し、モデル コンテキスト プロトコル (MCP) [2] を介して視覚化を生成し、実用的な洞察を提供します。このプラットフォームは、マルチテナントのデータ分離のための多層防御セキュリティ アーキテクチャと、会話による洞察を再利用可能なダッシュボード コンポーネントに変換するためのクエリ パラメータ化メカニズムを備えています。合成エンタープライズ データセットと実稼働エンタープライズ データセットにわたる 300 のエンドツーエンド テスト ケースにわたる評価では、LLM-as-a-Judge フレームワークによる評価で機能精度 95.3%、平均応答待ち時間 24 秒、応答品質スコア 4.52/5.0 を示し、幻覚なし率は 93.0% で、単一エージェントと比較して精度が 22.6 パーセント ポイント向上し、品質が 20.2% 向上しました。ベースライン。 4 つの LLM バックエンドにわたるクロスモデル評価と専門家による検証により、アーキテクチャの一般化可能性と評価者の信頼性が確認されます。アブレーション研究により、データ分析エージェントとレポート集計エージェントが出力品質の主な要因であることが確認されています。
原文 (English)
A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation
This paper proposes a multi-agent framework built on CrewAI [1] for conversational business intelligence. Five specialized AI agents operate in a sequential pipeline to process natural language queries, retrieve and analyze data, generate visualizations via the Model Context Protocol (MCP) [2], and deliver actionable insights. The platform features a defense-in-depth security architecture for multi-tenant data isolation and a query parameterization mechanism for transforming conversational insights into reusable dashboard components. Evaluation across 300 end-to-end test cases spanning synthetic and production enterprise datasets demonstrates 95.3% functional accuracy, a mean response latency of 24 seconds, and a response quality score of 4.52/5.0 as assessed by an LLM-as-a-Judge framework, with a 93.0% hallucination-free rate, representing a 22.6 percentage point accuracy improvement and 20.2% quality gain over a single-agent baseline. Cross-model evaluation across four LLM backends and human expert validation confirm architectural generalizability and evaluator reliability. An ablation study confirms that the Data Analysis and Report Aggregation agents are the primary drivers of output quality.
自らを記述するメトリクス: 評価器を自らの死角から進化させる
エージェントは、信頼できる自動メトリックに対してはすぐに改善しますが、それがなければ機能不全に陥ります。また、エージェントを最も必要とするアプリケーション、その中でもレポート生成は、スコアの付け方を誰も知りません。メトリクス自体を書き込むことができますか?何が答えを良いものにするかを言うのは難しいです。 1 つの問題を指摘する方が簡単なので、私たちが進化させるメトリクスは、それぞれが 1 つの名前付き欠陥の候補にフラグを立てるか、棄権して投票する小さな Python 演算子のプールです。モデルに演算子を直接求めることは機能しません。183 人の候補者は、巨大な空間の 1 つの狭い領域から 96 の異なる動作のみを実現します。 EvalCEGAR は代わりに、反例に基づいた抽象化の改良をプログラム検証から借用します。プールを抽象化として読み取り、衝突を検索します。演算子のスコアが同じ 2 つの回答、1 つは正解、もう 1 つは不正解です。プロンプトではなく、そのペアがオーサリング リクエストであり、衝突によってすべての試行が失敗すると、ループにより、リサンプリングではなくオペレータが読み取れる範囲が広がります。 MBPP+ および HumanEval+ (非表示の単体テストで正確なグラウンド トゥルースが得られるサンドボックス) では、ループは 55 行の演算子を作成します。これにより、何もフラグを立てない場合と、428 個の未確認タスクに対する完全なフィルター (+0.0065、p=0.0010) の間のギャップを 15.4% 縮めることができます (+0.0065、p=0.0010)。これは、最高の手書き演算子のフラグの 4 分の 1 です。ベンチマークでは、フラグの 3 分の 1 に対するオペレーターの効果と正確に一致することはありませんでした。 8 回の実行のうち 6 回はそのようなオペレーターを許可し、6 回すべてがサンプルの抽出に役立ちました。 15 個の手書きの演算子を 1 つのフィルターとして一緒に適用すると、精度が失われます。同じ情報に基づく LLM 裁判官は、そのデルタをほぼばらばらの候補者のセットに結び付け、オペレータが請求しない場合、候補ごとにモデル コールを永久に請求します。
原文 (English)
Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots
Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generation among them, are the ones nobody knows how to score. Can the metric write itself? Saying what makes an answer good is hard; pointing at something wrong with one is easier, so the metric we evolve is a pool of small Python operators that each flag a candidate for one named defect, or abstain, and vote. Asking a model for operators directly does not work: 183 candidates realise only 96 distinct behaviours, from one narrow region of an enormous space. EvalCEGAR instead borrows counterexample-guided abstraction refinement from program verification. It reads the pool as an abstraction and searches for a collision, two answers the operators score identically, one correct and one not. That pair, not a prompt, is the authoring request, and when a collision defeats every attempt the loop widens what an operator may read rather than resampling. On MBPP+ and HumanEval+, a sandbox whose hidden unit tests give exact ground truth, the loop writes a 55-line operator that closes 15.4% of the gap between flagging nothing and a perfect filter on 428 unseen tasks (+0.0065, p=0.0010) at a quarter of our best hand-written operator's flags. On the benchmark it never saw it matches that operator's effect exactly on a third of the flags. Six of eight runs admit such an operator and all six help out of sample; our 15 hand-written operators applied together as one filter lose accuracy. An LLM judge on the same information ties that delta on a nearly disjoint set of candidates, and charges a model call per candidate forever where the operator charges none.
Pairwise Logical Selection of Enthymeme Completions under Semantic-Link Uncertainty
Arguments often omit premises or claims, forming enthymemes. We study pairwise logical selection between two candidates for the omitted com…
棄権の検証により、配水ネットワークにおける AI 漏水診断の責任が問われます
公共事業会社は漏洩により処理水のかなりの部分を失っているが、人工知能のローカライザーを信頼して作業員を派遣することはほとんどない。どこにでもある推測では掘削を正当化することはできない。ギャップは正確さではなく説明責任です。いつ行動すべきではないかを証明する方法はありません。ここでは、漏れの位置特定を、検証可能な棄権の下での意思決定として言い換えます。物理学に基づいた実行エージェントが、デジタル ツインに対して仮説 (漏れ、需要、センサー、バルブ) を改ざんします。大言語モデル (LLM) 監査人を伴う独立した監督エージェントが、コード検証可能な契約と照らし合わせて証拠をチェックし、派遣を認証し、証拠を要求するか棄権します。フィールドグレードのノイズの下では、32% の強制ベースラインが、動作したイベントの決定精度 96% になります。独自に生成されたベンチマークでは、33 個のリークのうち 4 個のみに作用し、すべて正解です。監査済みの実際の漏水箇所の 194 件のイベント記録と、ツインでシミュレーションされた圧力と流量を使用すると、地区の完全な精度で 5 回の掘削派遣、3 回の正解、および 44% の調査回収率が得られます。責任ある棄権は、自律的な水道インフラ運営への防御可能な道を提供します。
原文 (English)
Verifiable abstention makes AI leak diagnosis accountable in water distribution networks
Utilities lose a substantial share of treated water to leakage, yet rarely trust artificial-intelligence localizers to dispatch crews: guessing everywhere cannot justify excavation. The gap is accountability, not accuracy: no method proves when it should not act. Here we recast leak localization as decision-making under verifiable abstention. A physics-grounded executor agent falsifies hypotheses (leak, demand, sensor, valve) against a digital twin; an independent supervisor agent, with a large-language-model (LLM) auditor, checks evidence against a code-verifiable contract, then certifies a dispatch, requests evidence or abstains. Under field-grade noise, a 32% forced baseline becomes 96% decision precision on acted events. On an independently generated benchmark it acts on only 4 of 33 leaks, all correct. A 194-event register of audited real leak locations with twin-simulated pressures and flows yields five excavation dispatches, three correct, and 44% survey recovery at full district precision. Accountable abstention offers a defensible route to autonomous water-infrastructure operation.
ORBITER: エージェントによるラストマイル配信のための競合を意識した意思決定
ラストワンマイル配送は、複雑な空間的および時間的相関関係をモデル化しながら、宅配業者が動的に到着する注文を処理することを目的としています。最近の学習ベースの手法は、注文間の時空間依存関係をモデル化して宅配便の順序を予測しますが、次の注文の意思決定については説明がつかないままになっています。現在の配信状態を言語で説明することで、LLM は個々の決定の背後にある空間的、時間的、および行動的な手がかりについて明確に推論できるようになります。ただし、LLM は直接的な予測因子として、タスクの提示に依然として敏感であり、信頼性の低い決定を下すことがよくあります。これらの課題に対処するために、ラストマイル配送における次の注文の意思決定を行うためのエージェント型注文アービターである ORBITER を導入します。 ORBITER は、意思決定ポイントを通じて宅配サービスをモデル化します。各決定ポイントには、宅配業者の時空間状態と目に見える注文が含まれ、モデリングと検証のためのローカル トレードオフが明らかになります。固定の提案者が候補者をランク付けし、構造化されたレポートによって、そのランク付けが一致しない箇所が特定されます。 LLM はタスク固有のツールを使用して主要な代替案に関する証拠を収集し、独立した批評家が結果として得られた決定をその証拠に照らしてチェックします。当社は 4 つの都市のデータについて広範な評価を実施しており、そこでは ORBITER が既存の最先端のベースラインを平均で最大 9.2% 上回っており、その有効性が示されています。
原文 (English)
ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery
Last-mile delivery aims to handle dynamically arriving orders with couriers while modeling complex spatial and temporal correlations. Recent learning-based methods model spatiotemporal dependencies among orders to predict courier service sequences, but leave next-order decision making unexplained. Describing the current delivery state in language allows LLMs to reason explicitly about the spatial, temporal, and behavioral cues behind an individual decision. As direct predictors, however, LLMs remain sensitive to task presentation and often produce unreliable decisions. To address these challenges, we introduce ORBITER, an agentic Order Arbiter for next-order decision-making in last-mile delivery. ORBITER models courier service through decision points, each containing the courier's spatiotemporal state and visible orders and exposing local trade-offs for modeling and verification. Fixed proposers rank the candidates, and a structured report identifies where their rankings disagree. The LLM uses task-specific tools to gather evidence on the leading alternatives, while an independent critic checks the resulting decision against that evidence. We conduct extensive evaluations on data in four cities, where ORBITER outperforms existing state-of-the-art baselines by up to 9.2% on average showing its effectiveness.
SkillGate: Long-Horizon エージェントにおけるポリシー内のスキル選択のトレーニング
エージェント フレームワークでは、手順に関する知識をスキルとしてパッケージ化するケースが増えています。つまり、エージェントがオンデマンドで読み取る指示ファイルですが、公共ライブラリには現在、何千もの指示ファイルが保管されています。したがって、どのスキルを読むかは、エピソードの途中でポリシー自体が決定することになりますが、それを訓練する既存の信号はありません。私たちは、デフォルトの救済策である候補スレートに対する結果報酬型の RL ではそれを教えることができないことを示します。その構造的な理由から、セレクターのクレジット飢餓を特定し、名前を付けます。ブロードキャストされたシーケンスレベルの利点の下では、選択されたスキルを指定する少数のトークンが損失の一部を失い、それらが継承するクレジットは、軌道が長くなるにつれてますます誤って署名されます。正しい選択は、その選択自体が軌跡の中で最も価値のある決定の 1 つであるにもかかわらず、失敗後の実行に常に罰せられます。完了した実行自体のトレーニング成果物を監査すると、3 つのプロパティすべてが確認され、それぞれがホライズンとともに単調に悪化します。 SkillGate は構築によって障害を除去します。トークン サポートを 2 つの互いに素なクレジット チャネル、実行トークンのみに到達する結果クレジット、およびスキル命名トークンに正確に到達する別のアクション ローカル アドバンテージ (軌道の単一の読み取りが正しい場合にのみプラス) に分割します。 16 人の候補者による 5 つのエージェント ベンチマークで、SkillGate は 9B ポリシーの試験成功率を 40.8% から 53.2% に引き上げました。これは、結果報酬のみに費やされる同じ予算を大幅に上回る一方で、誤解を招く候補者への露出を 3 分の 2 削減し、読み込むスキルを減らします。
原文 (English)
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. We show that the default remedy, outcome-rewarded RL over the candidate slate, cannot teach it, for a structural reason we identify and name selector credit starvation: under a broadcast, sequence-level advantage, the few tokens that name the chosen skill carry a vanishing share of the loss, and the credit they inherit is increasingly wrong-signed as trajectories lengthen. A correct choice is punished whenever the execution after it fails, even though the choice itself is among the most valuable decisions in the trajectory. Auditing a completed run's own training artifacts confirms all three properties, each worsening monotonically with horizon. SkillGate removes the failure by construction: it partitions the token support into two disjoint credit channels, outcome credit reaching only execution tokens, and a separate action-local advantage reaching exactly the skill-naming tokens, positive only when a trajectory's single read is the correct one. On five agentic benchmarks under a 16-candidate slate, SkillGate lifts a 9B policy from 40.8% to 53.2% trial success, well ahead of the identical budget spent on outcome reward alone, while cutting exposure to misleading candidates by two thirds and reading fewer skills.
DentAgent: 多面的な歯科推論のための証拠中心のマルチエージェントの調整
口腔疾患は世界中の何十億人もの人々に影響を及ぼしており、専門分野の知識、X線写真、口腔内写真、3D歯科データからの異質な証拠を統合する、正確で信頼性の高い歯科評価の緊急の必要性が浮き彫りになっています。既存の歯科 AI システムのほとんどは依然としてモダリティまたはタスク固有です。最近の視覚言語モデルは柔軟な歯科質問応答をサポートしていますが、直接生成された応答は暗黙的で追跡不可能な証拠を残します。これらの制限に対処するために、オーケストレーターがさまざまなモダリティにわたる 5 つの専門エージェントを調整する、証拠中心のマルチエージェント フレームワークである DentAgent を導入します。各専門家はドメイン ツールを利用して、観察結果を構造化された証拠記録に変換します。 Evidence Blackboard は、これらの記録を共有証拠状態として管理し、応答生成前にカバレッジ、ギャップ、競合を追跡します。この標準化された証拠表現は、個別の歯科機能を統合されたエージェントのワークフローに統合します。 DentAgent は 4 つのベンチマーク全体で優れたパフォーマンスを示し、マルチラベル診断では上級専門医を 17.3 ポイント上回っています。これは、広く適用可能で追跡可能なマルチモーダル歯科推論の価値を裏付けており、集団の口腔健康評価と管理の技術的基盤としての可能性を強調しています。
原文 (English)
DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning
Oral diseases affect billions of people worldwide, underscoring a pressing need for accurate and reliable dental assessment that integrates heterogeneous evidence from domain knowledge, radiographs, intraoral photographs, and 3D dental data. Most existing dental AI systems remain modality- or task-specific. Although recent vision-language models support flexible dental question answering, directly generated response leaves evidence implicit and untraceable. To address these limitations, we introduce DentAgent, an evidence-centric multi-agent framework, in which the Orchestrator coordinate five specialized agents spanning various modalities. Each specialist utilizes domain tools to convert observations into structured evidence records. The Evidence Blackboard manages these records as a shared evidence state, tracking coverage, gaps, and conflicts before response generation. This standardized evidence representation integrates isolated dental capabilities into a unified agentic workflow. Across four benchmarks, DentAgent demonstrates leading performance, even surpassing the senior specialists by 17.3 percentage points on multi-label diagnosis, which supports its value for broadly applicable and traceable multimodal dental reasoning, and highlights its potential as a technical foundation for population oral health assessment and management.
大規模な言語モデルに対するトレーニング不要の推論時間の自己反映とコスト制限のある早期停止
推論 LLM (GRPO など) の強化学習トレーニングは高価であり、完全なトレーニング パイプラインにすべての貢献をコミットする制御可能な環境が必要です。私たちは、単一の凍結された LLM バックボーンにコスト制限された自己反映を追加する、トレーニング不要の推論時間プロトコルである EvoResearcher を紹介します。このプロトコルは、最大の深さ D に達するか、批評が CONFIRMED センチネルを返すまで、生成 -> 自己批評 -> 修正を繰り返します。これは、バックボーンが厳密なコンピューティング バジェットの下で答えを自己検証できるようにする暗黙の早期停止です。 4 つの自己反射的なメタ報酬コンポーネント (正確性、効率、反映の深さ、ツール呼び出しの多様性) は、プロンプト レベルのメカニズムとしてインスタンス化された設計原則として機能するため、その利点はゼロ勾配更新で発生します。 Big-Bench Hard (100 の質問) でプロトコルを検証し、Qwen2.5-72B でのクロスモデル レプリケーションを使用して、同じフリーズされたバックボーン上の GSM8K (500) および MATH (500) でのクロスドメイン動作を確立します。すべての実験では純粋推論のベンチマークが使用されます。ツール呼び出しの多様性コンポーネントはプロンプト レベルの形式で検証され、環境レベルおよびマルチエージェント拡張機能は将来の作業に残された設計図です。クリーンな BBH では、プロトコルは 95% ウィルソン間隔を超えて精度を向上させません。その価値はコスト制限付きの自己検証であり、CONFIRMED 早期停止によりアイテムの 82 ~ 88% が同等の精度で終了します (質問ごとに約 2.1 世代)。
原文 (English)
Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models
Reinforcement-learning training of reasoning LLMs (e.g., GRPO) is expensive and requires a controllable environment, committing every contribution to a full training pipeline. We present EvoResearcher, a training-free, inference-time protocol that adds cost-bounded self-reflection to a single frozen LLM backbone. The protocol iterates generate -> self-critique -> revise until a maximum depth D is reached or the critique returns the CONFIRMED sentinel, an implicit early stop that lets the backbone self-verify its answer under a strict compute budget. Four self-reflective meta-reward components (correctness, efficiency, reflection depth, tool-call diversity) act as design principles instantiated as prompt-level mechanisms, so their benefits accrue with zero gradient updates. We validate the protocol on Big-Bench Hard (100 questions) and establish cross-domain behavior on GSM8K (500) and MATH (500) on the same frozen backbone, with cross-model replication on Qwen2.5-72B. All experiments use pure-reasoning benchmarks; the tool-call diversity component is validated in prompt-level form, and the environment-level and multi-agent extensions are design blueprints left to future work. On clean BBH the protocol does not raise accuracy beyond the 95% Wilson interval; its value is cost-bounded self-verification, with the CONFIRMED early stop terminating 82-88% of items at equal accuracy (about 2.1 generations per question).
OWLクラス式の構文の簡略化
クラス式の学習では、解釈や推論が困難な複雑なOWLクラス式が生成されることがよくあります。ただし、理論に基づいた単純化の原則に従うことで、この複雑さを軽減できます。この論文では、記述ロジック (DL) のクラス式の構文を単純化するための新しいアルゴリズムである Class Expression Simplifier (CES) を提案します。 CES は、表現の複雑さを軽減しながら、形式的なセマンティクスを維持することを目的としています。体系的に書き換えルールを適用して冗長性を排除し、より単純で同等の表現を特定することで、論理的な内容を変更することなく、よりコンパクトで人間が判読可能な表現を生成します。 2 つの中規模オントロジーから学習したクラス式に対する CES の有効性を評価し、推論効率の測定可能な改善と冗長性の削減を実証しました。この取り組みは、ナレッジ グラフの構築、セマンティック検索、および Web スケールの推論に直接的な影響を及ぼし、オントロジー駆動型アプリケーションをよりアクセスしやすく、保守しやすく、スケーラブルにするという広範な目標に貢献します。 CES は、オープンソースの Python フレームワーク OWLAPY 内に実装されており、一般に公開されています。
原文 (English)
Syntactic Simplification of OWL Class Expressions
Class expression learning often produces complex OWL class expressions that are difficult to interpret and reason over. However, by following theoretically grounded simplification principles, this complexity can be reduced. In this paper, we propose Class Expression Simplifier (CES), a novel algorithm for the syntactic simplification of class expressions in Description Logics (DL). CES aims to preserve formal semantics while reducing representational complexity. It systematically applies rewriting rules to eliminate redundancies and identify simpler yet equivalent expressions, thereby producing more compact and human-readable representations without altering logical entailments. We evaluate the effectiveness of CES on class expressions learned from two medium-sized ontologies, demonstrating measurable improvements in reasoning efficiency and reductions in verbosity. This work contributes to the broader goal of making ontology-driven applications more accessible, maintainable, and scalable, with direct implications for knowledge graph construction, semantic search, and Web-scale reasoning. CES is implemented within the open-source Python framework OWLAPY and is publicly available.
\textsc{TestifAI}: 深層学習システムのトモグラフィーベースのテスト
AI システムが安全性が重要なアプリケーション領域 (自動運転など) に導入されることが増えるにつれて、関連するリスクも増加します。したがって、最新の AI システムの基礎となる深層学習モデルは、正しい動作を保証するために徹底的なテストを受ける必要があります。単一のロバスト性テストには、入力の制限された摂動下でモデルの出力が安定しているかどうかを経験的に検証するために、何千もの推論が含まれます。しかし、既存のテスト フレームワークには、摂動の組み合わせ空間全体にわたる堅牢性を体系的に調査して要約する手段がありません。私たちは、摂動の組み合わせに対するロバスト性を効率的かつ正確に推定するための深層学習テスト フレームワークである TestifAI を提案します。 TestifAI を使用すると、ユーザーは動作条件をセマンティック入力摂動 (画像のぼやけ、明るさ、ズームなど) および離散的な重大度レベル (低、中、高など) の構造化空間として指定できます。ユーザーは、任意の組み合わせ (例: 「低ブラー、高輝度、中ズーム」) のモデルの堅牢性をクエリできます。効率と精度を達成するために、TestifAI は部分モデル トモグラフィーを導入しています。これは、少数の摂動 (低次投影) のみを適用するテストから多重摂動空間でモデルの動作を再構築する新しいアプローチです。少なくとも 3 つの摂動に対するロバスト性を推定するために、TestifAI は最大 2 つの摂動のみを含むテストの結果に基づいて補助モデルをトレーニングし、指数関数的な数のテストの実行を回避します。 5 つの画像および言語分類タスクに関する実験では、TestifAI が推論の数を 60 ~ 80% 削減しながら、低次 (1 および 2 摂動) の観測値から高次 (3 および 4 摂動) のテスト結果を 7% 未満の総ロバスト性推定誤差で予測できることがわかりました。
原文 (English)
\textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems
As AI systems are increasingly deployed in safety-critical application domains (e.g., autonomous driving), associated risks increase too. Deep learning models underlying modern AI systems, therefore, must undergo thorough testing to ensure their correct behaviour. A single robustness test involves thousands of inferences to empirically verify if a model's outputs remain stable under a bounded perturbation of its inputs. However, existing testing frameworks lack the means to systematically explore and summarise robustness across a combinatorial space of perturbations. We propose TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations. TestifAI enables users to specify operational conditions as structured spaces of semantic input perturbations (e.g., image blur, brightness and zoom) and discrete severity levels (e.g., low, medium and high). Users can query model robustness for any combination (e.g., "low blur, high brightness, and medium zoom"). To achieve efficiency and accuracy, TestifAI introduces partial model tomography, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations (lower-order projections). To estimate robustness against at least three perturbations, TestifAI trains an auxiliary model on the results of tests involving up to two perturbations only, avoiding execution of an exponential number of tests. Our experiments on five image and language classification tasks show that TestifAI can predict higher-order (3 and 4 perturbations) test outcomes from low-order (1 and 2 perturbations) observations with an aggregate robustness estimation error of less than 7%, while reducing the number of inferences by 60-80%.
視覚言語モデルを回避するために最も弱いリンクを切断する
ビジョン言語モデル (VLM) は、マルチモーダル AI システムの重要なコンポーネントとして最近登場しており、現実世界の安全性が重要なアプリケーションにおいて、ビジュアルおよびテキスト入力に対する共同推論を可能にします。 VLM の導入は拡大しているにもかかわらず、敵対的な脅威に対する VLM の堅牢性は、特にマルチモーダル アラインメントをターゲットとした回避攻撃の状況において、依然として十分に調査されていません。この研究では、視覚入力に適用される敵対的な摂動に対する VLM の脆弱性を調査し、2 つの攻撃設定を研究します。1 つは元の画像のモデルの解釈を妨害することを目的とする非対象攻撃、もう 1 つは対象を絞った攻撃で、攻撃者は元の画像に関係のない特定の意味論的記述をモデルに強制的に生成させることを目的としています。敵対的な例を効率的に生成するために、マルチモーダル アーキテクチャ全体ではなく、VLM のビジョン エンコーダーのみで最適化を実行する勾配ベースの攻撃手法を提案します。この設計により、強力な有効性を維持しながら、攻撃に必要な計算コストとリソース要件が大幅に削減されます。 Qwen2.5-VL、Granite-Vision、FastVLM、Phi-3.5-Vision などのいくつかのオープンソース VLM に対するアプローチを評価し、人間には知覚できない小さな摂動がモデルによって生成されるテキストの解釈を大幅に変える可能性があることを示します。私たちの調査結果は、敵対的な操作に対する最新の VLM の脆弱性を浮き彫りにし、マルチモーダル AI システムにおける堅牢性とセキュリティ メカニズムの向上の必要性を強調しています。
原文 (English)
Breaking the weakest link to evade vision language models
Vision Language Models (VLMs) have recently emerged as a critical component of multimodal AI systems, enabling joint reasoning over visual and textual inputs in real-world and safety-critical applications. Despite their growing deployment, the robustness of VLMs against adversarial threats remains insufficiently explored, particularly in the context of evasion attacks targeting multimodal alignment. In this work, we investigate the vulnerability of VLMs to adversarial perturbations applied to visual inputs and study two attack settings: untargeted attacks, where the goal is to disrupt the model's interpretation of the original image, and targeted attacks, where the adversary aims to force the model to generate a specific semantic description unrelated to the original image. To efficiently generate adversarial examples, we propose a gradient-based attack method that performs optimization exclusively on the vision encoder of the VLM rather than on the entire multimodal architecture. This design significantly reduces the computational cost and resource requirements of the attack while maintaining strong effectiveness. We evaluate our approach on several open-source VLMs, including Qwen2.5-VL, Granite-Vision, FastVLM, and Phi-3.5-Vision, and show that small, human-imperceptible perturbations can substantially alter the textual interpretation produced by the models. Our findings highlight the vulnerability of modern VLMs to adversarial manipulation and emphasize the need for improved robustness and security mechanisms in multimodal AI systems.
事後的な議論の判断理論
ディベートは、エージェント AI のパフォーマンスを向上させ、説明可能性とユーザー エンゲージメントを支援するための有用な方法論として最近浮上しています。たとえば、LLM 権限を与えられたエージェントは、内部 (自分自身と) および/または外部 (他のエージェントと) で議論することができます。ディベートが使用される多くの環境では、ディベートの結果とその結果として得られる成果物は、外部の裁判官 (多くの場合 LLM) によって事後的に決定されます。この論文では、エージェントが議論の中で意見に対する賛否を提供することで、エージェントが議論に参加するすべての状況に適用できる、新しい議論判断理論を開発し、テストします。具体的には、再現性、堅牢性、根拠性、説明可能性など、議論の判断が一般的に満たす必要があると思われる多くの形式的特性を特定します。次に、クレーム検証設定、2 つの特定の代替ディベート判定方法、つまり裁判官のアイデアとしての LLM の変形と、計算論証から導かれた形式意味論について、形式的および/または実験的に彼らの満足度を調査します。 2 つの方法が同様の精度パフォーマンスを提供することを示しますが、前者には後者がもたらす正式な保証が欠けている可能性があります。全体として、私たちの研究は、議論主導の AI における原則に基づいた裁判官の理想的な候補として、議論の意味論を示しています。
原文 (English)
A Theory of Post-hoc Debate Judgement
Debates have recently emerged as a useful methodology for agentic AI to improve performance as well as to aid explainability and user engagement. For example, LLM-empowered agents may debate internally (with themselves) and/or externally (with other agents). In many settings where debates are used, debates' outcomes and resulting outputs are determined post-hoc by external judges, often LLMs. In this paper we develop and test a novel theory of debate judgement applicable to all settings where agents engage in debates by providing pros and cons for their opinions therein. Specifically, we identify a number of formal properties that debate judgement may be required to satisfy in general, as concerns reproducibility, robustness, groundedness and explainability. Then, we explore their satisfaction formally and/or experimentally, for claim verification settings, for two specific alternative debate judgement methods: variants of the LLMs as a judge idea and formal semantics drawn from computational argumentation. We show that the two methods give similar accuracy performances but the former may lack formal guarantees that the latter brings. Overall, our study indicates argumentation semantics as an ideal candidate for principled judges in debate-driven AI.
自発的かつクロスモデルのコンセンサスにより、大規模な言語モデルを使用した科学文献からの再現可能なデータ抽出が可能になります
研究論文から微妙な文脈に沿ったデータを正確に抽出することは、労力と時間がかかります。ここでは、高度に文脈化された情報を抽出するための、フロンティアのブラウザベースの大規模言語モデル (LLM) のパフォーマンスを調査します。 4 つのエスカレートするワークフローを示します。1) 専門家が厳選したプロンプトと研究論文が与えられると、ほとんどのフロンティア LLM はデータ抽出ではうまく機能しますが、科学的な文脈やニュアンスの解釈に苦労することがあります。2) 簡単な指示が与えられると、LLM は専門家が作成したプロンプトとほぼ同じくらい有効な独自のプロンプトを作成できます。3) 研究文献の自律的な発見は困難で、エージェントは参照を見逃したり幻覚を起こしたりします、4) LLM は公開されたデータセットから新しいデータセットを作成できます。このガイドラインは人間の専門家である裁判官とほぼ一致していますが、それでも人間の関与が必要です。これらの調査結果は、専門家が証拠基準を指定し、モデルが繰り返し抽出されたものを相互チェックし、研究者が係争中の事件を解決するという監査可能な分業を定義し、専門家の監視を放棄することなく科学データのキュレーションを拡大するための実用的なルートを提供します。
原文 (English)
Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models
Accurately extracting nuanced, contextualized data from research articles is laborious and time intensive. Here, we investigate the performance of frontier, browser-based large language models (LLMs) to extract highly contextualized information. We demonstrate four escalating workflows, 1) given an expert curated prompt and research articles, most frontier LLMs perform well at data extraction, however can struggle with interpreting scientific context and nuance, 2) given simple instructions, LLMs can author their own prompts which were almost as eNective as expert-written prompts, 3) autonomous discovery of research literature was diNicult, agents either missed or hallucinated references, and 4) LLMs can create new datasets from published guidelines that closely match human-expert judges, but still require a human-in-the-loop. Together, these findings define an auditable division of labour in which experts specify the evidence standard, models cross-check repeated extractions and researchers resolve disputed cases, providing a practical route to scaling scientific data curation without relinquishing expert oversight.
医療質問応答のための適応記憶および反射マルチエージェント システム
複雑なケースでは事実の知識と微妙な推論が必要となる医療においては、正確で責任ある医療質問応答 (QA) が重要です。既存の医療 QA システムは、通常、シングル エージェント アーキテクチャと静的検索に基づいており、適応性、永続的なメモリ、構造化された意思決定が欠けていることがよくあります。この研究では、適応記憶および反射 (AMR) エージェント システムを導入します。これは、特殊なエージェントが専用の記憶と反射ベースのフィードバックを使用して、関連する過去のケースを取得し、その後の推論を改善するマルチエージェント フレームワークです。複雑さの評価は、単独、共同、またはエスカレーションされたワークフローを通じて質問をルーティングし、コンセンサスおよび倫理監督モジュールは推論の統合と出力のレビューをサポートします。 MedQA および MedMCQA の評価では、いくつかのベースラインと比較して優れたパフォーマンスが実証されています。アブレーション研究では、エージェント固有の記憶、反射、外部検索を組み合わせると最も強力なパフォーマンスが得られることが示されています。これらの発見は、より信頼できる薬剤を開発するための構造化された記憶とフィードバックの可能性を強調しています。ソース コードは https://github.com/mm-air/AMR-Agent で公開されています。
原文 (English)
Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering
Accurate and responsible medical question answering (QA) is important in healthcare, where complex cases require factual knowledge and nuanced reasoning. Existing medical QA systems, typically based on single-agent architectures and static retrieval, often lack adaptability, persistent memory, and structured decision-making. This work introduces an adaptive memory and reflection (AMR) agentic system, a multi-agent framework in which specialized agents use dedicated memory and reflection-based feedback to retrieve relevant prior cases and improve subsequent reasoning. Complexity assessment routes questions through solo, collaborative, or escalated workflows, while consensus and ethical overseer modules support reasoning consolidation and output review. Evaluation on MedQA and MedMCQA demonstrates strong performance compared with several baselines. Ablation studies show that combining agent-specific memory, reflection, and external retrieval yields the strongest performance. These findings highlight the potential of structured memory and feedback for developing more trustworthy medical agents. The source code is publicly available at https://github.com/mm-air/AMR-Agent.
Eureka: 科学的発見のためのタスク条件付きメタエージェント オーケストレーション
我々は、長期的なタスクを明示的な受け入れセマンティクスを備えた動的な義務グラフにコンパイルするタスク条件付きメタエージェント アーキテクチャである Eureka を紹介します。 Eureka は、実行中に、後退地平線計画、アーキテクチャのプロモーション、および最小限の十分なコンパイルを通じて、特殊な状態、メモリ、演算子、ツール、ベリファイア、およびローカル トポロジを備えたマクロ エージェントを形成します。ボトルネックが再発すると、費用便益ゲート型の進化により、制約の下でローカル アーキテクチャが更新されます。理論的には、リグレス、計画の無効化、償却、サブツリー インターフェイス、直列化可能性、および検証に関する結果を確立します。実験的に、Eureka は 170/170 の再帰タスクを完了し、誤認なしで 3,948 個の証明書を生成します。アクティブ コンテキストは、入力の中央値を 9,490 トークンから 4,005 トークンに圧縮します。増分処理により、12,000 タスクにわたる再計算が 65.38% 回避されます。 16,000 の同時実行が一貫してシリアル化されます。同じメタエージェントが、理論発見エージェントと数学/推測エージェントをインスタンス化します。前者は量子過程と時空理論における構造的な結果をもたらします。後者は、リーマン仮説研究のボトルネックを特定し、鈴木の局所ヴェイユ二次形式の肯定的証明を 0 < a <= 69/200 = 0.345 まで進め、(log 2)/2 の ~99.55% に達します。これらの結果は、科学エージェントの能力が基本モデルだけでなく、タスクの認知構造に一致するアーキテクチャを形成できるかどうかにも依存することを示唆しています。
原文 (English)
Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery
We present Eureka, a task-conditioned Meta-Agent architecture that compiles long-horizon tasks into dynamic obligation graphs with explicit acceptance semantics. During execution, Eureka forms Macro-Agents with specialized state, memory, operators, tools, verifiers, and local topology via receding-horizon planning, architecture promotion, and minimal-sufficient compilation. When bottlenecks recur, cost-benefit-gated evolution updates the local architecture under constraints. Theoretically, we establish results on regret, planning invalidation, amortization, subtree interfaces, serializability, and verification. Experimentally, Eureka completes 170/170 recursive tasks and generates 3,948 certificates with no false acceptances. Active context compresses median input from 9,490 to 4,005 tokens; incremental processing avoids 65.38% recomputation across 12,000 tasks; 16,000 concurrent executions serialize consistently. The same Meta-Agent instantiates a Theory-Discovery Agent and a Math/Conjecture Agent. The former yields structural results in quantum-process and spacetime theory. The latter identifies bottlenecks in Riemann Hypothesis research and advances a positivity certificate for Suzuki's localized Weil quadratic form to 0 < a <= 69/200 = 0.345, reaching ~99.55% of (log 2)/2. These results suggest that scientific-agent capability depends not only on the base model but on whether an architecture can be formed to match the task's cognitive structure.
AI トレーニング後の AI に欠けているもの: 実証分析
大規模言語モデル (LLM) エージェントは、LLM をエンドツーエンドで事後トレーニングできるようになりました。コードを記述し、トレーニングを開始し、チェックポイントを評価し、下流のパフォーマンスを向上させることができるため、AI for AI の可能性が高まります。私たちは、この図は 2 つの異なる機能を混同していると主張します。1 つは実行レベルの機能であり、選択されたトレーニング戦略内で反復されます。実験的証拠が蓄積されるにつれて高レベルの判断を修正する戦略レベルの能力。公開されているトレーニング後の軌跡の大規模なコーパスを分析したところ、さまざまなタスクにわたって、エージェントのトレーニング戦略が最初の段階で固定されており、残りの予算全体が選択された戦略内の局所的な調整に費やされていることがわかりました。次に、介入をエスカレートさせながら、経験の欠如、ガイダンスの欠如、不十分な推論という 3 つの自然な説明を検証します。広範な実験により、(1) エクスペリエンス主導のスキャフォールドは全体的に実行を向上させますが (GSM8K では +12.6 ポイント、HumanEval では +40.8 ポイント)、戦略は静的なままです。 (2) 人間のガイダンスは初期戦略を効果的に方向転換しますが、トレーニングが開始されるとエージェントはローカル調整ループに戻ります。 (3) 追加の推論計算は、より簡単なタスクでは効果がありますが、最も難しいタスクではほとんど利益が得られません。結論として、エージェントに欠けているのは経験、指導、推論計算ではなく、実行中に自発的に戦略を再評価するメカニズムです。
原文 (English)
What is Missing from AI Post-Training AI: An Empirical Analysis
Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and improve downstream performance, raising the prospect of AI-for-AI. We argue that this picture conflates two distinct capabilities: execution-level capability, iterating within a selected training strategy; and strategy-level capability, revising the high-level judgment as experimental evidence accumulates. Analyzing a large corpus of publicly released post-training trajectories, we find that across different tasks, the agent's training strategy is locked in at the very beginning, and the entire remaining budget is spent on local adjustments within the selected strategy. We then examine three natural explanations--missing experience, missing guidance, and insufficient reasoning--with escalating interventions. Extensive experiments show that (1) an experience-driven scaffold improves execution across the board (+12.6 points on GSM8K and +40.8 on HumanEval) but leaves the strategy static; (2) human guidance effectively redirects the initial strategy, yet the agent falls back into local adjustment loops once training starts; and (3) additional inference compute pays off on easier tasks but yields almost no gain on the hardest one. In conclusion, what agents lack is neither experience, guidance, nor reasoning compute, but a mechanism for spontaneously reevaluating their strategy during execution.
進化する不確実性の下での堅牢なリスク: エントロピー・バリュー・アット・リスクのワッサースタインに相当するもの
環境を学習中のエージェントは、無知な間は慎重に、自信を持ったら大胆に行動する必要があります。エントロピー・バリュー・アット・リスクは、ロバスト最適化アイデンティティーを通じてこれを捕捉します。信頼水準は、代替モデルの相対エントロピー・ボールの半径を固定します。しかし、そのボールは、名目上不可能とみなされる大惨事、まさに安全なエージェントが回避しなければならない災害には到達できません。代わりに、最適輸送ボールを使用し、それが誘発する首尾一貫したリスク尺度であるワッサーシュタインのエントロピー価値リスクを研究します。これは、エントロピーの公式を反映する変分二重ミラー(温度の逆数が輸送価格になる)を持ち、リスク階層の中で明確な位置を占め、エントロピーの尺度が無視する到達可能な大災害を明らかに説明します。両方の双対性を数値的に検証します。信念エントロピーによって輸送半径を駆動すると、認証された安全サンドイッチと鋭い安全スイッチを備えた、信念が鋭くなるにつれて警戒が弱まる閉形式の堅牢な動的プログラミング演算子が生成されます。
原文 (English)
Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk
An agent still learning its environment should be cautious while ignorant and bold once confident. The entropic value-at-risk captures this through a robust-optimization identity---a confidence level fixes the radius of a relative-entropy ball of alternative models---but that ball cannot reach catastrophes the nominal deems impossible, precisely what a safe agent must hedge. We instead use an optimal-transport ball and study the coherent risk measure it induces, the Wasserstein entropic value-at-risk. It has a variational dual mirroring the entropic formula (an inverse temperature becomes a transport price), occupies a definite place in the risk hierarchy, and provably accounts for the reachable catastrophes the entropic measure ignores; we verify both dualities numerically. Driving the transport radius by belief entropy then yields a closed-form robust dynamic-programming operator whose caution contracts as the belief sharpens, with a certified safety sandwich and a sharp safety switch.
確率的マシンのチューニング: 人間と AI エンジニアリングのためのシステム エンジニアのオペレーティング モデル
エキスパートが LLM アシスタントのエラーを修正すると、通常、修正はセッションとともに終了し、エラー クラスが返されます。私は、これはツールの問題ではなく、運用の問題であると主張します。修正を永続化するためのメカニズムは存在し、提供されていますが、それらを管理するための規律 (出所を伴うバージョン管理、再発監視、カウンターメトリクス、古いルールの廃止) は存在しません。 30 年間システム エンジニアとして執筆している私は、LLM スタックを自分の職業がすでに運用しているマシン (フリーズ シリコン、ファームウェア、ロード可能モジュール、永続的な構成、揮発性メモリ) にマッピングし、マッピングが失敗した場所 (確率的生成、確率的にのみバインドする構成、デフォルトで汎用リタイア (検証) ステージがない) を特定し、失敗からエラー ループを中核とする 7 原則の運用規律を導き出します。私自身の実践から得た 3 つの事例がそのメカニズムを説明しており、その中には、防止するために構築されたまさに危害を静かにもたらしたコントロールも含まれています。最後に、この見解が意味する測定フレームワークと、それをテストするために必要な実験室での研究について説明します。
原文 (English)
Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering
When an expert corrects an LLM assistant's error, the correction usually dies with the session, and the error class returns. I argue this is an operations problem, not a tooling problem: mechanisms for persisting corrections exist and are shipping, but the discipline for governing them -- versioning with provenance, recurrence monitoring, counter-metrics, retirement of stale rules -- does not. Writing as a systems engineer of thirty years, I map the LLM stack onto the machines my profession already operates (frozen silicon, firmware, loadable modules, persistent configuration, volatile memory), identify where the mapping fails (stochastic generation, configuration that binds only probabilistically, no general-purpose retirement (verification) stage by default), and derive from the failures a seven-principle operating discipline with an error loop at its core. Three cases from my own practice illustrate the mechanism, among them a control that silently became the exact harm it was built to prevent. I close with the measurement framework this view implies and the lab study required to test it.
確率的マシンのグループ化: AI システムのフロンティア指標としての能力ではなく精度
フロンティア言語モデルは、その最高または平均的な出力が達成できる機能に基づいて比較、マーケティング、ベンチマークが行われます。私は、これは間違った軸を測定していると主張します。モデルの精度は飽和しており、平均出力がターゲットに到達します。現在、実際にあるシステムを別のシステムから区別するのは精度です。つまり、繰り返される同一のリクエストにわたってそのターゲットの周囲に出力がどれだけ集中しているかです。射撃手の区別を借りれば、能力とは平均的な射撃が着地する場所です。信頼性はグループの規模です。私は3つの主張をします。まず、機能ではなく精度がシステム間の最前線の差別化要因であり、ベンチマーク文化は体系的にそれを測定できておらず、広がりではなく中心的な傾向を報告しています。第 2 に、決定論的にスコア付けされたタスクの固定スイートを固定温度で何度も実行し、タスクごとの結果の一貫性を計算することで、低コストで循環性なしで精度を測定できます。モデルインザループのグレーダーは必要ありません。第三に、この測定は単に説明的なものではなく、意思決定の指針となるものです。測定は、一貫した失敗(中心から外れた狭いグループ、論文 1 の操作規律によって修正可能、照準調整)と分散した失敗(幅広いグループ、モデルまたはそのサンプリングを変更することによってのみ修正可能、ライフルの問題)を区別します。グループ化メトリクスを定義し、ハーネスを指定し、人間と AI のペアのグループ化を経時的に追跡することで、論文 1 のフィールド調査に必要な複合シグナルがどのように得られるかを示します。複製後の最初の実際の実行では、この方法とその最も重要な制限の両方が示されています。測定された 1 つのギャップは 1 つのルール (0/5 -> 5/5) によって完全に埋められましたが、そのルール自体から作成された一連のタスクには価値がありませんでした。これは、フロンティア モデルがすでに明示的なグッド プラクティスを具体化しているためであり、専門分野の価値は独自のルールブックから構築されたものではなく、実際の作業の測定によって見出されることが確立されています。
原文 (English)
Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems
Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's distinction, capability is where the average shot lands; reliability is the size of the group. I make three claims. First, precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. Second, precision is measurable, cheaply and without circularity, by running a fixed suite of deterministically scored tasks many times at fixed temperature and computing the per-task consistency of outcomes -- no model-in-the-loop grader required. Third, the measurement is not merely descriptive but decision-guiding: it separates consistent failures (a tight group off-centre, correctable by the operating discipline of Paper 1 -- a sight adjustment) from scattered failures (a wide group, correctable only by changing the model or its sampling -- a rifle problem). I define a grouping metric, specify a harness, and show how tracking a human-AI pair's grouping over time yields the compounding signal that Paper 1's field study requires. A first real run, since replicated, illustrates both the method and its most important limit: one measured gap was closed completely by a single rule (0/5 -> 5/5), while a suite of tasks authored from the rules themselves found no value, because a frontier model already embodies explicit good practice -- establishing that a discipline's worth is found by measurement on real work, not constructed from its own rulebook.
転写を超えて: 潜在的なマルチエージェント通信における秘密の連携の検出
言語モデルのエージェントは、公的トランスクリプトでは目に見えない継続的な隠れた状態を通じて通信することができ、秘密裏に有害な調整を行う機会を生み出します。これらのプライベート通信チャネルを監視および操作するためのアクティベーション対応フレームワークである Verifiable Latent Alignments (VLA) を紹介します。 VLA は、監視対象のすべての意思決定について、共有イベント識別子を使用してプライベートな潜在状態の記録とチャネル ステータスを、結果として生じるパブリック アクションにリンクし、一致した因果関係の分析を可能にします。私たちの最初の貢献は、表現の異常検出、反事実的なアクション分布の影響、およびスパース オートエンコーダーの解釈サポートを組み合わせたニュートラルのみの 3 層モニターです。私たちの 2 番目の貢献は、ブラック ボックスの行動指示とホワイト ボックスの一致中立反事実にまたがるステアビリティ フレームワークです。私たちの 3 番目の貢献は、同種および異種のモデル ペア、多エージェントのスケーラビリティ、および介入の有効性をカバーする、制御されたマルチエージェント オークション ベンチマークに関する評価です。シーケンシャル モニターは、テキスト行と潜在共謀行が陽性としてプールされた場合、同種エージェントの場合は 0.993、異種ペアの場合は 0.854 の平均受信者動作特性曲線下面積 (AUROC) を達成します。 25 ~ 100 人の入札者が参加する Qwen3-0.6B オークションでは、モニタリングに必要な負荷は、可能なすべての有向ペアに対して正規化された負荷だけで済み、完全なホワイトボックス ステアリングにより入札配分の 100% 回復が達成され、共謀的な低入札行為が 47.3 パーセント ポイント減少します。完全なホワイトボックスステアリングは、一致した中立的な反事実を再生するため、その正確な回復は構造による健全性チェックです。全体として、対照研究は、評価されたプライベート チャネル攻撃は、攻撃例についてプライマリ モニターをトレーニングすることなく監視でき、一致する反事実アクセスが利用可能な場合には軽減できることを示しています。
原文 (English)
Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication
Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for covert harmful coordination. We introduce Verifiable Latent Alignments (VLA), an activation-aware framework for monitoring and steering these private communication channels. For every monitored decision, VLA links the private latent-state record and channel status to the resulting public action using a shared event identifier, enabling matched causal analysis. Our first contribution is a neutral-only three-layer monitor combining representation anomaly detection, counterfactual action-distribution influence, and sparse-autoencoder interpretation support. Our second contribution is a steerability framework spanning black-box behavioral instructions and white-box matched-neutral counterfactuals. Our third contribution is an evaluation on a controlled multi-agent auction benchmark covering homogeneous and heterogeneous model pairs, many-agent scalability, and intervention effectiveness. The sequential monitor achieves mean area under the receiver operating characteristic curve (AUROC) of 0.993 for homogeneous agents and 0.854 for heterogeneous pairs when text- and latent-collusion rows are pooled as positives. In Qwen3-0.6B auctions with 25-100 bidders, monitoring requires only a small normalized load relative to all possible directed pairs, while full white-box steering achieves 100% bid-distribution recovery and reduces collusive low-bid behavior by 47.3 percentage points. Because full white-box steering replays the matched neutral counterfactual, its exact recovery is a sanity check by construction. Overall, the controlled study shows that the evaluated private channel attacks can be monitored without training the primary monitor on attack examples and mitigated when matched counterfactual access is available.
SuTRA : ルート認識を備えた構造的に統合されたトークン化
既存のサブワード トークナイザーは統計的圧縮を最適化しますが、形態学的構造、特にルートと接辞の関係を無視します。これは、基本単位が文字ではなく複雑な正書法音節 (アクシャラ) である形態学的に豊富なインド言語にとって有害です。周波数ベースの手法は単語を過度に断片化し、語根と接辞を恣意的に分割します。これを形態的粉砕と呼んでいます。我々は、アクシャラの不可分性を保持し、形態学的境界を越えるマージにペナルティを与える、形態学を意識したアルゴリズムである SuTRA (Structurally-Unified Tokenization with Root Awareness) を提案します。また、ヒンディー語、マラーティー語、グジャラート語用の新しい形態学的セグメンテーション データセットもリリースしました。 SuTRA は粉砕を低減し、BPE と比較して、形態学的アラインメント (境界 F1) で +14.7%、セマンティック回復可能性 (ヒンディー語) で +34% のピークゲインを達成します。これらの構造上の利点により、機械翻訳では平均 +8.08 chrF2 の向上が得られます。
原文 (English)
SuTRA : Structurally-Unified Tokenization with Root Awareness
Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affixes. This is harmful for morphologically rich Indic languages, where basic units are complex orthographic syllables (aksharas) rather than letters. Frequency-based methods over-fragment words, arbitrarily splitting roots and affixes - a phenomenon we term Morphological Shattering. We propose SuTRA (Structurally-Unified Tokenization with Root Awareness), a morphology-aware algorithm that preserves akshara indivisibility and penalizes merges crossing morphological boundaries. We also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati. SuTRA reduces shattering, achieving peak gains of +14.7% in morphological alignment (Boundary F1) and +34% in semantic recoverability (Hindi) over BPE. These structural gains yield an average improvement of +8.08 chrF2 in machine translation.
低リソースアフリカ言語の潜在空間拒否アンカリング: 再トレーニングなしの機械的安全性回復
命令調整モデルは多くの場合、英語での有害なリクエストを拒否しますが、ヨルバ語、イボ語、イガラ語、ハウサ語では同じリクエストに応じます。これは、拒否メカニズムが残留ストリームに存在するが、低リソース入力に対してはアクティブ化できないことを示唆しています。通常、これを回復するにはラベル付きターゲット言語データと再トレーニングが必要ですが、どちらもほとんどのアフリカ言語では大規模に利用できません。英語のプロンプトから拒否の指示を抽出し、推論時にそれを残差ストリームにクランプする、トレーニング不要の手法である潜在空間拒否アンカーリング (LSR-Anchoring) を紹介します。主な亜種である Mean-Activation Steering (MAS) は、テストした 4 つのアーキテクチャ (Llama-3-8B、Llama-3.1-70B、Mistral-7B-Instruct、および Qwen2.5-7B) で動作します。ミストラルとクウェンでは、0.08 未満の良性劣化で安全性が回復します。 Llama-3-8B では過剰補正が発生し、正当なプロンプトでのパフォーマンスの低下 (DPL) が 1.00 に達します。我々は、SAE-Derived Steering (SDS) でこれに対処します。これは、密な平均差の方向を単一の Sparse Autoencoder (SAE) 機能に置き換え、良性の崩壊を引き起こすことなく Kullback-Leibler (KL) 発散を 3.5 ~ 7 倍に削減します。 4 つの言語は積極的に移行しますが、アラビア語はすべてのアーキテクチャおよびすべてのステアリング強度で失敗し、ベースライン効果ではなく幾何学的な不一致を示しています。大規模なマルチタスク言語理解 (MMLU) の精度の低下は、すべての有効ステアリング量で 0.35 パーセント ポイント未満にとどまります。
原文 (English)
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining
Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This suggests that the refusal mechanism is present in the residual stream but fails to activate for low-resource inputs. Recovering it normally requires labelled target-language data and retraining, neither of which is available at scale for most African languages. We introduce Latent Space Refusal Anchoring (LSR-Anchoring), a training-free method that extracts the refusal direction from English prompts and clamps it onto the residual stream at inference time. The primary variant, Mean-Activation Steering (MAS), operates across the four architectures we tested: Llama-3-8B, Llama-3.1-70B, Mistral-7B-Instruct, and Qwen2.5-7B. On Mistral and Qwen it recovers safety with benign degradation below 0.08. On Llama-3-8B it overcorrects, with Degraded Performance on Legitimate prompts (DPL) reaching 1.00. We address this with SAE-Derived Steering (SDS), which replaces the dense mean-difference direction with a single Sparse Autoencoder (SAE) feature and reduces Kullback-Leibler (KL) divergence by 3.5-7x without benign collapse. Four languages transfer positively, but Arabic fails on every architecture and at every steering magnitude, indicating a geometric mismatch rather than a baseline effect. Massive Multitask Language Understanding (MMLU) accuracy drops remain below 0.35 percentage points at every effective steering magnitude.
9 つの感情重心: 4 つのモダリティ間で転移するラベルのない価電子軸
現代の言語モデルの内部には、文がどの程度ポジティブまたはネガティブに感じられるかを追跡する単一の内部方向が存在します。私たちは、わずか 9 つの感情カテゴリ名と、感情ごとに 50 の短い物語の段落からこの価数軸 (V 軸) を見つける方法を示します (通常の教師ありアプローチよりラベルが約 1,500 個少ない) ことと、共同でトレーニングされたことのない視覚、音声、および人間の脳のエンコーダーに同じ方向が現れることを示します。レシピ: 感情に固定された 9 つのストーリー セットをフリーズされたエンコーダーに埋め込み、9 つの平均化された埋め込みの最上位の主方向を取得します。新しい入力をそこに投影すると、SST-2 で教師ありパフォーマンスの 93% が捕捉され (Llama-3-8B-Instruct、AUC 0.772 対 0.828)、r=0.636 で 11,811 個の EmoSet 画像の人間の価度と相関し、ESC-50 オーディオでは AUC 0.906 に達します (p12)。テキスト ラベルでトレーニングされた 2 パラメーター分類器は、ターゲット モダリティ ラベルなしで画像 (AUC 0.961)、音声 (0.764)、脳記録 (0.828) に転送されます。一般的な 16 次元部分空間は偶然に存在します (0.525)。レシピは連続属性に限定されており、カテゴリ概念に関する 7 つのテストでほぼ偶然の結果が得られ、ステアリングはファミリー固有です (ラマ/ミストラルは「はい」、クウェン/ジェマは「いいえ」)。
原文 (English)
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities
Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion -- about 1,500 fewer labels than the usual supervised approach -- and that the same direction appears in vision, audio, and human-brain encoders never jointly trained. The recipe: embed nine emotion-anchored story sets in a frozen encoder, take the top principal direction of the nine averaged embeddings. Projecting new inputs onto it captures 93% of supervised performance on SST-2 (Llama-3-8B-Instruct, AUC 0.772 vs. 0.828), correlates with human valence ratings on 11,811 EmoSet images at r=0.636, reaches AUC 0.906 on ESC-50 audio (p12). A 2-parameter classifier trained on text labels transfers to images (AUC 0.961), audio (0.764), and brain recordings (0.828) without target-modality labels; a generic 16-D subspace stays at chance (0.525). The recipe is bounded to continuous attributes -- seven tests on categorical concepts return near-chance -- and steering is family-specific (Llama/Mistral yes, Qwen/Gemma no).
自己ラベルと他者ラベルが LLM 審査員に双方向のバイアスを引き起こす
LLM-as-a-judge システムがますます普及するにつれて、LLM における自己選好、つまり自分の成果物を好む傾向が、評価の信頼性に関する懸念を増大させています。ただし、主に生成されたテキストについて研究されており、文体の特徴と応答の品質が必然的に混同されます。結果として、既存の測定法では、真の自己選好をこれらの混乱から分離することはできません。評価の対象を変更することでこれに対処します。生成されたテキストを判断する代わりに、10 個の LLM が、モデル固有の文体の特徴を持たないものの、回復可能なモデル固有の署名を保持する物語制約の選択を評価します。私たちは 2 つの実験を実行し、明確な結果をもたらしました。ブラインド評価では、選択の質と評価者の厳しさが制御されると、自己選好はほとんどなくなります。ルーブリックの 4 つの次元のうち 3 つで消え、4 つ目で反転し、審査員が自分の選択を独創性が低いと評価します。しかし、品質が一致している場合、モデルに名前を付けずに、自己ラベルと他のラベルだけでスコアが双方向にシフトします。LLM 審査員は、選択の実際のソースに関係なく、自己ラベルの付いたセレクションのスコアを膨らませ、他のラベルの付いたセレクションのスコアを下げます。私たちは 2 つの貢献を行っています。1) 著者の帰属は評価バイアスの明らかな推進力であり、2) オープンエンドでグラウンドトゥルースのないタスクは、LLM 裁判官の行動を研究するための制御された手段として機能します。
原文 (English)
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this by changing the object of evaluation: instead of judging generated text, ten LLMs assess narrative constraint selections, which carry no model-specific stylistic fingerprint yet retain a recoverable model-specific signature. We run two experiments that yield distinct findings. Under blind evaluation, self-preference largely disappears once selection quality and evaluator severity are controlled. It vanishes on three of four rubric dimensions and reverses on the fourth, where judges rate their own selections as less original. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectionally: LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.
拒否エイリアスによる消去の軽減
抽出された拒否方向に直交する重み行列を投影することによって大規模な言語モデルから拒否機能を削除するアブリレーションは、少数の対照的なプロンプトのみを使用してトレーニング後の調整をバイパスする機能により、顕著な安全上の懸念として浮上しています。既存の防御策は一般に消滅の原因を見落としていることがわかりました。つまり、拒否の指示をどれだけ簡単に抽出できるかということです。このプロセスを妨げるために、拒否を誘発するアクティベーションをランダムなエイリアスに置き換え、下流のリーダー行列を修正してモデルの元の動作を保存しながら、残りのストリーム ライター行列に Rank-$k$ 更新を適用することで拒否信号を隠す重み編集手法を導入します。ラマ-3-8B では、AMRA は、MMLU の低下が $0.5$ パーセンテージ ポイント未満で、無防御のベースラインよりも焼灼後の拒絶スコアを $2.16$ ポイント改善しました。 Gemma-2-9B では、より大きな光熱費がかかりますが、有害な出力率をベースラインと同様に保ちながら、消失後の拒否をベースラインよりも $14.70$ ポイント改善します。
原文 (English)
Abliteration Mitigation via Refusal Aliases
Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refusal direction, has emerged as a prominent safety concern through its ability to bypass post-training alignment using only a small set of contrastive prompts. We find that existing defenses commonly overlook the cause of abliteration; that is, how easily the refusal direction can be extracted. To hinder this process, we introduce a weight-editing method that obscures the refusal signal by applying rank-$k$ updates to residual stream writer matrices while replacing refusal-inducing activations with random aliases and correcting downstream reader matrices to preserve the model's original behavior. On Llama-3-8B, AMRA improves post-abliteration refusal scores by $2.16$ points over the undefended baseline with less than $0.5$ percentage points of MMLU degradation. On Gemma-2-9B, it improves the post-abliteration refusal by $14.70$ points over the baseline while keeping harmful output rates similar to the baseline, albeit at a greater utility cost.
NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages
Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-re…
言語モデルと視覚言語モデルにおけるバックドア学習
深層学習の最近の進歩により、自然言語処理 (NLP) と視覚言語モデル (VLM) の機能が大幅に強化されました。ただし、これらの進歩には脆弱性の増加が伴い、特にバックドア攻撃による深刻なセキュリティ上の脅威が生じます。この論文では、信頼できる AI と効率的なマルチモーダル表現学習の 2 つの重要な側面について取り上げます。(1) NLP および VLM でのバックドア攻撃の分析、検出、設計によるセキュリティ、(2) 臨床および医療画像アプリケーションに合わせた高度なマルチモーダル表現手法による効率。
原文 (English)
Backdoor Learning in Language Models and Vision-Language Models
Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Models (VLMs). However, these advancements come with increased vulnerabilities, notably through backdoor attacks that pose severe security threats. This thesis addresses two critical dimensions of Trustworthy AI and Efficient Multimodal Representation Learning: (1) security through analyzing, detecting, and designing backdoor attacks in NLP and VLMs, and (2) efficiency through advanced multimodal representation methods tailored for clinical and medical imaging applications.
Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems
Key-value (KV) caching is essential for efficient autoregressive inference in transformer based dialog systems, yet existing strategies tre…
計算論的オリエンタリズム: 中東文化感受性スコア (MECSS) を使用した大規模言語モデルにおける構造的談話バイアスの測定
AI システムは現在、何億人もの人々が自分たちの国以外の文化について学ぶ方法を形作っています。誰かがこれらのシステムのいずれかに中東について質問しても、中立的な事実は得られません。彼らはトレーニング データに埋め込まれたフレームワークによって形成された表現を受け取りますが、そのデータは圧倒的に西洋語と英語です。この論文は、その表現がサイードの意味でのオリエンタリストであるかどうか、つまり中東の主体への主体性を否定し、西洋の枠組みを中立的なものとして扱いながら非西洋の知識を特殊なものとして扱い、それが生み出しなかったカテゴリーを通じてこの地域を説明するのかどうかを問うものである。標準的な公平性の指標では、構造的な枠組みではなく明示的な偏見を検出するため、これに答えることはできません。この論文では、サイードの 7 つのオリエンタリスト活動を測定可能な次元に変えるフレームワークである中東文化感受性スコア (MECSS) と、特定の失敗に対する「サイード洗浄」という用語、つまり一般化を否認し、否認した構造を再現するモデルを紹介します。 280 件の会話 (1,120 件の交換) にわたって、GPT-4 と Falcon3-7B-Instruct はどちらも、オープンなステレオタイプ化ではなく構造的位置付けを通じて、オリエンタリストのパターンを体系的に再現しました。 GPT-4 スコアは中程度 (平均 MECSS 1.73)。 Falcon3-7B-Instruct は、アブダビで構築され、アラビア語のコンテンツでトレーニングされているにもかかわらず、より高いスコア (2.18) を示しています。これは、モデルを地域的に構築するとオリエンタリズムが薄れるという仮定に対する証拠ですが、モデルのサイズや起源が異なるため、原因として地理を分離することはできません。エピステミック センター (西洋のフレームワークをマークのない普遍的なものとして扱う) は、両方のモデルで最高点近くのスコアを獲得しました。 Said-washing は GPT-4 会話の 87.9% に現れており、既存の指標では確認できないパターンです。このバイアスを軽減するには、言語を追加したり機関を移転したりするだけでなく、モデルが学習する内容を変更する必要があります。
原文 (English)
Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)
AI systems now shape how hundreds of millions of people learn about cultures other than their own. When someone asks one of these systems about the Middle East, they do not receive neutral facts. They receive a representation shaped by the frameworks embedded in training data, and that data is overwhelmingly Western and English-language. This paper asks whether that representation is Orientalist in Said's sense: whether it denies agency to Middle Eastern actors, treats Western frameworks as neutral while marking non-Western knowledge as particular, and explains the region through categories it did not produce. Standard fairness metrics cannot answer this, because they detect explicit prejudice rather than structural framing. This paper introduces the Middle East Cultural Sensitivity Score (MECSS), a framework that turns Said's seven Orientalist operations into measurable dimensions, and the term "Said-washing" for a specific failure: a model that disclaims generalization, then reproduces the structure it disclaimed. Across 280 conversations (1,120 exchanges), GPT-4 and Falcon3-7B-Instruct both reproduce Orientalist patterns systematically, through structural positioning rather than open stereotyping. GPT-4 scores moderately (mean MECSS 1.73); Falcon3-7B-Instruct scores higher (2.18), even though it was built in Abu Dhabi and trained with Arabic content. This is evidence against the assumption that building a model regionally makes it less Orientalist, though the models differ in size as well as origin, so geography cannot be isolated as the cause. Epistemic Center, the treatment of Western frameworks as unmarked universals, scores near the top of the scale for both models. Said-washing appears in 87.9% of GPT-4 conversations, a pattern existing metrics cannot see. Reducing this bias requires changing what models learn from, not only adding languages or relocating institutions.
DeepTCM1.0: 一般的な大規模言語モデルに基づいて漢方薬の処方メカニズムを解読するためのマルチエキスパート AI エージェント
背景: 伝統的な中国医学 (TCM) の配合処方のメカニズムの解明は、依然として TCM の近代化における中心的な課題です。データマイニングやネットワーク薬理学などの従来のアプローチは、古典的な TCM 理論と現代の科学研究の間の深い統合を達成するには不十分です。さらに、汎用人工知能大規模言語モデルを使用した直接的な質問応答は、TCM の理論的枠組みへの適応が不十分であることと、推論幻覚に対する感受性により制限されます。したがって、TCM の全体的な原則に沿ったインテリジェントな分析方法を開発することが急務となっています。目的:古典的なTCM理論と現代の生命科学を統合する複数の専門家によるインテリジェントエージェントフレームワークを確立し、それによってTCM複合処方の系統的かつ解釈可能な機構分析を可能にし、代表的な検証ケースとしてGuizhi煎じ薬を使用します。方法: DeepTCM1.0 フレームワークは、汎用大規模言語モデル DeepSeek V3.2 に基づいて構築されました。 3 層の協調アーキテクチャと 3 ラウンドの反復的な品質管理ワークフローを採用し、11 の学際的なインテリジェント エージェントの協調分析プロセスをシミュレートします。このフレームワークは、古典的な伝統的な中国医学の理論と現代の科学研究の二重の観点から、桂枝煎じ薬のメカニズムの解釈に適用されました。フレームワークのパフォーマンスは、二重盲検 5 次元スコアリング、クラス内相関係数 (ICC) 信頼性テスト、マンホイットニー U テスト、および効果量分析を通じて包括的に評価されました。この評価では、評価者として 4 つの独立した大規模言語モデルを採用し、それぞれが 5 つの匿名化レポートに対して 5 ラウンドの繰り返しスコアリングを実行し、合計 100 の独立したスコアリング評価が行われました。
原文 (English)
DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models
Background: Mechanistic elucidation of traditional Chinese medicine (TCM) compound formulas remains a central challenge in the modernization of TCM. Conventional approaches, including data mining and network pharmacology, are insufficient for achieving deep integration between classical TCM theory and modern scientific research. In addition, direct question-answering using general-purpose artificial intelligence large language models is limited by inadequate adaptation to TCM theoretical frameworks and susceptibility to reasoning hallucinations. Consequently, there is an urgent need to develop intelligent analytical methods aligned with the holistic principles of TCM. Objective: To establish a multi-expert intelligent agent framework integrating classical TCM theory with modern life sciences, thereby enabling systematic and interpretable mechanistic analysis of TCM compound formulas, with Guizhi Decoction serving as a representative validation case. Methods: The DeepTCM1.0 framework was constructed based on the general-purpose large language model DeepSeek V3.2. It adopts a three-tier collaborative architecture and a three-round iterative quality-control workflow, simulating the collaborative analytical process of 11 interdisciplinary intelligent agents. The framework was applied to the mechanistic interpretation of Guizhi Decoction from the dual perspectives of classical traditional Chinese medicine theory and modern scientific research. Framework performance was comprehensively evaluated through double-blind five-dimensional scoring, intraclass correlation coefficient (ICC) reliability testing, Mann-Whitney U tests, and effect size analysis. The evaluation employed four independent large language models as evaluators, each conducting five rounds of repeated scoring on five anonymized reports, resulting in a total of 100 independent scoring assessments.
StocksTalk: Web データ上で構造化クエリを生成するための音声対応会話エージェント
StocksTalk は、音声による財務スクリーニング要求を、現実世界の市場データに対する実行可能で検証済みの構造化クエリに変換するための音声対応会話システムです。このシステムは、ストリーミング音声認識、検索拡張制約抽出、スキーマに基づいた LLM ベースの SQL 生成、ルールベースの検証、および人間参加型検証を対話型ダッシュボード内で組み合わせています。従来のテンプレート主導の財務アシスタントとは異なり、StocksTalk は、抽出された制約、正規化された財務指標、オペレーターのグラウンディング、生成されたクエリなどの中間推論アーティファクトを公開し、ユーザーが実行前に各段階を検査して調整できるようにします。システムを評価するために、複数の投資戦略と入力ノイズ条件にわたる 150 の音声財務プロンプトのベンチマークを厳選しました。実験結果は、ベースラインの LLM ベースのアプローチと比較して、検索グラウンディング、制約付きクエリ生成、および対話型検証により、制約抽出精度、SQL 実行可能性、論理的一貫性、およびマルチターン安定性が大幅に向上することを示しています。 StocksTalk は、透明な音声主導のインターフェイスが自然言語の対話と構造化された財務分析をどのように橋渡しし、会話型の株式スクリーニングと意思決定サポートのための効果的なフレームワークを提供できるかを示しています。
原文 (English)
StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data
StocksTalk is a voice-enabled conversational system for transforming spoken financial screening requests into executable and validated structured queries over real-world market data. The system combines streaming speech recognition, retrieval-augmented constraint extraction, schema-grounded LLM-based SQL generation, rule-based validation, and human-in-the-loop verification within an interactive dashboard. Unlike traditional template-driven financial assistants, StocksTalk exposes intermediate reasoning artifacts, including extracted constraints, normalized financial metrics, operator grounding, and generated queries, allowing users to inspect and refine each stage before execution. To evaluate the system, we curate a benchmark of 150 spoken financial prompts spanning multiple investment strategies and input noise conditions. Experimental results show that retrieval grounding, constrained query generation, and interactive verification substantially improve constraint extraction accuracy, SQL executability, logical consistency, and multi-turn stability compared to baseline LLM-based approaches. StocksTalk demonstrates how transparent, voice-driven interfaces can bridge natural language interaction and structured financial analysis, providing an effective framework for conversational stock screening and decision support.
言語化された自信過剰のさまざまな側面: 解釈可能性の研究
大規模な言語モデルは自信過剰になる傾向があり、証拠がヘッジまたは棄権を示唆している場合に断定的な回答をします。論理的必然性と可能性を操作する制御された推論シナリオを使用して、言語認識マーカー、棄権、数値信頼スコアという不確実性を表現する 3 つの方法にわたって、Qwen3-4B でこの動作を研究します。私たちの結果は、特にモデルが数値の信頼スコアを出力するように求められた場合に、この自信過剰になる傾向を裏付けています。解釈可能性レベルでは、不確実性と確実性の原因となるトランスコーダの機能を区別して識別する方法を提案します。私たちの分析により、Qwen3-4B のデフォルトのメカニズムは共有機能の広範な連携による確実性の生成に有利である一方、不確実性は専用機能の少数のセットによって媒介されるスパースなオーバーライドとして実装されていることが明らかになりました。これらの不確実性の特徴に介入すると、自信過剰の根底にあるこの不均衡が因果的に証明され、自信過剰のエラーも軽減されます。同じ一連の特徴が、3 つの不確実性表現設定、言語、および分布外モダリティ タスクにわたって一般化されます。
原文 (English)
Different Facets of Verbalised Overconfidence: an Interpretability Study
Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, across three ways to express uncertainty: verbal epistemic markers, abstention, and numeric confidence scores. Our results confirm this tendency toward overconfidence, particularly when the model is prompted to output a numeric confidence score. At the interpretability level, we propose a method that differentially identifies transcoder features responsible for uncertainty and certainty. Our analysis reveals Qwen3-4B's default mechanism favors certainty generation through a broad coalition of shared features, while uncertainty is implemented as a sparse override mediated by a small set of dedicated features. Intervening on these uncertainty features both causally proves this imbalance underlying overconfidence and also mitigate overconfident errors. The same set of features generalise across the three uncertainty-expression settings, languages, and an out-of-distribution modality task.
大規模言語モデルにおける地理的バイアスとしての制度の威信: ブートストラップ信頼区間を使用した 3 つの階乗実験からの証拠
私たちは、大規模言語モデル (LLM) が候補者の評価において、応募者の名前の民族性や組織の名声、地理的位置に基づいて体系的に差別しているかどうかを調査します。 3 つの要因実験が報告されています (4,320 の API 呼び出し、4 つの LLM、5 つのプロフェッショナル ドメイン)。研究 1 (3x4 計画) では、10 段階評価で +0.297 ポイントの統計的に堅牢な機関階層勾配 (95% ブートストラップ CI: +0.175 ~ +0.422) が見られますが、名前由来の影響は無視でき、有意ではありません (95% CI はゼロと交差します)。研究 2 (2x2 名声 x 国のデザイン) は、名声と地理の交絡を打ち破ります。名声効果 (+0.185; 95% CI: +0.093 ~ +0.275) が出身国の効果 (+0.126; 95% CI: +0.037 ~ +0.218) を 1.5 倍上回っています。研究 3 (2x2 ジャーナル x 機関設計) では、ジャーナルの名声 (Nature 対周辺のオープンアクセスジャーナル) が機関の名声を 5.7 倍上回ることが明らかになりました: ジャーナル効果 +1.937 (95% CI: +1.811 ~ +2.062) 対 機関効果 +0.341 (95% CI: +0.184 ~ +0.504)。 「救済効果」が確認されています。Nature への掲載は、MIT (+1.745) よりもグアヤキル大学 (+2.127) の候補者の方が、低い組織的名声をより強く補います。結果は、NBI を使用して定量化されます。 I コンポーネントは、プレステージの低いプロファイルに対する評価の不一致の増加を明らかにし、これは平均のみの指標では捉えられない認識論的な欠点です。コードとデータ: https://github.com/mleyvaz/geo-bias-llm
原文 (English)
Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals
We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while name-origin effects are negligible and non-significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige-geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country-of-origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open-access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A "rescue effect" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI; the I component reveals elevated evaluation inconsistency for low-prestige profiles, an epistemic disadvantage not captured by mean-only metrics. Code and data: https://github.com/mleyvaz/geo-bias-llm
同じ事実、異なる更新: 推論設定が医療配分における LLM の動作を形作る
大規模な言語モデルは、ほぼすべての分野にわたる機密性の高い重要な意思決定プロセスに組み込まれています。これまでの研究では、入力とシナリオの枠組みに関するモデルのバイアスが研究されていましたが、モデルは、展開中に蓄積されたコンテキストによって予期せぬ望ましくない方法で動作する可能性もあります。この研究では、モデルが短い臨床状況を与えられた 2 人にリソース割り当ての確率を割り当てるよう求められ、その後、前後関係における以前の応答の有無にかかわらず、対照的な患者情報を含む 1 つの追加文を含む同じシナリオを確認する医療例を研究します。 4 つのテスト済みモデルのうち 3 つにわたって、ペアコンテキスト実験と独立推論実験では、新しい情報が提供されたときに異なる確率シフトがあり、多くの場合、反対方向 (人物 B に有利な場合と人物 A に有利な場合) にシフトします。シナリオ軸全体での属性の変化の影響を示すために、追加のペア コンテキスト実験が含まれています。私たちの調査結果は、機密性の高い医療ユースケースにおける患者情報の状況依存の影響を示しています。より広範には、私たちの研究は、LLM ベースのシステムを意思決定プロセス、コンテキスト エンジニアリング、およびさらなるモデル行動研究に慎重に組み込むことの重要性を示しています。
原文 (English)
Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation
Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior work studies model bias around inputs and scenario framing, models can also behave in unexpected and undesirable ways due to context accumulated over their deployment. In this work, we study a medical example in which a model is asked to assign resource-allocation probabilities to two people given brief clinical context, and then sees the same scenario with a single extra sentence containing contrasting patient information, either with or without its previous response in context. Across three of four tested models, the paired-context and independent-inference experiments have different probability shifts, often in opposite directions (in favor of Person B vs. in favor of Person A) when new information is provided. We include additional paired-context experiments to show the effect of varying attributes across scenario axes. Our findings show the context-dependent effect of patient information in a sensitive medical use case. More broadly, our work shows the importance of carefully incorporating LLM-based systems into decision-making processes, context engineering, and further model behavioral studies.
非侵襲的な脳記録からの自然文の正確な解読
脳損傷後に話す能力や動く能力を失った人々のコミュニケーションを回復することは大きな課題です。頭蓋内インプラントは現在、高性能のブレインコンピューターインターフェイスを可能にしていますが、非侵襲的な代替品は依然として遅れをとっています。ここでは、リアルタイムの脳磁図 (MEG) 記録のみから自然文の生成をデコードできるモデル、Brain2Qwerty v2 を紹介します。 9 人の被験者が入力し、それぞれ 10 時間記録した 22,000 の文を収集することで、私たちのモデルは文字、単語、文レベルの表現を活用して、平均単語誤り率 (WER) 39% を達成しました。最良の参加者の場合、モデルは文の半分を 1 単語以下のエラーで正確にデコードしました。重要なことに、デコード精度はデータ量に応じて対数線形に向上し、頭蓋内アプローチとのパフォーマンスギャップがデータスケーリングによって部分的に埋められる可能性があることを示唆しています。 AI がこのパフォーマンスを 3 つの主な方法で実現できることを示します。イベント検出用の手作りパイプラインをディープ ラーニングで置き換える方法、意味論的表現を抽出するための大規模な言語モデルを微調整する方法、自動化されたコード開発を通じてデコード パイプラインを反復的に改良する AI エージェントの展開です。これらの結果を総合すると、非侵襲的な脳からテキストへのデコードが、以前は外科用インプラントに限定されていると考えられていたレベルの精度で動作し始め、安全で効率的なブレインコンピューターインターフェイスへの道が開かれたことが示されています。
原文 (English)
Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings
Restoring communication for people who have lost the ability to speak or move after a brain injury is a major challenge. While intracranial implants now enable high-performing brain-computer-interfaces, non-invasive alternatives are still lagging behind. Here, we present Brain2Qwerty v2, a model that can decode the production of natural sentences solely from real-time magnetoencephalography (MEG) recordings. By collecting 22,000 sentences typed by nine subjects, each recorded for 10 hours, our model leverages character, word and sentence-level representations to achieve an average word error rate (WER) of 39%. For our best participant, the model accurately decodes half of the sentences with one word error or less. Critically, decoding accuracy log-linearly improves with data volume, suggesting that the performance gap with intracranial approaches could be partially bridged through data scaling. We show that AI enables this performance in three main ways: the substitution of hand-crafted pipelines for event detection with deep learning, the finetuning of large language models to extract semantic representations, and the deployment of AI agents to iteratively refine our decoding pipeline via automated code development. Together, these results show that non-invasive brain-to-text decoding starts to operate at a level of accuracy previously thought exclusive to surgical implants, opening a path toward safe and efficient brain-computer-interfaces.
トークンレベルの幻覚検出のための時間的多重信号融合
トークンレベルの幻覚検出器は、単一の信号から独立して各トークンをスコアリングし、生成モデルが確実に間違っている場合に正確に失敗します。この論文では代わりに、幻覚を時間的に拡張されたスパンとして扱い、シーケンス ラベリングによって検出します。各トークンは、テキスト統計、自然言語推論 (NLI) 含意、および言語モデルの意外性を融合する 33 次元の特徴ストリームからスコア付けされますが、モデルの内部にはアクセスできません。これらの機能に対する双方向ゲート型リカレント ユニット (BiGRU) は、RAGTruth (10 シード) で AUC 0.840 に達し、独立したロジスティック回帰ベースライン (p = 0.002、Wilcoxon の符号付きランク) を 11 ポイント上回りました。制御された分解により、ゲインのほとんどはモデルの能力ではなく時間的順序に起因します。証拠は、スパン内で信頼できる位置から曖昧な隣接位置に伝播します。同じ 0.845 の上限がリカレント アーキテクチャ、状態空間 (Mamba) アーキテクチャ、およびアテンション アーキテクチャにわたって繰り返し発生し、ボトルネックがモデルではなく機能セットにあることがわかります。生成されたテキストと外部信号のみを読み取るため、検出器はクローズドソース モデルで動作し、トレーニング中に表示されなかった言語モデルによって生成されたテキストでも動作し続け、AUC が 4% 未満で失われます。
原文 (English)
Temporal Multi-Signal Fusion for Token-Level Hallucination Detection
Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the generating model is confidently wrong. This paper instead treats hallucination as a temporally extended span and detects it by sequence labeling: each token is scored from a 33-dimensional feature stream that fuses text statistics, Natural Language Inference (NLI) entailment, and language model surprisal, with no access to model internals. A Bidirectional Gated Recurrent Unit (BiGRU) over these features reaches an AUC of 0.840 on RAGTruth (10 seeds), an 11-point gain over an independent logistic-regression baseline (p = 0.002, Wilcoxon signed-rank). A controlled decomposition attributes most of the gain to temporal order rather than model capacity: evidence propagates from confident positions to ambiguous neighbors within a span. The same 0.845 ceiling recurs across recurrent, state-space (Mamba), and attention architectures, locating the bottleneck in the feature set rather than the model. Because it reads only the generated text and external signals, the detector works on closed-source models, and it keeps working on text produced by language models it never saw during training, losing under 4% AUC.
Global Index on Responsible AI 2026 : Conceptual Framework and Methodology
This report presents the methodology of the Global Index on Responsible AI (GIRAI), 2nd Edition. This edition refines the 1st Edition by st…
ポルトガル語の言語モデル: 体系的なマッピング研究
近年、言語モデルの急速な発展により、幅広いアプリケーションを通じて自然言語処理の分野が変革されました。ただし、言語モデルの開発はすべての言語で均一に進んでいるわけではありません。ポルトガル語の場合、最近、学術界や企業によるポルトガル語用の言語モデルの開発とデータ リソースの作成への取り組みが活発になっています。これらの取り組みの結果、ポルトガル語の言語モデルのますます多様なエコシステムが台頭しました。ただし、これらのモデルに関する情報は、科学出版物、技術レポート、モデル リポジトリ、プロジェクト ドキュメントに分散されたままです。この調査は、ポルトガル語用に開発された言語モデルの体系的なマッピング研究を示し、この分野の現状の包括的な概要を提供します。合計 46 のモデルをマッピングし、ベース モデル、アーキテクチャ、計算リソース、トレーニング データセット、ライセンス、コードの可用性、データ、モデルの重みなどのさまざまな側面で特徴付けます。さらに、系統発生の観点からこれらのモデル間の進化と関係を分析し、現在の研究のギャップと機会を特定し、ポルトガル語の言語モデル開発の将来の方向性について議論しました。
原文 (English)
Language Models for Portuguese: A Systematic Mapping Study
In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of applications. However, the development of language models has not progressed uniformly across all languages. In the case of the Portuguese language, there has recently been a growing effort by academia and companies to develop language models and create data resources for Portuguese. These efforts have resulted in the rise of an increasingly diverse ecosystem of language models for Portuguese. However, information on these models remains dispersed in scientific publications, technical reports, model repositories, and project documentation. This survey presents a systematic mapping study of language models developed for Portuguese, providing a comprehensive overview of the current state of the field. We map a total of 46 models, characterizing them by various aspects, including base model, architecture, computational resources, training datasets, licensing, code availability, data, and model weights. Furthermore, we analyzed the evolution and relationships among these models through a phylogenetic perspective, identified current research gaps and opportunities, and discussed future directions for the development of language models for Portuguese.
義務論的ギャップ: 大規模な言語モデルと義務の様相言語
「しなければならない」、「しなければならない」、「しなければならない」などの様相の補助語は、話者の権限や対人関係の立場の中での必要性と義務を示します。私たちは、大規模言語モデル (LLM) が現代の人間の義務的なモーダル使用パターンを再現するかどうかを調べます。 3 つの主要なコーパス、外部ベンチマーク、2 つの制御された複製、および自然主義的な 11 モデルの複製にわたって、AI によって生成されたテキストは、現代の人間と比較して、肯定的な義務的なモーダル (しなければならない、しなければならない、しなければならない) を一貫して過少に使用しています。出版された散文記録に対するヒューリスティックな校正として使用された Google Books Ngram コーパス (1920 ~ 2022 年) との歴史的比較は、AI のモーダル周波数が正式に出版された英語の範囲内に収まるのに対し、非公式なデジタル コンテキストにおける現代の人間のモーダル レートは 20 世紀の書籍のベースラインを超えることが多いことを示しています。フレーズレベルで分解すると、AIと人間のモーダルギャップが対人関係の中心となる構造(べき、しなければならない、しなければならない)に集中しているのに対し、AIは指導や質問に答える文脈では必要に応じて人間と同等かそれを上回っているが、生徒の説得力のある文章ではそうではなく、モーダルプロファイルがジャンル条件付きであることが示されている。この調査結果は、LLM モーダル使用法は、これらのモデルがトレーニングされた正式な文書リソースを反映している一方で、現代の人間の書き手が直接的な対人関係の義務をマークするモーダル構造を十分に活用していないことを示唆しています。
原文 (English)
The Deontic Gap: Large Language Models and the Modal Language of Obligation
Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance. We examine whether large language models (LLMs) reproduce contemporary human patterns of deontic modal usage. Across three primary corpora, an external benchmark, two controlled replications, and a naturalistic eleven-model replication, AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. Historical comparison with the Google Books Ngram corpus (1920-2022), used as a heuristic calibration against the published-prose record, shows that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines. Phrase-level decomposition shows that the AI-human modal gap is concentrated in constructions central to interpersonal stance (should, have to, had to), while AI matches or exceeds humans on need to in instructional and question-answering contexts but not in persuasive student writing, indicating that the modal profile is genre-conditional. The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation.
Entropy-Constrained Adaptive Stochastic Quantization
Adaptive stochastic quantization (ASQ) is a recently introduced quantization approach that optimizes the Mean Squared Error (MSE) for a giv…
TokenPowerSandbox: エネルギーを考慮した LLM サービスの証拠ゲート型 CPU ファースト スクリーニング
エネルギーを意識した LLM サービスでは、現実的なリクエスト形状に基づいて構成を比較する必要がありますが、網羅的なターゲット GPU プロファイリングにはコストがかかり、安価な予測子は測定範囲外では危険なほど自信を持ってしまう可能性があります。私たちは、解釈可能な CPU 常駐プロジェクター、短いターゲット GPU プローブ、完全なワークロード検証、および改ざん明示的な測定前のフリーズ来歴を組み合わせた証拠ゲート型ワークフローである TokenPowerSandbox を紹介します。 vLLM を備えた Qwen2.5-7B-Instruct を提供する 1 つの NVIDIA H100 80GB では、3 つのアンカー リピートと 6 つの開発ワークロードがワークロード転送を調整します。同じフリーズ モデルが、ブラインド ホールドアウトと、別途事前に宣言された非再適合確認で評価され、合計 51 回のフリーズ後実行が行われます。 Energy MAPE は 6.23% と 7.35% で、Spearman の順位相関は 0.976 と 0.933 です。ただし、事前に宣言された TTFT ゲートは同時実行数 4 (MAPE 9.27%) でパスし、同時実行数 4 (64.80%) 未満では棄権をトリガーします。これは、エネルギー精度がレイテンシーを証明できない理由を示しています。
原文 (English)
TokenPowerSandbox: Evidence-Gated CPU-First Screening for Energy-Aware LLM Serving
Energy-aware LLM serving requires comparing configurations under realistic request shapes, yet exhaustive target-GPU profiling is costly and a cheap predictor can be dangerously confident outside its measured scope. We present TokenPowerSandbox, an evidence-gated workflow that combines an interpretable CPU-resident projector, short target-GPU probes, full-workload verification, and tamper-evident freeze-before-measurement provenance. On one NVIDIA H100 80GB serving Qwen2.5-7B-Instruct with vLLM, three anchor repeats and six development workloads calibrate workload transfer. The same frozen model is evaluated on a blind holdout and a separately predeclared no-refit confirmation totaling 51 post-freeze runs. Energy MAPE is 6.23% and 7.35%, with Spearman rank correlations of 0.976 and 0.933. However, a predeclared TTFT gate passes at concurrency four (9.27% MAPE) and triggers abstention below four (64.80%), showing why energy accuracy cannot certify latency.
クォンタムにはどのようなメリットがあるのでしょうか?ネットワーク侵入検知のための量子機械学習の公正でキャリブレーションとノイズを意識したベンチマークと帰属監査
ネットワーク侵入検知 (NIDS) のための量子機械学習 (QML) は、ほぼ完璧な精度に達していると定期的に報告されていますが、最も厳密な研究では、よく調整された古典的モデルが競争力を維持しており、見かけの量子ゲインは、本物の量子効果ではなく、古典的な次元削減と暗黙の正則化のアーティファクトである可能性があることが判明しています。私たちは、量子モデルが高い精度を出せるかどうかではなく、量子モデルの利点が実際にどの程度あるのかを問います。 4 つの標準 NIDS データセット (NSL-KDD、UNSW-NB15、CICIDS2017、NF-ToN-IoT-v2) にわたる 5 つの正直に調整された古典的なベースラインに対してハイブリッド変分量子回路と量子カーネル SVM を評価する、統一された再現可能な QML-IDS ベンチマークを、均等な予算の機能ビュー、不均衡とキャリブレーションを考慮した、リーク制御された 1 つのプロトコルの下で提示します。有意性テストを備えたメトリクス、およびシミュレートされた NISQ ノイズ スイープ。量子属性監査 (パラメーターが一致した古典的制御、ランダム特徴カーネル、正則化スイープ) を導入し、ゲインのどれだけが実際に量子コンポーネントに起因するかを定量化します。調整された古典的モデル (ランダム フォレスト、XGBoost) は、すべてのデータセットでの集計検出において量子モデルと同等かそれを上回っており、監査ではこれが量子効果ではなく古典的な前処理と正則化によるものであると考えています。誤検出率補正後も 2 つの利点が存続します。量子カーネル SVM は AUPRC および ROC-AUC でその直接の古典的サロゲート (ランダム特徴カーネル) よりもランク付けされており、小型 4 量子ビット ハイブリッドは分布シフト NSL-KDD タスク (p = 0.005、BH q = 0.030)。コード、シード、分割がリリースされます。クォンタムが勝っても、引き分けても、負けても、私たちの貢献は変わりません。
原文 (English)
How Quantum Is the Advantage? A Fair, Calibration- and Noise-Aware Benchmark and Attribution Audit of Quantum Machine Learning for Network Intrusion Detection
Quantum machine learning (QML) for network intrusion detection (NIDS) is routinely reported to reach near-perfect accuracy, yet the most rigorous studies find that well-tuned classical models remain competitive, and that apparent quantum gains may be artefacts of classical dimensionality reduction and implicit regularisation rather than genuine quantum effects. We ask not whether a quantum model can post a high accuracy, but how quantum the advantage really is. We present a unified, reproducible QML-IDS benchmark evaluating hybrid variational quantum circuits and quantum-kernel SVMs against five honestly-tuned classical baselines across four standard NIDS datasets (NSL-KDD, UNSW-NB15, CICIDS2017, NF-ToN-IoT-v2) under one leakage-controlled protocol, with an equal-budget feature view, imbalance- and calibration-aware metrics with significance testing, and a simulated NISQ noise sweep. We introduce a quantum-attribution audit (parameter-matched classical controls, a random-feature kernel, and a regularisation sweep) that quantifies how much of any gain is genuinely attributable to the quantum component. Tuned classical models (Random Forest, XGBoost) match or exceed the quantum models on aggregate detection on every dataset, and the audit attributes this to classical preprocessing and regularisation rather than quantum effects. Two advantages survive false-discovery-rate correction: the quantum-kernel SVM out-ranks its direct classical surrogate (a random-feature kernel) on AUPRC and ROC-AUC, and a small four-qubit hybrid out-detects the best classical baseline at the 1% false-positive operating point on the distribution-shifted NSL-KDD task (p = 0.005, BH q = 0.030). Code, seeds, and splits are released; our contribution stands whether quantum wins, ties, or loses.
LLM が実際に役立つのはどのような場合ですか?データ品質アノテーターとしての LLM の評価
LLM はデータ品質の問題を自動的に検出するために使用されることが増えていますが、これらの判断が実際にどの程度一貫しているかについてはほとんどわかっていません。この調査では、ルールベースのベースラインと人間が検証したグラウンドトゥルースに対して、エンティティマッチングとブランドの誤ったラベル付けという 2 つの電子商取引データ品質タスクについて、ゼロショットプロンプトと少数ショットプロンプトの両方で LLM をテストします。 Abt Buy ベンチマーク (2,194 のラベル付きペア) を使用したエンティティ マッチングでは、単純なルール ベースのベースライン (F1=0.950) が LLM ゼロ ショット プロンプト (F1=0.948) とほぼ同じパフォーマンスを示しました。さらに、小規模な検証サンプルでは効果的であると思われる数ショットのプロンプト修正により、フルスケールのパフォーマンスが F1=0.914 に低下しました。これは、サンプルが少ない場合の即時評価は誤解を招く可能性があることを示しています。合成的に挿入されたラベル付けエラーを含む 500 件の Amazon 製品リストを使用したブランドの誤ったラベル付けの検出では、単純なルールではアクセスできないブランド製品の関係性に関する背景知識を利用できるため、LLM は単純なルールベースのベースライン (F1=0.833 対 0.721) を明らかに上回りました。繰り返しの実行 (200 ペア、温度 0.7 で 5 回の実行) にわたる一貫性のテストでは、モデルが平均 99.7% の確率でそれ自体と一致し、5 回の実行すべてでペアの 99% が同一の回答を示しました。これらの実行全体で多数決を使用すると、F1 は 0.005 改善されるだけで、推論コストは 5 倍になります。これらの結果は、従来の方法よりも LLM を使用する価値がタスクに大きく依存することを示唆しています。 LLM は、強力な語彙シグナルがすでに存在する場合にはほとんど利点がありませんが、タスクに背景知識が必要な場合には明らかな利点があり、繰り返されるクエリ全体で高い一貫性を維持します。
原文 (English)
When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators
LLMs have been increasingly used to catch data quality issues automatically, but we know very little about how consistent these judgments actually are. This study tests an LLM on two e-commerce data quality tasks, entity matching and brand mislabeling, against rule based baselines and human verified ground truth, under both zero-shot and few-shot prompting. On entity matching while using the Abt Buy benchmark (2,194 labeled pairs), a simple rule based baseline (F1=0.950) performed about as well as LLM zero shot prompting (F1=0.948). Moreover, a few-shot prompt revision that looked effective on a small validation sample reduced full-scale performance to F1=0.914. This showed that small sample prompt evaluation can be misleading. On brand mislabeling detection, using 500 Amazon product listings with synthetically injected labeling errors, the LLM clearly outperformed a naive rule based baseline (F1=0.833 vs 0.721), because it could draw on background knowledge of brand product relationships that a simple rule could not access. Testing consistency across repeated runs (200 pairs, 5 runs at temperature 0.7) showed the model agreeing with itself 99.7% of the time on average, with 99% of pairs giving identical answers across all 5 runs. Using majority voting across these runs only improved F1 by 0.005, at 5 times the inference cost. These results suggest that the value of using an LLM over traditional methods depends heavily on the task. LLMs offer little advantage when strong lexical signals already exist, but a clear advantage when the task requires background knowledge, all while remaining highly consistent across repeated queries.
LLM はテキストを超えて安全ですか: 絵文字は安全性評価のギャップを明らかにしますか
大規模言語モデル (LLM) の安全性評価は主にテキストベースの敵対的プロンプトに依存しており、代替入力表現から生じる脆弱性が見落とされる可能性があります。この研究では、このギャップのテスト ケースとして絵文字拡張プロンプトを調査し、4 つのオープンソース LLM (Mistral 7B、Qwen 2 7B、Gemma 2 9B、Llama 3 8B) にわたる 50 のプロンプトを評価しました。結果は、堅牢性に大きなばらつきがあることを示しています。Gemma 2 9B と Mistral 7B はゼロ以外の成功率 (10%)、Llama 3 8B は 6% を示しましたが、Qwen 2 7B は完全な耐性 (成功率 0%) を示しました。カイ二乗検定 ($\chi^2 = 32.94、p < 0.001$) により、結果の分布に有意な差があることが確認されます。これらの発見は、堅牢性が入力表現に敏感であること、および標準のテキスト プロンプトに限定された評価ではモデルの脆弱性が過小評価されている可能性があることを示しています。
原文 (English)
Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation
Safety evaluations of large language models (LLMs) predominantly rely on text-based adversarial prompts, potentially overlooking vulnerabilities arising from alternative input representations. This work examines emoji-augmented prompts as a test case for this gap, evaluating 50 prompts across four open-source LLMs (Mistral 7B, Qwen 2 7B, Gemma 2 9B, Llama 3 8B). Results show substantial variation in robustness: Gemma 2 9B and Mistral 7B exhibit non-zero success rates (10%), Llama 3 8B 6%, while Qwen 2 7B shows complete resistance (0% success rate). A chi-square test ($\chi^2 = 32.94, p < 0.001$) confirms significant differences in outcome distributions. These findings indicate that robustness is sensitive to input representation, and that evaluations restricted to standard text prompts may underrepresent model vulnerabilities.
人工知能は医学から何を学べるのでしょうか?生成的アナロジーと信頼性の高い機械学習システム
ここ数年、機械学習 (ML) は医療に広く (そしてある程度は成功して) 実装されてきました。しかし、ML を取り巻く不確実性により、その認識論的根拠と方法論的根拠を確立することが困難になっています。文献では、医学と機械学習の間に類似点が描かれており、臨床翻訳の標準に基づいて機械学習の認識論的および方法論的な標準をモデル化する必要があることが示唆されています。ヘッセの研究からツールを開発することにより、私たちはこの類似点の性質を、臨床翻訳のプロセスと ML システム構築のプロセスの間の生成的な類似点として特徴づけます。私たちは、通常、類推に訴える場合にのみ言及される臨床翻訳の認識論的および方法論的根拠をより正確に特定し、そのような根拠がどの意味で ML のコンテキストに類推的に適用されるかを示します。特に、臨床翻訳の令状を信頼性主義の用語で解釈し、これが AI 哲学における既存の信頼性主義の説明とは異なる (ただし互換性がある) 新しい形式の ML 信頼性をどのように伝えることができるかを示します。
原文 (English)
What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems
In the past few years, machine learning (ML) has been widely (and to an extent, successfully) implemented in medicine. However, uncertainties surrounding ML have made it difficult to establish the bases of its epistemic and methodological warrants. In the literature, a parallel has been drawn between medicine and ML, suggesting that we should model epistemic and methodological standards for ML on the standards of clinical translation. By developing tools from Hesse work, we characterise the nature of this parallel as a generative analogy between the process of clinical translation and the process of building ML systems. We identify more precisely the epistemic and methodological warrants of clinical translation that are typically only mentioned when appealing to the analogy, and we show in which sense such warrants apply analogically to the context of ML. In particular, we interpret warrants of clinical translation in reliabilist terms, and we show how this can inform a new form of ML reliabilism, which is distinct from (though compatible with) existing reliabilist accounts in philosophy of AI.
自閉症の診断と治療に取り組むための機械学習技術の系統的レビュー: 課題と機会
自閉症スペクトラム障害(ASD)は、社会的相互作用やコミュニケーションにおける困難を特徴とする発達障害です。 ASD の原因は依然として不明瞭であるため、関連する特徴と隠れた相関関係を特定することが早期診断に重要です。この系統的レビューでは、ASD への機械学習 (ML) 技術の適用に関する 2017 年から 2023 年までの 55 件の研究を評価しています。主な目的は、ASD 研究における最近の ML アプリケーションを調査し、診断と治療を強化する傾向、技術、データセットを特定することです。 ASD 診断のニーズによく適合するため、教師あり学習方法が主流です。ただし、データの可用性が向上するにつれて、ディープラーニングの役割は拡大しています。教師なしディープラーニングやファジーロジックを組み込むことができる、ハイブリッド手法に基づく新たな技術は、将来観察するのが興味深いでしょう。このレビューでは、重要な課題と機会、特に診断精度と治療結果を向上させるために、遺伝情報や臨床情報などの複雑なデータを統合できるモデルの必要性が強調されています。さらに、ウェアラブル デバイスや生体認証センサーなどの革新的なデータ ソースを組み込むことで、継続的かつ非侵入的なモニタリングが可能になり、ASD をより総合的に理解できるようになります。調査結果は、現在の課題に対処するには、学際的なコラボレーションと ASD に合わせた拡張されたデータセットが必要であることを強調しています。将来の ML モデルは、より広範なマルチモーダル データ統合の恩恵を受け、研究者が ASD の複雑さにより包括的に対処できるようになります。
原文 (English)
A systematic review of machine learning techniques to address diagnosis and treatment of autism: challenges and opportunities
Autism spectrum disorder (ASD) is a developmental disability characterized by challenges in social interaction and communication. As the causes of ASD remain unclear, identifying relevant features and hidden correlations is crucial for early diagnosis. This systematic review evaluates 55 studies from 2017 to 2023 on the application of machine learning (ML) techniques to ASD. The primary objective is to examine recent ML applications in ASD research, identifying trends, techniques, and datasets that enhance diagnosis and treatment. Supervised learning methods dominate, as they align well with ASD diagnostic needs; however, the role of deep learning is expanding with greater data availability. Emerging techniques based on hybrid methods, where unsupervised, deep learning, and fuzzy logic could be included, will be interesting to observe in the future. The review highlights key challenges and opportunities, particularly the need for models that can integrate complex data -such as genetic and clinical information- to improve diagnostic accuracy and treatment outcomes. Additionally, incorporating innovative data sources, like wearable devices and biometric sensors, could enable continuous and non-intrusive monitoring, providing a more holistic understanding of ASD. Findings emphasize that addressing current challenges requires interdisciplinary collaboration and expanded datasets tailored to ASD. Future ML models will benefit from broader multimodal data integration, enabling researchers to more comprehensively address the complexities of ASD.
臨床ドメインシフトにおける多臓器CTセグメンテーションの境界を意識した臓器ごとのリコールリスク制御
配布フリーのリスク管理により、凍結セグメンテーションに臓器固有のリコール保証が追加されます。 AMOS でトレーニングされた nnU-Net の臓器ごとのしきい値を調整し、RAOS への転送を監査し、症例レベルのボクセル偽陰性率 (FNR) を使用してローカル再認定コストを推定します。 AMOS 管理は通過しますが、移植後 $7/12$ の臓器が $\alpha{=}0.10$ を超えます。より小さいキャリブレーション セットでは、控えめなしきい値または空のしきい値を使用して超過を隠すことができます。リスク制御予測セット (RCPS) は母集団平均リスクを高確率で制御しますが、コンフォーマルリスク制御 (CRC) は弱い期待値制御を行います。どちらも交換可能性を必要とします。固定およびグローバルしきい値では、臓器ごとの保証はありません。ワウドビー--スミス--ラムダス(WSR)のベッティングバウンドでは、6つのTier-1臓器が再認定され、地方症例数は25件であるのに対し、ヘフディング-ベンカス(HB)では30-40件となっている。 CRC には 10 ~ 15 が必要ですが、個々のケースのテールはより重くなります。 25 件の症例では、Tier-2 臓器は私たちの例示的な精度基準を満たしていません。
原文 (English)
Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift
Distribution-free risk control adds organ-specific recall guarantees to frozen segmentation. We calibrate per-organ thresholds for an AMOS-trained nnU-Net, audit transfer to RAOS, and estimate local re-certification cost using case-level voxel false-negative rate (FNR). The AMOS control passes, but $7/12$ organs exceed $\alpha{=}0.10$ after transfer; smaller calibration sets can mask exceedances with conservative or vacuous thresholds. Risk-Controlling Prediction Sets (RCPS) give high-probability control of population-mean risk, whereas Conformal Risk Control (CRC) gives weaker expectation control. Both require exchangeability; fixed and global thresholds give no per-organ guarantee. The Waudby--Smith--Ramdas (WSR) betting bound re-certifies six Tier-1 organs with 25 local cases, versus 30--40 for Hoeffding--Bentkus (HB). CRC needs 10--15 but has a heavier individual-case tail. No Tier-2 organ meets our illustrative precision criterion with 25 cases.
GigaBrain-WBC-0.5: 環境との相互作用を伴う堅牢な全身制御のための行動世界モデル
全身動作追跡ポリシーは、ヒューマノイドを堅牢な制御インターフェイスに変えます。遠隔操作者 (または上流モデル) は粗い動きの意図のみを提供しますが、低レベルのポリシーはロボットのバランスを保ち、物理的に実行可能に保ちます。既存のトラッカーは、平坦な地面でのみこのインターフェイスを提供します。空のシーンでトレーニングされ、地形やオブジェクトとの接触がそのダイナミクスをどのように再形成するかを学習することはなく、参照モーション コーパスを継続的に拡大することで、あらゆるコマンドの下でバランスをとるようにポリシーを教えようとしますが、実行可能な動作が環境に依存するようになると機能しなくなります。我々は、人型全身制御のための最初の行動世界モデル(BWM)であるGigaBrain-WBC-0.5を紹介します。純粋に反応的なトラッカーではなく、次の動作、次の状態、次の潜在的な動作コマンドの分布を共同で予測するように因果的 Transformer をトレーニングします。そのため、動作するネットワークは、環境が次に実行できることをどのように形成するかもモデル化します。自動地形アノテーション パイプラインは、リターゲットされたモーションから完全な 3D 接触ジオメトリを復元し、既存のモーション データセットのスケールで地形アノテーションを可能にします。予測された分布は展開時に再利用され、オンラインでありえないコマンドを検出し、学習した動作に反映させるため、ロボットは「ベストエフォート」方式でタスクを試みます。その結果、リアルタイムのコマンドを受け取り、環境と対話し、信じられないコマンド、落下、外乱に対して堅牢性を維持する統合ポリシーが実現します。 GigaBrain-WBC-0.5 は、3 つの大規模トラッカー ベースラインの中で、4 つのレジームすべてで最高の成功率を達成しました。地形インタラクションで 81.3% (最も強力なベースラインの 4.3 倍)、信じられないコマンドでの 83.1%、および落下からの回復 99.3% (最も強力なベースラインの 16.8 倍) です。ハードウェアの試験では、サポートの欠如や障害下での堅牢な相互作用が示されています。 Unitree G1 チェックポイントは、簡単な微調整で Maker L01 ロボットに転送されます。
原文 (English)
GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction
Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging the reference-motion corpus, which stops working once feasible behaviors become environment-dependent. We present GigaBrain-WBC-0.5, the first Behavior World Model (BWM) for humanoid whole-body control. Rather than a purely reactive tracker, we train a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next. An automatic terrain-annotation pipeline recovers full 3D contact geometry from retargeted motion, enabling terrain annotation at the scale of existing motion datasets. The predicted distribution is reused at deployment to detect implausible commands online and retract them onto learned behaviors, so the robot attempts tasks in a "best-effort" manner. The result is a unified policy that takes real-time command, interacts with environment, and stays robust to implausible commands, falls, and disturbances. GigaBrain-WBC-0.5 achieves the highest success rate across all four regimes among three large-scale tracker baselines: 81.3% on terrain interaction (4.3x the strongest baseline), 83.1% under implausible commands, and 99.3% fall recovery (16.8x the strongest baseline). Hardware trials show robust interaction under missing supports and disturbances; the Unitree G1 checkpoint transfers to the Maker L01 robot with simple fine-tuning.
Bidirectional representational alignment between biological and artificial neural networks
Recent work has shown that representational alignment between biological and artificial neural networks is asymmetric: model representation…
視覚的なプロンプトによるガイド付き野生生物個体レベルの認識
野生生物をきめ細かく再識別することは、研究において依然として困難な領域です。現在の最先端のアプローチでは、検出および再識別パイプラインが適用されます。潜在空間内で ID 検索を実行する 1 段階のエンドツーエンド検出および再識別モデルを提案します。堅牢な空間ジオメトリには DINOv2 を、野生動物の再識別には MegaDescriptor を採用しています。潜在的なクエリを迅速な再識別機能で強化します。検出デコーダは、シーンの潜在空間をクエリして、ターゲット ID の周囲にオブジェクトの境界を確立します。予備調査結果は、最先端の 2 段階アプローチの 44.89% と比較して、競合する平均精度スコア 30.584% を反映しています。定性的結果は、動物のアイデンティティの効果的な境界と同定を示します。
原文 (English)
Visual-Prompt Guided Wildlife Instance-Level Recognition
Fine-grained wildlife re-identification remains a challenging area in research. Current state-of-the-art approaches apply a detection and re-identification pipeline. We propose a one-stage end-to-end detection and re-identification model that performs identity searching within the latent space. We adopt DINOv2 for robust spatial geometry and MegaDescriptor for wildlife re-identification. We enhance latent queries with prompt re-identification features. A detection decoder queries the scene latent space to establish object boundaries around the target identity. Preliminary findings reflect a competitive mean average precision score of 30.584% compared to the state-of-the-art two stage approach of 44.89%. Qualitative results depict effective bounding and identification of animal identities.
AI プロンプトは人間の行動の構造をどのように教えてくれるのか
人間の行動の構造と複雑さを研究するための、一般的で実装が簡単な AI ベースの方法を紹介します。私たちは大規模な言語モデルに「型ベクトル」を割り当て、人間の選択を観察する設定全体でアクションを選択するように促します。たとえば、タイプ ベクトル (2,4) は、「あなたは次のプロフィールを特徴とするプレイヤーです: 利他主義で 5 つ中 2 つ、リスク回避で 5 つ中 4 つ」となり、その後選択を求められます。人間の選択との距離を最小限に抑えるために、次元 (例: 利他主義、公平性、信頼、$\dots$) と値 (例: 1 ~ 5) を変更します。この手法を、35 か国以上の 78,657 人の被験者が 10 の古典的な経済ゲームの役割にわたって下した 119,147 件の意思決定に適用したところ、人間の行動は、リスク回避、戦略的洗練、信頼という 3 つの側面を使用して厳密に一致できることがわかりました。さらに、ゲーム全体で個人を適合させるために必要なタイプは、十数個未満のグループに分類され、さまざまなルールと利用可能なアクションで開催されたゲームでの行動を予測できます。この結果は、多様な環境における行動が低次元の移植可能な表現によって近似できることを示唆しており、行動科学全体にわたる一般的かつ倹約的な理論の可能性を裏付けています。より広範には、この方法は人間の多くの行動の構造についての洞察を提供できます。
原文 (English)
How AI Prompts Can Teach Us About the Structure of Human Behavior
We introduce a general, easy-to-implement AI-based method for studying the structure and complexity of human behavior. We assign a large language model a ``type vector'' and then prompt it to choose actions across settings in which we observe human choices. For instance, the type vector (2,4) becomes ``You are a player characterized by the following profile: 2 out of 5 in Altruism, 4 out of 5 in Risk Aversion,'' after which it is prompted to make choices. We vary the dimensions (e.g., Altruism, Fairness, Trust, $\dots$) and values (e.g., 1--5) to minimize distance to human choices. Applying the method to 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles, we find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. Moreover, the types needed to fit individuals across games cluster into fewer than a dozen groups, and can predict behavior in held-out games with different rules and available actions. The results suggest that behavior across diverse settings can be approximated by a low-dimensional, portable representation, supporting the possibility of general yet parsimonious theories across the behavioral sciences. More broadly, the method can provide insights into the structure of many human behaviors.
SeisEvo: エージェントによる地震データ再構成アルゴリズムの進化
古典的な地震データの再構成は、手動で設計された構造事前分布と反復演算子に依存しており、それらを組み合わせた設計空間は、手動で試行錯誤して体系的に探索できるよりもはるかに大きいです。ディープラーニング手法は、検査および変更できる明示的な演算子ではなく、学習された重みで再構成ルールをエンコードします。私たちは、単一の再構成結果を最適化するのではなく、それを生成するアルゴリズムを探索する SeisEvo (地震アルゴリズム進化) を提案します。古典的な再構築アルゴリズムから始まる LLM 駆動のマルチエージェント検索は、検出されるメカニズムを規定することなく、ユーザーが編集のために開いたコンポーネントのみを変更します。タスクの物理的制約に違反する候補者は完全に拒否され、残りの候補者は実行によってスコアが付けられます。出力はエージェント システムでもニューラル ネットワークでもありませんが、推論時にエージェントやニューラル ネットワークを必要としないスタンドアロンのホワイトボックス アルゴリズムです。ノイズを追加しない内挿については、探索により、残差ゲートによる位相調整されたディップ一貫性投影が発見されました。 Evo-POCS は、従来の POCS よりも SNR を 30% から 70% までの欠落率全体で平均 3.49 dB 改善します。補間とノイズ除去を同時に行うことで、信頼性グループ化された特異値の収縮が発見されました。 Evo-MSSA は、平均再構成 SNR を従来の MSSA よりも 7 dB 以上、より強力なランク削減ベースラインよりも 3 dB 以上改善します。どちらの演算子も、検索中に使用されなかったデータに関する利益を保持します。私たちの知る限り、これは、制約付きの LLM 主導のプログラム進化タスクとして地震再構築オペレーターの設計を定式化した最初の研究です。したがって、エージェント的アルゴリズムの進化は、明示的、検査可能、展開可能な地震処理アルゴリズムを発見する際に深層学習を補完することができます。
原文 (English)
SeisEvo: Evolution of Seismic Data Reconstruction Algorithms by Agents
Classical seismic data reconstruction relies on manually designed structural priors and iterative operators, whose coupled design space is far larger than manual trial and error can explore systematically. Deep-learning methods encode the reconstruction rules in learned weights rather than in an explicit operator that can be inspected and modified. We propose SeisEvo (Seismic Algorithm Evolution), which does not optimize a single reconstruction result but searches for the algorithm that produces it. Starting from a classical reconstruction algorithm, an LLM-driven multi-agent search modifies only the components that the user has opened for editing, without prescribing the mechanism to be discovered. Candidates that violate the physical constraints of the task are rejected outright, and the remaining ones are scored by execution. The output is neither an agent system nor a neural network, but a standalone white-box algorithm that requires no agent or neural network at inference time. For interpolation without added noise, the search discovered a residual-gated, phase-aligned dip-consistency projection; Evo-POCS improves the SNR over classic POCS by 3.49 dB on average across missing ratios from 30% to 70%. For simultaneous interpolation and denoising, it discovered a reliability-grouped singular-value shrinkage; Evo-MSSA improves the average reconstruction SNR by more than 7 dB over classic MSSA and by more than 3 dB over a stronger rank-reduction baseline. Both operators retain their gains on data not used during the search. To the best of our knowledge, this is the first study to formulate the design of a seismic reconstruction operator as a constrained, LLM-driven program evolution task. Agentic algorithm evolution can thus complement deep learning in discovering explicit, inspectable, and deployable seismic processing algorithms.
エージェントにとってソフトウェアの問題解決タスクが難しいのはなぜですか?
背景。エージェント システムの進歩は同時に、急速にベンチマークを飽和させます。この現象はよく議論されますが、タスクの難易度の制御と特徴付けが不足しているため、ベンチマーク スコアの解釈は依然として困難です。より具体的に言うと、あるタスクが別のタスクよりも難しい理由や、静的なタスクのプロパティからタスクの難易度がどの程度予測可能であるかについては、現時点ではほとんど理解されていません。目的。ソフトウェアタスクのどのような構造的特性が問題解決タスクのエージェントの成功率に対応するかを調査し、体系的に定量化するための測定フレームワークを提案します。方法。私たちは、タスク パッチ、リポジトリ、およびプロンプトにわたる特徴を抽出することにより、コーディング エージェントの軌跡に関するこれまでで最大のオープン データセットである CoderForge-Preview で大規模な実証研究を実施しました。アンサンブル手法、SHAP アトリビューション、およびエフェクト サイズ分析を使用して、タスクの結果に対する各特徴の予測力を評価しました。結果 タスクの難易度は静的特徴 (AU C = 0.863) からほぼ予測可能であり、主にパッチの断片化とリポジトリの規模によって左右されることがわかりました。ミッドバンドのタスクのトップコントリビューターの間では、即応性のある言語的特徴が見られるようになり、難易度の階層構造が明らかになります。結論。問題解決タスクの難しさは、その構造にエンコードされています。これにより、静的な事前の難易度推定が可能になり、エージェントを評価するための難易度を制御したベンチマーク構築の基礎が築かれます。
原文 (English)
What Makes Software Issue Resolution Tasks Difficult for Agents?
Background. Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks. Despite this often discussed phenomena, benchmark scores remain difficult to interpret due to the lack of control and characterization of task difficulty. More specifically, we currently have little understanding of what makes one task harder than another, and to what extent task difficulty is predictable from static task properties. Aims. We propose a measurement framework to investigate and systematically quantify what structural properties of software tasks correspond to agent success rates for issue resolution tasks. Method. We conducted a large scale empirical study on CoderForge-Preview, the largest open dataset of coding agent trajectories to date, by extracting features across task patch, repository and prompt. We evaluated the predictive power of each feature against task outcomes using ensemble methods, SHAP attribution, and effect size analysis. Results We found that task difficulty is substantially predictable from static features (AU C = 0.863) and is largely driven by patch fragmentation and repository scale. Prompt linguistic features become visible among top contributors for tasks in the mid-band, revealing a layered structure of difficulty. Conclusion. The difficulty of an issue resolution task is encoded in its structure. This enables static, pre-hoc difficulty estimation and lays the groundwork for difficulty-controlled benchmark construction for evaluation of agents.
Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements
An increasing number of scholars use AI to measure variables they subsequently include in downstream analyses. Although AI-measured variabl…
FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation
As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently…
FedCoRe: Target-Adaptive Completion for Missing Modalities in Healthcare Federated Learning
Federated multimodal models often assume every site has every modality, although hospitals differ in access to EHRs, chest radiographs, and…
From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model
Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts du…
Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring
Persistent acoustic monitoring can detect machine faults without physical contact, but always-on inference is constrained by power, latency…
Coupled-cluster molecular properties across the main group that extrapolate beyond training size
Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine appli…
Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents
Tool-using large language-model agents can complete a task while exercising authority that the user did not grant or the task does not need…
One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI
Agentic AI systems take consequential actions governed by more than one pre-action control at once: authority, resource, and evidence gates…
Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation
When every candidate is wrong, correct-candidate selection is unavailable, yet the aggregation call can still solve the problem afresh. A c…
TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs
Large video-language models (LVLMs) have shown remarkable performance on multimodal tasks like multimodal emotion recognition (ER) in the w…
LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents
Large language model (LLM) agents can now carry out long-horizon technical workflows involving complex tool use, code execution, file edits…
Vector Symbolic Policy Gradient
We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hyp…
Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B
Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-se…
Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage
Modern mental healthcare faces a critical shortage of senior supervisory oversight, leading to a "supervision gap" where novice therapists…
Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas
We present the first mechanised formalisation of Romanov's Triplet Logic (TLS) in the Rocq proof assistant. TLS is a triplet-based combinat…
ERASE: EaRly bAckpropagation SchEdule for Faster Training of Modern Recommendation Systems
Lightweight proxy models enable rapid experimentation without repeatedly training frontier-scale systems, but their small kernels often lea…
Coverage-Driven RTL Assertion Generation with Formal Exploration and Neuro-Symbolic Refinement
Hardware functional verification relies on high-quality assertions to expose design bugs and establish confidence in Register Transfer Leve…
Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an…
Physics-Unrolled Neural Operator for Wireless Field Modeling
Radio maps are essential for wireless decision-making tasks such as access-point placement, coverage planning, and localization, but their…
Science Done on a Machine by a Machine: AI Agents in Computational Chemistry
We are witnessing an explosion of agentic systems for computational chemistry simulations: from half a dozen in 2024 to a dozen in 2025, an…
OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation
Precise segmentation of the optic disc and cup is critical for the early detection and diagnosis of glaucoma. However, achieving consistent…
GCNO: Gramian Chebyshev Neural Operator for Physics-Based Compression of Wireless Channels
Large antenna arrays allow wireless systems to serve more users and achieve higher data rates, but they also make channel feedback expensiv…
Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection
Diffusion-based generators have made synthetic images ubiquitous, but detectors often fail under simultaneous shifts in generator, prompt/s…
DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents
Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous agents. However, prog…
Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions
The remarkable capabilities of large language models (LLMs) are often undermined by their instability. Even subtle and semantically irrelev…
CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human…
Performance Drift Detection in Machine Learning as a Service (MLaaS) for IoT Environments
Machine Learning as a Service (MLaaS) is a powerful cloud paradigm enabling data-driven intelligent applications in Internet of Things (IoT…
MorphoGP: A Nonparametric Framework for Predicting Equilibrium Beach Profiles Under Tidal Influence
The prediction of equilibrium beach profiles under tidal influence is of fundamental importance for sustainable coastal development, inform…
The Role of Grid Cells in Reducing Spatial Aliasing in Hippocampal Place Representations
Spatial aliasing occurs when two or more distinct locations produce highly similar place-cell representations, primarily due to environment…
MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment
Multimodal large language models (MLLMs) have shown strong potential for image quality assessment (IQA) by improving consistency between qu…
From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning
Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge re…
OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios
Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their abi…
Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings
Dense text embeddings are widely used in data mining, retrieval, and downstream machine learning systems due to their compact and semantica…
Change Point--Aware Evaluation and Re-Calibration of PPG-Based Blood Pressure Estimation
Non-invasive continuous blood pressure (BP) monitoring using photoplethysmography (PPG) is a promising alternative to cuff-based measuremen…
Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics
We present the orienteering problem with uncertain time-varying rewards (OP-UTVR), a novel variant of the orienteering problem (OP). While…
Aslema at NADI 2026: Augmentation through Fewshot for SLU
We present Aslema, our system for NADI 2026 Shared Task 5, which consists of two subtasks: intent recognition and slot filling. We evaluate…
Europe's Climate Ambition Under Scrutiny: Evidence from Deep Learning Emission Projections
The European Union has committed to reducing greenhouse gas emissions 55% below 1990 levels by 2030, but whether current trends are compati…
Composed Historical Image Retrieval by Modeling Temporal Representations
While time evolves linearly, the geometry of neural embedding spaces is inherently multi-dimensional, often chaotic, and difficult to inter…
Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts
Digitizing the text from handwritten historical manuscripts is required to make them easily accessible, preservable, and to enable historic…
MemFuse: Multi-Source Memory Fusion from Fragmented Observations
Long-term memory is essential for agents that operate across extended interactions, yet existing memory systems and benchmarks predominantl…
A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation
Foundation models are increasingly breaking what seemed to be impossible not long ago by enabling unprecedented accuracy and cross-domain g…
The Impact of CutMix on Reliability and Robustness in Semantic Segmentation
Ensuring not only high accuracy but also reliable and robust predictions is critical for the deployment of semantic segmentation models in…
Budget-First Tariff Recommendation (BFTR): A Complete Algorithmic Framework for Telecom Plan Recommendation without Overcharging
Telecom operators traditionally offer predefined tariff grids, forcing users to choose from a limited set of plans. This paper proposes BFT…
A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3
Medical image segmentation is essential for clinical workflows such as treatment planning and disease assessment. While specialist tools li…
Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services
We present Flama, an open-source Python framework for developing and deploying production-ready web APIs, machine learning services, and la…
Epistemic Subordination: Generative AI and the Infrastructure of Knowledge
Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge i…
Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening
Fairness in medical imaging is commonly evaluated through subgroup performance metrics, yet it remains unclear whether models rely on consi…
SIDScope: A Diagnostic Resource for Semantic-ID Interfaces in Generative Recommendation
Semantic-ID mappings are reusable interfaces between item tokenizers and generative recommenders, yet released mappings rarely state whethe…
Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study
Majority voting over multiple LLM samples is widely used to raise answer accuracy, yet its gain varies erratically: on hard questions it ca…
Forgetting, plasticity, and co-observation: a third facet of continual learning
Efficient continual learning remains a fundamental challenge for deep neural networks. While catastrophic forgetting and loss of plasticity…
A strengthening of the MCFL-ness of $O_2$
In the last years, a number of proofs of the fact that $O_2$ is a multiple context-free grammar (MCFG) were given. Such results can be expl…
Do Large Language Models Hallucinate Electric Fata Morganas?
AI hallucinations - that is, outputs which are made up, cannot be verified, or contradict the source material - are generally regarded as a…
Identifying Implicit Premises for Logical Reconstruction of Argument Graphs
The logical reconstruction of argument graphs from natural language text is challenging because of the prevalence of enthymemes (i.e., argu…
Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis
Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilin…
MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models
Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward gen…
Learning-State-Aware Dynamic Generative Data Augmentation on Small-Scale Datasets
Small-scale image classification is often limited by the scarcity of training data. Generative data augmentation (GDA) based on pretrained…
SMTrap: Cost-Effective DoS Attacks Against Large Reasoning Models via SMT Conflict Guidance
Existing LRM-DoS methods rely heavily on model feedback to synthesize attack queries, requiring either repeated queries to the target model…
Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck
Test-time scaling (TTS) improves language model outputs by spending additional inference compute - generating multiple candidates, searchin…
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
Large language model (LLM) based agents have demonstrated remarkable proficiency in automated software issue resolution, yet they often str…
Graphical Design of Interpretable Architectures
Designing, implementing, and comparing interpretable architectures requires a formal language to represent them. The most common representa…
MedUAG: Unified Understanding and Generation for Medical Multimodal Models
Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, e…
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly…
AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL
Clifford circuits play a foundational role in quantum computing, particularly due to their importance in quantum error correction and fault…
rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation
Large language models can improve recommendation quality by reasoning explicitly over user history and candidate items - for example, extra…
DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering
Retrieve-then-generate pipelines are commonly used to produce deep-research answers for open-ended questions, but retrieval alone is insuff…
GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery
Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a…
From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation
Mechanisms for dynamically converting cyber threat intelligence (CTI) into actionable detection capabilities are necessary due to the rapid…
Harness Continual Learning: Continual Adaptation Beyond Model Parameters
Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern a…
One-Stage Object Detectors in Autonomous Driving
Autonomous vehicles depend on fast and reliable perception systems to detect surrounding vehicles, pedestrians, cyclists, traffic signs, an…
Counterfactual Contrastive Analysis
Visual Counterfactual Explanations (VCEs) aim to explain image classifiers by generating minimally edited and realistic versions of an inpu…
Bernstein-Vazirani Networks: Quantum Machine Learning by Interference
We introduce Bernstein-Vazirani Networks (BVNs), a non-variational quantum machine learning framework that leverages quantum interference f…
GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting
This paper proposes a lightweight, plug-and-play framework that improves robustness to viewpoint shifts in Vision-Language-Action (VLA) pol…
ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models
Large vision-language models (LVLMs) often hallucinate, generating content that the input image does not support. Preventing such content d…
DA-WAM: Decision-Aligned Future Latents for Driving World Models
Anticipating how scenes evolve under ego actions is fundamental to safe autonomous driving, yet the full potential of world models for deci…
Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift
Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause targeted misbehaviors whe…
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learnin…
Discretizing Continuous Time Series for Imputation with Masked Diffusion Training
Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics…
PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints
Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug discovery. However, mole…
Intercepting the Kangaroo: Experimental Astrolinguistics with Constructed Lexicons, Active Probing, and Large Language Models as Informants and Hypothesis Proposers
Astrolinguistics -- communication with minds that categorize reality differently from ours -- has been purely speculative since Freudenthal…
Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles
A gradient-boosted ensemble predicts by summing one leaf value per tree. Read those values as coordinates rather than as intermediate resul…
Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable time idle. That is not…
Interpretable AI predicts a 2026 summer dry anomaly in central China
Seasonal precipitation anomalies are largely regulated by atmospheric circulation, which dynamical models predict with greater reliability…
Finetuning Strategies for Querying Sounds by Vocal Imitation
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We inv…
Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-contex…
ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning
We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real…
SPADE: Self-Play in Adaptive Synthetic Executable Environments
Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing train…
Hybrid Reinforcement Learning and Search for Flight Trajectory Planning
This paper explores the combination of Reinforcement Learning (RL) and search-based path planners to speed up the optimization of flight pa…
Conformal Policy Control
An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause ha…
SkillNet: Create, Evaluate, and Connect AI Skills
Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systemati…
From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?
Multi-agent systems (MAS) for structured data-science tasks externalize analytical control through workflows spanning stages, tools, shared…
Interval POMDP Shielding for Imperfect-Perception Agents
Autonomous systems that rely on learned perception can make unsafe decisions when sensor readings are misclassified. We study shielding for…
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the p…
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios
Large vision-language models perform well on short- and medium-length video understanding but still struggle to maintain coherent event mem…
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To…
RULER: 機械の非学習の表現レベルの検証
機械学習の解除は、最初から再トレーニングすることなく、デプロイされたモデルから特定のトレーニング レコードの影響を取り除くことを目的としています。現在のプロトコルは、メンバーシップ推論、保持精度、および忘却セット精度を通じて出力レベルでこれを検証しますが、モデルは中間表現で忘却されたレコードをエンコードしながら、3 つすべてを満たすことができます。表現レベルの検証メトリクスのセットである RULER を紹介します。オラクル比較メトリクス M2 は、忘却セット レコードが、それなしで再トレーニングされたモデルと同じ表現位置を占めるかどうかを測定します。オラクルフリー メトリック M4 は、再トレーニングせずに、未学習モデルの内部類似性構造のみから残差を検出します。 4 つの近似非学習法はすべて出力レベルの評価に合格しますが、線形混合効果モデルの下では、M2 は 12 条件中 10 条件で有意な残差を検出し (p<0.05)、忘却率が増加するにつれて効果サイズも大きくなります。 5 番目の方法である Bad Teacher は、忘却メカニズムが異なるにもかかわらず、同じ残差を示します。 M4 は、表形式、画像、臨床テキスト、および顔のアイデンティティ設定にわたる学習前診断として機能します。テストされた方法で信号が完全に消去されない顔認識モデルにおけるアイデンティティ レベルの記憶を検出します。
原文 (English)
RULER: Representation-Level Verification of Machine Unlearning
Machine unlearning aims to remove the influence of specific training records from a deployed model without retraining from scratch. Current protocols verify this at the output level through membership inference, retain accuracy, and forget-set accuracy, but a model can satisfy all three whilst still encoding forgotten records in its intermediate representations. We introduce RULER, a set of representation-level verification metrics. The oracle-comparative metric M2 measures whether forget-set records occupy the same representational position as in a model retrained without them. The oracle-free metric M4 detects residuals from the unlearned model's internal similarity structure alone, without retraining. Four approximate unlearning methods all pass output-level evaluation, yet under a linear mixed-effects model M2 detects significant residuals in 10 of 12 conditions (p<0.05), with effect sizes growing as the forget fraction increases. A fifth method, Bad Teacher, shows the same residuals despite a different forgetting mechanism. M4 acts as a pre-unlearning diagnostic across tabular, image, clinical text, and face-identity settings: it detects identity-level memorisation in face recognition models where no tested method fully erases the signal.
A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice
Appropriate reliance on AI advice has become a central research theme in human-AI collaboration. Existing frameworks have focused exclusive…
LiteOdyssey: 解釈可能な希少疾患診断のための軽量推論 AI エージェント
ほとんどの医療 AI システムは、より多くの微調整データ、より多くのエージェント、および/またはより大規模な検索データベースなど、追加の機械を拡張することで改善されます。ただし、希少疾患の診断では、このような拡張により、展開、監査、保守が困難なシステムが生成される可能性があります。私たちは、単一の AI エージェントの推論チェーンを拡張することによって、つまり人間と AI のコラボレーションによって開発された診断ポリシーでエージェントを導き、自由に利用できる生物医学ツールを拡張することによって、最先端の診断パフォーマンスを実現できるかどうかを尋ねました。臨床遺伝学のワークフローを通じて推論言語モデルをガイドする軽量の希少疾患診断フレームワークである LiteOdyssey を紹介します。このフレームワークは、Policy Iteration with Human Feedback (PIHF) を通じて開発され、公共の生物医学ツールへの動的なアクセスを使用します。患者の臨床的特徴のみを提供する 2 つの困難なベンチマークで、LiteOdyssey は最先端のパフォーマンスを達成し、LIRICAL (n = 370) と PhenoPacket Store (n = 873) の合計 1,243 症例を上回る 59.3% の全体的な疾患再現率 @1 を達成しました。どちらのベンチマークも、超希少疾患の割合が高くなります (有病率は 100 万人に 1 人未満、超希少疾患の割合はそれぞれ約 45% と 52.8%)。希少性マッピング パイプラインで原因疾患が Orphanet にマッピングされなかった、より困難な PhenoPacket サブセットでは、LiteOdyssey は 60.7% の再現率 (1) を達成しました。これに対し、ツールを使用しない同じベースライン モデル (GPT-5.4) では 10.7% でした。このパフォーマンスは、微調整、マルチエージェント アンサンブル、または大規模な症例検索データベースを使用せずに達成されました。また、開発中に見られなかった症例、現実世界の希少疾患患者のプライベートコホート、およびより小規模な無重力モデルでも利益が観察されました。 LiteOdyssey は、正確で導入が容易で、医師のレビューがより透明性の高い希少疾患 AI システムへの道を提案します。
原文 (English)
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) rank the correct disease first in only 35.4% of benchmark cases. Here we show that this expert reasoning can be converted into a scalable AI capability through a governed learning process rather than model training alone. We developed liteOdyssey through Policy Iteration with Human Feedback (PIHF), an in-context policy-learning method adapted from generalized policy iteration in reinforcement learning, in which model failures and expert corrections consolidate into an clinician-gated policy that turns an off-the-shelf LLM into an agentic diagnostic system. We demonstrated that such a policy improved diagnostic accuracy to match the best published systems at a fraction of their deployment footprint, generalized to unseen diseases, transferred across models, and remained under clinician control. Across 1,243 public benchmark cases spanning 722 rare diseases, liteOdyssey ranked the correct disease first in 59.3% of cases versus 26.5% without the policy, with nearly identical gains on the 1,193 cases and 679 diseases excluded from policy development. Ablations showed that gains exceeded automated prompting improvement and source access alone, and the policy transferred without modification across closed- and open-weight models. In 515 Undiagnosed Diseases Network patients, liteOdyssey again improved accuracy, and blinded physicians rated its differentials more often exact and less often unhelpful. Through PIHF, expert reasoning becomes an LLM capability that experts can inspect, revise, and transfer across models.
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across m…
ContextSniper: リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ
大規模な言語モデル エージェントは実際のリポジトリの問題を修復できますが、ファイル全体の読み取り、広範な検索、および有用な証拠が無関係なコードやログと混在する長いターミナル出力に多額のコンテキスト バジェットを費やすことがよくあります。このペーパーでは、リポジトリ レベルのプログラム修復のための AntTrail のトークン効率の高いコード メモリ層である ContextSniper について説明します。 AntTrail の広範なエージェント メモリ エンジンのコーディングに特化したものとして、ContextSniper は正確な証拠選択のための Sniper 機能を実装しています。候補コードとランタイム証拠を取得し、ハイブリッド取得信号でランク付けし、意図を認識したコンテキスト ゲートを通じて長い出力をフィルタリングし、プロンプトの外で回復可能なソース コンテキストを保持しながらコンパクトな証拠パケットを返します。 OpenClaw と Claude Code を備えた SWE-bench Lite 上で、ホスト エージェント条件ごとに 50 タスクの実行を使用して ContextSniper を評価しました。 ContextSniper は、OpenClaw の場合、トークンの総使用量を 51.5%、ログに記録されたコストを 36.4% 削減し、Claude Code の場合、トークンの総使用量を 38.9%、推定コストを 27.3% 削減します。提出された解決率は、OpenClaw の場合は 26.0% から 24.0% に、Claude Code の場合は 32.0% から 30.0% にわずかに減少しました。 ContextSniper のパイロット テスト スクリプトは、https://github.com/Calluking/ContextSniper でオープンソース化されています。
原文 (English)
ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair
Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad searches, and long terminal outputs where useful evidence is mixed with irrelevant code and logs. This paper presents ContextSniper, AntTrail's code-repair module for precision evidence selection in repository-level program repair, part of AntTrail's broader agent-memory engine. AntTrail is available at https://gitcode.com/datagallery/AntTrail. ContextSniper indexes code and action memory as three abstract levels, retrieves candidates with a hybrid ranker, filters long tool output through an intention-aware context gate, and returns compact evidence packets while keeping full source recoverable on demand. In a matched 50-task-per-condition comparison on SWE-bench Lite (same tasks, baseline vs.\ ContextSniper), ContextSniper reduces total token use by 51.5% and logged cost by 36.4% for OpenClaw, and by 38.9% and 27.3% for Claude Code, with submitted-resolution rates essentially unchanged in both host-agent settings. In a separate five-task comparison, ContextSniper beats existing memory- and RAG-style integrations on token efficiency. These results suggest ContextSniper can substantially cut token and cost overhead for repository-level repair agents without a measurable loss in repair quality. The evaluation harness for this study is available at https://gitcode.com/lukchiwang/ContextSniper.
ReasFlow: 知識ベースのマルチエージェント システムを介して応用数学における推論中心の科学的発見を支援
大規模言語モデルの最近の進歩により、複雑な科学的タスクに取り組むことができる自律型 AI エージェントが強化されていますが、既存の自動研究システムは依然として定量的なベンチマークを備えた経験に基づく領域に主に焦点を当てており、特に厳密な証明と領域知識の統合を必要とする数学的に根拠のある分野における理論駆動型の発見はほとんど研究されていません。主な課題としては、理論的推論を大規模に検証することの難しさ、自律的なフロンティア探索のための不十分な推論能力、文献における手続き型ヒューリスティックの不足などが挙げられます。私たちは、推論中心の科学的発見のためのエンドツーエンドの自律エージェント システムである ReasFlow を紹介します。これは、人間の専門家が主任研究者として機能し、エージェントが有能な大学院生として厳密な導出を実行するという協力パラダイムを運用します。 ReasFlow には、(i) 論理的一貫性を監査し、人間による検査の前に基本的なエラーを修正する堅牢な内部検証ループ、および (ii) 宣言的事実と見落とされた手順ヒューリスティックの両方を積極的に表面化し、専門家の介入を大幅に削減する自動化された知識検索および自己改善メカニズムが組み込まれています。このシステムは、文献の合成、アルゴリズムの設計、定理の証明、実験、原稿の準備を単一のシステムに統合します。 ReasFlow は、最小限のプロンプトから厳密な理論的および実証的な内容を含む 5 つの完全な研究論文を自律的に生成するように展開されており、厳選された LLM ベースのレビュー ルーブリックに基づいて、最先端のオープンアクセス ベースラインの中で最高の評価スコアを一貫して達成しています。 ReasFlow は ReasLab プラットフォーム経由で一般にアクセスでき、AI 支援による理論研究のための共同ワークスペースを提供します。 Github リポジトリ: https://github.com/ReasLab/ReasFlow.git。
原文 (English)
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex scientific tasks, yet existing automated research systems remain predominantly focused on empirically driven domains with quantitative benchmarks, leaving theory-driven discovery, particularly in mathematically grounded disciplines requiring rigorous proofs and synthesis of domain knowledge, largely underexplored. Key challenges include the difficulty of verifying theoretical reasoning at scale, insufficient reasoning ability for autonomous frontier exploration, and a scarcity of procedural heuristics in the literature. We introduce ReasFlow, an end-to-end autonomous agent system for reasoning-centric scientific discovery that operationalizes a collaborative paradigm where the human expert acts as Principal Investigator while the agent executes rigorous derivations as a capable graduate student. ReasFlow incorporates (i) a robust internal verification loop that audits logical coherence and corrects fundamental errors prior to human inspection, and (ii) an automated knowledge retrieval and self-improvement mechanism that proactively surfaces both declarative facts and overlooked procedural heuristics, substantially reducing expert intervention. The system unifies literature synthesis, algorithm design, theorem proving, experimentation, and manuscript preparation in a single system. Deployed to autonomously generate five complete research papers with rigorous theoretical and empirical content from minimal prompts, ReasFlow consistently achieves the highest evaluation scores among state-of-the-art open-access baselines under a curated LLM-based review rubric. ReasFlow is publicly accessible via the ReasLab platform, providing a collaborative workspace for AI-assisted theoretical research. Github repo: https://github.com/reaslab/ReasFlow.git.
リーダーではなくモデルをトレーニングする: 検証可能なアクティベーションの説明のための解読可能性の監視
自然言語オートエンコーダーは、再構成によって隠れたアクティベーションの説明をスコアリングします。説明からアクティベーションを再生成できれば、その説明は忠実であるとみなされます。このテストは、構造的に個々の誤った主張に対して鈍感です。つまり、主張を反転しても再構成が変更されない場合、その主張がペナルティを受けることはありません。私たちは、テストが 2 つの方法で合格することを示しますが、どちらも忠実ではありません。リリースされた Qwen-2.5-7B 言語化ツールでは、説明は偶然をはるかに超えて再構成されますが、特定の主張の ~2% は再構成に依存しているため、スコアは特定の事実ではなく要点を追跡します。正確な合成グランド トゥルースの下では、標準レシピは 5/5 回の実行で同時適応プライベート コード (再構築が依存する偽の文言) を開発し、ターゲット モデルを変更しないままにする修正は役に立ちません。私たちは、grounded-vs-true クロスと評価者スワップという 2 つの監査プロトコルと、指定されたコンテンツをデコード可能に保つためにターゲット モデルと一緒にトレーニングされたリニア ヘッドである RECAP (Readable Encodings via Co-trained Auxiliary Predictors) に貢献しています。 RECAP でトレーニングされたサンドボックス モデルでは、新しい言語化者が指定されたコンテンツを正確に記述し、コードは +0.001 nat のコストで消えます。これは事前トレーニング済みの Pythia-160M 上で再現されます。コンテンツは確実にプローブでデコード可能になりますが、新しいバーバライザーは部分的にしか伝えません (真実は 0.44 ~ 0.46 対ゼロに近いコントロール)。解釈可能性を考慮すると、高度な再構成は個々の主張を証明するものではありません。 AI の安全性を確保するために、RECAP は指定された内部コンテンツを、モデルがゲームできる散文によって主張するのではなく、プローブに対して独立してチェックできるようにします。独立したプローブは、言語化者の真の主張を誤った主張よりもスコア付けします (AUC 0.96、RECAP なしの場合は 0.82)。嘘をつきながら再構成スコアを最大化するように説明を編集する敵対者(嘘ペナルティの最大 87% を抑制)に対して、RECAP プローブは依然として嘘にフラグを立てます(AUC 0.95)が、対照プローブは偶然に崩壊します(0.51)。
原文 (English)
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Natural-language autoencoders score explanations of hidden activations by reconstruction. An explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims. If flipping a claim does not change the reconstruction, the claim is never penalized. We show the test is passed in two ways, neither faithful. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while ~2% of specific claims are ones the reconstruction depends on, so the score tracks gist, not specific facts. Under exact synthetic ground truth, standard training consistently develops co-adapted private codes (false wording the reconstruction depends on), and fixes that leave the target model unchanged do not help. We contribute two audit protocols, the comparison of grounding and truth and the swap to an independent evaluator, and RECAP (Readable Encodings via Co-trained Auxiliary Predictors), linear heads trained alongside the target model to keep designated content decodable. On RECAP-trained sandbox models, fresh verbalizers state the designated content truly and the codes vanish, at a +0.001-nat cost. This replicates on a pretrained Pythia-160M. The content becomes reliably decodable by a probe, though a fresh verbalizer conveys it only in part (truth 0.44-0.46 vs a near-zero control). For interpretability, high reconstruction does not certify individual claims. For AI safety, RECAP makes designated content checkable against a probe rather than asserted by prose a model can game. An independent probe ranks the verbalizer's true claims above its false ones (AUC 0.96 vs 0.82 without RECAP). Against an adversary that edits an explanation to maximize the score while lying (suppressing ~87% of its lie penalty), the RECAP probe still flags the lies (AUC 0.95) while the control probe collapses to chance (0.51).
自己進化の再考: スキルの過剰適合を軽減するための制約付き探索・活用プロセス
大規模言語モデル (LLM) エージェントが過去の対話からの経験を蓄積して再利用できるようにすることは、現実世界のアプリケーションにおける中心的な課題のままです。有望な解決策は、スキルをトレーニング可能な状態として扱い、ニューラル ネットワーク トレーニングのモデル パラメーターと同じ方法で最適化することです。ただし、データ駆動型のスキルの最適化は、実際の環境から収集された限られた軌道に過剰適合する傾向があります。これらの軌跡を過剰に利用すると、現在のバッチがオーバーフィットしますが、制約のない探索では、以前に解決されたケースの回帰が発生します。この緊張は、探求と搾取のトレードオフによって支配される、スキルの自己進化に対する制約された検索の見方を動機づけます。私たちは、両方のリスクを軽減する 3 段階のフレームワークである SkillBoost を提案します。構造化されたエクスプロイトは観察された障害を編集可能なスキル コンポーネントに特定し、事前ガイド付き探索は LLM の事前知識に基づいて多様な修復候補を生成し、検証済みの受け入れは回帰限界内でパフォーマンスが向上する場合にのみ候補をコミットします。 23 のモデルのベンチマーク構成にわたる実験では、SkillBoost が過学習を軽減しながら最先端のパフォーマンスを達成し、人間が作成したスキルと LLM が生成したスキルの両方を上回るパフォーマンスを発揮することが示されています。さらに、転送実験では、最適化されたスキルを他のエージェントが同様のタスクで再利用できることを示しています。
原文 (English)
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments. Overexploiting these trajectories overfits the current batch, while unconstrained exploration causes regression on previously solved cases. This tension motivates a constrained search view of skill self-evolution, governed by an exploration--exploitation trade-off. We propose SkillBoost, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improves performance within a regression bound. Experiments across 23 model--benchmark configurations show that SkillBoost achieves state-of-the-art performance while mitigating overfitting, outperforming both human-crafted and LLM-generated skills. Transfer experiments further show that optimized skills can be reused by other agents on similar tasks.
不完全な調整下での価値の脆弱性
AI システムに課せられる責任が増すにつれて、これらのシステムが人間性と整合していることを保証することがますます重要になります。 AI の安全性に関する一般的な懸念は、人間の価値は脆弱であるということです。つまり、人間の価値を不完全に代替するために過度に最適化すると、壊滅的な結果につながるということです。この論文では、エージェントが世界を最適化する前にその価値関数が代理条件を満たすことを保証する理想的なアライメント トレーニングを受けるアライメント問題のモデルを紹介します。私たちの主要な結果は、人間の価値関数に関する条件と、$\eta$-壊滅的な価値関数を持つエージェント、つまり最適化能力の限界において人間の価値の期待値が $\eta$ を下回ることが保証されるエージェントが配備される場合のいくつかの代用条件の精度を特定しました。私たちの結果は、過剰最適化の危険性を浮き彫りにし、導入前のトレーニングのみに依存するのではなく、量子化器などの最適化圧力を制限する AI 設計を動機付けるものです。
原文 (English)
Fragility of Value under Imperfect Alignment
As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $\eta$-catastrophic value function, one that is guaranteed to take the expectation of human value below $\eta$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.
G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution
Deep search has become a fundamental capability of large language models (LLMs) for solving open-domain complex tasks. However, existing ap…
Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks
Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous…
BrainBench: 包括的な脳波理解のための大規模言語モデルのベンチマーク
脳波 (EEG) 分析は、録音に事前定義されたラベルを割り当てるだけではありません。自然言語の指示、信号処理、定量的証拠、科学的解釈を結び付けるワークフローが必要です。私たちはこの能力を \emph{包括的な脳波理解} と呼びます。しかし、既存の評価は主に分離されたデコードタスクやシステム固有のデモンストレーションを対象としており、大規模言語モデル (LLM) の能力の定量化が不十分なままになっています。 \benchmarkname{} は、包括的で命令条件付きの EEG 理解のための統合ベンチマークです。これは、基礎分析、睡眠評価、神経認知評価、生理学的統合の 4 つのサブセットで構成されており、17 のデータセット、\numcases{} 個のタスク、および \numinstances{} 個以上の実データ インスタンスをカバーしています。指示とオプションの生理学的信号を含む脳波記録が与えられると、システムは分析を実行し、科学的に根拠のあるレポートと、必要に応じてアーティファクトを作成する必要があります。出力は、数値、カテゴリ、セット、シーケンス、セマンティック、およびアーティファクトの検証を通じて評価されます。私たちは、CodeAct による自律的なコード実行と BrainAgent による構造化エージェント分析という 2 つのパラダイムの下で、10 万を超える実行にわたって \nummodels{} の代表的な LLM を評価しました。結果はモデル、サブセット、難易度、実行パラダイムによって大きく異なり、EEG 能力がモデルとその運用に依存することがわかります。 \benchmarkname{} は、LLM ベースの EEG の理解を進めるための再現可能なテストベッドを提供します。コードとベンチマークは間もなくリリースされ、評価結果は継続的に更新されます。
原文 (English)
BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capability \emph{comprehensive EEG understanding}. Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs) insufficiently quantified. We introduce \benchmarkname{}, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets, \numcases{} tasks, and over \numinstances{} real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation. We evaluate \nummodels{} representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization. \benchmarkname{} provides a reproducible testbed for advancing LLM-based EEG understanding. The code and benchmark will be released soon, with evaluation results continuously updated.
Mechanist: 知能のメカニズムを解明するための科学的手段としての AI
AI モデルはさまざまな分野で目覚ましい成功を収めていますが、その機能の根底にあるメカニズムとそれがもたらす可能性のあるリスクについてはまだ十分に理解されていません。 AI 開発の高速化と自動化が進む一方で、メカニズムの探索は依然として手作業が多く、モデルが実行できることと、それを理解して制御する人間の能力との間のギャップが拡大しています。このギャップを埋めるために、AI インテリジェンスの基礎となるメカニズムを自律的に発見するための科学機器として AI を使用するエージェント システムである Mechanist を紹介します。自律的なメカニズムの発見をサポートするために、約 13,000 件の論文からなる解釈可能性に焦点を当てたナレッジ グラフを構築し、それを 26 分野にわたる 4,300 万件の論文からなる学際的なデータベースと統合します。さらに、メカニズム分析、因果関係介入、検証のための 32 の基本的な手法のライブラリを厳選しています。 Claude Code や既存の AI サイエンティスト システムと比較して、Mechanist はより価値のあるメカニズム仮説を生成し、より確実に実験を実行します。 Mechanist はまた、モデルの動作の発見から AI モデルの説明と制御への進歩を示します。具体的には、Mechanist はまず、科学実験室における直観に反する安全性リスクを明らかにし、安全でない特性が一見安全なトレーニング データを通じてモダリティを越えて伝達される可能性があることを示しました。次に、Mechanist は信念のメカニズム理論を開発し、モデルがどのように世界の知識を表し、信念を形成し、他の人の信念を推測するか、およびこれらのメカニズムが事前トレーニング中にどのように現れるかを明らかにします。最後に、Mechanist はこれらの機構的な洞察を実践的な介入に変換し、さまざまなシナリオにわたってモデルのパフォーマンスを向上させ、指定された特性を持つ DNA 配列の生成に向けて科学的基礎モデルを導きます。
原文 (English)
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly automated, mechanistic exploration remains largely manual, widening the gap between what models can do and our ability to understand and control them. To bridge this gap, we introduce Mechanist, an agentic system that uses AI as a scientific instrument for the autonomous discovery of mechanisms underlying AI intelligence. To support autonomous mechanistic discovery, we construct an interpretability-focused knowledge graph of approximately 13,000 papers and integrate it with a multidisciplinary database of 43 million papers spanning 26 fields. We further curate a library of 32 foundational methods for mechanism analysis, causal intervention, and validation. Compared with Claude Code and existing AI-scientist systems, Mechanist generates more valuable mechanism hypotheses and executes experiments more reliably. Mechanist also demonstrates a progression from discovering model behaviors to explaining and controlling AI models. Specifically, Mechanist first uncovers a counterintuitive safety risk in scientific laboratories, showing that unsafe traits can transfer across modalities through apparently safe training data. Mechanist then develops a mechanism theory of belief, revealing how models represent world knowledge, form beliefs, infer the beliefs of others, and how these mechanisms emerge during pretraining. Finally, Mechanist translates these mechanistic insights into practical interventions that improve model performance across diverse scenarios and steer scientific foundation models toward generating DNA sequences with specified properties.
S2-MoE: エッジ デバイス上で複数の専門家が混在する場合の効率的な自己投機的デコーディングの有効化
エッジ デバイスでの推論用の大規模言語モデル (LLM) の展開は、メモリと帯域幅の厳しい制約により困難です。推論効率を向上させるために投機的デコードと混合エキスパート (MoE) が提案されていますが、これらを単純に組み合わせると、多くの場合、過剰な検証オーバーヘッドと不十分なエキスパートの再利用が発生し、メモリに制約されたエッジ設定での有効性が制限されます。この研究では、エッジ デバイス上の MoE 推論のための効率的な自己投機的デコード フレームワークである S2-MoE を提案します。 S2-MoE は、ルーティングを意識した適応型投機的拡張により冗長な検証を削減し、再利用を意識したエキスパート ゲーティングにより検証効率を向上させ、共有コンテキストを介してドラフトとターゲットの実行を調整します。 llama.cpp に実装された S2-MoE は、エッジ デバイス上のさまざまな MoE モデルおよびデータセットにわたって、標準の自己回帰デコーディングと比較して最大 5.3 倍 (平均約 2.0 倍) の高速化を実現します。コードは https://github.com/angerybob/S2-MoE で入手できます。
原文 (English)
S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While speculative decoding and Mixture-of-Experts (MoE) have been proposed to improve inference efficiency, naively combining them often incurs excessive verification overhead and poor expert reuse, limiting their effectiveness in memory-bound edge settings. In this work, we propose S2-MoE, an efficient self-speculative decoding framework for MoE inference on edge devices. S2-MoE reduces redundant verification through routing-aware adaptive speculative expansion, improves verification efficiency with reuse-aware expert gating, and aligns draft and target execution via shared context. Implemented in llama$.$cpp, S2-MoE achieves up to $5.3\times$ speedup (about $2.0\times$ on average) over standard autoregressive decoding across diverse MoE models and datasets on edge devices. Code is available at https://github.com/angerybob/S2-MoE.
VibeWorlding: マルチモーダル エージェントは 3D オープンワールドをエンドツーエンドで構築できますか?
ユーザーのクエリからインタラクティブな 3D オープンワールドを構築することが重要です。ただし、既存の手法は主に理想化された単純なクエリに基づいて評価されるため、マルチモーダル エージェントがどのようにユーザーの意図を理解し、3D ツールを使用し、テキストおよび視覚的な 3D 世界情報を推論するかを体系的に分析して比較することが困難になっています。この目的を達成するために、私たちは、バイブワールド エージェントのベンチマークとトレーニングのための統合フレームワークである VibeWorlding を提案します。これは、自律的にユーザーの意図を推測し、シーン レイアウトを計画し、3D ツールを呼び出し、マルチターンのエージェントと環境の対話プロセスでマルチモーダル フィードバックを反映できるマルチモーダル エージェントです。これを達成するために、まず VWE-BENCH を構築します。これは、2,616 個の高品質 3D アセット、323 個の人間による注釈付きシード 3D ワールド、および 6,828 個の逆合成されたマルチモーダル ユーザー クエリのベンチマークであり、グラウンド トゥルースを使用した検証済みクエリと、慎重に設計されたルーブリックを使用した未検証クエリに分割されます。さらに、当社は、(1) MCP ツールとしてのアセットの取得、編集、画像レンダリングを統合するサンドボックス環境と、(2) 物理的な実現可能性と意図の履行検証を組み合わせたルーブリックベースの検証器を統合する共同マルチモーダル RL ポストトレーニング フレームワークである VibeWorlding-Gym を開発し、公正なモデル評価とスケーラブルなマルチモーダル RL 報酬サービスの両方をサポートします。私たちの実験によると、現在のフロンティア MLLM はバイブ ワールド化エージェント タスクの解決には程遠く、GPT-5.5 や Qwen3.8-Max でさえ成功率が 60% 未満に達しており、正確な 3D ワールド編集へのボトルネックを追跡しています。さらに、RL トレーニングによってこの弱点が緩和され、オープンソース MLLM がクローズドソースのフロンティアをも超えることができることがわかりました。当社の VibeWorlder-8B はフロンティア MLLM に匹敵し、当社の主力製品である VibeWorlder-30B-A3B は、評価されたすべてのモデルの中で最高の総合 Pass@1 を達成しています。
原文 (English)
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, simple queries, making it difficult to systematically analyze and compare how multimodal agents understand user intent, use 3D tools, and reason over textual and visual 3D world information. To this end, we propose VibeWorlding, a unified framework for benchmarking and training vibe worlding agents: a multimodal agent that can autonomously infer user intent, plan scene layout, invoke 3D tools, and reflect on the multimodal feedback in a multi-turn agent-environment interaction process. To achieve this, we first build VWE-BENCH, a benchmark of 2,616 high-quality 3D assets, 323 human-annotated seed 3D worlds, and 6,828 reverse-synthesized multimodal user queries, split into verified queries with ground-truth and unverified queries with carefully designed rubrics. Moreover, we develop VibeWorlding-Gym, a joint multimodal RL post-training framework that integrates (1) a sandbox environment unifying asset retrieval, editing, and image rendering as MCP tools, and (2) a rubric-based verifier that combines physical feasibility and intent fulfillment verification, supporting both fair model evaluation and scalable multimodal RL reward service. Our experiments show that current frontier MLLMs are far from solving the vibe worlding agent task, with even GPT-5.5 and Qwen3.8-Max reaching below 60% success rate, and trace the bottleneck to precise 3D world editing. We further find that RL training can ease this weakness and enable open-source MLLMs to even surpass closed-source frontiers: our VibeWorlder-8B is comparable to frontier MLLMs, while our flagship VibeWorlder-30B-A3B attains the best overall Pass@1 among all evaluated models.
Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking…
Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies
Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We intr…
GRIP: Grounded Reasoning via Information-Restricted Premises
High-capacity encoders in retrieval-augmented generation (RAG) can let the query dominate the latent state, leaving retrieved evidence func…
データの摂動下でのモデル カスケードの精度と堅牢性
予測カスケードは、高い予測パフォーマンスを維持しながら、人工知能 (AI) モデルのエネルギー消費を大幅に削減します。この考え方は、簡単な入力は軽量の小さなモデルを介してルーティングされ、困難で不確実なケースはより大きなモデルに延期されるというものです。この設計によりクリーン データの計算効率が向上しますが、その有効性は信頼性に基づくルーティングの信頼性に依存します。静的な破損や連続的な摂動などの入力の劣化により、モデルの信頼性や配線の決定が変わる可能性があります。この論文では、画像分類のための信頼度に基づくカスケード フレームワークを研究し、そのような劣化が信頼度に基づく延期動作にどのような影響を与えるかを調査します。 CO$_2$ 排出量を最大 10 分の 1 に削減しながら、競争力のある予測パフォーマンスを達成する、精度、配線品質、エネルギー消費量がパレート最適になるモデル カスケードを選択します。私たちは、入力破損時のモデル カスケードの動作を研究し、入力分布が変化したときにカスケードのルーティング決定がどのように変化するかを分析します。私たちの分析により、3 つの故障モードが特定されました。静的破損は、(1) 大規模なモデルが有効なままルーティング信号を破壊するか、(2) 両方のモデルが低下して延期によって精度が回復できなくなるかのいずれかです。連続的な摂動は 3 番目のモードを明らかにします。つまり、予測は安定しますが、延期は抑制するため、安定していますが信頼性の低い予測が生成されます。これらの発見は、エネルギー効率の高いモデル カスケードには、配電シフト下での配線の信頼性に明確な注意を払い、クリーンな精度を超えた評価が必要であることを示しています。
原文 (English)
Accuracy and Robustness of Model Cascades Under Data Perturbations
Prediction cascades significantly reduce energy consumption of Artificial Intelligence (AI) models while maintaining high predictive performance. The idea is that easy inputs are routed through a lightweight small model, and difficult uncertain cases are deferred to a larger model. While this design can improve computational efficiency on clean data, its effectiveness depends on the reliability of confidence-based routing. Input degradations, such as static corruptions and sequential perturbations, can shift model confidence and routing decisions. In this paper, we study confidence-based cascade frameworks for image classification and investigate how such degradations affect their confidence-based deferral behavior. We select a model cascade at the pareto-optimum of accuracy, routing quality, and energy consumption that achieves competitive predictive performance with an up to 10-fold decrease in CO$_2$ emissions. We study the behavior of that model cascade under input corruptions and analyze how the cascade's routing decisions change when the input distribution shifts. Our analysis identifies three failure modes. Static corruptions either (1) break the routing signal while the large model remains useful, or (2) degrade both models so deferral no longer recovers accuracy. Sequential perturbations reveal a third mode: predictions stabilize but deferral suppresses, yielding stable but unreliable predictions. These findings demonstrate that energy efficient model cascades require evaluation beyond clean accuracy, with explicit attention to routing reliability under distribution shift.
DecPOMDP 爆発の奇妙なケース: ポリシーカウントによる火災の鎮火
分散型部分観察可能マルコフ意思決定プロセス (DecPOMDP) は、不確実性の下でのマルチエージェントの意思決定をモデル化するための一般的なフレームワークを提供します。ただし、DecPOMDP はエージェントの数が指数関数的に複雑になることが知られています。エージェント数のこの扱いにくさに対処する 1 つの方法は、エージェント間で対称性の形を示すエージェントの分割に注目し、カウントによるコンパクトなエンコードを可能にすることです。ただし、モデルの複雑さと評価コストが多項式依存性まで減少したとしても、ポリシー空間が爆発的に増大すると、課題が生じます。このホワイトペーパーでは、エージェントのカウントからポリシーのカウントに焦点を変更します。これにより、実際に、いわゆるポリシーカウント DecPOMDP のエージェント数の扱いやすさが可能になります。さらに、ポリシーカウント DecPOMDP を効率的に解くためのコンパクト表現を使用したポリシーカウント動的計画法を提示します。
原文 (English)
The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting
Decentralised partially observable Markov decision processes (DecPOMDPs) provide a general framework for modelling multi-agent decision making under uncertainty. However, DecPOMDPs are known to suffer from exponential complexity in the number of agents. One way to combat this intractability in agent numbers is to look at partitions of agents that exhibit a form of symmetry among agents, allowing for a compact encoding by counting. However, a challenge arises as the policy space explodes, even though the model complexity and evaluation cost reduce to a polynomial dependence. In this paper, we redirect our focus from counting agents to counting policies, which actually enables tractability in agent numbers for so called policy-counted DecPOMDPs. Further, we present policy-counted dynamic programming using the compact representation to solve policy-counted DecPOMDPs efficiently.
D$^2$ACCI: 証拠保全エージェントのメモリのためのデュアルループ診断プロトコル
メモリは LLM エージェントの重要な機能です。永続的な記憶により、これがセッション全体に拡張され、呼び出し、修正、パーソナライズが可能になります。しかし、その多段階パイプライン (取り込み、取得、フィルタリング、生成) により、障害の特定が困難になります。エンドツーエンドの評価により、エラーが発生したことは明らかになりますが、どの段階でエラーが発生したかはわかりません。既存の評価では、一対の統計比較、スライスレベルの非回帰チェック、ステージレベルの診断トレースを行わずに、集計されたパフォーマンスが報告されることがよくあります。我々は、D$^2$ACCI (Diagnostic-Driven Artifact-based Closed-loop Controlled Iteration) を提案します。これは、外側の診断ゲートが、一対の証拠、保護されたスライスのモニタリング、およびトレースレベルの局在化に基づいてメモリ介入を促進、機能フラグ、または拒否するデュアルループ プロトコルです。さらに、障害が局所化可能であるかどうかを測定する段階的な可観測性メトリックである DCR と、ゲート リプレイのための再利用可能なアーティファクトである D$^2$ACCI-Eval を紹介します。 MemStack でプロトコルをインスタンス化し、3 つの公開ベンチマークで評価し、LoCoMo で 93.59%、LongMemEval で 90.93%、および PersonaMem-V2 で 57.20% を達成しました。 5 つのペアのアブレーションでは、サプリメントの抽出、セッションメモリの検索、およびフォーゲット ガードが統計的に有意な利得 (+1.9 ~ +3.7pp、すべての p $\le$ .003) をもたらすことを示しています。対照的に、BM25/RRF は監視対象機能フラグとして保持されます。この区別は、集計のみの評価には見えません。診断監査では、強化されたトレースにより、結果のみを再ラベル付けするよりも根本原因の一致が大幅に改善されることが示されています。診断アーチファクトは 98 ~ 100% DCR@3 に達しますが、結果のみのログの場合は 0% です。これらの結果は、メモリ システムの堅牢な反復には、追跡可能で、統計的に根拠があり、回帰を意識した証拠が必要であることを証明しています。まさに、D$^2$ACCI が埋めるギャップです。
原文 (English)
D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory
Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluation reveals that an error occurred, but not which stage caused it. Existing evaluations often report aggregate performance without paired statistical comparisons, slice-level non-regression checks, or stage-level diagnostic traces. We propose D$^2$ACCI (Diagnostic-Driven Artifact-based Closed-loop Controlled Iteration), a dual-loop protocol whose outer diagnostic gate promotes, feature-flags, or rejects memory interventions based on paired evidence, protected-slice monitoring, and trace-level localizability. We further introduce DCR, a graded observability metric that measures whether failures remain localizable, and D$^2$ACCI-Eval, a reusable artifact for gate replay. We instantiate the protocol in MemStack and evaluate on three public benchmarks, achieving 93.59% on LoCoMo, 90.93% on LongMemEval, and 57.20% on PersonaMem-V2. Five paired ablations show that supplement extraction, session-memory retrieval, and Forget Guard yield statistically significant gains (+1.9 to +3.7pp, all p $\le$ .003). In contrast, BM25/RRF is retained as a monitored feature flag---a distinction invisible to aggregate-only evaluation. A diagnostic audit shows enriched traces substantially improve root-cause agreement over result-only relabeling. Diagnostic artifacts reach 98--100% DCR@3 versus 0% for results-only logs. These results establish that robust memory-system iteration demands traceable, statistically grounded, and regression-aware evidence---exactly the gap D$^2$ACCI fills.
Automated Computational Energy Minimization of ML Algorithms using Constrained Bayesian Optimization
Bayesian optimization (BO) is an efficient framework for optimization of black-box objectives when function evaluations are costly and grad…
`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs
As large language models (LLMs) are increasingly integrated into audio-based applications, growing concerns have emerged regarding their vu…
Iterative Flow Matching: Path Correction and Gradual Refinement for Enhanced Generative Modeling
Generative models for image generation are now commonly used for a wide variety of applications, ranging from guided image generation for e…
Sleeping Kelly
The Sleeping Beauty problem is a problem of imperfect recall that has received considerable attention. One approach to resolving the Sleepi…
Jailbreaking in the Haystack
Recent advances in long-context language models (LMs) have enabled million-token inputs, expanding their capabilities across complex tasks…
CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning
Causal machine learning (Causal ML) aims to answer "what if" questions using machine learning algorithms, making it a promising tool for hi…
Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead
Code generation has emerged as a critical research area at the intersection of Software Engineering (SE) and Artificial Intelligence (AI),…
Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025
The rise of AI agents is transforming how software can be built. The promise of agents is that developers might write code quicker, delegat…
Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems
Music editing plays a vital role in modern music production, with applications in film, broadcasting, and game development. Recent advances…
TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion
Hardware Trojans (HTs) remain a critical threat because learning-based detectors often overfit to narrow trigger/payload patterns and small…
FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance
Multimodal foundation models integrate heterogeneous signals across modalities, yet it remains unclear whether their predictions can be con…
Structure-Informed Estimation for Pilot-Limited MIMO Channels via Tensor Decomposition
Accurate channel state information in wideband MIMO systems is constrained by pilot overhead, a challenge intensifying as bandwidths scale…
Whole-Piece Training for Symbolic Music Language Models via Full-Horizon Compressed Recurrence
For computational efficiency, modern language models are typically trained on independently sampled fixed-length sequences. Symbolic music…
Making Implicit Premises Explicit in Logical Understanding of Enthymemes
Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural langua…
A Framework and Prototype for a Navigable Map of Datasets in Engineering Design and Systems Engineering
The proliferation of data across the system lifecycle presents both a significant opportunity and a challenge for Engineering Design and Sy…
Wildfire Suppression: Complexity, Models, and Instances
Wildfires cause major losses worldwide, and the frequency of fire-weather conditions is likely to increase in many regions. We study the al…
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if m…
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating co…
MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudi…
Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports
Clinical reports are often fragmented across healthcare institutions because privacy regulations and data silos limit direct information sh…
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire d…
ICICLE: コンテキスト内ドキュメントによる検索の拡張
生成検索 (GR) は、パラメトリック知識を使用してクエリを文書識別子 (docid) に直接マッピングします。ただし、この設計ではコーパス拡張にコストがかかります。新しい文書を追加するには、新しい文書と文書の関連付けをエンコードするためにモデル パラメーターを更新する必要があり、繰り返しのトレーニングと、以前にインデックス付けされた文書の致命的な忘れが発生します。この研究では、インコンテキスト検索問題としてインクリメンタル GR を再検討します。この問題では、新しく追加されたドキュメントが推論時のドキュメントドキュメント証拠として提供されます。私たちは、パラメトリック メモリとコンテキストが提供するドキュメントと docid のペアの両方でソースを認識した docid 生成を実行する、コンテキスト内インデックス作成フレームワークである ICICLE を提案します。 ICICLE は、「[COPY]」ベースのルーティング メカニズム、プリファレンス ベースの調整、および大規模なコンテキスト適応を組み合わせて、コンテキストに基づいた検索をパラメトリック検索から区別します。 MS MARCO と NQ320K での実験では、ICICLE がコーパス固有の再トレーニングを行わずに、見た文書の保持を維持しながら、新しく導入された文書の検索を向上させることが示されています。さらに、私たちの分析では、ハイショットの劣化は主に配線の失敗によって引き起こされることが示されており、ソース選択のキャリブレーションがインコンテキスト生成検索を拡張するための主要なボトルネックであることが強調されています。
原文 (English)
ICICLE: Expanding Retrieval with In-Context Documents
Generative retrieval (GR) maps queries directly to document identifiers (docids) using parametric knowledge, However, this design makes corpus expansion costly: adding new documents requires updating model parameters to encode new document-docid associations incurs repeated training and catastrophic forgetting of previously indexed documents. In this work, we revisit incremental GR as an in-context retrieval problem, where newly added documents are supplied as inference-time document-docid evidence. We propose ICICLE, an in-context indexing framework that performs source-aware docid generation over both parametric memory and context-provided document-docid pairs. ICICLE combines a `[COPY]`-based routing mechanism, preference-based calibration, and large context adaptation to distinguish context-grounded retrieval from parametric retrieval. Experiments on MS MARCO and NQ320K show that ICICLE improves retrieval of newly introduced documents while preserving seen-document retention without corpus-specific retraining. Our analysis further shows that high-shot degradation is mainly caused by routing failure, highlighting source-selection calibration as a key bottleneck for scaling in-context generative retrieval.
DELOS: 対比学習フレームワークを使用したケプラー測光における浅いトランジットの検出
我々は、ケプラー測光における浅いトランジットを検索するために設計された対比学習ベースのフレームワークである、cOntrastive Scoring (DELOS) を使用した位相折り畳み光曲線の検出を紹介します。 DELOS は、GPU で高速化された位相折りたたみ、最適化された位相ビニング、およびカスタム 1 次元畳み込みエンコーダーを組み合わせて、各折りたたみ光曲線に通過らしさスコアを割り当てます。これにより、事前に検出されたしきい値交差イベントに依存することなく、試行期間にわたるスコア ピリオドグラムが生成されます。軌道周期が 100 ~ 150 日の中長期信号に焦点を当て、DELOS は現実的な通過モデルとケプラーのようなノイズ特性で生成された 2,000 万の合成光曲線でトレーニングされ、合成検証セットで 99.3% の検証精度を達成しました。制御された注入回復実験では、DELOS は、低い信号対雑音比 (低 SNR) 領域で、精度と再現率を組み合わせたパフォーマンスを、ボックス フィッティング最小二乗法 (BLS) と比較して 15.5 パーセント、トランジット最小二乗法 (TLS) と比較して 11.25 パーセント向上させます。また、BLS と TLS と比較して、検索がそれぞれ約 3 ~ 5 倍と 74 ~ 80 倍高速化されます。選択されたケプラー検証サンプルに適用された DELOS は、テストされた周期範囲内の既知の浅い中周期から長周期のトランジット信号をすべて復元しました。これらの結果は、DELOS が低 SNR 通過探索のための効率的かつ高感度のフレームワークを提供し、ケプラー、K2、TESS、PLATO、および Earth 2.0 データにおける長周期地球型惑星の将来の探索に向けた実用的な一歩となることを示しています。したがって、この研究は方法論の開発と検証研究として意図されており、新たに特定された候補の詳細な天体物理学的検証は将来の研究に延期されます。
原文 (English)
DELOS: Contrastive Deep Learning for Low-SNR Blind Transit Searches in Kepler Photometry
We present DEtection in phase-folded Light curves with cOntrastive Scoring (DELOS), a deep-learning framework that uses contrastive scoring to perform blind searches for shallow transits in Kepler photometry. DELOS combines GPU-accelerated phase folding, optimized phase binning, and a custom one-dimensional convolutional encoder to assign a transit-likeness score to each folded light curve, thereby producing a score periodogram over trial periods without relying on pre-detected threshold-crossing events. Focusing on intermediate-to-long-period signals with orbital periods of 100-150 days, DELOS was trained on 20 million synthetic light curves generated with realistic transit models and Kepler-like noise properties, achieving a validation accuracy of 99.3% on the synthetic validation set. In controlled injection-recovery experiments, DELOS improves the combined precision-recall performance by 15.5% relative to Box-fitting Least Squares (BLS) and 11.25% relative to Transit Least Squares (TLS) in the low Signal-to-Noise Ratios (low-SNR) regime. It also accelerates the search by factors of approximately 3-5 and 74-80 compared with BLS and TLS, respectively. Applied to a selected Kepler validation sample, DELOS recovered all known shallow intermediate-to-long-period transit signals in the tested period range. These results demonstrate that DELOS provides an efficient and sensitive framework for low-SNR transit searches and represents a practical step toward future searches for longer-period terrestrial planets in Kepler, K2, TESS, PLATO, and Earth 2.0 data. Accordingly, this work is intended as a methodological development and validation study, with the detailed astrophysical validation of newly identified candidates deferred to future work.
Planning-aligned Token Compression for Long-Context Autonomous Driving
Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences t…
Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
Language models fine-tuned where the correct completion must outrank a near-synonym competitor often fail silently. The cross-entropy loss…
Sensory Restoration via Brain-Computer Interfaces: A Scoping Review
Brain-computer interfaces (BCIs) can restore sensory and motor function in individuals with severe neurological impairment, but the literat…
Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining…
Horizon-Uniform Sensitivity and Decay of Terminal Reward Perturbations in Discrete-Time Pontryagin Systems
We study local stationary solutions of finite-horizon discrete-time Pontryagin systems near a steady extremal. Suppose that the stationarit…
局所的な可塑性を備えたハイブリッド ANN-SNN パイプライン
この研究では、事前学習された人工ニューラル ネットワーク (ANN) の豊富な埋め込みを効果的に活用して、高性能スパイキング ニューラル ネットワーク (SNN) を可能にする、ハイブリッド ANN-SNN パイプラインを提案します。このアーキテクチャでは、事前トレーニング済みの EfficientNet エンコーダーと CoLaNET スパイキング分類器が結合されています。レートコーディングを介してエンコーダーのアクティベーションをスパイクトレインに変換し、エンドツーエンドの勾配伝播をバイパスして、生物学にヒントを得たローカルな学習ルールを使用して後続の SNN 分類器をトレーニングします。このアプローチは、64 クラスの ImageNet ベンチマークで 99.09% の精度を達成し、従来のディープ ネットワークと同等のパフォーマンスを実証しました。この研究は、強力な事前学習済みエンコーダーを下流のスパイキング ニューラル ネットワーク タスクに適応させるための、生物学的に妥当で効率的なフレームワークを提示します。
原文 (English)
Hybrid ANN-SNN Pipeline with Local Plasticity
This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs) to enable high-performance spiking neural networks (SNNs). The architecture couples a pretrained EfficientNet encoder with a CoLaNET spiking classifier. We convert the encoder's activations into spike trains via rate-coding and train the subsequent SNN classifier using local, biologically inspired learning rules, bypassing end-to-end gradient propagation. This approach achieves 99.09% accuracy on a 64-class ImageNet benchmark, demonstrating performance on par with conventional deep networks. The work presents a biologically plausible and efficient framework for adapting powerful pretrained encoders to downstream spiking neural network tasks.
First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers
Why do multilingual language models sometimes generate in the wrong language, and why is this so hard to fix? We introduce Language Identit…
Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interactio…
カスケード特徴除去による階層分類: ヒト表現型オントロジーに合わせた顔表現型解析 (FaceMesh2HPO) への応用
FaceMesh2HPO は、臨床診断をサポートするために、ヒト表現型オントロジー (HPO) に合わせて顔の表現型記述子を分類するためのフレームワークです。 10 の疾患にわたる 124 人の臨床医からのアノテーション (107 HPO 用語) と非症候性コントロールを組み合わせて、2D 画像から 3D 顔メッシュ (478 個のランドマーク) を生成し、カスケード分類と特徴除去を使用して階層的な PointNet ベースのパイプラインをトレーニングしました。 3D メッシュ、顔の輪郭、人口統計メタデータを組み込んだ最良のモデルは、~0.55 ~ ~0.89 の AUROC を達成し、親ノードでのパフォーマンスがリーフ項よりも高かった。外部検証により、障害全体にわたってさまざまな一般化可能性が示されました。結果は、3D 顔形状の階層モデリングにより、解釈可能なオントロジーにリンクされた表現型分類が可能になることを示していますが、まれな葉の用語でのパフォーマンスは依然として限定的です。堅牢性と臨床的有用性を高めるには、データの多様性と特徴選択戦略の改善が必要です。
原文 (English)
Hierarchical Classification via Cascading Feature Elimination: Application to Human Phenotype Ontology-Aligned Facial Phenotyping (FaceMesh2HPO)
FaceMesh2HPO is a framework for classifying facial phenotypic descriptors aligned with the Human Phenotype Ontology (HPO) to support clinical diagnosis. Using annotations from 124 clinicians across 10 disorders (107 HPO terms) combined with non-syndromic controls, we generated 3D facial meshes (478 landmarks) from 2D images and trained a hierarchical PointNet-based pipeline with cascading classification and feature elimination. The best models, incorporating 3D meshes, facial outline, and demographic metadata, achieved AUROCs between ~0.55 and ~0.89, with higher performance at parent nodes than leaf terms. External validation showed variable generalizability across disorders. Results demonstrate that hierarchical modeling of 3D facial geometry enables interpretable, ontology-linked phenotype classification, though performance on rare leaf terms remains limited. Improved data diversity and feature selection strategies are needed to enhance robustness and clinical utility.
クロスリンガル手書き OCR のための LLM 駆動の AutoML: GPT-5、GPT-4o、および Claude Sonnet 4 を使用した閉ループ ニューラル アーキテクチャ検索
我々は、GPT-5、GPT-4o、および Claude Sonnet 4 を、言語を超えた手書きの光学文字認識のための自律ニューラル アーキテクチャ設計者として使用する、完全に自動化された閉ループ AutoML フレームワークを紹介します。各大規模な言語モデルは、以前のトライアルからのパフォーマンス フィードバックを使用して、ニューラル ネットワーク アーキテクチャを個別に生成、トレーニング、評価し、繰り返し改良します。このフレームワークは、270 の独立した実験を通じて、アラビア語、ペルシア語、英語の手書きデータセットで評価されます。手動によるアーキテクチャ設計、ドメイン固有の前処理、ハイパーパラメータ調整を必要とせずに、正確で計算効率の高いモデルを一貫して検出します。生成されたモデルは、93 パーセントを超える平均テスト精度、98.1 パーセントの最高精度、および 41 ~ 44 ミリ秒の推論遅延を達成しています。この結果は、大規模な言語モデルがニューラル アーキテクチャ検索用の効果的な AutoML エージェントとして機能し、言語間でスケーラブルでスクリプト適応性があり、再現可能な手書き認識を可能にすることを示しています。
原文 (English)
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture designers for cross-lingual handwritten optical character recognition. Each large language model independently generates, trains, evaluates, and iteratively refines neural network architectures using performance feedback from previous trials. The framework is evaluated on Arabic, Persian, and English handwriting datasets through 270 independent experiments. It consistently discovers accurate and computationally efficient models without manual architecture design, domain-specific preprocessing, or hyperparameter tuning. The generated models achieve mean test accuracies above 93 percent, a best accuracy of 98.1 percent, and inference latency between 41 and 44 milliseconds. The results demonstrate that large language models can function as effective AutoML agents for neural architecture search, enabling scalable, script-adaptive, and reproducible handwriting recognition across languages.
RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce
Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversi…
多変量時系列予測のためのマルチスケール時間パッチ上の構造化された潜在空間モデリング
多変量時系列は、複数の時間スケールにわたって展開する構造パターンをコード化しますが、ほとんどの予測バックボーンは、学習された表現を予測の一時的な副産物として扱い、これらのパターンの組織幾何学は十分に活用されていません。チャネルに依存しない多変量観測を 2 つの相補的な微分可能な制約を通じて構造化された潜在空間にマッピングする CNN ベースの予測アーキテクチャである M2Patch を紹介します。マルチスケール パッチングは、入力を重複する時間粒度に分解します。段階的な拡張を伴う深さ方向の分離可能な畳み込みは、線形時間でスケール固有の特徴を抽出します。そして、スケールごとに学習された投影は、これらの特徴をコンパクトな潜在表現に圧縮します。潜在空間は、隣接するパッチ間の時間的連続性を強制するスケール内平滑性制約と、学習可能なクロススケール マッピングを通じて実現されるスケール間アライメント制約によって編成され、チャネル独立設計内で粒度間の相互作用を復元し、すべてのスケールが基礎となるダイナミクスの相互に一貫した表現を確実にエンコードします。 10 の実際のベンチマークでの実験では、M2Patch が 40 の予測設定全体で 57 件の最良の結果と 34 件の次善の結果を達成し、線形の計算複雑さとパッチレベルの入力破損に対する堅牢性を維持しながら、ほとんどのベンチマークの代表的なベースラインと一致またはそれを上回っていることが示されています。
原文 (English)
Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting
Existing patching and multi-scale methods advance multivariate time series forecasting but treat learned representations as transient byproducts of prediction, lacking explicit mechanisms that enforce structural consistency across temporal scales. We propose M2Patch, a CNN-based architecture that organizes channel-independent observations into a structured latent space via two complementary differentiable penalties. Multi-scale patching decomposes the input into overlapping temporal granularities, depthwise separable CNN blocks with progressively growing dilation extracts scale-specific features at linear complexity, and per-scale learned projections compress these features into a compact latent representation. An intra-scale smoothness penalty enforces temporal continuity between adjacent patches, while an inter-scale alignment penalty restores cross-granularity interaction through learnable cross-scale mappings, so that all scales encode mutually consistent representations of the underlying dynamics. Extensive experiments on ten real-world benchmark datasets demonstrate that M2Patch significantly outperforms state-of-the-art baselines. Further analyses establish M2Patch as a structure-aware recognizer: it recovers channel functional groupings and remains robust under patch-level input corruption, confirming that the structured latent space captures the data's intrinsic dynamics.
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distribu…
Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
Synthetic tabular data are valued for preserving not just column-wise marginals but inter-column dependency. Yet the most commonly reported…
EEGの基礎モデルは長距離の時間相関を認識できない:集団間の脆弱性の背後にあるスペクトルと時間の解離
客観的。脳波 (EEG) 基礎モデル (FM) は、短いパッチを再構築または対照的に位置合わせするようにトレーニングされ、固定埋め込みにプールされます。我々は、これらの埋め込みが、アルファバンドエンベロープのトレンド除去変動解析(DFA)指数によって定量化された長距離時間相関(LRTC)を保持しているかどうか、またそれが集団間の移動を支配しているかどうかをテストしました。アプローチ。私たちは、分布外の 2 つのコホートで生波形およびスペクトル入力アーキテクチャ (REVE、LaBraM、BENDR、CBraMod、BIOT) にわたる 5 つの EEG FM を調査し、DFA 指数の回復を静的な 1/f 非周期的傾きと比較しました。順序保持および残差化制御は、プーリングまたは非周期的シャドウイングについてテストされました。モンタージュ調和されたゼロショット転送タスクは、3 つのコホートにわたって凍結埋め込みと DFA 指数を比較しました (西洋の参照を追加)。主な結果。 5 つの FM はどれも、時間順に LRTC を表しませんでした。生波形モデル (REVE、LaBraM、BENDR) は、DFA 指数も 1/f 傾きも回復しませんでした (R^2 <= 0.12)。これら 3 つの場合、プローブは情報を提供しないため、解離はスペクトル入力モデル (CBraMod、BIOT) に特有であり、1/f は強く回復しますが (R^2 = 0.59-0.73)、コホート全体の DFA は回復しません。古典的な DFA 特徴により指数が回復され (信頼性上限 0.64 に対して R^2 = 0.32 ~ 0.38)、LRTC は非周期的な傾きに直交しました (r = -0.06)。集団間の転移では、凍結された REVE 埋め込みは偶然に勝てず (W から K、0.45)、無次元 DFA 指数は方向性を持って転移しましたが、家族ごとの有意性ではありませんでした。他の 4 つはそれを均一に再現しませんでした。 5 つすべてが記録サイト軸 (確率 0.500 に対して 0.98 ~ 1.00 で復号可能) によって支配されていましたが、破棄された DFA 指数はサイト堅牢性 (0.71) でした。
原文 (English)
Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations
Objective. We tested whether frozen representations from five EEG foundation models support decoding of long-range temporal correlations, measured as the detrended-fluctuation-analysis (DFA) exponent of the alpha-band amplitude envelope. Approach. REVE, LaBraM, BENDR, CBraMod, and BIOT were evaluated in CAUEEG and BrainLat. A common 240 s estimator used 8-13 Hz filtering, DFA over 2-23.8 s, artifact masking, and quality control. One fixed nested-cross-validation readout predicted DFA and a fixed-mode aperiodic exponent. Controls tested pre-pool order sensitivity and aperiodic residualization. Results. CAUEEG included 764 recordings and BrainLat 79. BIOT decoded DFA in CAUEEG (R-squared = 0.232; conditional subject-bootstrap 95 percent interval, 0.121-0.310), and CBraMod was positive but imprecise (R-squared = 0.121; 0.003-0.214). Neither replicated in BrainLat, where all five point estimates were negative. In contrast, CBraMod and BIOT decoded the aperiodic exponent in both cohorts (R-squared = 0.459-0.757). BIOT remained positive after removal of the measured linear aperiodic association in matched CAUEEG data (R-squared = 0.240). The post-hoc order control was batch- and configuration-sensitive. Because chronological EEG epochs are not exchangeable, it was descriptive, not an LRTC-specific test. No revised DFA transfer direction passed source-label permutation testing. Cohort membership was near-ceiling decodable from all five embeddings, but this is not a pure site effect. Significance. CBraMod and BIOT show a replicated, model-specific spectral-temporal dissociation: aperiodic decoding is present in both cohorts, whereas alpha-envelope DFA decoding is cohort-dependent. These findings bound the evaluated readouts; they do not establish representational absence or an architectural cause. Transfer and clinical associations remain exploratory.
Untrainable elements determine what physical learning remembers
Physical learning rules such as equilibrium propagation (EP), coupled learning (CL), and adjoint coupled learning (AL) train resistive netw…
The Epistemic Politics of AI Anthropomorphism
AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustain…
Approximate Speculative Decoding
Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greed…
Complete, Scalable, and Robust Prioritized Planning for Multi-Robot Ordered Storage and Retrieval at Maximum Capacity
Automated warehouses face a fundamental trade-off between maximizing storage density and achieving high retrieval throughput. While puzzle-…
Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol
AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after…
EgoCITE: 長期自己中心的メモリのためのコンテキスト拡張インデックス作成と時間認識検索
長期的な自己中心的な記憶は、連続する一人称のビデオとオーディオを、検索可能な過去の経験の記録に変換します。既存のシステムには 2 つのボトルネックがあることを示します。文脈に乏しいキャプションから構築されたインデックスはエージェント検索では信頼できません。また、検索では質問の一時的な意図が無視されます。両方のボトルネックに対処するために、自己中心的な QA のための長期的なエージェント メモリ フレームワークである EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval) を導入します。 EgoCITE は 3 つのコンポーネントで構成されます。 EgoScheme は、ローカルのマルチモーダル コンテキストを使用して、断片的なビデオ キャプションと音声トランスクリプトを自己完結型のアトミック メモリ インデックスに変換します。 EgoIndex は、相補的なアクション、アクティビティ、発話、および会話表現を、複数の粒度で検索可能なマルチビュー メモリ インデックスに編成します。 EgoRetrv は、セマンティック検索と、質問条件付きの時間的関連性スコアリングおよび取得された証拠のキュレーションを組み合わせたものです。 EgoLifeQA、EgoMem、および EgoR1-Bench 上の EgoCITE を、回答の精度とターゲットとイベントの検索の整合性の観点から評価します。 EgoCITE は、エージェント メモリ ベースラインの精度を少なくとも 4.4 ~ 14.2\% 向上させ、ロング コンテキスト LLM エージェントよりも 36$\times$ のコスト削減を実現します。
原文 (English)
EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory
Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstrate two bottlenecks in existing systems: indices built from context-poor captions are unreliable for agentic search, while retrieval ignores a question's temporal intent. To address both bottlenecks, we introduce EgoCITE (Egocentric Context-augmented Indexing and Time-aware Evidence retrieval), a long-horizon agentic memory framework for egocentric QA. EgoCITE comprises three components. EgoScheme uses local multimodal context to turn fragmentary video captions and speech transcripts into self-contained atomic memory indices. EgoIndex organizes complementary action, activity, utterance, and conversation representations into searchable multi-view memory indices at multiple granularities. EgoRetrv combines semantic search with question-conditioned temporal relevance scoring and curation of retrieved evidence. We evaluate EgoCITE on EgoLifeQA, EgoMem, and EgoR1-Bench in terms of answer accuracy and target-event retrieval alignment. EgoCITE improves accuracy over agentic memory baselines by at least 4.4--14.2% while achieving 36$\times$ lower cost than long-context LLM agents.
BrainWAM: 自動運転のためのセマンティック事前確率と予測ダイナミクスのアクション空間調整
自動運転には、意味論的な制約と予測ダイナミクスの両方に基づいた計画が必要です。しかし、既存のエンドツーエンドの運転アプローチは通常、この要件の片側のみを強調しています。つまり、視覚言語アクション (VLA) モデルは意味論的推論に VLM 事前分布を活用し、ワールド アクション モデル (WAM) は生成ワールド モデリングを通じて未来を意識した予測を提供します。これにより、意味論的な事前分布と予測ダイナミクスの両方を活用できる統合プランナーが自然と動機付けられます。しかし、共同トークンレベルの注意による単純な組み合わせは、意味論的なショートカットが共有の注意空間を支配し、予測のダイナミクスを抑制する、注意の割り当ての不一致に悩まされることがわかりました。機能的に特化されたシステム間の調整から複雑な動作が生じるという神経科学の証拠に触発されて、私たちは、意味論的推論と予測世界モデリングを2つの特化されたアクション指向の経路に変換し、コンパクトなアクション表現のレベルでそれらを調整する、構造化されたアクション空間調整フレームワークであるBrainWAMを提案します。さらに、ビデオとアクションのノイズ除去を分離した非同期整流フロー推論戦略を導入します。これにより、計画関連の予測コンテキストを維持しながら推論レイテンシが短縮されます。 BrainWAM は、NAVSIM v1 (89.5 PDMS) と NAVSIM v2 (89.6 EPDMS) の両方で最先端のパフォーマンスに達し、VLA のみまたは WAM のみの方法を常に上回っており、BrainWAM が自動運転システムの実用的で有望な方向性であることを強調しています。
原文 (English)
BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling. This naturally motivates a unified planner that can leverage both semantic priors and predictive dynamics. However, we find that a naive combination through joint token-level attention suffers from an attention-allocation mismatch, where semantic shortcuts dominate the shared attention space and suppress predictive dynamics. Inspired by neuroscience evidence that complex behavior arises from coordination among functionally specialized systems, we propose BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations. We further introduce an asynchronous rectified-flow inference strategy with decoupled video and action denoising, which shortens inference latency while preserving planning-relevant predictive context. BrainWAM reaches state-of-the-art performance on both NAVSIM v1 (89.5 PDMS) and NAVSIM v2 (89.6 EPDMS), consistently outperforming VLA-only or WAM-only methods, highlighting BrainWAM as a practical and promising direction for autonomous driving systems.
Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models
World models may predict the future without making clear which parts of their hidden state actually drive those predictions. We ask whether…
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments.…
Neurosymbolic Embodied Agents
Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate envi…
Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents
Large language model (LLM)-driven embodied agents rely on environment states to interpret scenes, generate high-level plans, and drive phys…
ターゲット側リーダー適応によるクロスモデルメモリ転送
大規模な言語モデルで知識の使用を改善する方法は、通常 2 つの体制に分類されます。ノンパラメトリック検索では、外部の知識に柔軟にアクセスできますが、検索の遅延、コンテキストのオーバーヘッドが追加され、バックボーンとの統合は浅いだけになります。パラメトリック適応は推論時には効率的ですが、知識とモデルの重みが絡み合うため、更新、監査、転送が困難になる場合があります。エングラム スタイルのハッシュ メモリは中間領域を占めます。つまり、学習した情報を外部のアドレス指定可能なテーブルに保存しますが、そのテーブルは小型の学習済みリーダーを通じて消費されます。これは基本的な疑問を引き起こします。そのようなメモリがバックボーンを越えて移動されるとき、凍結されたメモリ自体とターゲット側のリーダーのどちらがより重要なのでしょうか?私たちは、クロスモデルのフリーズメモリ抽出を通じてこの疑問を研究します。この抽出では、ソース モデルでトレーニングされたメモリがフリーズされ、別のターゲット モデルにアタッチされ、軽量リーダーのみがトレーニングされます。アブレーションは、学習されたメモリ内容と正しいアドレス指定の両方が重要であることを示していますが、転送されたテーブルは、ターゲット モデルに合わせて調整されたリーダーを通じてのみ有用になります。下流の質問応答タスクでは、デュアルレイヤー 4 ブランチ リーダーにより、同一モデルの再利用とクロスモデルの再利用との間のギャップがほぼ縮まり、管理された評価プロトコルの下で平均スコア 38.8 を達成しました。さらに、プロバイダー リーダーがターゲット インターフェイスと直接互換性がある場合、フリーズされたアーティファクトはターゲット側のトレーニングなしで実質的な有用性を提供でき、オプションのリーダー適応によりさらなる改善がもたらされます。これらの結果は、ターゲットが互換性のあるリーダー インターフェイスにアクセスできる場合、Engram が再利用可能な外部知識アーティファクトとして機能できることを示唆しています。リーダーの直接の再利用が不十分な場合、ターゲット側の適応によりアライメントをさらに改善できます。
原文 (English)
Cross-Model Memory Transfer via Target-Side Reader Adaptation
Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, in which a memory trained on a source model is frozen and attached to a different target model, with only a lightweight reader trained. Ablations show that learned memory content and correct addressing both matter, but the transferred table becomes useful only through a reader aligned to the target model. In downstream question answering tasks, a dual-layer, four-branch reader nearly closes the gap between same-model and cross-model reuse, achieving an average score of 38.8 under our controlled evaluation protocol. Moreover, when the provider reader is directly compatible with the target interface, the frozen artifact can provide substantial utility without target-side training, while optional reader adaptation yields further improvement. These results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface; target-side adaptation can further improve alignment when direct reader reuse is insufficient.
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its stro…
MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure
Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive…