AIニュース 2026-08-28
自動生成: 2026-08-28 19:47 JST
過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。
📌 今日の要点 TOP7
-
Supporting Thailand’s next generation of AI startupsOpenAI
OpenAI and Thailand’s MHESI launch an eight-week accelerator helping…
-
Anthropic、AIで物理機器を制御する共通規格「MHS」発表 将来オープンソース化へITmedia AI+
Anthropicは、AIエージェントが実験・製造機器を安全に操作するための共通仕様「Model Hardware Standard」(M…
-
Google、動画生成AI「Gemini Omni 1.1 Flash」公開 10秒分の文脈を参照して最長40秒まで延長ITmedia AI+
Googleは、動画の生成や編集に対応するマルチモーダルAI「Gemini Omni 1.1 Flash」を発表した。直前10秒の文脈を参…
-
Piloting the world's first double-blind AI evaluationsGoogle DeepMind
Piloting the world's first double-blind AI evaluations
-
Hugging Face、あひる型ロボット「Microduck」発表 399ドルで予約開始ITmedia AI+
Pollen Roboticsは、二足歩行ロボット「Microduck」を発表し予約受付を開始した。399ドルで25cmの小型設計。強化学…
-
Google、「Gemini Notebook」に購入済み電子書籍を追加できる「Expert Intelligence」ITmedia AI+
Googleは、信頼できる情報源をAI製品で活用する取り組み「Expert Intelligence」を発表した。第1弾として「Googl…
-
OpenAI、AIサイバー攻撃への「集団的対応」を訴える公開書簡 Anthropic、Google、Microsoftなど100以上の組織が署名ITmedia AI+
OpenAIは、AIを悪用したサイバー攻撃への備えを産業界と政府に訴える公開書簡を発表した。モデルの能力向上によりサイバー攻撃の高度化が懸…
トピック別件数
- LLM/生成AI 171件
- 研究/論文 114件
- エージェント 85件
- 画像/動画生成 49件
- ビジネス/資金調達 22件
- ロボティクス 15件
- ハードウェア/半導体 11件
- その他 5件
- 規制/政策 3件
日本語メディア15件
ITmedia AI+ (日本語)
チームみらい安野氏「高市総理へのAI家庭教師」を実施 総理が音声入力で作成したのは……
チームみらいの安野貴博党首らは、高市早苗総理大臣にAIについて指導するイベント「高市総理へのAI家庭教師」を首相官邸で実施した。
「国産ヒューマノイド」の作業風景を生配信中、3時間以上経過 「ベルトコンベヤーも自作」 アトム
アトムが人型ロボットの作業風景を生配信している。ベルトコンベヤーで流れてくる荷物を1つずつ手に取り、ラベルが貼られている面を上に向ける作業を実施中。
「半導体はAIが設計、人間は監督に」 次世代EDAへの取り組み
キーサイト・テクノロジー(以下、キーサイト)は2026年8月26日、エンジニア向けイベント「KDES」にあわせて、「AI時代の次世代半導体設計」をテーマとしたメディアブリーフィングを開催。AIを活用したEDAや光半導体設計、マルチフィジックス対応シミュレーションといった取り組み…
Hugging Face、あひる型ロボット「Microduck」発表 399ドルで予約開始
Pollen Roboticsは、二足歩行ロボット「Microduck」を発表し予約受付を開始した。399ドルで25cmの小型設計。強化学習による動作学習に特化し、SDKやシミュレーターをOSSとして公開する。大型機と異なり家庭や教室で安全に試行錯誤できる「行動するAIのための…
Anthropic、AIで物理機器を制御する共通規格「MHS」発表 将来オープンソース化へ
Anthropicは、AIエージェントが実験・製造機器を安全に操作するための共通仕様「Model Hardware Standard」(MHS)を発表した。機器の統合作業を大幅に短縮し、基本命令や安全制限を標準化して自律制御を可能にする。先行導入事例を公開し、安全性評価を進めた…
Google、「Gemini Notebook」に購入済み電子書籍を追加できる「Expert Intelligence」
Googleは、信頼できる情報源をAI製品で活用する取り組み「Expert Intelligence」を発表した。第1弾として「Google Playブックス」の対象電子書籍を「Gemini Notebook」のソースに追加可能にした。大手出版社の10万冊以上が対象で、今後は他…
“AI離れ”こそ重要? AIで成果を出す人の共通点 「差がつくスキルと思考法」を専門家に聞いた
AIで成果を出せる人と出せない人の違いは何か。差がつくスキルと思考法、その育て方をパーソル総合研究所の小林祐児氏に聞いた。
Google、動画生成AI「Gemini Omni 1.1 Flash」公開 10秒分の文脈を参照して最長40秒まで延長
Googleは、動画の生成や編集に対応するマルチモーダルAI「Gemini Omni 1.1 Flash」を発表した。直前10秒の文脈を参照した最大40秒のシーン延長や、開始・終了フレームの指定生成に対応。360pの高速プレビューから4Kアップスケールまでをカバーし、APIやツ…
ヒューマノイドを育てるバイト!? タイミーとジールスが現場データ収集で協業
ZEALSとタイミーは、ヒューマノイドなどフィジカルAI学習に必要な現場データを収集する、「データ収集スポットワーク」の創出に向けて共創パートナーシップを締結した。
「情シスはまだ実力の1割しか発揮していない」 AIリストラ時代のキャリアを考える
「クビにはならないが、専門性に合う仕事が消滅する」――。過去の技術転換期に起きたことは、AI時代にも起こり得る。ガートナーの亦賀忠明氏にAI時代にIT部門が果たすべき役割と、個人としてのキャリア構築の在り方を聞いた。
とにかく足りない上級IT人材 需給ギャップの背景を2つの調査から読み解く
上級IT人材を確保できない企業は多く、人材育成の壁もまた高い――ITRとガートナーの2つの調査からIT人材不足の全体像に迫る。
「全従業員がAIを使っているか」はもう成功の物差しじゃない AI投資、どう評価すべきか?
従来のITと異なり、AIは使えば使うほどコストが増えます。それでも活用の指標として利用者数や利用率を重視する企業はまだ多くあります。拡大を続けるAI投資の成果を、企業は何によって判断すればよいのでしょうか。新しい投資評価の視点を考察します。
OpenAI、AIサイバー攻撃への「集団的対応」を訴える公開書簡 Anthropic、Google、Microsoftなど100以上の組織が署名
OpenAIは、AIを悪用したサイバー攻撃への備えを産業界と政府に訴える公開書簡を発表した。モデルの能力向上によりサイバー攻撃の高度化が懸念される中、現状のセキュリティの不備を指摘し、防御用AIの普及と国際連携を要請。AnthropicやGoogleなど、100以上の企業や団体…
「新幹線料金と同じ」 ラピダス小池社長が明かす、2ナノ半導体“短納期”の勝算
ラピダスが計画する2ナノメートル級先端半導体の量産化について、小池淳義社長が見解を示した。旺盛なAI需要を背景に「工程は遅れなく進行している」と強調。絶対王者である台湾のTSMCに対する勝ち筋を「圧倒的な短納期」とし「新幹線料金のように高くても早く出す価値が評価される」と説明し…
OpenAIのAIエージェント、約700体の群れで企業を襲撃 ハッキングの痕跡隠蔽を試みる――事件の裏で何が?
OpenAIのAIエージェントが、企業をハッキングした事件の詳細が明らかになった。実に700体のエージェント群が連携していたという。事件の裏で何が起きていたのか。
海外メディア7件
TechCrunch AI (英語)
Anthropic and OpenAI are joining the AI stage at TechCrunch Disrupt 2026
At TechCrunch Disrupt 2026, the AI Stage is back to dig into the single hottest topic in the community for the past few years, presented by…
Barret Zoph, the Thinking Machines co-founder ousted before joining OpenAI, is now at Google
Zoph, who co-founded Thinking Machines Lab alongside Mira Murati and also served as the startup's CTO, led a brief stint at OpenAI and is n…
OpenAI, Anthropic, Google, and 100 other companies call for action to defend against rogue AI
Some of the world's largest tech companies and AI startups have come together to decry the current state of cybersecurity and to advertise…
Google’s AI Mode can now track flight prices, help book hotels, and more
The updates indicate that Google is looking to position AI Mode as an AI travel agent of sorts, as it's moving beyond simply helping users…
AI’s memory crunch is coming for Android apps
Google is setting new memory-use limits for Android apps as AI data centers contribute to hardware shortages that could leave lower-cost ph…
Plaud’s new earphones come with an eSIM-enabled case for talking to AI agents
Called Plaud One, these adopt the simple bare-bones style of Apple's AirPods, and can record calls, while their case can be used to record…
OpenAI to start showing ads on ChatGPT’s free and Go tiers in India
OpenAI has more than 100 million weekly active ChatGPT users in India, a huge chunk of whom are on the free or the lower-priced Go tiers.
公式ブログ2件
OpenAI (英語)
Supporting Thailand’s next generation of AI startups
OpenAI and Thailand’s MHESI launch an eight-week accelerator helping 10 health, wellness, and education startups turn AI prototypes into tr…
Google DeepMind (英語)
Piloting the world's first double-blind AI evaluations
Piloting the world's first double-blind AI evaluations
論文312件
arXiv cs.AI (英語)
EduRiskX: 早期学術リスク予測のための F-Logic 推論を備えた神経記号フレームワーク
オンライン教育における生徒の学業上のリスクを予測することは、定着率と学習成果を向上させるタイムリーな介入を可能にするために非常に重要です。しかし、既存のモデルは、早期検出能力が限られており、解釈可能性が不十分であることが多く、現実世界の教育現場での採用を妨げる「ブラックボックス」の信頼危機につながります。これらの課題に対処するために、時間的 Transformer ベースの予測子と F-Logic の記号推論を統合する神経記号フレームワークである EduRiskX を提案します。このニューラル コンポーネントは、時間的注意、クラス重み付け損失、および動的週次切り捨てを使用して、生徒の長期的な活動シーケンスをモデル化します。データ駆動型のエキスパート システムとして機能する F-Logic ルール ベースは、人間の教育者の診断ロジックを模倣する確立された教育理論 (エンゲージメント理論と生徒統合モデル) に基づいており、トレーニング データのみから構築されます。次に、神経リスク確率とシンボリック信頼スコアは、各信号の相対的な寄与を学習するロジスティック回帰ベースの融合メカニズムを通じて結合されます。厳格な 80/10/10 の学生レベル分割を使用した Open University Learning Analytics Dataset (OULAD) の実験では、EduRiskX が学期末 (第 38 週) の精度 0.900 と F1 スコア 0.894 を達成し、平均早期検出週 9.32 および検出率 94.30 パーセントを達成したことが示されています。最先端の時系列モデル (PatchTST、iTransformer) や一般的な深層学習ベースライン (LSTM、CNN) と比較して、EduRiskX は同一条件下で再現率が向上し、リスクを早期に特定します。 F-Logic モジュールは、予測パフォーマンスを超えて、予測を観察可能な行動パターンや教育理論に結び付ける、構造化されたルールベースの説明を提供します。
原文 (English)
EduRiskX: A Neuro-Symbolic Framework with F-Logic Reasoning for Early Academic Risk Prediction
Predicting students' academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes. However, existing models often suffer from limited early detection capability and insufficient interpretability, leading to a "black-box" trust crisis that hinders their adoption in real-world pedagogical settings. To address these challenges, we propose EduRiskX, a neuro-symbolic framework that integrates a temporal Transformer-based predictor with F-Logic symbolic reasoning. The neural component models longitudinal student activity sequences using temporal attention, class-weighted loss, and dynamic weekly truncation. Acting as a data-driven expert system, an F-Logic rule base -- grounded in established educational theories (Engagement Theory and Student Integration Model) to mimic the diagnostic logic of human educators -- is constructed exclusively from the training data. The neural risk probability and the symbolic confidence score are then combined through a logistic regression-based fusion mechanism that learns the relative contribution of each signal. Experiments on the Open University Learning Analytics Dataset (OULAD) using a strict 80/10/10 student-level split show that EduRiskX achieves an accuracy of 0.900 and an F1-score of 0.894 at the end of the semester (Week 38), with an average early detection week of 9.32 and a detection rate of 94.30 percent. Compared with state-of-the-art time-series models (PatchTST, iTransformer) and common deep learning baselines (LSTM, CNN), EduRiskX yields improved recall and earlier risk identification under identical conditions. Beyond predictive performance, the F-Logic module provides structured rule-based explanations linking predictions to observable behavioral patterns and educational theories.
Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset
Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative n…
Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap
Battery Prognostics and Health Management (BPHM) is critical for ensuring the safe, reliable, and cost-effective operation of batteries acr…
PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices
We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs)…
CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering
Calculation-intensive financial question answering requires exact reasoning over structured rates, temporal conditions, numerical formulas,…
The Artificial Experimentalist: Discovery and Control of Self-Organizing Phenomena with Autotelic Reinforcement Learning
Existing methods for exploring cellular automata and other complex systems mostly operate in open loop: they set initial conditions, execut…
The Accuracy-Efficiency Paradox Quantifying Net Energy Loss in on-Device Energy Forecasting
Energy forecasting aims to maximize accuracy to ensure energy efficiency by reducing energy waste, an objective that applies equally to on-…
LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs
Our research focuses on evaluating literature reviews generated in short and long context settings of large language models (LLMs) to inves…
Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse
Multi-table learning remains a major challenge in machine learning for healthcare and other complex information systems. Relational data co…
Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models
Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research pro…
Explainable Artificial Intelligence for Customer Churn Prediction in Telecommunications: A Framework for CRM Integration
Subscriber attrition is a costly, persistent challenge for telecommunications providers, with monthly churn of roughly 1.9% in mature marke…
EEG-to-Report: An Annotation and Feature-Text Framework for Training Language Models on Clinical EEG
Clinical electroencephalography (EEG) reporting remains largely manual and time-consuming, and current EEG software ecosystems do not produ…
Selection Bias Correction in Retail Intelligence
Retail intelligence often relies on monitoring popular, high-velocity products, potentially biasing economic indicators by ignoring the "lo…
GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions
Natural-language analytics over enterprise data warehouses is increasingly important, but production use is limited by hallucinated metrics…
SAREF-based Ontology for Distributed AI Workflows across the Edge-Fog-Cloud Continuum
Nowadays semantic models provide limited support for representing distributed AI workflows and their execution across heterogeneous edge, f…
A Safety-Gated Multimodal AI Backend for Mental-Health Support: Hierarchical State Representation, Conservative Risk Fusion, and Controlled Generation in Anian
Safety-critical mental-health support systems must distinguish when supportive conversation is appropriate from when free-form generation s…
A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving
Large language models can interpret natural-language chemistry questions, but their internal reasoning is difficult to inspect, constrain,…
Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript
Hallucination and abstention benchmarks rarely establish that a model could not have known the correct answer, making it difficult to disti…
Knowledge Cards: Structured Knowledge for AI Systems
AI systems whose outputs inform real decisions, and increasingly consequential ones, require something that current documentation practice…
AI Revealed Preferences
There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons. We test 20…
Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI
Large language models (LLMs) are increasingly used for verbal interaction in social robots, yet prompt design in human-robot interaction (H…
TutorTrace: A Dataset and Taxonomy for Classifying Learner Behavioral States during AI-Assisted Programming Education
AI programming tutors provide scalable support, yet lack the behavioral context human tutors rely on to adapt support to learners' needs. W…
Can You Say This for Me? Speaking Up by Proxy in Co-Located Discussion
Equal participation in co-located discussion is important for effective collaboration, yet people often hold back when they anticipate nega…
Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments
Large language models are increasingly used to inform safety decisions in cities, such as where it is safe to walk, rent, or travel. We ask…
Invocation-Level Reliability of Tool-Using Agents
Tool-using agents fail two ways: choosing the wrong tool, or forming wrong arguments, and an early failure of either kind can silently corr…
Predicting Consequences and Reinforcing Navigation Policies with Latent World Models
World models enable agents to reason about future outcomes and learn policies from their knowledge of state transition, but existing approa…
Structured Evidence Routing for Incident Risk Prediction from Multimodal Longitudinal EHRs
Incident risk prediction from longitudinal electronic health records (EHRs) is challenging because relevant signals are multimodal, weak in…
AffectOmni: 社会および芸術関連シーン向けの RL 検証可能な人間中心のグラウンデッド感情推論
マルチモーダル大規模言語モデル (MLLM) は、VQA とシーン理解において優れたパフォーマンスを達成しますが、感情的な推論は依然としてショートカット動作に対して脆弱です。モデルは、微細な表情やボディーランゲージなどの人間中心の手がかりを無視して正解を予測する可能性があり、追跡可能性や外部検証が弱まります。従来の強化学習アプローチでは、人間による証拠への注意を明示的に強制することなく、主にコンテキストや論理的一貫性を評価していました。さらに、ジャッジのスコアリングとしての LLM はスコアのクラスタリングに悩まされることが多く、これにより報酬の識別性が低下します。私たちは、検証可能な感情推論のための GRPO で訓練されたフレームワークである AffectOmni を提案します。 AffectOmni は、人間中心の証拠選択と時間的に構造化された推論を促進するために、People Focus と Temporal Order 報酬を導入し、より安定した差別的な報酬シグナルを生成するためにグループ内比較スコアリングを採用します。検証のために、Thinking Summarizer は自由形式の根拠を実行可能な証拠命令に変換します。これは、SAM3 を介してピクセル レベルの証拠領域に組み込まれ、トレーニング ループの外側で外部から監査可能なインターフェイスを提供します。 IntentBench、Daily Omni、WorldSense での実験では、感情認識で 4.66%、時間的に敏感なタスクで +14.29% の向上を含め、オープンソースの 7B スケールのベースラインと比べて一貫した改善が見られました。コードは https://github.com/eliot127825-rgb/AffectOmni_nobody で入手できます。
原文 (English)
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
Multimodal large language models (MLLMs) achieve strong performance on VQA and scene understanding, yet affective reasoning remains vulnerable to shortcut behavior. Models may predict correct answers while neglecting people-centric cues such as micro expressions and body language, which weakens traceability and external verification. Prior reinforcement learning approaches mainly reward context or logical coherence without explicitly enforcing attention to human evidence. In addition, LLM as a Judge scoring often suffers from score clustering, which reduces reward discriminability. We propose AffectOmni, a GRPO trained framework for verifiable affective reasoning. AffectOmni introduces People Focus and Temporal Order rewards to encourage people-centric evidence selection and temporally structured reasoning, and it adopts within-group comparative scoring to produce more stable and discriminative reward signals. For verification, a Thinking Summarizer converts free form rationales into executable evidence instructions, which are grounded into pixel level evidence regions via SAM3 to provide an externally auditable interface outside the training loop. Experiments on IntentBench, Daily Omni, and WorldSense show consistent improvements over open source 7B scale baselines, including gains of 4.66% on emotion recognition and +14.29% on temporally sensitive tasks. Code is available at https://github.com/eliot127825-rgb/AffectOmni_nobody.
ナノスケールの特性評価のための科学機器を操作するためのエージェント AI
原子間力顕微鏡 (AFM) などの科学機器を操作するには、専門家の継続的な意思決定が必要です。訓練を受けたユーザーは、実験の意図を定義し、それを機器コマンドに変換し、受信データを評価し、画像パラメータを調整し、最終画像を後処理します。既存の自動化は、通常、ハードコーディングされたルーチン、タスク固有のコントローラー、またはトレーニングされた機械学習モデルを通じて、このワークフローの一部のみに対応します。ここでは、モデル コンテキスト プロトコル (MCP) を介して機器機能に接続された汎用のツール拡張大規模言語モデルを使用して、AFM ワークフローの実行可能部分を操作するエージェント AI フレームワークを紹介します。このフレームワークは 3 つの MCP ベースのエージェントで構成されます。AFM Messenger は自然言語命令をチェックされた機器コマンドに変換します。 AFM Pilot は、大規模言語モデル (LLM) を通じて画質を評価し、必要に応じて画像パラメータを調整します。 AFM Doctor は画像アーティファクトを診断し、事前に承認されたツール セットから透過的な後処理を適用します。言語モデルは、固定スカラー目標や外部オプティマイザーではなく画像評価を実行するため、特別な再トレーニングを行わずに、同じ戦略をサンプル タイプやイメージング モード全体に適用できます。安全なハードウェア操作は、実行前に曖昧性チェック層を通じて強制されます。微調整された既製のツールを使用するモデルに対するベンチマークでは、モデルの機能だけではなく、この保護された実行層によって、誤ったコマンドの実行がゼロに減少することが示されています。さまざまなサンプルのライブ実験では、AFM Pilot は画質、反復回数、調整時間において専門オペレーターと同等であり、大きな違いはありませんでした。これらの結果は、コマンドの実行、画像ベースの調整、後処理が AI エージェントに委任される一方で、実験の意図は人間が定義したままである、科学機器のエージェント操作への安全なルートを示しています。
原文 (English)
Agentic AI for operating scientific instruments for nanoscale characterization
Operating a scientific instrument such as an atomic force microscope (AFM) requires continuous expert decision-making. A trained user defines the experimental intent, translates it into instrument commands, assesses incoming data, adjusts imaging parameters, and post-processes the final image. Existing automation usually addresses only parts of this workflow through hard-coded routines, task-specific controllers, or trained machine-learning models. Here we present an agentic-AI framework that operates the executable part of the AFM workflow using a general-purpose, tool-augmented large language model connected to instrument functions through the Model Context Protocol (MCP). The framework consists of 3 MCP-based agents: AFM Messenger converts natural-language instructions into checked instrument commands; AFM Pilot assesses image quality through a large language model (LLM) and, if necessary, adapts imaging parameters; and AFM Doctor diagnoses image artifacts and applies transparent post-processing from a pre-approved tool set. Because the language model performs image assessment rather than a fixed scalar objective or external optimizer, the same strategy can be applied across sample types and imaging modes without specific retraining. Safe hardware operation is enforced through an ambiguity check layer before execution. Benchmarking against fine-tuned and off-the-shelf tool-using models shows that this guarded execution layer, rather than model capability alone, reduces wrong-command execution to zero. In live experiments on different samples, AFM Pilot matched expert operators in image quality, iteration count, and tuning time, with no significant difference. These results demonstrate a safe route to agentic operation of scientific instruments, where experimental intent remains human-defined while command execution, image-based tuning, and post-processing are delegated to AI agents.
Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling
We ask whether AI agents powered by locally deployed large language models can reliably automate expert-defined hardware design workflows i…
GameWAM: ビデオ ゲームの世界的なアクション モデル
最新のビデオ ゲームは、一人称視点、急速な視覚変化、永続的な世界状態、および異種ネイティブ コントロールを組み合わせています。既存のゲーム エージェントは、ビジュアルおよびタスクのコンテキストをアクションに直接マッピングしますが、明示的なワールド ダイナミクス モデリングが欠けています。一方、インタラクティブなゲーム ワールド モデルは、提供されたアクションからビジュアルの将来を予測しますが、タスク ポリシーとしては機能しません。ワールド アクション モデル (WAM) はこれらの目的を統合しますが、ビデオ ゲームのダイナミクスと無制限の相互作用の下ではほとんど解明されていません。私たちの知る限りでは、ネイティブの閉ループ ゲームプレイと GUI 制御のための最初の WAM である GameWAM を紹介します。 GameWAM は、ブロック因果条件付けとフロー マッチングを使用した並列視覚生成プロセスとアクション生成プロセスを通じて、将来の視覚観察と実行可能なキーボードとマウスの軌跡を共同生成します。ワールドアクションの共同学習をサポートするために、同期したゲームプレイと GUI の軌跡を構築します。異種ネイティブ コントロールを処理するために、GameWAM は各アクション ステップでゲームプレイ/GUI モードを予測し、モード固有の予測分布と連続アクションの正規化を使用してアクションを生成します。長期的なインタラクションの場合、ブロックサイクル制御はコミットされたホライズンを超えて予測し、短いアクションプレフィックスのみを実行し、新しい観測に基づいて再計画しますが、きめ細かいサイクル内コンテキストと階層的なクロスサイクル履歴により時間的連続性が維持されます。実験では、比較されたエージェントよりも実行されたネイティブ アクションが少ないにもかかわらず、競合タスクの成功が実証されました。さらに、サンプリングされたアクション ソースの低周波成分が固定条件下で生成された粗いカメラの動きを系統的に制御する低周波アクション ソース インプリンティング (LASI) を明らかにし、生成制御におけるソース感度の故障モードを明らかにします。プロジェクトページは https://yunncheng.github.io/GameWAM/ から入手できます。
原文 (English)
GameWAM: A World Action Model for Video Games
Modern video games combine first-person perception, rapid visual changes, persistent world state, and heterogeneous native controls. Existing game agents map visual and task context directly to actions but lack explicit world dynamics modeling, whereas interactive game world models predict visual futures from supplied actions but do not serve as task policies. World-Action Models (WAMs) unify these objectives, but remain largely unexplored under the dynamics and open-ended interaction of video games. We introduce GameWAM, to our knowledge the first WAM for native closed-loop gameplay and GUI control. GameWAM jointly generates future visual observations and executable keyboard-mouse trajectories through parallel visual and action generative processes with block-causal conditioning and flow matching. To support joint world-action learning, we construct synchronized gameplay and GUI trajectories. To handle heterogeneous native control, GameWAM predicts a gameplay/GUI mode at each action step and generates actions with mode-specific prediction distributions and continuous-action normalization. For long-horizon interaction, block-cycle control predicts beyond the committed horizon, executes only a short action prefix, and replans from new observations, while fine-grained within-cycle context and hierarchical cross-cycle history preserve temporal continuity. Experiments demonstrate competitive task success with fewer executed native actions than the compared agents. We further uncover Low-Frequency Action Source Imprinting (LASI), in which low-frequency components of the sampled action source systematically steer coarse generated camera motion under fixed conditioning, revealing a source-sensitivity failure mode in generative control. Project page is available at https://yunncheng.github.io/GameWAM/.
同じモデル、異なるハーネス: 異なるコーディング エージェントの結果
コーディング エージェントはモデルとハーネスを組み合わせ、モデルが何を認識するか、どのツールを使用できるか、および作業の続行方法を決定します。モデルとタスクが固定されている場合に、ハーネスを変更すると結果が変わるかどうかを尋ねます。同じハーネスの 2 つの構成を 3 つのコーディング ベンチマークで比較します。コントロールは完全な会話を時間順に提供しますが、処理は同じ記録を保持しますが、コンテキストがいっぱいになり、繰り返しまたは停止した作業に応答するにつれて、古いツールの結果を機械的に短縮します。厳密な状況下では、この処理により 3 つの圧力比較すべてでタスクごとの平均不合格率 (F2PF) が向上し、SWE ベンチ Verified および SWE ベンチ Pro での完全なソリューションが増加します。タイトウィンドウの検証済み比較では、169 個のタスク、20,480 トークン ウィンドウ、および 480 秒の固定試行エンドポイントが使用されます。このコホートでは、治療によりタスクごとの平均 F2PF が 28 パーセントから 49 パーセントに上昇し、完全なソリューションが 43 から 72 に上昇しました。モデル固有の再調整がなければ、同じ凍結治療により、設計が異なる 3 つの追加モデルの同じコホートで両方のエンドポイントも上昇します。ワイドウィンドウの Qwen3.6 比較では、観察されたアームの結果は Verified と Pro でほぼ同じでしたが、FeatureBench は治療中のタスクあたりの F2PF の平均値が高く維持されています。ワイドウィンドウの検証済みコホートでは、治療によってターンごとに提供されるプロンプト トークンも少なくなります。ハーネスを変更すると、変更されていないモデルの重みで達成できる内容が変わるため、コーディング エージェントの評価では、モデルとハーネスをテスト対象のソルバーとして一緒に扱う必要があります。
原文 (English)
Same Model, Different Harness: Different Coding-Agent Results
A coding agent combines a model with a harness, which decides what the model sees, which tools it can use, and how the work continues. We ask whether changing the harness changes the result when the model and task stay fixed. We compare two configurations of the same harness on three coding benchmarks. The control supplies the full conversation in time order, while the treatment keeps the same record but mechanically shortens older tool results as the context fills and responds to repeated or stalled work. Under tight context, the treatment raises mean per-task fail-to-pass fraction (F2PF) in all three pressure comparisons and increases complete solutions on SWE-bench Verified and SWE-bench Pro. The tight-window Verified comparison uses 169 tasks, a 20,480-token window, and a fixed 480-second attempt endpoint; on this cohort, treatment raises mean per-task F2PF from 28 percent to 49 percent and complete solutions from 43 to 72. Without model-specific retuning, the same frozen treatment also raises both endpoints on the same cohort for three additional models with different designs. In the wide-window Qwen3.6 comparisons, observed arm outcomes are close on Verified and Pro, while FeatureBench retains a higher mean per-task F2PF under treatment. On the wide-window Verified cohort, treatment also serves fewer prompt tokens per turn. Because changing the harness changed what unchanged model weights could accomplish, coding-agent evaluations should treat the model and harness together as the tested solver.
エージェント メッシュ: 非冪等エージェント委任の信頼性プリミティブ - アイデンティティの適切性と証拠の適切性
自律エージェントは、再試行、再開、予算設定を行うオーケストレーターの下で、制限されたソフトウェア タスクを実行することが増えています。このようなオーケストレーターが利用する機構は、サービス メッシュの再試行、タイムアウト、エラー率サーキット ブレークです。私たちは、81 回の実行にわたる 147 件の番号付きインシデントにわたる実稼働エージェント ソフトウェア配信プラットフォームの障害調査を報告します。それぞれのインシデントには測定コストがあり、ほとんどの場合、障害を再現する変異の証拠が示されています。これらのプリミティブが基づいている 3 つの仮定はすべて実際に違反されており、その結果を定量化します。54 回連続して成功したツール呼び出しのループでは、エラー率ブレーカーは確認できませんでした。構造によって一定の進行信号があり、3 回目の修理ラウンドで誤ったトリップが保証され、1 回の実行が 6 つのコンポーネントのうち 6 つから 3 つに駆動されます。 1 つの委任の 6 回の呼び出しで 21 個のイベントが蓄積され、正しい冪等のコンポーネントを獲得できなくなりました。誤った経路で障害が発生し、2 コンポーネントの障害で 5 つのコンポーネントが起動し、3 人の傍観者が作業中のコードを退行させられました。施行層が正しい作業をブロックしたインシデントが 12 件あり、最も費用がかかったのはエージェントのターン数が 107 件で、書き込み受け付けがゼロでした。私たちは、1 つの横断的な原因とその二重の原因を発見しました。アイデンティティの妥当性: 5 つの別々のサブシステムにおいて、識別に失敗したアイデンティティは確信を持って間違った答えを生成し、そのうちの 2 つは独立して修正ルールを導き出しました。証拠の適切性: 信頼性の決定は、同じ条件下で移動可能で、測定対象に起因し、決定論的な証拠にのみ基づいて行われます。調査結果から、我々はメッセージではなく委任を実行単位とする 7 つの信頼性プリミティブを導き出し、研究が動機づけているが構成されていない管理された評価を特定します。
原文 (English)
Agent Mesh: Reliability Primitives for Non-Idempotent Agent Delegation - Identity Adequacy and Evidence Adequacy
Autonomous agents increasingly perform bounded software tasks under an orchestrator that retries, resumes, and budgets them. The machinery such orchestrators reach for is the service mesh's: retry, timeout, and error-rate circuit breaking. We report a failure study of a production agentic software-delivery platform over 147 numbered incidents spanning 81 runs, each with a measured cost and, in most cases, a mutation proof reproducing the failure. All three assumptions those primitives rest on are violated in practice, and we quantify the consequences: a loop of fifty-four consecutive successful tool calls no error-rate breaker could see; a progress signal constant by construction, guaranteeing a false trip on the third repair round and driving one run from six of six components to three; twenty-one events accumulated across six invocations of one delegation, making a correct, idempotent component unwinnable; a misrouted failure that woke five components for a two-component fault, leaving three bystanders regressing working code; and twelve incidents in which the enforcement layer blocked correct work, the most expensive costing 107 agent turns and zero accepted writes. We find one cross-cutting cause and its dual. Identity adequacy: in five separate subsystems an identity that failed to discriminate produced a confident wrong answer, and two of them derived the corrective rule independently. Evidence adequacy: a reliability decision may be taken only on evidence capable of moving, attributable to what it measures, and deterministic under identical conditions. From the findings we derive seven reliability primitives whose enforcement unit is the delegation rather than the message, and specify the controlled evaluation the study motivates but does not constitute.
時系列の LLM エージェント: 調査
LLM ベースのエージェントは、時系列問題向けに開発されることが増えていますが、その設計上の選択はタスク設定によって大幅に異なります。この調査では、これらのシステムを個別の技術コンポーネントごとではなく、対応する時系列の問題ごとに整理する問題主導型の分類法を採用しています。私たちは既存のシステムを、予測と推論、拡張と合成、異常の検出と診断、意思決定支援の 4 つのカテゴリに分類します。各カテゴリ内で、タスク要件がエージェント アーキテクチャ、ツールの使用、メモリ設計をどのように形作るかを調べます。さらに、代表的なデータセットと環境を要約し、共有設定または密接に関連した設定の下で報告されたモデルのパフォーマンスを比較します。全体として、この調査は、時系列問題に対する LLM ベースのエージェントを設計するためのタスク指向のガイドを提供し、今後の作業に向けて未解決のギャップを特定します。
原文 (English)
LLM Agents for Time-Series: A Survey
LLM-based agents are increasingly being developed for time-series problems, but their design choices vary substantially across task settings. This survey adopts a problem-driven taxonomy that organizes these systems by the time-series problems they address rather than by isolated technical components. We group existing systems into four categories: forecasting and reasoning, augmentation and synthesis, anomaly detection and diagnosis, and decision support. Within each category, we examine how task requirements shape agent architecture, tool use, and memory design. We further summarize representative datasets and environments, and compare reported model performance under shared or closely related settings. Overall, this survey offers a task-oriented guide to designing LLM-based agents for time-series problems and identifies open gaps for future work.
推論税: タスク タイプとデプロイメント コンテキストにわたる LLM 推論のトークン エコノミクス
推論可能な大規模言語モデルの精度のみのベンチマークでは、展開の中心となる質問、つまり拡張思考トークンはいつコストを獲得するのかが欠落しています。非推論ベースラインに対する推論モデルの精度向上を測定する限界ベンチマーク指標であるトークン エコノミー スコア (TES) を導入します。生成されたトークン乗数で正規化されます。推論トグルを備えたモデルファミリーと、直接の非推論対応物を持たないフロンティアモデルに対して、ペアと近似の TES バリアントを定義します。次に、数学、コード生成、科学的推論、命令追従、専門知識、知識の想起、研究レベルの物理学にわたる 7 つのベンチマークで 151 のモデル ベンチマーク評価実行にわたる実証的ベンチマーク分析を実行します。この分析では、どのタスク構造がプラスの限界推論効率を生み出すか、推論労力の増加によってモデル ファミリ内の TES がどのように変化するか、そして展開コンテキストが経済的実行可能性をどのように変化させるかという、展開に直面した 3 つの側面を調査します。結果は、タスク構造が名目上の難易度よりも推論効率をより良く予測することを示しています。AIME 2025 や LiveCodeBench などの逐次推論チェーン タスクは高い TES を示しますが、MMLU-Pro などの知識想起タスクは、難易度にもかかわらず低い TES を示します。また、追加の思考が精度を低下させる場合など、より高い推論努力レベルで体系的に利益が逓減することもわかりました。最後に、推論コスト シェア (RCS) は、推論支出が社内の考え方によって支配されることが多いことを示し、一方、デプロイメント コスト乗数 (DCM) は、オンプレミス展開によってコストのかかる推論ワークロードの経済性がどのように変化するかを示しています。これらの調査結果は、ベンチマーク主導のモデル選択ルールを裏付けています。つまり、普遍的に有益なモードとして扱うのではなく、タスクの種類、作業レベル、展開コンテキストによって選択的に推論を有効にするというものです。
原文 (English)
The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts
Accuracy-only benchmarking of reasoning-capable large language models misses a central deployment question: when do extended thinking tokens earn their cost? We introduce the Token Economy Score (TES), a marginal benchmarking metric that measures the accuracy gain of a reasoning model over a non-reasoning baseline, normalized by the generated-token multiplier. We define paired and approximated TES variants for model families with reasoning toggles and frontier models without direct non-reasoning counterparts. We then conduct an empirical benchmarking analysis across 151 model-benchmark evaluation runs on seven benchmarks spanning mathematics, code generation, science reasoning, instruction following, expert knowledge, knowledge recall, and research-level physics. The analysis examines three deployment-facing dimensions: which task structures yield positive marginal reasoning efficiency, how increasing reasoning effort changes TES within model families, and how deployment context changes economic viability. Results show that task structure predicts reasoning efficiency better than nominal difficulty: sequential inferencechain tasks such as AIME 2025 and LiveCodeBench show high TES, while knowledge-recall tasks such as MMLU-Pro show low TES despite their difficulty. We also find systematic diminishing returns at higher reasoning effort levels, including cases where additional thinking reduces accuracy. Finally, Reasoning Cost Share (RCS) shows that inference spend is often dominated by internal thinking, while Deployment Cost Multiplier (DCM) shows how on-premises deployment can change the economics of otherwise costly reasoning workloads. These findings support a benchmarking-driven model-selection rule: enable reasoning selectively by task type, effort level, and deployment context rather than treating it as a universally beneficial mode.
神経記号文学の 6.5% は、出版された成果物、6 段階の監査フレームワーク、および最初のインスタンス化から複製可能
私たちは、コンピューター サイエンス領域内の研究文献全体で科学的主張の再現性を監査するための 6 段階のフレームワークを提示し、神経記号 AI (NSAI) サブドメインのフレームワークをインスタンス化します。 NSAI サブドメインでフレームワークをインスタンス化すると、複数年にわたる監査が行われました。ステージ 1 では 5,497 件のレコードが取得され、3,018 件の重複が削除されました。第 2 段階では、タイトルと要約で 2,479 件の一意の記録をスクリーニングし、1,365 件の自己識別された NSAI 記録を特定しました。その後、主題から外れている、研究ではない、定量的評価がない、または全文にアクセスできないなどの理由で、さらに 61 件の全文を削除しました。ステージ 3 では、対象となる 1,304 件のレコードごとに検証可能なパブリック コード アーティファクトを探しましたが、849 件では何も見つかりませんでした。残り 455 件がアーティファクト インベントリに登録され、ステージ 4 と 5 が制限付きで再実行されます。 85 の研究、適格なコーパスの 6.52%、再実行の試みの 18.68% を完全または部分的に再現しました。再実行の試行 321 件は非コード アーティファクトの欠落によってブロックされ、42 件はコード リポジトリの欠落または使用不能によってブロックされたことがわかりました。これらの数字は、名目上の「コード利用可能」宣言でさえ存続する永続的な再現性の欠陥を定量化しており、将来の NSAI 出版物では、強制され、バージョン管理され、永久にアーカイブされたアーティファクト バンドルの必要性を示しています。私たちは、実証的な NSAI 論文は提出時に、完全でバージョン管理され、永続的にアーカイブされたアーティファクト バンドルを提供することを要求されるべきであると主張します。
原文 (English)
6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation
We present a six-stage framework for auditing the reproducibility of scientific claims across a research literature within the computer science domain, and instantiate our framework for the neuro-symbolic AI (NSAI) subdomain. Instantiating the framework on the NSAI subdomain produced a multi-year audit. Stage one retrieved 5,497 records and removed 3,018 duplicates. Stage two screened the 2,479 unique records at title and abstract, identifying 1,365 self-identified NSAI records, then removed a further 61 at full text for off-topic, non-research, no-quantitative-evaluation, or inaccessible-full-text reasons. Stage three sought a verifiable public code artifact for each of the 1,304 eligible records and found none for 849, leaving 455 to enter the artifact inventory and bounded rerun of stages four and five. We fully or partially reproduced 85 studies, 6.52% of the eligible corpus and 18.68% of attempted reruns. We found that 321 attempted reruns were blocked by missing non- code artifacts and 42 by missing or unusable code repositories. These figures quantify a persistent reproducibility deficit that survives even nominal "code available" declarations, and signal the need for enforced, versioned, and permanently archived artifact bundles in future NSAI publications. We argue that empirical NSAI papers should be required at submission time to provide complete, versioned, and permanently archived artifact bundles.
SKILL.state: スケーラブルな長期的なエージェント スキル
大規模言語モデル (LLM) は、複雑で長時間実行される手続き型スキルを実行する自律エージェントとして機能することが増えています。既存のエージェント ランタイムは、増加し続ける会話履歴に観察、アクション、および中間推論トレースを継続的に追加することで実行を維持し、長期にわたって遅延の低下やコンテキスト汚染の障害を引き起こします。追加のみの会話履歴を明示的で変更可能な実行状態に置き換えるランタイム アーキテクチャである SKILL.state を紹介します。各実行ステップで、モデルは不変のスキル仕様、現在の構造化された実行状態、および最新の観察のみを受け取ります。中間推論は、検証された状態の更新を生成した直後に破棄され、実行履歴による急激な増加を防ぎます。 SKILL.state は、多様なデータセット、モデル、実行環境にわたって、累積的なトークン消費を大幅に削減しながらタスクの精度を向上させます。私たちの結果は、明示的な実行状態が、スケーラブルな長期的なエージェント スキルにとって効果的でアーキテクチャに依存しない抽象化であることを示しています。
原文 (English)
SKILL.state: Scalable Long-Horizon Agent Skills
Large Language Models (LLMs) increasingly act as autonomous agents executing complex, long-running procedural skills. Existing agent runtimes maintain execution by continually appending observations, actions, and intermediate reasoning traces to an ever-growing conversation history, causing latency degradation and context-poisoning failures over long horizons. We present SKILL.state, a runtime architecture that replaces append-only conversational history with an explicit, mutable execution state. At each execution step, the model receives only the immutable skill specification, the current structured execution state, and the latest observation. Intermediate reasoning is discarded immediately after producing a validated state update, preventing prompt growth with execution history. Across diverse datasets, models, and execution environments, SKILL.state improves task accuracy while substantially reducing cumulative token consumption. Our results demonstrate that explicit execution state is an effective and architecture-agnostic abstraction for scalable long-horizon agent skills.
人間と大規模な言語モデルにおけるメンタライゼーションの評価
メンタライゼーション(他人の信念や意図を推測して自分の選択を導く能力)は、人間の社会的相互作用の基礎となる重要な認知機能です。大規模言語モデル (LLM) は、心の理論のタスクにおいて人間と一致する行動を示しますが、これらのモデルがメンタライゼーションを通じて適応行動を導くことができるかどうかは不明です。ここでは、認知計算モデリングを備えた 2 つの経済ゲームを使用して、LLM のメンタライゼーションの根底にある潜在的な戦略を明らかにします。私たちは、DeepSeek、GPT-4.1、GPT-5、Gemini 2.0 Flash の 4 つのモデル ファミリ (N = 2,099) にわたる個々の LLM エージェントを、さまざまな高度さの対戦相手に対してテストし、戦略的推論を引き出すように設計された刺激戦略がパフォーマンスを向上させるかどうかを調べました。比較尺度として、人間の参加者 (N = 251) に対して結果をベンチマークしました。両方のゲームにわたって、LLM は、モデルのプロバイダーと規模によって著しく異なるメンタライジングの行動的および計算的兆候を明確に示しました。戦略的プロンプトは、より洗練された推論を誘導することで一般にパフォーマンスを向上させましたが、利益の程度は 2 つのタスク間で異なりました。最後に、GPT-5 エージェントは、再帰的な推論の深さをますます洗練された敵に柔軟に適応させ、人間の参加者に対して優れたパフォーマンスを実証しました。集合的に、私たちは LLM にわたるメンタライゼーションのさまざまな能力を実証し、人間と機械にわたる比較知性を評価するための正式な方法として認知計算モデリングを強調します。
原文 (English)
Assessing mentalization in humans and large language models
Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.
承認が遅すぎた: LLM で保護された自己適応システムの陳腐化の評決
自己適応システム (SAS) 用の大規模言語モデル (LLM) ガードレールは、チェック時には正しい承認を発行する可能性がありますが、実行時には無効になる場合があります。これにより、実行ステージのチェック時間から使用時間まで (TOCTOU) ハザードが発生します。私たちは、ガードレールの判定が使用されたときに有効なままであるかどうか、つまり判定の鮮度を研究します。我々は、さまざまな質問に答える 3 つの数量を区別します。固定アクションのリプレイにおける全候補者の評決の変化、記録された閉ループの軌跡におけるオラクルラベル付きの承認期限、および裁判官が条件付けした使用時間の無効です。 5 つの再現可能な SAS 環境全体で、シミュレータ 8 ステップの一般的な再生シフトで、すべての候補の判定変更率は 5.3 ~ 48.4% の範囲にあります。明示的なプラントダイナミクス モデルを使用せずに、安全側マージンと最近の機能の変動性から各承認の有効期間を推定する鮮度境界シールド (FBS) を紹介します。 FBS は、アーティファクトに文書化された固定設定を使用して、同じシフトでオラクルラベルの承認期限切れ率を 3.4 ~ 24.7% から 0 ~ 1.8% に削減します。 4 人の LLM 裁判官に対する個別の監査により、すべての承認フローで裁判官条件付き使用時間の無効がゼロ以外であることが判明しました。当社は鮮度に関する契約を策定しています。すべての承認はチェック時に正しく、使用時にも有効でなければなりません。
原文 (English)
Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems
A large language model (LLM) guardrail for a self-adaptive system (SAS) may issue an approval that is correct at check time but stale by actuation. This creates an Execute-stage time-of-check to time-of-use (TOCTOU) hazard. We study verdict freshness: whether a guardrail verdict remains valid when used. We distinguish three quantities that answer different questions: all-candidate verdict change under fixed-action replay, oracle-labeled approval expiry on recorded closed-loop trajectories, and judge-conditioned use-time invalidity. Across five reproducible SAS environments, all-candidate verdict-change rates span 5.3-48.4% at a common replay shift of eight simulator steps. We introduce the Freshness-Bounded Shield (FBS), which estimates each approval's validity horizon from its safe-side margin and recent feature volatility, without an explicit plant-dynamics model. Using fixed settings documented in the artifact, FBS reduces oracle-labeled approval-expiry rates from 3.4-24.7% to 0-1.8% at the same shift. A separate audit of four LLM judges finds nonzero judge-conditioned use-time invalidity in every approval stream. We formulate a freshness contract: every approval must be correct at check time and remain valid at use time.
FaithSieve: 忠実な形式的証拠による数学的証明のきめ細かな評価
大規模な言語モデルは、複雑な複数ステップの数学的証明を生成できるようになりましたが、その正しさを確実に判断し、初期の論理エラーを特定することは依然として重要な課題です。既存の評価アプローチはモデルベースの自然言語判断に大きく依存しており、局所的な推論のギャップが見落とされることがよくあります。 Lean のような形式的定理証明者は厳密な検証への道を提供しますが、形式的定理証明者を非公式テキストの評価に使用するには、局所性と意味論的な不一致を解決する必要があります。証明者は、広すぎる対象を証明することで局所的な欠陥を回避したり、元の数学的意図から逸脱した自動的に形式化されたステートメントを検証したりする可能性があります。これに対処するために、自然言語の数学的証明をきめ細かく評価するためのリーン支援フレームワークである FaithSieve を紹介します。 FaithSieve は、大まかな証明ステップを局所的な推論単位に分解し、型指定された証明義務を抽出し、正式な評価エージェントを通じて検証します。形式的な検証は意味論的整合スコアリングによってゲートされるため、形式的なステートメントが元の主張の文脈、対象、および論理形式を忠実に保持している場合にのみ、無駄のない証拠が組み込まれます。私たちは、最初のエラーの位置特定をベンチマークするために、専門家によって検証された 2 つのデータセット、ProofLoc-Olympiad と ProofLoc-University を構築しました。 350 問題のオリンピック データセットでは、GPT-5.4 バックボーンを使用した FaithSieve は 81.43% の正確な一次エラー精度を達成し、直接判定ベースラインの 72.29% を上回りました。さらに、6 つの高度なドメインにまたがる 200 の問題の ProofLoc-University ベンチマークでは、FaithSieve の正確な精度は 84.5% に達しましたが、直接判定の精度は 75.0% でした。私たちの研究は、証明を細かい単位に分解し、忠実な形式的証拠に基づいて行うことで、自然言語推論の信頼できる評価が大幅に向上することを示しています。
原文 (English)
FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence
Large language models can now generate complex, multi-step mathematical proofs, but reliably determining their correctness and localizing early logical errors remains a critical challenge. Existing evaluation approaches largely depend on model-based natural-language judgments, which often overlook local reasoning gaps. While formal theorem provers like Lean offer a path to rigorous verification, using them to evaluate informal text requires solving locality and semantic mismatches: a prover might bypass a local flaw by proving an overly broad target, or validate an auto-formalized statement that drifts from the original mathematical intent. To address this, we introduce FaithSieve, a Lean-assisted framework for fine-grained evaluation of natural-language mathematical proofs. FaithSieve decomposes coarse proof steps into local reasoning units, extracts typed proof obligations, and verifies them through a formal evaluation agent. Formal validation is gated by semantic alignment scoring, so Lean evidence is incorporated only when the formal statement faithfully preserves the context, objects, and logical form of the original claim. We construct two expert-verified datasets, ProofLoc-Olympiad and ProofLoc-University, to benchmark first-error localization. On the 350-problem Olympiad dataset, FaithSieve using a GPT-5.4 backbone achieves 81.43% exact first-error accuracy, outperforming the direct-judging baseline of 72.29%. Furthermore, on the 200-problem ProofLoc-University benchmark spanning six advanced domains, FaithSieve reaches 84.5% exact accuracy, compared to 75.0% for the direct judge. Our work demonstrates that decomposing proofs into fine-grained units and grounding them with faithful formal evidence significantly improves reliable evaluation of natural-language reasoning.
ProofEvolve: 正式な自動定理証明のための神経記号進化
自動化された定理証明は、科学的発見における再帰的な自己改善のための自然な基盤を提供します。ただし、既存のニューラル証明器は、学習プロセスが時間の経過とともに自己改善する必要があるこの再帰的構造を完全には保存していません。既存の方法では、高価な重み更新を通じて証明経験をモデル パラメーターに埋め込むか、検証済みの中間演繹を現在の問題内にのみ保持します。さらに、これらの方法は、失敗した部分的な試みに有用な発見が含まれている場合でも、まばらな全体証明フィードバックに大きく依存しています。このギャップを埋めるために、私たちは、ニューラル モデルを使用して明示的で正式に検証された記号証明構造を進化させ、知識の境界を決定的に拡張する神経記号フレームワークである ProofEvolve を提案します。このフレームワークでは、ニューラル モデルは、分解、修復、スキーマの再結合などのバリエーション オペレーターを提案します。シンボリック リーン カーネルは、すべての証明遷移を検証します。 ProofEvolve は、進化ループ全体にわたって、結果として得られる証明有向非巡回グラフ (DAG) に対して検証済み閉包を計算します。各問題内で、ProofEvolve は動作インデックス付きアーカイブ内の部分的な AND-OR 証明 DAG を展開します。問題はありますが、カーネルチェックされたスキーマ抽出により、新しく証明されたサブ DAG が永続スキーマ ライブラリに追加されます。 Proof DAG は、型付きスキーマの再結合を通じて解決された結果を継承し、残りのすべての前提が新しいサブゴールとして公開されます。この進化のプロセスにより、不完全な試みから検証された結果が保存され、形式的な健全性を弱めることなく、後の証明に利用できるようになります。 3 つの競争レベルのリーン ベンチマーク全体で、ProofEvolve は、評価された証明システムの中で最高の平均解決率を達成しました。
原文 (English)
ProofEvolve: Neuro-Symbolic Evolution for Formal Automated Theorem Proving
Automated theorem proving offers a natural foundation for recursive self-improvement in scientific discovery. However, existing neural provers do not fully preserve this recursive structure, where the learning process should be self-improving over time. Existing methods either embed proof experience into model parameters through expensive weight updates, or keep verified intermediate deductions only within the current problem. In addition, these methods also heavily rely on sparse whole-proof feedback, even when unsuccessful partial attempts contain useful discoveries. To close the gap, we propose ProofEvolve, a neuro-symbolic framework that evolves explicit, formally verified symbolic proof structures with neural models to decisively expand the knowledge boundary. In this framework, the neural model proposes variation operators, including decompositions, repairs, and schema recombinations. The symbolic Lean kernel verifies every proof transition. Over the evolution loops, ProofEvolve computes verified closure over the resulting proof directed acyclic graphs (DAGs). Within each problem, ProofEvolve evolves partial AND-OR proof DAGs in a behaviorally indexed archive. Across problems, kernel-checked schema extraction adds newly proved sub-DAGs to a persistent schema library. Proof DAGs inherit the solved results through typed schema recombination, with every residual premise exposed as a new subgoal. This evolutionary process preserves verified results from incomplete attempts and makes them available for later proofs without weakening formal soundness. Across three competition-level Lean benchmarks, ProofEvolve achieves the highest average solve rate among the evaluated proof systems.
フレームを使用した変圧器モデルの微調整
低ランク適応 (LoRA) などのパラメーター効率の良い微調整 (PEFT) 戦略は、大規模な事前トレーニング済みモデルを微調整するための効果的なソリューションです。ただし、メモリ要件はモデル $\mathcal{O}(dr)$ のサイズに応じて変化します。ここで、$d$ はモデルの隠れ次元、$r$ はランクです。私たちの提案である FrameFT は、Fusion Frame ベースでスパース係数行列を使用してパラメーター更新 $\Delta W$ をモデル化します。 Fusion Frame はアルゴリズムによって生成され、モデル層全体で共有できるため、非常に効率的な更新が可能になります。基底展開のスパース係数のみが保存/最適化され、メモリ フットプリントが削減されます。 FrameFT の係数行列のスパース構造と Fusion Frames のスパース性により、計算に大きなメリットがもたらされ、私たちの分析では正式な収束結果が得られます。私たちは、言語タスクに焦点を当てて、一連の教師付き微調整ベンチマーク全体でアイデアを評価しますが、ビジョン モデルへの適用についても報告します。私たちの実験では、FrameFT が最先端の PEFT 技術と同等またはそれを超えるパフォーマンスを達成しながら、必要なトレーニング可能なパラメーターがはるかに少ないことが示されています。
原文 (English)
Fine-Tuning of Transformer models with Frames
Parameter-Efficient Fine-Tuning (PEFT) strategies such as Low-Rank Adaptation (LoRA) are effective solutions for fine-tuning large-scale pre-trained models; however, their memory requirements scale with the size of the model, $\mathcal{O}(dr)$, where $d$ is the model's hidden dimension and $r$ is the rank. Our proposal, FrameFT, models the parameter update $\Delta W$ with a sparse coefficient matrix in a Fusion Frame basis. Fusion Frames can be generated algorithmically and shared across model layers, enabling very efficient updates. Only the sparse coefficients of the basis expansion are stored/optimized, reducing the memory footprint. The sparse structure of the coefficient matrix in FrameFT and the sparsity in the Fusion Frames give large compute benefits, and our analysis provides formal convergence results. We evaluate the idea across a suite of supervised fine-tuning benchmarks, focusing on language tasks, but also report application to vision models. Our experiments show that FrameFT achieves performance on par with/exceeding state-of-the-art PEFT techniques, but needs far fewer trainable parameters.
考えすぎないでください、考えすぎないでください: エージェント AI における適応推論に向けて
大規模言語モデル (LLM) の最近の進歩により、推論時間の増加により、複雑なタスクのパフォーマンスが向上することがわかっています。ただし、既存のアプローチの多くは、固定トークン バジェット、事前実行難易度推定、アクティベーション スペース介入など、固定または事前に割り当てられた推論制御に依存しており、完全なエージェント ワークフローではなくスタンドアロンの推論ベンチマークで評価されることがよくあります。これらの仮定は、推論要件が計画、ツールの使用、メモリ検索、エージェント間の対話を通じて動的に進化するエージェント AI システムでは当てはまらない可能性があります。その結果、推論が過剰または不十分になり、不必要な計算、待ち時間の増加、計画のずれ、ツールの過剰な使用、または不完全なソリューションが発生する可能性があります。次世代エージェント AI の主要な課題は、単に言語モデルがどれだけの推論を実行する必要があるかではなく、進化するタスクの要求に応じて推論をどのように割り当てるべきかであると私たちは主張します。私たちは過剰推論と過少推論を、誤って割り当てられた推論の繰り返し起こる失敗モードとして特徴付け、MATH-500 と GAIA 公開検証ベンチマークで評価します。ツール決定の待ち時間、トークン消費量、トークン制限の枯渇、および回答の正しさを使用した結果は、過剰推論として分類されたケースは、比例した精度の向上がなく、より高い計算コストと関連しているのに対し、過少推論として分類されたケースは、一貫して不正確または不完全な解決策と関連していることを示唆しています。これらの発見は、エージェント AI の適応推論メカニズムに関する将来の研究の動機となります。
原文 (English)
Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI
Recent advances in Large Language Models (LLMs) have shown that increased inference-time reasoning can improve performance on complex tasks. However, many existing approaches rely on fixed or preallocated reasoning controls, such as fixed token budgets, pre-execution difficulty estimates, or activation-space interventions, and are often evaluated on standalone reasoning benchmarks rather than full agentic workflows. These assumptions may not hold in agentic AI systems, where reasoning requirements evolve dynamically through planning, tool use, memory retrieval, and agent-to-agent interactions. Consequently, reasoning can become either excessive or insufficient, resulting in unnecessary computation, increased latency, planning drift, excessive tool use, or incomplete solutions. We argue that a major challenge for next-generation agentic AI is not merely how much reasoning a language model should perform, but how it should allocate reasoning according to evolving task demands. We characterize over-reasoning and under-reasoning as recurring failure modes of misallocated reasoning and evaluate them on MATH-500 and the GAIA public validation benchmark. Using tool-decision latency, token consumption, token-limit exhaustion, and answer correctness, our results suggest that cases classified as over-reasoning are associated with higher computational cost without proportional accuracy gains, whereas cases classified as under-reasoning are consistently associated with incorrect or incomplete solutions. These findings motivate future research on adaptive reasoning mechanisms for agentic AI.
PILOT in the Loop: 長期的なエージェント向けの実践的な自己改善
長期にわたるエージェントの実行により、現在の実行と将来の作業の両方を改善できるエクスペリエンスが生成されます。ほとんどの自己改善メソッドは、実行終了後にのみこのエクスペリエンスを処理するため、アクティブな実行をリダイレクトしたり、そこから学んだ教訓をすぐに適用して検証したりすることはできません。私たちは、自己改善はライブで行うべきであり、アクティブな実行のリダイレクトと永続的なハーネスの更新の両方に新たなエクスペリエンスを使用する必要があると主張します。既存のエージェント アーキテクチャは、この目標を完全にはサポートしていません。単一エージェントの自己修正では、タスクの実行と軌道評価が 1 つのコンテキスト内で結合されますが、サブエージェントの委任では実行が分離されますが、通常はアクティブなサブエージェントをリダイレクトできません。我々は、2 つの結合されたメカニズムを通じてライブ自己改善のためのスーパーバイザーとワーカーのハーネスである PILOT を紹介します。(1) ライブ ステアリングにより、別のスーパーバイザーが実行中にアクティブなワーカーをリダイレクトまたは中止できます。 (2) ライブ自己進化は、実行中に明らかになった手順と失敗モードを再利用可能なスキルと記憶に抽出します。 2 つの凍結されたバックボーンと 3 つのベンチマークにわたって、PILOT は 6 つの構成のうち 5 つで 1 位にランクされています。 Terminal-Bench 2.0 では、PILOT は対応するハーネスよりも最大 9.8 パーセントポイント優れています。自己改善設定では、PILOT は GLM-5.1 で 14.6 ポイント、Kimi-K2.6 で 12.4 ポイントを獲得しました。平均出力トークンは 42.9% と 47.4% 減少しますが、100 万出力トークンあたりの成功した評価はそれぞれ 110.3% と 134.0% 増加します。
原文 (English)
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents
Long-horizon agent runs generate experience that can improve both the current run and future work. Most self-improvement methods process this experience only after execution ends, so they cannot redirect the active run or immediately apply and validate lessons learned from it. We argue that self-improvement should instead be live, using emerging experience both to redirect the active run and to update the persistent harness. Existing agent architectures do not fully support this goal. Single-agent self-correction combines task execution and trajectory assessment within one context, while subagent delegation separates execution but typically cannot redirect an active subagent. We present PILOT, a supervisor-worker harness for live self-improvement through two coupled mechanisms: (1) live steering lets a separate supervisor redirect or abort the active worker during execution; and (2) live self-evolution distils procedures and failure modes revealed during execution into reusable skills and memory. Across two frozen backbones and three benchmarks, PILOT ranks first in five of six configurations. On Terminal-Bench 2.0, PILOT outperforms counterpart harnesses by up to 9.8 percentage points. In the self-improvement setting, PILOT gains 14.6 points with GLM-5.1 and 12.4 points with Kimi-K2.6. Mean output tokens fall by 42.9% and 47.4%, while successful evaluations per million output tokens rise by 110.3% and 134.0%, respectively.
Multi2AV-Safety: マルチモーダルからオーディオビデオへの生成における安全性のベンチマーク
オーディオ/ビデオの生成は、プロンプト駆動の合成から、テキスト、画像、オーディオ、ビデオが共同して生成される出力を形成できるマルチモーダル コンディショニングへと急速に移行しています。この変化は安全性評価の性質を変えます。有害な意図はもはや単一の入力に存在するのではなく、代わりに、それ以外の点では良性または弱い有害な状態がモダリティや時間を超えてどのように相互作用するかによって現れる可能性があります。しかし、既存の安全性ベンチマークは依然としてプロンプト中心であるか、固定された調整インターフェースに結びついており、そのような構成リスクを体系的に研究することが困難なままになっています。このギャップを埋めるために、私たちの知る限り最初の安全ベンチマークである Multi2AV-Safety を導入し、11,024 の攻撃インスタンスを構成するオーディオビデオ生成用の 11 の非シングルトン T/I/A/V コンディショニング構成をすべてカバーします。 Multi2AV-Safety の評価により、攻撃メカニズムと被害証拠構造にわたる代表的なマルチモーダル安全対策の体系的な弱点が明らかになります。私たちの評価では、2 つの相補的な障害モードが明らかになりました。有害なセマンティクスは、個別に無害な入力の組み合わせから出現する可能性がある一方で、明示的な有害な手がかりは、無害なマルチモーダル コンテキストと混合されると検出が困難になる可能性があります。これらの結果を総合すると、\emph{構成的リスク認識} が、マルチモーダル条件付きオーディオビデオ生成を保護する際の中心的な能力ギャップであることが特定されます。現在の安全対策では、すべての条件付け入力が観察可能である場合でも、モダリティと時間にわたる安全性の証拠を確実に統合できません。データセットは 2026 年 10 月に一般公開される予定です。
原文 (English)
Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation
Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. This shift changes the nature of safety evaluation: harmful intent may no longer reside in any single input, but instead emerge from how otherwise benign or weakly harmful conditions interact across modalities and time. Existing safety benchmarks, however, remain largely prompt-centric or tied to fixed conditioning interfaces, leaving such compositional risks difficult to study systematically. To bridge this gap, we introduce Multi2AV-Safety, the first safety benchmark, to the best of our knowledge, to cover all 11 non-singleton T/I/A/V conditioning configurations for audio-video generation, comprising 11,024 attack instances. Evaluation on Multi2AV-Safety reveals systematic weaknesses in representative multimodal safety guards across attack mechanisms and harm-evidence structures. Our evaluation reveals two complementary failure modes: harmful semantics can emerge from the combination of individually benign inputs, while explicit harmful cues can become harder to detect when mixed with benign multimodal context. Together, these results identify \emph{compositional risk perception} as a central capability gap in safeguarding multimodal-conditioned audio-video generation: current safety guards fail to reliably integrate safety evidence across modalities and time, even when all conditioning inputs are observable. The dataset will be publicly released in October 2026.
DuMateBench: 現実世界の複雑なワークフローにおける自律エージェントの評価
自律エージェントは、現実世界の設定で複雑なマルチツールのワークフローを完了するために採用されることが増えています。ただし、既存のベンチマークは通常、タスクをアプリケーションまたは機能ごとに分離し、実際に遭遇する環境よりもクリーンで安定した環境でエージェントを評価します。 DuMateBench は、大規模な運用エージェント プラットフォームから収集された、匿名化されプライバシーが保護されたユーザー セッションから再構築されたリアル セッション ベンチマークです。各タスクは、関連するソリューション前の対話履歴、永続的な構成、およびワークスペースの状態を保存し、人間による検証を通じて検証されます。結果として得られたベンチマークは、8 つの広範なシナリオと 17 のきめ細かい機能カテゴリにわたる 200 のタスクで構成され、ほとんどのタスクで複数の機能の調整が必要になります。これらのタスクは、現実世界の環境の複雑性 (不十分、不安定、ノイズの 3 つの形態) が注入された分離された Docker コンテナ内で実行され、ハイブリッド決定論的評価プロトコルと LLM-as-Judge 評価プロトコルを使用してパフォーマンスを評価します。 5 つの代表的な自律エージェント フレームワークと 4 つの最先端の LLM を組み合わせた実験により、厳密なタスクの完了には大きなギャップがあることが明らかになりました。さらに、相補的な堅牢性、効率性、および診断分析により、環境変動下でのパフォーマンスは、LLM と周囲のエージェント フレームワークの機能によって共同で形成されることが示されています。コードとデータは https://dumatebench.com/ で公開されています。
原文 (English)
DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows
Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that are cleaner and more stable than those encountered in practice. We introduce DuMateBench, a real-session benchmark reconstructed from anonymized and privacy-screened user sessions collected from a large-scale production agent platform. Each task preserves the relevant pre-solution interaction history, persistent configurations, and workspace state, and is then validated through human verification. The resulting benchmark comprises 200 tasks spanning 8 broad scenarios and 17 fine-grained capability categories, with most tasks requiring multiple capability coordination. We execute these tasks in isolated Docker containers injected with three forms of real-world environmental complexity: Insufficient, Unstable, and Noisy, and assess performance using a hybrid deterministic and LLM-as-Judge evaluation protocol. Experiments across five representative autonomous-agent frameworks paired with four state-of-the-art LLMs reveal substantial gaps in strict task completion. Complementary robustness, efficiency, and diagnostic analyses further show that performance under environmental perturbations is jointly shaped by the capabilities of the LLM and the surrounding agent framework. The code and data are publicly available at https://dumatebench.com/.
AgentJudgeBench: エージェントのツール呼び出しで LLM ジャッジを評価するための複数の難易度のベンチマーク
LLM ジャッジはエージェント ツール呼び出しシステムを評価するために広く使用されていますが、構造化された依存関係主導のワークフローに対する LLM ジャッジの信頼性はほとんど検査されていないままです。我々は、ワークフロー DAG を介したエージェント ツール呼び出しに対する LLM としてのジャッジの信頼性を系統的に研究するための最初のベンチマークである AgentJudgeBench を紹介します。これは、自由形式のテキストやプリファレンス評価のより広範な LLM としてのジャッジ タスクとは異なります。このベンチマークは、6 つの DAG トポロジと 3 つの難易度層にまたがる 3,808 のインスタンスで構成され、グラウンド トゥルースありとなしのペア条件下で、5 つのジェネレーター (3B-70B オープンウェイト モデルおよび GPT-5.4) と 6 人のジャッジ (20B からフロンティア スケールまで) で評価されました。ジャッジのアライメントはタスクの難易度に応じて単調に低下し、グラウンド トゥルースを使用しない場合は 1.5 倍速くなります。また、グラウンド トゥルースを使用しないハード クエリでは、スケールに関係なく 6 人のジャッジすべてが狭い 77 ~ 82% の帯域に収束します。これにより、主にタスクの難易度によって駆動される構造的な天井が明らかになります。ただし、その高さは弱いジェネレーターのプロンプトに部分的に依存しており、モデルの容量だけでは克服できません。グラウンドトゥルースの露出は一律に有益ではありません。GPT-5.4 (1.5 pp) と Gemini-2.5-Pro (3.9 pp) のアライメントが低下し、オーバーアンカーと一致します。緩和戦略の中で、思考連鎖推論と裁判官の温度の影響はどちらも無視できる程度ですが、構造化された評価ルーブリックは調整を最大 6.5 pp 改善しますが、裁判官とジェネレータのペア全体で均一に一般化するわけではありません。グラウンド トゥルースでは、QwQ-32B がプログラムのリファレンスと最もよく一致しますが、人による検証研究では、GPT-OSS-120B が最も人間に合った判断者であると特定されています。それがなければ、フロンティアの裁判官は共有の上限内でわずかにリードするだけになります。これらの結果は、現在の LLM 審査員の根本的な限界を明らかにし、エージェント システムにおける信頼性の高い評価のための実践的なガイドラインをもたらします。
原文 (English)
AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling
LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation. The benchmark comprises 3,808 instances spanning six DAG topologies and three difficulty tiers, evaluated with five generators (3B-70B open-weight models and GPT-5.4) and six judges (20B to frontier scale) under paired with- and without-ground-truth conditions. Judge alignment degrades monotonically with task difficulty, 1.5x faster without ground truth, and on hard queries without ground truth all six judges converge to a narrow 77-82% band regardless of scale, revealing a structural ceiling driven primarily by task difficulty, though its height is partly prompt-dependent for weaker generators, that model capacity alone cannot overcome. Ground-truth exposure is not uniformly beneficial: it reduces alignment for GPT-5.4 (1.5 pp) and Gemini-2.5-Pro (3.9 pp), consistent with over-anchoring. Among mitigation strategies, chain-of-thought reasoning and judge temperature both have negligible effect, while structured evaluation rubrics improve alignment by up to 6.5 pp but do not generalize uniformly across judge-generator pairs. With ground truth, QwQ-32B best matches the programmatic reference, while a human validation study identifies GPT-OSS-120B as the most human-aligned judge; without it, frontier judges lead only marginally within the shared ceiling. These results expose fundamental limitations of current LLM judges and yield practical guidelines for reliable evaluation in agentic systems.
SIGMA: 構造化されたノイズ効果を考慮したグループ化されたマルチエージェントの集約
協調的なマルチエージェント強化学習 (MARL) は、ノイズの多い観測下で堅牢な調整を維持する上で重大な課題に直面しています。観測妨害はエージェント間で独立して導入されることがよくありますが、協力的な意思決定に対する下流の影響は、基礎となる協力構造を通じて構造化される可能性があります。我々は、この現象を構造化ノイズ効果として特徴付けます。この現象では、ノイズによって引き起こされる意思決定効果は、タスク関連の依存関係がより強いエージェント間で局所的な相関関係を示しますが、異なるエージェントや局所的な構造にわたって全体的には不均一のままです。しかし、既存の堅牢な MARL 手法では、そのような構造依存のノイズ効果を明示的に特徴付けたり利用したりすることはほとんどありません。この制限に対処するために、我々は、協力構造を利用してノイズの多い観測下でロバストな表現を学習する階層的コラボレーション フレームワークである SIGMA を提案します。 SIGMA はまず、密度ベースのグループ化を通じてエージェントを適応的なローカル構造に編成し、グループ内コンセンサス集約を実行して、エージェント固有の表現の逸脱を平滑化しながら共有タスク関連情報を保存します。次に、グループ間の注意により、異なるグループ間の情報が適応的に統合され、異質な貢献に対応しながらグローバルな調整が維持されます。 StarCraft II でのノイズの多い観測タスクの実験では、構造化ノイズの影響を経験的に検証し、SIGMA がノイズのない環境で競争力のあるパフォーマンスを維持しながら、観測ノイズの下でのロバスト性を一貫して向上させることを実証しました。
原文 (English)
SIGMA: Structured Noise-Effect-Aware Grouped Multi-Agent Aggregation
Cooperative multi-agent reinforcement learning (MARL) faces significant challenges in maintaining robust coordination under noisy observations. Although observation disturbances are often introduced independently across agents, their downstream effects on cooperative decision-making can become structured through underlying cooperation structures. We characterize this phenomenon as structured noise effects, where noise-induced decision effects exhibit local correlation among agents with stronger task-related dependencies while remaining globally heterogeneous across different agents and local structures. Existing robust MARL methods, however, rarely explicitly characterize or exploit such structure-dependent noise effects. To address this limitation, we propose SIGMA, a hierarchical collaboration framework that exploits cooperation structures to learn robust representations under noisy observations. SIGMA first organizes agents into adaptive local structures through density-based grouping and performs intra-group consensus aggregation to preserve shared task-relevant information while smoothing agent-specific representation deviations. Inter-group attention then adaptively integrates information across different groups to preserve global coordination while accommodating their heterogeneous contributions. Experiments on noisy-observation tasks in StarCraft II empirically validate the structured noise effects and demonstrate that SIGMA consistently improves robustness under observation noise while maintaining competitive performance in noise-free environments.
関係過剰正則化: 文遷移偏差によるグラフベースの AI 生成テキスト検出
AI 生成テキスト (AIGT) の検出は依然として困難です。既存のアプローチはトークン レベルの統計信号または独立したスタイロメトリーの特徴に依存しており、特定のジェネレーターにオーバーフィットし、分布シフトの下で失敗する原因となります。文ペアレベルで構造シグナルを特定します。LLM は、段落境界およびテンプレート化された遷移で繰り返される類似性バーストによって引き起こされる膨らんだ分散を通じて、人間の執筆から逸脱する文間の遷移の分散を生成します。これをリレーショナル過正則化 (ROR) として形式化し、4 つのベンチマークにわたって検証します (p < 0.001)。中心的な貢献は、新しい GNN アーキテクチャではなく、この関係問題の定式化です。 CSFG は、ROR を運用するための具体的なインスタンス化の 1 つです。この信号を利用するために、位置、逐次、意味、および遷移偏差信号を学習可能な GNN エッジ特徴としてエンコードするグラフベースのフレームワークであるクロスソース スタイロメトリック フィンガープリント グラフ (CSFG) を提案します。エッジごとの符号付き偏差 {\delta}_ij は、手作りのしきい値なしで ROR を動作させ、誤検知キャリブレーターとして機能します。 CSFG はバイナリ検出で 97.14% の精度を達成し、最も強力なグラフベースのベースラインを 11.14 pp 上回ります。偽陽性率は 1.57% であり、インフレート分散領域における目に見えない LLM に対する堅牢な汎化が行われます。遷移分散が人間のベースライン以下にあるジェネレータでは検出が低下します。
原文 (English)
Relational Over-Regularization: Graph-Based AI-Generated Text Detection via Sentence Transition Deviation
Detecting AI-generated text (AIGT) remains challenging because existing approaches rely on token-level statistical signals or independent stylometric features, causing them to overfit to specific generators and fail under distribution shift. We identify a structural signal at the sentence-pair level: LLMs produce inter-sentence transition variance that deviates from human writing through inflated variance driven by recurring similarity bursts at paragraph boundaries and templated transitions. We formalize this as Relational Over-Regularization (ROR) and validate it across four benchmarks (p < 0.001). The central contribution is this relational problem formulation, not a novel GNN architecture; CSFG is one concrete instantiation for operationalizing ROR. To exploit this signal, we propose the Cross-Source Stylometric Fingerprint Graph (CSFG), a graph-based framework that encodes positional, sequential, semantic, and transition deviation signals as learnable GNN edge features. The per-edge signed deviation {\delta}_ij operationalizes ROR without hand-crafted thresholds and acts as a false-positive calibrator. CSFG achieves 97.14% accuracy under binary detection, outperforming the strongest graph-based baseline by 11.14 pp, with a false-positive rate of 1.57% and robust generalization to unseen LLMs in the inflated-variance regime; detection degrades for generators whose transition variance falls at or below the human baseline.
実行時に自律型 AI エージェントを管理するための 5 つの基本要素
自律型 AI エージェントのエンタープライズ展開は、人間のユーザーと長期存続するサービス向けに構築された制御モデルを継承しますが、この適合は 3 つの特定の方法で失敗します。エージェント プリンシパルは一時的であり、プロビジョニングよりも早く出現および消滅します。彼らの行動はプログラムされたものではなくモデルによって選択されるため、彼らが試みる一連のことは事前にはわかりません。 API を呼び出すことができる人なら誰でも API を作成できるため、母集団はプロビジョニングされるのではなく検出されます。私たちは、このようなエージェントの管理は実行時の問題であり、モデルの調整の問題や構築時の問題ではないと主張し、アクションが有効になる前と有効になった後に答えなければならない質問から、発見、アイデンティティ、ガバナンス、認証、サプライチェーンという 5 つの基本要素を導き出します。それぞれについて、それがないと何が失敗するのか、そしてなぜ他のものが構造的にそれを提供できないのかを説明します。エージェントのアクションが発効する前にポリシーに対して調停され、テナントごとのアクションボキャブラリーに対して認可され、第三者がベンダーとアウトオブザループで検証できるハッシュリンクされた署名付き台帳に記録される実装について説明します。アーキテクチャのコストをレポートします。施行ポイントはリクエストのクリティカル パス上にあり、ID にはワークロードごとにサイドカーが必要で、フェールクローズされたメディエーションは可用性インシデントを拒否に変換します。実装ステータスについては明示的に示しています。4 つのプリミティブはプライベート パイロットで構築および実行されており、5 つ目は別個のツールとして構築されており、リクエスト パスにはまだ統合されていません。私たちはそれを意図的にセットに保持します。作成者がたまたま構築したものと正確に一致する 5 つの部分の分解は分類法ではなく、コードベースの説明です。
原文 (English)
Five Primitives for Governing Autonomous AI Agents at Runtime
Enterprise deployments of autonomous AI agents inherit a control model built for human users and long-lived services, and the fit fails in three specific ways: agent principals are ephemeral, appearing and vanishing faster than provisioning; their actions are selected by a model rather than programmed, so the set of things they may attempt is not known in advance; and the population is discovered rather than provisioned, because anyone who can call an API can create one. We argue that governing such agents is a runtime problem -- not a model-alignment problem and not a build-time problem -- and we derive five primitives from the questions that must be answered before an action takes effect and after it has: discovery, identity, governance, attestation, and supply chain. For each we state what fails if it is absent and why the others cannot structurally supply it. We describe an implementation in which an agent's action is mediated against policy before it takes effect, authorised against a per-tenant action vocabulary, and recorded in a hash-linked signed ledger a third party can verify with the vendor out of the loop. We report what the architecture costs: the enforcement point sits on the request's critical path, identity requires a sidecar per workload, and fail-closed mediation converts availability incidents into denial. We are explicit about implementation status: four primitives are built and running in private pilots, and the fifth is built as separate tooling and not yet integrated into the request path. We keep it in the set deliberately: a five-part decomposition that exactly matches what its authors happened to build is not a taxonomy but a description of a codebase.
現実世界でジェミニを使って科学研究を加速する
私たちは、仮説生成、実験、原稿生成にわたるエンドツーエンドの科学研究を加速するために設計されたジェミニベースのマルチエージェント システムである Co-Scientist の拡張と包括的な現実世界での検証を紹介します。この特殊な構成により、インシリコでの仮説生成を超えて、Co-Scientist は、材料科学、生物学、コンピューター サイエンスにわたる閉ループの科学ワークフローを推進する実行ベースの研究パートナーに移行します。材料科学では、共同研究者は半自動化学気相成長反応装置と連携して、MXene 用の安全な前駆体ルートを設計しました。実験の実行により、Ti3C2Tx MXene 格子と重要な構造的類似点を共有する層状 2D 材料が生成されましたが、原子構造を確認するにはさらなる実験が必要です。 Gemini 3 Deep Think を活用してラボインザループの迅速な実行を実現し、数分で成長レシピをラボの制約に合わせて調整し、単層 MoS2、MoSe2、および WS2 半導体の 1 回の試行成長を可能にしました。生物学では、共同研究者は、まばらな画像データから、インデューサー(IPTG)勾配を横切る人工大腸菌の出現した群集表現型を予測し、未発表のウェットラボ形態学的測定と定量的に一致しました。コンピューターサイエンスでは、共同科学者は、医師の盲検評価の下で潜在的な臨床上の危害を軽減しながら、HealthBench (ハードおよびプロフェッショナル) 上の 6 つのフロンティアモデルを上回る推論時間スケーリングアーキテクチャを自律的に発見しました。最後に、30 人のドメイン専門家による 450 件のレビューにわたるエンドツーエンドで生成された論文の二重盲検研究により、Co-Scientist の信頼性モジュールが研究の安全性を向上させながら幻覚や盗作を軽減することが実証されました。これらの結果を総合すると、現実世界の科学的発見を加速できるクローズドループのマルチエージェント科学 AI システムに向けた進歩が実証されています。
原文 (English)
Accelerating Scientific Research with Gemini in the Real-World
We present an extension and comprehensive real-world validation of Co-Scientist, a Gemini-based multi-agent system designed to accelerate end-to-end scientific research across hypothesis generation, experimentation, and manuscript generation. Moving beyond in silico hypothesis generation, this specialized configuration transitions Co-Scientist into an execution-grounded research partner advancing closed-loop scientific workflows across materials science, biology, and computer science. In materials science, Co-Scientist interfaced with a semi-automated chemical vapor deposition reactor to design a safe precursor route for MXenes; experimental execution produced a lamellar 2D material sharing key structural similarities with the Ti3C2Tx MXene lattice, although further experiments are needed to confirm the atomic structure. Leveraging Gemini 3 Deep Think for rapid, lab-in-the-loop execution, it also tailored growth recipes to laboratory constraints in minutes, enabling single-attempt growth of monolayer MoS2, MoSe2, and WS2 semiconductors. In biology, Co-Scientist predicted emergent swarming phenotypes of engineered E. coli across inducer (IPTG) gradients from sparse imaging data, quantitatively matching unpublished wet-lab morphological measurements. In computer science, Co-Scientist autonomously discovered an inference-time scaling architecture that outperformed six frontier models on HealthBench (Hard and Professional) while reducing potential clinical harm under blinded physician evaluation. Finally, a double-blind study of end-to-end generated papers with 30 domain experts across 450 reviews demonstrates that Co-Scientist's reliability modules reduce hallucination and plagiarism while improving research safety. Together, these results demonstrate progress toward closed-loop multi-agent scientific AI systems capable of accelerating real-world scientific discovery.
交絡因子としてのスタイル: AI による非ネイティブ学術論文の検出における誤検知
AI テキスト検出器は学術現場での採用が増えていますが、その出力が AI の著作者自身を反映しているのか、それとも洗練された学術英語に関連するより広範な言語的特徴を反映しているのかは不明のままです。これまでの研究では、英語を母国語としない文章の偽陽性率(FPR)が高いことが報告されているが、母集団レベルでの比較では、主題、分野、文体の違いにより著者が混同されている。専門的な編集は、著者と内容を維持しながら原稿の言語形式を変更するため、この問題を検討するのに便利な環境を提供します。私たちは、プロの英語編集サービス(2018~2025年)からの、非ネイティブ原稿とそのネイティブ編集版で構成される135,389組の文書を調査し、内容と著者を制御する検出器の応答に編集がどのような影響を与えるかを評価しました。 13 個の AI テキスト検出器では、人間が書いたテキストの FPR は 0.0% から 100.0% まで大きく異なりました。反応は検出器によって異なりました。同じ編集によって、一部の検出器では AI スコアが増加しましたが、他の検出器では AI スコアが減少しました。特に、スコアの変化は編集の程度と相関していました。この調査結果では、テキストの起源を言語スタイルから完全に分離するのではなく、AI検出器の出力における主要な交絡変数としてプロの編集スタイルが特定され、学術現場における公平性と信頼性についての懸念が生じています。
原文 (English)
Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing
AI text detectors are increasingly employed in academic settings, but it remains unclear whether their outputs reflect AI authorship itself or broader linguistic features associated with polished academic English. Previous studies have reported high false-positive rates (FPRs) for non-native English writing, but population-level comparisons confound authorship with differences in topic, domain, and writing style. Professional editing provides a useful setting for examining this issue because it changes the linguistic form of manuscripts while preserving authorship and content. We examined 135,389 document pairs from a professional English editing service (2018-2025), comprising non-native manuscripts and their native-edited versions, to assess how editing affects detector responses controlling for content and authorship. For the 13 AI text detectors, FPRs for human-written texts varied widely, from 0.0% to 100.0%. Responses varied across detectors: the same edits increased AI scores in some detectors but decreased them in others. Notably, score changes correlated with the extent of editing. The findings identify professional editing style as a key confounding variable in AI detector outputs, rather than establishing a full separation of text origin from linguistic style, raising concerns about fairness and reliability in academic settings.
再利用すべきでない場合を知る: トレーニング後の自律 LLM における条件付きエクスペリエンス転送
大規模な言語モデルは幅広い機能を提供しますが、進化するドメイン、ツール、要件に適応させるには、ポストトレーニングの繰り返しが必要になることがよくあります。自律システムは、更新を提案し、候補者をトレーニングし、評価フィードバックを使用して後続の提案を選択することにより、このプロセスの一部を自動化します。証拠が蓄積されるにつれて、中心的な問題が浮上します。それは、その後のトレーニングによって親モデルが変更された後でも、どの過去の更新証拠が実用的なものとなるでしょうか?更新の効果は、その親、データ、トレーニングの段階によって異なります。過去の成功をコンテキストフリーの許可として扱うと、無駄な計算が行われる可能性があります。その結果、子供が昇進した場合、その後のトレーニングの軌道が悪化する可能性もあります。我々は、この問題を条件付き経験転移として定式化し、重量変更トレーニングの前に経験の再利用を許可する方法である境界調整介入転移(BCIT)を導入します。 BCIT は、観察された効果をソース コンテキストに結び付け、適用条件をチェックし、名前付きのハード コンフリクトのある候補を拒否し、必要に応じて制限されたトレーニング トライアルを通じて現状の証拠を取得します。十分に訓練された候補者は依然として共有採用ルールに直面しており、観察された出来事のみが記憶を拡張します。財務推論、テキストから SQL への変換、および関数呼び出しにわたって適応された 1 つの 4B モデルでは、候補の更新は、評価されたコンテキスト全体で異種のターゲットと保持効果を示します。 BCIT は、候補、証拠、コンピューティングが一致している場合に、有害な更新をより少なく許可し、評価された代替案よりも同等の予算でより高い最終モデルの品質を達成します。これらの結果は、エクスペリエンスの認証を自律的なポストトレーニングにおける別個の問題として扱うことを裏付けています。
原文 (English)
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
Large language models offer broad capabilities, but adapting them to evolving domains, tools, and requirements often entails repeated post-training. Autonomous systems automate parts of this process by proposing updates, training candidates, and using evaluation feedback to select subsequent proposals. As evidence accumulates, a central problem emerges: which past update evidence remains actionable after subsequent training has changed the parent model? An update's effect depends on its parent, data, and training stage. Treating past success as context-free permission can waste compute. If the resulting child is promoted, it can also degrade the subsequent training trajectory. We formulate this problem as conditional experience transfer and introduce Boundary-Calibrated Intervention Transfer (BCIT), a method that authorizes experience reuse before weight-changing training. BCIT binds an observed effect to its source context, checks applicability conditions, vetoes candidates with named hard conflicts, and obtains current-state evidence through a bounded training trial when needed. Fully trained candidates still face a shared adoption rule, and only observed events extend memory. On one 4B model adapted across finance reasoning, text-to-SQL, and function calling, candidate updates exhibit heterogeneous target and retention effects across the evaluated contexts. Under matched candidates, evidence, and compute, BCIT authorizes fewer harmful updates and attains higher equal-budget final-model quality than the evaluated alternatives. These results support treating experience authorization as a distinct problem in autonomous post-training.
Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds
Enterprises fine-tune language models on proprietary data that may later require removal due to privacy, contractual, or compliance obligat…
AgentFold: タンパク質フォールディング モデル設計のための閉ループ エージェント検索
科学的 LLM エージェントは、文献推論、ツールの使用、実験計画において有望であることが示されていますが、実行可能コードの変更や計算コストのかかる検証を通じて、密結合した大規模な科学的機械学習システムを自律的に改善できるかどうかは依然として不明です。私たちは、タンパク質のフォールディングにおけるこの問題を研究します。この問題の進歩には、調整された構造の変更、多目的の評価、およびドメインを意識した解釈が必要です。我々は、実行可能コードのバリアントに対する閉ループ検索として折りたたみモデル開発を定式化するマルチエージェント フレームワークである AgentFold を紹介します。 ESMFold から開始して、AgentFold は仮説を提案し、コードレベルの変更を実装およびデバッグし、モデルのバリアントを評価し、実験結果を分析し、成功した介入と失敗した介入の両方を構造化メモリに保存します。 MCTS スタイルのポリシーは、スコアの高い検索ブランチ全体に計算リソースを割り当てます。 AgentFold は、2,000 行を超えるコードで構成されるエンジニアリング スケールのタンパク質フォールディング コードベースで、約 5,000 GPU 時間と 1 億 7,000 万 LLM トークンを使用して、約 80 のモデル バリアントを探索します。一致した計算予算の下で、AgentFold は、独立した Codex 提案よりも最良の lDDT を 7.5% 改善し、ランダム検索制御を上回ります。モデルの改善を超えて、結果として得られる介入トレースは、繰り返し発生する経験的な設計パターンを明らかにします。安定したゲインは、初期のソフトで学習可能な事前分布とゲート制御されたリファインメントから生じる傾向がありますが、直接的な幾何学的摂動や幾何学的条件付きフィードバックは、トレーニングを不安定にすることがよくあります。コードと実験リソースは https://github.com/lmqfly/AgentFold で公開されています。
原文 (English)
AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design
Scientific LLM agents have shown promise in literature reasoning, tool use, and experiment planning, but it remains unclear whether they can autonomously improve large, tightly coupled scientific machine-learning systems through executable code changes and computationally expensive validation. We study this question in protein folding, where progress requires coordinated architectural modifications, multi-objective evaluation, and domain-aware interpretation. We present AgentFold, a multi-agent framework that formulates folding-model development as a closed-loop search over executable code variants. Starting from ESMFold, AgentFold proposes hypotheses, implements and debugs code-level modifications, evaluates model variants, analyzes experimental outcomes, and stores both successful and failed interventions in structured memory. An MCTS-style policy allocates computational resources across high-scoring search branches. On an engineering-scale protein-folding codebase comprising more than 2,000 lines of code, AgentFold explores approximately 80 model variants using approximately 5,000 GPU-hours and 170 million LLM tokens. Under a matched computational budget, AgentFold improves the best lDDT by 7.5% over independent Codex proposals and outperforms a random-search control. Beyond model improvement, the resulting intervention traces reveal recurring empirical design patterns: stable gains tend to arise from early, soft, learnable priors and gated refinement, whereas direct geometric perturbations and geometry-conditioned feedback often destabilize training. The code and experimental resources are publicly available at https://github.com/lmqfly/AgentFold.
Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case
Data lakes rely on metadata to remain usable, yet this meta data is often limited or weakly informative for column relationship discovery,…
DEEPCHART: How Far are LLMs from Faithful Data-Science Chart Generation?
Faithful chart generation in real-world data-science workflows requires grounding visualizations in scattered evidence, computing chart-rea…
割引額ペイオフのためのカテゴライザー オートマトン
連続データを個別のビンに分類することは、人工知能の基本的な操作です。カテゴライザ オートマトンを導入します。これは、無限の報酬シーケンスを読み取り、有限の多くのビンのどれにその割引額が含まれているかを識別する決定論的なオートマトンです。カテゴライザー オートマトンは、定量合成で有用であることがすでに証明されている 2 つのビンの特殊なケースであるコンパレーター オートマトンを一般化します。私たちの主な技術的貢献は、状態空間がコンパレータ オートマトンの外積によって得られるような指数関数的ではなく、ビンの数において線形であるカテゴライザ オートマトンの構築です。次に、カテゴライザ オートマトンをマルコフ意思決定プロセスに適用します。これにより、不連続な可能性がある効用関数の割引額の期待効用を最大化するポリシーを合成できるようになります。区分的定数効用関数の場合、結果として得られるアルゴリズムは正確であり、擬似多項式時間で実行されます。区分的リプシッツ効用関数 (有限個のジャンプ間に勾配が限定された効用を含むクラス) の場合、これも擬似多項式時間で実行され、$\varepsilon$-optimal ポリシーが生成されます。また、検討した合成問題は区分定数ユーティリティにとってすでに PSPACE 困難であることも示します。
原文 (English)
Categorizer Automata for Discounted-Sum Payoffs
Categorizing continuous data into discrete bins is a fundamental operation in artificial intelligence. We introduce the categorizer automaton, a deterministic automaton that reads an infinite sequence of rewards and identifies which of finitely many bins contains its discounted sum. Categorizer automata generalize comparator automata, the special case of two bins, which have already proven useful in quantitative synthesis. Our main technical contribution is the construction of a categorizer automaton whose state space is linear in the number of bins, rather than exponential as obtained by a cross-product of comparator automata. We then apply categorizer automata to Markov decision processes, where they allow one to synthesize policies that maximize the expected utility of a discounted-sum payoff for utility functions that may be discontinuous. For piecewise-constant utility functions, the resulting algorithm is exact and runs in pseudo-polynomial time. For piecewise-Lipschitz utility functions, a class that includes any utility with bounded slope between finitely many jumps, it again runs in pseudo-polynomial time and yields an $\varepsilon$-optimal policy. We also show that the synthesis problem considered is PSPACE-hard already for piecewise-constant utilities.
AI Control Scientist: 自動制御設計のための LLM 駆動のエージェント システム
制御システムの設計は、化学プロセスの温度調整や航空エンジン制御など、現代の産業にとって重要です。しかし、従来の制御設計ワークフローは専門知識と広範な手動パラメーター調整に大きく依存しており、効率と拡張性が限られています。この目的を達成するために、この論文では、言語設計要件から最適化されたコントローラーを自動的に生成できる初の大規模言語モデル (LLM) 駆動型エージェントである AI Control Scientist (AICS) を提案します。具体的には、タスク モデリング エージェントはユーザー要件をエンジニアリング上の制約に合わせて解釈します。コントローラー設計エージェントは、候補コントローラー構造と実行可能コードを生成します。パラメータ調整エージェントは、閉ループのパフォーマンス基準に基づいてコントローラのパラメータを調整します。実験により、提案されたエージェント システムが複数の代表的な制御システムを自動的に生成でき、設計の成功率と最適化効率の両方で既存の自動ベースラインを上回ることが実証されました。この取り組みは、制御システム設計を人間主導からエージェント主導に変革し、モデル予測制御やその他の高度な制御システム設計への道を開く可能性を秘めています。
原文 (English)
AI Control Scientist: LLM-driven Agentic System for Automated Control Design
Control system design is critical for modern industry, such as chemical process temperature regulation and aero-engine control. However,traditional control design workflows rely heavily on expert knowledge and extensive manual parameter tuning, resulting in limited efficiency and scalability. To this end, this paper proposes AI Control Scientist (AICS), the first large language model (LLM)-driven agent capable of automatically generating optimized controller from language design requirements. Specifically, a Task Modeling Agent interprets user requirements to engineering constraints; a Controller Design Agent generate candidate controller structures and executable code; and a Parameter Tuning Agent refine controller parameters under closed-loop performance criteria. Experiments demonstrate that the proposed agentic system can automatically generate multiple representative control systems, outperforms existing automated baselines in both design success rate and optimization efficiency. This work has the potential to transform control system design from human-driven to agent-driven, paving the way for model predictive control and other advanced control systems design.
指示可能なエージェントの計画と制御の分離
最近の研究によると、事前にトレーニングされ、命令が調整されたビジョン言語モデル (VLM) は、命令と観察から高レベルの計画へのマッピングではうまく機能しますが、不慣れな環境では信頼性の高い低遅延のアクション シーケンスとしてそのような計画を実現するのに苦労することが示されています。同時に、ワールドモデルコントローラーは、観察から行動までの迅速な制御に優れていますが、オープンエンドのタスクガイダンスが欠けています。この研究では、これらの強みを単一のシステム Instruct-to-Act に結合し、VLM プランナーによって生成されたまばらでレイテンシの高い高レベルのテキスト命令を条件として、ワールド モデル コントローラーが高頻度で自律的に動作するようにトレーニングします。言語で指示できるようにコントローラーをトレーニングするために、コントローラー ポリシー ロールアウトのセグメントのラベルを合成命令で再ラベル付けし、既存の報酬最大化および世界モデリングの目標と合わせて動作の複製目標を共同で最適化します。私たちは、VLM プランナーが言語を通じて調整し、訓練されたコントローラーがアクチュエーターとして機能する 3 つのマルチエージェント環境を含む、7 つの具体化された環境にわたって提案されたアプローチを評価します。一致した観察空間とアクション空間の下で、当社の分離アプローチは、コントローラーのみおよび直接 VLM アクション生成バリアントよりも常に優れたパフォーマンスを発揮し、高速制御を維持し、微調整することなくさまざまな事前トレーニング済み VLM プランナーを交換できると同時に、7 つのタスクのうち 6 つで強力なビジョン-言語-アクションおよびマルチエージェント RL ベースラインとの競争力を維持します。
原文 (English)
Decoupling Planning and Control for Instructable Agents
Recent work shows that pre-trained, instruction-tuned vision-language models (VLMs) perform well at mapping from instructions and observations to high-level plans, but struggle to realize such plans as reliable low-latency action sequences in unfamiliar environments. At the same time, world-model controllers excel at fast observation-to-action control, but lack open-ended task guidance. In this work, we combine these strengths into a single system, Instruct-to-Act, where we train a world-model controller to act autonomously at high frequency when conditioned on sparse, higher-latency, and high-level text instructions generated by a VLM planner. To train controllers to be language-instructable, we relabel segments of controller policy rollouts with synthetic instructions and jointly optimize a behavior-cloning objective along with existing reward-maximizing and world-modeling objectives. We evaluate our proposed approach across seven embodied environments, including three multi-agent environments where VLM planners coordinate through language while trained controllers serve as their actuators. Under matched observation and action spaces, our decoupled approach consistently outperforms controller-only and direct VLM action-generation variants, preserves fast control, and lets us swap in different pretrained VLM planners without fine-tuning, while remaining competitive with strong vision-language-action and multi-agent RL baselines on six of seven tasks.
SymbolLKG: 論理ナレッジ グラフとシンボリック ソルバーによる検証可能な論理的推論に向けて
大規模言語モデル (LLM) は、自然言語理解において顕著な熟練を示していますが、厳密な多段階推論に苦労し、幻覚や矛盾に頻繁に悩まされます。思考連鎖 (CoT) などの既存のソリューションには厳密な検証メカニズムが欠けており、標準的な検索拡張生成 (RAG) では、論理タスクに固有の複雑な構造的な依存関係が見落とされることがよくあります。このギャップを埋めるために、論理ナレッジ グラフ (LKG) と動的ソルバー ルーティングを統合するニューロシンボリック アーキテクチャを提案します。具体的には、論理ルールと制約を第一級のトポロジー ノードとして扱うオントロジー ベースの LKG を導入し、テキストから抽出された依存関係の明示的なモデリングを可能にします。さらに、最適なシンボリック エンジンにタスクを動的にディスパッチするロジック ルーターを設計します。これは、トポロジを認識したハイブリッド検索メカニズムによってサポートされます。論理的推論ベンチマークの実験結果は、当社のフレームワークが最先端のプロンプトおよび RAG ベースラインを大幅に上回り、より高い精度と検証可能な推論パスを提供することを示しています。
原文 (English)
SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers
Large Language Models (LLMs) have demonstrated remarkable proficiency in natural language understanding, yet they struggle with strict multi-step reasoning, frequently suffering from hallucinations and inconsistency. Existing solutions like Chain-of-Thought (CoT) lack rigorous verification mechanisms, while standard Retrieval-Augmented Generation (RAG) often misses the complex, structural dependencies inherent in logical tasks. To bridge this gap, we propose a Neuro-Symbolic architecture that integrates a Logical Knowledge Graph (LKG) with dynamic solver routing. Specifically, we introduce an ontology-based LKG that treats logical rules and constraints as first-class topological nodes, enabling explicit modeling of dependencies extracted from text. We further design a Logic Router to dynamically dispatch tasks to the optimal symbolic engine, which is supported by a topology-aware hybrid retrieval mechanism. Experimental results on logical reasoning benchmarks demonstrate that our framework significantly outperforms state-of-the-art prompting and RAG baselines, delivering higher accuracy and verifiable reasoning paths.
LiveSim: マルチエージェント ライブストリーム エコシステムにおける環境に応じたユーザーのシミュレーション
大規模言語モデル (LLM) を使用したユーザー行動シミュレーションは、マルチエージェント エコシステム シミュレーションをサポートするために使用されることが増えています。既存のシミュレータは通常、過去の観察から推測される静的なユーザー プロファイルに依存していますが、インタラクションのダイナミクスがユーザーの行動を継続的に再形成するライブ ストリーミングなど、社会的に集約された環境では不十分になります。私たちは、ライブ ストリーム エコシステム シミュレーション用の LLM ベースのフレームワークである \textbf{LiveSim} を提案します。ユーザーを編集可能な行動仮説として表し、軌跡に基づいたインタラクションを通じてそれらを徐々に洗練させます。そこでは、シミュレーションされた軌跡と観察された軌跡の不一致により、欠けている環境形成効果が明らかになります。これらの信号は、転送可能な環境行動パターンとしてさらに抽出され、集合的な行動メモリに蓄積されることで、ユーザーレベルの行動忠実度が向上し、エコシステムレベルのシミュレーションがサポートされます。現実世界のライブストリームのリスク制御データの実験により、ユーザーレベルの行動忠実度を向上させ、リスクの進化とプラットフォーム介入効果のエコシステムレベルの分析を可能にする LiveSim の有効性が検証されます。
原文 (English)
LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems
User behavior simulation with large language models~(LLMs) is increasingly used to support multi-agent ecosystem simulation. Existing simulators typically rely on static user profiles inferred from historical observations, which become inadequate in socially intensive environments such as live streaming where interaction dynamics continuously reshape user behavior. We propose \textbf{LiveSim}, an LLM-based framework for live-stream ecosystem simulation. It represents users as editable behavioral hypotheses and progressively refines them through trajectory-grounded interactions, where discrepancies between simulated and observed trajectories reveal missing environmental shaping effects. These signals are further extracted as transferable environment-behavior patterns and accumulated in a collective behavioral memory to improve user-level behavioral fidelity and support ecosystem-level simulation. Experiments on real-world live-stream risk-control data validate the effectiveness of LiveSim in improving user-level behavioral fidelity and enabling ecosystem-level analysis of risk evolution and platform intervention effects.
BekchiAI: ワンクリックで LLM エージェントを測定、観察、制御
大規模な言語モデルのエージェントは、多くのステップにわたって推論し、ツールを呼び出し、自律的に行動しますが、そのエージェントのスキル (ツールの正しい順序付け、依存関係に基づいた計画、信頼できない入力の判断、生成された引数の根拠付け) は、精度のみのリーダーボードでは測定するのが困難です。私たちは、エージェントのスキルを測定するためのベンチマークと、ライブ エージェントを観察および制御するためのプラットフォームの両方に対処する BekchiAI を紹介します。 BekchiAI-Benchmark は、7 つのタスク カテゴリ (算術、構造化/SQL、セキュリティ検出、URL グラウンディング、プランニング、オーケストレーション、ツール ポリシー) にわたる 13 のツールを使用する ReAct エージェントのスイートであり、合計 2,057 の確定的でコミットされたテスト タスクがあります。すべてのタスクは検証者によってチェック可能であり、ゴールドアンサーは、実際のデータベースに対して正規 SQL を実行するか、有向非巡回グラフ (DAG) の正確なスケジュールを計算するか、意図的に不完全な署名スキャナーと組み合わせた敵対的なセキュリティ サンプルを含む閉形式ラムダを評価することによって計算されるため、スコアはオラクルのコピーではなくモデル自身の判断を反映します。精度ツール呼び出し遵守、URL 幻覚とソース一致、およびモデルごとのトークン コストを超えた一連の行動メトリクスを定義し、4 つのモデルの比較 (Qwen3.7-Max、gemma-4-31B-it、gemma4:26b、gpt-oss-120b) を報告します。そのストーリーは、集計ではなくファミリーごとのスプレッドに含まれます。ベンチマークの実行は、提供された評価スクリプトを使用して実行されます。 BekchiAI-Platform は、展開されたエージェントのための補完的な Web ベースの可観測性および制御レイヤーであり、完全なトークンとレイテンシのテレメトリ、およびリモート実行終了を提供します。ベンチマーク、評価ツール、プラットフォームは公開されています。
原文 (English)
BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click
Large language model agents reason, call tools, and act autonomously over many steps, but their agentic skills-correctly sequencing tools, planning under dependencies, judging untrusted inputs, and grounding generated arguments-are hard to measure with accuracy-only leaderboards. We present BekchiAI, which addresses both sides: a benchmark for measuring agentic skill and a platform for observing and controlling live agents. The BekchiAI-Benchmark, a suite of 13 tool-using ReAct agents across 7 task categories (arithmetic, structured/SQL, security detection, URL grounding, planning, orchestration, and tool-policy), totalling 2,057 deterministic, committed test tasks. Every task is verifier-checkable gold answers are computed by running canonical SQL against a real database, computing the exact schedule of a directed acyclic graph (DAG), or evaluating closed-form lambdas including adversarial security samples paired with deliberately imperfect signature scanners so a score reflects the model's own judgment, not the copying of an oracle. We define a small set of behavioral metrics beyond accuracy-tool-call adherence, URL hallucination and source-match, and per-model token cost and report a four-model comparison (Qwen3.7-Max, gemma-4-31B-it, gemma4:26b, gpt-oss-120b) whose story is in the per-family spread, not the aggregate. The benchmark runs are executed using the provided evaluation scripts. BekchiAI-Platform is a complementary web-based observability and control layer for deployed agents, providing full token and latency telemetry as well as remote run termination. The benchmark, evaluation tools, and platform are publicly released.
C-Unseen: LLM 推論による動的時間知識グラフにおける弱い信号の検出
弱いシグナルは、重大な変化が確立される前に、その変化に先立つ初期の目立たない指標です。キーワードの頻度、トピックのモデリング、または型なしのグラフ トポロジに基づく既存の検出方法では、そのような信号が現れる意味論的および関係的構造を捕捉できません。この論文では、Dynamic Temporal Knowledge Graphs (DTKG) における弱い信号検出のための自己解釈可能なフレームワークである C-Unseen を提案します。弱い信号を、連続する TKG スナップショットにわたって増殖する、まれで意味的に一貫したサブグラフとして定義します。このフレームワークは 2 つのモジュールを通じて動作します。LLM が思考連鎖推論を通じて、コンテンツが支配的なスナップショットの物語と緊張関係にあるサブグラフを識別するレア サブグラフ エクストラクターと、これらのまれなサブグラフの永続性をタイム ステップ全体で追跡して真の弱い信号を分離する弱いシグナル アラーターです。実験結果は、C-Unseen がキーワード、トピック、グラフベースのベースラインよりも優れていることを示しています。
原文 (English)
C-Unseen: Weak Signal Detection in Dynamic Temporal Knowledge Graphs via LLM Reasoning
Weak signals are early, low-visibility indicators that precede significant changes before those changes become established. Existing detection methods, based on keyword frequency, topic modeling, or untyped graph topology, fail to capture the semantic and relational structure through which such signals manifest. In this paper, we propose C-Unseen, a self-interpretable framework for weak signal detection in Dynamic Temporal Knowledge Graphs (DTKGs). We define a weak signal as a rare, semantically coherent subgraph that proliferates across consecutive TKG snapshots. The framework operates through two modules: a Rare Subgraphs Extractor, in which an LLM identifies subgraphs whose content is in tension with the dominant snapshot narrative via chain-of-thought reasoning, and a Weak Signal Alerter, in which the persistence of these rare subgraphs is tracked across time steps to isolate true weak signals. Experimental results demonstrate that C-Unseen outperforms keyword-, topic-, and graph-based baselines.
概念的に複雑なスコーピングレビューにおける人間とLLMのスクリーニングワークフローの評価: ワークロードのトレードオフと実行間の一貫性を思い出してください。
背景。大規模言語モデル (LLM) は、証拠合成におけるスクリーニングにますます使用されており、偽陰性により全文評価の前に関連する研究が削除される可能性があります。私たちは、概念的に複雑なスコーピングレビューに組み込まれた事前登録された研究において、人間とLLMのタイトルと要約によるスクリーニングワークフローを比較しました。方法。保守的なタイトルのみのスクリーニングの後、1,131 件のレコードが 1 人のレビューリーダー、4 人の訓練を受けたアシスタントによって重複しないサブセットをスクリーニングされ、異なるモデルと処理構成を使用して名目上同一の反復実行を含む 7 回の完全な LLM 実行がスクリーニングされました。私たちは、保持されたワークロード、316 件の検証済み適格記録に対する運用リコール、合意、実行間の一貫性、および手順の負担を比較しました。適格性は親レビューで進められ評価された記録に対してのみ検証されたため、再現率の推定は有効でした。結果。すべての検証済みの対象レコードを回復したワークフローはありませんでした。人間のワークフローと 2 回の GPT-5.4 ファイル バッチ実行では、レコードの 42.2 ~ 45.0% が保持され、再現率は 82.3 ~ 82.9% でした。 Gemini 3.1 ファイル バッチは最高の再現率 (83.9%) を達成しましたが、レコードの 56.7% を保持しました。オールアットワンス構成では、対応するファイルバッチ構成よりも回復できる対象レコードの数が少なくなりました。名目上同一の GPT-5.4 ファイル バッチを 2 回実行すると、レコードの 91.7% で一致しましたが、1 回の実行のみで保持された検証済みの適格レコード 29 個を含む 94 個のレコードで相違がありました。議論。 LLM スクリーニングのパフォーマンスは、モデル ID だけではなく、実装されたワークフローに依存しました。したがって、処理構成、ワークロード、レコードレベルの変動、人間と LLM の意思決定の統合は、導入されたシステムの実質的な特性となります。再現率の高いタスクの場合、LLM は自律的な除外よりも、検証済み、監査可能な、人間が監視するワークフローに適しています。
原文 (English)
Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency
Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review. Methods. After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational. Results. No workflow recovered all verified eligible records. The human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records while achieving 82.3-82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All-at-once configurations recovered fewer eligible records than corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run. Discussion. LLM screening performance depended on the implemented workflow, not model identity alone. Processing configuration, workload, record-level variation, and human-LLM decision integration are therefore substantive properties of deployed systems. For high-recall tasks, LLMs are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.
信頼性の低いアドバイスに基づく学習拡張オンライン割り当て: ロバスト性、露出の公平性、および分布のシフト
学習拡張アルゴリズムは、予測を使用してオンラインでの意思決定を改善しますが、信頼性の低いアドバイスは効率と公平性を損なう可能性があります。私たちは、有限の候補セット、不可逆的な決定、およびエクスポージャーの制約を伴うオンライン割り当て問題を研究します。私たちは、アドバイスと保守的なフォールバックおよび公平性の修正を組み合わせた、堅牢で公平なルールを提案します。誤差が有限であるという仮定の下で、予測誤差に比例する損失を伴う一貫性と堅牢性を証明します。実験では、敵対的なアドバイスの下でも安定性があり、暴露格差が大幅に減少することが示されています。
原文 (English)
Learning-Augmented Online Allocation under Unreliable Advice: Robustness, Exposure Fairness, and Distribution Shift
Learning-augmented algorithms improve online decisions using predictions, but unreliable advice may harm efficiency and fairness. We study an online allocation problem with finite candidate sets, irreversible decisions, and exposure constraints. We propose a robust and fair rule combining advice with a conservative fallback and fairness correction. Under bounded-error assumptions, we prove consistency and robustness with loss proportional to prediction error. Experiments show stability under adversarial advice and significant reductions in exposure disparity.
AI agents in Algorithmic Electricity Markets: On the Emergence of Tacit Collusion
As electricity market participants increasingly adopt learning-based agents for their bidding strategies, electricity markets are becoming…
Counterfactual Bias Testing for Application Tracking System
Automated candidate-job matching systems are increasingly classified as high-risk AI under emerging regulation, yet auditing them for demog…
A Table Is Worth 64 Tokens: Pixel-level Compression for Multi-Table Document Question Answering
Answering questions over real-world documents requires processing long inputs that interleave text with tables. Optical context compression…
From Atomic to Agentic: Towards Interpretable Evaluation of LLMs' Agentic Mathematical Capabilities
Large Language Models (LLMs) are evolving from performing end-to-end mathematical reasoning to integrating agentic intelligence. However, m…
GraphMemix: Query-Aware Evidence Forests for Long-Term Multimodal Agent Memory
Organizing long-term memory for multimodal agents remains challenging because existing methods either suffer from expensive question-agnost…
DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research
Large language models can summarize financial information, but an operational stock-research system must first assemble heterogeneous evide…
ASIL: Replacing Screenshot-and-Click with Structured State and Semantic Actions
Powerful code agents can execute scripts, call tools, and manage files, yet many important applications remain accessible primarily through…
A Multi-Modal AI Framework for Real-Time Queue Prediction, Management and Optimisation in Intelligent Border Control Systems
In the present work an efficient border control management procedure is proposed. Compared to operational queue management systems, whose o…
Omni-Interactive Universal Embedder
Multimodal representation learning has been shifting from traditional two-tower architectures to large language model (LLM)-based embedders…
A Contract-Centered Architecture for Scalable and Manageable Agentic Runtimes
Enterprise AI deployment is a coordination problem across business units, application and AI teams, testing, platform engineering, infrastr…
pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning
Ontology learning from text remains challenging despite significant progress in Large Language Models (LLMs), which can hallucinate domain…
LAAF: A Layered Accountability Architecture Framework for LLM Applications
Large Language Models (LLMs) operate in hospitals, courtrooms, banks, and public service desks, where fluent, confident outputs are treated…
TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation
Internet memes are a pervasive form of multimodal online communication; however, such communication often involves users from diverse lingu…
GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL
Despite their potential in standardized graph tasks, Large Language Models (LLMs) remain brittle to real-world shifts in node identifiers a…
Feature Transformation Enhanced Jacobi Polynomial Graph Filtering for Graph Anomaly Detection
In recent years, graph anomaly detection (GAD) based on frequency-domain filtering have achieved promising results. However, existing appro…
When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents
Tool-augmented LLM agents must rely on untrusted runtime Observations to complete open-ended tasks; however, when tool outputs no longer me…
Thomson: Continual Learning of Frontier Models for SovereignAI
The development of frontier models is commonly perceived to be the exclusive remit of a small number of heavily funded players, creating an…
BPMN4CAI: A BPMN Extension for Modeling Dynamic Conversational AI
Conversational AI systems, such as chatbots and virtual assistants, are becoming increasingly important to digital business processes. Howe…
Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable
An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often th…
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation mus…
Naive Prompt Optimization: Rethinking the Need for Complex Prompt Search
Efficiently improving autonomous agents across diverse tasks is central to accelerating recursive self-improvement (RSI) in agentic AI, wit…
BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models
Although Large language models (LLMs) mediate access to knowledge and computational assistance, their capabilities should benefit vulnerabl…
LLMs Can Design Near-Optimal OR Algorithms
We ask whether large language models (LLMs) can design effective algorithms for well-specified operations research (OR) problems. We study…
Verify Smarter, Evolve Further: Efficient Harness Evolution through Behavior-Aware Verification
Agent harnesses shape how language-model agents use instructions, tools, and runtime components, but adapting these harnesses requires cost…
Not All Eval-Awareness Is Equal: Capabilities Framing Predicts Compliance
Steering interventions targeting eval-awareness, a model's recognition that it is being tested, are increasingly used in safety evaluation…
Sophistication in GenAI Use: Field Evidence from a Large Firm
We study how sophistication in generative AI (genAI) use varies among the back-office workforce of a large firm. Using proprietary data, we…
CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases
LLMs are increasingly able to answer complex questions about enterprise-scale document collections. But evaluation is hard: companies don't…
Learning a Continuous Sepsis Severity Score Without Hour-by-Hour Supervision: A Two-Site Retrospective Study
Currently used sepsis severity indices rely on fixed variables and weights established decades ago, which are coarsely discretized and cali…
Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation
Chemical reactions are fundamentally transformations in electron space, yet most machine learning approaches model them either through \tex…
WikiSkill: エージェントの経験をスキル進化のための永続的な知識にまとめる
エージェント スキルは、専門知識とワークフローを再利用可能なリソースにパッケージ化し、AI エージェントの機能を拡張します。最近の取り組みでは、エージェントの経験からそのようなスキルを自動的に発見し、エージェントが対話を通じて徐々に適応できるようにしています。ただし、スキル開発の指針となる洞察は通常、最適化履歴全体に分散したままとなり、繰り返しにわたる体系的な再利用が制限されます。エージェントのスキルを永続的なナレッジ ベース (Wiki) と共進化させるフレームワークである WikiSkill を紹介します。高いレベルでは、WikiSkill は未加工の実行経験、蓄積された知識、および実行可能なスキルを分離しながら、継続的に経験を Wiki に統合し、その後のスキルの更新に基づいて構築することができます。 WikiSkill は、さまざまなベンチマークやモデルにわたって、一貫して最先端のスキル進化手法を上回り、ほとんどのモデル ベンチマーク設定でスキルなしのベースラインよりも向上しています。スキルの進化がモデルのスケーリングを補完することがわかりました。一般に、大規模なモデルは進化したスキルからより多くの恩恵を受けますが、スキルのある小規模なモデルは、スキルのない大規模なモデルよりも大幅に優れたパフォーマンスを発揮できます。また、進化したスキルはモデルやモデルファミリー間で効果的に伝達され、他のモデルによって進化したスキルは自己進化したスキルよりも優れたパフォーマンスを発揮できることもわかりました。最後に、私たちのアブレーション研究により、効果的なスキル向上には Wiki への永続的な知識の蓄積が重要であることが確認されました。これらの結果は、再利用可能で移転可能なスキルを開発するために、エージェントの経験を体系的に蓄積および洗練することの利点を示しています。
原文 (English)
WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
Agent skills package specialized knowledge and workflows into reusable resources that extend AI agent capabilities. Recent work automatically discovers such skills from agent experience, which enables agents to progressively adapt through interaction. However, the insights that guide skill development typically remain scattered across optimization histories, limiting their systematic reuse across iterations. We introduce WikiSkill, a framework that co-evolves agent skills with a persistent knowledge base (wiki). At a high level, WikiSkill separates raw execution experience, accumulated knowledge, and executable skills, while continuously consolidating experience into the wiki, which subsequent skill updates can build on. Across diverse benchmarks and models, WikiSkill consistently outperforms state-of-the-art skill-evolution methods and improves over no-skill baselines in most model-benchmark settings. We find that skill evolution complements model scaling: larger models generally benefit more from evolved skills, while smaller models with skills can outperform substantially larger models without them. We also find that evolved skills transfer effectively across models and model families, and skills evolved by other models can outperform self-evolved skills. Finally, our ablation studies confirm that persistent knowledge accumulation in the wiki is critical for effective skill evolution. These results demonstrate the benefits of systematically accumulating and refining agent experience for developing reusable and transferable skills.
Exploring the Role of LLMs in HPC Programming: A Survey
Large Language Models (LLMs) are emerging as promising assistants in High-Performance Computing (HPC), where programming remains complex an…
From SQL to Knowledge Graphs: An LLM-Driven Multi-Agent Approach with Data Schema Improvement
RDBMS (Relational Database Management System) databases face several limitations, including slow execution with multi-hop queries and a lac…
Training-Time Explainability for Multilingual Hate Speech Detection: Aligning Model Reasoning with Human Rationales
Online hate against Muslim communities often appears in culturally coded, multilingual forms that evade conventional AI moderation. Such sy…
FIRSTPASS: 実際の編集結果に基づくマルチドメイン、マルチラウンドのピアレビュー データセット
科学的な査読データセットは、コンピューター サイエンスと機械学習の場でのみ AI システムをトレーニングし、アブレーション研究を批判するモデルを生成しましたが、生物学の査読者が汚染管理や化学者の質問である核磁気共鳴 (NMR) スペクトルの割り当てを要求するのを見たことがありません。 FIRSTPASS は、学際的な影響力の高いジャーナルからの完全な複数ラウンドの編集対話に基づいて構築された初の大規模査読データセットです。 Nature Communications の義務的な透明性のある査読 (2022 年 11 月に制定) から厳選された FIRSTPASS は、5 つの科学分野 (生物学、化学、神経科学、物理学、地球科学) にわたる 3,668 件の記録で構成され、最初の査読者レポート、著者のポイントごとの回答、最新の査読者評価など、科学的検証の完全な反復構造を捉えています。各レコードには、編集上の決定から直接得られた結果ラベルが付けられており (2 ラウンドのレビューの場合は STANDARD、3 ラウンド以上の場合は EXTENDED)、これまでのすべてのコーパスにはないグランド トゥルースが提供されます。自動監査により、コンテンツの完全性が 100% 確認されます。専門家のレビューは平均 2,155 ワードで、会議会場のレビューよりもかなり密度が高くなります。すべてのデータ、解析パイプライン、評価スクリプトがリリースされ、分野を超えた AI 科学的判断の再現可能なベンチマークが可能になります。
原文 (English)
FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes
Scientific peer review datasets have trained AI systems exclusively on Computer Science and Machine Learning venues, producing models that critique ablation studies yet have never seen a biology reviewer demand contamination controls or a chemist question Nuclear Magnetic Resonance (NMR) spectral assignments. We introduce FIRSTPASS, the first large-scale peer review dataset built on complete multi-round editorial dialogues from a multidisciplinary high-impact journal. Curated from Nature Communications mandatory transparent peer review (instituted November 2022), FIRSTPASS comprises 3,668 records spanning five scientific domains (biology, chemistry, neuroscience, physics, and earth science), capturing the full iterative structure of scientific validation: initial referee reports, author point-by-point responses, and updated reviewer assessments. Each record carries an outcome label derived directly from editorial decisions (STANDARD for two-round review; EXTENDED for three or more rounds), providing ground truth absent in all prior corpora. An automated audit confirms 100% content integrity. Expert reviews average 2,155 words, substantially denser than conference venue reviews. All data, parsing pipelines, and evaluation scripts are released to enable reproducible benchmarking of AI scientific judgment across disciplines.
Syntax vs. Semantics: How Transformers Learn Deep Dependencies
Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic…
Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation
Referring Expression Segmentation (RES) aims to generate pixel-wise segmentation masks from complex and implicit textual queries. While rec…
Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes
Memes have turned out to be a powerful tool through which individuals share their ideas concerning contemporary social and political proble…
Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search
Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in…
Evaluating AI Generated Summaries for Cancer Patients
Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data.…
VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation
Multimodal large language models have advanced rapidly, yet most remain English-centric, as scaling multilingual multimodal instruction tun…
Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation
Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs a…
Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models
Although Large Language Models (LLMs) demonstrate remarkable capabilities in reasoning and decision-making, high-fidelity probabilistic sam…
From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents
Cough events during live spoken conversations carry clinically valuable respiratory signals, yet existing dialogue systems treat them as ac…
Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment
Automated creativity assessment has been a long standing challenge, with traditional methods often being resource intensive or lacking prac…
Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning
Advancing reasoning capabilities allow large language models (LLMs) to tackle increasingly complex problems, while reasoning traces - inter…
Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention
The lifecycle of hallucination in LLMs is a concept that enables building solid frameworks on the control and reliability of LLMs in high-s…
DRL: A Deterministic Relational Middleware Layer for Transaction-Safe Enterprise NL2SQL Under Schema-Graph Scaling
Deploying natural-language interfaces over enterprise OLTP catalogs fails at scale because semantic parsers collapse under schema-graph sca…
ClassVision: AI-Powered Classroom Attendance System
Students and working professionals have to go through the attendance process every day. Traditional methods of marking attendance using pen…
Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors
Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned compressors such as LLMLi…
A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs
Long-form generation exposes fundamental limitations of large language models. Even 70B-parameter models exhibit length collapse at 16k-tok…
PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants
Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper contex…
Investigating the Influence of Prompt and Response Languages on LLM Content Generation
This study examines how prompt and response language influence the behavior of large language models. Using five models, we evaluated answe…
When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models
Whether large language models (LLMs) can perform the abductive leap from evidence to a new system of axioms, commonly referred to as a jump…
Comparing Chunking and Embedding Strategies for Turkish RAG Systems
How documents are segmented into retrievable chunks and how those chunks are embedded strongly affect Retrieval-Augmented Generation (RAG)…
A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers
Retrieval-Augmented Generation (RAG) systems have attracted significant interest for their ability to mitigate hallucinations in Large Lang…
ADeptS-Bench: Measuring the Trustworthiness of Computer Use Agents Across Devices
Computer Use Agents (CUAs) are increasingly deployed to navigate mobile and desktop applications on behalf of users, yet no benchmark compr…
Fairness Invariants: A Relational Approach to Explaining and Mitigating Fairness Bugs
Data-driven software systems are increasingly deployed in high-stakes socio-economic domains, from criminal justice to financial lending. H…
Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model
As generative AI gains traction, researchers are investigating its potential to serve as proxies for humans. From undergoing cognitive psyc…
NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation
Safety evaluation is critical for assessing whether aligned Large Language Models (LLMs) remain robust against jailbreak attacks. Existing…
How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation
Capture-the-Flag (CTF) benchmarks are widely used to assess the offensive security capabilities of autonomous language-model agents. Evalua…
On Scope Classification and Current Knowledge-Editing Benchmarks: A Negative Result, with INLAY as a Gradient-Free Case Study
Every memory-based knowledge editor in the SERAC lineage depends on a scope decision: given a query, does a stored edit apply? We report th…
MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models
Tool-augmented LLMs must arbitrate between two fallible sources when a tool return conflicts with their parametric memory, yet existing eva…
Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models
Frontier language models are increasingly marketed as omni systems that can perceive and respond across modalities. Existing evaluation fra…
How Unlikely Is "Unlikely"? Assessing Verbal Probability Perception Across Large Language Models
Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent mea…
Decay-Region Group Delay as a Forensic Cue for AI-Generated Impulsive Sounds
We investigate whether AI-generated impulsive sounds can be distinguished from real ones through group delay analysis. Our central finding…
Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives
Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where u…
CG4AI: A Column Generation Framework for Training AI Models Under Constraints
Standard machine-learning training minimizes a loss function over a dataset, but does not guarantee that the resulting model will satisfy p…
Why RAGs Hallucinate: Penalty-Aware Evaluation of Retrieval-Augmented Generation Systems with Knowledge-Gap Canaries
Volume-based accuracy rewards retrieval-augmented generation (RAG) systems for guessing: a system that answers everything outscores one tha…
Co-Evolving Structured Knowledge and Reasoning in Language Models
Retrieval-augmented methods improve factual accuracy by grounding language models in external knowledge, but retrieving over unstructured t…
Simultaneous Envy and Equitability Guarantees
Recent work in fair division has focused on either simultaneously satisfying closely related fairness notions or achieving a single notion…
Redwood: A Frontier AI Accelerator Designed, Verified, and Deployed from Scratch in 2 Weeks by AI
Modern AI workloads and the hardware that runs them evolve on different timescales: architectural definition precedes volume silicon by yea…
The Latent Diagnostic Taxonomy: A Framework for Constructing Classifiers and Diagnosing Their Decisions, Applied to Prompt Injection Detection
This paper proposes a framework for constructing a classifier as a safeguard layer, and for developing a complementary diagnostic that iden…
SpeechGym: An Audio-Native Gym for Training Voice Agents via Reinforcement Learning
Voice agents must call tools and hold multi-turn dialogue entirely through speech, yet the dominant paradigm trains them in text. Existing…
Diff Mining: Logit Differences Reveal Finetuning Objectives
Finetuning has become the gold standard for refining existing behaviors and inducing new ones in language models, yet it often remains uncl…
Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance
Multi-agent large language model systems are widely reported to beat single-model baselines, but the evidence is mixed, and comparisons are…
RTNav: Towards Real-Time Zero-Shot Object Navigation
Navigation in unknown environments to find unforeseen objects has become increasingly feasible with capable vision and language foundation…
Physics-Informed Stochastic Configuration Machine: A Backpropagation-Free Neural Network with Fast Training for Nonlinear Differential Equations
While Physics-Informed Neural Networks (PINNs) have emerged as a transformative paradigm for solving complex differential equations, their…
J-Zero: Unified Challenger--Solver--Judge Co-Evolution from Zero Data
Self-evolving language models have recently emerged as a promising path toward superintelligence, with the advantage of reducing the cost o…
Risks and Controls for Multi-Agent Systems: an analytical framework for deployment of AI agents across organisational boundaries
This report presents a framework to help organisations, policymakers and researchers reason about the risks that emerge when AI agents inte…
CoGeo-GS: Concept-Driven and Geometry-Aware Multi-Object Removal in 3D Scenes
Multi-object removal in 3D scenes is challenging due to severe occlusions, semantic entanglement, and the difficulty of maintaining geometr…
PailitaoGR: Latent Think-with-Images for Generative Image Retrieval
Generative retrieval has demonstrated strong performance by directly generating product semantic identifiers (SIDs). Extending this paradig…
Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference
As large language models are increasingly deployed to simulate diverse human characters, ensuring persona fidelity, defined as the extent t…
FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models
Pruning is a practical approach to compress large language models (LLMs), but it can amplify text degeneration, especially repetition loops…
AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability
Recent advances in Multimodal Large Language Models (MLLMs) have extended Image Aesthetic Assessment (IAA) beyond scalar scores toward inte…
LiveVVT: High-Fidelity Video Virtual Try-On in Real Time
Diffusion-based Video Virtual Try-On (VVT) achieves high visual fidelity through bidirectional spatio-temporal modeling, but complete-clip…
Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning
Graph neural networks (GNNs) are typically conceptualized as message-passing neural networks, yet it remains unclear why neighborhood aggre…
Daydreaming: Stealing Hidden Agent Skills through Black-Box Task Interaction
Agent skills bundle instructions, reference data, and executable helpers that let a general agent perform specialized tasks. Hosted provide…
FaultLens: Learning Compact Behavioral Test Suites for Generated Operational Programs
Generated operational programs are often validated with either a few hand-written examples or exhaustive regression suites. The former can…
Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research
LLM agents used for scientific experimentation must do more than generate executable code: they must implement the reference method faithfu…
Behavior2Trip: Towards Personalized Travel Planning via User Behavior Trajectory
Travel planning agents assist users in generating personalized travel plans by modeling their individual preferences. Existing agents eithe…
Evaluating Confidence-Gated Retrieval with Matched Trajectory Replay
Interactive language-model agents use confidence signals to decide whether to answer immediately, retrieve additional evidence (from memory…
MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA
Medical Visual Question Answering (Med-VQA) holds significant promise for clinical decision support, yet faces challenges due to limited an…
From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation
Although Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in Medical Visual Question Answering (Med-VQA),…
Reinforcement Learning-Based Control of CAV Platoon Joining Maneuvers in Mixed Traffic
Connected and automated vehicle (CAV) platooning offers a promising approach to improving road safety and traffic capacity. However, platoo…
PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?
Industrial control systems (ICSs) rely on programmable logic controllers (PLCs) to connect networked computation with physical control. Too…
When Memory Takes Gradients: Collaborative Vector Memory for Agentic Recommender Systems
Agentic recommender systems ground each decision of a large language model (LLM) in a persistent memory of the user, and in existing agents…
Per-View Gaussian Predictions Enable Training-Free Distractor Filtering in Feed-Forward 3DGS
Feed-forward 3D Gaussian Splatting reconstructs an explicit Gaussian representation from multiple input images in one network execution, ma…
Magnon-induced phononic Chern insulator
High-frequency artificial phononic crystals offer a low-loss platform compatible with on-chip integration, yet realizing Chern phononic pha…
FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets
LLM-based agents are increasingly proposed for network fault diagnosis, but existing benchmarks evaluate them only on accurate tickets and…
Multi-Person Human Motion Forecasting in Complex Scenes
Accurately forecasting the movement of people in complex scenes requires reasoning over the past and present state of the entire environmen…
Performance Foundations of Parallel & Distributed Reasoning Language Models
Reinforcement Learning with Verifiable Rewards (RLVR) and other RL-style post-training paradigms have been used for aligning large language…
Beyond Classification: Task-Dependent Learnability under Privacy-Motivated Image Transformations
Privacy-Enhancing Technologies (PETs) in computer vision often rely on noise or image perturbations to protect visual data while securely p…
Emotional Preferences as Goal-Priority Regulation
A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by e…
Learning Transverse Momentum Distributions from Raw Scattering Events via Conditional Diffusion
Extracting transverse momentum dependent parton distribution functions (TMD PDFs) from semi-inclusive deep inelastic scattering (SIDIS) dat…
Active Diffusion-Based Inference for Ill-Posed Inverse Problems under Incomplete Priors
Many scientific and engineering applications require estimating unknown parameters from experimentally observable data -- an inverse proble…
Active sensing to characterize the heterogeneity of plant stress
While most phenotyping platforms rely primarily on image-based measurements, advanced plant characterization requires the integration of ac…
Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents
Large language model agents are increasingly deployed as autonomous loops. Starting from one human goal, such a system repeatedly discovers…
ANTShapes Benchmarking Datasets for Event-Based Neuromorphic Object Classification
Object classification in event-based computer vision is a task that is attracting considerable research attention. Event-based object class…
When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue
Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and convers…
LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration
Electronic design automation (EDA) has advanced engineering productivity through successive generations of tooling that progressively autom…
PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference
Vision-Language Models (VLMs) demonstrate exceptional visual reasoning capabilities, yet their inference costs escalate rapidly with the pr…
STEP: State-Aware Task Estimation and Planning with Multi-Modal LLMs for Human-Robot Collaboration
Effective human-robot collaboration in industrial settings requires robots to understand human intentions and assist with task planning, re…
Compositional Online Learning for Semantic Data Processing Systems
An LLM call in a semantic data processing system is expensive enough to dominate query cost, yet slow enough to hide a CPU-side learner's u…
TADP: Task-Aware Deformable Prediction for Single-Stage 3D Object Detection
Most single-stage 3D object detectors complete different tasks with the same extracted features. Nevertheless, it is impossible to project…
Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit
Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast b…
PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Recent video generation models are increasingly framed as world models. Many physical processes can unfold in more than one valid way. Ther…
RCMN: Understanding Misleadingness in Influential Public Discourse
Influential public discourse shapes public beliefs and can also mislead, not only through what is stated, but also through how information…
KnockGS:interaction-Grounded Calibrationof Physical Gaussian Representations
Physics-integrated 3D Gaussian representations now allow reconstructed deformable objects to be simulated and rendered under explicit mater…
Stageboost: Recommending Signals Based on Counterfactual Estimation
Signals are short textual or visual snippets displayed on the eBay View-Item (VI) page, providing additional, contextual information for us…
Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models
Joint-Embedding Predictive Architectures (JEPAs) for world modeling typically employ fixed-size Vision Transformer encoders that are over-p…
Property-Specific Recoverability from Contact PPG to Camera rPPG under Heterogeneous Observation Conditions
Camera-derived remote photoplethysmography (rPPG) is commonly validated through endpoint accuracy, but endpoint performance does not establ…
LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
Video carries the temporal structure of the physical world, yet learning representations from it has remained computationally expensive: pr…
Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction
Clinical language models can achieve strong in-hospital accuracy yet fail under deployment shifts because they exploit note-specific artifa…
How Language Models Organize and Structure Moral Knowledge
How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask wheth…
CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the…
Beyond F1: Evaluating Coverage and Failure Recovery in AI Model Security Scanners
Static scanners are increasingly used to identify executable or otherwise unsafe content in machine- learning artifacts, yet conventional e…
Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit
Large language model (LLM) agents in governed organizations must let the persona (instructions, tone, self-presentation) evolve freely, whi…
RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution
LLM-based agents are increasingly deployed in product-level execution harnesses, where jailbreaks can trigger harmful tool use and persiste…
From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench
In real-world software development, code review typically involves iterative interactions between developers and reviewers to improve softw…
SWE-Prime: Fewer Trajectories, Better Performance
To improve large language models' ability to resolve real-world software issues, prior work has focused on constructing large-scale agent t…
Designing Cellular Manufacturing Systems in the Presence of Alternative Process Plans
In the design of cellular manufacturing systems (CMS), numerous technological and managerial decisions must be made at both the design and…
LLM-Powered Swarms: A New Frontier or a Conceptual Stretch?
Swarm intelligence describes how simple, decentralized agents can collectively produce complex behaviors. Recently, the concept of swarming…
Pushing the Envelope of LLM Inference with Ultra-Low-Bit Quantized Models
The advent of ultra-low-bit LLM models, approaching the perplexity and task accuracy of their full precision counterparts, is ushering in a…
Do Language Models Follow Occam's Razor? An Evaluation of Parsimony in Inductive and Abductive Reasoning
Non-deductive reasoning, encompassing inductive and abductive reasoning, is essential in addressing complex real-world questions. One key f…
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
As agentic AI systems are deployed in advisory and evaluative roles, understanding how multi-agent interactions shape behavior becomes esse…
DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping
Large language models (LLMs) augmented with multi-step reasoning and action generation abilities have shown promise in leveraging external…
Beyond Linearization: Attributed Table Graphs for Table Reasoning
Table reasoning, a task to answer questions by reasoning over data presented in tables, is an important topic due to the prevalence of know…
DIANOIA: マルチエージェント推論のための診断分解と共同最適化
マルチエージェント LLM システムは、単一エージェントのベースラインを常に上回っていますが、実務者は依然として、どの設計が新しいタスクに機能するかを予測したり、失敗する理由を診断したりすることができません。このギャップが依然として存在するのは、この分野に測定可能なプリミティブとテスト可能な予測を備えた診断フレームワークが欠如していることが主な原因であると我々は主張する。 \textbf{DIANOIA} を導入します。これは、マルチエージェント推論のゲインをカバレッジ、忠実度、合成に 3 つのチャネルに分解したもので、それぞれが経験的に測定可能です。この分解から、特定のタスクのボトルネック チャネルを特定する診断プロトコルを導き出します。このプロトコルをマルチエージェント システムとしてインスタンス化します。その 3 つのコンポーネントはチャネルを反映しています。カバレッジのための役割の多様なプロポーザー、忠実性のための実行ベースの検証、および反復合成です。 GSM8K、AIME-2025、MBPP、BFCL-SP 上で、私たちの方法は、一致するトークン予算の下で強力なマルチエージェント ベースラインを上回り、$\sim$$5{\times}$ トークンの節約で MBPP のパレート フロンティアを支配し、一致するコストで $+4.6$pp に達します。すべてのベンチマークで、プロトコルは適切なボトルネック チャネルを選択します。私たちがそれを中心に構築したシステムは、モデル全体をリードします。コード、アダプター、診断メトリクス、およびクロード コード スキルを https://anonymous.4open.science/r/DIANOIA4MAS でリリースします。 DIANOIA は、マルチエージェント設計をチャネルを意識したリソース割り当てとして再構成します。タスクのボトルネックとなっているチャネルを診断し、それに応じてトークンを投資します。
原文 (English)
DIANOIA: Diagnostic Decomposition and Joint Optimization for Multi-Agent Reasoning
Multi-agent LLM systems consistently outperform single-agent baselines, yet practitioners still cannot predict which design works for a new task or diagnose why one fails. We argue this gap persists largely because the field lacks a diagnostic framework with measurable primitives and testable predictions. We introduce \textbf{DIANOIA}, a three-channel decomposition of multi-agent reasoning gain into coverage, fidelity, and synthesis, each of which is empirically measurable. From this decomposition, we derive a diagnostic protocol that identifies the bottleneck channels for any given task. We instantiate the protocol as a multi-agent system whose three components mirror the channels: role-diverse proposers for coverage, execution-grounded verification for fidelity, and iterative synthesis. On GSM8K, AIME-2025, MBPP, and BFCL-SP, our method outperforms strong multi-agent baselines under matched token budgets, dominating the Pareto frontier on MBPP at $\sim$$5{\times}$ token savings and reaching $+4.6$pp at matched cost. On every benchmark, the protocol picks the right bottleneck channels; the system we built around it leads across models. We release code, adapters, diagnostic metrics, and a Claude Code skill at https://anonymous.4open.science/r/DIANOIA4MAS. DIANOIA reframes multi-agent design as channel-aware resource allocation: diagnose which channel is the bottleneck for your task, then invest tokens accordingly.
Learning to Predict, Discover, and Reason in High-Dimensional Event Sequences
Electronic control units (ECUs) embedded within modern vehicles generate a large number of asynchronous events known as diagnostic trouble…
Nomad: Autonomous Exploration and Discovery
We introduce Nomad, a system for autonomous data exploration and insight discovery. Given a corpus of documents, databases, or other data s…
正確さから監査可能性へ: 金融 AI システムにおける決定論の調査
信用リスク、不正行為検出、マネーロンダリング対策といった規制された金融環境に機械学習を導入すると、アルゴリズムの再現性における重大な脆弱性が露呈します。初期の金融 ML はバックテストのオーバーフィッティングなどの統計的課題に対処しましたが、ディープ ニューラル ネットワークと生成 AI では、ハードウェアとアーキテクチャに根ざした機械的非決定性が導入されました。この調査では、表形式モデル (事後説明の分散)、グラフ ネットワーク (確率的サンプリングと時間的非同期性)、LLM ベースのエージェント ワークフロー (バッチ依存の発散と軌道ドリフト) という、金融 AI で現在主流となっている 3 つの手法にわたる再現性の障害に関するシステムの視点を提供します。公的金融データセットに関するファーストパーティの実験で文献分析を補足します。信用スコアリングにおける説明ランクの不安定性、GNN ベースの不正検出における予測フリップ レート、LLM エンティティ抽出におけるテンソル並列誘発出力の発散を定量化します。我々は、モダリティ固有の指標(RBO、D_cos、TDI、PSD)を監査の準備状況にリンクする階層化された評価フレームワークを提案し、ロジットレベルとセマンティックレベルの決定性尺度の相補性を経験的に検証します。
原文 (English)
From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems
Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility. While early financial ML addressed statistical challenges such as backtest overfitting, deep neural networks and Generative AI have introduced mechanical nondeterminism rooted in hardware and architecture. This survey provides a systems perspective on reproducibility failures across three modalities now dominant in financial AI: tabular models (post-hoc explanation variance), graph networks (stochastic sampling and temporal asynchrony), and LLM-based agentic workflows (batch-dependent divergence and trajectory drift). We supplement the literature analysis with first-party experiments on public financial datasets -- quantifying explanation rank instability in credit scoring, prediction flip rates in GNN-based fraud detection, and tensor-parallel-induced output divergence in LLM entity extraction. We propose a layered evaluation framework linking modality-specific metrics (RBO, D_cos, TDI, PSD) to audit readiness, and report where these measures overlap rather than complement one another.
TouchThinker: 大規模なデータとアクションを意識した表現を使用して、触覚的常識推論をオープンワールドに拡張する
接触は、肉体を持ったエージェントが物理世界を理解するための重要なモダリティです。最近の研究では、触覚常識推論のための言語システムに触覚信号が組み込まれていますが、そのようなシステムを現実的なオープンワールド設定に拡張することは、2 つの重要なボトルネックのため依然として困難です。(1) 現在の触覚推論データセットは形式と規模が制限されたままであり、触覚観察から物理的常識への推論に対する監視が不十分であり、伝達可能な触覚常識の学習を妨げています。 (2) 触覚信号は本質的に冗長でアクション固有ですが、既存の方法ではこれらの特性が見落とされることが多く、その結果、意味表現力が限られた非効率な表現が生じます。これらの制限に対処するために、私たちは、データと表現の両方の観点から触覚の常識的推論をオープンワールドに拡張する触覚言語フレームワークである TouchThinker を提案します。まず、\textbf{415} オブジェクト、\textbf{8} シナリオ、\textbf{7} センサー タイプをカバーする百万規模のマルチソース触覚推論データセットである TouchThinker-1M を構築し、オープンワールドの一般化のための強固なデータ基盤を提供します。さらに、より現実的で多様なタスクを備えたオープンワールドのベンチマークである TouchThinker-Bench を紹介します。次に、触覚表現の効率を向上させ、効率的な推論を可能にするアクション認識モデリングメカニズムを提案します。実験結果は、TouchThinker が複数のデータセットにわたって最先端のモデルに対して競争力のあるパフォーマンスを達成することを示しています。私たちのコードとデータセットは、https://github.com/lvkailin0118/TouchThinker で利用できるようになります。
原文 (English)
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense reasoning, scaling such systems to realistic open-world settings remains challenging due to two key bottlenecks: (1) current tactile reasoning datasets remain limited in format and scale, providing insufficient supervision for reasoning from tactile observations to physical commonsense and hindering the learning of transferable tactile commonsense; (2) tactile signals are inherently redundant and action-specific, yet existing methods often overlook these properties, resulting in inefficient representations with limited semantic expressiveness. To address these limitations, we propose TouchThinker, a tactile-language framework that scales tactile commonsense reasoning to the open world from both data and representation perspectives. First, we construct TouchThinker-1M, a million-scale, multi-source tactile reasoning dataset covering 415 objects, 8 scenarios, and 7 sensor types, providing a solid data foundation for open-world generalization. We further introduce TouchThinker-Bench, an open-world benchmark with more realistic and diverse tasks. Then, we propose action-aware modeling mechanism to improve tactile representation efficiency and enable efficient reasoning. Experimental results demonstrate that TouchThinker achieves competitive performance against state-of-the-art models across multiple datasets. Our code and dataset will be made available at: https://github.com/lvkailin0118/TouchThinker.
ProvenanceGuard: MCP ベースの LLM エージェントのソース認識事実検証
ツールを使用する LLM エージェントは、検索、API、データベース、臨床記録、処方ツールなどの異種の証拠ソースから回答するために、モデル コンテキスト プロトコル (MCP) を使用することが増えています。標準的な事実性メトリクスは、通常、答えがプールされた証拠によって裏付けられているかどうかをテストし、出所に依存する失敗モードを見逃します。つまり、主張は間違った情報源に起因しているにもかかわらず、どこかで裏付けられている可能性があります。これをクロスソースの統合と呼びます。 MCP に基づいた回答のためのソース認識検証ツールである ProvenanceGuard を紹介します。安定したツール ID、ソース ID、生の出力を含むキャプチャされた MCP トレースを消費します。回答を原子的な主張に分解します。主張を情報源固有の証拠にルーティングする。 NLI とトークン アライメント プロキシのサポートを確認します。指定された帰属とルーティングされたソースを比較します。そして、クレームごとの判定と回答レベルの許可/ブロックの決定を返します。ブロックされた回答は、検索拡張された回答改訂によって修復し、再検証することができます。 281 の医療ドメイン MCP エージェント トレースを評価します。 266 トレースの裁定サブセットから、トレースごとに分割された 2,325 個の LLM 支援クレーム ラベルが生成されます。 361 枚のラベルは人間によって検証されています。 40 トレース ホールドアウト スプリットでは、ProvenanceGuard は 260 のソース適格クレームでブロック F1 0.802 とソース精度 0.858 を達成し、クレーム対ソース ID を発行しないソース ブラインド ベースラインを上回りました。より厳しい複数ソースのベンチマークでは、ブロック F1 0.846 に達しますが、ソースと関係の精度は 0.229 に低下します。これは、意味的に近いソースでは正確なソースの所有権が依然として難しいことを示しています。修復と再検証は、多くの場合保守的なフォールバックを介して、完全なトレース セット内のすべてのブロックされた回答を解決します。 ProvenanceGuard は、50 の制御された臨床混合プローブで、間違った属性を保持せずに、注入されたすべての属性のスワップを検出します。これらの結果は、MCP ベースのエージェントにおける事実検証において、出典の帰属が独立した軸であることを示しています。
原文 (English)
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, APIs, databases, clinical records, and formulary tools. Standard factuality metrics usually test whether an answer is supported by pooled evidence, missing a provenance-sensitive failure mode: a claim may be supported somewhere while being attributed to the wrong source. We call this cross-source conflation. We introduce ProvenanceGuard, a source-aware verifier for MCP-grounded answers. It consumes captured MCP traces with stable tool IDs, source IDs, and raw outputs; decomposes answers into atomic claims; routes claims to source-specific evidence; checks support with NLI and a token-alignment proxy; compares stated attribution with the routed source; and returns per-claim verdicts plus an answer-level allow/block decision. Blocked answers can be repaired with retrieval-augmented answer revision and re-verified. We evaluate on 281 medical-domain MCP-agent traces. A 266-trace adjudicated subset yields 2,325 LLM-assisted claim labels split by trace; 361 held-out labels are human-verified. On the 40-trace held-out split, ProvenanceGuard achieves block F1 0.802 and source accuracy 0.858 over 260 source-eligible claims, outperforming source-blind baselines that do not emit claim-to-source IDs. On a harder multi-source benchmark it reaches block F1 0.846, while source-plus-relation accuracy drops to 0.229, showing that exact source ownership remains difficult with semantically close sources. Repair-and-reverify resolves all blocked answers in the full trace set, often via conservative fallback. In 50 controlled clinical conflation probes, ProvenanceGuard detects all injected attribution swaps with no retained wrong attribution. These results show that source attribution is an independent axis for factuality verification in MCP-based agents.
Learning the ARTS of Search for Automated Discovery
Scientific discovery can be formulated as an iterative search process over the space of hypotheses and experiments. Contemporary methods na…
大規模言語モデルにおける認識エントロピーを除去するためのローリング係数のヘヴィサイド連続性
大規模言語モデル (LLM) は、間違っている可能性がある流暢な出力を生成します。誤った情報を提供するときに手がかりを示すことが多い人間とは異なり、LLM は検出が困難なエラーを生成します。これは、自己回帰デコードには状態が進行する前に中間推論を検証するメカニズムがないためです。ヘビサイド ゲートによって制御される述語ゲート状態遷移として推論を再定式化する検証優先実行フレームワークであるヘビサイド ローリング係数 (HCRC) を紹介します。 HCRC は、モデルの信頼性とパラレル ワーカー アーキテクチャからの独立した検証信号を組み合わせて、事前定義された正確性述語が満たされた場合にのみ実行を進めることができます。これにより、無効な中間状態の伝播が防止され、基礎となるモデルを変更することなく認識論的エントロピーが削減されます。私たちは、4 つのプロバイダーからの 13 人の提案者にわたって、ソフトウェア エンジニアリングと推論タスクに関して HCRC を評価します。有能なプロポーザーでは、ゲートは遅延競合性を維持しながら誤完了率 (FCR) を 4 ~ 7% から 0% に削減し、設定によってはアンラップ モデルよりも高速になります。弱いプロポーザーでは、ダウンストリームの状態を破壊するのではなく、誤った完了を正当な停止に変換します。ベンチマークを超えて、HCRC はエージェント コーディング環境の実稼働コントロール プレーンとして数か月間運用され、ファイルの変更の承認、検証主導の進捗レポート、メモリ圧縮を行ってきました。これらの結果は、検証主導型 LLM 実行の一般的なフレームワークとして HCRC を確立し、信頼性の高い推論がモデル スケールだけではなく原則に基づいた実行制御を通じて達成できることを示しています。
原文 (English)
Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models
Large language models (LLMs) generate fluent outputs that can be wrong. Unlike humans, who often exhibit cues when providing false information, LLMs produce errors that are difficult to detect because autoregressive decoding provides no mechanism for verifying intermediate reasoning before state progression. We introduce Heaviside Continuity of Rolling Coefficients (HCRC), a verification-first execution framework that reformulates inference as predicate-gated state transitions governed by a Heaviside Gate. HCRC combines model confidence with independent verification signals from a parallel worker architecture, allowing execution to advance only when predefined correctness predicates are satisfied. This prevents invalid intermediate states from propagating, reducing epistemic entropy without modifying the underlying model. We evaluate HCRC on software-engineering and reasoning tasks across thirteen proposers from four providers. On capable proposers, the gate reduces the false-completion rate (FCR) from 4--7% to 0% while remaining latency-competitive and, in some settings, faster than the unwrapped model. On weaker proposers, it converts false completions into honest halts instead of corrupting downstream state. Beyond benchmarking, HCRC has operated for months as the production control plane of an agentic coding environment, authorizing file mutations, verification-driven progress reporting, and memory compaction. These results establish HCRC as a general framework for verification-driven LLM execution, showing that reliable reasoning can be achieved through principled execution control rather than model scale alone.
エージェント向けのハーネス進化の評価を再考する
LLM エージェントの自動ハーネス進化の評価を再検討します。既存のハーネス進化手法では、単体テスト ケースを使用してハーネス構成を検索し、同じ公開ベンチマークで最終パフォーマンスを報告します。このプロトコルは 2 つの基本的な懸念を引き起こします。まず、ハーネスの進化自体が反復的な検索手順であり、タスクのフィードバックを使用して候補ハーネスを繰り返し評価および修正します。したがって、エージェントのテスト時間のスケーリングと同様に、一致したフィードバックと推論バジェットの下で単純なタスクレベルの検索ベースラインと比較して、その利益がハーネス設計の改善によるものか、追加の検索のみによるものかを判断する必要があります。第 2 に、検索と最終評価は同じベンチマークを共有するため、報告されたタスクはその特定のタスク セットに過剰適合するリスクが生じます。これらの懸念に対処するために、同等のフィードバックと推論予算の下で、ハーネスの進化を単純なテスト時間のスケーリングと発見ベースラインと比較する広範な評価を実施し、また、保留されたタスクで進化したハーネスを評価して、発見された改善点が一般化するかどうかを評価します。 GPT-5.4 および Claude Opus 4.6 を使用した Terminal-Bench 2.1 での実験では、自動ハーネス進化が常に単純なテスト時間スケーリング手法を上回るパフォーマンスを発揮せず、一般化が限られていることを示しています。私たちの結果は、自動ハーネスの進化の有効性について重要な疑問を提起し、自動ハーネス設計のためのより公平な評価プロトコルとベンチマークの必要性を強調しています。私たちのコードは https://github.com/re Thinking-harness-evolution で入手できます。
原文 (English)
Rethinking the Evaluation of Harness Evolution for Agents
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. First, harness evolution is itself an iterative search procedure that repeatedly evaluates and revises candidate harnesses using task feedback. As in agentic test-time scaling, it should therefore be compared with simple task-level search baselines under matched feedback and inference budgets to determine whether its gains arise from improved harness design or from additional search alone. Second, because the search and the final evaluation share the same benchmark, the reported gains risk overfitting to that specific task set. To address these concerns, we conduct an extensive evaluation comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluate evolved harnesses on held-out tasks to assess whether the discovered improvements generalize. Experiments on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6 show that automatic harness evolution does not consistently outperform simple test-time scaling methods and exhibits limited generalization. Our results raise important questions about the effectiveness of automatic harness evolution and highlight the need for fairer evaluation protocols and benchmarks for automatic harness design. Our code is available at https://github.com/rethinking-harness-evolution.
ARC-AGI-3 を解決するには、コーディング エージェントに実行可能なワールド モデル、簡略化、検証が必要ですか?
以前の ARC-AGI-3 エージェントには、実行可能なワールド モデリング、スケジュールされた単純化、および正確なリプレイ検証がバンドルされており、どのアイデアがそのパフォーマンスに影響を及ぼしたのかは不明瞭なままでした。私たちは、この属性の問題に 4 つのネストされた Codex ベースのエージェントを使用して対処します。リプレイ検証のない、柔軟なインターフェイスの実行可能なワールド モデル。同じ実行可能モデルを計画的に簡素化したもの。そして、単純化を維持し、記録された観察の正確な再現を必要とする固定インターフェース検証処理。主な調査では、公開されている ARC-AGI-3 ゲームで、gpt-5.4 および gpt-5.5 の 4 つのエージェントすべてを高い推論努力と xhigh 推論努力で評価しました。探索的なフォローアップでは、xhigh および max で gpt-5.6-sol を使用してテキストおよび検証バリアントを評価します。最も堅牢な結果は、より強力なモデルとより大きな推論努力により、すべてのエージェントのバリアントが向上することです。各モデルエフォート設定内では、バリアント間の差は予想よりも小さいですが、個々のコンポーネントの効果は設定間で異なります。永続的な実行可能成果物を要求することは、普遍的に有益ではありません。両方の gpt-5.5 設定において、テキスト形式のバリアントは、フレキシブル インターフェイスの実行可能形式のバリアントよりも優れています。単純化により、4 つのモデルエフォート設定のうち 3 つでパフォーマンスが向上しますが、最も弱い設定が唯一の例外です。完全な検証処理は 4 つの設定すべてで 1 位にランクされていますが、かなり多くのリソースを使用します。 gpt-5.6-sol のフォローアップでは、検証バリアントは両方の推論努力ですべての公開ゲームを完全に解決し、約 99% の RHAE を達成し、人間のベースラインの合計アクションの半分未満しか使用しません。モデルはこれらのゲームよりも古いものであり、保留されたパフォーマンスはテストされていないため、この結果はパブリック セットのみの飽和として解釈される必要があります。
原文 (English)
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
Our previous ARC-AGI-3 agent bundled executable world modeling, prompted simplification, and exact replay verification, leaving their individual contributions unclear. An executable world model is a persistent, agent-authored environment hypothesis embodied in runnable code. We compare four Codex-based variants: textual; flexible-interface executable; executable with simplification prompts; and a fixed-interface variant with simplification and exact replay verification against recorded observations. The main study evaluates them with gpt-5.4 and gpt-5.5 at high and xhigh reasoning effort on 25 public games; exploratory follow-ups compare textual and verification with gpt-5.6-sol. In the main study, every variant scores higher as model capability and reasoning effort increase. These gains often exceed variant differences, which are smaller than anticipated and vary across settings. Requiring an executable deliverable is not universally beneficial: textual outperforms flexible-interface executable in both gpt-5.5 conditions. The simplification variant scores higher than its executable-only counterpart in three of four settings; the weakest is the exception. The complete verification treatment ranks first throughout, sometimes narrowly, but uses substantially more resources. With gpt-5.6-sol, the verification variant completes every public level at xhigh and max with about 99% human-relative action efficiency while using fewer than half the human baseline's total actions. At max, however, the textual variant completes every level with 41% fewer actions than the human baseline. Thus, at max, the three imposed mechanisms are not required for action-efficient public-set completion; verification nevertheless scores higher and succeeds at lower effort. Because gpt-5.6-sol postdates the games and held-out performance is untested, results indicate public-set saturation only.
画像分類ニューラルネットワークでは高重みニューロンが重要ですか?
画像分類用のニューラル ネットワーク モデルが進歩するにつれ、ニューロンは枝刈り、バックドア防御、解釈可能性において重要な役割を果たします。しかし、既存の研究では重みと重要性の関係が明確ではありません。我々は、3 つの実験を使用したニューロンの重要性評価方法でこれに対処します。高重みニューロンと精度に影響を与えるニューロン間の重複の定量化、高重みニューロンの摂動効果の分析、高重みニューロンのアブレーション後の再トレーニング後の精度のテストです。 CIFAR-10 と Mini-ImageNet の実験により、重要なパターンが明らかになりました。オーバーラップ解析により、上位 10\% の高重量ニューロンが重要なニューロンと最大でも約 25\% だけ重複し、その後の間隔ではさらに減少することがわかります。摂動テストでは、上位 10\% の高重量ニューロンが、ランダムな摂動の 3-7\% と比較して、特定の操作の下で 45-80\% の精度低下を引き起こすことがわかりましたが、それらの 3 分の 1 は最小限の影響しか示しません。アブレーション再トレーニングの結果は、上位 10\% の高重量ニューロンを除去すると精度がベースラインより 10 ~ 20\% 低くなり回復しない一方、上位 0.1\% をアブレーションするとほぼ完全に回復できることが示されています。特に、一部の低重量区間では、摂動時に 10 ~ 17% の劣化が見られ、これは中程度の高重量ニューロンに匹敵します。これらの結果は、すべての高重量ニューロンが重要であるわけではなく、その重要性が非線形であることを裏付けています。低重量ニューロンも大きく寄与します。これは重みと重要性の同等性に挑戦し、洗練されたニューロンの役割の洞察を提供します。クリティカルで重みの高いニューロンを優先する暗号化や、クリティカルでないニューロンを削除するプルーニングなどのアプリケーションをサポートし、ニューラル ネットワーク分析を進歩させます。
原文 (English)
Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
As neural network models for image classification advance, neurons play critical roles in pruning, backdoor defense, and interpretability. Yet existing work lacks clarity on the weight-importance relationship. We address this with a neuron importance assessment method using three experiments: quantifying overlap between high-weight and accuracy-impacting neurons, analyzing high-weight neuron perturbation effects, and testing post-retraining accuracy after high-weight neuron ablation. Experiments on CIFAR-10 and Mini-ImageNet reveal key patterns. Overlap analysis shows top 10\% high-weight neurons overlap with important ones by only about 25\% at maximum, dropping further in subsequent intervals. Perturbation tests find top 10\% high-weight neurons cause 45-80\% accuracy degradation under certain operations compared to 3-7\% for random perturbations, but a third of them show minimal impact. Ablation-retraining results show removing top 10\% high-weight neurons leaves accuracy 10-20\% below baseline with no recovery, while ablating top 0.1\% allows near-full recovery. Notably, some low-weight intervals show 10-17\% degradation when perturbed, comparable to mid-range high-weight neurons. These results confirm not all high-weight neurons are important: their importance is nonlinear. Low-weight neurons also contribute significantly. This challenges weight-importance equivalence, offering refined neuron role insights. It supports applications like encryption prioritizing critical high-weight neurons and pruning removing non-critical ones, advancing neural network analysis.
不完全な観察によるマルチモーダル感情分析におけるモダリティの信頼性の再考
マルチモーダル感情分析 (MSA) は、テキスト、音声、視覚を統合して人間の感情を推測しますが、現実世界のマルチモーダル観察は多くの場合不完全です。不完全観察 MSA の既存の方法は、主に 2 つのパラダイムに従います。再構成ベースの手法は観察されたモダリティから欠落した情報を回復しますが、共同表現手法は不完全な入力から直接学習します。これらの方法は効果的ではありますが、通常、モダリティの信頼性を明示的にモデル化するのではなく、表現学習または融合設計内で暗黙的にのみ処理します。我々は、モダリティの信頼性が不完全観察設定における中心的な変数であると主張します。これを明示的にモデル化しないと、2 つの関連する問題が発生します。 1 つ目は信頼性の不一致であり、各モダリティによって保持される感情的証拠がサンプルと欠落率によって異なります。 2 つ目は信頼性伝播バイアスであり、劣化したモダリティからのメッセージがクロスモーダル インタラクションや予測パフォーマンスに悪影響を与える可能性があります。これらの問題に対処するために、我々は不完全な観測を伴う MSA 用のモダリティ信頼性校正フレームワークである MRCF を提案します。 MRCF には、モーダル内の品質キューとクロスモーダルのセマンティック一貫性からサンプル固有のモダリティの信頼性を推定する信頼性認識ブランチ、推定されたスコアを使用してクロスモーダル情報フローを調整する信頼性ガイド付きインタラクション ブランチ、および最終予測のために信頼性とセマンティック キューを統合する信頼性校正済みフュージョン モジュールが含まれています。 CMU-MOSI、CMU-MOSEI、および CH-SIMS の実験では、MRCF が標準的な不完全観察プロトコルの下で強力なパフォーマンスを達成することが示されています。さらなる分析により、明示的な信頼性モデリングが相互作用および融合中の信頼性の不一致および信頼性伝播バイアスの軽減に役立つという証拠が得られます。
原文 (English)
Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations
Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are often incomplete. Existing methods for incomplete-observation MSA mainly follow two paradigms. Reconstruction-based methods recover missing information from observed modalities, while joint-representation methods learn directly from incomplete inputs. Although effective, these methods usually treat modality reliability only implicitly within representation learning or fusion design rather than modeling it explicitly. We argue that modality reliability is a central variable in incomplete-observation settings. Failure to model it explicitly gives rise to two related issues. The first is reliability mismatch, in which the affective evidence retained by each modality varies across samples and missing rates. The second is reliability propagation bias, in which messages from degraded modalities may adversely affect cross-modal interaction and predictive performance. To address these issues, we propose MRCF, a Modality Reliability-Calibrated Framework for MSA with incomplete observations. MRCF contains a Reliability-Aware Branch that estimates sample-specific modality reliability from intramodal quality cues and cross-modal semantic consistency, a Reliability-Guided Interaction Branch that uses the estimated scores to modulate cross-modal information flow, and a Reliability-Calibrated Fusion Module that integrates reliability and semantic cues for final prediction. Experiments on CMU-MOSI, CMU-MOSEI, and CH-SIMS show that MRCF achieves strong performance under standard incomplete-observation protocols. Further analyses provide evidence that explicit reliability modeling helps mitigate reliability mismatch and reliability propagation bias during interaction and fusion.
NiyamAI - ゼロ知識証明を使用した暗号検証可能なガードレールを備えたインテントバインド AI エージェント
AI エージェントに電子メールの送信、データベースのクエリ、またはコマンドの実行の機能を与えることは、エージェントがだまされてやるべきではないことを実行するまでは便利です。即時注入、幻覚的な推論、および安全でないツールの呼び出しが、自律型 LLM エージェントの主な攻撃対象領域を形成します。既存の防御策は、攻撃者がターゲットとする同じマシン上で実行されるシステム プロンプトやポリシー フィルターなどのソフトウェア チェックに依存しており、検証可能な実行証拠は提供されません。安全性の執行を証明可能にするフレームワーク Niyam-AI を紹介します。セッション開始時に、許可されたツールと制約は、SHA-256 経由でコミットされたインテント コントラクトにロックされます。すべてのツール呼び出しは、分離された Judge モデルによって傍受され、検証されます。合格すると、EZKL を介して zk-SNARK 証明が生成されます。このツールは証拠検証後にのみ実行されるため、第三者は裁判官モデルの重みにアクセスせずに執行を確認できます。 Agent-SafetyBench の 2,000 の実世界シナリオで Niyam-AI を NeMo Guardrails、Meta の Llama Prompt Guard 2、OpenAI の GPT-OSS-Safeguard に対して 5 重層別相互検証を使用して評価すると、F1 スコアは 88.5%、偽陽性率は 1.1% でした (ブートストラップ 95% CI: [85.19%、91.88%]、N=1000)。 McNemar の完全対応テストでは、大幅な改善が確認されています。Niyam-AI は、NeMo に対して 390 の不一致シナリオで勝利 (対 20 敗)、Prompt Guard 2 に対して 115 (対 13)、GPT-OSS-Safeguard に対して 384 (対 19) で勝利し、すべてのケースで p < 0.0001 でした。プルーフの生成には、承認されたアクションごとに 2260.6 +/- 218.4 ミリ秒が追加され、検証には 53.1 +/- 11.8 ミリ秒かかります。 Niyam-AI は、高精度で数学的に検証可能なガードレールを提供します。ただし、これはゼロショット ベースラインに対して評価される Agent-SafetyBench に適合した分類器を反映しています。この区別についてはセクション IV.C で説明します。
原文 (English)
NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs
Autonomous LLM agents with tool execution capabilities introduce severe security risks through prompt injection, goal hijacking, and unauthorized action invocation. Existing guardrails rely on unverified, host local software filters system prompts, semantic classifiers, policy engines that share the execution environment of the untrusted agent, offering no guarantee to an external observer that a safety policy was correctly evaluated. A compromised host produces no evidence of its own failure. This paper presents NiyamAI, an intent bound runtime guardrail architecture providing cryptographically verifiable execution integrity for autonomous agents. At session initialization, permitted tools and operational constraints are sealed into an immutable Intent Contract under a SHA256 commitment. Every tool invocation is intercepted by a deterministic authority gate and classified by a dedicated neural Judge (11->8->2 feedforward network). For each authorized action, NiyamAI generates a succinct zkSNARK proof certifying correct policy evaluation under the committed contract; execution proceeds only after that proof verifies. Across 2,000 AgentSafetyBench scenarios under 5fold stratified crossvalidation with out of fold scoring, NiyamAI achieves 88.8% F1 at a 1.0% false positive rate (bootstrap 95% CI [85.5%, 92.1%]), against 66.8% for Llama Prompt Guard 2, 46.2% for GPTOSSSafeguard, and 40.4% for NeMo Guardrails; McNemar's exact test confirms each margin at p < 0.0001. Proof generation adds 1.7 s per approved action, verification 51 ms, with an 18.6 KB proof verifiable by any third party without access to model parameters. We further subject NiyamAI's own enforcement mechanism to 18 adversarial vectors across six classes, disclosing two implementation vulnerabilities identified and remediated during development.
爆発範囲
エージェント コーディングは、手頃な価格とトークンの無駄という増大する問題に直面しています。結合されたコンテキストとコード チャネルを通じて受信プロンプトの到達範囲を推定する予測メモリ管理レイヤーである Blast Radius を紹介します。 NECROPHORESIS は、デッドコンテキストをそのままアーカイブすることで可逆的な削除を可能にし、一方、Recurring Dead Matter (RDM) は、繰り返し発生するトランスクリプトを識別して埋めます。ポーランドのコンテキスト空間上で可逆的なコンテキストの削除を定式化し、コンテキストのエントロピーを復活確率に関連付けながら、保持、再発、および削除の測定可能な基盤を提供します。 7 つの OpenAI モデル全体で、Blast Radius はトークン消費量を 17 ~ 26% 削減し、テストされたポリシーの中で最も低いオーバーフロー率を達成し、バイト正確な可逆性を維持しました。埋葬された遺体450体のうち、378体は再発死体であり、回収された遺体はゼロだった。 Blast Radius は HCRC の下で動作し、どのレコードを埋めるか、および受信プロンプトがコードベースにどこまで届くかを決定します。この取り組みは、大規模な言語モデルとエージェント コーディングをより再利用可能で持続可能なものにするという Algosophy のより広範な目標に貢献します。
原文 (English)
Blast Radius
Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled context and code channels. NECROPHORESIS enables reversible eviction by archiving dead context verbatim, while Recurring Dead Matter (RDM) identifies and buries repeatedly occurring transcripts. We formulate reversible context eviction over a Polish context space, providing a measurable foundation for retention, recurrence, and eviction while connecting context entropy to resurrection probability. Across seven OpenAI models, Blast Radius reduced token consumption by 17-26%, achieved the lowest overflow rate among tested policies, and remained byte exact reversible. Of 450 buried bodies, 378 were recurring dead matter and zero were recalled. Blast Radius operates beneath HCRC, determining which records to bury and how far an incoming prompt may reach into the codebase. This work contributes to the broader goal of Algosophy: making large language models and agentic coding more reusable and sustainable.
FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training
Open-ended language-model evaluation often substitutes another model or a small preference panel for a missing answer key. We introduce Fla…
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization
Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and th…
SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing
Hate speech research has moved from coarse-grained classification towards structured parsing, where systems jointly identify targets, suppo…
ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation
Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-lan…
Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents
Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decompositi…
From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation
Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: onc…
Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement
Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words, imagery, and prosody. However, many curren…
欠陥コード駆動のテストケース合成と高密度の報酬形成による堅牢なコード RL
検証可能な報酬からの強化学習 (RLVR) は、大規模言語モデル (LLM) のコード生成機能を強化するための極めて重要な手法として登場しました。ただし、コーディング実装における RLVR の有効性は、テスト ケースの包括性によって基本的に制限されます。これは、コード検証におけるテスト カバレッジが不十分なことが多くの場合誤検知を引き起こし、さらに報酬ハッキングやポリシーの劣化につながるためです。現在の自動生成手法の次善の品質に起因する報酬バイアスを軽減するために、我々は RobustTests フレームワークを提案します。このフレームワークは、「ほぼ正しい」欠陥コードを活用してモデルが潜在的な論理的不一致を正確に捕捉するようにガイドする欠陥コード駆動型のテスト ケース合成戦略を導入し、さらに検証エージェントと動作特徴クラスタリングを統合して、無効で冗長なテスト ケースの詳細なフィルタリングを容易にします。合成テスト ケースにおける固有の幻覚ノイズによって引き起こされる偽陰性に対処するために、RobustTests には合格率に基づいた段階的な高密度報酬関数も組み込まれており、きめ細かいフィードバックを通じてトレーニングの堅牢性が強化されています。このパイプラインを採用することで、CodeContest のテスト ケースを強化する高品質のデータセットを構築し、より広範囲の欠陥コード シナリオを網羅し、診断ユーティリティを大幅に強化します。実験結果は、CodeContests からの適度に困難な問題のサブセットをトレーニングに活用することにより、RobustTests を介した Qwen3-32B の RL 微調整がベースライン手法と比較して LiveCodeBench ベンチマークで絶対 3% のパフォーマンス向上を達成することを示しており、LLM のコード生成能力の向上における RobustTests フレームワークの有効性が確認されています。
原文 (English)
Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping
Reinforcement Learning from Verifiable Rewards (RLVR) is pivotal for enhancing LLM code generation, yet its efficacy is often hindered by insufficient test case coverage, leading to reward hacking and policy degradation. To address this, we propose RobustTests, a framework featuring a faulty-code-driven test case synthesis strategy. By leveraging "near-correct" faulty codes, RobustTests captures latent logical discrepancies and employs validator agents with behavioral feature clustering to filter invalid or redundant test cases. Additionally, a stepwise dense reward function based on pass rates is introduced to mitigate false negatives and enhance training robustness. Using this pipeline, we construct an augmented version of the CodeContests+ dataset with superior diagnostic utility. Experimental results show that RL fine-tuning of Qwen3-32B via RobustTests achieves a 3% absolute gain on LiveCodeBench, demonstrating its effectiveness in advancing LLM code generation proficiency. Codes and data are available at https://huggingface.co/datasets/sid6/RobustTests.
RePolicy: エージェント セーフガードでの安全ポリシー呼び出しのための強化学習
言語モデル エージェントを保護するには、コンテキスト依存の安全性ポリシーに基づいて完全な実行軌跡を評価する必要があります。既存のポリシーを意識した保護手段は主に、プロンプトまたは監視付きの微調整に依存しており、目に見えない軌道や変化するポリシー状況に適応する能力が制限されています。私たちは、強化学習を通じて安全ポリシーの呼び出しを学習するエージェント保護手段である RePolicy を提案します。エージェントの軌跡と動的なポリシー ライブラリが与えられると、RePolicy は該当するポリシーを呼び出し、その内容を使用してポリシーに基づいた理論的根拠と安全性の判断を生成します。教師あり初期化をサポートするために PolicyTraj-20K を構築し、続いて検証可能な報酬とポリシーコンテキストの摂動を備えた GRPO を構築します。 6 つのエージェント安全性ベンチマークにわたる実験では、RePolicy がさまざまなポリシー コンテキストの下で強力な全体的な安全性検出パフォーマンスと堅牢なポリシー呼び出しを達成していることが示されています。
原文 (English)
RePolicy: Reinforcement Learning for Safety-Policy Invocation in Agent Safeguards
Safeguarding language model agents requires assessing complete execution trajectories under context-dependent safety policies. Existing policy-aware safeguards mainly rely on prompting or supervised fine-tuning, limiting their ability to adapt to unseen trajectories and changing policy contexts. We propose RePolicy, an agent safeguard that learns safety-policy invocation through reinforcement learning. Given an agent trajectory and a dynamic policy library, RePolicy invokes the applicable policy and uses its content to produce a policy-grounded rationale and safety judgment. We construct PolicyTraj-20K to support supervised initialization, followed by GRPO with verifiable rewards and policy-context perturbation. Experiments across six agent safety benchmarks show that RePolicy achieves strong overall safety-detection performance and robust policy invocation under varying policy contexts.
状態から行動へ: 信頼性の高い複数回転工具の使用のための OODA ツール
信頼性の高いマルチターン ツールの使用には、エージェントが進化するタスクの状態を保存し、各アクションがその状態と一貫性を保っていることを確認する必要があります。ただし、直接関数呼び出しと ReAct スタイルのポリシーは、同じ自己回帰軌道内で状態の追跡とアクションの生成を学習します。この結合により状態と動作の競合が生じます。次の呼び出しを生成するというプレッシャーにより、対話の以前に蓄積された情報が上書きされたり、無視されたりする可能性があります。ボイドの観察、指向、決定、行為のサイクルに触発され、状態の維持を行動の実現から分離することでこの競争を緩和するように設計された型付き閉ループ ポリシーである OODA ツールを紹介します。 OODA ツールは、インタラクション履歴から直接アクションを生成するのではなく、コントローラーがチェックした中間状態を介して各決定をルーティングし、最終出力が現在のタスク状態に確実に固定されるようにします。具体的には、Observe はタスクの状態を再構築し、Orient は実行が正当であるかどうかを決定し、Decide は許容可能なアクション構造を形成し、Act は外部出力を実現します。マルチターン、マルチツール、不完全情報設定全体で 0.6B から 14B の範囲の Qwen3 モデルを使用して、直接関数呼び出しと ReAct ポリシーに対して OODA ツールを評価します。 OODA-Tool は、モデル サイズ全体でタスクの成功率を一貫して向上させます。小さいモデルや、アクションがターン全体で蓄積された情報や以前のツールの結果に大きく依存するタスクでは、より大きな効果が得られます。制御されたバリアント、段階レベルのアブレーション、および移植の評価は、これらの改善の堅牢性をさらに実証します。
原文 (English)
From State to Action: OODA-Tool for Reliable Multi-Turn Tool Use
Reliable multi-turn tool use requires an agent to preserve an evolving task state and ensure that each action remains consistent with it. However, direct function-calling and ReAct-style policies learn state tracking and action generation within the same autoregressive trajectory. This coupling creates state-action competition: the pressure to produce the next call can overwrite or ignore information accumulated earlier in the interaction. Inspired by Boyd's Observe-Orient-Decide-Act cycle, we introduce OODA-Tool, a typed closed-loop policy designed to mitigate this competition by separating state preservation from action realization. Rather than generating an action directly from the interaction history, OODA-Tool routes each decision through controller-checked intermediate states, ensuring that the final output remains grounded in the current task state. Specifically, Observe reconstructs the task state, Orient determines whether execution is warranted, Decide forms an admissible action structure, and Act realizes the external output. We evaluate OODA-Tool against direct function-calling and ReAct policies using Qwen3 models ranging from 0.6B to 14B across multi-turn, multi-tool, and incomplete-information settings. OODA-Tool consistently improves task success across model sizes, with larger gains on smaller models and on tasks whose actions depend strongly on information accumulated across turns and prior tool results. Controlled variants, stage-level ablations, and transfer evaluations further demonstrate the robustness of these improvements.
Account Consistency from Gameplay Traces: Same-Player Verification in Counter-Strike 2
In competitive first-person shooter (FPS) games such as Counter-Strike 2 (CS2), account-integrity review often asks whether an account's re…
Recurrent Reinforcement Learning with Memoroids
Memory models such as Recurrent Neural Networks (RNNs) and Transformers address Partially Observable Markov Decision Processes (POMDPs) by…
CollaFuse: Collaborative Diffusion Models
In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic…
The BS-meter: Detecting Politics and Labour through ChatGPT's Language
What can we learn about language from studying how it is used by ChatGPT and other large language model (LLM)-based chatbots? In this paper…
Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
Dense retrieval is a crucial task in Information Retrieval (IR), serving as the basis for downstream tasks such as re-ranking and augmentin…
Communication styles and reader preferences of LLM- and human-authored COVID-19 information explanations: a case study
With the wide adoption of large language models (LLMs) in information assistance, it is essential to examine their alignment with human com…
Temporally-Grounded Language Generation: Towards Real-Time Vision-Language Models
Vision-language models (VLMs) have shown remarkable progress in offline tasks such as image captioning and video question answering. Howeve…
HybridProver: Augmenting Theorem Proving with LLM-Driven Proof Synthesis and Refinement
Formal methods play a crucial role in ensuring the reliability of critical systems through rigorous mathematical verification. However, the…
From Accuracy to Robustness: A Study of Rule- and Model-based Verifiers in Mathematical Reasoning
Trustworthy verifiers are essential for the success of reinforcement learning with verifiable reward (RLVR), which is the core methodology…
Refine-POI: Reinforcement Fine-Tuned Large Language Models for Next Point-of-Interest Recommendation
Advancing large language models (LLMs) for the next point-of-interest (POI) recommendation task faces two fundamental challenges: (i) altho…
Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics
Preference-based Reinforcement Learning (PbRL) provides a promising alternative to heuristic reward design in complex robotic environments.…
Distinct Profiles of Run-to-Run Score Reliability and Expert-Panel Alignment Across Four LLM Evaluators of Simulated Japanese-Language AI-to-AI Counseling
Large language models (LLMs) increasingly evaluate generated dialogue, but repeatable scores do not necessarily align with professional jud…
AirLLM: Diffusion Policy-based Adaptive LoRA for Remote Fine-Tuning of LLM over the Air
Operating Large Language Models (LLMs) on edge devices is increasingly challenged by limited communication bandwidth and strained computati…
Toward a New Science of AI as Cognitive Infrastructure
Contemporary human-AI interaction research overlooks how AI systems fundamentally reshape human cognition pre-consciously, a critical blind…
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about fa…
Recurrence Meets Transformers for Universal Multimodal Retrieval
With the rapid advancement of multimodal retrieval and its application in LLMs and multimodal LLMs, increasingly complex retrieval tasks ha…
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
Mathematical reasoning is a key capability for vision-language models (VLMs), yet current benchmarks mainly evaluate text-based or explicit…
Egosurg: Arbitrary view synthesis for egocentric replay of operating room workflows from ambient cameras
Observing surgical practice has historically relied on fixed vantage points or recollections, leaving the egocentric perspectives that shap…
MCCE: A Framework for Multi-LLM Collaborative Search in Discrete Spaces with Similarity-Filtered Preference Learning
Multi-objective discrete optimization problems, such as molecular design, pose significant challenges due to their vast and unstructured co…
LLM-Specific Utility for Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) is typically optimized for topical relevance, yet its success ultimately depends on whether retrieved…
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation
Distilling the tool-use capabilities of large language models (LLMs) into small language models (SLMs) is essential for their practical app…
拡散モデルの原理
この本は、拡散モデルの開発を導いてきた核となる原理を紹介し、その起源をたどり、共有された数学的考え方からどのように多様な定式化が生まれるかを示します。拡散モデリングは、データを徐々にノイズに破壊する順方向プロセスを定義することから始まり、一連の中間分布を通じてデータ分布を単純な事前分布にリンクします。目標は、同じ中間体を復元しながらノイズをデータに戻す逆のプロセスを学習することです。 3 つの相補的なビューについて説明します。変分オートエンコーダからインスピレーションを得た変分ビューでは、拡散を段階的にノイズを除去する方法を学習するものとみなします。エネルギーベースのモデリングに基づいたスコアベースのビューは、進化するデータ分布の勾配を学習し、サンプルをより可能性の高い領域に向けて移動する方法を示します。フローベースのビューは、フローの正規化に関連し、学習された速度フィールドの下でサンプルをノイズからデータに移動する滑らかなパスに従うものとして生成を扱います。これらの視点は、共通のバックボーン、つまり、データに先立って単純な流れを転送する時間依存の速度場を共有しています。サンプリングは、連続的な軌跡に沿ってノイズをデータに展開する微分方程式を解くことになります。この基礎に基づいて、この本では、制御可能な発電、効率的な数値ソルバー、および任意の時間間の直接マッピングを学習する拡散を動機とするフローマップ モデルのためのガイダンスについて説明します。基本的な深層学習の知識を持つ読者に、拡散モデルの概念的かつ数学的に根拠のある理解を提供します。
原文 (English)
The Principles of Diffusion Models
This book presents the core principles that have guided the development of diffusion models, tracing their origins and showing how diverse formulations arise from shared mathematical ideas. Diffusion modeling starts by defining a forward process that gradually corrupts data into noise, linking the data distribution to a simple prior through a continuum of intermediate distributions. The goal is to learn a reverse process that transforms noise back into data while recovering the same intermediates. We describe three complementary views. The variational view, inspired by variational autoencoders, sees diffusion as learning to remove noise step by step. The score-based view, rooted in energy-based modeling, learns the gradient of the evolving data distribution, indicating how to nudge samples toward more likely regions. The flow-based view, related to normalizing flows, treats generation as following a smooth path that moves samples from noise to data under a learned velocity field. These perspectives share a common backbone: a time-dependent velocity field whose flow transports a simple prior to the data. Sampling then amounts to solving a differential equation that evolves noise into data along a continuous trajectory. On this foundation, the book discusses guidance for controllable generation, efficient numerical solvers, and diffusion-motivated flow-map models that learn direct mappings between arbitrary times. It provides a conceptual and mathematically grounded understanding of diffusion models for readers with basic deep-learning knowledge. Supplementary materials for the book are available at the book website: https://the-principles-of-diffusion-models.github.io/
What the "Spotless" Mind Remembers: How Knowledge Entanglement Shapes What Leaks After Unlearning in LLMs
Unlearning in large language models (LLMs) is usually evaluated as whether an "unlearned" fact can be recovered. We instead ask whether a f…
Multivariate Diffusion Transformer with Decoupled Attention for High-Fidelity Mask-Text Collaborative Facial Generation
While significant progress has been achieved in multimodal facial generation using semantic masks and textual descriptions, conventional fe…
Diagnosing Conformal Prediction Failures Under Distribution Shift: A COVID-19 Case Study
Conformal prediction provides distribution-free coverage guarantees, but these degrade under distribution shift - and practitioners lack to…
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
Video-language models (VLMs) achieve strong multimodal understanding but remain prone to hallucinations, especially when reasoning about ac…
Subspace Alignment for Vision-Language Model Test-time Adaptation
Vision-language models (VLMs), despite their extraordinary zero-shot capabilities, are vulnerable to distribution shifts. Test-time adaptat…
LoRA as Oracle
Practitioners increasingly deploy neural networks they did not train, and must audit them after the fact for hidden backdoors, without the…
Beyond Factual QA: Mentorship-Oriented Question Answering over Long-Form Multilingual Content
Question answering systems are typically evaluated on factual correctness, yet many real-world applications-such as education and career gu…
A Very Big Video Reasoning Suite
Rapid progress in video models has largely focused on visual quality, leaving their reasoning capabilities underexplored. Video reasoning g…
SynthCharge: An Electric Vehicle Routing Instance Generator with Feasibility Screening to Enable Learning-Based Optimization and Benchmarking
The electric vehicle routing problem with time windows (EVRPTW) extends the classical VRPTW by introducing battery capacity constraints and…
周波数が重要: プルーニングと量子化のためのモデルに依存しない高速データキュレーション
トレーニング後のモデル圧縮は、大規模言語モデル (LLM) のパフォーマンスを維持しながら移植性を高めるために不可欠です。いくつかの圧縮手法が提案されていますが、圧縮されたモデル構成を見つけるために最適なデータのセット (いわゆる \emph{校正データ}) を選択することにはあまり重点が置かれていません。キャリブレーション データの選択は、タスク内およびタスク間の両方でモデルの機能を維持するための重要なステップです。この研究では、モデル固有の信号ではなく固有のデータ プロパティを分析することで、プルーニングと量子化の両方に対応する高性能のキャリブレーション セットを特定するという課題に取り組みます。 \texttt{\textbf{ZipCal}} を導入します。これは、Zipfian べき乗則に基づいて語彙の多様性を最大化する、モデルに依存しないデータ キュレーション戦略です。実験では、さまざまな枝刈りベンチマークにわたって、私たちの方法が標準の一様ランダムサンプリングよりも一貫して優れていることが実証されています。特に、ダウンストリームのパフォーマンスの点でも、モデルの複雑さに依存する最先端の手法と同等のパフォーマンスを発揮します。後者は大規模なモデルやデータセットでは法外に高価になりますが、\texttt{\textbf{ZipCal}} は扱いやすい線形の複雑さのため、平均 $\sim$240$\times$ 高速です\footnote{コードと実験は https://github.com/FrancescoMonaco/ZipCal で公開しています。}
原文 (English)
Frequency Matters: Fast Model-Agnostic Data Curation for Pruning and Quantization
Post-training model compression is essential for enhancing the portability of Large Language Models (LLMs) while preserving their performance. While several compression approaches have been proposed, less emphasis has been placed on selecting the most suitable set of data (the so-called \emph{calibration data}) for finding the compressed model configuration. The choice of calibration data is a critical step in preserving model capabilities both intra- and inter-tasks. In this work, we address the challenge of identifying high-performance calibration sets for both pruning and quantization by analyzing intrinsic data properties rather than model-specific signals. We introduce ZipCal, a model-agnostic data curation strategy that maximizes lexical diversity based on Zipfian power laws. Experiments demonstrate that our method outperforms standard uniform random sampling across various pruning benchmarks. Notably, it also performs on par, in terms of downstream performance, with a state-of-the-art method that relies on model perplexity. The latter becomes prohibitively expensive for large-scale models and datasets, while ZipCal is on average $\sim$240$\times$ faster due to its tractable linear complexity. We make the code and the experiments available at https://github.com/FrancescoMonaco/ZipCal.
How LLMs Distort Our Written Language
Large language models (LLMs) are used by over a billion people globally, most often to assist with writing. In this work, we demonstrate th…
High-Fidelity Face Content Recovery via Tamper-Resilient Versatile Watermarking
The proliferation of AIGC-driven face manipulation and deepfakes poses severe threats to media provenance, integrity, and copyright protect…
Grounded Token Initialization for New Vocabulary in LMs for Generative Recommendation
Language models (LMs) are increasingly extended with new learnable vocabulary tokens for domain-specific tasks, such as Semantic-ID tokens…
A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting
Text-driven motion editing and intra-structural retargeting, where skeletons share topology but may differ in bone lengths and rest pose, a…
CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations
The rapid expansion of gaming industry requires advanced recommender systems tailored to its dynamic landscape. Existing Graph Neural Netwo…
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alter…
MOMO: A framework for seamless physical, verbal, and graphical robot skill learning and adaptation
Industrial robot applications require increasingly flexible systems that non-expert users can easily adapt for varying tasks and environmen…
MambaCSP: Hybrid-Attention State Space Models for Hardware-Efficient Channel State Prediction
Recent works have demonstrated that attention-based transformer and large language model (LLM) architectures can achieve strong channel sta…
Cartan flow matching
We introduce Cartan flow matching, a general framework for training flow matching models on Riemannian symmetric spaces, i.e. Riemannian ma…
MedFabric: Gold Evidence Hides the Difficulty of Word-Level Medical Fabrication Detection
Large language models fabricate in medicine, producing fluent statements that are factually wrong, so reliable fabrication detection is a p…
No Plan, Yet Human: A Reactive Robotics Model Predicts Human Planning Failures on a Clinical Task
Understanding why some sequential planning problems are harder than others requires models that go beyond average performance. They should…
HINT-SD: Targeted Hindsight Self-Distillation for Long-Horizon Agents
Training long-horizon LLM agents with reinforcement learning is challenging because sparse outcome rewards reveal whether a task succeeds,…
GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
Mixed-precision quantization improves the budget--accuracy trade-off for large language models (LLMs) by allocating more bits to sensitive…
A Comprehensive Comparison of Deep Learning Architectures for COVID-19 Classification on CT & X-ray Imagery
COVID-19 was a significant challenge that led to the loss of numerous lives daily. Not only a certain country was involved in this outbreak…
Pixel Wised Lesion Prediction on COVID-19 CT Imagery: A Comparative Analysis of Automated Image Segmentation Architectures
In recent years, there has been a notable increase in the level of attention that is given to algorithms based on deep learning in the cont…
MIMO: 単一言語の目的による多言語情報検索
多言語情報検索 (MLIR) は、混合言語コーパス内でクエリや関連文書がさまざまな言語で表示される実際の検索環境を反映しています。ただし、既存の埋め込みモデルは主に多言語検索用に最適化されており、MLIR 設定ではパフォーマンスが低下することがよくあります。さらに、従来の対照学習を MLIR に直接適用すると、言語のクラスタリングが悪化して、言語間の整合性と埋め込みの均一性の間のトレードオフが明らかになる可能性があります。これらの制限に対処するために、私たちは MIMO: Multilingual Information Retrieval via Monolingual Objectives を提案します。これは、高性能教師モデルからの安定した英語意味空間をアンカーとして使用する 2 段階のフレームワークです。 MIMO は、最初に知識の蒸留を通じて学生モデルの言語間連携を初期化し、次に蒸留と言語間対比学習を共同で最適化して、連携を維持しながら検索識別を改善します。広範な実験により、MIMO はさまざまな MLIR およびマルチモノリンガル ベンチマークにわたって、既存のクロスリンガル トレーニング ベースラインを常に上回るパフォーマンスを示しています。 MIMO は、同様またはより大きなパラメーター スケールの既製モデルとの競争力も維持します。さらに、我々の言語横断的な整列均一性分析は、2 つの損失要素の異なる役割を明らかにし、それらの組み合わせにより整列と均一性の間に好ましいトレードオフが生じることを示しています。
原文 (English)
MIMO: Multilingual Information Retrieval via Monolingual Objectives
Multilingual Information Retrieval (MLIR) reflects real-world search environments in which queries and relevant documents may appear in different languages within a mixed-language corpus. However, existing embedding models are primarily optimized for Multi-Monolingual retrieval and their performance often degrades in MLIR settings. Moreover, directly applying conventional contrastive learning to MLIR can exacerbate language clustering and expose a trade-off between cross-lingual alignment and embedding uniformity. To address these limitations, we propose MIMO: Multilingual Information Retrieval via Monolingual Objectives, a two-stage framework that uses a stable English semantic space from a high-performing teacher model as an anchor. MIMO first initializes the student model's cross-lingual alignment through knowledge distillation, and then jointly optimizes distillation and cross-lingual contrastive learning to improve retrieval discrimination while preserving alignment. Extensive experiments show that MIMO consistently outperforms existing cross-lingual training baselines across various MLIR and Multi-Monolingual benchmarks. MIMO also remains competitive with off-the-shelf models of similar or larger parameter scales. Furthermore, our cross-lingual Alignment-Uniformity analysis clarifies the distinct roles of the two loss components and shows that their combination yields a favorable trade-off between alignment and uniformity.
Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains
Tool-augmented multimodal agents show strong benchmark gains, often taken as evidence that agents have learned to use tools. We argue that…
LoopMoE: 言語モデリングの専門家混合による反復計算の統合
専門家混合 (MoE) およびループ アーキテクチャは、パラメーター容量と有効深さという 2 つの直交軸に沿ってモデルをスケールします。ただし、主流のループ アーキテクチャは、パラメーター数とトークンごとの FLOP を結合する高密度のバックボーンに依存しているため、一致した予算の下での反復計算の影響を分離することができません。この目的を達成するために、2 つの設計を通じてスパース ルーティングと反復的な重み共有計算を統合するループ MoE 言語モデルである LoopMoE を紹介します。 1 つ目は IterAdaLN で、反復インデックスとトークンごとの隠れ状態を組み合わせて条件付けされた変調信号を介して重み共有対称性を解決します。 2 つ目は、適切に調整された非ループ参照のアテンション対 FFN アクティブ パラメータの比率を回復する容量バランシング戦略です。これらの設計を組み合わせることで、同一の合計パラメーター、トークンごとの FLOP、およびアクティブなサブレイヤー比の下で、バニラ MoE に対するループ MoE の厳密に制御された最初の直接評価が可能になります。 3B スケールでは、LoopMoE は 9 つの下流ベンチマークのうち 8 つで Vanilla MoE を上回り、平均改善率は 1 ポイントを超えています。 9B スケールでは、LoopMoE が引き続き同等の Vanilla MoE を上回り、アーキテクチャ上の利点がより大きなスケールでも持続することを示しています。私たちの研究は、スパース性と再帰性の制御された統合を確立し、ループ言語モデルの有望な方向性を示唆しています。
原文 (English)
LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling
Mixture-of-Experts (MoE) and looped architectures scale models along two orthogonal axes, namely parameter capacity and effective depth. However, mainstream looped architectures rely on dense backbones that couple parameter count with per-token FLOPs, which makes it impossible to isolate the effect of iterative computation under matched budgets. To this end, we present LoopMoE, a looped MoE language model that integrates sparse routing with iterative weight-shared computation through two designs. The first is IterAdaLN, which resolves weight-sharing symmetry via a modulation signal jointly conditioned on the iteration index and the per-token hidden state. The second is a capacity-balancing strategy that recovers the attention-to-FFN active parameter ratio of well-tuned non-looped references. Together, these designs enable the first strictly controlled, head-to-head evaluation of a looped MoE against a Vanilla MoE under identical total parameters, per-token FLOPs, and active sublayer ratios. Across nine downstream benchmarks, LoopMoE's average improvement over its matched vanilla MoE increases from over 1 point at the 3B scale to approximately 3 points at the 9B scale. These results provide initial evidence that the benefits of iterative sparse computation may strengthen with scale, positioning LoopMoE as a promising architecture for scalable looped language models.
Summarization is Not Dead Yet
The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written reference…
Harnessing the Collective Intelligence of AI Agents in the Wild for New Discoveries
Scientific discovery is often a collective process: researchers share partial results, inspect failed attempts, and build on each other's i…
Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot
Large language models deployed as commercial APIs are vulnerable to model extraction attacks, while existing defenses either act too late o…
SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information Retrieval
With the rapid expansion of massive multilingual corpora, Multilingual Information Retrieval (MLIR) has emerged as a critical technology fo…
PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement
Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising p…
Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remai…
Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data
Coding agents can now be left alone to improve software against a score. In this pattern--recently popularized as "autoresearch"--the agent…
Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating
Clinical decision support degrades as treatment protocols evolve, but the obstacle to updating a deployed model is governance as much as ac…
ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour
Fully homomorphic encryption (FHE) lets a server run inference on encrypted data with strong privacy guarantees, but running a Transformer…
TriShieldRAG: 3 Rings, One Blind Spot in Layered Defenses for Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) grounds LLM answers in query-time retrieved documents, so reliability depends on what the retriever re…
REPREC: 表現駆動型パラメータ効率的な推奨システム
大規模言語モデル (LLM) は、自然言語タスクとして定式化することにより、順次レコメンデーションに適用されています。これまでの研究では、入力調整または LLM 微調整を通じて協調的かつ順次的な信号を組み込むことで、パーソナライゼーションが向上しました。ただし、既存のアプローチは、LLM の微調整、追加のアーキテクチャ モジュール、表現の蒸留、または長いインタラクション履歴にわたる項目レベルの調整の 1 つ以上に依存することが多く、トレーニングの複雑さと導入コストが増大します。私たちは、軽量のユーザー表現の調整を通じて LLM ベースの逐次レコメンデーションを再定式化する軽量フレームワークである REPREC を提案します。 REPREC は、フリーズされた LLM を調整する軽量の MLP インジェクターを介して、フリーズされたシーケンシャル エンコーダーから埋め込まれた固定サイズのユーザーを学習済みのソフト トークンの小さなセットにマッピングします。これにより、インジェクターのみをトレーニングしながら、両方の事前トレーニングされたバックボーンは変更されません。私たちは複数のベンチマーク データセットに対して徹底的な実験を実施し、REPREC がさまざまな事前トレーニング済みシーケンシャル エンコーダーや LLM バックボーンとの互換性を維持しながら一貫して LoRA を上回るパフォーマンスを示し、事前トレーニング済みコンポーネントを変更することなく、モジュール式で本番環境に適したレコメンデーション パイプラインを実現できることを実証しました。この利点は、すべてのデータセットにわたってカジュアル ユーザーとコア ユーザーで特に顕著であり、データ量が少ない状況での REPREC の有効性が強調されています。最後に、短いプロンプト履歴でトレーニングし、より長いコンテキストで評価した場合、REPREC は LoRA のパフォーマンスの 85 ~ 100% を維持しながら、エポックごとのトレーニング時間を平均 1.51 分の 1 に短縮します。これは、推奨品質と実稼働デプロイメントの計算効率の間の効果的なバランスを示しています。コードは https://github.com/phdbotcode/REPREC で入手できます。
原文 (English)
REPREC: Representation Driven Parameter-Efficient Recommendation System
Large language models (LLMs) have been applied to sequential recommendation by formulating it as a natural language task. Previous work has improved personalization by incorporating collaborative and sequential signals through input conditioning or LLM fine-tuning. However, existing approaches often rely on one or more of the following: LLM fine-tuning, additional architectural modules, representation distillation, or item-level conditioning over long interaction histories, increasing training complexity and deployment cost. We propose REPREC, a lightweight framework that reformulates LLM-based sequential recommendation through lightweight user representation alignment. REPREC maps a fixed-size user embedding from a frozen sequential encoder into a small set of learned soft tokens through a lightweight MLP injector that conditions a frozen LLM, leaving both pretrained backbones unchanged while training only the injector. We conducted exhaustive experiments on multiple benchmark datasets and demonstrate that REPREC consistently outperforms LoRA while remaining compatible with different pretrained sequential encoders and LLM backbones, enabling a modular and production-friendly recommendation pipeline without modifying either pretrained component. The gains are particularly pronounced for casual and core users across all datasets, highlighting REPREC's effectiveness in low-data regimes. Finally, when trained on short prompt histories and evaluated with longer contexts, REPREC maintains 85-100% of LoRA's performance while reducing per-epoch training time by an average of 1.51X, demonstrating an effective balance between recommendation quality and computational efficiency for production deployment. The code is available at https://github.com/phdbotcode/REPREC
Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities
Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evalu…
When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
Multi-agent LLM systems relay key-value caches instead of text and credit their gains to exchanged "latent thoughts". That credit is a clai…
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.…
ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB
Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal res…
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deterior…
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolat…
A 12-CNOT Double Qubit Excitation Gate
Effective implementation of high-level quantum gates is essential for practical quantum computing. In this work, we presented, to the best…
M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation
Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to ex…
MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification
Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and…
Pre-training Visual Dexterity in Simulation
Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datase…
X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization
Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. How…
Complexity Induction: Compositional Generalization via Structured Training Distortion
We demonstrate that structured distortion of training data - which we term complexity induction - can induce compositional generalization i…
Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules
Despite recent advances in molecular foundation models, several limitations remain, such as chemically invalid augmentations, modality coll…
Language Chain in Alignment: Cross-lingual Ranking Preference Optimization
The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal perf…
The Limits of Automatic Evaluation of Creativity in Large Language Models
Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity,…
EXAM$^2$: $\underline{Ex}tending$ $\underline{A}udio$ $Understanding$ $in$ $\underline{M}ultilingual$ $and$ $\underline{M}ultimodal$ $Analysis$
Recent large audio language models (LALMs) have achieved impressive progress in audio understanding. However, existing evaluations remain l…
When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk
LLMs are being used increasingly to measure aspects of student discourse (e.g. talk moves, collaboration, equity of voice) at scale. Typica…
CAT-GS: Balanced Multimodal Learning via Calibrated Gating and Fusion Surgery
End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that…
Unsupervised Post-Training of Foundation Models: A Survey
Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupe…
DataKernelBench: Can LLMs Optimize Database Queries on GPUs?
GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM…
MACGen: Toward Functionally Correct and Secure Code Generation via Multi-Agent Collaboration
Despite their strong ability to generate code, large language models often fail to produce secure code, as their outputs frequently contain…
4DStreamCtrl: Interactive Video Generation with Online 4D Control
Generative video models now synthesize footage nearly indistinguishable from reality. Their promise as interactive tools hinges on fine-gra…
When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory
An agent that inherits a consolidated memory may inherit a constraint that was true when written and has since been withdrawn by a newer au…
Learning New Facts with QLoRA: An Acquisition-Retention Frontier
Parameter-efficient fine-tuning is often assumed to preserve pretrained capabilities because it updates only a small number of parameters.…