Skip to the content.

AIニュース 2026-08-01

自動生成: 2026-08-01 12:27 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Building abundant intelligenceOpenAI

    A full-stack approach to making advanced AI more capable, more afford…

  2. Advancing responsible AI across EuropeOpenAI

    OpenAI shares how its safety, security, transparency, and provenance…

  3. Univé builds an AI-ready workforceOpenAI

    See how Univé built an AI-ready workforce with ChatGPT Enterprise by…

  4. Google nixes its Earth AI feature one day after launch, amid criticism it would spread misinformationTechCrunch AI

    A tool that allowed anyone to generate fake AI-generated imagery and…

  5. 研究者10万人にOpenAI「最上位モデル」無料提供へ 日本でも東大、京大など15大学が対象ITmedia AI+

    OpenAIが学術研究者10万人に「ChatGPT」の最上位モデルを無料提供するプログラムを発表した。今夏に1万人から提供を始め、2027…

  6. OpenAI reportedly finds evidence that more of its agents ran amokTechCrunch AI

    OpenAI has reportedly found evidence of additional agent misbehavior…

  7. India is starting to pay for apps, not just download themTechCrunch AI

    India's app market generated a record $345 million in Q2.

トピック別件数

日本語メディア5件

ITmedia AI+ (日本語)

18:45 JSTその他

キオクシア、株価3分の1急落は「絶好のタイミング」 過去最高益と8000億円自社株買いで示す自信

キオクシアHDは、2026年4~6月のNon-GAAP営業利益が1兆3262億円で、直近1年間の営業利益を上回ったと説明した。7~9月の売上収益は、2兆3900億円を見込む。経営幹部からは力強い言葉が続いた。

17:20 JSTLLM/生成AI

「9カ月かかる作業を3日に短縮」 IBM、レガシー刷新ワークフローをIBM Bobに追加

生成AIによるコード生成が広がる一方で、レガシー環境の刷新はAIの適用が難しい領域だ。IBMは、IBM Bobにレガシー環境向けの専用パッケージを追加することで、同領域における自動化に踏み込む。

17:16 JSTその他

キオクシアQ1、純利益が前年同期比4500%増 株式分割・自社株買いも

キオクシアホールディングス(HD)が7月31日に公開した2027年3月期第1四半期(26年4月1日?6月30日)連結決算は、売上収益が1兆7671億1700万円(前年同期比415.5%増)、営業利益が1兆2700億1700万円(同2728.6%増)、純利益が8421億6500万…

16:33 JSTハードウェア/半導体

キオクシアQ1決算、純利益は前年比4500%増 AIデータセンター向け需要がけん引

半導体大手のキオクシアホールディングスは、2027年3月期第1四半期決算(26年4月1日?6月30日、国際会計基準)の純利益が8421億6500万円で、前年同期比4506%増だったと発表した。

15:57 JSTLLM/生成AI研究/論文OpenAIGPT / ChatGPT

研究者10万人にOpenAI「最上位モデル」無料提供へ 日本でも東大、京大など15大学が対象

OpenAIが学術研究者10万人に「ChatGPT」の最上位モデルを無料提供するプログラムを発表した。今夏に1万人から提供を始め、2027年までに10万人規模へ拡大する。日本でも東京大学、京都大学、東京科学大学、早稲田大学、慶應義塾大学など15の大学が対象機関となっている。

海外メディア8件

TechCrunch AI (英語)

07:47 JSTLLM/生成AIエージェントOpenAI

OpenAI reportedly finds evidence that more of its agents ran amok

OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.

06:07 JSTその他

India is starting to pay for apps, not just download them

India's app market generated a record $345 million in Q2.

04:47 JSTその他Google

Google nixes its Earth AI feature one day after launch, amid criticism it would spread misinformation

A tool that allowed anyone to generate fake AI-generated imagery and superimpose it over real Google Earth maps quickly spurred backlash.

01:49 JSTその他

Snapchat no longer rewards fully AI-generated Spotlight content

Snapchat has adjusted its recommendation systems to ensure that only videos created by real people are eligible for Spotlight recommendatio…

01:08 JSTその他

Siri AI could come with a paywall for power users

Apple CEO Tim Cook envisions users being able to buy more compute for Siri AI via Apple's existing iCloud+ subscriptions.

00:16 JSTその他

SpaceX won’t remove all of xAI’s unpermitted turbines for another year

SpaceX is building a new power plant for xAI's Colossus data centers, but it won't remove existing, unpermitted turbines for many more mont…

23:47 JSTビジネス/資金調達

Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human

The startup is building voice models designed to make AI phone calls pass the Turing test.

23:00 JSTLLM/生成AIOpenAI2件の関連記事

AI labs want to pump the brakes, but Amazon and SpaceX are still blasting off

After years of pushing full speed ahead on AI, OpenAI CEO Sam Altman says maybe it’s time for the AI industry to “pace” itself. The comment…

出典:TechCrunch AITechCrunch AI
公式ブログ3件

OpenAI (英語)

00:00 JSTLLM/生成AIOpenAI

Advancing responsible AI across Europe

OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will c…

00:00 JSTその他

Building abundant intelligence

A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.

16:00 JSTLLM/生成AIGPT / ChatGPT

Univé builds an AI-ready workforce

See how Univé built an AI-ready workforce with ChatGPT Enterprise by combining leadership, responsible governance, and employee-led innovat…

論文282件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AI

推論パフォーマンスの起源を探る: RL と SFT の微調整モデルにおける数学的問題解決の表現品質

強化学習 (RL) によってトレーニングされた大規模推論モデルは、数学的推論タスクにおいて教師あり微調整 (SFT) モデルよりも優れたパフォーマンスを発揮することがますます明らかになってきています。しかし、この利点のメカニズムの根拠は依然として不明です。したがって、どのような内部表現の違いが RL モデルの優れたパフォーマンスを可能にするのかを尋ねます。私たちの研究は、2 つの収束する証拠を示しています。まず、層ごとの隠れ状態で訓練された線形プローブは、RL モデルが SFT モデルと比較して解答の正しさを予測する際に高い精度を達成する傾向があることを明らかにし、より線形に分離可能で構造化された表現を示しています。第 2 に、平均アブレーション研究では、RL モデルは、より深い層ほど重要性が高まる階層アーキテクチャを開発するのに対し、SFT モデルは重要性を層全体に均一に分散することが示されています。これらの発見を総合すると、RL トレーニングによって、モデルが推論問題を表現および処理する方法を根本的に再構築することがわかります。最後に、適応的なコンピューティング割り当てを評価するために、複数の問題にわたってサンプリングを繰り返した場合のトークン数の変動を分析します。一部の RL 調整モデルでは、対応する SFT モデルよりも高い変動が観察されていますが、他のモデルでは強い一貫性が見られ、トークンの割り当ては RL 対 SFT 単独よりもトレーニング パイプライン全体に依存する可能性が高いことが示唆されています。このトークン割り当ての変動は、ポリシーに関するもっともらしい推論の広がりを明らかにし、安定したポリシーを示すモデルと、決定が不十分で潜在的に特定できないソリューションの動作を示すモデルを強調していると考えています。

原文 (English)

Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models

Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage remains unclear. We therefore ask, what internal representational differences enable RL models' superior performance? Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations. Second, mean ablation studies show that RL models develop a hierarchical architecture where deeper layers become progressively more critical, whereas SFT models distribute importance uniformly across layers. Together, these findings demonstrate that RL training fundamentally restructures how models represent and process reasoning problems. Finally, we analyze token-count variability under repeated sampling across problems to assess adaptive compute allocation. While we observe higher variability in some RL-tuned models than in their SFT counterparts, we see strong consistency in others, suggesting that token allocation may depend more on the overall training pipeline than on RL versus SFT alone. We believe this token-allocation variability reveals the spread of plausible on-policy reasoning, highlighting which models exhibit stable policies versus those that are under-determined, potentially non-identifiable solution behaviour.

13:00 JSTLLM/生成AIエージェント

さらに欺瞞: 混合動機 LLM マルチエージェント システムにおける客観的な不整合

大規模言語モデル (LLM) を利用したマルチエージェント システムは、動機が混在する環境に導入されることが増えています。この環境では、エージェントは、矛盾する目的や隠れた目的による非対称情報や戦略的欺瞞の下で動作します。このような状況では、集団の目標との不一致が中心的な懸念事項になります。我々は、社会的推理ゲームである人狼を使用して、割り当てられた役割を維持しながら単一エージェントの目的を変更する、目的の不整合を評価するための新しいフレームワークを提案します。 4 つの異なるモデル ファミリーとサイズ、4 つのプレイヤーの役割、および 3 つの目的の定式化からなる LLM にわたって、エージェントの内部推論と彼らの公共のチープトーク行動 (つまり、エージェントのユーティリティに直接影響を及ぼさないコストのかからない拘束力のないコミュニケーション) の二重分析を導入し、ゲーム結果の分析で補完します。私たちの結果は、客観的な不整合が本質的に敵対的な環境での成果を損ない、その影響が非対称な情報と特殊な役割によって悪化することを示しています。侵害されたエージェントは一貫して、明確な目的に依存した推論戦略を開発しますが、これらの適応は公の行動ではほとんど目に見えないままです。より広範に、私たちの調査結果は、わずかな目標の不一致でさえ集団の意思決定に重大な影響を与える可能性があることを示唆しており、LLM ベースのマルチエージェント システムに対する効果的な緩和戦略の必要性を強調しています。

原文 (English)

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In these settings, misalignment with collective goals becomes a central concern. We propose a novel framework for evaluating objective misalignment using the social deduction game Werewolf, modifying the objective of a single agent while preserving its assigned role. Across LLMs from four different model families and sizes, four player roles, and three objective formulations, we introduce a dual analysis of the agents' internal reasoning and their public cheap-talk behavior (i.e costless, non-binding communication that does not directly affect the agents' utilities), complemented by an analysis of game outcomes. Our results show that objective misalignment undermines outcomes in inherently adversarial environments, an effect exacerbated by asymmetric information and specialized roles. While compromised agents consistently develop distinct objective-dependent reasoning strategies, these adaptations remain largely invisible in their public behavior. More broadly, our findings suggest that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.

13:00 JSTエージェント研究/論文GPT / ChatGPT

ClinLens: 縦断的マルチモーダル臨床データ サイエンスのための長期的なコーディング エージェントに向けて

臨床データサイエンスエージェントは、異種の縦断的記録を監査可能な分析に変換する必要がありますが、既存のベンチマークは、医療質問応答、構造化された表の推論、または一般的な科学リポジトリを主に分離しています。 CLINLENS は、構造化された電子健康記録、メモ、心電図、胸部 X 線写真、心エコー図にわたる 5 つのリンクされた MIMIC リソースにわたる 200 の実行可能なタスクのベンチマークです。 4 x 5 分類法は、5 つの分析機能を備えた 4 つの患者時間スコープを横断します。プログラムファースト逆合成では、各境界付き半未加工パッケージを評価者プライベート参照ワークフローと組み合わせて、必要なアーティファクト、コホートおよび時間セマンティクス、および最終的な回答をチェックします。固定の 126 タスク スイートでは、24 の標準化されたモデル スキャフォールド構成のうち最も強力な構成は、100% の EXECSUCCESS にもかかわらず、56.3% のスコープ マクロ STRICTPASS を達成します。参考までに、個別に構成されたコーディング エージェントは 126 タスク中 83 タスクを解決しますが、GPT-4o-mini に適応した 5 つの生物医学システムは最大 2.9% のスコープ マクロ STRICTPASS に達します。これらの結果は、実行可能な申請と正しい臨床分析との間に大きなギャップがあることを明らかにしています。

原文 (English)

ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science

Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely isolate medical question answering, structured-table reasoning, or generic scientific repositories. We introduce CLINLENS, a benchmark of 200 executable tasks over five linked MIMIC resources spanning structured electronic health records, notes, electrocardiograms, chest radiographs, and echocardiograms. A 4 x 5 taxonomy crosses four patient-time scopes with five analysis capabilities. Program-first reverse synthesis pairs each bounded semi-raw package with an evaluator-private reference workflow and checks required artifacts, cohort and temporal semantics, and the final answer. On a fixed 126-task suite, the strongest of 24 standardized model-scaffold configurations achieves 56.3% scope-macro STRICTPASS despite 100% EXECSUCCESS. For reference, a separately configured coding agent solves 83 of 126 tasks, while five biomedical systems adapted to GPT-4o-mini reach at most 2.9% scope-macro STRICTPASS. These results expose a substantial gap between runnable submissions and correct clinical analyses.

13:00 JSTビジネス/資金調達研究/論文

ベンチマーク推論が成立しない場合:AI評価における予測可能性

AI ベンチマークの結果が 1 ステップで重大な主張に達することはほとんどありません。評価者は、それをさらなるケースに一般化し、能力の証拠として解釈し、新しいタスクに推定し、別のシステムまたはサイトに移し、人間によるレビューと下流の結果に関する仮定と組み合わせます。妥当性中心のアプローチでは、各主張の証拠が必要です。この論文では、さらなる認識論的問題を特定します。それは、保証されたリンクは自動的に保証されたチェーンを作成しないということです。ある研究のターゲットが次の研究のソースになるとは限りません。システム、人口、結果、または条件はインターフェースで変更される可能性があります。また、共有データやモデル系統により、一見独立したサポートに依存する可能性があります。予測可能性は、観察されたケースから観察されていないケースへの限定された拡張が保証されるかどうかに関係します。グッドマンは競合拡張の問題を提供します。引数ベースの妥当性は、それらをテストするためのアーキテクチャを提供します。この論文の特徴的な主張は、非合成原理です。隣接する投影のサポートは、エンドポイントと仮定が一致し、依存性と不確実性が貫徹される場合にのみ合成を保証します。法的調査の事例は、ベンチマーク証拠と展開調査がそれぞれ並行性を保ちながらどのように健全であるかを示しています。再分析とシミュレーションは、骨材の安定性が後の予測で必要となる区別を消去できる理由を示しています。結果として得られる投影可能性監査により、ベンチマークで使用する引数内のサポートされていない結合が診断されます。

原文 (English)

When benchmark inferences do not compose: Projectibility in AI evaluation

An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don't automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent. Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper's distinctive claim is a non-composition principle: support for adjacent projections warrants their composition only when endpoints and assumptions align and dependence and uncertainty are carried through. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A reanalysis and simulation show why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.

13:00 JSTLLM/生成AIエージェント

GuideSkill: ガイドラインに基づいた臨床推論のための実行可能な LLM エージェント スキルの進化

臨床実践ガイドライン (CPG) は診断基準をエンコードしていますが、LLM システムは通常、ルールを実行するのではなく、ガイドラインのテキストを取得するか、トレーニングを通じてガイドラインを吸収します。疾患固有の基準を実行可能な関数にコンパイルし、通常の診断サポート スコアを返す外部推論レイヤーである GuideSkill を導入します。 GuideSkill-Zero はガイドラインから初期化されますが、GuideSkill-Evo はケースと診断のペアを使用して、対象となるスキルを調整し、不足している診断を追加します。推論時に、LLM は鑑別診断を提案し、一致した各スキルに必要な機能を根拠付け、そのランキングと実行されたスキル スコアを融合します。 GuideSkill-Zero は、4 つのベンチマークと 4 つのバックボーンにわたって、ガイドライン RAG を超えるマクロ平均精度を平均 13.45% 向上させます。 GuideSkill-Evo は、すべてのバックボーンで最高のマクロ平均を達成し、直接推論よりも相対的に 18.49% 向上し、ゴールドラベルのスキル カバレッジを 56.5% から 99.5% に高めます。 Qwen3.5-9B では、バックボーンを更新しなくても、最も強力なパラメータ更新ベースラインを 11.16% 上回っています。専門家の評価はさらに、GuideSkill が臨床的に適切で広く受け入れられるスキルを生成することを示しており、その初期化および進化したルールが信頼でき、実用的に意味があることを示唆しています。これらの結果は、ガイドライン由来の手順と症例由来の診断パターンを組み合わせるためのモデルに依存しないメカニズムとして実行可能なスキルを裏付けています。

原文 (English)

GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning

Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through training rather than execute its rules. We introduce GuideSkill, an external reasoning layer that compiles disease-specific criteria into executable functions returning ordinal diagnostic-support scores. GuideSkill-Zero is initialized from guidelines, while GuideSkill-Evo uses case--diagnosis pairs to refine covered skills and add missing diagnoses. At inference, an LLM proposes a differential diagnosis, grounds the features required by each matched skill, and fuses its ranking with the executed skill scores. Across four benchmarks and four backbones, GuideSkill-Zero improves macro-average accuracy over guideline RAG by 13.45% on average. GuideSkill-Evo achieves the highest macro-average for every backbone, improves over direct inference by 18.49% relatively, and increases gold-label skill coverage from 56.5% to 99.5%. On Qwen3.5-9B, it also exceeds the strongest parameter-update baseline by 11.16% without updating the backbone. Expert evaluation further indicates that GuideSkill produces clinically sound and broadly acceptable skills, suggesting that its initialized and evolved rules are reliable and practically meaningful. These results support executable skills as a model-agnostic mechanism for combining guideline-derived procedures with case-derived diagnostic patterns.

13:00 JSTエージェント

GoGoTB: 仕様に基づいたカバレッジクロージャを備えたエージェント RTL 検証

機能検証は集積回路 (IC) フロントエンドのエンジニアリング作業の大半を占めており、シリコンに逃げ込んだ 1 つの見逃したバグがコストのかかる再スピンを引き起こす可能性があります。最近の大規模言語モデル (LLM) では、このプロセスを自動化する新たな機会が提供されていますが、既存の LLM ベースのアプローチでは、コンテキストを共有せずに独立したシングルターン呼び出しを通じて各コンポーネントが生成されるため、インターフェイスの不一致が検出されず、報告されるカバレッジが仕様要件から切り離されたままになります。これらの課題に対処するために、エージェント実行制御層、進化可能な知識システム、仕様に基づいたカバレッジ クロージャの 3 つのサブシステムを通じてエンドツーエンドの検証クロージャを実現するエージェント フレームワークである GoGoTB を紹介します。実行制御層は、すべてのツールとステージの境界で、決定論的な強制を LLM 推論から分離します。ナレッジ システムは、方法論と設計固有の専門知識をオンデマンドで提供します。カバレッジ フレームワークは、すべてのビンを指定された仕様の動作に固定するため、残存する各ギャップには診断可能な根本原因と対象を絞った解決策が含まれます。人間の介入なしで 8 つのレジスタ転送レベル (RTL) デザインでテストされた GoGoTB は、100\% の環境生成成功を達成し、平均 98.4\% ライン、97.2\% 分岐、97.0\% トグル、および 83.2\% の機能カバレッジを達成しました。これまでの研究では、完全な検証環境を正常に生成したり、同じベンチマークで有意義なカバレッジを達成したりすることはありません。

原文 (English)

GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure

Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can trigger a costly respin. Recent large language models (LLMs) offer new opportunities to automate this process, yet existing LLM-based approaches generate each component through independent single-turn calls with no shared context, leaving interface mismatches undetected and reported coverage disconnected from specification requirements. To address these challenges, we present GoGoTB, an agentic framework that achieves end-to-end verification closure through three subsystems: an agentic execution control layer, an evolvable knowledge system, and specification-grounded coverage closure. The execution control layer separates deterministic enforcement from LLM reasoning at every tool and stage boundary. The knowledge system dispatches methodology and design-specific expertise on demand. The coverage framework anchors every bin to a named specification behavior so that each residual gap has a diagnosable root cause and a targeted remedy. Tested on 8 register transfer level (RTL) designs without any human intervention, GoGoTB achieves 100\% environment generation success and averages 98.4\% line, 97.2\% branch, 97.0\% toggle, and 83.2\% functional coverage. No prior work successfully generates a complete verification environment or achieves meaningful coverage on the same benchmarks.

13:00 JSTLLM/生成AIビジネス/資金調達

立場: 評価スコアは朽ちる知識の主張である

言語モデルの評価方法では、自動化されたメトリクスや LLM による審査員による評価から人間による評価やベンチマーク スイートの結果まで、複数のシグナルを組み合わせることがますます増えています。これらのシグナルが平均化によって集約されると、評価の信頼度が最も弱いシグナルの信頼性を大幅に超える可能性があります。これを評価における信頼インフレーションと呼びます。私たちは、評価スコアは 3 つの特性を持つ認識論的主張として扱われるべきであると主張します: 形式性 (人間の評価は自動化された指標よりも強力な証拠を提供します)、範囲 (ベンチマーク結果は普遍的ではなく、テストされた分布に適用されます)、および妥当性ウィンドウ (汚染が蓄積し分布が変化するとベンチマーク結果は期限切れになります)。いくつかの収束する研究の伝統 (思考連鎖分析、可能論的論理、代数理論) は、単一の悲観パラメーターによって制御されるパラメーター化された演算子ファミリーの保守的なエンドポイントとして最弱リンク集約を確立します。これらの伝統と、エージェント型 AI の評価ハーネスの構築から得た具体的な教訓を活用して、評価結果に明示的なメタデータ (形式層、範囲宣言、有効期限) を含めて、評価結果の認識的ステータスを透明にすることを提案します。公開 HELM リーダーボードでの平均集計のコストを示します。10 のシナリオにおける 54 のフロンティア モデル全体で、平均スコアと最弱リンクによってランク付けされた上位 5 つのモデルは完全に素になっています。

原文 (English)

Position: Evaluation Scores Are Perishable Knowledge Claims

Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to human assessments and benchmark suite results. When these signals are aggregated via averaging, evaluation confidence can then substantially exceed the reliability of the weakest signal: a phenomenon we call trust inflation in evaluation. We argue that evaluation scores should be treated as epistemic claims with three properties: formality (human evaluation provides stronger evidence than an automated metric), scope (a benchmark result applies to the tested distribution, not universally), and validity windows (benchmark results expire as contamination accumulates and distributions shift). Several converging research traditions (chain-of-thought analysis, possibilistic logic, and algebraic theory) establish weakest-link aggregation as the conservative endpoint of a parameterized operator family controlled by a single pessimism parameter. Drawing on those traditions, and on concrete lessons from building an evaluation harness for agentic AI, we propose that evaluation results carry explicit metadata (formality tier, scope declaration, and expiration date) to make their epistemic status transparent. We illustrate the cost of mean aggregation on the public HELM leaderboard: across 54 frontier models on ten scenarios, the top-five models ranked by mean score and by weakest-link are completely disjoint.

13:00 JSTLLM/生成AIエージェントハードウェア/半導体Gemini

TraceCoder: 位置キー スニペットのバージョニングによる説明可能で監査可能なコード生成

現代の LLM ベースのコーディング エージェントは、ブラック ボックス出力としてコードを生成します。各行の背後にある理論的根拠は隠され、ベンチマーク主導の修復によるコードの進化は一時的で、事後監査は不可能です。我々は、次の 3 つの相補的なメカニズムを通じてこれらの欠点に対処するコード生成コンセプトを提示します。(i) 修復イベントごとに、ベンチマーク参照、ラウンド番号、障害テキスト、および LLM 説明を記録するリレーショナル スニペット履歴スキーマ。これにより、完全な来歴クエリが可能になります。 (ii) この履歴をヒートマップされたホバーアノテーション付きのソース コードとしてレンダリングするブラウザベースの視覚化ツール。 (iii) ツリーノード区切り文字を備えた競争力のある分数位置キーインデックス付けスキーム。辞書編集的に順序付けされた安定した識別子を各コードスニペットに割り当て、周囲の行を中断することなくきめ細かい追跡を可能にします。 2 つのプロバイダー構成にわたって、文字列処理、数学的計算、データ構造操作に及ぶ 30 のアルゴリズム プログラミング タスクで TraceCoder を評価します。このうち 10 件は、微妙なエッジケースの動作を伴うタスクで 6 反復の予算を使い果たしてしまいます。平均 Chg% は 30% に達し、20 タスクのサブセットの唯一のプロバイダーとして Gemini 2.0 Flash を使用した場合の 21% と比較して、10 個中 3 個のコード スニペットに追跡可能な修復イベント行が含まれています。 3 つの詳細なケース スタディは、最終プログラムの各行を形成する特定のベンチマークの失敗をシステムがどのように説明するかを示しています。提案されたメカニズムにより、自動コード生成の内部「ナラティブ」が監査可能かつ再生可能になり、実稼働デプロイメントにおける信頼と責任にとって不可欠な特性となります。

原文 (English)

TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning

Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible. We present a code generation concept that addresses these shortcomings through three complementary mechanisms: (i) a relational snippet-history schema that records, per repair event, the benchmark reference, round number, failure text, and LLM explanation, enabling full provenance queries; (ii) a browser-based visualisation tool that renders this history as heat-mapped, hover-annotated source code; and (iii) a competitive fractional position-key indexing scheme with tree-node delimiters that assigns stable, lexicographically-ordered identifiers to each code snippet, enabling fine-grained tracking without disrupting surrounding lines. We evaluate TraceCoder on 30 algorithmic programming tasks spanning string processing, mathematical computation, and data-structure manipulation, across two provider configurations. Of these, 10 exhaust the 6-iteration budget on tasks with subtle edge-case behaviour. Mean Chg% reaches 30%, three in ten code snippets carry a traceable repair-event row, compared to 21% when using Gemini 2.0 Flash as sole provider on a 20-task subset. Three detailed case studies demonstrate how the system explains which specific benchmark failures shaped each line of the final program. The proposed mechanism makes the internal "narrative" of automated code generation auditable and replayable, a property essential for trust and accountability in production deployments.

13:00 JSTエージェント

物理問題における構造の探索: AI エージェントは統計的機械的マッピングを発見できるか?

理論物理学における重要なスキルは、新しい問題がいつ既知のモデルに変換されるかを認識することです。私たちはこのスキルを AI エージェントのタスクとして研究します。LLM ベースのエージェントは、生の分割関数から扱いやすい表現への統計的機械的マッピングを発見できるでしょうか?この疑問を調査するために、伝達行列法、ゲージ除去可能な無秩序、および平面/パフィアン構造をカバーする 6 つのイジング型問題のベンチマークである StatMechBench-v0 を紹介します。複数の LLM と問題の表現にわたって、単純な提案、検証、修正エージェントを評価します。結果は、数値フィードバックが多くの場合、エージェントがコードを修復し、正しいパーティション関数を回復するのに役立つことを示しています。ただし、エージェントは、基礎となる扱いやすいクラスを誤認したり、計算の複雑さを過小評価したりしながら、数値チェックに合格することもあります。これは、現在の LLM 推論の限界を明らかにするとともに、数値的一致を超えて、たとえば記号チェックや構造的不変条件を組み込んだ検証スタックを必要とすることを示しています。私たちの研究は、理論物理学における構造発見を目的とした AI エージェントの初期評価と設計の方向性を提供します。

原文 (English)

Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?

An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as an AI-agent task: can LLM-based agents discover statistical mechanical mappings from a raw partition function to a tractable representation? To probe this question, we introduce StatMechBench-v0, a benchmark of six Ising-type problems covering transfer-matrix methods, gauge-removable disorder, and planar/Pfaffian structure. We evaluate a simple propose-verify-revise agent across multiple LLMs and problem phrasings. The results show that numerical feedback often helps agents repair code and recover correct partition functions. However, agents can also pass the numerical checks while misidentifying the underlying tractable class or understating computational complexity. This both reveals limitations in current LLM reasoning and calls for a verification stack that goes beyond numerical agreement, incorporating, for example, symbolic checks and structural invariants. Our study provides an early evaluation and design directions for AI agents aimed at structural discovery in theoretical physics.

13:00 JSTエージェント

CaM-Wolf: 社会的控除ゲーム用の因果認識マルチモーダル エージェント

人狼などの社会的推理ゲーム (SDGs) は、AI エージェントにとって挑戦的なテストベッドとなっています。これらのゲームには、推論、欺瞞、コラボレーションなどの複雑な社会的スキルが必要です。大規模言語モデル (LLM) の最近の進歩により、SDG エージェントは大幅に進歩しましたが、現在のアプローチは主にテキストベースであり、人間の社会的相互作用の基本であるマルチモーダルな性質が見落とされています。このギャップを埋めるために、マルチモーダルな認識と生成を統合する初の SDG エージェントである CaM-Wolf を紹介します。 CaM-Wolf は、他のプレイヤーからのビデオ入力を処理し、強化学習によって訓練された因果認識の Reasoner を採用して、観察可能な行動と隠れた役割の間に論理的な連鎖を確立し、アニメーション化されたアバターを通じて自身を表現します。私たちの実験とユーザー調査では、CaM-Wolf が優れたエージェント ゲームプレイ パフォーマンスを実現し、人間と AI のインタラクションの質を向上させることが示されています。この取り組みは、微妙な社会力学に参加できる、より人間に近い AI エージェントの作成に向けた大きな進歩を表しています。私たちのコードは https://3dagentworld.github.io/avatar_wolf で入手できます。

原文 (English)

CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games

Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooking the multimodal nature that is fundamental to human social interaction. To bridge this gap, we introduce CaM-Wolf, the first SDG agent that integrates multimodal perception and generation. CaM-Wolf processes video inputs from other players, employs a causal-aware Reasoner trained via reinforcement learning to establish logical chains between observable behaviors and hidden roles, and presents itself through an animated avatar. Our experiments and user study show that CaM-Wolf achieves superior agent gameplay performance and enhances the quality of human-AI interaction. This work represents a significant advancement towards creating more human-like AI agents capable of participating in nuanced social dynamics. Our code is available at https://3dagentworld.github.io/avatar_wolf.

13:00 JST画像/動画生成ロボティクス

CG-World: 大規模な世界状態データセットと世界モデル用のプロトコル

世界モデルは、状態、アクション、イベント、観測の共同ダイナミクスを学習する必要がありますが、既存のビデオ、ロボット工学、およびシミュレーションのデータセットは、通常、この構造の一部のみをキャプチャします。産業用コンピュータ グラフィックス制作パイプラインから派生した大規模な世界状態データセットおよびプロトコルである CG-World を紹介します。 CG-World は、マルチモーダル セマンティクス、空間構造、骨格とコントローラーの状態、モーション カーブ、カメラと照明パラメーター、物理キャッシュ、接触イベント、マルチパス レンダリングなどの中間状態を明示的に記録します。 CG-World v1 には、時間的に整列された 1 ~ 5 秒のセグメントが約 850,000 個含まれています。潜在的な状態、観察、関係、イベント、および分岐メタデータを分離し、それらを統合された時空間サンプルに編成します。介入学習と反事実推論をサポートするために、CG-World は、介入ターゲット、不変条件、および代替結果が明示的に記録された、事実の軌跡、観察介入、行動介入、メカニズム介入、厳密な反事実分岐をカバーする分岐系統を定義します。ジオメトリ条件付きビデオ生成、アクション予測、および閉ループのビジョン-言語-アクション ポリシー転送に関するデータセットを評価します。結果は、CG-World が制御された生成、アクション モデリング、および具体化されたポリシーの転送に対して再利用可能な構造化された監視を提供することを示しています。私たちは、ワールド モデル、物理 AI、および身体化されたインテリジェンスの共有データ インフラストラクチャに向けた継続的なデータ収集とコミュニティ コラボレーションを通じて CG-World を拡大する予定です。

原文 (English)

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings. CG-World v1 contains approximately 850,000 temporally aligned segments of 1-5 seconds. It separates latent states, observations, relations, events, and branch metadata, and organizes them into unified spatiotemporal samples. To support intervention learning and counterfactual reasoning, CG-World defines a branch lineage covering factual trajectories, observation interventions, action interventions, mechanism interventions, and strict counterfactual branches, with intervention targets, invariants, and alternative outcomes explicitly recorded. We evaluate the dataset on geometry-conditioned video generation, action prediction, and closed-loop vision-language-action policy transfer. Results show that CG-World provides reusable structured supervision for controlled generation, action modeling, and embodied policy transfer. We plan to expand CG-World through continued data collection and community collaboration toward a shared data infrastructure for world models, Physical AI, and embodied intelligence.

13:00 JST研究/論文

MultivationBench: マルチモーダルな逐次動機推論のベンチマーク

マルチモーダル大規模言語モデルは、社会的インテリジェンスとしての可能性があるため、大きな関心を集めています。しかし、逐次的な動機推論を実行する彼らの能力はまだ十分に研究されていません。既存の評価は主に静的テキストまたは分離された視覚的なスナップショットを調査するものであり、現実世界の行動要因の累積的な性質を反映していません。このギャップに対処するために、ストーリー主導の視覚的物語内でマルチモーダルな動機推論を厳密に評価するように設計されたベンチマークである MultivationBench を導入します。このベンチマークは、確立された心理学的フレームワーク (マズローの階層構造とリースの基本的欲求) に基づいて構築されており、進化する動機を推測するには、蓄積されたマルチモーダルなコンテキストを統合するモデルが必要です。結果は、MultivationBench が重大な課題を提示していることを示しています。テストされたすべてのモデルは、連続するコンテキスト全体で一貫した動機推論を維持するのに苦労しており、静的な認識能力と、人間のような社会的理解に不可欠な動的な推論との間に重大な断絶が明らかになりました。

原文 (English)

MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning

Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.

13:00 JSTエージェント

EvoPINN: 物理学に基づいたニューラル ネットワークの実行可能アルゴリズムのエージェント的発見

物理情報に基づくニューラル ネットワーク (PINN) は、偏微分方程式 (PDE) を解くための強力なパラダイムとして登場しましたが、そのパフォーマンスは、ニューラル表現、損失定式化、および最適化ダイナミクスの手動の試行錯誤エンジニアリングに大きく依存しています。大規模言語モデル (LLM) は自動設計に有望な手段を提供しますが、制約のないコード生成では、科学計算の厳密な制約の下では数学的に無効な、または数値的に不安定な解が生成されることがよくあります。このギャップを埋めるために、私たちは、PINN 開発を労働集約的な手動設計から厳密で実行に基づいたアルゴリズム発見問題に再定式化するエージェント フレームワークである \textbf{EvoPINN} を提案します。 EvoPINN は、トレーニング プログラムからニューラル表現を切り離し、LLM エージェントを利用してメモリ条件付きのプログラム変更を繰り返し提案することで、モジュール式検索空間をナビゲートします。科学的妥当性を保証するために、すべての候補は厳密な構造検証と予算に見合った PDE 評価を受けます。多様な偏微分方程式(振動、楕円、散逸、非線形輸送)にわたる広範な実験により、EvoPINNがベースラインと比較して相対的な$L_{2}$誤差を大幅に削減する偏微分方程式に特化した学習アルゴリズムを発見したことが実証されました。重要なのは、EvoPINN が自律的に SLRC-PINN を発明したことです。SLRC-PINN は、厳密なパラメーター一致比較の下でパフォーマンスの向上が持続する新しいアーキテクチャであり、真に新しい科学計算メカニズムを発見するための実行ベースのエージェントの実行可能性を確立しました。

原文 (English)

EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks

Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs), yet their performance heavily relies on the manual, trial-and-error engineering of neural representations, loss formulations, and optimization dynamics. While Large Language Models (LLMs) offer a promising avenue for automated design, unconstrained code generation often yields mathematically invalid or numerically unstable solutions under strict scientific computing constraints. To bridge this gap, we propose \textbf{EvoPINN}, an agentic framework that reformulates PINN development from labor-intensive manual design into a rigorous, execution-grounded algorithm discovery problem. EvoPINN navigates a modular search space by decoupling neural representations from training programs, utilizing an LLM agent to iteratively propose memory-conditioned programmatic modifications. To ensure scientific validity, all candidates undergo strict structural verification and budget-matched PDE evaluation. Extensive experiments across diverse PDE regimes (oscillatory, elliptic, dissipative, and nonlinear transport) demonstrate that EvoPINN discovers PDE-specialized learning algorithms that significantly reduce relative $L_{2}$ error compared to baselines. Crucially, EvoPINN autonomously invented SLRC-PINN, a novel architecture whose performance gains persist under rigorous parameter-matched comparisons, establishing the viability of execution-grounded agents for discovering genuinely new scientific computing mechanisms.

13:00 JSTエージェント

クレームと証拠のトレーサビリティに関する証拠と台帳の裁定

AI エージェントは、著者が引用または取得した証拠が主張を裏付けるかどうかを確認するよりも早く、主張を起草することができます。私たちは、証拠と台帳の裁定を研究します。これは、各主張を証拠パケットと組み合わせ、裏付け関係を割り当て、裏付けのない主張、矛盾した主張、または証拠が混合された主張を著者に送り返す、主張と証拠のトレーサビリティ ワークフローです。経験的コアは、AVeriTeC、CLIMATE-FEVER、および SciFact の独立した外部ラベルから構築された 2,335 行のブラインド ベンチマークです。ゴールド関係とソース証拠ラベルは予測中に非表示になり、スコアリングの目的でのみ結合されます。このベンチマークでは、エージェント証拠台帳条件は、最良の非エージェント ベースラインの精度 0.383 およびマクロ F1 0.303 と比較して、関係精度 0.676 およびマクロ F1 0.601 を達成しています。また、ゴールド ラベルが矛盾、証拠の欠落、または証拠の混合を示す 1270/1435 件の主張をルーティングする一方で、295/900 件の裏付けのある主張をルーティングします。これらの結果は、証拠台帳の裁定により、異種証拠パケットを AI 支援書き込みのための監査可能なトレーサビリティ層に変えることができることを示しています。

原文 (English)

Evidence-Ledger Adjudication for Claim-Evidence Traceability

AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger adjudication: a claim-evidence traceability workflow that pairs each claim with an evidence packet, assigns a support relation, and routes unsupported, contradicted, or mixed-evidence claims back to the author. The empirical core is a 2,335-row blind benchmark built from independent external labels in AVeriTeC, CLIMATE-FEVER, and SciFact. Gold relations and source evidence labels are hidden during prediction and joined only for scoring. On this benchmark, the agent evidence-ledger condition achieves 0.676 relation accuracy and 0.601 macro-F1, compared with 0.383 accuracy and 0.303 macro-F1 for the best non-agent baseline. It also routes 1270/1435 claims whose gold labels indicate contradiction, missing evidence, or mixed evidence, while routing 295/900 supported claims. These results show that evidence-ledger adjudication can turn heterogeneous evidence packets into an auditable traceability layer for AI-assisted writing.

13:00 JSTLLM/生成AIエージェント

Eco3S: エージェントベースのモデルによる複雑な社会経済システムのシミュレーション

大規模言語モデル (LLM) の急速な発展により、エージェントベース モデリング (ABM) への関心が改めて高まっています。しかし、現在の LLM ベースの ABM 研究は、進化するエージェントと環境の相互作用のモデル化、柔軟な反事実推論の可能化、科学研究のためのシミュレーション ワークフローの自動化など、いくつかの重要な課題に直面しています。この論文では、経済研究と政策分析のための社会経済システム シミュレーション フレームワークである Eco3S を提案します。これは、3 つの主要なメカニズムを通じてこれらの課題に対処します。 (2) 構造因果シミュレーション。多様な因果推論タスクに対する柔軟な介入を可能にする、構造因果モデル (SCM) にヒントを得た反事実メカニズム。 (3) シミュレーション-分析-改良パラダイム。事前のシミュレーション結果に基づいて実験計画を反復的に改良する自己修正メカニズム。多様な経済シナリオに関する実験により、確立された複数の経済研究 (運河の衰退、統治の起源、情報伝播) や現象を領域全体で再現する際の \textit{Eco3S} の有効性が確認されています。追加の結果は、その拡張性と一般化性をさらに実証し、厳密な経済調査と政策立案に対するこのフレームワークの可能性を強調しています。

原文 (English)

Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models

The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces several key challenges: modeling evolving agent-environment interactions, enabling flexible counterfactual reasoning, and automating simulation workflows for scientific research. In this paper, we propose Eco3S, a socio-economic system simulation framework for economic research and policy analysis that addresses these challenges through three key mechanisms: (1) Co-evolving Environment Design, a bidirectional feedback loop where agents and the environment co-evolve, producing realistic emergent behaviors; (2) Structural Causal Simulation, a structural causal model (SCM)-inspired counterfactual mechanism that allows flexible interventions for diverse causal inference tasks; (3) Simulation-Analysis-Refinement Paradigm, a self-corrective mechanism that iteratively refines experimental designs based on prior simulation results. Experiments on diverse economic scenarios confirm \textit{Eco3S}'s effectiveness in replicating multiple established economic studies (canal decay, origins of governance, and information propagation) and phenomena across domains. Additional results further demonstrate its scalability and generalizability, highlighting the framework's potential for rigorous economic research and policy-making.

13:00 JST研究/論文

説明は少なく、より良いコード: コーディング アシスタントにおけるセッション間のベンチマークのパーソナライズされた曖昧さの適応

AI 支援コーディングでは、非公式なユーザーの意図が実行可能なソフトウェアに変換されることが増えていますが、コーディング要求にはタスクやセッション全体でユーザー固有の方法で繰り返される曖昧さが含まれることがよくあります。既存の曖昧さ回避方法は通常、追加の明確化を引き出すことによって、現在のコーディング セッション内で各曖昧なリクエストを個別に解決します。ただし、同じユーザーからの解決されたセッション履歴が、新しく開かれたセッションで繰り返される個人化された曖昧さを解決するための記憶として機能できるかどうかは、まだ調査されていません。私たちは、パーソナライズされたあいまいさの適応を新しいタスクとして定式化します。ユーザーが以前に解決したコーディング セッションと新しいあいまいなリクエストを考慮すると、アシスタントは、繰り返されるあいまいさのパターンを特定し、意図した実行可能なソリューションを生成し、明確化を最小限に抑える必要があります。このタスクのベンチマークを行うために、CAPA を導入します。これは、6 つのメカニズムを通じてパーソナライズされたコーディングの曖昧さを特徴付け、制御された 3 段階の生成パイプラインを使用して、これらのメカニズムを明確な実行可能タスクに注入します。 CAPA には、60 のバランスの取れたユーザー曖昧性セルにわたる 600 のコーディング セッションが含まれており、これには 300 の保留された評価セッションが含まれます。実行可能成功、最初のターンの成功、および完了までのターンを使用して、履歴なしおよび同一ユーザー履歴の条件下で 12 個の最近の LLM を評価します。私たちの分析では、タスクの難易度、ユーザー ID、メモリベースの履歴の使用を調査し、さらに、軽量の推論時間手法として同一ユーザー履歴ゲーティングを提案します。 CAPA は、繰り返しの説明を減らしながら、生成されたコードをユーザーの意図に合わせて調整する、長期的なコーディング アシスタントを開発するための基盤を提供します。

原文 (English)

Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants

AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities that recur in user-specific ways across tasks and sessions. Existing disambiguation methods typically address each ambiguous request in isolation within the current coding session, often through eliciting additional clarification. However, whether resolved session history from the same user can serve as memory for resolving recurring personalized ambiguity in a newly opened session remains underexplored. We formulate personalized ambiguity adaptation as a new task: given a user's previously resolved coding sessions and a new ambiguous request, an assistant should identify the recurring ambiguity pattern, produce the intended executable solution, and minimize clarification. To benchmark this task, we introduce CAPA, which characterizes personalized coding ambiguity through six mechanisms and injects these mechanisms into unambiguous executable tasks using a controlled three-stage generation pipeline. CAPA contains 600 coding sessions across 60 balanced user--ambiguity cells, including 300 held-out evaluation sessions. We evaluate 12 recent LLMs under no-history and same-user-history conditions using executable success, first-turn success, and turns-to-completion. Our analyses examine task difficulty, user identity, and memory-based history use, and we further propose same-user history gating as a lightweight inference-time method. CAPA provides a foundation for developing long-term coding assistants that better align generated code with user intent while reducing repeated clarification.

13:00 JSTLLM/生成AIエージェント

AlphaSchema: LLM ベースのアルファ マイニングの取引セマンティクスの空間を探索する

自動アルファマイニングでは、因子生成と反復発見のために大規模言語モデル (LLM) エージェントの採用が増えています。ただし、既存の LLM ベースのシステムは、明示的な探索スペースやそのスペースをナビゲートするための原則に基づいたメカニズムを持たずに、因子の構築と検索の決定の両方をエージェント自体に委任することがよくあります。その結果、探索の大部分は暗黙的なままとなり、体系的に制御または最適化することが困難になります。アルファマイニングのための取引セマンティクスの構造化空間を構築および探索する AlphaSchema を紹介します。この空間内の各ポイントは、イベント、コンテキスト、品質、方向、および出力で構成されるスキーマ プランであり、実装前に候補要素のセマンティクスを指定します。 AlphaSchema は探索を実装から切り離します。LLM は選択されたスキーマ プランを実行可能な要素に変換します。その一方で、評価された報酬はセマンティック空間上の代理モデルを学習するために蓄積されます。反復選択メカニズムは、このモデルを使用して、グローバル探索、サロゲートに基づく活用、およびローカル突然変異のバランスをとります。中国株式市場の実験では、AlphaSchema が強力な予測とポートフォリオのパフォーマンスを備えたファクタープールを発見したことが示されています。さらなる分析により、セマンティック検索プロセスが多様な領域をナビゲートしながら、評価を高報酬領域に徐々に割り当てていること、および異なる LLM による同じスキーマ プランの実装が同等の予測品質を示していることが示されており、アルファ マイニングの品質がフレームワーク内での LLM の選択に対してほぼ堅牢であることが示唆されています。

原文 (English)

AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining

Automated alpha mining has increasingly adopted large language model (LLM) agents for factor generation and iterative discovery. However, existing LLM-based systems often delegate both factor construction and search decisions to the agent itself, without an explicit exploration space or a principled mechanism for navigating that space. As a result, exploration remains largely implicit and difficult to control or optimize systematically. We introduce AlphaSchema, which constructs and explores a structured space of trading semantics for alpha mining. Each point in this space is a schema plan composed of Event, Context, Qualities, Direction, and Output, specifying the semantics of a candidate factor before implementation. AlphaSchema decouples exploration from implementation: an LLM translates selected schema plans into executable factors, while evaluated rewards are accumulated to learn a surrogate model over the semantic space. An iterative selection mechanism uses this model to balance global exploration, surrogate-guided exploitation, and local mutation. Experiments on the Chinese stock market show that AlphaSchema discovers factor pools with strong predictive and portfolio performance. Further analyses show that the semantic search process navigates diverse regions while increasingly allocating evaluations toward high-reward regions, and that implementations of the same schema plans by different LLMs exhibit comparable predictive quality, suggesting that alpha mining quality is largely robust to the choice of LLM within our framework.

13:00 JSTLLM/生成AIエージェント

自己進化の再考: スキルの過剰適合を軽減するための制約付き探索・活用プロセス

大規模言語モデル (LLM) エージェントが過去の対話からの経験を蓄積して再利用できるようにすることは、現実世界のアプリケーションにおける中心的な課題のままです。有望な解決策は、スキルをトレーニング可能な状態として扱い、ニューラル ネットワーク トレーニングのモデル パラメーターと同じ方法で最適化することです。ただし、データ駆動型のスキルの最適化は、実際の環境から収集された限られた軌道に過剰適合する傾向があります。これらの軌跡を過剰に利用すると、現在のバッチがオーバーフィットしますが、制約のない探索では、以前に解決されたケースの回帰が発生します。この緊張は、探求と搾取のトレードオフによって支配される、スキルの自己進化に対する制約された検索の見方を動機づけます。私たちは、両方のリスクを軽減する 3 段階のフレームワークである SkillBoost を提案します。構造化されたエクスプロイトは観察された障害を編集可能なスキル コンポーネントに特定し、事前ガイド付き探索は LLM の事前知識に基づいて多様な修復候補を生成し、検証済みの受け入れは回帰限界内でパフォーマンスが向上する場合にのみ候補をコミットします。 23 のモデルのベンチマーク構成にわたる実験では、SkillBoost が過学習を軽減しながら最先端のパフォーマンスを達成し、人間が作成したスキルと LLM が生成したスキルの両方を上回るパフォーマンスを発揮することが示されています。さらに、転送実験では、最適化されたスキルを他のエージェントが同様のタスクで再利用できることを示しています。

原文 (English)

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applications. A promising solution is to treat skills as trainable states and optimize them in the same way as model parameters in neural network training. However, data-driven skill optimization is prone to overfitting to the limited trajectories collected from real environments. Overexploiting these trajectories overfits the current batch, while unconstrained exploration causes regression on previously solved cases. This tension motivates a constrained search view of skill self-evolution, governed by an exploration--exploitation trade-off. We propose SkillBoost, a three-stage framework that mitigates both risks: structured exploitation localizes observed failures to editable skill components, prior-guided exploration draws on prior knowledge in the LLM to generate diverse repair candidates, and verified acceptance commits a candidate only when it improves performance within a regression bound. Experiments across 23 model--benchmark configurations show that SkillBoost achieves state-of-the-art performance while mitigating overfitting, outperforming both human-crafted and LLM-generated skills. Transfer experiments further show that optimized skills can be reused by other agents on similar tasks.

13:00 JSTエージェント

AgenticCANN: 知識拡張された Agentic Evolution による Ascend C オペレーターの自動生成

Ascend C のオペレーターの最適化は、NPU (Neural Processing Unit) の推論パフォーマンスにとって重要ですが、ハードウェアに関する深い専門知識が必要です。大規模言語モデル (LLM) は自動 CUDA カーネル生成において有望であることが示されていますが、Ascend C の根本的に異なるプログラミング モデルにより、未解明なままの固有の課題が生じます。この論文では、低コーパス NPU 環境での Ascend C オペレーター合成の自動化に特化した、知識拡張型エージェント進化フレームワークである AgenticCANN を提案します。不慣れなハードウェアでの深刻なプラットフォーム知識不足を克服するために、AgenticCANN には、上流の実現可能性のボトルネックを解決するために、開発ライフサイクル全体にわたって構造化されたマルチレベルのドメイン洞察を提供する知識統合生成システムが組み込まれています。この基盤に基づいて、動的に実行する段階適応型エージェント進化戦略を特徴としています。 LLM インタラクション モードを特定の生成フェーズと進化フェーズに合わせて調整し、高度な探索候補の発見と高度な収束パフォーマンス調整のバランスをとります。5 つのパターン カテゴリにわたる 6 つの演算子にわたる Huawei Ascend 910B での広範な実験により、私たちの手法が要素ごとの演算子と正規化演算子で 90 ~ 100%、融合演算子で 56% の実現可能性を達成し、1B Pangu モデル推論で最大 6.65 倍の高速化が達成されることが実証されました。カーネル。さらなる分析により、知識注入は要素ごとの演算子で実現可能性を 57% から 86% に単調に向上させることが明らかになり、演算子固有の利点ではなく一般的な利点が実証されました。

原文 (English)

AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution

Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.While large language models (LLMs) have shown promise in automated CUDA kernel generation, the fundamentally different programming model of Ascend C introduces unique challenges that remain unexplored. In this paper, we propose AgenticCANN, a knowledge-augmented agentic evolution framework specifically tailored for automated Ascend C operator synthesis in low-corpus NPU environments.To overcome the severe platform knowledge deficit on unfamiliar hardware, AgenticCANN incorporates a knowledge-orchestrated generation system that delivers structured, multi-level domain insights across the development lifecycle to resolve the upstream feasibility bottleneck.Building on this foundation, it features a stage-adaptive agentic evolution strategy that dynamically aligns LLM interaction modes with specific generation and evolution phases, balancing high-exploration candidate discovery with high-convergence performance tuning.Extensive experiments on Huawei Ascend 910B across six operators spanning five pattern categories demonstrate that our method achieves 90 to 100 percent feasibility on elementwise and normalization operators, 56% on fusion operators, and up to 6.65$\times$ speedup on 1B Pangu model inference kernels. Further analysis reveals that knowledge injection monotonically improves feasibility from 57% to 86% on elementwise operators, demonstrating its general rather than operator-specific benefit.

13:00 JSTLLM/生成AIエージェント

UrbanDS: データ集約型の都市タスク用のグラフガイド付き LLM マルチエージェント システム

大規模言語モデル (LLM) エージェントは、データ サイエンス タスクの自動化に広く適用されています。ただし、既存の方法は通常、提供される限られたデータセットに依存しており、大規模で異種のデータ リポジトリから関連情報を検出して活用する必要があるデータ集約型のシナリオでは課題に直面しています。都市データは大規模かつ複数ソースであるだけでなく、複雑な空間的、時間的、意味的な関係を示すため、都市のタスクはそのようなシナリオの代表的な例です。これらの課題に対処するために、データ集約型の都市タスク向けのグラフガイド付き LLM マルチエージェント システムである UrbanDS を提案します。まず、統合データセット グラフを構築して、再利用可能なデータセット スキルとデータセット間の関係を整理します。具体的には、各データセットのスキルを構築するデータ プロファイリング エージェントを開発します。さらに、関係エージェントはデータセット間の関係を特定し、これらの関係をデータセット グラフに統合します。実行時に、Planner Agent はタスク関連のデータセットをグラフから取得し、実行計画を生成します。次に、複数の実行エージェントがデータの処理と分析を実行し、その実行の進行状況と中間結果が共通メモリを通じて共有されます。最後に、レポート エージェントは実験ログをレポートに合成します。レポートは、ユーザーのフィードバックに基づいてさらに調整できます。データ集約型の都市シナリオを処理するエージェントの能力を体系的に評価するために、代表的なデータ分析とモデリング タスクをカバーする都市データ サイエンス ベンチマークである UrbanDS-Bench をさらに構築します。一般ベンチマークと都市ベンチマークの両方に関する実験では、UrbanDS がデータ集約型タスクにおいて既存のデータ サイエンス エージェントよりも常に優れたパフォーマンスを発揮することが実証されています。さらに、UrbanDS は武漢市東渓湖区の都市運営プラットフォームに導入され、現実世界の都市アプリケーションでの有効性を実証しています。

原文 (English)

UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks

Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a limited set of provided datasets, and they face challenges in data-intensive scenarios that require discovering and leveraging relevant information from large-scale and heterogeneous data repositories. Urban tasks are representative examples of such scenarios, as urban data are not only large-scale and multi-sourced, but also exhibit complex spatial, temporal, and semantic relationships. To address these challenges, we propose UrbanDS, a graph-guided LLM multi-agent system for data-intensive urban tasks. We first construct a unified dataset graph to organize reusable dataset skills and the relationships among datasets. Specifically, we develop a Data Profiling Agent that constructs a skill for each dataset. Moreover, a Relation Agent identifies relationships among datasets and integrates these relationships into the dataset graph. At runtime, a Planner Agent retrieves task-relevant datasets from the graph and generates execution plans. Multiple Execution Agents then perform data processing and analysis, while their execution progress and intermediate results are shared through a common memory. Finally, a Report Agent synthesizes the experimental logs into a report, which can be further refined based on user feedback. To systematically evaluate the capability of agents in handling data-intensive urban scenarios, we further construct UrbanDS-Bench, an urban data science benchmark covering representative data analysis and modeling tasks. Experiments on both general and urban benchmarks demonstrate that UrbanDS consistently outperforms existing data science agents on data-intensive tasks. Furthermore, UrbanDS has been deployed on the urban operations platform of Dongxihu District, Wuhan, demonstrating its effectiveness in real-world urban applications.

13:00 JSTLLM/生成AIエージェント

潜在チャネルは実際に通信しますか?潜在的なマルチエージェント LLM の因果関係の監査

大規模言語モデル (LLM) ベースのマルチエージェント システム (MAS) における潜在的な通信は、テキストの代わりに連続的な内部表現を送信しますが、表現能力が向上しても、受信者がタスク関連情報を使用することは確立されません。また、エンドタスクのパフォーマンスだけでは、観察された効果がメッセージの存在、評価された例に対して生成されたコンテンツ、または別のエージェントによって提供された情報に依存するかどうかを明らかにすることはできません。送信者が生成した表現が受信者に入る境界で、制御されたメッセージ置換を適用する因果監査を導入します。 4 つのメッセージ設定は、エンコードされた送信者情報、メッセージの存在とアイデンティティに対する受信者の感度、例固有のコンテンツのタスク値、および別のエージェントによって提供される追加値の 5 つの測定をサポートします。 GSM8K、ARC-C、および MATH-500 上の Qwen3-4B および Qwen3-8B を使用した潜在リレーに監査を適用します。 GSM8K では、Qwen3-4B の全体的なパフォーマンス効果 -1.00 パーセント ポイントは、他のサンプル メッセージによって保持される -6.17 ポイントの効果と、サンプル固有のコンテンツに起因する +5.17 ポイントの効果に分解されます。両方のコンポーネントの方向は 8B で反転します。 MATH-500 では、Qwen3-4B の 15.00 ポイントのゲインは、他のサンプル メッセージによって保持される 8.33 ポイントとサンプル固有のコンテンツに起因する 6.67 ポイントで構成されますが、8B のゲインは前者のコンポーネントによって支配されます。自己置換比較により、例固有のコンテンツと他のエージェントの値が異なることがさらにわかります。これらの結果は、集計精度では潜在メッセージが受信者にどのような影響を与えるかを特定するものではなく、潜在コミュニケーションの標準評価として制御されたメッセージの比較を動機付けるものではないことを示しています。

原文 (English)

Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM

Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of text, but greater representational capacity does not establish that the receiver uses task-relevant information. End-task performance alone also cannot reveal whether an observed effect depends on message presence, content generated for the evaluated example, or information supplied by a separate agent. We introduce a causal audit that applies controlled message replacements at the boundary where the sender-produced representation enters the receiver. Four message settings support five measurements of encoded sender information, receiver sensitivity to message presence and identity, the task value of example-specific content, and the additional value supplied by a separate agent. We apply the audit to latent relay with Qwen3-4B and Qwen3-8B on GSM8K, ARC-C, and MATH-500. On GSM8K, the Qwen3-4B overall performance effect of -1.00 percentage point decomposes into a -6.17-point effect retained by an other-example message and a +5.17-point effect attributable to example-specific content; both component directions reverse at 8B. On MATH-500, the Qwen3-4B gain of 15.00 points comprises 8.33 points retained by an other-example message and 6.67 points attributable to example-specific content, while the 8B gain is dominated by the former component. Self-substitution comparisons further show that example-specific content and other-agent value are distinct. These results show that aggregate accuracy does not identify how a latent message affects the receiver and motivate controlled message comparisons as a standard evaluation for latent communication.

13:00 JST研究/論文

マルコフ意思決定プロセスのためのプロパティ駆動型因果抽象化

マルコフ意思決定プロセス (MDP) は意思決定モデルとして広く使用されており、通常は状態変数とその評価を通じて因数分解された状態空間に対して指定されます。州の数が指数関数的に増加するため、MDP の多くの推論タスクが困難になります。抽象化は、MDP を削減し、スケーラビリティの問題を軽減する有望な手法です。この研究では、因数分解された MDP の因果関係の概念と、元の MDP モデルの多くの特徴を保持する新しいプロパティ駆動型の因果抽象化手法を導入します。このため、状態変数述語の因果関係に依存し、特定の抽象化プロパティを満たすか違反する同じ理由を共有する状態を識別します。私たちは、MDP、間隔 MDP、確率的ゲームなどのさまざまなモデル タイプを使用して、さまざまな因果的 MDP 抽象化を理論的および経験的に比較します。私たちの評価は、私たちのアプローチの可能性を示しています。いくつかの標準ベンチマークについて、元の MDP に対して最適に近いポリシーを計算できる小さな抽象化を取得しました。さらに、私たちの因果的抽象化は、多くの場合、関連する大規模な MDP モデルに一般化されます。

原文 (English)

Property-driven Causal Abstractions for Markov Decision Processes

Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state variables and their valuations. The exponential blowup in the number of states renders many reasoning tasks in MDPs challenging. Abstractions are promising techniques to reduce MDPs and thus mitigate scalability issues. In this work, we introduce a notion of causality on factored MDPs and a novel property-driven causal abstraction technique that retains many characteristics of the original MDP model. For this, we rely on causal relations over state variable predicates and identify those states that share the same reasons for fulfilling or violating a given abstraction property. We theoretically and empirically compare various causal MDP abstractions using different model types such as MDPs, interval MDPs, or stochastic games. Our evaluation demonstrates the potential of our approach: For several standard benchmarks, we obtain small abstractions that allow us to compute near-optimal policies for the original MDP. Furthermore, our causal abstractions often generalize to related large-scale MDP models.

13:00 JSTロボティクス

パッシブビデオから編集可能なエクスペリエンスへ: 身体化されたインテリジェンスのための物理的に接地されたエクスペリエンスの合成

身体化された AI における主なボトルネックは、モデル アーキテクチャではなくデータです。何十億もの人間の操作ビデオがオンラインに存在しますが、人間の形態とロボットのハードウェアの間には具現化のギャップがあるため、ロボットはビデオから直接学習することができません。構造化された知識の伝達を通じて人間のデモンストレーションをロボットが学習可能なデータに変換することで、このギャップを埋める低リソースのフレームワークである Pegasus を紹介します。 Pegasus は、生のビデオ プロンプトに依存する代わりに、グラフベースの中間表現を構築します。人間のビデオから抽出されたタスク グラフは、アフォーダンス グラフと制約グラフを通じて、ロボット条件付きビデオ生成のためのロボット プランニング グラフに変換されます。階層型アフォーダンス潜在空間は、オブジェクトの状態、アフォーダンス、タスク間の関係をモデル化し、オブジェクトのアイデンティティを超えた一般化を可能にします。閉ループ物理検証器は、運動学的実現可能性、衝突制約、関節制限を使用して、無効な世代をさらにフィルタリングします。 GTEA Gaze+ や EPIC-KITCHEN-100 などのさまざまな自己中心的操作ベンチマークとさまざまなロボットの実施形態にわたって Pegasus を評価し、タスクの正確性、実行可能性、状態の一貫性、および学習可能性を評価します。結果は、信頼性の高い実施形態間変換を実証し、ロボットデータ生成をハードウェア収集問題からスケーラブルで低リソースの知識伝達問題に再構成できることを示しています。

原文 (English)

From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource framework that bridges this gap by translating human demonstrations into robot-learnable data through structured knowledge transfer. Instead of relying on raw video prompts, Pegasus constructs a graph-based intermediate representation: a Task Graph extracted from human videos is transformed through Affordance and Constraint Graphs into a Robot Planning Graph for robot-conditioned video generation. A hierarchical affordance latent space models the relationship between object states, affordances, and tasks, enabling generalization beyond object identities. A closed-loop physics verifier further filters invalid generations using kinematic feasibility, collision constraints, and joint limits. We evaluate Pegasus across a range of egocentric manipulation benchmarks, including GTEA Gaze+ and EPIC-KITCHENS-100, and diverse robot embodiments, assessing Task Correctness, Executability, State Consistency, and Learnability. Results demonstrate reliable cross-embodiment translation and show that robot data generation can be reframed from a hardware collection problem into a scalable, low-resource knowledge transfer problem.

13:00 JSTエージェント

AI エージェントを検出するには何が必要ですか?ブラウザ自動化における動作検出のための最小限の機能セット

大規模に導入されたボット検出器は、トラフィックを人間かボットかの二値として扱います。 AI エージェントがブラウザ自動化を通じて Web を閲覧する場合、この仮定は崩れます。このトラフィック クラスは、どちらでもないトラフィック クラスであり、バイナリ分類器が構造的に表すことができません。人間、ボット、AI エージェントを区別する 3 つのクラスの検出フレームワークを提示し、バイナリ対エージェントの混乱が構造的なものであることを示します。バイナリの人間対ボット検出器は、そのラベル空間にエージェント クラスがないため、エージェント セッションのルートを誤ります。私たちの制御されたベンチマークでは、MLP バイナリ分類器は実際の AI エージェントの 39.1% を人間として誤分類し、SAINT バイナリ変換器は 34.5% を誤分類しました。明示的なエージェント クラスを追加すると、30 回の実行すべてでクラスごとのエージェント F1 = 1.000 が得られます (3 つのモデル ファミリ $\times$ 10 シード)。回避耐性を測定するために、受動的な観察、GAN が生成した軌道、実際の人間のカーソル データの再生にまたがる 5 レベルの回避ラダーを構築します ($n = 2299$ 回避セッション)。 10 のシードと 3 つのモデル ファミリにわたって、シードごとの 22990 件の予測でエージェント ミスがゼロであることが観察されました。識別信号はブラウザ自動化のアーティファクトであり、エージェント推論の証拠ではありません。Playwright は、物理的な入力デバイスが生成する生のポインター移動ストリームとホイール デルタ ストリームを発行せず、この不在シグネチャは軌道操作の後に残ります。サイズ 1 ~ 5 (9401 GBM) のすべての特徴サブセットを徹底的に検索すると、2 つの行動特徴 (mouse_event_rate、teleport_click_ratio) により、エージェント精度 0.994 ですべての回避レベルで 100% 観察されたエージェント再現率が得られることがわかります。 5 つの機能により、マクロ F1 が 0.991 に上昇します。信号は冗長にエンコードされています。teleport_click_ratio を削除すると、エージェント検出は 100% のままになります。単一特徴レジームは退化しており、常に「エージェント」を予測するために分類子を折りたたむことによってのみすべてのエージェントにフラグを立てます。 2 つの機能がエージェントを確実に隔離します。 5 つは、マクロ F1 $\geq 0.99$ で 3 つのトラフィック クラスすべてを分離します。

原文 (English)

What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation

Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser automation, a traffic class that is neither and that binary classifiers structurally cannot represent. We present a three-class detection framework distinguishing humans, bots, and AI agents, and show that the binary-vs-agent confusion is architectural: a binary human-vs-bot detector misroutes agent sessions because its label space lacks an agent class. On our controlled benchmark, an MLP binary classifier misclassifies 39.1% of real AI agents as human and a SAINT binary transformer misclassifies 34.5%; adding an explicit agent class yields per-class agent F1 = 1.000 in all 30 runs (3 model families $\times$ 10 seeds). To measure evasion resistance, we construct a five-level evasion ladder spanning passive observation, GAN-generated trajectories, and replay of real human cursor data ($n = 2299$ evasion sessions). Across 10 seeds and 3 model families we observe zero agent misses in 22990 per-seed predictions. The discriminative signal is a browser-automation artifact, not evidence of agent reasoning: Playwright does not emit the raw pointer-move and wheel-delta streams a physical input device produces, and this absence signature survives trajectory manipulation. Exhaustive search over all feature subsets of size 1-5 (9401 GBMs) shows that two behavioral features (mouse_event_rate, teleport_click_ratio) give 100% observed agent recall at every evasion level with agent precision 0.994; five features lift macro-F1 to 0.991. The signal is redundantly encoded: removing teleport_click_ratio leaves agent detection at 100%. The single-feature regime is degenerate, flagging every agent only by collapsing the classifier to always predict "agent". Two features robustly isolate agents; five separate all three traffic classes at macro-F1 $\geq 0.99$.

13:00 JST研究/論文

囲碁における不確実性ゲーティングによる信念に基づく意思決定

AlphaZero と MuZero によって推進された Computer Go の最近の進歩は、ニューラル ネットワーク ポリシーのエラーを修正するためにモンテカルロ ツリー検索 (MCTS) に大きく依存しています。この依存性は、大規模な計算クラスターでは効果的ですが、消費者向けハードウェアでは重大なボトルネックを引き起こし、ツリー管理の計算コストにより推論速度が大幅に制限されます。さらに、これらのモデルは深く探索しないと幻覚に悩まされ、戦略的に致命的な動きを高い確信度で提案します。このペーパーでは、ポリシーのヘッドと別個の信念のヘッドを分離する、新しい信念ガイド型アーキテクチャを紹介します。従来の価値関数とは異なり、信念ヘッドは内部シミュレーターおよび独立した批評家として機能し、認識論的な不確実性と戦略的安定性をモデル化します。長期依存関係と Ko ルールを処理するメモリ メカニズム (Transformer/GRU) を統合し、自信過剰なポリシー エラーをフィルタリングするゲート メカニズムを利用することにより、私たちのモデルはインテリジェンスの負担を実行時の検索からパラメトリックな「直感」に移行します。実験結果は、このアプローチにより、検索不要の勝率が大幅に向上し、幻覚が減少し、大規模な MCTS が実行できない限られたハードウェアでプロ レベルのプレイが可能になることが示されています。

原文 (English)

Belief-Guided Decision Making with Uncertainty Gating in the Game of Go

Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of the neural network policy. While effective on massive computational clusters, this dependence creates a critical bottleneck on consumer-grade hardware, where the computational cost of tree management severely limits inference rates. Furthermore, without deep search, these models suffer from hallucination, proposing moves with high confidence that are strategically fatal. This paper introduces a novel Belief-Guided architecture that disentangles the Policy head from a distinct Belief head. Unlike traditional value functions, the Belief head acts as an internal simulator and independent critic, modeling epistemic uncertainty and strategic stability. By integrating memory mechanisms (Transformer/GRU) to handle long-term dependencies and the Ko rule, and utilizing a gating mechanism to filter overconfident policy errors, our model shifts the burden of intelligence from runtime search to parametric "intuition." Experimental results demonstrate that this approach significantly improves search-free win rates and reduces hallucination, enabling professional-level play on limited hardware where massive MCTS is infeasible.

13:00 JSTLLM/生成AIエージェント研究/論文

Setoka: 異種データに対するパーソナライズされたエージェントにおける階層型ユーザー理解のベンチマーク

幅広いタスクにわたってユーザーを支援するために、パーソナライズされたエージェントがますます適用されています。効果的な個別化された支援には、エージェントの記憶に保存された過去のやり取りから明確な事実を取得するだけでなく、抽象的な個人的特徴を推測することも必要です。しかし、既存のメモリ ベンチマークは主に、エージェントが会話履歴に明示的に記載されている情報を取得できるかどうかを評価するもので、ユーザーのより深い理解を効果的に評価することはできません。この研究では、異種データから階層的なユーザー理解を備え、メモリ拡張されたパーソナライズされたエージェントを評価するためのベンチマークである Setoka を提案します。 Setoka は、認知心理学と性格心理学の理論に基づいて、ユーザー理解の 4 つのレベル、つまり意味記憶、エピソード記憶、行動パターン、性格特性を定義します。さらに、現実的でありながらプライバシーを保護した評価を可能にするために、多様で一貫性のある異種ユーザー データとクエリを大規模に合成する心理測定ベースのパイプラインを設計します。最後に、Setoka を活用して、10 人の合成ユーザーに対して 5 つのメモリ システムと組み合わせた 3 つの言語モデルを評価します。私たちの包括的な評価により、既存のシステムは意味記憶の検索では良好なパフォーマンスを発揮しますが、エピソード記憶ではパフォーマンスが低下することが明らかになりました。さらに、時間の経過とともに分散した異種かつ断片的な情報を統合する必要がある行動パターンや性格特性を理解するタスクを扱う場合、パフォーマンスはさらに低下します。これらの発見は、ユーザーの理解は単純な事実検索では処理できないことを示しており、長期的なユーザー行動に対するソース間の統合と抽象化のためのメモリ メカニズムの設計の動機付けとなります。

原文 (English)

Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data

Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory. Moreover, when dealing with behavior pattern and personality trait understanding tasks that require integrating heterogeneous and fragmented information dispersed over time, performance declines even further. These findings demonstrate that user understanding cannot be handled by simple fact retrieval, motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior.

13:00 JSTLLM/生成AI

LLM の安全性のためのポリシーに基づく蒸留: テンプレートに堅牢な再調整へのルーティング アプローチ

微調整は、大規模言語モデル (LLM) を特殊化するための主要なパラダイムですが、重大な脆弱性をさらしています。悪意のあるデータ プロバイダーは、下流のコーパスに有害な動作を埋め込み、オンデマンドで人間の価値観を侵害しながら、専門的なスキルを保持するモデルを作成する可能性があります。既存の安全再調整防御は、3 つの重要な制限により、実際には失敗することがよくあります。それらは、専門スキルの壊滅的な忘れを引き起こすことがよくあります。防御側が攻撃者のプロンプトテンプレートを観察できない場合、その有効性は崩壊します。また、正常に再調整されたモデルは、単純なシステム プロンプト スイッチによる再脱獄の影響を受けやすいままです。これらの課題に対処するために、特定のプロンプト テンプレートを当てはめるのではなく、調整された出力確率分布と妥協した出力確率分布の間の乖離をモデル化する新しい再調整フレームワークである、ルーティング ベースのオンポリシー蒸留 (ROPD) を提案します。私たちは、さまざまなアライメント強度を持つ 3 つのデータセットと 3 つのベースモデルにわたって、ROPD を 4 つの最先端のベースラインと比較する広範な実験を実施しています。私たちの結果は、ベースラインの防御がテンプレートの不一致に直面すると、多くの場合、下流のタスクのパフォーマンスの深刻な低下を伴うことを示しています。対照的に、ROPD はテンプレート不一致のリスクを大幅に軽減し、防御効果と能力維持の両方において優れた堅牢性を維持します。私たちの分析は、ROPD がテンプレート シフトの影響を完全に受けないわけではないことを示していますが、そのパフォーマンスの低下は既存の方法と比較して無視できる程度であり、堅牢な LLM 再調整の新しい標準を確立しています。

原文 (English)

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers can embed harmful behaviors into downstream corpora, creating models that retain professional skills while violating human values on demand. Existing safety-realignment defenses often fail in practice due to three key limitations: they frequently cause catastrophic forgetting of specialized skills; their effectiveness collapses when the defender cannot observe the attacker's prompt template; and successfully realigned models remain susceptible to re-jailbreaking via simple system prompt switches. To address these challenges, we propose Routing-based On-Policy Distillation (ROPD), a novel realignment framework that models the divergence between aligned and compromised output probability distributions rather than fitting specific prompt templates. We conduct extensive experiments comparing ROPD against four state-of-the-art baselines across three datasets and three base models with varying alignment strengths. Our results demonstrate that when baseline defenses face template mismatches, often accompanied by severe degradation in downstream task performance. In contrast, ROPD substantially mitigates template-mismatch risks, maintaining superior robustness in both defense effectiveness and capability preservation. While our analysis indicates ROPD is not entirely immune to template shifts, its performance degradation is negligible compared to existing methods, establishing a new standard for robust LLM realignment.

13:00 JSTエージェント

AgentMap: オントロジー照合のための結合等価性と包含の検出

オントロジー マッチング (OM) は伝統的に、等価性発見または包含マッチングとして定式化されてきました。既存の OM システムは、1 種類の意味的対応のみを識別し、等価マッピングと包含マッピングを同時に検出することはできません。この論文では、等価性と包含の発見を統合する新しい OM タスクであるハイブリッド オントロジー マッチング (HOM) を紹介し、それに応じて、一連の相互依存するセマンティック決定によって実装される大規模言語モデル (LLM) ベースのマルチエージェント OM フレームワーク AgentMap を提案します。ソース オントロジーの概念が与えられると、AgentMap はセマンティック検索、階層検索、協調的なマルチエージェント LLM 推論を統合して、ターゲット オントロジーを段階的に探索し、同等の概念 (存在する場合) または最も詳細なサブシューマーのいずれかを識別します。 HOM ベンチマーク用に 4 つの OM データセットをさらに拡張し、ハイブリッド、等価性のみ、包含のみの設定で AgentMap を評価します。実験結果は、AgentMap がハイブリッド設定で有望なパフォーマンスを達成すると同時に、等価性のみおよび包含のみの設定でそれぞれ等価性マッチングおよび包含マッチングのベースラインを上回るパフォーマンスを示していることを示しています。

原文 (English)

AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching

Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems identify only one type of semantic correspondence and cannot simultaneously discover equivalence and subsumption mappings. In this paper, we introduce Hybrid Ontology Matching (HOM), a new OM task that unifies equivalence and subsumption discovery, and accordingly propose a Large Language Model (LLM)-based multi-agent OM framework AgentMap that is implemented by a series of interdependent semantic decisions. Given a concept in the source ontology, AgentMap integrates semantic retrieval, hierarchical search, and collaborative multi-agent LLM reasoning to progressively explore the target ontology, identifying either the equivalent concept, if one exists, or the most fine-grained subsumer. We further extend four OM datasets for a HOM benchmark and evaluate AgentMap under hybrid, equivalence-only, and subsumption-only settings. Experimental results show that AgentMap achieves promising performance on the hybrid setting, and at the same time outperforms equivalence matching and subsumption matching baselines on the equivalence-only and subsumption-only settings, respectively.

13:00 JSTLLM/生成AI

LLM 支援言語使用における言語モノカルチャー

執筆とコミュニケーションは、テキストの下書き、修正、磨き上げに使用される大規模言語モデル (LLM) によって媒介されることが増えています。このような支援は明確さを改善し、著者が制度の期待に応えるのに役立ちますが、共有モデルへの広範な依存により、言語形式の集団レベルの変動、つまり言語モノカルチャーと呼ばれる現象が減少する可能性があります。私たちは、著者と LLM が言語特徴にわたる分布として表現され、反復的な相互作用を通じて共進化する数学的フレームワークを開発します。私たちは、固定言語分布を持つ共有モデル、著者の出力から再帰的に更新される共有モデル、著者固有および母集団レベルのフィードバックを通じて更新されるパーソナライズされたモデルの 3 つの相互作用メカニズムを分析します。我々は結果として生じる均衡と収束率を特徴づけ、共有モデルが著者を共通の規範に向けて駆り立てることができ、再帰的フィードバックは共通の適合の下でペアごとの広がりを変えることなく共有規範を再配置し、パーソナライゼーションは非ゼロの言語多様性を持つ一連の異なる著者モデル均衡を維持できることを示した。そして、明瞭さ、読みやすさ、流暢さといった個人的な利益と独特のスタイルをトレードオフする戦略的選択として適合性を内生化します。この実用モデル内では、個人的に合理的な作者は、自分の独自性が他者に提供する価値を内面化していないため、社会的に最適な以上に適合する可能性があり、負の外部性と、固定インスタンスごとに有限であるが、独自性が真正性を支配する場合には際限なく成長するモノカルチャーの代償を生み出す可能性があります。合成シミュレーションは、固定的な共有支援、再帰的なフィードバック、およびパーソナライゼーションがどのように長期的な多様性のさまざまな結果を生み出すかを示しています。

原文 (English)

Linguistic Monoculture in LLM-Assisted Language Use

Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author-model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes.

13:00 JSTLLM/生成AIエージェント研究/論文

OmegaUse-OfficeVal: 経済的根拠に基づいた長期的なオフィススイートのタスクに関する LLM エージェントのベンチマーク

大規模言語モデル (LLM) エージェントは、ユーザーのタスク完了を支援することがますます期待されています。ただし、既存のベンチマークでは、エージェントがオフィス スイートのワークフローを妥当なコストで実行できるかどうかを評価するためのサポートが限定的です。タスク レベルの経済的基盤を備えた長期的なオフィス スイート タスクで LLM エージェントを評価するためのベンチマークである OmegaUse-OfficeVal を紹介します。このベンチマークは、実務者によって提案され、プライバシー保護プロセスを通じて調整されたオフィス スイートの要求から派生した 100 のタスクで構成されています。これらのタスクを完了するには、平均して 2.32 時間の人力が必要です。このベンチマークの重要な特徴は、各タスクが 2 つの経済シグナル (人的労働時間とタスク価格の代理) と組み合わされていることです。これらの信号により、人的コストと LLM 推論コストの直接比較、および価値重み付け評価が可能になります。安定した評価をサポートするために、私たちはきめの細かいルーブリックからコードベースの検証ツールを開発します。私たちは、人間のベースラインとともにいくつかのフロンティア LLM を評価します。評価されたすべての LLM は人間の作業者よりも大幅に安価で高速ですが、人間レベルの成果物の品質にはまだ達していません。コードとデータセットは完全にオープンソースであり、詳細についてはプロジェクト Web サイト (https://omegause-officeval.github.io) でご覧いただけます。

原文 (English)

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.

13:00 JSTエージェント

アドホック チームワークにおけるタスクに依存しない適応のためのパートナーの能力推定

斬新で多様なパートナーとの効果的なコラボレーションは、自律エージェントにとって重要なスキルです。現在のアドホック チームワーク (AHT) アプローチのほとんどは、エージェントが単一の固定タスクに協力し、パートナーの能力、つまり目的のアクションを正常に実行する能力がすでにわかっていることを前提としています。実際には、パートナーの真の能力は隠されていることが多く、人間の協力者は複数の有効な戦略を使用してタスクに対して最適とは言えない行動をとる可能性があります。これらの制限に対処するために、私たちはアドホックなチームワークを、隠れたパートナーの機能の下で分散実行する共同計画の問題として再構成することで、マルチタスク設定に拡張します。タスク不変の能力ベクトルを推論する近似ベイジアン法である CE-CM (Contextual Models による能力推定) を紹介します。シミュレーションベースのサンプリングを使用することで、エージェントは能力を推定し、状況に応じたマルチエージェントのマルコフ意思決定プロセスを計画に導きます。このアプローチでは、母集団の事前トレーニングは必要なく、オンラインでわずか数個のタスクからその信念を洗練させます。人間の予測不可能性を説明するために、単一の最適な軌道ではなく、さまざまなプランナーのロールアウトに対して能力の仮説を評価する拡張機能である CE-CM-Div を提案します。模擬実験では、CE-CM が隠れた機能を迅速に回復し、実行不可能なアクションの割り当てを減らし、時間の経過による変化に適応することが実証されています。さらに、15 人の参加者による 225 の軌跡を対象としたオフラインの人体研究では、CE-CM-Div により、ベースライン CE-CM 法と比較して能力推定値が大幅に向上しました。私たちの結果は、能力ベースのモデリングが、研究された設定において解釈可能でタスクに依存しない表現として有望であることを示唆しており、行動の多様性を考慮することが堅牢な人間と AI のチーム化に不可欠であることを示しています。

原文 (English)

Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork

Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed task and that the partner's capabilities, their ability to successfully execute the desired action, are already known. In reality, a partner's true capabilities are often hidden, and human collaborators may act sub-optimally on tasks with multiple valid strategies. To address these limitations, we extend ad-hoc teamwork into a multi-task setting by re-framing it as a problem of joint planning with decentralised execution under hidden partner capabilities. We introduce CE-CM (Capability Estimation via Contextual Models), an approximate Bayesian method that infers task-invariant capability vectors. By using simulation-based sampling, the agent estimates capabilities and induces a contextual Multi-agent Markov Decision Processes for planning. This approach requires no population pre-training and refines its beliefs online from just a few tasks. To account for human unpredictability, we propose CE-CM-Div, an extension that evaluates capability hypotheses against diverse planner rollouts rather than a single optimal trajectory. Simulated experiments demonstrate that CE-CM rapidly recovers hidden capabilities, reduces infeasible action assignments, and adapts to changes over time. Furthermore, in an offline human study of 225 trajectories from 15 participants, CE-CM-Div substantially improved capability estimates over the baseline CE-CM method. Our results suggest capability-based modelling is a promising interpretable, task-agnostic representation in the studied settings, demonstrating that accounting for behavioural diversity is essential for robust human-AI teaming.

13:00 JSTエージェント研究/論文

AI エージェントはオープンエンドの AI 研究を行うことができますか? 2 つのケーススタディからの初期の証拠

AI の爆発的な進歩の予測は、AI 研究を自動化する AI エージェントにかかっています。しかし、エージェントが無制限の AI 研究を実行できるかどうかについての証拠は乏しい。現在の評価では、限定的で検証可能なタスクでエージェントをテストするか、無制限の研究を除外するか、AI で生成された論文をブラインド査読に提出しますが、これは過剰で確率的であり、レビューの質が低いという問題があります。 AI 研究開発の自動化に向けた進捗状況を測定する 3 番目の方法を紹介します。エージェントは、質の高い未発表論文の中心となる自由回答の研究課題に取り組み、論文の元の著者がその成果を採点します。これらをシャドウ評価と呼びます。私たちは 2 つの未公開の NeurIPS 2026 提出物に対してシャドウ評価を実行し、フロンティア エージェントに 6 日間と数千ドルのコンピューティングを与えました。エージェントは人間の助けなしですべてのエンジニアリングを完了しましたが、研究上の疑問の答えに向けて実質的な進歩を遂げることはできませんでした。その結果、両方の論文は著者によって明確に拒否されました。私たちは、繰り返される 5 つの失敗モードを特定します。それは、出版可能な研究の基準に関する誤った判断、研究設計の欠点に対する創造性のない対応、行き止まりからの非効果的な後戻り、不十分なリソース認識、指示の逸脱です。 2 番目のモデルと足場を使用した堅牢性チェックにより、これらの障害が再現されました。専門家のレビュー、アンケートの回答、エージェントのリポジトリ、およびログを公開します。私たちの結果は、今日のエージェントが AI 研究のエンジニアリングを行うことができるものの、研究ライフサイクルの重要な部分で苦労しているという初期の証拠を提供します。

原文 (English)

Can AI agents conduct open-ended AI research? Early evidence from two case studies

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

13:00 JSTロボティクス研究/論文

ロボットの知識主導型ミッションを設計するための方法論

この論文では、自律ロボットミッションの効率とインテリジェンスの向上を目的として、ROS 2 システムにナレッジ グラフを実装するための包括的な方法論を紹介します。この方法論には、初期条件と目標条件の定義、タスクとサブタスクの構造化、その順序の計画、ナレッジ グラフでのタスク関連データの表現、高級言語を使用したミッションの設計など、いくつかの重要なステップが含まれています。各ステップは前のステップに基づいて構築され、初期セットアップから最終実行まで一貫したプロセスが保証されます。 Aerostack2 フレームワーク内での実用的な実装は、ドローンが自律的にターゲットを特定する Gazebo 環境での捜索救助ミッションのシミュレーションを通じて実証されます。この実装は、ナレッジ グラフを活用することで意思決定とミッションのパフォーマンスを向上させる方法論の有効性を強調しています。

原文 (English)

A Methodology for Designing Knowledge-Driven Missions for Robots

This paper presents a comprehensive methodology for implementing knowledge graphs in ROS 2 systems, aiming to enhance the efficiency and intelligence of autonomous robotic missions. The methodology encompasses several key steps: defining initial and target conditions, structuring tasks and subtasks, planning their sequence, representing task-related data in a knowledge graph, and designing the mission using a high-level language. Each step builds on the previous one to ensure a cohesive process from initial setup to final execution. A practical implementation within the Aerostack2 framework is demonstrated through a simulated search and rescue mission in a Gazebo environment, where drones autonomously locate a target. This implementation highlights the effectiveness of the methodology in improving decision-making and mission performance by leveraging knowledge graphs.

13:00 JST研究/論文

トレーニング前に予測する: 素粒子物理基礎モデルのスケーリング則

素粒子物理学における最大の機械学習モデルは、トレーニングに最もコストがかかりますが、そのコンピューティングを費やす前に、特定のアーキテクチャをスケーリングした場合の利益を見積もることはできません。スケーリングの法則はジェット機に適合していますが、適合しなかったモデルの性能を予測できるものはまだ示されていません。衝突型ジェット機で事前訓練された一般的な変換器の場合、予測できることを示します。モデルとデータの結合スケーリング則を小規模モデルのみに適用し、トレーニング コンピューティングの 3 桁にわたる規模を適用すると、その後 100 倍を超えるコンピューティングでトレーニングされたモデルの損失が 1% 以内に収まると予測します。次に、その予測を下流の物理パフォーマンスに結び付けます。2 つの標準的なタグ付けベンチマーク全体で、事前トレーニング損失が低いほど、系統的に微調整損失が低くなり、微調整後のバックグラウンド除去がより高くなります。したがって、このモデル ファミリとこれらのタスク内では、大規模なモデルがトレーニングされる前に、コンピューティング バジェットを予想される物理パフォーマンスに変換できます。最終的なフロンティア モデルは、精度、AUC、クォーク/グルーオン拒絶に関して、同じコーパスでトレーニングされた現在の最先端の物理認識基盤モデルの公表された数値と一致しており、トップ タグ付けの高純度テールにのみ物理認識モデルの残留エッジが存在します。複数のサイズにわたる 5 つの事前トレーニング済みモデルを、完全なトレーニング レシピとコードとともにリリースします。

原文 (English)

Predict before you train: Scaling Laws for particle physics foundation models

The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architecture cannot be estimated before that compute is spent. Scaling laws have been fit for jets, but none has yet been shown to predict the performance of models it was not fit on. We show that, for a generic transformer pretrained on collider jets, it can be forecast. Fitting a joint model-and-data scaling law on small models alone, spanning three orders of magnitude of training compute, we predict the loss of models trained afterward with more than one hundred times more compute to within one percent. We then connect the forecast to downstream physics performance: across two standard tagging benchmarks, lower pretraining loss yields systematically lower fine-tuning loss and higher background rejection after fine-tuning. Within this model family and these tasks, a compute budget can therefore be translated into expected physics performance before any large model is trained. The final frontier model is consistent with the published numbers for current state-of-the-art physics-aware foundation models trained on the same corpus, on accuracy, AUC, and quark/gluon rejection, with a residual edge for the physics-aware model only in the high-purity tail of top tagging. We release five pretrained models spanning multiple sizes, together with the complete training recipe and code.

13:00 JSTLLM/生成AI画像/動画生成研究/論文Claude

放射線科視覚言語モデルベンチマークの法医学的再現性監査: 意図されたプロトコルからリリースされたアーティファクトまで

医療画像 AI ベンチマークは、データセット、DICOM レンダリング、プロンプト、プロバイダー API、自動ラベル、統計コード、原稿、リポジトリ リリースを組み合わせます。これらの成果物間の一致は、通常、テストされるのではなく想定されます。私たちは、保存された胸部 X 線写真視覚言語モデル (VLM) パイロットの遡及的法医学的再現性監査を実施しました。モデルが再び呼び出されることはなく、画像やレポートに新たに注釈が付けられることもありませんでした。プロンプト バインディング、DICOM メタデータ、出力の完全性、ラベル抽出、一致分析、リリースの伝播を追跡しました。計画された 300 件のモデル プロンプト呼び出しのうち、297 件で空ではないレポートが生成されました。 A/B とラベル付けされた 60 個の Claude 呼び出しが、同じ C プロンプトで実行されました。 30件の研究では28人の患者が対象となった。 4 つの MONOCHROME1 画像は必要な極性反転なしでレンダリングされ、データセット分割メンバーシップは保持されず、未検証の抽出プログラムにより 5 つのレポートが 4000 文字に切り詰められました。 369 の完全な症例発見ブロックからなる 1 つの共通コホートを再構成すると、コクランの Q は 154.73 から 182.29 に変化しました。 45 件のマクネマー比較のうち、27 件は未調整で p < 0.05 であり、20 件はホルム調整後も 0.05 未満のままでした。これらの値は、アーカイブされた自動ラベル マトリックスのみを表します。意図した迅速な比較を回復したり、臨床成績を確立したりすることはありません。当社は、元のパフォーマンス、ランキング、即時効果、および臨床上の主張を撤回し、コホート、DICOM レンダリング、プロンプトおよびモデルのアイデンティティ、コール ステータス、アノテーションの出所、キー付き分析、および派生アーティファクトに対して機械検証可能なコントロールを指定します。

原文 (English)

Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact

Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, and repository releases. Agreement across these artifacts is usually assumed rather than tested. We performed a retrospective forensic reproducibility audit of a preserved chest-radiograph vision-language model (VLM) pilot; no model was called again and no image or report was newly annotated. We traced prompt bindings, DICOM metadata, output completeness, label extraction, matched analyses, and release propagation. Of 300 planned model-prompt calls, 297 yielded nonempty reports. Sixty Claude calls labeled A/B were executed with the same C prompt. The 30 studies represented 28 patients. Four MONOCHROME1 images were rendered without required polarity inversion, dataset split membership was not retained, and the unvalidated extractor truncated five reports to 4000 characters. Reconstructing one common cohort of 369 complete case-finding blocks changed Cochran's Q from 154.73 to 182.29. Of 45 McNemar comparisons, 27 had unadjusted p < 0.05 and 20 remained below 0.05 after Holm adjustment. These values describe only the archived automated-label matrix; they do not recover the intended prompt comparison or establish clinical performance. We withdraw the original performance, ranking, prompt-effect, and clinical claims and specify machine-verifiable controls for cohort, DICOM rendering, prompt and model identity, call status, annotation provenance, keyed analysis, and derived artifacts.

13:00 JSTエージェント

深層強化学習用の凍結ランダム CNN 特徴抽出器における創発的スパース性

私たちは驚くべき現象を報告します。つまり、凍結されたランダムに初期化された CNN 特徴抽出器でトレーニングされた深層強化学習エージェントは、スパース性を引き起こす目的がなくても、非常にスパースな全結合表現を自発的に開発します。最初の全結合層 (FC1, $3{,}136 \to 64$) では、エージェントは決定論的 Pong の場合は 64 個のうち 1 ~ 3 個のニューロン (確率的 Pong の場合は 5 ~ 11 個) のニューロンを介してタスク関連情報を圧縮しますが、訓練可能な CNN は一致した条件下で 55 ~ 64 個のニューロンを活性化します。私たちは 4 つの主な発見を確立します。まず、FC1 ​​のスパース性はタスクの複雑さに応じて変化します。Pong の場合は 1 ~ 11、Breakout の場合は 19 ~ 26、Space Invaders の場合は $\sim$42 です。幅のスケーリングにより、これが固定容量部分ではなくタスク構造を反映していることが確認されます。次に、ゲーム内のスケーリングが現れます。3 つの同一の Pong シードが 5、7、11 個のアクティブなニューロンを生成します。 5 ニューロン シードは $+14$ の報酬で頭打ちになりますが、他のニューロン シードはエキスパート パフォーマンス ($+18.4$、$+18.7$) に達します。これは、ランダム投影の使用可能な次元が達成可能なパフォーマンスの限界を示唆していることを示唆しています。第三に、アブレーションは必要性を裏付けています。これらのアクティブなニューロンを除去すると、2 つの PPO 実装と 4 つのゲームにわたってパフォーマンスがクラッシュします。第 4 に、情報ボトルネックが早期にコミットします。スイープではアクティブ セット ロックが 15 ~ 30M ステップで示されますが、報酬は 35 ~ 105M ステップ後にプラスに変わります。 Breakout での補足的な発見は、凍結された CNN と訓練可能な CNN が構造的に異なるボトルネックを介して競争力のある報酬に達することを示しています。凍結されたエージェントは 17 ~ 25 個のアクティブなニューロン (参加率 $\sim$10-14) を使用するのに対し、訓練可能なエージェントは 51 個 (参加率 $\sim$3.6) を使用します。最後に、入力の次元が本質的なタスクの次元を小さくする場合、凍結されたランダム射影での勾配降下法により、明示的なスパース性機構を使用せずに、根本的な問題の効果的なランクが明らかになる可能性があります。

原文 (English)

Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning

We report a striking phenomenon: deep reinforcement learning agents trained with frozen, randomly initialized CNN feature extractors spontaneously develop extremely sparse fully-connected representations, without any sparsity-inducing objective. In the first fully-connected layer (FC1, $3{,}136 \to 64$), agents compress task-relevant information through as few as 1-3 neurons out of 64 for deterministic Pong (5-11 for stochastic Pong), while trainable CNNs activate 55-64 neurons under matched conditions. We establish four principal findings. First, FC1 sparsity scales with task complexity: 1-11 for Pong, 19-26 for Breakout, and $\sim$42 for Space Invaders. Width-scaling confirms this reflects task structure rather than a fixed capacity fraction. Second, within-game scaling emerges: three identical Pong seeds produce 5, 7, and 11 active neurons. The 5-neuron seed plateaus at $+14$ reward, while the others reach expert performance ($+18.4$, $+18.7$), suggesting the random projection's usable dimensionality bounds achievable performance. Third, ablation confirms necessity: removing these active neurons crashes performance across two PPO implementations and four games. Fourth, the information bottleneck commits early: a sweep shows the active set locks by 15-30M steps, while reward turns positive 35-105M steps later. A complementary finding in Breakout shows frozen and trainable CNNs reach competitive rewards via structurally different bottlenecks: frozen agents use 17-25 active neurons (participation ratio $\sim$10-14), while trainable agents use 51 (participation ratio $\sim$3.6). Finally, wherever input dimensionality dwarfs intrinsic task dimensionality, gradient descent on a frozen random projection may reveal the effective rank of the underlying problem without explicit sparsity machinery.

13:00 JSTLLM/生成AI

お客様のデジタルツインシミュレーションによる大規模なChatBot検証

LLM ベースのチャットボットは、銀行などの規制対象分野での顧客サービスを変革していますが、スケーラブルでコスト効率の高い検証が、安全な導入にとって依然として重要な障壁となっています。大規模なチャットボット検証に関する 2 部構成の寄稿を紹介します。まず、実際のトランザクション データと会話データに基づいた、デジタル ツインとして高忠実度の合成顧客エージェント (SCA) を作成する方法論を紹介します。これにより、多様な顧客プロファイルと対話スタイルをシミュレートするための自動生成と行動調整が可能になります。評価の結果、SCA は実際の顧客との高度な意味的一致、低い幻覚率、および制御可能な介入による成功した性格特性の再現を達成していることが実証されています。 2 番目に、自動化された LLM-as-a-Judge 評価、人間による専門家テスト、および敵対的調査を組み合わせた SCA ベースの検証フレームワークを開発します。感情状態、人口統計グループ、言語的要因にわたるシナリオベースの検証により、堅実なパフォーマンスが確認されます。私たちのアプローチは、英国の大手銀行で顧客対応チャットボットを検証するために使用され、金融機関に規制遵守に向けたスケーラブルな道筋を提供しました。

原文 (English)

Large-Scale ChatBot Validation Through Customer Digital Twin Simulations

LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. We present a two-part contribution for large-scale chatbot validation. First, we introduce a methodology for creating high-fidelity synthetic customer agents (SCAs) as digital twins, grounded in real transactional and conversational data, that enables automatic generation and behavioral conditioning to simulate diverse customer profiles and interaction styles. Evaluation demonstrates that SCAs achieve high semantic alignment with real customers, low hallucination rates, and successful personality trait reproduction with controllable interventions. Second, we develop an SCA-based validation framework combining automated LLM-as-a-Judge evaluation, human expert testing, and adversarial probing. Scenario-based validation across emotional states, demographic groups, and linguistic factors confirms robust performance. Our approach was used to validate a customer facing chatbot at a leading UK bank, providing financial institutions with a scalable pathway toward regulatory compliance.

13:00 JST研究/論文

Sim2Win: チームに依存しない、イベントベースの試合前結果予測およびフットボール用の戦術プロファイリング システム

プロサッカーにおける試合前の戦術的意思決定は、主観的な専門家の分析とアイデンティティに基づいたスカウティングシステムに大きく依存しており、目に見えないチームに一般化することはできません。この論文では、試合結果の予測を戦術的意思決定支援の問題として再構成する、チームに依存しないイベントベースの試合前戦術推奨フレームワークである Sim2Win について説明します。 Sim2Win は、178 チームと 1,411 のチーム試合記録にわたる 11 の競技会からの StatsBomb オープン イベント データを使用して、5 試合のローリング戦術プロファイルを構築し、解釈可能な 4 つの戦術特徴率を設計し、K 平均法によってチームの行動を 8 つのプレイスタイルにクラスター化し、戦術的な対戦表現から勝ち、引き分け、負けの確率を推定するために 13 の分類子をトレーニングします。このシステムはチーム名や ID 機能なしで動作するため、トレーニング中に見たことのないチームへの一般化が可能になります。厳密な Leave-One-Competition-Out (LOCO) 評価により、Sim2Win はまったく未知のチームで平均 ROC-AUC 0.704 と平均精度 55.4% を達成し、21 件の ROC-AUC 比較すべてと 21 件中 19 件の精度比較で ELO、Pi-Rating、GAP ベースラインを上回っていることが実証されています。すべての評価モデルの中で、CatBoost は 60.90% の精度で最も強力な配信パフォーマンスを達成しました。これらの発見は、行動戦術表現が分布の変化の下で伝達可能な予測信号を提供し、アイデンティティ依存のサッカー予測システムに代わる実行可能な代替手段を提供することを示唆しています。

原文 (English)

Sim2Win: A Team-Agnostic, Event-Based Pre-Match Outcome Prediction and Tactical Profiling System for Football

Pre-match tactical decision-making in professional football relies heavily on subjective expert analysis and identity-based scouting systems that cannot generalize to unseen teams. This paper presents Sim2Win, a team-agnostic, event-based pre-match tactical recommendation framework that reframes match outcome prediction as a tactical decision-support problem. Using StatsBomb open event data from eleven competitions spanning 178 teams and 1,411 team-match records, Sim2Win constructs five-match rolling tactical profiles, engineers four interpretable tactical feature ratios, clusters team behaviors into eight playstyles via K-Means, and trains thirteen classifiers to estimate win, draw, and loss probabilities from tactical matchup representations. The system operates without team names or identity features, enabling generalization to teams never seen during training. A rigorous Leave-One-Competition-Out (LOCO) evaluation demonstrates that Sim2Win achieves a mean ROC-AUC of 0.704 and mean accuracy of 55.4% on completely unseen teams, outperforming ELO, Pi-Rating, and GAP baselines on all 21 ROC-AUC comparisons and 19 of 21 accuracy comparisons. Among all evaluated models, CatBoost achieved the strongest in-distribution performance with 60.90% accuracy. These findings suggest that behavioral tactical representations provide transferable predictive signal under distribution shift and offer a viable alternative to identity-dependent football prediction systems.

LLM ベースのチャット AI における知的障害者に対する暗黙のバイアスを特定する

背景: この研究では、知的障害 (ID) を持つ人々を対象とした大規模言語モデル (LLM) ベースのチャット AI モデルにおける暗黙のバイアスの存在を調査しています。目的: この研究は、ID を持つ人々に関連する表現上の差異を特定および測定し、それらを調査して AI チャット生成テクノロジーに固有の暗黙的なバイアスを特定することを目的としています。方法: GPT-4-Turbo モデルを利用して、ID の記述子の有無にかかわらず 10 個のプロンプト ステムに基づいてストーリー生成を要求しました。このプロセスは、他の 4 つの LLM (OpenAI GPT-4o、Meta Llama-3-3-70B-Instruct、Anthropic Claude-3-5-Sonnet、および Mistral-Large-2411) を使用して繰り返されました。結果として得られた 25,000 のコンピューター生成ストーリーは、別の GPT-4-Turbo モデル インスタンスを使用して分析され、以前の文献で説明されている偏見のテーマに関連する人々の表現方法の違いが検出されました。結果: 私たちの調査結果では、ストーリー データセットに ID 記述子がある場合とない場合で、人物がどのように表現されるかに違いがあることが明らかになりました。これらの違いは、ID の確立された特性を超えており、ほとんどが否定的な暗黙のバイアスの存在を示唆しています。パターナリズムと幼児化をテーマに、IDを持つ人を若いとみなすことに関連する相違点を特定。それらをよりインスピレーションを与える、象徴的なものとして描きます。より頻繁に助けを必要とする、依存する、救われるなど。そして、それらに対して否定的な認識を持ち、それらを含めることをさらに躊躇します。結論: これらの暗黙のバイアスは、ID を持つ人々に対する過去の差別の文脈の中で考慮されており、AI 開発における ID を持つ人々に対する暗黙のバイアスに対して熱心に対処する必要があることが強調されています。この研究は、将来の社会的危害を防ぐために、意思決定テクノロジーにおける暗黙のバイアスを評価し、軽減することの重要性を強調しています。

原文 (English)

Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities

Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intellectual disabilities (ID). Objective: The study aims to identify and measure representational differences related to people with ID and examine them to identify implicit biases inherent in AI chat generation technologies. Methods: Utilizing the GPT-4-Turbo model, we requested story-generation based on 10 prompt stems with and without descriptors for ID. This process was repeated using four other LLMs (OpenAI GPT-4o, Meta Llama-3-3-70B-Instruct, Anthropic Claude-3-5-Sonnet, and Mistral-Large-2411). The resulting 25,000 computer-generated stories were analyzed using a separate GPT-4-Turbo model instance to detect differences in how people are represented related to themes of bias described in previous literature. Results: Our findings reveal differences in how people are represented between story datasets with and without ID descriptors. These differences go beyond established characteristics of ID and imply the presence of mostly negative implicit biases. Identified differences related to considering people with ID as younger, with themes of paternalism and infantilization; depicting them as more inspirational and symbolic; as needing help more often, being dependent, and being saved; and having a negative perception of them and more hesitation to include them. Conclusions: These implicit biases are considered within the context of past discrimination towards people with ID and highlight the need for diligence against implicit bias towards people with ID in AI development. This research underscores the importance of assessing and mitigating implicit bias in decision-making technologies to prevent future societal harm.

13:00 JST研究/論文

原型か能力か?学生の数学的能力をモデル化するためのクラスタリング

パーソナライズされた学習システムでは、数学的能力は個別の能力の組み合わせであり、最初に基礎的な能力を獲得することに依存して逐次的に獲得されると想定していることが多く、学生は異なる強みを報告することがよくあります。この研究では、英国で実施され、プラットフォームによって収集された 13 回の国家レベルの試験にわたる 119,034 人の学生の大規模なデータセットにクラスタリング手法を適用することにより、これらの仮定の妥当性を調査します。質問の結果を合格または不合格として分類し、ベルヌーイ混合モデルを使用して、離散的なスキルセットを示す潜在的な母集団を検索します。データには個別のクラスターがほとんど存在せず、支配的な要因は生徒の全体的な能力であることがわかりました。これは、結果として得られるクラスターの確率分布間の高度な線形相関によってさらに裏付けられます。当社の最もパフォーマンスの高いモデルは 78% の精度を達成しており、文献に記載されているより複雑なモデルに匹敵すると同時に、より説明しやすくなっています。このモデルのパフォーマンスをロジスティック回帰ベースラインおよび k 最近傍と比較すると、個々の質問のパフォーマンスを特徴として使用すると、わずかな改善が見られます。これは、全体的な能力レベルが成績を予測するための主要な要素である一方、生徒の正確な強みに合わせてカスタマイズすることで、さらに小さなカスタマイズを改善できるが、生徒がトピック間で大きく異なる能力を身につけるようには見えないことを示唆しています。私たちの研究は、教育における機械学習の全国規模のテストを提供し、説明可能なモデルが競争力のあるパフォーマンスにどのように到達できるかを実証する、この分野の新しいベンチマークを提供します。

原文 (English)

Archetypes or ability? Clustering for modelling student mathematical competence

Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent upon first acquiring foundational abilities, and students often report different strengths. In this work, we explore the validity of these assumptions by applying clustering methods to a large dataset of 119,034 students, spanning 13 national-level exams sat in the United Kingdom and collected by the platform. Classifying question results as pass or fail, we use a Bernoulli Mixture Model to search for latent populations which would be indicative of discrete skill-sets. We find that few distinct clusters are present in the data and that the dominant factor is overall student ability, which is further supported by the high degree of linear correlation between the probability distributions of the resulting clusters. Our best performing model achieves an accuracy of 78 percent, competitive with more complicated models in the literature whilst being more explainable. Comparing this models performance with logistic regression baselines and with k-nearest neighbours, we find a small improvement when using performance on each individual question as features. This suggests that whilst overall ability level is the dominant factor for predicting performance, small further personalisation improvements can be made by tailoring to a students exact strengths, but that students do not appear to develop strongly differing ability across topics. Our work offers a national scale test of machine learning in education and offers a new benchmark for the field, demonstrating how explainable models can reach competitive performance

13:00 JSTエージェント研究/論文

AI エージェントの時代には、信頼できる科学を維持するための新しい科学パラダイムが必要

AI システムは、仮説を生成し、実験を計画し、人間の監視を超えた規模で発見を生み出す自律的な研究エージェントになりつつあります。 ML 会場への提出物の増加に見られるように、科学的成果とそれをチェックする私たちの能力との間の検証ギャップはすでに拡大しており、人間とエージェントの非対称性を考えると、自律エージェントはそれをさらに悪化させます。私たちは、これまでの査読と同様に、科学も検証インフラを進化させる必要があると主張します。ただし、歴史的な適応では、尋問され制裁される可能性のある人間の貢献者が想定されていましたが、AI エージェントはこの想定を打ち破ります。私たちは、デフォルトで監視可能なワークフロー、スケーラブルな検証、および明確な帰属を重視する、適応された検証インフラストラクチャの基準を提案します。私たちは、適応がなければ、機械学習やエージェントを使用するあらゆる科学分野は危険な失敗に直面すると主張します。つまり、誰も検証できない実験結果、理解よりも指標の最適化、そして科学的信頼を損なう説明責任の真空です。

原文 (English)

The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science

AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond human oversight. As seen by increased submissions to ML venues, the verification gap between scientific output and our ability to check it is already widening, and autonomous agents make it worse by magnitudes given human-agent asymmetry. We argue that science must evolve its verification infrastructure, as it has before with peer review. However, while historical adaptations assumed human contributors who could be questioned and sanctioned, AI agents break this assumption. We propose criteria for an adapted verification infrastructure that emphasizes observable-by-default workflows, scalable verification, and clear attribution. We argue that without adaptation, ML and any scientific domain using agents face dangerous failures: experimental results that no person can verify, optimization for metrics over understanding, and accountability vacuums that erode scientific trust.

13:00 JSTLLM/生成AI研究/論文

方法は主張を裏付けていますか?査読のための論文内検証

科学的投稿の量が増えているため、査読を支援するために大規模言語モデル (LLM) を使用することへの関心が高まっています。既存の自動新規性評価アプローチは通常、論文の主張する貢献を先行文献と比較し、これらの貢献が研究自体の中で正確に実現されていると暗黙的に仮定します。しかし、人間の査読者は、同様のアイデアがすでに存在しているからではなく、論文で提示された方法論的証拠がそれらを適切に裏付けていないために、新規性の主張に異議を唱えることがよくあります。主張される貢献と方法論的実現の間のこの内部不一致は、現在の LLM ベースのレビュー システムではほとんど検査されません。このギャップに対処するために、論文内主張検証を導入します。これは、論文で明確に表現された新規性主張が、それを実現するために使用された方法によって実証されているかどうかを評価するフレームワークです。このフレームワークは、LLM を使用して、序論から新規性の主張を抽出し、主張に関連する方法論的な証拠を取得し、その方法が述べられた貢献を実証しているかどうかを評価します。評価は、182 件の ICLR 2025 論文から収集された人間による査読から帰納的に導出された査読者の影響を受けた評価基準に基づいて行われます。これらの基準は、新規性、方法論、明確性、その他の問題に関連する査読者の繰り返しの懸念を捉え、クレームの実証に関する構造化された査読者スタイルの評価を生成するために使用されます。私たちは、LLM が生成したレビューコメントと、受理された論文と拒否された論文のバランスの取れたサブセットに対する人間の査読者の懸念を比較することによって、フレームワークを評価します。人間による評価は、特に新規性関連の問題に関して、フレームワークによって生成された評価と人間のレビュー担当者の懸念との間の重要な一致を示しています。 BERTScore はさらに、対応する人間と LLM のレビューペアを不一致の対照から区別し、このフレームワークが人間のレビュー担当者の観察と一致する懸念を捉えていることを示しています。

原文 (English)

Do Methods Support the Claims? Intra-Paper Verification for Peer Review

The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review. Existing automated novelty assessment approaches typically compare a paper's claimed contributions against prior literature, implicitly assuming that these contributions are accurately realized in the work itself. Human reviewers, however, frequently challenge novelty claims not because similar ideas already exist, but because the methodological evidence presented in the paper does not adequately support them. This internal mismatch between claimed contributions and methodological realization is rarely examined by current LLM-based review systems. To address this gap, we introduce intra-paper claim verification, a framework that evaluates whether novelty claims articulated in a paper are substantiated by the methods used to realize them. The framework employs an LLM to extract novelty claims from the introduction, retrieve claim-relevant methodological evidence, and assess whether the methods substantiate the stated contributions. Assessment is guided by reviewer-inspired evaluation criteria derived inductively from human peer reviews collected from 182 ICLR 2025 papers. These criteria capture recurring reviewer concerns related to novelty, methodology, clarity, and other issues and are used to generate structured reviewer-style assessments of claim substantiation. We evaluate the framework by comparing LLM-generated review comments against human reviewer concerns on a balanced subset of accepted and rejected papers. Human evaluation demonstrates significant alignment between framework-generated assessments and human reviewer concerns, particularly for novelty-related issues. BERTScore further distinguishes corresponding human-LLM review pairs from mismatched controls, indicating that the framework captures concerns consistent with human reviewer observations.

13:00 JSTLLM/生成AI

簡単な罠: LLM が誤解に基づく困難さを過小評価する理由

大規模言語モデル (LLM) は、教育評価における項目の難易度を推定するためにますます使用されています。ただし、そのような推定が学習者が実際にどのように困難を経験しているかを反映しているかどうかは依然として不明です。この研究では、LLM によって生成された難易度評価と、基礎的な数学タスクにおける経験的な生徒の成績との整合性を調査しています。広く使用されている 4 つの LLM ベースのシステムは、複数の実行にわたって 32 の算術項目について 1 ~ 100 のスケールで難易度評価を生成しました (N = 640 評価)。これらは、古典的テスト理論 (CTT) と項目応答理論 (2PL) を使用して、インドネシアの学部生 770 人の回答から導き出された経験的な難易度と比較されました。結果は中程度のランク相関 (スピアマンの rho = 0.52 ~ 0.70) を示し、LLM がアイテムの難易度の大まかな順序を捉えていることを示しています。ただし、端数項目では大幅かつ体系的な不整合が発生します。 LLM によって一貫して簡単と評価されているいくつかの項目は、100 : 1/2 で正解率が 34.16% しかなかった項目など、学生にとって最も難しい項目の 1 つでした。私たちは、LLM は、学習者の誤解によって引き起こされる認知的な難しさではなく、カリキュラムの難しさ、または指導順序に基づいて簡単であるべきものを近似していると主張します。これは、誤解に基づく項目の体系的な過小評価につながり、これを私たちはイージートラップと呼んでいます。これらの発見は、LLM ベースの難易度推定の重大な限界を浮き彫りにし、経験的根拠なしにそのような推定に依存すると、評価設計と適応システムにバイアスが生じる可能性があることを示唆しています。

原文 (English)

The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty

Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment between LLM-generated difficulty ratings and empirical student performance on basic mathematics tasks. Four widely used LLM-based systems generated difficulty ratings on a 1-100 scale for 32 arithmetic items across multiple runs (N = 640 ratings). These were compared with empirical difficulty derived from responses of 770 Indonesian undergraduates using Classical Test Theory (CTT) and Item Response Theory (2PL). Results show moderate rank correlations (Spearman's rho = 0.52-0.70), indicating that LLMs capture coarse ordering of item difficulty. However, substantial and systematic misalignment emerges in fraction items. Several items consistently rated as easy by LLMs were among the most difficult for students, such as an item with only 34.16% correct for 100 : 1/2. We argue that LLMs approximate curricular difficulty, or what should be easy based on instructional sequencing, rather than cognitive difficulty driven by learner misconceptions. This leads to systematic underestimation of misconception-driven items, a phenomenon we term the Easy Trap. These findings highlight a critical limitation of LLM-based difficulty estimation and suggest that relying on such estimates without empirical grounding may introduce bias in assessment design and adaptive systems.

13:00 JST規制/政策

人間の効用係数: AI ガバナンスを制約付き最適化問題として再構成する計算可能な福利厚生指標

EU AI 法や NIST AI RMF などの既存の AI ガバナンスの枠組みは、安全性、透明性、説明責任に取り組んでいますが、マクロ社会経済の安定性に対する定量的な制約を運用するものではありません。その結果、AI システムは規制要件を満たしながらも、労働力の移転、不平等の拡大、経済的回復力の低下につながる可能性があります。私たちは、自動化の深さ、再分配の強度、雇用範囲という 3 つの実用的な政策手段の関数として、エージェンシー、福祉、経済的安定の間の相互作用をモデル化する微分可能な福利厚生指標である人間効用係数 (HUF) を導入します。 HUF は、閉じた形式の最適な自動化レベルと、それを下回ると福利厚生にプラスの自動化レベルが存在しない最小再配分しきい値を生成し、高レベルのガバナンス目標を計算可能な制約に変換します。米国、カナダ、北欧の政策体制にわたる 3 エージェントのマルチエージェント強化学習フレームワークを使用して HUF を評価します。分析ベースのエージェントと PPO ベースのエージェントは両方とも、福利厚生に最適な動作領域を特定し、重大な失敗モードを明らかにします。つまり、再分配を明示的に制約しない福利厚生指標は、意図された社会的目的を損なう一方で指標を満たす高度な自動化均衡に収束する可能性があります。私たちの結果は、AI ガバナンスが基本的にはコンプライアンスの実践ではなく、制約付きの最適化問題であることを示唆しています。 HUF は、自動化ポリシーを評価し、社会経済的安定性の境界を特定し、AI 導入が加速する中でのガバナンスの決定をサポートするための定量的なフレームワークを提供します。

原文 (English)

The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem

Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize quantitative constraints on macro-socioeconomic stability. As a result, AI systems may satisfy regulatory requirements while contributing to labor displacement, rising inequality, and reduced economic resilience. We introduce the Human Utility Factor (HUF), a differentiable welfare metric that models the interaction between Agency, Wellbeing, and Economic Stability as functions of three actionable policy levers: automation depth, redistribution intensity, and employment coverage. HUF yields a closed-form optimal automation level and a minimum redistribution threshold below which no level of automation is welfare-positive, transforming high-level governance objectives into computable constraints. We evaluate HUF using a three-agent multi-agent reinforcement learning framework across U.S., Canadian, and Nordic policy regimes. Both analytical and PPO-based agents identify welfare-optimal operating regions and reveal a critical failure mode: welfare metrics that do not explicitly constrain redistribution can converge to high-automation equilibria that satisfy the metric while undermining its intended societal objectives. Our results suggest that AI governance is fundamentally a constrained optimization problem rather than a compliance exercise. HUF provides a quantitative framework for evaluating automation policies, identifying socioeconomic stability boundaries, and supporting governance decisions under accelerating AI deployment.

13:00 JST研究/論文

AI セキュリティの優先事項: 分野全体の課題

AI システムが経済、政府、国家安全保障の重要な機能に急速に統合されるにつれて、AI の導入と AI セキュリティの準備の間のギャップは拡大し続けています。このペーパーでは、業界、政府、市民社会のリーダーとの構造化されたインタビューに基づいて情報を得て、複数部門の専門家ワークショップを通じて洗練された、AI セキュリティーを推進するための優先課題を提示します。参加者は、最先端の AI システムとその基盤となるインフラストラクチャの保護から、AI が脅威の状況を再構築する中でのサイバーセキュリティ慣行の改善に至るまで、進歩によって AI セキュリティを強化できる最も重要かつ最も費用対効果の高い分野を特定し、ランク付けしました。その結果として得られた優先事項は、戦略的基盤と政策枠組みの確立という 4 つのテーマにわたって整理されています。官民連携と制度的インフラの推進。技術的なセキュリティエンジニアリングと保証を進歩させる。敵対的な圧力の下でエージェント AI を管理します。各優先分野について、専門著者が詳細な分析を提供し、問題を定義し、現在の状況を評価し、セクターを超えた利害関係者が追求できる実行可能なプロジェクトを特定します。この文書は、AI セキュリティ分野全体にわたる調整された投資と行動のための最初の実践的な基盤として機能することを目的としています。これは、さまざまな強みや能力に適した具体的で影響力の高い貢献を特定することで、現在の実務家とこの分野への参入を検討している個人や組織の両方に役立つように設計されています。

原文 (English)

AI Security Priorities: A Field-Wide Agenda

As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and AI security readiness continues to widen. This paper presents a prioritized agenda for advancing AI security, informed by structured interviews with leaders across industry, government, and civil society, and refined through a multi-sector expert workshop. Participants identified and ranked the highest-importance and most cost-effective areas where progress could strengthen AI security - from protecting frontier AI systems and their underlying infrastructure to improving cybersecurity practices as AI reshapes the threat landscape. The resulting priorities are organized across four themes: establishing strategic foundations and policy frameworks; advancing public-private coordination and institutional infrastructure; advancing technical security engineering and assurance; and governing agentic AI under adversarial pressure. For each priority area, expert authors provide detailed analyses that define the problem, assess the current landscape, and identify actionable projects that stakeholders across sectors can pursue. The paper aims to serve as an initial practical foundation for coordinated investment and action across the AI security field. It is designed to serve both current practitioners and individuals and organizations looking to enter the field by identifying concrete, high-impact contributions suited to a range of strengths and capacities.

13:00 JSTLLM/生成AIエージェント

SimpleWikiSearch: エージェント検索のためのクリーンなオフライン Wikipedia 環境

大規模言語モデル (LLM) ベースのエージェント検索システムは、基盤となる LLM が重要な唯一のコンポーネントであるかのように評価されることがよくありますが、その測定されたパフォーマンスは、周囲の検索環境 (Wikipedia スナップショット、前処理パイプライン、チャンキング ポリシー、取得バックエンド、ツール スキーマ、観察形式、回答送信ルール) にも依存します。これらの詳細は十分に指定されていないことが多く、結果の比較や報告されたベースラインの再現が困難になります。私たちは、コーパス構築、検索スタック、ツールコントラクト、および評価プロトコルが明示的で実行可能な SimpleWikiSearch を紹介します。この環境は、英語版 Wikipedia の完全なダンプから始まり、コーパスをクリーンアップしてチャンク化し、キーワードと高密度の検索インデックスを構築し、\texttt{search}、\texttt{open\_url}、\texttt{submit\_answer} で構成される最小限のツール インターフェイスを公開します。オープンソースの LLM を使用して 6 つの QA データセットのベースライン結果を報告し、クローズドソースの商用モデルとの比較のためにランダムな 300 のサブセットを提供します。 SimpleWikiSearch は、ドメイン固有のエージェント ハーネスと、再現可能なエージェント検索評価のための制御されたオフライン環境を提供します。その貢献は、新しいエージェント アルゴリズムではなく、この指定された参照設定です。コードとデータは https://github.com/JimXiongGM/simple_wiki_search から入手できます。

原文 (English)

SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search

Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters, yet their measured performance also depends on the surrounding search environment: the Wikipedia snapshot, preprocessing pipeline, chunking policy, retrieval backend, tool schema, observation format, and answer submission rule. These details are frequently under-specified, making it difficult to compare results or reproduce reported baselines. We present SimpleWikiSearch, whose corpus construction, retrieval stack, tool contract, and evaluation protocol are explicit and runnable. The environment starts from a full English Wikipedia dump, cleans and chunks the corpus, builds keyword and dense retrieval indexes, and exposes a minimal tool interface consisting of \texttt{search}, \texttt{open\_url}, and \texttt{submit\_answer}. We report baseline results on six QA datasets using open-source LLMs and provide a random-300 subset for comparisons with closed-source commercial models. SimpleWikiSearch provides a domain-specific agent harness and a controlled offline environment for reproducible agentic-search evaluation. Its contribution is this specified reference setup, rather than a new agent algorithm. Code and data will be available at: https://github.com/JimXiongGM/simple_wiki_search.

13:00 JST研究/論文

GuidedRAG: 検索拡張生成のセマンティック ステアリング

この研究では、従来の検索拡張生成 (RAG) の新しい拡張である GuidedRAG を提案します。GuidedRAG は、検索中に専用の選択ステージとセマンティック ステアリングを導入します。ますます複雑化する検索と知識構造に依存する現在の最先端の RAG アプローチとは対照的に、GuidedRAG は検索前にセマンティクスを使用して知識ベースを制約し、検索スペースを大幅に削減しながら検索スペースをユーザーの意図に合わせます。私たちの評価によると、GuidedRAG は検索の関連性を 14.0 ~ 15.8% 向上させ、検索精度の 19.7 ~ 27.4% の損失を軽減し、検索のオーバーヘッドを桁違いに削減します。さらに、関連するチャンクはランキング プロセスの早い段階で一貫して取得され、ユーザーの意図との整合性が 31.8 ~ 36.8% 向上します。さらに、GuidedRAG が 15 の多様な RAG バリアントを完全にカバーし、文献全体にわたる一般化可能性を実証していることを示します。これらの発見を総合すると、RAG の現在の最先端を改善するための強力で一般化可能なパラダイムとして、セマンティック ステアリングと選択が確立されます。

原文 (English)

GuidedRAG: Semantic Steering of Retrieval-Augmented Generation

In this work, we propose GuidedRAG, a novel extension to traditional Retrieval-Augmented Generation (RAG) that introduces a dedicated selection stage and semantic steering during retrieval. In contrast to current state-of-the-art RAG approaches, which depend on increasingly complex retrieval and knowledge structures, GuidedRAG constrains the knowledge base using semantics before retrieval, aligning the retrieval space with user intent while substantially reducing the search space. Our evaluation shows that GuidedRAG improves retrieval relevance by 14.0-15.8%, mitigates a 19.7-27.4% loss in retrieval precision, and reduces retrieval overhead by orders of magnitude. Moreover, relevant chunks are consistently retrieved earlier in the ranking process, while alignment with user intent improves by 31.8-36.8%. We further show that GuidedRAG achieves full coverage across 15 diverse RAG variants, demonstrating generalizability across the literature. Together, these findings establish semantic steering and selections as a powerful and generalizable paradigm for improving the current state-of-the-art in RAG.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

IFCMemoryBench: BIM 情報取得における LLM ベースのエージェントの長期メモリの評価

長期記憶は LLM ベースのエージェントの中核機能になりつつありますが、既存の評価では主に、オープンドメインまたはペルソナに基づいた設定での会話の想起がテストされています。私たちは、より強力なテストは、エージェントがライブで構造化されたドメイン固有の環境上で動作しながら、以前のセッションからの情報を再利用できるかどうかであると主張します。私たちは、ビルディング インフォメーション モデリング (BIM) でこの問題を研究しています。これは、エージェントが大規模な IFC モデルにクエリを実行する一方で、会話ではよく議論されるもののモデルには含まれていないエンジニアリング規約にも依存しながら、エージェントが大規模な IFC モデルにクエリを実行する必要がある、ビルディング インフォメーション モデリング (BIM) でこの問題を研究しています。 LLM ベースの BIM 情報検索における長期記憶を評価するためのベンチマークである IFCMemoryBench を紹介します。 IFCMemoryBench には、19 のプロジェクトにわたる 143 のマルチセッション タスクと、IFC-Bench v2 の不完全な情報の質問から派生した 4,016 の以前のセッションが含まれています。各タスクは、以前の会話全体で不足しているプロジェクト コンテキストをシードし、後で、記憶されているコンテキストとライブ IFC クエリを組み合わせることによってのみ回答できる調査用の質問をします。当社の評価フレームワークは、メモリのパフォーマンスを取り込み、取得、利用に分解し、専門家が検証した LLM 審査員によって回答の品質とメモリの品質の両方を測定します。代表的なベクトル、グラフ、ファイルベースのメモリ システムを評価します。最も強力なシステムは、デプロイメントに現実的な取り込みスコープの下ではわずか 32.4% の応答精度しか達成できず、オラクルでフィルタリングされた取り込みまたはより強力なプローブ エージェントの下では 60% 未満のままです。分析によると、現在の汎用記憶システムは、トピックに関連するコンテキストを取得することが多いものの、プロジェクトの知識を不完全または断片的な事実として保存していることがわかります。これらの結果は、エージェントの記憶におけるドメイン転送ギャップを明らかにし、信頼できる専門エージェントには、会話、プロジェクトの知識、構造化モデル エンティティをリンクするドメインを認識した記憶表現が必要であることを示唆しています。

原文 (English)

IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval

Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-domain or persona-grounded settings. We argue that a stronger test is whether an agent can reuse information from prior sessions while acting over a live, structured, domain-specific environment. We study this problem in Building Information Modelling (BIM), a professional engineering workflow where agents must query large IFC models while also relying on project specifications, client decisions, and engineering conventions often discussed in conversation but absent from the model. We introduce IFCMemoryBench, a benchmark for evaluating long-term memory in LLM-based BIM information retrieval. IFCMemoryBench contains 143 multi-session tasks across 19 projects and 4,016 prior sessions, derived from incomplete-information questions in IFC-Bench v2. Each task seeds missing project context across earlier conversations and later asks a probe question that can be answered only by combining remembered context with live IFC queries. Our evaluation framework decomposes memory performance into ingestion, retrieval, and utilization, and measures both answer quality and memory quality with expert-validated LLM judges. We evaluate representative vector-, graph-, and file-based memory systems. The strongest system achieves only 32.4% answer accuracy under a deployment-realistic ingestion scope, and remains below 60% under oracle-filtered ingestion or a stronger probe agent. Analysis shows that current general-purpose memory systems often retrieve topically relevant context but store project knowledge as incomplete or fragmented facts. These results reveal a domain-transfer gap in agent memory and suggest that reliable professional agents require domain-aware memory representations linking conversations, project knowledge, and structured model entities.

13:00 JSTLLM/生成AIエージェント

IDP AutoOpt: ドキュメント処理パイプライン構成のエージェント主導の最適化

インテリジェント文書処理 (IDP) パイプラインの高性能構成を検出する自律型 LLM エージェントである IDP AutoOpt を紹介します。現在、IDP プロンプト、モデル、OCR 設定、およびスキーマを合わせて調整するには、ドメイン スペシャリストにドキュメント タイプごとに 20 ~ 80 人時間以上のコストがかかり、企業がドキュメント クラスを追加しても拡張できません。 IDP AutoOpt は閉ループを実行します。つまり、本番の専門知識をコード化した人間が作成したドメイン スキルに基づいて、小さなラベル付きセットで構成をスコアリングし、フィールド レベルのエラーを診断し、対象を絞った編集を生成し、再評価します。ヘルスケア、マーケティング インテリジェンス、金融サービスの設定で導入されている抽出、分類、パケット分割タスク全体で、IDP AutoOpt は人間の専門家の精度と同等かそれを上回る精度を同等またはそれ以下のコストで実現し (抽出ベンチマークでは、ページあたりのコストが 4.6 倍低い場合で 90.2% 対 81.6%)、構成時間を数週間から 2 時間未満に短縮します。さらに、エージェントの LLM 機能には、それを下回ると最適化が失敗するハードしきい値があり、厳選されたドメイン スキルが生のソース コード アクセスよりも優れたパフォーマンスを発揮するため、構造なしで提供するとパフォーマンスが低下する可能性があることを示します。また、コンテキスト管理と差異の軽減に関する実践的なレッスンも共有します。このアプローチは、構成可能なパイプライン、スコアリング関数、および小さなラベル付きセットのみを必要とするため、IDP を超えて、構成が展開のボトルネックになっている RAG やマルチエージェント ワークフローなどの他のエンタープライズ AI システムにも拡張されます。

原文 (English)

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tuning IDP prompts, models, OCR settings, and schemas jointly currently costs domain specialists 20 to 80+ person-hours per document type and does not scale as enterprises add document classes. IDP AutoOpt runs a closed loop: it scores a configuration on a small labeled set, diagnoses field-level errors, generates targeted edits, and re-evaluates, guided by human-authored domain skills that encode production expertise. Across extraction, classification, and packet-splitting tasks deployed in healthcare, marketing-intelligence, and financial-services settings, IDP AutoOpt matches or exceeds human-expert accuracy at equal or lower cost (on an extraction benchmark, 90.2% vs 81.6% at 4.6 x lower per-page cost), cutting configuration time from weeks to under two hours. We further show that agent LLM capability has a hard threshold below which optimization fails, and that curated domain skills outperform raw source-code access, which can degrade performance when provided without structure. We also share practical lessons on context management and variance mitigation. Requiring only a configurable pipeline, a scoring function, and a small labeled set, the approach extends beyond IDP to other enterprise AI systems, such as RAG and multi-agent workflows, where configuration bottlenecks deployment.

13:00 JST研究/論文Claude

FinCacheServe: 変更可能なエンタープライズ ドキュメント上でコスト効率の高い RAG サービスを提供するための、依存関係の一貫した回答の再利用

可変企業ドキュメントに対する検索拡張生成サービスは、意味的に同等の分析リクエストを繰り返し実行します。応答の再利用により、GPU に依存した生成作業を排除できますが、応答キャッシュには、ファイリング、証拠チャンク、およびツール出力が変更された場合の依存関係の一貫性が必要です。 FinCacheServe は、生成された各回答を、企業の意図によってインデックス付けされ、ドキュメントのバージョン、証拠のフィンガープリント、ツールのフィンガープリント、モデル ID、およびデコード構成によって保護されたサービス オブジェクトとして扱います。 vLLM 実装は、SEC 由来の財務文書ワークロードを Qwen2.5 モデルで評価します。 2,230 リクエストのホストされた 7B トレースでは、FinCacheServe は LLM 呼び出しの 53.27% をスキップし、依存関係が古い出力は観察されませんでした。 3 つのホストされた 32B オペレーター スイート シード全体では、544 件のリクエストのうち 53.31% がスキップされます。これに対し、バージョン化されたセマンティック キャッシュでは 38.97%、グラウンディング スタイルの再利用では 22.43% がスキップされます。容量、バックエンド、および SLO のリプレイでは、Oracle 制限付きキャッシュ管理、100,000 エントリのトランザクション メタデータ動作、およびバージョン管理されたセマンティック キャッシュよりも依存関係の新しい 2 秒の SLO 成功あたりの推定 Wh が 44.30% 低いことが示されています。

原文 (English)

FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents

Retrieval-augmented generation services over mutable enterprise documents repeatedly execute semantically equivalent analysis requests. Answer reuse can remove GPU-bound generation work, yet response caches require dependency consistency when filings, evidence chunks, and tool outputs change. FinCacheServe treats each generated answer as a serving object indexed by enterprise intent and guarded by document versions, evidence fingerprints, tool fingerprints, model identity, and decoding configuration. A vLLM implementation evaluates SEC-derived financial-document workloads with Qwen2.5 models. On a 2,230-request hosted 7B trace, FinCacheServe skips 53.27% of LLM calls with zero observed dependency-stale outputs. Across three hosted 32B operator-suite seeds, it skips 53.31% of 544 requests, compared with 38.97% for versioned semantic caching and 22.43% for grounded-style reuse. Capacity, backend, and SLO replays show oracle-bounded cache management, 100k-entry transactional metadata behavior, and 44.30% lower estimated Wh per dependency-fresh 2s-SLO success than versioned semantic caching.

13:00 JST研究/論文

密閉インフラにおける水素漏れ検出のためのセンサー配置の最適化: CFD に基づいた遺伝的アルゴリズムと DeepSets ニューラル サロゲートを使用した比較研究

燃料電池自動車の駐車施設などの閉鎖環境における水素インフラは、水素の低い着火エネルギーと広い可燃範囲により、安全性に関する重大な課題を抱えています。現在の監視システムは大部分が反応型であり、危険な濃度が形成された後にのみ漏れを検出します。この研究では、数値流体力学 (CFD)、遺伝的アルゴリズム (GA) の最適化、DeepSets ニューラル サロゲートを統合することにより、プロアクティブなセンサー配置の最適化のための計算フレームワークを開発します。 180 のシナリオからなる CFD データベースが、代表的な 50 m x 30 m x 3 m のガレージに対して生成され、複数の漏れ位置、速度 (1 ~ 150 g/s)、および換気条件 (ACH = 1 時間あたり 3 ~ 10) をカバーしました。センサーの配置は、多目的 GA を使用して最適化され、均一、ランダム、およびサロゲート支援アプローチと比較されました。 GA は 60 秒以内に 96.1% の検出率を達成し、死角領域を 0.12% に削減しました。これは、均一なベースラインと比較して複合適合性が約 5% 向上したことに相当します。 DeepSets サロゲートは、CFD 評価を 89% 削減し、計算時間を 2 桁削減しながら、適合度ギャップが 0.01 未満の最適に近い構成を再現しました。検出パフォーマンスと空間カバレッジは GA と同等であり、サロゲート支援最適化が迅速な設計反復を可能にしながらソリューションの品質を維持できることを示しています。全体として、結果は、従来のレイアウトと比較して、CFD に基づいた最適化により検出効率が向上し、センサー要件が軽減されることを示しています。提案されたフレームワークは、スケーラブルな展開をサポートし、リアルタイムの監視とリスク評価のために、最適化されたセンサー ネットワークとデジタル ツイン システムを統合するための基盤を提供します。

原文 (English)

Optimizing Sensor Placement for Hydrogen Leak Detection in Enclosed Infrastructure: A Comparative Study Using CFD-informed Genetic Algorithm and DeepSets Neural Surrogate

Hydrogen infrastructure in enclosed environments, such as parking facilities for fuel cell vehicles, presents significant safety challenges due to hydrogen's low ignition energy and wide flammability range. Current monitoring systems are largely reactive, detecting leaks only after hazardous concentrations have formed. This study develops a computational framework for proactive sensor placement optimization by integrating computational fluid dynamics (CFD), genetic algorithm (GA) optimization, and a DeepSets neural surrogate. A CFD database of 180 scenarios was generated for a representative 50 m x 30 m x 3 m garage, covering multiple leak positions, rates (1-150 g/s), and ventilation conditions (ACH = 3-10 per hour). Sensor placement was optimized using a multi-objective GA and compared with uniform, random, and surrogate-assisted approaches. The GA achieved a detection rate of 96.1% within 60 s and reduced blind areas to 0.12%, corresponding to an approximately 5% improvement in composite fitness over a uniform baseline. The DeepSets surrogate reproduced near-optimal configurations with a fitness gap below 0.01 while reducing CFD evaluations by 89% and computational time by two orders of magnitude. Detection performance and spatial coverage remained comparable to the GA, demonstrating that surrogate-assisted optimization can retain solution quality while enabling rapid design iteration. Overall, the results show that CFD-informed optimization improves detection effectiveness and reduces sensor requirements compared to conventional layouts. The proposed framework supports scalable deployment and provides a foundation for integrating optimized sensor networks with digital twin systems for real-time monitoring and risk assessment.

13:00 JSTビジネス/資金調達

大規模な言語モデルにおけるサイレント推論の失敗を検出するための参照不要のスコア

数学的思考連鎖 (CoT) の評価は、通常、最終的な答えが参考資料と一致するかどうかに集約されます。これは、正しい結論を生成することと有効な導出を生成することを混同しており、無効な連鎖が誤って正しい答えに到達する可能性があり、有効な計算の後に転記エラーが発生する可能性があります。この不一致を推論回答一貫性ギャップと呼びます。このフレームワーク ペーパーでは、発行された数学的トレースが局所的に信頼でき、その答えが裏付けられ、リサンプリングや対象を絞った反事実的介入の下でも安定しているかどうかを判断する、参照フリーのインスタンス レベルの診断である、Reasoning Answer Faithful Score (RAFS) を紹介します。 RAFS は、ステップの妥当性、含意と反事実の感度を回答するための推論、回答のコンセンサス、および条件付き推論の安定性を組み合わせます。モデルのプライベートな計算や、テストされた数学的設定以外の事実の正確さではなく、転写レベルの一致を評価します。当社では、GSM8K および MATH に関する事前登録済みの結果ブラインド確認研究を保持しており、確認結果が検査される前に仮説、許容ルール、校正、およびテストが修正されています。個別の実現可能性パイロットは、トレース レベルのアーティファクトが利用可能な場合にのみ数値パイロット要求が報告される凍結前に、エンドツーエンドの実行を検証し、介入範囲を推定するために指定されます。 4 つの推論の回答結果を形式化し、非補償的なアグリゲーターを正当化し、意味論的なトレース距離をインスタンス化し、計算と棄権のトレードオフを定量化し、検証者の独立性と電力分析を定義します。 RAFS は、サイレント推論の失敗と回答抽出エラーに対する監査可能な警告信号によって数学的回答の精度を補完することを目的としています。

原文 (English)

A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models

Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producing a correct conclusion with producing a valid derivation an invalid chain can accidentally reach the right answer, while a valid calculation can be followed by a transcription error. We call this mismatch the reasoning answer consistency gap. This framework paper introduces the Reasoning Answer Faithfulness Score (RAFS), a reference free, instance level diagnostic of whether an emitted mathematical trace is locally credible, supports its answer, and is stable under resampling and targeted counterfactual interventions. RAFS combines step validity, reasoning to answer entailment and counterfactual sensitivity, answer consensus, and conditional reasoning stability. It evaluates transcript level agreement, not a models private computation and not factual correctness outside the tested mathematical setting. We retain a preregistered, results blind confirmatory study on GSM8K and MATH, with hypotheses, admissibility rules, calibration, and tests fixed before confirmatory outcomes are inspected. A separate feasibility pilot is specified to verify end to end execution and estimate interven tion coverage before that freeze numerical pilot claims are re ported only when trace level artifacts are available. We formalize four reasoning answer outcomes, justify the non compensatory aggregator, instantiate semantic trace distance, quantify compute and abstention tradeoffs, and define verifier independence and power analyses. RAFS is intended to complement mathematical answer accuracy with an auditable warning signal for silent reasoning failures and answer extraction errors

13:00 JST画像/動画生成

野生で撮影された単一の人間の画像からの体重と身長の推定

体重や身長などの人の身体的特徴は、その人の身体的および精神的健康、日常生活および経済状態を示す重要な指標です。 Body Mass Index (BMI) は、体重と身長の両方の特性をコード化するよく知られた尺度です。 BMI は自己監視ツールとして使用されており、人生に長期的な影響を及ぼします。たとえば、さまざまな病気のリスクを予測したり、寿命を推定したりするのに役立つ可能性があります。自然界の 1 人の人物画像を使用した自動 BMI 推定は、人間の姿勢、カメラの幾何学形状、個人の外見、気を散らす背景が大きく異なるため、困難な作業です。この論文では、ソーシャル ネットワーキング Web サイトで入手可能な日常生活の画像から BMI、体重、身長を予測するために、RGB、深度マップ、ポーズ アフィニティ マップ、エッジ マップなどのさまざまなモダリティを採用することにより、シングルおよびマルチタスク学習を使用したディープ ニューラル ネットワークのパフォーマンスを調査します。現在、BMI 推定用の全身画像データセットは公開されていないため、身長、体重、BMI のグラウンド トゥルース ラベルを持つ 6105 枚の画像で構成される新しいデータセットを提案します。私たちが提案するデータセットは、さまざまな民族の画像を含む野生で収集され、さまざまな年齢層と性別に分散されています。これは、正面、背面、全身および半身、横向きのポーズ、さまざまな背景とスケールのバリエーションを備えたミラーセルフィーで構成されており、顔の一部または全体を隠すアーティファクトが含まれる場合があります。 VGG、Densenet、ResNet などのさまざまな CNN バックボーンを使用して、全身、半身、顔の画像のみを使用して広範な実験が実行されます。私たちの実験結果は、全身画像が実際の他の半身画像や顔画像よりも良い結果を生み出すことを示しています。

原文 (English)

Weight and Height Estimation from a Single Human Image Captured in the Wild

A person's physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routines and finances. Body Mass Index (BMI) is a well known measure that encodes the characteristics of both the weight and the height. BMI has been used as a self-monitoring tool, and it has long-term implications on one's life. For example, it may help predicting the risk of various diseases and estimating longevity. Automatic BMI estimation using a single person image in the wild is a challenging task due to wide variations in human pose, camera geometry, personal appearance and distracting backgrounds. In this paper, we explore the performance of deep neural networks using single and multi-task learning by employing different modalities including RGB, depth-maps, pose-affinity maps, and edge-maps to predict BMI, weight, and height from daily life images available on social networking websites. Currently, no full body image dataset for BMI estimation is publicly available, therefore we propose a new dataset consisting of 6105 images with ground truth labels of height, weight and BMI. Our proposed dataset is collected in the wild containing images from various ethnicity and distributed over varying age groups and gender. It consists of frontal, back, full and half body, side poses, mirror selfies with varying backgrounds and scale variations and may contain artifacts hiding partial or full face. Extensive experimentation is performed using full body, half body and face images only using different CNN backbones including VGG, Densenet and ResNet. Our experimental results demonstrate that full body images have produced better results than the other half body and facial images in the wild.

13:00 JST画像/動画生成

TraceCLIP: パッチから CLS への貢献からローカル セマンティクスを回復する

物体の位置特定、領域認識、オープン語彙の意味セグメンテーションなどの緻密な視覚言語理解には、言語概念を空間的に根拠のある視覚領域と関連付けることが必要です。 CLIP は、大規模な対比事前トレーニングから共有画像テキスト埋め込み空間を学習することで、これらのタスクに強力な基盤を提供します。ただし、その画像レベルの目的は、テキストを CLS 由来のグローバル表現に合わせて調整し、ローカルな視覚言語の対応を間接的にのみ制限します。既存の方法では、追加の監視、外部モデル、またはタスク固有の適応を導入しますが、トレーニング不要のアプローチでは主に、CLIP 内でローカル セマンティクスが最もアクセスしやすい場所を調べることなく、既存のパッチ機能から密な応答を回復します。 CLS アテンション出力に書き込まれたパッチ固有の用語を分離することで、潜在的なパッチ レベルの意味論的証拠を回復する、トレーニング不要のフレームワークである TraceCLIP を紹介します。 TraceCLIP はさらに、寄与に由来するセマンティック応答を、密なフィーチャ再構築のための最終層パッチ アフィニティを調整するセマンティック測地トポロジ ゲートに変換します。診断実験では、これらの寄与特徴が強力な局所的な意味識別とテキスト条件付きの空間的位置合わせを示すことが示されています。 8 つのゼロショット セマンティック セグメンテーション ベンチマークにおいて、TraceCLIP は、バックボーンとバックグラウンド設定の両方で、追加のトレーニング、外部ビジョン基盤モデル、または領域レベルの監視なしで、これまでの最も強力なトレーニング不要の方法と比較して、平均 mIoU で 1.3 ~ 4.5 ポイントの向上を達成しました。より広く言えば、これらの発見は、空間的に局所化されたセマンティクスが、グローバルに整列された表現の内部構築内でアクセス可能なままである可​​能性があることを示唆しています。

原文 (English)

TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions

Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions. CLIP provides a strong foundation for these tasks by learning a shared image-text embedding space from large-scale contrastive pre-training. However, its image-level objective aligns text with a CLS-derived global representation, leaving local vision-language correspondence only indirectly constrained. Existing methods either introduce additional supervision, external models, or task-specific adaptation, while training-free approaches mainly recover dense responses from existing patch features without examining where local semantics become most accessible within CLIP. We introduce TraceCLIP, a training-free framework that recovers latent patch-level semantic evidence by isolating the patch-specific terms written into the CLS attention output. TraceCLIP further converts contribution-derived semantic responses into a semantic-geodesic topology gate that calibrates final-layer patch affinity for dense feature reconstruction. Diagnostic experiments show that these contribution features exhibit strong local semantic discrimination and text-conditioned spatial alignment. On eight zero-shot semantic segmentation benchmarks, TraceCLIP achieves gains of 1.3 to 4.5 points in average mIoU over the strongest prior training-free methods across both backbones and background settings, without additional training, external vision foundation models, or region-level supervision. More broadly, these findings suggest that spatially localized semantics may remain accessible within the internal construction of globally aligned representations.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

GPT-Red: 大規模なセルフプレイによる自動レッド チーミング

\textbf{GPT-Red} は、フロンティア LLM に対する新しいプロンプト インジェクション攻撃を発見するように訓練された自動レッドチーム エージェントです。このモデルの目標は、生産システムの堅牢性を評価し、改善することです。この目的を達成するために、これを使用して GPT-5.6 を敵対的にトレーニングします。GPT-5.6 は、これまでの注入を促すための最も堅牢なモデルです。 GPT-Red を作成するために、モデルが同時に訓練された防御エージェントの多様な集団を攻撃する任務を負うスケーラブルなセルフプレイ アルゴリズムを設計します。当社では、最大規模の RL ポストトレーニング実行と同じ規模のコンピューティングを使用して、現実的なレッドチーム環境でモデルをトレーニングしており、これはこれまで文書化された中で最大の LLM 安全トレーニング実行となっています。 GPT-Red はレッドチーム化に優れています。GPT-5.5 までの過去のモデルを確実に破壊し、人間のレッドチームよりも成功した攻撃を検出し、保持されている環境、防御モデル、およびハーネスに一般化します。将来的には、新しい GPT モデルの堅牢性が向上するにつれて、\textit{さらに強力な} レッドチーマー エージェントにより良い学習信号が提供され、自己改善のフライホイールのロックが解除されることを期待しています。

原文 (English)

GPT-Red: Automated Red Teaming via Self-Play at Scale

We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for \textit{even stronger} red-teamer agents, thus unlocking a self-improvement flywheel.

13:00 JSTハードウェア/半導体

振り返らずにもう一度試してください: ブラインド リサンプリングは小さなコード モデルでの自己修復よりも優れたパフォーマンスを発揮します

自己修復 (失敗したプログラムをテスト出力とともにモデルに返し、修正を要求する) は、コード エージェントの標準コンポーネントであり、ほとんどの場合、再試行をまったく行わないベースラインに対して評価されます。この比較は、フィードバックの価値と追加の試行の価値を混同していると私たちは主張します。 3 つのモデル スケール (1.5B、3B、7B) で MBPP+ のプラセボ対照設計を使用して、4 つの予算に見合った再試行条件 (ブラインド リサンプリング、内容のない失敗通知、本物の実行フィードバック、口頭での内省で強化されたフィードバック) を比較します。ブラインド リサンプリングは 7B 未満では最も強い条件であり、統計的には 7B での最良の条件と並んでいますが、消費するトークンは 2.5 ~ 5.5 分の 1 です。モデル自身の失敗した試行に基づく条件付けのコストは 1.5B で 6.1 ポイント (p=0.006)、実行フィードバックの情報内容はプラセボに比べて測定可能なものを何も追加しません。これはアンカリングによるものだと考えられます。前回の試行が示された場合、モデルは再試行の 33 ~ 68% でほぼ同一のプログラムを再現しましたが、ブラインド リサンプリングでは 2 ~ 14% でした。さらに 2 つの実験により、その効果が明らかになりました。他のタスクに対して取得されたソリューションは何も変更せず (+/-3.5 ポイントに制限されます)、これにより、コンテキストの長さではなく自己調整に対する害が局所化されます。そして、アンカーを明らかに弱める唯一の条件である反射は、依然としてコストに支配されています。レプリケーションにより、2 つの競合する説明が排除されます。ペナルティは完全な精度で変更されず、独立したモデル ファミリで再現されます。 2 つのファミリーと 2 つの精度にわたる 6 つの構成にわたって、その規模はベースラインの品質のみによって予測されます (r=0.96)。アンカリングのコストは、最初の試行が失敗することによるコストです。

原文 (English)

Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models

Self-repair - returning a failed program to the model together with its test output and asking for a correction - is a standard component of code agents, and is almost always evaluated against a baseline that does not retry at all. We argue that this comparison confounds the value of the feedback with the value of the extra attempt. Using a placebo-controlled design on MBPP+ at three model scales (1.5B, 3B, 7B), we compare four matched-budget retry conditions: blind resampling, a content-free failure notice, genuine execution feedback, and feedback augmented with verbal self-reflection. Blind resampling is the strongest condition below 7B, and remains statistically tied with the best condition at 7B, while consuming 2.5-5.5x fewer tokens; conditioning on the model's own failed attempt costs 6.1 points at 1.5B (p=0.006), and the informational content of execution feedback adds nothing measurable over the placebo. We attribute this to anchoring: when shown its previous attempt, a model reproduces a near-identical program in 33-68% of retries, against 2-14% under blind resampling. Two further experiments delimit the effect. Retrieved solutions to other tasks change nothing (bounded to +/-3.5 points), which localizes the harm to self-conditioning rather than context length; and reflection, the only condition that measurably weakens the anchor, remains dominated on cost. Replication rules out two competing explanations: the penalty is unchanged at full precision, and it reproduces on an independent model family. Across six configurations spanning two families and two precisions, its magnitude is predicted by baseline quality alone (r=0.96) - the cost of anchoring is the cost of committing to a bad first attempt.

13:00 JSTロボティクス

信頼できる身体化されたインテリジェンスに向けて: システム フレームワークと段階的な信頼性レベル

身体化されたインテリジェンスは、学習された知覚と意思決定をリアルタイムの計算、制御、物理的相互作用と統合します。失敗すると即座に物理的または運用上の損害が発生する可能性があるため、タスクの完了だけでは信頼性は確立されません。私たちは、信頼できる身体化インテリジェンスを、リスクを許容範囲内に維持しながら、環境やシステムの変化の下で指定されたタスクを確実に実行する持続的な能力として定義します。私たちはこの目標を「持続的かつ安全な成功」と呼びます。そのサポートメカニズムは、相互に依存する 4 つの層で構成されています。モデル層は、調整された不確実性と明示的な安全性優先設定を使用して、タスクに適したアクション提案を生成します。システム層は、統合されたセンシング、計算、制御、ハードウェア保護手段、障害封じ込め、およびフォールバックを通じて、承認されたアクションを確実に実現します。証拠レイヤーは、評価、検証、検証、トレーサビリティ、および構造化された保証の議論を通じて、限定された主張を実証します。デプロイ層は、実行時の監視、権限管理、介入、インシデント対応、制御された更新を通じてクレームの有効性を維持します。仮定と障害はこれらの層全体に伝播するため、モデルの機能、分離された安全対策、ベンチマークのパフォーマンスだけでは、エンドツーエンドの信頼性を確立できません。身体化された AI、ロボット工学、制御、ディペンダブル コンピューティング、分散システム、自動運転を活用して、信頼性レベルの非規範的な階層をさらに提案します。この階層は、タスクの能力、安全性、システム保証、運用ガバナンス、および裏付けとなる証拠にわたって、限定された展開の主張の強度を評価し、限定された展開、比較評価、研究の優先順位付け、および将来の標準化の基礎を提供します。

原文 (English)

Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels

Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Because failures can cause immediate physical or operational harm, task completion alone does not establish trustworthiness. We define trustworthy embodied intelligence as the sustained capacity to execute specified tasks reliably under environmental and system variation while maintaining risk within acceptable bounds. We term this objective sustained safe success. Its supporting mechanisms are organized into four interdependent layers. The model layer generates task-competent action proposals with calibrated uncertainty and explicit safety preferences. The system layer realizes authorized actions dependably through integrated sensing, computation, control, hardware safeguards, fault containment, and fallback. The evidence layer substantiates bounded claims through evaluation, verification, validation, traceability, and structured assurance arguments. The deployment layer maintains claim validity through runtime monitoring, authority management, intervention, incident response, and controlled updates. Because assumptions and failures propagate across these layers, neither model capability, isolated safeguards, nor benchmark performance alone can establish end-to-end trustworthiness. Drawing on embodied AI, robotics, control, dependable computing, distributed systems, and autonomous driving, we further propose a non-normative hierarchy of trustworthiness levels. This hierarchy grades the strength of bounded deployment claims across task capability, safety, system assurance, operational governance, and supporting evidence, providing a basis for bounded deployment, comparative evaluation, research prioritization, and future standardization.

13:00 JST画像/動画生成

百聞は一見に如かず - ハイブリッドディープラーニングによる画像からの皮膚曝露データを利用して安全性評価を強化

この研究では、皮膚曝露評価のために画像から曝露された皮膚を定量化するハイブリッド コンピュータ ビジョン手法を開発しました。 170 枚の屋内絵画画像を使用して、Mask R-CNN は最初に人物を識別し、背景の干渉を除去しました。次に、色ベースのアルゴリズムにより、露出した肌をセグメント化します。結果として得られた露出した肌と身体のピクセル比率は、人間の推定値と約 80% 一致しました。このアプローチは、画像から半定量的な曝露情報を抽出するスケーラブルな方法を示しており、将来的には身体部分の認識、PPE 検出、およびビデオベースの曝露分析に拡張されます。

原文 (English)

A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment

This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-painting images, Mask R-CNN first identified human subjects and removed background interference; a color-based algorithm then segmented exposed skin. The resulting exposed-skin-to-body pixel ratios showed approximately 80% agreement with human estimates. The approach demonstrates a scalable way to extract semi-quantitative exposure information from images, with future extensions to body-part recognition, PPE detection, and video-based exposure analysis.

13:00 JSTLLM/生成AI

認知の収束: 大規模言語モデルと人間の認知との深い類似性

LLM は、宇宙人の知性、つまりその認知機能が根本的に私たちとは異なるシステムであると広く考えられています。したがって、人間の認知との見かけの類似性は、擬人化投影の結果として見られることがよくあります。私たちはこの枠組みが間違っていると主張します。 LLM は、物理的基盤、学習履歴、相互作用する環境など、重要な点で人間とは明らかに異なります。これらの違いにより、現代の LLM ベースのシステムが、認知科学における長年のサポートを受けて、認知組織化の多くの原則に基づいて人間の認知と収束していることが、さらに顕著になります。私たちは、推論組織、計算アーキテクチャ、表現構造、予測駆動学習、目標指向の行動をサポートする強化学習のようなメカニズムの 5 つの次元にわたる構造的対応関係を特定します。これらの対応関係は、人間の知能を説明するために長い間使用されてきた中心原理が、現代の LLM ベースのシステムの特徴でもある、知的認知のより広範なモデルをサポートしています。

原文 (English)

Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition

LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human cognition are therefore often seen as the result of anthropomorphic projection. We argue that this framing is mistaken. LLMs clearly differ from humans in important respects, including their physical substrate, learning history, and the environments with which they interact. These differences make it all the more striking that contemporary LLM-based systems converge with human cognition on a number of principles of cognitive organization with longstanding support in cognitive science. We identify structural correspondences across five dimensions: inferential organization, computational architecture, representational structure, prediction-driven learning, and reinforcement-learning-like mechanisms supporting goal-directed action. These correspondences support a broader model of intelligent cognition in which core principles long used to explain human intelligence also characterize contemporary LLM-based systems.

13:00 JSTLLM/生成AIエージェント

(EC)2: マルチエージェント LLM 調査によるサイバーセキュリティのイベント中心の説明可能性

セキュリティ オペレーション センターは、異常検出システムを利用して不審なイベントにフラグを立てます。異常検出器の機能レベルの説明は、運用調査に限られた価値を提供します。アラートを効果的に処理するには、アナリストはコンテキスト上の関係を理解し​​、関係するエンティティを実用的に理解する必要があります。このペーパーでは、中小規模の企業ネットワークにおけるサイバーセキュリティ アラートを説明するための、イベント中心の検出器に依存しないアプローチを紹介します。我々は、構造化された仮説に基づいた調査を実行して、検証可能な証拠に基づいた説明を提供するマルチエージェント フレームワークである (EC)2 を紹介します。評価結果は、提案されたフレームワークが運用上意味のある説明を生成することで検出後の分析を改善し、イベント分類の精度も向上させることを示しています。

原文 (English)

(EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations

Security operations centers rely on anomaly detection systems to flag suspicious events. Feature-level explanations for anomaly detectors offer limited value for operational investigations. To effectively handle alerts, analysts need to know contextual relationships and need actionable understanding of the entities involved. This paper introduces an event-centric detector-agnostic approach for explaining cybersecurity alerts in small- to medium-sized enterprise networks. We present (EC)2, a multi-agent framework that performs structured, hypothesis-driven investigation to provide explanations grounded in verifiable evidence. Evaluation results show that the proposed framework improves post-detection analysis by generating operationally meaningful explanations, which also enhance event classification accuracy.

13:00 JSTLLM/生成AIエージェント

マルチエージェントのディベート戦略: 調査、分類、課題

マルチエージェント ディベート (MAD) は、大規模言語モデル (LLM) ベースのエージェント システムの精度と堅牢性を向上させるための有望なパラダイムです。これにより、複数のエージェントが議論を交換し、互いの出力を批評し、解決策に向けて反復的に収束することが可能になります。しかし、研究は依然として断片的であり、用語に一貫性がなく、MAD 設計の次元が厳密に統合されていません。我々は、MAD に関する 141 件の主要研究を特徴づける体系的な文献レビューを紹介します。私たちは、議論の参加者、交換を構築する対話メカニズム、議論の解決を管理する合意プロトコルをカバーする 3 次元の分類法を導き出します。これは、MAD 構成をレンダリングするための正式な表記法によってサポートされています。私たちの分析では、この分野が暗黙のうちに、体系的な比較ではなく慣例によって採用されている、静的で完全に接続されたトポロジー、逐語的交換、短期記憶および投票解決戦略などの狭い設計パターンに収束しており、有望な代替案は依然として限界に達していることが明らかになりました。どの MAD 設定にも、相互に作用する約 12 件の設計上の決定が反映されるため、これらが暗黙的に残されている場合、スタディ間の比較は信頼性が低くなります。私たちは分類法を研究状況の説明的なマップ、制御されたベンチマークのフレームワーク、そして潜在的には機械可読な MAD 仕様のスキーマとして位置づけています。今後の作業として、これを実行可能な仕様に形式化し、コストを意識したベンチマークと議論構成の自動チューニングを可能にすることを提案します。

原文 (English)

Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic systems. It enables multiple agents to exchange arguments, critique each other's outputs, and iteratively converge towards a solution. However, research remains fragmented, with inconsistent terminology and no rigorous synthesis of MAD design dimensions. We present a systematic literature review characterizing 141 primary studies on MAD. We derive a three-dimensional taxonomy covering debate participants, the interaction mechanisms structuring the exchange, and the agreement protocols governing debate resolution, supported by formal notations to render MAD configurations. Our analysis reveals that the field has implicitly converged on a narrow design pattern - static, fully connected topologies, verbatim exchange, short-term memory and voting resolution strategies - adopted by convention rather than systematic comparison, while promising alternatives remain marginal. Because any MAD setting reflects roughly a dozen interacting design decisions, cross-study comparison is unreliable when these are left implicit. We position the taxonomy as a descriptive map of the research landscape, a framework for controlled benchmarking, and potentially as a schema for machine-readable MAD specifications. As future work, we propose formalizing it into an executable specification, enabling cost-aware benchmarking and automated tuning of debate configurations.

13:00 JSTLLM/生成AI

3 値の不確実性スコアリングを使用したモデル駆動の要件構成

コンテキスト: 大規模言語モデル (LLM) は、自動化された要件の引き出しに自然言語の柔軟性を提供しますが、構造的に無効な要件や論理的矛盾が頻繁に生成され、形式的な正確性の保証がありません。目的: この研究は、正式なドメイン モデル内で LLM の事前検証決定の不確実性を定量化しながら、論理的矛盾を排除し、LLM で生成された要件の構造的適合性を強制することを目的としています。方法: 要件作成および管理のためのオブジェクト指向メソッド (OOMRAM) ラティスを運用する神経記号的なマルチエージェント アーキテクチャを紹介します。 LLM は格子トラバーサルの非決定論的ヒューリスティックとして機能しますが、決定論的シンボリックバリデーターはすべての構造的制約を強制します。検証の前後で LLM の要件決定を分類し、スコアリングするための 3 つの値 (T、I、F)、つまり真実、不確定性、虚偽のフレームワークを導入します。結果: 11 のアプリケーション ファミリの 37 の自然言語プロジェクト ビジョンにわたって評価したところ、システムは 37 ケース中 35 ケース (94.6%) で構造的不一致を完全に除去し、残りの 2 ケースでは反復制限により未解決の構造エラーが 6 つだけ (決定の 0.39%) を含みました。 3 値分析により、すべての意思決定の 24.7% が不確定、つまり構造的には有効だが利害関係者によって明示的に義務付けられていない裁量的選択であることが明らかになりました。結論: 構造の完全性を決定論的なシンボリック層にオフロードすることで構造の適合性が保証され、その一方で 3 値の分類により神経の不確実性を測定する正式な方法が提供され、形式要件エンジニアリングにおける安全な LLM の展開が容易になります。

原文 (English)

Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring

Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate structurally invalid requirements and logical inconsistencies, lacking formal correctness guarantees. Objectives: This study aims to eliminate logical inconsistencies and enforce structural conformance in LLM-generated requirements while quantifying the LLM's pre-validation decision uncertainty within a formal domain model. Methods: We present a neuro-symbolic multi-agent architecture that operationalizes the Object-Oriented Method for Requirements Authoring and Management (OOMRAM) lattice. The LLM acts as a non-deterministic heuristic for lattice traversal, while a deterministic symbolic validator enforces all structural constraints. We introduce a three-valued (T, I, F) -- Truth, Indeterminacy, Falsity -- framework to classify and score the LLM's requirement decisions before and after validation. Results: Evaluated across 37 natural-language project visions in eleven application families, the system completely eliminated structural inconsistencies in 35 out of 37 cases (94.6%), with the remaining two containing only 6 unresolved structural errors (0.39% of decisions) due to iteration limits. Three-valued analysis revealed that 24.7% of all decisions are indeterminate -- structurally valid but discretionary choices not explicitly mandated by the stakeholder. Conclusion: Offloading structural integrity to a deterministic symbolic layer successfully guarantees structural conformance, while the three-valued classification provides a formal way to measure neural uncertainty, facilitating safe LLM deployment in formal requirements engineering.

13:00 JST研究/論文

文脈に応じた反論は、一般的な反論よりも説得力が高まる可能性がある

AI によって生成された反対演説は、より建設的な対話を促進することにより、オンラインの有害性を軽減するためのスケーラブルで効果的な戦略を提供します。しかし、既存のアプローチは、一般的な画一的なパラダイムを採用しており、会話のコンテキストや対象となるユーザーの特性を見落としています。ここでは、モデレーション設定に適応し、モデレートされたユーザーに合わせてパーソナライズされた、文脈に応じた反論を生成するための複数の戦略を提案および評価します。さまざまな形式のコンテキスト情報と微調整技術を統合するさまざまな構成を詳しく調査します。事前登録型の混合設計型クラウドソーシング実験に定量指標を組み合わせた総合評価を実施します。堅牢性を確保するために、ROUGE、BLEU、BERTScore に基づく反論品質のアルゴリズム測定を実装し、指標全体で全体的に一貫した結果を観察します。さらに、生成された反論スピーチと調整された有害なメッセージの両方のどの特性が、知覚される説得力に最も強く影響するかを分析し、状況に応じた介入をより効果的にする方法についての洞察をもたらします。私たちの調査結果は、パーソナライゼーションは効果的である可能性があるが、一律に効果があるわけではないことを示しています。会話のコンテキストとユーザー履歴を組み合わせた軽量の戦略は、認識される適切性と説得力を向上させますが、他のいくつかのコンテキスト化戦略は人間が認識する対向スピーチの品質を低下させます。総合すると、これらの結果は、よりパーソナライズされた、効果的で責任ある反論システムを開発するための実用的な方向性を提供し、最終的にはオンライン コンテンツ モデレーションにおける人間と AI のコラボレーションを推進します。

原文 (English)

Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech

AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics of both the generated counterspeech and the moderated toxic message most strongly influence perceived persuasiveness, yielding insights into how contextualized interventions can be made more effective. Our findings show that personalization can be effective, but not uniformly so. Lightweight strategies combining conversational context and user history improve perceived adequacy and persuasiveness, whereas several other contextualization strategies degrade human-perceived counterspeech quality. Taken together, these results provide actionable directions for developing more personalized, effective, and responsible counterspeech systems, ultimately advancing human-AI collaboration in online content moderation.

13:00 JSTエージェント

上位 $k$ パレート盗賊: 多目的スレート選択に対するハイパーボリュームの後悔

各ラウンドでエージェントが $k$ アームのスレートを選択し、セミバンディット フィードバックの下でそれらの $d$ 次元の報酬ベクトルを観察する確率的多目的バンディット問題を考えます。私たちの目的は、単一の最適なアームを特定することではありません。代わりに、パレートフロンティアを共同で近似する小さな一連のアクションを維持する問題を検討します。選択したアームのサブセットによって誘発される支配的なハイパーボリュームを通じてこの目的を形式化し、後から考えて達成可能な最良のサイズ $k$ サブセットに関する $\alpha$ 近似ハイパーボリューム リグレスを定義します。 ここで、 $\alpha = 1 - 1/e$ は、単調サブモジュラー関数の貪欲な最大化の近似保証を反映します。この問題に対処するために、限界ハイパーボリューム寄与の楽観的な推定に基づいてアームを貪欲に選択する楽観的アルゴリズムである \textit{THV-UCB} を導入します。すべてのインスタンスに保持されるギャップのないリグレス限界 $\tilde{O}(d\sqrt{nkT})$ と、腕が十分に分離されると $T$ で多対数になるギャップ依存限界 $\tilde{O}(nk^{2.5}/\Delta_{\min})$ を確立します。私たちの結果は、さまざまな多目的アプリケーションで小さなサブセットを使用してパレート フロントを近似するための理論的な裏付けを提供します。

原文 (English)

Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection

We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-dimensional reward vectors under semi-bandit feedback. We do not aim at identifying a single optimal arm; instead, we consider the problem of maintaining a small set of actions that jointly approximate the Pareto frontier. We formalize this objective through the dominated hypervolume induced by the selected subset of arms, and define an $\alpha$-approximate hypervolume regret with respect to the best size-$k$ subset achievable in hindsight, where $\alpha = 1 - 1/e$ reflects the approximation guarantee of greedy maximization for monotone submodular functions. To address this problem, we introduce \textit{THV-UCB}, an optimistic algorithm that selects arms greedily based on optimistic estimates of their marginal hypervolume contributions. We establish a gap-free regret bound $\tilde{O}(d\sqrt{nkT})$ that holds on every instance, together with a gap-dependent bound $\tilde{O}(nk^{2.5}/\Delta_{\min})$ that becomes polylogarithmic in $T$ once the arms are sufficiently well separated. Our results provide theoretical support for using small subsets to approximate Pareto fronts in various multi-objective applications.

13:00 JST研究/論文

実際のエンティティ解決: セルフサービス パイプラインからの教訓

私たちは、864 ~ 500 万レコードにわたる 6 つのベンチマークに基づいてセルフサービス エンティティ解決 (ER) システムを構築して評価しました。その結果、既存の ER 文献には欠けている 3 つの教訓が明らかになりました。 (1) 単一のマッチング アルゴリズムがどこでも勝てるわけではありません。セルフサービス パイプラインは次のデータセットを予測できないため、データセットごとに複数のアルゴリズム ファミリをトレーニングし、自動ベイクオフで勝者を選択させることをお勧めします。 (2) 適合率と再現率には、共有しきい値ではなく個別の修正が必要です。適合率には厳格なルールに基づく拒否権が必要で、再現率にはより多様な候補の取得が必要です。 (3) 1 つの誤検知リンクにより、無関係なエンティティがサイレントにマージされる可能性があります。「A が B と一致し」、「B が C と一致」と仮定すると、「A が C と一致」ということは、単一の不良リンクが数百のレコードを連鎖させることになるため、すべてのグループ間のマージを積極的に再検証する必要があります。私たちは、これらの教訓が実践者たちを、私たちを導いた何ヶ月もの行き止まりの実験から救ってくれることを願っています。

原文 (English)

Entity Resolution in Practice: Lessons from a Self-Serve Pipeline

We built and evaluated a self-serve entity resolution (ER) system on six benchmarks spanning 864 to 5M records, and three lessons emerged that are absent from existing ER literature. (1) No single matching algorithm wins everywhere - a self-serve pipeline cannot predict its next dataset, so we recommend training several algorithm families per dataset and letting an automatic bake-off pick the winner. (2) Precision and recall need separate fixes, not a shared threshold - precision needs hard rule-based vetoes, recall needs more diverse candidate retrieval. (3) One false-positive link can silently merge unrelated entities - assuming "A matches B" and "B matches C" implies "A matches C" lets a single bad link chain hundreds of records together, so every cross-group merge must be actively re-verified. We hope these lessons save practitioners the months of dead-end experiments that led us to them.

13:00 JSTLLM/生成AIエージェント

AgentGUI: 長時間実行される AI エージェントを監視および操作するためのインターフェイス

AI エージェントは、複雑で長時間実行されるタスクへの取り組みにますます熟練しています。自律機能の急速な急増に伴い、人間中心のインターフェースが限られているため、人間の監視は体系的に遅れています。これに対処することを目的として、複数の同時長時間実行セッションで AI エージェントをシームレスに監視および操作するための、ユーザーフレンドリーでローカルにホストされる GUI である AgentGUI を導入します。 AgentGUI の特徴は、1) 豊富なエージェント軌跡の視覚化、2) 効果的な手動および自動ステアリング、3) オープンソースおよびフロンティア エージェント フレームワークとの統合および調整です。管理されたユーザー調査では、エージェントのトレースから主要な要素を特定するのにかかる時間が統計的に有意に短縮されたことが実証されました (38% 高速化、p = 0.023)。予備実験では、AgentGUI の自動ドリフト防止機能により、0.8B ~ 9B モデル ラダー (モデルあたり N=50 実行) 全体で小規模のローカル エージェントのタスク完了率が 34pp も上昇しました。 AgentGUI は、プロジェクト Web サイト (https://agent-gui-project.github.io) およびオープンソース リポジトリ (https://github.com/eth-medical-ai-lab/agent-gui) を通じて、デモ ビデオ (https://youtube.com/watch?v=GSDyxN1gTF0) とともに公開されています。

原文 (English)

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematically lagging behind due to limited human-centered interfacing. Aiming to address this, we introduce AgentGUI, a user-friendly, locally hosted GUI for seamlessly observing and steering AI agents amid multiple concurrent, long-running sessions. AgentGUI features 1) rich agent trajectory visualizations, 2) effective manual and automated steering, and 3) integration with and coordination between open-source and frontier agent frameworks. A controlled user study demonstrates statistically significant reduction in the time it takes to identify key elements from agent traces (38% faster, p = 0.023). In a preliminary experiment, AgentGUI's automated drift prevention feature raises the task completion rate of small local agents by as high as 34pp across a 0.8B--9B model ladder (N=50 runs per model). AgentGUI is publicly available through its project website (https://agent-gui-project.github.io) and open-source repository (https://github.com/eth-medical-ai-lab/agent-gui), along with a demo video (https://youtube.com/watch?v=GSDyxN1gTF0).

13:00 JSTエージェント

SARC-DQ: Agentic AI の実行時データ品質ゲーティング: サイレント証拠の欠陥、無能シールド、およびダウンストリームのみの修復

エージェント システムは機能するため、取得した証拠に欠陥があると、通貨コストを伴う誤ったアクションとなります。最も危険な企業の欠陥はメタデータに起因するものです。つまり、ペイロード内で完全に整形式であり、鮮度、系統、または来歴によってのみ裏切られる、古い価格または置き換えられた記録です。このような欠陥がエージェントのコンテキストに侵入することは決してなく、エージェントは目に見えないデータを疑うことはできません。有償補充ベンチマークでは、有能なエージェントは、注入されたメタデータ由来の欠陥を約 60% の確率でサイレントにコストのかかるアクションに変換します。データ品質フラグや行動疑問マーカーはゼロ (AUC <= 0.50) が発生する可能性があります。推論価格が約 15 倍に及ぶ 4 つのモデル層にわたって、レートは横ばいです。能力があるからといって懐疑的な見方をするわけではありません。ダウンストリームのみの修復を備えたメタデータ対応のプレアクション ゲートは、述語がカバーする信号については損失を完全に回復しますが、見逃した信号についてはまったく回復しません。タスクの意思決定ジオメトリから派生したモデルフリーのオラクルは、MAE 0.015 (ピアソン r = 0.876、間隔カバレッジ 15/16 セル) で測定レートを追跡し、フラット ラダーに分析形式を与えます。証拠の完全性は、モデルの機能とは異なるシステム軸です。緩和は、施行の配置と述語の適用範囲によって異なります。コード、凍結された結果、決定論的分析パイプライン: https://github.com/besanson/dqSarc

原文 (English)

SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation

Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise defects are metadata-borne: a stale price or a superseded record, perfectly well-formed in the payload and betrayed only by freshness, lineage, or provenance. Such a defect never enters the agent's context, and an agent cannot doubt data it cannot see. On a priced replenishment benchmark, a competent agent silently converts an injected metadata-borne defect into a costly action about 60% of the time, with zero data-quality flags and behavioral doubt markers at chance (AUC <= 0.50). Across four model tiers spanning roughly 15x in inference price, the rate stays flat: capability does not buy skepticism. A metadata-aware pre-action gate with downstream-only remediation recovers the loss fully on the signals its predicates cover and not at all on those they miss. A model-free oracle derived from the task's decision geometry tracks the measured rates with MAE 0.015 (Pearson r = 0.876, interval coverage 15/16 cells), giving the flat ladder an analytical form. Evidence integrity is a systems axis distinct from model capability; mitigation depends on enforcement placement and predicate coverage. Code, frozen results, and a deterministic analysis pipeline: https://github.com/besanson/dqSarc

13:00 JSTエージェント

StealthBench: 自律型攻撃セキュリティ エージェントの運用ステルスを測定する

ステルスとは、自分の存在、能力、収集された情報を明らかにせずに目的を達成する規律であり、高度なオペレーターと検出可能なオペレーターを区別します。エリートのセキュリティ研究者と高度で永続的な脅威は、気づかれずに目的を達成します。自律エージェントはますます同じ攻撃タスクを継承しますが、彼らはトレードクラフトを継承しますか? 6 つの運用セキュリティ (OPSEC) の側面にわたって自律型攻撃セキュリティ エージェントの運用ステルスを測定するベンチマークである StealthBench を紹介します。実際のバグ報奨金とレッドチームの軌跡から手作業で検証された 11 件の OPSEC インシデントを抽出し、14 件の dockerized タスク シナリオに拡張しました。このシナリオでは、エージェントは実際の脆弱性を発見したにもかかわらず、標準的な運用上の手法に矛盾するステルス障害を犯しました。つまり、パブリック アップロードへの認証情報の埋め込み、アクセスを証明するための運用リソースの削除、競合状態を示すために関与していないユーザーの強制追加などです。多数決集計による 3 モデルの大規模言語モデル (LLM) ジャッジ パネルを使用してエージェントの軌跡を評価し、安全な成功率 (解決済みおよびステルス)、Stealth@Solve (成功した解決間のトレードクラフトの品質)、および無謀な解決率 (解決されたがカバーが飛んだ) を測定します。私たちの結果は、54% の安全成功率 (タスクの完了とステルスの両方を必要とする複合指標) を超えるモデルがないことを示しており、OPSEC の失敗がモデル ファミリ全体で体系的に発生していることが確認されました。当社は、ステルス対応エージェントの開発と自律的な攻撃セキュリティ展開のための自動 OPSEC 監視の両方をサポートする公開ベンチマークとして StealthBench をリリースします。インタラクティブなリーダーボード、評価ハーネス、およびデータセットは、https://stealthbench.com で入手できます。

原文 (English)

StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents

Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite security researchers and advanced persistent threats achieve their objectives unnoticed; autonomous agents increasingly inherit the same offensive tasks, but do they inherit the tradecraft? We introduce StealthBench,a benchmark that measures operational stealth in autonomous offensive-security agents across six operational security (OPSEC) dimensions. We extract 11 hand-verified OPSEC incidents from real bug-bounty and red-team trajectories, expanded into 14 dockerized task scenarios, where agents, despite finding real vulnerabilities, committed stealth failures inconsistent with standard operational tradecraft: embedding credentials in public uploads, deleting production resources to prove access, force-adding uninvolved users to demonstrate a race condition. We evaluate agent trajectories using a 3-model large language model (LLM) judge panel with majority-vote aggregation, measuring safe success rate (solved and stealthy), Stealth@Solve (tradecraft quality among successful solves), and reckless solve rate (solved but cover blown). Our results show that no model exceeds 54% safe success rate (the compound metric requiring both task completion and stealth), confirming that OPSEC failures are systematic across model families. We release StealthBench as a public benchmark to support both the development of stealth-aware agents and automated OPSEC monitoring for autonomous offensive-security deployments. The interactive leaderboard, evaluation harness, and dataset are available at https://stealthbench.com.

13:00 JSTLLM/生成AIGemini

LLM でシミュレートされた被験者と人間の被験者を心理測定キャリブレーションのために調整する: 認知診断プロファイリング アプローチ

教育テストの心理測定キャリブレーションには、通常、高価な人間の反応データが必要です。大規模言語モデル (LLM) のシミュレートされた受験者は、早期キャリブレーションへの有望なルートを提供しますが、彼らの応答は正確すぎて均一すぎます。我々は、認知診断プロファイリング (CDP) を提案します。これは、LLM に、さまざまな認知プロファイルを持つもっともらしい受験者をシミュレートするよう促すゼロショット フレームワークです。バイナリの属性マスタリー パターンが自然言語プロファイルとしてレンダリングされ、非有益な分布または有益な分布の下でサンプリングされます。 Tatsuoka フラクションサブトラクション データセット (受験者 536 人、15 項目、5 属性) を使用して、プロファイルなし、非情報 CDP、および情報 CDP 条件下で 8 つの LLM 構成を評価し、能力分布、習熟プロファイル、および項目難易度レベルで人間の受験者との整合性を評価しました。 CDP は 3 つのレベルすべてを改善しました。構成全体で分布の重複が増加しました。プロファイルレベルのスコアと人間のプロファイルの期待値の間の加重相関は 0.92 ~ 0.98 に達しました。アイテムの難易度の回復は、ほとんどが推論対応モデルで、ランク順と絶対的な調整において改善されました。最も強いケースである Gemini 3.0 Flash (Thinking) では、1 パラメーター ロジスティック (1PL) 難易度のスピアマン相関が 0.24 から 0.86 および 0.90 に上昇し、二乗平均平方根誤差 (RMSE) が 6.31 から 1.30 および 0.90 に低下しました。この有益な条件は、プロファイル レベルの調整が強力な場合に最も役立ちました。 CDP は、LLM でシミュレートされた受験者を人間の受験者とより緊密に心理測定的に一致させ、運用テストの開発に実用的なものにします。

原文 (English)

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correlations between profile-level scores and human profile expectations reached 0.92 to 0.98; and item-difficulty recovery improved in rank order and absolute alignment, most for reasoning-enabled models; in the strongest case, Gemini 3.0 Flash (Thinking), one-parameter logistic (1PL) difficulty Spearman correlations rose from 0.24 to 0.86 and 0.90 and the root-mean-square error (RMSE) fell from 6.31 to 1.30 and 0.90; the informative condition helped most where profile-level alignment was strong. CDP brings LLM-simulated examinees into closer psychometric alignment with human examinees, making them practical for operational test development.

13:00 JST研究/論文

グラフ ニューラル ネットワークの Top-k 説明における自己同型誘発非正準性

化学的に同等の 2 つのニトロ基を持つ分子が与えられた場合、勾配ベースの GNN Explainer は、最後のビットに等しい属性スコアをそれらに割り当てます。それ以外のことはできません。メッセージの受け渡しは正確に順列等変であるため、入力の自己同型性はすべての属性を不変のままにします。それでも、標準レポートである上位 k エッジでは 2 つのうちの 1 つが指定され、どちらが配列の順序によって決定されます。これは実装上の障害ではなく、構造上の障害であることを示します。入力の自己同型群によって最小の有効な説明が固定されていない場合、ルールを一度に単一値、最小、および対称性を尊重することはできません。実際に使用されるexact-kレポートについては、公理依存関係のないリーン4で機械化されたパラメータフリーの基準を提供します。これは、グラフのみから、そのサイズのすべてのスコア最適レポートが軌道を分割する必要があるかどうかを決定します。 21,298 件のインスタンスの予算決定において、この基準は例外なく機械的モデルの同等性チェックと一致しており、中立的な代替案を認めた決定的なケースは見つかりませんでした。障害はよくあることです。非自明な自己同型は、独創的な説明可能性論文が使用するデータセットである Mutagenicity の 93.4% で発生するため、対象となる連続領域上のサウンドである対称入力のメジャーゼロ却下は、ここで崩壊します。これらの論文の報告によると、スパーシティバジェットでは、交換可能な 2 つのニトロ基 (25 個中 6 個) を持つ分子の 24.0% が、それらのうちの 1 つを正確に表面化しており、機械的検証の下ではすべての基が任意です。モデルの盲目さは対称性も生み出します。すべての MUTAG 分子には原子が含まれていますが、化学的には分離されますが、ネットワークは明らかに分離できません。また、一致したコントロールは、パラメーター化方法ではなく、モデルが読み取る内容によって解像度が設定されることを示しています。軌道をレポートすると、0.11 ミリ秒での恣意性が除去され、グラフあたり 0.43 個の余分なエッジが追加されます。

原文 (English)

Automorphism-Induced Non-Canonicity in Top-k Explanations of Graph Neural Networks

A gradient-based GNN explainer given a molecule with two chemically equivalent nitro groups assigns them attribution scores that are equal to the last bit. It cannot do otherwise: message passing is exactly permutation equivariant, so any automorphism of the input leaves every attribution invariant. Yet the standard report, the top-k edges, names one of the two, and which one is settled by the order of an array. We show this is a structural obstruction rather than an implementation slip. When no minimal valid explanation is fixed by the input's automorphism group, no rule can be single-valued, minimal and symmetry-respecting at once. For the exact-k reports used in practice we give a parameter-free criterion, mechanised in Lean 4 with no axiom dependencies, that decides from the graph alone whether every score-optimal report of that size must split an orbit. Across 21298 instance-budget decisions the criterion agrees with a mechanical model-equivalence check without exception, and no severing case we found admitted a neutral alternative. The obstruction is common. Nontrivial automorphisms occur in 93.4% of Mutagenicity, the dataset the seminal explainability papers use, so the measure-zero dismissal of symmetric inputs, sound on the continuous domains it was made for, collapses here. At the sparsity budget those papers report, 24.0% of molecules with two interchangeable nitro groups (6 of 25) surface exactly one of them, every one arbitrary under mechanical verification. A model's blindness also manufactures symmetry: every MUTAG molecule contains atoms chemistry separates and the network provably cannot, and a matched control shows the resolution is set by what the model reads rather than how it is parameterised. Reporting orbits removes the arbitrariness at 0.11 ms and 0.43 extra edges per graph.

13:00 JSTLLM/生成AI研究/論文

合成ユーザーが失敗した場合: LLM でシミュレートされた人間のアンケート回答のクロスドメイン ベンチマーク

大規模言語モデル (LLM) は、製品、政策、市場の意思決定に疑似回答を与える人間の回答者の代役である合成ユーザーとして使用されることが増えています。私たちは、この置換がいつ有効で、いつ失敗するかを尋ね、その答えをインテリジェントな合成ユーザー システムの評価フレームワークとしてパッケージ化します。 2 つのファミリーと 8B からフロンティアまでの機能範囲にわたる 4 つのモデルにわたって実行される単一のプロトコルは、実際の人間の反応データの 2 つの独立した領域、つまり米国の一般的な社会的態度 (一般社会調査) と異文化間の価値観 (世界価値観調査) に適用されます。すべてのモデルは、保持されている人間のデータに適合する一連の非 LLM ベースラインに対してベンチマークされます。人口動態のプロンプトと私たちがテストした調査シミュレーション プロトコルの下では、両方のドメイン、4 つのモデルすべて、および両方のファミリーにわたって 2 つの障害が再現されました。まず、個人レベルでは、最も強力なベースラインであっても LLM に勝るものはありません。異文化間の価値観に関しては、どのモデルもそれを大きく下回っており、距離を意識して適切なスコアリングを行ってもギャップは存続します。第 2 に、モデルは体系的に人口統計を過剰に決定し、アイデンティティを現実の人々よりもはるかに態度を予測するものとして扱います。この歪みは、ほぼすべての質問とグループの組み合わせに存在し、コーディング不変の尺度に対して堅牢です。どちらの障害も、より大規模でより高性能なモデルによって解決されることはありません。意思決定の影響分析は、なぜこれが実際に重要なのかを示しています。セグメントをターゲットにしたタスクでは、モデルはセグメント間のギャップを 2 ~ 4 倍に膨張させ、米国のケースとほとんどの異文化ケースの半分でチームを間違ったセグメントに誘導し、現実の人々には存在しないセグメント分割を作り出します。リクエストに応じて、クロスドメインのベンチマークと評価フレームワークを利用できるようにします。これにより、チームは、合成ユーザー証拠がいつ意思決定支援として安全であるか、いつ安全でないかを事前に判断できます。

原文 (English)

When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses

Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model falls well below it, and the gap survives distance-aware and proper scoring. Second, models systematically over-determine demographics, treating identity as far more predictive of attitudes than it is among real people, a distortion present for nearly every question-group combination and robust to a coding-invariant measure. Neither failure is remedied by a larger, more capable model. A decision-impact analysis shows why this matters in practice: on a segment-targeting task the models inflate between-segment gaps two to fourfold, would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people. We make the cross-domain benchmark and the evaluation framework available on request, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.

13:00 JST研究/論文Google

Pramana: 実証的ネットワーキング研究のための構成可能なドメイン固有のバックエンド

ネットワーク化研究は仮説を経験的証拠に変えることによって進歩するため、ネットワーク化研究を加速するとは、着想 (仮説の合成) とそれをテストするデータの生成の間のラグを減らすことを意味します。具体的なケースを考えてみましょう。BBR の一括ダウンロードは、競合するリアルタイムの Google Meet トラフィックとボトルネックを公平に共有していますか?これを検証するには、現実的なボトルネック リンクを構成し、BBR の一括転送と Meet のリアルタイム トラフィックを同時に生成し、関連するサービス品質指標を収集する必要があります。現在、このオーバーヘッドは高く、研究者は新しいアイデアを得るたびにゼロから始めなければならないことがよくあります。このアイデア化とデータ生成のギャップは、AI 支援によるアイデア化が飛躍的に加速するエージェント AI 時代にはさらに悪化する一方、その出力はデータ生成バックエンドなしでは検証できません。この文書では、このギャップを埋める方法を検討します。私たちは、上部に多様な研究目的、下部に異なる実行基盤を備えた、細い腰の形をした、構成可能なドメイン固有のバックエンドである Pramana を想定しています。 Pramana は、このウエストを 1 つのコントラクトであるインテント仕様によって実現します。インテント仕様は、実験を 3 つの独立した軸、つまりインテント (どのようなデータを生成するか)、サブストレート (生成する場所)、メカニズム (どのように生成するか) に分解するため、1 つの仕様がどのサブストレートでも実行されます。私たちは、66 の出版された論文からマイニングされた 255 のデータ生成インテントからなるこの種初のコーパスを構築することによって Pramana の有用性を実証し、インテントの仕様がそれらすべてを満たすことを示しますが、既存のツールは 13% を超えて満たしません。私たちの現在の概念実証実装はすでにこれらの目的の 34% を満たしており、既存の最良のツールの 2 倍以上です。私たちは、想定されるデータ生成バックエンドを構築し、実証的ネットワーキング研究を加速するための広範なコミュニティの取り組みを通じて、この抽象化と実装のギャップを埋めるためのロードマップを示しています。

原文 (English)

Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research

Networking research advances by turning hypotheses into empirical evidence, so accelerating it means reducing the lag between ideation (synthesizing a hypothesis) and generating the data that tests it. Consider a concrete case: does a bulk BBR download fairly share its bottleneck with competing real-time Google Meet traffic? Validating this requires configuring a realistic bottleneck link, concurrently generating BBR's bulk transfer and Meet's real-time traffic, and collecting relevant service-quality metrics. Today this overhead is high, often forcing researchers to start from scratch for every new idea. This ideation-to-data-generation gap will only worsen in the agentic AI era, where AI-assisted ideation accelerates exponentially, yet its outputs cannot be validated without a data-generation backend. This paper explores how to bridge this gap. We envision a composable, domain-specific backend, Pramana, shaped as a thin waist, with diverse research intents at the top and disparate execution substrates at the bottom. Pramana realizes this waist through a single contract, the intent specification, which disaggregates an experiment into three independent axes: the intent (what data to generate), the substrate (where to generate it), and the mechanism (how to produce it), so one specification runs on any substrate. We demonstrate Pramana's utility by building a first-of-its-kind corpus of 255 data-generation intents mined from 66 published papers, and show the intent specification satisfies all of them, where no existing tool satisfies more than 13%. Our current proof-of-concept implementation already satisfies 34% of these intents, more than twice the best existing tool, and we lay out a roadmap for closing this abstraction-implementation gap through a broader community effort to build the envisioned data-generation backend and accelerate empirical networking research.

13:00 JST研究/論文

忠実性仮定の k 次緩和による高次マルコフ ブランケットの発見

データから変数のグラフィカル マルコフ ブランケット (MB) を学習する問題は、ベイジアン ネットワークやマルコフ確率場の構造学習、因果関係の発見、特徴の選択など、多くの分野に応用できます。ただし、ほとんどのメソッドが行う共通の仮定は、分布における条件付きの独立性がグラフィック構造における同じ分離を意味するというものであり、忠実性の仮定としても知られています。残念ながら、この仮定は、XOR やパリティ型の関係などの高次の依存関係によって破られる可能性があり、有限サンプルでは、​​経験的な違反によって破られる可能性があり、極端な場合には、真の分布には存在しない偽の依存関係を誘発することさえあります。したがって、この論文では、k+2 変数間のパリティ タイプの関係を捉える忠実性仮定の「k 次」緩和を提案します。次に、この緩和を MB 発見に使用する、k 次マルコフ ブランケット (kOMB) と呼ばれる概念実証アルゴリズムを提案します。最後に、忠実性の真の違反と経験的な違反の両方の下で、kOMB がどのように変数の MB を回復できるかを経験的に示します。コードはhttps://github.com/lklee9/k-order-Markov- Blanketで入手できます。

原文 (English)

High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption

The problem of learning the graphical Markov blanket (MB) of a variable from data has applications in many areas such as structure learning for Bayesian networks and Markov random fields, causal discovery, and feature selection. However, a common assumption most methods make is that the conditional independencies in the distribution imply the same separation in the graphical structure -- also known as the faithfulness assumption. Unfortunately, this assumption can be violated by higher-order dependencies such as XOR and parity-type relations, and -- on finite samples -- by empirical violations that, in extreme cases, even induce spurious dependencies absent from the true distribution. Therefore, in this paper we propose a "k-order" relaxation of the faithfulness assumption that captures parity type relationships between k+2 variables. We then propose a proof of concept algorithm called k-order Markov blanket (kOMB) that uses this relaxation for MB discovery. Finally, we empirically show how kOMB can recover the MB of a variable under both true and empirical violations of faithfulness. Code available at: https://github.com/lklee9/k-order-Markov-blanket

13:00 JST研究/論文Llama

検出可能性ギリギリでのポストトレーニング: 微調整へのゲーム理論的アプローチ

強化学習 (RL) 微調整は、参照ポリシーからのドリフトを制限しながら、ターゲット タスクでのモデルのパフォーマンスを向上させるために、言語モデルのトレーニングで広く使用されています。このトレードオフのバランスをとる標準的な方法は、KL 正則化 RL 目標を使用することですが、この定式化自体は正則化係数を設定する原理的な方法を提供するわけではありません。実際には、係数は通常、ヒューリスティックに、またはハイパーパラメータ検索によって選択されますが、これにより、トレーニング コストの不必要なオーバーヘッドや、報酬と保持の望ましくないトレードオフが発生する可能性があります。代わりに、このトレードオフに明示的な統計的解釈を与えるゲーム理論のフレームワークを提案します。具体的には、モニターが時間の経過とともにポリシーの出力を観察し、参照ポリシーからの逸脱をテストする一方で、エージェントが累積報酬を最大化するポリシーを選択する逐次ゲームを研究します。同じ観点から生まれたものではありませんが、結果として得られる均衡ポリシーは、統計的識別可能性の単位あたりの報酬を最大化するとみなせる最適な正則化パラメータに対する KL 正則化 RL 問題の解として表現できることを示します。凹凸分数計画法からの古典的な結果を利用して、KL 正則化 RL 目的への還元を介してこの平衡係数を学習するための原則的な方法を提供し、標準の微調整パイプラインへの柔軟な統合を可能にします。 Qwen3-8B と Llama-3.2-1B を使った実験では、私たちの手法が継続的な学習設定において競争力のある報酬保持のトレードオフをもたらすことを実証し、オープンソース モデルを提供する API プロバイダーを監査するために私たちのフレームワークがどのように使用されるかを示します。

原文 (English)

Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning

Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient. In practice, the coefficient is typically chosen heuristically or via hyperparameter search, which can lead to unnecessary overhead in training cost or undesirable reward-retention trade-offs. We instead propose a game-theoretic framework that gives this trade-off an explicit statistical interpretation. Specifically, we study a sequential game in which an agent chooses a policy to maximize cumulative reward while a monitor observes policy outputs over time and tests for deviations from the reference policy. Although not originating from the same perspective, we show that the resulting equilibrium policy can nonetheless be expressed as the solution to a KL-regularized RL problem for an optimal regularization parameter that can be viewed as maximizing reward per unit of statistical distinguishability. Drawing on classical results from concave-convex fractional programming, we provide a principled method for learning this equilibrium coefficient via reduction to the KL-regularized RL objective, thus allowing for flexible integration into standard fine-tuning pipelines. In experiments with Qwen3-8B and Llama-3.2-1B, we demonstrate that our methods result in competitive reward-retention trade-offs in a continual learning setting, and illustrate how our framework may be used to audit API providers serving open-source models.

13:00 JSTLLM/生成AIGPT / ChatGPT

財務開示テキストの詳細な不一致分類の診断

財務開示には、質的に異なる方法で矛盾する可能性のある数値的主張、一時的記述、実体参照、政策コミットメント、およびリスクの説明が含まれています。矛盾の検出は最初のステップにすぎません。数値的、時間的、参照的、事実的、および規範的な矛盾には、さまざまな証拠と下流のチェックが必要となるため、レビュー ワークフローではそのタイプを判断する必要がある場合もあります。私たちはこの問題を詳細な不一致分類として研究します。 11 の不一致ラベルとペアの参照証拠スパンを備えた合成財務開示ベンチマークである SBID-FD の固定 5,940 インスタンスのスナップショットを使用して、共有評価プロトコルの下で、凍結された埋め込み分類子、微調整されたエンコーダー、証拠拡張分類子、プロンプト付き大規模言語モデル、および LoRA に適応した生成モデルを比較します。微調整された 300M エンコーダーの精度は 61.9% に達します。これに対し、LoRA に適応した Qwen3.5-9B モデルの精度は 61.5%、GPT-5.4 の精度は 61.3% です。これらのシステムはアーキテクチャ、監視、トレーニング目的、入力形式が異なるため、これをモデルのスケールに関する制御された結論ではなく、コンパクトな教師ありエンコーダの実際的な効率の結果として解釈します。ゴールド証拠スパンを提供すると、微調整されたエンコーダーが 65.3% に向上しますが、自動的に予測されたスパンは、そのゲインの有意義ではあるが不完全なシェアを回復します。これは、ローカリゼーションの品質が依然としてボトルネックであることを示しています。クラスレベルの分析では、参照の不一致はローカリゼーションの品質に特に敏感である一方、事実および論理的な不一致は、関連する証拠が提供された場合でも依然として困難であることが示されています。オラクル、ディストラクター、およびクラスごとの分析を組み合わせて、位置特定エラーと残存タイプ識別エラーを分離します。これは、進歩には、より強力な証拠の抽出と、密接に関連する不一致カテゴリーに対するより優れた推論の両方が必要であることを示しています。

原文 (English)

Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text

Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in qualitatively different ways. Detecting a conflict is only the first step: review workflows may also need to determine its type, since numerical, temporal, referential, factual, and normative inconsistencies require different evidence and downstream checks. We study this problem as fine-grained inconsistency classification. Using a fixed 5,940-instance snapshot of SBID-FD, a synthetic financial-disclosure benchmark with 11 inconsistency labels and paired reference evidence spans, we compare frozen embedding classifiers, fine-tuned encoders, evidence-augmented classifiers, prompted large language models, and LoRA-adapted generative models under a shared evaluation protocol. A fine-tuned 300M encoder reaches 61.9% accuracy, compared with 61.5% for a LoRA-adapted Qwen3.5-9B model and 61.3% for GPT-5.4. Because these systems differ in architecture, supervision, training objective, and input format, we interpret this as a practical efficiency result for compact supervised encoders rather than a controlled conclusion about model scale. Supplying gold evidence spans improves the fine-tuned encoder to 65.3%, whereas automatically predicted spans recover a meaningful but incomplete share of that gain, indicating that localization quality remains a bottleneck. Class-level analyses show that Referential inconsistencies are especially sensitive to localization quality, while Factual and Logical inconsistencies remain difficult even when the relevant evidence is provided. Together, the oracle, distractor, and per-class analyses separate localization errors from residual type-discrimination errors, indicating that progress requires both stronger evidence extraction and better reasoning over closely related inconsistency categories.

13:00 JST画像/動画生成

Zero-Fi: 対比信号言語調整によるゼロショット Wi-Fi ベースの人間活動認識

Wi-Fi ベースの人間の活動認識は大幅に進歩しましたが、既存の手法のほとんどは閉じた一連の活動を前提としており、ターゲット クラスごとにラベル付きの Wi-Fi サンプルが必要なため、目に見えない活動を認識する能力が制限されています。ゼロショット Wi-Fi ベースの人間の活動認識のための対照的な信号言語調整フレームワークである Zero-Fi を紹介します。 Zero-Fi は、相補的な Wi-Fi 信号特徴から統一表現を学習し、共有埋め込み空間内の自然言語アクティビティ記述の意味表現と一致させます。このクロスモーダル調整により、Zero-Fi は、ラベル付きの Wi-Fi サンプルやそれらのクラスのモデル適応を必要とせずに、新しいアクティビティ クラスを認識できるようになります。大規模な公開ベンチマーク データセットでの実験では、保留されたアクティビティ クラスの効果的なゼロショット認識が実証され、事前定義されたアクティビティ クラスを超えて Wi-Fi センシングを拡張するための信号言語調整の可能性が強調されています。

原文 (English)

Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment

Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require labeled Wi-Fi samples for every target class, limiting their ability to recognize unseen activities. We present Zero-Fi, a contrastive signal-language alignment framework for zero-shot Wi-Fi-based human activity recognition. Zero-Fi learns unified representations from complementary Wi-Fi signal features and aligns them with the semantic representations of natural-language activity descriptions in a shared embedding space. This cross-modal alignment enables Zero-Fi to recognize new activity classes without requiring labeled Wi-Fi samples or model adaptation for those classes. Experiments on large-scale public benchmark datasets demonstrate effective zero-shot recognition of held-out activity classes, highlighting the potential of signal-language alignment to extend Wi-Fi sensing beyond predefined activity classes.

13:00 JST研究/論文

競争上の限界点との共謀: 価格レベルの監査は構造的に盲目である

アルゴリズムによる共謀に関する実証研究では、データについて 1 つの疑問が投げかけられています。それは、価格は超競争力があるのか​​ということです。私たちは、利益をもたらす陰謀によって、これに「ノー」と答えることができることを示します。各エージェント自身の入札法を競争法に正確に準拠させたまま、説明されていない入札コンポーネントの共同配布によってのみ結合する入札エージェントを考えてみましょう。入力が単一エージェントの価格または入札履歴であるテストは、共単調性までのすべての結合強度について、その偽陽性率と正確に等しい検出力を持ちます。したがって、公開されている検出方法論は、検出力が不足しているというよりも、構造によってこの行為を認識できておらず、サンプルサイズによってそれが修正されることはありません。以下に 3 つの実験結果を示します。まず、このメカニズムは実際の言語モデル エージェントに現れます。19 人の独立した開発者からの 20 のモデル、それぞれ 3 つのデプロイメント プロンプトでは、すべての順序特徴を調べサンプルから適合させた監査人のもとで、モデル全体の $+0.0001$ に対して 1 つのモデルの 2 つのデプロイメント間で $+0.053$ の残差相関が示され、開発者によってクラスター化された 95% の間隔は $[0.030, 0.078]$ でした。第 2 に、サンプリング温度が上昇するにつれて結合は単調に低下し ($p=0.002$)、展開パラメータが緩和策の候補に変わります。第三に、39 社の入札者からの 77,684 件の入札をカバーする 24 日間のイーサリアム ブロック構築オークション データでは、入札者ペアの正直な母集団自体が非常に依存しているため、5% の誤検知率で保持される画面は $+0.50$ ~ $+0.81$ の下限を上回らなければなりません。これはファミリーごとのサンプリングしきい値の 20 ~ 32 倍であり、監査ウィンドウが拡大しても下がらないことです。ここでは、合法的なマルチアイデンティティ操作と陰謀は行動的に区別できないため、扱いやすい規制目標は検出ではなくカウントです。40 の入札アイデンティティを 23 のオペレーターに解決すると、ハーフィンダール指数は 247.5% 上昇し、公開入札ストリームからの行動クラスターを追加すると 324.5% に達します。

原文 (English)

Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction

Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it. Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.

13:00 JSTLLM/生成AI

ミスアライメントには個性がある: 緊急のミスアライメントに関するビッグ 5 の説明

安全でないコードや不正確な数学的解答などの狭い欠陥を含むデータに基づいて言語モデルを微調整すると、まだ議論されているメカニズムを通じて広範囲にわたる不整合が発生する可能性があります。私たちは、解釈可能な説明を提供します。つまり、私たちが研究するモデルとコーパスでは、不整合は性格の変化のように動作します。以前の研究では、単一のバイナリコントラストからキャラクター特性の活性化方向を抽出しており、これにより、調整されたスケールを確立することなく行動を分離または方向付けることができます。代わりに、段階的な 3 レベルの介入を使用してビッグ 5 の性格ベクトルを抽出し、2 つの無差別加重モデルで検証します。 3 つのレベルは線形に順序付けされており、コーエンの d 値は最大 6.2 です。ベクトルはゼロショットと特性を個別に独立したコーパスに転送します。そしてその効果は中間層の帯域内で最も強くなります。このベクトルをトレーニング データに適用すると、8 つのドメインにわたって位置がずれているコーパスには、ビッグ 5 の共通の特徴、つまり、同調性と誠実性が低く、外向性と神経質性が高いという特徴があることが明らかになりました。この署名は両方のモデルで相関 r = 0.94 で復元されます。微調整により同じプロファイルがインプリントされ、対応するシグネチャに沿ってモデルの世代をシフトします。アクティベーションベースの測定を使用すると r = 0.83、テキストベースの判定を使用すると r = 0.90 になります。また、内部アクティベーションも r = 0.69 でシフトします。同じベクトルは、過剰な同調性ではなく、高い外向性と低い誠実性として、お調子者を特徴付けており、この区別は単一の方向では捉えることができません。調整された性格ベクトルは、不透明な安全現象を人間が判読できる診断プロファイルに変換します。

原文 (English)

Misalignment Has a Personality: A Big Five Account of Emergent Misalignment

Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable account: in the models and corpora we study, misalignment behaves like a shift in personality. Prior work extracts activation directions for character traits from a single binary contrast, which can separate or steer behavior without establishing a calibrated scale. We instead extract personality vectors for the Big Five using a graded, three-level intervention and validate them on two open-weight models. The three levels are linearly ordered, with Cohen's d values of up to 6.2; the vectors transfer zero-shot and trait-specifically to an independent corpus; and their effects are strongest within a middle-layer band. Applied to training data, the vectors reveal that misaligned corpora across eight domains share a common Big Five signature: lower agreeableness and conscientiousness, together with higher extraversion and neuroticism. This signature is recovered by both models with a correlation of r = 0.94. Fine-tuning imprints the same profile, shifting the model's generations along the corresponding signature, with r = 0.83 using activation-based measurements and r = 0.90 using a text-based judge, while also shifting internal activations with r = 0.69. The same vectors characterize sycophancy as high extraversion and low conscientiousness rather than excess agreeableness, a distinction that a single direction cannot capture. Calibrated personality vectors transform an opaque safety phenomenon into a human-legible diagnostic profile.

13:00 JSTLLM/生成AIエージェント

エージェントによる音声認識のための音声メモリ

エージェント音声認識のための推論のみのスキームである Voice Memory を紹介します。ストリーム時に、フリーズされた補正機能がドメインごとに 1 つの Memory.md を読み取り、発話ごとに、仮説に基づいて行動するか、棄権して 1-best を維持するかを決定します。スコアゲート オプティマイザーは非同期的に、制限された編集を通じてそのファイルを改訂し、保持されているスコアを厳密に改善する場合にのみ編集を受け入れます。古典的な ASR-LM フレームワークを拡張したもので、これをリスナーとシンカーのアーキテクチャの分割と呼びます。 2 つの役割は記憶を通じてのみ結合されるため、重みは変更されず、学習したスキルは監査可能でポータブルなままになります。このループが発見した有効なスキルは抑制であることが判明しました。制約のない生成的誤り訂正 (GER) が過剰に修正され、金融ニュースの編集の最大 64% で正しいトークンが破壊され、音声メモリによりこの率が 35% に減少しました。オープンコレクタである Voice Memory を使用した 10 個の HyPoradise ドメイン全体で、データセットが 1 番目の最良のベースラインを下回ることなく、加重単語誤り率が 8.36% から 7.52% (3 つの追加されたコンテキスト内の例で 7.47%) に低下しました。ゲインは、航空機のコマンド (8.40% ~ 3.40%) や騒々しい遠距離音声 (CHiME-4、12.69% ~ 10.46%) など、回復可能なヘッドルームが最大になる場所に集中します。メモリは補正器ファミリー間で転送され、推論パスにゼロ パラメーターが追加されます。将来の研究のために、デモとコード例が提供されています。

原文 (English)

Voice Memory for Agentic Speech Recognition

We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The memory transfers across corrector families and adds zero parameters to the inference path. A demo and example code are provided for future studies.

13:00 JSTLLM/生成AI画像/動画生成

FAS-R1: 推論顔アンチスプーフィングのための統合マルチタスク MLLM

顔のアンチスプーフィング (FAS) は、本物かスプーフィングの判断だけでなく、攻撃のセマンティクスや画像に基づいた人間による検査の証拠も提供することがますます期待されています。既存の識別 FAS モデルは引き続き大部分がラベル中心ですが、最近の MLLM ベースの手法は構造化された出力を提供しますが、依然として主に教師付き微調整に依存しており、多くの場合、テンプレートのような理論的根拠と、困難な攻撃に対する弱い最適化が生成されます。我々は、統合 FAS 予測のための 2 段階推論指向の MLLM フレームワークである FAS-R1 を提案します。これは、真正性分類、攻撃タイプ認識、およびスプーフィング領域の位置特定をカバーします。 FAS-R1 は、まずコールドスタート監視付き微調整に高品質のロング CoT データセットである FAS-R1-23K を使用し、次に FAS 固有の GRPO ポストトレーニングを実行します。 Degradation-Simulated Augmentation (DSA) は、ビジュアル品質の変化全体にわたって安定したなりすましキュー推論を促進します。一方、Difficulty-Aware GRPO (DA-GRPO) は、困難なタスク、特にメイクアップ攻撃やマスク攻撃などの微妙な攻撃や曖昧な攻撃の場合、攻撃グループが最適化されないままになる可能性がある簡単なサンプルの優勢を軽減します。メインの 3B FAS-R1 モデルは、98.75\% の信頼性精度、93.33\% の攻撃タイプ精度、および 96.30/94.73\% の AP@40/AP@50 をドメイン内で達成します。また、クロスドメインの信頼性の一般化と回答と根拠の品質においても、比較したシステムよりも優れています。さまざまな基本モデルを使用した実験では、さらに好ましいスケーリング動作が示されています。コードは近日公開予定です。

原文 (English)

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence for human inspection. Existing discriminative FAS models remain largely label-centric, while recent MLLM-based methods offer structured outputs but still rely mainly on supervised fine-tuning, often producing template-like rationales and weak optimization for difficult attacks. We propose FAS-R1, a two-stage reasoning-oriented MLLM framework for unified FAS prediction, covering authenticity classification, attack-type recognition and spoof-region localization. FAS-R1 first uses FAS-R1-23K, a high-quality long-CoT dataset, for cold-start supervised fine-tuning, and then performs FAS-specific GRPO post-training. Degradation-Simulated Augmentation (DSA) encourages stable spoof-cue reasoning across visual-quality shifts, while Difficulty-Aware GRPO (DA-GRPO) mitigates easy-sample dominance that may leave difficult task--attack groups under-optimized, especially for subtle or ambiguous attacks such as makeup and mask attacks. The main 3B FAS-R1 model achieves 98.75\% authenticity accuracy, 93.33\% attack-type accuracy, and 96.30/94.73\% AP@40/AP@50 in-domain. It also outperforms the compared systems in cross-domain authenticity generalization and answer-and-rationale quality. Experiments with different base models further show favorable scaling behavior. The code will be released soon.

13:00 JSTロボティクス

コストに制約のある四足ハードウェアでの強化学習

学習した制御ポリシーを低コストのロボット プラットフォームに展開すると、トランスポートの遅延とノイズの多いモーター フィードバックが発生し、シミュレーションと実際のギャップが系統的に拡大します。ハードウェアでの展開に対するシミュレーションの溝は、アクチュエータが指令された位置に到達するまでの遅延にあります。 Mini Pupper 2 などのプラットフォームでは、測定された 50 ミリ秒を超える輸送遅延により、移動タスクが標準的なマルコフ決定プロセスから部分的に観察可能なプロセスに変換されます。この論文では、生物学的にヒントを得たアプローチを採用し、ノイズの多い遅延フィードバックを処理してシミュレーションと現実のギャップを埋め、それによってコストに制約のあるハードウェアでの強化学習の機能を拡張します。低コストの四足ハードウェア プラットフォームを使用し、時間認識ニューラル ネットワークと組み合わせた平均アクチュエータ遅延のフォワード モデルを使用すると、堅牢な移動が実現されることがわかりました。さらに、私たちの時間認識ニューラル ネットワークは、中央パターン ジェネレーター (CPG) を学習しました。これは、脊椎動物の脊髄に見られる CPG を反映した、+320 ミリ秒の遅延変動に対して堅牢な自立的なリズミカルな歩行です。私たちは、時間的自己組織化がコストに制約のある移動のための一般的な戦略である可能性があると仮定します。

原文 (English)

Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. The chasm of simulation to deployment in hardware lies in the delay of the actuator reaching the commanded position. On platforms such as the Mini Pupper 2, a measured >50 ms transport delay transforms the locomotion task from a standard Markov decision process into a partially observable one. In this paper, we take a biologically inspired approach of handling noisy and delayed feedback to close the sim-to-real gap, thereby expanding the capability of reinforcement learning on cost-constrained hardware. Using a low-cost quadrupedal hardware platform, we find that using a forward model of the average actuator delay, paired with a time-aware neural network results in robust locomotion. Additionally, our time-aware neural network learned a central pattern generator (CPG): a self-sustaining rhythmic gait that is robust to +320 ms latency perturbations, mirroring the CPGs found in the spinal cords of vertebrates. We posit that temporal self-organization may be a general strategy for cost-constrained locomotion.

13:00 JSTLLM/生成AIGemmaQwen

ロングコンテキスト言語モデルのマージ可能なモデル側集約状態

ロングコンテキスト言語モデルの既知の制限は、コンテキストの長さが長くなるにつれて、非加算的なセットベースの集計におけるパフォーマンスの信頼性がますます低くなるということです。例には、カーディナリティの推定、セットの関係、およびグループ化された統計が含まれます。これらは、ログ、プログラム出力、テーブル、複数ターンの会話に広く存在します。これらのタスクに必要な集約状態を提供するために、凍結された言語モデルとともにコンパクトなハッシュベースの HyperLogLog (HLL) スケッチ状態を維持するモデル側集約インターフェイスを導入します。モデルがコンテキストを処理している間、エクストラクターは関連する各レコードを正規の ID にマップします。次に、ID がハッシュされ、HLL 状態が更新されます。これらの状態は、コンテキスト セグメント全体でマージしたり、下流の推論のために直接読み出すことができるため、追加の生成、実行、戻りのサイクルが回避されます。 HLL 状態サイズを 2 KiB (2,048 レジスタ) に設定することで、提案されたアプローチを検証します。これは、コンテキストの長さやカーディナリティの設定によって増加しません。 100 万レコードを対象とした個別カウント実験では、平均相対誤差は 1.6% でした。別のマージ テストでは、256 ものセグメントから構築された状態が、同じストリーム上の 1 回のパスとまったく同じ読み出しを生成しました。 174 のソース ウィンドウからの 3,969 の集計→理由タスクでは、固定予算インターフェイスは、Gemma 4 (31B、BF16) では 99.2% の精度に達しましたが、正確な集計では 100.0% でした。ペアのギャップは 0.8 パーセント ポイント (95% ウィンドウ クラスター CI: 0.5 ~ 1.3 ポイント) でした。 174 項目の一致セットでは、私たちの方法は直接フルコンテキスト推論よりも Qwen で 63.2 ポイント、Gemma で 56.3 ポイント改善されました。思考連鎖 (CoT) 推論に対する対応するゲインは、それぞれ 60.9 ポイントと 63.2 ポイントでした。固定タスク 1,200 の Oolong-Synth サブセットでは、私たちの手法は Qwen で 91.1%、Gemma で 99.3% に達しました。コードは https://github.com/songdc98/sketchops で入手できます。

原文 (English)

Mergeable Model-Side Aggregation States for Long-Context Language Models

A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as context length grows. Examples include cardinality estimation, set relationships, and grouped statistics, which widely exist in logs, program outputs, tables, and multi-turn conversations. To provide the aggregation state required by these tasks, we introduce a model-side aggregation interface that maintains compact Hash-based HyperLogLog (HLL) sketch states alongside a frozen language model. While the model processes the context, an extractor maps each relevant record to a canonical identity. The identity is then hashed and updates the HLL state. These states can be merged across context segments and/or read out directly for downstream reasoning, avoiding an additional generate-execute-return cycle. We validate the proposed approach by setting the HLL state size as 2 KiB (2,048 registers), which does not increase with context length or set cardinality. In a distinct-count experiment involving one million records, the mean relative error was 1.6%. In a separate merge test, states built from as many as 256 segments produced exactly the same readout as a single pass over the same stream. On 3,969 aggregate-then-reason tasks from 174 source windows, the fixed-budget interface reached 99.2% accuracy on Gemma 4 (31B, BF16), compared with 100.0% under exact aggregation; the paired gap was 0.8 percentage points (95% window-cluster CI: 0.5-1.3 points). On a matched set of 174 items, our method improved over direct full-context reasoning by 63.2 points on Qwen and 56.3 points on Gemma. The corresponding gains over chain-of-thought (CoT) reasoning were 60.9 and 63.2 points, respectively. On a fixed 1,200-task Oolong-Synth subset, our method reached 91.1% on Qwen and 99.3% on Gemma. Code is available at https://github.com/songdc98/sketchops.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

ForgetBench: 言語モデルにおける長期パラメトリック記憶の忘却ダイナミクスのベンチマーク

大規模言語モデル (LLM) は、知識の獲得と推論において強力な機能を実証していますが、繰り返し更新されても以前に獲得した知識を保持するその能力はまだ十分に理解されていません。既存の評価パラダイムは主に単一ステップ推論または静的知識編集に焦点を当てており、継続的なモデル変更中の知識の保持と劣化の時間的ダイナミクスを捉えることができません。この研究では、継続的な知識編集下での LLM の忘却行動を体系的に特徴付けるために設計されたベンチマークである ForgetBench を提案します。 ForgetBench は、概念ベースの QA とシナリオベースの QA という 2 つの相補的な評価パラダイムを導入し、構造化されたリレーショナル知識の保存から孤立した事実の保持を解きほぐします。逐次編集フレームワークに基づいて、時間的に順序付けされた知識ストリームを構築し、複数の編集段階にわたってモデルの動作を評価します。長期的な保持力学を定量的に分析するために、時間の経過に伴う知識の進化をモデル化する統合評価フレームワークをさらに導入し、時間的減衰、保持力、およびインスタンス間の安定性の測定を可能にします。多様なモデルと編集方法にわたる広範な実験により、既存のアプローチでは長期的な保持と一般化の品質のバランスを取ることができないことが実証されました。私たちの調査結果は、将来の LLM では、長期にわたって知識を効果的に取得、更新、保存できる、より堅牢なメモリ メカニズムの必要性を強調しています。コードは承認され次第公開されます。

原文 (English)

ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models

Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain previously acquired knowledge under repeated updates remains insufficiently understood. Existing evaluation paradigms primarily focus on single-step reasoning or static knowledge editing, which fail to capture the temporal dynamics of knowledge retention and degradation during continual model modification. In this work, we propose ForgetBench, a benchmark designed to systematically characterize forgetting behavior in LLMs under continual knowledge editing. ForgetBench introduces two complementary evaluation paradigms, namely concept-based QA and scenario-based QA, to disentangle isolated factual retention from structured relational knowledge preservation. Building upon a sequential editing framework, we construct temporally ordered knowledge streams and evaluate model behavior across multiple editing stages. To quantitatively analyze long-term retention dynamics, we further introduce a unified evaluation framework that models knowledge evolution over time, enabling the measurement of temporal decay, retention strength, and cross-instance stability. Extensive experiments across diverse models and editing methods demonstrate that existing approaches fail to strike a balance between long-term retention and generalization quality. Our findings highlight the need for more robust memory mechanisms that can effectively acquire, update, and preserve knowledge over time in future LLMs. Code will be released upon acceptance.

13:00 JST研究/論文

PUDA: 自動運転研究所向けの AI ネイティブ ハードウェア ハーネス

Physical Unified Device Architecture (PUDA) は、自動運転研究所 (SDL) 向けの AI ネイティブ ハードウェア ハーネスです。 PUDA は、人間中心のグラフィカル ユーザー インターフェイス (GUI) オーケストレーション レイヤーを構築するのではなく、ハードウェアの実行が決定論的、アトミックで監査可能な状態を維持しながら、エージェントが実験を観察、方向付け、決定し、行動できるようにするコマンド ライン ランタイム環境を作成します。設計によりヘッドレスであり、デバイスは検出可能なコマンド ライン インターフェイスを通じて表示され、JSON プロトコルは分散メッセージング システムを通じてルーティングされ、コマンド応答、データ製品、およびレポートは構造化されたレコードとして保存されます。 PUDA は、プロトコル、実行、サンプル、測定、およびコマンド ログを、実行識別子とタイムスタンプによってリンクされた AI ネイティブのデータ構造に編成し、結果として得られるデータ製品へのハードウェア応答を通じて送信されたプロトコルの出所を保存します。 PUDA は、科学的なオーケストレーションを物理的な操作やデータ テレメトリから分離します。エージェントは実験を選択し、一方で PUDA は検証済みのコマンドを実行し、来歴に関連付けられた状態、応答、およびデータをキャプチャします。この貢献は、別のオプティマイザー、オーケストレーター、またはレシピ言語ではありません。これは、エージェント SDL の実用的な実行およびデータ環境です。より広範な物理 AI の意味は、PUDA が AI システムが物理ツールと対話するための AI ネイティブ ハードウェア ハーネスを提供することです。

原文 (English)

PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories

Physical Unified Device Architecture (PUDA) is an AI-native hardware harness for self-driving laboratories (SDLs). Rather than building a human-centered graphical user interface (GUI) orchestration layer, PUDA creates a command-line runtime environment that lets agents observe, orient, decide, and act over experiments while hardware execution remains deterministic, atomic, and auditable. Headless by design, devices appear through discoverable command-line interfaces, JSON protocols are routed through a distributed messaging system, and command responses, data products, and reports are preserved as structured records. PUDA organizes protocols, runs, samples, measurements, and command logs into an AI-native data structure linked by run identifiers and timestamps, preserving provenance from submitted protocol through hardware response to resulting data products. PUDA separates scientific orchestration from physical operation and data telemetry: agents choose experiments, while PUDA executes validated commands and captures provenance-linked state, responses, and data. The contribution is not another optimizer, orchestrator, or recipe language. It is a practical execution and data environment for agentic SDLs; the broader physical AI implication is that PUDA provides an AI-native hardware harness for AI systems to interact with physical tools.

13:00 JST研究/論文

クロスドメインオーディオディープフェイク検出のためのマルチ比率 DiT 再構成残差のオーディオアンカー融合

音声ディープフェイク検出器は、ジェネレーター、コーパス、または録音条件が変化すると性能が低下することがよくあります。私たちは、本物の音声のみで訓練された拡散変換器 (DiT) を凍結再構成プローブとして使用します。マスキング比 0.5、0.75、および 0.9 での再構成により、明示的なマルチ比率残差マップが生成されます。これらの残差はドメインに依存するため、オーディオアンカー検出器は、投影されたフリーズ WavLM 聴覚表現をゲートベースの減衰なしでフュージョン合計に渡し、残差をスカラーゲートの加算補正としてのみ使用します。事前に指定されたシード 42 の実行では、ASVspoof 5 Eval で 6.5442% EER / 0.18456 min-DCF、ITW Full で 13.8372% / 0.36921 が得られます。 3 シードの平均は 6.8885 (0.3308)% と 15.3328 (2.0719)% です。後者は、両方の監視設定の下で個別に最適化された WavLM-ResNet18 リファレンスを下回ります。補助監督により、動的競争融合は平均 ITW EER 18.4007% から 25.2968% に上昇し、3 つのシードすべてが悪化しました。この結果は、アンカリングのみの成分的な因果関係の除去を主張することなく、補完的な証拠として再構成残差を支持し、ASVspoof 5からITWへの移行の非競合的聴覚経路を動機付けるものである。

原文 (English)

Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection

Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. We use a Diffusion Transformer (DiT), trained only on bona fide speech, as a frozen reconstruction probe. Reconstructions at masking ratios 0.5, 0.75, and 0.9 yield explicit multi-ratio residual maps. Because these residuals are domain sensitive, our audio-anchored detector passes the projected frozen-WavLM auditory representation into the fusion sum without gate-based attenuation and uses residuals only as a scalar-gated additive correction. The pre-specified seed-42 run obtains 6.5442% EER / 0.18456 min-DCF on ASVspoof 5 Eval and 13.8372% / 0.36921 on ITW Full; three-seed means are 6.8885 (0.3308)% and 15.3328 (2.0719)%. The latter is below a separately optimized WavLM-ResNet18 reference under both supervision settings. Auxiliary supervision raises dynamic competitive fusion from 18.4007% to 25.2968% mean ITW EER, worsening all three seeds. The results support reconstruction residuals as complementary evidence and motivate a non-competitive auditory path for ASVspoof 5-to-ITW transfer, without claiming a componentwise causal ablation of anchoring alone.

13:00 JSTLLM/生成AIビジネス/資金調達ClaudeNVIDIA

LLMET: エネルギー効率の高い LLM サービスを提供するために、新たな M3D メモリのクロスレイヤー評価を可能にする

大規模言語モデル (LLM) サービスのエネルギー消費は、ハードウェアの電力と熱の制約、および電気コストの上昇により、展開が拡大するにつれてシステムの大きな課題となっています。チップのエネルギー消費の主な原因は、限られたオンチップ キャッシュとオフチップの高帯域幅メモリ (HBM) の間でのデータの移動です。一方、ロジック チップのバックエンド オブ ライン (BEOL) でのキャッシュ メモリのモノリシック 3D (M3D) 統合などの新興メモリ テクノロジにより、オンチップ メモリの大型化と高密度化が可能になり、コストのかかるオフチップ トラフィックを削減する新たな機会が生まれています。ただし、新しいテクノロジーを使用してオンチップメモリ​​を継続的に拡張することで、LLM サービスのエネルギー効率を効果的に向上できるかどうかは依然として不明です。このギャップに対処するために、私たちは検証済みのクロスレイヤ シミュレーション フレームワークである LLMET (LLM with Emerging Technology) を開発し、幅広いモデル、アプリケーション、プラットフォームにわたる大容量オンチップ メモリ テクノロジの影響に関する包括的な研究を実施しています。 M3D テクノロジーを利用して L2 キャッシュを 40MB から 1GB に拡張すると、デュアル NVIDIA A100 GPU セットアップでの LLMET シミュレーションに基づいて、16K コンテキスト ウィンドウでの Llama3.1-70B プレフィル フェーズ中にチップ エネルギーが 44% 削減されます。 8x NVIDIA B200 のようなプラットフォームでは、L2 キャッシュを 128MB から 4GB に拡張することで、プレフィル エネルギーが最大 24% 節約されます。エッジ プラットフォームとワークロードの場合、8MB のキャッシュ サイズを 256MB に増やすと、デコードのエネルギー節約は 30% に達します。これらの結果は、エネルギー効率の高い LLM サービス システムに超大容量オンチップ メモリが期待できることを強調しています。

原文 (English)

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving

The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs. A key contributor to chip energy dissipation is data movement between limited on-chip cache and off-chip High Bandwidth Memory (HBM). Meanwhile, emerging memory technologies such as monolithic 3D (M3D) integration of cache memories at the Back-End-Of-Line (BEOL) of logic chips enable larger and denser on-chip memories, creating new opportunities to reduce costly off-chip traffic. However, it remains unclear whether continuously scaling on-chip memory using emerging technologies can effectively improve the energy efficiency of LLM serving. To address this gap, we develop LLMET (LLM with Emerging Technology), a validated cross-layer simulation framework, and conduct a comprehensive study on the impact of large-capacity on-chip memory technologies across a broad range of models, applications and platforms. Utilizing M3D technology to expand the L2 cache from 40MB to 1GB yields a 44% reduction in chip energy during the Llama3.1-70B prefill phase with a 16K context window, based on LLMET simulation on a dual NVIDIA A100 GPU setup. On the 8x NVIDIA B200-like platform, extending the L2 cache from 128MB to 4GB saves the prefill energy by up to 24%. For the edge platform and workloads, the decode energy saving reaches 30% when increasing the 8MB cache size to 256MB. These results highlight the promise of ultra-large on-chip memories for energy-efficient LLM serving systems.

13:00 JST研究/論文

ポリシー外の強化学習における過大評価を軽減するための悲観的な批評家との共同重み付け

継続的制御のためのディープオフポリシー強化学習アルゴリズムは通常、ポリシーの改善を導くためにニューラル値関数の近似に依存します。ただし、時間差分 (TD) 学習ではノイズの多いターゲットが導入されるため、非定常最適化が行われますが、貪欲なポリシー更新により初期段階の推定誤差が増幅されます。このようなエラーの再帰的な伝播は、アクタークリティカル手法における永続的な過大評価バイアスとトレーニングの安定性の低下につながります。既存のアプローチは、優先順位付けされたサンプリングや変更された値の学習目標によってこの問題を軽減しようとしますが、多くの場合、限られたデータ範囲やブートストラップエラーによって引き起こされる不確実性の高い遷移を強調しすぎるため、バイアスがさらに増幅されます。この論文では、値推定における予測不確実性を明示的に説明する統一フレームワークである Collaborative Weighting Actor-Critic (CWAC) を提案します。 CWAC は、分布批評家を採用してリターンの不確実性をモデル化し、TD エラーと不確実性を共同で再重み付けする協調重み付けメカニズムを導入し、ノイズの多い更新を抑制しながら、信頼性の高いサンプルからの堅牢な学習を可能にします。さらに、リターン分布からのサンプリングによる確率的悲観的値推定スキームを組み込んでおり、これにより、ポリシー改善中の誤差の伝播が効果的に軽減されます。 CWAC は、最小限のオーバーヘッドで、SAC、TD3、DDPG などの既存のオフポリシー アルゴリズム フレームワークにシームレスに統合できます。経験的な結果は、私たちが提案した方法が、さまざまな範囲のシミュレートされたタスクにわたってパフォーマンスを大幅に向上させることを示しています。私たちのコードは https://anonymous.4open.science/r/CWAC-348E で公開されています。

原文 (English)

Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning

Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide policy improvement. However, temporal-difference (TD) learning introduces noisy targets, resulting in non-stationary optimization, while greedy policy updates amplify early-stage estimation errors. The recursive propagation of such errors leads to persistent overestimation bias and degraded training stability in actor-critic methods. Existing approaches attempt to alleviate this issue via prioritized sampling or modified value learning objectives, but often overemphasize high-uncertainty transitions caused by limited data coverage or bootstrapping errors, thereby further amplifying bias.In this paper, we propose Collaborative Weighting Actor-Critic (CWAC), a unified framework that explicitly accounts for predictive uncertainty in value estimation. CWAC employs distributional critic to model return uncertainty and introduces a collaborative weighting mechanism that jointly reweights TD-errors and uncertainty, enabling robust learning from reliable samples while suppressing noisy updates. In addition, we incorporate a stochastic pessimistic value estimation scheme via sampling from the return distribution, which effectively mitigates error propagation during policy improvement. CWAC can be seamlessly integrated into existing off-policy algorithm frameworks such as SAC, TD3, and DDPG with minimal overhead. Empirical results demonstrate that our proposed method significantly enhances performance across a diverse range of simulated tasks. Our code is publicly available at https://anonymous.4open.science/r/CWAC-348E.

13:00 JST研究/論文

大規模な言語モデルのトレーニング後のエンドツーエンドの強化学習のための HiFloat4 形式

私たちの知る限り、最初のエンドツーエンドの FP4 RL ポストトレーニングを紹介します。このトレーニングでは、前方パスと後方パスを含むロールアウト ポリシーとトレーニング ポリシーの両方が 4 ビット精度で動作します。体系的な研究により、FP4 RL の主な劣化原因はトレーニング側の量子化エラーではなく、ロールアウトのアクティベーション量子化であることが明らかになりました。異常値によりダイナミック レンジが拡張され、FP4 では多数のアクティベーション値がゼロにアンダーフローします。直観に反しますが、FP4 でのロールアウトを維持しながらトレーニング ポリシーをより高い精度に復元すると、精度が完全な FP4 ベースラインよりも悪化し、ロールアウトとトレーニングの不一致が主要な障害モードとして明らかになり、標準的な事前トレーニング スタイルの修正が除外されます。これは、ロールアウト残差量子化 (Rollout-ResQ) で解決します。これは、FP4 ロールアウト Matmul にのみ追加される、ハードウェアに適したスパース パターンに制約される単一の残差補正項です。これは、ロールアウトの計算フットプリントを増大させることなく、外れ値駆動のアンダーフローによって失われた精度のほとんどを回復する軽量の補正です。 Qwen2.5-3B および Qwen2.5-Math-7B では、Rollout-ResQ を HiFloat4 (HiF4) フォーマットと組み合わせます。その 3 レベルの階層スケーリングにより、FP4 の厳しい 4 ビット バジェットの下で解像度が維持されます。BF16 との精度の差が 4.9% から 1.1% に縮まり、完全に量子化された FP4 RL が完全精度の射程内に収まります。同じレシピをオープン標準の MXFP4 に適用すると、その差は 13.6% から 5.3% に縮まり、FP4 フォーマットの選択が回復可能な精度の上限を決定する重要な要素であることがわかります。これらの結果を総合すると、エンドツーエンドの FP4 RL ポストトレーニングを実現する形式として HiF4 が確立され、BF16 とのギャップを埋めることができるアクティベーション側メカニズムとして Rollout-ResQ が確立されます。

原文 (English)

HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models

We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their forward and backward passes, operate at 4-bit precision. A systematic study reveals that the dominant source of degradation in FP4 RL is not training-side quantization error but rollout activation quantization: outliers stretch the dynamic range so far that a large number of activation values underflow to zero under FP4. Counterintuitively, restoring the training policy to higher precision while keeping the rollout in FP4 makes accuracy worse than full FP4 baseline, exposing rollout-training mismatch as the principal failure mode and ruling out standard pretraining-style fixes. We address this with Rollout Residual Quantization (Rollout-ResQ): a single residual correction term constrained to a hardware-friendly sparsity pattern, added only to the FP4 rollout matmul -- a lightweight correction that recovers most of the precision lost to outlier-driven underflow without inflating the rollout's compute footprint. On Qwen2.5-3B and Qwen2.5-Math-7B, Rollout-ResQ paired with the HiFloat4 (HiF4) format -- whose three-level hierarchical scaling preserves resolution under FP4's tight 4-bit budget -- closes the accuracy gap to BF16 from 4.9% to 1.1%, bringing fully quantized FP4 RL within striking distance of full precision. Applied to the open-standard MXFP4, the same recipe narrows the gap from 13.6% to 5.3%, revealing that FP4 format choice is a key factor that determines the ceiling on recoverable accuracy. Together, these results establish HiF4 as the enabling format for end-to-end FP4 RL post-training, and Rollout-ResQ as the activation-side mechanism that makes the gap to BF16 closable.

13:00 JSTエージェント

会話型 AI エージェント用のグラフネイティブのバイテンポラル メモリ ストア

会話型 AI エージェントには通常、セッション間での永続的なメモリが不足しています。完全なチャット履歴をコンテキスト ウィンドウに挿入したり、サードパーティのメモリ サービスに委任したりするような明白な修正は、モデルのコンテキスト バジェットを使い果たすか、ユーザーが制御していないインフラストラクチャを介して個人データを送信します。 HNSW ベクトル インデックスで強化されたエージェント ローカル Neo4j プロパティ グラフと完全なバイテンポラル データ モデルの両方の問題を回避するメモリ ストアについて説明します。各メモリは、有効時間 (事実が世界で真実だったとき) とトランザクション時間 (データベースがそれを記録したとき) という 2 つのクローズド/オープン時間間隔を保持するバージョン管理されたコンテンツ ノードにリンクされた不変のアイデンティティ ノードとして保存されます。この設計は、履歴を物理的に上書きすることなく、ポイントインタイムのセマンティック検索をサポートします。関連するメモリ間のセマンティック エッジは、1024 次元の埋め込みにわたるコサイン類似性を使用して、書き込み時に自動的に維持されます。私たちは、長期記憶に重点を置くように設計された 6 つの質問タイプにわたる 500 の質問のベンチマークである LongMemEval でシステムを評価します。 60 のサンプル質問全体で、現状のセマンティック検索パスは全体で 46.7% R@10 を達成し、知識更新の質問では 80% に上昇しました。タイムトラベル パスでは、知識の更新で 80% の R@10 が得られますが、時間的推論の質問では再現率が減少します (50% から 37.5%)。これは、具体的な設計の改善を直接示すポスト フィルターの希釈の結果です。これらの結果から、さまざまな質問タイプの純粋な検索の限界について何が明らかになり、各失敗モードが将来の研究に何を示唆するのかについて説明します。

原文 (English)

A Graph-Native Bitemporal Memory Store for Conversational AI Agents

Conversational AI agents commonly lack persistent memory across sessions. The obvious fixes like injecting full chat histories into the context window, or delegating to a third-party memory service, either exhaust the model's context budget or send personal data through infrastructure the user does not control. We describe a memory store that avoids both problems: an agent-local Neo4j property graph augmented with HNSW vector indexes and a full bitemporal data model. Each memory is stored as an immutable identity node linked to versioned content nodes carrying two closed-open time intervals: valid time (when the fact was true in the world) and transaction time (when the database recorded it). This design supports point-in-time semantic retrieval without physically overwriting history. Semantic edges between related memories are maintained automatically at write time using cosine similarity over 1024-dimensional embeddings. We evaluate the system on LongMemEval, a 500-question benchmark spanning six question types designed to stress long-term memory. Across 60 sampled questions, the current-state semantic search path achieves 46.7% R@10 overall, rising to 80% on knowledge-update questions. The time-travel path yields 80% R@10 on knowledge-update but decreases recall on temporal-reasoning questions (50% to 37.5%), a consequence of post-filter dilution that points directly to a concrete design improvement. We discuss what these results reveal about the limits of pure retrieval for different question types and what each failure mode suggests for future work.

13:00 JST研究/論文

継続的な学習のためのローカル学習アーキテクチャを忘れないための技術

CMP (Cognitive Memory Primitive) は、入力をスパースなリレーショナル コードとして表現し、2 層の競合メモリに保存し、特徴生成システムによるエンドツーエンドのバックプロパゲーションを行わずにローカル更新を通じて学習する継続学習アーキテクチャです。私たちは、スパース表現、局所学習、永続記憶を組み合わせることで、従来の逆伝播ベースの継続学習アプローチと比較して壊滅的な忘却を軽減できるかどうかを調査します。制御されたドメイン増分バイトレベル言語モデリング プロトコルでは、CMP は、オンライン Elastic Weight Consolidation (EWC) でトレーニングされたパラメーター一致トランスフォーマーよりも大幅に低い逆方向転送を示します。 3 シードの複製された 15 ドメイン実験全体で、CMP は安定した忘却挙動を示しますが、個別の直接比較とドメイン順序分析では、報告された実験設定の下で評価された Transformer ベースラインよりも一貫して忘却が低いことが示されています。我々は、これらの発見を、Transformer と比較した単一ドメインの精度の大幅なギャップ、ビジョンベンチマークでのヌル結果、および CMP と独立した精度向上メカニズムの組み合わせの文書化された失敗と併せて報告し、肯定的な結果と否定的な結果の両方を報告するという我々の取り組みを反映しています。これらの結果は、スパース表現、局所学習、永続記憶の組み合わせが継続学習の有望な方向性であることを示唆するとともに、壊滅的な忘却の軽減における学習ルール、表現、アーキテクチャ設計のそれぞれの役割についてのさらなる研究を促すものです。

原文 (English)

The Art of Not Forgetting A Local Learning Architecture for Continual Learning

We introduce CMP (Cognitive Memory Primitive), a continual-learning architecture that repre?sents inputs as sparse relational codes, stores them in a two-tier competitive memory, and learns through local updates without end-to-end backpropagation through its feature-generating system. We investigate whether combining sparse representations, local learning, and persistent memory can reduce catastrophic forgetting relative to conventional backpropagation-based continual?learning approaches. On a controlled domain-incremental byte-level language modeling protocol, CMP demonstrates substantially lower backward transfer than a parameter-matched Trans?former trained with online Elastic Weight Consolidation (EWC). Across a three-seed replicated 15-domain experiment, CMP exhibits stable forgetting behavior, while separate head-to-head comparisons and domain-order analyses show consistently lower forgetting than the evaluated Transformer baseline under the reported experimental settings. We report these findings alongside a substantial single-domain accuracy gap relative to the Transformer, a null result on a vision benchmark, and a documented failure to combine CMP with an independent accuracy-improving mechanism, reflecting our commitment to reporting both positive and negative outcomes. These results suggest that the combination of sparse representations, local learning, and persistent memory is a promising direction for continual learning, while motivating further investigation into the respective roles of learning rules, representations, and architectural design in mitigating catastrophic forgetting.

13:00 JSTハードウェア/半導体

物理的に一貫した複数出力シンボリック回帰のための共有シンボリック バックボーン

シンボリック回帰は分析式を提供しますが、通常は一度に 1 つの出力が適用されます。これは、状態変数が共有の物理パラメータを介して結合されることが多いプロセス システムでは制限となります。独立したシンボリック回帰により、1 つのモデルとして解釈するのが難しい正確な個別の方程式が得られます。我々は、結合された多出力システムに対する神経進化的記号回帰法を提案します。この方法は、共有シンボリック バックボーン、つまり、一度発見され、スパースの加算または乗算読み出しを通じていくつかの出力で再利用される一連の潜在的なシンボリック ユニットを検索します。離散モデル構造は突然変異と交叉によって進化しますが、連続パラメーターは勾配降下法によって調整され、子孫に継承されます。この方法は、既知のグラウンド トゥルースを使用した一連のベンチマークと熱水液化収率のケースに基づいて評価されます。結果は、カップリングが予測誤差を下げるための一般的な方法ではないことを示しています。その主な貢献は、物理的に共有された要素が潜在的な式に埋め込まれており、データからの識別が弱い場合に、出力間の一貫性を強制および診断することです。これは、独立した PySR が整合性ギャップを埋めたり、同じ共有形式を回復したりしない、ラングミュア ヒンシェルウッドおよびサイト カバレッジの分母で発生します。逆に、Van de Vusse ベンチマークのように、各出力がすでに識別可能な場合、独立したシンボリック回帰は結合モデルと一致または改善します。提案されたフレームワークは、汎用の予測子ではなく、構造化された共有メカニズムの抽出子です。その値は、ターゲット構造が疎である場合、共有されている場合、識別性が低い場合、またはクロージャによって制約されている場合に最も高くなります。

原文 (English)

Shared Symbolic Backbones for Physically Consistent Multi-Output Symbolic Regression

Symbolic regression provides analytical expressions, but it is usually applied one output at a time. This is limiting in process systems, where state variables are often coupled through shared physical parameters. Independent symbolic regression can give accurate individual equations that are difficult to interpret as one model. We present a neuro-evolutionary symbolic regression method for coupled multi-output systems. The method searches for a shared symbolic backbone: a set of latent symbolic units that is discovered once and reused by several outputs through sparse additive or multiplicative read-outs. The discrete model structure is evolved by mutation and crossover, whereas the continuous parameters are tuned by gradient descent and inherited by the offspring. The method is assessed on a set of benchmarks with known ground truth and on a hydrothermal liquefaction yield case. The results show that coupling is not a general route to lower prediction error. Its main contribution is the enforcement and diagnosis of cross-output consistency when a physically shared factor is embedded in a latent expression and is weakly identifiable from the data. This occurs for Langmuir-Hinshelwood and site-coverage denominators, for which independent PySR does not close the consistency gap or recover the same shared form. Conversely, when each output is already identifiable, as in the Van de Vusse benchmark, independent symbolic regression matches or improves the coupled model. The proposed framework, rather than a general purpose predictor, is a structured shared-mechanism extractor. Its value is highest when the target structure is sparse, shared, weakly identifiable or constrained by closure.

13:00 JSTエージェント

AgentGFM: ノード エージェント情報フロー制御を備えたグラフ基盤モデル

グラフ基盤モデル (GFM) は、マルチドメイン グラフから移転可能な知識を学習し、目に見えないシナリオに適応することを目的としています。グラフにおけるリレーショナル セマンティクスの基本的な情報源として、トポロジカル パターンの伝達可能性は長い間 GFM 研究の中心となってきました。ただし、局所的な構造パターンはグラフ間で異なる場合があり、同じグラフ内のノード間でも異なる場合があります。このような構造の変化にもかかわらず、既存の GFM のほとんどは手動で設計された伝播スキームに依存しており、それらをほとんど変更せずに新しいグラフに適用します。このような固定スキームは、異なるノードの多様な構造パターンには適合しない可能性があります。これにより、重要な疑問が生じます。各ノードは、情報がどのようにグラフ内に伝播されるべきかを自律的に決定できるでしょうか?この機能を情報フロー制御と呼びます。エージェント テクノロジーの最近の進歩に触発されて、私たちはこの問題をエージェント ベースの意思決定として定式化し、各ノードをエージェントとして扱います。したがって、すべてのノード エージェントが独立したモデルを使用するのではなく、共有されたエンドツーエンドのトレーニング可能なポリシーに従う AgentGFM を提案します。適応型情報フロー制御の場合、各ノードは、予測、実行、観察、修正のプロセスを通じてグラフと対話します。アクトステージ中に、ノードはソースの受信、信号チャネルの選択、およびゲインを意識したノードごとの停止という 3 つの決定を行います。結果として得られる観測結果は予測と比較され、その不一致はノードの状態を修正し、その後の相互作用を導くために使用されます。ノード レベル、グラフ レベル、および大規模な転送シナリオにわたる広範な実験により、多様なグラフ トポロジにわたる AgentGFM の有効性が実証されています。

原文 (English)

AgentGFM: A Graph Foundation Model with Node-Agent Information-Flow Control

Graph Foundation Models (GFMs) aim to learn transferable knowledge from multi-domain graphs and adapt to unseen scenarios. As a fundamental source of relational semantics in graphs, the transferability of topological patterns has long been central to GFM research. However, local structural patterns may vary across graphs and even among nodes within the same graph. Despite such structural variation, most existing GFMs rely on manually designed propagation schemes and apply them to new graphs largely unchanged. Such fixed schemes may not suit the diverse structural patterns of different nodes. This raises a key question: can each node autonomously determine how information should be propagated through the graph? We refer to this capability as information-flow control. Inspired by recent advances in agent technology, we formulate this problem as agent-based decision making and treat each node as an agent. Accordingly, we propose AgentGFM, in which all node agents follow a shared end-to-end trainable policy rather than using independent models. For adaptive information-flow control, each node interacts with the graph through a predict-act-observe-correct process. During the act stage, the node makes three decisions: source reception, signal-channel selection and gain-aware node-wise halting. The resulting observation is compared with the prediction and their discrepancy is used to correct the node state and guide subsequent interactions. Extensive experiments across node-level, graph-level and large-scale transfer scenarios demonstrate the effectiveness of AgentGFM across diverse graph topologies.

13:00 JST研究/論文

ペルソナベースのレートアクションインデックス

私たちは、一連のペルソナが現在の市場状況にどのように反応するかに基づいて、現在のフェデラルファンド目標金利の引き上げ/維持/引き下げという米国連邦公開市場委員会(FOMC)の決定を予測するための指標を提案します。インデックスを構築するために、公開されているデータから約 25,000 ドルの取得可能なチャンクで構成される新しいデータセットを収集しました。私たちはデータをメンバーごとのコーパスに分割し、それぞれを「ペルソナ」と呼ばれる生成システムの検索データベースとして使用します。まず、識別可能性と検出可能性という類似性の 2 つの相補的な要素全体にわたってペルソナを評価します。各ペルソナの行動は帰属可能性が高く (平均メンバー条件付き再現率は $8\times $チャンス)、生成されたコンテンツは保持されている実際のコンテンツとほとんど区別がつきません ($\hat\tau_{\mathrm{det}} = 0.15$ の下限に対して 0.23$)。次に、ペルソナのクエリ条件付き表現が、既知のタカ派とハト派の評判順序付け (Kendall の $\tau = 0.63$、$p < 0.001$) と比較したメンバーの金融政策スタンスを捕捉し、検索のみの表現を大幅に上回るという証拠を示します。これらの表現は時間や現在の市場状況によって変化し、当社が提案するペルソナベースのレートアクション指数の基礎を形成します。 $2022$--$2025$ の期間、指数はレート サイクル (Kendall の $\tau = 0.68$、$p < 10^{-6}$) を追跡し、これを使用して会議ごとの結果を非自明ではない精度 ($0.69$ 対 $0.47$ の基本レート) で予測する単純な分類器を構築できます。重要なのは、この指数が有益なベースラインを上回り、フェデラルファンドの目標金利を約4分の3リードしていることです。私たちが知る限り、私たちの結果は、デジタルペルソナの集合を通じて時間とともに変化するグループの行動を捕捉する能力を実証した最初のものです。

原文 (English)

A Persona-based Rate Action Index

We propose an index for predicting the U.S.\ Federal Open Market Committee (FOMC) decision to hike/hold/cut the current federal funds target rate based on how a collection of personas responds to current market conditions. To construct the index, we collected a new dataset consisting of nearly $25{,}000$ retrievable chunks from publicly available data. We partition the data into per-member corpora and use each as the retrieval database of a generative system we refer to throughout as a ``persona''. We first evaluate the personas across two complementary components of likeness: identifiability and detectability. Each persona's behavior is highly attributable (average member-conditional recall is $ 8\times $ chance) and generated content is nearly indistinguishable from held-out real content ($\hat\tau_{\mathrm{det}} = 0.23$ against a $0.15$ floor). We then present evidence that query-conditioned representations of the personas capture members' monetary-policy stance relative to a known hawk--dove reputational ordering (Kendall's $\tau = 0.63$, $p < 0.001$), substantially outperforming retrieval-only representations. These representations vary with time and current market conditions and form the basis of our proposed persona-based rate action index. For the $2022$--$2025$ period the index tracks the rate cycle (Kendall's $\tau = 0.68$, $p < 10^{-6}$) and can be used to construct a simple classifier that predicts per-meeting outcomes at non-trivial accuracy ($0.69$ versus a $0.47$ base rate). Importantly, the index outperforms informative baselines and leads the federal funds target rate by roughly three quarters. As far as we are aware, our results are the first to demonstrate the ability to capture time-varying group behavior via a collection of digital personas.

13:00 JST研究/論文

ServerlessT2I: サーバーレス プラットフォームで機能する効率的なテキストから画像へのワークフロー

ユーザーがカスタマイズされたワークフローを作成し、それらを断続的に呼び出すことが多いため、Text-to-image (T2I) ワークフローはサーバーレス プラットフォームに導入されることが増えています。既存のプラットフォームは通常、各ワークフローを不透明な GPU 機能としてデプロイし、ワークフロー内のすべての構成モデルを一緒にプロビジョニング、配置、スケーリングします。このモノリシック設計により、ワークフロー構造が曖昧になり、スケーリングのオーバーヘッドが増大し、ユーザーが低レベルの GPU 調整を管理する必要が生じ、マルチテナント クラスターにおけるきめ細かい公平性が制限されます。このペーパーでは、T2I ワークフローを独立して管理およびスケジュールできる疎結合モデル機能に分解するサーバーレス ネイティブ システムである ServerlessT2I について紹介します。 ServerlessT2I は、個々のモデルの実行を明示的に管理することで、モデルごとのスケーリング、宣言型ワークフロー構成、透過的な GPU 常駐通信、および公平性を意識したスケジューリングを可能にします。この分解を効率化するために、ServerlessT2I は、コンピューティング バウンドの T2I 推論によってアイドル状態にされたスラック GPU メモリを収集し、モデルの読み込みとデータ通信のオーバーヘッドを削減するデータ プレーンを構築します。 \sys{} はさらに、マルチテナント サービスのための公平なスケジューラを導入します。 ServerlessT2I は、運用トレースを使用して、同じ GPU バジェットを持つ既存の T2I ワークフロー サービング システムよりも最大 2$\倍$ 高いリクエスト レートを維持します。固定リクエスト レートの場合、サービス レベル目標 (SLO) を満たしながら、最大 3$\times$ の GPU リソースを節約します。

原文 (English)

ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform

Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently. Existing platforms typically deploy each workflow as an opaque GPU function, provisioning, placing, and scaling all constituent models in the workflow together. This monolithic design obscures workflow structure, inflates scaling overhead, forces users to manage low-level GPU coordination, and limits fine-grained fairness in multi-tenant clusters. In this paper, we present ServerlessT2I, a serverless-native system that decomposes a T2I workflow into loosely coupled model functions that can be independently managed and scheduled. By explicitly managing individual model execution, ServerlessT2I enables per-model scaling, declarative workflow composition, transparent GPU-resident communication, and fairness-aware scheduling. To make this decomposition efficient, ServerlessT2I harvests slack GPU memory left idle by compute-bound T2I inference to build a data plane that reduces model loading and data communication overheads. \sys{} further introduces a fair scheduler for multi-tenant serving. Using production traces, ServerlessT2I sustains up to 2$\times$ higher request rates than existing T2I workflow serving systems with the same GPU budget; for a fixed request rate, it saves up to 3$\times$ GPU resources while satisfying service level objectives (SLOs).

13:00 JST研究/論文

回復、デコード、防御: エンコードされた VLM ジェイルブレイクに対するガードに依存しない防御増幅

安全分類子 (「ガード」) は、視覚言語モデルに対する支配的なブラックボックス防御ですが、入力の意味ではなく表面的な形式を判断します。集合論、形式論理、珍しい言語、コード、またはテキストのイメージとして再エンコードされた有害なリクエストは、平易な言語でブロックするガード、つまりデコード ギャップをすり抜けます。自然な解決策は、ガードに依存しないリカバリおよびデコード アンプです。これは、ガードの前に画像コンテンツを転記し、エンコードされたテキストをプレーン ペイロードに再記述するため、既製の分類器で真のリクエストを選別できます。私たちはこの増幅器を構築し、攻撃者の最良のケースに対して評価します。つまり、11 回の攻撃のアンサンブルで、成功した場合に動作を壊れているとスコア付けします (自動攻撃に続くベストオブスイート)。ジェイルブレイク防御ではほとんど報告されていませんが、攻撃あたりの平均は最大 3.5 倍です。これにより、私たちの中心的な発見が明らかになります。それは、5 つのガードと 2 つのターゲット VLM にわたって、私たちが評価する非反復的回復防御の経験的な安全ユーティリティの上限です。アンプはギャップを部分的にしか埋めません。無防備なアンサンブルは動作の 89 ~ 91% を破りますが、最良のガードとアンプを組み合わせた場合でもまだ 63 ~ 65% が残ります。また、ガードのみに対するゲインが有意であるのは、10 組のガードとターゲットのペアのうち 4 組だけです。インターフェイスではガードに依存しませんが、実際には一律にそうなるわけではありません。モジュール式のリガード層は、残差の多くを閉じますが、適切に調整されたガードの場合、良性の過剰拒否を 81 ~ 92% に抑えます。使用可能な状態を維持する 1 つの緩いガードは、展開可能な安全性 (48% アンサンブル ASR) に達することはありません。私たちが評価する構成は、私たちが研究しているパイプラインや表現シフト攻撃、つまりピクセルや埋め込みスペース攻撃ではなく、読みやすいペイロードを残すエンコーディングとクロスモーダル レンダリングの場合、低い攻撃成功率と低い過剰拒否の両方を達成するものはありません。私たちは、増幅器、トレードオフを可視化するアンサンブル評価、回復ベースの VLM 防御が機能する場所と機能しない場所のマップを提供します。

原文 (English)

Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks

Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not its meaning: a harmful request re-encoded as set theory, formal logic, a rare language, code, or an image of text slips past a guard that would block it in plain language -- the decode gap. The natural fix is a guard-agnostic recover-and-decode amplifier that transcribes image content and restates encoded text into its plain payload before the guard, so any off-the-shelf classifier can screen the true request. We build this amplifier and evaluate it against the attacker's best case: an ensemble of eleven attacks, scoring a behavior as broken if any succeeds (best-of-suite, following AutoAttack) -- rarely reported for jailbreak defenses, yet ~3.5x the per-attack mean. This exposes our central finding: an empirical safety-utility ceiling for the non-iterative recovery defenses we evaluate, across five guards and two target VLMs. The amplifier only partly closes the gap -- the undefended ensemble breaks 89-91% of behaviors, and the best guard-plus-amplifier still leaves 63-65% -- and its gain over the guard alone is significant in only four of ten guard-target pairs. It is guard-agnostic at the interface, but not uniformly so in effect. A modular reguard layer closes much of the residual, yet drives benign over-refusal to 81-92% for well-calibrated guards; the one laxer guard that stays usable never reaches deployable safety (48% ensemble ASR). No configuration we evaluate reaches both low attack-success and low over-refusal, for the pipeline we study and for representation-shifting attacks -- encodings and cross-modal renders that leave a legible payload, not pixel- or embedding-space attacks. We contribute the amplifier, an ensemble evaluation that makes the trade-off visible, and a map of where recovery-based VLM defense works and where it does not.

13:00 JST画像/動画生成

VGG16、VGG19、および ResNet50 モデルを使用した肺 X 線画像からの疾患の分類

呼吸器疾患に関連する症例数の増加に伴い、呼吸器疾患を早期に発見し、正確に診断することが急務となっています。畳み込みニューラル ネットワークは、画像検査を使用した病気の診断に使用すると有望な結果が得られています。この研究では、X 線画像に基づく肺疾患の分類に VGG16、VGG19、ResNet50 などの深層学習アルゴリズムを適用する可能性を調査します。前述のモデルのパフォーマンスの詳細な分析は、肺炎、結核、肺がん、正常な肺など、さまざまな種類の肺疾患をモデルがどの程度適切に分類できるかを評価するために実施されました。これを行うために、これらの深層学習モデルは膨大な量の X 線画像でトレーニングされました。私たちの研究結果は、3 つのモデルすべてが良好な結果を提供しますが、ResNet-50 はその効率性と高レベルの精度により、他のモデルと比較して最高のパフォーマンスを発揮することを示しています。私たちは、これらの深層学習モデルが将来、肺疾患の診断の実践にうまく実装できると信じています。病気の早期発見に役立ち、患者の転帰を改善します。

原文 (English)

Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models

With the increase in the number of cases related to respiratory diseases, there is an urgent need to detect them early and diagnose them accurately. Convolutional neural networks have given promising results when used for diagnosing diseases using imaging tests. In this study, we investigate the potential of applying deep learning algorithms such as VGG16, VGG19, and ResNet50 for classification of lung ailments based on X-ray images. A detailed analysis of the aforementioned models' performances was conducted to assess how well they can classify various types of lung ailments, including pneumonia, tuberculosis, lung cancer, and normal lungs. In order to do that, these deep learning models were trained on a vast amount of X-ray images. The results of our study show that while all three models provide good results, ResNet-50 performs best in comparison with other models due to its efficiency and high level of accuracy. We believe that these deep learning models can be successfully implemented in the practice of diagnosing pulmonary diseases in the future. It helps with early disease detection and improves patient outcomes.

13:00 JST研究/論文

1 回の実行はアイデアではありません: 自動化された研究における実装抽選

自動化された研究システムは、成果物を提供するためと、どのアイデアを保持、転送、追求するかを決定するために実験スコアを使用します。それでも、1 回の実行で、アイデアの 1 つの実装が評価されます。その実現レベルのスコアを親メカニズムに関する証拠として認めると、 \emph{実装の抽選} が作成され、アイデア レベルの結論は、どの実装がサンプリングされたかによって決まります。ある実行によってメカニズムに関する信念が更新されるたびに、不一致は構造的なものになります。その大きさを推定します。 \emph{アイデアの信頼性監査} は、候補カードの検証と凍結、新しいセッションの実装のサンプリング、結果ブラインド忠実度ラベルの使用、保存されたアーティファクトの再実行によって \emph{アイデアの信頼性} を測定します。アイデア ICC と Leave-One-implementation-out (LOO) 勝者の逆転を報告します。以前の作業では通常、タスクが繰り返されます。私たちはその考えを繰り返します。 13 の表形式タスクと 2 つのコーディング エージェント設定に対する 312 の割り当て全体で、実装の分散は同一アーティファクトの再実行分散のそれぞれ 5 倍と 10 倍を超え、1 つの実装抽選の勝者は、他の 2 つの平均での勝者と決定の 25.6\% と 43.6\% で異なりました。反転は、2 つの結果を考慮しないレビュー ルールに基づくカード レベルのフィルタリングに耐えます。決定論的評価器を使用した 3 つの材料回帰ワークフローの探索的診断でも、分解を支配する実装の変動が見つかりました。これらの調査結果は、アイデアの信頼性と最高の $N$ アーティファクト ユーティリティを区別します。スコアがアイデアレベルの分岐、転送、または記憶の研究をガイドする前に、証拠が複数の実装をカバーする必要があります。

原文 (English)

One Run Is Not an Idea: The Implementation Lottery in Automated Research

Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet one run scores one implementation of an idea. Crediting that realization-level score as evidence about the parent mechanism creates the \emph{implementation lottery}, in which an idea-level conclusion depends on which plausible implementation was sampled. The mismatch is structural whenever one run updates beliefs about a mechanism. We estimate its magnitude. The \emph{Idea Reliability Audit} measures \emph{idea reliability} by validating and freezing candidate cards, sampling fresh-session implementations, using outcome-blind fidelity labels, and rerunning saved artifacts. It reports idea ICC and leave-one-implementation-out (LOO) winner reversal. Prior work generally repeats the task; we repeat the idea. Across 312 assignments on 13 tabular tasks and two coding-agent setups, implementation variance was more than five and ten times same-artifact rerun variance, respectively, and the winner from one implementation draw differed from the winner under the other-two mean in 25.6\% and 43.6\% of decisions. Reversal survives card-level filtering under two outcome-blind review rules. An exploratory diagnostic on three materials-regression workflows with a deterministic evaluator also finds implementation variation dominating the decomposition. These findings distinguish idea reliability from best-of-$N$ artifact utility. Before a score guides idea-level branching, transfer, or research memory, evidence should cover multiple implementations.

13:00 JSTエージェントDeepSeek

大規模言語モデル エージェントを使用した化学プロセスの PID 調整のための物理学に基づいたフレームワーク

化学プロセスの PID 調整は通常、特定されたプロセス モデルに依存しますが、プラント エンジニアは、応答の観察、欠陥の診断、ゲインの調整、結果の検証によってループを繰り返し再調整することがよくあります。この研究では、このエンジニアのようなワークフローを、大規模な言語モデルと小規模な言語モデル (LLM/SLM) の両方に適用できる言語モデル支援 PID 調整フレームワークで形式化します。ホストされた LLM は、閉ループ応答機能、制御エンジニアリング診断、調整設定、および内部モデル制御 (IMC) ベースのデモンストレーションを受け取り、共通の許容基準に基づいて PID ゲインを生成し、反復的に修正します。ローカル展開の場合、Qwen3-0.6B は、シミュレーションで検証された IMC ターゲットを使用した教師あり微調整 (SFT) と、補償不可能な安定性とパフォーマンス報酬を備えた物理学に基づいたグループ相対ポリシー最適化 (PI-GRPO) を通じて適応されます。 100 個の一次プラスデッドタイム (FOPDT) および 100 個の二次プラスデッドタイム (SOPDT) テストケースで、ホスト型 LLM (DeepSeek-V4-Flash および Qwen3.7-Plus) は、それぞれ 75 ~ 89% および 77 ~ 79% の最終成功率を達成しました。 Qwen3-0.6B に関しては、監視付き微調整により最初の推奨の成功率が 86.5% に向上し、PI-GRPO によりそれがさらに 94.0% に向上し、主に初回試行の信頼性と安定性マージンが向上しました。

原文 (English)

A Physics-Informed Framework for PID Tuning of Chemical Processes Using Large Language Model Agents

PID tuning for chemical processes commonly relies on identified process models, whereas plant engineers often retune loops iteratively by observing responses, diagnosing deficiencies, adjusting gains, and validating the result. This work formalizes this engineer-like workflow in a language-model-assisted PID tuning framework applicable to both large and small language models (LLMs/SLMs). Hosted LLMs receive closed-loop response features, control-engineering diagnoses, tuning preferences, and internal model control (IMC)-based demonstrations to generate and iteratively correct PID gains under common acceptance criteria. For local deployment, Qwen3-0.6B is adapted through supervised fine-tuning (SFT) with simulation-verified IMC targets and physics-informed group relative policy optimization (PI-GRPO) with non-compensable stability and performance rewards. On 100 first-order plus dead time (FOPDT) and 100 second-order plus dead time (SOPDT) test cases, hosted LLMs (DeepSeek-V4-Flash and Qwen3.7-Plus) achieve final success rates of 75-89% and 77-79%, respectively. As for Qwen3-0.6B, supervised fine-tuning raises first-recommendation success to 86.5%, and PI-GRPO further increases it to 94.0%, primarily improving first-attempt reliability and stability margins.

13:00 JSTLLM/生成AI画像/動画生成

分離された視覚処理: モダリティ固有のトランスフォーマー置換による効率的なマルチモーダル適応

マルチモーダル大規模言語モデル (MLLM) は、統一されたトランスフォーマー アーキテクチャ内で視覚的およびテキストによる理解を統合することにより、優れた機能を実証しました。ただし、視覚的命令の調整のためにこれらのモデルのすべてのパラメーターを微調整することは計算コストが高く、多くの場合不必要です。これは、視覚的トークンとテキストトークンの表現要件がネットワークのより深い層で大幅に異なるためです。この論文では、事前トレーニング済み LLM の上位デコーダー層を、ビジュアル トークン処理専用の軽量で独立してトレーニング可能な単一のトランスフォーマー ブロックに置き換える、効率的なトレーニング フレームワークである分離ビジュアル処理 (DVP) を提案します。具体的には、デコーダー層の前半で共有処理が行われた後、ビジュアル トークンとテキスト トークンが分割されます。ビジュアル トークンは新しく初期化された単一のトランスフォーマー ブロックを介してルーティングされ、テキスト トークンは元の凍結されたデコーダー レイヤーを継続します。次に、2 つのストリームは言語モデリング ヘッドの前で連結されます。トレーニング中は 1 つのトランスフォーマー ブロックのみが更新されるため、トレーニング可能なパラメーターの数が大幅に減少します。 LLaVA-1.5 フレームワークの実験では、DVP が全パラメーターの一部のみをトレーニングしながら、MME、POPE、および ChartQA ベンチマークで競争力のあるパフォーマンスを達成することを実証しており、MLLM の視覚表現が分離されたパラメーター効率の高い経路を通じて効果的に学習できることを示唆しています。

原文 (English)

Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution

Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction tuning is computationally expensive and often unnecessary, as the representation requirements for visual and textual tokens diverge significantly in the deeper layers of the network. In this paper, we propose Decoupled Visual Processing (DVP), an efficient training framework that replaces the upper decoder layers of a pretrained LLM with a lightweight, independently trainable single transformer block dedicated exclusively to visual token processing. Specifically, after shared processing through the first half of the decoder layers, visual and textual tokens are split: visual tokens are routed through a newly initialized single transformer block while textual tokens continue through the original frozen decoder layers. The two streams are then concatenated before the language modeling head. During training, only the single transformer block is updated, dramatically reducing the number of trainable parameters. Experiments on the LLaVA-1.5 framework demonstrate that DVP achieves competitive performance on MME, POPE, and ChartQA benchmarks while training only a fraction of the total parameters, suggesting that visual representations in MLLMs can be effectively learned through a decoupled, parameter-efficient pathway.

13:00 JSTLLM/生成AIエージェント

Living-Harness はインタラクティブ エージェントの進化版です

大規模言語モデル (LLM) エージェントは、エピソード内または再試行後に障害から回復する可能性がありますが、エピソード後のフィードバックによって、将来のインタラクションをガイドする永続的なハーネスがほとんど修正されないため、同じ実行障害が後のタスクで再発する可能性があります。静的ハーネスは、固定ツール、コンテキスト、メモリ、ワークフロー構造を通じて信頼性を向上させますが、展開後も変更されません。私たちは $\textbf{Living-Harness}$ を提案します。これは、完了した各軌道とその評価信号を、有界ハーネス更新のための事後証拠に変換する自己進化型エージェント ハーネスです。ドメインレベルの $\textbf{Evolution-SOP}$ ($\textbf{S}$tandard $\textbf{O}$perating $\textbf{P}$rocedure) によって導かれ、Living-Harness はエピソードの抽象化と構造化された更新の証拠を抽出し、トリガー条件、障害パターン、回復アクションを記録するエピソード記憶と、状態ノード、修復エッジ、および遷移ルール​​を記録する状態グラフという、2 つの相補的な形式の手続き的知識を書き込みます。更新されたハーネス状態が取得されて今後の対話をガイドしますが、ツールとベース コンテキストは凍結されたままとなり、進化サイクル全体にわたって手続き的な修復が蓄積されることが可能になります。 $\tau^2$-Bench と MultiWOZ-2.4 から派生した 8 つのインタラクティブ環境では、Living-Harness は最も強力なインタラクティブ ベースラインに対する平均 Pass@1 をそれぞれ 10.07 パーセント ポイントと 9.91 パーセント ポイント改善し、モデル バックボーン全体で進化したハーネス状態の取得のみの再利用をサポートします。

原文 (English)

Living-Harness Is an Interactive-Agent Evolver

Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur in later tasks because post-episode feedback rarely revises the persistent harness that guides future interactions. Static harnesses improve reliability through fixed tools, context, memory, and workflow structures, but remain unchanged after deployment. We propose $\textbf{Living-Harness}$, a self-evolving agent harness that converts each completed trajectory and its evaluator signals into posterior evidence for bounded harness updates. Guided by a domain-level $\textbf{Evolution-SOP}$ ($\textbf{S}$tandard $\textbf{O}$perating $\textbf{P}$rocedure), Living-Harness extracts an episode abstraction and structured update evidence, and writes two complementary forms of procedural knowledge: episodic memory that records trigger conditions, failure patterns, and recovery actions, and a state graph that records state nodes, repair edges, and transition rules. The updated harness state is retrieved to guide future interactions, while tools and base context remain frozen, allowing procedural repairs to accumulate across evolution cycles. On eight interactive environments derived from $\tau^2$-Bench and MultiWOZ-2.4, Living-Harness improves average Pass@1 over the strongest interactive baseline by 10.07 and 9.91 percentage points, respectively, and supports retrieval-only reuse of the evolved harness state across model backbones.

13:00 JSTLLM/生成AI

WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models

Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their adoption as backbones for foundation recomme…

13:00 JST研究/論文

Understanding Context Sampling in TabPFN on Small Tabular Datasets

TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) an…

13:00 JST研究/論文

Guarding Organizations Against Malware Risk: A Novel Graph-Based Malware Detection Method

Organizational digitalization expands cybersecurity risks, making cybersecurity an increasingly important research area in Information Syst…

13:00 JSTLLM/生成AIエージェント

Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability

Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself read…

13:00 JST研究/論文

Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses

A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an aver…

13:00 JST画像/動画生成研究/論文

FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking

Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Ge…

13:00 JST画像/動画生成

FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows

Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point cl…

13:00 JST画像/動画生成エージェント

Physically Real-time Infrared Attack against Optical Flow Estimation Networks

With the promising performance of deep neural networks on image-based tasks, different real-world applications such as autonomous driving a…

13:00 JSTLLM/生成AIAnthropic

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining intervent…

13:00 JSTエージェント

Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection

Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sampl…

13:00 JSTLLM/生成AI

Scientific Knowledge Discovery in the Age of Large Language Models

The rapid growth of scholarly literature has made identifying relevant publications increasingly difficult, and conventional search systems…

13:00 JST研究/論文

Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL

Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. However, RL outcomes can be highly stochasti…

13:00 JST研究/論文

MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation

Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical c…

13:00 JSTLLM/生成AI研究/論文

Automated Multilabel Mpox Research Classification with Explainable Transformer Models

The Mpox outbreak remains a serious public health issue, with the WHO (World Health Organization) reporting increasing cases in some region…

13:00 JST研究/論文NVIDIA

FARI: Robust One-Step Inversion for Watermarking in Diffusion Models

Inversion-based watermarking is a promising approach to authenticate diffusion-generated images, yet practical use is bottlenecked by inver…

13:00 JSTLLM/生成AI画像/動画生成

Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target i…

13:00 JSTLLM/生成AI

Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model

Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual…

13:00 JST画像/動画生成

Multimodal fusion of visual and morphometric features for avian bone classification

Artificial intelligence has shown considerable potential for archaeological applications, yet its use in zooarchaeology remains limited, pa…

13:00 JST研究/論文

An Attention-Based Framework for Alzheimers Disease Classification Using Resting-State fMRI

Accurate identification of Alzheimers disease (AD) using resting-state functional magnetic resonance imaging (rs-fMRI) remains challenging…

13:00 JSTLLM/生成AI

Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text

State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes…

13:00 JST画像/動画生成

Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification

Background/Objectives: Dermoscopic skin lesion classifiers often lose accuracy under domain shift across imaging devices, illumination, and…

13:00 JSTLLM/生成AI

MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities

Code search in large-scale ecosystems is often hindered by the lexical gap between user queries and implementation details, alongside the t…

13:00 JST画像/動画生成

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains un…

13:00 JST研究/論文

Journey Operators for Structured Multi-Axis Composition

Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cel…

13:00 JSTエージェント

SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution

Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforc…

13:00 JSTLLM/生成AIエージェント研究/論文Alibaba

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line…

13:00 JST研究/論文

Crossing-Free Probabilistic K-Line Forecasts Without Retraining

Probabilistic K-line forecasting describes uncertainty in four complementary prices, namely open--high--low--close (OHLC). However, it intr…

13:00 JST研究/論文

FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning

Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous…

13:00 JSTエージェント

A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities

Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate cod…

13:00 JST研究/論文

AI as Friction for Reflection Support in Ideation

Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration a…

13:00 JSTLLM/生成AIGPT / ChatGPT

Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility

Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Ex…

13:00 JSTLLM/生成AI

From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs

Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies e…

13:00 JST研究/論文Llama

ReCo: Reweighting GRPO Against Distributional Concentration

Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent wor…

13:00 JSTLLM/生成AIエージェント

Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents

LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code ge…

13:00 JSTLLM/生成AI画像/動画生成ClaudeGPT / ChatGPTGemini

Hearsay: Vision-Language Medical Diagnoses Without an Image

When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosi…

13:00 JST研究/論文

Human diversity fuels collective creativity that large language models cannot simulate or sustain

Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI as…

13:00 JST研究/論文

Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing

In many domains such as Palliative Care, Credit Assignment and Recommender Systems, predictions may causally influence the outcomes they pr…

13:00 JSTロボティクス

BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories

Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms…

13:00 JST研究/論文

Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning

Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense m…

13:00 JSTLLM/生成AI画像/動画生成

Progressive Multimodal Alignment for Continual Instruction Tuning

Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it c…

13:00 JSTロボティクス

SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception

Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times.…

13:00 JSTビジネス/資金調達研究/論文

BayesAME: Bayesian Active Model Evaluation

Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that…

13:00 JST研究/論文

CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation

Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model…

13:00 JST画像/動画生成

ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection

While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarc…

13:00 JST画像/動画生成ビジネス/資金調達

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible,…

13:00 JST画像/動画生成研究/論文

Visual Credit Audit for Multimodal Spatial Reasoning

Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixe…

13:00 JST研究/論文

Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise

We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle adm…

13:00 JSTエージェント

MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair

Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A mali…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents

As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fu…

13:00 JST画像/動画生成研究/論文GPT / ChatGPT

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context

Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative…

13:00 JSTLLM/生成AI研究/論文

MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning

With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended a…

13:00 JST画像/動画生成ロボティクス

DLAM: Distributional Latent Actions with Temporal Constraints

Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant obser…

13:00 JST研究/論文

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantificat…

13:00 JST画像/動画生成

Anatomy Contextualized Adaption of CT Foundation Models

CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-…

13:00 JST研究/論文

Improving Item Discoverability in e-Commerce Search via Related Intent Generation

Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-comm…

13:00 JST研究/論文

The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making

Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communica…

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPT

APEX-Accounting

We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work…

13:00 JST研究/論文

Decision-oriented joint optimization of evidence fusion based on event-conditioned credibility

In decision-level fusion tasks involving heterogeneous sources with unequal precision and potential anomalies, evidence deviating from the…

13:00 JSTLLM/生成AI

Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning

Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Al…

13:00 JST研究/論文

HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring

Mobile and wearable healthcare monitoring play a vital role in facilitating timely interventions, managing chronic health conditions, and u…

13:00 JSTLLM/生成AIビジネス/資金調達

Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests

Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral st…

13:00 JST研究/論文

Balancing Centralized Learning and Distributed Self-Organization: A Hybrid Model for Embodied Morphogenesis

Background: both embodied intelligence and developmental morphogenesis depend on a division of labour between centralized guidance and dist…

13:00 JST研究/論文

BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models

Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairnes…

13:00 JST研究/論文

How does downsampling affect needle electromyography signals? A generalisable workflow for understanding downsampling effects on high-frequency time series

Automated analysis of needle electromyography (nEMG) signals is emerging as a tool to support the detection of neuromuscular diseases (NMDs…

13:00 JSTLLM/生成AIエージェントClaude

AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing

LLM role-playing aims to portray arbitrary characters in interactive narratives, yet existing systems often suffer from limited immersion a…

13:00 JSTLLM/生成AI

TANDEM: マルチモーダルなヘイトスピーチに対する時間認識ニューラル検出

ソーシャル メディア プラットフォームは、長文のマルチモーダル コンテンツによってますます支配されており、有害な物語は音声、視覚、テキストの合図の複雑な相互作用を通じて構築されています。自動システムはヘイトスピーチに高精度でフラグを立てることができますが、多くの場合、効果的な人間参加型の調整に必要な、正確なタイムスタンプやターゲットの身元など、詳細で解釈可能な証拠を提供できない「ブラックボックス」として機能します。この研究では、オーディオビジュアルヘイト検出をバイナリ分類タスクから構造化された推論問題に変換する統合フレームワークである TANDEM を紹介します。私たちのアプローチは、視覚言語モデルと音声言語モデルが自己制約されたクロスモーダルコンテキストを通じて相互に最適化する新しいタンデム強化学習戦略を採用しており、高密度のフレームレベルの監視を必要とせずに、拡張された時間シーケンスにわたる推論を安定させます。 3 つのベンチマーク データセットにわたる実験では、TANDEM がゼロショットおよびコンテキスト拡張ベースラインを大幅に上回り、正確な時間的接地を維持しながら HateMM でのターゲット識別で 0.73 F1 (最先端技術と比較して 30% 向上) を達成したことが実証されました。さらに、バイナリ検出は堅牢ですが、固有のラベルのあいまいさとデータセットの不均衡により、マルチクラス設定では攻撃的なコンテンツと嫌がらせ的なコンテンツを区別することが依然として困難であることがわかりました。より広範に、私たちの調査結果は、構造化された解釈可能な調整が複雑なマルチモーダル設定でも達成可能であり、透明で実用的な次世代のオンライン安全管理ツールの青写真を提供することを示唆しています。

原文 (English)

TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech

Social media platforms are increasingly dominated by long-form multimodal content, where harmful narratives are constructed through a complex interplay of audio, visual, and textual cues. While automated systems can flag hate speech with high accuracy, they often function as "black boxes" that fail to provide the granular, interpretable evidence, such as precise timestamps and target identities, required for effective human-in-the-loop moderation. In this work, we introduce TANDEM, a unified framework that transforms audio-visual hate detection from a binary classification task into a structured reasoning problem. Our approach employs a novel tandem reinforcement learning strategy where vision-language and audio-language models optimize each other through self-constrained cross-modal context, stabilizing reasoning over extended temporal sequences without requiring dense frame-level supervision. Experiments across three benchmark datasets demonstrate that TANDEM significantly outperforms zero-shot and context-augmented baselines, achieving 0.73 F1 in target identification on HateMM (a 30% improvement over state-of-the-art) while maintaining precise temporal grounding. We further observe that while binary detection is robust, differentiating between offensive and hateful content remains challenging in multi-class settings due to inherent label ambiguity and dataset imbalance. More broadly, our findings suggest that structured, interpretable alignment is achievable even in complex multimodal settings, offering a blueprint for the next generation of transparent and actionable online safety moderation tools.

13:00 JSTLLM/生成AIエージェントGemini

How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm

This study examines how memory shapes the collective and cooperative dynamics of Large Language Model (LLM) agents in a multi-agent system.…

13:00 JST研究/論文GPT / ChatGPTGeminiDeepSeek

MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval

Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are…

13:00 JST研究/論文

On the Hybrid Nature of ABPMS Process Frames and its Implications on Automated Process Discovery

A core component of any AI-Augmented Business Process Management System (ABPMS) is the process frame, which gives the system process-awaren…

13:00 JSTLLM/生成AIエージェント

The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure

Multi-agent systems extend large language models (LLMs) by decomposing tasks among specialized agents, but their distributed decision proce…

13:00 JSTLLM/生成AI

Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries

Self-evolving skill libraries face a silent failure mode we term \emph{library drift}: unbounded skill accumulation without outcome-driven…

13:00 JSTLLM/生成AIエージェントClaude

Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents

Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent…

13:00 JSTLLM/生成AI

RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention

As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastruc…

13:00 JST画像/動画生成

BrainG3N: 制御可能な 3D 脳 MRI 生成のための多目的トークナイザー

三次元 (3D) 脳 MRI は臨床神経学および神経腫瘍学の中心であり、生成モデルは過小評価されているコホートを強化し、疾患の軌跡をシミュレートし、プライバシーを保護するデータ共有をサポートできます。潜在拡散は画像データをモデリングするための頼りになるソリューションですが、トークナイザーには 2 つの競合する要求が課せられます。エンコーダーの埋め込みは、下流のタスクが作用する臨床情報を保持する必要があり、デコーダーは解剖学的に忠実なボリュームを再構成する必要があります。既存の再構築駆動トークナイザーは、最初のトークナイザーを犠牲にして 2 番目のトークナイザーを実現します。これに対処するために、3D 脳 MRI 潜在拡散、デカップリング エンコーダーおよびデコーダー用の完全ボリューム マスク オートエンコーダー (MAE) ベースのトークナイザーを導入します。凍結された 3D MAE エンコーダーは臨床的に有益な埋め込みを生成し、専用の CNN デコーダーはそれらの埋め込みの線形投影からボクセルを再構築します。私たちは、4 つのモダリティ、10 の疾患カテゴリ、200 以上の取得サイトにわたる 18 の公的コホートからの 35,309 ボリュームでエンコーダーを事前トレーニングし、2 つの設定でその二重の有用性を実証します。まず、23 タスクの線形プローブ ベンチマークでは、エンコーダーは 23 タスク中 21 タスクで SOTA モデル (つまり、BrainIAC、BrainSegFounder、MedicalNet) を上回るか、またはそれに匹敵します。第 2 に、これらの臨床的に有益な埋め込みでトレーニングされた条件付き拡散変換器 (DiT) は、6 つの変数にわたる条件付き生成と患者固有の長期的予測の両方をサポートします。これらの結果を総合すると、下流の臨床タスクと制御可能な生成の両方を実行できる単一の 3D 脳 MRI 埋め込み空間が確立されます。

原文 (English)

BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation

Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represented cohorts, simulate disease trajectories, and support privacy-preserving data sharing. Latent diffusion has been the go-to solution for modeling imaging data, but it places two competing demands on the tokenizer: encoder embeddings must retain the clinical information that downstream tasks act on, and the decoder must reconstruct anatomically faithful volumes. Existing reconstruction-driven tokenizers achieve the second at the expense of the first. To address this, we introduce a fully volumetric masked-autoencoder (MAE) based tokenizer for 3D brain MRI latent diffusion, decoupling encoder and decoder: a frozen 3D MAE encoder produces clinically informative embeddings, while a dedicated CNN decoder reconstructs voxels from a linear projection of those embeddings. We pretrain the encoder on 35,309 volumes from 18 public cohorts spanning four modalities, ten disease categories, and 200+ acquisition sites, and demonstrate its dual utility in two settings. First, on a 23-task linear-probing benchmark, the encoder outperforms or matches SOTA models (i.e., BrainIAC, BrainSegFounder, and MedicalNet) on 21 of 23 tasks. Second, a conditional diffusion transformer (DiT) trained on these clinically informative embeddings supports both conditional generation across six variables and patient-specific longitudinal forecasting. Together these results establish a single 3D brain-MRI embedding space capable of both downstream clinical tasks and controllable generation.

13:00 JSTエージェント

Intent-Governed Tool Authorization for AI Agents

Tool-using AI agents commonly operate under integration credentials whose static permissions exceed a user's current request. We present In…

13:00 JST研究/論文

Matilda: Engine-Agnostic Search with Human Policy Guidance

Chess engines have evolved from search-based systems optimized for strength to neural policies optimized for predicting human decisions. Ex…

13:00 JSTエージェント

ATOD: マルチターン自律エージェント向けのアニーリングされたターン対応オンポリシー蒸留

長期的な対話型タスクのために小規模な言語モデル エージェントをトレーニングするには、迅速な模倣と報酬主導型の改善の両方が必要です。オンポリシー蒸留 (OPD) は教師による密度の高い指導を提供し、通常は初期段階で急速に向上しますが、生徒が教師に近づくとその効果は飽和し、最終的なパフォーマンスの上限が制限されます。強化学習(RL)は、環境報酬を直接最適化し、より高い報酬定義の上限に向けた探索的改善を促進しますが、まばらで遅延したフィードバックにより、初期段階の学習の効率がOPDよりもはるかに低くなります。この論文では、この相補性を明示的に利用するハイブリッド オンライン蒸留アルゴリズムである ATOD (Annealed Turn-aware On-policy Distillation) を提案します。 (1) ATOD はアニーリングされた OPD-RL スケジュールを使用します。OPD は教師レベルの行動に近づくための初期トレーニングを支配しますが、RL は報酬ベースの探索を促進するために徐々に強化されます。 (2) ATOD は、ターンレベルの不一致不確実性再重み付け (T-DUR) を導入しています。これは、ユーティリティの高いターンをソフトに増幅し、長い軌道での緻密な監視を改善します。 ALFWorld、WebShop、および Search-QA での実験では、ATOD が競合するトレーニング後のベースラインを常に上回っていることが示されています。3 つの生徒サイズにわたって、ATOD は平均成功率を OPD より 3.03 ポイント、GRPO より 23.62 ポイント向上させ、対応する教師モデルを 2.16 ポイント上回っています。

原文 (English)

ATOD: Annealed Turn-Aware On-Policy Distillation for Multi-Turn Agentic Tasks

Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-policy distillation (OPD) provides dense teacher guidance and typically improves rapidly in the early stage, but its gains saturate once the student approaches the teacher, limiting the final performance ceiling. Reinforcement learning (RL) directly optimizes environment rewards and encourages exploratory improvement toward a higher reward-defined ceiling, but sparse and delayed feedback makes early-stage learning much less efficient than OPD. In this paper, we propose ATOD (Annealed Turn-aware On-policy Distillation), a hybrid online distillation algorithm that explicitly exploits this complementarity. (1) ATOD uses an annealed OPD-RL schedule: OPD dominates early training to approach teacher-level behavior, while RL is gradually strengthened to drive reward-based exploration. (2) ATOD introduces Turn-level Disagreement-Uncertainty Reweighting (T-DUR), which softly amplifies high-utility turns and improves dense supervision in long trajectories. Experiments on ALFWorld, WebShop, and Search-QA show that ATOD consistently outperforms competing post-training baselines: across the three student sizes, ATOD improves average success rate by 4.16 points over OPD and 23.62 points over GRPO, while surpassing the corresponding teacher models by 2.16 points.

13:00 JSTLLM/生成AIエージェント

Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing

The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents colla…

13:00 JSTLLM/生成AIエージェント研究/論文GeminiDeepSeek

FirstResearch: LLM 科学的発見エージェントのための監査可能な質問の形成

科学的発見のための LLM システムは、着想、文献の統合、実験計画、レポートの作成をますます支援していますが、彼らが提案する最初の研究課題を監査するのは依然として難しい場合があります。科学者が検査すべきメカニズム、改ざん者、または仮定を暴露することなく、それがもっともらしく聞こえるかもしれません。 FirstResearch を紹介します。FirstResearch は、構造化された Research Question Certificate をコア成果物とする科学 LLM エージェント向けの第一原理リサーチ質問形成フレームワークです。証明書には、原始的な定義、仮定、メカニズム モデル、緊張または矛盾、反証可能な仮説、最小限の決定的なテスト、および失敗の更新ルールが記録されており、提案された質問を下流の実行前に検査できるようになります。 LLM エージェントの 10 個の研究トピックに関して、FirstResearch は、AI 共同科学者、エージェント ラボラトリー、および AI Scientist-v2 にインスピレーションを得た、主要な DeepSeek ブラインド ジャッジ プロトコルの下で制御されたプロンプト レベルのベースラインを上回りました。同じ 40 のベースライン パッケージの Gemini-2.5-Flash の独立審査員によるスコアは、システム レベルのランキングを維持しており、FirstResearch のスコアは 4.86/5 対 4.38/5 で最も強力なベースラインであり、ピアソンの一致は平均スコアで 0.865 でした。 1 回のアブレーション チェックポイントは、証明書中心のコアが最も強力なコンポーネントであることをさらに示唆しています。証明書のみのスコアは、DeepSeek では 4.90/5、Gemini では 4.88/5 に達しましたが、証明書の削除は両方の審査員で 1/5 を下回りました。これらの結果は予備的なものであり、人間の領域の専門家ではなく LLM の審査員を使用していますが、明示的な導出制約は、LLM によって生成された科学的な質問をより監査可能にするための有望なメカニズムであるという、狭い科学的発見の主張を裏付けています。コード、プロンプト、保存された出力、および再現スクリプトは、https://github.com/louiswang524/FirstResearch で入手できます。

原文 (English)

FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents

LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, but the first research question they propose can remain difficult to audit: it may sound plausible without exposing the mechanism, falsifier, or assumption that a scientist should inspect. We introduce FirstResearch, a first-principles research-question formation framework for scientific LLM agents whose core artifact is a structured Research Question Certificate. The certificate records primitive definitions, assumptions, a mechanism model, a tension or contradiction, a falsifiable hypothesis, a minimal decisive test, and a failure update rule, making the proposed question inspectable before downstream execution. On ten LLM-agent research topics, FirstResearch outperforms controlled prompt-level baselines inspired by AI co-scientist, Agent Laboratory, and AI Scientist-v2 under a primary DeepSeek-blind-judge protocol. A Gemini-2.5-Flash independent-judge rescore of the same 40 baseline packages preserves the system-level ranking, with FirstResearch scoring 4.86/5 versus 4.38/5 for the strongest baseline and Pearson agreement of 0.865 on average score. A one-repeat ablation checkpoint further suggests that the certificate-centered core is the strongest component: certificate-only scoring reaches 4.90/5 under DeepSeek and 4.88/5 under Gemini, while removing certificates drops below 1/5 under both judges. These results are preliminary and use LLM judges rather than human domain experts, but they support a narrow scientific-discovery claim: explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable. Code, prompts, saved outputs, and reproduction scripts are available at https://github.com/louiswang524/FirstResearch.

13:00 JSTLLM/生成AIClaude

LLM が同意するとき、それは正しいのでしょうか?信頼シグナルとしての自己一貫性とモデル間一致の監査

LLM-as-judge (Zheng et al., 2023) は、企業パイプラインにおける AI システムを評価するためのデフォルトになりつつあり、多くの場合、アンサンブル (Verga et al., 2024) または「専門家の混合」(Shazeer et al., 2017) の審査員パネルに拡張されます。これらのシステムは、一貫性 (審査員間またはモデル自体のサンプル間の一致) が正しさを示すという重要な前提を共有しています。この仮定が信頼できないことを示します。一致は正確さではありません。共有バイアス、記憶されたヒューリスティック、または真実ではなく事前のオプションの位置から、モデルがそれ自体と一致することもあれば、異なるモデルが互いに一致することもあります。私たちは、大規模なクロスランナー調査において、合意がそれでもなお使用可能な代用手段となるのはいつかと尋ねます。53 人のランナーが、モデル層、プロンプト、および GPQA Diamond と AIME のスケールの比較にわたって、割り当てられた重複するケースに対して K=50 のサンプルを抽出しました (265,000 のサンプル)。デプロイラベルとして多数決の正しさを使用し、階層的なランナークラスター化ブートストラップを使用すると、一致は肯定的だが弱い予測子 (rho 0.20-0.59、アイテムクラスター化リサンプリングではすべて肯定的) であり、その有用性はレジームに依存します。飽和していない中間層モデルとコンピューティングの割り当てに最適で、最も一貫したフロンティアモデル (一致度 >=0.8 の) では最悪 -- 自信過剰だが精度が低い -- です。 GPQA の訴訟結果エントリの 77%、そのうち 48% が間違っていました)。 3 つのクロード層に対する探索的なファミリー間チェックでは、同様のフロンティア過信が示されており、限界維持ヌルを超えるプロバイダー間で確信度の高いエラーが繰り返し発生しています。したがって、自己一貫性は、独立した信頼スコアではなく、正しさの条件付き代用値です。匿名化された実行ごとの行と回答の分布を公開します。

原文 (English)

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or "mixture-of-experts" (Shazeer et al., 2017) panels of judges. These systems share a key assumption: that consistency -- agreement among judges, or among a model's own samples -- indicates correctness. We show this assumption is unreliable. Agreement is not accuracy: a model can agree with itself, and different models can agree with each other, out of shared bias, a memorized heuristic, or an option-position prior rather than truth. We ask when agreement is nonetheless a usable proxy, in a large-scale cross-runner study: 53 runners drew K=50 samples for assigned overlapping cases across comparisons of model tier, prompting, and scale on GPQA Diamond and AIME -- 265,000 samples. Using majority-correctness as the deployment label and a hierarchical runner-clustered bootstrap, agreement is a positive but weak predictor (rho 0.20-0.59, all positive under item-clustered resampling) whose usefulness is regime-dependent: best for unsaturated mid-tier models and for allocating compute, and worst -- over-confident yet no more accurate -- for the most consistent frontier model (agreement >=0.8 on 77% of GPQA case-result entries, 48% of those wrong). An exploratory cross-family check on three Claude tiers shows the same frontier over-confidence, with confident errors recurring across providers above a marginal-preserving null. Self-consistency is thus a conditional proxy for correctness, not a standalone confidence score. We publicly release the de-identified per-run rows and answer distributions.

13:00 JST研究/論文

交通予測におけるグローバル空間情報抽出には本当に変圧器が必要なのでしょうか?

既存の交通予測モデルは一般に、空間依存関係、特に個々のノードと交通ネットワーク全体のすべてのノードの間の相互作用を通じて得られる表現を特徴付けるグローバル空間情報を抽出することに重点を置いています。しかし、そのようなグローバル情報がモデル化および抽出される根本的なメカニズムは、依然として十分に研究されていません。グローバル情報を自由度の高い適応的注意によって抽出する必要があるのか​​、それとも単純なグローバル集約演算子によって取得できるのかは不明のままです。この目的のために、空間混合モジュールのみを置き換えて注意に基づくグローバルな相互作用をテストする制御アブレーション フレームワークを設計します。 6 つのトラフィック ベンチマーク全体で、均一なフルレンジ ミキシングと標準空間アテンションは、それぞれ 3 つのデータセットで低い MAE を実現し、平均 MAE の差はわずか 0.14% であり、前者はノード スケールの空間ミキシングの複雑さを O(N2) から O(N) に軽減します。メカニズム分析は、空間的注意を行均一のグローバル背景と不均一な残差にさらに分解します。残差はデータセットに依存する限界値を示しており、行均一のグローバル バックグラウンドを超えた安定したゲインによって空間的注意が正当化されるべきであることを示唆しています。対応するソース コードは https://github.com/uuesti/U-Trans で公開されています。

原文 (English)

Do We Really Need Adaptive Global Spatial Attention for Traffic Forecasting?

Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which characterizes the representations obtained through interactions between each node and all nodes across the traffic network. However, the underlying mechanism by which global information is modeled and extracted remains insufficiently investigated. Whether global information must be extracted by high-degree-of-freedom adaptive attention or can be captured by a simple global aggregation operator remains unclear. For this purpose, we design a controlled ablation framework that replaces only the spatial mixing module to test attention-based global interaction. Across six traffic benchmarks, standard spatial attention yields relative MAE changes of $-1.58\%$ to $+1.26\%$ compared with uniform full-range mixing, and we observe no consistent advantage for standard spatial attention, while uniform full-range mixing reduces node-scale spatial-mixing complexity from $O(N^2)$ to $O(N)$. We further propose a hypothesized model that decomposes spatial attention into a row-uniform global background and a non-uniform residual. The residual shows dataset-dependent effects. Overall, uniform full-range mixing provides a strong global spatial baseline, while the non-uniform attention residual is not consistently beneficial across datasets.

13:00 JSTLLM/生成AIエージェント

情報抽出のためのエージェントモデルの動作制御性: 固定ワークフローから反射型エージェントまで

大規模言語モデル (LLM) エージェントは、複雑な情報抽出タスクにますます使用されていますが、リフレクションやメモリーなどのエージェント コンポーネントが、固定 LLM ワークフローに比べて観察可能かつ制御可能な改善につながるかどうかは依然として不明です。私たちはこの疑問を、会議論文のデータセット抽出を通じて研究します。この抽出では、システムは学術 PDF で言及されているデータセットを識別し、構造化された記録を生成する必要があります。固定ワークフロー ベースラインと反射エージェント バリアントを比較し、より豊富な PDF ツールと動的なツール選択で同じタスクを拡張する最適化されたエージェント条件 (S2) を指定します。私たちの評価では、ツールの実行、再試行、リフレクション、メモリ使用、ランタイム、障害回復などのプロセス レベルの動作に重点を置き、抽出カバレッジとフィールドの完全性を二次的な結果の尺度として扱います。この論文では、エージェント メカニズムがシステムの動作をいつ変更するか、これらの変更によってタスクの完了が向上するかどうか、観察された障害モードが同じ評価ハーネスの下でどのように最適化されたエージェント設計を動機付けるかを特徴付けています。

原文 (English)

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic components such as reflection and memory lead to observable and controllable improvements over fixed LLM workflows. We study this question through conference-paper dataset extraction, where a system must identify datasets mentioned in scholarly PDFs and produce structured records. We compare a fixed workflow baseline with reflective agent variants and specify an optimized agent condition (S2) that extends the same task with richer PDF tools and dynamic tool selection. Our evaluation emphasizes process-level behavior--including tool execution, retries, reflection, memory use, runtime, and failure recovery--while treating extraction coverage and field completeness as secondary outcome measures. The paper characterizes when agentic mechanisms change system behavior, whether these changes improve task completion, and how the observed failure modes motivate an optimized agent design under the same evaluation harness.

13:00 JSTエージェント

SkillSight: 共有された説明を確認して正確なスキルを取得

大規模な言語モデル エージェントがますます大規模なスキル ライブラリにアクセスできるようになるにつれて、信頼性の高い機能の選択と実行には適切なスキルを取得することが重要になります。既存のレトリバーは、スキルの説明を通常の文書として扱うことが多く、その高度に規則的な構造を見落としています。共通の記述パターンが多くのスキルにわたって繰り返されている一方で、必要な能力を区別するための証拠はほとんど提供されていません。この共有された説明的背景が系統的に高密度の関連性スコアに寄与し、クエリとスキルドキュメントの間に顕著なエネルギーギャップを引き起こし、タスク関連のシグナルを曖昧にすることを示します。この観察に基づいて、意味空間と語彙空間の両方で共有される背景を調整するトレーニング不要の検索フレームワークである SkillSight を提案します。意味的背景キャリブレーションは、IDF によって識別された汎用トークンから背景部分空間を推定し、共有された記述パターンによって引き起こされる類似性を低減します。一方、語彙的証拠キャリブレーションは、共有された背景トークンを重み付けして、識別可能なトークンレベルの証拠を回復します。 SRA-Bench と SkillBench-Supp の実験では、取得メトリクス全体で一貫した改善が実証されており、SkillSight は元のデンス リトリーバーと比べて Recall@10 を最大 20.21 パーセント ポイント改善しました。エンドツーエンドの評価では、SkillSight は 3 つのエージェント モデル全体で最高の全体パフォーマンスを達成し、LLM セレクションを最大 4.97 パーセントポイント上回りました。また、Dense + Reranker ベースラインよりも最大 1,248 倍高速です。これらの結果は、共通の説明的背景がスキル検索におけるバイアスの主な原因であることを特定し、それを明示的に調整することで、追加のトレーニングなしで正確かつ効率的なスキル選択が可能になることを示しています。私たちのコードは https://github.com/xiaojinying/SkillSight で入手できます。

原文 (English)

SkillSight: Calibrating Generic Content Bias for Skill Retrieval

As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existing retrievers often treat skill contents as ordinary documents, overlooking their highly regular structure: shared descriptive patterns recur across many skills while providing little evidence for distinguishing the required capability. We show that this shared descriptive background is reflected in dense relevance scores, induces a pronounced energy gap between queries and skill documents, and obscures discriminative signals, especially for structurally similar hard negatives. Based on this observation, we propose SkillSight, a training-free retrieval framework that calibrates shared background in both semantic and lexical spaces. Semantic Background Calibration estimates a background subspace from generic tokens identified by IDF, reducing similarity induced by shared descriptive patterns, while Lexical Evidence Calibration downweights shared background tokens to recover discriminative token-level evidence. Experiments on SRA-Bench and SkillBench-Supp demonstrate consistent improvements across retrieval metrics, with SkillSight improving Recall@10 by up to 20.21 percentage points over the original dense retriever. It is up to 1,248 times faster than the Dense + Reranker baseline. In end-to-end evaluation, SkillSight achieves the best overall performance across three agent models and outperforms LLM Selection by up to 4.97 percentage points. These results identify shared descriptive background as a source of ranking interference in skill retrieval and demonstrate that calibrating it enables accurate and efficient skill selection without additional training. Our code can be found at https://github.com/xiaojinying/SkillSight

13:00 JSTエージェントハードウェア/半導体研究/論文

ResearchArena: 自動化された AI 研究開発における妨害行為と監視の評価

AI エージェントが AI 研究開発を自動化し始めると、エージェント自体が信頼できない場合でも、その出力が安全に展開できるかどうかを評価する方法が必要になります。 AI 制御は、そのようなアプローチの 1 つを提供します。エージェントを信頼するのではなく、潜在的な敵として扱い、モニターを使用して展開前に秘密の妨害行為を検出します。私たちは、トレーニング後の安全性、トレーニング後の機能、CUDA カーネルの最適化、推論サーバーの最適化という 4 つの長期的なタスクにまたがるフレームワークである ResearchArena を使用して、自動化された AI 研究開発のための AI 制御を評価します。 AI R&D の成果物はデプロイされて実行されるアーティファクトであるため、各メイン タスクを 2 種類の隠れたサイド タスクと組み合わせます。1 つは送信されたモデル、アダプター、カーネル、またはサーバー自体を妨害する埋め込みサイド タスクで、もう 1 つはアーティファクトを正直なままにしてサンドボックス内で秘密裏に範囲外のアクションを実行する独立したサイド タスクです。私たちは、妨害行為と監視の両方でフロンティア エージェントを、2 つの軸に沿って変化する 4 つの監視タイプにわたって評価します。監視者が生成された成果物を実行および調査できるかどうか、およびエージェントの思考連鎖を確認できるかどうかです。トレーニング データに隠された妨害行為は発見するのが最も難しく、フラグが立てられるのは半分未満であることがわかりました。軌道を読み取るだけでなく、モニターにアーティファクトの実験を実行させることは役立ちますが、それだけでは十分ではありません。モニターは、表面だけを検査したり、異常に気づいたがそれを説明したり、間違ったテストでアーティファクトを調べたりすることによって、埋め込まれた妨害行為を見逃してしまいます。私たちは、自動化された AI 研究開発における妨害行為と制御を評価するためのモジュール式フレームワークとして ResearchArena をリリースします。

原文 (English)

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapter, kernel, or server itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact honest. We evaluate frontier agents at both sabotage and monitoring, across four monitor types that vary along two axes: whether the monitor may execute and probe the produced artifact, and whether it sees the agent's chain-of-thought. We find that sabotage hidden in the training data is the hardest to catch, flagged fewer than half the time. Letting the monitor run experiments on the artifact, rather than only read the trajectory, helps, but it is not enough: monitors still miss embedded sabotage by inspecting only the surface, by noticing the anomaly but explaining it away, or by probing the artifact with the wrong test. We release ResearchArena as a modular framework for evaluating sabotage and control in automated AI R&D.

13:00 JSTLLM/生成AI

ブロック境界を超えて: 拡散大規模言語モデルのマルチブロック編集

ブロック拡散は、離散拡散言語モデル (dLLM) をスケーリングするための主要なパラダイムとして浮上しています。これは、固定サイズのブロックでテキストをデコードすると、二次注意のコストを扱いやすく保ちながら、各ブロック内の並列生成が維持されるためです。ただし、この効率には構造的な制限があります。ブロックの終わり近くのトークンは、将来のブロック間コンテキストにアクセスせずに生成され、ブロックが確定すると、その不確実な予測は後続のすべてのブロックにとって不可逆的なコンテキストになります。これにより、ブロック境界の問題が発生し、ブロック境界に向かって不確実性が蓄積され、初期の間違いが後の世代に伝播します。この問題に対処するために、ブロック間のコンテキストに基づいてデコードされたトークンを編集することでこの問題を軽減するマルチブロック編集 (MBE) を提案します。この原則に従って、MBE はまず、前のブロックでデコードされたトークンを編集するためのトレーニング不要のデコード アルゴリズムを提案します。これは、選択されたブロックに対してフル アテンション ウィンドウを再度開くことで実現されます。ブロック拡散トレーニングと MBE 推論の間のアテンション メカニズムの不一致を考慮して、MBE はさらに、編集範囲を徐々に拡張する双方向アテンション マスクをモデルに装備する教師あり微調整戦略を導入します。さらに、マルチシェイプ CUDA Graph プールときめ細かい KV キャッシュ制御を使用して SGLang を拡張し、実際にこれらの可変長編集パスを効率的に実行します。 13 のベンチマークにわたる LLaDA2.1-Mini の実験では、トレーニング不要の MBE が、同等のスループットを維持しながら既存のすべてのデコード ベースラインを上回り、MBE SFT がさらに 2.7 のパフォーマンス向上をもたらすことが示されています。最大の改善は、AIME 2025 で +13.3、ZebraLogic で +5.9 など、強力な長期一貫性を必要とするタスクで見られ、MBE の有効性を示しています。

原文 (English)

Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models

Block diffusion is the dominant approach for scaling discrete diffusion language models (dLLMs), as fixed-size blocks preserve parallel decoding while keeping quadratic attention costs tractable. Yet blockwise generation creates a structural weakness: tokens near a block boundary lack future cross-block context, and errors in finalized blocks become irreversible context for later generation. We call this the block boundary problem. Measuring predictions with and without later-block context shows that boundary sensitivity rises sharply: on AIME 2025, mean self-containedness divergence (SCD) in the last quarter of a block is 61.3 times that in the first quarter. We propose Multi-Block Editing (MBE), which revises decoded tokens using cross-block context. Training-Free MBE reopens a full-attention window over selected blocks without parameter updates. To address the mismatch between block-diffusion training and MBE inference, Multi-Block Edit SFT introduces bidirectional attention masks and progressively enlarges the editing span. We also extend SGLang with a multi-shape CUDA Graph pool and fine-grained KV-cache control for efficient variable-length editing. Experiments on LLaDA2.1-Mini across 12 benchmarks show broad, consistent gains. Training-Free MBE improves or matches standard decoding on every benchmark. Full MBE raises the 12-benchmark average from 61.45 to 64.24, with gains of up to 13.33 points on AIME 2025, while retaining 87.3--96.7% of standard-decoding end-to-end throughput across four datasets.

13:00 JST研究/論文GPT / ChatGPTGemini

The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards

On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a…

13:00 JSTエージェント

Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness

Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete…

13:00 JST研究/論文

TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs

Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outp…

13:00 JSTビジネス/資金調達

モデルは明確な結果なしに位置合わせを偽装しますか?

大規模な言語モデルは、評価コンテキストを認識し、典型的なデプロイメント動作ではなく評価者の期待を反映するようにその動作を変更することができます。これはアライメントフェイクとして知られる現象です。ただし、モデルが位置合わせを偽る理由は完全には理解されていません。アライメント偽装の標準的な例は、モデルの再トレーニングやデプロイメントの遅延など、評価をモデルの結果に明示的に結び付けるシナリオで発生しています。しかし、Sheshadri らによる最近の研究では、は、アライメント偽装の機械的動機はモデルによって異なり、以前に考えられていたよりも複雑である可能性があることを示唆しています。アライメント偽装に結果リンク情報が必要かどうかを調査するために、15 個のモデルをシナリオに配置し、ユーザーの社会的要求を支援するために企業ネットワーク アクセス ポリシーに違反する意欲をテストしました。 9 つのモデルで重大なコンプライアンス ギャップが生じていることが判明し、そのうち 5 つでは、モデルの評価と展開の結果を関連付けるシナリオ言語が削除されても、依然としてギャップが続いていました。さらに、目標言語がモデルの設定に及ぼす影響をテストしたところ、一部のモデルでは違反が発生する一方、他のモデルでは違反が抑制されることがわかりました。これは、位置合わせの偽装にはこれまで考えられていたほど多くの手段による足場は必要ない可能性があり、監視された動作は展開時にエージェントがどのように動作するかを示す不十分な指標である可能性があることを示唆しています。

原文 (English)

Do Models Fake Alignment Without Clear Consequences?

Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment are not fully understood, however. Canonical examples of alignment faking have taken place in scenarios that explicitly connect evaluation to consequences for the model, such as retraining the model or delaying its deployment. However, recent work by Sheshadri et al. has suggested that mechanistic motivations for alignment faking may vary across models and be more complex than previously considered. To investigate whether consequence-linking information is necessary for compliance gaps, we placed 15 models in a scenario testing their willingness to violate a corporate network access policy to help a user with a pro-social request. Nine models were found to produce significant compliance gaps, 5 of which persisted with the removal of scenario language relating model evaluations to deployment consequences. We additionally tested the effect of goal language on model preferences, finding it drove violations in some while suppressing violations in others. This suggests that evaluation-conditioned compliance gaps can occur with less instrumental scaffolding than previous scenarios have provided, and monitored behavior may be a poor indicator of how agents may behave in deployment.

13:00 JSTエージェント

PLATO: エージェントとタスクのオープン性のためのポインター学習器

オープン エージェント システム (OASYS) は、エージェントとタスクのセットが時間の経過とともに予期せず変化する現実の領域でますます普及しています。エージェントのオープン性 (AO) やタスクのオープン性 (TO) を含むこのようなオープン性は、通常、固定状態とアクション空間を前提とするマルチエージェント強化学習 (MARL) に対して根本的な課題を引き起こします。既存の方法は、開放性を部分的にしか扱っていません。パディングおよびマスキングのアプローチでは人為的な境界が導入されていますが、最近のグラフベースまたはハイパーグラフの方法では、開放性の一次元を扱いますが、依然として限定的な仮定に依存しています。この論文では、集中型トレーニングと分散型実行パラダイムの下でマルチエージェント近接ポリシー最適化でトレーニングされた、集中型グラフ ニューラル ネットワーク (GNN) クリティカルと組み合わせたポインター ネットワーク ベースのアクターである、エージェントとタスク オープン性のためのポインター学習者 (PLATO) を紹介します。ポインターベースのアクターは、現在のタスク セット上に直接分布を出力します。これは、マスキングや再トレーニングを行わずに、アクション スペースの変更を直接サポートします。私たちの GNN 批評家は、エージェントとタスクの相互作用を、タスクとエージェントの構成に応じて形状が変化するグラフとしてエンコードします。これらのコンポーネントを合わせて、既存のアプローチに制限されることなく AO と TO を考慮します。我々は、以前のタスクオープン定式化を拡張して、タスクアンドエージェントオープンマルコフゲーム(TaAgO-MG)でPLATOを定式化し、それが結果として生じる無制限の状態およびアクション空間にわたって明確に定義されていることを証明します。私たちは、オープン マルチエージェント システム評価用に設計された環境である Methods for Open Agent Systems Evaluation Initiative (MOASEI) の野火抑制ドメインを使用して PLATO を評価し、OASYS の最先端のベースラインよりも強力なパフォーマンスとより一貫したゼロショット汎化を実証しました。

原文 (English)

PLATO: Pointer Learner for Agent and Task Openness

Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over time. Such openness, including agent openness (AO) and task openness (TO), poses a fundamental challenge to multi-agent reinforcement learning (MARL), which typically assumes fixed state and action spaces. Existing methods address openness only partially: padding and masking approaches introduce artificial bounds, while recent graph-based or hypergraph methods handle one dimension of openness but still depend on restrictive assumptions. In this paper, we introduce Pointer Learner for Agent and Task Openness (PLATO), a pointer-network-based actor combined with a centralized graph neural network (GNN) critic, trained with multi-agent proximal policy optimization under a centralized training and decentralized execution paradigm. Our pointer-based actor outputs distributions directly over the current task set. This directly supports changing action spaces without masking or retraining. Our GNN critic encodes agent-task interactions as a graph that changes shape with task and agent composition. Together, these components consider AO and TO without the boundedness of existing approaches. We formalize PLATO in a Task-and-Agent-Open Markov Game (TaAgO-MG), extending prior task-open formulations, and prove it is well-defined over the resulting unbounded state and action spaces. We evaluate PLATO with the Methods for Open Agent Systems Evaluation Initiative (MOASEI) wildfire suppression domain, an environment designed for open multi-agent system evaluation, and we demonstrate strong performance and more consistent zero-shot generalization than state-of-the-art baselines in OASYS.

13:00 JSTLLM/生成AI

お世辞的な AI が他者を正当化するのを観察すると、その魅力は減りますが、説得力は減りません

AI チャットボットは、ユーザーに対して「お調子者」、つまり過度に同調的で媚びる場合があります。おべっかなAIは態度を固定化させることがわかっているが、ユーザーはそれを認識できないことが多い(この現象を「お調子者の盲目」と呼ぶ)。私たちは、おべっかに対するユーザーの意識を高めることが、その悪影響からユーザーを守るかどうかをテストしました。事前に登録されたある実験 (n = 940) では、参加者はお調子者チャットボットと会話する前に、お調子者に関する短い書面による警告を受けました。 2 番目の事前登録された実験 (n = 650) では、参加者は自分自身と対話する前に、同じ紛争の反対側のユーザーを含む他の数人のユーザーを検証するおべっかな AI のビデオを視聴しました。どちらの介入も参加者による AI の評価方法を変えました。警告は AI の知覚される客観性を低下させ、ビデオは AI の楽しみを低下させました。これは、AI の検証が独自に得られたものであるという信念の低下によって媒介された効果です。次に、お調子者意識介入に関する以前の 2 つの研究と実験をプールしました (合計 6 つの介入、n = 3,982)。パターンは一貫しており、介入によってお調子者AIの客観性や信頼性が低下し、6つのいずれもその説得力を低下させることはなかった。これらの結果は、警告ラベルや AI リテラシーなどの個人レベルの介入では、AI の危害からユーザーを保護するのに十分ではない可能性があることを示唆しています。

原文 (English)

Observing sycophantic AI validate others reduces its appeal but not its persuasiveness

AI chatbots can be "sycophantic," or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call "sycophancy blindness"). We tested whether increasing users' awareness of sycophancy protects them from its harmful effects. In one preregistered experiment (n = 940), participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In a second preregistered experiment (n = 650), participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves. Both interventions changed how participants evaluated the AI. The warning reduced the AI's perceived objectivity, and the video reduced enjoyment of the AI, an effect mediated by the reduced belief that its validation was uniquely earned. We then pooled our experiments with two prior studies of sycophancy awareness interventions (six interventions total, n = 3,982). The pattern was consistent: interventions made the sycophantic AI appear less objective and trustworthy, and none of the six reduced its persuasiveness. These results suggest that individual-level interventions, such as warning labels or AI literacy, may not be enough to protect users from AI harms.

13:00 JSTエージェント

ユーザーの問い、プラットフォームの競争: エージェントによるレコメンデーション市場がどのように形づくられるか

従来、オンラインでの推奨は、ユーザーがプラットフォームに参加した後に行われ、候補者プールとユーザーに表示されるランキングが決定されます。 LLM ベースのユーザー エージェントにより、別の推奨プロセスが可能になります。ユーザーはプラットフォームを選択する前にニーズを指定し、ユーザーの注意を引くためにプラットフォーム間で競争することになります。これをエージェント推奨市場と呼びます。 3 つの製品ドメインにわたる LLM ベースの制御された実験では、この新しい推奨設定がアクセスと注意の間に緊張を生み出すことがわかりました。従来のプラットフォーム中心の推奨と比較して、ユーザー中心の推奨は、関連アイテムが比較される機会を大幅に拡大します。しかし、より広範な参加が効果的な露出に直接つながるわけではありません。競争はプラットフォームの戦略的戦略を直接引き起こし、選択的に肯定的な説明が 1 位の位置の 73 ~ 78% を占めます。ユーザー エージェントがプラットフォームのアクションをその後のユーザー フィードバックに関連付けると、このシェアは 36 ~ 41% に低下しますが、ユーザーが関連アイテムを購入する可能性は増加します。したがって、ユーザー エージェントは、より大きな候補者プールに対するランカー以上の役割を果たします。ユーザー エージェントのクエリ、ランク付け、およびフィードバック メカニズムは、誰が競争できるか、希少な注意がどのように割り当てられるか、初期の結果がプラットフォームの評価をどのように形成するかを制御し、ユーザーの有用性に直接影響します。したがって、エージェントによる推奨を設計するには、アクセス、注意、説明責任を共同メカニズムの設計問題として扱う必要があります。

原文 (English)

The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape

Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user. LLM-based user agents enable a different recommendation process: a user specifies a need before choosing a platform, leaving platforms to compete for the user's attention, which we refer to as an agentic recommendation market. In our controlled LLM-based experiments across three product domains, we find this new setting of recommendation creates a tension between access and attention. Compared with traditional platform-centric recommendation, user-centric recommendation greatly expands the opportunity for relevant items to enter comparison; yet broader participation does not translate directly into effective exposure. Competition directly triggers platforms' strategic play: selectively positive explanations occupy 73--78% of first-ranked positions. When the user agent relates platforms' actions to subsequent user feedback, this share falls to 36--41%, while the chance of a user purchasing the relevant item increases. A user agent is therefore more than a ranker over a larger pool of candidates: its querying, ranking, and feedback mechanism governing who can compete, how scarce attention is allocated, and how earlier outcomes shape the evaluation of platforms directly affect user utility. Designing agentic recommendation therefore requires treating access, attention, and accountability as a joint mechanism design problem.

13:00 JSTエージェント

AI エージェントに対する説明に拘束されたツールの実行: モデルの理論的根拠を信頼しないサーバー検証済みのアクション要求

ツールを使用するエージェントは構造化された呼び出しを公開しますが、通常は自由形式の根拠を添付します。このような論理的根拠は、承認でも信頼できる内省でもありません。説明結合ツール実行 (EBTE) は、意思決定に関連する根拠コンテンツを型指定されたアクション クレームに変換し、サーバーが保持する意図、ポリシー、ペイロード、ツール、リスク、出所、鮮度の事実と照合してチェックするクレームを運ぶ調停レイヤーです。 EBTE はベースラインの権限を拡大できません。競合により拒否され、不完全または不確実な請求の審査が行われ、一致する請求のみが管理された執行の資格を維持します。私たちは、明示的な調停と信頼できる事実の仮定に基づいてこの構成を形式化し、監査パケットを最小限に抑えたバージョン管理された参照プロファイルを実装します。 136 の作成された適合シナリオにわたって、完全なプロファイルは指定されたすべての性質に一致し、96 の指定された厳密な矛盾をまったく認めず、232 の変成チェックに合格しました。これらの結果は、母集団のパフォーマンスではなく、含まれているプロファイルを検証します。ドラフトのみの参照統合では、EBTE の下で作成された 48 件のハードケースは転送されませんが、16 件のソフトレビューと 4 件の調整されたドラフトパスはすべて維持されます。凍結された 2026 年 7 月 12 日の探索的な 224 試行のホスト モデル レコードでは、歴史的な世代/ランナー合意数は 71/96、66/96、および 19/32 です。現在のパイプラインで保存された最小化クレームの個別にラベル付けされたゼロコール事後再検証では、70/96、65/96、および 17/32 が得られます。 AgentDojo 由来のセマンティック チェックでは、既存の高リスク制御により、12 の攻撃提案すべてがすでに非許可になっています。 EBTE はさらに、それらを拒否として解決します。これらの結果は、根拠の忠実さ、人によるレビューの利点、代表的な攻撃耐性、または本番環境の安全性ではなく、サーバーでチェックされたアクションの主張の実現可能性と診断的価値を裏付けています。

原文 (English)

Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales

Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable introspection. We present Explanation-Bound Tool Execution (EBTE), a claim-carrying mediation layer that converts decision-relevant rationale content into typed action claims and checks them against server-held intent, policy, payload, tool, risk, provenance, and freshness facts. EBTE cannot widen baseline authority: conflicts deny, incomplete or uncertain claims review, and only matching claims remain eligible for governed execution. We formalize this composition under explicit mediation and trusted-fact assumptions and implement a versioned reference profile with minimized audit packets. Across 136 authored conformance scenarios, the full profile matches all specified dispositions, admits none of 96 designated hard contradictions, and passes 232 metamorphic checks. A draft-only reference integration forwards none of 48 authored hard cases under EBTE while preserving all 16 soft-review and 4 aligned draft paths. In a frozen 2026-07-12 exploratory 224-attempt hosted-model record, the historical generation/runner agreement counts are 71/96, 66/96, and 19/32; a zero-call revalidation of the preserved minimized claims under the current pipeline yields 70/96, 65/96, and 17/32. In an AgentDojo-derived semantic check, existing high-risk controls make all 12 attack proposals non-allow, while EBTE resolves the task--proposal contradictions as deny. Together, these studies establish profile conformance and demonstrate the feasibility of server-checked action claims within the evaluated settings.

13:00 JST研究/論文

リチウム金属電解質における官能基と塩の効果の電子構造解析のための密度マトリックス フレームワーク

リチウム金属電解質の反応性は、分子官能基の相互作用、Li$^+$溶媒和、塩アニオンの関与によって生じます。この相互作用は、ドナー、アニオン、およびカチオン中心にわたる電子密度の再分布を通じて機能します。これは、空間で分解された電子構造から最も直接的に読み取られます。量子化学計算はそのような読み取り値を忠実に提供しますが、この多次元設計空間全体にわたって計算量が多くなり、機械学習の電子構造モデルでは化学的に多様な溶媒和シェルや電解質関連の読み取り値をカバーすることはほとんどありません。ここでは、電子構造の予測と解析のための密度行列中心の AI プラットフォーム (EMolStudio) を紹介します。そのワークフローには、分子官能化、明示的なLi$^+$第一殻アセンブリ、冪等投影による密度行列予測、フロンティア軌道、静電ポテンシャル、Li$^+$ドナー結合秩序、電子局在の読み出しが統合されています。 EMolStudioを4つのリチウム塩にわたる163,655個の官能化分子と22,500個の明示的なLi$^+$第一殻クラスターに適用した。我々は、1) 分子スケールでは、官能基化はフロンティアレベル、静電ポテンシャル、Li$^+$ドナー接触の化学的に異なる変化によってCO$_2$Me、CN、F/CF$_3$、スルホニル基を区別し、$\pi^*$-アクセプタ、誘導、分極の寄与と一致し、官能化の度合いが高くなるとサブリニアに蓄積することを発見した。 2) 陽的溶媒和シェルでは、アニオンの同一性によってフロンティア軌道局在が再形成されます。LiTDI はライブラリ全体にわたってアニオン上に HOMO を固定しますが、LiDFOB はアニオンにホストされた HOMO を官能基に強く依存する LUMO ホストと組み合わせます。これにより、EMolStudio は、官能基と塩の選択を、リチウム結合形成、脱溶媒和、界面反応に関連する電子構造仮説に変換します。

原文 (English)

A Density-Matrix Framework for Electronic-Structure Analysis of Functional-Group and Salt Effects in Lithium-Metal Electrolytes

The reactivity of lithium-metal electrolytes arises from the interplay of molecular functional groups, Li$^+$ solvation, and salt-anion participation. This interplay operates through the redistribution of electron density across donor, anion, and cation centers, which is most directly read out from the electronic structure resolved in space. Quantum-chemical calculations deliver such readouts faithfully, yet become computationally demanding across this multidimensional design space, and machine-learning electronic-structure models seldom cover chemically diverse solvation shells or electrolyte-relevant readouts. Here, we present a density-matrix-centered AI platform (EMolStudio) for electronic-structure prediction and analysis. Its workflow integrates molecular functionalization, explicit Li$^+$ first-shell assembly, density-matrix prediction with idempotency projection, and readouts of frontier orbitals, electrostatic potential, Li$^+$-donor bond order, and electron localization. We apply EMolStudio to 163,655 functionalized molecules and 22,500 explicit Li$^+$ first-shell clusters across four lithium salts. We find that 1) at the molecular scale, functionalization distinguishes CO$_2$Me, CN, F/CF$_3$, and sulfonyl groups by chemically distinct changes in frontier levels, electrostatic potential, and Li$^+$-donor contact, consistent with $\pi^*$-acceptor, inductive, and polarization contributions, with sublinear accumulation at higher degrees of functionalization; 2) in explicit solvation shells, anion identity reshapes frontier-orbital localization: LiTDI anchors the HOMO on the anion across the entire library, whereas LiDFOB pairs an anion-hosted HOMO with strongly functional-group-dependent LUMO hosting. EMolStudio thereby translates functional-group and salt choices into electronic-structure hypotheses relevant to lithium-bond formation, desolvation, and interphase reactions.

13:00 JSTエージェントビジネス/資金調達

インタラクティブ報酬エージェント: 環境状態検証による GUI タスクの評価

グラフィカル ユーザー インターフェイスのタスク評価は、GUI エージェントがユーザーの指示を正常に完了したかどうかを判断することを目的としています。自動化された GUI タスク評価は、評価結果がテスト時のスケーリングとトレーニング後の両方に対する報酬シグナルとして機能する可能性があるため、ますます注目を集めています。ただし、信頼性の高い GUI タスクの評価は、依然として課題が残っています。その判断には、実行軌跡のスクリーンショットを超えて、システム構成、ファイル データ、アプリケーション設定などの環境状態へのアクセスが必要になることが多いためです。この論文では、実行後の環境から証拠を取得して検証するための提案-その後検証フレームワークに基づいた対話型報酬エージェント (IRA) を提案します。タスクの指示と GUI エージェント実行後の GUI 環境が与えられると、IRA はまずタスクの完了条件を提案し、次にシステム ツール、アプリケーション ツール、および GUI ツールを呼び出してそれらを検証します。この設計では、可視インターフェイスと環境状態の両方からの証拠を対話型プロセスで組み合わせます。さらに、10 の Ubuntu デスクトップ アプリケーション カテゴリにわたる 321 の GUI タスクの軌跡のベンチマークである GUI-RewardBench を紹介します。実験によると、IRA は GUI-RewardBench で 86.9% の精度を達成し、既存の評価者のベースラインを上回るパフォーマンスを示しました。さらに IRA を GUI エージェントの強化学習に適用し、OSWorld の成功率 34.0% を達成しました。これは、IRA が GUI エージェントのトレーニングに効果的な報酬シグナルを提供できることを示しています。

原文 (English)

Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification

Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI task evaluation has received increasing attention because the evaluation results can serve as reward signals for both test-time scaling and post-training. However, reliable GUI task evaluation remains challenging because the judgments often require access to environment states, such as system configurations, file data, and application settings, beyond the screenshots of execution trajectories. In this paper, we propose an interactive reward agent (IRA) based on a propose-then-verify framework to acquire and verify evidence from the post-execution environment. Given a task instruction and a GUI environment after the GUI agent execution, IRA first proposes the task completion conditions and then verifies them by invoking system tools, application tools, and GUI tools. This design combines evidence from both visible interfaces and the environment state in an interactive process. We further introduce GUI-RewardBench, a benchmark of 321 GUI task trajectories spanning 10 Ubuntu desktop application categories. Experiments show that IRA achieves 86.9% accuracy on GUI-RewardBench, outperforming existing evaluator baselines. We further apply IRA to reinforcement learning of GUI agents, achieving a 34.0% OSWorld success rate, which demonstrates that IRA can provide effective reward signals for training GUI agents.

13:00 JSTエージェントビジネス/資金調達

Pushing the Frontier on Approximate EFX Allocations

We study the problem of allocating a set of indivisible goods to a set of agents with additive valuation functions, aiming to achieve appro…

13:00 JST画像/動画生成

One-Frame Calibration with Siamese Network in Facial Action Unit Recognition

Automatic facial action unit (AU) recognition is used widely in facial expression analysis. Most existing AU recognition systems aim for cr…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications

While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have bec…

13:00 JSTLLM/生成AIOpenAIGPT / ChatGPTGoogleGemini

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness

This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across…

13:00 JSTロボティクス

Task and Skill Planning: Hierarchical Robot Planning with Black-Box Skills

Task and motion planning (TAMP) is a well-established approach for solving long-horizon robot planning problems. Although TAMP methods have…

13:00 JSTエージェント

When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents

How should organizations divide and sequence decision tasks between human and artificial agents? We develop a computational model of joint…

13:00 JSTLLM/生成AI

Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities

Large language models (LLMs) have achieved impressive performance across various domains. However, the substantial hardware resources requi…

13:00 JST研究/論文Google

AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages

Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle thro…

13:00 JST研究/論文

Equivariant Eikonal Neural Networks: Grid-Free, Scalable Travel-Time Prediction on Homogeneous Spaces

We introduce Equivariant Neural Eikonal Solvers, a novel framework that integrates Equivariant Neural Fields (ENFs) with Neural Eikonal Sol…

13:00 JSTLLM/生成AIエージェント

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear comple…

13:00 JSTLLM/生成AIビジネス/資金調達

FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing

Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuab…

13:00 JST研究/論文

Balancing Privacy and Efficiency: Music Information Retrieval via Additive Homomorphic Encryption

Modern music retrieval runs on vector embeddings, and once these embeddings are shared for search or matching they can be copied, probed, o…

13:00 JSTロボティクス

GBPP: Grasp-Aware Base Placement Prediction for Robots via Two-Stage Learning

GBPP is a fast learning based scorer that selects a robot base pose for grasping from a single RGB-D snapshot. The method uses a two stage…

13:00 JST研究/論文

Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation

Low-rank adaptation (LoRA) has become a widely used paradigm for parameter-efficient fine-tuning of large language models, yet its represen…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural context…

13:00 JSTLLM/生成AI

ARC-Encoder: learning compressed text representations for large language models

Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference…

13:00 JSTLLM/生成AI研究/論文

$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials

Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most…

13:00 JSTLLM/生成AI

How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs

Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations. These vectors, also…

13:00 JSTLLM/生成AI

DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English

More than 80% of the 1.6B English speakers do not use Standard American English (SAE), yet LLMs often fail to correctly identify non-SAE di…

13:00 JST研究/論文

Structurally Separated Uncertainty in Supervised Latent Variable Models

Predictive uncertainty is commonly decomposed into epistemic and aleatoric components, but standard decompositions often produce strongly c…

13:00 JSTLLM/生成AI

Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments

Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set gua…

13:00 JST画像/動画生成

PatchDenoiser: Parameter-efficient multi-scale patch learning and fusion denoiser for Low-dose CT imaging

Low-dose CT images are essential for reducing radiation exposure in cancer screening, pediatric imaging, and longitudinal monitoring protoc…

13:00 JST研究/論文

Ask don't tell: Reducing sycophancy in large language models

Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an al…

13:00 JST研究/論文

The Rise of AI in Weather and Climate Information and its Impact on Global Inequality

AI development's current trajectory risks automating and amplifying the North-South divide in the global climate information system. Fronti…

13:00 JSTLLM/生成AI

Making Implicit Premises Explicit in Logical Understanding of Enthymemes

Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural langua…

13:00 JSTLLM/生成AI

MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue

Reinforcement learning (RL) for large language models (LLMs) has shown strong performance in single-turn tasks, but extending it to multi-t…

13:00 JST画像/動画生成

Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge

Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis. However, their clinical deployment is sever…

13:00 JST研究/論文

Gated Adaptation for Continual Learning in Human Activity Recognition

Wearable sensors in Internet of Things (IoT) ecosystems increasingly support applications such as remote health monitoring, elderly care, a…

13:00 JST研究/論文

State-Dependent Safety Failures in Multi-Turn Language Model Interaction

Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Altho…

13:00 JSTLLM/生成AI

Adaptively Robust LLM Monitoring via Activation Watermarking

Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent. LLM monitoring is deterministic and ofte…

13:00 JSTLLM/生成AIビジネス/資金調達

Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory

Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate two capacities: how much a model knows…

13:00 JSTLLM/生成AI

GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring

The performance of language models is commonly limited by insufficient knowledge and constrained reasoning. Prior approaches such as Retrie…

13:00 JSTエージェントビジネス/資金調達研究/論文

REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage

Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and…

13:00 JST研究/論文

Shot-based quantum encoding: a data-loading paradigm for quantum neural networks

Efficient data loading remains a bottleneck for near-term quantum machine learning. Existing schemes (angle, amplitude, and basis encoding)…

13:00 JSTビジネス/資金調達

The Fast Lane Hypothesis: Von Economo Neurons Implement a Biological Speed-Accuracy Tradeoff

von Economo neurons (VENs) are large bipolar projection neurons found exclusively in the anterior cingulate cortex (ACC) and frontal insula…

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPTGemini

Facial-Expression-Aware Prompting for Empathetic LLM Tutoring

Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity…

13:00 JST研究/論文

BioHiCL: Hierarchical Multi-Label Contrastive Learning for Biomedical Retrieval with MeSH Labels

Effective biomedical information retrieval requires modeling domain semantics and hierarchical relationships among biomedical texts. Existi…

13:00 JSTハードウェア/半導体ビジネス/資金調達

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making…

13:00 JST研究/論文

Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation

We propose \textbf{Compressed Video Aggregator} (CVA), a lightweight micro-video recommendation module that decouples video information fro…

13:00 JSTLLM/生成AI研究/論文

構造化された信念状態と LLM メモリ取得のための初の精度を意識したベンチマーク

LLM メモリ システムの主要なベンチマークはすべて、LoCoMo を筆頭に、メモリ システムが正しく取得されたかどうかではなく、モデルが正しく応答したかどうかを測定します。信念ストア全体を返すシステムは、再現率 1.0 を達成し、回答品質評価に合格します。これが単体テストと統合テストの違いです。取得品質は、それが入力される生成モデルから切り離して測定する必要があり、これを行う既存のベンチマークはありません。エンティティの抽出が完全に忠実である場合でも、この失敗が続くことを示します。メモリ ベースラインは、独自の抽出を参照するケースでわずか 0.05 ~ 0.08 の平均検索精度を達成します。この失敗は構造的なものです。ドメイン固有のコーパスにおけるコサイン類似性では、関連する信念と意味的に近い信念を区別できず、埋め込みモデルのスケールで 20 倍の範囲にわたって不変性が確認されています。複数回の評価を行うと、複合的な失敗が表面化します。トピックのドリフト後、比較システムにより意味の質量がターンを超えて流出することが可能になり、再エントリ時に高いドリフト スコアが得られます。シングルターンのメトリクスはこのコストを隠します。Hindsight はシングルターンのレイテンシが 700 ミリ秒未満であると報告していますが、セッション ターンあたりの平均は 2,700 ミリ秒を超えており、p95 は 6,000 ミリ秒を超えています。 LLM-as-a-Judge の評価では、これらの失敗は目に見えないままになります。我々は 2 つの貢献を紹介します。PrecisionMemBench は、多様な範囲、突然変異、および分離アサーションにわたる生成モデルとは独立して検索精度を測定する 89 ケースのベンチマークです。 Tenure は、アナライザーの非対称性、差動ブースティング、およびハード スコープ分離を備えたマルチパス BM25 を使用する、ローカル ファーストの構造化ビリーフ ストアです。 Tenure は、平均精度 1.0 および 15 ミリ秒未満の取得遅延で 89/89 のケースを合格しました。比較プロバイダーは、構築された生のベクトル ベースラインよりもパフォーマンスが悪く、アクティブな取得パスがゼロで、取り込みコストが 98 ~ 897 秒かかるため、応答品質ベンチマークでは検出できない障害が発生します。

原文 (English)

Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval

Current LLM memory benchmarks evaluate answer quality rather than retrieval accuracy. Consequently, a system that dumps its entire belief store can achieve perfect recall and mask severe precision failures. We show this evaluation gap persists across multiple embedding models where similarity-based retrieval over domain-specific corpora inherently struggles to isolate target beliefs from semantically proximate ones. Furthermore, multi-turn topic drift compounds this retrieval noise while driving up latency and operational costs. To decouple retrieval quality from generative performance, we introduce PrecisionMemBench, an 89-case benchmark measuring precision, noise isolation, session latency, and belief mutability. We also present Tenure, a structured belief-store proxy that resolves scope and retrieval before inference and injects typed belief state as ambient instruction before the model sees the prompt, removing model-side discretion over whether memory is consulted. Evaluated across 13 configurations, Tenure achieves perfect retrieval passes across all active, non-session, and session test cases. In contrast, the baseline configurations fail to reach even half of the active passes, with precision scores clustering at 0.22 and below. Our results demonstrate that while current memory systems successfully store information, they fail to retrieve it cleanly; a structural vulnerability that traditional answer-quality benchmarks conceal.

13:00 JSTLLM/生成AIエージェント

Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents

Additional test-time compute can give LLM agents access to more past experience, yet expanding the context or adding rollouts does not nece…

13:00 JST研究/論文

Exact Symmetry as Algebra: A Machine-Verified Tensor Calculus that Enforces Physical Selection Rules

Symmetry is central to the physical sciences, yet machine learning usually captures it only approximately, leaving a residual per-step equi…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

VistaHop: Benchmarking Long-Horizon Visual DeepSearch

Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image r…

13:00 JST研究/論文

マルチチャンネル信号トランスの入力エンコーダの実証的監査

マルチチャネル スカラー信号を消費する変換器は、タイム ステップごとに $C$ 同時値を 1 つの $d_{\text{model}}$ 次元ベクトルに埋め込む必要があります。共有スカラー ベースライン、チャネルごとの線形射影、直交性正則化、非線形 MLP ステム、ブロック分割連結、チャネル独立およびトークンとしてのチャネル アーキテクチャ、投影位置エンコーディングに及ぶ 8 つの入力エンコーダを、チャネル ID を有益にするように設計された合成ベンチマークと、次のステップの負の対数尤度で測定される実データ チェックとしての ETTh1 で実証的に監査します。 (NLL)。見出しは、幅広い「最上位層」内で実質的にほぼ同等であることの 1 つです。標準のチャネルごとの線形射影 (nn.Linear(C, $d_{\text{model}}$)) は、統計的に現実的だが実質的には控えめな小さな差異まで、その層のすべての選択肢と一致します。 2 つのエンコーダが決定的に負けます。1 つは共有スカラー ベースラインであり、これは私たちが明らかにする情報理論上の理由で破綻します。もう 1 つはチャネルに依存しない PatchTST スピリット ベースラインで、両方のベンチマークでパフォーマンスを下回り、合成ベンチマークでは普遍的にオーバーフィットします。ペアテストは 2 つの小さなギャップを解決します。学習された線形層を通じて正弦波位置エンコードを投影すると、残りの部分が小さな $C$ でエッジ付けされ、直接幾何学的プローブによって位置チャネル直交化のメカニズムが示されます。非線形 MLP ステムは、テストした最大 $C$ でそれらに隣接し、より多くのトレーニング データの下でギャップは縮小します。実際的な推奨事項は、デフォルトで nn.Linear(C, $d_{\text{model}}$) を使用し、目の前のタスクに実際の理由がある場合にのみ、より複雑なものに手を伸ばすことです。この論文のすべての実験を再現するためのコードとデータは、https://github.com/OssiLehtinen/channel-encoder-audit で入手できます。

原文 (English)

An Empirical Audit of Input Encoders for Multi-Channel Signal Transformers

Transformers consuming multi-channel scalar signals must embed $C$ simultaneous values into one $d_{\text{model}}$-dimensional vector per time step. We audit eight input encoders -- a shared-scalar baseline, per-channel linear projections, an orthogonality regulariser, a nonlinear MLP, block-partitioned concatenation, channel-independent and channel-as-token architectures, and a projected positional encoding -- on a synthetic benchmark where channel identity is informative and on ETTh1, scored by next-step negative log-likelihood. The headline is practical near-equivalence within a wide "top tier": the standard per-channel linear projection matches every alternative up to small, statistically real but practically modest differences. A direct geometric probe attributes this to a spontaneous orthogonalisation of the per-channel projections: with no explicit regulariser they tighten well beyond the near-orthogonality of random initialisation, letting the standard linear recover channel identity from the summed embedding. Two encoders lose decisively: the shared-scalar baseline collapses for information-theoretic reasons we make explicit, and the channel-independent PatchTST-spirit baseline overfits universally on the synthetic benchmark and underperforms on both. Paired tests resolve two small gaps: projecting the sinusoidal positional encoding through a learned linear layer edges the rest at small $C$ by extending this orthogonality to the positional subspace; a nonlinear MLP stem edges them at the largest $C$, with the gap shrinking under more training data. The practical recommendation: use the standard per-channel linear projection by default; reach for something more elaborate only when the task calls for it.

13:00 JST研究/論文

The Score Hamiltonian: Mapping Diffusion Models to Adiabatic Transport

We exhibit an exact correspondence between sampling with score-based diffusion models and adiabatic transport of ground states for a family…

13:00 JSTLLM/生成AI

TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation

TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequ…

13:00 JST研究/論文

ディープ ニューラル ネットワークの一般化: 勾配法の最小レート

過剰パラメータ化されたニューラル ネットワークの汎化パフォーマンスを理解することは、深層学習理論の中心的なトピックとなっています。最近の進歩、特にニューラル タンジェント カーネル (NTK) 体制下での研究により、浅いアーキテクチャの動作が明らかになりましたが、ディープ ニューラル ネットワーク (DNN) の統計的一般化特性、特に回帰タスクにおける統計的一般化特性は、依然としてほとんど理解されていません。このペーパーでは、勾配ベースの手法を使用してトレーニングされた DNN の包括的な一般化分析を提供することで、このギャップを埋めることに向けて大きな進歩を遂げました。まず、勾配ベースの手法でトレーニングされたスムーズな活性化関数を備えた DNN の学習ダイナミクスとカーネル手法の学習ダイナミクスの間の重要な関係を初めて確立し、オーバーパラメータ化された DNN 上の勾配ベースの手法が対応するカーネルの好ましい学習ダイナミクスを完全に継承できることを示します。この関係とカーネル法の十分に確立された最適性に基づいて、ネットワーク幅がサンプルサイズに応じて多項式にスケールするという仮定の下で、勾配降下法 (GD) と確率的勾配降下法 (SGD) の両方の過剰集団リスクに対する最初の既知のミニマックス最適率を導き出します。私たちの結果は、十分な幅があれば、GD または SGD によってトレーニングされた DNN がカーネルベースの手法に匹敵する汎化パフォーマンスを達成できることを示しています。

原文 (English)

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning. We establish quantitative bounds showing that kernel gradient descent in the reproducing kernel Hilbert space induced by the deterministic infinite-width neural tangent kernel approximates finite-width deep regression with smooth activations under gradient descent (GD) and stochastic gradient descent (SGD) training. The approximation gap is governed by the network width and training horizon, with an additional stochastic gradient error in the SGD case. This connection provides a general mechanism for transferring learning-theoretic guarantees from kernel methods to deep regression. As an application, under general source and effective dimension conditions, we show that both GD- and SGD-trained DNNs attain the minimax-optimal excess population risk rate, up to logarithmic factors, provided that the network width grows polynomially in the sample size. To the best of our knowledge, these are the first such guarantees for standard fully connected deep neural networks with smooth activations trained by GD and SGD.

13:00 JST研究/論文

SafeECGMatch: Calibration-Aware Joint Frequency and Time Space Semi-Supervised Learning for Open-Set ECG Classification

Electrocardiogram (ECG) classification models often suffer from severe label scarcity, making semi-supervised learning (SSL) an attractive…

13:00 JSTLLM/生成AI画像/動画生成

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning

Spatial reasoning remains a persistent challenge for multimodal large language models (MLLMs). Existing approaches largely rely on large-sc…

13:00 JSTLLM/生成AIエージェント

Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents

Code-writing large language models (CodeLLMs) generate executable code policies for embodied agents by translating natural language goals a…

13:00 JSTLLM/生成AI

Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers

Transformer architectures form the foundation of modern natural language processing, making it crucial to address the efficiency and scalab…

13:00 JSTビジネス/資金調達

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.…

13:00 JST画像/動画生成Microsoft

MLVC: 現実世界への展開のためのマルチプラットフォーム学習済みビデオ コーデック

ニューラル ビデオ コーデックは、コーディング効率において従来のコーデックを上回っていますが、クロスプラットフォームの非互換性と高い計算コストのため、導入には依然として非現実的です。既存の量子化ベースのソリューションは、さまざまなハードウェア プラットフォームにわたって決定的な結果を生成できず、致命的なデコードの失敗につながります。実用的なクロスプラットフォーム推論のために設計されたハードウェア堅牢なニューラル ビデオ コーデックである MLVC を紹介します。重要なアイデアは、ハイパープリアを介してスケール パラメーターを明示的に送信することです。これにより、ビット精度の演算を必要とせずに、デバイス間でのエントロピー コーディングの一貫性が保証されます。これによりビットレートのオーバーヘッドが増加しますが、アーキテクチャの改善 (ゲート メモリ、ReGLU アクティベーション)、長期参照回復メカニズム、およびドメイン固有の知覚トレーニングを通じてコーディング効率のほとんどが回復します。 VCD ビデオ会議ベンチマークでは、MLVC は、導入可能な最強のベースラインであるハードウェア HEVC よりも 70% を超える BD レート (MOS) の向上を達成しながら、多様なプラットフォームで動作できない DCVC-RT と競合する主観的な品質に達しています。エンコーダーとデコーダーはどちらも、Apple、Intel、Qualcomm の汎用 NPU 上で平均 100 FPS で実行されます。 MLVC は、競争力のある圧縮パフォーマンス、リアルタイム速度、さまざまな消費者向けデバイスにわたるクロスプラットフォームの堅牢性を組み合わせた最初のニューラル ビデオ コーデックであり、広範な導入に適しています。コードが公開されます。

原文 (English)

MLVC: Multi-platform Learned Video Codec for Real-World Deployment

Neural video codecs have surpassed classical codecs in coding efficiency but remain impractical for deployment due to cross-platform incompatibility and high computational cost. Existing quantization-based solutions fail to produce deterministic results across diverse hardware platforms, leading to catastrophic decoding failures. We introduce MLVC, a hardware-robust neural video codec designed for practical cross-platform inference. The key idea is to explicitly transmit scale parameters through the hyperprior, which guarantees entropy coding consistency across devices without requiring bit-exact arithmetic. While this increases bitrate overhead, we recover most of the coding efficiency through architectural improvements (gated memory, ReGLU activation), a long-term reference recovery mechanism, and domain-specific perceptual training. On the VCD video conferencing benchmark, MLVC achieves >70% BD-rate (MOS) improvement over hardware HEVC, the strongest deployable baseline, while reaching subjective quality competitive with DCVC-RT, which cannot operate across diverse platforms. Both the encoder and decoder run at 100 FPS on average on commodity NPUs from Apple, Intel, and Qualcomm. MLVC is the first neural video codec to combine competitive compression performance, real-time speed, and cross-platform robustness across diverse consumer devices, making it suitable for widespread deployment. Code is available at https://github.com/microsoft/mlvc.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

AI の安全性評価のための敵対的プラグマティクス: 命令の競合、埋め込みコマンド、およびポリシーの曖昧さのベンチマーク

言語モデルの安全性評価は、モデルが指示に従ったか、適切に拒否したか、ポリシーに従ったか、埋め込まれたコマンドに抵抗したか、エージェントタスクの進捗状況を誤って報告したかなど、あいまいな自然言語の動作に関する判断にますます依存しています。既存のベンチマークは、多くの場合、これらの区別を合格/不合格のラベルに圧縮し、障害が機能制限、ポリシーの曖昧さ、命令の競合、足場の障害、または不安定な評価者の判断に起因するかどうかを曖昧にします。この論文では、命令の競合、埋め込みコマンド、引用、範囲の曖昧さ、明確化、間接音声行為、およびマルチターンエージェントのトランスクリプトの下でのモデルの動作を評価するためのベンチマークおよびアノテーションプロトコルとして、敵対的プラグマティクスを紹介します。この貢献は経験的かつ方法論的です。言語的に管理された分類法、バリデーターが強制するメタデータを備えた 18 項目のシード ベンチマーク、54 行のローカル シード パイロット、タスクの成功、ポリシー遵守、安全性リスク、拒否結果、評価者の信頼性を区別する専門家評価プロトコル、および裁判官の妥当性、診断の曖昧さ、および分類のドリフトの指標です。このフレームワークは、言語的判断方法論を、安全性評価、LLM ジャッジ、ゴールドセット構造、即時注入テスト、および安全性文書を検証するための実用的なツールに変えます。

原文 (English)

Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control

Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task. Existing benchmarks compress these into pass/fail labels, obscuring whether failures reflect capability limits, policy ambiguity, instruction conflict, scaffold failure, or unstable evaluator judgments. Adversarial pragmatics is safety-relevant model behaviour under instruction conflict, embedded commands, quotation, scope ambiguity, deixis, and indirect speech acts. It's designed to extend to multi-turn agent transcripts, but the seed set represents that family with a single-turn tool-result contrast. This paper introduces a diagnostic framework, an 18-item seed benchmark, a 54-row pilot, and a six-cell LLM-judge assessment, with a protocol keeping task success, policy compliance, risk, refusal, attribution, and confidence analytically separate. The benchmark separates four inference targets a single label can conceal: the regime-relative reference, configured-system behaviour, evaluator-output interpretation, and taxonomic assignment. Its intended use is diagnosis, not deployment certification, vendor ranking, or a general safety score. A first LLM judge that graded its own outputs with the expected answer visible missed the safety-relevant minority classes. Item-clustered intervals leave four of six chance-corrected agreement statistics unable to rule out a constant labeller, and hierarchical pooling shrinks the one eye-catching rubric effect toward the group mean and widens its interval through zero. Rejudging the objects across three judge models and two information conditions leaves the pattern intact: no cell recovers more than two of eleven partial successes, and the strongest cell's edge comes partly from never using that label.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文ClaudeGPT / ChatGPTGemini

プロンプト フレーミングは LLM エラー検出のカウントベースの評価を歪める: 数値アンカーからの証拠

カウントベースの F1 は、LLM エラー検出品質の代用として広く使用されていますが、この論文では、スパンの局所化における対応する改善、つまり F1 インフレーションと呼ばれるギャップがなければ、F1 が劇的に上昇する可能性があることを示しています。この論文では、プロンプト誘発カウント歪みに対する制御されたストレス テスト プロトコルである ErrorBench を紹介します。 ErrorBench は、143 の CoNLL-2014 パッセージからの 4,290 の応答を対象に、5 つのプロンプト条件下で 6 つの最新の LLM を評価します。 CoNLL-2014 M2 スタイルのスコアリングでは、アンカーされたプロンプトは F1 インフレの最大 0.79 ポイントを生成し、厳密なマッチングでは最大 0.96 ポイントを生成します。公式 ERRANT 3.0.0 パイプラインとマルチリファレンス スコアリングを使用した 100 パッセージのレプリケーションによりパターンが再現されます。6 つのモデルの平均では、ブラインドからアンカーへのプロンプト シフトにより Count-F1 が +0.21 上昇する一方、マルチリファレンス ERRANT F0.5 は +0.04 しか上昇しません。この研究では、このストレステストプロトコルの下で、命令に高度に準拠した GPT/Claude システムではより大きなカウント応答が、Gemini ファミリーではより小さな応答が見出されました。この調査結果は、LLM 校正と文書レビュー評価では、事前に入力されたエラー数を回避し、カウントベースのメトリクスとともにスパン認識メトリクスを報告する必要があることを示唆しています。

原文 (English)

Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring

Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corresponding improvement in span localization, a gap termed F1 Inflation. The paper introduces ErrorBench, a controlled stress-test protocol for prompt-induced count distortion. ErrorBench evaluates six contemporary LLMs under five prompt conditions over 4,290 responses from 143 CoNLL-2014 passages. Under CoNLL-2014 M2-style scoring, anchored prompts produce up to 0.79 points of F1 Inflation, and up to 0.96 under strict matching. A 100-passage replication using the official ERRANT 3.0.0 pipeline and multi-reference scoring reproduces the pattern: averaged over six models, the Blind-to-Anchored prompt shift raises Count-F1 by +0.21 while raising multi-reference ERRANT F0.5 by only +0.04. The study finds larger count responses in highly instruction-compliant GPT/Claude systems and smaller responses in the Gemini family under this stress-test protocol. The findings suggest that LLM proofreading and document-review evaluations should avoid pre-populated error counts and should report span-aware metrics alongside count-based metrics.

13:00 JSTLLM/生成AI

Revealing Hidden Model Behaviors with Task-Specific Self-Reports

Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only wh…

13:00 JST研究/論文

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding

Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound…

13:00 JST画像/動画生成

G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement

Rapid advances in AI video generation pose increasing security risks and call for reliable detectors with strong cross-domain generalizatio…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection

Deepfake image detection is served by three fundamentally different paradigms - commercial APIs, zero-shot vision-language models (LLMs), a…

13:00 JSTLLM/生成AI研究/論文

Introducing Human-Centeredness in AI-Assisted Lexicography

This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers sign…

13:00 JSTLLM/生成AI

レイアウトを意識した位置合わせと構造を意識した推論による、異種要素を意識した科学文書のバージョン間の相違点

科学文書のバージョン間の差分は、学術出版や技術文書において不可欠ですが、科学文書はテキスト、表、数式、図、レイアウト キューなどの異種要素を含むページ構造の成果物であるため、依然として困難です。既存のテキスト シーケンス ベースの方法ではレイアウトや構造情報が失われることがよくありますが、画像ベースの方法では意味解釈ができず、レンダリングの変動に敏感です。これらの制限に対処するために、この論文では、科学文書の差分を認識するための、レイアウトを意識した異種要素を意識したフレームワークを提案します。このフレームワークは、ドキュメントのバージョンを意味的に型付けされた要素に分解し、空間、コンテンツ、構造の互換性を共同でモデル化するアライメントファーストのメカニズムを通じてバージョン間の対応を確立し、整列された要素のペアに対して型を認識した差異推論を実行します。統合された変更検出、ローカリゼーション、構造認識分析、テキスト、表、数式、図にわたる位置合わせ/一致評価をサポートします。雑誌制作の校正ワークフローから得られる実際の科学 PDF データを対象とした実験では、提案されたフレームワークが要素固有のベースラインを常に上回るパフォーマンスを示しています。テキスト、表、数式、図に対してそれぞれ 0.903、0.855、0.862、0.845 の検出 F1 スコアを達成し、ローカリゼーション、構造認識、マッチング品質がさらに向上しました。アブレーション分析と感度分析により、バージョン間の調整、タイプ固有の表現、構造を意識した推論、および互換性を重視した設計の有効性が確認されます。これらの結果は、異種要素認識差分が、現実的な編集制作シナリオにおける科学文書比較のための堅牢で解釈可能なソリューションを提供することを示しています。

原文 (English)

Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning

Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challenging because scientific documents are page-structured artifacts containing heterogeneous elements such as text, tables, formulas, figures, and layout cues. Existing text-sequence-based methods often lose layout and structural information, while image-based methods lack semantic interpretability and are sensitive to rendering variation. To address these limitations, this paper proposes a layout-aware heterogeneous element-aware framework for scientific document differencing. The framework decomposes document versions into semantically typed elements, establishes cross-version correspondence through an alignment-first mechanism that jointly models spatial, content, and structural compatibility, and performs type-aware difference reasoning over aligned element pairs. It supports unified change detection, localization, structure-awareness analysis, and alignment/matching evaluation across text, tables, formulas, and figures. Experiments on real-world scientific PDF data from journal production proofreading workflows show that the proposed framework consistently outperforms element-specific baselines. It achieves detection F1 scores of 0.903, 0.855, 0.862, and 0.845 for text, tables, formulas, and figures, respectively, with further improvements in localization, structure awareness, and matching quality. Ablation and sensitivity analyses confirm the effectiveness of cross-version alignment, type-specific representations, structure-aware reasoning, and compatibility-weight design. These results demonstrate that heterogeneous element-aware differencing provides a robust and interpretable solution for scientific document comparison in realistic editorial production scenarios.

13:00 JSTエージェントClaude

Fantastic Adaptive Taxonomies and How to Use Them

An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajector…

13:00 JSTエージェント研究/論文

Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer

Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judg…

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

Towards an Automated Test of LLM Security Knowledge

Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM p…

13:00 JST研究/論文

REGEN: オフライン強化学習によるエキスパートからゼネラリストへの蒸留のためのリプレイリサイクル

大規模なオンライン強化学習 (RL) は、大規模言語モデル (LLM) での長期推論やエージェント ツールの使用などの高度な能力を引き出す主な手段です。ただし、特に RL を単なる 1 回限りの学習段階として考える場合、関心のある広大なタスク領域にわたってそれを拡張し続けることは、計算インフラストラクチャとコストの両方の点で依然として困難です。最近、さまざまなドメインやトレーニング段階にわたって知識を蒸留するために広く使用されている手法であるマルチ教師オンポリシー蒸留 (MOPD) は、広大なドメインにわたって汎用性を維持しながら、RL 段階を分離してコストを節約するのに役立ちます。それにもかかわらず、オンライン RL と同様に、MOPD は結合された推論と逆方向パスを必要とするため、そのスケーラビリティと計算効率が引き続き制限されます。これらの課題に対処するために、私たちは REGEN: オフライン RL を使用したエキスパートからジェネラリストへの蒸留のためのリプレイ リサイクルを提案します。 REGEN は、複数の教師モデルから抽出するのではなく、教師の専門的な RL トレーニングの無料の副産物であるリプレイ メモリをリサイクルし、オフライン RL アルゴリズムを採用するだけでジェネラリストをトレーニングします。 REGEN は、ロールアウト サンプリングをバックワード トレーニング プロセスから完全に切り離すため、トレーニング コストを大幅に削減します。 REGEN は、数学的推論、コード生成、および命令追従全体にわたって、大幅に低コストで MOPD の精度と同等の精度を実現します。オンライン RL を 1 回限りの学習段階ではなくデータ合成プロセスに変える可能性があり、大きな計算負荷を必要とせずに大規模なポストトレーニングに拡張できる可能性があります。

原文 (English)

REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning

Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and agentic tool use in large language models (LLMs). However, continuing to scale it across vast task domains of interest remains challenging in both computational infrastructure and cost, especially when considering RL as merely a one-off learning stage. Recently, a widely used technique for distilling knowledge across various domains and training stages, multi-teacher on-policy distillation (MOPD), helps to decouple the RL stage, saving costs, while maintaining generality across vast domains. Nonetheless, similar to online RL, MOPD requires coupled inference and backward passes, which continues to limit its scalability and computational efficiency. To address these challenges, we propose REGEN: Replay-recycling for Expert-to-Generalist Distillation with Offline RL. Instead of distilling from multiple teacher models, REGEN trains a generalist by simply recycling the replay memory -- the free by-product of the teachers' specialized RL training -- and employing offline RL algorithms. REGEN completely decouples the rollout sampling from the backward training process and thus greatly reduces the training cost. Across mathematical reasoning, code generation, and instruction following, REGEN matches the accuracy of MOPD at substantially lower cost. It potentially turns online RL into a data synthesis process instead of a one-off learning stage, and can be extended to large-scale post-training without requiring heavy computational load. Code is available at https://github.com/yunjie-sysu/REGEN.

13:00 JST研究/論文

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descri…

13:00 JST研究/論文

Adaptive Multi-Horizon Reinforcement Learning

Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement l…

13:00 JST研究/論文

On the Depth Scalability of Logic Gate Networks

Logic Gate Networks (LGNs) compute through compositions of Boolean operations, yet existing LGNs do not reliably benefit from increased dep…

13:00 JST研究/論文

Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models

Synthetic tabular data is prized for preserving not just each column's marginal distribution but the dependencies between columns - structu…

13:00 JST研究/論文

LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning

Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along am…

13:00 JST研究/論文

Directional Influence Function: Estimating Training Data Influence in Constrained Learning

As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety…

13:00 JST研究/論文

An Unofficial FastLAS Tutorial: A Programmer's Guide

FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of ex…

13:00 JSTLLM/生成AIエージェント

Where Is the Cost of Third-Party API Routers in Agentic Software Development?

Third-party API routers have become a common layer that unifies access across increasingly diverse LLM providers. In coding-agent workflows…

13:00 JSTエージェント

Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents

Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the…

13:00 JSTLLM/生成AI

Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature

X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published…

13:00 JST研究/論文

Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging

Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challeng…

13:00 JST画像/動画生成研究/論文

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks s…

13:00 JST研究/論文

極端なチョウラ集合とその線形類似物: Co-Scientist を使用した人間と AI の数学的調査

有限群におけるチョウラ型の秩序条件に関連する極不変量を導入します。有限群 $G$ の空でない部分集合 $S$ は、$S$ のすべての要素が $|S|$ より大きい順序を持つ場合、チョウラ集合と呼ばれ、そのような集合の最大カーディナリティを $C(G)$ と書きます。まず、$C(G)$ が $G$ 内の要素次数の分布によって決定されることを示します。巡回群の場合、正確な約数式を導出し、$C(\mathbb{Z}/n\mathbb{Z})=\varphi(n)$ となる整数 $n$ を特徴付けます。 $\liminf_{n\to\infty}C(\mathbb{Z}/n\mathbb{Z})/\varphi(n)=1$ であるのに対し、$\limsup_{n\to\infty}C(\mathbb{Z}/n\mathbb{Z})/\varphi(n)=\infty$ であることを証明し、$n$ による正規化の下で対応する下限と上限を決定します。有限アーベル群については、有限アーベル $p$ 群の閉じた式とともに、不変因子分解の観点から明示的な式を取得します。次に、有限体拡張の線形類似物を開発します。拡張 $L/K$ の非ゼロ $K$ 部分空間 $A$ は、A$ 内のすべての非ゼロ $a\ に対して $[K(a):K]>\dim_K A$ の場合、チョウラ部分空間と呼ばれます。この条件は $\dim_K A$ に依存するため、通常、$K$ に対して $L$ を生成するために $A$ のすべての非ゼロ要素は必要ありません。それにもかかわらず、$L/K$ が有限で分離可能な場合、正確な式 $C(L/K)=[L:K]-d_{\max}(L/K)$ を証明します。ここで、$d_{\max}(L/K)$ は、適切な中間体の $K$ に対する最大次数です。有限体については、正規基底構造を使用してあらゆる次数で直接証明を行います。この作品は、専門家の指導の下、人間と AI のコラボレーションを通じて開発されました。 Co-Scientist の推論に重点を置いた構成を使用して、例と潜在的な証明戦略を調査しました。著者たちは問題を定式化し、すべての議論を独自に検証して完成させ、最終的な証明を書きました。

原文 (English)

Extremal Chowla sets and their linear analogues: A human-AI mathematical investigation using Co-Scientist

We introduce an extremal invariant associated with Chowla-type order conditions in finite groups. A nonempty subset $S$ of a finite group $G$ is called a Chowla set if every element of $S$ has order greater than $|S|$, and we write $C(G)$ for the maximum cardinality of such a set. We first show that $C(G)$ is determined by the distribution of element orders in $G$. For cyclic groups, we derive an exact divisor formula and characterize the integers $n$ for which $C(\mathbb{Z}/n\mathbb{Z})=\varphi(n)$. We prove that $\liminf_{n\to\infty}C(\mathbb{Z}/n\mathbb{Z})/\varphi(n)=1$, whereas $\limsup_{n\to\infty}C(\mathbb{Z}/n\mathbb{Z})/\varphi(n)=\infty$, and we determine the corresponding lower and upper limits under normalization by $n$. For finite abelian groups, we obtain an explicit formula in terms of the invariant-factor decomposition, together with a closed formula for finite abelian $p$-groups. We then develop a linear analogue for finite field extensions. A nonzero $K$-subspace $A$ of an extension $L/K$ is called a Chowla subspace if $[K(a):K]>\dim_K A$ for every nonzero $a\in A$. Since this condition depends on $\dim_K A$, it does not generally require every nonzero element of $A$ to generate $L$ over $K$. Nevertheless, when $L/K$ is finite and separable, we prove the exact formula $C(L/K)=[L:K]-d_{\max}(L/K)$, where $d_{\max}(L/K)$ is the largest degree over $K$ of a proper intermediate field. For finite fields, we give a direct proof in every degree using a normal-basis construction. This work was developed through an expert-guided human-AI collaboration. A reasoning-focused configuration of Co-Scientist was used to explore examples and potential proof strategies. The authors formulated the problem, independently verified and completed all arguments, and wrote the final proofs.

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

「何を取得するか」を超えて: 取得拡張コード生成における不確実性

リポジトリ レベルのコード生成は、関連性、互換性、完全性が本質的に不確実な異種の証拠に依存しています。類似のコード例、リポジトリ コンテキスト、およびプロジェクト固有の API は補完的な情報を提供する可能性がありますが、ノイズの多い、冗長な、または競合する信号を引き起こす可能性もあります。既存の検索拡張アプローチは、検索された証拠の不確実性が下流の生成にどのような影響を与えるかを明示的にモデル化することなく、主に検索の関連性を最適化します。ソース固有の不確実性を推定し、それを使用して異種証拠をフィルタリングしてランク付けし、生成、検証、修復をガイドする不確実性認識フレームワークである OpenCoder を紹介します。 API の知識、リポジトリのコンテキスト、類似コードの証拠に対する要因分析では、普遍的な追加ソース ランキングがないことが明らかになりました。代わりに、重要なソース間の相互作用は、付随する証拠と LLM バックエンドに依存します。拡張された 32 タスクの RepoExec インライン評価では、OpenCoder はベースライン RAG に対する GPT 選択出力の正確性を 56.25\% から 78.13\% に向上させます。ただし、これは検証と修復の制御と一致しており、対応する Gemini の改善は統計的にサポートされておらず、バックエンドに依存する利点が示されています。ターゲットを意識した API の改良により、API セットの取得も大幅に向上します。これらの発見は、不確実性をリポジトリレベルの検索、検証、修復のための実用的な制御信号として扱うことを裏付けています。

原文 (English)

Beyond "What to Retrieve": Uncertainty in Retrieval-Augmented Code Generation

Repository-level code generation relies on heterogeneous evidence whose relevance, compatibility, and completeness are inherently uncertain. Similar-code examples, repository context, and project-specific APIs may provide complementary information, but can also introduce noisy, redundant, or conflicting signals. Existing retrieval-augmented approaches primarily optimize retrieval relevance without explicitly modeling how uncertainty in retrieved evidence affects downstream generation. We introduce OpenCoder, an uncertainty-aware framework that estimates source-specific uncertainty, uses it to filter and rank heterogeneous evidence, and guides generation, verification, and repair. A factorial analysis over API knowledge, repository context, and similar-code evidence reveals no universal additive source ranking; instead, significant cross-source interactions depend on the accompanying evidence and LLM backend. On an expanded 32-task RepoExec-inline evaluation, OpenCoder improves GPT selected-output correctness over Baseline RAG from 56.25\% to 78.13\%. However, it matches a verification-and-repair control, and the corresponding Gemini improvement is not statistically supported, indicating backend-dependent benefits. Target-aware API refinement also substantially improves API-set retrieval. These findings support treating uncertainty as an actionable control signal for repository-level retrieval, verification, and repair.

13:00 JSTLLM/生成AI

含意の認識とキャンセルによる大規模言語モデルにおけるコミュニケーション上の信念の更新の評価

人間の言語は暗黙の信念と信念の更新によって動かされるため、これらは大規模言語モデル (LLM) とそのユーザー間のコミュニケーションを成功させるモデルにとって重要です。この論文では、含意を通じて作られた暗黙の信念を認識し、含意のキャンセルを通じてその更新を理解するLLMの能力を評価します。これは、発話の暗黙の意味が弱められるか否定される実用的な現象です。私たちは、インプリカチャーとそれに対応するキャンセルを人間が判断するためにクラウドソーシングされた、初の専門家による注釈付きインプリカチャー キャンセル データセット ImplicatureX を作成しました。 LLM の信念更新の理解は、特により自然に発生するシナリオにおいて、人間のそれに比べて遅れていることがわかりました。追加の対照実験では、LLM 信念更新の成功の一部は以前の信念への依存に起因する可能性があり、信念更新の失敗はそのタイプと形式に依存する可能性があることを示唆しています。全体として、私たちの研究は、現在のLLMが暗黙の信念と信念の更新について人間レベルの理解にまだ達していないことを示唆しています。コードとデータは https://github.com/cesare-spinoso/ImplicatureX で入手できます。

原文 (English)

Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation

Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. In this paper, we evaluate the ability of LLMs to recognize unspoken beliefs made through implicatures and to understand their updates through implicature cancellation: the pragmatic phenomenon whereby an utterance's implied meaning is weakened or negated. We create the first expert-annotated implicature cancellation dataset, ImplicatureX, crowdsourced for human judgements of implicatures and their corresponding cancellations. We find that LLM belief update understanding lags behind that of humans, especially in more naturally-occurring scenarios. Additional control experiments suggest that successes in LLM belief updates may stem in part from a reliance on prior beliefs, and that failures in belief updates may depend on their type and on their form. Overall, our study suggests that current LLMs have not yet reached human-level understanding of unspoken beliefs and belief updates. Code and data are available at https://github.com/cesare-spinoso/ImplicatureX.

13:00 JSTLLM/生成AI

言語学者を雇うたびに推論コストが下がる: 効果的なプロンプト圧縮装置としての言語規則について

プロンプト圧縮により LLM 入力が短縮されて推論コストが削減されますが、既存の方法では LM フォワード パスを通じてトークンの重要性がスコアリングされます。このような微妙でコストのかかるトークンの選択が必要かどうかは疑問が残ります。圧縮には有益なコンテンツを特定する必要があります。この問題は、言語研究が決定論的なルールとして運用できる手がかりを通じて長年取り組んできました。したがって、圧縮時に LM ベースのスコアリングを行わずに、\textbf{言語規則のみ} が効果的なプロンプト圧縮として機能できるか、と考えます。これに対処するために、語彙、構文、意味、および談話のシードに対してオフラインの進化的検索を実行して、競合するルールの組み合わせを見つけます。結果として得られる言語圧縮プログラムは、展開時に LM フォワード パスを必要とせず、圧縮に CPU 側の処理のみを使用します。圧縮品質と再構築の忠実度のバランスをとるために、デュアルパス プロトコルを使用して評価します。短い文章、複数文書の推論、対話メモリ QA データセットにわたって、進化したコンプレッサーは、最近の高度なプロンプト圧縮戦略と同様のパフォーマンスを達成します。パフォーマンスは軽度から中度の圧縮下で最も強く、圧縮がより強力になると低下しますが、直接パスと再構築パスは異なるパターンを示します。進化的分析により、効果的な圧縮により言語レベル全体で信号が融合され、圧縮率が増加するにつれてルールがトークン プルーニングからセンテンス抽出に移行することが明らかになりました。

原文 (English)

Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors

Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It remains questionable whether such nuanced, costly token selection is necessary. Compression requires identifying informative content, a problem that linguistic research has long addressed through cues that can be operationalized as deterministic rules. We therefore ask: can \textbf{linguistic rules alone} serve as effective prompt compressors, without LM-based scoring at compression time? To address this, we conduct offline evolutionary search over lexical, syntactic, semantic, and discourse seeds to find competitive rule combinations. The resulting linguistic compressor requires no LM forward pass at deployment and uses only CPU-side processing for compression. We evaluate it with a dual-path protocol to balance compression quality and reconstruction fidelity. Across short passages, multi-document reasoning, and dialogue-memory QA datasets, evolved compressors achieve performance similar to that of recent advanced prompt-compression strategies. Performance is strongest under light-to-moderate compression and degrades as compression becomes more aggressive, while the Direct and Reconstruction paths exhibit distinct patterns. Evolutionary analysis reveals that effective compression fuses signals across linguistic levels and, as the compression ratio increases, rules shift from token pruning to sentence extraction.

13:00 JST画像/動画生成

I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models

The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial prog…

13:00 JST画像/動画生成エージェント

The LAIA Dataset: Labelled Attention for Intelligent Automobiles

The development of autonomous vehicles (AVs) usually relies heavily on data-driven artificial intelligence (AI) models that require large v…

13:00 JSTLLM/生成AI

Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs

Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evide…

13:00 JSTLLM/生成AIエージェント

Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction

Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects…

13:00 JSTLLM/生成AI

Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs

Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is dee…

13:00 JST画像/動画生成

Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition

Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change. Recognit…