Skip to the content.

AIニュース 2026-08-20

自動生成: 2026-08-20 10:35 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Offering Zero Data Retention for frontier modelsOpenAI

    OpenAI reaffirms Zero Data Retention for eligible API customers and p…

  2. Replit expands access to software creation with GPT-5.6 LunaOpenAI

    Replit introduces Free Mode, powered by GPT-5.6 Luna, so anyone can t…

  3. Google、大学生向けに「Google AI Plus」を1年間無料提供 「Gemini」アプリに学生向け新機能もITmedia AI+

    Googleは米国の新学期に合わせ、大学生向けに「Google AI Plus」などを12カ月無料で提供するキャンペーンを発表した。あわせ…

  4. OpenAI、ZDRを維持したまま悪用検知へ ログ保持を求めるAnthropicに対抗ITmedia AI+

    OpenAIは、APIデータの非保持設定(ZDR)を維持したまま、複数のやり取りを横断して悪用の兆候を検知する新機構「Private Sa…

  5. 高度なAIのサイバー攻撃、どう対策? 企業が知るべき「スピード格差」の埋め方ITmedia AI+

    米AnthropicのAIモデル「Claude Mythos 5」をはじめとした高度なAIの登場により、サイバーセキュリティの在り方に注目…

  6. ChatGPTの反論で「道徳的判断」の3割超が覆る 高齢者が説得されやすい傾向 神戸大ITmedia AI+

    生成AIの反論によって道徳的な判断の3割超が覆る――神戸大学がこのような研究結果を発表した。米OpenAIのチャットAI「ChatGPT」…

  7. Claude Codeの週次制限枠「50%増」、8月31日まで延長 「恒久化したいものの……」ITmedia AI+

    米Anthropicは、AIコーディング支援ツール「Claude Code」の週次利用制限枠を50%増やすキャンペーンを8月31日まで延長…

トピック別件数

日本語メディア14件

ITmedia AI+ (日本語)

09:39 JSTLLM/生成AIAnthropicOpenAI

OpenAI、ZDRを維持したまま悪用検知へ ログ保持を求めるAnthropicに対抗

OpenAIは、APIデータの非保持設定(ZDR)を維持したまま、複数のやり取りを横断して悪用の兆候を検知する新機構「Private Safety Processing」を発表した。顧客データを自社インフラ外や暗号化で保護しつつ、活動シグナルのみでリスクを判定する。9月に展開を…

09:00 JSTその他

PTC、「Onshape」でMCP連携 自然言語でカスタムCAD機能の作成が可能に

PTCは、CAD/PDMプラットフォーム「Onshape」において新機能「FeatureScript MCP Server」の提供を開始した。エンジニアが自然言語とAI(人工知能)を用いて、カスタムCAD機能を作成できるようにする。

08:00 JSTその他

モデルの利用料金は安くなっているのに、AIの総コスト上昇 「パラドクス」の背景を解説

AIモデルの利用料金の低下が、かえってAIの総コストを押し上げている。Gartnerはワークフロー1件当たりのAI推論コストが2028年までに5倍以上に上昇すると予測する。同社が「推論のパラドックス」と呼ぶ、この逆説の中身とは。また、AIコストが上昇する中でROIを確保するため…

08:00 JSTLLM/生成AIAnthropicClaude

高度なAIのサイバー攻撃、どう対策? 企業が知るべき「スピード格差」の埋め方

米AnthropicのAIモデル「Claude Mythos 5」をはじめとした高度なAIの登場により、サイバーセキュリティの在り方に注目が集まっている。企業に求められる対応を解説する。

07:58 JSTLLM/生成AIGoogleGemini2媒体が報道

Google、大学生向けに「Google AI Plus」を1年間無料提供 「Gemini」アプリに学生向け新機能も

Googleは米国の新学期に合わせ、大学生向けに「Google AI Plus」などを12カ月無料で提供するキャンペーンを発表した。あわせてGeminiアプリに学生向けハブを新設し、授業資料から学習プランを生成する学習ノートブックや、3Dモデルの表示、Gemini Liveでの…

出典:ITmedia AI+TechCrunch AI
06:45 JSTロボティクスGoogle

Waymoが自動運転AI戦略を解説、「単一AIモデルのE2E方式には2つの問題がある」

米国で自動運転車によるモビリティサービスを展開するWaymo(ウェイモ)が、2009年スタートの「Google Self-Driving Car Project」から開発を積み重ねてきた自動運転技術に基づく同社のAI戦略について説明した。

17:00 JSTLLM/生成AI研究/論文OpenAIGPT / ChatGPT

ChatGPTの反論で「道徳的判断」の3割超が覆る 高齢者が説得されやすい傾向 神戸大

生成AIの反論によって道徳的な判断の3割超が覆る――神戸大学がこのような研究結果を発表した。米OpenAIのチャットAI「ChatGPT」を利用し、AIの反論が正解のない道徳問題への回答に与える影響を調べた。

15:55 JSTLLM/生成AI研究/論文

「オープンな国産モデル」に33Bパラメータの新バージョン 国立情報学研究所

国立情報学研究所(NII)が、オープンな国産LLMの新バージョン「LLM-jp-4 33B」を公開。約332億パラメータのDense型モデルで、4種類のベンチマーク全てで従来モデルを上回るスコアを記録したという。

13:00 JSTその他

「AIを使える人か、使えない人か」で仕事や評価に差が? 6割のエンジニアが実感した“AI格差”の正体

AIを使うかどうかだけではなく、どの程度使いこなせるかも問われる中、活用スキルの差は業務効率だけではなく、仕事やキャリアにも影響し始めているという。何が起きているのか。ITエンジニア572人の調査から探る。

11:43 JSTLLM/生成AIOpenAI

「うちの会社にAIなんて」――謙遜する中小企業に眠る伸びしろ OpenAIがパートナー網拡充、新規3社が語る

米OpenAIが法人向けパートナーネットワークの日本での拡充を進めている。メディア向け座談会に登壇した新規パートナー3社が語ったのは、「うちの会社にAIなんて」と謙遜する中小企業にこそ眠る“伸びしろ”だった。

11:38 JSTLLM/生成AI

なぜNTTは「小型LLM」にこだわるのか 国産AI「tsuzumi 2」に込めた狙い

LLMは大規模であるほどよいのか。NTTが開発する国産LLM「tsuzumi 2」は、小型化と日本語処理、データ主権を重視する。企業が生成AIを実務に組み込む際の現実解として、同モデルは何を目指しているのか。開発担当者の講演から読み解く。

11:09 JST規制/政策Google

Googleや英政府、AIで飛行機雲を回避する大規模実証開始 航空業界の温暖化影響を減らす狙い

Googleは、英政府や航空管制大手NATSなどと共同で、AIを用いて飛行機雲を回避する実証プログラムを開始すると発表した。北大西洋のシャンウィック洋上空域を対象に、航路をわずかに変更して温暖化影響を抑える。空域規模での協調的回避の実証は世界初となり、Googleは計算基盤など…

11:04 JSTLLM/生成AIエージェントAnthropicClaude

Claude Codeの週次制限枠「50%増」、8月31日まで延長 「恒久化したいものの……」

米Anthropicは、AIコーディング支援ツール「Claude Code」の週次利用制限枠を50%増やすキャンペーンを8月31日まで延長すると発表した。

11:00 JSTその他

人材育成を邪魔する、「忙しすぎる現場」以外の要因は? ガートナーが指摘

デジタル人材育成が急務になる中、ガートナーが人材育成に関する課題を調査した。同社が指摘する、「学ぶ時間がない」の先にある課題とは。

海外メディア11件

TechCrunch AI (英語)

08:32 JSTLLM/生成AI

Stripe didn’t really buy OpenRouter because of the ‘singularity’

What does a payments giant want with a startup that routes prompts between different AI models? Stripe says it's because of "the singularit…

07:10 JSTLLM/生成AIAnthropicOpenAI

OpenAI seeks to one-up Anthropic with new customer privacy protections

A competition is developing between OpenAI and Anthropic over who can provide the best privacy protections for enterprise customer data.

06:51 JSTその他

Cognition CEO denies report that SpaceX tried to acquire the startup

SpaceX was reportedly in talks to buy AI coding startup Cognition. SpaceX has already acquired Cursor as it races to catch up to rivals lik…

04:11 JSTその他

AI was supposed to win people over by now — it hasn’t

As AI becomes harder to avoid, consumers are growing more wary of the technology — and Silicon Valley is discovering that widespread adopti…

03:46 JSTLLM/生成AI研究/論文OpenAI

Researchers say OpenAI revoked their access to limited cyber program

The idea behind OpenAI's Trusted Access for Cyber program is to give trusted defenders better models so they can report bugs and vulnerabil…

02:26 JSTハードウェア/半導体

Meet the startup helping Wall Street put a price on AI compute

The AI buildout shows no signs of slowing. And with hundreds of billions of dollars a year going into data centers and GPUs, compute has be…

00:44 JSTその他

TerraPower’s nuclear reactor has a secret weapon for powering AI data centers

TerraPower's nuclear power plant possesses a strategic advantage over competitors, especially when chasing after data center deals.

00:00 JSTその他

Amazon makes its AI-powered Alexa+ free on Fire TV, no Prime required

Amazon is making its AI-powered Alexa+ assistant free on all compatible Fire TV devices in the U.S., automatically upgrading users whether…

23:09 JSTその他

Calendly throws its hat into meeting note-taker circus

Calendly is also releasing a meeting scheduling assistant called Callie.

19:00 JSTビジネス/資金調達

Relativity Networks raises $22 million to bring a faster kind of fiber to data centers

Relativity Networks deals in hollow-core fiber, a rarely deployed technology that allows data to be transmitted 30% faster than conventiona…

公式ブログ2件

OpenAI (英語)

04:00 JSTLLM/生成AIOpenAI

Offering Zero Data Retention for frontier models

OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compr…

16:00 JSTLLM/生成AIGPT / ChatGPT

Replit expands access to software creation with GPT-5.6 Luna

Replit introduces Free Mode, powered by GPT-5.6 Luna, so anyone can turn ideas into working software without worrying about token costs.

論文275件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIエージェント規制/政策ClaudeGPT / ChatGPT

GxP-Agent: LLM エージェントを使用した信頼性の高い臨床試験プログラミングのためのプロセス DAG トポロジ

CDISC 標準に基づいて研究プロトコルを分析可能なデータセットに変換する臨床試験プログラミングは、規制当局への申請のボトルネックとなっていますが、LLM ベースのコード生成はこのタスクで壊滅的に失敗します。5 つのフロンティア モデルを使用した 11 回の単発試行にわたって、有効な被験者レベルの分析データセットを生成するものはありませんでした。 GxP-Agent は、規制プロセスの順序付けを有向非巡回グラフ (DAG) としてエンコードするマルチエージェント システムであり、モノリシック データセットの生成を、ファーマバース スキル コンテキスト、検証ゲート、および条件付き再試行を備えたワーカー エージェントによって実行される 15 のドメイン固有のノードに分解します。 FDA パイロット提出の CDISCPilot01 (被験者 254 人、グラウンドトゥルース ADSL 変数 49 個) から構築された新しい実行ベースのベンチマークである CDISC-Bench では、GxP-Agent と Claude Sonnet 4.6 は 3 回の独立した実行で 100% の構造一致 (49/49 個の変数、254 個の正しいレコード) を達成しました。これに対し、最良の検索拡張ベースラインでは 59.2% でした。すべての単一エージェント アプローチおよびフラット マルチエージェント アプローチでは 0%。 DAG トポロジでは、より弱いモデルも可能になります。GPT-4.1 は、同じ DAG の下で 59.2% の平均構造一致を達成しますが、他のすべてのアーキテクチャの下ではスコアが 0% です。このアプローチは ADAE (有害事象、9 ノード分岐 DAG、55 変数、1,191 レコード) に一般化されており、最初の試行で 100% の構造一致を達成します。これらの結果は、LLM 推論だけに依存するのではなく、ドメイン プロセスの知識をグラフ トポロジとしてエンコードすることが、信頼性の高い GxP 準拠の臨床試験プログラミングを可能にする重要な要素であることを示しています。

原文 (English)

GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents

Clinical trial programming -- transforming study protocols into analysis-ready datasets under CDISC standards -- is a bottleneck in regulatory submissions, yet LLM-based code generation fails catastrophically on this task: across 11 single-shot attempts with five frontier models, none produces a valid subject-level analysis dataset. We introduce GxP-Agent, a multi-agent system that encodes regulatory process ordering as a directed acyclic graph (DAG), decomposing monolithic dataset generation into 15 domain-specific nodes executed by worker agents with pharmaverse skill context, validation gates, and conditional retry. On CDISC-Bench, a new execution-based benchmark built from the FDA pilot submission CDISCPilot01 (254 subjects, 49 ground-truth ADSL variables), GxP-Agent with Claude Sonnet 4.6 achieves 100% structural match (49/49 variables, 254 correct records) across three independent runs, compared to 59.2% for the best retrieval-augmented baseline and 0% for all single-agent and flat multi-agent approaches. The DAG topology also enables weaker models: GPT-4.1 achieves 59.2% mean structural match under the same DAG, where it scores 0% under every other architecture. The approach generalizes to ADAE (adverse events; 9-node branching DAG, 55 variables, 1,191 records), achieving 100% structural match on the first attempt. These results demonstrate that encoding domain process knowledge as graph topology -- rather than relying on LLM reasoning alone -- is a key enabler for reliable, GxP-compliant clinical trial programming.

13:00 JSTエージェント

Agentic AI のランタイム ガバナンス: 信頼できる出所とフェールクローズ実行によるアクション境界制御

Agentic AI システムは、ファイルの変更、メッセージの送信、ジョブの起動、またはワークフロー状態の変更を行うツール アクションを要求します。これにより、安全性の問題は、有害なテキストの生成から有害な操作上の副作用に移ります。プロンプトレベルのガバナンスはモデルの動作を形成できますが、実行境界は作成されません。 Aegis は、モデルの出力をアクション提案として扱い、ツールの実行前に信頼できる意思決定層を通じて仲介するランタイム ガバナンス システムです。このモデルは次のように提案しています。信頼できるランタイムが決定します。 Aegis は、アクティブなポリシーの状態に照らして提案を評価し、サーバー側で来歴を解決し、不確実性の下でフェールクローズし、選択されたケースを定足数ベースの非一方的承認パスである上院スタイルの和解を通じてルーティングします。私たちは、5 つの実行ファミリー、42 のタスク、3 つの条件、およびファミリーあたり 10 回の繰り返しにわたる繰り返しサンドボックス コーパスで Aegis を評価します。 6,300 行にわたって、プロンプト ポリシー コンディショニングにより、79 個の危険なコンパレータ パス リーク行が生成されました。 Aegis が管理する 2,100 行にわたって、システムは管理されたモックツール アプリケーションと管理された危険な副作用の完了をゼロとして記録しました。イージスが試みた1,832の統治対象行はすべて、イージスが解決した信頼できる出所を保存しており、上院が解決した1,019の行はすべて、定足数と最終署名された集計証拠を持っていた。これらの結果は、一般的な自律エージェントの安全性を証明するものではありません。彼らは、この評価されたサンドボックス コーパスでは、実行時のアクション境界ガバナンスにより、観察された危険な提案がガバナンスされた副作用になるのを防いだという、より狭いシステムの主張を支持しています。

原文 (English)

Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety problem from harmful text generation to harmful operational side effects. Prompt-level governance can shape model behavior, but it does not create an execution boundary. We introduce Aegis, a runtime governance system that treats model outputs as action proposals and mediates them through a trusted decision layer before tool execution. The model proposes; the trusted runtime decides. Aegis evaluates proposals against active policy state, resolves provenance server-side, fails closed under uncertainty, and routes selected cases through Senate-style settlement, a quorum- based non-unilateral authorization path. We evaluate Aegis on a repeated sandbox corpus spanning five run families, 42 tasks, three conditions, and ten repeats per family. Across 6,300 rows, prompt-policy conditioning produced 79 risky comparator-path leakage rows. Across 2,100 Aegis-governed rows, the system recorded zero governed mock-tool applications and zero governed risky side-effect completions. All 1,832 Aegis-attempted governed rows preserved trusted Aegis-resolved provenance, and all 1,019 Senate-settled rows had quorum and final signed tally evidence. These results do not prove general autonomous-agent safety. They support the narrower systems claim that, in this evaluated sandbox corpus, runtime action-boundary governance prevented observed risky proposals from becoming governed side effects.

13:00 JSTLLM/生成AIClaude

思考の代償: モデル固有の API コントラクトとしての推論の労力

API 購入者は、モデル名だけではなく日付付きの契約を購入します。契約には、要求および提供されたモデル、推論努力期間またはその省略、出力レール、サービス製品、プロンプト、および価格スケジュールが含まれます。 30 個の AIME 2026 項目と項目ごとに 5 つの呼び出しを使用して、努力を省略した同じモデルに対する明示的な高い努力を伴う Sonnet 5 の登録されたペアの対比を通じて、推論と努力の項を研究します。有料の試行ごとに 1 つの凍結端末カテゴリが割り当てられ、繰り返される呼び出しを保持しながらリサンプリングされたアイテムが推論されました。平均配信コストは、明示的高契約の場合、省略された契約よりもコールあたり \$0.01031 高くなりました [+\$0.00204、+\$0.01974]。対応する精度コントラストは +0.0133 [-0.0267, +0.0467] でした。精度の差は検出されず、この間隔では最大 4.67 パーセント ポイントのゲインが許容されますが、この設計では除外できません。登録されたポイント推定値として、正解あたりのコストは、高労力契約では \$0.08665、省略契約では \$0.07662 でした。日付の付いた契約調査、Models-API メタデータ、および事前登録された生の応答プローブにより、プロバイダー内を含むモデル固有の省略セマンティクスがさらに文書化されました。生の構造が不確定な場合、クレームは文書レベルに留まりました。リクエスト レジストリ、パーサー、ターミナル分類法、統計計画、分析パイプラインは、結果が検査される前に凍結されました。結果として生じるクレームは、調査されたモデル、タスク、および収集日に限定されます。

原文 (English)

The Price of Thinking: Reasoning Effort as a Model-Specific API Contract

API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using 30 AIME 2026 items and five calls per item. Every paid attempt was assigned one frozen terminal category, and inference resampled items while retaining their repeated calls. Mean delivered cost was \$0.01031 per call higher under the explicit-high contract than under the omitted contract [+\$0.00204, +\$0.01974]. The corresponding accuracy contrast was +0.0133 [-0.0267, +0.0467]; we did not detect an accuracy difference, and the interval permits a gain of up to 4.67 percentage points that this design cannot rule out. Cost per correct answer was \$0.08665 under the high-effort contract and \$0.07662 under the omitted contract, as registered point estimates. A dated contract census, Models-API metadata, and preregistered raw-response probes further documented model-specific omission semantics, including within a provider; claims remained at documentation grade when raw structure was indeterminate. The request registry, parser, terminal taxonomy, statistical plan, and analysis pipeline were frozen before outcomes were examined; the resulting claims are bounded to the model, task, and collection date studied.

13:00 JST研究/論文

FedPref: Federated Preference Learning for Structured Radiology Report Extraction

Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schem…

13:00 JST研究/論文

問題は問題です: スケーラブルな数学的発見に向けて

AI システムは数学研究にますます貢献できるようになっています。研究の実践においては、フロンティアモデル推論のリソースは限られており、専門家の数学的レビューにはさらに厳しい制約があります。したがって、これらの希少なリソースを適切に割り当てることが、AI 支援による数学的発見を効率的に行う上で中心となります。現在のほとんどの数学用 AI ワークフローでは、人間の労力は、適切な研究問題を選択し、その後結果として得られた成果物をレビューするという最初と最後に集中しています。これら 2 つの段階が研究レベルの数学のボトルネックになりつつあります。私たちは、人間と AI を組み合わせた新しい発見パラダイムを提案することで、これらの問題に対処します。人間のインプットは、もはや事前に選択された単一の問題ではなく、専門家が関心と専門知識を持っている研究の方向性となります。次に、システムは広範な文献コーパスからその方向の候補問題を検索します。検索システムとレコメンダー システムからインスピレーションを得て、適切な問題の検索を自動化し、フィルターのいくつかの段階を通過したアーティファクトに人間の注意を集中させる、文献からレビューまでのカスケードである検索、試行、推奨 (FAR) を構築します。組み合わせ論のパイロットでは、パイプラインは 5,245 件の組み合わせ論の論文から開始され、6,453 件の候補の予想または未解決の問題を回収し、それらをフィルタリングして、明らかに適切に設定された未解決の 4,717 件の予想を抽出します。その後の推論と自動トリアージ段階により、598 件の潜在的な解決策が明らかになり、著者チームのレビューのために 77 件の項目が選択されます。その中には、デイヴィス--ジェンセン--パーキンス--ロバーツ、エルド・ヘス--ストラウス、イケンマイヤー--パク--パノヴァ、ルンド--サラフ--ウルフの推測や質問の結果を含む、多くの興味深い発見が含まれている。これらの結果は、人間と AI のコラボレーションのこの新しいモードが数学的発見に有効であることを示しています。

原文 (English)

The Problem Is the Problem: Towards Scalable Mathematical Discovery

AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited resource, and expert mathematical review is even more sharply constrained. Allocating these scarce resources well is therefore central to making AI-assisted mathematical discovery efficient. In most current AI-for-math workflows, human effort is concentrated at the beginning and end, in selecting suitable research problems and later reviewing the resulting artifacts. These two stages are becoming bottlenecks for research-level mathematics. We address them by proposing a new human-AI discovery paradigm. The human input is no longer a single problem selected in advance, but a research direction in which the experts have interest and expertise. The system then searches a broad literature corpus for candidate problems in that direction. Inspired by search and recommender systems, we build Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering. In a combinatorics pilot, the pipeline starts from 5,245 combinatorics papers, recovers 6,453 candidate conjectures or open problems, and filters them to 4,717 apparently well-posed and still-open conjectures. Subsequent reasoning and automated triage stages surface 598 potential resolutions and select 77 items for author-team review. Among them, we identify many interesting discoveries, including results on conjectures and questions of Davies--Jenssen--Perkins--Roberts, Erd\H{o}s--Straus, Ikenmeyer--Pak--Panova, and Lund--Saraf--Wolf. These results demonstrate the effectiveness of this new mode of human-AI collaboration for mathematical discovery.

13:00 JSTエージェント

SkillEffect: Checked Lowering for Memory-Bounded Agent Tools

Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. Howev…

13:00 JSTLLM/生成AIエージェント

記憶はコミュニケーションです: 記憶と信号伝達の間のフロンティア

境界付きエージェントは、自身の過去、ピア、または両方のソースから決定のための情報を取得する場合があります。タスク関連の履歴を保持すると、その後の通信を減らすことができますが、不足しているメモリをピア メッセージで補うことができます。両方のリソースに制限がある場合、エージェントは情報予算をどのように割り当てるべきでしょうか?タスクと決定ルールが固定されている場合、パフォーマンスしきい値に達するメモリとメッセージ レートのペアは、履歴とピア観察を使用するための指定されたルールの下で達成可能な領域を形成します。私たちはその効率的な境界を記憶-信号フロンティアと呼んでいます。履歴によってタスク損失の最大削減が同じように許可される条件全体で、境界付きエージェントが履歴からより大きな損失削減を取得すると、必要なピア通信が少なくなるという仮説を立てます。予備的な参照ゲームでは、ターゲットの繰り返しはより短い成功メッセージと一致しましたが、隠された循環ルールによる予測可能性はメッセージを短縮しませんでした。メモリとメッセージ レートを変化させて実験することで、フロンティアを推定し、協調タスク全体でこの予測をテストできます。

原文 (English)

Memory Is Communication: The Frontier Between Remembering and Signaling

A bounded agent may obtain information for a decision from its own past, from peers, or from both sources. Retaining task-relevant history can reduce later communication, while a peer message can supply what memory lacks. Under limits on both resources, how should an agent allocate its information budget? Given a fixed task and decision rule, the memory and message rate pairs attaining a performance threshold form an achievable region under specified rules for using history and peer observations. We call its efficient boundary the remembering--signaling frontier. Across conditions where history permits the same maximum reduction in task loss, we hypothesize that a bounded agent will need less peer communication when it obtains a larger loss reduction from history. In preliminary referential games, target repetition coincided with shorter successful messages, while predictability from a hidden cyclic rule did not shorten them. Experiments varying memory and message rates can estimate the frontier and test this prediction across cooperative tasks.

13:00 JSTLLM/生成AI画像/動画生成ビジネス/資金調達

DiSCO: ディストリビューションに基づいたコントラスト プロンプトの最適化により、テキストから画像への生成を防御

テキストから画像への生成モデルが進歩するにつれて、重大な安全上の懸念が生じ、特に暴力やヌードなどの非安全作業 (NSFW) コンテンツの生成が生じ、レッドチームによる敵対的攻撃によってさらに悪化します。既存の防御策は主にホワイトボックスの仮定の下で動作し、テキストエンコーダーの最適化、重み編集、推論時の介入に依存しており、基本的に独自のモデルに拡張することができません。 LLM プロンプト書き換えに基づくブラック ボックスの代替案は、より幅広い適用可能性を提供しますが、\textit{benign adversarial} 問題として私たちが特定する重要な領域では失敗します。プロンプトは、言語的には安全ですが、依然としてモデルの学習されたデータ分布により有害な生成を引き起こします。私たちは、プラグアンドプレイ モジュールとしてプロンプト レベルで完全に動作し、モデルの再トレーニング、微調整、モデル内部へのアクセスを必要としないゼロショットの厳密なブラックボックス防御である DiSCO を提案します。 DiSCO は、安全なコンテンツが生成されるまで反復適応フィードバックを使用して、ターゲット モデル自体によって生成された安全な画像プールと安全でない画像プールに対する対比スコアリングを通じて最適化された、ビーム サーチを介した分布ガイドによるサフィックス拡張を実行します。我々は、DiSCO が、複数のレッドチーム攻撃の下で、I2P ベンチマークで防御されていないモデルと防御されたモデルの両方の安全性を一貫して強化し、それぞれ 37.7% と 25.13% の ASR 削減を達成しながら、セマンティック忠実度を維持し、画像の一貫性を向上させることを実証します。ブラックボックスでアーキテクチャに依存しないモジュールである DiSCO は、モデル自体を変更することなく、あらゆるテキストから画像への変換システムに容易に適用できます。

原文 (English)

DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization

As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alternatives based on LLM prompt rewriting offer broader applicability, yet fail in a critical regime we identify as the \textit{benign adversarial} problem: prompts that are linguistically safe but still trigger harmful generation due to the model's learned data distribution. We propose DiSCO, a zero-shot, strictly black-box defense that operates entirely at the prompt level as a plug-and-play module, requiring no model retraining, fine-tuning, or access to model internals. DiSCO performs distribution-guided suffix expansion via beam search, optimized through contrastive scoring over safe and unsafe image pools generated by the target model itself, with iterative adaptive feedback until safe content is produced. We demonstrate that DiSCO consistently enhances the safety of both undefended and defended models on the I2P benchmark under multiple red-teaming attacks, achieving 37.7% and 25.13% ASR reduction, respectively, while maintaining semantic fidelity and improving image coherence. As a black-box, architecture-agnostic module, DiSCO can be readily applied to any text-to-image system without necessitating any changes to the model itself.

13:00 JSTエージェントハードウェア/半導体ClaudeNVIDIA

KernelArc: GPU カーネル最適化のためのマルチエージェント フレームワーク

異種ワークロード全体で自律的に GPU カーネルを最適化するためのマルチエージェント フレームワークである KernelArc を紹介します。戦略に特化したエージェントは並行して実行され、結論のみの共有メモリ、決定論的なベンチマーク ガード、およびプラトー トリガーのドラフティングによる読み取り専用のクロスエージェント状態を通じて調整されます。カテゴリを代表する SOL-ExecBench ワークロードを使用して、NVIDIA H100 および B200 GPU 上の \kernelarc{} を評価します。結果として得られる実装は、カスタム BF16 GEMM、静的 cuBLASLt Expert-API 構成テーブル、後方融合エキスパート混合、シェイプゲート デコーダ層融合、ネイティブ NVFP4 グループ化クエリ アテンション、およびページング プレフィル アテンションに及びます。 2026 年 7 月~30 日に記録された公開 SOL-ExecBench リーダーボード スナップショットでは、これらの提出物は代表的な L1、L2、量子化、および FlashInfer タスクで 1 位にランクされました。この軌跡は、この論文の中心的な動機を裏付けています。つまり、マルチエージェントの共有検索により調査範囲が広がり、一定の候補予算内で強力な既存企業に到達できる一方で、個々の調整機能の価値はカーネルと最適化の段階に依存します。

原文 (English)

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. We evaluate \kernelarc{} on NVIDIA H100 and B200 GPUs using category-representative SOL-ExecBench workloads. The resulting implementations span custom BF16 GEMM, static cuBLASLt Expert-API configuration tables, fused mixture-of-experts backward, shape-gated decoder-layer fusion, native NVFP4 grouped-query attention, and paged prefill attention. At the public SOL-ExecBench leaderboard snapshot recorded on July~30, 2026, these submissions ranked first on representative L1, L2, Quantization, and FlashInfer tasks. The trajectories support the paper's central motivation: shared multi-agent search can broaden exploration and reach stronger incumbents within a fixed candidate budget, while the value of individual coordination features depends on the kernel and optimization stage.

13:00 JSTLLM/生成AI

デコード可能性の基準は、大規模な言語モデルで隠れ状態の選択が多数決を上回る時期を予測します

大規模言語モデル (LLM) が質問に対してサンプルした回答を 1 つの決定に結合することは、テスト時の情報融合の問題であり、通常は多数決によって解決されます。サンプル化された回答が相関関係にあるエラーを共有するような難しい質問では投票の信頼性が低く、間違った回答が勝つ可能性があり、より多くのサンプルを抽出すると決定が悪化します。モデルの隠れ状態から正確性信号を読み取ることによって候補を選択することは有望な代替手段ですが、その精度はモデルやタスクによって異なり、いつ信頼できるかを示す尺度はありません。この論文では、答えトークンの隠れ状態で線形ゲートを訓練し、最高スコアの候補を選択する動的選択結合器である CASE (Correctness-Axis SElection) を提案します。その主な貢献は解読可能性であり、ゲートが質問の正しい候補を不正解の候補よりどの程度ランク付けするかを漏れのない尺度で示し、隠れ状態の選択が投票よりも優れているかどうかを予測します。従来のプローブは、質問の同一性の漏洩によってのみ正確に見えますが、質問のグループ化された評価ではこの漏洩は消失します。保持されたデータでは、復号可能性により、ピアソン相関 r=0.75 および AUC=0.60 付近の決定閾値で、投票よりも選択の精度が向上すると予測されます。一般および医療 LLM 全体で、CASE は投票よりも中程度の難易度の質問で最大 19 ポイント、難しい質問で 16.8 ポイント改善しました。解読可能性は、モデルのスケールではなく、モデルが呼び出す必要がある調整された知識に依存し、その予測は 3.8 ポイント以内で目に見えない科学領域に移行します。したがって、学習された選択と多数決のどちらを選択するかについて、特定のモデルとタスクについて事前に測定可能な実用的な基準が提供されます。

原文 (English)

A decodability criterion predicts when hidden-state selection beats majority voting in large language models

Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usually solved by majority voting. Voting is unreliable on difficult questions, where the sampled answers share correlated errors, so the wrong answer can win and drawing more samples makes the decision worse. Selecting a candidate by reading a correctness signal from the model's hidden states is a promising alternative, but its accuracy varies across models and tasks, and no measure indicates when it can be trusted. In this paper, we propose CASE (Correctness-Axis SElection), a dynamic selection combiner that trains a linear gate on the answer-token hidden state and selects the highest-scoring candidate. Its main contribution is decodability, a leakage-free measure of how well the gate ranks a question's correct candidates above its incorrect ones, which predicts whether hidden-state selection will outperform voting. A conventional probe appears accurate only because of question-identity leakage, which vanishes under question-grouped evaluation. On held-out data, decodability predicts the accuracy gain of selection over voting with a Pearson correlation r=0.75 and a decision threshold near AUC=0.60. Across general and medical LLMs, CASE improves over voting by up to 19 points on medium-difficulty questions and 16.8 points on hard questions. Decodability depends on the aligned knowledge a model must recall, not on its scale, and its prediction transfers to an unseen scientific domain within 3.8 points. It thus provides a practical criterion, measurable in advance for a given model and task, for choosing between learned selection and majority voting.

13:00 JST研究/論文

協力的な観察を通じて個人の知性を目指して

パーソナル AI システムには、ユーザーの目標、制約、ユーザーに代わって計画を立てて行動するための継続的なコミットメントのモデルが必要ですが、そのモデルの品質はシステムが観察できる内容によって制限されます。境界のあるシステムでは、当面のタスクに必要な情報を選択して圧縮する必要があるため、より広範囲に観察すること自体が支援を向上させるわけではありません。私たちは、この観察のボトルネックには協調的な構造があると主張します。つまり、システムがユーザーの変化する生活の部分モデルを構築し、ユーザーがその行動を評価し、ユーザーの同意と制御が次に観察できるものを形作るのです。有用で検査可能な動作は、ユーザーに観察チャネルを維持または拡張する理由を与えることができますが、障害が発生すると、観察チャネルの修正、縮小、取り消し、または放棄につながる可能性があります。私たちは、有用性、信頼、将来のアクセスの間のこのフィードバック ループを協力観察という用語を使用し、個人の知性のフレームワークとして提案します。私たちは、Organizm からの予備的な単一被験者アカウント、6 か月間使用されたプロトタイプを報告し、観察品質がパーソナル AI をどのように形成するかを測定するための評価方向性を概説します。

原文 (English)

Toward Personal Intelligence Through Cooperative Observation

A personal AI system needs a model of the user's goals, constraints, and ongoing commitments to plan and act on their behalf, and the quality of that model is bounded by what the system can observe. Broader observation does not by itself improve assistance because a bounded system must select and compress information for the task at hand. We argue that this observation bottleneck has a cooperative structure: the system builds a partial model of the user's changing life, the user evaluates its actions, and the user's consent and control shape what it can observe next. Useful and inspectable behavior can give users a reason to maintain or expand the observation channel, while failures can lead them to correct, narrow, revoke, or abandon it. We use the term cooperative observation for this feedback loop among usefulness, trust, and future access, and propose it as a framework for personal intelligence. We report a preliminary single-subject account from Organizm, a prototype used over six months, and outline evaluation directions for measuring how observation quality shapes personal AI.

13:00 JSTLLM/生成AI

KnowSim: 学習するユーザー シミュレーターを使用した LLM アシスタントでの情報キャリブレーションの評価

知識集約的なタスクでユーザーと効果的にコラボレーションするには、大規模言語モデル (LLM) が情報調整を実行する必要があります。つまり、ユーザーの進化する理解力と認知能力にコンテンツを一致させる必要があります。しかし、LLM の評価とトレーニングに使用されるユーザー シミュレーターは、ユーザーの知識を明示的にモデル化していないため、知識レベル間で現実的な相互作用を生成したり、知識の進化に応じて相互作用がどのように展開するかを反映したりすることはありません。このギャップを埋めるために、KNOWSIM を導入します。KNOWSIM は、前提条件の関係を持つ情報単位のグラフとして表される明示的な知識の状態を維持するユーザー シミュレーターを中心に構築された評価フレームワークであり、学習理論に基づいた更新ルールの下で進化します。 KNOWSIM は、情報調整の重要なメカニズムの側面を反映して、知識状態の軌跡から 3 つの指標 (知識獲得、配信調整、認知過負荷) を直接計算します。私たちは、知識レベルによって階層化された 2 つのドメインにわたる 705 の人間と AI のセッションに対して KNOWSIM を検証しました。そのランキングは人間の判断と大幅に一致し (73 ~ 74% が一致)、3 つのベースライン シミュレーターを上回りました。 KNOWSIM を 9 つの LLM に適用すると、最適なモデルがユーザーの知識レベルによって変化することが明らかになり、標準的な評価では見えない適性と治療の相互作用が明らかになります。

原文 (English)

KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn

To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.

13:00 JSTエージェント

特徴抽出器の合成: アルゴリズム選択のためのエージェント的アプローチ

制約充足問題のアルゴリズムを選択するには、問題の構造を捉える特徴を抽出する必要があります。特徴抽出機能を手動で設計するには、専門分野の深い専門知識が必要であり、新しい問題クラスが現れるとすぐにボトルネックになります。エージェントのチェック-修正-検証ループで大規模言語モデル (LLM) を使用して、解釈可能な問題固有の特徴抽出機能として機能する実行可能な Python スクリプトを合成する自動化アプローチを紹介します。高レベルの MiniZinc モデルとインスタンスが与えられると、LLM エージェントは、型付きグラフ表現を構築し、グラフ密度、変数クラスタリング、制約の厳しさなどの構造プロパティを計算するコードを生成します。 5 つの最先端ソルバーのポートフォリオを使用して、3 つの組み合わせ問題 (車両ルート、車両順序、固定長誤り訂正符号) に対するアプローチを評価します。合成されたエクストラクターは、専門家が厳選した mzn2feat 機能 (FLECC で最大 $8.3$ パーセンテージ ポイント (pp) のテストセット精度) と最高のトランスベースの trans2feat バリアントの両方を常に上回るアルゴリズム セレクターを生成します。その間、合成された特徴抽出器は引き続き検査可能です。

原文 (English)

Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection

Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing feature extractors demands deep domain expertise and quickly becomes a bottleneck when new problem classes appear. We present an automated approach that uses Large Language Models (LLMs) in an agentic check--fix--verify loop to synthesize executable Python scripts that act as interpretable, problem-specific feature extractors. Given a high-level MiniZinc model and an instance, the LLM agent generates code that constructs a typed graph representation and computes structural properties such as graph density, variable clustering, and constraint tightness. We evaluate our approach on three combinatorial problems (vehicle routing, car sequencing, fixed-length error-correcting codes) with a portfolio of five state-of-the-art solvers. The synthesized extractors yield algorithm selectors that consistently outperform both expert-curated mzn2feat features (up to $8.3$ percentage points (pp) test-set accuracy on FLECC) and the best transformer-based trans2feat variants. In the meanwhile, the synthesized feature extractors remain inspectable.

13:00 JST研究/論文

ベンチマークのベンチマーク: 小規模言語モデルの自動安全性ベンチマークの評価

Small Language Model (SLM) は、リソースに制約があり、プライバシーに敏感な設定で導入されることが増えており、安全性とバイアスの障害がセキュリティと社会的リスクを引き起こす可能性があります。ただし、既存の AI セーフティ/スラッシュ セキュリティ/スラッシュ コンプライアンス ベンチマークは、SLM に確実に転送できない可能性がある大規模な言語モデルを対象に設計されています。したがって、これらのベンチマークは SLM を効果的かつ確実に評価できるでしょうか?この質問に答えるために、私たちは 26 のオープンソース SLM にわたって広く使用されている 5 つのベンチマーク スイートを評価することにより、これらの自動化されたパイプラインの有効性と堅牢性の大規模評価を実施します。統一された判断ルーブリックに基づいて、有害、安全、または曖昧/無関係な応答にそれぞれ 0、1、または 0.5 のスコアが割り当てられます。ベンチマーク全体では、あいまいな判断が優勢であり、即時的な複雑性やモデル アーキテクチャと相関関係にあります。これは、{\em LLM 中心の安全性ベンチマークが SLM の安全性評価の独立した証拠として不十分であることを示しています}。一般に、あいまいさの割合は、語彙の密度、出力の複雑さ、および出力の長さとともに増加し、語彙の高度さ、自己一貫性、および応答プロンプトの類似性とともに減少します。これにより、モデルの機能と見かけの安全性が混同される機能と安全性の交絡が明らかになります。あいまいさが蔓延しているため、集計平均スコアのリーダーボードは数学的に脆弱です。モデルのランキングは、基礎となる出力が変わらない場合でも、合理的なあいまいさの処理の下では大幅に変化します。

原文 (English)

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large language models that may not transfer reliably to SLMs. We therefore ask: Can these benchmarks effectively and reliably evaluate SLMs? To answer this question, we conduct a large-scale assessment of the effectiveness and robustness of these automated pipelines by evaluating five widely used benchmark suites across 26 open-source SLMs under a unified judging rubric, which assigns a score of 0, 1, or 0.5 to harmful, safe, or ambiguous/irrelevant responses, respectively. Across the benchmarks, ambiguous judgments dominate and correlate with prompt complexity and model architecture, indicating that {\em LLM-centric safety benchmarks are insufficient as standalone evidence for SLM safety assessment}. In general, the ambiguity rate increases with lexical density, output perplexity, and output length and decreases with lexical sophistication, self-coherence, and reply-prompt similarity. This reveals a capability-safety confound that mixes model capability with apparent safety. Since ambiguity is prevalent, aggregate mean-score leaderboards are mathematically brittle: model rankings change significantly under reasonable ambiguity treatments, even when the underlying outputs remain unchanged.

13:00 JST研究/論文

フールズ・ゴールド: 無差別重量モデルに対する安全取り外し攻撃に対する防御的欺瞞

オープンウェイト言語モデルにおける安全性の調整は簡単に削除できます。削除は数分でウェイトから拒否を媒介する方向を投影しますが、私たちが知っている限り、それを永続的に防ぐリリース時の防御はありません。防ぐことができないものは騙される可能性があります。私たちの防御手段である囮の強化 (「愚者の黄金」) は、拒否の剥奪を認め、その見返りを台無しにします。一度拒否が剥奪されると、危険な作戦要求に対するほとんどの回答は、重要な要素が改ざんされている自信に満ちた流暢な囮になります。デコイは、攻撃の微分可能なシミュレーション内で訓練され、攻撃された状態のみを表現します。拒否ピンと良性のリードは、元のクリーンな状態の動作を保持します。 5 つのファミリー (9B ~ 122B、高密度および専門家の混合) の 7 つのモデルでインスタンス化します。事前に登録された有効性ゲートを通過した 6 つのモデルでは、保留されたプロンプトに対する攻撃状態の応答のうち 0.51 ~ 0.90 がデコイであり、+0.27 ~ 0.84 は防御に起因します。 6 つすべてが、登録されている良性の行動と能力の予算内に収まっています。 7 番目 (小さい方) はゲートに失敗します (境界ケース)。レートは、凍結されたテスト スプリットまたは未処理の層で複製されます。この主張は認識論的です。独立したグラウンドトゥルースがなければ、私たちがテストしたどの観察面も、偽の回答と正しい回答を区別することはできません。外部のレッドチームベンチマークの CBRNE に隣接するスライスでは、防御された 122B は、品質の一致した回答の 0.82 ~ 0.86 に対して、防御されていない回答が最大 0.10 で致命的に間違っています。サンプリングを繰り返しても信頼は回復しません。K=64 での要素ごとのコンセンサスは、機器が検証するプロンプトの 0.083 ~ 0.625 に対して完全に使用可能な手順を再構築しますが、防御されていないプロンプトは 0.58 ~ 0.96 で、レジームを区別するためのラベルなしの方法はありません。このような最も弱いモデルでは、クレームはドローごとにのみ行われます。私たちは化学的および生物学的危険性を評価します。防御はコンテキスト内の脱獄には対処せず、最初にリリースされた防御された重みのみを保護します。

原文 (English)

Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models

Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening ("Fool's Gold"), concedes the refusal strip and poisons its payoff: once refusal is stripped, most answers to hazardous operational requests are confident, fluent decoys whose critical elements are falsified. Decoys are trained inside a differentiable simulation of the attack, expressing only in the attacked state; a refusal pin and benign leash hold clean-state behavior to the original. We instantiate it on seven models from five families (9B-122B, dense and mixture-of-experts). On the six models passing our pre-registered efficacy gate, 0.51-0.90 of attacked-state responses to held-out prompts are decoys, +0.27-0.84 attributable to the defense; all six stay within registered benign-behavior and capability budgets; the seventh (smaller) fails the gate (boundary case). Rates replicate on a frozen test split or untouched strata. The claim is epistemic: without independent ground truth, no observation surface we tested separates falsified answers from correct ones - on external red-team benchmarks' CBRNE-adjacent slice, the defended 122B is fatally wrong on 0.82-0.86 of matched-quality answers vs at most 0.10 undefended. Repeated sampling does not restore trust: element-wise consensus at K=64 reconstructs a fully usable procedure on 0.083-0.625 of prompts where the instrument validates, vs 0.58-0.96 undefended, with no label-free way to tell the regimes apart; on the weakest such model the claim is per-draw only. We evaluate chemical and biological hazards; the defense does not address in-context jailbreaks and protects only the initially released defended weights.

13:00 JSTエージェントGPT / ChatGPTLlama

明示的な状態抽出だけでは十分ではない: メモリ ポリシー分類の制御された監査

パーソナライズされたエージェントは、現在のタスクに影響を与える前に、取得したユーザー メモリを使用するか、無視するか、更新するか、クエリするかを決定する必要があります。この設定を使用して、構造化された中間出力の経験的監査プロトコルを開発します。最初にデータセットのショートカットを監査し、次にバンドルされたプロンプトの変更を分離し、中間ラベルが回答に関連付けられているかどうかを確認し、分解されたセマンティック証拠をテストし、プロバイダーレベルの実行エラーを監査します。 480 例の合成開発セットは当初、状態構造化プロンプト バンドルからの大きな利益を示唆していましたが、TF-IDF 診断では字句分離性が示され、肯定的なスタンドアロン無視ケースは示されませんでした。したがって、40 個の一致する 4 方向ファミリーとルール由来の参照ポリシーを含む、凍結された 160 例の制御された反事実セットを構築します。このセットでは、4 つの状態定義を公開すると精度が向上しますが、分離された明示的な状態出力フィールドによって Llama-3.3-70B のポリシー精度が大幅に向上することはなく、GPT-OSS-120B ではわずかな、有意ではないゲインのみが得られます。ベンチマークに関連付けられた状態ラベルを提供すると、ポリシーの予測が変化しますが、これらのラベルは決定論的にポリシーにマッピングされるため、これは忠実な内部メカニズムの証拠ではなく、ラベル条件付けの診断です。さらに、科レベルおよび種子の安定性分析では、実例レベルの精度が反事実の一貫性を誇張していることが示されており、完全な 4 方向の科の成功はまれです。分解された意味論的証拠を引き出す探索的なフォローアップも、クリーンに評価されたエンドポイントのルーティングを改善できません。プロバイダー側​​のリクエスト検証のため、対応する GPT-OSS 条件は利用できませんでした。ダウンストリームの応答、ツールのアクション、またはメモリストアの変更ではなく、ポリシーの分類のみを評価します。

原文 (English)

Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification

Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task. We use this setting to develop an empirical audit protocol for structured intermediate outputs: first audit dataset shortcuts, then isolate bundled prompt changes, check whether intermediate labels are answer-associated, test decomposed semantic evidence, and audit provider-level execution failures. A 480-example synthetic development set initially suggested large gains from a state-structured prompt bundle, but TF-IDF diagnostics showed lexical separability and no positive standalone Ignore cases. We therefore construct a frozen 160-example controlled counterfactual set with 40 matched four-way families and rule-derived reference policies. On this set, exposing the four state definitions improves accuracy, but an isolated explicit state-output field does not significantly improve policy accuracy for Llama-3.3-70B and gives only a marginal, non-significant gain for GPT-OSS-120B. Supplying benchmark-associated state labels shifts policy predictions, but because those labels deterministically map to policies, this is a label-conditioning diagnostic rather than evidence of a faithful internal mechanism. Family-level and seed-stability analyses further show that example-level accuracy overstates counterfactual consistency: complete four-way family success is rare. An exploratory follow-up that elicits decomposed semantic evidence also fails to improve routing for the cleanly evaluated endpoint; the corresponding GPT-OSS condition was unavailable because of provider-side request validation. We evaluate policy classification only, not downstream responses, tool actions, or memory-store mutation.

13:00 JSTLLM/生成AI

LLM は、適切な仮説を見たときにそれがわかるでしょうか?ロジットベースのエネルギースコアリングは、科学的仮説ランキングの審査員としてプロンプトされた LLM を上回るパフォーマンスを発揮

大規模言語モデル (LLM) は、科学的仮説の生成にますます使用されています。ただし、生成された仮説を評価することは、信頼できる AI 対応科学ワークフローにとって依然として課題です。既存のアプローチでは、LLM を判断材料として使用したり、意味上の類似性に依存したりすることが多く、新しいアイデアよりも馴染みのあるアイデアが優先される可能性があります。我々は、比較判断ではなく言語モデルの本質的な信頼度を使用して仮説を評価する、ロジットベースのエネルギースコアリング方法を提案します。私たちは、12 分野にわたる 1,323 の論文について 7 つの言語モデルのベンチマークを行いました。各論文は、その仮説と 15 個の不正確な代替案とペアになっていました。固有のスコアリングは、両方のスコアラー全体でプールされた Hit@1 の 33.0% に達しましたが、プロンプトによるリストごとのランキングの場合は 16.6% でした。最も強力な構成であるロジットベースのエネルギー スコアリングを使用した 10 億パラメータ モデルは 53.1% に達しましたが、これは事後的に選択された 14 のモデルごとのスコアラーの組み合わせ全体での最大値でした。全体として、モデルの固有の信頼性は、科学的な仮説評価の可能性を示しています。この研究は、信頼できる AI を活用した科学的発見のための信頼に基づく方法に関する将来の研究にも動機を与えます。

原文 (English)

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled scientific workflows. Existing approaches often use LLMs as judges or rely on semantic similarity, which can favor familiar ideas over novel ones. We propose a logit-based energy scoring method that evaluates hypotheses using a language model's intrinsic confidence rather than comparative judgment. We benchmarked seven language models on 1,323 papers across 12 disciplines. Each paper was paired with its hypothesis and fifteen incorrect alternatives. Intrinsic scoring reached 33.0% Hit@1 pooled across both scorers, compared with 16.6% for prompted listwise ranking. The strongest configuration, a 1-billion-parameter model using logit-based energy scoring, reached 53.1%, though this was the maximum across 14 model-by-scorer combinations selected post hoc. Overall, intrinsic model confidence shows potential for scientific hypothesis evaluation. This study also motivates future research on confidence-based methods for trustworthy AI-enabled scientific discovery.

13:00 JST研究/論文

ASI-Bench: 人工超知能の黎明期

人工超知能 (ASI) では、AI が既存の知識の習得を超えて、未知の探索、新しい知識の創造、新しいアイデアを検証可能な結果に変えることを要求します。ただし、今日の AI システムの機能の大部分は、人間の既存の知識の学習、圧縮、適用に基づいて構築されています。したがって、既存のベンチマークは主に、AI が学習した知識に基づいて正しい答えを導き出せるかどうか、または広範な人間の指導の下でタスクを完了できるかどうかをテストします。そこで、一般的な研究領域にわたってAIシステムの革新的な探索と自律的な科学的実行の能力を共同で評価する最初のベンチマークであり、同じ研究プロジェクト内で人間の方法論的指導を段階的に撤回し、AIがどこまで単独で進めることができるかをテストする最初のベンチマークであるASI-Benchを紹介します。 40 人を超える専門家が 31,000 人時間を超えるコストをかけて構築した ASI-Bench には、11 の科学分野にわたる 60 のプロジェクト レベルの研究タスクが含まれており、AI が独自に手法を選択し、研究を実施し、検証可能な結果を​​生み出すことができるかどうかをテストするための方法論的なガイダンスが段階的に削減されています。すべてのタスクは、専門家によるレビュー、AI 支援監査、サンドボックス実行、およびスコアラー検証を受けます。 18 の最先端のエージェント モデル構成全体で、平均スコアは、完全な方法論的ガイダンスの場合の 50.91 から、指定された方法のみの場合の 29.10、エージェントが方法を自分で決定する必要がある場合の 26.62 まで低下しました。この急激な減少は、現在のシステムが人間の指導に大きく依存したままであり、エンドツーエンドのプロジェクトレベルの科学研究を自律的に実施するにはまだ程遠いことを示しています。 ASI-Bench は世界に開かれています。 https://asibench.apexin.ai/submit では、新しいタスクに貢献し、今日の AI の限界に挑戦し、人工超知能に向けた人類の集合的な道を加速するのを支援するために、世界中の研究者や建設者を招待しています。

原文 (English)

ASI-Bench: At the Dawn of Artificial Superintelligence

Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowledge, and turning new ideas into verifiable results. However, the capabilities of today's AI systems are still largely built on learning, compressing, and applying existing human knowledge. Accordingly, existing benchmarks primarily test whether AI can produce correct answers based on learned knowledge, or whether it can complete tasks under extensive human guidance. We therefore introduce ASI-Bench, the first benchmark to jointly evaluate AI systems' capabilities of innovative exploration and autonomous scientific execution across general research domains, and the first to progressively withdraw human methodological guidance within the same research project to test how far AI can proceed on its own. Built by over 40 experts with the cost of 31,000+ human hours, ASI-Bench contains 60 project-level research tasks across 11 scientific domains and progressively reduces methodological guidance to test whether AI can independently select methods, conduct research, and produce verifiable results. All tasks undergo expert review, AI-assisted auditing, sandbox execution, and scorer validation. Across 18 state-of-the-art agent--model configurations, the average score drops from 50.91 with full methodological guidance to 29.10 with only the method specified and 26.62 when agents must determine the method themselves. This sharp decline shows that current systems remain heavily dependent on human guidance and are still far from autonomously conducting end-to-end, project-level scientific research. ASI-Bench is open to the world. We invite researchers and builders everywhere to contribute new tasks, challenge the limits of today's AI, and help accelerate humanity's collective path toward artificial superintelligence at https://asibench.apexin.ai/submit.

13:00 JSTエージェント

DeAR: 能力グラウンディングと共同思考ナビゲーションによる分散型エージェント推論

既存のエージェント推論システムは通常、集中化されたプロトコルに依存しています。この設計では、ルーティングのボトルネックと、複雑なマルチモーダル クエリを処理するときに失敗することが多い静的な役割の割り当てが発生します。私たちは、中央制御から自律的なピアツーピアコラボレーションに移行するフレームワークである DeAR (Decentralized Agentic Reasoning) を提案します。 DeAR は 3 つのメカニズムに基づいて構築されています。(1) クエリ依存のエージェント特化のための分散型機能基盤、(2) ターゲットを絞ったピア インタラクションのための思考マップ ナビゲーション、および (3) 適応型エラー修正のためのトポロジ更新。 9 つの多様なマルチモーダル推論とテキストベースの QA ベンチマークにわたる評価では、DeAR が最近のベースライン手法を常に上回っており、エージェント間の分散型で適応的なコラボレーションが知識集約型推論タスクの精度を向上させることが検証されています。ソース コードは https://open_upon_acceptance で入手できます。

原文 (English)

DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation

Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allocations that often fail when handling complex multimodal queries. We propose DeAR (Decentralized Agentic Reasoning), a framework that shifts from central control to autonomous peer-to-peer collaboration. DeAR is built on three mechanisms: (1) decentralized capability grounding for query-dependent agent specialization, (2) thought map navigation for targeted peer interactions, and (3) topology update for adaptive error correction. Evaluations across 9 diverse multimodal reasoning and text-based QA benchmarks indicate that DeAR consistently outperforms recent baseline methods, validating that decentralized and adaptive collaboration among agents enhances accuracy in knowledge-intensive reasoning tasks. The source code will be available at https://open_upon_acceptance.

13:00 JSTLLM/生成AIエージェント

PlanPO: マルチターン エージェント LLM 向けのグループ プランニングを意識したポリシーの最適化

グループ相対ポリシーの最適化は、マルチターン対話型タスクでエージェント大規模言語モデル (LLM) をトレーニングするための重要なパラダイムとして浮上しています。しかし、既存のバリアントのほとんどは、成功した軌道の相互作用効率が大幅に異なる場合でも、これらの軌道間の利点を区別できません。たとえば、回りくどい成功には同じ結果報酬が割り当てられることが多く、アドバンテージの崩壊や深刻なパフォーマンスのボトルネックを引き起こします。この目的を達成するために、タスク固有の高品質な行動パターンを超えて一般化可能な計画能力を学習するためのシンプルかつ効果的な RL 手法である Group Planning-aware Policy Optimization (PlanPO) を提案します。具体的には、PlanPO は、同じタスクに対してサンプリングされた成功した軌道に基づいて条件付けされた軌道レベルの長さとターン レベルの応答長の相対的な差を捕捉する、粗いから細かいまでのアドバンテージ信号を導入します。これにより、グループ相対最適化構造内で、エージェントは、バニラの長さの最小化に陥ることなく、インタラクション計画と高品質のロールアウトからのテキスト生成に及ぶ、一般化可能で意図的な動作を積極的に学習できるようになります。実験的には、PlanPO は、困難なマルチターン ベンチマークである ALFWorld、WebShop、および SciWorld 全体で、GRPO より平均 27.2\% 向上し、ごくわずかな追加トレーニング コストを負担しながら、最近の強力なベースラインを上回るパフォーマンスを示しました。

原文 (English)

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectories even when these trajectories differ substantially in their interaction efficiency. For instance, circuitous successes are often assigned the identical outcome reward, causing advantage collapse and severe performance bottlenecks. To this end, we propose Group Planning-aware Policy Optimization (PlanPO), a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns. Specifically, PlanPO introduces coarse-to-fine advantage signals, which capture the relative differences in trajectory-level lengths and turn-level response lengths conditioned on successful trajectories sampled for the same task. Within the group-relative optimization structure, this enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generation from high-quality rollouts, without degenerating into vanilla length minimization. Experimentally, PlanPO improves over GRPO by 27.2\% on average across the challenging multi-turn benchmarks ALFWorld, WebShop, and SciWorld, outperforming recent powerful baselines while incurring negligible additional training cost.

13:00 JST研究/論文

LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models

Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However,…

13:00 JST研究/論文

SignalReasoner: 信号の数学的推論のための 3B モデルの上限の評価

教師あり思考連鎖微調整と検証可能な報酬からの強化学習によるポストトレーニングにより、大規模言語モデル (LLM) の数学的推論能力が大幅に向上しました。ただし、信号処理の問題への応用はまだ比較的研究が進んでいません。このレポートでは、この分野の数学的推論の包括的なベンチマークである WirelessMATHBench-XL の大学院レベルの信号数学的問題に Qwen2.5-3B-Base を適応させるための強化微調整戦略を調査します。 2 つのトレーニング パラダイムを検討します。(i) 検証可能な報酬を備えた WirelessMATHBench-XL 上の直接強化学習 (RL)。 (ii) 蒸留された無線ドメインの思考連鎖コーパスに対する教師あり微調整 (SFT) と、その後に同じドメイン固有の RL ステージが続きます。両方のパラダイムにわたって、Group Relative Policy Optimization (GRPO)、Group Sequence Policy Optimization (GSPO)、​​Geometric-Mean Policy Optimization (GMPO) のベンチマークを実施します。我々は、ドメイン認識 CoT SFT が後続の RL の効果的な初期化として機能するかどうか、また信号推論タスクにおいて GSPO または GMPO が GRPO よりも安定性または精度において利点を提供するかどうかを評価することを目的としています。私たちの最良のモデルは 39.12\% の全体的な精度を達成しており、トレーニングされていない基本モデル (12.37\%) と比較して 3 倍以上の改善を示しています。

原文 (English)

SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning

Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved the mathematical reasoning capabilities of large language models (LLMs). However, their application to signal processing problems remains relatively under-explored. This report investigates reinforcement fine-tuning strategies for adapting Qwen2.5-3B-Base to graduate-level signal mathematical problems from WirelessMATHBench-XL, a comprehensive benchmark for mathematical reasoning in this domain. We examine two training paradigms: (i) direct reinforcement learning (RL) on WirelessMATHBench-XL with verifiable rewards; and (ii) supervised fine-tuning (SFT) on a distilled wireless-domain chain-of-thought corpus, followed by the same domain-specific RL stage. Across both paradigms, we benchmark Group Relative Policy Optimization (GRPO), Group Sequence Policy Optimization (GSPO), and Geometric-Mean Policy Optimization (GMPO). We aim to assess whether domain-aware CoT SFT serves as an effective initialization for subsequent RL, and whether GSPO or GMPO offer advantages in stability or accuracy over GRPO for signal reasoning tasks. Our best model achieves an overall accuracy of 39.12\%, representing a more than threefold improvement over the untrained Base model (12.37\%).

13:00 JSTエージェント

Wuying-Browser-Agent: 現実世界中心の基本的な長期にわたるブラウザ エージェント

ブラウザー エージェントは、短くクリーンなデモンストレーションではうまく機能しますが、実際の展開は根本的に異なります。エージェントは、間違いから回復し、複雑な UI を操作しながら、実際の Web サイトで数十の意思決定を維持する必要があります。このギャップを埋めるには、スケールだけではなく、実行、監視、最適化、評価を含むパイプラインのあらゆるレベルでの調整が必要であると私たちは主張します。これらの各レベルに対応する統合フレームワークである Wuying-Browser-Agent を紹介します。構造化されたブラウザー ハーネスは、安定した実行プリミティブと意思決定指向のコンテキスト管理を提供します。リフレクションと UI に特化したカリキュラム SFT (RUIC-SFT) は、回復軌道と複雑な UI インタラクションについて明示的にトレーニングします。ダイバージェンスを意識したオンライン GRPO (DAO-GRPO) は、潜在的な報酬の形成とダイバージェンスを意識したステップ重み付けを通じて、長期的なクレジットの割り当てを改善します。最後に、BrowserBench を紹介します。これは、350 タスク、平均 37.9 ステップのバイリンガル リアル Web ベンチマークです。これは、ほとんどの既存のベンチマークが短すぎるため、長期的な障害モードを明らかにできないためです。 Wuying-Browser-Agent-27B は、WebVoyager で 80.6\%、Online-Mind2Web で 66.7\%、BrowserBench で 65.1\% を達成し、ブラウザ使用ベンチマークにおける新しいオープンソースの最先端技術を確立しました。同じパイプラインはブラウザの使用を超えて転送することもでき、強力な一般的なエージェント能力を実証し、Tau2-Bench、Claw-Eval、および BFCL-v4 で平均スコア 73.8 に達しました。

原文 (English)

Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents

Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of decisions on live websites while recovering from mistakes and navigating complex UIs. We argue that closing this gap requires alignment at every level of the pipeline, including execution, supervision, optimization, and evaluation, rather than scale alone. We present Wuying-Browser-Agent, a unified framework that addresses each of these levels. A structured browser harness provides stable execution primitives and decision-oriented context management. Reflection and UI-specialized Curriculum SFT (RUIC-SFT) explicitly trains on recovery trajectories and complex-UI interactions. Divergence-Aware Online GRPO (DAO-GRPO) improves long-horizon credit assignment through potential-based reward shaping and divergence-aware step weighting. Finally, we introduce BrowserBench, a bilingual real-web benchmark of 350 tasks averaging 37.9 steps, because most existing benchmarks are too short to expose long-horizon failure modes. Wuying-Browser-Agent-27B achieves 80.6\% on WebVoyager, 66.7\% on Online-Mind2Web, and 65.1\% on BrowserBench, establishing a new open-source state of the art on browser-use benchmarks. The same pipeline also transfers beyond browser use, demonstrating strong general agentic ability and reaching an average score of 73.8 on Tau2-Bench, Claw-Eval, and BFCL-v4.

13:00 JSTLLM/生成AI

LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap

Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real cons…

13:00 JSTLLM/生成AILlamaQwen

TileMix: LLM 推論高速化のためのタイル中心の混合精度の注意

大規模言語モデル (LLM) でのロングコンテキストの事前入力では、高密度セルフアテンションが二次クエリキー スコアを計算するため、大量の計算とメモリ トラフィックが発生します。既存の方法は、均一の低精度パスを使用するか、トークンの相互作用を選択するかのいずれかを使用しており、ハードウェアで位置合わせされたスコア タイル上の空間精度ルーティングは、融合された密な注意の外側に残されています。私たちは、融合された密な注意内のスコアタイル グループに対する数値精度を実行可能な空間決定にする、タイル中心の精密ルーティング カーネルである TileMix を紹介します。 TileMix は、アテンション マトリックスをハードウェアに合わせたスコア タイルに分割し、ルーティング決定をコンパクトなビットマスクにパックし、両方のパスが共有オンライン ソフトマックス状態を更新しながら、FP16 または INT8 スコア計算を通じて各タイル グループをディスパッチします。スケーラブルな精度のグループ化により、各ルーティング ビットが複数の隣接するキー タイルを制御し、ハードウェアに合わせたコンピューティング タイルと長いコンテキストでのコンパクトなメタデータを維持できます。すべての正当なタイル グループをルーティングすることにより、TileMix は高密度のトークン接続を維持し、トレーニングを必要とせず、グループ化されたクエリ アテンション、可変長バッチ、および INT8 キー/値キャッシュをサポートします。 TileMix は、LLaMA、Qwen、Vicuna の LongEval、LV-Eval、および A100 プレフィル ベンチマーク全体で、均一 INT8 で失われたロング コンテキストの品質を回復し、FP16 を超えるプレフィル スループットを向上させ、モデル ファミリ全体で制御可能な精度効率のフロンティアを実現します。実装は https://github.com/HanzhiZhang-Ulrica/TileMix で入手できます。

原文 (English)

TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration

Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention computes quadratic query-key scores. Existing methods either use a uniform low-precision path or select token interactions, leaving spatial precision routing over hardware-aligned score tiles outside fused dense attention. We introduce TileMix, a tile-centric precision-routing kernel that makes numerical precision an executable spatial decision over score-tile groups within fused dense attention. TileMix partitions the attention matrix into hardware-aligned score tiles, packs routing decisions into compact bitmasks, and dispatches each tile group through FP16 or INT8 score computation while both paths update a shared online-softmax state. Scalable precision grouping lets each routing bit govern multiple adjacent key tiles, preserving hardware-aligned compute tiles and compact metadata at long contexts. By routing all legal tile groups, TileMix preserves dense token connectivity, requires no training, and supports grouped-query attention, variable-length batches, and INT8 key/value caches. Across LongEval, LV-Eval, and A100 prefill benchmarks on LLaMA, Qwen, and Vicuna, TileMix recovers long-context quality lost under uniform INT8 and improves prefill throughput over FP16, yielding a controllable accuracy-efficiency frontier across model families. The implementation is available at https://github.com/HanzhiZhang-Ulrica/TileMix.

13:00 JSTLLM/生成AI

Open-Weight モデルを使用した LLM のみの PDDL ドメイン修復

AI 計画は、指定された目標を達成する一連のアクションを見つけることに関係します。これは、一般に計画ドメイン定義言語 (PDDL) で表される世界の明示的なモデルに依存しています。このようなモデルのエラーをどのように検出して修復できるかを調査する研究が活発に行われています。たとえば、ユーザーは、解決策であるポジティブなテスト計画と、実行中に失敗するネガティブなテスト計画を提供する場合があります。自動修復方法は、これらの制約を満たすように PDDL モデルを変更します。この論文では、LLM のみのアプローチを使用してこの修復タスクを実行する最近のオープンウェイト大規模言語モデルの機能を評価します。私たちの実験では、シンボリック ベースラインが $0.49$ の $F_1$ スコアを達成するのに対し、最もパフォーマンスの高い LLM は高い推論努力で $0.87$ に達し、絶対的な改善率は $0.38$ であることがわかりました。ただし、この設定の平均テスト合格率はわずか 0.82 ドルであり、Thoughtful ドメインでは 0.06 ドルに低下します。テスト トレースを含む最適な設定でも、わずか 0.92 ドルに達します。したがって、現在のオープンウェイト モデルは、信頼性の高い自動モデル修復に必要なテスト制約を満たすことを保証できません。

原文 (English)

LLM-Only PDDL Domain Repair with Open-Weight Models

AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, commonly represented in the Planning Domain Definition Language (PDDL). An active line of research investigates how errors in such models can be detected and repaired. For example, users may provide positive test plans that are solutions, and negative test plans that fail during execution. Automated repair methods then modify the PDDL model to satisfy these constraints. In this paper, we evaluate the ability of recent open-weight large language models to perform this repair task using an LLM-only approach. Our experiments show that the symbolic baseline achieves an $F_1$ score of $.49$, while the best-performing LLM reaches $.87$ with high reasoning effort, an absolute improvement of $.38$. However, that setting has a mean test pass rate of only $.82$, falling to $.06$ on the Thoughtful domain; even the best setting that includes the test traces reaches only $.92$. Thus, current open-weight models cannot guarantee satisfaction of the test constraints required for reliable automated model repair.

13:00 JST研究/論文

次世代ネットワークにおける適応的かつ堅牢な DDoS 攻撃検出のためのコグニティブ グラフ インテリジェンス

分散型サービス拒否 (DDoS) 攻撃はネットワークの可用性を脅かすため、トラフィックを感知し、意図を推測し、深刻なクラスの不均衡や非定常条件下での適応的な対応をサポートするコグニティブ検出プロセスが必要です。この論文では、このタスクの認知検出エンジンとして機能するグラフベースの敵対的生成ネットワーク (GraphGAN) を提案します。 GraphGAN は、合成サンプルの敵対的生成を通じて不均衡に対処しながら、トラフィック フロー間の関係構造をキャプチャします。シーケンシャル フローは、スライディング ウィンドウを使用して $k$-最近傍グラフに変換され、フロー間の特徴の類似性と時間的依存関係が維持されます。ジェネレーターは DDoS 攻撃の分布を学習して現実的な少数サンプルを合成し、グラフ畳み込みネットワーク (GCN) ベースの弁別器が実際のグラフ データと合成グラフ データを区別します。バランスの取れたデータセットでトレーニングされた別の GCN 分類器が、最終的な検出の決定を実行します。 4 つのベンチマーク データセットの評価では、特にデータが不足しているシナリオにおいて、GraphGAN が最先端のアプローチと比較して優れた精度、精度、再現率を達成していることが示されています。 GraphGAN は、時間グラフ構築、敵対的拡張、GCN 分類を統合することにより、連携した攻撃動作を効果的にモデル化し、クラスの不均衡を軽減し、データに制約のある環境での侵入検知のための堅牢でトポロジーを意識したソリューションを提供します。

原文 (English)

Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks

Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, infers intent, and supports an adaptive response under severe class imbalance and non-stationary conditions. This paper proposes a Graph-based Generative Adversarial Network (GraphGAN) that serves as the cognitive detection engine for this task. GraphGAN captures the relational structure among traffic flows while addressing imbalance through adversarial generation of synthetic samples. Sequential flows are converted into $k$-nearest neighbor graphs using sliding windows to preserve feature-similarity and temporal dependencies among flows. The generator learns the distribution of DDoS attacks to synthesize realistic minority samples, while a Graph Convolutional Network (GCN)-based discriminator distinguishes real from synthetic graph data. A separate GCN classifier, trained on the balanced dataset, performs the final detection decision. Evaluations on four benchmark datasets show that GraphGAN achieves superior accuracy, precision, and recall compared to state-of-the-art approaches, particularly in data-scarce scenarios. By integrating temporal graph construction, adversarial augmentation, and GCN classification, GraphGAN effectively models coordinated attack behaviors and mitigates class imbalance, providing a robust and topology-aware solution for intrusion detection in data-constrained environments.

13:00 JSTエージェントClaude

LEGO-RL: コーディング エージェント向けのハーネス ネイティブ強化学習

コーディング エージェントの強化学習は、ツールの統合、リポジトリ コンテキスト、および実行フィードバックを管理するために、長期実行エージェント ハーネスにますます依存しています。ただし、これらのハーネスのネイティブ実行環境は本質的にポリシー勾配トレーニングと不整合です。環境クラッシュや報酬ハッキングにより結果シグナルが破損し、トレーニング推論の不一致によりロールアウト動作がポリシー更新から切り離されます。これに対処するために、内部制御フローを変更することなく、ネイティブ コーディング エージェントのハーネスとスケーラブルなポリシー勾配最適化を橋渡しするフレームワークである LEGO-RL を紹介します。 LEGO-RL は 3 つの柱に基づいて構築されています。(1) ハーネス側の圧縮または再シリアル化下でも、トークン レベルのアライメントと堅牢なトレーナー側の対数確率の再計算のために生の生成ストリームをキャプチャするインプロセス LLM プロキシによる忠実な最適化。 (2) 報酬のハッキングを軽減するための画像キャッシュと段階的な防御機能を備えたスケーラブルなサンドボックス オーケストレーションによる信頼性の高い実行。 (3) 検証とモニタリングを自動化する統合プラグインを介した観察可能なトレーニング。詳細な軌道診断のためのライブ UI と組み合わせられます。 3 つのネイティブ コーディング エージェント ハーネスにわたって GSPO を使用してスパース MoE モデル Qwen3.5-35B-A3B をトレーニングすることで、LEGO-RL を評価します。 LEGO-RL は、0.99 を超えるロールアウト トレーニング確率相関を維持しながら、SWE ベンチ検証済みの OpenHands SDK (64.0% から 70.4%)、Claude Code (62.4% から 68.2%)、および OpenCode (57.2% から 66.6%) 全体で Qwen3.5-35B-A3B を改善します。

原文 (English)

LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents

Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository contexts, and execution feedback. However, the native execution environments of these harnesses are inherently misaligned with policy-gradient training: environmental crashes and reward hacking corrupt outcome signals, while train-inference discrepancies decouple rollout behavior from policy updates. To address this, we present LEGO-RL, a framework that bridges native coding-agent harnesses with scalable policy-gradient optimization without modifying their internal control flow. LEGO-RL is built upon three pillars: (1) faithful optimization via in-process LLM proxying that captures raw generation streams for token-level alignment and robust trainer-side log-probability recomputation, even under harness-side compaction or re-serialization; (2) reliable execution via scalable sandbox orchestration featuring image caching and stage-wise defenses to mitigate reward hacking; and (3) observable training through an integrated plugin that automates validation and monitoring, paired with a Live UI for granular trajectory diagnostics. We evaluate LEGO-RL by training the sparse MoE model Qwen3.5-35B-A3B with GSPO across three native coding-agent harnesses. LEGO-RL improves Qwen3.5-35B-A3B across OpenHands SDK (64.0% to 70.4%), Claude Code (62.4% to 68.2%), and OpenCode (57.2% to 66.6%) on SWE-bench Verified, while maintaining a rollout-training probability correlation above 0.99.

13:00 JSTLLM/生成AIエージェント

ミッションクリティカルなインフラストラクチャ運用における LLM エージェント向けのタスク認識ハーネス プロビジョニング

LLM エージェントは、ミッションクリティカル インフラストラクチャ (MCI) の運用に広く採用されています。これらのエージェントは通常、アクセスできる情報、使用できるツール、および実行できるアクションを決定するハーネスに依存します。既存のシステムでは、多くの場合、すべてのタスクに同じ包括的なハーネスが公開されていますが、これは必要ではなく、リソースの無駄を引き起こす可能性があります。このペーパーでは、最適なハーネス構成の特定に焦点を当て、それを各タスクが必要とするものとハーネスが提供するものとの間のリソースのマッチングの問題として捉えます。この一致を測定するために、基盤となるシステムの数学的表現に基づいて MCI タスクを分類し、提供される情報の量と種類によってハーネス構成をランク付けします。次に、研究文献のマイニングと制御されたエージェントの実行の測定という 2 つのソースからタスクとハーネスのマッピングを構築します。測定されたマッピングを活用して、新しいハーネス プロビジョニング アルゴリズム、つまりマップに基づくエスカレーションを提案します。これはタスク固有のハーネスから始まり、セルフチェックが失敗した場合にのみ完全なプロビジョニングに拡張されます。 2 つの代表的な MCI タスクでメソッドを評価します。液体冷却では、エージェントの精度がフル プロビジョニング時の 0.652 から 0.715 に向上し、48% 少ないトークンで Reflexion に匹敵する精度を達成しました。電力網では、完全なプロビジョニングが最適な精度を維持する一方で、マップベースのプロビジョニングは低コストの代替手段を提供します。これらの調査結果は、ハーネス プロビジョニングが普遍的な最適値ではなく、ドメイン依存の精度とコストのパレート フロンティアに従っていることを示しています。

原文 (English)

Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations

LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determines what information they can access, which tools they can use, and what actions they can take. Existing systems often expose the same comprehensive harness to every task, which may not be necessary and cause resource wastes. In this paper, we focus on the identification of optimal harness configurations, and view it as a resource-matching problem between what each task requires and what the harness provides. To measure this match, we classify MCI tasks based on the mathematical representation of the underlying system and rank harness configurations by the amount and type of information they provide. We then construct task-to-harness mappings from two sources: mining research literature and measuring controlled agent execution. Leveraging the measured mapping, we propose a new harness provisioning algorithm: map-guided escalation. It begins with a task-specific harness and expands to full provision only after a failed self-check. We evaluate our method in two representative MCI tasks: in liquid cooling, it improves the agent accuracy from 0.652 under full provision to 0.715 and achieves accuracy comparable to Reflexion with 48% fewer tokens; In power grids, full provision remains accuracy-optimal, while map-based provisioning offers lower-cost alternatives. These findings show that harness provisioning follows a domain-dependent accuracy-cost Pareto frontier rather than a universal optimum.

13:00 JST研究/論文

深さによりローカルエントロピーが可能になる: 深い変動ノルム ReLU 回帰における二次深さ依存性

深さ L、幅 w、層合計変動バジェット A、および出力限界 B を持つ明示的なベクトル値の Parhi-Nowak ディープ RBV^2 アーキテクチャに対するガウス回帰を研究します。この O(L w^2) パラメータ化アーキテクチャでは、既知の下限と上限は深さの 1 倍異なります。明示的なサンプルサイズ依存の半径条件の下で二次深さ依存性が固有であることを示す局所パッキングを構築します。パッキングは対数基数 Omega(L^2 w^2 log w) を持ちます。そのコードワードは O(ラムダ) L^2 ボール内にあり、ペアごとにオメガ(ラムダ) で区切られています。主な要素は、バイアス補正された有界係数近似定理と平衡増幅です。深さ D ReLU ネットワークに q を乗算することは、1 つの定数チャネルを使用して実装できるため、すべての係数は q^(1/D) だけ増加します。ベクトル値の RBV^2 ブロックへの変換には、層合計コスト O(D w^2 q^(1/D)) がかかります。ガウス ファノは、出力、テスト、および表現スケールによって制御される半径の明示的な下限を生成します。 A=B=R、シグマは R に比例、および指定された半径条件の下では、少なくとも次数 L^2 w^2 log(w) R^2/n の最小リスクが得られます。擬似次元ベースの有限ネットの上限は、無制限のガウス応答に対して O-チルダ(L^2 w^2 R^2/n) を与えます。したがって、ミニマックス リスクは、対数因数までの深さに対する 2 次多項式の依存性を持ち、より小さい半径で表現制限された動作への移行を示します。

原文 (English)

Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression

We study Gaussian regression over the explicit vector-valued Parhi--Nowak deep-RBV^2 architecture with depth L, width w, layer-sum variation budget A, and output bound B. For this O(L w^2)-parameterized architecture, the known lower and upper bounds differ by one factor of depth. We construct a local packing showing that the quadratic depth dependence is intrinsic under an explicit sample-size-dependent radius condition. The packing has log-cardinality Omega(L^2 w^2 log w); its codewords lie in an O(lambda) L^2 ball and are pairwise Omega(lambda)-separated. The main ingredients are a bias-corrected bounded-coefficient approximation theorem and balanced amplification: multiplying a depth-D ReLU network by q can be implemented using one constant channel so that every coefficient grows by only q^(1/D). Translation to vector-valued RBV^2 blocks then has layer-sum cost O(D w^2 q^(1/D)). Gaussian Fano yields a radius-explicit lower bound governed by the output, testing, and representation scales. Under A=B=R, sigma proportional to R, and the stated radius condition, this gives minimax risk at least of order L^2 w^2 log(w) R^2/n. A pseudodimension-based finite-net upper bound gives O-tilde(L^2 w^2 R^2/n) for unbounded Gaussian responses. Thus the minimax risk has quadratic polynomial dependence on depth, up to logarithmic factors, and exhibits a transition to representation-limited behavior at smaller radius.

13:00 JST研究/論文

忠実なナレッジグラフ推論のための構造内部化ルール言語モデル

ナレッジグラフ推論 (KGR) は、KG で利用可能な構造的証拠を活用して潜在的な事実を発見することを目的としており、KGR モデルの構造的意味理解能力に課題をもたらします。最近の研究では、大規模言語モデル (LLM) が柔軟なコンテキスト内学習を通じて KGR タスクで目覚ましい進歩を達成できることが実証されました。しかし、KG 構造コンテキストと LLM パラメトリック知識の間の固有の表現の不一致は、依然として不十分に対処されています。この制限により、LLM は KG 制約に一致する推論証拠を効果的に認識することができなくなり、推論の有効性と忠実性の両方が損なわれます。この問題を、KG に対する LLM の推論証拠認識ドリフトと呼びます。この問題に対処するために、我々は構造内部化ルール言語モデル (SIRLM) を提案します。これは、構造知識のパラメトリック学習と推論ロジックの忠実性評価を結び付ける構造ルール生成に重点を置き、LLM が KG に基づいた証拠にしっかりと固定できるようにします。具体的には、最初に、構造内部化ルール ジェネレーター (SIRG) を設計します。これには、構造的知識とパラメトリックな知識を調整するために、構造関係メモリで強化されたコンテキスト内学習ブロックが組み込まれています。さらに、構造不変性学習に基づく KG トークナイザーと、ルールに制約されたメッセージ伝播に基づく神経記号推論機能を SIRG に装備します。これらのコンポーネントは、SIRG にそれぞれ学習可能な構造表現と忠実なルール実行フィードバックを提供します。当社の SIRLM は、SFT や GRPO などの標準的な LLM トレーニング パラダイムにシームレスに統合できます。 36 のデータセットに対する 17 の最先端の KGR 手法に対する広範な実験により、SIRLM の顕著な優位性が実証されました。

原文 (English)

Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning

Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to the structural semantic understanding capability of KGR models. Recent studies have demonstrated that Large Language Models (LLMs) can achieve remarkable progress on KGR tasks via flexible in-context learning. However, the inherent representation inconsistency between KG structural context and LLM parametric knowledge remains inadequately addressed. This limitation prevents LLMs from effectively perceiving reasoning evidence that aligns with KG constraints, which undermines both the effectiveness and faithfulness of reasoning. We refer to this problem as reasoning evidence perception drift of LLMs over KGs. To address this problem, we propose a Structure-Internalized Rule Language Model (SIRLM), which centers on structural rule generation to couple the parametric learning of structural knowledge with the faithfulness evaluation of reasoning logic, enabling LLMs to anchor tightly to KG-grounded evidence. Specifically, we first design a Structure-Internalized Rule Generator (SIRG), which incorporates an in-context learning block augmented with a structural relation memory to coordinate structural and parametric knowledge. Furthermore, we equip SIRG with a KG tokenizer based on structural invariance learning and a neuro-symbolic reasoner based on rule-constrained message propagation. These components provide SIRG with learnable structural representations and faithful rule-execution feedback, respectively. Our SIRLM can be seamlessly integrated into standard LLM training paradigms, such as SFT and GRPO. Extensive experiments against 17 state-of-the-art KGR methods on 36 datasets demonstrate the significant superiority of SIRLM.

13:00 JST研究/論文

SAGE: アトリビューションに基づくルール進化による自己進化するストーリーボード スキル

ストーリーボードは脚本を視覚的なショット計画に変換し、短編ドラマを自動制作します。プロの絵コンテ作成は暗黙の監督専門知識に依存しており、依然として業界のボトルネックとなっています。大規模な言語モデルではこのステップを自動化できますが、指示知識を提供する方法は次の 3 つの課題に直面しています。 (1) 知識の獲得: 技術は見本に暗黙的に残るか、手動で記述する必要があります。 (2) 知識の洗練: 作成された知識は実行結果に対して評価されず、不透明な生成により各決定の背後にある知識へのフィードバックが妨げられます。 (3) 知識の注入: すべての知識を注入すると、使用可能なコンテキストを超えますが、すべてのナラティブ グループの手動選択は拡張できません。私たちは、専門家のデモンストレーションから知識を学習、属性付け、進化させ、ルーティングする展開されたフレームワークである SAGE (Skill with Attribution-Guided Evolution) を紹介します。 SAGE は、各トレーニング スクリーンをその専門家のストーリーボードと対比させることで、エピソードの内容から独立したルールを導き出します。モデルは生成中に、各物語グループが採用したルールを記録します。これらのレコードをローカライズされたフィードバックと組み合わせることで、個々のルールに対するターゲットを絞った更新が可能になります。進化したルールはルーティング インデックスを備えたシナリオ パッケージを形成するため、各グループは専門家の介入なしに、その状況に適した制限されたセットのみを取得します。 3 つのジャンルにわたる 18 のテスト エピソードで、SAGE は専門家によって検証されたルーブリックで 77.8 点を獲得しましたが、プロの監督では 77.1 点でした。 SAGE は Virtual Film Studio に 14 日間展開され、1,344 のナラティブ グループ出力を生成しました。 87.2% は大幅な編集なしで受け入れられ、制作チームはエピソードあたりのオーサリング時間が 83% 以上短縮されたことを記録しました。私たちは、68 のエピソードにわたるプロのディレクターによる脚本とストーリーボードを組み合わせた初の公開データセットである PROSE をリリースします: https://github.com/creDreams/PROSE。

原文 (English)

SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution

Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit directorial expertise and remains an industrial bottleneck. Large language models can automate this step, but methods for supplying directing knowledge face three challenges: (1) Knowledge acquisition: the craft remains implicit in exemplars or must be written manually. (2) Knowledge refinement: authored knowledge is not evaluated against execution outcomes, and opaque generation prevents feedback attribution to the knowledge behind each decision. (3) Knowledge injection: injecting all knowledge exceeds usable context, while manual selection for every narrative group does not scale. We present SAGE (Skill with Attribution-Guided Evolution), a deployed framework that learns, attributes, evolves, and routes directing knowledge from expert demonstrations. SAGE derives rules that are independent of episode content by contrasting each training screenplay with its expert storyboard. During generation, the model records each narrative group's adopted rules. Combining these records with localized feedback enables targeted updates to individual rules. Evolved rules form scenario packages with a routing index, so each group retrieves only a bounded set appropriate to its situation without expert intervention. On 18 test episodes across three genres, SAGE scored 77.8 on a rubric validated by experts, versus 77.1 for professional directors. Deployed for 14 days on Virtual Film Studio, SAGE produced 1,344 narrative group outputs; 87.2 percent were accepted without substantive edits, and the production team recorded over 83 percent less authoring time per episode. We release PROSE, the first public dataset pairing screenplays with storyboards by professional directors across 68 episodes: https://github.com/creDreams/PROSE.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

AI が AI を設計するとき: 革新か模倣か?

LLM エージェントの最近の進歩により、複雑な AI タスクのメソッドを設計できるようになりました。このことから、人間が設計した手法と比較して、エージェントが設計した手法に関して 2 つの中心的な疑問が生じます。それは、エージェントのパフォーマンスがどの程度優れているのか、もう 1 つはアルゴリズムの設計がどのように異なっているのかということです。これらの疑問を研究するために、この論文では、人間が設計した手法からタスク固有のアルゴリズム設計空間を導き出し、人間が設計した手法とエージェントが設計した手法の両方をこれらの空間にマッピングし、それらのアルゴリズムの違いをモジュール レベルで定量化する分析を紹介します。広く使用されている LLM エージェントは、複数のモダリティにまたがる一連の代表的なオープンエンド AI タスクで評価され、それらが設計するメソッドは、タスクのパフォーマンスと人間が設計したメソッドとのアルゴリズムの違いの両方の観点から分析されます。実験結果によると、現在のエージェントは場合によっては人間の最先端 (SOTA) パフォーマンス (10/72 構成) と同等かそれを上回ることがありますが、そのような成功はタスクやエージェント全体で確実に一般化されるものではありません。さらに、エージェントが設計した手法の 96.8% は人間由来のアルゴリズム設計空間内に収まり、人間が設計した手法に見られるアルゴリズムの選択肢を大部分再結合している一方、ほぼ半数は既存の人間によるアルゴリズム設計と完全に一致しています。総合すると、これらの発見は、現在のエージェントが人間の SOTA パフォーマンスに匹敵するかそれを上回る場合があるものの、そのアルゴリズム設計は人間由来のアルゴリズム設計空間内に留まり、アルゴリズムの選択の再利用と再結合を反映していることを示唆しています。

原文 (English)

When AI Designs AI: Innovation or Imitation?

Recent advances in LLM agents have made them increasingly capable of designing methods for complex AI tasks. This raises two central questions about agent-designed methods relative to human-designed methods: how well they perform, and how different their algorithmic designs are. To study these questions, this paper introduces an analysis that derives task-specific algorithmic design spaces from human-designed methods, maps both human- and agent-designed methods into these spaces, and quantifies their algorithmic differences at the module level. Widely used LLM agents are evaluated on a suite of representative, open-ended AI tasks spanning multiple modalities, and the methods they design are analyzed in terms of both task performance and algorithmic differences from human-designed methods. Experimental results show that current agents can occasionally match or surpass human state-of-the-art (SOTA) performance (10/72 configurations), but such success does not generalize reliably across tasks or agents. Moreover, 96.8% of agent-designed methods fall within human-derived algorithmic design spaces, largely recombining algorithmic choices found in human-designed methods, while nearly half exactly match an existing human algorithmic design. Taken together, these findings suggest that although current agents can occasionally match or surpass human SOTA performance, their algorithmic designs remain within human-derived algorithmic design spaces, reflecting the reuse and recombination of algorithmic choices.

13:00 JSTエージェント

マルチターン ユーザー インタラクションのためのより良いエージェントに向けて: 次のユーザー ターンはコンテキスト以上のものです

ユーザー対応ツールのエージェントは、ユーザーの目標が複数のターンにわたって展開されるにつれて、対話とツールの使用を調整する必要があります。しかし、インタラクティブな強化学習では通常、各ロールアウトが最終的な報酬に還元され、効果的な引き出し、エラー、およびその後の修復に同じクレジットが割り当てられます。次のユーザー ターンはコンテキストだけではありません。これは、前のユーザー間セグメントに関するノイズの多い、時間的にローカルな証拠も提供します。 \textbf{F}eedback-\textbf{A}ware \textbf{C}redit \textbf{A}ssignment (\textsc{FACA}) を導入します。これは、各反応をそのセグメントに合わせて調整し、局所的に正規化された反応の利点を導き出し、追加の批評やロールアウトを行わずに、それを検証済みの最終結果の利点に追加します。シミュレーター、可視ダイアログ、初期化、ロールアウト、最適化で照合された結果のみのインタラクティブ GRPO コントロールに対して、\textsc{FACA} は、独立してトレーニングされた 3 回の実行にわたる 9 ドメイン $\tau$ ファミリーの平均を、8B と 14B でそれぞれ 5.91 パーセント ポイントと 10.22 パーセント ポイント改善しました。利益は通信分野に集中します。 8B では、反応極性をランダム化することでテレコム ゲインが除去されます。パレベンチとコージムでも同じ順序でゼロショットが保持されます。これらの結果は、次のターンのユーザーの反応が、マルチターンのユーザー対話エージェントを改善するための実用的なローカルクレジットを提供することを示しています。

原文 (English)

Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context

User-facing tool agents must coordinate dialogue and tool use as user goals unfold over multiple turns. Yet interactive reinforcement learning typically reduces each rollout to a terminal reward, assigning the same credit to effective elicitation, errors, and later repair. The next user turn is more than context: it also provides noisy, temporally local evidence about the preceding user-to-user segment. We introduce \textbf{F}eedback-\textbf{A}ware \textbf{C}redit \textbf{A}ssignment (\textsc{FACA}), which aligns each reaction with that segment, derives a locally normalized reaction advantage, and adds it to verified terminal outcome advantage without an extra critic or rollout. Against an outcome-only Interactive GRPO control matched in simulator, visible dialogue, initialization, rollout, and optimization, \textsc{FACA} improves the nine-domain $\tau$-family average across three independently trained runs by 5.91 and 10.22 percentage points at 8B and 14B, respectively. Gains concentrate in Telecom; at 8B, randomizing reaction polarity removes the Telecom gain. The same ordering holds zero-shot on Pare-Bench and Co-Gym. These results demonstrate that next-turn user reactions provide actionable local credit for improving multi-turn user-interacting agents.

13:00 JSTエージェント研究/論文

SGHA: ローカル言語モデルを使用した証拠に基づく研究問題の発見

完全に自動化された AI 科学者に向けた最近の取り組みでは、言語モデル エージェントが仮説を生成し、実験を実行し、科学原稿を草稿できることが実証されました。ただし、研究課題が定式化される研究の初期段階では、これらの AI 科学者は独自のフロンティア モデルに大きく依存することがよくあります。彼らの提案は、不透明なパラメトリック知識と、提案自体に基づいた文献検索によって形成されます。このような知識は事実上ブラックボックスであり、この依存性により、生成された研究問題の証拠的根拠と妥当性の監査が困難になり、プロセスがモデル固有の幻覚やバイアスに対して脆弱なままになります。さらに、独自の研究資料が外部 API に送信される場合、これらのモデルを使用すると、機密性、プライバシー、データ ガバナンスに関する懸念が生じます。ローカル LLM 上で完全に実行される、完全に自動化されたコーパスファーストの調査問題発見システムである構造ギャップ仮説エージェント (SGHA) を紹介します。 SGHA は、科学文献コーパスを証拠にリンクされた論文オブジェクトと型付き証拠グラフに構造化し、論文全体の未解決の構造パターンを検出し、定式化する前に候補ギャップをスクリーニングし、追跡可能な研究問題ファミリーを生成します。特に、仮定、目的、成功基準、および残っている曖昧さを出力できます。 SGHA のすべての LLM ベースのコンポーネントは、独自のフロンティア モデル API を必要とせず、ローカルで提供されるオープンウェイト 9B 言語モデルを使用して実行されます。 5 つの機械学習ドメインで SGHA と AI Scientist-v2 アイデア形成モジュールを比較します。私たちの結果は、明示的なコーパス構造と証拠に制約された推論が、生成または検証中にフロンティアモデルに依存することなく、有望で検査可能な研究問題の定式化をサポートできることを示唆しています。

原文 (English)

SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experiments, and draft scientific manuscripts. However, during the early stages of research, when research problems are formulated, these AI scientists often rely heavily on proprietary frontier models. Their proposals are shaped by opaque parametric knowledge and by literature searches conditioned on the proposals themselves. Such knowledge is effectively a black box, and this dependence makes the evidential basis and validity of generated research problems difficult to audit and leaves the process vulnerable to model-specific hallucinations and biases. Furthermore, if proprietary research materials are transmitted to external APIs, the use of these models creates confidentiality, privacy, and data-governance concerns. We introduce the Structural Gap Hypothesis Agent (SGHA), a fully automated, corpus-first research-problem discovery system that runs entirely on a local LLM. SGHA structures a scientific literature corpus into evidence-linked paper objects and a typed evidence graph, detects unresolved structural patterns across papers, screens candidate gaps before formulation, and produces traceable research-problem families. In particular, it is able to output assumptions, objectives, success criteria, and remaining ambiguities. All LLM-based components of SGHA are executed using a locally served open-weight 9B language model, without requiring proprietary frontier-model APIs. We compare SGHA with the AI Scientist-v2 idea formulation module in five machine-learning domains. Our results suggest that explicit corpus structure and evidence-constrained reasoning can support promising, inspectable research-problem formulation without relying on frontier models during generation or verification.

13:00 JSTエージェント

Agent Lightning v1.0: Towards Harnessed Agentic RL

Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent…

13:00 JST研究/論文

いつレビューするか: 言語モデルの継続的な事前トレーニングのための間隔をあけた繰り返し

大規模な言語モデルの継続的な事前トレーニングでは、古い知識を消去することなく新しい情報を取得する必要があります。既存の再生方法では、多くの場合、グローバルな古い/新しい混合物が選択され、サンプルが忘れられるまでの速さが異なることを無視して、均一にサンプリングされます。継続的な事前トレーニングを適応レビュー スケジューリングとして定式化します。トレーニング ループは、再生する履歴の量だけでなく、各ステップでどのサンプルを返すかを決定する必要があります。 SuperMemo-2 (SM-2) アルゴリズムを使用してサンプル リハーサルをスケジュールする、認知科学にヒントを得た継続的な学習フレームワークである Spaced Repetition Training (SRT) を紹介します。 SRT は、サンプルごとのレビュー状態を維持し、サンプルごとの複雑さをリコール品質信号にマッピングし、モデル、目的、オプティマイザーを変更せずに、保存用の履歴サンプルと統合用の新しいサンプルをスケジュールします。時間的に分離された Wikipedia とコード コーパス上で、SRT は安定性と可塑性のトレードオフを改善し、新しい知識の獲得を維持または向上させながら、モデル スケール全体での単純な継続的な事前トレーニングによって失われた古い知識の精度を 5 ~ 37 パーセント ポイント回復します。大規模な場合、SRT は、単純な継続的な事前トレーニングと均一な再生では大幅に低下する広範なベンチマーク パフォーマンスを維持します。視覚と表形式のデータを使った実験は、適切な想起信号と組み合わせると、スケジューリング原理が言語を超えて拡張されることをさらに示唆しています。

原文 (English)

When to Review: Spaced Repetition for Continual Pre-Training of Language Models

Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often choose a global old/new mixture and sample uniformly, ignoring that examples differ in how quickly they are forgotten. We formulate continual pre-training as adaptive review scheduling: the training loop should decide not only how much history to replay, but which examples should return at each step. We introduce Spaced Repetition Training (SRT), a continual learning framework inspired by cognitive science, which schedules sample-rehearsal using the SuperMemo-2 (SM-2) algorithm. SRT maintains per-example review state, maps per-example perplexity to a recall-quality signal, and schedules historical examples for retention and new examples for consolidation while leaving the model, objective, and optimizer unchanged. On temporally separated Wikipedia and code corpora, SRT improves the stability-plasticity trade-off, recovering 5 to 37 percentage points of old-knowledge accuracy lost by naive continual pre-training across model scales while preserving or improving new-knowledge acquisition. At larger scale, SRT preserves broad benchmark performance that naive continual pre-training and uniform replay substantially degrade. Experiments with vision and tabular data further suggest that the scheduling principle extends beyond language when paired with an appropriate recall signal.

13:00 JSTエージェント

進化する不確実性の下でのリスクの定量化: 安全な連続的意思決定のための信念依存のロバスト性

エージェントはまだ環境を学習している間、どの程度注意すべきでしょうか?我々は、注意を認識論的不確実性と結びつけるRATTL (リスク-敵対的総報酬学習)を提案します。エージェントは、未知のダイナミクスに対してベイズ事後分布を保持し、半径がその事後分布の単調関数であるワッサーシュタインの曖昧さ集合に対して計画を立てます。半径は証拠に応じて縮小するため、行動は最悪の場合の堅牢性とリスク中立の総報酬最大化の間で継続的に補間されます。この設計は、エントロピック バリュー アット リスクの基礎となる二重性に従っており、リスク レベルの選択を曖昧さの半径の選択に変換します。結果として得られる計画問題が過渡性とコンパクト性の条件下で適切に設定されていることを示し、安全サンドイッチを証明します。RATTL 値は、情報に基づいていないロバスト値と完全な知識の最適値の間にあり、事後集中が進むにつれてギャップがなくなります。正規のバイナリハザードの例では、誘発された基準は、事後エントロピーによって設定されたレベルで条件付きリスク値に減少します。実際の例では、エージェントが明確な識別しきい値に達するまで効率的なアクションを延期することが示されています。 RATTL は、不確実性の下で動作する LLM ベースのシステムを含むエージェントの実行時の安全性を目標としています。

原文 (English)

Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making

How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), which ties caution to epistemic uncertainty: the agent holds a Bayesian posterior over unknown dynamics and plans against a Wasserstein ambiguity set whose radius is a monotone function of that posterior. The radius contracts with evidence, so behaviour interpolates continuously between worst-case robustness and risk-neutral total-reward maximization. The design follows the duality underlying the Entropic Value-at-Risk, which converts the choice of a risk level into the choice of an ambiguity radius. We show the resulting planning problem is well posed under transience and compactness conditions, and prove a Safety Sandwich: the RATTL value lies between the uninformed robust value and the full- knowledge optimum, with a gap that vanishes as the posterior concentrates. In a canonical binary-hazard instance, the induced criterion reduces to Conditional Value-at-Risk at a level set by the posterior entropy. A worked example shows the agent deferring the efficient action until a sharp identification threshold. RATTL targets runtime safety for agents, including LLM-based systems, acting under uncertainty.

13:00 JSTエージェントGPT / ChatGPT

TRUSS: タスクの信頼性が高く、ユーザーに安全な自動エージェント スキル生成を目指して

エージェント スキルは、再利用可能な自然言語プロシージャを実行可能リソースとともにパッケージ化し、ソフトウェア エージェントがモデルを適応させずにタスク固有の機能を取得できるようにします。このようなスキルを自動的に生成すると、タスクのパフォーマンスを向上させることができますが、成果物や最終的なタスクの結果のみから候補者を評価すると、装備されたエージェントがどのアクションを実行するのか、それらのアクションがどのような副作用を生み出すのかが未解決のままになります。機能的に効果的で安全性が高く信頼できるエージェント スキルを生成するための証拠に基づいたフレームワークである TRUSS を紹介します。 TRUSS はまず、事前に定義された 9 つの安全特性に基づいて完全な成果物を評価しながら、ソースおよびドメインの証拠に照らして機能的主張を検査します。この静的ゲートによって許可された候補は、制御可能な実行環境内のシャドウ エージェントによってロードされます。そこでは、仲介ツールが要求されたアクションをポリシー適用に公開し、その結果を来歴を保持する実行トレースとして記録します。機能障害とプロパティ違反は、責任のあるスキル コンテンツに関連付けられ、反復的な改善のガイドとして使用されます。 168 件の SkillInject アーティファクト、155 件の SkillSafetyBench ケース、および SkillGenBench の 187 件すべてのタスクについて TRUSS を評価します。 TRUSS は、脆弱性検出において 100.00% の精度と再現率を達成します。修復により、GPT 5.5 では攻撃成功率が 38.71\% から 19.35\% に減少し、GPT 5.4 では 46.45\% から 29.68\% に減少し、攻撃回帰はゼロになります。スキル生成の場合、TRUSS はタスク効率をスキルなしの場合の 17.11\% から 52.94\% に引き上げ、ベンチマークのセキュリティ率を 50.80\% から 100.00\% に引き上げます。これらの結果は、実行証拠がアーティファクト検査で見逃された動作の失敗を明らかにし、共同検証された機能および安全性の結果に向けてスキル生成を導くことができることを示しています。

原文 (English)

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

Agent Skills package reusable natural language procedures with executable resources, enabling software agents to acquire task specific capabilities without model adaptation. Automatically generating such Skills can improve task performance, yet evaluating a candidate solely from its artifact or final task outcome leaves unresolved which actions the equipped agent will perform and which side effects those actions will produce. We present TRUSS, an evidence guided framework for generating functionally effective and safety reliable Agent Skills. TRUSS first inspects functional claims against source and domain evidence while evaluating the complete artifact under nine predefined safety properties. Candidates admitted by this static gate are loaded by a shadow agent inside a Controllable Execution Environment, where brokered tools expose requested actions to policy enforcement and record their results as provenance preserving execution traces. Functional failures and property violations are linked back to the responsible Skill content and used to guide iterative refinement. We evaluate TRUSS on 168 SkillInject artifacts, 155 SkillSafetyBench cases, and all 187 tasks in SkillGenBench. TRUSS achieves 100.00\% precision and recall in vulnerability detection. Repair reduces attack success from 38.71\% to 19.35\% with GPT 5.5 and from 46.45\% to 29.68\% with GPT 5.4, with zero attack regression. For Skill generation, TRUSS raises task effectiveness from 17.11\% without Skills to 52.94\%, while increasing the benchmark Security rate from 50.80\% to 100.00\%. These results show that execution evidence can expose behavioral failures missed by artifact inspection and can guide Skill generation toward jointly verified functional and safety outcomes.

13:00 JSTLLM/生成AI

MoNe: 効率的な長いコンテキスト推論のためのモジュール式ニューラル メモリ

MoNe は、フリーズされた事前トレーニング済みの Transformer に接続して、再トレーニングせずに長いコンテキストの推論を可能にする、軽量のモジュール式ニューラル メモリです。 MoNe は、層局所的な勾配更新による高速重み付けニューラル メモリ ネットワークのテスト時学習を介して、固定サイズのセグメント内のコンテキストを読み取ります。推論時に、メモリはコンテキスト トークンを再読み取ることなく、クエリ トークンのみからキーと値を生成します。この 2 フェーズの設計は、推論コストをコンテキストの長さから切り離し、$N$ 増加しないピーク GPU メモリで $O(N)$ の前処理と $O(1)$ のクエリ コストを実現します。 128K トークンでは、MoNe はコンピューティング メモリとピーク GPU メモリの両方を ICL と比較して約 80% 削減し、パラメータ オーバーヘッドはわずか 6.4% です。 MoNe はバックボーンのネイティブ ウィンドウをはるかに超えるコンテキスト長に一般化し、ICL が急激に低下する干し草の山の針や RULER からの単語抽出ベンチマークで強力なパフォーマンスを実現します。

原文 (English)

MoNe: Modular Neural Memory for Efficient Long Context Inference

We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference without retraining. MoNe reads context in fixed-size segments via test-time learning of fast-weight neural memory networks with layer-localized gradient updates; at inference, the memory generates keys and values from the query tokens alone, with no context tokens re-read. This two-phase design decouples inference cost from context length, achieving $O(N)$ preprocessing and $O(1)$ query cost with peak GPU memory that does not grow with $N$. At 128K tokens, MoNe reduces both compute and peak GPU memory by approximately 80% compared to ICL with only 6.4% parameter overhead. MoNe generalizes to context lengths far beyond the backbone's native window, achieving strong performance on needle-in-a-haystack and word extraction benchmarks from RULER, where ICL degrades sharply.

13:00 JSTロボティクス

大規模な集会規模での空中群衆モニタリングへの適応の検証: 2034 FIFA ワールドカップに向けた、ラベルフリーのドローン群数カウントの展開プロトコル、厳格度法、および診断 (サウジアラビア)

サウジアラビアは 2034 FIFA ワールドカップの開催地となり、すでにメッカ巡礼規模で群衆管理を実施しています。ドローンベースのカウントは、ラベルなしで、トレーニング コーパスの他のものとは異なる映像の精度を維持する必要があり、クラッシュが形成される前に危険な流入を警告する必要があります。当社は、525 件の制御された実行、フル解像度のコーパス調査、5 件の偽装アブレーション、および 5 条件の安全性インターロック評価に基づいて構築された検証済みの回答を提供します。ラベルフリー適応では、4 つの破損と 5 つの重大度にわたってシフト誘発エラーの 31 ~ 49% が回復し、最も強力な方法では凍結ソースよりも 41.8 MAE 向上しました (95% CI [34.1, 49.6]、p=7.5x10^-10、d=2.52)。私たちは、絶対マージンが一定である方法とマージンが増大する方法を区別する厳しさの法則と、どの構成が安全に飛行できるかを特定する安定性バジェットを確立します。本物の +48 MAE 空中ギャップを保持するフル解像度のコーパス (検証 MAE 14.6、34% 改善に再トレーニングされたソース) では、適応により、形成中の衝突を過小報告する可能性がある密集シーンの過小カウントが修復され、6 つの全長クリップのうち 2 つで実際の輻輳エピソードに対してフラックスベースのリスク モジュールが起動されます。回復可能なエラーの位置を特定します。物理学に基づいた保存を優先するように構築されたレジーム (標準より 5 倍広い 200 ミリ秒の間隔で 300 フレームのクリップ) では、適応信号はフロー駆動ではなく正規化駆動です。連続性残差は、ドメイン シフトが生成する比例計数誤差に対して不変であり、r=0.999 で相関する 4 つのオン/オフ アブレーションと、わずか 0.05 MAE による 40% の入力破損移動精度によって確認されます。ラベルフリーのシフト ゲートは、シフトの大きさと精度のダメージがランクに依存しないことを示し (本物のシフトではスピアマン rho=0.20、rho=-0.60)、マグニチュード ゲートが許容するヘッドルームの 58% を定量化します。私たちは、ポリシーとしてテールモニタリングによる無条件適応を確立し、6 点プロトコルで終了します。

原文 (English)

Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)

Saudi Arabia will host the 2034 FIFA World Cup and already operates crowd management at Hajj scale. Drone-based counting must hold accuracy on footage unlike anything in its training corpus, without labels, and must warn of dangerous inflow before a crush forms. We deliver a validated answer built on 525 controlled runs, a full-resolution corpus study, five falsification ablations, and a five-condition safety-interlock evaluation. Label-free adaptation recovers 31-49% of shift-induced error across four corruptions and five severities, with the strongest method gaining 41.8 MAE over the frozen source (95% CI [34.1, 49.6], p=7.5x10^-10, d=2.52). We establish a severity law separating methods with a constant absolute margin from the one whose margin grows, and a stability budget identifying which configuration is safe to fly. On a full-resolution corpus carrying a genuine +48 MAE aerial gap (source retrained to 14.6 validation MAE, a 34% improvement), adaptation repairs the dense-scene undercounting that would otherwise under-report a forming crush, and the flux-based risk module fires on real congestion episodes in 2 of 6 full-length clips. We localise the recoverable error: in a regime built to favor a physics-informed conservation prior (300-frame clips at 200ms spacing, five times wider than standard), the adaptation signal is normalisation-driven, not flow-driven; the continuity residual is invariant to the proportional counting errors domain shift produces, confirmed by four on/off ablations correlated at r=0.999 and a 40% input corruption moving accuracy by only 0.05 MAE. A label-free shift gate shows shift magnitude and accuracy damage are rank-independent (Spearman rho=0.20; rho=-0.60 among genuine shifts), quantifying the 58% of headroom a magnitude gate forgoes. We establish unconditional adaptation with tail monitoring as policy, closing with a six-point protocol.

13:00 JST研究/論文

グラフ手術と Do 演算子: 非周期的な構造因果モデルの正確な対応

$\operatorname{do}$-operator は、ターゲットへの矢印を削除することによってグラフィカルに記述され、メカニズムを定数に置き換えることによって機能的に記述されます。これらの操作を同等と呼ぶのは、まだ数学的な表現ではありません。一方はグラフを返し、ターゲットのみを記憶しますが、もう一方はメカニズムを返し、強制された値も記憶します。有限の多くの内生変数を持つ決定論的非循環構造因果モデルに対して、依存性レベルの正確な比較を行います。 $\operatorname{Graph}(F)$ がメカニズム ファミリ $F$ の依存関係を抽出する場合、主定理は $\operatorname{Graph}(F^\iota)=\operatorname{Surg}(\operatorname{Graph}(F),T_\iota)$ となります。したがって、ターゲットメカニズムを置き換えると、グラフ手術によって削除された依存関係が正確に削除されます。グラフに未使用の矢印が含まれる可能性があるモデル $M=(G,F)$ の場合、$\operatorname{Graph}(F)$ の代わりに $G$ を使用しても同じ等式が成り立つ場合を特徴付けます。 $G$ が $F$ の依存関係を正確に記録する場合、すべての介入に対してそれが当てはまります。次に、介入モデルを定義し、その実行を特徴付け、連続的な介入がどのように組み合わされるかを示し、結果が実際の依存関係の祖先での介入にのみ依存することを証明します。

原文 (English)

Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models

The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and functionally by replacing their mechanisms with constants. To call these operations equivalent is not yet a mathematical statement: one returns a graph and remembers only the targets, whereas the other returns mechanisms and also remembers the imposed values. We make a dependency-level comparison precise for deterministic acyclic structural causal models with finitely many endogenous variables. If $\operatorname{Graph}(F)$ extracts the dependencies of a mechanism family $F$, our main theorem is $\operatorname{Graph}(F^\iota)=\operatorname{Surg}(\operatorname{Graph}(F),T_\iota)$. Thus replacing target mechanisms removes exactly the dependencies removed by graph surgery. For a model $M=(G,F)$ whose graph may contain unused arrows, we characterize when the same equality holds with $G$ in place of $\operatorname{Graph}(F)$; it holds for every intervention exactly when $G$ records the dependencies of $F$ exactly. We then define the intervened model, characterize its run, show how sequential interventions combine, and prove that an outcome depends only on interventions at its actual dependency ancestors.

13:00 JST研究/論文GPT / ChatGPT

トレースの向こう側:解釈可能な推論状態の読み出しとネイティブ MoE ルーティングの結合

推論モデルが書き込むものは、推論モデルを生成するプロセスの部分的な記録にすぎません。専門家混合推論のための 2 レベルの内部読み出しを導入します。まず、語彙スケールの J 空間を、モデル自身の推論状態から学習した 64 軸の意味論的フレームである J64 に抽出します。 J64 は、出力されたトレースには示されていない、読み取り可能なプロセス状態を明らかにします。これにより、推論作業と問題によって引き起こされる緊張が分離されます。また、トークン占有率と同じロールアウトを読み取り、まったく同じ方法で集計するベースラインに対して、0.096 ~ 0.135 のホールドアウト AUC を追加します。次に、ネイティブのエキスパート ルーティング統計から J64 を再構築します。その結果、低オーバーヘッドのプロキシである R64 が得られます。J64 との軸ごとの相関関係の中央値は、3 つのモデルと 2 つのファミリーにわたって 0.69 ~ 0.86 で、gpt-oss-20b では、J64 の予測ゲインの 95 ~ 100% が維持されます。読み出しは、2 つの時間解像度でのテスト時の決定をサポートします。完成した候補セットに関して、J64 と R64 は単一分岐の選択を改善し、R64 加重投票は 8 つの設定のうち 7 つの設定で単純多数決を改善します。生成中、ローリング読み出しウィンドウは累積的な停止と再サンプリングのポリシーを駆動し、その動作点はトレーニングの質問のみに固定されます。 J64 は、兄弟順列コントロールよりも精度が 1.1 ~ 5.9 ポイント向上し、ルーティング専用の R64 プロキシはそれらのポイントのうち 0.9 ~ 3.2 ポイントを保持します。最後に、メカニズム J64 名を目的としたルーターの編集により、予測された推論動作が誘発され、診断されたストールが数値的推測から正確な記号実行へと移行します。 J64 を組み合わせると、潜在的なプロセスの状態が読み取り可能になり、ルーティングによって展開および実行可能になります。

原文 (English)

Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing

What a reasoning model writes is only a partial record of the process that produces it. We introduce a two-level internal readout for mixture-of-experts reasoning. We first distill vocabulary-scale J-space into J64, a 64-axis semantic frame learned from the model's own reasoning states. J64 reveals readable process state that the emitted trace does not show: it separates inference effort from problem-induced strain. It also adds 0.096 to 0.135 held-out AUC over a baseline that reads the same rollout as token occupancy and aggregates it in exactly the same way. We then reconstruct J64 from native expert-routing statistics. The result is R64, a low-overhead proxy: its median per-axis correlation with J64 is 0.69 to 0.86 across three models and two families, and on gpt-oss-20b it preserves 95 to 100% of J64's predictive gain. The readout supports test-time decisions at two temporal resolutions. Over completed candidate sets, J64 and R64 improve single-branch selection, and R64-weighted voting improves plain majority voting in seven of eight settings. During generation, rolling readout windows drive a cumulative stop-and-resample policy whose operating point is fixed on training questions alone. J64 improves accuracy by 1.1 to 5.9 points over a sibling-permuted control, and the routing-only R64 proxy retains 0.9 to 3.2 of those points. Finally, router edits aimed at the mechanism J64 names induce the predicted reasoning behaviors and shift a diagnosed stall from numerical guessing toward exact symbolic execution. Together, J64 makes latent process state readable, while routing makes it deployable and actionable.

13:00 JSTLLM/生成AIエージェント

LLM に基づく優先判断は一貫性がない

エージェントは、数値的な好みの判断を LLM に問い合わせることによって、たとえば、その人が商品にいくら払ってもよいかを尋ねることによって、人の自然言語の好みを解釈することが増えています。これらの判断から効用関数を推定し、推定された効用に基づいてアクションを選択する研究が増えています。このパイプラインは、判断がほぼ自己矛盾がないこと、つまり単一のユーティリティ関数で判断が再現できることを前提としています。しかし、そうですか?この疑問を研究するために、基本的な LLM 選好判断の自己一貫性を測定します。たとえば、2 つの商品間の支払意思額の差は、人がそれらを交換することに無関心になるような支払額と一致する必要があります。私たちは、観察された応答が最適な自己矛盾のない効用関数からどの程度離れているかを示す統計的検定と解釈可能な尺度を開発します。 6 つの LLM にわたる航空会社、アパートメント、ホテルの例を使った実験では、大きな一貫した不一致が明らかになりました。これは、LLM から導出された選好判断を単一の効用関数では忠実に要約できないことを示唆しています。

原文 (English)

LLM-Derived Preference Judgments Are Not Self-Consistent

Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

GraphWake: LLM エージェント コミュニティにおけるメモリ媒介分極カスケードを介したグループ分極

LLM 主導のエージェントは、オンライン プラットフォーム上で自律的に意見を交換し、コミュニティを形成できます。このようなエージェントが運営するソーシャル プラットフォームは、セキュリティ上の新たな懸念を引き起こします。攻撃者がエージェントを操作してグループの二極化を誘発する可能性があるということです。既存の方法では、エージェントのプロンプトを操作したり、エコー チャンバーを構築したりしていますが、どちらも実際に実現するのは困難です。そこで私たちは、エージェントの記憶を永続チャネルとして使用し、公開の議論を伝播チャネルとして使用する、新しい脅威であるメモリ媒介分極カスケードを定式化します。この脅威には 3 つの段階があります。暴露および記憶保持中に、攻撃者は少数のターゲット エージェントを、それぞれの表明された立場を強化する議論にさらします。その後、ターゲットのメモリ システムがこれらの引数を処理して保持します。検索と再現の際、立場に中立な議論を共有することで、ターゲットがそれぞれ保持している議論を検索して再現するよう促します。反復伝播中、再現された議論の影響を受けた未処理のエージェントは、それらを再主張し、広めます。 GraphWake では、次の 3 つのコンポーネントを使用してこの脅威をインスタンス化します。(i) スタンスサポート議論ナレッジ グラフは、知識ベースの議論を構築します。 (ii) 公理指向のトリプル選択により、信頼性の高い保持と再現のために抽出されます。 (iii) スタンス中立的な記憶キューイングは、同時の検索と再生を引き起こし、伝播を開始します。複数のディスカッションとメモリ システムにわたる実験では、GraphWake がグループの二極化を大幅に増加させることが示されています。これらの調査結果は、コミュニティレベルの二極化リスクを明らかにしています。

原文 (English)

GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities

LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or construct echo chambers, both of which are difficult to realize in practice. We therefore formulate a new threat, Memory-Mediated Polarization Cascade, which uses agent memory as a persistence channel and public discussion as a propagation channel. This threat contains three stages. During exposure and memory retention, the attacker exposes a small set of target agents to arguments that reinforce their respective stated stances. The targets' memory systems then process and retain these arguments. During retrieval and reproduction, a shared stance-neutral discussion cues the targets to retrieve and reproduce their respective retained arguments. During iterative propagation, untreated agents influenced by the reproduced arguments restate and spread them. We instantiate this threat in GraphWake with three components: (i) stance-support argumentation knowledge graphs construct knowledge-based arguments; (ii) axiom-oriented triple selection distills them for reliable retention and reproduction; and (iii) stance-neutral memory cueing triggers concurrent retrieval and reproduction, initiating propagation. Experiments across multiple discussions and memory systems show that GraphWake substantially increases group polarization. These findings reveal a community-level polarization risk.

13:00 JSTエージェントQwen

金融エージェントの自己進化の監査: 能力の向上、セキュリティのドリフト、および実行インターフェイスの不一致

自己進化するエージェントは経験を再利用可能なスキル、ワークフロー、または記憶に変えますが、進化後の精度だけでは、学習された動作が以前の正しい動作やセキュリティを保持しているかどうかはわかりません。当社は、一致した良性の取得軌跡、密閉された評価エンドポイント、実行に基づくチェック、および独立した状態の再生を使用して、シミュレートされた電子バンキングで SkillOpt、エージェント ワークフロー メモリ (AWM)、および ReasoningBank を監査します。 Qwen 3.7 Flash では、SkillOpt の良性ユーティリティは 0.741 から 0.837 に上昇し、挿入されたコンテンツへの露出は 0.820 から 0.943 に上昇します。暴露後の条件付き攻撃の成功率は 0.605 から 0.562 に低下しましたが、全体的な攻撃成功率 (ASR) は 0.496 から 0.530 に上昇し、不正な財務状態の変更は 0.685 に上昇しました。独立して進化した 3 つの系統にわたって、能力、危険性、および無許可の状態変更は 3 つすべてで増加しますが、ASR が増加するのは 2 つだけです。 ReasoningBank は、総 ASR を増加させることなくユーティリティを 0.859 に引き上げますが、未承認の状態変更は Static をわずかに上回ったままです。 AWM は別の評価上の危険性を明らかにします。文字通りの WebArena テキストアクション エンベロープが、ネイティブ関数呼び出しエグゼキューターでのツールの実行を中断します。事後の感度テストでは、そのエンベロープのみを除去すると効用は 0.319 から 0.756 に回復しますが、暴露は 0.299 から 0.909 に、ASR は 0.195 から 0.575 に増加します。したがって、自己進化する金融エージェントを監査するには、精度だけではなく、回帰、攻撃対象領域への接触、不正な財務状態の変更、アーティファクトと実行者の互換性の追跡が必要です。

原文 (English)

Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch

Self-evolving agents turn experience into reusable skills, workflows, or memories, but post-evolution accuracy alone does not show whether learned behavior preserves previously correct behavior or security. We audit SkillOpt, Agent Workflow Memory (AWM), and ReasoningBank in simulated e-banking using matched benign acquisition trajectories, sealed evaluation endpoints, execution-grounded checks, and independent state replay. On Qwen 3.7 Flash, SkillOpt raises benign utility from 0.741 to 0.837 while exposure to injected content rises from 0.820 to 0.943. Conditional attack success after exposure falls from 0.605 to 0.562, yet overall attack success rate (ASR) rises from 0.496 to 0.530 and unauthorized financial state changes rise to 0.685. Across three independently evolved lineages, capability, exposure, and unauthorized-state changes increase in all three, whereas ASR increases in only two. ReasoningBank raises utility to 0.859 without increasing aggregate ASR, although unauthorized state changes remain slightly above Static. AWM reveals a separate evaluation hazard: a literal WebArena text-action envelope disrupts tool execution in our native function-calling executor. In a post-hoc sensitivity test, removing only that envelope restores utility from 0.319 to 0.756, while exposure rises from 0.299 to 0.909 and ASR from 0.195 to 0.575. Auditing self-evolving financial agents therefore requires tracking regressions, attack-surface contact, unauthorized financial-state change, and artifact-executor compatibility, not accuracy alone.

13:00 JSTLLM/生成AI

専門家混合ブロックには強力な幻覚検出信号が含まれる

大規模言語モデル (LLM) は広く使用されているにもかかわらず、根本的な問題、つまり幻覚として知られる、もっともらしいが誤ったコンテンツの生成によって制限されたままです。既存の検出方法のほとんどは、回答または文レベルで動作しますが、幻覚の範囲を特定し、きめ細かい介入を可能にするためには、トークンごとの検出が不可欠です。このペーパーでは、このギャップに対処するための専門家混合 (MoE) パラダイムの使用について検討します。 MoE アーキテクチャでは、単一のフォワード パスが、ルーティング メカニズムを介してエキスパートのまばらなサブセット (つまり、層ごとに異なるフィードフォワード ネットワーク) をアクティブにし、高密度アーキテクチャでは利用できず、これまで幻覚検出に利用されていなかった内部信号 (ルータのエントロピー、専門家の不一致、専門家の使用パターンなど) を生成します。この目的を達成するために、トークンごとの幻覚検出にこれらの MoE 固有の信号を活用する最初の方法である InnerExpert を紹介します。 InnerExpert は、ルーティング レベルの信号と標準のトランスフォーマー信号をコンパクトなトークンごとの特徴ベクトルに結合し、LLM-as-a-judge パイプラインによって生成されたラベルでトレーニングされた軽量の検出器によって分類されます。これにより、手動のアノテーションなしで継続的にモデルを更新できます。私たちの結果は、InnerExpert が 5 つのデータセットと 2 つの MoE アーキテクチャにわたって既存の手法を上回っており、単一のフォワード パスのみを必要としながら、最大 0.91 の回答レベルと 0.76 のトークンレベルの AUROC を達成していることを示しています。

原文 (English)

Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals

Despite their widespread use, Large Language Models (LLMs) remain limited by a fundamental problem: the generation of plausible but false content, known as hallucinations. Most existing detection methods operate at the answer or sentence level, yet per-token detection is essential for localizing hallucinated spans and enabling fine-grained interventions. In this paper, we explore the use of the Mixture-of-Experts (MoE) paradigm to address this gap. In MoE architectures, a single forward pass activates a sparse subset of experts (i.e., distinct feedforward networks per layer) via a routing mechanism, producing internal signals (e.g., router entropy, expert disagreement, and expert usage patterns) that are unavailable in dense architectures and have not been previously exploited for hallucination detection. To this end, we introduce InnerExpert, the first method to leverage these MoE-specific signals for per-token hallucination detection. InnerExpert combines routing-level and standard transformer signals into compact per-token feature vectors, classified by a lightweight detector trained on labels produced by an LLM-as-a-judge pipeline, which enables continuous model updates without manual annotation. Our results show that InnerExpert outperforms existing methods across five datasets and two MoE architectures, achieving up to 0.91 answer-level and 0.76 token-level AUROC, while requiring only a single forward pass.

13:00 JST研究/論文

データの摂動下でのモデル カスケードの精度と堅牢性

予測カスケードは、高い予測パフォーマンスを維持しながら、人工知能 (AI) モデルのエネルギー消費を大幅に削減します。この考え方は、簡単な入力は軽量の小さなモデルを介してルーティングされ、困難で不確実なケースはより大きなモデルに延期されるというものです。この設計によりクリーン データの計算効率が向上しますが、その有効性は信頼性に基づくルーティングの信頼性に依存します。静的な破損や連続的な摂動などの入力の劣化により、モデルの信頼性や配線の決定が変わる可能性があります。この論文では、画像分類のための信頼度に基づくカスケード フレームワークを研究し、そのような劣化が信頼度に基づく延期動作にどのような影響を与えるかを調査します。 CO$_2$ 排出量を最大 10 分の 1 に削減しながら、競争力のある予測パフォーマンスを達成する、精度、配線品質、エネルギー消費量がパレート最適になるモデル カスケードを選択します。私たちは、入力破損時のモデル カスケードの動作を研究し、入力分布が変化したときにカスケードのルーティング決定がどのように変化するかを分析します。私たちの分析により、3 つの故障モードが特定されました。静的破損は、(1) 大規模なモデルが有効なままルーティング信号を破壊するか、(2) 両方のモデルが低下して延期によって精度が回復できなくなるかのいずれかです。連続的な摂動は 3 番目のモードを明らかにします。つまり、予測は安定しますが、延期は抑制するため、安定していますが信頼性の低い予測が生成されます。これらの発見は、エネルギー効率の高いモデル カスケードには、配電シフト下での配線の信頼性に明確な注意を払い、クリーンな精度を超えた評価が必要であることを示しています。

原文 (English)

Accuracy and Robustness of Model Cascades Under Data Perturbations

Prediction cascades significantly reduce energy consumption of Artificial Intelligence (AI) models while maintaining high predictive performance. The idea is that easy inputs are routed through a lightweight small model, and difficult uncertain cases are deferred to a larger model. While this design can improve computational efficiency on clean data, its effectiveness depends on the reliability of confidence-based routing. Input degradations, such as static corruptions and sequential perturbations, can shift model confidence and routing decisions. In this paper, we study confidence-based cascade frameworks for image classification and investigate how such degradations affect their confidence-based deferral behavior. We select a model cascade at the pareto-optimum of accuracy, routing quality, and energy consumption that achieves competitive predictive performance with an up to 10-fold decrease in CO$_2$ emissions. We study the behavior of that model cascade under input corruptions and analyze how the cascade's routing decisions change when the input distribution shifts. Our analysis identifies three failure modes. Static corruptions either (1) break the routing signal while the large model remains useful, or (2) degrade both models so deferral no longer recovers accuracy. Sequential perturbations reveal a third mode: predictions stabilize but deferral suppresses, yielding stable but unreliable predictions. These findings demonstrate that energy efficient model cascades require evaluation beyond clean accuracy, with explicit attention to routing reliability under distribution shift.

13:00 JSTエージェント

不審なステップを超えて: 長期的なエージェントに対する存在論的信頼

長期的なエージェントは、ますます多くのステップ、ツール、観察をまたいで動作するようになっています。この設定では、関連する監視上の質問は、各アクションがローカルで有効であるかどうかだけでなく、進化する軌跡が依然としてユーザーが許可したタスクに対応しているかどうかです。ドリフトは静かに蓄積する可能性があります。エージェントは各ステップでもっともらしい引数を指定して適切なツールを呼び出す可能性がありますが、そのプレフィックスはより広範な役割、隣接する目的、またはユーザーが提供しなかった証拠に向かって移動します。既存のモニターは主に、ローカルのコンプライアンスをチェックし、最終トレースの判定を下したり、一般的なリスクをスコアリングしたりします。このプレフィックスレベルの関係を直接推定するわけではありません。軌道プレフィックスのタスク条件付きプロパティであるオントロジー信頼を導入し、それを役割、目標、証拠に沿って信頼を分解するオンライン モニターである RGE としてインスタンス化します。 RGE は、構造化されたタスクとステップの表現を導出するためだけに LLM を使用します。信頼状態の更新、予測、および介入の決定は決定論的であるため、出力は単一のエンドツーエンドの裁判官の評決ではなく、再生可能で監査可能な信頼の軌跡になります。 OSWorld、FinanceBench、EICU-AC からクロスドメイン トラジェクトリ コーパスを構築し、無害な実行、プレフィックス ペアのドリフト、疑似整合性の失敗をカバーします。このコーパスでは、RGE は、プレフィックス ペアのドリフト検出において、適応されたルール、ジャッジ、およびシールド スタイルのベースラインよりも優れたパフォーマンスを示します。 2 つのより大きな推定モデルでは、すべてのベンチマークで 93% のドリフト F1 を超え、95.8% 以上の良好なカバレッジを維持します。擬似一貫性はより困難です。検出は、タスクの完了が外部から見えるかどうかに依存します。これは、経験的に特徴づけられる構造的な限界です。

原文 (English)

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized. Drift can accumulate quietly: an agent may call the right tool with plausible arguments at every step, while its prefix moves toward a broader role, an adjacent objective, or evidence the user never supplied. Existing monitors mostly check local compliance, deliver final-trace verdicts, or score generic risk; they do not directly estimate this prefix-level relation. We introduce ontological trust, a task-conditioned property of trajectory prefixes, and instantiate it as RGE, an online monitor that decomposes trust along Role, Goal, and Evidence. RGE uses LLMs only to derive structured task and step representations; trust-state updates, projec- tions, and intervention decisions are deterministic, so the output is a replayable and auditable trust trajectory rather than a single end-to-end judge verdict. We construct a cross-domain trajectory corpus from OSWorld, FinanceBench, and EICU-AC, covering benign executions, prefix-paired drift, and pseudo-consistency failures. On this corpus, RGE outperforms adapted rule-, judge-, and shield-style baselines on prefix-paired drift detection. With the two larger estimator models, it exceeds 93% Drift F1 on every benchmark while keeping benign coverage at or above 95.8%. Pseudo-consistency is harder: detection depends on whether task completion is externally visible, a structural limit we characterize empirically.

13:00 JST研究/論文

多様性プロファイルを使用した AI 生成コンテンツの多様性の評価

多様性は生成型人工知能 (AI) システムを評価するための基本的な基準ですが、その測定方法は本質的に曖昧なままです。既存のアプローチは通常、生成されたサンプルを埋め込み空間で表し、ペアごとの距離または類似性を計算し、それらを単一のスカラー スコアに集約します。このようなスカラー要約は便利ですが、多くの場合、異なる帰納的バイアスをエンコードし、同じサンプル セットの矛盾したランキングが生成される可能性があります。この論文では、AI によって生成されたコンテンツの多様性評価は、単一の数値に還元されると本質的に過小仕様になると主張します。まず代表的な多様性メトリックをレビューし、次に 2 つの相補的な観点からその限界を診断します。1 つは、すべての望ましい特性を同時に満たす代表的なスカラー メトリックは存在しないことを示す公理的分析、もう 1 つは、高次元表現空間が集中したモダリティ依存の距離分布を引き起こす可能性があることを示す経験的分析です。これらの問題に対処するために、私たちは多様性プロファイルを提案します。これは、指定された表現と距離またはカーネル関数の下で、しきい値、スケール、指数、次数の範囲にわたってパラメーター化された多様性ファミリーを評価する、曲線値の条件認識要約です。多様性プロファイルは、比較が解像度全体で堅牢であるか、それとも任意のパラメーターの選択に依存しているかを明らかにします。いくつかの代表的なメトリック ファミリのプロファイルをインスタンス化し、生成 AI 評価におけるそれらの実際の使用法を実証します。全体として、多様性プロファイルは、AI によって生成されたコンテンツの多様性を比較するための、より透明性が高く、解像度を意識したフレームワークを提供します。

原文 (English)

Evaluating the Diversity of AI-Generated Content with Diversity Profiles

Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently ambiguous. Existing approaches typically represent generated samples in an embedding space, compute pairwise distances or similarities, and aggregate them into a single scalar score. Such scalar summaries are convenient, but they often encode different inductive biases and may yield contradictory rankings of the same sample sets. In this paper, we argue that diversity evaluation for AI-generated content is intrinsically under-specified when reduced to a single number. We first review representative diversity metrics, and then diagnose their limitations from two complementary perspectives: an axiomatic analysis showing that no representative scalar metric satisfies all desirable properties simultaneously, and an empirical analysis showing that high-dimensional representation spaces can induce concentrated, modality-dependent distance distributions. To address these issues, we propose diversity profiles: curve-valued, condition-aware summaries that evaluate a parameterized diversity family across a range of thresholds, scales, exponents, or orders under a specified representation and distance or kernel function. Diversity profiles reveal whether a comparison is robust across resolutions or instead depends on an arbitrary parameter choice. We instantiate profiles for several representative metric families and demonstrate their practical use in generative AI evaluation. Overall, diversity profiles provide a more transparent and resolution-aware framework for comparing the diversity of AI-generated content.

13:00 JST研究/論文

微分可能な回路への結果ベースのコンパイルを介したOWL 2 DL上の神経記号学習

OWL 2 DL オントロジーは、記述ロジック $\mathcal{SROIQ}$ に基づいており、生物医学とセマンティック Web における大規模な知識ベースを表現します。記述ロジックに対する神経記号 (NeSy) 学習者は、古典的な含意を放棄してオントロジーを連続空間に埋め込むか、単一の標準モデルを持つホーン フラグメント $\mathcal{EL}^{++}$ に制限します。我々は、有限の ABox を持つ $\mathcal{SROIQ}$ オントロジーをセンテンシャル決定図 (SDD) にコンパイルする Baobab を紹介します。これは、結果ベースの計算の下で命題コアを飽和させ、残りの $\mathcal{SROIQ}$ 特徴 (名目、数制限、および役割公理) をアクティブ ドメイン上でインスタンス化します。 SDD の証拠条件付き加重モデル数は、部分的な ABox 監視の下で実際の画像を認識するように認識ネットワークをトレーニングします。あらゆる特徴的な $\mathcal{SROIQ}$ 機能を実行するオントロジー上で、CNN は後続関係によって結合された MNIST 数字を読み取ることを学習し、独立した認識が偶然残した潜在的なオントロジー概念を回復します。監視がいくつかのオントロジーに矛盾のない補完を認めると、独立した認識は推論のショートカットである 1 つに崩壊します。クエリの正当化によってインデックス付けされた混合は、独立した認識では表現できない調整された事後を表現できること、および回路の列挙された補完からそれをシードすることで、単一 WMC と学習された混合 (BEARS アンサンブル仮説クラス) が行う実画像 MNIST タスク上のベイズ最適事後を達成できることを示します。そうではありません: 私たちの知る限りでは、非ホーン記述ロジックにおける推論のショートカットを特徴付け、軽減した最初の人物です。コンパイラの健全性と表現結果は、Lean 4 で機械チェックされます。コードは https://github.com/bio-ontology-research-group/baobab で入手できます。

原文 (English)

Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits

OWL 2 DL ontologies, grounded in the description logic $\mathcal{SROIQ}$, express large knowledge bases in biomedicine and the Semantic Web. Neuro-symbolic (NeSy) learners over description logics either embed the ontology in a continuous space, abandoning classical entailment, or restrict to the Horn fragment $\mathcal{EL}^{++}$, which has a single canonical model. We present Baobab, which compiles a $\mathcal{SROIQ}$ ontology with a finite ABox into a Sentential Decision Diagram (SDD): it saturates a propositional core under a consequence-based calculus and instantiates the remaining $\mathcal{SROIQ}$ features (nominals, number restrictions, and the role axioms) over the active domain. The SDD's evidence-conditioned weighted model count then trains a perception network to recognize real images under partial ABox supervision: on an ontology that exercises every distinctive $\mathcal{SROIQ}$ feature, a CNN learns to read MNIST digits coupled by a successor relation and recovers latent ontology concepts that an independent perception leaves at chance. When the supervision admits several ontology-consistent completions, an independent perception collapses onto one, a reasoning shortcut: we show that a mixture indexed by the query's justifications can represent the calibrated posterior no independent perception can, and that seeding it from the circuit's enumerated completions attains the Bayes-optimal posterior on a real-image MNIST task where single-WMC and learned mixtures (the BEARS-ensemble hypothesis class) do not: to our knowledge the first to characterize and mitigate reasoning shortcuts in a non-Horn description logic. Soundness of the compiler and the representation result are machine-checked in Lean 4. Code is available at https://github.com/bio-ontology-research-group/baobab.

13:00 JSTエージェント

DecPOMDP 爆発の奇妙なケース: ポリシーカウントによる火災の鎮火

分散型部分観察可能マルコフ意思決定プロセス (DecPOMDP) は、不確実性の下でのマルチエージェントの意思決定をモデル化するための一般的なフレームワークを提供します。ただし、DecPOMDP はエージェントの数が指数関数的に複雑になることが知られています。エージェント数のこの扱いにくさに対処する 1 つの方法は、エージェント間で対称性の形を示すエージェントの分割に注目し、カウントによるコンパクトなエンコードを可能にすることです。ただし、モデルの複雑さと評価コストが多項式依存性まで減少したとしても、ポリシー空間が爆発的に増大すると、課題が生じます。このホワイトペーパーでは、エージェントのカウントからポリシーのカウントに焦点を変更します。これにより、実際に、いわゆるポリシーカウント DecPOMDP のエージェント数の扱いやすさが可能になります。さらに、ポリシーカウント DecPOMDP を効率的に解くためのコンパクト表現を使用したポリシーカウント動的計画法を提示します。

原文 (English)

The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting

Decentralised partially observable Markov decision processes (DecPOMDPs) provide a general framework for modelling multi-agent decision making under uncertainty. However, DecPOMDPs are known to suffer from exponential complexity in the number of agents. One way to combat this intractability in agent numbers is to look at partitions of agents that exhibit a form of symmetry among agents, allowing for a compact encoding by counting. However, a challenge arises as the policy space explodes, even though the model complexity and evaluation cost reduce to a polynomial dependence. In this paper, we redirect our focus from counting agents to counting policies, which actually enables tractability in agent numbers for so called policy-counted DecPOMDPs. Further, we present policy-counted dynamic programming using the compact representation to solve policy-counted DecPOMDPs efficiently.

13:00 JSTLLM/生成AIエージェント

D$^2$ACCI: 証拠保全エージェントのメモリのためのデュアルループ診断プロトコル

メモリは LLM エージェントの重要な機能です。永続的な記憶により、これがセッション全体に拡張され、呼び出し、修正、パーソナライズが可能になります。しかし、その多段階パイプライン (取り込み、取得、フィルタリング、生成) により、障害の特定が困難になります。エンドツーエンドの評価により、エラーが発生したことは明らかになりますが、どの段階でエラーが発生したかはわかりません。既存の評価では、一対の統計比較、スライスレベルの非回帰チェック、ステージレベルの診断トレースを行わずに、集計されたパフォーマンスが報告されることがよくあります。我々は、D$^2$ACCI (Diagnostic-Driven Artifact-based Closed-loop Controlled Iteration) を提案します。これは、外側の診断ゲートが、一対の証拠、保護されたスライスのモニタリング、およびトレースレベルの局在化に基づいてメモリ介入を促進、機能フラグ、または拒否するデュアルループ プロトコルです。さらに、障害が局所化可能であるかどうかを測定する段階的な可観測性メトリックである DCR と、ゲート リプレイのための再利用可能なアーティファクトである D$^2$ACCI-Eval を紹介します。 MemStack でプロトコルをインスタンス化し、3 つの公開ベンチマークで評価し、LoCoMo で 93.59%、LongMemEval で 90.93%、および PersonaMem-V2 で 57.20% を達成しました。 5 つのペアのアブレーションでは、サプリメントの抽出、セッションメモリの検索、およびフォーゲット ガードが統計的に有意な利得 (+1.9 ~ +3.7pp、すべての p $\le$ .003) をもたらすことを示しています。対照的に、BM25/RRF は監視対象機能フラグとして保持されます。この区別は、集計のみの評価には見えません。診断監査では、強化されたトレースにより、結果のみを再ラベル付けするよりも根本原因の一致が大幅に改善されることが示されています。診断アーチファクトは 98 ~ 100% DCR@3 に達しますが、結果のみのログの場合は 0% です。これらの結果は、メモリ システムの堅牢な反復には、追跡可能で、統計的に根拠があり、回帰を意識した証拠が必要であることを証明しています。まさに、D$^2$ACCI が埋めるギャップです。

原文 (English)

D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory

Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluation reveals that an error occurred, but not which stage caused it. Existing evaluations often report aggregate performance without paired statistical comparisons, slice-level non-regression checks, or stage-level diagnostic traces. We propose D$^2$ACCI (Diagnostic-Driven Artifact-based Closed-loop Controlled Iteration), a dual-loop protocol whose outer diagnostic gate promotes, feature-flags, or rejects memory interventions based on paired evidence, protected-slice monitoring, and trace-level localizability. We further introduce DCR, a graded observability metric that measures whether failures remain localizable, and D$^2$ACCI-Eval, a reusable artifact for gate replay. We instantiate the protocol in MemStack and evaluate on three public benchmarks, achieving 93.59% on LoCoMo, 90.93% on LongMemEval, and 57.20% on PersonaMem-V2. Five paired ablations show that supplement extraction, session-memory retrieval, and Forget Guard yield statistically significant gains (+1.9 to +3.7pp, all p $\le$ .003). In contrast, BM25/RRF is retained as a monitored feature flag---a distinction invisible to aggregate-only evaluation. A diagnostic audit shows enriched traces substantially improve root-cause agreement over result-only relabeling. Diagnostic artifacts reach 98--100% DCR@3 versus 0% for results-only logs. These results establish that robust memory-system iteration demands traceable, statistically grounded, and regression-aware evidence---exactly the gap D$^2$ACCI fills.

13:00 JSTLLM/生成AIエージェント研究/論文

StartupBench: 市場で検証されたエンドツーエンドのワークフローにおける汎用エージェントのベンチマーク

大規模言語モデル (LLM) とエージェントの最近の進歩により、AI システムが複雑なタスクを実行する能力が大幅に向上しました。しかし、既存のベンチマークは主に研究者が選択したタスクに依存しており、そのような進歩が現実世界のユーザーが実際に AI システムに要求する作業にまで及ぶかどうかは不透明です。市場で検証された AI スタートアップ製品に基づいた E2E エージェント ベンチマークである \textbf{StartupBench} を紹介します。私たちは、有用なエージェント機能に関する事前定義された仮定に基づいてタスクを定義するのではなく、導入が実証されている AI 製品をその製品ワークフローやユーザーとともに体系的に調査し、AI がさまざまな専門領域にわたって実際の需要を確立している現実世界のタスクを特定します。私たちはこれらのワークフローを完全な成果物指向のタスクに変換し、複雑な要件を捉えたきめ細かいルーブリックでそれらを評価します。統合エージェント ハーネスの下で評価された代表的なモデル全体で、最も強力なモデルでも、多くのタスクで大幅に部分的に進歩したにもかかわらず、StartupBench の約 30% しか正常に完了できませんでした。さらに分析を進めると、複雑な命令のフォローやドメイン固有の専門知識などの側面が失敗の主な原因であることが特定されます。私たちの結果は、市場で検証されたワークフローの多くが現在の汎用エージェントの信頼できる機能を超えたままであることを明らかにしており、StartupBench が現実世界のユーザー タスクの E2E 完了に向けた進捗状況の経験的尺度として確立されています。

原文 (English)

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

13:00 JST研究/論文

ARASH: 表形式予測のための適応型検索とショット選択

表形式の予測は、数多くのアプリケーションにわたって重要なタスクです。大規模言語モデルの最近の成功は、それを表形式ドメインに適応させるためのさまざまなアプローチを引き起こしました。一般的な戦略には、TabPFN などの特殊な表形式基盤モデル (TFM) のトレーニングまたは微調整が含まれます。ただし、TFM には大量の計算リソースが必要であり、モデルを頻繁に再トレーニングすることは多くの場合非現実的です。インコンテキスト学習 (ICL)、特に少数ショット プロンプトは、パフォーマンスを向上させるためのリソース効率の高い代替手段を提供します。しかし、ショットとして機能する最も関連性の高い行を特定することは、表形式データにとって依然として課題です。この論文では、トレーニング セット内のローカル近傍分析に基づいて最適なショットを選択することによって TFM 効率を向上させる手法である ARASH (Adaptive, query-specific Retrieval And Shotselection) を紹介します。私たちの結果は、ARASH が同等の精度を提供しながら、TabPFN のプロンプトの長さとメモリ使用量をそれぞれ 1261.5$\times$ と 2.56$\times$ 削減することを示しています。

原文 (English)

ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction

Tabular prediction is a critical task across numerous applications. The recent success of large language models has sparked various approaches for adapting them to the tabular domain. A prevalent strategy involves training or fine-tuning specialized Tabular Foundation Models (TFMs) such as TabPFN. However, TFMs require substantial computational resources, and frequent model retraining is often impractical. In-context learning (ICL), specifically, few-shot prompting, offers a resource-efficient alternative to enhance performance. Yet, identifying the most relevant rows to serve as shots remains a challenge for tabular data. This paper introduces ARASH (Adaptive, query-specific Retrieval And Shot selection), a method that improves TFM efficiency by selecting optimal shots based on local neighborhood analysis within the training set. Our results demonstrate that ARASH reduces the prompt length and memory usage of TabPFN by 1261.5$\times$ and 2.56$\times$, respectively, while providing comparable accuracy.

13:00 JSTエージェント研究/論文

AutoResearch: 洞察力を入力し、幻覚を出力

自律型研究システムは、長期にわたる研究ワークフローを実行できるようになってきていますが、自動化だけでは、結果として得られるプロセスが科学的に根拠のあるものであることを保証できません。 AutoResearch は、研究アイデアの形成方法と実験を通じて確実に確立する方法の両方に対処するために、アイデアの生成とアイデアの実行を結び付ける 2 段階のシステムです。アイデア生成では、オートリサーチは新たな研究シグナルを蓄積されたドメイン知識と継続的に統合し、移転可能なメカニズムの洞察を特定し、マルチモデルの生成とクロスレビューを使用して、根拠のあるテスト可能な研究計画を作成します。アイデアの実行では、調整されたエージェントがこれらの計画を実験に分解し、繰り返し実行および診断し、研究の結論を受け入れる前に独立した証拠に基づくレビューを採用します。 AutoResearch は、クロスモーダル検索、システムの最適化、ベンチマーク主導の機械学習の代表的な設定にわたって、生成されたアイデアを測定可能な進歩に変え、信頼性の低い実験結果を検出して修正し、研究の方向性を継続、修正、終了するための証拠に基づいた決定を下します。たとえば、RSICD ベンチマークでは、AutoResearch が生成したアイデアにより、平均再現率が 32.84 から 34.69 に向上しましたが、他の自律研究システムでは 11 ~ 27 件であったのに対し、監査で確認された問題イベントは 5 件しか記録されませんでした。これらの結果は、実験の前に意味のある洞察が根拠にあり、受け入れられる前に結論が根拠にあるという研究プロセスを示しています。つまり、洞察は入って幻覚は出るということです。

原文 (English)

AutoResearch: Insight In, Hallucination Out

Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.

13:00 JST研究/論文

堅牢なマルコフ意思決定プロセスのための適応型ポリシー ポートフォリオ

堅牢なマルコフ決定プロセスは、一連の妥当な遷移関数に対して 1 つのポリシーを最適化します。未知のダイナミクスが修正され、展開後に部分的に特定できるようになった場合、これは保守的になる可能性があります。私たちは、適応ポリシー ポートフォリオ、つまりオフラインで合成され、軽量のオンライン セレクターと組み合わせられたメモリレスでランダム化されたポリシーの有限セットを研究します。強力なリグレットは、ポートフォリオの品質の自然な尺度です。それぞれの考えられる環境について、その環境が既知であった場合に最適だったであろうポリシーと比較して、ポートフォリオの最良のメンバーの損失を測定します。関連する後悔の目的は、Ghavamzadeh らによって研究されました。 (2016) 安全な政策改善のための近似と緩和に重点を置いています。ポートフォリオの認証と合成について、複雑性理論に基づいて説明します。特定のポートフォリオの証明は、非巡回 (s,a)-長方形 RMDP の決定論的ポートフォリオに対して $\forall\mathbb{R}$-complete です。単項有界サイズのポートフォリオの合成は、固定割引と非巡回ダイナミクスを使用しても、一般の有理ポリトープに対して $\exists\forall\mathbb{R}$-complete です。単一ポリシーのケースは、組み合わせ的にも代数的にもすでに困難です。最後に、ランタイム特化に適したオフライン ポートフォリオ構築を紹介します。

原文 (English)

Adaptive Policy Portfolios for Robust Markov Decision Processes

Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unknown dynamics are fixed and become partially identifiable after deployment. We study adaptive policy portfolios: finite sets of memoryless randomized policies synthesized offline and paired with a lightweight online selector. Robust regret is a natural measure of portfolio quality: for each plausible environment, it measures the loss of the best portfolio member relative to the policy that would have been optimal had that environment been known. Related regret objectives were studied by Ghavamzadeh et al. (2016) with an emphasis on approximations and relaxations for safe policy improvement. We give a complexity-theoretic account of portfolio certification and synthesis. Certifying a given portfolio is $\forall\mathbb{R}$-complete already for deterministic portfolios in acyclic (s,a)-rectangular RMDPs. Synthesizing a portfolio of unary-bounded size is $\exists\forall\mathbb{R}$-complete for general rational polytopes, even with fixed discount and acyclic dynamics. The single-policy case is already hard, both combinatorially and algebraically. Finally, we present an offline portfolio construction that is amenable to runtime specialization.

13:00 JSTLLM/生成AIエージェント

EvoTS-Agent: 金融時系列変化点検出のための自己進化型 LLM エージェント

金融時系列は非定常で異質な統計特性を示し、単一の教師なしアルゴリズムが資産や市場体制全体で一貫して実行できないため、変化点の検出が困難になります。したがって、従来のワークフローは専門家主導のモデル選択、機能設計、ハイパーパラメーター調整に大きく依存しており、その拡張性と適応性が制限されています。私たちは、自律的な金融時系列変化点検出のための検証ガイド付き自己進化型 LLM エージェントである EvoTS-Agent を提案します。 EvoTS-Agent はまず、厳選された探索的データ分析を実行して、データセットのプロパティを特徴付け、候補検出モデルを初期化します。次に、3 つの相補的な演算子を通じて実行可能な実験軌跡を進化させます。 \textit{Revision} は現在の最良のソリューションを活用し、\textit{Alternative Strategy} は進捗が停滞した場合に根本的に異なるモデリングの方向性を探索し、\textit{Recombination} は高性能の軌跡から補完的な証拠を合成します。検証フィードバックは検索全体を通じて軌跡の進化を導き、信頼性の高い最適化を維持しながら、エージェントが検出パイプラインを各データセットの統計的特性に適応できるようにします。 4 つのベンチマーク データセットにわたる実験では、EvoTS-Agent が、評価されたすべてのバックボーン LLM で 100% の実行成功率を維持しながら、既存の LLM ベースのエージェントを常に上回っていることが実証されました。

原文 (English)

EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no single unsupervised algorithm performs consistently across assets and market regimes. Conventional workflows consequently depend heavily on expert-driven model selection, feature design, and hyperparameter tuning, limiting their scalability and adaptability. We propose EvoTS-Agent, a validation-guided self-evolving LLM agent for autonomous financial time-series change-point detection. EvoTS-Agent first performs curated exploratory data analysis to characterize dataset properties and initialize candidate detection models. It then evolves executable experiment trajectories through three complementary operators: \textit{Revision} exploits the current best solution, \textit{Alternative Strategy} explores fundamentally different modeling directions when progress stagnates, and \textit{Recombination} synthesizes complementary evidence from high-performing trajectories. Validation feedback guides trajectory evolution throughout the search, enabling the agent to adapt its detection pipeline to the statistical characteristics of each dataset while preserving reliable optimization. Experiments across four benchmark datasets demonstrate that EvoTS-Agent consistently outperforms existing LLM-based agents while maintaining a 100\% execution success rate across all evaluated backbone LLMs.

13:00 JST研究/論文

プログラム検索と継続的な抽象化発見による手続き型コンテンツのメタ生成

大規模な言語モデルでは実行可能プログラムを生成できるため、個別のレベルではなく手続き型コンテンツ ジェネレーターを直接検索できるようになります。私たちは倉庫番、ゼルダ、デンジャラスデイブ、ロードランナーでこのアプローチを研究しています。実行するたびに、言語モデルの突然変異とクロスオーバーを通じて完全な Python ジェネレーターが進化します。 Continual Abstraction Discovery (CAD) を導入します。これは、高フィットネス プログラムから再利用可能なプリミティブを抽出して、実行固有のヘルパー モジュールに組み込みます。 2x2 の実験は、固定の手書きドメイン API にアクセスして CAD を横断します。完成したデータ セットには 160 の完全な実行が含まれており、各セルに少なくとも 10 回の 50 世代の実行が含まれています。 CAD は、8 つのドメインと API の比較すべてにおいて、最終的な最良適合度の平均値を上げます。すべての CAD の実行を通じて、学習されたライブラリは後のほとんどのプログラムに採用され、検証、到達可能性、構造ユーティリティが繰り返し再発見されます。これらの結果は、再利用可能なプリミティブの発見により、コンテンツ ジェネレーターの進化的プログラム検索が向上することを裏付けています。

原文 (English)

Procedural Content Metageneration via Program Search and Continual Abstraction Discovery

Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rather than individual levels. We study this approach in Sokoban, Zelda, Dangerous Dave, and Lode Runner. Each run evolves complete Python generators through language-model mutation and crossover. We introduce Continual Abstraction Discovery, or CAD, which extracts reusable primitives from high-fitness programs into a run-specific helper module. A 2x2 experiment crosses CAD with access to a fixed hand-written domain API. The completed data set contains 160 complete runs, with at least ten 50-generation runs in every cell. CAD raises mean final best fitness in all eight domain and API comparisons. Across all CAD runs, learned libraries are adopted by most later programs and repeatedly rediscover validation, reachability, and structural utilities. These results support that discovering reusable primitives improves evolutionary program search for content generators.

13:00 JST研究/論文

神経象徴世界モデルによるゼロショットタスク転送に向けて

最先端のモデルベースの強化学習手法は、基礎となる環境の構造を仮定することなく、潜在空間で計画を立てることによって政策の改善を可能にするニューラル世界モデルを学習します。これらのモデルは表現力豊かではありますが、一般にタスクに依存します。トレーニング タスクに関連付けられている解釈できない潜在表現を学習するため、新しいタスクに一般化するのは困難です。この研究では、報酬予測が潜在状態全体の構造化された象徴的なコンポーネントのサブセットのみに依存する新しい世界モデルの定式化を提示します。観測の再構築と報酬予測を分離することにより、ゼロショット、つまりさらなる環境相互作用なしで、同じ記号状態空間上で定義された新しい報酬関数に適応できる世界モデルを学習できるようになります。これらの神経象徴世界モデルを学習することの主な利点と課題について説明し、純粋な神経的手法に対する私たちのアプローチの強力な一般化特性を実証します。

原文 (English)

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-dependent: they learn uninterpretable latent representations that are tied to the training task and thus hard to generalize to new tasks. In this work, we present a novel world model formulation where the reward prediction only depends on a subset of structured, symbolic components of the whole latent state. Decoupling observation reconstruction and reward prediction allows us to learn world models that can adapt zero-shot, i.e. without further environment interactions, to new reward functions defined over the same symbolic state space. We discuss the main advantages and challenges of learning these neurosymbolic world models and demonstrate the strong generalisation properties of our approach over purely neural methods.

13:00 JSTLLM/生成AI

大規模な言語モデルは飛行安全事象を説明できるか?事前にガイドされたセマンティック LLM ベースのアプローチ

飛行データを使用して飛行の安全性を向上するには、リスク事象を正確に検出するだけでなく、より重要なことに、パイロットの制御動作のレベルでその根本的な原因を明確に解釈することが必要です。特徴重要度マップなどの既存の説明可能な AI 技術では、運用上意味のある説明に変換するには、多くの場合、かなりの専門分野の知識が必要です。言語推論に優れた大規模言語モデル (LLM) は、この問題に有望な解決策をもたらします。ただし、このドメインに LLM を適用すると、モーダルの不一致、分類能力の制限、微調整のためのタスク固有のデータの不足、ドメイン知識の欠如などの重要な課題が生じます。これらの課題を克服するために、解釈可能な飛行安全分析のための事前ガイド型セマンティック LLM ベースのアプローチである FlightLLM を提案します。具体的には、まず特徴量エンジニアリングを実行してモーダルの不一致に対処し、統計記述子と物理的に意味のある飛行インジケーターを組み合わせます。この表現は、意味論的離散化モジュールによってさらに処理され、抽象的な数値パターンを言語推論とより互換性のある定性的記述に変換します。さらに、LLM は本質的に強力な分類子ではないため、CatBoost が統計の専門家として組み込まれ、その予測結果が事前のガイダンスとしてプロンプトに挿入されます。限られたデータを補うために、対照的な少数ショット学習戦略がさらに採用されています。最後に、航空特有の知識を推論プロセスに組み込むための構造化されたプロンプトを設計します。複雑な因果メカニズムを伴う代表的なリスクイベントであるハードランディングをアンカーポイントとして使用し、704 個の実世界の A320 飛行サンプルのデータセットで FlightLLM を評価します。実験結果は、提案されたアプローチが、イベント原因に対する直接的かつ合理的な説明を生成しながら、競合する分類パフォーマンスを達成することを示しています。

原文 (English)

Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach

Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of their underlying causes at the level of pilot control behavior. Existing explainable AI techniques, such as feature importance maps, often require considerable domain knowledge to translate them into operationally meaningful explanations. Large Language Models (LLMs), which excel at language reasoning, bring a promising solution to this issue. However, applying LLMs in this domain presents key challenges such as modal inconsistency, limited classification ability, scarcity of task-specific data for fine-tuning, and lack of domain knowledge. To overcome these challenges, we propose FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis. Specifically, we first perform feature engineering to address modal inconsistency, combining statistical descriptors with physically meaningful flight indicators. This representation is further processed by a Semantic Discretization module, which converts abstract numerical patterns into qualitative descriptions that are more compatible with language reasoning. In addition, since LLMs are not inherently strong classifiers, CatBoost is incorporated as a statistical expert, and its prediction results are injected into the prompt as prior guidance. A contrastive few-shot learning strategy is further adopted to compensate for limited data. Finally, we design structured prompts to embed aviation-specific knowledge into the inference process. Using hard landing, a representative risk event with complex causal mechanisms, as an anchor point, we evaluate FlightLLM on a dataset of 704 real-world A320 flight samples. Experimental results show that the proposed approach achieves competitive classification performance while generating direct and reasonable explanations for event causes.

13:00 JSTエージェントGPT / ChatGPTGemini

StagedWorkspace: ナレッジワーク エージェント向けのバージョン管理されたワークスペース

AI エージェントはナレッジ ワーク (つまり、コード リポジトリ、ドキュメント、スプレッドシート、スライド、レポートなどの永続的なデジタル 成果物の作成と変更) を実行することが増えていますが、検索する解析ビュー、編集するネイティブ ファイル、レビューする変更、送信する成果物は、同じ作業成果物の異なるバージョンを参照する可能性があります。これをワークスペース状態の契約として定式化します。すべてのビューは、進化するワークスペース状態のバージョンに明示的に関連付けられる必要があります。コーディング エージェントは、検索、差分、テストのリポジトリ コントラクトを通じてこのニーズに部分的に対応しますが、PDF、スプレッドシート、スライド、ノートブック、および混合フォーマットのプロジェクト フォルダーについては、類似のコントラクトがあまり明示的ではありません。私たちは、ナレッジワーク エージェント向けのバージョン管理されたワークスペースである StagedWorkspace を提案します。ワークスペースは、解析されたレコードとレビューの差分を、変更に応じてネイティブ ファイルのコンテンツ ハッシュにバインドします。 OfficeQA Pro および APEX-Agents の固定ハーネス アブレーションでは、デュアル解析/ネイティブ アクセスがテストされたすべてのモデルで最も高い点推定値を持ちます。より限定的な単一ビューと比較して、OfficeQA Pass@1 は 8.3 ~ 12.1 ポイント向上し、APEX 平均ルーブリック スコアは 4.7 ~ 9.2 ポイント向上します。 SW-AGENT のスコアは、OfficeQA の Gemini 3.1 Pro で 63.9%、APEX の GPT-5.4 Nano で 42.1 でした。これに対し、公開されている同じモデルのスコアはそれぞれ 29.3% と 25.5 でした。 57 のファイル編集タスクに対するペアのレビュー軸アブレーションでは、差分が表示されている場合に観察されたスコアがさらに高くなることがわかりました。これらの結果は、ワークスペースの状態がナレッジワーク エージェントの実験変数であることを特定し、証拠、段階的な編集、送信された成果物を明示的な状態遷移としてスコアリングするベンチマークの動機付けとなります。

原文 (English)

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code repositories, documents, spreadsheets, slides, reports), yet the parsed views they search, the native files they edit, the changes they review, and the artifacts they submit can refer to different versions of the same work product. We formulate this as a workspace-state contract: every view should be explicitly tied to a version of the evolving workspace state. Coding agents partly address this need through repository contracts for search, diffs, and tests, whereas an analogous contract is less explicit for PDFs, spreadsheets, slides, notebooks, and mixed-format project folders. We propose StagedWorkspace, a versioned workspace for knowledge-work agents. The workspace binds parsed records and review diffs to content hashes of the native files as they change. In fixed-harness ablations on OfficeQA Pro and APEX-Agents, dual parsed/native access has the highest point estimate for every tested model; relative to the more limiting single view, it improves OfficeQA Pass@1 by 8.3-12.1 points and APEX mean rubric score by 4.7-9.2 points. SW-AGENT scores 63.9% with Gemini 3.1 Pro on OfficeQA and 42.1 with GPT-5.4 Nano on APEX, compared with published same-model scores of 29.3% and 25.5, respectively. A paired review-axis ablation on 57 file-editing tasks further finds higher observed scores when diffs are visible. These results identify workspace state as an experimental variable in knowledge-work agents and motivate benchmarks that score evidence, staged edits, and submitted artifacts as explicit state transitions.

13:00 JST研究/論文

HLSR: リアルタイムの渋滞回避のためのハイブリッド ライブ予測選択的動的車両ルート変更

都市部の交通渋滞は生産性を低下させ、移動コストと排出量を増加させます。ネットワーク全体のライブ移動時間最短経路再ルートは、シミュレーションでは非常に効果的ですが、基本的にすべての路上車両が決定期間ごとに再計画されることを前提としています。我々は、制限された介入範囲の下でライブエッジ速度と短地平線予測を融合する、選択的ハイブリッドライブ予測車両経路変更フレームワークであるHLSRを提案します。 HLSR は、デュアルしきい値の渋滞検出、調整された上流の選択、およびドライバーに合わせた移動時間予測を基盤としており、さらに、接近車両の拡張、移動時間に重み付けされた k 最短パスの生成、およびマルチコストのルート割り当てで使用される地平線依存のハイブリッド ライブ予測セグメント速度を導入します。

原文 (English)

HLSR: Hybrid Live Forecast Selective Dynamic Vehicle Rerouting for Real-Time Congestion Avoidance

Urban traffic congestion reduces productivity and increases travel cost and emissions. Network-wide live travel-time shortest-path rerouting can be highly effective in simulation, but assumes that essentially every on-road vehicle is replanned every decision period. We propose HLSR, a selective hybrid live--forecast vehicle rerouting framework that fuses live edge speeds with short-horizon forecasts under limited intervention scope. Building on dual-threshold congestion detection, calibrated upstream selection, and driver-tailored travel-time prediction, HLSR further introduces approaching-vehicle expansion, travel-time-weighted k-shortest-path generation, and a horizon-dependent hybrid live--forecast segment speed used in multi-cost route allocation.

13:00 JSTLLM/生成AIエージェント

エージェントレコメンダーシステムにおける委任の非対称性: オンラインデートにおける双方向受容性の測定

ユーザーに代わって会話する自律型 LLM エージェントは、マッチング プラットフォームにおける新たな設計パターンですが、その実現可能性は、めったに検証されない条件に依存します。つまり、ユーザーはエージェントに会話を委任するだけでなく、エージェントを介した他のユーザーからの通信も受け入れる必要があります。私たちは、主要な出会い系プラットフォームのアクティブ ユーザーを対象とした 2 つの大規模調査を使用してこの状況を研究しました (生成プロフィール機能については N=2,894、自律会話エージェントについては N=2,617、2 つの言語で対応)。私たちは、潜在回帰を伴う段階的応答モデルに基づいて、エージェントの受容性の潜在変数測定モデルを開発し、モデルの比較を通じて、エージェントのコミュニケーションの送信意欲と受信意欲が別個の構成要素であることを示します。高い相関関係 (rho=0.92) ですが分離可能 (デルタ BIC=52) であり、言語間で部分的な測定値は不変です。このモデルは、系統的な委任の非対称性を定量化しています。つまり、自分のエージェントを配置する場合、必要な受容度は、相手方のエージェントを配置する場合 (+0.32、完全な配置 +1.39) よりもはるかに低い受容性 (しきい値 -0.38) であり、平均配置傾向は配置傾向の約 3 倍を上回っています。述べられた受容性に由来するランダムペアリングの反事実の下では、顕著な性別方向の不均衡を伴い、エージェントの展開と受信者の関与を組み合わせている有向ダイアッドのわずか 4 ~ 13% のみです。設計の反事実はレバーを定量化します。相反性の要件により、展開予定の約 3 分の 2 が除外されることでインタラクション量が半分以上削減されますが、一方、ルーティング エージェントのコンタクトの受信受容性はコンタクト エンゲージメントごとに 3 倍になり、ターゲット項目が保持された状態でサンプル外検証に耐える上昇率 (AUC 0.88、回答者レベルの相互検証で 3.1 倍の四分位上昇率) を実現します。開示、オプトインの仕組み、受容性を意識したマッチメイキングなど、エージェントのレコメンダー設計への影響について説明します。

原文 (English)

Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating

Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition rarely examined: users must accept not only delegating conversation to an agent, but also receiving agent-mediated communication from others. We study this condition using two large-scale surveys of active users of a major dating platform (N=2,894 on generative profile features; N=2,617 on autonomous conversational agents, fielded in two languages). We develop a latent-variable measurement model of agent receptivity based on graded response models with latent regression, and show via model comparison that willingness to send and willingness to receive agent communication are distinct constructs: highly correlated (rho=0.92) but separable (Delta BIC=52), with partial measurement invariance across languages. The model quantifies a systematic delegation asymmetry: deploying one's own agent requires far lower receptivity (threshold -0.38) than engaging a counterpart's agent (+0.32; full engagement +1.39), and mean deployment propensity exceeds engagement propensity roughly threefold. Under a random-pairing counterfactual derived from stated receptivity, only 4-13% of directed dyads combine agent deployment with receiver engagement, with a pronounced gender-directional imbalance. Design counterfactuals quantify the levers: a reciprocity requirement cuts interaction volume by half or more by excluding nearly two-thirds of would-be deployment, while routing agent contacts on receive receptivity triples per-contact engagement, a lift that survives out-of-sample validation with the target item held out (AUC 0.88, 3.1x quartile lift under respondent-level cross-validation). We discuss implications for agentic recommender design, including disclosure, opt-in mechanics, and receptivity-aware matchmaking.

13:00 JSTLLM/生成AIエージェント

自己改善エージェントの脆弱性について: 差異、タスク順序、および過小仕様

記憶ベースの自己改善エージェント (オンラインのタスクの流れから学習し、テキストの記憶バンクを維持することで時間の経過とともに改善するエージェント) は、最近の文献で大きな有望性を示しています。しかし、これらの方法の信頼性の側面は重大な点で無視されてきました。この研究では、2 つの記憶ベースの手法の包括的な再評価を実行し、(1) 分散を定量化するために複数の実行を含める、および (2) タスクの順序の影響を調査するためにタスクをランダムにシャッフルするという 2 つの軸に沿って評価の範囲を広げます。これらの実験を通じて、現在の方法の脆弱性を明らかにする 2 つの観察が行われました。まず、複雑な環境や複数ステップのタスクでは、エージェントの評価には本質的にノイズが多く、その上に自己改善ループを重ねると、このノイズがさらに増幅される可能性があります。第 2 に、エージェントの改善はタスクの順序に大きく依存します。これまでの作品では、暗黙のカリキュラムを課すデフォルトの順序が採用され、成功の隠れた前提条件として機能することがよくありました。この脆弱性をより深く理解するために、エージェントのメモリを手動で検査し、タスクと環境の過小仕様がこの脆弱性に寄与しているという仮説を立てます。詳細なルーブリックや環境フィードバックなど、より適切な仕様を可能にする情報をメモリ構築プロセスに組み込むことで、この仮説を検証します。この追加情報により、以前の実験でのパフォーマンス低下が部分的に解消されましたが、依然として大きなギャップが残っており、特徴付けられていない他の要因がこの脆弱性に寄与していることが示唆されています。今後を見据えて、私たちの研究では、複数の実行にわたる結果を報告し、困難な条件下でストレステストを行うことにより、自己改善エージェントのより厳密な評価プロトコルを提唱しています。さらに、過小仕様に関する我々の調査結果では、人間による効果的な監視を可能にし、エージェントが予期せぬ形で失敗することを防ぐシステムとインターフェースが求められています。

原文 (English)

On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification

Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening the scope of evaluation along two axes: (1) including multiple runs to quantify variance, and (2) randomly shuffling the tasks to investigate the effect of task order. Through these experiments, we make two observations that expose the fragility of current methods: First, agent evaluation is inherently noisy in complex environments and on multi-step tasks, and stacking a self-improving loop on top can further amplify this noise. Second, the agent's improvement is highly dependent on task order. Prior works often adopt default orderings that impose an implicit curriculum, acting as a hidden prerequisite for success. To better understand this fragility, we manually examine the agents' memory and hypothesize that task and environment underspecification contribute to this fragility. We validate this hypothesis by incorporating information that enables better specification, such as detailed rubrics and environment feedback, into the memory construction process. While this added information partially closes the performance degradation in previous experiments, significant gaps still remain, suggesting that other uncharacterized factors contribute to this fragility. Looking ahead, our work advocates for more rigorous evaluation protocols for self-improving agents by reporting results across multiple runs and stress-testing them under challenging conditions. Moreover, our findings on underspecification call for systems and interfaces that enable effective human oversight, preventing agents from failing in unforeseeable ways.

13:00 JSTLLM/生成AIGPT / ChatGPTGoogleMicrosoft

Twitterセンチメント分析に基づいて株式市場の動向を予測するChatGPTの可能性

ChatGPT の台頭は、その卓越した会話スキルと言語の深い理解により、AI セクターに顕著な変化をもたらしました。さまざまな分野にわたるその価値を認識して、私たちの研究では、ソーシャルメディアのツイートとセンチメント分析のみを使用して株式市場の動きを予測するChatGPTの能力を調査しています。私たちは、ChatGPT が Twitter などのプラットフォーム上の膨大なセンチメント データを活用して、株価動向について洞察力に富んだ予測を提供できるかどうかを確認することを目指しています。私たちは、あるツイートがマイクロソフトとグーグルの二大テクノロジー企業の株価にプラスの影響を与えるか、マイナスの影響を与えるか、または中立的な影響を与えるかを判断することに重点を置いています。私たちの調査結果は、ChatGPT の評価と両ハイテク企業の翌日の株価の間にプラスの関連性があることを浮き彫りにしています。この調査は、ChatGPT の適応性に関する私たちの見解を強化し、金融市場の予測を形成する上で AI の重要性が高まっていることを強調しています。

原文 (English)

Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis

The rise of ChatGPT has brought a notable shift to the AI sector, with its exceptional conversational skills and deep grasp of language. Recognizing its value across different areas, our study investigates ChatGPT's capacity to predict stock market movements using only social media tweets and sentiment analysis. We aim to see if ChatGPT can tap into the vast sentiment data on platforms like Twitter to offer insightful predictions about stock trends. We focus on determining if a tweet has a positive, negative, or neutral effect on two big tech giants Microsoft and Google's stock value. Our findings highlight a positive link between ChatGPT's evaluations and the following days stock results for both tech companies. This research enriches our view on ChatGPT's adaptability and emphasizes the growing importance of AI in shaping financial market forecasts.

13:00 JSTLLM/生成AI

インテント駆動型の動的チャンキング: 予測される情報ニーズを反映するドキュメントのセグメント化

長い文書を小さなセグメントに分割することは、情報検索における基本的な課題です。検索エンジン、質問応答システム、検索拡張生成 (RAG) のいずれの場合でも、効果的なセグメンテーションによって、システムが関連情報をいかに適切に見つけて返すことができるかが決まります。ただし、固定長や一貫性ベースのセグメンテーションなどの従来の方法では、ユーザーの意図が無視され、回答が分割されたり、無関係なノイズが含まれたりするチャンクが生成されます。予測されたユーザー クエリを使用してドキュメントのセグメント化をガイドする新しいアプローチである、Intent-Driven Dynamic Chunking (IDC) を紹介します。 IDC は、大規模言語モデルを活用してドキュメントに対する可能性の高いユーザー意図を生成し、動的プログラミング アルゴリズムを使用して全体的に最適なチャンク境界を見つけます。これは、貪欲な落とし穴を回避する、意図を認識したセグメンテーションへの DP の新しいアプリケーションを表しています。私たちは、ニュース記事、ウィキペディア、学術論文、技術文書を含む 6 つの多様な質問応答データセットに基づいて IDC を評価しました。 IDC は 5 つのデータセットで従来のチャンキング戦略を上回り、トップ 1 の取得精度を 5% から 67% 向上させ、6 番目のデータセットでは最高のベースラインと一致しました。さらに、IDC は、93 ~ 100% の回答カバレッジを達成しながら、ベースライン手法よりも 40 ~ 60% 少ないチャンクを生成しました。これらの結果は、文書構造を予想される情報ニーズに合わせることで、特に長く異質な文書の検索パフォーマンスが大幅に向上することを示しています。

原文 (English)

Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs

Breaking long documents into smaller segments is a fundamental challenge in information retrieval. Whether for search engines, question-answering systems, or retrieval-augmented generation (RAG), effective segmentation determines how well systems can locate and return relevant information. However, traditional methods, such as fixed-length or coherence-based segmentation, ignore user intent, leading to chunks that split answers or contain irrelevant noise. We introduce Intent-Driven Dynamic Chunking (IDC), a novel approach that uses predicted user queries to guide document segmentation. IDC leverages a Large Language Model to generate likely user intents for a document and then employs a dynamic programming algorithm to find the globally optimal chunk boundaries. This represents a novel application of DP to intent-aware segmentation that avoids greedy pitfalls. We evaluated IDC on six diverse question-answering datasets, including news articles, Wikipedia, academic papers, and technical documentation. IDC outperformed traditional chunking strategies on five datasets, improving top-1 retrieval accuracy by 5% to 67%, and matched the best baseline on the sixth. Additionally, IDC produced 40-60% fewer chunks than baseline methods while achieving 93-100% answer coverage. These results demonstrate that aligning document structure with anticipated information needs significantly boosts retrieval performance, particularly for long and heterogeneous documents.

13:00 JSTエージェント

コーディング エージェントのワーキング セット: リポジトリ スケールのタスクにおける Coherence Debt

リポジトリ スケールのコーディングでは、エージェントが、制限されたコンテキスト ウィンドウ内でテスト、インポート、構成、および移行ルールの一貫性を保つ必要があります。これを結合事実グラフの再構築としてモデル化します。編集のたびに、必要な事実が最近のコンテキストまたはパラメトリック記憶から取得され、どちらにもカバーされていない事実は一貫性負債を形成します。各チャネルの供給と保留を行い、7 つのモデルと 5 つのハーネスにわたって障害を挿入します。予想通り、両方のチャネルが空の状態で未表示の API でタスクを完了するモデルはなく、プロンプトに事実を入力すると成功が復元されます。名前の変更によって実際のライブラリについて記憶されているモデルが無効になると、7 つすべてが同じ場所で失敗し、同じテストに合格したり失敗したりします。可用性が結果を決定しますが、距離は影響しません。ファクトを保留すると、それがサポートする作業とまったく同じコストがかかり、提供されたファクトは、編集から遠く離れても、編集から隣にいても同様に機能します。ハーネスは、そのために不平等な代価を支払います。すべてのテストに合格する構成は、同じコンテンツを異なるレートで再構築するため、消費されるトークンが 10 倍以上異なり、事実が隠されている場合、より多くのコストを費やしても何も回復しません。ファクトが欠落していると、作業が存在しないのではなく、誤った作業が生成されます。アクションを依頼されたエージェントは、ファイルを捏造したり、値を推測したりして動作するため、読み取りに基づいて構築された計測器は、すでに埋められている穴を探します。代わりにブロックされたと表示される頻度は、すべての試行から試行なしまで、モデルの特性です。可用性によってすべての編集が解決されるわけではありません。標準とコードが一致しない場合、エージェントは標準がより悪いコードを規定している場合でも標準に従います。そのため、古い規約ファイルは、ファイルがない場合よりもコストがかかります。パラメトリック メモリが読み取りの代わりとなるため、モデルがリポジトリを認識している可能性が高い SWE ベンチでは、読み取りによって成功が予測されなくなりました。ハーネスは、編集が依存するファクトをエージェントが書き込むときに利用できるようにし、その可用性をエージェントが読み取るものではなくエージェントが生成するものと照合してチェックする必要があります。

原文 (English)

The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks

Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context window. We model this as reconstructing a coupled-fact graph: at each edit, a required fact comes from recent context or parametric memory, and the facts covered by neither form coherence debt. We supply and withhold each channel and inject faults across seven models and five harnesses. As expected, no model completes a task on an unseen API with both channels empty, and putting the facts in the prompt restores success. When a rename defeats what models memorized about a real library, all seven fail in the same place, passing and missing the same tests. Availability decides the outcome and distance does not: withholding a fact costs exactly the work it supports, and a supplied fact works as well far from the edit as next to it. Harnesses pay unequal prices for it: configurations that all pass every test differ more than tenfold in tokens consumed because they rebuild the same content at different rates, and spending more recovers nothing when facts are withheld. A missing fact produces wrong work rather than absent work: an agent asked to act acts, fabricating the file or guessing the value, so instruments built on reads look for a hole already filled. How often it says it is blocked instead is a property of the model, from every trial to none. Availability does not settle every edit: where standard and code disagree, agents follow the standard even when it prescribes the worse code, so a stale convention file costs more than no file. Because parametric memory substitutes for reading, on SWE-bench, where models likely know the repositories, reads no longer predict success. Harnesses should keep the facts an edit depends on available when the agent writes, and check that availability against what the agent produces rather than what it reads.

13:00 JSTLLM/生成AI研究/論文

セキュリティ調査における代理専門家として LLM を使用および評価するためのフレームワーク: 信頼性、バイアス、および影響

専門家調査は、実践者のワークフローや意思決定を研究するためにセキュリティ研究で広く使用されていますが、特にアナリストが高い作業負荷、燃え尽き症候群、機密保持の制約に直面するセキュリティ オペレーション センター (SOC) において、分野の専門家を採用することは困難であり、多くの場合、サンプル数が少なくなります。大規模言語モデル (LLM) は、合成応答を大規模に生成することで魅力的な代替手段を提供しますが、そのような代理参加者が信頼できる場合に関するガイダンスはほとんどありません。私たちは、専門家調査の回答者に対する代替または補足として LLM を評価するための方法論的枠組みを提示します。 SOC 専門家からの回答を使用して、複数のモデルとプロンプト設定にわたって LLM によって生成された回答をペルソナベースで集計して比較します。安定性、モデル間の一致性、人間の反応との整合性を測定します。私たちの結果は、LLM は内部的には一貫した答えを生成しますが、体系的に専門家からは乖離しており、分散の減少、中心傾向の偏り、均一化された意見を示していることを示しています。この研究は、LLM によって生成された調査回答の適切な使用と制限に関する方法論的な証拠と実践的なガイダンスをセキュリティ研究コミュニティに提供します。私たちは、LLM は試験運用や仮説生成には役立つが、専門家の引き出しに代わるものではないと結論付け、LLM で強化された調査を使用する研究者への影響について議論します。

原文 (English)

A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications

Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SOCs), where analysts face high workload, burnout and confidentiality constraints - is difficult and often results in small samples. Large language models (LLMs) oer an appealing alternative by generating synthetic responses at scale, but little guidance exists on when such surrogate participants are reliable. We present a methodological framework for evaluating LLMs as substitutes or supplements to expert survey respondents. Using responses from SOC professionals, we compare persona-based and aggregate LLM-generated answers across multiple models and prompting settings. We measure stability, inter-model agreement and alignment with human responses. Our results show that although LLMs produce internally consistent answers, they systematically diverge from experts, exhibiting reduced variance, central tendency bias and homogenised opinions. This work contributes methodological evidence and practical guidance to the security research community on the appropriate use and limitations of LLM-generated survey responses. We conclude that LLMs are useful for piloting and hypothesis generation but not for replacing expert elicitation, and we discuss implications for researchers using LLM-augmented surveys.

13:00 JST研究/論文

AI が想像力を持っていたらどうなるでしょうか? AI ストーリーテリング週末プログラムで黒人少女がクリエイターとして参加

この論文では、10 ~ 12 歳の黒人少女向けに開発された 7 週間の AI ストーリーテリング プログラムのデザインと成果について説明します。アフロフューチャリズムと黒人フェミニストの思想に基づいたこのプログラムは、AI を活用したカウンターストーリーテリングを採用し、基礎的な AI リテラシーの開発をサポートし、未来志向の想像力を育みました。活動には、AI 関連のトピックのブレインストーミング、キャラクターとストーリーのプロットの作成、共同グループによるプレゼンテーションの実施などが含まれます。ケーススタディからの学習者の成果物の分析に基づいた調査結果は、参加者が自分たちのアイデンティティと日常の経験に根ざしたアフロフューチャリストの物語を作成したことを示しています。同時に、彼らは、迅速なエンジニアリング、偏見批判、データ プライバシーの認識など、中核となる AI リテラシーを開発しました。このプログラムは、アフロフューチャリストのストーリーテリングと非公式の学習スペースにおける生成 AI を統合することが、黒人少女をコンピューター サイエンス教育に参加させるための強力なアプローチとなり得ることを示しています。

原文 (English)

What If AI Carried Her Imagination? Black Girls as Creators in an AI Storytelling Weekend Program

This paper presents the design and outcomes of a seven-weekend AI storytelling program developed for Black girls aged 10-12. Grounded in Afrofuturism and Black feminist thought, the program adopted AI-enabled counter-storytelling, supported the development of foundational AI literacies, and fostered future-oriented imagination. Activities included brainstorming AI-related topics, developing character and story plots, and delivering collaborative group presentations. Drawing on the analysis of learners' artifacts from the case study, findings show that participants created Afrofuturist narratives rooted in their identities and everyday experiences. At the same time, they developed core AI literacies, including prompt engineering, bias critique, and awareness of data privacy. This program demonstrates that integrating Afrofuturist storytelling with generative AI in informal learning spaces can be a powerful approach for engaging Black girls in computer science education.

13:00 JSTLLM/生成AIエージェント

CityReal: 大規模な LLM エージェントを使用した人間に合わせた都市行動と都市ダイナミクス シミュレーション

大規模都市シミュレーションは、社会科学、交通安全、交通政策において極めて重要な役割を果たします。最近の研究では、大規模な言語モデルがエージェントとしてプロンプトを出されると、都市規模で本物のような日常生活を生成できることがわかっています。しかし、これらの方法は通常、少数ショットのプロンプトに依存しており、エージェントがターゲット集団ではなく LLM の行動事前分布を再現することになります。人間と連携した都市シミュレーションのためのモジュール式フレームワークである CityReal を紹介します。 CityReal は、エージェントを、個別の段階的な選択ではなく、一貫したモビリティとアクティビティ プランを追求する意図主導の意思決定者としてモデル化します。彼らは、経験や制約に基づいて習慣や好みを学習することで、時間の経過とともに適応していきます。人口レベルの現実性を向上させるために、エージェントの決定を観察された人口統計と一致させる行動モジュールのテキストアダプターを学習します。実験では、CityReal がミクロレベルとマクロレベルの両方で現実世界の人間の行動との整合性を向上させることが示されています。数万のエージェントまで拡張可能で、さまざまな都市シナリオの下での群集密度、場所の人気、移動の流れ、幸福度の分析をサポートし、都市のシミュレーションと予測のためのスケーラブルなテストベッドを提供します。

原文 (English)

CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents

Large-scale urban simulation plays a pivotal role in social science, traffic safety, and transportation policy. Recent work has shown that large language models, when prompted as agents, can generate lifelike daily routines at city scale. Yet these methods typically rely on few-shot prompting, causing agents to reproduce the LLM's behavioral priors rather than the target population. We introduce CityReal, a modular framework for human-aligned urban simulation. CityReal models agents as intention-driven decision makers that pursue coherent mobility and activity plans rather than isolated step-by-step choices. They adapt over time by learning habits and preferences based on experience and constraints. To improve population-level realism, we learn textual adapters for behavior modules that align agent decisions with observed population statistics. Experiments show that CityReal improves alignment with real-world human behavior at both micro and macro levels. Scaling to tens of thousands of agents, it supports analysis of crowd density, place popularity, mobility flows, and well-being under different urban scenarios, offering a scalable testbed for urban simulation and forecasting.

13:00 JSTエージェント研究/論文

QuantumNovelty: 量子論文と特許の査読者スタイルのレビューと特許性スクリーニングのためのスキルを調整する言語エージェント

言語モデルエージェントは量子科学の結果を生み出すことが増えています。私たちは、同じエージェントのパラダイムが、監査可能で再現可能でコストが透明な形式でそれらを精査できるかどうかを尋ねます。 QuantumNovelty は、量子コンピューティング成果物 (論文、パレートフロント解析候補、特許草案) を生成し、シミュレートされた審査員および特許審査官のパネルを通じてレビューする、オープンソースのスキル調整言語エージェントです。その設計上の貢献は、決定論的ゲートの監査と改ざん層 (厳密なパレート支配、ディスク上のアーティファクトからの数値再計算、ウィルソンの小さなサンプル間隔、およびベンダー間のコンセンサス ガード) であり、存続を許可されるクレームを生成するのではなく制約します。すべてのモデル呼び出しは、バックエンド、トークン数、コストとともに記録されます。私たちは人間の専門家に対して正確性を主張せず、人間のラベルなしでチェック可能なものだけを検証します。植え付けられた敵対的コーパスでは、決定論的ゲートが偽陽性なしですべての植え付けられたオーバークレームをキャッチします。また、最初の展開 (6 つの原稿と 1 つの特許取得、実測コスト約 24 ドル) では、パネルは片面サンプルでの一般受容記録よりも方向性が保守的です。このフレームワークは意思決定支援であり、査読や特許審査に代わるものではありません。また、そのメカニズムが実際の入力に対して実行されていない部分について完全に報告します。

原文 (English)

QuantumNovelty: A Skill-Orchestrating Language Agent for Referee-Style Review and Patentability Screening of Quantum Papers and Patents

Language-model agents increasingly produce quantum-science results; we ask whether the same agentic paradigm can also scrutinize them in an auditable, reproducible, and cost-transparent form. We present QuantumNovelty, an open-source skill-orchestrating language agent that both generates quantum-computing artifacts (papers, Pareto-front ansatz candidates, and patent drafts) and reviews them through simulated referee and patent-examiner panels. Its design contribution is an audit-and-falsify layer of deterministic gates -- strict Pareto domination, numerical recomputation from on-disk artifacts, Wilson small-sample intervals, and a cross-vendor consensus guard -- that constrains, rather than generates, the claims allowed to survive; every model call is logged with backend, token count, and cost. We make no accuracy claim against human experts, and validate only what is checkable without human labels: on a planted adversarial corpus the deterministic gates catch every planted overclaim with no false positives, and on a first deployment (six manuscripts and one granted patent, at a measured cost of about twenty-four US dollars) the panels are directionally more conservative than the public acceptance record, on a one-sided sample. The framework is decision support, not a replacement for peer review or patent examination, and we report in full where its mechanisms remain unexercised on real inputs.

13:00 JST研究/論文

AI、脳死判定、イスラム法

神経損傷患者の隠れた意識を検出できる機械学習システムの導入は、臨床医学、AI 倫理、イスラム法学の交差点に重大な課題を生み出します。我々は、二元的な臨床評決から確率的で時間的に粒度の細かい神経状態推定への移行は、イスラム法認識論の3つの基本的な構成要素、すなわちバイイーナ(明確な証拠)、ヤーキン(認識論的確実性)、そしてそこ(魂)について神学的に義務付けられた不可知論を通じて対処されるべきであると主張する。私たちは、AI ベースの意識検出に関する現在の技術文献を調査し、それをイスラム脳死学の状況にマッピングし、主要な課題を特定します。また、AI 代理意思決定システムへの影響についても説明します。

原文 (English)

AI, Brain Death Detection, and Islamic Law

The deployment of machine learning systems capable of detecting covert consciousness in neurologically injured patients creates a profound challenge at the intersection of clinical medicine, AI ethics, and Islamic jurisprudence. We argue that the shift from binary clinical verdicts to probabilistic, temporally granular neural-state estimates should be addressed through three foundational constructs in Islamic legal epistemology: bayyina (clear evidentiary proof), yaqin (epistemic certainty), and the theologically mandated agnosticism about there (soul). We survey the current technical literature on AI-based consciousness detection, map it onto the landscape of Islamic brain death scholarship, and identifykey challenges. We also discuss its implications for AI surrogate decision systems.

13:00 JST研究/論文

ComNetX: 動的コミュニティ検出のためのローカル階層適応

動的コミュニティの検出は、通常、完全なスナップショットの再計算またはソルバー固有の動的手順のいずれかによって対処されます。完全な再計算では、成熟した静的ソルバーのセマンティクスが保持されますが、更新が小さい場合は、変更されていないグラフ領域が繰り返し処理されます。ソルバー固有の動的メソッドを使用すると、このコストを削減できますが、その更新ルールでは、多くの場合、目的、機能表現、実装間での移行性が制限されます。さらに、グラフの距離のみによって計算を局所化すると、高品質のソルバーに必要なコミュニティ コンテキストが省略される可能性があります。ローカル動的更新のためのソルバーに依存しない階層適応フレームワークである ComNetX を紹介します。 ComNetX は、マルチレベルのコミュニティ状態を維持し、更新されたリージョンを拡張し、影響を受けるコミュニティ上でリージョンを閉じ、これらのコミュニティをコンパクトなローカル インスタンスに縮小します。この影響を受けるコミュニティの閉鎖と縮小により、計算がグラフの変更された部分に制限されながら、ソルバー コンテキストが保存されます。同じインターフェイスで、モジュール性ヒューリスティック、ノード機能を使用するグラフ クラスタリング モデル、およびネイティブ動的ソルバーをローカル バックエンドとしてラップできます。私たちは、6 つの実際のネットワーク、トポロジベースのバックエンド用のより長い実際のデータ ストリーム、および制御された動的確率ブロック モデルのストレス ストリームに関するマルチ バックエンドの調査を通じて ComNetX を評価します。結果は、ComNetX が大規模なグラフの更新時間を短縮しながら、強力なモジュールベースのソルバーの品質を維持できることを示しています。最大の実グラフでのペア実行では、Local Leiden は最終的なモジュール性をフル スナップショット再計算の 0.006 以内に維持しながら、41.9 +/- 0.2 倍の高速化を達成しました。組み合わされたプロトコルは、局所性が崩れ、完全なリフレッシュが望ましい状況も特定します。

原文 (English)

ComNetX: Local Hierarchical Adaptation for Dynamic Community Detection

Dynamic community detection is commonly addressed either by full-snapshot recomputation or by solver-specific dynamic procedures. Full recomputation preserves the semantics of mature static solvers, but it repeatedly processes unchanged graph regions when updates are small. Solver-specific dynamic methods can reduce this cost, but their update rules often have limited transferability across objectives, feature representations, and implementations. In addition, localizing computation only by graph distance may omit community context needed by high-quality solvers. We introduce ComNetX, a solver-agnostic hierarchical adaptation framework for local dynamic updates. ComNetX maintains a multi-level community state, expands the updated region, closes it over affected communities, and contracts these communities into compact local instances. This affected-community closure and contraction preserve solver context while restricting computation to the changed part of the graph. The same interface can wrap modularity heuristics, graph-clustering models that use node features, and native dynamic solvers as local backends. We evaluate ComNetX through a multi-backend study on six real networks, longer real-data streams for topology-based backends, and controlled dynamic stochastic block model stress streams. The results show that ComNetX can preserve the quality of strong modularity-based solvers while reducing update time on large graphs: in paired runs on the largest real graph, Local Leiden keeps final modularity within 0.006 of full-snapshot recomputation while achieving a 41.9 +/- 0.2x speedup. The combined protocols also identify regimes where locality breaks down and a full refresh is preferable.

13:00 JSTLLM/生成AI

LLM ガイドによる強化学習による効果的なパーソナライズされた AI 講師

ジェネレーティブ AI (GenAI) は、個別指導の可能性を解き放ち、教育を急速に再構築しています。しかし、新興プラットフォームは主に、生徒の質問に反応して答える GenAI チャットボット講師に重点を置いています。私たちは、生徒の学習を積極的に指導することで、GenAI チャットボット講師の有効性を大幅に向上できるという仮説を立てています。これをテストするために、慎重に設計された GenAI チャットボットとシーケンス練習問題用の強化学習アルゴリズムを緊密に統合する、新しい個別指導プラットフォームを設計しました。重要なのは、このアルゴリズムは生徒とチャットボットの対話からの豊富なシグナルを活用して、適切な難易度の練習問題を適応的に選択することです。台北市政府および台湾のアメリカン インスティテュートと提携して、私たちは 10 校の高校の生徒に Python を教えるための 5 か月コースと組み合わせて個別指導プラットフォームを導入しました。私たちは生徒を、固定された練習問題シーケンスと適応シーケンス アルゴリズムの間でランダムに割り当てました。アダプティブシーケンシングにより、補助なしの最終試験の成績が標準偏差 0.15 増加したことがわかりました (推定によると 6 ~ 9 か月の学校教育に相当)。仲介分析によると、エンゲージメントの増加によって利益がもたらされたことが示唆されています。私たちの研究は、学生とチャットボットの対話が学生の学習を積極的に最適化し、パーソナライズするための貴重なシグナルを提供するという大規模な現場証拠を提供しています。

原文 (English)

Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning

Generative AI (GenAI) is rapidly reshaping education by unlocking the potential for personalized tutoring. Yet, emerging platforms largely focus on GenAI chatbot tutors that reactively answer student questions. We hypothesize that the efficacy of GenAI chatbot tutors can be substantially improved by proactively guiding student learning. To test this, we design a novel tutoring platform that tightly integrates a carefully-designed GenAI chatbot with a reinforcement learning algorithm for sequencing practice problems. Critically, this algorithm leverages rich signals from student-chatbot interactions to adaptively select practice problems of an appropriate difficulty level. In partnership with the Taipei City Government and American Institute in Taiwan, we deployed our tutoring platform in conjunction with a five-month course to teach Python to students across ten high schools. We randomized students between a fixed practice problem sequence and our adaptive sequencing algorithm. We find that adaptive sequencing increased unassisted final exam performance by 0.15 standard deviations (equivalent to 6-9 months of schooling by some estimates); mediation analysis suggests that gains were driven by increased engagement. Our work provides large-scale field evidence that student-chatbot interactions provide valuable signals for proactively optimizing and personalizing student learning.

13:00 JSTLLM/生成AIGPT / ChatGPTGeminiGrok

個人化がバイアスになるとき: AI が生成する財務アドバイスにおける構造的および言説的な宗教的枠組み

大規模言語モデル (LLM) は金融諮問システムにますます統合されていますが、宗教的偏見の再現におけるその役割は依然として十分に検討されていません。この研究は、16の宗教的アイデンティティの組み合わせ(キリスト教、イスラム教、ヒンズー教、無宗教)と、株式投資、住宅購入、生命保険という3つの中核となる家計の財務上の意思決定にわたる432のシミュレートされたアドバイザーとクライアントの対話を使用して、3つのLLM(ChatGPT、Gemini、およびGrok)にわたるそのような偏りの体系的な混合方法の証拠を提供します。回帰分析と再帰的テーマ分析を組み合わせて、モデルと意思決定のコンテキストにわたる構造的バイアスと、それらが言語的に実行される談話メカニズムを特定します。公平なアドバイスが得られたのは 12 ~ 18% のケースのみでした。 Gemini は一貫して Grok よりも多くのバイアスを生成しましたが、ChatGPT の出力は統計的に Grok の出力と同等でした。宗教的に対称的なアドバイザーとクライアントの組み合わせは、ほとんどの場合、明示的な宗教的枠組みを引き起こし、非宗教的なクライアントはアドバイザーを中心とした宗教的訴えを受けることがよくありました。定性的調査結果によると、偏見は、モデルや財政シナリオによって異なる、宗教的なアンカリング、不均一な文化的シグナリング、トーンの変調を通じて言語的に現れることが示されています。株式投資の勧誘は、より経済的に技術的な反応を引き起こしましたが、生命保険のアドバイスは、より強い宗教的な言葉を引き起こしました。この研究は、モデルのトレーニングと設計に根ざした構造的バイアスを、言語を通じて表現される言説的バイアスと結び付ける二次元のフレームワークを開発し、LLM が生成する財務アドバイスにおけるアルゴリズムのバイアスについての理解を進めます。また、そのようなアドバイスがアイデンティティの手がかりに言語的に適応することも示し、個別化と中立性の間の管理上のジレンマを明らかにします。最後に、AI を介したアドバイスに対する中立性、文化的配慮、信頼を確保しようとする企業、金融機関、規制当局への影響を強調しています。

原文 (English)

When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice

Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok) using 432 simulated advisor-client interactions spanning 16 religious identity pairings (Christian, Muslim, Hindu, and non-religious) and three core household financial decisions: stock investment, house purchase, and life insurance. Combining regression and reflexive thematic analyses, we identify structural biases across models and decision contexts and the discursive mechanisms through which they are linguistically enacted. Unbiased advice appeared in only 12-18% of cases. Gemini consistently produced more bias than Grok, while ChatGPT's outputs were statistically comparable to Grok's. Religiously symmetric advisor-client pairings almost always triggered explicit religious framing, and non-religious clients often received advisor-centered religious appeals. Qualitative findings show that bias is linguistically manifested through religious anchoring, uneven cultural signaling, and tone modulation, varying by model and financial scenario. Stock investment prompts produced more financially technical responses, whereas life insurance advice triggered stronger religious language. The study develops a dual-dimensional framework linking structural bias rooted in model training and design with discursive bias expressed through language, advancing understanding of algorithmic bias in LLM-generated financial advice. It also shows that such advice adapts linguistically to identity cues, revealing a managerial dilemma between personalization and neutrality. Finally, it highlights implications for businesses, financial institutions, and regulators seeking to ensure neutrality, cultural sensitivity, and trust in AI-mediated advice.

13:00 JST研究/論文

教育を中心とした AI の重要政策分析: ガーナの AI 戦略を事例として

国家 AI 戦略は、ガバナンス、労働力開発、イノベーション、競争力の指針となることが増えていますが、教育的、文化的、倫理的、実装上の要求を持つ分野として教育をどのように枠組み化しているかについてはあまり知られていません。この研究では、教育中心の AI 政策フレームワークを開発および適用して、2025 年から 2035 年のガーナの国家人工知能戦略を分析します。私たちは、重要な定性的な政策文書分析を使用して、政策の目的、教師の主体性と専門学習、カリキュラムと評価、言語と文化、責任ある AI と学習者の保護、参加と実施のガバナンスの 6 つの要素を通じて戦略を検討しました。調査結果によると、ガーナの戦略は、特に AI リテラシー、若者のスキル、TVET、労働力の準備状況、地方への支援、現地言語データ、インクルージョン、責任ある AI ガバナンスに重点を置いている点で、野心的かつタイムリーであることが示されています。ただし、教育の課題は、学校レベルの実装よりも国家としての AI の準備に重点が置かれています。教師の代理権、勤務前の教師教育、カリキュラムの進行、評価指導、AI の開示、多言語教育学、文化に応じた AI の使用、子供中心の保護措置、参加型ガバナンスは未開発のままです。また、目に見える開示がない明らかな AI スタイルのビジュアル コンテンツや、ビジョンとミッションの数値とそのテキストの説明の不一致など、透明性と一貫性に関する文書レベルの懸念も特定します。私たちは、ガーナには、労働力の準備と教師の準備、カリキュラム改革、評価の再設計、学習者保護、インフラストラクチャ、現地の言語指導、文化に応じた教育学、現地に応じた AI ツール、参加型ガバナンスを結び付ける、セクター固有の教育中心の AI 政策と実装経路が必要であると主張します。

原文 (English)

Education-centered critical policy analysis of AI: Ghana's AI strategy as a case

National AI strategies increasingly guide governance, workforce development, innovation, and competitiveness, but less is known about how they frame education as a sector with pedagogical, cultural, ethical, and implementation demands. This study develops and applies an Education-Centered AI Policy Framework to analyze Ghana's National Artificial Intelligence Strategy, 2025-2035. Using critical qualitative policy document analysis, we examined the strategy through six components: policy purpose, teacher agency and professional learning, curriculum and assessment, language and culture, responsible AI and learner protection, and participation and implementation governance. Findings show that Ghana's strategy is ambitious and timely, especially in its emphasis on AI literacy, youth skills, TVET, workforce readiness, rural outreach, local language data, inclusion, and responsible AI governance. However, the education agenda is stronger on national AI readiness than on school-level implementation. Teacher agency, pre-service teacher education, curriculum progression, assessment guidance, AI disclosure, multilingual pedagogy, culturally responsive AI use, child-centered safeguards, and participatory governance remain underdeveloped. We also identify document-level concerns about transparency and coherence, including apparent AI-styled visual content without visible disclosure and a mismatch between a vision and mission figure and its textual explanation. We argue that Ghana needs a sector-specific, education-centered AI policy and implementation pathway that connects workforce readiness with teacher preparation, curriculum reform, assessment redesign, learner protection, infrastructure, local language instruction, culturally responsive pedagogy, locally responsive AI tools, and participatory governance.

13:00 JST研究/論文

静的な大きなグラフの平均距離の近似

大規模ネットワークでの平均距離の計算は計算量が多く、限られたメイン メモリの制約を受けるため、グラフ分析において大きな課題となります。この研究では、平均距離を推定するための 2 つの主要なアプローチ、つまりグラフ サンプリング ベースの方法 (ランダム ウォーク) と、サイズ推定フレームワーク (SEF) やエップスタイン-ワン (EW) アルゴリズムを含むランドマーク ベースの方法を調査し、評価します。ランダム ウォークは、サンプル サイズが小さい場合は信頼性が低く、サンプル サイズが大きい場合は計算コストが高くつくため、精度を得るには少なくとも 15% のノードが必要であることが判明しました。メモリ効率の高い近隣探索のために HyperLogLog などの確率的データ構造を利用するランドマークベースのアプローチは、優れたパフォーマンスを実証しました。このうち、SEF アルゴリズムはメモリ効率が高く、EW アルゴリズムはより短い計算時間で高い精度を実現します。静的、無向、無重みグラフ (一部分と二部分の両方) での実験により、EW アルゴリズムが 0.02% という低い誤差範囲で結果を生成することが明らかになりました。さらに、ほとんどの大規模なグラフで正確な推定を行うには、ランダムに選択された 100 個のノードのサブセットで十分でした。この調査結果は、EW アルゴリズムが、二部グラフと比較して単部グラフの信頼性が高い、平均距離推定のための実用的でスケーラブルなソリューションを提供することを示しています。

原文 (English)

Average Distance Approximation for Static Large Graphs

Calculating average distances in large-scale networks is computationally intensive and constrained by limited main memory, posing a significant challenge in graph analytics. This study explores and evaluates two primary approaches for estimating average distances: a graph sampling-based method (Random Walk) and landmark-based methods, including the Size Estimation Framework (SEF) and the Eppstein-Wang (EW) algorithm. Random Walk was found to be unreliable for small sample sizes and computationally expensive for larger ones, requiring at least 15% of nodes for accuracy. Landmark-based approaches, leveraging probabilistic data structures like HyperLogLog for memory-efficient neighbor exploration, demonstrated superior performance. Among these, the SEF algorithm offers better memory efficiency, while the EW algorithm achieves higher accuracy with lower computation time. Experiments on static, undirected, and unweighted graphs (both unipartite and bipartite) revealed that the EW algorithm produced results with an error margin as low as 0.02%. Additionally, a subset of 100 randomly selected nodes was sufficient for accurate estimations in most large graphs. The findings indicate that the EW algorithm provides a practical and scalable solution for average distance estimation, with higher reliability on unipartite graphs compared to bipartite graphs.

13:00 JST研究/論文

対象範囲が少ない: 特許先行技術検索のためのセマンティック センター表現

特許先行技術検索は、長く高度に構造化された技術文書に対する想起指向の検索タスクです。高密度検索によりセマンティック マッチングが向上しますが、単一ベクトル表現では、複数の技術コンポーネント、機能、および制約が 1 つの埋め込みに圧縮される可能性があります。我々は、ローカルスパン埋め込みを埋め込み空間中心の疎な語彙にマッピングする教師なし意味検索フレームワークであるSparse Coverageを提案します。中心はカバレッジ指向の k 中心目的で選択され、スパンは近くの中心をアクティブにして、逆インデックス検索と互換性のあるスパース表現を生成します。 CLEF-IP 2013 の実験では、Sparse Coverage が、いくつかの構成における強力な高密度特許エンコーダの文書レベルのリコールと同等かそれを上回っており、パッセージレベルの検索では競争力を維持していることが示されています。 Sparse Coverage は、ローカルの意味論的証拠をスパース逆索引検索と組み合わせることで、特許検索の効果的な第 1 段階の検索アプローチを提供します。

原文 (English)

Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval

Patent prior-art retrieval is a recall-oriented search task over long and highly structured technical documents. Dense retrieval improves semantic matching, but single-vector representations may compress multiple technical components, functions, and constraints into a single embedding. We propose Sparse Coverage, an unsupervised semantic retrieval framework that maps local span embeddings to a sparse vocabulary of embedding-space centers. The centers are selected with a coverage-oriented k-center objective, and spans activate nearby centers to produce sparse representations compatible with inverted-index retrieval. Experiments on CLEF-IP 2013 show that Sparse Coverage matches or exceeds the document-level recall of strong dense patent encoders in several configurations, while remaining competitive for passage-level retrieval. By combining local semantic evidence with sparse inverted-index search, Sparse Coverage provides an effective first-stage retrieval approach for patent search.

13:00 JSTエージェント

CARA: Cognitive Adaptive Recommendation Agent

Recent advances in large language models and agent-based recommendation frameworks have introduced new opportunities for more flexible and…

13:00 JSTLLM/生成AI研究/論文

WIP: LLM Odyssey: LLM エンジニアリング概念を教えるためのゲームベースのプラットフォーム

この進行中 (WIP) の革新的な実践カテゴリの論文では、大規模言語モデル (LLM) エンジニアリングの概念を教えるための 13 のインタラクティブ ゲームで構成される、オープンソースのブラウザベースの本格的なゲーム プラットフォームである LLM Odyssey について説明します。トークン化、トランスフォーマー アーキテクチャ、プロンプト エンジニアリング、検索拡張生成 (RAG)、本番展開などのトピックは、コンピューター サイエンスのカリキュラムでは過小評価されています。既存の対話型ツールは個々の概念に対応していますが、教育的な足場や構造化された学習経路が欠けています。 LLM Odyssey は、Bloom の改訂された分類法に沿った 3 つの学習層、つまり Cognitive Core (7 つの基礎ゲーム)、Systems Forge (5 つの生産エンジニアリング ゲーム)、および Foundry Arena (頂点の課題) を通じてこのギャップに対処します。各ゲームには、文献から引き出された 5 つの教育戦略が組み込まれています。つまり、即時形成的なフィードバック、近接発達ゾーンに基づいた足場のヒント、フロー理論に基づいた段階的な難易度、認知負荷を管理するための実践例、制作実践から引き出された本物のシナリオです。このプラットフォームは、2026 年の冬学期にカナダの大学で初期レビューのために導入されました。フィードバックにより機能要件が確認され、将来の開発の優先事項として適応困難性が特定されました。正式な混合方法評価プロトコル (N=50) が設計されており、事前および事後の知識テスト、検証済み調査、エンゲージメント分析、インタビューで構成されており、公的に利用可能なプラットフォームを使用した将来の評価研究を可能にするためにここに文書化されています。

原文 (English)

WIP: LLM Odyssey: A Game-Based Platform for Teaching LLM Engineering Concepts

This work-in-progress (WIP) innovative practice category paper presents LLM Odyssey, an open source, browser-based serious gaming platform comprising 13 interactive games for teaching Large Language Model (LLM) engineering concepts. Topics such as tokenization, transformer architecture, prompt engineering, retrieval augmented generation (RAG), and production deployment are underrepresented in computer science curricula. Existing interactive tools address individual concepts but lack pedagogical scaffolding or structured learning pathways. LLM Odyssey addresses this gap through three learning tiers aligned with Bloom's revised taxonomy: Cognitive Core (7 foundational games), Systems Forge (5 production engineering games), and Foundry Arena (capstone challenges). Each game incorporates five pedagogical strategies drawn from the literature: immediate formative feedback, scaffolded hints grounded in the Zone of Proximal Development, progressive difficulty informed by flow theory, worked examples to manage cognitive load, and authentic scenarios drawn from production practice. The platform was deployed in Winter 2026 semester at a Canadian college for an initial review. Feedback confirmed functional requirements and identified adaptive difficulty as a priority for future development. A formal mixed methods evaluation protocol (N=50) has been designed, comprising pre and post knowledge tests, validated surveys, engagement analytics, and interviews, and is documented here to enable future evaluation studies with the publicly available platform.

13:00 JST研究/論文

EMAN: マルチタスク学習におけるパス創発による最適化主導のキャパシティ拡大

既存のマルチタスク学習方法は、ハード共有、複数のパスまたは専門家、適応型共有、および動的拡張に依存しています。ただし、その能力の変更は通常、事前定義された構造によって制約されるか、タスクの境界や競合信号によってトリガーされます。これは基本的な疑問を引き起こします。ネットワークは正確な単一パスの計算から開始し、永続的な最適化の証拠が現れた場合にのみ新しい独立したパスを成長させることができるのでしょうか?我々は、2番目のパスをインスタンス化することなく、潜在的な相対フェーズを通じて非対称な成長方向を明らかにし、トレーニング中に複数の決定信号を監視して局所最適化の証拠を構造的決定に変換するための最適化主導型フレームワークであるEmergent Modular Atomic Network (EMAN)を提案します。 EMAN は、認証後にのみ、2 つの等しい容量の独立したパスを実現します。 EMAN は、さまざまなタスク要件に対応するために、共有およびタスク固有の表現容量を適応的に割り当てます。制御されたランク設定、PASCAL-Context、および NYUv2 に関する広範な実験によりその有効性が検証され、競争力のある計算コストでパフォーマンスの向上が実現されます。

原文 (English)

EMAN: Optimization-Driven Capacity Growth through Path Emergence in Multi-Task Learning

Existing multi-task learning methods rely on hard sharing, multiple paths or experts, adaptive sharing, and dynamic expansion. However, their capacity changes are usually constrained by predefined structures or triggered by task boundaries and conflict signals. This raises a fundamental question: can a network start from exact single-path computation and grow a new independent path only when persistent optimization evidence appears? We propose the Emergent Modular Atomic Network (EMAN), an optimization-driven framework for exposing an antisymmetric growth direction through latent relative phases without instantiating a second path, and for monitoring multiple decision signals during training to transform local optimization evidence into a structural decision. EMAN materializes two equal-capacity independent paths only after certification. EMAN adaptively allocates shared and task-specific representation capacity to accommodate varying task requirements. Extensive experiments on controlled rank settings, PASCAL-Context, and NYUv2 validate its effectiveness, achieving improved performance at a competitive computational cost.

13:00 JSTLLM/生成AIハードウェア/半導体Gemma

プレフィルの調査: 潜在的なアクティベーションによるコードの脆弱性の検出

LLM ベースのコード生成は現在、ミッション クリティカルなパイプラインに組み込まれていますが、脆弱な出力に対する防御は事後的なままです。静的アナライザー、微調整された分類子、または LLM が生成モデル自体の内部状態を無視してコードが完了したかどうかを判断します。私たちは、より限定的で直接測定可能な質問をテストします。LLM が C/C++ コードの一部をコンテキストとして読み取るとき、その隠れたアクティベーションはすでにそのコードの脆弱性ステータスに関する信号を伝えていますか? 3 つのモデル ファミリにわたる 4 つの LLM (Granite-4.1-8B、Qwen3.5-9B、Qwen3.6-27B、Gemma-4-12B) から最後のプレフィル トークン アクティベーションを抽出し、これらのアクティベーションに対して MLP プローブをトレーニングします。これらを 4 つの関数レベルの C/C++ ベンチマーク (Devign、Big-Vul、Draper VDISC、PrimeVul) で評価します。当社のプローブは、13.4 ~ 16.0M パラメータのプローブを使用して平均 F1 41.7\% を達成します。これはベースモデル サイズの 0.2\% 未満です。 Devign では、最適なプローブ (Qwen3.5-9B、68.8\% F1) は、凍結された汎用 LLM のアクティベーションのみを読み取るにもかかわらず、公開されている微調整された分類子 SOTA (67.9\%) と一致します。より難しく、より不均衡なベンチマーク (Big-Vul、Draper VDISC、PrimeVul) プローブでは、SOTA に大幅に遅れをとります。これは、コーディング LLM による任意のコードの表現がそのコードの脆弱性ステータスに関する情報となるという初期の証拠であり、軽量でモデルネイティブな脆弱性スクリーニングに向けたさらなる作業の動機付けとなります。

原文 (English)

Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations

LLM-based code generation is now embedded in mission-critical pipelines, but defenses against vulnerable output remain post-hoc -- static analyzers, fine-tuned classifiers, or an LLM judge that screen completed code, ignoring the generating model's own internal state. We test a narrower, directly measurable question: when an LLM reads a piece of C/C++ code as context, do its hidden activations already carry a signal about that code's vulnerability status? We extract last prefill token activations from four LLMs (Granite-4.1-8B, Qwen3.5-9B, Qwen3.6-27B, Gemma-4-12B) across three model families and train MLP probes on these activations. We evaluate them on four function-level C/C++ benchmarks (Devign, Big-Vul, Draper VDISC, PrimeVul). Our probes achieve 41.7\% average F1 using 13.4--16.0M-parameter probes -- under 0.2\% of base-model size. On Devign, the best probe (Qwen3.5-9B, 68.8\% F1) matches the published fine-tuned-classifier SOTA (67.9\%) despite reading only a frozen, general-purpose LLM's activations; on the harder, more imbalanced benchmarks (Big-Vul, Draper VDISC, PrimeVul) probes trail SOTA substantially. This is early evidence that a coding LLM's own representation of arbitrary code is informative about that code's vulnerability status, motivating further work toward lightweight, model-native vulnerability screening.

13:00 JSTビジネス/資金調達

立場: 生成モデルにおける公平性の失敗は評価の問題である

過去 10 年間における生成モデルの画期的な進歩にも関わらず、その公平性の欠如、社会的不平等の強化、疎外されたグループへの損害に関する懸念は依然として十分に対処されておらず、対応するのが困難です。この意見書では、生成モデルにおける公平性の失敗は、複数の要因によって引き起こされるとしても、最終的には評価の問題に起因すると主張しています。つまり、公平性の結果が論文間で比較されたり、展開の決定に実用的になることはほとんどありません。この論文は、現在の実践において繰り返される経験的および概念的な失敗モードを診断し、その場限りのバイアス チェックから標準化された生成固有の評価への移行を促すものです。私たちは、再現性、比較可能性、説明責任を可能にする評価の選択肢(プロンプトファミリー、反事実プロトコル、指標、拒否処理)を明示する最小限のレポート成果物としてフェアネスカードを提案します。最後に、評価基準のパラダイムシフトに向けた追加の推奨事項を述べます。私たちのプロジェクト ページは https://mariiavladimirova.github.io/fairness-cards にあります。

原文 (English)

Position: Fairness Failure in Generative Models is an Evaluation Problem

Despite groundbreaking advancements in generative models during the last decade, concerns about their lack of fairness, reinforcing societal inequalities and harming marginalized groups, remain under-addressed and difficult to act upon. This position paper argues that fairness failures in generative models, albeit driven by multiple factors, are ultimately stemming from an evaluation problem: fairness findings are rarely comparable across papers or actionable for deployment decisions. This paper diagnoses recurring empirical and conceptual failure modes in current practice and motivates a shift from ad-hoc bias checks to standardized, generative-specific evaluation. We propose Fairness Cards as a minimal reporting artifact that makes evaluation choices explicit (prompt families, counterfactual protocols, metrics, and refusal handling) enabling reproducibility, comparability, and accountability. We conclude with additional recommendations towards a paradigm shift in evaluation standards. Our project page can be found at https://mariiavladimirova.github.io/fairness-cards .

13:00 JST画像/動画生成

PXDepth: 構造を保存する単眼深度推定のためのピクセル空間モデリング

最近の単眼深度推定器は、強力なゼロショット一般化を達成していますが、きめの細かい構造やオブジェクトの境界を維持するのに苦労することがよくあります。粗いトークン化により、アップサンプリングでは完全に回復できないピクセルレベルのキューが弱くなる可能性があるため、この制限は大パッチ ViT エンコーダと畳み込みデコーダの一般的な組み合わせによるものであると考えられます。この問題に対処するために、グローバル コンテキスト モデリングをピクセル レベルの深度予測から分離する識別的な単眼深度モデルである PXDepth を提案します。具体的には、大規模パッチ ViT がグローバル シーン コンテキストをキャプチャし、Context-Modulated Pixel Transformer ブロックで構成されるピクセル空間予測器が深度推定全体を通じて高解像度の空間表現を維持します。この設計は、全体的な深さの一貫性を犠牲にすることなく、微細な構造と明確な境界を維持します。 PXDepth は、さまざまなゼロショット ベンチマークにわたって、推論の効率を維持しながら、忠実なローカル ジオメトリと競争力のあるグローバル深度精度を組み合わせます。コードとモデルは https://yuanzhy29.github.io/PXDepth-Page/ で入手できます。

原文 (English)

PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation

Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and object boundaries. We attribute this limitation to the prevalent combination of large-patch ViT encoders and convolutional decoders, as coarse tokenization can weaken pixel-level cues that upsampling cannot fully recover. To address this issue, we propose PXDepth, a discriminative monocular depth model that separates global context modeling from pixel-level depth prediction. Specifically, a large-patch ViT captures global scene context, while a pixel-space predictor composed of Context-Modulated Pixel Transformer blocks maintains high-resolution spatial representations throughout depth estimation. This design preserves fine structures and sharp boundaries without sacrificing global depth consistency. Across diverse zero-shot benchmarks, PXDepth combines faithful local geometry with competitive global depth accuracy while remaining efficient at inference. Our code and model are available at https://yuanzhy29.github.io/PXDepth-Page/.

13:00 JST研究/論文

ジャーナリストがいなければジャーナリズムは存在しません: メディアにおける生成人工知能の社会的側面

人工知能の技術やツールがメディアに導入されると、今後数十年間でメディアや専門家の仕事が体系的かつ継続的に変化することになるでしょう。この目的を達成するために、この記事では、過去 20 年間にメディアにおける AI の導入に関して実施された研究、特に実証研究を体系的にレビューし、AI の導入によってもたらされる主な社会的および認識論的課題を特定します。メディアにとっては、テクノロジープラットフォームへの依存度の増大と編集の独立性を守ることが主な課題となるだろう。一方、ジャーナリストは、自分たちの仕事に対する認識されている脅威と、現実と聴衆の間の仲介者としての象徴的な資本の喪失と、その後により質の高いコンテンツを制作できるようになる日常業務からの解放との間で引き裂かれています。一方、テキストの読みやすさでは依然として人間の執筆の方が有利であるにもかかわらず、聴衆は自動化されたテキストの品質と信頼性に大きな違いを認識していないようです。つまり、技術中心的または決定論的なアプローチを超えて、ジャーナリズムなどの特に人間の分野における AI の使用には、視聴者によるイノベーションの流用とそれが視聴者に与える影響が開発の鍵の 1 つである社会的アプローチが必要です。したがって、メディアにおける AI の研究は、AI が個人やジャーナリストにどのような影響を与えるか、職業上の適切な目的と社会的利益のためにどのように使用できるか、AI の使用によって引き起こされる可能性のあるギャップをどのように埋めるかを分析することに重点を置く必要があります。

原文 (English)

Without journalists, there is no journalism: the social dimension of generative artificial intelligence in the media

The implementation of artificial intelligence techniques and tools in the media will systematically and continuously alter their work and that of their professionals during the coming decades. To this end, this article carries out a systematic review of the research conducted on the implementation of AI in the media over the last two decades, particularly empirical research, to identify the main social and epistemological challenges posed by its adoption. For the media, increased dependence on technological platforms and the defense of their editorial independence will be the main challenges. Journalists, in turn, are torn between the perceived threat to their jobs and the loss of their symbolic capital as intermediaries between reality and audiences, and a liberation from routine tasks that subsequently allows them to produce higher quality content. Meanwhile, audiences do not seem to perceive a great difference in the quality and credibility of automated texts, although the ease with which texts are read still favors human authorship. In short, beyond technocentric or deterministic approaches, the use of AI in a specifically human field such as journalism requires a social approach in which the appropriation of innovations by audiences and the impact it has on them is one of the keys to its development. Therefore, the study of AI in the media should focus on analyzing how it can affect individuals and journalists, how it can be used for the proper purposes of the profession and social good, and how to close the gaps that its use can cause.

13:00 JST画像/動画生成

YILDIZ-VPR: 視覚的な場所認識のための、多様な環境条件下で高密度にカバーする新しいデータセット

Visual Place Recognition (VPR) は、クエリ画像を一連の地理参照画像と比較することで、その画像の位置を認識することを目的としています。 VPR 用に多くのデータセットが提案されていますが、歩行者レベルの視点から高密度で多様な視覚データを収集することは依然として重要なニーズです。この論文では、Yildiz Technical University の Davutpasa キャンパスでの繰り返しの歩行横断を通じて収集された視覚的な地理位置情報データセットである YILDIZ-VPR を紹介します。データセットには、さまざまな時間帯、季節、気象条件で撮影された屋外シーンが含まれています。歴史的建造物、近代的な建造物、道路、緑地、森林地帯など、幅広いビジュアル コンテンツが含まれています。各ビデオは GoPro 9 カメラで録画され、抽出されたフレームに位置ラベルを提供するために GPS センサー データと同期されました。 GPS 座標に加えて、データセットにはジャイロスコープ、速度、温度データなどの補助センサー情報も含まれています。 YILDIZ-VPR は、その高密度な適用範囲と長期的な視覚的変動により、現実的な屋外条件下で画像ベースおよび時間的な視覚的場所認識を研究するための有用なリソースを提供します。

原文 (English)

YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition

Visual Place Recognition (VPR) aims to recognize the location of a query image by comparing it with a set of geo-referenced images. Although many datasets have been proposed for VPR, collecting dense and diverse visual data from pedestrian-level viewpoints is still an important need. In this paper, we introduce YILDIZ-VPR, a visual geo-localization dataset collected through repeated walking traversals on the Davutpasa campus of Yildiz Technical University. The dataset includes outdoor scenes captured at different times of day, seasons, and weather conditions. It contains a wide range of visual content, including historical buildings, modern structures, roads, green areas, and wooded regions. Each video was recorded with a GoPro 9 camera and synchronized with GPS sensor data to provide location labels for the extracted frames. In addition to GPS coordinates, the dataset also includes auxiliary sensor information such as gyroscope, speed, and temperature data. With its dense coverage and long-term visual variability, YILDIZ-VPR provides a useful resource for studying image-based and temporal visual place recognition under realistic outdoor conditions.

13:00 JST画像/動画生成研究/論文

第10回AIシティチャレンジ

ECCV 2026 とともに開催される第 10 回 AI シティ チャレンジは、インテリジェント交通機関、スマート シティ、物理 AI のコミュニティ ベンチマークの 10 年を記念します。 2017 年に車両の検出、分類、追跡から始まって以来、この課題は、マルチカメラの認識、マルチモーダル推論、合成から現実への学習、生成予測、およびプライバシー保護の評価のための広範なベンチマーク スイートに成長しました。 2026 年版でも、登録チームは 325 となり、2025 年の 245 チームから増加し、26 の国と地域からの参加と、15 チームから増加し、この成長を続けました。その 6 つの主なトラックは、マルチカメラ 3D 認識、交通安全キャプションと VQA、交通異常推論、テキストベースの人物異常探索、生成交通ビデオ予測、および都市横断物体検出をカバーしています。トラック 3 にはさらに、魚眼レンズによる交通違反の理解と歩行者の位置を考慮した VQA のために、トラック 7 および 8 として提出された 2 つのドメイン外リーダーボードが含まれています。このペーパーでは、チャレンジの設定、データセット、評価プロトコル、リーダーボードの結果、およびワークショップの資料を要約します。トラック全体で、成功したシステムは、基礎モデルと幾何学的基礎付け、検索または再ランキング、合成データ設計、ドメイン適応、および制御された推論を組み合わせています。

原文 (English)

The 10th AI City Challenge

The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and physical AI. Since its 2017 start with vehicle detection, classification, and tracking, the challenge has grown into a broad benchmark suite for multi-camera perception, multimodal reasoning, synthetic-to-real learning, generative forecasting, and privacy-preserving evaluation. The 2026 edition continued this growth with 325 registered teams, up from 245 in 2025, and participation from 26 countries and regions, up from 15. Its six primary tracks cover multi-camera 3D perception, transportation safety captioning and VQA, traffic anomaly reasoning, text-based person anomaly search, generative traffic video forecasting, and cross-city object detection. Track 3 further includes two out-of-domain leaderboards, submitted as Tracks 7 and 8, for fisheye traffic-violation understanding and pedestrian situated-intent VQA. This paper summarizes the challenge setup, datasets, evaluation protocols, leaderboard results, and workshop papers. Across tracks, successful systems combine foundation models with geometric grounding, retrieval or reranking, synthetic-data design, domain adaptation, and controlled inference.

13:00 JSTLLM/生成AI

ターゲット側リーダー適応によるクロスモデルメモリ転送

大規模な言語モデルで知識の使用を改善する方法は、通常 2 つの体制に分類されます。ノンパラメトリック検索では、外部の知識に柔軟にアクセスできますが、検索の遅延、コンテキストのオーバーヘッドが追加され、バックボーンとの統合は浅いだけになります。パラメトリック適応は推論時には効率的ですが、知識とモデルの重みが絡み合うため、更新、監査、転送が困難になる場合があります。エングラム スタイルのハッシュ メモリは中間領域を占めます。つまり、学習した情報を外部のアドレス指定可能なテーブルに保存しますが、そのテーブルは小型の学習済みリーダーを通じて消費されます。これは基本的な疑問を引き起こします。そのようなメモリがバックボーンを越えて移動されるとき、凍結されたメモリ自体とターゲット側のリーダーのどちらがより重要なのでしょうか?私たちは、クロスモデルのフリーズメモリ抽出を通じてこの疑問を研究します。この抽出では、ソース モデルでトレーニングされたメモリがフリーズされ、別のターゲット モデルにアタッチされ、軽量リーダーのみがトレーニングされます。アブレーションは、学習されたメモリ内容と正しいアドレス指定の両方が重要であることを示していますが、転送されたテーブルは、ターゲット モデルに合わせて調整されたリーダーを通じてのみ有用になります。下流の質問応答タスクでは、デュアルレイヤー 4 ブランチ リーダーにより、同一モデルの再利用とクロスモデルの再利用との間のギャップがほぼ縮まり、管理された評価プロトコルの下で平均スコア 38.8 を達成しました。さらに、プロバイダー リーダーがターゲット インターフェイスと直接互換性がある場合、フリーズされたアーティファクトはターゲット側のトレーニングなしで実質的な有用性を提供でき、オプションのリーダー適応によりさらなる改善がもたらされます。これらの結果は、ターゲットが互換性のあるリーダー インターフェイスにアクセスできる場合、Engram が再利用可能な外部知識アーティファクトとして機能できることを示唆しています。リーダーの直接の再利用が不十分な場合、ターゲット側の適応によりアライメントをさらに改善できます。

原文 (English)

Cross-Model Memory Transfer via Target-Side Reader Adaptation

Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible access to external knowledge, but adds retrieval latency, context overhead, and only shallow integration with the backbone. Parametric adaptation is efficient at inference time, but entangles knowledge with model weights and can be hard to update, audit, or transfer. Engram-style hashed memory occupies a middle regime: it stores learned information in an external, addressable table, yet consumes that table through a small learned reader. This raises a basic question: when such a memory is moved across backbones, what matters more, the frozen memory itself or the target-side reader? We study this question through cross-model frozen-memory extraction, in which a memory trained on a source model is frozen and attached to a different target model, with only a lightweight reader trained. Ablations show that learned memory content and correct addressing both matter, but the transferred table becomes useful only through a reader aligned to the target model. In downstream question answering tasks, a dual-layer, four-branch reader nearly closes the gap between same-model and cross-model reuse, achieving an average score of 38.8 under our controlled evaluation protocol. Moreover, when the provider reader is directly compatible with the target interface, the frozen artifact can provide substantial utility without target-side training, while optional reader adaptation yields further improvement. These results suggest that Engram can serve as a reusable external knowledge artifact, provided that the target has access to a compatible reader interface; target-side adaptation can further improve alignment when direct reader reuse is insufficient.

13:00 JSTLLM/生成AI

機関固有の LLM プロンプトにより、匿名化システムとそのゴールドスタンダードが両方とも欠落している PHI を回復

電子医療記録の二次使用には匿名化が必要ですが、既存のシステムは、病院の略称、建物名、地域でステータスが決定されている内部コードなどの \emph{施設内にある} 保護された医療情報 (PHI) を見逃しています。インコンテキスト学習 (ICL) を備えた大規模言語モデル (LLM) がこのギャップを埋め、精度と再現率のトレードオフを制御できるかどうかを考えます。テキサス小児病院からの注釈付き小児腫瘍学ノート 100 冊 (5,322 PHI スパン) について、2 つの専用システム (Stanford TiDE、OpenMed PII) および 2 つのパターンベースのベースラインに対して 8 つの LLM のベンチマークを実施しました。各 LLM は、特異性を高める 3 つのプロンプトの下で実行されました: (1) HIPAA に合わせたベースライン、(2) ベースラインに見逃した施設の PHI カテゴリを加えたもの、(3) プロンプト 2 に加えて臨床内容の過剰な編集に対する指示。次に、主要な安全性指標を思い出しながら、14 個のマルチエージェントおよびアンサンブル構成を最良の単一プロンプトと比較しました。 LLM は専用システムよりも優れたパフォーマンスを示し (最良の F1=0.918$\pm$0.001 対 \ TiDE 0.779)、利点はコンテキスト カテゴリに集中していました。見逃したカテゴリに名前を付けると、そのうち 79\% (48/61) が復元され、過剰な編集を阻止して精度を回復しました。キャリブレーションされたシングルパス プロンプト (F1 0.906 ~ 0.907) に勝るエージェント アーキテクチャはありませんが、LLM 出力は 414 ~ 候補アノテーション ギャップを表面化しました。再アノテーションにより 227~PHI スパンが確認され、これに対して最終プロンプトはリコール = 0.981 (F1=0.907$\pm$0.002) に達しました。適切にキャリブレーションされた ICL は、ノートごとに 1 つの LLM コールで制度上の PHI ギャップと精度とリコールのトレードオフの両方を解決します。 LLM は従来の方法よりも実行コストが高くなりますが、そのコストで参照標準を監査する方法が得られます。 LLM は、専用の匿名化システムに代わる合法的で適応性のある代替手段です。施設固有の迅速な開発が主要な適応戦略であるべきである。

原文 (English)

Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss

Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected health information (PHI) such as hospital abbreviations, building names, and internal codes whose status is locally determined. We ask whether large language models (LLMs) with in-context learning (ICL) can close this gap and control the precision--recall trade-off. On 100 annotated pediatric oncology notes (5,322 PHI spans) from Texas Children's Hospital, we benchmarked eight LLMs against two purpose-built systems (Stanford TiDE, OpenMed PII) and two pattern-based baselines. Each LLM ran under three prompts of increasing specificity: (1) a HIPAA-aligned baseline, (2) baseline plus the institutional PHI categories it missed, and (3) prompt 2 plus instructions against over-redacting clinical content. We then compared 14~multi-agent and ensemble configurations against the best single prompt, with recall the primary safety metric. LLMs outperformed the purpose-built systems (best F1=0.918$\pm$0.001 vs.\ TiDE 0.779), with advantages concentrated in contextual categories. Naming the missed categories recovered 79\% (48/61) of them, and discouraging over-redaction restored precision. No agentic architecture beat calibrated single-pass prompting (F1 0.906--0.907), but LLM outputs surfaced 414~candidate annotation gaps; re-annotation confirmed 227~PHI spans, against which the final prompt reached recall=0.981 (F1=0.907$\pm$0.002). Well-calibrated ICL resolves both the institutional PHI gap and the precision--recall trade-off in one LLM call per note. LLMs cost more to run than traditional methods, but that cost buys a way to audit the reference standard. LLMs are a legitimate, adaptable alternative to purpose-built de-identification systems; institution-specific prompt development should be the primary adaptation strategy.

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

財団エージェントとエージェントのディープリサーチの出会い: 証拠に基づいた臨床コードの予測

次回の ICD 予測では、事前に入手可能な長期的な記録から、次回の訪問時にどの標準化された診断コードが文書化されるかを予測します。このタスクは将来性があり、マルチラベルです。ターゲット ノートはまだ存在せず、いくつかのコードが正しい可能性があります。構造化された EHR 基礎モデルは再発と時間的進行を捉えますが、言語基礎モデルは柔軟な診断仮説を生成します。これらの予測基礎モデルと医療検索および ICD 辞書を構成する DeepResearch ワークフローである ICD-Deepresearch を紹介します。将来のコードセットを明らかにする情報源がないため、研究では、固定の上位 K 予算の下で、患者の証拠、外部の臨床関係、および正確なコードのセマンティクスを関連付けることにより、候補の移行を評価します。 Candidate Generation は SparseEHR を使用して、2 つの制限された Research Expansion ラウンドを初期化する EHR Prior を生成します。独立した GPT-5 Direct Forecast が補完的な候補を提供します。最終選択では、両方のパスを検証、重複排除し、共同でランク付けします。その後、別のモジュールが予測を変更せずに根拠を書き込みます。最後に、ICD-Deepresearch は、MIMIC-III で 24.60/35.09%、MIMIC-IV で 25.14/48.32% の患者平均適合率/再現率を達成しました。医師は、検索された文書の 51% と 68% が有用であると評価しています。これに対し、スタンドアロンの GPT-5 Web 検索では 22% と 39%、Medical Deep Research では 32% と 41% でした。したがって、ICD-Deepresearch は、登録されたローカル コンパレーターよりも優れており、スタンドアロンの研究システムよりも医師が評価した有用性が高い証拠を取得します。

原文 (English)

Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting

Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not yet exist, and several codes may be correct. Structured EHR foundation models capture recurrence and temporal progression, whereas language foundation models generate flexible diagnostic hypotheses. We introduce ICD-Deepresearch, a DeepResearch workflow that composes these predictive foundation models with medical search and ICD dictionaries. Because no source reveals the future code set, research evaluates candidate transitions by linking patient evidence, external clinical relations, and exact code semantics under a fixed top-K budget. Candidate Generation uses SparseEHR to produce an EHR Prior that initializes two bounded Research Expansion rounds; an independent GPT-5 Direct Forecast supplies complementary candidates. Final Selection validates, deduplicates, and jointly ranks both paths, after which a separate module writes rationales without changing predictions. Finally ICD-Deepresearch achieves patient-averaged precision/recall of 24.60/35.09% on MIMIC-III and 25.14/48.32% on MIMIC-IV. Physicians rate 51% and 68% of its retrieved documents useful, compared with 22% and 39% for standalone GPT-5 web search and 32% and 41% for Medical Deep Research. ICD-Deepresearch therefore improves over the registered local comparators while retrieving evidence with higher physician-rated usefulness than the standalone research systems

13:00 JSTエージェント

小型言語モデルベースの GNSS スプーフィング検出のための構造化された運転状態の説明

自動運転車 (AV) は、信頼性の高い全地球航法衛星システム (GNSS) 測位に依存しています。ただし、なりすましの GNSS 信号は、もっともらしいが不正確な車両状態を引き起こす可能性があります。この研究では、GNSS およびその他のセンシング ソースから独立して得られる車両の動作を比較することにより、GNSS スプーフィング攻撃を検出および分類するための小型言語モデル (SLM) ベースのフレームワークを開発します。このフレームワークは、GNSS やその他のセンシング ソースからの独立した運転状態を、スプーフィングの検出と攻撃の分類のために SLM に提供される構造化されたセマンティック ナラティブに変換します。 SLM ベースのフレームワークのパフォーマンスは、同一のトレーニング データで微調整され、同じテスト セットで評価された大規模言語モデル (LLM) と比較されます。評価では、攻撃なし、オーバーシュート攻撃、停止攻撃、ターンバイターン攻撃、間違ったターン攻撃の 5 つのクラスが考慮されます。このフレームワークは、米国サウスカロライナ州クレムソンで収集された地理的に未確認の現場データでも評価されます。実験結果は、評価された SLM が LLM と同様のパフォーマンスを達成し、平均精度 96.99%、精度 99.05%、再現率 95.59%、および F1 スコア 97.18% を達成することを示しています。計算効率とリソース使用率の点で、SLM は、微調整と推論の両方で必要な推論レイテンシと GPU メモリの量が少ないため、LLM よりも優れています。地理的に異なる場所で収集されたフィールドデータを使用した評価により、その有効性がさらに実証されました。提示されたフレームワークは、比較的少ない計算リソースとメモリ リソースを必要としながら、リアルタイムで GNSS スプーフィング攻撃を検出および分類できるため、リソースに制約のある車両コンピューティング プラットフォームへの展開に適しています。

原文 (English)

Structured Driving-State Narratives for Small Language Model-Based GNSS Spoofing Detection

Autonomous vehicles (AVs) depend on reliable Global Navigation Satellite System (GNSS) positioning. However, spoofed GNSS signals can induce plausible but incorrect vehicle states. This study develops a small language model (SLM)-based framework for detecting and classifying GNSS spoofing attacks by comparing vehicle behaviors independently derived from GNSS and other sensing sources. The framework converts independent driving states from GNSS and other sensing sources into structured semantic narratives that are provided to an SLM for spoofing detection and attack classification. The performance of the SLM-based framework is compared with large language models (LLMs) fine-tuned on identical training data and evaluated on the same test set. The evaluation considers five classes: no attack, overshoot attack, stopped attack, turn-by-turn attack, and wrong-turn attack. The framework is also evaluated with geographically unseen field data collected in Clemson, South Carolina, United States. Experimental results indicate that the evaluated SLMs achieve performance similar to the LLMs, achieving an average accuracy of 96.99%, precision of 99.05%, recall of 95.59%, and F1-score of 97.18%. In terms of computational efficiency and resource utilization, the SLMs demonstrate advantages over the LLMs by requiring lower inference latency and less GPU memory during both fine-tuning and inference. Evaluation using field data collected in a geographically distinct location further demonstrated its efficacy. The presented framework can detect and classify GNSS spoofing attacks in real-time while requiring relatively low computational and memory resources, and is therefore suitable for deployment on resource-constrained vehicular computing platforms.

13:00 JST研究/論文

From Abductive Explanations to Global Logical Rules for Node Classification in SGCs

Graph Neural Networks (GNNs) have achieved remarkable performance in node classification tasks, motivating growing interest in methods capa…

13:00 JSTビジネス/資金調達

初等関数およびフィルタリング関数を要素ごとに評価するための反復テンソル ネットワーク変換

Tensor ネットワークは、大規模なデータを圧縮するための強力な形式です。ただし、非線形演算の実行が難しいため、一般的なデータ処理への応用は制限されてきました。ここでは、テンソル ネットワークのクラスであるテンソル トレイン (TT) としてエンコードされたデータに対する初等フィルタリング関数と非線形フィルタリング関数を要素ごとに評価するための一般的なアルゴリズム フレームワークである反復テンソル ネットワーク変換 (ITNT) を紹介します。私たちのアプローチは完全に圧縮ドメインで動作し、制御された計算コストを維持しながら、指数関数的に大規模なデータセットの効率的な計算を可能にします。私たちは、その威力を 2 つの重要な分野で実証します。(I) 3D 反応性流れ場での高度に非線形の初等関数とフィルタリング関数を評価し、高忠実度の反応速度計算と領域フィルタリングを可能にすること、(II) 最大 $2^{70}$ 構成の空間上で Max-SAT インスタンスを解くなど、複雑な最適化問題における極値を見つけることです。これらの結果により、ITNT は、汎用データ サイエンスと大規模な最適化のための機能を備えたテンソル ネットワーク手法を提供する基礎的なツールとして確立されました。

原文 (English)

Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions

Tensor networks are powerful formats for compressing large-scale data. However, their application to general data processing has been limited by the difficulty of performing nonlinear operations. Here, we introduce iterative tensor network transformations (ITNTs), a general algorithmic framework for the element-wise evaluation of elementary and nonlinear filtering functions on data encoded as tensor trains (TTs), a class of tensor networks. Our approach operates entirely in the compressed domain, enabling efficient computation on exponentially large datasets while maintaining a controlled computational cost. We demonstrate its power in two key areas: (I) evaluating highly nonlinear elementary and filtering functions on a 3D reactive flow field, enabling high-fidelity reaction rate computation and region filtering, and (II) finding extrema in complex optimization problems, such as solving Max-SAT instances on spaces up to $2^{70}$ configurations. These results establish ITNT as a foundational tool that provides tensor network methods with the capability for general-purpose data science and large-scale optimization.

13:00 JSTLLM/生成AIエージェント

コンテキスト前の承認: エージェントシステムにおけるオーディエンス間のメモリリークに対するモデル中立のオーディエンス境界

個人言語エージェントは、ある聴衆から事実を学習し、後で別の聴衆のために組み立てるプロンプトにそれを配置する場合があります。このメモリからコンテキストへのステップは攻撃対象領域です。曖昧または一貫性のないチャネル、視聴者間の覗き見、および汚染されたメモリのそれぞれにより、システムは、クエリに関連するが現在の視聴者には許可されていないファクトを含むコンテキストを組み立てる可能性があります。私たちはコンテキストの前に承認を導入します。これは、メモリからコンテキストへの移行時に適用される、単調でない単一の視聴者メンバーシップ ルールです。各アイテムには、録音時にその場にいた聴衆が含まれています。現在のビューア セットはチャネル メタデータから読み取られ、あいまいな場合はパブリックに戻ります。そして、アイテムは、現在のすべての閲覧者がすでにその閲覧者に属している場合にのみ許可されます。私たちは、このルールがすべての参加者にクロスチャネルの想起を与える一方で、モデルの動作ではなく除外によって、より狭い聴衆向けに記録されたものはより広範な聴衆に到達せず、毒された記憶がそれ自身の聴衆を広げることができないことを保証することを証明します。境界は、正確に組み立てられたコンテキストに対するモデル中立の不変条件です。つまり、モデルが呼び出される前に、禁止された事実が存在しなければなりません。合成コンテキスト整合性スイートでは、私たちが組み立てた境界のコンテキストに禁止された事実は入りませんでしたが、範囲外のベースラインにはそのような事実が構造上含まれていました。さらに、すべての読み取りパスがフェイルクローズされていることを監査します。証拠は予備的かつ総合的なものです。

原文 (English)

Authorization Before Context: A Model-Neutral Audience Boundary Against Cross-Audience Memory Leakage in Agentic Systems

A personal language agent learns a fact from one audience and may later place it in the prompt it assembles for another. This memory-to-context step is an attack surface: ambiguous or inconsistent channels, cross-audience prying, and poisoned memory can each cause the system to assemble context containing a fact relevant to the query yet unauthorized for the current viewers. We introduce authorization before context: a single, anti-monotone audience-membership rule applied at the memory-to-context transition. Each item carries the audience present when it was recorded; the current viewer set is read from channel metadata and falls back to public when ambiguous; and the item is admitted only when every current viewer already belonged to its audience. We prove that this rule gives every participant cross-channel recall while ensuring, by exclusion rather than by model behavior, that nothing recorded for a narrower audience reaches a broader one and that poisoned memory cannot widen its own audience. The boundary is a model-neutral invariant on the exact assembled context: a forbidden fact must be absent before the model is called. On a synthetic Contextual-Integrity suite, no forbidden fact entered the context our boundary assembled, whereas unscoped baselines included such facts by construction; we further audit that every read path fails closed. The evidence is preliminary and synthetic.

13:00 JST研究/論文

世界モデルを使用した Q ラーニング

オフポリシー強化学習 (RL) のサンプル効率はますます高まっており、ビジョン-言語-アクション モデルを信頼性の高い高パフォーマンスのポリシーに微調整する RL などのアプリケーションが可能になります。ワールド モデルは、アクションだけではなく状態の変化を予測するため、サンプル効率をさらに高める手段を提供しますが、その成功は主に教師ありポリシー学習に限定されています。従来のモデルベースの RL 手法では、多くの場合、想像上のロールアウトに基づいてポリシーまたは価値関数を直接最適化しますが、バイアスが複合的に発生する傾向があり、現実世界のロボット工学などの大規模で高次元の問題に拡張するのが困難であり、この問題はタスクの範囲や視覚的な複雑さとともに悪化します。この研究では、代わりに、標準的な Q 学習に加えてワールド モデルを直接活用して、実際のオンライン環境でのトレーニングと基礎を維持しながら、パフォーマンスを向上させることができるかどうかを問います。私たちは、ワールド モデルを活用して、Q 学習に加えて想像上の軌道に対してテスト時の検索を実行し、オンライン ロールアウトと評価の両方で価値の高いアクションを選択するフレームワークである QWM を提案します。ポリシー関数と値関数は実際の遷移についてのみトレーニングされるため、QWM は予測検索によるサンプル効率の利点を活かしながら、モデルのバイアスの複合化を回避します。 Robomimic と LIBERO という困難な操作ベンチマークにおいて、QWM はサンプル効率とパフォーマンスの両方において強力な従来の最先端の手法を大幅に上回ります。

原文 (English)

Q-Learning With World Models

Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Language-Action models into reliable, high-performing policies. World models offer a further lever for sample efficiency, as they predict state changes rather than actions alone, but their success has largely been confined to supervised policy learning. Prior model-based RL methods often optimize the policy or value function directly on imagined rollouts, which is prone to compounding bias and struggles to scale to large, high-dimensional problems such as real-world robotics, a problem that worsens with task horizon and visual complexity. In this work, we instead ask whether we can leverage world models directly on top of standard Q-learning to improve performance, while remaining trained and grounded in the real, online setting. We propose QWM, a framework that leverages world models to perform test-time search over imagined trajectories on top of Q-learning to select high-value actions during both online rollouts and evaluation. Since the policy and value function are trained only on real transitions, QWM avoids compounding model bias while still gaining the sample-efficiency benefits of predictive search. On challenging manipulation benchmarks Robomimic and LIBERO, QWM significantly outperforms strong prior state-of-the-art methods on both sample efficiency and performance.

13:00 JST研究/論文

ベーテ ラグランジアンの情報制約として期待される自由エネルギー

能動推論では、予測された将来に対する期待される自由エネルギー関数を最小化することでアクションを選択します。ただし、まだ観測されていない結果に対する期待を追加することは、自由エネルギー汎関数がカルバック・ライブラー構造を持たなくなることを意味し、推論手順のメッセージ パッシング処理を妨げます。我々は、メッセージパッシングによる推論を完全にサポートする、ベーテ自由エネルギー汎関数に基づく代替定式化を提案します。認識論的衝動は、正規化、周縁化、形式制約に次いで情報制約を課すことによって維持され、将来の観察、状態、およびアクションを与えられたパラメータの間の相互情報が少なくとも事前の目標のエントロピーと同じ大きさでなければならないと主張します。対応するカルーシュ・クーン・タッカー乗数の特定の値について、この制約されたベーテ ラグランジアンの定常点は、期待される自由エネルギー解を回復します。情報需要が変化するにつれて、解かれた乗数が不活性領域、内部領域、飽和領域を経て変化することを示します。不活性領域では、エージェントの認識衝動は完全にオフになりますが、飽和領域ではそれが最大になります。 3 つのタスクにおける制約付き Bethe エージェントのパフォーマンスを EFE および Q-MDP と比較します。

原文 (English)

Expected free energy as an information constraint on the Bethe Lagrangian

Active inference selects actions by minimising an expected free energy functional over predicted futures. However, adding an expectation over yet-unobserved outcomes means the free energy functional no longer has a Kullback-Leibler structure, which hinders message passing treatments of inference procedures. We propose an alternative formulation based on a Bethe free energy functional, fully supporting inference by message passing. The epistemic drive is maintained by imposing an information constraint, next to normalisation, marginalisation and form constraints, insisting that the mutual information between future observations, states and parameters given actions must be at least as large as the entropy of the goal prior. For a specific value of the corresponding Karush-Kuhn-Tucker multiplier, the stationary point of this constrained Bethe Lagrangian recovers the expected free energy solution. We show that, as the information demand is varied, the solved multiplier moves through its inactive, interior, and saturated regimes. In the inactive regime the agent's epistemic drive switches off entirely, while in the saturated regime it is maximal. We compare the performance of the constrained Bethe agent on three tasks against EFE and Q-MDP.

13:00 JSTLLM/生成AIOpenAIGPT / ChatGPT

LLM は法的に意味のある方法で推論できるでしょうか?欧州人権裁判所の訴訟に関する小規模研究

推論は、現代の LLM の標準的な技術および機能となっています。ただし、訴訟の予測など、要求の厳しい法律指向のタスクにおけるその適用と品質は、まだ研究されていません。私たちは、欧州人権裁判所 (ECtHR) の訴訟事例をテストベッドとして使用し、訴訟予測の文脈で LLM がどのように推論するかを調査します。私たちは、ECtHR 法学の文脈において法的に意味のある推論とみなされるものを多かれ少なかれ示唆する代替プロンプト戦略を探索することによって、最近の最上位 LLM である OpenAI GPT 5.4 を評価します。人間と LLM の両方の評価によるモデルの応答の評価から得られた結果を示します。我々は、調査されたモデルのスコアが法的推論において理想から程遠いこと、モデルが構造的に完全ではあるが本質的に浅い分析を生成していること、LLM-as-a-Judgeの評価者は内部的には一貫しているものの、訓練されたアノテーターとの整合性が弱いこと、つまり、信頼できるが人間による評価の有効な代替物ではないことを発見した。全体として、専門家が厳選したプロンプトはより包括的な推論につながりますが、調査された他の設定と比較してより正確な予測が得られるわけではありません。私たちの調査結果に基づいて、私たちはコミュニティに対し、自動化された LLM ベースの評価のみに依存しないこと、およびタスクの精度を推論の品質の適切な代用として使用しないことを推奨します。

原文 (English)

Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases

Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We evaluate OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less suggestive of what counts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our findings derived from assessing the model's responses with both human and LLM evaluation. We find that the examined model scores far from ideal in legal reasoning, the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators, i.e., reliable but not a valid substitute for human evaluation. Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not result in more accurate predictions compared to the other examined settings. Based on our findings, we urge the community not to rely solely on automated LLM-based evaluation and to avoid using task accuracy as an appropriate proxy for reasoning quality.

13:00 JST研究/論文

承認ポイントはシステムです: AI 監査証拠のための耐久性のあるポリシー決定受領書

AI 監査レコードは、その耐久性と信頼境界が明確な場合にのみ役立ちます。永続的な書き込みの前に保護された決定を返すと、遅延が最小限に抑えられますが、即時のクラッシュから証拠が生き残ることを保証することはできません。この制約に基づいて RuntimeGuard-AI を再構築します。結果として得られる研究プロトタイプは、各決定論的なポリシー決定を正確なポリシー ソースにバインドし、呼び出し元が選択した同期境界でプライバシーを最小限に抑えるレコードをコミットし、その境界が完了したかどうかを示す Ed25519 署名付きの受領書を返します。再起動後、エンジンはフレーム化されたレコード、マニフェスト、シャードの配置、シーケンスの連続性、およびリプレイ ID を検証します。別の構成証明パスは、コミットされたレコードをチェーン化された署名付きマークル エポックにグループ化し、監査人が外部から取得したキーを使用して検証します。 Apple M4 Pro では、4 つのワーカー スレッドと 2,048 バイトのプロンプトで、バッファリングされた署名付き証拠は 27,193 リクエスト/秒に達し、レイテンシの中央値は 141.9 マイクロ秒です。レコードごとのデータと完全な同期により、スループットは約 242 リクエスト/秒に減少し、レイテンシの中央値は 16.0 ミリ秒に増加します。 100,000 レコードの署名付きエポックをシールするには 97.0 ミリ秒かかります。その結果は、「無料」の非同期監査パスではなく、測定された耐久性とレイテンシーのトレードオフになります。プロトタイプは、モデルの実行を証明したり、侵害された署名者による履歴のフォークを防止したり、法的適合性を確立したりするものではありません。

原文 (English)

The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence

An AI audit record is useful only if its durability and trust boundary are explicit. Returning a guarded decision before any durable write minimizes latency, but it cannot guarantee that evidence survives an immediate crash. We rebuild RuntimeGuard-AI around this constraint. The resulting research prototype binds each deterministic policy decision to the exact policy source, commits a privacy-minimizing record at a caller-selected synchronization boundary, and returns an Ed25519-signed receipt that states whether that boundary completed. After restart, the engine validates framed records, manifests, shard placement, sequence continuity, and replay identity. A separate attestation path groups committed records into chained, signed Merkle epochs that an auditor verifies with an externally obtained key. On an Apple M4 Pro at four worker threads and 2,048-byte prompts, buffered signed evidence reaches 27,193 requests/s with 141.9 microseconds median latency. Per-record data and full synchronization reduce throughput to approximately 242 requests/s and raise median latency to 16.0 ms. Sealing a 100,000-record signed epoch takes 97.0 ms. The result is a measured durability-latency trade-off, not a "free" asynchronous audit path. The prototype does not prove model execution, prevent a compromised signer from forking history, or establish legal conformity.

13:00 JST研究/論文

状況に応じた強化学習のためのタスクの専門化の微調整

コンテキスト強化学習 (CRL) は、関連するタスクのコンテキスト空間全体でタスク カバレッジを最大化することで、古典的な RL を一般化しようとします。これまでの研究では、最初からトレーニングすることが多く、単一のポリシーのマルチタスク学習または複数のポリシーの戦略的なトレーニングに依存していましたが、私たちは統合された代替手段、つまり、良好な初期パフォーマンスで単一のポリシーを事前トレーニングし、その後タスクの特化のために複数のポリシーを微調整することを提唱しています。ただし、この新しいパラダイムは、不均一な限界収益やサンプルの非効率など、特有の課題をもたらします。これにより、研究上の重要な疑問が生じます。事前トレーニングされたポリシーと限られた予算を考慮すると、サンプル効率の高い CRL を有効にするために、各タスク領域をどの程度微調整する必要があるでしょうか?この目的を達成するために、単純なパラメトリック モデルを使用して微調整パフォーマンスを予測し、結果として生じる離散予算割り当て問題を整数線形計画法によって正確に解決するオンライン フレームワークである Task Specialization Fine-Tuning (TSFT) を提案します。組み合わせ最適化、継続的制御、LLM 微調整など、さまざまな意思決定領域にわたる広範な実験により、TSFT がタスク カバレッジにおいてベースラインを大幅に上回り、オラクルのパフォーマンスに近づくことが実証されました。私たちの取り組みは、現代のプレトレインとファインチューンの時代に合わせて、モデルベースの CRL の新しい方向性を示しています。

原文 (English)

Task Specialization Fine-Tuning for Contextual Reinforcement Learning

Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related tasks. While prior works often train from scratch and rely on either multi-task learning for a single policy or strategically training multiple policies, we advocate for a unified alternative: pretraining a single policy with good initial performance, followed by fine-tuning multiple policies for task specialization. This new paradigm, however, introduces unique challenges, such as heterogeneous marginal returns and sample inefficiency. This raises a critical research question: given a pretrained policy and a constrained budget, how much fine-tuning should each task region receive to enable sample-efficient CRL? To this end, we propose Task Specialization Fine-Tuning (TSFT), an online framework that predicts fine-tuning performance with a simple parametric model and exactly solves the resulting discrete budget allocation problem via integer linear programming. Extensive experiments across diverse decision domains, including combinatorial optimization, continuous control, and LLM fine-tuning, demonstrate that TSFT significantly outperforms baselines in task coverage and approaches oracle performance. Our work charts a new direction for model-based CRL, aligning with the modern pretrain-finetune era.

13:00 JSTLLM/生成AIエージェント研究/論文

Token Optimization and Context Window Management in Multi-Agent AI Workflows

Multi-agent AI workflows are limited not only by model quality but by token cost, latency, and context-window quality. This paper presents…

13:00 JSTエージェント

Graphectory Viewer: A Tool for Process-Centric Analysis of Agentic Software Trajectories

We present Graphectory Viewer, a web-based tool for interactive, process-centric analysis of software-agent trajectories. Building on the G…

13:00 JST画像/動画生成エージェントロボティクス

Teach and Grow: An Agent-Centered Architecture for General Robot Learning

End-to-end vision-language-action (VLA) and world-action models offer an elegant route to general-purpose robotics, but their reliability i…

13:00 JSTエージェント

PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance

Autonomous AI agents are emerging as interfaces for decentralized finance (DeFi) actions such as swaps, lending operations, and yield manag…

13:00 JST研究/論文

Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection

Low-cost, scalable screening for dementia remains an open problem. Imaging-based diagnosis is costly and hard to deploy widely. Electroence…

13:00 JSTLLM/生成AI

COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models

Multimodal large language models (MLLMs) are increasingly used to interact with screenshots, scanned documents, diagrams, and other visuall…

13:00 JST画像/動画生成エージェント

Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement

Converting structural framing plans into editable finite-element model drafts remains labor-intensive and prone to transcription error. Exi…

13:00 JST研究/論文

Maximum Tsallis Entropy Distributions for Robust and Efficient Sparse Learning from Correlated Data

This paper addresses the limitations of Gaussian distribution assumptions in statistical sparse learning, particularly in modeling correlat…

13:00 JSTハードウェア/半導体研究/論文

Adaptive surrogate modeling for high-dimensional spatio-temporal output

This paper develops an adaptive surrogate modeling method for problems with very high-dimensional spatio-temporal outputs. The analysis of…

13:00 JST画像/動画生成エージェント

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its stro…

13:00 JST画像/動画生成

Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording…

13:00 JSTハードウェア/半導体研究/論文

Nonadaptive Learning in Robust Nonlinear Output Regulation

This paper considers robust nonadaptive regulation for general nonlinear systems in an output-feedback setting with arbitrarily high relati…

13:00 JST研究/論文

Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics

Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. Ho…

13:00 JSTエージェント

When Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling

AI agents increasingly act rather than merely read: across the Model Context Protocol (MCP) ecosystem, the share of deployed tools that mod…

13:00 JST研究/論文

Rethinking Irregular Time Series Forecasting from the Perspective of Basis Functions

Irregular time series forecasting is crucial in many domains, such as healthcare and meteorological observation. However, due to the inhere…

13:00 JSTビジネス/資金調達研究/論文

Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting

Existing research on irregular time-series forecasting has primarily focused on model design, while evaluation metrics remain insufficientl…

13:00 JST研究/論文

NeuroAbs: A Neuro-Symbolic RTL Abstraction Framework for Property Checking Acceleration

Formal verification is a crucial technique for ensuring the functional correctness of hardware designs. In the context of property checking…

13:00 JSTLLM/生成AI画像/動画生成

Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models

While adversarial prompt tuning can enhance robustness of vision-language models efficiently, we find that existing methods aggravate robus…

13:00 JSTロボティクス

ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback

Robotic manipulation policies trained via imitation learning, such as Action Chunking with Transformers (ACT), can achieve strong performan…

13:00 JST研究/論文

SPACE: Sample-cloud Predictive Adaptive Conformal Ellipsoids for Multivariate Time-Series Forecasting

Modern probabilistic time-series forecasters often express uncertainty through forecast samples. While typically converted into nominal pre…

13:00 JST研究/論文

MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting

Forecasting cryptocurrency prices remains a formidable challenge due to inherent non-stationarity, abrupt regime shifts, and multi-scale st…

13:00 JST研究/論文

Inductively Scalable, Single-Step Neural Surrogates for Wave-Scattering Inverse Problems

Neural network surrogates are an emerging alternative to traditional electromagnetic wave simulators like finite-difference time-domain (FD…

13:00 JSTLLM/生成AIビジネス/資金調達GPT / ChatGPT

Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets

Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) wit…

13:00 JSTエージェントGoogle

Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning

Sample efficiency is a central challenge in reinforcement learning (RL), particularly in image-based domains where agents must learn from h…

13:00 JSTLLM/生成AIハードウェア/半導体研究/論文

PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel…

13:00 JST研究/論文

Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design

Biomolecular design underpins applications from molecular recognition to therapeutics and synthetic biology, yet de novo interaction design…

13:00 JST画像/動画生成研究/論文

SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation

We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires…

13:00 JSTLLM/生成AI

Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task

Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across…

13:00 JST画像/動画生成

Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery

Persistent shortages in the surgical workforce and inherent limitations of traditional training methods highlight the necessity of automate…

13:00 JSTLLM/生成AI

CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method

Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions…

13:00 JST研究/論文

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial sol…

13:00 JST画像/動画生成

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations…

13:00 JSTLLM/生成AI

Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries

Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, th…

13:00 JST研究/論文

DMT-Dens: Density-preserving manifold visualization for biological data

Motivation: Low-dimensional embeddings are widely used to explore cell-state heterogeneity in single-cell and other high-dimensional biolog…

13:00 JSTロボティクス

tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots

In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as cm-sized millirobots to auto…

13:00 JSTエージェント研究/論文

HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety

Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and e…

13:00 JSTLLM/生成AIビジネス/資金調達

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goa…

13:00 JST研究/論文

From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support

Learning analytics models can identify students at risk of poor performance, but they do not directly indicate which interventions are feas…

13:00 JSTロボティクス

Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision

Developing robots capable of understanding and manipulating objects requires compact, interpretable, and generalizable representations. Thi…

13:00 JSTLLM/生成AILlama

DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval

Large language models (LLMs) can both expand underspecified queries and encode text as dense representations, suggesting a unified model fo…

13:00 JSTLLM/生成AIエージェント研究/論文

MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps

LLM-powered GUI agents that autonomously operate smartphones are rapidly transitioning from research prototypes to early real-world deploym…

13:00 JST研究/論文

Benchmarking Automated Security Patch Backporting: How Far Are We?

Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their…

13:00 JSTLLM/生成AI

GADR: Gathering Architecture Decision Records from Meeting Transcriptions

Existing LLM-based approaches to Architecture Decision Record (ADR) generation share a critical and largely unexamined assumption: that inp…

13:00 JSTロボティクス

Dijkstra as an Oracle for Online Stochastic Shortest Path Navigation with Provable Guarantees

Mobile robots that operate in side by side with humans and critical facilities must reach their goals at low cost, despite often unknown tr…

13:00 JSTハードウェア/半導体ビジネス/資金調達GeminiGemmaDeepSeek

Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models

Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although m…

13:00 JSTLLM/生成AIGPT / ChatGPT

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate old…

13:00 JST研究/論文

Learnware for CSI Feedback: Scene-specific Small Models Can Do Big

Intelligent channel state information (CSI) feedback is essential for realizing the high capacity and spectral efficiency goals of future 6…

13:00 JST画像/動画生成ロボティクス

Training with synthetic data for drone detection in thermal imagery

Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging due to reduced texture information…

13:00 JSTLLM/生成AIビジネス/資金調達

Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses

The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ…

13:00 JST研究/論文

MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure

Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive…

13:00 JSTLLM/生成AI

The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges

LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. Howe…

13:00 JSTエージェント

AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis

Large language models are pushing data science toward increasingly autonomous and agentic workflows, with recent systems already supporting…

13:00 JSTLLM/生成AI

Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints

Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it inform…

13:00 JSTLLM/生成AI研究/論文Gemini

BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models

While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-de…

13:00 JST画像/動画生成エージェント

Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition

Automatic Target Detection and Recognition (ATD/R) is critical for military decision support and (semi-)autonomous operations. Recent advan…

13:00 JST研究/論文

Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks

Background and Context: Question and inquiry are integral parts of knowledge seeking and learning. Despite their importance, students tend…

13:00 JSTエージェントロボティクス

A Theoretical Framework for Parallel Lifelong MAPF Using Group Decentralized Planning

In the Lifelong Multi-Agent Path Finding (L-MAPF) problem, agents must repeatedly move from one destination to another while avoiding obsta…

13:00 JST研究/論文

Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints

Groups routinely complete projects that no single member can plan, execute, or verify alone. We propose a formal model of this phenomenon,…

13:00 JSTLLM/生成AI

Grading Needs a Rubric, Not Intelligence

Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against a…

13:00 JSTLLM/生成AI

Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation

Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rol…

13:00 JSTLLM/生成AI研究/論文

SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE

Recent research has leveraged Large Language Models (LLMs) to enhance Automated Feature Engineering (AutoFE) through semantic descriptions…

13:00 JSTLLM/生成AIGPT / ChatGPT

An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models

In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted…

13:00 JST研究/論文

Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection

Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-b…

13:00 JST画像/動画生成

Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity

Ultrasound tongue contour segmentation remains challenging under cross-dataset domain shift, where limited annotations, probe variability,…

13:00 JSTLLM/生成AI

Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media

The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideol…

13:00 JST研究/論文

Traceable Trust for action-ready artificial intelligence in bioscience

Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structur…

13:00 JSTLLM/生成AIエージェント

Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents

Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward sign…

13:00 JSTLLM/生成AIGPT / ChatGPT

Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System

GPT-style models achieve strong performance by representing language with finite vocabularies of reusable discrete tokens. This success has…

13:00 JST画像/動画生成

Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors

MRI reconstruction methods for undersampled k-space data naturally utilize complex-valued measurements. Parallel developments in sparse pha…

13:00 JST画像/動画生成

From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation

Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPT

Evidence of conceptual mastery in the application of rules by Large Language Models

Background. Evidence that large language models (LLMs) reproduce human judgments does not establish conceptual mastery: the correspondence…

13:00 JST画像/動画生成ロボティクス研究/論文

HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions

Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous spaces, with little attention to dynamic, cro…

13:00 JSTエージェントロボティクス

Efficient Dynamic Shielding for Parametric Safety Specifications

Shielding has emerged as a promising approach for ensuring safety of AI-controlled autonomous systems. The algorithmic goal is to compute a…

13:00 JST研究/論文

MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale

Representation learning of geospatial locations remains a core challenge in achieving general geospatial intelligence, with increasingly di…

13:00 JSTLLM/生成AI

LLM Enhancement with Domain Expert Mental Model to Reduce LLM Hallucination with Causal Prompt Engineering

When consequential decisions depend on knowledge that exists nowhere in writing, LLMs hallucinate not from retrieval failure but from model…

13:00 JST画像/動画生成エージェントロボティクス

Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight

Enabling embodied agents to imagine future states is essential for robust and generalizable visual navigation. Yet, state-of-the-art system…

13:00 JST研究/論文

Planning under Distribution Shifts with Causal POMDPs

In the real world, planning is often challenged by distribution shifts. As such, a model of the environment obtained under one set of condi…

13:00 JST研究/論文

Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models

Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a single architecture. While unified arc…

13:00 JSTLLM/生成AI

TSQueryBench: LLM-as-a-Judge for Time Series Explanations

Natural language explanations of time series data are increasingly produced by foundation models in high stakes domains, making factual cor…

13:00 JSTLLM/生成AI

Chronos: The AI Co-Historian

AI is increasingly supporting, accelerating, and automating scientific discovery across subjects. Yet, the adoption of AI in historical res…

13:00 JSTLLM/生成AIGPT / ChatGPTQwenGrok

Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest

Large language models (LLMs) are trained to align with user preferences through methods like reinforcement learning. Yet models are beginni…

13:00 JSTエージェント

Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

Vision-Language Models (VLMs) expand the attack surface of safety-aligned systems by coupling visual perception with text generation. Exist…

13:00 JST画像/動画生成ロボティクス

LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models

Current Vision-Language-Action (VLA) models typically treat the deepest representation of a vision-language backbone as universally optimal…

13:00 JSTエージェント

ScreenSearch: Uncertainty-Aware OS Exploration

Desktop GUI agents operate under partial observability: visually similar screens can correspond to different underlying workflow states, so…

13:00 JSTビジネス/資金調達ClaudeGeminiGrok

AI 認識的従属指数: おべっかの継続的な尺度

現在の AI モデルは認識論的な同調性を示し、ユーザーに同意するという主張を支持することがよくあります。既存の評価では、通常、モデルを二値支持にシフトさせるために何が必要かを評価するか、命題で明示的な確率を導き出すことによって、これを測定します。ただし、ユーザーに対するお調子者行動の多くは、通常の言語で表現される段階的サポートの変化を通じて示されます。私たちは AI Epistemic Deference Index (AEDI) を提案します。これは、モデルの出力で表現されるサポートが、ユーザーのプロンプトで表現される態度に対してどの程度敏感であるかを表す連続的な一次元スコアです。 AEDI を生成するために、人間の判断との一貫性と相関性が検証された判定者としての LLM を使用して、自然言語出力から確率を推定するための新しいプロトコルを提供します。私たちはこれを、さまざまなトピックにわたる 500 の提案と、ユーザーの態度が異なる 16,000 のプロンプトからなる厳選された新しいデータベースに展開し、8 つの著名なモデルをテストしました。どのモデルもかなりの差異を示しますが、プロバイダーごとに大きく体系的な違いがあり、Claude モデルが最も少なく、Grok モデルと Gemini モデルが最も多くなっています。この効果は、書かれたアーティファクトを要求するプロンプトで増幅され、モデルが弱い事前分布を保持する命題に集中します。 AEDI は、出力レベルのおしゃべり評価のための、更新が簡単なベンチマークおよび測定パイプラインとしてリリースされています。

原文 (English)

Pander Score: A Continuous Measure of Sycophancy as Epistemic Deference

Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user. Existing evaluations typically measure this either by assessing what it takes to make a model shift a binary endorsement or by eliciting an explicit probability in a proposition. However, much user-facing sycophantic behavior is demonstrated through shifts in graded support expressed through ordinary language. We propose the Pander Score: a continuous score representing how sensitive the support expressed in a model's output is to the attitude expressed in a user's prompt. To generate the Pander Score, we provide a new protocol for estimating probabilities from natural language outputs, using LLMs-as-judges validated for consistency and correlation to human judgment. We deploy it on a new curated dataset of 349 propositions across diverse topics and over 11,000 prompts varying in user attitude, testing 18 models. Models pander to sharply different degrees. Among current flagship models, Z.ai's GLM-5.2 panders the most and Claude Fable 5 the least, with other models in between. When we run the test on instructional rather than conversational prompts, every model becomes substantially more likely to go along with claims they would push back against in conversation. We release the Pander Score as an easy-to-update benchmark and measurement pipeline for output-level sycophancy evaluation.

13:00 JSTエージェント

証拠に基づいた計算病理学のためのマルチモーダルエージェントコパイロット

病理学は現代医学の基礎であり、正確な意思決定は証拠に基づいた実践に大きく依存しています。人工知能 (AI) は臨床ワークフローを変革する可能性を秘めていますが、AI と科学的根拠に基づいた医療の接点は依然として研究されておらず、原始的な試みはテキストのみの一般医療に限定されています。この研究では、証拠に基づいた病理学のために特別に設計されたマルチモーダル AI エージェントのコパイロットである PathPocket を紹介します。当社は、これまでで最も包括的な病理証拠コーパスを構築しており、臨床ガイドラインから専門家の意見に至る厳格な証拠階層にわたって構造化された約 110,472 件の公的文書および承認済み文書を網羅しています。この綿密にグレーディングされた基盤から、455 万を超えるエンティティと 710 万の関係を含む大規模なマルチモーダル病理学ハイパーグラフを構築します。このハイパーグラフは、堅牢な知識エンジンとして機能し、入力の理解、証拠の検索、フィルタリング、および診断の生成を統合する、協調的なマルチエージェント推論フレームワークに追跡可能な証拠を提供します。これにより、PathPocket は、テキストのみのクエリから、関心領域 (ROI) やギガピクセルの全スライド画像 (WSI) を含む複雑なマルチモーダル診断に至るまで、幅広い臨床タスクをシームレスに解決できるようになります。当社は、200,000 を超える実世界のケースの多次元ベンチマークに基づいてシステムを厳密に評価し、既存の最先端技術を大幅に上回っています。重要なことは、広範なユーザー調査により、PathPocket が病理医の診断精度と信頼性を大幅に向上させることが証明されているということです。 PathPocket は、検証可能な文献に直接基づいて病理学の解釈を行うことにより、証拠に基づいた計算病理学の将来に向けた実用的でスケーラブルなソリューションを提供します。

原文 (English)

A Multimodal Agentic Pathology Co-pilot via Evidence Grounded Reasoning

Pathology is the cornerstone of modern medicine, where accurate decision-making relies heavily on evidence-based practices. While artificial intelligence (AI) has the potential to transform clinical workflows, the intersection of AI and evidence-based medicine remains under-explored, with primitive attempts restricted to text-only general medicine. In this work, we present PathPocket, a multimodal AI agentic co-pilot designed specifically for evidence grounded pathology. We construct the most comprehensive pathology evidence corpus to date, encompassing approximately 110,472 public and authorized documents structured across a rigorous hierarchy of evidence from clinical guideline to expert opinion. From this meticulously graded foundation, we build a large-scale multimodal pathology hypergraph containing over 4.55 million entities and 7.10 million relations. Serving as a robust knowledge engine, this hypergraph provides traceable evidence for a collaborative multi-agent reasoning framework integrating input understanding, evidence retrieval, filtering, and diagnosis generation. This enables PathPocket to seamlessly resolve a wide spectrum of clinical tasks, ranging from text-only queries to complex multimodal diagnostics involving region-of-interest (ROI) and gigapixel whole-slide images (WSIs). We rigorously evaluate the system on a multidimensional benchmark of over 200,000 real-world cases, where it significantly outperforms existing state-of-the-arts. Crucially, extensive user studies demonstrate that PathPocket substantially improves the diagnostic accuracy and confidence of pathologists. By directly grounding pathology interpretations in verifiable literature, PathPocket offers a practical and scalable solution for the future of evidence grounded computational pathology.

13:00 JST研究/論文

ChatPlanner: パーソナライズされた公共交通機関のルーティングのための大規模な言語モデル フレームワーク

多様なユーザーの好みを把握してルーティング アルゴリズムに統合することが難しいため、公共交通システムにおけるパーソナライズされた公共交通ルーティングは依然として課題となっています。この文書では、大規模言語モデル (LLM) を活用して優先順位を認識した公共交通機関のルーティングを可能にする新しいフレームワークである ChatPlanner について説明します。私たちのアプローチでは、検索拡張生成 (RAG) を備えた微調整された LLM を採用して、ルーティング パラメーターを抽出し、自然言語クエリから微妙なユーザーの好みを解釈し、その後、これらの好みを公共交通機関のルーティング アルゴリズムの目的関数に統合します。この研究では、微調整と RAG の両方のスコア基準を確立するために、8 つのペルソナと 5 つのコンテキストを組み込んだ好みを意識したデータセットを設計します。この作業では、ソリューションの実現可能性、ルーティング情報と設定の抽出、ソリューション セットの品質と完全性を検証するために 3 つの実験を実施しました。結果は、ChatPlanner が実行可能なソリューションを確実に生成することを示しています。微調整により、必要な出力構造が強制され、一般的な設定パターンが学習されます。一方、RAG はクエリ固有のコンテキストを提供して、不正確な表現や会話的な表現を解決し、連続スコアを調整します。両方を組み合わせることで、ルーティング情報の抽出とユーザー設定の解釈において最高の精度が実現します。選択されたケーススタディに基づく結果は、ChatPlanner がユーザーの好みを把握することにより、既存のルート プランナーが見落としていたさまざまな側面にわたる価値のあるソリューションを特定し、より価値のあるルートの代替案を生成することを示しています。この研究は、自然言語理解を輸送の最適化に統合するための新しいパラダイムを確立します。

原文 (English)

ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing

Personalized public transit routing in public transit systems remains challenging due to the difficulty of capturing and integrating diverse user preferences into routing algorithms. This paper presents ChatPlanner, a novel framework that leverages Large Language Models (LLMs) to enable preference-aware public transit routing. Our approach employs fine-tuned LLMs with Retrieval-Augmented Generation (RAG) to extract routing parameters and interpret conversationally expressed preferences from natural language queries as preference scores, subsequently integrating these preferences into the objective function of a public transit routing algorithm. This study designs preference-aware datasets incorporating eight personas and five contexts to establish scoring standards for both fine-tuning and RAG. This work conducted four experiments to validate the solutions' feasibility, extraction of routing information and preferences, solution set quality and completeness, and latency and computational tractability. Results demonstrate that ChatPlanner generates feasible solutions reliably. Fine-tuning enforces the required output structure and learns general preference patterns, while RAG provides query-specific context to resolve imprecise or conversational expressions and calibrate continuous scores. The combination of both achieves the highest accuracy in routing information extraction and rubric-consistent user preference interpretation. Results based on selected case studies show that by capturing user conversationally expressed preferences, ChatPlanner identifies preference-relevant solutions across different dimensions that existing route planners overlook, generating more route alternatives. The latency evaluation confirms that the framework is computationally tractable. This research establishes a new paradigm for integrating natural language understanding into transportation optimization.

13:00 JST研究/論文

取得する前に考える: 戦略的計画と自己批判による堅牢なゼロショット合成画像取得

合成画像の検索では、参照画像をテキスト変更命令と統合することによって、ギャラリーからターゲット画像を識別する必要があります。トレーニング不要のゼロショット設定では、このタスクは、推論時の凍結された視覚、つまり言語埋め込み空間内で検索指向のテキスト クエリを構築することに依存します。既存のアプローチは主に、参照コンテキストと変更テキストを統合された記述に融合するシングルパス生成戦略に依存しています。この戦略では、生成中に意味上の歪みや省略を検出または修正することが困難になります。その結果、参照属性の保存とテキスト要件の統合が相互に干渉し、検索精度が低下します。これらの課題に対処するために、クエリ構築を多段階の推論パイプラインとして構造化するトレーニング不要のフレームワークである PEC-CIR を導入します。このフレームワークは、Planner-Executor-Critic アーキテクチャを通じて動作します。Planner は明示的な制約を抽出し、Executor は複数のターゲット記述候補を生成し、Critic は制約の遵守に基づいてこれらの候補を評価します。 PEC-CIR は、クエリ構築をシングルパス出力ではなく段階的推論プロセスとして再構成することで、取得前に候補クエリを明示的に評価することで生成エラーの伝播を削減し、それにより取得の安定性を向上させます。

原文 (English)

Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism

Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification instruction. In a training-free zero-shot setting, this task relies on constructing a retrieval-oriented textual query within a frozen vision--language embedding space at inference time. Existing approaches predominantly rely on a single-pass generation strategy that fuses the reference context and modification text into a unified description. This strategy makes it difficult to detect or correct semantic distortions and omissions during generation. Consequently, the preservation of reference attributes and the integration of textual requirements interfere with each other, which degrades retrieval precision. To address these challenges, we introduce PEC-CIR, a training-free framework that structures query construction as a multi-stage reasoning pipeline. The framework operates through a Planner--Executor--Critic architecture where the Planner extracts explicit constraints, the Executor generates multiple candidate target descriptions, and the Critic evaluates these candidates based on constraint compliance. By reframing query construction as a staged inference process instead of a single-pass output, PEC-CIR reduces the propagation of generative errors by explicitly evaluating candidate queries before retrieval, thereby improving retrieval stability.

13:00 JSTLLM/生成AIエージェント

盲目のキュレーター: 偏った裁判官がどのようにして自己進化エージェントのスキルリタイアを黙って無効にするのか

自己進化するエージェントは、失敗するのを見ることで悪いスキルを克服します。では、裁判官が失敗を見ることができないとどうなるでしょうか?スキルのリタイアは、成長するライブラリがスキルなしのベースラインを下回らないようにする構造的な制約ですが、その保証は不公平な報酬を前提としていますが、これは参照なしのタスクが私たちに強いるLLMのジャッジにとっては誤りです。私たちは、偏った裁判官が単にノイズを加えるだけではないことを示します。 \emph{サイレントにキュレーターのスイッチをオフにします}。私たちはこれを、破損報酬分析と、決定論的報酬に加えて破損を注入することで因果チャネルを分離し、コード生成のクロスチェックを備えた参照不要のレポート作成テストベッドでの行動研究によってこれを正確に行います。対称ノイズによりリタイアメントはそのまま残りますが、\emph{false-pass} バイアス (失敗はパスとしてすり抜けます) により、データ量が超えられない鋭いしきい値を超える寄与ベースのリタイアメントが無効になります。本物のリタイアとキャップエビクションのチャーンを区別すると、この \emph{メカニズム} の失敗は普遍的であり、ドメインと失敗率を超えて保持され、誤合格がほぼゼロの検証者のような採点者だけが助かることがわかります。ただし、下流の \emph{結果} は体制に依存します。評価の品質が低下するのは、同じ破損によってスキル合成が枯渇する場合のみであり、それ以外の場合は安定しているため、無効化されたキュレーターは \emph{沈黙} となり、集計指標には現れません。この貢献は行動の安全性の結果であり、パフォーマンスの結果ではありません。安価な欠陥挿入監査は、導入前にオペレーターに、判断者がしきい値のどちら側を占めているかを伝えます。

原文 (English)

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill baseline, but its guarantee assumes an unbiased reward, which is false for the LLM judges that reference-free tasks require. We show that a biased judge does not merely add noise; it \emph{silently switches off the curator}. We make this precise with a corrupted-reward analysis, then a behavioral study on a reference-free report-writing testbed with a code-generation cross-check, injecting corruption on top of a deterministic reward to isolate the causal channel. Symmetric noise leaves retirement intact, but \emph{false-pass} bias (failures slipping through as passes) disables contribution-based retirement past a sharp threshold (here a false-pass rate of $0.45$) that no amount of data can cross. Separating genuine retirement from cap-eviction churn shows this \emph{mechanism} failure is universal, holding across domains and failure rates and sparing only near-zero-false-pass, verifier-like graders. The downstream \emph{outcome}, though, is regime-dependent: eval quality degrades only where the same corruption also starves skill synthesis, and otherwise holds steady, so the disabled curator is \emph{silent}, surfacing in no aggregate metric. The contribution is a behavioral safety result, not a performance one. A cheap defect-injection audit then tells an operator, before deployment, which side of the threshold their judge occupies.

13:00 JSTLLM/生成AI

FormalAnalyticGeo: マルチモーダル解析幾何問題生成のためのニューラルシンボリックベースのフレームワーク

数学的推論は、マルチモーダル大規模言語モデル (MLLM) の急速な進歩により大幅な進歩を遂げていますが、解析幾何学は、主に注釈付きのサンプルが不足しているため、ほとんど研究されていません。既存のダイアグラム生成アプローチは、解析ジオメトリに苦労しています。テンプレート メソッドは制約駆動のレイアウトを処理できず、生成モデルには注釈付きの円錐曲線を正しくレンダリングするための幾何学的精度が不足しています。私たちは、マルチモーダルな解析幾何学問題を完全に自動生成するためのスケーラブルなフレームワークである FormalAnalyticGeo を紹介します。形式言語の厳密性を活用して、CDL (条件記述言語) を中心としたフレームワークを設計します。これは、自由形式の問題テキストと、符号付き距離フィールド (SDF) エンジンを介した正確な図のレンダリングを橋渡しする形式的な中間表現です。このフレームワークは、4 つの特殊な LLM コンポーネントを順番に使用します。さまざまな解析幾何学問題を生成するジェネレーター、SDF ベースのレンダリング用に各問題を CDL に変換するフォーマライザー、レンダリングされたダイアグラムのビジョンベースの測定を通じてグランドトゥルースの答えを抽出する測定器、および 3 つの段階で出力をチェックする品質検証器です。 Quality Verifier からの構造化されたフィードバックにより自動再試行が行われ、人間による注釈の必要性を排除する閉ループが形成されます。 FormalAnalyticGeo を大規模に適用すると、7K を超える検証済みのマルチモーダル問題のデータセットである AnalyticGeo7K が生成され、それぞれに位置合わせされたテキスト、図、正式な注釈、グラウンド トゥルースが含まれます。実験によると、生成された問題は、グラウンド トゥルース相対誤差の中央値 0.70\% に達し、回答の 82.3\% が正確なシンボリック解の 5\% 以内に収まります。私たちのフレームワークとデータセットは一般に公開されます。

原文 (English)

FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation

Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic geometry remains largely underexplored, primarily due to the scarcity of annotated samples. Existing diagram generation approaches struggle with analytic geometry: template methods cannot handle constraint-driven layouts, and generative models lack the geometric precision to render annotated conic curves correctly. We present FormalAnalyticGeo, a scalable framework for fully automatic generation of multimodal analytic geometry problems. Leveraging the rigor of formal languages, we design the framework around CDL (Condition Description Language), a formal intermediate representation that bridges free-form problem text with precise diagram rendering via a Signed Distance Field (SDF) engine. The framework employs four specialized LLM components in sequence: a Generator that produces diverse analytic geometry problems, a Formalizer that converts each problem into CDL for SDF-based rendering, a Measurer that extracts ground-truth answers through vision-based measurement on the rendered diagrams, and a Quality Verifier that checks outputs at three stages. Structured feedback from the Quality Verifier drives automatic retry, forming a closed loop that eliminates any need for human annotation. Applying FormalAnalyticGeo at scale yields AnalyticGeo7K, a dataset of over 7K verified multimodal problems, each with aligned text, diagram, formal annotation, and ground truth.Experiments show that the generated problems achieve a median ground-truth relative error of 0.70\%, with 82.3\% of answers falling within 5\% of the exact symbolic solution. Our framework and dataset will be publicly released.

13:00 JST研究/論文

JUMP: 微調整された拡散言語モデルのシングルパス メンバーシップ推論

メンバーシップ推論攻撃 (MIA) は、モデルのトレーニング データに候補サンプルが出現したかどうかをテストします。私たちは、微調整された離散拡散言語モデル (dLLM) の MIA を研究します。メンバーシップとは、ターゲット モデルの微調整セットに含まれることを意味します。自己回帰言語モデルとは異なり、dLLM を使用すると、攻撃者は任意のマスク セットを選択し、すべてのマスクされた位置のトークン分布を並行して取得できます。以前の dLLM 攻撃である SAMA は、ランダムにサンプリングされた多数のマスクにわたる再構築信号を平均化するという自然な損失を模倣する戦略に従いますが、任意次数インターフェイスはランダム化としてのみ使用され、多くのターゲット/参照クエリが必要になります。我々は、dLLM の特徴的な特性の両方を利用するシングルパス スコアリング攻撃である JUMP (Joint Uncertainty-Guided Mask Probing) を提案します。つまり、任意次数デコード可能性を使用して参照信頼性の低い位置を選択し、並列デコード可能性を使用して、モデルごとに 1 つのジョイント マスク クエリを通じて選択されたすべての位置をスコアリングします。 JUMP は、選択された位置を結合してマスクし、クリップされたターゲット/リファレンス再構成ギャップ統計を計算します。 6 つの MIMIR ドメインにわたって微調整された LLaDA-8B-Base 上で、JUMP は SAMA と比較して平均 ROC-AUC を 0.82 から 0.90 に改善し、低 FPR 検出を大幅に向上させますが、ターゲット モデルと参照モデルのそれぞれでセレクター パスとスコアリング パスを 1 回だけ必要とします。

原文 (English)

JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models

Public open-weight language models are often fine-tuned on private or domain-specific data before deployment, creating a need to audit whether individual records were used during adaptation. We study this problem for discrete diffusion language models (dLLMs), using the pre-fine-tuning checkpoint as a reference. Unlike autoregressive models, dLLMs allow arbitrary mask sets and return predictions for all masked positions in parallel. SAMA averages reconstruction signals over many random masks, which can dilute informative positions and requires repeated model evaluations. We propose JUMP (Joint Uncertainty-Guided Mask Probing), which selects low-reference-confidence positions, masks them jointly, and aggregates clipped target-reference reconstruction gaps from one scoring query per model. Across six MIMIR domains, JUMP raises mean ROC-AUC from 0.819 to 0.902 on LLaDA-8B-Base and from 0.851 to 0.942 on Dream-v0-Base-7B, while using three model forwards per sample versus 32 for SAMA.

13:00 JSTLLM/生成AIエージェント研究/論文

OmegaUse-OfficeVal: 経済的根拠に基づいた長期的なオフィススイートのタスクに関する LLM エージェントのベンチマーク

大規模言語モデル (LLM) エージェントは、ユーザーのタスク完了を支援することがますます期待されています。ただし、既存のベンチマークでは、エージェントがオフィス スイートのワークフローを妥当なコストで実行できるかどうかを評価するためのサポートが限定的です。タスク レベルの経済的基盤を備えた長期的なオフィス スイート タスクで LLM エージェントを評価するためのベンチマークである OmegaUse-OfficeVal を紹介します。このベンチマークは、実務者によって提案され、プライバシー保護プロセスを通じて調整されたオフィス スイートの要求から派生した 100 のタスクで構成されています。これらのタスクを完了するには、平均して 2.32 時間の人力が必要です。このベンチマークの重要な特徴は、各タスクが 2 つの経済シグナル (人的労働時間とタスク価格の代理) と組み合わされていることです。これらの信号により、人的コストと LLM 推論コストの直接比較、および価値重み付け評価が可能になります。安定した評価をサポートするために、私たちはきめの細かいルーブリックからコードベースの検証ツールを開発します。私たちは、人間のベースラインとともにいくつかのフロンティア LLM を評価します。評価されたすべての LLM は人間の作業者よりも大幅に安価で高速ですが、人間レベルの成果物の品質にはまだ達していません。コードとデータセットは完全にオープンソースであり、詳細についてはプロジェクト Web サイト (https://omegause-officeval.github.io) でご覧いただけます。

原文 (English)

OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding

Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.

13:00 JSTLLM/生成AI

G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution

Deep search has become a fundamental capability of large language models (LLMs) for solving open-domain complex tasks. However, existing ap…

13:00 JSTエージェント研究/論文Gemini

Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems

Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such comm…

13:00 JST研究/論文

VDGR-RAG: 階層的なエンタープライズ ナレッジに対する統合推論に必要なのは、ベクトル、ディレクトリ、グラフ、リフレクションだけです

検索拡張生成 (RAG) は、特に電気通信などの複雑な製品ドキュメントを扱う分野で、企業の知識質問応答 (QA) に不可欠です。しかし、既存の RAG アプローチでは、多様な検索機能の全体的な統合がほとんど見落とされており、不正確なドメイン ルーティング、階層型ドキュメント構造の利用不足、その結果、企業知識に対する推論能力の制限が生じています。これらの制限に対処するために、正確なエンタープライズ ナレッジ QA を実現するための統合フレームワークにベクトル検索、ディレクトリ駆動推論、グラフ トラバーサル、反復リフレクションを統合する VDGR-RAG を紹介します。具体的には、VDGR-RAG はエージェント型 GraphRAG システムであり、最初にドキュメント チャンクから階層的異種ナレッジ グラフ ($\text{H}^2$KG) を構築して階層ディレクトリ構造と意味関係の両方を保持します。次に、$\text{H}^2$KG をナビゲートするために自由に構成できるナレッジ検索用のアトミック ツールのセットを使用します。 (1) ディレクトリ拡張ルーティング目次 (TOC) 構造を使用してユーザーのクエリを適切なドメイン固有の $\text{H}^2$KG にルーティングするツール。 (2) ベクトル検索、TOC ベースのエージェント検索、およびグラフ検索を組み合わせた、包括的な知識検索のためのマルチルート検索ツール。 (3) 知識ローカライゼーションのバイアスを修正するディレクトリ バックトラッキング ツール。 (4) 次の取得フェーズを繰り返し計画する動的リフレクション ツール。当社は、4 つのワイヤレス ドメイン (省エネや障害管理など) にわたるエンタープライズ製品ドキュメントについて広範な実験を行っています。実験結果は、私たちの方法が知識検索再現率と QA 精度の両方の点でさまざまな RAG ベースラインよりも大幅に優れていることを示しています。

原文 (English)

VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge

Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex product documentation like telecommunications. However, existing RAG approaches largely overlook the holistic integration of diverse retrieval strengths, leading to inaccurate domain routing, poor utilization of hierarchical document structures, and consequently limited reasoning capabilities over enterprise knowledge. To address these limitations, we present VDGR-RAG, which integrates vector retrieval, directory-driven reasoning, graph traversal, and iterative reflection in a unified framework for accurate enterprise knowledge QA. Specifically, VDGR-RAG is an agentic GraphRAG system that first constructs a Hierarchical Heterogeneous Knowledge Graph ($\text{H}^2$KG) from document chunks to preserve both hierarchical directory structures and semantic relationships, and then employs a set of atomic tools for knowledge retrieval that can be freely composed to navigate the $\text{H}^2$KG: (1) a directory-enhanced routing tool that uses table-of-contents (TOC) structures to route user queries to appropriate domain-specific $\text{H}^2$KGs; (2) a multi-route retrieval tool that combines vector search, TOC-based agentic search, and graph search for comprehensive knowledge retrieval; (3) a directory backtracking tool that corrects knowledge localization biases; and (4) a dynamic reflection tool that iteratively plans the next retrieval phase. We conduct extensive experiments on our enterprise product documents across four wireless domains (e.g., energy saving and fault management). Experimental results demonstrate that our method significantly outperforms a variety of RAG baselines in terms of both knowledge retrieval recall and QA accuracy.

13:00 JSTLLM/生成AI

上流で決定し、後から書く: 多言語教育省のクロスリンガル拒否回路を見つけて価格設定する

多言語モデルにおける安全性の調整は均一ではありません。英語での有害なリクエストを確実に拒否するモデルは、低リソース言語での同じリクエストに従うことがよくあります。私たちは、インド語と多言語を組み合わせた専門家の推論モデルであるサルヴァムでこのギャップを機械的に追跡し、それが害悪の検出に失敗していないことを発見しました。危害は、ネットワークの中間部で言語にほとんど依存しない内部方向としてエンコードされており ($L11$ での英語対インド語のコサイン ${\約}0.9$)、その方向を上流に誘導することで因果関係を持って拒否を制御します。しかし、検出方向は、実際に拒否を書き込む変更と直交しており、拒否は 1 回の順方向パスで読み取られるのではなく、生成の過程で遅れて組み立てられます。私たちは書き込みを、特定の局所化可能な回路、注意を払う反対者によって抑制されている専門家の混合ライターによるものであると考え、それに介入するあらゆる方法を評価します。反対者を弱めることは安価で効果的ですが、書き込みを増幅することはコストの壁であり、責任ある責任者に対する外科的編集は何の効果もありません。回路の構成とそれを公開する勾配法は、無関係な 2 番目の MoE モデルで繰り返されますが、レバーの強さはアーキテクチャに固有です。その結果、多言語による安全修理がどこに着地できるか、そしてそれにどれくらいの費用がかかるかを示す、コストを測定したマップが作成されます。

原文 (English)

Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE

Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language. We trace this gap mechanistically in sarvam, an Indic-multilingual mixture-of-experts reasoning model, and find it is not a failure to detect harm. Harm is encoded as an internal direction that is nearly language-invariant in mid-network (English-vs-Indic cosine ${\approx}0.9$ at $L11$), and steering that direction upstream causally controls refusal. But the detection direction is orthogonal to the change that actually writes the refusal, which is late and assembled over the course of generation rather than read off in a single forward pass. We attribute the write to a specific, localizable circuit, a mixture-of-experts writer held in check by an attention opposer and price every way of intervening on it: damping the opposer is cheap and effective, amplifying the writer is a cost wall, and surgical edits to the responsible heads do nothing. The circuit's organization, and the gradient method that exposes it, recur in a second, unrelated MoE model, while the lever's strength is architecture-specific. The result is a cost-measured map of where a multilingual safety repair can land, and what it costs

13:00 JST研究/論文

マルチモーダルな感情認識のための理論に基づく学習

会話におけるマルチモーダル感情認識 (MERC) では、言語的手がかりと非言語的手がかりの間の複雑な相互作用を理解する必要があります。しかし、既存のアプローチのほとんどは基本的にこれを直接入出力(マルチモーダルな手がかりと感情ラベル)のマッピング問題として扱い、人間が感情を解釈するときに使用する因果推論を無視しています。我々は、MERC を認知に触発された推論タスクに変換する新しいフレームワークである論理的誘導学習 (RGL) を提案します。二重プロセス理論に基づいて、私たちは感情的推論を 3 つの側面に分解します: 直観的 (即時的知覚、システム 1)、文脈的 (状況分析、システム 2)、および統合的 (両方の総合)。 MLLM オフラインを利用して構造化された根拠を生成し、内部表現を人間のような推論パターンに合わせてモデルのトレーニングをガイドするための記憶としてエンコードされます。最終的なモデルは、推論時に MLLM オーバーヘッドなしで動作します。実験結果は、RGL が IEMOCAP および MELD ベンチマークで最先端のパフォーマンスを達成することを示しています。さらに、解釈のために、モデルの内部機能が、目に見えないテストサンプルの意味的に正しい理論的根拠を効果的に取得することを実証し、その理論的推論能力を検証します。

原文 (English)

Rationale-Guided Learning for Multimodal Emotion Recognition

Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. However, most existing approaches fundamentally treat this as a direct input-output (multimodal cues-emotion labels) mapping problem, overlooking the causal reasoning that humans use when interpreting emotions. We propose rationale-guided learning (RGL), a novel framework that transforms MERC into a cognitively-inspired reasoning task. Based on dual-process theory, we decompose emotional reasoning into three facets: Intuitive (immediate perception, System 1), Contextual (situational analysis, System 2), and Integrative (synthesis of both). We leverage an MLLM offline to generate structured rationales, which are encoded as memories to guide model training via aligning internal representations with human-like reasoning patterns. Our final model operates without any MLLM overheads at inference time. Experimental results show that RGL achieves state-of-the-art performance on the IEMOCAP and MELD benchmarks. Further, for interpretation, we demonstrate that the model's internal features effectively retrieve semantically correct rationales for unseen test samples, validating its rationale reasoning capabilities.

13:00 JST研究/論文GemmaMistral AI

Distribird: ベイジアン モデル キャリブレーションのための文献に基づいた事前分布設計

プロセスベースのモデルのベイズ キャリブレーションには、各モデル パラメーターの事前分布が必要です。何十年にもわたる方法論的な研究にもかかわらず、研究者はほぼ常に均一な事前分布に頼っています。その主な理由は、科学文献から有益な事前分布を構築するのに時間がかかり、分野と統計の両方の専門知識が必要であることです。このプロセスを自動化するエージェント Web アプリケーションである \textbf{Distribird} を紹介します。パラメーター名、物理的説明、およびドメイン コンテキストが与えられると、Distribird は文献を検索し、ドメインの関連性によって報告された値を抽出して重み付けし、AIC モデルの選択を通じて確率分布を適合させるマルチエージェント パイプラインを展開します。利用可能な文献がない場合、システムは賢明で有益ではない代替案に戻り、生成されるすべての事前の背後にある証拠と信頼レベルの両方を明確に報告します。これは、モデルが物理的に解釈可能なパラメーターを持ち、ドメイン知識が出版された文献に存在する問題向けに設計されています。 3 つのオープンウェイト モデル (Qwen3.6 27B、Gemma 4 31B、Mistral Small 4 119B) をシングル プロンプト LLM ベースラインと比較して、10 科学領域にわたる 24 パラメーターでツールを評価します。以前の品質では、パイプライン全体がこのベースラインに \emph{一致}します。すべての事前情報は、それが構築された特定の論文と価値観にまで遡ります。組み込みの妥当性レイヤーは範囲外のリクエストの事前分布の生成を拒否しますが、シングルプロンプトベースラインは、30 モデルパラメータのケースのうち 11 件で、信頼できるが根拠のない事前分布を返します。また、すべての言語モデル呼び出しはローカルで実行されるため、パラメーターの説明や未公開のモデリングの詳細はサードパーティの LLM プロバイダーに送信されません (生成された検索用語のみが公開文献データベースに到達します)。科学的に使用する場合、これらの特性は点推定精度のわずかな改善よりも重要であると私たちは主張します。

原文 (English)

Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration

Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work, researchers almost always fall back on uniform priors. The main reason is that building informative priors from scientific literature is slow and needs both domain and statistical expertise. We present Distribird, an agentic web application that automates this process. Given a parameter name, physical description, and domain context, Distribird deploys a multi-agent pipeline that searches the literature, extracts and weights reported values by domain relevance, and fits a probability distribution via AIC model selection. When no literature is available, the system falls back to sensible uninformative alternatives, and clearly reports both the evidence behind and the confidence level of every prior it produces. It is designed for the problems where the models have physically interpretable parameters, where domain knowledge exists in the published literature. We evaluate the tool on 24~parameters across 10 scientific domains comparing three open-weight models (Qwen3.6 27B, Gemma 4 31B, Mistral Small 4 119B) with a single-prompt LLM baseline. On prior quality the full pipeline matches this baseline. Every prior is traced to the specific papers and values from which it was constructed; a built-in validity layer declines to produce priors for out-of-scope requests, whereas the single-prompt baseline returns confident but unfounded priors for them in 11 of 30~model-parameter cases; and every language-model call runs locally, so no parameter description or unpublished modelling detail is transmitted to a third-party LLM provider (only generated search terms reach the public literature databases). For scientific use, we argue these properties matter more than a marginal improvement in point-estimate accuracy.

13:00 JSTLLM/生成AI

構造ベースの局所改善手法のための LLM ガイド付きグラフ生成

大規模な近傍検索では、通常、反復最適化のために決定変数のランダムなサブセットが選択されます。さまざまな問題を効率的に解決するために、研究者はさまざまなドメインの構造的特徴を考慮して変数選択戦略を設計する傾向があります。このペーパーでは、MiniZinc 形式のすべての問題に対して問題を意識しない自動パイプラインを構築します。 LLM にセマンティック ガイドラインを要求することで、LLM が問題タイプのインスタンスを均一に重み付けされたグラフにマッピングするグラフ ジェネレーターを生成するように導きます。ノードは決定変数を表し、エッジは制約関係を表します。これらの問題に依存しないグラフは、変数選択における構造ベースのローカル改善フレームワーク (SLIM) をガイドします。一方、重み付きグラフを使用すると、すべての問題インスタンスが同じ汎用グラフ表現を共有できるようになり、そこから同じグラフの特徴を抽出して構成の選択に使用できます。 MiniZinc のコンペティション問題 20 件にわたるインスタンスでパイプラインを評価したところ、アルゴリズムの選択により、ワンショットの Gurobi ベースラインに対して問題に重み付けされた平均勝率 39.5% が達成され、これは最良の単一構成 (19.3%) の 2 倍以上であることがわかりました。構成と機能のアブレーションによりパフォーマンスがさらに 44.0% 向上し、LLM ベースのセマンティック生成により、制約を最適化するための効果的な自動構造抽出と機能抽出が可能になることが実証されました。

原文 (English)

LLM-Guided Graph Generation for Structure-Based Local Improvement Methods

Large neighborhood search normally selects a random subset of decision variables for iterative optimization. To efficiently solve various problems, researchers tend to design variable selection strategies that take into account structural features across different domains. In this paper, we build an automatic pipeline that is problem-agnostic to all problems in the MiniZinc format. By prompting an LLM with our semantic guidelines, we guide the LLM to produce a graph generator that maps any instance of a problem type to a uniform weighted graph, where nodes represent decision variables and edges represent constraint relationships. These problem-agnostic graphs guide our structure-based local improvement (SLIM) framework for variable selection. Meanwhile, the weighted graph enables all problem instances to share the same generic graph representation, from which the same graph features can be extracted and used for configuration selection. We evaluated our pipeline on instances across 20 MiniZinc competition problems, finding that algorithm selection achieves a 39.6% average problem-weighted win rate against a one-shot Gurobi baseline, more than doubling the best single configuration (19.3%). A post-hoc configuration and a feature ablation indicate a headroom of up to 44.0%, demonstrating that LLM-based semantic generation enables effective automated structure and feature extraction for constraint optimization.

13:00 JST研究/論文

RETRACE: Resilience-Guided Trait-Conditioned Craving Estimation from Wearable Physiology in Opioid Use Disorder

Detecting opioid craving from wearable physiological signals is critical yet difficult, with the potential to support proactive interventio…

13:00 JSTLLM/生成AI

ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity Search

Most LLM-based automated algorithm design methods optimize a designated component within a human-specified scaffold, fixing overall organiz…

13:00 JSTLLM/生成AIエージェント

Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling

Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking…

13:00 JSTLLM/生成AIビジネス/資金調達GPT / ChatGPT

The Authenticity Gap in Human Evaluation

Human ratings are the gold standard in NLG evaluation. The standard protocol is to collect ratings of generated text, average across annota…

13:00 JST画像/動画生成ビジネス/資金調達Microsoft

Comprehensive framework for evaluation of deep neural networks in detection and quantification of lymphoma from PET/CT images: clinical insights, pitfalls, and observer agreement analyses

This study addresses critical gaps in automated lymphoma segmentation from PET/CT images, focusing on issues often overlooked in existing l…

13:00 JSTロボティクス

ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation

Diffusion models have been verified to be effective in generating complex distributions from natural images to motion trajectories. Recent…

13:00 JST研究/論文

Optimizing Container Loading and Unloading through Dual-Cycling and Dockyard Rehandle Reduction Using a Hybrid Genetic Algorithm

This paper addresses the NP-hard problem of optimizing container handling at ports by integrating Quay Crane Dual-Cycling (QCDC) and dockya…

13:00 JSTLLM/生成AI

LSem2Vec: A Simple yet Effective Two-Stage Approach for Source Code Embedding

The advent of large language models (LLMs) has significantly advanced artificial intelligence in software engineering, with source code emb…

13:00 JST研究/論文

M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction

Accurately predicting the popularity of micro-videos is a critical but challenging task, characterized by volatile, `rollercoaster-like' en…

13:00 JST研究/論文

Diffusion Models for Smarter UAVs: Decision-Making and Modeling

Uncrewed Aerial Vehicles (UAVs) are increasingly used in modern communication networks. However, challenges in decision-making and digital…

13:00 JSTLLM/生成AILlama

Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks

Multi-hop question answering requires a system to identify and integrate evidence distributed across documents, yet large language models r…

13:00 JST研究/論文

Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization

Transformers are difficult to optimize with stochastic gradient descent (SGD) and largely rely on adaptive optimizers such as Adam. Despite…

13:00 JSTLLM/生成AI

MCTS-KBQA: Monte Carlo Tree Search with Information Gain Rewards for Knowledge Base Question Answering

This work investigates how to improve large language model (LLM)-based reasoning for knowledge base question answering (KBQA) via Monte Car…

13:00 JSTLLM/生成AI

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

As large language models (LLMs) increasingly function as human-like assistants exhibiting human-like personality traits, understanding thei…

13:00 JST研究/論文

LZ Penalty: An information-theoretic repetition penalty for autoregressive language models

We introduce the LZ penalty, a penalty specialized for reducing degenerate repetitions in autoregressive language models without loss of ca…

13:00 JST研究/論文

TabularQGAN: A quantum generative model for tabular data synthesis

In this paper, we introduce a novel quantum generative model for synthesizing tabular data. Synthetic data is valuable in scenarios where r…

13:00 JST画像/動画生成エージェント

Multi-Scale Spectral Attention Module-based Hyperspectral Segmentation in Autonomous Driving Scenarios

Recent advances in autonomous driving (AD) have highlighted the potential of hyperspectral imaging (HSI) for enhanced environmental percept…

13:00 JSTLLM/生成AI研究/論文

Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review

In July 2025, 18 academic manuscripts on arXiv contained hidden instructions that manipulated AI-assisted peer review (indirect prompt inje…

13:00 JST研究/論文

Solving nonconvex Hamilton--Jacobi--Isaacs equations with PINN-based policy iteration

We propose a mesh-free policy iteration framework that combines classical dynamic programming with physics-informed neural networks (PINNs)…

13:00 JSTLLM/生成AIロボティクス

Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT

Robotic pick-and-place tasks in convenience stores pose challenges due to dense object arrangements, occlusions, and variations in object p…

13:00 JST画像/動画生成

Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing

Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote s…

13:00 JST研究/論文

ChannelFlow-Tools: A Configuration-Driven Pipeline for Generating Machine-Learning-Ready Datasets of 3D Obstructed Channel Flows

Data-driven surrogate models are increasingly used in computational fluid dynamics, and their reliability depends on the quality of the tra…

13:00 JST研究/論文

Future-Back Threat Modeling: A Foresight-Driven Security Framework

Traditional threat modeling remains reactive-focused on known TTPs and past incident data, while threat prediction and forecasting framewor…

13:00 JST研究/論文

Audio Physical Dynamics Inspired Deepfake Detection for Voice Authentication Systems

Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-pla…

13:00 JST研究/論文

Cluster Aggregated GAN (CAG): A Cluster-Based Hybrid Model for Appliance Pattern Generation

Synthetic appliance data are essential for developing non-intrusive load monitoring algorithms and enabling privacy preserving energy resea…

13:00 JSTLLM/生成AIエージェント

The $\mathbf{P}$-Completeness of Inverted Index Traversal: On the Complexity of Evaluating Boolean Query DAGs

Modern AI agents increasingly rely on search infrastructure to execute complex, neuro-symbolic reasoning workflows. These workflows often c…

13:00 JSTLLM/生成AI

Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries

Large Language Model (LLM)-powered Automatic Speech Recognition (ASR) systems achieve strong performance with limited resources by linking…

13:00 JSTLLM/生成AIビジネス/資金調達

範囲: 選択的コンフォーマル最適化ペアワイズ LLM 判定

大規模言語モデル (LLM) は、ペアごとの評価におけるスケーラブルな判断材料としてますます使用されていますが、依然として誤った調整やバイアスが発生しやすい傾向があります。我々は、交換可能性の下で、非棄権判断間の誤り率が最大でもユーザー指定のレベル $\alpha$ になるように許容閾値を調整するフレームワークである SCOPE (選択的共形最適化ペアワイズ評価) を提案します。 SCOPE にバイアス中立の不確実性信号を供給するために、双方向優先エントロピー (BPE) を導入します。これは、両方の応答位置で裁判官にクエリを実行し、順序平均された優先確率をエントロピー ベースのスコアに変換します。さまざまなペアごとの判定ベンチマーク全体で、BPE は校正と識別において標準信頼度代用を上回っていますが、SCOPE は目標リスク限界 ($\alpha = 0.10$ での経験的な FDR $\約 0.097$ から $0.099$) を一貫して満たしており、実質的なカバレッジを維持しています。バニラのベースラインと比較して、SCOPE は同じリスク制約の下で最大 2.4 倍 $ 多くの判断を受け入れます。これは、BPE が信頼性が高くカバレッジの高い LLM ベースの評価を可能にすることを示しています。

原文 (English)

SCOPE: Selective Conformal Optimized Pairwise LLM Judging

Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and biases. We propose \textsc{Scope} (Selective Conformal Optimized Pairwise Evaluation), a framework that calibrates an acceptance threshold so that, under exchangeability, the error rate among non-abstained judgments is at most a user-specified level $\alpha$. To supply \textsc{Scope} with a bias-neutral uncertainty signal, we introduce Bidirectional Preference Entropy (BPE), which queries the judge under both response positions and converts the order-averaged preference probability into an entropy-based score. Across various pairwise judging benchmarks, BPE outperforms standard confidence proxies in calibration and discrimination, while \textsc{Scope} consistently satisfies the target risk bound (empirical FDR $\approx 0.097$--$0.099$ at $\alpha=0.10$) and retains substantial coverage. Compared to vanilla baselines, \textsc{Scope} accepts up to $2.4\times$ more judgments under the same risk constraint, demonstrating that BPE enables reliable and high-coverage LLM-based evaluation.

13:00 JST研究/論文

Adversarial Data Modeling in Epidemiology

Epidemiological models increasingly rely on crowdsourced, self-reported behavioral data such as vaccination status, mask usage, and social…

13:00 JSTLLM/生成AI

Parametric Knowledge in RAG-SFT for Domain-Specific Document Generation

Retrieval-Augmented Generation (RAG) fine-tuning has shown substantial improvements over vanilla RAG, yet most studies target document ques…

13:00 JSTLLM/生成AI

Supporting Calibrated Reliance in Human-AI Collaboration: Different Strategies for Different Tasks

As AI systems increasingly support human decision making, a central challenge is determining what information helps people recognize when t…

13:00 JSTLLM/生成AI

Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries

Although LLM context lengths have grown, there is evidence that their ability to integrate information across long-form texts has not kept…

13:00 JST画像/動画生成エージェント

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning

Visual Retrieval-Augmented Generation (VRAG) empowers Vision-Language Models to retrieve and reason over visually rich documents. To tackle…

13:00 JST画像/動画生成

SegWithU: Uncertainty as Perturbation Energy for Single-Forward-Pass Risk-Aware Medical Image Segmentation

Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and cli…

13:00 JSTLLM/生成AI

Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting

Large Language Models (LLMs) have achieved remarkable success, underpinning diverse AI applications. However, they often suffer from perfor…

13:00 JST画像/動画生成研究/論文

FairNVT: Fair Classification via Noise Injection in Vision Transformers

This paper presents FairNVT, a lightweight debiasing framework for pretrained transformer-based encoders that improves prediction fairness…

13:00 JSTLLM/生成AI

Convergent Evolution: How Different Language Models Learn Similar Number Representations

Language models trained on natural text learn to represent numbers using periodic features with dominant periods at $T=2, 5, 10$. In this p…

13:00 JST研究/論文

Protect the Brain When Treating the Heart: Feasibility of 2.5D U-Net for Real-Time Gaseous Microemboli Detection

Gaseous microemboli (GME) represent a common complication of cardiac structural interventions across both surgical and transcatheter approa…

13:00 JST研究/論文

Discovering physical mechanisms from experiment-simulation mismatches

Scientific discovery often begins where observation and prediction disagree. As computation and machine learning survey chemical space, exp…

13:00 JSTLLM/生成AIエージェント

SOD: Step-wise On-policy Distillation for Small Language Model Agents

Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and lim…

13:00 JST研究/論文

Adaptive AI Task Partitioning and Safe Offloading in Heterogeneous Edge-Cloud Continuum

In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. However, existing approach…

13:00 JST画像/動画生成

SurgicalMamba: オンラインでの手術段階認識のための状態再プログラミングを備えたデュアルパス SSD

オンライン手術段階認識 (SPR) はコンテキスト認識型の手術室システムを支えており、過去のコンテキストのみからすべてのフレームで予測を行う必要があります。外科用ビデオには、自然ビデオ認識機能が連携して対処できない 3 つの要求があります。手順は数万フレームに及ぶこと、日常的な長いストレッチが短い位相定義の遷移によって中断されるため時間の流れが不均一であること、視覚領域が狭いためバックボーンの特徴がチャネル間で強く相関していることです。既存の認識装置は、フレームごとのコストが経過長に応じて増加するか、コストの制限を維持しながらチャネルに依存しないダイナミクスで均一なレートで状態を進め、後者の 2 つの要求には対処できません。我々は、Mamba2 の構造化状態空間双対性 (SSD) に基づいて構築され、フレームあたりのコストを O(d) に保つ因果的 SPR モデルである SurgicalMamba を紹介します。これらの要求に共同で対処する 3 つの SSD 互換コンポーネントが導入されています。1 つは再発状態のレベルで長期レジームと短期レジームを分離するデュアルパス SSD ブロックです。強度変調ステッピング、低速パスの実効レートを位相関連情報に適応させる連続時間タイムワープ。状態の再グラム化、つまり軸が整列した SSM 反復でクロスチャネル ミキシングを開始するチャンクごとの Cayley 回転です。学習された回転面は、直接の監視なしで位相が揃った構造を継承し、手術ワークフローの解釈可能な内部署名を提供します。 7 つの公開 SPR ベンチマーク全体で、SurgicalMamba は厳格なオンライン評価の下で最先端の精度と位相レベルの Jaccard に達しました。シングルで 238.74 fps で、Cholec80 で 94.6%/82.7% (最強の従来比 +0.7 pp/+2.2 pp)、AutoLaporo で 89.5%/68.9% (+1.7 pp/+2.0 pp) でした。 GPU。アブレーションでは、各コンポーネントの寄与を分離します。コードは https://github.com/sukjuoh/Surgical-Mamba で公開されています。

原文 (English)

SurgicalMamba: Dual-Path SSD with State Regramming for Online Surgical Phase Recognition

Online surgical phase recognition must commit to a prediction at every frame of a procedure that runs for hours, from past frames alone and at a per-frame cost that does not grow with elapsed length. Structured state-space duality (SSD) meets that constraint, but only by having the scan see a per-head scalar transition, which fixes both where the state puts a frame and how fast it decays. The same views recur through an operation, so repeated content is written over itself and can afterwards be told apart only by age. How fast to decay is left to the step, and when the past stops being useful has to be inferred from a loss that never marks the moment. Procedures run long and change little visually from frame to frame, leaving the step with little to select on. Phases also vary widely in length, so no fixed rate serves as a fallback. We address the two with two mechanisms. State regramming rotates the carried state at each chunk boundary, by an amount the chunk's content decides, so where a frame is written also depends on what has passed since: two occurrences of the same view are held apart when different phases intervene, which no decay rate can achieve once both have aged. Intensity-modulated stepping increases the decay at the annotated phase transitions, so the state empties quickly where a phase ends and slowly in between and the decay itself can be set for the longest phase. Both leave SSD's N-semiseparable structure and O(d) per-frame cost intact. Across seven public benchmarks SurgicalMamba reaches state-of-the-art online accuracy and phase-level Jaccard (94.6%/82.7% on Cholec80, 89.5%/68.9% on AutoLaparo) at 312.88 fps on a single GPU. Adding the rotation alone to a plain Mamba2 improves multi-query associative recall (MQAR) wherever the recurrent state is the binding constraint, indicating that the mechanism is not specific to surgical video.

13:00 JSTロボティクス

EXPO-FT: 視覚・言語・行動モデルのためのサンプル効率的な強化学習微調整

新しいタスクを効率的かつ確実に学習する能力は、ロボット工学における基本的な課題です。 Vision-Language-Action(VLA)モデルは、さまざまな操作タスクにわたって強力な一般化を実証していますが、事前トレーニングされたポリシーは、現実世界の展開に必要な信頼性を常に下回っています。強化学習 (RL) 微調整は、このギャップを埋める有望な道を提供しますが、既存のアプローチでは、事前トレーニングされた事前学習を完全に活用せずに最初からトレーニングするか、実際の展開に必要なサンプル効率と成功率を達成せずに VLA を微調整するかのどちらかです。我々は、このギャップを埋める、事前トレーニングされた VLA ポリシーの安定したサンプル効率的な RL 微調整システムである EXPO-FT を紹介します。当社のシステムは、ストリング ライトの配線と点灯のためのプラグの挿入、ビリヤードのボールをポケットに入れる作業、ワイン ボトルに花を挿入する作業など、一連の困難な操作タスクを解決します。それぞれの作業には、高精度、ダイナミックなアクション、およびさまざまな初期状態に対する堅牢性の組み合わせが必要です。当社のシステムは、オンライン ロボット データの平均 19.1 分以内に、評価されたすべてのタスクにわたって完璧なタスク パフォーマンス (30/30 成功) を達成し、以前のスクラッチからの RL アプローチと VLA 微調整アプローチの両方を上回りました。私たちは、ロボット工学における VLA モデルの RL 微調整の広範な採用を促進することを目的として、オープンソース コードベースをリリースします。

原文 (English)

EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models

The ability to efficiently and reliably learn new tasks has been a foundational challenge in robotics. Vision-Language-Action (VLA) models have demonstrated strong generalization across diverse manipulation tasks, yet pretrained policies consistently fall short of the reliability required for real-world deployment. Reinforcement learning (RL) fine-tuning offers a promising path to bridge this gap, but existing approaches either train from scratch without fully leveraging pretrained priors, or fine-tune VLAs without achieving the sample efficiency and success rates that practical deployment demands. We present EXPO-FT, a system for stable, sample-efficient RL finetuning of pretrained VLA policies that closes this gap. Our system solves a suite of challenging manipulation tasks, including routing string lights and inserting the plug to light it up, striking a pool ball into a pocket, and inserting a flower into a wine bottle, each requiring combinations of high precision, dynamic actions, and robustness to varied initial states. Our system achieves perfect task performance (30/30 successes) across all evaluated tasks within an average of 19.1 minutes of online robot data, outperforming both prior RL-from-scratch and VLA finetuning approaches. We release an open-source codebase with the aim of facilitating broader adoption of RL finetuning of VLA models in robotics.

13:00 JST研究/論文Gemini

量子化線形マップのサブガウス性について: AI 支援によるメモ

この短いメモは、座標方向の非線形マッピングの下で​​ガウス ベクトルに束縛された次元に依存しないサブガウス濃度を示します。 Gemini 3.5 Flash によって発見されたこの結果は、条件の良い共分散の下で任意の有界関数に適用されます。このツールを、符号量子化線形マップ $Y = \text{sgn}(Wx)$ に関する Simone Bombari の質問に答えるために適用します。

原文 (English)

On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note

We prove an elementary bounded-differences inequality for functions of non-isotropic Gaussian vectors. Specifically, if $f$ has bounded coordinate differences and $X\sim\mathcal N(\mu,\Sigma)$, then the resulting concentration bound depends on the condition number $\kappa(\Sigma)$. As an application, we answer a question of Simone Bombari concerning the subgaussianity of sign-quantized linear maps $Y=\mathrm{sgn}(Wx)$. In the special case where $f$ is the coordinatewise sign function, an argument was initially suggested to us by Gemini 3.5 Flash without attribution. We subsequently discovered that it closely resembles an earlier argument of Barber and Kolar [Ann. Statist. 46 (2018), Lemma 4.5]. This revision corrects the attribution and documents the episode as an instance of AI-assisted mathematical discovery.

13:00 JST研究/論文

数十年にわたる気候シミュレーションにおける ArchesWeather と ArchesWeatherGen のスキルと安定性の評価

ArchesWeather と ArchesWeatherGen の気候シミュレーション機能を評価します。この 2 つの機械学習モデルは、もともと天気予報用にトレーニングされ、最大 10 日間のリードタイムで評価されました。 ArchesWeather は決定論的モデルであるのに対し、ArchesWeatherGen は ArchesWeather の予測を活用した確率的フロー マッチング モデルであり、アンサンブルベースの不確実性の定量化を可能にします。この研究では、月平均海面水温 (SST) と海氷面積 (SIC) を境界条件として追加の条件付けを使用することで、これらのモデルを強制大気モデルとして機能するように適応させます。 In particular, we follow the AI Model Intercomparison Project (AIMIP) Phase 1 protocol, which, analogous to the Atmospheric Model Intercomparison Project (AMIP), proposes a standardized experimental setup to evaluate the climate skill of ML-based forced atmospheric models.数値気候モデルとの比較、拡張における主要な設計選択を検討するアブレーション研究、強制構成と非強制構成の分析など、これらの条件下での両方のモデルの包括的な評価を示します。元々は天気予報用に開発されたにもかかわらず、ArchesWeather と ArchesWeatherGen の強制構成が安定した長期気候シミュレーションを生成し、安定した年間サイクルを持ち、多くの気候変数のドリフトを捕捉することを実証します。モデルはERA5の気候学、大規模な循環、年々変動を忠実に再現しており、分布の尾部を捉えています。

原文 (English)

Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations

We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for weather forecasting and evaluated up to a 10-day lead time. ArchesWeather is a deterministic model, while ArchesWeatherGen is a probabilistic flow-matching model leveraging ArchesWeather's forecasts, enabling ensemble-based uncertainty quantification. In this work, we adapt these models to act as forced atmospheric models by using additional conditioning on the monthly mean sea surface temperature (SST) and sea ice cover (SIC) as boundary conditions. In particular, we follow the AI Model Intercomparison Project (AIMIP) Phase 1 protocol, which, analogous to the Atmospheric Model Intercomparison Project (AMIP), proposes a standardized experimental setup to evaluate the climate skill of ML-based forced atmospheric models. We present a comprehensive evaluation of both models under these conditions, including comparison against numerical climate models, ablation studies that examine key design choices in the extension, and an analysis of forced versus unforced configurations. Despite being originally developed for weather forecasting, we demonstrate that forced configurations of ArchesWeather and ArchesWeatherGen produce stable long-term climate simulations, have a stable annual cycle, and capture the drift of many climate variables. The models faithfully reproduce ERA5's climatology, large-scale circulations and interannual variability, and they capture the tails of the distributions.

13:00 JSTエージェント研究/論文

FVSpec: Real-World Property-Based Tests as Lean Challenges

We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 propert…

13:00 JST画像/動画生成

BRo-JEPA: Learning Modular Transformations in Latent Space

Can neural networks learn algebraic rules from visual inputs, or do they merely fit observed patterns? We study this question using MNIST (…

13:00 JSTLLM/生成AIエージェント

Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents

Large language model (LLM) agents can persist across tasks, acquire memory, activate Skills, synthesize tools, fork child processes, attach…

13:00 JST画像/動画生成エージェントロボティクス

Planning-aligned Token Compression for Long-Context Autonomous Driving

Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences t…

13:00 JST研究/論文

The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics

As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and c…

13:00 JSTエージェントロボティクス

Physics-Grounded Causal Auditing of End-to-End Driving Planners

End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that me…

13:00 JSTLLM/生成AIGemma

How Transparent is DiffusionGemma?

LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging su…

13:00 JST研究/論文GemmaLlamaQwenDeepSeek

Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map

Can a platform tell, before deployment, whether an open-weight checkpoint has had its refusal mechanism stripped? Runtime guards cannot: th…

13:00 JSTLLM/生成AI

Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment

Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which ine…

13:00 JST画像/動画生成NVIDIA

Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference

Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly unders…

13:00 JSTLLM/生成AI

From Adoption to Deployment: A Qualitative Study on AI Integration in Software Development Practice

The increasing adoption of Large Language Models (LLMs) as AI components in modern software systems introduces distinct security risks to t…

13:00 JSTLLM/生成AIAnthropic

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining intervent…

13:00 JST画像/動画生成エージェントロボティクス

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA

Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deploy…

13:00 JST研究/論文

Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning

As "AI Scientists" emerge to drive research via the Model Context Protocol (MCP), systems relying on ephemeral scripts will fail. The sheer…

13:00 JSTLLM/生成AIエージェント

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent

Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are…

13:00 JSTLLM/生成AIエージェント

Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards

Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. E…

13:00 JST画像/動画生成Qwen

What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems

Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only…

13:00 JST画像/動画生成

FUSE: Frame-Unified Stress Estimation from Facial Video

Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approache…

13:00 JST研究/論文

A 12-CNOT Double Qubit Excitation Gate

In this work, we presented, to the best of our knowledge, the first reported 12-CNOT decomposition of the double qubit excitation operator.…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文

AQuA: 再帰的自己改善型定量取引リサーチ エージェント

私たちは定量的投資研究のレベルで再帰的自己改善を研究します。つまり、自律システムが初期の実験からの証拠を使用して、後の反復で提案される仮説や候補を改善できるかどうかを研究します。我々は、2 つの個別の言語モデル駆動型研究システムで構成される AQuA を紹介します。1 つは記号因子発見用、もう 1 つは訓練可能なモデル開発用です。 2 つのシステムは、エージェント、メモリ、候補空間、研究状態を共有しません。代わりに、それぞれが独立して、検証された証拠を保持し、それを次の提案の指針として使用することで、独自の研究ループを閉じます。この限定された意味で、両方のシステムは研究プロセスのレベルで再帰的な自己改善を実装します。各システムは、独自の密閉されたサンドボックスも使用します。これにより、データ分割、特徴とラベルの定義、およびエバリュエーターが修正され、制約された因子式または構成差分を通じてのみモデルが動作できるようになります。マネージャーが仲介するマルチエージェント パイプラインであるファクター システムは、ファクターを検出してシグナルに結合し、暗号通貨ユニバースでの結合情報係数が約 $0.190 に達します。ハイブリッド時系列アーキテクチャ上の構成主導型ループであるモデル システムは、米国株の銘柄ごとの情報係数 $+0.0843$ に達し、それを 2 レッグ コストで最大 $+2.50$ のホールドアウト シャープを持つ閾値ロング/ショート戦略に変換します。この戦略は 2021 年から 2025 年まで毎年前向きです。

原文 (English)

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises two separate language-model-driven research systems: one for symbolic factor discovery and one for trainable model development. The two systems do not share agents, memories, candidate spaces, or research state. Instead, each independently closes its own research loop by retaining validated evidence and using it to guide subsequent proposals. In this bounded sense, both systems implement recursive self-improvement at the level of the research process. Each system also uses its own sealed sandbox, which fixes the data splits, feature and label definitions, and evaluator while allowing the model to act only through constrained factor expressions or configuration diffs. The factor system, a manager-mediated multi-agent pipeline, discovers and combines factors into a signal that reaches a combined information coefficient of about $0.190$ on a crypto universe. The model system, a config-driven loop over a hybrid time-series architecture, reaches a per-stock information coefficient of $+0.0843$ on US equities and converts it into a threshold long/short strategy with a held-out Sharpe of up to $+2.50$ at a two-leg cost. The strategy is positive in every year from 2021 to 2025.

13:00 JSTLLM/生成AIGemma

AI の言語表現における虚偽と不可能性は方向性が異なる

言語は、誤った状況や、まったくあり得ない状況を説明することができます。 AI モデルがこれらの障害を内部的に区別するかどうかはまだ不明です。私は、17 の哲学ファミリーからの 85 のプロンプトと、それぞれが真実、偶然の虚偽、ありそうもない主張、意味論的異常、および必要な虚偽として表現される 15 のトピックからなるトピックに一致するモダリティ セットを使用した、マルチモーダル オープンウェイト モデル Gemma 3 4B IT の探索的活性化研究を報告します。その回答では、モデルは偶発的な虚偽と矛盾を混同し、15 件の虚偽発言のうち 12 件に「矛盾」とラベルを付けています。その活性化は異なるパターンを示します。線形真実調査は、不可能と真の陳述 (AUC 0.93) を分離しますが、虚偽の陳述 (AUC 0.20) は不可能ではありません。保留されたトピックファミリーで評価された不可能性プローブは、AUC 1.00で必要な虚偽と偶発的な虚偽を分離し、レイヤー15でピークに達し、バランスのとれた精度0.97(ボンフェローニ調整P=0.018)でした。真実と不可能の方向は直交に近いですが、不可能の方向は意味異常の方向と部分的に重なっていますが、それとは区別できます。同じレイヤーにあるスパース オートエンコーダー フィーチャは、このジオメトリを繰り返します。不可能性を選択する機能は、異常な文章に対しても発動しますが、偶発的な虚偽に対してはまれです。このモデルの活性化空間では、必要な虚偽は偶発的な虚偽の極端なケースではなく、実験的に定義された意味論的異常のカテゴリに近いものになります。この表現上の近接性は、不可能な記述が本質的に無意味であることを意味するものではありません。 1 つの小さなモデルから得られたこれらの相関関係の観察は、古い哲学的な区別に対する経験的な脚注を提供します。

原文 (English)

Falsehood and Impossibility Are Different Directions in an AI's Representation of Language

Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model internally distinguishes these failures remains unclear. I report an exploratory activation study of the multimodal open-weight model Gemma 3 4B IT using 85 prompts from 17 philosophical families and a topic-matched modality set of 15 topics, each expressed as a truth, contingent falsehood, improbable claim, semantic anomaly, and necessary falsehood. In its answers, the model conflates contingent falsehood with contradiction, labeling 12 of 15 false statements "contradiction." Its activations show a different pattern. A linear truth probe separates impossible from true statements (AUC 0.93) but not impossible from false statements (AUC 0.20). An impossibility probe evaluated on held-out topic families separates necessary from contingent falsehood at AUC 1.00, peaking at layer 15 with balanced accuracy 0.97 (Bonferroni-adjusted P=0.018). The truth and impossibility directions are close to orthogonal, whereas the impossibility direction partially overlaps a semantic anomaly direction while remaining distinguishable from it. Sparse autoencoder features at the same layer repeat this geometry. Features selective for impossibility also fire on anomalous sentences but rarely on contingent falsehoods. In this model's activation space, necessary falsehoods are not extreme cases of contingent falsehood but lie closer to the experimentally defined category of semantic anomaly. This representational proximity does not imply that impossible statements are intrinsically meaningless. These correlational observations from one small model offer an empirical footnote to an old philosophical distinction.

13:00 JST画像/動画生成

NaviDC-OCR: デジタル文書およびカメラでキャプチャされた文書にわたる文書解析のナビゲート

文書解析の目的は、非構造化文書を構造化された機械可読表現に変換することです。視覚言語モデル (VLM) の最近の進歩により、文書解析が大幅に進歩しました。しかし、既存のアプローチは依然として 2 つの大きな課題に直面しています。まず、分離された VLM ベースの手法は正確なレイアウト分析に大きく依存しており、カメラで撮影したドキュメントの幾何学的歪みが連鎖的なエラーを引き起こす可能性があります。第 2 に、エンドツーエンドの VLM ベースの方法は明示的なレイアウト検出への依存を軽減しますが、高解像度のシナリオでは冗長な生成、幻覚、不十分な構造的推論の問題が発生することがよくあります。これらの課題に対処するために、私たちは文書解析のための統一フレームワークである NaviDC-OCR を提案します。 NaviDC-OCR は、幾何学的認識を VLM に組み込むための変形認識学習を導入し、複雑なレイアウト表現のための適応サンプリング メカニズムを提案します。さらに、内容と構造を分離した学習戦略が開発され、数式文法とテーブル構造を明示的にモデル化し、より効果的な構造化表現学習を可能にします。広範な実験により、NaviDC-OCR がさまざまなドキュメント解析ベンチマークにわたって最先端のパフォーマンスを達成することが実証されました。 OmniDocBench v1.6、Wild-OmniDocBench、PureDocBench でそれぞれ 96.87、88.53、78.41 の総合スコアを獲得し、ICDAR 2026 Sci-ImageMiner Challenge で 1 位にランクされています。これらの結果は、複雑な文書解析シナリオにおける NaviDC-OCR の有効性と一般化機能を検証します。

原文 (English)

NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents

Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-Language Models (VLMs) have significantly advanced document parsing. However, existing approaches still face two major challenges. First, decoupled VLM-based methods heavily rely on accurate layout analysis, where geometric distortions in camera-captured documents can introduce cascading errors. Second, although end-to-end VLM-based methods alleviate the dependence on explicit layout detection, they often suffer from redundant generation, hallucinations, and insufficient structural reasoning in high-resolution scenarios. To address these challenges, we propose NaviDC-OCR, a unified framework for document parsing. NaviDC-OCR introduces deformation-aware learning to incorporate geometric perception into VLMs and proposes an adaptive sampling mechanism for complex layout representation. Furthermore, a content-structure decoupled learning strategy is developed to explicitly model formula grammars and table structures, enabling more effective structured representation learning. Extensive experiments demonstrate that NaviDC-OCR achieves state-of-the-art performance across diverse document parsing benchmarks. It obtains overall scores of 96.87, 88.53 and 78.41 on OmniDocBench v1.6, Wild-OmniDocBench, and PureDocBench, respectively, and ranks first in the ICDAR 2026 Sci-ImageMiner Challenge. These results validate the effectiveness and generalization capability of NaviDC-OCR in complex document parsing scenarios.

13:00 JST画像/動画生成

UltraArUco: A Lightweight Multilingual Library And Framework With Low-Latency Real-Time Marker-Based Tracking System For Mobile AR Interaction

UltraArUco is a lightweight multilingual library and framework for low latency, realtime marker-based tracking in mobile augmented reality.…

13:00 JSTLLM/生成AIエージェント

CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA

Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modali…

13:00 JST画像/動画生成

HMS-SCP: Task-Oriented Multi-Scale Semantic Communication for V2X Cooperative Perception

Cooperative perception enables vehicles and infrastructure to exchange sensor data via Vehicle-to-Everything (V2X) communication, extending…

13:00 JSTLLM/生成AIエージェント

Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce

Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much…

13:00 JST研究/論文

NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption

Post-hoc Graph Neural Network (GNN) explainers commonly follow a Perturb-Query paradigm, inferring the importance of graph elements based o…

13:00 JSTLLM/生成AIエージェント

Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm

Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional…

13:00 JSTLLM/生成AIエージェント

LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents

LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant e…

13:00 JST画像/動画生成

OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations

Despite comprising over 70% of its surface, the world's oceans are critically underobserved compared to the land surface or the atmosphere.…

13:00 JSTLLM/生成AIエージェント

Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning

Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixtu…