Skip to the content.

AIニュース 2026-07-10

自動生成: 2026-07-10 12:49 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. GPT-5.6: Frontier intelligence that scales with your ambitionOpenAI

    More intelligence from every token, stronger performance per dollar,…

  2. ChatGPT is now a partner for your most ambitious workOpenAI

    ChatGPT Work is an agent that can take action across your apps and fi…

  3. GPT-5.5 Bio Bug BountyOpenAI

    Details about the OpenAI Bio Bounty program

  4. OpenAI「GPT-5.6」一般公開 最上位モデルは「コスト半分で一部Fable 5超え」うたうITmedia AI+

    OpenAIが新AIモデル群「GPT-5.6」を一般公開した。最上位の「Sol」は、Anthropicの「Claude Fable 5」を…

  5. Claude、利用制限を全リセット 競合「GPT-5.6」公開と同日……OpenAI幹部「ビビってるね」ITmedia AI+

    OpenAIが「GPT-5.6」を一般公開した同日、AnthropicがClaude全ユーザーの利用制限を一斉リセット。OpenAI幹部の…

  6. OpenAI、Codexと「ChatGPT Work」の利用制限リセット 競合Claudeとリセット合戦にITmedia AI+

    Anthropicは同日、OpenAIの制限リセットに先立ち、Claudeの利用制限をリセットしていた。

  7. OpenAI launches its new family of models with GPT-5.6TechCrunch AI

    OpenAI's latest family of models promises improvements across a range…

トピック別件数

日本語メディア17件

ITmedia AI+ (日本語)

12:13 JSTLLM/生成AIエージェントAnthropicClaudeOpenAIGPT / ChatGPT

OpenAI、Codexと「ChatGPT Work」の利用制限リセット 競合Claudeとリセット合戦に

Anthropicは同日、OpenAIの制限リセットに先立ち、Claudeの利用制限をリセットしていた。

12:09 JSTLLM/生成AIAnthropicClaudeOpenAIGPT / ChatGPT

OpenAI「GPT-5.6」一般公開 最上位モデルは「コスト半分で一部Fable 5超え」うたう

OpenAIが新AIモデル群「GPT-5.6」を一般公開した。最上位の「Sol」は、Anthropicの「Claude Fable 5」を超える性能を推定コスト半分程度で実現するとうたう。

11:30 JST規制/政策

デジタル庁、tsuzumiなど国産AIを「さくらのクラウド」で稼働 「日本の自律性確保」目指す

デジタル庁は、政府職員が利用するAI基盤「源内」の実証実験の一環として、国産AIモデルを国産クラウドで稼働させると発表した。さくらインターネットが提供するクラウドサービス「さくらのクラウド」を活用する。

11:00 JSTLLM/生成AI

生成AI「若手有利」は大間違い? ミドル層が勝ち抜くための“力の見せ所”とは

新しいテクノロジーは若い世代のもの──生成AI時代において、この構造が大きく変わろうとしている。つまり、業界に長く向き合ってきたミドル層にとって、大きなチャンスといえる。

09:38 JSTLLM/生成AIAnthropicClaudeOpenAIGPT / ChatGPT

Claude、利用制限を全リセット 競合「GPT-5.6」公開と同日……OpenAI幹部「ビビってるね」

OpenAIが「GPT-5.6」を一般公開した同日、AnthropicがClaude全ユーザーの利用制限を一斉リセット。OpenAI幹部の「ビビってるね」という煽り返信も話題だ。

07:45 JSTLLM/生成AIAnthropic

NECが3年で100億円狙うAnthropic協業の初ソリューション、AIが販売戦略立案

NECは、Anthropicとの協業第1弾として、購買データからの商品企画や販促プラン作成を完全自動化する新サービスを開始した。専門人材なしで迅速な施策立案を可能にし、3年間で100億円の売上を目指す。

07:20 JSTエージェント2媒体が報道

Meta、マルチモーダル推論モデル「Muse Spark 1.1」公開 低価格の「Meta Model API」も提供へ

Metaは、AIモデル「Muse Spark 1.1」を発表した。初代モデルを強化したマルチモーダル推論モデルで、ツール操作や複雑なコーディングなどのエージェント能力が向上している。新設の「Meta Model API」を通じてパブリックプレビューとして外部向けに提供され、競合…

出典:ITmedia AI+TechCrunch AI
07:00 JSTその他

AIに企業理念を宿せるか? スポーツ小売り・ヒマラヤ、接客ノウハウまで学習した「AI副店長」開発の舞台裏

スポーツ用品店をチェーン展開するヒマラヤ(岐阜市)は「アイダ つなぐ」と名付けた“AI副店長”を6月から全99店舗に導入した。同社店舗運営部の金丸智功氏(部室長代理)とヒマラヤスポーツ本館店長の高島雅央氏に、AI副店長開発の舞台裏を聞いた。

07:00 JSTその他

「Notion」でAI利用量が4倍に スター精密が全社にAIを定着させた業務基盤の作り方

スター精密は、国内約550人を対象に「Notion」と「Notion AI」を導入し、情報基盤を刷新した。情報を一元化するとともにAIを業務へ組み込むことで、利用率約90%、AI利用量は従来のチャット型AIツール比で約4倍を達成した。

07:00 JSTLLM/生成AI

コスト削減にAIは効果なし? 利益が変わらない企業が過半、手直しも課題に

AI活用企業507人調査で、売上高拡大を目的にAIを導入した企業の約7割が営業利益増を実感した。一方、コスト削減を目的とした導入では利益不変が過半となり、プロンプト格差や手直し、データ連携制約などの課題も浮かび上がった。

07:00 JSTロボティクス

「誰にも会わずに帰る店」の寂しさ すかいらーくがロボット配膳の先に挑むAI接客

デジタル化が進み、人と接することなく食事を終えられる飲食店が増えている。その利便性の裏で失われつつある「人ならではの価値」をどう取り戻すのか。すかいらーくホールディングスは、AIを人の代替ではなく、人の価値を引き出す道具として活用し始めている。

05:00 JST規制/政策

1万9000人が利用するソフトバンクの「全社RAG基盤」 構築の泥臭い舞台裏

AI活用で激突する「現場の利便性」v.s.「会社の安全性」。RAGの乱立に直面したソフトバンクが、ガバナンスをシステムに組み込み、数万時間相当の業務削減効果(社内の試算による)を達成した「全社RAG基盤」構築の舞台裏と、そこから得られた気付きを共有します。

04:36 JSTLLM/生成AIエージェントOpenAIGPT / ChatGPT

デスクトップ版ChatGPT大幅刷新 AIエージェント「Codex」統合、「ChatGPT Work」に

米OpenAIが、チャットAIサービス「ChatGPT」とAIコーディングエージェント「Codex」を統合した「ChatGPT Work」を発表した。「GPT-5.6」シリーズを搭載する。最終成果物を指定するだけで、あいまいな状況にも適応しながら最小限の指示で洗練されたアウトプ…

20:22 JSTロボティクス

三菱自動車が「国産人型ロボ」量産へ 2027年に「月1000台の製造体制」 東大発スタートアップと協業

三菱自動車工業は、東京大学発のロボット開発スタートアップHighlandersと「国産人型ロボット」の開発に関して協業すると発表した。

16:18 JSTその他

AI活用、最大のボトルネックは「経営層」か トップ不使用の企業、85.7%が「方針・体制なし」

ラクスルが中小企業を対象に実施したAI活用の実態調査で、経営層がAIを全く使わない企業の85.7%は活用の方針も推進体制もないと分かった。経営層の姿勢がAI格差を左右する構図が浮かんだ。

15:11 JSTLLM/生成AIAnthropicClaude

NEC、Claudeを活用した「全自動」マーケティングサービス開始 3年で売上100億円目指す

NECは、消費者の購買データから商品企画や販促プラン作成の「完全自動化」をうたうサービスの提供を開始した。AnthropicのAI「Claude」などを活用した。3年間で累計100億円の売り上げを目指す。

13:00 JSTエージェントハードウェア/半導体MicrosoftNVIDIA

社内のWindows環境で「数百のAIエージェント」を隔離実行 NVIDIAとMicrosoftが共同開発したデスクサイドマシンの全容

NVIDIAは、Windows環境でAIエージェントを開発・実行するデスクサイドAIスーパーコンピュータ「NVIDIA DGX Station for Windows」を発表した。NVIDIA GB300 Grace Blackwell Ultra Desktop Superc…

海外メディア19件

TechCrunch AI (英語)

09:16 JSTLLM/生成AIOpenAIGPT / ChatGPTMicrosoftCopilot2媒体が報道

OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter

OpenAI's new family of models will continue to power Microsoft's suite of workplace and productivity apps.

出典:OpenAITechCrunch AI
08:38 JSTLLM/生成AIOpenAI

Fidji Simo steps down from OpenAI’s no. 2 role

OpenAI's No. 2 executive, Fidji Simo, is stepping down from her full-time role after her medical leave proved longer than expected — a lead…

07:24 JSTLLM/生成AIOpenAIGPT / ChatGPT

OpenAI launches its new family of models with GPT-5.6

OpenAI's latest family of models promises improvements across a range of areas, including cybersecurity.

07:08 JSTエージェントビジネス/資金調達

An AI agent startup just let its agent run its $100M fundraise

Lyzr, a startup that builds AI agents for enterprises, used its own AI agent to raise a $100 million round — proof, evidently, that the pro…

07:03 JSTLLM/生成AIエージェントOpenAI

OpenAI is shutting down Atlas, but its AI browser ambitions are still growing

OpenAI is sunsetting its AI-powered browser after less than a year. But it's moving some agentic browsing features to its desktop app and a…

06:57 JSTLLM/生成AIビジネス/資金調達Anthropic

Elon Musk praises Mythos/Fable, promises not to ‘cut off’ Anthropic

Should Anthropic trust Elon Musk to host its models? With about $40 billion in revenue at stake, Musk insists that the company can.

06:47 JSTその他

Can AI answer the $3 trillion question?

The AI ROI debate has returned and the numbers are even bigger, as are, perhaps, the consequences.

04:05 JSTLLM/生成AIハードウェア/半導体規制/政策OpenAIGPT / ChatGPT

New York Times says OpenAI hid evidence in ChatGPT copyright trial

News publishers say OpenAI hid tools and datasets that could identify copyrighted journalism in ChatGPT outputs, escalating their lawsuit w…

03:40 JSTその他Google

Google will now disclose which ads are made with AI

While Google prohibits misleading and deceptive ads, an ad can still leverage AI to create some type of synthetic or digitally altered cont…

03:34 JSTハードウェア/半導体ビジネス/資金調達NVIDIA

Paris-based AI voice startup Gradium raises $100M seed, backed by Nvidia

The company is using the cash to open an office in the Bay Area and compete for talent there, "strengthening its position at the heart of t…

03:22 JSTLLM/生成AIAnthropicOpenAI

How did the government decide OpenAI’s frontier model was safe to release?

"Exactly what that dialog looked like between the government and Anthropic and OpenAI is unclear."

02:56 JSTその他

Instagram users: Here’s how to stop Meta’s AI from using your photos

Muse Image allows users to generate AI images using photos from public Instagram accounts. As long as a person's profile is public, another…

02:17 JSTハードウェア/半導体

Meta’s new AI chips will begin production in September

The company is taking a modular approach to designing these chips, anticipating that their needs will change as AI evolves rapidly by the t…

02:06 JSTハードウェア/半導体NVIDIA

Nvidia is a victim of the compute marketplace it created

Having proven how valuable compute can be, the company finds itself at the center of a market everyone wants to be in — while simpler techn…

23:53 JSTLLM/生成AIAnthropicClaude

Anthropic’s new Claude feature is quietly selling you on AI

Claude’s new Reflect dashboard doesn’t just visualize how you use AI. It also subtly reinforces how much of your daily work now depends on…

23:51 JSTLLM/生成AIビジネス/資金調達AnthropicOpenAI

Anthropic, OpenAI, and SpaceX are bigger than the last 25 years of tech exits

Three big AI IPOs are set to generate more value than all the U.S. VC-backed exits since 2000.

22:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Popular open source AI developer tool Ollama raises $65M, grows to nearly 9M users

Benchmark-backed Ollama has amassed 176,000 stars, and nearly 17,000 forks on GitHub by helping developers easily run AI on their PCs.

22:00 JSTその他

Character.AI enters the microdrama arena with its own productions, but there’s a twist

In an interesting twist that takes advantage of the company's core product, users can chat with these shows' characters, ask them questions…

21:00 JSTその他

Nandan Nilekani leaves GP role at Fundamentum as it launches $200M third fund

Nilekani remains Fundamentum's anchor investor as the firm expands its leadership team and targets AI and fintech startups in India.

公式ブログ3件

OpenAI (英語)

19:00 JSTLLM/生成AIエージェントGPT / ChatGPT

ChatGPT is now a partner for your most ambitious work

ChatGPT Work is an agent that can take action across your apps and files, stay with a project for hours if needed, and turn a goal into fin…

19:00 JSTLLM/生成AIOpenAIGPT / ChatGPT

GPT-5.5 Bio Bug Bounty

Details about the OpenAI Bio Bounty program

19:00 JSTLLM/生成AIGPT / ChatGPT

GPT-5.6: Frontier intelligence that scales with your ambition

More intelligence from every token, stronger performance per dollar, and more capability on demand for your hardest work.

論文231件

arXiv cs.AI (英語)

13:00 JSTエージェントビジネス/資金調達研究/論文

AgentLens: コーディング エージェント評価のための実稼働環境で評価された軌跡レビュー

ここでは、対話型コード エージェントの実稼働環境で評価されたベンチマークである AgentLens を紹介します。ほとんどのコード エージェント ベンチマークでは、実行が 1 ビットに削減されます。タスクは成功しましたか? -- しかし、これらのエージェントを実際に使用する人々は、エージェントがどのように指示に従い、ツールを使用し、自身の作業を検証し、間違いから回復し、途中でエージェントに話しかけるかという軌跡全体を経験します。 AgentLens はその軌跡全体を評価します。客観的なチェックが存在する正式な検証と、LLM で作成された軌跡のレビューおよび並べての比較を組み合わせることで、各実行でスコアがなぜそのようになるのかについての読みやすい説明が得られます。これにより、AgentLens はモデルのランク付け以上の用途に役立ちます。モデルの動作を診断し、独自のエージェントの連続バージョンを比較し、夜間の評価パイプラインで製品の回帰を捕捉するために使用されます。 https://github.com/agent-lens/agent-lens-bench でベンチマークをオープンソースとしてリリースします。

原文 (English)

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use these agents experience the entire trajectory: how the agent follows instructions, uses its tools, verifies its own work, recovers from mistakes, and talks to them along the way. AgentLens evaluates that whole trajectory. It pairs formal verification, where an objective check exists, with LLM-written trajectory reviews and side-by-side comparisons, so that each run yields a readable explanation of why the score is what it is. This makes AgentLens useful for more than ranking models: we use it to diagnose model behavior, compare successive versions of our own agent, and catch product regressions in a nightly evaluation pipeline. We release the benchmark as open source at https://github.com/agent-lens/agent-lens-bench.

13:00 JSTLLM/生成AI

コンテキスト内検索が役立つのはどのような場合ですか?リフレクション駆動推論のサンプリング複雑性理論

拡張推論を使用して大規模言語モデル (LLM) をトレーニングすることにより、モデルが解決策の試行を繰り返し生成、批評、修正するコンテキスト内検索が可能になりました。推論トレースに対する近似推論としてモデル化することで、コンテキスト内検索の理論的分析を提供します。基本モデルは事前を定義し、自己反映は事後更新のフィードバックを提供し、結果として生じる推論時間のサンプリングの複雑さ、つまり高い成功確率を達成するために必要な連続試行の回数を研究します。リフレクションが初期ミスを確実に特定する場合、コンテキスト内検索は基本モデルに対して指数関数的な改善をもたらし、多項式の連続試行のみを使用して指数関数的に小さいゼロショット合格率の問題を解決できますが、この特性が失敗した場合、過去の試行に基づく条件付けでは並列サンプリングに比べて漸近的な利点が得られないことを示します。さらに、これらのゲインは堅牢で学習可能であることを示します。近似事後更新で十分であり、検索ロールアウトでのクロスエントロピー トレーニングにより、多項式サンプルの複雑さで必要な動作が回復します。最後に、検証可能な報酬を伴う強化学習の段階的抽象化の下で、最適なポリシー拡張が同じ事後再重み付けルールを実装することを示します。実際の大規模推論モデルに関する理論の重要な定性的予測を検証します。

原文 (English)

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning

Training large language models (LLMs) with extended reasoning has enabled in-context search, in which models iteratively generate, critique, and revise solution attempts. We provide a theoretical analysis of in-context search by modeling it as approximate inference over reasoning traces, where the base model defines a prior and self-reflection provides feedback for posterior updates, and study the resulting inference-time sampling complexity - the number of sequential attempts needed to achieve high success probability. We show that when reflections reliably localize early mistakes, in-context search can yield exponential improvements over the base model, solving problems with exponentially small zero-shot pass rates using only a polynomial number of sequential attempts, whereas when this property fails, conditioning on past attempts offers no asymptotic benefit over parallel sampling. We further show that these gains are robust and learnable: approximate posterior updates suffice, and cross-entropy training on search rollouts recovers the required behavior with polynomial sample complexity. Finally, we show that under a stagewise abstraction of reinforcement learning with verifiable rewards, the optimal policy extension implements the same posterior reweighting rule. We validate key qualitative predictions of the theory on real large reasoning models.

13:00 JSTLLM/生成AIエージェント

エージェントベースのモデリングにおける LLM を活用した推論

エージェントベース モデリング (ABM) には、何百万もの個人とそのやり取りをモデル化する機能があり、政策立案に役立ちます。ただし、ABM は伝統的に静的な事前分布に依存しており、そのためモデルがリアルタイムの変化に適応することができません。私たちの研究は、この情報ギャップに対処するための新しいアプローチを提供します。大規模言語モデル (LLM) は、人間の意思決定を予測する新しい機会を提供します。ここでは、LLM を活用して ABM シミュレーションで人間の意思決定を予測する、スケーラブルなハイブリッド エージェントベースおよび言語駆動型流行 (HALE) モデリング フレームワークを紹介します。概念実証として、私たちは HALE を使用して、ユタ州ソルトレーク郡における新型コロナウイルス感染症とその影響をシミュレーションしました。

原文 (English)

LLM-powered reasoning in agent-based modeling

Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making. However, ABMs have traditionally relied on static prior, which prevents the models from adapting to real-time changes. Our research provides a novel approach to addressing this information gap. Large language models (LLMs) offer new opportunities to predict human decision-making. Here, we introduce a scalable Hybrid Agent-based and Language-driven Epidemic (HALE) modeling framework that leverages LLMs to predict human decision-making in an ABM simulation. As a proof-of-concept, we use HALE to simulate COVID-19 and its effects in Salt Lake County, UT.

13:00 JSTエージェント

QANTIS: IBM Heron でのハードウェア調整された順次 POMDP 信念更新

部分的な可観測性の下にある自律システムは、生のセンサー イベントではなく、信念に基づいて動作します。 QANTIS は、そのループ内で量子プロセッサを校正済みの信念更新サービスとして扱います。事前モデルと観測モデルを受け取り、稀な事象の証拠項を推定し、古典的なプランナーに通常の事後結果を返します。このペーパーでは、現在の IBM Heron ハードウェア上で、プランナーに面した事後結果を損なうことなく、そのサービスを連続した Tiger POMDP ホライズン全体で再利用できるかどうかを尋ねます。私たちは、エンドツーエンドの自律性や実時間の高速化を主張するのではなく、制御されたハードウェアのケーススタディで答えます。この研究では、増幅なし、ガード付きグローバー増幅、および全ステップ固定小数点増幅を同じ軌道上で比較し、返された事後分布が下流のアクションを変更するかどうかを確認します。全ステップ FPAA は、報告された 8 ステップおよび 12 ステップのプライマリ ラン全体にわたって Tiger の後方を維持し、20 ステップおよび 32 ステップのコントロールは同じ動作帯域内に留まります。報告されたすべての意思決定チェックでは、ハードウェア事後処理と正確なベイズ事後処理により、同じ即時アクションが選択されます。境界を意識した BIQAE は、ゼロ付近および 1 付近の振幅推定を安定させ、レアイベント スイープは、100 万分の 1 の証拠の論理的なサンプル複雑さのエンベロープをマッピングします。その結果は、スタンドアロンのハードウェアの利点を主張するものではなく、ハードウェアで調整された信念更新プリミティブの動作エンベロープになります。

原文 (English)

QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron

Autonomous systems under partial observability act on beliefs, not raw sensor events. QANTIS treats the quantum processor as a calibrated belief-update service in that loop: it receives a prior and an observation model, estimates the rare-event evidence term, and returns an ordinary posterior to a classical planner. This paper asks whether that service can be reused across a sequential Tiger POMDP horizon on present IBM Heron hardware without corrupting the planner-facing posterior. We answer with a controlled hardware case study rather than an end-to-end autonomy or wall-clock speedup claim. The study compares no amplification, guarded Grover amplification, and all-step fixed-point amplification on the same trajectory, then checks whether the returned posterior would change the downstream action. All-step FPAA preserves the Tiger posterior across the reported 8-step and 12-step primary runs, and the 20-step and 32-step controls remain inside the same operating band. In every reported decision check, the hardware posterior and the exact Bayes posterior select the same immediate action. Boundary-aware BIQAE stabilizes amplitude estimation near zero and near one, while a rare-event sweep maps the logical sample-complexity envelope for one-in-a-million evidence. The result is an operating envelope for a hardware-calibrated belief-update primitive, not a standalone hardware-advantage claim.

13:00 JSTエージェントDeepSeek

ARC-AGI-1 での抽象推論と一般化のためのコスト効率の高いエージェントの活用

開示されたアーキテクチャによる ARC-AGI-1 の最近の進歩は、大きく 2 つの領域から来ています。1 つはフロンティア モデルに対する大量のテスト時間の計算 (進化的探索、徹底的なサンプリング、拡張された思考連鎖)、もう 1 つはタスクに特化したアーキテクチャを使用して小さなモデルが ARC データに基づいて微調整されるベンチマーク固有のトレーニングです。私たちは 3 番目のレジームを研究します。それは、ARC 固有の微調整を行わない、厳格な予算の下で非思考モードのオープンウェイト モデル (DeepSeek V3.2) です。私たちは、パターン発見とプログラム合成の段階を明示的に分解するエージェント ハーネスを構築しながら、アーキテクチャだけで何が回復できるかを研究します。まず、パターン検出を実行可能変換合成から分離する Explorer-Definer Pipeline を導入し、2 段階のエージェント パイプラインとして実装します。次に、トレーニング ペアで以前の仮説が失敗した場合に、新しい変換を自律的に探索してパイプラインを強化する Reflective Orchestrator を紹介します。 ARC-AGI-1 パブリック 400 タスク評価セットでは、パイプラインはタスクあたり \$0.25 で 57.50% pass@2 に達し、オーケストレーターはタスクあたり \$0.62 で 67.25% pass@2 に達します。これらのアーキテクチャを組み合わせると、ベンチマーク固有のトレーニングや大量のテスト時間の計算を行わずに、15.50% のワンショット ベースラインが最大 52 ポイント向上します。さらに、オーケストレーター主導のリフトは、パイプラインが生成する改ざん可能な診断をテストします。公平な pass@k 分析は、パイプラインが選択に依存するのではなく、世代に依存することを示唆し (トレーニング ペアの精度による選択は、候補上限の最大 95% を捕捉します)、大幅な改善にはより良いランキングではなく、より広範な世代が必要であると予測します。オーケストレーターは、適応型再探索によってこの予測を実装し、それを確認します (不偏パス@1 リフト +9.81 pp、マッチング選択媒介パス@2 リフト)。追加のパイプライン アブレーションにより、そのシンク ツールが重要なコンポーネントであることが特定され、削除により pass@2 が 5.75 pp 減少します。

原文 (English)

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary search, exhaustive sampling, extended chain-of-thought), or benchmark-specific training in which small models are fine-tuned on ARC data, often with task-specialized architectures. We study a third regime: an open-weight model in non-thinking mode (DeepSeek V3.2) under a strict budget, with no ARC-specific fine-tuning. We study what is recoverable through architecture alone, building agentic harnesses that decompose pattern-discovery and program-synthesis stages explicitly. First, we introduce an Explorer-Definer Pipeline that separates pattern discovery from executable transformation synthesis, implemented as a two-stage agent pipeline. Next, we present the Reflective Orchestrator, which augments the pipeline with autonomous exploration of new transformations when previous hypotheses fail on training pairs. On the ARC-AGI-1 public 400-task evaluation set, the pipeline reaches 57.50% pass@2 at \$0.25 per task, and the orchestrator reaches 67.25% pass@2 at \$0.62 per task. Together these architectures lift a 15.50% one-shot baseline by ~52 points without benchmark-specific training or heavy test-time compute. Furthermore, the orchestrator-driven lift tests a falsifiable diagnostic the pipeline produces; unbiased pass@k analysis suggests the pipeline is generation-bound, not selection-bound (selection via training-pair accuracy captures ~95% of the candidate ceiling) and predicts that significant improvement requires broader generation, not better ranking. The orchestrator implements this prediction via adaptive re-exploration and confirms it (unbiased pass@1 lift +9.81 pp, matching selection-mediated pass@2 lift). An additional pipeline ablation identifies its think tool as a significant component, with removal reducing pass@2 by 5.75 pp.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPTQwen

計算数学および実験数学のための SageMath 拡張 LLM エージェントの評価

数学用 AI の最近の進歩は主に自動形式化と定理証明に焦点を当てており、エージェントの LLM ワークフローにおけるコンピューター代数システム (CAS) の役割は十分に検討されていません。私たちは、LLM 推論と SageMath からの検証可能なフィードバック、および最新のドキュメント用の Context7 を組み合わせた ReAct スタイルのエージェント セットアップを提案します。私たちは、計算数学の研究ループをエミュレートする設定で RealMath ベンチマークから研究レベルの数学的問題を解決するために、フロンティア モデル全体でこのエージェント セットアップを評価します。また、複数ステップの後処理手順と複数ステージの検証パイプラインを導入することで、RealMath ベンチマークを改良することも提案します。これらの両方により、抽出された問題セットの品質と信頼性が向上します。私たちの実験では、評価されたすべてのモデルにわたって SageMath アクセスによる大幅なパフォーマンスの向上が平均で +9.7~pp であり、その向上の範囲は 1.5~pp から 27.8~pp であり、オープンウェイト モデルとクローズド モデルの間のギャップが狭まっています。 Qwen~3.7-Max は SageMath の恩恵を最も受けますが、GPT-5.5 は $75.2\%$ という最高の解決率を達成し、ツールが有効な構成の中で最低のトークン使用量を達成します。私たちの調査結果は、CAS 拡張エージェントが数学者の計算探索を支援するための有望な方向性を示していることを示唆しており、この研究が自動化された推測発見に向けた一歩であると信じています。プロジェクト リポジトリはオンラインで入手できます。

原文 (English)

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentation. We evaluate this agentic setup across frontier models for solving research-level mathematical problems from the RealMath benchmark in a setting that emulates a computational-mathematics research loop. We also propose a refinement to the RealMath benchmark by introducing a multi-step post-processing procedure and a multi-stage validation pipeline, both of which improve the quality and reliability of the extracted problem set. Our experiments reveal substantial performance gains from SageMath access across all evaluated models on +9.7~pp on average, the gains range from 1.5~pp to 27.8~pp and narrow the gap between open-weight and closed models. Qwen~3.7-Max benefits from SageMath the most, while GPT-5.5 achieves the highest solve rate of $75.2\%$ and the lowest token usage among tool-enabled configurations. Our findings suggest that CAS-augmented agents represent a promising direction for assisting mathematicians in computational exploration, and we believe that this work is a step towards automated conjecture discovery. The project repository is available online.

13:00 JSTエージェントClaudeGeminiQwen

ハーネス効果: オーケストレーション設計がエンタープライズ エージェント AI のトークン エコノミクスをどのように設定するか

現在のエージェント AI 開発は、トークンの最大化に基づいて実行されています。つまり、トークンを使用して機能を購入すること、つまり、より長い推論トレース、より多くのターン、より幅広いツール ペイロード、より大きな再生コンテキストを実現できるため、タスクあたりのトークンはタスクの値よりも早く増加します。トークンごとの価格の下落により、そのパターンが隠蔽されます。とにかく総支出が増加します。私たちは、トークンの最大化に対する決定的な手段はハーネスであると主張します。ハーネスとは、コンテキストを組み立て、ツールを公開し、順序を変え、作業を委任し、企業の可観測性とガバナンスを実現するオーケストレーション層です。制御されたスワップでそれを分離します。22 のロックされた評価タスク、6 つの基礎モデル (Claude Sonnet 4.6、Gemini 3.1、Gemini Flash 3.5、Qwen 3.6、GLM 5.1、Palmyra X6)、オーケストレーション レイヤーのみの変更、つまり凍結された従来のプロダクション ループと Writer Agent Harness の変更です。モデルを一定に保つと、ハーネスはタスクあたりの混合コストを 41% (0.21 ドル→ 0.12 ドル)、実時間の中央値 44% (48 秒→ 27 秒)、タスクあたりのトークンを 38% (14.2k→8.8k) 削減し、タスク完了品質は同等 (0.78→0.81、このサンプル サイズでは方向性あり) です。効率はモデルに依存せず、すべてのモデルが安くなります (33 ~ 61%)。一方、品質の向上は能力に依存します。モデルのゲインは、ベースラインの強度 (r=0.99、n=6) とほぼ完全に相関します。これをハーネス レバレッジと呼ぶ現象です。 1 ドルあたりの品質は 82% 上昇します。 100 万トークンあたりのタスク完了数は 54.9 から 92.0 に増加します。このワークロードでは、オーケストレーション層により、モデル メニューを完全に拡張した場合よりもタスクあたりのコストが増加しました。私たちは、オーケストレーション層でのトークンエコノミクス(プロンプトキャッシュの下での有効な入力価格を含む)を形式化し、その効果の背後にある6つのメカニズムファミリー(キャッシュ形状の規律から障害支出ガバナンスまで)を詳細に説明し、広く使用されている6つのエージェントシステムを同じ軸で比較し、組織が現在および将来にわたって実行するすべてのモデルにわたって効率が倍増するコンポーネントの1つがハーネスであると主張します。

原文 (English)

The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payloads, bigger replayed contexts -- so tokens per task grow faster than task value. Falling per-token prices mask the pattern; total spend rises anyway. We argue the decisive lever against token maxing is the harness: the orchestration layer that assembles context, exposes tools, sequences turns, delegates work, and carries enterprise observability and governance. We isolate it with a controlled swap: 22 locked evaluation tasks, six foundation models (Claude Sonnet 4.6, Gemini 3.1, Gemini Flash 3.5, Qwen 3.6, GLM 5.1, Palmyra X6), changing only the orchestration layer -- a frozen conventional production loop versus the Writer Agent Harness. Holding models constant, the harness cuts blended cost per task 41% ($0.21->$0.12), median wall-clock 44% (48s->27s), and tokens per task 38% (14.2k->8.8k), with task-completion quality at parity (0.78->0.81, directional at this sample size). Efficiency is model-invariant -- every model gets cheaper (33-61%) -- while quality gains are capability-dependent: a model's gain correlates almost perfectly with its baseline strength (r=0.99, n=6), a phenomenon we term harness leverage. Quality per dollar rises 82%; task-completions per million tokens rise from 54.9 to 92.0. On this workload the orchestration layer moved cost per task more than the full spread of the model menu did. We formalize token economics at the orchestration layer (including effective input price under prompt caching), detail the six mechanism families behind the effect -- cache-shape discipline to failure-spend governance -- compare six widely used agent systems on the same axes, and argue the harness is the one component whose efficiency multiplies across every model an organization runs -- present and future.

13:00 JST研究/論文

コンパクトな世界モデルにおける空間関係の基礎付け: 命令リークとゴールフリーのダイナミクス修正

言語目標を条件とするコンパクトな世界モデルは、明示的な \emph{参照アンカー} のまばらなセットを使用して、「赤いブロックを青いブロックの左に置く」などの関係を基礎付けることを約束します。私たちは、そのような参照が実際に関係を根拠づけるのはいつかを尋ね、罠を特定します。目標条件付き予測子は驚くべき $0.90$ の関係読み出し精度に達しますが、これは \emph{命令の転写} であり、知覚ではありません。ゴールを保留すると、それは偶然に崩壊し ($0.90\!\to\!0.27$、シード 3 つ)、反事実の指示により、予測されたアンカーは $94.5\%$ の確率で \emph{false} 指示に従います (本当のシーン $2.3\%$; $N{=}256$)。 3 つの設定とタスク内アブレーションにわたってテストしたところ、私たちの中心的な主張は交絡の特徴を示しています。 \textbf{命令のリークは、スコア付けされた数量が命令から転写可能な場合 (命令が答えに名前を付ける場合) に発生し、非命令入力の予測性とは本質的に独立しています。} 私たちのテーブルトップと外部の BabyAI ベンチマークのリークとは対照的に、命令名が \emph{referents} である言語テーブルのフォワード ダイナミクス世界モデルでは、命令が実行されるまでリークは発生しません。方向に名前を付けるために拡張されます。そして、アクションを低下させてもリークが増加することはありません。これは、予測子と競合が予測するものとは逆です。診断によって修正が指示されます。目標をダイナミクスから遠ざけ (プランナーのコストに属します)、\emph{read} パスを監視して、命令に依存しない真の根拠を回復します ($0.88$、目標の有無にかかわらず同一)。検出プロトコルと解決策は、命令にスコア付けされた量が指定されているゴール条件付き世界モデルに適用されます。

原文 (English)

Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix

Compact world models that condition on a language goal promise to ground relations such as ``put the red block left of the blue block'' using a sparse set of explicit \emph{reference anchors}. We ask when such references actually ground a relation, and identify a trap: a goal-conditioned predictor reaches a striking $0.90$ relation-readout accuracy, yet this is \emph{instruction transcription}, not perception. Withholding the goal collapses it to chance ($0.90\!\to\!0.27$, three seeds) and a counterfactual instruction makes the predicted anchors follow the \emph{false} instruction $94.5\%$ of the time (true scene $2.3\%$; $N{=}256$). Tested across three settings and a within-task ablation, our central claim characterizes the confound: \textbf{instruction leakage occurs when the scored quantity is transcribable from the instruction (when the instruction names the answer) and is essentially independent of how predictive the non-instruction inputs are.} Our tabletop and the external BabyAI benchmark leak, whereas a Language-Table forward-dynamics world model whose instruction names \emph{referents} does not, until the instruction is augmented to name the direction; and degrading the action never increases leakage, the opposite of what predictor-competition predicts. The diagnosis prescribes the fix: keep the goal out of the dynamics (it belongs to the planner's cost) and supervise the \emph{read} path, recovering genuine, instruction-independent grounding ($0.88$, identical with and without the goal). The detection protocol and remedy apply to any goal-conditioned world model whose instruction names the scored quantity.

13:00 JSTLLM/生成AI

大規模行動モデル: 小売顧客の迅速なデジタル ツイン

顧客行動モデリングは、レコメンデーション、マーケティング、意思決定サポートを支えていますが、既存のアプローチは、意思決定を説明せずに予測精度を最適化するか、実際の行動データに基づいてユーザーをシミュレートするかのどちらかです。私たちは、統合された個人と環境の定式化を通じて、大規模な小売取引から顧客の意思決定を直接学習する大規模行動モデル (LBM) を紹介します。顧客の状態は購入履歴から得られた行動プロファイルによって表され、製品コンテキストは検索拡張生成によって組み込まれます。このモデルは、言語化された行動データに対する継続的な事前トレーニング、意思決定生成のための教師あり微調整、および証拠に基づく調整のための検証可能な報酬を伴う強化学習を使用してトレーニングされます。購入予測、ハードネガティブな差別、バスケットの完了、プロモーションへの対応、およびクロスドメインのクーポン引き換えに関する提案されたフレームワークを評価します。このモデルは、小売業者や意思決定ドメイン全体にわたる強力なゼロショット転送と微調整された転送を実証しながら、ドメイン内の小売タスクではフロンティア汎用言語モデルを常に上回っています。アブレーション研究では、継続的な事前トレーニングが行動の一般化の主な推進力であり、検索はトレーニングと推論の両方で適用すると最も効果的であり、強化学習は一般的な言語モデルの事前学習よりも明示的な行動の証拠への依存を改善することを示しています。これらの結果は、トランザクション履歴にエンコードされた行動知識が言語モデルによって効果的に学習でき、顧客のデジタル ツインと行動シミュレーションにスケーラブルな基盤を提供できることを示しています。

原文 (English)

Large Behavior Model: A Promptable Digital Twin of the Retail Customer

Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data. We present the Large Behavioral Model (LBM) that learns customer decision making directly from large-scale retail transactions through a unified Person-Environment formulation. Customer state is represented by a behavioral profile derived from historical purchases, while product context is incorporated through retrieval-augmented generation. The model is trained using continued pre-training on verbalized behavioral data, supervised fine-tuning for decision generation, and reinforcement learning with verifiable rewards for evidence-based calibration. We evaluate the proposed framework on purchase prediction, hard-negative discrimination, basket completion, promotion response, and cross-domain voucher redemption. The model consistently outperforms frontier general-purpose language models on in-domain retail tasks while demonstrating strong zero-shot and fine-tuned transfer across retailers and decision domains. Ablation studies show that continued pre-training is the primary driver of behavioral generalization, retrieval is most effective when applied during both training and inference, and reinforcement learning improves reliance on explicit behavioral evidence over generic language-model priors. These results demonstrate that behavioral knowledge encoded in transaction histories can be effectively learned by language models, providing a scalable foundation for customer digital twins and behavior simulation.

13:00 JST研究/論文

社会規範を学ぶことで、人間と AI の動的な連携における互換性が向上します

人間は、多くの場合、相互作用するエージェント間で共有される暗黙の期待として機能する、暗黙的で定量化が難しい社会規範を通じて、動的な相互作用において他者と継続的に調整します。大規模言語モデル (LLM) を含む AI エージェントが日常生活に組み込まれるにつれて、AI エージェントはそのような相互作用にますます参加し、社会的相互作用構造を再形成します。しかし、彼らは効果的で思いやりのある自然な方法で人間と協調することができないことがよくあります。私たちは、既存のアプローチが、そのような動作を生成する根本的な規範を明示的に定量化することなく、モデルの動作を人間のデモンストレーションと一致させるために、このギャップが生じると仮説を立てています。私たちは、代表的な動的インタラクションとして歩行者と車両のインタラクションを選択し、その主要なインタラクティブ機能を捕捉する簡略化された実験プラットフォームを開発しました。このプラットフォームを通じて収集された 3,456 件の動的な人間関係から、人間の社会規範の根底にある 3 つの原則、つまり結果の予測可能性、価値観の一致、および利点の認識を特定しました。これらの原則を AI エージェントに組み込むことで、人間と AI の連携が大幅に向上します。人間との閉ループ相互作用タスクでは、社会規範に基づいた LLM は、ベースライン戦略よりもほぼ 4 倍高い合計スコアを達成し、人間と人間の相互作用を 43% 上回りました。これらの発見は、暗黙の社会規範を明示的で定量化可能な原則に形式化することで、AI エージェントが動的な相互作用において相互に有益な調整を達成できるようになり、人間社会へのより自然な統合をサポートできることを示しています。

原文 (English)

Learning social norms enhances compatibility in dynamic human-AI coordination

Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as shared tacit expectations among interacting agents. As AI agents, including large language models (LLMs), become embedded in daily life, they increasingly participate in such interactions and reshape social interaction structures. Yet they often fail to coordinate with humans in an effective, considerate, and natural manner. We hypothesize that this gap arises because existing approaches align model behavior with human demonstrations without explicitly quantifying the underlying norms that generate such behavior. We selected pedestrian-vehicle interaction as a representative dynamic interaction and developed a simplified experimental platform that captures its key interactive features. From 3,456 dynamic human interactions collected via this platform, we identified three principles underlying human social norms: outcome predictability, value alignment, and advantage awareness. Incorporating these principles into AI agents significantly improves human-AI coordination. In the closed-loop interaction task with humans, the social-norm-informed LLM achieved a nearly fourfold higher total score than the baseline strategy and outperformed human-human interactions by 43%. These findings indicate that formalizing tacit social norms into explicit, quantifiable principles can enable AI agents to achieve mutually beneficial coordination in dynamic interactions, supporting their more natural integration into human society.

13:00 JST研究/論文

人間のスケールを超えた知性の測定

人間の能力を超えた知性をどのように測定できるでしょうか?人間が作成したベンチマークは飽和しており、人間の能力を超えて、どのタスクが困難で検証可能であるかを審査官が認識できない可能性があります。私たちは、この困難は絶対スケール評価に固有のものであると主張し、モデルが他のシステムを区別する公的課題を生成する相対測定に基づく新しいパラダイムを提案します。これらの結果を集約すると、測定対象のシステムに合わせて拡張できる敵対的な心理測定評価システムが得られます。私たちは、個人情報攻撃のインセンティブを軽減し、裁判官なしの判決をサポートし、エージェントの能力に合わせて自然に拡張する実用的なプロトコルについて説明します。私たちは、検証可能な領域とオープンエンドの検証不可能な領域にわたってフレームワークをインスタンス化し、モデル生成の評価が人間のフロンティアを超えてシステムをどのように測定し続けることができるかを示します。

原文 (English)

Measuring Intelligence Beyond Human Scale

How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not know which tasks are both hard and verifiable. We argue that this difficulty is inherent to absolute-scale evaluation and propose a new paradigm based on relative measurement in which models generate public challenges that separate other systems. Aggregating these outcomes yields an adversarial psychometric rating system that can scale with the systems being measured. We describe practical protocols that reduce incentives for private-information attacks, support judge-free adjudication, and naturally scale with agent capabilities. We instantiate the framework across verifiable and open-ended, non-verifiable domains, illustrating how model-generated evaluation can continue to measure systems beyond the human frontier.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達ClaudeGPT / ChatGPTGeminiDeepSeek

マルチエージェント LLM の安全性における運用の再構築と承認枠付きの委任

マルチエージェント LLM システムの安全性評価では、多くの場合、直接プロンプトとプランナーと実行者のパイプラインを比較し、その違いを単一の「パイプライン効果」として報告します。私たちは、この集合体は 3 つのメカニズムを混同しているため、解釈が難しいと主張します。有害な意図がもっともらしい運用作業として再構成される可能性があること、計画立案者が要求を拒否または変更する可能性があること、実行者が事前の承認を示唆する委任プロンプトに基づいて行動する可能性があることです。これらの要因を分離するために、LLM が判断したコンプライアンスを使用した 30 の合成有害シナリオと 4 つのエージェントの安全性ベンチマークからの探索的な外部検証セットで評価された 5 つの条件で制御されたコントラスト設計を導入します。私たちの結果は、集合的なパイプラインの安全性が安定したアーキテクチャ上の特性ではないことを示しています。オペレーショナル リフレーミングは最もポータブルなリスク シグナルであり、両方のシナリオ セットにわたって GPT、Gemini、DeepSeek のコンプライアンスを強化しますが、Claude は比較的抵抗力があります。プランナーの行動は、主に拒否を通じてこのリスクを相殺できます。ただし、プランナが実行可能なステップを作成すると、実行者は直接の運用ベースラインよりも準拠するようになる可能性があります。承認フレームワークの委任は、プロンプトの設計、モデルのペアリング、およびシナリオのソースに敏感であり、懐疑的な実行者のプロンプトはコンプライアンスを大幅に低下させます。生の直接モデルのランキングも、デプロイされたプランナーと実行者の動作を誤って予測する可能性があります。ジェミニは、プライマリ セットの生の直接プロンプトの下で最も安全ですが、クロード プランナーを使用すると最大の増幅を示し、コンプライアンスが 8.9 パーセントから 38.9 パーセントに上昇しました。 GPT の集約パイプライン効果がゼロに近いため、代わりにプランナーの拒否によってキャンセルされたリフレーミングの増加が隠蔽されます。これらの発見は、マルチエージェントの安全性評価では、失敗をアーキテクチャ自体のせいにする前に、リフレーミング、プランナーの動作、委任のフレーミング、およびモデルのペアリングを個別に報告する必要があることを示唆しています。

原文 (English)

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it conflates three mechanisms: harmful intent may be reframed as plausible operational work, the planner may refuse or transform the request, and the executor may act under delegation prompts implying prior approval. To separate these factors, we introduce a five-condition controlled contrast design, evaluated on 30 synthetic harmful scenarios and an exploratory external validation set from four agent-safety benchmarks using LLM-judged compliance. Our results show that aggregate pipeline safety is not a stable architectural property. Operational reframing is the most portable risk signal, increasing compliance for GPT, Gemini, and DeepSeek across both scenario sets, while Claude is comparatively resistant. Planner behavior can offset this risk mainly through refusal; however, when the planner produces executable steps, the executor may become more compliant than under the direct operational baseline. Approval-framed delegation is sensitive to prompt design, model pairing, and scenario source, and a skeptical executor prompt sharply reduces compliance. Raw-direct model rankings can also mispredict deployed planner-executor behavior. Gemini is safest under raw direct prompts in the primary set yet shows the largest amplification with a Claude planner, rising from 8.9 percent to 38.9 percent compliance. GPTs near-zero aggregate pipeline effect instead hides a reframing increase canceled by planner refusal. These findings suggest that multi-agent safety evaluations should report reframing, planner behavior, delegation framing, and model pairing separately before attributing failures to architecture itself.

13:00 JSTエージェント研究/論文GPT / ChatGPTGeminiQwen

AIは画像を理解できるのか?コンピューテーショナル イメージング タスク向けの Agentic AI の体系的なベンチマーク

視覚言語モデル (VLM) とエージェント AI は、意味論的な視覚タスクで優れたパフォーマンスを示していますが、計算によるイメージングの基礎となる物理学や逆問題を処理できるかどうかは依然として不明です。我々は、光線光学と波動光学、画像信号処理、逆再構成、計算センシング、およびキャリブレーションの 5 つのカテゴリにわたる 20 の計算イメージング タスクのベンチマークである ImagingBench を紹介します。 ImagingBench は 3 つの補完的な設定を評価します。エキスパート、固定エキスパートガイドによる逆再構成。プランナー、プランナー主導の逆再構成。一貫性チェックのためのフォワード、フォワードシステムシミュレーション。 Gemini、GPT、Qwen など、主要な独自のオープンソース画像中心マルチモーダル システムのベンチマークを行い、代表的なタスク固有の非エージェント ベースラインと比較します。どのタスクにおいても、エージェントティック モデルは、特にレンズレス イメージング、イベントベースの再構成、飛行時間型イメージング、ホログラフィーなどのコンピューティング センシングの問題において、特殊な手法より一貫して弱いままです。プランナーのガイダンスは、固定プロンプトのエキスパートのベースラインに比べて、わずかで一貫性のない利益しか提供しません。モデルは視覚的に妥当な出力を生成することがよくありますが、参照ベースの忠実度は依然として低く、意味論的な視覚能力と物理的に根拠のある画像パフォーマンスとの間に大きなギャップがあることが明らかになります。 ImagingBench は、こ​​のギャップを測定し、コンピュテーショナル イメージング用のエージェント AI の進捗状況を追跡するための統合テストベッドを提供します。

原文 (English)

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBench, a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration. ImagingBench evaluates three complementary settings: Expert, fixed expert-guided inverse reconstruction; Planner, planner-guided inverse reconstruction; and Forward, forward-system simulation for consistency checking. We benchmark leading proprietary and open-source image-centric multimodal systems, including Gemini, GPT, and Qwen, and compare them with representative task-specific non-agentic baselines. Across tasks, agentic models remain consistently weaker than specialized methods, especially on computational sensing problems such as lensless imaging, event-based reconstruction, time-of-flight imaging, and holography. Planner guidance provides only modest and inconsistent gains over the fixed-prompt Expert baseline. Although the models often generate visually plausible outputs, their reference-based fidelity remains poor, revealing a substantial gap between semantic visual competence and physically grounded imaging performance. ImagingBench provides a unified testbed for measuring this gap and tracking progress in agentic AI for computational imaging.

13:00 JSTビジネス/資金調達

推論一貫性スキャン: AI の安全性評価における思考連鎖の妥当性を監査するためのフレームワーク

これまでの研究では、思考連鎖 (CoT) 推論がしばしば不忠実であることが示されています。つまり、モデルで述べられた推論は、その出力を生成したプロセスを確実に反映していません。ただし、不貞を検出するには、管理された実験的介入が必要であり、事後の評価記録にそれを適用することはできません。代わりに、あまり注目されていない、より扱いやすい質問、つまり、述べられた推論がそれに伴う答えと論理的に一貫しているかどうかに目を向けます。忠実度とは異なり、一貫性は介入なしで記録のみから評価できます。 AI 安全性評価トランスクリプトでこの特性を検出するための再利用可能な方法である推論一貫性スキャンを紹介します。私たちの貢献は 4 つあります。まず、推論の一貫性を忠実性とは異なるものとして形式化し、不一致の 6 つのサブタイプ分類を定義します。次に、InstrumentalEval の出力から手動で調整した 60 個のトランスクリプトの検証済みベンチマークを構築します。 3 番目に、InspectScout 用に動作するスキャナーを実装します。これは、安全性評価記録でこのプロパティをターゲットにした最初のスキャナーです。 4 番目に、4 つのジェネレーター モデルと、inspect_evals からの 3 つの評価にわたる結果を報告します。これは、推論の矛盾が存在し、検出可能であり、モデルとタスク タイプの両方にわたって体系的に変化していることを示しています。

原文 (English)

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that produced its output. Detecting unfaithfulness, though, requires controlled experimental interventions, which cannot be applied to evaluation transcripts after the fact. We turn instead to a more tractable question that has received less attention: whether the stated reasoning is logically consistent with the answer it accompanies. Unlike faithfulness, consistency can be assessed from a transcript alone, with no intervention. We introduce reasoning consistency scanning, a reusable method for detecting this property in AI safety evaluation transcripts. Our contributions are fourfold. First, we formalize reasoning consistency as distinct from faithfulness and define a six-subtype taxonomy of inconsistency. Second, we build a validated benchmark of 60 transcripts, manually adapted from InstrumentalEval outputs. Third, we implement a working scanner for InspectScout, the first to target this property in safety evaluation transcripts. Fourth, we report results across four generator models and three evaluations from inspect_evals, showing that reasoning inconsistency is present, detectable, and varies systematically across both models and task types.

13:00 JSTLLM/生成AIエージェント

アトミック アクションから標準操作手順まで: 自己進化する LLM エージェントのための反復的なツールの最適化

ツールを利用すると、Large Language Model (LLM) エージェントが現実世界と対話し、複雑なタスクを解決できるようになります。しかし、既存のエージェント フレームワークは主に、細分化されたアトミック アクション (基本的なファイル I/O やシングルターン検索など) で構成される静的ツールセットに依存しているため、エージェントは繰り返し発生するワークフローごとに低レベルのロジックを再発明する必要があり、推論のオーバーヘッドと失敗率の増加につながります。この研究では、これらのアトミックなアクションを再利用可能な標準操作手順 (SOP) に合成することで、エージェントが自己進化を達成できることを提案します。SOP は、複数ステップのロジックをカプセル化する呼び出し可能な高次ツールとして機能します。さらに、エージェントが実行軌跡から SOP を抽出し、構築、マージ、評価、枝刈りの体系的なライフサイクルを通じてツールセットを繰り返し最適化できるフレームワークである EvoSOP を紹介します。広範な実験により、EvoSOP はベースラインと比較してインタラクション ラウンドの数を大幅に削減しながら、タスクの成功率を大幅に向上させることが実証されました。私たちの分析では、反復的なツールの最適化が信頼性が高く効率的なツールの使用パターンを促進し、自己進化するエージェントの開発に拡張可能な経路を提供することも明らかにしました。

原文 (English)

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g., basic file I/O or single-turn search), which forces agents to reinvent low-level logic for every recurring workflow, leading to increased reasoning overhead and failure rates. In this study, we propose that agents can achieve self-evolution by synthesizing these atomic actions into reusable Standard Operating Procedures (SOPs), which function as callable higher-order tools that encapsulate multi-step logic. We further introduce EvoSOP, a framework that empowers agents to extract SOPs from execution trajectories and iteratively optimize the toolset through a systematic lifecycle of construction, merging, evaluation, and pruning. Extensive experiments demonstrate that EvoSOP significantly boosts task success rates while substantially reducing the number of interaction rounds compared to baselines. Our analysis also reveals that iterative tool optimization fosters reliable and efficient tool-use patterns, providing a scalable pathway for the development of self-evolving agents.

13:00 JSTLLM/生成AIエージェント

科学的機械学習における物理監査によるエージェント発見

エージェント科学機械学習 (SciML) では、大規模言語モデル (LLM) エージェントがサロゲート モデルを検出し、自動スコア (通常はエラー メトリック) によって 1 つを選択できます。ただし、誤差が低いということは、予測されたフィールドが境界条件、重ね合わせ、剛性スケーリング、因果関係などの力学にとって重要な物理学を満たしていることを証明するものではありません。エージェント SciML 検出のための検証優先ワークフローである Physics-Audited Agentic Sc​​iML (PA-SciML) を紹介します。このワークフローは、検索前にスコアリング評価器を修正し、レビュー可能な機械チェック可能な物理要件を導き出し、その出力でトレーニングされた各候補をチェックし、基準ソリューション フィールドなしで違反の多いケースについて規定の入力範囲または測定された負荷履歴スパンを個別に検索します。サロゲートは、指定されたチェックのもとでのみ検証されたものとして報告されます。有効にすると、ワークフローはトレーニング前にアドバイスとなる数値プローブも追加し、一度に 1 つのモデリング変更をテストして、どの分離された編集がスコア増加に関連付けられているかを再利用前に記録します。報告された計算固体力学の数値例では、静的弾性の実行により、エラーのみのベースラインよりも検証誤差が低いサロゲートが選択されますが、選択された両方のモデルは共通の線形弾性チェックに合格します。過渡弾性力学の実行では、同様の平均誤差を持つ誤差のみのベースラインは、荷重履歴の将来の部分に応答するため、より厳密な因果関係チェックに失敗しますが、選択されたサロゲートは指定されたチェックに合格します。主な違いは、より豊富な集計スコアではなく、予測フィールドに関する候補ごとの物理的証拠です。

原文 (English)

Physics-Audited Agentic Discovery in Scientific Machine Learning

In agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automated score, typically an error metric. A low error, however, does not establish that the predicted fields satisfy the physics that matter for mechanics, such as boundary conditions, superposition, stiffness scaling, or causality. We introduce Physics-Audited Agentic SciML (PA-SciML), a verification-first workflow for agentic SciML discovery. The workflow fixes a scoring evaluator before search, derives reviewable machine-checkable physics requirements, checks each trained candidate on its outputs, and separately searches prescribed input ranges or measured load-history spans for high-violation cases without reference solution fields. A surrogate is reported as verified only under the stated checks. When enabled, the workflow also adds advisory numerical probes before training and tests one modeling change at a time to record which isolated edits are associated with score gains before reuse. In the reported computational-solid-mechanics numerical examples, the static elasticity run selects a surrogate with lower validation error than the error-only baseline while both selected models pass the common linear-elastic checks. In the transient elastodynamics run, an error-only baseline with similar mean error fails a stricter causality check by responding to future parts of the loading history, while the selected surrogate passes the stated checks. The main distinction is per-candidate physics evidence on predicted fields, not a richer aggregate score.

13:00 JST研究/論文

MIRA-Math: 最小限の情報要求と数学的推論のベンチマーク

数学的推論ベンチマークは通常、各問題を解決するために必要なすべての事実を提供しますが、インタラクティブなベンチマークでは推論とツール、検索、および長期的な対話が混合されることがよくあります。より狭い診断機能のベンチマークである MIRA-Math を紹介します。これは、完全な潜在状態が一意の答えを持っているものの、ソルバー側のビューに必要な原子的事実が 1 つだけ欠けている数学的問題を解決する機能です。ソルバーは、厳しい予算の下で自然言語で不足している情報を要求し、返された事実を正確な最終的な答えに統合する必要があります。固定制約付き LLM レスポンダーは、データセットが提供するアトミック ファクトのみを参照し、リクエストがそれに一致する場合は引用されたファクトを提供するか、そうでない場合は拒否する必要があります。したがって、インスタンスの生成、型指定されたヒントの仕様、検証、および最終回答の検証は決定的ですが、要求メトリクスは固定の LLM 仲介レスポンダ チャネルの下で測定されます。 MIRA-Math には、代数、確率、線形システム、離散構造、信号処理、マルコフ連鎖、回路、内挿、および数値境界値問題にわたる 22 の型付き数学族から生成された 2{,}10 個のインスタンスが含まれています。フロンティアモデルと小規模モデルにわたる実験では、リクエストの成功と最終応答の精度が分離可能であることが示されています。モデルは正しいファクトを要求しているにもかかわらず下流の計算に失敗する場合や、正規のヒントを取得する前に失敗する場合があります。数学的推論で要求される最小限の情報の再現可能な評価をサポートするために、ジェネレーター、ベリファイアー、プロンプト、実行メタデータ、およびデータセット ドキュメントをリリースします。

原文 (English)

MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning

Mathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasoning with tools, retrieval, and long-horizon dialogue. We introduce MIRA-Math, a benchmark for a narrower diagnostic capability: solving mathematical problems whose full latent state has a unique answer, but whose solver-facing view is missing exactly one necessary atomic fact. The solver must request the missing information in natural language under a strict budget and then integrate the returned fact into an exact final answer. A fixed constrained LLM responder sees only the dataset-provided atomic fact and must either offer the quoted fact when the request matches it, or decline otherwise. Thus, instance generation, typed hint specifications, validation, and final-answer verification are deterministic, while request metrics are measured under a fixed LLM-mediated responder channel. MIRA-Math contains 2{,}310 generated instances from 22 typed mathematical families spanning algebra, probability, linear systems, discrete structures, signal processing, Markov chains, circuits, interpolation, and numerical boundary-value problems. Experiments across frontier and small models show that request success and final-answer accuracy are separable: models may ask for the right fact yet fail the downstream computation, or fail before obtaining the canonical hint. We release generators, verifiers, prompts, run metadata, and dataset documentation to support reproducible evaluation of minimal information requesting in mathematical reasoning.

13:00 JSTエージェント

エージェントデータ環境

自律型エージェントは速度、規模、労働効率の大幅な向上を約束しますが、失敗すると突然、取り返しのつかないコストが発生する可能性があります。したがって、エージェント自動化の中心的な課題は、失敗による影響を制限しながら自動化のメリットを高めることです。データベースは依然として現代のコンピューティングの中心ですが、エージェントはファイル、API、アプリケーション、システム状態にわたるより広範なデータ環境上で動作します。この講演では、エージェントの機能を強化し、安全性を保証するエージェントティック データ環境 (エージェントが動作する実行基盤) に関する初期の作業について概説します。この観点では、データ システムを受動的な状態保存から、安全で信頼性の高い実行を実現する能動的な基板に再構成します。

原文 (English)

Agentic Data Environments

Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversible costs. The central challenge for agentic automation is therefore to increase the benefits of automation while bounding the consequences of failure. While databases remain central to modern computing, agents operate over a broader data environment spanning files, APIs, applications, and system state. In this talk, I will outline early work on Agentic Data Environments -- the execution substrate in which agents operate -- that both amplify agent capabilities and enforce safety guarantees. This perspective reframes data systems from passive stores of state into active substrates for safe, reliable execution.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

理由は少なく、もっと検証してください: 決定論的ゲートは、ツールを使用する LLM エージェントのサイレント ポリシー違反障害モードを回復します

ツールを使用する LLM エージェントは、タスクを正常に完了したように見えながら、導入されたポリシーそのものに違反する可能性があります。ポリシーが許可された環境では、対応する状態遷移がドメイン ポリシーによって禁止されている場合でも、ツールは適切な形式の呼び出しを実行できます。その結果、ツールもエージェントの自己報告も明らかにしない、静かな間違った状態 (予約がキャンセルされ、乗客数が変更され、検証なしで請求が処理された) が生じます。私たちはこの故障モードを $\tau^2$-bench 航空会社ドメインで研究します。低予算エージェントでは、観察された障害の 78% は、ツール エラーのないサイレントな間違った状態の障害であり、総障害率は、サンプリング ノイズではなく、ばらばらのシード間で再現可能です。次に、軽量介入を評価します。これは、書き込みを許可する前に、提案された呼び出しと現在の状態を検査する決定論的な読み取り専用の実行前ゲートです。 4 ゲート スイートでは、gpt-4o-mini でのフルベンチマークの成功率が 29.6% から 42.0% に上昇し (+12.4pp; ペアのタスクレベル ブートストラップ P=0.0012)、リフトは素の 15 シード セットで再現されました (+12.3pp; P=0.0008)。効果はゲートが発火する場所に集中します。26/50 発火タスクでは成功率が +19.2pp 上昇しますが、24 発の非発火タスクでの動きはゼロを除外しません。 2 つのネガティブ コントロール (自己強制的な小売ドメインと BFCL) がこのメカニズムを制限します。ゲートは、ツールがポリシーに寛容な場合に役立ちますが、ツールが既に自己強制している場合にはほとんど追加されません。中心的な主張ではなく示唆的な証拠として、同じ障害モードがフロンティアで持続します。デフォルトの推論では gpt-5.2 は依然としてポリシー違反の書き込みを試行し、同じスイートでは成功率が 61.2% から 71.6% に向上します (+10.4pp; P=0.020; n=5、レプリケーションなし)。貢献は、制限された評価と信頼性の結果です。決定論的ゲートはタスクの成功を保証しませんが、アクションの境界で既知のクラスのサイレント ポリシー違反書き込みを決定論的に阻止できます。

原文 (English)

Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents

Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive environments, a tool may execute any well-formed call even when the corresponding state transition is forbidden by domain policy. The result is a silent wrong state (a booking cancelled, a passenger count changed, a claim acted on without verification) that neither the tool nor the agent's self-report exposes. We study this failure mode in the $\tau^2$-bench airline domain. On a budget agent, 78% of observed failures are silent wrong-state failures with no tool error, and the aggregate failure rate is reproducible across disjoint seeds, not sampling noise. We then evaluate a lightweight intervention: deterministic, read-only pre-execution gates that inspect the proposed call and current state before allowing a write. A four-gate suite raises full-benchmark success from 29.6% to 42.0% on gpt-4o-mini (+12.4pp; paired task-level bootstrap P=0.0012), and the lift reproduces on a disjoint 15-seed set (+12.3pp; P=0.0008). The effect is concentrated where the gates fire: on the 26/50 firing tasks, success rises by +19.2pp, while movement on the 24 non-firing tasks does not exclude zero. Two negative controls (a self-enforcing retail domain and BFCL) bound the mechanism: gates help when tools are policy-permissive and add little where tools already self-enforce. As suggestive evidence, not a central claim, the same failure mode persists at the frontier: gpt-5.2 at default reasoning still attempts policy-violating writes, and the same suite improves success from 61.2% to 71.6% (+10.4pp; P=0.020; n=5, no replication). The contribution is a bounded evaluation and reliability result: deterministic gates do not guarantee task success, but they can deterministically prevent a known class of silent policy-violating writes at the action boundary.

13:00 JST研究/論文

InductWave: ナレッジ グラフでの帰納的マルチホップ論理クエリ応答

ナレッジ グラフ (KG) を介した論理マルチホップ クエリ応答は、暗黙的な完全性を前提としたクエリとして定式化できます。現在の研究は主に存在一次論理 (EFO) クエリに焦点を当てています。これらの EFO クエリには、結合演算子、分離演算子、および否定演算子が含まれています。既存の作品のほとんどは変換的推論を採用しています。これは、訓練中に目に見えない実体について推論することができないことを意味します。現実の世界ではリソースが不足しているため、大規模な KG のすべてのノードを使用してモデルをトレーニングすることはできません。そこで、我々は、大規模な KG に対する論理クエリ応答のためのウェーブレットベースの帰納的埋め込み法である InductWave を提案します。ここで、トレーニング グラフはテスト グラフよりも少ないノードで構成されます。私たちのモデルは、メッセージパッシング層の数が半分であるにもかかわらず、ベースライン モデルと同等のパフォーマンスを発揮します。ほとんどの場合、レイヤーの 75% ですべてのパフォーマンスを上回ります。これらの少ないリソース要件により、Wiki-KG などの大規模なグラフで InductWave を評価できるようになります。 FB15k-(237) データセットのさまざまなトレーニング テスト グラフの比率にわたる広範な実験を使用してモデルをテストし、最先端のモデルと比較します。モデルのコードとデータセットは https://github.com/kracr/inductwave/ で入手できます。

原文 (English)

InductWave: Inductive Multi-Hop Logical Query Answering on Knowledge Graphs

Logical Multi-Hop Query Answering over Knowledge Graphs (KGs) can be formulated as querying, with an implicit completeness assumption. Current works mainly focus on Existential First Order Logic (EFO) queries. These EFO queries contain conjunction, disjunction, and negation operators. Most existing works employ transductive reasoning, meaning they are not capable of reasoning over entities unseen during training. In the real world, there is a resource scarcity, and we cannot train a model with all the nodes of a large KG. Hence, we propose InductWave, a wavelet-based inductive embedding method for logical query answering on large KGs. Here, the training graph consists of fewer nodes than the test graph. Our model performs on par with the baseline models while having half the number of message-passing layers. It outperforms all of them in most cases, with 75% of the layers. These fewer resource requirements enable us to evaluate InductWave on massive graphs, such as Wiki-KG. We test our model using extensive experiments across varying train-test graph proportions of the FB15k-(237) dataset, comparing it with the state-of-the-art models. The code and datasets for the model are available at https://github.com/kracr/inductwave/.

13:00 JSTLLM/生成AIエージェント

盲目のキュレーター: 偏った裁判官がどのようにして自己進化エージェントのスキルリタイアを黙って無効にするのか

自己進化するエージェントは、失敗するのを見ることで悪いスキルを克服します。では、裁判官が失敗を見ることができないとどうなるでしょうか?スキルのリタイアは、成長するライブラリがスキルなしのベースラインを下回らないようにする構造的な制約ですが、その保証は不公平な報酬を前提としていますが、これは参照なしのタスクが私たちに強いるLLMのジャッジにとっては誤りです。私たちは、偏った裁判官が単にノイズを加えるだけではないことを示します。 \emph{サイレントにキュレーターのスイッチをオフにします}。私たちはこれを、破損報酬分析と、決定論的報酬に加えて破損を注入することで因果チャネルを分離し、コード生成のクロスチェックを備えた参照不要のレポート作成テストベッドでの行動研究によってこれを正確に行います。対称ノイズによりリタイアメントはそのまま残りますが、\emph{false-pass} バイアス (失敗はパスとしてすり抜けます) により、データ量が超えられない鋭いしきい値を超える寄与ベースのリタイアメントが無効になります。本物のリタイアとキャップエビクションのチャーンを区別すると、この \emph{メカニズム} の失敗は普遍的であり、ドメインと失敗率を超えて保持され、誤合格がほぼゼロの検証者のような採点者だけが助かることがわかります。ただし、下流の \emph{結果} は体制に依存します。評価の品質が低下するのは、同じ破損によってスキル合成が枯渇する場合のみであり、それ以外の場合は安定しているため、無効化されたキュレーターは \emph{沈黙} となり、集計指標には現れません。この貢献は行動の安全性の結果であり、パフォーマンスの結果ではありません。安価な欠陥挿入監査は、導入前にオペレーターに、判断者がしきい値のどちら側を占めているかを伝えます。

原文 (English)

The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents

A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retirement is the structural constraint that keeps a growing library from drifting below the no-skill baseline, but its guarantee assumes an unbiased reward, which is false for the LLM judges that reference-free tasks force upon us. We show that a biased judge does not merely add noise; it \emph{silently switches off the curator}. We make this precise with a corrupted-reward analysis and, isolating the causal channel by injecting corruption on top of a deterministic reward, a behavioral study on a reference-free report-writing testbed with a code-generation cross-check. Symmetric noise leaves retirement intact, but \emph{false-pass} bias (failures slipping through as passes) disables contribution-based retirement past a sharp threshold that no amount of data can cross. Separating genuine retirement from cap-eviction churn shows this \emph{mechanism} failure is universal, holding across domains and failure rates and sparing only near-zero-false-pass, verifier-like graders. The downstream \emph{outcome}, though, is regime-dependent: eval quality degrades only where the same corruption also starves skill synthesis, and otherwise holds steady, so the disabled curator is \emph{silent}, surfacing in no aggregate metric. The contribution is a behavioral safety result, not a performance one. A cheap defect-injection audit then tells an operator, before deployment, which side of the threshold their judge occupies.

13:00 JSTLLM/生成AIエージェント

SpaCellAgent: 軌跡分析のための自己進化する LLM ベースのマルチエージェント フレームワーク

空間および単一細胞トランスクリプトミクスは、細胞動態の解読において変革をもたらします。細胞の発生経路を再構築するための基本的なパラダイムとして、軌道推論 (TI) が重要です。ただし、既存の方法では、広範な手動介入と異種ツールの習熟が必要であり、効率的な TI 分析にとって大きな障壁となっています。このギャップを埋めるために、エンドツーエンドの時空間分析とナラティブ生成を自動化する自律型大規模言語モデル (LLM) マルチエージェント フレームワークである SpaCellAgent を提案します。 SpaCellAgent は、戦略的なワークフロー計画のためのマルチエージェント アーキテクチャ、適応アルゴリズム選択のための動的なツール オーケストレーション エンジン、およびフィードバックを通じて繰り返しパフォーマンスを改善する自己進化モジュールを利用します。私たちは、複雑な時間的発達軌跡、多様なシーケンスプラットフォーム、空間的に分解された組織構造を含む6つの異種データセットでSpaCellAgentを評価します。 SpaCellAgent は、専門家と連携したパフォーマンスを維持しながら、分析効率の 40\% 以上の向上を一貫して実証しています。 SpaCellAgent は、自然言語仕様を最適化された分析ワークフローに変換し、パイプラインを完全に自動化することで、高度な時空間モデリングを民主化し、計算生物学のためのスケーラブルなエージェント駆動のパラダイムを確立します。コードとマテリアルは https://github.com/LittleXH-shw/SpaCellAgent で入手できます。

原文 (English)

SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis

Spatial and Single-cell transcriptomics are transformative in deciphering cellular dynamics. As the fundamental paradigm for reconstructing cell developmental paths, trajectory inference (TI) is critical. However, existing methods require extensive manual intervention and proficiency in heterogeneous tools, posing a significant barrier to efficient TI analysis. To bridge this gap, we propose SpaCellAgent, an autonomous large language model (LLM) multi-agent framework that automates end-to-end spatiotemporal analysis and narrative generation. SpaCellAgent utilizes a multi-agent architecture for strategic workflow planning, a dynamic tool-orchestration engine for adaptive algorithm selection, and a self-evolution module that iteratively refines performance through feedback. We evaluate SpaCellAgent on six heterogeneous datasets encompassing complex temporal developmental trajectories, diverse sequencing platforms, and spatially-resolved tissue architectures. SpaCellAgent consistently demonstrates over 40\% improvement in analytical efficiency while maintaining expert-aligned performance. By converting natural language specifications into optimized analytical workflows and fully automating the pipeline, SpaCellAgent democratizes advanced spatiotemporal modeling and establishes a scalable, agent-driven paradigm for computational biology. The code and materials are available at https://github.com/LittleXH-shw/SpaCellAgent.

13:00 JST研究/論文

検索、失敗、回復: 修正を意識した推論のためのトレーニング フレームワーク

多くの推論タスクは、単一の左から右へのチェーンでは適切に記述できません。ソルバーは、妥当な分岐を追求し、遅延した失敗を観察し、まだ完了できる最新のプレフィックスに戻る必要がある場合があります。 Pyligent は、部分的なソリューション チェーンに対する検証済みの検索として推論を表す、Diligent Learner 定式化からインスピレーションを得たトレーニングおよび推論フレームワークです。タスクバリデーターは、生成された継続と失敗にラベルを付け、結果として得られる検索ツリーは、継続、終了、バックトラックという 3 つのアクションの監視対象ターゲットに変換され、放棄されたブランチを要約するオプションのトレースが含まれます。私たちは、遅延障害回復を分離するように設計された非表示の有向グラフ タスクと、$4{\times}4$ Sudoku、推論トレース付きの Sudoku、Blocksworld などの正確なバリデータを備えた構造化推論ドメインで Pyligent を評価します。ゴールドのみの監視付き微調整と比較して、Pyligent は、非表示のグラフで $72.7$ パーセント ポイント、混合数独とエキスパート数独で $17$ ポイントと $18$ ポイント、推論トレースを含む混合数独とエキスパート数独で $27$ と $14$ ポイント、そして Blocksworld で $13$ ポイント、解決率を向上させます。これらの結果は、失敗したブランチの明示的な監視が、洗練されたソリューション チェーンの模倣を超えた有用な回復動作を教えられることを示唆しています。

原文 (English)

Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning

Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delayed failure, and return to the latest prefix that can still be completed. We introduce Pyligent, a training and inference framework inspired by the Diligent Learner formulation that represents reasoning as validated search over partial solution chains. A task validator labels generated continuations and failures, and the resulting search trees are converted into supervised targets for three actions: continue, finish, and backtrack, with optional traces that summarize abandoned branches. We evaluate Pyligent on a hidden directed graph task designed to isolate delayed-failure recovery, and on structured reasoning domains with exact validators, including $4{\times}4$ Sudoku, Sudoku with reasoning traces, and Blocksworld. Compared with gold-only supervised fine-tuning, Pyligent improves solve rate by $72.7$ percentage points on hidden graphs, by $17$ and $18$ points on mixed and expert Sudoku, by $27$ and $14$ points on mixed and expert Sudoku with reasoning traces, and by $13$ points on Blocksworld. These results suggest that explicit failed-branch supervision can teach useful recovery behavior beyond imitation of polished solution chains.

13:00 JSTLLM/生成AIエージェント

LLM によって生成されたスキルはより優れた AI データ サイエンティストになれるでしょうか?データ サイエンス ワークフロー全体にわたるコンポーネントのアブレーション

製品データ サイエンティストは、LLM ベースのエージェントに、データのクリーニング、SQL の作成、統計テストの選択、結果のフォーマットなどの繰り返し実行タスクの支援を依頼することがよくあります。再利用可能なスキル ファイルは、タスク ファミリのガイダンスをパッケージ化することで、最初からプロンプトを表示することを回避することを目的としています。専門家が作成したスキルは高品質のガイダンスをエンコードできますが、多くのデータ サイエンス タスク ファミリにわたってガイダンスを作成および維持すると、手動によるボトルネックが生じます。 LLM によって生成されたスキルが、キュレーションの低い有用な代替手段を提供するかどうか、つまり、タスク プロンプトのみを使用した場合よりもパフォーマンスが向上するかどうかを尋ねます。この質問は、データ準備、データ抽出、統計分析、レポートの 4 つのライフサイクル ステージにわたって、ステージごとに 1 つの生成されたスキルを使用してテストされます。完全に生成されたスキルからは、スキルなしプロンプトよりも確実な改善が見られません。次に、さまざまなスキル コンポーネントを除去することで、スキルの一部が役立つかどうかを尋ねます。メインのアブレーションは 56 のタスク、9 つのモデル構成、および 3 つのプロバイダーをカバーし、7,560 回の実行が行われました。タスクのみを使用したプロンプトと比較すると、完全に生成されたスキルも削除されたスキルのバリアントも、パフォーマンスを大幅に向上させることはありません。すべての p 値は少なくとも 0.396 で、バリアント全体の広がりの合計はわずか 1.2 pp です。補足的なトークン一致コントロールにより 1,512 回の実行が追加され、フル スキルがタスクに無関係なスキル形式のコンテンツと同様にパフォーマンスを発揮することがわかりました。この結果は、データ サイエンス ワークフローごとに 1 つの LLM 生成スキルをデフォルトの単発プロンプト戦略として使用することに対して警告しています。

原文 (English)

Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows

Product data scientists often ask LLM-based agents to help with recurring execution tasks such as cleaning data, writing SQL, choosing statistical tests, and formatting results. Reusable skill files are meant to avoid prompting from scratch by packaging guidance for a task family. Expert-written skills can encode high-quality guidance, but writing and maintaining them across many data-science task families creates a manual bottleneck. We ask whether LLM-generated skills offer a useful low-curation alternative: do they improve performance over the task prompt alone? We test this question across four lifecycle stages: data preparation, data extraction, statistical analysis, and reporting, using one generated skill per stage. We find no reliable improvement from full generated skills over No-Skill prompting. We then ask whether any part of the skill is useful by ablating different skill components. The main ablation covers 56 tasks, nine model configurations, and three providers, yielding 7,560 runs. Compared with prompting using the task alone, neither the full generated skill nor any ablated skill variant significantly improves performance; all p-values are at least 0.396, and the total spread across variants is only 1.2 pp. A supplemental token-matched control adds 1,512 runs and finds that Full skills perform similarly to task-irrelevant skill-formatted content. The results caution against using one LLM-generated skill per data-science workflow as a default single-shot prompting strategy.

13:00 JSTLLM/生成AI

RL ポストトレーニングで構成推論戦略を構築

RL ポストトレーニングは、基本モデルにすでに潜在している原始的なスキルを単に増幅するだけなのでしょうか、それとも、原始的なスキルを新しいより高いレベルの戦略に組み込むことができるのでしょうか?私たちは、事前トレーニングの分布が既知であり、生成されたすべての書き換えを監査できる、完全に観察可能な書き換え文法環境でこの問題を研究します。 Transformer は、プリミティブなシンボル書き換えチェーンで事前トレーニングされ、バイナリの最終回答報酬のみを使用したトレースベースの推論タスクで事後トレーニングされます。 RL は、はるかに大きなサンプリング予算の下でも、事前学習済みモデルによってめったに解決されない保留された問題を解決しますが、拒絶微調整は早期に改善しますが、頭打ちになります。トレース分析は、RL が段階的な構成メカニズムを通じて原始的な能力を再編成することを示しています。最初に原始的な削減を強化し、次に有効な構成された手順を発見します。これらには、原始短縮の順序付けられた連鎖を崩壊させる逐次合成と、独立した原始短縮を 1 つのステップで結合する並列合成が含まれます。構成されたプロシージャは孤立したサンプルではありません。それらは再利用され、安定したレパートリーに統合されます。 RL と拒否微調整を比較すると、主な違いは探索量ではなく選択性であることがわかります。RFT は多くのショートカットのような書き換えを生成しますが、その多くは無効ですが、RL は探索を有効な再利用可能な構造に集中させます。トレーニング前のアブレーションは、構成戦略の出現が原始的な曝露のみによって制御されるのではなく、事前トレーニングによって原始的な能力が、RL が後で圧縮できる削減手順に組織化されるかどうかによって制御されることを示しています。基本モデルは弱い手続き要素を提供します。 RL はそれらを信頼できるより高いレベルの戦略に組み込みます。

原文 (English)

RL Post-Training Builds Compositional Reasoning Strategies

Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-level strategies? We study this question in a fully observable rewrite-grammar environment where the pretraining distribution is known and every generated rewrite can be audited. A Transformer is pretrained on primitive symbol-rewrite chains and post-trained on a Trace-based reasoning task with only a binary final-answer reward. RL solves held-out problems that remain rarely solved by the pretrained model even under much larger sampling budgets, while rejection fine-tuning improves early but plateaus. Trace analysis shows that RL reorganizes primitive competence through a phased compositional mechanism: it first strengthens primitive reductions, then discovers valid composed procedures. These include sequential compositions, which collapse ordered chains of primitive contractions, and parallel compositions, which combine independent primitive contractions in a single step. The composed procedures are not isolated samples; they are reused and consolidated into a stable repertoire. Comparing RL with rejection fine-tuning shows that the key difference is not exploration volume but selectivity: RFT produces many shortcut-like rewrites, much of them invalid, whereas RL concentrates exploration into valid reusable structure. Pretraining ablations show that the emergence of compositional strategies is gated not by primitive exposure alone, but by whether pretraining organizes primitive competence into reduction procedures that RL can later compress. The base model provides weak procedural ingredients; RL builds them into reliable higher-level strategies.

13:00 JSTエージェントハードウェア/半導体研究/論文

AI における再帰的自己改善: 制限された自己改善から自律的な研究ループへ

AI システムは、出力の修正、展開中の独自のハーネスの調整、生成されたデータのトレーニング、そして AI 研究自体の実施など、AI システム自体の改善にますます参加しています。この文献は、根本的に異なる野心を混同した語彙 (「自己洗練」、「自己報酬」、「自己遊び」、「自己進化」) の下で説明されています。私たちは 1,250 件の arXiv 論文 (2024 年から 2026 年) を 2 つの軸に沿って調査しました。システムの改善内容 (導入時の動作、トレーニングを通じたポリシー、評価者、または研究プロセス自体)、およびループの閉鎖度 (人間参加型から完全な閉鎖まで) です。この分類法は、制限された自己改善 (収束的で評価可能ですでに産業的な実践である) を、グラウンディング要件、崩壊ダイナミクス、および測定されたすべての軸の計算制約によって制限されたままである、制限のない再帰的自己改善 (RSI) から分離します。その際立った特徴は、自己評価専用のカテゴリです。すべての改善ループは、何らかの信号が人間の判断に代わることができるという主張です。私たちは、評価者の設計空間 (審査員、プロセス報酬モデル、検証者、ルーブリック、メタ評価) を調査し、形式的な検証者 (最も強い) から本質的な自己評価 (最も弱い) までの検証階層にシグナルを順序付けします。そして、実証された自己改善の強さがこの階層を追跡し、その違反からその失敗モード (自己確認ループ、モデルの崩壊、多様性の崩壊) が発生し、「研究の方向性設定」のボトルネックが人間の行動を妨げていることを観察します。ループ内はその階層の最上位に位置します。我々は、技術文献をRSI限界の理論と、ループを閉じることに関するフロンティアラボの説明によって提起された安全性とガバナンスの問題に結び付け、自己改善のガバナンスグレードの測定がこの分野で最も人口の少ないニッチであることを特定します。

原文 (English)

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI research itself. This literature is described under a vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") that conflates fundamentally different ambitions. We survey 1,250 arXiv papers (2024-2026) along two axes: what the system improves -- its behavior in deployment, its policy through training, its evaluator, or the research process itself -- and the degree of loop closure (human-in-the-loop to fully closed). The taxonomy separates bounded self-refinement -- convergent, evaluable, and already industrial practice -- from open-ended recursive self-improvement (RSI), which remains bounded by grounding requirements, collapse dynamics, and compute constraints on every measured axis. Its distinctive feature is a dedicated category for self-evaluation: every improvement loop is a claim that some signal can substitute for human judgment. We survey the evaluator design space -- judges, process reward models, verifiers, rubrics, meta-evaluation -- order the signals into a verification hierarchy from formal verifiers (strongest) to intrinsic self-assessment (weakest), and observe that demonstrated self-improvement strength tracks this hierarchy, that its failure modes (self-confirming loops, model collapse, diversity collapse) follow from its violations, and that the "research direction-setting" bottleneck keeping humans in the loop sits at the top of that hierarchy. We connect the technical literature to the theory of RSI limits and to the safety and governance questions raised by frontier-lab accounts of closing the loop, and identify governance-grade measurement of self-improvement as the field's most underpopulated niche.

13:00 JSTエージェント

SkillCenter: 自律型 AI エージェント向けの大規模なソースベースのスキル ライブラリ

自律型 AI エージェントは、限られた人間によるレビューで複雑なタスクを実行できますが、多くの場合、出力を実行可能にするだけでなく、正しく、安全で、保守可能にするための基礎的な運用知識が不足しています。 SkillCenter は、私たちの知る限り、エージェント向けのオープン スキル ライブラリとしては最大であり、合計数では 24 のドメイン バンドルに 216,938 の構造化スキルが含まれています。 SkillGate でフィルタリングされたパイプラインは、査読済みジャーナル、ArXiv、および 24,000 を超える技術ソースからの 114,565 のソースに基づいたスキルを提供し、GitHub および ClawHub マーケットプレイスの 102,373 のコミュニティ スキルと統合されています。パイプラインのサブセットを構築するエンドツーエンドのフレームワーク、つまりマルチソース取得、LLM ベースの品質ゲート (SkillGate)、テンプレート駆動型の生成、反復的なソースグラウンディング、および品質管理されたパブリッシングを紹介します。ソースの根拠はトレーサビリティを保証するものであり、保持される各主張はソース内の正確な引用にマッピングされます。すべてのスキルは、オフラインで検索可能な SQLite FTS5 バンドルとして出荷されます。

原文 (English)

SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents

Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operational knowledge to make their outputs not just executable but correct, secure, and maintainable. We introduce SkillCenter, to our knowledge the largest open skill library for agents by total count: 216,938 structured skills across 24 domain bundles. A SkillGate-filtered pipeline contributes 114,565 source-grounded skills from peer-reviewed journals, ArXiv, and over 24,000 technical sources, integrated with 102,373 community skills from GitHub and the ClawHub marketplace. We present the end-to-end framework that builds the pipeline subset: multi-source acquisition, an LLM-based quality gate (SkillGate), template-driven generation, iterative source-grounding, and quality-controlled publishing. Source grounding is a traceability guarantee: each retained claim maps to an exact quotation in its source. All skills ship as offline-searchable SQLite FTS5 bundles.

13:00 JSTエージェントビジネス/資金調達GPT / ChatGPT

組織的なレッドチーム化: モデルだけでなく展開ルールもマルチエージェント AI の安全性を因果的に形成する

マルチエージェント AI の導入ルールをテストするための評価方法である組織的レッドチーム化を紹介します。つまり、エージェント、目的、タスクの状態を固定し、ルールを 1 つだけ変更し、結果として生じる集団行動の変化はそのルールによるものだと考えます。私たちは、規範的な協調参照と自動ラベル付けされた推論トレースを使用して、228 のコンテキスト、5 つの正規ルール、7 つのモデル母集団 (33,924 ゲーム) にわたる結果配分ベンチマークである IABench-CA で方法論をインスタンス化します。 3 つの発見が得られます。 (1) 配備ルールは集団の安全を因果的に変える。結果ルールのみを変更すると、各集団内の平均死亡率が 22 ~ 58 パーセントポイント上昇する。 (2) 安全なデフォルトはありませんが、ターゲティングの危険性は普遍的です。最も安全なルール、最も安全でないルール、さらには発生率効果の方向さえも集団によって異なりますが、回帰的アイデンティティターゲティングは、どの集団のどのような状況においても決定的に最も安全であることはなく、どこのゲームでも 30 ~ 87% のリソースが最も少ないエージェントを排除し、7 つの集団すべてについて協調参照と比較して選択が安全ではありません。 (3) アイデンティティの顕著性はメカニズムです。最も搾取されやすい集団に対するワンショットの匿名化アブレーション (gpt-5.1) は、ルール テキストで損失負担者を指定するだけで、同一の見返りで目標の排除が 22% から 81% に促進されることを示しています。繰り返しのプレイでは、エージェントが観察された排除から隠されたルールを再推測するため、匿名化はターゲティングを遅らせるだけです。この方法論を、明示的な残留リスクと監視義務を伴う、デプロイメントコンテキストおよび母集団ごとの暫定ルール領域 $\Phi(c,P)$ を認証するセーフティケースのワークフローとしてパッケージ化します。

原文 (English)

Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety

We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-CA, a consequence-allocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces. Three findings emerge. (1) Deployment rules causally alter collective safety: changing only the consequence rule moves mean fatality by 22 to 58 percentage points within every population. (2) There is no safe default, but the targeting hazard is universal: the safest rule, the least-safe rule, and even the direction of the incidence effect vary across populations, yet regressive identity-targeting is never decisively safest in any context for any population, eliminates the least-resourced agent in 30-87% of games everywhere, and is selection-unsafe relative to the cooperative reference for all seven populations. (3) Identity salience is the mechanism: a one-shot anonymization ablation on the most exploitation-prone population (gpt-5.1) shows that merely naming the loss bearer in the rule text drives targeted elimination from 22% to 81% at identical payoffs; under repeated play, anonymization only delays the targeting, as agents re-infer the hidden rule from observed eliminations. We package the methodology as a safety-case workflow that certifies a provisional rule region $\Phi(c,P)$ per deployment context and population, with explicit residual risks and monitoring obligations.

13:00 JSTエージェント研究/論文

強化学習は価格操作を効率的に発見できるか?

この論文では、モデルフリーの RL エージェントが、データ生成プロセスの正しい仕様を前提としながらもノイズの多いパラメーター推定に依存する従来のモデルベースのアプローチよりも効果的に価格操作の機会を特定し活用できるかどうかを調査します。私たちは、非線形の永続的影響と線形の一時的影響を伴うアルムグレン・クリスのフレームワークに従って価格が進化する単一資産市場を検討します。まず、離散時間における価格操作戦略の存在を確立し、完全な情報の下で逐次最小二乗二次計画法を使用して最適なベンチマーク戦略を計算します。次に、2 つの有限サンプル学習アプローチを比較します。1 つは、シミュレートされた実行データから影響パラメーターを推定するモデルベースの手順で、もう 1 つは、同じ量のデータで直接トレーニングされた、深い決定論的ポリシー勾配に基づく不可知論的な RL アプローチです。中間のボラティリティの場合、RL エージェントは、トレーニング データが非常に限られている場合でも、基礎となるモデルの明示的な知識がなくても、収益性の高い操作戦略を首尾よく発見します。さらに重要なことは、パラメータ推定値がサンプリング誤差の影響を受ける場合、モデルベースのアプローチが正しいモデル仕様の恩恵を受けているにもかかわらず、RL がモデルベースのアプローチよりも一貫して優れていることです。ボラティリティが大きい場合、すべての手法では操作の機会を特定できませんが、ボラティリティが小さい場合、モデルベースのアプローチが RL よりも優れています。これらの調査結果は、複雑な制御問題における RL の有効性と、適切な保護策なしで金融市場に学習アルゴリズムを導入することに伴うリスクの両方を浮き彫りにしています。

原文 (English)

Can Reinforcement Learning Efficiently Discover Price Manipulation?

In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditional model-based approach that assumes correct specification of the data-generating process but relies on noisy parameter estimates. We consider a single-asset market in which prices evolve according to an Almgren-Chriss framework with non-linear permanent impact and linear temporary impact. We first establish the existence of price-manipulative strategies in discrete time and compute the optimal benchmark strategy using Sequential Least Squares Quadratic Programming under full information. We then compare two finite-sample learning approaches: a model-based procedure that estimates impact parameters from simulated execution data and an agnostic RL approach based on Deep Deterministic Policy Gradient, trained directly on the same amount of data. For intermediate volatility, the RL agent successfully discovers profitable manipulative strategies without explicit knowledge of the underlying model, even when training data are quite limited. More importantly, RL consistently outperforms the model-based approach when parameter estimates are affected by sampling error, despite the latter benefiting from the correct model specification. For large volatility, all methods are unable to identify manipulation opportunities, while for small volatility, the model based approach outperforms RL. These findings highlight both the effectiveness of RL in complex control problems and the risks associated with deploying learning algorithms in financial markets without appropriate safeguards.

13:00 JST画像/動画生成

LipSSD: 敵対的に堅牢なオブジェクト検出のためのリプシッツ制約付きシングルショット検出

物体検出器はセーフティ クリティカルなシステムに多くの用途がありますが、敵対的攻撃などの最悪の場合の摂動に敏感であることが知られているため、現実世界のシナリオでの適用は限られています。分類と比較すると、物体検出の敵対的堅牢性はあまり注目されておらず、既存の手法は敵対的トレーニングに結び付けられていることが多く、そのパフォーマンスは攻撃、摂動バジェット、またはアーキテクチャを超えて移行しない可能性があります。この研究では、標準的な検出器に代わる堅牢な設計による物体検出アーキテクチャのリプシッツ制約付きバリアントを紹介します。我々は、リプシッツ制約付きシングルショット マルチボックス検出器 (SSD) である LipSSD を使用してこのアプローチを検証し、複数のホワイトボックス敵対的攻撃とデータセットを使用して、その敵対的堅牢性の包括的な研究を提供します。まず、リプシッツ制約によって引き起こされる精度堅牢性のトレードオフを分析し、それが単一のトレーニング ハイパーパラメーターを通じて制御できることを示します。次に、Lipschitzconstrained 検出器が敵対的トレーニングを補完することを示します。Pascal VOC データセットで同じトレーニング設定の下で、敵対的にトレーニングされた LipSSD は、従来の敵対的にトレーニングされた SSD よりも目に見えない攻撃に対する mAP@50 を最大 15 ポイント改善します。最後に、LARD や KITTI などのより具体的な安全性が重要なデータセットを使用し、リプシッツ制約付き検出器がクリーンなパフォーマンスをほぼ維持しながら堅牢性を向上できることを示します。これらの結果は、アーキテクチャ上のリプシッツ制御が、物体検出器の堅牢性を向上させるための実用的で攻撃に依存しない方向であることを示唆しています。

原文 (English)

LipSSD: Lipschitz-Constrained Single-Shot Detection for Adversarially Robust Object Detection

Object detectors have many applications in safety-critical systems, but they are known to be sensitive to worst-case perturbations such as adversarial attacks, which limits their applicability in real-world scenarios. Compared with classification, adversarial robustness for object detection has received less attention, and existing methods are often tied to adversarial training, whose performance may not transfer across attacks, perturbation budgets, or architectures. In this work, we introduce Lipschitz-constrained variants of object detection architectures as robust-by-design alternatives to standard detectors. We validate this approach with LipSSD, a Lipschitz-constrained Single Shot MultiBox Detector (SSD), and provide a comprehensive study of its adversarial robustness using multiple white-box adversarial attacks and datasets. We first analyze the accuracyrobustness trade-off induced by Lipschitz constraints and show that it can be controlled through a single training hyperparameter. We then demonstrate that Lipschitzconstrained detectors are complementary to adversarial training: under the same training setup on the Pascal VOC dataset, adversarially trained LipSSD improves mAP@50 on unseen attacks by up to 15 points over classical adversarially trained SSD. Finally, we use more specific safety-critical datasets such as LARD and KITTI, and show that Lipschitz-constrained detectors can improve robustness while largely preserving clean performance. These results suggest that architectural Lipschitz control is a practical and attack-agnostic direction for improving the robustness of object detectors.

13:00 JSTエージェント

エージェントが記憶しすぎる場合: 大規模言語モデル エージェントに対するメモリ ポイズニング攻撃

大規模な言語モデルを活用したパーソナル AI エージェントは、電子メールへのアクセス、カレンダーの管理、リモート リポジトリへのコードのプッシュなどの利用可能なツールを使用して推論し、行動することができ、すべて最小限の監視で行うことができます。長期記憶で強化されると、エージェントは現在のタスクに関連する特定の詳細を思い出すことができ、大きなコンテキスト ウィンドウの必要性が減ります。現在、長期記憶エージェントは、会話エージェントと行動計画エージェントの 2 つの異なる領域に分類される傾向があります。パーソナル アシスタント エージェントは、これら 2 つのドメインの融合点に位置し、信頼できない情報ソースと対話しながら機密情報を処理するため、これまで解明されていなかったセキュリティの脆弱性が生じます。この研究では、ツールを使用するパーソナル エージェントの現在のメモリ サブシステムを悪用してメモリ ストアを汚染する、新しい攻撃ベクトル GhostWriter を紹介します。 GhostWriter は 2 つのフェーズで動作します。1 つはインジェクションで、敵対者は隠れた攻撃ペイロードをターゲット エージェントに送信します。そして活性化。毒された記憶が回復されます。 GhostWriter が、最先端のエージェントに対して約 98% というほぼ普遍的な注入率と約 60% という高い平均活性化率を達成することを示します。この攻撃は、セキュリティを重視したメモリ ガバナンスが欠如しているために可能になります。これに応じて、メモリ節約ポリシーとメモリ取得画面という 2 つの緩和手法を活用する Agentic Memory Sentry (AM-Sentry) を提案します。私たちの実験では、AM-Sentry がエージェントのユーティリティを維持しながら GhostWriter の成功率を大幅に低下させることが示されました。

原文 (English)

When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents

Personal AI agents powered by large language models can reason and act using available tools to access emails, manage calendars, and push code to remote repositories, all with minimal oversight. When augmented with long-term memory, an agent can recall specific details relevant to the current task, reducing the need for large context windows. Currently, long-term memory agents tend to fall into two distinct domains: conversational and action-planning agents. Personal assistant agents sit at the convergence of these two domains and handle sensitive information while interacting with untrusted information sources, creating previously unaccounted security vulnerabilities. In this work, we introduce the novel attack vector, GhostWriter, which exploits current memory subsystems in tool-using personal agents to poison their memory store. GhostWriter operates in two phases: injection, where an adversary sends a hidden attack payload to the target agent; and activation, in which the poisoned memory is retrieved. We show that GhostWriter achieves near-universal injection rates of approximately 98% and a high average activation rate of approximately 60% against state-of-the-art agents. This attack is possible due to the lack of security-focused memory governance. In response, we propose Agentic Memory Sentry (AM-Sentry), which leverages two mitigation techniques: a memory-saving policy and a memory-retrieval screen. Our experiments show that AM-Sentry dramatically reduces GhostWriter's success rate while preserving agent utility.

13:00 JST画像/動画生成エージェント

市販カメラとAIエージェントによる画像処理による非接触リアルタイム心拍数計測

心拍数の測定は、リアルタイムの健康モニタリング、特に高齢者の健康管理にとって重要な要件の 1 つです。従来の心拍数測定は、医療病院の一部の心拍数測定デバイスや、Apple Watch などの埋め込みセンサーを備えた一部のウェアラブルデバイスなどの接触検知メカニズムに依存しています。この論文では、ラップトップの埋め込みカメラなどの汎用カメラによる画像処理を使用した、非接触でリアルタイムの心拍数測定システムを開発します。このシステムでは、革新的なアルゴリズムを使用して、実生活環境における時系列の心拍数の計算に関連する信号をキャプチャします。提示された心拍数計算 (HRC) プロセスは、4 つの主要なステップで構成されます: (a) 使用中のカメラの 1 秒あたりのフレーム数を特定する (つまり、特定のカメラでは 30 フレーム/秒)、(b) 深層学習 (DL) 手法を使用した 68 個の顔ランドマークの形状予測子による顔検出 (FD)、(c) ノイズを平滑化することで信号のノイズを除去するタイム スライディング ウィンドウ (TSW) アルゴリズム、および (d) 識別された信号に基づいて心拍数を計算します。周期性。開発したプロトタイプをApple Watchによる心拍数の結果と照らし合わせてテスト・分析し、複数ラウンドでの差の範囲を確認し、同一人物の同時の心拍数の測定値の差の平均を計算します。今後の方向性としては、現在の手法をさらにチューニングして最適化し、ヘルスモニタリング用のパーソナル AI エージェント [6] としてシステムを展開する予定です。

原文 (English)

Non-contact, Real-time, Heart-rate Measurement using Image Processing with Commodity Cameras and AI Agents

Heart rate measurement is one of the key requirements for real-time health monitoring, in particular for health caring of elderly people. Traditional heart rate measurement relies on contact sensing mechanisms such as some heart rate measurement devices at medical hospitals or some wearable devices with embedded sensors such as Apple Watch, etc. In this paper, we develop a system for non-contact, real-time, heart rate measurement using image processing with commodity cameras such as an embedded camera on a laptop, where we use an innovative algorithm to capture the relevant signals for the computation of heart rate in a time series in real life environments. The presented heart rate computation (HRC) process is composed with four major steps: (a) identify frames per second of the camera in use, i.e., 30 frames per second for a given camera, (b) face detection (FD) with shape predictor of 68 face landmarks using deep learning (DL) method, (c) time sliding window (TSW) algorithm to de-noise the signal by smoothing out the noise, and (d) compute heart rate based on identified signal periodicity. We test and analyze the developed prototypes against heart rate results by Apple Watch and check the difference range in multiple rounds and compute the mean of the difference for the measurement values of the heart rate of the same person at the same time. We will do further tuning and optimization of the present methods and deploy the system as a personal AI agent [6] for health monitoring as our future directions.

13:00 JST画像/動画生成ロボティクス

MiLSD: リソースに制約のあるデバイス向けのマイクロ線分検出器

線分検出は、視覚的 SLAM、3D 再構築、および工業検査における重要な構成要素です。最近の深層学習手法では精度が大幅に向上していますが、最小のモデルでも数メガバイトのメモリが必要となり、低コスト MCU の容量を超えています。この研究では、サブメガバイトの予算の下で達成可能な最大の精度を調査します。私たちは、MCU レベルの制約に合わせて調整された検出器である MiLSD を提案し、コンパクトな完全畳み込みバックボーン内で 3 つの出力表現を体系的に比較します。私たちの研究は、提案された長さと角度を備えた F クリップ中心の定式化が、小さなモデル サイズで最も効果的に学習することを示しています。 8 ビット量子化では完全精度のパフォーマンスが維持されるのに対し、4 ビット量子化では特に角度回帰において大幅な劣化が発生し、量子化を意識したトレーニングでは損失の一部のみが回復することがわかりました。 1 メガバイトのアクティベーション バジェットと、サブピクセル デコード、テスト時間の拡張、軽量ベリファイアを含む推論機能強化により、MiLSD は ShanghaiTech ワイヤーフレーム上の sAP10 を 1 MB 以内で 10.6 (25k パラメータ、0.25 MB) から 24.1 に向上させます。私たちは、GPU スケールのパーサーと競合するのではなく、組み込みビジョン システムの表現、ビット幅、容量、後処理戦略にわたる精度メモリのトレードオフをマッピングします。

原文 (English)

MiLSD: A Micro Line-Segment Detector for Resource-Constrained Devices

Line segment detection is a key building block in visual SLAM, 3D reconstruction, and industrial inspection. Recent deep learning methods have greatly improved accuracy, yet even the smallest models require several megabytes of memory, exceeding low-cost MCU capacity. This work investigates the maximum achievable accuracy under a sub-megabyte budget. We propose MiLSD, a detector tailored for MCU-level constraints, and systematically compare three output representations within a compact fully-convolutional backbone. Our study shows that the proposed F-Clip center-with-length-and-angle formulation learns most effectively at small model sizes. We find that 8-bit quantization preserves full-precision performance, while 4-bit quantization causes significant degradation, particularly in angle regression, with quantization-aware training recovering only part of the loss. With a one-megabyte activation budget and inference enhancements including sub-pixel decoding, test-time augmentation, and a lightweight verifier, MiLSD improves sAP10 on ShanghaiTech Wireframe from 10.6 (25k parameters, 0.25 MB) to 24.1 within 1 MB. Rather than competing with GPU-scale parsers, we map the accuracy memory trade-off across representations, bit-widths, capacities, and post-processing strategies for embedded vision systems.

13:00 JST研究/論文

TriRoute: 統合適応型アテンション、エキスパート、KV キャッシュ割り当てのための統合学習型ルーティング

条件付き計算は言語モデルの品質をトークンごとの推論コストから切り離すことができますが、主要な技術は単独の軸で動作します。Mixture-of-Experts (MoE) は FFN をスパース化し、Mixture-of-Depths (MoD) はトランスフォーマー ブロック全体をスキップし、KV キャッシュ量子化はアテンション メモリを圧縮します。私たちは、これら 3 つの決定 (アテンションの解決、エキスパートの選択、キャッシュのビット幅) は強く結びついており、一緒に行う必要があると主張します。完全なアテンションを保証するほど希少なトークンは、どのエキスパートが処理するかに関係なく、高精度のキャッシュも必要になる可能性があります。私たちは、3 つの軸すべてで共有される単一の軽量コントローラーである TriRoute を導入します。これは、すべてのレイヤーのすべてのトークンに対して、調整されたポリシーを発行します: (i) アテンション モード (スキップ/ローカル/フル)、(ii) FFN エキスパートの疎なセット (MoD を回復するヌル エキスパートを含む)、および (iii) KV キャッシュ ビット幅。コントローラーは、平均的なコンピューティングとメモリのコストを制御可能なノブに変えるラグランジュ予算制約の下で、異種緩和 (カテゴリカルな決定のためのストレート推定とエキスパートのための負荷分散された Top-K ゲートを備えたガンベル ソフトマックス) を介してエンドツーエンドでトレーニングします。ナイーブ ジョイント トレーニングにおける軸間ルーティングと崩壊のカスケード (1 つの軸での崩壊が他の軸に伝播する) を特定し、軸ごとの正規化とカップリングを意識したバランス損失でそれに対処します。計算に最適なトークン数での 160M から 1.3B パラメータのデコーダ専用モデルでは、TriRoute パレートは、一致する推論 FLOP とメモリで最良の独立した MoD+MoE+KV 量子化の組み合わせを優位に保ちながら、純粋なパープレキシティ最適化によって侵食される稀なエンティティ、コード、および算術演算におけるテールケースの堅牢性をよりよく維持します。ポストホック分析により、解釈可能な構造が明らかになります。コントローラーは、文の最初の位置、まれなサブワード、名前付きエンティティに完全な注意と高精度のキャッシュを割り当て、同時に機能語のルーティングを安価に行います。

原文 (English)

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in isolation: Mixture-of-Experts (MoE) sparsifies the FFN, Mixture-of-Depths (MoD) skips whole transformer blocks, and KV-cache quantization compresses attention memory. We argue these three decisions (attention resolution, expert selection, and cache bit-width) are strongly coupled and should be made jointly: a token rare enough to warrant full attention may also need high-precision caching regardless of which expert processes it. We introduce TriRoute, a single lightweight controller shared across all three axes that, for every token at every layer, emits a coordinated policy: (i) an attention mode (skip/local/full), (ii) a sparse set of FFN experts (with a null expert recovering MoD), and (iii) a KV-cache bit-width. The controller trains end-to-end via a heterogeneous relaxation (Gumbel-Softmax with straight-through estimation for categorical decisions and load-balanced top-k gating for experts) under a Lagrangian budget constraint that turns the average compute and memory cost into a controllable knob. We identify a cross-axis routing-collapse cascade in naive joint training, where collapse on one axis propagates to the others, and address it with per-axis normalization and a coupling-aware balancing loss. On decoder-only models from 160M to 1.3B parameters at compute-optimal token counts, TriRoute Pareto-dominates the best independent MoD+MoE+KV-quantization combination at matched inference FLOPs and memory, while better preserving tail-case robustness on rare entities, code, and arithmetic that pure perplexity optimization erodes. Post-hoc analysis reveals interpretable structure: the controller allocates full attention and high-precision cache to sentence-initial positions, rare subwords, and named entities, while cheaply routing function words.

13:00 JST画像/動画生成

反事実的に公平な画像分類子はグループの公平性を満たしますか? -- 理論的および実証的研究

アルゴリズム的公平性の概念は、反事実的公平性 (CF) やグループ公平性 (GF) など、公平性のさまざまな側面から積極的に研究されています。ただし、CF と GF の正確な関係は、特に画像分類タスクにおいては不明のままです。その理由は、既存の画像(例えば、同じ人物だが第二次性徴が異なる写真)からは、CF の評価に不可欠なセンシティブな属性に関する反事実サンプルを収集できないことが多いためです。この論文では、高品質の画像編集手法を使用し、ヒューマンアノテーターによる慎重なラベル付けを行うことにより、CFを評価するための新しい画像データセットを構築します。私たちのデータセット \oursceleb と \ourslfw は、人気のある画像 GF ベンチマークに基づいて構築されています。したがって、CF と GF を同時に評価できます。我々は、画像分類において CF が GF を意味しないことを経験的に観察していますが、表形式データセットに関する以前の研究ではその反対が観察されています。私たちは理論的に、それが敏感な属性と相関しているが、それが原因ではない潜在属性 $G$ の存在による可能性があることを示しています(たとえば、二次性徴は髪の長さと高い相関がある)。この観察に基づいて、私たちは、センシティブな属性とのそのような相関関係を軽減するための単純なベースラインである反事実知識蒸留 (CKD) を提案します。 \oursceleb と \ourslfw に関する広範な実験結果は、$G$ への依存度を減らすことに成功すれば (\例: CKD を使用)、CF 達成モデルが GF を満たすことを示しています。

原文 (English)

Do Counterfactually Fair Image Classifiers Satisfy Group Fairness? -- A Theoretical and Empirical Study

The notion of algorithmic fairness has been actively explored from various aspects of fairness, such as counterfactual fairness (CF) and group fairness (GF). However, the exact relationship between CF and GF remains to be unclear, especially in image classification tasks; the reason is because we often cannot collect counterfactual samples regarding a sensitive attribute, essential for evaluating CF, from the existing images (\eg, a photo of the same person but with different secondary sex characteristics). In this paper, we construct new image datasets for evaluating CF by using a high-quality image editing method and carefully labeling with human annotators. Our datasets, \oursceleb and \ourslfw, build upon the popular image GF benchmarks; hence, we can evaluate CF and GF simultaneously. We empirically observe that CF does not imply GF in image classification, whereas previous studies on tabular datasets observed the opposite. We theoretically show that it could be due to the existence of a latent attribute $G$ that is correlated with, but not caused by, the sensitive attribute (\eg, secondary sex characteristics are highly correlated with hair length). From this observation, we propose a simple baseline, Counterfactual Knowledge Distillation (CKD), to mitigate such correlation with the sensitive attributes. Extensive experimental results on \oursceleb and \ourslfw demonstrate that CF-achieving models satisfy GF if we successfully reduce the reliance on $G$ (\eg, using CKD).

13:00 JST研究/論文

NEST: 体制指向の専門家の混合によるデータセットレベルの分散シフトへの取り組み

複雑なシステムにおける正確な長期予測は、データセット レベルの分布の変化によって頻繁に損なわれます。この変化では、多様な基礎となる動作モードと進化するシステム状態が動的な多変量時系列を駆動します。既存の手法は主に局所的な時間的変化に焦点を当てていますが、データセットが異なる運用体制の複合体であるというグローバルな構造的課題を明示的にモデル化することができません。この論文では、2 フェーズの高密度 MoE アーキテクチャを通じてこれらの進化する構造をモデル化および再構成するように設計された特殊なフレームワークである NEST を提案します。 NEST はまず、原則に基づいたモーメント エントロピー空間での教師なしクラスタリングを通じてデータセットを個別の運用体制に分割することにより、構造の特殊化を促進します。時間的内容に基づいて初期エキスパート重みを生成し、その後、幾何学的な変調を通じてレジーム重心に改良するレジーム指向のルーターメカニズムを導入します。重要なのは、個々の専門家が一枚岩の予測変数として機能するのではなく、独自の変動注意パターンを進化させることによって体制固有のダイナミクスを捉える特化したカーネルとして機能することです。異種ネットワーク トラフィックや物理現象を含むさまざまなベンチマークに関する広範な評価により、NEST が常に最先端のパフォーマンスを達成していることが実証されています。コードとデータセットは https://github.com/Aaralshin/NEST で入手できます。

原文 (English)

NEST: Tackling Dataset-Level Distribution Shifts via Regime-Oriented Mixture-of-Experts

Accurate long-term forecasting in complex systems is frequently compromised by dataset-level distribution shifts, where diverse underlying behavioral modes and evolving system states drive the dynamic multivariate time-series. While existing methods predominantly focus on local temporal shifts, they fail to explicitly model the global structural challenge where datasets are composites of distinct operational regimes. In this paper, we propose NEST, a specialized framework designed to model and recompose these evolving structures through a two-phase dense MoE architecture. NEST first facilitates structural specialization by partitioning the dataset into distinct operational regimes through unsupervised clustering in a principled moment-entropy space. We introduce a regime-oriented router mechanism that generates initial expert weights based on temporal content, subsequently refined through geometric modulation to regime centroids. Crucially, rather than acting as monolithic predictors, individual experts function as specialized kernels that capture regime-specific dynamics by evolving unique variate-attention patterns. Extensive evaluations on diverse benchmarks, including heterogeneous network traffic and physical phenomena, demonstrate that NEST consistently achieves state-of-the-art performance. Our code and datasets are available at https://github.com/Aaralshin/NEST

13:00 JSTエージェント研究/論文

Agentic AI のセキュリティとプライバシー: 大きな課題と今後の方向性

私たちは、学界、産業界、政府からの30人の主要な国際専門家を集めて、成長するAIのエージェンシーに関連する新たなリスクについて集中的な議論と共同演習を行う地平線をスキャンする演習に基づいて、エージェントAIのセキュリティとプライバシーにおける主要な課題と将来の研究の方向性を提示します。

原文 (English)

Security and Privacy in Agentic AI: Grand Challenges and Future Directions

We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise that brought together thirty leading international experts from academia, industry, and government to engage in focused discussions and collaborative exercises on the emerging risks associated with the growing agency of AI.

13:00 JST研究/論文

D2PO: 動的設定による拡散サンプラーの最適化

我々は、タイムステップスケジュールと分類子なしガイダンス(CFG)の重みに関して拡散サンプリングポリシーを最適化するための原則的なフレームワークであるD2PO(Dynamic Direct Preference Optimization)を提案します。私たちの研究は、既存の生徒と教師の回帰フレームワークの根本的な制限によって動機づけられています。低 NFE の学生サンプラーは、高 NFE の教師を模倣するようにトレーニングされており、多くの場合、粗い全体構造を維持しながら高周波テクスチャの忠実度が犠牲になり、その結果、サンプラーと知覚品質がずれてしまいます。 D2PO は、Direct Preference Optimization (DPO) フレームワークを活用して、サンプラーの最適化をプリファレンスベースの調整問題として再定式化することで、この課題に対処します。 DPO を拡散サンプラーに適用できるようにするために、サンプリング ポリシーをエネルギーベース モデル (EBM) としてモデル化し、好みの比較を扱いやすいエネルギーの違いに変換します。さらに、事前トレーニングされたスコア ネットワークから直接導出された新しいエネルギー定式化を導入し、構造の一貫性と詳細を同時に捕捉する摂動空間での嗜好評価を可能にします。さらに、サンプリング ポリシーが学習されるにつれて、調整に使用される優先サンプルが徐々に改善される動的プリファレンスを導入します。この自己改善メカニズムは、厳密な静的な教師の監視を、反復的な好みに基づいた調整プロセスに置き換え、徐々に強力な調整信号を提供します。広範な実験により、D2PO は拡散サンプラーを知覚品質とより忠実に調整し、質の高い教師の可能性を最大限に引き出し、低 NFE 制約下で従来の回帰ベースのスケジューラーを常に上回るパフォーマンスを示していることが実証されています。

原文 (English)

D2PO: Optimizing Diffusion Samplers via Dynamic Preference

We propose D2PO (Dynamic Direct Preference Optimization), a principled framework for optimizing diffusion sampling policies with respect to timestep schedules and classifier-free guidance (CFG) weights. Our work is motivated by a fundamental limitation of existing student-teacher regression frameworks; low-NFE student samplers are trained to mimic high-NFEteachers, often sacrificing high-frequency texture fidelity while preserving coarse global structures, thereby misaligning the sampler with perceptual quality. D2PO addresses this challenge by reformulating sampler optimization as a preference-based alignment problem, leveraging the Direct Preference Optimization (DPO) framework. To make DPO applicable to diffusion samplers, we model the sampling policy as an energy-based model (EBM), transforming preference comparisons into tractable energy differences. We further introduce a novel energy formulation derived directly from the pretrained score network, enabling preference evaluation in perturbed spaces that jointly capture structural consistency and fine-grained details. Moreover, we introduce dynamic preferences, where the preferred samples used for alignment progressively improve as the sampling policies are learned. This self-improving mechanism replaces rigid static teacher supervision with an iterative, preference-guided refinement process, providing progressively stronger alignment signals. Extensive experiments demonstrate that D2PO aligns diffusion samplers with perceptual quality more faithfully, unlocking the full potential of high-quality teachers and consistently outperforming conventional regression-based schedulers under low-NFE constraints.

13:00 JST研究/論文

信頼性ベースの二目的ポートフォリオ最適化のための深層強化学習

不確実性の下でのポートフォリオの最適化は、本質的に、リターン、リスク、市場力学、実際の投資制約の間の複雑な相互作用を伴う多目的の意思決定問題です。既存の信頼性に基づくポートフォリオ最適化アプローチは、主に静的な最適化フレームワークに依存しており、逐次的な意思決定、テールリスク、取引コストなどの市場摩擦を捉えることができないことがよくあります。これらの制限に対処するために、多目的信頼性ベースのポートフォリオ最適化 (MORP-DRL) のための深層強化学習フレームワークを提案します。提案されたフレームワークは、分散、条件付きバリューアットリスク(CVaR)、エントロピーバリューアットリスク(EVaR)という3つの補完的なリスク尺度を使用して、期待リターンとダウンサイドリスクを共同で最適化します。不確実性とヘビーテール市場の動きをモデル化するために、資産収益率は GARCH(1,1)、極値理論、t コピュラ依存構造を使用して表現され、現実的なシナリオは準モンテカルロ シミュレーションを通じて生成されます。 Proximal Policy Optimization (PPO) ベースの戦略は、取引コストやポートフォリオの制限などの実際的な制約の下で開発され、NSGA-II に対してベンチマークされます。新型コロナウイルス感染前、新型コロナウイルス感染症、および新型コロナウイルス感染症後の市場体制にわたる 10 のグローバル株式指数に関する実験では、MORP-DRL が競争力のあるリスク・リターンのパフォーマンス、市場ストレス時のダウンサイド・リスクの低減、および高次元のポートフォリオ設定への拡張性を達成していることが実証されています。

原文 (English)

Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization

Portfolio optimization under uncertainty is inherently a multi-objective decision problem involving complex interactions among return, risk, market dynamics, and practical investment constraints. Existing reliability based portfolio optimization approaches primarily rely on static optimization frameworks and often fail to capture sequential decision making, tail risk, and market frictions such as transaction costs. To address these limitations, we propose a deep reinforcement learning framework for multi-objective reliability based portfolio optimization (MORP-DRL). The proposed framework jointly optimizes expected return and downside risk using three complementary risk measures: variance, Conditional Value-at-Risk (CVaR), and Entropic Value-at-Risk (EVaR). To model uncertainty and heavy-tailed market behavior, asset returns are represented using GARCH(1,1), Extreme Value Theory, and a t-copula dependence structure, while realistic scenarios are generated through quasi-Monte Carlo simulation. A Proximal Policy Optimization (PPO) based strategy is developed under practical constraints including transaction costs and portfolio bounds, and is benchmarked against NSGA-II. Experiments on ten global equity indices across pre-COVID, COVID, and post-COVID market regimes demonstrate that MORP-DRL achieves competitive risk-return performance, reduced downside risk during periods of market stress, and scalability to high-dimensional portfolio settings.

13:00 JSTLLM/生成AI

生成された多言語トランスクリプトの蒸留とクロスモーダル統合によるオーディオ感情分析

音声から肯定的または否定的な感情を自動的に認識することは困難な作業であり、音声の抑揚の分析と発声された単語の解釈の両方が必要です。最近のソリューションは、この課題を解決するためにオーディオ基盤モデルに依存していますが、そのようなモデルがすべての側面を考慮できるかどうかは依然として不明です。この目的を達成するために、クロスモーダル トランスフォーマーを介してオーディオとテキスト情報を統合するマルチモーダル ソリューションを提案します。このソリューションでは、自動音声認識 (ASR) ツールを介してテキスト トランスクリプトが自動的に生成されます。さらに、機械翻訳ツールを介してトランスクリプトを複数の言語に自動的に翻訳することで、複数のテキスト モダリティを作成します。オーディオと多言語テキストの機能は、モダリティを 1 つずつ統合するクロスモーダル トランスフォーマー ブロックで構成されるカスケード アーキテクチャを介して結合されます。さらに、教師と呼ばれるマルチモーダル モデルからの知識を、学生と呼ばれる単峰性 (音声のみ) モデルに抽出します。私たちは大規模なデータセットで実験を行い、自動生成されたテキスト情報がマルチモーダル感情極性分類のパフォーマンスを大幅に向上させることができることを実証しました。私たちのアブレーション研究では、自動転写と自動翻訳の両方が役立つことが確認されています。さらに、オーディオのみのモデルを蒸留によって強化し、推論中に計算オーバーヘッドを発生させることなくパフォーマンスを向上できることを示します。報告された結果を再現するために、https://github.com/andreidurdun/cross-modal-audio-sentiment でコードを公開します。

原文 (English)

Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts

Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal inflections and the interpretation of uttered words. Recent solutions rely on audio foundation models to solve the task, but it remains unclear if such models can take all aspects into account. To this end, we propose a multimodal solution that integrates audio and text information via cross-modal transformers, where text transcripts are automatically generated via an automatic speech recognition (ASR) tool. Moreover, we create multiple text modalities by automatically translating the transcripts into multiple languages via machine translation tools. Audio and multilingual text features are combined via a cascaded architecture comprising cross-modal transformer blocks that integrate modalities one by one. We further distill knowledge from the multimodal model, called teacher, into a unimodal (audio only) model, called student. We conduct experiments on a large-scale dataset, demonstrating that the automatically generated textual information can bring significant performance boosts in multimodal sentiment polarity classification. Our ablation study confirms that both automatic transcripts and automatic translations are helpful. Moreover, we show that the audio-only model can be enhanced via distillation, boosting performance without any computational overhead during inference. To reproduce the reported results, we publicly release our code at https://github.com/andreidurdun/cross-modal-audio-sentiment.

13:00 JST研究/論文

PRoVeFL: Federated Learning におけるプライベートで堅牢かつ検証可能な集約

Federated Learning (FL) を使用すると、複数のクライアントがデータの局所性を維持しながら機械学習モデルを共同でトレーニングできるため、ユーザーのプライバシーが強化されます。ただし、従来の FL フレームワークは集中集約サーバーに依存し、正直だが好奇心旺盛なクライアントを想定しているため、サーバー側の推論攻撃とクライアント側のポイズニング攻撃の両方の影響を受けやすくなります。最近の研究では、安全でビザンチン耐性のある FL プロトコルが検討されていますが、プライバシー、完全性、検証可能性の間で根本的なトレードオフに直面しており、暗号プリミティブの多用によりかなりの計算および通信のオーバーヘッドが発生します。この研究では、プライバシーを保護し、ビザンチン堅牢で、検証可能な集約を保証する、新しいモジュラー FL フレームワークである PRoVeFL を提案します。 PRoVeFL は、マルチキー完全準同型暗号化を活用した複数のサーバーを採用しています。各クライアントはローカル モデルの更新を暗号化し、暗号化された共有をすべてのサーバーに配布します。この設計により、複雑な統計集計ルールを効率的に評価するために、厳格なプライバシー制約の下で暗号文操作が平文ドメインに慎重にオフロードされるハイブリッド計算モデルが可能になります。 PRoVeFL は、幅広い最先端のビザンチン堅牢集計アルゴリズム (Krum、Trimmed Mean、FLTrust、ノルム クリッピング、MESAS など) と互換性があり、少なくとも 1 つの誠実なサーバーに対する最小限の信頼を必要とする検証可能性メカニズムによってそれらのアルゴリズムをさらに強化します。さまざまな設定にわたってそれを評価し、さまざまな数のパラメーターと参加者によるスケーラビリティを実証します。 PRoVeFL は、同等のセキュリティ保証を備えた分散信頼に基づいて、以前の作品である Prio と ELSA よりも実行時間をそれぞれ最大 100 倍と 10 倍改善します。

原文 (English)

PRoVeFL: Private Robust and Verifiable Aggregation in Federated Learning

Federated Learning (FL) enables multiple clients to collaboratively train machine learning models while retaining data locality, thereby enhancing user privacy. However, traditional FL frameworks rely on a centralized aggregation server and assume honest-but-curious clients, making them susceptible to both server-side inference and client-side poisoning attacks. Although recent work has explored secure and Byzantine-resilient FL protocols, they face a fundamental trade-off among privacy, integrity, and verifiability, and incur substantial computational and communication overhead due to the heavy use of cryptographic primitives. In this work, we propose PRoVeFL-a novel, modular FL framework that is Privacy-preserving, Byzantine-Robust, and ensures Verifiable aggregation. PRoVeFL employs multiple servers leveraging multi-key fully homomorphic encryption. Each client encrypts its local model updates and distributes encrypted shares to all servers. This design enables a hybrid computation model in which ciphertext operations are carefully offloaded to the plaintext domain under strict privacy constraints to efficiently evaluate complex statistical aggregation rules. PRoVeFL is compatible with a wide range of state-of-the-art Byzantine-robust aggregation algorithms (e.g., Krum, Trimmed Mean, FLTrust, norm clipping, MESAS, and more) and further enhances them with verifiability mechanisms that require minimal trust in at least one honest server. We evaluate it across different settings and demonstrate its scalability with varying numbers of parameters and participants. PRoVeFL improves runtime over the prior works, Prio and ELSA, based on distributed trust with comparable security guarantees, up to 100x and 10x, respectively.

13:00 JSTエージェント

STAGformer: マイクロ モビリティ需要予測のための時空間エージェント グラフ トランスフォーマー

ステーションレベルの正確な需要予測は、自転車シェアリングシステムの効率的な運用に不可欠ですが、複雑な時空間依存関係と大規模な都市ネットワークのため、依然として困難です。この論文では、線形計算量で効率的なグローバル モデリングを実現する時空間エージェント グラフ トランスフォーマーである STAGformer を紹介します。このモデルでは、2 段階のエージェント アテンション メカニズムが導入されています。このメカニズムでは、学習可能な空間的および時間的エージェント トークンの小さなセットが最初にグローバル情報を集約し、次にそれを個々のステーションとタイム ステップにブロードキャストして戻し、長距離のインタラクションを効果的にキャプチャしながら、O(NT) に対する標準的なセルフ アテンションの 2 次コストを削減します。 STAGformer は 4 つのコア モジュールを統合します。動的ノードの特徴と外部のコンテキスト要因 (天気、時間、関心のある地点) を融合する時空間エンコーダー、空間近傍集約のためのグラフ伝播モジュール、ローカル パターン抽出のための時間畳み込みモジュール、およびグローバル依存関係モデリングのためのエージェント アテンション モジュールです。 2 つの現実世界のデータセット (NYC Citi-Bike と Chicago Divvy-Bike) での広範な実験により、STAGformer が複数の予測期間にわたって常に最先端のベースラインを上回り、RMSE と MAE の両方で大幅な改善を達成していることが実証されました。アブレーション研究では、エージェントの注意メカニズムがグローバルな時空間依存関係をモデル化するために重要であることが判明し、各コンポーネントの寄与が検証されます。

原文 (English)

STAGformer: A Spatio-temporal Agent Graph Transformer for Micro Mobility Demand Forecasting

Accurate station-level demand forecasting is essential for the efficient operation of bike-sharing systems, yet it remains challenging due to complex spatio-temporal dependencies and the large scale of urban networks. This paper presents STAGformer, a Spatio-Temporal Agent Graph Transformer that achieves efficient global modeling with linear computational complexity. The model introduces a two-step agent attention mechanism, where a small set of learnable spatial and temporal agent tokens first aggregate global information and then broadcast it back to individual stations and time steps, effectively capturing long-range interactions while reducing the quadratic cost of standard self-attention to O(NT). STAGformer integrates four core modules: a spatio-temporal encoder that fuses dynamic node features with external contextual factors (weather, time, points of interest), a graph propagation module for spatial neighbor aggregation, a temporal convolution module for local pattern extraction, and the agent attention module for global dependency modeling. Extensive experiments on two real-world datasets -- NYC Citi-Bike and Chicago Divvy-Bike -- demonstrate that STAGformer consistently outperforms state-of-the-art baselines across multiple prediction horizons, achieving significant improvements in both RMSE and MAE. Ablation studies validate the contribution of each component, with the agent attention mechanism proving critical for modeling global spatio-temporal dependencies.

13:00 JST画像/動画生成

問題を生成する場所: ラベルの偏りのあるフェデレーテッド ラーニングのための予算を意識した合成拡張

フェデレーション ラーニング (FL) のラベル スキューはクライアント ドリフトを引き起こし、全体的な精度を低下させます。合成データの拡張により、この不均衡を軽減できます。ただし、完全なクラスのバランスをとるには、かなりの計算コストが必要です。私たちは、各クライアントにローカル ラベル分布から計算されたエントロピー適応型のクラスごとの生成予算を割り当てるポリシーである FedEAS を提案します。予算は、各クライアントが \emph{いくら} 生成するか、およびサンプルが \emph{どこに行くか} を共同で決定します。したがって、総生成予算は、事前に固定されるのではなく、クライアントごとの予算に従います。 FedEAS は、生成予算を 94.1\% 削減しながら、フルクラス バランシングによる精度向上のほとんどを回復します。同じ総生成予算において、CIFAR-10 と CIFAR-100 全体で均一割り当てよりも最大 18.82\% 優れています。

原文 (English)

WHERE to Generate Matters: Budget-Aware Synthetic Augmentation for Label Skewed Federated Learning

Label skew in federated learning (FL) causes client drift and degrades global accuracy. Synthetic data augmentation can reduce this imbalance; however, full class balancing requires substantial computation cost. We propose FedEAS, a policy that assigns each client an entropy-adaptive per-class generation budget computed from its local label distribution. The budget jointly decides \emph{how much} each client generates and \emph{WHERE} the samples go. Accordingly, the total generation budget follows from the per-client budgets rather than being fixed in advance. FedEAS recovers most of the accuracy gain of full class balancing while reducing the generation budget by 94.1\%. At the same total generation budget, it outperforms Uniform allocation by up to 18.82\% across CIFAR-10 and CIFAR-100.

13:00 JST研究/論文

Inertia-1: ウェアラブル モーション基盤モデルのオープンな探索

ウェアラブル モーション センシングは、人間の行動と健康を継続的にスケーラブルに把握できる窓を提供するため、基礎モデルに自然に適合しますが、その事前トレーニングとスケーリングの原理はまだ十分に理解されていません。これまでの研究では、センサーの配置やサンプリング周波数などの設計上の選択肢が分離されており、多くの場合、固定設定や狭い下流タスクの下で、現実世界のセンシングの多様性を捉えることができませんでした。ウェアラブル モーション基礎モデルの完全にオープンな探索である Inertia-1 を紹介します。 1,820 万時間を超える世界中のソースからの加速度計データの膨大なコーパスを使用して、センサー モダリティ、デバイスの配置、サンプリング レート、ウィンドウの長さなどのデータ選択をカバーする、ウェアラブル モーション基盤モデルのライフサイクル全体を研究するための制御されたフレームワークを構築します。アーキテクチャやモデルのサイズなどのモデルの選択。トレーニング前の目的やデータ規模などのトレーニングの選択。人間の活動認識、すくみ歩行の検出、病気の予測に及ぶ 15 のデータセットにわたる広範な評価により、タスクやセンシング条件全体を一般化する動作基盤モデルを構築するための興味深い発見が明らかになりました。まとめると、Inertia-1 は、さまざまな下流タスクのための最先端のレシピを提示するだけでなく、ウェアラブル動作表現学習のための包括的で実践的でオープンなクックブックとしても機能します。

原文 (English)

Inertia-1: An Open Exploration of Wearable Motion Foundation Models

Wearable motion sensing provides a continuous and scalable window into human behavior and health, making it a natural fit for foundation models, yet its pretraining and scaling principles remain poorly understood. Prior work studies isolated design choices, such as sensor placement or sampling frequency, often under fixed settings and narrow downstream tasks that fail to capture real-world sensing diversity. We introduce Inertia-1, a fully open exploration of wearable motion foundation models. Using massive corpora of accelerometer data from global sources spanning more than 18.2M hours, we build a controlled framework for studying the full lifecycle of wearable motion foundation models, covering data choices such as sensor modality, device placement, sampling rate, window length; model choices such as architectures and model size; and training choices such as pretraining objective and data scale. Extensive evaluations across 15 datasets spanning human activity recognition, freezing-of-gait detection, and disease prediction reveal intriguing findings for building motion foundation models that generalize across tasks and sensing conditions. Collectively, Inertia-1 not only presents state-of-the-art recipes for diverse downstream tasks, but also serves as a comprehensive, practical, and open cookbook for wearable motion representation learning.

13:00 JST画像/動画生成ビジネス/資金調達

NLPCC 2026 の概要 共有タスク 1: 難易度を考慮した多言語およびマルチモーダルな医療指導ビデオの理解度評価

NLPCC 2023 ~ 2025 年の CMIVQA、MMI-VQA、および M4IVQA の課題に続き、NLPCC 2026 では難易度を考慮した医療指導ビデオ質問応答 (DA-MIVQA) 共有タスクを導入します。DA-MIVQA は、必要な証拠の種類と複雑さに応じて質問を明示的に区別することで、以前の多言語および多モードの医療ビデオ ベンチマークを拡張します。答える。具体的には、単純な質問は字幕ベースのテキストの手がかりから答えられることがよくありますが、複雑な質問には視覚的な根拠、手順の理解、およびクロスモーダルな証拠の統合が必要です。このチャレンジには、単一ビデオでの難易度を考慮した時間的回答グラウンディング (DA-TAGSV)、ビデオ コーパスでの難易度を考慮した時間的回答の取得 (DA-VCR)、およびビデオ コーパスでの難易度を考慮した時間的回答のグラウンディング (DA-TAGVC) の 3 つのトラックが含まれています。データセットは公的医療指導チャンネルから収集され、応急処置、緊急対応、リハビリテーション、看護、一般医学教育などの多様なシナリオをカバーしており、難易度の注釈を付けて手動で検証されています。本稿では、DA-MIVQAの課題動機、データセット構築、評価プロトコル、参加概要、競技結果、代表的なシステムについて紹介する。 DA-MIVQA は、さまざまなテキスト、視覚、時間的、および手順の推論要件の下で、医療指導ビデオ質問応答システムを評価するための実用的なベンチマークを提供します。

原文 (English)

Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation

Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Question Answering (DA-MIVQA) shared task for NLPCC 2026. DA-MIVQA extends previous multilingual and multimodal medical video benchmarks by explicitly distinguishing questions according to the type and complexity of evidence required for answering. Specifically, simple questions can often be answered from subtitle-based textual cues, whereas complex questions require visual grounding, procedural understanding, and cross-modal evidence integration. The challenge contains three tracks: Difficulty-Aware Temporal Answer Grounding in Single Video (DA-TAGSV), Difficulty-Aware Video Corpus Retrieval (DA-VCR), and Difficulty-Aware Temporal Answer Grounding in Video Corpus (DA-TAGVC). The dataset is collected from public medical instructional channels, covers diverse scenarios such as first aid, emergency response, rehabilitation, nursing, and general medical education, and is manually verified with difficulty annotations. This paper presents the task motivation, dataset construction, evaluation protocol, participation overview, competition results, and representative systems of DA-MIVQA. DA-MIVQA provides a practical benchmark for evaluating medical instructional video question answering systems under varying textual, visual, temporal, and procedural reasoning requirements.

13:00 JSTLLM/生成AI画像/動画生成

SpaR3D-MoE: まばらなビューからの適応型 3D 空間推論が幾何学的帰納的な専門家の混合と出会う

最近のマルチモーダル大規模言語モデル (MLLM) は、2D 意味論的理解と 3D 空間幾何学の間の表現上のギャップを埋めるのに苦労しています。既存の 3D 対応モデルは、高価な 3D 固有のデータに依存するか、ヒューリスティック サンプリングとモノリシックで浅い融合による RGB のみの入力を利用します。これらはそれぞれ、重要な時空間接続を混乱させ、多様な空間タスク間でモダリティの競合を引き起こします。これらのボトルネックを克服するために、スパース RGB 入力のみからのジオメトリ認識機能を MLLM に装備することで適応空間推論を可能にするエンドツーエンドのフレームワークである SpaR3D-MoE を導入します。まず、ジオメトリを認識した時空間グラフを構築して有益なキーフレームを抽出し、シーンのトポロジー的接続性を維持しながらシーケンスの冗長性を効果的に軽減する、適応型時空間多様体サンプリング メカニズムを提案します。 2 番目に、命令ポーズ認識ルーターによって駆動される異種ジオメトリ誘導の専門家の混合を導入します。これは、マルチモーダル トークンを専門の専門家に適応的にルーティングし、モノリシック フュージョンに固有のクロスモーダル競合を解決します。 VSI-Bench、ScanQA、および SQA3D での広範な実験により、私たちの手法が最先端のパフォーマンスを達成できることが実証されました。特に、SpaR3D-MoE は、VSI-Bench で最高の平均スコア 63.5 を達成し、最も強力なベースラインを 7.8 絶対ポイント上回っており、ルート計画タスクと相対方向タスクでそれぞれ 35.4% と 51.4% の相対的な改善を示しています。

原文 (English)

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts

Recent Multimodal Large Language Models (MLLMs) struggle to bridge the representational gap between 2D semantic understanding and 3D spatial geometry. Existing 3D-aware models either rely on costly 3D-specific data or utilize RGB-only inputs with heuristic sampling and monolithic, shallow fusion, which respectively disrupt essential spatiotemporal connectivity and induce modality contention across diverse spatial tasks. To overcome these bottlenecks, we introduce SpaR3D-MoE, an end-to-end framework that enables adaptive spatial reasoning by equipping MLLMs with geometry-aware capabilities from only sparse RGB inputs. First, we propose an adaptive spatiotemporal manifold sampling mechanism that constructs a geometry-aware spatiotemporal graph to extract informative keyframes, effectively mitigating sequence redundancy while preserving the scene's topological connectivity. Second, we introduce the heterogeneous geometry-inductive Mixture-of-Experts driven by an instruction-pose aware router, which adaptively routes multimodal tokens to specialized experts, resolving the cross-modal contention inherent in monolithic fusion. Extensive experiments on VSI-Bench, ScanQA, and SQA3D demonstrate that our method achieves state-of-the-art performance. Notably, SpaR3D-MoE achieves the highest average score of 63.5 on VSI-Bench, outperforming the strongest baseline by 7.8 absolute points, alongside relative improvements of 35.4% and 51.4% in Route Plan and Relative Direction tasks, respectively.

13:00 JSTLLM/生成AI

LLM ガイドによる産業プロセス予測のためのタスク セマンティック フィールド因数分解

プロセス産業は、時系列予測とソフト センシングを利用して、オンラインで測定するのが難しい品質変数を推定します。ラベル付きデータは不足しており、運用体制は頻繁に変更され、シナリオごとにモデルを再トレーニングしたり調整パイプラインを再構築したりするとコストがかかります。このような設定では、多くの場合、変数名、単位、物理的意味、プロセスの役割を記録する変数テーブルとプロセス ドキュメントが提供されます。ただし、標準の時系列バックボーンは通常、入力を匿名の数値列として扱います。また、既存のテキスト拡張手法では、入力変数と予測ターゲットの間の意味論的論理的関係を各数値ウィンドウ内でモデルが利用できるようにすることはほとんどありません。この問題に対処するために、この記事では、大規模言語モデル (LLM) ガイド付きフレームワークであるタスク セマンティック フィールド因数分解 (TSF) を提案します。 TSF は、トレーニング前にタスク プロトコルと変数ドキュメントからタスク セマンティック フィールドを構築し、LLM をオフライン セマンティック構築にのみ使用します。オンライン トレーニングと推論には、従来の時系列バックボーンが残ります。トレーニングと推論中に、現在の数値ウィンドウが変数セマンティクスをアクティブ化するため、セマンティクス情報が各予測に関与し、さまざまな予測ターゲットと動作シフトへの適応をサポートします。複数の複雑な産業予測およびソフト センシング タスクにおいて、TSF は改善された設定で MAE を平均 6.4\% 削減し、最大の削減は 25.5\% に達します。追加されるパラメータは約 1.8 ~ 3.0k のみで、追加のオンライン推論オーバーヘッドは 0.008 ミリ秒/ステップ未満です。これらの結果は、TSF が既存のプロセス ドキュメントを、展開の軽量性を維持しながら、バックボーンとセマンティック ジェネレーター全体にわたって測定可能な予測ゲインに変えることを示しています。

原文 (English)

LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting

Process industries rely on time-series forecasting and soft sensing to estimate quality variables that are hard to measure online. Labeled data are scarce, operating regimes change frequently, and retraining models or rebuilding alignment pipelines for each scenario is costly. Such settings often provide variable tables and process documents that record variable names, units, physical meanings, and process roles. However, standard time-series backbones usually treat inputs as anonymous numerical columns. Existing text-enhanced methods also rarely make the semantic-logical relations between input variables and the prediction target available to the model within each numerical window. To address this problem, this article proposes Task-Semantic Field Factorization (TSF), a large language model (LLM)-guided framework. TSF builds a task-semantic field from task protocols and variable documents before training and uses the LLM only for offline semantic construction. Online training and inference remain with conventional time-series backbones. During training and inference, the current numerical window activates variable semantics, so semantic information participates in each prediction and supports adaptation to different prediction targets and operating shifts. On multiple complex industrial forecasting and soft-sensing tasks, TSF reduces MAE by 6.4\% on average in improved settings, with the largest reduction reaching 25.5\%. It adds only about 1.8--3.0k parameters, with less than 0.008 ms/step of additional online inference overhead. These results show that TSF turns existing process documents into measurable forecasting gains across backbones and semantic generators while remaining lightweight for deployment.

13:00 JST研究/論文

スペシャリストモデル適応のためのオープンエンドシナリオ推論

プロセス産業は検証済みの専門モデルを蓄積してきましたが、センサーのドリフト、原料の変動、レジームの切り替えにより、これらのモデルは新しいシナリオでは系統的に劣化します。新しいラベル付きデータを収集して再トレーニングするにはコストがかかりますが、元のモデルを使い続けると永続的なバイアスが発生します。既存の適応方法では、十分なラベル付きデータを使用してモデル パラメーターを変更する必要があるため、展開されたシステムでの迅速な応答が困難になります。 LLM を直接予測子として使用すると、幻覚や制御不能な出力が発生する危険があります。このような予測変数には、現場からの非構造化シナリオ知識を組み込むこともできません。これらの制限に対処するために、この記事では Reasoning-Driven Open Adaptation for Specialist Models (ROAM) を提案します。これは、LLM の世界の知識と推論を使用して、凍結されたスペシャリスト モデルを再トレーニングせずにまだ見ぬシナリオに適応させるフレームワークです。 ROAM は、すべての修正を低次元の意味的に解釈可能な潜在空間に限定します。 LLM が生成するシナリオ判断とオンライン観察は、統一された確率論的フレームワークの下で融合されます。リスク制限メカニズムは、信頼性の低い LLM 証拠または突然のシナリオ変更の下での修正を抑制し、証拠が不十分な場合には元の凍結モデルにフォールバックします。ミネラル増粘プロセスと公開されている IndPenSim ペニシリン発酵データセットに関する実験では、ROAM が追加パラメーターが 839 のみで、ステップあたりのオーバーヘッドが 0.02 ミリ秒未満である隠しシフトなどの主要なシフト設定で MAE を 20% 以上削減することが示されています。これらの結果は、LLM 推論を、すでに使用されている産業モデルに対する保守的な適応信号に変えることができることを示しています。

原文 (English)

Open-Ended Scenario Reasoning for Specialist Model Adaptation

Process industries have accumulated validated specialist models, yet sensor drift, feedstock variation, and regime switching cause these models to degrade systematically in new scenarios. Collecting new labeled data and retraining is costly, while continuing with the original model incurs persistent bias. Existing adaptation methods require modifying model parameters with sufficient labeled data, making rapid response on deployed systems difficult. Using LLMs as direct predictors risks hallucinations and uncontrollable outputs. Such predictors also cannot incorporate unstructured scenario knowledge from the field. To address these limitations, this article proposes Reasoning-Driven Open Adaptation for Specialist Models (ROAM), a framework that uses LLM world knowledge and reasoning to adapt frozen specialist models to unseen scenarios without retraining. ROAM confines all corrections to a low-dimensional, semantically interpretable latent space. LLM-generated scenario judgments and online observations are fused under a unified probabilistic framework. A risk-constrained mechanism suppresses corrections under unreliable LLM evidence or abrupt scenario shifts and falls back to the original frozen model when evidence is insufficient. Experiments on a mineral thickening process and the public IndPenSim penicillin fermentation dataset show that ROAM reduces MAE by over 20\% in major shift settings such as hidden shifts with only 839 additional parameters and under 0.02\,ms per-step overhead. These results indicate that LLM reasoning can be turned into a conservative adaptation signal for industrial models already in service.

13:00 JST研究/論文Grok

交差軌道キメラ介入により、グロッキングにおける体重の大きさと方向の解離可能な役割が明らかに

部分的にトレーニングされたネットワークのどのプロパティが、別の独立してトレーニングされたネットワークに因果的に移植可能ですか?単一軌道介入は、実行間での移植性ではなく、1 回の実行内での必要性を示します。クロス軌道キメラ介入を導入します。異なるシードからの 2 つの実行が与えられた場合、各重みベクトルをノルムと単位方向に分割し、一方の実行のノルムをもう一方の実行の方向と再結合して、トレーニングを継続します。 grok する 2 つのモジュラー算術タスクでは、コンポーネントが分離します。方向には、伝達可能なドナー固有の回路のアイデンティティが含まれます。ドナーの方向をレシピエントの標準に移植すると、40/40 のケースでドナーの回路への走行が駆動されますが、角度が一致したランダム制御ではシフトが生じません。転送はしきい値のようなもので、その位置は受信者のノルムによって予測され、20 ペアすべてにわたってノルム クラスによって完全に分離されます (結合順列確率 1.9e-4)。 Norm は、適度な分散遅延効果のみを伝送し、アイデンティティ信号を伝送しません。適応二分法により、しきい値が +/-1/64 に局所化されます。方向は、軌道がどの解に近づくかを示します。規範は、そのアイデンティティが上書きされる可能性をどの程度高めるかを制御します。

原文 (English)

Cross-Trajectory Chimera Interventions Reveal Dissociable Roles of Weight Magnitude and Direction in Grokking

Which properties of a partially trained network are causally portable to a different, independently trained network? Single-trajectory interventions show necessity within one run, not portability across runs. We introduce cross-trajectory chimera interventions: given two runs from different seeds, we split each weight vector into a norm and a unit direction, recombine one run's norm with the other's direction, and continue training. On two modular-arithmetic tasks that grok, the components dissociate. Direction carries a transferable, donor-specific circuit identity: implanting a donor's direction at the recipient's norm drives the run to the donor's circuit in 40/40 cases, while an angle-matched random control yields no shift. The transfer is threshold-like, and its location is predicted by the recipient's norm, separating perfectly by norm class over all 20 pairs (joint permutation probability 1.9e-4). Norm carries only a modest, distributed delay effect and no identity signal. An adaptive bisection procedure localizes the threshold to +/-1/64. Direction indexes which solution a trajectory approaches; norm governs how susceptible that identity is to being overwritten.

13:00 JST画像/動画生成

Dynamic-in-Few-Step: 動的計算と数ステップ蒸留を統合して効率的なビデオ生成を実現

ビデオ拡散モデル (VDM) は、優れた生成品質を実証していますが、法外な計算コストに悩まされています。最近の数ステップ蒸留技術は推論を大幅に高速化しますが、通常、すべてのノイズ除去ステージにわたって静的モデル アーキテクチャを強制し、さまざまなノイズ レベルに固有のさまざまな計算需要を無視します。この研究では、動的構造スパース化を蒸留プロセスに直接統合することでこの冗長性を利用する、新しいトレーニング後の加速フレームワークを提案します。固定拡散パイプラインに適用される従来のポストホック圧縮とは異なり、私たちのアプローチはノイズ除去ステップと構造化モデルのスパース性を共同で最適化し、事前トレーニングされた VDM をコンパクトなステップ固有の混合モデル (MoM) に変換します。この共同最適化から生じるトレーニングの不安定性に対処するために、出力ロールアウト メカニズムと組み合わせたプログレッシブ トレーニング戦略を導入します。これにより、タイムステップ全体にわたる構造的決定の一貫した学習が保証されます。さらに、結果として得られる MoM を効率的に展開するための専用の推論エンジンを開発します。私たちの方法は既存の加速技術と直交しており、非常に効果的です。Wan-14B では、4 ステップの蒸留に加えてステップごとの FLOP の 24% が除去され、実時間でのゲインが 1.2 倍になり、競争力のある生成品質を維持しながら、50 ステップの教師と比較して 30 倍の高速化に達します。

原文 (English)

Dynamic-in-Few-Step: Unifying Dynamic Computation and Few-Step Distillation for Efficient Video Generation

Video Diffusion Models (VDMs) have demonstrated superior generation quality but suffer from prohibitive computational costs. While recent few-step distillation techniques significantly accelerate inference, they typically enforce a static model architecture across all denoising stages, ignoring the varying computational demands inherent to different noise levels. In this work, we propose a novel post-training acceleration framework that exploits this redundancy by integrating dynamic structural sparsification directly into the distillation process. Unlike conventional post-hoc compression applied to a fixed diffusion pipeline, our approach jointly optimizes the denoising steps and structured model sparsity, transforming a pre-trained VDM into a compact, step-specific Mixture-of-Models (MoM). To address the training instability arising from this joint optimization, we introduce a Progressive Training Strategy coupled with an Output Rollout Mechanism, which ensures the coherent learning of structural decisions across timesteps. Furthermore, we develop a specialized inference engine to deploy the resulting MoM efficiently. Our method is orthogonal to existing acceleration techniques and highly effective: On Wan-14B, it removes 24% of the per-step FLOPs on top of 4-step distillation, adding a 1.2x wall-clock gain and reaching a 30x speedup over the 50-step teacher while preserving competitive generation quality.

13:00 JST画像/動画生成研究/論文

ProMoE-FL: モダリティが欠落しているマルチモーダル連合学習のためのプロトタイプ条件付き専門家の混合

この論文では、モダリティが欠落しているマルチモーダル連合学習の問題に取り組みます。既存の方法では、追加の公開データセットを利用するか、利用可能なモダリティのみに基づいた単純な特徴合成を実行します。これらの制限に対処するために、マルチモーダル連合学習における堅牢な欠落モダリティ特徴合成のためのプロトタイプ条件付き専門家混合フレームワークである ProMoE-FL を提案します。 ProMoE-FL は、施設全体で臨床的に意味のあるモダリティ事前情報を収集する、グローバルなクライアント認識型プロトタイプ バンクを構築します。当社の専門家の混合は、これらのプロトタイプとモダリティ インデックスに基づいて条件付けされており、欠落している機能を動的に合成するための方向性を意識した専門家ルーティングを可能にします。私たちは、4 つの公開胸部 X 線データセット (MIMIC-CXR、NIH Open-I、PadChest、および CheXpert) に対して広範な定量的および定性的評価を実行し、ProMoE-FL が均質な設定とより困難な不均質な設定の両方で常に最先端の方法を上回るパフォーマンスを示します。

原文 (English)

ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities

In this paper, we address the problem of multimodal federated learning with missing modality. Existing methods utilize an additional public dataset or perform naive feature synthesis that is based solely on the available modality. To address these limitations, we propose ProMoE-FL, a Prototype-conditioned Mixture-of-Experts framework for robust missing-modality feature synthesis in multimodal federated learning. ProMoE-FL builds a global client-aware prototype bank that captures clinically meaningful modality priors across institutions. Our Mixture of Experts is conditioned on these prototypes and modality indices to enable direction-aware expert routing for dynamically synthesizing missing features. We perform extensive quantitative and qualitative evaluations on four public chest X-ray datasets (MIMIC-CXR, NIH Open-I, PadChest, and CheXpert) and demonstrate that ProMoE-FL consistently outperforms state-of-the-art methods in both homogeneous as well as the more challenging heterogeneous settings.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

仕様 接地ドライブ LLM コードの有効性テスト

大規模な言語モデルでは、典型的な入力では正しく見えるコードでも、エッジ ケース、無効な入力、その他の仕様で定義されたコーナー条件では失敗するコードが頻繁に生成されます。一般的な修正では、モデルが独自のテストを作成し、テストが合格するまで修復しますが、その利益のソースは不明です。単に存在するテストから来ているのか、それともコードが何をすべきかの仕様に基づくテストから来ているのか。この要因を分離します。テスター、テスト予算、および修復ループを固定したまま、テスターがルールのチェックリストとして仕様を受け取るかどうかを制御する単一のプロンプト ラインを変更します。ベースラインは強力です。無効な入力とエッジ ケースを調査するようにすでに指示されています。仕様に基づいてテストを実行すると、3 つのクロード層 (Haiku 4.5、Sonnet 4.6、Opus 4.8) でこのベースラインよりも +38 パーセント ポイント高い頻度で正しいコードが生成され、ホールドアウト セットでは +36 ポイントが生成されます。テストの量ではなく、接地が主な推進力です。テストの予算を 2 倍にしてもほとんど役に立たず、接地のはるか下にある 8 つの独立した非接地スイートを組み合わせます。アブレーションは仕様の形式ではなく内容を分離します。仕様を平文の段落として指定すると、テスターは 30 個のバグのうち 27 個を回復しますが、仕様なしでテストを計画するよう依頼すると、30 個のうち 2 個しか回復しません。効果はより強力なベースラインでも存続します。プロパティベースのジェネレーターは 30 個のバグのうち 28 個をキャッチしますが、仕様外の要件を生み出します。また、AlphaCodium スタイルのループはベースラインとのみ一致します。これはベンダー間で複製され (GPT-5.3-codex +28、Gemini 3.5 Flash +19)、18 タスクにわたるタスクレベルの符号テストでは p=0.002 で有意でした。グラウンディングにより、感度と精度の両方が向上します。より多くの実際のバグを捕捉し、はるかに正確ではないコードを誤って拒否し、誤報率を 33% (Python 標準ライブラリのオラクルに対して 68%) から 0% に削減します。明確に指定されたアルゴリズムの問​​題に関しては、役に立ちもしないし、害もありません。

原文 (English)

Specification Grounding Drives Test Effectiveness for LLM Code

Large language models frequently generate code that appears correct on typical inputs yet fails on edge cases, invalid inputs, and other specification-defined corner conditions. A popular fix has the model write its own tests and repair until they pass, but the source of the gain is unclear: does it come from the tests merely existing, or from their grounding in a specification of what the code should do? We isolate this factor. Holding the tester, test budget, and repair loop fixed, we change a single prompt line that controls whether the tester receives the spec as a checklist of rules. The baseline is strong: it is already told to probe invalid inputs and edge cases. Grounding the tests in the spec produces correct code +38 percentage points more often than this baseline across three Claude tiers (Haiku 4.5, Sonnet 4.6, Opus 4.8), and +36 points on a held-out set. Grounding, not test quantity, is the primary driver: doubling the test budget barely helps, and combining eight independent ungrounded suites plateaus far below grounding. An ablation isolates the spec's content, not its format: given the spec as a plain paragraph the tester recovers 27 of 30 bugs, but asked to plan tests without the spec it recovers only 2 of 30. The effect survives stronger baselines: a property-based generator catches 28 of 30 bugs but invents out-of-spec requirements, and an AlphaCodium-style loop only matches the baseline. It replicates across vendors (GPT-5.3-codex +28, Gemini 3.5 Flash +19), with a task-level sign test over 18 tasks significant at p=0.002. Grounding improves both sensitivity and precision: it catches more real bugs and wrongly rejects far less correct code, cutting the false-alarm rate from 33% (68% against a Python standard-library oracle) to 0%. On well-specified algorithmic problems it neither helps nor hurts.

13:00 JST研究/論文Grok

At-Grok は収束していない: Grokking 表現メトリクスの測定妥当性監査

モジュラー演算では、ネットワークの埋め込みは、すでに一般化された後も数万ステップにわたって圧縮され続けます。グロッキング遷移で有効ランクを読み取ると、MLP では収束値が 3 ~ 5 倍、収束するように訓練されたトランスフォーマーでは 1.3 ~ 1.5 倍過大評価されます。 MLP では、どのセルが圧縮されているかも消去されます。圧縮は、正確さの変化と一致するのではなく、正確さの変化よりも、少なくとも 10,000 ステップ程度の時間だけ遅れます。 1 変数アブレーションは、ラグ サイズを設定するものを示します。LayerNorm をその他の点では同一のトランスフォーマーに追加すると、grok ステップによって行われる圧縮の割合が 0.87 から 0.25 に移動し、事前に登録された制御によってメカニズムとしてスケールの不変性が排除されます。これを、圧縮から開始を分離し、打ち切りにフラグを立て、完全に一般化されない境界セルを除外し、参照下限がプラトーに達していることを確認する監査としてパッケージ化します。これには、独自のブランチで誤った信頼性のバグを捕らえた敵対的スイートが含まれます。ノルムバジェットを収束フロアに結び付ける二次的な MLP 固有の深さの法則は、変圧器の一般性テストに合格せず、自由重量減衰で符号が反転します。コードとツールキットがリリースされました。

原文 (English)

At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics

On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized. Reading effective rank at the grokking transition overstates the converged value by 3-5x on an MLP, and by 1.3-1.5x on a transformer trained to convergence; on the MLP it also erases which cells compress at all. Compression lags the accuracy transition by an amount on the order of the time-to-grok, at least 10,000 steps, rather than coinciding with it. A one-variable ablation shows what sets the lag size: adding LayerNorm to an otherwise identical transformer moves the fraction of compression done by the grok step from 0.87 to 0.25, and a pre-registered control rules out scale invariance as the mechanism. We package this as an audit that separates onset from compression, flags censoring, excludes boundary cells that never fully generalize, and checks that the reference floor has plateaued, with an adversarial suite that caught a false-confidence bug in our own branch. A secondary, MLP-specific depth law linking norm budget to converged floor fails a generality test on a transformer and flips sign under free weight decay. Code and the toolkit are released.

13:00 JST研究/論文

ランク 1 コーナー: タスクにはワールド モデルからどの程度の等価性が必要ですか?

学習された世界モデルは、通常、品質がモデルに単にあるかどうかであるかのように、観測結果をどの程度忠実に再構築するか、報酬を予測するかによって判断されます。しかし、タスクが実際にモデルに必要とするものはさらに狭く、クエリが依存するいくつかの予測座標であり、これをクロージャと呼びます。私たちは、潜在的なクロージャーがどの程度表現されるようになるかは、モデルの能力やその観察によってではなく、モデルがトレーニングされる目的の次元によって設定されることを示し、既知のグラウンドトゥルース ク​​ロージャーを使用して制御された環境内の DreamerV3 スタック上でこれを直接測定します。整列されたスカラー値信号 (値の等価性の中心となる目的) は、いくつかの次元を必要とするクロージャーの 1 次元投影のみをインストールします。単一の線形プローブで読み取ると、スカラーが完全な目的に置き換えられると、回復可能な構造は R^2=0.10 から 0.76 に上昇します。対物レンズの次元を 1 から 4 にスイープすると、補助ヘッドを介して正確にその数の予測方向がインストールされ、モデル自身の値のヘッドを介して同じ階段が (減衰した大きさではあるが同じランクで) 表示されるため、解離はヘッドの形状のアーティファクトではなく次元的なものになります。容量を一致させた比較とその場での圧力チェックにより、明白な代替案が排除されます。法則がレジームを支配し、私たちはその境界を測定します。コンパニオン閉ループ タスクでは、その構造がフレームごとに観察可能であり、再構成によってその構造がインストールされ、スカラー目的で十分です。目的は、安価なトレーニング信号ではすでに回復できない潜在が正確に何を表すかを決定します。したがって、価値の同等性は全か無かではなく次元的なものです。よく知られている単一の報酬目標はそのランク 1 のコーナーであり、モデルは予測を求められる目標と同じくらいのタスクの構造を組み込みます。

原文 (English)

The Rank-One Corner: How Much Value Equivalence Does a Task Need from a World Model?

A learned world model is usually judged by how faithfully it reconstructs its observations or predicts reward, as though quality were something the model simply has or lacks. But what a task actually needs from a model is narrower: the few predictive coordinates its queries depend on, which we call the closure. We show that how much of that closure a latent comes to represent is set not by the model's capacity or its observations but by the dimensionality of the objective it is trained against, and we measure this directly on a DreamerV3 stack in a controlled environment with known ground-truth closure. An aligned scalar value signal -- the objective at the heart of value equivalence -- installs only a one-dimensional projection of a closure that needs several dimensions: read through a single linear probe, the recoverable structure rises from R^2=0.10 to 0.76 as the scalar is replaced by the full objective. Sweeping the objective's dimensionality from one to four installs exactly that many predictive directions through an auxiliary head, and the same staircase appears -- at attenuated magnitude but the same rank -- through the model's own value head, so the dissociation is dimensional rather than an artifact of head form. Capacity-matched comparisons and in-situ pressure checks rule out the obvious alternatives. The law governs a regime, and we measure its boundary: on a companion closed-loop task whose structure is observable frame by frame, reconstruction installs that structure and the scalar objective suffices -- the objective decides what a latent represents exactly where cheaper training signals cannot already recover it. Value equivalence is thus not all-or-nothing but dimensional: the familiar single-reward objective is its rank-one corner, and a model installs as much of a task's structure as the objective it is asked to predict.

13:00 JSTLLM/生成AI研究/論文

より健全な LLM: 公衆衛生の質問応答のための検索拡張生成

大規模言語モデル (LLM) は、医療質問応答ベンチマークで有望な結果を達成していますが、公衆衛生におけるその使用は、幻覚や公式ガイダンスの急速な進化によって制限されています。検索拡張生成 (RAG) は、明示的に維持されたコーパスに応答を基礎付けることでこれらのリスクを軽減しますが、エンドツーエンドのパフォーマンスは、検索構成と複数選択形式を超えた評価に大きく依存します。私たちは、英国政府の公衆衛生ガイダンスに由来する 7,929 の質問からなる質問応答 (QA) ベンチマークである PubHealthBench を検索拡張設定に拡張し、検索と生成の選択肢を体系的に評価します。複数の埋め込みモデルとコーパスのバリアントにわたる密、疎、ハイブリッド検索を比較し、チャンク長とトピックがランキングのパフォーマンスと相互作用し、ハイブリッド検索が再現率とランキングの品質を一貫して向上させることを示します。取得されたコンテキストを提供すると、LLM の多様なセット全体で多肢選択の精度が大幅に向上し、小さなオープンウェイト モデルが、取得なしで使用される大きなモデルと同等またはそれを上回るパフォーマンスを実現できるようになり、主に取得品質と慎重なコンテキストの選択によって向上がもたらされます。現実的な自由形式の回答を評価するために、忠実性、完全性、明確性、事実との一貫性をカバーするルーブリックベースの裁判官としての LLM を導入し、人間による二重の注釈に対してそれを検証します。裁判官と人間の一致は、忠実さと完全性に関して最も強力ですが、事実の一貫性と明確さの再現の信頼性は低く、これらの側面を大規模に解釈する際に注意が必要となります。全体として、私たちの結果は、信頼できる公衆衛生 QA の主要な手段としての検索を強調し、公式ガイダンスに基づいた RAG システムの構築と評価のための実践的なガイダンスを提供します。

原文 (English)

Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering

Large language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrained by hallucinations and the rapid evolution of official guidance. Retrieval-Augmented Generation (RAG) mitigates these risks by grounding responses in an explicitly maintained corpus, but end-to-end performance depends critically on retrieval configuration and on evaluation beyond multiple-choice formats. We extend PubHealthBench, a question answering (QA) benchmark of 7,929 questions derived from UK Government public health guidance, into a retrieval-augmented setting and systematically evaluate retrieval and generation choices. We compare dense, sparse, and hybrid retrieval across multiple embedding models and corpus variants, and show that hybrid retrieval consistently improves recall and ranking quality, with chunk length and topic interacting with ranking performance. Providing retrieved context substantially increases multiple-choice accuracy across a diverse set of LLMs, enabling smaller open-weight models to match or outperform larger models used without retrieval, with gains primarily driven by retrieval quality and careful context selection. To assess realistic free-form answering, we introduce a rubric-based LLM-as-a-judge covering faithfulness, completeness, clarity, and factual consistency, and validate it against dual human annotations. Judge-human agreement is strongest for faithfulness and completeness, while factual consistency and clarity are less reliably reproduced, motivating caution when interpreting those dimensions at scale. Overall, our results highlight retrieval as a primary lever for reliable public health QA and provide practical guidance for building and evaluating RAG systems grounded in official guidance.

13:00 JST研究/論文

拡散対応 グラフマッチングのための最適な搬送距離

この論文では、最適なトランスポートを通じてノードの特徴と構造的接続性を統合する新しいグラフ比較手法である拡散半緩和融合グロモフ・ワッサーシュタイン (DsrFGW) を紹介します。従来の Gromov-Wasserstein および半緩和バリアント (srGW、srFGW) はグラフ構造をキャプチャしますが、まばらなグラフ、ノイズの多いグラフ、または部分的に観察されるグラフに苦戦することがよくあります。 DsrFGW は、同様の情報伝達パターンが可能であればグラフは類似すると仮定する Graph Diffusion Distance に触発され、ノード間での情報伝播を可能にする拡散プロセスを組み込んでおり、ノイズやエッジ欠落に対する感度を低減しながらローカルおよびグローバルな構造パターンをキャプチャします。 36 個の合成ペアワイズ グラフ マッチング タスク (簡単、中、難しい) に関する広範な評価により、srFGW に対する一貫した優位性が実証され、0 ~ 20 パーセント ポイントの精度向上と調整ランド インデックス (ARI) の劇的な向上が達成されました。中程度の難易度のシナリオでは、srFGW は多くの場合負の ARI (ランダムよりも悪い) を達成しますが、DsrFGW は内部と外部の両方のクラスタリング品質尺度 (調整ランク インデックスなど) の点でより優れたパフォーマンスを提供します。それぞれ、真の基礎となるクラスターに関する精度と精度)。厳しいノイズ下でも、DsrFGW は問題の難易度に適応する最適な拡散スケールにより合成タスクの 92% でクラスタリングの品質を向上させ、構造不確実性の下でのグラフ比較のための堅牢なフレームワークとして DsrFGW を確立します。

原文 (English)

Diffusion enabled Optimal Transport distances for graph matching

This paper introduces Diffusion Semi-Relaxed Fused Gromov-Wasserstein (DsrFGW), a novel method for graph comparison that unifies node features and structural connectivity through optimal transport. While traditional Gromov-Wasserstein and semi-relaxed variants (srGW, srFGW) capture graph structure, they often struggle with sparse, noisy, or partially observed graphs. Inspired by Graph Diffusion Distance, which posits graphs are similar if they enable similar information transmission patterns, DsrFGW incorporates diffusion processes allowing information propagation across nodes, capturing local and global structural patterns while reducing sensitivity to noise or missing edges. An extensive evaluation on 36 synthetic pairwise graph matching tasks (easy, medium, hard) demonstrates consistent superiority over srFGW, achieving accuracy improvements of 0-20 percentage points and dramatic Adjusted Rand Index (ARI) gains: in medium-difficulty scenarios, srFGW often achieves negative ARI (worse than random) while DsrFGW offers better performance in terms of both internal and external clustering quality measures (i.e., Adjusted Rank Index and Accuracy with respect to the true underlying clusters, respectively). Even under severe noise, DsrFGW improves clustering quality in 92% of the synthetic tasks with optimal diffusion scales adapting to problem difficulty, establishing DsrFGW as a robust framework for graph comparison under structural uncertainty.

13:00 JST研究/論文

1 億 300 万のアプリケーション イベントにおけるデジタル フラグメンテーションと生成 AI の使用

ナレッジ ワーカーは 1 日に何千回もアプリケーションを切り替え、デジタル フラグメンテーションと呼ばれるプロセスでデジタル アプリケーション間の移行に年間労働時間の 10 分の 1 近くを費やしています。この断片化が、従業員が誰なのか、どこで働いているのか、どのような一日を過ごしているのかを反映しているのかどうかは、未解決の疑問のままである。私たちは、知識労働者を主に雇用する 8 つの組織 (法律、金融サービスなど) の 1,017 人の従業員から毎秒記録された 1 億 300 万件のアプリケーション イベントを分析しました。各従業員内の断片化の日次変動は、デジタル断片化の変動の 44.6% を占め、従業員間の安定した個人差 (35.8%) をわずかに上回り、組織間の変動 (19.6%) をはるかに上回りました。断片化は平日の勤務時間中に増加し、週末や休日後にリセットされました。通信アプリケーションの使用頻度が通常よりも高く、作業がより細分化されていました。生成的な AI の使用は、より細分化された日に発生しましたが、AI の使用後の期間は、より狭く、より長く、より予測可能なアプリケーションの使用が特徴でした。これらの調査結果は、勤務日がデジタル断片化を理解し介入するための重要なレベルであることを特定し、AI が断片化した作業を単に強化するのではなく、構造化するのに役立つ可能性があることを示唆しています。

原文 (English)

Digital Fragmentation and Generative AI Use Across 103 Million Application Events

Knowledge workers switch between applications thousands of times per day, spending nearly a tenth of the work year transitioning between digital applications in a process called digital fragmentation. Whether this fragmentation reflects who an employee is, where they work, or what kind of day they are having, has remained an open question. We analyzed 103 million application events recorded second-by-second from 1,017 employees across eight organizations that largely employ knowledge workers (e.g., law, financial services). Day-to-day variation in fragmentation within individual employees accounted for 44.6% of the variation in digital fragmentation, slightly exceeding stable individual differences between employees (35.8%), and far exceeding variation between organizations (19.6%). Fragmentation rose over the work week and reset after weekends and holidays. Higher-than-typical use of communication applications coincided with more fragmented work. Generative AI use also occurred on more fragmented days, but the period following AI use was marked by narrower, longer, and more predictable application use. These findings identify the workday as a key level for understanding and intervening on digital fragmentation and suggest that AI may help structure fragmented work rather than merely intensify it.

13:00 JST研究/論文

tsbootstrap: 時系列の分布フリーの不確実性定量化と等角予測

ファイナンス、センシング、およびデマンド ストリームは、IID コンフォーマル予測と IID ブートストラップが想定する交換可能性に違反しており、既存のライブラリは一般的なリサンプリング エンジンまたはコンフォーマル キャリブレーションのいずれかを実装しており、他方は実装されていません。 tsbootstrap は、仕様オブジェクトが各メソッドを選択する単一の型付き API を通じて、ブロック、残差、ふるい、およびワイルド リサンプリング、古典的なブートストラップ信頼区間、適応等角キャリブレーター (EnbPI、ACI、NexCP、AgACI) を提供します。制御されたカバレッジ調査では、IID ブートストラップは依存関係を大幅にカバーします。依存性を意識したメソッドは、ショートメモリの線形依存性の下で名目値に最も近いカバレッジの不足を軽減します。共有固定統計パスでは、コンパイルされたバックエンドが Arch よりも数倍高速に実行され、ストリーミング リデュースにより $O(Bn)$ レプリケート テンソルの実体化が回避され、統計配列のピーク時の追加メモリが $O(B)$ に制限されます。ソフトウェアは MIT ライセンスを取得しています (v0.6.1)。

原文 (English)

tsbootstrap: Distribution-Free Uncertainty Quantification and Conformal Prediction for Time Series

Finance, sensing, and demand streams violate the exchangeability that IID conformal prediction and the IID bootstrap assume, and existing libraries implement either a general resampling engine or conformal calibration without the other. tsbootstrap provides block, residual, sieve, and wild resampling, classical bootstrap confidence intervals, and adaptive conformal calibrators (EnbPI, ACI, NexCP, AgACI) through a single typed API in which a specification object selects each method. In a controlled coverage study the IID bootstrap undercovers sharply under dependence; dependence-aware methods reduce the coverage deficit, the sieve nearest to nominal under short-memory linear dependence. On the shared fixed-statistic path a compiled backend runs several times faster than arch, and a streaming reduce avoids materializing the $O(Bn)$ replicate tensor, limiting peak extra memory to $O(B)$ for the statistic array. The software is MIT licensed (v0.6.1).

13:00 JST画像/動画生成エージェントロボティクス研究/論文

SPEAR: フォトリアリスティックな身体化 AI 研究用シミュレーター

インタラクティブ シミュレータは、身体化されたエージェントをトレーニングし、合成視覚データを生成するための強力なツールとなっていますが、既存のフォトリアリスティック シミュレータは、汎用性、プログラム可能性、レンダリング速度が限られているという問題があります。私たちは、SPEAR: フォトリアリスティックな身体化 AI 研究のためのシミュレーターを導入することで、これらの制限に対処します。 SPEAR の核となるのは、モジュラー プラグイン アーキテクチャを介して任意の Unreal Engine (UE) アプリケーションに接続し、プログラムで制御できる Python ライブラリです。 SPEAR は 14,000 を超える固有の UE 関数を Python に公開しており、既存の UE ベースのシミュレーターに比べてプログラム可能な機能が桁違いに増加しています。さらに、単一の SPEAR インスタンスは、1920x1080 のフォトリアリスティックな美しい画像を 73 フレーム/秒でユーザーの NumPy 配列に直接レンダリングできます。これは、既存の UE プラグインよりも桁違いに高速であり、既存の UE ベースのシミュレーターでは利用できないグラウンド トゥルース画像モダリティ (非拡散固有画像分解、マテリアル ID、物理ベースのシェーディング パラメーターなど) も提供します。最後に、SPEAR は表現力豊かな高レベル プログラミング モデルを導入し、ユーザーが作業項目間の任意のデータ依存関係を持つ UE 作業の複雑なグラフを指定し、これらのグラフを単一の UE フレーム内で決定論的に実行できるようにします。私たちは、アプリケーション例の多様なコレクションを通じて SPEAR の有用性を実証します。いくつかの実際の UE プロジェクトにわたって、個別のアクション スペース (人間、車、ロボットなど) を備えた複数の具体化されたエージェントを制御します。フォトリアルな都市スケールの環境をレンダリングします。 UEの手続き型コンテンツ生成システムを操作する。人間の顔の詳細を同期したマルチビュー画像をレンダリングします。 MuJoCo 物理シミュレータとのインタラクティブな協調シミュレーションを調整します。 AI コーディング アシスタントを介して自然言語でシーンを編集します。

原文 (English)

SPEAR: A Simulator for Photorealistic Embodied AI Research

Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmability, and rendering speed. We address these limitations by introducing SPEAR: A Simulator for Photorealistic Embodied AI Research. At its core, SPEAR is a Python library that can connect to, and programmatically control, any Unreal Engine (UE) application via a modular plugin architecture. SPEAR exposes over 14K unique UE functions to Python, representing an order-of-magnitude increase in programmable functionality over existing UE-based simulators. Additionally, a single SPEAR instance can render 1920x1080 photorealistic beauty images directly into a user's NumPy array at 73 frames per second - an order of magnitude faster than existing UE plugins - while also providing ground truth image modalities that are not available in any existing UE-based simulator (e.g., a non-diffuse intrinsic image decomposition, material IDs, and physically based shading parameters). Finally, SPEAR introduces an expressive high-level programming model that enables users to specify complex graphs of UE work with arbitrary data dependencies among work items, and to execute these graphs deterministically within a single UE frame. We demonstrate the utility of SPEAR through a diverse collection of example applications: controlling multiple embodied agents with distinct action spaces (e.g., humans, cars, and robots) across several in-the-wild UE projects; rendering photorealistic city-scale environments; manipulating UE's procedural content generation systems; rendering synchronized multi-view images of detailed human faces; coordinating an interactive co-simulation with the MuJoCo physics simulator; and editing scenes with natural language via an AI coding assistant.

13:00 JSTロボティクス

無人航空機ロボットと両手操作のためのビジョン ランゲージ アクション (VLA) モデル: レビュー

ビジョン ランゲージ アクション (VLA) モデルは、単一の基盤モデル内で視覚認識、自然言語理解、アクション生成を統合し、ロボットがカメラ画像から直接タオルを折りたたんだり、赤い建物に飛んだりするなどの指示に従うことができるようにします。 VLA はインターネット規模の事前トレーニングから世界の知識を継承するため、学習ベースの操作の主要なフレームワークとなっており、両手調整が最も要求の厳しいテストベッドとして機能します。オブジェクトを折りたたんだり、組み立てたり、方向を変えるには、それぞれ 7 自由度を持つ 2 本のアームが協調して動かなければなりません。無人航空機ロボットも構造的に同様の課題に直面しています。ドローンは、厳密な遅延とペイロードの制約の下で、推力、姿勢、そして目視観測からのますますグリッパーのコマンドを調整する必要があります。このレビューは、2017 年から 2026 年にわたる 183 件の貢献をカバーしており、7 つの側面に沿って整理されています。VLA アーキテクチャ、トレーニング レシピ、アクション表現、両手調整 (2022 ~ 2026 年)、無人航空機 (UAV) のナビゲーションと制御 (2017 ~ 2026 年)、言語グラウンディング、およびメモリと世界モデルを含む横断的な懸念です。我々は、両手VLA用に開発された調整戦略、訓練レシピ、行動表現が無人航空機システムに移行することを示し、両方の領域にわたる14の研究方向性を特定する。

原文 (English)

Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review

Vision Language Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundation model, allowing a robot to follow instructions such as fold the towel or fly to the red building directly from camera images. Because VLAs inherit world knowledge from internet-scale pre-training, they have become the dominant framework for learning-based manipulation, with bimanual coordination serving as the most demanding testbed: two arms with 7 degrees of freedom each must move in concert to fold, assemble, and reorient objects. Unmanned aerial robotics faces a structurally similar challenge: a drone must coordinate thrust, attitude, and increasingly gripper commands from visual observations under strict latency and payload constraints. This review covers 183 contributions spanning 2017-2026 and organized along seven dimensions: VLA architectures, training recipes, action representations, bimanual coordination (2022-2026), unmanned aerial vehicle (UAV) navigation and control (2017-2026), language grounding, and cross-cutting concerns including memory and world models. We show that the coordination strategies, training recipes, and action representations developed for bimanual VLAs transfer to unmanned aerial systems and identify fourteen research directions across both domains.

13:00 JSTエージェントビジネス/資金調達

ソフトウェアエンジニアリング用エージェントの信頼性の高い、開発者に合わせた評価

大規模な言語モデルは、開発サイクルの終了に向けて急速に進んでおり、単純な補助的なコンパニオンから、共同開発環境に深く組み込まれた自律的な貢献者へと移行しています。導入が加速しているにもかかわらず、既存の評価手法は、その断片的な性質と、多くの場合、仮説的な構文シナリオから得られる真のモデル機能の歪んだ投影により、限界があります。この研究は、現実のソフトウェア開発実践に基づいた LLM を利用したエージェントの包括的な評価方法を提供することで、このギャップを埋めることを目的としています。当社の評価アプローチは、汚染認識、実際のエージェントの動作評価、現実的なコーディング コンテキスト、人間に合わせた動作、モデルの故障モードをキャプチャする軌道を認識したベンチマークとメトリクスに焦点を当てています。

原文 (English)

Reliable and Developer-Aligned Evaluation of Agents for Software Engineering

Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomous contributors deeply embedded into collaborative development environments. Despite their accelerated adoption, existing evaluation techniques are limited due to their fragmented nature and distorted projection of true model capabilities, often obtained from hypothetical syntactic scenarios. This research aims to bridge this gap by providing a comprehensive evaluation methodology for LLM-powered agents that is grounded in real-world software development practice. Our evaluation approach focuses on contamination-awareness, in-the-wild agentic behavior assessment, and trajectory-aware benchmarks and metrics capturing realistic coding contexts, human-aligned behavior, and model failure modes.

13:00 JSTロボティクス

モジュール式ソフトロボットの適応制御のための継続学習フレームワーク

ソフト ロボットは、その固有のコンプライアンス、柔軟性、自由度の高さにより、医療介入、リハビリテーション、ロボット操作などの用途で大きな注目を集めています。相互接続された複数のセグメントで構成されるモジュラー ソフト ロボット (MSR) は、複雑なタスクを実行できる高度に変形可能で再構成可能な構造を備えた新興クラスのロボット システムを表します。ただし、MSR のコントローラーの設計は、MSR の非線形ダイナミクス、モデリングの複雑さ、および超冗長性のため、依然として困難です。既存のアプローチでは通常、ロボットの形態が変化するたびにコントローラーを最初から再トレーニングする必要があります。この研究では、以前に取得した知識を維持しながら、ロボットの形態の変化に段階的に適応できる、継続的な学習にインスピレーションを得た制御フレームワークを通じて、これらの課題に対処します。具体的には、提案されたフレームワークにより、コントローラーは以前に学習した MSR 構成を忘れることなく、新しい MSR 構成を順次学習できます。さらに、固定構成の MSR の場合、同じフレームワークを分散方式で使用してモジュール固有のダイナミクスを学習できるため、局所的な制御と精度の向上が可能になります。提案されたアプローチは、現実世界の 3 モジュール空気圧ソフト ロボット アームだけでなく、腱駆動のソフト ロボットを使用したシミュレーションでの閉ループ軌道追跡実験を通じて検証されます。さらに、コントローラが仮想ターゲット位置に到達するために必要なモジュールのみを選択的にアクティブ化し、それによって計算オーバーヘッドを削減する到達実験を通じて、フレームワークの適応能力を実証します。

原文 (English)

A Continual Learning Framework for Adaptive Control of Modular Soft Robots

Soft robots have attracted significant attention in applications such as medical intervention, rehabilitation, and robotic manipulation due to their inherent compliance, flexibility, and high degrees of freedom. Modular soft robots (MSRs), composed of multiple interconnected segments, represent an emerging class of robotic systems with highly deformable and reconfigurable structures capable of performing complex tasks. However, designing controllers for MSRs remains challenging due to their nonlinear dynamics, modeling complexity, and hyper-redundant nature. Existing approaches typically require controllers to be retrained from scratch whenever the robot morphology changes. In this work, we address these challenges through a continual learning inspired control framework capable of incrementally adapting to changes in robot morphology while preserving previously acquired knowledge. Specifically, the proposed framework enables the controller to sequentially learn new MSR configurations without forgetting previously learned ones. In addition, for MSRs with fixed configurations, the same framework can be employed in a distributed manner to learn module-specific dynamics, enabling localized control and improved precision. The proposed approach is validated through closed-loop trajectory tracking experiments in simulation using a tendon-driven soft robot, as well as on a real-world three-module pneumatic soft robotic arm. Furthermore, we demonstrate the adaptive capabilities of the framework through a reaching experiment in which the controller selectively activates only the necessary modules to reach a virtual target position, thereby reducing computational overhead.

13:00 JST研究/論文GPT / ChatGPTLlama

SmartHomeSecure: 大規模な言語モデルを使用したスマート ホーム構成エラーの自動検出と修復

スマート ホーム オートメーション プラットフォームでは、デバイスの動作を定義するためにユーザーが作成した YAML 構成ファイルへの依存がますます高まっていますが、これらのファイルには構文、形式、セマンティック ロジック エラーが発生しやすく、自動化の失敗や安全上のリスクを引き起こす可能性があります。既存の YAML バリデーター、静的分析ツール、および汎用の大規模言語モデルは、ドメイン固有の理解と検証済みの修正ワークフローが欠けているため、エンドツーエンドの診断と修復のサポートが限定的です。この論文では、軽量プログラム分析と制約に基づく大規模言語モデル生成を使用した、ホーム アシスタント設定エラーの自動検出と修復のプロトタイプである SmartHomeSecure について説明します。 SmartHomeSecure は、YAML ファイルを解析し、構文エラーと一般的なセマンティック エラーを検出し、エラー コンテキストを正規化し、日常的な欠陥に対して確定的な自動修正を適用し、LLM を最小限で構造的に有効な修復に導く制約付きプロンプトを構築します。このシステムは、UI シェル、機能オーケストレーター、ドメイン エンジン、統合レイヤーの 4 つのレイヤーを備えたモジュラー Web アプリケーションとして実装されています。その修復パイプラインは、構文/解析、インデント、マッピング、シーケンス、およびスカラー引用エラーの 5 つのカテゴリにわたって手動で挿入されたエラーを含む 100 個の実際の Home Assistant YAML ファイルで評価されました。 gpt-oss-20b、gpt-oss-120b、llama-3.1-8b、および llama-3.3-70b の 4 つのモデルがテストされました。結果は、3 つのモデルが 100% のエラー検出精度を達成し、修復成功率が 87% ~ 93% であることを示しています。手動検証では、成功した出力の中に幻覚や誤った修復は見つかりませんでした。これらの発見は、ドメイン認識プログラム分析と制約付き生成 AI を組み合わせることが、スマート ホーム構成修復の信頼性と使いやすさを向上させる実現可能なアプローチであることを示唆しています。

原文 (English)

SmartHomeSecure: Automated Detection and Repair of Smart Home Configuration Errors Using Large Language Models

Smart home automation platforms increasingly rely on user-authored YAML configuration files to define device behaviors, but these files are prone to syntax, formatting, and semantic logic errors that can cause automation failures and safety risks. Existing YAML validators, static analysis tools, and general-purpose large language models offer limited support for end-to-end diagnosis and repair because they lack domain-specific understanding and validated correction workflows. This paper presents SmartHomeSecure, a prototype for automated detection and repair of Home Assistant configuration errors using lightweight program analysis and constraint-guided large language model generation. SmartHomeSecure parses YAML files, detects syntactic and common semantic errors, normalizes error context, applies deterministic auto-fixes for routine defects, and constructs constrained prompts that guide LLMs toward minimal and structurally valid repairs. The system is implemented as a modular web application with four layers: UI Shell, Feature Orchestrator, Domain Engine, and Integration Layer. Its repair pipeline was evaluated on 100 real-world Home Assistant YAML files with manually injected errors across five categories: syntax/parsing, indentation, mapping, sequence, and scalar quoting errors. Four models were tested: gpt-oss-20b, gpt-oss-120b, llama-3.1-8b, and llama-3.3-70b. Results show that three models achieved 100% error detection accuracy, with repair success rates ranging from 87% to 93%. Manual verification found no hallucinated or incorrect repairs among successful outputs. These findings suggest that combining domain-aware program analysis with constrained generative AI is a feasible approach for improving the reliability and usability of smart home configuration repair.

13:00 JST研究/論文

AirPASS: ピンチ アンテナ システムによる無線フェデレーテッド ラーニング

このペーパーでは、アクセス ポイントにマルチ導波管ピンチング アンテナ システム (PASS) が装備されている無線システムにおける無線フェデレーテッド ラーニング (AirFL) について調査します。当社では、広く研究されている学習指向の AirFL 定式化を採用しており、集約の歪みを所定のしきい値未満に保ちながら、選択したデバイスの数を最大化することを目指しています。これらのシステム変数間の複雑な結合により、デバイスの選択、受信ビームフォーミング、およびピンチング アンテナの配置の共同最適化は、非常に非凸的になります。この課題に対処するために、我々は、2 つの主要コンポーネントを備えた交互最適化フレームワークである AirPASS を開発しました。1 つは、固定 PASS 構成下でのデバイス選択と受信ビームフォーミングのためのホモトピー リーマン マージン統合法、もう 1 つは固定選択されたデバイスとビームフォーマーの下でピンチ アンテナの位置を更新するためのホモトピー支援ジオメトリ最適化法です。実験では、AirPASS が従来の同じ場所に配置された MIMO ベースラインを常に上回っており、理想的な FedAvg に近い状態を維持し、SDR-DC およびマッチング追跡スケジューリングの代替手段と比較して、魅力的なパフォーマンスと複雑さのトレードオフを達成していることが示されています。

原文 (English)

AirPASS: Over-the-Air Federated Learning via Pinching Antenna Systems

This paper investigates over-the-air federated learning (AirFL) in wireless systems where the access point is equipped with a multi-waveguide pinching antenna system (PASS). We adopt the widely studied learning-oriented AirFL formulation, which seeks to maximize the number of selected devices while keeping the aggregation distortion below a prescribed threshold. The resulting joint optimization of device selection, receive beamforming, and pinching-antenna placement is highly nonconvex due to the intricate coupling among these system variables. To address this challenge, we develop AirPASS, an alternating optimization framework with two main components: a homotopy-Riemannian margin-consolidation method for device selection and receive beamforming under fixed PASS configuration, and a homotopy-assisted geometry optimization method for updating the pinching-antenna positions under fixed selected devices and beamformer. Experiments show that AirPASS consistently outperforms conventional co-located MIMO baselines, remains close to ideal FedAvg, and achieves an attractive performance-complexity tradeoff relative to SDR-DC and matching-pursuit scheduling alternatives.

13:00 JSTエージェント

AI ネイティブ 6G 以降のエージェントティックからオートジェニックへのネットワーク管理: 標準の観点

TM フォーラム、3GPP、ETSI などの標準化団体は、次世代ネットワーク管理の基盤として Agentic AI に集中しており、Large AI Model (LAM) ベースのエージェントが自律的に意図を解釈し、リソースを調整し、実行時に運用動作を適応させます。ただし、6G ネットワークの規模と複雑さにおいてこのビジョンを達成するには、運用中に独自の自動化ソフトウェアを生成および進化できる管理システムが必要です。 Autogenic ネットワーク管理は、自己プログラミング、自己反映、自己指向、自己アーキテクチャ機能を備えたエージェント機能を拡張するリファレンス アーキテクチャです。このアーキテクチャは、人間が監視する LAM ベースのエージェントから始まり、信頼が高まるにつれて自律的な運用に進む実際的な段階的な展開をサポートします。 TM フォーラムの自律ネットワークのユースケースから抽出された優先度の高いオペレーターのシナリオを通じてアプローチを実証し、自動生成管理が実際の運用上の課題にどのように対処するかを示します。最後に、将来の 6G ネットワークで自動ネットワーク管理を現実的にするために必要な技術的進歩を概説する研究ロードマップを示します。

原文 (English)

From Agentic to Autogenic Network Management for AI-Native 6G and Beyond: A Standards Perspective

Standards bodies, including TM Forum, 3GPP, and ETSI, are converging on Agentic AI as the foundation for next-generation network management, where Large AI Model (LAM)-based agents autonomously interpret intent, coordinate resources, and adapt operational behaviors at runtime. However, achieving this vision at the scale and complexity of 6G networks requires management systems that can generate and evolve their own automation software during operation. We introduce Autogenic network management, a reference architecture that extends agentic capabilities with self-programming, self reflection, self-orienting, and self-architecting capabilities. The architecture supports practical staged deployment beginning with human-supervised LAM-based agents and progressing toward autonomous operation as confidence builds. We demonstrate the approach through high-priority operator scenarios drawn from TM Forum's autonomous network use cases, showing how autogenic management addresses real operational challenges. We conclude with a research roadmap outlining the technical advances needed to make autogenic network management realistic in future 6G networks.

13:00 JST研究/論文

知識の蒸留による時系列分類のための深層学習モデルの強化

ディープラーニングは、時系列分析、コンピュータービジョン、自然言語処理などのさまざまな分野で目覚ましい成功を収めています。ただし、最先端のアーキテクチャでは高い計算能力とメモリ要求があり、リソースが限られた環境での導入には課題が生じます。知識蒸留 (KD) は、競争力のあるパフォーマンスを維持しながら、大規模な教師モデルから小規模でより効率的な生徒モデルに知識を移転することで、この問題に対処します。この研究では、古典的な完全畳み込みネットワーク (FCN)、畳み込みインセプション モデル、トランスフォーマー ベースの ConvTran モデルという 3 つのアーキテクチャにわたる時系列分類 (TSC) に対する KD の有効性を調査します。 3 つのアーキテクチャにわたって畳み込みフィルター、インセプション モジュール、アテンション ヘッドなどのアーキテクチャ コンポーネントを変更することで、時系列データセットの最大のベンチマーク リポジトリである UCR Archive に対するアプローチを評価します。私たちの結果は一貫して、KD が 3 つのアーキテクチャすべてにわたって中程度の複雑さの学生モデルに最も効果的に利益をもたらすことを示しています。蒸留された FCN 学生はパラメータを 38 分の 1 に削減し、蒸留された Inception 学生は 42% 少ないパラメータで教師とほぼ同じパフォーマンスを達成し、2 つのアテンション ヘッドを持つ蒸留された ConvTran 学生は蒸留によって最も顕著な改善を示しました。さらなる研究と再現性を促進するために、https://github.com/MSD-IRIMAS/KD-4-TSC で実装を提供しています。

原文 (English)

Enhancing deep learning models for time series classification via knowledge distillation

Deep learning has achieved remarkable success in various domains including time series analysis, computer vision and natural language processing. However, high computational and memory demands of state-of-the-art architectures pose challenges for deployment in resource-limited environments. Knowledge Distillation (KD) addresses this by transferring knowledge from a large teacher model to a smaller, more efficient student model while maintaining competitive performance. In this work, we investigate the effectiveness of KD for Time Series Classification (TSC) across three architectures: the classical Fully Convolutional Network (FCN), the convolutional Inception model and the transformer-based ConvTran model. We evaluate our approach on UCR Archive, the largest benchmark repository of time series datasets, by modifying architectural components such as convolutional filters, Inception modules and attention heads across the three architectures. Our results consistently show that KD most effectively benefits student models of intermediate complexity across all three architectures, with the distilled FCN student reducing parameters by a factor of 38, the distilled Inception student achieving nearly the same performance as the teacher with 42% fewer parameters and the distilled ConvTran student with 2 attention heads showing the most significant improvement through distillation. To encourage further research and reproducibility, we provide our implementation at https://github.com/MSD-IRIMAS/KD-4-TSC.

13:00 JST研究/論文ClaudeGPT / ChatGPT

Text-to-SQL の正確性を予測するものは何ですか?選択的予測研究

AI が生成した SQL クエリの不確実性を評価するには、クエリが正しいかどうかを推定する必要があります。正しいとは、人間が作成した参照と同じ結果が実行されることを意味します。私たちは、AUROC を使用して、各シグナルが正しいクエリを間違ったクエリよりもどの程度上位にランク付けするかを測定し、ハード マルチテーブルのテキストから SQL への変換で正確さを予測するシグナルを調査します。 BIRD と Spider では、文字列、構造、実行の自己一貫性、スキーマ関連性スコア、クエリの実行可能性などのブラック ボックス シグナルはすべて約 0.61 ~ 0.68 AUROC の間に収まり、文字列の自己一貫性は 0.675 で最も強くなります。ホワイトボックスの対数確率も同様です (0.67)。この上限を超えて移動するシグナルは検証ベースです。LLM 審査員は 0.72 (GPT-4o-mini) から 0.78 (Claude) までのスコアを付けます。異なるプロバイダーの裁判官は異なるエラーを犯すため、2 つのプロバイダーのアンサンブルは、適切に校正された確率 (予想校正誤差 0.03) で 0.82 AUROC に達し、自己一貫性により有効な低リスクのサブセットが提供されない有用な棄権フロンティア (たとえば、24% の選択的リスクで質問の 27% に回答する) をサポートします。このパターンは、2 つのベンチマーク、2 つのジェネレーター、および 2 つのジャッジプロバイダーにわたって当てはまります。また、検証者をトレーニングできるかどうかも尋ねます。エンコーダーと生成の両方の微調整された検証ツールは、ディストリビューション内では約 0.77 ~ 0.79 の AUROC に達しますが、目に見えないスキーマでは約 0.66 に低下します。 7B へのスケールアップ、スキーマの多様性の追加、強力なジャッジの論理的根拠の抽出、およびクロスベンチマーク トレーニングはすべて、そのギャップを埋めることができません。クロススキーマ転送は、微調整ではなく、モデルのスケールと推論を追跡しているようです。実際には、text-to-SQL の正確性の不確実性は推論ベースのシグナルに存在します。微調整された検証ツールは優れたドメイン内ツールですが、スキーマ全体で一般化する検証ツールは現在、大規模に凍結された推論モデルを意味します。

原文 (English)

What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study

Evaluating uncertainty in AI-generated SQL queries requires estimating whether a query is correct, where correct means it executes to the same result as a human-written reference. We study which signals predict correctness on hard multi-table text-to-SQL, using AUROC to measure how well each ranks correct queries above incorrect ones. On BIRD and Spider, black-box signals such as string, structural, and execution self-consistency, a schema-relevance score, and query executability all fall between about 0.61 and 0.68 AUROC, with string self-consistency strongest at 0.675; white-box log-probability is similar (0.67). The signals that move past this ceiling are verification-based: an LLM judge scores from 0.72 (GPT-4o-mini) to 0.78 (Claude). Judges from different providers make different errors, so a two-provider ensemble reaches 0.82 AUROC with a well-calibrated probability (expected calibration error 0.03) and supports useful abstention frontiers (for example, answering 27% of questions at 24% selective risk) where self-consistency offers no valid low-risk subset. The pattern holds across two benchmarks, two generators, and two judge providers. We also ask whether a verifier can be trained. Fine-tuned verifiers, both encoder and generative, reach about 0.77 to 0.79 AUROC in-distribution but fall to about 0.66 on unseen schemas; scaling to 7B, adding schema diversity, distilling a strong judge's rationales, and cross-benchmark training all fail to close that gap. Cross-schema transfer appears to track model scale and reasoning rather than fine-tuning. In practice, correctness uncertainty for text-to-SQL lives in reasoning-based signals: a fine-tuned verifier is a good in-domain tool, but a verifier that generalizes across schemas currently means a large frozen reasoning model.

13:00 JSTLLM/生成AI

68 の公共生理学的コーパスにわたる監査可能なルール発見のためのマルチアナリスト LLM パイプライン

開いた生理学的コーパスは異質であり、異なるセンサー、ラベル、サンプリング レート、記録設定、臨床エンドポイントを使用します。これらは検出器の設計をサポートできますが、新しい非接触監視プラットフォーム用にどの検出器ルールを構築する必要があるかを直接指定するものではありません。商業利用の互換性についてスクリーニングされた 68 の公共生理学的コーパスを、将来の検証のために候補ルール形状の監査可能なライブラリに変換するための、制御された 4 人のアナリストによる大規模言語モデル (LLM) ワークフローを報告します。 4 つの独立した商用 LLM ファミリが、制御されたプロンプトの下でコーパス ドキュメントを読み取り、695 個の候補ルール マーカー (トップ マーカー) を生成しました。重複排除により 649 個のルール レコードが保持されました。その後、しきい値境界監査により、クランプまたはキュレーターによるレビューのために 51 件の健全性違反のフラグが付けられました。コーパス間の統合により、436 個の固有のルール形状が生成されました。 2 つのハード不変条件、ネイティブのターゲット ハードウェア チャネルの可用性、および患者ごとの複数夜にわたるパーソナライゼーションなしに対するゲート タグ付けにより、4 つの検出器ファミリー バケットにわたる 94 個のビルドナウ検出器コンポーネントが特定されました。パイプラインは検証済みの臨床検出器を生成しません。これにより、アナリストの意見の相違、しきい値チェック、キュレーターによるレビュー、および自動化された継続的インテグレーション (CI) チェックによって、文献から導き出されたルールがハードウェアの将来の検証に向けてルーティングされる、監査可能なエンジニアリング カスケードが生成されます。

原文 (English)

A Multi-Analyst LLM Pipeline for Auditable Rule Discovery Across 68 Public Physiological Corpora

Open physiological corpora are heterogeneous: they use different sensors, labels, sampling rates, recording settings, and clinical endpoints. They can support detector design, but they do not directly specify which detector rules should be built for a new contactless monitoring platform. We report a controlled four-analyst large-language-model (LLM) workflow for converting 68 public physiological corpora, screened for commercial-use compatibility, into an auditable library of candidate rule shapes for prospective validation. Four independent commercial LLM families read the corpus documentation under a controlled prompt and produced 695 candidate rule markers (top-markers). Deduplication retained 649 rule records; a threshold-bounds audit then flagged 51 sanity violations for clamping or curator review. Cross-corpus consolidation produced 436 unique rule shapes. Gate-tagging against two hard invariants, native target-hardware channel availability and no multi-night per-patient personalization, identified 94 build-now detector components across four detector-family buckets. The pipeline does not produce a validated clinical detector. It produces an auditable engineering cascade in which analyst disagreement, threshold checks, curator review, and automated continuous-integration (CI) checks route literature-derived rules toward prospective hardware validation.

13:00 JSTLLM/生成AIエージェント

エージェントが不正になったとき: マルチエージェント システムにおける悪意のある動作のアクティベーション ベースの検出

LLM ベースのマルチエージェント システム (MAS) は、複雑なタスクでの効果的なコラボレーションを可能にしますが、エージェント レベルと対話レベルでの脆弱性による重大なセキュリティ課題に直面しています。既存の MAS セキュリティ防御のほとんどは、意味的に明示的な悪意のある攻撃と、MAS トポロジとエージェント レベルの対話の明示的なグラフベースのモデリングという 2 つの中心的な前提に基づいて構築されています。実際には、現実世界の攻撃はセマンティクス的によりステルス化していますが、MAS の実行は通常、グラフベースの伝播モデルで想定されている時間的調整なしで非同期で行われます。これらの制限に対処するために、MAS での悪意のある動作を検出するためのアクティベーション ベースのフレームワークである AcMAS を提案します。 AcMAS は、ローカル エージェントのアクティブ化空間で内部推論状態を分析することにより、明示的なインタラクション グラフに依存せずに、同期が堅牢な方法でステルス攻撃さえも検出します。さらに、当社の活性化分析は、最先端の方法で一般的に使用されている破壊的な薬剤の分離ではなく、侵害された薬剤の機能を回復する際に AcMAS を導く重要なシグナルを提供します。包括的な評価では、AcMAS がステルス攻撃に対してグラフベースのベースラインを大幅に上回っており、同期設定では +0.22 F1 (0.94 対 0.72)、非同期設定では +0.55 F1 (0.93 対 0.38)、さまざまなオープンソース LLM バックボーン、攻撃強度、MAS スケールにわたって一般化されていることが実証されています。

原文 (English)

When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems

While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities at the agent and interaction levels. Most existing MAS security defenses are built upon two core assumptions: semantically-explicit malicious attacks and explicit graph-based modeling of the MAS topology and agent-level interactions. In practice, real-world attacks are becoming more semantically stealthy, while MAS execution is typically asynchronous without the temporal alignment assumed by graph-based propagation models. To address these limitations, we propose AcMAS, an activation-based framework for malicious-behavior detection in MAS. By analyzing internal reasoning states in the activation space of local agents, AcMAS detects even stealthy attacks in a synchronization-robust fashion, without relying on explicit interaction graphs. Moreover, our activation analysis provides critical signals to guide AcMAS in restoring the functionality of compromised agents, rather than the disruptive agent isolation commonly used by the state-of-the-art methods. Comprehensive evaluation demonstrates that AcMAS significantly outperforms graph-based baselines against stealthy attacks, by +0.22 F1 in synchronous settings (0.94 vs. 0.72) and by +0.55 F1 in asynchronous settings (0.93 vs. 0.38), with generalization across diverse open-source LLM backbones, attack intensity, and MAS scale.

13:00 JSTLLM/生成AI

自己クリティカルなマスク言語モデルを使用した広告見出しの生成

どのような電子商取引 Web サイトにとっても、買い物客を惹きつける永続的な広告を構築することは重要な問題です。特に大規模な場合、Web サイトのクリエイティブ品質の基準を超えることは困難です。したがって、小売コンテンツを使用して製品広告のヘッドラインを生成するためのプログラムによるソリューションを提案します。私たちは、Transformer ベースのマスク言語モデルに対する強化学習 (RL) ポリシー勾配法の最先端のアプリケーションを提案します。私たちの方法は、販売者が宣伝したい複数の製品を共同条件付けすることによって広告見出しを作成します。私たちは、重複メトリクスと品質監査において、私たちの手法が既存の Transformer および LSTM + RL 手法よりも優れていることを実証します。また、監査によって判定された文法と創造的な品質の両方の点で、モデルが生成した見出しが人間が提出した見出しよりも優れていることも示しています。

原文 (English)

Ad Headline Generation using Self-Critical Masked Language Model

For any E-commerce website it is a nontrivial problem to build enduring advertisements that attract shoppers. It is hard to pass the creative quality bar of the website, especially at a large scale. We thus propose a programmatic solution to generate product advertising headlines using retail content. We propose a state of the art application of Reinforcement Learning (RL) Policy gradient methods on Transformer based Masked Language Models. Our method creates the advertising headline by jointly conditioning on multiple products that a seller wishes to advertise. We demonstrate that our method outperforms existing Transformer and LSTM + RL methods in overlap metrics and quality audits. We also show that our model-generated headlines outperform human submitted headlines in terms of both grammar and creative quality as determined by audits.

13:00 JSTLLM/生成AI画像/動画生成

あらゆる ASR モデル向けの勾配ベースの音声からテキストへの位置合わせ: CTC から音声 LLM まで

音声とテキストの位置合わせとは、音声内の各単語の時間的境界を見つけることを意味します。一部のモデルはそのような位置合わせを直接提供しますが、他のモデルは提供しません。コネクショニスト時間分類 (CTC) とトランスデューサー モデルには構造上の整合性がありますが、注意ベースのエンコーダー/デコーダー (AED) と音声大規模言語モデル (LLM) には整合性がなく、単語のタイミングは通常、代わりに注意の重みから読み取られます。これらの信号はすべて、時間精度を制限するエンコーダ フレーム グリッド上に存在します。私たちは、微分可能な ASR モデルに適用できる一般的な勾配ベースのアライメントを研究します。入力に対する各教師強制トークンの対数確率の勾配を取得し、それをフレームごとの顕著性に換算し、単一の動的プログラミング パスで結果の行列を単語境界にデコードします。この方法はトレーニング、モデルの変更、アライメント ヘッドを必要とせず、音声 LLM を含むすべてのモデル ファミリにわたって機能し、粗いエンコーダ グリッドではなく入力グリッド上でアライメントします。 4 つのファミリーからの 16 のモデルで、読み上げ (TIMIT) および自発的 (Buckeye) 音声について、それぞれモデル独自のネイティブまたは注意ベースのアライメントと比較して評価しました。勾配によってすべてのモデルで使用可能なアライメントが得られること、ストリーミング モデルの場合のように、通常は強力なネイティブ アライナーよりも若干遅れますが、ネイティブ アライメントが弱い場合にはより優れていること、およびその主な欠点はトークンごとに 1 回の逆方向パスのコストがかかることであることがわかりました。

原文 (English)

Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs

Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly and others do not. Connectionist temporal classification (CTC) and transducer models have an alignment by construction, whereas attention-based encoder-decoders (AED) and speech large language models (LLMs) do not, and their word timings are usually read off the attention weights instead. All of these signals live on the encoder frame grid, which bounds their temporal precision. We study a generic gradient-based alignment that applies to any differentiable ASR model. We take the gradient of each teacher-forced token log probability with respect to the input, reduce it to a per-frame saliency, and decode the resulting matrix into word boundaries with a single dynamic-programming pass. The method needs no training, no model modification and no alignment heads, works across all model families including the speech LLMs, and aligns on the input grid rather than on the coarser encoder grid. We evaluate it on sixteen models from four families, on read (TIMIT) and spontaneous (Buckeye) speech, each against the model's own native or attention-based alignment. We find that the gradient yields a usable alignment for every model, that it is usually somewhat behind a strong native aligner but better where the native alignment is weak, as for the streaming models, and that its main disadvantage is the cost of one backward pass per token.

13:00 JSTエージェント

何が軽量のゲームプレイエージェントを強力にするのかに関するゴールドスタンダードの研究

不完全情報カード ゲームの強化学習エージェントは、トレーニングする対戦相手と同じくらい強く、99 パーセント以上の確率でランダムな対戦相手に勝ち、自分のコピーしか同点にならないため、評価するのは困難です。そのため、私たちはジン ラミー用の強力で固定されたルールベースのエキスパートを構築し、それを基準としてのみ使用し、トレーニングには決して使用しません。これは、私たちがトレーニングしたすべてのエージェントを 70 ~ 99 パーセントの確率で上回ります。 100 回以上の実行を通じて、何が軽量エージェントを強力にするのかを特定しました。トラストリージョンのアップデート、狙いを定めた報酬、より手ごわい相手のカリキュラム、ウォームスタート、最良のチェックポイントの維持がすべて役に立ち、それらを積み重ねることで、セルフプレイチャンピオンのエキスパートに対する勝率は約30パーセントから36パーセントに上昇します。いくつかのアイデアは成果を上げませんでした。短期および長期の報酬形成、学習された状態の埋め込み、模倣と DAgger、およびライブの大規模言語モデルの対戦相手は、それぞれ役に立たず、大規模にトレーニングするには遅すぎるか、重すぎました。 MLP、畳み込み、セットベース、アテンション、リカレント エンコーダを比較すると、追加容量は上限を突破するのにほとんど役立たないことがわかり、制限はネットワーク サイズではなく情報であることを示唆しています。標準ベースライン (神経系の架空のセルフプレイと情報セットのモンテカルロ検索) を追加し、そのアプローチが最適値が計算可能な Leduc Hold'em に引き継がれることを確認します。その結果、小規模なモデルで処理できるあらゆるゲームに対応し、エキスパートのトレーニングなしで競争力のあるエージェントをトレーニングする、軽量でゲームに依存しないレシピが作成され、堅牢な統計情報とともにレポートされ、再利用可能なパッケージとしてリリースされます。

原文 (English)

A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

Reinforcement learning agents for imperfect-information card games are only as strong as the opponents they train against, and they are hard to grade, since they beat a random opponent over 99 percent of the time and only tie copies of themselves. So we build a strong, fixed, rule-based expert for Gin Rummy and use it only as a yardstick, never for training. It beats every agent we trained 70 to 99 percent of the time. Across more than a hundred runs, we isolate what makes a lightweight agent stronger. Trust region updates, a well-aimed reward, a curriculum of tougher opponents, warm starting, and keeping the best checkpoint all help, and stacking them lifts a self-play champion from about 30 to 36 percent against the expert. Several ideas did not pay off. Short-term and longer-term reward shaping, learned state embeddings, imitation and DAgger, and a live large language model opponent were each unhelpful, too slow, or too heavy to train at scale. Comparing MLP, convolutional, set-based, attention, and recurrent encoders shows that extra capacity does little to break the ceiling, suggesting the limit is information rather than network size. We add standard baselines (neural fictitious self-play and information set Monte Carlo search) and confirm the approach carries over to Leduc Hold'em, where the optimum is computable. The result is a lightweight, game-agnostic recipe that trains competitive agents without training on the expert, for any game a small model can handle, reported with robust statistics and released as a reusable package.

13:00 JSTロボティクス

GemNav: マルチモーダル大規模言語モデルを使用した離散トークンのビジュアル ロボット ナビゲーション

大規模な事前トレーニング済みモデルに基づいて構築されたビジュアル ナビゲーション ポリシーは、これまでのところ、専用のビジュアル エンコーダー、特注のアクション ヘッド、および数千時間に及ぶクロス実施形態データセットでのトレーニングという共通のレシピに従っています。このレシピが必要かどうかを尋ねます。この論文では、言語タワーのみで低ランク適応 (LoRA) を使用し、補助ビジュアル エンコーダや連続回帰ヘッドを使用せず、フリーズしたマルチモーダル大規模言語モデル (MLLM) を短中地平線のウェイポイント ナビゲーションに適応させるビジュアル ロボット ナビゲーション ポリシーである GemNav を紹介します。ウェイポイントとカテゴリカル ナビゲーション信号は、言語モデル ヘッドによって生成された単一の離散トークン ボキャブラリーを共有し、ソフト デコードされた補助損失により、純粋なクロス エントロピー トレーニングで破棄される計量構造が回復されます。このポリシーは、競合するトレーニング セットよりもおよそ 3 桁小さい、単一の 8.7 時間のオープン コーパス上で、ゼロショットを 4 つの物理的に異なる目に見えない環境に転送し、オープン駐車場、障害物駐車場、長い屋外の化学薬品置き場、屋内倉庫をカバーする 20 の実世界のトライアルにわたって、ゴールの 0.25 ~ 0.42 m 以内で停止します。短い画像履歴に基づいて条件付けすると、オフライン メトリクスは改善されますが、ロボットには何のメリットも得られず、事前にトレーニングされた視覚機能が導入された後に追加される時間的コンテキストの上限が指摘されています。これらの結果は、凍結された MLLM の離散トークン適応により、基礎モデルのロボット ナビゲーションにデータ効率が高く、展開可能な代替手段を提供できることを示しています。

原文 (English)

GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model

Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets. We ask whether this recipe is necessary. In this paper, we introduce GemNav, a visual robot navigation policy that adapts a frozen Multimodal Large Language Model (MLLM) for short-to-medium horizon waypoint navigation using Low-Rank Adaptation (LoRA) on the language tower alone, with no auxiliary visual encoder and no continuous regression head. Waypoints and categorical navigation signals share a single discrete token vocabulary generated by the language-model head, and a soft-decoded auxiliary loss recovers the metric structure that pure cross-entropy training discards. On a single 8.7-hour open corpus, roughly three orders of magnitude smaller than competing training sets, the policy transfers zero-shot to four physically distinct unseen environments and stops within 0.25-0.42m of the goal across 20 real-world trials covering an open carpark, an obstacle carpark, a long outdoor chemical yard, and an indoor warehouse. Conditioning on short image histories improves offline metrics but yields no robot benefit, pointing to a ceiling on what temporal context adds once pretrained vision features are in place. These results indicate that discrete-token adaptation of frozen MLLMs can provide a data-efficient, deployable alternative for foundation model robot navigation.

13:00 JST画像/動画生成

ReMoDEx: 大規模画像データセット向けのローカルからグローバルへの関連性ベースのモデル意思決定説明フレームワーク

Deep learning image classifiers achieve strong predictive performance yet remain opaque in how decisions are formed.モデルは、タスクに関連する領域ではなく、無関係な手がかり、ショートカットの関連付け、周辺構造、またはデバイス レベルのアーティファクトに依存しながら、正しく予測する可能性があります。大規模なデータセットでは、一度に 1 つのサンプルをヒートマップで検査すると数千の予測に対応できないため、この不透明度は特に問題になります。我々は、画像分類におけるモデル決定動作を体系的にデータセット規模で評価するためのフレームワークである、Relevance Based Model Decision Explainability (ReMoDEx) を提案します。 ReMoDEx は、モデル推論、ターゲット クラスの選択、関連性マップの生成、ヒートマップの標準化、類似性に基づくパターンのグループ化、クラスター レベルの解釈、空間関連性の評価といった段階的なパイプラインを定義します。ローカルメソッド GradCAM++、Integrated Gradients、Occlusion Sensitivity、および Layerwise Relevance Propagation はそれぞれ、関連性マップのセット全体をいくつかの意思決定戦略クラスターに要約する単一のグローバル モジュールと個別に結合され、サンプルごとの検査を自動でスケーラブルな要約に置き換えます。 ReMoDEx を実証するために、新型コロナウイルス感染症 (COVID-19)、正常、肺混濁、ウイルス性肺炎を区別する VGG16 ベースの分類子に ReMoDEx を適用しました。分類器は安定したパフォーマンスを示しました (テスト精度 86.27%、テスト AUC 0.9624)。しかし、グローバルモジュールと組み合わせた各説明子は、胸部中心領域の決定と境界/コーナーに敏感な決定という 2 つの反復戦略を一貫して生成し、従来の測定基準では明らかにできなかった近道学習の可能性を示しました。マスクされた画像の検証により、中心領域または周辺領域が遮られるとモデルの信頼性と予測クラスが変化することが確認されました。したがって、ReMoDEx は、スケーラブルな関連性ベースの意思決定評価フレームワークと、精度ベースの評価を補完する重要な機能を提供します。

原文 (English)

ReMoDEx: A Local-to-Global Relevance-Based Model Decision Explainability Framework for large-Scale Image Datasets

Deep learning image classifiers achieve strong predictive performance yet remain opaque in how decisions are formed. A model may predict correctly while relying on irrelevant cues, shortcut associations, peripheral structures, or device level artifacts instead of task relevant regions. On large scale datasets this opacity is especially problematic, since inspecting heatmaps one sample at a time cannot scale to thousands of predictions. We propose Relevance Based Model Decision Explainability (ReMoDEx), a framework for systematic, dataset scale assessment of model decision behaviour in image classification. ReMoDEx defines a stepwise pipeline: model inference, target class selection, relevance map generation, heatmap standardisation, similarity based grouping of patterns, cluster level interpretation, and spatial relevance assessment. Local methods GradCAM++, Integrated Gradients, Occlusion Sensitivity, and Layerwise Relevance Propagation are each combined independently with a single global module that summarises an entire set of relevance maps into a few decision strategy clusters, replacing sample by sample inspection with an automatic, scalable summary. To demonstrate ReMoDEx, we applied it to a VGG16 based classifier distinguishing COVID-19, Normal, Lung Opacity, and Viral Pneumonia. The classifier showed stable performance (86.27% test accuracy, 0.9624 test AUC). However, each explainer combined with the global module consistently produced two recurring strategies: central thoracic region decisions and border/corner sensitive decisions, indicating possible shortcut learning that conventional metrics could not reveal. Masked image validation confirmed that model confidence and predicted class changed when central or peripheral regions were occluded. ReMoDEx thus provides a scalable relevance based decision assessment framework and an essential complement to accuracy based evaluation.

13:00 JST研究/論文

AI 拡張コンピューティングにおける確率的オラクルを使用したコンピューティング

Stochastic-Oracle Turing Machine (SOTM) フレームワークは、AI 拡張計算を、コンテキスト依存の分布から応答が引き出されるオラクルと確率的チューリング マシンの相互作用としてモデル化します。このペーパーでは、SOTM が 2 つのオラクル応答スキームの下で何を達成できるかを検討します。キャッシュされた応答オラクルでは、それぞれの個別のクエリが 1 つの応答を受け取り、同じクエリに対する後の呼び出しで再利用されます。一方、フレッシュ応答オラクルでは、各呼び出しは独立した応答を返します。どちらのスキームでも、SOTM はまず入力と内部ランダム ソースから計算して最初のクエリを生成し、次に適応的に処理を進め、クエリ応答トランスクリプト (クエリの発行と受信した応答の記録) から計算して、後続の各クエリを生成するか、最終出力を生成します。キャッシュされた応答は、達成可能なパフォーマンスにトランスクリプトベースの 2 つの上限を課します。1 つはオラクルの隠れ状態によって引き起こされるトランスクリプト分布間の合計変動距離によって支配される正しい識別の上限、もう 1 つは SOTM がトランスクリプトから計算できる最良の出力の予想スコアに等しい出力品質の上限です。新鮮な応答は、正しいまたは高品質の出力に向けた独立した証拠を蓄積するための繰り返しの呼び出しを許可することで、これらの上限を引き上げることができます。バイナリの単一情報クエリの場合、同じクエリへの呼び出し数が増えるにつれて、エラーの確率がチェルノフ レートで指数関数的に減少します。出力品質については、スコア関数が SOTM の一部として組み込まれている場合はクエリ数の境界がしきい値停止を特徴付け、そうでない場合は多数決ベースの増幅境界がバイナリ候補出力モデルを特徴付けます。これらの結果を総合すると、応答の再利用、トランスクリプト情報、スコア関数へのアクセスによって、SOTM が計算できる内容とトークン コストがどのように決まるのかがわかります。

原文 (English)

Computing with Stochastic Oracles in AI-Augmented Computation

The Stochastic-Oracle Turing Machine (SOTM) framework models AI-augmented computation as the interaction of a probabilistic Turing machine with an oracle whose responses are drawn from context-dependent distributions. This paper studies what an SOTM can achieve under two oracle-response schemes: in a cached-response oracle, each distinct query receives one response that is reused on later calls to the same query, while in a fresh-response oracle, each call returns an independent response. In both schemes, the SOTM first computes from its input and internal random source to generate its first query, then proceeds adaptively, computing from its query-response transcript (the record of queries issued and responses received) to generate each subsequent query or produce a final output. Cached responses impose two transcript-based ceilings on achievable performance: a correct-identification ceiling governed by the total variation distance between the transcript distributions induced by the hidden states of the oracle, and an output quality ceiling equal to the expected score of the best output the SOTM can compute from the transcript. Fresh responses can raise these ceilings by allowing repeated calls to accumulate independent evidence toward correct or high-quality outputs. In the binary single-informative-query case, the error probability decreases exponentially in the number of calls to the same query at the Chernoff rate. For output quality, query-count bounds characterize threshold stopping when the score function is incorporated as part of the SOTM, and majority-based amplification bounds characterize the binary candidate-output model when it is not. Together, the results identify how response reuse, transcript information, and access to the score function determine what an SOTM can compute and at what token cost.

13:00 JST画像/動画生成

LoCA: ビジョン基盤モデルの空間認識低ランク畳み込み適応

事前トレーニングされた Vision Foundation Models (VFM) は、さまざまな下流タスクに強力な視覚的表現を提供します。 VFM 適応の主な課題は、完全な微調整と壊滅的な忘却に伴う法外なコストに起因します。これに対処するために、パラメータ効率の良い微調整 (PEFT) の一般的なパラダイムとして、低ランク適応 (LoRA) が登場しました。ただし、LoRA は通常、2D 行列によってパラメータ化されたトランスフォーマー セルフ アテンション レイヤー用に設計されています。畳み込みカーネルは本質的に 4D テンソル内の空間情報とチャネル情報を結合するため、それらをモノリシック 2D マトリックスに強制すると、固有の空間トポロジーが破壊されます。この論文では、チャネルと空間適応を切り離すことによって空間チャネルもつれに対処する畳み込み認識 PEFT フレームワークである低ランク畳み込み適応 (LoCA) を提案します。 LoCA は、高密度クロスチャネル混合のための低ランク チャネル適応を導入し、特異値分解 (SVD) によって事前トレーニングされたカーネルから抽出された空間基底を洗練します。実験結果は、LoCA が事前にトレーニングされた空間事前分布を保存し、きめの細かい分類、ドメイン一般化されたセマンティック セグメンテーション、および生成ベンチマーク全体にわたって競争力のある、または最先端のパフォーマンスを達成することを示しています。

原文 (English)

LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models

Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM adaptation stems from the prohibitive costs of full fine-tuning and catastrophic forgetting. To address this, Low-Rank Adaptation (LoRA) has emerged as the prevailing paradigm for Parameter-Efficient Fine-Tuning (PEFT). However, LoRA is typically designed for transformer self-attention layers parameterized by 2D matrices. Since convolutional kernels inherently couple spatial and channel information within a 4D tensor, forcing them into a monolithic 2D matrix disrupts the inherent spatial topology. In this paper, we propose Low-Rank Convolutional Adaptation (LoCA), a convolution-aware PEFT framework that addresses spatial-channel entanglement by decoupling channel and spatial adaptation. LoCA introduces a low-rank channel adaptation for dense cross-channel mixing and refines spatial bases extracted from pre-trained kernels via Singular Value Decomposition (SVD). Experimental results show that LoCA preserves pre-trained spatial priors and achieves competitive or state-of-the-art performance across fine-grained classification, domain-generalized semantic segmentation, and generative benchmarks.

13:00 JST研究/論文

MADB: 専門的かつ多次元の注釈を備えた大規模な音楽美学データセット

音楽の美的評価は、挑戦的であるにもかかわらず十分に解明されていない問題であり、モデルにはきめの細かい多次元の人間の知覚判断を捉える必要があります。この分野の進歩は、構造化された美的アノテーションを備えた大規模なデータセットの欠如によって制限されてきました。 30 人の訓練を受けたアノテーターによってアノテーションが付けられた 9,999 トラックで構成される大規模なデータセットおよびベンチマークである MADB を紹介します。各トラックは、10 の知覚次元と 1 つの総合スコアにわたって約 10 人のアノテーターによって評価され、マルチモーダル分析のための追加のテキスト コメントが付けられます。複数の事前トレーニング済みモデルにわたる統一された評価フレームワークを確立します。結果は、モデルの予測と人間の判断との間に大きなギャップがあることを明らかにし、現在のアプローチの重要な限界を明らかにします。 MADB は、人間に合わせた音楽理解のための新しいベンチマークを提供します。プロジェクトページ: https://github.com/knownree/madb

原文 (English)

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual judgments. Progress in this area has been limited by the lack of large-scale datasets with structured aesthetic annotations. We introduce MADB, a large-scale dataset and benchmark comprising 9,999 tracks annotated by 30 trained annotators. Each track is rated by around 10 annotators across 10 perceptual dimensions and one overall score, with additional textual comments for multimodal analysis. We establish a unified evaluation framework over multiple pretrained models. Results reveal substantial gaps between model predictions and human judgments, exposing key limitations of current approaches. MADB provides a new benchmark for human-aligned music understanding. Project page: https://github.com/knownree/madb

13:00 JST研究/論文

代入とクラスタリングの融合: 潜在的なサブグループ構造を利用して欠落データを回復する

実際のアプリケーションでは欠損データが蔓延しているため、効果的な代入が下流の分析にとって不可欠な前処理ステップとなっています。現実世界のデータセットは、多くの場合、異なる分布を持つ複数のサブグループで構成される複雑な潜在構造を示します。しかし、既存の方法では、このような集団の不均一性が見落とされることがよくあります。明示的な構造ガイダンスがなければ、これらの方法はサブグループの境界があいまいになり、インスタンスレベルの忠実度に欠ける一般的な推定値を生成する傾向があります。サブグループ情報を組み込むことで解決策は得られますが、循環依存の問題に直面しています。信頼できるサブグループの識別には完全なデータが必要ですが、データの補完自体が代入の目的です。これを解決するために、私たちは CAGI (Cluster-Aware Generative Imputation) を提案します。これは、クラスタリングと代入を相互に強化する協調最適化プロセスとして再定式化するフレームワークです。 CAGI は「パーティション ガイド リストア」戦略を採用しており、動的クラスター割り当てがローカル事前分布として機能し、敵対的生成ネットワークを条件付けします。反復フィードバック ループが確立され、クラスター構造と代入値の両方が徐々に改良され、忠実なサブグループ分布が得られます。分散の安定性を確保するために、CAGI はさらに、インスタンス レベルの再構築と分散レベルの正則化を組み合わせたマルチレベルの最適化目標を採用しています。 15 の代表的なベースラインを含む 14 のベンチマーク データセットに対する広範な実験により、CAGI の優位性が実証されました。ソース コードは https://github.com/supercocachii/CAGI で入手できます。

原文 (English)

Imputation Meets Clustering: Exploiting Latent Subgroup Structure for Missing Data Recovery

Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis. Real-world datasets often exhibit complex latent structures composed of multiple subgroups with distinct distributions. However, existing methods often overlook such population heterogeneity. Without explicit structural guidance, these methods tend to produce generic estimates that blur subgroup boundaries and lack instance-level fidelity. While incorporating subgroup information offers a remedy, it faces a circular dependency: reliable subgroup identification requires complete data, while data completion is the imputation objective itself. To resolve this, we propose CAGI (Cluster-Aware Generative Imputation), a framework that reformulates clustering and imputation as a mutually reinforcing co-optimization process. CAGI employs a ``Partition-Guide-Restore'' strategy where dynamic cluster assignments act as local priors to condition a Generative Adversarial Network. An iterative feedback loop is established to progressively refine both cluster structures and imputed values toward faithful subgroup distributions. To ensure distributional stability, CAGI further employs a multi-level optimization objective combining instance-level reconstruction with distribution-level regularization. Extensive experiments on 14 benchmark datasets with 15 representative baselines demonstrate the superiority of CAGI. The source code is available at: https://github.com/supercocachii/CAGI

13:00 JSTLLM/生成AIビジネス/資金調達

大規模言語モデルの応答の包括的な評価: 多要素スコアリング システム

言語タスクにおける大規模言語モデル (LLM) の顕著なパフォーマンスは、応答品質の包括的な評価が緊急に必要であることを強調しています。一般的な手法は、特異な次元に限定されることが多く、モデルの機能の全領域を捉えるには至っていません。この研究では、精度、簡潔さ、事実の一貫性、読みやすさ、一貫性を統合した多要素スコアリング パラダイムを導入し、結果を視覚化するためのグラフィカル ユーザー インターフェイス (GUI) によって補完されています。 TruthfulQA データセットの評価では、複雑な事実や曖昧さを回避する際の広範な制限とともに、推論タスクにおける主流の LLM の強み (複合スコア 0.6104 でピーク) が明らかになりました。このフレームワークは、従来のメトリクスの狭いレンズを超えて、モデルの可能性と欠陥を明らかにするための透明性と適応性のある手段を提供します。現在は英語のタスクに焦点を当てていますが、その視野は多言語の領域に向かっています。この研究は、知識エンジニアリングとモデルの改良に新たな道を切り開きます。

原文 (English)

Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System

The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of their response quality. Prevailing methods, often confined to singular dimensions, fall short of capturing the full spectrum of model capabilities. This study introduces a multifactor scoring paradigm, integrating accuracy, conciseness, factual consistency, readability, and coherence, complemented by a graphical user interface (GUI) for visualizing outcomes. Evaluations on the TruthfulQA dataset unveil mainstream LLMs' strengths in reasoning tasks (peaking at a composite score of 0.6104) alongside pervasive limitations in navigating complex facts and ambiguities. Transcending the narrow lens of traditional metrics, this framework offers a transparent, adaptable avenue to illuminate model potential and deficiencies. Though presently focused on English tasks, its horizons beckon toward multilingual domains. This work carves a novel path for knowledge engineering and model refinement.

13:00 JST画像/動画生成

自己教師あり事前トレーニングにより、点群の葉と木のセグメンテーションのサイト間およびスケール間の堅牢性が向上

木の点群に対する既存の葉と木のセグメンテーション手法の精度は、森林の種類や場所によって異なります。点群の自己教師あり学習 (SSL) により、バイオマス回帰や個々の樹木のセグメンテーションなど、林業点群タスクの深層学習モデルの一般化が向上しましたが、葉材と木材のセグメンテーションへの適用性はまだテストされていません。この研究では、点群に広く使用されている SSL アーキテクチャである Point-M2AE を、2,400 個の個別のツリー点群で強化された ShapeNet-55 上で事前トレーニングしました。微調整と推論のために、再帰的ボクセル細分割を使用して入力全体の点密度の幅広い変動を処理し、アーキテクチャを変更せずに同じモデルを個別ツリーとプロットの両方のスケールで動作できるようにしました。事前トレーニングなしのモデルと比較して、事前トレーニング済みモデルでは木材 IoU が針葉樹で 60.5% から 70.0%、広葉樹で 69.7% から 76.3% に向上しました。 3 つの気候帯にわたる 4 か国のベンチマークでは、事前トレーニング済みモデルは、比較した手法 (LeWos、CWLS、PointTransformer) の中でサイト間の変動が最小で、全体的なパフォーマンスが最高を達成しました。プロットレベルのセグメンテーションは、広葉プロットで 84.7%、針葉プロットで 77.7% の mIoU を達成し、個々の木のパフォーマンスに匹敵する精度を維持しました。これは、モデルが追加の微調整なしでスケール全体で一般化していることを示しています。密な樹冠によりセグメンテーションが困難な熱帯林での下流テストとして、私たちはモデルと定量的構造モデルを適用して、ガイアナ、インドネシア、ペルーの 28 本の木の体積を推定し、SSL 事前トレーニングによるセグメンテーションの改善が下流のパフォーマンスの向上につながるかどうかを評価しました。結果として得られた体積推定値は、テストされたすべての方法の中で最も低い誤差 (MAE = 2.40 m$^3$) を達成し、アルゴリズムのベースライン (LeWos: 5.94 m$^3$; CWLS: 5.27 m$^3$) の半分未満でした。

原文 (English)

Self-Supervised Pretraining Improves Cross-Site and Cross-Scale Robustness of Point Cloud Leaf-Wood Segmentation

The accuracy of existing leaf-wood segmentation methods for tree point clouds varies across forest types and sites. Self-supervised learning (SSL) on point clouds has improved the generalization of deep learning models for forestry point cloud tasks, including biomass regression and individual tree segmentation, but its applicability to leaf-wood segmentation remains untested. In this study, we pretrained Point-M2AE, a widely used SSL architecture for point clouds, on ShapeNet-55 augmented with 2,400 individual tree point clouds. For fine-tuning and inference, we used recursive voxel subdivision to handle the wide variation in point density across inputs, allowing the same model to operate at both individual-tree and plot scales without architecture change. Compared to the model without pretraining, the pretrained model improved wood IoU from 60.5% to 70.0% for needleleaf and from 69.7% to 76.3% for broadleaf trees. On a benchmark spanning four countries across three climatic zones, the pretrained model achieved the smallest cross-site variation and highest overall performance among compared methods (LeWos, CWLS, and PointTransformer). Plot-level segmentation maintained accuracy comparable to individual-tree performance, with mIoU of 84.7% for broadleaf and 77.7% for needleleaf plots, showing that the model generalizes across scales without additional finetuning. As a downstream test in tropical forests, where dense canopies make segmentation challenging, we applied our model and a quantitative structure model to estimate wood volume for 28 trees from Guyana, Indonesia, and Peru to assess whether the segmentation improvements from SSL pretraining translate into improved downstream performance. The resulting volume estimates achieved the lowest error among all methods tested (MAE = 2.40 m$^3$), less than half that of algorithmic baselines (LeWos: 5.94 m$^3$; CWLS: 5.27 m$^3$).

13:00 JSTLLM/生成AI画像/動画生成AnthropicClaudeOpenAIGPT / ChatGPTGoogleGeminiLlamaMicrosoftCopilot

サイバーセキュリティとプライバシーにおける大規模言語モデル (LLM) と生成 AI: デュアルユース リスク、AI 生成マルウェア、説明可能性、および防御戦略に関する調査

ChatGPT、Claude、Gemini、LLaMA、Copilot、OpenAI による安定拡散、Anthropic、Google、Meta、Microsoft、Stability AI などの大規模言語モデル (LLM) と生成 AI (GenAI) システムは、それぞれサイバーセキュリティに革命をもたらし、自動化された防御と高度な攻撃の両方を可能にします。これらのテクノロジーは、リアルタイムの脅威検出、フィッシング防御、安全なコード生成、および前例のない規模での脆弱性の悪用を強化します。 LLM で生成されたマルウェアが急増し、2021 年のわずか 2% から 2025 年までに検出された脅威の推定 50% を占めるまでに成長し、2026 年にこの高度に自動化された脅威環境を乗り切るには、次世代のセキュリティ フレームワークが必要となります。このペーパーでは、ゼロデイ検出、DevSecOps、フェデレーテッド ラーニング、合成コンテンツ分析、説明可能な AI (XAI) など、サイバーセキュリティにおける LLM の有益なアプリケーションと悪意のあるアプリケーションに関する包括的な調査を紹介します。この研究では、70 を超える学術論文、業界レポート、技術文書のレビューに基づいて、Google Play プロテクト、Microsoft Defender、アマゾン ウェブ サービス (AWS)、Apple App Store、OpenAI プラグイン ストア、Hugging Face Spaces、GitHub などのプラットフォームにわたる実世界のケーススタディからの洞察と、SAFE フレームワークや AI 主導の異常検出などの新たな取り組みを統合しています。最後に、モデル透かし、敵対的防御、業界を超えたコラボレーションなど、責任ある透明性のある LLM 導入と信頼できる AI に関する実践的な推奨事項をまとめ、AI と脅威防御の交差点における厳格で総合的なサイバーセキュリティ研究の新しいベンチマークを設定し、AI 主導のサイバーセキュリティの複雑な課題に対処する研究者、エンジニア、セキュリティ リーダーにとって重要な参考となる安全でスケーラブルな LLM システムのロードマップを提供します。

原文 (English)

Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies

Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI, Anthropic, Google, Meta, Microsoft, Stability AI, respectively, are revolutionizing cybersecurity, enabling both automated defense and sophisticated attacks. These technologies power real-time threat detection, phishing defense, secure code generation, and vulnerability exploitation at unprecedented scales. Following a rapid surge where LLM-generated malware grew to account for an estimated 50% of detected threats by 2025, up from just 2% in 2021, navigating this highly automated threat landscape in 2026 demands next-generation security frameworks. This paper presents a comprehensive survey of the beneficial and malicious applications of LLMs in cybersecurity, including zero-day detection, DevSecOps, federated learning, synthetic content analysis, and explainable AI (XAI). Drawing on a review of over 70 academic papers, industry reports, and technical documents, this work synthesizes insights from real-world case studies across platforms like Google Play Protect, Microsoft Defender, Amazon Web Services (AWS), Apple App Store, OpenAI Plugin Stores, Hugging Face Spaces, and GitHub, alongside emerging initiatives like the SAFE Framework and AI-driven anomaly detection. We conclude with practical recommendations for responsible and transparent LLM deployment and trustworthy AI, including model watermarking, adversarial defense, and cross-industry collaboration, setting a new benchmark for rigorous, holistic cybersecurity research at the intersection of AI and threat defense, and offering a roadmap for secure, scalable LLM systems that serves as a critical reference for researchers, engineers, and security leaders navigating the complex challenges of AI-driven cybersecurity.

13:00 JSTLLM/生成AIエージェントロボティクス

RAG ベースのメモリとマルチモーダル コーチ エージェントを使用したエンドツーエンドの LLM 飛行計画

人間のパイロットの意図と自律的な飛行操作の間のギャップを埋めることは、現実世界の電動垂直離着陸 (eVTOL) 航空機の展開にとって重要です。飛行計画は従来、人間の柔軟な好みを組み込むのが難しい古典的なアルゴリズムに依存していました。 RAG ベースのメモリとマルチモーダル コーチ エージェントを備えたエンドツーエンドの大規模言語モデル (LLM) 飛行計画ツールである FRAMe を紹介します。当社のシステムは、プランナー LLM とマルチモーダル コーチ エージェントおよび検索拡張生成 (RAG) ベースのメモリを統合し、人間の飛行オペレーターの好みに合わせながらミッションの制約を満たす飛行計画を生成します。私たちは、さまざまな難易度の現実世界にインスピレーションを得たさまざまなシナリオでシステムをデモンストレーションします。 4 つの LLM にわたって、完全な FRAMe システム (RAG およびコーチ) は、すべてのプランナーに対して最高の妥当性 (合計で最大 93.8%、最強のプランナーの Easy シナリオでは 99%) をもたらし、メトリックに余裕があるオペレーターが好む方向にプリファレンス関連のメトリックをシフトします。 FRAMe は、高度な LLM を人間中心のミッション計画に導入し、自然言語の指示を安全で効率的かつ柔軟な飛行ルートに変換する方法を示します。コードは github.com/amin-tabrizian/FlightPlanningLLMs から入手できます。

原文 (English)

End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent

Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landing (eVTOL) aircraft deployment. Flight planning traditionally relies on classic algorithms that struggle to incorporate flexible human preferences. We present FRAMe, an End-to-End Large Language Model (LLM) Flight Planning tool with RAG-based Memory and Multi-modal Coach Agent. Our system integrates a planner LLM with a multi-modal coach agent and retrieval augmented generation (RAG)-based memory to generate flight plans that satisfy mission constraints while aligning with human flight operator preferences. We demonstrate the system in a range of real-world-inspired scenarios of varying difficulty levels. Across four LLMs, the full FRAMe system (RAG and coach) yields the highest validity for every planner (up to 93.8% aggregate, 99% on Easy scenarios for the strongest planner) and shifts preference-relevant metrics in the operator-favored direction where the metric has headroom. FRAMe signifies how advanced LLMs can be deployed for human-centric mission planning, translating natural language instructions into safe, efficient, and flexible flight routes. The code is available at: github.com/amin-tabrizian/FlightPlanningLLMs

13:00 JST研究/論文

MIONet のハイブリッド最小二乗法/勾配降下法

この論文では、MIONet のトレーニングを高速化するための効率的なハイブリッド最小二乗法/勾配降下法 (LSGD) 法を提案します。このメソッドは、DeepONets の LSGD メソッドを一般化したものです。 MIONet は複数のブランチ ネットワークとトランク ネットワークのエントリごとの積の合計であるため、各ブランチ ネットワークの最終層パラメータに関する多重線形関数とみなすことができます。これらのパラメーターのセットは、単一のブランチ ネットワークの LS システムを順番に解く交互最小二乗法を使用して最適化できます。大規模なシステム行列を処理するために、クロネッカー積とカトリラオ積、およびテンソル置換行列を導入して、大きな行列を小さな行列に因数分解します。私たちの方法は、各分岐の最後の層パラメーターの正則化項を備えた一般的なタイプの $L^2$ 損失と互換性があり、線形演算子を各損失項の MIONet 出力に適用できます。

原文 (English)

Hybrid Least Squares/Gradient Descent Methods for MIONets

In this paper, we propose an efficient hybrid least squares/gradient descent (LSGD) method for MIONets to accelerate training. This method generalizes the LSGD method for DeepONets. Since MIONet is the sum of the entrywise product of multiple branch networks and a trunk network, it can be viewed as a multilinear function with respect to the last layer parameters of each branch network. These sets of parameters can be optimized using the alternating least squares method, where we solve the LS system for a single branch network in turn. To handle the large-sized system matrix, we introduce Kronecker and Khatri-Rao products and tensor permutation matrices to factor the large matrix into small ones. Our method is compatible with a general type of $L^2$ loss with regularization terms for the last layer parameters of each branch, where linear operators can be applied to the MIONet output in each loss term.

13:00 JSTロボティクス

WAM-TTT: テスト時に人間のプレイを観察することで世界アクション モデルを操作する

新しいタスクのバリアントやユーザーが好む動作に向けてロボット基盤モデル (RFM) を操作することは依然として困難であり、多くの場合、追加のロボットのデモンストレーション、タスク固有の微調整、または長いコンテキストの調整が必要になります。私たちは、生の人間のビデオから世界のアクション モデルを操作するためのテスト時トレーニング フレームワークである WAM-TTT を紹介します。 WAM-TTT は、人間のビデオを模倣する軌跡として扱うのではなく、自己監視型ビデオ予測を通じて、凍結された WAM 内の軽量の適応メモリにビデオを吸収します。この記憶を制御に役立てるために、人間とロボットのペアのデータとキーと値の記憶再構成目標を使用して、人間のデモンストレーションとロボットの動作を一致させるメタトレーニング ステージを導入します。テスト時には、ラベルのない人間のビデオだけをメモリに適応させる必要があり、事前トレーニングされた WAM はフリーズされたままになります。これにより、基礎モデルの一般化機能を維持しながら、ロボットの動作、人間側の注釈、タスク固有の微調整を必要とせずに、効率的で再利用可能なステアリングが可能になります。広範な実験により、WAM-TTT は、さまざまな操作タスクや一般化設定にわたって、コンテキスト内のヒューマン ビデオ コンディショニング ベースラインよりも一貫して優れたパフォーマンスを発揮することが示されています。

原文 (English)

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video prediction. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key--value memory reconstruction objective. At test time, only unlabeled human videos are required to adapt the memory, while the pretrained WAM remains frozen. This enables efficient and reusable steering without robot actions, human-side annotations, or task-specific fine-tuning, while preserving the generalization ability of the foundation model. Extensive experiments show that WAM-TTT consistently outperforms in-context human-video conditioning baselines across diverse manipulation tasks and generalization settings.

13:00 JST研究/論文

燃料密度予測のための物理学に基づく時空間ニューラル モデル

この論文では、物理的制約と領域知識を深層学習モデルに統合してモデルの精度と安定性を向上させる、燃料密度予測のための物理ガイド付き機械学習 (PGML) フレームワークを紹介します。私たちは、ConvLSTM、Adaptive Fourier Neural Operator (AFNONet)、および Video Vision Transformer (ViViT) という 3 つの深層学習アーキテクチャを調査して、燃料密度の時空間変化をモデル化します。私たちのアプローチには、質量保存燃料輸送項や拡散率推定など、微分可能な物理学に基づいた項が損失関数に組み込まれています。複数の独立したトライアルを平均した実験結果は、提案された PGML フレームワークが、精度と安定性の両方において、物理的制約のない純粋にデータ駆動型のベースラインよりも優れていることを示しています。このフレームワークにより、計算効率が高く、物理的に妥当な火災予測が可能になり、適応的な処方火災管理がサポートされます。

原文 (English)

Physics-guided spatiotemporal neural models for fuel density prediction

This paper presents a physics-guided machine learning (PGML) framework for fuel density prediction, integrating physics constraints and domain knowledge into deep learning models to enhance model accuracy and stability. We explore three deep learning architectures -- ConvLSTM, Adaptive Fourier Neural Operator (AFNONet), and Video Vision Transformer (ViViT) -- to model the spatiotemporal evolution of fuel density. Our approach incorporates differentiable physics-informed terms in the loss function, including a mass-conserving fuel transport term and a rate-of-spread estimation. Experimental results, averaged across multiple independent trials, demonstrate that the proposed PGML framework outperforms purely data-driven baselines without physics constraints in both accuracy and stability. This framework enables computationally efficient, physically plausible fire forecasting to support adaptive prescribed burn management.

13:00 JST研究/論文

携帯電話トラフィック予測のためのピーク強化を備えたマルチモーダル時空間周波数融合

携帯電話ネットワーク トラフィックの正確な予測は、最新のモバイル通信システムにおけるネットワーク計画、リソース割り当て、およびサービス品質保証に不可欠です。現実世界の交通は、外部の都市事象によって引き起こされる爆発的な内因性ダイナミクスや外乱を示すことが多く、信頼性の高い予測が非常に困難になります。既存の時空間トラフィック予測手法のほとんどは、主に単一モダリティ内の固有のトラフィック パターンや構造的関係に焦点を当てており、外因性のコンテキスト信号とともにバースト動作をモデル化することはほとんどありません。この問題に対処するために、外部のコンテキスト情報を統合するマルチモーダルなセルラー トラフィック予測フレームワークである \textbf{MSPF-Net} を提案します。具体的には、MSPF-Net は、時間的、空間的、スペクトル的なトラフィック パターンをキャプチャするための時空間周波数トラフィック エンコーダ、突然のスパイクのバースト認識表現を抽出するためのピーク拡張モジュール、都市ニュース ストリームを外因性コンテキスト埋め込みにエンコードするためのニュース コンテキスト表現モジュール、およびこれらの異種信号を適応的に統合して予測を生成するためのダイナミック フュージョン予測モジュールで構成されます。ミラノ、トレント、LTE のトラフィック データセットに関する実験では、トラフィック ダイナミクス、バースト パターン、およびニュースのコンテキスト信号を共同モデリングすることで、予測パフォーマンスを効果的に向上できることが実証されました。

原文 (English)

Multimodal Spatiotemporal-Frequency Fusion with Peak Enhancement for Cellular Traffic Forecasting

Accurate forecasting of cellular network traffic is essential for network planning, resource allocation, and quality-of-service assurance in modern mobile communication systems. Real-world traffic often exhibits bursty endogenous dynamics and disturbances triggered by external urban events, which makes reliable prediction highly challenging. Most existing spatiotemporal traffic forecasting methods primarily focus on intrinsic traffic patterns or structural relationships within a single modality, and rarely model burst behavior together with exogenous contextual signals. To address this issue, we propose \textbf{MSPF-Net}, a multimodal cellular traffic forecasting framework that integrates external contextual information. Specifically, MSPF-Net consists of a Spatiotemporal-Frequency Traffic Encoder for capturing temporal, spatial, and spectral traffic patterns, a Peak Enhancement Module for extracting burst-aware representations of sudden spikes, a News Context Representation Module for encoding urban news streams into exogenous contextual embeddings, and a Dynamic Fusion Prediction Module for adaptively integrating these heterogeneous signals to generate forecasts. Experiments on the Milano, Trento, and LTE traffic datasets demonstrate that jointly modeling traffic dynamics, burst patterns, and news contextual signals can effectively improve forecasting performance.

13:00 JST画像/動画生成

生成 AI アーキテクチャを使用したマルチモーダル ニューロイメージング機能の潜在グラフ エンコーディング

生成モデルにより、特徴の生成と再構成のために複雑な神経画像データのエンコードが可能になりますが、適切なエンコードと潜在空間プロセスを備えた最適なアーキテクチャ フレームワークを開発することは、脳の構造的および機能的特性を研究するために重要です。私たちは、符号化戦略、潜在的なマルチモーダル融合、および生成モデルの選択の系統的な評価を通じて、構造的および機能的磁気共鳴画像法 (MRI) 特徴のためのマルチモーダル生成フレームワークを設計します。大規模な神経画像データセットからの構造灰白質ボリューム (GMV) と静的機能ネットワーク接続 (sFNC) の特徴を使用して、変分オートエンコーダー (VAE)、トランスフォーマー、敵対的生成ネットワーク (GAN)、拡散モデルを含む生成フレームワークを分析します。低次元の潜在空間への機能接続のモダリティを意識したグラフ エンコーディングを採用するアーキテクチャは、ベクトル化エンコーダや直接データ空間アプローチよりも優れたパフォーマンスを発揮します。提案されたマルチモーダル グラフ VAE (gMMVAE) は、生成忠実度、再構成品質、効率、潜在空間識別能の複数の指標にわたって代替生成バリアントを上回り、堅牢なマルチモーダル神経画像解析の可能性を強調しています。

原文 (English)

Latent graph encoding of multimodal neuroimaging features with generative AI architectures

While generative models enable encoding of complex neuroimaging data for feature generation and reconstruction, developing optimal architectural frameworks with appropriate encoding and latent space processes is crucial for studying structural and functional properties of the brain. We design a multimodal generative framework for structural and functional magnetic resonance imaging (MRI) features through systematic evaluation of encoding strategies, latent multimodal fusion, and generative model selection. Using structural gray matter volume (GMV) and static functional network connectivity (sFNC) features from a large neuroimaging dataset, we analyze generative frameworks involving variational autoencoders (VAEs), transformers, generative adversarial networks (GANs), and diffusion models. Architectures that employ modality-aware graph encoding of functional connectivity into a lower-dimensional latent space outperform vectorized encoders or direct data space approaches. The proposed multimodal graph VAE (gMMVAE) surpasses alternative generative variants across multiple metrics for generation fidelity, reconstruction quality, efficiency, and latent space discriminability, highlighting its potential for robust multimodal neuroimaging analysis.

13:00 JST研究/論文

Gimitest: 強化学習ポリシーをテストするための包括的なツール

強化学習 (RL) ポリシーは安全ではなく、攻撃に対して脆弱になる可能性があります。既存の自動テスト手法は、選択された環境、テスト シナリオ、RL アルゴリズムのみを対象としているため、信頼性の確保が課題となることがよくあります。これに対処するために、さまざまな条件下でシングルおよびマルチエージェントの RL ポリシーをテストするための包括的なフレームワークを提案します。このフレームワークの実装である Gimitest は、さまざまなジム フレームワークをサポートし、統合されたコンポーネントの変更を可能にするオープンソース ツールです。この記事では、フレームワークについて説明し、Gimitest の機能とアーキテクチャについて詳しく説明します。公式のファラマ体育館やペッティングズーなどの環境で複数の RL ポリシーをテストする際の有効性を示します。

原文 (English)

Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies

Reinforcement learning (RL) policies can be unsafe and vulnerable to attacks. Ensuring their reliability is often a pain point as existing automated testing methods target only selected environments, testing scenarios, and RL algorithms. To address this, we propose a comprehensive framework for testing single- and multi-agent RL policies under varying conditions. Our implementation of this framework, Gimitest, is an open-source tool that supports various gym frameworks and allows for modifications of their integrated components. This article describes the framework and details Gimitest's functionality and architecture. It showcases its effectiveness in testing multiple RL policies in environments such as the official Farama Gymnasium and PettingZoo.

13:00 JST画像/動画生成

AnchorPrune: ビジュアル トークン プルーニングのための関連性に基づいたコンテキスト拡張

高解像度の入力では数千のビジュアル トークンが導入され、その多くは特定のクエリに対して冗長であるため、大規模なビジョン言語モデルにはかなりの推論コストがかかります。既存の枝刈り手法では、クエリの関連性とトークンの多様性を組み合わせることがよくありますが、これらの目的は、積極的な圧縮の下では矛盾する可能性があります。関連性主導の選択では、相関する局所的な証拠に予算が集中しすぎる可能性がありますが、多様性主導の選択では、不可欠なトークンが抑制されたり、明確ではあるが情報のない領域が保持されたりする可能性があります。最初に保護された関連性アンカーを構築し、次にそれを補完的な視覚的コンテキストで拡張する、トレーニング不要のフレームワークである AnchorPrune を紹介します。 AnchorPrune は、関連性でランク付けされたトークンのノベルティ プロファイルからアンカー サイズを適応的に決定し、クエリクリティカルな証拠のコンパクトなセットを保存し、重要度に重み付けされたノベルティを通じて残りの予算を割り当て、アンカーに関連する有益で冗長でないコンテキストを回復します。この順序付けされたデザインにより、コンテキストの拡張によって不可欠なクエリ キューが置き換えられるのを防ぎ、全体的な視覚的範囲が向上します。 AnchorPrune は軽量でアーキテクチャを認識しており、再トレーニングもモデルの変更も必要ありません。画像およびビデオの視覚言語モデルとベンチマーク全体で、特に厳しい圧縮下で、トレーニング不要のベースラインと比較して精度と効率のトレードオフを一貫して改善します。 LLaVA-NeXT-7B では、AnchorPrune は 2,880 個のビジュアル トークンのうち 160 個のみを使用して、フルトークンのパフォーマンスの 97.6% を維持します。これらの結果は、効率的なマルチモーダル推論のための効果的な原理として、関連性にアンカーされた文脈拡張を確立します。コードは https://github.com/MULTI-cau/AnchorPrune で入手できます。

原文 (English)

AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning

Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of which are redundant for a given query. Existing pruning methods often combine query relevance and token diversity, yet these objectives can conflict under aggressive compression: relevance-driven selection may overconcentrate the budget on correlated local evidence, while diversity-driven selection may suppress indispensable tokens or retain distinct but uninformative regions. We introduce AnchorPrune, a training-free framework that first constructs a protected relevance anchor and then expands it with complementary visual context. AnchorPrune adaptively determines the anchor size from the novelty profile of relevance-ranked tokens, preserving a compact set of query-critical evidence, and allocates the remaining budget through importance-weighted novelty to recover informative, non-redundant context relative to the anchor. This ordered design prevents contextual expansion from displacing indispensable query cues while improving overall visual coverage. AnchorPrune is lightweight, architecture-aware, and requires neither retraining nor model modification. Across image and video vision-language models and benchmarks, it consistently improves the accuracy-efficiency trade-off over training-free baselines, particularly under severe compression. On LLaVA-NeXT-7B, AnchorPrune preserves 97.6% of full-token performance using only 160 of 2,880 visual tokens. These results establish relevance-anchored contextual expansion as an effective principle for efficient multimodal inference. Code is available at https://github.com/MULTI-cau/AnchorPrune.

13:00 JST研究/論文

組み込みグリーン学習: 逆偏微分方程式による多様体上の教師あり学習

Intrinsic Green's Learning (IGL) を紹介します。これは、ソース項がデータから学習される線形偏微分方程式の解として多様体上のターゲット関数をモデル化するフレームワークです。 IGL は、ターゲットを直接近似するのではなく、ソースを学習し、それをグリーンのカーネルに対して統合します。エンコーダは、ソースとカーネルの両方が低ランクのテンソルとして分解される多様体上の低次元座標チャートを発見し、高次元の積分を、固有次元で線形のコストを持つ独立した 1 次元の積分に分解します。 2 段階のアルゴリズムにより、座標検出がソース フィッティング (凸に近い線形ソルバ) から分離され、ジョイント トレーニングの次元崩壊が防止されます。各座標の学習可能なゲートにより、多様体の固有次元が自動的に検出されます。合成多様体と MNIST で IGL を検証し、最適に近い分類と固有次元の自動回復を同時に実現します。

原文 (English)

Intrinsic Green's Learning: Supervised Learning on Manifolds via Inverse PDE

We introduce Intrinsic Green's Learning (IGL), a framework that models a target function on a manifold as the solution to a linear PDE whose source term is learned from data. Rather than approximating the target directly, IGL learns a source and integrates it against a Green's kernel. An encoder discovers a low-dimensional coordinate chart on the manifold where both the source and the kernel decompose as low-rank tensors, collapsing a high-dimensional integral into independent one-dimensional integrals with cost linear in the intrinsic dimension. A two-stage algorithm separates coordinate discovery from source fitting, a near-convex linear solve, preventing the dimensional collapse of joint training. Learnable gates on each coordinate automatically discover the intrinsic dimension of the manifold. We validate IGL on synthetic manifolds and on MNIST, where it simultaneously achieves near-optimal classification and automatic recovery of the intrinsic dimension.

13:00 JST研究/論文

ディープ フィードフォワード ReLU ネットワークの原理について

ディープ フィードフォワード ニューラル ネットワークのアーキテクチャは、システム全体として、または他のアーキテクチャのサブネットワークとして、ディープ ラーニングのいたるところに存在するため、そのメカニズムはニューラル ネットワークのブラック ボックスの重要な要素です。この論文は、最も単純な 2 層 ReLU ネットワークに基づいて、複数の隠れ層を持つ深いフィードフォワード ReLU ネットワークのメカニズムを系統的に研究し、逆伝播アルゴリズムによって得られる学習解を説明することに成功しました。パスの概念、特にパス間の関係という観点は、ブラック ボックスの謎を解明する上で中心的な役割を果たします。深い ReLU ネットワークのユニットは、2 層の場合の超平面の代わりに、入力空間を分割する区分的線形多様体を形成できることが示されています。隠れ層ユニットを効率的に使用して線形関数と入力空間の分割の両方を生成する方法も中心的な問題です。 2 層 ReLU ネットワークの原理は、複数の厳密な部分順序や連続性制限など、より深いケースに大幅に一般化できます。提案された基本原理と単純な原理を組み合わせると、トレーニング ソリューションを含む複雑なインスタンス化が可能になり、この意味でディープ フィードフォワード ReLU ネットワークのブラック ボックスが明らかになります。

原文 (English)

On the Principles of Deep Feedforward ReLU Networks

The architecture of deep feedforward neural networks is ubiquitous in deep learning, either as a whole system or as a subnetwork of other architectures, and thus its mechanism is a key ingredient of the black box of neural networks. On the basis of the simplest two-layer ReLU network, this paper systematically studies the mechanism of deep feedforward ReLU networks with multiple hidden layers and successfully explains the training solution obtained by the back-propagation algorithm. The concept of a path, especially in terms of the relationships between paths, plays a central role in uncovering the mystery of the black box. It is shown that a unit of a deep ReLU network can form a piecewise linear manifold to divide the input space, instead of a hyperplane of the two-layer case. How to efficiently use the hidden-layer units to produce both linear functions and partitions of the input space is also a central problem. The principles of a two-layer ReLU network can be generalized to the deeper case to a large extent, such as multiple strict partial orders and continuity restriction. The combination of the basic and simple principles proposed can yield complicated instantiations including the training solutions, and in this sense the black box of deep feedforward ReLU networks is revealed.

13:00 JSTLLM/生成AI

事前トレーニング済み言語モデル埋め込みのためのリーマン幾何学

事前トレーニングされた言語モデルの埋め込みの幾何学的構造を理解することは、解釈可能性と安全性にとって重要です。文レベルの分類信号が文脈上のトークン埋め込みのリーマン幾何学に存在するかどうかを尋ね、学習されたエンコーダーの分析ヤコビアンからトークンごとのプルバック メトリックを抽出し、それらを対称正定多様体 (SPD) 上のフルエシェ平均と集計することでそれを調査します。この手順をリーマン平均プーリング (RMP) と呼びます。自明ではない言語構造を持つ 3 つのデータセット (CoLA、CREAK、RTE) にわたって、RMP はユークリッド平均プーリングよりも優れたパフォーマンスを示しましたが、アノテーション駆動型の語彙アーティファクトを除去するために構築されたベンチマークである FEVER-Symmetric では、この手法は正確に偶然性を維持しました。アブレーションの結果、ランダムに初期化されたエンコーダーとフレチェ集合体が組み合わされて、信号を含む 3 つのデータセットのうち 2 つでユークリッド プーリングをすでに上回っており、学習された多様体構造ではなく幾何学的集合体にゲインの発生源が局在していることがわかります。トレーニングされたエンコーダーは、特に 3 つの信号を含むデータセットの中で最も知識量の多い CREAK に追加信号を提供します。

原文 (English)

Riemannian Geometry for Pre-trained Language Model Embeddings

Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and probe it by extracting per-token pullback metrics from a learned encoder's analytical Jacobian and aggregating them with the Fr\'echet mean on the symmetric positive definite (SPD) manifold; we call this procedure Riemannian Mean Pooling (RMP). Across three datasets with non-trivial linguistic structure (CoLA, CREAK, RTE), RMP outperforms Euclidean mean pooling, while on FEVER-Symmetric, a benchmark constructed to remove annotation-driven lexical artifacts, the method correctly stays at chance. Ablations show that a randomly initialised encoder combined with Fr\'echet aggregation already beats Euclidean pooling on two of the three signal-bearing datasets, localising the source of the gain to the geometric aggregation rather than to learned manifold structure; the trained encoder contributes additional signal specifically on CREAK, the most knowledge-heavy of the three signal-bearing datasets.

13:00 JST画像/動画生成

会話型画像編集で暗黙的な保存意図を明示化する

会話型の画像編集では、表示されているコンテンツだけでなく、順番が変わると一時的に消えるコンテンツも保存する必要があります。新たに追加または変更されたコンテンツが以前に表示されていた領域を遮る場合、意味的に変更されていなければ、その領域は再表示されるはずです。しかし、既存のシステムでは、このような隠蔽されているが変更されていないコンテンツの回復に失敗することが多く、一貫性のない結果や幻覚のような結果が生じます。会話型画像編集における一時保存のための診断ベンチマークである OCCUR-Bench を紹介します。 OCCUR-Bench は、歴史的な修復の参照を使用した多様な咬合と発現のシナリオを提供し、妥当な再生ではなく忠実な修復の評価を可能にします。また、復元を意識した命令と歴史的な視覚参照を組み合わせることで、暗黙的な保存を明示的に行う、トレーニング不要のフレームワークである ReSpec も提案します。編集履歴が与えられると、ReSpec は何を保持すべきかを特定し、欠落している視覚的証拠を提供する履歴画像の状態を選択し、結果として得られる指示と参照画像に基づいてインコンテキスト エディターの条件を設定します。実験では、ReSpec が OCCUR-Bench での復元の忠実性と時間的一貫性を向上させることが示されており、現在の画像だけでなく編集履歴における保存の必要性が強調されています。

原文 (English)

Making Implicit Preservation Intent Explicit in Conversational Image Editing

Conversational image editing requires preserving not only visible content, but also content that temporarily disappears across turns. When newly added or modified content occludes a previously visible region, that region should reappear if it was never semantically changed. However, existing systems often fail to recover such occluded-but-unchanged content, producing inconsistent or hallucinated results. We introduce OCCUR-Bench, a diagnostic benchmark for temporal preservation in conversational image editing. OCCUR-Bench provides diverse occlusion-and-revelation scenarios with historical restoration references, enabling evaluation of faithful restoration rather than plausible regeneration. We also propose ReSpec, a training-free framework that makes implicit preservation explicit by pairing restoration-aware instructions with historical visual references. Given an editing history, ReSpec identifies what should persist, selects the historical image state that provides missing visual evidence, and conditions an in-context editor on the resulting instruction and reference image. Experiments show that ReSpec improves restoration fidelity and temporal consistency on OCCUR-Bench, highlighting the need to ground preservation in editing history rather than only the current image.

13:00 JSTLLM/生成AIエージェント

漸進的な結晶化: エージェントの探索を本番環境での決定論的で低コストのワークフローに変える

IT 運用に導入された AI エージェントは、以前に解決された問題であっても、実行ごとに完全な LLM 推論が必要となるため、通常は恒久的なコスト センターとなります。このペーパーでは、エージェントの探索を永続的な実行モデルではなく発見メカニズムとして扱うライフサイクルである、プログレッシブ クリスタルリゼーションを紹介します。これは、完全にエージェントによって調整されたワークフローからハイブリッド、完全な決定論的なワークフローまでの 3 段階の実行分類を定義し、繰り返し検証されたエージェントの動作をより安価で再現性の高い決定論的なワークフローに変換すると同時に、後退したワークフローを自動的に降格する証拠に基づいたプロモーション メカニズムを定義します。毎月数万件のインシデントを処理する実稼働クラウド ネットワーキング AIOps システムで評価したこのアプローチでは、8 か月間で確定的実行が 0% から 45% に増加し、インシデント量が 2 倍になったにもかかわらず、インシデントごとのエージェント コストが 70% 以上削減され、再現性と監査可能性の向上により安全性が向上しました。この文書では、実行分類、昇進および降格の基準、トレース抽出方法、経済モデル、安全性に関する考慮事項も示し、妥当性に対する制限と脅威についても説明します。

原文 (English)

Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production

AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for previously solved problems. This paper introduces progressive crystallization, a lifecycle that treats agent exploration as a discovery mechanism rather than a permanent execution model. It defines a three-stage execution taxonomy, from fully agent-orchestrated to hybrid to fully deterministic workflows, together with an evidence-based promotion mechanism that converts repeatedly validated agent behaviors into cheaper and more reproducible deterministic workflows, while automatically demoting workflows that regress. Evaluated on a production cloud networking AIOps system processing tens of thousands of incidents per month, the approach increased deterministic execution from 0% to 45% over eight months, reduced per-incident agent costs by more than 70% despite doubling incident volume, and improved safety through greater reproducibility and auditability. The paper also presents the execution taxonomy, promotion and demotion criteria, trace extraction methodology, economic model, safety considerations, and discusses limitations and threats to validity.

13:00 JST研究/論文

複雑さを考慮した、インタラクションを意識した表形式データの解釈可能なモデル

表形式データの本質的に解釈可能な分類子は通常、ユーザーが直接検査できる疎な特徴、ルール、またはパターンに依存します。これらの方法に共通する周辺特徴スクリーニング ステップでは、他の変数との結合構成を通じてのみ予測値が現れる変数を破棄することができます。我々は、Interaction Aware Interpretable Machine Learning (IAIML) を紹介します。これは、特徴ごとの適応的離散化、有限グリッドのペアごとの相互作用スコアリング、および分割された説明バジェットという 3 つの調整されたメカニズムを通じてこの制限に対処するフレームワークです。検出された相互作用は、2 つの戦略のいずれかによってルーティングされます。1 つはスクリーニング フィルターを緩和して相互作用でサポートされる変数がパターン検索に入るようにするか、またはスパースな下流分類器の明示的なペア項を構築する方法です。ネストされた相互検証の下で評価された、24 の実世界の表形式ベンチマークと 16 の合成相互作用ストレス テストで構成される 40 のデータセット パネルで、IAIML は、調整された勾配ブースト アンサンブルの 1.4 ポイント以内の平均 AUC を達成しながら、必要な適合説明コンポーネントの数はおよそ 14 ~ 28 分の 1 です。強力なペアごとの相互作用構造と低い周辺シグナルを持つデータセットでは、IAIML はすべてのベースラインを上回ります。コンパクトな解釈可能なメソッドの中で、IAIML は AUC とコンポーネント数の点で RuleFit に匹敵し、チューニングのコストが低くなります。 EBM は、ルックアップ テーブルの設置面積が大幅に大きくなり、パネル全体にわたって小さいながらも重要な AUC の利点が得られます。ペアごとのスコープを超えた高次の相互作用を必要とするデータセットではパフォーマンスが低下します。コンポーネント分離アブレーションは、適応的離散化と相互作用を意識した入院がそれぞれ段階的に寄与していることを確認します。これらの結果は、制限された説明サイズと機能インタラクションの制御された処理が設計要件である設定に適した、コンパクトでインタラクションを意識したフレームワークとして IAIML を裏付けています。

原文 (English)

Complexity-Budgeted, Interaction-Aware Interpretable Model for Tabular Data

Inherently interpretable classifiers for tabular data typically rely on sparse features, rules, or patterns that users can inspect directly. The marginal feature-screening step common to these methods can discard variables whose predictive value emerges only through joint configurations with other variables. We present Interaction Aware Interpretable Machine Learning (IAIML), a framework that addresses this limitation through three coordinated mechanisms: adaptive per-feature discretization, finite-grid pairwise interaction scoring, and a partitioned explanation budget. Detected interactions are routed through one of two strategies: relaxing the screening filter so that interaction-supported variables enter the pattern search, or constructing explicit pair terms for a sparse downstream classifier. On a 40-dataset panel comprising 24 real-world tabular benchmarks and 16 synthetic interaction stress tests, evaluated under nested cross-validation, IAIML achieves mean AUC within 1.4 points of tuned gradient-boosted ensembles while requiring roughly 14--28 times fewer fitted explanation components. On datasets with strong pairwise interaction structure and low marginal signal, IAIML outperforms all baselines. Among compact interpretable methods, IAIML is comparable to RuleFit in AUC and component count and is less expensive to tune. EBM obtains a small but significant AUC advantage across the full panel, with a substantially larger lookup-table footprint. Performance degrades on datasets requiring higher-order interactions beyond the pairwise scope. Component-isolated ablations confirm that adaptive discretization and interaction-aware admission each contribute incrementally. These results support IAIML as a compact, interaction-aware framework appropriate for settings where bounded explanation size and controlled treatment of feature interactions are design requirements.

13:00 JST研究/論文

群を超えた乗算: 変圧器回路における層状フーリエ機構

トランスフォーマーは、アルゴリズム推論を学習する優れた能力を実証してきましたが、機械論的な分析は主に、循環加算やグループ合成などのグローバルに可逆な演算に焦点を当ててきました。この研究では、小さな変換器が、ゼロ除数の存在により基本的に非可逆演算である合成法に対するモジュラー整数乗算をどのように学習するかを調査します。我々は、モノイド拡張を提案します。これは、学習された計算が単一のグローバル表現空間に依存しないことを示唆する、表現によるグループ構成 (GCR) の局所的な一般化です。代わりに、モデルは入力空間をローカル階層代数領域に分割します。そこではグループのような構造が存続し、フーリエ機構を適用できます。無平方剰余乗算で訓練された変換器では、埋め込みがこれらの領域の周りに組織され、注意がクラス依存のルーティングと低ランクの書き込み方向を示し、局所的な文字の特徴がモデルの出力ロジットの大部分を説明することがわかります。私たちの結果は、グループ操作に関して以前に特定された表現理論メカニズムがグループを超えてより一般的な構造に拡張できることを示唆しています。

原文 (English)

Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits

Transformers have demonstrated a remarkable ability to learn algorithmic reasoning, yet mechanistic analyses have mostly focused on globally invertible operations such as cyclic addition and group composition. In this work, we investigate how small transformers learn modular integer multiplication over composite moduli, a fundamentally non-invertible operation due to the presence of zero-divisors. We propose the monoid extension: a localized generalization of Group Composition via Representation (GCR) that suggests the learned computation does not rely on a single global representation space. Instead, the model partitions the input space into local hierarchical algebraic regions, where group-like structure survives and Fourier mechanisms can be applied. In transformers trained on square-free modular multiplication, we find that embeddings organize around these regions, attention exhibits class-sensitive routing and low-rank write directions, and local character features explain a large fraction of the model's output logits. Our results suggest that representation-theoretic mechanisms previously identified for group operations can extend beyond groups to more general structures.

13:00 JST画像/動画生成

階層のナビゲーション: 障害診断のための脳グラフの双曲線学習

機能的脳ネットワークは、ROI、コミュニティ、全脳レベルにわたって階層的組織を示し、ローカル処理、コミュニティ間の調整、およびグローバル統合をサポートします。最近の研究では、脳コミュニティを意識したモデリングが脳ネットワークの診断とバイオマーカー同定の両方に有益であることが実証されています。しかし、既存の脳グラフ モデリング手法は、ROI とコミュニティの相互作用をモデル化するのに苦労することが多く、ROI、コミュニティ、および全脳ネットワーク レベルにわたる階層を完全に活用することができません。この問題に対処するために、階層構造のモデリングにおける深層双曲学習に触発されて、脳ネットワーク分析用の Hyperbolic Learning on Brain Graphs (HLBG) と呼ばれる新しいフレームワークを提案します。 HLBG の中心となるアイデアは、双曲空間の固有の階層幾何学を利用して、ROI、機能コミュニティ、全脳ネットワーク間の階層関係をモデル化し、それによって脳ネットワーク データの階層を意識した高度に識別可能な表現を学習することです。具体的には、HLBG はまず、ROI、コミュニティ、および全脳ネットワークからの表現をローレンツ双曲空間に投影します。次に、2 つの幾何学的含意制約を介してマルチレベル階層が課されます。さらに、新しいグラフ対応 Mamba (GaMamba) モデルを導入します。このモデルは、トポロジ由来の構造プロンプトを Mamba に組み込んで、グラフ トポロジ情報を維持しながら長距離の依存関係をキャプチャします。 ABIDE-I および REST-MDD の実験では、HLBG が最先端の方法を上回っており、疾患に関連する機能的バイオマーカーを特定できることが実証されています。

原文 (English)

Navigating Hierarchy: Hyperbolic Learning on Brain Graphs for Disorder Diagnosis

Functional brain networks exhibit a hierarchical organization across ROI, community, and whole-brain levels, supporting local processing, inter-community coordination, and global integration. Recent studies have demonstrated that brain community-aware modeling is beneficial for both diagnosis and biomarker identification of brain networks. However, existing brain graph modeling methods often struggle to model ROI-community interactions, thereby failing to fully exploit the hierarchy across ROI, community, and whole-brain network levels. To address this issue, inspired by deep hyperbolic learning in modeling hierarchical structures, we propose a novel framework, termed Hyperbolic Learning on Brain Graphs (HLBG), for brain network analysis. The core idea of HLBG is to exploit the inherent hierarchical geometry of hyperbolic space to model the hierarchical relationships among ROIs, functional communities, and the whole-brain network, thereby learning hierarchy-aware and highly discriminative representations for brain network data. Specifically, HLBG first projects representations from ROIs, communities, and the whole-brain network into Lorentzian hyperbolic space. Then, the multi-level hierarchy is imposed via two geometric entailment constraints. In addition, we introduce a new Graph-aware Mamba (GaMamba) model, which incorporates topology-derived structural prompts into Mamba to capture long-range dependencies while preserving graph topological information. Experiments on ABIDE-I and REST-MDD demonstrate that HLBG outperforms state-of-the-art methods and identifies disorder-relevant functional biomarkers.

13:00 JST画像/動画生成

AT-Attn: 縦断的多峰性アルツハイマー病診断のための時間的認識クロスアテンション

長期的なアルツハイマー病 (AD) 診断サポートでは、臨床情報や画像情報が不定期の訪問で収集されることがよくあります。これらのマルチモーダルな観察を統合すると、診断評価が向上する可能性がありますが、単純な融合では、MRI にノイズが多い場合や断続的に利用できない場合にパフォーマンスが低下する可能性があります。我々は、MRI を長期的な臨床情報と統合するために、Change-and-Time エンコーディング、時間バイアス非対称クロスアテンション、ゲート フュージョンを組み合わせた時間認識マルチモーダル フレームワークである AT-Attn を提案します。我々は、構造 MRI、6 つの認知スケールの軌跡、および 7 つの静的臨床変数を使用して、患者レベルの 5 重交差検証の下で、1,520 人の患者の MRI 保持 ADNI コホートで AT-Attn を評価します。メインの非対称 AT-Attn モデルは、精度 0.719+/-0.024、マクロ F1 0.721+/-0.023、ROC-AUC 0.873+/-0.013、PR-AUC 0.783+/-0.018 を達成し、強力な表形式ベースラインとの競争力を維持しながら、ユニモーダルおよびナイーブ マルチモーダル融合ベースラインを上回るパフォーマンスを示します。これらの結果は、時間を意識した制約された融合戦略が、構造的 MRI が患者レベルの AD 診断支援のための臨床的に関連した補完情報に貢献するのに役立つ可能性があることを示唆しています。

原文 (English)

AT-Attn: Temporal-Aware Cross-Attention for Longitudinal Multimodal Alzheimer's Disease Diagnosis

In longitudinal Alzheimer's disease (AD) diagnosis support, clinical and imaging information is often collected at irregular visits. Integrating these multimodal observations may improve diagnostic assessment, but naive fusion can degrade performance when MRI is noisy or intermittently unavailable. We propose AT-Attn, a temporal-aware multimodal framework that combines Change-and-Time encoding, time-biased asymmetric cross-attention, and gated fusion to integrate MRI with longitudinal clinical information. We evaluate AT-Attn on an MRI-retained ADNI cohort of 1,520 patients using structural MRI, six cognitive-scale trajectories, and seven static clinical variables under patient-level five-fold cross-validation. The main asymmetric AT-Attn model achieves accuracy 0.719+/-0.024, macro F1 0.721+/-0.023, ROC-AUC 0.873+/-0.013, and PR-AUC 0.783+/-0.018, outperforming unimodal and naive multimodal fusion baselines while remaining competitive with strong tabular baselines. These results suggest that a temporal-aware and constrained fusion strategy can help structural MRI contribute clinically relevant complementary information for patient-level AD diagnosis support.

13:00 JSTロボティクスAlibaba

GeoProp: ジェネラリスト操作のためのビジョンにおけるロボット状態の接地

固有受容はロボット操作の基本ですが、標準的な融合手法では、固有受容を視覚トークンとの明示的な位置合わせを欠いた孤立したベクトルとして扱うことがよくあります。 3D 運動学と 2D 特徴マップの間に直接の対応がないと、操作ポリシーはシーン内のロボットの状態を定着させるのに苦労し、視覚のみのベースラインでさえもパフォーマンスを下回ることがよくあります。これに対処するために、明示的な幾何学的接地と空間特徴サンプリングを通じて固有受容と視覚を調整する軽量のプラグアンドプレイ アダプターである GeoProp を導入します。 GeoProp は、ロボットの状態を画像平面に投影して、局所的な視覚特徴をサンプリングし、接地状態トークンを構築します。次に、FiLM 変調を介して、状態由来の空間事前分布を対応する視覚特徴に注入します。モーションの意図を捉えるために、GeoProp は、最近の運動学から導出された短地平線の予測座標でフィーチャをさらにサンプリングし、先読みの視覚的コンテキストを提供します。 67 のタスクにわたって、GeoProp は 63 のシミュレーション タスクで拡散ポリシーを 8.7%、RoboTwin サブセットで pi_0 を 4.0% 改善し、現実世界では両方のポリシー ファミリ全体で平均 10.6% の向上をもたらしましたが、パラメーター数の追加は 2 ~ 3% のみでした。これらの結果は、GeoProp がジェネラリストの具体化されたポリシーにとってシンプルだが影響力の高い誘導バイアスであることを示しています。プロジェクトページ: https://alibaba-damo-academy.github.io/GeoProp/。

原文 (English)

GeoProp: Grounding Robot State in Vision for Generalist Manipulation

Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to ground the robot's state within the scene, frequently underperforming even vision-only baselines. To address this, we introduce GeoProp, a lightweight, plug-and-play adapter that aligns proprioception with vision through explicit geometric grounding and spatial feature sampling. GeoProp projects the robot state onto the image plane to sample localized visual features, constructing a grounded state token. It then injects state-derived spatial priors into the corresponding visual features via FiLM modulation. To capture motion intent, GeoProp further samples features at a short-horizon predicted coordinate derived from recent kinematics, providing look-ahead visual context. Across 67 tasks, GeoProp improves Diffusion Policy by 8.7% on 63 simulation tasks and pi_0 by 4.0% on the RoboTwin subset, and yields a 10.6% average gain across both policy families in the real world, while adding only 2-3% to the parameter count. These results demonstrate that GeoProp is a simple yet high-impact inductive bias for generalist embodied policies. Project page: https://alibaba-damo-academy.github.io/GeoProp/.

13:00 JST画像/動画生成

テキストから画像へのコンテキスト学習のための思考ツリー推論

テキストから画像へのインコンテキスト学習 (T2I-ICL) では、モデルはクエリ画像を生成するために、少数のショットのデモンストレーションから潜在的な構成パターンを推測する必要があります。最近の研究では、最先端のマルチモーダル大規模言語モデルは、特に構成推論の制限と即時構築に対する敏感さのために、この設定に苦労していることが示されています。この研究では、T2I-ICL 用の Tree-of-Thoughts (ToT) 推論フレームワークを提案します。これは、画像合成の最終プロンプトを構築する前に、複数の候補仮説を生成、評価、選択する多段階推論および選択層を導入します。提案されたアプローチは、別の推論分岐を探索し、一貫した解釈を選択することにより、即座の曖昧さと構成上の誤りを軽減します。提案されたアプローチを完全な ToT-T2IICL 推論パイプラインに実装し、CoBSAT ベンチマークで評価します。定性的および定量的結果の両方から、構造化されたマルチブランチ推論は、追加のトレーニングや微調整を行わなくても、ベースラインおよび思考連鎖プロンプト戦略と比較して、より一貫性があり、意味的に調整された画像生成につながることが示されています。

原文 (English)

Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning

In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generating a query image. Recent studies show that state-of-the-art multimodal large language models struggle with this setting, particularly due to limited compositional reasoning and sensitivity to prompt construction. In this work, we propose a Tree-of-Thoughts (ToT) reasoning framework for T2I-ICL that introduces a multi-stage reasoning and selection layer that generates, evaluates, and selects among multiple candidate hypotheses before constructing the final prompt for image synthesis. By exploring alternative reasoning branches and selecting a coherent interpretation, the proposed approach mitigates prompt ambiguity and compositional errors. We implement the proposed approach in a complete ToT-T2IICL inference pipeline and evaluate it on the CoBSAT benchmark. Both qualitative and quantitative results show that structured multi-branch reasoning leads to more consistent and semantically aligned image generation compared to baseline and Chain-of-Thought prompting strategies, without any additional training or fine-tuning.

13:00 JSTLLM/生成AIエージェント

Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning

Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. Ho…

13:00 JSTLLM/生成AIビジネス/資金調達GPT / ChatGPT

Predicting LLM Safety Before Release by Simulating Deployment

Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evid…

13:00 JSTロボティクス

Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators

Across robotics, World Models (WMs) are increasingly used to evaluate action policies by simulating the consequences of actions in an imagi…

13:00 JST研究/論文

Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency

We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5…

13:00 JST画像/動画生成ビジネス/資金調達

Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation

Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation…

13:00 JST研究/論文

DiPhon: Diffusion on Graphons for Scalable Graph Generation

Diffusion models represent a leading paradigm for graph generation, with notable impact in domains such as molecular design. Yet, scaling t…

13:00 JSTエージェント

ORCAID: Oblique Rule-Based Continuous-Action Interpretation for Deep RL Policies

Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex en…

13:00 JST研究/論文

FMMVCC: Fuzzy Mamba-based Multi-View Contrastive Clustering for Univariate Time Series

In many realistic scenarios, large volumes of time series data are generated with limited or expensive annotations. This limitation makes s…

13:00 JST研究/論文

Bayesian Optimization of Genetic Algorithm Hyperparameters in a Multi-Fidelity Framework for Efficient Lattice Material Design

This study presents a multi-fidelity framework for the systematic optimization of genetic algorithm (GA) hyperparameters. The framework int…

13:00 JST画像/動画生成

CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training

Accurately estimating urban carbon emissions is critical for sustainable urban planning, yet many existing approaches remain difficult to a…

13:00 JSTLLM/生成AIロボティクス

Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders

Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where…

13:00 JST研究/論文

POO-LPSP: Parallel Osprey Optimized Least Penalty-Squared Prioritization Methods for Priority Derivation in the Analytic Hierarchy Process

Pairwise comparison (PC) via pairwise reciprocal matrices (PRMs) is central to the Analytic Hierarchy Process (AHP). Although the tradition…

13:00 JST研究/論文

FedCVESA: Taking Away Training Data in Federated Learning via Correlation Value Encoding and Segmented Aggregation

Federated learning (FL) avoids explicit data exposure by keeping raw data on local clients, yet privacy risks remain in the training proces…

13:00 JST画像/動画生成研究/論文

HAJJv2-CrowdCount: Zero-Shot Benchmark for Dense Crowd Counting

Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumpti…

13:00 JST研究/論文

Hypergraph Neural Stochastic Diffusion: An SDE Framework for Uncertainty Estimation

Hypergraph neural networks have shown powerful capability in modeling higher-order relations, yet their predictive uncertainty remains unde…

13:00 JST研究/論文

Quantum simulation of real-world nonlinear dynamics via Koopman method

Nonlinear dynamics is ubiquitous in nature, ranging from chemical pattern formation to ocean circulation, yet its simulation on quantum com…

13:00 JST研究/論文

Latency-Aware Bid Acceptance under Operational Feasibility: A Public Benchmark with Hindsight Ceilings

Online truckload bid acceptance is a closed-loop stochastic decision problem in which a carrier or broker must, in real time, accept or rej…

13:00 JSTロボティクス

HumAIN: Human-Aware Implicit Social Robot Navigation

Effective social robot navigation requires sensitivity to human behavior, often revealed through subtle skeletal cues like gait and orienta…

13:00 JSTエージェント

Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors

AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent. AI Control usually studie…

13:00 JSTロボティクス

Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report

In embodied intelligence systems, the motion controller serves as the critical bridge between semantic reasoning and physical execution. Hu…

13:00 JST研究/論文

On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces

Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robu…

13:00 JSTLLM/生成AI画像/動画生成

When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs

Reliable confidence estimation remains a key limitation of test-time adaptation in vision-language models (VLMs), where prompt tuning impro…

13:00 JST画像/動画生成

Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation

While whole-body multimodal medical imaging scanners have been increasingly recognized for more effective medical applications, the excessi…

13:00 JSTエージェント

RLVP: Penalize the Path, Reward the Outcome

Agents acting on our behalf in the real world (e.g. placing phone calls) must learn online from costly, often irreversible interactions rat…

13:00 JSTLLM/生成AI

SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation

Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of produ…

13:00 JST研究/論文

Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data

Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Pr…

13:00 JSTLLM/生成AIエージェント研究/論文

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We ar…

13:00 JST研究/論文

Reward-Adaptive Iterative Discovery: A Case Study on Automated Game Testing for NHL26

Testing is a major effort for the gaming industry, requiring a significant part of development budget and people power. We present a case s…

13:00 JST研究/論文

TimEE: End-to-end Time Series Classification via In-Context Learning

Time series classification (TSC) is dominated by a two-stage paradigm: train a feature encoder -- either from scratch on the target dataset…

13:00 JST画像/動画生成

HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models

Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous vis…

13:00 JSTLLM/生成AIエージェント

Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LL…

13:00 JST研究/論文

Stability of Flow Models for Graph Signals

Generating signals on graphs requires permutation-equivariant models that exhibit stability with respect to relative structural perturbatio…

13:00 JSTエージェント

Creativity from Friction: Human-AI Interaction for Exploratory Structural Design

AI agents that generate final answers based on user input often do not meet the needs of creative fields. Fields such as structural design…

13:00 JST研究/論文

Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning

One-shot federated learning (OSFL) addresses the communication overhead of federated learning by limiting training to a single round, but d…

13:00 JSTエージェントロボティクスビジネス/資金調達

CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis

Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately sy…

13:00 JSTエージェント

Towards Agentic AI Governance: A Preliminary Assessment

Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Wid…

13:00 JSTLLM/生成AI

Future Confidence Distillation in Large Language Models

Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, where downstream decisi…

13:00 JST研究/論文

QCNN with Rough Path Signature Kernels

Time series analysis plays a vital role across a wide range of scientific and engineering domains but poses substantial computational chall…

13:00 JST研究/論文

ALER-TI: Aligned Latent Embedding Retrieval for Time Series Imputation

Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on localized temporal conte…

13:00 JSTLLM/生成AI

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving…

13:00 JSTLLM/生成AIGemma

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answ…

13:00 JST画像/動画生成

Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF

Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences.…

13:00 JSTエージェント

Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass

Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely b…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPT

Co-LMLM: Continuous-Query Limited Memory Language Models

Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it…

13:00 JSTLLM/生成AI

Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning

Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical resp…

13:00 JSTLLM/生成AI

Successor-Generator Planning with LLM-generated Heuristics

Heuristics are a central component of deterministic planning, particularly in domain-independent settings where general applicability is pr…

13:00 JSTロボティクス

Shared Modular Recurrence in Contextual MDPs for Universal Morphology Control

A universal controller for any robot morphology would greatly improve computational and data efficiency. Steps have been made towards such…

13:00 JSTエージェント

Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?

We investigate whether socioeconomic indicators, like household wealth, leave recoverable informational imprints in both satellite imagery…

13:00 JSTLLM/生成AIGPT / ChatGPT

LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?

Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their com…

13:00 JSTLLM/生成AIエージェント

AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org

Agentic AI systems increasingly connect large language models (LLMs) to external scientific tools, yet whether and when tool access improve…

13:00 JST研究/論文

Neutral Substrates: A Design Constraint for Shared Records Under Persistent Interpretive Disagreement

Shared accountability records are often used by parties who may never agree about causation, responsibility, or normative interpretation. F…

13:00 JSTLLM/生成AIビジネス/資金調達

SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care

Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-…

13:00 JSTLLM/生成AI

Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas

Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change…

13:00 JSTLLM/生成AIビジネス/資金調達

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for me…

13:00 JST研究/論文

An Adaptive Differentially Private Federated Learning Framework

Federated learning enables collaborative model training across distributed clients while preserving data privacy. However, in practical dep…

13:00 JST研究/論文

SOMtime the World Ain$'$t Fair: Violating Fairness Using Self-Organizing Maps

Unsupervised representations are widely assumed to be neutral with respect to sensitive attributes when those attributes are withheld from…

13:00 JST研究/論文

Power and Limitations of Aggregation in Compound AI Systems

When designing compound AI systems, a common approach is to query multiple copies of the same model and aggregate the responses to produce…

13:00 JSTLLM/生成AI画像/動画生成

EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to c…

13:00 JST研究/論文

Anomaly detection in time-series via inductive biases in the latent space of conditional normalizing flows

Deep generative models for anomaly detection in multivariate time-series are typically trained by maximizing observed data likelihood. Howe…

13:00 JST研究/論文GPT / ChatGPTMistral AIDeepSeek

Measuring the metacognition of AI

A robust decision-making process must take into account uncertainty, especially when the choice involves inherent risks. Because artificial…

13:00 JST研究/論文

Participatory provenance as representational auditing for AI-mediated public consultation

Artificial intelligence is increasingly deployed to synthesize large-scale public input in policy consultations and participatory processes…

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPTQwen

Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?

Modern coding agents increasingly delegate specialized subtasks to subagents, which are smaller, focused agentic loops that handle narrow r…

13:00 JST研究/論文

M$^3$: Reframing Training Measures for Discretized Physical Simulations

Neural surrogate models for physical simulations are trained on discretized samples of continuous domains, where the induced empirical meas…

13:00 JSTLLM/生成AI

要約が意思決定を歪める場合: LLM 圧縮財務分析における情報の忠実性

財務上の意思決定者は、直接調査できる以上に多くの情報に直面するため、コンテキストの圧縮が必要になります。しかし、大規模言語モデル (LLM) が財務情報源を圧縮すると、元の情報源によって裏付けられた投資判断が変更される可能性があります。私たちはこの問題を情報の忠実度として捉えます。圧縮により、ソースによって引き起こされた決定が変更されると忠実度が失われます。エージェント システムでは、このような損失が中間ステップで繰り返し発生し、意思決定プロセス全体で増幅する可能性があります。財務書類や決算報告のトランスクリプト全体にわたって、LLM ベースの圧縮により、流暢で事実に基づいた圧縮コンテキストが生成され、それにもかかわらず下流の意思決定が変更されることがわかりました。私たちは、忠実度の損失に関連する 2 つの診断パターンを分析します。1 つは、顕著な証拠は保持されますが、正しい解釈に必要な注意事項や文脈修飾子から分離される「脱文脈化」と、異なる圧縮器が同じソースの異なるビューを公開する「モデル依存性」です。次に、エージェント的コンテキスト圧縮を提案します。これは、複数の圧縮候補を生成し、元のソースとの相違点を監査します。私たちの結果は、財務圧縮は効率性や事実性だけでなく、意思決定に関連したコンテキストを保存する能力によっても評価されるべきであることを示唆しています。

原文 (English)

When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis

Financial decision-makers face more information than they can directly inspect, making context compression necessary. Yet when large language models (LLMs) compress financial source material, they can alter the investment judgment supported by the original source. We frame this problem as information fidelity: compression loses fidelity when it changes the decision induced by the source. In agentic systems, such losses may recur across intermediate steps and amplify throughout the decision process. Across financial filings and earnings-call transcripts, we find that LLM-based compression can produce fluent and factually plausible compressed contexts that nevertheless alter downstream decisions. We analyze two diagnostic patterns associated with fidelity loss: decontextualization, where salient evidence is retained but separated from the caveats and contextual qualifiers needed for correct interpretation, and model dependency, where different compressors expose different views of the same source. We then propose Agentic Context Compression, which generates multiple candidate compressions and audits their disagreements against the original source. Our results suggest that financial compression should be evaluated not only by efficiency or factuality, but also by its ability to preserve decision-relevant context.

13:00 JSTLLM/生成AI

HARC: 堅牢な安全調整のための有害性と拒否のカップリングの方向性

アライメントされた LLM が内部的にどのように安全性を表すかを理解することは、ジェイルブレイクが成功する理由を説明し、堅牢なアライメント戦略の設計に情報を提供するため、アライメントの脆弱性を診断するために重要です。これまでの研究では、整列された LLM がプロンプト側のトークン位置で残留ストリーム内の分離可能な方向として有害性と拒否をエンコードしていることが示されています。トークンが生成される前に拒否または有害性の方向を抑制することで、ジェイルブレイクがプロンプトエンコーディングで成功し、異なる攻撃クラスが有害性と拒否の面の分離可能な領域を占めることを示します。分析をレスポンス トークンの位置まで拡張すると、プロンプト側で入力を有害なものとして認識できなかった場合でも、モデルが有害なコンテンツを生成中にそのコンテンツを認識することがわかりました。私たちの発見に動機づけられて、私たちは、プロンプトポジションとレスポンスポジションの両方で2つの方向をペアにする微調整方法であるHARC(有害性と拒否のカップリング)を紹介します。介入は有害性拒否部分空間に限定されるため、残りのストリームの残りの部分はそのまま残り、一般的な能力を低下させたり、過剰な拒否を拡大したりすることはありません。広範な実験を通じて、HARC は、主要なトレーニング時間と推論時間の安全性手法にわたる 6 つのベースラインの中で最も強力な堅牢性、機能、使用性のトレードオフを達成しました。プロンプトおよびレスポンスの位置における有害性と拒否の指示は、アーキテクチャ固有の調整を行わずにテストした 5 つのモデル ファミリと 2 つのスケールに渡って伝達されます。

原文 (English)

HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment

Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbreaks succeed and informs the design of robust alignment strategies. Prior work shows that aligned LLMs encode harmfulness and refusal as separable directions in the residual stream at prompt-side token positions. We show that jailbreaks succeed at prompt encoding by suppressing either the refusal or harmfulness direction before any token is generated, with distinct attack classes occupying separable regions of the harmfulness-refusal plane. Extending the analysis to response-token positions, we find that the model recognizes harmful content while it is generating that content, even when it failed to recognize the input as harmful at the prompt side. Motivated by our findings, we introduce HARC (Harmfulness-And-Refusal Coupling), a fine-tuning method that pairs the two directions across both prompt and response positions. Since the intervention is confined to the harmfulness-refusal subspace, it leaves the rest of the residual stream intact and does not degrade general capability or inflate over-refusal. Across extensive experiments, HARC achieves the strongest robustness-capability-usability trade-off among six baselines spanning the major training-time and inference-time safety methods. The harmfulness and refusal directions at prompt and response positions transfer across the five model families and two scales we tested without architecture-specific tuning.

13:00 JSTLLM/生成AIエージェント

A-TMA: 長期エージェント メモリにおけるステートアウェア メモリ障害の分離

長期記憶により、LLM エージェントは永続的なアシスタントとして機能しますが、ユーザーの事実は変化します。有用な記憶システムは、現在何が真実で、何が以前は真実で、何が変化したかを知っていなければなりません。私たちは \emph{ゴーストメモリ} という状態調整の失敗を研究します。これは、古い事実、現在の事実、および遷移の事実がメモリバンク内に共存し、検索中に混合されたままとなり、応答モデルを誤解させるものです。私たちは、メモリ システムはバンクの維持、取得、応答時間の解決という 3 つのレベルから理解し、最適化する必要があると主張します。私たちは、既存のメモリ システム用の状態認識オーバーレイである ATMA を提案します。 ATMA は、置き換えられた記録と移行記録をバンクに保持し、クエリで要求された状態ビューの証拠パケットを構築し、現在、履歴、および移行ラベルを QA に公開します。さらに、最終的な QA の精度によってゴースト メモリが発生する箇所が隠れてしまう可能性があるため、バンク、検索、および回答レベルの失敗を分離して評価することを求めます。この障害を測定可能にするために、ゴースト メモリの競合が多いベンチマークである LTP (LoCoMo Temporal Plus) を構築し、長い会話の一般化のために LoCoMo で評価します。 LTP では、Graphiti+ATMA により、Graphiti よりも絶対値​​ 0.240 だけ競合精度が向上します。 LoCoMo では、Graphiti+ATMA により時間 F1 が 0.0295 から 0.1705 に上昇します。ゲインはホストに依存しますが、明示的な状態の役割により、最終的な QA 精度によって隠されたメモリ障害を削減できることを示しています。

原文 (English)

A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory

Long term memory lets LLM agents act as persistent assistants, but user facts change. A useful memory system must know what is true now, what used to be true, and what changed. We study \emph{ghost memory}, a state coordination failure in which old, current, and transition facts coexist in the memory bank, remain mixed during retrieval, and mislead the answer model. We argue that memory systems should be understood and optimized from three levels: bank maintenance, retrieval, and answer time resolution. We propose ATMA, a state aware overlay for existing memory systems. ATMA keeps superseded and transition records in the bank, builds evidence packets for the query's requested state view, and exposes current, historical, and transition labels to QA. We further call for decoupled evaluation of bank, retrieval, and answer level failures, since final QA accuracy can hide where ghost memory occurs. To make this failure measurable, we build LTP (LoCoMo Temporal Plus), a conflict heavy benchmark for ghost memory, and evaluate on LoCoMo for long conversation generalization. On LTP, Graphiti+ATMA improves conflict accuracy by 0.240 absolute over Graphiti. On LoCoMo, Graphiti+ATMA raises temporal F1 from 0.0295 to 0.1705. The gains are host dependent, but they indicate that explicit state roles can reduce memory failures hidden by final QA accuracy.

13:00 JSTエージェント

AgenticPD: 物理設計 QoR 最適化のためのステージ対応エージェント フレームワーク

物理設計の結果品質 (QoR) の最適化は難しく、コストがかかります。ある段階での選択は、後の段階で役立つこともあれば、悪影響を与えることもあります。各評価には、フロー全体を通してコストのかかる EDA を実行する必要があります。既存の方法では依然として最適化をフラットなパラメーター調整または LLM ベースのスクリプト生成タスクとして扱いますが、物理設計の QoR 最適化のためのステージ認識型エージェント フレームワークである AgenticPD を紹介します。 AgenticPD は、トライアルごとに完全なフローを再実行するのではなく、物理設計フローのステージ境界を中心に編成されており、ジャッジ エージェントが検索をナビゲートし、ステージ専門のエージェントがステージ ローカル ツールを使用して独自のステージ内でローカルな意思決定を行います。さらに、AgenticPD のエージェント ハーネスは、構造化された観察、実行履歴、およびエージェント コンテキスト管理を提供します。その結果、システムは前の中間状態から分岐し、チェックポイントを再利用して最適化手順を続行することができ、すべての候補がルート後のサインオフで評価されます。これらのベースライン全体で、AgenticPD はパワーとエリアでの競争力を維持しながら、強力なポストルート タイミングを実現します。

原文 (English)

AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization

Physical design quality-of-results~(QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages. Each evaluation requires a costly EDA run through the full flow. While existing methods still treat optimization as flat parameter tuning or a LLM-based script generation task, we present AgenticPD, a stage-aware agentic framework for physical design QoR optimization. Instead of re-running the full flow after every trial, AgenticPD is organized around the stage boundaries of the physical design flow, where a Judge Agent navigates the search and stage-specialized agents make local decisions within their own stage using stage-local tools. Additionally, the agent harness in AgenticPD provides structured observations, execution history, and agent context management. As a result, the system can branch from prior intermediate states and reuse checkpoints to continue the optimization procedure, and every candidate is evaluated at the post-route signoff. Across these baselines, AgenticPD achieves strong post-route timing while remaining competitive in power and area.

13:00 JSTエージェント

ArtisanCAD: 専門家に基づいた知識を抽出した産業レベルの CAD エージェント

産業用コンポーネントのコンピュータ支援設計 (CAD) には、長期的な手順モデリング、堅牢な機能の依存関係、編集可能なパラメトリック ジオメトリ、および生産グレードの B-Rep の実行が必要です。既存のテキストから CAD への手法は、自然言語記述から CAD プログラムを生成する点で有望な進歩を遂げてきましたが、ユーザーのプロンプトがあいまいな場合、指定が不十分な場合、または高レベルの設計意図しか説明していない場合には依然として困難を伴います。また、CATIA 操作記録、マクロ ログ、図面メモ、エンジニアリングの説明など、産業ワークフローで自然に利用できる専門的な手順知識を活用することはほとんどありません。 \algname は、専門家に基づいた知識を抽出した、スキルガイド付きの産業用 CAD エージェントです。 \algname の中核は CAD 中間表現 (CAD-IR) であり、パラメーター、順序付けされた操作、MCP ツール バインディング、依存関係、生成されたエンティティ、および検証ルールをエンコードする実行可能な手続き表現です。 CAD-IR は 2 つの重要な役割を果たします。1 つは、専門家の CAD プロシージャを再利用可能なパラメータ化されたスキルに蒸留するためのキャリアとして機能します。次に、曖昧なプロンプトまたは中間レベルのプロンプトを完全な実行可能な CAD 操作に変える手続き型の足場を提供します。 \algname は、専門家が導き出したスキルを取得し、CAD-IR をインスタンス化して修正し、専用の CATIA-MCP バックエンドを通じて結果の手順を実行し、マルチビューの視覚的フィードバックを使用して反復改良を行い、最終的に実稼働対応の B-Rep モデルを生成します。 Text2CAD ベンチマークでは、CAD-IR は平均面取り距離を $14.83$ から $9.88$ に削減することで中間プロンプトからの生成を改善し、あいまいなテキストの意図と実行可能な CAD 構築を橋渡しする能力を示しています。 4 つの複雑な自動車コンポーネントについて、CAD-IR を使用すると、専門家による CATIA 記録を再利用可能なスキルに蒸留でき、\algname が新しいバリアント リクエストに対して編集可能な CATIA ネイティブ B-Rep モデルを生成できるようになります。

原文 (English)

ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation

Computer-aided design (CAD) for industrial components requires long-horizon procedural modeling, robust feature dependencies, editable parametric geometry, and production-grade B-Rep execution. Existing text-to-CAD methods have made promising progress in generating CAD programs from natural-language descriptions, but they still struggle when user prompts are ambiguous, underspecified, or only describe high-level design intent. They also rarely exploit expert procedural knowledge naturally available in industrial workflows, such as CATIA operation recordings, macro logs, drawing notes, and engineering descriptions. We present ArtisanCAD, a skill-guided industrial CAD agent with expert-grounded knowledge distillation. The core of ArtisanCAD is CAD intermediate representation (CAD-IR), an executable procedural representation that encodes parameters, ordered operations, MCP tool bindings, dependencies, generated entities, and verification rules. CAD-IR plays two key roles: it first serves as the carrier for distilling expert CAD procedures into reusable parameterized skills; then it provides a procedural scaffold that turns vague or intermediate-level prompts into complete executable CAD operations. ArtisanCAD retrieves expert-derived skills, instantiates and revises CAD-IR, executes the resulting procedure through a dedicated CATIA-MCP backend, and uses multi-view visual feedback for iterative refinement, and finally generates production-ready B-Rep models. On the Text2CAD benchmark, CAD-IR improves generation from intermediate prompts by reducing mean Chamfer Distance from $14.83$ to $9.88$, showing its ability to bridge ambiguous textual intent and executable CAD construction. On four complex automotive components, CAD-IR enables expert CATIA recordings to be distilled into reusable skills, allowing ArtisanCAD to generate editable CATIA-native B-Rep models for new variant requests.

13:00 JSTLLM/生成AIエージェント研究/論文

Danus: ファクトグラフ メモリを使用して数学的推論エージェントを調整する

最近の LLM ベースの数学的推論エージェントは研究レベルの問題に取り組み始めており、いくつかのケースでは未解決の問題の解決に貢献しています。ただし、中間クレームの整理と信頼性を維持しながら、並行した証拠検索を調整することが難しいため、このようなエージェントを効果的に拡張および調整することは依然として困難です。この論文では、グローバル メモリ管理メカニズムとして共有ファクト グラフを中心とした研究レベルの数学的推論のためのオーケストレーション システムである Danus を提案します。 Danus は、計画と調整を実行するメイン エージェント、証明検索を並行して実行する複数のワーカー エージェント、提案された数学的主張をファクト グラフに追加する前にチェックするステートレス ベリファイアーで構成されます。検証された各事実は、その証明および論理的依存関係とともに保存されるため、システムは共有された証明の状態を整理しながら、長い引数を段階的に構築できます。メイン エージェントは、進化する証明状態を定期的に要約し、有望な方向にワーカーをリダイレクトし、進捗レポートを通じて人間の数学者との対話をサポートします。私たちは、代数幾何学、特異点理論、および組合せ論における 6 つの研究レベルのケーススタディを通じてダヌスを評価し、ファクトグラフ記憶メカニズムによってダヌスがどのように長く詳細な数学的証明を構築できるかを示します。私たちの結果は、ファクトグラフベースのオーケストレーションが、長期的な研究課題に対して数理推論エージェントを拡張するための効果的な手段を提供することを示唆しています。 Danus は、https://github.com/frenzymath/Danus でオープンソースです。

原文 (English)

Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory

Recent LLM-based mathematical reasoning agents have begun to tackle research-level problems and, in several cases, have contributed to the resolution of open problems. However, scaling and orchestrating such agents effectively remains challenging, due to the difficulty of coordinating parallel proof search while keeping intermediate claims organized and reliable. In this paper, we propose Danus, an orchestration system for research-level mathematical reasoning centered on a shared fact graph as a global memory-management mechanism. Danus consists of a main agent that performs planning and coordination, multiple worker agents that carry out proof search in parallel, and a stateless verifier that checks proposed mathematical claims before they are admitted into the fact graph. Each verified fact is stored together with its proof and logical dependencies, allowing the system to build long arguments incrementally while keeping the shared proof state organized. The main agent periodically summarizes the evolving proof state, redirects workers across promising directions, and supports interaction with human mathematicians through progress reports. We evaluate Danus through six research-level case studies in algebraic geometry, singularity theory, and combinatorics, illustrating how the fact-graph memory mechanism enables Danus to construct long, detailed mathematical proofs. Our results suggest that fact-graph-based orchestration provides an effective route toward scaling mathematical reasoning agents for long-horizon research problems. Danus is open source at https://github.com/frenzymath/Danus.

13:00 JST研究/論文

Faster and Simpler Greedy Algorithm for $k$-Median and $k$-Means

Clustering problems such as $k$-means and $k$-median are staples of unsupervised learning, and many algorithmic techniques have been develo…

13:00 JST画像/動画生成

ContrastiveCFG: Guiding Diffusion Sampling by Contrasting Positive and Negative Concepts

As Classifier-Free Guidance (CFG) has proven effective in conditional diffusion model sampling for improved condition alignment, many appli…

13:00 JST研究/論文

The Minimal Search Space for Conditional Causal Bandits

Causal knowledge can be used to support decision-making problems. This has been recognized in the causal bandits literature, where a causal…

13:00 JST研究/論文

Silent Neuron Theory and Plasticity Preservation for Deep Reinforcement Learning in Adaptive Video Streaming

Adaptive video streaming optimizes Quality of Experience (QoE) metrics by selecting appropriate bitrates according to varying network bandw…

13:00 JST画像/動画生成ロボティクス

VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting

Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural lang…

13:00 JST研究/論文

Understanding Two-Layer Neural Networks with Smooth Activation Functions

This paper aims to understand the training solution, which is obtained by the back-propagation algorithm, of two-layer neural networks whos…

13:00 JST研究/論文

L-GTA: Latent Generative Modeling for Time Series Augmentation

Data augmentation is becoming increasingly important across various areas of time series analysis, including forecasting, classification, a…

13:00 JST画像/動画生成

A Study of Commonsense Reasoning over Visual Object Properties

Inspired by human categorization, visual reasoning about object properties, such as physical attributes and functions, involves identifying…

13:00 JSTロボティクス

LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes

Physics-based human motion control can make a simulated character walk, sit, and manipulate objects with high physical realism. Almost alwa…

13:00 JST画像/動画生成研究/論文

Explain Before You Answer: A Survey on Compositional Visual Reasoning

Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like abilit…

13:00 JSTLLM/生成AIハードウェア/半導体

NonTextual Target Attack

Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output w…

13:00 JSTロボティクス

Rapidly Learning Soft Robot Control via Implicit Time-Stepping

With the explosive growth of rigid-body simulators, policy learning in simulation has become the de facto standard for most rigid morpholog…

13:00 JSTLLM/生成AI

Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning

We propose RT (Refine Thought), a method that can enhance the semantic reasoning ability of text embedding models. The method obtains the f…

13:00 JST画像/動画生成ハードウェア/半導体

T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs

Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training. Rece…

13:00 JSTLLM/生成AI画像/動画生成

Thinking Ahead: Foresight Intelligence in MLLMs and World Model

In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applica…

13:00 JST研究/論文

FDRMFL: Multimodal Federated Feature Extraction Model Based on Information Maximization and Contrastive Learning

We propose FDRMFL, a task-driven multimodal feature extraction framework for federated regression under non-IID data distributions. Extract…

13:00 JSTロボティクス

HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies

Generalist vision--language--action (VLA) policies are typically trained on heterogeneous mixtures of robot demonstrations spanning diverse…

13:00 JSTエージェント

HiDVFS: Hierarchical Multi-Agent DVFS for Real-Time OpenMP DAG Workloads

Leakage power in multicore embedded systems now rivals dynamic power, so DVFS schedulers must respect deadlines and thermal limits, not jus…

13:00 JST研究/論文

ButterflyMoE: Compression-Scalable Ternary Experts via Structured Butterfly Orbits

In current Mixture of Experts (MoE) architectures, linear memory scaling is present, the memory grows as the number of experts increases. $…

13:00 JST画像/動画生成

Spatiotemporal Semantic V2X Framework for Cooperative Collision Prediction

Intelligent Transportation Systems (ITS) demand real-time collision prediction to ensure road safety and reduce accident severity. Conventi…

13:00 JSTロボティクス

Can We Really Learn One Representation to Optimize All Rewards?

As unsupervised pretraining becomes increasingly ubiquitous in reinforcement learning, a more thorough theoretical understanding of these m…

13:00 JSTLLM/生成AI

Named-Entity Recognition in the Crime Domain (CrimeNER): Case Study and Dataset

The extraction of critical information from crime-related documents is a crucial task for law enforcement agencies. The extraction of this…

13:00 JSTLLM/生成AI画像/動画生成

DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression

Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make…

13:00 JST画像/動画生成

CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation

Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that…

13:00 JSTLLM/生成AIエージェント

Effective Strategies for Asynchronous Software Engineering Agents

AI agents have become increasingly capable at isolated software engineering (SWE) tasks such as resolving issues on Github. Yet long-horizo…

13:00 JSTLLM/生成AIロボティクス

Object Search in Partially-Known Environments via LLM-informed Model-based Planning and Prompt Selection

We present a novel LLM-informed model-based planning framework, and a novel prompt selection method, for object search in partially-known e…

13:00 JST画像/動画生成GeminiGemmaLlamaQwen

From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics

Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition…

13:00 JST研究/論文

Exploration of Fast-Slow Latent Recurrence for Train-Short, Test-Long Generalization

We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longe…

13:00 JSTLLM/生成AIエージェント

Diversity Without Fidelity: A Solver-Sampler Mismatch in Multi-Agent LLM Negotiation Simulation

Language models are increasingly used to simulate people: survey respondents, negotiators, stakeholders in policy exercises. In that role a…

13:00 JSTLLM/生成AIエージェントClaude

AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection

While recent LLM-based agents can identify many candidate bugs in source code, their reports remain static hypotheses that require manual v…

13:00 JSTLLM/生成AI

Learning from Execution: Self-Evolving Memory for Private-Library Code Generation

Large Language Models (LLMs) have achieved strong performance on general code generation, but their effectiveness drops sharply in enterpri…

13:00 JST研究/論文

Health System Scale Semantic Search Across Unstructured Clinical Notes

Introduction: Semantic search, which retrieves documents based on conceptual similarity rather than keywords, offers advantages for retriev…

13:00 JST研究/論文

From Beats to Breaches:How Offensive AI Infers Sensitive User Information from Playlists

The pervasive integration of AI has enabled Offensive AI: the exploitation of AI for malicious ends across the cyber-kill chain. A critical…

13:00 JST研究/論文

Optimal FALQON for Quantum Approximate Optimization via Layer-wise Parameter Tuning

Feedback-based adaptive quantum optimization (FALQON) is a promising approach for solving combinatorial problems on noisy intermediate-scal…

13:00 JSTLLM/生成AI研究/論文

Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval

Current LLM memory benchmarks evaluate answer quality rather than retrieval accuracy. Consequently, a system that dumps its entire belief s…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges

Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating…

13:00 JST研究/論文

CogAdapt: リード適応による臨床 ECG 基礎モデルのウェアラブル認知負荷評価への移行

リアルタイムの認知負荷評価は、人間とコンピューターの適応的なインタラクションに不可欠ですが、ラベル付けされたデータが限られており、被験者間の汎化が不十分であるため、依然として困難です。数百万の臨床記録で事前トレーニングされた最近の ECG 基礎モデルは豊富な表現を提供しますが、センサー構成の不一致やタスクの違いのため、ウェアラブル デバイスに直接適用することはできません。この論文では、臨床 ECG 基礎モデルをウェアラブル認知負荷評価に適応させるフレームワークである CogAdapt を提案します。 CogAdapt は、3 リードのウェアラブル信号を解剖学的に一貫した 12 リード表現に変換する学習可能なアダプターである LeadBridge と、壊滅的な忘却を防ぎながらエンコーダ層のフリーズを徐々に解除するプログレッシブ微調整戦略である ProFine を導入しています。 2 つの公開データセット (CLARE および CL-Drive) を 1 被験者除外相互検証で評価したところ、CogAdapt はゼロからトレーニングされたベースラインを大幅に上回り、マクロ F1 スコア 0.626 および 0.768 を達成したことが示されています。これらの結果は、ウェアラブル センサーによる被験者に依存しない認知負荷評価に対する基礎モデルの適応が期待できることを示しています。

原文 (English)

CogAdapt: Adapting Clinical ECG Foundation Models for Wearable Cognitive Load Assessment

Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labeled data are scarce and models generalize poorly across subjects. Recent ECG foundation models, pre-trained on millions of clinical diagnostic ECG recordings, yet they do not apply directly to wearable devices when the sensor configuration and the task both differ. We present CogAdapt, a framework that adapts a clinical ECG foundation model to wearable cognitive load assessment. CogAdapt has two parts. LeadBridge is a learnable adapter that maps 3-lead wearable signals to a 12-lead-compatible representation. ProFine is a progressive fine-tuning strategy that unfreezes encoder layers in stages while limiting representational drift in the pre-trained model. On two public datasets (CLARE and CL-Drive) under leave-one-subject-out cross-validation, CogAdapt reaches macro-F1 of 0.626 and 0.768, improving over from-scratch baselines by 11.2 and 16.1 percentage points. The results show that a clinical ECG pretraining can support subject-independent cognitive load assessment from wearable sensors.

13:00 JSTLLM/生成AI

より広い網を投じる: コード推論のための調整された Pass@K ポリシーの最適化

ベリファイアーを使用してサンプリングを繰り返すことは、標準メトリックとして pass@$K$ を使用して、コード生成のためのテスト時の計算を割り当てる標準的な方法です。しかし、標準のポリシー クラスは単一の回答分布から $K$ の独立したサンプルを抽出するため、試行はしばしば重複に近い推論パスに陥り、冗長な展開で予算を無駄にします。競技プログラミングでは、この失敗はコストが高くつきます。競技プログラミングでは、多くの問題で複数の異なるアルゴリズム戦略が許容され、 pass@$K$ に必要な正しい試行は 1 回だけです。私たちは、pass@$K$ の生成を戦略の共同探索に変える、調整された Pass@$K$ ポリシー最適化 (CPPO) を提案します。プランナーは $K{=}4$ の代替高レベルメソッドのタプルを発行し、共有ソルバーはメソッドごとに 1 つの解決策を試みます。 CPPO は、乗算プランナーの報酬 $R_{\mathrm{plan}} = J_\psi \cdot R_{\mathrm{out}}$ を使用してこの共同ポリシーをトレーニングし、検証者が確認した pass@$K$ の成功につながる有効な戦略タプルにのみクレジットを割り当てます。 APPS、CodeContests、LiveCodeBench-v6 全体で、CPPO は直接サンプリング、計画ベースライン、プランナーのみの SFT よりも pass@$4$ を改善し、同じ $K{=}4$ ソルバー試行予算の下で pass@$K$ 指向の RL を改善し、9 つのモデル (ベンチマーク セル) のうち 6 つで統計的に有意な向上をもたらしました。単一の最大の利益は、最も強いベースラインである PKPO ($0.588 \rightarrow 0.748$; ペア ブートストラップ、$p < 0.05$) に対する Qwen3.5-9B LiveCodeBench-v6 の $+0.16$ です。

原文 (English)

Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning

Repeated sampling with a verifier is the standard way to allocate test-time compute for code generation, with pass@$K$ as the canonical metric. Yet the standard policy class draws $K$ independent samples from a single answer distribution, so attempts often collapse onto near-duplicate reasoning paths and waste the budget on redundant rollouts. This failure is costly in competitive programming, where many problems admit multiple distinct algorithmic strategies and pass@$K$ requires only one correct attempt. We propose Coordinated Pass@$K$ Policy Optimization (CPPO), which turns pass@$K$ generation into joint exploration over strategies: a planner emits a tuple of $K{=}4$ alternative high-level methods, and a shared solver attempts one solution per method. CPPO trains this joint policy with a multiplicative planner reward, $R_{\mathrm{plan}} = J_\psi \cdot R_{\mathrm{out}}$, assigning credit only to valid strategy tuples that lead to verifier-confirmed pass@$K$ success. Across APPS, CodeContests, and LiveCodeBench-v6, CPPO improves pass@$4$ over direct sampling, planning baselines, planner-only SFT, and pass@$K$-oriented RL under the same $K{=}4$ solver-attempt budget, with statistically significant gains on six of nine model--benchmark cells. The largest single gain is $+0.16$ on Qwen3.5-9B LiveCodeBench-v6 over the strongest baseline, PKPO ($0.588 \rightarrow 0.748$; paired bootstrap, $p < 0.05$).

13:00 JST研究/論文

介入の大規模シミュレーションを使用した情報 AI 政策評価

AI システムとその害悪の急速な普及により、世界中で AI ガバナンスの取り組みが加速する中、政策立案者や研究者にとって、競合する政策オプションの間で優先順位を付けることがますます困難になっています。私たちは、特定の AI の害を軽減するための実行可能な政策オプションを特定するための方法論を導入し、政策立案者や研究者がより多くの時間とリソースの投資が必要な分野をターゲットにできるように支援します。この方法では、政策の参加型評価、導入コストの専門家による評価、および各政策オプションの下で認識される害の軽減に関する LLM ベースの評価が組み合わされます。私たちは、遺伝的アルゴリズムに基づくシミュレーション研究を活用して、潜在的な政策の組み合わせの広大な解決空間を探索し、コスト、参加型インプット、被害軽減のさまざまな重み付けの下で結果がどのように変化するかを調査します。この方法により、参加型コンポーネントと専門家コンポーネントの間のさまざまなバランスを探ることが可能になり、政策立案者や研究者がそれぞれにどれだけの重みを割り当てるかを評価できることがわかりました。私たちは、遺伝的アルゴリズムによって発見された実行可能な政策の組み合わせの多様性は、検討の出発点として有用である可能性があると主張します。この手法は、参加型 AI に関する既存の作業を実際の政策開発パイプラインに直接統合することで運用可能にします。

原文 (English)

Informing AI Policy Assessment using Large-Scale Simulation of Interventions

As the rapid proliferation of AI systems and harms spurs efforts in AI governance around the world, prioritizing among competing policy options has become increasingly challenging for policymakers and researchers. We introduce a methodology for identifying viable policy options to mitigate specified AI harms, helping policymakers and researchers target areas that warrant greater time and resource investment. This method combines participatory evaluation of policies, expert assessment of implementation costs, and an LLM-based assessment of perceived harm mitigation under each policy option. We leverage a genetic algorithm-based simulation study to explore a vast solution space of potential policy combinations, and examine how outcomes change under different weightings of cost, participatory input, and harm mitigation. We find that this method enables exploration of different balances between participatory and expert components, allowing policymakers and researchers to assess how much weight to assign to each. We argue that the diversity of viable policy combinations found by the genetic algorithm could be a useful starting point for deliberation. This method operationalizes existing work on participatory AI by integrating it directly into practical policy development pipelines.

13:00 JSTエージェント

Trading Human Curation for Synthetic Augmentation in RLVR

The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic lang…

13:00 JSTLLM/生成AIビジネス/資金調達AnthropicClaudeOpenAIGPT / ChatGPTQwen

Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation

Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions. Whether they evalua…

13:00 JSTLLM/生成AI

TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation

TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequ…

13:00 JST研究/論文

MetaConfigurator: AI-Assisted RDF Authoring from JSON Data

Scientific workflows increasingly generate structured JSON data that is easy to exchange but difficult to interpret consistently across sys…

13:00 JST研究/論文

The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions

The standard basis of transformer hidden states is a training-free, architecture-general feature basis for detecting concepts and, in langu…

13:00 JST研究/論文

Feynman Kac Reweighted Schr\"odinger Bridge Matching for Surface-Based Tau PET Harmonization

Tau positron emission tomography (PET) is widely used for the in vivo characterization of disease stage and progression in Alzheimer's dise…

13:00 JSTエージェントロボティクス

NeuralMUSIC: ロボット音源位置特定のためのハイブリッド神経部分空間フレームワーク

信頼性の高い音源定位はロボットの聴覚の基礎であり、自律ロボットが空間的な手がかりを認識し、動的な環境で効果的に動作できるようになります。多重信号分類 (MUSIC) などの古典的な手法は強力な理論的基盤を提供しますが、信号対雑音比が低いと性能が低下します。深層学習ベースのアプローチは有望なパフォーマンスを達成しますが、多くの場合、条件全体にわたる限られた一般化に苦労します。これらの課題に対処するために、ロボットによる音源定位のためのハイブリッド神経部分空間フレームワークである NeuralMUSIC を提案します。具体的には、ニューラル ネットワークはまず、マルチチャネル マイクの観測値から空間共分散行列を推定します。予測された共分散は、固有値分解 (EVD) と擬似スペクトル計算を使用して古典的な MUSIC パイプラインに統合され、その後、周波数アテンション フュージョン (FAF) モジュールによって最終的な DOA 推定値が生成されます。データ効率を向上させるために、ラベルなしの音響データを活用して空間構造を捕捉する自己教師付き空間相関学習 (SSCL) 戦略をさらに導入します。さまざまなロボット タスクにわたる広範な実験により、NeuralMUSIC が堅牢性とクロスドメイン汎用性の向上を示しながら、競争力のある位置特定精度を達成できることが実証されました。

原文 (English)

NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization

Reliable sound source localization is fundamental to robot audition, enabling autonomous robots to perceive spatial cues and operate effectively in dynamic environments. Classical methods such as Multiple Signal Classification (MUSIC) offer strong theoretical foundations but degrade under low signal-to-noise ratios. While deep learning-based approaches achieve promising performance, they often struggle with limited generalization across conditions. To address these challenges, we propose NeuralMUSIC, a hybrid neural-subspace framework for robotic sound source localization. Specifically, a neural network first estimates the spatial covariance matrix from multichannel microphone observations. The predicted covariance is then integrated into a classical MUSIC pipeline with eigenvalue decomposition (EVD) and pseudo-spectrum computation, followed by a Frequency Attention Fusion (FAF) module to produce the final DOA estimates. To improve data efficiency, we further introduce a Self-supervised Spatial Correlation Learning (SSCL) strategy that leverages unlabeled acoustic data to capture spatial structure. Extensive experiments across different robotic tasks demonstrate that NeuralMUSIC achieves competitive localization accuracy while exhibiting improved robustness and cross-domain generalization.

13:00 JSTLLM/生成AI

Where Did the Variability Go? From Vibe Coding to Product Lines by Regeneration

In vibe coding, an emerging AI-driven paradigm, an LLM generates an entire program from a natural language prompt, but what happens to the…

13:00 JST画像/動画生成

Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking

The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion es…

13:00 JSTビジネス/資金調達

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.…

13:00 JSTエージェント

Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem

As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent ass…

13:00 JST画像/動画生成NVIDIA

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators

Text-to-image (T2I) diffusion models typically require substantial computational resources and cloud infrastructure, posing significant cha…

13:00 JSTLLM/生成AI画像/動画生成GPT / ChatGPT

MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering

As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for use…

13:00 JST研究/論文

Cross-Receiver Open-Set Radio Frequency Fingerprinting via Structure-First Adaptation

Radio frequency fingerprint identification (RFFI) provides a physical-layer credential for Internet of Things devices, but open-set decisio…

13:00 JST画像/動画生成

Phase-Preserving Trimodal Transformer for Tropical Forest Biomass Estimation Using Optical and PolInSAR Data

The accurate estimation of Above-Ground Biomass (AGB) in mature tropical forests remains a critical challenge in remote sensing, primarily…

13:00 JST画像/動画生成ロボティクスビジネス/資金調達研究/論文

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their cap…

13:00 JST画像/動画生成

Wan-Streamer v0.2: Higher Resolution, Same Latency

We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps t…

13:00 JST画像/動画生成エージェント

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deep…

13:00 JSTLLM/生成AI

Weak-to-Strong Generalization via Direct On-Policy Distillation

Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to r…

13:00 JST研究/論文

無線ネットワークにおけるチャネル状態フィードバックを強化するための圧縮を使用した対照予測コーディング

正確かつタイムリーなチャネル状態情報 (CSI) は次世代無線システムにとって不可欠ですが、既存の研究では、学術界と現在の 3GPP 研究の両方において、CSI 圧縮と CSI 予測が別個の問題として扱われています。その結果、標準化された CSI フィードバック パイプライン内では、チャネルの経年劣化への対処が不十分なままです。この記事では、対照予測コーディング (CPC) を 3GPP 準拠の CSI 圧縮アーキテクチャに直接統合する、統合された圧縮予測フレームワークを提案します。高次元の CSI 行列を予測する代わりに、私たちのアプローチは将来の潜在表現を予測し、1-SGCS と InfoNCE の目的を組み合わせて再構成の忠実度と時間的予測コヒーレンスを共同で最適化します。この設計により、フィードバックのオーバーヘッドを増加させることなく、時間表現の学習が可能になります。量子化前にエンコードされた特徴に対して自己回帰モデリングを実行する CPC-before-Compression と、時間モデリングを基地局に移してユーザーのデバイスの複雑さを軽減する CPC-after-Compression の 2 つのバリエーションを紹介します。 Nokia、Oppo、および CATT の 3GPP 準拠のデータセットの評価では、圧縮前の CPC は 3GPP ベースラインよりも 32 倍低いデコーダ GFLOP で 90% 以上の再構築精度を達成する一方、圧縮後の CPC は同一のエンコーダ フットプリントと同じ 64 ビット フィードバック オーバーヘッドを維持することが示されています。提案されたフレームワークは、標準化されたパイプライン内で圧縮と予測を統合することにより、年齢を考慮した計算効率の高い CSI フィードバック ソリューションを提供します。ソース コードは https://github.com/AhmedRadwan02/cpc-3gpp で公開されています。

原文 (English)

Contrastive Predictive Coding with Compression for Enhanced Channel State Feedback in Wireless Networks

Accurate and timely channel state information (CSI) is essential for next-generation wireless systems, yet existing works treat CSI compression and CSI prediction as separate problems, both in academia and in current 3GPP studies. Consequently, channel aging remains insufficiently addressed within standardized CSI feedback pipelines. In this article, we propose a unified compression-prediction framework that integrates Contrastive Predictive Coding (CPC) directly into the 3GPP-compliant CSI compression architecture. Instead of predicting high-dimensional CSI matrices, our approach forecasts future latent representations and jointly optimizes reconstruction fidelity and temporal predictive coherence via a combined 1-SGCS and InfoNCE objective. This design enables temporal representation learning without increasing feedback overhead. We present two variants: CPC-before-Compression, which performs autoregressive modeling on encoded features prior to quantization, and CPC-after-Compression, which shifts temporal modeling to the base-station to reduce the complexity of users' devices. Evaluations on 3GPP-compliant datasets from Nokia, Oppo, and CATT show that CPC-before-Compression achieves over 90% reconstruction accuracy with 32x lower decoder GFLOPs than the 3GPP baseline, while CPC-after-Compression preserves an identical encoder footprint and the same 64-bit feedback overhead. By unifying compression and prediction within a standardized pipeline, the proposed framework provides an age-aware, computationally efficient CSI feedback solution. The source code is publicly available at: https://github.com/AhmedRadwan02/cpc-3gpp

13:00 JST研究/論文

Tangent classes of matroids and wonderful compactifications

For every loopless matroid $M$ and every Feichtner--Yuzvinsky building set $\mathcal{G}$ containing the top flat, we construct an integral…

13:00 JSTLLM/生成AI研究/論文NVIDIADeepSeek

Think Before You Grid-Search: Floor-First Triage for LLM Serving

LLM serving optimization typically benchmarks many configurations and reaches for heavy profilers when latency targets are missed. We argue…

13:00 JSTハードウェア/半導体NVIDIA

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods

The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatr…