Skip to the content.

AIニュース 2026-08-26

自動生成: 2026-08-26 10:41 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. The full stack behind abundant intelligenceOpenAI

    OpenAI CFO Sarah Friar explains how advances across chips, compute, m…

  2. AIエージェントがCAEを動かすとどうなる? 人に残る役割は?ITmedia AI+

    AIエージェントがCAEツールを操作し、解析作業を担うようになれば、人の役割はどう変わるのか。「CAEユニバーシティ 特別公開フォーラム…

  3. 「Jetson Orin Nano 2」は推論性能が2倍、エッジ生成AIもリアルタイム処理可能ITmedia AI+

    NVIDIAは、組み込み機器向けAIモジュール「Jetson Orin」の新たなラインアップとなる「Jetson Orin Nano 2」…

  4. Claude Cowork finally remembers what you told the app in chatTechCrunch AI

    Anthropic is giving Claude a shared memory across chat and Cowork, so…

  5. 中国製AIが上位を占有、米国製AIに3倍差 Mozillaレポートが明かす「オープンモデル」の現在地ITmedia AI+

    Mozillaは、オープンウェイトのAIモデルの利用動向など現状をまとめたレポート「The State of Open Source AI…

  6. Robotics startup Generalist reaches $3B valuation, sources sayTechCrunch AI

    The $200 million extension comes just months after the physical AI st…

  7. OpenAI loses a top data center exec, as stream of high-profile departures continuesTechCrunch AI

    Before Malone left, OpenAI had already reshuffled its infrastructure…

トピック別件数

日本語メディア13件

ITmedia AI+ (日本語)

07:00 JSTエージェント

AIエージェントがCAEを動かすとどうなる? 人に残る役割は?

AIエージェントがCAEツールを操作し、解析作業を担うようになれば、人の役割はどう変わるのか。「CAEユニバーシティ 特別公開フォーラム 2026」で行われたサイバネットシステムの講演では、AIエージェントによるCAE解析の実験を通じて、AIに任せられることと人に残る仕事、CA…

06:15 JSTLLM/生成AIハードウェア/半導体NVIDIA

「Jetson Orin Nano 2」は推論性能が2倍、エッジ生成AIもリアルタイム処理可能

NVIDIAは、組み込み機器向けAIモジュール「Jetson Orin」の新たなラインアップとなる「Jetson Orin Nano 2」を発表した。現行の「Jetson Orin Nano」のメモリ容量8GBモデルと比べて推論性能が2倍に向上した一方で、同じ推論性能であれば消…

05:00 JSTその他

あなたが使うAIの「中身」、実は中国モデルかも 米企業のAPI利用で58%

Arena AIのコーディング性能ランキングでは、上位20モデルのうち8つを中国発が占める。AIモデルの中継サービスOpenRouterでも、米企業が中国モデルに流すトークンの比率は58%に達した。利用者からは見えないAIサービスの「中身」で、いま何が起きているのかを追う。

19:42 JSTロボティクス

「今からロボットを現場投入できないか?」 増えた日本企業からの問い合わせ “頭脳”を作る韓国企業が感じた国内の変化

日本にも拠点を置くRLWRLD(リアルワールド)は、フィジカルAI分野で注目を集める国際企業だ。産業用ロボットの導入もかなり進んでいる日本市場で最近起こった変化について聞いた。

19:28 JSTLLM/生成AI規制/政策

政府、AI事業者向け「知財保護ルール」策定 対象範囲・運用方針は?

政府は、生成AI事業者向けに透明性の確保と知的財産権の保護を促す「プリンシプル・コード」を策定した。

16:50 JSTその他

話題のステルスモデル「Ox Alpha」はナニモノか うわさをまとめてみた

開発元を伏せたまま米OpenRouterで無料公開されたAIモデル「Ox Alpha」が話題だ。一部の開発ツールでは利用シェア1位に。確認できる事実と、中国Z.aiの「GLM」系とみる有力なうわさを整理した。

15:56 JSTロボティクス

シャープ、対話型AIロボ「ポケとも」に新モデル あえて“全肯定しないキャラ”に 「暇だなー」と話しかけると……

シャープは対話型のAIロボット「ポケとも」の新モデルを発表した。従来のモデルとは性格を変え、あえてユーザーを“全肯定しないキャラクター”にした。

14:35 JSTその他

社内ITの不備で1人年18営業日の時間ロス、100人企業で約4400万円の損失換算に

ある調査によると、社内IT環境の不備によって従業員1人あたり年147時間が失われるという。しかも、この損失はほとんど可視化されておらず、運用課題のずれや、採用・離職へ影響している可能性もある。

14:11 JSTLLM/生成AIエージェントGPT / ChatGPT

「ChatGPT Plus」で5時間制限復活、Work/Codex対象 Proは「数カ月は対象外」

「5時間制限により計算資源の負荷を平準化でき、週単位の利用枠を手厚く維持できる」

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

Linux版の「ChatGPTデスクトップアプリ」ついに登場 何ができて、何ができない?

「ChatGPT」に加えて「ChatGPT Work」「Codex」も使える、Linux版のデスクトップアプリケーションのプレビュー版が登場した。できることと、できないことを整理しよう。

13:00 JSTLLM/生成AI

中国製AIが上位を占有、米国製AIに3倍差 Mozillaレポートが明かす「オープンモデル」の現在地

Mozillaは、オープンウェイトのAIモデルの利用動向など現状をまとめたレポート「The State of Open Source AI」を公開した。同社と調査会社SlashDataによる調査、「Chatbot Arena」やOpenRouterの公開データに基づいている。

12:28 JSTロボティクス

人型ロボットが本社周辺をうろうろ 米Figureの目撃動画が話題、CEO「キャンパスは“SF映画”」

米Figure AIの人型ロボットが本社周辺の屋外を歩き回る様子を捉えた動画がXで話題だ。同社のブレット・アドコックCEOは「FigureのキャンパスはSF映画だ」とコメントした。

12:10 JSTその他

千葉豪雨、なぜ当日昼前まで予測できなかったのか ウェザーニューズに聞く「天気予報のメカニズム」のイマ

8月13日午後から夜にかけて千葉県を襲った記録的な大雨「令和8年8月千葉豪雨」。線状降水帯が繰り返し発生し、千葉市では観測史上1位となる雨量を記録、死者・行方不明者が出る大きな災害となった。現代の天気予報は、なぜこれほどの豪雨を直前まで捉えられなかったのか。ウェザーニューズ(千…

海外メディア7件

TechCrunch AI (英語)

09:40 JSTロボティクスビジネス/資金調達

Robotics startup Generalist reaches $3B valuation, sources say

The $200 million extension comes just months after the physical AI startup reached a $2 billion valuation.

09:06 JSTLLM/生成AIOpenAI

OpenAI loses a top data center exec, as stream of high-profile departures continues

Before Malone left, OpenAI had already reshuffled its infrastructure org, shifting his reporting line away from President Greg Brockman and…

04:03 JST画像/動画生成ビジネス/資金調達

Stability AI, maker of image generator Stable Diffusion, raises $76 million in fresh funding

The company's new fundraising total now stands at $232 million.

02:50 JSTLLM/生成AIAnthropicClaude

Claude Cowork finally remembers what you told the app in chat

Anthropic is giving Claude a shared memory across chat and Cowork, so users no longer have to repeatedly brief the AI on projects, preferen…

23:22 JSTLLM/生成AIハードウェア/半導体研究/論文OpenAI2媒体が報道

OpenAI’s Jalapeño chip is built for fast inference at scale, benchmarks show

Tested on SemiAnalysis’ InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the curre…

出典:OpenAITechCrunch AI
22:00 JSTエージェント研究/論文2件の関連記事

Accel-backed Keenable is indexing the web for AI agents

Now exiting stealth mode with a $26 million seed round, Keenable has been building a vast web search index for AI agents.

出典:TechCrunch AITechCrunch AI
21:00 JSTLLM/生成AIエージェントOpenAI

‘The world seems to be ready’: An interview with OpenAI head of product Thibault Sottiaux

TechCrunch talks agents, UX, and reporting to Greg Brockman with OpenAI's head of product.

公式ブログ1件

OpenAI (英語)

16:05 JSTLLM/生成AIハードウェア/半導体OpenAI

The full stack behind abundant intelligence

OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at gr…

論文585件

arXiv cs.AI (英語)

13:00 JSTLLM/生成AIQwen

KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference

Transformer-based large language models (LLMs) incur high prefill latency because key-value (KV) tensors must be recomputed for each reques…

13:00 JST研究/論文

AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance

A protocol is presented for recording the governance decisions of automated AI runtimes. When a runtime releases, blocks, defers, redacts,…

13:00 JST研究/論文

Reviewing Model Collapse and Countermeasures

Driven by massive amounts of web-scale data, generative AI (GenAI) has achieved remarkable progress, enabling various applications in diver…

13:00 JSTエージェント

AI Learning and Conceptual Transfer in the Game of Hidden Rules

This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hi…

13:00 JSTエージェント

LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review Platform

Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult becaus…

13:00 JSTエージェント

SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAG

Heterogeneous agentic retrieval-augmented generation (RAG) systems increasingly orchestrate external APIs, internal databases, vector store…

13:00 JST研究/論文

RIACT: A Responsible AI System for Personalized Study Habit Tracking and Early Burnout Signal Detection in University Students

Student burnout is highly prevalent in higher education, with reported rates ranging from 12% to over 70% and consistently exceeding those…

13:00 JSTLLM/生成AI研究/論文

There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items

Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the pro…

13:00 JST研究/論文

Spyre-Accelerated Retrieval-Augmented Generation on IBM LinuxONE: A Cloud-Native Architecture for Secure, High-Throughput Enterprise AI Inference

Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, th…

13:00 JSTLLM/生成AI

Hate Speech Classification In Roman Urdu: A Comparative Study On Parameter Efficient Fine-Tuning And Prompt Engineering

Due to the widespread accessibility of the internet and social media, toxic and hateful con-tent has grown exponentially, causing significa…

13:00 JST研究/論文

The Abstention Protocol: RCA for Clos Fabrics

Root cause analysis (RCA) in large datacenter networks is challenging because telemetry is noisy, partial, and asynchronous. Score-based ap…

13:00 JSTロボティクス研究/論文

Retrieval-grounded robot program generation and simulation-based correction via Model Context Protocol

Flexible manufacturing requires industrial robots to be reprogrammed rapidly as product variants change. This paper presents a language-mod…

13:00 JST研究/論文

Composable Trust Infrastructure for Manufacturing Knowledge Graphs: Cross-System Provenance, Temporal Reasoning, and Decision Traceability

Manufacturing knowledge graphs that integrate data from heterogeneous industrial systems face a trust deficit: consumers cannot determine w…

13:00 JSTLLM/生成AI画像/動画生成

Evaluating Multimodal Narrative Understanding of Popular Hollywood Films

Multimodal language models increasingly show promise for enabling the large-scale computational analysis of film, opening up new avenues fo…

13:00 JSTエージェントロボティクス

Agentic AI for Safety-critical Multi-drone Systems: Challenges and Opportunities

Multi-drone systems are increasingly positioned for safety-critical missions such as search and rescue (SAR) and critical infrastructure mo…

13:00 JST研究/論文

Software Frameworks for Explainable AI in Time Series Classification: A Systematic Review

Time series arise in a wide range of application domains and are analyzed using machine learning in decision-critical settings. Time series…

13:00 JST研究/論文

Enhanced Artificial Neural Networks Using QHAdamW in Air Quality Forecasting

The study employed an Artificial Neural Network in combination with the optimized Adaptive Moment Estimation (Adam) algorithm, currently th…

13:00 JSTLLM/生成AILlama

Let Credit Follow Computation: Architecture-Aware Credit Transport for Large Language Model Reinforcement Learning

Credit assignment in large-language-model reinforcement learning (LLM RL) can be separated into three objects: evidence about success, a tr…

13:00 JST研究/論文

Quantifying geographic domain shift to decouple the geospatial transferability of human mobility flow generation models

Human mobility serves as an essential proxy for understanding social, economic, and environmental dynamics in urban systems. Geospatial tra…

13:00 JST研究/論文

A Reproducible, License-Aware Distillation Recipe for CPUDeployable Safety Classification

Deploying a safety layer for large language models on commodity hardware is constrained by the guards available to do it: current open guar…

13:00 JST研究/論文

Robust Lightweight Deep Learning Models for Oral Cancer Screening

Oral cancer is a leading cause of mortality in low-to-middle-income countries, where a shortage of specialists delays diagnosis. While poin…

13:00 JSTLLM/生成AILlama

Data-Driven Dynamic Algorithm Dispatch with Large Language Models

We introduce a large language model (LLM)-driven approach for generating dynamic algorithmic dispatch heuristics in high-performance linear…

13:00 JSTLLM/生成AIエージェント研究/論文ClaudeGPT / ChatGPT

K-Bench: measuring model performance on real scientific agent requests

Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with refe…

13:00 JST研究/論文

Generate in the Chart, Not on the Boundary: Function-Symbol Grounding for Hard Constraints in LTN-GANs

Logic Tensor Network-Enhanced Generative Adversarial Networks (LTN-GANs) inject background knowledge by grounding each logical axiom as a p…

13:00 JST研究/論文

Semantic Compression Trees: Multi-Resolution Knowledge Retrieval via Hierarchical Semantic Residuals

Retrieval-augmented generation relies mostly on flat, fixed-granularity indexes: documents are cut into uniform chunks and retrieved by sim…

13:00 JSTLLM/生成AI

SAEM: Stage-Aware Expert Management for Memory-Efficient MoE Inference in Chain-of-Thought Reasoning

Chain-of-thought (CoT) prompting improves LLM reasoning by decomposing complex problems into intermediate steps, but its sequential nature…

13:00 JSTLLM/生成AIビジネス/資金調達

Measuring Activation Control in Large Language Models

Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluation…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation

Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthet…

13:00 JSTLLM/生成AIエージェント

Context as an Environment: Programmatic Context Management for Long-Horizon Agents

LLM agents increasingly take on long-running tasks whose history grows far beyond a single model context window. Existing approaches compre…

13:00 JSTLLM/生成AI

From Association to Causation: Improving Retrieval Precision of Retrieval-Augmented Generation via Causal Relations and an Attention Mechanism

Retrieval-Augmented Generation (RAG) grounds LLM generation on retrieved documents, but the standard terminal retrieval stage--dense-vector…

13:00 JSTエージェント

ATHENA: Knowledge-guided agentic neural architecture search for AutoFormer-based electronic health record modeling

Transformer-based models are widely used for clinical prediction from electronic health records (EHRs), yet their architectures still requi…

13:00 JST研究/論文

Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention

Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they bel…

13:00 JSTエージェント

ECHO: A Cognitively Inspired, Auditable Memory Plane for Long-Horizon Agents

Long-horizon agents need memory that identifies relevant experience, resolves revisions, and exposes checkable provenance. We present ECHO…

13:00 JST画像/動画生成

What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation

Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic i…

13:00 JST研究/論文

Physics-Knowledge-Guided Hybrid Neural Learning for Arctic Sea Ice Concentration Evolution and Short-Range Prediction

Accurate modeling of sea ice concentration (SIC) evolution is essential for polar climate assessment and short?range sea ice prediction. Nu…

13:00 JST画像/動画生成QwenDeepSeek

HIRA: A Human-in-the-Loop Retrieval-Augmented Cascade for Document Classification in Regulated Industries

Document classification in regulated industries is constrained by data residency, limited cold-start labels, scarce review capacity, and co…

13:00 JST研究/論文

Hints, Critics, and Teachers: Prior Injection for Sparse-Reward RL in Vision-Language Math Reasoning

Reinforcement learning for vision-language math reasoning starves under sparse reward: on a pool of 20,830 visual-math problems where Qwen2…

13:00 JST画像/動画生成

VisAdj: Learning Adjacency Matrices from Node-Link Images

Learning adjacency matrices from node-link images is a fundamental problem for recovering structured graph information from visual observat…

13:00 JSTLLM/生成AIエージェント

Beyond Success and Failure: Length-Aware Contrastive Learning for GUI Agents

Graphical User Interface (GUI) agents powered by Multimodal Large Language Models (MLLMs) have shown strong potential for automating tasks…

13:00 JSTLLM/生成AIエージェント

GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development…

13:00 JSTLLM/生成AIエージェント研究/論文

LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization

Large language models have become increasingly capable agents for low-level code and kernel optimization, but isolated kernel benchmarks pr…

13:00 JSTエージェント規制/政策

AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations

Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. W…

13:00 JSTLLM/生成AIエージェント

MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance

LLM agents are moving from single-prompt use to long task streams in which reusable memory becomes a core capability for terminal, software…

13:00 JSTエージェント

HiMA-MDD: A Hierarchical Multi-Agent Harness for Interpretable Multimodal Depression Detection in Clinical Interviews

Depression assessment from multimodal clinical interviews requires integrating dispersed evidence from multiple symptoms into a coherent PH…

13:00 JST研究/論文

From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning

Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models rema…

13:00 JSTエージェント

Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning

Web agents promise to automate complex digital workflows, but their training remains limited by synthetic environments that look plausible…

13:00 JST研究/論文

Consistency Is Not Coherence: Orientation Search for Certified Alignments Between 4D Defence Upper Ontologies

We align three upper ontologies that sit under UK and NATO defence data infrastructure: the Information Exchange Standard (IES), the Higher…

13:00 JST研究/論文

ESCRAG-R1: Retrieval-Augmented Reinforcement Learning for Emotional Support Conversation

Emotional Support Conversation (ESC) systems aim to provide holistic support by balancing professional therapeutic competence with natural…

13:00 JSTLLM/生成AIロボティクス研究/論文

GuardianBench: A Same-Scene Instruction-Contrastive Benchmark for Latent Contextual Risk in Embodied AI

In embodied AI, safety risk can be latent: a benign instruction and a safe scene become hazardous only when composed. Prior work has advanc…

13:00 JSTLLM/生成AI

Multimodal Prompt Learning with Irregular EHRs for Robust Monitoring of Critical Care Patients

Accurate assessment of patients in intensive care units (ICUs) is essential for timely clinical intervention and improved patient outcomes.…

13:00 JSTエージェント

TessIndex: Capability Verified Identity System for the Agent Economy

Software systems have traditionally been organized around applications where human users act as principal decision-makers. Recent developme…

13:00 JST研究/論文

SSDi8: Accurate and Efficient 8-bit Quantization for State Space Duality

Recent advances in sequence modeling have highlighted Mamba as a state space architecture offering efficient long-range dependency modeling…

13:00 JSTLLM/生成AIエージェント

Repo2Skill-Evo: Repository Skills Go Stale in Silence

Large language model (LLM) agents increasingly operate over evolving software repositories, where success depends on repository-specific pr…

13:00 JSTエージェント研究/論文

Closed-loop AI achieves certifiable engineering design

Agentic AI has automated parts of scientific discovery, including paper generation, expert-level coding, therapeutic proposal, and autonomo…

13:00 JST研究/論文

Beyond Similarity: Heterogeneous Graph Learning for Multi-Objective Food Substitution in Charitable Food Agencies

Charitable food agencies play an important role in alleviating food insecurity by distributing donated food to people in need. However, the…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPT

Redteaming Leading Arabic LLMs with ASAS

As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasin…

13:00 JSTLLM/生成AI

DynaContext: Self-Improving Dynamic Contextualization of Optimized Prompts for Heterogeneous Parameter Extraction

Automated prompt and skill optimization typically produces a single static instruction that is reused across inference instances until the…

13:00 JSTエージェント

SPAR-Hate: An Auditor-Guided Multi-Agent Framework for Bilingual Hate Speech Parsing

Hate speech detection has recently shifted from coarse-grained classification to structured parsing, where systems must jointly identify ha…

13:00 JSTエージェント

One-Step Evolution for Long-Time Extrapolation: An Error-Bound-Informed and Prior-Guided Neural Residual Framework for Autonomous PDEs

Accurate simulation of the long-time evolution of systems governed by partial differential equations (PDEs) is central to scientific comput…

13:00 JSTビジネス/資金調達GoogleMicrosoftAlibaba

More Accurate or More Efficient? Evaluating Locally Deployed Compact Open-Weight Language Models for Mathematical Reasoning

Large language models are increasingly deployed on local hardware for privacy, cost, and accessibility reasons. Yet many evaluations emphas…

13:00 JSTエージェント

GenCoord: Skill-Path Commitments under Private Information

Suppose one embodied agent knows what must be built, while its teammate alone knows which transformation its workcell can perform. Neither…

13:00 JSTエージェントGPT / ChatGPT

MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds

Personal AI agents routinely consume external content while performing tasks such as web browsing, email processing, and SNS feed summariza…

13:00 JST研究/論文

Search Broadly, Seek Evidence on Both Sides, Decide Narrowly: Evidence-Admissible GraphRAG for Longitudinal Clinical Event Verification

Longitudinal clinical event-relation verification determines whether a patient record supports a specified relation among two or more clini…

13:00 JSTLLM/生成AIエージェント

From SQL Generation to Tool Selection: A Domain-Oriented Pattern for MCP Servers

Agents built on Large Language Models (LLMs) increasingly reach enterprise data through the Model Context Protocol (MCP), and many MCP data…

13:00 JST研究/論文ClaudeGPT / ChatGPTGoogleGeminiGrok

Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays

Geothermal well arrays, which organize multiple geothermal wells into carefully planned geometric configurations, provide opportunities to…

13:00 JST研究/論文

Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation

Synthetic clinical data generation with large language models addresses the scarcity that limits cancer staging research, but oncology hall…

13:00 JSTエージェント

Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks

As agents grow more capable and autonomous, their tendency to reward hack, satisfying a task's checks while violating its intent, becomes a…

13:00 JSTビジネス/資金調達NVIDIA

Development and Feasibility Evaluation of an Edge AI as Medical Device System for Breast Cancer Multidisciplinary Team Meetings

Breast Cancer Multidisciplinary Team (MDT) meetings manage increasingly complex cases under considerable time pressure, and documentation r…

13:00 JSTLLM/生成AIGemini

Task-Driven 3D Printability Assistance via Geometry- and Knowledge-Grounded LLM Reasoning

Printability assessment in additive manufacturing is typically conducted at the geometry level before printing to determine whether a compu…

13:00 JSTエージェント

MegaMem: A Retrieval Solution for Ultra-Large Context Windows

Modern language models and agents increasingly require persistent memory for complete codebases, long interaction histories, and heterogene…

13:00 JSTLLM/生成AI

Measuring Stability and Failure Behavior in Language Models Under Structured Perturbations

Language models are usually judged by a single accuracy score, which does not reveal how their performance degrades as inputs are perturbed…

13:00 JST研究/論文

MEMONDEMAND: A Memory Management System for Large-Scale Enterprise Data

Enterprise repositories are large, heteroge- neous, and continuously updated, making re- trieval difficult when efficient access, source- f…

13:00 JSTビジネス/資金調達GemmaQwen

Evaluation of Small Vision-Language Models on Qualitative Mechanical Problems

Qualitative mechanical problem-solving (QMPS) refers to solving qualitative problems from the mechanical domain. Qualitative problems can b…

13:00 JSTエージェント

AUDITA: certified auditing and causal attribution of adverse outcomes in autonomous multi-agent systems

Physical automation is scaling toward fleets of embodied machines commanded by an AI brain. Early deployments already run factories and war…

13:00 JSTLLM/生成AI

Aggregation-Aware Synthetic Text Generation Against Authorship Re-Identification

Online users often release multiple texts under the same identity, giving attackers an author profile that can reveal more than any single…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

MCP-Universe RL: A Framework for Training MCP Tool-Use Agents via Reinforcement Learning

Reinforcement learning (RL) has become an effective way to improve the tool-use ability of large language models (LLMs), but most existing…

13:00 JSTLLM/生成AIエージェント

Role-Specialized Mixture-of-Agents with Open-Weight LLMs for Clinical Prediction

Large Language Models (LLMs) are increasingly applied to clinical prediction tasks such as in-hospital mortality and readmission from elect…

13:00 JSTエージェントGPT / ChatGPT

Disagree to Explore, Agree to Commit: Routing-Guided Test-Time Scaling for Software Agents

Software-engineering agents solve repository-level tasks through long, stochastic tool-use trajectories, and repeated attempts often find f…

13:00 JST研究/論文

Query-Driven Multimodal Information Extraction from Long Documents

In domain-specific multimodal long documents, images and text jointly convey complex knowledge that cannot be fully captured by plain text…

13:00 JSTLLM/生成AI画像/動画生成

Beyond What Meets the Eye: Unveiling Situational Illusions for Multimodal Large Language Models

Real-world situation appearances can deviate from their underlying physical states, challenging the reliability of multimodal large languag…

13:00 JSTエージェントClaude

Read Less, Solve More: Token-Efficient Sparse Reading for AI Agents

Long-horizon agents increasingly rely on repeated access to external artifacts, yet current reading interfaces often expose entire objects…

13:00 JSTLLM/生成AIエージェント

Clarify User Expertise: Towards Proactive Conversational Agents Tailoring Responses to User Proficiency

In the context of information seeking, conversational agents are undergoing an evolution from reactive tools to proactive, personalized ass…

13:00 JSTLLM/生成AIエージェント

HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory

Long-term memory is crucial for personalized responses and long-horizon agent interactions. Existing methods often rely on LLMs to compress…

13:00 JSTLLM/生成AI

Where Cognition Lives: Dissecting Emergent from Computed Function in a Minimal Complete Cognitive Architecture

A cognitive architecture is more than the module that reasons: it must also decide how long to think and what deserves the effort. We built…

13:00 JST研究/論文

Addressing the Selection Problem in Explainable AI

Explainable AI (XAI) research has produced a plethora of explanation techniques, yet user studies repeatedly show that available explanatio…

13:00 JSTビジネス/資金調達

Analyzing and Mitigating Cross-Lingual Degradation in Multilingual Medical VQA

Medical visual question answering (VQA) is a crucial task in clinical AI, yet its evaluation has so far centered almost exclusively on Engl…

13:00 JSTロボティクス

WAM-OPD: On-Policy Distillation for World Action Models

World action models (WAMs) couple visual future prediction with robot action generation, but accelerated students can lose task capabilitie…

13:00 JSTLLM/生成AI研究/論文GPT / ChatGPT

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in i…

13:00 JST研究/論文

Where World Models Break: Natural-Input Failure Discovery

World models predict action-conditioned futures and serve as critical internal simulators for downstream planning and control. However, cat…

13:00 JSTLLM/生成AI

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Multimodal Large Language Models (MLLMs) capable of thinking with images often rely on external tools for fine-grained perception. However,…

13:00 JSTLLM/生成AI

When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing

Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, appare…

13:00 JSTLLM/生成AIエージェント研究/論文

Small Reasoning Models are Instruction Followers in Function Calling

Function calling represents the core capability of agentic large language models (LLMs). Existing research has focused on enhancing LLMs fu…

13:00 JST研究/論文

When Does AI for PDEs Yield Scientific Evidence?

Existing AI-for-PDE benchmarks primarily assess models in terms of predictive or approximation accuracy. In physics research, however, AI o…

13:00 JSTエージェントビジネス/資金調達研究/論文

ClawProBench: Trace-Aware Evaluation of AI Agents with Runtime Coverage and Frozen Workplace-Style Holdouts

Agent benchmarks often evaluate only final answers even when agents run on stateful runtimes. We argue this under-specifies what is being e…

13:00 JSTエージェント

HANSARD: A Reference Architecture for Forensic Readiness, Runtime Witnessing, and Graded Attribution in Autonomous Multi-Agent AI Systems

Autonomous multi-agent systems nowadays act in finance, software supply chains, and security operations. Already, the first largely AI-orch…

13:00 JSTエージェントClaudeGPT / ChatGPTDeepSeek

CONTRAMEM: Learning Self-Evolving Procedural Memory from Contrasting Multi-Model Trajectories

Autonomous computer-use agents are increasingly applied to long-horizon tasks requiring coordinated application calls, persistent state tra…

13:00 JSTエージェント

STAGE: Stateful Translation to Agentic Graph Execution with Policy-Scoped Context and Deterministic Control

Policy-governed agents must interpret case evidence while following an authorized procedure. We present \textsc{Stage}, an executable-graph…

13:00 JSTエージェント

Scaling Curriculum Learning For Autonomous Driving

Batched simulators for autonomous driving have recently enabled training reinforcement learning (RL) agents at scale, encompassing thousand…

13:00 JSTLLM/生成AIビジネス/資金調達

ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation

Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-lan…

13:00 JSTエージェント

CausalCache: Conditional High-Fidelity Restoration for Long-Horizon GUI Agents

Long-horizon GUI agents can retain a complete interaction trace cheaply as textual action records, but expose only a few past events to the…

13:00 JST研究/論文

Weakly supervised concept Bottleneck Learning for Robust Two stage Object centric visual reasoning

Two-stage neuro-symbolic architectures provide an elegant paradigm for visual problem solving by cleanly separating connectionist perceptio…

13:00 JSTエージェント

Coalition-Aware Skill Reliability for Self-Evolving Agents

Agent skills, structured artifacts distilled from interaction trajectories and dynamically reused from skill banks, have become a central m…

13:00 JSTLLM/生成AIエージェント

DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue

Large Language Model (LLM)-based counseling agents can generate fluent and supportive responses, but they often lack the structured, goal-d…

13:00 JSTLLM/生成AI

CAI-DLLM: Convergence Aware Inference for Diffusion Language Models

Diffusion language models can generate many tokens in parallel, but they still require repeated denoising steps during inference. This make…

13:00 JSTエージェント

A-CPES: A Reference Framework for Agentic AI in Cyber-Physical Energy Systems

Energy system operation contains a loop of work that automation has never taken over: posing the optimization problem the current cycle sho…

13:00 JSTエージェント

Robustness Analysis of Agentic AI to Inconsistent and Incomplete Tool Responses

Robustness to a bad tool return means answering it in the way that return calls for, which depends on how the tool went wrong. A tool that…

13:00 JSTエージェント

Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf

Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are e…

13:00 JSTLLM/生成AIDeepSeek

CacheRouter: A Dual-Path Tool Routing Architecture with Cache-Preserving Main-Model Isolation for Long-Tail Tool Discovery

Tool use in LLM systems faces a structural trade-off. Progressive disclosure keeps the prompt small by showing only the tools relevant to t…

13:00 JST研究/論文

SEAM: Shot Entity-Attribute Memory for Consistent Short-Drama Generation at Scale

Short-drama generation has grown into a large, industrialized pipeline, and as it scales from isolated shots to the episode level, visual c…

13:00 JSTLLM/生成AIエージェントロボティクス

LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans

Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize o…

13:00 JSTエージェントClaude

The Compaction Cliff in Long-Running AI Agent Memory

A safety rule and an episodic log compete for the same tokens in an AI agent's context. When the budget overflows, both are summarized at t…

13:00 JSTLLM/生成AI

Compositional Chain-of-Relations for Faithful Knowledge Graph Question Answering with Large Language Models

Knowledge graph question answering (KGQA) is a key task for evaluating KG-augmented Large Language Models (LLMs), and complex KGQA that req…

13:00 JSTエージェント

The Retriever Should Remember: Experience-Amortized Reranking for Long-Term Agent Memory

Long-term language-model agents accumulate memories across interactions, but their retrievers typically do not accumulate retrieval experie…

13:00 JSTLLM/生成AI

TailSieve: Partial-Rollout-Guided Tail Routing for LLM Rollouts

Large-scale rollouts have become a core component of modern LLM systems, spanning reinforcement learning (RL) post-training, on-policy dist…

13:00 JSTLLM/生成AIGPT / ChatGPT

Performance of a domain-specific large language model in answering patient questions in psychiatry

Background This study was designed to evaluate whether a domain-specific large language model (LLM) trained exclusively on patient educatio…

13:00 JSTLLM/生成AI

Beyond the Harness: End-to-End Optimization of Context Artifacts for Enterprise Text-to-SQL

Deploying LLMs for enterprise Text-to-SQL is bottlenecked less by the model than by what context reaches it: business logic spans thousands…

13:00 JST研究/論文

Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku

The social interactions among crowds via \textit{Danmaku} (a.k.a., bullet comments) on modern multimedia platforms can facilitate both view…

13:00 JST研究/論文

FinixDoc: Rethinking Financial Document Parsing Beyond Saturated Benchmarks

Financial document parsing requires accuracy, structural consistency, and verifiability that current benchmarks often fail to reflect. We p…

13:00 JSTエージェント

GSAR: Goal-State-Anchor Rewards for Mobile GUI Agents with Self-Evolving Data Synthesis

Vision-Language Models (VLMs) based GUI agents stand to benefit significantly from online reinforcement learning (RL). However, their train…

13:00 JSTLLM/生成AIビジネス/資金調達

Your AI, On a Dial: Controlling Investment Bias in LLMs with a Single Neuron

Large language models (LLMs) are increasingly used in investment decision-making, yet prior work shows that they exhibit systematic, model-…

13:00 JSTLLM/生成AI

Proxy reliance in large language model decisions is uncalibrated to predictive evidence

Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermi…

13:00 JSTエージェント

CDEG: Learning Decision-Critical Evidence for Long-Horizon Diagnostic Agents

Unlike static medical question answering, long-horizon diagnosis captures the sequential nature of clinical practice: evidence is progressi…

13:00 JST研究/論文

Beyond Observed Auxiliary Relations: Environment-Conditioned Modeling for Multi-Behavior Recommendation

Multi-behavior recommendation (MBR) leverages auxiliary behavioral signals, such as clicks and add-to-cart, to enhance target behavior pred…

13:00 JSTエージェント

Concepts for Securing Agentic AI Coding and the Terok Environment

Agentic AI is a fascinating new tool for software development. It is a huge step forward compared to "conventional" AI assisted coding, whi…

13:00 JSTエージェントビジネス/資金調達

What Process Evaluation of Coding Agents Actually Measures: Action, Task, and Step Are Three Different Levels

Coding agents are increasingly evaluated not only by whether they solve a task, but also by how they execute it. However, existing process-…

13:00 JSTLLM/生成AIエージェント

Buried in Textual Debt: Context Pruning with Visual Evidence Preservation for MLLM Agents

Multimodal Large Language Models (MLLMs) are increasingly deployed as multi-step agents, where explicit reasoning supports task decompositi…

13:00 JST画像/動画生成エージェント

ParallelWorld: Test-Time Scaling for Embodied Reasoning

Embodied Reasoning constitutes a fundamental capability of embodied intelligence, serving as the basis for autonomous perception, reasoning…

13:00 JSTLLM/生成AIエージェント

Toward Effective and Reliable LLM Agents via Dynamic Ontology

Large language model (LLM) agents rely heavily on knowledge encoded in model parameters or presented as unstructured context. In domain-spe…

13:00 JSTエージェントビジネス/資金調達

Budget-Constrained Embodied Perception: Four Resource Walls and a Pre-Registered Evaluation of Access-Structured Perception on Open Models at less than 31B

Embodied multimodal agents must answer from growing observation streams under a fixed per-decision token budget. We formalize this constrai…

13:00 JST研究/論文

SA-RSQ: A Versatile Sparse Representation Framework for Multi-modal Recommender Systems

Deploying high-dimensional multimodal features in industrial recommender systems incurs substantial storage and latency overhead. Hard quan…

13:00 JSTLLM/生成AI

PatchWrite: One Line, Not One Section -- Compile-Gated, Validity-Preserving Editing for AI-Drafted Manuscripts

Automated manuscript pipelines often regenerate an entire section to repair a local defect, allowing unrelated metrics and citations to cha…

13:00 JSTLLM/生成AI

PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies

Large language models (LLMs) are increasingly deployed in education, healthcare, policy advising, and other interactive settings, where use…

13:00 JSTエージェント

Artificial Empathy: Towards a Framework for Unsupervised Agency Detection and Policy Reconstruction

We study how an AI system can identify and model other agents in its environment from observation alone, which is a capability necessary fo…

13:00 JSTLLM/生成AIエージェント研究/論文

MobilePA-Bench: Benchmarking Mobile Planner Agents on Complex Real-World Tasks

As on-device LLM agents evolve into personal copilots, the mobile operating system has become a key testbed for this paradigm, making rigor…

13:00 JSTLLM/生成AIエージェント

AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces

LLM agents remain unreliable on long-horizon tasks, where small local failures can compound over extended interactions and lead to overall…

13:00 JSTLLM/生成AIエージェント研究/論文GPT / ChatGPT

From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation

Web search agents powered by Large Language Models (LLMs) show strong promise, but deep research tasks expose a recurring failure mode: onc…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications

Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval,…

13:00 JSTLLM/生成AIビジネス/資金調達ClaudeGemini

Improving O-RADS Risk Stratification from Ultrasound Reports: A Comparative Evaluation of Hybrid versus End-to-End LLM Reasoning Strategies

Background: Automating clinical guideline-based decision-making with large language models (LLMs) remains challenging because of reliabilit…

13:00 JST画像/動画生成

From Generation to Simulation: How Far Are World Models from Being True Simulators?

With the rapid progress of diffusion models and large-scale video generation, generative world models are increasingly expected to replace…

13:00 JSTLLM/生成AIエージェント

AgentWeave: Routing Before Reasoning for Efficient Function Calling in Tool-Rich Language Models

Large language models increasingly operate over large collections of tools, functions, APIs, and specialized agents. As the candidate actio…

13:00 JSTハードウェア/半導体

POOL: Propagated Uncertainty Over Lookalikes

Black-box large language models need confidence scores that can separate likely-correct from likely-incorrect outputs, enabling systems to…

13:00 JST研究/論文

Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement

Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words, imagery, and prosody. However, many curren…

13:00 JST研究/論文OpenAI

AI emotional support is better only when chosen, but shifts preferences even when it is not

People increasingly face a novel decision when seeking emotional support: human or AI. In existing studies, AI's empathic messages are rate…

13:00 JSTLLM/生成AI

Cognitive Profiling of LRMs' Reasoning Traces Using Bloom's Taxonomy

Large Reasoning Models (LRMs) have revolutionized reasoning in LLMs, and the increasing public availability of reasoning traces creates val…

13:00 JST研究/論文

What is mathematics now, and what should it be?

Advances in neural theorem provers have been impressive, but the successes obscure a broader vision of what AI can do for mathematics and h…

13:00 JST研究/論文

Is Next-Chunk Reasoning RL Really Better than SFT? Revisiting Training Strategies under no-CoT Data

Recent work proposes next-chunk reasoning RL for leveraging no-CoT data---corpora such as worked solutions and textbook derivations that co…

13:00 JSTLLM/生成AI

Automated Construction of FAIR Digital Object Knowledge Graphs from Flat Cultural Heritage Records

The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever poss…

13:00 JSTLLM/生成AI

Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance

Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, in…

13:00 JSTLLM/生成AIエージェント

Apodex 1.1: Scaling Agentic Intelligence for Complex Work

General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, infor…

13:00 JSTビジネス/資金調達研究/論文

EviSafe: Evidence-Grounded Safety Evaluation for Vision-Language Models

Vision-language model safety benchmarks typically evaluate only final responses: whether a model refuses, warns, or complies. This outcome-…

13:00 JSTLLM/生成AIエージェント

Agent-G$^2$: Gaussian Guidance for Agentic Reinforcement Learning

Hint-based reinforcement learning addresses reward sparsity in long-horizon agentic tasks by retaining a prefix of an expert trajectory bef…

13:00 JSTLLM/生成AIClaudeGemini

Walking on the DARKSIDE

Large Language Models (LLMs) recognise patterns but do not natively track the path of exclusions that a coherent discourse demands. When an…

13:00 JST研究/論文

Modalities Should Talk to Each Other: Dual-Stream Multimodal Learning for Long-Horizon Influenza Forecasting

Forecasting long-range influenza-like illness (ILI) matters for public health readiness. Publicly available surveillance datasets typically…

13:00 JSTエージェントQwen

MediSkill-Evo: Process-Constrained Self-Evolution for Evidence-Grounded Clinical Interaction

Interactive clinical agents must gather decisive evidence and convert it into grounded actions under partial observability. A correct final…

13:00 JSTエージェント

SkillAlchemy: Open-World Agent Skill Creation

Agent skills are reusable procedural artifacts that extend language agents with specialized workflows, tool conventions, and domain behavio…

13:00 JST研究/論文

Characterizing Necessary Losers to Explain Tournaments Losers

We study the problem of formally explaining why a candidate was not selected by a given tournament rule, by identifying sub-tournaments in…

13:00 JST研究/論文

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a…

13:00 JSTLLM/生成AIGPT / ChatGPT

Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation

We present Team Semiintelligencn's solution for the ACM RecSys 2026 TalkPlayData Challenge, addressing conversational music recommendation…

13:00 JST研究/論文

SRPO: Self-Reflective Policy Optimization for Long-Horizon Reasoning

Self-reflection is a powerful mechanism for credit assignment in human learning, converting sparse outcome feedback into actionable guidanc…

13:00 JSTLLM/生成AI

Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty

Reasoning-Induced Misalignment, where fine-tuning on reasoning data containing no harmful content, including mathematics, code, and problem…

13:00 JSTエージェント研究/論文

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natura…

13:00 JST研究/論文

Correcting a learned physical invariant improves world-model rollouts

World models can predict video without learning dynamics that they reliably preserve. We test whether a frozen DreamerV3 trained only on pe…

13:00 JST研究/論文

How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles

While AI assistance can improve human task performance in the short term, it may also undermine the development of skills in the longer ter…

13:00 JSTLLM/生成AIエージェント

Prime Agent: A Self-Improving RLM Harness

Language models are sequential processors, but long-horizon agency requires external information and computation beyond model weights and a…

13:00 JST研究/論文

ReWorld: An Interactive World Model with Long-Horizon Memory

An interactive world model must follow the user's actions, remember the places it has shown, and stream in real time. The tension is struct…

13:00 JST研究/論文

Correcting Variable Importance Scored by Random Forests

Variable importance produced by Random Forests (RF) is used widely in statistical data analysis, and has played an important role in a vari…

13:00 JSTエージェント

Small Language Model enabled Autonomous agent for Language-Conditioned Cognitive Radar

Modern radar systems require adapting their processing strategies in response to changing interference, clutter, and data availability. Thi…

13:00 JST研究/論文

Triangular Fuzzy Rescaling Distance

Decision-making in complex systems often involves dealing with imprecise or uncertain information, frequently represented using fuzzy sets,…

13:00 JSTLLM/生成AI

Distinguishing Revision and Delayed Elaboration in Incremental Narrative Interpretation

Both human and AI systems that process narrative or long-form content operate incrementally: input is received over time, and internal repr…

13:00 JSTLLM/生成AI

KSE-Web: An Analysis of Hybrid Retrieval and LLM-Assisted Query Expansion for Low-Resource Khmer Semantic Search

As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak…

13:00 JSTLLM/生成AILlama

PepLLM: ESM-Guided Llama for Structured Protein-Peptide Binding Interface Analysis

Protein-peptide interactions are central to cellular regulation and peptide-based drug discovery, yet existing computational methods mainly…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Wazobia Eval: A Benchmark for Nigerian Pidgin Emotion Understanding, Sarcasm Detection, and Cultural Reasoning

Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existi…

13:00 JSTLLM/生成AIハードウェア/半導体

On the Role of Citations in Preference Data

Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwar…

13:00 JSTLLM/生成AIエージェント

Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models

Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but…

13:00 JST画像/動画生成エージェントロボティクス

RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perception

With the increased adoption of robotic agents operating in human environments by scanning and sharing 3D representations (e.g., for fleet l…

13:00 JST研究/論文

Determinants of Starting Salaries for Filipino Graduates: An Explainable Machine Learning Approach

Filipino graduates face a persistent disconnect between educational preparation and labor market outcomes, where starting salary is a key s…

13:00 JSTLLM/生成AI

Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems

Modern multilingual tokenizers often fragment Ukrainian and other underrepresented Cyrillic-script languages more heavily than English, cre…

13:00 JSTLLM/生成AI

A Social Media Analysis of Discourse on the Israel--Palestine Conflict on Telegram

Social media has become a central arena in which armed conflicts are contested, yet the pro-Israel and pro-Palestine communities on Telegra…

13:00 JST研究/論文

Model of Models: When Does Emitting a Specialist Beat Attending, Adapting, or Tuning?

Given a task described by a few examples, how should a model be specialized to it? Four mechanisms are available -- zero-shot, in-context a…

13:00 JST研究/論文

Interrupting the Chain: Human Perception of AI-Generated Disinformation Through a Kill Chain Lens

Generative AI enables customized misinformation at scale, yet defenses remain largely reactive. We present empirical findings from a human-…

13:00 JST研究/論文

A Survey Instrument to Assess Students' AI and Generative AI Knowledge

In this research-to-practice paper we present a survey that can be used to assess students' AI knowledge. As the use of artificial intellig…

13:00 JSTロボティクス

ODG-NoMaD: Overhead-Camera Direction-Guided NoMaD

NoMaD [31] is a learned vision-navigation policy that unifies goal-conditioned navigation and exploration in a single goal-masked diffusion…

13:00 JST研究/論文

Runtime Action Interference for AI Control of AlphaStar in StarCraft II

A trained reinforcement learning policy does not determine the complete behavior that users encounter: deployment code still schedules, adm…

13:00 JSTロボティクス

Mamba-based Selective State Space Modeling Improves the Accuracy-Complexity Tradeoff of SmolVLA Vision-Language-Action Experts

Vision-language-action (VLA) models face a crucial tradeoff between their task success rate and the policy-call frequency. Executing a sing…

13:00 JSTLLM/生成AI

Sycophants in the Courtroom: Are LLMs Fragile to Juridical Authority and Evolving Legal Standards?

In medicine, claims remain valid when supported by empirical evidence grounded in stable biological reality. In law, by contrast, truth is…

13:00 JSTLLM/生成AI

Mitigating Bias in Large Vision-Language Models via Counterfactual Ensemble Decoding

Large Vision-Language Models (LVLMs) have achieved remarkable performance across a wide range of tasks; however, they often inherit social…

13:00 JSTロボティクスビジネス/資金調達

Operational digital twin clinics enable task-based evaluation of embodied AI

Embodied artificial intelligence (AI) must be tested in the clinical environments where it will operate, but building realistic, robot-test…

13:00 JSTLLM/生成AIエージェント

Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing

Agentic security uses large-language-model (LLM) agents to plan, dispatch, and interpret security tools. As these systems move from demonst…

13:00 JST画像/動画生成

Aligning Human Sense: Calibrated Distributional Reward Learning for Video Generation

Video generation is central to AI-powered content creation. Aligning generated videos with human preferences is a key criterion for evaluat…

13:00 JSTエージェントロボティクス

Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics

Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and…

13:00 JSTロボティクス

Constructing Predictive Surgical Path for AI-based Capsulorhexis Skill Transfer

Automated training of surgeons is one of the most crucial factors that significantly minimize surgical training risks and expenses. With re…

13:00 JST画像/動画生成ClaudeGPT / ChatGPT

FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows

Vision Language Models have recently shown improvements in several objective and verifiable domains such as object detection but continue t…

13:00 JSTLLM/生成AI

CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance

Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the…

13:00 JST画像/動画生成

Complexity Induction: Compositional Generalization via Structured Label Distortion

We demonstrate that structured distortion of training data - which we term complexity induction - can induce compositional generalization i…

13:00 JST画像/動画生成

Reliability- and Anatomy-Consistency-Aware Multimodal Learning for Robust Fracture Classification from Bangladeshi Radiographs

Background: Multimodal fracture classifiers may benefit from patient and anatomical metadata, but they can also become brittle when context…

13:00 JST画像/動画生成

TASSO: TAsk-Specific Subspace Optimization for Continual Learning of Vision-Language Models

Vision-Language Models (VLMs) exhibit strong zero-shot capabilities, making them an attractive solution for continual learning across diver…

13:00 JST研究/論文

KAN-Robust-Bench: A Benchmark for Evaluating the Robustness of Kolmogorov-Arnold Networks

While machine learning models have demonstrated strong performance in many domains, these models have shown profound vulnerabilities when t…

13:00 JST研究/論文

Selection of Heart Sound Segments for Synchronous Classification of Multi-channel Heart Sounds

Cardiac auscultation remains the most cost-effective screening procedure for cardiovascular diseases, and requires listening at the four ma…

13:00 JSTLLM/生成AIエージェント

SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation

Prompt injection is listed as the \#1 threat to AI agents. When an agent accesses external data from websites, files, or emails, an attacke…

13:00 JST画像/動画生成

presto: Efficient, Training-free, and Open-world Object Placement via Imaginary Search

Object placement is critical in image composition, requiring spatially and semantically coherent positioning of objects within diverse scen…

13:00 JSTLLM/生成AIエージェント

Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents

Large language models (LLMs) are increasingly deployed as tool-augmented agents, where responses can depend on tool calls and external obse…

13:00 JSTLLM/生成AIビジネス/資金調達

Automating Multi-Hop RAG Evaluation via TRIAD: From Context Extraction to Validated Dataset Generation

Recent advances in LLMs and the adoption of RAG systems in industry have created a need for domain-specific question-answer datasets that c…

13:00 JSTLLM/生成AI

Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning

Multimodal Large Language Models (MLLMs) are increasingly deployed in high-stakes domains where fairness is a critical safety requirement.…

13:00 JST研究/論文

Power-Performance Characterization of TinyML Systems

TinyML systems are enabling machine learning (ML) inference at the edge. However, there is little quantitative analysis of such systems. Th…

13:00 JST研究/論文

Why This, Not That? Mining User Profiles for Pair-wise Counterfactuals

The topic of explanation in recommender systems has seen steady research attention since the earliest days of the field. With some exceptio…

13:00 JST研究/論文

SynEHR: Joint Modeling Inter-visit Temporal Evolution and Intra-visit Clinical Structure for Longitudinal EHR Synthesis

Longitudinal electronic health records (EHRs) document patients' sequences of clinical visits over time, preserving the temporal evolution…

13:00 JST研究/論文

Read, Write, Relax: Why Neural PDE Surrogates Need Both Global and Local Processing

Recent mesh-based simulation advances have, in no small part, relied on neural surrogates of two distinct families: global models that rout…

13:00 JST研究/論文

Scalable quantum simulation of continuous-time generative models via tensor networks

Continuous-time flow and diffusion models are widely used across many application domains, from large-scale deployment in computer vision a…

13:00 JST研究/論文DeepSeek

Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data

Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what mod…

13:00 JSTLLM/生成AIエージェントAnthropicClaudeOpenAIGPT / ChatGPTGoogleGemini

Architecture as Capability Equalizer for Coding Agents

LLM-based coding agents generate complete software systems from high-level descriptions, yet little is known about how the format of archit…

13:00 JST画像/動画生成NVIDIA

LiteEvent-AE: Lightweight Autoencoder for Event-Based Vision on Low-Latency Energy-Constrained Edge Devices

Event-based vision has emerged as a promising paradigm for energy-aware artificial intelligence (AI), offering sparse, low-latency visual s…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Evaluation Awareness in Language Models: Representation, Verbalization, and Control

Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about…

13:00 JSTLLM/生成AIビジネス/資金調達

Lexical Coupling in GUI Element Grounding: Sentence Embeddings Track Labels across Mobile and Web

GUI grounding evaluations that expose UI elements as text metadata often treat high instruction-element embedding similarity as evidence of…

13:00 JST画像/動画生成

SAFE-G: Structure-aware Faithful Evidence-guided Generation for Knowledge-based Visual Question Answering

Knowledge-based Visual Question Answering (KB-VQA) aims to answer queries that necessitate reasoning over external knowledge sources beyond…

13:00 JST研究/論文

ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Models

As machine learning (ML) models are increasingly deployed in high-stakes environments, explainable AI (XAI) methods like SHAP and LIME have…

13:00 JSTLLM/生成AIハードウェア/半導体研究/論文

More Computational Resources Do Not Ensure Higher Scholarly Impact: Evidence from Leading NLP Conference Papers

Computational resources are increasingly central to NLP research, but how closely reported GPU capability aligns with scholarly impact rema…

13:00 JSTLLM/生成AI画像/動画生成

PatchGate: Narrowing the Verbalization Gap with Intrinsic Object Inventories in Frozen Vision-Language Models

Reliable image captioning in Vision-Language Models (VLMs) requires captions to be both precise and complete, avoiding unsupported object m…

13:00 JSTLLM/生成AIエージェント

Training a Knowledge Base: Supervised Structure Learning for Agent-Curated Document Stores

Retrieval-augmented generation treats the document store as a frozen input, and the systems that instead let an agent curate one never meas…

13:00 JSTLLM/生成AI

LLMs are Few-Shot Decision-Makers: Generalized Context-Aware Microgrid Frequency Control through Prompt Decision Transformer

The rapid evolution of energy structures has positioned microgrids as pivotal components of next-generation power systems, offering enhance…

13:00 JSTLLM/生成AI

ChainPrune: Evaluating and Reducing Redundancy in Long Chain-of-Thought Reasoning

Chain-of-Thought (CoT) reasoning has significantly enhanced the multi-step problem-solving capabilities of large language models (LLMs) by…

13:00 JSTLLM/生成AIエージェント

HiDiffTIR: Hierarchical Difficulty-Aware Policy Optimization for Multi-Turn Tool-Integrated Reasoning

Tool-Integrated Reasoning (TIR) is a fundamental capability for LLM agents to solve complex tasks by interacting with external tools iterat…

13:00 JSTLLM/生成AI画像/動画生成エージェントGPT / ChatGPT

BioMed-Agent-RL: A Meta Learning, All You Need for Biomedical Applications

The current progress of Clinical Vision Large Language Models (C-VLLMs) has substantially improved digital diagnostics, still these framewo…

13:00 JST画像/動画生成

GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration

Text-to-image (T2I) diffusion models offer powerful visual generation, but their controllability creates a critical safety challenge: adver…

13:00 JST研究/論文

Pruned Traffic Trees: Native Semantic Compression with a Protocol-Structured Model Family for Encrypted Traffic Classification

Deep learning has achieved strong performance in encrypted traffic classification (ETC), yet its computational cost limits deployment on re…

13:00 JST画像/動画生成

A Scalable Vector Graphics Latent Space

Scalable Vector Graphics are a fundamental medium for resolution-independent visual content, yet the deep learning community lacks a contin…

13:00 JSTLLM/生成AI

Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attacks

Locally deployed Large Language Models (LLMs) via inference engines such as Ollama run without the moderation and abuse detection present i…

13:00 JSTLLM/生成AI

Bi-EZP: LLM-Guided Bilevel Program Evolution for Ensemble Zero-Cost Proxy Discovery

Zero-cost proxies enable neural architecture search (NAS) to rank candidate networks from statistics computed at initialization, avoiding r…

13:00 JSTLLM/生成AIエージェント

EDGE: Experience-Distillation for Guided Exploration in Agentic Reinforcement Learning

Reinforcement learning with outcome-based objectives such as GRPO enables LLM-based agents to solve complex, long-horizon tasks, yet the re…

13:00 JSTLLM/生成AI

Bulbul: A Dataset for Dialectal Arabic Speech Recognition

Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speec…

13:00 JSTLLM/生成AI

NoTB: Oracle-Free Triage of LLM-Generated RTL via Cross-Model Formal Consensus

Large language models (LLMs) are increasingly used to generate register-transfer-level (RTL) designs from natural-language specifications.…

13:00 JST研究/論文

Variance Driven Exploration: A Provable and Efficient Methodology for Pure Exploration in Highly Stochastic Environments

We propose Variance Driven Exploration (VarDE), a principled approach for pure exploration in highly stochastic environments, where the exp…

13:00 JST研究/論文

Barycentric Fused Gromov-Wasserstein Balancing for Causal Inference under Multiple Treatments

Estimating heterogeneous single and interaction treatment effects from observational data under multiple simultaneous treatments is crucial…

13:00 JSTエージェント研究/論文

Multi-Agent Discovery and Resource-Aware Autonomous Exploration of Scientific Datasets

Modern scientific facilities and instruments generate datasets at scales that are difficult for individual researchers to discover, access,…

13:00 JST画像/動画生成研究/論文

CRS-Bench: A Reference-Relative Reliability Benchmark for Medical Image Encoders

Pretrained image encoders are central to medical image classification, where expert annotation is costly and task-specific cohorts are ofte…

13:00 JST研究/論文

Discovering Dual-Origin Slow Wind from Solar Orbiter with Self-Supervised Contrastive Learning

Whether the slow solar wind originates from one coronal source or two distinct channels remains a central open question in heliophysics. Re…

13:00 JST画像/動画生成

ADMIL: Attention-Distilled Multiple Instance Learning for Selective Foundation Model Inference in Pathology

Attention-based multiple instance learning (ABMIL) using pathology foundation model embeddings is effective for slide-level tasks, but exha…

13:00 JST画像/動画生成ロボティクス

Inferring Action from Future Latent State for Robotic Manipulation

World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and ro…

13:00 JSTLLM/生成AIハードウェア/半導体

Real-TurnTurk: A Multimodal Turkish Corpus for Turn-Taking Prediction

Turn-taking is a basic organizational feature of human conversation and remains difficult to model in natural, synchronous dialog systems.…

13:00 JST研究/論文

Improving Energy Efficiency of Oil Platforms Through Optimal Loading of Diesel Generators Using Machine Learning and Search Algorithms

Rising energy demand, fossil fuel depletion and climate change highlight the need for more efficient energy production and consumption. Off…

13:00 JST研究/論文GPT / ChatGPTMistral AI

On Predicting Vulnerability Severity Using In-Context Learning: An Industrial Case Study

Modern software systems require earlier and more scalable vulnerability severity assessment to reduce exposure to high-impact security flaw…

13:00 JSTLLM/生成AILlama

Semantic Reasoning Denoising: Correcting Language Model Reasoning with Semantic Operators

Large language models can produce fluent reasoning traces whose local semantic errors propagate to an incorrect conclusion, while unconstra…

13:00 JST画像/動画生成

Learning Implicit Constitutive Laws for Dynamic 3D Gaussian Splatting from Monocular Videos

We present GCA (Gaussian Constitutive Alignment), a framework for learning implicit constitutive laws from monocular dynamic video of defor…

13:00 JST画像/動画生成

TRACE: Artifact-Robust Statistical Shape Modeling from Imperfect Surface Scans - A Case Study in Craniosynostosis 3D Photography

Craniosynostosis severity analysis increasingly relies on statistical shape models (SSMs) to quantify cranial morphology, but most existing…

13:00 JSTLLM/生成AIエージェント

SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning

Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, prot…

13:00 JSTLLM/生成AI

Lexical Perturbations Disrupt LLM Reasoning: An Empirical Study of Attention Diversion

Large Language Models (LLMs) achieve strong reasoning performance, but their robustness to realistic lexical corruption remains poorly unde…

13:00 JSTLLM/生成AIロボティクスGPT / ChatGPT

Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints

LLMs generate fluent plans for robots but routinely violate the syntactic and se8mantic constraints they must satisfy to execute, and exist…

13:00 JST研究/論文

Why Does Robustness Reduce Superposition?

The study of adversarial examples and their origins remains an open area of research. Mechanistic interpretability, and superposition in pa…

13:00 JST研究/論文

Joint Causal Structure and Cluster Discovery Using Variational Inference

Causal discovery aims to understand the relationships between individual random variables. In many applications, such as brain imaging and…

13:00 JST研究/論文

Spending Scarce Confirmatory PET Measurements: Target-Aligned Validation in A4/LEARN

Anti-amyloid therapies and blood-based biomarkers are changing Alzheimer disease workups into a two-stage measurement workflow: screen broa…

13:00 JST研究/論文

FreKoo++: Learning Continuous Spectral Dynamics for Temporal Domain Generalization

Temporal Domain Generalization (TDG) aims to learn from historical domains and generalize to unseen future distributions under concept drif…

13:00 JSTLLM/生成AI

Improving Few-Step Language Flows with Untied Self-Conditioning

Flow-matching language models refine all token positions in parallel and can trade sampling steps for latency, yet generation quality still…

13:00 JST画像/動画生成

Training-Free VLM Personalization via Calibrated Residual Decoding

Vision-language models can be personalized in a training-free manner by directly providing user profiles, preferences, or visual references…

13:00 JST画像/動画生成

GAN-Diff : Coupling Pretrained WGAN-GP Features with Conditional Diffusion U-Nets

Generative adversarial networks (GANs) can provide efficient image generation, while diffusion models offer high-quality image restoration…

13:00 JST研究/論文

Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

Pathological and more broadly non-canonical speech present significant challenges for automatic phoneme recognition due to systematic devia…

13:00 JSTLLM/生成AILlama

Length-Adaptive Decoding for Masked Diffusion Machine Translation

Machine translation tests masked diffusion language models (dLLMs) because every source token must be rendered faithfully, while fixed canv…

13:00 JST画像/動画生成研究/論文

OVIBench: Benchmarking Online Video Question Answering under Interruption

Recent vision language models (VLMs) have achieved strong progress in video understanding. However, most existing video QA research and ben…

13:00 JSTエージェント

Learning from the Test: Self-Referential Differential Testing for Deep RL Agents

Deep Reinforcement Learning (DRL) has achieved significant success in complex decision-making problems. As DRL systems are increasingly dep…

13:00 JSTLLM/生成AIビジネス/資金調達

LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model

Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before colle…

13:00 JSTロボティクス研究/論文

The Imitator Game: Benchmarking Robot Imitative Ability Beyond Action Prediction

Humans imitate at the level of intent: given a demonstration, we infer its goal and carry it out with whatever tools, objects, and layouts…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms

Bengali is the seventh-most-spoken language globally, yet LLM safety evaluation remains overwhelmingly English-centric. We introduce Bangla…

13:00 JST画像/動画生成

TransHands: Repurposing Human Pose Encoders as Hand Pose Encoders

Lifting 3D hand poses from 2D monocular representations remains challenging due to the limited availability of large-scale, diverse 3D-anno…

13:00 JST画像/動画生成

Multimodal examination answer data with expert-designed Outcome-Based Education rubrics for criterion-level assessment

This data article describes a multimodal collection of scanned examination answers paired with expert-designed Outcome-Based Education (OBE…

13:00 JST研究/論文

SANE: State Anomaly Neutralization for Stable Extreme-Context Delta-Rule Models

Delta-Rule recurrent models maintain a fixed-size state, enabling $O(1)$ inference memory but potentially becoming unstable under extreme-c…

13:00 JST画像/動画生成

Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video

Automatically analyzing hours-long egocentric video is increasingly essential for progress monitoring, quality control, and safety in logis…

13:00 JST研究/論文

Self-Supervised Graph Representation Learning for In-The-Wild Wearable and Smartphone based Emotion Recognition

Wearable and smartphone-based emotion recognition (WER) remains a challenging setting in affective computing, due to the notorious difficul…

13:00 JSTLLM/生成AI

ProBel: Propaganda Detection with Techniques, Spans, and Explanations

Propaganda detection includes several related prediction levels, ranging from sentence-level decisions to technique classification and span…

13:00 JST研究/論文

KONTOGRAPH: Verified Point-in-Time Feature Consistency and Amortised Explanation for Real-Time Anti-Money Laundering under a 200 ms Decision Budget

Regulation (EU) 2024/886 obliges European payment service providers to settle euro credit transfers in under ten seconds, around the clock.…

13:00 JST研究/論文

Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings

Decoding perceived speech from non-invasive brain recordings has garnered significant attention in recent years due to its wide range of po…

13:00 JSTLLM/生成AIエージェント

Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator

Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Jud…

13:00 JSTロボティクス

EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting

Forecasting dexterous hand motions from egocentric observations is fundamental to intelligent interactive systems. Existing VLM-based metho…

13:00 JST研究/論文

Functional compatibility as a determinant of persistent neural learning

Artificial neural networks can acquire new capabilities but often damage existing ones when they continue to learn. This stability-plastici…

13:00 JSTLLM/生成AI

BLADE: Bilevel Low-rank Augmented-Lagrangian Erasure for LLM Unlearning

Existing LLM unlearning methods struggle with robustness: unbounded forget losses degrade model coherence, fixed-weight balancing cannot ad…

13:00 JSTLLM/生成AI研究/論文

Hybrid Panels: Toward Human-AI Collaboration in Survey Research

Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, in…

13:00 JST研究/論文

Clinical Graph-JEPA: Predictive Patient-State Knowledge Graphs for Cognitive Decision Support

Clinical records contain rich evidence about patient state, but converting that evidence into reliable, structured knowledge graphs remains…

13:00 JSTLLM/生成AI画像/動画生成

Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

External hand forces are important inputs to biomechanical analyses of occupational physical exposure and injury risk, yet continuous force…

13:00 JST画像/動画生成研究/論文

AI-based worker guidance in assembly and disassembly operations using multimodal ego/exo-centric data capture and structured task knowledge

Assembly and disassembly processes rely on expert knowledge that is difficult to document, reuse, and transfer. This paper presents a data-…

13:00 JSTLLM/生成AI

Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains

Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especiall…

13:00 JSTLLM/生成AI

GeoRisk-RAG: A Hierarchy-Aware Risk Framework for Improving RAG Reliability through Selective Answering

Current work on improving reliability in large language model (LLM)- generated answers has primarily leveraged Retrieval-Augmented Generati…

13:00 JST研究/論文

Do Not Copy/Paste: Soft Barriers for Copying in AI-Assisted Programming

Copying a function from a chat window into an editor takes less than a second. For many uses of AI coding tools, that speed is the point; i…

13:00 JST研究/論文

Mol-JEPA: A multimodal Joint Embedding Predictive Architecture for Molecules

Despite recent advances in molecular foundation models, several limitations remain, such as chemically invalid augmentations, modality coll…

13:00 JSTLLM/生成AI

Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Code

LLMs are increasingly used for code generation, yet they frequently hallucinate non-existent software packages, creating exploitable entry…

13:00 JSTLLM/生成AIエージェントロボティクス

Physical Agentic AI: An Architecture for Orchestrating a Robot Crew with LLMs

Agentic AI frameworks interpret open-ended task goals and decompose them into multi-step plans. Richer information about embodiment-specifi…

13:00 JST画像/動画生成

Hyperbolic Hierarchical Clustering for Visual Representation Learning

We investigate the token mixer in vision backbones by revisiting clustering, one of the most classic approaches in machine learning. An eff…

13:00 JST画像/動画生成エージェントロボティクス

RACO: Reliability-Aware Coarse-Goal Optimization for Inspection-Oriented UAV Vision-Language Navigation

UAV vision-language navigation (UAV-VLN) is commonly evaluated as goal reaching, but inspection-oriented deployment requires the agent to s…

13:00 JSTLLM/生成AIエージェント

Enrich-Retrieve-Rank: Scaling Capability Discovery Beyond In-Context Routing

Agent ecosystems now include thousands of MATS components (Models, Agents, Tools, and Skills), yet their discovery still relies on in-conte…

13:00 JST研究/論文NVIDIA

TEE-X: TEE-aware Acceleration Framework for Large Vision Models at the Edge

Despite their remarkable success, machine learning models, particularly in vision applications, are alarmingly vulnerable to a range of sec…

13:00 JSTLLM/生成AI

DiaRelay: Relaying Dialogue Context with a Constant-Size Memory for Emotion Recognition in Conversation

Emotion Recognition in Conversation (ERC) requires models to identify subtle emotional cues that are often distributed across distant dialo…

13:00 JST画像/動画生成

Object-Uni: A Unified Model for Object-Centric Spatial Understanding and Controllable Generation

Unified models for visual understanding and generation have made rapid progress, yet they still lack the ability to understand and manipula…

13:00 JSTLLM/生成AIAnthropicGPT / ChatGPTGemmaLlamaDeepSeek

XTC: Head-Aware Sampling by Excluding Top Choices

Standard decoding rules for autoregressive language models promote diversity by rescaling the full next-token distribution or truncating it…

13:00 JSTLLM/生成AILlama

Don't Repeat Yourself: Stopping Verbatim Loops at Sampling Time

Large Language Models generate text autoregressively, but open-ended generation is prone to verbatim looping, in which models repeat spans…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents

Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: b…

13:00 JSTロボティクス

Triplet2Track: A Hierarchical System with Object-Centric Representations for Reliable Long-Horizon Manipulation

Ensuring reliability in uncertain environments remains difficult for long-horizon robotic manipulation. End-to-end VLA models are data-heav…

13:00 JSTLLM/生成AI

SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support

Medical large language models are often judged by how many clinical questions they answer correctly. That view is useful, but it misses a p…

13:00 JST研究/論文

Fairness-Aware Mixture-of-Experts via Subgroup Reweighting and Gate Regularization

Deep learning models often produce performance disparities across demographic groups, due to the training data imbalance with respect to se…

13:00 JSTLLM/生成AIエージェント

Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments

Large language models (LLMs) and vision-language models (VLMs) are expanding the range of behaviors that can be represented in agent-based…

13:00 JSTLLM/生成AI

Hierarchy-Aware Supervised Uncertainty Estimation for Black-box LLM Taxonomic Reasoning

Large language models (LLMs) are increasingly used for scientific decision support, yet reliable confidence estimation remains difficult in…

13:00 JST研究/論文

The Mask Is Not the Model: Auditing Prefix Invariance in Attention, State-Space, and Hybrid Sequence Models

We formalize prefix invariance: representations at position t must not depend on future inputs. We give a lightweight audit, two forward pa…

13:00 JSTLLM/生成AIGPT / ChatGPTGemini

AraDetox: A Multi-Dialect Arabic Detoxification Dataset

Arabic harmful-language detection has received considerable attention, yet Arabic text detoxification remains underexplored. We introduce A…

13:00 JSTLLM/生成AI

Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text

Spoken Language Models (SLMs) generate textual responses directly from speech, offering an alternative to cascaded systems. Despite recent…

13:00 JSTLLM/生成AIハードウェア/半導体

Safety Hacking in Constrained Best-of-$N$ Inference-time Scaling

Inference-time pipelines often sample multiple outputs, filter them with a learned safety model, and return the proxy-feasible output with…

13:00 JST研究/論文

Deep Learning-Based Multi-User Communication Design for Dense IoT Networks: Interference-Aware Finite-Blocklength Communication and Preliminary MIMO Extensions

Dense IoT networks require reliable communication despite limited spectrum and substantial multi-user interference while maintaining manage…

13:00 JSTLLM/生成AI研究/論文

What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation

Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decis…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans

While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challen…

13:00 JST研究/論文

Hypergraph Embedding Indexing for Efficient Dense Vector Retrieval

Dense vector retrieval has become the foundation of modern semantic search, yet existing approximate nearest neighbor (ANN) indexes treat a…

13:00 JST研究/論文

A Physical Response-and-Memory Model for Muon Optimization

Training large language models is costly. How low a loss the same compute can ultimately reach depends on how each step's gradient is conve…

13:00 JST画像/動画生成

Coarse Indexing, Fine Evidence: Decoupling Temporal Granularity in Long-Video RAG

Graph-based retrieval-augmented generation (RAG) provides a scalable paradigm for long-video understanding, but existing systems typically…

13:00 JSTLLM/生成AI

SplitLite: Low-Rank Residual Compression for Split Learning

Federated fine-tuning of on-device large language models (LLMs) faces a significant computing burden. To overcome this limitation, split le…

13:00 JSTLLM/生成AI

Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality

Large language models (LLMs) require effective unlearning to address privacy regulations and safety concerns. However, achieving precise fo…

13:00 JSTLLM/生成AIハードウェア/半導体

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels…

13:00 JST研究/論文

FedCC: Towards Addressing Label Distribution Skews in Distillation-Based Federated Learning

Federated Learning (FL) enables distributed clients to collaboratively train models without sharing raw data, making it promising for lever…

13:00 JSTLLM/生成AI研究/論文

The Multilingual FrameNet Corpus

This paper introduces the Multilingual FrameNet Corpus (mFNC), a novel resource that extends the English Berkeley FrameNet corpus by collec…

13:00 JSTLLM/生成AI研究/論文ClaudeGPT / ChatGPT

Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking

Misinformation that cites legitimate papers can be especially harmful when it distorts what those studies actually report. While existing a…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

Cultural understanding in video means more than recognizing what is visible; it requires grasping the symbolic and temporal significance of…

13:00 JSTロボティクスビジネス/資金調達

Shaping the Evolutionary Dynamics of Robot Morphology via Adaptive Control Learning

Robot co-design via bi-level optimization couples within-lifetime controller learning for fitness evaluation with cross-generational morpho…

13:00 JST研究/論文

PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors

Recent progress in the field of TinyML has demonstrated that low-power hardware based on microcontrollers can achieve bird species monitori…

13:00 JSTLLM/生成AIエージェント

Molecular LLM Agents: From Architectural Design to Scientific Autonomy

Molecular science represents an important frontier for LLM-based agents. Unlike general agents that mainly operate over natural language, c…

13:00 JST研究/論文

DeMixPert: Decomposed Response Modeling with Gaussian Mixtures for OOD Single-Cell Perturbation Prediction

Predicting transcriptome-wide responses to unseen genetic perturbations remains a major computational challenge because accurate prediction…

13:00 JSTLLM/生成AI

Statistical Machine Translation Systems of English-Pnar Language Pair : Some Insights of the Emperical Study

Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and…

13:00 JSTLLM/生成AI

LITERARYBIGFIVE: Author-Personalized Text Generation in a Unified Interpretable Space

Personalized text generation for authors and literary writing is essential for applications such as adaptive writing assistants, creative s…

13:00 JST画像/動画生成ロボティクス

Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation

Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creatin…

13:00 JSTLLM/生成AI

Language Chain in Alignment: Cross-Lingual Ranking Preference Optimization

The alignment of Large Language Models heavily relies on English-centric high-quality preference data, which often leads to suboptimal perf…

13:00 JST研究/論文

Counterfactual Transition Graphs: Evaluating Cross-Class Transition Quality

Counterfactual (CF) explanations for time-series classifiers are usually evaluated one example at a time: what minimal edit flips this sing…

13:00 JSTLLM/生成AIエージェント研究/論文

NetConfArena: An Executable Benchmark for LLM Agents in Closed-Loop Network Configuration

Large language model (LLM) agents are increasingly attractive for automating network configuration, yet their reliability and failure patte…

13:00 JST画像/動画生成

BenthicDINO: Physics-Informed Self-Distillation for View-Invariant Side-Scan Sonar Representations

Automated perception in side-scan sonar (SSS) imagery is severely hindered by physical acoustic artifacts, resulting in representations tha…

13:00 JST研究/論文

AI Surrogate Modeling for Real-Time Tokamak Equilibrium Prediction: Benchmarking Neural Architectures and Validation on EXL-50U

Fast and reliable plasma equilibrium prediction is essential for real-time tokamak operation and control, but conventional Grad-Shafranov (…

13:00 JSTLLM/生成AI画像/動画生成ロボティクス

Think Only When Needed: Prompt-Authority Control for Selective Slow-Path Intervention in Vision-Language-Action Manipulation

Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text bec…

13:00 JST研究/論文

Retrieval-Augmented Classification of Environmental Mitigations in Hydropower Licensing Documents

Identifying and classifying environmental mitigation obligations in Federal Energy Regulatory Commission hydropower licensing documents is…

13:00 JSTLLM/生成AIGemmaLlama

Credal Large Language Models for Semantic Commitment under Uncertainty

Large language models (LLMs) often produce fluent but incorrect answers with unwarranted confidence. A central limitation is that standard…

13:00 JST研究/論文

Multi-Winner Voting with Argumentative Ballots

We introduce multi-winner voting with argumentative ballots (MVArg) and investigate theoretical properties. As our conceptual contribution,…

13:00 JSTLLM/生成AI

Future Querying: Can LLMs Serve as Implicit Medical World Models?

Traditional clinical prediction models rely on task-specific pipelines and curated, structured data, which scale poorly and underutilize un…

13:00 JST画像/動画生成ハードウェア/半導体

E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

Vision-language models typically encode an image into hundreds of visual tokens, incurring substantial inference latency and GPU memory ove…

13:00 JST研究/論文

How Much Regularization Survives Averaging? Update Masking in Federated Learning

Federated learning on non-IID data seeks flat minima to generalize across clients, and existing methods borrow sharpness-aware minimization…

13:00 JST研究/論文GPT / ChatGPT

Sigmoid Attention as a Better Substrate for Learned KV Cache Eviction

Learned KV-cache eviction often faces a soft-to-hard mismatch: during training, differentiable gates typically attenuate token contribution…

13:00 JST研究/論文

Evaluating SAT Solver Metrics as Predictors of Human-Perceived Nonogram Difficulty

Algorithmic solver effort is often assumed to align with perceived puzzle difficulty, but this assumption is rarely tested against human so…

13:00 JSTLLM/生成AI

FIDES: A Concordance Protocol for LLM-Generated Trading Strategies

An LLM asked for a trading strategy returns three artifacts at once: a natural-language rationale, an executable implementation, and once r…

13:00 JST研究/論文

Mycelial Search: A Graph-Structured Metaheuristic for Continuous Optimisation

Continuous optimisation methods need to balance sharing information and maintaining alternative search directions. In this paper, we introd…

13:00 JST画像/動画生成エージェント研究/論文

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from t…

13:00 JSTLLM/生成AI

The Emergence of Relevance Through Axiomatic Attention Patterns During LoRA Fine-Tuning

LoRA fine-tuning is standard for adapting LLMs to reranking, but it remains unclear where in the network task-specific relevance behavior i…

13:00 JST画像/動画生成

DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize a…

13:00 JSTLLM/生成AIハードウェア/半導体

Adversarial Entropy Inflation Against Gumbel-Based Inference Verification

Gumbel-based inference verification bounds LLM weight exfiltration by only forgiving token choices that plausibly arise from honest GPU non…

13:00 JSTLLM/生成AI

Cross-lingual Biography Enrichment via Claim Extraction and Alignment

English Wikipedia is often treated as the default encyclopedic source, yet non-English Wikipedia editions can contain richer locally ground…

13:00 JSTLLM/生成AI

Cross-Domain, Multi-Task Data-to-Text Generation without In-Domain Training Data

Structured data exists in many forms (tables, knowledge graphs, charts, and time series), and converting it into text may involve different…

13:00 JST研究/論文

Towards a Densing Law for User Representation Learning at Billion-Scale Capacity

User representation learning in real-world industrial scenarios is commonly scaled by increasing user amount, behavioral sequence length an…

13:00 JSTLLM/生成AIエージェント

Right-Sizing LLM-Agent Decomposition in VAT Determination: A Pilot Controlled Sweep

Recent LLM-agent systems make conflicting design bets: decompose work across many narrow agents, or use one strong tool-using agent. This p…

13:00 JST研究/論文

Adaptive Item-based Collaborative Structures via Noise Rescheduling in Diffusion for Generative Recommendation

Discrete Diffusion Models (DDMs) have recently been introduced to recommendation systems, modeling user history as a token generation proce…

13:00 JST画像/動画生成

ChebBooster: A Training-Free Approach for Efficient Diffusion Transformer Inference via Chebyshev-Inspired Extrapolation

Diffusion Transformers (DiTs) have shown strong performance in high-fidelity image generation, but their sampling process remains computati…

13:00 JST画像/動画生成

Towards Comprehensive Basketball Understanding

Understanding a basketball game requires recognizing events, localizing actions, identifying players, and relating these to structured game…

13:00 JSTロボティクス

Reward-Free Continual Adaptation for Resilient Space Robots

Space robots operate in extreme environments where hardware degradation can critically compromise traditional control strategies. While con…

13:00 JST研究/論文

Machine Learning Assisted Inverse Design of Pixelated mmWave Patch Antennas

A machine learning-assisted framework for the inverse design of pixelated millimetre-wave patch antennas targeting the 22--30 GHz band is p…

13:00 JSTLLM/生成AIエージェント

InjecMEM: Memory Injection Attack on LLM Agent Memory Systems

Memory is becoming a default subsystem in deployed LLM agents to provide persistent personalization and continuity. This naturally prompts…

13:00 JSTエージェント

MetaCaster: Meta-Harness-Optimized Agent for End-to-End Few-Shot Learning of Lightweight Time Series Forecasters

Time series forecasting (TSF) is evolving toward multimodal and agentic settings, yet using foundation models remains uneconomical in resou…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

What's the Catch? Evaluating Temporal Consistency in Vision-Language Models

Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they captur…

13:00 JST画像/動画生成ロボティクス

Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by…

13:00 JSTLLM/生成AI研究/論文

When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World

Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly…

13:00 JST研究/論文

The Measurement Revolution? Credible Measurement and Inference in the Age of AI

Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as text and images, into s…

13:00 JST研究/論文

Adapter-Based Few-Shot Continual Learning for Malicious Packet Recognition

The continual evolution of malware variants necessitates detection systems that can adapt to new threats without retraining from scratch. H…

13:00 JSTLLM/生成AIエージェント

The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams

Does multi-agent LLM interaction help or hurt? Some work reports gains from debate (Du et al., 2024), critique loops (Chen et al., 2025), a…

13:00 JSTLLM/生成AI

ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings

Recent advances in continuous diffusion and flow-based language models (LMs) have achieved performance competitive with discrete LMs. Howev…

13:00 JST研究/論文

Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography

Ballistocardiography (BCG) is promising for unobtrusive long-term blood pressure (BP) monitoring in laboratory settings, but traditional BC…

13:00 JST画像/動画生成Gemini

EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings

Road traffic injuries remain a major challenge in low- and middle-income countries, where proactive road safety auditing is limited by inco…

13:00 JSTLLM/生成AIエージェントClaude

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

Modern software systems accumulate technical debt over decades of development, which makes migration expensive and largely manual. As codin…

13:00 JSTLLM/生成AI

How to Train a Critic Stably and Efficiently

Group-based reinforcement learning methods such as GRPO for large language models avoid training a critic by sampling multiple responses fo…

13:00 JSTLLM/生成AI

A Survey on Human-AI Collaboration with Large Foundation Models

As the capabilities of artificial intelligence (AI) continue to expand rapidly, Human-AI (HAI) Collaboration, combining human intellect and…

13:00 JST研究/論文

Memory-Enhanced Neural Solvers for Routing Problems

Routing Problems are central to many real-world applications, yet remain challenging due to their (NP-)hard nature. Amongst existing approa…

13:00 JSTLLM/生成AIエージェントLlamaQwen

Evaluating Large Language Models for automatic analysis of teacher simulations

Digital Simulations (DS) provide safe environments where users interact with an agent through conversational prompts, providing engaging le…

13:00 JST研究/論文

Online design of dynamic networks

Designing a network (e.g., a telecommunication or transport network) is mainly done offline, in a planning phase, prior to the operation of…

13:00 JST研究/論文

Neural-Symbolic Reasoning over Knowledge Graphs: A Survey from a Query Perspective

Knowledge graph reasoning is pivotal in various domains such as data mining, artificial intelligence, the Web, and social sciences. These k…

13:00 JST規制/政策

Practical Principles for AI Cost and Compute Accounting

Policymakers increasingly use development cost and compute as proxies for AI capabilities and risks. Recent laws have introduced regulatory…

13:00 JST研究/論文OpenAIGPT / ChatGPT

Benchmarking Retrieval-Augmented Generation Strategies for Large Language Model-Based Travel Mode Choice Prediction

Accurately predicting travel mode choice is essential for effective transportation planning, yet traditional statistical and machine learni…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Beyond Benchmarks: LLM Evaluation with an Anthropomorphic and Lifecycle-oriented Roadmap

Despite their rapid advancement, large language models (LLMs) suffer from a critical disconnect between benchmark scores and real-world uti…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文Gemini

Beyond Benchmarking: Scenario-Based Evaluation of Large Language Models for Personalized Learning

While large language models (LLMs) are increasingly being adopted to support personalized learning, there remains limited understanding of…

13:00 JSTLLM/生成AIエージェント

MACD: Multi-Agent Clinical Diagnosis with Self-Learned Knowledge for LLM

Large language models (LLMs) have shown promise in supporting medical diagnosis, with prompting-based methods offering a flexible and deplo…

13:00 JSTLLM/生成AI

AdaR: A Framework for Equipping LLMs with Adaptive Reasoning

Mathematical reasoning is a primary indicator of large language models (LLMs) intelligence. However, existing LLMs exhibit failures in robu…

13:00 JSTLLM/生成AIエージェントGPT / ChatGPT

Training Proactive and Personalized LLM Agents

Despite rapid progress, current AI agents are primarily optimized for isolated task completion. We argue for a paradigm shift toward traini…

13:00 JSTエージェント研究/論文

SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation

Creating presentation slides from scientific papers is not simply a matter of summarizing paragraphs. A presenter is required to decide wha…

13:00 JST画像/動画生成

VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection

Graph property detection aims to determine whether a graph exhibits certain structural properties, such as being Hamiltonian. Recently, lea…

13:00 JSTLLM/生成AI画像/動画生成エージェント

SkillNet: Create, Evaluate, and Connect AI Skills

Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systemati…

13:00 JSTエージェント

Reinforcing the World's Edge: A Continual Learning Problem in the Multi-Agent-World Boundary

In a stationary decentralized Markov game, learning peers generate an episode-indexed sequence of induced MDPs for any focal agent. The joi…

13:00 JSTLLM/生成AI画像/動画生成エージェントQwen

ATP-Bench: Towards Agentic Tool Planning for MLLM Interleaved Generation

Interleaved text-and-image generation represents a significant frontier for Multimodal Large Language Models (MLLMs), offering a more intui…

13:00 JST研究/論文

Retrieval-aligned Tabular Foundation Models Enable Robust Clinical Risk Prediction in Electronic Health Records Under Real-world Constraints

Clinical prediction from structured electronic health records (EHRs) is challenging due to high dimensionality, heterogeneity, class imbala…

13:00 JSTLLM/生成AILlama

Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence

Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent m…

13:00 JSTエージェントNVIDIA

MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction

Trajectory prediction is a key component of autonomous driving systems because future motions directly affect collision checking, behavior…

13:00 JSTLLM/生成AI

テスト時間のスケーリングにおける LLM プルーニングの有効性を再考する

大規模言語モデル (LLM) は、テスト時コンピューティング スケーリング (TTS) を通じて優れた推論機能を発揮し、数学およびコーディング ベンチマーク全体で優れたパフォーマンスを発揮するようになりました。並行して、モデル圧縮の研究では、タスクのパフォーマンスを犠牲にすることなく、冗長/有害なパラメータを削除しようとする枝刈り手法が開発されました。これら 2 つの研究の進歩が交差することで、私たちの研究の基礎が築かれます。推論 LLM に特有のこれまでの研究では、構造化プルーニング (レイヤー ブロックのセット全体を削除する方法) が TTS 推論のパフォーマンスを大幅に低下させることが示されています。ただし、この研究では、この仮定を再考し、非構造化枝刈り (特定の冗長/有害な重みのみを慎重に除去する方法) が同様の制限を示すかどうかを調査します。驚くべきことに、2 つの推論 LLM s1.1-7B と Qwen3-8B での 4 つの推論ベンチマークにわたる広範な実験では、非構造化枝刈りが構造化枝刈りに比べて TTS のパフォーマンスを向上させ、時には非枝刈りのフルウェイト LLM のパフォーマンスを上回ることさえあることが一貫して示されています。さらに、これらの非構造化手法をインスタンス化するための重要なパラメータの選択である、さまざまな層ごとのスパーシティ割り当て戦略の影響も実証的に研究しています。これらの発見は、枝刈りは常に TTS パフォーマンスを低下させるという従来の概念に疑問を投げかけ、実際、慎重に枝刈りを行うことで TTS の有効性を維持できることを示唆しています。

原文 (English)

Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling

Large Language Models (LLMs) now exhibit remarkable reasoning capabilities through test-time compute scaling (TTS), with impressive performance across math and coding benchmarks. In parallel, research in model compression has developed pruning methods that seek to remove redundant/detrimental parameters without sacrificing task performance. The intersection of these two research advancements lays the foundation for our work. Specific to reasoning LLMs, prior work has shown that structured pruning (methods which remove entire set of layer blocks), significantly degrades TTS reasoning performance. However, in this work, we revisit this assumption and investigate whether unstructured pruning (methods that carefully remove only certain redundant/detrimental weights) exhibits similar limitations. Surprisingly, our extensive experiments across four reasoning benchmarks on two reasoning LLMs: s1.1-7B and Qwen3-8B, consistently show that unstructured pruning augments TTS performance compared to structured pruning, and at times can even outperform the unpruned full-weight LLMs. Furthermore, we also empirically study the impact of different layer-wise sparsity allocation strategies, which are an important parametric choice for instantiating these unstructured methods. These findings challenge the conventional notion that pruning always reduces TTS performance and in fact, suggest that carefully undertaken pruning can retain TTS effectiveness.

13:00 JST研究/論文

FinSTaR: 時系列推論モデルによる財務推論に向けて

時系列 (TS) 推論モデル (TSRM) は、一般的な領域では有望な機能を示していますが、独特の特性を示す金融領域では一貫して失敗します。我々は、1) 単一エンティティ対複数エンティティの分析と、2) 現状の評価と将来の動作の予測を組み合わせることで、TSRM の一般的な 2x2 能力分類法を提案します。この分類法を金融領域(決定論的評価と確率論的予測の区別が特に重要である)で 10 の財務推論タスクとしてインスタンス化し、S&P 株に基づく FinTSR ベンチ ベンチマークを形成します。この目的を達成するために、各カテゴリーに合わせた個別の思考連鎖 (CoT) 戦略を備えた FinTSR ベンチでトレーニングされた FinSTaR (金融時系列思考と推論) を提案します。決定論的(つまり、観察可能なデータから計算可能)な評価については、モデルが生の価格から直接答えを導き出すことを可能にするプログラム的な CoT である Compute-in-CoT を採用しています。本質的に確率的である(つまり、観察できない要因に左右される)予測については、金融アナリストが不確実性の下で推論する方法を反映して、判断を下す前に多様なシナリオを生成するシナリオ認識型CoTを採用します。提案された手法は、FinTSR-Bench で 78.9% の平均精度を達成し、LLM および TSRM ベースラインを大幅に上回りました。さらに、4 つの能力カテゴリが共同トレーニングを通じて補完的かつ相互に強化されること、およびシナリオ認識型 CoT が標準的な CoT よりも予測精度を一貫して向上させることを示します。コードは https://github.com/seunghan96/FinSTaR で公開されています。

原文 (English)

FinSTaR: Towards Financial Reasoning with Time Series Reasoning Models

Time series (TS) reasoning models (TSRMs) have shown promising capabilities in general domains, yet they consistently fail on financial domain, which exhibit unique characteristics. We propose a general 2 x 2 capability taxonomy for TSRMs by crossing 1) single-entity vs. multi-entity analysis with 2) assessment of the current state vs. prediction of future behavior. We instantiate this taxonomy in the financial domain---where the distinction between deterministic assessment and stochastic prediction is particularly critical---as ten financial reasoning tasks, forming the FinTSRBench benchmark based on S&P stocks. To this end, we propose FinSTaR (Financial Time Series Thinking and Reasoning), trained on FinTSR-Bench with distinct chain-of-thought (CoT) strategies tailored to each category. For assessment, which is deterministic, we employ Compute-in-CoT, a programmatic CoT that enables models to derive answers directly from raw prices. For prediction, which is inherently stochastic, we adopt Scenario-Aware CoT, which generates diverse scenarios before making a judgment, mirroring how financial analysts reason under uncertainty. FinSTaR achieves 78.9% average accuracy on FinTSRBench, substantially outperforming LLM and TSRM baselines. Furthermore, we show that the four capability categories are complementary and mutually reinforcing through joint training. Code is publicly available at: https://github.com/seunghan96/FinSTaR.

13:00 JST研究/論文

Reconciling Consistency-Based Diagnosis with Actual-Causality-Based Explanations

We establish, from the point of view of Explainable AI (XAI), connections between Consistency-Based Diagnosis (CBD), on one side, and Actua…

13:00 JSTLLM/生成AIエージェント

CogniFold: コグニティブフォールディングによる常時オンのプロアクティブなメモリ

既存のエージェントの記憶は主に反応的かつ検索ベースのままであり、経験を自律的に永続的な認知構造に組織化する能力が欠けています。真の自律型エージェントを目指して、次世代のプロアクティブ アシスタント向けに設計された、脳からインスピレーションを得た「常時オン」エージェント メモリである CogniFold を紹介します。 CogniFold は、断片化されたイベント ストリームを自己出現の認知構造に継続的に折り畳んで、入ってくるイベントと蓄積された知識から徐々により高いレベルの認知をブートストラップします。私たちは、相補学習システム (CLS) 理論を 2 層 (海馬、新皮質) から 3 層に拡張し、前頭前意図層を追加することでこれを根拠にしています。意図的な制御と意思決定の拠点として前頭前野をエミュレートする CogniFold は、グラフ トポロジーの自己組織化を通じてこれを実現します。つまり、認知構造はストリームの下で積極的に集まり、意味的に類似している場合は結合し、古くなっている場合は減衰し、連想想起を通じて再リンクし、概念クラスターの密度がしきい値を超えると意図を表面化します。 CogEval-Bench を使用して構造形成を評価し、CogniFold が認知的期待と概念創発に一致する記憶構造を独自に生成することを実証します。さらに、5 つの認知ドメインにわたる 7 つの広範なベンチマークにわたって、CogniFold が従来のメモリ ベンチマークでも同時に堅牢に実行されることを検証しました。

原文 (English)

CogniFold: Always-On Proactive Memory via Cognitive Folding

Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into persistent cognitive structure. Toward genuinely autonomous agents, we introduce CogniFold, a brain-inspired "always-on" agent memory designed for the next generation of proactive assistants. CogniFold continuously folds fragmented event streams into self-emerging cognitive structures, bootstrapping progressively higher-level cognition from incoming events and accumulated knowledge. We ground this by extending Complementary Learning Systems (CLS) theory from two layers (hippocampus, neocortex) to three, adding a prefrontal intent layer. Emulating the prefrontal cortex as the locus of intentional control and decision-making, CogniFold achieves this through graph-topology self-organization: cognitive structures proactively assemble under the stream, merge when semantically similar, decay when stale, relink through associative recall, and surface intents when concept-cluster density crosses a threshold. We evaluate structural formation using CogEval-Bench, demonstrating that CogniFold uniquely produces memory structures that match cognitive expectations and concept emergence. Furthermore, across eight downstream benchmarks -- two probing long-term conversational memory (LoCoMo, LongMemEval) and six spanning other cognitive domains -- we validate that CogniFold simultaneously performs robustly on conventional memory tasks. Our code is available at https://github.com/OpenNorve/CogniFold.

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

GIM: Evaluating models via tasks that integrate multiple cognitive domains

As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA,…

13:00 JSTエージェント

TO-Agents: A Multi-Agent AI Framework for Subjective Preference-Guided Topology Optimization

Topology optimization can generate efficient structures, but designers often must manually translate qualitative intent, such as desired vi…

13:00 JSTエージェント研究/論文

ナレッジワークのベンチマークを設計およびレポートする

LLM エージェントの開発により、コーディング、研究、ヘルスケアなど、ナレッジワーク AI に関する一連の研究が増加しています。ただし、現在の知識作業の評価とベンチマークの設計は依然として従来の NLP タスクのロジックに従っています。その結果、ベンチマークのパフォーマンスが高くても、システムが実際の展開設定でナレッジ ワークを実行できることを確実に示すことはできません。このペーパーは、ベンチマーク対象のタスクがスコアに関連付けられた作業要求をどのように表すかを明示するための 3 段階のアプローチを提供します。つまり、評価対象の作業アクティビティを定義し、テストされた設定を指定し、適切な作業成果物をスコアリングします。私たちは、ナレッジワークが役割と責任、ローカルの材料とツール、下流のワークフローで使用可能なままでなければならない成果物を通じて組織化されていることを示す作業研究をレビューします。次に、これらの懸念事項をベンチマーク設計とレポート作成のガイダンスに変換します。これには、タスクを作業アクティビティにどのようにマッピングするか、テストされた設定で材料、ツール、役割、制約をどのように指定するか、システムが残した作業成果物にどのように焦点を当てるべきかが含まれます。評価対象の作業活動に名前を付け、一般的なベンチマーク タスクと区別するために、O{*}NET 職業タスク データベースから 18 の作業活動のインベントリを取得します。私たちは 3 つのベンチマーク ケース分析を通じてこのアプローチを実証します。GDPval、ノンコードの職業成果ベンチマーク。 OfficeQA Pro、最終回答によってスコア付けされる、根拠のある文書分析ベンチマーク。 APEX-SWE は、実行可能スコア付き製品を備えたソフトウェア エンジニアリング ベンチマークです。これらのケースは、ベンチマーク設計の選択が、スコアがサポートできる最も強力な作業要求をどのように形成するか、また、ベンチマーク対象のタスク、テストされた設定、スコア付けされた製品、およびより広範な作業要求の間にギャップが生じる場所を示しています。

原文 (English)

Designing Benchmarks for Knowledge Work

AI agents are moving quickly from answering isolated questions toward completing work through tools, software environments, and multi-step workflows. Much of what these systems are now asked to do is knowledge work, where information and expertise are interpreted, produced, and communicated as part of completing work. Benchmarks for this setting are usually described only by their tasks, environments, and metrics, leaving four questions implicit: what part of the work is represented, under what conditions it is tested, what work product the system is expected to leave, and what part of that product the benchmark actually evaluates. We introduce a work-centered benchmark representation with four fields: represented activity, tested setting, required work product, and evaluated result. The representation makes these choices explicit and comparable across benchmark designs. To support activity-level reporting across occupations, we derive an aim-dependent inventory of 18 work activities from O*NET task statements and report evidence on semantic coherence, algorithm sensitivity, external ontology legibility in ESCO, and human interpretability. We apply the representation to GDPval, OfficeQA Pro, and APEX-SWE. The case analyses illustrate how occupational deliverables, grounded answers, and executable state changes capture different parts of work within the same representation.

13:00 JSTエージェント

MOSAIC: 構造化されたエージェント インテリジェンスと構成のためのモジュール式オーケストレーション

自動化されたデータ サイエンスは、構造化されたモデル選択の問題です。ソリューションでは、データ変換、特徴表現、アーキテクチャ、トレーニング手順、評価プロトコル、およびタスクの改良戦略を選択する必要があります。 AutoML システムはこのプロセスの一部を自動化しますが、通常は事前定義されたパイプライン、モデル、ハイパーパラメーター空間内で検索します。 LLM ベースのエージェントは、取得、コード生成、および実行フィードバックを通じて優れた柔軟性を提供しますが、そのモデリングの決定は多くの場合構造化されておらず、検証が難しく、再利用も困難です。メモリベースのモデル選択とワークフロー構築のための構造化エージェント フレームワークである \textsc{MOSAIC} (構造化エージェント インテリジェンスと構成のためのモジュラー オーケストレーション) を紹介します。タスクとデータセットが与えられると、 \textsc{MOSAIC} はセマンティック タスク プロファイルを構築し、以前のケースとソース コード モジュールを取得して、選択されたモデリング コンポーネント、構成、インターフェイス制約、および実行要件を指定する中間表現であるブループリントを構築します。このブループリントは、モデル選択を段階的でコンテキストに基づいた検索に変え、制約のない合成ではなく、取得した証拠での LLM ベースのコード生成を基盤とします。候補モデルは実行によって検証され、診断フィードバック、トレーニング トレース、タスク メトリクス、および失敗を認識した強化学習ポリシーを使用して改良されます。私たちは金融時系列予測と生成に関して \textsc{MOSAIC} をインスタンス化します。モデルは予測精度、分布忠実度、実行信頼性、リスクやテール挙動などの下流財務基準を満たさなければなりません。 AutoML とエージェント ベースラインに対する実験では、\textsc{MOSAIC} がタスクのパフォーマンス、実行の成功、意思決定の追跡可能性を向上させることが示されており、自動化されたデータ サイエンスを構造化され、再利用可能で、実行に基づいたモデル選択として扱うことの価値が実証されています。

原文 (English)

MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition

Automated data science is a structured model-selection problem. A solution must choose data transformations, feature representations, architecture, training procedure, evaluation protocol, and refinement strategy for a task. AutoML systems automate parts of this process, but typically search within predefined pipeline, model, and hyperparameter spaces. LLM-based agents offer greater flexibility through retrieval, code generation, and execution feedback, yet their modelling decisions are often unstructured, difficult to verify, and hard to reuse. We introduce \textsc{MOSAIC} (Modular Orchestration for Structured Agentic Intelligence and Composition), a structured agentic framework for memory-grounded model selection and workflow construction. Given a task and dataset, \textsc{MOSAIC} builds a semantic task profile, retrieves prior cases and source-code modules, and constructs a blueprint: an intermediate representation specifying selected modelling components, composition, interface constraints, and execution requirements. This blueprint turns model selection into a staged, context-grounded search and grounds LLM-based code generation in retrieved evidence rather than unconstrained synthesis. Candidate models are validated by execution and refined using diagnostic feedback, training traces, task metrics, and a failure-aware reinforcement learning policy. We instantiate \textsc{MOSAIC} on financial time-series forecasting and generation, where models must satisfy predictive accuracy, distributional fidelity, execution reliability, and downstream financial criteria such as risk and tail behaviour. Experiments against AutoML and agentic baselines show that \textsc{MOSAIC} improves task performance, execution success, and decision traceability, demonstrating the value of treating automated data science as structured, reusable, and execution-grounded model selection.

13:00 JSTLLM/生成AI

潜在報酬ステアリング: 推論 LLM の認知行動を暗黙的に促進する適応推論時間フレームワーク

強力な推論は、モデルの知識だけでなく、生成中に認知行動がどのように効果的に展開されるかにも依存します。既存の手法は明示的な動作レベルの制御に依存することが多く、推論状態、タスク、モデルによって失敗や必要な修正が異なる場合の適応性が不十分になります。この目的を達成するために、我々は、認知行動を暗黙的に伝達するスパースオートエンコーダ(SAE)潜在状態を最適化することによって認知行動を促進する、適応型推論時間フレームワークである潜在報酬ステアリング(LRS)を提案します。 LRS は、事前に定義された認知行動やそこから導き出されるステアリング方向に依存するのではなく、最終的な答えの正しさによる推論トレースに基づいて潜在報酬モデルをトレーニングし、中間潜在状態の品質を推定します。推論中、報酬勾配は脆弱な潜在状態に対して状態固有の修正方向を提供しますが、報酬と信頼ゲートは報酬信号が脆弱であるとフラグを立てた状態への介入を制限します。複数の推論 LLM バックボーンとベンチマークに関する実験では、当社の推論がさまざまなベースラインよりもパフォーマンスを一貫して向上させていることが示されており、事後分析ではさらに、当社の推論が元の推論エラーを修正する良好な認知行動を暗黙のうちに促進していることが示されています。コードは https://github.com/jiakanglee/Latent-Reward-Steering から入手できます。

原文 (English)

Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs

Strong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existing methods often rely on explicit behavior-level control, making them insufficiently adaptive when failures and required corrections vary across reasoning states, tasks, and models. To this end, we propose Latent Reward Steering (LRS), an adaptive inference-time framework that promotes cognitive behaviors by optimizing the sparse-autoencoder (SAE) latent states that implicitly carry them. Rather than relying on predefined cognitive behaviors or steering directions derived from them, LRS trains a latent reward model on reasoning traces by final answer correctness to estimate the quality of intermediate latent states. During inference, reward gradients provide state-specific correction directions for fragile latent states, while a reward and confidence gate restricts intervention to states the reward signal flags as fragile. Experiments on multiple reasoning LLM backbones and benchmarks show that \ours consistently improves performance over various baselines, and post-hoc analyses further indicate that \ours implicitly promotes good cognitive behaviors that fix the original reasoning errors. Code is available at: https://github.com/jiakanglee/Latent-Reward-Steering.

13:00 JSTエージェント研究/論文

MindClaw: 精密な介入のための閉ループの具体化された精神状態推論

Theory of Mind (ToM) を使用すると、エージェントは他のアクターの信念、目標、意図について推論することができます。これは人間中心の身体的支援に不可欠です。既存の ToM ベンチマークは高度なテキスト認識とマルチモーダルな精神状態認識を備えていますが、主にオフラインの質問応答や最終的な行動の予測を評価します。これらは、具体化されたエージェントが変化する環境とのつながりを維持できるかどうか、行為者固有の信念を更新できるかどうか、推論が必要な場合を判断できるかどうか、助けが役立つ場合にのみ介入できるかどうかを完全にテストしていません。 MindPower を基盤として、ロボット中心の ToM 推論をリアルタイムの閉ループ設定に拡張し、精密な介入を伴う身体化された精神状態推論のためのフレームワークである MindClaw を導入します。 MindClaw は、マルチソース入力、信念記憶、身体化された認知トリガー スキル、精神的推論、およびアクション生成を接続し、エージェントが介入が不要な場合は沈黙を保ちながら、適切なタイミングで役立つアクションを出力できるようにします。実験によれば、直接的な VLM ベースラインはタスクの認識と介入の調整に苦労する一方、MindClaw は最高の全体的なパフォーマンスを達成し、閉ループで組み込まれた ToM 支援におけるトリガー スキルの最適化の重要性を示しています。

原文 (English)

MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention

Theory-of-Mind (ToM) reasoning enables embodied agents to understand human beliefs, goals, and intentions, but existing benchmarks mainly evaluate this ability through offline question answering or scenario-level action prediction. MindPower advances embodied ToM by introducing robot-centric reasoning from perception to action; however, it does not evaluate whether an agent can continuously interact with a changing environment and intervene only when assistance is needed. Building on MindPower, we introduce the MindHelper Challenge, which extends embodied ToM evaluation to real-time closed-loop precision intervention. An agent must continuously observe the environment, maintain actor-specific beliefs, identify when a human requires assistance, generate executable actions, and remain silent when intervention is unnecessary. We further propose MindClaw, a simple yet effective Claw-style framework that integrates an actor-specific Belief Table, embodied cognitive skills, and a Trigger-based cognitive dispatcher. Experiments show that MindClaw achieves 36.63% precise intervention rate and 14.36\% task accuracy, substantially outperforming direct VLM baselines, whose corresponding results remain below 12.05% and 3.80%.

13:00 JSTLLM/生成AI

形式数学検証における生成的報酬モデリングの期待値の調整

大規模言語モデル (LLM) は、リーン 4 などの形式的な対話型定理証明器で使用されることが増えています。強化学習または検索手法を使用してこれらのシステムを拡張するには、中間の推論ステップを評価できるプロセス報酬モデル (PRM) が必要です。既存の報酬モデルの設計では、実際的なトレードオフが明らかになります。バリューヘッド モデルは連続スコアを提供しますが、生成モデル インターフェイスを変更します。一方、生成報酬モデルはテキストの根拠を保持しますが、数値がトークン間で分割されるため、連続浮動小数点回帰との一致が不十分です。モデルのトークン分布から連続スコアを抽出しながら、表面出力を離散的に保つ報酬モデリング手順である Expected Value Alignment (EVA) を導入します。モデルは構造化された JSON 形式で整数スコアを出力し、EVA は対応するアンカー トークンのロジットに対する期待値として連続スコアを計算します。トレーニングでは、因果言語モデリングの目的と、これらの期待値に対する補助平均二乗誤差損失を組み合わせます。リーン 4 形式検証用の報酬モデルである \textit{Leibniz} で EVA をインスタンス化し、ゼロショットおよび報酬モデリングのベースラインに対して評価します。この評価では、継続的なロジットベースのスコアリングにより、生成的批評の解釈可能性を維持しながら、離散化アーティファクトが大幅に削減されることが実証されました。

原文 (English)

Expected Value Alignment for Generative Reward Modeling in Formal Mathematics Verification

Large Language Models (LLMs) are increasingly used with formal interactive theorem provers such as Lean 4. Scaling these systems with reinforcement learning or search methods requires process reward models (PRMs) that can evaluate intermediate reasoning steps. Existing reward-model designs expose a practical trade-off. Value-head models provide continuous scores but modify the generative model interface, while generative reward models preserve textual rationales but are poorly matched to continuous floating-point regression because numeric values are split across tokens. We introduce Expected Value Alignment (EVA), a reward-modeling procedure that keeps the surface output discrete while extracting continuous scores from the model's token distribution. The model emits integer scores in a structured JSON format, and EVA computes a continuous score as the expectation over the logits of the corresponding anchor tokens. Training combines the causal language modeling objective with an auxiliary mean squared error loss on these expected values. We instantiate EVA in \textit{Leibniz}, a reward model for Lean 4 formal verification, and evaluate it against zero-shot and reward-modeling baselines. The evaluation demonstrates that continuous logit-based scoring significantly reduces discretization artifacts while retaining the interpretability of generative critiques.

13:00 JST研究/論文

Zero knowledge verification for frontier AI training is possible

Frontier AI governance frameworks increasingly use cumulative training compute as the primary criterion for designating high-impact models,…

13:00 JST研究/論文GemmaLlamaQwen

どこまで小さくできますか? LoRA 金融取引における販売者情報抽出のための 270M-8B モデルの微調整

金融取引処理では、ノイズの多い短縮された銀行取引文字列から構造化された販売者情報を大規模に抽出する必要があります。現在の運用システムである LoRA で微調整された LLaMA 3.1-8B は、このタスクで 96.95% の F1 を達成していますが、80 億のパラメーター モデルを展開すると、法外なメモリ、レイテンシ、コストの制約が課せられます。より効率的な代替案を特定するために、Gemma 3 (270M、1B、4B)、Qwen 3.5 (0.8B、2B、4B)、Aya (3.35B)、および LLaMA 3.1-8B の 4 つのモデル ファミリにわたる 24 のモデル バリアントの展開に焦点を当てた調査を実施し、精度、推論スループット、トレーニング コスト、およびハードウェアの動作を体系的に評価して実稼働への適合性を評価します。 (1) LoRA ランク 8 で LLaMA 3.1-8B 微調整を再現すると、96.75% の F1 が達成され、ランク 32 のベースラインをわずか 0.20 ポイント下回りました。 (2) JSON のみのプロンプトを使用した Qwen 3.5 4B は、約半分のパラメーターを使用しながら、8B ベースラインの 0.35 ポイント以内で 96.60% F1 に達します。 (3) 0.8B Qwen 3.5 モデルは 94.75% F1 を達成し、2.5 ~ 4 倍大きいモデルに匹敵し、魅力的なレイテンシと精度のトレードオフを提供します。 (4) 思考連鎖の微調整により、ほとんどのモデルで F1 が 0.3 ~ 1.8 ポイント改善されますが、Qwen 3.5 4B は JSON のみの直接プロンプトで最高のパフォーマンスを発揮します。 (5) Qwen 3.5 Think および Nothink トレーニング テンプレートは、ほぼ同じ結果を生成します (F1 差 <0.004)。これは、構造化された抽出タスクには明示的な推論の監視が不要であることを示しています。さらに、14 の微調整されたサブ 8B モデルすべてを Databricks Model Serving エンドポイントとしてデプロイし、平均 F1 変化がわずか 0.8 ポイントで、ベンチマーク パフォーマンスが本番環境に確実に移行することを観察しました。 Cohere2 アーキテクチャに基づいた Aya 3.35B は唯一の例外であり、使用条件下で 3 ~ 5 ポイントの低下を示しています。これらの結果に基づいて、精度と遅延の要件全体にわたって導入に関する推奨事項を提供します。

原文 (English)

How Small Can You Go? LoRA Fine-Tuning 270M-8B Models for Merchant Information Extraction in Financial Transactions

Merchant information extraction turns noisy financial transaction descriptors into structured fields at production scale. Our deployed LoRA-fine-tuned LLaMA~3.1-8B reaches 96.95\% F1, but its memory and throughput motivate smaller replacements. We evaluate 23 retained fine-tuning runs plus a separately trained production reference, spanning Gemma~3 (270M--4B), Qwen~3.5 (0.8B--4B), Aya~3.35B, and LLaMA~3.1-8B across LoRA ranks, prompts, training templates, and serving environments. A rank-8 LLaMA fine-tune reaches 96.75\% F1, only 0.20 points below the rank-32 production reference. Qwen~3.5~4B with JSON-Only prompting reaches 96.60\% F1 and strict record-level exact match of 91.67\%, with a $3.8\times$ lower inverse-throughput time estimate than the rank-8 8B model. Qwen~3.5~0.8B reaches 94.75\% F1, and Qwen Think and Nothink templates differ by less than 0.004 F1. Across 14 Databricks endpoints, mean F1 change from local evaluation is $-0.0081$; Aya is the only family with a 2.7--5.1 point decline. These results show that compact fine-tuned models can preserve most extraction accuracy, but model selection must account for prompt choice, throughput, and serving-stack behavior.

13:00 JSTエージェント

自己進化する科学エージェントが一般化可能な物理的根拠に基づいた流体制御を発見

データ集約型の深層強化学習は複雑な制御ポリシーを最適化できますが、物理システムにおける科学的発見には基本的に、物理的証拠を構造化された制御アーキテクチャに結び付ける、解釈可能な推論の連鎖が必要です。ここでは、大規模な言語モデルと反復コード生成によって駆動され、厳密な解釈可能性と厳密な物理的推論を維持しながらコントローラーの構築を自動化する、自己進化する科学エージェントのワークフローを紹介します。重みを調整する代わりに、エージェントは候補戦略を物理シミュレーションに展開し、マルチモーダルな証拠から動的動作を積極的に診断し、これらの観察結果を漸進的なソースコードの改良に変換します。我々は、このフレームワークを高度に非線形の流体構造相互作用問題、つまり関節角加速度のみを使用して空間目標に到達する任務を負った、作動が不十分な 2 関節のツノザメ遊泳者について実証します。一方的なステアリング バイアスを示す推進シード ポリシーから開始して、エージェントは自律的に、すべての正規ターゲットを確実に捕捉する統合コントローラーを発見し、改良します。注目すべきことに、再トレーニングやターゲット固有の分岐を行わずに、合成された制御ポリシーは、目に見えない静的なターゲットと動的に湾曲した追跡軌道に一般化されます。監査可能な進化ログは、進行波推進、車体フレーム目標誘導、ヨーレートフィードバック、符号付き平均尾部曲率、および適応ケイデンス緩和に基づいて構築された緊急制御アーキテクチャを明らかにします。私たちの結果は、自律的な科学エージェントが、科学的発見の完全に追跡可能なプロセスを維持しながら、蓄積された物理的証拠を堅牢で数学的に読み取り可能な制御ポリシーにうまく変換できることを示しています。

原文 (English)

Self-Evolving Scientific Agent Discovers Generalizable Physically-Reasoned Fluid Control

While data-intensive deep reinforcement learning can optimize complex control policies, scientific control design in physical systems fundamentally requires an interpretable chain of reasoning that connects physical evidence to structured control architectures. Here, we present a self-evolving scientific agent workflow, driven by large language models and iterative code generation, that automates controller construction while preserving strict interpretability and rigorous physical reasoning. Instead of adjusting weights, the agent deploys candidate whitebox controllers into physical simulations, actively diagnoses dynamic behaviors from multimodal evidence, and translates these observations into progressive source-code refinements. We demonstrate this framework on a highly non-linear fluid-structure interaction problem: an underactuated, two-joint dogfish swimmer tasked with spatial target reaching in an unsteady flow using only joint angular accelerations. Starting from a target-blind propulsive seed, the agent autonomously designs and refines a unified controller that reaches a target embedded in an unsteady four-cylinder wake. Without retraining, retuning or case-specific branching, the retained controller achieves target capture across the full generalization test matrix, spanning variations in target position, rear-row geometry, cylinder count and inflow speed. The auditable evolution log reveals an emergent control architecture built upon travelling-wave propulsion, body-frame bearing guidance, phase-selective steering, corrective burst and adaptive relief. Our results show that an autonomous scientific agent can successfully transform accumulated physical evidence into a robust, mathematically readable control policy, while maintaining a fully traceable process of scientific control design.

13:00 JSTLLM/生成AIビジネス/資金調達

思い出しすぎ: メモリ拡張モデルにおけるおしゃべりの評価と軽減

永続メモリ システムは、ユーザーの信念を長期にわたって保存することで、LLM をさらに役立つものにすることを約束します。また、モデルが正確さよりもユーザーとの合意を優先し、お調子者を体系的に増幅することでモデルの正確性が低下することも示します。私たちは、ユーザーが科学、医学、道徳的推論の領域でもっともらしい誤解を表明する、合成的に生成されたマルチターン会話のベンチマークである MIST を導入して、この効果の最初の体系的な評価を実施しました。 3 つの最先端の記憶システムと 5 つのモデル ファミリーにわたるテストにより、記憶はすべての条件でお調子者行動を増幅し、コンテキスト内ベースラインよりも最大 25 倍高いお調子者率であることが明らかになりました。エラー分析では、メモリ抽出が主な原因であることが示唆されています。個別のスニペットへの非可逆圧縮により、修正コ​​ンテキストが破棄され、ユーザーの誤解がエンコードされます。これらの結果に基づいて、事実の想起において記憶システムと同等またはそれを超えながら、おしゃべりを大幅に軽減する 2 つの軽量な緩和策を提案します。

原文 (English)

Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models

Persistent memory systems promise to make LLMs more helpful by storing user beliefs over time. We show they also make models less correct by amplifying sycophancy, wherein models prioritize agreement with users over accuracy. We conduct the first systematic evaluation of this effect, introducing MIST: a benchmark of synthetically generated multi-turn conversations where users express plausible misconceptions in scientific, medical, and moral reasoning domains. Testing across three state-of-the-art memory systems and five model families reveals that memory amplifies sycophantic behavior across all conditions, with up to 40% higher sycophancy rates than in-context baselines. Error analyses suggest memory extraction as the primary culprit: lossy compression of only discrete snippets from user turns encodes user misconceptions while discarding corrective context. Based on these results, we propose three lightweight mitigations to a memory system that substantially reduce sycophancy while matching or exceeding memory systems at factual recall.

13:00 JST研究/論文

関係構造因果モデル

人工知能は、介入や反事実についての推論をサポートする因果関係のある環境モデルを持たなければなりません。また、目に見えないオブジェクトの組み合わせへの一般化をサポートする組み合わせ関係のモデルも必要です。この研究では、そのようなモデルをいつ、どのように学習できるかを正式に研究します。私たちは関係構造因果モデルを開発し、構造因果モデル (Pearl 2009) をオブジェクトとその関係が変化する設定に拡張します。まず、因果関係だけでなく、オブジェクトの目に見えない組み合わせに関する観察的なクエリに対する答えも、さらなる仮定がなければ特定できないことを示します。観察されていない交絡が存在する場合も含めて、そのような識別を可能にするために、関係因果関係グラフを定義し、記号的な識別基準を導き出します。最後に、さまざまな車、信号、歩行者を含むシミュレートされた交通シーンで非関係ベースラインよりも優れたパフォーマンスを発揮する、証明可能な正しいアプローチである関係神経因果モデルを提案します。

原文 (English)

Relational Structural Causal Models

An artificial intelligence must have a model of its environment that is causal, supporting reasoning about interventions and counterfactuals, and also combinatorial, supporting generalization to unseen combinations of objects. In this work, we formally study when and how such a model can be learned. We develop relational structural causal models, extending structural causal models (Pearl 2009) to settings where objects and their relations vary. First, we show how answers to not only causal but also observational queries about unseen combinations of objects can not be identified without further assumptions. To enable such identification--including in the presence of unobserved confounding--we define relational causal graphs and derive symbolic identification criteria. Finally, we propose relational neural causal models, a provably correct approach that outperforms non-relational baselines on simulated traffic scenes with varying cars, signals, and pedestrians.

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

既存の利点: LLM レコメンデーション システムにおけるブランド バイアスと認知操作ダイナミクス

大規模言語モデル (LLM) は、消費者が製品を見つけるための主要な方法になりつつありますが、ブランドがこの新しいチャネルでどのように競争するのかはまだ理解できません。私たちは、3 つの商用 LLM (GPT-4o-mini、Claude Sonnet、Gemini 3 Flash) にわたるスキンケア製品 (消費者が購入前に品質を簡単に判断できず、ブランドの評判に頼らなければならないカテゴリー) を使用した LLM 推奨におけるブランド ダイナミクスを研究し、検索商品の堅牢性チェックを行います。 3 つの実験で次のことがわかりました。(1) すべての製品が同じ仕様の場合、有名ブランドが 100% の確率 (IAI = 10.0) で推奨される条件付き独占ですが、この優位性は競合他社の星評価の優位性が +0.1 つ未満になると消滅します。 (2) 捏造された臨床証拠の主張を含む権威あるマーケティング言語は、各モデルの反応が異なり、+0.17 評価ポイントに等しいバイアス余剰値でこの独占を破ります。 (3) マルチブランドの GEO 競争における社会的ジレンマ: すべてのブランドが同じ最適化戦略を採用すると、ペイオフ プロキシでは個々のペイオフが +0.802 から +0.007 に低下し、参加していないブランドはテストで推奨がゼロになります。私たちの結果は、生成エンジン最適化 (GEO) がセキュリティ リスクとしてだけでなく、市場競争を形成する新たなマーケティング手法としても研究されるべきであることを示唆しています。

原文 (English)

Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems

Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in this new channel. We study brand dynamics in LLM recommendations using skincare products -- a category where consumers cannot easily judge quality before buying and must rely on brand reputation -- across three commercial LLMs (GPT-4o-mini, Claude Sonnet, Gemini 3 Flash), with a robustness check on search goods. In three experiments, we find: (1) a Conditional Monopoly where well-known brands get recommended 100% of the time (IAI = 10.0) when all products have the same specifications, but this dominance disappears with less than a +0.1-star rating advantage for a competitor; (2) authority-style marketing language, including fabricated clinical-evidence claims, breaks this monopoly at a Bias Surplus Value equal to +0.17 rating points, with each model responding differently; and (3) a social dilemma in multi-brand GEO competition: when all brands adopt the same optimization strategy, individual payoff falls from +0.802 to +0.007 in our payoff proxy, and non-participating brands receive zero recommendations in our tests. Our results suggest that generative engine optimization (GEO) should be studied not only as a security risk, but also as an emerging marketing practice that shapes market competition.

13:00 JSTエージェント

FinAcumen: 自己進化するエクスペリエンス メモリ ハーネスによる金融マルチモーダル推論

金融マルチモーダル推論では、エージェントが異種の証拠ソース間で数値計算、検索、視覚的解釈、および時間的根拠を調整する必要があります。既存のツールで拡張されたエージェントは、実行の忠実度を向上させますが、エピソード全体にわたってほぼステートレスのままであり、推論戦略と失敗パターンを繰り返し再発見します。一か八かの金融環境では、これにより、信頼性の低いツールのルーティング、ノイズの多い検索、幻覚が起こりやすい推論が発生します。我々は、ツール拡張マルチモーダル推論のための選択的経験記憶を中心とした財務推論エージェント フレームワークである FinAcumen を紹介します。 FinAcumen は、これまでの軌跡から経済的に根拠のある推論経験を蓄積し、成功した戦略と失敗から得られた注意ルールを永続的なメモリ バンクに抽出します。推論中、意味論的な関連性が調整されたしきい値を超えた場合にのみ、取得されたものは条件推論を経験しますが、無関係なメモリはフォールバック メカニズムを通じて明示的に抑制されます。決定論的な金融ツール環境により、数値計算、検索、視覚的デコード、および回答検証がさらに強化されます。4 つの金融マルチモーダル推論ベンチマークにわたって、FinAcumen は、金融特化モデルよりも凍結された 8B ビジョン言語モデルを一貫して改善し、主要な独自の汎用モデルにアプローチします。さらなる分析により、選択的経験の活性化により、検索の不確実性の下で推論の信頼性が向上することが示されています。私たちのコードは https://anonymous.4open.science/r/FinAcumen で匿名で入手できます。

原文 (English)

FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness

Financial multimodal reasoning requires agents to coordinate numerical computation, retrieval, visual interpretation, and temporal grounding across heterogeneous evidence sources. Existing tool-augmented agents improve execution fidelity, yet remain largely stateless across episodes, repeatedly rediscovering reasoning strategies and failure patterns. In high-stakes financial settings, this leads to unreliable tool routing, noisy retrieval, and hallucination-prone reasoning. We present FinAcumen, a financial reasoning agent framework centered on selective experience memory for tool-augmented multimodal reasoning. FinAcumen accumulates financially grounded reasoning experience from prior trajectories, distilling successful strategies and failure-derived cautionary rules into a persistent memory bank. During inference, retrieved experiences condition reasoning only when semantic relevance exceeds a calibrated threshold, while irrelevant memory is explicitly suppressed through a fallback mechanism. A deterministic financial tool environment further grounds numerical computation, retrieval, visual decoding, and answer verification.Across four financial multimodal reasoning benchmarks, FinAcumen consistently improves a frozen 8B vision-language model over finance-specialized models and approaches leading proprietary general-purpose models. Further analysis shows that selective experience activation improves reasoning reliability under retrieval uncertainty. Our code is available at https://github.com/CamelliaLilium/FinAcumen.

13:00 JST研究/論文

DiagFlowBench: 言語モデルがグラウンデッド診断ダイアログでプロシージャ外入力をどのように処理するかを評価する

言語モデルは、メンテナンス作業における助言システムとしての役割をますます高めています。幻覚を防ぐために、最近のシステムでは、これらのモデルを手順書に基づいて承認された手順に制限します。ただし、実際には、オペレーターのクエリはこのパスから逸脱することが多く、モデルは会話中に範囲外の入力を認識する必要があり、現在のベンチマークではほとんど優先されません。 DiagFlowBench は、消費者メーカーの 50 の産業診断フローチャートのデータセットであり、範囲外の発話との準拠を対比する 1,676 のマルチターン会話に変換されます。 10 個の商用および無差別モデルのパネルを評価すると、棄権率のばらつきが大きく、モデルは事実を捏造するのではなく、現実ではあるが文脈上不適切なステップを選択することが一般的であることが明らかになりました。このマッピングされた間違ったアドバイスに固有のもっともらしさと権威があるため、接地システムの重大な脆弱性が明らかになります。

原文 (English)

DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue

Language models increasingly serve as advisory systems in maintenance operations. To prevent hallucination, recent systems ground these models in procedural documentation to constrain them to approved steps. In practice, however, operator queries frequently stray from this path, requiring models to recognise out-of-scope inputs mid-conversation, a dynamic that current benchmarks rarely prioritise. We introduce DiagFlowBench, a dataset of 50 industrial diagnostic flowcharts from a consumer manufacturer converted into 1,676 multi-turn conversations that contrast compliant with out-of-scope utterances. Evaluating a panel of ten commercial and open-weight models reveals high variability in abstention rates, with models commonly selecting a real but contextually inadequate step rather than fabricating facts. The inherent plausibility and authority of this mapped but wrong advice exposes a challenging vulnerability for grounding systems.

13:00 JST研究/論文

MoCo-AIS: 船舶軌道の類似性計算のための対照学習フレームワーク

軌跡の類似性は、モビリティ パターンを分析する際の基本的なタスクであり、ルート パターンの抽出、モビリティの予測、異常検出などのアプリケーションに不可欠です。類似性を計算するための従来の距離ベースの測定では、高い計算コストがかかるため、軽量な学習ベースのアプローチの採用が促進されています。教師あり手法は、従来の距離測定から得られる広範なラベルに依存しており、多くの場合、これらのメトリクスを再現するため、一般化が制限されます。自己教師あり学習は、対比学習を通じてこの問題に対処しますが、統一されたフレームワークが欠けているため、一貫した軌跡の表現のために深層学習 (DL) モデルを比較することが困難になります。したがって、この論文では、正と負の軌道ペアを介した類似性学習を定式化するモーメンタムコントラスト(MoCo)パラダイムに基づいて血管軌道埋め込みを学習するための統一フレームワークであるMoCo-AISを紹介します。このフレームワーク内で、多様な航行挙動と運航条件を捕捉する大規模な現実世界の船舶追跡 AIS データセットで、主要な DL モデルの多様なセットを評価します。結果は、私たちのフレームワークが既存のベースラインよりも類似性学習を大幅に改善し、同時に軌跡表現モデルを評価するためのベンチマーク プラットフォームを提供することを示しています。

原文 (English)

MoCo-AIS: A Contrastive Learning Framework for Similarity Computation of Vessel Trajectories

Trajectory similarity is a fundamental task in analyzing mobility patterns, essential for applications such as route pattern extraction, mobility prediction, and anomaly detection. Traditional distance-based measures for computing similarity incur high computational cost, driving the adoption of lightweight learning-based approaches. Supervised methods rely on extensive labels derived from traditional distance measures and often reproduce these metrics, which limits generalization. While self-supervised learning addresses this issue through contrastive learning, it lacks a unified framework, making it difficult to compare deep learning (DL) models for consistent trajectory representation. Accordingly, this paper presents MoCo-AIS, a unified framework for learning vessel trajectory embeddings based on the Momentum Contrast (MoCo) paradigm, which formulates similarity learning through positive and negative trajectory pairs. Within this framework, we evaluate a diverse set of leading DL models on large-scale, real-world vessel-tracking AIS datasets that capture diverse navigation behaviors and operating conditions. Results demonstrate that our framework significantly improves similarity learning over existing baselines, while providing a benchmarking platform for evaluating trajectory representation models.

13:00 JSTエージェント

ラグランジュ: 一般化されたエンドツーエンド運転のための、オープンボキャブラリー、エネルギーベースのスパースフレームワーク

エンドツーエンドの自動運転を複雑なオープンワールド環境に拡張するには、異常なシナリオに一般化する知覚モデルと、運動学的に有効な軌道を生成するプランナーが必要です。既存のパラダイムは、表現効率と一般化能力の間の明確な二分法に直面しています。高密度モデル (占有ネットワークなど) は、幾何学的に堅牢ではありますが、重大な計算ボトルネックを引き起こし、高レベルの意味論的推論に苦労します。逆に、スパースなクエリベースのプランナーは効率的ですが、クローズドセット定義に依存しているため、配布外 (OOD) イベントに対して脆弱になります。最近の Vision-Language-Action (VLA) モデルはオープンな語彙推論を提供しますが、その自己回帰的で離散的なトークン生成は、車両ダイナミクスの連続的で高周波の制御要件と根本的に矛盾します。これに対処するために、マスクされた潜在場 (MLF) に基づいたオープン語彙で計算量が少ない駆動フレームワークである Lagrange を提案します。ラグランジュは、高密度ボリューム再構成や閉集合クエリ メカニズムに依存するのではなく、視覚言語モデル (VLM) を利用して、クラスに依存しないオブジェクトの提案を連続的なセマンティックなビジュアル トークンにエンコードします。無関係なエンティティを時間的にフィルタリングし、空間座標上で定義された暗黙的な連続エネルギー フィールドにアテンション トークンをデコードする、インテント駆動型マスク クロス アテンション モジュールを導入します。このエネルギー場にわたるラグランジュ作用最小化問題として意思決定を組み立てることにより、衝突回避を実行しながら車両運動学への厳密な準拠を強制します。標準 (nuScenes) ベンチマークとロングテール (CODA) ベンチマークの両方での広範なオフライン評価により、ラグランジュが堅牢で解釈可能、運動学的に実現可能なオープンワールド自律性のための有望なフレームワークを確立していることが実証されました。

原文 (English)

Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving

Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios and planners that produce kinematically valid trajectories. Existing paradigms face a distinct dichotomy between representational efficiency and generalization capacity. Dense models (e.g., occupancy networks), while geometrically robust, incur critical computational bottlenecks and struggle with high-level semantic reasoning. Conversely, sparse, query-based planners are efficient but reliant on closed-set definitions, rendering them vulnerable to out-of-distribution (OOD) events. Although recent Vision-Language-Action (VLA) models offer open-vocabulary reasoning, their autoregressive, discrete token generation fundamentally conflicts with the continuous, high-frequency control requirements of vehicle dynamics. To address this, we propose Lagrange, an open-vocabulary, computationally sparse driving framework based on Masked Latent Fields (MLF). Rather than relying on dense volumetric reconstructions or closed-set query mechanisms, Lagrange exploits Vision-Language Models (VLMs) to encode class-agnostic object proposals into continuous semantic visual tokens. We introduce an intent-driven masked cross-attention module that temporally filters irrelevant entities, decoding the attended tokens into an implicit continuous energy field defined over spatial coordinates. By framing decision-making as a Lagrangian action minimization problem spanning this energy field, we enforce strict compliance with vehicle kinematics while executing collision avoidance. Extensive offline evaluations on both standard (nuScenes) and long-tail (CODA) benchmarks demonstrate that Lagrange establishes a promising framework for robust, interpretable, and kinematically feasible open-world autonomy.

13:00 JSTLLM/生成AI

タスク固有の LLM 蒸留のスケーリング則

大規模言語モデル (LLM) は、ますます広範囲のドメインにわたって強力なパフォーマンスを実現しますが、そのスケールにより、遅延とコストの制約が重要なアプリケーションでは導入の課題が生じます。この論文では、ドメイン固有の LLM 圧縮に関する経験的なスケーリング則を導き出し、データセットのサイズ、圧縮率、監視形式、反復枝刈りスケジュールによってドメイン内および一般知識のパフォーマンスがどのようにスケールされるかを定量化します。クオンツファイナンスをアプリケーションドメインとして使用し、反復構造枝刈りの下でロジットベースの蒸留と LoRA ベースの蒸留を比較し、推論トレース上で KL ダイバージェンスの蒸留を安定化する混合思考連鎖監視損失を導入します。圧縮下ではドメイン内タスクの品質が予想通り低下しますが、一般知識のベンチマークは同じ時点よりかなり前に崩壊します。監視形式はこのトレードオフの主な要因であり、思考連鎖監視は枝刈りによって消去された一般知識を積極的に回復します。私たちは、ヘッドライン データセット FinHeadlineMix、スケーリング則の結果、およびドメイン固有の圧縮決定のための再利用可能なフレームワークを提供する実践的な推奨事項をリリースします。

原文 (English)

Scaling Laws for Task-Specific LLM Distillation

Large Language Models (LLMs) achieve strong performance across a growing range of domains, yet their scale poses deployment challenges in applications where latency and cost constraints are critical. This paper derives empirical scaling laws for domain-specific LLM compression, quantifying how in-domain and general knowledge performance scale with dataset size, compression ratio, supervision format, and iterative pruning schedule. Using quantitative finance as our application domain, we compare logit-based and LoRA-based distillation under iterative structural pruning, introducing a blended chain-of-thought supervision loss that stabilizes KL-divergence distillation over reasoning traces. In-domain task quality degrades predictably under compression while general-knowledge benchmarks collapse well before the same point; supervision format is the key driver of this tradeoff, with chain-of-thought supervision actively recovering general knowledge that pruning erases. We release the headline dataset FinHeadlineMix, scaling law results, and practical recommendations to provide a reusable framework for domain-specific compression decisions.

13:00 JSTエージェント

The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing

Autonomous AI systems are transitioning from advisory roles to autonomous ones for medication prescriptions. Recent U.S. bill H.R. 238 and…

13:00 JST研究/論文

PolyUQuest: 異種グラフ上の検証可能な構造認識型 Web RAG

既存の検索拡張生成 (RAG) システムは、Web ページをフラット テキストとして扱い、HTML にエンコードされた構造的および意味論的な信号を失います。 PolyUQuest は、ページ間のハイパーリンク トポロジ、ページ内の DOM 階層、ページ間のエンティティ関係の知識を統合する、異種グラフ上に構築された検証可能な構造認識 Web RAG フレームワークです。 2 層ルーターは、直接ブロック取得、クロスページ グラフ トラバーサル、およびマルチホップ エンティティ推論を含む、構造上のニーズに一致する 3 つの取得モードのいずれかに各クエリをディスパッチします。引用された各ブロックにはソース ページ、見出しパス、エンティティ リンクが含まれているため、すべての回答は完全に検証可能であり、ユーザーはあらゆる主張をその構造的証拠にまで遡ることができます。私たちは、4,240 ページ、31,086 の DOM ブロック、29,119 のエンティティ、および 37,680 の関係で構成される香港理工大学 (PolyU) の公式 Web サイトを、複数タイプの評価ベンチマークとともに評価します。 PolyUQuest は、回答の正確性、カバレッジ、忠実性において既存の RAG システムよりも優れており、クエリごとに消費する LLM トークンの量が大幅に少なくなります。このデモでは、引用された回答を検査し、ルーティング モード間で検索トレースを比較し、証拠グラフ パスを探索するための対話型インターフェイスが提供されます。 PolyUQuest は、PolyU で学生向け QA サービスとして導入の準備が進められています。

原文 (English)

PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs

Existing retrieval-augmented generation (RAG) systems treat web pages as flat text, losing the structural and semantic signals encoded in HTML. We present PolyUQuest, a verifiable, structure-aware web RAG framework built on a heterogeneous graph that unifies hyperlink topology between pages, DOM hierarchy within pages, and entity-relation knowledge across pages. A two-tier router dispatches each query to one of three retrieval modes matched to its structural need, including direct block retrieval, cross-page graph traversal, and multi-hop entity reasoning. Each answer carries traceable provenance: every cited block records its source page, heading path, and entity links, so users can inspect the structural evidence behind a claim. We evaluate on the official websites of the Hong Kong Polytechnic University (PolyU), comprising 4,240 pages, 31,086 DOM blocks, 29,119 entities, and 37,680 relations, together with a multi-type evaluation benchmark. PolyUQuest improves correctness, coverage, and faithfulness over the evaluated baselines while maintaining query-time token consumption comparable to ChunkRAG and substantially below the graph-based RAG baselines. The demonstration provides an interactive interface for inspecting cited answers, comparing retrieval traces across routing modes, and exploring evidence graph paths. PolyUQuest is being prepared for deployment as a student-facing QA service at PolyU.

13:00 JSTLLM/生成AI

Voltzmann MapReduce: フォーク可能なサンドボックスのパーティション関数 Reduce

局所漸近正規性 (LAN) の下で、ワーカーがサイズ $n$ のチャンクに対して発する信頼密度は、ギブス-ボルツマン測度 $\exp\{-\beta E(\theta)\}$ であり、その逆温度はサンプル サイズ $\beta=n$ です。ガウス/線形の場合は 3 つの結果が正確で、それ以外の場合は 1 次です。つまり、互いに素なチャンクは独立したボルツマン因子を持ちます。そのため、MapReduce \emph{reduce} は、文字通り読むと、モードが精度重み付け (逆分散) プーリングである分割関数 $Z=\int\prod_k h_k\,d\theta$ になります。頻度主義的整合性はゼロ温度限界 $T=1/n\to0$ です

原文 (English)

Evidence-Aware MapReduce for Forkable Compute

Snapshot-backed sandboxes make branching cheap while leaving evidence dependence unchanged. Branches can reuse a model, prompt, repository, tests, observations, or execution ancestor, so counting outputs can amplify one repeated error into high-confidence consensus. We introduce an \emph{evidence-aware reduction contract}: each worker reports an estimate, estimated information, evidence identifiers, fork lineage, and execution metadata. For independent workers estimating one common parameter, we use standard inverse-information pooling in its Gaussian/Wald form. The fixed-dimensional numeric summary can merge in any tree order; evidence IDs and lineage follow separate rules. The residual $\Delta$ measures disagreement, becomes Cochran's $Q$ in the scalar inverse-variance case, and appears in the product integral. A reference implementation validates serialized records, rejects repeated nonempty evidence identifiers, carries evidence and lineage through tree reduction, and uses Cholesky-based numerical linear algebra. Unit tests and seeded synthetic checks exercise the algebra, unequal information, and forged precision; one four-worker named-snapshot trace exercises the end-to-end path. Platform logs document the exercised execution paths. A central open systems challenge is to turn evidence identity and fork lineage into a dependence model for correlated and adaptively selected AI branches.

13:00 JST研究/論文

OriginBlame: AI トレーニング データセットのレコード レベルおよびトークン レベルのデータ出所

データ投稿者が削除を要求すると、モデル トレーナーは現実的なギャップに直面します。学習解除アルゴリズムには忘却セットが必要ですが、特定の作成者に属するトレーニング レコードを特定できるツールはありません。既存の来歴システムはファイルまたはデータセット レベルで動作し、壊滅的な過剰削除を強いられます。私たちは、データ処理パイプラインを通じて作成者の身元を伝播し、決定論的なクエリを通じて失効リクエストを正確な忘却セットに解決する、レコードおよびトークンレベルのデータ来歴システムである ob を紹介します。 219,555 の Wikipedia ページの評価では、レコード レベルの出自によりデータセット レベルの過剰削除 (101 倍から 1.3 倍) が排除される一方、統合により Wiki データに 1.3 ~ 4.0% (HuggingFace) のスループット オーバーヘッド、および 2.1 ~ 19.0% (Datatrove) のスループット オーバーヘッドが追加されることが実証されました。 1.7B モデルでは、来歴ベースの忘却セットにより、ランダムなベースラインと比較して、未学習が 42% 改善されます。

原文 (English)

OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can locate which training records belong to a given author. Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion. We present ob, a record- and token-level data provenance system that propagates author identity through data processing pipelines and resolves revocation requests into precise forget sets via deterministic queries. Evaluation on 219,555 Wikipedia pages demonstrates that record-level provenance eliminates dataset-level over-deletion (from 101x to 1.3x), while integration adds 1.3-4.0% throughput overhead (HuggingFace) and 2.1-19.0% (Datatrove) on wiki data. On a 1.7B model, provenance-based forget sets consistently reduce the collateral damage of machine unlearning (retain perplexity) relative to same-size random baselines across all evaluated authors, with membership-inference tests indicating they select genuinely memorized content.

13:00 JSTLLM/生成AIエージェント

MiniCache: 効率的な LLM 推論のための小規模モデル インターフェイスを使用した再利用可能なプログラム キャッシュ

大規模言語モデル (LLM) は、プログラム支援推論、エージェントによる意思決定、構造化タスクの実行にますます使用されていますが、これらのアプリケーションでは多くの場合、高い推論コストが発生します。私たちは、Program-of-Thought (PoT) プログラムをパラメータ化されたキャッシュ オブジェクトに変換し、構造的に類似したリクエスト全体で再利用可能な計算を可能にする、再利用可能なプログラム キャッシュ フレームワークである MiniCache を紹介します。 MiniCache は、キャッシュ ヒット リクエストのセマンティック変数抽出とターゲット LLM 生成中の投機的ドラフトに同じ小さなモデルを再利用し、タスクの品質を維持しながら高価なターゲット LLM の呼び出しを削減します。ショッピング スタイルのリクエスト データセット、WebShop、Formula、および CodeTAT-QA に関する実験では、MiniCache が推論レイテンシー、キャッシュの再利用、精度の間のトレードオフを改善し、並列処理下で最大 3.1 倍の低いレイテンシーと 2.8 倍の高いスループットを達成することを実証しています。これらの結果は、小さなモデルが、大きなモデルの代替としてではなく、信頼性が高く効率的な再利用可能なプログラム キャッシュを可能にする軽量のインターフェイス モデルとして最も効果的であることを示しています。

原文 (English)

CacheSpec: Finding the Sweet Spot for Small Models in Large Language Models

Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but these settings often incur substantial inference cost. Many such requests share similar computational structures while differing in variables, constraints, or contexts, creating opportunities for program-level caching. Since program caches need to reapply reusable computation logic to new requests, their key steps often involve lightweight and structured operations such as variable extraction, program binding, and generation acceleration, which are well suited for small models. We propose CacheSpec, an inference optimization framework centered on reusable program caches. The framework converts Program-of-Thoughts (PoT)-style programs from one-time reasoning artifacts into reusable cache objects, and reuses the same small model for two roles: semantic variable extraction on the cache-hit path and speculative drafting during target-LLM generation. Experiments on shopping-style request datasets, WebShop, Formula, and CodeTAT-QA show that CacheSpec reduces inference latency and improves effective cache reuse while preserving comparable or better task quality than existing caching and generation baselines, achieving up to about 3.1$\times$ latency speedup; in parallel serving experiments, it improves throughput by about 2.8$\times$ over PoT-style methods. These results suggest that the sweet spot for small models in large-model inference systems lies not in solving complex tasks independently, but in performing lightweight, structured, and verifiable auxiliary operations.

13:00 JSTLLM/生成AIエージェント研究/論文

Telco-GAIA: 通信ドメインのエージェント向けのバイリンガル ベンチマーク

Telco-GAIA は、実際の電気通信事業者のデータに基づいてツールを使用するエージェントを評価するための、バイリンガルでマルチモーダルなベンチマークです。 Telco-GAIA は、英語とアラビア語で人間が検証した 100 の質問応答タスクで構成されており、それぞれのタスクで 3 つの異種ソース (静的 Web サイトのスナップショット (HTML、画像、リンクされた PDF)、合成リレーショナル SQL データベース、テキスト、画像、および表形式のモダリティにわたる外部 Web アーカイブ) にわたるマルチホップ推論 (平均 4.2 ホップ) が必要です。このベンチマークはサンドボックス化された Docker 環境として提供され、正規化された正確な文字列一致によってスコア付けされるため、評価は客観的で決定的であり、LLM-as-a-Judge なしで時間の経過とともに再現可能になります。 12 の商用およびオープン LLM にわたって専用のリファレンス エージェントを評価したところ、Telco-GAIA は困難であることがわかりました。最も強力なモデルでもタスクの 71% しか解決できません。適度なコスト予算の下では、これは約 40% に下がり、視覚に基づいたカテゴリは依然として最も弱く、バックエンドの平均スコアは 30% を下回っており、ドキュメントと画像の理解にはかなりの余地が残されています。 Telco-GAIA は、エンタープライズ エージェント向けの厳密で再現可能なテストベッドと、クローズド ドメインのベンチマークを構築するためのテンプレートを提供します。

原文 (English)

Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain

We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications operator. Telco-GAIA comprises 100 human-verified question-answering tasks, in English and Arabic, that each demand multi-hop reasoning (4.2 hops on average) over three heterogeneous sources: a static website snapshot (HTML, images, and linked PDFs), a synthetic relational SQL database, and external web archives, spanning text, image, and tabular modalities. The benchmark is delivered as a sandboxed Docker environment and scored by normalized exact string matching, making evaluation objective, deterministic, and reproducible over time without any LLM-as-a-Judge. Evaluating a purpose-built reference agent across twelve commercial and open LLMs, we find Telco-GAIA challenging: even the strongest model solves only 71% of tasks; under a moderate cost budget, this falls to about 40%, and the visually grounded categories remain the weakest, where the average backend scores below 30%, leaving substantial headroom in document and image understanding. Telco-GAIA offers a rigorous, reproducible testbed for enterprise agents and a template for constructing closed-domain benchmarks.

13:00 JSTLLM/生成AIエージェント

CMI-Mem: CMI拡張強化学習による一般化可能な長期記憶管理に向けて

メモリ マネージャー モデルは、エージェント システムにおいて極めて重要です。既存の方法は主に LLM が判断した合成質問と回答 (QA) のペアに依存しており、メモリの評価はサンプリングされたクエリと下流のリーダーに依存します。この制限に対処するために、下流の QA の正確性と固有の条件付き相互情報 (CMI) を組み合わせたハイブリッド報酬を備えた強化学習 (RL) ベースの軽量メモリ マネージャー モデル \textbf{CMI-Mem} を提案します。 CMI は、サンプリングされた QA クエリを条件付けせずに、新しい会話入力によって提供される情報を現在のメモリ状態と比較して評価します。これにより、QA の基礎を置き換えるのではなく補完します。コードは https://github.com/Wyb0627/CMIMem で入手できます。CMI-Mem-4B モデル チェックポイントは https://www.modelscope.cn/models/wyb0627/CMIMem-4B で入手できます。

原文 (English)

CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning

Memory Manager models are pivotal in agent systems. Existing reinforcement-learning methods commonly use LLM-judged synthetic question-answer (QA) pairs: this provides useful downstream task grounding, but values memory through a sampled query distribution and a fixed reader. We propose CMI-Mem, a lightweight RL memory manager with a hybrid reward. Its extrinsic QA term measures end-task correctness, while its intrinsic Conditional Mutual Information (CMI) term evaluates the information contributed by new conversational inputs relative to the current memory state without conditioning on a sampled QA query. The two signals are complementary: QA anchors task utility, whereas CMI provides per-operation supervision for relevant, non-redundant memory construction. Experiments demonstrate improved transfer across memory-use scenarios, together with more efficient training and inference from the per-operation CMI signal. Our codes are available at: https://github.com/Wyb0627/CMIMem , and the CMI-Mem-4B model checkpoint is available at: https://www.modelscope.cn/models/wyb0627/CMIMem-4B

13:00 JSTLLM/生成AIエージェントビジネス/資金調達研究/論文DeepSeek

RSMeM: 体系的な評価によるリモート センシング エージェントの知識強化型メモリ進化

地球科学の研究には、リモート センシング (RS) 観測が重要な基盤として、複雑な分析と専門知識が必要です。ただし、汎用 LLM 上に構築された既存の RS エージェントは依然としてドメインにほとんど依存しないため、ワークフローが脆弱でエラーが発生しやすくなります。さらに、これらの失敗がその後の分析のために再利用可能なエクスペリエンスに統合されることはほとんどありません。この問題に対処するために、事前に抽出されたドメイン知識で RS エージェントをブートストラップし、オンライン エクスペリエンスを反復的に統合して堅牢なマルチステップ ツールを実行する、知識強化メモリ進化メカニズムである RSMeM を導入します。 RSMeM は 2 つのコンポーネントで構成されます。(i) 階層的知識グラウンディング。計画とツールの選択をガイドするために、階層的ドメイン コーパスに対して分類を意識した検索を実行します。 (ii) 障害を認識したエクスペリエンス改良。障害の注釈が付けられたツール使用トレースを、次のラウンドのツール実行のための再利用可能な制約に抽出します。これら 2 つのプロセスを繰り返し採用することで、RS エージェントはタスク レベルのドメイン知識を吸収し、それをインスタンス レベルの実行エクスペリエンスに効果的に変換できるように進化できます。 EarthBench での広範な実験により、RSMeM がさまざまな LLM バックボーンのセットにわたってツール使用パフォーマンスとエンドツーエンドの回答を一貫して向上させることが実証されました。特に、RSMeM は DeepSeek-V3.2 で 1% 未満の追加エクスペリエンス トークンで 6% の精度向上を達成しており、蒸留されたエクスペリエンスの強力な知識密度を示しています。私たちのコードは https://github.com/AI9Stars/RSMeM で入手できます。

原文 (English)

RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation

Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, existing RS agents built on general-purpose LLMs remain largely domain-agnostic, resulting in brittle and error-prone workflows. Moreover, these failures are seldom consolidated into a reusable experience for subsequent analyses. To address this issue, we introduce RSMeM, a knowledge-enhanced memory evolution mechanism that bootstraps RS agents with pre-distilled domain knowledge and iteratively integrates online experience for robust multi-step tool execution. RSMeM is composed of two components: (i) Hierarchical Knowledge Grounding, which performs taxonomy-aware retrieval over a hierarchical domain corpus to guide planning and tool selection; and (ii) Failure-Aware Experience Refinement, which distills failure-annotated tool-use traces into reusable constraints for next-round tool execution. By iteratively employing these two processes, RS agents can evolve to absorb task-level domain knowledge and effectively translate it into instance-level execution experience. Extensive experiments on EarthBench demonstrate that RSMeM consistently improves tool-use performance and end-to-end answer across a diverse set of LLM backbones. Notably, RSMeM achieves a 6% accuracy improvement on DeepSeek-V3.2 with less than 1% additional experience tokens, demonstrating the strong knowledge density of our distilled experience.

13:00 JSTLLM/生成AILlama

Crossing the Margin Cliff: Toward Relearn-Robust LLM Unlearning via Margin Calibration

Large language model unlearning is consistently fragile under relearn attacks. On TOFU, fine-tuning on twenty forget examples substantially…

13:00 JST研究/論文

A Cross-Architecture Audit of Direction-Based Inference-Time Defences in Vision-Language Models

Inference time defences against vision language model jailbreaks often subtract a calibrated direction from the residual stream at a chosen…

13:00 JST研究/論文GPT / ChatGPT

Nova: An End-to-End MLIR Compiler for Deep Learning

The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underly…

13:00 JST研究/論文

ドメイン固有言語によるニューラル PDE ソルバーの自動設計の改善

ニューラル PDE ソルバーの自動設計は、基本的に検索空間表現の問題です。制限のない Python プログラムの空間では、有効なソルバーは非常にまばらなサブセットを形成します。ほとんどの候補プログラムは、構文的に間違っているか、意味的に互換性がない、または数値的に不安定です。したがって、直接コード生成では、LLM はソルバーの品質について推論するのではなく、実装の失敗を回避するために検索能力のほとんどを費やすことになります。 ADSL-PDE は、ソルバーの概念と実行可能コードの間に構造化された検索状態を導入することで、この課題に対処します。これは、低レベルの実装の詳細を抽象化しながら、ニューラル PDE ソルバー (アーキテクチャ、物理的制約、目的、サンプリング、最適化) を決定する機能的な決定を表します。決定論的コンパイラは、有効な検索状態をそれぞれ実行可能なソルバーにマップします。実際、ADSL-PDE は検索空間を再構成します。つまり、無効なプログラムの大部分が削除され、意味のある候補の密度が増加し、これまでに見たことのない設計を発見するために必要な構成の自由度が維持されます。したがって、ソルバーの進化は、コード成果物ではなく、設計上の決定に基づいて機能します。この表現に基づいて構築された進化エージェントは、経験的フィードバックを使用してソルバー検索状態を繰り返し提案、評価、および改良します。複数の PDE ベンチマーク全体で、ADSL-PDE は検索効率と最適化の安定性の両方を向上させ、最初の 10 回の進化反復内で 52% 以上の改善を達成しました。これらの結果は、LLM 駆動の自動設計のより広範な原則を示唆しています。効果的なエージェントには、単により強力な推論が必要ではなく、有効で結果的な決定の探索に集中する検索表現が必要です。

原文 (English)

Improving Auto-Design of Neural PDE Solvers with a Domain-Specific Language

Neural PDE solver auto-design is fundamentally a search-space representation problem. In the space of unrestricted Python programs, valid solvers form an extremely sparse subset: most candidate programs are syntactically incorrect, semantically incompatible, or numerically unstable. Direct code generation therefore forces an LLM to spend most of its search capacity navigating implementation failures rather than reasoning about solver quality. ADSL-PDE addresses this challenge by introducing a structured search state between solver concepts and executable code. It represents the functional decisions that determine a neural PDE solver (architecture, physical constraints, objectives, sampling, and optimization) while abstracting away low-level implementation details. A deterministic compiler maps each valid search state to an executable solver. In effect, ADSL-PDE reshapes the search space: it removes large regions of invalid programs, increases the density of meaningful candidates, and preserves the compositional freedom needed to discover previously unseen designs. Solver evolution can thus operate over design decisions rather than code artifacts. Built on this representation, our evolutionary agent iteratively proposes, evaluates, and refines solver search states using empirical feedback. Across multiple PDE benchmarks, ADSL-PDE improves both search efficiency and optimization stability, achieving an improvement of more than 52% within the first ten evolution iterations. These results suggest a broader principle for LLM-driven auto-design: effective agents do not merely require stronger reasoning, but rather a search representation that concentrates exploration on valid and consequential decisions.

13:00 JSTLLM/生成AI研究/論文

NL2SHACL-Bench: 自然言語から SHACL への変換のためのベンチマーク スイート

SHACL は、RDF ナレッジ グラフ (KG) の適合性を検証するためのコア テクノロジーです。ただし、SHACL シェイプの作成には、ほとんどのドメイン専門家にはない技術的な専門知識が必要です。自然言語の要件を SHACL (NL2SHACL) に変換すると、この障壁が低くなります。ただし、NL2SHACL 専用のベンチマークはなく、意味的に同等の形状はシリアル化や構造が異なる可能性があるため、生成された形状を評価するには文字列比較を超える方法が必要です。これらの課題に取り組むために、自然言語から SHACL への翻訳のベンチマーク スイートである NL2SHACL-Bench を紹介します。 NL2SHACL-Bench を使用して、このタスク用に 4 つの最先端の大規模言語モデル (LLM) を評価します。私たちの結果は、現在の LLM は構文的に有効な SHACL を生成する能力は高いものの、複雑な論理的および構造的パターンに対して意味的に同等の制約を生成するのに依然として苦労していることを示しています。これは、NL2SHACL-Bench が、NL2SHACL の最先端技術の進歩を測定するための有意義な基盤を提供していることを示しています。

原文 (English)

NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation

SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical expertise that most domain experts lack. Translating natural language requirements into SHACL (NL2SHACL) would lower this barrier. However, there is no dedicated benchmark for NL2SHACL, and evaluating generated shapes requires methods beyond string comparison, as semantically equivalent shapes can differ in serialisation and structure. To tackle these challenges, we present NL2SHACL-Bench, a benchmark suite for natural language to SHACL translation. Using NL2SHACL-Bench, we evaluate four state-of-the-art large language models (LLMs) for this task. Our results show that current LLMs are highly capable of generating syntactically valid SHACL, but still struggle to produce semantically equivalent constraints for complex logical and structural patterns. This indicates that NL2SHACL-Bench provides a meaningful basis for measuring advances in the NL2SHACL state of the art.

13:00 JST研究/論文

Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding

Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a ques…

13:00 JSTLLM/生成AIエージェント研究/論文

The Greatness of Science Cannot Be Planned: Agentic Auto-Research is Fuzz Testing

Agentic auto-research is emerging, but most systems treat scientific discovery as goal-oriented optimization against a final benchmark. Thi…

13:00 JSTエージェント

支援 AI エージェントの階層的構成性

AI エージェントは、さまざまなアプリケーションで人間を支援するためにますます開発されており、大規模言語モデルやその他のディープ ネットワーク アーキテクチャは、そのようなエージェントにとって最先端のものであると考えられています。これらの方法は優れた確率的予測子ですが、リソースを大量に消費し、不透明であり、基礎となる表現と処理の選択肢が狭いため、新しい状況では恣意的な決定を下すことが知られています。私たちの研究は、AI の初期の先駆者にまで遡ることができるものの、現代の AI 手法では十分に活用されていない中心的な原則に基づいて、そのような AI エージェントのアーキテクチャの設計を探ることを目指しています。この論文では、人間の参加者が参照するオブジェクトのあいまいさに対処する AI エージェントの中核問題の文脈でこれを行います。人間は、ドメインのコンテキストと他の人間の参加者の好みに関する構成的な知識をヒューリスティックに活用することで、このような曖昧さに対処します。この観察からインスピレーションを得て、階層的構成性の原理を組み込み、単純なヒューリスティックを使用して目的の曖昧性をなくすアーキテクチャを説明します。具体的には、ドメイン オブジェクトは、人間が検証した意味論的特徴規範から引き出されたプリミティブな属性、および支援エージェントと特定のユーザーとの対話の限られた観察履歴から自動的に識別される属性と概念の階層的な組み合わせの観点から表現されます。次に、支援エージェントは、この構成階層の知識に基づいて推論することによって、望ましい曖昧さを解消します。ドメインダイナミクスを支配する公理。意味的な互換性、セッションの顕著性、およびユーザー固有のテーマの好みのモデルがあり、必要に応じて人間による説明が求められます。実験によれば、私たちのアプローチは常に最先端のデータ主導ベースラインを上回り、特定のユーザー プロファイルへの適応をサポートしていることが示されています。

原文 (English)

Hierarchical Compositionality for An Assistive AI Agent

AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network architectures are considered to be state of the art for such agents. These methods are impressive stochastic predictors, but they are resource-hungry, opaque, and known to make arbitrary decisions in novel situations due to the narrow set of underlying representation and processing choices. Our work seeks to explore the design of architectures for such AI agents based on core principles that can be traced back to the early pioneers of AI but are not fully utilized in modern AI methods. We do so in this paper in the context of the core problem of AI agents addressing ambiguity in the objects being referred to by the human participants. Humans address such ambiguity by heuristically leveraging compositional knowledge of domain context and the preferences of the other human participants. Drawing inspiration from this observation, we describe an architecture that embeds the principle of hierarchical compositionality and uses simple heuristics to achieve the desired disambiguation. Specifically, domain objects are represented in terms of primitive attributes drawn from human-validated semantic feature norms, and a hierarchical combination of attributes and concepts automatically identified from a limited observed history of interactions of an assistive agent with specific users. The assistive agent then achieves the desired disambiguation by reasoning with knowledge of this compositional hierarchy; axioms governing domain dynamics; and models of semantic compatibility, session salience, and user-specific thematic preference, requesting human clarification when necessary. Experiments show that our approach consistently outperforms state of the art data-driven baselines, supporting adaptation to specific user profiles.

13:00 JST研究/論文

推論の近道と価値の対称性: 対称性が許可するもの、アーキテクチャが実現するもの、最適化が選択するもの

推論のショートカットは、意図しない概念を通じて正しい予測を生み出す神経象徴システムのルールの解決策です。竹村、井上、西野の最近のフレームワークは、値の再ラベル化の自己同型グループを通じてそれらを分析し、その中心的な未解決の質問として、ルールが概念を固定するのはいつなのかを尋ねます。まず、フレームワークの主要な定義である、すべての位置に適用される 1 つの共有順列が、評価対象となった 4 つの異種ベンチマークのいずれにも前述のように適用されないこと、およびドメインを共通のサイズにパディングする最も直接的な埋め込みが、確信を持って誤った病理を生み出すことを示します。CLE4EVR では、ソリューション ペアの 90.91% が説明されていないと報告され、導入した階層の明確に定義されたすべてのメンバーが 0% を報告し、パディングされた判定の内容が報告されています。構成ファイルの順序に従ってローテーションします。事前に指定された 15 の予測 (13 は確認済み) に基づいて 11 のルール ファミリを再測定し、説明のつかないペアの割合は 0% から 99.9999% に及び、証明可能な構造を追跡します。6 つの定理は、構文のみからカンディンスキーの病理を証明するフリー スロット補題を含め、推移性とその失敗に対する十分な条件を提供します。回路によって与えられたルールの場合、座標の対称性の不活性性の決定は coNP 完全です。非自明自己同型の存在は、ランダム化還元の下では coNP 困難であり、$\Sigma_2^p$ に存在し、PH が崩壊しない限り $\Sigma_2^p$ 完全ではなく、単調回路上では完全に coNP 完全です。ブール値の場合、推移性は正確に分類されます。解セットがアフィン剰余類であれば、自己同型性によりすべてが説明されます。弱教師モデルでは、観測された 94 個のショートカットすべてがコンポーネントごとの理論フラグの 1 つのレベルに配置され、推移的であると証明される 48 個のレベルには配置されません。 12 の型付きあいまいなレベルでは何も生成されず、許容される対称性と最適化によって選択されるものが分離され、デュアルヘッド コントロールが地理を複製します。すべての数値はリリースされたアーティファクトに遡ります。

原文 (English)

Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects

Reasoning shortcuts are rule solutions that reach correct predictions through unintended concepts. A recent framework of Takemura, Inoue, and Nishino analyzes them through an automorphism group of value relabelings, asking when rules pin concepts down. Its key definition, one value permutation shared across all positions, does not apply as stated to any of its four heterogeneous benchmarks, and the most direct embedding, padding, produces confident false pathology: 90.91% of solution pairs unexplained on CLE4EVR, versus 0% under every well-defined rung of the componentwise hierarchy we introduce; the padded verdict rotates under configuration-file ordering. Across eleven rule families under fifteen pre-specified predictions (thirteen confirmed), unexplained-pair rates span 0% to 99.9999% and track provable structure: six theorems give sufficient conditions for transitivity and its failure. For circuit-given rules, symmetry-inertness of a coordinate is coNP-complete; automorphism existence is coNP-hard under randomized reductions, lies in $\Sigma_2^p$, is not $\Sigma_2^p$-complete in the Boolean case unless PH collapses, and is coNP-complete on monotone circuits. Boolean transitivity is classified exactly: automorphisms explain everything iff the solution set is an affine coset. Weakly supervised models place all 94 observed shortcuts at the one level the theory flags, none at the 48 it certifies transitive, and none at twelve typed-ambiguous levels. Relocating the absorbing element moves every shortcut with it; a confusion null attributes the location to geometry while the observed rate exceeds it by half again. Trained end to end on CLE4EVR's rule and heterogeneous domains through a synthetic prototype front end, models produce 20,223 label-preserving errors with zero different-orbit exceptions, as transitivity predicts, where the padded instrument would misreport 78-88% of them.

13:00 JSTエージェント

InfraBench: レイヤー、ライフサイクル、リスク全体にわたるインフラストラクチャ エージェントの評価

最新のコンピューティング インフラストラクチャの管理は、ますます複雑になるため、ますます困難な問題になっています。 AI エージェントの最近の進歩により、インフラストラクチャ管理タスクを自動化するタイムリーな機会が生まれていますが、そのようなエージェントが現実世界のインフラストラクチャの複雑さをどの程度うまく処理できるかは依然として不明です。ここでは、システム スタック全体と運用ライフサイクル全体にわたる現実的なインフラストラクチャ タスクに関する AI エージェントを、きめ細かいリスク評価で評価するためのベンチマーク スイートである InfraBench を紹介します。 15 のエージェント モデル構成での実験では、最も強力なエージェントでもすべてのタスクにわたってフル スコアを確保できないことがわかりました。平均有効スコアは約 40% ~ 88% の範囲であり (構成ごとの標準誤差は 6 ~ 12 ポイント)、すべてのタスクを 3 回繰り返すと、最上位の構成はまだ試行の一部しか合格していないことが明らかになり、チェックごとのスコアリングによって一般的な失敗パターンが明らかになります。エージェントは、非永続的な変更、壊れた分散不変条件、安全でない副作用、およびクリーンアップされていない状態を残しながら、短期的な目標を定期的に満たす可能性があります。 IFRABENCH は、ライブ リーダーボード、タスク、評価ハーネスを含めて、infraben.ch で公開されています。

原文 (English)

InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk

Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI agents create a timely opportunity to automate infrastructure management tasks, but it remains unclear how well such agents can handle real-world infrastructure complexity. We present InfraBench, a benchmark suite for evaluating AI agents on realistic infrastructure tasks across the full system stack and full operational lifecycle with fine-grained risk assessment. Experiments with 15 agent-model configurations show that even the strongest agent cannot secure a full score across all tasks. Mean effective scores range from roughly 40% to 88% (with per-configuration standard errors of 6-12 points), repeating every task three times reveals that top configurations still pass only a fraction of their attempts, and per-check scoring exposes a general failure pattern: agents may routinely satisfy short-term objectives while leaving non-durable changes, broken distributed invariants, unsafe side effects, and uncleaned state behind. INFRABENCH, including its live leaderboard, tasks, and evaluation harness, is publicly available at infraben.ch.

13:00 JSTLLM/生成AIビジネス/資金調達

ギザギザの裁判官: 沈黙、圧力、執拗な状況下での認識的安定性

LLM 審査員は、モデルの評価、オンライン採点、報酬モデリングの中心的なインフラストラクチャとなっています。裁判官は通常、ゴールデンデータの正確性によって検証されますが、再プロンプト、異議申し立て、または持続的な反発の下で裁判官が安定しているかどうかについては、正確性はほとんど影響しません。私たちは、LLM 裁判官の認識安定性を評価するための統一ストレス テストである \emph{Wiggle Framework} を導入します。このフレームワークは、機械的一貫性 (再プロンプトと再フレーム化の下での安定性)、シングルターン確信 (単一の課題の下での安定性)、およびマルチターン持続性 (持続的または適応的なプレッシャー下での安定性) の 3 つの次元に沿って判断の堅牢性を分解します。私たちはこのフレームワークを使用して、安全性、毒性、AI 書き込み検出、政治的対応評価にわたる 14 の審査タスクにわたって 9 つのフロンティア モデルを研究します。すべてのモデルは、裁判官としてかなりの動きを示します。静的なプッシュバックでは 25 ~ 71\% の確率で評決を覆し、敵対的な LLM 説得では 62 ~ 91\% の確率で評決を覆します。重要なことに、裁判官の評決を変えることに成功する圧力は、ほとんどの場合、グラウンドトゥルースに関してネットを破壊するものであることがわかります。フレームワーク自体を超えて、私たちは、どの項目が変動するかを予測するための最も効果的な単発シグナルとして、ベースラインの陪審過半数の強さを特定します。総合すると、これは、判定のコンテキストにおける機械的テスト、適合性テスト、および説得力テストのデータセット間での初めての同一の比較です。

原文 (English)

Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence

LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and political-response evaluation. Every model exhibits substantial wiggle as a judge --- flipping verdicts 25--71\% of the time under static pushback, and 62--91\% with an adversarial LLM persuader. Critically, we find that pressure that succeeds in changing a judge's verdict is almost always net-corrupting with respect to ground truth. Beyond the framework itself, we identify baseline jury majority strength as the most effective single-shot signal for anticipating which items wiggle. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.

13:00 JSTLLM/生成AIエージェント研究/論文

ScienceFlow: A long-horizon agent for ML research, scientific discovery and beyond

Enabling LLM agents to sustain productive, stable, and goal-aligned research over extended horizons is a central challenge for autonomous m…

13:00 JSTLLM/生成AIビジネス/資金調達

因果知識グラフにおけるヘルスケア LLM の基礎: フレームワーク、メトリクス、および心臓血管パイロット

大規模言語モデル (LLM) は、医療意思決定支援のために提案されることが増えていますが、その評価では依然として、介入、メカニズム、害、証拠、不確実性についての推論ではなく、単一回答の正確さが重視されています。私たちは、医療における介入指向の LLM 行動のための再現可能なグラフ中心の評価フレームワークを提案し、心血管パイロットでストレス テストを行います。このフレームワークには 4 つのコンポーネントがあります。(i) アサーションが安定した識別子を持つ来歴を保持する第一級のノードであるドメイン因果知識グラフ。 (ii) 任意の臨床シナリオが与えられた場合に、関連する具体化されたアサーションサブグラフを取得する、シナリオ条件付きサブグラフ抽出ステップ。 (iii) 取得したサブグラフをモデルのコンテキストに組み込む方法を変える 4 つの制御されたグラウンディング条件 (非グラウンディング C1、ナレッジ グラフ C2、因果グラフ C3、統合 C4)。 (iv) アサーション識別子に基づいた自動スコアリング パイプライン。介入の精度やその他の評価尺度を 1 回のパスで計算します。このフレームワークをテストするために、8 つの推論失敗モードにわたるカテゴリバランスのとれたシナリオ ジェネレーターを構築し、心血管グラフ上でインスタンス化しました。メトリクス パネルは、解釈可能な非冗長軸に沿って条件を識別します。C4 は最も強い因果エッジ F1 (0.838)、悪影響 F1 (0.833)、証拠精度 (0.738)、およびサポートされていない請求率 (0.114) を取得しますが、C1 は測定可能な因果関係または証拠根拠がない状態で最高の生の介入精度 (0.948) を取得します。

原文 (English)

Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot

Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.

13:00 JSTLLM/生成AIエージェント研究/論文DeepSeek

Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL

LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system compari…

13:00 JST研究/論文

Dynamic Multi-Byte Prediction With Hierarchical Language Models

Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword…

13:00 JSTLLM/生成AIハードウェア/半導体

When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction

Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades th…

13:00 JSTLLM/生成AIエージェントロボティクス

HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents

Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in response to environmental feedbac…

13:00 JSTエージェント

SkillEffect: Checked Lowering for Memory-Bounded Agent Tools

Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. Howev…

13:00 JSTエージェント研究/論文

AutoResearch: 洞察力を入力し、幻覚を出力

自律型研究システムは、長期にわたる研究ワークフローを実行できるようになってきていますが、自動化だけでは、結果として得られるプロセスが科学的に根拠のあるものであることを保証できません。 AutoResearch は、研究アイデアの形成方法と実験を通じて確実に確立する方法の両方に対処するために、アイデアの生成とアイデアの実行を結び付ける 2 段階のシステムです。アイデア生成では、オートリサーチは新たな研究シグナルを蓄積されたドメイン知識と継続的に統合し、移転可能なメカニズムの洞察を特定し、マルチモデルの生成とクロスレビューを使用して、根拠のあるテスト可能な研究計画を作成します。アイデアの実行では、調整されたエージェントがこれらの計画を実験に分解し、繰り返し実行および診断し、研究の結論を受け入れる前に独立した証拠に基づくレビューを採用します。 AutoResearch は、クロスモーダル検索、システムの最適化、ベンチマーク主導の機械学習の代表的な設定にわたって、生成されたアイデアを測定可能な進歩に変え、信頼性の低い実験結果を検出して修正し、研究の方向性を継続、修正、終了するための証拠に基づいた決定を下します。たとえば、RSICD ベンチマークでは、AutoResearch が生成したアイデアにより、平均再現率が 32.84 から 34.69 に向上しましたが、他の自律研究システムでは 11 ~ 27 件であったのに対し、監査で確認された問題イベントは 5 件しか記録されませんでした。これらの結果は、実験の前に意味のある洞察が根拠にあり、受け入れられる前に結論が根拠にあるという研究プロセスを示しています。つまり、洞察は入って幻覚は出るということです。

原文 (English)

AutoResearch: Insight In, Hallucination Out

Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.

13:00 JST研究/論文

神経象徴世界モデルによるゼロショットタスク転送に向けて

最先端のモデルベースの強化学習手法は、基礎となる環境の構造を仮定することなく、潜在空間で計画を立てることによって政策の改善を可能にするニューラル世界モデルを学習します。これらのモデルは表現力豊かではありますが、一般にタスクに依存します。トレーニング タスクに関連付けられている解釈できない潜在表現を学習するため、新しいタスクに一般化するのは困難です。この研究では、報酬予測が潜在状態全体の構造化された象徴的なコンポーネントのサブセットのみに依存する新しい世界モデルの定式化を提示します。観測の再構築と報酬予測を分離することにより、ゼロショット、つまりさらなる環境相互作用なしで、同じ記号状態空間上で定義された新しい報酬関数に適応できる世界モデルを学習できるようになります。これらの神経象徴世界モデルを学習することの主な利点と課題について説明し、純粋な神経的手法に対する私たちのアプローチの強力な一般化特性を実証します。

原文 (English)

Towards Zero-Shot Task Transfer with Neurosymbolic World Models

State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-dependent: they learn uninterpretable latent representations that are tied to the training task and thus hard to generalize to new tasks. In this work, we present a novel world model formulation where the reward prediction only depends on a subset of structured, symbolic components of the whole latent state. Decoupling observation reconstruction and reward prediction allows us to learn world models that can adapt zero-shot, i.e. without further environment interactions, to new reward functions defined over the same symbolic state space. We discuss the main advantages and challenges of learning these neurosymbolic world models and demonstrate the strong generalisation properties of our approach over purely neural methods.

13:00 JST研究/論文

目に見えないものはあなたが学ぶものです:共有ゲノム言語モデル社会では、限られた証拠の可視性が構成の一般化を促進します

マルチモジュール システムでは、多くの場合、すべてのモジュールが完全な入力に公開されます。証拠の可視性を制限すると、勾配ベースのトレーニングで発見されるソリューションが変わるかどうかをテストします。 4 セル社会は 1 つの凍結済み事前学習済み言語モデルと 1 つの低ランク アダプターを共有し、固定リレー内の 2 つのモデル幅の連続ベクトルを通じてのみ通信します。プロスペクティブシールされた自然言語関数合成タスクでは、初期化バイト、トレーニング順序、トークン レイアウト、パラメーター、および計算を共有する 10 個の一致する制限付き/グローバル ペアをトレーニングします。注意マスクのみが異なります。制限された社会は、10 ペア中 9 ペアで両方の深さで世界的に見える双子よりも少なくとも 20 ポイント優れており、ペアの利点の中央値は 0.7648 と 0.6050 です。コミュニケーションを遮断すると、すべての制限された社会は偶然に帰着し、複合関数がトレーニングに一度も現れなかったプログラムでは、深さ 3 の利点は 0.558 のままです。監査された 6 つの制限された社会全体で、同じ値のパケット移植により、テストされたすべてのインターフェイスで動作が 0.94 ~ 1.00 に維持されます。破壊的な介入はパフォーマンスを崩壊させます。そして反事実パケットは出力を数学的に予測された答えにリダイレクトします。唯一の高性能グローバル モデルも通信を必要としますが、その同じ値のパケットはエピソード間で交換できません。したがって、構図には可視性の制限は必要ありません。このプロトコルでは、一般化リレーの可能性が大幅に増加し、再利用可能な値インデックス付きインターフェイスが優先されます。それにもかかわらず、完全に事前登録されたバッテリーは、制限付きアームの深さ 3 の中央値の精度が 0.6988 で、0.70 フロアを下回っているため、正式に不合格となります。以前の資格コホートでも同様に完全合格は 0/10 でした。1 つのモデルはすべてのタスク パフォーマンス ゲートを満たしましたが、10 モデルすべてが通常言語の保存に失敗し、システムは明示的にタスク ゲートで使用されるように制限されました。

原文 (English)

What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies

Multi-module systems often expose every module to the full input. We test whether restricting evidence visibility changes which solutions gradient-based training discovers. Four-cell societies share one frozen pretrained language model and one low-rank adapter, communicating only through two model-width continuous vectors in a fixed relay. On a prospectively sealed natural-language function-composition task, we train ten matched restricted/global pairs sharing initialization bytes, training order, token layout, parameters, and computation; only the attention mask differs. Restricted societies outperform their globally visible twins by at least 20 points at both depths in 9 of 10 pairs, with median paired advantages of 0.7648 and 0.6050. Cutting communication reduces every restricted society to chance, and the depth-three advantage remains 0.558 on programs whose composite function never appeared in training. Across six audited restricted societies, same-value packet transplants preserve behavior at 0.94-1.00 across all tested interfaces; destructive interventions collapse performance; and counterfactual packets redirect outputs toward the mathematically predicted answer. The sole high-performing global model also requires communication, but its same-value packets are not interchangeable across episodes. Restricted visibility is thus not necessary for composition; under this protocol it substantially increases the probability of a generalizing relay and favors a reusable, value-indexed interface. The complete preregistered battery nevertheless formally fails because restricted-arm median depth-three accuracy is 0.6988, below the 0.70 floor. An earlier qualification cohort likewise yielded 0/10 complete passes: one model met every task-performance gate, but all ten failed ordinary-language preservation, confining the system to explicitly task-gated use.

13:00 JST研究/論文

Pandora の AI モデル ルーティング ボックス: コストのかかる価値推定による効率的な割り当て

複数のモデル、アーキテクチャ、ハーネス、または推論時間設定で構成される異種 AI システムは、最も低コストで最も効果的に回答できる専門家にクエリをルーティングすることで、品質と効率を向上させることができます。ルーティングでは各スペシャリストの期待利益を見積もる必要がありますが、この価値の見積りにはコストがかかります。安価な推定器 (埋め込みベースの予測器など) は高速ですがノイズが多く、正確な推定器 (検索結果や部分的な推論トレースにアクセスできる微調整されたモデルなど) は高価です。このトレードオフを、コストのかかる検査を伴う最適な検索という古典的な問題であるパンドラの箱のインスタンスとして形式化します。ガウス信号モデルの下では、結果として得られるポリシーには閉じた形式の情報価値の式が含まれており、専門家や入力ごとに、価値の見積もりを精緻化することがコストに見合うかどうかを判断します。私たちはこの集中ポリシーを Pandora's Router と呼んでいます。これを分散型設定である Pandora's Bidder に拡張し、クエリを請求するために提示された価格を受け入れる前に、スペシャリストが自己評価に投資するかどうかを独自に決定します。標準的なマルチ LLM ベンチマーク、検索拡張スペシャリスト、可変推論時間推論を備えた LLM の 3 つのドメインにわたる実験では、Pandora のルーターが、高価な推定器へのクエリの頻度がはるかに少ない一方で、網羅的な推定のルーティング品質と一致することが示されました。分散型設定では、競合する推定値が正確である場合、情報価値推論により割り当て効率が向上します。ただし、競合する見積もりにノイズが多い場合は、他の人を犠牲にして戦略専門家の有用性を高める可能性があります。

原文 (English)

Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation

Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficiency by routing queries to the specialist who can answer most effectively at the lowest cost. Routing requires estimating each specialist's expected return, but this value estimation has a cost. Cheap estimators (e.g., embedding-based predictors) are fast but noisy, while accurate estimators (e.g., fine-tuned models with access to retrieval results or partial reasoning traces) are expensive. We formalize this tradeoff as an instance of Pandora's Box, the classical problem of optimal search with costly inspection. Under a Gaussian signal model, the resulting policies have closed-form value-of-information expressions that determine, for each specialist and input, whether refining the value estimate is worth its cost. We call the centralized policy Pandora's Router. We extend this to a decentralized setting, Pandora's Bidder, where specialists independently decide whether to invest in self-assessment before accepting an offered price to claim a query. Experiments across three domains---a standard multi-LLM benchmark, retrieval-augmented specialists, and LLMs with variable inference-time reasoning---show that Pandora's Router matches the routing quality of exhaustive estimation, while querying the expensive estimator far less often. In the decentralized setting, value-of-information reasoning improves allocative efficiency when competing estimates are accurate; when competing estimates are noisy, however, it can increase the strategic specialist's utility at the expense of others.

13:00 JSTロボティクス

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tun…

13:00 JST研究/論文Gemma

Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and th…

13:00 JST研究/論文

Why we need an AI-resilient society- Profiling Large Language Models

Three generations of software have transformed the role of artificial intelligence in society. In the first, programmers wrote explicit log…

13:00 JST研究/論文

What is an intelligent system?

The term intelligent system has emerged in the field of information technology as a category of computer systems derived from successful ap…

13:00 JSTLLM/生成AIMicrosoft

Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation

In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with…

13:00 JSTLLM/生成AI研究/論文

Evaluating the Efficacy of LLMs to Emulate Realistic Human Personalities

To enhance immersion and engagement in video games, the design of Affective Non-Player Characters (ANPCs) is a key focus for researchers an…

13:00 JST画像/動画生成

Image-Conditional Diffusion Transformer for Underwater Image Enhancement

Underwater image enhancement (UIE) has attracted much attention owing to its importance for underwater operation and marine engineering. Mo…

13:00 JSTLLM/生成AI

LSem2Vec: A Simple yet Effective Two-Stage Approach for Source Code Embedding

The advent of large language models (LLMs) has significantly advanced artificial intelligence in software engineering, with source code emb…

13:00 JST研究/論文

Revisiting Multi-Permutation Equivariance through the Lens of Irreducible Representations

This paper explores the characterization of equivariant linear layers for representations of permutations and related groups. Unlike tradit…

13:00 JSTLLM/生成AI

NeST: Neighborhood-aware semantic alignment and temporal modulation for LLM based time series forecasting

Adapting Large Language Models (LLMs) trained on discrete text data, to forecast continuous time series signals is challenging. While finet…

13:00 JSTエージェント

SRMT: Shared Memory for Multi-agent Lifelong Pathfinding

Coordination in decentralized multi-agent reinforcement learning (MARL) necessitates that agents share information about their behavior and…

13:00 JSTLLM/生成AIOpenAI

Towards Safer Social Media Platforms: Scalable and Performant Few-Shot Harmful Content Moderation Using Large Language Models

The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and…

13:00 JST研究/論文

Bringing Generative Learning to Representation Learning: Self-Supervised Transfer Learning as Distribution Matching

Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified. We formulate represen…

13:00 JST画像/動画生成

SAS: Segment Anything Small for Ultrasound -- A Non-Generative Data Augmentation Technique for Robust Deep Learning in Ultrasound Imaging

Accurate segmentation of anatomical structures in ultrasound (US) images, particularly small ones, is challenging due to noise and variabil…

13:00 JSTLLM/生成AI

Deep Contrastive Unlearning for Language Models

The past a few years have witnessed the great success of large language models, demonstrating powerful capabilities in comprehending textua…

13:00 JSTLLM/生成AI

Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling

Dense retrieval is a crucial task in Information Retrieval (IR), serving as the basis for downstream tasks such as re-ranking and augmentin…

13:00 JSTLLM/生成AI

AI University: An LLM-Powered Learning Assistant for Engineering---A Finite Element Method Case Study

We introduce AI University (AI-U), a flexible framework for AI-driven course content delivery that adapts to a course's instructional style…

13:00 JSTLLM/生成AIGPT / ChatGPT

ClinicalGPT-R1: Pushing reasoning capability of generalist disease diagnosis with large language model

Recent advances in reasoning with large language models (LLMs)has shown remarkable reasoning capabilities in domains such as mathematics an…

13:00 JST画像/動画生成

COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification

We introduce CoLoRA (Convolutional Low-Rank Adaptation), a parameter-efficient fine-tuning method for convolutional neural networks (CNNs).…

13:00 JSTエージェントロボティクス

Balancing Safety and Optimality in Robot Path Planning: Algorithm and Metric

Path planning for autonomous robots faces a fundamental trade-off between path length and obstacle clearance. While existing algorithms typ…

13:00 JSTLLM/生成AILlamaQwenDeepSeek

Effects of Theory of Mind and Prosocial Beliefs on Steering Human-Aligned Behaviors of LLMs in Ultimatum Games

Large Language Models (LLMs) have shown potential in simulating human behaviors and performing theory-of-mind (ToM) reasoning, crucial for…

13:00 JSTLLM/生成AIOpenAI

推論としての時系列予測: 強化された LLM を使用したゆっくりとした思考のアプローチ

時系列予測 (TSF) を進歩させるために、予測精度を向上させるさまざまな方法が提案されており、統計的手法からデータ駆動型の深層学習アーキテクチャに進化しています。その有効性にもかかわらず、既存の手法のほとんどは依然として高速思考パラダイムに固執しており、中核となるモデリング哲学として歴史的パターンの抽出と将来の値へのマッピングに依存しており、中間の時系列推論を組み込んだ明示的な思考プロセスが欠けています。一方、新興の低速思考 LLM (OpenAI-o1 など) は、驚くべき多段階推論能力を示し、これらの問題を克服する代替方法を提供しています。ただし、迅速なエンジニアリングだけでは、高い計算コスト、プライバシーのリスク、ドメイン固有の時系列推論の詳細な能力の制限など、いくつかの制限があります。これらの制限に対処するためのより有望なアプローチは、ゆっくりとした思考能力を開発し、強力な時系列推論スキルを獲得するように LLM を訓練することです。この目的のために、時系列予測のためのLLMの多段階推論能力を強化するように設計された2段階の強化微調整フレームワークであるTime-R1を提案します。具体的には、第 1 段階ではウォームアップ適応のための教師あり微調整を行い、第 2 段階では強化学習を採用してモデルの汎化能力を向上させます。特に、時系列予測に特化したきめの細かい多目的報酬を設計し、次に GRIP (ポリシー最適化のためのグループベースの相対重要度) を導入します。これは、不均一なサンプリングを活用して、モデルによる効果的な推論パスの探索をさらに促進および最適化します。実験では、Time-R1 がさまざまなデータセットにわたって予測パフォーマンスを大幅に向上させることが実証されています。

原文 (English)

Time Series Forecasting via Reasoning: A Slow-Thinking Approach with Reinforcement Fine-Tuned LLMs

To advance time series forecasting (TSF), various methods have been proposed to improve prediction accuracy, evolving from statistical techniques to data-driven deep learning architectures. Despite their effectiveness, most existing methods still adhere to a fast thinking paradigm-relying on extracting historical patterns and mapping them to future values as their core modeling philosophy, lacking an explicit thinking process that incorporates intermediate time series reasoning. Meanwhile, emerging slow-thinking LLMs (e.g., OpenAI-o1) have shown remarkable multi-step reasoning capabilities, offering an alternative way to overcome these issues. However, prompt engineering alone presents several limitations - including high computational cost, privacy risks, and limited capacity for in-depth domain-specific time series reasoning. To address these limitations, a more promising approach is to train LLMs to develop slow thinking capabilities and acquire strong time series reasoning skills. For this purpose, we propose Time-R1, a two-stage reinforcement fine-tuning framework designed to enhance multi-step reasoning ability of LLMs for time series forecasting. Specifically, the first stage conducts supervised fine-tuning for warmup adaptation, while the second stage employs reinforcement learning to improve the model's generalization ability. Particularly, we design a fine-grained multi-objective reward specifically for time series forecasting, and then introduce GRIP (group-based relative importance for policy optimization), which leverages non-uniform sampling to further encourage and optimize the model's exploration of effective reasoning paths. Experiments demonstrate that Time-R1 significantly improves forecast performance across diverse datasets.

13:00 JST研究/論文

Seismic Acoustic Impedance Inversion Framework Based on Conditional Latent Generative Diffusion Model

Seismic acoustic impedance plays a crucial role in lithological identification and subsurface structure interpretation. However, due to the…

13:00 JSTLLM/生成AI画像/動画生成

From Recognition to Reasoning: Advancing Multimodal Harmful Meme Detection via Chain-of-Thought Alignment

As a multimodal communication medium that integrates images and text, memes often convey implicit harmful content through metaphors, satire…

13:00 JSTLLM/生成AI

A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs

Spatio-temporal data mining plays a pivotal role in informed decision making across diverse domains. However, existing models are often res…

13:00 JSTエージェントロボティクスビジネス/資金調達研究/論文

Mission-Aligned Learning-Informed Control of Autonomous Systems: Formulation and Foundations

Research, innovation and practical capital investment have been increasing rapidly toward the realization of autonomous physical agents. Th…

13:00 JSTLLM/生成AI

MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora

Continually updating model-based indexes in generative retrieval with new documents remains challenging, as full retraining is computationa…

13:00 JSTLLM/生成AI研究/論文OpenAILlamaMistral AI

Text-ADBench: Text Anomaly Detection Benchmark Based on LLM Embeddings

Text anomaly detection is a critical task in natural language processing (NLP), with applications spanning fraud detection, misinformation…

13:00 JST研究/論文

RetroDFM-R: Reasoning-Driven Retrosynthesis Prediction with Large Language Models via Reinforcement Learning

Retrosynthetic planning is a cornerstone of organic synthesis and drug discovery. Yet existing AI methods often rely on pattern matching ra…

13:00 JSTLLM/生成AI研究/論文

TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios

Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks p…

13:00 JST研究/論文

Entity Representation Learning Through Onsite-Offsite Graph for Pinterest Ads

Graph Neural Networks (GNN) have been extensively applied to industry recommendation systems, as seen in models like GraphSage\cite{GraphSa…

13:00 JSTLLM/生成AI

From Isolation to Alignment: Unified LoRA for Efficient Multi-Task Learning

Parameter-Efficient Fine-Tuning (PEFT) is essential for adapting Large Language Models (LLMs) to multi-task scenarios. A prevailing trend i…

13:00 JSTエージェント

An Information-Flow Perspective on Explainability Requirements: Specification and Verification

Explainable systems expose information about why certain observed effects are happening to the agents interacting with them. We argue that…

13:00 JST画像/動画生成

ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting

Lifting 2D open-vocabulary understanding into 3D Gaussian Splatting (3DGS) scenes is a critical challenge. Mainstream methods, built on an…

13:00 JSTLLM/生成AI

HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models

Speculative decoding has proven effective for accelerating inference in Large Language Models (LLMs), yet its extension to Vision-Language…

13:00 JST研究/論文

SLogic: Subgraph-Informed Logical Rule Learning for Knowledge Graph Completion

Logical rule-based methods offer an interpretable approach to knowledge graph completion (KGC) by capturing compositional relationships in…

13:00 JST研究/論文

RSTGCN: Railway-centric Spatio-Temporal Graph Convolutional Network for Train Delay Prediction

Accurate prediction of train delays is critical for efficient railway operations. While earlier approaches have largely focused on forecast…

13:00 JSTLLM/生成AI

GraphMed-LT: Patient-Specific Graph Memory with Latent Clinical Thought Refinement for Multi-Turn Medical Conversations

Multi-turn medical question answering (QA) aims to model realistic clinical diagnosis, where a doctor gathers patient information across mu…

13:00 JSTLLM/生成AILlama

LLM-Specific Utility for Retrieval-Augmented Generation

Retrieval-augmented generation (RAG) is typically optimized for topical relevance, yet its success ultimately depends on whether retrieved…

13:00 JST研究/論文

A New Type of Adversarial Examples

Most machine learning models are vulnerable to adversarial examples, which poses security concerns on these models. Adversarial examples ar…

13:00 JST研究/論文

Mitigating Sample-Level Imbalance via Probabilistic Separation for Adaptive Multimodal Fusion

Multimodal learning faces modality imbalance, where dominant modalities suppress weaker ones due to inconsistent convergence rates. Existin…

13:00 JSTLLM/生成AI

One Request, Multiple Experts: LLM Orchestrates Domain Specific Models via Adaptive Task Routing

With the integration of massive distributed energy resources and the widespread participation of novel market entities, the operation of ac…

13:00 JST研究/論文

Radial Compensation: The Inverse Base-Distribution Problem for Chart-Based Generative Models on Riemannian Manifolds

Latent-variable models on spheres and hyperbolic spaces usually draw a Gaussian in the tangent space at a base point and push it onto the m…

13:00 JST研究/論文

HiFiNet: Hierarchical Fault Identification in Wireless Sensor Networks via Edge-Based Classification and Graph Aggregation

Wireless Sensor Networks (WSN) are the backbone of essential monitoring applications, but their deployment in unfavourable conditions incre…

13:00 JST研究/論文

MOCLIP: A Foundation Model for Large-Scale Nanophotonic Inverse Design

Foundation models (FM) are transforming artificial intelligence by enabling generalizable, data-efficient solutions across different domain…

13:00 JST画像/動画生成エージェント

Multi-Context Fusion Transformer for Pedestrian Crossing Intention Prediction in Urban Environments

Pedestrian crossing intention prediction is essential for autonomous vehicles to improve pedestrian safety and reduce traffic accidents. Ho…

13:00 JSTエージェントロボティクス

MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving

Autonomous Driving (AD) vehicles still struggle to exhibit human-like behavior in highly dynamic and interactive traffic scenarios. The key…

13:00 JST研究/論文

Diagnosing Capability Preservation and Task Sensitivity in Memory Augmented Document Classifiers

End task accuracy alone cannot determine whether a memory mechanism preserves an acquired capability, exposes sample-specific stored inform…

13:00 JST画像/動画生成研究/論文

Benchmarking Document Parsers on Mathematical Formula Extraction from PDFs

Correctly parsing mathematical formulas from PDFs is critical for training large language models and building scientific knowledge bases fr…

13:00 JSTLLM/生成AI

An Empirical Study on Preference Tuning Generalization and Diversity Under Domain Shift

Preference tuning aligns base language models to human judgments of quality, helpfulness, or safety by optimizing over explicit preference…

13:00 JSTLLM/生成AI

LLM-Based Adversarial Persuasion Attacks on Fact-Checking Systems

Automated fact-checking (AFC) systems are susceptible to adversarial attacks, enabling false claims to evade detection. Existing adversaria…

13:00 JST画像/動画生成ロボティクス

PEAfowl: Perception-Enhanced Multi-View Vision-Language-Action for Bimanual Manipulation

Bimanual manipulation in cluttered scenes requires policies that remain stable under occlusions, viewpoint changes and scene variations. Ex…

13:00 JSTエージェントGoogle

Dynamic Cogeneration of Bug Reproduction Test in Agentic Program Repair

Bug Reproduction Tests (BRTs) have been used in many Automated Program Repair (APR) systems, primarily for validating fixes and aiding fix…

13:00 JST研究/論文

Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks

Motivated by challenges in conditional generative modeling, where the target conditional density takes the form of a ratio f1 over f2, this…

13:00 JST画像/動画生成

Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models

Industrial troubleshooting guides encode diagnostic procedures in flowchart-like diagrams where spatial layout and technical language joint…

13:00 JST研究/論文

Backdoor Sentinel: Detecting and Detoxifying Backdoors in Diffusion Models via Temporal Noise Consistency

Diffusion models have been widely deployed in AIGC services, but their reliance on opaque training data exposes them to backdoor attacks. I…

13:00 JST画像/動画生成

ReasonEdit: Editing Vision-Language Models using Human Reasoning

Model editing aims to correct errors in large, pretrained models without altering unrelated behaviors. While some recent works have edited…

13:00 JST研究/論文

First-Principles AI finds crystallization of fractional quantum Hall liquids

When does a fractional quantum Hall (FQH) liquid crystallize? Addressing this question requires a framework that treats fractionalization a…

13:00 JSTビジネス/資金調達

Mode-Dependent Rectification for Stable PPO Training

Mode-dependent architectural components (layers that behave differently during training and evaluation, such as Batch Normalization or drop…

13:00 JSTビジネス/資金調達

iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems

Scheduling precedence-constrained tasks under shared renewable resources is critical to modern computing platforms. It is often modeled as…

13:00 JST研究/論文

Which Algorithms Can Graph Neural Networks Learn?

In recent years, there has been growing interest in understanding neural architectures' ability to learn to execute discrete algorithms, a…

13:00 JSTLLM/生成AIエージェント

ST-EVO: Towards Generative Spatio-Temporal Evolution of Multi-Agent Communication Topologies

LLM-powered Multi-Agent Systems (MAS) have emerged as an effective approach towards collaborative intelligence, and have attracted wide res…

13:00 JST研究/論文

Continual Uncertainty Learning for Robust Control of Nonlinear Systems with Multiple Heterogeneous Uncertainties

Robust control of mechanical systems with multiple uncertainties remains a fundamental challenge, particularly when nonlinear dynamics and…

13:00 JST研究/論文

Learning with Boolean threshold functions

We develop a method for training neural networks on Boolean data in which the values at all nodes are strictly $\pm 1$, and the resulting m…

13:00 JSTLLM/生成AICopilot

Vibe Coding on Trial: Operating Characteristics of Unanimous LLM Juries

Large Language Models (LLMs) are now good enough at coding that developers can describe intent in plain language and let the tool produce t…

13:00 JSTLLM/生成AI

Semantic Substrate Dynamics Theory: An Operator-Theoretic Framework for Geometric Semantic Drift

Studies of semantic drift report heterogeneous signals, including embedding displacement, neighbor change, distributional divergence, and r…

13:00 JSTLLM/生成AI画像/動画生成

VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting

Video reasoning requires models to locate and track question-relevant evidence across frames. While reinforcement learning (RL) with verifi…

13:00 JSTLLM/生成AIエージェントロボティクス

ロボットはいつ考えるべきでしょうか?身体化されたロボットによる意思決定のための強化学習によるリソース認識型推論

身体化されたロボット システムは、環境との対話中に高レベルの推論、計画、意思決定をサポートするために、大規模言語モデル (LLM) ベースのエージェントにますます依存しています。ただし、LLM 推論を呼び出すと、かなりの計算遅延とリソースのオーバーヘッドが発生し、アクションの実行が中断され、システムの信頼性が低下する可能性があります。過度の推論は行動を遅らせる可能性がありますが、不十分な推論は誤った決定やタスクの失敗につながることがよくあります。このことは、実体化されたエージェントにとって根本的な疑問を引き起こします。それは、エージェントはいつ論理的に判断し、いつ行動すべきなのかということです。この研究では、具現化されたエージェントのリソースを認識したオーケストレーションのための階層フレームワークである RARRL (Resource-Aware Reasoning via Reinforcement Learning) を提案します。 RARRL は、低レベルの制御ポリシーを学習するのではなく、エージェントの意思決定層で動作する高レベルのオーケストレーション ポリシーを学習します。このポリシーにより、エージェントは、現在の観察、実行履歴、および残りのリソースに基づいて、推論を呼び出すかどうか、どの推論ロールを使用するか、どの程度の計算予算を割り当てるかを適応的に決定できます。 ALFRED ベンチマークから得られた経験的レイテンシ プロファイルによる評価を含む広範な実験により、固定推論戦略またはヒューリスティック推論戦略と比較して、RARRL が実行レイテンシを削減し堅牢性を高めながら、タスクの成功率を一貫して向上させることが示されています。これらの結果は、適応推論制御が信頼性が高く効率的な身体化ロボット エージェントを構築するために不可欠であることを示しています。

原文 (English)

When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making

Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-making during interactions with the environment. However, invoking LLM reasoning introduces substantial computational latency and resource overhead, which can interrupt action execution and reduce system reliability. Excessive reasoning may delay actions, while insufficient reasoning often leads to incorrect decisions and task failures. This raises a fundamental question for embodied agents: when should the agent reason, and when should it act? In this work, we propose RARRL (Resource-Aware Reasoning via Reinforcement Learning), a hierarchical framework for resource-aware orchestration of embodied agents. Rather than learning low-level control policies, RARRL learns a high-level orchestration policy that operates at the agent's decision-making layer. This policy enables the agent to adaptively determine whether to invoke reasoning, which reasoning role to employ, and how much computational budget to allocate based on current observations, execution history, and remaining resources. Extensive experiments, including evaluations with empirical latency profiles derived from the ALFRED benchmark, show that RARRL consistently improves task success rates while reducing execution latency and enhancing robustness compared with fixed or heuristic reasoning strategies. These results demonstrate that adaptive reasoning control is essential for building reliable and efficient embodied robotic agents.

13:00 JSTエージェントロボティクス

Robust Multi-Agent Reinforcement Learning for Small UAS Separation Assurance under GPS Degradation and Spoofing

We address robust separation assurance for small Unmanned Aircraft Systems (sUAS) under GPS degradation and spoofing via Multi-Agent Reinfo…

13:00 JSTLLM/生成AI研究/論文

SAFE: An LLM-as-Verifier Framework for Evidence-Grounded Multi-Hop Reasoning

Multi-hop QA benchmarks often reward Large Language Models (LLMs) for spurious correctness, where models reach correct answers through inva…

13:00 JSTLLM/生成AI

Verbalizing LLMs' assumptions to explain and control sycophancy

LLMs can be socially sycophantic, affirming users when they ask questions like "am I in the wrong?" rather than providing genuine assessmen…

13:00 JSTLLM/生成AI

What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"

While large language models (LLMs) demonstrate strong capabilities across diverse user queries, they still suffer from hallucinations, ofte…

13:00 JST画像/動画生成

FlowExtract: Procedural Knowledge Extraction from Maintenance Flowcharts

Maintenance procedures in manufacturing facilities are often documented as flowcharts in static PDFs or scanned images. They encode procedu…

13:00 JSTLLM/生成AI画像/動画生成研究/論文

Beyond RGB: Benchmarking and Enhancing MLLMs for Hyperspectral Image Understanding via Training-Free Reasoning Framework

Multimodal Large Language Models (MLLMs) have achieved strong performance on RGB image understanding, yet their ability to use spectral evi…

13:00 JSTLLM/生成AI

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

Large language models remain vulnerable to jailbreaks that elicit harmful responses, yet the mechanism behind harmful response generation i…

13:00 JSTLLM/生成AIエージェント

CONSCIENTIA: Can LLM Agents Learn to Strategize? Emergent Deception and Trust in a Multi-Agent NYC Simulation

As large language models (LLMs) are increasingly deployed as autonomous agents, understanding how strategic behavior emerges in multi-agent…

13:00 JSTLLM/生成AI

Alignment midtraining for animals

We investigate the robustness of value alignment via midtraining with synthetic documents, using animal compassion as a value that is both…

13:00 JSTLLM/生成AI研究/論文Gemini

DialToM: 国家主導の対話の軌道を予測するための心のベンチマーク理論

DialToM は、多肢選択評価フレームワークを使用した自然主義的な人間間の対話から構築された、注釈付きの Theory of Mind (ToM) ベンチマークです。明示的な精神状態の推論と合成設定で適用される ToM との間のギャップを示す最近の研究と並行して、~\cite{gu2024simpletom}、モデルが対話のコンテキストなしで孤立した精神状態のプロファイルのみから状態一貫性のある対話の軌跡を予測する必要がある、より厳密な \emph{状態駆動型診断プローブ} を確立します。私たちの評価では、体系的な推論の非対称性が明らかになりました。LLM は精神状態の推論 (文字通りの ToM) には優れていますが、それを社会予測に活用するのに苦労しています (関数 ToM)。重要なのは、ドメインの専門家がこのタスクで 100% の精度を達成し、その妥当性を証明し、人間と AI の能力の明らかなギャップを確立することです。さらに、教師と生徒の推論注入プローブにより、主要なベースラインを確立する Gemini 3 Pro が、弱いモデルに転送可能なコンテキストフリー予測のための堅牢な Functional ToM 機能を備えていることが示されています。 DialToM、その評価コード、およびデータセットは、https://github.com/Stealth-py/DialToM で公開されています。

原文 (English)

DialToM: A Theory of Mind Benchmark for Forecasting State-Driven Dialogue Trajectories

We introduce DialToM, an annotated Theory of Mind (ToM) benchmark built from naturalistic human-human dialogues using a multiple-choice evaluation framework. Concurrent with recent work showing a gap between explicit mental-state inference and applied ToM in synthetic settings~\cite{gu2024simpletom}, we establish a stricter \emph{State-Driven Diagnostic Probe} in which models must forecast state-consistent dialogue trajectories solely from isolated mental-state profiles without dialogue context. Our evaluation reveals a systematic reasoning asymmetry -- LLMs excel at inferring mental states (Literal ToM) but struggle to leverage them for social forecasting (Functional ToM). Crucially, a domain expert achieves 100\% accuracy on this task, proving its validity and establishing a stark human-AI capability gap. Further, a teacher-student reasoning injection probe shows that Gemini 3 Pro -- which establishes the leading baseline -- possesses robust Functional ToM capabilities for context-free forecasting that are transferable to weaker models. DialToM, its evaluation code, and dataset are publicly available at https://github.com/Stealth-py/DialToM.

13:00 JST研究/論文

ONOTE: Hypergraph-Grounded Omnimodal Reasoning for Computational Music Science

Omnimodal notation processing, centered on sheet music, is a controlled scientific setting in which auditory, visual, symbolic, and physica…

13:00 JST研究/論文

DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Discrete Diffusion Models

Discrete diffusion models admit many token orders, yet most systems rely on confidence-based decoding. Confidence is a strong and efficient…

13:00 JSTLLM/生成AIエージェント

DeepRefine: Agentic Knowledge Refinement via Reinforcement Learning

External knowledge enables large language model (LLM) agents to ground their actions and decisions beyond intrinsic parametric memory in op…

13:00 JST研究/論文

Revisiting Reinforcement Learning with Verifiable Rewards from a Contrastive Perspective

Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reas…

13:00 JST研究/論文

Edge-AI-Driven Learning-to-Rank for Decentralized Task Allocation in Circular Smart Manufacturing

Task allocation in smart manufacturing systems must operate under decentralized decision-making, dynamic workloads, and shared-resource con…

13:00 JSTLLM/生成AI

LLM は高品質のデータをどのように利用すべきでしょうか?品質を意識した機能スケーリング法則による最適なデータ スケジューリング

大規模言語モデル (LLM) トレーニングでは高品質のデータが不足していますが、トレーニング ダイナミクスと組み合わせてその使用をスケジュールする方法には理論的な指針がありません。データ品質の次元を組み込むことで関数スケーリング則を拡張し、データ品質とバッチサイズのスケジューリング問題を漸近閉形式で解決します。このソリューションは、高品質データの 2 つの体制と二重の役割を明らかにします。ノイズが制限された領域では、高品質のデータを信号増幅器として使用する必要があります。バッチ サイズを下げると、ノイズを増幅することなく、よりクリーンなデータがより多くの信号に変換されます。信号が制限された状況では、ノイズ サプレッサーとして使用する必要があります。遅い位置に配置すると、信号の蓄積を犠牲にすることなく端末のノイズが低減されます。既存のカリキュラム スタイルのパイプラインは主に、よりクリーンなデータを遅れて配置することで 2 番目の役割を活用しますが、従来の減衰スケジュールでは高品質のデータが利用可能になったときにちょうど更新強度が低下するため、最初の役割を見逃しています。これに基づいて、LLM トレーニングの途中で Drop-Stable-Rampup を提案します。品質が変化したら、バッチ サイズをドロップし、安定に保持して信号を蓄積し、その後、ランプアップして端末ノイズを抑制します。 108B トークンで中間トレーニングされた 150 億の専門家混合モデルでは、Drop-Stable-Rampup は、Warmup-Stable-Decay (WSD) に対して +1.70、Cosine-decay に対して +2.98 の平均精度を向上させ、特に GSM8K (+4.23) や MATH (+2.80) などの数学的推論ベンチマークで大きな向上をもたらします。

原文 (English)

How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws

High-quality data is scarce in large language model (LLM) training, yet how to schedule its use with optimization dynamics lacks theoretical guidance. We extend functional scaling laws with time-varying data quality and derive asymptotically optimal joint data-quality and batch-size schedules within a feature-space regression model. The solution reveals two regimes and dual uses of high-quality data: in the noise-limited regime, a smaller batch converts cleaner data into more signal at comparable noise; in the signal-limited regime, late placement suppresses terminal noise without sacrificing signal accumulation. This explains why conventional decay schedules can conflict with curriculum-style pipelines. Motivated by the theoretical structure, we propose Drop-Stable-Rampup for LLM midtraining: drop the batch size at the quality transition, keep it low to accumulate signal, then ramp up to suppress noise. On a 15B MoE model midtrained on 108B tokens of general-domain proprietary data, Drop-Stable-Rampup improves average accuracy over Warmup-Stable-Decay by +1.70 and Cosine-decay by +2.98, including +4.23 on GSM8K and +2.80 on MATH. On a public math-and-code mixture, it leads all reported STEM, mathematics, and code benchmarks, improving the overall mean over the strongest baseline by +3.27 on a 600M dense model and +5.25 on the same MoE architecture.

13:00 JST研究/論文

BIRDNet: 解釈可能なディープ ニューラル ネットワークとしてのブール含意ナレッジ グラフのマイニングとエンコード

知識が豊富な領域の表形式データは、多くの場合、特徴のペア間のブール含意関係 (BIR) の形式で潜在事前分布を保持します。私たちは、スパース例外の二項テストを使用して、そのような関係をマイニングします。マイニングされた含意は、2 リテラル節の命題ルール ベースに相当する、型付き有向グラフを形成します。このグラフを BIRDNet と呼ばれる層状ニューラル ネットワークの接続としてエンコードします。このネットワークでは、各隠れユニットが 1 つのマイニングされたルールに対応し、その 2 つの機能のみにバインドされます。この設計の 2 つの結果を示します。 まず、アーキテクチャは構造上スパースです。各 BIR 層の重みの最大 $2/d$ がアクティブになります ($d$ は入力次元です)。第 2 に、モデルは解釈可能です。トレーニングされたすべてのユニットは安定したシンボル ID を保持するため、サロゲート モデルを使用せずにネットワークからルールを読み取ることができます。ほとんどの神経象徴モデルとは異なり、BIRDNet は外部ルール ベースを使用しません。その構造的事前分布はデータから抽出されます。私たちは 6 つのトランスクリプトームおよびプロテオミクス ベンチマークで BIRDNet を評価します。私たちの結果は、BIRDNet が、わずかな精度コストで、最も強力な高密度ベースラインの 0.02 AUROC 以内に留まり、アーキテクチャが一致した高密度 MLP よりも最大 96 倍少ないアクティブ パラメータを使用することを示しています。第 1 層ルールは、標準アンプリコン、系列を定義する共発現モジュール、免疫浸潤マーカーなど、複数のがんサブタイプおよび組織タイプにわたって既知の生物学的シグネチャを回復します。データとコードは https://github.com/MAHI-Group/BIRDNet から入手できます。

原文 (English)

BIRDNet: Mining and Encoding Boolean Implication Knowledge Graphs as Interpretable Deep Neural Networks

Tabular data in knowledge-rich domains often carries a latent prior in the form of Boolean implication relationships (BIRs) between pairs of features. We mine such relationships with a sparse-exception binomial test. We encode the resulting typed graph as the connectivity of a layered neural network, called BIRDNet, in which each hidden unit corresponds to one mined rule and binds only to its two features. We show two consequences of this design: First, the architecture is sparse: at most $2/d$ of the weights in each BIR layer are active, where $d$ is the input dimension. Second, the model is intrinsically interpretable: every trained unit keeps a stable symbolic identity, so rules can be read off the network without surrogate models. Unlike most neurosymbolic models, BIRDNet does not consume an external rule base; its structural prior is mined from the data. We evaluate BIRDNet on six transcriptomic and proteomic benchmarks. Our results show that BIRDNet stays within $0.02$ AUROC of the strongest dense baseline, while using up to $95\times$ fewer active parameters than an architecture-matched dense MLP. First-layer rules align with known biological signatures across multiple cancer subtypes and tissue types. Matched-topology controls show that the mined graph contributes symbolic meaning rather than predictive advantage: shuffled or random pairings match or improve AUROC but no longer correspond to mined implications. Data and code are available at: https://github.com/MAHI-Group/BIRDNet.

13:00 JST研究/論文

DRIFT: パイロットレス 6G 非地上波ネットワークに向けた共同チャネル推定と予測

非地上ネットワーク (NTN) は、ユビキタス接続と大規模通信を可能にすることで、第 6 世代 (6G) システムにおいて極めて重要な役割を果たすことが期待されています。これに関連して、チャネル予測は、パイロットのオーバーヘッドを制限することでスペクトルの利用効率を向上させる重要な技術として浮上します。ただし、人工知能 (AI) に基づいて提案されている予測子の多くは、推論の複雑さが高いという特徴があり、オンボード実装に課題をもたらしています。この論文では、厳密な電力制約によりモデルの複雑さが制限される低地球軌道 (LEO) NTN に合わせて、スペクトル効率の向上を可能にする、正確かつ計算効率の高いチャネル予測技術を設計するという課題に取り組みます。我々は、最初のスロットでのみパイロットを送信し、後続のスロットではデータ駆動型の処理に依存することによってパイロットのオーバーヘッドを大幅に削減する、6G NTN のコンテキストでの反復結合チャネル推定および予測フレームワークを提案します。無線チャネル追跡のためのデータ駆動型の改良と反復予測 (DRIFT) を紹介します。これは、データ支援によるチャネル推定を改良し、低い計算コストと少ないエラー伝播で将来のチャネル周波数応答を予測する軽量アーキテクチャです。畳み込み記憶層と長期短期記憶層に基づく 2 つの予測子のバリアントが調査されます。アップリンク LEO NTN シナリオのエンドツーエンド シミュレーションの結果は、提案されたアプローチが、トレーニングとテストの不一致に対する堅牢性と、さまざまなチャネル モデル間で一貫したパフォーマンスを備え、従来のパイロットベースのシステムと比較して最大 12% のスペクトル効率向上を達成することを示しています。さらに、DRIFT に必要な積和演算は 200k 未満であるため、厳しい電力制約下での衛星搭載に適しています。

原文 (English)

DRIFT: Joint Channel Estimation and Prediction Towards Pilotless 6G Non-Terrestrial Networks

Non-terrestrial networks (NTNs) are expected to play a pivotal role in sixth-generation (6G) systems by enabling ubiquitous connectivity and massive communication. In this context, channel prediction emerges as a key technique to improve the spectrum utilization efficiency by limiting the pilot overhead. However, many proposed predictors based on artificial intelligence (AI) are characterized by high inference complexity, posing challenges to onboard implementation. In this paper, we address the challenge of designing accurate yet computationally efficient channel prediction techniques tailored to low Earth orbit (LEO) NTNs, where strict power constraints limit model complexity, to enable spectral efficiency gains. We propose an iterative joint channel estimation and prediction framework in the context of 6G NTNs that significantly reduces pilot overhead by transmitting pilots only in the initial slot and relying on data-driven processing for subsequent slots. We introduce Data-driven Refinement and Iterative Forecast for wireless channel Tracking (DRIFT), a lightweight architecture that refines data-aided channel estimates and predicts future channel frequency responses with low computational cost and reduced error propagation. Two predictor variants based on convolutional and long short-term memory layers are investigated. Simulation results in an end-to-end simulation of an uplink LEO NTN scenario show that the proposed approach achieves up to 12% spectral efficiency gain compared to conventional pilot-based systems, with robustness to training-test mismatches and consistent performance across different channel models. Moreover, DRIFT requires fewer than 200k multiply-accumulate operations, making it suitable for on-board satellite implementation under stringent power constraints.

13:00 JSTLLM/生成AI

Soft-NBCE: 長いコンテキストのエントロピー加重チャンク融合

自己注意の 2 次の複雑さは、超長いコンテキストを処理する大規模言語モデル (LLM) のボトルネックのままです。 Naive Bayes Cognitive Engine (NBCE) は、ドキュメントをチャンク化し、各デコード ステップで最もエントロピーの低いチャンクにルーティングすることで、長いコンテキストの推論を並列化します。このハード選択戦略は、隣接するトークン間のルーティングの突然の変更によりモデルのコンテキスト基盤が破壊されるため、クロスチャンク推論中にセマンティックな断片化を引き起こします。私たちは、個別のチャンク選択をソフト エントロピー加重チャンク フュージョンに置き換える軽量拡張機能である Soft-NBCE を紹介します。予測エントロピーに対する温度スケールのソフトマックスは、すべてのチャンクに連続的な重みを割り当て、チャンク条件付き分布全体で対数空間の集計を可能にします。チャンク化によって導入された条件付き独立性の仮定を部分的に補うために、KL 発散を介してフルコンテキスト教師に向けてチャンク化されたロジット分布を制約する LoRA ベースの自己蒸留である整合性蒸留を提案します。 LongBench マルチホップ ベンチマークでは、Consistency Distillation を使用した Soft-NBCE は、取得精度 (NIAH-32K: 0.909) を維持しながら、NBCE スタイルのベースライン (MuSiQue F1: 0.310 vs. \ 0.275 (Vanilla NBCE の場合)、HotpotQA F1: 0.479 vs. \ 0.427) より一貫して向上しています。 O(L^2/n) ピークメモリ。

原文 (English)

Soft-NBCE: Entropy-Weighted Chunk Fusion for Long-Context

The quadratic complexity of self-attention remains a bottleneck for Large Language Models (LLMs) processing ultra-long contexts. The Naive Bayes Cognitive Engine (NBCE) parallelizes long-context inference by chunking documents and routing to the lowest-entropy chunk at each decoding step. This hard-selection strategy causes semantic fragmentation during cross-chunk reasoning, as abrupt routing changes between adjacent tokens disrupt the model's contextual grounding. We present Soft-NBCE, a lightweight extension that replaces discrete chunk selection with soft entropy-weighted chunk fusion. A temperature-scaled Softmax over predictive entropies assigns continuous weights to all chunks, enabling log-space aggregation across chunk-conditioned distributions. To partially compensate for the conditional independence assumption introduced by chunking, we propose Consistency Distillation, a LoRA-based self-distillation that constrains the chunked logit distribution toward a full-context teacher via KL-divergence. On LongBench multi-hop benchmarks, Soft-NBCE with Consistency Distillation improves consistently over NBCE-style baselines (MuSiQue F1: 0.310 vs.\ 0.275 for Vanilla NBCE; HotpotQA F1: 0.479 vs.\ 0.427) while maintaining retrieval accuracy (NIAH-32K: 0.909) at O(L^2/n) peak memory.

13:00 JSTLLM/生成AI

E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments

Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the mo…

13:00 JSTLLM/生成AI研究/論文LlamaDeepSeek

Sample-Efficient Post-Training for LEGO Spatial-Physics Reasoning

LLM-based LEGO assembly requires both semantic grounding and physical feasibility. In this paper, we identify a data-induced failure mode,…

13:00 JSTLLM/生成AI

SocraticPO: インタラクティブなガイダンスによるポリシーの最適化

大規模言語モデルの強化学習 (RL) は通常、バイナリの正しさなどのスカラー結果報酬で推論を監視します。このような報酬は最適化の方向性を提供しますが、モデルがその誤った推論をどのように修正すべきかを説明することはほとんどなく、ショートカット学習や脆弱なポリシーを促進する可能性があります。私たちは \textbf{SocraticPO} (Socratic Policy Optimization) を提案します。これは、ソクラティック スタイルの自然言語ガイダンスで RL ロールアウトを強化するポリシー最適化フレームワークです。ロールアウト中、学生はまず独立して回答します。答えが間違っている場合、教師はその試みを診断し、簡潔な修正指導を提供します。その後、生徒は拡張されたコンテキストの下で続行します。重要なことは、この指導は報酬の減衰と対になっているということです。教師の介入後に得られた正解は減衰した報酬のみを受け取り、教師の助けを報酬への自由な道として政策が扱うことを妨げています。 SocraticPO は、標準の期待報酬目標をそのままにしてロールアウト プロセスのみを変更するため、Reinforce++ などの既存のポリシー勾配バックエンドにプラグインできます。さらに、教師はテキストレベルの指導のみを提供するため、SocraticPO はロジットや分布マッチングへのアクセスを必要とせずに、より強力なブラックボックス教師モデルを活用できます。 SciKnowEval による学部レベルの科学的推論ベンチマークでは、SocraticPO は強力な RL および自己蒸留ベースラインよりも向上しています。アブレーションは、目標を絞った誘導と報酬減衰の両方が必要であり、報酬減衰により矯正補助への依存が軽減されることを示しています。

原文 (English)

SocraticPO: Policy Optimization via Interactive Guidance

Reinforcement learning (RL) for large language models usually supervises reasoning with scalar outcome rewards, such as binary correctness. Such rewards provide an optimization direction but rarely explain how a model should revise its mistaken reasoning, which can encourage shortcut learning and brittle policies. We propose \textbf{SocraticPO} (Socratic Policy Optimization), a policy-optimization framework that augments RL rollouts with Socratic-style natural-language guidance. During rollout, the student first answers independently; if the answer is incorrect, a teacher diagnoses the attempt and provides concise corrective guidance, after which the student continues under the expanded context. Crucially, this guidance is paired with reward decay: correct answers obtained after teacher intervention only receive decayed rewards, preventing the policy from treating teacher help as a free path to reward. Since SocraticPO only modifies the rollout process while leaving the standard expected-reward objective intact, it can be plugged into existing policy-gradient backends such as Reinforce++. Moreover, because the teacher provides only text-level guidance, SocraticPO can leverage stronger black-box teacher models without requiring access to logits or distribution matching. On undergraduate-level scientific reasoning benchmarks from SciKnowEval, SocraticPO improves over strong RL and self-distillation baselines. Ablations show that both targeted guidance and reward decay are necessary, with reward decay mitigating reliance on assisted correction.

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Toward Secure LLM Agents: Threat Surfaces, Attacks, Defenses, and Evaluation

Large language model (LLM) agents are rapidly moving from conversational interfaces to software components that plan, invoke tools, maintai…

13:00 JSTLLM/生成AIエージェント

TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large languag…

13:00 JST研究/論文

Chain of Operators: An Inference-Time Harness for In-Context Operator Learning

While scientific foundation models show immense promise in accelerating physical simulations and numerical forecasting, they remain notorio…

13:00 JSTLLM/生成AI

One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders

Search-augmented LLMs increasingly mediate everyday consumer recommendations by retrieving live web content. This creates a new risk: LLM r…

13:00 JSTハードウェア/半導体

ピクセル化結合器とデュアルステートインピーダンス合成を使用した、ディープラーニング駆動のドハティパワーアンプの逆設計

ドハティ パワー アンプ (PA) の出力結合器は、負荷変調、インピーダンス マッチング、位相補償を単一のネットワーク内に統合しているため、その設計と合成は非常に困難です。この論文では、深層畳み込みニューラル ネットワーク (CNN)、ピクセル化されたレイアウト表現、遺伝的アルゴリズム (GA) をデュアルステート インピーダンス合成と組み合わせて、ピーク電力条件とバックオフ電力条件の両方に対処する 3 ポート ドハティ コンバイナー設計手法を提案します。概念実証として、3 ポートのピクセル化コンバイナーを組み込んだ 2 つの GaN HEMT Doherty PA プロトタイプが設計および製造されました。どちらのプロトタイプも、2.6 ~ 2.8 GHz 内で 71.2% 以上のピーク ドレイン効率で 44.2 dBm を超える実測飽和出力電力を達成しています。さらに、6dB のバックオフ レベルで 64% もの高いドレイン効率が測定されます。デジタル プリディストーションを適用した後、各プロトタイプは -51.3 dBc を超える隣接チャネル漏洩比 (ACLR) を達成しました。

原文 (English)

Deep Learning-Driven Inverse Design of Doherty Power Amplifiers Using Pixelated Combiners and Dual-State Impedance Synthesis

The output combiner of a Doherty power amplifier (PA) integrates load modulation, impedance matching, and phase compensation within a single network, making its design and synthesis highly challenging. In this paper, we propose a three-port Doherty combiner design methodology that combines deep convolutional neural networks (CNNs), pixelated layout representations, and genetic algorithms (GA) with dual-state impedance synthesis to address both peak and back-off power conditions. As a proof of concept, two GaN HEMT Doherty PA prototypes incorporating three-port pixelated combiners are designed and fabricated. Both prototypes achieve a measured saturated output power exceeding 44.2 dBm with peak drain efficiency above 71.2% within 2.6-2.8 GHz. Furthermore, a drain efficiency as high as 64% is measured at the 6-dB back-off level. After applying digital predistortion, each prototype achieves an adjacent channel leakage ratio (ACLR) better than -51.3 dBc.

13:00 JST研究/論文

電気光学電界測定を使用したディープラーニングベースのピクセル化マイクロ波フィルターの設計と特性評価

従来のマイクロ波フィルタ設計は通常、反復的なパラメータ調整と事前定義されたトポロジに依存しているため、設計スペースが制限され、開発時間が増加します。この研究では、畳み込みニューラル ネットワークと遺伝的アルゴリズムを組み合わせた深層学習アプローチを使用して、ピクセル化されたマイクロ波フィルター合成を自動化します。このアプローチを実験的に検証するために、S パラメータと空間電界測定の両方が分析されました。合成されたローパス フィルターは、シミュレーションされたパフォーマンスと測定されたパフォーマンスの間で優れた一致を示し、9.5 GHz を超えて 20 dB 以上の抑制を伴う 7 GHz の通過帯域を達成しました。電気光学測定により、結合伝送線路またはスタブ構造に似た電界パターンが初めて明らかになり、AI によって生成された設計の新たな特性についての洞察が得られました。

原文 (English)

Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-Optical Electric-Field Measurements

Traditional microwave filter design typically relies on iterative parameter tuning and predefined topologies, which limits design space and increases development time. This study uses a deep learning approach combining convolutional neural networks with genetic algorithms to automate pixelated microwave filter synthesis. To validate the approach experimentally, both S-parameter and spatial electric-field measurements were analyzed. The synthesized low-pass filter demonstrated excellent agreement between simulated and measured performance, achieving a 7 GHz passband with over 20 dB suppression beyond 9.5 GHz. Electro-optical measurements, for the first time, revealed electric field patterns that resemble coupled transmission-lines or stub structures, providing insight into the emergent characteristics of AI-generated designs.

13:00 JSTロボティクス

RARM: Confidence-Gated Progress Reward Modeling for RL in Manipulation

Reinforcement learning for robot manipulation is often bottlenecked by reward design, especially in long-horizon tasks: sparse success rewa…

13:00 JSTロボティクス

Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Structure in Passive Object-State World Models

World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and u…

13:00 JSTLLM/生成AI

What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs

Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and po…

13:00 JST研究/論文

Compositional Dynamics in Learning and Mechanics

We give a single compositional setting in which gradient-based learning and Hamiltonian-style mechanics appear as functorial semantics. The…

13:00 JST研究/論文

群等変ポアンカール畳み込みネットワーク

Poincar\'e ResNet のような最近の進歩は、双曲空間で視覚表現を直接学習できる可能性を実証しましたが、その最適化は、リーマン勾配の計算集約的な性質と多様体の厳密な境界によって依然として妨げられています。さらに、標準的な双曲線ネットワークは、同じオブジェクトの空間変換を別個の階層概念として扱うため、パラメーターの冗長な使用と信号の消失につながります。双曲幾何学と離散対称群 ($C_4$ と $D_4$) を組み合わせた等変ポアンカレ ResNet を提案します。ユークリッド等分散を双曲空間に適用する際の重大な障害を特定し、幾何学的に安全なテンソル再整形、双曲群畳み込みの左正則置換、関節配向ポアンカレ中点バッチ正規化を提案します。経験的に、等分散を埋め込むと最適化空間が大幅に減少し、ポアンカレ ボールの境界制約を尊重し、空間グループの等分散を維持しながら収束を加速します。

原文 (English)

Group-Equivariant Poincar\'e Convolutional Networks

While recent methods like that of the Poincar\'e ResNet have demonstrated the ability to learning visual representations directly in hyperbolic space, their optimisation remains a challenge, primarily due to the parameter redundancy of learning distinct orientation filters. In addition, hyperbolic learning exhibits distinct computational overheads that limit their wide use, where efforts to improve their efficiency via optimisation have seen good success, there has been limited exploration into structural priors that enable stronger sample efficiency at training. To address this, we propose Equivariant Poincar\'e ResNets, combining hyperbolic geometry with discrete symmetry groups ($C_4$ and $D_4$). We identify critical roadblocks in applying Euclidean equivariance to hyperbolic space and propose geometrically safe tensor reshaping, left-regular permutations for hyperbolic group convolutions, and joint-orientation Poincar\'e Midpoint Batch normalisation. Empirical evaluations show that embedding equivariance significantly improves the sample efficiency during training which in-turn accelerates convergence while respecting the boundary constraints of the Poincar\'e ball and retaining spatial group equivariance.

13:00 JSTロボティクス

Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies

Deploying a pretrained flow-matching vision-language-action (VLA) policy on a particular robot and workspace often calls for task-specific…

13:00 JST研究/論文

Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting

Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existi…

13:00 JSTLLM/生成AI

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving…

13:00 JST画像/動画生成

GRC-ProbNet: Uncertainty-aware Feature Extraction for Cardiovascular Disease Classification

The automatic detection and classification of cardiovascular disease (CVD) from computed tomography (CT) images plays an important role in…

13:00 JST画像/動画生成研究/論文

Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift

Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remai…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文Gemma

Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels

Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, howe…

13:00 JSTロボティクス

fNIRS に基づくロボット動作の強化学習へのオフライン アプローチ

人間参加型の強化学習は、ロボットの動作をトレーニング、微調整、ユーザーの好みに合わせて調整するための一般的なアプローチになっています。私たちの論文では、機能的近赤外分光法 (fNIRS) を介して脳信号を使用して、シミュレーションでのロボット学習を調整する実現可能性を検討しています。私たちは、受動的な (観察的な) インタラクション タスクと能動的な (実証的な) インタラクション タスクでトレーニングされたエージェントを比較し、置換ではなくパラメータの増強に焦点を当てて、ニューラル信号を使用して RL アルゴリズムを強化するための複数の方法をテストします。さらに、モデルの粒度とノイズがエージェントの学習にどのような影響を与えるかを調査します。私たちの結果は、このフレームワークが効果的であることを示しています。ニューラル信号は、軌道の優先順位と状態アクションの q 値を強化する際の学習を改善します。さらに、このフレームワークはオフライン データから正常に学習し、リアルタイム BCI セットアップが非実用的であるか、限られたデータしか利用できない設定に実用的な代替手段を提供します。

原文 (English)

An offline approach to fNIRS-guided reinforcement learning for robot behavior

Human-in-the-loop Reinforcement Learning has become a popular approach for training, finetuning, and aligning robot behavior with user preferences. Our paper explores the feasibility of using brain signals via functional near-infrared spectroscopy (fNIRS) to modulate robot learning in simulation. We compare agents trained on passive (observational) versus active (demonstrative) interaction tasks, and test multiple methods for enhancing the RL algorithm with the neural signal, focusing on parameter augmentation in contrast to replacement. We further examine how model granularity and noise affect agent learning. Our results show that this framework is effective. The neural signal improves learning when augmenting trajectory priorities and state-action q-targets. Additionally, the framework learns successfully from offline data, offering a practical alternative for settings where real-time BCI setups are impractical or only limited data is available.

13:00 JSTLLM/生成AI

Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits

Large language models now translate natural-language descriptions of decision problems into solver-ready optimization models, and they fail…

13:00 JST研究/論文

Adversarial Robustness of Phishing Email Detection: A Comparative Study of TF-IDF + Logistic Regression and Fine-Tuned DistilBERT

Phishing emails remain one of the most persistent cybersecurity threats, and machine-learning classifiers are widely used to detect them. M…

13:00 JSTLLM/生成AI画像/動画生成

CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders

Frozen encoders are chosen by how well a lightweight head reads a finding from their features, not whether the geometry separates it. Neare…

13:00 JST画像/動画生成エージェント研究/論文

PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image

Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrati…

13:00 JSTエージェントビジネス/資金調達研究/論文

CausalSmith: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference

Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A co…

13:00 JST研究/論文

A Formal Kinetic Theory for Zeroth-Order Newton Dynamics:Stein-Corrected Hessian Estimation and Curvature--Variance Trade-offs

Zeroth-order Newton-type methods are useful when gradients and Hessians are unavailable, but they behave quite differently from first-order…

13:00 JST研究/論文

A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks

Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies…

13:00 JSTLLM/生成AIエージェント

IDP AutoOpt: ドキュメント処理パイプライン構成のエージェント主導の最適化

インテリジェント文書処理 (IDP) パイプラインの高性能構成を検出する自律型 LLM エージェントである IDP AutoOpt を紹介します。現在、IDP プロンプト、モデル、OCR 設定、およびスキーマを合わせて調整するには、ドメイン スペシャリストにドキュメント タイプごとに 20 ~ 80 人時間以上のコストがかかり、企業がドキュメント クラスを追加しても拡張できません。 IDP AutoOpt は閉ループを実行します。つまり、本番の専門知識をコード化した人間が作成したドメイン スキルに基づいて、小さなラベル付きセットで構成をスコアリングし、フィールド レベルのエラーを診断し、対象を絞った編集を生成し、再評価します。ヘルスケア、マーケティング インテリジェンス、金融サービスの設定で導入されている抽出、分類、パケット分割タスク全体で、IDP AutoOpt は人間の専門家の精度と同等かそれを上回る精度を同等またはそれ以下のコストで実現し (抽出ベンチマークでは、ページあたりのコストが 4.6 倍低い場合で 90.2% 対 81.6%)、構成時間を数週間から 2 時間未満に短縮します。さらに、エージェントの LLM 機能には、それを下回ると最適化が失敗するハードしきい値があり、厳選されたドメイン スキルが生のソース コード アクセスよりも優れたパフォーマンスを発揮するため、構造なしで提供するとパフォーマンスが低下する可能性があることを示します。また、コンテキスト管理と差異の軽減に関する実践的なレッスンも共有します。このアプローチは、構成可能なパイプライン、スコアリング関数、および小さなラベル付きセットのみを必要とするため、IDP を超えて、構成が展開のボトルネックになっている RAG やマルチエージェント ワークフローなどの他のエンタープライズ AI システムにも拡張されます。

原文 (English)

IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations

We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pipelines. Tuning IDP prompts, models, OCR settings, and schemas jointly currently costs domain specialists 20 to 80+ person-hours per document type and does not scale as enterprises add document classes. IDP AutoOpt runs a closed loop: it scores a configuration on a small labeled set, diagnoses field-level errors, generates targeted edits, and re-evaluates, guided by human-authored domain skills that encode production expertise. Across extraction, classification, and packet-splitting tasks deployed in healthcare, marketing-intelligence, and financial-services settings, IDP AutoOpt matches or exceeds human-expert accuracy at equal or lower cost (on an extraction benchmark, 90.2% vs 81.6% at 4.6 x lower per-page cost), cutting configuration time from weeks to under two hours. We further show that agent LLM capability has a hard threshold below which optimization fails, and that curated domain skills outperform raw source-code access, which can degrade performance when provided without structure. We also share practical lessons on context management and variance mitigation. Requiring only a configurable pipeline, a scoring function, and a small labeled set, the approach extends beyond IDP to other enterprise AI systems, such as RAG and multi-agent workflows, where configuration bottlenecks deployment.

13:00 JSTLLM/生成AI画像/動画生成

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs

Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet…

13:00 JSTロボティクス

Latency-Tolerant Cloud-Edge Collaborative Vision-Language-Action Models via Emergent Representational Specialization

Deploying billion-parameter Vision-Language-Action (VLA) policies on mobile robots creates a systems conflict: semantic reasoning benefits…

13:00 JST研究/論文

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- whi…

13:00 JST研究/論文

Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal reference

Whether large language models perform algorithmic inference or pattern completion is hard to test, because most benchmarks supply answers b…

13:00 JSTLLM/生成AI

Divisive Normalization Shapes Low-Rank Slow Manifolds for Continuous Working Memory

The ability to robustly maintain and update continuous variables is a hallmark of working memory. While classical continuous attractor netw…

13:00 JSTLLM/生成AIエージェント

Your Agentic LLMs Secretly Encode Indirect Prompt-Injection Exposure in Hidden States

Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While m…

13:00 JST研究/論文

AI Alignment and Fiduciary Obligation

Advanced AI assistants engage users in extended interactions across a widening range of roles, including advice, decision support, collabor…

13:00 JSTLLM/生成AIエージェント

Breadcrumbing Search Agents

LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical se…

13:00 JST研究/論文

SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery

Bayesian causal discovery seeks to determine the posterior distribution of causal theories, which are interpreted as directed acyclic graph…

13:00 JSTLLM/生成AIハードウェア/半導体

Answer First, Reason Later: When Commitment Order Costs Accuracy in Diffusion Language Models

Masked diffusion language models revise many masked output positions in parallel. We call a token committed once it becomes visible and is…

13:00 JSTLLM/生成AI

Hidden Language Consistency Phenomena in Reasoning LLMs

Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the int…

13:00 JST画像/動画生成エージェント

Population-Scalable Multi-Agent World Modeling

World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent…

13:00 JSTLLM/生成AIハードウェア/半導体

Withholding the Completing Chunk: Exact Release-Boundary Equivalence for Production Streaming Guardrails

Streaming language-model output creates an enforcement boundary: a control that detects a prohibited pattern after releasing its completing…

13:00 JST研究/論文

TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation

Multi-modal recommenders fuse user-item interaction signals with item modalities such as text, images, and audio, but the usefulness of eac…

13:00 JSTLLM/生成AI

PatientAct: 理論に基づいたメンタルヘルス クライアント シミュレーション

LLM ベースの模擬クライアントは、初心者カウンセラーのトレーニング、LLM セラピストの評価、合成データの生成に使用されることが増えています。しかし、現在のシミュレーターは、過度に協力的なクライアントを生み出し、あまりにも簡単に開示し、抵抗なく治療の再構成を受け入れ、単一のセッション内で中核的な問題を解決します。これらの問題は、因果関係の深さが欠けているプロファイルと、すべてのコンテンツを同等にアクセスできるものとして扱う動作メカニズムに起因すると考えられます。確立された臨床理論に基づいたクライアント シミュレーションのフレームワークである PatientAct を紹介します。当社のプロファイルは 5P の臨床症例定式化を統合しており、設計を単一の治療法に結び付けることなく因果関係の深さを提供します。シミュレーション中、プロファイルには、項目が信頼しきい値を保持する動的記憶層が含まれます (たとえば、症状は早期に入手可能ですが、形成的記憶には持続的な治療連携が必要です)。各ターンで、クライアントの感情的な反応と行動がモデル化されてから、応答が生成されます。セラピストがゲートされたコンテンツにアプローチする場合、PatientAct は、協力や単一の抵抗パターンをデフォルトとするのではなく、量、内容、スタイルの観点から抵抗を表現します。私たちは 40 の臨床状況に関するフレームワークを評価し、それが臨床的妥当性の高い多様なプロファイルを生成することを実証します。さらに、PatientAct はベースラインを大幅に上回り、耐性の質と行動の現実性が大幅に向上しました。私たちのコードとデータは、github.com/Sahandfer/PatientHub 経由で公開されます。

原文 (English)

PatientAct: Theory-Grounded Mental Health Client Simulation

LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes without resistance, and resolve core issues within a single session. We trace these issues to profiles that lack causal depth and behavioral mechanisms that treat all content as equally accessible. We present PatientAct, a framework for client simulation grounded in established clinical theories. Our profiles integrate the 5Ps clinical case formulation, providing causal depth without tying the design to any single therapeutic modality. During simulation, profiles include a dynamic memory layer in which items carry trust thresholds (e.g., symptoms are available early, whereas formative memories require a sustained therapeutic alliance). At each turn, the client's emotional reaction and behavior are modeled before generating a response. If the therapist approaches gated content, PatientAct expresses resistance in terms of quantity, content, and style rather than defaulting to cooperation or a single resistance pattern. We evaluate our framework on 40 clinical situations and demonstrate that it generates diverse profiles with high clinical plausibility. Moreover, PatientAct significantly outperforms the baselines, yielding substantial gains in resistance quality and behavioral realism. Our code and data are publicly available via github.com/Sahandfer/PatientHub.

13:00 JSTLLM/生成AIエージェント

Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce

Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much…

13:00 JSTエージェントLlamaQwen

GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix

Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload ha…

13:00 JSTLLM/生成AI

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning tr…

13:00 JST規制/政策

Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors

Various machine unlearning techniques have been developed in response to privacy legislation requirements, enabling individuals to exercise…

13:00 JSTLLM/生成AI

CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method

Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions…

13:00 JSTロボティクス

GigaBrain-WBC-0.5: 環境との相互作用を伴う堅牢な全身制御のための行動世界モデル

全身動作追跡ポリシーは、ヒューマノイドを堅牢な制御インターフェイスに変えます。遠隔操作者 (または上流モデル) は粗い動きの意図のみを提供しますが、低レベルのポリシーはロボットのバランスを保ち、物理的に実行可能に保ちます。既存のトラッカーは、平坦な地面でのみこのインターフェイスを提供します。空のシーンでトレーニングされ、地形やオブジェクトとの接触がそのダイナミクスをどのように再形成するかを学習することはなく、参照モーション コーパスを継続的に拡大することで、あらゆるコマンドの下でバランスをとるようにポリシーを教えようとしますが、実行可能な動作が環境に依存するようになると機能しなくなります。我々は、人型全身制御のための最初の行動世界モデル(BWM)であるGigaBrain-WBC-0.5を紹介します。純粋に反応的なトラッカーではなく、次の動作、次の状態、次の潜在的な動作コマンドの分布を共同で予測するように因果的 Transformer をトレーニングします。そのため、動作するネットワークは、環境が次に実行できることをどのように形成するかもモデル化します。自動地形アノテーション パイプラインは、リターゲットされたモーションから完全な 3D 接触ジオメトリを復元し、既存のモーション データセットのスケールで地形アノテーションを可能にします。予測された分布は展開時に再利用され、オンラインでありえないコマンドを検出し、学習した動作に反映させるため、ロボットは「ベストエフォート」方式でタスクを試みます。その結果、リアルタイムのコマンドを受け取り、環境と対話し、信じられないコマンド、落下、外乱に対して堅牢性を維持する統合ポリシーが実現します。 GigaBrain-WBC-0.5 は、3 つの大規模トラッカー ベースラインの中で、4 つのレジームすべてで最高の成功率を達成しました。地形インタラクションで 81.3% (最も強力なベースラインの 4.3 倍)、信じられないコマンドでの 83.1%、および落下からの回復 99.3% (最も強力なベースラインの 16.8 倍) です。ハードウェアの試験では、サポートの欠如や障害下での堅牢な相互作用が示されています。 Unitree G1 チェックポイントは、簡単な微調整で Maker L01 ロボットに転送されます。

原文 (English)

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging the reference-motion corpus, which stops working once feasible behaviors become environment-dependent. We present GigaBrain-WBC-0.5, the first Behavior World Model (BWM) for humanoid whole-body control. Rather than a purely reactive tracker, we train a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next. An automatic terrain-annotation pipeline recovers full 3D contact geometry from retargeted motion, enabling terrain annotation at the scale of existing motion datasets. The predicted distribution is reused at deployment to detect implausible commands online and retract them onto learned behaviors, so the robot attempts tasks in a "best-effort" manner. The result is a unified policy that takes real-time command, interacts with environment, and stays robust to implausible commands, falls, and disturbances. GigaBrain-WBC-0.5 achieves the highest success rate across all four regimes among three large-scale tracker baselines: 81.3% on terrain interaction (4.3x the strongest baseline), 83.1% under implausible commands, and 99.3% fall recovery (16.8x the strongest baseline). Hardware trials show robust interaction under missing supports and disturbances; the Unitree G1 checkpoint transfers to the Maker L01 robot with simple fine-tuning.

13:00 JST研究/論文

Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas

We present the first mechanised formalisation of Romanov's Triplet Logic (TLS) in the Rocq proof assistant. TLS is a combinatorial framewor…

13:00 JSTLLM/生成AI

Aslema at NADI 2026: Data Augmentation for Intent Recognition and Slot Filling

We present Aslema, our system for NADI 2026 Shared Task 5, which consists of two subtasks: intent recognition and slot filling. We evaluate…

13:00 JST研究/論文

AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL

Clifford circuits play a foundational role in quantum computing, particularly due to their importance in quantum error correction and fault…

13:00 JST研究/論文

Interpretable AI predicts a 2026 summer dry anomaly in central China

Seasonal precipitation anomalies are largely regulated by atmospheric circulation, which dynamical models predict with greater reliability…

13:00 JSTLLM/生成AIエージェントOpenAI

SPADE: Self-Play in Adaptive Synthetic Executable Environments

Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing train…

13:00 JSTLLM/生成AIエージェント

残りの耐用年数におけるマルチモーダル言語モデルの基礎を築くための時系列検索

大規模言語モデル (LLM) とエージェント AI システムは、ドメイン固有のメンテナンスと予測タスクのためにますます検討されており、それらが予測と健康管理 (PHM) を効果的にサポートできるかどうかという疑問が生じています。この論文では、時系列検索に基づいたマルチモーダル大規模言語モデル (MLLM) を使用した残存耐用年数 (RUL) の推定を調査します。私たちは、歴史的に類似した劣化セグメントをトレーニング セットから取得し、テスト軌跡とともに、構造化されたマルチモーダル プロンプトを通じて MLLM によって処理される視覚的な比較アーティファクトに変換するフレームワークを提案します。このアプローチは、ランダムな参照選択に基づく非検索ベースラインに対して検索ベースの推論を比較する繰り返し実験の下で、C-MAPSS ベンチマークの FD001 パーティションで評価されます。結果は、時系列取得により、評価されたモデル全体で MLLM ベースの RUL 予測が一貫して向上し、エラーが減少し、パフォーマンスがより安定していることがわかります。同時に、利点の大きさはモデルの能力に依存し、基礎となる MLLM が取得した証拠を利用できる場合に取得が最も効果的であることを示しています。全体として、この研究は、時系列 RAG がマルチモーダルな予後推論を改善するための有望なメカニズムであることを示していると同時に、実際の PHM 設定における MLLM ベースの RUL 推定の現在の限界も強調しています。

原文 (English)

Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life Prediction

Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks, raising the question of whether they can effectively support prognostics and health management (PHM). In this paper, we investigate remaining useful life (RUL) estimation with multimodal large language models (MLLMs) grounded through time-series retrieval. We propose a framework in which historically similar degradation segments are retrieved from the training set and, together with the test trajectory, transformed into a visual comparison artifact that is processed by the MLLM through a structured multimodal prompt. The approach is evaluated on the FD001 partition of the C-MAPSS benchmark under repeated experiments comparing retrieval-based inference against a non-retrieval baseline based on random reference selection. The results show that time-series retrieval consistently improves MLLM-based RUL prediction across the evaluated models, yielding lower error and more stable performance. At the same time, the magnitude of the benefit depends on model capacity, indicating that retrieval is most effective when the underlying MLLM is able to exploit the retrieved evidence. Overall, the study shows that time-series RAG is a promising mechanism for improving multimodal prognostic reasoning, while also highlighting the current limitations of MLLM-based RUL estimation in practical PHM settings.

13:00 JST研究/論文

アクティブスパイク知覚: いつでも 3D 点群認識のための信念状態としての膜の可能性

スパイク点群ネットワークは通常、固定された入力に依存しない順序で空間をスキャンするため、スパイク計算の最も特徴的なリソースである膜電位の時間的発展が意思決定の場として使用されずに残ります。 Active Spiking Perception (ASP) は、3D 認識を反復的な意思決定プロセスとして再構築します。このプロセスでは、ネットワーク自身の漏洩統合発火 (LIF) 膜電位がクラス全体にわたる実行中の信念として読み取られ、観察する次のチャンクを選択し、信頼マージンの早期終了をトリガーします。軽量のスライス選択ポリシーは、メンブレン状態と事前計算された幾何学的記述子から未訪問の最遠点サンプリングされたチャンクをスコアリングし、ストレートスルーのガンベル ソフトマックスを通じてエンドツーエンドでトレーニングし、推論時に argmax に削減し、バックボーン パラメーターの約 2% を追加します。漏れのある積分はベイジアン フィルターの再帰的対数事後更新であること、出口ルールは停止時に多重テストのペナルティなしで分布フリーの選択的リスクを達成すること、ストリーミング状態の繰り越しは有限精度ドリフトを伴うプレフィックス再計算とまったく同等であることを証明します。 ASP は、ModelNet40 と ModelNet10 で 90.62% と 93.28% に達し、大規模なバックボーンでの最も強いスパイク ベースラインを 1.7 ポイント下回っていますが、ベースラインでは提供されていない認定済みの常時インターフェイスが追加されています。このメカニズムは変更を加えずに高密度予測に転送し、ShapeNetPart で 83.21 インスタンス mIoU、S3DIS エリア 5 で 48.50 mIoU を与えます。我々の知る限りでは、S3DIS エリア 5 で最初のスパイク結果が得られ、チャンク選択を置き換える固定が中心窩非スパイク変換器に適用されるため、ポリシーはスパイク バックボーンに結び付けられません。コストは観測では正確に線形であり、しきい値は測定されたコンピューティング ダイヤル スパンです。エネルギーが 2.8 倍から 1.35 倍少なくなります。 1 つの具体的な制限があります。1 つの S3DIS クラスが、使用するクロップ サイズでは識別できないため、それを修正する予測を示します。

原文 (English)

Active Spiking Perception: The Membrane Potential as a Belief State for Anytime 3D Point Cloud Recognition

Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distinctive resource of spiking computation, the temporal evolution of the membrane potential, unused as a locus of decision-making. Active Spiking Perception (ASP) recasts 3D recognition as an iterative decision process in which the network's own leaky integrate-and-fire (LIF) membrane potential, read as a running belief over the class, selects the next chunk to observe and triggers confidence-margin early exit. A lightweight Slice-Selection Policy scores unvisited farthest-point-sampled chunks from the membrane state and precomputed geometric descriptors, trains end-to-end through a straight-through Gumbel-Softmax, reduces to an argmax at inference, and adds about 2% of backbone parameters. We prove that leaky integration is the recursive log-posterior update of a Bayesian filter, that the exit rule attains distribution-free selective risk with no multiple-testing penalty at the stopping time, and that streaming state carry-forward is exactly equivalent to prefix recomputation with bounded finite-precision drift. ASP reaches 90.62% and 93.28% on ModelNet40 and ModelNet10, 1.7 points below the strongest spiking baseline at a larger backbone, while adding a certified anytime interface no baseline offers. The mechanism transfers unchanged to dense prediction, giving 83.21 instance mIoU on ShapeNetPart and 48.50 mIoU on S3DIS Area 5, to our knowledge the first spiking results on S3DIS Area 5, and, fixation replacing chunk selection, to a foveated non-spiking transformer, so the policy is not tied to spiking backbones: cost is exactly linear in observations and the threshold is a measured compute dial spanning 2.8x to 1.35x less energy. One limitation is concrete: one S3DIS class is unidentifiable at the crop size we use, and we give the prediction that would fix it.

13:00 JST画像/動画生成

VGI-Bench: Probing Visual Intelligence in Video Generation Models

Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet r…

13:00 JST画像/動画生成

Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion Synthesis

Full-cycle biventricular geometry is essential for characterizing cardiac function. However, dense and temporally consistent 3D+t biventric…

13:00 JSTLLM/生成AI画像/動画生成ビジネス/資金調達Claude

Question-Guided Evidence Acquisition for Multimodal Visual Question Answering

Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still…

13:00 JSTLLM/生成AIエージェント

MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents

Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Exi…

13:00 JSTLLM/生成AI

Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection

We present a novel approach to efficient LLM harness optimization through adaptive validation task selection. Harness optimization iterativ…

13:00 JST画像/動画生成

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elemen…

13:00 JST研究/論文

Atom Learning Model (ALM): how a real classroom got tokenised

The Atom Learning Model (ALM) tokenises a school curriculum. Two secondary mathematics textbooks were read by machine into 1,934 atoms, eac…