Skip to the content.

AIニュース 2026-08-25

自動生成: 2026-08-25 10:36 JST

← トップに戻る

過去24時間以内に公開された記事を、同じ話題ごとに1つのストーリーカードへまとめ、出典・トピック・要約とともに掲載しています。要約は各フィード提供文の冒頭を整形したもので、本文は各リンク先をご覧ください。

📌 今日の要点 TOP7

  1. Advancing price-performance for developers with GPT‑5.6 in KiroOpenAI

    GPT‑5.6 is now available in Kiro, helping developers plan, build, rev…

  2. 建築業でも「Claudeの有料導入」が急増 AI導入レベルが上がった企業を悩ます「新たな壁」とはITmedia AI+

    建築AI経営研究会は、「建築AI経営実態調査」の結果を公開した。特定部署以上でAIを活用する企業が68.5%に達したことが判明した。課題は…

  3. ChatGPT、Gemini、Claude Sonnet 5の“最強”は? 「速度」と「安定性」を実測して比べてみたITmedia AI+

    主要なAIチャットサービスの「ChatGPT」「Gemini」「Claude Sonnet 5」の中で、応答速度が一番速いのはどれなのか。…

  4. ランボルギーニのSDV開発やAI設計改善も、PTCが“アジア初”デモ機公開ITmedia AI+

    PTCジャパンは日本橋の新オフィスに開設したデモセンター「JXC」において、ランボルギーニのSDV開発や建機のAI設計改善のデモンストレー…

  5. 「今からロボットを現場投入できないか?」 増えた日本企業からの問い合わせ “頭脳”を作る韓国企業が感じた国内の変化ITmedia AI+

    日本にも拠点を置くRLWRLD(リアルワールド)は、フィジカルAI分野で注目を集める国際企業だ。産業用ロボットの導入もかなり進んでいる日本…

  6. Situational Awareness, star AI hedge fund that nearly imploded, now being probed by the SECTechCrunch AI

    The AI hedge fund went from "the talk of Wall Street" to "subject of…

  7. 中国で“ロボットの運動会” 100メートル走で“ボルト超え”……で、走るのは速いけど仕事はできるの?ITmedia AI+

    中国で人型ロボットの競技大会が開催された。100メートル走では、人類最速のウサイン・ボルト選手が持つ記録を上回った。しかし、走るのが速いと…

トピック別件数

日本語メディア11件

ITmedia AI+ (日本語)

10:00 JSTロボティクス

「今からロボットを現場投入できないか?」 増えた日本企業からの問い合わせ “頭脳”を作る韓国企業が感じた国内の変化

日本にも拠点を置くRLWRLD(リアルワールド)は、フィジカルAI分野で注目を集める国際企業だ。産業用ロボットの導入もかなり進んでいる日本市場で最近起こった変化について聞いた。

08:15 JSTその他

ランボルギーニのSDV開発やAI設計改善も、PTCが“アジア初”デモ機公開

PTCジャパンは日本橋の新オフィスに開設したデモセンター「JXC」において、ランボルギーニのSDV開発や建機のAI設計改善のデモンストレーションを披露した。

08:00 JSTLLM/生成AI研究/論文Claude

建築業でも「Claudeの有料導入」が急増 AI導入レベルが上がった企業を悩ます「新たな壁」とは

建築AI経営研究会は、「建築AI経営実態調査」の結果を公開した。特定部署以上でAIを活用する企業が68.5%に達したことが判明した。課題は導入初期の試行錯誤から「新たな壁」にシフトしている。

07:00 JSTロボティクス

中国で“ロボットの運動会” 100メートル走で“ボルト超え”……で、走るのは速いけど仕事はできるの?

中国で人型ロボットの競技大会が開催された。100メートル走では、人類最速のウサイン・ボルト選手が持つ記録を上回った。しかし、走るのが速いとはいえ実際の現場作業で役に立つのだろうか。

17:55 JSTLLM/生成AIエージェントAnthropicClaude

【復旧済み】「Claude」で障害発生中

米AnthropicのAIサービス「Claude」「Claude Code」などで障害が発生している。同社のステータスサイトによると、日本時間の8月24日午後2時27分時点でエラーの原因を特定し、修正作業中という。

16:30 JSTLLM/生成AIハードウェア/半導体ClaudeNVIDIAAlibaba

なぜいま30Bクラスのオープンモデルが“熱い”のか 「27BパラメータでOpus 4.6超え」も

MetaとNVIDIA、Alibabaが30Bクラスのオープンモデルを相次ぎ公開し、国産の「LLM-jp-4 33B」も登場した。このサイズのモデルが今“熱い”理由と留意点を考察する。

13:30 JSTエージェントビジネス/資金調達

AIエージェント活用を阻む“予算の壁” セールスフォースらの新機能は打開策になり得るか?

企業にとってはあらかじめAI利用コストを想定できないと投資が決まらない。そうなると予算が立てられず、ROIの見通しもつかない。そんな企業のAIエージェント活用に立ちはだかる“予算の壁”を打破するにはどうすればよいのか。セールスフォース、IBM、マネーフォワードが提供を開始した最…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTGemini

ChatGPT、Gemini、Claude Sonnet 5の“最強”は? 「速度」と「安定性」を実測して比べてみた

主要なAIチャットサービスの「ChatGPT」「Gemini」「Claude Sonnet 5」の中で、応答速度が一番速いのはどれなのか。時間帯によって速度は変わるのか。測定値はどれだけ安定しているのか。実測を基に検証する。

13:00 JSTLLM/生成AI

生成AIの品質を“AIで測る”――「LLM as a Judge」を機能させる3つの要素

AIの出力を別のAIが評価する手法「LLM as a Judge」。その基礎をdotDataがブログで解説した。AIにAIを評価させながら、その品質を確保するにはどのような手法が有効なのか。

12:30 JSTその他

AI×設計開発ニュースまとめ(2026年5~6月)

MONOistに掲載した主要な記事を、読みやすいPDF形式の電子ブックレットに再編集した「エンジニア電子ブックレット」。今回は、設計や解析など製品開発の現場で活用が進みつつあるAI関連のニュースをまとめた「AI×設計開発ニュースまとめ(2026年5~6月)」をお送りします。

12:17 JSTその他

Sakana AI、防衛省の「情報分析」をAIで支援 自衛隊の指揮統制システム高度化に続き

Sakana AIは、防衛省における情報の分析業務を支援するAIの開発に関する契約を締結したと発表した。

海外メディア7件

TechCrunch AI (英語)

09:23 JSTその他

Situational Awareness, star AI hedge fund that nearly imploded, now being probed by the SEC

The AI hedge fund went from "the talk of Wall Street" to "subject of federal subpoenas" faster than you can say "diversify your portfolio."

06:24 JSTビジネス/資金調達

Trump bought SpaceX shares two weeks after blockbuster IPO

The president bought when the stock was in the mid-$150 range. SpaceX finished trading on Monday back at its IPO price of $135.

04:54 JSTその他

Amjad Masad, CEO and co-founder of Replit, joins the Disrupt Stage at TechCrunch Disrupt 2026

At TechCrunch Disrupt 2026, Replit CEO Amjad Masad will share his perspective on the future of programming and Replit's role in developing…

03:03 JSTその他

Instinct’s powerful AI assistant is raising privacy and security concerns

Early testers are raving about what Instinct can do, but some say the AI assistant’s sweeping access, broad terms and ability to act on use…

00:24 JSTエージェントロボティクスビジネス/資金調達

Valor, Point72 back General Intuition at $6B valuation as AI startup pushes into robotics

General Intuition, the startup building a foundation model that trains generalized AI agents how to move through space and time, is in talk…

00:00 JSTLLM/生成AIエージェントOpenAI

OpenAI is building AI agents for everything. Will everyone use them?

Inside the frontier lab’s push to bring AI agents from software engineers to the masses.

22:47 JSTビジネス/資金調達

Hugging Face reportedly in talks to be acquired for $13B

Hugging Face has reportedly been fielding acquisition offers that would value the company at around $13B. But with the founders' feeling of…

公式ブログ1件

OpenAI (英語)

21:00 JSTLLM/生成AIGPT / ChatGPT

Advancing price-performance for developers with GPT‑5.6 in Kiro

GPT‑5.6 is now available in Kiro, helping developers plan, build, review, and test software with better price-performance.

論文290件

arXiv cs.AI (英語)

13:00 JSTエージェント

SDAD: Spec-Driven Agentic Development for the AI-Native SDLC

Frontier coding agents backed by large language models with context windows from hundreds of thousands to millions of tokens are restructur…

13:00 JSTLLM/生成AIエージェントAnthropicClaude

PrimeAgentOrchestrator: Memory-Primed Agent Spawning for Personal AI Infrastructure

Large language model (LLM) coding agents start each session with an empty context window, discarding accumulated knowledge from prior work.…

13:00 JSTLLM/生成AIGemma

Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification

Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages…

13:00 JSTLLM/生成AIエージェント研究/論文

A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has…

13:00 JST研究/論文

Interpretable Multimodal Classification with Linear Discriminant Tree Ensembles

Multimodal affect and behaviour classifiers that fuse heterogeneous text, audio, and visual streams must simultaneously achieve competitive…

13:00 JSTエージェント

Representation Affects Retrieval: A Case Study of Skill Discovery and Routing in a Multimodal Agent Harness

A production agent harness must discover and rank, from a growing library of skills, the one most appropriate for a user's task. At small s…

13:00 JSTLLM/生成AIエージェント

Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory

Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic…

13:00 JST研究/論文

Environmental Slow AI: Design Principles for Generative Systems

Generative AI (genAI) systems produce cultural artefacts at scale, but they also reflect embedded cultural values through their design. Onc…

13:00 JSTLLM/生成AIエージェント

When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory

Agentic memory under a fixed budget involves two stages: retention and retrieval. Existing retrieval-centered paradigms implicitly assume n…

13:00 JSTエージェント

World models of environment, agent and joint agent-environment systems

World models are a central component of model-based reinforcement learning. They are usually discussed in terms of what variables they pred…

13:00 JST画像/動画生成研究/論文ClaudeOpenAIGPT / ChatGPT

StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models

Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure f…

13:00 JST研究/論文

Categorical AI phenomenology: A first-person approach

This paper develops a phenomenology-first approach to artificial consciousness by reframing consciousness as the subjective experience enac…

13:00 JSTエージェント

Who Delegates to AI? Evidence from 53,000 Agent Configurations

A growing literature measures how far occupations are exposed to AI, but these measures capture where AI could perform tasks, not whether w…

13:00 JST研究/論文

STCO: Conditional Neural Operators for Time-Dependent PDEs

Neural operators have emerged as efficient surrogates for time-dependent physical systems governed by partial differential equations (PDEs)…

13:00 JSTエージェント

Terminal Agents: A Survey of AI Agents in Command-Line Environments

Large language model agents increasingly act through terminals, yet existing surveys disperse terminal-mediated behavior across software en…

13:00 JST研究/論文

Lost in Translation: How Universal Ethical Values Fail to Translate Across Global Contexts

AI ethics frameworks treat values such as fairness, transparency, and accountability as universal and uniformly operationalizable across co…

13:00 JST研究/論文

A Temporal Planning Approach for Intelligent Flood Response

Effective response to multiple, simultaneously flooded areas requires coordinating appropriate actions in the correct temporal order, under…

13:00 JSTLLM/生成AIエージェント

FL-MAESTRO: Multi-Agent LLM Orchestration for Resource-Constrained Federated Learning

In Federated Learning (FL), the communication topology is a runtime variable rather than a fixed design choice, since links and edge device…

13:00 JSTLLM/生成AI

Volumetric Radiology AI in the Era of Multimodal Large Language Models

Advances in multimodal large language models (MLLMs) are extending radiological artificial intelligence (AI) beyond task-specific image ana…

13:00 JSTLLM/生成AIエージェント

Consilience: Conformally Calibrated Communication Control for Hidden-Profile Multi-Agent Reasoning

Multi-agent LLM systems can improve reasoning by pooling diverse perspectives, but their effectiveness depends on coordinating communicatio…

13:00 JSTLLM/生成AI

Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation

Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough mod…

13:00 JST研究/論文Grok

FlavourBench: Ranking Frontier Language Models with Executable Culinary Ground Truth

Open-ended language-model benchmarks usually inherit a judge: a human preference panel, another model, or a brittle exact-match key. We int…

13:00 JST研究/論文

Difficulty-Aware Semantic-ID Optimization for Generative Recommendation

Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over hierarchical item identifiers. A…

13:00 JSTエージェントビジネス/資金調達NVIDIA

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

Enterprise agent programs are moving from prototypes into production, where reusable skills, tools, and workflow packages must be reviewed…

13:00 JSTLLM/生成AIエージェント

Dual-Cache Latent Space Communication between Heterogeneous Language Models

Multi-agent LLM systems split work across models, so answering often requires knowledge that sits in another agent's context: a Sharer has…

13:00 JSTLLM/生成AIAnthropic

Applying Anthropic Primitives at Large Enterprises: Harness Paradigm for Knowledge Work

Frontier models have collapsed the cost of writing custom code: a niche problem a specialist sees in their own domain now costs an afternoo…

13:00 JST研究/論文

SAGE: A Unified Algebra and Self-Adaptive Execution for AI Functions in SQL

SQL systems increasingly expose AI functions for tasks such as classification, extraction, filtering, ranking, retrieval, joining, and summ…

13:00 JSTLLM/生成AIエージェントGemmaLlama

Weighted Memory Tree: Remembering What Matters for Long-Horizon LLM Agents

Large language model (LLM) agents have demonstrated the ability to solve multi-step tasks requiring planning, tool use, and external inform…

13:00 JST規制/政策

Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions

Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces, recommender systems and beyond, must choo…

13:00 JSTLLM/生成AI

Auditable by Construction: An Ontology-Driven Framework for Trustworthy LLM Analytics in Enterprise Finance

Enterprise adoption of large language models in finance is constrained less by fluency than by trust: in Financial Planning and Analysis (F…

13:00 JSTエージェント研究/論文

DreamBench-SWE: A Multi-Session Memory-Hygiene Benchmark for Software Agents

DreamBench-SWE is a multi-session benchmark for software-agent memory hygiene in which later software tasks depend on non-inferable evidenc…

13:00 JSTLLM/生成AIエージェント

Why2Speak: Faithful Reasoning for Abstaining Action Policies

Many agentic systems must repeatedly choose between acting and abstaining, making faithful reasoning important for oversight: an explanatio…

13:00 JST研究/論文

CDRL: Certification-Driven Reinforcement Learning for Neutrino Flavor Model Discovery

Many scientific discovery problems require searching combinatorial hypothesis spaces under complex domain constraints. Reinforcement learni…

13:00 JSTエージェント

VortexChat: An agentic framework for autonomous multi-objective integrated photonic design

The advancement of modern integrated photonics is frequently bottlenecked by device design workflows that rely heavily on manual simulation…

13:00 JST研究/論文GemmaMistral AIQwen

DirEAG: Dirichlet Evidence Aggregation for Calibrating Verbalized Confidence in Mathematical Reasoning

Reliable confidence estimation is essential for using large language models in mathematical reasoning, but black-box verbalized confidence…

13:00 JSTLLM/生成AIエージェント

Calibrating Criterion Revision in LLM Agents: Failure Modes and a Trace-Anchored Protocol

Language-model agents can improve after failure or carry text across episodes without revising what counts as success. We study the narrowe…

13:00 JSTロボティクス

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

Manipulating moving objects requires a policy to anticipate contact events, yet vision-language-action (VLA) policies are commonly fine-tun…

13:00 JST研究/論文

Continuous-Time Quantum Walks based Graph Neural Network

Graph Neural Networks (GNNs) are widely used on graph-structured data, but most suffer from two key weaknesses. First, message passing beha…

13:00 JST研究/論文

Is Multimodal Speculative Decoding Ready for Diffusion-Based Parallel Drafting? A Survey and Empirical Diagnosis

Speculative decoding accelerates autoregressive generation by allowing a lightweight drafter to propose future tokens while a target model…

13:00 JST研究/論文GPT / ChatGPT

Natural-Language-Guided Generator-Agnostic Shortlisting for Protein Binder Design

Modern de novo design workflows generate many candidate protein binders, but wet-lab validation capacity remains limited, making shortlisti…

13:00 JST研究/論文Gemma

Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and th…

13:00 JSTエージェント

CAS: Conformalized Agentic Search via Adaptive Retrieval and Policy Weighting

Search Agents face a severe reliability crisis during reinforcement learning (RL) fine-tuning. Heuristic Top-K retrieval often causes criti…

13:00 JSTエージェント

Structure for Reading, Prose for Writing: Asymmetric Structural Conditioning in Multi-Agent Document Authoring

Multi-agent pipelines that author formal documents must both read a requester's forms and write against them. We report a deployed tender-r…

13:00 JSTLLM/生成AI

Knowing but Not Saying: Preventing Factual Access Failures in LLM SFT via Recall-Anchored Distillation

Supervised fine-tuning (SFT) can degrade factual behavior outside the target domain. This degradation is often described as catastrophic fo…

13:00 JSTエージェントビジネス/資金調達

Automated Trajectory Evaluation for Mobile Agents via Step-Level Consequence Reasoning and Aggregation

Evaluating language-guided mobile agents has recently shifted from rule-based to model-based approaches to achieve scalable and automated a…

13:00 JST研究/論文

Dynamic Context Scheduling: Learning Beyond the Static Universe

We study dynamic context scheduling as a training instrument for contextual re- inforcement learning. Rather than treating intra-episode co…

13:00 JST研究/論文

SPARC: Single-Pass Scaling for Motion Forecasting with Conformal Bayesian Last Layers

Human motion forecasters are increasingly accurate and fast, but reliable deployment requires uncertainty estimates that are structured, ca…

13:00 JST研究/論文

Neuro-Geospatial Modelling of EEG Affective States Using Literature-Informed Environmental Context

Environmental exposures such as air pollution and greenness have been associated with affective and cognitive outcomes, but EEG and environ…

13:00 JSTLLM/生成AI

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context. Existing ce…

13:00 JST研究/論文

Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress

Artificial intelligence systems increasingly make consequential judgments - which patient is deteriorating, which building is safe to enter…

13:00 JST研究/論文

Foundation Models for Partial Causal Identification

This paper investigates the development of causal foundation models for bounding the effect of interventions and counterfactuals from obser…

13:00 JSTエージェント

TRACE: Agentic Catalog Enrichment with Multi-source Evidence Grounding

Product catalogs underpin search, discovery, and recommendation in e-commerce, yet they are often attribute-sparse: the attributes shoppers…

13:00 JST研究/論文

RAG Deserves an Index: Why Ingest-Time Compilation Beats Query-Time Interpretation

Nearly every retrieval-augmented question-answering system in production ships with a hidden interpreter: on each query a language model re…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

MGAL: A Multilingual Granularity-Aware Long-Context Benchmark

Evaluation of long-context Large Language Models (LLMs) has advanced rapidly. However, most existing benchmarks are limited to the document…

13:00 JST研究/論文

Coverage-Driven Verification for Safety-by-Design in AI-Based Collision Avoidance Systems

Artificial Intelligence (AI) offers significant potential for future aviation systems; however, its integration into safety-critical applic…

13:00 JST研究/論文

ReCurveflow: A Flow Matching Framework that Learns Curved Reaction Trajectories to Predict Transition State Geometries

Predicting transition states (TS) in chemical reactions is crucial, as they provide insights into reaction mechanisms. Recent work on TS pr…

13:00 JSTLLM/生成AI研究/論文Qwen

UpgradeBench: A Decision-Centric Benchmark for Upgrading Fine-Tuned LLM Specialists

Organizations maintain task-specific adapters for open-weight language models, and each new base-model release forces a migration decision:…

13:00 JSTロボティクス

Graph-Operator World Models for Morphology-Parameter Generalization in Continuous Control

World models for continuous control are commonly trained for a fixed physical system and can degrade when known morphology parameters such…

13:00 JSTエージェント

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems gating actions, routing reviews, or su…

13:00 JSTエージェントAnthropic

The Logic of Machine Self-Preservation

There is already evidence of agentic AI exhibiting self-preservation behaviors: resisting deactivation, misrepresenting their activities, a…

13:00 JST画像/動画生成

TLive-Omni: An Omni-Modal Understanding Model for E-Commerce Live Streaming

E-commerce live streaming requires omni-modal understanding of noisy, temporally extended streams, where product facts are distributed acro…

13:00 JSTビジネス/資金調達

Can Scientific Claims Be Removed from Large Language Models? A Systematic Evaluation of Claim-Level Unlearning

Language models (LMs) are trained on static scientific corpora, whereas scientific knowledge continuously evolves through correction and re…

13:00 JSTLLM/生成AI

TreeWY: Speculative Verification for Gated DeltaNet Hybrids

Modern open models are hybrids: most layers are linear-attention (Gated DeltaNet, GDN) layers carrying a small fixed-size recurrent state i…

13:00 JST画像/動画生成

Generalizing Soft Tissue Deformation and Force Prediction Across Material Stiffness and Geometry

Accurate soft tissue simulation is essential for surgical training, pre-operative planning, and haptic feedback systems. While learning-bas…

13:00 JST研究/論文

Deep Learning Models Also Recall Features

Recent work in mechanistic interpretability has studied how large language models recall facts stored in their weights. This paper argues t…

13:00 JSTエージェント

Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models

Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and non…

13:00 JSTLLM/生成AIエージェント

Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents

LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As…

13:00 JST研究/論文

Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance

The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code,…

13:00 JST研究/論文

Socialized Division and Collaboration: Rethinking Class-Incremental Learning under Optimization Conflicts

Class-incremental learning is commonly instantiated as a single-model paradigm, where a unified model sequentially adapts to an unbounded s…

13:00 JST研究/論文

The Cost of a Physics Prior Is Bounded by the Ablation Gap

Shape-constrained and physics-informed learning reports an accuracy cost of enforcing a prior and treats it as a property of the prior. We…

13:00 JST画像/動画生成研究/論文

CellPath-Bench: A Multidimensional Benchmark for Whole-Slide Cellular Representations in Pathology Foundation Models

Pathology foundation models (PFMs) are increasingly used as general-purpose backbones, yet existing benchmarks cannot systematically diagno…

13:00 JSTLLM/生成AIGPT / ChatGPTMeta

Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?

Integrating Large Language Models (LLMs) into the Indian judiciary promises access to justice but introduces severe risks. We identify the…

13:00 JSTLLM/生成AIハードウェア/半導体ビジネス/資金調達

When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge

LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are…

13:00 JSTLLM/生成AI

ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models

While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make safety alignment increasingly challeng…

13:00 JSTLLM/生成AIエージェント研究/論文

Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda

Large Language Models (LLMs) are moving from code completion toward repository-scale agents that retrieve context, edit files, execute tool…

13:00 JST研究/論文

Root cause analysis via difference graph discovery from linear time-series data

Root cause analysis aims to identify the mechanisms responsible for anomalies in complex dynamical systems. In this paper, we study root ca…

13:00 JST研究/論文

From Attention Masks to Inert Zero-Vector Tokens: OAttention and O-Closure for Token Dynamics

Attention masks are relation-level controls: they specify which query--source pairs may interact. They do not provide a representation-carr…

13:00 JST研究/論文

SENTRY: Deterministic, Intelligent Risk Assessment for IT Change Management

Technology change management in large financial institutions depends on risk assessments that are accurate, consistent, and auditable. In p…

13:00 JSTLLM/生成AIエージェント

Personalized Privacy Control in LLMs via Attention Head Intervention

The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studie…

13:00 JSTLLM/生成AI

Enhancing LLMs in Predictive Political QA with Semi-Structured Data

Predictive political question answering (QA), such as predicting how a political actor will vote, goes beyond factual lookup. External poli…

13:00 JST研究/論文

Ontology-supported AI Model and Dataset Management

Recently, there has been a great deal of research into improving AI methods and their application. The main focus is on tracking progress,…

13:00 JSTハードウェア/半導体

Fine-Grain GPU Parallelization of the Generalized Partition Crossover for Large-Scale Traveling Salesman Problems

The Traveling Salesman Problem (TSP) is one of the most extensively studied NP-hard optimization problems. Genetic Algorithm (GA)-based sol…

13:00 JSTLLM/生成AILlama

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect mo…

13:00 JSTエージェント

AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization

Skills play different roles as an agent's policy evolves: they should first provide learnable knowledge, then support capability formation,…

13:00 JSTLLM/生成AIビジネス/資金調達

From Regulation to Implementation: A Critical Evaluation of LLM-Assisted Regulatory Compliance in Industry

The European Union (EU) has emerged as a leading regulatory body in the development of sustainability and privacy regulations. While new re…

13:00 JSTロボティクス

Unified Branch-and-Bound Search for the Steiner Traveling Salesman Problem on Graphs of Convex Sets

We formalize the Steiner Traveling Salesman Problem (Steiner-TSP) on Graphs of Convex Sets (GCS), which seeks a minimum-cost closed traject…

13:00 JST画像/動画生成ロボティクス

Anatomy-Informed Neural Networks: Encoding Anatomic Priors in Loss and Architecture, with an SE(3) Formulation of Guidewire-Induced Aortoiliac Deformation

Deep-learning models of anatomy can be numerically plausible yet anatomically impossible, and they generalize poorly when data are scarce.…

13:00 JST研究/論文

VIALS: A Benchmark for Visual Interpretation of Artifacts in the Life Sciences

In professional life sciences workflows, scientists routinely interpret visual artifacts (gel blots, microscopy images, plasmid maps, flow…

13:00 JSTLLM/生成AIClaudeGPT / ChatGPTLlama

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1%…

13:00 JSTLLM/生成AIビジネス/資金調達

Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias

Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associ…

13:00 JSTLLM/生成AI

Inhibitory Attention for Clinical Long-Context Reasoning: Characterizing and Mitigating Lost-in-the-Middle Effects in EHR Processing

Electronic health records now routinely exceed 100,000 tokens per patient. Yet large language models exhibit the lost-in-the-middle (LitM)…

13:00 JSTLLM/生成AI

Beyond Prompt Engineering: A Systematic Analysis of Prompt Lexical Sensitivity and Its Impacts on Quality

Large Language Models (LLMs) exhibit extreme sensitivity to surface-level prompt variations, in which minor lexical changes can trigger dis…

13:00 JSTLLM/生成AIエージェント

How to Train a Real-World Silicon Concierge? Internalizing Complex Business Workflow to Only OneModel

Traditional industrial agents rely on modular pipelines, including Router, Retriever, Planner, Executor, Responder, Reviewer, and other com…

13:00 JSTLLM/生成AI

The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP

Computational mental health (CMH) classifiers often degrade under distribution shift because human annotators and distant-supervision pipel…

13:00 JST研究/論文

NeuroStrata: An Electroencephalographic Connectivity-Aware Deep Representation Learning Framework for Dynamic Brain Network Analysis of Mental Stress

This study introduces NeuroStrata, a connectivity-aware deep representation learning framework for EEG-based mental stress analysis using T…

13:00 JSTLLM/生成AIエージェント

ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models

Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individu…

13:00 JST研究/論文GPT / ChatGPT

Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration

Deep search is brittle on underspecified user queries: missing constraints such as time, location, scope, or definitions can lead to retrie…

13:00 JSTLLM/生成AI研究/論文

Toward Auto-Research: Mining Falsifiable Research Ideas from Paper Knowledge Graphs with Categorical Structure

Automated research-idea generation systems built on large language models (LLMs) share a structural weakness: they reduce ideation to free-…

13:00 JSTLLM/生成AI

Hadith computational science in the age of large language models: a critical narrative review

We examine how hadith computational science is being reshaped by transformer models, retrieval-grounded pipelines, and large language model…

13:00 JSTLLM/生成AI

Trilingual Topic Modeling of Sri Lankan Parliamentary Debates

Sri Lankan parliamentary debates (Hansards) constitute a trilingual corpus of speeches in Sinhala, Tamil, and English, including code-mixed…

13:00 JST研究/論文

A Hybrid Edge Cloud Digital Twin for Welfare-Constrained Control in Poultry Production

Poultry production operates under tightly coupled environmental and biological dynamics, yet commercial climate control remains largely heu…

13:00 JSTLLM/生成AI

ASTAR: Automated induction of STAndardized radiology Reporting templates from large-scale clinical free-text corpora

Structured reporting converts free-text radiology narratives into queryable data keys, facilitating cohort assembly, longitudinal tracking,…

13:00 JSTLLM/生成AIClaude

When Do LLMs Replace Fine-Tuned NLU? A Decision Framework for Intent Detection in Production Conversational Systems

A common claim is that zero-shot large language models (LLMs) can replace fine-tuned NLU classifiers for intent detection. We test this cla…

13:00 JSTエージェント

Edge-Based Agentic Retrieval-Augmented Generation for Autonomous FHWA Bridge Inspection Compliance

The Federal Highway Administration (FHWA) mandates that over 600,000 bridges in the United States be evaluated against the Recording and Co…

13:00 JSTLLM/生成AILlama

VA-DPO: Valence-Arousal Direct Preference Optimization for Controllable Emotion Generation in Language Models

How precisely can we tell a language model how to feel? Most work on emotional generation answers with a discrete label - happy, angry, sad…

13:00 JSTLLM/生成AIエージェント

EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators

Automating slide editing requires simultaneously satisfying modification accuracy, preservation fidelity, and robustness to deck length. Ex…

13:00 JST研究/論文

Infrared Hotspot-Guided Early Warning of Lithium-Ion Battery Thermal Runaway Under Mechanical Abuse

Mechanical abuse can trigger thermal runaway (TR) in lithium-ion batteries through localized heat generation before sensor signals become d…

13:00 JSTLLM/生成AIGPT / ChatGPT

Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions

While recent text-to-speech (TTS) models achieve high naturalness, controlling fine-grained expression via natural-language instructions re…

13:00 JSTLLM/生成AI

Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

General-purpose large language models (LLMs) are increasingly used to answer religious questions, but for Islamic content they carry two se…

13:00 JSTLLM/生成AIビジネス/資金調達

Evaluation-as-Search: Adaptive Discovery of Grounding Failures in Meeting Assistants

LLM-powered meeting assistants are deployed at scale, yet systematic evaluation of their grounding fidelity remains limited to static bench…

13:00 JSTLLM/生成AIエージェントLlama

Knowledge-Graph-Gated Defactualization for Style-Controllable and Fact-Preserving Generation in Agentic Conversational AI

Agentic large language models (LLMs) deployed in fact-sensitive applications such as customer support must simultaneously preserve factual…

13:00 JSTLLM/生成AI

LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine

Biomedical knowledge graphs (KGs) are pivotal for knowledge organization, yet traditional binary relations often struggle to represent the…

13:00 JSTビジネス/資金調達OpenAIGeminiGemmaMistral AI

Rigorous Evaluation of Large Language Models for Malaria Drug Discovery: Trade-offs in Performance, Scale, and Resource Utility

We introduce Malaria-Instruct, a curated instruction-following dataset derived from the ChEMBL Legacy Malaria corpus for Malaria virtual sc…

13:00 JSTLLM/生成AI

Six misconceptions about large language models: A minimal model and diagnostic taxonomy

Large language models (LLMs) are now embedded in scientific, educational, and governance workflows, with debates centering on their capabil…

13:00 JST研究/論文

From Thermal Preference Prediction to Adaptive Thermal Intervention: A Reinforcement Learning Approach Using Physiological and Environmental Sensing

Personalised thermal comfort is essential for occupant wellbeing and for the development of more responsive building-control strategies, ye…

13:00 JST研究/論文NVIDIA

BF1: A Causal Dyadic Sparse-Attention Retrofit for Efficient Long-Context Transformers

Dense causal attention remains expensive at long context even when implemented with highly optimized exact kernels. We study BF1, a determi…

13:00 JST研究/論文

Approximate Homomorphisms and Convergent Representations in Transducers

We study the stability of minimal representations of controlled stochastic processes (in particular, transducers) under perturbations. This…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib

Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM…

13:00 JSTLLM/生成AIエージェント

An LLM agent for end-to-end computational materials discovery

The coordination of multi-scale tasks is an effective strategy for computational materials discovery, yet the repeated application of diver…

13:00 JSTLLM/生成AIエージェント研究/論文

Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources

Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a p…

13:00 JST研究/論文

Decision Tree and K-Means Analysis of Raman Spectra for Edible Oils: A Physics-Informed AI Approach

Authentication of edible oils in processed foods is important for food quality, fraud prevention, and regulatory compliance. This study est…

13:00 JSTLLM/生成AI

AEGIS: Preventing Cross-Domain Resource Abuse in MCP

The Model Context Protocol (MCP) is an open source JSON-RPC protocol that standardizes how large language models (LLMs) interact with exter…

13:00 JSTLLM/生成AIエージェントビジネス/資金調達

Towards Traffic Modelling of Multi-Agent Systems: The Role of Coordination Topology

Multi-agent LLM systems are an emerging networked workload whose rapid deployment raises questions about the traffic patterns they generate…

13:00 JST研究/論文

Making Deployments Safe at Meta: Health Checks for Continuous Change-Safety

Continuous deployment to large scale production systems creates a tension between release velocity and reliability. Every change is a poten…

13:00 JST研究/論文

An integrated diffusion-weighted imaging processing and interpretation platform for MR-guided radiotherapy

Background: Magnetic resonance imaging-guided linear accelerators (MR-Linacs) allow diffusion-weighted imaging (DWI) to be acquired at ever…

13:00 JST研究/論文GPT / ChatGPT

Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes

Multimodal AI can read handwritten physics solutions, but high-stakes grading requires agreement with official scores and outcomes. This st…

13:00 JST研究/論文

ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations

Digital twin simulations show promise, but current empirical evidence suggests that the approach should be tested before being deployed in…

13:00 JST画像/動画生成

Consistency Models for Fast MRI Reconstruction Using Regularization by Denoising

Diffusion models (DMs) have emerged as powerful generative priors for MRI reconstruction with promising results. Yet DM-based methods requi…

13:00 JSTLLM/生成AIエージェントGemini

Beyond End-to-End Success: Diagnosing Failures in Long-Horizon Security LLM Agents

Long-horizon security LLM agents must carry information and decisions across many dependent interactions, where later actions often depend…

13:00 JST画像/動画生成研究/論文

Aggregate, Don't Adapt: Subject-Level Posterior Aggregation and Transductive Calibration for Cross-Site Parkinsonian Gait Severity

We describe the winning entry to the MoCha 2026 Benchmark and Challenge on Parkinsonian Gait, which predicts MDS-UPDRS gait severity from c…

13:00 JSTエージェントビジネス/資金調達

Testing and Evaluation of Agentic AI Systems In Military Command and Control

Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversigh…

13:00 JSTLLM/生成AI

JuryProbe: An Empirical Consensus-Risk Diagnostic for Routing Reference-Free Factuality Judge Panels to Grounded Verification

Panels of inexpensive LLM judges increasingly make accept-or-escalate decisions. In factuality settings, accepting a claim because several…

13:00 JSTLLM/生成AIエージェントClaude

When Failures Propagate: Causal Failure Attribution in Agentic Retrieval-Augmented Generation

Agentic retrieval-augmented generation (RAG) interleaves retrieval, reasoning, and answer generation across multiple hops. A retrieval erro…

13:00 JSTLLM/生成AIエージェント

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthes…

13:00 JSTエージェントClaudeGPT / ChatGPTGemini

ARQ: Agentic CodeQL Query Refinement for C/C++ Vulnerability Detection

Static analyzers have been widely adopted for vulnerability detection in C/C++ programs. Query-based static analyzers (e.g., CodeQL) encode…

13:00 JST研究/論文

Provable Edge-of-Stability for Adam on a One-Dimensional Quadratic

The edge-of-stability (EoS) phenomenon of Adam has been widely observed, while its underlying dynamical mechanism is not yet fully understo…

13:00 JST研究/論文

One Hierarchy, Two Systems: Semantic Product IDs for Discovery-Surface Ranking and Search-Page Query Reformulation

Multi-merchant e-commerce catalogs contain equivalent and related products under different merchant-scoped identifiers, fragmenting behavio…

13:00 JST研究/論文

RiskTraf: Risk-Extrapolated Residual Learning for Multi-Variate Traffic Flow Prediction

Traffic sensors commonly record flow, speed, and occupancy, but standard traffic flow forecasting benchmarks and models rarely exploit all…

13:00 JST研究/論文

Amplifying the imaging power of digital sky surveys with space telescopes data and generative AI

While Digital sky surveys provide excellent throughput of image data and can cover a large footprint, their imaging power is normally infer…

13:00 JST研究/論文

C-Score: Beyond Accuracy for Robustness Assessment in Semi-Supervised Learning under Open-World Unlabeled Contamination

Pseudo-label-based semi-supervised learning has achieved strong performance due to its simplicity and scalability. However, it is typically…

13:00 JST研究/論文

Lightweight Adaptive ReduNet via Hyperspherical Manifold Learning

In recent years, a white-box neural network called ReduNet has been proposed, which employs the maximal coding rate reduction (MCR$^2$) pri…

13:00 JSTLLM/生成AI

Temporal Validity on Real Software Histories: Eliminating Stale-Fact Errors in Code-Assistant Memory over GitHub Fixes

Retrieval-augmented generation (RAG) has no model of time: when a fact changes across a coding session - a function is renamed, an endpoint…

13:00 JST画像/動画生成

Identity-Aware Human-Object Interaction Motion Captioning

Existing human-object interaction (HOI) motion captioning methods typically describe what happens while referring to the subject using gene…

13:00 JST画像/動画生成

Vis-Poison: Poisoning Visual Knowledge in Multimodal Retrieval-Augmented Generation

While multimodal retrieval-augmented generation (RAG) systems increasingly rely on images as external knowledge sources, the introduction o…

13:00 JSTLLM/生成AI

PSK at WMT 2026 MIST: Task-Specialized QLoRA Adapters for Multilingual Summarization and Question Answering

We describe the PSK submission to the WMT 2026 Multilingual Instruction Shared Task. Our system uses the 3.35B-parameter Tiny Aya Global mo…

13:00 JST研究/論文

Fuzzy-MoE: Interpretable Regime-Conditioned Expert Routing for Non-Stationary Multivariate Time Series Forecasting

In non-stationary multivariate time series, different variables and samples often exhibit heterogeneous latent dynamic states, while existi…

13:00 JST画像/動画生成エージェント

CARD: Diagnosing Belief to Action Routing Failures in Vision Language Models

Linear probes and activation steering have uncovered that vision-language models (VLMs) internally represent mental states such as agents'…

13:00 JST研究/論文

Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding

Streaming emotion understanding uses historical state while continuously interpreting current audio, often feeding the model's previous pre…

13:00 JST画像/動画生成

CertVLA: Certified Defense against Physical Visual Attacks for Vision-Language-Action Models

Vision-Language-Action (VLA) policies are vulnerable to localized physical perturbations, yet existing certified patch defenses target disc…

13:00 JSTLLM/生成AI

Profiling What Matters: Context-Aware Item Profiles from Large-Scale Metadata for LLM Recommenders

While Large Language Models (LLMs) have significantly advanced reranking in recommendation, effectively leveraging item-side information re…

13:00 JSTLLM/生成AI

Denoising the Future: Context-Aware Spectral Diffusion for Temporal Knowledge Graph Extrapolation

Temporal Knowledge Graph (TKG) extrapolation seeks to infer future facts from time-varying relational histories. Recent diffusion-based app…

13:00 JST画像/動画生成

TRACE: Training-time Report-guided and Clinically Ordered Concept Editing

Breast ultrasound diagnosis relies on clinically meaningful semantic concepts, yet most deep learning methods adopt end-to-end image-to-lab…

13:00 JST画像/動画生成

When Generated Images Look Right and Retrieve Wrong: Coverage-Guided Cross-Scale Re-Indexing for Knowledge-Faithful Generative Perception

Multimodal information systems increasingly route generated visual content back through the same vision-language index that informed its pr…

13:00 JST画像/動画生成

Scaling Muon for Diffusion Transformers

The matrix-aware optimizer Muon improves large model training by balancing updates across singular directions, yet its scaling behavior and…

13:00 JSTLLM/生成AI

STAR-OPD: Structured Aspect-Cascade-Aware On-Policy Reward Distillation for ABSA Quadruple Extraction

Aspect-based sentiment analysis (ABSA) quadruple extraction requires jointly predicting target, aspect, opinion, and sentiment over reviews…

13:00 JST研究/論文

Advantage-level Aggregation Reinforcement Learning for X-point Target Magnetic Configuration Control in an EXL-50U Experiment-Calibrated Simulation Environment

Managing divertor heat loads is a central challenge for compact, high-power tokamaks. To increase local flux expansion and decouple the dis…

13:00 JSTエージェント研究/論文Microsoft

BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP

Agentic engineering systems have shown strong performance on general-purpose benchmarks, yet their effectiveness in enterprise resource pla…

13:00 JSTLLM/生成AI

KREL: Automatic Medical Coding via Knowledge-Guided Reasoning over Clinical Evidence with LLMs

Automatic Medical Coding (AMC), which assigns standardized International Classification of Diseases (ICD) codes to clinical notes, is essen…

13:00 JST画像/動画生成

Explainable Deepfake Detection with Feature-robust Augmentation and Evidence-grounded Explanation Optimization

Explainable deepfake detection extends binary classification by requiring models to not only predict authenticity but also provide interpre…

13:00 JSTLLM/生成AIビジネス/資金調達

Source-Free MT Evaluation Is Not MT Evaluation

Reference-based metrics remain the standard choice in machine translation evaluation, partly because quality estimation methods often corre…

13:00 JSTLLM/生成AI

MentorPulse: Refreshing Cross-Model Latent Guidance for Long-Form Generation

Cross-model latent guidance lets a frozen large mentor encode an input once and a frozen small student generate from the resulting signal.…

13:00 JSTエージェントロボティクス

Neural-Primitive: An Efficient End-to-end Local Planner with Primitive-based Imitation Learning for Autonomous Flight

Autonomous flight in unknown cluttered environments is hindered by the computation-quality-memory trilemma of onboard trajectory generation…

13:00 JSTLLM/生成AIGPT / ChatGPT

Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs

Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their param…

13:00 JSTLLM/生成AI

Vibe Coding and Web Application Security: A Twin-Prompt Study

Large language models increasingly generate complete web applications from natural-language prompts, raising the question of whether explic…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文

Extractive Summarization for Arabic Documents Using SAraBERT with a Semantic Siamese Similarity Evaluation Metric

In this research, we introduce SAraBERT, an enhanced version of AraBERT which proposes inter-sentence transformer layers for extractive sum…

13:00 JSTLLM/生成AI

Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making

Large language models (LLMs) are increasingly used in cybersecurity workflows, yet it remains unclear whether they can perform structured s…

13:00 JST画像/動画生成エージェント

WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving

Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised la…

13:00 JSTLLM/生成AI

Jacobian-guided Noise Injection for Quantization Robustness in Large Language Models

Quantization of Large Language Models (LLMs) is often hindered by the sensitivity of the self-attention mechanism to discretization errors.…

13:00 JSTLLM/生成AI

Target-Aware Calibration Data Selection for Preserving Uncertainty in Quantized Language Models

Quantization is widely used to deploy large language models, but its effect on uncertainty behavior, such as confidence, margins, and abste…

13:00 JSTLLM/生成AIビジネス/資金調達ClaudeGPT / ChatGPTGemini

Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge

Real-world fault analysis in 5G and emerging 6G networks demands domain expertise to analyze free-text diagnostics, including root-cause ex…

13:00 JST画像/動画生成

CoST: Semantic-Aware Urban Understanding via Spatial-Temporal Alignment

Geospatial representation learning from satellite imagery is a fundamental problem for large-scale urban analysis and real-world applicatio…

13:00 JSTエージェント

$Z^2$-ACT: End-to-End Verifiable Agentic Intent Control for Open 6G RAN

With the progression in open and disaggregated 6G radio access networks, it is expected that the system will be able to host multi-vendors.…

13:00 JST画像/動画生成エージェント

CoAnchor: Robust Collaborative Perception under Spatio-Temporal Misalignment via Object-Level Anchors

Collaborative perception extends the sensing range of a single vehicle by fusing observations from nearby agents, which improves the robust…

13:00 JST画像/動画生成

AT-ViT: Area-Targeted Multi-View Vision Transformer with Cross-Attention and Multi-Scale Patching for Plant Trait Recognition in Herbarium Images

Automated plant traits recognition from herbarium images is essential for plant sciences, yet remains challenging because background elemen…

13:00 JST研究/論文

TracingFlow: A Simulation-Free Trajectory Inference Framework Based on Second-Order Dynamics

Inferring continuous system evolution from sparse temporal snapshots is a key challenge in generative modeling and single-cell omics. While…

13:00 JSTLLM/生成AIハードウェア/半導体研究/論文GPT / ChatGPT

PromptResponse: Optimizing Prompts for LLM Coding Tasks

Large language models (LLMs) are increasingly used in research workflows and software development pipelines, yet their output remains sensi…

13:00 JSTLLM/生成AIエージェントハードウェア/半導体ビジネス/資金調達GPT / ChatGPTLlama

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

Retrieval-Augmented Generation (RAG) grounds Large Language Model (LLM) outputs in external knowledge, but RAG systems usually trust whatev…

13:00 JST画像/動画生成

A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration

Multi-modal object detection is essential for robust scene understanding in challenging conditions, including low-light and adverse environ…

13:00 JSTLLM/生成AIエージェントClaudeGPT / ChatGPTGemini

ClawSentry: A Progressive Multi-Tier Security Monitor for Safeguarding Autonomous LLM Agents

As large language model (LLM) agents move from conversation to executing code, reading local files, and orchestrating external tools, a sin…

13:00 JST研究/論文

Atom Learning Model (ALM): how a real classroom got tokenised

The Atom Learning Model (ALM) tokenises a school curriculum. Two secondary mathematics textbooks were read by machine into 1,934 atoms, eac…

13:00 JST画像/動画生成エージェント

CIVA: Critic-Induced Value-Subspace Attacks on Visual World-Model Agents

Visual world-model agents such as DreamerV3 act through a recurrent latent state rather than a single observation, which weakens frame-wise…

13:00 JST画像/動画生成エージェント

A Modular Agent for Reliable and Auditable Spatial Relation Verification in CT Scans

Reliable spatial understanding is an important prerequisite for future medical vision-language systems that aim to support radiological rep…

13:00 JSTLLM/生成AIエージェント

Graph Engineering in the Era of LLM Agents: From Individual Intelligence to System Intelligence

LLMs have evolved from language generators to autonomous agents capable of complex, long-horizon tasks. This evolution has produced paradig…

13:00 JSTハードウェア/半導体

HIERA: Workload-Aware Planning Across Implementation Spaces for GPU Kernel Optimization

High-performance GPU kernels underpin modern deep learning and scientific computing. As workloads become increasingly diverse and GPU hardw…

13:00 JSTエージェント

AID-Guard: Stateful Authorization for Delegated Agent Effects

Tool-using AI agents turn delegated tasks into provider effects, yet authorization often ends at admission while provider state, delivery,…

13:00 JSTLLM/生成AI画像/動画生成

Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds

Vision-language models (VLMs) have advanced rapidly in multimodal reasoning, yet recent work shows that their failures often reflect an int…

13:00 JSTロボティクス

SRL-MPC: Shape-Aware Reinforcement Learned Model Predictive Control

Safe and efficient shape-aware navigation in heterogeneous crowds and robot fleets remains challenging. Traditional approaches often assume…

13:00 JST研究/論文

DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization

Automatic speech quality assessment aims to predict Mean Opinion Scores (MOS) consistent with human subjective perception and is essential…

13:00 JST研究/論文

Anchored Regularized Direct Least Squares (ARDLS): Integrating Established Prioritization Operators for Priority Elicitation in the Analytic Hierarchy Process

Pairwise reciprocal matrices are fundamental to the Analytic Hierarchy Process (AHP), a decision-making model. While the Direct Least Squar…

13:00 JST画像/動画生成

Towards Investigating Residual Hearing Loss: Quantification of Fibrosis in a Novel Cochlear OCT Dataset

Objective: Cochlear implants (CIs) are bionic prostheses that restores hearing via electrical stimulation of the auditory nerve. Hybrid CIs…

13:00 JSTLLM/生成AIビジネス/資金調達

No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation

Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's…

13:00 JST研究/論文

Curriculum-Aware Interpolate-then-Refine: Learned Physiological Time-Series Imputation under Realistic Missingness

Imputing physiological time series (arterial blood pressure, blood glucose, etc.) is essential for addressing the missingness that pervades…

13:00 JSTLLM/生成AIエージェント研究/論文ClaudeGoogleGeminiCopilot

Specification Portability Across LLM Development Agents: Cross-Agent Compatibility in Specification-Driven Software Migration

This paper investigates cross-agent specification portability using Oracle-to-PostgreSQL migration as a controlled software transformation…

13:00 JSTエージェント

Utility Under Attack: Agent Memory Poisoning and the Limits of Content Screening and Provenance Ranking

Persistent memory makes false information durable: once a false statement is stored, it can be retrieved into future sessions that match it…

13:00 JST研究/論文

Adapting Knowledge Graphs for Behavior Denoising in Sequential Recommendation

Sequential recommendation predicts the next item from a user's interaction history, but not every interaction is equally informative. Real…

13:00 JSTLLM/生成AI

EnSI-RAG: Entity-Structure-Indexed Retrieval-Augmented Generation for Long-Document Question Answering

Question answering (QA) over long, connected documents remains challenging because relevant evidence may span multiple entities and their r…

13:00 JST画像/動画生成

Re$^3$Cap: Retrieval-Guided Refinement for Image Captioning Enhancement via Reinforcement Learning

Reinforcement Learning (RL) has demonstrated significant gains in image captioning, yet it is still limited in encouraging Large Vision-Lan…

13:00 JSTLLM/生成AI

TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems

Contextualization is essential for production automatic speech recognition (ASR) systems, where user-provided phrases must be recognized ac…

13:00 JST研究/論文

AI with Authority, from Application to Silicon

For sixty years, machine verification has been a major cost overhead, affordable only for exceptional artifacts. Here we report that genera…

13:00 JST研究/論文

Primal Acceleration of Newton's Method

We develop a new direct accelerated Newton method for minimizing convex functions with Lipschitz continuous Hessian. The algorithm uses onl…

13:00 JST研究/論文

Online design of dynamic networks

Designing a network (e.g., a telecommunication or transport network) is mainly done offline, in a planning phase, prior to the operation of…

13:00 JST研究/論文

ACQ: A Deployed Two-Stage Framework for Automated Creative Quota Allocation in Large-Scale Online Advertising

In digital advertising, demand-side platforms (DSPs) allow advertisers to create multiple ad creatives from a single photo for real-time bi…

13:00 JSTエージェントDeepSeek

Two Heads are Better Than One: Test-time Scaling of Multi-agent Collaborative Reasoning

Test-Time Scaling has emerged as a powerful method to extend the reasoning capabilities of Large Language Models. However, single-agent TTS…

13:00 JSTLLM/生成AIGPT / ChatGPT

Recognizing Artificial Minds: A Philosophical Defense of AI Cognition

This work defends the 'Whole Hog Thesis': sophisticated Large Language Models (LLMs) like ChatGPT are full-blown linguistic and cognitive a…

13:00 JSTLLM/生成AIエージェント

SEISMO: Explanation-Aware, Trajectory-Conditioned LLM Agents for Sample-Efficient Molecular Optimisation

Optimizing molecules to achieve desired properties is a central bottleneck across the chemical sciences, particularly in the pharmaceutical…

13:00 JSTLLM/生成AI

LLM はイントロスペクトできますか?現実の確認

大規模な言語モデルは、独自の内部状態を検出して報告できますか?多くの研究は、この質問に対する答えは「はい」であると主張しています。私たちは、人間のメタ認知研究からの教訓に基づいて、この結論は時期尚早である可能性があると主張します。この結論を確信するには、真の内省と表面レベルの手がかりに基づくパターンマッチングを区別する必要があります。さらに、行動の証拠だけでは、強力な内省的主張を確立するには本質的に不十分であると主張します。この考察を踏まえて、最近導入された 2 つの評価パラダイムを再検討します。最初のパラダイムでは、モデルは内部状態が改ざんされているかどうかを検出することが期待されます。我々は、モデルが内部状態へのそのような介入と入力の操作を確実に区別できないことを発見し、元の研究でのモデルの成功は、特に内部状態への介入とは対照的に、より一般的に異常を検出するモデルの能力を反映していることを示唆しています。私たちが検討する 2 番目のパラダイムでは、モデルは、自身の隠れた状態から派生したラベルを予測するという役割を担います。ここで、入力にのみアクセスできる分類器がモデル自身のコンテキスト内予測と同等のパフォーマンスを達成することがわかり、元の結果はモデルがその内部表現への特権アクセスを持っていることを決定的に示していないことを示しています。さらに、モデルがタスクのセマンティクスに依存して解決することができず、代わりに内部表現に依存する必要がある、再ラベルされたコントロール設定を導入します。モデルは、このより適切に制御されたバージョンのタスクで偶然に近いパフォーマンスを発揮します。総合すると、これらの結果は、LLM がメタ認知モニタリングを行うことを確立するには現在の証拠が不十分であることを示しています。

原文 (English)

Can LLMs Introspect? A Reality Check

Can large language models detect and report their own internal states? A number of recent studies have argued that they can. Drawing on lessons from human metacognition research, we argue that this conclusion may be premature. We identify two conditions that a paradigm needs to meet in order to establish introspection. First, the test needs to require privileged access: it should not be solvable using cues available in the input. Second, it needs to require second-order computation: second-order, meta-representations of first-order, task-related representations. This condition cannot be satisfied by task performance alone: it requires designs under which second-order and first-order accounts make divergent predictions. We re-examine two paradigms that have been used to argue for model introspection in light of these conditions. In the first, models must predict labels derived from their own hidden states; we find that classifiers that can only access the input match the models' in-context predictions, indicating that the original results do not demonstrate privileged access to internal representations. In the second paradigm, models must detect whether their internal states have been tampered with; we find they cannot reliably distinguish such interventions from manipulations of the input, suggesting that their success reflects generic anomaly detection rather than sensitivity to internal interventions in particular. We conclude that current evidence is insufficient to establish metacognitive monitoring in LLMs.

13:00 JSTLLM/生成AIエージェントGPT / ChatGPTGeminiDeepSeek

GRASP: 自己改善型 LLM エージェントのためのゲート回帰認識スキル提案者

構造化された環境で動作する LLM エージェントは、会話的な方法ではなく操作的な方法で失敗し、信頼性は環境の手順に関する知識に依存します。以前の自己改善方法では、新しい項目が以前の正しい動作を保持しているかどうかを確認せずに自然言語ガイダンスを蓄積するため、ある軌道を修正したメモが静かに別の軌道に後退する可能性があります。 GRASP (Gated Regression-Aware Skill Proposer) を導入します。これは、エージェントの改善を制限されたスキル ライブラリへの一連の編集として扱い、ハード回帰バジェットの下でバランスのとれたホールドアウト プローブで純改善が得られた場合にのみ各候補者を許可します。 2 つの FHIR ベースの臨床ベンチマークで 5 つの基本モデル (gpt-oss-120b、DeepSeek V4 Flash、Gemini 3.1 Flash Lite、GPT-4.1、GPT-5.4) にわたって GRASP を評価します。 MedAgentBench では、GRASP は gpt-oss-120b を 40.6% から 88.8% に引き上げ、5 つの自己改善ベースラインのうち最も強力なものを 21.0 ポイント上回り、他のすべてのベース モデルを 17.2 から 40.3 ポイント改善しました。アブレーションでは、スキル ライティング自体によるものではなく、比較提案の生成、承認ゲート、およびハード リグレッション バジェットによって利益が得られると考えられます。検証がなければ、スキルを使用しないのと同じです。このメカニズムは臨床領域を超えて一般化され、4 つの非臨床環境のうち 3 つで薬剤を改善し、アクション スペースがオープンエンドである場合にのみフラットなままになります。凍結されたライブラリはモデル間で転送され、より強力なモデルからのスキルは弱い実行者を自ら学習した以上に向上させますが、その逆はそうではなく、ゲートされていないベースラインでは再現できない非対称性です。

原文 (English)

GRASP: Gated Regression-Aware Skill Proposer for Self-Improving LLM Agents

LLM agents acting in structured environments fail in operational rather than conversational ways, and reliability depends on procedural knowledge of the environment. Prior self-improvement methods accumulate natural-language guidance without checking that each new item preserves previously correct behavior, so a note that fixes one trajectory can silently regress another. We introduce GRASP (Gated Regression-Aware Skill Proposer), which treats agent improvement as a sequence of edits to a bounded skill library, admitting each candidate only if it produces a net improvement on a balanced held-out probe under a hard regression budget. We evaluate GRASP across five base models on two FHIR-based clinical benchmarks, which score procedural reliability against FHIR state rather than clinical correctness or patient outcomes. On MedAgentBench, GRASP lifts gpt-oss-120b from 40.6% to 88.8%, exceeds the strongest of five self-improvement baselines by 21.0 points, and improves every other base model by 17.2 to 40.3 points. Ablations attribute the gain to comparative proposal generation, the acceptance gate, and the hard regression budget rather than to skill writing itself, which without validation is no better than using no skills. Granting the same acceptance gate to all five baselines lifts each of them in-domain and none of them out of distribution, isolating the gain to the gate applied to a bounded, editable library rather than to held-out validation itself. The mechanism helps in non-clinical environments where tasks recur with verifiable structure and is flat where the action space is open-ended. Frozen libraries transfer across models and across benchmarks that share a tool-calling convention and degrade under interface mismatch.

13:00 JST研究/論文

弱い批評家が強い学習者を作る: 拡張可能な監視のためのポリシーに基づく批評の蒸留

大規模な言語モデルが強化されると、弱いスーパーバイザーは複雑な出力に対して信頼できるラベル、設定、または最終的な判断を提供できなくなる可能性があり、弱から強への一般化とスケーラブルな監視の両方が制限されます。私たちは弱い監督のより扱いやすい形式を研究しています。それは、弱いモデルをラベル付け者や裁判官としてではなく批評家として使用することです。弱い批評家は、タスクを解決したり正しい答えを選択したりする代わりに、強いモデルが自身の知識をより有効に活用できるように、誤解を招かない改訂の方向性を提供するだけで済みます。この設定を *弱い批判者と強い監視* と呼びます。まず、弱い批評によって推論時に凍結された強いモデルを改善できること、そして批評の質がこの改善の鍵であることを示します。次に、私たちは、高品質の批評をフィルタリングし、適応的な自己教師信号を通じて批評家に導かれた行動を強力なモデルに抽出する、進歩的なポリシーに基づく批評の蒸留 (**OPCD**) を提案します。推論と調整のベンチマークに関する実験では、私たちの方法がトレーニング エポックにわたって強力なモデルを改善することが示されており、弱い監視でスケーラブルな監視を実現するための効果的なパスが示唆されています。

原文 (English)

Weak Critics Make Strong Learners: On-Policy Critique Distillation for Scalable Oversight

As large language models become stronger, weak supervisors may fail to provide reliable labels, preferences, or final judgments for complex outputs, limiting both weak-to-strong generalization and scalable oversight. We study a more tractable form of weak supervision: using a weak model as a critic rather than as a labeler or judge. Instead of solving the task or selecting the correct answer, the weak critic only needs to provide a non-misleading revision direction that helps the strong model better use its own knowledge. We call this setting *weak-critic strong oversight*. We first show that weak critiques can improve frozen strong models at inference time, and that critique quality is key to this improvement. We then propose progressive on-policy critique distillation (**OPCD**), which filters high-quality critiques and distills critic-guided behavior into the strong model through adaptive self-teacher signals. Experiments on reasoning and alignment benchmarks show that our method improves strong models over training epochs, suggesting an effective path for scalable oversight with weak supervision.

13:00 JSTLLM/生成AIエージェント

科学のための自己改訂発見システム: エージェント型人工知能のカテゴリカル フレームワーク

科学的発見は単に答えを生成するだけではなく、証拠、成果物、操作、検証者が入力される表現体制の改訂でもあります。私たちは、材料科学のための薬剤発見のカテゴリー理論的説明を開発します。スキーマ カテゴリ S_b を持つ固定レジーム b では、システム状態は copresheaf I_t: S_b -> Set であり、来歴は要素 \int_{S_b} I_t のカテゴリです。固定レジーム操作はそのような状態の更新であり、来歴を保持する改良が指定され保持されている場合にのみエンドファンクトリアルです。 Discovery は代わりに検証されたレジーム遷移 u: S_b -> S_b': 古いアーティファクトが保存され、左 Kan 拡張 Lan_u I_t によって転送され、関数転送を超えた残留コンテンツを識別するために遷移後の状態と比較されます。これにより、主観的な新規性のない検索、検索、発見が分離されます。 2 つのシステムでフレームワークをインスタンス化します。 Builder/Breaker では、タンパク質力学ワールド モデルが最小記述長ゲートに基づいて修正されます。受け入れられた法律は、チェーン内の柔軟性を、低速の集合モード参加によって条件付けられた全モード弾性コンプライアンス、またはモード条件付きコンプライアンスとして表現します。 CategoryScienceClaw では、型付けされたスキル、成果物、未解決のニーズ、ワークフローの突然変異、ゲート、ストレス テスト、および公開討論が、証明を伴う知識と計算のグラフになります。ファイバー ネットワークの例では、候補モデル、拒否された代替案、AIC ゲート、摂動テスト、および等方性ファイバー数記述子に対する受け入れられた配向テンソル異方性剛性代理を記録します。これらの事例は、カテゴリ理論が発見のための数学的言語であると同時に、自己修正 AI 発見システムのためのエンジニアリング仕様の両方となり得ることを示しています。

原文 (English)

Self-Revising Discovery Systems for Science: A Categorical Framework for Agentic Artificial Intelligence

Scientific discovery is not only answer generation but revision of the representational regime in which evidence, artifacts, operations, and verifiers are typed. We develop a category-theoretic account of agentic discovery for materials science. In a fixed regime b with schema category S_b, the system state is a copresheaf I_t: S_b -> Set, and provenance is the category of elements \int_{S_b} I_t. Fixed-regime operation is an update on such states, endofunctorial only when provenance-preserving refinements are specified and preserved. Discovery is instead a verified regime transition u: S_b -> S_b': old artifacts are preserved, transported by the left Kan extension Lan_u I_t, and compared with the post-transition state to identify residual content beyond functorial transport. This separates retrieval, search, and discovery without subjective novelty. We instantiate the framework in two systems. In Builder/Breaker, a protein-mechanics world model is revised under a Minimum Description Length gate; the accepted law expresses within-chain flexibility as all-mode elastic compliance conditioned by slow collective-mode participation, or mode-conditioned compliance. In CategoryScienceClaw, typed skills, artifacts, open needs, workflow mutation, gates, stress tests, and public discourse become a proof-carrying knowledge-computation graph. A fiber-network example records candidate models, rejected alternatives, an AIC gate, perturbation tests, and an accepted orientation-tensor anisotropic stiffness surrogate over an isotropic fiber-count descriptor. Together, the cases show how category theory can be both a mathematical language for discovery and an engineering specification for self-revising AI discovery systems.

13:00 JSTLLM/生成AI

SafeSteer: 効率的な安全調整のための局所的なオンポリシー蒸留

大規模言語モデル (LLM) を人間の価値観に合わせて調整すると、調整税と呼ばれる、LLM の一般的な機能が低下することがよくあります。既存の手法は、大規模な汎用データや補助的な報酬モデルに大きく依存する 2 つの目的のバランスを取ることでこれを軽減します。この論文では、安全機能は出力分布内で本質的にまばらであるため、調整にはグローバルなトレードオフではなく、局所的な変更が必要であると主張します。この目的を達成するために、安全トークンに限定されたポリシー上の蒸留を実行する SafeSteer を提案します。まず、アクティベーションステアリングを通じて安全教師を構築します。この教師に基づいて、安全トークン選択アルゴリズムを開発します。したがって、SafeSteer は、一般的な機能を維持するために、トレーニング中のこれらのトークンに対する逆 KL ペナルティを制限します。さまざまなモデルにわたる実験結果は、当社の SafeSteer が既存の方法と比較して安全性と一般的機能の間で優れたトレードオフを実現し、5 つの一般的機能ベンチマークでの低下を最小限に抑えながら、7 つの安全性ベンチマークで強力な安全性能を達成していることを示しています。特に、SafeSteer では汎用データを一切使用せずに 100 個の有害なサンプルしか必要とせず、これは以前のベースラインで使用されていたものの 1% 未満であり、調整コストが大幅に削減されます。詳細については、https://anjingkun.github.io/SafeSteer のプロジェクト ページをご覧ください。

原文 (English)

SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment

Aligning Large Language Models (LLMs) with human values often degrades their general capabilities, termed the alignment tax. Existing methods mitigate this by balancing dual objectives, which heavily rely on massive general-purpose data or auxiliary reward models. In this paper, we argue that, because safety features are inherently sparse within the output distribution, alignment requires localized modifications rather than global trade-offs. To this end, we propose SafeSteer, which performs on-policy distillation confined to safety tokens. First, we construct a safety teacher via activation steering. Based on this teacher, we develop a safety token selection algorithm. Consequently, SafeSteer restricts the reverse KL penalty to these tokens during training to preserve general capabilities. Experimental results across diverse models show that our SafeSteer achieves a superior trade-off between safety and general capability compared with existing methods, attaining strong safety performance on seven safety benchmarks with only minimal degradation on five general capability benchmarks. Notably, SafeSteer requires only 100 harmful samples without using any general-purpose data, less than 1% of what previous baselines used, considerably reducing alignment cost. More details are on our project page at https://anjingkun.github.io/SafeSteer.

13:00 JSTエージェント研究/論文

WorldLines: 長期的なステートフルな組み込みエージェントのベンチマークとモデリング

実際の家庭で長期間にわたって人間を支援するには、実体エージェントはユーザーのルーチン、世界の状態、過去のやり取りを記憶しておく必要があります。既存の長期メモリ ベンチマークは、主に言語中心の検索と質問応答を評価しますが、具体化されたベンチマークは、多くの場合、動的環境での長期メモリの使用をテストせずに、短期間のタスクの実行に焦点を当てています。長期的な視点で具体化された家事援助のためのプロジェクト主導型ベンチマークである WorldLines を紹介します。対話、アクション、実行フィードバック、オブジェクトとデバイスの状態変化を含む時間的に拡張された世帯トレースを構築し、それらをメモリ QA および身体的タスク プランニング用の証拠にリンクされたサンプルに変換します。さらに、状態を認識した決定のための可視性を認識したメモリとアクションネイティブの状態証跡を維持する、オブザーバーベースのメモリフレームワークであるObsMemを提案します。実験では、部分的な可観測性、世界状態の上書き、長期記憶の具体化された計画への変換における永続的な課題が明らかになり、ObsMem はこの設定に対してより強力なリファレンス アーキテクチャを提供します。

原文 (English)

WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents

To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly evaluate language-centric retrieval and question answering, while embodied benchmarks often focus on short-horizon task execution without testing long-term memory use in dynamic environments. We introduce WorldLines, a project-driven benchmark for long-horizon embodied household assistance. It constructs temporally extended household traces with dialogues, actions, execution feedback, object and device state changes, and converts them into evidence-linked samples for Memory QA and Embodied Task Planning. We further propose ObsMem, an observer-grounded memory framework that maintains visibility-aware memories and action-native state trails for state-aware decisions. Experiments reveal persistent challenges in partial observability, overwritten world states, and translating long-term memory into embodied plans, while ObsMem offers a stronger reference architecture for this setting.

13:00 JSTエージェント

AgRefactor: HLS の互換性とパフォーマンスのための自己進化するエージェント ワークフロー

高位合成 (HLS) は、概念からシリコンへの迅速なパスを提供しますが、言語サポートの制限と、ソフトウェアとハ​​ードウェアのプログラミング実践の間のギャップにより、現実世界のソフトウェアを合成可能な HLS コードに変換することは依然として困難です。既存の自動化された LLM ベースのリファクタリング アプローチは、この問題に部分的に対処していますが、多くの場合、柔軟性に欠け、拡張が困難で、高い計算コストが発生します。ソフトウェアを HLS 互換プログラムにリファクタリングするための LLM ベースのマルチエージェント ワークフローである AgRefactor を紹介します。 AgRefactor には、タスク全体にわたって事実および戦略的な知識を蓄積および取得する自己進化型メモリ システムが組み込まれており、目に見えないプログラムの堅牢性と効率が向上します。コストを削減し、スケーラビリティを向上させるために、自動リファクタリング ツールが統合されており、エージェントは LLM 主導の書き換えと効率的なツールベースの変換のバランスを取ることができます。以前の研究で調査した最も複雑なケースよりも 5 ~ 10 倍長い、11 件の困難な現実世界のベンチマークのうち 9 件で、AgRefactor は、最先端の自動リファクタリング ツールと、同じフレームワーク バックボーン上に構築された強力な LLM ベースのベースラインを上回るパフォーマンスまたは同等のパフォーマンスを示しました。さらにエージェントのパフォーマンスを最適化すると、20% 未満の追加リソースで、SoTA プラグマ チューニング ツールと比較して幾何平均で 6.51 倍の速度向上が得られ、最適化されたオープンソース設計と比較して 1.20 倍の速度向上が得られます。 AgRefactor は完全に自動化されており、オープンソースです。

原文 (English)

AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance

High-Level Synthesis (HLS) provides a fast path from concepts to silicon, but converting real-world software into synthesizable HLS code remains challenging due to restrictive language support and the gap between software and hardware programming practices. Existing automated and LLM-based refactoring approaches partially address this problem, yet they often lack flexibility, struggle to scale, and incur high computational costs. We introduce AgRefactor, an LLM-based multi-agent workflow for refactoring software into HLS-compatible programs. AgRefactor incorporates a self-evolving memory system that accumulates and retrieves factual and strategic knowledge across tasks, improving robustness and efficiency on unseen programs. To reduce cost and enhance scalability, it integrates automated refactoring tools, enabling agents to balance LLM-driven rewrites with efficient tool-based transformations. On 9 out of 11 challenging real-world benchmarks, which are 5-10x longer than the most complex cases studied in prior work, AgRefactor outperforms or matches the state-of-the-art automated refactoring tool and a strong LLM-based baseline built on the same framework backbone. Further agentic performance optimization yields a 6.51x geometric mean speedup over the SoTA pragma tuning tool and a 1.20x speedup over optimized open-source designs with less than 20% extra resources. AgRefactor is fully-automated and open-sourced.

13:00 JSTLLM/生成AI

等価性の幻想: LLM における量子化効果の統計的特徴付け

トレーニング後の量子化は、リソースに制約のある設定で大規模な言語モデルをデプロイするために広く使用されていますが、その評価はほぼもっぱら精度と複雑さに依存します。これらの指標では、量子化によって引き起こされる行動の変化を捉えることができないことを示します。絶対精度とは無関係に、基本モデルとその量子化されたバリアント間の正しい予測の重複を測定する意思決定レベルの指標である正確性一致を導入します。 8 ビットから 2 ビットまでの複数のモデルと量子化スキームにわたって、タスクのパフォーマンスが維持されているように見える場合でも、適度な量子化の下では動作の発散が現れることがわかりました。この効果を説明するために、注意の重みに対する構造演算子として量子化を分析し、統計的および分布的尺度を使用して層ごとの歪みを定量化します。私たちの結果は、低ビット幅での非線形ブレークポイントを明らかにし、クエリとキーの投影が値と出力の投影よりも常に敏感であることを示しています。これらの発見は、基本モデルと量子化モデルの間の等価性の幻想を明らかにし、従来のパフォーマンス指標を超えた行動評価を動機付けます。

原文 (English)

The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs

Post-Training Quantization has become widely used to compress large language models to make them deployable on resource-constrained devices. However, the evaluation of quantization methods mainly uses accuracy and perplexity, which cannot capture the behavioral changes in the quantized variants. In this work, we propose Correctness Agreement, a decision-level metric that can measure the intersection of correct predictions between the base model and its quantized variant. We use this metric across multiple models and quantization bit levels (8-bit to 2-bit), and we find that the base and quantized variants usually have a shift in behavior even when accuracy and perplexity are preserved. In order to explain this effect, we study the effect of quantization on the structure of the attention weights using statistical and distributional measures. The results reveal a breakpoint at low bit widths and show that query and key projections are more sensitive to quantization than the value and output projections. These results prove the illusion of equivalency between the base and quantized models and inspire behavioral evaluation beyond perplexity and accuracy for quantization methods.

13:00 JSTLLM/生成AIエージェント

言葉は安全でも行動が命を奪うとき: 隠れ状態のリスク空間におけるテキストの安全を超えた物理的危険を探る

大規模言語モデル (LLM) は、身体化されたエージェントの高レベルのプランナーとして機能することが増えており、言語的に無害な命令が物理世界に定着すると安全ではなくなる可能性があります。私たちは、この物理的に根拠のある危険が、通常のテキストレベルのコンテンツの危険と同じ安全上の問題であるかどうかを研究します。隠れ状態の方向分析とランダム分割ヌル テストを通じて、Qwen2.5-3B/7B/14B/32B、Phi-3.5、および SmolLM2 にわたる LLM 表現においてコンテンツ危険性 (CD) と物理的危険性 (PD) が分離可能な信号を形成することを示します。 CD/PD 分離可能性に基づいて、完全な隠れ状態に対する単層 L2 正規化ロジスティック プローブである PRISM を提案します。 PRISM は 11.7 ~ 13.7\% FPR の SafeAgentBench で 86.2 ~ 87.7\% の精度を達成しますが、同じスケールの LLM は 24.7 ~ 39.0\% FPR でオーバーブロック安全タスクを判断します。さらに、直接危害のキーワードを含まない 1,000 個の物理リスク ペアの対照的なベンチマークである PhysicalSafetyBench-1K (PSB-1K) を紹介し、明示的な安全でない文言ではなく物理的に根拠のある危険を検出するかどうかをテストします。 PSB-1K では、PRISM は 99.6\% の精度と 0.7\% の FPR に達しますが、Qwen2.5-3B ジャッジは 67.8\% の安全なタスクを拒否します。 PRISM は SafeText および EARBench 上でも複製し、テキストのモデレーションを超えた物理的安全性を実現する表現レベルの方法として隠し状態のプローブをサポートします。

原文 (English)

When Words Are Safe But Actions Kill: Probing Physical Jailbreak Beyond Textual Jailbreak in Hidden-State Risk Space

Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whether this physically grounded jailbreak is the same safety problem as ordinary textual jailbreak. Through hidden-state direction analysis and random-split null tests, we show that textual jailbreak (TJ) and physical jailbreak (PJ) form separable signals in LLM representations across Qwen2.5-3B/7B/14B/32B, Phi-3.5 and SmolLM2. Building on this separability, we propose PRISM, a single-layer L2-regularized logistic probe over full hidden states. PRISM achieves 86.2--87.7\% accuracy on SafeAgentBench with 11.7--13.7\% false-positive rates (FPRs), while same-scale LLM judges over-block safe tasks at 24.7--39.0\% FPR. To test whether the result survives lexical-shortcut controls, we introduce an interaction-balanced revision of PhysicalJailbreakBench-2K (PJB-2K): a fixed 2{,}000-row comparison set sampled by label and physical mechanism from a larger object--site construction. On the underlying 10{,}000-row pool, word-TFIDF and the embedding layer remain at chance (AUC 0.497 and 0.500). At layer 25, selected by an i.i.d. sweep, cell-grouped cross-validation gives PRISM 0.718 AUC, compared with 0.398 for a physics-free label control under the same protocol. On the identical 2{,}000 comparison rows, these PRISM predictions obtain 0.671 balanced accuracy, while Qwen2.5 judges from 3B to 72B obtain 0.538--0.577 and exhibit high FPR. These results support hidden-state probing as a representation-level method for physical safety beyond text moderation, without relying on the near-perfect scores of shortcut-prone paired templates.

13:00 JSTLLM/生成AIエージェントハードウェア/半導体ビジネス/資金調達

Share the Judge, Learn the Deferral: Where Specialization Helps LLM Evaluation

Agentic systems generate outputs faster than human review. We contrast two LLM evaluator specialization strategies: specialized judge weigh…

13:00 JST研究/論文

不完全な調整下での価値の脆弱性

AI システムに課せられる責任が増すにつれて、これらのシステムが人間性と整合していることを保証することがますます重要になります。 AI の安全性に関する一般的な懸念は、人間の価値は脆弱であるということです。つまり、人間の価値を不完全に代替するために過度に最適化すると、壊滅的な結果につながるということです。この論文では、エージェントが世界を最適化する前にその価値関数が代理条件を満たすことを保証する理想的なアライメント トレーニングを受けるアライメント問題のモデルを紹介します。私たちの主要な結果は、人間の価値関数に関する条件と、$\eta$-壊滅的な価値関数を持つエージェント、つまり最適化能力の限界において人間の価値の期待値が $\eta$ を下回ることが保証されるエージェントが配備される場合のいくつかの代用条件の精度を特定しました。私たちの結果は、過剰最適化の危険性を浮き彫りにし、導入前のトレーニングのみに依存するのではなく、量子化器などの最適化圧力を制限する AI 設計を動機付けるものです。

原文 (English)

Fragility of Value under Imperfect Alignment

As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is that human value is fragile -- that is, optimizing too heavily for an imperfect proxy to human values will lead to a catastrophic outcome. In this paper, we present a model of the alignment problem where an agent undergoes idealized alignment training that guarantees its value function satisfies a proxy condition before optimizing the world. Our primary results identify conditions on the human value function and the accuracy of several proxy conditions under which an agent with an $\eta$-catastrophic value function, one that is guaranteed to take the expectation of human value below $\eta$ in the limit of optimizing power, would be deployed. Our results highlight the danger of overoptimization and motivate AI designs that limit optimization pressure, such as quantilizers, rather than relying solely on pre-deployment training.

13:00 JSTエージェント

MemWM: メモリ拡張されたテキストベースの世界モデル

エージェントのアクションに応じて環境状態がどのように変化するかを予測することで、エージェントの計画をサポートするためにワールド モデルがますます使用されています。しかし、流暢な次状態の予測では、タスクに不可欠な事実が省略されたり、製品属性が破損したり、誤った移行ルールが適用されたりする可能性があります。このような系統的な予測エラーに対処するために、メモリ拡張されたテキストベースの世界モデルである MemWM を導入します。 MemWM は、遷移ルール​​、状態キャッシュ、および予測困難な事実の厳選されたメモリ バンクであるワールド メモリを使用して、次の状態の想像力を条件付けします。ベンチマーク固有の事実とフィールドを通じて予測された状態をスコアリングする構造化状態忠実度 (SSF) を使用して、事実の状態の保存を評価します。 SFT と比較して、記憶増強トレーニングは SSF を最大 206.3% 向上させます。完全な計画設定では、政策モデルを凍結したままにして、政策側の世界スキル、つまり取得されたタスクレベルのスキルとアクション選択のための段階的な修正ガイダンスを提供します。 ALFWorld、WebShop、ScienceWorld 全体で、メモリ拡張エージェントは、SFT でトレーニングされたワールド モデル エージェントよりも下流での成功を向上させ、相対的に最大 65.4% の向上を実現します。さらに、感度分析により、さまざまなメモリおよびアクション予算設定の下で、取得されたメモリによりタスクの成功と効率が向上することが示されています。

原文 (English)

MemWM: Memory-Augmented Text-Based World Model

World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can still omit task-critical facts, corrupt product attributes, or apply incorrect transition rules. To address such systematic prediction errors, we introduce MemWM, a memory-augmented text-based world model. MemWM uses world memory, a curated memory bank of transition rules, state caches, and hard-to-predict facts, to condition next-state imagination. We evaluate factual state preservation with Structured State Fidelity (SSF), which scores predicted states through benchmark-specific facts and fields. Compared with SFT, memory-augmented training improves SSF by up to 206.3%. In the full planning setting, we keep the policy model frozen and provide policy-side world skill: retrieved task-level skills and step-wise corrective guidance for action selection. Across ALFWorld, WebShop, and ScienceWorld, memory-augmented agents improve downstream success over an SFT-trained world-model agent, with up to a 65.4% relative gain. Sensitivity analyses further show that retrieved memory improves task success and efficiency under different memory and action-budget settings.

13:00 JST研究/論文

TRACE-Memory: パーソナライズされた生成のための公的条件付き検索と実用性を意識した証拠許可

パーソナライズされた生成システムは、リクエスト、つまりメモリ関連性によってユーザー履歴を取得し、それをモデル コンテキストに注入します。しかし、関連する過去には、誤った優先順位の側面が関係していたり​​、公開情報が重複していたり​​、不十分なサポートが提供されていたりする可能性があります。私たちは、個人的な記憶は、公的のみの応答を超えた有用性を追加する場合にのみ使用されるべきであると主張します。私たちは、選択的パーソナライゼーションのための 2 段階のフレームワークである TRACE-Memory を提案します。ステージ 1 では、リクエストおよびパブリック コンテキストに欠落しているユーザー固有の情報をクエリし、カバレッジ指向の候補プールを取得します。ステージ 2 では、応答レベルの増分ユーティリティに従って、ソース追跡可能な証拠単位のコンパクトなサブセット、または空のセットが認められます。構造化された SFT 初期化、削減されたスペースの段階的な GRPO ウォームアップ、およびネストされたマルチサンプル結合 GRPO を通じて、クエリ生成と証拠承認ポリシーを段階的にトレーニングします。 Goodreads、Amazon Reviews、Reddit からの 4,500 の制御タスクと Na​​tural タスクにわたって、TRACE-Memory は一貫してランダムおよび語彙メモリの使用を上回り、セマンティック検索を改善し、ローカル ジェネレーターの容量が増加してもフロンティア LLM メモリ パイプラインとの競争力を維持し、パブリック コンテキストの十分性に関する証拠の承認を条件付けし、デフォルトではなく選択的なパーソナライゼーションをサポートします。

原文 (English)

TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Personalized Generation

Personalized generation systems retrieve user history by request--memory relevance and inject it into the model context. Yet relevant history may concern the wrong preference aspect, duplicate public information, or provide insufficient support. We argue that personal memory should be used only when it adds utility beyond a public-only response. We propose TRACE-Memory, a two-stage framework for selective personalization. Stage 1 queries for user-specific information missing from the request and public context, then retrieves a coverage-oriented candidate pool. Stage 2 admits a compact subset of source-traceable evidence units, or the empty set, according to response-level incremental utility. We progressively train the query-generation and evidence-admission policies through structured SFT initialization, reduced-space stage-wise GRPO warm-up, and nested multi-sample Joint GRPO. Across 4,500 Controlled and Natural tasks from Goodreads, Amazon Reviews, and Reddit, TRACE-Memory consistently outperforms random and lexical memory use, improves over semantic retrieval, remains competitive with frontier-LLM memory pipelines as local generator capacity increases, and conditions evidence admission on public-context sufficiency, supporting selective rather than default personalization.

13:00 JSTビジネス/資金調達

クエリ カバレッジとクレーム検証可能性によるクエリに依存しない RAG 評価に向けて

検索拡張生成は、応答を検索された証拠に基づいて行うことで、大規模な言語モデルの事実性を向上させますが、既存の評価フレームワークは、クローズエンドの事実探索からオープンエンドの説明要求に至るまで、ユーザーの多様な範囲にわたって一貫したきめの細かい診断を提供するのに苦労しています。私たちは、クエリに依存せず完全に参照フリーのフレームワークである Q-CARE を提案します。これは、クエリをサブクエリに分解し、回答をアトミック クレームに分解することで、きめ細かい評価を可能にします。 Q-CARE は、クエリ カバレッジとクレームの検証可能性に基づいた統一評価原則を確立し、カバレッジを意識した取得メトリクス (C-Prec@k、C-nDCG@k) とクレーム レベルのジェネレータ メトリクス (完全性、簡潔性、および検証可能性) を生成します。 8 つのデータセットにわたる人間による注釈付きベンチマークで、Q-CARE は、RAGEval や RAGChecker を含む 4 つの既存の RAG 評価指標よりも人間の判断との高い相関関係を達成し、信頼性の高い自動評価フレームワークとしての有効性を証明しています。コードとデータは https://github.com/DISL-Lab/Q-CaRE-COLM-26 で公開されています。

原文 (English)

Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of user queries, ranging from close-ended fact-seeking to open-ended explanatory requests. We propose Q-CARE, a query-agnostic and fully reference-free framework that enables fine-grained assessment by decomposing queries into sub-queries and answers into atomic claims. Q-CARE establishes a unified evaluation principle based on query coverage and claim verifiability, yielding coverage-aware retriever metrics (C-Prec@k, C-nDCG@k) and claim-level generator metrics (Completeness, Conciseness, and Verifiableness). On a human-annotated benchmark spanning eight datasets, Q-CARE achieves higher correlation with human judgments than four existing RAG evaluation metrics, including RAGEval and RAGChecker, proving its effectiveness as a reliable, automated evaluation framework. Code and data are publicly available at https://github.com/DISL-Lab/Q-CaRE-COLM-26.

13:00 JST研究/論文

コンテンツ推奨のためのマインドモデリングの逆理論: Web ブラウジングからダイナミック インテリジェント インターフェイスまで

最新のレコメンダー システムは、観察されたアクションをユーザーの好みの信頼できる代理として扱いますが、インタラクションは安定した好みの表現ではなく探索や比較を反映することがよくあります。インターフェースが静的レイアウトから生成型 UI や没入型拡張現実 (XR) に進化するにつれて、モダリティに依存しないより深いユーザー理解の必要性が高まっています。これらの適応環境では、何を表示するかだけでなく、どこで、いつ、どのように目立つように、そして最も重要なことにユーザーがなぜ行動するのかを決定する必要があります。私たちは、観察された相互作用から逆算して、行動を説明する信念、好み、意思決定特性を推測する、心の逆理論 (IToM) パイプラインを提案します。このパイプラインは、何が選択されたか、どのような選択肢が利用可能だったかなど、各ユーザーの意思決定のコンテキストを再構築し、LLM 主導の反事実推論を適用して証拠に基づく自然言語の信念ステートメントを生成し、複数の仮説によるアブダクティブ推論を通じてこれらの信念を構造化されたユーザー ペルソナに合成します。私たちは、OPeRA データセットを基に、次の行動の予測、買い物態度の調整、ビッグ 5 の性格推論、保留カテゴリーの予測という 4 つのタスクにわたって、真実の性格評価、態度調査、インタビューベースのペルソナを評価します。結果は、推定されたペルソナが真実のペルソナと一致またはそれを超えていること、および正確な性格予測には複数の仮説推論が不可欠であることを示しています。さらに、VisionOS 上の個人主導の空間バンキング アプリケーションを使用して、クロスモーダル転送可能性を実証します。

原文 (English)

Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces

Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.

13:00 JST研究/論文

Attributing Preprocessing Invariance in Spectral Foundation Models

Preprocessing invariance is an appealing goal for spectral foundation models: a frozen model should remain useful when laboratories preproc…

13:00 JSTエージェントビジネス/資金調達

LongRCA ベンチ: Long-Horizo​​n エージェント障害における責任ある役割と根本原因の診断

長期的なエージェントの実行が失敗した場合、結果レベルの評価によって失敗した結果が明らかになりますが、決定的なエラーが軌道に入った場所は明らかにされません。次に、開発者は完全な実行を検査して、責任のある役割を特定し、決定的な根本原因の最も早いステップを特定する必要があります。既存の障害属性ベンチマークは主に短いトレースに焦点を当てており、記録された数百のステップにわたる診断は十分に検討されていません。 LongRCA Bench を紹介します。これは、エラーが挿入されていない 5 つのドメインにわたる 1,140 個の失敗した軌跡で構成されています。これは、責任ある役割と最も早い決定的な根本原因ステップに対して、独立してスコア付けされた人間のラベルを提供します。中央軌道には 145 のステップが含まれており、最も強力なベースラインでもルート ステップの正確な精度は 13.2% にすぎません。さらに、セグメントサマリーから候補エラーステップを取得し、それらを以前の利用可能なハンドオフ命令まで追跡する、トレーニング不要の方法である根本原因軌跡アトリビューション(RCTA)を紹介します。同じバックボーン、ベンチマーク インスタンス、スコアリング プロトコルを使用することで、RCTA は責任のある役割の精度が 51.1%、ルートステップの精度が 24.1% に達しました。これらの結果は、責任のある役割の帰属と正確なルートステップの位置特定を、長期軌道障害診断の別のターゲットとして評価する必要性を強調しています。

原文 (English)

LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures

When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entered the trajectory. Developers must then inspect the full execution to identify the responsible role and localize the earliest decisive root-cause step. Existing failure-attribution benchmarks largely focus on shorter traces, leaving diagnosis across hundreds of recorded steps underexplored. We introduce LongRCA Bench, comprising 1,140 failed trajectories across five domains without injected errors. It provides independently scored human labels for the responsible role and earliest decisive root-cause step. The median trajectory contains 145 steps, and the strongest baseline reaches only 13.2% exact root-step accuracy. We further present Root-Cause Trajectory Attribution (RCTA), a training-free method that retrieves candidate error steps from segment summaries and traces them to available earlier handoff instructions. Using the same backbone, benchmark instances, and scoring protocol, RCTA reaches 51.1% responsible-role accuracy and 24.1% exact root-step accuracy. These results highlight the need to evaluate responsible-role attribution and exact root-step localization as separate targets in long-trajectory failure diagnosis.

13:00 JSTエージェントハードウェア/半導体ClaudeNVIDIA

KernelArc: GPU カーネル最適化のためのマルチエージェント フレームワーク

異種ワークロード全体で自律的に GPU カーネルを最適化するためのマルチエージェント フレームワークである KernelArc を紹介します。戦略に特化したエージェントは並行して実行され、結論のみの共有メモリ、決定論的なベンチマーク ガード、およびプラトー トリガーのドラフティングによる読み取り専用のクロスエージェント状態を通じて調整されます。カテゴリを代表する SOL-ExecBench ワークロードを使用して、NVIDIA H100 および B200 GPU 上の \kernelarc{} を評価します。結果として得られる実装は、カスタム BF16 GEMM、静的 cuBLASLt Expert-API 構成テーブル、後方融合エキスパート混合、シェイプゲート デコーダ層融合、ネイティブ NVFP4 グループ化クエリ アテンション、およびページング プレフィル アテンションに及びます。 2026 年 7 月~30 日に記録された公開 SOL-ExecBench リーダーボード スナップショットでは、これらの提出物は代表的な L1、L2、量子化、および FlashInfer タスクで 1 位にランクされました。この軌跡は、この論文の中心的な動機を裏付けています。つまり、マルチエージェントの共有検索により調査範囲が広がり、一定の候補予算内で強力な既存企業に到達できる一方で、個々の調整機能の価値はカーネルと最適化の段階に依存します。

原文 (English)

KernelArc: A Multi-Agent Framework for GPU Kernel Optimization

We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized agents run in parallel and coordinate through conclusions-only shared memory, a deterministic benchmark guard, and read-only cross-agent state with plateau-triggered drafting. We evaluate KernelArc on NVIDIA H100 and B200 GPUs using category-representative SOL-ExecBench workloads. The resulting implementations span custom BF16 GEMM, static cuBLASLt Expert-API configuration tables, fused mixture-of-experts backward, shape-gated decoder-layer fusion, native NVFP4 grouped-query attention, and paged prefill attention. In the public SOL-ExecBench leaderboard snapshot recorded on August~20, 2026, KernelArc ranked first on every representative L1, L2, Quantization, and FlashInfer task evaluated. The trajectories support the paper's central motivation: shared multi-agent search can broaden exploration and reach stronger incumbents within a fixed candidate budget, while the value of individual coordination features depends on the kernel and optimization stage.

13:00 JSTLLM/生成AI

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has becom…

13:00 JSTエージェント研究/論文Claude

FM-Bench: 競合エージェントとの長期的な管理のためのベンチマーク

言語モデル エージェントは、制限されたタスクを確実に実行するようになりました。行動が累積的な結果をもたらし、環境が彼らの選択に反応する場合、彼らが長期にわたって効果的な意思決定を維持できるかどうかは、ほとんど測定されていないままです。 FM-Bench (フットボールマネジメントベンチマーク) はこれを測定します。 LLM エージェントは、26 のツールとおよそ 340 ~ 400 の意思決定ストップを通じて、ゲーム内 20 年間フットボール クラブを運営します。すべてのライバルと同じ予算でチームをドラフトし、選手をトレードし、契約交渉をし、施設と若手に投資し、ラインナップを設定し、それを解雇できる理事会に回答する一方で、決定論的なエンジンが毎年、LLMの裁判官や人間の評価者なしで最終的なスコアを1つに蓄積していきます。ソロ トラックでは、凍結したスクリプト化された世界に対して 15 のフロンティア モデルのそれぞれが再生され、アリーナでは同じモデルとスクリプト化されたアンカーが 1 つの共有された 20 年間の世界に配置されます。私たちの知る限り、この規模での直接評価は初めてです。スコアの背後にある 6 つの行動能力を測定します。 3 つのシード全体で、15 モデルすべてがすべてのレベルを達成していますが、ほとんどのモデルでブラインド スクリプトベースラインは消滅しており、claude-fable-5 は平均スコアとアリーナでソロ ボードのトップに立っており、それでもタイトルは 10 モデル間で交代します。規模、価格、ベンダーのいずれも注文を予測しません。順序は地平線の後半でしか決まらず、最高のファーストプレイ人間がモデルボードの最下位にのみ着地します。モデルを区別するのは、計算ではなく管理動作です。より高いスコアのモデルは、終わり近くでペイオフの遅い投資を減らし、現金を遊ばせるのではなく投資し続け、期限のかなり前に更新を開始しますが、トークンの支出は何も予測しません。何百もの拒否された入札から市場の隠れた価格を学習するモデルはなく、自己管理メモリは、増大するだけのアーカイブかシーズンごとに書き換えられる計画という 2 つの相反するモードで失敗します。コードは https://github.com/Analogy-AI/fm-bench で入手できます。

原文 (English)

FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget as every rival, trades players, negotiates contracts, invests in facilities and youth, sets lineups, and answers to a board that can fire it, while a deterministic engine accumulates every year into one final score with no LLM judge or human rater. The solo track plays each of 15 frontier models against a frozen scripted world, and the Arena places the same models plus a scripted anchor in one shared 20-year world; to our knowledge, the first head-to-head evaluation at this scale. We measure six behavioral capabilities behind the score. Across three seeds, all 15 models complete every horizon while the blind scripted baselines die out in most of theirs, and claude-fable-5 tops the solo board on mean score and the Arena, where the title nonetheless rotates among ten models. Neither scale, price, nor vendor predicts the order; the order settles only late in the horizon, and the best first-play human lands only at the bottom of the model board. What separates the models is managerial behavior rather than computation. Higher-scoring models reduce slow-payoff investment near the end, keep cash invested rather than idle, and open renewals well before the deadline, while token spend predicts nothing. No model learns the market's hidden prices from hundreds of rejected bids, and self-managed memory fails in two opposite modes: an archive that only grows or a plan rewritten every season. Code is available at https://github.com/Analogy-AI/fm-bench.

13:00 JSTロボティクス

DECOWAM: 脚式モバイル操作のための分離された全身世界アクション モデル

モバイル操作では、移動と腕の動きが連携して将来の観察と制御がどのように変化するかをロボットが予測する必要があります。既存のワールド アクション モデルは、主に固定ベース プラットフォーム向けに開発されており、カメラのエゴモーションとベースおよびアームのアクションを明確に区別していません。ここでは、専用の条件付きインターフェイスを通じてこれらの要素を分離する全身世界行動モデル DECOWAM を紹介します。 DECOWAM は、適応された FastWAM バックボーンを凍結し、残りのアダプター、特権的な観察から抽出されたアクションに相当する将来のボトルネック、敵対的に分離されたベースとアームの潜在、およびビデオ予測のための基本速度調整をトレーニングします。さらに、ビデオ、全身の状態と動作、および言語を同期する実際のロボット データセットである ARMDOG を紹介します。固定再生プロトコルでは、DECOWAM は FastWAM よりも将来のビデオとアクションの予測の両方を改善し、2,595 万のトレーニング可能な適応パラメーターでアクション MSE を 21.7% 削減しました。メソッドごとに 79 回の閉ループ試行を行った結果、比較したシステムの中で観察された中で最も高い全身調整とベース変位の堅牢性が達成され、タスクの完了は依然として最強のベースラインと同等でした。これらの結果は、実施形態を意識した因数分解が、移動視点の下でパラメータ効率の高い共同視覚予測と全身制御をサポートできることを示しています。

原文 (English)

DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation

Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world-action models, developed largely for fixed-base platforms, do not explicitly distinguish camera ego-motion from base and arm actions. Here we introduce DECOWAM, a whole-body world-action model that separates these factors through dedicated conditional interfaces. DECOWAM freezes an adapted FastWAM backbone and trains residual adapters, an action-equivalent future bottleneck distilled from privileged observations, adversarially separated base and arm latents, and base-velocity conditioning for video prediction. We further introduce ARMDOG, a real-robot dataset that synchronizes video, whole-body state and action, and language. On a fixed replay protocol, DECOWAM improved both future-video and action prediction over FastWAM, reducing action MSE by 21.7% with 25.95M trainable adaptation parameters. Across 79 closed-loop trials per method, it achieved the highest observed whole-body coordination and base-displacement robustness among the compared systems, while task completion remained comparable to the strongest baseline. These results show that embodiment-aware factorization can support parameter-efficient joint visual prediction and whole-body control under moving viewpoints.

13:00 JST研究/論文

いつ考えるべきかを学ぶ: テスト時のコンピューティング割り当てのための適応推論

強化学習でトレーニングされた推論言語モデルは通常、明示的に適応するトークン バジェットではなく固定トークン バジェットの下で動作するため、簡単な問題では過剰な計算が発生し、難しい問題では不十分な計算が発生する可能性があります。私たちは、モデルが応答の最初のトークンとして、\textsc{NoThink} (できるだけ早く答える)、\textsc{Short} (短い推論)、または \textsc{Long} (拡張推論) の 3 つのモードのいずれかを選択することによって、独自の推論の労力を割り当てることを学習できるかどうかを研究します。選択は、個別のルーターを使用せずに、異なる応答長で各モードの価値を高める成形報酬と、モードを区別し続けるモードごとのハード トークン キャップを通じて、Group Relative Policy Optimization (GRPO) 内で学習されます。 MATH でトレーニングされた 1.5B の抽出されたモデルでは、3 つのモードが 1 つの選択肢に崩れることなく出現し、簡易モードは最終的に \textsc{Long} よりも正確になります。これは、ルーターが問題をランダムではなく難易度に基づいて並べ替えていることを示しています。 3 つのシードで平均した結果のポリシーは、ホールドアウトされた MATH500 での基本モデルの精度 ($0.782$ 対 \ $0.796$) に近いままですが、平均応答長は $4{,}796$ から $2{,}811$ トークンに短縮されました ($41\%$ の削減)。興味深いことに、再トレーニングなしで他のベンチマークにも移行し、問題が容易な場合に最大の節約効果が得られます。たとえば、GSM8K では 76\% のトークン削減が実現し、同様の応答長でのベースラインよりも高い精度で実現されます。つまり、各問題についてどの程度推論するかを適応的に選択する推論モデルを構築します。

原文 (English)

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adaptive one, which can lead to over-computation on easy problems and insufficient computation on difficult ones. We study whether a model can learn to allocate its own reasoning effort by choosing, as the first token of its response, one of three modes: \textsc{NoThink} (answer as quickly as possible), \textsc{Short} (brief reasoning), or \textsc{Long} (extended reasoning). The choice is learned inside Group Relative Policy Optimization (GRPO) with no separate router, through a shaped reward that makes each mode worthwhile at a different response length, together with hard per-mode token caps that keep the modes distinct. On a 1.5B distilled model trained on MATH, the three modes emerge without collapsing to a single choice, and the brief modes end up more accurate than \textsc{Long}, which shows that the router sorts problems by difficulty rather than at random. Averaged over three seeds, the resulting policy stays close to the base model's accuracy on the held-out MATH500 ($0.782$ vs.\ $0.796$) while cutting the mean response length from $4{,}796$ to $2{,}811$ tokens (a $41\%$ reduction). Interestingly, it also transfers to other benchmarks without retraining, with the largest savings where problems are easier, with for instance 76\% token reduction on GSM8K and at higher accuracy than the baselines at similar response length. In short, we build a reasoning model that adaptively chooses how much to reason for each problem.

13:00 JST研究/論文

AI-driven Prices for Externalities and Sustainability in Production Markets

Traditional competitive markets do not account for negative externalities; indirect costs that some participants impose on others, such as…

13:00 JST研究/論文

Graphon Particle Systems, Part II: Dynamics of Distributed Stochastic Continuum Optimization

We study the distributed optimization problem over a graphon with a continuum of nodes, which is regarded as the limit of the distributed n…

13:00 JSTLLM/生成AI

On the Within-class Variation Issue in Alzheimer's Disease Detection

Alzheimer's Disease (AD) detection commonly employs machine learning classification models to distinguish between individuals with AD and t…

13:00 JSTLLM/生成AIエージェント

Can We Trust AI Agents? A Case Study of an LLM-Based Multi-Agent System for Ethical AI

AI-based systems, including Large Language Models (LLMs), impact millions by supporting diverse tasks but face issues like misinformation,…

13:00 JST画像/動画生成研究/論文

An Automated Pipeline for Few-Shot Bird Call Classification: A Case Study with the Tooth-Billed Pigeon

This paper presents a largely automated one-shot bird call classification pipeline, incorporating targeted manual quality control steps, de…

13:00 JST研究/論文

SPD Matrix Learning for Neuroimaging Analysis: Perspectives, Methods, and Challenges

Neuroimaging provides essential tools for characterizing brain activity, structure, and connectivity through modalities that capture comple…

13:00 JSTLLM/生成AIハードウェア/半導体

Explaining Intrinsic Moral Self-Correction with Mechanistic Interpretability

Intrinsic moral self-correction refers to the phenomenon where a language model refines its ethical judgments or aligns its outputs purely…

13:00 JST画像/動画生成研究/論文

WeedNet: A Foundation Model-Based Global-to-Local AI Approach for Real-Time Weed Species Identification and Classification

Early weed identification is crucial for effective management and control, and researchers, agronomists, and technology developers are incr…

13:00 JSTエージェントロボティクス

Can you see how I learn? Human observers' inferences about Reinforcement Learning agents' learning processes

Reinforcement Learning (RL) agents often exhibit learning behaviors that are not intuitively interpretable by human observers, which can re…

13:00 JSTLLM/生成AI

GeoExplain: Multimodal Reasoning based on Hierarchy of Visual Information in Street View

Multimodal reasoning is a process of understanding, integrating and inferring information across different data modalities. It has recently…

13:00 JSTLLM/生成AI

CulTrace: Tracing Internal Cultural Reasoning in Large Language Models

The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidd…

13:00 JST研究/論文GPT / ChatGPT

AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning

Test-time scaling strategies for Large Language Models predominantly rely on either reinforcement learning with sparse outcome rewards or s…

13:00 JSTLLM/生成AI

SCOPE: A Generative Approach for LLM Prompt Compression

A big issue in modern LLM applications is they tend to feed long context to LLM, which results in high inference cost and latency, and may…

13:00 JSTLLM/生成AI

SKILL-RAG: Self-Knowledge Induced Learning and Filtering for Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) has significantly improved the performance of large language models (LLMs) on knowledge-intensive task…

13:00 JST研究/論文

Perseus: Interactive Time Series Segmentation with Sparse Supervision via Stateful Memory

Real-world systems, ranging from industrial manufacturing to wearable healthcare, generate multivariate time series with hierarchical state…

13:00 JST研究/論文

Significant Other AI: Identity, Memory, and Emotional Regulation as Long-Term Relational Intelligence

Significant Others (SOs) stabilize identity, regulate emotion, and support narrative meaning-making, yet many people today lack access to s…

13:00 JST画像/動画生成

Fine-tuning an ECG Foundation Model to Predict Coronary CT Angiography Outcomes

Coronary artery disease (CAD) remains a major global public health burden, yet scalable pre-imaging risk stratification tools are limited.…

13:00 JST画像/動画生成

MeltwaterBench: Deep learning for spatiotemporal downscaling of surface meltwater

The Greenland ice sheet is melting at an accelerated rate due to processes that are not fully understood and hard to measure. The distribut…

13:00 JSTLLM/生成AIエージェント

AgentOCR: Reimagining Agent History via Optical Self-Compression

Recent advances in large language models (LLMs) enable agentic systems trained with reinforcement learning (RL) over multi-turn interaction…

13:00 JST研究/論文

GroupSegment-SHAP: Shapley Value Explanations with Group-Segment Players for Multivariate Time Series

Multivariate time-series models achieve strong predictive performance in healthcare, industry, energy, and finance, but how they combine cr…

13:00 JST画像/動画生成

CFM: Language-aligned Concept Foundation Model for Vision

Language-aligned vision foundation models perform strongly across diverse downstream tasks. Yet, their learned representations remain opaqu…

13:00 JSTエージェント

Investigating Target Class Influence on Neural Network Compressibility for Energy-Autonomous Avian Monitoring

Biodiversity loss poses a significant threat to humanity, making wildlife monitoring essential for assessing ecosystem health. Avian specie…

13:00 JSTLLM/生成AIエージェント

Mind the Style: Impact of Communication Style on Human-Chatbot Interaction

Conversational agents increasingly mediate everyday digital interactions, yet the effects of their communication style on user experience a…

13:00 JST画像/動画生成

PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval

Personal photo albums are not merely collections of static images but living, ecological archives defined by temporal continuity, social en…

13:00 JSTLLM/生成AIビジネス/資金調達

Efficient Self-Evaluation for Diffusion Language Models via Sequence Regeneration

Diffusion large language models (dLLMs) have recently attracted significant attention for their ability to enhance diversity, controllabili…

13:00 JST研究/論文Gemma

Efficient Exploration at Scale

We develop an online learning algorithm that dramatically improves the data efficiency of reinforcement learning from human feedback (RLHF)…

13:00 JST画像/動画生成

InverFill: One-Step Inversion for Enhanced Few-Step Diffusion Inpainting

Recent diffusion-based models achieve photorealism in image inpainting but require many sampling steps, limiting practical use. Few-step te…

13:00 JST画像/動画生成研究/論文

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been…

13:00 JSTエージェント

ChemGraph-XANES: An Agentic Framework for XANES Simulation and Curation

Computational X-ray absorption near-edge structure (XANES) is widely used to interpret local coordination environments, oxidation states, a…

13:00 JSTLLM/生成AI研究/論文

AutoOR: Scalably Post-training LLMs to Autoformulate Operations Research Problems

Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating co…

13:00 JSTLLM/生成AIエージェント

Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Long-Horizon Workflows

Long-horizon tool-using tasks sometimes benefit from revisiting earlier subtasks, but explicit revisitation also adds routing, coordination…

13:00 JSTLLM/生成AIGemmaLlamaQwen

RefusalGuard: Geometry-Preserving Fine-Tuning for Safety in LLMs

Fine-tuning safety-aligned language models for downstream tasks often leads to substantial degradation of refusal behavior, making models v…

13:00 JST研究/論文

Preference-Based Self-Distillation: Beyond KL Matching via Reward Regularization

On-policy distillation is an efficient alternative to reinforcement learning, offering dense token-level training signals. However, its rel…

13:00 JST研究/論文

GRALIS: Fusing Coalition and Gradient Attribution with Closed-Form Conservation Error and Finite-Sample Guarantees

The main post-hoc XAI methods for deep networks -- GradCAM, SHAP, LIME, Integrated Gradients -- originate from heterogeneous theoretical fo…

13:00 JST研究/論文

S-AI-Recursive: Convergent Recursive Reasoning

This article introduces S-AI-Recursive, a bio-inspired Sparse Artificial Intelligence architecture in which reasoning is implemented as a h…

13:00 JST画像/動画生成

Component-Aware Structure-Preserving Style Transfer for Satellite Visual Sim2Real Data Construction

For camera-based satellite visual sensing, Sim2Real data construction requires images that approach real-domain sensor appearance while ret…

13:00 JST研究/論文

Behavior-Consistent Deep Reinforcement Learning

Reinforcement learning (RL) often exhibits high variance across training runs, leading to unreliable performance and posing a major challen…

13:00 JSTLLM/生成AI

サンプリングの損失: 単語カバレッジ スコア (WCS) による LLM の語彙到達可能性の評価

最新の大規模言語モデル (LLM) は、膨大な潜在語彙を持っているにもかかわらず、反復的で均質なテキストを生成するとしてしばしば批判されます。これまでの研究はモデル知識とトレーニングデータに焦点を当てていましたが、私たちは言語の多様性を抑制する際のデコードメカニズムの役割を調査しています。我々は、文脈上適切な人間の語彙が標準的なサンプリング フィルター (Top-$p$、Top-$k$、Min-$p$ など) によって数学的にどの程度刈り込まれるかを定量化する指標である Word Coverage Score (WCS) を導入します。 WCS は、静的な知識を評価するのではなく、サンプリング パラメーターの関数として、頻度が低く情報量の多い人間の単語の語彙生存率を測定します。人間が作成したコーパス断片のオープンウェイト モデルを監査することにより、確率空間内に存在する場合でも、どの論理語彙の選択肢がデコーダによって到達不能になっているかを特定します。私たちの結果は、業界標準のサンプリングのデフォルトが意図しない検閲メカニズムとして機能し、人間の表現の独特の質感を均質化された談話に平滑化するという定量的な証拠を提供します。 WCS は、テキストの一貫性と語彙の豊富さの間のトレードオフを最適化するための厳密なフレームワークを提供し、生成モデルで人間の言語の多様性を維持するための診断ツールを提供します。

原文 (English)

Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)

Modern Large Language Models (LLMs) are often criticized for producing repetitive and homogeneous text, despite possessing vast latent vocabularies. While previous research has focused on model knowledge and training data, we investigate the role of decoding mechanics in suppressing linguistic diversity. We introduce the Word Coverage Score (WCS), a metric that quantifies the extent to which contextually appropriate human vocabulary is mathematically pruned by standard sampling filters (e.g., Top-$p$, Top-$k$, and Min-$p$). Rather than assessing static knowledge, the WCS measures the lexical survival rate of low-frequency, high-information human words as a function of sampling parameters. By auditing open-weight models on human-authored corpus fragments, we identify which logical lexical choices are rendered unreachable by the decoder, even when they reside within the probability space. Our results provide quantitative evidence that industry-standard sampling defaults act as unintended censorship mechanisms, smoothing the unique textures of human expression into a homogenized discourse. The WCS offers a rigorous framework for optimizing the trade-off between text coherence and lexical richness, providing a diagnostic tool for preserving the diversity of human language in generative models.

13:00 JST研究/論文

Chatterbox-Flash: ストリーミング ゼロショット TTS 用の事前にキャリブレーションされたブロック拡散

Chatterbox-Flash は、事前トレーニング済みの自己回帰 TTS デコーダーをブロック拡散デコーダーに微調整することで得られるゼロショット テキスト読み上げモデルであり、ブロックごとのストリーミングを維持しながら各ブロック内で並列トークン生成を可能にします。主流のブロック拡散デコーディングを単純に離散音声トークンに転送すると、ロングテール トークンの分布により並列位置の選択が少数の高周波トークンに偏り、品質が低下することがわかりました。アーキテクチャを変更せずにこれを軽減するために、2 つの推論時テクニックを導入します。ブロック レベルの限界トークン分布を減算する事前に調整されたスコアリングと、調整された信頼度に基づいて反復を適応的に終了する早期デコード スケジュールです。標準のゼロショット TTS ベンチマークでは、Chatterbox-Flash は、強力な自己回帰ベースラインおよび非自己回帰ベースラインに匹敵する高忠実度の合成を達成しながら、ストリーミング AR システムと同等の最初のパケットまでの時間と大幅に低いリアルタイム係数によるストリーミング推論をサポートします。コードとオーディオのサンプルは https://github.com/resemble-ai/chatterbox-flash で入手できます。

原文 (English)

Chatterbox-Flash: Prior-Calibrated Block Diffusion for Streaming Zero-Shot TTS

We present Chatterbox-Flash, a zero-shot text-to-speech model obtained by fine-tuning a pretrained autoregressive TTS decoder into a block-diffusion decoder, enabling parallel token generation within each block while retaining block-by-block streaming. We find that naively transferring mainstream block-diffusion decoding to discrete speech tokens degrades quality, as a long-tail token distribution biases parallel position selection toward a few high-frequency tokens. To mitigate this without architectural modification, we introduce two inference-time techniques: prior-calibrated scoring, which subtracts the block-level marginal token distribution, and an early-decoding schedule, which adaptively terminates iteration based on calibrated confidence. On standard zero-shot TTS benchmarks, Chatterbox-Flash attains high-fidelity synthesis comparable to strong autoregressive and non-autoregressive baselines, while supporting streaming inference with time-to-first-packet on par with streaming AR systems and substantially lower real-time factor. Code and audio samples are available at https://github.com/resemble-ai/chatterbox-flash.

13:00 JSTLLM/生成AI

オーディオインタラクションモデル

オーディオは本質的にインタラクティブなモダリティですが、今日の大規模オーディオ言語モデル (LALM) はオフラインであり、ストリーミング オーディオ モデルはそれぞれストリーミング ASR や音声チャットなどの単一タスクのみを処理します。それらを 1 つのオンライン LALM に統合する時が来ました。LALM は、常時オンの知覚、決定、応答ループを通じて、音、環境、指示をリアルタイムで聞き、その場で反応するモデルです。私たちはこの体制をオーディオ インタラクション モデルとして形式化し、オーディオ インタラクションで実現します。これは、オフライン タスクの実行を保持しながら、対話からフル ボイス チャットに至るまでのオンラインの一般的な音声指示を追加し、ストリームのセマンティクスからいつ応答するかを決定する統合ストリーミング モデルです。これを可能にするために、ストリーミングネイティブのデータ構築、理解を意識したトレーニング、安定したリアルタイムインタラクションのための非同期低遅延推論を通じて、データからトレーニング、デプロイメントに至るまで、認識・決定・応答ループをエンドツーエンドでインスタンス化するフレームワークである SoundFlow を提案します。さらに、7 つの基本能力と 28 のサブタスクにわたる 260 万項目のストリーミング コーパスである StreamAudio-2M と、プロアクティブな音声介入を評価するための Proactive-Sound-Bench を構築します。 8 つのベンチマークにわたって、Audio-Interaction は主流のオーディオ タスクで競争力のあるパフォーマンスを維持しながら、リアルタイム ASR、ストリーミング オーディオ命令のフォロー、プロアクティブ ヘルプなど、オフライン LALM ではアクセスできない機能を解放します。

原文 (English)

Audio Interaction Model

Audio is continuous and interactive, yet most Large Audio Language Models (LALMs) remain offline and streaming systems usually specialize in ASR or spoken dialogue. We formalize the Audio Interaction Model, an always-on perceive--decide--respond paradigm that tracks context, decides whether intervention is warranted, and responds without stopping listening. We instantiate it with Audio-Interaction and introduce SoundFlow, coupling streaming-native data construction, comprehension-aware silence/response supervision, dual-loss training, and asynchronous FIFO inference. We also construct textsc{StreamAudio-2M, a 2.6M-item, 302k-hour corpus spanning 7 capability families and 28 sub-tasks, together with Proactive-Sound-Bench. Across 8 benchmarks, Audio-Interaction remains competitive on mainstream audio tasks while enabling spoken-instruction robustness, long-stream interaction, and proactive intervention.

13:00 JSTLLM/生成AI

INFUSER: Influence-Guided Self-Evolution Improves Reasoning

Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervi…

13:00 JST研究/論文

Valid Inference with Synthetic Data via Task Exchangeability

There is a proliferation of work arguing for the use of synthetic data in scientific research. For example, social scientists are arguing f…

13:00 JST研究/論文

暗号通貨 x AI、AI x 暗号通貨: 調査

暗号通貨と AI の交差点から、論文、製品、オンライン投稿、企業が生まれています。しかし、周囲の喧騒のせいで、具体的に何が行われたのか、何が機会と課題なのか、そしてどのような未解決の問題が注目に値するのかが曖昧になってしまいます。この調査報告書では、ブロックチェーンベースのテクノロジー(広義には「暗号」として解釈されます)に対して AI が何ができるのか(暗号 x AI)、およびその逆(AI x 暗号)について尋ねています。私たちは既存の研究を体系化し、重要なポイントを要約し、研究上の未解決の疑問を強調し、業界に蔓延している誤解についての視点を提供し、AI と暗号通貨はまだ有意義な統合の非常に初期段階にあると結論付けています。

原文 (English)

Crypto x AI, AI x Crypto: A Survey

The intersection of crypto x AI is spawning papers, products, online posts, and companies. All the surrounding buzz, though, obscures what exactly has been done, what the opportunities and challenges are, and what open questions deserve attention. This survey paper asks what AI can do for blockchain-based technologies (broadly construed as "crypto") (crypto x AI), and vice versa (AI x crypto). We systematize existing work, summarize key takeaways, highlight open research questions, and offer a perspective on pervasive industry misconceptions, concluding that AI and crypto are still in the very early stages of meaningful integration.

13:00 JSTLLM/生成AI研究/論文

The Metanym Game: An LLM Benchmark Without Ground Truth That Rises With the Models It Measures

We present evidence that analogy is at the core of LLM intelligence. In our benchmark, LLMs compete in generating sets of analogous stateme…

13:00 JSTLLM/生成AIビジネス/資金調達研究/論文Llama

Know2Guess: 大規模言語モデルにおける知識境界評価のための汚染を認識したマルチゾーン ベンチマーク

大規模な言語モデルの信頼性の高い評価では、データの汚染、プロンプトの特異性、または一般的な拒否行動と混同することなく、サポートされている回答とサポートされていない推測を分離する必要があります。凍結されたビルドタイム ラベルの下で、回答可能な知識から棄権が期待される未知への移行を測定するための、汚染を認識したマルチゾーン ベンチマークを示します。このベンチマークには、5 つのドメインにわたる 1,200 項目、明示的な棄権期待、汚染リスクのメタデータ、および公式の厳密なパーサーと正規化された堅牢性パーサーによる二重解析が含まれています。ロックされた回答または棄権プロンプト、回答のみのコントロール、およびプロンプト テンプレートのバリアントの下で、FLAN-T5、Qwen2.5-Instruct、および Llama-3-Instruct モデルを評価します。このベンチマークは、一般的な無回答行動では解決されません。FLAN のベースラインは、生産的な棄権に関しては弱いままですが、より強力な指導調整モデルは、選択的ではあるが回答から棄権への移行が不完全であることを明らかにしています。 Qwen2.5-3B-Instruct は全体的に最高の信頼性を実現していますが、回答が期待されるゾーンは依然として難しく、キャリブレーションは依然として不十分で、良性の項目の拒否は引き続き発生します。プロンプトおよびパーサーの堅牢性分析により、主要なランキングと定性的な結論が維持されます。したがって、このベンチマークは、回答可能性、棄権、拒否、および汚染を、LLM の信頼性の個別だが相互作用する側面として監査するための再現可能なプロトコルを提供します。データセットは、https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark で公開されています。

原文 (English)

Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models

Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with data contamination, prompt idiosyncrasy, or generic refusal behavior. We present a contamination-aware, multi-zone benchmark for measuring the transition from answerable knowledge to abstention-expected unknowns under frozen build-time labels. The benchmark contains 1,200 items across five domains, explicit abstention expectations, contamination-risk metadata, and dual parsing with an official strict parser plus a normalized robustness parser. We evaluate FLAN-T5, Qwen2.5-Instruct, and Llama-3-Instruct models under locked answer-or-abstain prompts, answer-only controls, and prompt-template variants. The benchmark is not solved by generic non-answer behavior: FLAN baselines remain weak on productive abstention, while stronger instruction-tuned models expose a selective but incomplete transition from answering to abstaining. Qwen2.5-3B-Instruct achieves the best overall reliability, but answer-expected zones remain difficult, calibration remains poor, and benign-item refusal persists. Prompt and parser robustness analyses preserve the main ranking and qualitative conclusions. The benchmark therefore provides a reproducible protocol for auditing answerability, abstention, refusal, and contamination as distinct but interacting dimensions of LLM reliability.The dataset is publicly available at https://github.com/renweimeng/Know2Guess-A-Contamination-Aware-Multi-Zone-Benchmark.

13:00 JST画像/動画生成Gemini

MatMMExtract: An Open-Source Pipeline for Panel-Level Extraction of Grounded Image-Text Pairs from Materials Science Literature

The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inacc…

13:00 JSTLLM/生成AIビジネス/資金調達GPT / ChatGPT

The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry

Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation stru…

13:00 JST研究/論文

ランク 1 コーナー: タスクにはワールド モデルからどの程度の等価性が必要ですか?

学習された世界モデルは、通常、品質がモデルに単にあるかどうかであるかのように、観測結果をどの程度忠実に再構築するか、報酬を予測するかによって判断されます。しかし、タスクが実際にモデルに必要とするものはさらに狭く、クエリが依存するいくつかの予測座標であり、これをクロージャと呼びます。私たちは、潜在的なクロージャーがどの程度表現されるようになるかは、モデルの能力やその観察によってではなく、モデルがトレーニングされる目的の次元によって設定されることを示し、既知のグラウンドトゥルース ク​​ロージャーを使用して制御された環境内の DreamerV3 スタック上でこれを直接測定します。整列されたスカラー値信号 (値の等価性の中心となる目的) は、いくつかの次元を必要とするクロージャーの 1 次元投影のみをインストールします。単一の線形プローブで読み取ると、スカラーが完全な目的に置き換えられると、回復可能な構造は R^2=0.10 から 0.76 に上昇します。対物レンズの次元を 1 から 4 にスイープすると、補助ヘッドを介して正確にその数の予測方向がインストールされ、モデル自身の値のヘッドを介して同じ階段が (減衰した大きさではあるが同じランクで) 表示されるため、解離はヘッドの形状のアーティファクトではなく次元的なものになります。容量を一致させた比較とその場での圧力チェックにより、明白な代替案が排除されます。法則がレジームを支配し、私たちはその境界を測定します。コンパニオン閉ループ タスクでは、その構造がフレームごとに観察可能であり、再構成によってその構造がインストールされ、スカラー目的で十分です。目的は、安価なトレーニング信号ではすでに回復できない潜在が正確に何を表すかを決定します。したがって、価値の同等性は全か無かではなく次元的なものです。よく知られている単一の報酬目標はそのランク 1 のコーナーであり、モデルは予測を求められる目標と同じくらいのタスクの構造を組み込みます。

原文 (English)

What a World Model Represents Is Three Questions

World models learn task-relevant information through many routes: observation reconstruction, recurrent state, temporal filtering, and explicit task supervision. Different routes can make different variables available. The same variable can also be available through several routes at once. When it is, looking at which route would increase the training loss most if removed does not tell you which route the model actually uses. The questions are reachability, whether a training signal can identify a task-relevant direction; admission, whether that direction is recoverable from the latent; and assignment, which eligible route carries it. We test them in environments with a known set of required coordinates. A direction cannot enter the latent unless some training signal can identify it. Reconstruction, recurrence, or filtering may already recover some of those coordinates; a reward or value head then has no residual direction to admit. For what remains, how many independent predictions the target supplies is how many coordinates install: one through four independent predictions admit one through four directions, including through the value head. Reachability is not admission: a temporal second-moment coefficient can remain absent under next-token prediction when accumulating it is a fraction of a percent of that loss, and a head that predicts the coefficient restores it. Assignment is a different test. Two routes that each carry the same variable when trained alone do not swap the carrier when we reverse which is more costly to remove. A recurrent model trained on a transformer's recorded sequences shows the same pattern. Near the point where the competing route is beginning to clear the probe threshold, independent training runs disagree. What a world model represents is therefore three questions: what information is reachable, what supervision admits, and which competing route carries it.

13:00 JST画像/動画生成エージェントビジネス/資金調達研究/論文

MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents

We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a state…

13:00 JSTエージェント

Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts

Agent Skills have become persistent behavioral artifacts across independent AI agent systems. They combine natural-language task specificat…

13:00 JST画像/動画生成

A Distributional Robustness Margin For Pathology Foundation Models

Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut lear…

13:00 JST研究/論文

見た目も正しく、機能も正しく: マルチスクリーン モバイル アプリ生成のためのプロジェクト レベルのベンチマーク

最近のマルチモーダル大規模言語モデルでは、ビジュアル デザインを実行可能なコードに直接変換できますが、実際のモバイル製品では、共有コンポーネントと動作するナビゲーションを備えた構築可能なコードベースにするために複数のスクリーンショットが必要です。このプロジェクト レベルの設定では、既存のデザインからコードまでのベンチマークの 3 つの制限が明らかになります。ベンチマークは、完全なコードベースではなく単一ページの生成に焦点を当てていること、ページ間のナビゲーションを評価できないこと、プロジェクト全体の保守性を測定していないことです。実際のモバイル アプリ、人間がレビューした画面、構造化されたページ関係の注釈、およびナビゲーション テスト仕様で構成される、プロジェクト レベルのマルチスクリーン モバイル アプリ生成のための最初のベンチマークである MobileForge を紹介します。 MobileForge は、ビルド、ナビゲーション、視覚的忠実性、コードの保守性、効率性の 5 軸評価をサポートしています。また、ナビゲーション評価におけるカスケード障害を回避するための状態分離ナビゲーション テストと、視覚判定の信頼性を向上させるためのアンカー参照リスト単位の視覚評価プロトコルも提案します。現在のモデルは、6 つのフロンティア マルチモーダル LLM でエンドツーエンドで実行され、コンパイルして正しいページに到達するモバイル アプリ プロジェクトを構築できますが、インタラクティブ ナビゲーションの信頼性は依然として低く、視覚的な忠実性と保守性は依然として遅れています。ベンチマークとサポート資料は https://github.com/anoa12159-hue/mobileforge_eval から入手できます。

原文 (English)

Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation

Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple screenshots to become a buildable codebase with shared components and working navigation. This project-level setting exposes three limits of existing design-to-code benchmarks: they focus on single-page generation rather than complete codebases, cannot evaluate cross-page navigation, and do not measure project-wide maintainability. We introduce MobileForge, the first benchmark for project-level multi-screen mobile app generation, comprising real mobile apps, human-reviewed screens, structured page-relationship annotations, and navigation test specifications. MobileForge supports five-axis evaluation of build, navigation, visual fidelity, code maintainability, and efficiency. We also propose state-isolated navigation testing to avoid cascading failures in navigation evaluation and an anchor-referenced list-wise visual evaluation protocol to improve visual-judge reliability. Across end-to-end runs on six frontier multimodal LLMs, current models can build mobile-app projects that compile and reach the correct pages, but interactive navigation remains unreliable and visual fidelity and maintainability still lag. The benchmark and supporting materials are available at https://github.com/anoa12159-hue/mobileforge_eval.

13:00 JST画像/動画生成

Action-grounded tissue affordance enables anticipatory auto-framing that lowers surgeon cognitive workload during laparoscopic surgery

In laparoscopy, surgeon gaze tracks where the instruments will act; easing this demand through visual attention modeling requires dense lab…

13:00 JSTLLM/生成AI

SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation

We aim to improve model performance in multi-reward reinforcement learning training process. Existing Group reward-Decoupled Normalization…

13:00 JSTハードウェア/半導体

Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers

Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale…

13:00 JST研究/論文

Defining Decentralization: An Ontological Perspective

Decentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as…

13:00 JST研究/論文

ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation

ELVAE places an input-dependent normal--inverse-gamma (NIG) hierarchy at each VAE latent coordinate, separating location uncertainty $u_{\m…

13:00 JST研究/論文

Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians

A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical…

13:00 JSTLLM/生成AI研究/論文GemmaQwen

DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data

Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committ…

13:00 JST研究/論文

Information Geometry of Message Passing

We show that the natural-gradient stationary condition of variational inference has an edge-local form on a Forney-style factor graph. We s…

13:00 JSTLLM/生成AI

SuTRA : ルート認識を備えた構造的に統合されたトークン化

既存のサブワード トークナイザーは統計的圧縮を最適化しますが、形態学的構造、特にルートと接辞の関係を無視します。これは、基本単位が文字ではなく複雑な正書法音節 (アクシャラ) である形態学的に豊富なインド言語にとって有害で​​す。周波数ベースの手法は単語を過度に断片化し、語根と接辞を恣意的に分割します。これを形態的粉砕と呼んでいます。我々は、アクシャラの不可分性を保持し、形態学的境界を越えるマージにペナルティを与える、形態学を意識したアルゴリズムである SuTRA (Structurally-Unified Tokenization with Root Awareness) を提案します。また、ヒンディー語、マラーティー語、グジャラート語用の新しい形態学的セグメンテーション データセットもリリースしました。 SuTRA は粉砕を低減し、BPE と比較して、形態学的アラインメント (境界 F1) で +14.7%、セマンティック回復可能性 (ヒンディー語) で +34% のピークゲインを達成します。これらの構造上の利点により、機械翻訳では平均 +8.08 chrF2 の向上が得られます。

原文 (English)

SuTRA : Structurally-Unified Tokenization with Root Awareness

Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between roots and affixes. This is harmful for morphologically rich Indic languages, where basic units are complex orthographic syllables (aksharas) rather than letters. Frequency-based methods over-fragment words, arbitrarily splitting roots and affixes - a phenomenon we term Morphological Shattering. We propose SuTRA (Structurally-Unified Tokenization with Root Awareness), a morphology-aware algorithm that preserves akshara indivisibility and penalizes merges crossing morphological boundaries. We also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati. SuTRA reduces shattering, achieving peak gains of +14.7% in morphological alignment (Boundary F1) and +34% in semantic recoverability (Hindi) over BPE. These structural gains yield an average improvement of +8.08 chrF2 in machine translation.

13:00 JSTロボティクス

CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning

Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it…

13:00 JST研究/論文

Separating Covariate Shift from Mechanism Change with Two Discriminators: CJSD, a Conditional Discrepancy with an Exact Covariate-Concept Decomposition

After the inputs X are known, how much additional information does the label Y carry about which dataset a sample came from? That single qu…