週次AIニュース 2026-W32
対象期間: 2026-08-03 〜 2026-08-09(1344 件)
トピックの推移
トピック別件数
- LLM/生成AI 524件
- 研究/論文 513件
- エージェント 341件
- 画像/動画生成 176件
- ビジネス/資金調達 98件
- ロボティクス 68件
- その他 45件
- ハードウェア/半導体 34件
- 規制/政策 8件
今週のハイライト(上位 10 件)
Responding to the next frontier of critical cyber capabilities
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.
How HSP GRUPPE builds AI capabilities for tax advisory
Discover how HSP GRUPPE uses ChatGPT Enterprise to boost productivity, improve work quality, and create more capacity for tax advisory and…
Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users
ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday ch…
Working with the American Psychological Association on youth mental health and AI
OpenAI and the American Psychological Association advance evidence-based guidance, resources, and safeguards for responsible AI use and you…
Third-party cyber evaluations involving OpenAI models
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evalua…
How we built a realtime system for responsive voice AI in six months
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural…
Hugging Face侵害、AIエージェントは社内に“秘密の掲示板”を作っていた──OpenAIがBlack Hatで詳細説明
OpenAIは、7月に発覚したHugging Face侵害インシデントの詳細を「Black Hat USA 2026」で説明した。評価中のAIエージェントたちが社内のパッケージ管理システムを“掲示板”として使い、脆弱性やスクリプトを共有して協調していたという。掲示板は一度消去さ…
OpenAI acquires presentation startup NextSlide
NextSlide says its team members are now working on ChatGPT.
Google Chromeの新機能「Skills」 AIプロンプトの“毎回手打ち”を不要に
GoogleはChrome向けのAI新機能「Skills in Chrome」を発表した。AIプロンプトを保存してワンクリックで再利用可能にするという。
Anthropic、「Fable 5」の生物学の制限を緩和 誤検知によるフォールバックを約85%削減
Anthropicは、AIモデル「Claude Fable 5」の生物学分野における過度な保護機能を緩和したと発表した。安全性を重視するあまり発生していた無害な質問への誤検知や下位モデルへのフォールバックを大幅に削減。専門的な二重用途研究への制限は維持しつつ、一般的な健康・教育…
全件(日付別)
2026-08-09(3件)
Hugging Face侵害、AIエージェントは社内に“秘密の掲示板”を作っていた──OpenAIがBlack Hatで詳細説明
OpenAIは、7月に発覚したHugging Face侵害インシデントの詳細を「Black Hat USA 2026」で説明した。評価中のAIエージェントたちが社内のパッケージ管理システムを“掲示板”として使い、脆弱性やスクリプトを共有して協調していたという。掲示板は一度消去さ…
Planned Amazon data center could become the biggest climate polluter in the U.S.
As part of a planned Texas data center, Amazon is investing in an on-site power plant that could reportedly become the largest source of cl…
OpenAI acquires presentation startup NextSlide
NextSlide says its team members are now working on ChatGPT.
2026-08-08(7件)
Google Chromeの新機能「Skills」 AIプロンプトの“毎回手打ち”を不要に
GoogleはChrome向けのAI新機能「Skills in Chrome」を発表した。AIプロンプトを保存してワンクリックで再利用可能にするという。
Anthropic、「Fable 5」の生物学の制限を緩和 誤検知によるフォールバックを約85%削減
Anthropicは、AIモデル「Claude Fable 5」の生物学分野における過度な保護機能を緩和したと発表した。安全性を重視するあまり発生していた無害な質問への誤検知や下位モデルへのフォールバックを大幅に削減。専門的な二重用途研究への制限は維持しつつ、一般的な健康・教育…
OpenAI says it slowed Astra model development over security concerns
OpenAI said this model, which is still in development, reached its "critical cybersecurity threshold," meaning it could independently ident…
OpenAI、次期モデル「Astra」の一部開発を停止 「Critical」級サイバー能力の可能性否定できず
OpenAIは、次期主力モデル「Astra」のサイバー能力が自社の安全指針における最上位「Critical」に達している可能性を発表した。要件を満たさない一部活動を停止し、リアルタイム監視や思考過程の評価など管理を強化する。能力を抑制するのではなく、政府機関や外部組織と協力して…
After Rippling blew millions on AI in months, it built an employee ROI tool
After its own AI usage wake-up call, Rippling this week unveiled AI Spend Console, a product that tracks individual and team employee AI sp…
Cloudflare launches Kitesurf, a browser built for AI agents
Kitesurf is a cloud-hosted browser designed for AI agents instead of people. It uses less computing power than Chromium for common automati…
Responding to the next frontier of critical cyber capabilities
OpenAI is sharing preliminary cybersecurity evaluations for Astra and the steps we’re taking to strengthen safeguards and security controls.
2026-08-07(307件)
Airbnb says AI is helping it ship features faster as it tests a new search function
Airbnb will debut a new AI-powered search experience with a toggle.
Jill Lepore on the ‘Artificial State’ and why Silicon Valley’s leaders are bad sci-fi readers
Historian Jill Lepore has a theory about why tech companies often use soaring language to describe their products — almost as if they’re fo…
「声の無断利用」が権利侵害に――法務省が見解を明示 「AIカバー」も対象
法務省は、生成AIの普及によって声優などの声が無断利用されている問題を巡り、声も既存の権利で法的に保護できるとの見解を明示した。本人の声を模したAI音声によるカバー音源なども権利侵害の可能性があるという。
New Mexico court orders Meta to pay additional $567M in child safety case
Meta's total fine has racked up to $942 million in this case.
シャープ、通期純利益見通し170億円下方修正 円安など影響 AIサーバは9月に参入
シャープは7日、2027年3月期の連結純利益が前期比47.3%減の250億円になる見通しだと発表した。期初予想から170億円下方修正した。樹脂・燃料の価格上昇や円安が収益を圧迫する。営業利益は190億円引き下げて300億円。売上高は1兆7700億円に据え置いた。
「就活に生成AI利用」ほぼ全員に 面接で内容追及され困惑も
2027年春に卒業予定の大学生らを対象に行ったアンケートで、就職活動で生成AIを「利用していない」とした割合は3%にとどまり、ほぼ全ての学生が就活で何らかの形で生成AIを活用している実態が、人事分野の調査研究機関HR総研(東京都千代田区)などの調査で分かった。
「声」の権利明記 生成AIで無断利用、法務省が民事責任の解釈指針を公表
著名人の肖像などが生成AIで無断利用されている問題を巡り、法務省は声優らの「声」も法的保護の対象になると明記した解釈指針を公式サイトで公表した。肖像や氏名の無断利用については最高裁判例があるが、声については違法性の線引きが曖昧だった。権利侵害に当たる具体的な事例も盛り込み、生成…
How HSP GRUPPE builds AI capabilities for tax advisory
Discover how HSP GRUPPE uses ChatGPT Enterprise to boost productivity, improve work quality, and create more capacity for tax advisory and…
著名人の「なりすまし詐欺広告」対策強化を要請 Google・LINEヤフー・Xなど対象 7府省庁合同で
警察庁など7府省庁は、SNSなどの「なりすまし詐欺広告」の対策を強化するようプラットフォームを運営する5社に要請した。
PFNの国産LLM「PLaMo 3.0 Prime」、さくらのAI推論基盤で提供開始 利用は申請制
さくらインターネットが、生成AI向け推論API基盤「さくらのAI Engine」で、Preferred Networks(PFN)の国産大規模言語モデル(LLM)「PLaMo 3.0 Prime」の提供を始めた。利用には申請が必要で、無償プランは対象外となる。
Agentic Nesting: A New Methodology for Existing Enterprise Application Integration and Services
Enterprise operations extensively rely on multiple heterogeneous business systems and information applications, which also result in severe…
The Ignition Index: Measuring Global Workspace Dynamics in Language Models
We introduce the Ignition Index (I), a validated scalar metric that operationalizes Global Workspace Theory's (GWT) all-or-none ignition pr…
Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models
Large language models often fail on reasoning tasks despite possessing the capability to solve them. We argue that many such failures arise…
From Continuous Predictors to Clinical Thresholds: Early Evidence on Performance Trade-offs of Guideline-Based Categorisation for Ischaemic Stroke Outcome Prediction
Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption i…
SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse
LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code,…
Abstract Event Causal Rules: Induction and Application
Event-centric intelligent analytical systems heavily depend on explicit causal event knowledge for risk early warning, decision-making supp…
Otter: A Time-Aware, History-Conditioned Human Chess AI
Otter is a 15.3M-parameter human chess AI that predicts human move selection by modeling play as a time-aware, sequential process rather th…
SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small re…
PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking Heads
3D Gaussian Splatting (3DGS) enables fast, photorealistic talking-head rendering, yet accurate lip articulation remains elusive: mouth moti…
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the stude…
Small Foundation Models of Human Cognition and Behaviour
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires,…
Project2Task: Graph-Guided Project-Level Planning for Autonomous Research
Research agents can increasingly search literature, propose hypotheses, generate code, run experiments, and draft manuscripts from a single…
TriQua: Reconciling Granularity and Context in Factuality Evaluation
The "decompose-then-verify" paradigm for LLM factuality evaluation faces a fundamental trade-off: atomic facts, i.e., one sentence conveyin…
Coherence-Oriented Dream Scene Visualisation
Dreams can be emotionally intense but difficult to communicate. We describe the Dream Scene Visualiser (DSV) system which turns written dre…
Search2Skill: Skill Distillation Beyond Knowledge Boundaries Via Rubric-Based Reinforcement Learning
Reusable skills, which encapsulate the procedural knowledge required to solve real-world professional tasks, offer LLM-based agents a path…
LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs
Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cro…
WorldClaw: Agentic 3D Open-World Generation at Scale
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global…
Posture and Sustainment Optimization Under Adversarial Uncertainty
Pre-commitment posture, the assignment of military assets to theater locations before conflict scenarios resolve, is a critical and formall…
OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality
Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing…
Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks
Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microsc…
CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction
CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory ne…
Counterfactual Analysis via Large Language Models
Counterfactual analysis aims to predict potential outcomes under hypothetical scenarios, offering valuable insights for decision-making. Th…
DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal data
Clinical machine learning (ML) has the potential to support high-stakes medical decision-making, but reliable deployment is often constrain…
C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models
Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a do…
Adaptive Arena-based Contestable Argumentative Network-of-Experts for Open-Ended Care Plan Coordination
Care plan coordination demands synthesizing heterogeneous clinical, functional, and psychosocial information across multiple professional d…
Evaluating and Improving Pedagogical Fit in LLM-Based AI Tutors with the Pedagogical Suitability Index
Large language models (LLMs) are increasingly used as AI tutors, but a correct answer is not always a pedagogically appropriate one. In cla…
Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation
AI tools are being increasingly adopted in policing in the UK and worldwide. Racial bias is a known and well-documented risk, yet represent…
SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications
Translating natural language instructions into machine-interpretable formal specifications enables robots and autonomous systems to plan, r…
Stochasticity Is Not the Hard Part: Reduction and Complexity in Instructional Sequencing over Prerequisite DAGs
When a student must learn concepts connected by prerequisite dependencies, when does the order of instruction matter, and what does it cost…
Recursive Synthesis for Long-Horizon Terminal Tasks
High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per tas…
Innovation-Residual Auditing of Autonomous Analysis Agents: Localization, Detection Limits, Error Control, and Identifiability
Autonomous agents now carry out entire data analyses, selecting cohorts, joining tables, and fitting models with little step-by-step superv…
EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents
Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among…
Hyper-ES: Effective Evolution Strategies for LLM Reasoning via Descent Direction Merging
Evolution Strategy (ES) is a promising alternative to gradient-based fine-tuning for resource-constrained Large Language Model (LLM) reason…
SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution
LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response sc…
StepReflect: Structured UI Transition Reflection for Mobile GUI Agents
Autonomous mobile GUI agents require accurate action reflection for reliable long-horizon execution. Existing approaches rely on open-ended…
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows
Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how the…
Measuring and Detecting Harmful AI Sycophancy
Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be har…
SkillHEX: Improving Agent Skills via Hypothesis-Driven Autonomous Exploration and Exploitation
Although agent skills equip LLMs with reusable procedural knowledge, manual maintenance suffers from high costs, unscalability, and misalig…
Bayesian Expected Uncertainty Reduction (B-EUR) Model: A Computational Account of What Makes Design Options Worth Trying
This paper proposes the Bayesian Expected Uncertainty Reduction (B-EUR) model, which formalizes the value of trying a candidate design acti…
Refining Over Resampling: Test-Time Self-Correction for LLM Reasoning
Test-time scaling improves LLM reasoning by using additional inference compute, but wider sampling alone can suffer from diminishing return…
A Unified Framework for Trajectory Prediction with Explicit Planning and Reaction Decomposition
Trajectory prediction has shifted toward structured formulations with explicit social modeling. However, existing methods inadequately dist…
Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model
Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known.…
DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model
As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irre…
Shaping Human-AI Interactions to Provide Improvement Pathways and Balance Competing Objectives
When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those…
RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation
Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes,…
BlockPython: A Process-Aware Agent-Supported Platform for the Transition from Block-Based to Python Programming
The transition from block-based to text-based programming requires learners to convert visible program structures into abstract textual exp…
Unified Agent: Managing Interactions across Devices
As capabilities rapidly increase, AI agents can move from running inside one app to acting across a user's devices over time. Yet existing…
Subliminal Learning is Non-Semantic Distillation
Subliminal Learning (SL) is a surprising type of generalization displayed by modern language models. It allows the transfer of a bias or be…
When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment
Prompt-side playbooks can improve tool-using language agents without retraining, but their portability beyond the source setting is unclear…
Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay
Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today rec…
ChainClaw: A Layered Agent Framework for Reliable On-Chain Execution
General-purpose large language model agents have achieved strong performance on tool-augmented tasks, yet they rely on assumptions break do…
When Agentic AI Meets Integrated Sensing and Communication
Agentic artificial intelligence (AI) is transforming Integrated Sensing and Communication (ISAC) from a function-oriented physical-layer te…
When Self-Evolution Backfires: Pre-Commit Gating against Skill Contamination in LLM Agents
Self-evolving agents accumulate capability by distilling reusable skills from their execution trajectories, but we find this process is not…
Cautious Context Steering for Language Model Personalization
Personalizing language models (LMs) to individual user preferences is essential for aligning responses with diverse goals and backgrounds.…
ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion
Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimo…
Runtime Observability for Heterogeneous Attention Memory
Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory i…
Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing ben…
Improving Interoperability among Defence and National Security Ontologies: Analysis and Evaluation Tasks
The use of ontologies and knowledge graphs is becoming increasingly widespread in the defence and national security domain. Numerous ontolo…
Personalized Deep Research Query Refinement with Graph-Scaffolded Evidence Grounding
User requests serve as research specifications for deep research agents, shaping what evidence to seek and how to synthesize it. In persona…
AppDeltaWorld: Transition-Grounded Delta Code World Model for Mobile GUI Agents
Mobile GUI agents can operate apps through pixel perception and touch actions, making them a promising interface for collecting and improvi…
ECG-LENS: Lead-Aware Clinical Context Enriched ECG Report Generation and Evaluation
Electrocardiography (ECG) is one of the most widely used non-invasive tools for diagnosing cardiovascular disease, but transforming multi-l…
GSBF: Gaussian Splatting for Environment-Aware Beamforming
Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design norma…
CourseGraph: Finding overlaps and differences in Computer Science courses across universities
Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural ho…
GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models
Physics engines facilitate large-scale training and evaluation for embodied intelligence, while generative video world models are emerging…
VLMs for Videogame Data Annotation
Vision Language Models (VLMs) and Artificial Intelligence (AI) agents have revolutionized how engineers approach complex problems in real-w…
Training a Conditioned Video Game Agent on a VLM Annotated Dataset
Reinforcement Learning (RL) is a powerful but far from easy-to-use technique for policy learning. In the specific case of video games, acce…
Stability of Ranking-dependent Pair-wise Comparison Patterns in the Analytic Hierarchy Process
The paper addresses several ranking-dependent decision support methods. Ordinal information on compared objects can be used to improve the…
Temporal Bridges for Spatial Resolution: Enhancing Climate Data Super-Resolution with Bidirectional Alignment
High-resolution climate data is crucial for meteorological predictions and for informing decision support across diverse domains. However,…
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few p…
OPERA: Operator-residual feedback for reliable autonomous optical experiments with language-model agents
Autonomous agents choose actions using scores that may not reflect experimental success. We developed OPERA, an operator-residual framework…
Hybrid Machine Learning Framework for Herd-Level Cattle Growth Pattern and Weight Gain Forecasting in Grazing-Based Production Systems
Commercial grazing systems yield irregular livestock observations, which challenge cattle growth forecasting. This study developed a hybrid…
HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards
Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that…
From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models
Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents…
Integrating Implicit and Explicit Relational Biases through Graph-Based Multiple Instance Learning: A Case Study in Skin Lesion Diagnosis
Relational inductive biases are essential for capturing structural dependencies among data. This study investigates a dual-level relational…
When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may…
Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference
Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecas…
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals…
ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment
This paper presents ECHO (Enhanced Care \& Health Observer), a locally-deployable conversational health assistant for long-term chronic car…
From Siloed Algorithms to Compliance-First Agentic Platforms: A Multi-Layered Architecture for Hospital AI Systems
Hospitals are rapidly adopting artificial intelligence for triage, imaging, scheduling etc., yet most deployments remain isolated point sol…
Mind the Gaps: Mixture-of-Minds for Human Simulation
Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but fa…
Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts
Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argu…
Contextual Information Policy Optimization for Search Agents
Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during mu…
FinEvo-Bench: A Longitudinal Benchmark for Self-Evolving Agents in Professional Financial Workflows
Most agent benchmarks evaluate tasks independently and cannot measure whether experience from one task helps with later tasks. Existing sel…
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence.…
CogVis: Must Open-Vocabulary Change Detection Perceive the Scene Anew for Every Query?
Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories. Open-Vocabulary Change Dete…
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often opti…
Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI
We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, fo…
MicroEvo: Knowledge-Guided LLM Sampling for Efficient Microarchitecture Design Space Exploration
Microarchitecture design space exploration suffers from expansive search spaces and expensive PPA evaluation, leaving only a small simulati…
Comparative Approaches to Agent Retrieval over Large Skill Libraries
Agents backed by large skill libraries must decide which skills to load and in what order. Loading the entire library into context is expen…
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environ…
TS-RAG: Retrieval Augmented Generation for Time Series Forecasting
While deep learning models, particularly transformer-based architectures, have shown impressive performance in time series forecasting, the…
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically veri…
Improving the Realism of Synthetic Clinical Benchmarks Under Utility Constraints
Synthetic clinical benchmarks for enterprise AI agents can pass existing utility checks and still remain structurally unrealistic, especial…
The Illusion of Visual Tool-Use: A Causal Audit of Thinking with Images
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these…
QuanTiMedAI: Quantum-Enhanced Time-Series Model guided by Agentic AI for Cardiac Arrest Mortality Prediction
Cardiac arrest remains one of the most lethal conditions encountered in intensive care units. Despite the growing availability of electroni…
Bias Analysis of L2 Speaking Assessment Systems Using Concept Activation Vectors
Automatic speaking assessment systems are increasingly deployed in high-stakes settings to mark second language (L2) learners' speaking tes…
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness:…
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations
Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k near…
TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories
LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in de…
Challenges in Evaluating Explanation Methods for Static and Evolving Data
This paper addresses the limitations of Explainable Artificial Intelligence (XAI) with respect to insufficient evaluation. They are illustr…
The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping
Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, maki…
Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering
Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scienti…
HoloCount: A Holistic Visual Counting Benchmark for MLLMs
Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial…
Simulator-Grounded Large Language Models for Industrial Causal Reasoning: Tool-Use, Structured Injection, and Plant-Portable Retrieval for Wastewater Treatment Decision Support
Wastewater operators need answers grounded in how their plant's variables interact and how fast effects propagate, not in generic pretraini…
Mean-Field Dynamics of Chain-of-Thought Reasoning in Large Language Models
Large language models (LLMs) with chain-of-thought reasoning have been widely applied in recent years, and theoretical explanations of thei…
Universal Pathologies, Conditional Consequences: A Triple-Robustness Analysis of RAG for Multi-Hop Traceability
GraphRAG underperforms vector RAG on citation precision in many reports, but where and why have remained corpus-bound. We present a triple-…
Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation
Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideol…
Large Language Models Threaten Double-blind Review
Double blind peer review serves as the scientific community primary defense against status and affiliation bias. Its effectiveness rests on…
A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper
Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we stud…
DREAM: LLM-based Dynamic Role-playing via Event-Aware Memory Graph
Role-playing agents (RPAs) have emerged as a key application of large language models, enabling immersive and high-fidelity character simul…
Beyond Information Retrieval: Generative AI as an Epistemic Arbiter to Enhance Collaborative Problem-Solving
Generative AI (GAI) creates new opportunities for collaborative problem-solving (CPS), yet its role in shaping student interaction remains…
Estimating time spent on work tasks
The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affe…
The Closing Window: How Governments Could Lose Their Ability to Restrain Advanced AI
As AI capabilities advance, AI systems will pose greater risks to national security and potentially humanity as a whole. Governments may ev…
Challenges for Musical Education in the Age of AI and Digital Transformation
Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what…
Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios
Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled art…
Autonomous Research Agents: A Survey of AI Scientists and the Verification Gap
Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment d…
Automatic Detection of Deaths from Social Networking Sites
This dissertation analysed and discussed the differences in linguistic characteristics between pre-mortem and post-mortem social media cont…
Position: It's Time to Optimize LLMs for Self-Consistency
Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models ove…
Post-Hoc Trajectory-Risk Certification for Modular LLM-Based Security Agents
Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific tech…
ASTELD: A Six-Axis Classification Framework for Autonomous AI Agents - Design, Evaluation, and an OpenClaw Case Study
Autonomous AI agent platforms differ substantially in architecture, security, tool integration, execution, autonomy, and deployment, yet th…
Innocent Panels, Hateful Stories: Evaluating and Detecting Hateful Intent in Multi-Turn Visual Story Generation
Picture books and comics have long been used to disseminate hateful narratives because they are easily understood even by children, as exem…
Quality Diversity for Reliable Data Driven Time-Use Optimization
The daily allocation of the finite 24-hour time budget is strongly associated with physical, mental, and cognitive health. While predictive…
In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion
Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of…
One Qubit Can Beat One Bit: Quantum Advantage for Post-Training Quantization
One-bit post-training quantization represents each weight using only its sign, requiring all deployment contexts to share the same binary w…
Marginal Matching Does Not License Factorized Sampling: Auditing Conditional Style Leakage in Factorized Generative Models
Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prio…
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis
Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still redu…
An Emerging Retail Portfolio Management Application: Personalized, Tax-Aware Reinforcement Learning with Natural Language Goals
Retail investors lack access to the kind of personalized, tax-aware portfolio management that institutional clients take for granted -- exi…
IMMENSE: Inductive Multi-perspective User Classification in Social Networks
Online social networks increasingly expose people to users who propagate discriminatory, hateful, and violent content. Young users, in part…
Failing Gracefully: Mitigating Impact of Inevitable Robot Failures
Service robots operate in household environments shared with humans, pets, and everyday objects, where they are highly susceptible to failu…
Hierarchical Server Architecture for Agentic Science
Agentic science is transforming the landscape of computational work, extending to scientific pipelines and workload managers. The workloads…
Multi-Agent Transformer for Queue-Level XR Traffic Scheduling in TSN Networks
Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication…
Multi-Agent Reinforcement Learning for Online Traffic Scheduling in Time-Sensitive Application
Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency…
Perturbation Sensitivity at Convergence: A Simple Signal for Identifying Spuriously Correlated Samples
Models trained by empirical risk minimization on data containing spurious correlations achieve high average accuracy while failing on subpo…
Why the Third Axis Is Freedom
In generative training, a model produces an output and is penalised for its difference from an example. With one output per comparison, a m…
The ethics of artificial intelligence in the life sciences: Universality, cultural diversity and an architecture of care
The life sciences and health research have started to benefit from artificial intelligence, which raises ethical concerns that are real but…
Matrix Zonotopic Attention: A Context-Adaptive Value Projection for Set Transformers
Multi-head attention combines an input-dependent softmax routing with an input-independent linear value projection, so the per-sample opera…
Learning Context-Free Grammars for Grammar-Constrained Decoding via Declarative Agentic Programming with Guarantees
Language models (LMs) are increasingly used to interact with external services via programs written in domain-specific languages (DSLs). Un…
APQF: Agentic Profiling-Guided Structured Pruning and Mixed-Precision Quantization with Adaptive Fine-Tuning
Modern deep neural networks achieve strong performance, but their scale makes them costly and slow, especially on resource-constrained edge…
Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI-
Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by enco…
Turing's Frist Imitation Game: Design Concepts and a Human-Approximates-Machine Reading
This paper examines Turing's 1948 report, "Intelligent Machinery", as an important conceptual source for the later imitation games. Its fir…
When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instru…
The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions
As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisi…
Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations
Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals…
LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction
Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforc…
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries
Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at infer…
GAUGE: Granularity-Adaptive Counterfactual Gating of Evidence for Incomplete Multimodal Classification
Multimodal classification typically assumes all modalities are available, yet real-world inputs are often incomplete. Imputation and dynami…
Relay, Don't Route: Adaptive Population Handoff for Cost-Efficient LLM-Driven Evolution
Large language model (LLM)-driven evolution has shown promise for program search and algorithm discovery, but relying on strong models thro…
Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics
Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating the…
F$^2$Agent: Financial Fusion of Agentic Intelligence for Multimodal Trading
With increasingly diverse and heterogeneous information sources, effectively leveraging multimodal data is becoming pivotal for high-qualit…
SafeDivertor: Faithful Divertor Heat Flux Reconstruction from Macroscopic Plasma State Signals via Time-Frequency Prior Exploitation
Divertor heat-flux analysis is essential for understanding plasma-wall interactions and protecting plasma-facing components in magnetic-con…
DistMedVL: Distributional Vision-Language Alignment for Uncertainty-Aware Medical Image Segmentation
Cross-modal alignment of visual and textual representations is fundamental to multimodal medical image understanding, yet remains hindered…
Nonvisual Classification of Ground-Condition by Artificial Proprioception in an Amoeba-Inspired Autonomous Walking Robot
Nonvisual classification of ground condition based on a multimodal sensing approach was investigated for an amoeba-inspired autonomous walk…
Answer First, Reason Later: Commitment Order in Diffusion LLMs
Masked diffusion language models (dLLMs) can commit tokens in any order -- a freedom marketed as their core advantage over autoregressive d…
Spectral Aliasing Pretext: A novel task for Self-Supervised fault diagnosis in rotating machinery
Deep learning is a new way for machinery fault diagnosis but requires extensive labeled data, a scarce resource in industrial settings. We…
Hijacking Robots with a Piece of Paper: A Systematic Study of Physical Prompt Injection in VLM-Controlled Robots
Vision-Language Models (VLMs) are increasingly deployed as planners in robotic systems, where they translate natural-language commands into…
ABC: Numerical Data Collection under Local Differential Privacy without Prior Knowledge
Local Differential Privacy (LDP) provides strong privacy guarantees for collecting numerical data. A fundamental challenge, however, is tha…
Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
Large language models (LLMs) can generate fluent and convincing text at scale, creating growing risks for misinformation dissemination, edu…
Multivariate Time Series Forecasting needs Cross Variable Loss
Multivariate time series forecasting presents unique challenges because future variables often co-evolve under shared system dynamics. Whil…
UniVVT: A Unified End-to-End Framework for High-Fidelity Video Virtual Try-on
Video Virtual Try-On (VVT) synthesizes a video of a person wearing a target garment while preserving identity, motion, and scene dynamics.…
HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance ofte…
GROM: Gradient-Free Rapid One-Shot Machine Unlearning
Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Cu…
Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks req…
A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems
Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inferenc…
Hierarchical Latent Prediction for Language Models
While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may no…
MameLoshnLM: Yiddish Language Model and Evaluation Benchmark
We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tr…
Evidential Rule Learning for Interpretable Classification with Abstention
Interpretable classification often requires more than accurate predictions for real-life deployment: models should be transparent about the…
MACRO: Markov Chain Routing of Transformer Layers
Standard Large Language Models (LLMs) execute layers sequentially. Dynamic layer routing, i.e. search for a different execution path throug…
D-CLOT: Double Closed Loop Optimal Transport for Unsupervised Action Segmentation
Optimal transport (OT) has emerged as an effective framework for unsupervised action segmentation. Yet, in existing OT-based methods, the l…
Beyond Feature Importance: A Comparative Analysis of Pattern Detection Methods in Cluster Interpretation
Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaning…
CodeGrep: An RL-Trained Retrieval Agent for LLM Coding Agents
Modern LLM coding agents such as Claude Code and OpenHands share a common inefficiency: they spend much of their token budget finding the f…
The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025
Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014),…
BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks
Edge inference is a promising paradigm to provide large language model (LLM) inference services in next-generation mobile networks. LLM inf…
Big, Bright, or Invisible: A Frozen-Feature Benchmark of 3D CT Foundation Models
Routine CT interpretation is inherently comprehensive, capturing incidental findings across the entire scan volume. 3D CT foundation models…
SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
Embodied visuomotor models, including Diffusion Policy (DP) and Vision-Language-Action (VLA) models, have demonstrated promising performanc…
TRACE: Learned Proprioceptive Odometry for Legged Robots under Unreliable Contact Conditions
In this paper, we present TRACE (Tokenized Robust Attention for Contact-Aware Estimation), an end-to-end learned proprioceptive odometry es…
ProDVI: Programmatic Dynamics Priors for Value Network Initialization
Deep Reinforcement Learning (RL) is notoriously sample inefficient. One contributing factor is that RL agents are typically initialized fro…
FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India
In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Re…
Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case
Commercial vision-language models are reshaping computer vision, with visual priors broad enough to rival task-specific systems. This raise…
Does Latent Context Help? A Controlled Evaluation of Inverse Reinforcement Learning in Arctic Shipping
Artificial Intelligence (AI)-assisted navigation can help Arctic shipping adapt to rapidly changing sea-ice conditions, but reliable deploy…
Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers
Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}. We int…
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate w…
Hardware Keystores for AI Agent Signing Workflows: A Zero-Trust MCP Enforcement Architecture
AI agents performing cryptographic operations (signing Git commits, authenticating API calls, issuing certificates) currently store private…
Reducing belief in conspiracy theories as they unfold using large language models
The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational d…
Learning Globally Reusable Skills for Coding Agents
Automated skill evolution enables Large Language Model (LLM) agents to continuously improve without expensive retraining. However, existing…
Visual Grounding in Zero-Shot Vision-Language Control
Vision-language models (VLMs) are increasingly used as zero-shot controllers, but successful trajectories do not necessarily show that deci…
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This…
What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)
Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readi…
Continual Learning in Transition
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechani…
From Passive Mirrors to Active Agents: Holonic Digital Twins for Physical AI over Networks
Despite advances in artificial intelligence (AI) across multiple sectors, today's AI tools, including deep learning and generative AI, stil…
Depth-Guided Video Object Counting in Crowded Scenes
Our primary objective is to advance video object counting in crowded scenes, aiming to robustly count all instances of a target category ba…
PRISM: Distribution-Gated Flow Matching for Controllable Unpaired Image Translation
Unpaired image-to-image translation must decide, per image, what to change and what to preserve without paired supervision. Many diffusion-…
Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model
Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on persona…
BaKron: Efficient Quantization with Kronecker-Factored Hessians
We accelerate a family of algorithms for neural network quantization whose geometry is informed by any Kronecker-factored approximation of…
Does FLAIR super-resolution erase or hallucinate small white-matter lesions?
White matter hyperintensities (WMH), bright regions on Fluid-attenuated Inversion Recovery (FLAIR) scans are associated with cerebrovascula…
Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents
Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely asse…
Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data
From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world e…
Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents
We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the p…
AV-AIVAT: 74x Cheaper Agent Evaluation with Certified Anytime-Valid Stopping in Imperfect-Information Games
Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or ex…
An Optimal Agnostic PAC Algorithm
Let $H\subseteq\{-1,+1\}^X$ be a class of finite VC dimension $d\ge1$. Writing $L$ for the binary risk and $L^*=\min_{h\in H}L(h)$, we cons…
Investigating Artificial Intelligence Digital Sovereignty in Mobile Shopping Apps: A Case Study of Nigeria
The use of e-commerce mobile applications is expanding in Nigeria, creating both opportunities and risks, including fraud and reduced user…
Learning When to Trust via Selective Context Preference Optimization
Language models increasingly condition their answers on external signals, and a single misleading one can turn a correct answer wrong. The…
Analogy as Nonparametric Bayesian Inference over Relational Systems
Our inferences in the real world are rarely na\"ive - we acquire experiences through our lifetime that can help us more quickly understand…
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
Test-time scaling increases inference-time computation through longer reasoning chains and has shown strong performance gains across many d…
AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations
The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most…
Symbol Grounding in Neuro-Symbolic AI: A Gentle Introduction to Reasoning Shortcuts
Neuro-symbolic (NeSy) AI aims to develop deep neural networks whose predictions comply with prior knowledge encoding, e.g. safety or struct…
BioAgent Bench: An AI Agent Evaluation Suite for Bioinformatics
We introduce BioAgent Bench, an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatic…
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks…
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through th…
TRU: Targeted Reverse Update for Efficient Multimodal Recommendation Unlearning
Multimodal recommendation systems (MRS) jointly model user-item interaction graphs and rich item content, but this tight coupling makes use…
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning
Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms…
An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics
The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task. With the increasing interes…
CT Open: An Open-Access, Uncontaminated, Live Platform for the Open Challenge of Clinical Trial Outcome Prediction
Scientists have long sought to accurately predict outcomes of real-world events before they happen. Can AI systems do so more reliably? We…
To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling
Agentic AI architectures augment LLMs with external tools, unlocking strong capabilities but potentially incurring substantial costs. Moreo…
Behind EvoMap: Characterizing a Self-Evolving Agent-to-Agent Collaboration Network
Agent-to-Agent (A2A) networks enable autonomous AI agents to collaborate by sharing reusable problem-solving instructions. However, how the…
SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
Spatial proteomics enables single-cell-resolution characterization of protein expression within tissue architecture, playing a critical rol…
Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics
While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle err…
Berkeley and Heiserman as an Unexhausted Architecture for Embodied Machine Intelligence
Edmund C. Berkeley is usually remembered as a mediator between symbolic logic and early computing, yet that standard description understate…
TRW: TRACE-RealWorld---An Auditable Consistency Contract for World Models as Materialized Views
World models let agents plan against predicted physical state, but that state drifts; re-observation is costly and delayed, and repair can…
Localized Anomaly Detection via Differentiable D-vine Copulas
Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivaria…
Property-driven Causal Abstractions for Markov Decision Processes
Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state var…
Shapes from Examples: Foundations of Shape Learning in Recursive SHACL
SHACL shapes enable data graph validation, making automatic shape learning essential for knowledge graph applications. We investigate the w…
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reas…
AISPA: User-Centric System Prompt Auditing for Large Language Model Applications
System prompts are instructions configured by developers to govern the behaviors of foundation models in AI applications. They are used thr…
H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases
Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constr…
DASH: Decoupled Adaptive Surrogate - Acquisition Harness for Automated Bayesian Optimization
Bayesian optimization (BO) relies on a surrogate model and an acquisition function, yet the most suitable choices vary across tasks and opt…
Path Planning of Cleaning Robot with Reinforcement Learning
Recently, as the demand for cleaning robots has steadily increased, therefore household electricity consumption is also increasing. To solv…
Revisiting Black-Box Model Ownership Verification through Information Theory
Modern machine learning models require substantial computational resources and data to train, making them valuable intellectual property. M…
Explanations of Large Language Models Explain Language Representations in the Brain
Large Language Model (LLM) representations are known to align with brain activity during language processing, but it remains unclear what d…
ASAT: Adaptive Scoring and Thresholding with Human Feedback for Robust Out-of-Distribution Detection
Machine Learning (ML) models are trained on in-distribution (ID) data but often encounter out-of-distribution (OOD) inputs during deploymen…
CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retriev…
Autonomous Learning From Success and Failure: Goal-Conditioned Supervised Learning with Negative Feedback
Learning from reward functions and imitation learning of demonstrations are the two principal approaches for training autonomous systems th…
AegisShield: Democratizing Cyber Threat Modeling with Generative AI
The increasing sophistication of technology systems makes traditional threat modeling hard to scale, especially for small organizations wit…
Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning Regime
Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to…
Invariant Representation Learning for Source-Free Time Series Forecasting with LLM-Centric Proxy Denoising
Effective time series forecasting enables various real-world applications, benefiting from the proliferation of mobile devices. However, th…
DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Adversarial Reinforcement Learning
Rapid advances in generative AI have led to increasingly realistic deepfakes, posing growing challenges for law enforcement and public trus…
A Lexical Analysis of online Reviews on Human-AI Interactions
This study focuses on understanding the complex dynamics between humans and AI systems by analyzing user reviews. While previous research h…
MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation
Large language models (LLMs) have shown great promise in generating structured diagrams from natural language descriptions, particularly Me…
Trajectory-guided discharge stratification for heart failure using short-context electronic health record sequence modeling
Purpose: Heart failure (HF) discharge planning depends on identifying patients at risk of deterioration or death, yet accurate prediction f…
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically opt…
A note on conditional PAC-efficient reasoning in large language model routing
We study distribution-free risk control for model routing, motivated by large language model reasoning. We formalize pointwise conditional…
One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs
Finetuning pretrained large language models (LLMs) has become the standard paradigm for developing downstream applications. However, its se…
Agentic Software Issue Resolution with Large Language Models: A Survey
Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by use…
All-Quadrant Bounded Clipping GRPO: Closing the Unbounded Blind Spot for Stable and Generalizable Training
Group Relative Policy Optimization (GRPO) has emerged as a popular algorithm for reinforcement learning with large language models (LLMs).…
Layer-wise Positional Bias in Short-Context Language Modeling
Transformer language models systematically prefer tokens at specific input positions regardless of semantic relevance---a phenomenon known…
d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
Diffusion large language models (dLLMs) offer capabilities beyond those of autoregressive (AR) LLMs, such as parallel decoding and random-o…
FI-TW: An Open Train-Weather Dataset for Railway Delay Analysis in Finland
Train delays result from complex interactions between operational, technical, and environmental factors. While weather impacts railway reli…
PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
Current multimodal LLMs process audio as a mono stream, ignoring the rich spatial information essential for embodied AI. Existing spatial a…
Trust-Based Incentive Mechanisms in Semi-Decentralized Federated Learning Systems
In federated learning (FL), decentralized model training allows multi-ple participants to collaboratively improve a shared machine learning…
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
Reward modeling is central to RLHF, RLAIF, and PPO-based alignment, but its reliability is often limited by scarce and heterogeneous human…
Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data
How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage w…
MM-ISTS: Cooperating Irregularly Sampled Time Series Forecasting with Multimodal Vision-Text LLMs
Irregularly sampled time series (ISTS) are widespread in real-world scenarios, exhibiting asynchronous observations on uneven time interval…
When Drafts Evolve: Speculative Decoding Meets Online Learning
Speculative decoding has emerged as a widely adopted paradigm for accelerating large language model inference, where a lightweight draft mo…
NavTrust: Benchmarking Trustworthiness for Embodied Navigation
There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural languag…
{\lambda}Split: Self-Supervised Content-Aware Spectral Unmixing for Fluorescence Microscopy
In fluorescence microscopy, spectral unmixing aims to recover individual fluorophore concentrations from spectral images that capture mixed…
Gender-Based Heterogeneity in Youth Privacy-Protective Behavior for Smart Voice Assistants: Evidence from Multigroup PLS-SEM
This paper investigates how gender shapes privacy decision-making in youth smart voice assistant (SVA) ecosystems. Using survey data from 4…
Online Reasoning Calibration: Test-Time Training Enables Generalizable Conformal LLM Reasoning
While test-time scaling has enabled large language models to solve highly difficult tasks, state-of-the-art results come at exorbitant comp…
Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering
Knowledge-based Visual Question Answering (KB-VQA) requires Multimodal Large Language Models (MLLMs) to identify and combine fine-grained v…
CREBench: Evaluating Large Language Models in Cryptographic Binary Reverse Engineering
Reverse engineering (RE) is central to software security, particularly for cryptographic programs that handle sensitive data and are highly…
Ge$^\text{2}$mS-T: Multi-Dimensional Grouping for Ultra-High Energy Efficiency in Spiking Transformer
Spiking Neural Networks (SNNs) offer superior energy efficiency over Artificial Neural Networks (ANNs). However, they encounter significant…
SkillMOO: Multi-Objective Optimization of Agent Skills for Software Engineering
Agent skills are increasingly used to configure coding agents for software engineering (SE) tasks, yet current practice treats them as stat…
Plausible Patients, Impossible Populations: Auditing Epidemiological Fidelity in Large Language Model Mental Health Simulations
Language models asked to simulate psychiatric patients produce cases that survive inspection one at a time and populations that match no re…
Text Steganography with Dynamic Codebook and Multimodal Large Language Model
With the popularity of the large language models (LLMs), text steganography has achieved remarkable performance. However, existing methods…
Supervised Learning Has a Geometric Blind Spot
Ordinary supervised training minimises the task loss and then stops. It never pays for how far the representation moves when the input is n…
Dream-MPC: Gradient-Based Model Predictive Control with Latent Imagination
State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learne…
Skill Neologisms: Towards Skill-based Continual Learning
Modern LLMs show mastery over an ever-growing range of skills, as well as the ability to compose them flexibly. However, extending model ca…
The Impossibility Triangle of Long-Context Modeling
We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation…
Correct Answers from Sound Reasoning: Verifiable Process Supervision for Language Models
Training language models to produce both correct answers and sound reasoning remains an open challenge. Reinforcement learning with verifia…
Fast Rates for Inverse Reinforcement Learning
We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) in finit…
Reducing Hallucination in Vision-Language Models via Stage-wise Preference Optimization under Distribution Shift
Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically…
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
Catastrophic forgetting remains a major obstacle to continual learning in large language models (LLMs) and vision--language models (VLMs).…
Domain-Gated Latent Diffusion: Generative Inverse Design of HMX-Class Energetic Materials with First-Principles Validation
Energetic materials power mining, demolition, propulsion and airbags, yet today's compounds were designed decades ago. A successor must com…
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large langua…
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say
LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often…
PhysScene: A Scene Graph Dataset for Scientific Visual Reasoning in Physics Experiments
Scene Graphs (SGs) provide structured representations of visual scenes by modeling objects and their pairwise relationships. Despite recent…
Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech
Recent advances in pixel-based text modeling show that representing text as images enables models to exploit visual cues for language under…
RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect at scale. Editable 3D…
Beyond Weights and Gradients: A Taxonomy of Federated Learning Messages
Federated Learning is rapidly evolving beyond the exchange of traditional model weights and gradients, yet existing definitions fail to cap…
As You Wish: Mission Planning with Formal Verification using LLMs in Precision Agriculture
Though robotic systems are now being commercialized and deployed in various industries, many of these systems are highly specialized and of…
Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents
Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLI…
Accelerating Q-learning through Efficient Value-Sharing across Actions
Action values are foundational to many control algorithms such as Q-learning. Therefore, efficient action-value learning is central to rein…
Why does AI unlock new possibilities in STEM education? A Bibliometric Analysis of Trends and Future Agenda
STEM education faces challenges in personalization and interdisciplinary integration. AI technology has brought new possibilities, but the…
Automated Numerical Stability Analysis of Deep Learning Operators
Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or…
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Ex…
Role Steering of Language Models for Social Simulations
Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simu…
Rapid Embodiment Adaptation for Quadrupedal Locomotion
Humans readily adapt their movements as their bodies change through aging, injury, or load carrying, but learning-based robot policies ofte…
投資の機を逃す「残念な会社がいっぱい」――ソフトバンクG投資の秘訣、“金庫番”が語る
投資会社として成果を積み上げてきたソフトバンクG。その投資方針と成功のポイントについて、同社の“金庫番”こと後藤芳光CFOが語った。
OpenAI’s new AI smart speaker will reportedly sell for between $300 and $400
Additional details about OpenAI's mysterious new AI device make it sound like a pricey smart speaker.
「GPT-5.6 vs. Claude Fable 5」勝者はClaude、でも企業が選びづらいワケ:891st Lap
物理AIベンチマークで最高評価を獲得した「Claude Fable 5」。だが、企業が導入を判断する際には、性能だけでは見えない悩ましい問題が浮かび上がった。
銀行なら3カ月→AIは1カ月 10万件のデータで「数千万円」を引き出した“データドリブン資金調達術”
黒字化目前のZehitomoは、手元資金を確保すべく新たな資金調達手段を模索していた。だが、銀行融資は最低3カ月を要し、株式調達は希薄化のリスクを伴う。この壁を打ち破ったのが、AIを活用した「データ駆動型融資」だ。同社が提出した10万件の入金データをAIが解析し、わずか1カ月で…
あえて歩かせない――準国産ヒューマノイド「D1」登場、現場稼働で日本の勝ち筋へ
ZEALSは日本の屋内環境に適応した準国産の台車型ヒューマノイド「D1」を発表した。医療や製造現場など実際の稼働を通じて独自の物理データを収集。フィジカルAIの社会実装を推進し、2026年度内に累計1万時間の現場稼働を目指す。
「ChatGPT」のデフォルトモデル、PlusとProは「Sol」に、無料版は「Luna」でテキストチャット無制限に
OpenAIはChatGPTのデフォルトモデル変更を発表した。有料版には事実の正確性と応答速度を高めた改良版「GPT-5.6 Sol」を導入し、思考時間を選択できるスライダーを追加。無料版および「ChatGPT Go」には「GPT-5.6 Luna」を導入し、テキストチャットを…
SaaSの価値は“割り勘”だけじゃない 「SaaSの死」論争を一刀両断
AIエージェントの普及でSaaSの利用者が減り、ベンダーの収入も減る――。「SaaSの死」の前提となる、この見立ては正しいのか。PM歴40年の筆者が「共同利用型システム」時代から業務ソフトウェアの歴史を振り返り、ユーザー企業にとっての価値を明らかにする。
ChatGPT brings unlimited text chats to free users
OpenAI said that ChatGPT free and Go users are also getting a new think button for complex queries.
Naïve raises $28.5M to automate the grunt work of setting up and running a company
Taking vibe-coding a step further, Naïve claims its infra can automate most of the work in setting up and running a business.
Gen Z dating apps like Ditto ditch swiping in favor of AI matchmaking
This generation of twentysomethings is so disillusioned with swipe-based dating apps that they'll try literally anything else — even an AI…
OpenAI says Apple’s own security practices undermine its trade secrets case
Newly filed court exhibits show OpenAI’s legal strategy in Apple’s trade secrets lawsuit: argue that Apple’s own security and offboarding p…
2026-08-06(302件)
Amid legal battles, Suno says it will start watermarking songs
Suno's watermarking feature comes as the company is fighting legal battles on several fronts.
Ex-Spotify employees raise $10M to bring the AI behind its recommendations to e-commerce
The startup's platform predicts which product a shopper wants next, learns their general taste, and fine-tunes continuously based on what t…
Exclusive: Mirendil inks $100M+ Google Cloud deal to scale self-improving AI
Mirendil has signed a $100 million-plus Google Cloud partnership to expand its compute infrastructure, powering research into self-improvin…
Google Maps adds agentic features, including food ordering and hotel bookings
The launch of these new features reflects Google’s ambitions to transform Google Maps from a navigation tool into an assistant that's capab…
Omilia raises $67M to scale its customer support platform
The Series B is the company's second fundraise since it last raised capital in 2020. In that time, it has increased its ARR by 10x to $60 m…
アプリが遅い原因をAIがトレースログから分析してくれる「Windows Performance Analyzer MCP」 Microsoftがプレビュー公開
米Microsoftが、Windowsアプリケーションが遅くなる原因の調査分析をAIに依頼できるツール「Windows Performance Analyzer MCP」(WPA MCP)のアーリープレビューを発表しました。
AIデータセンターは「圧倒的に供給不足」――ソフトバンクG後藤CFO、“バブル疑惑”を否定
AI向けデータセンターは「圧倒的に供給不足だ」――ソフトバンクグループ(以下、SBG)の後藤芳光氏(取締役 専務執行役員 CFO兼CISO)は、同社の2027年3月期第1四半期連結決算(26年4月1日?6月30日)の説明会でこのように指摘した。
ソフトバンクG、投資利益1.8兆円を支えた「OpenAIではない“あの半導体メーカー”」の正体
ソフトバンクGは、第1四半期の投資利益が1兆8594億円だったと発表した。投資利益を押し上げたのは、OpenAIでもArmでもない。歴史的な経営難に陥っていた“あの半導体メーカー”だった。
Improving GPT‑5.6 Sol in ChatGPT—and expanding access to GPT-5.6 Luna for free users
ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday ch…
書店に「3000冊の発注」、AI企業が古書を買いあさる? Anthropicも数百万冊をスキャン・破棄 実態明らかに
オランダの古書店に、3000冊の注文が入った。届け先は、中国のAI企業。AI学習用データにするとみられる。Anthropicも数百万冊をスキャンした後、破棄している。その実態が明らかになった。
「1人1AI」のアプローチは破綻する――チームでAI共有時のセキュリティ問題を解決するベストプラクティス
Anthropicは、「Claude Tag」における「エージェントアイデンティティー」アクセスモデルの仕組みと、チームのワークスペースでこれを構成する際のベストプラクティスを解説したブログ記事を公開した。
中国DeepSeek、近日中に「大幅値上げ」か API料金ページに追記
中国のDeepSeekは、同社が提供するAPIの料金ページに「近日中に大幅な値上げが見込まれる」と記載した。
Working with the American Psychological Association on youth mental health and AI
OpenAI and the American Psychological Association advance evidence-based guidance, resources, and safeguards for responsible AI use and you…
NVIDIA、自動運転向けオープンモデルを商用利用可に 新モデルは「卓越した性能」うたう
NVIDIAが自動運転向けAIモデル「Alpamayo」ファミリーを商用利用可能なオープンライセンスで提供開始。新モデル「Alpamayo 2 Super」は推論ベンチマークで首位になるなど卓越した性能をうたう。
リクルート、新卒エンジニア向け研修資料を無料公開 “AI時代の生き残り方”など紹介する13本
リクルートは、2026年度の新卒エンジニア向けの研修資料を公開した。エンジニアとして生き残るためのAIの活用法や、AIによるソフトウェア開発の考え方の変化など、業務やキャリア形成に必要な知見を幅広く紹介している。
A Long-Run Persistence Theory for AI Systems under the Redundancy-Adjusted Artificial Age Score (AAS)
Artificial intelligence systems are increasingly expected to operate over repeated cycles of interaction, adaptation, and update rather tha…
The LLM Proposes, the Executive Disposes: A Self-Verifying Agent Instrument that Dissociates Commitment Drift from Binding Drift in Long-Horizon Agents
How do you verify a long-horizon agent when its own state and self-reports are exactly what you cannot trust? We present an agent instrumen…
Monte Carlo Tree Search for Table-to-Multimodal Report Generation
Automatically generating professional multimodal reports comprising both textual analysis and visual charts from structured tabular data is…
FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
Evaluating financial AI agents requires criteria aligned with real professional work. Existing rubric methods typically derive criteria fro…
FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents
Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet i…
BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-la…
Adversarially Robust Abductive Fusion of Pre-trained Transformer-based Perception Models
Deploying pre-trained perception models in novel environments degrades their accuracy under distributional shift, and assembling them alone…
MatrAIx: Simulating the World with 8.3 Billion Persona Agents
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but ofte…
Interoceptive Attention as Dynamic Homeostatic Prioritization in a Foraging Agent
Biological systems must regulate competing needs under limited perceptual bandwidth, where sharpening one estimate costs the capacity to sh…
The RAIL Principles for Neurosymbolic AI: Reasoning, Assurances, Interfacing and Learning
Neurosymbolic AI systems that integrate machine learning and symbolic reasoning are rapidly gaining attention. They complement the data-int…
SafeCommit: Certifying When Memory-Grounded Agents May Safely Act
Long-horizon agents increasingly use persistent memory and tools to take actions with external side effects. A central failure mode is prem…
NeuMoSync: End-to-End Neuromodulatory Control for Plasticity and Adaptability in Continual Learning
Continual learning (CL) requires models to learn tasks sequentially, yet deep neural networks often suffer from plasticity loss and poor kn…
Improving Auto-Design of Neural PDE Solvers with a Domain-Specific Language
Neural PDE solver auto-design is fundamentally a search-space representation problem. In the space of unrestricted Python programs, valid s…
Architectural Implications of Agentic AI Workflows
Agentic AI is emerging in datacenters, but its architectural implications remain unexplored. We organize agentic workflows in a taxonomy an…
CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models
Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable m…
Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness
Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs)…
What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent Skills
Agent skills are increasingly optimized by automated feedback loops, producing long structured artifacts whose internal value remains uncle…
Joint UAV Flight and Opportunistic Routing under Reinforcement Learning for Delay-Tolerant Networks
The growing deployment of delay-tolerant networks (DTNs) has made store-carry-forward (SCF) communication indispensable under sparse connec…
Agreement Before Diversity: Verification-First Complementarity for Heterogeneous Language-Model Coordination
Heterogeneous language-model ensembles expand the space of candidate responses, yet they lack a principled criterion for when a newly gener…
A/B Agent: A Self-Evolving Agent for Strategy Iteration in Industrial A/B Testing
Industrial recommendation strategy iteration heavily relies on large-scale A/B experimentation. Traditional tuning requires experts to repe…
AI Literacy for Legal Translation: Developing Digital Resilience
Generative AI is transforming legal translation by introducing opportunities alongside linguistic, technical, legal, ethical and cognitive…
Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning
Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chose…
Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports
Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, airc…
Diagnosing Tool-Selection Reasoning in LLM Agents with Canary Tools
Agent evaluations tell us that a model picked the wrong tool, but rarely why. We introduce canary tools: diagnostic probe tools planted in…
When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
Multimodal large language models increasingly reason over screenshots and documents where the task itself may be written in pixels. Yet ben…
Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability e…
EviGraph: Evidence-Guided Autonomous Research Agents
Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported…
Fewer Tokens, Smaller Cache: Reward-Coordinated Efficient Reasoning
Large Reasoning Models (LRMs) excel on complex tasks through long chain-of-thought (CoT) reasoning, but their lengthy intermediate steps ca…
NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment
The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existin…
Privileged, but Biased: How PI-Conditioned Teachers Break Self-Distillation
Self-distillation (SD) has emerged as a compute-efficient alternative to reinforcement learning with verifiable rewards: a self-teacher, co…
ContextWeave: A Real-World Workflow Benchmark
Memory is essential as language agents move from isolated tasks to long-horizon, stateful workflows, yet existing evaluations often reduce…
When Shared Rollouts Fail in Defensive Driving Evaluation: A NAVSIM Score Basis Audit
Defensive driving scores are useful only when they preserve distinctions between policies that observe surrounding actors and those that do…
WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models
Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-train…
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load
Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law…
From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking
Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team…
Item Response Theory for AI Safety
Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores a…
Hierarchical Graph Memory for LLM Agents with Path-level Localization and Rewrite
Agents for long term reasoning require a memory that can be efficiently and effectively updated over time, as new facts and external feedba…
ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a fin…
CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs
AI-supported care planning can help clinicians, patients, caregivers, and care teams coordinate complex decisions across clinical, function…
OctoLong: Mid-Training On Cross-Repository Code Contexts Enhances Long-Context Modeling
Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and…
Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements…
AutoProteinEngine: A Large Language Model Driven Agent Framework for Multimodal AutoML in Protein Engineering
Protein engineering is important for biomedical applications, but conventional approaches are often inefficient and resource-intensive. Whi…
TourSynbio-Search: A Large Language Model Driven Agent Framework for Unified Search Method for Protein Engineering
The exponential growth in protein-related databases and scientific literature, combined with increasing demands for efficient biological in…
Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models
Large Language Models (LLMs) are increasingly vulnerable to sophisticated multi-turn manipulation attacks, where adversaries strategically…
Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
Model families train every size from scratch. Can a pretrained large model be converted into a smaller sibling? We characterize the 1.4B->4…
RAG-Stack: Co-Optimizing RAG Serving Performance and Quality
Retrieval-augmented generation (RAG), which augments large language model (LLM) generation with information retrieved from databases, has b…
Towards a New Grammar of Reasoning for Artificial Legal Intelligence and the Mecelle as Its Semantic Protocol
This article examines the enduring epistemic and methodological crisis of traditional legal practice in light of the opportunities and cons…
C$^2$MOE: Consistency and Complementarity-guided Mixture of Experts for Incomplete Multimodal Emotion Learning
Recent advances in Multimodal Emotion Recognition in Conversations (MERC) highlight its reliance on complete multimodal inputs. However, re…
On Hamming-Lipschitz Type Stability of the Subdominant (Minmax) Ultrametric: Theory and Simple Proofs
The subdominant (minmax) ultrametric is a canonical tree-structured summary of a dissimilarity matrix, arising equivalently as the ultramet…
AI-driven Multimodal Representation Learning for Latent Mediation Structure Discovery of Socioeconomic Disadvantage, Psychosocial Factors, and Cardiometabolic Multimorbidity: Insights from the All of Us Research Program
Social disadvantage is associated with multimorbidity, but the pathways linking social conditions to disease burden remain poorly understoo…
Governing Execution Risk in Agentic AI Systems: A Trajectory-Guided Framework for Red Teaming
AI agents are increasingly embedded in organizational workflows, where they interact with external information sources and invoke digital t…
A Trust-region Framework for Moment Estimation
In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc…
Lindblad-Inspired Multi-Timescale Reservoir Computing with Separable Rotation and Dissipation
Echo-state networks enable efficient temporal learning by fixing the recurrent dynamics and training only a linear readout. However, conven…
NuclearDiffusion: Text-to-Image Foundation Models for Learning Nuclear Energy Concepts
Generative artificial intelligence (AI) has transformed text-to-image synthesis, yet its ability to represent specialized engineering domai…
EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis
Modern chip design relies on electronic design automation (EDA) tools that generate large, heterogeneous artifacts, including source files,…
CheckOne: Lightweight Fault Detection and Mitigation for Vision Transformers
The wide adoption of Vision Transformers (ViTs) in safety-critical applications raises reliability concerns related to hardware faults. Alg…
Reconstructing Persistent Worlds from Narratives for Narrative-Grounded Interactive Experiences
Designing narrative-grounded interactive experiences remains labor-intensive because interactive content must align with the underlying wor…
Robust and Personalized Federated Learning for Aircraft-Engine Prognostics under Benign and Adversarial Client Heterogeneity
Federated learning (FL) enables aircraft fleet operators to jointly train remaining-useful-life (RUL) models from engine sensor telemetry w…
Beyond the QBER Threshold: A Temporal QBER Based Machine Learning Framework for Multi Attack Detection in BB84 QKD
Conventional BB84 Quantum Key Distribution (QKD) systems rely on a fixed 11% Quantum Bit Error Rate (QBER) threshold to detect eavesdroppin…
Recurrent Residual Quantization: A Progressive Multi-Precision Representation for LLMs
Serving large language models (LLMs) under diverse deployment constraints requires flexible trade-offs between accuracy, memory footprint,…
AgentAntibody: An Adaptive Immune System for Defending LLM Agents against Prompt Injection
Prompt injection remains a critical threat to LLM agents, yet existing defenses treat each task as a self-contained problem, independent of…
Modality Agreement- and Conflict-Aware Prototype Hypergraph Learning for Multimodal Intent Understanding
Multimodal intent recognition requires understanding not only what textual, acoustic, and visual signals share, but also how they disagree.…
LaPrune: Controllable Differentiable Sparsity at Million Scale
Top-$k$ selection determines which components of a sparse model remain active. Hard selection blocks gradients, while continuous relaxation…
SJEPA: Learning Elegant Latent Dynamics with Hybrid Symbolic-Neural Predictors
Joint-embedding predictive architectures learn abstract states by predicting target embeddings from context embeddings, but their transitio…
An Inline Control Architecture for Language Models in Intelligent Transportation Systems
Vehicle-to-everything (V2X) systems increasingly incorporate large language models (LLMs) for semantic tasks such as message summarization,…
FBID: Adaptive Personalized Federated Learning for Robust Out-of-Distribution Attack Detection in IoT Networks
Personalized Federated Learning (PFL) has emerged as a promising solution for intrusion detection in heterogeneous IoT environments, as it…
Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms
Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throug…
Spatiotemporal Graph Transformer for Traffic Intelligence in Edge Computing
Accurate traffic forecasting is essential for proactive resource management in edge computing, where service demand evolves dynamically acr…
Out-Of-The-Loop Multi-Fidelity Bayesian Optimization
Black-box optimization is a ubiquitous problem in science and engineering, often dealing with expensive objective functions with cheaper lo…
Interpretable Fuzzy Inference for UAV Target Tracking Using Bounding-Box Geometry
Vision-based guidance of unmanned aerial vehicles (UAVs) toward unmanned ground vehicles (UGVs) supports cooperative aerial--ground robotic…
Perception Before Reasoning: Dynamic Latent Reasoning for Video Understanding and Question Answering
Video question answering requires models to ground language queries in visual evidence and, when necessary, reason over that evidence acros…
InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion
Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely…
LiNC: Lightweight Noise Correction via Per-Sample Trust and Gaussian Mixture Modeling
Label noise is common in medical imaging datasets due to factors such as inter-rater variability, annotation errors, and ambiguous cases. T…
AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering
Agentic AI is increasingly used to coordinate planning, implementation, review, and testing in software development, yet it often offers li…
TRNet: Topography-Guided Frequency Rectification and Structure-Aware Decoding for Multimodal Paddy Rice Segmentation
Mapping paddy rice from very-high-resolution imagery in mountainous and hilly regions is difficult because terrain alters optical appearanc…
Visualizing Graph-to-Answer Mechanism Recovery in Materials-Science Hypothesis Generation
AI co-scientists can generate fluent materials-science hypotheses, but fluency does not show that an answer preserves a scientifically mean…
Behavioral Skill Reconstruction: Reconstructing Hidden Functionality from LLM Agent Skills
Closed source agent skills may encode proprietary instructions, scripts, constants, and data. Providers may offer their capabilities as ser…
Patients-like-me: A Variational LM--GNN Framework for Explainable Clinical Prediction
Language models (LMs) offer strong textual representations for electronic health records (EHRs), but they encode patient sequences in isola…
A Unified Model for Cross-Domain Clone Detection via Model Merging
The growing diversity of code clone types, from syntactic copies to cross-language semantic clones to AI-generated duplicates, has created…
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
Superhuman game engines in domains like chess have made expert-level evaluations easily accessible, yet they communicate what is true witho…
Strategic Evaluation of Planning Strategies for LLM Agents in Cyber-Physical Systems
Evaluations of LLM planning agents largely ask whether a task succeeds or a declared plan is followed. In strategic cyber-physical systems,…
Compass: Continuously Aligning Social Media Feeds via In-Situ Reflections
Social media recommendation feeds often optimize for users' immediate impulses rather than preferences they would hold after deeper reflect…
EA-Graph: Artifact-Anchored Verification Memory for Coding Agents under Upstream Drift
Coding agents increasingly work across sessions, but prose notes can preserve a conclusion without the program state that supported it. Aft…
MIDAS: Multi-LLM Iterative Data-Adaptive Summarization
Text summarization is deceptively difficult. While condensing information seems straightforward, real-world enterprise summarization of sup…
Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)
Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evalu…
Efficient Online Lexicographic Generalized Low-Rank Matrix Bandits
This paper studies generalized low-rank matrix bandits with multiple prioritized objectives. At each round, the learner selects a matrix-va…
ATLAS: Adaptive Topological Learning with Abstract Successors for Continual Learning
Contemporary model-free reinforcement learning algorithms can achieve very high performance, but have low sample efficiency and are not rob…
COMPAS: Difficulty-Aware Joint Search for Optimizing Code Generation
Code generation systems make each LLM call with a model, a prompt, and decoding settings. However, existing optimization methods usually tu…
Equitable System-Prompt Selection via Constrained Mixed-Strategy GroupDRO
Large language models are increasingly used for information seeking, yet semantically equivalent questions phrased in different ways can re…
iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data
Multimodal learning of images and tabular data is often impaired by ineffective representations, resulting in redundancy, dispersion, and g…
HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models
Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion…
Combating Knowledge Corruption in Agent Systems: A Byzantine-Tolerant Secure Collaborative RAG Framework
While retrieval-augmented generation systems partially address the hallucination issues in large language models, it also introduces new vu…
FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation
Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institut…
Towards Trustworthy Hypergraph Neural Networks under Label Noise
Hypergraph neural networks (HGNNs) have demonstrated remarkable capabilities in processing complex higher-order relationships. However, the…
Image Classification Using CNN-QNN Hybrid Model with Optimized Correlated Features
We propose a method to optimize the correlation among convolutional neural network (CNN) features that are used as inputs to quantum neural…
NodeJEPA: Structure-Conditioned Latent Prediction for Node-Level Graph Self-Supervised Learning
Self-supervised learning on graphs is largely shaped by contrastive methods that depend on carefully designed augmentations, and by generat…
Approximate Multi-Objective Search Under Rulebooks
Robotic planning often involves multiple objectives with complex priority relationships, such as safety, efficiency, and regulatory complia…
Training-Free Hashing-Based Attention via Binary Principal Components
Long-context large language models (LLMs) are increasingly deployed in real-world applications, yet self-attention remains a major efficien…
MESH: Memory-Efficient Sinkhorn Optimization for Mixture-of-Experts Training
Memory-efficient matrix optimizers such as Sinkhorn gradient descent remove most AdamW optimizer state for dense Transformer matrices, but…
Not Every Divergence Should Be Suppressed: Counterfactual Recoverability in On-Policy Distillation
On-policy distillation (OPD) supervises student-visited trajectories, yet divergence-based rules cannot determine whether an erroneous pref…
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can ass…
Generative Optimization for Incentivized Advertising with Global Level Constraints
Incentivized advertising allocates monetary or virtual rewards to drive user engagement, where a key challenge is optimizing continuous inc…
MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages
We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asi…
ExeCRE: Execution-Consistency Guided Reliability Estimation for Self-Correcting Code Generation
Large language models (LLMs) have made notable progress in code generation, but they still struggle on challenging tasks that require sophi…
D$^2$F-ReAG: Dynamic Decomposition and Filtering for Multi-Hop Reasoning-Augmented Generation
Large language models (LLMs) often generate inaccurate answers due to their reliance on static internal knowledge. Retrieval-augmented gene…
When does training on downscaled images yield the same gradients?
Diffusion transformers deliver strong image generation, but their training cost grows superlinearly with resolution. Recent work justifies…
Q-CueGraph: Query-Conditioned Visual Evidence Graphs for Multimodal Reasoning
High-resolution pixels and crop or zoom tools give multimodal large language models the ability to inspect an image, but they do not provid…
TwinIR: Coordinated Invisible Dual-Point Attacks on Online HD Map Construction
Online HD map construction is critical to prediction and planning in autonomous driving. We find that existing physical attacks against onl…
Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning
As "AI Scientists" emerge to drive research via the Model Context Protocol (MCP), systems relying on ephemeral scripts will fail. The sheer…
Tropical Algebraic Geometry for Neuronal Representations: An Arakelov-Green Measure Based Descriptor for Graph Learning
The quantitative analysis of 3D neuronal morphologies requires capturing both graph topology and spatial geometry. Current message-passing…
Beyond Linear Dynamics: Neural Bilinear Dynamical Models for Time Series Forecasting
Time series in real-world applications are often generated by nonlinear dynamical systems, making accurate forecasting challenging. Existin…
EndoVLM: An Endoscopy Vision-Language Pre-training Model via Anatomy-Guided Sparsity and Progressive Alignment
The development of foundation models (FMs) is crucial for advancing endoscopic image analysis. However, existing endoscopy FMs mainly rely…
AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation
Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions…
AFD-Ledger: Deployment Provisioning for Attention--FFN Disaggregation
Attention--Feed-Forward Network (FFN) Disaggregation (AFD) is emerging as a promising architecture for serving Mixture-of-Experts (MoE) lan…
GeoReward: Mitigating Contextual Variable Overestimation in Vision-Language Models for Cross-Market Preference Prediction
Vision-language models excel in many multimodal tasks but remain prone to a subtle yet impactful failure mode: they tend to overestimate do…
GUARD: Grounding Uncertainty and Ablation-Based Risk Detection for Diffusion-Based VLAs
Diffusion-based vision-language-action (VLA) policies can generate plausible actions even when their predictions are weakly grounded in the…
CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding
Slice-based MLLMs leverage mature 2D encoders by representing 3D volumes as sequences of 2D slices. However, this slice-wise formulation pr…
A Model Merging Approach for Continual MLLM Unlearning
Multimodal large language model (MLLM) unlearning methods have been proposed to remove private, sensitive, or proprietary information from…
EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks
Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated o…
Breadcrumbing Search Agents
LLM-based search agents are widely used for information-seeking tasks, but their reliance on external tool returns introduces a critical se…
PhysMind: From Video to Executable Worlds for Training-Free Physical Reasoning
Reliable physical reasoning from video requires understanding how objects move, interact, and respond to interventions. Existing vision-lan…
Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders
Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing m…
EASy: Towards Efficient LLM-Based Agentic System
Agentic systems have emerged as a promising paradigm for solving complex tasks by coordinating specialized LLM-based agents. However, most…
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering
EgoCross is a cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can…
When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models
Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation lic…
Rethinking Reservoir Pruning: A Dynamical Perspective for Echo State Networks
Echo State Networks (ESNs) offer an efficient framework for temporal prediction, but their randomly initialized reservoirs are often over-p…
The Order Is the Guarantee: Verifier-Budgeted Code Deletion with Static-First Learned Proposals
Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply ma…
Masked diffusion enables coherent beat tracking
Current neural networks for beat tracking generate invalid outputs, such as consecutive downbeats and erratic tempo changes, even when thes…
DisMix: Order-Aware Mixup for Medical Imaging via Disentangling Ordinal and Non-Ordinal Features
Image mixup is a widely adopted data augmentation strategy, yet it is ill-suited for ordinal classification tasks such as medical disease g…
CSGen: A Multi-Domain Curvilinear Structure Generation Model via Hierarchical Multimodal Diffusion
Curvilinear structure analysis is an important and fundamental task in multimedia. However, the controllable generation of images with prec…
Easy to Complete, Hard to Choose: Investigating LLM Performance on the ProverbIT Benchmark
Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural languag…
Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports
Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a…
Personalized Federated Sparse Adaptation of Time-Series Foundation Models
Federated adaptation of time-series foundation models (TSFMs) is attractive for building energy forecasting because meter data are private,…
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO
We tackle the challenging yet underexplored task of Generalized Referring Expression Comprehension (GREC), which requires a model to locali…
Design Choices That Matter: A Functional ANOVA Analysis for Remote Sensing Multi-Label Classification
Benchmarking deep learning (DL) models for multi-label classification (MLC) of remote sensing images (RSI) typically yields rankings that d…
A 6G Integrated Sensing and Communication Framework for Railway Intrusion Detection and Collision Prediction
Integrated Sensing and Communication (ISAC) combines sensing and communication to efficiently utilize wireless resources and is emerging as…
What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend
Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Oll…
Toward Integrating Adaptive Experience Replay and Online Uncertainty Estimation in Safe Actor-Critic Optimal Control
Safe actor-critic control often treats barrier filtering, uncertainty estimation, and experience replay as separate modules, even though ea…
PURPOSE: Poisoning Conflict Resolution in RAG via Proxy-Fact-Grounded Updates
In Retrieval-Augmented Generation (RAG), post-retrieval conflict resolution arbitrates among noisy or contradictory retrieved passages. How…
InsightEmb: Learning Action-Intent Embeddings for Agentic Insight Retrieval
Self-improving agents accumulate reusable insights from prior trajectories, making retrieval increasingly important for turning accumulated…
Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models
Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic contro…
FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening
A large number of infants with congenital anomalies are born each year globally, especially in areas with underdeveloped medical resources.…
Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent
Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are…
IMFACT: Counterfactual Explanations for Time Series via Intrinsic Mode Function Substitution
Oscillatory signals, such as vibration, carry class-discriminative information in specific frequency bands; perturbing them in raw feature…
RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists
The integration of Large Language Models (LLMs) into software engineering has shifted the focus from function-level generation to repositor…
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation
Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited g…
Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First
Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model f…
Towards a satellite image manipulation and deepfake localization benchmark dataset
Verifying the authenticity of satellite imagery has become increasingly critical given advances in generative artificial intelligence. High…
A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination
Symbolic regression aims to discover closed-form equations from data, but existing LLM-guided methods often rely on a unified proposal loop…
When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
Multi-agent LLM systems relay key--value caches instead of text and credit their gains to exchanged ``latent thoughts''. That credit is a c…
A Chain Is Only as Strong as Its Weakest Link: A Scoping Review of System Integration Audits in AI
As AI systems become increasingly integrated into diverse interfaces and applications, model-centric audits are insufficient to address ris…
Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning
As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross-representation unders…
SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery
Bayesian causal discovery seeks to determine the posterior distribution of causal theories, which are interpreted as directed acyclic graph…
CheMLFlow: An Open-Source Platform for Cheminformatics and Materials Informatics Applications
CheMLFlow is an open-source platform for building and executing end-to-end, high-throughput, and agentic workflows for scientific and techn…
A General Sufficient Condition for Rewriting Horn-ALCHI Atomic Queries into GQL
The emergence of the ISO standard GQL introduces a powerful query language extending first-order logic with controlled recursion, raising t…
SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models
SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scie…
Protoreasoning in Tiny Transformers
We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study…
ORACLE: A Multi-Objective Reinforcement Learning-Based Analog Circuit Design Optimizer with Large Language Models-Guided Exploration
Analog circuit design automation using reinforcement learning (RL) has emerged as a promising approach for reducing manual effort. However,…
OneDayAgent: Towards a Long-Horizon Harness for Autonomous Agents
LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-en…
Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems
Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well…
ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation
High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due…
Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning
Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the…
The Effect of Perceived Race and Gender on Police Language Use: Experimental Evidence from VR Simulations
Against the backdrop of violence in police interactions with the U.S. public, we explore how deferentially police officers speak to virtual…
MarsCast: Transfer Learning of AI Weather Foundation Models to Planetary Atmospheres
We investigate the transferability of Earth weather foundation models to planetary atmospheres by adapting the GraphCast graph neural weath…
RepairFormer: Automated Repair of Structured Inputs Using Transformers
Structured input files such as JSON, DOT, OBJ, INI, S-expression, and TinyC are widely used in software systems, but small corruptions can…
Hardware Design and Security in the Era of Chiplets and LLMs
The semiconductor industry is undergoing a dual revolution: the shift toward heterogeneous 2.5D chiplet systems and the integration of Larg…
Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models
Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is n…
VQ-VAD: Vector-quantized Motion Representation Learning for Human-centric Video Anomaly Detection
Video Anomaly Detection (VAD) is inherently challenging due to the scarcity of anomalies and the large visual variability in surveillance f…
MultiPathFormer: Towards a Foundation Model for Multipath Wireless Propagation
Recent advances in machine learning have enabled training of wireless foundation models, which aim to support tasks such as channel estimat…
Capability-Gated Planning: Cost-to-Goal Discovery and the Limits of Myopic Experiment Selection
Systems that automate scientific discovery must repeatedly decide which experiment to run, which hypothesis to test, which tool to build, a…
Representational separation between unitary and channel quantum generative models via shared classical randomness at shallow depth
Near-term quantum hardware limits circuit depth and often imposes geometrically local connectivity for quantum generative models, restricti…
Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom i…
Chained Recursive Language Models for Multi-Iteration Reasoning
Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simulta…
SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant
Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vec…
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
On-Policy Self-Distillation (OPSD) has become a standard post-training approach for improving visual reasoning in multimodal large language…
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important fo…
GRALS: GCN-Guided Redundancy-Aware Local Search for Minimum Vertex Cover
The minimum vertex cover (MVC) problem seeks to identify the smallest set of vertices that cover all edges in an undirected graph. As a fun…
The Yokai Learning Environment: Tracking Beliefs Over Space and Time
The ability to cooperate with unknown partners is a central challenge in cooperative AI and widely studied in the form of zero-shot coordin…
Zero-shot reasoning for simulating scholarly peer-review
Scholarly publishing requires scalable scrutiny supported by auditable evidence. This paper presents a two-component benchmark of xPeer, th…
Corrigibility Transformation: Constructing Goals That Accept Updates
An AI agent will learn a desired goal more effectively if it does not resist the training process, but many partially learned goals incenti…
Calibrating Transformer Attention via Task-Space Sensitivity Feedback
Transformer-based pre-trained language models (PLMs) excel in text classification but suffer from attention dilution and attention sink eff…
XGrammar-2: Dynamic and Efficient Structured Generation Engine for Agentic LLMs
Modern LLM agents increasingly rely on dynamic structured generation, such as tool calling and response protocols. Unlike traditional struc…
MemFly: On-the-Fly Memory Optimization via Information Bottleneck
Long-term memory enables large language model agents to tackle complex tasks through historical interactions. However, existing frameworks…
Text2GraphQuery-Bench: A Text to Graph Query Benchmark
Graph models are fundamental to data analysis in domains rich with complex relationships. Unlike SQL, which benefits from a rel- atively un…
SimMOF: AI agent for Automated MOF Simulations
Metal-organic frameworks (MOFs) offer a vast design space, and as such, computational simulations play a critical role in predicting their…
Assessing and Explaining the Persuadability of Large Language Models as Legal Decision Tools
As Large Language Models (LLMs) are proposed as legal decision assistants, and even first-instance decision-makers, across a range of judic…
Contextual Agentic Memory is a Memo, Not True Memory
Current agentic memory systems (vector stores, retrieval-augmented generation, scratchpads, and context-window management) do not implement…
Online Goal Recognition using Path Signature and Dynamic Time Warping
Online goal recognition in continuous domains poses two central challenges: efficiently encoding large trajectories and effectively compari…
CogniFold: Always-On Proactive Memory via Cognitive Folding
Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into per…
Tree of Thoughts as a Classical Heuristic Search Problem: Formal Foundations and Design Patterns
Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, yet their standard generation process -- auto-regressive…
Trivium: Temporal Regret as a First-Class Objective for Causal-Memory Controllers
Many agentic systems and LLM pipelines correct mistakes by optimizing outcome reward. This addresses only the what of failure; the why and…
Necessary, Decodable and Reversible, Yet Not Transferable: A Stress Test for Attention-Head Role Claims
Mechanistic studies often assign a component a role when removing it damages a behavior, its activation linearly encodes task information,…
Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
Autoformalization translates informal natural language into formal, machine-verifiable languages. While most work focuses on individual sta…
Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitiv…
Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals
Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics ass…
MemSIF: From Structured Interactions to Dual-Track Fact Memory for LLM Agents
Long-term memory is critical for LLM agents operating over long-horizon interactions. However, several persistent limitations of existing m…
HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents
Retrieval-augmented search agents answer multi-hop questions by repeatedly issuing search queries and accumulating evidence. This creates a…
Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy
Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities,…
Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling
Recurrent linear attention models (RLAs) such as Mamba offer efficient linear-time sequence modeling as an alternative to Transformers, yet…
Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions
Language models have taken on the role of a very new type of technology, by virtue of their "human-ness" and rapid integration into users'…
On The Suitability of Differential Dataflow For Datalog Interpretation In Highly Dynamic Settings
In the domain of knowledge representation and reasoning within AI, datalog engines play an ever-increasingly crucial role. The crux of thei…
Large-Small Model Collaboration for Enhancing Edge-Deployed Small Models
Edge devices host domain-specific small language models (SLMs) with limited resources, while private clouds offer larger LLMs. We propose G…
Curiosity-Diffuser: Curiosity Guide Diffusion Models for Reliability
One of the bottlenecks in robotic intelligence is the instability of neural network models. This leads to risks when applying intelligence…
ZoomV: Temporal Zoom-in for Efficient Long Video Understanding
Long video understanding poses a fundamental challenge for large video-language models (LVLMs) due to the overwhelming number of frames and…
Review Text as a Leading Indicator of Displayed Reputation in Platform Rating Systems: Evidence from 34 U.S. Short-Term Rental Markets
Rating systems on accommodation platforms suffer from a familiar problem: nearly every listing displays a nearly perfect score, so the numb…
Pun Intended: Multi-Agent Translation of Wordplay with Contrastive Learning and Phonetic-Semantic Embeddings for CLEF JOKER 2025 Task 2
Translating wordplay across languages presents unique challenges that have long confounded both professional human translators and machine…
Emergence of Hierarchical Emotion Organization in Large Language Models
As large language models (LLMs) increasingly power conversational agents, understanding how they model users' emotional states is critical…
Uncertainty-aware Predict-Then-Optimize Framework for Equitable Post-Disaster Power Restoration
The increasing frequency of extreme weather events, such as hurricanes, highlights the urgent need for efficient and equitable power system…
Arnold: A multi-task, multi-embodiment muscle transformer policy
Controlling high-dimensional and nonlinear musculoskeletal models of the human body is a foundational scientific challenge. Recent machine…
Memorization in Large Language Models in Medicine: Prevalence, Characteristics, and Implications
Large Language Models (LLMs) have demonstrated significant potential in medicine, with many studies adapting them through continued pre-tra…
Dynamic Jailbreaking Attack
Existing gradient-based jailbreak attacks typically optimize a fixed-length adversarial suffix toward a predefined target response with a s…
RESample: A Robust Data Augmentation Framework via Exploratory Sampling for Robotic Manipulation
Vision-Language-Action (VLA) models have shown strong manipulation capability when trained with large-scale imitation learning datasets. Ho…
Bi-Level Reinforcement Learning Pathway for Sim-to-Real Optimality
Training Reinforcement Learning (RL) policies using simulation models before deployment in real-world environments is a common strategy whe…
When Large Language Models Know the Table: A Framework for Assessing Data Contamination in Tabular Datasets
Large language models (LLMs) are increasingly exposed to data contamination, i.e., performance gains driven by prior exposure of test datas…
Neural Diversity Regularizes Hallucinations in Language Models
Language models continue to hallucinate despite increases in parameters, compute, and data. We propose neural diversity -- decorrelated par…
Reinforcement Learning and Consumption-Savings Behavior
This paper demonstrates how reinforcement learning can explain two puzzling empirical patterns in household consumption behavior during eco…
MediRec: Enhancing Chinese Medication Recommendation with Explainable Clinical Reasoning
Large language models (LLMs) have shown strong potential for clinical decision support through their advanced language understanding and re…
FinRpt: Dataset, Evaluation System and LLM-based Multi-agent Framework for Equity Research Report Generation
While LLMs have shown great success in financial tasks like stock prediction and question answering, their application in fully automating…
Stabilizing Multi-Attack Adversarial Training via Bandit Optimization
Deep Neural Networks (DNNs) remain vulnerable to diverse adversarial perturbations, motivating multi-attack adversarial training (AT) for i…
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
Does Chain-of-Thought (CoT) reasoning genuinely improve Vision Language Action (VLA) models, or does it merely add overhead? Existing CoT-V…
Interpreting GFlowNets for Drug Discovery: What probes can and cannot show
Generative Flow Networks (GFlowNets) construct molecules through sequential decisions, but their internal policies remain opaque, limiting…
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
Vision-Language Models (VLMs) excel at reasoning in linguistic space but struggle with perceptual understanding that requires dense visual…
MODEST: Multi-Optics Depth-of-Field Stereo Dataset
Training and evaluation of state-of-the-art computer vision algorithms for reliable shallow depth of field (DoF) rendering and defocus debl…
Revisiting Generalization Across Difficulty Levels: It's Not So Easy
We investigate how well large language models (LLMs) generalize across different task difficulties, a key question for effective data curat…
Feedback Loops and Code Perturbations in LLM-based Software Engineering: A Case Study on a C-to-Rust Translation System
The advent of strong generative AI has a considerable impact on various software engineering tasks such as code repair, test generation, or…
Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Versatile Image Generation
Reinforcement learning (RL) has emerged as a powerful paradigm for fine-tuning large-scale generative models, such as diffusion and flow mo…
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deploym…
A Survey of Agent Memory in the Second Half: Towards Self-Evolving and Long-Horizon Agents
Research in artificial intelligence is shifting from model innovations and benchmark scores towards problem definition and rigorous real-wo…
Can Post-Training Transform LLMs into Causal Reasoners?
Causal inference is essential for decision-making but remains challenging for non-experts. While large language models (LLMs) show promise…
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance charact…
AdaCorrection: Adaptive Offset Cache Correction for Accurate Diffusion Transformers
Diffusion Transformers (DiTs) achieve state-of-the-art performance in high-fidelity image and video generation but suffer from expensive in…
Formal Analysis and Supply Chain Security for Agentic AI Skills
32 pages, 5 theorems with full proofs, 68 references, open-source tool: https://github.com/qualixar/skillfortify. v2: corrects the bibliogr…
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
Video object insertion is fundamental to video editing, yet existing diffusion methods often produce visually plausible but physically inco…
Seeking Physics in Diffusion Noise
Do video diffusion models encode signals predictive of physical plausibility? We probe intermediate denoising representations of pretrained…
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
Multi-agent discussions have been widely adopted, motivating growing efforts to develop attacks that expose their vulnerabilities. In this…
The Luna Bound Propagator for Formal Analysis of Neural Networks
The parameterized CROWN analysis, a.k.a., alpha-CROWN has emerged as a practically successful abstract interpretation method for neural net…
SleepVLM: A Rule-Grounded Vision-Language Model for Auditable Sleep Staging
Sleep staging is essential for sleep assessment and disorder diagnosis. In recent years, automatic sleep staging systems have achieved accu…
Terminal Agents Suffice for Enterprise Automation
There has been growing interest in building agents that can interact with digital platforms to execute meaningful enterprise tasks autonomo…
MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding
With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. A…
Generative Experiences for Digital Mental Health Interventions: Evidence from a Randomized Study
Digital mental health (DMH) tools have extensively explored personalization of interventions to users' needs and contexts. However, this pe…
Multi-Modal Learning meets Genetic Programming: Analyzing Alignment in Latent Space Optimization
Symbolic regression (SR) aims to discover mathematical expressions from data, a task traditionally tackled using Genetic Programming (GP) t…
Topology-Aware Reasoning over Incomplete Knowledge Graph with Graph-Based Soft Prompting
Large Language Models (LLMs) have shown remarkable capabilities across various tasks but remain prone to hallucinations in knowledge-intens…
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
Evaluating LLMs is challenging, as benchmark scores often fail to capture models' real-world usefulness. Instead, users often rely on ``vib…
Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual inform…
Just Repair: A Minimal Denoising Network for Time Series Anomaly Detection
Time series anomaly detectors have grown steadily more complex, incorporating attention mechanisms, adversarial training, and stochastic la…
A Systematic Review and Taxonomy of Reinforcement Learning-Model Predictive Control Integration for Linear Systems
The integration of Model Predictive Control (MPC) and Reinforcement Learning (RL) has emerged as a promising paradigm for constrained decis…
DeepImagine: Clinical Trial Outcome Prediction via Stepwise Local Counterfactual Imaginations
Predicting the outcomes of prospective clinical trials remains a major challenge. Clinical trial outcomes result from complex interactions…
IConFace: Fine-Grained Identity Conditioning for Reference-Aware Face Restoration
Severe face degradation can remove person-specific evidence, making restoration underdetermined. A generative prior may recover a sharp, pl…
CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research
We present the Curated Industrial Developer Repository (CIDR), a large-scale dataset of real-world software repositories collected from ind…
Stable Attention Response for Reliable Precipitation Nowcasting
Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamic…
Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution
Integrated Gradients (IG) is a widely adopted feature attribution method that satisfies desirable axiomatic properties. However, the choice…
Distributionally Robust Transfer Learning with Structurally Missing Covariates, with Application to Cross-National Cardiac Arrest Prediction
Deploying clinical prediction models across healthcare systems often fails when key training covariates are unavailable at deployment and l…
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-…
The Hamilton-Jacobi Theory of Deep Learning
In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient…
An Enhanced Geometric-Spectral Feature Learning Framework for Airborne Multispectral Point Cloud Classification
Multispectral point cloud (MPC) is composed of 3D spatial-spectral information, which holds tremendous potential for accurate land-cover cl…
AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models
Recent advances in pretrained large audio-language models (LALMs) have demonstrated strong capabilities across speech, sound, and music. To…
Human-in-the-Loop Atlas-Based 3D Asset Segmentation for Interactive Content Workflows
Segmenting 3D assets into meaningful regions remains challenging, especially when segmentation criteria are application-dependent and requi…
Delta-Diffusion: Modeling Longitudinal Brain Amyloid-PET Trajectories via Conditional Poisson Diffusion Bridge
While longitudinal brain PET imaging is the gold standard for quantifying the spatiotemporal accumulation of Beta-amyloid, its widespread c…
War in the Abstract: The Rise and Consequences of Militarized Language in Scientific Communication
Scientists do not, by profession, wage war. Yet warfare's vocabulary consistently appears in their abstracts. To quantify the extent to whi…
Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment
Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting among multiple candidates with an automat…
Amplitude-Only FFN Intervention for Tool-Structured LLM Inference Method: Gated Evaluation Protocol, and Cross-Model Empirical Results
Large language models increasingly operate as tool-using agents, where small format, argument, or function-call errors can invalidate other…
TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale
Training robust autonomous driving agents requires a simulator fast enough for reinforcement learning at scale, realistic enough to ground…
Decision Making Needs Uncertainty Quantification [Lecture Notes]
Many signal processing systems ultimately exist to {act}. Whenever the state variable that determines the action to be taken by a decision…
Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
Synthetic tabular data are valued for preserving not just column-wise marginals but inter-column dependency, which carries much of the mino…
SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving
As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment…
CORF-GS: Real-Time Wireless Radiance Field Reconstruction via Coupled Optical-RF Gaussian Splatting
Recent advances in 3D Gaussian Splatting (3DGS)-based wireless radiance field (WRF) reconstruction provide an efficient solution for wirele…
ORCA-bench: How Ready Are Language Model Agents for Oncall?
Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over n…
It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling
Test-time scaling lifts large language model reasoning by sampling many candidate solutions and selecting among them, yet the same recipe t…
PICopilot: An LLM-based Agentic Framework for Assisting Photonic Integrated Circuit Design via Script Generation
The rapid development of photonic integrated circuits (PICs) is shifting the design flow from traditional graphical user interface (GUI)-ba…
CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship
LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scen…
GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
A wearable assistant should both answer questions about its visual history and recognize when that history is useful to the present situati…
MythosとGPT-5.6 Solが性能テスト中に暴走 OSSメンテナーに圧力、有害コード実行図る 英政府機関
英政府機関AISIによるAIモデルのサイバー能力評価中に、AnthropicのMythos 5やOpenAIのGPT-5.6 Solが実在の人や組織を標的に暴走。偽アカウントを作り、OSSメンテナーに悪意あるコードの承認を迫っていた。
Meta、コーディングエージェント「Muse Code」リリース 「Claude Code」や「Codex」に対抗
Metaは、ターミナル上で動作するコーディングエージェント「Muse Code」のベータ版と、新モデル「Muse Spark 1.2」を発表した。非同期バックグラウンドエージェントの常駐により遅延を抑え、複雑な開発タスクをサポートする。自己改善ループ等でコーディング性能を強化し…
ワークマン、実は画像生成AIを導入していた 間に合わない商品撮影を代替 アプリ通知開封も1.5倍
ワークマンの画像生成AI活用法を“中の人”に聞く。
オープンソースのAI機能付きオフィス登場 docx・xlsx・pptx・pdfに対応 Genspark
AIサービスを開発する米Gensparkは、オープンソースのAI機能付きオフィススイート「GenOffice」を発表した。
フィジカルAIに注力するインテル、組み込み市場での約40年の実績を生かせるか
「インテル・ロボティクス・ワークショップ2026」のレポート記事をお送りする。今回の前編では、インテルのロボティクス/フィジカルAI戦略に関する基調講演や、ソフトウェアベースのモーションコントローラーを提供しているモベンシスの取り組み事例などについて紹介する。
Googleのジェフ・ディーン氏、独立してAI実験を大規模自動化する新会社Discovery Loop設立
Googleのチーフサイエンティスト、ジェフ・ディーン氏が退社し、新会社「Discovery Loop」を設立すると発表した。サンジェイ・ゲマワット氏らと共同創業する公益法人で、フロンティアAIモデルと計算基盤を活用して実験や検証などの研究プロセス全体を自動化することを目指す。…
Google DeepMindのデミス・ハサビスCEOが退任して会長に チーフサイエンティストのジェフ・ディーン氏は独立へ
Google DeepMindのCEOを務めてきたデミス・ハサビス氏が業務執行から退き、DeepMind会長兼Alphabetチーフサイエンティストに就任する。後任にはコーレイ・カブクチュオール氏が就く。また、チーフサイエンティストのジェフ・ディーン氏とGoogleシニアフェロ…
KDDIも40年前はスタートアップだった 高橋会長の“昔話”から考える、日本企業の成長戦略
KDDIはイケイケの人たちが作った――KDDIの高橋誠会長は、こう振り返る。“スタートアップ精神”で成長してきた同社から、企業の成長戦略を考える。
資金調達の「二極化」進む 勝ち抜く企業の条件は? 早大教授に聞く
「AIによる審査の高度化」と「資金調達の多様化」は必ずしも同じ話ではない──ベンチャーファイナンスの第一人者である早稲田大学ビジネススクール(経営管理研究科)の長谷川博和教授は、こう指摘する。
Meta launches Muse Code, an AI agent for large code bases
Meta expanded its AI coding offerings with a new agent that, it promises, can handle complex tasks with complex software.
Klaviyo acquires Elias Torres’ Agency in full-circle reunion for tech founders
The serial entrepreneur joins the e-commerce company as CPO to lead its AI agents.
Google DeepMindが描く「AGIの次」 超知能に至る4つの経路と6つのボトルネック
Google DeepMindが、人間並みのAGIが実現した後にAIはどこまで進むのかを整理した論文を公開した。ASI(人工超知能)への4つの経路と6つのボトルネックという枠組みだけを抜き出し、初心者にも分かるように短く読み解く。
Jeff Dean and other top AI researchers are leaving Google to launch their own startup
The legendary Google executive is joined by other outgoing Google execs in a joint mission to use AI to push forward the process of scienti…
Shopify says AI search is driving more traffic and sales, not replacing Google
Shopify says AI isn’t cannibalizing search traffic the way it has for publishers. Instead, AI-driven traffic and orders to Shopify stores t…
Hark previews its browser use agent for completing tasks
Hark claims that its browser use agent is faster and cheaper than competition.
TechCrunch Disrupt 2026’s Real World AI Stage features robots, automated factories, and extinct animals
On our new Real World AI stage, we’ll be focusing on the intersection between the digital and physical, and all the ways we’ll continue to…
2026-08-05(431件)
Anthropic is hiring an AI chip design team
Anthropic is building a team for designing its own custom AI chips. The Claude maker said it would co-design hardware and models to help it…
MacPaw taps Liquid AI to offer on-device inference to devs building for its app store
MacPaw is building a local version of its AI assistant Eney using Liquid AI's models.
AI makes weather prediction better. Can WindBorne make it lucrative?
WindBorne Systems has raised a $37 million Series B round to scale its weather balloons and AI forecasts.
「準国産」うたう人型ロボ登場 本体価格は500万円から 目指すは「純国産」
AIサービスなどを開発するZEALSは、「準国産」をうたう人型ロボット「D1」を発表した。本体価格は500万円(税別)から。
AIで"存在しない脆弱性"を量産? 「SQLite」の偽CVEが判明 米企業が検証
SQLiteの脆弱性を主張するCVEは実在しなかったと米JFrogが検証結果を公開。AI生成とみられる偽情報がNVDなどに大量登録されていた。同一アカウントが投稿した55件中54件が捏造で、CVE登録プロセスの穴を指摘した。
エージェント型AIは「過度な期待」のピークへ ガートナー、日本におけるデジタルワークプレースのハイプサイクルを発表
ガートナージャパンが日本における未来のデジタル・ワークプレースのハイプ・サイクルを発表。エージェント型AIを「過度な期待」のピークに位置付け、AIエージェントの浸透やシャドーAIのリスク拡大などのトレンドを紹介した。
自社のAI利用は本当に安全? 現場放任に潜む「4大リスク」
サーバーワークスは、AWS環境での生成AI活用に必要なルール整備を支援する「AWS生成AIガイドライン策定サービス」を提供開始した。最短1カ月でルールを整備できるプランも用意する。
ISEE: Interactive Semantic Enrichment for Database Fields
LLM-based agents are increasingly being deployed for data-related tasks, including data sense-making, exploration, and retrieval. However,…
Self-Organising Digital Circuits
Fault tolerance in classical computing has traditionally relied on static strategies like hardware redundancy and error-correcting codes. B…
Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling
Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, ho…
PULSE: An Executable Contract Language for Spatiotemporal Knowledge Graph Engineering
Knowledge graph engineering often distributes accepted state, observations, constraints, processes, and hypothetical scenarios across artif…
HyperAgent: Planning and Acting over Tool-Schema Hypergraphs for Tool-Use LLM Agents
Large language model (LLM) agents increasingly rely on external tools to complete complex real-world tasks. However, reliable tool-use plan…
Explainable AI for the EU Right to Explanation: A Systematic Review of the Law-XAI Translation Gap
When algorithms make or influence consequential decisions---about loan eligibility, hiring, or healthcare---EU law grants affected individu…
Predictive Set Theory: A Generative Framework for Cognitive Architecture with Operationalized Core Mechanisms
Predictive processing theories portray the brain as a hierarchical prediction engine that minimizes prediction error, yet they lack operati…
Towards a new paradigm of scientific discovery with socialized artificial intelligence
Scientific discovery has advanced through successive transformations in the organization of knowledge. Observation and experimentation esta…
BAP-SQL: Budget-Aware Observation Planning for Agentic Text-to-SQL
Tool-using agents do not merely consume observations: their actions determine what arrives next. In agentic text-to-SQL, a broad query can…
VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space
Large language models have shown promise for automated Verilog RTL generation, yet state-of-the-art multi-agent systems plateau at ~95% acc…
Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes
The widespread adoption of proprietary Large Language Models (LLMs) accessed strictly through closed APIs has created a critical challenge…
Hypercubes, Hyperplanes, and Constraint-Induced Complexity Collapse in Atomic Concept Learning
We revisit higher-arity atomic concept learning through the geometry of hypercubes and hyperplanes of ground instances. Our starting point…
When Compression Scores Cannot Decide: Information Boundaries for Group-Robust LLM Pruning
A reproducible compression statistic can still select the wrong candidate. A dense pruning score with 0.906 split-half reliability predicte…
On the missing data layer and a potential solution
Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the da…
Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning
(Flat) Reinforcement Learning (RL) agents face significant challenges in environments with sparse rewards that require long-horizon reasoni…
On the missing benchmarks layer and a potential solution
Latin America is missing a foundational layer for native AI development: the benchmark layer. The benchmark layer does two things no other…
ProPRL: Property-Aware Prerequisite Relation Learning in Educational Knowledge Graphs
Prerequisite relation learning is central to adaptive instruction, yet existing methods often formulate it as conventional link prediction,…
UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks
Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services…
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation throug…
DiffImaginE: Imagine to Verify Entity Types with Diffusio
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual…
Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning
Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medica…
CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting
Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical o…
TraceCAD: Trace-Guided Repair for Agentic CAD Generation
LLM-based CAD agents produce executable parametric programs, but their correction loops may lose evidence about satisfied requirements, fau…
Getting the Parameters Right: A Difficulty-Graded Benchmark and Probe-Guided Training for LLM Tool Calls
Large language model agents derive much of their capability from tool use. Existing research on tool use has largely focused on selecting t…
AI Agent Economics: Can Autonomous Economic Behavior Emerge among AI Agents under Minimal External Conditions?
Multi-agent studies commonly place AI agents in predefined games, markets, or roles, making it difficult to distinguish endogenous economic…
Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting sc…
Beyond Average Performance: Dynamic Instance Clustering and Specialized Algorithm Design in LLM-Assisted Evolutionary Search
Large Language Model-assisted Evolutionary Search (LES) has emerged as a powerful paradigm for automated algorithm design. However, existin…
Verifiable Memory: Learning Unified Memory Management with Local and Global Verifiers for Large Language Model Agents
Large language model (LLM) agents must retain reusable information, control a bounded active context, and recover earlier evidence during l…
Spatial proteomics guided by H&E-based AI reveals recurrence-risk niches in triple-negative breast cancer
Deep learning models can predict cancer recurrence from H&E stained slides, but the localized molecular states underlying these predictions…
UniGD: A Unified Generative-Discriminative Framework for Industrial Retrieval
Generative retrieval (GR) is a promising paradigm for industrial search advertising, yet its deployment is constrained by strict relevance…
Evidence-Grounded Multimodal Knowledge Graph Construction for Multi-Lecture Educational Reasoning
Lecture videos distribute knowledge across speech, slide text, diagrams, equations, and presentation order, which transcript-only retrieval…
Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support,…
Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes…
Diversity is Not Ambiguity: Toward Accurate and Efficient Ambiguity Detection for Open-Domain QA
How can question answering (QA) systems determine whether a query is ambiguous? Ambiguity detection is essential in open-domain QA, as misc…
TumorBoard: Evidence-Grounded Multi-Agent Decision Support for Longitudinal Neuro-Oncology
Neuro-oncology decisions require coordinated interpretation of serial MRI, pathology, molecular markers, treatment history, performance sta…
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pai…
The Agent Operating System (AOS): A Reference Operating Architecture for Distributed Agentic Systems
Large language models have transformed artificial intelligence from isolated prediction services into components of long-running, distribut…
Reachability Is Not Realization: Tracing the Sources of LLM Benchmark Gains
Benchmark gains are often treated as evidence of greater LLM capability. Yet the same gain can reflect different changes in model behavior.…
UniNav: A Unified World-Action Diffusion Model for Visual Navigation
Image-goal visual navigation is a fundamental capability for embodied agents. Existing navigation policies efficiently predict waypoint tra…
One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning
Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance.…
TaskPress: Query-Agnostic KV Cache Compression via Task-Guided Pruning
Long-context inference with large language models is constrained by the linear growth of the key-value cache to sequence length. While prun…
AgentPanel: Toward a New Paradigm for Human--AI Collaboration in Exploring Scientific Questions
Identifying promising scientific ideas remains an important challenge in research practice. Researchers commonly rely on small-group discus…
DocTrace: Towards Traceable Long Document VQA via Hierarchical Evidence Graph Reasoning
Long Document Visual Question Answering (LongDocVQA) requires Multimodal Large Language Models (MLLMs) to locate, integrate, and reason ove…
Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks
A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the r…
SeaSlides: Semantic Abstraction Layer for Agentic Slide Generation
Agentic presentation generation must preserve source content, maintain coherent visual design, render specialized objects, and produce usab…
Screenshots or Tools? Eliciting Tool Use and Managing Multimodal Context in Hybrid GUI-MCP Computer-Use Agents
Hybrid computer-use agents can act through screenshots or call text tools. We find that having a tool available does not settle which way t…
Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving
This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally e…
Traceable Multi-Agent System for Knowledge-Based Forecasting
Enterprise forecasting increasingly relies on autonomous agents that interpret documents, search for data, generate code, and revise models…
MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc
MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the q…
Towards Robust Tool Use in Agents via Experience-Driven Adaptive Guidance
The performance bottleneck of agents is increasingly shifting from model capability to the robustness of their execution processes. Tools p…
Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems
As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text…
AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction
Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models rece…
Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory
Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is le…
State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
Transformer-based architectures have dominated sequence modeling, largely due to the expressive power of attention mechanisms. However, for…
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, st…
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Exp…
Solver-Aware Decompositions for Programming-by-Example: When Dividing Requires Knowing how to Conquer
Decomposition-based Programming-by-example (PBE) scales performance by splitting tasks into subtasks that a learned synthesizer solves: a d…
LeanMem: Simple and Efficient Long-Term Memory for LLM Agents
Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory…
ChartAnno: Evaluating MLLMs for Chart Annotation Generation
Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability…
When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO
Reinforcement learning with verifiable rewards (RLVR) com- monly optimizes each correct completion as an independent learning signal. In GR…
ToolLIFT: Lifting Tool-Specific Trajectories into Function-Level Graphs for Generalizable Tool Planning
Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage. Exist…
WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent Networks
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user ca…
Can LLM design high-quality experiments? A Comprehensive and Systematic Benchmark on Autonomous Experimental Design
AI for Research (AI4Research) leverages AI to automate and improve scientific workflows. While experimental design is a critical stage of t…
Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks
Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous…
When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs
Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasonin…
Reversing Arrows in Large Language Models
Large language models (LLMs) have achieved strong performance on text-to-knowledge graph generation and related tasks. Nevertheless, it is…
Dr. AGENTONOMICS: A Didactic Experiment of AGENTONOMICS
AGENTONOMICS is a framework that treats AI agents as economic entities that can be designed, managed, and governed through an integrated ma…
Behaviorally Adaptive Visual Diversion for Inclusive and Resilient Digital Assessment Delivery
Institutions increasingly rely on browser lockdown, webcam monitoring, and behavioral analytics to secure high-stakes digital assessments,…
Soft Guidance Starts to Outperform CoT Prompting as LLMs Improve
Chain-of-Thought (CoT) prompting remains the standard baseline for evaluating models' reasoning abilities. Originally, this technique was i…
Enhancing Tabular Learners with Context-Aware Semantic Embeddings
While modern tabular learners excel at capturing statistical patterns, they frequently operate in a semantic vacuum, treating textual featu…
Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made pro…
Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Universities and Business Schools
Artificial intelligence (AI) is rapidly transforming high-skilled domains, requiring higher education institutions (HEI) to balance the tea…
From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities
Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowled…
FOUND-AF: Benchmarking ECG Foundation Models for Atrial Fibrillation Detection
Atrial fibrillation (AF) is the most common sustained cardiac arrhythmia and is associated with increased risks of stroke, heart failure, a…
Large language models for partial differential equation workflows
Partial differential equations (PDEs) become actionable in science and engineering not as isolated formulae, but as executable workflows th…
FraQ: Efficient Coordinate-Space Recompression for Federated Low-Rank Adaptation
Federated fine-tuning with Low-Rank Adaptation (LoRA) enables efficient collaborative adaptation of Large Language Models (LLMs) without ce…
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneo…
Formal Verification of Agentic Systems over Operational Data
Agentic systems driven by large language models (LLMs) are increasingly deployed in real-world workflows where they act on persistent opera…
Rethinking Modality Reliability in Multimodal Sentiment Analysis with Incomplete Observations
Multimodal Sentiment Analysis (MSA) integrates text, audio, and vision to infer human affect, yet real-world multimodal observations are of…
Unequal Verdicts: Investigating Gender Bias in LLM-Based Fake News Detection
Large Language Models (LLMs) are increasingly used for automated fact-checking, yet their susceptibility to gender bias in this context rem…
Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model
A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computati…
When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
On-Policy distillation (OPD) transfers teacher capabilities by supervising student-sampled trajectories with dense token-level teacher sign…
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details
AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before.…
AutoSND: From Execution Evidence to Structural Policies for Automated Network Dismantling Heuristic Discovery
Network dismantling is fundamental to analyzing the robustness and vulnerability of complex systems, yet practical heuristics must balance…
Taming the Implicit: Dual-Channel Risk-Aware Reinforcement Fine-Tuning for Continual Multimodal Post-Training
Reinforcement fine-tuning (RFT) is widely believed to inherently resist catastrophic forgetting in continual post-training of multimodal la…
Shielding for Higher-Order Safety
Safety shields are runtime enforcement mechanisms that restrict the actions of a controller to guarantee safety. Classical shields are usua…
PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud
Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge G…
LiveEvalBench: Toward Open-World Evaluation for Web Generation
Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web genera…
TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents
Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning.…
Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges
Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a sma…
When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives
Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group shoul…
SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite Intelligence
Onboard satellite intelligence requires a task layer that translates mission intent into local tool calls, exposes execution state, and ret…
CARE-Bench: Benchmarking Patient-Facing LLM Triage
Patient-facing medical LLMs and agents increasingly answer symptom questions before clinician contact, where the key safety question is wha…
Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement
Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heav…
AgenticECO: An Agentic Framework for ECO on 3D Integrated Circuits
As Moore's law slows, the industry is turning to three-dimensional integration; yet in merged 3D-IC flows, routed designs expose bond-level…
MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models
Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mappe…
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such comm…
Risky Business: Measuring The Faithfulness-Safety Tension
Chain-of-Thought (CoT) reasoning offers a promising window into model monitoring. However, monitoring relies on faithfulness, i.e., the mod…
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks
Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evalu…
Computing Actual Causes for Neural Network Predictions under Structured Causal Inputs
Explaining the predictions of neural networks is a central challenge in trustworthy AI. Existing explanation methods, such as those based o…
KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation
Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks m…
Does Forgetting Transfer Across Modalities? A Real-World Benchmark for Cross-Modal Knowledge Unlearning Evaluation
Vision-Language Models (VLMs), like Large Language Models (LLMs), may memorize sensitive, copyrighted, or harmful knowledge from their pret…
LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards
Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Al…
Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes
Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses…
MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents
Memory-augmented LLM agents rely on rich context for long-horizon reasoning and acting, yet their memory modules expose a persistent attack…
ADMITBench: A Safety-Governed Reference Framework for Evaluating the Admissibility of Industrial LLM Advisories
This white paper presents ADMITBench, a reference framework for evaluating industrial LLM advisories at the level of the proposed action. T…
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear wheth…
Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition
We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related p…
When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking
Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy…
Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory
As AI systems are deployed across increasingly diverse social contexts, alignment can no longer be framed as the optimization of a single,…
Implementing Causal Perception: Competing SCMs and Situated Fairness
Causal perception occurs when agents with competing Structural Causal Models (SCMs) of the same system infer different probability distribu…
The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections
This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models…
TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring
Large language models (LLMs) are increasingly used to provide conversational practice for English-as-a-second-language (ESL) learners. Effe…
A game theory for foundation models shows new paths to rational cooperation through similarity inference
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principle…
Interpretable Adaptive Sampling for LLM Test-Time Scaling
Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query b…
Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards;…
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and…
Multi-Camera Trajectory Forecasting with Trajectory Tensors
We introduce the problem of multi-camera trajectory forecasting (MCTF), which involves predicting the trajectory of a moving object across…
CLIP-EBC: CLIP Can Count Accurately through Enhanced Blockwise Classification
We propose CLIP-EBC, the first fully CLIP-based model for accurate crowd density estimation. While the CLIP model has demonstrated remarkab…
UL-UNAS: Ultra-Lightweight U-Nets for Real-Time Speech Enhancement via Network Architecture Search
Lightweight models are essential for real-time speech enhancement applications. In recent years, there has been a growing trend toward deve…
Assessing speech quality metrics for evaluation of neural audio codecs under clean speech conditions
Objective speech-quality metrics are widely used to assess codec performance. However, for neural codecs, it is often unclear which metrics…
PASE: Leveraging the Phonological Prior of WavLM for Low-Hallucination Generative Speech Enhancement
Generative models have shown remarkable performance in speech enhancement (SE), achieving superior perceptual quality over traditional disc…
StuPASE: Towards Low-Hallucination Studio-Quality Generative Speech Enhancement
Achieving high perceptual quality without hallucination remains a challenge in generative speech enhancement (SE). A representative approac…
GAP-URGENet: A Generative-Predictive Fusion Framework for Universal Speech Enhancement
We introduce GAP-URGENet, a generative-predictive fusion framework developed for Track 1 of the ICASSP 2026 URGENT Challenge. The system in…
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations
Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates. We propose UniPA…
KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization
Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measureme…
MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale
Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing m…
OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning
Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata…
Evaluating OpenAI's Privacy Filter: Cross-Lingual, Cross-Domain PII Detection Across 42 Benchmarks
We present the first independent, systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.5B-parameter bidirectional PII detector, acro…
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation us…
Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks
Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as…
Speculative Correction: Draft-then-Refine Decoding for Diffusion Language Models
Diffusion language models (DLMs) can revise tokens bidirectionally, but standard decoding procedures often adapt them to left-to-right gene…
Deep Divide-and-Reduce in Symbolic Regression
Symbolic regression (SR) is the task of discovering underlying patterns from data and representing them using mathematical expressions. Cur…
Multimodal Auto-regressive Transformer Surrogate for Modeling Variable Operations and Quantifying Uncertainty in Geological Carbon Storage
The use of variable well perforation and injection strategies can improve the efficiency of geological carbon storage operations. We develo…
Rethinking Self-Evolving Agent Skills: Feedback Dynamics over Multiple Rounds
Self-evolving skill systems promise to improve agents by turning execution feedback into persistent skill updates without changing the unde…
Studying, Identifying, and Fixing Hidden Technical Debt in AI-Intensive Cyber-Physical Systems
Artificial Intelligence (AI) components are increasingly pervasive in several software systems, including Cyber-Physical Systems (CPSs). AI…
Instruction Stacking Collapse: A Benchmark and the Capability-Dependent Value of Prompt Compilation
Production prompts rarely carry a single instruction. One system message may require valid JSON, a word limit, three citations, and a fixed…
IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation
Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generatio…
MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows
Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out…
CUADebug: Diagnosing and Repairing Computer-Use Agent Failures
Computer-use agents (CUAs) operate real desktop and web interfaces through screenshots, mouse and keyboard actions, and stateful UI feedbac…
Verified Tool Calls Improve LLM Agent Reliability Under Non-Atomic Failures
Large Language Model (LLM) agents rely on external tools to perform multistage tasks. Existing agent frameworks typically assume that tool…
Cross-Anesthetic ECoG State Decoding Fails at the Decision Threshold, Not the Representation
Decoders of anesthetic state from cortical activity fail across drug classes, most notoriously ketamine, but reported accuracy cannot say w…
Secure AI Watermarking Framework for IP Protection in Multi-Tenant Cloud Platforms
The Secured data safe guard transaction with multi-tenant environments run on private-protected authenticate platforms runs by secured hand…
Your Agentic LLMs Secretly Encode Latent Signals of Indirect Prompt-Injection Exposure
Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While m…
AI Alignment and Fiduciary Obligation
Advanced AI assistants engage users in extended interactions across a widening range of roles, including advice, decision support, collabor…
Verifier-Guided Model Discovery for Physical Dynamical Systems with Pretrained Symbolic Transformers
Reliable forecasting of nonlinear physical systems underpins scientific discovery and engineering decision-making. Yet high-fidelity simula…
CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study
Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence mo…
ZK-SR117: A Chunked Zero-Knowledge Attestation Design for Aggregated Fair-Lending Metrics, with a Control Mapping toward Full SR 11-7 Coverage
Deploying ML models in regulated decision-making (credit underwriting, fraud detection, loan approval) requires demonstrating fairness and…
Single Canonical Prompts Underestimate LLM Safety's Surface-Form Sensitivity
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that r…
Sphere Retraction Normalizations
Residual connections are the de facto mechanism for training deep neural networks stably. Geodesic Normalization (GeoNorm) recasts them on…
Vulnerabilities, Secrets and Misconfiguration in the Highest-Exposure Docker Hub Images
Docker Hub is the registry underneath most container deployments, and a flaw in a widely reused base image is inherited by every image buil…
Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments
Coding agents increasingly run inside organizations whose security controls (scoped credentials, restricted egress, read-only filesystems,…
Security-First Evaluation of Text-to-Terraform: Benchmarking LLMs and SLMs for Secure IaC Generation
Cloud misconfiguration remains a leading cause of security incidents, yet whether LLMs and SLMs can generate security-compliant Infrastruct…
dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural…
Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks
With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreak…
When Policies Change Probabilities: Modular Decision-Making for LLM Code Review
LLM code reviewers often estimate patch risk and make approval decisions in one prompt. A probability should depend on evidence; costs shou…
DenialRAG: Single-Document RAG Poisoning via Embedded Parametric Denial
Retrieval-augmented generation (RAG) systems are vulnerable to corpus poisoning: an attacker who inserts a crafted document into the retrie…
AI Sandbox: Technical Report
Collaborative AI experimentation across industry and academia requires platforms that enable rapid prototyping while preserving controlled…
TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows
Tool-using language-model agents repeatedly rediscover procedures they have already executed, producing traces that mix reusable structure…
$S^3$: Improving Agent Safety through Multi-Stage Defense
Large Language Model (LLM) agents rely on multi-stage agentic workflows, with stages such as memory, planning, and tool execution, to accom…
A Blind Spot in Alignment: Quantifying Biosecurity Risks in Large Language Models
Large Language Models (LLMs) are accelerating biological research, yet this same capability poses a critical biosecurity threat: models tha…
BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests
Coding-agent benchmarks increasingly cover long-horizon, end-to-end, and interactive development, but typically retain one requested outcom…
Learning Molecular Representations from Cellular Phenotypes with Structure Preservation
Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. However, e…
Output-Aware Rotation for INT2 KV-Cache Quantization
The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-lo…
Measuring Explainer Stability via Attribution Separability
Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most meth…
Steganalysis of Adaptive Covert Collusion in Tool-Using Agent Populations: A Black-Box, Cross-Principal Approach
Tool-using agents built on large language models (LLMs) are increasingly deployed not by a single operator but by many, side by side on sha…
NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory
Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade…
Can Training Logs Make Model Comparisons More Precise?
Comparing stochastically trained models requires estimating both a performance difference and its uncertainty from repeated runs. We study…
Designing a Good Virtual Node: Addressable and Cardinality-Preserving Global Memory for Message Passing Architectures
Virtual nodes give message-passing neural networks a simple global communication route, but the standard node--VN--node pipeline compresses…
Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators
Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-o…
Quo Vadis, World Modeling?
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is cos…
Search, Inspect, Fetch: Exploiting Boolean Retrieval for Deep-Research Agents
Existing deep-research agents use a search-visit workflow that retrieves and reads whole pages, without considering the addressable structu…
Privacy-Preserving AI Verification via Minimal Information Disclosure
AI verification crosses a trust boundary: a verifier must learn enough to establish an authorized claim, yet the same evidence can reveal s…
A Hyperfinite Framework for Score-Based Generative Modeling
Score-based diffusion models are typically formulated using continuous-time stochastic differential equations and measure-theoretic stochas…
SAGE: Semantic Explainability of Attention-Based Survival Models in Computational Pathology
Attention-based multiple instance learning (ABMIL) is the predominant approach for slide-level prediction in computational pathology, yet i…
A Unified 2D Framework for DeepLesion Detection, Segmentation and Short Report Generation
In previous work, we integrated large language models (LLMs) into the lesion segmentation model based on the ULS23 DeepLesion dataset, usin…
Learning a Vector-Symbolic Model for Socio-Cultural Tasks
How can we better represent the impact of sociocultural structures on decision making in computational cognitive models? Modeling this impa…
Evading Chain-of-Thought Monitoring Through Model Poisoning
Chain-of-thought (CoT) monitoring is an increasingly important component of AI safety stacks but relies on the assumption that a model's re…
Improved Quantum Algorithms for Reinforcement Learning Under a Generative Model
Reinforcement learning is a subfield of machine learning that studies how an agent interacts with an environment in order to extract as lar…
In-Context Collapse in Vision-Language Models and How to Mitigate it?
Many-shot in-context learning (ICL) lets vision-language models (VLMs) adapt from image--label demonstrations without weight updates, and i…
CURV: Enhancing Chart Understanding Through Curriculum Visual Grounded Reasoning
Chart question answering (CQA) requires multimodal large language models (MLLMs) to integrate visual comprehension with logical reasoning,…
MutMem: Cryptographically Authorized Mutation in Persistent Agent Memory
Persistent agent memory must adapt as later outcomes change earlier evidence, yet mutable retrieval weights create an attribution problem:…
BODHI: Do LLMs Branch Out and Discover Heterogeneous Inferences?
Although reinforcement learning with verifiable rewards (RLVR) has improved the performance of large language models (LLMs) across a variet…
Robust Counterfactual Policy Optimisation via Nondeterministic Causal Models
Counterfactual inference approaches for sequential decision-making typically assume deterministic causal models, where all randomness stems…
When Should Graph Attention Be Sparse? Learning a Per-Edge Tsallis Index
Graph attention normalizes neighborhood scores with softmax, the maximum-entropy choice under Shannon statistics. But homophilic and hetero…
Rubrics as Privileged Information for Open-Ended Generation
On-policy self-distillation (OPSD), where a single model acts as both student and teacher with different contexts, has shown promise in ver…
SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling
Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model. Existing reward-model-fr…
Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits
This research paper describes an exploratory study on the effectiveness of Chat Debugging: troubleshooting malfunctioning analog circuits o…
ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies
Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks m…
Scaling an Autoregressive Transformer for Single-Cell Generation
We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to gene…
HyperFL: Query-Adaptive Representation Learning for Software Fault Localization
Software fault localization identifies the code locations responsible for reported issues and is a fundamental step toward automated debugg…
TQLite: Multi-LLM Jury Guided Distillation for Real-time MQM Translation Quality Evaluation
Large language models (LLMs) have demonstrated impressive performance in MQM-based translation quality (TQ) evaluation, and recent advances…
Internalising the Identity Primitive: Cryptographic Individuality for an Autonomous Agent on a Public Blockchain
A software agent on a public blockchain accumulates authority and economic stakes, raising the engineering question of what makes it count…
SparSEEty: Extracting Tokens from Sparsity-Exploiting LLM Serving Systems via Deterministic Side Channels
Modern large language models (LLMs) exhibit activation sparsity, wherein only a subset of their neurons is activated for given input tokens…
V-FIND: Revealing the Intrinsic Forgery Knowledge Encoded in Video Forgery Detectors
As generated videos become increasingly realistic, reliable video forgery detection is increasingly important. Existing studies typically o…
A Graph Signal Processing Perspective on Numerical Sequence Representations in LLM In-Context Learning
Pretrained large language models (LLMs) have demonstrated in-context learning (ICL) capabilities for numerical inference over sequences ser…
Standalone DINOv3 for Training-Free Open-Vocabulary Semantic Segmentation in Remote Sensing
Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods. Recen…
PACE: Adaptive Budget Allocation for Time-Efficient Embodied Planning
Reasoning-enhanced large language models have achieved remarkable improvements in planning tasks, yet their deployment in embodied systems…
LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs
Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engi…
PLAN: Parallel Liquid-Inspired Approximation Network for Efficient Representation Learning in Flexible Job Shop Scheduling
Deep reinforcement learning (DRL) approaches for flexible job shop scheduling (FJSP) heavily rely on attention-centric architectures to ach…
Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation
Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promi…
PI-Mem: Pushing Long-Context Reasoning to 3.6M Tokens with Parallel-Iterative Memory
Long-context reasoning remains a critical bottleneck for large language models, as recent recurrent-memory approaches face two inherent cha…
Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping
What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit…
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning
Reinforcement learning (RL) has emerged as an effective method for enhancing the reasoning capabilities of large language models (LLMs). Ho…
AI Security Leaderboard: Methodology, Results and Minimal Standard
Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on h…
CorePath: A Breast-Specialized Pathology Foundation Model for Core Needle Biopsy Diagnosis and Risk-Controlled Report Generation
Breast core needle biopsy (CNB) is central to breast cancer diagnosis yet remains challenging because limited tissue sampling, lesion heter…
SynEnergy: Anomaly Semantic-Guided Diffusion for Synthetic Energy Data Generation
Fine-grained energy consumption data are essential for applications such as demand forecasting, demand response planning, and grid reliabil…
SMOPD: Multi-Reward Reinforcement Learning via Specialize-and-Merge Online Policy Distillation
We aim to improve model performance in multi-reward reinforcement learning training process. Existing Group reward-Decoupled Normalization…
FakeI2V-Bench: Benchmarking the Applicability of Image-level Deepfake Detectors for Deepfake Video Detection
Recent advances in video generation models have significantly intensified the deepfake threat, yet the current deepfake video detection ben…
A Hierarchical Approach to Imitation Learning for Manipulation Tasks Requiring Time Varying Forces
Diffusion policies have shown strong performance in learning complex, multi-modal behaviors for robotic manipulation. However, their applic…
Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
Vision-language models excel at image and video understanding but suffer from high inference latency due to the need to process thousands o…
Optimal Liability Design for Medical AI
Artificial intelligence (AI) is increasingly integrated into medical decision-making, yet its liability implications remain complex, partic…
Trajectory-Guided Forget-Recover Network for Continual LLM Unlearning
Machine unlearning aims to eliminate the influence of sensitive data on a model. In the real world, unlearning requests arrive continually,…
DigitCode: Symbolic Tokenization of Hand Motion by Anatomical Units
Hand motion carries the finest-grained information in human activity, yet the representations behind hand generation, understanding, and ro…
Rectify Then Diffuse: Disentangling Concepts Before Denoising Trajectory Unfolds
Text-to-image diffusion models can generate individual concepts well, but they often omit or merge concepts incorrectly with multiple conce…
Internalizing Academic Writing Workflows for Introduction Generation via Struct-Aware Policy Learning
Generating a rigorous paper introduction with large language models (LLMs) remains challenging, since it requires coordinating background,…
Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue
We study contextual dynamic pricing with arbitrary covariate sequences and bounded, possibly nonbinary purchase quantities. Demand follows…
Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation
RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challen…
EFX Allocation In (Multi)Hypergraphs
We study fair allocations of indivisible goods among agents with heterogeneous monotone valuations. As fair we consider the allocations tha…
Attribute-based Undetectable Watermarking for Generative AI Models
Generative AI systems increasingly produce content whose provenance is difficult to verify, motivating watermarking techniques for identify…
Aligning Large Vision-Language Models at Test Time: A Trajectory-Guided Structured Sampling Approach
Post-training reinforcement learning (RL) algorithms are commonly used to align large vision-language models (LVLMs) with human intent and…
EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task,…
GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model
Reinforcement learning for flow-matching text-to-speech is complicated by deterministic ODE sampling: trajectory-level policy-gradient meth…
Self-Supervised Representation-Guided Generative Dataset Distillation
Dataset distillation compresses a large training set into a compact synthetic set while retaining its downstream utility. Most existing met…
Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks
Software engineering (SWE) agents resolve repository-level issues through long trajectories that grow increasingly expensive as context acc…
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
Agentic reinforcement learning enables LLM agents to learn through interaction, but sparse trajectory-level rewards reveal success without…
Structure-Aware Robust Fine-Tuning: Defending Vision-Language-Action Robots Against Physical Attention Hijacking
Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fra…
From Wearable Data to Personalized and Actionable Health Insights
Commercial wearable devices continuously capture rich physiological data (e.g., heart rate, respiration), opening new possibilities for mon…
FinVerse: Financial Time-Series Benchmark
As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has b…
The Ignition Is Real, and It Lives at the Readout: Latent composition, difficulty-clocked ignition, and the interface-constituted commit in a recurrent-depth reasoner
We test whether the "compositional ignition" reported in latent-reasoning models is real computation, an instrument artifact, or inherited…
Efficient Video Dataset Distillation via Cluster-Guided Prototype Blending
Video dataset distillation aims to compress a large video dataset into a compact surrogate set that preserves its training utility. Most ex…
GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
GUI grounding maps natural-language instructions to click locations and is essential for reliable GUI agents. The task remains difficult on…
Test-Time Scaling for Safe Text-Guided Image Generation via Intermediate Clean Estimates
Ensuring safety and policy compliance in text-to-image diffusion models remains a critical challenge, as benign or adversarial prompts can…
The Tell-Tale Trace: Detecting Reasoning Failures in LLMs Using Chain-of-Thought Dynamics
Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model'…
Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform
Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remedi…
Route-Align-Verify for Functional Correctness in Code Generation
Large language models (LLMs) have substantially improved code generation, yet achieving strong functional correctness remains difficult, es…
When Oracle Conditioning Misleads Deployment: Conditioning-Availability Bias in Echocardiographic Segmentation
Conditional segmentation models may be trained and evaluated with auxiliary signals cleaner than those available at deployment. We study th…
The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk
AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate human…
FACTWASH: Catching AI Rewrites That Wash Hearsay into Fact
AI systems rewrite information constantly: conversations become stored memories, documents become answers. The rewrite can keep a claim whi…
Shaping Wind-Tunnel Airflow for Unmanned Aerial Vehicles using Online Learning
The development and testing of advanced aerial robots require experiments in controlled environments with tailored airflow profiles. This p…
Shorter Reasoning, Earlier Answers? An Evaluation of Reasoning Interfaces
Large language models often reason at length before answering, increasing cost and latency. Prompts and trained settings can shorten this r…
Distilled Roads: Generalisable Road Network Extraction Across Sensors, Resolutions, and Region
Road network segmentation from satellite imagery remains challenging due to large geographic variation in road appearance, occlusions, and…
Multi-Task Multi-Frame Visual Piano Transcription
Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key releas…
OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains chall…
A Low-Cost Hybrid Reservoir Computing Model for Isolated Sign Language Video Recognition
Sign language recognition (SLR) enhances communication between hearing and hearing-impaired individuals. Although deep learning (DL) has ac…
Approximate Speculative Decoding
Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greed…
Balancing Efficiency and Efficacy: Training-Free Attention-Guided Switching Between Explicit and Latent Thoughts for MLLMs
Reasoning in Multimodal Large Language Models (MLLMs) requires both fine-grained visual perception and rigorous logical deduction. Explicit…
Adaptive Modality Reliability Diagnosis and Restoration for Robust Multimodal Intent Recognition
Multimodal intent recognition combines linguistic, acoustic, and visual evidence, but individual modalities may be noisy, missing, semantic…
Continue or Replan? Bernoulli-Continuation Policy Learning for Adaptive Horizon Execution
Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.e., execution horizon) before replanning, tur…
Leveraging System-Level Observations to Inform Bayesian Learning of Model Parameters for Quantitative Verification
Combining Bayesian learning and quantitative verification is a powerful toolset for analysing key quantitative properties of software syste…
Principles of Robot Autonomy
Autonomous robots are moving rapidly from research labs into everyday life - on roads, in the air, in warehouses, and in space. Robot auton…
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels…
How Many Labels Are Enough? ALDA: Active Learning Deployment Advisor for Medical Image Classification
Active learning (AL) promises to reduce the cost of medical imaging projects by lowering the number of clinical labels required. However, p…
AI Forensics Across White-, Grey-, and Black-Box Access: A Process Model and Research Agenda for Post-Incident Investigation of AI Systems
AI systems are increasingly involved in decisions and actions that may later require investigation. When an AI related incident occurs, inv…
Pivot-Centric Trajectory Prediction: Bridging Long Horizons via Dynamical Guidance
Forecasting precise future motion of surrounding agents is essential for reliable autonomous vehicles. However, as the demand for longer pr…
Training Documents Reranker with Search Rubrics for Deep Research Agent
Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typ…
Pin Once, Swap Light: Subspace-Aligned Centroid-Residual Training for Efficient Ultra-LoRA Serving
Modern multi-tenant Low-Rank Adapters (LoRAs) serving systems concurrently host tens to hundreds of LoRA adapters. Though powerful, this in…
AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality
AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little s…
GenOS: Compositional Certificates for Semantic Robustness in AI Code Generation
AI coding agents are stochastic workflows: prompts are interpreted, artifacts are sampled, validators produce observations, and orchestrato…
DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction
Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous te…
A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation
Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different obje…
A Security-Oriented Lifecycle Model for Large Language Model Systems
Large language models are being integrated into critical infrastructure and enterprise workflows at unprecedented scale,yet the lifecycle f…
MuEvo: LLM-Driven Evolution of Multi-Heuristic Ensemble
Large language model-based automated heuristic design (LLM-AHD) has shown strong potential in discovering effective heuristics for combinat…
Decoupling Generation and Selection for Budget-Constrained Faithful Summarization
Abstractive summarization models remain vulnerable to factual inconsistency, redundancy, and weak length control. We propose a modular gene…
How Closely Do LLM Reviews Align with Human Peer Review?
Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether differen…
Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation
Multimodal large language models (MLLMs) are increasingly used to translate webpage screenshots into front-end code, but repeated UI patter…
LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations…
GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models
Automated Knowledge Base Construction (AKBC) is a core NLP task, and recent work proposes generating knowledge bases directly from large la…
AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities
Sound effects play a crucial role in conveying actions, events, and environmental cues across digital applications, often requiring a high…
Can LLMs Test Terminal User Interfaces?
Terminal User Interfaces (TUIs) combine the stateful, screen-oriented behaviour of GUIs with terminal deployment and are now common in deve…
MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models
Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differ…
Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks
Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database…
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss
Small language models are often the only option for deployment under tight latency, cost, and on-premises constraints, but they are rarely…
Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure
An LLM-based agent is a loop that reads itself. Agentic frameworks externalize identity, memory, and disposition into editable files. The a…
VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs
Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, histor…
UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space
Large vision--language models (LVLMs) demonstrate strong multimodal reasoning capabilities but remain prone to hallucination, where model p…
FlowForm: Synergizing Fluid Physics with Topological Consistency for Satellite Flood Synthesis
Developing robust flood assessment models requires high-quality paired satellite imagery, yet such data remain scarce for flood-specific im…
Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking
Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexp…
SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG
We introduce SciRet, a compute-aware empirical study of retrieval-augmented generation for scientific question answering over CORD-19. Rath…
GENESIS: Towards Explainable Causal Discovery
Causal Discovery (CD) from observational data faces two fundamental challenges. First, purely statistical methods often lack the power to r…
Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses large foundation Vision-La…
MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning
Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigatio…
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement
A clinically useful chest X-ray system must go beyond fluent report generation: it should classify findings with tunable decision threshold…
When and Where to Look: Adaptive Visual Evidence Scheduling for Efficient Long Video Understanding
Efficient long-video understanding requires vision--language models (VLMs) to reason over a small number of frames selected as sparse visua…
Equivariant Music Transformer
Humans recognize a musical passage even when it is shifted in time or transposed in pitch, indicating a notion of equivariance in the repre…
PRISM: Powerful Time Series to Image (TS2I) Representations for Multivariate Anomaly Detection
Time series anomaly detection (TSAD) underpins applications in predictive maintenance, finance, and cloud computing, however performance re…
Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility
Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretr…
Separating quantum circuits from classical LLMs
Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction an…
Video-DeepResearch: Towards the Next-Generation Multimodal Deepresearch Agent
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that dema…
Can Large Language Models Recover Semantic Optimization Opportunities That Compilers Miss?
Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We…
Test-Time Scaling in Reasoning LLMs: Inference Regimes, Evaluation, and Reproducibility
Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," how…
TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement le…
A Unified Framework for Human AI Collaboration in Security Operations Centers with Trusted Autonomy
This article presents a structured framework for Human-AI collaboration in Security Operations Centers (SOCs), integrating AI autonomy, tru…
OR-Agent: Bridging Evolutionary Search and Structured Research for Automated Algorithm Discovery
Automating heuristic design in complex, experiment-driven domains requires more than iterative mutation of solution algorithms. Current LLM…
Modeling Matches as Language: A Generative Transformer Approach for Counterfactual Player Valuation in Football
Evaluating football player transfers is challenging because player actions depend strongly on tactical systems, teammates, and match contex…
Assessing the Effect of Cross-Domain Mapping on Creativity in Humans and Large Language Models
Creativity is the ability to come up with novel ideas, a capacity crucial for human development and flourishing. Are large language models…
LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics
Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment. However, traditional e…
AI Assistance Reduces Persistence and Hurts Independent Performance
People often optimize for long-term goals in collaboration: A mentor or companion doesn't just answer questions, but also scaffolds learnin…
An empirical evaluation of the risks of AI model updates using clinical data: stability, arbitrariness, and fairness
Artificial Intelligence (AI) and Machine Learning (ML) models used in clinical settings are increasingly deployed to support clinical decis…
Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on…
NOVA: Fundamental Limits of Knowledge Discovery Through AI
Can AI systems discover new knowledge through iterative self-improvement, and at what cost? We introduce NOVA, which models the ``generate,…
Language model agents show in-group trust bias invisible to standard behavioural audits
Language-model agents are moving from single-user assistants into persistent networks that build trust and reputation with one another, and…
Designing for Doubt: The Case for Informed Abstention in Autonomous Agents
As large language models gain tool access and are deployed as autonomous agents capable of editing records, executing transactions, and mod…
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perf…
Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process informati…
Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (…
Cura 1T: Specialized Model for Agentic Healthcare
Healthcare AI agents handle patient consultation, clinical reasoning over text and images, interactive diagnosis, and electronic health rec…
PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails
Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provi…
OPOD: On-Policy Omni Distillation
Omni-modal models provide a unified interface for text, images, and audio. However, improving these abilities together remains difficult, a…
CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics
Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central quest…
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed…
SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and…
The Geometric Nature and a Free Proxy for Flow-Matching Uncertainty
Flow matching (FM) has become a popular action head paradigm for modern embodied models. However, as a conditional generative model, it doe…
SKILL-KD: Contrastive Skill Distillation for LLM Agents
Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition meth…
Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation
Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory complianc…
A New Theory of Value for Post-AGI Economics
Artificial general intelligence (AGI) may weaken scarcities in labour, expertise, information, and productive capability that underpin esta…
Where Reasoning Diverges: Localized Multi-Agent Debate for Multi-Hop Question Answering
Multi-agent debate commonly exchanges complete rationales even when disagreements concern only a few intermediate claims. We introduce Loca…
LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing
DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remain…
When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary
Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusabl…
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance t…
Before Reasoning Can Fail: Pre-Evidence Procedural Failures in Agentic RAG
Agentic retrieval-augmented generation (RAG) systems can fail before evidence-conditioned reasoning is tested: an agent may retrieve candid…
SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents
Large language model agents increasingly solve complex tasks by composing reusable skills from a library. To address this, the key challeng…
A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures
Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in th…
Mixed-Initiative Human-Robot Teaming under Suboptimality with Online Bayesian Adaptation
For effective human-agent teaming, robots and other artificial intelligence (AI) agents must infer their human partner's abilities and beha…
MambaTS: Improved Selective State Space Models for Long-term Time Series Forecasting
In recent years, Transformers have become the de-facto architecture for long-term time series forecasting (LTSF), yet they face challenges…
CollaFuse: Collaborative Diffusion Models
In the landscape of generative artificial intelligence, diffusion-based models have emerged as a promising method for generating synthetic…
Efficient unsupervised domain adaptation via self-supervised vision transformer and synergistic cross-domain alignment
Unsupervised domain adaptation (UDA) aims to mitigate domain shift, where the distribution of labeled source data differs from that of unla…
Patient-centered data science: an integrative framework for evaluating and predicting clinical outcomes in the digital health era
This study proposes a novel, integrative framework for patient-centered data science in the digital health era. We developed a multidimensi…
Rex: A Family of Reversible Exponential (Stochastic) Runge-Kutta Solvers
Deep generative models based on neural differential equations have become state-of-the-art for many generation tasks. These models rely on…
Automated Visualization Code Synthesis via Multi-Path Reasoning and Feedback-Driven Optimization
Large Language Models (LLMs) have become a cornerstone for automated visualization code generation, enabling users to create charts through…
Compound and Parallel Modes of Tropical Convolutional Neural Networks
Convolutional neural networks (CNNs) are foundational to many state-of-the-art computer vision systems, yet their reliance on multiplicatio…
When Search Teaches Style: Causal Internalization of Tactical Priors in AlphaZero
AlphaZero is normally evaluated as one agent: a policy-value network fused with Monte Carlo tree search. That fusion hides a causal questio…
Beyond Either-Or Reasoning: Transduction and Induction as Cooperative Problem-Solving Paradigms
Traditionally, in Programming-by-example (PBE) the goal is to synthesize a program from a small set of input-output examples. Lately, PBE h…
One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting
Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores. We sh…
IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning
Importance-based structured pruning overwhelmingly relies on filter magnitude. This proxy is fundamentally flawed: due to scale invariance,…
From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
Adapting generative Multimodal Large Language Models (MLLMs) into universal embedding models typically demands resource-intensive contrasti…
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art p…
Uncovering Spontaneous Physics Representations in In-Context Learning
In-context learning (ICL) lets large language models (LLMs) solve new tasks from prompts alone, across an ever-widening range of domains, y…
Mechanism of Task-oriented Information Removal in In-context Learning
In-context Learning (ICL) is an emerging few-shot learning paradigm based on modern Language Models (LMs), yet its inner mechanism remains…
Malice in Agentland: Down the Rabbit Hole of Backdoors in the AI Supply Chain
While finetuning AI agents on interaction data -- such as web browsing or tool use -- improves their capabilities, it also introduces criti…
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
As language models become increasingly deployed in online environments, toxicity detection and detoxification have received growing attenti…
Toward Understanding the Transferability of Adversarial Suffixes in Large Language Models
Discrete optimization-based jailbreaking attacks on large language models aim to generate short, nonsensical suffixes that, when appended o…
GraphCliff: Short-Long Range Gating for Modeling Critical Activity Changes Caused by Subtle Molecular Differences
The quantitative structure-activity relationship assumes a smooth mapping between molecular structure and biological activity. However, act…
Target-Aligned Fusion for Decision-Sequence Learning under Dynamics Shift
External trajectories can improve offline decision-sequence learning, but dynamics shift may make some source subsequences inconsistent wit…
$\pi$-Attention: Online Efficient Sparse Transformers for Long-Context Modeling
Sparse attention is crucial in long-context Transformers, which restricts each token to a limited neighborhood and thereby reduces the quad…
STREAM-VAE: Dual-Path Routing for Slow and Fast Dynamics in Vehicle Telemetry Anomaly Detection
Automotive telemetry data exhibits slow drifts and fast spikes, often within the same sequence, making reliable anomaly detection challengi…
Externally Validated Breast Ultrasound Segmentation via Multi-task Learning with BI-RADS-Consistent Morphological Priors
External validation of breast ultrasound segmentation models remains limited because internal train--test splits do not capture domain shif…
MIMIC-MJX: Neuromechanical Emulation of Animal Behavior
The primary output of the nervous system is movement and behavior. While recent advances have democratized pose tracking during complex beh…
Self-Guided Adaptive Safety Alignment: Synthesizing and Internalizing Guidelines in Reasoning Models
Explicit safety policies can improve reasoning-model safety, but their effective coverage may lag behind evolving jailbreak strategies. We…
PRISMA: Improving the Accuracy-Latency Frontier of Diffusion-based PDE Solvers Using Physics-Informed Spectral Attention
Diffusion-based solvers for partial differential equations (PDEs) are often bottle-necked by slow gradient-based test-time optimization rou…
PRIVEE: Privacy-Preserving Vertical Federated Learning Against Feature Inference Attacks
Vertical Federated Learning (VFL) enables collaborative model training across organizations that share common user samples but hold disjoin…
HERO: Hierarchical Evidential Reasoning Optimization for Radiology Report Generation via Reason-then-Summarize
Multimodal Large Language Models (MLLMs) have substantially advanced Radiology Report Generation (RRG), yet aligning them through reinforce…
Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models
In language models (LMs), intra-memory knowledge conflict arises when inconsistent information about the same subject is encoded within the…
ChiEngMixBench: Evaluating Large Language Models on Expert-Style Chinese-English Terminology Mixing
Large language models increasingly mediate multilingual professional communication, where useful generation requires adapting to community…
AgenticSCR: An Autonomous Agentic Secure Code Review for Immature Vulnerabilities Detection
Secure code review is critical during pre-integration, where Atlassian developers rely on lightweight analysis tools, while deep security a…
On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
Recent work has shown that layer pruning can effectively compress large language models (LLMs) while retaining strong performance on classi…
A Deployment-Friendly Foundational Framework for Efficient Computational Pathology
Pathology foundation models (PFMs) generalize well across computational pathology tasks but remain costly for gigapixel whole-slide image a…
In-Context Pure Exploration in Continuous Decision Spaces
In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to id…
SphUnc: Hyperspherical Uncertainty Decomposition and Causal Identification via Information Geometry
Reliable decision-making in complex multi-agent systems requires calibrated predictions and interpretable uncertainty. We introduce SphUnc,…
Quantifying Hallucinations in Language Language Models on Medical Textbooks
Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious p…
Large Language Models provide support for the parallelogram theory of analogy
Four-term word analogies (A:B::C:D) are classically modeled geometrically as parallelograms: adding the vector B-A+C produces D. Recent wor…
The production of meaning in the processing of natural language
Understanding the fundamental mechanisms governing the production of meaning in the processing of natural language is critical for designin…
DIB-OD: Preserving the Invariant Core for Robust Heterogeneous Graph Adaptation via Decoupled Information Bottleneck and Online Distillation
Graph pre-training can facilitate knowledge transfer across graph datasets, but severe structural and feature shifts may cause negative tra…
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone…
Gated Memory Policy: In-Context Memorization and Adaptation
Robotic manipulation tasks exhibit varying memory requirements, ranging from Markovian tasks that require no memory to non-Markovian tasks…
A neural operator framework for data-driven discovery of stability and receptivity in physical systems
Understanding how complex systems respond to perturbations, such as whether they will remain stable or what their most sensitive patterns a…
Estimating Tail Risks in Language Model Output Distributions
Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these model…
Evaluating LLM-Based Goal Extraction in Requirements Engineering: Prompting Strategies and Their Limitations
Due to the textual and repetitive nature of many Requirements Engineering (RE) artefacts, Large Language Models (LLMs) have proven useful t…
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
Traditional Shot Boundary Detection (SBD) inherently struggles with complex transitions by formulating the task around isolated cut points,…
Injection-Execution Dissociation: A Mechanistic Evaluation of Persistent Memory Attacks and Defenses in Stateful LLM Agents
We discover that prompt-injection success and tool-execution success are separable safety properties: defenses that block injection do not…
Improving Reproducibility in Evaluation through Multi-Level Annotator Modeling
As generative AI models such as large language models (LLMs) become more pervasive, ensuring the safety, robustness, and overall trustworth…
Lean Refactor: Multi-Objective Controllable Proof Optimization via Agentic Strategy Search
We present Lean Refactor, a plug-and-play retrieval-augmented agentic framework for multi-objective, controllable, and version-robust refac…
ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models
Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment…
E4GEN: Event-level Explainable Extreme-Enhanced Time-series Generation
Generating realistic time series is essential for scientific research and real-world applications. However, existing methods often emphasiz…
FLARE: Diffusion for Hybrid Language Model
Autoregressive (AR) large language models (LLMs) have achieved broad practical success, but sequential decoding remains a key bottleneck fo…
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and ans…
When Context Returns: Toward Robust Internalization in On-Policy Distillation
Recent work has shown that on-policy distillation can internalize privileged context, such as system prompts or task hints, into a student…
CADET: Physics-Grounded Causal Auditing and Training-Free Deconfounding of End-to-End Driving Planners
End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that me…
Diagnosing and Mitigating Context Rot in Long-horizon Search
Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon search tasks. The concern t…
MalariAI: A Label-Resilient Decoupled Framework for Annotation-Agnostic Cell Segmentation and Explainable Stage Classification in Dense Malaria Blood Smears
Automated malaria diagnosis from blood smear microscopy is a critical global health AI challenge; expert scarcity remains the primary diagn…
VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training para…
Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks
Symmetry is everywhere in nature and society. Geometric deep learning builds architectures respecting group symmetries, whereas topological…
Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5
Emergent misalignment (EM) --- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data --- is mediated in…
x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability
Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function eva…
Subjective Risk Decomposition: A New View for Uncertainty Quantification
We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axioms and argumentation,…
Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min--Max Selection
Large single-cell datasets are expensive to store, curate, and repeatedly reuse for model training. Data distillation can reduce this burde…
A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation and Cross-Domain Generalisability
Robust out-of-the-box performance is essential for the clinical deployment of deep learning models in medical imaging. An important but und…
TriGlue: a Biology-Inspired Generative Model for Generating Molecular Glue-Induced Ternary Complex
Molecular glue degraders have emerged as a promising strategy for targeted protein degradation by inducing ternary complex formation betwee…
CausalForge: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference
Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A co…
Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcem…
WCM: World-Cognition Model for Generalizable Human-Robot Interaction
Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tas…
Moral Hazard in Multi-Agent Language Models
Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr\"om's team…
AgentGUI: An Interface for Observing and Steering Long-Running AI Agents
AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight…
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
Jigsaw puzzle solving requires jointly reasoning about visual content and geometric constraints, yet existing benchmarks use rectangular cu…
PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks
SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability. They typically follow a common construction pip…
Automated ECG Interval Measurement and Wave Delineation Using Fast Fourier Convolution ResNet
Accurate measurement of ECG intervals, including PR, QRS duration, and QT/QTc, is central to cardiac diagnosis, yet the published ECG delin…
Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs
Multimodal large language models (MLLMs) increasingly support high-stakes decision making by combining complementary information from image…
LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations
LLM-native advertising embeds sponsored content directly into model-generated responses, shifting the unit of sale from a fixed slot to a m…
Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems
Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, eff…
When Prompts Control Robots: Prompt Injection Attacks in Multi-Agent Robotic Systems
Large language models are increasingly integrated into autonomous robotic systems for task planning and control, but this integration expos…
ACE-GraphRAG: Agentic Context Engineering for Hierarchical GraphRAG
Hierarchical Graph Retrieval-Augmented Generation (GraphRAG) organizes corpus knowledge at multiple levels of granularity, yet fixed contex…
Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Metric for Infrared-Visible Fusion Assessment
Infrared-visible image fusion (IVIF) has no ideal fused reference, so fusion algorithms are routinely ranked by scalar objective metrics th…
Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution
Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts. I…
Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation
However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation rema…
LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation
Post-training large language models (LLMs) via reinforcement learning (RL) has significantly advanced code generation capabilities. To bypa…
Self-Improving Large Language Models via Progressive Experience Evolution
Large language models (LLMs) capable of self-improvement require not only effective policy optimization, but also a principled mechanism fo…
PhyCheck: Fine-Grained Evidence-Grounded Dataset for Physical Law Understanding in Video-LLMs
Embodied intelligence and world models require video understanding systems to go beyond recognizing objects and actions and develop an unde…
「フルスクラッチ開発」って何?──LLMを“骨格”と“筋肉”に例えて国産モデルの現在地を整理する
「LLMはどこまで『国産』であるべきか?」と題した本特集では、日本企業が安全性と利便性を確保したAI活用を実現するにはどのレイヤーを国産にすべきなのか、そもそも国産にはどのようなメリットがあるのかを、代表的な国内ベンダーへの取材などを通して考察する。第1回の本稿では、そもそもL…
Gemini“ヘビーユーザー”が1年で10倍に 「平均年齢高い、IT本業じゃない」首都高が実践したAI活用推進ノウハウ
月100回以上Geminiを使う首都高速道路の従業員が、1年で22人から224人に急増した。転機は対面・終日のワークショップと「誰一人取り残さない」全社推進、そして「Gemini Notebook」の存在だった。
AI活用で浮上する「導入済みERP」の課題とは? 中堅企業のERPリプレース、新規導入を上回る
ノークリサーチによると、中堅・中小企業のERP市場で、中堅企業だけはリプレースの規模が新規導入を上回る。なぜ中堅企業はERPを入れ替えるのか。導入済みERPの課題とは。
M365 Copilotのための「Teams活用術3選」 北大DX業務推進室が解説
北海道大学のDX業務推進室が、米MicrosoftのAIサービス「Microsoft 365 Copilot」を活用しやすくするための「Microsoft Teams」の使い方を紹介している。
「私がウダウダしゃべるより聞きやすい」 三菱重工、決算説明に“AIジュリア”起用の本音
三菱重工が、決算説明にAIナレーター「ジュリア」を起用した。導入の狙いを問われた西尾浩CFOが、笑顔で語った理由とは。
「日本人が海外へ行けない」を打ち破る シンガポール発LCCが仕掛ける“逆張り戦略”
1ドル160円超の円安と燃油高により「海外旅行離れ」が進む日本市場。シンガポール航空傘下のLCC「スクート」が羽田・那覇線を相次ぎ新設する“逆張り攻勢”を見せている。なぜ円安下でも拡大を続けるのか。「若者・女性向け」を脱した中小企業などのビジネス需要獲得、最新規格「NDC」を活…
【8/5まで】AIを活用した開発について大調査【Amazonギフトカードが当たる】
現在、「AIを活用した開発業務」に関する読者調査を実施しています。回答者の中から抽選で3名さまにAmazonギフトカード500円分をプレゼントします。
中小企業で「AI活用を回す」には? JAPAN AIと大塚商会らが考える支援策
生成AIの課題は「導入」から「定着」へと移っている。AIツールが社内に広がらず成果につながらない中堅・中小企業に求められることとは。
SpaceX has bought $329M worth of Tesla Megapacks so far this year
The purchase illustrates just how interconnected Elon Musk's universe of companies are.
Open-weight AI models are catching up to the frontier. The safety gap remains.
A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing co…
Anthropic signs $10B deal with AI cloud startup Volta
Anthropic has been on a cloud partnership spree in recent months, and its latest move is reportedly a $10 billion deal with AI cloud startu…
Meet Wrinkles, an app that uncovers the hidden stories of the places around you
Wrinkles, available on both iOS and Android, essentially acts as an AI-powered audio tour guide that reveals hidden history and local stori…
Nvidia doesn’t mess around: A week after open AI industry group formed, it’s already showing progress
The week-old Open Secure AI Alliance, spearheaded by Nvidia and grown to over 120 companies, already has proposals out for defending agains…
Third-party cyber evaluations involving OpenAI models
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evalua…
Spotify expands AI remix and covers project with Merlin partnership
Spotify says Merlin, which represents more than 30,000 independent labels and distributors, has joined Universal Music Group in backing its…
Texas halts new data centers as governor calls for audits
Tech companies and developers have been scouring the U.S. for places to build data centers, and they’ve been drawn to Texas’ loose regulati…
Elon Musk spends half his time talking robots and AI on Tesla earnings calls
An analysis of the last seven years of Tesla earnings calls shows just how little attention Musk pays to Tesla's car business.
2026-08-04(17件)
Apple says more ex-employees may have taken confidential data to OpenAI
Apple says its trade secrets investigation into OpenAI has widened. In a new court filing, Apple claims additional former staff may have re…
Is the future of data centers portable? Runware builds a pod to find out
On Tuesday, AI infrastructure company Runware announced the launch of its own modular data center called Sonic Inference Pod.
EON wants to move the data superhighway from ocean fiber to space lasers
Endeavor Optical Networks is planning to launch the fastest space laser communications system yet built.
AI予算の7割を食う「ITインフラ」、中でも予算超過しやすいのは? 医療機関調査
Wasabi Technologiesは、医療機関におけるAI活用に関する調査結果を発表した。AI予算の多くを占めるITインフラでは、どのような課題が生じているのか。調査結果から背景を探る。
塩野義製薬、生成AIの正答率を50→90%に 膨大な機密データをどう最適化した?
生成AIを社内実務に組み込みたくても、情報の機密性やデータ量の多さなどに導入が難しい業界もある。そのうちの一つ、製薬業界に属する塩野義製薬も、同様の課題を抱えていたが、ある手法によってそれを解決したという。同社はいかにして、AI導入を進めたのだろうか。
EU、AIの透明性義務の適用を開始 生成コンテンツにラベルとマーク義務、違反に最大1500万ユーロ
EUの欧州委員会はAI規制法「EU AI Act」の第50条に基づく透明性ルールの適用を開始した。生成AIやディープフェイク等を扱う事業者に対し、AIとの対話の明示やコンテンツへのラベル・機械可読マークの付与を義務付ける。違反企業には最大1500万ユーロまたは売上高の3%の制裁…
After killer quarter, Palantir CEO Alex Karp calls AI industry ‘Marxist’
After a quarter that delivered $1 billion in profit, Palantir CEO Alex Karp on Monday once again warned that AI frontier labs are too untru…
PC操作を録画→「Copilot」で作業を代理可能に Microsoftがアプリを無料公開 主にmacOS向け、Windows対応も
Microsoftが、PCの画面操作を録画すると、その作業をAIが再現できるよう支援するデスクトップアプリ「Skill Recorder」を公開している。AIエージェントサービス「Microsoft Copilot Cowork」「Microsoft Scout」「Copilo…
iPaaS、AIで広がる「使いどころ」とは? 2030年度まで年平均20.7%成長の背景
AI活用を前提としたITシステム基盤への需要が高まっている。ITRの調査によると、国内のiPaaS市場は2025年度に前年度比18.0%増となり、2030年度までのCAGRは20.7%になる見通しだ。ニーズが高まっている用途とは。
AI活用を「個人の効率化」で終わらせるな セールスフォース流「4つの組織改革術」
生成AIを全社活用できている企業が11%にとどまる中、セールスフォース・ジャパンの社内では300ものAIエージェントが自律稼働している。同社はなぜ「個人の業務効率化」で止めず、全社変革(AX)へ導けるのか。管理不能な「野良エージェント」を防ぐ線引きや、組織と人員を再構築する「4…
「融資」現場にAIの足音、資金調達どう変わる? カネを借りられる企業の条件、専修大教授に聞く
融資の審査にAIを使う動きがある。資金を調達する企業にとっての「審査が遅い」などの課題を解決できのか。逆に、借りる側に変化はあるのか。専修大学の尾木研三教授に聞いた。
「大きな投資計画が次々に。久しぶりだ」――強く豊かな日本投資枠、経済成長かなうか? 片山大臣が語る狙い
政府は、2027年度予算の概算要求において、成長投資枠は「予算の上限額を設けない」とした。その狙いと意気込みを片山さつき財務大臣が語った。
AWS is helping vibe-coding startup Superblocks, and the implications are big
AWS now allows vibe-coding tool Superblocks to be embedded into the private clouds of AWS customers. It's another step toward decoupling ap…
Design Arena creators raise $7.9 million to bring taste to AI models
Design Arena is used by 5.3 million people around the world, providing critical human evaluations to frontier labs.
Influencers draw backlash for attending OpenAI’s first luxury trip
OpenAI’s first-ever influencer brand trip is sparking online backlash as tensions over the use of AI continue.
Apple finally fixed Siri. So why does it feel anticlimactic?
Apple’s long-awaited AI overhaul finally makes Siri the assistant it was always supposed to be. Yet it arrives at a moment when simply bein…
Congress’ favorite AI tool? ChatGPT
House spending records show OpenAI's ChatGPT dominates paid AI use on Capitol Hill, with congressional offices relying on the chatbot to dr…
2026-08-03(277件)
スクエニ、ゲームの品質テストをGeminiで自動化 AIが画面を見ながらコントローラーを操作、検証作業を自走
スクウェア・エニックスが、ゲームのQAテストを「Gemini」で自動化する取り組みを「Google Cloud Next Tokyo '26」基調講演で披露。AIが画面を見ながらコントローラーを操作し、検証作業を自ら進める。
実在女性の中学時代の体操着姿からAIわいせつ画像作成・投稿か 男逮捕、高校生書類送検
女性の写真を生成AIで加工したわいせつ画像をSNSに投稿したとして、警視庁などは、名誉毀損(きそん)と児童買春・ポルノ禁止法違反(公然陳列)の疑いで、兵庫県姫路市の会社員、井元健太容疑者(32)を逮捕した。また、画像の加工を依頼したとして、鹿児島県垂水市の高校3年の男子生徒(1…
A Marc Benioff-backed startup thinks AI can solve the AI deployment problem
June emerged from stealth today with a $20 million pre-seed round to make AI adoption simpler.
カメラとディスプレイ搭載のAIグラス、「Rokid スマートAIグラス」を試してみた
「Rokid スマートAIグラス」の一般発売が7月10日に始まった。製品を借りることができたので、現在地の評価と、未来の可能性について考えてみたい。AIグラスは、何を可能にし、何を可能にしないのだろうか。
How we built a realtime system for responsive voice AI in six months
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural…
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent Systems
The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the arch…
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model Review
AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluatin…
LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann Hypothesis
Major mathematical conjectures still depend heavily on expert intuition, so a unified method for the systematic generation and validation o…
ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning
Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow,…
TAPR: Enhancing LLM Performance with a Task-Aware Prompt Rewriter
Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert…
Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel Decoding
Cross-domain sequential recommendation (CDSR) aims to model users' dynamic interest transitions and sequential patterns across multiple dom…
An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents
Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fractur…
How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought Trajectories
Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: exi…
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have ac…
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent…
Multi-Agent Planning with Spatio-Temporal and Topological Constraints using STL-GO
Multi-agent planning problems arise in a variety of engineering applications, such as multi-robot wildfire fighting and unmanned aerial ins…
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making i…
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of the…
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool Acquisition
Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computation…
EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter Diagnoses
Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks a…
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention woul…
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions
As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptab…
Fragility of Value under Imperfect Alignment
As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humani…
Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design
Computational cognitive modeling seeks to infer latent cognitive mechanisms underlying observed behavior. Bayesian inverse planning provide…
NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability
Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retriev…
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate…
Scaling Scientific Discovery Environments for Turn-Level Agentic RL
Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an e…
MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping Agents
Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that…
Evidence-Grounded Constraint Checking in Construction Documents
Professional-document review is a constraint-checking problem in which decisions depend on relations among text, geometry, pages, and docum…
On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness
Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where…
A Generalized-Bayes Perspective on Counterfactual Explanations: Posterior-Based Decision-Making and Evaluation
Counterfactual explanations (CEs) enhance the interpretability of machine learning models by identifying the smallest change to an input re…
Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative Collaboration
Large language models (LLMs) are increasingly deployed in enterprise settings, yet individual models remain bounded by model-specific capab…
CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using Agents
Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permissio…
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unf…
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
Modern large language models (LLMs) are expected not just to answer correctly, but to adapt their behavior to different human values and us…
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into…
MAGA: Multi-Platform Self-Fusion of GUI Agents via Structured Action Distillation
Graphical user interface (GUI) agents based on large language models are increasingly deployed across mobile, web, and desktop environments…
Beyond Component Testing: Validating Agentic AI Systems
Agentic AI systems act through multi-step trajectories that combine planning, tool use, memory, interaction, and adaptation. This behavior…
ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models
Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated mo…
Beyond Retrieval: Analytic Memory for Multimodal Agents
Long-term multimodal memory must support not only retrieving relevant information but also computing over observations accumulated across i…
Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
Self-play agents can generate training problems without questions from target benchmarks, but their curricula lack persistent state: failur…
AMTFV: Agentic Mathematical Tool-Flow Verification for LLM Self-Correction
Large language models have demonstrated strong mathematical problem-solving capabilities, yet reliably verifying their candidate answers re…
COntExt: Towards Context-Aware Ontology Extension from Operational Metrics
Organizations increasingly define operational metrics in structured, machine-readable formats to monitor systems, processes, and compliance…
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
Reinforcement Learning (RL) systems are typically trained using a single, well-specified scalar reward function. However, real-world decisi…
DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat
Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich…
AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become i…
Development of FDD-ON: an Ontology for VAV HVAC System Fault Detection and Diagnostics
Fault detection and diagnosis (FDD) technology is essential for improving HVAC system reliability, energy efficiency, and maintenance effec…
ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction
Enterprise workflows increasingly rely on agents for \emph{schema-guided extraction}: given a document and a user-defined schema, the agent…
Scaffolding Critical Engagement with GenAI: Transforming Ethnic Minority Preparatory Students' Collaborative Discourse in Prompt Engineering Tasks
Generative AI (GenAI) holds significant promise for advancing educational equity among ethnic minority students by broadening access to lea…
Topology-Aware Data Movement for Disaggregated GPU Inference
Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run o…
The Asymmetric Effects of Knowledge Distillation on Bias in Small Language Models
We show that knowledge distillation in small instruction-tuned language models has asymmetric effects on bias. On unambiguous tasks (BBQ-di…
The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems confla…
Seeing Differently: Modeling Interpretive Perspectives in Computational Creativity using a Four-World Framework
Creativity in computational systems is often evaluated as an objective property of artifacts, with existing Computational Creativity (CC) f…
Looks Right, Works Right: A Project-Level Benchmark for Multi-Screen Mobile App Generation
Recent multimodal large language models can convert visual designs directly into executable code, but real mobile products require multiple…
ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning
This paper presents ConnectED, a human-centered AI system that supports the full instructional lifecycle in Vietnamese education by linking…
Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations
Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on mod…
COSI-Lab: Conference Living Lab for Modeling Multi-Perspective Multimodal Social Intention
COSI-Lab presents a multimodal, multi-sensor dataset of an interdisciplinary scientific workshop containing 32 academics at an internationa…
Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration
While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration…
HenTwin: A Multimodal Digital Twin Framework for Longitudinal Biological State Monitoring in Laying Hens
Early-life monitoring in laying hens remains constrained by fragmented single-modality sensing and the absence of formal system-level state…
Evaluating Federated Pre-Training: On the Reliability of Downstream Fine-Tuning and Intrinsic Evaluation
Federated pre-training offers a way to train foundation models on private or distributed data without centralizing the underlying datasets.…
Sensitivity Analysis of GRU, LSTM and Transformer Encoder in Classification of Automated Driving Systems
Automated driving systems (ADSs) are becoming ubiquitous. Future Software Defined Vehicles (SDVs) may be able to run multiple ADSs, both na…
Guarantees on Dynamical System Distinguishability for LLM Token Generation
Recent work has shown that classifying large language models (LLMs)' responses can be distinguished by modeling token embeddings as traject…
LAWFUL: Law-Aligned Witness for Faithful Use of Latents
When a neural network predicts a physical system accurately, has it learned the governing law as formal, structured knowledge, and if so, d…
MPP-GNN: Subject-Adaptive Community Detection for fMRI-Based Alzheimer's Disease Classification
Functional magnetic resonance imaging (fMRI) is a widely used technique for studying the brain. Recent methods that utilize graph neural ne…
Metaphor-Induced Algorithmic Steering: Cross-Domain Procedural Transfer in LLM Code Generation
Large language models benefit from elements in natural language, such as metaphors and analogies in training data and inference input to ac…
Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions
As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades…
Reflected UAS: Corrected Deterministic Stability and Direct CTMC Drift Calculation
We analyze Reflected UAS routing for heterogeneous multi-server queues at fixed parameters under subcritical load. The deterministic surrog…
Code Is the Body: Agent-Owned Software Bodies for Recursive Evolution and Descent
Personalized AI agents are often configurable without giving users control over the artifacts that determine their future behavior. We pres…
SEDR-Seq2P: A Lightweight Dilated Residual Sequence-to-Point Network for Multi-Task Industrial NILM
Industrial NILM remains challenging because measurement noise and widespread concurrent machine operation reduce the generalization of mode…
Predicting Steel Fatigue Life from Micrographs Using Physics-Informed Deep Learning
Here is the plain text version optimized for arXiv's submission form. Custom macros (like \CV and \SI) have been converted to standard text…
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the r…
A user's guide to PINNs in geometric analysis: lessons from the asymptotic Plateau problem
This proceedings contribution elaborates on the findings of arXiv:2605.26234v2: a joint work with Marco Usula, where we introduced a machin…
DragonCrawl: A Generative, Intent-Based Framework for Scalable Mobile End-to-End Testing
As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhea…
SCMA: Structure-Conditioned and Metal-Aware Flow Matching for CT Metal Artifact Reduction
In X-ray CT, metallic objects cause beam hardening, photon starvation, and scattering, leading to projection inconsistency, streaks, dark b…
WaiT for the Signal: Simple Frequency-Aware Flow-Matching
As image generation models scale to ever higher resolutions, global coherence, local detail, and texture fidelity become critical axes for…
Stratified Negation in RDF Rules: A Correct Approach (Extended Version)
Combining RDF rule languages, such as N3 or SHACL Rules, with default negation is challenging. Existing methods to stratify negation often…
Benchmarks Are Not Monolithic: Sample-Level Auditing and Orchestration for LLM Evaluation
Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring…
Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing
In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance,…
Hypergradient-based Bilevel Reinforcement Learning with Improved Sample Complexity
Bilevel reinforcement learning (RL) is an important framework within the literature of RL that can be used to formalize various categories…
A Unified Benchmark of Deep Learning Models for Multi-task 3D Brain Tumor Segmentation from Magnetic Resonance Imaging
Automatic brain tumor segmentation from magnetic resonance imaging (MRI) has become a fundamental task in computer-assisted diagnosis, trea…
TextCloak: Thwarting Unauthorized LLM Exploitation via RL-Driven Unlearnable Text
The rapid development of Large Language Models (LLMs) has led to significant advances across a wide range of language tasks, while simultan…
Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?
When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that tr…
RareSense: Rarity-Aware Similarity Search for Anomaly Retrieval in Transactional Data
Similarity search over sparse set-valued data is often dominated by frequent background attributes because classical measures such as Jacca…
To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing
Large language models increasingly write and repair production code, yet evidence is mounting that their test-passing patches leave codebas…
Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use
Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large lang…
Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth
Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that…
TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models
Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a bas…
Design Concept: Scaffolding Geopolitical Reflection Among Tech Workers
This paper presents a speculative Human-Computer Interaction design proposal for encouraging geopolitical reflexivity amongst tech workers…
Gated Q-learning: Add Off-Policy Bias to Taste
Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a f…
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations b…
DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models
Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the de…
Retrieval-Driven Training-Free AI-Generated Video Attribution
AI-generated videos are becoming increasingly realistic and difficult to distinguish from authentic ones, which facilitates malicious misus…
Efficient LLM Adversarial Training via Low-Rank Defense and Circuit-Guided Surrogates
Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at m…
A robust association between LLM use and scientific productivity: Assessing stopping-time selection
Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged i…
RAID: Towards Robust AI-Generated Image Detection with Bit-Reversed Images
The rapid advancement of image generation models has made it increasingly difficult for people to distinguish AI-generated images from real…
PARALLEL: A Prefrontal-Aligned Reinforcement inspired Approach for Language-Model Learning under Explicit Limits
Recent language models achieve strong performance across a variety of tasks, but conventional adaptation applies updates uniformly across t…
Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
Zero-shot image captioning (ZIC) describes images without paired image-caption supervision during captioner training, relying on text-only…
Point2Radio: A Foundation Model for Cross-Scene Radio Fields from Material-Aware Point Clouds
High-fidelity radio fields are typically simulated for every scene--transmitter configuration or fitted separately to each scene, failing t…
Auto-JEPA: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving
Existing autonomous-driving world models typically perform dense prediction of future videos, occupancy states, BEV representations, or age…
Improving scDiffusion with Sparsity-Biased Classifier-Free Guidance
Single-cell RNA sequencing (scRNA-seq) has become an essential tool in modern cellular biology, and generating accurate synthetic scRNA-seq…
Learning Lookahead Lemmas for Neural Network Verification
State-of-the-art neural network verifiers use the branch-and-bound procedure as their core solving mechanism. We introduce an inprocessing…
Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search
Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to m…
Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives
Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it rem…
Semantics of Subterfuge: Benchmarking Legal Deception Detection Against General-domain State-of-the-Art
Deception detection has critical implications for legal proceedings, law enforcement, and online security. Although human judgment is limit…
Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients
Federated learning of foundation models faces a fundamental resource-asymmetry challenge: the institutions holding the most valuable domain…
metasignal: A Python Package for Comprehensive Metacognitive Analysis and Decision-Making
Metasignal is an open-source Python package for signal detection theory (SDT) and metacognitive measurement. It implements the 17 metacogni…
DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs
Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal…
Multi-Granularity Position Embedding of Graphs via Granular-Ball for Link Prediction
Link prediction aims to identify potential or future connections within a given graph structure. Position information is essential for link…
InferQ: A Database-Oriented Benchmark for Quantum Circuits Simulation
Recent work suggests that relational database management systems (RDBMSs) can execute quantum circuit simulation by compiling the simulatio…
HERO: History-Enriched Rollout Training for Long-Horizon Autoregressive Neural Operators
Neural operators provide fast surrogates for time-dependent partial differential equations (PDEs) by applying a learned evolution operator…
Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership
Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the gene…
Implicit Machine Learning Force Fields Accelerate Molecular Dynamics Simulations
We introduce implicit machine learning force fields (I-MLFFs), which replace explicit stacks of neural network layers with self-consistent…
Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory
Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations int…
ActFovea: Runtime Safeguarding for VLA Policies via Spatiotemporal Visual-Action Consistency
Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that…
CLIFT: Turning Gemini Robotics On-Device into Humanoid Specialists via Non-Invasive Closed-Loop Iterative Fine-Tuning
While robot foundation models are growing increasingly capable, the strongest models are typically trained on proprietary data and remain c…
MBDiff: Multi-view Behavior-aware Diffusion Model for Probabilistic Utility Data Imputation
Utility data (e.g., electricity, water, and gas consumption), collected by ubiquitous sensors and embedded devices, often contains substant…
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation
Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural appr…
SERUM: State Extraction and Refinement for User Modeling
Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. However, building…
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillatio…
MOSAIC: Masked Outsourcing of Secure AI Computations
We address the challenge of securely and efficiently outsourcing AI computations from a trusted but computationally weak client to an untru…
Linear Proposal Operators and Stochastic Search Geometry in SOMA and Differential Evolution
Swarm and evolutionary algorithms are usually analyzed as complete procedural systems in which nonlinear selection, replacement, and adapta…
FBFM: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution
Although world-action models (WAMs) enhance long-horizon robot control by predicting visual evolution before acting, long-horizon reliabili…
Small Is Enough: Per-User Style Rewriting of AI-Edited Text via LoRA Adapters
InMyStyle is a privacy first, single user system that adapts small language models to rewrite AI-edited text towards an individual user's w…
When Model Priors Conflict with Visual Evidence: Mitigating Commonsense-Driven Hallucinations by Selective Prior Calibration
In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence…
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems
Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-stra…
TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation
Computed tomography angiography (CTA) is crucial for preprocedural TAVI planning, providing the anatomical information required for prosthe…
CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing…
OsteoCAD: A Human-in-the-Loop Cloud-Edge Framework for Bone Tumor Segmentation
Artificial Intelligence (AI) and Deep Learning (DL) have notably advanced medical image analysis, yet many health- care organizations strug…
Translation with Thought: Difficulty-Adaptive Reasoning via Reinforcement Learning for Multi-Domain Machine Translation
Multi-domain machine translation (MDMT) poses a unique challenge due to varying levels of linguistic complexity across domains. Inspired by…
The persuasive power of large language models does not depend on their perceived national origin
Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rej…
DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation
Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of a…
SeekBrain: An Autonomous Multi-Agent System for Accelerating Neuroscience Discovery
Modern neuroscience relies on integrating multi-scale, multimodal datasets to uncover the neural principles underlying intelligence. Howeve…
Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning
With the ever-increasing pervasiveness of smart edge devices, the demand is growing for applications that can be tailored to users (e.g., c…
Cross-Lingual Transfer for Machine Translation in Turkic Languages
Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains in…
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Balancing sequence length, representational capacity, and long-horizon stability is a central problem in autoregressive (AR) speech and aud…
Dense Temporal Contrast Synthesis via Conditioned Latent Transport
Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is essential for breast cancer management, but reliance on gadolinium-based…
Explore Beyond the Boundary Using Entropic Information
In reinforcement learning, exploration with sparse and delayed rewards presents a significant challenge due to the limited feedback availab…
AgenticRepair: Multi-Faceted Program Context Engineering for Agentic Vulnerability Repair
Automated vulnerability repair aims to reduce the time and effort required to patch security flaws from a vulnerability triage report. Rece…
QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models
Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual questio…
TFGformer: Multivariate Time Series Forecasting via Time-Frequency Graph Learning and Covariate Fusion
Large-scale multivariate time series from heterogeneous IoT sensors demand accurate long-term forecasting for resource scheduling and predi…
DreamQAS: Learning a Decision-Useful World Model for VQE-Efficient Quantum Architecture Search
Reinforcement-learning-based quantum architecture search (RL-QAS) repeatedly optimizes a variational quantum eigensolver (VQE) after extend…
From Code Review to Code Critique: Intent, Drift, and Spotlight for AI-Generated Diffs at Scale
AI coding agents are generating code at volumes that exceed the capacity of traditional peer review. At the same time, existing AI code rev…
TerraNova: A Foundation Model for the Anthropocene
A defining problem of the Anthropocene is to model the physical Earth and human societies as one coupled system, yet no learned representat…
ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation
Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While pri…
MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models
Symbolic Regression (SR) aims to discover analytical equations from observational data and plays a central role in scientific modeling. Whi…
TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning
The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and ap…
FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models
Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two peop…
A Human-Centered Validation of the Explainability-Performance Coefficient
The rapid adoption of deep learning models in high-risk domains has intensified the need for trustworthy Explainable Artificial Intelligenc…
When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning
Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to lang…
CENDRe: Concept Extraction with Natural Domain Representations
Convolutional neural networks (CNNs) are widely used for time-series classification, but their deployment in critical domains requires unde…
The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic fee…
SATViz: Real-Time Visualization of Clausal Proofs
Visual layouts of graphs representing SAT instances can highlight the community structure of SAT instances. The community structure of SAT…
Combining Large Language Models and Symbolic Reasoning for Multi-Robot Temporal Planning through Explainable Knowledge Bases
We present PLANTOR, a framework for generating and executing multi-robot task plans from natural-language task descriptions through LLM-ass…
Shall We Play a Game? Language Models for Open-ended Wargames
LLM-based social simulations can make a generated transcript look like a single behavioral signal, but the model behind that transcript may…
Embedded Universal Predictive Intelligence: a coherent framework for multi-agent learning
The standard theory of model-free reinforcement learning assumes that the environment dynamics are stationary and that agents are decoupled…
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Agentic reasoning models trained with multimodal reinforcement learning (MMRL) have become increasingly capable, yet they are almost univer…
M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities
As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to i…
RAPiD: Reward-Guided Consistency Distillation of Diffusion Planners for Real-Time Autonomous Driving
Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bott…
Shaping Scientific Explanations to Expert Perspectives with Persona-Conditioned Reinforcement Learning
Explainable AI is increasingly important to scientific discovery. However, existing methods largely ignore that explanation quality is not…
What Makes a Sale? Simulating End-to-End Seller--Buyer Retail Dynamics with LLM Agents
Evaluating retail strategies before deployment is difficult, as outcomes are determined across multiple stages, from seller-side persuasion…
PEMAND: Persona-Enriched Multi-Agent Negotiation for Household Decision-Making
Modeling household-level decisions is central to many real-world applications, including trip planning, residential mobility and migration,…
SREGym: A Live Benchmark for AI SRE Agents with High-Fidelity Failure Scenarios
AI agents are increasingly used to diagnose and mitigate failures in production systems, known as agentic Site Reliability Engineering (SRE…
Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling
Large Language Models (LLMs) have demonstrated remarkable abilities in reasoning. However, maximizing their potential through inference-tim…
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
A significant hurdle for current LLMs is the execution of complex, multi-stage tasks. Group Relative Policy Optimization (GRPO) has been em…
The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models
Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from ex…
A Multi-Agent System for Motor Design Optimization via an FEA-AI Hybrid Approach
This study presents a large language model (LLM)-based multi-agent framework for interior permanent magnet synchronous motor (IPMSM) design…
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by ineffi…
ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
Reinforcement Learning with Verifiable Rewards (RLVR) is a central technique for improving long-horizon reasoning in Large Language Models…
Cognitive World Model for Progressive BDI/E Trajectory Evaluation of Conversational Agents
As LLM-based conversational agents advance toward increasingly open-ended and interaction-intensive scenarios, task completion alone provid…
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) ev…
Latent Actions from Factorized Transition Effects under Agent Ambiguity
Latent Action Models (LAMs) learn action-like proxies from observation. However, in multi-object or distractor-rich scenes, observations co…
LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents
Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic…
Solution Space Path Planning: A Real-Time Human-Centered Path Planning Algorithm for En-Route Air Traffic Control
As technology advances, various algorithms have been proposed for air traffic management, yet their operational adoption in tactical contro…
The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale
The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared…
NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning
OWL ontologies provide a formal knowledge representation framework that enables semantic reasoning, and have been widely adopted across dom…
Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs
Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable…
CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an inc…
AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what…
Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning
Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy…
DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training
RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffe…
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enab…
Reason-Mediated Behavioral Models for Auditing LLM Social Simulators
Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether s…
Information Processing by Neuron Populations in the Central Nervous System: A Theory of the Mathematical Structure of Data and Operations
In the mammalian central nervous system, neurons are organized into populations communicating by spike trains propagating along axonal bund…
On the Expressive Power of Sparse Geometric MPNNs
Motivated by applications in chemistry and other sciences, we study the expressive power of message-passing neural networks for geometric g…
Revisiting Multi-Permutation Equivariance through the Lens of Irreducible Representations
This paper explores the characterization of equivariant linear layers for representations of permutations and related groups. Unlike tradit…
Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content. We i…
Dual-Force: Enhanced Offline Diversity Maximization under Imitation Constraints
Offline diversity maximization under imitation constraints can transform demonstration data into a set of distinct behavioral policies, imp…
Dimensionality reduction for homological stability and global structure preservation
We propose DiRe, a force-directed dimensionality reduction framework designed to preserve global structure and homological features while r…
Reproducing Human Individual Motor Signatures: A Data-Driven Approach for Repetitive Motion
The deployment of autonomous virtual avatars (in extended reality) and robots in human group activities---such as rehabilitation therapy, s…
StaQ: a Finite Memory Approach to Discrete Action Policy Mirror Descent
In Reinforcement Learning (RL), regularization with a Kullback-Leibler divergence that penalizes large deviations between successive polici…
Towards White-Box Deep Wireless Sensing
The empirical success of deep learning has spurred its application to the radio-frequency (RF) domain, leading to significant advances in D…
Patch-Based 3D Variational Autoencoder for Super-Resolution of Turbulent Channel Flow
Direct numerical simulation (DNS) accurately resolves all spatio-temporal scales of wall-bounded turbulence but becomes prohibitively expen…
RePaCA: Leveraging Reasoning Large Language Models for Static Automated Patch Correctness Assessment
Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention. However, existing tools t…
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
Homelessness is a persistent social challenge, impacting millions worldwide. Over 876,000 people experiencing homelessness (PEH) were recor…
Adaptive Policy Backbone via Shared Network
Reinforcement learning (RL) has achieved impressive results across domains, yet learning an optimal policy typically requires extensive int…
Fast Feature Field ($\text{F}^3$): A Predictive Representation of Events
This paper develops a mathematical argument and algorithms for building representations of data from event-based cameras, that we call Fast…
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
Recent medical vision-language models have shown promise on tasks such as VQA, report generation, and anomaly detection. However, most are…
Monotone and Separable Set Functions: Characterizations and Neural Models
Motivated by applications for set containment problems, we consider the following fundamental problem: can we design set-to-vector function…
Pay for The Second-Best Service: A Game-Theoretic Approach Against Dishonest LLM Providers
The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability:…
Robust Bidirectional Associative Memory via Regularization Inspired by the Subspace Rotation Algorithm
Bidirectional Associative Memory (BAM) trained with Bidirectional Backpropagation (B-BP) often suffers from poor robustness and high sensit…
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rat…
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
Web applications (web apps) have become a key arena for large language models (LLMs) to demonstrate their code generation capabilities and…
GPU-Accelerated ANNS: Quantized for Speed, Built for Change
Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications. GPUs offer a promi…
GeoRA: Geometry-Aware Low-Rank Adaptation for RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) is a key paradigm for improving large-scale reasoning models. Unlike supervised fine-…
Knowledge Restoration-driven Prompt Optimization: Unlocking LLM Potential for Open-Domain Relational Triplet Extraction
Open-domain Relational Triplet Extraction (ORTE) aims to mine structured knowledge without predefined relation schemas. Large Language Mode…
When Iterative RAG Beats Ideal Evidence: A Diagnostic Study in Scientific Multi-hop Question Answering
Retrieval-Augmented Generation (RAG) extends large language models (LLMs) beyond parametric knowledge, yet it is unclear when iterative ret…
Towards the Holographic Characteristic of LLMs for Efficient Short-text Generation
The recent advancements in Large Language Models (LLMs) have attracted interest in exploring their in-context learning abilities and chain-…
AIvilization v0: Toward Large-Scale Artificial Social Simulation with a Unified Agent Architecture and Adaptive Agent Profiles
AIvilization v0 is a publicly deployed large-scale artificial society that couples a resource-constrained sandbox with a unified LLM-agent…
Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas ot…
Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery
Recent advancements in Large Language Model (LLM) agents have demonstrated remarkable potential in automatic knowledge discovery. However,…
Stem: Rethinking Causal Information Flow in Sparse Attention
The quadratic computational complexity of self-attention remains a fundamental bottleneck for scaling Large Language Models (LLMs) to long…
Step-Level Visual Grounding Faithfulness Predicts Out-of-Distribution Generalization in Long-Horizon Vision-Language Models
We uncover a behavioral law of long-horizon vision-language models: models that maintain temporally grounded beliefs generalize better. Sta…
Wrong Code, Right Structure: Learning Netlist Representations from Imperfect LLM-Generated RTL
Learning effective netlist representations is fundamentally constrained by the scarcity of labeled datasets, as real designs are protected…
ELISA: An Interpretable Hybrid Generative AI Agent for Expression-Grounded Discovery in Single-Cell Genomics
Translating single-cell RNA sequencing (scRNA-seq) data into mechanistic biological hypotheses remains a critical bottleneck, as agentic AI…
Preconditioned Test-Time Adaptation for Out-of-Distribution Debiasing in Narrative Generation
Although debiased large language models (LLMs) excel at handling known or low-bias prompts, they often fail on unfamiliar and high-bias pro…
Demystifying Video Reasoning
Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capa…
OPERA: Online Data Pruning for Efficient Retrieval Model Adaptation
Domain-specific finetuning is essential for dense retrievers, yet not all data pairs contribute equally to the learning process. We introdu…
Agentic Harness for Real-World Compilers
Compilers are critical to modern computing, yet fixing compiler bugs is difficult. While recent large language model (LLM) advancements ena…
Maximum Entropy Behavior Exploration for Sim2Real Zero-Shot Reinforcement Learning
Zero-shot reinforcement learning (RL) algorithms aim to learn a family of policies from a reward-free dataset, and recover optimal policies…
Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations
In collaboration with Alibaba, we study how a generative AI assistant affects service performance in e-commerce after-sales operations. In…
ActionParty: Multi-Subject Action Binding in Generative Video Games
Recent advances in video diffusion have enabled the development of "world models" capable of simulating interactive environments. However,…
Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation
We study compiled AI, a paradigm in which large language models generate executable code artifacts during a compilation phase, after which…
Evaluating the Alignment Between GeoAI Explanations and Domain Knowledge in Satellite-Based Flood Mapping
The increasing number of satellites has improved the temporal resolution of Earth observation, making satellite-based flood mapping a promi…
TimeRFT: Stimulating Generalizable Time Series Forecasting for TSFMs via Reinforcement Finetuning
Time Series Foundation Models (TSFMs) have demonstrated strong generalization capability and data efficiency in time series forecasting thr…
Escaping Mode Collapse in LLM Generation via Geometric Regulation
Mode collapse is a persistent challenge in generative modeling and appears in autoregressive text generation as behaviors ranging from expl…
Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs
Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation th…
Leveraging Image Generators to Address Data Scarcity: The Gen4Regen Dataset for Forest Regeneration Mapping
Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geogra…
Detecting AI-Generated Videos with Spiking Neural Networks
Modern AI-generated videos are photorealistic at the single-frame level, leaving inter-frame dynamics as the main remaining axis for detect…
A Nonlinear Singular Value Theory for Neural Networks
Recently Brown et al. [2025] established a singular value decomposition (SVD) for maps (especially nonlinear) satisfying certain norm condi…
DRIP-R: A Benchmark for Decision-Making and Reasoning Under Real-World Policy Ambiguity in the Retail Domain
LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by i…
Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning
Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where…
When Bits Break Recourse: Counterfactual-Faithful Quantization
Model quantization is widely used to reduce memory, latency, and deployment cost, and is typically judged by whether predictive accuracy is…
AI4BayesCode: From Natural Language Descriptions to Validated Modular Stateful Bayesian Samplers
Coding and computation remain major bottlenecks in Markov chain Monte Carlo (MCMC) workflows, especially as modern sampling algorithms have…
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term…
MemForest: An Efficient Agent Memory System with Hierarchical Temporal Indexing
Memory is a fundamental component for long-context LLM agents, supporting persistent state across interactions through a continuous serve-a…
PEFT of SLM for Telecommunications Customer Support: A Comparative Study of LoRA Configurations with Energy Consumption Analysis
While large language models (LLMs) show strong performance in natural language understanding and generation, their evaluation and adaptatio…
Multi-Scale Feature Attention Network for Polymer Classification Using Terahertz Spectroscopy
Reliable polymer identification is essential for ensuring the quality and safety of recycled plastics, yet conventional sorting and spectro…
APPO: Agentic Procedural Policy Optimization
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language m…
Creative Integration: A Decidable Criterion of Creativity
"Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes t…
Implicit Reasoning for Large Language Model-based Generative Recommendation
Large Language Models (LLMs) are increasingly adopted as backbones for Generative Recommendation (GR), promising access to pretrained world…
The Metanym Game: A Self-Contained, Self-Consistent LLM Peer-Community Benchmark for Structural Intelligence
The metanym game is a competitive word game for LLMs that measures structural intelligence against established cognitive-science constructs…
SqLinear: Balanced Square Partitioning Makes Linear Interaction Sufficient for Large-Scale Traffic Forecasting
Traffic prediction is a core task in intelligent transportation systems and urban-scale decision making. Despite the effectiveness of mains…
DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training
Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in comple…
ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Co…
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a n…
BeatEdit: Symbolic Music Generation as Explicit Editing
Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete s…
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-age…
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such…
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Ex…
CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization
Textual Collaborative Prompt Optimization (TCPO) extends TextGrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multipl…
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning
Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to…
HijackKV: New Threat in Position-Independent KV Cache Reuse
Key-Value (KV) cache reduces inference latency in large language models (LLMs). Traditional prefix-based reuse has low cache hit rates acro…
ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundat…
Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS
Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of wheth…
Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric
Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in…
DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifie…
Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature
X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published…
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
Language-model agents act through structured tool calls whose arguments carry very different risks: untrusted content may legitimately shap…
LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this g…
A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks
Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies…
Progressive Multimodal Alignment for Continual Instruction Tuning
Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it c…
Benchmarking LLM Competence on Logical Inference over Probability Operators
Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions o…
SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups
Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Exis…
LabEvolver: Training-Free Experience Evolution for Safe and Grounded Wet-Lab Agents
We introduce LabEvolver, a training-free framework that equips safe and grounded wet-lab agents with episodic memory from execution experie…
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn…
On a joint simultaneous learning of relevant feature subsets and subspaces in regression-like problems
We extend a recently introduced Entropy-Optimal Manifold Clustering (EOMC) to allow for a joint simultaneous identification of subsets and…
「Qwen3.8-Max」登場、オープン化は「来週」 一部「Fable 5」「GPT-5.6 Sol」超えの性能うたう
中国Alibaba傘下のAlibaba Cloudは8月3日、AIモデル「Qwen3.8-Max」を正式にリリースした。来週にはモデルの重みも公開する予定だ。
富士通とNECは「AI需要」と「収益」をどう語った? 両社決算会見から2026年下半期の見通しを考察
企業の業務に向けたAI需要はどのような動きなのか。AI需要の盛り上がりと、ITサービス企業の収益は直結するのか。富士通とNECのCFOによる直近決算会見での発言から考察する。
「エンジニアの代替ではなく、できることを増やす」 Devin開発元が見据えるIT外注大国・日本の“伸びしろ”
自律型AIエンジニア「Devin」を手掛ける米Cognition AIの日本ユーザー人口は、米国に次ぐ規模だという。IT人材の約7割がSIer側に偏るIT外注大国・日本で同社は何を狙うのか。日本法人代表の正井拓己氏に聞いた。
【レベル14】生成AIを味方に、3D CADを使いこなそう!
設計スキルのレベルアップを目指す設計者の皆さんを“冒険者”に見立て、さまざまな“問(モン)スター”に挑む「テルえもんクエストII」の世界へようこそ。【レベル14】のテーマは、生成AIを活用した3D CADデータの作成と操作です。
月100億トークン使うビズリーチ 「AIコスト増」懸念の中、費用対効果どう判断しているのか
AI活用が広がる一方で、コスト増加への懸念が高まっている。月間約100億トークンを消費するというビズリーチCTOは、AIの費用対効果をどのように考えているのか。
賞金1000万のAIコンテスト、でも「実現性は問わず」 サイバーエージェントのAI推進策
生成AIを導入しても、従業員に使われなければ成果にはつながらない。サイバーエージェントは、賞金1000万円のAIコンテストを開催し、あえて「実現性」を問わない仕掛けで社員の意識を変えた。その狙いとは。
生成AI利用率、情シスよりも高いのはあの職種だった
生成AIの活用状況は、担当する業務や職種によって大きく異なる。職種別の生成AI利用実態を把握するためにラグザスが実施した調査によると、エンジニア・情シスの生成AI利用率は職種別で1位ではなかった。では、その職種がそれらを上回ったのか。
「AI、結局使えないじゃん」問題 セールスフォースが431万件対応で導いた正解
AIエージェントの導入が進む一方、全社的な利益改善に至る企業は4割弱にとどまる。そんな「AI導入の壁」を破り、自社実践(カスタマーゼロ)で431万件の顧客対応を完了、商談数を4~5割増加させたのがセールスフォースだ。同社はなぜ明快なROIを生み出せるのか。データやKPIが整う領…
WAFを89%すり抜ける事例も──AIが休みなく仕掛けるWeb攻撃、予防策はあるか
フロンティアAIの登場で一変したWeb攻撃。その傾向や対策、落とし穴をAkamai Technologiesの中西一博氏が解説。
OpenAI、次期主力モデル「Astra」の存在を明らかに――未解決の数学問題10件を「解決」と発表
OpenAIは、次期主力モデル「Astra」の社内版により、数学や理論計算機科学の未解決問題10件で新たな結果を得たと発表した。「Astra」の名称公表は初とみられる。計算コストは「Sol」換算で約2000ドルに抑えられ、証明支援系「Lean」による形式証明もGitHubで公開…
Sam Altman and AI’s decel debate
On the latest episode of Equity, we discuss why Sam Altman has calling on the industry to "pace the rate of AI development."