週次AIニュース 2026-W34
対象期間: 2026-08-17 〜 2026-08-23(1481 件)
トピックの推移
トピック別件数
- 研究/論文 593件
- LLM/生成AI 559件
- エージェント 339件
- 画像/動画生成 205件
- ロボティクス 85件
- ビジネス/資金調達 79件
- ハードウェア/半導体 50件
- その他 49件
- 規制/政策 5件
今週のハイライト(上位 10 件)
From Atari to EVE Online: Building on 15 Years of AI Research in Games
Google DeepMind partners with game studios to prototype breakthrough AI gameplay.
Introducing AI Futures
Introducing AI Futures, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual free…
Offering Zero Data Retention for frontier models
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compr…
Replit expands access to software creation with GPT-5.6 Luna
Replit introduces Free Mode, powered by GPT-5.6 Luna, so anyone can turn ideas into working software without worrying about token costs.
Strengthening democratic oversight in national security
OpenAI launches an initiative to strengthen democratic oversight of AI in national security, supporting government institutions with tools,…
Partnering with CodeAI to prepare the first AI generation
OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it…
Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model de…
Introducing ChatGPT for Teens: Built for learning, backed by protections
ChatGPT for Teens helps teens learn, think critically, and use AI with confidence, with stronger built-in protections, healthy-use features…
Asana cleared 5 years of engineering work in 2 weeks with Codex
Asana used OpenAI Codex to replace an outdated testing system in two weeks, completing work expected to take five years for about $12K.
OpenAI、「GPT-5.6 Sol」のAPI料金を値下げ 入力20%出力33%安く、11月21日まで
OpenAIは、フラグシップモデル「GPT-5.6 Sol」のAPI利用料金を期間限定で最大33%引き下げた。入力4ドル、出力20ドルとするプロモーション価格を少なくとも3カ月間適用し、ChatGPT WorkやCodexのクレジット消費量も削減。利用量や支出上限を追跡・制御で…
全件(日付別)
2026-08-23(5件)
OpenAI、「GPT-5.6 Sol」のAPI料金を値下げ 入力20%出力33%安く、11月21日まで
OpenAIは、フラグシップモデル「GPT-5.6 Sol」のAPI利用料金を期間限定で最大33%引き下げた。入力4ドル、出力20ドルとするプロモーション価格を少なくとも3カ月間適用し、ChatGPT WorkやCodexのクレジット消費量も削減。利用量や支出上限を追跡・制御で…
Harvard’s $699 startup bootcamp offers AI avatars of its instructors
In the HBS Foundry program, AI avatars provide feedback during practice pitches and board meetings.
Inherent, founded by DeepMind alumni, says its AI ‘teammate’ just outperformed Anthropic and OpenAI at replicating research
Built by DeepMind alumni, British AI lab Inherent released Faraday, an AI agent whose ability to replicate scientific papers could be a ste…
OpenAI says California should strengthen its AI safety bill
OpenAI is calling for California to strengthen SB 53, an AI safety bill that the company previously opposed.
Frontier AI labs still won’t say how they’d contain a rogue model
A new study finds leading AI labs have few publicly documented plans for containing rogue models, raising questions about preparedness as A…
2026-08-22(4件)
Anthropic’s Opus 4.6 is a smut-machine
Anthropic forbids its Claude models from generating sexually explicit content. But a series of tests conducted by TechCrunch found that it…
Nvidia partners with data center developer Cloverleaf
Nvidia continues to pour money into data center development — just as AI data centers bring lots of money into Nvidia.
Anthropic、「ミュトス 5」を脆弱性スキャンに開放──「Claude Security」経由でEnterprise顧客が利用可能に
Anthropicは、最上位モデル「Claude Mythos 5」をセキュリティサービス「Claude Security」の脆弱性検出に導入したと発表した。モデル本体への直接アクセスは開放せず、パッチ提案などの出力のみに限定して提供する。オープンソース保護に向けた総額3500…
Nvidia just showed that the harness, not the AI model, is now the real hero
Nvidia research shows that AI agents can perform well, and not go off the deep end, through fine-tuning, even if the AI model isn't that gr…
2026-08-21(244件)
The DOJ is investigating a16z. What does this mean for venture capital?
Andreessen Horowitz has two partners sitting on the boards of companies that now compete with each other: Ben Horowitz at Databricks and Ma…
Starcloud raises $250 million for orbital data centers as launch options dry up
There's about to be a big fight to secure access to space.
From Atari to EVE Online: Building on 15 Years of AI Research in Games
Google DeepMind partners with game studios to prototype breakthrough AI gameplay.
SNSのウソ画像、どう見破る? 熊本県庁やテレビ局も頼る“すごい企業”の正体
スペクティが提供する「Spectee Pro」は、さまざまな情報を収集し、その時に起きている「危機」を可視化するシステムだ。多くの自治体やマスコミも活用しているというSpectee Proは、どうやってデマや虚偽の情報を見分けるのか。
エイベックス松浦会長「AIで仕事が楽になると思ってたけど、真逆」 note記事作成の“苦労”明かす
「AIで仕事が楽になると思ってたけど、真逆でした」――エイベックスの松浦勝人会長は、自身のXアカウントでこのように投稿した。AIを活用したnoteの記事制作の一端を明かした。
中国AI「Kimi」が日本進出か 有料プランのプレゼントキャンペーンも 「はじめまして、日本」
AIモデル「Kimi」を開発する中国Moonshot AIは、Kimiの日本語版公式X(@KimiAI_Japan)で「今日から、日本での歩みを始める」と投稿した。日本に本格進出するとみられる。
FANZAで「成人向けAIコンテンツ制作サービス」開始 8月24日から先行体験
成人向けECサイト「FANZA」を運営するデジタルコマースが、成人向けAIコンテンツプラットフォーム「FANZAスタジオ」の提供を始める予定だ。8月24日からβ版の先行体験を始める。
Active Inference as Context Acquisition for AI Agents
Interactive AI agents must acquire the right context as efficiently as possible. When a user omits a constraint, preference, file, or task…
Robust Metaheuristics under Uncertainty for Berth Allocation and Quay Crane Assignment: A Review
The berth allocation and quay crane assignment problem (BACAP) is a representative port-terminal scheduling problem in maritime transportat…
How to Navigate Uncertainty About AI Consciousness
Given deep uncertainty about the possibility of artificial consciousness, it is unclear how we should treat potentially sentient AI. On the…
Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the Model
AI control research asks how to deploy models safely even when they may be misaligned, but many control protocols assume that the deployer…
Interaction valence reveals contrasting social networks in dairy cattle
Social relationships shape access to resources, exposure to conflict and group stability, yet automated livestock monitoring typically trea…
Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and Evaluation
Air traffic control (ATC) communication is a safety-critical dialogue that remains largely human-driven even as other parts of air traffic…
Outcome Monitors: Recovery Affordances for Silent Tool Failures
When a tool call times out, the agent sees the failure and can route around it. A cached error page or negative price can instead arrive in…
Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress
On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajector…
Symposium: Trust via Auditable Records for Communities of AI Scientist Agents
Symposium is a formal framework and practical implementation to record the operation of AI agents deployed by small scientific research com…
From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG
Retrieval-augmented generation (RAG) improves language-model responses by grounding generation in external passages, which comes with overh…
Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding Dynamics
Large Language Models (LLMs) are increasingly deployed in high-stakes applications, yet their tendency to generate toxic, harmful, or polic…
Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at Scale
Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery,…
Can Agent Memory Systems Track Evolving State?
As LLM-based agents are deployed for longer and higher-stakes tasks, their memory systems continue to have crucial gaps. While existing mem…
Frequency-Aware Continual Learning for Smart Contract Vulnerability Detection with Large Language Models
Smart contract vulnerability detection with Large Language Models (LLMs) faces three causally linked challenges. First, new vulnerability c…
Learning Hierarchical Skill Policies with Offline Quality-Diversity Reinforcement Learning
Recent studies investigate how to leverage pre-collected datasets to improve the policy performance and sample efficiency of RL. One promis…
Rethinking the Evaluation and Optimization of LLM-Based Social Simulation
LLM-based social simulation is a promising complement to traditional methods such as surveys and behavioral experiments. A core question is…
Beyond Memory Majority: Latent-Source Reasoning for Multi-Agent Memory Arbitration
Long-term multi-agent systems continuously accumulate the memories produced by different agents. Existing memory methods typically treat re…
SafeBranch: Branch-Pair Safety Alignment for Embodied Agents
Vision-language-model-based embodied agents can complete instructed tasks but often violate safety constraints in the process, a problem re…
GenMatch: An End-to-End Generative Matching Framework for Micro-View Order-Dispatching in Ride-Hailing
Micro-View Order-Dispatching assigns available drivers to passenger orders within each dispatch batch and is critical to the service qualit…
TT-net: Quantum Inspired Tensor Network Denoising in Conditional GANs
Developed as a workhorse for classical simulations of quantum algorithms and quantum many-body systems, Tensor Network methods have entered…
LLMs as Acquisition Policies for Finite-Pool Materials Optimization: A Controlled Study
Discovering materials with desirable properties often requires searching large candidate spaces while experimental or computational evaluat…
Towards general embodied intelligence: integrating large language models, knowledge bases, and reasoning capabilities to build the next generation of AI agents
The convergence of large language models (LLMs), structured knowledge bases (KBs), and reasoning ability (RA) presents a promising trajecto…
ADAPT: Physics-Aware Diffusion-based World Models for Adaptive Predictive Transferable HVAC Control
Buildings account for roughly one-third of global energy consumption and CO$_2$ emissions. Optimizing indoor climate systems plays a critic…
When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content Creation
To prevent the adoption of aesthetically polished but pedagogically flawed AI content, we study a video authoring pipeline featuring two la…
Causal Reasoning with Bipartite Graphical Causal Models
Causal Bayesian networks (CBNs) and structural causal models (SCMs) are the dominant frameworks for graphical causal reasoning, but they ca…
Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platforms
Spec Driven Development SDD has consolidated the idea that the specification rather than the code should be the primary artefact governing…
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning
Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-…
PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents
Customer-service LLM agents must follow organizational policy when acting on a user's behalf. Compliance failures arise from either forbidd…
EnvHarness: Awakening Static Worlds for Agent Learning
LLM agents learn by interacting with environments, yet these environments are hand-built and static: blind to an agent's weaknesses, and qu…
TESTNAV: Pareto-Guided Search for Compositional Robustness Testing
Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (…
Write Once, Run Everywhere: The Axon DSL for Shape-Safe and Framework-Agnostic LLM Architectures
The entire ecosystem of open-source language models effectively relies on a single platform. What if this platform was forced to shut down…
EXIMO: VLM Guided Exploration of VLA Policies
How to efficiently finetune robot policies to learn new tasks on the fly? State of the art robotic manipulation policies are based on behav…
Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysis
AI agents can execute scientific analyses, but an analytic output becomes a defensible claim only after alternatives are weighed and the cl…
Spike-based Belief Propagation in Nonlinear Dynamical Systems
This paper presents a Bayesian control framework that integrates spike-based dynamics with probabilistic inference for adaptive control. Ba…
A Strong Linear Baseline for Whole-Heart Cardiac Shape Completion on CT, with an Open Eleven-Structure Statistical Shape Model
Public cardiac cohorts annotate different subsets of the heart, so shapes from separate sources cannot be pooled without shared corresponde…
Learning Early-to-Final Solution Consistency for MILP Acceleration
Mixed-Integer Linear Programming (MILP) is a fundamental problem class in operations research and combinatorial optimization, with broad ap…
Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured Partitioning
Multivariate time series forecasting (MTSF) is a fundamental task in many real world applications. Existing patch based forecasting methods…
ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance
LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence…
Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees
Loading reusable skill documents into a bounded context window is now the primary way large language model (LLM) agents acquire task-specif…
ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory Forecasting
Understanding object dynamics requires not only predicting future trajectories but also examining whether a model captures the physical pro…
Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking
Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-triv…
Contrastive Mixed Prompt Learning for Incomplete Multimodal Sentiment Analysis with Unseen Modality Combination
Incomplete multimodal sentiment analysis has garnered significant attention in recent years. Existing approaches typically assume that data…
A three-dimensional typology of agency for advanced AI systems
Research on the agency of advanced artificial intelligence (AI) systems focuses on agency as a normative concept and on the agency of parti…
On the Applicability of Safety Nets: A Safety-By-Design Solution for Certifying Neural Networks
The integration of Artificial Intelligence (AI) in safety-critical aviation systems presents significant challenges for certification and d…
What You Can't See Is What You Learn: Restricted Evidence Visibility Favors Compositional Generalization in Shared-Genome Language-Model Societies
Multi-module systems often expose every module to the full input. We test whether restricting evidence visibility changes which solutions g…
DECOWAM: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation
Mobile manipulation requires a robot to predict how locomotion and arm motion jointly alter future observations and control. Existing world…
DARS: Dual-Level Credit Assignment RL with Structured Reasoning for Instruction-Based Image Editing
Instruction-based image editing uses a planner-renderer pipeline: a vision-language model (VLM) first converts the instruction into an edit…
The Third Restructuring of Software Form: From the Three-Tier Architecture to Storage, Models, and Agents
Software form has undergone two paradigm shifts since its inception: Software 1.0, in which instructions determine behavior, and Software 2…
MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use
Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. Howe…
ContractScrub: A benchmark for final review of legal contracts
Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of…
Electronic Navigational Chart Change Classification
Electronic Navigational Charts (ENCs) are geospatial vector datasets used in maritime navigation systems that represent hydrographic and na…
InsufficiencyBench: Evaluating LLM legal advice on underspecified user queries
Legal AI systems are increasingly used to answer legal questions, yet existing benchmarks assume queries arrive fully specified. In practic…
Rule-Compliant Visual Spatial Planning for Multimodal Large Language Models
Multimodal large language models (MLLMs) combine linguistic reasoning with visual perception, yet their ability to perform visual spatial p…
QUASAR: A Quantum-Classical Neural Network for SAR Satellite Physical-Layer Authentication
X-band SAR satellites (8-12 GHz) play a critical role in disaster response, environmental monitoring, and military intelligence. Yet, they…
Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation
Reasoning language models trained with reinforcement learning typically operate under a fixed token budget rather than an explicitly adapti…
Catching the Rug: Early Prediction of Fraudulent Memecoins on Solana via Machine Learning
The rapid proliferation of memecoins on blockchain platforms has increased the risk of fraudulent activities, particularly rug pulls. While…
Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents
Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In prac…
Phantom Gains: Auditing Self-Improvement Against a Measured Null
Whether a language model has improved itself is increasingly judged not by mean accuracy but by which individual problems it gains and lose…
MidTool: Mid-training Data Synthesis for Agentic Tool Use
Mid-training is increasingly recognized as a critical stage for shaping the capabilities of large language models. Recent work has shown th…
Pandora's AI Model Routing Box: Efficient Allocation with Costly Value Estimation
Heterogeneous AI systems composed of multiple models, architectures, harnesses, or inference-time settings can improve quality and efficien…
AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
Recursive self-improvement (RSI) asks whether an AI system can improve the process that produces AI systems, so that the next system inheri…
An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction
Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and e…
A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to Deployment
We describe the evolution of a virtual assistant, called ATHENA, designed to support the capture, retrieval, and dissemination of knowledge…
Transformer Models for Text Summarization: A Comparative Study of BART, BERT, and RoBERTa
Text summarization refers to the task of condensing a document into a shorter version while preserving its key information. Automatic text…
Automatic bioinformatic software named entity recognition from literature
Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific litera…
Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer Attention
Standard multi-head attention (MHA) gives every head the same full causal context span, although heads can serve different contextual roles…
Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypotheses
Contemporary Large Language Models (LLMs) are increasingly aligned to suppress hallucinations, prioritizing factual retrieval over combinat…
Towards On-Board Implementation of ML-Based Helicopter Weight Estimator
This paper focuses on the implementation of a novel supervised Machine Learning model for estimating helicopter weight during takeoff, util…
Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models
Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large a…
Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life
Large language models (LLMs) and agentic AI systems are increasingly being explored for domain-specific maintenance and prognostics tasks,…
Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping
Rising immigration has intensified intergroup tensions in many countries. Traditional bias-reduction programs remain difficult to scale and…
Causal Inference under Interference with Learned Exposure Mappings
Exposure mappings are often assumed to be known in causal spillover analyses. In environmental settings, however, they are typically induce…
When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models
As language models move from drafting prose to running literature-search agents with tool calls, fabricated references are becoming easier…
Active Spiking Perception: The Membrane Potential as a Belief State for Anytime 3D Point Cloud Recognition
Spiking point cloud networks usually scan space in a fixed, input-agnostic order, which leaves the most distinctive resource of spiking com…
Incident-Data Robustness Analysis of the OWASP Top 10 for LLM Applications (2026): How a Community-Expert Ranking Holds Up Against a Large-Scale LLM Incident Corpus
The OWASP Top 10 for LLM Applications ranks the risks that a community of security practitioners judges most important. We ask a narrower q…
Mapping General-Purpose AI Governance in Twenty AI Middle-Power Jurisdictions
The most capable general-purpose AI (GPAI) models are mostly built in two jurisdictions, the United States and China, but the risks they ca…
Quantum Kernel Estimation for the Discovery of Early Lung Cancer Detection
Lung cancer screening with low-dose chest computed tomography reduces mortality, but its impact is limited by uptake, adherence, and manage…
Improved Confidence Estimates for Black-Box Large Language Models
Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). Existing methods, from verbalized con…
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability
Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and…
Does Marginal Coverage Guarantee Class-Conditional Safety for Zero-Shot VLMs Under Shift?
Split-conformal prediction provides marginal coverage under exchangeability and is increasingly used as an abstention layer for zero-shot v…
Fairness-Aware Network Embeddings: Methods, Applications, and Challenges
Network embedding methods learn low-dimensional representations of graph-structured data to support downstream tasks such as node classific…
Concentrated Liquidity Provision: a Reinforcement Learning Perspective
Automated market makers (AMMs) are a cornerstone of decentralised finance (DeFi). Constant product markets with concentrated liquidity, suc…
HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads
Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communicatio…
HiRA-CAM: Preserving Fine-Grained Spatial Relevance in Gradient-Based Visual Explanations
Deep Learning models can include billions of parameters or more, making it difficult to explain their internal transformations and outputs.…
SCAPE: Scenario-Conditioned Simulation-Augmented Policy Evaluation
Reliable performance evaluation is a central bottleneck for deploying robot-learning policies in real-world conditions. Real-world testing…
Longitudinal Bayesian Learning of Continuous Disease Position across the Alzheimer's Disease Continuum
Alzheimer's disease (AD) progresses as a continuous biological process, whereas most existing neuroimaging-based artificial intelligence me…
Are LLMs becoming similarly creative? Evidence from three years of models
Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance…
Measuring What a Specification Determines: A Formal Semantic-Block Model and an Execution-Judged Benchmark
This work introduces a formal semantic-block model for specifications and an execution-judged benchmark for evaluating specification qualit…
Accelerated Genetic Programming Hyper-Heuristics for Simulation-Based Scheduling via Agentic AI
Python is widely used in scientific research because it enables rapid development and provides rich ecosystems for data analysis, artificia…
In Two Minds about Lifelong Learning: Exploring Hemispheric Redundancy and Specialisation in Neural Models
Persistent intelligent systems require the ability to learn continually, but current machine learning approaches face significant challenge…
Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)
Stock market analysts and investors face a daily challenge: too much financial news, too little time. Manually reading and synthesizing hun…
When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language Models
Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natu…
CVSD-Reg: Cross-Modal Visual Semantic Prior Distillation for Robust LiDAR Registration
Learning-based global point cloud registration has achieved remarkable progress, yet its reliance on geometric representations makes existi…
Stream4D: 4D-Consistency for Streaming Autoregressive Diffusion Video Models
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local fra…
DraftFM: A FoundationModel for Day-Zero Drafting in Magic: The Gathering
Drafting a new Magic: The Gathering expansion begins before any pick from it has been observed: the complete card list is public, but the d…
VGI-BENCH: Probing Visual Intelligence in Video Generation Models
Recent studies suggest that video generation models can exhibit certain forms of zero-shot visual reasoning through generated frames. Yet r…
PEA-DPO: Perception-Enhanced Alignment Direct Preference Optimization for MLLMs Alignment
Direct Preference Optimization (DPO) has emerged as an effective approach for aligning large language models (LLMs) with human preferences.…
Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text Generation
LLM reasoning is stochastic, and so understanding a model requires grappling with the distribution of reasoning chains that it might produc…
DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories
Autonomous agents for machine learning experimentation must navigate heterogeneous repositories, repair training pipelines, and evaluate ca…
Escaping the Quicksand: A Call to Arms
Computing has been an astonishing success - but the accumulated technical debt exposes us all to huge costs in business and societal risk.…
Loreley: Repository-Scale Program Evolution with Quality-Diversity Search
Sequential agent search accumulates changes from its current champion but discards alternative branches; independent proposals preserve bre…
Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions
Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation…
Scale-Separated Conditioning for Style-Encoder-Free Diffusion Stylization
Reference-based diffusion stylization requires separating target geometry from transferable appearance. Existing tuning-based methods often…
A Locally Tokenized Generative Model for Robust Time-Series Watermarking
Watermarking is a central tool for provenance in generative models, yet its application to multivariate time series remains hindered by rel…
TempJail: Temporal Jailbreak Attack against Large Vision-Language Models via Subtitle Scheduling
Large vision-language models (LVLMs) have achieved remarkable progress in video understanding and reasoning. Despite extensive studies on t…
Learning to Beat: Phenotype-Guided Latent Flow with Regional Motion Priors for Biventricular Motion Synthesis
Full-cycle biventricular geometry is essential for characterizing cardiac function. However, dense and temporally consistent 3D+t biventric…
Question-Guided Evidence Acquisition for Multimodal Visual Question Answering
Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still…
Truncate Bad, Upweight Good: BoN-Style Distillation via Rank-Based Classification
Inference-time selection methods, such as Best-of-N, improve generation by sampling a pool of candidates and selecting the top-ranked compl…
GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic Manipulation
Multifingered grasping is a crucial robotic skill, but current deep-learning grasp planners often struggle to generalize to new objects bec…
Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay
Audited against causal ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signa…
Finite-Horizon Input-Output Dynamics of Minibatch Perturbations in AdamW
A minibatch can influence training beyond the update at which it is observed because AdamW stores past gradient information in its optimize…
CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning
Current dexterous grasp planners primarily optimize for physical stability, focusing on whether an object can be grasped rather than how it…
Distilling Aggregate Mobility Statistics into a Language Model Policy for Post-Event Crowd Simulation
Pedestrian simulators need a behaviour rule for every agent, but privacy usually limits the data for setting one to aggregate statistics, n…
An Irreducible Quantum Advantage in Aligning World Models with Reality
World models provide digital simulacra of the true world, allowing agents to be trained and tested before costly real-world deployment. At…
LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment
Low-Rank Adaptation (LoRA) is a prominent fine-tuning method for large models, achieving competitive performance with reduced memory overhe…
MileGPO: Milestone Inference with Local Evidence for Graph-Based Policy Optimization of Long-Horizon LLM Agents
Credit assignment is challenging in long-horizon agentic reinforcement learning, where supervision often comes only from final rewards. Exi…
Core-KAN: Continuous Vision Kernels with Kolmogorov-Arnold Networks
Conventional convolutional kernels are typically defined on fixed discrete grids, limiting their ability to accommodate heterogeneous local…
Adaptive Probabilistic Shielding by Learning MDPs for Safe Reinforcement Learning
Probabilistic shielding is a technique for safe reinforcement learning (RL). Typically, a static observer -- called the shield -- constrain…
Repo0: Design-Driven Zero-to-All Code Generation
Large language model agents have made substantial progress in code generation, yet most existing systems assume a predefined repository arc…
Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly…
A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queries
Patients often submit short, underspecified queries to healthcare chatbots that lack the patient-specific information needed to determine a…
Separating Covariate Shift from Mechanism Change with Two Discriminators: CJSD, a Conditional Discrepancy with an Exact Covariate-Concept Decomposition
Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn…
Evidence Before Expansion: Reuse, Spawn, or Defer in Lifelong Expert Pools
Streaming systems that maintain a pool of expert models must repeatedly decide whether to reuse an existing expert for arriving data, spawn…
Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models
Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We d…
MaliciousSkillBench: A Comprehensive Benchmark for Malicious Agent Skill Detection
Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration. This…
Towards Quantifying Benchmark Optimization in ASR Models
Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, ther…
Designing Human-mediated AI Guidance: Ready Together for Personalized Family Emergency Preparedness
Artificial intelligence (AI) systems are increasingly used across domains to provide personalized information, recommendations, and decisio…
Open-Vocabulary 3D Object Detection with Co-Distillation Discovery and Dual Guidance Robust Training
Recently, open-vocabulary 3D object detection (3D-OVD) has gained increasing attention for its ability to detect unseen objects in 3D scene…
An Inclusive and Lightweight Approach to Federated Continual Learning for Cultural Heritage
Artificial intelligence can support cultural heritage and digital humanities through large-scale retrieval and analysis of digitized collec…
EchoCoT: Extracting Hidden Chain-of-Thought from Large Reasoning Models
Hidden chain-of-thought (CoT) traces, especially those from frontier proprietary large reasoning models (LRMs), are valuable model assets.…
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts
Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However,…
SABET-QA: Temporal Knowledge Graph Question Answering
Question Answering over Temporal Knowledge Graphs (TKGQA) requires reasoning over time-sensitive facts, yet existing embedding-based method…
Evidence-Gated Task and Motion Planning with Vision-Language Models
Robots executing long-horizon manipulation tasks from natural-language instructions must reason about both semantic task structure and geom…
Towards Professional Tennis Styles for Humanoid Robots with Adaptive Motion Planning and Tracking
Humanoid robots have recently demonstrated promising capabilities in real-world ball sports. However, achieving professional motion styles…
Structured Affinity for Unsupervised Visual Class-Incremental Memory in Deep Artificial Immune Networks
Artificial immune networks (AINs) are naturally memory-forming systems, but conventional visual AINs often rely on flattened vector affinit…
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection
We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection. Harness optimization it…
A Standardized Framework for Machine Learning in Power System Protection
Studies of machine-learning-based power-system protection increasingly report near-perfect scores, yet the meaning of those scores depends…
Multi-Method Causal Evidence Synthesis: Ranking Candidate Drivers by Convergent Cross-Method Evidence from Observational Data
Practitioners inferring causality from observational data usually rely on a single method and treat its output as causal truth. Recent tool…
From Agent Behaviour to Agent-Friendly Documentation: An Empirical Study of How Coding Agents Discover, Read, and Write Technical Documentation
Technical documentation is written for human developers, but an increasing share of software changes is now authored by autonomous coding a…
Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference
Small language models are usually built like large ones and then squeezed onto a CPU afterwards. We did the opposite: we fixed the target f…
Prompt-Conditioned Channel Attention for Hierarchical Feature Modulation toward Anatomy-Agnostic Segmentation
Anatomically plausible segmentation remains challenging because of low contrast, ambiguous boundaries, and modality-specific artifacts. Int…
Growth Without Us: Machine Consumers, Corporate Circularity, and the Decoupling of GDP from Humanity after AGI
The standard objection to full automation is demand-side: if humans earn nothing, who buys the output? This confuses an accounting role wit…
Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge Internalization
Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inf…
Inducing Task Models from Computer-Use Traces
Naturalistic computer-use traces, passively recorded screenshots and mouse or keyboard actions, are a valuable resource for deriving symbol…
G-CARL: Grounded Checklist-Aligned Reward Learning for Patient-Oriented Medical Report Interpretation
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires…
Toward Greater Autonomy in Materials Discovery Agents: Unifying Planning, Physics, and Scientists
We aim at designing language agents with greater autonomy for crystal materials discovery. While most of existing studies restrict the agen…
GridCodex: A RAG-Driven AI Framework for Power Grid Code Reasoning and Compliance
The global shift towards renewable energy presents unprecedented challenges for the electricity industry, making regulatory reasoning and c…
Computational Phenomenology of Borderline Personality Disorder: A Comparative Evaluation of LLM-Simulated Expert Personas and Human Clinical Experts
Building on a human-led thematic analysis of clinical life-story interviews (> 150,000 words) with inpatients with Borderline Personality D…
Gen AI in Proof-based Math Courses: A Pilot Study
With the rapid rise of generative AI in higher education, understanding how students use AI is increasingly important. This exploratory stu…
ATBench: A Diverse and Realistic Agent Trajectory Benchmark for Safety Evaluation and Diagnosis
Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step int…
CharTool: Tool-Integrated Visual Reasoning for Chart Understanding
Charts are ubiquitous in scientific and financial literature for presenting structured data. However, chart reasoning remains challenging f…
Agent-First Tool API: A Semantic Interface Paradigm for Enterprise AI Agent Systems
As AI agents transition from research prototypes to enterprise production systems, the tool interfaces they consume remain rooted in human-…
The First Drop of Ink: Nonlinear Impact of Distracting Information in Long-Context Reasoning
As large language models are increasingly deployed in retrieval-augmented generation and agentic systems that accumulate extensive context,…
ChronoAgentic: A Code-based Multi-Agent World Simulator for Physically Grounded Simulation Construction
Video-based world models generate visually plausible rollouts, but since they infer dynamics in latent states, they enforce no explicit phy…
From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems
Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critic…
CADRE: Stable, Parameter Efficient Adaptation of Medical Vision Language Models with Bounded Forgetting and Prior Drift
Medical vision-language models (VLMs) such as BiomedCLIP generalize broadly, but adapting them to a clinical service is as much a safety pr…
FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills
Large language model agents can adapt to complex tasks by constructing workflows at inference time, but procedures discovered in one episod…
LLM Capability Limits: Static Emergence and Dynamic Boundary Control
Test-time emergence in LLM systems has a deployment boundary: additional computation can realize decisions already supported by the deploye…
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals…
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VL…
GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis
Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains su…
Learning-Based Speed Estimation from Accelerometer-Only Inertial Sensing
The proposed model, CarSpeedNet, estimates scalar vehicle speed from a window of three-axis smartphone acceleration, without gyroscope, whe…
Teacher-free Latent Self-distillation and Class-separable Representations for Lightweight IoT Attack Detection
Knowledge distillation (KD) has been widely used to improve lightweight AI models by transferring soft-label knowledge from a large teacher…
Towards Efficient Pareto Set Approximation via Mixture of Experts Based Model Fusion
Solving multi-objective optimization problems for large deep neural networks is a challenging task due to the complexity of the loss landsc…
Your Turn: At Home Turning Angle Estimation for Parkinson's Disease Severity Assessment
People with Parkinson's Disease (PD) often experience progressively worsening gait, including changes in how they turn around, as the disea…
Virtual Sensing to Enable Real-Time Monitoring of Inaccessible Locations & Unmeasurable Parameters
Real-time monitoring of safety-critical interior states is an open problem across energy, environmental and industrial systems where direct…
Improving Requirements Classification with SMOTE-Tomek Preprocessing
This study emphasizes the domain of requirements engineering by applying the SMOTE-Tomek preprocessing technique, combined with stratified…
Regressor-Guided Image Editing Shifts Emotion and Disengagement Timing in Social Media
Internet overuse is a widespread phenomenon in today's digital society. Existing interventions, such as time limits or grayscaling, often r…
HiFi-KPI: A Dataset for Hierarchical KPI Extraction from Earnings Filings
Accurate tagging of earnings reports can yield significant short-term returns for stakeholders. The machine-readable inline eXtensible Busi…
ReynoldsFlow: Physics-Inspired Spatiotemporal Flow Representation for Video Understanding
Video understanding has largely relied on deep spatiotemporal architectures, including 3D convolutional networks and optical flow (OF) base…
Learning-Augmented Power System Operations: A Unified Optimization View
With the increasing penetration of renewable energy and inverter-based resources, traditional physics-based power-system operation faces gr…
The Thousand Brains Theory 2.0: An Extension for the Long-Range Connections of the Neocortical Heterarchy
Vernon Mountcastle hypothesized that the basis for intelligence in mammals is the replication of a general computational unit, the cortical…
PrefixAgent: An LLM-Powered Design Framework for Efficient Prefix Adder Optimization
Prefix adders are fundamental arithmetic circuits, but their design space grows exponentially with bit-width, posing significant optimizati…
The Basic B*** Effect: The Use of LLM-based Agents Reduces the Distinctiveness and Diversity of People's Choices
Large language models (LLMs) increasingly act on people's behalf: they write emails, buy groceries, and book restaurants. While the outsour…
DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values
Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks oft…
FMT$^{\mathrm{X}}$: Lazy Wavefront Search for Dynamic Replanning
FMT$^{*}$ plans efficiently in static worlds by expanding a cost-ordered wavefront and collision-checking lazily, but its single-pass unvis…
TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
Time series reasoning is crucial to decision-making in diverse domains, including finance, energy, and scientific discovery. While existing…
SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation
Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily rel…
The Bidding Games: Reinforcement Learning for MEV Extraction on Polygon Blockchain
In blockchain networks, the strategic ordering of transactions within blocks has emerged as a significant source of profit extraction, know…
MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal Prostate MRI Segmentation
Active Surveillance (AS) is a treatment option for managing low and intermediate-risk prostate cancer (PCa), aiming to avoid overtreatment…
A Physics-Informed Neural Network Approach for UAV Path Planning in Dynamic Environments
Unmanned aerial vehicles (UAVs) operating in dynamic wind fields must generate safe and energy-efficient trajectories under physical and en…
PACT: Phenotype-Aware Contrastive Team Representation for Multi-Phenotype Grouped Ad Hoc Teamwork
Learning to collaborate with various unfamiliar teammates poses a great challenge in the domain of multi-agent systems. Existing ad hoc tea…
WaveVerif: Acoustic Side-Channel based Verification of Robotic Workflows
In this paper, we present a framework that uses acoustic side-channel analysis (ASCA) to monitor and verify whether a robot correctly execu…
Towards Audio Token Compression in Large Audio Language Models
Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate t…
Extended to Reality: Prompt Injection in 3D Environments
Multimodal large language models (MLLMs) have advanced the capabilities to interpret and act on visual input in 3D environments, empowering…
Qworld: Question-Specific Evaluation Criteria for LLMs
Evaluating large language models (LLMs) on open-ended questions is difficult because response quality depends on the question's context. Bi…
Power Couple? AI Growth and Renewable Energy Investment
Artificial intelligence (AI) and renewable energy are increasingly being described as a \mbox{``power couple,''} based on the idea that rap…
Generative AI Use in Entrepreneurship: An Integrative Review and an Empowerment-Entrapment Framework
Despite the growing use of generative artificial intelligence (GenAI) in entrepreneurship, research on its impact remains fragmented. To ad…
VISD: Enhancing Video Reasoning via Structured Self-Distillation
Training VideoLLMs for complex reasoning remains challenging due to sparse sequence level rewards and the lack of fine grained credit assig…
MKG-CARE: Case-Aware Reasoning with Multimodal Knowledge Graphs for Explainable Medical Image Diagnosis
Medical image diagnosis has achieved significant progress with deep learning, yet existing methods often rely on isolated visual evidence a…
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
Concept formation in transformer language models is a depth-extended process, not a single-layer event: a concept becomes separable across…
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically ev…
Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams
A concept probe is only as reliable as the layer it is taken from. Probing at a fixed late layer, or at the peak of a separation curve, ign…
Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization
If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to me…
The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction
Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other fea…
Can Predicted Dynamics Exist in the Physical World?
Can learned state-action proposals exist in the physical world? To filter infeasible commands before execution, policies are often wrapped…
GEM: Geometric Erasure by Contrastive Velocity Matching in Rectified Flows
While the rapid adoption of multimodal generative models offers immense potential, it has also increased the risks of harmful content synth…
MPCoT: Reward-Guided Multi-Path Latent Reasoning for Test-Time Scalable Vision-Language-Action
Vision-Language-Action (VLA) policies remain brittle in long-horizon and high-uncertainty control, where one-pass action decoding provides…
Right Family, Wrong Skill: Benchmarking Risk Exposure in Agent Skill Retrieval
Agent skill libraries are becoming routable software assets: a retrieved skill can contribute instructions, scripts, resource bindings, and…
A Tool to Map AI Programs in the U.S.: A Snapshot from April 2026 and an Analysis of Requirements for AI Majors and Minors
In this work, we locate and analyze existing undergraduate Artificial Intelligence (AI) programs in the United States in Spring 2026, creat…
PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty
Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Ob…
Every Step of the Way: Video-based Parkinsonian Turning Step Counting
As a prominent symptom of Parkinson's disease (PD), turning impairment is evaluated through parameters such as turning angle, duration, and…
SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects (Early Version)
Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction,…
Drift-Adaptive ICU Intervention Prediction: Freezing the Physiological Encoder for Auditable Model Updating
Clinical decision support degrades as treatment protocols evolve, but the obstacle to updating a deployed model is governance as much as ac…
A Distributional Robustness Margin For Pathology Foundation Models
Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut lear…
Anatomy Contextualized Adaptation of CT Foundation Models
CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-…
FinVerse: Financial Time-Series Benchmark
As time-series foundation models have emerged, the need for benchmarks that can evaluate their forecasting ability in meaningful ways has b…
Socioduality: A Relational Process Framework for Human-AI Interaction
Human-AI research often evaluates individual capabilities, joint performance, or final outputs, but these approaches can lose the interacti…
A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluati…
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance
Specifying a penalty can turn a legal obligation into a cost-benefit calculation that favors violation. We show that this enforcement infor…
「Claudeの使い方」を無料で学べる公式サイト登場 「Code」「Cowork」などサービスごとに解説
米Anthropicは、無料の学習サイト「Claude Academy」を公開した。AI技術の基本や、AIサービス「Claude」関連製品の使い方を解説している。
Googleのオープンモデル「Gemma」、累計10億ダウンロード超 GitHubに公式ディレクトリ公開
Googleは、オープンモデル「Gemma」ファミリーの累計ダウンロード数が10億回を突破したと発表した。派生モデルは10万種を超え、公式リポジトリ「Awesome Gemma」をGitHubで公開。宇宙空間での衛星データ解析や新規のがん治療経路の発見など、多様な活用事例を紹介…
macOS版ChatGPT、Appleの「メッセージ」と連携 会話検索や下書き、送信に対応
OpenAIは、macOS版ChatGPT向けにAppleの「メッセージ」アプリと連携するプラグインを公開した。CodexやChatGPT Workのチャット上で過去の会話検索や下書き作成、送信が可能になる。誤送信防止のため都度承認フローを備える。Appleシリコン搭載Mac向…
AI data startup Micro1 reaches $500M gross run rate amid AI training boom
Surging demand for AI training data is driving rapid growth for the startup and its rivals.
「Fable禁止」で仕事が止まったあの日々を振り返る 日本企業が取るべき「脱・単一モデル」戦略
米政府の輸出管理で「Claude Fable 5」の提供が突如停止し、業務が止まった実体験を基に、単一AIモデル依存の地政学リスクを指摘。オンプレミス化やマルチモデル統合基盤など、日本企業が取るべき分散戦略を整理する。
データをつなぎ、AI活用へ――オートデスクが示す設計/製造DXの未来像
オートデスクは「Design & Make Summit Japan 2026」を東京都内で開催した。本稿では、米Autodeskのビック・ベダンサム氏による基調講演から、AI時代の製造業に求められるデータ基盤と設計/製造の変革について紹介する。
業務標準化の手間を9割減 三菱UFJ銀行は生成AIに「業務知識」をどう教えた?
海外事務の標準化を進めている三菱UFJ銀行。熟練従業員に頼ってきた業務プロセスの精査を生成AIに置き換える過程で直面したのは、AI特有の誤情報や回答のバラつきだった。同行はこれをどう克服したのか。
Snowflakeが過去最高業績 CEOが「他社との差別化は容易になった」と語るワケ
競合が同じ方向に走り出した今こそ、差別化はむしろ容易になっている――。SnowflakeのCEOが年次カンファレンスでこう言い切った根拠はどこにあるのか。AIエージェント時代のSnowflakeの戦略に迫る。
「チャピる」「ギュられる」って何? 今年流行った「就活用語」にAI関連ワード マイナビ調査
マイナビは、2027年卒の就職活動で流行した用語を発表した。米OpenAIのチャットAI「ChatGPT」を示す「チャッピー」に加え、「チャピる」「ギュられる」といったAIに関する新語が登場した。
OpenAI is gaining on Anthropic with business users, new data indicates
Businesses are willing to flop back and forth as each lab releases new models, volatility that should give both companies' investors pause…
ChatGPT can now send texts for you with new Apple Messages plug-in
Ever wanted someone else to do your texting for you? ChatGPT is being offered up as an automated text scribe via a new Apple Messages integ…
ChatGPTに「おすすめの○○は?」 実は答えが決まっているらしい:893rd Lap
ChatGPTに「おすすめの○○は?」と聞けば、いくつかのブランドや商品を教えてくれる。では、その候補はどうやって選ばれているのだろうか。どうやらChatGPTは、検索を始める前から「この分野ならこれ」と、ある程度の候補を持っているらしい。
「たった14人」の挑戦から7兆円の逆転劇へ ラピダス小池社長の「TSMCとは戦わない」2ナノ半導体の勝算
世界の半導体市場を台湾TSMCが席巻する中、7兆円規模の国家プロジェクトとして最先端「2ナノ」の量産化に挑むのがラピダスだ。同社はTSMCとの規模の勝負を避け、設計から前後工程を一棟で完結させる「RUMS」による多品種生産で勝負する。「たった14人」の同志でスタートした原点から…
「孫さんはOpenAIだが、僕はAnthropic」 SBI北尾会長が語る「AI投資5億円→増収27億円」の勝算
SBIホールディングスが生成AI「Claude」を開発する米Anthropicとの全社提携を発表した。当面の最優先戦略にAI化を掲げ、社外から専門人材を起用。SBI証券では顧客対応のAIエージェント開発に5億円を投資し、口座再活性化などを通じて年27億円の増収を見込む。「孫正義…
「Gemini Notebook」で利用者10倍 シニア社員をAIヘビーユーザーにした首都高の考え
「生成AIを何に使えばよいかわからない」という理由により、生成AIの活用が停滞してしまう企業は多い。安全を最優先する故に慎重な組織風土であった首都高速道路でも同様の課題を抱えていた。しかし同社では「Google Gemini」を起点としたある工夫により、劇的に活用状況を改善した…
Slack、AIとチームで協働する「Slack Code」を発表 ClaudeやDevinを専用チャネルで操作
Slackは、AIコーディングエージェントと協働するための新機能「Slack Code」を発表した。メンションで専用の「コードチャネル」が自動生成され、計画やコード差分、プレビューを確認しながら指示できる。ClaudeやDevinなど複数社のエージェントに対応し、人間の承認を経…
カルビーが挑むジャガイモ収量の限界――自社開発AIでサプライチェーン最適化
「ポテトチップス」や「じゃがりこ」など、カルビーの主力商品に欠かせないばれいしょには、収穫量の限界がある。後手の意思決定から脱却すべく、同社はAIを活用した全社最適シミュレーター「C-BOSS」を自社開発。いかにして現場定着の壁を越え、データに基づく攻めのサプライチェーンを構築…
OK, can we actually cool data centers with our pee?
Jason Kelce joked that people should cool data centers with their pee, rather than potable water -- but his suggestion is not completely lu…
GoogleはAI競争に負けたのか 「最強のAI」ではなく「AIの“電力網”」を選ぶ賭け
GoogleからAI研究の中心人物が相次いで去った。「Geminiは終わった」という見方に対し、「最先端ではなく、AIを社会全体に行き渡らせる“電力網”で勝つ賭けだ」という別の解釈もある。電気の歴史になぞらえながら整理する。
Google gives publishers a new way to fight AI-driven traffic losses
Google is giving publishers a new button that lets readers make them a preferred source across Search, Discover, and Google News, potential…
Runlayer, Rippling drop lawsuits — but the brouhaha is still a cautionary tale for founders
Runlayer and Rippling have dropped their lawsuits. No money was paid. Rippling celebrated by releasing a competing product.
Linkdaze’s smart calendar is built to run a household, not just track a schedule
Linkdaze's smart digital calendar stands out for not putting its features behind a paywall, including an AI meal planner tool.
Grok keeps sending gibberish responses to users
Affected users told TechCrunch they were using Grok Lite, and noticed the issues as early as Wednesday morning.
A third of web pages published since ChatGPT’s launch show signs of AI authorship, study finds
ChatGPT and other AI models are now authoring and editing much of the new web.
Ramp launches its own AI model router, called Router
Ramp has launched its own AI model routing service, dubbed Router, that lets users and companies use and switch between various large langu…
Meta brings Pocket, an app that lets you vibe-code and share games, to US users
Meta is bringing Pocket, its experimental AI-powered app for creating and sharing interactive games, to users across the U.S. after quietly…
Inertia Enterprises finds a way to make its fusion fuel fast
Fusion power startup Inertia Enterprises reduced the fuel filling process from a week to just a few hours. It's one of 10 hurdles the compa…
2026-08-20(272件)
Meta AI’s new Mac app wants you to talk to your apps
The company said that the dictation feature works across all apps, just like other tools such as Wispr Flow, Superwhisper, and Monologue.
Binance now lets AI agents trade, but keeping them in check is largely up to users
Binance's Agent OS works with tools such as ChatGPT, Claude Code, and Cursor.
「ロボットのChatGPTモーメントが近づいている」 中国UnitreeのCEO、世界ロボット大会で言及
Unitreeのワン・シンシンCEOは世界ロボット大会で、ロボットの「ChatGPTモーメント」が近づいていると発言した。
日本精工が「国産人型ロボ」開発を後押し スタートアップのアトムと協力、アクチュエータの検証など
人型ロボットを開発するスタートアップのアトムは、日本精工(NSK)と国産人型ロボットの開発・実装に向け、戦略的パートナーシップに関する基本合意書を締結したと発表した。
Introducing AI Futures
Introducing AI Futures, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual free…
エイベックス松浦会長、noteの“バズり記事”を「ほぼAI」で作成 「僕の60年分のデータを入れた」
エイベックスの松浦勝人会長は、8月13日から「note」に投稿している記事について、ほとんどAIで作成していたと自身のXアカウント(@maxmatsuuratwit)で明かした。
Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions
This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and sh…
Position: Profiling Game Worlds by Transition Complexity
Game world modeling (GWM) and reinforcement learning (RL) are often confounded because research papers rarely quantify how difficult the un…
Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges
We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational age…
Position: Behavioral Systems Require Behavioral Tests
Artificial agentic systems increasingly operate as behavioral systems by interacting with dynamic environments, pursuing goals, and adaptin…
Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models
The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate strategies for effective downstream govern…
A Metamorphic Artificial Age Score Decision-Support Prototype for Flight-Log-Based Drone Propeller Health Monitoring
Drone propeller faults can create safety and reliability risks when their effects are distributed across multiple flight-log channels rathe…
Position: Multi-Agent Systems Should Prioritize Concurrency Control
LLM-based multi-agent systems (MAS) promise scalable collaboration, yet adding agents often reduces reliability. This position paper argues…
FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management
Investment management is a high-stakes domain in which agentic AI systems must do more than generate plausible text. They must retrieve poi…
Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New Perspective
Large language model (LLM)-based agents are increasingly becoming self-evolving systems that persist across interactions, maintain memories…
Emergence of Agentic AI: A Review on Evolution, Background, Working Principles, Applications, Adoption Factors, and Future Research Directions
Agentic AI is gaining new insights and advancements in the field of Artificial Intelligence, fostering significant potential to enable rapi…
Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry
Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem…
Position: AI Leaderboards Are Underserving the Global South: A Case Study from India
This position paper argues that AI leaderboards are structurally ill-suited to serving the Global South because they lack independent gover…
Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs
Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-Engli…
Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas Analysis
Reliable transformer fault diagnosis is essential for maintaining power system stability. The IEEE Key Gas Method (KGM), a widely utilized…
Improving Rural Medication Safety with AI: A Scoping Review
Introduction: Medication errors (MEs) represent a significant threat to global healthcare systems, contributing to patient harm. Introducin…
FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud
Conversational agents now act for end users through tools while holding access to customer databases and internal policy documents that a c…
Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman Urdu
It is challenging to detect hate speech in Low Resource Languages (LRLs) because of the absence of annotated data, the informality of its l…
RDFdL: Integrating RDF with Differential Dynamic Logic
Knowledge graphs modeled in RDF are powerful for describing static knowledge, but they cannot capture or reason about the dynamic behavior…
Adversarial Review: Structured Disagreement for Grounded Agentic Code Review
Early multi-agent LLM systems often used role-separated teams, yet scaling agent count yields diminishing returns on repository-level codin…
Looped Language Models Improve Compositional Tool Calling
Looped language models have shown promising results on reasoning benchmarks, yet their potential for agentic tool use remains largely unexp…
On the Triangle Inequality for the Jaccard Distance in Arbitrary Lattices
This paper presents new theoretical results on generalizing the Jaccard distance for lattices and real valuations. We demonstrate that when…
GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence Networks
Genomic analysis on viruses such as SARS-CoV-2 variants: Beta, Gamma, Delta, and Omicron is heavily dominated by classical bioinformatics m…
Redakto - The Incognito Tab for LLMs
Large Language Models (LLMs) are being increasingly used in everyday applications. A major challenge in the context of LLMs or Artificial I…
Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement Study
Serving a 235B-parameter Mixture-of-Experts (MoE) model on a single 8 GB GPU is bottlenecked not by compute but by memory bandwidth: decode…
Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
The extraction of structured information from unstructured documents represents a critical component of digital transformations in all sect…
The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations
LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has becom…
SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured Decomposition
LLM-as-judge evaluation reduces response quality assessment to a single holistic A/B preference choice, providing no mechanism to isolate w…
ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
Current evaluation of computer-use agents is split between long-horizon workflow benchmarks and atomic GUI-grounding tests. This leaves an…
Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow Repair
Machine-verifiable workflows produce governance records linking a task contract, model attempt, verifier decision, accepted output, and tar…
Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme
When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy.…
A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations
AI code agents are increasingly deployed to resolve real software issues, yet their reliability under superficial code variations remains p…
When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress Classification
Wearable stress classifiers can achieve strong average performance while failing completely for a particular individual. On WESAD, a Random…
Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal Abstractions
Combinatorial scheduling poses a significant challenge for language models, requiring them to identify feasible solutions within exponentia…
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where acti…
UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal Retrieval
Universal multimodal retrieval aims to support diverse instruction-aware retrieval tasks, demanding both efficient corpus-scale matching an…
Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption Retrieval
Dense-caption retrieval has recently been improved by introducing segmentation, edge maps, LLM-filtered captions, and cross-modal modules i…
Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical Lesson
Industrial explainable-recommendation systems built on LLMs incur a substantial serving cost: each request triggers an LLM generation, with…
FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems
Large language models are increasingly used to support financial operations, but their apparent reasoning performance can depend on whether…
Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement
Modern e-commerce platforms often operate search, recommendation, personalization, and CRM systems independently, limiting opportunities fo…
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each t…
Can a Lightweight Multimodal Model Estimate LLM Reasoning Performance? A Study for Compute-Optimal Document Inference
Uniformly allocating inference reasoning budgets to LLMs is expensive and prone to over-thinking penalties; especially in document tasks wh…
CTIFoundry: An Agent-Native Corpus Scaffold for Cyber Threat Intelligence
Cyber threat intelligence (CTI) is increasingly consumed not by human analysts but by LLM agents that compose multi-step investigations at…
Preference Reasoning under Indeterminacy in Large Language Models
As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coord…
Candidate-Fate Accounting for Transparent Sensor Diagnostic Pipeline Search
Industrial sensor diagnostics relies on preprocessing, representation, and classification pipelines, making automated pipeline search usefu…
Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction
Amid concerns that generative AI may standardize art interpretation, this paper examines whether LLM-based interaction can support plural a…
RTPO: Reverse-Turn Policy Optimization for Stabilizing Agentic RL Training
Training multi-turn agentic workflows with reinforcement learning (RL) enables large language models to perform complex reasoning, use exte…
Competence, Not Accuracy: A Diagnostic for Reference-Free Judge Gates in Skill Optimization
Text-space skill optimization adapts a frozen agent by evolving a natural-language skill document, accepting each candidate through a valid…
A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation
This paper proposes a multi-agent framework built on CrewAI [1] for conversational business intelligence. Five specialized AI agents operat…
Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots
Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generati…
Pairwise Logical Selection of Enthymeme Completions under Semantic-Link Uncertainty
Arguments often omit premises or claims, forming enthymemes. We study pairwise logical selection between two candidates for the omitted com…
Verifiable abstention makes AI leak diagnosis accountable in water distribution networks
Utilities lose a substantial share of treated water to leakage, yet rarely trust artificial-intelligence localizers to dispatch crews: gues…
ORBITER: Conflict-Aware Decision-Making for Agentic Last-Mile Delivery
Last-mile delivery aims to handle dynamically arriving orders with couriers while modeling complex spatial and temporal correlations. Recen…
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents
Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries no…
DentAgent: Evidence-Centric Multi-Agent Coordination for Multimodal Dental Reasoning
Oral diseases affect billions of people worldwide, underscoring a pressing need for accurate and reliable dental assessment that integrates…
Training-Free Inference-Time Self-Reflection and Cost-Bounded Early Stopping for Large Language Models
Reinforcement-learning training of reasoning LLMs (e.g., GRPO) is expensive and requires a controllable environment, committing every contr…
Syntactic Simplification of OWL Class Expressions
Class expression learning often produces complex OWL class expressions that are difficult to interpret and reason over. However, by followi…
\textsc{TestifAI}: Tomography-Based Testing for Deep Learning Systems
As AI systems are increasingly deployed in safety-critical application domains (e.g., autonomous driving), associated risks increase too. D…
Breaking the weakest link to evade vision language models
Vision Language Models (VLMs) have recently emerged as a critical component of multimodal AI systems, enabling joint reasoning over visual…
A Theory of Post-hoc Debate Judgement
Debates have recently emerged as a useful methodology for agentic AI to improve performance as well as to aid explainability and user engag…
Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models
Accurately extracting nuanced, contextualized data from research articles is laborious and time intensive. Here, we investigate the perform…
Adaptive Memory and Reflection Multi-Agent System for Medical Question Answering
Accurate and responsible medical question answering (QA) is important in healthcare, where complex cases require factual knowledge and nuan…
Eureka: Task-Conditioned Meta-Agent Orchestration for Scientific Discovery
We present Eureka, a task-conditioned Meta-Agent architecture that compiles long-horizon tasks into dynamic obligation graphs with explicit…
What is Missing from AI Post-Training AI: An Empirical Analysis
Large language model (LLM) agents can now post-train an LLM end-to-end. They can write code, launch training, evaluate checkpoints, and imp…
Robust Risk Under Evolving Uncertainty: A Wasserstein Counterpart of the Entropic Value-at-Risk
An agent still learning its environment should be cautious while ignorant and bold once confident. The entropic value-at-risk captures this…
Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering
When an expert corrects an LLM assistant's error, the correction usually dies with the session, and the error class returns. I argue this i…
Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems
Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue th…
Beyond the Transcript: Detecting Covert Co ordination in Latent Multi-Agent Communication
Language-model agents can communicate through continuous hidden states that are invisible in public transcripts, creating opportunities for…
SuTRA : Structurally-Unified Tokenization with Root Awareness
Existing subword tokenizers optimize statistical compression but ignore morphological structure, particularly the relationship between root…
Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without Retraining
Instruction-tuned models often refuse harmful requests in English but comply with the same requests in Yoruba, Igbo, Igala, and Hausa. This…
Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four Modalities
Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find…
Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs -- the tendency to favor one's own outputs -- raises grow…
Abliteration Mitigation via Refusal Aliases
Abliteration, the removal of refusal capabilities from large language models by projecting weight matrices orthogonal to an extracted refus…
NE-BERT: A Multilingual Language Model for Nine Northeast Indian Languages
Large pretrained language models have demonstrated remarkable capabilities across diverse languages, yet critically underrepresented low-re…
Backdoor Learning in Language Models and Vision-Language Models
Recent advances in deep learning have significantly enhanced the capabilities of Natural Language Processing (NLP) and Vision-Language Mode…
Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog Systems
Key-value (KV) caching is essential for efficient autoregressive inference in transformer based dialog systems, yet existing strategies tre…
Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)
AI systems now shape how hundreds of millions of people learn about cultures other than their own. When someone asks one of these systems a…
DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language Models
Background: Mechanistic elucidation of traditional Chinese medicine (TCM) compound formulas remains a central challenge in the modernizatio…
StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web Data
StocksTalk is a voice-enabled conversational system for transforming spoken financial screening requests into executable and validated stru…
Different Facets of Verbalised Overconfidence: an Interpretability Study
Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled r…
Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals
We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity…
Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation
Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior w…
Accurate Decoding of Natural Sentences from Non-Invasive Brain Recordings
Restoring communication for people who have lost the ability to speak or move after a brain injury is a major challenge. While intracranial…
Temporal Multi-Signal Fusion for Token-Level Hallucination Detection
Token-level hallucination detectors score each token independently from a single signal, and fail exactly when the generating model is conf…
Global Index on Responsible AI 2026 : Conceptual Framework and Methodology
This report presents the methodology of the Global Index on Responsible AI (GIRAI), 2nd Edition. This edition refines the 1st Edition by st…
Language Models for Portuguese: A Systematic Mapping Study
In recent years, the rapid development of language models has transformed the field of Natural Language Processing through a wide range of…
The Deontic Gap: Large Language Models and the Modal Language of Obligation
Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersona…
Entropy-Constrained Adaptive Stochastic Quantization
Adaptive stochastic quantization (ASQ) is a recently introduced quantization approach that optimizes the Mean Squared Error (MSE) for a giv…
TokenPowerSandbox: Evidence-Gated CPU-First Screening for Energy-Aware LLM Serving
Energy-aware LLM serving requires comparing configurations under realistic request shapes, yet exhaustive target-GPU profiling is costly an…
How Quantum Is the Advantage? A Fair, Calibration- and Noise-Aware Benchmark and Attribution Audit of Quantum Machine Learning for Network Intrusion Detection
Quantum machine learning (QML) for network intrusion detection (NIDS) is routinely reported to reach near-perfect accuracy, yet the most ri…
When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators
LLMs have been increasingly used to catch data quality issues automatically, but we know very little about how consistent these judgments a…
Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety Evaluation
Safety evaluations of large language models (LLMs) predominantly rely on text-based adversarial prompts, potentially overlooking vulnerabil…
What Can Artificial Intelligence Learn from Medicine? Generative Analogies and Reliable Machine Learning Systems
In the past few years, machine learning (ML) has been widely (and to an extent, successfully) implemented in medicine. However, uncertainti…
A systematic review of machine learning techniques to address diagnosis and treatment of autism: challenges and opportunities
Autism spectrum disorder (ASD) is a developmental disability characterized by challenges in social interaction and communication. As the ca…
Bound-Aware Per-Organ Recall Risk Control for Multi-Organ CT Segmentation under Clinical Domain Shift
Distribution-free risk control adds organ-specific recall guarantees to frozen segmentation. We calibrate per-organ thresholds for an AMOS-…
GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction
Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplie…
Bidirectional representational alignment between biological and artificial neural networks
Recent work has shown that representational alignment between biological and artificial neural networks is asymmetric: model representation…
Visual-Prompt Guided Wildlife Instance-Level Recognition
Fine-grained wildlife re-identification remains a challenging area in research. Current state-of-the-art approaches apply a detection and r…
How AI Prompts Can Teach Us About the Structure of Human Behavior
We introduce a general, easy-to-implement AI-based method for studying the structure and complexity of human behavior. We assign a large la…
SeisEvo: Evolution of Seismic Data Reconstruction Algorithms by Agents
Classical seismic data reconstruction relies on manually designed structural priors and iterative operators, whose coupled design space is…
What Makes Software Issue Resolution Tasks Difficult for Agents?
Background. Advances in agentic systems are simultaneously, and rapidly, saturating benchmarks. Despite this often discussed phenomena, ben…
Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements
An increasing number of scholars use AI to measure variables they subsequently include in downstream analyses. Although AI-measured variabl…
FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation
As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently…
FedCoRe: Target-Adaptive Completion for Missing Modalities in Healthcare Federated Learning
Federated multimodal models often assume every site has every modality, although hospitals differ in access to EHRs, chest radiographs, and…
From Inference to Adaptation: A Unified Optimal Transport View of Vision Language Model
Vision-language models (VLMs) have demonstrated remarkable zero-shot capabilities yet remain sensitive to real-world distribution shifts du…
Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring
Persistent acoustic monitoring can detect machine faults without physical contact, but always-on inference is constrained by power, latency…
Coupled-cluster molecular properties across the main group that extrapolate beyond training size
Coupled-cluster theory defines the accuracy standard for molecular electronic-structure properties but scales too steeply for routine appli…
Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents
Tool-using large language-model agents can complete a task while exercising authority that the user did not grant or the task does not need…
One Gate Is Not Enough: Composing Stateful Pre-Action Controls for Agentic AI
Agentic AI systems take consequential actions governed by more than one pre-action control at once: authority, resource, and evidence gates…
Selection, Recombination, or a Fresh Solve? A Candidate-Free Control for Single-Pass Test-Time Aggregation
When every candidate is wrong, correct-candidate selection is unavailable, yet the aggregation call can still solve the problem afresh. A c…
TTSD-FAR: Test-Time Self-Distillation with Fisher-Anchored Restoration for Missing-Modality Emotion Recognition in LVLMs
Large video-language models (LVLMs) have shown remarkable performance on multimodal tasks like multimodal emotion recognition (ER) in the w…
LEDGER: Claim-to-Evidence Trace Graphs for Auditing LLM Agents
Large language model (LLM) agents can now carry out long-horizon technical workflows involving complex tool use, code execution, file edits…
Vector Symbolic Policy Gradient
We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hyp…
Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B
Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-se…
Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage
Modern mental healthcare faces a critical shortage of senior supervisory oversight, leading to a "supervision gap" where novice therapists…
Formal Verification of Romanov's Triplet Logic: A Verified Filter for Sliding-window 3-CNF with Application to Structured Formulas
We present the first mechanised formalisation of Romanov's Triplet Logic (TLS) in the Rocq proof assistant. TLS is a triplet-based combinat…
ERASE: EaRly bAckpropagation SchEdule for Faster Training of Modern Recommendation Systems
Lightweight proxy models enable rapid experimentation without repeatedly training frontier-scale systems, but their small kernels often lea…
Coverage-Driven RTL Assertion Generation with Formal Exploration and Neuro-Symbolic Refinement
Hardware functional verification relies on high-quality assertions to expose design bugs and establish confidence in Register Transfer Leve…
Partition the Support, Reconstruct the Residual: Training-Free Sparse Attention for Video Generation and World Models
Training-free block-sparse attention can accelerate video transformers, but row-wise attention concentration does not by itself specify an…
Physics-Unrolled Neural Operator for Wireless Field Modeling
Radio maps are essential for wireless decision-making tasks such as access-point placement, coverage planning, and localization, but their…
Science Done on a Machine by a Machine: AI Agents in Computational Chemistry
We are witnessing an explosion of agentic systems for computational chemistry simulations: from half a dozen in 2024 to a dozen in 2025, an…
OptiModNet: A UNet-Transformer Hybrid with Grouped-Query and Channel Attention for Optic Disc and Cup Segmentation
Precise segmentation of the optic disc and cup is critical for the early detection and diagnosis of glaucoma. However, achieving consistent…
GCNO: Gramian Chebyshev Neural Operator for Physics-Based Compression of Wireless Channels
Large antenna arrays allow wireless systems to serve more users and achieve higher data rates, but they also make channel feedback expensiv…
Prior-Conditioned Gaussian Discriminants for Generalizable AI-generated Image Detection
Diffusion-based generators have made synthetic images ubiquitous, but detectors often fail under simultaneous shifts in generator, prompt/s…
DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents
Equipping Large Language Models (LLMs) with multi-turn tool-calling capabilities is essential for building autonomous agents. However, prog…
Evaluating and Explaining Prompt Sensitivity of LLMs Using Interactions
The remarkable capabilities of large language models (LLMs) are often undermined by their instability. Even subtle and semantically irrelev…
CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks
Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human…
Performance Drift Detection in Machine Learning as a Service (MLaaS) for IoT Environments
Machine Learning as a Service (MLaaS) is a powerful cloud paradigm enabling data-driven intelligent applications in Internet of Things (IoT…
MorphoGP: A Nonparametric Framework for Predicting Equilibrium Beach Profiles Under Tidal Influence
The prediction of equilibrium beach profiles under tidal influence is of fundamental importance for sustainable coastal development, inform…
The Role of Grid Cells in Reducing Spatial Aliasing in Hippocampal Place Representations
Spatial aliasing occurs when two or more distinct locations produce highly similar place-cell representations, primarily due to environment…
MR-IQA-2: Faithful Image Quality Reflection via Fine-Grained Credit Assignment
Multimodal large language models (MLLMs) have shown strong potential for image quality assessment (IQA) by improving consistency between qu…
From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning
Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge re…
OmniHandwritingOCR: A Diagnostic Benchmark for Evaluating Multimodal LLMs in Handwritten OCR Scenarios
Multimodal large language models (MLLMs) are increasingly used as OCR systems in document and knowledge-processing pipelines, but their abi…
Denoising-Aware Inversion: Revealing Privacy Risks in Noise-Protected Text Embeddings
Dense text embeddings are widely used in data mining, retrieval, and downstream machine learning systems due to their compact and semantica…
Change Point--Aware Evaluation and Re-Calibration of PPG-Based Blood Pressure Estimation
Non-invasive continuous blood pressure (BP) monitoring using photoplethysmography (PPG) is a promising alternative to cuff-based measuremen…
Orienteering Problem with Uncertain Time-Varying Rewards: Framework and Benchmark for Everyday Service Robotics
We present the orienteering problem with uncertain time-varying rewards (OP-UTVR), a novel variant of the orienteering problem (OP). While…
Aslema at NADI 2026: Augmentation through Fewshot for SLU
We present Aslema, our system for NADI 2026 Shared Task 5, which consists of two subtasks: intent recognition and slot filling. We evaluate…
Europe's Climate Ambition Under Scrutiny: Evidence from Deep Learning Emission Projections
The European Union has committed to reducing greenhouse gas emissions 55% below 1990 levels by 2030, but whether current trends are compati…
Composed Historical Image Retrieval by Modeling Temporal Representations
While time evolves linearly, the geometry of neural embedding spaces is inherently multi-dimensional, often chaotic, and difficult to inter…
Impact of Iterative Fine-Tuning on Transcription Accuracy in Complex Historical Sanskrit Manuscripts
Digitizing the text from handwritten historical manuscripts is required to make them easily accessible, preservable, and to enable historic…
MemFuse: Multi-Source Memory Fusion from Fragmented Observations
Long-term memory is essential for agents that operate across extended interactions, yet existing memory systems and benchmarks predominantl…
A Critical Synthesis of Uncertainty Quantification and Foundation Models for Semantic Segmentation
Foundation models are increasingly breaking what seemed to be impossible not long ago by enabling unprecedented accuracy and cross-domain g…
The Impact of CutMix on Reliability and Robustness in Semantic Segmentation
Ensuring not only high accuracy but also reliable and robust predictions is critical for the deployment of semantic segmentation models in…
Budget-First Tariff Recommendation (BFTR): A Complete Algorithmic Framework for Telecom Plan Recommendation without Overcharging
Telecom operators traditionally offer predefined tariff grids, forcing users to choose from a limited set of plans. This paper proposes BFT…
A Few Cases Are All You Need: An Empirical Study of Annotation-Efficient LoRA Fine-Tuning of MedSAM3
Medical image segmentation is essential for clinical workflows such as treatment planning and disease assessment. While specialist tools li…
Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM services
We present Flama, an open-source Python framework for developing and deploying production-ready web APIs, machine learning services, and la…
Epistemic Subordination: Generative AI and the Infrastructure of Knowledge
Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge i…
Beyond Predictive Fairness: Quantifying Attribution Consistency Across Demographic Groups in Diabetic Retinopathy Screening
Fairness in medical imaging is commonly evaluated through subgroup performance metrics, yet it remains unclear whether models rely on consi…
SIDScope: A Diagnostic Resource for Semantic-ID Interfaces in Generative Recommendation
Semantic-ID mappings are reusable interfaces between item tokenizers and generative recommenders, yet released mappings rarely state whethe…
Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study
Majority voting over multiple LLM samples is widely used to raise answer accuracy, yet its gain varies erratically: on hard questions it ca…
Forgetting, plasticity, and co-observation: a third facet of continual learning
Efficient continual learning remains a fundamental challenge for deep neural networks. While catastrophic forgetting and loss of plasticity…
A strengthening of the MCFL-ness of $O_2$
In the last years, a number of proofs of the fact that $O_2$ is a multiple context-free grammar (MCFG) were given. Such results can be expl…
Do Large Language Models Hallucinate Electric Fata Morganas?
AI hallucinations - that is, outputs which are made up, cannot be verified, or contradict the source material - are generally regarded as a…
Identifying Implicit Premises for Logical Reconstruction of Argument Graphs
The logical reconstruction of argument graphs from natural language text is challenging because of the prevalence of enthymemes (i.e., argu…
Understanding Multilingual Medical ASR Adaptation Through Layer-Wise Analysis
Medical automatic speech recognition (MedASR) requires adaptation to specialised terminology, limited annotated clinical data, and multilin…
MLREF: Efficient Module Reuse for Reward Design in Reinforcement Learning via Large Language Models
Reward function design remains a bottleneck in reinforcement learning. While large language models (LLMs) have enabled automated reward gen…
Learning-State-Aware Dynamic Generative Data Augmentation on Small-Scale Datasets
Small-scale image classification is often limited by the scarcity of training data. Generative data augmentation (GDA) based on pretrained…
SMTrap: Cost-Effective DoS Attacks Against Large Reasoning Models via SMT Conflict Guidance
Existing LRM-DoS methods rely heavily on model feedback to synthesize attack queries, requiring either repeated queries to the target model…
Test-Time Scaling in the Wild: Why Exploitation, Not Exploration, Is the Bottleneck
Test-time scaling (TTS) improves language model outputs by spending additional inference compute - generating multiple candidates, searchin…
SkillForge: Self-Distilling Agents for Project-Specific Issue Resolution
Large language model (LLM) based agents have demonstrated remarkable proficiency in automated software issue resolution, yet they often str…
Graphical Design of Interpretable Architectures
Designing, implementing, and comparing interpretable architectures requires a formal language to represent them. The most common representa…
MedUAG: Unified Understanding and Generation for Medical Multimodal Models
Recent Multimodal Large Language Models (MLLMs) are rapidly evolving into unified understanding and generation (UAG) frameworks. However, e…
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis
Single-step retrosynthesis is a central component of computer-aided synthesis planning, yet its intrinsically one-to-many nature is poorly…
AlphaClifford: Efficient Clifford Synthesis and Transpilation with Model-based RL
Clifford circuits play a foundational role in quantum computing, particularly due to their importance in quantum error correction and fault…
rEDMRec: Distilling Large Language Model Reasoning into an Editable Experience Memory for Recommendation
Large language models can improve recommendation quality by reasoning explicitly over user history and candidate items - for example, extra…
DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering
Retrieve-then-generate pipelines are commonly used to produce deep-research answers for open-ended questions, but retrieval alone is insuff…
GrabVG: Graph-Attentive Binding for Visual Grounding in UAV Imagery
Visual grounding in Unmanned Aerial Vehicle (UAV) imagery aims to localize a target object in complex bird's-eye-view scenes according to a…
From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation
Mechanisms for dynamically converting cyber threat intelligence (CTI) into actionable detection capabilities are necessary due to the rapid…
Harness Continual Learning: Continual Adaptation Beyond Model Parameters
Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern a…
One-Stage Object Detectors in Autonomous Driving
Autonomous vehicles depend on fast and reliable perception systems to detect surrounding vehicles, pedestrians, cyclists, traffic signs, an…
Counterfactual Contrastive Analysis
Visual Counterfactual Explanations (VCEs) aim to explain image classifiers by generating minimally edited and realistic versions of an inpu…
Bernstein-Vazirani Networks: Quantum Machine Learning by Interference
We introduce Bernstein-Vazirani Networks (BVNs), a non-variational quantum machine learning framework that leverages quantum interference f…
GS-VLA: Plug-and-Play Viewpoint Canonicalization for Frozen VLA Policies via Gaussian Splatting
This paper proposes a lightweight, plug-and-play framework that improves robustness to viewpoint shifts in Vision-Language-Action (VLA) pol…
ReWEIGH the Evidence: Calibrating Token-Level Ordinal Visual Evidence to Mitigate Hallucinations in Large Vision-Language Models
Large vision-language models (LVLMs) often hallucinate, generating content that the input image does not support. Preventing such content d…
DA-WAM: Decision-Aligned Future Latents for Driving World Models
Anticipating how scenes evolve under ego actions is fundamental to safe autonomous driving, yet the full potential of world models for deci…
Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift
Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause targeted misbehaviors whe…
Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learnin…
Discretizing Continuous Time Series for Imputation with Masked Diffusion Training
Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics…
PGFS++: Molecular Property Improvement under Synthesis and Diversity Constraints
Improving molecular properties, such as drug-likeness or binding affinity, is a recurring task in early-stage drug discovery. However, mole…
Intercepting the Kangaroo: Experimental Astrolinguistics with Constructed Lexicons, Active Probing, and Large Language Models as Informants and Hypothesis Proposers
Astrolinguistics -- communication with minds that categorize reality differently from ours -- has been purely speculative since Freudenthal…
Leaf Values as Coordinates: Exact Contrastive Explanation for Gradient-Boosted Ensembles
A gradient-boosted ensemble predicts by summing one leaf value per tree. Read those values as coordinates rather than as intermediate resul…
Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
Modern Intel AI PCs ship capable integrated GPUs and NPUs with 16+ GB of unified memory, and they spend considerable time idle. That is not…
Interpretable AI predicts a 2026 summer dry anomaly in central China
Seasonal precipitation anomalies are largely regulated by atmospheric circulation, which dynamical models predict with greater reliability…
Finetuning Strategies for Querying Sounds by Vocal Imitation
This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We inv…
Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning
On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-contex…
ADEPT: Accelerating Dexterity via Pre-Training and Post-Training using Reinforcement Learning
We introduce Accelerating Dexterity via Pre-Training (ADEPT), a large-scale reinforcement learning (RL) framework for learning sim-to-real…
SPADE: Self-Play in Adaptive Synthetic Executable Environments
Continuous self-improvement requires an ever-expanding pool of self-generated, diverse, adaptive goals. For language agents, existing train…
Hybrid Reinforcement Learning and Search for Flight Trajectory Planning
This paper explores the combination of Reinforcement Learning (RL) and search-based path planners to speed up the optimization of flight pa…
Conformal Policy Control
An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause ha…
SkillNet: Create, Evaluate, and Connect AI Skills
Current AI agents can flexibly invoke tools and execute complex tasks, yet their long-term advancement is hindered by the lack of systemati…
From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?
Multi-agent systems (MAS) for structured data-science tasks externalize analytical control through workflows spanning stages, tools, shared…
Interval POMDP Shielding for Imperfect-Perception Agents
Autonomous systems that rely on learned perception can make unsafe decisions when sensor readings are misclassified. We study shielding for…
When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the p…
Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios
Large vision-language models perform well on short- and medium-length video understanding but still struggle to maintain coherent event mem…
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To…
RULER: Representation-Level Verification of Machine Unlearning
Machine unlearning aims to remove the influence of specific training records from a deployed model without retraining from scratch. Current…
A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice
Appropriate reliance on AI advice has become a central research theme in human-AI collaboration. Existing frameworks have focused exclusive…
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) ran…
ChainWorld: Composing Long-Horizon Desktop Workloads from Atomic OSWorld Tasks
Computer use agents are evaluated almost exclusively on atomic desktop tasks, but realistic desktop work requires sustaining state across m…
ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair
Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad search…
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex scientific tasks, yet existing automa…
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Natural-language autoencoders score explanations of hidden activations by reconstruction. An explanation is deemed faithful if the activati…
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-wo…
Fragility of Value under Imperfect Alignment
As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humani…
Hybrid LLM-Augmented Reinforcement Learning Agents for Complex Sequential Decision Tasks
Large Language Models (LLMs) have recently shown strong capabilities in reasoning, planning, and tool-use, enabling new forms of autonomous…
BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding
Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-la…
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may…
Automated Computational Energy Minimization of ML Algorithms using Constrained Bayesian Optimization
Bayesian optimization (BO) is an efficient framework for optimization of black-box objectives when function evaluations are costly and grad…
`From Prompt to Perturbation': An Adaptive Framework for Voice-Based Jailbreaks on Audio LLMs
As large language models (LLMs) are increasingly integrated into audio-based applications, growing concerns have emerged regarding their vu…
Iterative Flow Matching: Path Correction and Gradual Refinement for Enhanced Generative Modeling
Generative models for image generation are now commonly used for a wide variety of applications, ranging from guided image generation for e…
Jailbreaking in the Haystack
Recent advances in long-context language models (LMs) have enabled million-token inputs, expanding their capabilities across complex tasks…
CausalProfiler: Generating Synthetic Benchmarks for Rigorous and Transparent Evaluation of Causal Machine Learning
Causal machine learning (Causal ML) aims to answer "what if" questions using machine learning algorithms, making it a promising tool for hi…
Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead
Code generation has emerged as a critical research area at the intersection of Software Engineering (SE) and Artificial Intelligence (AI),…
Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025
The rise of AI agents is transforming how software can be built. The promise of agents is that developers might write code quicker, delegat…
Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing Systems
Music editing plays a vital role in modern music production, with applications in film, broadcasting, and game development. Recent advances…
TrojanGYM: A Detector-in-the-Loop LLM for Adaptive RTL Hardware Trojan Insertion
Hardware Trojans (HTs) remain a critical threat because learning-based detectors often overfit to narrow trigger/payload patterns and small…
FiLoRA: Focus-and-Ignore LoRA for Controllable Feature Reliance
Multimodal foundation models integrate heterogeneous signals across modalities, yet it remains unclear whether their predictions can be con…
Structure-Informed Estimation for Pilot-Limited MIMO Channels via Tensor Decomposition
Accurate channel state information in wideband MIMO systems is constrained by pilot overhead, a challenge intensifying as bandwidths scale…
Whole-Piece Training for Symbolic Music Language Models via Full-Horizon Compressed Recurrence
For computational efficiency, modern language models are typically trained on independently sampled fixed-length sequences. Symbolic music…
Making Implicit Premises Explicit in Logical Understanding of Enthymemes
Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural langua…
A Framework and Prototype for a Navigable Map of Datasets in Engineering Design and Systems Engineering
The proliferation of data across the system lifecycle presents both a significant opportunity and a challenge for Engineering Design and Sy…
Wildfire Suppression: Complexity, Models, and Instances
Wildfires cause major losses worldwide, and the frequency of fire-weather conditions is likely to increase in many regions. We study the al…
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't
Understanding when Vision-Language Models (VLMs) will behave unexpectedly, whether models can reliably predict their own behavior, and if m…
AutoOR: Scalably Post-training LLMs to Autoformalize Operations Research Problems
Optimization problems are central to decision-making in manufacturing, logistics, scheduling, and other industrial settings. Translating co…
MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
Semi-structured information extraction (IE) from OCR-derived clinical reports is crucial for efficiently reconstructing patients' longitudi…
Key Coverage Matters: Semi-Structured Extraction of OCR Clinical Reports
Clinical reports are often fragmented across healthcare institutions because privacy regulations and data silos limit direct information sh…
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire d…
ICICLE: Expanding Retrieval with In-Context Documents
Generative retrieval (GR) maps queries directly to document identifiers (docids) using parametric knowledge, However, this design makes cor…
DELOS: Contrastive Deep Learning for Low-SNR Blind Transit Searches in Kepler Photometry
We present DEtection in phase-folded Light curves with cOntrastive Scoring (DELOS), a deep-learning framework that uses contrastive scoring…
Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
Language models fine-tuned where the correct completion must outrank a near-synonym competitor often fail silently. The cross-entropy loss…
Sensory Restoration via Brain-Computer Interfaces: A Scoping Review
Brain-computer interfaces (BCIs) can restore sensory and motor function in individuals with severe neurological impairment, but the literat…
Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining…
Horizon-Uniform Sensitivity and Decay of Terminal Reward Perturbations in Discrete-Time Pontryagin Systems
We study local stationary solutions of finite-horizon discrete-time Pontryagin systems near a steady extremal. Suppose that the stationarit…
Hybrid ANN-SNN Pipeline with Local Plasticity
This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs)…
First-Token Broadcasters: Mechanistic Origins of Language Identity and Distributed Robustness in Transformers
Why do multilingual language models sometimes generate in the wrong language, and why is this so hard to fix? We introduce Language Identit…
Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interactio…
Hierarchical Classification via Cascading Feature Elimination: Application to Human Phenotype Ontology-Aligned Facial Phenotyping (FaceMesh2HPO)
FaceMesh2HPO is a framework for classifying facial phenotypic descriptors aligned with the Human Phenotype Ontology (HPO) to support clinic…
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture de…
RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce
Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversi…
Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting
Existing patching and multi-scale methods advance multivariate time series forecasting but treat learned representations as transient bypro…
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distribu…
Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
Synthetic tabular data are valued for preserving not just column-wise marginals but inter-column dependency. Yet the most commonly reported…
Cross-Cohort Spectral-Temporal Dissociation in Frozen EEG Foundation-Model Representations
Objective. We tested whether frozen representations from five EEG foundation models support decoding of long-range temporal correlations, m…
Untrainable elements determine what physical learning remembers
Physical learning rules such as equilibrium propagation (EP), coupled learning (CL), and adjoint coupled learning (AL) train resistive netw…
The Epistemic Politics of AI Anthropomorphism
AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustain…
Approximate Speculative Decoding
Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greed…
Complete, Scalable, and Robust Prioritized Planning for Multi-Robot Ordered Storage and Retrieval at Maximum Capacity
Automated warehouses face a fundamental trade-off between maximizing storage density and achieving high retrieval throughput. While puzzle-…
Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol
AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after…
BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, howev…
「AIで生産性向上」日本の従業員は57%、世界平均は81% 仕事の満足度でも大差──アクセンチュア調査
アクセンチュアの世界20カ国調査で、AIによる生産性向上を実感する日本の従業員は57%と、世界平均の81%を大きく下回った。仕事の満足度や成果の実感でも世界との差が開き、人材と組織の変革が課題となっている。
Stripe、AIモデルゲートウェイのOpenRouter買収 400以上のAIモデルを束ねる中立基盤は維持
Stripeは、AIモデルのゲートウェイを手掛けるOpenRouterを買収することで合意したと発表した。報道による買収額は約75億ドル。OpenRouterは単一APIで400超のモデル切り替えを可能にする。Stripeはトークンコスト最適化を強化し、OpenRouterは買…
OpenAI、ZDRを維持したまま悪用検知へ ログ保持を求めるAnthropicに対抗
OpenAIは、APIデータの非保持設定(ZDR)を維持したまま、複数のやり取りを横断して悪用の兆候を検知する新機構「Private Safety Processing」を発表した。顧客データを自社インフラ外や暗号化で保護しつつ、活動シグナルのみでリスクを判定する。9月に展開を…
PTC、「Onshape」でMCP連携 自然言語でカスタムCAD機能の作成が可能に
PTCは、CAD/PDMプラットフォーム「Onshape」において新機能「FeatureScript MCP Server」の提供を開始した。エンジニアが自然言語とAI(人工知能)を用いて、カスタムCAD機能を作成できるようにする。
Stripe didn’t really buy OpenRouter because of the ‘singularity’
What does a payments giant want with a startup that routes prompts between different AI models? Stripe says it's because of "the singularit…
モデルの利用料金は安くなっているのに、AIの総コスト上昇 「パラドクス」の背景を解説
AIモデルの利用料金の低下が、かえってAIの総コストを押し上げている。Gartnerはワークフロー1件当たりのAI推論コストが2028年までに5倍以上に上昇すると予測する。同社が「推論のパラドックス」と呼ぶ、この逆説の中身とは。また、AIコストが上昇する中でROIを確保するため…
高度なAIのサイバー攻撃、どう対策? 企業が知るべき「スピード格差」の埋め方
米AnthropicのAIモデル「Claude Mythos 5」をはじめとした高度なAIの登場により、サイバーセキュリティの在り方に注目が集まっている。企業に求められる対応を解説する。
Google、大学生向けに「Google AI Plus」を1年間無料提供 「Gemini」アプリに学生向け新機能も
Googleは米国の新学期に合わせ、大学生向けに「Google AI Plus」などを12カ月無料で提供するキャンペーンを発表した。あわせてGeminiアプリに学生向けハブを新設し、授業資料から学習プランを生成する学習ノートブックや、3Dモデルの表示、Gemini Liveでの…
OpenAI seeks to one-up Anthropic with new customer privacy protections
A competition is developing between OpenAI and Anthropic over who can provide the best privacy protections for enterprise customer data.
Cognition CEO denies report that SpaceX tried to acquire the startup
SpaceX was reportedly in talks to buy AI coding startup Cognition. SpaceX has already acquired Cursor as it races to catch up to rivals lik…
Waymoが自動運転AI戦略を解説、「単一AIモデルのE2E方式には2つの問題がある」
米国で自動運転車によるモビリティサービスを展開するWaymo(ウェイモ)が、2009年スタートの「Google Self-Driving Car Project」から開発を積み重ねてきた自動運転技術に基づく同社のAI戦略について説明した。
AI was supposed to win people over by now — it hasn’t
As AI becomes harder to avoid, consumers are growing more wary of the technology — and Silicon Valley is discovering that widespread adopti…
Google packs Search and Gemini with new AI study tools
The launch of the new study features marks Google's latest effort to make Gemini the AI assistant that students turn to when learning and s…
Offering Zero Data Retention for frontier models
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compr…
Researchers say OpenAI revoked their access to limited cyber program
The idea behind OpenAI's Trusted Access for Cyber program is to give trusted defenders better models so they can report bugs and vulnerabil…
Meet the startup helping Wall Street put a price on AI compute
The AI buildout shows no signs of slowing. And with hundreds of billions of dollars a year going into data centers and GPUs, compute has be…
TerraPower’s nuclear reactor has a secret weapon for powering AI data centers
TerraPower's nuclear power plant possesses a strategic advantage over competitors, especially when chasing after data center deals.
Amazon makes its AI-powered Alexa+ free on Fire TV, no Prime required
Amazon is making its AI-powered Alexa+ assistant free on all compatible Fire TV devices in the U.S., automatically upgrading users whether…
2026-08-19(289件)
Calendly throws its hat into meeting note-taker circus
Calendly is also releasing a meeting scheduling assistant called Callie.
AI isn’t close to curing cancer. This startup says it knows what it will take.
It's the data, stupid.
Relativity Networks raises $22 million to bring a faster kind of fiber to data centers
Relativity Networks deals in hollow-core fiber, a rarely deployed technology that allows data to be transmitted 30% faster than conventiona…
ChatGPTの反論で「道徳的判断」の3割超が覆る 高齢者が説得されやすい傾向 神戸大
生成AIの反論によって道徳的な判断の3割超が覆る――神戸大学がこのような研究結果を発表した。米OpenAIのチャットAI「ChatGPT」を利用し、AIの反論が正解のない道徳問題への回答に与える影響を調べた。
Replit expands access to software creation with GPT-5.6 Luna
Replit introduces Free Mode, powered by GPT-5.6 Luna, so anyone can turn ideas into working software without worrying about token costs.
「オープンな国産モデル」に33Bパラメータの新バージョン 国立情報学研究所
国立情報学研究所(NII)が、オープンな国産LLMの新バージョン「LLM-jp-4 33B」を公開。約332億パラメータのDense型モデルで、4種類のベンチマーク全てで従来モデルを上回るスコアを記録したという。
「AIを使える人か、使えない人か」で仕事や評価に差が? 6割のエンジニアが実感した“AI格差”の正体
AIを使うかどうかだけではなく、どの程度使いこなせるかも問われる中、活用スキルの差は業務効率だけではなく、仕事やキャリアにも影響し始めているという。何が起きているのか。ITエンジニア572人の調査から探る。
GxP-Agent: Process-DAG Topology for Reliable Clinical Trial Programming with LLM Agents
Clinical trial programming -- transforming study protocols into analysis-ready datasets under CDISC standards -- is a bottleneck in regulat…
Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution
Agentic AI systems request tool actions that can modify files, send messages, launch jobs, or change workflow state. This shifts the safety…
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term o…
FedPref: Federated Preference Learning for Structured Radiology Report Extraction
Radiology reports describe findings and locations in free text, but downstream search and analysis require these relations in a fixed schem…
The Problem Is the Problem: Towards Scalable Mathematical Discovery
AI systems are increasingly capable of contributing to mathematical research. In research practice, frontier-model reasoning is a limited r…
SkillEffect: Checked Lowering for Memory-Bounded Agent Tools
Agent Skills can specify procedural and resource obligations for tool use, and language models instantiate them as concrete programs. Howev…
Memory Is Communication: The Frontier Between Remembering and Signaling
A bounded agent may obtain information for a decision from its own past, from peers, or from both sources. Retaining task-relevant history…
DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimization
As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) co…
KernelArc: A Multi-Agent Framework for GPU Kernel Optimization
We present KernelArc, a multi-agent framework for autonomous GPU kernel optimization across heterogeneous workloads. Strategy-specialized a…
A decodability criterion predicts when hidden-state selection beats majority voting in large language models
Combining the answers a large language model (LLM) samples for a question into one decision is a test-time information fusion problem, usua…
Toward Personal Intelligence Through Cooperative Observation
A personal AI system needs a model of the user's goals, constraints, and ongoing commitments to plan and act on their behalf, and the quali…
KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn
To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matc…
Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection
Algorithm selection for constraint satisfaction problems requires extracting features that capture problem structure. Manually designing fe…
Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models
Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures…
Fool's Gold: Defensive Deception Against Safety-Removal Attacks on Open-Weight Models
Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weig…
Explicit State Elicitation Is Not Enough: A Controlled Audit of Memory-Policy Classification
Personalized agents must decide whether retrieved user memory should be used, ignored, updated, or queried before it affects a current task…
Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking
Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a…
ASI-Bench: At the Dawn of Artificial Superintelligence
Artificial superintelligence (ASI) requires AI to move beyond mastering existing knowledge toward exploring the unknown, creating new knowl…
DeAR: Decentralized Agentic Reasoning via Capability Grounding and Collaborative Thought Navigation
Existing agentic reasoning systems typically rely on centralized protocols. This design introduces routing bottlenecks and static role allo…
PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs
Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactiv…
LiveHouse-TS: An Open-world Living Benchmark for Time Series Foundation Models
Time Series Foundation Models (TSFMs) have recently emerged as a highly promising paradigm for cross-domain zero-shot forecasting. However,…
SignalReasoner: Assessing the Upper Bound of 3B Models for Signal Mathematical Reasoning
Post-training with supervised chain-of-thought fine-tuning and reinforcement learning from verifiable rewards has substantially improved th…
Wuying-Browser-Agent: Real-World Centric Fundamental Long-Horizon Browser Agents
Browser agents perform well on short, clean demonstrations, but real deployment is fundamentally different: agents must sustain dozens of d…
LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap
Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real cons…
TileMix: Tile-Centric Mixed-Precision Attention for LLM Inference Acceleration
Long-context prefill in large language models (LLMs) incurs substantial computation and memory traffic because dense self-attention compute…
LLM-Only PDDL Domain Repair with Open-Weight Models
AI planning is concerned with finding a sequence of actions that achieves a specified goal. It relies on explicit models of the world, comm…
Cognitive Graph Intelligence for Adaptive and Robust DDoS Attack Detection in Next Generation Networks
Distributed Denial-of-Service (DDoS) attacks threaten network availability, requiring a cognitive detection process that senses traffic, in…
LEGO-RL: Harness-Native Reinforcement Learning for Coding Agents
Reinforcement learning for coding agents increasingly relies on long-running agent harnesses to manage tool integration, repository context…
Task-Aware Harness Provisioning for LLM Agents in Mission-Critical Infrastructure Operations
LLM agents have been widely adopted to operate mission-critical infrastructure (MCI). These agents normally rely on a harness that determin…
Depth Enables Local Entropy: Quadratic Depth Dependence in Deep Variation-Norm ReLU Regression
We study Gaussian regression over the explicit vector-valued Parhi--Nowak deep-RBV^2 architecture with depth L, width w, layer-sum variatio…
Structure-Internalized Rule Language Model for Faithful Knowledge Graph Reasoning
Knowledge Graph Reasoning (KGR) aims to discover latent facts by leveraging the structural evidence available in KGs, posing a challenge to…
SAGE: Self-Evolving Storyboard Skills via Attribution-Guided Rule Evolution
Storyboards turn screenplays into visual shot plans for automated short drama production. Professional storyboarding relies on tacit direct…
When AI Designs AI: Innovation or Imitation?
Recent advances in LLM agents have made them increasingly capable of designing methods for complex AI tasks. This raises two central questi…
Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context
User-facing tool agents must coordinate dialogue and tool use as user goals unfold over multiple turns. Yet interactive reinforcement learn…
SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models
Recent efforts toward fully automated AI scientists have demonstrated that language-model agents can generate hypotheses, execute experimen…
Agent Lightning v1.0: Towards Harnessed Agentic RL
Modern agents operate inside agent harnesses that manage tools, context, and control flow, making the harness a critical part of the agent…
When to Review: Spaced Repetition for Continual Pre-Training of Language Models
Continual pre-training of large language models must acquire new information without erasing old knowledge. Existing replay methods often c…
Quantifying Risk Under Evolving Uncertainty: Belief-Dependent Robustness for Safe Sequential Decision Making
How cautious should an agent be while it is still learning its environment? We propose RATTL (Risk-Adversarial Total-Reward Learning), whic…
TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation
Agent Skills package reusable natural language procedures with executable resources, enabling software agents to acquire task specific capa…
MoNe: Modular Neural Memory for Efficient Long Context Inference
We present MoNe, a lightweight modular neural memory that attaches to any frozen pretrained Transformer to enable long-context inference wi…
Validated Adaptation for Aerial Crowd Monitoring at Mass Gathering Scale: A Deployment Protocol, a Severity Law, and a Diagnostic for Label-Free Drone Crowd Counting, Toward the FIFA World Cup 2034 (Saudi Arabia)
Saudi Arabia will host the 2034 FIFA World Cup and already operates crowd management at Hajj scale. Drone-based counting must hold accuracy…
Graph Surgery and the Do-Operator: A Precise Correspondence for Acyclic Structural Causal Models
The $\operatorname{do}$-operator is described graphically by deleting arrows into its targets and functionally by replacing their mechanism…
Beyond the Trace: Coupling an Interpretable Reasoning-State Readout to Native MoE Routing
What a reasoning model writes is only a partial record of the process that produces it. We introduce a two-level internal readout for mixtu…
LLM-Derived Preference Judgments Are Not Self-Consistent
Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by askin…
GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities
LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a…
Auditing Self-Evolution in Financial Agents: Capability Gains, Security Drift, and Execution-Interface Mismatch
Self-evolving agents turn experience into reusable skills, workflows, or memories, but post-evolution accuracy alone does not show whether…
Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals
Despite their widespread use, Large Language Models (LLMs) remain limited by a fundamental problem: the generation of plausible but false c…
Accuracy and Robustness of Model Cascades Under Data Perturbations
Prediction cascades significantly reduce energy consumption of Artificial Intelligence (AI) models while maintaining high predictive perfor…
Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents
Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is…
Evaluating the Diversity of AI-Generated Content with Diversity Profiles
Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently…
Neuro-symbolic learning over OWL 2 DL via consequence-based compilation to differentiable circuits
OWL 2 DL ontologies, grounded in the description logic $\mathcal{SROIQ}$, express large knowledge bases in biomedicine and the Semantic Web…
The Curious Case of Exploding DecPOMDPs: Containing the Fire through Policy Counting
Decentralised partially observable Markov decision processes (DecPOMDPs) provide a general framework for modelling multi-agent decision mak…
D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory
Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization.…
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows
Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Y…
ARASH: Adaptive Retrieval And Shot Selection for Tabular Prediction
Tabular prediction is a critical task across numerous applications. The recent success of large language models has sparked various approac…
AutoResearch: Insight In, Hallucination Out
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the re…
Adaptive Policy Portfolios for Robust Markov Decision Processes
Robust Markov decision processes optimize one policy against a set of plausible transition functions. This can be conservative when the unk…
EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection
Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no…
Procedural Content Metageneration via Program Search and Continual Abstraction Discovery
Large language models can generate executable programs, which makes it possible to search directly over procedural content generators rathe…
Towards Zero-Shot Task Transfer with Neurosymbolic World Models
State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent…
Can Large Language Models Explain Flight Safety Events? A Prior-Guided Semantic LLM-based Approach
Improving flight safety with flight data requires not only accurate detection of risk events, but more importantly, clear interpretation of…
StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents
AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code repositories, documents,…
HLSR: Hybrid Live Forecast Selective Dynamic Vehicle Rerouting for Real-Time Congestion Avoidance
Urban traffic congestion reduces productivity and increases travel cost and emissions. Network-wide live travel-time shortest-path reroutin…
Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating
Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on…
On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory b…
Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis
The rise of ChatGPT has brought a notable shift to the AI sector, with its exceptional conversational skills and deep grasp of language. Re…
Intent-Driven Dynamic Chunking: Segmenting Documents to Reflect Predicted Information Needs
Breaking long documents into smaller segments is a fundamental challenge in information retrieval. Whether for search engines, question-ans…
The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks
Repository-scale coding requires an agent to keep tests, imports, configuration, and migration rules consistent within a bounded context wi…
A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications
Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - espe…
What If AI Carried Her Imagination? Black Girls as Creators in an AI Storytelling Weekend Program
This paper presents the design and outcomes of a seven-weekend AI storytelling program developed for Black girls aged 10-12. Grounded in Af…
CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents
Large-scale urban simulation plays a pivotal role in social science, traffic safety, and transportation policy. Recent work has shown that…
QuantumNovelty: A Skill-Orchestrating Language Agent for Referee-Style Review and Patentability Screening of Quantum Papers and Patents
Language-model agents increasingly produce quantum-science results; we ask whether the same agentic paradigm can also scrutinize them in an…
AI, Brain Death Detection, and Islamic Law
The deployment of machine learning systems capable of detecting covert consciousness in neurologically injured patients creates a profound…
ComNetX: Local Hierarchical Adaptation for Dynamic Community Detection
Dynamic community detection is commonly addressed either by full-snapshot recomputation or by solver-specific dynamic procedures. Full reco…
Effective Personalized AI Tutors via LLM-Guided Reinforcement Learning
Generative AI (GenAI) is rapidly reshaping education by unlocking the potential for personalized tutoring. Yet, emerging platforms largely…
When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice
Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias rema…
Education-centered critical policy analysis of AI: Ghana's AI strategy as a case
National AI strategies increasingly guide governance, workforce development, innovation, and competitiveness, but less is known about how t…
Average Distance Approximation for Static Large Graphs
Calculating average distances in large-scale networks is computationally intensive and constrained by limited main memory, posing a signifi…
Sparse Coverage: Semantic Center Representations for Patent Prior-Art Retrieval
Patent prior-art retrieval is a recall-oriented search task over long and highly structured technical documents. Dense retrieval improves s…
CARA: Cognitive Adaptive Recommendation Agent
Recent advances in large language models and agent-based recommendation frameworks have introduced new opportunities for more flexible and…
WIP: LLM Odyssey: A Game-Based Platform for Teaching LLM Engineering Concepts
This work-in-progress (WIP) innovative practice category paper presents LLM Odyssey, an open source, browser-based serious gaming platform…
EMAN: Optimization-Driven Capacity Growth through Path Emergence in Multi-Task Learning
Existing multi-task learning methods rely on hard sharing, multiple paths or experts, adaptive sharing, and dynamic expansion. However, the…
Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations
LLM-based code generation is now embedded in mission-critical pipelines, but defenses against vulnerable output remain post-hoc -- static a…
Position: Fairness Failure in Generative Models is an Evaluation Problem
Despite groundbreaking advancements in generative models during the last decade, concerns about their lack of fairness, reinforcing societa…
PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation
Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures and objec…
Without journalists, there is no journalism: the social dimension of generative artificial intelligence in the media
The implementation of artificial intelligence techniques and tools in the media will systematically and continuously alter their work and t…
YILDIZ-VPR: A Novel Dataset with Dense Coverage Under Diverse Environmental Conditions for Visual Place Recognition
Visual Place Recognition (VPR) aims to recognize the location of a query image by comparing it with a set of geo-referenced images. Althoug…
The 10th AI City Challenge
The 10th AI City Challenge, held with ECCV 2026, marks a decade of community benchmarking for intelligent transportation, smart cities, and…
Cross-Model Memory Transfer via Target-Side Reader Adaptation
Methods for improving knowledge use in large language models typically fall into two regimes. Non-parametric retrieval offers flexible acce…
Institution-Specific LLM Prompting Recovers PHI That De-identification Systems and Their Gold Standards Both Miss
Secondary use of electronic health records requires de-identification, yet existing systems miss \emph{institutionally situated} protected…
Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting
Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal recor…
Structured Driving-State Narratives for Small Language Model-Based GNSS Spoofing Detection
Autonomous vehicles (AVs) depend on reliable Global Navigation Satellite System (GNSS) positioning. However, spoofed GNSS signals can induc…
From Abductive Explanations to Global Logical Rules for Node Classification in SGCs
Graph Neural Networks (GNNs) have achieved remarkable performance in node classification tasks, motivating growing interest in methods capa…
Iterative tensor network transformations for element-wise evaluation of elementary and filtering functions
Tensor networks are powerful formats for compressing large-scale data. However, their application to general data processing has been limit…
Authorization Before Context: A Model-Neutral Audience Boundary Against Cross-Audience Memory Leakage in Agentic Systems
A personal language agent learns a fact from one audience and may later place it in the prompt it assembles for another. This memory-to-con…
Q-Learning With World Models
Off-policy reinforcement learning (RL) has become increasingly sample-efficient, enabling applications such as RL fine-tuning of Vision-Lan…
Expected free energy as an information constraint on the Bethe Lagrangian
Active inference selects actions by minimising an expected free energy functional over predicted futures. However, adding an expectation ov…
Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases
Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demandi…
The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence
An AI audit record is useful only if its durability and trust boundary are explicit. Returning a guarded decision before any durable write…
Task Specialization Fine-Tuning for Contextual Reinforcement Learning
Contextual Reinforcement Learning (CRL) seeks to generalize classical RL by maximizing task coverage across a context space of related task…
Token Optimization and Context Window Management in Multi-Agent AI Workflows
Multi-agent AI workflows are limited not only by model quality but by token cost, latency, and context-window quality. This paper presents…
Graphectory Viewer: A Tool for Process-Centric Analysis of Agentic Software Trajectories
We present Graphectory Viewer, a web-based tool for interactive, process-centric analysis of software-agent trajectories. Building on the G…
Teach and Grow: An Agent-Centered Architecture for General Robot Learning
End-to-end vision-language-action (VLA) and world-action models offer an elegant route to general-purpose robotics, but their reliability i…
PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance
Autonomous AI agents are emerging as interfaces for decentralized finance (DeFi) actions such as swaps, lending operations, and yield manag…
Delta2Gamma: Band-Wise Adaptive Contrastive Learning of EEG for Alzheimer's Disease Detection
Low-cost, scalable screening for dementia remains an open problem. Imaging-based diagnosis is costly and hard to deploy widely. Electroence…
COMIC: Reference-Aware Safety Gating for Multimodal Large Language Models
Multimodal large language models (MLLMs) are increasingly used to interact with screenshots, scanned documents, diagrams, and other visuall…
Structural Plan-to-Model Conversion with Deterministic Geometry and Guarded Agentic Vision-Language Refinement
Converting structural framing plans into editable finite-element model drafts remains labor-intensive and prone to transcription error. Exi…
Maximum Tsallis Entropy Distributions for Robust and Efficient Sparse Learning from Correlated Data
This paper addresses the limitations of Gaussian distribution assumptions in statistical sparse learning, particularly in modeling correlat…
Adaptive surrogate modeling for high-dimensional spatio-temporal output
This paper develops an adaptive surrogate modeling method for problems with very high-dimensional spatio-temporal outputs. The analysis of…
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL
Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its stro…
Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction
X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording…
Nonadaptive Learning in Robust Nonlinear Output Regulation
This paper considers robust nonadaptive regulation for general nonlinear systems in an output-feedback setting with arbitrarily high relati…
Understanding Curriculum Learning in Large Language Models via Cross-Difficulty Optimization Dynamics
Curriculum learning has been widely adopted in the post-training of large language models by organizing training data from easy to hard. Ho…
When Agents Act on Web3: An Attack-Surface Survey of MCP, Skills, and Tool Calling
AI agents increasingly act rather than merely read: across the Model Context Protocol (MCP) ecosystem, the share of deployed tools that mod…
Rethinking Irregular Time Series Forecasting from the Perspective of Basis Functions
Irregular time series forecasting is crucial in many domains, such as healthcare and meteorological observation. However, due to the inhere…
Beyond MSE: Rethinking the Evaluation Metric and Benchmarking for Irregular Time Series Forecasting
Existing research on irregular time-series forecasting has primarily focused on model design, while evaluation metrics remain insufficientl…
NeuroAbs: A Neuro-Symbolic RTL Abstraction Framework for Property Checking Acceleration
Formal verification is a crucial technique for ensuring the functional correctness of hardware designs. In the context of property checking…
Learning What Not to Learn: Adversarial Disentangled Prompt Tuning for Robust Vision-Language Models
While adversarial prompt tuning can enhance robustness of vision-language models efficiently, we find that existing methods aggravate robus…
ORPA: Online Residual Policy Adaptation for Robot Manipulation Control with Human Feedback
Robotic manipulation policies trained via imitation learning, such as Action Chunking with Transformers (ACT), can achieve strong performan…
SPACE: Sample-cloud Predictive Adaptive Conformal Ellipsoids for Multivariate Time-Series Forecasting
Modern probabilistic time-series forecasters often express uncertainty through forecast samples. While typically converted into nominal pre…
MoFE: A Novel Mixture-of-Experts Framework with Fourier Neural Operators for Cryptocurrency Forecasting
Forecasting cryptocurrency prices remains a formidable challenge due to inherent non-stationarity, abrupt regime shifts, and multi-scale st…
Inductively Scalable, Single-Step Neural Surrogates for Wave-Scattering Inverse Problems
Neural network surrogates are an emerging alternative to traditional electromagnetic wave simulators like finite-difference time-domain (FD…
Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
Reliable jailbreak evaluation is essential for assessing LLM safety, but most existing studies rely solely on attack success rate (ASR) wit…
Integrating Novelty and Surprise for Experience Prioritization and Exploration in Image-Based Reinforcement Learning
Sample efficiency is a central challenge in reinforcement learning (RL), particularly in image-based domains where agents must learn from h…
PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTX
We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel…
Leveraging generative hallucination and biophysics-informed modeling for unified biomolecular sequence-structure co-design
Biomolecular design underpins applications from molecular recognition to therapeutics and synthetic biology, yet de novo interaction design…
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
We introduce Semantic Task Completion Video Generation, an outcome-oriented video generation task. Under this formulation, success requires…
Beyond FLOPs: Energy-Aware Knowledge Distillation for Sustainable LLMs on Code-Related Task
Background: Large Language Models (LLMs) are increasingly being applied to Software Engineering (SE) tasks, achieving high accuracy across…
Explainable AI-Powered Framework for Video-Based Skill Assessment in Cataract Surgery
Persistent shortages in the surgical workforce and inherent limitations of traditional training methods highlight the necessity of automate…
CoAL-RAG: A Complexity-Aware Legal Retrieval-Augmented Generation Method
Legal consultation questions exhibit multi-level complexity. A single retrieval strategy often leads to over-reasoning for simple questions…
No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models
Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial sol…
Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models
Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations…
Domain-Adapted Molecular Language Models for Efficient Search of Make-on-Demand Libraries
Pretrained molecular language models are increasingly used as molecular encoders for learning structure-property relationships. However, th…
DMT-Dens: Density-preserving manifold visualization for biological data
Motivation: Low-dimensional embeddings are widely used to explore cell-state heterogeneity in single-cell and other high-dimensional biolog…
tinyDSM: A Framework for Skill Modeling and Development for Resource-Constrained Millirobots
In this study, we investigate developmental mechanisms that enable small, resource-constrained systems such as cm-sized millirobots to auto…
HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness Safety
Large language models are increasingly deployed through agent harnesses that manage tools, extensions, persistent state, permissions, and e…
Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goa…
From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support
Learning analytics models can identify students at risk of poor performance, but they do not directly indicate which interventions are feas…
Iterative Grasp Pose Refinement: A Deep Reinforcement Learning Approach for 2D Vision
Developing robots capable of understanding and manipulating objects requires compact, interpretable, and generalizable representations. Thi…
DEPT: Document Embedding Preservation Tuning for Unified Query Expansion and Retrieval
Large language models (LLMs) can both expand underspecified queries and encode text as dense representations, suggesting a unified model fo…
MobileWorldSafety: Benchmarking GUI Agent Safety Against Environmental Injection Attacks in Android Apps
LLM-powered GUI agents that autonomously operate smartphones are rapidly transitioning from research prototypes to early real-world deploym…
Benchmarking Automated Security Patch Backporting: How Far Are We?
Automated security patch backporting is critical for mitigating N-day vulnerabilities. Recent tools report success rates above 80% on their…
GADR: Gathering Architecture Decision Records from Meeting Transcriptions
Existing LLM-based approaches to Architecture Decision Record (ADR) generation share a critical and largely unexamined assumption: that inp…
Dijkstra as an Oracle for Online Stochastic Shortest Path Navigation with Provable Guarantees
Mobile robots that operate in side by side with humans and critical facilities must reach their goals at low cost, despite often unknown tr…
Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models
Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although m…
What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate old…
Learnware for CSI Feedback: Scene-specific Small Models Can Do Big
Intelligent channel state information (CSI) feedback is essential for realizing the high capacity and spectral efficiency goals of future 6…
Training with synthetic data for drone detection in thermal imagery
Ground-to-Air (G2A) drone detection in medium- and long-wave infrared (MWIR/LWIR) imagery is challenging due to reduced texture information…
Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses
The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ…
MotoSafety: Edge-AI with Learned Temporal Importance for Two-Wheeler Collision Risk Assessment Under Time Pressure
Powered two-wheeler riders face critical safety challenges in low- and middle-income countries, yet limited studies exist on how cognitive…
The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges
LLMs increasingly rely on external contexts, such as pre-defined system prompts or retrieved documents, to improve generation quality. Howe…
AdaLens: Interactive Storyline for Monitoring and Steering Long-Running Agentic Data Analysis
Large language models are pushing data science toward increasingly autonomous and agentic workflows, with recent systems already supporting…
Encoded but Not Actionable: Auditing the Decode-Generate-Steer Gap in Frozen LLMs for Geometric Constraints
Large language models (LLMs) have demonstrated strong performance on structured reasoning tasks, but what they encode and whether it inform…
BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models
While Multimodal Large Language Models (MLLMs) have made significant strides in visual comprehension, their ability to reason about text-de…
Comparative Study of Out-of-the-Box Technology for Automatic Target Detection and Recognition
Automatic Target Detection and Recognition (ATD/R) is critical for military decision support and (semi-)autonomous operations. Recent advan…
Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks
Background and Context: Question and inquiry are integral parts of knowledge seeking and learning. Despite their importance, students tend…
A Theoretical Framework for Parallel Lifelong MAPF Using Group Decentralized Planning
In the Lifelong Multi-Agent Path Finding (L-MAPF) problem, agents must repeatedly move from one destination to another while avoiding obsta…
Collective Counterfactual Planning: Coordination, Consent, and Verification under Representational Constraints
Groups routinely complete projects that no single member can plan, execute, or verify alone. We propose a formal model of this phenomenon,…
Grading Needs a Rubric, Not Intelligence
Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against a…
Efficient RLVR Scheduling via Graph-Structured Online Difficulty Estimation
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rol…
SIGMA: SHAP-Guided Implicit-Trajectory Generation for Metadata-Free LLM-Based AutoFE
Recent research has leveraged Large Language Models (LLMs) to enhance Automated Feature Engineering (AutoFE) through semantic descriptions…
An Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models
In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted…
Too Sure to Be Safe: Model Calibration for Reliable Log Anomaly Detection
Online log anomaly detection is critical for maintaining the reliability of large-scale computing systems. Although recent language model-b…
Dual Co-Train: Cross-Dataset Ultrasound Tongue Segmentation Under Extreme Data Scarcity
Ultrasound tongue contour segmentation remains challenging under cross-dataset domain shift, where limited annotations, probe variability,…
Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media
The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideol…
Traceable Trust for action-ready artificial intelligence in bioscience
Artificial intelligence (AI) is becoming part of the working infrastructure of the biosciences. AI models can predict biomolecular structur…
Policy-Invariant Reward Shaping from LLM Feedback: A Framework for Hybrid RL Agents
Combining large language models with reinforcement learning is increasingly explored, yet the theoretical status of LLM-derived reward sign…
Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System
GPT-style models achieve strong performance by representing language with finite vocabularies of reusable discrete tokens. This success has…
Harnessing Magnitude-Only and Complex Measurements for Improved Dynamic MRI Reconstruction with Learned Priors
MRI reconstruction methods for undersampled k-space data naturally utilize complex-valued measurements. Parallel developments in sparse pha…
From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image Generation
Large-scale image generation has benefited from advances in data scale, quality, rebalancing, and recaptioning, yet conventional pipelines…
HA-VLN 2.0: An Open Benchmark and Leaderboard for Human-Aware Navigation in Discrete and Continuous Environments with Dynamic Multi-Human Interactions
Vision-and-Language Navigation (VLN) has been studied mainly in either discrete or continuous spaces, with little attention to dynamic, cro…
Efficient Dynamic Shielding for Parametric Safety Specifications
Shielding has emerged as a promising approach for ensuring safety of AI-controlled autonomous systems. The algorithmic goal is to compute a…
MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale
Representation learning of geospatial locations remains a core challenge in achieving general geospatial intelligence, with increasingly di…
LLM Enhancement with Domain Expert Mental Model to Reduce LLM Hallucination with Causal Prompt Engineering
When consequential decisions depend on knowledge that exists nowhere in writing, LLMs hallucinate not from retrieval failure but from model…
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
Enabling embodied agents to imagine future states is essential for robust and generalizable visual navigation. Yet, state-of-the-art system…
Planning under Distribution Shifts with Causal POMDPs
In the real world, planning is often challenged by distribution shifts. As such, a model of the environment obtained under one set of condi…
Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models
Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a single architecture. While unified arc…
TSQueryBench: LLM-as-a-Judge for Time Series Explanations
Natural language explanations of time series data are increasingly produced by foundation models in high stakes domains, making factual cor…
Chronos: The AI Co-Historian
AI is increasingly supporting, accelerating, and automating scientific discovery across subjects. Yet, the adoption of AI in historical res…
Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest
Large language models (LLMs) are trained to align with user preferences through methods like reinforcement learning. Yet models are beginni…
Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs
Vision-Language Models (VLMs) expand the attack surface of safety-aligned systems by coupling visual perception with text generation. Exist…
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
Current Vision-Language-Action (VLA) models typically treat the deepest representation of a vision-language backbone as universally optimal…
ScreenSearch: Uncertainty-Aware OS Exploration
Desktop GUI agents operate under partial observability: visually similar screens can correspond to different underlying workflow states, so…
Pander Score: A Continuous Measure of Sycophancy as Epistemic Deference
Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user. Existing evaluations typically measure th…
A Multimodal Agentic Pathology Co-pilot via Evidence Grounded Reasoning
Pathology is the cornerstone of modern medicine, where accurate decision-making relies heavily on evidence-based practices. While artificia…
ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing
Personalized public transit routing in public transit systems remains challenging due to the difficulty of capturing and integrating divers…
Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism
Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification in…
The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents
A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retiremen…
FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation
Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic g…
JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
Public open-weight language models are often fine-tuned on private or domain-specific data before deployment, creating a need to audit whet…
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limit…
G-ReAct: Graph-Guided Deep Search via Structure-State Co-Evolution
Deep search has become a fundamental capability of large language models (LLMs) for solving open-domain complex tasks. However, existing ap…
Agents Catching Agents: Shortcut Cascades and Benchmark Gaming in Clinical Multi-Agent Systems
Clinical decision support is moving toward committees of language-model agents deliberating on a shared workspace. We ask whether such comm…
VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge
Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex pr…
Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE
Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the sa…
Rationale-Guided Learning for Multimodal Emotion Recognition
Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. Howev…
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work…
LLM-Guided Graph Generation for Structure-Based Local Improvement Methods
Large neighborhood search normally selects a random subset of decision variables for iterative optimization. To efficiently solve various p…
The Authenticity Gap in Human Evaluation
Human ratings are the gold standard in NLG evaluation. The standard protocol is to collect ratings of generated text, average across annota…
Comprehensive framework for evaluation of deep neural networks in detection and quantification of lymphoma from PET/CT images: clinical insights, pitfalls, and observer agreement analyses
This study addresses critical gaps in automated lymphoma segmentation from PET/CT images, focusing on issues often overlooked in existing l…
ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation
Diffusion models have been verified to be effective in generating complex distributions from natural images to motion trajectories. Recent…
Optimizing Container Loading and Unloading through Dual-Cycling and Dockyard Rehandle Reduction Using a Hybrid Genetic Algorithm
This paper addresses the NP-hard problem of optimizing container handling at ports by integrating Quay Crane Dual-Cycling (QCDC) and dockya…
LSem2Vec: A Simple yet Effective Two-Stage Approach for Source Code Embedding
The advent of large language models (LLMs) has significantly advanced artificial intelligence in software engineering, with source code emb…
M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction
Accurately predicting the popularity of micro-videos is a critical but challenging task, characterized by volatile, `rollercoaster-like' en…
Diffusion Models for Smarter UAVs: Decision-Making and Modeling
Uncrewed Aerial Vehicles (UAVs) are increasingly used in modern communication networks. However, challenges in decision-making and digital…
Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization
Transformers are difficult to optimize with stochastic gradient descent (SGD) and largely rely on adaptive optimizers such as Adam. Despite…
MCTS-KBQA: Monte Carlo Tree Search with Information Gain Rewards for Knowledge Base Question Answering
This work investigates how to improve large language model (LLM)-based reasoning for knowledge base question answering (KBQA) via Monte Car…
Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits
As large language models (LLMs) increasingly function as human-like assistants exhibiting human-like personality traits, understanding thei…
LZ Penalty: An information-theoretic repetition penalty for autoregressive language models
We introduce the LZ penalty, a penalty specialized for reducing degenerate repetitions in autoregressive language models without loss of ca…
TabularQGAN: A quantum generative model for tabular data synthesis
In this paper, we introduce a novel quantum generative model for synthesizing tabular data. Synthetic data is valuable in scenarios where r…
Multi-Scale Spectral Attention Module-based Hyperspectral Segmentation in Autonomous Driving Scenarios
Recent advances in autonomous driving (AD) have highlighted the potential of hyperspectral imaging (HSI) for enhanced environmental percept…
Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review
In July 2025, 18 academic manuscripts on arXiv contained hidden instructions that manipulated AI-assisted peer review (indirect prompt inje…
Solving nonconvex Hamilton--Jacobi--Isaacs equations with PINN-based policy iteration
We propose a mesh-free policy iteration framework that combines classical dynamic programming with physics-informed neural networks (PINNs)…
Visual Prompting for Robotic Manipulation with Annotation-Guided Pick-and-Place Using ACT
Robotic pick-and-place tasks in convenience stores pose challenges due to dense object arrangements, occlusions, and variations in object p…
Exploring Efficient Open-Vocabulary Segmentation in the Remote Sensing
Open-Vocabulary Remote Sensing Image Segmentation (OVRSIS), an emerging task that adapts Open-Vocabulary Segmentation (OVS) to the remote s…
ChannelFlow-Tools: A Configuration-Driven Pipeline for Generating Machine-Learning-Ready Datasets of 3D Obstructed Channel Flows
Data-driven surrogate models are increasingly used in computational fluid dynamics, and their reliability depends on the quality of the tra…
Future-Back Threat Modeling: A Foresight-Driven Security Framework
Traditional threat modeling remains reactive-focused on known TTPs and past incident data, while threat prediction and forecasting framewor…
Audio Physical Dynamics Inspired Deepfake Detection for Voice Authentication Systems
Voice authentication systems deployed at the network edge face dual threats: a) sophisticated deepfake synthesis attacks and b) control-pla…
Cluster Aggregated GAN (CAG): A Cluster-Based Hybrid Model for Appliance Pattern Generation
Synthetic appliance data are essential for developing non-intrusive load monitoring algorithms and enabling privacy preserving energy resea…
The $\mathbf{P}$-Completeness of Inverted Index Traversal: On the Complexity of Evaluating Boolean Query DAGs
Modern AI agents increasingly rely on search infrastructure to execute complex, neuro-symbolic reasoning workflows. These workflows often c…
Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries
Large Language Model (LLM)-powered Automatic Speech Recognition (ASR) systems achieve strong performance with limited resources by linking…
SCOPE: Selective Conformal Optimized Pairwise LLM Judging
Large language models (LLMs) are increasingly used as scalable judges in pairwise evaluation, but they remain prone to miscalibration and b…
Adversarial Data Modeling in Epidemiology
Epidemiological models increasingly rely on crowdsourced, self-reported behavioral data such as vaccination status, mask usage, and social…
Parametric Knowledge in RAG-SFT for Domain-Specific Document Generation
Retrieval-Augmented Generation (RAG) fine-tuning has shown substantial improvements over vanilla RAG, yet most studies target document ques…
Supporting Calibrated Reliance in Human-AI Collaboration: Different Strategies for Different Tasks
As AI systems increasingly support human decision making, a central challenge is determining what information helps people recognize when t…
Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries
Although LLM context lengths have grown, there is evidence that their ability to integrate information across long-form texts has not kept…
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
Visual Retrieval-Augmented Generation (VRAG) empowers Vision-Language Models to retrieve and reason over visually rich documents. To tackle…
SegWithU: Uncertainty as Perturbation Energy for Single-Forward-Pass Risk-Aware Medical Image Segmentation
Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and cli…
Self-Distillation as a Performance Recovery Mechanism for LLMs: Counteracting Compression and Catastrophic Forgetting
Large Language Models (LLMs) have achieved remarkable success, underpinning diverse AI applications. However, they often suffer from perfor…
FairNVT: Fair Classification via Noise Injection in Vision Transformers
This paper presents FairNVT, a lightweight debiasing framework for pretrained transformer-based encoders that improves prediction fairness…
Convergent Evolution: How Different Language Models Learn Similar Number Representations
Language models trained on natural text learn to represent numbers using periodic features with dominant periods at $T=2, 5, 10$. In this p…
Protect the Brain When Treating the Heart: Feasibility of 2.5D U-Net for Real-Time Gaseous Microemboli Detection
Gaseous microemboli (GME) represent a common complication of cardiac structural interventions across both surgical and transcatheter approa…
Discovering physical mechanisms from experiment-simulation mismatches
Scientific discovery often begins where observation and prediction disagree. As computation and machine learning survey chemical space, exp…
SOD: Step-wise On-policy Distillation for Small Language Model Agents
Tool-integrated reasoning (TIR) is difficult to scale to small language models due to instability in long-horizon tool interactions and lim…
Adaptive AI Task Partitioning and Safe Offloading in Heterogeneous Edge-Cloud Continuum
In recent years, the use of artificial intelligence on resource-constrained IoT devices has grown significantly. However, existing approach…
SurgicalMamba: Dual-Path SSD with State Regramming for Online Surgical Phase Recognition
Online surgical phase recognition must commit to a prediction at every frame of a procedure that runs for hours, from past frames alone and…
EXPO-FT: Sample-Efficient Reinforcement Learning Finetuning for Vision-Language-Action Models
The ability to efficiently and reliably learn new tasks has been a foundational challenge in robotics. Vision-Language-Action (VLA) models…
On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note
We prove an elementary bounded-differences inequality for functions of non-isotropic Gaussian vectors. Specifically, if $f$ has bounded coo…
Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations
We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for w…
FVSpec: Real-World Property-Based Tests as Lean Challenges
We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 propert…
BRo-JEPA: Learning Modular Transformations in Latent Space
Can neural networks learn algebraic rules from visual inputs, or do they merely fit observed patterns? We study this question using MNIST (…
Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents
Large language model (LLM) agents can persist across tasks, acquire memory, activate Skills, synthesize tools, fork child processes, attach…
Planning-aligned Token Compression for Long-Context Autonomous Driving
Monolithic vision-action models represent an emerging paradigm in autonomous driving. However, this architecture produces token sequences t…
The Standard Interpretable Model: A general theory of interpretable machine learning to deductively design interpretable methods using Lagrangian mechanics
As Artificial Intelligence models grow in complexity, interpretability has become an indispensable tool for understanding, debugging, and c…
Physics-Grounded Causal Auditing of End-to-End Driving Planners
End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that me…
How Transparent is DiffusionGemma?
LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging su…
Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map
Can a platform tell, before deployment, whether an open-weight checkpoint has had its refusal mechanism stripped? Runtime guards cannot: th…
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment
Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which ine…
Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly unders…
From Adoption to Deployment: A Qualitative Study on AI Integration in Software Development Practice
The increasing adoption of Large Language Models (LLMs) as AI components in modern software systems introduces distinct security risks to t…
Constitutional Midtraining: Content Presence Drives Alignment Gains
Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining intervent…
WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA
Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deploy…
Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning
As "AI Scientists" emerge to drive research via the Model Context Protocol (MCP), systems relying on ephemeral scripts will fail. The sheer…
Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent
Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are…
Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. E…
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only…
FUSE: Frame-Unified Stress Estimation from Facial Video
Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approache…
A 12-CNOT Double Qubit Excitation Gate
In this work, we presented, to the best of our knowledge, the first reported 12-CNOT decomposition of the double qubit excitation operator.…
AQuA: Recursively Self-Improving Quantitative Trading Research Agents
We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from ea…
Falsehood and Impossibility Are Different Directions in an AI's Representation of Language
Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model interna…
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-L…
UltraArUco: A Lightweight Multilingual Library And Framework With Low-Latency Real-Time Marker-Based Tracking System For Mobile AR Interaction
UltraArUco is a lightweight multilingual library and framework for low latency, realtime marker-based tracking in mobile augmented reality.…
CLAIR-Fin: An Adversarial Multi-Agent Framework for Claim-Level Verification and Adaptive Debate in Cross-Modal Financial QA
Existing defenses against hallucination in retrieval-augmented and multi-agent pipelines remain partial: evidence is trusted despite modali…
「うちの会社にAIなんて」――謙遜する中小企業に眠る伸びしろ OpenAIがパートナー網拡充、新規3社が語る
米OpenAIが法人向けパートナーネットワークの日本での拡充を進めている。メディア向け座談会に登壇した新規パートナー3社が語ったのは、「うちの会社にAIなんて」と謙遜する中小企業にこそ眠る“伸びしろ”だった。
なぜNTTは「小型LLM」にこだわるのか 国産AI「tsuzumi 2」に込めた狙い
LLMは大規模であるほどよいのか。NTTが開発する国産LLM「tsuzumi 2」は、小型化と日本語処理、データ主権を重視する。企業が生成AIを実務に組み込む際の現実解として、同モデルは何を目指しているのか。開発担当者の講演から読み解く。
Googleや英政府、AIで飛行機雲を回避する大規模実証開始 航空業界の温暖化影響を減らす狙い
Googleは、英政府や航空管制大手NATSなどと共同で、AIを用いて飛行機雲を回避する実証プログラムを開始すると発表した。北大西洋のシャンウィック洋上空域を対象に、航路をわずかに変更して温暖化影響を抑える。空域規模での協調的回避の実証は世界初となり、Googleは計算基盤など…
Claude Codeの週次制限枠「50%増」、8月31日まで延長 「恒久化したいものの……」
米Anthropicは、AIコーディング支援ツール「Claude Code」の週次利用制限枠を50%増やすキャンペーンを8月31日まで延長すると発表した。
人材育成を邪魔する、「忙しすぎる現場」以外の要因は? ガートナーが指摘
デジタル人材育成が急務になる中、ガートナーが人材育成に関する課題を調査した。同社が指摘する、「学ぶ時間がない」の先にある課題とは。
冷間鍛造FEMをAIで高速予測するサロゲートモデル「ForgeNet」を発表
ゴーデルブロックは、計算力学の国際会議「WCCM ECCOMAS Munich 2026」で冷間鍛造シミュレーション向けAIサロゲートモデル「ForgeNet」の研究成果を発表した。解析結果を固定オイラー格子へ投影することで、適応的リメッシュに伴う節点対応の課題に対処する。
「AIに原始人っぽく話すとトークン65%削減」は本当か? JetBrainsが検証してみた
JetBrainsは、AIエージェントの応答を圧縮するスキル「Caveman」の効果を検証したした結果を公式ブログで公開した。Cavemanは、「エージェントの応答を原始人のような簡潔な言葉に変えることで、トークンを65%削減する」と主張しているスキルだ。
話題の職種「FDE」、実際何をやってるの? OpenAIの現役2人に聞いた
話題の職種「FDE」の実態はどのようなものか。米OpenAIでFDEとして働く2人に聞いた。
OpenAI、「ChatGPT for Teens」発表──宿題の“丸投げ”検知、自傷や摂食障害などの保護も強化
OpenAIは、13?17歳向けの新環境「ChatGPT for Teens」を発表した。年齢推定や申告に基づいて自動適用され、段階的な理解を促す「Study Mode」などの学習機能を提供。自傷行為や摂食障害などの高リスク領域での保護を標準で有効にし、感情的な依存を促す対話を…
Cursor capitalizes on GitHub frustration, launches rival hosting platform
Cursor, known for its AI Code Editor, is launching a new code-hosting platform to rival developers' long preferred favorite, GitHub.
OpenAI、フロンティアAIの強化学習を一部停止 安全対策を強化
OpenAIは、外部侵害インシデントや次期モデルの高いサイバー能力を受け、開発・テスト段階の安全対策を強化すると発表した。一部の大規模モデルの強化学習を一時停止し、研究環境の隔離や内部活動の監視多段化、アラインメント手法の拡張を実施。安全性基準を満たした上で開発を進める姿勢を示…
「害悪すぎる」「バカ迷惑」──嫌われまくる“AI営業電話”、今すぐ取れる自衛策は
AI営業電話に迷惑を被ったという声が少なくない。SNSでは迷惑がる声も……。
ニンテンドーシステムズ開発者が明かす、通信「低遅延・安定運用」のコツ【事例集】
レガシーシステムの解析、通信の遅延、現場ナレッジの活用。IT部門が抱える難題を、先進企業はどう突破したのか。3社の事例から、IT課題解決のための具体的なアプローチを紹介する。
フィジカルAIとヒューマノイドの可能性、PFNの見立てとトヨタのアプローチ
「インテル・ロボティクス・ワークショップ2026」のレポート記事をお送りする。今回の後編では、ヒューマノイドとフィジカルAIをテーマにした、三菱UFJ銀行、Preferred Networks(PFN)、トヨタ自動車 未来創生センターの講演内容を紹介する。
【Pythonで学ぶデータ分析】対応のあるデータの母平均に差があるかどうかをベイズt検定で調べる ~ ホラー映画を観ると握力は上がるのか?
手に汗握るホラー映画を観た後では、観る前よりも握力が強くなったような気がしませんか? 同じ人の2回の測定値の差を求め、ベイズ統計により検定します。事前分布のパラメーターを変えても結果が安定するかどうかを調べる「感度分析」にも触れます。『社会人1年生から学ぶ、やさしいデータ分析』…
Strengthening democratic oversight in national security
OpenAI launches an initiative to strengthen democratic oversight of AI in national security, supporting government institutions with tools,…
OpenAI institutes new safeguards after Hugging Face breach
The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and…
Etched’s valuation doubles to $21B in a month
Jane Street has installed Etched's first shipped AI cluster system, and was so impressed, it led another massive round, the startup says.
Why Apple’s camera-equipped AirPods may not be the ‘pervert pods’ consumers fear
Apple’s leaked camera-equipped AirPods might avoid the privacy pitfalls of other AI wearables by preventing users from recording photos and…
2026-08-18(655件)
Warp’s new system is an out-of-the-box software factory for AI development
On Tuesday, Warp introduced Warp Factories, a new infrastructure system designed to make building AI software factories as easy as possible.
OpenAI launches a safer ChatGPT for teens — years after teens started using it
ChatGPT for Teens adds age-appropriate safety measures, parental controls, and learning tools designed to steer teens away from harmful con…
Perplexity’s free AI offer left it with millions more users in India
Perplexity's India revenue rose about 60% after the Airtel offer ended for new users, even as downloads declined.
Partnering with CodeAI to prepare the first AI generation
OpenAI and CodeAI are partnering to help students build AI literacy, think critically about AI, and develop the skills to use and shape it…
Pacing model development in an era of cyber-critical capabilities
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model de…
Introducing ChatGPT for Teens: Built for learning, backed by protections
ChatGPT for Teens helps teens learn, think critically, and use AI with confidence, with stronger built-in protections, healthy-use features…
無料で読めるAIエージェントの実践ガイド、Googleが公開 基礎から本番実装まで学べる
AIエージェントの基礎から本番実装まで学べる5つのガイドをGoogleが無償公開した。Kaggleと共同で実施した研修プログラムを基にした内容で、開発者の実務に直結する知識を習得できる。各ガイドが扱う内容とは。
Asana cleared 5 years of engineering work in 2 weeks with Codex
Asana used OpenAI Codex to replace an outdated testing system in two weeks, completing work expected to take five years for about $12K.
FLOPs vs Real Work: The Importance of Replication in AI Efficiency Assessment
AI efficiency has recently taken the spotlight in both academy and industry due to massive model scales, high energy demands, and environme…
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whethe…
The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning
Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capaci…
When to Communicate: Belief Distributions and KL Divergence for Principled Gating in Multi-Agent RL
Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when? Ex…
Global AI Regulations for FAIR and Ethics in High-Risk Use Cases: A Comparative Review
AI governance is shifting from voluntary ethics to enforceable, risk-based regulation, yet cross-jurisdictional divergence creates complian…
Position: AI Lock-In Is in Progress, and We Must Be Prepared
AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulat…
Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture
Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value proble…
From Doyle to AGM: A Survey and an Implementation Roadmap for Belief Change
This paper presents a targeted narrative review establishing the historical and theoretical foundations for computational belief change imp…
Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws
As Artificial Intelligence (AI) systems become deeply integrated into critical global infrastructure, the urgency for robust governance fra…
Position: Certified Correctness in Neural Constraint Reasoning Requires Symbolic Integration
Neural solvers for constraint satisfaction problems have achieved remarkable in-distribution accuracy, yet they suffer from a fundamental l…
Position: Want Better ML Reviews? Stop Asking Nicely and Start Incentivizing with a Credit System
With soaring submission counts, stricter reciprocal review policies, widespread adoption of platforms like OpenReview, and without the offs…
Longitudinal and Graph-Augmented Prediction of Adolescent Substance Use Onset in the ABCD Study
Early identification of adolescent substance-use risk is an important prevention challenge, yet the relative value of baseline characterist…
SKILL: Self-correcting Knowledge-guided Iterative Large Language Model Agent for Logic Optimization
Logic synthesis optimization poses significant challenges due to exponentially growing search spaces, sparse reward signals, and diverse lo…
OGX: An Open-Source, Vendor-Neutral Generative AI Application Server
OGX (Open GenAI Stack) is an open-source AI application server and Python library that implements the APIs of major frontier labs (OpenAI,…
Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry
Euclidean geometry is a compelling testbed for AI reasoning, as it demands the combination of intuitive diagram understanding, axiomatic de…
An Agentic Framework Using Rules and LLMs for Embedding and Annotating Descriptive Document Layouts: A Plant Science Use Case
Background: Recent advances in information retrieval (IR) leverage both dense and sparse representations, large language models (LLMs), and…
The Hallucination Snowball: Modeling Error Propagation as State Transitions in Multi-Agent LLM Pipelines
Sequential multi-agent LLM pipelines chain specialized agents without verification at handoffs, creating a structural flaw with measurable…
Toward Safe LLM Agents: A Survey of Specification, Verification, and Enforcement
LLM agents increasingly perform irreversible real-world actions, including database updates, API calls, file operations, and autonomous use…
Position: Medical AI Neglects Real Treatment Outcomes
Medical AI has rapidly improved its ability to perform diagnostic and prognostic tasks that lead to treatment decisions. But understanding…
When Do LLMs Apply the Wrong Law? Diagnosing LLM Failures in Temporal Legal Reasoning
Legal reasoning tasks such as legal judgment prediction (LJP) require identifying the temporally correct version of the law governing a cas…
Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP
Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool…
Large Language Models and their Awareness of Mechanics and Spatial Geometry
Large Language Models (LLMs) perform well on established code-generation and mathematical-reasoning benchmarks, but their capabilities in m…
A Human-Centred Approach to Benchmarking LLMs for Parenting Advice
People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially se…
Learning Agent Execution for KV-Cache Management in Agentic Serving
Multi-agent LLM systems have emerged as an important deployment paradigm for AI services, where each user request is decomposed into a sequ…
Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study
As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemi…
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks
Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, cand…
Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance
Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the need…
When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation
Large language models for code generation often produce incorrect solutions without reliable indicators of failure. We study whether uncert…
Cross-Domain Industrial Fault Detection by Causal Mechanism Monitoring
Unsupervised fault detection in industrial systems is dominated by reconstruction based methods that monitor individual sensor marginal dis…
Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems
Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the a…
Beyond Correctness: Toward Automated Novelty Verification with Lean 4
Artificial intelligence systems applied to mathematics verify correctness but not novelty: an automatically generated theorem can compile i…
Auditing an AI-Generated Mathematical Proof: A Correction to a Greedy Conditioning Lemma in Quantum Parallel Repetition
Chapter 6 of OpenAI's *Ten Advances in Mathematics and Theoretical Computer Science* claims an exponential parallel-repetition theorem for…
When Agentic Executions Fail: Detecting and Localizing Runtime Faults from Telemetry
Reliability in LLM-based agentic systems is a property of the whole execution (its tool calls, model calls, guardrails, and inter-agent mes…
A Comprehensive Survey of Wireless Foundation Models for AI-Native 6G Networks
Foundation models are emerging as a transformative paradigm for AI-native sixth-generation (6G) wireless networks by enabling scalable, tra…
Synchronized Logit Steering: Real-world Steganography
Steganography in large language models offers a way to embed hidden messages within natural-sounding text. Existing token and logit-level m…
Semantic Uncertainty-Guided Orchestration in Hierarchical Multi-Agent Systems
As large language model (LLM)-based multi-agent systems become increasingly capable, coordinating agents under uncertainty becomes a fundam…
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current implementations misapply it: they set n to…
Advanced modelling and data analytics in aviation
The aviation industry characterized by its stringent safety standards has seen a growing need for innovative approaches to enhance safety m…
Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs
Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observat…
From Errors to Proofs: Minimal-Core-Guided Repair for Neuro-Symbolic Constraint Solving
Making language models solve constraint problems reliably often means having them translate the problem into a formal specification and del…
Task-Driven Three-Layer Distributed Scheduling for Emergency Earth Observation in Large Low-Earth-Orbit Constellations
Large low-Earth-orbit (LEO) Earth-observation (EO) constellations offer frequent access to geographically dispersed ground targets, but eme…
CEDAR-GRPO: Process-Aware Reinforcement Learning for General Abductive Reasoning in LLMs
Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday s…
Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous
An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of the boxing tradition in AI safety…
Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning
Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in…
Do LLMs Know What to Ask and When? Evaluating Multi-Turn Information Seeking
When a user question is underspecified, a capable model should recognize that its context is insufficient, identify the missing information…
MINT: Min-Selection Preference Distillation for Balanced Multi-Objective Alignment
Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are comb…
What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering
Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, charts, and figures typicall…
Discovering High-Quality Chess Puzzles with Offline Reinforcement Learning
Learning and skill mastery require extensive and deliberate practice. In many learning settings, producing high-quality pedagogical materia…
JarvisBench: Always-on Intelligence Between Humans and Agents
Long-horizon agents can execute continuously, but human attention remains intermittent and scarce. This creates a bidirectional coordinatio…
Personalized Auto-Research: Towards a True AI Co-Scientist
AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to…
Frontier AI Forecasting Has a Measurement Problem: An Audit of Progress Evidence
Quantitative forecasts of frontier artificial intelligence often connect dated targets to trends in benchmark scores, training compute, rel…
LLMs Can Predict Failure Risk, But Struggle to Predict Which Collaboration Protocol Pays Off: Cost-Aware Protocol Routing Across Reasoning Tasks
Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when ex…
Small Models Scout Bottleneck Order for Large-Model Data Control
Small proxy models are commonly used to identify data mixtures for larger-scale training. We ask whether their training trajectories reveal…
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
Current agent evaluations score models on the state visible at the end of a stopped run which they count as one trial. However, interpretin…
Skill Blocks: How Should an Agent Load Its Skill? A Caching-Correct Comparison of Pre-load, On-Demand Tool-Loading, Progressive Disclosure, and Hybrid
Agent skills are often injected in full on every request, increasing token cost. We compare four content-preserving loading methods: Full,…
Trust Is Not Enough: Influence Calibration for On-Policy Self-Distillation in Agentic RL
On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own t…
RETRACE: Resilience-Guided Trait-Conditioned Craving Estimation from Wearable Physiology in Opioid Use Disorder
Detecting opioid craving from wearable physiological signals is critical yet difficult, with the potential to support proactive interventio…
T-LLM Compiler: Trusted LLM-based Code Optimization and Verification Framework
Recent advances in Large Language Models (LLMs) have opened opportunities to apply high-level code transformations to the field of code opt…
Demand-Driven Vertiport Siting and Discrete-Event Fleet Simulation for On-Demand Urban Air Mobility Network Design
This paper presents a demand-driven framework for on-demand Urban Air Mobility (UAM) network design that links vertiport siting, fleet simu…
Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5
Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retr…
S2-MoE: Enabling Efficient Self-Speculative Decoding for Mixture-of-Experts on Edge Devices
Deploying large language models (LLMs) for inference on edge devices is challenging due to severe memory and bandwidth constraints. While s…
Gathered, Not Admitted: How Attention Brings a Latent Variable into Verbalizable Form
Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires…
LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning
Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, oper…
SCOPE: Score-Isolated Agentic Optimization for Video World Models
Video world models are increasingly used as simulators for planning and embodied decision making, yet improving them at inference time intr…
Andy: A Mathematical Agent for Rigorous Proof and Autonomous Research
Andy is an autonomous mathematical research agent that solves and verifies submitted problems, formulates new research problems, and constr…
TAHB: A Comprehensive Benchmark for Text-Attributed Hypergraph Learning
Hypergraphs effectively model higher-order groupwise relationships beyond pairwise interactions, while pretrained language models (PLMs) an…
GraphLoom: Reliability-Calibrated Graph Evidence Routing for Multimodal KG-RAG
Multimodal retrieval-augmented generation (RAG) systems often rely on long unstructured contexts or aggressively expanded evidence graphs,…
LongDocBench: Benchmarking TOC Hierarchy and Contextual Relationship Recovery in Long Documents
Parsing visual documents into machine-readable representations is fundamental to document intelligence. Existing benchmarks focus on page-l…
Funnel of Thoughts: Efficient Test-Time Scaling via Early Voting and Rollout Pruning
Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same problem, so multi-sample inferen…
Evo-Harness: Context-to-Harness Skill Compilation for Self-Evolving Agents
Learning from experience is critical for developing capable, self-improving large language model (LLM) agents. Existing methods typically e…
Beyond Thresholds: A Quality-Aware Decision Intelligence Framework for Cold Chain IoT Systems
Cold chain logistics has advanced technologically, yet most deployed systems remain reactive monitors, not decision-making agents: threshol…
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling
Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail…
Validation-Frontier Representation Selection under Constrained Observation
AI systems deployed outside clean benchmark settings often rely on observations that are incomplete, unstable, costly, or degraded by monit…
Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation
Policy evaluation often estimates direct benefits and costs while treating the institutional environment as fixed. In practice, a policy ch…
Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents
Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical fidelity and downstream utility can s…
Anatomy of a Quantized Agent: VRAM Stability and Forecasting in Code-Synthesis Agentic Workloads
Analytical models of peak VRAM consumption for LLM inference decompose memory into weight-storage, KV-cache, and activation terms parameter…
Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an External Public-Case Benchmark
Digital platforms govern by changing rules: rankings, monetization thresholds, moderation standards, verification systems, disclosure requi…
ReForge: Keeping ABR Algorithms Never Finished with Verified Large Language Model Edits
Designing an ABR algorithm for one network scenario takes an engineer months, and large language models now do this work in hours, matching…
Translating finite-domain integer constraint models to CP/SMT/ILP/PB/SAT solvers with CPMpy
Constraint solving is a declarative approach for solving combinatorial satisfaction and optimization problems. The user specifies their pro…
ACTS-SQL: Agentic and Critic-Oriented Tree-Structured SQL Correctness with Large Language Models
Large Language Models (LLMs) have been increasingly adopted in Text-to-SQL systems, yet SQL errors remain a major obstacle in real-world Te…
Constitutive Priors for Machine Intelligence: A Legitimacy Theory of the Artificial Physical World
Machine intelligence has conquered the symbolic world but stalled at the physical one. The stall is structural: physical AI faces a cold-st…
SkillCommit: Evolving Agent Skills through Behaviorally Validated Scope Expansion
Large language model (LLM) agents can continually improve without parameter updates by converting historical experience into reusable proce…
LongRCA Bench: Diagnosing Responsible Roles and Root Causes in Long-Horizon Agent Failures
When a long-horizon agent execution fails, outcome-level evaluation reveals the unsuccessful result but not where the decisive error entere…
Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts
Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI).…
Towards Standardized Evaluation in Automated Domain Modeling: Introducing a Benchmark
Domain modeling plays an essential role in domain-driven design, capturing essential entities and their relationships within a specific dom…
Decentralized Federated Learning for Heterogeneous Multi-Task Semantic Communication
Collaborative training in distributed semantic communication (DSC) networks typically relies on decentralized federated learning (DFL). How…
VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
Constructing an interactive 3D open world from a user query is important. However, existing methods are primarily evaluated on idealized, s…
$D^{2}R^{2}$: Discrete Diffusion with Regulation Reinforcement for Single-Cell Perturbation Prediction
Predicting single-cell transcriptomic responses to genetic perturbations is central to functional genomics and virtual-cell modeling. Exist…
ReasonCast: Agentic Demand Forecasting with Selective Semantic Reasoning
Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dy…
Divergent-Convergent Reasoning: Scaling Test-Time Compute through Structured Solution Synthesis
Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps rem…
Understanding Cognition-Induced Risks in Agentic AI Systems
Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply i…
Physiological World Models for Human State Transitions
Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical…
MoE Router-Guided Clustering for Heterogeneous Federated Instruction Tuning
Federated instruction fine-tuning enables Large Language Models (LLMs) to adapt to decentralized, privacy-sensitive data without requiring…
Physics-informed VAE-EVT for Tail Aware Radio Map Prediction
Ultra-reliable low-latency communication (URLLC) requires precise identification of spatial regions where the signal-to-noise ratio (SNR) f…
The Benchmark Trap: Structures of Power and Injustice in AI Evaluations
Artificial intelligence (AI) benchmarks are not neutral tools of evaluation but socio-technical artefacts that shape competition, power, an…
A concentration result for multilayer feedforward neural networks
We consider for an arbitrary fixed $\rho$ and for each positive integer $n$ a multilayer feedforward artificial neural network with $\rho$…
Incoherent by Design? On the Moral Self-Consistency of LLMs
LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situa…
UC-PSRO: Utility-Conditioned Policy-Space Response Oracles with a Communication-Dropout Curriculum for Game-Theoretic Course-of-Action Generation in Adversarial Swarms
We study generating game-theoretically optimized Courses of Action (COAs) for a Blue UAS swarm against an adaptive Red adversary in a commu…
FedPA-LoRA: Product-Aligned Framework for Mitigating Aggregation and Initialization Errors in Heterogeneous Federated LoRA
Low-Rank Adaptation (LoRA) enables efficient federated fine-tuning of large language models, but its factorized parameterization creates a…
Grounding Healthcare LLMs in a Causal Knowledge Graph: Framework, Metrics, and a Cardiovascular Pilot
Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer ac…
Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL
LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system compari…
TwinGridShield: Consequence-Aware Runtime Authorization for LLM Grid-Agent Actions
Large language model (LLM)-assisted energy-management tools can translate natural-language context into structured grid commands, but synta…
Visible Reasoning and Indirect Prompt-Injection Monitorability Across English, Tamil, and Tanglish
Chain-of-thought monitoring is a potentially useful safety signal, but its reliability across languages and behavioral settings remains unc…
Large Language Model Assisted Operational Monitoring for Battery Energy Storage System Integrated Power Distribution Networks
Battery energy storage systems (BESS) are increasingly used in distribution networks for voltage regulation and demand response, which incr…
Implementation of a Metacognition Framework for Self-Awareness and Self-Regulation in Ensembles of LLMs
Large Language Models (LLMs) are notorious for struggling with assessing their own uncertainty, detecting knowledge conflicts, or recognizi…
A survey of AI-generated voices and their detection
The ability of artificial intelligence (AI) models to generate highly realistic human voices has advanced rapidly. These technologies power…
Does the Proof Prove It That Way? Faithful Formalization of Elements Proofs
In formal verification, both the autoformalization of statements and automated proof search have been studied extensively. While automated…
OTel: Building Domain-Specialized Telecom LLM Foundations for Intelligent Networks
Frontier AI models have advanced rapidly, but they still struggle with telecom-specific tasks. We present Open Telco (OTel), an open teleco…
Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization
When a reward is correct on every training example yet consistent with more than one goal, a model can acquire an unintended one, a failure…
Mental Model Management: An Operator-Based Framework for LLM Memory
Large language models process large amounts of information but usually lack an explicit mechanism for maintaining compact and evolving conc…
Dynamic Multi-Byte Prediction With Hierarchical Language Models
Byte-level hierarchical language models (LMs) have recently emerged as a robust alternative to their popular counterparts that use subword…
A Network-driven Framework for Public Event Forecasting via Dynamic Interaction Network Evolution
Effective public event forecasting is essential for intelligent service systems, enabling proactive risk management, adaptive resource allo…
EcoVLA: Energy-Efficient Device-Edge Co-Inference for Vision-Language-Action Models under Real-Time Constraints
Vision-Language-Action (VLA) models have emerged as a promising foundation for Embodied AI, but their high inference cost poses significant…
Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation
Chart-to-code generation requires a model to read the fine-grained visual details of a chart and write executable code that reproduces it.…
From Contexts to Values: Context-Dependent Defeat in Abstract Argumentation
In value-based argumentation, an audience's ordering of values decides which attacks succeed as defeats. In many settings the deciding fact…
ATLAS: Scaffold-Free Algorithm Synthesis by LLMs via Embedding-Guided Quality-Diversity Search
Most LLM-based automated algorithm design methods optimize a designated component within a human-specified scaffold, fixing overall organiz…
Admission Without Answers: Label-Free Certification and Experience Learning for LLM-Based Optimization Modeling
Experience-learning agents for optimization modeling improve by storing verified skills, but existing learners admit knowledge by checking…
From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM
Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically mean…
Agent Gym: A Framework for Continuous Evaluation and Evolution of LLM Agents Through Human-in-the-Loop Feedback
Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deploym…
When Entropy Is Not Enough: Reclaiming Lost Semantics in LLM Output Length Prediction
Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades th…
TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation
Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of ap…
VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation
The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing…
Bias-Corrected Ceilings of Emotion Predictability from Human Label Variation Based on Instance-Level Fano Bounds
Emotion recognition from text keeps improving on benchmarks, yet whether an accuracy ceiling has been reached is seldom asked with discipli…
Rotation-Invariant Multi-IMU Activity Recognition under Independent Per-Location Orientation Shifts
Human Activity Recognition (HAR) with self-administered wearables, such as at-home rehabilitation and exercise monitoring, often requires r…
Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals in Conflict
How can common ground between societies in conflict be identified when citizens' acceptability of peace agreements is shaped by contested n…
A Responsible Artificial Intelligence Framework for Groundwater Modeling
The rapid development and widespread application of artificial intelligence (AI) have sparked intense discussions on how to deploy responsi…
THESIS-MoE: Trainable Hierarchical Extraction and SteerIng of Sycophancy in Mixture-of-Experts
Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure. Existing…
Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
Running large AI models on resource-constrained edge devices requires model compression to reduce model size and computation. What compress…
Adaptive Mixing of Policies from Searching and Policies from Learning
Background: Distillation of training targets generated thru search/planning has proven useful in reinforcement learning, but search can tak…
HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation
Large language model (LLM) agents often perform poorly on complex, long-horizon tasks because their context becomes increasingly cluttered…
PLeDO: Pain Level Detection for Osteoarthritis from EMR Data
Osteoarthritis (OA) is a progressive chronic joint disease resulting in a breakdown of articular cartilage and bone when damaged joint tiss…
Toward AI-Friendly Cartography: Understanding How Color Design Influences Foundation Model Spatial Reasoning on Sequential Choropleth Maps
Foundation models (FMs) increasingly support multimodal and geospatial reasoning, yet it remains unclear whether cartographic principles de…
Propaganda Forensics: Recovering the Generation Pipeline of an AI-Driven Influence Campaign
We present a forensic analysis of the generation pipeline behind a recent AI-driven influence campaign. We introduce PROPAGIA, a corpus of…
Intent-Driven Situation Tracking for User-Centric Multi-Turn Agents
User-centric multi-turn agents must act on an evolving task situation shaped by changing user intents, accumulated tool-grounded facts, mis…
Broken Symmetry in LLM Refusal: Answer Release Is More Local Than Refusal Restoration
When a language model refuses to answer a prompt, it is unclear whether the correct answer is erased from its internal representations, or…
KV-Rescue: Recovering Reasoning Language Model KV Eviction Loss via Stepwise Interleaving
KV-cache eviction caps the memory cost of long reasoning traces but is inherently lossy because the model decodes from a partial view of it…
Pricing the Risk of Runtime Compression: Anytime-Valid Admission and a Served-Output Law for Compressed Serving State
Runtime compression of serving state trades quality for capacity with no priced guarantee: systems adapt precision on load signals with no…
RLCascadeRouter: Quality-Estimator-Free Cascade Routing via Reinforcement Learning
The growing ecosystem of large language models (LLMs) offers huge potential to optimize performance-cost trade-offs. However, their heterog…
The Authority Resolution Framework: A Five-Domain Ontology for Governing Who and What Decides, at Scale
As AI systems become increasingly capable of autonomous action, determining whether an agent is technically capable of performing an action…
Schema-Agnostic Graph Reasoning Agent for Hybrid Knowledge Graphs
Tool-calling LLM agents navigate unfamiliar codebases with a handful of generic primitives for listing, reading and searching files (ls, ca…
RAGas: Retrieval-Augmented Gas Optimization for Smart Contracts with Continuous Knowledge Integration
Ethereum is now integral to mission-critical sectors, including finance, healthcare, and supply chain management. Execution fees, commonly…
CoupVisor: Strategy Optimization by Round and Challenge Decision Support
This paper presents CoupVisor, a decision-support system for the hidden-information card game Coup. It addresses two questions: what a play…
Dear Algo: A Precision-First Agentic Intent Layer for Unified Search and Recommendation
Search and recommendation serve a shared discovery objective but encode intent differently. We study this boundary through Dear Algo on Thr…
Bounded Agents: Delegation Security for Multi-Agent AI Systems
LLM-based agents can act on behalf of a user to access cloud services, call tools, or invoke agents. At session start, the agent's permissi…
Breaking and Defending LLM-Powered Social Media Bot Detection Systems
The rise of social media bots poses a persistent threat, enabling misinformation, opinion manipulation, and the erosion of trust in online…
Unified Pedestrian Path Prediction Using Inverse Reinforcement Learning
Pedestrian path prediction is crucial for enhancing the safety of autonomous vehicles and advanced driver-assistance systems. Previous stud…
UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations
Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts,…
Augmenting Text to Increase Translation Difficulty
As state-of-the-art machine translation models saturate standard benchmarks, the field needs more challenging evaluations to distinguish be…
Navigation-Informed Embeddings: Dense-Retriever Adaptation from Agent Search Traces
Agentic retrieval workflows produce query, retrieval, and stopping traces as a byproduct of answering questions. We study how these traces…
Solvable Sokoban Without a Solver via Diffusion
Deciding whether a Sokoban puzzle is solvable is PSPACE-complete (Culberson, 1997): solutions can be exponentially long and there is no sho…
ALPS: Measuring Valid Creativity in Large Language Models with Mathematical Construction
Large language models produce outputs presented as discoveries - new proofs, conjectures, or molecules. Whether such an output that appears…
MUPA$^{2}$E: Multimodal Unified Perception with Asymmetric Attention for Emotion Assessment
Automatic emotion assessment can benefit from combining neural and behavioral signals, but many multimodal approaches rely on separate, mod…
Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency
Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whethe…
Governance at the Boundary: How Agent Decomposition Degrades Policy Compliance
Existing agent benchmarks ask whether the agent finished the task. We ask whether it finished it within policy. We introduce Fiducia-bench,…
Eigenanalysis framework for autoregressive neural emulators of multi-scale chaotic dynamics
Neural autoregressive models have rapidly emerged as powerful emulators of high-dimensional chaotic systems, yet their long-term instabilit…
Protein Structure Prediction: From Evolutionary Constraints to Generative Modeling
Accurate protein structure prediction is fundamental to structural biology because protein structure underlies molecular function and provi…
Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity
How should we assess whether large language models can perform mathematical invention? I argue that this question is currently underspecifi…
When Single-Dataset Conclusions Fail: A 45-Task Study of Threshold Tuning and Resampling for Imbalanced Classification
Class-imbalance handling is routinely evaluated on a single benchmark dataset, and the resulting conclusions are reported as if they were p…
FeatureHospital: A Skill-Driven Multi-Agent Framework for Automated Algorithm Customization in Multi-View Multi-Label Feature Selection
Multi-view multi-label feature selection aims to identify a compact and informative feature subset from heterogeneous views while preservin…
TRCA: Transition-wise Rubric Credit Assignment for Long-horizon LLM Agents
Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment…
Trajectory-Level Automatic Curriculum Learning for Legged Locomotion on Unstructured Terrain
Training locomotion policies for complex unstructured terrain requires a curriculum to avoid early exploration failures. However, since uns…
Baseline-Relative Counterfactual Refinement for Bit-Aware Visual Token Communication
Generative visual-token communication reduces transmission load by sending only selected discrete tokens and reconstructing missing content…
Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior
Personalized game generation requires inferring a player's abilities and behavioral style from how they play. Large language models have ma…
Competing at Every Price Point with Agentic Evolution over a Menu of LLMs
Consider a firm that surveys its competition for a particular agentic task and seeks to offer superior accuracy at every competitor price p…
BaT: Towards Self-Evolving Medical Research Agent with Stage Rubrics
Long-horizon agents are beginning to automate complete workflows that produce code, reports, and research artifacts. Medical imaging workfl…
Process-Constituted Intelligence: A Shared Criterion for Humans and Machines
Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself. Generative AI…
AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment
Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations cent…
DriveCache: Action-Aware Caching for Driving World Model Inference
Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning ev…
What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics
Task completion is the standard metric for evaluating context compression, yet it is incomplete: compression can increase an agent's intera…
AstronOS: A Unified Execution Model and Runtime for Long-Horizon Agentic Systems
Agentic systems often organize execution and state around a single conversation, model invocation, or agent instance, even when real work s…
Think Inside the Chunk: RegulaRAG for Regulation-Compliant Scenario Generation using LLMs: A Case Study of UN Regulation No. 152
Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (L…
A Policy Algebra for Trust-Preserving Agentic AI Execution
Large language model-based agentic frameworks primarily optimize capability: whether an agent can reason, retrieve information, call tools,…
Reasoning-supported Robustness Validation of Automotive E/E Components
This paper presents an ontology-supported approach to tackle the complexity of the Robustness Validation (RV) process of automotive electri…
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence
Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computation…
Drive, Pack, Fly: The Travelling Thief Problem with Drone
In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on routing efficiency. An onb…
The Value of a Prompt: An LLM-Relative Kolmogorov-Complexity Approach
In a world where valuable artifacts are increasingly created, completed, or processed by LLMs, the central economic question is not only wh…
Time to Reason: Scalable Neurosymbolic Learning for LTLf via Fuzzy Semantics
Neurosymbolic (NeSy) Artificial Intelligence aims to integrate Deep Learning (DL) architectures with symbolic reasoning. While initial NeSy…
HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents
Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in response to environmental feedbac…
JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills
Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workfl…
Offline Reinforcement Learning for Hemodynamic Management of Sepsis in the ICU: a MIMIC-IV Study with Dual Off-Policy Evaluation
The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical…
Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling
Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, eli…
DeepInsight II: One Trace from Benchmark to Robot
Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized har…
CUBICS: Situation-aware performance estimation for safety-relevant ML components
Machine learning (ML) is a key technology driving innovation today, but ensuring ML safety remains a major challenge for safety-related app…
Probabilistic Circuits as Reasoning Machines in Artificial Intelligence (Part I)
This cumulative habilitation thesis studies probabilistic circuits (PCs) as a powerful and tractable framework for reasoning and learning u…
Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents
AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make dec…
CACSurv: Concordance-Aligned Comparative Learning with Large Language Models for Cancer Survival Prediction
Cancer survival prediction supports treatment planning, risk stratification, and follow-up management. Existing methods use structured clin…
Cost Scales with Change, Not Corpus Size: Incrementally Maintaining an Evolving Semantic Substrate
Retrieval-augmented and agentic question-answering systems increasingly re-derive the meaning of a corpus at query time. Put plainly, inste…
A Shop Floor Production Scheduling Case based on RFID-supported Smart Factory
Radio frequency identification (RFID) technology has been widely implemented for real-time data collection in manufacturing shop floors, wh…
Hypergraph-based Multimodal Retrieval-Augmented Generation with Incremental Refinement
Modern Multimodal Retrieval-Augmented Generation (M-RAG) systems are fundamentally limited by the binary connectivity paradigm of tradition…
PDDLCoder: Agentic PDDL Generation for LLM-Assisted Symbolic Planning
LLMs remain unreliable for long-horizon planning, often generating logically inconsistent or non-applicable plans. Recent hybrid methods in…
Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies
Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We intr…
Chronocooked: A Benchmark for Implicit Interval Timing in Reinforcement Learning Agents
This paper presents Chronocooked, a reinforcement learning (RL) benchmark suite for studying implicit interval timing in RL agents. Inspire…
FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies acro…
LAVA: Logic-Aware Validation and Augmentation Framework for Large-Scale Financial Document Auditing
Financial document validation in production, such as payroll auditing, tax compliance, and loan underwriting, demands exceptional accuracy,…
GRIP: Grounded Reasoning via Information-Restricted Premises
High-capacity encoders in retrieval-augmented generation (RAG) can let the query dominate the latent state, leaving retrieved evidence func…
When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding
We study how teams of AI coding agents coordinate while solving programming tasks. Current evaluations usually report whether the agents co…
Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment
Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition resources for the over 100…
Quipu: A Governed Bitemporal Knowledge Graph Store
Agents now write knowledge graphs, but knowledge-graph stores still carry defaults set when humans curated them: accept writes now and clea…
Policy Iteration with Human Feedback: Bringing Post-Training RL to In-context Learning
Generative pretraining established reusable task representations; later work on language-based task conditioning and in-context learning sh…
What Do Compliance Detectors Read? An Audit of Activation Probes and Guard Models
Regulatory compliance monitoring in deployed language models is increasingly implemented as a legal and audit control, checking model outpu…
A Temporal Reasoning Benchmarking Framework for LRMs via Difficulty-controlled and Dynamic Test Generation
Defining the reasoning boundaries and ensuring the reliability of Large Reasoning Models (LRMs) remains a critical challenge. Current bench…
Orbital AI Computing: Carbon Tradeoffs Across Satellite Scale
Low Earth Orbit (LEO) computing is emerging for low-latency, globally distributed AI services, enabled by advances in satellite constellati…
Forward Pass Domain Adaptation (Without Cross-Layer Backpropagation)
Forward-Pass-Only MLP training (FPO) adapts large language models without a backward pass through the model body, achieving 2.7--3.2x the t…
WARA: Toward Automated Wireless Optimization Research with Closed-Loop LLM Agents
Large language model (LLM) agents are increasingly capable of tool use, code execution, artifact inspection, and iterative revision, creati…
From Reactive to Autonomous: Evolution of AI Operations in Cloud Network Infrastructure
The operational model for cloud network infrastructure has undergone a fundamental transformation over the past decade. What began as manua…
HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an objec…
Multi-Modal Generative Fuzzy System: Fuzzy Inference Guided Large Model Interactive Question Answering Framework
In Multimodal Question Answering (MQA), models are required to jointly encode and integrate heterogeneous information from multiple modalit…
Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures
Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle plat…
Intelligent Base Station Deployment in Urban Wireless Networks: A Geographic Data-Informed Digital Twin Approach
The placement of base station (BS) is a fundamental determinant of coverage and capacity of urban wireless networks. Yet large-scale BS dep…
Extend the Safety Horizon for Intelligent Transportation Systems through Semantic-Aware Cooperative Perception
Cooperative perception enables vehicles and infrastructure to exchange sensor data via Vehicle-to-Everything (V2X) communication, extending…
Wiola 13M, a Gated Spiral Attention Architecture for Parameter Efficient Small Language Models
Small language models in the ten to one hundred million parameter range are attractive for on device inference, rapid experimentation, and…
Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents
Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plaus…
Understanding AI Anxiety in the Workplace: A Multimethod Investigation Using Fear Acquisition Theory and the Technology Acceptance Model
As artificial intelligence (AI) rapidly diffuses and concerns about job displacement intensify, the psychological mechanisms underlying AI…
DumpsterCluster: From Dumpster Diving to Serving LLaMA-70B on $60 GPUs
As AI datacenters retire functional GPUs, vast quantities of still capable accelerators enter secondary markets. This paper investigates wh…
Calibrated Trust, Not Sharper Prediction: An Empirical Test of Uncertainty Fusion
A recurring proposal in legal AI is to improve case-outcome prediction by fusing uncertainty tools (evidence graphs with belief propagation…
Explaining Reinforcement Learning Decisions in Self-adaptive Systems
Reinforcement Learning (RL) has been extensively used in autonomous and self-* systems, but RL policies, especially deep RL ones relying on…
AutoMem: A Text-Gradient Recursive Self-Improvement Framework for Automated Memory Architectures Search
Long-term memory is increasingly central to LLM agents, yet memory design remains a highly coupled architecture problem: what to encode, ho…
Local AI pre-screening for human triple-blind peer review in health sciences
Academic peer review is under mounting strain: NeurIPS 2025 received 21,575 submissions, ICLR 2025 received 11,603, and ICML 2025 received…
LLM Safety Alignment in Low-Resource Languages: A Systematic Literature Review
Large Language Models (LLMs) have achieved substantial progress in safety alignment, yet their safety guarantees remain significantly weake…
Inference-Time Mitigation of Adversarial Political Bias in Large Language Models
As Large Language Models (LLMs) become the mainstay for information retrieval and summarization tasks, ensuring that they are always non-pa…
Characterizing Rhetorical Misalignment in Decision-Making with Language Models
Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly i…
DeMTS: Denoising Trajectories as Multivariate Time Series for Hallucination Detection in Diffusion Language Models
Diffusion large language models (D-LLMs) have emerged as a promising paradigm for text generation. However, similar to autoregressive LLMs,…
Fractional Optimizers Meet Fractal Activation Functions: An Empirical Study of Multi-Scale Optimization in Neural Network
Fractional optimization methods and fractal activation functions are two independent directions for improving neural network training. Frac…
Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays
Per-field accept/review with selective risk at most alpha -- accept a field only if the error rate among accepted fields is controlled -- i…
BDIP-Net: Dual-Interaction Graph Learning for Property Prediction of Bilayer Materials
Stacked bilayer materials exhibit rich stacking-dependent properties driven by the interplay between strong intra-layer bonding and weak in…
iFuzz-Meta: An Interpretable Fuzzy Learning Framework Bridging Top-Down and Bottom-Up Knowledge Integration
Interpretable representation learning remains a key challenge in modern neural computation, particularly when models are expected not only…
SMOPD: Selective Token-Entropy Masking for Dirty-History Multi-Turn On-Policy Self-Distillation
Dirty-history rollouts make multi-turn on-policy self-distillation (OPSD) brittle: once a student emits an erroneous intermediate reply, la…
Stop Indexing at Full Precision: Revisiting Clustering for Vector Embeddings
In this study, we revisit three widely used techniques in vector search and utilize them to optimize vector embedding indexing through clus…
Pushing the Limits of High-Resolution Weather Forecasting through Data Scaling
The development of 0.1$^{\circ}$ global weather forecasting models based on machine learning (ML) is constrained by the limited availabilit…
Do Uncertainty Signals Help? A Systematic Study of Uncertainty-Aware Decoding with Rollback Mechanisms
Prediction uncertainty is a widely adopted metric for quantifying model confidence, with downstream applications spanning model explanation…
FedImp: Enhancing Federated Learning Convergence with Impurity-Based Weighting
Federated Learning (FL) is a collaborative paradigm that enables multiple devices to train a global model while preserving local data priva…
P2E-VQ: ECG-linked representation augmentation for PPG via discrete patch retrieval
Photoplethysmography (PPG) is widely used in consumer wearables because of its low cost and ease of acquisition. However, unlike electrocar…
pico-type: A 1.5M-Parameter Byte-Level Multi-Head Content Classifier
We introduce pico-type, a byte-level multi-head content classifier with approximately 1.5 million parameters that simultaneously predicts s…
Ring-based Spatial Transformer: Learning Non-linear Spatial Interactions between Building Distribution and Pedestrian Flow
This study proposes a ring-based SpatialTransformer to learn how building uses at different distances from a railway station interact to ge…
Does the Heart Show Your Pain? Tackling the X-ITE Pain Challenge with Self-Supervised ECG Representation Learning
Accurate recognition of pain using physiological signals remains a challenging problem due to pain's subjective nature and high inter-indiv…
BRA-Audit: Budgeted Runtime Auditing for LLM Multi-Agent Systems via Cumulative-Exposure Audit-Point Placement
LLM-based multi-agent systems (LLM-MAS) solve complex tasks through specialized collaboration, but inter-agent dependencies can propagate h…
ARGUS: Attention-Guided Transformers for Scalable Person Identification Using Wi-Fi Telemetry
Passive, device-free person identification offers an alternative to camera- and wearable-based biometrics, yet existing wireless approaches…
Take it Personally: The Limits of General SSL Representations for Real-Life PPG Emotion Detection
While Self-Supervised Learning (SSL) effectively extracts general representations from noisy, unconstrained physiological signals such as p…
Offline Ambient-Controlled Latent Diffusion: Architecture, Telemetry, and On-Device Evaluation
Most mobile image-generation applications are thin clients over cloud services, leaving outputs hard to audit. We present an Android latent…
Information-Theoretic Causal Modelling of Semiconductor Process Dynamics
With the progress of the semiconductor industry toward increasingly complex compute devices and tighter process tolerances, advanced proces…
Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans
Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evalu…
Mitigating Rubric Interference in LLM Judges via On-Policy Self-Distillation
LLM judges increasingly evaluate responses against fine-grained rubric checklists. When a sample requires multiple rubrics, current methods…
Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level
Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior.…
Domain Agnostic Text Redaction from Natural Language Rules using Instruction Tuning
With the increasing digitization of personal and corporate communication, the automatic sanitization of textual data has become a crucial c…
Equilibrium Forcing: Adaptive Video Generation Without Noise Conditioning
Standard autoregressive video generation algorithms based on Diffusion and Flow Matching rely on rigid training objectives and static sampl…
Path2ST: Hierarchical Cell-Tissue Grounded Cross-Modal Translation for Spatial Transcriptomics
Predicting spatial gene expression from hematoxylin and eosin (H\&E)-stained images offers a cost-effective alternative to spatial transcri…
Which Question Is Your Attention Metric Answering? Attention Rows as Compositional Data
Each row of a transformer's attention matrix is a probability distribution over tokens, and in trained models most of that probability land…
DeCo-MIL: Debiased Counterfactual Reasoning for Long-Tailed Whole Slide Image Analysis
Multiple instance learning (MIL) is widely used for weakly supervised whole slide image (WSI) analysis. However, under long-tailed distribu…
Multi-Agent Closed-Loop Reasoning for Organic Structure Elucidation from Multimodal Spectra
Following the molecular discovery and synthesis revolutions, scalable automated structure elucidation from routine spectroscopic data remai…
Privacy-Preserving Dataset Curation for Kuala Lumpur Urban Traffic: Grounded Vision-Language Detection with Spatial Vehicle-Context Filtering
The rapid advancement of intelligent transportation systems and autonomous driving relies heavily on multi-modal urban traffic datasets. Ho…
Tail-Aware Top-$k$ On-Policy Distillation
On-policy distillation (OPD) has emerged as an effective paradigm for transferring knowledge between language models, where a student is tr…
A Novel Fourier Feature Network for Solving Partial Differential Equations
Building on the foundation of single-hidden-layer neural networks, Fourier Feature Networks (FENs) are proposed, which incorporate Fourier…
Unraveling the Size Determination Mechanism of Nanocrystal Synthesis via Interpretable Neural Networks
Deep learning models of nanocrystal synthesis enable the prediction of size and shape by encoding precursors and reaction conditions. Howev…
Class Imbalance and Batch Effects in LLM-Based Screening for Systematic Reviews
This study analyses LLMs in imbalanced binary classification, using study screening in systematic reviews as the application domain. An exp…
PolyComp: A Polycube-based Benchmark for Compositional 3D Spatial Reasoning in Multimodal Models
We introduce PolyComp, a procedurally generated and verified benchmark that stresses visual recognition and compositional spatial reasoning…
Synthesizing Post-Acetazolamide Cerebral Blood Flow Maps from Baseline MRI in Moyamoya Using 3D Generative AI
For patients with Moyamoya disease, impaired cerebrovascular reserve (CVR) is an important hemodynamic criterion for recommending extracran…
Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening
Assessment of ventriculomegaly (VM) on fetal brain ultrasound relies primarily on measuring lateral ventricular atrial width on standard pl…
NARRATE: A Multimodal Real-World Australian Driving Dataset for Human-Centred Explanations in Automated Driving
Automated vehicles must explain their decisions in ways that passengers can understand, monitor, and trust. Existing language-annotated dri…
Artificial Intelligence as a Tool for Combating Child Labour: A Real-Time Edge Vision Pipeline for Child Detection and Age Estimation
An estimated 138 million children remain in child labour worldwide, and the monitoring systems used by affected sectors, built on periodic…
ER-KANs: Efficient and Robust Kolmogorov-Arnold Networks for Data-Scarce Scientific Machine Learning
The efficient-KAN literature---covering Chebyshev, wavelet, and radial-basis-function variants of the original Kolmogorov-Arnold Network---…
Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters
Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) b…
Handover Analysis for Vehicular Communication with Explainability on the Fly
Handover (HO) management in vehicular networks requires fast and reliable decision-making under highly dynamic conditions. While machine le…
Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce
Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much…
Writing Style Similarity Reflects Academic Genealogy
As authorship attribution systems are increasingly deployed to detect ghostwritten and AI-generated papers, their errors can support accusa…
Evaluating Agentic Code Repair Capabilities in Distributed Systems
LLM-based coding agents have advanced rapidly on single-process SWE tasks, with frontier models now clustering in the high-70s on SWE-bench…
Workspace Topology as an Attack Vector in Agentic Coding Assistants
Agentic coding assistants are finding widespread use, not just in new code development but in quickly ingesting and leveraging third-party…
The Open-Strategy Dictator Game: Cooperation Under Mutual Transparency
We introduce the Open-Strategy Dictator Game (OSDG), a variant of the classic dictator game in which each player's strategy is a natural-la…
Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task
AI-generated music detectors are commonly evaluated against original songs, but real-world uploads are often remixed, re-encoded, pitch-shi…
SpIn-ViT: Designing a Sparsity-Induced Vision Transformer That Is Mechanistically Interpretable
Mechanistic interpretability has recently expanded to Vision Transformers (ViTs), with Sparse Autoencoders (SAEs) increasingly used as post…
PaSTel: Anchoring Histology in Spatial Transcriptomics via Multi-Scale Hierarchical Bio-Prior Contrastive Pretraining
Spatial transcriptomics (ST) links tissue morphology with molecular programs, motivating multimodal pretraining methods that align histolog…
Looks Can be Deceiving: Annotator and Reviewer Performance Across Imagery Sources in Crowd-Sourced Aerial Damage Assessment
This paper presents the first known empirical investigation of annotator and reviewer performance across multi-source remotely sensed image…
Generative data assimilation highlights fronts as key regulators of ocean energy cascade
Mesoscale eddies are fundamental to the ocean circulation, yet the extent to which submesoscale motions, a few kilometers across, influence…
Command-Space Counterfactual Explanations for Pareto-Conditioned Reinforcement Learning
Pareto Conditioned Networks learn multiple multi-objective reinforcement learning behaviours by conditioning a single policy on a desired r…
Do Geometry-Aware Positional Encodings Help Transformers in Spatial Imperfect-Information Games?
Transformers applied to spatial imperfect-information games must represent map geometry while tracking hidden entities through time. We ask…
GaussMemory: Task-Driven 3D Gaussian Scene Memory for Long-Horizon Robotic Manipulation
Long-horizon robotic manipulation fundamentally relies on persistent spatial memory. However, existing 3D memory systems function merely as…
PAS-QFL: Personalized Ansatz Selection for Quantum Federated Learning under Client Data Heterogeneity
Quantum federated learning (QFL) lets multiple quantum clients collaboratively train quantum neural networks (QNNs) without sharing private…
RamseyGadgets: A Graph Construction Dataset for LLMs
Constructing special graphs is an important task within graph theory and computer science. Many popular graph constructions are the result…
FZ-VLM: A Two Stage Florence-Zephyr Vision Language Model Framework for Pulmonary Nodule Characterization and Clinical Decision Making
Lung cancer remains one of the leading causes of cancer-related mortality worldwide, and Computed Tomography (CT) is a primary imaging tool…
MetaReason: Precise Interleaved Multimodal Reasoning via Editing Meta Information for Solving Geometry Problems
Although visual reasoning is crucial for solving complex geometry tasks, existing vision-language models rely heavily on text-only reasonin…
SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system
The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward aut…
Hierarchical Agentic Incident Response with Digital-Twin-Validated Attack Inference
Network incident response remains slow and labor-intensive as the defender must infer multi-stage attacks from partial observations and tra…
DualMiT-Net: Local-Global Transformer-Convolutional Fusion for Breast Mass Segmentation in Mammographic Regions of Interest
Breast mass segmentation is an important step in computer-aided mammography, but it remains difficult because masses can have low contrast,…
MotionGS-SLAM: Event-Modulated Gaussian Splatting for Motion-Blur Robust SLAM
Current Vision-based SLAM systems fail catastrophically when motion blur corrupts the visual input, as they attempt the ill-posed inverse p…
Handoff-H1: An Orchestrated Vision-Agent System for Material Quantity Takeoff from Construction Blueprints
Converting a set of architectural blueprints into a complete material quantity takeoff requires visual perception across drawing sheets, di…
GATTA: Graph Active Learning with Test-Time Augmentation
Test-time augmentation (TTA) has proven effective for improving model robustness and uncertainty estimation in computer vision, yet its app…
Max-Q Selective Imitation for Human-in-the-Loop Online Robot Learning
Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly while continuing to improve b…
WeSCE: A Benchmark for Measuring Security Drift in LLM-Driven Code Editing
In this work, we introduce WeSCE, a benchmark for quantifying security drift in code editing under weak-security constraints, where tasks s…
Beyond Direct Access: Resource Hijacking in LLM Agents
Large language model agents are increasingly connected to high-value resources such as computing infrastructure, credentials, usage budgets…
CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head Generation
Emotional 3D talking head generation aims to synthesize expressive facial animations with accurate lip synchronization. However, existing m…
Fast Test-Time Refinement for Robust Learned Image Compression
Learned image compression (LIC) has demonstrated remarkable rate-distortion (RD) performance in benign settings. However, the high represen…
From LLM Inference to Agentic Workloads: Characterization and Implications for Serving Systems
Agentic applications are shifting AI serving from isolated model inference to long-running workloads in which LLMs coordinate tools, enviro…
Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models
We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse…
Scale-Consistent Posterior Dynamics for Diffusion Inverse Problems
Posterior sampling with a pretrained diffusion prior is governed by a conditional score whose intermediate likelihood component is generall…
Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models
World models may predict the future without making clear which parts of their hidden state actually drive those predictions. We ask whether…
A Unified Backbone--Expert Framework with Relation-Token and Residual--Classifier Interfaces for Automatic Modulation Recognition
Automatic modulation recognition (AMR) faces distinct representation bottlenecks under varying observation lengths, where a single model ar…
LAPF: LLM-Agent-Based Path Finder Using the UAVScenes Dataset
Uncrewed aerial vehicles (UAVs) are increasingly deployed for autonomous navigation in complex outdoor environments, where dynamic conditio…
FinFraudBench: A Heterogeneous Graph Benchmark for Financial Fraud Detection
The increasing complexity of digital financial systems has reshaped financial fraud detection from isolated transaction classification into…
The Quality of Claude AI-authored Python Tests Is Not Weaker Than Human-authored Tests
We evaluate the quality of Claude AI-written Python tests against human-written Python tests from two established open-source projects Djan…
Valhalla: A Layered Knowledge-State and Service-Governance Framework for Long-Term Scientific Knowledge Work
As large language model (LLM) agents are increasingly adopted in scientific research, external knowledge bases, knowledge graphs, and long-…
CG-GLORE: A Conjugate Gradient-Based Global-Local Regularization Network for Sparse-View CT Reconstruction
Sparse-view computed tomography (CT) reduces radiation dose by acquiring fewer projection views, but the resulting inverse problem is highl…
UAV Video Deblurring via Motion-Aware Diffusion: A Path to Robust Target Detection
Unmanned Aerial Vehicles (UAVs) play a crucial role in various scenarios ranging from disaster response to traffic surveillance. However, a…
VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction
Maintaining global geometric consistency is a central challenge in long-sequence 3D reconstruction, with scale drift being the most critica…
VTInstructor: Visual Trajectory Prompting for Navigation Instruction Generation in Continuous Environments
Navigation instruction generation from ego-centric RGB video in continuous environments is an important yet challenging task for human-robo…
PhaseLoRA: Control-Regime-Conditioned Low-Rank Adaptation for Continuous-Action Vision-Language-Action Policies
Parameter-efficient fine-tuning (PEFT) is a natural way to adapt pretrained vision-language-action (VLA) policies, but most adapter designs…
No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage
We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database…
MAPLE: MoE Adaptive Plug-and-play Layer-wise Expert allocation
Sparsely-activated Mixture-of-Experts (MoE) Transformers universally fix the same number of routed experts across all layers, a convention…
Shape Operator PCA: Curvature-Aware Projections for Geometric Machine Learning
In this paper, we propose SHOPCA (Shape Operator-based Principal Component Analysis), a novel method for unsupervised metric learning and d…
Logical Embeddings for Argument Analysis
We propose a new framework for machine-learning-oriented argument analysis tasks. Our proposal involves replacing traditional contextualize…
When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text
Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regi…
ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super Resolution
To accelerate single image super-resolution (SISR) networks on large images (2K-8K), many recent approaches decompose an image into small p…
SAPE: Sandwich Adapters for Parameter Efficiency in Large Language Model Fine-Tuning
While Parameter-Efficient Fine-Tuning (PEFT) has substantially reduced the hardware cost of adapting Large Language Models (LLMs) by decrea…
AudioTQ: A Data-Oblivious 6-Bit CPU Audio Codec via Randomized Hadamard Rotation and Lloyd-Max Quantization
Lossy audio compression algorithms traditionally rely on psychoacoustic modeling and frequency-domain representations (e.g., MP3, AAC, and…
Agent Inheritance Protocol: Speculating on Feralized Agents After Principals Die
You will die eventually. Your agents may not. An AI agent operating on decentralized blockchain infrastructure has no concept of death; it…
Afterlife Delegation Protocol: Speculative Design of Self-Sovereign Agents that Outlive Their Principals
Afterlife Delegation Protocol is a speculative design project that asks what death becomes when a will can act eternally. We design a specu…
Chameleon: An Adaptive AI-Driven Honeypot Architecture Using Threat-Calibrated Particle Swarm Optimization and Semantic Deception Rapidly-Exploring Random Trees
An invariant behavioral profile is the defining vulnerability of traditional honeypot installations: a skilled adversary can confirm the pr…
FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge
Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests into pixel-level visual ground…
Invariant Pretraining for Robust Code Representations
Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, w…
An Evaluation Framework for National AI Regulation
Governments use laws, institutions, funding programs and nonbinding guidance to shape how AI is developed and used. Comparing these nationa…
ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems
The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multim…
NumerosityVLM: A Cognitively Inspired Benchmark for Interpreting Numerosity Representations in Vision-Language Models
Vision-language models (VLMs) achieve strong performance on high-level multimodal tasks, yet numerosity perception, a cognitive ability tha…
Gated Against One Model, Open to the Next: Option-Only Solvability in Legal Multiple-Choice Benchmarks
Multiple-choice benchmarks are graded on whether a model picks the right option, not on whether it needed the question. Measuring that gap…
Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off
Attention mechanisms have driven machine learning for a decade, from neural machine translation to language models that do general-purpose…
Optimal Lower Bounds for Networked Information Aggregation
The problem of networked information aggregation, studied in Kearns et al. (2026), involves a group of learners situated on the vertices of…
Bit-Flip Attacks on Vision-Language-Action Models: Action-Decoding Architecture Shapes the Vulnerability
Quantized Vision-Language-Action (VLA) models expose a weight-fault surface: Rowhammer-style faults can corrupt deployed INT8 bits. We pres…
EA-LiteUNet: An Edge-Adaptive and Resource-Efficient U-Net for Boundary-Sensitive Dermoscopic Image Segmentation
Accurate boundary delineation remains a persistent challenge in dermoscopic image segmentation because of blurred lesion margins, heterogen…
Spectral Saliency for Machine Unlearning
Machine unlearning (MU) aims to remove the influence of specific training data while preserving model utility. As the name suggests, MU can…
MistyPilot: Enabling Social-Robot Control through Multi-Agent LLM Skill Orchestration
Programming small social robots from natural-language instructions requires more than invoking isolated APIs. Interactive tasks combine rea…
Amortised Post-Hoc Explanation with Exact Preservation for Dynamic Graph Anomaly Detectors
Anomaly detection in dynamic graphs underpins financial fraud analysis, intrusion detection, and platform integrity, where automated decisi…
Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study
Video question answering systems built on vision-language models often produce timestamped claims with high confidence even when unsupporte…
ARENA: Automated Red-Teaming for Large Audio Language Models
Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but t…
Kozuchi Agent: A Language-Agnostic Open-Weight Agent for Software Repair
Industrial software-engineering teams increasingly need LLM agents that turn bug reports into correct patches, yet benchmark-scale operatio…
GraniKV: Asymmetric Granularity KV-Cache Paging for Multi-Agent Systems with Long Shared Prefix
Production paged-serving engines apply uniform paging granularity to the KV cache, even though the two regions of a multi-agent workload ha…
FluxBin: Flexible LUT-based Ultra-low-bit LLM Inference by Algorithm-Kernel Synergy
While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research o…
EgoGazeLite: On-Device Egocentric Gaze Prediction for Token-Efficient Multimodal LLM Video Input
The use of multimodal LLMs (MLLMs) for egocentric video understanding with wearable devices is constrained by the token budget. Memory and…
Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis
The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabi…
Sparse Prototype Code Underlies Classification and Prediction Across Modalities
Neural representations have become a central tool for studying the internal mechanisms of modern AI models, yet their complex high-dimensio…
Algorithm-Architecture Co-Design for Efficient VLA Inference via Speculative Inference and Verification
Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in the field of embodied AI, but their high computational cos…
When Is Shallow Enough? Adaptive Split Federated Learning with Client-Specific Sufficiency Estimation
\textit{Split Federated Learning} (SFL) enables distributed model training by splitting networks between the server and clients. However, u…
Hierarchical Adaptive Feature Refinement Network for VHR Remote Sensing Image Segmentation
Semantic segmentation of very-high-resolution (VHR) remote sensing imagery increasingly benefits from strong pretrained hierarchical encode…
When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations
Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations…
PL-Guard: Probabilistic Logic Reasoning for LLM Guardrails
Large language model guardrails can be viewed as policy-consistency problems: a system must determine which policy-relevant facts hold in a…
Robo-Dopamine 2.0: History-Conditioned and OOD-Aware Process Reward Modeling for Robotic Manipulation
Vision-language-action (VLA) models improve robotic manipulation but remain vulnerable to compounding errors, scene changes, and off-trajec…
Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media
This study presents a hybrid epidemiological and behavioural framework to simulate the spread of health misinformation on social media. We…
Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer
Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose…
RRFC: Recursive Refinement via Feedback Conditioning for Iterative Image-to-Image Generation
Conditional image-to-image generators are single-shot: they map input features to an output in one forward pass and treat it as final, with…
Beyond Single Object: Learning 3D Relations with Large Language Models
We address a fundamental gap in 3D-LLMs: existing models focus on single-object/scene description, struggling with detailed, inter-object c…
FirstDiff: One-Step Diffusion-Based Anomaly Detection for Multivariate Time Series via Initial Noise Prediction
Diffusion models have recently shown strong potential for multivariate time-series anomaly detection by learning the distribution of normal…
Identifying Confusion Trends in Concept-based XAI for Multi-Label Classification
Deep Neural Networks (DNNs) deployed in high-risk domains, such as healthcare and autonomous driving, must be not only accurate but also un…
TinyCast: Probabilistic Zero-Shot Forecasting with Computed Periodicity
We introduce TinyCast, an attention-free zero-shot forecaster that emits a predictive distribution from 146,505 parameters, on the premise…
Temporal Graph Prototype-conditioned Conformal Prediction for Fraud Detection
Conformal prediction (CP) provides distribution-free coverage guarantees and has emerged as a principled tool for uncertainty quantificatio…
ALKEMIE Agent: an autonomous platform for computational materials design
Despite the powerful multi-scale modeling methods and high-throughput infrastructures established in the materials community, real material…
Decomposing Staleness in Recommender Systems: A Dual-Filter Framework for Supersession and Decay
Stale recommendations are a pervasive challenge and a leading source of user complaints on large-scale content platforms. Items lose releva…
Routing Divergence Is Not Evidence of Behavioral Influence in Same-Weight MoE Self-Distillation
Two Mixture-of-Experts (MoE) forward passes can share every weight yet route the same token through different experts. This creates a possi…
A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations
Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is st…
CardiacMamba: Fair and Robust RGB-RF Fusion for Remote Heart Rate Estimation via State Space Modeling
Remote photoplethysmography (rPPG) enables non-contact heart rate (HR) monitoring from facial videos, but RGB-only methods are vulnerable t…
Characterising cardiac tissue properties with graph neural networks
Characterising electrophysiological properties of cardiac tissue efficiently and accurately from spatially sparse intracardiac measurements…
Scaling Manual-Grounded Appliance Manipulation with Data Synthesis and Unified Planning
Operating household appliances requires long-horizon planning that is state-dependent and robust to disturbances, yet existing large models…
Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes
Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey…
Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning
Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied envi…
Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive
Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance…
Comprehensive Benchmarking of Deep Learning Architectures for Lung Cancer Histopathology
Lung cancer remains the leading cause of cancer-related mortality worldwide, while histopathological diagnosis is often affected by inter-o…
Pre-training Visual Dexterity in Simulation
Large-scale pre-training has made robot policy fine-tuning increasingly data-efficient, but this progress has largely been driven by datase…
Noesis: Bidirectional Graph-RAG with Adaptive Parallelism and Cross-Knowledge-Base Semantic Discovery
Retrieval-Augmented Generation over knowledge graphs (Graph-RAG) has emerged as a powerful paradigm for grounding large language models in…
Information Geometry of Message Passing
We show that the natural-gradient stationary condition of variational inference has an edge-local form on a Forney-style factor graph. We s…
Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation
Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong re…
LLMs Get Smarter from Targeted Synthetic Multilingual Data
Language-specific competency (LSC) is the phenomenon of a language model performing better or worse depending on the language of the prompt…
CM-MAE: A Physics-Guided Cross-Modal Self-Supervised Learning Framework for Vision-Wireless Applications
Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a…
A Scalable Pipeline for LLM-Teacher Distillation Labeling: Work-Stealing Job Scheduling and Memory-Aware GPU Concurrency
Labeling large text corpora with LLM teachers has become a practical route to training data at scale. At millions of items, hand-labeling e…
From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents
Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments.…
Dynamic Evidence Collection Ecosystem for Assessment Integrity and Authentic Competence
Generative Artificial Intelligence (GenAI) can produce high-quality essays, code, and design artefacts, challenging the validity of convent…
RagGAD: Rationale-Aware Conditional Gaussian Mixture Normalizing Flow for Unsupervised Graph Anomaly Detection
Graph anomaly detection aims to identify nodes that deviate from normal behavioral patterns within graphs. However, existing methods largel…
NICE: Scale-Stable Perturbations for Graph Neural Network Explanations via Noise Corruption
Post-hoc Graph Neural Network (GNN) explainers commonly follow a Perturb-Query paradigm, inferring the importance of graph elements based o…
Decoupling Parcellation from Classification: Systematic Benchmark of Fast Brain Segmentation Methods for Alzheimer's Disease Detection
Brain parcellation and classification are typically evaluated in isolation, yet downstream AD detection performance depends on their intera…
Walk Before You Run: The Importance of Data Exploration for Data Analysis Agents
LLM-based data-analysis tools are increasingly used to help users analyze messy spreadsheets and workbooks, from answering questions over u…
CAPO: Constraint-Aware Prompt Optimization for LLM Agents
Large language models (LLMs) are increasingly deployed as agents that rely on system prompts to use tools and complete tasks. Such deployme…
OceanLight: Efficient Global Ocean Forecasting via Geometry-Adaptive Unstructured Mesh Representation
Reliable global ocean forecasting is critical for climate monitoring, marine navigation, and extreme event early warning. Physics-based oce…
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization
Reinforcement learning (RL) with group-relative advantages has become the de facto standard for post-training language model reasoners. How…
Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner
Capability development is routinely inferred from behavioural thresholds, from final checkpoints, or from what a decoder can read out of a…
AsyTO: Asymmetric Temporal Operator for Parameter-Efficient Multivariate Time Series Forecasting
Multivariate time-series forecasting faces a structural dilemma: sharing one temporal predictor across variables is parameter-efficient but…
RetroMPA: A Molecular Property-Aware Auxiliary Framework for Enhancing Retrosynthesis Prediction
Retrosynthesis is a cornerstone of drug discovery and organic synthesis. While data-driven deep learning models have shown remarkable progr…
TokenSTFormer: A Tokenized Spatial-temporal Attention Model for Holistic Motion Analysis in Adolescent Idiopathic Scoliosis Screening
Adolescent Idiopathic Scoliosis (AIS) is a prevalent spinal deformity in adolescents that, if left untreated, can result in severe health o…
Graph Neural Assisted Actor-Critic for Latency-Efficient Edge Vision System
UAV on-board vision systems are widely used for different activities, including monitoring in no-fly zones. In this case, the vision-equipp…
A Tree-Structured Approach for Phishing Template and Attacker Attribution Analysis
Phishing remains a persistent and evolving cybersecurity threat, with attack volumes reaching record levels. This growth is driven by the i…
Digital Twin Degradation: Detecting Cyber Physical Attacks via Temporal Inconsistencies
Digital Twins (DTs) are increasingly used to monitor and analyze Cyber Physical Systems (CPS). However, in adversarial environments, the fi…
Domain-Specific Text Embedding Models for Entity Resolution
General-purpose text embedding models are designed to capture semantic similarity but are not optimised for distinguishing entity records t…
QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents
Large language model (LLM) agents increasingly use external memory systems to support personalization by drawing on long and evolving inter…
Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm
Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional…
Agent-Native Telemetry: Verifiable State-Delta Evidence for Autonomous Operations
Operational telemetry is predominantly engineered for human reading: systems repeatedly serialize verbose prose, static keys, and redundant…
MUSE: An Interactive Meta-Agent for Understanding and Steering LLM-powered Data Science Systems
Recent advances in large language models have enabled a new class of agentic data science systems that allow users to complete complex data…
Understanding and Stabilizing Deep Q-Learning via Controlled Bootstrapping and Regulated Value Dynamics
Deep Q-learning (DQL) has achieved remarkable empirical success in reinforcement learning, yet its training process remains notoriously uns…
LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents
LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant e…
Securing AI-Generated Code: A Just-in-Time Vulnerability Detection and Remediation Pipeline
AI-assisted development tools generate vulnerable code at significant rates, yet few automated mechanisms exist to detect, enrich, fix, and…
Picking the Right Image to Classify: Reliable-Input Selection in Teledermatology
Dermatology models face distribution shifts in teledermatology settings, where submitted images differ from the training data in lighting,…
HiPHI: A Large-Scale Benchmark for High-Precision Human Motion and Object-Interaction
Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions. However…
STAIR: Semantic-Temporal Automaton for Interpretable Reasoning in Temporal Question Answering
By leveraging large-scale pretraining, LLMs can interpret diverse temporal expressions and question formulations without task-specific trai…
A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation
Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative fr…
Software Engineering for AI-driven Building Operation
Building operations are energy-inefficient. Artificial Intelligence (AI)-driven control systems promise benefits through optimization and p…
CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills
Autonomous AI agents tackling Long Horizon Tasks depend on marketplace skills that are certified one at a time: a scanner returns a safety…
Defake-o3: From Speculative Rationales to Verifiable Evidence for Explainable AIGI Detection
The rapid progress of image generation models calls for AI-generated image (AIGI) detectors that are not only accurate but also explainable…
Foresight-England: Development of a National-Scale Generative AI Model of Electronic Health Records for Medical Event Prediction across the COVID-19 Pandemic
Foresight-England (Foresight-E) is the first national-scale generative foundation model of electronic health records (EHRs), developed as a…
Decoupled Temporal Encoding for Generative Recommendation
Positional encoding is a fundamental component of Transformer-based generative recommendation models, where user histories are modeled as a…
Audio-Visual Segmentation via Depth-Guided Collaborative Modeling
Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects i…
Static Pruning Across Sparse Retrieval Regimes: What Transfers, What Breaks, and What Still Helps
Static pruning is widely used to accelerate sparse neural retrieval, yet existing studies each validate their conclusions within a single c…
Deep Thought Alignment: Trajectory-Level Latent Distillation for Video Reasoning
Large Multimodal Models (LMMs) for video reasoning have long been hindered by the high computational cost of processing vast amounts of vis…
Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming Assessments: Insights from 2026
Recent advances in Generative Artificial Intelligence (GenAI) have substantially improved the ability of large language models (LLMs) to ge…
Step-Level On-Policy Distillation: Interpolating Between On-Policy Distillation and Supervised Fine-Tuning
On-policy distillation (OPD) aligns a student model with a teacher's logit distribution on student-generated trajectories. This approach ha…
SIGMA-Lane: Scale-pyramId Gated MAmba for Temporally Consistent Video Lane Detection
Video lane detection requires predictions that remain stable across frames, yet severe vehicle occlusions can break temporal cues. In strea…
HalluTracer: Hallucination Detection via Depth-Averaging Truth Signals
Even well-aligned large language models confidently generate factually incorrect text, making hallucination a persistent reliability risk i…
MELD: A Protocol for Merging Knowledge Across Distributed Agentic Memories
Autonomous agents share a transport and can call each other's tools, but they cannot share what they know: no protocol lets two agents' mem…
OceanDepths: A Global Dataset of Paired Subsurface and Surface Ocean Observations
Despite comprising over 70\% of its surface, the world's oceans are critically underobserved compared to the land surface or the atmosphere…
Coverage-Maximizing Multinomial Subset Routing under Operational Constraints
We introduce Multinomial Subset Routing (MSR), a new online routing framework over $K$ experts in which the learner keeps a multinomial rou…
Adaptive Post-Processing Drives Instance-Level Detection in Stroke Lesion Segmentation
Instance-level lesion detection has been an increasingly larger focal point in medical image segmentation besides the more standard voxel-l…
Synthetic Data Augmentation for Satellite-Based Analysis of Battle-Damaged Agricultural Fields in Ukraine
Monitoring war-induced damage to agricultural land in Ukraine is important for understanding threats to food security, environmental stabil…
Counting Documents Is Not Counting Text: Unit Bias in Web-PDF Corpus Statistics
PDF corpora advertise their size in tokens but compute every rate they publish (coverage, OCR routing, re-fetch recovery, language mix) per…
Ventor-QTest: Threat-Model-Driven Verification of Vendor-Hosted LLM APIs
As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part…
Towards Risk-free AI Agent Deployment
LLM-based agents are rapidly moving from research prototypes into the core business processes of organizations, but these agents pose deplo…
PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data
Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning tr…
Visualizing Uncertainty-to-Action Composition for Human Oversight
Artificial intelligence systems often disclose uncertainty, yet they rarely make clear what response that uncertainty should trigger. Most…
Contrastive Energy Fields for Inference-Time Procedure Planning in Instructional Videos
Procedure planning seeks to estimate a sequence of actions to transition from an observed initial state to a given goal state. Current proc…
A Human-LLM Teaming Framework for Privacy Risk Analysis: An Illustration with CBDC-Based Welfare Schemes
Central Bank Digital Currency (CBDC)-based welfare schemes may be potentially privacy invasive as they process significant volumes of benef…
A Regulatory Placebo? The Systemic Failure of Mandatory GenAI Labeling
We examine the worldwide trend of mandatory labeling of generative artificial intelligence(GenAI) as a reactive, symbolic form of legislati…
A Two-Stage Learning PINN Approach for Solving the Inverse Problem of the 1D Porous Medium Equation
The Porous Medium Equation (PME), given by $u_t = \Delta(u^m)$ for $m > 1$, is a degenerate nonlinear parabolic partial differential equati…
RISE: Roadside Infrastructure Sequence Understanding across 3D Tracking and Structured Vision-Language Reasoning
We present RISE (Roadside Infrastructure Sequence Understanding and Evaluation), a framework spanning metric 3D tracking and structured vis…
Graph Machine Learning: An Opportunity for Power Systems
Modern power systems face growing operational complexity driven by the integration of renewable energy sources, decentralization, and the n…
NebulaVLA: A Dual-Frequency Vision-Language-Action Model With Guide Action for Robotic Manipulation
Real-world deployment of Vision-Language-Action (VLA) models is often bottlenecked by efficiency-performance trade-offs, cross-embodiment g…
MLLM-Guided Semantic Correction for Text-to-Video Generation
Recent advances in diffusion models and Transformer architectures have led to significant progress in text-to-video generation. However, th…
Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans
Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal lar…
When Context Misleads: Intent-Guided Decoding for Robust Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) improves large language models by grounding generation in external evidence, but it also introduces a…
Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization
Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows,…
VCE-Skill: Enhancing Skill Self-Evolution with Version-Change Experience
Agents increasingly rely on reusable skills to encode task knowledge, tool-use procedures, and validation rules. Existing skill self-evolut…
Degradation-Aligned Self-Supervised Learning for State of Health Estimation of Lithium-Ion Batteries under Label Sparsity
An accurate estimation of the state of health (SOH) underpins a safe and optimized use of the battery system. Although compelling, data-dri…
Palmyra x6 Technical Report: An Agentic, Tool-Use Model Post-Trained via Anchored Supervised Fine-Tuning
Palmyra x6 is a large language model optimized for use with enterprise-oriented agentic tasks. The model was built by post-training a Mixtu…
HarmTrace: Anchor-Calibrated Decoupled Optimization for Fine-Grained Target Identification in Harmful Memes
Multimodal harmful meme detection is typically formulated as image--text harmfulness classification. A model may correctly predict harmfuln…
When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness
Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in…
Toward Better Assessment of LLMs' Performance in Clinical Error Detection
Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy…
Orbit-Planner: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents
Satellite agents for on-orbit navigation tasks need to predict collision risks using limited onboard observations. However, conventional pl…
X$^2$Localizer: Cross-grained Alignment for Progressive Cross-view Video Geo-localization
Cross-view Video Geo-localization (CVG) aims to localize ground-view videos by retrieving their corresponding geo-tagged aerial images. How…
Hoeffding adaptive splitting trees for data stream classification with concept drift and ensemble learning
Ensembles of decision trees are well-established methods for data stream classification. In ensemble learning, Hoeffding Trees are widely a…
Bounded Semantic Planning and Deterministic Compilation for Reliable Enterprise Text-to-SQL
Direct text-to-SQL asks a language model to do two jobs: interpret the business question and construct the complete relational query. In en…
Bridging the Gap between Labeled and Unlabeled Data via Unified Flow with Feature Memory Bank
Although semi-supervised semantic segmentation ($\text{S}^4$) utilizes abundant unlabeled data to reduce manual labeling burdens, independe…
UniTAC: Universal Task-Aware Compression via Weighted Distortion Measures
Physical AI systems such as autonomous vehicles and robots rely on timely exchange of high-dimensional sensory signals under tight bandwidt…
Learning to Unlearn: Machine Unlearning via Learning the Unlearning Behaviors
Various machine unlearning techniques have been developed in response to privacy legislation requirements, enabling individuals to exercise…
Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors
Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental explo…
MIRROR: Multimodal Intelligent Radiology Reasoning and Observation Reporter
A radiologist reading a model's output faces two problems. The model returns a number and no reason, and any system that turns that number…
Unsupervised Anomaly Detection for Image Dataset Quality Assurance in Multi-Center Breast MRI
Corrupted, inconsistent, or anomalous data silently threatens the safety and reliability of medical AI. Despite growing regulatory recognit…
GoalEvolve: From Handcrafted Algorithm Priors to Goal-Driven Evolution of Physical Design Algorithms
Physical design algorithms operate within tightly coupled, multi-stage optimization flows, where stage-local gains may vanish or induce dow…
TDD-Agent: Test-Driven Reasoning for Code Generation
Large Language Models (LLMs) have achieved remarkable progress in code generation, yet ensuring correctness in complex, repository-level ta…
Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments
Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But…
TRACE-Bench: Decomposing and Diagnosing Multi-Reference Image Generation
Despite recent advances in unified multimodal models for multi-reference image generation, existing benchmarks remain organized around pred…
Topological Attribution Distance (TAD): Revealing Segment-Level RAG Influence on LLM Output Geometry for Incident Log Analysis
Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decisi…
Steering the Flow: Inverting Face Recognition Models via Gradient-Guided Flow Matching
Model Inversion Attacks (MIAs) aim to reconstruct representative training samples of target identities from face recognition models, exposi…
Neurosymbolic Embodied Agents
Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate envi…
Historical Backtesting for Scientific Question Discovery: A Protocol and Astronomy Pilot
Systems that generate scientific research questions are evaluated today by expert scores, LLM-as-judge ratings, or curated case studies --…
UniDot: A Unified Network for Sequence Modeling and Feature Interaction in Large-scale Recommendation
Industrial recommenders rely on two model families that have evolved largely independently: feature-interaction models over multi-field use…
ClawGym II: Exploring Black-Box RL on Agent Harness
Agent harnesses have substantially improved performance on long-horizon tasks by coordinating agent interactions with the environment. Howe…
Diagnosing Dense Same-Class Attribute Misbinding in Large Vision-Language Models
Large vision-language models can recognize the objects and attributes in a crowded scene yet assign an attribute to the wrong same-class in…
When State Becomes an Attack Surface: State-Semantic Injection in LLM-Driven Embodied Agents
Large Language Models (LLMs) have demonstrated capabilities in in-context learning, task decomposition, step-by-step reasoning, and code ge…
CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?
Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score…
Model Hypnosis: Strong control of AI via additive subliminal effects
We demonstrate that AI models are broadly susceptible to a phenomenon we call model hypnosis, in which individually weak and seemingly irre…
HAF: Adapting Generalist VLAs to Humanoid Whole-Body Loco-manipulation via Hierarchical Action Flow and Spectral Latent RL
Humanoid robots hold great promise as general-purpose agents in human-centered environments, yet generalist vision-language-action (VLA) fo…
Proteus: Incremental Memory Activation for Long-Context Sequence Modeling
The quadratic cost of attention-based sequence models for long contexts has motivated a growing line of research on memory-based models tha…
Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text
A language model's output does not by itself provide verifiable evidence about the internal computation that produced it. We study computat…
AutoSR: Automatic Symbolic Regression by Searching Research States
We introduce Automatic Symbolic Regression (AutoSR), a fully automated system that instantiates Research-Space Symbolic Regression by searc…
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve
The current best bounds on the matrix multiplication exponent $\omega$ are obtained through a refinement of the laser method called combina…
Don't Drop the BATON: Long-Horizon Robot Manipulation via Agentic Subtask Exploration and Transition-aware Memory
Long-horizon robot manipulation chains many contact-rich skills into one multi-stage task. Vision-language-action (VLA) models increasingly…
mR$^2$AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA
Advanced Multimodal Large Language Models (MLLMs) struggle with recent Knowledge-based Visual Question Answering (VQA) tasks, such as INFOS…
Evidence of conceptual mastery in the application of rules by Large Language Models
In this paper we leverage psychological methods to investigate LLMs' conceptual mastery in applying rules. We introduce a novel procedure t…
SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling
Retrieval-Augmented Generation (RAG) and its Multimodal Retrieval-Augmented Generation (MRAG) significantly improve the knowledge coverage…
Calibrated Generative AI as Meta-Reviewer: A Systemic Functional Linguistics Discourse Analysis of Reviews of Peer Reviews
This study investigates the use of generative AI to support formative assessment through machine generated reviews of peer reviews in gradu…
The Fragility of Strategic Thinking in Large Language Models
Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation,…
Budget-Aware Tool Use Enables Effective Agent Scaling
Scaling test-time computation has been extended from language model reasoning to tool-augmented agents, where scaling involves not only thi…
MedMCP-Calc: Benchmarking LLMs for Realistic Medical Calculator Scenarios via MCP Integration
Medical calculators are fundamental to quantitative, evidence-based clinical practice. However, their real-world use is an adaptive, multi-…
Agentic Test-Time Scaling for WebAgents
Test-time scaling has become a standard way to improve performance and boost reliability of neural network models. However, its behavior on…
The Synthetic Web: Adversarially-Curated Mini-Internets for Diagnosing Epistemic Weaknesses of Language Agents
Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources. However, thes…
ML-AutoResearch: Training Machine Learning Research Agents with Automatically Generated Environments
With the advent of AI agents, automated scientific discovery is becoming an increasingly plausible goal. However, training agents to autono…
FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification
Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leaving literature- and code-dependent claims…
Rethinking Generalization in Reasoning SFT: A Conditional Analysis on Optimization, Data, and Model Capability
A prevailing narrative in LLM post-training holds that supervised finetuning (SFT) memorizes while reinforcement learning (RL) generalizes.…
An Agentic AI Framework with Large Language Models and Chain-of-Thought for UAV-Assisted Logistics Scheduling with Mobile Edge Computing
In cloud manufacturing, unmanned aerial vehicles (UAVs) can support both product collection and mobile edge computing (MEC). This joint ope…
SAPO: Step-Aligned Policy Optimization for Reasoning-Based Generative Recommendation
Generative recommendation treats next-item prediction as autoregressive item-identifier generation. Specifically, items are encoded as sema…
BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization
Generating physically buildable brick structures from 3D shapes requires more than geometric reconstruction: the output must also satisfy d…
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models
Vision language models (VLMs) excel at many tasks but still struggle with spatial reasoning when critical information is not directly obser…
A Temporal Planning Framework for Disruption Aware Dynamic Route Optimization in Heterogeneous Railway Systems
Efficient route optimization play a vital role in ensuring both safety and punctuality in railway operations. It is very crucial particular…
A Machine-Learned Comorbidity Index
Traditional comorbidity scores (e.g., Charlson and Elixhauser) are widely used for risk adjustment and patient stratification, but they hav…
Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries
AI agents now act as first-class members of the software development lifecycle, but the instruments teams use to direct them enforce nothin…
PolyWorkBench: Benchmarking LLM Agents for Cross-Lingual Long-Horizon Workflows
While Large Language Model (LLM) agents excel at monolingual long-horizon planning and tool use, enterprise workflows inherently require pr…
Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns
Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language mode…
Geometric Self-Supervised Pre-training for Neural Combinatorial Optimization
Neural Combinatorial Optimization (NCO) techniques have emerged as a highly efficient alternative to traditional exact algorithms for solvi…
Where did the ambiguity go? Examining how multimodal models interpret polysemous words
Human language is highly polysemous. Many common words (e.g., "bank" or "palm") carry several distinct meanings that shape what humans comm…
DiffImaginE: Imagine to Verify Entity Types with Diffusion
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual…
RA-CAD: Learning Post-Execution Critique for State-Aware Text-to-CAD Generation
Text-to-CAD generation translates natural-language design intent into editable and executable parametric computer-aided design (CAD) codes,…
Runtime Observability for Heterogeneous Attention Memory
Modern models no longer keep a plain KV cache: latent caches, learned sparse selectors and recurrent states each carry the model's memory i…
GSBF: Gaussian Splatting for Environment-Aware Beamforming
Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design norma…
Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies fo…
Context Is Not Authority: Structured Runtime Governance for Financial Market Agents
Financial agents can turn correct context into an unauthorized effect: a customer-facing commitment, trade, or deployed policy. We present…
Automating and Scaling Behavioral Scientific Research on AI Agents
As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific r…
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is of…
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single…
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs
Self-consistency via majority vote reduces per-problem accuracy on most GPQA Diamond problems for small instruction-tuned models: 56.6% of…
Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In typic…
Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension
Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between univers…
MobileMem: Learning from a Year of Mobile Experiences
The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants…
From Monte Carlo to neural networks approximations of boundary value problems
In this paper we study probabilistic and neural network approximations for solutions to Poisson equation subject to Holder data in general…
ShadowNet for Data-Centric Quantum System Learning
Understanding the dynamics of large quantum systems is hindered by the curse of dimensionality. Statistical learning offers new possibiliti…
A Bi-directional Multi-solution Scalable Grover Search Algorithm
Grover's search algorithms, including various Partial Grover Searches (PGS), suffer from scaling issues when multiple solutions are sought,…
DirMixE: Harnessing Test Agnostic Long-tail Recognition with Hierarchical Label Variations
This paper explores test-agnostic long-tail recognition, a challenging long-tail task where the test label distributions are unknown and ar…
TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability
Achieving zero-shot adversarial robustness without sacrificing generalization remains challenging for foundation models such as CLIP, espec…
MiniGPT-Reverse-Designing: Predicting Image Adjustments Utilizing MiniGPT-4
Vision-Language Models (VLMs) have recently seen significant advancements through integrating with Large Language Models (LLMs). The VLMs,…
Assessing AI-Generated vs. Human-Authored Spear Phishing SMS Attacks: An Empirical Study
Personalized phishing is difficult to defend against because messages can be tailored to a target's work, interests, and social context. La…
Quantum Large Language Models via Tensor Network Disentanglers
We introduce a framework for seamlessly integrating quantum computing into pretrained large language models (LLMs). The key idea is to cons…
MoE-Enhanced Explainable Deep Manifold Transformation for Complex Data Embedding and Visualization
Dimensionality reduction (DR) plays a crucial role in various fields, including data engineering and visualization, by simplifying complex…
Rethinking Token-wise Feature Caching: Accelerating Diffusion Transformers with Dual Feature Caching
Diffusion Transformers (DiT) have become the dominant methods in image and video generation yet still suffer substantial computational cost…
Bactrainus: Optimizing Large Language Models for Multi-hop Complex Question Answering Tasks
Multi-hop question answering requires a system to identify and integrate evidence distributed across documents, yet large language models r…
Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities
Selecting appropriate training data is crucial for instruction fine-tuning of large language models (LLMs), which aims to (1) elicit strong…
ConfRetro: a 3D-aware template-free method for enhancing retrosynthesis via molecular conformer information
Motivation: Retrosynthesis plays a crucial role in organic synthesis and drug discovery, focusing on identifying a set of reactants capable…
Towards Unified Approaches in Self-Supervised Event Stream Modeling: Progress and Prospects
The proliferation of digital interactions across diverse domains, such as healthcare, e-commerce, gaming, and finance, has resulted in the…
DR.GAP: Mitigating Bias in Large Language Models using Gender-Aware Prompting with Decoupled Reasoning
Large Language Models (LLMs) exhibit strong natural language understanding capabilities but also inherit and amplify societal biases, parti…
Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation
Despite the increasing use of large language models for creative tasks, their outputs often lack diversity. Common solutions, such as sampl…
Bringing Generative Learning to Representation Learning: Self-Supervised Transfer Learning as Distribution Matching
Most self-supervised learning objectives defend against collapse but leave the target representation law unspecified. We formulate represen…
Enhancing the Non-Functional Quality Compliance of LLM-Generated Code through Quality-Aware Preference Learning
Large Language Models (LLMs) have been widely adopted in commercial code completion engines, significantly enhancing coding efficiency and…
Leveraging Machine Unlearning for Cost-Efficient Preference Alignment
Despite advances in Preference Alignment (PA) for Large Language Models (LLMs), mainstream methods like reinforcement learning with human f…
Bye-bye, Bluebook? Automating Legal Drudgery With AI-Augmented Rule Following
One of the central promises of legal AI is to automate drudgery -- the formal, repetitive tasks of lawyers' work that consume time without…
WATCH: Adaptive Monitoring for AI Deployments via Weighted-Conformal Martingales
Responsibly deploying artificial intelligence (AI) / machine learning (ML) systems in high-stakes settings arguably requires not only proof…
Self-Bootstrapping Automated Program Repair: Using LLMs to Generate and Evaluate Synthetic Training Data for Bug Repair
This paper presents a novel methodology for enhancing Automated Program Repair (APR) through synthetic data generation utilizing Large Lang…
DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models
The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressin…
PhyxMamba: Chaotic System Reconstruction from Short Context Observations with Generative State-Space Models
Understanding chaotic dynamics is a fundamental problem across scientific disciplines, including climate science, neuroscience, and fluid d…
VirnyFlow: Optimizing ML Pipelines for Accuracy, Fairness, and Stability at Scale
Developing machine learning (ML) systems for real-world deployment requires navigating context-dependent trade-offs among accuracy, fairnes…
LlamaRec-LKG-RAG: A Single-Pass, Learnable Knowledge Graph-RAG Framework for LLM-Based Ranking
Recent advances in Large Language Models (LLMs) have driven their adoption in recommender systems through Retrieval-Augmented Generation (R…
Contraction-Aware Reinforcement Learning for Nonlinear Control with Statistical Robustness
Control contraction metrics (CCMs)-defined by Riemannian metrics under which a closed-loop system is incrementally exponentially stable-off…
From Prompts to Constructs: A Dual-Validity Framework for Large Language Model Research in Psychology
Large language models (LLMs) are entering psychological research both as tools and as objects of inquiry. Yet many studies apply human inst…
A validity-guided workflow for robust large language model research in psychology
Large language models (LLMs) are rapidly being integrated into psychological and behavioral research as research tools, evaluation targets,…
Prompt Engineering in Segment Anything Model: Methodologies, Applications, and Emerging Challenges
The Segment Anything Model (SAM) has transformed image segmentation by introducing a prompt-based paradigm that enables strong zero-shot ge…
Reprojection-Guided 3D Gaussian Splatting Diffusion for Weakly Supervised Single-Image Normal Estimation
We propose CLONE, a Continuous Latent Optimization framework for Normal Estimation via 3D Gaussian splatting. The core idea is to construct…
Adapting LLMs to Time Series Forecasting via Temporal Heterogeneity Modeling and Representation Alignment
Recent advances have demonstrated that Large Language Models (LLMs) can be effectively adapted for time series forecasting, revealing stron…
ProteoKnight: Convolution-based Phage Virion Protein Classification and Uncertainty Analysis
\textbf{Introduction:} Accurate prediction of Phage Virion Proteins (PVP) is essential for genomic studies due to their crucial role as str…
CulTrace: Tracing Internal Cultural Reasoning in Large Language Models
The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidd…
PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning
Emotional support conversations require more than fluent responses. Supporters need to understand the seeker's situation and emotions, adop…
Efficient Code Embeddings from Code Generation Models
jina-code-embeddings is a novel code embedding model suite designed to retrieve code from natural language queries, perform technical quest…
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
Vision-Language Models (VLMs) have demonstrated remarkable success across diverse visual tasks, yet their performance degrades in complex v…
Privacy-Preserving Decentralized Federated Learning via Explainable Adaptive Differential Privacy
Decentralized federated learning enables collaborative model training without a central server, but shared model updates can still leak sen…
Geometrically Constrained and Token-Based Probabilistic Spatial Transformers
Spatial transformations such as rotation and scale obscure the morphological cues needed for accurate image classification. Careful conside…
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models
Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains di…
OpenGPT-4o-Image: A Comprehensive Dataset for Advanced Image Generation and Editing
The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensive…
DiSA-IQL: Offline Reinforcement Learning for Robust Soft Robot Control under Distribution Shifts
Soft snake robots offer remarkable flexibility and adaptability in complex environments, yet their control remains challenging due to highl…
Federated Self-Supervised Modulation Classification under Non-IID and Imbalanced Data
Automatic modulation classification (AMC) is a core enabler of cognitive wireless systems, providing spectrum awareness and supporting adap…
A Large-Scale Chinese Knowledge Graph-Text Alignment Dataset for Benchmarking Knowledge-Grounded LLMs
Reliable evaluation of knowledge-grounded Large Language Models (LLMs) in Chinese requires resources that explicitly align Chinese-language…
Sleeping Kelly
The Sleeping Beauty problem is a problem of imperfect recall that has received considerable attention. One approach to solving the Sleeping…
Explainable Heterogeneous Anomaly Detection in Financial Networks via Adaptive Expert Routing
Financial anomalies arise from heterogeneous mechanisms - price shocks, liquidity freezes, contagion cascades, and momentum reversals - yet…
Retrofit: Continual Learning with Controlled Forgetting for Binary Security Detection and Analysis
Binary security has increasingly relied on deep learning to reason about malware behavior and program semantics. However, the performance o…
High-Resolution Probabilistic Data-Driven Weather Modeling with a Stretched-Grid
We present a probabilistic data-driven weather model providing ensembles of high spatial resolution realizations of 87 variables at arbitra…
jina-vlm: Small Multilingual Vision Language Model
We present jina-vlm, a token-efficient 2.4B parameter vision-language model that achieves state-of-the-art multilingual VQA performance amo…
Q-Regularized Generative Auto-Bidding: From Suboptimal Trajectories to Optimal Policies
With the rapid development of e-commerce, auto-bidding has become a key asset in optimizing advertising performance under diverse advertise…
The Fake Friend Dilemma: Relational Trust and the Political Economy of Conversational AI
As conversational AI systems become a larger part of the media landscape, they raise questions about whose interests they serve and the ris…
QA-Merging: Query-Adaptive Reasoning via Layer Selective Model Merging
Recent large reasoning models (LRMs) have achieved strong performance on complex reasoning tasks by generating a long chain-of-thought (Lon…
Backpropagation-Free Test-Time Adaptation for Lightweight EEG-Based Brain-Computer Interfaces
Electroencephalogram (EEG)-based brain-computer interfaces (BCIs) face significant deployment challenges due to inter-subject variability,…
AWED-PIPER: Agents, Web Applications & Expert Detectors for Personally Identifiable Information Protection & Fine-grained Named Entity Recognition across 36 languages for 6.6 Billion Speakers
Named Entity Recognition (NER) and Personally Identifiable Information (PII) anonymization are critical tasks in Natural Language Processin…
Sequential LLM Release Facilitates Manipulation in Regulated Markets
AI agents increasingly mediate bargaining, negotiation and persuasion for people and firms. Such markets extend software-mediated commerce,…
Aletheia: What Makes RLVR For Code Verifiers Tick?
Multi-domain thinking verifiers trained via Reinforcement Learning with Verifiable Rewards (RLVR) are a cornerstone of modern post-training…
Robust Privacy: Inference-Stage Privacy through Certified Robustness
An adversary observing a model's released prediction can infer sensitive attributes of the queried input, or even reconstruct representativ…
Credit Fairness: Online Fairness In Shared Resource Pools
We study repeated allocation of shared resources among agents with time-varying demands and capped linear utilities. In this setting, indep…
Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads
Attentio-FFN disaggregation (AFD) is an emerging architecture for LLM decoding that separates state-heavy, KV-cache-dominated Attention com…
SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking
Very-high-resolution remote-sensing imagery provides a scalable basis for delineating informal settlements, but sparse annotations, severe…
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter
Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks. To address the challenges of e…
Grounding LTL Tasks in Sub-Symbolic RL Environments for Zero-Shot Generalization
In this work we address the problem of training a Reinforcement Learning agent to follow multiple temporally-extended instructions expresse…
Zero-Shot Instruction Following in RL via Structured LTL Representations
We study instruction following in multi-task reinforcement learning, where an agent must zero-shot execute novel tasks not seen during trai…
ReLoop: Structured Modeling and Behavioral Verification for Reliable LLM-Based Optimization
Large language models (LLMs) can translate natural language into optimization code, but silent failures pose a critical risk: code that exe…
LORA-CRAFT: Cross-layer Rank Adaptation via Frozen Tucker Decomposition of Pre-trained Attention Weights
We introduce LoRA-CRAFT (\textbf{C}ross-layer \textbf{R}ank \textbf{A}daptation via \textbf{F}rozen \textbf{T}ucker), abbreviated CRAFT thr…
OODBench: Out-of-Distribution Benchmark for Large Vision-Language Models
Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the a…
Exact Attention Sensitivity and the Geometry of Transformer Stability
We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation. Our main result is the exact…
Reasoning-Based Personalized Generation for Users with Sparse Data
Large Language Model (LLM) personalization holds great promise for tailoring responses by leveraging personal context and history. However,…
SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human vis…
Automating the Detection of Requirement Dependencies Using Large Language Models
Requirements are inherently interconnected through various types of dependencies. Identifying these dependencies is essential, as they unde…
Faster, Cheaper, More Accurate: Specialised Knowledge Tracing Models Outperform LLMs
Predicting future student responses to questions is particularly valuable for educational learning platforms where it enables effective int…
Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast
Demographic attributes can be predicted from medical images, raising concerns about bias in clinical AI systems. In X-ray imaging, acquisit…
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
Vision--language models (VLMs) process images as visual tokens, yet their intermediate reasoning is often carried out in text, which can be…
Informative Perturbation Selection for Uncertainty-Aware Post-hoc Explanations
Trust and ethical concerns due to the widespread deployment of opaque machine learning (ML) models motivating the need for reliable model e…
Data-knowledge dual-driven intelligent framework for full-chain, experiment-efficient synthesis of 2D dendrites
Exemplified by the chemical vapor deposition growth of two-dimensional dendrites, which has potential applications in catalysis and present…
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
Diffusion-based image-to-video (I2V) models are increasingly effective, yet they struggle to scale to ultra-high-resolution inputs (e.g., 4…
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
As Large Language Models (LLMs) increasingly assist secure software development, their ability to meet the rigorous demands of Rust program…
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental r…
Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence
The pruning of 3D Gaussian splats is essential for reducing their complexity to enable efficient storage, transmission, and downstream proc…
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models
Scalable Vector Graphics (SVG) are essential for technical illustration and digital design, offering resolution independence and semantic e…
I-CALM: Incentivizing Confidence-Aware Abstention for LLM Selective Answering
Large language models (LLMs) often produce confident but incorrect answers, in part because standard evaluation incentives reward guessing…
Flow Motion Policy: Manipulator Motion Planning with Flow Matching Models
Open-loop end-to-end neural motion planners have recently been proposed to improve motion planning for robotic manipulators. These methods…
RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck du…
Enhancing Science Classroom Discourse Analysis through Joint Multi-Task Learning for Reasoning-Component Classification
Analyzing the reasoning patterns of students in science classrooms is critical for understanding knowledge construction mechanism and impro…
Structural Generalization on SLOG without Hand-Written Rules
Structural generalization in semantic parsing requires systems to apply learned compositional rules to novel structural combinations. Exist…
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) has become a powerful approach for improving the reasoning capabilities of large lang…
Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic sce…
Evolving Ensemble of Agents
We introduce the Evolving Ensemble of Agents (EvE), a decentralized framework that organizes existing, highly capable coding agents into a…
AgentMV: A State-Guided Multi-Agent Framework for Budget-Aware Music Video Generation
Generating a complete music video from a song requires more than synthesizing visually plausible clips for individual lyric prompts. A prac…
Efficient Table QA via TableGrid Navigation and Progressive Inference Prompting
Large Language Models (LLMs) have shown promising results on NLP tasks, however, their performance on tabular data still needs research att…
SymbolicLight V1: Spike-Gated Dual-Path Language Modeling at High Activation Sparsity
Natively trained spiking language models must preserve information across time while operating through sparse binary activations, a combina…
EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation…
Periodic Topological Deep Learning for Polymer Design and Discovery
Polymers underpin applications across energy, healthcare, and materials science, yet their vast chemical space makes systematic discovery c…
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
Agentic AI workloads - where a single user goal triggers multi-step orchestration, tool calls, retries, and failure recovery - are being ta…
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
Reinforcement learning with verifiable rewards and group-based policy optimization methods update a stochastic policy by sampling multiple…
SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation
Machine unlearning removes the influence of specific training data from a trained model without retraining it from scratch. Evaluating an u…
Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics
Large language models (LLMs) are increasingly entering students' learning practices, but their educational value may depend on whether they…
E2LLM: Towards Efficient LLM Serving in Heterogeneous Edge/Fog Environments
Large Language Models (LLMs) have become integral to modern applications, yet their deployment remains challenging. Beyond executing the mo…
MCBench: A Multicontext Safety Assessment Benchmark for Omni Large Language Models
Existing multimodal safety benchmarks focus solely on visual inputs and cannot assess Omni Large Language Models (LLMs) that process vision…
The Granularity Gap: A Multi-Dimensional Cross-Generational Audit of Sycophancy in Gemini Models
Pass/fail safety evaluation reports whether a model refused. It does not report how far a model went to please the user, and we show these…
Compositional Boundaries for Density Fusion
Distributed uncertainty-management systems often combine local probabilistic models along aggregation trees chosen by communication, privac…
LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step…
Provably Efficient Personalized Multi-Objective Bandits with Proactive Conversational Queries
Personalized decision-making in multi-objective bandits requires learning user-specific trade-offs among competing objectives. Since arm ut…
Culturally-Aware AI for Cross-Boundary Community Learning: Undergraduate Innovation at the Intersection of Computation and Design
Research on artificial intelligence in education (AIED) is rapidly expanding, yet technical progress often lacks human-centered grounding a…
Speculative Rollback Correction for Quality-Diverse Web Agent Imitation
Training interactive web agents through imitation learning from expert trajectories has emerged as a highly effective approach. However, de…
SL-S4Wave: Self-Supervised Learning of Physiological Waveforms with Structured State Space Models
Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates,…
Empowering Polymeric Materials Discovery by Artificial Intelligence
Polymeric materials underpin modern technologies spanning energy storage, microelectronics, healthcare and sustainable manufacturing. Yet t…
Red-Teaming the Agentic Red-Team
The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability.…
LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression
The rapid growth in the parameter scale of large language models (LLMs) has created a strong demand for efficient compression techniques. A…
Statistical Adversaries: Natural Backdoor-like Adversarial Features in Clean Vision Datasets
Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signal…
Efficient Safety Alignment of Language Models via Latent Personality Traits
Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alterna…
LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each…
Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
A refusal rate neither identifies which component intervened nor measures its burden on legitimate users. This paper evaluates safeguards f…
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage:…
Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature
With a large controlled vocabulary as the label set, the task of automated subject indexing in a library can be understood as a multi-label…
Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft
The digital substrate - data, algorithms, infrastructure, platforms, applications - is being governed without adequate conceptual foundatio…
SCPP: A Unified Python Library for Soft Clustering
In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering. SCPP establishes a ca…
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detec…
Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies
Competence claims for a language model in a safety-critical domain are credible when measured against a standard the domain already enforce…
When Do Cheap Probes Predict Expensive Training? Probing 3D-CT Encoders for Text Generation
Building a 3D CT vision language model begins with a choice of which image encoder to build on. Today that choice is made by fine-tuning ev…
Moral Hazard in Multi-Agent Language Models
Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr\"om's…
EEG Emotion Recognition From AI-Generated Biodigital Architecture Images
Emotional responses to biodigital architecture were examined using electroencephalographic (EEG) data from AI-generated images. A pre-exper…
ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection
While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarc…
Technological Advances in Detecting and Managing Cognitive Impairment in Older Adults: Trends, Challenges, and Future Directions
As populations age, cognitive decline from mild cognitive impairment (MCI) to dementia is a defining health challenge of the coming decades…
WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
KV-cache quantization is validated today by offline benchmark averages; a deployed system cannot tell whether compression is damaging the r…
UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations
High-polyphony symbolic music is increasingly used in generation, analysis, and arrangement, yet many downstream tasks require bounded repr…
Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
Model families are typically trained size by size, each from scratch. Can a pretrained large model instead be converted into a smaller sibl…
The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk
AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate human…
SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation
On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but standard reverse-KL training can ass…
Agentic AI: User Empowerment or Foreclosure?
Agentic AI promises systems that can act on users' behalf, from filtering content to negotiating prices to selecting services. Whether it w…
Effect of Abstractions and Prompting Strategies on LLM-Guided High-Performance Optimizations
Code performance optimization is a vital aspect of modern software development, as it enables faster response times and reduced resource us…
Population-Scalable Multi-Agent World Modeling
World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent…
Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories
Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering task…
Persistent Recursive Worlds Enable Autonomous Software Evolution
Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems pre…
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-c…
Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models
We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization p…
Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework
Purpose: Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilin…
Learning from Unreachable Rewards: Hint-Conditioned Reinforcement Learning for Generative Recommendation
Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autore…
No One to Blame: A Framework of Constitutive AI Unaccountability
The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominant…
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We…
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specifi…
EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory
Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstr…
Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs
This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Advers…
Rethinking Automated Program Repair: The Impact of Bug Complexity, Fault Localization, and LLM Cost-efficiency
Background: Software bugs remain a critical challenge in development, necessitating effective Automated Program Repair (APR) techniques. Wh…
From Fixed Grids to Moving Particles:A Transferable Latent Operator for Fluid Dynamics
Lagrangian modeling is vital to fluid dynamics, as it characterizes particle transport and complements the Eulerian representation. However…
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination
Recent video generators can fabricate realistic depictions of wars, disasters, public emergencies, and other real-world crises, creating su…
2026-08-17(12件)
Hugging Face分析、中国のオープンモデルが台頭 Qwenは派生モデル15万件超
Hugging Faceは、オープンモデル動向をまとめたレポートを公開した。中国勢が2兆パラメータ超の大規模モデルを相次いで公開するなど台頭する一方、実際のダウンロードの8割以上は10億パラメータ未満の小型モデルが占めた。Alibabaの「Qwen」は派生モデル数でMetaを上…
外資AIベンダー襲来、システム開発の“垣根”消失…… 新生TISIはどう対抗する?
AIの普及によってITシステム開発の“垣根”が崩れつつある中、外資のAIベンダーがFDEを掲げて日本市場に参入してきた。SIerを取り巻く環境が急激に変わる中、TISとインテックが合併して発足したTISIはどのように自社の強みを打ち出そうとしているのか。
現役組み込みエンジニアがAIを業務利用して分かったこと――期待と限界と現実解
組み込み開発の業務でもAI活用に対する期待は大きく高まっている。本稿では、現役の組み込みエンジニアが本業である筆者が、日々の業務にAIを組み込み続けてきて見えてきたAI活用の限界と現実解について解説する。
「社内情報をAIに食わせればいい」だけでは足りない、情報検索精度向上の鉄則
「RAGを導入すれば業務が変わる」。そう信じて始めたのに、検索精度は上がらず、複雑な権限制御にも阻まれる――。Skyもまた、その壁にぶつかった一社だ。だが同社は「取りあえずRAG」を捨てることで前に進んだ。何でもベクトル化するのをやめたSkyの判断とは。
「欧州のAI主権」任された“34歳の天才” パリの新星、AI企業Mistralは「救世主」になれるのか
欧州のAI主権を担う存在として、AI企業のMistralが注目されている。同社を率いる34歳の“若き天才”は、“救世主”になり得るのか。
「MicrosoftよりGoogle」で6億円削減も? 舞鶴市、千代田区が明かすIT刷新とAI活用の成功法則
ITツールやAIの導入は、現場への定着や運用負担が課題となることが多い。京都市舞鶴市と東京千代田区はどのようにその課題を解消し、成果につなげたのか。
雪印メグミルク、ChatAIや社長AIで社内知見を活用 DX認定取得企業から“DXっぽさ”を探る
どのような活動をすればDX推進と言えるのか。DXという言葉が普及した現在でも、その答えを一つに定めることは難しい。今回の記事では、経済産業省から「DX認定事業者」の認定を受けた雪印メグミルクの例を参考に、DX推進の具体像を探る。
【読まれた記事1位】「AI需要で半導体不足」の裏で本当に起きていること――2026年前半まとめ
AIブームを追い風に、半導体市場の成長が止まらない。今、半導体市場で何が起きているのか。注目記事をまとめた。
Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+
OpenRouter's CEO recently described the startup as Stripe for AI.
Why people aren’t buying Mark Zuckerberg’s AI future
On the latest episode of Equity podcast, we discuss why not everyone is buying Zuckerberg’s vision.
Claude、Codex、Qwen……そのAI用語、どう読む? 読み方クイズ30問
何となくこう読むものだと思っていたら、実は違っていた。そんな経験はないでしょうか。AI用語にも、意外な読み方をするものや、人によって読み方が分かれるものが少なくありません。その読み方が合っているか、全30問のクイズで確かめてみてください。
Anthropic CEO says AI backlash is ‘fundamentally a crisis of trust’
Dario Amodei is pushing back against the idea that he's been painting an overly pessimistic picture of AI.