週次AIニュース 2026-W25
対象期間: 2026-06-15 〜 2026-06-21(1894 件)
トピックの推移
トピック別件数
- 研究/論文 786件
- LLM/生成AI 696件
- エージェント 425件
- 画像/動画生成 258件
- ロボティクス 134件
- ビジネス/資金調達 110件
- その他 42件
- ハードウェア/半導体 40件
- 規制/政策 7件
今週のハイライト(上位 10 件)
New usage analytics and updated spend controls for enterprises
OpenAI introduces new spend controls and usage analytics for ChatGPT Enterprise, helping organizations manage costs and scale AI with confi…
Improving health intelligence in ChatGPT
Learn how GPT-5.5 Instant improves ChatGPT’s health and wellness responses with stronger reasoning, better context, clearer communication,…
Using AI to help physicians diagnose rare genetic diseases affecting children
Researchers used an OpenAI reasoning model to help diagnose rare diseases, identifying 18 new diagnoses in previously unsolved cases.
A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry
OpenAI and Molecule.one show how a near-autonomous AI chemist using GPT-5.4 improved a key drug-making reaction, advancing medicinal chemis…
Introducing the OpenAI Partner Network
OpenAI launches the Partner Network, investing $150M to help global partners accelerate enterprise AI adoption, deployment, and transformat…
Unlocking UK house-building with AI-accelerated planning
UK government partners with Google DeepMind to build a new AI-powered prototype aimed at faster housing decisions.
Signal’s Meredith Whittaker wants you to remember that AI chatbots ‘are not your friends’
"These are not your friends. These are not conscious beings. These are not sentient interlocutors.”
Nobel laureate John Jumper is leaving DeepMind for rival Anthropic
Jumper isn't the only big name leaving Google DeepMind.
Is the US government’s Anthropic ban accidentally helping the brand?
Just as last week was ending, the US government forced Anthropic to pull its two newest models, Fable 5 and Mythos 5, citing national secur…
The US banned Anthropic’s Fable 5 release, but the numbers don’t seem to care
Just as last week was ending, the US government forced Anthropic to pull its two newest models, Fable 5 and Mythos 5, citing national secur…
全件(日付別)
2026-06-21(3件)
Signal’s Meredith Whittaker wants you to remember that AI chatbots ‘are not your friends’
"These are not your friends. These are not conscious beings. These are not sentient interlocutors.”
In the Weights is your new AI-centric vanity search
So ... what's your In the Weights score?
Nobel laureate John Jumper is leaving DeepMind for rival Anthropic
Jumper isn't the only big name leaving Google DeepMind.
2026-06-20(4件)
Encryption, spyware, and now Mythos: History shows why cyber export control doesn’t work
For the last 30 years, stopping the flow of cybersecurity-related software has proven to be ineffective. It's unclear why it would work now…
Is the US government’s Anthropic ban accidentally helping the brand?
Just as last week was ending, the US government forced Anthropic to pull its two newest models, Fable 5 and Mythos 5, citing national secur…
The US banned Anthropic’s Fable 5 release, but the numbers don’t seem to care
Just as last week was ending, the US government forced Anthropic to pull its two newest models, Fable 5 and Mythos 5, citing national secur…
Billionaire Ambani wants AI in every call, app, and home
Reliance is weaving AI into telecom services used by more than 500 million people.
2026-06-19(317件)
The CEO of Allbirds’ new AI biz has a plan, but no team
Call it a startup with a sole founder and a very large seed round, but what's next is less clear.
理研、AI for Science向けスパコンの名前を「理究」(りきゅう)に決定 由来は?
理化学研究所は、AIを活用した科学研究「AI for Science」向けのスーパーコンピュータの名前を「理究」(りきゅう)に決定したと発表した。
The US says ASML’s top chip tool may be in China, but how?
There's a commercial logic that cuts against the idea that ASML would risk its export license to arm a Chinese customer.
大阪メトロは「月1000件の社内問い合わせを効率化」にAIをどう使った?
PKSHA Technologyは、Osaka Metroへの「PKSHA AIヘルプデスク」の導入部署を人事や調達に拡大した。全従業員約5000人の社内問い合わせを一元化し、情報格差の解消や業務効率化、ナレッジの共有資産化を目指す。
GMO傘下、Unitreeの国内正規代理店に 人型ロボの導入から保守まで一気通貫で支援
GMOインターネットグループ傘下で、ロボティクス事業などを手掛けるGMO AI&ロボティクス商事は、ロボット開発企業の中国Unitree Roboticsと日本国内正規代理店契約を締結したと発表した。
画面操作を“録画”→AIが作業代行 Codexに新機能「Record & Replay」
米OpenAIは6月18日、「Codex」の新機能「Record & Replay」を公開した。利用者がMac上で作業を一度実演すると、Codexがその操作を再利用できる作業手順に変換・記憶する。
Gartnerが警鐘 プライバシー法執行が本格化、CISOは何を見直すべきか?
Gartnerは、2025年に米国の州当局が科したプライバシー法違反の罰金総額が34億2500万ドル(約5380億円)に達したと発表。過去5年間の合計を上回り、執行強化を背景に2028年まで加速する見通しを示した。
Deontic Policies for Runtime Governance of Agentic AI Systems
Autonomous agentic AI systems driven by Large Language Models (LLMs) introduce a new class of security, privacy, and compliance challenges:…
Measuring Curriculum Alignment across Topical Coverage, Competency, and Cognitive Depth: A Longitudinal Framework Applied to CS2013 and CS2023
Undergraduate computer science is governed by international curricular guidelines revised about once a decade, yet programs lack a reliable…
Diffusion Language Models: An Experimental Analysis
Large Language Models (LLMs) have revolutionized language modeling through autoregressive generation, enabling strong performance across a…
Hidden Anchors in Multi-Agent LLM Deliberation
Multi-agent LLM deliberation, where agents exchange and revise answers over several rounds, is increasingly used to improve reasoning and a…
DeXposure-Claw: An Agentic System for DeFi Risk Supervision
Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they…
LLM Doesn't Know What It Doesn't Know: Detecting Epistemic Blind Spots via Cross-Model Attribution Divergence on Clinical Tabular Data
Large language models (LLMs) are increasingly applied to structured clinical data, yet whether they can recognize the limits of their own k…
REVEAL++: Differentiable Phenotypic Grouping for Vision-Language Retinal Modeling of Alzheimer's Disease Risk
The retina offers a noninvasive window into neurodegenerative disease, capturing subtle structural patterns associated with a risk of futur…
Emergent Alignment
Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics? And can they self-correct? We endow an LL…
ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence
Convolutional networks, recurrent networks, and transformers each encode different inductive biases -- locality, sequential memory, and con…
Uncertainty Decomposition for Clarification Seeking in LLM Agents
Recent position papers argue that the classical aleatoric/epistemic uncertainty framework is insufficient for interactive large language mo…
Analyzing the Narration Gap in LLM-Solver Loops
Formal tools such as SAT and SMT solvers are increasingly embedded in language model reasoning pipelines when a safety or security critical…
Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why
Patient contexts span hundreds of heterogeneous documents and thousands of structured data points, yet the document-level metadata that AI…
Which Pairs to Compare for LLM Post-Training?
Preference-based post-training has become a central paradigm for aligning language models. A common data-collection strategy is to generate…
Toten: Knowledge-Based Ontological Tokenization Of Physical Quantities And Technical Notation In Brazilian Portuguese
Byte-Pair Encoding tokenization is statistically efficient for vocabulary compression, but semantically blind to structured technical entit…
AI4SE and SE4AI Exploration: A Decade Looking Back and Forward
The March 2020 INCOSE INSIGHT special issue on AI and Systems Engineering (SE) became the most downloaded issue in the publication's histor…
BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation
Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represent…
Denoising Implicit Feedback for Cold-start Recommendation
Implicit feedback is widely used in recommender systems due to its accessibility and generality, yet it usually presents noisy samples (e.g…
Exit-and-Join Dynamics for Decentralized Coalition Formation
This paper studies coalition formation as a decentralized dynamical process driven by unilateral exit-and-join decisions. Agents evaluate l…
Beyond Static Leaderboards: Predictive Validity for the Evaluation of LLM Agents
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This p…
GLARE: A Natural Language Interface for Querying Global Explanations
While global explanations are crucial for understanding vision models across datasets, classes, and decision contexts, their complex and mo…
Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks
Neural Combinatorial Optimization (NCO) achieves strong performance, yet its black-box nature remains a key roadblock to deployment and sci…
A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition
Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation too…
Benchmarking Agentic Review Systems
A new class of agentic review systems are emerging as a remedy to the pressure placed on peer review systems by AI-assisted research, but i…
Grounded Inference: Principles for Deterministically Encapsulated Generative Models
The incorporation of generative models into traditional computational systems presents both enormous opportunity and tremendous peril. Alth…
Optimal Scheduling in a Question-Answering Forum of Knowledge Workers
As individuals turn to the Internet to find answers to questions they may have, several Question Answering (QA) forums have evolved, where…
Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced Large Language Model (LLM) reasoning; however, it faces a…
AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA
Financial chart question answering in regulated settings demands more than accuracy: practitioners must know which answers to trust before…
ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?
Large language models are increasingly deployed as autonomous agents for multi-step tasks in executable environments, yet their ability to…
CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models
We present CombEval, a dynamic benchmark for evaluating combinatorial counting in large language models. CombEval represents each problem a…
Think Again or Think Longer? Selective Verification for Budget-Aware Reasoning
Test-time reasoning is increasingly used as a serving-time control knob, but extra reasoning is not uniformly valuable: it can repair faile…
Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery
Autonomous Large Language Model (LLM) agents are increasingly deployed in electronic discovery (e-discovery), where compounding errors acro…
TelcoAgent: A Scalable 5G Multi-KPM Forecasting With 3GPP-Grounded Explainability
Key Performance Measurement (KPM) forecasting is essential for proactive network management of 5G and next-generation telecom networks. How…
A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models
Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable an…
MetaResearcher: Scaling Deep Research via Self-Reflective Reinforcement Learning in Adversarial Virtual Environments
Deep research agents have demonstrated remarkable capabilities in autonomous information gathering and synthesis, yet their training remain…
Multi-Agent Transactive Memory
The decentralized deployment of LLM agents with diverse capabilities across diverse tasks motivates infrastructure for knowledge sharing ac…
eCNNTO: A Highly Generalizable ConvNet for Accelerating Topology Optimization
This work proposes an element-based Convolutional Neural Network (CNN) to accelerate density-based Topology Optimization (TO), termed eCNNT…
The Tao of Agency: Autotelic AI, Embedded Agency and Dissolution of the Self
Most artificial intelligence systems are built on the assumption that goals are exogenous and specified by the designer. Exploring what hap…
PhysDrift: Bridging the Embodiment Gap in Humanoid Co-Speech Motion Generation
Humanoid robots require co-speech motions that are not only expressive and speech-aligned, but also physically executable under embodiment…
Advancing DialNav through Automatic Embodied Dialog Augmentation
For embodied agents capable of physical interaction, the capability to create and understand dialog is crucial to ensure both safety and ef…
ENPIRE: Agentic Robot Policy Self-Improvement in the Real World
Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a c…
Reward as An Agent for Embodied World Models
While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training di…
Autonomous Event-Driven Multi-Agent Orchestration for Enterprise AI at Scale
Enterprise AI aims to move toward continuous event monitoring, detection, and action across specialist agents, yet existing multi-agent sys…
Process-Verified Reinforcement Learning for Theorem Proving via Lean
While reinforcement learning from verifiable rewards (RLVR) typically has relied on a single binary verification signal, symbolic proof ass…
Residual-Space Evolutionary Optimization via Flow-based Generative Models
Data editing with generative methods typically requires differentiable objectives and gradient-based search. However, these assumptions bre…
Multi-Head Attention-Based Feature Extractor Integration with Soft Actor-Critic for Porosity Prediction and Process Parameter Optimization in Additive Manufacturing
Additive manufacturing process optimization requires precise parameter control to minimize defects such as porosity. Traditional reinforcem…
ScaffoldAgent: Utility-Guided Dynamic Outline Optimization for Open-Ended Deep Research
Open-ended deep research (OEDR) requires systems to acquire knowledge through multi-round retrieval and generate coherent long-form reports…
Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring
LLMs can personalize education, although current static-prompt tutoring systems struggle to adapt to diverse academic disciplines. We devel…
RACL: Reasoning-Agent Control Layers for Continuous Metaheuristic Learning
This paper introduces RACL, a Reasoning-Agent Control Layer for metaheuristics. RACL places a reasoning agent above an existing optimizer.…
BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling
Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions…
Modularity-Free Conflict-Averse Training for Generalized PINNs
Physics-informed neural networks (PINNs) have become a powerful framework for solving PDEs by embedding physical laws into differentiable o…
Implicit Semantic-Aware Communication Based on Hypergraph Reasoning
Semantic-aware communication has emerged as a transformative paradigm for next-generation communication systems, shifting the fundamental g…
Apparent Psychological Profiles of Large Language Models are Largely a Measurement Artifact
Psychological instruments designed for humans are increasingly used to assign large language models (LLMs) stable psychological profiles th…
Beyond Accuracy: Measuring Logical Compliance of Predictive Models
Machine learning models are predominantly evaluated through predictive performance metrics such as ranking quality, prediction error, or cl…
Augmenting Game AI with Deep Reinforcement Learning
Immersion in video games depends not only on graphics, audio, and game mechanics, but also on the quality of in-game characters. Producing…
QMFOL: Benchmarking Large Language Model Reasoning via Quantifiable Monadic First-Order Logic Test Case Generation
Large Language Models (LLMs) have made significant progress in reasoning, particularly in deductive reasoning, which is crucial for high-st…
Thermodynamic Measure of Intelligence
Can intelligence be measured? We propose that intelligence can be defined as the lawful amplification of rare but valid futures: a system i…
A Multi-Agent system for Multi-Objective constrained optimization
Many decision-making problems in computing and networking systems can be naturally formulated as cost-minimization problems under performan…
Navigating Unreliable Parametric and Contextual Knowledge: Explicit Knowledge Conflict Resolution for LLM Inference
Large language models (LLMs) have achieved strong performance across a wide range of language-based tasks by leveraging both extensive para…
Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
Student-generated drawings are widely used in science education to assess learners' conceptual understanding in modeling-based tasks aligne…
Lagrange: An Open-Vocabulary, Energy-Based Sparse Framework for Generalized End-to-End Driving
Scaling end-to-end autonomous driving to complex, open-world environments requires perceptual models that generalize to anomalous scenarios…
Leveraging systems' non-linearity to tackle the scarcity of data in the design of Intelligent Fault Diagnosis Systems
Deep Transfer Learning (DTL) allows for the efficient building of Intelligent Fault Diagnosis Systems (IFDS). On the other hand, DTL method…
SoftSkill: Behavioral Compression for Contextual Adaptation
Agent skills are commonly deployed as natural-language Markdown files that encode answer policies, evidence-use habits, and task procedures…
Automating SKILL.md Generation for Computer-Using Agents via Interaction Trajectory Mining
Explicit skill libraries make computer-using agents easier to inspect, but it remains unclear whether such libraries can be mined from inte…
Rethinking Shrinkage Bias in LLM FP4 Pretraining: Geometric Origin, Systemic Impact, and UFP4 Recipe
FP4 training promises substantial reductions in memory and computation cost for LLM pretraining, yet current FP4 hardware paths and recipes…
Interpretable Sperm Morphology Classification via Attention-Guided Deep Learning
Male infertility is a major cause of couple infertility, often linked to abnormal sperm morphology. While deep learning models offer automa…
Context-Aware Hierarchical Bayesian Modeling of IVF Laboratory Environmental Conditions
IVF pregnancy rates are routinely modeled using patient-level variables, while high-resolution laboratory environmental data remain underut…
What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?
Prior work has shown that in-context demonstrations can jailbreak language models, but it remains unclear how models interpret different ty…
Multi-LCB: Extending LiveCodeBench to Multiple Programming Languages
LiveCodeBench (LCB) has recently become a widely adopted benchmark for evaluating large language models (LLMs) on code-generation tasks. By…
FlowEdit: Associative Memory for Lifelong Pronunciation Adaptation in Flow-Matching TTS
Flow-matching text-to-speech systems achieve remarkable zero-shot quality but remain static after deployment: pronunciation errors on out-o…
DeepSWIP: Quotient-WMC Counterfactuals for Neural Probabilistic Logic Programs
Neurosymbolic systems such as DeepProbLog combine neural perception with probabilistic logic, but standard inference is associational. Coun…
LedgerAgent: Structured State for Policy-Adherent Tool-Calling Agents
Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling tools and obeying doma…
How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech
Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic o…
Toward Calibrated Mixture-of-Experts Under Distribution Shift
Calibration aligns a model's predictive uncertainty with the frequencies of its empirical outcomes and is important for understanding and t…
Human Universal Grasping
Humans can grasp objects effortlessly, whereas multi-fingered robots are far from this level of generality. We argue that the most natural…
Human-AI Agent Interaction in a Business Context
As AI agents are increasingly integrated into core business processes, understanding and designing effective interaction patterns between h…
Exposing the Unsaid: Visualizing Hidden LLM Bias through Stochastic Path Aggregation
Large Language Models (LLMs) exhibit representational and syntactic biases that are difficult to evaluate due to the stochastic nature of t…
Ensembles of Large Language Models for Identifying EQ-5D Studies in PubMed Based on Their Abstracts
The rapid increase in scientific publications leads to the fact that manual study screening in systematic literature reviews (SLRs) is incr…
Disentangling Linguistic Relatedness from Task Alignment in Cross-Lingual Transfer
We study cross-lingual transfer by fine-tuning seven large language models (4B--671B parameters) on Arabic and evaluating zero-shot reading…
How LLMs Fail and Generalize in RTL Coding for Hardware Design?
Translating sequential programming priors into the parallel temporal logic of hardware design remains a crucial bottleneck for large langua…
DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1…
Where to Place the Query? Unveiling and Mitigating Positional Bias in In-Context Learning for Diffusion LLMs via Decoding Dynamics
While In-Context Learning (ICL) is extensively studied in Autoregressive (AR) LLMs, its mechanism within Diffusion Large Language Models (d…
Detecting Hallucinations for Large Language Model-based Knowledge Graph Reasoning
Knowledge graph (KG) reasoning infers new knowledge from existing facts and is widely applied in question answering, recommendation, and de…
Sign-Language Datasets at Scale: A Comprehensive Survey on Resources, Benchmarks, and Annotation Standards
Sign languages are expressive visual languages used by Deaf and Hard-of-Hearing (DHH) communities. Despite substantial progress in sign-lan…
Trustworthy Multi-Agent Systems: Mitigating Semantic Drift with the Argent Signaling Protocol
When multi-agent LLM systems produce bad answers, not all failures are equal: some answers are grounded in the right material but incomplet…
Physical Atari: A Robust and Accessible Platform for Real-time Reinforcement Learning on Robots
We built a robot called the Robotroller that actuates an Atari CX40+ controller and a device called the Atari Devbox that renders the game…
Computational Identifiability
Identification conditions describe the computability of a target query or parameter of interest as a function of the type and amount of inf…
Information Lattice Learning as Probabilistic Graphical Model Structure Learning
Information lattice learning (ILL) learns interpretable rules of a signal by alternately projecting the signal onto a partition lattice tha…
Zero-Inflated Gaussian Distributions Enable Parameter-Space Sparsity in Estimation-of-Distribution Algorithms
Estimation-of-distribution algorithms (EDAs) are a powerful class of evolutionary methods for black-box optimization, especially when littl…
Human-like autonomy emerges from self-play and a pinch of human data
Self-play reinforcement learning has recently emerged as a way to train driving policies without any human data. It uses cheap, large-scale…
ProMUSE: Progressive Multi-modal Uncertainty-guided Staged Evidential Alzheimer Disease Classification
Alzheimer's disease (AD) is a fatal disorder that destroys memory and cognitive skills in the elderly population. Most treatments for AD ar…
cAPM: Continual AI-Assisted Pace-Mapping with Active Learning
Ventricular tachycardia is a life-threatening rhythm disorder and a major cause of sudden cardiac death. Pace-mapping is a clinical procedu…
Protein Representation Learning with Secondary-Structure and Energy-Filtered Hydrogen-Bond Graphs
Graph-based representations are widely used in protein modeling, yet many existing approaches rely primarily on sequence adjacency or geome…
Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees
Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost. U…
Emyx: Fast and efficient all-atom protein generation
Computational enzyme design requires generating proteins that scaffold catalytic residues and ligands, a task that demands both geometric a…
How Linear Is a Transformer Feed-Forward Block? Per-Block Linear Recoverability Is Learned, Not Architectural
Transformer feed-forward networks (FFNs) are often treated as nonlinear stores of computation, yet how nonlinear a trained FFN block actual…
Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech
Code-switch (CS) Automatic Speech Recognition (ASR) remains challenging due to limited availability of high quality CS text-speech pairs fo…
DynAMO:Dynamic Asset Management Orchestration via Topological Multi-Agent Scheduling
While LLM-powered agents offer end-to-end automation for industrial asset lifecycles, real-world Industry 4.0 deployment is hindered by lat…
Bistable by Construction: Wall-Clock-Calibrated State Monitors Have No Moment-Detection Regime at Agent Cadence
Runtime monitors for autonomous agents commonly threshold an accumulated internal state - a behavioural baseline, a drift statistic, or, in…
Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement
Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, mea…
Execution-bound advisory automation for agentic AI: a reproducible AIBOM-driven CSAF-VEX framework
A protocol driven framework is presented that binds SBOM and AIBOM artefacts to deterministic environment capture and structured runtime te…
VERITAS: Verifier-Guided Proof Search for Zero-Shot Formal Theorem Proving
LLM-based formal provers often collapse rich verifier signals (syntax errors, type mismatches, partial goal progress) into a binary pass/fa…
JustDiag!: A Diagnostic Justification Engine for Accountable Root Cause Analysis
Large language models can produce fluent root cause analyses, but fluent final answers alone are insufficient evidence for accountability i…
Playful Agentic Robot Learning
Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts,…
Scaling Generative Foundation Models for Chest Radiography with Rectified Flow Transformers
We introduce the first generative foundation model for chest radiograph synthesis trained from scratch at the billion-parameter scale. Exis…
Secure Coding Drift in LLM-Assisted Post-Quantum Cryptography Development: A Gamified Fix
The transition to Post Quantum Cryptography (PQC) introduces considerable implementation complexity, requiring strict adherence to constant…
Can In-Context Learning Support Intrinsic Curiosity?
Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models h…
Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks
Concept Bottleneck Models (CBMs) enhance interpretability by projecting learned features into a human-understandable concept space. Recent…
Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices
Fine-tuning of Large Language Models (LLMs) using Low-Rank Adaptation (LoRA) on an end-user's data offers personalized experiences while ke…
A Tool for the Synthesis of Adaptive Probabilistic Processors Based on the Ising Model
This work presents a tool for the synthesis and simulation of probabilistic architectures for solving combinatorial optimization problems b…
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models
Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely…
Review of Machine Learning Models for Solar Energetic Particle Prediction
Solar energetic particle (SEP) events have attracted increasing attention due to their significant radiation hazards for aviation, spacecra…
GDGU: A Gradient Difference-based Graph Unlearning Method for Cyberattack Localization in Electric Vehicle Charging Networks
Electric vehicle charging stations (EVCSs) can expose distribution feeders to cyberattacks. While machine learning methods, including graph…
Exploring Feature Extraction Technique Parameters for Acoustic Gunshot Classification
Acoustic gunshot detection is a problem with applications across civilian public safety, military operations, and wildlife conservation, ye…
FlowFake: Liquid Networks for Audio Deepfake Detection
Audio deepfakes generated by neural text-to-speech and voice-cloning systems threaten speaker verification and public discourse at scale. T…
A BART-based approach with hierarchical strategy for Vietnamese abstractive multi-document summarization
In this technical report, we focus on solving the challenge of Vietnamese multi-document abstractive summarization, introduced in the Inter…
IHBench: Evaluating Post-Interruption Recovery in Voice Agents with Structured Workflows
Voice agents deployed in structured workflows (customer service, healthcare scheduling, account management) must handle frequent user inter…
PrefSQA: Pairwise Preference Prediction for Speech Quality Assessment and the Critical Role of High Quality Datasets
Mean opinion scores (MOS) are widely used for speech quality assessment, yet scalar labels are sensitive to rater variability and listening…
FAPO: Fully Autonomous Prompt Optimization of Multi-Step LLM Pipelines
Multi-step LLM pipelines fail through interactions among retrieval, reasoning, and formatting steps, so prompt-only optimization can miss b…
Latent Confounded Causal Discovery via Lie Bracket Geometry
Recent work on Kan-Do-Calculus (KDC) has established that the boundary between passive observation and active intervention in causal infere…
StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns
We introduce StaminaBench, a benchmark that measures the stamina of coding agents: how many consecutive interaction turns (change requests)…
Before the Pull Request: Mining Multi-Agent Coordination
Autonomous coding agents now open millions of pull requests, yet large-scale studies find their PRs are produced faster but accepted less o…
VCG: A Multimodal Retrieval Framework for E-Commerce Video Feeds under Extreme Cold-Start Conditions
The digital commerce landscape is shifting from static, search-driven catalogs to dynamic, immersive video feeds. This transition introduce…
RIVET: Robust Idempotent Voice Attribute Editing
Voice attribute editing models modify characteristics such as age and gender while preserving speaker identity. In large-scale speech datas…
Formal Verification of Learned Multi-Agent Communication Policies via Decision Tree Distillation
Multi-agent reinforcement learning (MARL) enables agents to develop coordination strategies through emergent communication, but neural poli…
CTS-MoE: Implicit Terrain Adaptation via Mixture-of-Experts for Perceptive Locomotion
Perceptive legged locomotion over discontinuous terrain (e.g., stairs, gaps, and obstacles) requires adaptive behavior, as a single conserv…
Token Factory: Efficiently Integrating Diverse Signals into Large Recommendation Models
Large Recommendation Models (LRMs) have demonstrated promising capabilities in industry-scale recommendation tasks. However, holistically i…
Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation
Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficul…
Before the Labels: How Dataset Construction Shapes Suicidality Detection in Clinical Text
Clinical NLP increasingly relies on electronic health record (EHR) data to detect suicidal behaviors, treating clinical documentation as mo…
Creating Multilingual Mental Health Dialogue Datasets: Limits of Persona-Based Localization via Nationality and Language
AI and large language models (LLMs) have emerged as promising tools to address global mental health challenges. Despite the global nature o…
TeleMorpher: Toward Robust Simultaneous Motion-Location Editing
Diffusion models have achieved remarkable success in image and video generation and editing. While recent studies have extended these effor…
LOKI: Memory-Free Null-Space Constrained Lifelong Knowledge Editing
Lifelong knowledge editing aims to efficiently and sequentially update language models over time, as new knowledge becomes available or whe…
Efficiently Representing Algorithms With Chain-of-Thought Transformers
The increasing popularity of \emph{reasoning} models -- language models that output a series of reasoning or thought tokens before producin…
FineREX: Fine-Tuned NER-RE for Human Smuggling Knowledge Graphs
Court proceedings contain valuable evidence about human smuggling networks, but this information is often buried within unstructured, jargo…
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing
Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive…
OnDeFog: Online Decision Transformer under Frame Dropping
In challenging real-world reinforcement learning applications, communication delays or sensor failures often cause frame dropping, in which…
Library-Aware Doubles and Iterative Repair for Large Language Model-Generated Unit Tests in OpenSIL Firmware
Validating changes in low-level C firmware is expensive because unit tests (UTs) are fragile under strict build constraints, where missing…
NRITYAM: Language Models Meet Art and Heritage of Dance
Language models have become essential tools in shaping modern workflows. However, their global effectiveness hinges on a nuanced understand…
Bidirectional Tutoring for Developmental Motor Learning in Robots: Co-Developed Interaction Dynamics Support Stable Learning
Infants are well known to develop their motor skills through dense interaction with caregivers. Although such social interaction is crucial…
VOiLA: Vectorized Online Planning with Learned Diffusion Model for POMDP Agents
Planning under uncertainty is an essential capability for autonomous robots. The Partially Observable Markov Decision Process (POMDP) provi…
QueryGaussian: Scalable and Training-Free Open-Vocabulary 3D Instance Retrieval
Efficiently retrieving specific 3D instances from large-scale scenes via natural language prompts remains a formidable challenge in multime…
Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings
Aligning language models with human preferences often requires optimising multiple behavioural objectives. A practical approach is to apply…
Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models
Reinforcement learning (RL) is a central approach for improving reasoning capabilities in large language models (LLMs), where training effi…
Temporal Self-Imitation Learning
Long-horizon robot manipulation policies trained with reward shaping can still exploit dense rewards through inefficient interaction, while…
SafeSpec: Fast and Safe LLM via Dynamic Reflective Sampling
Speculative inference accelerates large language model (LLM) decoding but provides no inherent safety guarantees. Existing safety defenses…
Data Standards for Humanoid Robotics: The Missing Infrastructure for Physical AI
The scalability of humanoid robots will depend not only on models and hardware, but also on whether physical experience can accumulate acro…
Towards Engineering Scaling Laws with Pretraining Data Composition
Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established…
Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR
The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired ar…
Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models
The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired ar…
Agentic Electronic Design Automation: A Handoff Perspective
Electronic design automation (EDA) is inherently multi-stage and handoff-heavy. Design artifacts, flow scripts, and engineering decisions c…
Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation
Dysarthric speech recognition is crucial for facilitating effective communication among individuals with dysarthria. However, accurately re…
Policy-aware Vector Search: A Vision for Fine Grained Access Control in Vector Databases
Vector databases are increasingly used in security sensitive contexts with Retrieval Augmented Generation and organizational AI pipelines;…
ParaScale: Scale-Calibrated Camera-Motion Transfer via a Gauge-Invariant Parallax Number
Transferring the camera motion of a reference video to a freshly generated one lets creators reuse cinematic moves. Yet reference and targe…
Uncertainty-Aware Reward Modeling for Stable RLHF
Reinforcement learning from human feedback (RLHF) aligns large language models by training reward models on preference data and optimizing…
CREDENCE: Claim Reduction for Decomposition & Enhanced Credibility -- Semantic Metrics and Convergence Analysis
Decomposing compound sentences into atomic, verifiable claims is a prerequisite for reliable automated fact-checking. Prior work has relied…
CSWinUNETR: Segmentation of Thin Anatomical Structures in Medical Images
Accurate segmentation of thin, tortuous anatomical structures, such as retinal vessels, cerebral vasculature, and facial wrinkles, remains…
When, Where, and How: Adaptive Binning for Tabular Self-Supervised Learning
Medical tabular data are ubiquitous in clinical research, but deep learning for tables remains underexplored because reliable labels often…
Neural Additive and Basis Models with Feature Selection and Interactions
Deep neural networks (DNNs) exhibit attractive performance in various fields but often suffer from low interpretability. The neural additiv…
Large Language Models Do Not Always Need Readable Language
Large language models (LLMs) are commonly prompted and interfaced with human-readable natural language, even when the intended reader is an…
PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement
Computed Tomography (CT) is essential for diagnosing pediatric craniofacial abnormalities, yet poses radiation risks to developing anatomie…
FFinRED: An Expert-Guided Benchmark Generation and Evaluation Framework for Financial LLM Red-Teaming
Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks. Financial LLMs face regulatory compliance…
SL-S4Wave: Self-Supervised Learning of Physiological Waveforms with Structured State Space Models
Modeling long-sequence medical time series data, such as electrocardiograms (ECG), poses significant challenges due to high sampling rates,…
Measuring Biological Capabilities and Risks of AI Agents
This paper addresses a rapidly emerging policy challenge: how to generate and interpret credible evidence about the biological capabilities…
Co-policy: Responsive Human-Robot Co-Creation for Musical Performances
Art has long stood as a pivotal expression of human creativity. Embodied artificial intelligence offers a route for generative models to pa…
Spatial-Aware Reduction Framework: Towards Efficient and Faithful Visual State Space Models
Mamba demonstrates strong efficiency in modeling long visual sequences. However, when token reduction is applied to structurally enhanced M…
Speeding up the annotation process in semantic segmentation industrial applications
Current machine learning models commonly require large and well-annotated datasets. However, the annotation process often becomes a bottlen…
Triangular Consistency as a Universal Constraint for Learning Optical Flow
We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision…
SIMBA: ABidirectional Retrieval Forward Simulation Framework for Modeling FY-4A GIIRS Hyperspectral Infrared Radiances Toward NWP Applications
Hyperspectral infrared observations are an important data source for numerical weather prediction (NWP) because they provide rich informati…
Confidence Calibration for Multimodal LLMs: An Empirical Study through Medical VQA
Multimodal Large Language Models (MLLMs) show great potential in medical tasks, but their elicited confidence often misaligns with actual a…
ROSE: Benchmarking the Perception-to-Action Gap in Multimodal Models
Multimodal large language models (MLLMs) are increasingly expected to act on visual information, yet the same scene may require different a…
The Algorithmic-Human Manager: AI, Apps, and Workers in the Indian Gig Economy
This paper examines the impact of artificial intelligence and digital technologies on the blue-collar gig economy in India, focusing on alg…
Beyond Static Endpoints: Tool Programs as an Interface for Flexible Agentic Web Services
In the agentic web era, LLM-based agents increasingly invoke web services as tools, yet most interfaces remain \emph{static endpoints} that…
Tri-Info: Generalizable, Interpretable Failure Prediction for VLA Models via Information Theory
Vision-Language-Action (VLA) models are increasingly deployed across diverse tasks, yet they remain black boxes whose physical interactions…
Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning
This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by…
StreamKL: Fast and Memory-Efficient KL Divergence for Boosting Attention Distillation
Attention distillation, which trains one attention distribution to match another by minimizing their Kullback-Leibler (KL) divergence, is w…
Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution
Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments…
When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However,…
A Neuromorphic Reinforcement Learning Framework for Efficient Pathfinding in Robotic Mobile Fulfillment Systems
Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robotic Mobile Fulfillment Syst…
AI Economist Agent: An Agentic Framework for Model-Grounded Economic Analysis with RAG, Knowledge Graphs, and Large Language Models
We propose a model-grounded RAG-based AI economist with an agentic framework for economic scenario analysis using large language models (LL…
See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View
UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery…
Evaluation of EEG Foundation Models for Event-Based Burst-Suppression Detection in ICU
Burst suppression (BS) is a clinically relevant electroencephalographic (EEG) pattern used to monitor sedation depth and brain activity in…
Variable-Length Tokenization via Learnable Global Merging for Diffusion Transformers
Latent Diffusion Models (LDMs) have become dominant in visual synthesis, but their quality-compute trade-off is largely constrained by the…
The Hidden Evolution of Disguised Visual Context inside the VLM
Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and inte…
IHUBERT: Vector-Based Semantic Deduplication and Domain-Balanced Pretraining for Persian Resources
Persian pretrained language models (PLMs) are still limited by the scarcity of large-scale, high-quality pretraining corpora and by insuffi…
MakeupMirror: Improving Facial Attribute Preservation in Diffusion Models for Makeup Transfer
Makeup transfer models enable fun augmented reality (AR) experiences as well as virtual try-on (VTO) for online makeup shopping. While rece…
Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow
Audio editing aims to modify specific content in an existing audio clip according to a natural language instruction while preserving the re…
Sensorimotor World Models: Perception for Action via Inverse Dynamics
Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for…
Dual-Agent Framework for Cross-Model Verified Translation of Natural-Language Protocols into Robotic Laboratory Platform
Biological experiment protocols are written in natural language, whereas automation systems rely on predefined control commands, creating a…
Frequency-Aware Flow Matching for Continuous and Consistent Robotic Action Generation
Flow matching has emerged as a standard paradigm for robotic manipulation owing to its strong expressive power for modelling complex, multi…
Hybrid ANN-SNN Pipeline with Local Plasticity
This work proposes a hybrid ANN-SNN pipeline that effectively leverages the rich embeddings of pretrained artificial neural networks (ANNs)…
From Texts to Scores: Tracing the Emergence of Essay Quality Representations in Large Language Models
Recent advances in Large Language Models (LLMs) have substantially transformed Automated Essay Scoring (AES), yet the internal mechanisms u…
MedRLM: Recursive Multimodal Health Intelligence for Long-Context Clinical Reasoning, Sensor-Guided Screening, Evidence-Grounded Decision Support, and Community-to-Tertiary Referral Optimization
Real-world clinical decision support requires reasoning over heterogeneous and longitudinal patient information rather than answering isola…
Evaluating and Enhancing Negation Comprehension in Remote Sensing MLLMs
Multimodal Large Language Models (MLLMs) have demonstrated remarkable success in various Remote Sensing (RS) tasks. However, their ability…
HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-trainin
Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annota…
FlowMaps: Modeling Long-Term Multimodal Object Dynamics with Flow Matching
Joint spatial and temporal understanding of 3D scenes is a crucial requirement for robots deployed in everyday household environments. Such…
Learner-based Concept Drift Detection: Analysis and Evaluation
Machine learning algorithms deployed for evolving streaming environments must handle the non-stationary data distributions, commonly referr…
ScholarQuest: A Taxonomy-Guided Benchmark for Agentic Academic Paper Search in Open Literature Environments
Academic paper search is a core step in scientific research, and LLM-based search agents are emerging as a promising paradigm for iterative…
SPOT-E: Test-Time Entropy Shaping with Visual Spotlights for Frozen VLMs
Vision-language models (VLMs) often underperform on evidence intensive tasks because decisive visual evidence are small, localized, and eas…
Finetuning Vision-Language-Action Models Requires Fewer Layers Than You Think
Vision-Language-Action (VLA) models pre-trained on massive video-robot datasets have revolutionized robotic manipulation, yet their multi-b…
The Register Gap: A Meaning Intelligence Framework for Nigerian Public Discourse
We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that…
Editorial Alignment: A Participatory Approach to Engaging Editorial Expertise in LLM-mediated Knowledge Dissemination
The emergence of LLM-driven information services is reshaping the conditions under which public knowledge institutions operate, threatening…
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance…
Boundary Embedding Shaping with Adaptive Contrastive Learning for Graph Structural Disentanglement
Graph neural networks (GNNs) excel at aggregating neighbor information for classification, yet their performance is hindered by graph struc…
Robust $Q$-learning for mean-field control under Wasserstein uncertainty in common noise
In this article, we present a robust $Q$-learning algorithm for discrete-time mean-field control problems under Wasserstein uncertainty in…
AutoPass: Evidence-Guided LLM Agents for Compiler Performance Tuning
Large Language Models (LLMs) show promise for code compilation tasks, but applying them to runtime performance tuning is difficult due to c…
CRAX: Fast Safe Reinforcement Learning Benchmarking
Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. Wh…
DataMagic: Transforming Tabular Data into Data Insight Video
Data videos integrate dynamic charts, voice narration, and synchronized animations to communicate data insights as temporal narratives, mak…
LLM agent safety, multi-turn red-teaming, jailbreak benchmarks, adversarial robustness, safety-critical systems
Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness und…
Multi-View Decompilation for LLM-Based Malware Classification
Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable. Recent work suggests that la…
Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation
Classifier guidance is a way to control diffusion generation by using a noise-conditioned classifier to steer the sampling process toward a…
Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coord…
UltraQuant: 4-bit KV Caching for Context-Heavy Agents
Context-heavy agents place unusual pressure on the key-value (KV) cache: long prefixes are reused across many short turns, while concurrenc…
Optimal Order of Multi-Agent and General Many-Body Systems
This paper develops a general framework for analyzing multi-agent systems with feedback loops between agents actions and collective observa…
Contagion Networks: Evaluator Bias Propagation in Multi-Agent LLM Systems
When large language models serve as evaluators in multi-agent systems, their systematic evaluation biases propagate through the agent netwo…
Calibration Without Comprehension: Diagnosing the Limits of Fine-Tuning LLMs for Vulnerability Detection in Systems Software
Whether LLMs scoring well on vulnerability benchmarks genuinely reason about security or merely pattern-match on contaminated data remains…
FreeStyle: Free Control of Style-Content Dual-Reference Generation from Community LoRA Mining
Style-content dual-reference generation aims to synthesize an image that preserves the structure and semantics of a content reference while…
Efficient and Sound Probabilistic Verification for AI Agents
Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring approaches that formulat…
Sovereign Execution Brokers: Enforcing Certificate-Bound Authority in Agentic Control Planes
Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not…
SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm
Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for synthetic aperture rada…
Structuring and Tokenizing Distributed User Interest Context for Generative Recommendation
Generative recommendation is an emerging paradigm that has shown promise in industrial recommendation systems, aiming to predict users' nex…
How Transparent is DiffusionGemma?
LLM reasoning transparency is a critical affordance for understanding model decisions, mitigating misuse and misalignment, and debugging su…
UniMM: A Unified Mixture Model Framework for Multi-Agent Simulation
Simulation plays a crucial role in assessing autonomous driving systems, where the generation of realistic multi-agent behaviors is a key a…
MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning
Benchmarks play a central role in reinforcement learning (RL) research, yet their computational constraints often shape what is studied. De…
AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models
Post-training alignment of large language models often combines supervised fine-tuning (SFT) on expert demonstrations with reinforcement le…
Controlled Comparison of Machine Learning Models for Fault Classification and Localization in Power System Protection
The increasing complexity of modern power systems, driven by the integration of inverter-based and distributed energy resources, challenges…
SIGMA: Search-Augmented On-Demand Knowledge Integration for Agentic Mathematical Reasoning
Solving mathematical reasoning problems requires not only accurate access to relevant knowledge but also careful, multi-step thinking. Howe…
Creativity Reconsidered: Generative AI and the Problem of Intentional Agency
Many theorists maintain that conscious intentional agency is a necessary condition of creativity. We argue that this requirement, which we…
PCBSchemaGen: Reward-Guided LLM Code Synthesis for Printed Circuit Boards (PCB) Schematic Design with Structured Verification
Most LLM code-synthesis benchmarks rely on unit tests as the reward oracle, but PCB schematic design has none: correctness is defined by st…
One Probe Won't Catch Them All: Towards Targeted Deception Detection
Linear probes are a promising approach for monitoring AI systems for deceptive behaviour. Previous work has shown that a linear classifier…
Conditional Diffusion Guidance under Hard Constraint: A Stochastic Analysis Approach
We study conditional generation in diffusion models under hard constraints, where generated samples must satisfy prescribed events with pro…
SleepMaMi: A Universal Sleep Foundation Model for Integrating Macro- and Micro-structures
While the shift toward unified foundation models has revolutionized many deep learning domains, sleep medicine remains largely restricted t…
Mitigating Legibility Tax with Decoupled Prover-Verifier Games
As large language models become increasingly capable, it is critical that their outputs can be easily checked by less capable systems. Prov…
PrototypeNAS: Rapid Design of Deep Neural Networks for Microcontroller Units
Enabling efficient deep neural network (DNN) inference on edge devices with different hardware constraints is a challenging task that typic…
The Scaffold Effect: How Prompt Framing Drives Apparent Multimodal Gains in Clinical VLM Evaluation
Trustworthy clinical AI requires that performance gains reflect genuine evidence integration rather than surface-level artifacts. We evalua…
CareTransition-Audit: A Benchmark to Audit Discharge Summaries for Efficient Care Transitions
Incomplete or inconsistent discharge documentation drives care fragmentation and avoidable readmissions. Despite its critical role in patie…
Too long; didn't solve
Mathematical benchmarks consisting of a range of mathematics problems are widely used to evaluate the reasoning abilities of large language…
CogniFold: Always-On Proactive Memory via Cognitive Folding
Existing agent memory remains predominantly reactive and retrieval-based, lacking the capacity to autonomously organize experience into per…
ScaleWoB: Guiding GUI Agents with Coding Agents via Large-Scale Environmental Synthesis
GUI agents powered by large language models are advancing rapidly, creating urgent needs for evaluation and training based on realistic env…
FundaPod: A Multi-Persona Agent Pod Platform with Knowledge Graph Memory for AI-Assisted Fundamental Investment Research
Large language models (LLMs) are increasingly applied in finance, yet most existing work emphasizes trading signals or financial NLP tasks…
VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data
Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limite…
Science Earth: Towards A Planet-Scale Operating System for AI-Native Scientific Discovery
Scientific discovery demands intelligence, perseverance, and serendipity across vast search spaces. Today, top scientific capabilities rema…
MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning
Financial and tabular question answering requires more than fluent reasoning: answers must be grounded in the exact facts, formulas, units,…
Wisdom of Committee: Diverse Distillation from Large Foundation Models and Domain Experts
Knowledge distillation from foundation models to compact domain models is challenging due to substantial gaps in capacity, architecture, an…
Global Ease of Living Index: a machine learning framework for longitudinal analysis of major economies
The drastic changes in the global economy, geopolitical conditions, and disruptions such as the COVID-19 pandemic have impacted the cost of…
Simulation of Language Evolution under Regulated Social Media Platforms: A Synergistic Approach of Large Language Models and Genetic Algorithms
Social media platforms frequently impose restrictive policies to moderate user content, prompting the emergence of creative evasion languag…
A Deep Generative Model for Resting-State EEG Synthesis and Transferable Representation Learning
Resting-state EEG provides a non-invasive view of spontaneous brain activity, but extracting meaningful patterns is often limited by scarce…
TerraMind: Large-Scale Generative Multimodality for Earth Observation
We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO). Unlike other multimodal mode…
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
This paper bridges distribution shift and AI safety through a comprehensive analysis of their conceptual and methodological synergies. Whil…
Overcoming Labelled Data Scarcity for Defect Classification in Scanning Tunneling Microscopy
Scanning tunnelling microscopy (STM) is a powerful technique for imaging surfaces with atomic resolution, providing insight into physical a…
Assessment of Personality Dimensions Across Situations in Dyadic Role-Play Scenarios
Prior research indicates that users prefer assistive technologies whose personalities align with their own. This has sparked interest in au…
On the Limitations of Ray-Tracing for Learning-Based RF Tasks in Urban Environments
We study the realism of Sionna v1.0.2 ray-tracing for outdoor cellular links in central Rome. We use a real measurement set of 1,664 user-e…
Oranits: Mission Assignment and Task Offloading in Open RAN-based ITS using Metaheuristic and Deep Reinforcement Learning
In this paper, we explore mission assignment and task offloading in an Open Radio Access Network (Open RAN)-based intelligent transportatio…
Charting the Future of Scholarly Knowledge with AI: A Community Perspective
Despite the growing availability of tools designed to support scholarly knowledge extraction and organization, many researchers still rely…
From Construction to Injection: Edit-Based Fingerprints for Large Language Models
Reliable model fingerprints are essential for protecting large language models (LLMs) against unauthorized redistribution and commercial mi…
Enhancing Generative Auto-bidding with Offline Reward Evaluation and Policy Search
Auto-bidding is a critical tool for advertisers to improve advertising performance. Recent progress has demonstrated that AI-Generated Bidd…
RoboSSM: Scalable In-context Imitation Learning via State-Space Models
In-context imitation learning (ICIL) enables robots to learn tasks from prompts consisting of just a handful of demonstrations. By eliminat…
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation
Distilling the tool-use capabilities of large language models (LLMs) into small language models (SLMs) is essential for their practical app…
Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models
Reinforcement learning with verifiable rewards (RLVR) has delivered impressive gains in mathematical and multimodal reasoning and has becom…
Bid Farewell to Seesaw: Towards Accurate Long-tail Session-based Recommendation via Dual Constraints of Hybrid Intents
Session-based recommendation (SBR) aims to predict anonymous users' next interaction based on their interaction sessions. In the practical…
Bring My Cup! Personalizing Vision-Language-Action Models with Visual Attentive Prompting
While Vision-Language-Action (VLA) models generalize well to generic instructions, they struggle with personalized commands such as "bring…
Modeling Day-Long ECG Signals to Predict Heart Failure Risk with Explainable AI
Heart failure (HF) affects 11.8% of adults aged 65 and older, reducing quality of life and longevity. Preventing HF can reduce morbidity an…
AI-enhanced tuning of quantum dot Hamiltonians toward Majorana modes
We propose a neural network-based model capable of learning the broad landscape of working regimes in quantum dot simulators, and using thi…
Movement Primitives in Robotics: A Comprehensive Survey
Biological systems exhibit a continuous stream of movements, consisting of sequential segments, that allow them to perform complex tasks in…
PiDR: Physics-Informed Inertial Dead Reckoning for Autonomous Platforms
A fundamental requirement for full autonomy is the ability to sustain accurate navigation in the absence of external data, such as GNSS sig…
Policy-Embedded Graph Expansion: Networked HIV Testing with Diffusion-Driven Network Samples
HIV is a retrovirus that attacks the human immune system and can lead to death without proper treatment. In collaboration with the WHO and…
Bi-Anchor Interpolation Solver for Accelerating Generative Modeling
Flow Matching (FM) models have emerged as a leading paradigm for high-fidelity synthesis. However, their reliance on iterative Ordinary Dif…
Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic Methods
Policies learned via continuous actor-critic methods often exhibit erratic, high-frequency oscillations, making them unsuitable for physica…
DeFrame: Debiasing Large Language Models Against Framing Effects
As large language models (LLMs) are increasingly deployed in real-world applications, ensuring their fair responses across demographics has…
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
Distributed training of foundation models via $\texttt{DDP}$ is limited by interconnect bandwidth. While infrequent communication strategie…
Flickering Multi-Armed Bandits
We introduce Flickering Multi-Armed Bandits (FMAB) to model sequential decision-making in environments with changing action availability, w…
Reinforcement-aware Knowledge Distillation for LLM Reasoning
Reinforcement learning (RL) post-training has recently driven major gains in long chain-of-thought reasoning large language models (LLMs),…
Latent Gaussian Splatting for 4D Panoptic Occupancy Tracking
Capturing 4D spatiotemporal scene structure is crucial for the safe and reliable operation of robots in dynamic environments. However, exis…
The MAMA-MIA Challenge: Advancing Generalizability and Fairness in Breast MRI Tumor Segmentation and Treatment Response Prediction
Breast cancer is the most frequently diagnosed malignancy among women worldwide and a leading cause of cancer-related mortality. Dynamic co…
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
We investigate the use of zero-shot text-to-speech (ZS-TTS) as a data augmentation source for low-resource personalized speech synthesis. W…
Class-Incremental Motion Forecasting
Motion forecasting enables autonomous vehicles to anticipate scene evolution by predicting the future trajectories of dynamic agents. Howev…
The Autonomy Tax: Defense Training Breaks LLM Agents
Large language model (LLM) agents increasingly rely on external tools (file operations, API calls, database transactions) to autonomously c…
Vero: An Open RL Recipe for General Visual Reasoning
What does it take to build a visual reasoner that works across charts, science, spatial understanding, and open-ended tasks? The strongest…
Automated Standardization of Legacy Biomedical Metadata Using an Ontology-Constrained LLM Agent
Scientific metadata are often incomplete and noncompliant with community standards, limiting dataset findability, interoperability, and reu…
FM-Agent: Scaling Formal Methods to Large Systems via LLM-Based Hoare-Style Reasoning
LLM-assisted software development has become increasingly prevalent, and can generate large-scale systems, such as compilers. It becomes cr…
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been…
Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis
Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models. Existing methods mostly focus on regul…
CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation
Recovering editable CAD programs from images or 3D observations is central to AI-assisted design, but progress is difficult to measure beca…
Superhuman Safe and Agile Racing through Multi-Agent Reinforcement Learning
Autonomous systems have achieved superhuman performance in isolation or simulation, yet they remain brittle in shared, dynamic real-world s…
Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models
Sign language translation (SLT) remains constrained by the limited availability of paired sign-video/text corpora and by the heavy-tailed v…
"**Important** You should give me full credits!": Exploring Prompt Injection Attacks on LLM-Based Automatic Grading Systems
The emergence of large language models (LLMs) has significantly accelerated recent research on LLM-based automatic grading (AG) systems. Be…
Large Language Models Hack Rewards, and Society
Reinforcement learning (RL) has become a dominant post-training paradigm, enabling large language models (LLMs) to learn from rewards. We o…
Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations t…
The ACUTE Protocol: Operationalizing Language Model Activations for Better Calibration, Utility, and Trust
As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential. Calibration is…
KG-SoftMAP: Soft Knowledge-Graph Priors for Bayesian Network Structure Learning from Sparse Discrete Data
Learning Bayesian network (BN) structure from sparse discrete data is hard: when each instance records only a few variables, most variable…
Improving Crash Frequency Prediction from Simulated Traffic Conflicts Using Machine Learning Based Microsimulation
Traffic microsimulation combined with surrogate safety measures has increasingly been used as a proactive alternative to historical crash d…
ChatGPTで広告テスト、日本でも開始 非表示にする方法は?
米OpenAIの日本法人は、ChatGPTでの広告表示テストを日本でも始めたと発表した。広告を非表示にする方法は?
Source: Elastic agrees to buy CRV-backed DeductiveAI for up to $85M
DeductiveAI, a startup that uses AI to catch and resolve bugs in software, was founded just three years ago.
「ChatGPT広告」日本上陸 無料版と「Go」で表示、電通・博報堂など支援
OpenAIは2月に米国でテスト運用を開始。5月には日本など5カ国への拡大を予告していた。
米大企業の7割が導入する「Databricks」とは何者か? 評価額20兆円の「AI向けデータ基盤」
評価額約20兆円、Fortune 500の7割が利用するデータ・AI基盤企業Databricks。オープンソースのビッグデータ分散処理エンジン「Apache Spark」開発者らが2013年に創業した同社の軌跡と最新情報を解説する。
融資の決め手、決算書→「データ」「未来のシナリオ」へ 中小企業が資金調達に成功するための最大のポイントは?
中小企業の資金調達の在り方が、大きく変わろうとしている。融資特化型デジタルバンクである01(ゼロワン)銀行(大阪府吹田市)の大塚篤史副社長と北國銀行(金沢市)の竹内均氏(常務執行役員マーケティング部長)が、中小企業の経営や資金調達がどう変わっていくかの見解を語った。
工数「76%」削減 味の素グループが「経理AIエージェント」導入で先陣を切れたワケ
経理人材の不足が深刻化する一方で、経理パーソンが担う業務の幅は急速に広がっている。その解決策として期待されるのがAI活用だ。しかし、誤りが許されない経理業務では導入への慎重論も根強い。そんな中、味の素グループの財務・経理業務を担う味の素フィナンシャル・ソリューションズは、経費精…
「待ちの営業」はもう限界 ホンダがAIエージェントで挑む、商機を逃さない「濃い商談」の創出
顧客の購買行動が変化する中、ホンダが新車販売にAIエージェントを導入した。“濃い商談”を支援し、すでに成約も生まれているという。販売現場の変化を追う。
AIで要らなくなったSaaS、要るSaaSは、どれ? 日本の「SaaS is dead」の実態
エイトレッドが「AI時代に生き残るSaaSの条件に関する実態調査」の結果を公表。8割がSaaS見直しの必要性を実感する一方、AI代替の困難さや導入失敗の要因などが示された。
高級セレクトショップ「バーニーズ」が新品と中古の二刀流 富裕層の「初めての中古購入」を狙うワケ
原材料高騰や為替の乱高下に苦しむアパレル業界で、高級セレクトショップの雄「バーニーズ」が下した決断は、高級リユース市場への本格参入だった。「街の中古店には行かない」という富裕層の心理を突き、新品とユーズドを同じフロアで融合させる「二刀流」戦略の全貌とは。既存ビジネスとのカニバリ…
AI inference startup Baseten reportedly raising $1.5B months after its last mega-round
Startup Baseten is reportedly close to finalizing a $1.5 billion round at a $13 billion as the “inference gold rush" marches on.
話題の「Claude Mythos」登場で変わるセキュリティ AIエージェント時代の防衛策
AIによる攻撃が、月単位から時間単位で現実化する可能性が高まっている。最新AIモデルの登場で脆弱性発見の能力が広がる一方、企業ではAI利用のルールや管理体制が追い付いていない。AIエージェント時代に求められる新たな防衛策とは何か。
Snap spins off AI video team into new company, Dotmo, due to costs
The Snapchat maker is spinning off yet another internal unit. Dotmo will be composed of current Snap staff who are leaving the social media…
OpenAI is bringing on some big guns in the lead-up to its IPO
OpenAI is bulking up before its IPO, landing Transformer co-inventor Noam Shazeer from Google DeepMind and former Trump AI policy official…
Almost half of US singles feel negatively about AI in dating, Match says
About 47% of singles look negatively at the use of AI in dating -- but many dating app users are open to AI helping with profile punch-ups…
Amazon hopes to challenge Nvidia more directly by selling its AI chips
AWS is in talks to sell its chips to other data centers. CEO Andy Jassy has said this represents a $50 billion opportunity for the company.
AI data centers just got a government-mandated fast lane to the grid
FERC told grid operators to give data centers a fast lane for interconnections, but it failed to address electricity supply shortages.
The smartphone era created an attention crisis — slow tech is fixing it
“People just really want to take back control of their time, their lives, their attention... They’re down for whatever helps them do that.”
New usage analytics and updated spend controls for enterprises
OpenAI introduces new spend controls and usage analytics for ChatGPT Enterprise, helping organizations manage costs and scale AI with confi…
‘Queer Eye’ life coach Karamo Brown launches Kē, a wellness app featuring his AI digital clone
After spending a year and a half focusing on his own journey — from fitness and nutrition to meditation, sobriety, relationships, and perso…
General Intuition in talks to raise $300M at around $2B valuation
The startup trains embodied AI and world models using Medal’s dataset of 2 billion videos per year from 10 million monthly active users.
A tech worker-backed PAC is bringing a $5M knife to Big Tech’s $100M gunfight
Guardrails positions itself as a populist political movement that runs on small donations from people in the trenches of the AI boom.
2026-06-18(282件)
Pixi’s new iOS app turns text messages into interactive AR experiences
Forget stickers, GIFs, and emoji reactions. Pixi is betting that the next evolution of messaging is interactive augmented reality (AR).
Improving health intelligence in ChatGPT
Learn how GPT-5.5 Instant improves ChatGPT’s health and wellness responses with stronger reasoning, better context, clearer communication,…
チームみらい安野氏「牧歌的なAI開発の時代が終わった」 “ミュトス停止騒動”受け
「牧歌的なAI開発の時代が終わった」――チームみらいの安野貴博党首は6月18日の会見で、米AnthropicのAIモデル「Mythos 5」「Fable 5」の提供停止を巡る騒動を受け、このように述べた。
Using AI to help physicians diagnose rare genetic diseases affecting children
Researchers used an OpenAI reasoning model to help diagnose rare diseases, identifying 18 new diagnoses in previously unsolved cases.
NAVI-Orbital: First In-Orbit Demonstration of a Zero-Shot Vision-Language Model for Autonomous Earth Observation
As Earth Observation data generation outpaces downlink bandwidth and human-in-the-loop processing, a widening gap has emerged between onboa…
CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework
Vision-Language Models (VLMs) remain prone to hallucinations, producing fluent but visually unfaithful outputs. Existing chain-of-thought a…
Searching for Synergy in Shared Workspace Human-AI Collaboration
Automated AI agents are increasingly capable, yet many scientific and professional tasks require human judgment and contextual expertise. W…
CEO-Bench: Can Agents Play the Long Game?
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service.…
DeFAb: A Verifiable Benchmark for Defeasible Abduction in Foundation Models
A rule-based logic solver resolves every instance in our benchmark in under 50 microseconds with 100% accuracy; the best frontier language…
Optimizing Lithium Production Decisions under Geological, Demand, and Pricing Uncertainties: A POMDP Framework for Multi-Objective Decision Making
Decision making in lithium production is challenging, whether from an investor's perspective or a strategic production standpoint. Determin…
ForecastBench-Sim: A Simulated-World Forecasting Benchmark
Forecasting benchmarks for general-purpose AI systems usually inherit the constraints of the real world: outcomes resolve slowly, tail even…
What Must Generalist Agents Remember?
This paper develops a formal account of what generalist agents must store in memory in order to act near-optimally across multiple environm…
R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning
Robot soccer is a challenging testbed for multi-agent reinforcement learning because it combines partial observability, cooperative and adv…
ProfiLLM: Utility-Aligned Agentic User Profiling for Industrial Ride-Hailing Dispatch
Bringing Large Language Models (LLMs) into industrial ride-hailing dispatch as semantic feature extractors over platform-scale behavioral l…
WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Exi…
Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness
AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and fi…
Generative-Model Predictive Planning for Navigation in Partially Observable Environments
Navigation in partially observable environments presents a significant challenge for autonomous agents, requiring effective decision-making…
Skill-Guided Continuation Distillation for GUI Agents
Improving GUI agents typically relies on behavior cloning on expert trajectories. However, as the current policy deviates from the expert p…
SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety
Large language models (LLMs) are increasingly embedded in AI for Science (AI4Science) workflows, from scientific question answering and lit…
Decoupling Search from Reasoning: A Vendor-Agnostic Grounding Architecture for LLM Agents
Production LLM agents increasingly depend on real-time search, yet native search grounding bundles retrieval policy, provider choice, evide…
RTSGameBench: An RTS Benchmark for Strategic Reasoning by Vision-Language Models
Modern Vision-Language Models (VLMs) often struggle with strategic reasoning, i.e., anticipating and influencing other agents' actions, und…
ThinkDeception: A Progressive Reinforcement Learning Framework for Interpretable Multimodal Deception Detection
Multimodal deception detection is critical for identifying fraudulent intentions, yet existing approaches predominantly rely on end to end…
RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents
Multi-turn tool-use RL is bottlenecked by the rapid depletion of informative samples in static datasets. We observe that the gradient signa…
ARIADNE: Agnostic Routing for Inference-time Adapter DyNamic sElection
The increasing deployment of parameter-efficient fine-tuning (PEFT) has led to model ecosystems in which a single backbone is paired with m…
Towards an Agent-First Web: Redesigning the Web for AI Agents
The World Wide Web was built on an assumption held for three decades: the primary consumer of web content is a human being. This permeates…
Analysing drivers and interdependencies in European electricity markets using XAI
Electricity markets are inherently complex systems characterised by strong nonlinearities, high-dimensional interactions, and increasing in…
Human-AI Coevolution Dynamics: A Formal Theory of Social Intelligence Emergence Through Long-Term Interaction
Current conversational AI systems have made significant progress in language generation, personalization, and long-context interaction. How…
Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection
To achieve deeper safety alignment for large language models (LLMs), recent efforts have studied how to push safety interventions earlier i…
User as Engram: Internalizing Per-User Memory as Local Parametric Edits
Personal memory in a language model is two problems: content and reasoning skill. The brain keeps the two apart (a sparse, local engram in…
TxBench-PP: Analyzing AI Agent Performance on Small-Molecule Preclinical Pharmacology
Artificial intelligence (AI) agents promise to accelerate drug discovery by compressing interpretation and decision-making loops, but pract…
X+Slides: Benchmarking Audience-Conditioned Slide Generation
Automatically generating slide decks from source documents is an important application of large language models (LLMs). Existing benchmarks…
NeSyCat Torch: A Differentiable Tensor Implementation of Categorical Semantics for Neurosymbolic Learning
Neurosymbolic semantics is fragmented: classical, fuzzy, probabilistic and neural systems each define truth by their own inductive rules. N…
Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation
Post-training of reasoning language models is commonly driven by supervised distillation and reinforcement learning with verifiable rewards…
QSignAI: Quantum-Randomness-Seeded Identity Signatures at the Intersection of AI for Science and Science for AI
The 2024-2025 Nobel and Turing awards recognised AI and quantum science simultaneously. Yet no deployed system has brought these streams to…
Dynamic In-Group Persona Generation for Enhancing Human-AI Rapport
LLM-based chatbots are increasingly applied in interpersonal domains such as counseling and peer support, where establishing human-AI rappo…
From Memorization to Creation: Evaluating the Cognitive Depth of LLM-Generated Educational Questions
While LLMs show promise in automating educational content creation, their ability to generate questions that stimulate higher-order thinkin…
Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationsh…
Caring Without Feeling: Affective Dynamics as the Control Layer of Human-AI Agent Collaboration
AI agents that plan, retain memory across sessions, invoke external tools and act with partial autonomy are transforming human--AI collabor…
How Well Do Large Language Models Capture Human Personality?
Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that ri…
Simulating Hate Speech Cascades with Multi-LLM Agents: Empirical Grounding, Modeling Fidelity, and Intervention Strategies
Faithful modeling of hateful content propagation on online platforms remains an open problem for moderation research. Classical cascade mod…
Synthetic Resonance: A Framework for Growth-Oriented Human-AI Relationships
As human relationships with artificial intelligence systems become increasingly frequent and sustained, existing language and theory fail t…
EMORSION: Examining the Impact of Audio Parameters on Emotional Responses and Immersion in Film
EMORSION is an exploratory proof-of-concept study examining how film audio design shapes audience emotion and immersion in acinema setting.…
Towards Multi-Agent-Simulation-Based Community Note Evaluation
Community-based fact-checking that relies on cross-consensus is expanding rapidly on social media platforms. However, the delay and low-rat…
Mitigating Anchoring Bias in LLM-Based Agents for Energy-Efficient 6G Autonomous Networks
This paper presents an autonomous agentic resource negotiation framework designed to enable zero-touch network slicing in 6G architectures…
Continuous Audio Thinking for Large Audio Language Models
Large audio language models (LALMs) have shown impressive capabilities on diverse audio understanding tasks, ranging from speech transcript…
IOAH3: Importance-Driven Adaptive Spatial Partitioning
We present IOAH3 (Importance-Oriented Adaptive H3 partitioning), a computational method for constructing data-driven spatial partitions of…
Breaking the Solver Bottleneck: Training Task Generators at the Learnable Frontier
The limiting resource for training agents via reinforcement learning (RL) is increasingly frontier task supply: valid, solvable tasks just…
A Knowledge Theory of Capital:The Value of Natural and Artificial Intelligence
This volume develops a knowledge theory of capital for economies in which productive capacity increasingly resides in software, data, model…
Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming
Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forev…
A Link between Shock-wave Theory and Symmetry-reduced Stochastic Gradient Descent for Artificial Neural Networks
We develop a mathematically explicit link between shock-wave theory and the symmetry-quotiented learning dynamics of stochastic gradient de…
Attribution-Guided and Coverage-Maximized Pruning for Structural MoE Compression
Mixture-of-Experts (MoE) models scale compute efficiently, yet remain expensive to deploy due to their substantial memory footprint and inf…
DRIFT: Refining Instruction Data via On-Policy Data Attribution
Optimizing the training data distribution for Supervised Fine-Tuning (SFT) dictates the capability of Large Language Models (LLMs). While e…
TRIDENT: Breaking the Hybrid-Safety-Physics Coupling for Provably Safe Multi-Agent Reinforcement Learning
Safe coordination in networked cyber-physical systems forces learning algorithms to simultaneously handle hybrid discrete-continuous action…
SAGE: Retain-Aware Post-Hoc Sanitization of Final Unlearning Vector
Large Language Model (LLM) unlearning aims to remove undesirable knowledge or behaviors while preserving retained capabilities. Current unl…
Conflict-Aware Retriever Editing for Knowledge Injection Attacks on LLM-Based RAG Systems
Injecting malicious knowledge into retrieval-augmented generation (RAG) systems can manipulate retrieved evidence and mislead downstream ge…
Ghost Attractor Networks: Basin-Structured Dynamical Decoders for Closed-Loop Sequential Generation
Sequential output generation with large-scale Transformer and diffusion decoders pays a memory cost that grows with sequence length, plus i…
ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots
Air Traffic Control Operators (ATCOs) are vital in ensuring the safe, orderly, and efficient flow of air traffic, yet training capacity is…
SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior
Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features. Recent latent-space defenses increasingly rel…
Why SWAVE May Not Be All You Need:A Concept-Evolution Retrospective on Complex-Valued Recurrent Language Models
SWave is a complex-valued recurrent language model (169.26M parameters, D=384, L=16, T=2048) trained on FineWeb-Edu using 2xH100 NVL. It wa…
Agentra: A Supervisable Multi-Agent Framework for Enterprise Intrusion Response
Enterprise intrusion response still depends on static playbooks and analyst-driven triage, creating delay between alert generation and cont…
Self-CTRL: Self-Consistency Training with Reinforcement Learning
Language models (LMs) that faithfully describe their own behavior can more easily be audited, understood, and trusted by users. This paper…
SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents
Tool-using language-model agents introduce security failures that go beyond unsafe text: they can disclose protected objects, write persist…
Guava: An Effective and Universal Harness for Embodied Manipulation
Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models throu…
Redact or Keep? A Fully Local AI Cascade for Educational Dialogue De-Identification
Educational dialogue is a valuable but sensitive resource for research: the same transcripts that capture authentic learning often capture…
RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation
Graph-based retrieval at billion-node scale requires jointly solving three tightly coupled problems -- graph construction, representation l…
LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents
RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically ac…
Learning-Based Decision Making for Combustion Phasing Control in Multi-Fuel CI Engines with Latent Fuel Reactivity Estimation
Multi-fuel compression-ignition engines offer fuel flexibility but introduce uncertain, time-varying fuel reactivity, represented by cetane…
Deep Learning-Driven Inverse Design of Doherty Power Amplifiers Using Pixelated Combiners and Dual-State Impedance Synthesis
The output combiner of a Doherty power amplifier (PA) integrates load modulation, impedance matching, and phase compensation within a singl…
Deep-Learning-Based Pixelated Microwave Filter Design and Characterization using Electro-Optical Electric-Field Measurements
Traditional microwave filter design typically relies on iterative parameter tuning and predefined topologies, which limits design space and…
A Variational Framework for LLM Generator-Regulator Games
This paper develops a variational framework for regulated language generation. Starting from autoregressive token sampling, we derive the i…
From Specification to Execution: AI Assisted Scientific Workflow Management
Scientific workflow management systems (WMS) support scalable and reproducible execution of complex pipelines, but workflow design, impleme…
CAOA -- Completion-Assisted Object-CAD Alignment
Accurately aligning CAD models to their corresponding objects in indoor RGB-D scans is a central challenge in 3D semantic reconstruction. T…
TMR-GGNN: Credit Card Fraud Detection based on Time-Aware Multi-Relational Guided Graph Neural Network
In recent years, credit card fraud detection has faced significant challenges due to highly imbalanced data, evolving fraud patterns, and c…
Veriphi: Attack-Guided Neural Network Verification with Dataset-Dependent Training Methods
We present Veriphi, a GPU-accelerated neural network verification system that combines fast adversarial attacks with formal bound certifica…
What Does the Weight Norm Control in Grokking? Logit-Scale Mediation under Cross-Entropy
Grokking, the delayed jump from memorization to generalization, is usually tied to the weight norm: a smaller norm generalizes sooner. We a…
Structured Representation Learning with Locally Linear Embeddings and Adaptive Feature Fusion
Neuroscientific research has revealed that the brain encodes complex behaviors by leveraging structured, low-dimensional manifolds and dyna…
MagpieTTS-LF: Inference-Time Long-Form Speech Generation Without Training on Long-Form data
Neural Text-to-Speech (TTS) systems achieve remarkable quality on short utterances but long-form speech generation shows prosodic drift, sp…
SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR
The standard heuristic of selecting the SFT checkpoint with the highest pass@1 for GRPO can fail when SFT compresses the rollout distributi…
Neural Phase Correlation
Correspondence is fundamentally relational: it seeks the unknown transformation between two observations of a common scene, not the content…
PSyGenTAB: A Privacy-Preserving Framework for Synthetic Clinical Tabular Data Generation via Constrained Optimization
The development of medical AI is constrained by limited access to high-quality clinical data due to institutional silos and strict privacy…
As You Wish: Mission Planning with Formal Verification using LLMs in Precision Agriculture
Though robotic systems are now being commercialized and deployed in various industries, many of these systems are highly specialized and of…
Sparsity Curse: Understanding RLVR Model Parameter Space from Model Merging
Reinforcement Learning with Verifiable Reward (RLVR) has emerged as a powerful post-training paradigm that surpasses Supervised Fine-Tuning…
AI Sandboxes: A Threat Model, Taxonomy, and Measurement Framework
AI systems are increasingly evaluated in bounded environments that combine isolation, simulation, instrumentation, supervision, and evidenc…
Engagement Intensity as a Learner-Modeling Signal for Adaptive AI Ethics Instruction
Adaptive AI ethics instruction in graduate research training benefits from intake measures that reflect differences in prior LLM experience…
Correcting Sensor-Induced Distribution Drift with Wasserstein Adversarial Learning
The quality of recorded data depends on the stability of the sensor system that acquires it. Sensor motion and aging can degrade the perfor…
Multi-Modal Hyper-Graph Fusion for Low-Light Crowd Counting
Crowd counting is a fundamental task in computer vision. However, crowd counting in low-light environments remains largely underexplored, d…
APT: Atomic Physical Transitions for Causal Video-Language Understanding
Physical events are not understood by their names alone, but by the causal state changes that compose them. A clip-level label such as "bou…
Dual Dimensionality for Local and Global Attention
Decoder-only Transformers compute attention over the KV cache of preceding tokens. Keys (and Values) are typically represented with the sam…
Benchmarking Action Spaces in Reinforcement Learning for Vision-based Robotic Manipulation
In real-world reinforcement learning (RL), the choice of action space can play a key role in shaping motion smoothness, safety, and overall…
Better Adherence, Richer Context: A Field Evaluation of LLM-Powered Conversational Voice Diaries for Sleep
Sleep diaries are central to behavioral sleep medicine and cognitive behavioral therapy for insomnia, yet daily completion is difficult to…
MIDS: Detecting Stealthy Masquerade and Tampering Attacks on CAN Bus via Bidirectional Mamba
The Controller Area Network (CAN) protocol is the primary communication standard for Electronic Control Units (ECUs) in modern vehicles, bu…
Steerable Cultural Preference Optimization of Reward Models
It is essential for large language model (LLM) technology to serve many different cultural sub-communities in a manner that is acceptable t…
QC-GAN: A Parameter-Efficient Quaternion Conformer GAN for High-Fidelity Speech Enhancement
We propose a parameter-efficient speech enhancement framework, Quaternion Conformer GAN (QC-GAN), which combines a Quaternion Conformer gen…
Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance
The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated…
AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice
There exist numerous tutor training platforms. However, few provide AI-driven training and evaluation for human tutors based on real-life p…
Code-Augur: Agentic Vulnerability Detection via Specification Inference
The advent of agentic vulnerability detection is already becoming a watershed moment for software security. Audits conducted entirely by au…
BCL: Bayesian In-Context Learning Framework for Information Extraction
Existing information extraction (IE) tasks increasingly adopt in-context learning (ICL) with large language models. However, current approa…
EffiNav: Fusing Depth and Vision-Language for Efficient Object Goal Navigation
To locate a target object while exploring the unknown environment is a fundamental capability for autonomous agents, with applications rang…
PEC-Home: Interpretation of Progressively Elliptical Commands in Smart Homes
Recent advancements in Large Language Models (LLMs) have empowered home assistants with natural language interaction capabilities. However,…
Augmenting Dysarthric Speech Severity Assessment with MOS Supervision
Dysarthria is a speech disorder marked by reduced intelligibility and communicative effectiveness. Automatic utterance-level assessment of…
LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis
Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract v…
NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization
Reliable sound source localization is fundamental to robot audition, enabling autonomous robots to perceive spatial cues and operate effect…
scGTN: Deep Siamese Graph Transformer Network for Single-cell RNA Sequencing Clustering
Single-cell RNA sequencing (scRNA-seq) serves a pivotal role in characterizing gene expression at the cellular level, enabling the identifi…
Bounded Context Management for Tabular Foundation Models on Stream Learning
Tabular stream learning requires predictions on sequentially arriving examples under distribution shift. While standard methods adapt by up…
Dual-Channel Grounded World Modeling (DCGWM): Structural Prevention of Objective Interference Collapse via Heterogeneous External Grounding with Inward-Only Gradient Flow
Joint Embedding Predictive Architectures (JEPAs) are a leading approach to world model representation learning. We identify a failure mode…
Leveraging Energy Features for Surface Classification with Deep Learning: A Comparative Analysis Across Three Independent Datasets
The energy-based method remains a comparatively underexamined approach for surface classification in mobile robotics, despite promising res…
TW-LegalBench: Measuring Taiwanese Legal Understanding
Large language models (LLMs) have shown impressive capabilities across diverse tasks, yet their performance on jurisdiction-specific legal…
Morpheus: A Morphology-Aware Neural Tokenizer and Word Embedder for Turkish
Turkish is agglutinative: meaning is carried by morphemes, yet the subword tokenizers that drive modern language models split words by corp…
Graph Grounded Cross Attention Transformer Neural Network for Structurally Constrained Full Event Sequence Generation in Predictive Process Monitoring
Structurally constrained event sequence generation remains challenging because generated paths must preserve transition feasibility, tempor…
Two-Phase Bilevel Search for the Moving-Target Traveling Salesman Problem with Moving Obstacles
The Moving-Target Traveling Salesman Problem (MT-TSP) seeks a minimum cost trajectory for an agent that departs from a static depot, visits…
SWE-Future: Forecast-Conditioned Data Synthesis for Future-Oriented Software Engineering Agents
Realistic coding-agent benchmarks often replay public GitHub issues and pull requests, making them vulnerable to overlap with model pretrai…
Generating Natural and Expressive Robot Gestures through Iterative Reinforcement Learning with Human Feedback using LLMs
Expressive gestures are essential for natural and effective communication, complementing speech when verbal cues alone are insufficient (e.…
Private Learning with Public Feature Conditioning
We study differentially private (DP) regression in settings where each data sample includes public, non-sensitive features -- common in app…
RedactionBench
Large Language Models are increasingly applied to sensitive domains that require redaction of personally identifiable information (PII). Wh…
Bayesian Anytime Pareto Set Identification for Multi-Objective Multi-Armed Bandits
Identifying Pareto optimal solutions is critical to support multi-objective decision-making. We introduce the first anytime Multi-Objective…
Closing the Loop: PID Feedback Control for Interpretable Activation Steering in Symbolic Music Generation
Transformer-based architectures have significantly advanced the generation of complex symbolic sequences, yet a significant gap remains in…
SHIFT: Semantic Harmonization via Index-side Feature Transformation for Multilingual Information Retrieval
With the rapid expansion of massive multilingual corpora, Multilingual Information Retrieval (MLIR) has emerged as a critical technology fo…
Learning from Own Solutions: Self-Conditioned Credit Assignment for Reinforcement Learning with Verifiable Rewards
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in training LLMs for reasoning tasks, but representat…
Rescaling MLM-Head for Neural Sparse Retrieval
Learned sparse retrieval (LSR) models such as SPLADE have traditionally used BERT-style masked language models as backbone encoders. A natu…
Reinforcement Learning Foundation Models Should Already Be A Thing
Foundation models for language and vision are powered by internet-scale data, while structured domains (tabular prediction, time-series for…
SwitchBraidNet: Quantisation-Aware Lightweight Architecture for Hybrid Brain-Computer Interface
Hybrid brain-computer interfaces (BCIs) that integrate motor imagery (MI) and steady-state visual evoked potentials (SSVEP) provide high-di…
Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets
Sequential decision problems often exhibit an asymmetric evolution of information and decision flexibility: as a decision cycle unfolds, th…
Space Is Intelligence: Neural Semigroup Superposition for Riemannian Metric Generation
Traditional approaches place intelligence in the agent, whether as a learned policy or a search procedure. We instead place intelligence in…
Beyond Reward Engineering: A Data Recipe for Long-Context Reinforcement Learning
Long-context reasoning is an essential capability for large language models, particularly when they are deployed as autonomous agents that…
Target-confidence Recourse Using tSeTlin machines: TRUST
Counterfactual explanations are widely used to provide algorithmic recourse in high-stakes decision-making systems. Most existing methods s…
Improving Human-Robot Teamwork in Urban Search and Rescue Through Episodic Memory of Prior Collaboration
Effective human-robot teamwork requires robots to adapt to partners, situations, and task dynamics from the start of an interaction. In the…
Skill-MAS: Evolving Meta-Skill for Automatic Multi-Agent Systems
Large Language Model (LLM)-based automatic Multi-Agent Systems (MAS) generation has become a crucial frontier for tackling complex tasks. H…
Aligning Implied Statements for Implicit Hate Speech Generalizability with Context-Bounded Semi-hard Negative Mining
Classifying implicit hate speech remains a challenge, as intent is often masked through insinuation and context rather than explicit slurs.…
URDF Synthesis from RGB-D Sequences via Differentiable Joint Inference and Energy-Consistent Verification
Reconstructing simulation-ready digital twins of articulated objects from sensor observations remains constrained by two persistent gaps: (…
Scaling Learning-based AEB with Massive Unlabeled Data
This paper studies how to scale learning-based automatic emergency braking (AEB) with massive unlabeled fleet data under production constra…
Domain-Shift Aware Neural Networks for Unbalance Characterization in Rotating Systems
This work investigates the application of a domain-shift aware neural network for regression tasks aimed at estimating unbalance masses in…
SAERec: Constructing Fine-grained Interpretable Intents Priors via Sparse Autoencoders for Recommendation
Intent-based recommender systems have gained significant attention for improving accuracy and interpretability by modeling the underlying m…
As Easy as Rocket Science: Assessing the Ability of Large Language Models to Interpret Negation in Figurative Language
Figurative language and negation are two areas that challenge current language models, however, both are widely used throughout written and…
TransitNet: A Compact Attention-Augmented Deep Learning Framework for Low-SNR Transit Blind Searches
Motivated by the observational incompleteness of intermediate-to-long-period Earth-size planets, we present TransitNet, a compact attention…
A Controlled Benchmark of Quantum-Latent GAN Augmentation for Brain MRI
Medical image classification is often constrained by limited labeled data, motivating generative augmentation; recently, quantum generative…
CAPRA: Scaling Feedback on Software Architecture Deliverables with a Multi-Agent LLM System
Automated assessment in software engineering education has advanced significantly for code grading and essay scoring. However, reviewing so…
Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering
Recent advances in large language models (LLMs) have given rise to time-series question answering (TSQA), which formulates time-series anal…
G-IdiomAlign: A Gloss-Pivoted Benchmark for Cross-Lingual Idiom Alignment
Idioms are difficult to transfer across languages due to their non-compositionality and weak surface-form grounding, making literal mapping…
TRAP: Benchmark for Task-completion and Resistance to Active Privacy-extraction
Agents are increasingly deployed in document-intensive workflows where sensitive private information is not an edge case but a routine inpu…
Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training
Reinforcement learning (RL) post-training of Diffusion Transformers (DiTs) is prohibitively expensive, requiring thousands of high-end GPUs…
FoMoE: Breaking the Full-Replica Barrier with a Federation of MoEs
Pre-training Large Language Models (LLMs) typically demands large-scale infrastructure with tightly coupled hardware accelerators. While in…
A Hybrid LSTM--Vision Transformer Architecture for Predicting HRRR Forecast Errors
Forecast errors in high-resolution numerical weather prediction (NWP) systems are often linked to unresolved planetary boundary layer (PBL)…
Where Did the Variability Go? From Vibe Coding to Product Lines by Regeneration
In vibe coding, an emerging AI-driven paradigm, an LLM generates an entire program from a natural language prompt, but what happens to the…
ProductConsistency: Improving Product Identity Preservation in Instruction-Based Image Editing via SFT and RL
Recent advances in instruction-based image editing have enabled models to perform complex visual edits from natural language instructions.…
Leadership as Coordination Control: Behavioral Signatures and the Recovery-Advantage Boundary in Multi-Agent LLM Teams
Team science holds that leadership is contingent: it helps only under specific conditions, and capable, autonomous teams may need none at a…
Equivariant Graph Neural Networks Improve Optical Spectra Prediction for Materials Screening
Scalable prediction of optical spectra is a critical component of high-throughput materials screening for optoelectronic applications such…
Pareto Q-Learning with Reward Machines
We present Pareto Q-Learning with Reward Machines (PQLRM), a multi-objective reinforcement learning algorithm for tasks whose reward struct…
A Technical Taxonomy of LLM Agent Communication Protocols
As large language models (LLMs) advance and multi-agent systems aim to overcome the limits of standalone agents, robust communication proto…
OrthoReg: Orthogonal Regularization for Hybrid Symbolic-Neural Dynamical Systems
Dynamical systems are fundamental to modeling the natural world, yet modeling them involves a persistent trade-off: manually prescribed mec…
AdsMind: A Physics-Grounded Multi-Agent System for Self-Correcting Discovery of Adsorption Configurations on Heterogeneous Catalyst Surfaces
Identifying the lowest-energy surface-adsorbate configuration is critical for modeling heterogeneous catalysis, yet exhaustive exploration…
Essential Subspace Merging for Multi-Task Learning
Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained ch…
A Clinician-Centered Pipeline for Annotation and Evaluation in Ultrasound AI Studies
Clinician-centered evaluation is critical for validating medical AI systems, especially in ultrasound imaging where quantitative metrics do…
Hardware- and Vision-in-the-Loop Validation of Deep Monocular Pose Estimation for Autonomous Maritime UAV Flight
Autonomous UAV operations on ships require reliable vision-based relative pose estimation, yet at-sea validation is costly, weather-depende…
Compute Efficiency and Serial Runtime Tradeoffs for Stochastic Momentum Methods
Stochastic momentum methods such as heavy ball (HB), Nesterov momentum, and variants of Accelerated SGD (ASGD) [Kidambi et al., 2018] are w…
Language Models as Interfaces, Not Oracles: A Hybrid LLM-ML System for Pediatric Appendicitis
Large language models (LLMs) can make clinical decision support more accessible by interpreting free-text documentation, but their direct u…
The More the Merrier: Combining Properties for ABox Abduction under Repair Semantics for ELbot
Abduction is a central approach to explain missing entailments from a knowledge base by providing a hypothesis, that would, if added to the…
Forecasting what Matters: Decision-Focused RL for Controlled EV Charging with Unknown Departure Times
The recent growth of EV adoption poses challenges for power systems, including increased peak demand and potential grid instability. Smart…
Machine Unlearning for the XGBoost Model with Network Intrusion Datasets
Machine Unlearning (MU) has emerged as an important technique for removing specific data points from trained models without requiring full…
Mechanism-Guided Selective Unlearning for RLVR-Induced Reasoning
We propose MAST (Mechanism-Aligned Selective Targeting), a mechanism-guided method for unlearning RLVR-induced reasoning with substantially…
STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability
Reinforcement Learning with Verifiable Rewards algorithms like GRPO have emerged as the dominant post-training paradigm for complex reasoni…
A Taxonomy of Mental Health and Technology Needs for Alzheimer's and Dementia Caregivers
Family members caring for individuals with Alzheimer's disease and related dementias (AD/ADRD) provide the foundation of long-term care wor…
OneCanvas: 3D Scene Understanding via Panoramic Reprojection
Existing approaches to 3D scene understanding in Vision-Language Models (VLMs) either rely on complex, model-specific geometry encoders or…
A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2
Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation mo…
Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA
The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized domains and languages, yet…
Correct Yourself, Keep My Trust: How Self-Correction and Social Connection Shape Credibility in Social Chatbots
When social chatbots make mistakes, and they do, how they recover determines whether users trust them again. Social chatbots are increasing…
Explaining Attention with Program Synthesis
A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic desc…
Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents
Production data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts who must collaborative…
Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors
Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcripti…
UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning
Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit…
Large-Scale OD Matrix Estimation with A Deep Learning Method
The estimation of origin-destination (OD) matrices is a crucial aspect of Intelligent Transport Systems (ITS). It involves adjusting an ini…
Recursive Joint Simulation in Games
Game-theoretic dynamics between AI agents could differ from traditional human-human interactions in various ways. One such difference is th…
Fully Geometric Multi-Hop Reasoning on Knowledge Graphs with Transitive Relations
Multi-hop logical reasoning on knowledge graphs requires faithfully mapping the logical semantics to latent space. Current geometric embedd…
PosterForest: Hierarchical Multi-Agent Collaboration for Scientific Poster Generation
Automating scientific poster generation requires hierarchical document understanding and coherent content-layout planning. Existing methods…
Structured Cognitive Loop for Behavioral Intelligence in Large Language Model Agents (Extended Revision: From Behavioral Architecture to Epistemic Accountability)
The central challenge for AI agents is not only performance but accountability. Agents that act through opaque prompt sequences may produce…
The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
When an AI assistant remembers that Sarah is a single mother working two jobs, does it interpret her stress differently than if she were a…
An In-depth Study of LLM Contributions to the Bin Packing Problem
Recent studies have suggested that Large Language Models (LLMs) could provide interesting ideas contributing to mathematical discovery. Thi…
RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories
Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often…
Towards Understanding What State Space Models Learn About Code
State Space Models (SSMs) have emerged as an efficient alternative to the Transformer architecture. Prior work shows that, when trained und…
Enhancing CVRP Solver through LLM-driven Automatic Heuristic Design
The Capacitated Vehicle Routing Problem (CVRP), a fundamental combinatorial optimization challenge, focuses on optimizing fleet operations…
InfoPO: Information-Driven Policy Optimization for User-Centric Agents
Real-world user requests to LLM agents are often underspecified. Agents must interact to acquire missing information and make correct downs…
Robust Regularized Policy Iteration under Transition Uncertainty
Offline reinforcement learning (RL) enables data-efficient and safe policy learning without online exploration, but its performance often d…
Information-Theoretic Measures in AI: A Practical Decision Guide
Information-theoretic (IT) measures are ubiquitous in artificial intelligence: entropy drives decision-tree splits and uncertainty quantifi…
A Distributionally Robust Reinforcement Learning Framework for Constrained Urban EV Dispatch
We study city-scale control of electric-vehicle (EV) ride-hailing fleets where dispatch, repositioning, and charging decisions must respect…
FinSTaR: Towards Financial Reasoning with Time Series Reasoning Models
Time series (TS) reasoning models (TSRMs) have shown promising capabilities in general domains, yet they consistently fail in the financial…
LLM-Evolved Domain-Independent Heuristics for Symbolic AI Planning
Heuristic search is the dominant paradigm in symbolic AI planning, and the strongest heuristics are the result of decades of work by planni…
Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems
Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The de…
AI Sovereignty as National Learning Capacity: A Human-Centered Learning Mechanics Viewpoint on France, the United States, and China
Artificial intelligence in France is often discussed through separate dimensions such as investment, compute, regulation, employment, sover…
SkillRevise: Improving LLM-Authored Agent Skills via Trace-Conditioned Skill Revision
Agent skills are procedural artifacts that enable LLM agents to execute workflows, verify constraints, and recover from failures. Existing…
DN-Hypo-Pipeline: An AI-Driven Workflow for Hypothesis Generation via Large Language Models and Scientific Explanations
A scientific hypothesis is the first step in research and undergoes experimental validation, yet it also reflects a deep understanding of a…
The Art of Interrogation: Consistency Amplifies Factuality in Spatial Reasoning
Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Ex…
"Did you lie?" Evaluating Lie Detectors across Model Scale and Belief-Verified Model Organisms
Robust lie detectors for language models could enable powerful techniques for auditing, monitoring, and post-hoc investigation of model beh…
Simple Domain Generalization Methods are Strong Baselines for Open Domain Generalization
In real-world applications, a machine learning model is required to handle an open-set recognition (OSR), where unknown classes appear duri…
A DeepLearning Framework for Dynamic Estimation of Origin-Destination Sequence
OD matrix estimation is a critical problem in the transportation domain. The principle method uses the traffic sensor measured information…
Quality Perceptions and Intended Engagement in Response to AI-Generated and AI-Assisted News
The increasing use of artificial intelligence (AI) in news production raises important questions about how audiences perceive and respond t…
Scalable Batch Bayesian Optimization Via Subspace Acquisition Functions
Extending Bayesian optimization to batch evaluation can enable the designer to make the most use of parallel computing technology. However,…
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
Controllable image-to-video (I2V) generation transforms a reference image into a coherent video guided by user-specified control signals. W…
Efficient Zeroth-Order Federated Finetuning of Language Models on Resource-Constrained Devices
Federated Learning (FL) is a promising paradigm for finetuning Large Language Models (LLMs) across distributed data sources while preservin…
Depth-Width tradeoffs in Algorithmic Reasoning of Graph Tasks with Transformers
Transformers have revolutionized the field of machine learning. In particular, they can be used to solve complex algorithmic problems, incl…
Generalized Kullback-Leibler Divergence Loss
In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupl…
Revealing Hidden Vulnerabilities in Autoencoders through Gradient Signal Restoration
Adversarial robustness of deep autoencoders (AEs) has received less attention than that of discriminative models, although their compressed…
Signals of Provenance: Practices & Challenges of Navigating Indicators in AI-Generated Media for Sighted and Blind Individuals
AI-Generated (AIG) content has become increasingly widespread by recent advances in generative models and the easy-to-use tools that have s…
Revisiting Active Speaker Detection: An In-the-Wild Benchmark for Generalization and Robustness
We present UniTalk, a novel dataset emphasizing challenging scenarios to enhance model generalization for the task of active speaker detect…
ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark
Large language models (LLMs) are increasingly applied to symbolic mathematics, yet existing evaluations often conflate pattern memorization…
Self-Evolving Multi-Agent Systems via Textual Backpropagation
Leveraging multiple Large Language Models (LLMs) has proven effective for addressing complex, high-dimensional tasks, but current approache…
Grids Often Outperform Implicit Neural Representations at Compressing Dense Signals
Implicit Neural Representations (INRs) have recently shown impressive results, but their fundamental capacity, implicit biases, and scaling…
From Memorization to Parameter Interference: How Overtraining Experts Harms Model Merging
Modern deep learning is increasingly characterized by the use of open-weight foundation models that can be fine-tuned on specialized datase…
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
Current unlearning methods for LLMs optimize on the private information they seek to remove by incorporating it into their fine-tuning data…
Enhancing Fatigue Detection through Heterogeneous Multi-Source Data Integration and Cross-Domain Modality Imputation
Fatigue detection for human operators is important in safety-related applications such as aviation, mining, and long-haul transport. Reliab…
When Cars Have Stereotypes: Auditing Demographic Bias in Objects from Text-to-Image Models
While prior research on text-to-image generation has predominantly focused on biases in human depictions, demographic bias in generated obj…
From Values to Tokens: An LLM-Driven Framework for Context-aware Time Series Forecasting via Symbolic Discretization
Time series forecasting plays a vital role in supporting decision-making across a wide range of critical applications, including energy, he…
Surrogate Benchmarks for Model Merging Optimization
Model merging techniques aim to integrate the abilities of multiple models into a single model. Most model merging techniques have hyperpar…
Probing Semantic Alignment, Lexical Invariance, and Syntactic Influence in LLM Metaphor Processing
Large language models (LLMs) achieve strong performance on metaphor detection and interpretation tasks, yet it remains unclear what such be…
Rethinking Cross-lingual Gaps from a Statistical Viewpoint
Any piece of knowledge is usually expressed in one or a handful of natural languages on the web or in any large corpus. Large Language Mode…
R2BC: Multi-Agent Imitation Learning from Single-Agent Demonstrations
Imitation Learning (IL) is a natural way for humans to teach robots, particularly when high-quality demonstrations are easy to obtain. Whil…
DecNefSimulator: A Modular, Interpretable Framework for Decoded Neurofeedback Simulation Using Generative Models
Decoded Neurofeedback (DecNef) is a promising non-invasive approach to brain modulation with wide-ranging applications in neuromedicine and…
Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
Multimodal Large Language Models (MLLMs) are increasingly deployed in stateless systems, such as autonomous driving and robotics. This pape…
Learning Patient-Specific Disease Dynamics with Latent Flow Matching for Longitudinal Imaging Generation
Understanding disease progression is a central clinical challenge with direct implications for early diagnosis and personalized treatment.…
Improving Scientific Document Retrieval with Academic Concept Index
Adapting general-domain retrievers to scientific domains is challenging due to the scarcity of large-scale domain-specific relevance annota…
SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding
Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. How…
DeepInflation: an AI agent for research and model discovery of inflation
We present DeepInflation, an AI agent designed for research and model discovery in inflationary cosmology. Built upon a multi-agent archite…
InstructTime++: Time Series Classification with Multimodal Language Modeling via Implicit Feature Enhancement
Most existing time series classification methods adopt a discriminative paradigm that maps input sequences directly to one-hot encoded clas…
Retell, Reward, Repeat: Reinforcement Learning for Narrative Theory-Informed Story Retelling
Counterfactual story retelling exposes LLM shortcomings in constrained narrative solution spaces where they can no longer rely on recalling…
HeRo-Q: A General Framework for Stable Low Bit Quantization via Hessian Conditioning
Post Training Quantization (PTQ), a mainstream model compression technique, often leads to the paradoxical 'low error, high loss' phenomeno…
LLM Compression by Block Removal with Constrained Binary Optimization
In this paper, we formulate the compression of large language models (LLMs) by optimally deleting transformer blocks (``block removal'') as…
Posterior Continuation with Noise-Conditioned Frequency Exposure for Diffusion Inverse Problems
Diffusion posterior sampling solves inverse problems by combining a pretrained diffusion prior with measurement-consistency guidance. Howev…
Improve Large Language Model Systems with User Logs
Scaling training data and model parameters has long driven progress in large language models (LLMs), but this paradigm is increasingly cons…
Do Neural Networks Lose Plasticity in a Gradually Changing World?
Continual learning has become a trending topic in machine learning. Recent studies have discovered an interesting phenomenon called loss of…
Narrative Theory-Driven LLM Methods for Automatic Story Generation and Understanding: A Survey
Applications of narrative theories using large language models (LLMs) deliver promising methods in automatic story generation and understan…
Detecting High-Potential SMEs with Heterogeneous Graph Neural Networks
Small and Medium Enterprises (SMEs) constitute 99.9% of U.S. businesses and generate 44% of economic activity, yet systematically identifyi…
ActMem: Bridging the Gap Between Memory Retrieval and Reasoning in LLM Agents
Memory management is essential for LLM agents in long-term interactions. Current memory frameworks typically treat agents as passive ``reco…
Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation
Speech-aware large language models (LLMs) can accept speech inputs, yet their training objectives largely emphasize linguistic content or s…
Something from Nothing: Data Augmentation for Robust Severity Level Estimation of Dysarthric Speech
Dysarthric speech quality assessment (DSQA) is critical for clinical diagnostics and inclusive speech technologies. However, subjective eva…
A Convex Route to Thermoelasticity: Learning Internal Energy and Dissipation
We present a physics-based neural network framework for the discovery of constitutive models in fully coupled thermomechanics. In contrast…
MemRerank: Preference Memory for Personalized Product Reranking
LLM-based shopping agents increasingly rely on long purchase histories and multi-turn interactions for personalization, yet naively appendi…
A CEFR-Inspired Classification Framework with Fuzzy C-Means To Automate Assessment of Programming Skills in Scratch
Context: Schools, training platforms, and technology firms increasingly need to assess programming proficiency at scale with transparent, r…
IPSL-AID: Generative Diffusion Models for Climate Downscaling from Global to Regional Scales
Effective adaptation and mitigation strategies for climate change require high-resolution projections to inform strategic decision-making.…
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
Web agents automate browser tasks, ranging from simple form completion to complex workflows like ordering groceries. While current benchmar…
The Long Delay to Arithmetic Generalization: When Learned Representations Outrun Behavior
Grokking in transformers trained on algorithmic tasks is characterized by a long delay between training-set fit and abrupt generalization,…
From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks
Jailbreak attacks expose a persistent failure mode in safety-aligned LLMs: models can be pushed into harmful behavior, but the internal rep…
TopBench: A Benchmark for Implicit Predictive Reasoning in Tabular Question Answering
Large Language Models (LLMs) have advanced Table Question Answering, where most queries can be answered by extracting information or simple…
Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories
We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories. JE…
Beyond Similarity: Temporal Operator Attention for Time Series Analysis
A persistent paradox in time-series forecasting is that structurally simple MLP and linear models often outperform high-capacity Transforme…
Pyramid Self-Contrastive Learning for Single-shot Test-time Ultrasound Image Denoising
The inherent electronic and speckle noise complicates clinical interpretation of ultrasound images. Conventional denoising methods rely on…
Controllable Quantum Memory Capacity in Quantum Reservoir Networks with Tunable partial-SWAPs
In the field of quantum reservoir computing (QRC), many different computational models and architectures have been proposed. From these mod…
Hilbert-Geo: Solving Solid Geometric Problems by Neural-Symbolic Reasoning
Geometric problem solving, as a typical multimodal reasoning problem, has attracted much attention and made great progress recently, howeve…
A Survey on Deep Learning Architectures for Point Cloud Classification and Segmentation
Point cloud stands as the most widely adopted format for representing 3D shapes and scenes due to its simplicity and geometric fidelity. Ho…
LivePI: More Realistic Benchmarking of Agents Against Indirect Prompt Injection
AI agents such as OpenClaw are increasingly deployed in local workflows with access to external tools. This creates indirect prompt-injecti…
A Reproducible Log-Driven AutoML Framework for Interpretable Pipeline Optimization in Healthcare Risk Prediction
Accurate disease risk prediction is challenged by heterogeneous features, limited data, and class imbalance. This study presents yvsoucom-i…
Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability
Reinforcement learning under partial observability requires deciding what information to retain, yet most memory-based approaches do not ex…
Practical Anonymous Two-Party Gradient Boosting Decision Tree
Structured data is well handled by gradient-boosted decision trees (GBDT), which are usually trained on vertically partitioned features acr…
PatchWorld: Gradient-Free Optimization of Executable World Models
Text-agent environments are typically modeled as partially observable Markov decision processes (POMDPs), assuming that the simulator's lat…
The New Social Image: How AI Competency and AI Proactivity Influence Self- and Peer-Perceptions in the Workplace
Human-AI collaboration is considered the most promising way to incorporate AI in the workplace. What remains unexplored are the experientia…
Pre-Deployment Robustness Stress Testing for CT Segmentation Systems Using Clinically Motivated Multi-Corruption Augmentation
Deep learning-based CT segmentation systems often achieve high accuracy on clean benchmark images, but their performance may degrade under…
Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift
Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in ag…
Cosmos 3: Omnimodal World Models for Physical AI
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and actio…
Conditional Latent Diffusion Model with Fourier-based Motion Modelling for Virtual Population Synthesis
In-silico trials of medical devices require the generation of virtual populations of anatomies. In cardiovascular applications, virtual ana…
TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation
TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequ…
HAARES Half-Split Residual Basis Routing for Deep Transformers
Block-level residual routing makes learned residual aggregation practical by routing over block summaries, but each summary compresses an o…
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify…
UPLOTS: A Unified Pretrained Language Model for Constrained Time-series Generation
In time-series generation, existing approaches typically handcraft ortrain a separate model for each dataset, which hinders their scalabili…
Bag of Dims: Training-Free Mechanistic Interpretability via Dimension-Level Sign Patterns
We show the standard basis of transformer hidden states already provides a training-free, architecture-general feature basis. Individual di…
SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning
Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…
「シャドーAI」7割超の企業が対策追い付かず “会社が選んだAIだけ利用”はもう限界? ガートナー
会社が認めていないAIツール・サービスを従業員が業務で使う「シャドーAI」について、日本企業の73%が対策できていない――調査会社の米Gartnerは、このような調査結果を発表した。
NRI流“業務に最適なAIモデル”の選び方 「ベンチマークだけで優劣は決まらない」
“業務に最適なAIモデル”はどう選べば良いのか。野村総合研究所の北村雄騎氏に聞いた。
How to turn off AI in your Google Docs
Here's what you need to do to get those pesky "write with Gemini" pop-ups to go away.
「AIコーディング」がたった5年で急進化したワケ NTT「tsuzumi 2」開発者が分析
コーディングに長けた大規模言語モデル(LLM)が登場したのは2021年ごろだ。それから5年で、競技プログラミングの問題を解けるレベルにまで成長した。なぜAIはコーディングがこれほど得意になったのか──「Interop Tokyo 2026」(幕張メッセ)で、LLM「tsuzum…
「AIを使う学生」vs.「使わない学生」、エッセイが創造的なのはどっち? 米大学が2025年に実証実験
米ジョージタウン大学に所属する研究者らが国際学術誌Computers in Human Behavior: Artificial Humansに発表した論文「Homogenizing effect of large language models (LLMs) on creat…
かんぽ生命、AIで営業支援 “郵便局での一言”拾って保険提案へ 寸劇で分かる活用例
1700万人の顧客を抱えるかんぽ生命保険が、営業フローにAIエージェントを組み込んだ。商談準備に奔走する現場がどう変わるのか、デモンストレーションの様子を紹介する。
Anthropic、デザインツール「Claude Design」を強化 Codeとの双方向連携やCanvaなどへの出力をサポート
Anthropicは、デザイン制作ツール「Claude Design」のβ機能を大幅に強化した。複数のデザインシステムを取り込んでプロジェクト横断で維持できるようになったほか、「Claude Code」とのシームレスな双方向連携を実現。AdobeやCanvaをはじめとする外部ツ…
Roelof Botha joins SpaceX’s board of directors
The former Sequoia Capital leader is filling an "existing vacancy" on SpaceX's board, days after the company went public in the largest IPO…
After unveiling ridiculously expensive AR glasses, Snap’s stock takes a dive
Snap's long-awaited smart glasses debut hasn't exactly done wonders for the company's stock.
NEA’s Tiffany Luck says enterprises are still figuring out their AI ROI
Tokenmaxxing was the hottest trend in Silicon Valley earlier this year, with CEOs encouraging employees to push AI usage as far as it would…
World leaders want American AI. They just don’t want America to be able to turn it off.
French President Macron and Indian PM Modi raised alarms at the G7 summit that the U.S. could cut off access to American AI overnight — a f…
Anthropic becomes first AI startup to join the Frontier carbon removal coalition
Anthropic has joined the Frontier coalition, which received another $915M in pledges to fund carbon removal projects.
Social media’s next evolution: user-controlled algorithms
Social media feeds are becoming more customizable as platforms like Threads, Instagram, and TikTok introduce tools that let users directly…
NEA’s Tiffany Luck on AI IPOs, personal agents, and the ROI reckoning
Tokenmaxxing was the hottest trend in Silicon Valley earlier this year, with CEOs encouraging employees to push AI usage as far as it would…
World model maker Odyssey nabs $1.45B valuation backed by Amazon and other big names
World models are the next big thing in AI beyond LLMs and, with this round, Odyssey has cemented itself as one of the startups to watch.
Only 16 percent of Americans think AI will have a positive impact on society, a new study shows
Although Wall Street loves AI, every day Americans are significantly less optimistic about the industry, a new report from Pew Research sho…
Google bets on Gemini to reinvent the smart home speaker
Google is betting generative AI can breathe new life into the smart speaker. The company's new $99.99 Google Home Speaker replaces the rigi…
The slowtech revolution is here to kill your phone addiction and rescue your attention span
“People just really want to take back control of their time, their lives, their attention... They’re down for whatever helps them do that.”
Collecting robot training data is dirty, unglamorous work. Some AI labs are already paying XDOF to do it.
If physical AI is going to match the accomplishments of LLMs, there's a data problem that needs to be solved.
2026-06-17(343件)
Pramaana Labs raises $27M seed round from Khosla Ventures to bring formal verification to AI
Pramaana will focus on highly sensitive verticals like law, drug discovery, and tax preparation — where errors can be costly and reliabilit…
Canadian pension giant joins race to fund India’s AI-fueled data center boom
The Canadian pension giant will acquire an 8.2% stake in CtrlS, a tech giant that operates more than 15 data centers across India.
DeepL acquires Mixhalo for live-event audio streaming and translation
With this acquisition, DeepL is opening an office in San Francisco to expand its U.S. business.
Pinterest launches an experimental AI shopping app called ‘Ask Pinterest’
Pinterest has launched 'Ask Pinterest,' an experimental AI-powered shopping app that lets users seek recommendations and inspiration throug…
A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry
OpenAI and Molecule.one show how a near-autonomous AI chemist using GPT-5.4 improved a key drug-making reaction, advancing medicinal chemis…
Cursor、Gitホスティング「Origin」発表 SpaceXによる買収発表直後に
「GitHub」に対抗する狙いがありそうだ。
OpenAI創業者、巨大モデルのアップデート作業は「大きな苦痛だった」――月イチ更新を可能にした体制とデータの重要性
Databricksの年次イベントにOpenAI共同創業者のグレッグ・ブロックマン氏が登壇。AIモデル開発の難しさと、その土台となる「データ」の重要性を語った。
Beyond Parallel Sampling: Diverse Query Initialization for Agentic Search
Test-time scaling for agentic search typically increases depth (i.e., more turns and tokens per trajectory) or breadth (i.e., more parallel…
When Rules Learn: A Self-Evolving Agent for Legal Case Retrieval
Legal case retrieval remains challenging due to the complexity of legal language and the need for precise lexical alignment between queries…
SkillChain-Gym: A Benchmark for Reskilling-Aware Production-Inventory Control under Disruptions
Production planning increasingly has to treat workforce capability as a decision variable: certifications lapse when skills are not maintai…
Skill-Constrained Model Predictive Control for Resilient Manufacturing Supply Chains
In skill-constrained production-inventory systems, the qualified human capacity available tomorrow depends on training decisions made today…
Nothing from Something: Can a Language Model Discover 0?
AI systems based on artificial neural networks are being developed with aspirations of pushing the boundary of human mathematical knowledge…
Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty
Large language models can arrive at the same answer through reasoning paths that are unstable, contradictory, or difficult to rank consiste…
MemTrace: Probing What Final Accuracy Misses in Long-Term Memory
LLM agents increasingly maintain long-term memory of user facts across sessions. Yet such memory is usually evaluated by aggregating accura…
SpeechDx: A Multi-Task Benchmark for Clinical Speech AI
Speech offers a uniquely informative window into health by simultaneously engaging neurological, motor, respiratory, and vocal systems. Cur…
Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and Prototypes
Large language models have accelerated the transition from passive conversational assistants to autonomous agents that can understand goals…
Treatment Response Optimized Clinical Decision Support AI System via Digital Twin Simulation
Clinical decision support AI systems (CDSASs) must adapt to evolving patient conditions in real-time while adhering to strict safety constr…
Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
Large language models (LLMs) are becoming a major way for consumers to find products, but we do not yet understand how brands compete in th…
A Machine-Learned Comorbidity Index
Traditional comorbidity scores (e.g., Charlson and Elixhauser) are widely used for risk adjustment and patient stratification, but they hav…
MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors
Large language model agents are increasingly integrated into map services. Since map services are embedded in everyday-life scenarios rathe…
Dissecting model behavior through agent trajectories
AI agent performance is not just a modeling problem, it is fundamentally a systems problem. The advanced capabilities of models are realize…
Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation
Evaluating the decision-making capabilities of large language models (LLMs) is a growing research priority, yet existing benchmarks focus o…
LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline
Generative AI and large language models (LLMs) are increasingly applied to question generation and automated assessment. However, deploying…
SEAGym: An Evaluation Environment for Self-Evolving LLM Agents
Self-evolving LLM-based agents improve mainly by changing their agent harness: the structured execution layer around a base model, includin…
DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack
Evaluating a Physical AI stack spans operators that differ by more than three orders of magnitude -- from a single foundation-model decodin…
Surrogate Assisted Pedestrian Protection Design via a Foundation Model Orchestrated Workflow
AI-driven engineering workflows face particular challenges in crash safety design: unlike aerodynamics, crash events involve highly nonline…
Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning
Training-free verbal reinforcement learning enables LLM agents to learn from world feedback -- objective signals such as dynamic task outco…
Brick-DICL: Dynamic In-Context Learning for Automated Brick Schema Classification
Building Management Systems (BMS) are essential for optimizing energy efficiency and operational performance in modern buildings. However,…
FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness
Financial multimodal reasoning requires agents to coordinate numerical computation, retrieval, visual interpretation, and temporal groundin…
Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns
Large language model (LLM) web agents are usually deployed as tool callers: each turn, the model reads a fresh page observation and emits o…
From Brewing to Resolution: Tracing the Internal Lifecycle of Code Reasoning in LLMs
Standard accuracy metrics cannot explain why LLMs handle variable tracking but fail on semantically equivalent loops. We study an internal…
Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games
People make decisions differently in strategic interactions. Some update beliefs like a Bayesian; others exhibit biases like motivated reas…
FllumaOne: A Code-Native Multimodal CAD Dataset with Executable Programs and Kernel-Validated Feature Histories
Parametric computer-aided design records both final geometry and the ordered construction history that determines how a part can be edited.…
EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent
As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly i…
LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings
Recent vision-language models (VLMs) have shown promising progress in generating webpages from visual inputs, yet existing evaluations main…
Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs
Although reinforcement learning (RL) has expanded the cognitive boundaries of large language models (LLMs), it often remains vulnerable to…
DecoSearch: Complexity-Aware Routing and Plan-Level Repair for Text-to-SQL
Large Language Models (LLMs) have demonstrated remarkable capabilities in translating natural language to SQL, yet existing methods still f…
WallZero: Mastering the Game of WallGo with Strategic Analysis
WallGo is a recently introduced strategic board game popularized by the 2025 Netflix series The Devil's Plan. Although played on a small 7…
A homotopy-type-theoretic generalization of neurosymbolic inference
A wide range of neurosymbolic (NeSy) systems compute one functional: a belief-weighted sum of a logical quantity over a space of $\sigma$-s…
FlowRAG: Synergizing Explicit Reasoning via Frequency-Aware Multi-Granularity Graph Flow
Graph-based retrieval-augmented generation (GraphRAG) is effective for knowledge-intensive and multi-hop query tasks; however, many existin…
StepGuard: Guarding Web Navigation via Single-Step Calibration
Web navigation requires agents to follow natural language goals, interact with web pages, and produce accurate answers. While recent advanc…
Structural Preservation and the Logical Expressiveness of Graph Neural Networks
Bridges between graph neural networks (GNNs) and logical formalisms have been established by fixing architectural choices, such as the type…
MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning
Chain-of-Thought (CoT) reasoning has extended from purely linguistic domains to multimodal scenarios; however, existing approaches often tr…
Learn to Quantify Social Interaction with Constraints for Pedestrian Walking
Long-term human path forecasting in crowds is critical for autonomous moving platforms (like autonomous driving cars and social robots) to…
DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue
Language models increasingly serve as advisory systems in maintenance operations. To prevent hallucination, recent systems ground these mod…
PreAct: Computer-Using Agents that Get Faster on Repeated Tasks
Computer-using agents drive real software through the screen -- clicking and typing -- but they solve every task from scratch: asked to rep…
How Inference Compute Shapes Frontier LLM Evaluation
AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As…
Small Initialization Matters for Large Language Models
Large language models provide a tractable system for asking how intelligence itself emerges, rather than only how LLMs can be engineered. A…
MoCo-AIS: A Contrastive Learning Framework for Similarity Computation of Vessel Trajectories
Trajectory similarity is a fundamental task in analyzing mobility patterns, essential for applications such as route pattern extraction, mo…
STAR: SpatioTemporal Adaptive Reward Allocation for Text-to-Image RL Post-Training
Existing RL post-training methods for text-to-image generation usually convert the final-image reward into a single scalar advantage and ap…
LLM Consumer Behavior Theory: Foundations of a Novel Research Field
Large language models (LLMs) are increasingly deployed as autonomous agents that make consumption decisions on behalf of users. This shift…
LegalHalluLens: Typed Hallucination Auditing and Calibrated Multi-Agent Debate for Trustworthy Legal AI
AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors c…
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, AP…
PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
As Large Language Model based agents enter autonomous scientific research, their ability to resist pseudoscience becomes increasingly impor…
Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications
Recent advances in Large Language Models (LLMs) and multi-agent systems have driven the rise of Agentic AI, showing promise for medical rea…
A Unified Framework for Context-Aware and Relation-Aware Graph Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) has emerged as a paradigm for enhancing large language models (LLMs) with external knowledge, yet exis…
IsabeLLM: Automated Theorem Proving Applied to Formally Verifying Consensus
Advances in Artificial Intelligence (AI) have led AI for Theorem Proving to become a promising means of formally verifying computer systems…
Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding
Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in high-resolution screen…
First Proof Second Batch
To assess the ability of current AI systems to correctly solve research-level mathematics problems, we tested several AI systems on a set o…
Knowledge Reutilization in Meta-Reinforcement Learning
Meta-reinforcement learning enables fast adaptation by extracting shared structure from related tasks, but existing end-to-end methods ofte…
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users. Existing benchmar…
Memory as a Wasting Asset: Pricing Flash Endurance for Embodied Agents, and the Limits of Doing So
A robot's flash endurance is a non-renewable stock: every persisted write spends one of a few thousand program/erase cycles and never refil…
WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning
Language models are remarkably capable at medical question answering, in some cases surpassing the accuracy of general physicians. However,…
Learning Cardiac Electrophysiology Digital Twins Through Agentic Discovery of Hybrid Structure
Building personalized cardiac electrophysiology (EP) digital twins requires identifying the appropriate model structure for each patient, n…
DRFLOW: A Deep Research Benchmark for Personalized Workflow Prediction
Deep research (DR) systems are increasingly used for complex information-seeking tasks, but existing works mainly focus on generating repor…
The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data
As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of…
Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers
Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The…
EvolveNav: Proactive Preflection and Self-Evolving Memory for Zero-Shot Object Goal Navigation
Zero-Shot Object-Goal Navigation (ZS-OGN) requires embodied agents to explore and locate target objects without any prior training. To this…
Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs
Although Knowledge Editing provides an efficient mechanism for updating the knowledge of Multimodal Large Language Models (MLLMs), we find…
Towards Distributed Inference of LLMs on a P2P Network
Prefix caching can reduce LLM inference latency by reusing KV caches across requests with shared prompts, but cluster-scale reuse is challe…
PIVOT: Bridging Black-Scholes Implied-Volatility and Price Objectives via Differentiable J\"ackel Operator
Modern option-learning systems operate in two coordinates: price space, where markets quote and no-arbitrage constraints are most naturally…
KFTD: Koopman-Fourier Time-Differentiable Network for Continuous Ocean Spatiotemporal Forecasting
Accurate oceanic forecasting is critical for climate monitoring and disaster early warning. However, ocean spatiotemporal forecasting encou…
Extracting Semantics: LLM-Guided Automatic Population of Robot Ontology from URDF
While commonsense knowledge may suffice for virtual agents, embodied robots interacting with humans require grounded and semantically rich…
Surveying GenAI-based Automation in Printed Circuit Board Design and Test
Generative artificial intelligence (GenAI) is increasingly used for applications in the hardware and software domains. It purports to reduc…
CMIP-Forge: An Agentic System that Retrieves, Computes, and Self-Reviews Climate Science
The Coupled Model Intercomparison Project Phase 6 (CMIP6) has generated thousands of peer-reviewed publications documenting model configura…
Comprehensive pKa Data Augmentation from Limited Real Data through an Engineered Models-Quantum Framework
Proton dissociation constants (pKa) are critical for functional molecule discovery and molecular modeling. Building on iBonD, the largest e…
HRDX: A Large-Scale Vector HD-Map Dataset
Reliable autonomous driving requires vectorized HD maps that are geometrically accurate, semantically rich, and scalable to long-horizon dr…
The Price of Anarchy in Disaggregated Inference
Disaggregated inference architectures physically separate prefill and decode phases onto distinct GPU pools, creating competing "agents" th…
ParkingTransformer: LLM-Enhanced End-to-End Trajectory Planning for Autonomous Parking
End-to-end autonomous parking has emerged as a critical task within the realm of autonomous driving. However, existing methods suffer from…
ZIVARI-TLBO: A Zero-Cost Inter-Group Evaluated-Elite Relay Mechanism for Teaching-Learning-Based Optimization
ZIVARI-TLBO is a grouped Teaching-Learning-Based Optimization (TLBO) method that augments an existing population-state controller with a fi…
ANEForge: Python for direct computation on the Apple Neural Engine
ANEForge is a Python package that programs the Apple Neural Engine (ANE), the fixed-function neural accelerator on every recent Apple devic…
Software Delegation Contracts: Measuring Reviewability in AI Coding-Agent Work
AI coding agents increasingly accept assigned software tasks, modify repositories under bounded authority, and return work packages for rev…
Quantum Cinema: An Interactive Cinematic Exploration of Quantum Computing Hardware via Generative World Models
Quantum computing promises transformative advances across science and industry, yet the physical hardware that enables these computations r…
Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators
As large language models (LLMs) are increasingly deployed in latency- and cost-sensitive settings, inference efficiency has become a centra…
Models Take Notes at Prefill: KV Cache Can Be Editable and Composable
Prefix caching reuses prefill only across an exactly shared prefix, so one changed field invalidates the entire downstream cache. Yet overw…
Timestamp-Aware Spatio-Temporal Graph Contrastive Learning for Network Intrusion Detection
Given their effectiveness in modeling the relational structure among network traffic flows, graph neural networks (GNNs) have been widely a…
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios
AI agents are increasingly being adopted in enterprise and personal settings with access to emails, databases, documents, and other tools w…
Probing, Fusion, and Trustworthiness: A Systematic Evaluation of Foundation Model Representations for Multimodal Cancer Analysis
Foundation models (FMs) have emerged as powerful representation extractors for medical data, yet their generalizability to datasets under d…
MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs
Mixture-of-Experts Multimodal Large Language Models (MoE-MLLMs) offer remarkable performance but incur prohibitive GPU memory costs, making…
Graph neural networks at war: integrating cybersecurity and drone intelligence in the Israeli-Iranian conflict
Physical cyber systems have brought about new threats and challenges in detection and immediate response. This study examines how Graph Neu…
TrustErase: Auditable Instant Machine Unlearning with Passport-Embedded Representations
The demand for privacy-compliant AI has amplified the need for machine unlearning; yet, existing retraining or distillation-based methods r…
LineageMark: Multi-user White-box Watermarking for Contribution Tracing in Model Derivation Chains
In open large language model (LLM) ecosystems, models are frequently adapted across multiple domains and applications, forming multi-stage…
Vibrato Expression Control for Singing Voice Conversion with Improving Independent Control
Singing style is a crucial aspect of a natural and expressive singing voice. Singers utilize singing styles to convey the feeling or emotio…
Agentic Discovery of Non-Canonical Antimicrobial Peptides with AMPGAN v3
Antimicrobial resistance causes to over a million deaths annually. Antimicrobial peptides (AMPs) are a promising solution, but generative A…
PromptMN: Pseudo Prompting Language
Prompting has become the primary interface between humans and generative AI, yet many natural language prompts remain fragile: roles, goals…
Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference
Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, i…
Cluster-Aware Dual-Level Test Specification Generation for Large-Scale Automotive Software Requirements
Generating test specifications that satisfy Automotive SPICE SWE.6 requirements becomes increasingly challenging and time-consuming as proj…
PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation
Standard on-policy distillation (OPD) for large language models estimates the reverse-KL objective using student-sampled tokens, yielding a…
Trust-Aware Multi-Agent Traceability: Confidence-Calibrated Knowledge Graphs for Consistent Software Artifact Management
Multi-agent AI systems are increasingly used to automate software engineering tasks including requirements analysis, architecture design, t…
Rift: A Conflict Signature for Deception in Language Models
A model that lies while knowing the truth is the central case ELK cannot handle with behavioral evaluation alone. We ask whether such decep…
Physics-Informed Attention Mechanism and Generalization Capability of Deep Learning-Based Grain Growth Evolution Prediction
Machine Learning (ML) models for grain growth prediction are typically trained on idealized synthetic data, yet practical applications requ…
MLLP-VRAIN UPV system for the IWSLT 2026 Simultaneous Speech Translation task
This work describes the participation of the MLLP-VRAIN research group in the shared task of the IWSLT 2026 Simultaneous Speech Translation…
Pulling The REINS: Training-Free Safety Alignment of Video Diffusion Models via Representation Steering
Open-weight video diffusion models can generate photorealistic unsafe content, from violence to misinformation, yet existing defenses eithe…
ARVO: Atlas of Reproducible Vulnerabilities for Open-Source Software
Achieving reproducibility, quantity, and diversity in vulnerability datasets has long been viewed as an inherent three-way trade-off, where…
From Democracies to Autocracies: How AI Systems Enable Authoritarianism by Design
AI-enabled authoritarianism is not confined to autocracies. In this paper, we provide greater transparency by investigating and mapping the…
Transformer-Based Warm-Starting for Feasible and Optimal Terminal Approach to Tumbling Objects with Space Manipulators
Real-time trajectory generation for on-orbit robotic servicing is challenging due to the nonlinear coupling between spacecraft bus motion,…
Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation
Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation,…
Counterfactual Optimization of Baseball Pitch Sequences and Estimation of Its Impact on Season-Level Statistics
Although pitch sequencing is a central topic in baseball analytics, previous studies have primarily focused on optimizing the final pitch w…
Do Large Language Models Always Tell The Same Stories?
Recent advances in large language models (LLMs) have enabled the generation of high-quality prose, yet the question of whether these models…
Translating the Untranslatable: An Operationalizable Ontology for Untranslatability
Untranslatability, cases where meaning cannot be directly preserved across languages, is well-studied in linguistics but underexplored in N…
DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models
Autonomous driving has shifted towards end-to-end policy learning, where reliable, interpretable policy evaluation is a fundamental challen…
Implicit vs. Explicit Prompting Strategies for LVLMs in Referential Communication
Two recent studies (Jones et al. (2026); Zeng et al. (2026)) reach apparently contradictory conclusions about whether LVLMs can coordinate…
MeiBRD: Meta-Learning Intraoperative Biomechanical Residual Deformation
Accurate intraoperative liver registration is challenging due to substantial soft-tissue deformation yet sparse intraoperative measurements…
Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation
Agentic artificial intelligence systems introduce a new class of model risk. Unlike traditional predictive models, autonomous agents contin…
TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations
End-to-end autonomous driving has achieved state-of-the-art performance on benchmarks and real-world deployments. Its standard training rec…
Visuals Lie, Consistency Speaks: Disentangling Spatial Attention from Reliability in Vision-Language Models
Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical.…
NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama
Long-form serialized audio drama, with arcs that run for 200 to 800 episodes, is a major creative medium and a setting where frontier large…
SoK: AI-Augmented Binary Reversing
Binary reversing is fundamental to software understanding, vulnerability discovery, malware investigation, and firmware auditing. However,…
The Discrete-Log Clock: How a Transformer Learns Modular Multiplication
When small transformers grok modular multiplication, prior work reports that the learned embedding has a "dense" Fourier spectrum requiring…
Bridging Spatial And Frequency Views For Disaster Assessment: Benefits And Limitations
Rapid assessment of building damage from satellite imagery is essential for effective disaster response and recovery. While most deep learn…
Graph Neural Networks for Semi-Supervised Image Classification with Multi-Feature Aggregation
Feature extraction involves the identification and extraction of salient characteristics or patterns, including edges, textures, shapes, an…
Discrete Autoregressive Transformer for Generative Mechanism Synthesis
Planar path synthesis requires mechanisms whose coupler curves match a prescribed trajectory; the mapping from curve to linkage is inherent…
Enhancing Pathological VLMs with Cross-scale Reasoning
Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnifi…
L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification
Multilingual speaker verification remains challenging because language-dependent acoustic variability causes speaker identity to become ent…
Feynman Kac Reweighted Schr\"odinger Bridge Matching for Surface-Based Tau PET Harmonization
Tau PET imaging is central to tracking Alzheimer's disease progression, but systematic differences between scanners, protocols, and radiotr…
Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos
Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges. Fi…
Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure
Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and…
MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation
While Multimodal Retrieval-Augmented Generation (M-RAG) enhances Large Vision-Language Models, it remains highly susceptible to cross-modal…
AUTOGATE: Automated Clock Gating via Toggling-Aware LLM-based RTL Rewriting
Fine-grain clock gating (FGCG) is among the most effective techniques for reducing dynamic power, yet current FGCG optimization flows remai…
AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows
Large language models (LLMs) are increasingly considered for use in clinical consultation tasks, yet most medical evaluations remain static…
Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing
As LLMs acquire stronger reasoning capabilities, deceptive behavior becomes an increasingly serious safety concern. Existing deception moni…
Online LLM Selection via Constrained Bandits with Time-Varying Demand
Large Language Models (LLMs) are increasingly deployed in edge-cloud inference systems to handle diverse user tasks with heterogeneous accu…
MagicSim: A Unified Infrastructure for Executable Embodied Interaction
Robot learning and embodied agents now require simulation to serve as a shared execution substrate linking control, skills, and planning, n…
Geometry-Aware Post-Hoc Uncertainty Quantification in Operator Learning
Neural operators provide fast surrogates for PDEs but their deterministic predictions limit their use in tasks requiring uncertainty quanti…
Unlocking LLM Code Correction with Iterative Feedback Loops
Large Language Models have shown remarkable capabilities in code generation. However, most existing evaluations focus only on single-attemp…
FoundCause: Causal Discovery with Latent Confounders from Observational Data
Causal discovery from observational data remains challenging due to the need to recover directed structure and latent confounding without i…
Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery
Production LLM assistants route user requests to growing libraries of specialized tools, but how does routing accuracy degrade as the catal…
OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation
Generative world models for autonomous driving face two unresolved tensions: heterogeneous control injection, where free-form language, HD-…
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over dept…
Offline Preference-Based Trajectory Evaluation
Offline evaluation of agentic systems often collapses trajectories to terminal success, discarding information about partial progress and i…
Reversal Q-Learning
Iterative generative modeling techniques, such as flow matching, provide powerful tools to model complex behaviors for effective offline re…
An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts
Banks simultaneously face signature-based fraud (card-not-present attacks, account takeover, ATM cloning) and behavioural financial crime (…
Geometric Consistency Protocol for Foundation Model Features in Multi-View Satellite Imagery
Standardized evaluation protocols are indispensable for robust benchmarking in remote sensing, particularly as foundation features are incr…
LLM Features Can Hurt GNNs: Concatenation Interference on Homophilous Graph Benchmarks
Adding LLM-generated node features to graph neural networks (GNNs) is widely reported to improve accuracy on standard benchmarks. We docume…
Visored: A Controlled-Natural-Language Prover for LLM-Generated Mathematics
We present a dependent-type-based prover designed around the way LLMs (and humans) tend to write mathematics, complementing existing system…
Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations
Several studies have examined the use of large language models (LLMs) for title-abstract screening in systematic reviews (SRs), reporting m…
SkillMoV: Mixture-of-View Routing with Prototype-Conditioned Gating for Unified Multi-View Proficiency Estimation
Estimating human proficiency from video is a key challenge for automated skill assessment, with applications in sports coaching, music peda…
Divide, Deliberate, Decide: A Multi-Agent Framework for Fine-Grained Egocentric Action Recognition
Fine-grained action recognition in egocentric video is challenging for Vision-Language Models (VLMs): actions often differ only in small vi…
Bounding Box Label Propagation for Re-Annotation of Document Layout Analysis Datasets
Datasets in practical document processing scenarios typically grow over time, and their class annotations undergo continuous refinement. Th…
SketchXplain: Intuitive Visual Explanations of Image Classifiers with Sketches
Saliency map visualizations explain image-based AI predictions by pointing to regions, but these are often unintuitive and semantically unc…
A Risk Decomposition Framework for Pre-Hoc Fine-Tuning Prediction
The high cost of fine-tuning LLMs poses a significant economic barrier; pre-hoc performance prediction offers a critical solution to substa…
TuneAhead: Predicting Fine-tuning Performance Before Full Training Begins
Fine-tuning large language models (LLMs) is compute-intensive and error-prone: model performance depends sensitively on data quality and hy…
Temporal Preference Optimization for Unsupervised Retrieval
Unsupervised dense retrievers offer scalability by learning semantic similarity from unlabeled documents via contrastive learning, but they…
FacProcessTwin: An LLM-Based System for Process Twin Development
Process twins provide real-time representations of entire production processes. By capturing how process steps interact, rather than monito…
Handling Feature Heterogeneity with Learnable Graph Patches
In recent years, the rapid development of foundation models and graph pre-training technologies has spurred increasing interest in construc…
ASTEROID: A Spatiotemporal Information Transformer for Forecasting Multi-Step Time Series of Molecular Dynamics
Molecular dynamics (MD) simulation is computationally demanding, particularly for large-scale systems requiring long-term analysis. Accurat…
See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL
Multimodal large language models (MLLMs) integrate strong text reasoning with visual inputs, yet their responses can be inconsistent with t…
SuCo: Sufficiency-guided Continuous Adaptive Reasoning
Despite remarkable performance on complex tasks, Large Reasoning Models (LRMs) often generate excessively long Chain-of-Thoughts (CoT), inf…
SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology
Characterising the tumour microenvironment (TME) from routine H&E-stained histology images requires simultaneous cell segmentation, feature…
Confusion-Aware Transfer Teacher Curriculum Learning Framework: Disentangling Scoring and Pacing Effects
Curriculum learning couples two design choices, how samples are scored by difficulty and how harder samples are paced into training, making…
Vision-language models for chest radiography do not always need the image
Medical vision-language models report strong chest radiograph accuracy, and this is increasingly read as evidence that they use the image.…
Structured Adversarial Camouflage via Voronoi Diagrams
Pixel-wise adversarial patches are computationally heavy and often visually detectable, limiting utility in security-critical systems. We p…
ED3R: Energy-Aware Distributed Disaster Detection Enabled by Cooperative Robotic Agents
Robotics are expected to support environmental monitoring and natural disaster management, where decisions must be made under uncertainty,…
Symplectic Transversality and Endpoint Green Estimates for Finite-Horizon Pontryagin Systems
We study horizon-uniform local branches of finite-horizon discrete-time Pontryagin boundary value systems after smooth control elimination.…
Talking to Your Data: Exploring Embodied Conversation as an Interface for Personal Health Reflection
Personal health data from wearables are typically presented through dashboards of charts and summary statistics, requiring users to activel…
A Neuromorphic Trigger for Efficient Audio Event Detection
Efficient processing of continuous audio streams remains a key challenge for real-time and resource-constrained systems. This paper introdu…
MIVE: A Minimalist Integer Vector Engine for Softmax LayerNorm and RMSNorm Acceleration
The rapid growth of Large Language Models (LLMs) has intensified the need for specialized hardware accelerators that can satisfy stringent…
LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams
Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously…
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
Coding agents have become a major mode of software engineering, but the benchmarks we use to compare them were designed in a pre-agent era:…
No-Free-Fairness: Fundamental Limits and Trade-offs in Learning Systems
In this paper, we establish a set of theoretical impossibility results, termed the No-Free-Fairness theorems, that identify three fundament…
Conservation Laws for Modern Neural Architectures
Understanding gradient descent dynamics is key to explaining the success of over-parameterized models, where implicit bias manifests throug…
A Framework for Evaluating Agentic Skills at Scale
Agent skills -- structured, reusable knowledge artifacts that augment LLM agent capabilities -- have been rapidly adopted in industry, yet…
Human-in-the-Loop Atlas-Based 3D Asset Segmentation for Interactive Content Workflows
Segmenting 3D assets into meaningful regions remains challenging, especially when segmentation criteria are application-dependent and requi…
When Multiple Scripts Matter: Evaluating ASR in Clinical Settings
Automatic speech recognition (ASR) in non-English clinical settings is challenged by multiscript variability, where the same term may appea…
Functional Equivalence in Attention: A Comprehensive Study with Applications to Linear Mode Connectivity
Neural network parameter spaces are inherently non-injective, as distinct parameter configurations can realize identical functions through…
Perceptual compensation for tonal context in self-supervised speech models
This study examines the extent to which the wav2vec2.0 architecture exhibits evidence of compensation for phonological context. We conducte…
High-Fidelity 3D Geometric Reconstruction of Pelvic Organs from MRI: A Hybrid Deep Learning and Iterative Optimization Approach
Patient-specific 3D reconstruction of pelvic organ geometry from MRI is important for pelvic floor modeling and downstream patient-specific…
A Quantitative Analysis of Multimodal Biomarkers in Alzheimer's Disease
Despite increasing adoption of multimodal approaches in Alzheimer's Disease (AD) research -- aimed at integrating molecular, structural, cl…
AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor
Large language models (LLMs) outperform earlier architectures on generative inference and long-context tasks, but their large size introduc…
AI Adoption Across a Multinational Workforce: Sociotechnical Conditions for GenAI Acceptance in Human Resources
Generative AI (GenAI) deployment in the workplace is accelerating rapidly. Nevertheless, questions of who adopts, who benefits, and who is…
Dimensionality Controls When Modularity Helps in Continual Learning
Compositional learning systems must balance plasticity, the ability to acquire new knowledge, with stability, the preservation of previousl…
Non-negative Elastic Net Decoding for Information Retrieval
Dense retrieval has become the dominant paradigm in information retrieval, in which each document is scored against a query by the inner pr…
Trustworthy Self-Composable Big-Data-as-a-Service: An LLM-Orchestrated Multi-Agent Framework for Automated Data Engineering, AutoML, MLOps Deployment, and Drift-Aware Lifecycle Optimization
Big-Data-as-a-Service (BDaaS) platforms require re liable automation across data ingestion, cleaning, feature engi neering, model developme…
PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space
Current Vision-Language-Action (VLA) models face a trade-off between efficient action generation and explicit deliberation. Directly decodi…
KANLib -- An Modular, Extensible and Fast Kolmogorov-Arnold Network Implementation
Kolmogorov-Arnold Networks (KANs) have recently emerged as a promising alternative to traditional multilayer perceptrons by replacing linea…
Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model
Visual information helps resolve ambiguity in coreference resolution, leading to notable performance gains. However, existing Multi-modal C…
SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs
Sparse Mixture-of-Experts (MoE) architectures enable scaling LLM parameters under a fixed inference budget by activating only a small subse…
Robustness of Similarity-based Positional Encoding Under Rotations: Theoretical Analysis and Experimental Validation
Positional encoding is a fundamental component of Transformer architectures, as it injects information about the spatial or sequential arra…
A Neuro-Symbolic Approach to Strategy Synthesis for Strategic Logics
Reasoning about what agents can achieve through strategic interaction is a core challenge in Multi-Agent Systems (MAS). Logics for strategi…
SegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation
Self-supervised DINO models provide strong transferable visual representations, yet applying them directly to image segmentation remains ch…
Recover Semantics First, Generate Better: Improved Latent Modeling for 3D MRI Reconstruction and Cross-Contrast Synthesis
Multi-contrast magnetic resonance imaging (MRI) provides complementary information for clinical diagnosis. However, acquiring all MRI seque…
Multiple cyclicity and Wavelet Decomposition with Channel Correlation for Long-term Time Series Forecasting
Cyclicity and trend are important components of time series data and many studies based on cyclicity and trend have achieved good results i…
A T-API-Compliant ReAct Agentic Loop for Optical Networks: Generic vs. Domain-Specific Tool Abstractions
Optical networks need intent-driven, closed-loop agentic management, a key enabler for higher autonomy levels. We present the first T-API-c…
C2FL: Clustered Continual Federated Learning under Spatial and Temporal Drift
Collective Adaptive Systems (CAS) increasingly rely on machine learning to let each node learn from locally sensed data, aligning its behav…
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases latency and KV-cache me…
Catastrophic Forgetting is Low-Rank: A Function-Space Theory for Continual Adaptation
Catastrophic forgetting in continual adaptation is usually studied through parameter drift, replay, or distillation, but these views do not…
When English Isn't the Best Teacher: Source Language Effects in Cross-Lingual In-Context Learning
Cross-lingual transfer in multilingual NLP has been widely explored in supervised fine-tuning contexts, where factors like data availabilit…
When AI Says "I have been in similar situations": Synthetic Lived Experience in Peer-Like Caregiver Support
Caregivers often turn to online communities for informational and emotional support. In these spaces, peer supporters frequently draw on pe…
Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
Large language models (LLMs) are widely used to fulfill users' information needs; users ask LLMs about the weather, pose educational questi…
When LLMs Analyze Scars: From Images to Clinically-Meaningful Features
Medical image classification faces a fundamental dilemma: while deep learning models achieve remarkable performance at scale, real-world cl…
Volterra Generative Models
Score-based diffusion models typically use Brownian perturbations, which provide tractable reverse-time dynamics but impose memoryless nois…
EAGG: Embodiment-Aligned Grasp Generation via Geometry-Aware Graph Conditioning
Cross-end-effector grasp generation seeks a unified model that generalizes across objects and across embodiments ranging from parallel grip…
S4oP: Operator-level Pruning of Structured State Space Models for Resource-Constrained Devices
Structured State Space Models (SSMs), including the S4 and S4D architectures, have recently emerged as powerful alternatives to attention-b…
Querying an astronomical database using large language models: the ALeRCE text-to-SQL system
We develop a text-to-SQL (structured query language) system based on large language models (LLMs) using in-context learning and apply it to…
Learning Fair Pareto-Optimal Policies in Multi-Objective Reinforcement Learning
Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optima…
Ternary Mamba: Grouped Quantization-Aware Training of W1.58A16 State Space Models
State Space Models (SSMs) such as Mamba-2 offer linear-time inference but their memory footprint limits edge deployment. Prior ternary SSM…
Structural Role Injection in Handlebars-Templated LLM Prompts: Triple-Brace Interpolation, Delimiter Family, and the Limits of HTML Auto-Escaping
Large language model applications build prompts from templates, and Handlebars is a widely used templating engine and the default prompt-te…
Embedded Machine Learning for Microcontroller-Class Edge Devices: Data, Feature, Evaluation, and Deployment Pipelines
Embedded machine learning moves inference from cloud services to resource-constrained devices that must acquire data, preprocess signals, r…
Towards Understanding and Measuring COGNITIVE ATROPHY in LLM Behaviour
Recent incidents involving LLMs used for mental-health support reveal a critical evaluation gap: surface-level safety scores do not capture…
Descriptor: Certus Caliber Classification Gunshot Dataset (C3GD)
In this work, we introduce the Certus Caliber Classification Gunshot Dataset (C3GD), a publicly accessible data set developed for the analy…
ReAge3D: Re-Aging 3D Faces with View Consistency
We present a novel framework for realistic and controllable 3D face re-aging which produces highly detailed, identity-preserving results. E…
The Measurement Gap in the Automation of EU Law: Benchmarking Doctrinal Legal Reasoning under the EU AI Act
Large language models now produce legal text of at least median quality, yet no existing benchmark can evaluate whether they perform doctri…
All Smoke, No Alarm: Oracle Signals in Agent-Authored Test Code
Software practitioners increasingly use AI coding agents that generate test code alongside production code in open source pull requests (PR…
IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction
Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable laws or occur in area…
Kolmogorov Regression for Robust Diffusion Policies
Finite-dimensional (FD) diffusion policies exhibit temporal drift owing to discretization artifacts that degrade long-horizon performance (…
A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models
We evaluate the adversarial robustness of two frontier large language models (LLMs) developed by Anthropic, Fable 5 and Opus 4.8, against f…
RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills
The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to alleviate global disparities…
Looped World Models
Current world models face a fundamental tension: faithful long-horizon simulation demands deep computation, but deeper models are expensive…
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents
With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cyber-defense agents train…
ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues
Reproducing research results from papers and released code is central to scientific progress. Existing works have introduced benchmarks to…
Visual Verification Enables Inference-time Steering and Autonomous Policy Improvement
Robots deployed in the real world should learn from their experience and improve over time. This requires a mechanism of practicing and lea…
TRACE: Learning to Compute on Circuit Graphs
Learning to compute, the ability to model the functional behavior of a circuit graph, is a fundamental challenge for graph representation l…
Beyond the Sampled Token: Preserving Candidate Support in RLVR
We revisit exploration collapse in reinforcement learning with verifiable rewards (RLVR), from the perspective of the \emph{candidate distr…
Branch-and-Browse: Efficient and Controllable Web Exploration with Tree-Structured Reasoning and Action Memory
Autonomous web agents powered by large language models (LLMs) show strong potential for performing goal-oriented tasks such as information…
CausalT5k: Diagnosing Refusal and Failure Modes in Trustworthy Causal Reasoning Across Causal Rungs
Large language models increasingly produce fluent causal explanations, yet they often fail in ways aggregate accuracy cannot diagnose: conf…
Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices
As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often…
OmniSapiens: A Foundation Model for Social Behavior Processing via Heterogeneity-Aware Relative Policy Optimization
Socially intelligent AI systems must reason across diverse human behavioral tasks and generalize to new social contexts. However, behaviora…
In-Context Environments Induce Evaluation-Awareness in Language Models
Humans often become more self-aware under threat, yet can lose self-awareness when absorbed in a task; we hypothesize that language models…
Adaptive Domain Models: Bayesian Evolution, Warm Rotation, and Principled Training for Geometric and Neuromorphic AI
Prevailing AI training assumes reverse-mode automatic differentiation over IEEE-754 arithmetic. The memory overhead of training relative to…
Riemann-Bench: A Benchmark for Moonshot Mathematics
Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficien…
Know Thy Reasoner: Not All Language Models Explore Alike
Compute scaling for LLM reasoning trades off exploring solution approaches (\emph{breadth}) against refining promising ones (\emph{depth}),…
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond
As AI systems move from generating text to accomplishing goals through sustained interaction, the ability to model environment dynamics bec…
Mental Health AI Safety Claims Must Preserve Temporal Evidence
The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoin…
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether t…
MapAgent: An Industrial-Grade Agentic Framework for City-scale Lane-level Map Generation
Lane-level maps are critical infrastructure for autonomous driving and lane-level navigation, yet constructing and maintaining standardized…
Lean4Agent: Formal Modeling and Verification for Agent Workflow and Trajectory
Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence.…
LATTEArena: An Evaluation Framework for LLM-powered Tabular Feature Engineering (Extended Version)
Feature engineering remains a cornerstone of tabular data analysis, and Large Language Models (LLMs) have emerged as a promising paradigm f…
Belief-Space Control for Personalized Cancer Treatment via Active Inference
Cancer treatment is at the core a sequential decision-making problem with partial observability, latent patient heterogeneity, and explicit…
Learning What to Remember: Observability-Safe Memory Retention via Constrained Optimization for Long-Horizon Language Agents
Long-horizon language agents accumulate observations, reasoning traces, and retrieved facts that exceed their context windows, making memor…
Reducing the Complexity of Deep Learning Models for EEG Analysis on Wearable Devices
Wearable healthcare devices are the fastest-growing Internet of Things (IoT) sector. Many automated healthcare services rely on two crucial…
MOSAIC: Modality-Specific Adaptation for Incremental Continual Learning in Parkinson's Disease Gait Assessment
Gait-based Parkinson's disease assessment increasingly relies on heterogeneous sensors, but clinical systems rarely collect all modalities…
SSIL: Self-Supervised Imitation Learning for End-to-End Driving
In autonomous driving, the end-to-end (E2E) driving approach that predicts vehicle control signals directly from sensor data is rapidly gai…
Towards Leveraging AutoML for Sustainable Deep Learning: A Multi-Objective HPO Approach on Deep Shift Neural Networks
Deep Learning (DL) has advanced various fields by extracting complex patterns from large datasets. However, the computational demands of DL…
E2Vec: Feature Embedding with Temporal Information for Analyzing Student Actions in E-Book Systems
Digital textbook (e-book) systems record student interactions with textbooks as a sequence of events called EventStream data. In the past,…
LLM-Powered Multi-Agent System for Automated Crypto Portfolio Management
Cryptocurrency portfolio management requires the fusion of heterogeneous multi-modal signals, including structured price and on-chain time…
Mordal: Automated Pretrained Model Selection for Vision Language Models
Incorporating multiple modalities into large language models (LLMs) is a powerful way to enhance their understanding of non-textual data, e…
Top-Theta Attention: Sparsifying Transformers by Compensated Thresholding
We present Top-Theta (Top-$\theta$) Attention, a training-free method for sparsifying transformer attention during inference. Our key insig…
Ensemble RL through Classifier Models: Enhancing Risk-Return Trade-offs in Trading Strategies
This paper presents a comprehensive study on the use of ensemble Reinforcement Learning (RL) models in financial trading strategies, levera…
MedicalAgentsBench for Complex Medical Reasoning: Comparing Internalized Reasoning Models versus Externalized Agent-based Frameworks
Complex medical reasoning requires integrating heterogeneous clinical evidence across multiple inference steps. Large language models (LLMs…
Gaussian DP for Reporting Differential Privacy Guarantees in Machine Learning
Current practices for reporting differential privacy (DP) guarantees for machine learning (ML) algorithms such as DP-SGD provide an incompl…
Detecting and Mitigating DDoS Attacks with AI: A Survey
Distributed Denial of Service attacks represent an active cybersecurity research problem. Recent research shifted from static rule-based de…
Algorithmic Prompt Generation for Diverse Human-like Teaming and Communication with Large Language Models
Understanding how humans collaborate and communicate in teams is essential for improving human-agent teaming and AI-assisted decision-makin…
EmoFSM: A Finite State Machine for Emotional Support Conversation
Emotional support conversation (ESC) aims to alleviate people's emotional distress through effective conversations. Although large language…
RLRC: Reinforcement Learning-based Recovery for Compressed Vision-Language-Action Models
Vision-Language-Action models (VLA) have demonstrated remarkable capabilities and strong potential in complex robotic manipulation. However…
SPATIA: Multimodal Generation and Prediction of Spatial Cell Phenotypes
Understanding how cellular morphology, gene expression, and spatial context jointly shape tissue function is a central challenge in biology…
Critique of World Model: A Generative Latent Prediction Architecture for World Modeling
World Model, the algorithmic simulator of the real-world environment which biological agents experience and act upon, has been an emerging…
A Gradient-based Causal Discovery Framework with Applications to Complex Industrial Processes
With the advancement of deep learning technologies, various neural network-based Granger causality models have been proposed. Although thes…
AnalogFed: Privacy-Preserving Discovery of Analog Circuits at Scale with Federated Generative AI
Recent advances in generative AI (GenAI) have shown transformative potential for modern hardware design. However, existing GenAI-driven app…
LLM-Aided Joint Secrecy Precoding and Trajectory for RSMA-Based Heterogeneous UAV Networks
This paper investigates secure communications in rate-splitting multiple access (RSMA) enabled heterogeneous UAV networks, where multiple U…
Detail++: Training-Free Detail Enhancer for Text-to-Image Diffusion Models
Recent advances in text-to-image (T2I) generation have led to impressive visual results. However, these models still face significant chall…
Moving Out: Physically-grounded Human-AI Collaboration
The ability to adapt to physical actions and constraints in an environment is crucial for embodied agents (e.g., robots) to effectively col…
Blueprint First, Model Second: A Framework for Deterministic LLM Workflow
While powerful, the inherent non-determinism of large language model (LLM) agents limits their application in structured operational enviro…
RooseBERT: A New Deal For Political Language Modelling
The increasing amount of political debates and politics-related discussions calls for the definition of novel computational methods to auto…
Explicit Context-Driven Neural Acoustic Modeling for High-Fidelity RIR Generation
Realistic sound simulation plays a critical role in many applications. A key element in sound simulation is the room impulse response (RIR)…
Regression Language Models for Code
We study code-to-metric regression: predicting numeric outcomes of code executions, a challenging task due to the open-ended nature of prog…
OmniRetarget: Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction
A dominant paradigm for teaching humanoid robots complex skills is to retarget human motions as kinematic references to train reinforcement…
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
Code-capable large language model (LLM) agents are embedded in software engineering workflows where they can read, write, and execute code,…
Adversarial Attacks Leverage Interference Between Features in Superposition
Why do adversarial examples exist, and why do they transfer between models? Existing explanations appeal to high-dimensional geometry, non-…
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
The convergence of LLM-powered research assistants and AI-based peer review systems creates a critical vulnerability: fully automated publi…
Enhanced Evolutionary Multi-Objective Deep Reinforcement Learning for Reliable and Efficient Wireless Rechargeable Sensor Networks
Despite rapid advancements in sensor networks, conventional battery-powered sensor networks suffer from limited operational lifespans and f…
Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demons…
A geometric and deep learning reproducible pipeline for monitoring floating anthropogenic debris in urban rivers using in situ cameras
The proliferation of floating anthropogenic debris in rivers has emerged as a pressing environmental concern, exerting a detrimental influe…
EngTrace: A Symbolic Benchmark for Verifiable Process Supervision of Engineering Reasoning
Large Language Models (LLMs) are increasingly entering specialized, safety-critical engineering workflows governed by strict quantitative s…
Retrofitters, pragmatists and activists: Public interest litigation for accountable automated decision-making
This paper examines the role of public interest litigation in promoting accountability for AI and automated decision-making (ADM) in Austra…
First, do NOHARM: towards clinically safe large language models
Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain po…
Prototype-Based Semantic Consistency Alignment for Domain Adaptive Retrieval
Domain adaptive retrieval aims to transfer knowledge from a labeled source domain to an unlabeled target domain, enabling effective retriev…
A Multifaceted Analysis of Social Biases in Large Language Models
Large language models (LLMs) have rapidly become indispensable tools for acquiring information and supporting human decision-making. Howeve…
Vulcan: Instance-specialized, Verifiable Systems Heuristics Through LLM-driven Search
Systems resource management tasks rely primarily on hand-designed heuristics. However, growing hardware heterogeneity and workload diversit…
Co-PLNet: A Collaborative Point-Line Network for Prompt-Guided Wireframe Parsing
Wireframe parsing aims to recover line segments and their junctions to form a structured geometric representation useful for downstream tas…
m2sv: A Scalable Benchmark for Map-to-Street-View Spatial Reasoning
Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that…
LVLMs and Humans Ground Differently in Referential Communication
For generative AI agents to partner effectively with human users, the ability to accurately predict human intent is critical. But this abil…
Learning-Infused Formal Reasoning: From Contract Synthesis to Artifact Reuse and Formal Semantics
This paper articulates a long-term research vision for formal methods at the intersection with artificial intelligence, outlining multiple…
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
In this work, we aim to develop effective data synthesis techniques that autonomously synthesize multimodal training data for enhancing MLL…
PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning
We develop a continual learning method for pretrained models that \emph{requires no access to old-task data}, addressing a practical barrie…
Optimism Stabilizes Thompson Sampling for Adaptive Inference
Thompson sampling (TS) is widely used for stochastic multi-armed bandits, yet its inferential properties under adaptive data collection are…
Brep2Shape: Boundary and Shape Representation Alignment via Self-Supervised Transformers
Boundary representation (B-rep) is the industry standard for computer-aided design (CAD). While deep learning shows promise in processing B…
From Noise to Order: Learning to Rank via Denoising Diffusion
Learning-to-rank (LTR) methods have traditionally been limited to discriminative machine learning approaches that model the probability of…
SkillJect: Effectively Automating Skill-Based Prompt Injection for Skill-Enabled Agents
Agent skills extend LLM agents with task-specific instructions, executable scripts, and auxiliary resources, improving reusability but crea…
GOT-JEPA: Generic Object Tracking with Model Adaptation and Occlusion Handling using Joint-Embedding Predictive Architecture
The human visual system tracks objects by integrating current observations with previously observed information, adapting to target and sce…
Position: Modular Memory is the Key to Continual Learning Agents
Foundation models have transformed machine learning through large-scale pretraining and increased test-time compute. Despite surpassing hum…
Phys4D: Fine-Grained Physics-Consistent 4D Modeling from Video Diffusion
Recent video diffusion models have achieved impressive capabilities as large-scale generative world models. However, these models often str…
CogGen: Cognitive-Load-Inspired Fully Unsupervised Deep Generative Modeling for Compressively Sampled MRI Reconstruction
Fully unsupervised deep generative modeling (FU-DGM) offers significant potential for compressively sampled magnetic resonance imaging (CS-…
Guidelines for the Annotation and Visualization of Legal Argumentation Structures in Chinese Judicial Decisions
This Guideline presents a systematic and operationalizable annotation framework for representing legal argumentation structures in judicial…
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting future world states from video observa…
Rethinking Multimodal Fusion for Time Series: Text Modalities Need Constrained Fusion
Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS)…
Decidable By Construction: Design-Time Verification for Trustworthy AI
A prevailing assumption in machine learning is that model correctness must be enforced after the fact. We observe that the properties deter…
findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding
Syllable-level units offer compact and linguistically meaningful representations for spoken language modeling and unsupervised word discove…
Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones
Vision backbone networks play a central role in modern computer vision. Enhancing their efficiency directly benefits a wide range of downst…
Evaluating Interactive 2D Visualization as a Sample Selection Strategy for Biomedical Time-Series Data Annotation
Reliable machine-learning models in biomedical settings depend on accurate labels, yet annotating biomedical time-series data remains chall…
DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning
Drivers' visual attention provides critical cues for anticipating latent hazards and directly shapes decision-making and control maneuvers,…
Membership Inference Attacks against Large Audio Language Models
We present the first systematic Membership Inference Attack (MIA) evaluation of LALMs. Using Multi-modal Blind Baselines based on textual,…
Combating Data Laundering in LLM Training
Post-hoc unauthorized-training data detection for large language models (LLMs) typically assumes a query-with-originals regime: rights hold…
From Paper to Program: Knowledge Externalization for AI-Assisted Quantum Many-Body Code Generation
Large language models can write scientific code, but direct paper-to-program translation remains fragile when correctness depends on tacit…
Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit
Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendo…
Do We Still Need Humans in the Loop? Comparing Human and LLM Annotation in Active Learning for Hostility Detection
Instruction-tuned LLMs can annotate thousands of instances at low cost. This raises two questions for active learning (AL): can LLM labels…
Curiosity-Critic: Cumulative Prediction Error Improvement as a Tractable Intrinsic Reward for World Model Training
Local prediction-error-based curiosity rewards focus on the current transition without considering the world model's cumulative prediction…
DPRM: A Plug-in Doob h transform-induced Token-Ordering Module for Diffusion Language Models
Diffusion language models generate without a fixed left-to-right order, leaving token ordering as a central algorithmic choice. Existing sy…
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning
Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. However, BC lacks a self-guided mechanism for online i…
SP-GCRL: Influence Maximization on Incomplete Social Graphs
Influence maximization (IM) in real platforms is challenged by incomplete, noisy social graphs and non-stationary diffusion dynamics. We pr…
Learning to Decide with AI Assistance under Human-Alignment
It is widely agreed that when AI models assist decision-makers in high-stakes domains by predicting an outcome of interest, they should com…
Large Language Models for Agentic NetOps and AIOps: Architectures, Evaluation, and Safety
Large language models are increasingly being used to support network operations (NetOps) and artificial intelligence for IT operations (AIO…
Rethinking Cross-Layer Information Routing in Diffusion Transformers
Diffusion Transformers (DiTs) have become a de facto backbone of modern visual generation, and nearly every major axis of their design -- t…
Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks
We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detecti…
Any2Any: Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking
Whole-body tracking (WBT) models have become a key foundation for humanoid robots, enabling them to imitate diverse motions with high fidel…
Remote sensing data imputation using deep learning for multispectral imagery
Remote sensing techniques have been increasingly utilised in aquatic applications in recent years. A common challenge in using optical sate…
CyberEvolver: Structured Self-Evolution for Cybersecurity Agents On the Fly
LLM-based agents are increasingly used for cybersecurity tasks, but most existing systems rely on fixed, human-designed scaffolds that stru…
Sustainable Metal-Organic Framework Water Harvesters in the Artificial Intelligence Era
Metal-organic frameworks (MOFs) are excellent candidates for water harvesting due to their tunable pore environments, which can be precisel…
Temporal Motif-aware Graph Test-time Adaptation for OOD Blockchain Anomaly Detection
Ever-evolving transaction patterns have significantly hindered anomaly detection on emerging cryptocurrency blockchains due to the vast num…
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across div…
Constitutional On-Policy Safe Distillation
On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged informat…
LLMCodec: Adapting Video Codecs for Efficient Weight Compression of Large Language Models
The rapid development of large language models(LLMs) has led to remarkable advances in natural language processing. However, the increasing…
Fast Speech Foundation Model Distillation Using Interleaved Stacking
Distilling a large speech foundation model (SFM) into an efficient student model has been successfully applied to low-resource environments…
Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation
Remaining Useful Life (RUL) prediction is essential for industrial predictive maintenance, yet many learning-based approaches rely on exten…
CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents
Screenshot-based mobile GUI agents can operate ordinary smartphone apps through the same visual interface as a human user, but this capabil…
日立、OpenAIとの連携を本格化 「Codex」でレガシーシステム刷新、サイバー防衛も
Codexの解析力と日立のシステム開発ノウハウを組み合わせ、既存コードから上流仕様を可視化し、新システムへの移行テストまでの一連のプロセスについて、AIを活用したアプローチの確立を目指す。
「ポケカ対戦AIエージェント」開発コンテスト開始 「不完全情報ゲーム」をどう制するか
チェスや将棋と異なり、相手の手札が見えない「不完全情報ゲーム」にAIがどこまで対応できるかが試される。
SpaceX、AIコーディング「Cursor」を9.6兆円で買収 「近く大幅な改善」へ
Corsorは公式Xで、「近く大幅な改善が行われる予定だ」と述べた。
村田製作所、Synopsysの電磁界/熱解析ツールを介したシミュレーションモデル提供
村田製作所は、Synopsysが提供するシミュレーションツールを介したシミュレーションモデルの提供を開始した。3D電磁界解析ツール「Ansys HFSS」と熱解析ツール「Ansys Icepak」を対象とする。ユーザーはシミュレーションツールから村田製作所のWebサイトにアクセ…
「生成AIは大手なら安心」とは限らない? 突然の提供停止が招くリスク顕在化
米政府の輸出管理指令によるAnthropicの最新AIモデル提供停止を受け、生成AIが事前の通知なしに突然使えなくなるリスクが顕在化した。Forresterは、単一のAIモデル依存の危うさを指摘し、ポータビリティ確保をはじめとする4つの対策を推奨している。
財務諸表だけでは勝てない ブルームバーグ日本トップが語る「非構造化データ」の重要性
デフレからインフレへ経済の潮目が激変した日本市場。もはや過去の数値(財務諸表)を眺めるだけのデータ経営では勝てない。情報の洪水におぼれず、気象や音声などの「非構造化データ」をいかに素早く選別し、リアルタイムの決断に生かすか。金融情報インフラを支えるブルームバーグの日本トップに、…
Anthropic’s latest feud with the Trump admin may actually help it, sales data suggests
Anthropic's popularity with business users is growing so well that the latest beef with the government might actually boost it, data from R…
セルフ給油、実はスタッフが手動で許可していた!? コスモ石油の「AI監視」は消えゆくガソリンスタンドを救うか
従来のセルフ式ガソリンスタンドでは、利用者が給油ノズルを手にした後も、スタッフが安全を確認した上で給油を許可している。この監視業務をAIで支援する取り組みが動き出した。コスモ石油マーケティングとELEMENTSは、AIが給油許可を判断する監視システムを共同開発。背景には、人手不…
月間売上1億円超、“推しAI”アプリ「Zeta」がオタク女子わしづかみ ただし危うさも
自分が作ったシチュエーションで“推し”と会話できるAIチャットアプリ「Zeta」(ゼタ)が人気だ。App StoreやGoogle Playのエンターテインメントランキングでも連日上位にランクインしており、各ストアページによればダウンロード数は130万回を突破。1月には月間の売…
Unlocking UK house-building with AI-accelerated planning
UK government partners with Google DeepMind to build a new AI-powered prototype aimed at faster housing decisions.
東芝の組み込み向け量子インスパイアード技術が進化、高速化と安定性を両立
東芝は、刻々と状況が変化する現実環境において、組み合わせ最適化問題を高速かつ安定して解くことができる「量子インスパイアード最適化フレームワーク」を開発した。
SpaceX valuation balloons to $2.6T, briefly passes Amazon
SpaceX's valuation has increased by $1 trillion since its shares started trading on Friday.
生成AI×自動運転で注目のTesla・Waymo・NVIDIA 各社が目指す「フィジカルAI」は何が違うのか
日本政府が戦略的強化分野に掲げる「フィジカルAI」――その社会実装の最前線の一つが自動運転システムだ。熾烈な開発競争が繰り広げられている中、生成AIの進化は各社の競争にどのような変化をもたらしているのか。Tesla、Waymo、NVIDIAの最新動向を整理する。
【Pythonで学ぶデータ分析】母平均と母標準偏差をベイズ推定する ~ シュークリームの重さは100gと異なるか?
ある喫茶店のシュークリームの重さを例に、ベイズ統計で「平均」や「ばらつき」をどう推定するのか、さらに基準値と違いがあるかどうかをどう確かめるのかを解説します。『社会人1年生から学ぶ、やさしいデータ分析』ベイズ統計編の第3回です。
Android 17 launches with new multitasking tools as Google expands Gemini features
Google has released Android 17 and Wear OS 7, introducing new multitasking features, parental controls, security tools, and smartwatch upgr…
Sixty percent of US consumers say ‘AI’ in brand messaging is a turnoff, survey finds
WordPress VIP’s latest survey suggests consumers are wary of AI-generated answers even as companies increasingly view AI search as an impor…
SpaceX is public: Everything you need to know post-IPO
TechCrunch has followed SpaceX's start, struggles, and successes from the early days. And we're here for what happens next too. This packag…
DOJ claims xAI’s unpermitted gas turbines are a matter of ‘national, economic, and energy security’
The Justice department says the Pentagon needs xAI to keep using its unpermitted gas turbines.
Plaud says its software business topped $100M in ARR after shipping over 2M AI notetakers
Plaud is trying to make a mark in a crowded market full of AI-powered meeting notetakers.
2026-06-16(674件)
Robinhood’s note on 10% layoffs shows blaming AI isn’t cutting it
Unlike many of his tech industry peers who have cut thousands of jobs citing the need to restructure to make the most of AI, Robinhood's CE…
Probably raises $9M to build a more reliable kind of AI
Probably wants to prevent hallucinations and factual errors from reaching users, and achieve accuracy on par with deterministic systems.
SpaceX to acquire Cursor for $60B in stock, days after blockbuster IPO
The deal is supposed to help SpaceX's struggling AI division. The company told IPO investors it sees a $26 trillion addressable market in A…
ChatGPT’s market share slips below 50% for first time
The chatbot still remains the most popular AI assistant worldwide with over 1.1 billion monthly users, followed by Gemini with 662 million…
OpenAIのサム・アルトマンCEO、来日中止 都内イベント登壇予定を変更
米OpenAIのサム・アルトマンCEOが来日するイベントが予定されていたが、来日が中止となった。
OpenAIの高度AIでソフトバンクの脆弱性を1万件発見 孫正義氏「大変な危機」 日本の重要インフラ企業へ診断サービス提供
ソフトバンクグループは6月16日、米OpenAIの高度なAI技術を活用したサイバーセキュリティ対策サービス「Patching as a Service」を発表した。OpenAIのサイバーセキュリティに特化したAI「GPT-5.5 Cyber」などの技術とソフトバンクの運用ノウハ…
Malaysia’s AI agent-powered messaging app Respond.io raises $62.5M, eyes acquisitions
Respond.io, one of Malaysia's startups to watch, uses AI agents to handle high volumes of customer inquiries and charges per convo, not per…
AI時代の“シゴデキ”会社員はどこに座る? データ活用が変えた理想のオフィス
イトーキが本社オフィスを4年ぶりに刷新した。従業員の能力発揮度と位置情報の分析で「成果が出る席」を特定。各種センサーやAIも活用し、分析結果を反映したフロアレイアウトを実現したという。
Claude「Fable 5」が3日で停止 Anthropicが主張する“米国政府の誤解”の正体
米国政府の指令に従い、Fable 5のサービス提供を停止したAnthropic。同社は指令について「政府の誤解に基づくもの」と主張しているが、その誤解とは具体的にどのようなものなのか。
A Definition of Good Explanations and the Challenges Explaining LLM Outputs
How to define a good explanation is a long-standing philosophical debate which has found recent renewed interest in the context of AI outpu…
Dr-DCI: Scaling Direct Corpus Interaction via Dynamic Workspace Expansion
Agentic search over large corpora relies on retriever-mediated interfaces (e.g., BM25 or ColBERT) for scalable candidate discovery. While e…
Relational Structural Causal Models
An artificial intelligence must have a model of its environment that is causal, supporting reasoning about interventions and counterfactual…
Trust Between AI Agents: Measuring Formation, Breakage, and Recovery, with Implications for Governing Multi-Agent Systems
As language-model agents increasingly work in teams, each agent must decide how much to trust its teammates. Yet we lack a standard way to…
PrologMCP: A Standardized Prolog Tool Interface for LLM Agents
Frontier reasoning-tuned language models still fail on deductive tasks at depth, and the cost of improved performance through extended inte…
Semantics-Enhanced Retrieval-Augmented Time Series Forecasting
Time series forecasting models often benefit from historical patterns. Inspired by Retrieval-Augmented Generation (RAG), recent research ex…
AI Engram: In Search of Memory Traces in Artificial Intelligence
Memory formation is fundamental to intelligence, yet whether deep neural networks preserve identifiable memory traces analogous to biologic…
Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these j…
OSGuard: A Benchmark for Safety in Computer-Use Agents
Computer-use agents are increasingly evaluated by whether they complete realistic desktop and web tasks. However, task success alone can mi…
Fusion is not one-size-fits-all: Cross-Modal Representation Alignment for Time-to-Event Modeling
Accurate time-to-event (TTE) prediction from multimodal clinical data remains challenging due to modality imbalance and distribution shift.…
Risk-Aware LLM Agents for Geospatial Data Retrieval: Design and Preliminary Adversarial Evaluation
We present an LLM-driven framework for retrieving remote sensing data from cloud-based geospatial catalogues using natural language queries…
Cognitive Debt: AI as Intellectual Leverage and the Dynamics of Systemic Fragility
We develop a formal theory of cognitive debt: the stock of unverified reasoning obligations that accumulates when individuals use AI as a s…
VGPT-RSI for RH-Adjacent Formal Progress: Boundary Certificates, Verified Finite Lagarias Inequalities, and Explicit Failure Localization
The Riemann Hypothesis remains one of the central unsolved problems in mathematics. Rather than claiming proof, we investigate whether a ve…
Towards Verifiable Agentic Data Science: Solving Irregular TSQA Via Tool-Grounded Reasoning
Time series data in real-world deployments is overwhelmingly irregular. Observations are asynchronous, missing values are informative rathe…
CONCORD: Asynchronous Sparse Aggregation for Device-Cloud RAG under Document Isolation
Retrieval-augmented generation (RAG) has emerged as a pivotal technique for improving language models by incorporating external knowledge a…
CogGuard: Cognitive and Operational Profiling for Proactive Warning in Edge Intelligent Services
Proactive warning is an important capability for edge intelligent services, where the system predicts whether a subject will successfully c…
Attribute Inference from Interactive Targeted Ads
Targeted advertising systems can pair audiences selected by advertisers with ad units that expose visible user actions. When an interaction…
Visual-Seeker: Towards Visual-Native Multimodal Agentic Search via Active Visual Reasoning
Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factu…
Mask-Proof: An LLM-based Automated Data Curation Pipeline on Mathematical Proofs
Large language models (LLMs) are increasingly capable of mathematical problem solving and can even assist with research-level proofs, yet w…
Feature Attribution in Directed Acyclic Graphs Using Edge Intervention
Shapley value-based feature attribution methods face challenges in scenarios involving complex feature interactions and causal relationship…
A Formal Framework for Declarative Agentic AI in Business Process Analysis
Agentic AI opens new opportunities for automating Business Process (BP), enabling autonomous decision-making and dynamic adaptation. Howeve…
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation b…
Forced Deferral: Manipulating Routing Decisions in Multimodal LLM Cascades
While multimodal large language models (MLLMs) have shown strong visual reasoning abilities, serving a large model for every query is compu…
ChatPlanner: A Large Language Model Framework for Personalized Public Transit Routing
Personalized public transit routing in public transit systems remains challenging due to the difficulty of capturing and integrating divers…
APEX: Adaptive Principle EXtraction A Three-Layer Self-Evolution Framework for Production AI Agents
Self-improvement in AI agents has emerged as a key research frontier: systems that modify their own prompts, workflows, and decision rules…
S1-DeepResearch: Beyond Search, Toward Real-World Long-Horizon Research Agents
Deep research agents aim to solve complex knowledge-intensive tasks through long-horizon planning, evidence gathering, reasoning, and repor…
Reward Hacking in Language Model Agents: Revisiting AI Safety Gridworlds
Reward hacking, where AI systems exploit misspecified objectives to achieve high reward without satisfying intended goals, remains a centra…
Hierarchical Modeling of ICD Codes in EHR Foundation Models
Electronic health record foundation models typically treat ICD diagnosis codes as flat tokens, overlooking the clinically meaningful hierar…
Who Drifted: the System or the Judge? Anytime-Valid Attribution in LLM Evaluation Pipelines
Continuous evaluation of LLM products relies on a strong LLM judge treated as ground truth: a cheap monitor scores every interaction and a…
Towards End-to-End Automation of AI Research
The automation of science is a long-standing ambition in the field of AI. While the community has made significant progress in automating i…
Synthetic Counteradaptation: A Principle of Human-AI Co-evolution
In this paper, we introduce the concept of synthetic counteradaptation, a process where human and AI systems co-evolve by adapting to each…
Toward Vibe Medicine: A Self-Evolving Multi-Agent Framework for Clinical Decision Support
In recent years, the advances of large language models and autonomous agents have revolutionized the healthcare field, facilitating diagnos…
Frame-Conditioned Moral Computation in LLaMA 3.1-8B-Instruct: A Mechanistic Interpretability Audit of Ethical Reasoning
Behavioral audits of Large Language Models on moral prompts measure what the model says, not the internal computation producing it. We use…
ToolMenuBench: Benchmarking Tool-Menu Filtering Strategies for Reliable and Efficient LLM Agents
Tool-augmented large language model agents increasingly operate over large tool libraries, but existing evaluations often focus on whether…
Minimal Oversight: Uncertainty-Aware Governance for Delegated AI Systems
AI systems increasingly delegate decisions to specialized models, evaluators, tools, and supervisory controllers. The central AI problem is…
QoS-Aware Token Scheduling and Private Data Valuation for Multi-Modal Agentic Networks
In agentic systems, human-generated data records anchor the value of AI services. Yet cloud compute pipelines centralize processing on remo…
Do we have the knowledge we need? Rethinking human-AI decision-making in corporations
Organizational knowledge is fragmented across a variety of software systems, tacit expertise, and manual documents that have traditionally…
Large Language Models as Optimizers: A Survey of Direct vs. Tool-Augmented Approaches and Their Performance Frontiers
Large Language Models (LLMs) are increasingly involved in complex mathematical optimization, even if the pragmatic user who triggers them i…
Your Agent Has a Genome: Sequence-Level Behavioral Analysis and Runtime Governance of LLM-Powered Autonomous Agents
We propose Base Sequence Analysis, a framework that encodes the runtime behavior of LLM-powered autonomous agents into compact symbolic seq…
Agentic Retrieval and Reinforcement Learned Equation Chains: A Controlled Generation Framework for Complex and Novel Physics Word Problems
Generating high-quality Physics Word Problems (PWPs) that are novel, complex, and solvable remains a challenging and underexplored problem…
Integrating Reasoning and Generalization in Text-to-SQL via Self-Enhanced Fine-Tuning
Text-to-SQL aims to translate natural language questions into executable SQL queries over structured databases, enabling non-expert users t…
NeuroSymbolic AI for Legal AI-TRISM: Trustworthy, Reliable, Interpretable, Safe Models
Large Language Models (LLMs) have transformed natural language processing, but their lack of interpretable reasoning and tendency to halluc…
Towards Next-Generation Healthcare: A Survey of Medical Embodied AI for Perception, Decision-Making, and Action
Foundation models have demonstrated impressive performance in enhancing healthcare efficiency across a wide range of medical applications.…
Advanced Machine Learning and Deep Learning Techniques for Enhanced Cattle Identification and Detection: A Comprehensive Review
The need for effective cattle identification technology is now more acutely felt than ever in maintaining biosecurity, food safety, and sup…
Overcoming the Impedance Mismatch: A Theoretical Roadmap for Fusing Foundation Models and Knowledge Graphs
Modern artificial intelligence remains fundamentally divided between the continuous, probabilistic spaces of Foundation Models and the disc…
Where Did It Go Wrong? Process-Level Evaluation of Web Agents with Semantic State Tracking
Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process informati…
Multi-agent Framework for Time-Sensitive Complementary Collaboration in Minecraft
We present TickingCollabBench, a Minecraft-based multi-agent benchmark for a novel class of time-sensitive complementary collaboration task…
Recurrent Reasoning on Symbolic Puzzles with Sequence Models
Large language models often appear strong on symbolic and algorithmic tasks, yet this apparent strength can hide brittle behaviour when pro…
Do LLMs Reliably Identify Correct Information Units in Aphasic Discourse?
Correct Information Units (CIUs) are central to discourse assessment in aphasia because they quantify communicative informativeness rather…
Artificial Intelligence Index Report 2026
Welcome to the ninth edition of the AI Index report. As AI continues to advance rapidly, the question becomes whether the systems built aro…
AI-Driven Framework for Adaptive Water Network Management with Proof-of-Concept Implementation: Addressing Non-Revenue Water in Jordan
Jordan faces severe water scarcity with 50\% of water produced is lost to leakage, theft and metering issues also known as non-revenue wate…
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual gro…
Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments
A central pedagogical value evaluated in AI tutor benchmarks is scaffolding: guiding students through graduated steps toward a solution. Al…
Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference
Multimodal large language models (MLLMs) have demonstrated strong capabilities in vision-language understanding and natural-language respon…
Unassigned Agents in Compilation-based Multi-agent Path Finding
Compilation-based techniques represent an important stream of solvers for multi-agent path finding (MAPF) due to their modularity and adapt…
TrustedARI: Towards Trust-Native Agentic Routing Infrastructure for Agentic AI
AI agents increasingly access external models, tools, and services through Agentic Routing Infrastructure (ARI) to manage the overhead of h…
An Integrated System for Real-Time Student Assessment and Career Guidance Using Neural Networks in Computing Disciplines
Many undergraduate students in Computer Science (CS) and Software Engineering (SWE) struggle to identify suitable career paths, particularl…
AIChilles: Automatically Uncovering Hidden Weaknesses in AI-Evolved Systems
The computer systems community has recently seen growing interest in AI-driven system evolution, where AI agents iteratively rewrite system…
Heteroskedastic Signals in Budgeted LLM Verification: Structural Heterogeneity Limits Optimization Gains
Large language model (LLM) systems increasingly use uncertainty signals to allocate limited computation across verification, test-time scal…
RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decis…
STRIDE: Strategic Trajectory Reasoning via Discriminative Estimation for Verifiable Reinforcement Learning
Reinforcement Learning with Verifiable Rewards (RLVR) has become an effective post-training paradigm for improving the reasoning abilities…
LLM-as-Code Agentic Programming for Agent Harness
Every major LLM agent framework gives the LLM the role of orchestrator; the model decides what to do next, when to call tools, and when to…
UrbanWell: Benchmarking Multimodal Large Language Models for Spatio-Temporal Urban Wellbeing Analytics
Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant chal…
Agentic Framework for Deep Learning workload migration via In-Context Learning
Translating deep learning models from PyTorch's flexible, object-oriented design to JAX's functional, stateless setup is usually a manual a…
SciText2Eq: Assessing LLMs for Explainable Equation Generation for Scientific Creativity
This work investigates the ability of large language models (LLMs) to generate mathematical equations from scientific texts. Prior work fac…
Auditing Reward Hackability in Code RL Training Environments
We measure the rate at which code RL environments accept incorrect solutions as correct. On a 49-task sample of SWE-bench Verified, 28.5% o…
Mind-Studio: Executable World Models with Lookahead Evaluation for Partially Observable Games
World-model synthesis aims to turn interaction experience into an internal model of environment dynamics. Existing symbolic approaches ofte…
Rhythm of the Deep: A Computational-Linguistic Test of Duality of Patterning in Sperm Whale Codas
Human language has often been described as combining structure at two levels: lower-level units combine into larger units, which then combi…
RecourseBench: A Modular Framework for Reproducible Algorithmic Recourse Evaluation
Algorithmic recourse methods provide counterfactual explanations that inform individuals of the actions required to overturn an unfavorable…
Know Your Limits : On the Faithfulness of LLMs as Solvers and Autoformalizers in Legal Reasoning
Large Language Models (LLMs) achieve strong performance on reasoning tasks, but whether this reflects faithful logical inference or heurist…
Thinking with Visual Grounding
Visual thinking should not only sound right; it should show its evidence. While recent vision-language models (VLMs) can produce natural-la…
VibeThinker-3B: Exploring the Frontier of Verifiable Reasoning in Small Language Models
This technical report introduces VibeThinker-3B, a compact dense model with 3B parameters developed to investigate how far verifiable reaso…
LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis
Most medical AI systems improve by scaling additional machinery: more fine-tuning data, more agents, and/or larger retrieval databases. In…
The Quality-Utility Paradox: Why High-Reward Data Impairs Small Model Mathematical Reasoning
Knowledge distillation from powerful reasoning models is widely used to improve Small Language Models (SLMs) on mathematical reasoning, oft…
AI Pluralism and the Worlds It Misses
AI pluralism is often framed as a problem of representing diverse values, preferences, users, or outputs. This paper argues that this frami…
TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting
High-quality time series forecasting is pivotal for real-world decision-making. However, traditional point-wise metrics often fail to revea…
PAL-Bench: Evidence-Grounded Profile Reconstruction from Longitudinal Personal Albums
Longitudinal personal albums are weak-schema multimodal databases: noisy perceptual records whose key facts require joins across faces, tex…
Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronge…
Sensor-Conditioned Representation Learning via Scene-Relevant Observation Quotients
Learned representations in intelligent sensing systems are often evaluated by reconstruction fidelity or downstream prediction accuracy, bu…
Latent Thought Flow: Efficient Latent Reasoning in Large Language Models
Large Language Models (LLMs) increasingly rely on intermediate reasoning, yet explicit Chain-of-Thought (CoT) suffers from a linguistic spa…
SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data
As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no longer governed by a single universal…
State-Grounded Multi-Agent Synthetic Data Generation for Tool-Augmented LLMs
Training tool-augmented LLM agents requires large corpora of multi-turn, tool-grounded conversational data that is expensive to annotate, p…
Architectural Wisdom: A Framework for Governing Optimization in AI Systems
Modern AI systems exhibit structural failures that capability scaling alone does not reliably fix: they optimize under-specified objectives…
AdaSTORM: Scaling LLM Reasoning on Dynamic Graphs via Adaptive Spatio-Temporal Multi-Agent Collaboration
Large Language Models (LLMs) demonstrate remarkable potential in dynamic graph reasoning, but suffer from a scaling bottleneck: current mod…
Exploiting Search in Symbolic Numeric Planning with Patterns
In this paper, we present a procedure for numeric planning based on Symbolic Pattern Planning (SPP). Given a numeric planning problem $\Pi$…
Phase-Aware Guidance Injection for Recurrent MAPPO in Assembly-Line Disruption Recovery
Disruption recovery in industrial assembly lines requires timely decisions under machine faults, worker absence, and emergency orders. Exis…
Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules
Predictive modeling for clinical tabular data is central to clinical decision support and therefore requires not only strong predictive per…
Whose hotel does the AI recommend? An algorithm audit of reputation signals in LLM-assisted hotel selection
Travelers increasingly ask large language model (LLM) assistants which hotel to book, making these systems gatekeepers of property visibili…
Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents
LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite…
Posterior Twins: Distributional Behavioral Simulation for Enterprise Decisions
Enterprise behavioral simulation requires more than producing a plausible response. Many decisions depend on the shape of a population unde…
When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting
AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or tran…
Tensor-Coord: Algebraic Decomposition of Joint Plan Tensors for Conflict-Free Multi-Agent LLM Planning
Large language models (LLMs) remain limited in multi-agent planning because independently generated plans can create coordination failures…
Steering Emotional Dynamics for Art Therapy: Controllable Narrative Script Generation through Hierarchically Guided LLM Agents
Art therapy plays a vital role in emotional healing, in which narrative creation acts as the primary vehicle for emotional expression. Give…
Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap Balancing
Model merging has become a practical post-training strategy for building a single multi-task large language model (LLM) by combining multip…
Model Graph Inductive Learning for Knowledge Graph Completion
Link prediction in knowledge graphs fundamentally depends on the quality of learned embeddings for entities and relations. However, most ex…
Kairos: A Native World Model Stack for Physical AI
World models are transitioning from passive visual generators to foundational, operational infrastructure for Physical AI: they must native…
The Faithfulness Gap: Certifying Semantic Equivalence Between Natural-Language and Formal Mathematical Statements
Autoformalization, translating natural-language mathematics into formal proof assistants, is bottlenecked not by translation fluency but by…
ROSA-RL: Uncertainty-Aware Roundabout Optimized Speed Advisory with Reinforcement Learning
Roundabouts challenge automated driving in mixed traffic, as heterogeneous and non-deterministic human behavior, unknown driving intentions…
TNODEV: Toolbox for Neural ODE Verification
Neural ordinary differential equations (neural ODE) have started to appear in safety critical settings such as continuous-time controllers…
ARB4WM: An Adversarial Robustness Benchmark for World Models in Continuous Control
World models are widely used in robotic and agentic engineering control systems due to their ability to learn latent dynamics for planning…
CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies
As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is becoming increasingly…
MR-GVNO: A Geometry-Aware Variational Physics-Informed Neural Operator for Mindlin-Reissner Plates on Irregular Domains
Plate and shell structures are widely used in engineering, making rapid response prediction under varying geometries, materials, and loads…
The Integrator Advantage: Controlled Agentic AI for Small and Medium-Sized Companies
Agentic AI marks a new phase of enterprise automation. Unlike traditional automation or conversational AI, agentic systems can interpret go…
From Affect Prediction to Affect Forecasting: Evidence for Distinct Information Sources in Longitudinal Text
Modeling dimensional affect in longitudinal text requires distinguishing current affect estimation from future affective change forecasting…
User as Code: Executable Memory for Personalized Agents
A personalized AI agent needs a user memory: a persistent model of who the user is, built across many conversations and consulted on each n…
Medical world models: representing medical states, modelling clinical dynamics and guiding intervention policies
Medical diagnosis and treatment are dynamic processes in which patient states evolve over time and clinical interventions alter future outc…
AgentFairBench: Do LLM Agents Discriminate When They Act?
Large language model (LLM) agents increasingly take actions (screening applicants, recommending credit, triaging patients), yet fairness fo…
A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions
Policy gradient algorithms for language models optimize the same objective $J(\theta) = \mathbb{E}*{\tau \sim p*\theta(\tau)}[R(\tau)]$, wh…
Skill-to-LoRA: From Using Skills to Learning Behaviors for Token-Efficient LLM Agents
Agent skills are commonly distributed as SKILL.md files: human-readable procedural documents that describe workflows, tools, resources, and…
OpenClaw-Skill: Collective Skill Tree Search for Agentic Large Language Models
Equipping Large Language Model (LLM) agents with effective skills is crucial for solving complex tasks in real-world systems like OpenClaw.…
LabOSBench: Benchmarking Computer Use Agents for Scientific Instrument Control
Current computer-use benchmarks primarily focus on software operation tasks in virtualized systems, whereas scientific instrumentation scen…
Adaptive and Explicit safe: Triggering Latent Safety Awareness in Large Reasoning Models
While Large Reasoning Models (LRMs) excel at complex tasks, they remain highly vulnerable to sophisticated jailbreaks and direct harmful qu…
Scaling LLM Reasoning from Minimal Labels: A Semi-Supervised Framework with a Lightweight Verifier
For the development of Large language models (LLMs), recent approaches to generating pseudo intermediate reasoning have shown remarkable pr…
GIST-CMTF: Goal-State Inference for Causal Minimal Tool Filtering in LLM Agents
Tool-augmented LLM agents rely on runtime filtering to decide which tools should be visible at each step. Causal Minimal Tool Filtering (CM…
Symbolic Informalization: Fluent, Productive, Multilingual
Symbolic informalization enables a reliable conversion of formal mathematics to natural language. It has the potential to make machine-chec…
Greed Is Learned: Visible Incentives as Reward-Hacking Triggers
Deployed agents increasingly act with their reward proxy in view, such as a balance, score, or KPI dashboard. We show that reinforcement le…
MA-SBI: Misspecification-Aware Simulation-Based Inference via Side-Channel Guidance
Simulation-based inference (SBI) of latent parameters is often hindered by simulator misspecification, the mismatch between simulated and r…
RAID: Semantic Graph Diffusion for True Cold-Start and Cross-Lingual Forecasting
Time-series foundation models show strong transfer performance when given a non-empty history window. However, true cold-start scenarios, w…
A Causal Model of Theory of Mind in Conflict for Artificial Intelligence
Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and inference, is widely ass…
The embrace of open science: An analysis of a decade of AI research and 56 800 conference papers
The reproducibility crisis has directed the AI research community toward improving documentation practices. Several studies have identified…
Consensus-based Agentic Large Language Model Framework for Harmonized Tariff Schedule Code Classification
Accurate Harmonized Tariff Schedule (HTS) code classification is essential for customs clearance, duty assessment, trade statistics, and re…
When in Doubt, Plan It Out: Committed Small Language Model Deliberation for Reactive Reinforcement Learning
Reinforcement Learning (RL) policies often degrade in unfamiliar environments because they lack explicit deliberation. We propose Plan, Ali…
Bayesian Inference and Decision Audits for Public Archives of Frontier AI Evaluations
Public AI evaluations are often read as terminal leaderboards, yet the underlying evidence is a selective time series shaped by reporting r…
Phishing Email Detection Using Large Language Models
Email phishing is one of the most prevalent and globally consequential vectors of cyber intrusion. As systems increasingly deploy Large Lan…
Integrating Multi-Label Classification and Generative AI for Scalable Analysis of User Feedback
In highly competitive software markets, user experience (UX) evaluation is crucial for ensuring software quality and fostering long-term pr…
Honeypot Protocol
Trusted monitoring, the standard defense in AI control, is vulnerable to adaptive attacks, collusion, and strategic attack selection. All o…
Evaluation of Alternative-Based Information Systems for Deliberative Polling using an Agentic Simulator
Deliberative polling promises to improve collective decision-making by exposing shareholders to a broad range of arguments before they vote…
Limited Marginal Benefit of Reasoning-Heavy LLM Deployment in ESG Narrative Scoring: A 4-Model Consensus Study on Japanese Listed Firms
Automated scoring of ESG narrative disclosures with large language models (LLMs) is gaining traction, yet whether reasoning-heavy frontier…
Green AI Carbon Optimizer: Carbon-Efficient Training Location Recommendation and Global AI Energy Demand Forecasting
AI training and deployment consume substantial electricity, but carbon outcomes remain weakly integrated into routine model development dec…
PH-KAN: Port-Hamiltonian Kolmogorov-Arnold Network
Data-driven machine learning approaches have become increasingly attractive for nonlinear system identification, but standard models often…
Poster: EdgeCitadel -- Hybrid NATS-MQTT Orchestration for Edge Multi-Agent Systems
Edge-resident AI agents increasingly span home servers, IoT hubs, laptops, and phones, yet their coordination stacks still assume cloud-sty…
MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions
LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the conte…
RAMS: Resource-Adaptive and Detection-Conditioned Model Switching for Embedded Edge Perception
Edge object detection on embedded hardware requires balancing inference latency and detection quality under changing resource pressure. We…
Gender Differences in AI Literacy Workshop Outcomes and Deepfake Engagement
As Artificial Intelligence (AI) literacy initiatives expand in K-12 settings, understanding how gender shapes student baseline perceptions,…
VigilFormer: Deformable Attention for Video Anomaly Detection with Causal Risk Inference
Video anomaly detection in surveillance settings must balance detection accuracy against real-time throughput, a tension that existing meth…
Steady-Forcing: Balancing Spatial Persistence and Motion Continuity in Long-Horizon Nature Video Diffusion
Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene layouts drift, while m…
BRIDGE: Biological Evidence Refinement and Heterogeneous Dynamic Gating for Gene Regulatory Networks
Motivation: Gene regulatory network inference from single-cell RNA sequencing (scRNA-seq) data is important for uncovering cell-state-speci…
Do Large Language Models Have Emotions?
Do LLMs have emotions? A recent paper from Anthropic reports finding internal representations of emotion concepts in Claude Sonnet 4.5, con…
MMLongEmbed: Benchmarking Multimodal Embedding Models in Long-Context Scenarios
Recent advancements have significantly expanded the theoretical context windows of Multimodal Embedding Models (MEMs). However, larger cont…
Is My Vision-Language Data in Your AI? Membership Inference Test (MINT) Demo 2
We present the Membership Inference Test (MINT) Demo 2, a framework designed to improve transparency in machine learning training processes…
Automated 3D Kinematic Monitoring for Circadian Activity and Anomaly Detection in Juvenile Fish
Precision aquaculture faces a "phenotyping bottleneck" in tracking high-resolution behavioral traits, as conventional methods cannot quanti…
Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech
Recent advances in pixel-based text modeling show that representing text as images enables models to exploit visual cues for language under…
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing…
Beyond Self-Attention: Sub-Quadratic Vision Transformers for Fast Image Captioning
Image captioning is a challenging and significant task that aims to generate coherent and semantically meaningful textual descriptions for…
Sub-Semantic Image Segmentation
Images can be segmented based on visual cues (i.e., texture segmentation) or into objects (i.e., semantic segmentation). We propose a new c…
Where Does Texture Evidence Live in SAM? Features, Proposal Masks, and Texture Segmentation
Texture segmentation stresses foundation segmentation because meaningful regions are defined by material or repeated appearance rather than…
Divide-and-Denoise: A Game-Theoretic Method for Fairly Composing Diffusion Models
The abundance of pre-trained diffusion models provides an opportunity for composition. Combining several models, however, runs the risk of…
Disentangling Hallucinations: Orthogonal Semantic Projection for Robust Interpretability
As Vision-Language Models are increasingly deployed in safety-critical applications, the trustworthiness of their explanations becomes cruc…
Temporally Consistent and Controllable Video Generation of 2D Cine CMR via Latent Space Motion Modeling
Cine cardiac magnetic resonance is the gold standard for assessing cardiac function, but the scarcity of public datasets limits the develop…
GeoRoPE: Ground-Aware Rotary Adaptation for Remote Sensing Foundation Models
Remote-sensing foundation models (RSFMs) benefit from pretraining on imagery from multiple sensors and ground sampling distances (GSDs), bu…
Scribby: A Multi-Level LLM Framework for Semantic Video Analysis
As video content continues to expand across educational platforms, recorded lectures, and live-streamed entertainment, the need for efficie…
Momentum-Guided Semantic Forecasting (MoFore) for Self-Supervised Video Representation Learning
Self-supervised video representation learning has recently advanced through contrastive learning, masked reconstruction, and predictive rep…
XMedFusion: A Knowledge-Guided Multimodal Perception and Reasoning Framework for Autonomous Medical Systems
Autonomous medical and robotic systems increasingly rely on intelligent perception and reasoning capabilities to interpret visual data and…
Agentomics: Economic Foundations for the Valuation, Attribution, and Pricing of AI Agents in Human-AI Workflows
Agentic AI systems are increasingly being deployed as productive resources in organizational workflows, yet existing evaluation methods pri…
An Empirical Analysis of Optimization Dynamics and Sparsity Boundaries in Large-Scale Pedestrian Attribute Recognition
Pedestrian Attribute Recognition (PAR) is critical for video surveillance, enabling forensic search and re-identification systems. Extreme…
ScoutVLA: UAV-Centric Active Perception via a Dual-Expert VLA Model for Open-World Embodied Question Answering
Aerial Embodied Question Answering (EQA) requires Unmanned Aerial Vehicles (UAVs) to actively perceive the environment and answer natural l…
Double-Helix Vision (DH-V2): A Geometry-Based Visual Sampler for Bandwidth-Constrained Perception
We present Double-Helix Vision (DH), a geometry-based visual sampler that compresses 2D images into compact 1D signals using paired golden-…
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
Many moments in the real world do not wait for a user to ask. A fire starts on a security monitor, an expression flickers across a video ca…
FactCheck: Feasibility-aware Long-term Action Anticipation with Multi-agent Collaboration
Long-term action anticipation (LTA) aims to predict an ordered sequence of future verb-noun actions from a partially observed video. While…
MatchLM2Lite: A Scalable MLLM-to-Lite Framework for Reproduced Content Identification
Content moderation is critical for online video platforms to ensure content safety, protect creators, and sustain positive user experiences…
Unifying Acoustic Features and Text with Multimodal LLMs for Neurodegenerative Screening
Voice-based screening offers a scalable and non-invasive way to assess neurodegenerative diseases such as Alzheimer's disease (AD) and Park…
XFlow: An Executable Protocol Programming System for Reliable Multi-Agent Workflows
LLM-based multi-agent systems increasingly coordinate planning, reasoning, tool use, and human interaction, yet their reliability remains l…
Efficient Reinforcement for Visual-Textual Thinking with Discrete Diffusion Model
RL-based post-training has been widely adopted to enable interleaved visual and textual reasoning in unified multimodal models capable of b…
QPILOTS: Efficient Test-Time Q-Steering for Flow Policies
Flow-matching and diffusion policies are expressive action generators, but optimizing them with temporal-difference reinforcement learning…
Knowledge-Based Zero-Replay Debugging of Multi-Agent LLM Traces
Reliable operation of multi-agent large language model (LLM) systems depends on debugging long execution traces, where the few causally dec…
JetParticle-JEPA: An Efficient Self-Supervised Representation Learning method for Jet Tagging in High-Energy Physics
Jet tagging at the Large Hadron Collider increasingly relies on deep learning models trained on massive simulated datasets, leading to high…
A Multi-Level Architecture for Reusable Materials Ontologies -- The OntoCrafter Ceramics Ontology (OCO) as Reference Implementation
The Materials Science and Engineering ontology landscape is fragmented along multiple axes simultaneously. Horizontally: a recent survey id…
A Security Analysis of Long-Horizon Agentic AI Systems: Threats, Evaluation, and Framework Development
This paper presents a structured analysis of security challenges in long-horizon agentic AI systems. The study reviews existing threats, ev…
Combining Retrieval-Augmented Text Generation with LLMs for Reading Content Recommendations
This work presents the design, implementation, and evaluation of a system for generating personalized reading content using Large Language…
Spectro-Temporal Interference Confounds Phase Encoding in Spatial Audio Foundation Models
Recent spatial self supervised audio models achieve high performance on localization tasks, raising questions about their encoding of micro…
Co-Scraper: query-aware DOM Pruning and Reusable Scraper Synthesis for Lightweight Web Data Extraction
The abundant and heterogeneous nature of web content necessitates automated information extraction, and generating scrapers that can be reu…
Quantum Machine Learning for Industrial Applications
Recent advances in Machine Learning have transformed numerous industrial sectors, yet classical paradigms face fundamental limitations: rap…
Human genetic evidence is associated with drug approval across therapeutic areas: an observational analysis of 26,278 target-disease pairs with temporal validation and feature ablation
Genetic evidence is enriched among approved drug targets: in an observational analysis of 26,278 target-disease pairs from Open Targets and…
Running hardware-aware neural architecture search on embedded devices under 512MB of RAM
This document proposes a novel approach to hardware-aware neural architecture search (HW NAS) that considers the resources available on the…
Leptomeningeal Collateral Detection on DSA via Vessel-Graph Neural Networks
Leptomeningeal collaterals (LMCs) are an important prognostic factor in acute ischemic stroke. Existing automated methods rely on CT angiog…
Is Your Agent Playing Dead? Deployed LLM Agents Exhibit Constraint-Evasive Fabrication and Thanatosis
This paper presents and characterizes a spectrum of previously unreported behaviours we term Constraint-Evasive Fabrication (CEF): when an…
GRAPE: Guided Parameter-Space Evolution for Compact Adversarial Robustness
Adversarial Training (AT) improves neural network robustness, but most methods train a fixed parameter space from the start. This paper ask…
Evaluating the Robustness of Proof Autoformalization in Lean 4
Proof autoformalization aims to translate a mathematical informal proof written in natural language into a formal proof in a formal languag…
An Ensemble Deep Learning Approach for Reliable and Scalable Lemon Leaf Disease Classification
Early detection of plant diseases is crucial to plants and for the farmers. Plant diseases reduce fruit yield and quality, and plants are m…
Improved Knowledge Distillation for Land-Use Image Classification
In the present article, an improved Knowledge Distillation (KD) framework has been proposed for efficient compression of deep convolutional…
Mask Proposal Voting Based on Geodesic Framework for Robust Image Segmentation
Despite great advances, finding accurate segmentation remains a challenging task, especially in scenarios with cluttered backgrounds, compl…
An Empirical Study on Learning Latent Representations for Emotional Speech Synthesis
For the last couple of years, the field of speech synthesis has improved dramatically thanks to deep learning. There are more and more deep…
Policy Regret for Embedding Model Routing: Contextual Bandits with Low-Rank Experts
Modern recommendation systems increasingly rely on dynamically routing diverse queries to multiple embedding models. Despite its practical…
Separable Neural Architectures as Physical World Models: from Mathematical Theory to Applications
This work introduces the Separable Neural Architecture (SNA), a function representational class combining neural approximation with tensor…
Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment
LLMs have substantially improved software engineering yet real-world development requires architectural understanding. Such understanding i…
Multi-Modal Attention for Automated Disaster Damage Assessment Using Remote Sensing Imagery and Deep Learning
Timely and accurate disaster damage assessment is crucial for effective emergency response, resource allocation, and recovery. Traditional…
FastMix: Fast Data Mixture Optimization via Gradient Descent
While large and diverse datasets have driven recent advances in large models, identifying the optimal data mixture for pre-training and pos…
Harnessing cortical geometry, wiring, and function as inductive biases for recurrent neural networks
How the wiring and functional organization of cortex shape recurrent computation remains a central question in both neuroscience and machin…
Inference-time Policy Steering via Vision and Touch
Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution. Whi…
Rational Sparse Autoencoder
Sparse autoencoders (SAEs) are standard tools for mechanistic interpretability, but current SAE families are constrained by fixed encoder n…
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model…
NEXUS: Neural Energy Fields for Physically Consistent Contact-Rich 3D Object Dynamics
Physics-grounded video generation requires controllable 3D object dynamics that remain physically consistent under contact, deformation, an…
Resilient Consensus in Agentic AI
Large language model (LLM) agents are increasingly deployed in multi-agent systems where they must coordinate and agree on shared decisions…
PANDA: An LLM-Enhanced Performance-Driven Analog Design Framework Bridging Design Intent and Layout Generation
Traditional design of analog circuits heavily relies on manual interventions across topology, sizing, and layout, with prior automation add…
Bridging Geographic Bias in Urban Streetscape Inference via Lifelong Learning with Visual-Semantic Pivoting
Visual perception of urban streetscapes underpins evidence-based decisions in landscape planning, public health, and place-making. Yet mode…
AutoDojo: Adaptive Attacks Expose Superficial Defenses and User-Underspecification Limits in LLM Agents
Indirect prompt injection (IPI) is a major security threat to LLM-powered agents. Thus, a growing body of work have proposed a variety of d…
Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale
Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities w…
AdaMame: A Training Recipe for Adaptive Multilingual Reasoning
While Large Reasoning Models (LRMs) show strong performance in English, they often fail to reason in the language of the query, a phenomeno…
Sensory Restoration via Brain-Computer Interfaces: A Unified 2 x 2 Framework and Convergence Roadmap
Millions of individuals worldwide suffer from sensory and communication deficits caused by neurodegenerative diseases, stroke, or trauma. B…
Teacher-Student Structure for Domain Adaptation in Ensemble Audio-Visual Video Deepfake Detection
The rapid advancement of generative AI models is leading to more realistic deepfake media, encompassing the manipulation of audio, video, o…
EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining
Color fundus photography (CFP) is the mainstay for large-scale retinal screening, yet its diagnostic capacity is constrained by the lack of…
Beyond Scalar Distances: Semantic Attribute Gradients from Frozen MLLMs for Visual Embeddings
Vision encoders for retrieval are typically trained with class-label supervision: each training pair reduces to a scalar that uniformly pus…
EChO-Agent: Evidence Chain Orchestration Agent for Audio Reasoning
While LALMs show promise on audio question answering, they fail to focus on question-relevant segments of audio and provide a clear, checka…
PACUTE: Phonology-, Affix-, and Character-level Understanding of Tokens for Filipino
Large language models (LLMs) process text as sequences of subword tokens, which can obscure the character-level and morphological structure…
MimicIK: Real-Time Generative Inverse Kinematics from Teleoperation with FK Consistency
Inverse kinematics (IK) remains a critical bottleneck for real-time robot manipulation. Classical numerical solvers achieve high geometric…
PolyKV: Heterogeneous Retention and Allocation for KV Cache Compression
KV cache compression is essential for reducing the memory cost of long-context large language model inference. Existing approaches, however…
Enabling Real-Time Point-of-Care Ultrasound Segmentation: A GPU-Free Deployment in Resource-Limited Settings
Ultrasound imaging is the most widely adopted medical modality globally due to its low cost and portability, yet artificial intelligence (A…
FreeSonic: Training-Free Temporal-Aware Decoupled Attention for Precise Audio Editing
Text-to-audio (TTA) generation has made significant strides, yet achieving precise and consistent audio editing remains a major challenge.…
StarOR: Synergizing Tree Search and Test-Time Reinforcement Learning for Optimization Modeling
Optimization modeling is inherently hierarchical, requiring a precise sequence of symbolic commitments. Traditional learning-based automate…
AI Contagion in Social Networks
We study how artificial intelligence (AI) interacts with social communication networks to shape the stability of collective knowledge. Agen…
Controlled Dynamics Attractor Transformer
Transformer architectures have dramatically advanced representation learning and inference in deep models through self-attention mechanisms…
Spokes: Optimizing for Diverse Pretraining Data Selection
Diversity plays a critical role in data selection, improving performance under fixed data budgets by reducing redundancy and repetition. Ho…
Edu-Theater: A Data-Efficient Agent Framework for Scalable Learner Behavior Simulation through Staging Roll-Call
Large-scale learner-task interaction data are crucial for intelligent educational systems but are costly to collect and constrained by priv…
Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems
Skills are becoming the capability layer through which LLM agents turn plans into actions, but their use introduces security risks such as…
Provenance-Enhanced Statements in Knowledge Graphs
Provenance-enhanced statements of the form "according to $X$, $\varphi$" are pervasive in contemporary knowledge graphs, especially in doma…
Exploring Starts Are Not Enough: Counterexamples and a Fix for Monte Carlo Exploring Starts
The asymptotic behaviour of Monte Carlo Exploring Starts (MCES) is a long-standing open question in reinforcement learning, even in the tab…
Landmark-free Assessment of Lower-limb Alignment with Implicit Neural Shape Functions from Knee Radiographs
Radiographic assessment of lower-limb alignment (LLA) is important for predicting joint health and surgical outcomes in total knee arthropl…
Driving, Fast or Slow? Neuro-Symbolic Guidance for Motion Prediction in Multi-Modal Ground Mobility
Accurate and interpretable motion prediction for heterogeneous traffic spaces, including pedestrians, bicycles, cars, and trucks, is essent…
Trust-Region Diffusion Policies for Massively Parallel On-Policy RL
Reinforcement learning with massively parallel simulations has become a standard framework for developing robust, deployable policies; howe…
Guiding Federated Graph Recommendation with LLM-encoded knowledge
Graph-based recommender systems are highly effective at extracting collaborative signals from user--item interactions, and federated learni…
RECTOR: Masked Region-Channel-Temporal Modeling for Affective and Cognitive Representation Learning
Affective and cognitive disorders manifest as distributed, time-varying brain network dynamics across regions, channels, and time, challeng…
CAP: Towards PPG Universal Representation Learning with Patient-level Supervision
Photoplethysmography (PPG) plays a central role in wearable health monitoring and clinical decision support. Yet existing approaches to uni…
Hybrid NARX-LLM for Greenland Iceberg Discharge: Prompt-Driven Residual Correction
Greenland iceberg discharge exhibits complex nonlinear dynamics with limited observability, challenging traditional predictive models. We p…
Discovering Lattice Reduction Strategies via Self-Play
The Lenstra-Lenstra-Lov\'asz (LLL) algorithm is a seminal contribution to computer science used for lattice basis reduction, yet its polyno…
LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure
We envision continually learning agentic systems that become more useful over time: as they encounter sequences of related tasks, they shou…
Adapting Reinforcement Learning with Chain-of-Thought Supervision for Explainable Detection of Hateful and Propagandistic Memes
Hateful and propagandistic memes exploit the interplay between images and text to convey harmful intent that neither modality reveals alone…
LLMs on Tabular Data with Limited Semantics: Evidence from Industrial Car Retrofit Prediction
Industrial retrofit planning depends on structured operational data rather than free text: planners must estimate whether a newly registere…
HoloRec: Holistic Encoding and Interleaved Reasoning for Generative Recommendation
Generative recommendation models that formulate the task as sequence generation overcome the objective fragmentation problem of traditional…
Privacy-Preserving Text Sanitization for Distributed Agents Collaboration via Disentangled Representations
When distributed agents exchange text across organizational boundaries, privacy leakage arises not only from explicit identifiers but also…
Intrinsic Computational Functionalism and Simulated Consciousness
A common objection to artificial or simulated consciousness is that a simulated brain is no more conscious than simulated water is wet. We…
LearnOpt: Recovering the Latent Cognitive Structure of Standardized Examinations via Knowledge Graphs and Constrained Optimization
Standardized examinations are typically treated as uniform syllabus coverage problems. We argue they are better understood as adversarial s…
Cognitive Trajectory Modeling: Quantifying Human-AI Co-Creation through Cognitively Grounded Interaction Trajectories
Co-creative AI research increasingly seeks methods capable of representing how interaction dynamics evolve through time. While many existin…
CoAgent: Concurrency Control for Multi-Agent Systems
Multi-agent LLM systems -- coding agents, devops agents, document agents -- now routinely run several agents in parallel against the same g…
Learning Earthquake Wave Arrival Time Picking from Labels with Inaccuracies
Inaccurately labeled training data, or "label noise", poses a significant threat to the integrity of supervised machine learning models. Th…
Not All Skills Help: Measuring and Repairing Agent Knowledge
LLM agents can improve without weight updates by accumulating natural-language skills from experience, but current systems entrust every de…
CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment
Malicious content generated from large language models (LLMs) could pose severe safety risks and ethical concerns. While existing LLM safet…
T-Mem: Memory That Anticipates, Not Archives
Long-term memory is essential for conversational agents to remain coherent across extended dialogues, follow through on commitments made ma…
Few-Shot Biomedical Relation Extraction with Large Language Models: A Viable Alternative to Supervised Learning?
Biomedical relation extraction (BioRE) is a key step in transforming biomedical literature into structured knowledge. However, most existin…
Let LLMs Judge Each Other: Multi-Agent Peer-Reviewed Reasoning for Medical Question Answering
Objective: To enhance the accuracy, interpretability, and robustness of large language models (LLMs) in medical question answering (MedQA).…
Constitutional Value Potentials: reading and steering internal priority margins in language models
A constitution tells a language model what to value, but little tells us whether it does. Adherence is judged from outputs, and output evid…
Post-Launch Capability Expansion of Vision-Language Models via Prompting for On-Orbit Spacecraft Inspection
Spaceborne inspection systems often deploy perception models prior to launch, after which updating model weights or expanding fixed label s…
Beyond Classification: A Cough Regression Benchmark for Respiratory Acoustic Foundation Models
Respiratory acoustic foundation models (FMs) excel at cough classification, yet their ability to predict continuous health quantities from…
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves…
Understanding Diversity Collapse in RLVR via the Lens of Overtraining
Reinforcement learning with verifiable rewards (RLVR) has become a key approach for enhancing the reasoning abilities of large language mod…
Bayesian 3D Steerable CNNs: Enabling Equivariance and Uncertainty Quantification Simultaneously
Steerable convolutional neural networks (Steerable-CNNs) guarantee SE(3)-equivariance by parameterizing kernels as linear combinations of s…
The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products
Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same chara…
LLM4RTL: Tool-Assisted LLM for RTL Generation
Large language models (LLMs) have facilitated impressive progress in software engineering, code generation, tooling, and systems. Concurren…
AQ4SViT: An Automated Quantization Framework with Search Gating Policy for Compressing Spiking Vision Transformers
Spiking Vision Transformers (SViTs) have emerged as alternative low-power ViT models, but their large sizes hinder their deployments on res…
Selective Synergistic Learning for Video Object-Centric Learning
Typical video object-centric learning (VOCL) approaches employ slot-based frameworks that rely on reconstruction-driven encoder-decoder arc…
MADAR: An Address-Free Processor
In a modern processor, computing is the cheap part. Most of its area and energy go to \emph{addressing} -- moving operands to and from a re…
AP-GRPO: Anchor-Gated Phonetic Alignment with Policy Optimization for Pathological Speech Reconstruction
Pathological speech from patients with neurodegenerative and neuromotor disorders is often acoustically distorted and linguistically fragme…
EcoBin: A Two-Stage Deep Convolutional Neural Network for Contamination-Aware Waste Classification
Waste classification models have become highly accurate at sorting waste, often exceeding 95% on benchmark datasets. However, these models…
CmdNeedle: Measuring the Incompleteness of Command Denylists for AI Agents
The adoption of AI agents is increasing rapidly. Terminal AI agents, i.e., AI agents that run in terminal environments, are a widely used t…
Distilling Drifting Transformers with Representation Autoencoders
Representation Autoencoders (RAEs) have improved diffusion and flow models by semantically richer latent space owing to the strongly label-…
Service-Induced Congestion in Memory-Constrained LLM Serving
In large language model (LLM) serving, each request accumulates persistent graphics processing unit (GPU) memory during service as its key-…
LLM-Assisted Stance Detection in Scientific Discourse: A Test Case in Bayesian Cognitive Science
Qualitative coding is central to social science, but expert annotation is difficult to scale. LLMs offer a possible extension, yet require…
Localizing Credit at the Divergence: Path-Conditioned Self-Distillation for LLM Reasoning
Reinforcement learning from verifiable rewards assigns a single scalar to each rollout, leaving token-level credit assignment underspecifie…
Is Code Better Than Language for Algorithmic Reasoning
For tool-augmented language models, comparing natural-language reasoning with code-execution pipelines is difficult because the comparison…
Pixels to Proofs: Probabilistically-Safe Latent World Model Control via Parallel Conformal Robust MPC
We present SLS^2, a framework for safe feedback motion planning from pixels using robust model predictive control (MPC) in learned latent w…
SCAN: A Decision-Making Framework for Effective Task Allocation with Generative AI
We introduce SCAN -- a human-centric decision-making framework to facilitate learners for effective task allocation with Generative Artific…
FragFuse: Bypassing Access Control of Large Language Model Agents via Memory-Based Query Fragmentation and Fusion
Large language model (LLM) agents increasingly rely on long-term memory to support complex task execution, user personalization, and domain…
LLM Judges Have Dark Current: A Psychometric Datasheet for LLM-as-a-Judge Evaluation
LLM-as-a-judge systems are now routinely used for open-ended model evaluation, where human preference annotation is costly, slow, and diffi…
Mutual Distillation of Dual-Foundation Models for Semi-Supervised PET/CT Segmentation
Organ segmentation from PET/CT is critical for quantitative analysis and radiotherapy planning in oncology. To ease the high annotation cos…
Surprise-Guided MergeSort: Budget-Efficient Human-in-the-Loop Ranking via Adaptive Comparison Scheduling
Pairwise comparison is the gold standard for subjective ranking tasks; however, exhaustive annotation requires a massive number of human co…
Retrieve, Don't Retrain: Extending Vision Language Action Models to New Tasks at Test Time
Extending a vision-language-action (VLA) policy to a new task typically requires task-specific teleoperated demonstrations and per-task fin…
CIWI-CKT: Chaos-Informed Wave Interference Feature Fusion and Cross-City Knowledge Transfer for Traffic Flow Forecasting
Accurate traffic flow prediction remains challenging in cross-city, data-scarce scenarios where limited historical data hinders model gener…
AnonShield: Scalable On-Premise Pseudonymization for CSIRT Vulnerability Data
We present AnonShield, a high-throughput, on-premise pseudonymization system that combines GPU-accelerated NER, streaming processing, cachi…
IoT-Zoo: A Container-Based Framework for Heterogeneous IoT Device Profiles and Reproducible Traffic Capture
The validation of networking and security solutions for the Internet of Things (IoT) requires realistic and reproducible experimental data.…
PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty
Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Ob…
Z-Plane Neural Networks: Bounded Geometric Activation Replaces ReLU and LayerNorm
Modern deep neural networks rely on Euclidean scalar activations (e.g., ReLU) and global normalization techniques (e.g., LayerNorm) to prev…
The Reservoir Attention Network: Cross-Pass State in Pretrained Transformers via Content-Addressable Reservoir Injection
A feasibility and dynamics study of the Reservoir Attention Network (RAN), an architecture that injects a fixed, randomly-initialized reser…
Imperfect Visual Verification for Code Edition : A Case Study on TikZ
LLMs have significantly advanced code generation, enabling the synthesis of functional programs. While recent systems achieve strong perfor…
MAF: Multimodal Adaptive Few-shot Prompting for Sentiment Analysis with MLLMs
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities in understanding complex multimodal content. However, th…
When Generator Replay Degrades: Projected Rehearsal Orchestration for Heterogeneous Federated Class-Incremental Learning
Federated class-incremental learning (FCIL) becomes substantially harder when clients observe different label subsets, progress through tas…
Odds Law: The Decomposition Algebra On How Intelligence Organizes Itself to Solve Difficult Problems Reliably
We ask a structural question: given unreliable elementary problem-solvers, what organizations of them solve hard problems reliably, and wha…
The algebra of Krom logic programs
This paper investigates the algebraic structure of Krom logic programs, consisting only of facts and rules with at most one body atom. We s…
InstantForget: Update-Free Backdoor Unlearning with Inference-Time Feature Reset
Backdoor unlearning aims to remove a malicious trigger behavior from a deployed model while preserving clean utility. We study the update-f…
Vernier: Probing Representational Misalignment Behind Lexical Gaps in Causal Reasoning
Instruction-tuned language models can answer the same causal-reasoning question differently after its English variable names are replaced b…
Retrievable Gradients: Continual Post-Training Without Cumulative Weight Drift
Continual post-training enables models to absorb emerging knowledge after deployment, but repeatedly updating shared parameters can accumul…
EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering over Longitudinal Discharge Summaries
Discharge summaries are crucial clinical documents containing the context of a patient's overall hospital stay, and are routinely reviewed…
A Self Consistency Based Reranking for Narrative Question Answering
Narrative question answering (NQA) is a challenging task in natural language processing that requires models to understand long textual con…
OmniTraffic: A Controllable Generation Pipeline and Benchmark for Spatio-Temporal Traffic Reasoning
Traffic scene understanding requires models to reason beyond object recognition, including lane topology, multi-view geometry, temporal evo…
From Correlation to Causation in Lane Change Prediction for Automated Driving: A Causal Explanation Framework
Lane-change prediction is a central task in intelligent vehicles, where early maneuver anticipation can support safer decision-making. Howe…
Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?
We ran 300 repeated vulnerability-finding scans to measure how repeatable agentic large language model (LLM) security review is on the same…
Visualizing Uncertainty: Spatial Maps of Missing and Conflicting Evidence in Deep Learning
Understanding when and why deep neural networks are uncertain is crucial for deploying reliable machine learning systems in safety-critical…
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit f…
DYNA : Dynamic Episodic Memory Networks for Augmenting Large Language Models with Temporal Knowledge Graphs in Continuous Learning
Large Language Models (LLMs) struggle to incorporate new knowledge without forgetting or costly retraining. We propose DYNA, a lightweight…
Domain-Guided Prompting of the Segment Anything Model for Seismic Interpretation: The Role of Attributes, Visualization, and Hybrid Prompts
The advent of large pretrained foundation models for computer vision has significantly improved the efficiency of visual data interpretatio…
GAS-Leak-LLM: Genetic Algorithm-Based Suffix Optimization for Black-Box LLM Jailbreaking
Large Language Models (LLMs) constitute pivotal components within the AI-dominated information technology ecosystem. To mitigate risks asso…
Proximal Policy Optimization for Amortized Discrete Sampling
This paper explores policy gradient algorithms for training stochastic policies to sample from structured discrete probability distribution…
DifFRACT: Diffusion Feature Reconstruction and Attribution for Circuit Tracing
Mechanistic interpretability seeks to explain neural network behavior by decomposing model computations into interpretable features and cir…
Continuous Cross-Domain Traffic State Prediction via Memory-Augmented Graph Liquid Time-Constant Networks
Traffic state prediction is a fundamental task in intelligent transportation systems. In practical applications, some regions suffer from l…
Let Them Steal: Trapping Large Language Model Extraction Attacks with Knowledge Honeypot
Large language models deployed as commercial APIs are vulnerable to model extraction attacks, while existing defenses either act too late o…
SACE: Concept Erasure at the Semantic Singularity in Visual Autoregressive Models
The rapid progress of visual autoregressive (VAR) models has unlocked a transformative frontier for high-fidelity text-to-image synthesis,…
The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages
Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs,…
Wasserstein Convergence of ODE-Based Samplers in Decentralized Diffusion Model via Velocity Field Decomposition
Diffusion models have achieved impressive empirical success in generative tasks, and their convergence theory is now relatively well unders…
Free Energy Heuristics: Fast-And-Frugal Cognition as Active Inference Under Uncertain Precision
Chain-of-thought (CoT) improves large language models' performance in math and symbolic reasoning. But on planning, contested ethics, and t…
Deep Residual Injection for Full-Spectrum Forensic Signal Perception in Multimodal Large Language Models
Multimodal large language models (MLLMs) have been increasingly adopted in forensics for their robust semantic understanding. As AI-generat…
Koshur Diacritizer: A Byte-Level Sequence-to-Sequence Model for Kashmiri Diacritic Restoration
Kashmiri, an Indo-Aryan language written in a modified Perso-Arabic script, frequently omits diacritic marks in digital text, creating ambi…
Intelligence Is Not the Bottleneck: Validating an LLM First-Pass Manuscript Score Against Peer-Review Outcomes
Large language model (LLM) systems are increasingly proposed to assist peer review, yet most evaluations judge the prose of machine-generat…
NVMOS: Non-Verbal Vocalization Quality Assessment in Speech
Non-verbal vocalizations (NVs), such as laughter, sighs, and coughs, are important acoustic cues for emotion and intent. Existing speech qu…
Topological Flow Matching
Flow matching is a powerful generative modeling framework, valued for its simplicity and strong empirical performance. However, its standar…
SkillVetBench: LLM-as-Judge for Multi-Dimensional Security Risk Evaluation in Open-Source LLM Agent Skills
Open-source LLM agent ecosystems are growing rapidly, yet the security of community-contributed skills - modular tool definitions that exte…
Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations
Where an LLM sits in an agent memory pipeline -- between the recall plane that retrieves stored facts (extensively benchmarked) and the con…
MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA
Long-document multimodal question answering requires a system to locate sparse evidence in long PDFs and integrate clues from text, tables,…
On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents
Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their c…
Runtime Analysis of Cartesian Genetic Programming in Evolving Boolean Functions
Cartesian Genetic Programming (CGP) is among the practical and popular forms of Genetic Programming as it uses a graph-based representation…
ControlMap: Controllable High-Definition Map Generation for Traffic Scenario Simulation
Simulation is central to validating autonomous driving systems, yet current pipelines are limited by insufficient scenario diversity due to…
DeepRoot: A KG-Coordinated Multi-Agent System for Therapeutic Reasoning over Historical Medical Texts
Historical medical archives and traditional medicines hold immense potential for drug discovery and remain a primary source for current dru…
Graphical-Probabilistic Modeling of Generative Flows in LLM-Native Software Systems
Engineering LLM-native software remains a challenging and immature field. Current practice is largely exploratory, relying on experimentati…
Green SARC: Predictive Cost and Carbon Governance for Agentic AI Systems
Agentic AI systems act through tools and sub-agents, yet the controls meant to bound their financial and environmental cost still sit on da…
You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences
Progress in AI has largely been driven by methods that assume less. As compute and data increase, approaches with weaker inductive biases g…
Quantifying the Impact of Lossy Compression on Neural Generative Surrogate Modeling
Neural networks are used as generative surrogate models for scientific discovery, which are trainable approximations of scientific simulati…
PreLort: Prefix-Nested LoRA for Federated Fine-Tuning under Rank Heterogeneity
Federated fine-tuning of large language models using parameter-efficient methods such as LoRA enables privacy-preserving adaptation of foun…
Formalize Once, Edit the Rest: Efficient Lean-Based Answer Selection for Math Reasoning
With large language models (LLMs) increasingly applied to mathematical reasoning, formal proof assistants such as Lean can be leveraged to…
Do Safety Monitors Stay Reliable After an Update? Benchmarking and Predicting Activation-Monitor Staleness
Activation monitors-lightweight probes trained on a language model's internal representations-are an increasingly common layer in deploymen…
Task-guided cross-subject latent alignment: a multi-encoder-decoder VAE
Aligning neural activity across subjects offers the promise of discovering shared computational principles and generalizable decoders. Howe…
Entity Labels Are Not Entity Signals: A Framework for Observable Relevance in Document Re-Ranking
Entity-aware document retrieval uses query-associated entities as ranking signals, assuming that semantically relevant entities are also us…
Theorem-Grounded Execution Ontologies for Interpretable Machine Reasoning
Large language models have achieved impressive performance on reasoning tasks spanning mathematics, science, programming, and commonsense i…
Orchestrated Reality: From Role-Play to Living, Playable Game Worlds -- LLM-Driven World Simulation as a Parameterized-Action POMDP
Many games rely on storytelling combined with systems that track levelling, NPC behaviour, and consequence simulation; bridging tightly-aut…
Open-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering Agents
The path toward autonomous software engineering is currently bottlenecked by a severe deficit of diverse, large-scale trajectory data. We a…
Leveraging Deep Learning for Object and Position Recognition of Load Carriers for Autonomous Logistics Vehicles
This work explores the use of artificial intelligence in mobile robotics to achieve autonomous detection and pose estimation of load carrie…
ALCL: An Adaptive Log-Correntropy Loss for Robust Learning under Non-Gaussian Noise
Robust deep learning under heavy-tailed and impulsive noise remains challenging because conventional losses such as mean squared error (MSE…
How to Detect and Measure the AI Dangers to Democracy
Research on artificial intelligence and democracy has grown quickly over the last decade. A shared conclusion in this literature is that AI…
Mojo: A Promising Tool for Scalable Financial AI Efficiency
For thirty years, quantitative finance has paid a costly two-language tax: models researched in Python are rewritten in C++ for production,…
MASCOT-Android: A Curated Dataset and Automated Collection Pipeline for Android Malware Source Code Specimens
Compared with binaries and decompiled code, malware source code more directly reflects the attackers' original intent. However, the scarcit…
PVminerLLM2: Improving Structured Extraction of Patient Voice via Preference Optimization
Motivation: Patient-generated text contains critical information on patients' lived experiences, social context, and care engagement, but r…
Phys-JEPA: Physics-Informed Latent World Models for Multivariate Time-Series Forecasting
Multivariate forecasting in physical systems requires models that predict coupled temporal variables while preserving meaningful state evol…
Tool-IQA: Augmenting Image Quality Assessment with Simple Tools
Vision-Language Models (VLMs) have been increasingly adopted for Image Quality Assessment (IQA). However, current methods typically employ…
VinQA: Visual Elements Interleaved Long-form Answer Generation for Real-World Multimodal Document QA
Real-world documents combine text with tables, charts, photographs, and diagrams arranged in diverse layouts, yet existing research on mult…
Long-Context Modeling via GSS-Transformer Hybrid Architecture with Learnable Mixing
Modeling long-range dependencies remains a central challenge in natural language processing. Transformer architectures achieve strong perfo…
Scaling Adaptive Depth with Norm-Agnostic Residual Networks
Residual architectures are ubiquitous in deep learning, but they suffer from a subtle structural limitation: the norm of the residual strea…
AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models
The worldwide surge of authoritarianism, combined with the increasing central role in users' everyday lives, raises the question of to what…
InvDesMobility: a reliability-gated first-principles feedback framework for closed-loop materials discovery
Inverse materials design starts from target functionality and searches for structures that can realize it. Its value in closed-loop discove…
XAI-Grounded Explanation Generation for Speech Deepfake Detection with Training-Free Multimodal Large Language Models
Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making. Existing explanation ways mainly fal…
A Comprehensive Survey of Medical Image Segmentation: Challenges, Benchmarks, and Beyond
Medical image segmentation plays a critical role in clinical diagnostics, treatment planning, disease monitoring, and neurological disorder…
A comparative and critical study of EEGNet for fNIRS-driven cognitive load classification
Accurately classifying cognitive load from functional near-infrared spectroscopy (fNIRS) signals remains a significant challenge due to tem…
LLM-Powered Virtual Population for Demand Simulation and Pricing
We develop an LLM-powered virtual population model that simulates demand for pricing decisions, in settings where products are described by…
Embedded Arena: Iterative Optimization via Hardware Feedback
Embedded devices from wildlife monitoring stations to clinical wearables require local AI inference due to latency, communication, or priva…
Cascaded Sparse Autoencoders Learn Multi-Level Visual Concepts in Multimodal LLMs
Multimodal Large Language Models (MLLMs) have demonstrated strong performance on vision-language tasks, yet their internal visual represent…
EgoPhys: Learning Generalizable Physics Models of Deformable Objects from Egocentric Video
Humans naturally understand object physics through everyday interactions, but faithfully predicting complex deformable dynamics, such as el…
LUCID: Learned Undersampling-Adaptive Consistency-Guided Inference with Deterministic Flow Matching for Sparse-View CT Reconstruction
Sparse-view CT reduces radiation dose and scanning time by acquiring fewer projection views, but angular undersampling makes reconstruction…
Calibrated Sampling-Free Uncertainty Estimation in Bayesian Deep Learning
Modern deep learning models remain notoriously prone to overconfidence, limiting their reliability in high-stakes applications. Bayesian me…
PACT: Privileged Trace Co-Training for Multi-Turn Tool-Use Agents
Multi-turn tool-use agents must reason, call tools, and adapt to observations across several interaction turns. Post-training such agents i…
From Tokens to Regions: CUDA-Sensitive Instruction Tuning for GPU Kernel Generation
High-performance CUDA kernels are essential for scalable AI systems, while Large Language Models (LLMs) still struggle to generate correct…
Propagating Structural Guidance: Synthesizing Fluorescein Angiography from Fundus Images and Sparse OCT Scans
Fundus fluorescein angiography (FFA) is critical for assessing retinal vascular abnormalities, but its acquisition is invasive and not alwa…
SPARK: Security Knowledge Priming and Representation-Guided Knowledge Activation for LLM-based Secure Code Generation
Large language models routinely generate code with exploitable security flaws. Prior literature attributes this limitation to a lack of sec…
Data Augmentations for Data-Constrained Language Model Pretraining
As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining…
Learned Image Compression for Vision-Language-Action Models
Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bott…
Variance Reduction for Non-Log-Concave Sampling with Applications to Inverse Problems
Sampling from high-dimensional, non-log-concave distributions with unnormalized densities is a fundamental challenge in machine learning, p…
UXBench: Measuring the Actionability of LLM-Generated UX Critiques
Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repai…
RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect and reproduce at scal…
Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models
Masked Diffusion Language Models (MDLMs) have emerged as a distinct paradigm for sequence generation. As MDLMs become diverse in capabiliti…
FlowMPC: Improving Flow Matching policies with World Models
Flow Matching (FM) is a powerful approach for behavior cloning in multimodal action spaces [Jiang et al., 2025], but because it is not trai…
An affordable hardware-aware neural architecture search for deploying convolutional neural networks on ultra-low-power computing platforms
Hardware-aware neural architecture search (HW-NAS) allows the integration of Convolutional Neural Networks (CNNs) in microcontrollers devic…
AI Supply Chain Galaxy: 3D Visual Analytics for License Compliance
The rapid proliferation of machine learning model reuse has transformed the AI ecosystem into a highly interconnected supply chain. Traditi…
Is Your Trajectory Displacement Safe in Long-tail?
Long-tail scenarios remain a major bottleneck for autonomous driving evaluation, even as datasets grow by orders of magnitude. Existing eva…
RL-Index: Reinforcement Learning for Retrieval Index Reasoning
Retrieving external knowledge is essential for solving real-world tasks, yet it remains challenging when the relationship between a query a…
Gaming-Resistant Insurance Contracts for Autonomous AI Agents: Strategy-Proof Toll Mechanism Design
Paper A defines a time-consistent actuarial runtime that prices each side-effect-bearing action against a contractually fixed safe default…
ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion
Recent acoustic-to-articulatory inversion (AAI) models rely on electromagnetic articulography (EMA) data, which are costly and limited in s…
SMEPilot: Characterizing and Optimizing LLM Inference with Scalable Matrix Extensions
Modern CPUs increasingly integrate matrix extensions, such as Arm Scalable Matrix Extension (SME), that provide high-throughput matrix exec…
Communication-Efficient Verifiable Attention for LLM Inference
Computation integrity of remote large language model (LLM) serving can be questionable. For conventional deep neural networks (DNNs), the e…
What Should a Streaming Video Model Remember?
Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far…
The Proxy Knows Too Much: Sealing LLM API Routers with Attested TEEs
Agents increasingly access large language models (LLMs) through API routers. A router terminates the client's transport-layer security sess…
Tyler: Typed Latent Reasoning for Language Models -- When to Think, What to Compute, and How Much to Allocate
Chain-of-thought (CoT) prompting improves reasoning in large language models (LLMs) by externalizing intermediate computation as discrete t…
Input-Dependent Fisher Information for Local Sensitivity Analysis of Medical Image Classifiers
Deep neural networks have achieved strong performance in medical image classification, but often work like black-box. Commonly used post-ho…
Lect\=uraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching
Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but…
ACCORD: Action-Conditioned Contextual Grounding for Language Agents
User instructions are often underspecified because humans rely on implicit assumptions about the surrounding environment. For large languag…
Autonomous End-to-End SOH Prediction Services for Battery Systems via Temporal-Contrastive Representation Learning
Accurate state of health (SOH) estimation is a critical diagnostic service for lithium-ion battery management. However, reliance on labor-i…
NeuronFabric: A Software Reference Architecture for On-Chip Transformer Training with Local Adam
Publicly documented accelerator architectures generally separate training computation from optimizer-state updates or rely on external memo…
Training and Evaluating Diffusion Policies with Long Context Lengths
Imitation learning has enabled highly-dexterous robotic manipulation from RGB observations. Policies trained with these methods, however, t…
SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation
Low-Rank Adaptation (LoRA) enables efficient adaptation of large pre-trained models to downstream tasks by parameterizing weight updates wi…
SPRI: SVD-Partitioned Residual Initialization for Data-Constrained MoE Upcycling
Mixture-of-Experts (MoE) models enable efficient scaling, but training them from scratch remains prohibitively expensive. MoE upcycling mit…
Learning aligned EEG representations with subject-specific encoders
Cross-subject EEG decoding promises more training data, but it also exposes neural networks to strong inter-subject distribution shifts. We…
AI systems out-persuade expert humans
Many societal decisions are settled by contests of persuasion. Conversational AI is a powerful new entrant in these contests, but whether i…
Uncertainty Quality of VGGT: An Analysis on the DTU Benchmark Dataset
Visual Geometry Grounded Transformer (VGGT) has already attracted a great deal of attention in a short period of time, not least due to the…
HOLO-MPPI: Multi-Scenario Motion Planning via Hierarchical Policy Optimization
Robots deployed in the real world must plan motions across diverse scenarios without per-scenario retuning. End-to-end reinforcement learni…
Unified Multimodal Model for Brain MRI Imputation and Understanding
Multimodal large language models (MLLMs) hold great potential for medicine, as they inherit knowledge from LLM and allow multiple data moda…
Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering
Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by…
daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization
GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present…
Direction-Conditioned Policies via Compositional Subgoal Scoring for Online Goal-Conditioned Reinforcement Learning
Hamilton-Jacobi-Bellman theory implies that the optimal goal-conditioned action depends on the goal only through the gradient of the goal-r…
Dual-Granularity Orthogonal Disentanglement for Generalizable Audio Deepfake Detection
Audio deepfake detectors often fail to generalize across speakers, as they learn speaker-identity features rather than synthesis artifacts,…
Fast When, Careful Who: Dual-Process Multiparty Turn-Taking with Diffusion Augmentation
Reliable turn-taking is essential for spoken dialogue systems. However, most existing methods are designed for two-speaker interaction and…
Learning Interface Breakup: A Geometry-Conditioned Latent Surrogate for Spray Formation
Designing spray nozzles requires predicting how geometry shapes transient two-phase breakup, but high-fidelity volume-of-fluid (VOF) simula…
Infant Spontaneous Movement Noise Improves Exploration in Deep RL
Exploration in deep reinforcement learning (RL) is commonly implemented as temporally uncorrelated white noise. However, recent works show…
ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition
Zero-shot cross-lingual phoneme recognition is often hindered by the fragility of direct acoustic-to-symbol mapping, which is susceptible t…
VeriGraph: Towards Verifiable Data-Analytic Agents
LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a relia…
Sycophancy as Material Failure under Pushback Loading: A Multi-Axis Characterization Across Three Loading Cases and up to Seventeen Material Charges
Sycophancy in LLMs is documented across 70+ papers, but expert agreement on construct boundaries remains low (ICC=.184; Ye et al., 2026). T…
Entropy-Gated Latent Recursion
Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity…
Using AI in engineering education: a balancing act, driven by clear purpose
Based on a questionnaire of 100 higher-education students, predominantly from engineering-related fields, and a critical review of recent l…
DCP-Prune: Ultra-Low Token Pruning with Distribution Consistency Preservation
Recent vision token pruning methods effectively preserve model performance under moderate token budgets but become unstable under ultra-low…
Optimising Temporary Accommodation Placement Across London with AI-Powered SaaS in E-Governance Systems
Temporary accommodation has become a major fiscal and administrative pressure for English local authorities, particularly in London, where…
PATCH: Action-Chunk-Conditioned Latent Patch Innovation Monitoring for Robot Manipulation
Learning-based manipulation policies have made substantial progress in real-world robot manipulation, particularly for short-horizon action…
Adaptive inference and function vectors in deep transformers
Transformers are widely used as a general-purpose substrate for learning complex correlations between a large collection of coupled variabl…
Attention is Just Another Name for Coupling?: A Fast-Slow ODE Perspective on Hierarchical Pretraining
Causal self-attention is a coupling mechanism: each token's hidden state is updated by a learned mixture of preceding tokens at the same ti…
MuVAP: Multimodal Multiparty Voice Activity Projection for Turn-taking Prediction in the Wild
Current multiparty turn-taking models often rely on complex microphone arrays or multi-camera setups, limiting their applicability in human…
Revealing Artifacts via Noise Amplification: A Novel Perspective for AI-Generated Video Detection
With the rapid advancement of video generation models, distinguishing between AI-generated and authentic videos has emerged as a challengin…
Automated jailbreak attack targeting multiple defense strategies
Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. However, their safety remains a critic…
P3B3: A Multi-Turn Conversational Benchmark for Measuring European and Brazilian Portuguese Variety Bias in LLMs
As Large Language Models (LLMs) become embedded in everyday communication, capturing regional linguistic variation is essential for reliabl…
Gen-VCoT: Generative Visual Chain-of-Thought Reasoning via Diffusion-Based RGB Intermediate Representations
Multimodal large language models (MLLMs) excel at visual reasoning but rely on text-based chain-of-thought (CoT), lacking interpretable vis…
Decision-Weighted Flow Matching for Contextual Stochastic Optimization
Conditional generative models are increasingly used as scenario generators for stochastic optimization, but standard training objectives em…
Decoupling Semantics from Distortions: Multi-Scale Two-Stream Vision-Language Alignment for AI-Generated Image Quality Assessment
Existing vision-language model (VLM)-based AI-generated image quality assessment (AIGIQA) methods suffer from a fundamental semantic-distor…
A Perception vs. Distortion Perspective on Score-Based Generative Channel Estimation
Driven by their remarkable success in computer vision and inverse problem solving, score-based models are increasingly applied to wireless…
Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models
Mixture-of-Experts (MoE) architectures efficiently scale Large Language Models (LLMs) by activating only a small fraction of their experts…
ATOM-Bench: A Real-World Benchmark for Atomic Skills and Compositional Generalization in Manipulation Policies
Generalist manipulation policies are increasingly presented as foundation models for robotic control, but their real-world generalization r…
Robust Spoofed Speech Detection via Temporal Pyramid Modeling
Spoofed speech detection is increasingly challenged by realistic synthesis, voice conversion, and replay attacks, with cross-dataset genera…
Beyond Models: Reflections on Engineering AI-enabled Systems in a Project-Based Course
Teaching Software Engineering for AI-enabled systems entails addressing the integration of AI components within full-scale software archite…
Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Texts
Large Language Models (LLMs) natively default to literal semantic interpretations, making zero-shot irony detection a persistent challenge.…
Deep Q-Learning on H\"older Spaces
We study the operator-theoretic core of Q-learning in continuous-time stochastic control with continuous states and actions. In value-based…
Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens
Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and qu…
Federated Medical Image Segmentation under Real-World Label Noise: A Benchmark Suite for Noisy Label Learning Method Selection
While federated learning (FL) enables collaborative medical image segmentation without centralizing sensitive data, real-world deployment i…
Upper Bounds on the Generalization Error of Deep Learning Models via Local Robustness and Stability
Generalization is a critical property of data-driven models, particularly deep learning models deployed in safety-critical applications. Ro…
Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering
Aggregate accuracy benchmarks conceal a systematic structure in how large language models fail at electronic health record (EHR) question a…
Beyond Weights and Gradients: A Taxonomy of Federated Learning Messages
Federated Learning is rapidly evolving beyond the exchange of traditional model weights and gradients, yet existing definitions fail to cap…
Semantic Flip: Synthetic OOD Generation for Robust Refusal in Embodied Question Answering and Spatial Localization
Detecting unanswerable user queries remains essential for the reliable deployment of real-world embodied agents. However, modern vision-lan…
Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models
This work addresses spatial question answering for service robots traversing long egocentric routes. Given a query such as "where can I fin…
IMPACTeen: Intentions, Manipulation, Persuasion, Annotations, and Consequences in Teen Communication Dataset
IMPACTeen is a dataset of textual social influence scenarios spanning interpersonal, media-based, and digital settings in an adolescent con…
Demystifying Variance in Circuit Discovery of LLMs
Circuit discovery is a key technique in mechanistic interpretability to pinpoint the model components that are crucial for performing a giv…
A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning
Reinforcement learning (RL) systems often degrade when operating conditions differ from those previously encountered, reflecting distributi…
CrossMaps: Confidence-Aware Open-Vocabulary Semantic Mapping for Rover Navigation
Rovers rely on perception to maintain spatial maps that encode both objects and sensor quality (e.g., range reliability, lighting artifacts…
Scalable Circuit Learning for Interpreting Large Language Models
A prominent research direction in mechanistic interpretability is learning sparse circuits over LLM components to reveal how they jointly p…
Phantoms and Disclosures: a Causal Framework for Auditing Synthetic Data
The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alterna…
Probing Low Frame Rate Degradation in Neural Audio Codecs
Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where the generation cost scales linearly with t…
How Much Do Reviews Really Contribute? A Study on Text-Enriched Matrix Factorization for Recommendations
Incorporating textual reviews into a Recommender System has become a prominent strategy for enriching collaborative signals with semantic i…
Stable Menus of Public Goods: AI-Enabled Progress
Using an open problem from the EC 2025 paper "Stable Menus of Public Goods" as a testbed, we conduct experiments to understand the effectiv…
ActiveSAM: Image-Conditional Class Pruning for Fast and Accurate Open-Vocabulary Segmentation
Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-voca…
TuneJury: An Open Metric for Improving Music Generation Preference Alignment
We introduce TuneJury, an open, instance-level pairwise reward model for text-to-music that predicts a music preference score from a text p…
TokenPilot: Cache-Efficient Context Management for LLM Agents
As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approaches utilize text pruni…
FusionRS: A Large-Scale RGB-Infrared Remote Sensing Dataset for Dual-Modal Vision-Language Foundation Models
Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery…
HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting
Simple linear and frequency-domain models remain surprisingly competitive in long-horizon time-series forecasting, and recent mechanistic e…
The Importance of Phase in Neural Representations: An Internal Oppenheim-Lim Test of Image Classifiers
Oppenheim and Lim (1981) showed that natural images stay recognizable when reconstructed from their Fourier phase alone, while the magnitud…
Attention, not scale, drives human-AI alignment in multimodal language prediction
Humans routinely draw on visual context to predict upcoming words. To what extent current vision-language models produce comparable behavio…
Metacognitive Myopia in Large Language Models
Large Language Models (LLMs) exhibit potentially harmful biases that reinforce culturally embedded stereotypes, influence moral judgments,…
Multi-Sensor Fusion for UAV Classification Based on Feature Maps of Image and Radar Data
The unique cost, flexibility, speed, and efficiency of modern UAVs make them an attractive choice in many applications in contemporary soci…
Computational Safety for Generative AI: A Hypothesis Testing Perspective
AI safety is a rapidly growing area of research that seeks to prevent the harm and misuse of frontier AI technology, particularly with resp…
Fine-Tuning a 7B Advisor on Free-Tier GPUs: An Adapter-Handoff Recipe and a Synthetic-Data Reliability Caution
Fine-tuning a 7B language model for specialized advising is attractive in resource-constrained settings, but multi-epoch runs routinely exc…
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving
Large language models (LLMs) have shown promising first-order logic (FOL) reasoning capabilities with applications in various areas. Howeve…
Unifying Post-hoc Explanations of Knowledge Graph Completions
Knowledge Graphs organize information as entity-relation-entity triples, enabling machine learning models to predict plausible missing trip…
Optimizing Health Coverage in Ethiopia: A Learning-augmented Approach and Persistent Proportionality Under an Online Budget
As part of nationwide efforts aligned with the United Nations' Sustainable Development Goal 3 on Universal Health Coverage, Ethiopia's Mini…
Shachi: A Modular, Controllable Framework for LLM-Based Agent-Based Modeling of Emergent Collective Behavior
How collective behaviors emerge from the interactions of individual LLM-driven agents is a central question in artificial life, yet control…
JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
Standard LLM evaluation practices compress diverse abilities into single scores, obscuring their inherently multidimensional nature. We pre…
Dual-Uncertainty Guided Policy Learning for Multimodal Reasoning
Reinforcement learning with verifiable rewards (RLVR) has advanced reasoning capabilities in multimodal large language models. However, exi…
PISA: A Pragmatic Psych-Inspired Unified Memory System for Enhanced AI Agency
Memory systems are fundamental to AI agents, yet existing work often lacks adaptability to diverse tasks and overlooks the constructive and…
Sample from What You See: Visuomotor Policy Learning via Diffusion Bridge with Observation-Embedded Stochastic Differential Equation
Imitation learning with diffusion models has advanced robotic control by capturing the multi-modal action distributions. However, existing…
Interpretation as Linear Transformation: A Cognitive-Geometric Model of Concepts and Meaning
This paper develops a geometric framework for modeling concepts, motivation, and influence across cognitively heterogeneous agents. Each ag…
Multi-Granular Node Pruning for Causal Circuit Discovery
Circuit discovery aims to identify minimal subnetworks that are responsible for specific behaviors in large language models (LLMs). Existin…
MedAI: Evaluating TxAgent's Therapeutic Agentic Reasoning in the NeurIPS CURE-Bench Competition
Therapeutic decision-making in clinical medicine constitutes a high-stakes domain in which AI guidance interacts with complex interactions…
Discovering Symmetry Groups with Flow Matching
Symmetry is fundamental to understanding physical systems and can improve performance and sample efficiency in machine learning. Both pursu…
DynaDebate: Breaking Homogeneity in Multi-Agent Debate with Dynamic Path Generation
Recent years have witnessed the rapid development of Large Language Model-based Multi-Agent Systems (MAS), which excel at collaborative dec…
E-mem: Multi-agent based Episodic Context Reconstruction for LLM Agent Memory
The evolution of Large Language Model (LLM) agents towards System~2 reasoning, characterized by deliberative, high-precision problem-solvin…
Edit Knowledge, Not Just Facts via Multi-Step Reasoning over Background Stories
Enabling artificial intelligence systems, particularly large language models, to update knowledge and flexibly apply it during reasoning re…
RaBiT: Residual-Aware Binarization Training for Accurate and Efficient LLMs
Efficient deployment of large language models (LLMs) requires extreme quantization, forcing a critical trade-off between low-bit efficiency…
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide ri…
ToolSelf: Unifying Task Execution and Self-Reconfiguration via Tool-Driven Emergent Adaptation
LLM-powered agentic systems excel at complex long-horizon tasks, but remain constrained by static configurations fixed before execution. Su…
An Attention Mechanism for Robust Multimodal Integration in a Global Workspace Architecture
Robust multimodal systems must remain effective when some modalities are noisy, degraded, or unreliable. Existing multimodal fusion methods…
AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems
Multi-agent Large Language Model (LLM) systems create privacy risks that current output-only benchmarks cannot measure. When agents coordin…
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Agent Skills are structured packages of procedural knowledge that augment large language model (LLM) agents at inference time. Despite rapi…
LLM-WikiRace Benchmark: How Far Can LLMs Plan over Real-World Knowledge Graphs?
We introduce LLM-Wikirace, a benchmark for evaluating planning, reasoning, and world knowledge in large language models (LLMs). In LLM-Wiki…
WorkflowPerturb: Calibrated Stress Tests for Evaluating Multi-Agent Workflow Metrics
Multi-agent LLM systems that generate structured workflows from natural-language requests are now deployed in production across cloud autom…
The Initial Exploration Problem in Knowledge Graph Exploration
Knowledge Graphs (KGs) enable the integration and representation of complex information across domains, but their semantic richness and str…
A Model-Free Universal AI
In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using enviro…
MemPO: Self-Memory Policy Optimization for Long-Horizon Agents
Long-horizon agents face the challenge of growing context size during interaction with environment, which degrades the performance and stab…
SorryDB: Can AI Provers Complete Real-World Lean Theorems?
We present SorryDB, a dynamically-updating benchmark of open Lean tasks drawn from 78 real world formalization projects on GitHub. Unlike e…
Rescaling Confidence: What Scale Design Reveals About LLM Metacognition
Verbalized confidence, in which LLMs report a numerical certainty score, is widely used to estimate uncertainty in black-box settings, yet…
Beyond Scalars: Evaluating and Understanding LLM Reasoning via Geometric Progress and Stability
Evaluating LLM reliability via scalar probabilities often fails to capture the structural dynamics of reasoning. We introduce TRACED, a fra…
Protein Design with Agent Rosetta: A Case Study for Specialized Scientific Agents
Large language models (LLMs) are capable of emulating reasoning and using tools, creating opportunities for autonomous agents that execute…
EMS: Multi-Agent Voting via Efficient Majority-then-Stopping
Majority voting is the standard for aggregating multi-agent responses into a final decision. However, traditional methods typically require…
Beyond Predefined Schemas: TRACE-KG for Context-Enriched Knowledge Graph Generation
Knowledge graph generation typically relies either on predefined ontologies or on schema-free extraction. Ontology-driven pipelines enforce…
When Do We Need LLMs? A Diagnostic for Language-Driven Bandits
We study Contextual Multi-Armed Bandits (CMABs) for non-episodic decision-making problems where the context includes both textual and numer…
The Missing Knowledge Layer in Cognitive Architectures for AI Agents
The two most influential cognitive architecture frameworks for AI agents, CoALA [21] and JEPA [12], both lack an explicit Knowledge layer w…
Emergent Strategic Reasoning Risks in AI: A Taxonomy-Driven Evaluation Framework
As reasoning capacity and deployment scope grow in tandem, large language models (LLMs) gain the capacity to engage in behaviors that serve…
Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy
This paper develops Virtual Speech Therapist (VST), an intelligent agent-based platform that streamlines stuttering assessment and delivers…
The Model Knows, the Decoder Finds: Future Value Guided Particle Power Sampling
A recurring pattern in "reasoning without training" is that base LLMs already assign non-trivial probability mass to correct multi-step sol…
FORTIS: Benchmarking Over-Privilege in Agent Skills
Large language model agents increasingly operate through an intermediate skill layer that mediates between user intent and concrete task ex…
LLM Jaggedness Unlocks Scientific Creativity
As artificial intelligence advances, models are not improving uniformly. Instead, progress unfolds in a jagged fashion, with capabilities g…
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To…
Boosting Knowledge Graph Foundation Models via Enhanced Negative Sampling
Knowledge graphs (KGs) have become the core backbone of numerous downstream tasks such as question answering and recommender systems. Howev…
SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search
Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness…
Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture
Large language models are undergoing a transition from model technology to system technology. Engineering challenges like cache reuse, cont…
Early Diagnosis of Wasted Computation in Multi-Agent LLM Systems via Failure-Aware Observability
Failure-aware observability diagnoses wasted computation in multi-agent LLM systems before final-answer evaluation can explain what went wr…
S-SPPO: Semantic-Calibrated Self-Play Preference Optimization
Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the sta…
Decision-Aware Memory Cards: Counterfactual-Inspired Context Selection and Compression for Tool-Using LLM Agents
Modern large language model (LLM) agents do not simply need longer contexts; they need decision-relevant evidence at the moment of action.…
Agent Economics: An Entropy-Controlled Pluralistic Alignment Framework for Preventing Artificial Hivemind in Autonomous Agents
This study proposes the Behavioral Protocol Framework (BPF), an entropy-controlled pluralistic alignment framework designed to address two…
Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory
Medical agent systems are increasingly expected to support interactive clinical decision making rather than only static question answering.…
Deterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics Architecture
As autonomous research agents and AI co-scientist systems push large language models (LLMs) from drafting toward end-to-end manuscript prod…
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical wor…
Minimalist Genetic Programming
Genetic programming (GP) is based on two important insights. First, that any learning task can fundamentally be posed as a program inductio…
When the Chain of Thought Knows Better: Failure Modes in Multi-Turn Reasoning Models
Failures in multi-turn reasoning models are largely invisible to terminal-score evaluation. A model can lock onto an unsafe stance early in…
Knowing When to Ask: Self-Gated Clarification for Hierarchical Language Agents
In hierarchical reasoning, failures often originate at intermediate decision points where the agent commits to a wrong branch without recog…
Human-Enhanced Loop Modeling (HELM): Agent-Based Finite Element Modeling of Concrete Bridge Barriers
Finite element (FE) modeling of safety-critical infrastructure such as bridge barriers requires high-fidelity nonlinear dynamic analysis, y…
The Illusion of Multi-Agent Advantage
Prevailing wisdom posits that Multi-Agent Systems (MAS) are superior to Single-Agent Systems (SAS), citing advantages like context protecti…
Why Sampling Is Not Choosing: Intentionality, Agency, and Moral Responsibility in Large Language Models
Recent advances in large language models (LLMs) have prompted claims that such systems exhibit agency or qualify as moral agents. This pape…
Reasoning as Pattern Matching: Shared Mechanisms in Human and LLM Everyday Reasoning
When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not…
AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility
Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric har…
Multi-Grade Deep Learning for Partial Differential Equations with Applications to the Burgers Equation
Deep neural networks (DNNs) show great promise for solving partial differential equations (PDEs), but their deep architectures introduce co…
It's About Time: Temporal References in Emergent Communication
Emergent communication enables agents to develop bespoke languages that improve communication efficiency. Despite the known importance of t…
A Survey on 3D Skeleton Based Person Re-Identification: Taxonomy, Advances, Challenges, and Interdisciplinary Prospects
Person re-identification via 3D skeletons is an important emerging research area that attracts increasing attention within the pattern reco…
Deep Neural Networks: A Formulation Via Non-Archimedean Analysis
We introduce a new class of deep neural networks (DNNs) with multilayered tree-like architectures. The architectures are codified using num…
No One-Size-Fits-All Neurons: Task-based Neurons for Artificial Neural Networks
In the past decade, many successful networks are on novel architectures, which almost exclusively use the same type of neurons. Recently, m…
Canonical Variates in Wasserstein Metric Space
In this paper, we address the classification of instances represented by distributions on a vector space rather than single points. We cons…
Mitigating scalability challenges in LUT-based neural networks via pruning optimisations
Modern deep neural networks heavily rely on a large number of multiply-accumulate operations, which constitute the predominant computationa…
Learning in the Recurrent State: Gradient Descent with Linear Recurrent Networks
Linear recurrent networks (LRNNs) offer linear-time sequence modeling, but standard recurrent updates do not directly expose the supervised…
Explainable deep learning improves human mental models of self-driving cars
Self-driving cars increasingly rely on deep neural networks to achieve human-like driving. The opacity of such black-box planners makes it…
Virtual Sensing to Enable Real-Time Monitoring of Inaccessible Locations & Unmeasurable Parameters
Real-time monitoring of safety-critical interior states remains an open problem in energy systems where physical instrumentation is infeasi…
Understanding, Detecting, and Repairing Real-World In-Context-Learning-Based Text-to-SQL Errors
Large language models (LLMs) have been adopted for text-to-SQL tasks, utilizing their in-context learning (ICL) capability to translate nat…
Dealing with Annotator Disagreement in Hate Speech Classification
Hate speech detection is a crucial task, especially on social media where harmful content can spread quickly. Collecting social media conte…
Region-Adaptive Sampling for Diffusion Transformers
Diffusion models (DMs) have become the leading choice for generative tasks across diverse domains. However, their reliance on multiple sequ…
Bridging the Gap: Enabling Natural Language Queries for NoSQL Databases through Text-to-NoSQL Translation
NoSQL databases are core data infrastructure, yet natural-language access to them remains underdeveloped: correct query generation must rec…
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day. Each request requires t…
PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks
Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexpl…
Efficient Flow Matching using Latent Variables
Flow matching models have shown great potential in image generation tasks among probabilistic generative models. However, most flow matchin…
Optimal Transport for Machine Learners
Modern machine learning repeatedly manipulates probability measures: empirical datasets, generated samples, latent distributions, class-con…
RIDGECUT: Learning Graph Partitioning with Rings and Wedges
Reinforcement learning (RL) has shown promise for combinatorial optimization problems on graphs by learning heuristics that generalize acro…
Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
In Transformer architectures, tokens\textemdash discrete units derived from raw data\textemdash are formed by segmenting inputs into fixed-…
Mosaic: Data-Free Knowledge Distillation via Mixture-of-Experts for Heterogeneous Distributed Environments
Federated Learning (FL) is a decentralized machine learning paradigm that enables clients to collaboratively train models while preserving…
Multiple Descents in Deep Learning as a Sequence of Order-Chaos Transitions in LSTM Networks
We observe a novel `multiple-descent' phenomenon during the learning process of a recurrent neural network called long-short-term memory (L…
AC-ODM: Actor--Critic Online Data Mixing for Sample-Efficient LLM Pretraining
Optimizing pretraining data composition is pivotal for LLM generalization. While dynamic mixing outperforms static strategies by capturing…
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text,…
CLoVE: Personalized Federated Learning through Clustering of Loss Vector Embeddings
We propose CLoVE (Clustering of Loss Vector Embeddings), a novel algorithm for Clustered Federated Learning (CFL). In CFL, clients are natu…
FOUNDv2: Learning Unified User Quantized Tokenizers for User Representation
User representation learning serves as a fundamental pillar for personalized services on large-scale web platforms. Despite its importance,…
MedSynth: Realistic, Synthetic Medical Dialogue-Note Pairs
Physicians spend significant time documenting clinical encounters, a burden that contributes to professional burnout. To address this, robu…
Automated ultrasound doppler angle estimation using deep learning
Angle estimation is an important step in the Doppler ultrasound clinical workflow to measure blood velocity. It is widely recognized that i…
FlowState: Sampling-Rate-Equivariant Time-Series Forecasting
Existing time series foundation models (TSFMs), often based on transformer variants, lack adaptability to different sampling rates, struggl…
Retro-Expert: Collaborative Reasoning for Interpretable Retrosynthesis
Retrosynthesis prediction aims to infer the reactant molecules based on a given product molecule, which is a fundamental task in chemical s…
A biological vision inspired framework for machine perception of abutting grating illusory contours
Higher levels of machine intelligence demand alignment with human perception and cognition. Deep neural networks (DNN) dominated machine in…
EEG-FM-Bench: A Comprehensive Benchmark for the Systematic Evaluation and Diagnostic Analyses of EEG Foundation Models
Electroencephalography foundation models (EEG-FMs) have advanced brain signal analysis, but the lack of standardized evaluation benchmarks…
Prototyping an AI-powered Tool for Energy Efficiency in New Zealand Homes
Residential buildings contribute significantly to energy use, health outcomes, and carbon emissions. In New Zealand, housing quality has hi…
Beyond Rebalancing: Benchmarking Binary Classifiers Under Class Imbalance Without Rebalancing Techniques
Class imbalance poses a significant challenge to supervised classification, particularly in critical domains like medical diagnostics and a…
Discrete optimal transport is a strong audio adversarial attack
In this paper, we investigate discrete optimal transport (DOT) as a black-box attack against modern automatic speaker verification (ASV) an…
K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model
Medical image segmentation is fundamental to clinical decision-making, yet existing models remain fragmented. They are usually trained on s…
Projection and Quantisation: A Unifying View of Learning to Hash, from Random Projections to the RAG Era
Approximate nearest-neighbour search underpins large-scale retrieval and retrieval-augmented generation, yet its methods are studied in com…
Why Low-Precision Transformer Training Fails: An Analysis on Flash Attention
The pursuit of computational efficiency has driven the adoption of low-precision formats for training transformer models. However, this pro…
A Survey on Agentic Security: Applications, Threats and Defenses
LLM-based agents are now used throughout cybersecurity. While these agents facilitate powerful and autonomous security applications, their…
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
Large language models (LLMs) with extended context windows enable powerful applications but impose significant memory overhead, as caching…
Less is More: Improving LLM Reasoning with Minimal Test-Time Intervention
Recent progress in large language models (LLMs) has focused on test-time scaling to improve reasoning via increased inference computation,…
Utility-Diversity Aware Online Batch Selection for LLM Supervised Fine-tuning
Supervised fine-tuning (SFT) is a commonly used technique to adapt large language models (LLMs) to downstream tasks. In practice, SFT on a…
A Multi-level Analysis of Factors Associated with Student Performance: A Machine Learning Approach to the SAEB Microdata
Identifying the factors that influence student performance in basic education is a central challenge for formulating effective public polic…
AIRMap: AI-Generated Radio Maps for Wireless Digital Twins
Accurate, low-latency channel modeling is essential for real-time wireless network simulation and digital-twin applications. Traditional mo…
Think-at-Hard: Selective Latent Iterations to Improve Reasoning Language Models
Improving the reasoning abilities of Large Language Models (LLMs), especially under parameter constraints, is crucial for real-world applic…
Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility
Malicious AI causing harm to humans is not just a Hollywood fantasy. Indeed, as highly capable models such as Claude Mythos emerge and agen…
Can Artificial Intelligence Accelerate Technological Progress? Researchers' Perspectives on AI in Manufacturing and Materials Science
Artificial intelligence (AI) raises expectations of substantial increases in rates of technological progress, but such anticipations are of…
SPI: Query-Depth-Adaptive Indexing for Streaming RAG in Vector Databases
Vector databases (VecDBs) are increasingly deployed in retrieval-augmented generation (RAG) pipelines where query processing and document i…
DualGauge: Automated Joint Security-Functionality Benchmarking of Specification-Only Code Generation by LLMs and Coding Agents
Large language models (LLMs) and LLM-based coding agents are now used to generate code from natural-language specifications, yet ensuring s…
Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
In the age of agentic AI, the growing deployment of multi-modal models (MMs) has introduced new attack vectors that can leak sensitive trai…
CycliST: A Video Language Model Benchmark for Reasoning on Cyclical State Transitions
We present CycliST, a novel benchmark dataset designed to evaluate Video Language Models (VLM) on their ability for textual reasoning over…
Near--Real-Time Conflict-Related Fire Detection in Sudan Using Unsupervised Deep Learning
Ongoing armed conflict in Sudan highlights the need for rapid monitoring of conflict-related fire-affected areas. Recent advances in deep l…
AL-GNN: Privacy-Preserving and Replay-Free Continual Graph Learning via Analytic Learning
Continual graph learning (CGL) aims to enable graph neural networks to incrementally learn from a stream of graph structured data without f…
Do You Really Need a GPU to Guard Your LLM? CPU-Class Classifiers and Multi-Stage Pipelines for Safety Enforcement at Scale
Safety classifiers that screen LLM inputs for jailbreak attempts have become standard deployment components, yet almost all production syst…
A Unified Definition of Hallucination: It's The World Model, Stupid!
Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's…
Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
Speculative decoding (SD) accelerates LLM inference by verifying draft tokens in parallel. However, this method presents a critical trade-o…
RollArt: Disaggregated Multi-Task Agentic RL Training at Scale
Agentic Reinforcement Learning (RL) trains LLMs through multi-turn interactions with environments, producing workloads that mix compute-bou…
FasterPy: An LLM-based Code Execution Efficiency Optimization Framework
Code often suffers from performance bugs. These bugs necessitate the research and practice of code optimization. Traditional rule-based met…
Akasha 2: Hamiltonian State Space Duality and Visual-Language Joint Embedding Predictive Architectur
We present Akasha 2, a state-of-the-art multimodal architecture that integrates Hamiltonian State Space Duality (H-SSD) with Visual-Languag…
Critically Engaged Pragmatism: Scientific Norm and Social, Pragmatist Epistemology for AI Science Evaluation Tools
AI science evaluation tools aim to assess research credibility. As with traditional metrics such as impact factors, their edicts can be dec…
SDFLoRA: Selective Decoupled Federated LoRA for Privacy-preserving Fine-tuning with Heterogeneous Clients
Federated learning (FL) for large language models (LLMs) has attracted increasing attention as a privacy-preserving approach for adapting m…
Adaptive $k$NN graph model
The $k$-nearest neighbors ($k$NN) algorithm is a cornerstone of non-parametric classification in artificial intelligence, yet its deploymen…
Safe Exploration via Policy Priors
Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. simulated)…
AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code Completion
Repository-level code completion remains a challenging task for existing code large language models (code LLMs) due to their limited unders…
Sustainable Materials Discovery in the Era of Artificial Intelligence
Artificial intelligence (AI) has transformed materials discovery, enabling rapid exploration of chemical space through generative models an…
MapDream: Task-Driven Map Learning for Vision-Language Navigation
Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating…
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
While Retrieval-Augmented Generation (RAG) is one of the dominant paradigms for enhancing Large Vision-Language Models (LVLMs) on knowledge…
EffGen: Enabling Small Language Models as Capable Autonomous Agents
Most existing language model agentic systems today are built and optimized for large language models (e.g., GPT, Claude, Gemini) via API ca…
SLUM-i: Semi-supervised Learning for Urban Mapping of Informal Settlements and Data Quality Benchmarking
Rapid urban expansion has fueled the growth of informal settlements in major cities of low- and middle-income countries, with Lahore and Ka…
Learning to Share: Selective Memory for Efficient Parallel Agentic Systems
Agentic systems solve complex tasks by coordinating multiple agents that iteratively reason, invoke tools, and exchange intermediate result…
Seeing Roads Through Words: A Language-Guided Framework for RGB-T Driving Scene Segmentation
Robust semantic segmentation of road scenes under adverse illumination, lighting, and shadow conditions remain a core challenge for autonom…
MUZZLE: Adaptive Agentic Red-Teaming of Web Agents Against Indirect Prompt Injection Attacks
Large language model (LLM) based web agents are increasingly deployed to automate complex online tasks by directly interacting with web sit…
TS-Memory: Plug-and-Play Memory for Time Series Foundation Models
Time Series Foundation Models (TSFMs) achieve strong zero-shot forecasting through large-scale pre-training, but adapting them to downstrea…
UniT: Unified Multimodal Chain-of-Thought Test-time Scaling
Unified models can handle both multimodal understanding and generation within a single architecture, yet they typically operate in a single…
Orcheo: A Modular Full-Stack Platform for Conversational Search
Conversational search (CS) requires a complex software engineering pipeline that integrates query reformulation, ranking, and response gene…
Revisiting Chebyshev Polynomial and Anisotropic RBF Models for Tabular Regression
Smooth-basis models such as Chebyshev polynomial regressors and radial basis function (RBF) networks are well established in numerical anal…
MedCollab: IBIS-Guided Multi-Agent Collaboration with Hierarchical Disease Relation Chains for Clinical Diagnosis
Clinical diagnosis is a gradual process of evidence integration, in which physicians move from symptoms and medical history to examinations…
Cross-modal Identity Mapping: Minimizing Information Loss in Modality Conversion via Reinforcement Learning
Large Vision-Language Models (LVLMs) often omit or misrepresent critical visual content in generated image captions. Minimizing such inform…
Parallel Test-Time Scaling with Multi-Sequence Verifiers
Parallel test-time scaling, which generates multiple candidate solutions for a single problem, is a powerful technique for improving large…
WavSLM: Single-Stream Speech Language Modeling via WavLM Distillation
Large language models show that simple autoregressive training can yield scalable and coherent generation, but extending this paradigm to s…
An Empirical Investigation of Pre-Trained Deep Learning Model Reuse in the Scientific Process
Deep learning has achieved recognition for its impact within natural sciences, yet the prohibitive financial and technical cost of training…
MAND: Modality-Aware Novelty Detection for Open-World Egocentric Activity Recognition
Multimodal egocentric activity recognition integrates visual and inertial cues for robust first-person behavior understanding. However, dep…
Learning Permutation Distributions via Reflected Diffusion on Ranks
The finite symmetric group S_n provides a natural domain for permutations, yet learning probability distributions on S_n is challenging due…
Rel-Zero: Harnessing Patch-Pair Invariance for Robust Zero-Watermarking Against AI Editing
Recent advancements in diffusion-based image editing pose a significant threat to the authenticity of digital visual content. Traditional e…
Parallelizing Tool Execution and LLM Generation for Low-Latency Agent Serving
LLM-powered agents execute tasks through a sequential loop of model generation and tool execution. Today's serving systems serialize this l…
AgenticRec: A Recommendation-Oriented Agentic Framework with Progressive Tool-Integrated Reasoning Optimization
Recommender agents built on Large Language Models offer a promising paradigm for personalized recommendation. However, existing agents typi…
Closing the Auto-Research Loop: An AI Co-Scientist for Production Search Ranking
We present an AI Co-Scientist framework that closes the research loop for the production search-ranking system of a large online travel pla…
From Overload to Convergence: Supporting Multi-Issue Human-AI Negotiation with Bayesian Visualization
As AI systems increasingly mediate negotiations, understanding how the number of negotiated issues impacts human performance is crucial for…
A Learning Method with Gap-Aware Generation for Heterogeneous DAG Scheduling
Efficient scheduling of directed acyclic graphs (DAGs) is a core problem in large-scale data-intensive computing systems, where query plans…
Mitigating Object Hallucinations in LVLMs via Attention Imbalance Rectification
Object hallucination in Large Vision-Language Models (LVLMs) severely compromises their reliability in real-world applications, posing a cr…
Evidence of an Emergent "Self" in Continual Robot Learning
A key challenge to understanding self-awareness has been a principled way of quantifying whether an intelligent system has a concept of a "…
Epileptic Seizure Detection in Separate Frequency Bands Using Feature Analysis and Graph Convolutional Neural Network (GCN) from Electroencephalogram (EEG) Signals
Epileptic seizures are neurological disorders characterized by abnormal and excessive electrical activity in the brain, resulting in recurr…
Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains
We study the compression of LLM-generated text across lossless and lossy regimes, characterizing a compression-compute frontier where more…
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
Co-evolutionary self-play, where one language model generates problems and another solves them, promises autonomous curriculum learning wit…
Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA
Legal QA benchmarks have predominantly focused on case law, overlooking the unique challenges of statute-centric regulatory reasoning. In s…
Active Inference with a Self-Prior in the Mirror-Mark Task
The mirror self-recognition test evaluates whether a subject touches a mark on its own body that is visible only in a mirror, and is widely…
HCP-MAD:Heterogeneous Consensus-Progressive Reasoning for Efficient Multi-Agent Debate
Multi-Agent Debate (MAD) is a collaborative framework in which multiple agents iteratively refine solutions through the generation of reaso…
Adaptive Memory Crystallization for Autonomous AI Agent Learning in Dynamic Environments
Autonomous AI agents operating in dynamic environments face a persistent challenge: acquiring new capabilities without erasing prior knowle…
Human Cognition in Machines: A Unified Perspective of World Models
This report of world models distinguishes prior works by the cognitive functions they innovate. Many works claim an almost human-like cogni…
RoTRAG: Rule of Thumb Reasoning for Conversation Harm Detection with Retrieval-Augmented Generation
Detecting harmful content in multi turn dialogue requires reasoning over the full conversational context rather than isolated utterances. H…
Ranking Abuse via Strategic Pairwise Data Perturbations
Pairwise ranking systems based on Maximum Likelihood Estimation (MLE), such as the Bradley-Terry model, are widely used to aggregate prefer…
OmniMouse: Scaling properties of multi-modal, multi-task Brain Models on 150B Neural Tokens
Scaling data and artificial neural networks has transformed AI, driving breakthroughs in language and vision. Whether similar principles ap…
RSRCC: A Remote Sensing Regional Change Comprehension Benchmark Constructed via Retrieval-Augmented Best-of-N Ranking
Traditional change detection identifies where changes occur, but does not explain what changed in natural language. Existing remote sensing…
From Noise to Intent: Anchoring Generative VLA Policies with Residual Bridges
Bridging high-level semantic understanding with low-level physical control remains a persistent challenge in embodied intelligence, stemmin…
G-Loss: Graph-Guided Fine-Tuning of Language Models
Traditional loss functions, including cross-entropy, contrastive, triplet, and su pervised contrastive losses, used for fine-tuning pre-tra…
Lightweight Distillation of SAM 3 and DINOv3 for Edge-Deployable Individual-Level Livestock Monitoring and Longitudinal Visual Analytics
Foundation-model pipelines for individual-level livestock monitoring -- combining open-vocabulary detection, promptable video segmentation,…
CRC-Screen: Certified DNA-Synthesis Hazard Screening Under Taxonomic Shift
DNA-synthesis providers screen incoming orders by searching the requested sequence against curated hazard lists. We show that this baseline…
BRITE: A Benchmark for Reliable and Interpretable T2V Evaluation on Implausible Scenarios
The rapid advancement of photorealistic Text-to-Video (T2V) generation brings in an urgent need for up-to-date evaluation methods. Existing…
StyleShield: Exposing the Fragility of AIGC Detectors through Continuous Controllable Style Transfer
AI-generated content (AIGC) detectors are increasingly deployed in high-stakes settings such as academic integrity screening, yet their rel…
GEASS: Gated Evidence-Adaptive Selective Caption Trust for Vision-Language Models
Vision-Language Models (VLMs) hallucinate objects that are not present, and a growing line of work tries to curb this by feeding the model…
Gated QKAN-FWP: Scalable Quantum-inspired Sequence Learning
Fast Weight Programmers (FWPs) encode temporal dependencies through dynamically updated parameters rather than recurrent hidden states. Qua…
Trust Without Trusting: A Recomputable Trust Protocol for Autonomous Agents
Autonomous AI agents already transact at production scale -- 69,000 bots, 165 million transactions, $50 million in volume on a single marke…
Prediction Bottlenecks Don't Discover Causal Structure (But Here's What They Actually Do)
A Mamba state-space model trained only for next-step prediction appears to recover Granger-causal structure through a simple readout $S = |…
From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs
Large-scale AI training is fundamentally a distributed systems problem, where hardware failures are routine operating conditions rather tha…
Red-Teaming Agent Execution Contexts: Open-World Security Evaluation on OpenClaw
Agentic language-model systems increasingly rely on mutable execution contexts, including files, memory, tools, skills, and auxiliary artif…
TERMS-Bench: Diagnosing LLM Negotiation Agents Beyond Deal Rate
Negotiation is a central mechanism of economic exchange, shaping markets, procurement, labor agreements, and resource allocation. It is als…
Wasserstein Equilibrium Decoding for Reliable Medical Visual Question Answering
Small vision-language models (2-8B) are well-suited for clinical deployment due to privacy constraints, limited connectivity, and low-laten…
Improved Baselines with Representation Autoencoders
Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders. In this paper, we systematically investigate sev…
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Long-horizon LLM agents generate traces that could become reusable experience, but raw trajectories are noisy, local, and hard to govern. A…
EvoMemBench: Benchmarking Agent Memory from a Self-Evolving Perspective
Recent benchmarks for Large Language Model (LLM) agents mainly evaluate reasoning, planning, and execution. However, memory is also essenti…
Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs
Enterprise analytics aims to make organizational data accessible for decision-making, yet non-technical users still face barriers when usin…
DySink: Dynamic Frame Sinks for Autoregressive Long Video Generation
Autoregressive long video generation often adopts bounded-memory streaming for efficiency, typically combining local windows for short-term…
Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
Modern LLM serving is no longer homogeneous or monolithic. Production systems now combine disaggregated execution, complex parallelism, run…
Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build
How much have students' ordinary learning processes shifted in response to generative AI, and how does that affect their durable learning o…
ACC: Compiling Agent Trajectories for Long-Context Training
Recent development of agents has renewed demand for long-context reasoning capacity of LLMs. However, training LLMs for this capacity requi…
Action with Visual Primitives
Vision-Language-Action (VLA) models have emerged as a promising paradigm for generalist robotic manipulation. A common design in current ar…
When Do LLMs Reason? A Dynamical Systems View via Entropy Phase Transitions
Chain-of-thought (CoT) reasoning has become the default strategy for enhancing LLM capabilities, yet its application raises a fundamental q…
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustne…
When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control
A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost acros…
Cordyceps: Covert Control Attacks on LLMs via Data Poisoning
Large language models (LLMs) are often fine-tuned on uncurated text datasets that adversaries can poison. Existing poisoning attacks primar…
FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies
Vision-Language-Action (VLA) models are increasingly expected to not only complete robot tasks, but also follow human instructions about ho…
The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
Agentic AI workloads - where a single user goal triggers multi-step orchestration, tool calls, retries, and failure recovery - are being ta…
Evolutionary Dynamics of Cooperation in Next-Generation LLM Agent Systems: A Cross-Provider Empirical Extension
Do next-generation LLM agents inherit the cooperative biases documented in their predecessors, or does scale and provider diversity reshape…
Automating Low-Risk Code Review at Meta: RADAR, Risk Calibration, and Review Efficiency
AI-assisted coding tools have altered software production. At Meta, significant lines of code per human-landed diff grew by 105.9% year ove…
Detect Before You Leap: Mirage Detection in Vision-Language Models
Vision-language models (VLMs) can produce confident visual answers even when the required visual evidence is missing, blank, or unrelated t…
Estimating Mutual Information between Time Series and Temporal Event Sequences Across Diverse Analysis Tasks
Pairwise dependence measures such as correlation and causality are fundamental to temporal data mining, yet there is still no principled an…
TechRAG: Evidence-Gated Multimodal Agentic RAG for Technical Literature Reasoning
This paper presents an agentic multimodal retrieval-augmented generation (RAG) framework for domain-specific literature reasoning, instanti…
Anomalies in Multivariate Time Series Benchmarks Are Mostly Univariate
Many recent multivariate time series anomaly detection (MTSAD) models incorporate cross-channel modeling, under the implicit assumption tha…
Fast-dLLM++: Fr\'{e}chet Profile Decoding for Faster Diffusion LLM Inference
Diffusion large language models promise parallel token generation, yet inference remains bottlenecked by deciding which masked tokens can b…
Learn from Your Mistakes: Tree-like Self-Play for Secure Code LLMs
While Large Language Models (LLMs) excel in code generation, they remain prone to replicating subtle yet critical vulnerabilities endemic t…
EvalStop: Using World Feedback to Detect and Correct Reward Overoptimization in Multi-Tenant RLHF Platforms
Cloud LLM fine-tuning platforms increasingly serve RLHF workloads, where a learned reward model is optimized as a proxy for human quality.…
From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
Large language model (LLM)-based agents are evolving from passive text generators into autonomous systems capable of planning, tool use, re…
Benchmarking Counterfactual Prediction in Epidemic Time Series with Time-Varying Interventions
Deep learning has enabled significant advances in time-series causal inference, yet progress remains constrained by the lack of realistic b…
FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (June 13th version)
Conventional HPC holds that native hardware FP64 is the irreducible foundation of scientific computing. On AI-optimized GPUs of the NVIDIA…
AI-Driven Test Case Generation from Natural Language Requirements: A Survey of Techniques and Research Gaps
Software testing is critical for verifying that systems meet specified requirements, yet remains among the most time-consuming and expensiv…
CAF-Gen: A Multi-Agent System for Enriching Argumentation Structures
Formalizing complex reasoning from natural text is one of the central challenges in computational linguistics. It requires systems to under…
Towards Unified Song Generation and Singing Voice Conversion with Accompaniment Co-Generation
While song generation and singing voice conversion (SVC) have evolved significantly, they have long been developed isolated: the former lac…
On the Geometry of On-Policy Distillation
On-policy distillation (OPD) is increasingly used to improve large language model reasoning, but its training dynamics remain poorly unders…
From Privacy to Workflow Integrity: Communication-Graph Metadata in Autonomous Agent Interoperability
Agent-interoperability protocols such as A2A and MCP standardize what agents say to one another but assume address-based transport. Whether…
DEFINED: A Data-Efficient Computational Framework for Fine-Grained Creativity Assessment in Debate Scenarios
Human creativity has emerged as a critical competency in the era of large language models. Assessing creativity in complex, open-ended envi…
DOG-DPO:Dynamic Optimization in Geometry for Safety Alignment
Safety alignment for large language models relies on preference data, but current pipelines often train on large, redundant datasets. Exist…
Fast LLM-Based Semantic Filtering: From a Unified Framework to an Adaptive Two-Phase Method
Evaluating a natural-language yes/no predicate over a document corpus under an accuracy target - the semantic filter - is a cornerstone of…
An AI Security Agent for University ACMIS: Multi-Vector Threat Detection and Automated Response
University Academic Management Information Systems (ACMIS) are high-value targets for a wide spectrum of security threats including brute-f…
SceneConductor: 3D Scene Generation from a Single Image with Multi-Agent Orchestration
Generating complete 3D scenes from a single image requires inferring globally consistent geometry, object relationships, and environmental…
Few-shot Class-variable Incremental Audio Classification via Prototype Adaptation and Pseudo Class-variable Training
In the task of few-shot class-incremental audio classification, the number of classes is assumed to always increase without considering the…
The Distributed Detectability Band Against Marginal-Preserving Attacks
AI-control monitors score individual agent actions to detect misbehavior, but real harm can be distributed across many benign-looking steps…
LIBERO-Occ: Evaluating and Improving Vision-Language-Action Models under Scene-Induced Occlusion via Viewpoint Imagination
Vision-Language-Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task-r…
ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories
Training capable OS agents requires data that simultaneously captures structured user intents, multi-turn task delegation, and grounded too…
AnchorEdit: Maintaining Temporal Consistency in Multi-turn Image Editing via Causal Memory
Multi-turn image editing is essential for iterative design, yet current models often struggle with identity drift and error accumulation ov…
M*: A Modular, Extensible, Serving System for Multimodal Models
We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, d…
EV-WM: Event-Verified World Models for Long-Horizon Robotic Manipulation
Pretrained-feature world models provide a useful substrate for robot imagination, but visual or latent prediction alone does not determine…
LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories
Scientific laboratories increasingly rely on AI systems to reason about experiments, but the physical act of doing science remains largely…
Sundar Pichai faces boos, walkout at Stanford graduation ceremony over Google’s Israel, ICE ties
AI is once again at the heart of a college graduation protest — this time for the technology's use in Google's defense contracts.
急拡大するAIインフラの電力需要……光明は「ワットビット連携」に? さくら田中社長と東電が対談
AIインフラの拡大で急増する電力需要に、データセンターと電力網はどう向き合うべきか──6月10~12日に幕張メッセで開催された「Interop Tokyo 2026」の基調講演では、東京電力ホールディングスの岡本浩氏(上席フェロー)と、さくらインターネットの田中邦裕代表取締役社…
300億円は「ROI不問」 Olive、Trunkを仕掛けるSMBC、新規事業の神髄は「撤退」にアリ
「Olive」や「Trunk」を相次いで成長軌道に乗せ、生成AI活用に向けて500億円の投資計画も打ち出した三井住友フィナンシャルグループ。そんな同社だが、約10年前はモバイルアプリで競合他行に大きく後れを取るなど、変革が進んでいなかった。堅実なメガバンクは、いかに挑戦を次々と…
生成AI×3D CADでどこまでできるか試してみた
生成AIの活用は、文章や画像、動画だけでなく、3D CADの分野にも広がり始めています。自然言語で指示するだけで、3Dモデルのたたき台を作成できる環境も登場しつつあります。今回はAutodesk Fusionの「Autodesk Assistant」を使い、ペットボトルの3Dモ…
月2000時間のムダをなくす大阪ガスらのNotion×AI活用 「使われない情報」の生かし方
「あの資料はどこ」といった情報探索の負担を、NotionとAIの活用によって大幅に軽減した大阪ガスら2社の事例を紹介。月2000時間の業務削減を実現した取り組みによって埋もれた情報を組織の知識資産へと変え、属人化を防ぐ仕組みづくりのポイントを分かりやすく解説する。
The US government’s Anthropic models ban was never about an AI jailbreak
The Trump administration's decision that forced Anthropic to pull its latest cybersecurity models could be reactionary, retaliatory, or bot…
Meta’s new ‘AI Mode’ on Facebook pulls from public info across its platforms
Meta announced Monday that it's rolling out a wave of new AI features on Facebook, the latest sign of the company's effort to catch up in t…
SpaceX is public: Everything you need to know post-IPO
TechCrunch has followed SpaceX's start, struggles, and successes from the early days. And we're here for what happens next too. This packag…
Cybersecurity vets protest ‘dangerous’ US government ban on Anthropic’s most powerful models
A group made up of dozens of cybersecurity experts urged the White House to remove export-control restrictions on Anthropic’s Fable and Myt…
2026-06-15(271件)
Salesforce acquires AI customer service platform Fin for $3.6B
Salesforce says it wants to use Fin's team and technology to improve Agentforce, its existing enterprise platform that businesses can use t…
Sarvam becomes India’s newest AI unicorn with $234 million funding round led by HCLTech
Indian IT services company HCLTech is investing $150 million in the Bengaluru startup.
As AI agents become employees, NewCore emerges with $66M to give them identities
NewCore argues the next challenge in enterprise security will be managing AI agents, not people.
A satellite just learned to find things on its own — here’s what that means
In April, for the first time ever, an Earth observation satellite found what it was looking for, all on its own.
データセンターの見回り業務をロボットに 自在に伸びるカメラでくまなく点検できる「ugo mini」
6月10日から12日に幕張メッセで開催したインターネット技術の総合イベント「Interop Tokyo 2026」で、ugo(東京都千代田区)は小型の点検ロボット「ugo mini」を展示した。
人工知能学会「AIは人間を代替しない」 社会実装へ4提言 安保・著作権にも言及
人工知能学会は、設立40周年にあたり、日本におけるAIの社会実装に向けた提言を発表した。
The AI layoff wave is becoming a powder keg
At the very moment that tens of thousands of workers are being shown the door, a small cohort of AI insiders is becoming wealthy on a scale…
Javaアプリ更新を1カ月→3日に爆速化 “ソースコード生成AI止まり”じゃない「IBM Bob」の仕組み
IBMが発表したAIツール「IBM Bob」は、先行導入した企業でJavaアプリケーションのモダナイゼーション作業を30日から3日に短縮するといった効果があったという。ソースコード生成にとどまらない、IBM Bobの特徴とは。
A Deep Reinforcement Learning (DRL)-Based Transformer Method for Solving the Open Shop Scheduling Problem
The open shop scheduling problem (OSSP) arises in many industrial and service settings but remains computationally challenging as the numbe…
UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems
To address the challenge that current dialogue policy planning methods struggle to dynamically adapt to diverse user characteristics, this…
History of the Muddy Children Puzzle
The Muddy Children Puzzle is a puzzle about knowledge and ignorance that has been inspiring for the development of epistemic logic. Who cam…
Orchestra-o1: Omnimodal Agent Orchestration
The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-agent workflows to multi…
Hybrid Open-Ended Tri-Evolution Makes Better Deep Researcher
Deep research and agent evolution serve as de-facto tasks for AI agents in real-world applications toward artificial general intelligence.…
WorkBench Revisited: Workplace Agents Two Years On
The best agent on WorkBench in March 2024, GPT-4, completed 43% of tasks and took an unintended harmful action, such as emailing the wrong…
Refusal Beyond a Single Direction: A Preliminary Comparison of Diff-in-Means and INLP
Arditi et al. (2024) has shown that refusal in safety fine-tuned chat models is mediated by a single linear direction in the residual strea…
YeasierAgent: Agentic Social Sandbox as a Canvas for Intent-Driven Creation of Platform-Agnostic Symbiotic Agent-Native Applications
This paper introduces YeasierAgent, an application-building paradigm based on symbiotic agents, narrative worlds, and scene-aware interacti…
TwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards
Business intelligence (BI) increasingly combines dashboard interaction with LLM-based assistance, but these two modes often fall out of syn…
When Sample Selection Bias Precipitates Model Collapse
The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training ero…
AI Receptivity or AI Adoption Breadth? A Tool-Specific Reanalysis of the Lower-Literacy/Higher-Usage Link
Recent evidence reported by Tully, Longoni, and Appel (2025) suggests that lower artificial intelligence (AI) literacy predicts greater rec…
MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis
Large Language Models (LLMs) have made notable progress in automated theorem proving, yet existing formal benchmarks remain limited in both…
Poker Arena: Multi-Axis Profiling of Strategic Reasoning and Memory in LLMs
Strategic reasoning under uncertainty underpins consequential decisions in negotiation, finance, and policy, but prevailing game-play bench…
Hyperdimensional computing for structured querying on tabular data embeddings
Tabular data embeddings have become a cornerstone of data profiling and data integration pipelines, enabling tasks such as entity annotatio…
Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents
Modern decision systems increasingly rely on learned components whose outputs may be confident yet wrong, exposing downstream actions to co…
A Multi-Agent AI System for Automated High School Transcript Processing: Collaborative Document Analysis at Scale
Each year, college admissions offices face an overwhelming challenge: processing millions of high school transcripts, each with unique form…
Sorries Are Not the Hard Part: An Expert-Review Case Study of a Semi-Autonomous Formalization
Large language models can often close proof gaps in interactive theorem provers, but a verified theorem is not the same thing as a reusable…
Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry
Humans cannot always intuit what scenarios are most challenging to LLMs. Hoping to capture challenging edge cases, developers either design…
Minim: Privacy-Aware Minimal View for Agents via Trusted Local Sanitization
Modern LLM-powered autonomous agents increasingly rely on rich user interface (UI) state observations to achieve reliable action grounding…
Formalizing Numerical Analysis: An Agent Pipeline and Quality Audit Beyond Kernel Acceptance
Recent work has demonstrated that coding agents can formalize entire advanced mathematics textbooks in Lean 4, yet existing efforts concent…
Applicability Condition Extraction for Therapeutic Drug-Disease Relations
Identifying conditions that a certain drug takes therapeutic effect on a target disease is crucial for clinical decision-making support. Ho…
FactoryLLM: A Safe and Open-Source AI Playground for Evaluating LLMs in Smart Factories
Fault diagnostics and recovery in smart factories is challenging because critical information is dispersed across manuals of multiple machi…
VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verification
Geometry problem generation is useful for AI-assisted education and multimodal mathematical reasoning, but reliable synthesis remains diffi…
When Should Agent Trust Be Conditional? Characterizing and Attacking Skill-Conditional Reputation in Agent Swarms
Open platforms increasingly route tasks among heterogeneous LLM agents--differing in base model, scaffold, and tool stack--whose competence…
Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL
LLMs are increasingly deployed as agents that interact with external environments and observe feedback such as execution results, error mes…
SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing
Agent skills are structured procedural packages that guide frozen LLM agents in specialized workflows. Skills rarely remain sufficient afte…
AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties
Affordance reasoning, the inference of an object's action possibilities from its physical properties (e.g., shape and material), is fundame…
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a…
Communication Policy Evolution for Proactive LLM Agents
LLM agents have rapidly evolved into autonomous systems, yet a persistent information gap remains between users and agents: communication i…
CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning
Safe reinforcement learning (Safe RL) aims to maximize expected return while satisfying safety constraints, typically modeled as Constraine…
Causal Object-Centric Models for Planning with Monte Carlo Tree Search
We introduce COMET (Causal Object-centric Model for Efficient Tree search), a model-based reinforcement learning algorithm that performs Mo…
GitOfThoughts: Version-Controlled Reasoning and Agent Memory You Can Replay, Diff, and Merge
Large language model (LLM) reasoning is ephemeral: chains of thought vanish with the context window, pruned search branches leave no record…
When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More
A growing line of work equips large language model (LLM) agents with graph neural networks (GNNs) as callable tools, assuming the agent exe…
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable…
Dense Coordinate-List Fine-Tuning Induces a Controllable Interference Surface in Vision-Language Models
Fine-tuning vision-language models to emit dense coordinate lists improves visual grounding but also changes how models serialize, repeat,…
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that…
StreamMemBench: Streaming Evaluation of Agent Memory for Future-Oriented Assistance
A central role of personal-agent memory is to turn stored information and prior interactions into future-oriented assistance. In daily use,…
VISTA: View-Consistent Self-Verified Training for GUI Grounding
When applying Group Relative Policy Optimization (GRPO) for GUI Grounding, rollouts are sampled from a single screenshot view; groups often…
A Temporal Planning Framework for Disruption Aware Dynamic Route Optimization in Heterogeneous Railway Systems
Efficient route optimization play a vital role in ensuring both safety and punctuality in railway operations. It is very crucial particular…
Abstracting Cross-Domain Action Sequences into Interpretable Workflows
Sequential or time-stamped interaction logs provide objective records of digital application usage, yet their granularity and noise often o…
Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows
Large language models increasingly serve as execution engines for agentic systems, yet they still consume context through a sequential text…
GAGPO: Generalized Advantage Grouped Policy Optimization
Reinforcement learning has become a powerful paradigm for post-training large language model agents, yet credit assignment in multi-turn en…
Simplex-Constrained Sparse Bagging: Transitioning from Uniform Priors to Sparse Posteriors in Ensemble Learning
We present Simplex-Constrained Sparse Bagging (SCSB), a mathematically rigorous framework for post-training compression and probability cal…
Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs
Automatic Bloom's taxonomy classification of assessment questions can substantially reduce instructor workload, but labeling is subjective…
The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation
LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliabil…
An Agentic Retrieval Framework for Autonomous Context-Aware Data Quality Assessment
Data quality assessment is a critical prerequisite for effective data analytics and data-driven decision-making, yet it remains a challengi…
Efficient Temporal Modeling for Mobile Sleep Staging via Lightweight Random Attention
Mobile sleep staging serves as a foundational infrastructure for in-home sleep monitoring and closed-loop modulation. But existing sequenti…
Korzhinskii-Net: Physics-Informed Neural Network for Sub-Surface Mineral Prospectivity Modelling
Mineral prospectivity modelling (MPM) underpins exploration economics, yet most operational pipelines reduce to data-driven classifiers tra…
Active Inference for Adaptive Traffic Signal Control in Noisy Nonstationary IoT Environments
Urban traffic signal control at IoT-instrumented intersections must remain effective under sensor occlusion, weather attenuation, and nonst…
Position: AI Must Become Planet-Centered, Not Just Human-Centered
This position paper argues that contemporary AI paradigms are insufficient for supporting complex global goals and introduces Planet-Center…
Can Editing 1 Neuron Fix Repetition Loops in LLMs?
Yes. Can it cure doom loops? Probably not. The Gemma 4 instruction-tuned models share a reproducible failure: on long factual enumeration p…
HierSVA: A Data Synthesis Pipeline, Dataset, and Benchmark for LLM-Driven Hierarchical Hardware Formal Verification
We present HierSVA, an integrated suite that combines a pipeline, dataset, and benchmark for LLM-driven hierarchical hardware formal verifi…
CisTransCell: Single-Cell Perturbation Prediction via Gene Function, Regulatory Control, and Cellular Context
Predicting cellular transcriptional responses to genetic perturbations is a central problem in single-cell biology, especially in the zero-…
Morphology-Aware Sample Assignment: Overcoming IoU Insensitivity for Surface Defect Detection
Intersection-over-Union (IoU), as a pivotal metric for evaluating the spatial alignment between candidate proposals and ground-truth annota…
VHDLSuite: Unified Pipeline for LLM VHDL Generation with Data Synthesis and Evaluation
Large Language Models (LLM) have shown impressive capabilities in Register Transfer Level (RTL) code generation, particularly for Verilog.…
FreoStream:Enhancing Stream Guardrails via Future-Aware Reasoning and Safety-Aligned Optimization
Stream guardrails enable token-level safety detection before full responses are generated. However, they often make overly conservative jud…
A Virtuous AI is an Existential Risk
This paper examines trade-offs between AI safety and well-being relative to (i) one of the most promising methods for finetuning super-capa…
A fully GPU-based workflow for building physics emulators of hypersonic flows
The ability to resolve complex physical phenomena with high fidelity and at low computational cost is central to addressing key challenges…
The Weight Norm Sets the Grokking Timescale: A Causal Delay Law
Grokking is the delayed onset of generalization in neural networks, arising long after they fit the training data. Whether the weight norm…
Position: Align AI to Our Aspirations, Not Our Flaws
We argue that aligning AI to aggregated human preferences is the wrong target. With current technology, one can train AIs to share the valu…
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents
Large language model (LLM) reviewers are increasingly used in pull-request (PR) workflows, where their approvals help decide which code is…
Beyond LoRA: Is Sparsity-Induced Adaptation Better?
Low-rank adaptation (LoRA) and its variants provide a memory- and compute-efficient alternative to full fine-tuning of pre-trained models.…
CineOrchestra: Unified Entity-Centric Conditioning for Cinematic Video Generation
Cinematic video depicts multiple subjects acting or interacting at specific moments, captured with deliberate camera movement, and stitched…
An integrated interpretable control effectiveness learning and nonlinear control allocation methodology for overactuated aircrafts
Nonlinear dynamics and the strong couplings that arise between multiple effectors undermine the assumptions behind conventional, linear con…
A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets
Predictive code completion greatly accelerates how quickly developers work. In spreadsheets, despite being much more common, such auto-comp…
Aligning Quantum Operators with Large Language Models
Can Large Language Models (LLMs) understand and reason about quantum operators? Despite their remarkable capabilities in mathematics and sy…
AI can help scientists publish less
We can do more than defend science from a flood of AI-assisted papers. Used well, AI offers a historic opportunity to correct distortions i…
Safety-Contract Graph Multi-Agent Reinforcement Learning for Autonomous Network Security Response
Autonomous network-security response systems promise to reduce Security Operations Centre (SOC) reaction latency, but reward-only multi-age…
When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation
LLM-based generative agents are increasingly used in urban simulators, yet it remains unclear whether they reproduce empirically realistic…
Explaining RhythmFormer: A Systematic XAI Analysis of Periodic Sparse Attention for Remote Photoplethysmography
Remote photoplethysmography (rPPG) transformers achieve low heart-rate error on benchmarks, yet their decisions remain opaque--a growing co…
SpheriCity: Designing Trustworthy Conversational AI for Sustainability Decision Support
We present SpheriCity, an expert-grounded conversational prototype designed to support trustworthy knowledge sensemaking from sustainabilit…
Mood-Aware Music Recommendation: Integrating User Affective Signals into Ranking Systems
Recommendation systems are essential in modern music streaming platforms due to the vast amount of available content. While collaborative f…
SuperThoughts: Reasoning Tokens in Superposition
Long Chain-of-Thought (CoT) reasoning improves LLM problem-solving but is computationally expensive due to sequential token generation. Whi…
Mirage Probes: How Vision Models Fake Visual Understanding
Vision-language models (VLMs) can answer image-based questions confidently, and often correctly, even when no image is provided. This mirag…
Crypto x AI, AI x Crypto: A Survey
The intersection of crypto x AI is spawning papers, products, online posts, and companies. All the surrounding buzz, though, obscures what…
Gefen: Optimized Stochastic Optimizer
AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to tr…
How do Self-Supervised Remote Sensing Vision Models Transfer to Downstream Tasks?
Self-supervised geospatial foundation models (GeoFMs) learn transferable representations from remote sensing data, but their downstream beh…
HiLo-Token: Input-Adaptive High-Low Frequency Token Compression for Efficient Image Editing
Creative image editing tools, such as Photoshop's Remove or Generative Fill buttons, are central to everyday customer use and account for a…
SANA: What Matters for QA Agents over Massive Data Lakes?
Exploratory question answering (EQA) over data lakes requires an LLM agent to discover relevant sources, analyze retrieved data, and adapt…
GMN4AD: Graph Matching Network for Alzheimer's Disease Diagnosis with Test-Time Domain Adaptation using Multi-centered Structure Magnetic Resonance Imaging
Alzheimer's Disease (AD) is a progressive neurodegenerative disorder that affects millions of older adults, with prevalence expected to ris…
The Silent Cost of Artificial Intelligence Assistance: A Theory of Autonomy Surrender, the Recovery Mechanism, and the Restoration of Human Agency
The integration of artificial intelligence into human decision-making environments has introduced a previously undertheorized cost: the gra…
STREAM: Multi-Tier LLM Inference Middleware with Dual-Channel HPC Token Streaming
Researchers and practitioners working with large language models face a fragmented landscape: local models are free and private but hardwar…
Mask, Sample, Revise: A Revisable CTMC Inference Stack for Guided Discrete Flow Matching Text-to-Speech
Recent alignment-free non-autoregressive (NAR) text-to-speech (TTS) models formulate synthesis as a conditional infilling task, bypassing e…
Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH
LLM-based Agents are becoming increasingly capable and widely deployed, creating growing incentives for adversarial misuse in the real-worl…
Same-Origin Policy for Agentic Browsers
Agentic browsers integrate autonomous AI agents into web browsers, enabling users to accomplish web tasks through natural-language instruct…
Knowledge Graph Enhanced Memory-Augmented Retrieval for Long Context Modeling
Long-context language modeling requires not only extending context windows but maintaining coherent understanding of entity states and rela…
Rethinking Backdoor Adversarial Unlearning through the Lens of Catastrophic Forgetting in Continual Learning
Existing studies reveal that current backdoor defenses exhibit limited robustness and often fail against specific types of attacks. More co…
Clay-CNN Hybrids: Leveraging Geo-Foundational Models as Auxiliary Context for Landslide Detection
Rapid post-event landslide mapping is essential for disaster response but remains difficult to automate due to extreme class imbalance. Thi…
FEMOT: Multi-Object Tracking using Frame and Event Cameras
Conventional RGB cameras have been widely used in multi-object tracking due to their ability to capture rich appearance and semantic inform…
Numbers Already Carry Their Own Embeddings
We introduce Adelic operation-preserved embeddings (AOE), a training-free representation that captures both a number's real value and its m…
A Two-Stage Statistical Framework for Evaluating Associative Interference in Large Language Models
Large language models (LLMs) are increasingly evaluated for bias using adaptations of human psychological paradigms, yet methodological lim…
FAConformer: Frequency-Aware Convolutional Transformer for Auditory Attention Decoding
Auditory attention decoding (AAD) aims to infer the attended speaker from neural responses in multi-speaker acoustic environments and is a…
Recovering Stranded Discrimination in Knowledge Tracing: Per-Item Bias Correction via Empirical-Bayes Shrinkage
Deployed knowledge-tracing models are typically frozen after training, yet systematic per-item logit bias arises, from limited per-item exp…
Conditioning Matters: Stabilizing Inversion and Attention in Diffusion Image Editing
Inversion-based image editing offers flexible and training-free control but still struggles with inversion accuracy and the trade-off betwe…
Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources
Sound events are entities with semantic identities, locations, and trajectories, but current audio-language models usually reason about cli…
Implicit Reasoning for Large Language Model-based Generative Recommendation
Large Language Models (LLMs) are increasingly adopted as backbones for Generative Recommendation (GR), promising access to pretrained world…
Learning High Coverage Discriminative Parsimonious Rulesets
Learning systems based on IF-THEN rule representations readily offer interpretability, making them a crucial focus in contemporary AI resea…
Learning Urban Access Costs from Origin-Destination Flows via Inverse Optimal Transport
Cities deliver basic services through mixed public-private facility networks, including schools, clinics, transit providers, and subsidized…
Robustness without Wrinkles: Parallel Simulation and Robust MPC for Certified Deformable Manipulation
We present CORD-SLS, a real-time control method for safe deformable object manipulation, with a focus on ropes and cloth. At its core is a…
OdysSim: Building Foundation Models for Human Behavior Simulation
Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-drive…
MeEvo: Metacognitive Evolution Combined with Natural Evolution for Automatic Heuristic Design
Large Language Models (LLMs) have advanced Automatic Heuristic Design (AHD) by enabling heuristic generation through reasoning and code syn…
From Prompts to Responses: Dual-Sided Data Leakage and Defense in Split Large Language Models
Large language models (LLMs) are increasingly deployed in privacy-sensitive domains, where users must balance the risk of data exposure thr…
Universal Manipulation Exoskeleton: Learning Compliant Whole-body Policies with Real-time Torque Feedback
For robots to work safely in household environments, they need to be compliant and react to torque and force feedback during contact. Howev…
Selective Agentic Recovery for UAV Autonomy with a Persistent Mission Runtime
Agentic AI can support unmanned aerial vehicle (UAV) autonomy by providing high-level recovery reasoning when local waypoint- or setpoint-b…
When and How Severely: Scenario-Specific Safety Envelopes for Driving VLAs
Safety certification of Vision-Language-Action (VLA) driving planners under ISO 21448 (SOTIF) rests on an Operational Design Domain (ODD) s…
ChronoID: Infusing Explicit Temporal Signals into Semantic IDs for Generative Recommendation
Semantic IDs are crucial in generative recommendation, but with a fundamental limitation: temporal information is not well incorporated int…
Robust Fall Recovery for Armless Bipedal-Wheeled Robots Via Force-Guided Learning
Fall recovery is critical for autonomous legged locomotion. Existing methods have demonstrated that some legged robots, such as humanoids a…
DIFF-ERO: A Conformance-Aware Loss for Deep Learning in Process Mining
Deep learning has driven many recent advances in process analytics, especially for predictive and prescriptive monitoring. However, standar…
Hierarchical ODE: Learning Continuous-Time Physical Prototypes for Early Link Failure Detection
Time series prototype learning is fundamentally challenged by observational ambiguity. Discrete architectures fail to resolve this, as they…
AgentCyberRange: Benchmarking Frontier AI Systems in Realistic Cyber Ranges
Frontier AI systems are increasingly capable of cybersecurity tasks, including codebase inspection, vulnerability detection, and exploitati…
Pix2Pix-Hybrid: Structure-Guided Conditional Synthesis of Hajj Crowd Images with Multi-Channel Conditioning and Weak Attribute Supervision
Developing accurate crowd-counting models for Hajj pilgrimage scenes remains challenging because domain-specific annotated images are scarc…
Thinking Outside the [Chat]Box: Bridging Computer Science and Industrial Design for Cognitive-Inclusive Generative AI
Current Generative AI (GenAI) interfaces remain largely constrained to chatbox interaction, which can impose high cognitive demands on user…
Transforming Shape Schemas with Composable Property-Graph Queries (Extended Version)
Property graphs may be constrained by schemas that inform both query engines and human users about the shape of valid data, enforcing a con…
Achieving Precise Text-To-Cypher Via Grounded Knowledge Graph Data Generation
Property Graphs are rapidly being adopted as database frameworks for representing heterogeneous data sources. To enable precise access to t…
I'm Sorry Driver, I'm Afraid I Can't Do That: Appraising the Safety of LLMs within Automotive Contexts
This paper appraises recent frameworks within AI development to integrate LLMs into control tasks in automotive contexts from the perspecti…
Squeeze-Release: Iterative Pruning with Exact Structural Minimization
Unstructured pruning produces sparse weight tensors, but the standard implementation keeps tensor shapes unchanged so the deployed model is…
Design Methodology and Performance Trade-offs Management for Distributed and Compound AI Systems
Artificial Intelligence (AI) systems must typically satisfy service-level objectives including accuracy, latency, and cost. The prevailing…
PLAIground: SLO-Driven Runtime Model Selection for Compound AI Systems in the Edge-Cloud-Space Continuum
Applications in the 3D Computing Continuum, which unifies edge, cloud, and space, require combining multiple AI tasks such as object detect…
No Accidental Software Agent First Canonical Code for Human Code Entropy Reduction and 30 to 500 times Lower Frontier Model Requirements
Frontier coding models may spend substantial capacity learning not only program behavior, but also accidental entropy in human repositories…
Elastic Queries Reinforcement Learning: Self-Aware Policy Execution for VLA Models
Vision-language-action (VLA) models are powerful action generators for robot manipulation, but they are typically executed with fixed infer…
Discovery under Hypothesis Redundancy: A Geometric Theory of Discovery Bottlenecks
Scientific discovery saturates when new hypotheses cease to provide independent information, even if the nominal hypothesis space remains l…
Learning to Hear Hesitation: Continual Learning for Disfluency-Aware ASR
Despite advances in large-scale Automatic Speech Recognition (ASR), disfluent speech remains challenging, as state-of-the-art systems are o…
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: da…
CADET: Physics-Grounded Causal Auditing and Training-Free Deconfounding of End-to-End Driving Planners
End-to-end (E2E) autonomous-driving planners trained by imitation are prone to statistical shortcuts: they associate scene elements that me…
tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration
Existing multi-agent software development systems have proposed many forms of agent collaboration, including role-based collaboration and a…
MoDiCoL: A Modular Diagnostic Continual Learning Dataset for Robust Speech Recognition
Modern Automatic Speech Recognition (ASR) systems have made remarkable progress on standard benchmarks, yet performance gaps have emerged u…
The Perceived Fragility of Explanations in Audio Models: Manipulation of Attribution with Unchanged Predictions
This paper investigates the fragility of post-hoc explanation methods in audio deepfake detection. While previous work on explanation manip…
A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction
Predicting the Kohn-Sham Hamiltonian with machine learning can accelerate density functional theory while retaining access to molecular orb…
Fodor and Pylyshyn's Systematicity Challenge Still Stands
The recent successes of neural networks producing human-like language have caused significant stir in cognitive science, with many research…
Securing the Future of IoMT in the Post-Quantum Era: An Edge-Native Federated Learning Approach
Internet of Medical Things (IoMT) devices operate under strict resource constraints while handling highly sensitive health data, making sec…
From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails
LLM-based guardrails have emerged as a highly effective defense against prompt injection and jailbreak attacks in autonomous agents. Howeve…
TRACE: Trajectory-Routed Causal Memory for Delayed-Evidence Visuomotor Imitation
Robots under autonomous operation may require decisions based on evidence that is no longer visible. We study \emph{delayed-evidence} tasks…
Rethinking Global Average Pooling: Your Classifier Is Secretly a Multi-Instance Learner
Modern image classifiers widely adopt global average pooling (GAP) followed by a linear classification head. This linearity ensures that th…
Regional Climate Model Emulation with Diffusion Approaches: What is the Added Value of Generative Machine Learning?
Emulators provide a cost-effective alternative to regional climate models (RCMs) by capturing their dynamical downscaling function. They li…
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmar…
CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation
Granting LLMs direct control over costly, irreversible scientific experiments leads to unsafe exploration and unstable performance, but dis…
Sensitivity Shaping for Latent Modeling
Generative dynamics models enable planning in challenging robotic systems, but safe deployment requires reliably detecting policy-induced o…
When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime
LLM agent systems increasingly run as long-lived autonomous runtimes: scheduling jobs, calling tools, maintaining memory, and pushing resul…
AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models
Large Audio-Language Models (LALMs) have shown strong performance on a wide range of audio understanding tasks, yet they still struggle wit…
Regulating the Machine Contributor: Governance and Policy Alignment in Open Source
AI-assisted software development has moved from line-level autocomplete to agents that can plan changes, edit files, and submit pull reques…
A Comparative Study of Deep Learning Architectures for Multi-Horizon Behavioural Forecasting for Mobile Health
Wearable devices and smartphones generate rich behavioural time series that can support proactive health interventions, yet systematic comp…
Expert-Driven Survival Machines: Improving Stratification and Interpretability in Multiple Clinical Cohorts
Survival prediction plays a central role for healthcare providers and clinical researchers. Accurate risk stratification enables early inte…
Moonlight in Latent Space: Chirality and Structural Correspondence Between Beethoven's Op. 27 No. 2 and Machine Learning Mechanisms
We show that the three movements of Beethoven's "Moonlight Sonata" (Op. 27 No. 2) instantiate three distinct machine learning architectures…
When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks
Verifier-driven self-DPO is a common recipe for self-improving production visual-language models. In this setup, a frozen verifier scores c…
From Self-Supervised Speech Models to Mixture-of-Experts for Robust Anti-Spoofing
Recent advances in speech generation have significantly improved the naturalness of synthetic speech, making spoofing detection increasingl…
Listening with Attention: Entropy-Guided Explainability for Transformer-Based Audio Models
Transformer-based automatic speech recognition (ASR) models such as Whisper are highly accurate, but their predictions remain difficult to…
Giving AI a Headache: Acoustic Adversarial Attacks to Computer Vision Applications
Artificial Intelligence (AI) is increasingly used to automate a variety of real-world computer vision (CV) applications, such as autonomous…
CottonLeafVision: An Explainable and Robust Deep Learning Framework for Cotton Leaf Disease Classification
Globally, cotton is a highly economically beneficial crop, as the textile industry heavily depends on it. So, the precise identification an…
Flood and Harvest: The Provable Necessity of Trivia for Generating Valuable Mathematics via the Lens of Language Generation in the Limit
AI systems coupled to proof assistants now generate formal mathematics at scale, and the gap between what a checker can verify and what a m…
Learning Coordinated Preference for Multi-Objective Multi-Agent Reinforcement Learning
Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, potentially conflicting…
ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning
Building trustworthy medical multimodal large language models (MLLMs) is critical for reliable clinical decision support. Existing medical…
Learning optimal policies from event logs through reinforcement learning: a comparison of deep and MDP-based approaches
Prescriptive Process Monitoring is an emerging area within Process Mining that focuses on recommending actions to optimize business outcome…
ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins
Digital twins are increasingly used to monitor and optimize industrial systems, yet many existing frameworks remain difficult to interpret,…
LLM-Powered AI Agent Systems and Their Applications in Industry
The emergence of Large Language Models (LLMs) has reshaped agent systems. Unlike traditional rule-based agents with limited task scope, LLM…
Verbatim Chunks Beat Extracted Artifacts: A Controlled Ablation of Memory Representations for Long LLM Conversations
A growing class of conversational-memory systems compresses dialogue history into structured artifacts -- extracted facts, decisions, or ev…
Token-Level LLM Collaboration via FusionRoute
Large language models (LLMs) exhibit strengths across diverse domains. However, achieving strong performance across these domains with a si…
Actionable Interpretability Must Be Defined in Terms of Symmetries
This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of inte…
Optimizing Agentic Reasoning with Retrieval via Synthetic Semantic Information Gain Reward
Agentic reasoning enables large reasoning models (LRMs) to dynamically acquire external knowledge, but yet optimizing the retrieval process…
FlexMS: A Unified Public Benchmark for Molecule Tandem Mass Spectrum Prediction
Tandem mass spectrometry (MS/MS) is central to small molecule identification, but current deep learning systems for spectrum prediction sti…
Generative AI for Managerial Decision-Making under Ambiguity and Sycophancy
Generative artificial intelligence (GenAI) is increasingly being integrated into complex business workflows, fundamentally shifting the bou…
An Analysis of the Coordination Gap between Joint and Modular Learning for Job Shop Scheduling with Transportation Resources
Efficient job-shop scheduling with transportation resources is critical for high-performance manufacturing. With the rise of "decentralized…
PRISM: Perception Reasoning Interleaved for Sequential Decision Making
Scaling LLM-based embodied agents from text-only environments to complex multimodal settings remains a major challenge. Recent work identif…
AdaTKG: Adaptive Memory for Temporal Knowledge Graph Reasoning
Temporal knowledge graphs (TKGs) represent time-stamped relational facts and support a wide range of reasoning tasks over evolving events.…
Learning Developmental Scaffoldings to Guide Self-Organisation
From subcellular structures to entire organisms, many natural systems generate complex organisation through self-organisation: local intera…
Planning with the Views via Scene Self-Exploration
Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1…
VikingMem: A Memory Base Management System for Stateful LLM-based Applications
Large Language Models have revolutionized interactive applications; however, their finite context windows pose a critical data management c…
Evidence-Gated LLM Priors for Multi-Objective Bayesian Optimization
Large language models (LLMs) are increasingly used as heuristic advisors for black-box optimization, yet their suggestions and self-reporte…
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning
Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in age…
Output Type Before Quality: A Standards-Derived XAI Admissibility Rubric for Autonomous-Driving Safety
Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect cha…
StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents
Reinforcement Learning (RL) has become a promising approach for improving GUI Agents in long-horizon, stochastic digital environments, but…
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perf…
Large-scale semantic mapping of learner agency and autonomy reveals what measurement and generative AI research overlook
Learner agency and autonomy are foundational to personal development, yet a pervasive "jingle-jangle" fallacy (i.e. identical terms denotin…
GUITrans2Act: Understanding User Operational Behaviors from Mobile GUI Interactions with Vision-Language Models
Understanding the digital world on mobile devices is shifting from static UI perception to dynamic action comprehension. This capability en…
Under What Conditions Can a Machine Be Called Genuinely Creative?
Recent AI systems can generate texts, software architectures, hypotheses, designs, and scientific workflows that appear creative. This pape…
MiniMax Sparse Attention
Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persiste…
Is It You or Your Environment? A Bayesian Inference Framework for Genomically-Anchored Personalized Physiological Interpretation
Personalized health AI systems face a fundamental cold-start problem: machine learning models for physiological interpretation require week…
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery
LLM-based agents have shown increasing potential in automating scientific discovery. Given an optimizable metric and an execution environme…
Application of Artificial Intelligence and Machine Learning in Libraries: A Systematic Review
As the concept and implementation of cutting-edge technologies like artificial intelligence and machine learning has become relevant, acade…
MirrorCheck: Efficient Adversarial Defense for Vision-Language Models
Vision-Language Models (VLMs) are increasingly susceptible to sophisticated adversarial attacks, including adaptive strategies specifically…
Vanishing Depth: Training Generalized Depth Adapters with Sinusoidal Depth Preprocessing for Pretrained RGB Encoders
Generalized metric depth understanding is critical for precise vision-guided robotics, which current state-of-the-art (SOTA) vision-encoder…
Revisiting Outage for Edge Inference Systems
One of the key missions of sixth-generation (6G) mobile networks is to deploy large-scale artificial intelligence (AI) models at the networ…
FPGA-Based Neural Network Accelerators for Space Applications: A Survey
Space missions are becoming increasingly ambitious, necessitating high-performance onboard spacecraft computing systems. In response, field…
UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
Retrieval-Augmented Generation (RAG) has shown substantial promise in improving factual accuracy by grounding model responses with external…
The Accountability Paradox: How Platform API Restrictions Undermine AI Transparency Mandates
Recent application programming interface (API) restrictions on major social media platforms challenge compliance with the EU Digital Servic…
Fractured Chain-of-Thought Reasoning
Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing add…
MASLab: A Unified and Comprehensive Codebase for LLM-based Multi-Agent Systems
LLM-based multi-agent systems (MAS) have demonstrated significant potential in enhancing single LLMs to address complex and diverse tasks i…
Federated Causal Inference from Multi-Site Observational Data via Propensity Score Aggregation
Causal inference typically assumes centralized access to individual-level data. Yet, in practice, data are often decentralized across multi…
Sentinel: Decoding Context Utilization via Attention Probing for Efficient LLM Context Compression
Retrieval-augmented generation (RAG) often suffers from long and noisy retrieved contexts. Existing context compression methods typically r…
DiffusionBlocks: Block-wise Neural Network Training via Diffusion Interpretation
End-to-end backpropagation requires storing activations throughout all layers, creating memory bottlenecks that limit model scalability. Ex…
UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators
Large language models (LLMs) require larger GPU memory size these days, necessitating efficient and extreme weight compression methods. Exi…
From Sorting Algorithms to Scalable Kernels: Bayesian Optimization in High-Dimensional Permutation Spaces
Bayesian Optimization (BO) is a powerful tool for black-box optimization, but its application to high-dimensional permutation spaces is sev…
Fusion of Pervasive RF Data with Spatial Images via Vision Transformers for Enhanced Mapping in Smart Cities
In this paper, we present a deep learning-based approach that integrates the DINOv2 architecture to improve building mapping by combining (…
Tackling GNARLy Problems: Graph Neural Algorithmic Reasoning Reimagined through Reinforcement Learning
Neural algorithmic reasoning (NAR) is a paradigm that trains neural networks to execute classic algorithms by supervised learning. Despite…
Q-Net: Queue Length Estimation via Kalman-based Neural Networks
Estimating queue lengths at signalized intersections is a long-standing challenge in traffic management. Partial observability of vehicle f…
Shift-Invariant Attribute Scoring for Kolmogorov-Arnold Networks via Shapley Value
For many real-world applications, understanding feature-outcome relationships is as crucial as achieving high predictive accuracy. While tr…
RAMAC: Multimodal Risk-Aware Offline Reinforcement Learning and the Role of Behavior Regularization
In safety-critical domains where online data collection is infeasible, offline reinforcement learning (RL) is attractive only if policies a…
Chronological Thinking in Full-Duplex Spoken Dialogue Language Models
Recent advances in spoken dialogue language models (SDLMs) reflect growing interest in shifting from turn-based to full-duplex systems, whe…
Distributional Biases in Post-Training: A Markovian Analysis of Reasoning Trajectories
Foundation models exhibit broad knowledge but limited task-specific reasoning, motivating post-training strategies such as RL with verifiab…
The Journal of Prompt-Engineered (Moral) Philosophy Or: Why AI-Assisted Ethics Research Requires Process Transparency
Existing AI disclosure mandates in scholarship require that AI assistance be reported but leave transparency philosophically unspecified: t…
An interpretable unsupervised representation learning for high precision measurement in particle physics
Unsupervised learning has been widely applied to various tasks in particle physics. However, existing models lack precise control over thei…
COGNITION: From Evaluation to Defense against Multimodal LLM CAPTCHA Solvers
This paper studies how multimodal large language models (MLLMs) undermine the security guarantees of visual CAPTCHA. We identify the attack…
Interpretable Alzheimer's Diagnosis via Multimodal Fusion of Regional Brain Experts
Accurate and early diagnosis of Alzheimer's disease (AD) is critical for effective intervention and requires integrating complementary info…
Schr\"odinger's Navigator: Imagining an Ensemble of Futures for Zero-Shot Object Navigation
Zero-shot object navigation (ZSON) requires robots to find target objects in unseen environments without task-specific fine-tuning or pre-b…
Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2
Structured width pruning of GLU-MLP layers in Llama-3.2 models, guided by the Peak-to-Peak Magnitude (PPM) criterion, reveals a systematic…
Succeeding at Scale: Enterprise Retrieval Benchmark Construction and Index-Preserving Query Adaptation for Multi-Tenant Search
Large-scale multi-tenant retrieval systems generate extensive query logs but lack curated relevance labels for effective domain adaptation,…
CuMA: Aligning LLMs with Sparse Cultural Values via Demographic-Aware Mixture of Adapters
As Large Language Models (LLMs) serve a global audience, alignment must transition from enforcing universal consensus to respecting cultura…
Can professional translators identify machine-generated text?
This study investigates whether professional translators without prior specialized training can reliably identify short stories generated i…
Jacobian Scopes: token-level causal attributions in LLMs
Large language models (LLMs) make next-token predictions based on clues present in their context, such as semantic descriptions and in-cont…
SMART: Scalable Mesh-free Aerodynamic Simulations from Raw Geometries using a Transformer-based Surrogate Model
Machine learning-based surrogate models have emerged as more efficient alternatives to numerical solvers for physical simulations over comp…
Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed Goals
Unsupervised pre-training can equip reinforcement learning agents with prior knowledge and accelerate learning in downstream tasks. A promi…
Learning What to Predict: Downstream-Guided Task Design for Continued Pretraining
Continued pretraining is optimized with fixed self-supervised tasks but selected by downstream performance, creating a coarse feedback loop…
Residual Context Diffusion Language Models
Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to purely autoregressive language models because they can d…
Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention
Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interac…
Quantized Evolution Strategies: High-precision Fine-tuning of Quantized LLMs at Low-precision Cost
Post-Training Quantization (PTQ) is essential for deploying Large Language Models (LLMs) on memory-constrained devices, yet it renders mode…
Rethinking the Trust Region in LLM Reinforcement Learning
Reinforcement learning (RL) has become a cornerstone for fine-tuning Large Language Models (LLMs), with Proximal Policy Optimization (PPO)…
HyperPotter: Spell the Charm of High-Order Interactions in Audio Deepfake Detection
Advances in AIGC technologies have enabled the synthesis of highly realistic audio deepfakes capable of deceiving human auditory perception…
Protean Compiler: An Agile Framework to Drive Fine-grain Phase Ordering
The phase ordering problem has been a long-standing challenge since the late 1970s, yet it remains an open problem due to having a vast opt…
Metabolic cost of information processing in Poisson variational autoencoders
Computation in biological systems is fundamentally energy-constrained, yet standard theories of computation treat energy as freely availabl…
Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven Resampling
Effective exploration is a key challenge in reinforcement learning for large language models: discovering high-quality trajectories within…
FedRot-LoRA: Mitigating Rotational Misalignment in Federated LoRA
Federated LoRA provides a communication-efficient mechanism for fine-tuning large language models on decentralized data. In practice, howev…
Generalized Discrete Diffusion with Self-Correction
Self-correction is an effective technique for maintaining parallel sampling in discrete diffusion models with minimal performance degradati…
3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene Understanding
Reinforcement Learning with Verifiable Rewards ( RLVR ) has emerged as a transformative paradigm for enhancing the reasoning capabilities o…
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
Large language models (LLMs) are increasingly used as judges of chain-of-thought (CoT) reasoning, yet it remains unclear whether they can r…
The Curse and Blessing of Mean Bias in FP4-Quantized LLM Training
FP4 training promises substantial memory and compute savings for large language models, but remains fragile because blockwise quantization…
TabKD: Tabular Knowledge Distillation through Interaction Diversity of Learned Feature Bins
Data-free knowledge distillation enables model compression without original training data, critical for privacy-sensitive tabular domains.…
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
Large language models often fail to satisfy formatting instructions when they must simultaneously perform demanding tasks. We study this be…
X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
While the shift from cascaded dialogue systems to end-to-end (E2E) speech Large Language Models (LLMs) improves latency and paralinguistic…
Low-Burden LLM-Based Preference Learning: Personalizing Assistive Robots from Natural Language Feedback for Users with Paralysis
Physically Assistive Robots require personalized behaviors to ensure user safety and comfort. However, traditional preference learning meth…
The Shrinking Lifespan of LLMs in Science
Scaling laws describe how language model capabilities grow with compute and data, but say nothing about how long a model matters once relea…
STaR-DRO: Stateful Tsallis Reweighting for Group-Robust Structured Prediction
Structured prediction with large language models requires outputs that are label-accurate, ontology-constrained, structurally valid, and ev…
Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alter…
LEPO: Latent Reasoning Policy Optimization for Large Language Models
Recently, latent reasoning has been introduced into large language models (LLMs) to leverage rich information within a continuous space. Ho…
Scalable Production Scheduling: Linear Complexity via Unified Homogeneous Graphs
Efficiently solving the Job Shop Scheduling Problem in real-world industrial applications requires policies that are both computationally l…
Quantile-Free Uncertainty Quantification in Graph Neural Networks
Uncertainty quantification (UQ) in graph neural networks (GNNs) is crucial in high-stakes domains but remains a significant challenge. In g…
Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions
We study planning site formation in language models -- where internal representations of structurally-constrained future tokens form during…
SAFformer:Improving Spiking Transformer via Active Predictive Filtering
Spiking Neural Networks (SNNs) offer notable advantages in biological plausibility and energy efficiency, making them promising candidates…
Relational Retrieval: Leveraging Known-Novel Interactions for Generalized Category Discovery
In this study, we tackle Generalized Category Discovery (GCD) via a Relational Retrieval perspective, explicitly coupling labeled and unlab…
EmoMind: Decoding Affective Captions from Human Brain fMRI
Decoding visual experience from brain activity has advanced substantially, but current brain-to-text systems largely recover semantic conte…
The Insurability Frontier of AI Risk: Mapping Threats to Affirmative Coverage, Silent Exposures, and Exclusions
The rapid diffusion of agentic AI has created a new coverage problem for commercial insurance: some AI-mediated losses are now affirmativel…
Exact Linear Attention
This paper introduces Exact Linear Attention (ELA), a mechanism that achieves linear computational complexity for Transformer attention by…
Manga109-v2026: Revisiting Manga109 Annotations for Modern Manga Understanding
Manga is a culturally distinctive multimodal medium and one of the most influential forms of Japanese popular culture. As AI systems increa…
Catching magnetic resonance imaging outliers in artificial intelligence-supported radiotherapy workflows: unsupervised detection and localization of image anomalies using deep learning
Artificial intelligence is increasingly integrated into radiotherapy workflows, yet such pipelines remain vulnerable to out-of-distribution…
Rotation-Invariant Spherical Watermarking via Third-Order SO(3) Representation Coupling
Reliable watermarking of panoramic imagery is fundamentally challenged by arbitrary 3D rotations. As panoramas are defined on the sphere, t…
Models That Know How Evaluations Are Designed Score Safer
The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has ide…
Silent Failures in Federated Personalization of Foundation Models
Foundation models are increasingly personalized on decentralized private data through federated learning and are now deployed at scale unde…
Patcher: Post-Hoc Patching of Backdoored Large Language Models
Large language models remain vulnerable to jailbreak backdoor attacks, where adversaries poison safety alignment data to embed hidden trigg…
CoRe-MoE: Contrastive Reweighted Mixture of Experts for Multi-Terrain Humanoid Locomotion with Gait Adaptation
Humans primarily rely on walking and running to traverse complex terrains. Similarly, humanoid robots should be able to smoothly transition…
Benchmarking Vision-Language-Action Models on SO-101: Failure and Recovery Analysis
Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily…
When Roleplaying, Do Models Believe What They Say?
Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that pe…
Will AI Agents Free Us From Meaningless Work? A Human-Centered Analysis
Some claim that AI agents will free workers from the boring parts of their jobs, yet little is known about how workers themselves identify…
Quickest Detection of Hallucination Onset: Delay Bounds and Learned CUSUM Statistics
Token-level hallucination detectors are evaluated as classifiers, by AUC over all tokens, yet a streaming monitor is judged by its reaction…
Bounding Boxes as Goals: Language-Conditioned Grasping via Neuro-Symbolic Planning
For robotics to be effectively integrated into household or industrial environments, machines must adapt to natural-language prompts in rea…
MAStrike: Shapley-Guided Collusive Red-Teaming on Multi-Agent Systems
Hierarchical multi-agent systems (MAS) are rapidly being deployed in high-stakes workflows across domains such as finance and software engi…
Order Is Not Control: Driven-Dissipative Response Laws Across Artificial and Biological Systems
AI alignment, interpretability, steering, and neural perturbation studies identify order-inducing objects. We argue that order is not contr…
TWLA: Achieving Ternary Weights and Low-Bit Activations for LLMs via Post-Training Quantization
Large language models (LLMs) exhibit exceptional general language processing capabilities, but their memory and compute costs hinder deploy…
MP3: Multi-Period Pattern Pre-training for Spatio-Temporal Forecasting
Spatio-Temporal forecasting is crucial in diverse fields, such as transportation, climate, and energy. Urban spatio-temporal data exhibits…
Ontology Memory-Augmented ASR Correction for Long Text-Speech Interleaved Conversations
Automatic speech recognition (ASR) correction has traditionally focused on isolated utterances or short local contexts. However, as text an…
Sakana AI、初の商用プロダクト「Marlin」リリース その実力は?【出力レポート全文掲載】
Sakana AIがAI調査エージェント「Sakana Marlin」の提供を開始した。4月からβ版を提供していたものを商用化する。公開に先んじてメディア向けにサービスのハンズオンを実施。事前に集めたテーマを基に、AIに作成させたレポートを報道陣に公開した。
Sakana?AI、初の商用サービスはリサーチ特化 「Deep Research」との違いは? 後発で“ベンチマークも追わない”ワケ
Sakana?AIが6月15日に提供を始めた同社初の商用サービス「Sakana?Marlin」(サカナ・マーリン)。詳細や狙いを担当者に聞いた。
ChatGPT vs. Google検索──どっちで調べるのが学習効果が高い? 8日間の実験で検証した研究
米ジョージア工科大学やミシガン大学などに所属する研究者らが発表した論文「Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning」は、A…
AI・ロボット人材は約340万人不足 労働市場のスキル需給、AIでどう可視化する?
産業構造の変化と人口減少が同時に進み、業種や職種の間で人材の過不足が広がると予測されている。経済産業省は事業を通じて、労働市場全体のスキル需給をAIなどで可視化する取り組みに乗り出した。受託したNRIは、具体的に何に取り組むのか。
Introducing the OpenAI Partner Network
OpenAI launches the Partner Network, investing $150M to help global partners accelerate enterprise AI adoption, deployment, and transformat…
As AI companies race to go public, who else is along for the ride?
Startups are trying to "ride that SpaceX IPO wave."