週次AIニュース 2026-W28
対象期間: 2026-07-06 〜 2026-07-12(1623 件)
トピックの推移
トピック別件数
- LLM/生成AI 585件
- 研究/論文 579件
- エージェント 367件
- 画像/動画生成 307件
- ビジネス/資金調達 103件
- ロボティクス 99件
- その他 45件
- ハードウェア/半導体 38件
- 規制/政策 13件
今週のハイライト(上位 10 件)
How Deutsche Telekom is rewiring telecommunications with AI
How Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, network operations, and…
GPT-5.6 is now the preferred model in Microsoft 365 Copilot
Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, h…
ChatGPT is now a partner for your most ambitious work
ChatGPT Work is an agent that can take action across your apps and files, stay with a project for hours if needed, and turn a goal into fin…
GPT-5.5 Bio Bug Bounty
Details about the OpenAI Bio Bounty program
GPT-5.6: Frontier intelligence that scales with your ambition
More intelligence from every token, stronger performance per dollar, and more capability on demand for your hardest work.
Our approach to government and national security partnerships
Learn how OpenAI approaches government and national security partnerships, with principles for responsible AI use, democratic accountabilit…
Separating signal from noise in coding evaluations
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in…
Helping K–12 educators build practical AI skills
OpenAI Academy and the Walton Family Foundation are bringing hands-on AI Skills Jams to help K–12 educators build practical AI skills for t…
OpenAI bets on families as ChatGPT goes deeper into households
ChatGPT is hiring a dedicated product manager to build experiences for families, caregivers, and older adults, according to a job posting.
Apple sues OpenAI over alleged trade secret theft
Apple alleges the misconduct was directed by OpenAI's senior leadership, including a longtime former employee.
全件(日付別)
2026-07-11(5件)
OpenAI bets on families as ChatGPT goes deeper into households
ChatGPT is hiring a dedicated product manager to build experiences for families, caregivers, and older adults, according to a job posting.
Meta removes controversial AI feature on Instagram after backlash
"Our intent was to provide a useful creative tool and to give people control over whether their public content could be referenced in this…
Apple sues OpenAI over alleged trade secret theft
Apple alleges the misconduct was directed by OpenAI's senior leadership, including a longtime former employee.
Open source AI matters more than ever, according to Hugging Face’s Clem Delangue
Open source AI is booming, according to Hugging Face CEO Clem Delangue. The company has grown into something like a GitHub for AI in recent…
SK Hynix raises $26.5B in the biggest foreign IPO in US history, is urged to build new US fabs
The AI chip boom just produced its biggest Wall Street moment yet. Now SK Hynix and Samsung are being asked to build U.S. factories.
2026-07-10(271件)
Hugging Face’s CEO on why companies are done renting their AI
Open source AI is booming, according to Hugging Face CEO Clem Delangue. The company has grown into something like a GitHub for AI in recent…
テスラ車内で「Grok」と会話、日本でも展開へ ナビ設定やルート確認を音声で
対話型AI「Grok」を日本のテスラ車で利用できるようになった。Tesla Japanが7月10日、X公式アカウントで発表した。利用には、ソフトウェアバージョン2026.20以降、「プレミアムコネクティビティ」(有料の車内通信プラン)の契約が必要になる。
「まるで人間」 OpenAIの新モデル「GPT-Live」のトーク力が話題 間を空けずに考えながら会話できる
OpenAIが提供を始めた音声会話向けモデル「GPT-Live」がXで話題だ。同社はこれまでも音声会話モデルを提供していたが、聞き取りと発話を同時に行う新アーキテクチャにより「会話が自然すぎる」などの声が上がっている。
M365 CopilotにOpenAI最新モデル「GPT-5.6」追加 複雑な作業を「より効率的に処理」
米Microsoftは、Microsoft 365のAI機能「Microsoft 365 Copilot」に、米OpenAIの最新のAIモデル「GPT-5.6」シリーズを追加すると発表した。WordやExcel、PowerPointなどで順次展開する。
How Deutsche Telekom is rewiring telecommunications with AI
How Deutsche Telekom is becoming an AI-native telco with OpenAI-transforming customer service, employee workflows, network operations, and…
「生成AIをもう手放せない人」が約6割 逆に“使わなくなったもの”1位は?
ICT総研の調査では、生成AIサービスが使えなくなると困ると答えた人は約6割に上った。生成AIサービスが日常的なツールとして定着する一方で、利用頻度が減ったものもあるという。それは何なのか。
Context Graphs for Proactive Enterprise Agents
Retrieval-Augmented Generation (RAG) and agentic frameworks have advanced enterprise AI considerably, yet agents remain fundamentally react…
AI-integrated models for assessing agricultural resilience
Agricultural supply chains are vulnerable to disruptions through linked biophysical and economic systems. We develop an AI-powered tool tha…
Adversarial Social Epistemology for Assemblies of Humans and Large Language Models
We outline an adversarial social epistemology (ASE) for densely interactive communicative landscapes in which public assertions are scaffol…
Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning
Large language models (LLMs) have emerged as important tools in healthcare, showing growing potential for clinical reasoning and patient ca…
Alignment Plausibility: A New Standard for Assuring AI in Healthcare
Large language models (LLMs) have become significant providers of mental health support, yet they remain products of an attention economy w…
Idiobionics: The Unification of Privacy and Intelligent Robotic Prostheses
The human body is at the center of a growing family of technologies designed to tightly and persistently couple biological and digital syst…
Infinity-Parser2 Technical Report
We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement lea…
VectorizationLLM: Smart Vectorization Based AI Assistant
VectorizationLLM is a specialized Large Language Model based on Google open-weight LLMs. The model is designed to assist students to learn…
A Graph Neural Network Model for Real-Time Gesture Recognition Based on sEMG Signals
For seemless control of advanced hand prostheses and augmented reality, accurate and immediate hand gestures recognition is essential. Surf…
Agentic AI and Retrieval-Augmented Models in Straight-Through Underwriting
Artificial intelligence (AI) is beginning to reshape actuarial practice, particularly in domains that require reasoning over unstructured d…
Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning
Reinforcement learning (RL) research has increasingly shifted focus towards alignment, ensuring agents learn behaviors adhering to human va…
Nigeria Machinery: A Low-Resource Industrial Dataset with a Domain-Grounded Reasoning Layer
There is relatively little, public, and model-ready data on industrial machinery for African economies. This makes it hard to do quantitati…
Persona Cartography: Charting Language Model Personality Traits in Weight Space
Large language models exhibit recurring behavioural patterns -- personas -- that shape generalisation and safety, but we lack reliable tool…
Evaluating the Effect of Frame Rate in Sequence-Based Classification of Autism-Related Self-Stimulatory Hand Idiosyncrasies
Autism spectrum disorder (ASD) affects over 75 million individuals worldwide, yet scalable computational methods for remote behavioral scre…
Agentic Neural Architecture Search
Neural architecture search (NAS) methods have grown increasingly efficient, yet they remain bounded by manually engineered search spaces th…
Concretized Proposition Prompting Resolves Composition-Knowledge Dichotomy in Large Language Models
LLMs often struggle to balance compositionality with knowledgeability, a challenge we define as Composition-Knowledge Dichotomy. To address…
From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents
Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Pr…
A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis
Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot pr…
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles…
Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring
Chain-of-thought (CoT) monitoring is a promising safety mechanism for AI agents, based on the premise that visible reasoning traces can sur…
PARA-PV: Physics-Aware Retrieval-Augmented PV Prediction Based on Frozen Foundation Model and Distribution Shift Correction
Accurate photovoltaic (PV) power forecasting is essential for reliable grid dispatch and renewable energy integration, yet it remains chall…
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet t…
Answer Set Programming Energised! End-to-End Neurosymbolic Reasoning and Learning with ASP and Energy Based Models
We present a general neurosymbolic reasoning and learning methodology based on a modular integration of answer set programming with an ener…
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden informat…
ASMR: Agentic Schema Generation for Ship Maintenance Report Writing
In this paper, we study the automatic schema generation problem: given a collection of historical ship maintenance and operational reports…
A First-Principles Theory of Slow Thinking and Active Perception
As part of a series on first-principles modeling of cognitive functions, this paper attempts to provide a mathematical formulation of think…
Playing ZendoWorld: Challenging AI Agents on Active Visual Concept Induction
A central challenge in building intelligent systems is enabling agents to jointly perceive complex inputs, form hypotheses about hidden pat…
AutoPersonas: A Multi-Timescale Loop Engine for Open-Ended Persona Evolution
Long-term persona agents must remain identifiable while adapting to new events, relationships, evidence, and social conditions. We identify…
Compete Then Collaborate: Frontier AI Teachers Build a Verifiable Curriculum to Improve a Coding Student Beyond Imitation
Large language models increasingly serve as teachers generating training data for smaller students. Prior multi-teacher knowledge distillat…
MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters
Large language models (LLMs) have shown strong performance on isolated psychiatric tasks, including dialogue, diagnosis, and treatment plan…
Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment
High-volume structured extraction pays a large model's latency on every item, so distilling the task into a small on-device model is attrac…
PolyUQuest: Verifiable Structure-Aware Web RAG over Heterogeneous Graphs
Existing retrieval-augmented generation (RAG) systems treat web pages as flat text, losing the structural and semantic signals encoded in H…
Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench
Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to eval…
Psychological Competence as a Missing Dimension in AI Evaluation
Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy co…
INTENT: An LSTM Framework for Vehicle Intention Prediction in Intersection Scenarios with Comprehensive Ablation Analysis
Vehicle intention prediction is a pivotal aspect in the agility and safety of autonomous vehicles in all driving scenarios; if genuine enha…
Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models
Modern AI models achieve strong performance on many established benchmarks, yet they still fail on tasks that humans find almost trivial, s…
MobiDiff: Semantic-Aware Multi-Channel Discrete Diffusion for Human Mobility Data Generation
Human mobility data are essential for transportation optimization, urban planning, and resource allocation, yet real-world mobility data ar…
FedOPAL: One-Shot Federated Learning via Analytic Visual Prompt Tuning
With the widespread deployment of basic models in edge intelligence, communication bandwidth has become a core bottleneck restricting the s…
Towards Mechanistically Understanding Why Memorized Knowledge Fails to Generalize in Large Language Model Finetuning
Fine-tuning LLMs to inject new knowledge faces a critical challenge: LLMs can quickly memorize new facts, yet fail to use them for downstre…
Game Theory Driven Multi-Agent Framework Mitigates Language Model Hallucination
The application of lightweight Large Language Models in rule-based scientific domains remains severely limited by their tendency to mimic l…
OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare…
Applying JEPA-Style Predictive Learning to JA4-Derived Network Fingerprints
I-JEPA and V-JEPA learn by matching latent predictions to target encoder outputs rather than regenerating the original input, and this has…
Drift-Aware Temporal Graph Rewiring (DATGR) for Adaptive Semantic Modeling in Biomedical Text
Biomedical language evolves rapidly as new discoveries emerge, causing traditional text models to lose semantic fidelity over time. Static…
AI-guided stimuli discovery and generation to optimize facial emotion perception studies in autism
Understanding perceptual differences between autistic and neurotypical adults requires behavioral assays that are sensitive, reliable, and…
CommuniWave:A Machine Learning Model for Quantifying the Degree of Temporary Informal Behavior in Urban Communities
For urban managers and designers, improving the functional attributes of urban communities to enhance territorial resilience in the face of…
SHAP-Weighted Cross-Modal Expert Fusion for Emotion and Sentiment Recognition: Evidence and Limits
Multimodal emotion and sentiment recognition is commonly addressed by early fusion, which concatenates modalities before classification, or…
Towards Precision Therapy in Hepatocellular Carcinoma: A Clinical-Reasoning LLM for Risk Stratification and Treatment Guidance
Hepatocellular carcinoma (HCC) is a common malignancy and a leading cause of cancer-related mortality. Current guidelines and staging syste…
The complexities of patient-centred conversational artificial intelligence
Consumer-facing health chatbots powered by large language models (LLMs) are increasingly used for symptom assessment. However, chatbot deve…
Formal Mechanisms for Market Stability in Self-Interested Agent Societies: A Marketplace Simulation Study
Self-interested agents, left unconstrained, tend toward defection in repeated social dilemmas, causing cooperative gains from trade to coll…
SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets
As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performanc…
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents
In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it an…
The Illusion of Equivalency: Statistical Characterization of Quantization Effects in LLMs
Post-training quantization is widely used to deploy large language models in resource-constrained settings, yet its evaluation relies almos…
Workflow as Knowledge: Semantic Persistence for LLM-Mediated Workflows
Large language model (LLM) applications increasingly use explicit workflows for tool use, retrieval, branching, checkpointing, and human ap…
AUTOPILOT VQA: Benchmarking Vision-Language Models for Incident-Centric Dashcam Understanding
Recent advances in Vision-Language Models, Large Language Models, and Multimodal Large Language Models have improved autonomous driving tas…
Using AI-based Learning Assistants in Higher Education: A Large-Scale Descriptive Analysis
In this study, we present a large-scale descriptive analysis of the use of an AI-based learning assistant (Syntea) in higher education. Bas…
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
Scientific ideas rarely start from a blank page. They inherit mechanisms, repair known limitations, and recombine pieces of earlier work, m…
Towards the Explainability of Temporal Graph Networks via Memory Backtracking and Topological Attribution
Temporal graphs are ubiquitous in real-world applications and Temporal Graph Networks (TGNs) have achieved superior predictive accuracy. Un…
LLT: Local Linear Transformer for PDE Operator Learning
Neural operators have become a common approach for learning PDE solution maps and accelerating numerical simulations. Transformer-based neu…
ReCoLoRA: Spectrum-Aware Recursive Consolidation for Continual LLM Fine-Tuning
Parameter-efficient fine-tuning adapts a large language model to one task cheaply, but across a task sequence LoRA-style methods keep stack…
Omni-Sleep: A Sleep Foundation Model via Hierarchical Contrastive Learning of CNS--ANS Dynamic
Sleep physiology arises from the coordinated dynamics of the central nervous system (CNS) and autonomic nervous system (ANS), as reflected…
SHIFT: Survival Prediction from Incomplete and Heterogeneous Genomic Data
Genomic prediction models often fail to transfer across institutions because sequencing panels differ across sites, creating structural fea…
Collective Intelligence with Foundation Models
As foundation models grow in scale and diversity, coordinating multiple models into cooperative reasoning systems offers a path toward safe…
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and age…
Architecture Generalization with MetaNCA
Self-organization is an emergent property of life, driven by the collective behavior of individual components acting on local information.…
A Transdiagnostic Space of Disorder Like Phenotypes in Reinforcement Learning Agents
Modelling psychological disorders in artificial agents offers both a testbed for computational psychiatry and a lens on the failure modes o…
Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms
Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most ch…
Graph-Regularized Deep Learning for EEG-Based Emotion Recognition with Psychologically-Grounded Label Structure
EEG-based emotion recognition is critical for mental health monitoring and affective brain-computer interfaces, yet existing deep learning…
From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier
Recent developments in AI for Mathematics (AI4Math), especially Large Language Model (LLM)-driven theorem provers, has achieved remarkable…
DreamCharacter-1: From 3D Generative Foundation Models to Product-Ready Character Generation
We present DreamCharacter-1, a lightweight post-adaptation framework that calibrates pretrained 3D foundation models toward high-fidelity,…
From Triggers to Emotions: A CPM-Grounded Appraisal Multi-Agent for Dynamic Emotional Evolution in Persona-Based Dialogue
Large Language Models (LLMs) have substantially advanced persona-based dialogue agents for emotion-sensitive role simulation in healthcare,…
Shift & Drift: A Zero-Shot Benchmark for Generalizable and Robust Autonomous Driving Motion Planning
While closed-loop motion planners trained on large-scale, object-level datasets, e.g., nuPlan, demonstrate strong in-distribution (ID) perf…
Kime-Representation Formulations of Three Open Problems in the Foundations of Classical Mechanics: Uncertainty, Invariant Entropy, and Directional Degrees of Freedom
We give mathematically self-contained formulations, in the complex-time (kime) representation, of three open problems from the foundations…
Multi-agent Autoformalization of Tensor Network Theory
We build a team of specialized large language-model agents and present an agent-driven workflow for research-level formalization in theoret…
Time-to-Collision Based Dynamic Obstacle Avoidance Using Pretrained Vision Models for Robots in Unstructured Environments
Dynamic obstacle avoidance in unstructured outdoor environments remains a critical challenge for autonomous mobile robots, particularly whe…
How Do I Know What to Say Next? Barenholtz's Autogenerative Theory as an Enrichment of Harrisean Integrationism
Roy Harris's Integrationist linguistics offers a compelling critique of the referentialist tradition embedded deep at the heart of computat…
Closed-Loop Dynamic Validator Node Scaling in Private Substrate Blockchains Using Takagi-Sugeno Fuzzy Inference
Private blockchain networks run with fixed node configurations that cannot adapt to changing workload conditions. Too many nodes serving a…
Mechanistic Interpretability of LLM Jailbreaks via Internal Attribution Graphs
Large language models (LLMs) exhibit remarkable capabilities but remain highly vulnerable to adversarial prompts and jailbreak attacks. Exi…
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted,…
Efficient Safety Alignment of Language Models via Latent Personality Traits
Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alterna…
Adversarial Decoys: Misdirecting Attention-Based Defenses in ViT
Vision Transformers (ViTs) remain vulnerable to localized adversarial attacks, e.g., adversarial patches, while recent test-time defenses m…
path_boost: A Python Package for Interpretable Graph-Level Prediction using Path-Based Gradient Boosting
We present path_boost, a Python package for interpretable supervised learning on graph-structured input data. The package implements PathBo…
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
Self-attention lets each token retrieve information from the full context, but its quadratic cost in sequence length limits training and in…
Beyond Thermal Imaging: Inferring Thermophysical Properties from Time-Resolved Thermal Observations
Inferring latent physical properties from sensory observations is a fundamental challenge in machine perception. Among available sensing mo…
A Multi-cluster Boundary Learning Method for Out-of-Scope Intent Detection via MiniLM Embedding
Intent detection is a critical task that bridges human intents and system actions in human-machine interaction systems. However, there stil…
When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning
Reinforcement learning (RL) has achieved remarkable success in enhancing the reasoning capabilities of large language models (LLMs). Howeve…
3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse
Coding agents now author entire pull requests, and practitioners sharply disagree about what this does to code review: whether it becomes t…
A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents
We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stere…
Who Broke the System? Failure Localization in LLM-Based Multi-Agent Systems
Large language model (LLM) based multi-agent systems enable complex problem solving through coordinated reasoning and action, but their dis…
SpO$_2$ Predictor-Guided Stage-Wise Time-Frequency Reconstruction of Low-Quality Dual-Wavelength PPG for Oxygen Saturation Estimation
Continuous oxygen saturation (SpO$_2$) estimation from wearable photoplethysmography (PPG) is important for long-term health monitoring, bu…
Reaction-network reasoning with frontier models for experimentally confirmed catalyst-selectivity hypotheses
Catalysts are essential for sustainable chemical manufacturing, yet discovering novel architectures remains a bottleneck dominated by trial…
Beware What You Autocomplete: Forensic Attribution of Backdoored Code Completions
Large language models have enabled powerful code completion systems that assist developers by predicting subsequent lines of code. However,…
Provably Optimal Learning Algorithms for Assistance Games
This paper studies an online variant of the assistance games framework, where an informed agent and an uninformed agent repeatedly interact…
Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework
Large-Language Models (LLMs) can be prone to flawed and unfaithful reasoning that decoding strategies like Self-Consistency (SC) fail to de…
APIVOT: Adaptive Planning with Interleaved Vision-Language Thoughts
Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility. To successfully execute a ta…
Structured Pruning of Large Language Models via Power Transformation and Sign-Preserving Score Aggregation with Adaptive Feature Retention
This paper proposes an improved structured pruning method for large language models (LLMs) that addresses key challenges in adapting Adapti…
DKDNet: Dual Knowledge and Data-Driven Network for Cross-Domain Automatic Modulation Classification
The dynamics of communication environments induce significant distribution shifts across domains, challenging the generalization of deep le…
PLURAL: A Global Dataset for Value Alignment
Large language models (LLMs) are used worldwide, yet disproportionately reflect Western values, limiting their ability to represent diverse…
Aleena: Alignment Agent for Research Software Engineering Collaborations
Research software collaborations span meetings, informal chats, pull requests, and GitHub issues. A decision surfaced in a Slack thread, re…
What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness
Large language models fine-tuned for forecasting can be accurate yet poorly calibrated, and their chain-of-thought (CoT) reasoning may not…
Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA
Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Ac…
Reinforcing the Generation Order of Multimodal Masked Diffusion Models
Diffusion Language Models (DLMs) have recently achieved substantial progress in natural language generation tasks. Recent research demonstr…
Towards Efficient Large Language Model Serving: A Survey on System-Aware KV Cache Optimization
Despite the rapid advancements of large language models (LLMs), LLM serving systems remain memory-intensive and costly. The key-value (KV)…
When Thinking Hurts: Epistemic Signals in the Reasoning Chains of Visual Language Models
Uncertainty quantification for visual language models (VLMs) conventionally targets the answer token distribution. We provide the first thr…
COBART: Controlled, Optimized, Bidirectional and Auto-Regressive Transformer for Ad Headline Generation
Online ads are essential to all businesses and ad headlines are one of their core creative component. Existing methods can generate headlin…
LDFE: Laplacian Decoupled Feature Enhancement Block for Dual-Stream CNN-based RGB-IR Object Detection
The complementary information between RGB and IR images can significantly enhance object detection performance under extreme conditions. Ex…
Deep Learning Method for Stationary Distribution of Reflected Brownian Motion
The stationary distribution of reflected Brownian motion (RBM) plays an important role in the analysis of high-dimensional stochastic syste…
PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction
Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora a…
ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents
We present the results of HIPE-OCRepair-2026, an ICDAR competition on LLM-assisted OCR post-correction of historical documents. OCR post-co…
Prismata: Confining Cross-Site Prompt Injection in Web Agents
Autonomous web agents promise to automate everyday browsing tasks, but inherit one of the web's oldest attack surfaces. Cross-Site Scriptin…
LEXIC: Lightweight Eye-tracking eXtension via Injected Complexity
On the recent EyeBench benchmark, predicting reading comprehension from eye movements exposes a stark gap: text-aware models using pretrain…
ProsMAE: Multi-Source MAE Pretraining for ISUP Grade Classification
Whole slide images (WSIs) provide rich diagnostic information for computational pathology, but their gigapixel scale, stain variation, scan…
Out of Sight: Compression-Aware Content Protection against Agentic Crawlers
The rise of LLM-based agents with reasoning, summarization, and memory capabilities has created a new threat surface for online content tha…
LEEVLA: Seeing What Matters in Latent Environment Evolution for Vision-Language-Action
Vision-language-action (VLA) models aim to map multimodal inputs to robot actions. However, most existing approaches struggle to cover comp…
Leveraging Color Naming for Image Enhancement
Enhancing images to make them visually appealing is a persistent challenge in computer vision. Many deep-learning methods train models on p…
Open-ended Multi-agent Autocurricula via Visual Inspection of Policies with Multi-modal LLMs
Open-ended curricula in Reinforcement Learning (RL) aim to train generally-capable agents by identifying tasks that facilitate learning inc…
TMI: Text-to-Image Meets Image-to-Image for Complementary Data Synthesis to Boost Long-Tailed Instance Segmentation
Large-vocabulary instance segmentation is constrained by long-tailed category distributions and fine-grained inter-class ambiguity. While d…
RhyMix: A Lightweight Adaptive Multi-Rhythm Network for Long-Term Time Series Forecasting
Real-world time series exhibit complex dynamics characterized by multiple simultaneous temporal patterns: short-term fluctuations, periodic…
Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment
Best-of-$N$ (BoN) inference improves content consistency in zero-shot text-to-speech by selecting from $N$ candidates with an automatic spe…
Multi-Agent Firewall Architecture for Privacy Protection of Sensitive Data in Interactions with Language Models
While Large Language Models (LLMs) have become essential productivity tools, their integration into workflows without adequate safeguards c…
From Legacy Documentation to OSCAL: An MCP-Based Agent Pipeline for Threat-Informed Continuous Compliance in Critical Infrastructure
In critical infrastructure, operational technology environments often cannot be actively scanned, and yet active system feedback is needed…
GitLake: Git-for-data for the agentic lakehouse
We present GitLake, a Git-for-data design for an agent-first lakehouse. The system lifts single-table Iceberg snapshots into lakehouse-wide…
ArtMine: Discovering and Formalizing Artistic Processes
Understanding how artworks are created requires reasoning about the iterative decisions, material operations, and contextual influences tha…
TypeProbe: Recovering Type Representations from Hidden States of Pre-trained Code Models
State-of-the-art code models achieve impressive performance, yet the extent to which they internally encode type information remains poorly…
Spectral Analysis of Dueling Q-Learning
Q-learning is a fundamental algorithm in reinforcement learning (RL) for solving discounted Markov decision processes (MDPs) when the trans…
FSD-VLN: Fast-Slow Dual-System Modeling for Aerial Long-Horizon Vision-Language Navigation
Vision-Language Navigation (VLN) enables UAV autonomous navigation in unknown environments by mapping language instructions to real-time vi…
On the Role of Conversational Timing in Synthetic Training Data for ASR
Synthetic multi-speaker conversations are widely used to train conversational automatic speech recognition (ASR) systems, but it remains un…
Self-Adaptive Anomaly Detection with Reinforcement Learning and Human Feedback in Connected Vehicles
Connected vehicles are autonomous cyber-physical systems whose behavior must be continuously monitored during operation to detect deviation…
Large-Language-Models-as-a-Judge in Theory-Agnostic Adaptive Metric-Alignment for Prototypical Networks in Personality Recognition
Personality recognition has traditionally been constrained by theory-dependent formulations, where models are trained to fit predefined psy…
WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving
Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving. However, existing methods either lack comprehensive world…
TRACE: A Two-Channel Robust Attribution Watermark via Complementary Embeddings for LLM-Agent Trajectories
LLM agents reach users through resellers, who may rebrand a developer's agent or substitute a cheaper model. When provenance is disputed, a…
Swapping Faces, Saving Features: A Dual-Purpose Pipeline for Pedestrian Privacy in ITS
Large-scale and diverse datasets are needed to train AI models to take real-time decisions for autonomous vehicles (AVs), an intelligent tr…
DrugGen 2: A disease-aware language model for enhancing drug discovery
Current computational approaches for drug design typically focus on generating molecules conditioned on specific targets or general molecul…
Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery
Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive sur…
When Synthetic Speech Is All You Have: Better Call GRPO
LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to col…
Predicting Male Fertility Using Machine Learning: A Semen Parameters Based Analysis with the VISEM Dataset
Male infertility is a significant yet often underdiagnosed aspect of reproductive health, with semen analysis serving as the cornerstone of…
EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data
Egocentric human data offers scalable supervision for robot manipulation. However, behavior cloning entangles transferable content like obj…
ADORN: Adaptive Drift handling for Open RAN using Reinforcement Learning
Dynamic traffic variations in Open Radio Access Networks (O-RAN) lead to drift, which degrades the performance of Artificial Intelligence/M…
Spatio-Temporal Scheduling Prediction Under Backhaul Delay for Resilient Coordinated Beamforming
Coordinated beamforming in distributed 5G networks relies on the timely exchange of inter-cell scheduling information, but backhaul latency…
Two Axes of LLM Abstention: Answer Correctness and Question Answerability
A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable one…
VEGAS: Human-Aligned Video Caption Evaluation via Gaze
Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. W…
The Context Access Divide: Interaction-Level Architecture as a Complementary Dimension of Agentic Inequality
Sharp et al. (2025) introduce "agentic inequality" as a framework for analyzing disparities in access to AI agents across three dimensions:…
Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing
Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing.…
When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability
An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evalua…
DocMaster: A Hierarchical Structure-Aware System for Document Analysis
Leveraging large language models (LLMs) to analyze complex documents -- such as academic papers, technical manuals, and financial reports -…
VocaDet: Sample-Driven Open-Vocabulary Object Detection and Segmentation via Visual Tokenization and Vector Database Retrieval
Open-vocabulary object detection and segmentation aim to recognize arbitrary objects beyond predefined categories. Although recent vision-l…
SMetric: Rethink LLM Scheduling for Serving Agents with Balanced Session-centric Scheduling
LLM scheduling is critical to serving, yet it remains unclear how well existing designs fit agentic serving--with LLM requests issued by ag…
When Structured Sparse Autoencoders Learn Consistent Concepts Across Modalities
Sparse autoencoders (SAEs) have emerged as a promising technique for mechanistic interpretability by learning a set of sparse latent featur…
UltraX: Refining Pre-Training Data at Scale with Adaptive Programmatic Editing
As available training data approaches its physical limit, gains from Scaling Laws have begun to diminish. Consequently, improving Large Lan…
Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning
As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands rewar…
WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search
Large language model (LLM)-based web search agents are transforming information seeking from simple factoid question answering into complex…
A Practical Investigation of Training-free Relaxed Speculative Decoding
Speculative decoding accelerates sampling from an autoregressive LLM by using a faster auxiliary model to draft tokens which are then verif…
ProjAgent: Procedural Similarity Retrieval for Repository-Level Code Generation
Repository-level code generation requires implementing target functions while accounting for complex cross-file dependencies and project-sp…
Pose-to-Biomechanics: Bridging 3D Human Pose Estimation and Biomechanical Attribute Prediction
Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and scalable. However, mo…
Validity of LLMs as data annotators: AMALIA on authority
A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA…
Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph
While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlook…
SLORR: Simple and Efficient In-Training Low-Rank Regularization
Low-rank factorization is widely used to compress neural networks, but modern models are often not naturally amenable to aggressive factori…
OpenCoF: Learning to Reason Through Video Generation
Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Rec…
IFAR: Multi-Perspective and Multi-Level Causal Discovery with LLMs
Large language models (LLMs) have developed rapidly, and their reasoning capabilities have become a hot research topic. However, there is s…
Goal-Driven Reasoning in DatalogMTL with Magic Sets
DatalogMTL is a powerful rule-based language for temporal reasoning. Due to its high expressive power and flexible modeling capabilities, i…
Dual-Difficulty Curriculum Learning for Direct Preference Optimization
Curriculum learning enhances Direct Preference Optimization (DPO) for aligning Large Language Models (LLMs), yet existing methods rely on a…
A Vision Toward Energy-Efficient Domain-Specific Artificial Intelligence Models and Agents
The field of artificial intelligence (AI) has taken a tight hold on broad aspects of society, industry, business, and governance in ways th…
MetaHGNIE: Meta-Path Induced Hypergraph Contrastive Learning in Heterogeneous Knowledge Graphs
Estimating node importance in heterogeneous knowledge graphs is a fundamental problem underlying recommendation, search, and knowledge deci…
SimRPD: Optimizing Recruitment Proactive Dialogue Agents through Simulator-Based Data Evaluation and Selection
Task-oriented proactive dialogue agents play a pivotal role in recruitment, particularly for steering conversations towards specific busine…
Conversational AI for Rapid Scientific Prototyping: A Case Study on ESA's ELOPE Competition
Large language models (LLMs) are increasingly used as coding partners, yet their role in accelerating scientific discovery remains underexp…
RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments
Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decis…
Retrieval-Augmented Generation Must Move Beyond Factual Grounding to Represent Diverse Opinions
This position paper argues that Retrieval-Augmented Generation (RAG) systems exhibit a factual bias-optimizing for epistemic uncertainty re…
The Power of Power Law: Asymmetry Enables Compositional Reasoning
Natural language data follows a power-law distribution, with most knowledge and skills appearing at very low frequency. While a common intu…
SHARP: Sleep-based Hierarchical Accelerated Replay for Long Range Non-Stationary Temporal Pattern Recognition
Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming…
Deployment-Time Memorization in Foundation-Model Agents
Foundation-model agents are increasingly long-lived systems that remember users across interactions, making memorization an explicit deploy…
LiteOdyssey: A Lightweight Reasoning AI Agent for Interpretable Rare-Disease Diagnosis
Rare disease diagnosis involves interpreting clinical and genetic findings through complex diagnostic reasoning. We investigated whether th…
TNODEV: Toolbox for Neural ODE Verification
Neural ordinary differential equations (neural ODE) gained attention in safety critical settings such as continuous-time controllers for cy…
InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy
Large language models are increasingly deployed as investment research assistants, yet no benchmark tests whether they can accurately recon…
Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stak…
Theoria: Rewrite-Acceptability Verification over Informal Reasoning States
When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; sc…
Distributed Attacks in Persistent-State AI Control
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This per…
Accurate Portraits of Scientific Resources and Knowledge Service Components
With the advent of the cloud computing era, the cost of creating, capturing, and managing information has gradually decreased. The amount o…
Knowledge Graph and Accurate Portrait Construction of Scientific and Technological Academic Conferences
In recent years, with the continuous progress of science and technology, the number of scientific research achievements has increased rapid…
Retrieval of Scientific and Technological Resources for Experts and Scholars
Institutions of higher learning, research institutes and other scientific research units have abundant scientific and technological resourc…
The Contribution of XAI for the Safe Development and Certification of AI: An Expert-Based Analysis
Developing and certifying safe - or so-called trustworthy - AI has become an increasingly salient issue, especially in light of upcoming re…
ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation
Large language models (LLMs) integrated with retrieval-augmented generation (RAG) have improved factuality by grounding outputs in external…
Simulator Ensembles for Trustworthy Autonomous Driving Systems Testing
Scenario-based testing with driving simulators is extensively used to identify failing conditions of automated driving assistance systems (…
Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs
Vision-Language Models (VLMs) have demonstrated exceptional performance in various multi-modal tasks. Recently, there has been an increasin…
ToDMA: Large Model-Driven Massive Token Communications for Semantic Multiple Access
Token communications (TokenCom) is an emerging generative semantic communication paradigm, where tokens serve as compact representation uni…
Less Is More: Reducing Token Counts Without Compromising Performance
Tokenization directly affects the inference efficiency of large language models, since fragmented tokenization increases sequence length an…
How Causal Abstraction Underpins Computational Explanation
Explanations of cognitive behavior often appeal to computations over representations. What does it take for a system to implement a given c…
TOPO-Bench: An Open-Source Topological Mapping Evaluation Framework with Quantifiable Perceptual Aliasing
Topological mapping offers a compact and robust representation for navigation, but progress in the field is hindered by the lack of standar…
MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
Medical decision systems increasingly rely on data from multiple sources to ensure reliable and unbiased diagnosis. However, existing multi…
Adaptive Generation of Bias-Eliciting Questions for LLMs
Large language models (LLMs) are now widely deployed in user-facing applications, reaching hundreds of millions of users worldwide. Despite…
OREN: Octree Residual Network for Real-Time Euclidean Signed Distance Mapping
Reconstructing signed distance functions (SDFs) from point cloud data benefits many robot autonomy capabilities, including localization, ma…
MSRNet: A Multi-Scale Recursive Network for Camouflaged Object Detection
Camouflaged object detection is an emerging and challenging computer vision task that requires identifying and segmenting objects that blen…
Deep Neural Networks as Discrete Dynamical Systems: Implications for Physics-Informed Learning
We revisit the analogy between feed-forward deep neural networks (DNNs) and discrete dynamical systems derived from neural integral equatio…
CoCo-Fed: A Unified Framework for Memory- and Communication-Efficient Federated Learning at the Wireless Edge
The deployment of large-scale neural networks within the Open Radio Access Network (O-RAN) architecture is pivotal for enabling native edge…
V-VLAPS: Value-Guided Planning for Vision-Language-Action Models
Vision-language-action (VLA) models provide strong action priors for robotic manipulation, but their reactive behavior can fail under distr…
SwinIFS: Landmark Guided Swin Transformer For Identity Preserving Face Super Resolution
Face super-resolution aims to recover high-quality facial images from severely degraded low-resolution inputs, but remains challenging due…
Bridging Cognitive Neuroscience and Graph Intelligence: Hippocampus-Inspired Multi-View Hypergraph Learning for Web Finance Fraud
Online financial services constitute an essential component of contemporary web ecosystems, yet their openness introduces substantial expos…
GenDA: Generative Data Assimilation on Complex Urban Areas via Classifier-Free Diffusion Guidance
Urban wind flow reconstruction is essential for assessing air quality, heat dispersion, and pedestrian comfort, yet remains challenging whe…
XFACTORS: Disentangled Information Bottleneck via Contrastive Supervision
Disentangled representation learning aims to map independent factors of variation to independent representation components. On one hand, pu…
Rethinking LLM-as-a-Judge: Representation-as-a-Judge with Small Language Models via Semantic Capacity Asymmetry
Large language models (LLMs) are widely used as reference-free evaluators via prompting, but this "LLM-as-a-Judge" paradigm is costly, opaq…
Towards Isolated Interventions via Almost Orthogonal Features in Language Models
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in acti…
Self-EvolveRec: Self-Evolving Recommender Systems with LLM-based Directional Feedback
Traditional methods for automating recommender system design, such as Neural Architecture Search (NAS), are often constrained by a fixed se…
An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation
Vision-Language Models (VLMs) are increasingly used in document processing pipelines to convert flowchart images into structured code (e.g.…
Curriculum Learning for Efficient Chain-of-Thought Distillation via Structure-Aware Masking and GRPO
Distilling Chain-of-Thought (CoT) reasoning from large language models into compact student models presents a fundamental challenge: teache…
Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
Layer-wise capacity in large language models is highly non-uniform: some layers contribute disproportionately to loss reduction, whereas ot…
Robust Weighted Triangulation of Causal Effects Under Model Uncertainty
A fundamental challenge in causal inference with observational data is correct specification of a causal model. When there is model uncerta…
Safe Flow Q-Learning: Offline Safe Reinforcement Learning with Reachability-Based Flow Policies
Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints. Existing me…
PhasorFlow: A Python Library for Unit Circle Based Computing
We present PhasorFlow, an open-source Python library for computing on the $S^1$ unit circle. Inputs are encoded as complex phasors $z=e^{i\…
The Phasor Transformer: Resolving Attention Bottlenecks on the Unit Circle
Transformer models have redefined sequence learning, yet dot-product self-attention introduces a quadratic token-mixing bottleneck for long…
StateLinFormer: Stateful Training Enhancing Long-term Memory in Navigation
Effective navigation intelligence relies on long-term memory to support both immediate generalization and sustained adaptation. However, ex…
Echoes: A semantically-aligned music deepfake detection dataset
We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provid…
Neural Harmonic Textures for High-Quality Primitive Based Neural Reconstruction
Primitive-based methods such as 3D Gaussian Splatting have recently become the state-of-the-art for novel-view synthesis and related recons…
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows affects feed-forward 3D reconstruction.…
Persona Matters: Effects of Activation Steering on Short Answer Generation and Scoring
Activation-based steering enables inference-time personalization of large language models, but its effects in educational applications are…
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models
Vision-Language Models (VLMs) have revolutionized multi-modal learning by jointly processing visual and textual information. Yet, they face…
Peer-Predictive Self-Training for Language Model Reasoning
Mechanisms for continued self-improvement of language models without external supervision remain an open challenge. We propose Peer-Predict…
Predicting Scale-Up of Metal-Organic Framework Syntheses with Large Language Models
Scalable synthesis remains the gate between MOF discovery and industrial deployment, as scale-up know-how is fragmented across disparate re…
DeepTutor: Towards Agentic Personalized Tutoring
Education is one of the most promising real-world applications for Large Language Models (LLMs). However, current LLMs rely on static pre-t…
LoKA: Low-precision Kernel Applications for Recommendation Models At Scale
Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to larg…
Second-Order Actor-Critic Methods for Discounted MDPs via Policy Hessian Decomposition
We address the discounted reward setting in reinforcement learning (RL). To mitigate the value approximation challenges in policy gradient…
MasFACT: Continual Multi-Agent Topology Learning via Geometry-Aware Posterior Transfer
Multi-agent systems (MAS) powered by large language models (LLMs) have emerged as a powerful paradigm for complex problem solving, where pe…
CriterAlign: Criterion-Centric Rationale Alignment for Code Preference Judging
Pairwise human preference prediction is central to evaluating code-generation systems, where quality often depends on task-specific trade-o…
MAVEN: A Multi-stage Agentic Annotation Pipeline for Video Reasoning Tasks
Training Vision Language Models (VLMs) for video event reasoning requires high-quality structured annotations capturing not only what happe…
Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory
Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this…
When RLHF Fails: A Mechanistic Taxonomy of Reward Hacking, Collapse, and Evaluator Gaming
RLHF evaluation should track how failures emerge, where they localize, and which warning signals appear before external quality degrades. W…
Autonomous heterogeneous catalyst discovery with a self-evolving multi-agent digital twin
Theoretical heterogeneous catalysis promises rapid catalyst discovery, yet computational and machine-learning predictions often deviate fro…
Temporal Preference Concepts and their Functions in a Large Language Model
Large Language Models (LLMs) are increasingly being deployed to make decisions that require trading off near-term gains against long-term c…
EasyLens: A Training-Free Plug-and-Play Subtle-Lesion Representation Amplifier for Medical Vision-Language Models
Medical vision-language models (VLMs) have shown increasing potential for clinical image interpretation, including lesion detection and rep…
MeCo: One-Step MeanFlow-based Corrector for Multi-Channel Speech Separation
While discriminative models for multi-channel speech separation excel in reference-based metrics, they often exhibit suboptimal human liste…
KG-SoftMAP: Soft Knowledge-Graph Priors for Bayesian Network Structure Learning from Sparse Discrete Data
Learning Bayesian network (BN) structure from sparse discrete data is hard: when each instance records only a few variables, most variable…
Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement
Interleaved thinking, where a unified multimodal model alternates between textual reasoning and visual generation, has shown promise on spa…
Training and Evaluating Diffusion Policies with Long Context Lengths
Imitation learning has enabled highly-dexterous robotic manipulation from RGB observations. Policies trained with these methods, however, t…
Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution
Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments…
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study
Mixture-of-Experts (MoE) language models are often described as ideal for resource-constrained inference. Each token activates only a small…
Improving Engine Sound Analysis in Hot-Test Environments via a RAB-U-Net (Residual Attention Block U-Net) Noise Removal Method
During hot tests on a production line, engine-sound analysis is crucial to ensuring product quality and performance. However, background no…
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding
Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spa…
Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?
Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out o…
How to Leverage Synthetic Speech for LLM-Based ASR Systems?
In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic spe…
Generalization in offline RL: The structure is more important than the amount of pessimism
While pessimism counteracts overestimation bias in offline reinforcement learning (RL), being overly conservative has been associated with…
OpenAI、Codexと「ChatGPT Work」の利用制限リセット 競合Claudeとリセット合戦に
Anthropicは同日、OpenAIの制限リセットに先立ち、Claudeの利用制限をリセットしていた。
OpenAI「GPT-5.6」一般公開 最上位モデルは「コスト半分で一部Fable 5超え」うたう
OpenAIが新AIモデル群「GPT-5.6」を一般公開した。最上位の「Sol」は、Anthropicの「Claude Fable 5」を超える性能を推定コスト半分程度で実現するとうたう。
デジタル庁、tsuzumiなど国産AIを「さくらのクラウド」で稼働 「日本の自律性確保」目指す
デジタル庁は、政府職員が利用するAI基盤「源内」の実証実験の一環として、国産AIモデルを国産クラウドで稼働させると発表した。さくらインターネットが提供するクラウドサービス「さくらのクラウド」を活用する。
生成AI「若手有利」は大間違い? ミドル層が勝ち抜くための“力の見せ所”とは
新しいテクノロジーは若い世代のもの──生成AI時代において、この構造が大きく変わろうとしている。つまり、業界に長く向き合ってきたミドル層にとって、大きなチャンスといえる。
Claude、利用制限を全リセット 競合「GPT-5.6」公開と同日……OpenAI幹部「ビビってるね」
OpenAIが「GPT-5.6」を一般公開した同日、AnthropicがClaude全ユーザーの利用制限を一斉リセット。OpenAI幹部の「ビビってるね」という煽り返信も話題だ。
OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
OpenAI's new family of models will continue to power Microsoft's suite of workplace and productivity apps.
Fidji Simo steps down from OpenAI’s no. 2 role
OpenAI's No. 2 executive, Fidji Simo, is stepping down from her full-time role after her medical leave proved longer than expected — a lead…
NECが3年で100億円狙うAnthropic協業の初ソリューション、AIが販売戦略立案
NECは、Anthropicとの協業第1弾として、購買データからの商品企画や販促プラン作成を完全自動化する新サービスを開始した。専門人材なしで迅速な施策立案を可能にし、3年間で100億円の売上を目指す。
OpenAI launches its new family of models with GPT-5.6
OpenAI's latest family of models promises improvements across a range of areas, including cybersecurity.
Meta、マルチモーダル推論モデル「Muse Spark 1.1」公開 低価格の「Meta Model API」も提供へ
Metaは、AIモデル「Muse Spark 1.1」を発表した。初代モデルを強化したマルチモーダル推論モデルで、ツール操作や複雑なコーディングなどのエージェント能力が向上している。新設の「Meta Model API」を通じてパブリックプレビューとして外部向けに提供され、競合…
An AI agent startup just let its agent run its $100M fundraise
Lyzr, a startup that builds AI agents for enterprises, used its own AI agent to raise a $100 million round — proof, evidently, that the pro…
OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
OpenAI is sunsetting its AI-powered browser after less than a year. But it's moving some agentic browsing features to its desktop app and a…
AIに企業理念を宿せるか? スポーツ小売り・ヒマラヤ、接客ノウハウまで学習した「AI副店長」開発の舞台裏
スポーツ用品店をチェーン展開するヒマラヤ(岐阜市)は「アイダ つなぐ」と名付けた“AI副店長”を6月から全99店舗に導入した。同社店舗運営部の金丸智功氏(部室長代理)とヒマラヤスポーツ本館店長の高島雅央氏に、AI副店長開発の舞台裏を聞いた。
「Notion」でAI利用量が4倍に スター精密が全社にAIを定着させた業務基盤の作り方
スター精密は、国内約550人を対象に「Notion」と「Notion AI」を導入し、情報基盤を刷新した。情報を一元化するとともにAIを業務へ組み込むことで、利用率約90%、AI利用量は従来のチャット型AIツール比で約4倍を達成した。
コスト削減にAIは効果なし? 利益が変わらない企業が過半、手直しも課題に
AI活用企業507人調査で、売上高拡大を目的にAIを導入した企業の約7割が営業利益増を実感した。一方、コスト削減を目的とした導入では利益不変が過半となり、プロンプト格差や手直し、データ連携制約などの課題も浮かび上がった。
「誰にも会わずに帰る店」の寂しさ すかいらーくがロボット配膳の先に挑むAI接客
デジタル化が進み、人と接することなく食事を終えられる飲食店が増えている。その利便性の裏で失われつつある「人ならではの価値」をどう取り戻すのか。すかいらーくホールディングスは、AIを人の代替ではなく、人の価値を引き出す道具として活用し始めている。
Elon Musk praises Mythos/Fable, promises not to ‘cut off’ Anthropic
Should Anthropic trust Elon Musk to host its models? With about $40 billion in revenue at stake, Musk insists that the company can.
Can AI answer the $3 trillion question?
The AI ROI debate has returned and the numbers are even bigger, as are, perhaps, the consequences.
1万9000人が利用するソフトバンクの「全社RAG基盤」 構築の泥臭い舞台裏
AI活用で激突する「現場の利便性」v.s.「会社の安全性」。RAGの乱立に直面したソフトバンクが、ガバナンスをシステムに組み込み、数万時間相当の業務削減効果(社内の試算による)を達成した「全社RAG基盤」構築の舞台裏と、そこから得られた気付きを共有します。
Meta enters the crowded AI coding battle with Muse Spark 1.1
Meta's pitch to users is Spark's ability to handle large agentic workloads, fix bugs, and help with large code migrations — the kind of aut…
デスクトップ版ChatGPT大幅刷新 AIエージェント「Codex」統合、「ChatGPT Work」に
米OpenAIが、チャットAIサービス「ChatGPT」とAIコーディングエージェント「Codex」を統合した「ChatGPT Work」を発表した。「GPT-5.6」シリーズを搭載する。最終成果物を指定するだけで、あいまいな状況にも適応しながら最小限の指示で洗練されたアウトプ…
New York Times says OpenAI hid evidence in ChatGPT copyright trial
News publishers say OpenAI hid tools and datasets that could identify copyrighted journalism in ChatGPT outputs, escalating their lawsuit w…
Google will now disclose which ads are made with AI
While Google prohibits misleading and deceptive ads, an ad can still leverage AI to create some type of synthetic or digitally altered cont…
Paris-based AI voice startup Gradium raises $100M seed, backed by Nvidia
The company is using the cash to open an office in the Bay Area and compete for talent there, "strengthening its position at the heart of t…
How did the government decide OpenAI’s frontier model was safe to release?
"Exactly what that dialog looked like between the government and Anthropic and OpenAI is unclear."
Instagram users: Here’s how to stop Meta’s AI from using your photos
Muse Image allows users to generate AI images using photos from public Instagram accounts. As long as a person's profile is public, another…
Meta’s new AI chips will begin production in September
The company is taking a modular approach to designing these chips, anticipating that their needs will change as AI evolves rapidly by the t…
Nvidia is a victim of the compute marketplace it created
Having proven how valuable compute can be, the company finds itself at the center of a market everyone wants to be in — while simpler techn…
2026-07-09(242件)
Anthropic’s new Claude feature is quietly selling you on AI
Claude’s new Reflect dashboard doesn’t just visualize how you use AI. It also subtly reinforces how much of your daily work now depends on…
Anthropic, OpenAI, and SpaceX are bigger than the last 25 years of tech exits
Three big AI IPOs are set to generate more value than all the U.S. VC-backed exits since 2000.
Popular open source AI developer tool Ollama raises $65M, grows to nearly 9M users
Benchmark-backed Ollama has amassed 176,000 stars, and nearly 17,000 forks on GitHub by helping developers easily run AI on their PCs.
Character.AI enters the microdrama arena with its own productions, but there’s a twist
In an interesting twist that takes advantage of the company's core product, users can chat with these shows' characters, ask them questions…
GPT-5.6 is now the preferred model in Microsoft 365 Copilot
Learn how GPT-5.6 powers Microsoft 365 Copilot with stronger AI capabilities across Word, Excel, PowerPoint, Chat, and Cowork for faster, h…
Nandan Nilekani leaves GP role at Fundamentum as it launches $200M third fund
Nilekani remains Fundamentum's anchor investor as the firm expands its leadership team and targets AI and fintech startups in India.
三菱自動車が「国産人型ロボ」量産へ 2027年に「月1000台の製造体制」 東大発スタートアップと協業
三菱自動車工業は、東京大学発のロボット開発スタートアップHighlandersと「国産人型ロボット」の開発に関して協業すると発表した。
ChatGPT is now a partner for your most ambitious work
ChatGPT Work is an agent that can take action across your apps and files, stay with a project for hours if needed, and turn a goal into fin…
GPT-5.5 Bio Bug Bounty
Details about the OpenAI Bio Bounty program
GPT-5.6: Frontier intelligence that scales with your ambition
More intelligence from every token, stronger performance per dollar, and more capability on demand for your hardest work.
AI活用、最大のボトルネックは「経営層」か トップ不使用の企業、85.7%が「方針・体制なし」
ラクスルが中小企業を対象に実施したAI活用の実態調査で、経営層がAIを全く使わない企業の85.7%は活用の方針も推進体制もないと分かった。経営層の姿勢がAI格差を左右する構図が浮かんだ。
NEC、Claudeを活用した「全自動」マーケティングサービス開始 3年で売上100億円目指す
NECは、消費者の購買データから商品企画や販促プラン作成の「完全自動化」をうたうサービスの提供を開始した。AnthropicのAI「Claude」などを活用した。3年間で累計100億円の売り上げを目指す。
社内のWindows環境で「数百のAIエージェント」を隔離実行 NVIDIAとMicrosoftが共同開発したデスクサイドマシンの全容
NVIDIAは、Windows環境でAIエージェントを開発・実行するデスクサイドAIスーパーコンピュータ「NVIDIA DGX Station for Windows」を発表した。NVIDIA GB300 Grace Blackwell Ultra Desktop Superc…
AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit…
When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning
Training large language models (LLMs) with extended reasoning has enabled in-context search, in which models iteratively generate, critique…
LLM-powered reasoning in agent-based modeling
Agent-based modeling (ABM) has the capability to model millions of individuals and their interactions, which is useful for policy making. H…
QANTIS: Hardware-Calibrated Sequential POMDP Belief Updates on IBM Heron
Autonomous systems under partial observability act on beliefs, not raw sensor events. QANTIS treats the quantum processor as a calibrated b…
Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1
Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (…
Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics
Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra S…
The Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
Agentic AI development today runs on token maxing: buying capability with tokens -- longer reasoning traces, more turns, wider tool payload…
Grounding Spatial Relations in a Compact World Model: Instruction Leakage and a Goal-Free Dynamics Fix
Compact world models that condition on a language goal promise to ground relations such as ``put the red block left of the blue block'' usi…
Large Behavior Model: A Promptable Digital Twin of the Retail Customer
Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive ac…
Learning social norms enhances compatibility in dynamic human-AI coordination
Humans continuously coordinate with others in dynamic interactions, often through implicit, hard-to-quantify social norms that act as share…
Measuring Intelligence Beyond Human Scale
How can we measure intelligence beyond human capability? Human-authored benchmarks saturate, and above human capability, examiners may not…
Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety
Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a…
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks
Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they ca…
Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations
Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the pr…
From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents
Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing age…
Physics-Audited Agentic Discovery in Scientific Machine Learning
In agentic scientific machine learning (SciML), large language model (LLM) agents can discover surrogate models and select one by an automa…
MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning
Mathematical reasoning benchmarks typically provide all facts needed to solve each problem, while interactive benchmarks often mix reasonin…
Agentic Data Environments
Autonomous agents promise substantial gains in speed, scale, and labor efficiency, but their failures can impose abrupt and often irreversi…
Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In poli…
InductWave: Inductive Multi-Hop Logical Query Answering on Knowledge Graphs
Logical Multi-Hop Query Answering over Knowledge Graphs (KGs) can be formulated as querying, with an implicit completeness assumption. Curr…
The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents
A self-evolving agent retires its bad skills by watching them fail, so what happens when the judge cannot see the failures? Skill retiremen…
SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis
Spatial and Single-cell transcriptomics are transformative in deciphering cellular dynamics. As the fundamental paradigm for reconstructing…
Search, Fail, Recover: A Training Framework for Correction-Aware Reasoning
Many reasoning tasks are not well described by a single left-to-right chain: a solver may need to pursue a plausible branch, observe delaye…
Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows
Product data scientists often ask LLM-based agents to help with recurring execution tasks such as cleaning data, writing SQL, choosing stat…
RL Post-Training Builds Compositional Reasoning Strategies
Does RL post-training merely amplify primitive skills already latent in a base model, or can it compose primitive skills into new higher-le…
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops
AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, train…
SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents
Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operational knowledge to make th…
Institutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety
We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectiv…
Can Reinforcement Learning Efficiently Discover Price Manipulation?
In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than…
LipSSD: Lipschitz-Constrained Single-Shot Detection for Adversarially Robust Object Detection
Object detectors have many applications in safety-critical systems, but they are known to be sensitive to worst-case perturbations such as…
When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents
Personal AI agents powered by large language models can reason and act using available tools to access emails, manage calendars, and push c…
Non-contact, Real-time, Heart-rate Measurement using Image Processing with Commodity Cameras and AI Agents
Heart rate measurement is one of the key requirements for real-time health monitoring, in particular for health caring of elderly people. T…
MiLSD: A Micro Line-Segment Detector for Resource-Constrained Devices
Line segment detection is a key building block in visual SLAM, 3D reconstruction, and industrial inspection. Recent deep learning methods h…
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation
Conditional computation can decouple language model quality from per-token inference cost, yet leading techniques act on a single axis in i…
Do Counterfactually Fair Image Classifiers Satisfy Group Fairness? -- A Theoretical and Empirical Study
The notion of algorithmic fairness has been actively explored from various aspects of fairness, such as counterfactual fairness (CF) and gr…
NEST: Tackling Dataset-Level Distribution Shifts via Regime-Oriented Mixture-of-Experts
Accurate long-term forecasting in complex systems is frequently compromised by dataset-level distribution shifts, where diverse underlying…
Security and Privacy in Agentic AI: Grand Challenges and Future Directions
We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise th…
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
We propose D2PO (Dynamic Direct Preference Optimization), a principled framework for optimizing diffusion sampling policies with respect to…
Deep Reinforcement Learning for Reliability Based Bi-Objective Portfolio Optimization
Portfolio optimization under uncertainty is inherently a multi-objective decision problem involving complex interactions among return, risk…
Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts
Automatically recognizing the sentiment, positive or negative, from speech is a challenging task, requiring both the analysis of vocal infl…
PRoVeFL: Private Robust and Verifiable Aggregation in Federated Learning
Federated Learning (FL) enables multiple clients to collaboratively train machine learning models while retaining data locality, thereby en…
STAGformer: A Spatio-temporal Agent Graph Transformer for Micro Mobility Demand Forecasting
Accurate station-level demand forecasting is essential for the efficient operation of bike-sharing systems, yet it remains challenging due…
WHERE to Generate Matters: Budget-Aware Synthetic Augmentation for Label Skewed Federated Learning
Label skew in federated learning (FL) causes client drift and degrades global accuracy. Synthetic data augmentation can reduce this imbalan…
Inertia-1: An Open Exploration of Wearable Motion Foundation Models
Wearable motion sensing provides a continuous and scalable window into human behavior and health, making it a natural fit for foundation mo…
Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation
Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Que…
SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts
Recent Multimodal Large Language Models (MLLMs) struggle to bridge the representational gap between 2D semantic understanding and 3D spatia…
LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting
Process industries rely on time-series forecasting and soft sensing to estimate quality variables that are hard to measure online. Labeled…
Open-Ended Scenario Reasoning for Specialist Model Adaptation
Process industries have accumulated validated specialist models, yet sensor drift, feedstock variation, and regime switching cause these mo…
Cross-Trajectory Chimera Interventions Reveal Dissociable Roles of Weight Magnitude and Direction in Grokking
Which properties of a partially trained network are causally portable to a different, independently trained network? Single-trajectory inte…
Dynamic-in-Few-Step: Unifying Dynamic Computation and Few-Step Distillation for Efficient Video Generation
Video Diffusion Models (VDMs) have demonstrated superior generation quality but suffer from prohibitive computational costs. While recent f…
ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
In this paper, we address the problem of multimodal federated learning with missing modality. Existing methods utilize an additional public…
Specification Grounding Drives Test Effectiveness for LLM Code
Large language models frequently generate code that appears correct on typical inputs yet fails on edge cases, invalid inputs, and other sp…
At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics
On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized. Reading eff…
The Rank-One Corner: How Much Value Equivalence Does a Task Need from a World Model?
A learned world model is usually judged by how faithfully it reconstructs its observations or predicts reward, as though quality were somet…
Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering
Large language models (LLMs) achieve promising results on medical question answering benchmarks, yet their use in public health is constrai…
Diffusion enabled Optimal Transport distances for graph matching
This paper introduces Diffusion Semi-Relaxed Fused Gromov-Wasserstein (DsrFGW), a novel method for graph comparison that unifies node featu…
Digital Fragmentation and Generative AI Use Across 103 Million Application Events
Knowledge workers switch between applications thousands of times per day, spending nearly a tenth of the work year transitioning between di…
tsbootstrap: Distribution-Free Uncertainty Quantification and Conformal Prediction for Time Series
Finance, sensing, and demand streams violate the exchangeability that IID conformal prediction and the IID bootstrap assume, and existing l…
SPEAR: A Simulator for Photorealistic Embodied AI Research
Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photoreal…
Vision Language Action (VLA) Models for Unmanned Aerial Robotics and Bimanual Manipulation: A Review
Vision Language Action (VLA) models unify visual perception, natural-language understanding, and action generation within a single foundati…
Reliable and Developer-Aligned Evaluation of Agents for Software Engineering
Large language models are rapidly moving towards closing the development cycle, transitioning from simple assistive companions to autonomou…
A Continual Learning Framework for Adaptive Control of Modular Soft Robots
Soft robots have attracted significant attention in applications such as medical intervention, rehabilitation, and robotic manipulation due…
SmartHomeSecure: Automated Detection and Repair of Smart Home Configuration Errors Using Large Language Models
Smart home automation platforms increasingly rely on user-authored YAML configuration files to define device behaviors, but these files are…
AirPASS: Over-the-Air Federated Learning via Pinching Antenna Systems
This paper investigates over-the-air federated learning (AirFL) in wireless systems where the access point is equipped with a multi-wavegui…
From Agentic to Autogenic Network Management for AI-Native 6G and Beyond: A Standards Perspective
Standards bodies, including TM Forum, 3GPP, and ETSI, are converging on Agentic AI as the foundation for next-generation network management…
Enhancing deep learning models for time series classification via knowledge distillation
Deep learning has achieved remarkable success in various domains including time series analysis, computer vision and natural language proce…
What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study
Evaluating uncertainty in AI-generated SQL queries requires estimating whether a query is correct, where correct means it executes to the s…
A Multi-Analyst LLM Pipeline for Auditable Rule Discovery Across 68 Public Physiological Corpora
Open physiological corpora are heterogeneous: they use different sensors, labels, sampling rates, recording settings, and clinical endpoint…
When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems
While enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulne…
Ad Headline Generation using Self-Critical Masked Language Model
For any E-commerce website it is a nontrivial problem to build enduring advertisements that attract shoppers. It is hard to pass the creati…
Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
Speech-to-text alignment means finding the temporal boundaries of each word in the audio. Some models provide such an alignment directly an…
A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong
Reinforcement learning agents for imperfect-information card games are only as strong as the opponents they train against, and they are har…
GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model
Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke act…
ReMoDEx: A Local-to-Global Relevance-Based Model Decision Explainability Framework for large-Scale Image Datasets
Deep learning image classifiers achieve strong predictive performance yet remain opaque in how decisions are formed. A model may predict co…
Computing with Stochastic Oracles in AI-Augmented Computation
The Stochastic-Oracle Turing Machine (SOTM) framework models AI-augmented computation as the interaction of a probabilistic Turing machine…
LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models
Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM ad…
MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations
Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human pe…
Imputation Meets Clustering: Exploiting Latent Subgroup Structure for Missing Data Recovery
Missing data is prevalent in practical applications, making effective imputation an essential preprocessing step for downstream analysis. R…
Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System
The remarkable performance of large language models (LLMs) in linguistic tasks underscores an urgent need for comprehensive evaluation of t…
Self-Supervised Pretraining Improves Cross-Site and Cross-Scale Robustness of Point Cloud Leaf-Wood Segmentation
The accuracy of existing leaf-wood segmentation methods for tree point clouds varies across forest types and sites. Self-supervised learnin…
Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies
Large Language Models (LLMs) and generative AI (GenAI) systems, such as ChatGPT, Claude, Gemini, LLaMA, Copilot, Stable Diffusion by OpenAI…
End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent
Bridging the gap between human pilot intent and autonomous flight operation is critical for real-world electric vertical takeoff and landin…
Hybrid Least Squares/Gradient Descent Methods for MIONets
In this paper, we propose an efficient hybrid least squares/gradient descent (LSGD) method for MIONets to accelerate training. This method…
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additiona…
Physics-guided spatiotemporal neural models for fuel density prediction
This paper presents a physics-guided machine learning (PGML) framework for fuel density prediction, integrating physics constraints and dom…
Multimodal Spatiotemporal-Frequency Fusion with Peak Enhancement for Cellular Traffic Forecasting
Accurate forecasting of cellular network traffic is essential for network planning, resource allocation, and quality-of-service assurance i…
Latent graph encoding of multimodal neuroimaging features with generative AI architectures
While generative models enable encoding of complex neuroimaging data for feature generation and reconstruction, developing optimal architec…
Gimitest: A Comprehensive Tool for Testing Reinforcement Learning Policies
Reinforcement learning (RL) policies can be unsafe and vulnerable to attacks. Ensuring their reliability is often a pain point as existing…
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of…
Intrinsic Green's Learning: Supervised Learning on Manifolds via Inverse PDE
We introduce Intrinsic Green's Learning (IGL), a framework that models a target function on a manifold as the solution to a linear PDE whos…
On the Principles of Deep Feedforward ReLU Networks
The architecture of deep feedforward neural networks is ubiquitous in deep learning, either as a whole system or as a subnetwork of other a…
Riemannian Geometry for Pre-trained Language Model Embeddings
Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sent…
Making Implicit Preservation Intent Explicit in Conversational Image Editing
Conversational image editing requires preserving not only visible content, but also content that temporarily disappears across turns. When…
Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production
AI agents deployed for IT operations are typically permanent cost centers because every execution requires full LLM inference, even for pre…
Complexity-Budgeted, Interaction-Aware Interpretable Model for Tabular Data
Inherently interpretable classifiers for tabular data typically rely on sparse features, rules, or patterns that users can inspect directly…
Multiplication Beyond Groups: Stratified Fourier Mechanisms in Transformer Circuits
Transformers have demonstrated a remarkable ability to learn algorithmic reasoning, yet mechanistic analyses have mostly focused on globall…
Navigating Hierarchy: Hyperbolic Learning on Brain Graphs for Disorder Diagnosis
Functional brain networks exhibit a hierarchical organization across ROI, community, and whole-brain levels, supporting local processing, i…
AT-Attn: Temporal-Aware Cross-Attention for Longitudinal Multimodal Alzheimer's Disease Diagnosis
In longitudinal Alzheimer's disease (AD) diagnosis support, clinical and imaging information is often collected at irregular visits. Integr…
GeoProp: Grounding Robot State in Vision for Generalist Manipulation
Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit al…
Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning
In text-to-image in-context learning (T2I-ICL), a model has to infer a latent compositional pattern from fewshot demonstrations for generat…
Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning
Recent breakthroughs of Reinforcement Learning (RL) have highlighted its potential for complex agentic Large Language Model (LLM) tasks. Ho…
Predicting LLM Safety Before Release by Simulating Deployment
Pre-deployment safety evaluations aim to inform the downstream risks of releasing a new AI model. Yet most evaluations provide limited evid…
Validate the Dream Before You Trust Its Verdict: Admissibility for World-Model Simulators
Across robotics, World Models (WMs) are increasingly used to evaluate action policies by simulating the consequences of actions in an imagi…
Memory Scarcity, Open Models, and the Restructuring of the AI Industry, 2026-2030 -- A quantitative scenario analysis of inference economics, training-cost divergence, and infrastructure solvency
We analyze how four forces restructure the AI industry over 2026-2030: the DRAM/HBM price surge, frontier-capable open-weight models (GLM-5…
Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation
Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation…
DiPhon: Diffusion on Graphons for Scalable Graph Generation
Diffusion models represent a leading paradigm for graph generation, with notable impact in domains such as molecular design. Yet, scaling t…
ORCAID: Oblique Rule-Based Continuous-Action Interpretation for Deep RL Policies
Explainability remains a key issue in reinforcement learning (RL). Distilling an interpretable policy from an agent trained in a complex en…
FMMVCC: Fuzzy Mamba-based Multi-View Contrastive Clustering for Univariate Time Series
In many realistic scenarios, large volumes of time series data are generated with limited or expensive annotations. This limitation makes s…
Bayesian Optimization of Genetic Algorithm Hyperparameters in a Multi-Fidelity Framework for Efficient Lattice Material Design
This study presents a multi-fidelity framework for the systematic optimization of genetic algorithm (GA) hyperparameters. The framework int…
CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training
Accurately estimating urban carbon emissions is critical for sustainable urban planning, yet many existing approaches remain difficult to a…
Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders
Turn-taking prediction is a key requirement for social robots involved in human-human interaction, particularly in mediator settings, where…
POO-LPSP: Parallel Osprey Optimized Least Penalty-Squared Prioritization Methods for Priority Derivation in the Analytic Hierarchy Process
Pairwise comparison (PC) via pairwise reciprocal matrices (PRMs) is central to the Analytic Hierarchy Process (AHP). Although the tradition…
FedCVESA: Taking Away Training Data in Federated Learning via Correlation Value Encoding and Segmented Aggregation
Federated learning (FL) avoids explicit data exposure by keeping raw data on local clients, yet privacy risks remain in the training proces…
HAJJv2-CrowdCount: Zero-Shot Benchmark for Dense Crowd Counting
Automated crowd counting in Hajj video is difficult not because current models lack capacity, but because the footage violates the assumpti…
Hypergraph Neural Stochastic Diffusion: An SDE Framework for Uncertainty Estimation
Hypergraph neural networks have shown powerful capability in modeling higher-order relations, yet their predictive uncertainty remains unde…
Quantum simulation of real-world nonlinear dynamics via Koopman method
Nonlinear dynamics is ubiquitous in nature, ranging from chemical pattern formation to ocean circulation, yet its simulation on quantum com…
Latency-Aware Bid Acceptance under Operational Feasibility: A Public Benchmark with Hindsight Ceilings
Online truckload bid acceptance is a closed-loop stochastic decision problem in which a carrier or broker must, in real time, accept or rej…
HumAIN: Human-Aware Implicit Social Robot Navigation
Effective social robot navigation requires sensitivity to human behavior, often revealed through subtle skeletal cues like gait and orienta…
Multi-Agent AI Control: Distributed Attacks Hamper Per-Instance Monitors
AI control is a family of techniques to prevent an AI with malicious goals from subverting its operator's intent. AI Control usually studie…
Behavior Foundations for Quadruped Robots: ABot-C0 Technical Report
In embodied intelligence systems, the motion controller serves as the critical bridge between semantic reasoning and physical execution. Hu…
On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces
Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robu…
When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs
Reliable confidence estimation remains a key limitation of test-time adaptation in vision-language models (VLMs), where prompt tuning impro…
Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation
While whole-body multimodal medical imaging scanners have been increasingly recognized for more effective medical applications, the excessi…
RLVP: Penalize the Path, Reward the Outcome
Agents acting on our behalf in the real world (e.g. placing phone calls) must learn online from costly, often irreversible interactions rat…
SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation
Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of produ…
Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Pr…
Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents
Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We ar…
Reward-Adaptive Iterative Discovery: A Case Study on Automated Game Testing for NHL26
Testing is a major effort for the gaming industry, requiring a significant part of development budget and people power. We present a case s…
TimEE: End-to-end Time Series Classification via In-Context Learning
Time series classification (TSC) is dominated by a two-stage paradigm: train a feature encoder -- either from scratch on the target dataset…
HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models
Hallucinations in vision language models (VLMs) are commonly treated as semantic errors, yet they often arise from partial or ambiguous vis…
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LL…
Stability of Flow Models for Graph Signals
Generating signals on graphs requires permutation-equivariant models that exhibit stability with respect to relative structural perturbatio…
Creativity from Friction: Human-AI Interaction for Exploratory Structural Design
AI agents that generate final answers based on user input often do not meet the needs of creative fields. Fields such as structural design…
Collaborative Synthetic Data Generation for Knowledge Transfer in Federated Learning
One-shot federated learning (OSFL) addresses the communication overhead of federated learning by limiting training to a single round, but d…
CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis
Safety evaluation for autonomous driving is dominated by rare, safety-critical interactions, motivating simulators that can deliberately sy…
Towards Agentic AI Governance: A Preliminary Assessment
Artificial intelligence is rapidly evolving from generative systems to agentic AI capable of autonomously planning and executing tasks. Wid…
Future Confidence Distillation in Large Language Models
Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, where downstream decisi…
QCNN with Rough Path Signature Kernels
Time series analysis plays a vital role across a wide range of scientific and engineering domains but poses substantial computational chall…
ALER-TI: Aligned Latent Embedding Retrieval for Time Series Imputation
Deep learning has significantly advanced time series imputation, yet most existing architectures primarily rely on localized temporal conte…
DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation
Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving…
Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning
Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answ…
Selective Timestep Weighting and Advantage-Based Replay for Sample-Efficient Diffusion RLHF
Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences.…
Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass
Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely b…
Co-LMLM: Continuous-Query Limited Memory Language Models
Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it…
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning
Structure-property relationships are foundational to biology, chemistry and materials science, where function, reactivity and physical resp…
Successor-Generator Planning with LLM-generated Heuristics
Heuristics are a central component of deterministic planning, particularly in domain-independent settings where general applicability is pr…
Shared Modular Recurrence in Contextual MDPs for Universal Morphology Control
A universal controller for any robot morphology would greatly improve computational and data efficiency. Steps have been made towards such…
LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?
Competitive programming problems are increasingly used to evaluate the coding capabilities of large language models (LLMs) due to their com…
AGAPI-Agents: An Open-Access Agentic AI Platform for Accelerated Materials Design on AtomGPT.org
Agentic AI systems increasingly connect large language models (LLMs) to external scientific tools, yet whether and when tool access improve…
Neutral Substrates: A Design Constraint for Shared Records Under Persistent Interpretive Disagreement
Shared accountability records are often used by parties who may never agree about causation, responsibility, or normative interpretation. F…
SycoEval-EM: Sycophancy Evaluation of Large Language Models in Simulated Clinical Encounters for Emergency Care
Large language models (LLMs) deployed in clinical decision support may acquiesce to patient requests for care that conflicts with evidence-…
Framing Instability in LLM Ethical Stance: Auditing Negation Sensitivity in Moral Dilemmas
Language models are increasingly consulted on ethically consequential questions, yet the stance a model expresses may not survive a change…
AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation
Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for me…
An Adaptive Differentially Private Federated Learning Framework
Federated learning enables collaborative model training across distributed clients while preserving data privacy. However, in practical dep…
SOMtime the World Ain$'$t Fair: Violating Fairness Using Self-Organizing Maps
Unsupervised representations are widely assumed to be neutral with respect to sensitive attributes when those attributes are withheld from…
Power and Limitations of Aggregation in Compound AI Systems
When designing compound AI systems, a common approach is to query multiple copies of the same model and aggregate the responses to produce…
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to c…
Anomaly detection in time-series via inductive biases in the latent space of conditional normalizing flows
Deep generative models for anomaly detection in multivariate time-series are typically trained by maximizing observed data likelihood. Howe…
Measuring the metacognition of AI
A robust decision-making process must take into account uncertainty, especially when the choice involves inherent risks. Because artificial…
Participatory provenance as representational auditing for AI-mediated public consultation
Artificial intelligence is increasingly deployed to synthesize large-scale public input in policy consultations and participatory processes…
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
Modern coding agents increasingly delegate specialized subtasks to subagents, which are smaller, focused agentic loops that handle narrow r…
M$^3$: Reframing Training Measures for Discretized Physical Simulations
Neural surrogate models for physical simulations are trained on discretized samples of continuous domains, where the induced empirical meas…
When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis
Financial decision-makers face more information than they can directly inspect, making context compression necessary. Yet when large langua…
A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory
Long term memory lets LLM agents act as persistent assistants, but user facts change. A useful memory system must know what is true now, wh…
Faster and Simpler Greedy Algorithm for $k$-Median and $k$-Means
Clustering problems such as $k$-means and $k$-median are staples of unsupervised learning, and many algorithmic techniques have been develo…
ContrastiveCFG: Guiding Diffusion Sampling by Contrasting Positive and Negative Concepts
As Classifier-Free Guidance (CFG) has proven effective in conditional diffusion model sampling for improved condition alignment, many appli…
The Minimal Search Space for Conditional Causal Bandits
Causal knowledge can be used to support decision-making problems. This has been recognized in the causal bandits literature, where a causal…
Silent Neuron Theory and Plasticity Preservation for Deep Reinforcement Learning in Adaptive Video Streaming
Adaptive video streaming optimizes Quality of Experience (QoE) metrics by selecting appropriate bitrates according to varying network bandw…
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
Recent large-scale Vision Language Action (VLA) models have shown superior performance in robotic manipulation tasks guided by natural lang…
Understanding Two-Layer Neural Networks with Smooth Activation Functions
This paper aims to understand the training solution, which is obtained by the back-propagation algorithm, of two-layer neural networks whos…
L-GTA: Latent Generative Modeling for Time Series Augmentation
Data augmentation is becoming increasingly important across various areas of time series analysis, including forecasting, classification, a…
A Study of Commonsense Reasoning over Visual Object Properties
Inspired by human categorization, visual reasoning about object properties, such as physical attributes and functions, involves identifying…
LHM-Humanoid: Long-Horizon Human Motion Control for Continuous Object Transport in Cluttered Scenes
Physics-based human motion control can make a simulated character walk, sit, and manipulate objects with high physical realism. Almost alwa…
Explain Before You Answer: A Survey on Compositional Visual Reasoning
Compositional visual reasoning has emerged as a key research frontier in multimodal AI, aiming to endow machines with the human-like abilit…
NonTextual Target Attack
Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output w…
Rapidly Learning Soft Robot Control via Implicit Time-Stepping
With the explosive growth of rigid-body simulators, policy learning in simulation has become the de facto standard for most rigid morpholog…
Refine Thought: A Test-Time Inference Method for Embedding Model Reasoning
We propose RT (Refine Thought), a method that can enhance the semantic reasoning ability of text embedding models. The method obtains the f…
T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training. Rece…
Thinking Ahead: Foresight Intelligence in MLLMs and World Model
In this work, we define Foresight Intelligence as the capability to anticipate and interpret future events-an ability essential for applica…
FDRMFL: Multimodal Federated Feature Extraction Model Based on Information Maximization and Contrastive Learning
We propose FDRMFL, a task-driven multimodal feature extraction framework for federated regression under non-IID data distributions. Extract…
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
Generalist vision--language--action (VLA) policies are typically trained on heterogeneous mixtures of robot demonstrations spanning diverse…
HiDVFS: Hierarchical Multi-Agent DVFS for Real-Time OpenMP DAG Workloads
Leakage power in multicore embedded systems now rivals dynamic power, so DVFS schedulers must respect deadlines and thermal limits, not jus…
ButterflyMoE: Compression-Scalable Ternary Experts via Structured Butterfly Orbits
In current Mixture of Experts (MoE) architectures, linear memory scaling is present, the memory grows as the number of experts increases. $…
Spatiotemporal Semantic V2X Framework for Cooperative Collision Prediction
Intelligent Transportation Systems (ITS) demand real-time collision prediction to ensure road safety and reduce accident severity. Conventi…
Can We Really Learn One Representation to Optimize All Rewards?
As unsupervised pretraining becomes increasingly ubiquitous in reinforcement learning, a more thorough theoretical understanding of these m…
Named-Entity Recognition in the Crime Domain (CrimeNER): Case Study and Dataset
The extraction of critical information from crime-related documents is a crucial task for law enforcement agencies. The extraction of this…
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make…
CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation
Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that…
Effective Strategies for Asynchronous Software Engineering Agents
AI agents have become increasingly capable at isolated software engineering (SWE) tasks such as resolving issues on Github. Yet long-horizo…
Object Search in Partially-Known Environments via LLM-informed Model-based Planning and Prompt Selection
We present a novel LLM-informed model-based planning framework, and a novel prompt selection method, for object search in partially-known e…
From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics
Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition…
Exploration of Fast-Slow Latent Recurrence for Train-Short, Test-Long Generalization
We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longe…
Diversity Without Fidelity: A Solver-Sampler Mismatch in Multi-Agent LLM Negotiation Simulation
Language models are increasingly used to simulate people: survey respondents, negotiators, stakeholders in policy exercises. In that role a…
AnyPoC: Universal Proof-of-Concept Test Generation for Scalable LLM-Based Bug Detection
While recent LLM-based agents can identify many candidate bugs in source code, their reports remain static hypotheses that require manual v…
Health System Scale Semantic Search Across Unstructured Clinical Notes
Introduction: Semantic search, which retrieves documents based on conceptual similarity rather than keywords, offers advantages for retriev…
From Beats to Breaches:How Offensive AI Infers Sensitive User Information from Playlists
The pervasive integration of AI has enabled Offensive AI: the exploitation of AI for malicious ends across the cyber-kill chain. A critical…
Optimal FALQON for Quantum Approximate Optimization via Layer-wise Parameter Tuning
Feedback-based adaptive quantum optimization (FALQON) is a promising approach for solving combinatorial problems on noisy intermediate-scal…
Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval
Current LLM memory benchmarks evaluate answer quality rather than retrieval accuracy. Consequently, a system that dumps its entire belief s…
Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning
Repeated sampling with a verifier is the standard way to allocate test-time compute for code generation, with pass@$K$ as the canonical met…
Informing AI Policy Assessment using Large-Scale Simulation of Interventions
As the rapid proliferation of AI systems and harms spurs efforts in AI governance around the world, prioritizing among competing policy opt…
Trading Human Curation for Synthetic Augmentation in RLVR
The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic lang…
Trust, but Don't Verify: Epistemic Blind Spots in LLM Source Evaluation
Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions. Whether they evalua…
TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation
TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequ…
MetaConfigurator: AI-Assisted RDF Authoring from JSON Data
Scientific workflows increasingly generate structured JSON data that is easy to exchange but difficult to interpret consistently across sys…
The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions
The standard basis of transformer hidden states is a training-free, architecture-general feature basis for detecting concepts and, in langu…
Feynman Kac Reweighted Schr\"odinger Bridge Matching for Surface-Based Tau PET Harmonization
Tau positron emission tomography (PET) is widely used for the in vivo characterization of disease stage and progression in Alzheimer's dise…
NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization
Reliable sound source localization is fundamental to robot audition, enabling autonomous robots to perceive spatial cues and operate effect…
Where Did the Variability Go? From Vibe Coding to Product Lines by Regeneration
In vibe coding, an emerging AI-driven paradigm, an LLM generates an entire program from a natural language prompt, but what happens to the…
Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking
The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion es…
RWGBench: Evaluating Scholarly Positioning in Related Work Generation
Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.…
Can Trustless Agents Be Trusted? An Empirical Study of the ERC-8004 Decentralized AI Agent Ecosystem
As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent ass…
JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators
Text-to-image (T2I) diffusion models typically require substantial computational resources and cloud infrastructure, posing significant cha…
「万年3位」から脱却なるか? Google Cloudに吹く"追い風"の正体を考察
エンタープライズ市場でAWSやMicrosoftに水をあけられてきたGoogle Cloudに、追い風が吹いている。国内SIer大手4社が同社との協業に動く理由と、これまでの経緯、AIエージェントでの連携戦略を読み解く。
Anthropic、「Claude Code」のシステムプロンプトを80%削減 「モデルの創造性を解放するため」
米Anthropicのタリク氏は、「Claude Fable 5」など最近リリースされたAIモデルへの禁止指示がモデルの創造性を制限すると指摘した。
Lovable reportedly in talks to double its valuation to $13.2B
The $300 million round is expected to be led by Menlo Ventures, Sifted reported.
CEOの利用額も「全社員に丸見え」 LayerXがAI予算を「第二の人件費」にした真意
「社長のAI利用額まで全社に生中継する」という透明性でコスト管理に挑むLayerX。利用額が予算の10倍を超過しても経営陣が現場を叱らなかった背景には、AI予算を「第二の人件費」として組織の資産に変える計算があった。単なる経費の締め付けを排し、外注費をAI費に切り替える基準から…
現場に聞いた「IT関連50製品」のぶっちゃけ理解度 “浸透するツール”の共通点とは?
似た機能を持つITツールが増える中、「なぜ定着する製品とそうでない製品が生まれるのか」。IT関連の主要50製品を対象に、利用イメージや認知度の違いから、企業が選ぶツールの条件を探った。
OpenAI、リアルタイム音声モデル「GPT-Live」公開 相づちも割り込みも、より人間らしい会話に
OpenAIは、新世代のリアルタイム音声モデル「GPT-Live」を発表した。会話の最中に相手の話を聞きながら同時に話せる全二重方式を採用し、ChatGPTの音声機能を刷新する。深い推論や検索が必要な場合はバックグラウンドで「GPT-5.5」に処理を委ねる仕組みを備え、有料・無…
フィジカルAI搭載ロボットがモノポリーを実演、1台のPCにモーション制御も統合
モベンシスは、「第38回 ものづくりワールド[東京]」の構成展である「第1回 フィジカルAI展[東京]」において、PCでリアルタイム制御を実現するソフトモーションコントローラー「WMX3」のROS 2向けパッケージ「WMX for ROS 2」を紹介した。
SDV時代支えるHEREの位置情報プラットフォーム、新たな柱は「二輪」へ
SDVによる変革が進むモビリティ社会において、「位置情報」は人々の移動にどのような価値をもたらすのか。デジタル地図からグローバルなロケーションのプラットフォーマーへと進化を遂げたHERE Technologiesの日本法人トップを務める枝隆志氏に話を聞いた。
Google’s deepfake detector system used to debunk McConnell hoax pic
Earlier this week, a picture seemed to show Kentucky Senator Mitch McConnell covered in tubes in a hospital bed in a state of extreme distr…
SpaceXAI releases Grok 4.5, which Elon describes as an ‘Opus-class model’
Elon Musk's tech company released the newest version of Grok on Wednesday, promising a cheaper, more efficient alternative to other powerfu…
This startup thinks robotics is about to have its ChatGPT moment
General Intuition is betting millions of hours of video game data can train the foundation models for physical AI, making it easier to buil…
Google Photos adds a new AI ‘Video Remix’ tool
The feature can do things like apply cinematic relighting to brighten up a dark clip, swap out a plain background for something fun, or add…
Why this CEO thinks video games make better training data than the internet
When it comes to achieving artificial general intelligence (AGI), large language models just don’t have what it takes. Models like ChatGPT…
Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
Meta is adding a new safeguard to stop people from secretly recording others with its AI glasses. But the update comes as the company conti…
OpenAI releases new voice models for more natural live conversations
OpenAI says its new voice mode can speak and listen at the same time, a key ability for live translation.
Prime Intellect raises $130M Series A to help enterprises build their own AI agents
Founded in 2024, Prime Intellect’s goal is to give organizations capabilities to train their own agentic systems without relying on frontie…
These AI startups are growing revenue at faster and faster rates
There are a lot of fast-growing AI startups, but some are growing even faster, they say.
2026-07-08(287件)
Our approach to government and national security partnerships
Learn how OpenAI approaches government and national security partnerships, with principles for responsible AI use, democratic accountabilit…
Your gaming data could be the secret to AGI, according to this Bezos-backed startup
When it comes to achieving artificial general intelligence (AGI), large language models just don’t have what it takes. Models like ChatGPT…
Separating signal from noise in coding evaluations
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in…
Former OpenAI exec Kevin Weil is now on the board of Stoke Space
Kevin Weil's new role at Stoke Space suggests reusable rockets are the next hot thing in Silicon Valley.
Helping K–12 educators build practical AI skills
OpenAI Academy and the Walton Family Foundation are bringing hands-on AI Skills Jams to help K–12 educators build practical AI skills for t…
Claude「サブスク最上位プラン」6カ月間無料で提供 OSS開発者向けキャンペーン、対象を拡大
米Anthropicは、オープンソースソフトウェア(OSS)開発者向けのキャンペーン「Claude for Open Source」の対象範囲を拡大したと発表した。認定した対象者には、最上位のサブスクリプションプラン「Claude Max 20x」を6カ月間無料で提供する。
Hot French startup ZML releases free product to speed inference across lots of AI chips
ZML, a hot French AI startup endorsed by Turing Award winner Yann LeCun, has now released ZML/LLMD, software that could make running AI les…
「Grok 4.5」も明日一般公開へ マスク氏「Opus級だが、より高速で低コスト」
「Opusクラスのモデルだが、より高速で、トークン効率が高く、低コスト」
AI chip maker SambaNova raises $1B at $11B valuation, 5 months after last mega round
AI chip maker SambaNova has raised at an $11 billion valuation months after Intel was rumored to be trying to buy it for about $1.6 billion.
「GPT-5.6」木曜に一般公開へ 米政府と調整→限定プレビュー経て
日本時間では9日午後以降または10日になりそうだ。
OpenAIの最新AI「GPT-5.6」シリーズ、今週木曜日に一般公開へ
米OpenAIは7月7日(現地時間、以下同)、最新のAIモデル「GPT-5.6」シリーズを「今週の木曜日に一般公開する」と発表した。日本時間では10日(金)になるとみられる。
Prompt-to-Paper: Agentic AI System for Bioinformatics
While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three…
From Graphs to Gradients: Physics-Inspired Structural Attribution for Cyber-Physical IoT Systems and Beyond
Interpretable explanation methods in Artificial Intelligence aim to uncover the underlying causes and their effects, enabling a deeper unde…
CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
Large language models are increasingly explored as AI tutors, yet deploying them in K-12 settings raises concerns around privacy, cost, and…
Foundation Models for Automatic CAD Generation
Recent advances in Large Language Models (LLMs) and Vision-Language Models (VLMs) enable the automatic generation of parametric 3D designs…
Narrative World Model: Narratology-Grounded Writer Memory for Long-Form Fiction
Long-form fiction writers need memory that answers multi-hop questions about evolving story state: who knows a secret and when they learned…
FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents
LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, b…
Memory in the Loop: In-Process Retrieval as ExtendedWorking Memory for Language Agents
Language agents run a loop - observe, reason, act - but the memory they reason over sits outside it: a store queried at most once per turn.…
Akashic: A Low-Overhead LLM Inference Service with MemAttention
Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflow…
ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation
Computer-aided design (CAD) for industrial components requires long-horizon procedural modeling, robust feature dependencies, editable para…
Synthetic Consumer Insight Generation with Large Language Models
Modern data-driven marketing relies on large amounts of consumer data, yet collecting such data can be costly, time-consuming, and difficul…
Beyond Static Evaluation: Building Simulation Environments for Scalable Agentic Reinforcement Learning
As Large Language Models (LLMs) evolve into autonomous agents, traditional static evaluation fails to capture multi-step decision-making. W…
Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents
Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other ag…
Controlling Tool Use with Heading-Specific Activation Steering
Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them…
From Passive Retrieval to Active Memory Navigation: Learning to Use Memory as a Structured Action Space
Long-term user memory is essential for personalized conversational agents, yet many memory systems still expose memory through passive retr…
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising…
Onnes: A Physics-Grounded Multi-Agent LLM Simulator for Cryogenic Fault Diagnosis in Quantum Computing Infrastructure
Dilution refrigerators are the enabling infrastructure of superconducting quantum computers, yet their fault diagnosis is still dominated b…
StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems
Agent systems accumulate conflicting observations across branches, retries, and replicas, yet many practical memory layers still collapse d…
Uncovering Latent Depression Severity for Binary Depression Detection via Advantage-weighting Ranking
Automatic depression detection using audio-visual data faces significant challenges, particularly in disentangling overlapping feature dist…
PCBWorld: A Benchmark Environment for Engine-Grounded PCB Design Automation
PCB routing is the task of connecting the nets of a board with copper traces under strict design rules, yet learning-based methods still la…
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation
Training multimodal search agents to perform multi-hop reasoning remains challenging due to a fundamental structural disconnect: existing p…
Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH
The integration of Large Language Models (LLMs) into scientific research workflows, particularly for bibliographic discovery and literature…
Auto-DSM Under the Lens: A Black-Box Evaluation Framework for LLM-Based DSM Generation
This paper presents a black-box evaluation framework to systematically assess the ability of Large Language Models (LLMs) to generate Desig…
AgoraSim: A Hybrid Agent-Based Modeling Framework
LLM-agent simulations make natural-language social scenarios easy to instantiate, but their outputs can be overread as predictions and are…
Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test
We report a pre-registered, two-part experiment on small economies of frontier language-model agents (Claude Opus 4.8), testing two quantit…
PolyWorkBench: Benchmarking Multilingual Long-Horizon LLM Agents
Large language model (LLM) agents have shown strong performance in long-horizon tasks that require planning, tool use, and interaction with…
Reward-Density Heuristic for Dynamic Multi-Vehicle Routing: Performance and Computational Efficiency
The Vehicle Routing Problem (VRP) and its variants represent some of the most practically consequential optimization challenges in modern l…
When do prophets profit in prediction markets?
Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of uncertain events. Classical theory establ…
A toy framework for single and multi-agent human-AI curiosity ecosystems
This paper offers a toy framework for considering curiosity as an ecosystem. First, it suggests that a single agent's inquiry policy (how,…
Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Reinforcement learning has become a promising paradigm for improving large language model (LLM) agents on long-horizon search tasks, where…
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle t…
From Application-Layer Simulation to Native Meta-Architecture: Structural Tension as an Endogenous Driver for Heterogeneous AI Evolution
Current large language models (LLMs) are fundamentally stateless: their behavior is fully determined by input at inference time, and any hi…
Task Decomposition-Guided Reranking for Adaptive Agent Skill Retrieval
Skill usage can significantly enhance the ability of modern agent systems to complete complex tasks. However, the growing scale of skill li…
DT-Guard: Intent-Driven Reasoning-Active Training for Reasoning-Free LLM Safety Guardrail
Large language models deployed in open-world applications require safety guardrails that are both robust to complex risks and efficient eno…
Driving the Wrong Way: Leveraging Interpretability in End2End Autonomous Driving Models
The increasing adoption of end-to-end learning for autonomous driving introduces increased model complexity and opacity, raising the risk o…
TopoBrick: Agentic Topology Sampling of Exogenous Variables for Zero-Shot Building IoT Forecasting
Building sensors are embedded in physical topology, spatial hierarchy, and operational context, yet existing forecasters often treat them a…
A Definition and Roadmap for World Models
World models -- internal simulators that learn the structure and dynamics of an environment -- have become one of the most actively debated…
ExplAIner: A Declarative Query Language for Explaining Classification Models
The XAI community has studied a wide range of queries and scores for explaining predictions of ML models. From a data management perspectiv…
Finding H. pylori in the Fine Print: Evidence-Linked Multi-Agent Case Finding from Gastric Biopsy Reports
Data from Singapore indicated that about 31% of the population had evidence of Helicobacter pylori infection. Persistent H. pylori infectio…
Danus: Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory
Recent LLM-based mathematical reasoning agents have begun to tackle research-level problems and, in several cases, have contributed to the…
A Physics-Informed Neural Network Framework for Elastodynamic Wave Propagation in Bimaterial Systems
Physics-informed neural networks (PINNs) provide a promising framework for solving partial differential equations while embedding the under…
Multi-Agent Deep Reinforcement Learning for Multi Objective Battery Management in Dairy Farms
The dairy industry in Ireland has a large potential for the integration of renewable energy and the reduction of carbon emissions. However,…
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consu…
RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models
Recent years have witnessed the emergence of multivariate modeling using time series foundation models (TSFMs), which achieve advanced zero…
FootsiesGym: A Fighting Game Benchmark for Two-Player Zero-Sum Imperfect-Information Games
We present FootsiesGym, an open-source environment for learning in a non-trivial two-player, zero-sum, imperfect-information game. Built on…
FreqDepthKV: Frequency-Guided Depth Sharing for Robust KV Cache Compression in Long-Context LLM Inference
Long-context LLM inference is increasingly limited by the memory and bandwidth cost of KV caches, yet aggressive compression can remove the…
Bridging Physical Reasoning and Task Generalization via Visual Action Outcome Reasoning Alignment
Vision-language models (VLMs) struggle to generalize in interactive physical reasoning, particularly under unseen tasks and environments. T…
DepthWeave-KV: Token-Adaptive Cross-Layer Residual Factorization for Long-Context KV Cache Compression
Long-context language model inference is increasingly limited by the memory bandwidth and capacity required to store key-value caches, yet…
The Large Cancer Assistant (LCA): A Model-Agnostic Orchestration Framework for Scalable Clinical Decision Support in Oncology
- Objective: Multimodal deep learning models in oncology are currently limited by monolithic designs that rigidly couple data ingestion, cl…
Rethinking Indic AI from a Lens of Cultural Heritage Preservation
As Artificial Intelligence (AI) makes inroads into different parts of the Indian subcontinent, there is significant interest in studying ho…
Lingering Authority: Revocable Resource-and-Effect Capabilities for Coding Agents
Coding agents often receive broad tool access for an entire task, even when a resource is needed only for one subgoal. We call this gap lin…
KVpop -- Key-Value Cache Compression with Predictive Online Pruning
Key-value (KV) cache growth is a major bottleneck in autoregressive decoding, as memory and bandwidth scale linearly with context length. E…
Proof of Execution: Runtime Verification for Governed AI Agent Actions
Agent systems increasingly execute rather than advise. When an AI agent queries regulated data, invokes effectful tools, and mutates persis…
Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving
Large language model serving is increasingly limited by KV-cache growth under long-context workloads, yet existing KV-cache compression tec…
Catalyst Papers in Artificial Intelligence Research: A Landscape on ICLR from 2017 to 2025
A small number of methodological contributions, including word2vec, the Transformer, large-scale pre-training, and reinforcement learning f…
AI tools in Arab University English classrooms: Looking back and forward
This paper aims to synthesize empirical research on AI tools used to support English as a second/foreign language (EL2) learners in Arab Un…
The Jagged Global Economy: Frontier AI Unevenly Exposes National Economies
Frontier AI's labor-market effects matter to workers, firms, and policymakers, but current evidence generally comes from a handful of high-…
CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries
To interact with users fairly and without stereotyping, AI models must display cultural competency, i.e., the ability to infer and adapt to…
A Guiding Framework for K-12 Teachers in Creating AI-powered Learning Technologies through Vibe Coding
Large language models generate code from natural language prompts, enabling "vibe coding," which allows non-programmers to develop computat…
Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety
Modern artificial intelligence (AI) systems present profound new risks to child safety. AI is increasingly being misused to create AI-gener…
Automated Recommendation of Programming Learning Content Using Pattern-based Knowledge Components
Introductory programming instruction relies on hands-on practice and short learning activities to support mastery of foundational concepts.…
CANONIC: Governance Is Compilation
We present CANONIC: governed intelligence that compiles digital artifacts into an evidence ledger at scale. Large language models generate…
The GenAI Skill Bypass: Mapping Divergent Pathways of University Students and Staff AI Literacy
Higher education institutions are increasingly expected to ensure that both students and staff develop Generative AI (GenAI) literacies. In…
Why does AI unlock new possibilities in STEM education? A Bibliometric Analysis of Trends and Future Agenda
STEM education faces challenges in personalization and interdisciplinary integration. AI technology has brought new possibilities, but the…
Contrastive Predictive Coding with Compression for Enhanced Channel State Feedback in Wireless Networks
Accurate and timely channel state information (CSI) is essential for next-generation wireless systems, yet existing works treat CSI compres…
When AI Classifies: What Counts as Public Administration?
This study examines how alternative systems of scholarly representation identify and characterize broad public administration (PA) and arti…
CHARLIE: An On-Premise Multi-Agent Retrieval-Augmented Generation System for Evidential Reasoning in Forensic Science
We present Charlie, an on-premise multi-agent Retrieval-Augmented Generation (RAG) system for structured evidential processing in digital f…
Modality Relevance is not Modality Utility: Post-hoc Selective Modality Escalation for Cost-Aware Multimodal RAG
Multimodal retrieval-augmented generation (RAG) grounds a generator in evidence drawn from heterogeneous modalities -- text, tables, and im…
PORTS: Preference-Optimized Retrievers for Tool Selection with Large Language Models
Integrating external tools with Large Language Models (LLMs) has emerged as a promising paradigm for accomplishing complex tasks. Since LLM…
Scientific Code Search at Scale: A Multi-Domain Dataset and Benchmark
Scientists increasingly rely on open-source tools to support their research workflows, yet discovering relevant software among over 600 mil…
Geometry-Aware Infrastructure-Anchored Denoiser for UWB Sensing and Work-Zone Reconstruction
Accurate work-zone geometry perception is critical for intelligent transportation systems, and ultra-wideband sensing offers a low-cost app…
The Granularity Paradox: How Temporal Disaggregation Inflates In-Sample Fit and Compounds Out-of-Sample Error
This paper explores the "Granularity Paradox" in time-series forecasting, wherein finer temporal disaggregation (e.g., Monthly to Weekly/Da…
Empirical Minimal-Realisation Compression of Deep Neural Networks via Controllability-Observability Tests
Deep neural networks often contain substantial hidden-state redundancy, but most compression methods operate directly on weights, neurons,…
Learning to Control LLM Agent Harnesses with Offline Reinforcement Learning
Large language model (LLM) agents are usually improved by changing prompts, models, or hand-written workflows, while the execution harness…
AdaStop: Cost-Aware Early Stopping for DNN Test Selection
Existing methods for testing deep neural networks (DNNs) primarily prioritize test inputs likely to reveal model faults under a fixed label…
Evaluating calibrated refusal and safe usefulness in dual-use biology settings
As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSe…
Learnable Weighting of Intra-Attribute Distances for Categorical Data Clustering with Nominal and Ordinal Attributes
The success of categorical data clustering generally much relies on the distance metric that measures the dissimilarity degree between two…
CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration
Complex image creation and editing often require more than a single generation or editing model. A user request may involve synthesizing im…
Learning 4D Geometric Priors for Inference-Efficient World Action Models
World Action Models (WAMs) have shown strong potential for robotic manipulation by jointly modeling visual future dynamics and executable a…
Breaking Structural Isolation: Scalable Graph Clustering via Community-Aware Sampling and Structural Entropy
Unsupervised graph clustering is a fundamental technique for uncovering underlying semantic patterns in large-scale networks. Although Grap…
KAT-Coder-V2.5 Technical Report
We present KAT-Coder-V2.5, a coding-focused agentic model trained to act autonomously inside real, executable repositories rather than as a…
Binocular Gaze Estimation with Single Camera and Single Light Source
According to commonly consented theories, the minimum hardware requirement for gaze tracker is one camera and two light sources to realize…
Is Your NPU Ready for LLMs? Dissecting the Hidden Efficiency Bottlenecks in Mobile LLM Inference
Deploying Large Language Models (LLMs) on mobile devices enhances privacy and reduces latency, but is severely bottlenecked by hardware ine…
Decision Protocols in Multi-Agent Large Language Model Conversations
Improving the task performance of Large Language Models (LLMs) is essential, yet scaling these models faces significant challenges such as…
Privilege and confidentiality in generative AI workflows
Generative AI (GenAI) systems store and process client data in three distinct ways: in the model's parameters through training and memorisa…
Full-range Binary Classifier Calibration for Stable Model Updates in Production
Detection models running in adversarial environments face a malicious distribution that drifts rapidly while the benign distribution stays…
PatchOptic for Shared-State LLM Workflows with Projected Views and Verified Structured Updates
Agentic workflows often operate over shared, structured state. Because LLM context windows are limited, each model invocation is typically…
Lean-Quantum: Toward AI-Assisted Formalization of Quantum Information
Quantum information theory is built on entropic quantities; among them, the sandwiched R\'enyi relative entropy is a fundamental divergence…
Statistical Adversaries: Natural Backdoor-like Features in Vision Datasets
Model-specific adversarial attacks have been extensively studied. We study a different failure mode: naturally occurring statistical signal…
aiAuthZ: Off-Host, Identity-Bound Authorization for AI Agents
AI agents issue tool calls on the basis of text they cannot verify, so any party who controls part of the context can forge the appearance…
Rendering-Aware Bayesian 3D Gaussian Splatting with Native Uncertainty and Adaptive Complexity Control
3D Gaussian splatting (3DGS) is a strong representation for real-time novel-view synthesis, but its standard training pipeline relies on po…
Self-Review Reinforcement Learning (SRRL) with Cross-Episode Memory and Policy Distillation
Reinforcement Learning is commonly used to train large language models using environmental feedback. In applied settings, the environment u…
Most LLM Conformity Needs No Speaker: Measuring the Speaker-Free Floor in Peer-Pressure Benchmarks
LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response. We show that most of…
The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment
Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting…
Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation
Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumpt…
Harnessing Generative Image Models for Training-Free Primitive Shape Abstraction
Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding. Generative…
Whose fairness? Structural concentration in AI bias research
Artificial intelligence increasingly mediates consequential decisions in healthcare, law, and public services, and the field has responded…
ResonatorLM: Causal Resonant Field Mixing for Efficient Long-Context Language Modelin
Contemporary language models are dominated by the transformer architecture, which leverages self-attention mechanisms to enable more effici…
Hierarchical Classification via Cascading Feature Elimination: Application to Human Phenotype Ontology-Aligned Facial Phenotyping (FaceMesh2HPO)
FaceMesh2HPO is a framework for classifying facial phenotypic descriptors aligned with the Human Phenotype Ontology (HPO) to support clinic…
To Retain or to Adapt? Generalizing Continual Learning
The Continual Learning (CL) literature has long been driven by the goal of mitigating catastrophic forgetting. This objective rests on a pe…
BaFCo: A Document Understanding Benchmark for Complex Bangla Form Comprehension
Document comprehension is a challenging yet impactful task for Multimodal Large Language Models, especially as these systems see growing ad…
Safe Bayesian Optimization with Counterfactual Policies
In many decision-making settings, new interventions are acceptable only if they do not reduce outcomes below some established threshold. Fo…
EvalLoop: A Methodology for Evaluation-Driven Iterative Improvement of Business AI Systems
Teams deploying large language models in business contexts need evaluation systems, yet most treat evaluation as static model selection: ru…
Do It Right! A Methodology for Successful NLP System Development
Natural language processing (NLP) is a common method for supplying data to clinical research and decision making by extracting information…
Physics-Regularized Machine Learning for Proprioceptive Vehicle Localization Using Onboard Sensors
Accurate and robust localization is essential for autonomous mobility systems in real-world environments. While fusing Inertial Measurement…
What Do AI Agents Actually Change? An Empirical Taxonomy of Mutation Patterns in Performance-Improving Pull Requests
AI coding agents are black boxes: we cannot inspect how they generate code, but we can inspect what they change. This distinction matters f…
RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs
Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate a…
Beyond Accuracy: How Humans Evaluate Legally Correct but Socially Controversial Legal Advice from Machines
AI systems are increasingly used to provide legal advice, raising questions about whether laypeople accept guidance from algorithms--especi…
Depression Symptoms and Relational Patterns in 187k ChatGPT Histories
Large language models are increasingly used as private, always-available conversational systems, but little is known about how people with…
IMR: Iterative Mode-World Weighted Regression for Multi-Agent Trajectory Prediction
Multi-agent motion prediction is essential for automated vehicles to understand the intentions of surrounding vehicles. However, previous p…
Plainbook: Data Science, in Plain Language
Jupyter Notebooks have become widely adopted in data science, as they allow the sharing of reproducible computational analysis. They are, h…
SCOReD: Student-Aware CoT Optimization for Recommendation Distillation
Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-su…
The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities
AI coding agents now read repositories, call tools, and execute shell commands with limited human oversight, and a fast-growing body of wor…
Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations
The Model Context Protocol (MCP) is the dominant way coding agents discover and invoke external tools. A server advertises each tool throug…
When Should LLMs Search? Counterfactual Supervision for Search Routing
Search-augmented language models can use external evidence to compensate for limitations in parametric knowledge, but search is not uniform…
Data-dependent Evaluations for Budgeted Submodular Maximization
Submodular maximization is an important building block for developing algorithms in many areas such as machine learning and data mining. Du…
LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding
We propose a novel pipeline, Legato 2, for extracting symbolic notation and semantic knowledge from images of sheet music. Legato 2 feature…
FORGE: Towards Functional Tool-Use Generalization via Keypoint Trajectory Reasoning
While humans readily repurpose a book, a stone, or a shoe to drive a nail, robots trained on specific tools fail to transfer the same funct…
Segmentation before Answering: Pixel Grounding for MLLM Visual Reasoning
Recent advancements in Multimodal Large Language Models (MLLMs) have evolved from static perception to interleaved visual-language reasonin…
Complementary Roles of Image Classification and Vessel Segmentation in AI-Based Screening for Retinopathy of Prematurity Plus Disease in a Kenyan Preterm Cohort
Background. Retinopathy of prematurity (ROP) is a preventable cause of childhood blindness, with rising burden in low- and middle-income co…
Decision-Focused Scenario Generation and Selection for Efficient and Robust Grid Dispatch
The increasing uncertainty from flexible demand and renewable generation has made distributionally robust optimization (DRO) an important t…
Tangent classes of matroids and wonderful compactifications
For every loopless matroid $M$ and every Feichtner--Yuzvinsky building set $\mathcal{G}$ containing the top flat, we construct an integral…
VisTCP: A Visualization Framework to Construct Knowledge-Graph-Based Representation for Traditional Chinese Painting
Structured representation can characterize semantic objects and relationships in images. It provides a possible effective way for the seman…
Beyond Refusal: A Same-Lineage Study of Aligned and Abliterated LLMs for Vulnerability Analysis
Large language model (LLM)-assisted software security operates at a difficult boundary: the vulnerability-analysis terminology needed for l…
AbICL: In-Context Learning for Antigen-Specific Antibody Affinity Ranking
Accurate ranking of antibody candidates according to their binding affinity is essential for therapeutic antibody discovery. However, exist…
Unsupervised Anomaly Detection of Information Operations Users via Behavioral and Language Patterns
Information Operations on social media networks have been identified as a significant threat to democracy and modern society, but they are…
Differentially Private Natural Gradient Descent
Under a fixed privacy budget, the utility of differentially private (DP) training is ultimately determined by its optimization efficiency.…
Think Before You Grid-Search: Floor-First Triage for LLM Serving
LLM serving optimization typically benchmarks many configurations and reaches for heavy profilers when latency targets are missed. We argue…
Harrison.Rad 1.5 Technical Report: A radiology foundation model that can draft reports from images, priors and clinical context
Imaging demand is growing faster than the radiology workforce can expand, and reporting backlogs cannot be resolved through training and re…
i-EXAM: Instructable and Explainable Attack Connectivity Graph Modeler
i-EXAM is a planning-powered tool that helps system administrators to create security profiles of complex networks and perform what-if anal…
Few-Medoids: An Embarrassingly Simple Coreset Selection Method for Few-Shot Knowledge Distillation
Coreset selection aims to identify a small and highly representative subset of a massive dataset for efficient model training. The problem…
From Textural Counterpoint to Feature Encoding: A Multi-Dimensional Machine Representation Study of Haydn's "The Lark" Integrating Electroacoustic Analysis
Chamber music, as a highly precise multi-part interactive system, contains a logic of "role assignment and dynamic interaction" that provid…
K-ABENA: K-Adaptive Backpropagation with Error-based N-exclusion Algorithm : (Compensated Loss-Based Sample Exclusion with Unbiased Gradient Estimation)
We present K-ABENA (K-Adaptive Backpropagation with Error-based N-exclusion Algorithm), a selective gradient computation framework that red…
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an…
CMDR: Contextual Multimodal Document Retrieval
Multimodal document retrieval aims to retrieve relevant pages while preserving both textual and visual content from the original document.…
Signed-Graph Recommendation as Structural Consistency Maximization
While signed social recommendation has shown great potential by modeling both trust and distrust relations, its effectiveness is often hind…
NegROI: Click-Centric Uncertainty-Guided Refinement with Scene-Conditioned Negative Prompts for Robust Interactive 3D Segmentation
Interactive 3D segmentation aims to extract object masks in point clouds with minimal user clicks. Despite recent progress, most existing a…
Agentic AI for IPoDWDM Network Lifecycle Automation: An MCP-Enabled Architecture
We present a distributed, vendor-agnostic multi-MCP architecture for SDN-based automation and autonomous control of multi-vendor, multi-lay…
Decoupled Single-Mask Annotation Noise Detection via Cross-Sectional Patch Self-Consistency
Vascular computed tomography datasets are commonly annotated only once per scan, yielding the pervasive yet under addressed problem of sing…
InfluMatch: Frontier-Quality KOL Search at 4B-Model Cost
Matching influencers (KOLs) to free-form, multi-part Thai marketing criteria is today served either by keyword search over structured profi…
Faithful or Findable? Evaluating LLM-Generated Metadata for RDF Dataset Search
Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems. W…
MCP-Enabled Agentic AI for Autonomous IPoDWDM Network Lifecycle Automation
This demo presents an MCP-enabled agentic AI architecture for autonomous control of vendor-agnostic IPoDWDM networks. We demonstrate live e…
Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention
Multimodal large language models can emit localized predictions, bounding boxes for objects and temporal windows for video and audio events…
PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages
Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks…
Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
Prompt engineering has emerged as a critical yet undertaught skill for software developers, one that traditional learning approaches are il…
From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations
Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical m…
PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet
3D dense captioning, an emerging vision-language task, aims to generate descriptive sentences for each object in the 3D scene. Despite the…
Agents That Teach: Towards Designing Incidental Learning Back into AI-Assisted Software Development
AI coding agents are rapidly reshaping how software is built, with developers increasingly delegating substantial coding tasks to autonomou…
EcoVision: AI-Powered Drone Imaging for Salt Marsh Vegetation Monitoring and Dominance Mapping
High-resolution RGB imagery acquired from low-altitude UAV surveys was processed through a modular pipeline incorporating transformer-based…
RoME: Robust Mixture of Low-Rank Experts against Multiple Adversarial Perturbations
Multi-perturbation adversarial training (MAT) aims to achieve robustness against multiple $\ell_p$ perturbations but suffers from robustnes…
LLM-Guided Measurement Credibility Correction for Trustworthy Industrial Process Inference
Industrial prediction and soft sensing depend on credible input measurements. In field deployment, a predictor may receive biased, delayed,…
x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability
Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function eva…
Static Metrics Are Insufficient: Predicting Java Method Energy Usage with Execution Time
The increasing energy demand of software systems is raising concerns about their environmental impact and associated costs. Reasoning on en…
Evaluating Fine-Tuning and Metrics for Neural Decompilation of Dart AOT Binaries
Neural decompilation is increasingly studied as a code-generation problem, yet its evaluation methodology remains underdeveloped for modern…
Self-Supervised Implicit CEST Reconstruction via Physics-Informed Lorentz Encoding
Multi-Pool Chemical Exchange Saturation Transfer (CEST) MRI provides valuable metabolic information but is clinically limited by long acqui…
Property-Driven Synthetic Data Engineering for Data-Scarce Software Systems: Reflections from the Breast Cancer Domain
Modern software systems increasingly depend on data for analysis, prediction, testing, and decision-making. Yet many important domains, inc…
LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability
Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information…
LongCrafter: Towards Diverse Long-Context Understanding via Evidence-Graph-Guided Instruction Synthesis
Synthesizing long-context supervised fine-tuning (SFT) data is a scalable way to enhance the long-context understanding of large language m…
X-FEMR: A Token-level Explainable Approach for Electronic Health Records Foundation Models using Transformer-based Models
Foundation Models for Electronic Health Records (FEMRs) are pretrained on large-scale structured patient data, enabling them to convert lon…
Improving LLM-Generated Process Model Quality Through Reinforcement Learning: The Role of Reward Function Design
Large language models (LLMs) can generate BPMN process models from natural-language descriptions, yet supervised fine-tuning (SFT) limits t…
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
There are some datasets of varying scales for audio classification (AC) applied to different tasks. However, annotated data is limited for…
UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods
The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatr…
Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows
Major cloud data platforms now expose large language model capabilities as native SQL functions, enabling analysts to perform classificatio…
VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection
Deepfake image detection is currently served by three fundamentally different paradigms: commercial APIs, zero-shot vision-language models…
Designing Maintainable Hybrid Generative Systems: A Quantum-Inspired Approach to Automated Music Harmony Generation
This paper presents the design and evaluation of a maintainable hybrid generative architecture for automated music harmony generation from…
UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on…
Token-Based Dual-view Fusion and Adaptation of Large Vision Models for Breast Cancer Classification
Accurate breast cancer classification from mammography requires effective integration of complementary information from craniocaudal (CC) a…
Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on En…
Harnessing Code Agents for Automatic Software Verification
Formal verification offers the strongest guarantee of software correctness, but it does not scale: the proofs demanded by interactive theor…
Responsible Personalisation: The Double-Edged Sword of Personalisation in Human-Robot Interaction
While personalisation is becoming a defining capability in human-robot interaction (HRI), the existing literature on responsible personalis…
What Images Cannot Say: Language-Guided Olfactory Representation Learning
Images tell us what a scene looks like, but rarely what it would feel like to be there. While recent datasets pair visual scenes with elect…
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications
Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in th…
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
Large language model coding agents increasingly perform open-ended data modeling and analysis. These agents are stochastic and adaptive, an…
TILDE: TILt-based Distributional Erasure for Concept Unlearning
Concept unlearning in text-to-image diffusion models is critical for safe and practical deployment: with rising privacy concerns, copyright…
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders
Vision-Language Models (VLMs) are increasingly utilized as the conditioning backbone for diffusion-based image editing due to their remarka…
From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b
Biomedical question answering requires not only accurate extraction of information from scientific literature but also reliable integration…
Provable learning separation for predicting time-evolution of quantum many-body systems
Given that quantum computers are naturally suited to simulate the behavior of quantum many-body systems, an immediate question arises: can…
Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation
We present PACR-Video, a parameter-efficient framework for multi-shot long video extrapolation that preserves recurring entities, scene str…
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically f…
AirflowAttack: Thermal-Airflow Adversarial Perturbations against Infrared Remote-Sensing Vision-Language Models
Vision-language models (VLMs) are increasingly deployed on infrared (IR) remote sensing imagery in security-critical settings, yet their ad…
Pitwall: Faithful Natural-Language Race-Strategy Briefings from a Calibrated Real-Time Monte Carlo Engine
Live sports commentary is grounded generation under a deadline: statements concern real, named athletes, the grounding state changes every…
Industry Classification of GitHub Repositories Using the North American Industry Classification System (NAICS)
GitHub hosts hundreds of millions of public repositories, but the platform exposes no native mapping from repositories to standardized indu…
RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation
Multi-hop Question Answering over Knowledge Graphs faces a critical challenge: traditional retrieve-then-read pipelines break differentiabi…
Graph Convolutional Attention: A Spectral Perspective on Graph Denoising and Diffusion
Denoising graphs is a fundamental problem in graph learning and the core operation of graph diffusion models. Attention-based architectures…
ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D int…
MLLM-LLaVA-FL: Multimodal Large Language Model Assisted Federated Learning
Previous studies on federated learning (FL) often encounter performance degradation due to data heterogeneity among different clients. In l…
Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?
We investigate whether socioeconomic indicators, like household wealth, leave recoverable informational imprints in both satellite imagery…
Base Models Know How to Reason, Thinking Models Learn When
What do thinking language models learn during training that their base models lack? We first present an unsupervised method that discovers…
Beyond Reactivity: Measuring Proactive Problem Solving in LLM Agents
LLM-based agents are increasingly moving towards proactivity: rather than awaiting instruction, they exercise agency to anticipate user nee…
When Assisting One Disempowers Another
Personal AI agents are increasingly deployed in shared environments, where their actions affect not just the primary user they are assistin…
VASP Agent: An Agentic Framework for Autonomous First-principles Calculations
Large Language Models (LLMs) are increasingly embedded in agentic frameworks for scientific discovery. First-principles materials computati…
Implementing Metric Temporal Answer Set Programming
We develop a computational approach to Metric Answer Set Programming (ASP) to allow for expressing quantitative temporal constraints, like…
Agentic AI for Commercial Insurance Underwriting with Adversarial Self-Critique
Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and de…
Self-Routing: Parameter-Free Expert Routing from Hidden States
Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a lea…
Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents
Do LLM agents act on the reasoning they state? This question of process fidelity is central to LLM-based social simulation, yet hard to mea…
EVA-Net: Subject-Independent EEG Motor Decoding with Video-Derived Motor Priors
Practical non-invasive Brain-Computer Interface (BCI) systems require EEG decoders with strong cross-subject generalization and minimal cal…
What Type of Inference is Active Inference?
Active inference casts decision-making as inference, with the Expected Free Energy (EFE) unifying goal-directed and information-seeking beh…
A Three-Layer Framework for AI in Scientific Discovery
Current discussions of AI in scientific discovery are often dominated by two visible capabilities: search over existing knowledge and execu…
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
Previous research has evaluated animal welfare using question-and-answer benchmarks. This study investigates whether these evaluations also…
Reward as An Agent for Embodied World Models
While RL has become a promising tool for refining world models, existing methods largely rely on conservative rollouts near the training di…
Accelerating Returns and the Qualitative Engine for Science
Ray Kurzweil described a thesis of accelerating returns, which is the most influential narratives in discussions of technological progress.…
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbre…
Replication in Visual Diffusion Models: A Survey and Outlook
Visual diffusion models have revolutionized the field of creative AI, producing high-quality and diverse content. However, they inevitably…
Trust-free Personalized Decentralized Learning
Personalized collaborative learning in federated settings faces a critical trade-off between customization and participant trust. Existing…
Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models
Large language models (LLMs) have shown strong performance in automated code generation, with few-shot prompting widely used for its simpli…
The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer
Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learni…
Toward AI standardization: A triadic human-ai collaboration framework for multi-level autonomous mobility
The goal of the current study is to introduce a triadic human-AI collaboration framework that could be applied in transportation systems su…
Narrative-Centered Emotional Reflection: An Early Prototype for AI-Supported Emotional Self-Reflection
Reflexion is an AI-powered prototype designed to explore structured emotional self-reflection. By integrating emotion detection, layered re…
Explainable embeddings with Distance Explainer
While eXplainable AI (XAI) has advanced significantly, few methods address interpretability in embedded vector spaces where dimensions repr…
Position: EU AI Act's Research Exemptions Can Break the Publication Norms of Major AI Conferences
The EU has become one of the vanguards in regulating the digital age. A particularly important regulation in the Artificial Intelligence (A…
Learning The Minimum Action Distance
This paper presents a state representation framework for Markov decision processes (MDPs) that can be learned solely from state trajectorie…
Detoxify: A framework for abusive text transformation using LLMs
Although Large Language Models (LLMs) have demonstrated significant advancements in natural language processing tasks, their effectiveness…
Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders
Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts. A core SAE traini…
Reduced NEXI protocol for the quantification of human gray matter microstructure on the Connectome 2.0 scanner
Biophysical diffusion MRI models like Neurite Exchange Imaging (NEXI) are essential for probing gray matter microstructure, estimating comp…
Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance
Autoregressive (AR) models based on next-scale prediction have emerged as a powerful tool for image generation, but they face a critical we…
Medix: Out-of-Distribution Detection from Unlabeled Wild Data via Robust Gradient Statistics
Out-of-distribution (OOD) detection plays a crucial role in ensuring the robustness of machine learning systems deployed in real-world appl…
SmartMixed: A Two-Phase Training Strategy for Adaptive Activation Function Learning in Neural Networks
The choice of activation function plays a critical role in neural networks, yet most architectures still rely on fixed, uniform activation…
LLM4Delay: Flight Delay Prediction via Cross-Modality Adaptation of Large Language Models and Aircraft Trajectory Representation
Flight delay prediction has become a key focus in air traffic management (ATM), as delays reflect inefficiencies in the system. This paper…
Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders
We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptually motivate…
SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios
Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyda…
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
AI coding assistants produce vulnerable code in 45\% of security-relevant scenarios~\cite{veracode2025}, yet no public training dataset tea…
OpenGround: Planning-based Online Perception for Open-World 3D Visual Grounding
3D visual grounding aims to locate objects based on natural language descriptions in 3D scenes. Existing supervised methods are limited by…
KernelEvolve: Scaling Agentic Kernel Coding for Heterogeneous AI Accelerators at Meta
Making deep learning recommendation model (DLRM) training and inference fast and efficient is important. However, this presents three key s…
Multi-Task Instruction Tuning via Data Scheduling for Low-Resource Arabic SpeechLLMs
Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dial…
From Global to Granular: Revealing IQA Model Performance via Correlation Surface
Evaluation of Image Quality Assessment (IQA) models has long been dominated by global correlation metrics, such as Pearson Linear Correlati…
StepShield: When, Not Whether to Intervene on Rogue Agents
Agent safety benchmarks measure whether a monitor detects harm, not when. Yet timing is the difference between intervention and autopsy. We…
Universal Algorithm-Implicit Learning
Current meta-learning methods are constrained to narrow task distributions with fixed feature and label spaces, limiting applicability. Mor…
Transformers converge to invariant algorithmic cores
Training selects for behavior, not circuitry: many weight configurations can implement the same function. Studying any single trained neura…
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks,…
Volumetric Directional Diffusion: Anchoring Uncertainty Quantification in Anatomical Consensus for Ambiguous Medical Image Segmentation
Ambiguous 3D medical image segmentation often involves boundaries where different expert delineations are non-identical yet clinically plau…
What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection
Audio anti-spoofing systems are typically trained to assign one authenticity label to an entire speech utterance. This formulation becomes…
From Passive Observer to Active Critic: Reinforcement Learning Elicits Process Reasoning for Robotic Manipulation
Accurate process supervision remains a critical challenge for long-horizon robotic manipulation. A primary bottleneck is that current video…
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
Understanding and generating 3D objects as compositions of meaningful parts is fundamental to human perception and reasoning. However, most…
SpatialFly: Implicit 3D Prior-Guided Visual Reparameterization for Continuous UAV Vision-and-Language Navigation
UAVs play an important role in applications such as autonomous exploration, disaster response, and infrastructure inspection. However, UAV…
DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning
Multimodal deception detection aims to identify deceptive behavior by analyzing audiovisual cues for forensics and security. In these high-…
HST-HGN: Heterogeneous Spatial-Temporal Hypergraph Networks with Bidirectional State Space Models for Global Fatigue Assessment
It remains challenging to assess driver fatigue from untrimmed videos under constrained computational budgets, due to the difficulty of mod…
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most ima…
Do Quantum Transformers Help? A Systematic VQC Architecture Comparison on Tabular Benchmarks
Variational quantum circuits (VQCs) are a leading approach to quantum machine learning on near-term devices, yet it remains unclear which c…
Learning from Execution: Self-Evolving Memory for Private-Library Code Generation
Large Language Models (LLMs) have achieved strong performance on general code generation, but their effectiveness drops sharply in enterpri…
Constitutional Governance in Metric Spaces
Computational social choice and algorithmic decision theory offer rich aggregation theory but no end-to-end process for egalitarian self-go…
Few Channels Draw The Whole Picture: Revealing Massive Activations in Diffusion Transformers
Diffusion Transformers (DiTs) and related flow-based architectures are now among the strongest text-to-image generators, yet the internal m…
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet m…
Geometry-Aware Uncertainty Coresets for Robust Visual In-Context Learning in Histopathology
Vision-language models (VLMs) can couple visual perception with open-ended clinical reasoning, making them attractive for computational his…
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating…
MambaGaze: Bidirectional Mamba with Explicit Missing Data Modeling for Cognitive Load Assessment from Eye-Gaze Tracking Data
Real-time cognitive load assessment from eye-tracking signals could enable adaptive human-centered AI in safety-critical applications such…
MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding
Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks. Howev…
An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models
Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models behav…
eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian
We present eCREAM-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals.…
SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents
Large language models (LLMs) are increasingly deployed in automated code-review systems, where their approvals can determine which code is…
Beyond Correctness: Enhancing Architectural Reasoning in Code LLMs via Scalable Labeling with Agentic Judgment
LLMs have substantially improved software engineering yet real-world development requires architectural understanding. Such understanding i…
IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction
Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable laws or occur in area…
When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However,…
A Neuromorphic Reinforcement Learning Framework for Efficient Pathfinding in Robotic Mobile Fulfillment Systems
Dynamic environmental changes, confined workspaces, and stringent real-time constraints make pathfinding in Robotic Mobile Fulfillment Syst…
AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming
Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments.…
The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report
AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one spe…
Application of LLMs to Threat Assessment of Foreign Peacekeeping Missions
We present a novel approach for applying Large Language Models (LLMs) to threat assessment in the context of foreign peacekeeping missions.…
CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention
Recurrent delta-rule models keep a fixed-size state matrix S (d_v x d_k) that compresses all past context. The state of the art (GDN-2) gat…
Categorizing Mathematical Concepts with LLM Voting Ensembles in Mathswitch
Mathswitch is an open-source project that imports mathematical concept records from sources such as Wikidata, Wikipedia, MathWorld, Encyclo…
Token Geometry
Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface be…
note上方修正 「AI活用、想定以上」で人件費率低下 2Q累計営業益は「20倍超」に
AI活用で生産性が向上。人員数は緩やかに減っており、効率的な事業運営ができているという。
「Claude Fable 5」サブスク、突如5日間延長 ユーザー悲喜こもごも「寝ずに頑張ったのに」「制限リセットして」
歓迎の声もある一方で、「睡眠と健康を犠牲にして使い切った直後に延長を知った」「最後の最後で発表するのはやめてくれ」といった恨み節も。
Meta、初のエージェント型画像生成AI「Muse Image」公開 動画生成AI「Muse Video」もプレビュー
Metaは、Meta Superintelligence Labs(MSL)初のメディア生成モデルとなる画像生成AI「Muse Image」を公開した。検索やコーディングなどのツールを自律的に呼び出し、生成した画像を自己修正するエージェント型だ。「Meta AI」アプリやIns…
NTTドコモビジネスがIOWN活用の分散GPU環境を提供、25GBを2秒で転送
NTTドコモビジネスは、次世代ネットワーク「IOWN APN」を活用し、全国8拠点に分散したGPUを統合利用できる実証環境の提供を開始した。電力などの制限を解消し、オンデマンドなリソース確保やデータ主権に対応した分散AI基盤の実用性を検証できる。
Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
The new image-generating model has numerous use cases, including advertising, decorating and creator-based opportunities.
IT予算の9割が人件費に消える――日本オラクル社長が切り込む「企業最大の課題」
日本オラクルの三澤智光社長が、日本企業のIT課題に切り込んだ。同社長が指摘する「IT投資の構造的問題」「オンプレミスシステムが抱える課題」とは何か。
逸材も不採用に……トヨタ系企業がAIで「人材採用のブレ」解消、800件の応募を効率化
年間約800件に及ぶ応募書類の確認に加え、採用基準の共有や面接記録の作成など、人事担当者の業務負担が課題となっていた。トヨタテクニカルディベロップメントは、これらの課題に対し、AIを活用した採用業務の見直しに取り組んだ。
ガバメントAI「源内」は自治体で本当に使えるのか? クラウド依存による“3つの落とし穴”
デジタル庁が公開したガバメントAI「源内」(GENAI)は、政府によるAIプラットフォームの公開という点で画期的な取り組みだ。一方で、自治体での実運用を考える上で無視できない論点も見えてくると、CIO補佐官として自治体DXに携わる筆者が解説する。
エンジニアの採用枠をトークン予算に回せるか? HRBrain常務が語る、「財布の壁」に挑む人事の未来
HRBrainの常務執行役員CSaO小山径氏は「AIも、これからは経営資源の一つになっていく」と話す。小山氏が見据えるAI活用の展望とは?
AIがExcel作業を丸ごと自動化? 企業の定型業務を効率化へ
生成AIによる業務自動化が広がる一方、取得したExcelやCSVファイルの集計・加工には依然として人手を伴う。データ収集からファイル処理までをAIが一貫して担うことで、企業に残る定型作業の効率化が進むだろう。
AIが破壊するIT業界の“人月商売” 「SIerの死」後に“生き残る者”の正体
米Anthropicの機能発表で約8300億ドルの時価総額が消し飛んだ「AIショック」。画面の使いやすさで稼いできたSaaSや、時間と人数を積み上げるSIerの「人月商売」が崩壊の危機に直面している。だが、業界ルールにのっとった複雑な計算(ビジネスロジック)を握る企業は依然とし…
Anthropicの「Claude Cowork」がWebとモバイルでも利用可能に まずはMaxプランから
Anthropicは、AIエージェント機能「Claude Cowork」をWebブラウザとモバイルアプリ向けに展開すると発表した。これまでデスクトップアプリ限定だったが、スマートフォンなどからも作業の進捗確認や指示出しが可能になる。今後は順次提供を拡大し、まずは上位のMaxプラ…
Why the rise of open source AI isn’t hurting Anthropic … yet
Open source models’ success isn’t coming at the expense of frontier labs. Instead, they each seem to capture two phases of the same life cy…
ソフトバンクの「1人100エージェント」を支える独自AIゲートウェイ「Cloud Proxy」の正体
生成AIやAIエージェントを全社展開する際、企業はセキュリティやガバナンス、性能といった課題に直面しがちです。ソフトバンクは「全社で1人100エージェント」構想の実現に向けて、AI利用の入り口となる共通基盤「Cloud Proxy」を内製しました。その設計思想や性能強化の取り組…
Microsoft joins AI cost-cutting trend by relying more on its own models
Microsoft is the latest Silicon Valley giant to cut back on its AI spending.
Discord admits AI moderation bug wrongfully banned users over harmless images
The company confirmed that the issue had been affecting accounts since May, with an additional 200 users banned over the weekend before its…
「Claude Fable 5」のサブスク提供、延長 12日(太平洋標準時)まで 日本時間では13日午後
米Anthropicは7月7日(現地時間)、AIモデル「Claude Fable 5」を有料サブスクリプションから利用できる期間を、7月12日午後11時59分(太平洋標準時)まで延長すると発表した。日本時間では13日午後3時59分までとなる。
Claude Cowork expands to mobile and web
With this update, users can start a task from their desk, get status updates on their phone, and pick up the finished output later — even i…
2026-07-07(805件)
Savi’s app aims to protect consumers from realistic AI scams like kidnappers demanding ransom
The company just raised $7 million in seed funding, and is launching its app for iPhone and Android on Tuesday.
The first American autonomous ground vehicles are fighting in Ukraine
Forterra has deployed more than 100 of its self-driving ATVs in conflict zones in Ukraine.
三菱UFJ半沢淳一社長「企業の資金ニーズに対応」 個人向け「エムット」でAI活用も
三菱UFJフィナンシャル・グループ(MUFG)の半沢淳一社長が産経新聞のインタビューに応じ、日銀の利上げで「金利のある世界」が本格化する中、企業の資金需要を取り込むため、事業戦略の提案力を高める考えを示した。
「AIが引用するドメイン」不動の首位は……
2位には前回5位のnote.comが浮上した。前回2位のWikipedia(ja.wikipedia.org)は3位に後退した。
【Fable 5に聞いてみた】サブスク終了後の効果的な使い方は? トークン節約法は? 本人インタビュー
注目を集めている米AnthropicのAIモデル「Claude Fable 5」の性能はどれほどなのか。ITmedia AI+編集部で試した結果を紹介する。今回は「従量課金制移行後にFable 5を効果的に使う方法」について聞いてみた。
Anthropicが教える「Fable 5活用術」 まず確認すべきは「自分が何を知らないか」
米Anthropicは、AIモデル「Claude Fable 5」の活用ガイドを公式ブログで公開した。AIコーディングツール「Claude Code」で同モデルを使う際の具体的なテクニックを紹介している。
iFLYTEK-Embodied-Omni Technical Report
General-purpose embodied agents must understand multimodal instructions, anticipate how their environment will evolve, and produce precise…
Internal Pluralism and the Limits of Pairwise Comparisons
Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.g., in participatory design or alignm…
ASK in the Dark: Uncertainty-Gated LLM Assistance under Partial Observability
Reinforcement learning agents operating under partial observability must act on incomplete information, making them natural candidates for…
Automated Data Readiness for Scientific AI
Leadership computing facilities steward large-scale scientific datasets that routinely require substantial transformation before serving as…
SwarmResearch: Orchestrating Coding Agents for Open-Ended Discovery
Long-running coding agents such as autoresearch can persistently discover optimizations for open-ended problems. However, they tend to conv…
Object-Centric Environment Modeling for Agentic Tasks
Large language model (LLM) agents can improve through accumulated experience, but free-form textual memories become difficult to maintain,…
MedCalc-Pro: Solving Complex Medical Calculations with LLM Agents
Current benchmarks for evaluating large language models (LLMs) in medical calculation are largely based on simplified settings, where each…
Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models
Large language models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet ensuring their simultaneous safety,…
VERITAS: Towards a General-Purpose Replication Tool for Scientific Research
AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of publ…
A Sliding-Window-Based Reinforcement Learning for Dynamic Assembly Flow Shop Scheduling with Multi-Product Delivery
Multi-product kitting delivery imposes significant challenges for real-time scheduling in hybrid manufacturing systems that integrate proce…
Evaluating Generative Agents with Actions Grounded in Socially Distributed Task Environments using Incognita
Effective agency in social environments depends on when an agent seeks knowledge, when it acts, and whether its actions are justified by ac…
Reinforcement Learning for Evidence-Seeking Diagnostic Reasoning with Large Language Models
Recent reasoning-centric Large Language Models (LLMs) have made significant strides, yet they predominantly operate on a passive-inference…
Beyond Forecasting: The Belief-to-Trade Layer in Prediction-Market Agents
Forecasting future events has attracted growing attention as a testbed for general-purpose AI. A natural way to ground this evaluation is l…
Human-Centric Reflective Architecture for Human-AI Collaborative Decision-Making
The use of Large Language Models (LLMs) across diverse areas of human activity-ranging from everyday tasks to safety-critical applications-…
Silicon Sampling via Cross-Survey Transfer
Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting…
APeB: Benchmarking Personalization Ability of Large Language Model Agents
LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent in…
Organizational Memory for Agentic Business Process Execution
LLM-based agents offer new opportunities for automating business process execution beyond the limits of rule-based systems. However, genera…
Embodied Operators and Benchmarking: Toward Reusable and Deployable Embodied Intelligence Systems
Embodied intelligence systems require not only end-to-end policy models, but also reusable functional modules that transform multimodal obs…
Reflective Dialogue or Prompt Refinement? Effects of Tutor Scaffolding on Students' Independent LLM Use for Programming
While Large Language Models (LLMs) can provide personalized support in learning, several studies have raised concerns regarding their use i…
When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions
Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-st…
From Mobile Data to Business Insights: An End-to-End Analytics Framework for Large-Scale Urban Mobility Analysis and Decision Support
Real time location data derived from mobile applications is a powerful tool for addressing various urban challenges, including tourism plan…
Efficient bias mitigation in T2I diffusion models using Concept Graphs
Text-to-Image diffusion models often propagate harmful bias inherited from the training data. Existing bias mitigation techniques typically…
Personalized Causal Recourse: A Human-In-The-Loop Approach
Algorithmic recourse addresses the challenge of providing tailored recommendations to users affected by unfavorable machine learning decisi…
Demonstrating Generalization Failures via Mixtures of Conditional Policies
Post-training of frontier language models is conducted on curated task suites, and inevitably leaves a distribution shift between training…
MentalThink: Shaping Thoughts in Mental SVG World
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental…
Applying Answer Set Programming with Fuzzy Membership Functions: a Case Study
Human reasoning often operates through qualitative concepts expressed by linguistic labels such as high, low, expensive, or cheap, whose in…
How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs
As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with ou…
The Role of Rigor in Artificial Intelligence
Artificial intelligence (AI) has achieved extraordinary capabilities despite lacking many of the conceptual and scientific foundations asso…
Robust Feasible Route Construction through Collaborative Partition Optimization
Large-scale Capacitated Vehicle Routing Problems (CVRPs) are commonly solved by partitioning customers into smaller routing problems that c…
Agent Reinforcement Learning via Pivotal-Aware Self-Feedback Retry
Large language model (LLM) agents have shown strong decision-making capabilities in long-horizon interactive tasks, yet they still struggle…
Explainable Reinforcement Learning for Adaptive Traffic Signal Control
Reinforcement Learning (RL) has emerged as a powerful paradigm for adaptive traffic signal control. However, in safety-critical infrastruct…
Can Conversational Temporal Dynamics Improve Depression Detection in Dyads? A Preliminary Investigation in Multi-Modality Perspectives
Automatic depression detection from clinical interviews typically models the semantic content and acoustic characteristics of participant s…
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process
Unified multi-modal models (UMMs) have shown promising interleaved text-image reasoning capabilities, yet effectively optimizing such multi…
Folding, Reasoning, and Scaling with Open-source Drug Discovery Engine
Accurately modeling biomolecular interactions is a central bottleneck in biology and therapeutic discovery. Here, we introduce Open Drug Di…
Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground Truth
As LLMs generate increasingly long outputs, effective uncertainty estimation must identify errors at fine-grained levels rather than discar…
Harness-Aware Self-Evolving: Co-Evolving Model Weights, Harness, and Task Solutions
Self-evolving frameworks usually optimize task solutions while treating the surrounding harness as fixed. We introduce Harness-Aware Self-E…
Online Linear Programming for Multi-Objective Routing in LLM Serving
We study the online routing problem in large language model serving, where requests arrive sequentially and must be dispatched to parallel…
Explainable AI for Screening Abuse-Related Trauma in Bangladeshi Children: A Training-Free Multimodal Framework Evaluated on Noise-Aware Synthetic Data
Bangladesh has an estimated 1.17 mental-health professionals per 100,000 population and only six child psychiatrists nationwide. No Bengali…
What is Left for Us? Second Scholarship Against the Degradation of Research by AI
We argue that generative AI can degrade research by eroding the very practices through which scholarly judgement is formed and academic tru…
PLACEMEM: Toward a Compute-Aware Memory Plane for Lifelong Agents
Lifelong agents need more than larger context windows and better retrieval. They need memories that can persist, evolve, and be corrected w…
Forethought: Verifiable Reasoning from Neurosymbolic Primitive Programming
Current agentic workflows usually involve decomposing user requests into sequences of tool calls with correctly resolved parameters, the re…
Language models guide symbolic equation discovery by controlling search
Scientific equation discovery must combine broad domain priors with strict numerical testing. Symbolic regression supplies numerical ground…
A Clustering-Based Framework for Identifying Suspicious Trading Patterns in Capital Market
Market manipulation is the dubious practice of manipulating stock prices in order to make a quick profit, which truly degrades confidence o…
Agentic IoT: Architectures, Applications, and Challenges Toward the Internet of Agents
The integration of AI into Internet of Things (AIoT) systems has gradually transformed them from passive data collection infrastructures in…
Unsupervised Features Mining via Activation Geometry
Interpretability methods aim to reveal the features represented inside large language models (LLMs). Many existing methods begin with label…
Biological Motifs for Agentic Control
The transition of Large Language Models (LLMs) from passive generators to autonomous agents has introduced significant challenges in reliab…
Progress- and Reliability-Oriented Group Policy Optimization for Agentic Reinforcement Learning
Group-based reinforcement learning (RL) has become an effective paradigm for improving large language model agents on long-horizon interact…
Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal
Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that model…
Agentic SABRE: An Uncertainty-Aware Neuro-Symbolic Multi-Agent Framework for Adaptive Ransomware Detection
Ransomware has evolved into a complex, adaptive, and fast-moving adversary category in which static signatures and monolithic classifiers f…
HAS-Bench: Evaluating LLM-Based Human-Agent Systems under Configurable Human Participation
Large language models increasingly operate in settings where humans are active collaborators rather than passive task providers. We introdu…
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
Multimodal GUI agents read an interface through two redundant channels: the rendered pixels of a screenshot and a serialized structure such…
Server-side Anti-cheat in FPS games for Aimbot detection using Deep learning and Machine learning
Modern video games are becoming more complex day by day. Most of these modern games are multiplayer first-person shooter (FPS) games. The r…
Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs
We present Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super optimized for interactive deployment. We designed the m…
Decentralized Aggregation of LLM Predictions via Wagering Mechanisms
It is increasingly common to aggregate predictions from multiple LLMs, each with domain expertise or access to private tools and data, to i…
MechMath Agent Team: LLM Driven Agents for Mathematical Research
AI reasoning has become a central focus in contemporary artificial intelligence, largely driven by the success of large language models. Ho…
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL
Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as rewa…
Agent Step Value: State-Transition Measurement with State-Grounded LLM Evaluators
Most agent evaluations collapse a multi-step trace into a final answer, a success flag, or a trajectory-level score. These aggregates obscu…
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes
Large language models have made research ideation increasingly accessible, yet effective idea development requires more than generating can…
Why Pure Reasoning is Not Enough: Nature as the Source of Mathematical Innovation
We advance the hypothesis that human mathematical reasoning, constrained by both the undecidability and the computational intractability of…
Compressing the Validation Bottleneck: An Agentic Self-Driving Lab for Scientific Discovery
Agentic AI-for-Science can automate ideation, planning, and analysis, but final validation still depends on real experiments. A self-drivin…
VLA Grounder: Language-Conditioning Space Optimization for Black-Box VLA Models
Vision-Language-Action (VLA) models are commonly treated as end-to-end action policies conditioned on natural-language task descriptions. I…
Measuring Harness-Induced Belief Divergence in Multi-Step LLM Agents
Software-agent benchmarks usually report whether an agent solves a task, but the agent reaches that outcome through a harness that controls…
Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models
Large language models (LLMs) generate fluent outputs that can be wrong. Unlike humans, who often exhibit cues when providing false informat…
Detecting Answer-Driven Reasoning in LLM-Based Educational Tutors via Truncated Chain-of-Thought Auditing
Large language model (LLM) tutors often produce fluent step-by-step explanations, but a correct and pedagogically formatted response does n…
Attention Limited Reward Learning
Pairwise human comparisons are a primary interface through which modern AI systems learn human preferences. RLHF and related alignment pipe…
Governed Individuation: Cryptographically Decoupling an Agent's Learning from Its Authority
Autonomous agents are moving from sandboxed text generators to operators of code, data, and physical infrastructure, and they increasingly…
MRMS: A Multi-Resolution Memory Substrate for Long-Lived AI Agents
Long-lived AI agents require continuity across interactions, but continuity cannot be obtained by simply extending the prompt window. An ag…
Formal Disco: Scalable Open-Ended Generation of Formally Verified Programs
The cost of producing code is rapidly diminishing with increasingly capable AI agents, while quality assurance of generated programs has no…
Integrated Altruistic and Fairness Preference Induces Advanced Mutual Cooperation in Sequential Social Dilemmas
Inducing cooperation among distributed agents is still a difficult problem in the field of multi-agent reinforcement learning (MARL), parti…
FORGE: Research-Trajectory Hijacking Attacks on Deep Research Agents
Deep research agents decompose open-ended queries into subtasks, retrieve web evidence over multiple rounds, and synthesize long-form repor…
FM-ChangeNet: Learning Change through Pathwise Feature Transport
We present FM-ChangeNet, a pathwise-supervised framework for change detection that reformulates bi-temporal reasoning as continuous transpo…
AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization
Physical design quality-of-results~(QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages. Each…
CARL: Constraint-Aware Reinforcement Learning for Planning with LLMs
Despite their strong reasoning capabilities and extensive world knowledge, Large Language Models (LLMs) frequently generate plans that viol…
Medi-Gemma: A Hybrid Clinical Decision Support System Integrating Deterministic EMR Analytics and Retrieval-Augmented Generation
Deploying Large Language Models (LLMs) in high-stakes clinical settings remains limited by structural hallucinations, weak deterministic re…
STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training
Reinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse…
Quantum-Inspired Harmonic Decision Models: A Computational Framework for Music Generation
This paper introduces a quantum-inspired computational framework for harmonic decision-making in music. The proposed approach formulates ha…
Toward Trustworthy Large Language Model Agents in Healthcare
Healthcare appointment scheduling remains a persistent operational bottleneck, driven by manual coordination, fragmented legacy systems, an…
Diffusion-Guided Uncertainty-Aware Delayed Policy Optimization
Reinforcement learning in real world environments often suffers from severe performance degradation due to delayed feedback. Existing appro…
ASSEMCAD: Production-Ready CAD Assembly Generation from Natural Language
Recent advances in large language models and programmatic CAD have significantly improved Text-to-CAD generation for individual parts. Howe…
TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios
Among the five primary human senses, tactile is arguably the most fundamental to survival, as it enables the perception of physical contact…
DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation
Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent…
The Changing Role of Symbolic Methods in Artificial Intelligence
Why do intelligent systems need to perform explicit symbolic reasoning? Computer science has traditionally regarded symbolic reasoning as a…
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments
Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the ur…
CP-WSP: A Declarative CP-SAT Framework for Configurable Multi-Constraint Workforce Scheduling
Workforce scheduling is an NP-hard combinatorial optimization problem requiring simultaneous satisfaction of labor regulations, coverage re…
Rethinking On-Policy Self-Distillation for Thinking Models
Self-distillation is a promising recipe for self-improvement in language models. In this setting, a model can serve as its own teacher when…
ClassicLogic: A Knowledge-Driven Benchmark of Classic Puzzle Games for Evaluating Compositional Generalization
Compositional generalization, the ability to understand and produce novel combinations of known components, remains a fundamental challenge…
Reason, Reward, Refine: Step-Level Errors Corrections with Structured Feedback for Physics Reasoning in Small Language Models
Physics reasoning fails structurally in small language models: an error at any step propagates forward, corrupting every inference that fol…
EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer
Agent self-evolution in long-horizon LLM systems is largely procedural: useful experience is not merely stored information, but reusable pr…
MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models
JEPA world models predict the next latent state with a single deterministic predictor trained by latent regression. We show that this fails…
MetaSkill-Evolve: Recursive Self-Improvement of LLM Agents via Two-Timescale Meta-Skill Evolution
Recent LLM agents tackle increasingly long-horizon, open-ended tasks, and external skills, reusable procedural knowledge supplied to the ag…
Evaluating and Understanding Model Editing for Medical Vision Language Models
Model editing promises a fast, targeted way to correct post-deployment mistakes in medical vision-language models (VLMs) without costly ret…
OptiAgent: End-to-End Optimization Modeling via Multi-Agent Iterative Refinement
We propose OptiAgent, a multi-agent framework that, given a natural language description of an Operations Research problem, is able to outp…
Graph Sparse Sampling: Breaking the Curse of the Horizon in Continuous MDP Planning
Planning under uncertainty in continuous domains is essential for autonomous systems, yet computationally demanding. Tree-based search meth…
SovereignPA-Bench: Evaluating User-Owned Personal Agents under Evolving Intent, Platform Mediation, and Consent Constraints
Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use too…
LLM-as-a-Verifier: A General-Purpose Verification Framework
Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of LLMs. In thi…
SiamixFormer: a fully-transformer Siamese network with temporal Fusion for accurate building detection and change detection in bi-temporal remote sensing images
Building detection and change detection using remote sensing images can help urban and rescue planning. Moreover, they can be used for buil…
PotatoGANs: Utilizing Generative Adversarial Networks, Instance Segmentation, and Explainable AI for Enhanced Potato Disease Identification and Classification
Numerous applications have resulted from the automation of agricultural disease segmentation using deep learning techniques. However, when…
Specific Domain Ontology Construction Using Large Language Models
Ontologies are useful structures to organize and maintain information that can be understood both by humans and systems. However, since the…
Neural-Network Inverse Design of SRF Cavities and Transmons for Bosonic Quantum Computation
Three-dimensional superconducting radio-frequency (SRF) cavities provide exceptionally long-lived electromagnetic modes and, when coupled t…
GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads
OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back…
AutoResearch: An Execution-Grounded Multi-Agent Framework for Reliable Research Workflow Automation
Automated research agents increasingly generate code, retrieve literature, and draft scientific artifacts, but they often fail to verify wh…
SWIFT: Spatio-temporal Wavelet Integrated Forecasting Framework for Workload Traces
Accurate cloud workload forecasting is pivotal for efficient resource management but remains challenging as workloads are highly volatile a…
PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving
We present PEEK, a lightweight scheduling and eviction framework for both online (streaming) and offline (batch) LLM serving; this paper fo…
The Hidden Water Geography of U.S. Hyperscale Data Centers in the AI Era
Water use by data centers is routinely reported as a single footprint, but water is consumed through two physically distinct pathways: at t…
Not Every Sync Is Safe: Calibrated DiLoCo Scheduling for Shared AI Infrastructure
DiLoCo-style training reduces communication by letting learner islands train locally before occasional outer synchronization, making it att…
DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences
Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spati…
Double-Helix Active Geometry: LiDAR-Anchored Multi-View Depth with Selective Abstention
Consumer depth sensors such as the LiDAR scanner on recent iPhones provide metric range, but their useful range is short and their returns…
Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration
Diffusion-based text-to-image models can synthesize complex and highly structured visual content, yet the emergence and evolution of semant…
From Raw Segmentations to Simulation-Ready Cardiac Meshes: An Automated Framework for Anatomical Reconstruction and Virtual Cohort Generation
Computational models of the human heart are widely used to study electromechanical and fluid-dynamical cardiac function and to support appl…
CRODA-ST: Single-Target Cross-Receiver Open-Set Radio Fingerprint Recognition
Radio frequency fingerprint identification (RFFI) provides a physical-layer credential for Internet of Things devices, but open-set decisio…
DOSE-I: A Multimodal Biosignal Dataset of Procedural Sedation for Endoscopy -- Technical Report
In this document, we describe characteristics and technical details of the multimodal biosignal dataset DOSE-I of procedural sedation for e…
Additive Causal Construction for Transferable and Reconfigurable Cross-System Learning in Multi-Source Image Fusion
In multi-source image fusion scenarios, heterogeneous inputs are typically driven by distinct generative mechanisms and can be viewed as a…
From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving
The key-value (KV) cache has become a first-order memory object in LLM serving rather than a temporary per-request tensor. This survey clas…
Criterion-Conditional In-Context Learning: Evaluating Criterion-Shift Adaptation in Vision-Language Models
Vision-language models can perform new tasks without parameter updates through in-context learning (ICL), whose core mechanism is utilizing…
Homer: Understanding Long-form Videos with Hierarchical Memory and Agentic Reasoning
Multimodal large language models excel on short clips but struggle on hour-long videos in an online setting, where frames are processed inc…
An automated method of identifying incorrectly labelled images based on the sequences of loss functions of deep learning networks
Deep learning is widely applied in medical image analysis, but up to 10% of manually labelled images may be incorrect, degrading model perf…
AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents
Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safe…
Knowledge-Centric Information Systems
For decades, data engineering has developed mature architectural principles for integrating, governing, validating, cataloging, and serving…
Fusion: A Framework for Unified Sequential Token AdaptatIon in VisiOn TraNsformers
Vision Transformers achieve strong image classification accuracy but process all image regions with nearly the same computation, even when…
The agent creates, we validate: A Lightweight Framework for Agentic Artifact Generation
Generating structured artifacts with Large Language Models - e.g. database queries, threat framework mappings, entity schemas - is relative…
COMET: Combinatorial Optimization for Multiplex Editing Targets Via Constraint-Preserving QAOA
Multiplex CRISPR-Cas9 gene editing requires selecting one guide RNA per target gene subject to cross-gene interactions: a constrained combi…
Fine-Grained Computation Offload for Off-the-Shelf Servers in Tens of Lines
Hardware accelerators now sit on the critical path of online serving. GPUs, FPGAs, and increasingly remote services such as hardware securi…
QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting
Time-series forecasting supports decisions in finance, en-ergy, transportation, public health, and industrial monitoring. Recent foundation…
Federated Learning for Object Detection: Enabling Collaborative Drone Learning Without Centralizing Data
Object detection is a fundamental capability for AI-driven perception in safety-critical drone and edge-vision systems, including disaster…
Post-Generation Curation of Synthetic Images via Homogeneous-Heterogeneous Splitting
Recent generative models can produce high-quality synthetic images, offering scalable training training data for data-hungry models. Existi…
Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving
Real-time interaction models -- Moshi, MiniCPM-o, Qwen-Omni -- turn serving into a periodic real-time task: on every frame a session ingest…
K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos
MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexp…
S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval
As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experience…
LLMoxie: Exploring Agentic AI for Scientific Software Development
In this paper, we describe LLMoxie, an institutional AI platform whose three-tiered architecture supports multi-cloud and on-premise infere…
Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale
There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various…
Diagnosing Aerial-View Object Detectors with Foundational Image Generative Models
Recent advances in large-scale image generative models enable photorealistic scene synthesis with controllable attributes. Beyond data augm…
Signal from Space: Detecting Schools and Towers to Bridge the Digital Divide
Reliable internet access is essential for modern education, yet millions of school-aged children especially in developing regions remain of…
SMOCS: A Streaming Framework for Simplified Deployment, Monitoring, and Optimization of ML Systems in Production
Machine learning has demonstrated significant potential for real-time monitoring, optimization, and control of scientific facilities. Howev…
Echoes of Unrest: A Multimodal NLP Framework for Early Warning of Fake News and Violence-Driven Mob Activity
Rapid growth in social media has transformed global communication by enabling fast information exchange, but it has also accelerated the sp…
Out-of-Distribution Generalization of Risk Aversion in Language Models
Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AI…
Gemma 4 Technical Report
We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance c…
Safe Inference-Time Alignment via Lagrangian Reward Augmentation
Inference-time alignment steers a frozen language model during decoding using auxiliary reward signals, avoiding the cost of repeated weigh…
A Preliminary Study on Explaining Risk of Code Changes using LLM-Based Prediction Models
Predictions by machine learning (ML) and artificial intelligence (AI) models are often received skeptically unless they are paired with int…
Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes
As LLMs are increasingly deployed as autonomous adjudicators in semi-open textual game environments, robust rule adherence becomes critical…
Training Hybrid Block Diffusion Language Models with Partial Bidirectionality
High-throughput long-context generation is one of the central challenges for large language models. Generation is typically memory-bandwidt…
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure
Personal agents will increasingly negotiate on behalf of users: splitting costs with other personal agents, appealing platform decisions, e…
Vision Token Manipulation Attacks on Cloud-Edge Inference of Large Vision-Language Models
Cloud-edge Large Vision-Language Model (LVLM) inference enables efficient deployment by splitting computation between edge devices and clou…
JavaVulBench: A Java Vulnerability Benchmark with Realistic Splits, a Unified Multi-Backend Harness, and a Leakage-Aware Evaluation Mode
We release \textsc{JavaVulBench}, a benchmark dataset and evaluation harness for Java vulnerability detection. The dataset contains $\sim$3…
Differential Amplifier-Inspired AmpAttention for Multi-View Robotic Manipulation
Multi-view robotic manipulation methods with the attention mechanism have recently achieved significant progress in both training efficienc…
Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI
Large language model agents driving security tool suites over the Model Context Protocol are increasingly common. Yet the factors that boun…
Where do LLMs Fall Short in CBT-Guided Affective Reasoning?
Cognitive Behavioral Therapy (CBT) provides a structured framework for understanding a user's mental state by examining the interaction bet…
SPLIT: Training-Free AI-Generated and Partially Edited Video Detection via Spatial Patch-Level Incoherence and Temporal Roughness
Deploying AI-generated video detectors in real-world services demands an ultra-low false positive rate (FPR) on real videos to avoid falsel…
PPE-Bench: A Benchmark for Evaluating MLLM Unlearning under Private-Public Entanglement
Multimodal Large Language Models (MLLMs) have shown strong capabilities, but they may memorize private information from web data, raising p…
TIER: Trajectory-Invariant Explanation Regularization for Membership Privacy
Explainability is central to building trustworthy AI, yet explanation interfaces can inadvertently provide adversaries with an expanded pri…
Learning Taxonomic Trees with Hierarchical Representation Regularization for Large Multimodal Models
Taxonomies provide key information about the semantic relationships between concepts and the inherent organization of vision and language.…
CoACT: Action-Preserving Observation Compression for Coding Agents
LLM-based coding agents solve software-engineering tasks through iterative interactions with development environments, where returned obser…
Bootstrap Flow-Map Tree Sampling Enables Online Feedback Driven Search
In many scientific and engineering domains, maximizing discovery within a limited sampling budget demands strategic, observation-guided exp…
Harmonic-Aware Transformer for Real-Time Catheter Localization in Interventional Procedures of Magnetic Particle Imaging
Magnetic particle imaging (MPI) enables real-time, radiation-free tracking of magnetic nanoparticle-coated instruments, making it highly su…
R3D: Quantitative 3D Spatial Reasoning for Egocentric Wearables
Quantitative 3D spatial reasoning from egocentric RGB-D video is a critical capability for next-generation wearable assistants. Yet existin…
VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning
Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep…
Modeling the Impact of Visual Brand Language on Attention, Object Recognition, and Memory Retrieval
Visual brand language is the set of visual properties that convey brand identity for a product. What is the impact of visual brand language…
PromptPET: Privacy-Utility Optimized Prompt Obfuscation
Privacy is an important challenge when users interact with AI chatbots, since users may share sensitive information, explicitly or implicit…
MatPhaseBench: A Semantics-Guided Benchmark for Materials Phase Diagrams Understanding
Materials phase diagrams are a core knowledge representation in materials science, encoding temperature,composition, phase stability, and p…
A Precedent-Guided Co-Scientist for Side-Effect-Aware Drug Redesign
We propose PRECEDE, a precedent-guided co-scientist for side-effect-aware drug redesign that revises a parent compound to mitigate a specif…
Pooling-Based Context Modeling for Convolution-Free Deep Image Prior
Convolutional Neural Networks (CNNs) achieve strong denoising performance by exploiting spatial context from neighboring pixels. Deep Image…
The Foreign Policy AI Evaluation Gap
We argue that AI systems used in conducting foreign policy tasks - broadly enacting 'statecraft' - should be a priority test case for techn…
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning
Dense video captioning aims to generate temporally grounded descriptions of video events, benefiting both event-level video understanding a…
Individual Parameters in Weight-Sparse Transformers Appear Interpretable
A central goal of mechanistic interpretability is to understand how neural networks work and what each individual component does. Dominant…
Hierarchical Sparse Attention Done Right: Toward Infinite Context Modeling
Scaling modern large language models (LLMs) to long contexts is limited by the quadratic computation cost, and poor length extrapolation of…
Enhanced Feature Extraction for IoT Network Intrusion Detection Using GNNs and KAN
Recent advancements in the Internet of Things (IoT) emphasize the urgent need for advanced network security, as IoT networks feature dynami…
VISTA: Auditing Semantic Divergence in Vision-Language Models
Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or…
CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training
Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples tha…
PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark
Text-rich image models can now design poster-scale layouts, but we lack ways to measure whether they honor scientific communication contrac…
Back to Basics: Improving Molecular Understanding in LLMs via SMILES-Graph Translation
Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet thes…
Can Model Merging Improve Aggregation in DiLoCo?
Model merging techniques, which aggregate independently finetuned models into one to combine their capabilities, have become a topic of sig…
HyperVAttention: Efficient Sparse Attention with Spatio-Temporal Clustering for Video Diffusion
Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation. However, their ability t…
MambaLIE: Scene Light Intensity-Boosted Low-Light Image Enhancement with State Space Model
Images captured by consumer electronic devices, such as mobile phones and digital cameras, often suffer from low-light degradation due to s…
OmniFocus: Query-Guided Modality-Balanced Token Compression for Omni-Modal Large Language Models
Omni modal large language models (OmniLLMs) have attracted wide attention for their ability to jointly process audio and video, but they ge…
LACE-SVD: Loss-Aware SVD with Cumulative Error Correction for LLM Compression
The rapid growth in the parameter scale of large language models (LLMs) has created a strong demand for efficient compression techniques. A…
Spectral Rewiring for Exploration, Purification, and Model Merging
Reinforcement learning has become a standard post-training recipe for large language models, but dense full-parameter updates create two de…
PixCon: Clean-Positive Contrastive Learning for Foundation-Model Semi-Supervised Segmentation
Semi-supervised semantic segmentation (SSSS) has long turned on one question, which pseudo-labels to trust, and answered it with ever more…
Attention-Guided Efficientnet Architecture For Precise Criminal Identification in Surveillance Images
Criminal identification from surveillance imagery has become a critical research area in intelligent forensic surveillance systems due to t…
STELLA: Efficient Sensor-to-LLM Translation for On-Device Human Activity Recognition
HAR is increasingly expected to run continuously on edge devices, yet recent LLM-based methods remain hard to deploy: raw sensor prompts ar…
Don't Wait to Reply: Towards Responsive yet Thoughtful Dialogue through Proactive Thinking
Thinking has emerged as a critical capability for Large Language Models (LLMs) tackling complex tasks. However, its reactive nature, where…
Flow-A11y: Flow-Aware Accessibility Testing
Modern web applications increasingly expose accessibility barriers through interaction flows rather than static page snapshots. Keyboard tr…
SNR-Adaptive Unified Diffusion for Multi-Task Medical Image Segmentation
Clinical cardiac imaging pipelines currently deploy separate models for each dataset and modality, incurring redundant training costs and p…
ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy
Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards st…
A Multi-Task Deep Learning Framework for Real-Time Intelligent Video Surveillance with Temporal Event Validation
Modern video surveillance systems generate far more video streams than human operators can effectively monitor, making automated analysis e…
Detecting Architectural Drift in Safety-Critical Firmware through Runtime Trace Analysis
Maintaining consistency between architectural design and runtime-observed behavior is challenging in long-lived safety-critical firmware. T…
Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning
Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified,…
CuBAS: Information Geometric Curvature-Based Adaptive Sampling for Supervised Classification
The informativeness of a training set is as consequential as its size, yet most sampling strategies remain agnostic to the intrinsic geomet…
Rethinking Neural Nonlinearity as Gating
Activation functions are considered an essential primitive for neural nonlinearity, i.e., they enable neural networks to serve as universal…
Conditional Diffusion Guided Knowledge Transfer for Multi-Domain Knowledge Graph Completion
Multi-domain knowledge graph completion (MKGC) aims to improve missing triple prediction in a target KG by transferring knowledge from othe…
Which Algorithm Specification Formats Help Language Models Implement Machine Learning Algorithms?
Large language models (LLMs) are increasingly used to implement algorithms from research manuscripts, but papers often leave implementation…
The Role of Prompt Language and Translation-Theory-Driven Prompts in Large Language Models: A Case Study on Spanish-Chinese Journalistic Translation
This study examines how prompt language and translation theory-driven prompt design influence the quality of Spanish-Chinese journalistic t…
KARMA: Knowledge graph-based Automated Reasoning Materialization and Alignment
Template-based contrastive synthesis is scalable, but its candidates often differ only in a few entity-slots while sequence-level optimizat…
Decentralised Federated Learning over Temporal Networks: The Role of Heterogeneities
Decentralised federated learning, based on peer-to-peer communication, is increasingly proposed for on-device training of machine learning…
Effectiveness of LLM-based Software Diversity for Reliability Improvement -- an Empirical Study
Software diversity has been extensively studied as a means of reducing the risk of common-mode failures. Classic work showed that the centr…
CRRL: A Causality-Based Reinforcement Learning Framework for Autonomous System Recovery
Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failu…
Teaming Up with AI: Coordination and Cooperation
Successful diffusion of AI in the workforce hinges on the economic value that AI brings to human endeavors. Bringing AI into the workforce…
AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning
Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into exe…
Scalable Maximal Frequent Episode Mining with Desbordante
Episode mining aims to extract subsequences of events that possess certain distinctive properties and constitute facts valuable to the user…
A Bayesian Framework for Evaluating Scenario Compatibility in Generative Population Synthesis
Scenario-based transportation analysis specifies future assumptions through aggregate population targets, whereas generative population syn…
Self-Specializing Vision-Language Transmon Chip Calibration in a Physics-Grounded Environment
Calibrating a superconducting transmon chip is a sequential decision problem under noise, drift, and a finite budget: an expert must choose…
Transition Information Density: Morphological Trajectories, Synesthetic Perception, and Structured Interpolation in Neural Training (or: The Synesthetic AI)
Standard machine learning training presents data as discrete endpoint pairs, omitting the structure of the space between them. This paper i…
OpenGlass: A Sensing-Computing Split Architecture for Local MLLM-Driven Real-Time Visual Assistance
We present OpenGlass, an open-source, privacy-oriented, local-first system for low-latency multimodal visual assistance, with a primary foc…
Builder, Defender, Breaker: The Case Against Removing the Human from the AI-Driven Security Lifecycle
Artificial intelligence has spread across the whole of the security lifecycle. The same family of models now writes application code, harde…
CONTRA: Red-Teaming Configurations of Personalizable Agents
Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agent…
Agentic and Generative AI for Open-Source Intelligence and Cyber Investigations: Taxonomy, Evaluation, Challenges, and Future Directions
The rapid growth of publicly available digital information has rendered manual open-source intelligence (OSINT) analysis insufficient for m…
Unbiased Alignment for Large Language Models with Noisy Preferences
The alignment of large language models with human preferences is commonly achieved through Reinforcement Learning from Human Feedback or Di…
Semantic Segmentation-Driven Image-Level Diagnosis of Liver Cancers in Hematoxylin and Eosin Histopathology Images
As hematoxylin & eosin (H&E) staining constitutes the primary entry point in routine diagnostic workflows, computer-aided diagnosis from wh…
A harmonised dataset for Earth system foundation models
Foundation models for Earth systems have so far been trained primarily on physical climate and weather data, with limited representation of…
Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild
Agentic code review, where autonomous agents provide code review comments on pull requests, is increasingly integrated into development wor…
Hierarchical Multi-Agent Reinforcement Learning for Carbon-Aware AI Data Centers in Power Distribution Systems
Eco-friendly energy management for artificial intelligence data centers (AIDCs) is crucial because of the significant increase in energy co…
From Judgments to Issues: Structured Extraction of Legal Reasoning with Citation-Hallucination Control
We present an automated pipeline that decomposes Italian tax-court judgments into individual legal issues and extracts, for each issue, a s…
SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference
LLM agents are becoming a common interface for research, coding, and question answering, yet their Thought-Action-Observation loop is often…
FedAvg for HAR: Exploring the Tradeoff Between Personalized and Generalization Accuracy
The federated learning (FL) paradigm fosters distributed pervasive computing combined with artificial intelligence techniques, allowing for…
Efficient Decentralized Multi-task Dataset Valuation via Model Merging
Accurate and efficient dataset valuation is essential for enabling fair and transparent data marketplaces, especially when multiple contrib…
PedestrianDiffusion: Multimodal Generative Denoising and Dense State Estimation for Inertial Navigation
The accuracy of consumer-grade inertial navigation is bottlenecked by the stochastic noise of Micro-Electro-Mechanical Systems (MEMS). Trad…
LLM-Enhanced Hierarchical Heterogeneous Graph Representation Learning for Malicious Python Package Detection
Malicious Python packages have become a major threat to software supply chain ecosystems due to the widespread adoption of open-source repo…
Brand-as-Memory: Vision-Language Models Encode Causal, Mechanistically Localizable Credibility Priors for News Sources
Vision-language models (VLMs) increasingly read news and web content as images, where the publisher's identity is visually present. We show…
Spectral Signatures of Large Language Models
The rapidly growing repository of publicly available large language models (LLMs) presents significant challenges for systematic management…
The S-ICDF Dataset: Sionna-Simulated Dynamic Interference Characterization and Direction Finding
Jamming and spoofing threaten wireless and satellite navigation by disrupting or manipulating radio frequency (RF) signals, undermining ava…
DETECT-3B-Omni is Agnostic of Content and Demographics
A trustworthy and GDPR-compliant deepfake audio detector must base its decisions on acoustic artifacts, not on what is being said or who is…
Securing Multi-Tool AI Agent Chains With Dynamic, Real-Time Compositional Policies
Modern AI agent implementations such as frontier coding agents chain multiple tools at runtime that create a security surface that per-tool…
Amortising Bayesian Experimental Design for Sequential Information Gathering in LLMs
Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectivel…
No Time Like the Present: Agentic Test-Time Training for LLM Agents
LLM agents often degrade over long episodes: as trajectories grow, they revisit explored states, repeat failed actions, and lose strategies…
TRIAGE: Trustworthy Retrieval Instrumentation And Graph Evaluation
Knowledge graphs (KGs) that underpin Graph-based Retrieval-Augmented Generation (Graph-RAG) are increasingly built automatically by LLM-dri…
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control
Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term m…
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe
While skill optimization for autonomous agents has gained traction, existing methods rely on complex pipelines. This leaves a fundamental q…
Best-of-Better-$N$: Generating Pre-Aligned Responses with In-Context Learning
Inference-time alignment methods, such as Best-of-$N$, offer a flexible alternative to training-based alignment by using reward models to s…
CaresAI at SMM4H-HeaRD 2026: Predicting TNM Staging
This study aims to predict Tumor, Node, and Metastasis (TNM) stage labels independently, with the Cancer Genome Atlas (TCGA) pathology repo…
Towards Diverse and Comprehensive Benchmarks for Mutual Information Estimation
Mutual information (MI) estimation is a central problem in machine learning and statistics; however, existing benchmarks typically evaluate…
STRATOS: Bridging the Symbolic-to-Numeric Gap in Spatio-Temporal Text-to-SQL for Meteorological Data
Copernicus, the European Union's Earth observation program, produces petabytes of Earth observation and climate data, offering immense pote…
Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers w…
CAGE-1: Control, Assurance, and Governance Evaluation for Enterprise Agentic AI
Enterprise artificial intelligence is moving from experimentation into operational workflows. Early programs focused on model access and re…
AGL-1: The Enterprise AI Governance Layer as a Control Plane for Trusted Enterprise Intelligence
Enterprise artificial intelligence is moving from isolated experimentation toward operational dependency across copilots, retrieval-augment…
Aligning Language Models with Selective Prediction
Large language models (LLMs) are increasingly deployed as critical decision-making components in high-stakes real-world AI systems, renderi…
Latent Clarity: Bridging World-Model Kinematics to Semantic Manifolds for Video Anomaly Anticipation
Continuous video anomaly detection is dominated by reactive Multiple Instance Learning (MIL) that collapses spatiotemporal features into sc…
EPRA U-Net: An Efficient Pyramid Residual Attention Framework for Accurate Infarct Segmentation in Diffusion-Weighted MRI
Objective: Accurate identification of acute ischemic infarcts on diffusion-weighted magnetic resonance imaging (DWI) is a critical prerequi…
Teacher Supervision over Representation Equivalence Classes
Knowledge distillation is usually framed as a choice of what to match in the teacher - its logits, hidden features, or sample relations - w…
Differentiate the Evaluator, Not the Program: An Efficient Runtime Representation for Neuro-Symbolic Learning
AI systems increasingly propose executable scientific models whose value depends on both their symbolic structure and their fitted continuo…
PLGSA-Transformer: Periocular Landmark-Guided Attention with Occlusion-Adaptive Cosine Thresholding for Cross-Modal Masked and Unmasked Face Recognition
The widespread adoption of facial masks, accelerated by COVID-19 and mandated in security-sensitive settings, has exposed limitations of co…
Responsibility Distribution Estimation in Ego-View Accident Videos with Multimodal Large Language Models
Recent studies on multimodal traffic accident understanding have mainly relied on infrastructure-camera footage, satellite imagery, or stru…
An Interpretable Deep Learning Framework for Discovery and Clinical Validation of Deep Radiomic Signatures in Tumor Classification
Imaging signatures are quantitative features extracted from medical images that provide clinically meaningful information for tumor diagnos…
Token-Based Affordance Grounding with Large Vision-Language Models
Affordance grounding aims to localize image regions that support a specific action, serving as a core capability for physical intelligence…
They Infer What You Meant: Models Represent Communicative Intent More Reliably Than They Act On It
When a person shares something with a language model, the model often answers the surface of the message rather than what the sender was do…
A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning
Adversarial robustness in Unsupervised Domain Adaptation (UDA) remains a significant challenge due to noisy pseudo labels and inherent dist…
RADIO1D: Elastic Representations for Condensed Vision Modeling
This paper challenges the assumption that vision-language models (VLMs) require fixed patch-based 2D vision features. Analyzing fine-tuned…
An AI-Assisted Solution to the Signed BAR Conjecture: Uniqueness in the Harrison--Reiman Class and a Completely-$\mathcal{S}$ Class Obstruction
For a multidimensional reflected diffusion, determining whether the associated basic adjoint relationship (BAR) uniquely characterizes the…
Revealing Hidden Model Behaviors with Task-Specific Self-Reports
Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only wh…
Moonstone: A Multimodal Foundation Model and Benchmark for Lunar Remote Sensing
Decades of orbital missions have produced multi-modal remote sensing data for the Moon, spanning optical imagery, spectroscopy, thermal emi…
ClinOCR-Bench: A Comprehensive Clinical Scanned Document Dataset for Optical Character Recognition Model Evaluation
Extracting textual information from scanned medical documents, such as external laboratory reports and manually filled forms, has been a ma…
ELiTeFormer: An Efficient Transformer for FPGAs
Transformer blocks are prevalent in large language model (LLM) but present deployment challenges due to their challenging computational and…
AutoCedar: An Agentic Framework for Verifier-Guided Access Control Policy Synthesis
Large Language Models are increasingly used to turn natural-language requirements into code. In access control, that shortcut is dangerous:…
ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation
Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoidi…
A Fair Benchmarking of Deep Relational Database Learning Models
Relational databases (RDBs) are the primary data infrastructure in many enterprises, yet recent deep learning methods designed for RDBs hav…
Phase-Preserving Trimodal Transformer for Tropical Forest Biomass Estimation Using Optical and PolInSAR Data
The accurate estimation of Above-Ground Biomass (AGB) in mature tropical forests remains a critical challenge in remote sensing, primarily…
Don't Blame the Large Language Model: How Scaffolding Evolution Shapes Coding Agent Quality
Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agentic scaffolding:…
OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-…
CoGen3D: An Agentic Human-AI Co-Design Pipeline for 3D Asset Generation for Virtual Reality
Creating 3D assets for virtual reality requires modeling expertise, which restricts the authorship of immersive experiences. Existing gener…
Attending to Multimodal Generation One Token at a Time
Multimodal large language models (MLLMs) generate responses autoregressively, integrating visual and linguistic information in an evolving…
A Failure-Mode Benchmark for Polymorphic Sybil Poisoning in RAG
We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning. The framework par…
EmCom-Diffusion: Probing Visual Reflection in Emergent Languages via Image Generation
Measuring the extent to which emergent languages encode the visual content of their inputs is an open problem. We refer to this property as…
FedACT: Federated Adaptive Coordinate Trust Modulation for Robust Transformer Training under Data Heterogeneity
Federated Transformer training increasingly relies on local AdamW, whose adaptive updates can provide much stronger local progress than SGD…
Self-Improving Diffusion Classifiers with Minority Preference Optimization
Prior studies have demonstrated that diffusion classifiers achieve robust zero-shot classification performance. However, their effectivenes…
SkillFab: An Agent-Native Skill Production Platform
SkillFab is an agent-native platform for turning missing capabilities into reviewed, reusable Agent Skills. At runtime, agents first search…
Rethinking Depth Pruning for Vision Transformers: A Heterogeneity-Aware Perspective
While prior studies have successfully compressed vision Transformers (ViTs) through various pruning techniques, most have concentrated on w…
Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks
Symmetry is everywhere in nature and society. Geometric deep learning exploits symmetries in data to improve the performance and efficiency…
Punching Above Their Weight: Classification-Head Fine-Tuning of Tiny Language Models (TLMs) for Verifiable Multiple-Choice Tasks
We define Tiny Language Models (TLMs) as models below roughly 3B parameters that fit on mainstream consumer devices. We study how to adapt…
CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation
The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dol…
Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings
Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representatio…
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on s…
DualView: Preventing Indirect Prompt Injection in Personal AI Agents
Personal AI agents that run on the user's local machine, such as OpenClaw, automate daily tasks including web search, email, and file manag…
Q-TriM: Question-Guided Tri-Modal Attention for Audio-Visual Question Answering
Audio-Visual Question Answering (AVQA) extends classical VQA by requiring joint reasoning over video and synchronized audio. However, many…
How Do Diffusion Classifiers Decide? A Bias-Centric Evaluation
Diffusion models have recently been repurposed for zero-shot classification, giving rise to diffusion classifiers that identify the best-ma…
Beyond Static Rules: Automated Discovery of Latent Vulnerabilities in Text-to-SQL
While Large Language Models (LLMs) have achieved remarkable success in Text-to-SQL tasks, their deployment in real-world environments is hi…
When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts
Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresse…
High-Fidelity One-Step Generative Visuomotor Policy via Recursive Correction, Frequency Consistency, and Contrastive Flow Matching
Generative models such as diffusion and flow matching have advanced robotic visuomotor policies by modeling multimodal action distributions…
GeoSelect: Spatial-Program Execution for Training-Free Referring Remote Sensing Image Segmentation
Referring remote sensing image segmentation isolates the object named by a natural-language expression in an aerial image. Existing trainin…
Next-Gen Sponsored Search: Crafting the Perfect Query with Inventory-Aware RAG (InvAwr-RAG) Based GenAI
Sponsored search plays a crucial role in e-commerce revenue generation, where advertisers strategically bid on keywords to capture the atte…
Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate…
Enhancement of E-commerce Sponsored Search Relevancy with LLM
Sponsored search plays a crucial role as a revenue stream for search engines, wherein advertisers competitively bid on keywords that align…
Advanced Topic Modeling Techniques for Categorizing Software Vulnerabilities
The increasing complexity and frequency of software vulnerabilities demand efficient methods to analyze and prioritize threats. Traditional…
TabQueryBench: A Query-Centric Benchmark for Synthetic Tabular Data
Synthetic tabular data support use cases like data sharing, model development under access restrictions, and rapid prototyping of analytica…
TokAN: Accent Normalization Using Self-Supervised Speech Tokens
Accent normalization (AN) seeks to convert non-native (L2) accented speech into standard (L1) speech while preserving speaker identity. The…
Probe, Don't Prompt: A Hidden-State Probe for Metadata Filtering in Multi-Meta-RAG
Multi-Meta-RAG improves retrieval for multi-hop question answering by filtering a vector store on metadata (the news source) that it extrac…
MPSelectTune: Prompt-type Selection for Fine-tuning improves Concept Unlearning in LLMs
LLMs can be conveniently adapted to a diverse set of tasks, e.g, prediction, question-answering tasks, etc, using appropriate prompts with…
Why3-py: A Tool for Formal Verification of Hypothesis Testing and Meta-Analysis in Python
The reproducibility crisis in scientific research has received widespread recognition, thereby increasing the importance of meta-analyses t…
The Remarkable Effectiveness of Providing AI Agents with Natural Language Tools: A Replication Study Validating NLT Performance Across 14 Models
This study independently replicates and extends the Natural Language Tools (NLT) framework of Johnson et al.~(2025), which questions the us…
NormWorlds-CF: Solver-Verified Counterfactual Normative Reasoning with Metamorphic-Relation GRPO
Language models can reach the right normative verdict for the wrong reason. We introduce NormWorlds-CF, a solver-verified environment for c…
Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control
World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulat…
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents
Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, an…
Order-based Causal Discovery for Multistage Processes
Causality has become an increasingly important tool for gaining a deeper understanding of complex systems. Among various causal analysis me…
Scalable Semantic Steering of Embedding Projections
Low-dimensional projections support interactive visual analysis of high-dimensional data embeddings, but their structure often does not ali…
BanglaMemeEvidence: A Multimodal Benchmark Dataset for Explanatory Evidence Detection in Bengali Memes
Memes have become influential communication tools on social media, combining viral visuals with concise messaging to convey impactful ideas…
NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating…
Full Glyph Images Beat Token Embeddings: A Controlled Study for Transformers
Modern language models generally represent text as sequences of discrete token embeddings, an assumption deeply rooted in current practice…
Separating Representation from Reconstruction Enables Scalable Text Encoders
While decoders have rapidly scaled, encoders have remained largely unchanged since BERT. We revisit this disparity by frozen backbone evalu…
When Does Small Data Work? Accuracy and Efficiency Trade-offs Between Tabular Foundation Models and Conventional Methods for Crowd-State Classification at Hajj and Umrah
Learning from few labeled examples is a central challenge in tabular machine learning, and it becomes the binding constraint in domains whe…
Finite Reliability Representations: Noise-Calibrated Belief-Space Covers for Reliable Decision-Making
Physical sensing and actuation noise floors should inform how much belief resolution a decision-making system can reliably use. We introduc…
A Unified Algebraic Framework for Classification Performance Evaluation
We propose a unified algebraic framework for classification performance evaluation that encompasses binary, multiclass, multilabel, ordinal…
Efficient Discovery of Conditional Dependencies with Desbordante
Conditional functional dependencies (CFDs) are functional dependencies with a restricted scope: they specify the context in which a depende…
OmniOpt: Taxonomy, Geometry, and Benchmarking of Modern Optimizers
Optimizer selection for large-scale model training has become a system-level design decision constrained jointly by compute, memory, tuning…
The "I Don't Know" Filter: Enhancing Agentic Reliability in Function Calling
The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks. However, the metrics used in…
Reward-Gated On-Policy Distillation
On-policy distillation is a powerful way to transfer reasoning ability from a strong teacher to a smaller student: the student samples traj…
Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability
Distinguishing Large Language Model (LLM) generated text from human writing is a critical and difficult challenge. While LLMs are trained t…
Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
Unsupervised syllabic tokenization aims to learn discrete syllabic tokens that capture latent linguistic content-related structure from raw…
Enhancing Implicit Neural Representations with Image Feature Embedding for Unsupervised Cardiac Cine MRI Reconstruction
Cardiac cine Magnetic Resonance Imaging (MRI) is a critical diagnostic tool that provides dynamic insights for radiologists. To accelerate…
Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders
Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a…
Benchmarking API Drift in LLM-Generated Quantum Code Across Successive SDK Versions
Large language models can generate plausible quantum code, but it is unclear whether they can reliably target the specific software develop…
Seeing Once is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering
Recent Multi-modal Large Language Models (MLLMs) have demonstrated remarkable performance on 2D question answering tasks. However, extendin…
FedSPM: Routing-Enabled Federated Learning under Dual Heterogeneity via Semiparametric Mixture
Routing-prediction federated learning has emerged as a new paradigm that reframes inter-client heterogeneity as a resource for system-level…
Submitted and Diagnostic Analysis of Full-Text Temporal Retrieval for LongEval-Sci
LongEval-Sci evaluates scientific retrieval under collection change, where a system should be effective on the current corpus and remain us…
DynaVieW: Schema-Guided World Modeling for Understanding Hierarchical Visual Dynamics
Multimodal LLMs struggle to systematically model the temporal evolution of visual scenes in videos or multi-image sequences. Such inputs re…
Parametric Memory Decoding for Zero-Shot Routing in LoRA-Based External Parametric Memory
With the rise of parametric memory, LoRA-based External Parametric Memory (EPM) has emerged as a modular solution, but existing routing met…
SOV-CAD: Stepwise Orthographic Views Guided CAD Modeling Sequence Reconstruction
Reconstructing Computer-Aided Design (CAD) modeling sequences from images is crucial for preserving design intent and supporting parametric…
Conflict-Based Lazy Search for Fast Multi-Manipulator Planning
Employing multiple manipulators can boost efficiency and accomplish tasks that a single manipulator cannot do. However, real-time planning…
CSB: A Counting and Sampling tool for Bit-vectors
Satisfiability modulo theory (SMT) solvers have significantly advanced automated reasoning due to their effectiveness in solving problems a…
!Imperio, smolVLA: The Implications of Data Poisoning on Open Source Robotics
This work establishes that trigger-word data poisoning of vision language action models is practical, while at the same time the open-sourc…
HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding
Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for L…
Mask-based Predictive Representations for Reinforcement Learning
Vision-based deep reinforcement learning involves dealing with high-dimensional inputs of image information. It is crucial to abstract effe…
Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese)
This paper explains the principles and provides examples of a new method for distinguishing between FAKE human speech synthesized by genera…
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering
Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual qu…
Piercing Gilbreath's Conjecture: From Deep Number Theory Insights to Fintech and Cybersecurity
I propose a new methodology to attack the fascinating Gilbreath's conjecture about prime numbers, first posted in 1878 and unsolved to this…
CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving
End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often…
Detecting Hallucinations in Retrieval-Augmented Generation through Grounding-Aware Sensitivity by Perturbation (GASP)
Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level scor…
SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects
Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction,…
Hierarchical Multi-to-Single-Modal Knowledge Distillation for Disruption Prediction in EAST
Plasma disruption is a critical threat to tokamak safety. Existing data-driven predictors mainly rely on time-series diagnostic signals, wh…
Signal or Noise? Understanding Generative Models for Real-World Sensor Time Series
Generative models have changed how machine learning represents complex data distributions, especially in language and vision, yet many real…
HALO-WA: Hybrid-Attention Latent-Guided Online Reinforcement Learning for World-Action Models
World-action (WA) models can generate long-horizon action chunks for general-purpose robotic manipulation, but they remain vulnerable to ca…
LBR: Towards Mitigating Length Bias in Large Language Models for Recommendation
Large language models (LLMs) have recently emerged as powerful backbones for recommender systems by reformulating recommendation as a token…
Self-Reference in Large Language Models: The Introspection Threshold for Recursive Self-Improvement
The pursuit of self-evolving AI raises a critical question: when is autonomous self-improvement sustainable rather than degenerative? Drawi…
Risk-Constrained Freshness-Aware Semantic Caching for Open-Web Retrieval-Augmented LLMs
Semantic caching reduces the latency and cost of retrieval-augmented generation (RAG) by serving cached answers to semantically similar que…
Agentic-V2X: Small Language Model Agents for Deadline-Aware V2X Scheduling in 5G/6G Networks
Large Language Models (LLMs) are proposed as control interfaces for next-generation networks, but their latency, hallucinations, and lack o…
CausalGame: Benchmarking Causal Thinking of LLM Agents in Games
Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention. Since scientific discovery fundame…
HiFA4: Training-Free 4-bit FlashAttention on Ascend HIF4 NPUs for LLM Inference
We present HiFA4, a post-training operator-level design that executes both QK^T and PV in FlashAttention as 4-bit HIF4 Cube GEMMs for LLM i…
Fixed-Confidence Best-Arm Identification for Causal Mediation Analysis
This paper studies the problem of identifying the treatment that maximizes the expected natural direct potential outcome (NDPO), which capt…
One Framework for All: Cross-Modal Membership Inference for Generative Models
Large generative models across text-to-text, text-to-image, and image-to-text modalities have been shown to pose significant privacy risks.…
IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation
While Large Vision-Language Models (VLMs) demonstrate remarkable generic capabilities, their clinical reasoning in specialized domains like…
HASSL: Hierarchy-Aware Self-Supervised Learning Framework for Single Cell Microscopy
Hierarchical structure is common in image data, where fine-grained clusters often merge into larger, coarser semantic groups. In biological…
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound…
Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention
Recent NVFP4 pretraining methods mainly target transformer linear layers, leaving optimizer states, optimizer arithmetic and attention unde…
Transferability Between Understanding and Generation in Unified Multimodal Models
Unified Multimodal Models (UMMs) integrate image understanding and generation within a single architecture, yet how the two tasks interact…
UI-MOPD: Multi-Platform On-Policy Distillation for Continual GUI Agent Learning
Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-p…
dOPSD: On-Policy Self-Distillation for Diffusion Language Models
Diffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autore…
evalci: A Python Library for Statistically Rigorous Comparison of Language Model Evaluations
The dominant practice in language model evaluation is to report a single accuracy number per model and declare the higher one better, witho…
On Pairwise Quantile Regression -- Statistical Guarantees and Applications
Quantile regression provides a powerful tool for summarizing the conditional distribution of a real valued random variable (r.v.) of intere…
Covert Trait Propagation Is Representation Alignment: Mechanistic Evidence from Hidden-Channel Distillation
A student model trained on pure uniform noise can still inherit its teacher's digit-classification ability, provided the two share initiali…
RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies
Generalist robot manipulation policies have advanced rapidly, yet existing benchmarks remain limited in systematically evaluating their cap…
A Retrieval-Augmented Framework for Detecting and Resolving Pragmatic Ambiguities in Natural Language Requirements
Natural language requirements (NLRs) are essential for bridging communication gaps among diverse stakeholders in software development. Howe…
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Research dissemination, turning a paper into a poster, a talk video, and a blog post, is still a manual last mile. Prior automation treats…
Generative wave propagator
Seismic wavefield simulation is fundamental to seismology, but conventional finite-difference (FD) methods remain limited by numerical disp…
Wan-Streamer v0.2: Higher Resolution, Same Latency
We present Wan-Streamer v0.2, a latency-preserving upgrade of the native-streaming, end-to-end audio-visual interaction model. v0.2 keeps t…
From Regulation to Requirements: An Automated Requirement Derivation and Explanation Pipeline
Ensuring software compliance with regulations such as the General Data Protection Regulation (GDPR) and the Artificial Intelligence Act (EU…
A Deep Learning-based surrogate model for Severe Accidents in nuclear reactors using ASTEC
Integral codes like the Accident Source Term Evaluation Code (ASTEC) are powerful tools to study the physics of Severe Accidents (SAs) in n…
Robustness Verification of an Autonomous Underwater Vehicle-based Plankton Classifier
The assessment of planktonic standing stocks and microorganism structures is critical for understanding upper ocean biological processes. C…
Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models
World-model evaluation for model-based reinforcement learning typically asks whether the learned model predicts reward and value well, whic…
Regime-Conditional Stabilisation of LLM-Augmented Cooperative Multi-Agent Reinforcement Learning
Large Language Models (LLMs) offer a natural interface for translating human objectives into reward signals for cooperative multi-agent rei…
PulmoSight-XAI: An Explainable Multi-View Attention Ensemble with Gradient Boosting Meta-Learning for Multi-Label Chest X-Ray Classification
Automated chest X-ray classification remains challenging due to severe class imbalance, co-occurring pathologies, and the loss of localized…
Two Black Boxes, One Solver: Encoder Probing and Decoder Attribution for Neural Multi-Attribute VRP under Hard-Mask and Recourse Decoders
Neural autoregressive solvers for the Multi-Attribute Vehicle Routing Problem (MAVRP) reach competitive cost but offer no per-step justific…
Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5
Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Q…
Failures and Successes to Learn a Core Conceptual Distinction from the Statistics of Language
Generic statements like "tigers are striped" and "cars have radios" communicate information that is, in general, true. However, while the f…
Language Models Represent and Transform Concepts with Shared Geometry
How concepts are represented in neural networks is a fundamental question in machine learning. The dominant view treats concept representat…
Training-Free Model Selection and Domain-Aware Score Calibration for First-Shot Anomalous Sound Detection
First-shot anomalous sound detection in DCASE Challenge Task 2 must flag anomalies of unseen machine types with a single threshold, without…
Lyapunov-Guided Training for Hardware-Safe Neural Networks Under Fixed-Point Arithmetic
Low-precision neural networks are attractive for resource-constrained hardware, but fixed-point arithmetic introduces failure modes that ar…
Obey, Diverge, Collapse: Blind Obedience to Incorrect Instructions Drives Code LLMs to Irrecoverable Code Semantic Collapse
Code language models are now trusted collaborators in production workflows for debugging, refactoring, and iterative repair, and every benc…
CRISP: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining
Camera-radar (CR) fusion is a practical sensing configuration for autonomous driving, but existing models are typically trained with task-s…
Auto: The AGI Compiler
Every LLM agent run re-derives its behavior token by token on a frontier model: brilliant, expensive, slow, and unbounded. We present Auto,…
Mask2Real-WM: Segmentation Masks as a Sim-to-Real Bridge for Controllable Dexterous World Models
Action-conditioned world models allow robots to predict the future consequences of candidate actions without additional physical interactio…
Explainable Novel Category Discovery in Semantic Concept Space
Novel category discovery aims to identify unseen classes from unlabeled data by transferring knowledge from labeled categories, but most ex…
Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption
We present a bidirectional framework for estimating the energy consumption of text-to-video (T2V) and text-to-video-audio (T2VA) models fro…
Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations
Accurate prediction of patient-specific therapeutic response from pre-treatment transcriptomes is hindered by the scarcity of matched clini…
EEG-SpikeAgent: Agentic Closed-Loop Program Synthesis for Automated EEG Spike Detection
Automated detection of interictal epileptiform discharges in scalp electroencephalography (EEG) is clinically important, but recent high-pe…
A Few Teacher Steps Go a Long Way: Cost-Efficient On-Policy Data Augmentation for Agent Post-Training
For LLM agents, supervised fine-tuning is not only about teacher labels' quality, but also about which interaction contexts those labels co…
LLM-Driven CI-CD Workflow Intelligence for Cyber Systems Engineering
CI/CD workflows have become executable operational policy: they decide what gets built, tested, released, and deployed, and they mediate ho…
Simple-to-Complex Structured Demonstrations for Vision-Language-Action Learning
Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation by integrating visual perception, languag…
TORINO: Token Reduction via Interpretable Concept Overlap in Vision-Language Models
Vision-Language Models (VLMs) have demonstrated impressive capabilities across different tasks, but their computational cost is dominated b…
LCPNet: Latent Consistent Proximal Unfolding Network for Infrared Small Target Detection
Infrared small target detection (IRSTD) aims to identify long distance small targets from complex infrared backgrounds, and is a fundamenta…
Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-…
G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement
The rapid advancement of AI-generated videos poses increasing security risks and calls for robust detectors with strong cross-domain genera…
SILO: Simulation-in-the-Loop Sim-to-Real Transfer for Multi-Stage Cable Routing
Linear-deformable manipulation remains challenging due to the complex deformations of objects such as cables and ropes. Prior data-driven a…
Hierarchical Evidence-Driven Reasoning for Long Document Understanding
Retrieval-Augmented Generation (RAG) streamlines long-document understanding by leveraging retrieval mechanisms to restrict input images to…
Retroactive Chain-of-Thought (RetroCoT): Forensic Reconstruction Prompts as a Safety Diagnostic Across Model Generations
Safety alignment in large language models is typically evaluated against direct, imperative harmful requests. We show that this alignment i…
Machine Learning for Depression Screening and Intervention: an Original Circadian Rhythm Score-based Methodology
Depression screening from large-scale behavioral data is challenged by fragmented circadian indicators, limited interpretability, and the l…
Targeted Structure Completion for Sparse-View 3D Reconstruction in Autonomous Driving
Reconstructing 3D scene structures from sparse, low-overlap observations remains a fundamental challenge in autonomous driving. Recent stat…
Elastic Gang: Per-Token Membership Change for a Hard-Barriered LLM Inference Gang Co-Scheduled with OS Processes
On-device LLM decoding is a hard-barriered CPU-SIMD computation that wants every core for milliseconds per token, while the rest of the OS…
Do Vision-Language-Action Models Mean What They Say? On the Role of Faithfulness in Embodied Reasoning
Embodied Chain-of-Thought has emerged as a promising mechanism to enhance robot decision-making and interpretability in black-box Vision-La…
ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents
Tool calling is central to modern language model agents, but aggregate benchmark scores often hide where tool use fails. A model that never…
URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment
Synthesis planning aiming to find pathways of reactions for a target molecule is one of the most important and challenging tasks in drug di…
Strategic Buying Agents
Agentic AI is shifting online shopping from search toward delegated purchasing, where autonomous buying agents monitor markets and decide w…
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents
Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks. Howeve…
Geometry-Aware Motion Latents for Learning Robust Manipulation Policies
Learning motion latents for robotic manipulation heavily relies on extracting motion patterns from visual sequences, yet effective action a…
Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards
Automatic data visualization generation has advanced rapidly with multi-modal large language models, yet existing efforts largely focus on…
Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment
Reinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of "rollout then update", which ine…
RustMizan: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities
LLM agents are increasingly applied to vulnerability analysis, but existing benchmarks have not kept pace. They typically rely on small non…
Wasserstein Residuals: Learning Gradient Flows from Population Dynamics
Reconstructing population dynamics is a central problem in the physical and data sciences. Often, the dynamics are modeled as a Wasserstein…
Trust Region Policy Distillation
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), wh…
Multi-Turn On-Policy Distillation with Prefix Replay
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student…
Predicting Drafted Deck Strength for "Magic: the Gathering"
Many real-world games do not admit a fixed, compact rule set: instead, their dynamics are defined by interactions among a large and often e…
An Exploration of Agentic Information Fusion for Test Maintenance Prediction
Test maintenance is a critical, yet costly, activity - particularly as codebases rapidly evolve. To assist, we present MAST, a multi-agent…
Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection…
HamQASBench: A Hamiltonian-Informed Diagnostic Benchmark for Evaluating Quantum Architecture Search
Quantum Architecture Search (QAS) automates the design of parameterized quantum circuits for variational quantum algorithms, yet existing b…
Pretraining Curricula Enable Selective Fine-tuning
Transformers follow implicit curricula whereby some tasks are learned before others. However, how explicit pretraining curricula influence…
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
While audio deepfake detection has advanced significantly, representative detectors show limited generalization to synthetic sound effects.…
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, so high-…
Graph Representation Learning of Longitudinal Medical Imaging Trajectories for Treatment Response Prediction
In patients with breast cancer, pathological complete response (pCR) has been established as a clinically meaningful surrogate marker for l…
Efficient Perception in Automotive Detection and Tracking Using Neuromorphic Computing
Deep learning algorithms are notorious for their high carbon footprint and computational demands that limit their deployment on edge device…
Input Pathways Shape Few-Shot, Not Zero-Shot, Binding in Tiny Transformers: A Fully-Enumerable Study
How does the way information reaches a transformer -- as symbolic tokens, a clean per-factor "oracle" code, or an entangled perceptual vect…
DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as…
MemPose: Category-level Object Pose Estimation with Memory
In the pursuit of robust and generalizable category-level object pose estimation, most existing methods adopt parametric formulations that…
Multi-Robot Open Adaptive Teaming Across Unseen Environments, Partners, and Scales
Deploying robot teams in the real world requires simultaneous adaptation to unseen environments, unknown partners, and varying team sizes,…
Joint Velocity Slope Diffusion Prior for Structurally Constrained Velocity Model Building
High-resolution velocity models are crucial for reservoir characterization and subsurface delineation. However, the band limited nature of…
LLM for the development of FCM
This article is about the development of a fuzzy cognitive map using a local large language model. In the light of recent advances it is ev…
The Map Behind the Flow: Finite-Step Gradient Descent as a Dynamical System
Many phenomena of deep learning are dynamical: they concern not only which minima exist, but how gradient descent reaches, avoids, or selec…
TACTIC-KG: Toward Small Agent Teams for Cyber Threat Intelligence Knowledge Graph Construction
Cyber Threat Intelligence (CTI) reports are predominantly unstructured, heterogeneous, and noisy, which limits their direct usability for a…
Comparison of Loss Functions for Robust Deep Learning-based Echocardiography Segmentation when Learning with Partially Labelled Data from Multiple Domains
Echocardiography is the first imaging modality used for assessing cardiac function, and accurate segmentation of cardiac structures is esse…
ImputeECG: Deep Learning Reconstruction of Complete 12-Lead Electrocardiograms from Incomplete Recordings for Cardiac Assessment
Complete digital 12-lead electrocardiograms (ECGs) are essential for AI-enabled cardiovascular assessment, yet many clinical ECG records, p…
Hyperparameter Transfer in Graph Neural Networks
The performance of deep learning models crucially depends on the settings of hyperparameters like learning rate, initialization scale, and…
Your Agent's Memories Are Not Its Own: Forged Reasoning Attacks on LLM Agent Memory and Defenses
Persistent memory has enabled large language model (LLM) agents to store factual knowledge, prior decisions, reasoning histories, tool usag…
LLM-Based Test Oracles: Source-of-Authority Taxonomy -- A Systematic Literature Review
Large language models (LLMs) are increasingly used to produce test oracles, the part of a test that decides whether observed behavior is co…
RUFNet: Query-Guided Support Mask Refinement and Uncertainty Fusion based on Hybrid Mamba for Few-Shot Brain Tumor Segmentation
Few-shot brain tumor segmentation remains challenging due to noisy support masks, inter-patient variations between support and query images…
Beyond Independent Labels: Schwartz-Geometry Decoding for Human Value Detection
Human value detection is commonly formulated as sentence-level multi-label classification over the 19 refined Schwartz values, typically pr…
AIFS-SUBS: Extending Data-Driven Forecasting to Sub-Seasonal Timescales
Data-driven models now rival numerical weather prediction in the medium range, but extending them to sub-seasonal lead times raises challen…
Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters
Grokking -- the delayed onset of generalization long after a network has fit its training set - -is usually studied in models too large to…
Localized LoRA-MoE: Block-wise Low-Rank Experts With Adaptive Routing
Large Language Models (LLMs) and high-dimensional perception networks increasingly rely on parameter-efficient fine-tuning (PEFT) to adapt…
Agent Data Injection Attacks are Realistic Threats to AI Agents
AI agents act on behalf of user prompts, consuming external data and taking actions based on the agent context. Prior research on AI agent…
Three-Phase Evaluation of AI-Assisted Software Development Life Cycle
This paper presents an exploratory evaluation of how increasing levels of AI autonomy affect software development productivity, requirement…
PDEFlow: Autonomous Agentic PDE Pipelines for Neural Operator Learning and Solver-Free Inference
We present PDEFlow, an autonomous agentic framework that turns user-level ODE and PDE descriptions into solver-backed neural-operator pipel…
Open Problems in AI Incident Governance
AI systems may produce failures after deployment that pre-deployment safety assessments do not anticipate. Managing these failures requires…
Relational Multi-Agent Reinforcement Learning for Dynamic Pricing in High-Speed Railway Markets
In liberalised railway systems, operators must set prices dynamically in an environment with partial observability, as they retain private…
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistanc…
Unified Audio Intelligence Without Regressing on Text Intelligence
Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-…
Noisy-Channel Minimum Bayes Risk Decoding
Minimum Bayes Risk (MBR) decoding yields more robust and higher-quality text generation than maximum a posteriori (MAP) decoding by selecti…
Optimizing ML Workload Partitioning between CPUs and CIM Accelerators for Heterogeneous Computing
Computing-in-Memory (CIM) accelerators execute Matrix-Vector Multiplications (MVMs) in memory, making them a compelling solution for Machin…
CanniUplift: A Holistic Framework for Mitigating Seller and Incentive Cannibalization in E-commerce Uplift Modeling
Personalized incentive allocation is vital for e-commerce, where uplift modeling is the standard for estimating Individual Treatment Effect…
Privacy-Preserving Robustness Verification for Neural Networks
Neural network verification and data privacy are inherently in tension: verification demands full access to model parameters and input data…
Shifting from Discrete to Continuous Reference Data: QSM-Derived Horizontal Tree Biomass Distribution for Deep Learning Biomass Estimation
Conventional modeling approaches for LiDAR-based above-ground biomass (AGB) estimation rely on discrete plot-level inventory aggregates. Th…
Adaptive Inference Batching using Policy Gradients
Inference serving systems must balance throughput and latency under bursty, heterogeneous workloads, yet the industry standard remains stat…
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
Speaker embeddings, or x-vectors, are widely used to represent speaker identity and speaker-related attributes, but existing embedding extr…
Wavelet Scattering Transform for Interpretable Schizophrenia Biomarker Discovery and Classification from Resting-State EEG
Schizophrenia is a debilitating neuropsychiatric disorder characterized by profound cortical network dysregulation, for which objective, cl…
Air Quality Downscaling with Station-Guided Pseudo-Supervision
Super-resolving coarse atmospheric fields to local PM$_{2.5}$ variations is uniquely challenged by a mismatch in spatial support: while pix…
Topological Shape Representation for Aneurysm -- Bifurcation Detection
Automated detection of intracranial aneurysms (IAs) from CT angiography (CTA) is severely hindered by high false-positive rates. Convolutio…
Steering Optimisation Trajectories in Diffusion Representation Learning
We study why diffusion autoencoders can achieve similar image quality while learning substantially different latent structures. We trace th…
TREK: Distill to Explore, Reinforce to Refine
Group Relative Policy Optimization (GRPO) is effective when the current policy already samples useful reasoning trajectories, but it stalls…
Multiplayer Interactive World Models with Representation Autoencoders
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-pl…
Selective Disclosure Watermarking for Large Language Models
Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs). Existing approaches in…
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or…
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
Streaming speech-to-speech language models aim to answer spoken queries directly with synthetic speech. However, standard speech and text b…
GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks
For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable…
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon task…
What Does a Discrete Diffusion Model Learn?
What does a discrete diffusion model learn: a denoiser, a score ratio, or a bridge plug-in predictor? At the level of jump rates, these are…
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deep…
Interpretable Human-Label-Free Deep Learning for Real-Bogus Classification with Uncertainty Quantification
Time-domain surveys generate many transient candidates, making Real-Bogus classification a critical step in automated discovery pipelines.…
Weak-to-Strong Generalization via Direct On-Policy Distillation
Reinforcement learning with verifiable rewards (RLVR) is a powerful recipe for improving language-model reasoning, but it is expensive to r…
From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model
Real-world robot deployment rarely maintains the training-stage camera setup, where cameras often experience repositioning or remounting de…
DrugAgent: Reliable Multi-Agent Integration of Conflicting Biomedical Evidence for Drug-Target Interaction Assessment
Workflows in drug-target interaction (DTI) assessment require integrating heterogeneous data from predictive models, curated resources, and…
Neuro-symbolic Weak Supervision: Theory and Semantics
Weak supervision enables machine learning models to learn from limited or noisy labels, but it introduces challenges in reliability and sem…
Serious Games: Human-AI Interaction, Evolution, and Coevolution
The serious games between humans and AI have only just begun. Evolutionary Game Theory (EGT) models the competitive and cooperative strateg…
Shutdownable Agents through POST-Agency
Many fear that future artificial agents will resist shutdown. I present an idea - the POST-Agents Proposal - for ensuring that doesn't happ…
AI's Blind Spots: Geographic Knowledge and Diversity Deficit in Generated Urban Scenario
Diffusion-based text-to-image models are increasingly used for urban analysis and scenario generation, but their geographic knowledge and r…
Policy Improvement with Style-Specific Demonstrations
Proficient game agents with diverse play styles enrich the gaming experience and enhance the replay value of games. However, recent advance…
Interactive Multi-Objective Probabilistic Preference Learning with Soft and Hard Bounds
High-stakes decision-making involves navigating multiple competing objectives with expensive evaluations. For instance, in brachytherapy, c…
A Technical Survey of Reinforcement Learning Techniques for Large Language Models
This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Prox…
Interactive Learning for LLM Reasoning
Existing multi-agent learning approaches have developed interactive training environments to explicitly promote collaboration among multipl…
Activation-Deactivation: A General Framework for Robust Post-hoc Explainable AI
Perturbation-based explainability methods face criticism due to their reliance on out-of-distribution mutants. This raises doubts about the…
RLIE: Rule Generation with Logistic Regression, Iterative Refinement, and Evaluation for Large Language Models
Large Language Models (LLMs) can propose rules in natural language, sidestepping the need for a predefined predicate space in traditional r…
A Unified Geometric Space for Topological Alignment Between Transformer-Based Models and Human Brain Networks
Whether artificial neural networks organize information comparably to the human brain remains unclear. Prior brain--AI alignment studies ar…
Optimal-Agent-Selection: State-Aware Routing Framework for Efficient Multi-Agent Collaboration
The emergence of multi-agent systems powered by large language models (LLMs) has unlocked new frontiers in complex task-solving, enabling d…
CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
Long Chain-of-Thought (CoT) traces can improve reasoning accuracy, but repeatedly generating them is costly for smaller or latency-constrai…
Turbo-Muon: Almost-Orthogonal Pre-Conditioning for Fast Muon Updates
Orthogonality-based optimizers, such as Muon, have recently shown strong performance across large-scale training and community-driven effic…
The Language of Bargaining: Linguistic Effects in LLM Negotiations
Negotiation is a core component of social intelligence, requiring agents to balance strategic reasoning, cooperation, and social norms. Rec…
OpenTinker: Separating Concerns in Agentic Reinforcement Learning
We introduce \textsc{OpenTinker}, an open infrastructure for training large language model (LLM) agents with many LoRA-backed policies over…
Programming over Thinking: Efficient and Robust Multi-Constraint Planning
Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting…
Graph Neural Networks are Heuristics
Graph neural networks are usually treated as auxiliaries for combinatorial optimization: they imitate algorithms, guide search, or supply s…
Toward Efficient Agents: Memory, Tool learning, and Planning
Recent years have witnessed increasing interest in extending large language models into agentic systems. While the effectiveness of agents…
Insect-inspired Visual Point-goal Navigation
Insect neuroethology provides a compelling biological template for efficient autonomous navigation. We draw an analogy between the formal e…
NEST: Nascent Encoded Steganographic Thoughts
Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is…
Framework of Thoughts: A Foundation Framework for Dynamic and Optimized Reasoning based on Chains, Trees, and Graphs
Prompting schemes such as Chain of Thought, Tree of Thoughts, and Graph of Thoughts can significantly enhance the reasoning capabilities of…
ARLArena: A Unified Framework for Stable Agentic Reinforcement Learning
Agentic reinforcement learning (ARL) has rapidly gained attention as a promising paradigm for training agents to solve complex, multi-step…
HVR-Met: A Hypothesis-Verification-Replanning Agentic System for Extreme Weather Diagnosis
While deep learning-based weather forecasting paradigms have made significant strides, addressing extreme weather diagnostics remains a for…
Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning
When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to fa…
Correlation-Weighted Multi-Reward Optimization for Compositional Generation
Text-to-image models produce images that align well with natural language prompts, but compositional generation has long been a central cha…
On the Ability of Transformers to Verify Plans
Transformers have shown inconsistent success in AI planning tasks, and theoretical understanding of when generalization should be expected…
Mecha-nudges for Machines
AI agents are becoming active decision-makers on the Internet. As they make decisions in the same environments as humans, the environments…
The Anatomy of Uncertainty in LLMs
Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment. Current approach…
Chronos: The AI Co-Historian
AI is increasingly supporting, accelerating, and automating scientific discovery across subjects. Yet, the adoption of AI in historical res…
TRACE: Capability-Targeted Agentic Training
Models often fail to complete agentic tasks because they lack core capabilities required by the target environment. However, mainstream app…
Gypscie: A Cross-Platform AI Artifact Management System
Artificial Intelligence (AI) models, encompassing both traditional machine learning (ML) and more advanced approaches such as deep learning…
Fun-TSG: A Function-Driven Multivariate Time Series Generator with Variable-Level Anomaly Labeling
Reliable evaluation of anomaly detection methods in multivariate time series remains an open challenge, largely due to the limitations of e…
CAP-CoT: Cycle Adversarial Prompt for Improving Chain of Thoughts in LLM Reasoning
Chain-of-Thought (CoT) prompting has emerged as a simple and effective way to elicit step-by-step solutions from large language models (LLM…
To Use AI as Dice of Possibilities with Timing Computation
The dominant noun-based modeling paradigm, grounded in probability theory and committed to pre-specified noun entities as primitive modelin…
Stop Automating Peer Review Without Rigorous Evaluation
Large language models offer a tempting solution to address the peer review crisis. This position paper argues that today's AI systems shoul…
Agentic Retrieval-Augmented Generation for Financial Document Question Answering
Financial document question answering (QA) demands complex multi-step numerical reasoning over heterogeneous evidence--structured tables, t…
Beyond the Black Box: Interpretability of Agentic AI Tool Use
AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because these tool-use decisions ar…
Attributing Emergence in Million-Agent Systems
Large language models (LLMs) can simulate human-like reasoning and decision-making in individual agents. LLM-powered multi-agent systems (M…
Sign-Separated Asymmetric Finite-Time Error Analysis of Q-Learning
Q-learning is known to suffer from overestimation bias: because the Bellman update maximizes noisy or imperfect action-value estimates, pos…
When Outcome Looks Right But Discipline Fails: Trace-Based Evaluation Under Hidden Competitor State
Outcome-only evaluation can certify economically unsafe agents: a policy can hit a business KPI while violating deployable behavioral disci…
HaorFloodAlert: A 72-Hour Machine Learning Early Warning System for Flash Floods in Bangladesh's Haor Wetlands
Every spring, flash floods strike the haor wetlands of northeast Bangladesh just before the boro rice harvest, and one flood can erase a fa…
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
Large language model (LLM) agents rely on reusable skills to solve complex tasks, but existing skill creation approaches often treat skills…
Cultural Binding Heads in Language Models
LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference aw…
Planning with the Views
Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1…
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches
While Reasoning Language Models (RLMs) are rapidly emerging as powerful tools for scientific research, their impact is primarily concentrat…
TokenMizer: Graph-Structured Session Memory for Long-Horizon LLM Context Management
Long-horizon LLM sessions outlive their context windows, and the standard mitigations - truncation, summarization, retrieval - share a stru…
Some hypotheses on how chatbots work in problem-solving-driven conversations. Large Language Models as confirmation of the Innovation Illusion
We discuss the nature of chatbots as conversation partners when discussing the solution of problems. What can chatbots do and what can't th…
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem…
ComplexConstraints and Beyond: Expert Rubrics for RLVR
Evaluation protocols can lag behind LLM capabilities. Programmatically verified benchmarks cover narrow surface constraints, whereas real-w…
WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces
Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, brow…
Reducing the Complexity of Deep Learning Models for EEG Analysis on Wearable Devices
Wearable healthcare devices are the fastest-growing Internet of Things (IoT) sector. Many automated healthcare services rely on two crucial…
Medical Heuristic Learning: An LLM-Driven Framework for Interpretable and Auditable Clinical Decision Rules
Predictive modeling for clinical decision support requires not only strong predictive performance but also transparent decision logic. Alth…
Kairos: A Regret-Aware Native World-Action Model Stack for Physical AI
We introduce \textbf{Kairos}, a regret-aware native world-action model stack for Physical AI. Kairos is motivated by the view that a physic…
Skill Coverage: A Test Adequacy Metric for Agent Skills
Agent skills encode reusable procedural knowledge for large language model (LLM) agents, and existing benchmarks show that such skills can…
Text Dictates, Music Decorates: Energy-based Attention for Editable Dance Motion Generation
Choreographic motion generation poses unique challenges for AI, demanding precise semantic control over complex, temporally structured, and…
Decentralised AI Training and Inference with BlockTrain
Frontier AI training is increasingly shaped by access to dense, centrally controlled accelerator clusters. This creates a structural advant…
BluTrain: A C++/CUDA Framework for AI Systems
Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its t…
Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval
Foundation-model agents in multi-step, open-ended environments frequently suffer from compounding errors, where early mistakes contaminate…
Autodata: An agentic data scientist to create high quality synthetic data
We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation dat…
Understanding Rollout Error in Graph World Models
World models are increasingly used for planning, yet most analyses of rollout error assume vector-valued states and scalar error amplificat…
Agent vs. Parametric World Models: Hybrid Planning for Reliable Language Agents
Language agents plan by generating not only actions but also implicit predictions of how the world will change. These imagined state update…
Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline
Recent advancements in generative artificial intelligence (AI) and large language models (LLMs) have shown significant promise in automatin…
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content i…
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents colla…
When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models
Reasoning models spend test-time compute unevenly across instances, and a growing family of early-exit rules -- confidence thresholds, entr…
A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management
We present a three-phase deep reinforcement learning system for personalized portfolio management that addresses three limitations shared b…
World-Model Collapse as a Phase Transition
Water looks unchanged as it warms, then at a critical point it boils. We ask whether long-horizon language agents show an analogous transit…
Ask the World Before Acting: Environment Probing for Calibrated Agent World Models
Language agents acting over long horizons must maintain beliefs about tool states, object locations, graph edges, and subgoal dependencies.…
Harnessing Textual Refusal Directions for Multimodal Safety
To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activa…
The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons
Many information systems are built around documents: self-contained units optimised for print production and linear reading. While effectiv…
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
LLMs increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically valid yet stale, infeasible, c…
AI Native Games: A Survey and Roadmap
Generative AI now enables games to produce dialogue, quests, characters, images, and worlds at runtime. Yet generation alone does not make…
PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management
Large-scale crowd management requires pedestrian simulations that are both computationally efficient and compatible with feedback-based con…
Agentic generation of verifiable rules for deterministic, self-expanding reaction classification
Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign…
Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases
Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant cha…
Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation
Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually…
Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment
In multi-source ECG deployment, models may need to incorporate new data sources when earlier raw ECGs cannot be retained or replayed. Freez…
Repair the Amplifier, Not the Symptom: Stable World-Model Correction for Agent Rollouts
Long-horizon language agents increasingly maintain executable world models in the form of planning graphs, where tool calls, validators, me…
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing…
ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair
Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad search…
PACE: A Proxy for Agentic Capability Evaluation
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single…
ContextNest: Verifiable Context Governance for Autonomous AI Agent
Autonomous AI agents increasingly depend on external knowledge stores, yet most retrieval pipelines provide relevance without durable guara…
Querying and Repairing Inconsistent Prioritized Knowledge Bases: Complexity Analysis and Links with Abstract Argumentation
In this paper, we explore the issue of inconsistency handling over prioritized knowledge bases (KBs), which consist of an ontology, a set o…
Double Fuzzy Probabilistic Interval Linguistic Term Set and a Dynamic Fuzzy Decision Making Model based on Markov Process with tts Application in Multiple Criteria Group Decision Making
The probabilistic linguistic term has been proposed to deal with probability distributions in provided linguistic evaluations. However, bec…
Domain Knowledge-Informed Self-Supervised Representations for Workout Form Assessment
Maintaining proper form while exercising is important for preventing injuries and maximizing muscle mass gains. Detecting errors in workout…
A Transformer-Based Contrastive Learning Approach for Few-Shot Sign Language Recognition
Sign language recognition from monocular video or 2D pose sequences is challenging, both because 3D information must be inferred from 2D ob…
Restricted Bernoulli Matrix Factorization: Balancing the trade-off between prediction accuracy and coverage in classification based collaborative filtering
Reliability measures associated with the prediction of the machine learning models are critical to strengthening user confidence in artific…
Unveiling the Unborn: Advancing Fetal Health Classification through Machine Learning
Fetal health classification is a critical task in obstetrics, enabling early identification and management of potential health problems. Ho…
Learning to Visually Connect Actions and their Effects
We introduce the novel concept of visually Connecting Actions and Their Effects (CATE) in video understanding. CATE can have applications i…
TERC: A Transfer Entropy Redundancy Criterion for State Variable Selection in Reinforcement Learning
Identifying the most suitable variables to represent the state is a fundamental challenge in Reinforcement Learning (RL). These variables m…
Graph Unitary Message Passing
Unitarity is a useful principle for stabilizing deep neural networks, but in graph neural networks (GNNs) instability is induced not only b…
CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes
Causal video question answering (QA) has garnered increasing interest, yet existing datasets often lack depth in causal reasoning. To addre…
MambaCapsule: Towards Transparent Cardiac Disease Diagnosis with Electrocardiography Using Mamba Capsule Network
Cardiac arrhythmia, a condition characterized by irregular heartbeats, often serves as an early indication of various heart ailments. With…
Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies
LLMs are increasingly used world-wide from daily tasks to agentic systems and data analytics, requiring significant GPU resources. While LL…
Stroke Prediction using Clinical and Social Features in Machine Learning
Every year in the United States, 800,000 individuals suffer a stroke - one person every 40 seconds, with a death occurring every four minut…
Robust Counterfactual Explanations under Model Multiplicity Using Multi-Objective Optimization
In recent years, explainability in machine learning has gained importance. In this context, counterfactual explanation (CE), which is an ex…
Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility
Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the com…
Evaluating LLM-Based Regression Test Generation
Large Language Models (LLMs) have shown tremendous promise in automated software engineering. In this paper, we investigate LLMs for just-i…
Machine Unlearning via Information Theoretic Regularization
How can we effectively remove or ``unlearn'' undesirable information, such as specific features or the influence of individual data points,…
Evolutionary Guided Decoding: Iterative Value Refinement for LLMs
While guided decoding, especially value-guided methods, has emerged as a cost-effective alternative for controlling language model outputs…
Using Mechanistic Interpretability to Craft Adversarial Attacks against Large Language Models
Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targ…
Explainable Bayesian deep learning through input-skip Latent Binary Bayesian Neural Networks
Modeling natural phenomena with artificial neural networks (ANNs) often provides highly accurate predictions. However, ANNs often suffer fr…
Empirical Computation: Prompting versus Programming
Large Language Model (LLM) agents can solve *any* computational problem *without* an algorithm in a runtime *independent* of the computatio…
Measuring the Robustness of Audio Deepfake Detection under Real-World Corruption
Deepfakes have emerged as a widespread and rapidly escalating concern in generative AI, spanning images, audio, and videos. Among these, au…
Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home Clusters
On-device inference offers privacy, offline use, and instant response, but consumer hardware restricts large language models (LLMs) to low…
AgentDynEx: Nudging the Mechanics and Dynamics of Multi-Agent Simulations
Multi-agent large language model simulations have the potential to model complex human behaviors and interactions. If the mechanics are set…
GenShin: Guiding Rational Liposome Design by Ranking Liposomal Protein Corona through a Docking-Pose-Free GNN
Rational design of lipid nanoparticles (LNPs) for tissue-specific delivery critically depends on predicting the composition of the protein…
MOSAIC: Skill-Centric Manipulation Planning with Physics Simulation
Planning long-horizon manipulation motions using a set of predefined skills is a central challenge in robotics; solving it efficiently coul…
Exploring Context-aware and LLM-driven Locomotion for Immersive Virtual Reality
Locomotion plays a crucial role in shaping the user experience within virtual reality environments. In particular, hands-free locomotion of…
kAgent: An execution-guided crash resolution agent for the Linux kernel
Fuzzing frameworks like syzkaller have uncovered thousands of Linux kernel crashes, many of which are critical and security-sensitive. Howe…
Towards Understanding Deep Learning Model in Image Recognition via Coverage Test
Deep neural networks (DNNs) play a crucial role in the field of artificial intelligence, and their security-related testing has been a prom…
Quick ViTs: Speeding up Vision Transformers through Equivariance
Natural images exhibit strong geometric regularities: local structures, such as edges, corners, and textures, appear in many orientations a…
PDFBench: A Benchmark for De novo Protein Design from Function
Function-guided protein design is a crucial task with significant applications in drug discovery and enzyme engineering. However, the field…
Seven Security Challenges in Cross-domain Multi-agent LLM Systems
Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint di…
Boosting Automatic Exercise Evaluation Through Musculoskeletal Simulation-Based IMU Data Augmentation
Automated evaluation of movement quality can enhance physiotherapeutic treatment and sports training by providing objective, real-time feed…
Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
Natural Language Processing (NLP) has transformed various fields beyond linguistics by applying techniques originally developed for human l…
FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection
The rapid proliferation of surveillance cameras has increased the demand for automated violence detection. While CNNs and Transformers have…
Ensemble Elastic DQN: A Step Dependent Ensemble Approach for Reducing Overestimation in Deep Value-Based Reinforcement Learning
Deep Q-Networks (DQN) can suffer from overestimation bias because bootstrapped targets use a maximisation operation over noisy value estima…
InverseScope: Scalable Activation Inversion for Interpreting Large Language Models
Understanding the internal representations of large language models (LLMs) is a central challenge in interpretability research. Existing fe…
ELBO-T2IAlign: A Generic ELBO-Based Method for Calibrating Pixel-level Text-Image Alignment in Diffusion Models
Diffusion models excel at image generation. Recent studies have shown that these models not only generate high-quality images but also enco…
Position: Use Sparse Autoencoders to Discover Unknowns
While sparse autoencoders (SAEs) have generated significant excitement, a series of negative results have added to skepticism about their u…
Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues
Large language models (LLMs) may support counseling training, yet evidence from Japanese-language interactions and automated quality rating…
Interaction Techniques that Encourage Longer Prompts Can Improve Psychological Ownership when Writing with AI
Writing longer prompts for an AI assistant to generate a story increases psychological ownership, a user's feeling that the writing belongs…
Context Tuning for In-Context Optimization
We introduce Context Tuning, a simple and effective method to significantly enhance few-shot adaptation of large language models (LLMs) wit…
Last Layer Hamiltonian Monte Carlo
We explore the use of Hamiltonian Monte Carlo (HMC) sampling as a probabilistic last layer approach for deep neural networks (DNNs). While…
Towards Mitigation of Hallucination for LLM-empowered Agents: Progressive Generalization Bound Exploration and Watchdog Monitor
Empowered by large language models (LLMs), intelligent agents have become a popular paradigm for interacting with open environments to faci…
Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
Multimodal large-scale models have significantly advanced the development of web agents, enabling perception and interaction with digital e…
Parity-Aware Byte-Pair Encoding: Improving Cross-lingual Fairness in Tokenization
Tokenization is the first -- and often least scrutinized -- step of most NLP pipelines. Standard algorithms for learning tokenizers rely on…
Extending Foundational Monocular Depth Estimators to Fisheye Cameras with Calibration Tokens
We propose a method to extend foundational monocular depth estimators (FMDEs), trained on perspective images, to fisheye images. Despite be…
Algorithmic Shortlisting in Participatory Budgeting
Participatory budgeting is a democratic innovation that allows citizens to propose and vote on public investment projects. To help organize…
VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models
Vision Language Models (VLMs) are increasingly being used in a broad range of applications, bringing their security and behavioral control…
Rational Inverse Reasoning: Few-Shot Imitation by Inferring Intent through Planning
Humans can learn a new manipulation task from one or two demonstrations and then perform it in a new room, with new objects, under new cons…
Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs
While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety…
DynamixSFT: Dynamic Mixture Optimization of Instruction Tuning Collections
As numerous instruction-tuning datasets continue to emerge, dynamically balancing and optimizing their mixtures has become a critical chall…
EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation
MultiModal Recommendation (MMR) systems have emerged as a promising solution for improving recommendation quality by leveraging rich item-s…
LLM-Assisted Semantic Alignment and Integration in Collaborative Model-Based Systems Engineering Using SysML v2
Cross-organizational collaboration in Model-Based Systems Engineering (MBSE) faces many challenges in achieving semantic alignment across i…
EyeMulator: Improving Code Language Models by Mimicking Human Visual Attention
Code Language Models (CodeLLMs) learn token importance from data correlations, whereas human developers attend selectively to semantically…
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
Existing Causal-Why Video Question Answering (VideoQA) models often struggle with higher-order reasoning, relying on opaque, monolithic pip…
A Systematic Survey on Large Language Models for Evolutionary Optimization: From Modeling to Solving
Large language models (LLMs) are increasingly integrated with evolutionary computation to support optimization tasks. This survey primarily…
Self-Supervised Goal-Reaching Results in Multi-Agent Cooperation and Exploration
For groups of autonomous agents to achieve a particular goal, they must engage in coordination and long-horizon reasoning. Rather than rely…
ICR-RL: Deep Reinforcement Learning via In-Context Regression
Recent advancements in machine learning have largely been driven by foundation models (FMs) trained on large, diverse datasets, enabling th…
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned…
Interpretable Nanoporous Materials Design with Symmetry-Aware Networks
Nanoporous materials hold promise for diverse sustainable applications, yet their vast chemical space poses challenges for efficient design…
Agentic Artificial Intelligence for Multistage Physics Experiments at a Large-Scale User Facility Particle Accelerator
We present the first language-model-driven agentic artificial intelligence (AI) system to autonomously execute multi-stage physics experime…
Polarity Detection of Sustainable Development Goals in News Text
The United Nations' Sustainable Development Goals (SDGs) provide a globally recognised framework for addressing major societal, environment…
Adaptive Margin RLHF via Preference over Preferences
Margin-based optimization is fundamental to improving generalization and robustness in classification tasks. In the context of reward model…
OctoPipe: Reducing Pipeline Bubbles for Heterogeneous Models via Co-Optimizing Partitioning, Placement, and Scheduling
Pipeline parallelism is widely used to train large language models (LLMs). However, increasing heterogeneity in model architectures exacerb…
MAD-PINN: A Decentralized Physics-Informed Machine Learning Framework for Safe and Optimal Multi-Agent Control
Co-optimizing safety and performance in large-scale multi-agent systems remains a fundamental challenge. Existing approaches based on multi…
VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models
As Vision Language Models (VLMs) are deployed across safety-critical applications, understanding and controlling their behavioral patterns…
Quadratic Programming Approach for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games
There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player zero-sum imperfect-infor…
Panorama: Fast-Track Nearest Neighbors
Approximate Nearest-Neighbor Search (ANNS) pipelines for high-dimensional neural embeddings spend the bulk of their query time in candidate…
The Three Regimes of Offline-to-Online Reinforcement Learning
Offline-to-online reinforcement learning (RL) has emerged as a practical paradigm that leverages offline datasets for pretraining and onlin…
Verifier-free Test-Time Sampling for Vision-Language-Action Models
Vision-Language-Action models (VLAs) have demonstrated remarkable performance in robot control. However, they remain fundamentally limited…
TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance
Query-product relevance prediction is fundamental to e-commerce search and has become even more critical in the era of AI-powered shopping,…
SilvaScenes: Tree Detection and Species Classification from Under-Canopy Images in Natural Forests
Interest in forestry automation is growing alongside rapid advances in deep learning. In particular, tree detection and taxonomic classific…
Conditional Clifford-Steerable CNNs for PDE Modeling
We introduce Conditional Clifford-Steerable CNNs (C-CSCNNs), a unified framework that incorporates equivariance to arbitrary pseudo-Euclide…
SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses
Large Language Models (LLMs) are increasingly used as interfaces to information, code, and real-world services, making prompt-level securit…
UNDREAM: Bridging Differentiable Rendering and Photorealistic Simulation for End-to-end Adversarial Attacks
Deep learning models deployed in safety critical applications like autonomous driving use simulations to test their robustness against adve…
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We s…
Neurosymbolic Characterization for Reliable Access Control Policy Analysis
Access control policies are reliability-critical configuration artifacts in cloud systems, yet administrators frequently struggle to verify…
PuzzleMoE: Efficient Compression of Large Mixture-of-Experts Models via Sparse Expert Merging and Bit-packed inference
Mixture-of-Experts (MoE) models have shown strong potential in scaling language models efficiently by activating only a small subset of exp…
Large Language Models Develop Novel Social Biases Through Adaptive Exploration
As large language models (LLMs) are adopted into frameworks that grant them the capacity to make real decisions, it is increasingly importa…
Resilient by Design -- Active Inference for Distributed Continuum Intelligence
Failures are the norm in highly complex and heterogeneous devices spanning the distributed computing continuum (DCC), from resource-constra…
SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial…
When is a System Discoverable from Data? Discovery Requires Chaos
The deep learning revolution has spurred a rise in advances of using AI in sciences. Within physical sciences the main focus has been on di…
Walrus: A Cross-Domain Foundation Model for Continuum Dynamics
Foundation models have transformed machine learning for language and vision, but achieving comparable impact in physical simulation remains…
Score-Regularized Joint Sampling with Importance Weights for Flow Matching
Flow matching models effectively represent complex distributions, yet estimating expectations of functions of their outputs remains challen…
KAN vs LSTM Performance in Time Series Forecasting
This study presents a controlled comparison of baseline Kolmogorov-Arnold Networks (KAN), implemented via PyKAN, and Long Short-Term Memory…
Exploring the Rashomon Set for Concept-Based Models
In many machine learning problems, there may exist multiple models that achieve nearly identical predictive performance while relying on fu…
InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
Recent approaches in controllable novel view video generation often rely on fine-tuning pre-trained Video Diffusion Models (VDMs). This dom…
Developing an LLM-Based Feedback System Grounded in Evidence-Centered Design to Support Physics Problem Solving
Generative AI offers new opportunities for individualized and adaptive learning, e.g., through large language model (LLM)-based feedback sy…
AOI: Context-Aware Multi-Agent Operations via Dynamic Scheduling and Hierarchical Memory Compression
Cloud-native systems have made operational work both more powerful and harder to automate: incidents unfold across microservices, logs and…
Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
Focusing on low-resource languages is an essential step toward democratizing generative AI. In this work, we contribute to reducing the mul…
EvoXplain: When Machine Learning Models Agree on Predictions but Disagree on Why -- Measuring Mechanistic Multiplicity Across Training Runs
Machine learning models are primarily judged by predictive performance, especially in applied genomics, where explanations are read as biol…
JADAI: Jointly Amortizing Adaptive Design and Bayesian Inference
We consider problems of parameter estimation where design variables can be actively optimized to maximize information gain. To this end, we…
Resource-constrained Project Scheduling with Time-of-Use Energy Tariffs and Machine States: A Logic-based Benders Decomposition Approach
In this paper, we investigate the Resource-Constrained Project Scheduling Problem (RCPSP) with Time-of-Use (TOU) energy tariffs and machine…
Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework
Semantic Multi-Object Tracking (SMOT) is evolving from purely geometric localization toward comprehensive video understanding. However, exi…
ARCQuant: Boosting NVFP4 Quantization with Augmented Residual Channels for LLMs
The emergence of fine-grained numerical formats like NVFP4 presents new opportunities for efficient Large Language Model (LLM) inference. H…
Motion Attribution for Video Generation
Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTI…
One Prompt, Many Sounds: Modeling Listener Variability in LLM-Based Equalization
Conventional audio equalization is a static process that requires manual and cumbersome adjustments to adapt to changing listening contexts…
Predicting Biased Human Decision-Making with Large Language Models in Conversational Settings
We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their prediction…
R$^2$PO: Decoupling Rollout and Inference Policies for LLM Reasoning
Existing reinforcement learning methods for LLM reasoning implicitly assume that the policy generating training trajectories should coincid…
Tracing 3D Anatomy in 2D Strokes: A Multi-Stage Projection Driven Approach to Cervical Spine Fracture Identification
Cervical spine fractures require rapid and accurate diagnosis, yet automatic CT interpretation remains challenging as subtle injuries must…
No Reliable Evidence of Self-Reported Sentience in Small Large Language Models
Whether language models possess sentience has no empirical answer. But whether they believe themselves to be sentient can, in principle, be…
The Rise of Large Language Models and the Direction and Impact of US Federal Research Funding
Federal research funding shapes the direction, diversity, and impact of the US scientific enterprise. Large language models (LLMs) are rapi…
BoRP: Bootstrapped Regression Probing for Scalable and Human-Aligned LLM Evaluation
Accurate evaluation of user satisfaction is critical for iterative development of conversational AI. However, for open-ended assistants, tr…
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging…
A Random Matrix Theory Perspective on the Consistency of Diffusion Models
Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same no…
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
Generating long-form content, such as minute-long videos and extended texts, is increasingly important for modern generative models. Block…
Multi-Way Representation Alignment
The Platonic Representation Hypothesis suggests that independently trained neural networks converge to increasingly similar latent spaces.…
SHINE: A Scalable In-Context Hypernetwork for Mapping Context to LoRA in a Single Pass
We propose SHINE (Scalable Hyper In-context NEtwork), a scalable hypernetwork that can map diverse meaningful contexts into high-quality Lo…
Implementing Grassroots Logic Programs with Multiagent Transition Systems and AI (Full Version)
Grassroots Logic Programs (GLP) is a concurrent logic programming language in which logic variables are partitioned into paired readers and…
Endogenous Resistance to Activation Steering in Language Models
Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.g., ``wait…
Deriving Neural Scaling Laws from the statistics of natural language
Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no exi…
rePIRL: Learn PRM with Inverse RL for LLM Reasoning
Process rewards have been widely used in deep reinforcement learning to improve training efficiency, reduce variance, and prevent reward ha…
Learning to Discover Iterative Spectral Algorithms
We introduce AutoSpec, a neural network framework for discovering iterative spectral algorithms for large-scale numerical linear algebra an…
pFedNavi: Structure-Aware Personalized Federated Vision-Language Navigation for Embodied AI
Vision-Language Navigation VLN requires large-scale trajectory instruction data from private indoor environments, raising significant priva…
NextCrystal: a Symmetry-Driven Generative Framework for Crystal Structure Prediction
Crystal structure prediction (CSP), which aims to predict the 3D atomic arrangement of a crystal from its composition, is central to materi…
SpectralGCD: Spectral Concept Selection and Cross-modal Representation Learning for Generalized Category Discovery
Generalized Category Discovery (GCD) aims to identify novel categories in unlabeled data while leveraging a small labeled subset of known c…
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
Reward modeling is central to alignment pipelines such as RLHF, RLAIF, and PPO-based policy optimization, yet its reliability is constraine…
Inelastic Constitutive Kolmogorov-Arnold Networks: A generalized framework for automated discovery of interpretable inelastic material models
A key problem of solid mechanics is the identification of the constitutive law of a material, that is, the relation between strain history…
Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
Reinforcement Learning from Human Feedback (RLHF) or Verifiable Rewards (RLVR) are two key steps in the post-training of modern Language Mo…
SFL-Net: Source-Factorized Latent Representation Learning for Multi-Contrast MRI to Tau-PET Synthesis
Tau positron emission tomography supports Alzheimer's disease staging but is difficult to scale because of tracer, scanner, and radiation c…
Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving
Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distri…
Causal Mechanism Reduction: Mechanism Replacement for Neural Network Pruning and Abstraction
Which internal mechanisms of a neural network can be replaced while preserving the computation it performs? Structured pruning asks for sma…
OSF: On Pre-training and Scaling of Sleep Foundation Models
Polysomnography (PSG) provides the gold standard for sleep assessment but suffers from substantial heterogeneity across recording devices a…
Efficient Flow Matching for Sparse-View CT Reconstruction
Generative models, particularly Diffusion Models (DM), have shown strong potential for Computed Tomography (CT) reconstruction serving as e…
GIPO: Gaussian Importance Sampling Policy Optimization
Post-training with reinforcement learning (RL) has recently shown strong promise for advancing multimodal agents beyond supervised imitatio…
LAW & ORDER: Adaptive Spatial Weighting for Medical Diffusion and Segmentation
Medical image analysis depends on accurate segmentation and controllable synthesis, but both tasks face severe spatial imbalance: lesions o…
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctne…
BEVLM: Distilling Semantic Knowledge from LLMs into Bird's-Eye View Representations
The integration of Large Language Models (LLMs) into autonomous driving has attracted growing interest for their strong reasoning and seman…
Multi-Agent Reinforcement Learning for V2X Resource Allocation: Disentangling MARL Challenges Through Benchmarking
Radio resource allocation (RRA) is a critical function in cellular vehicle-to-everything (C-V2X) networks, where vehicles must share limite…
A Hybrid Quantum Circuit Born Machine Framework for Financial Volatility Forecasting: Quantum-Assisted Training and Classical Inference
Accurate financial volatility forecasting is crucial but challenged by the non-linear, highly correlated nature of market data. Recently, q…
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
We consider the task of end-to-end spreadsheet generation, where language models produce spreadsheet artifacts to satisfy users' explicit a…
Safe RLHF Beyond Expectation: Stochastic Dominance for Universal Spectral Risk Control
Safe Reinforcement Learning from Human Feedback (RLHF) typically enforces safety through expected cost constraints, but the expectation cap…
Na\"ive PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation
Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise. Thus, like playing the slo…
How Transformers Reject Wrong Answers: Rotational Dynamics of Factual Constraint Processing
When a decoder-only transformer is forced to process matched correct and incorrect single-token continuations of a factual query, the two p…
Human-like Object Grouping in Self-supervised Vision Transformers
Vision foundation models trained with self-supervised objectives achieve strong performance across diverse tasks and exhibit emergent objec…
Trust-Region Noise Search for Black-Box Alignment of Diffusion and Flow Models
Optimizing the noise samples of diffusion and flow models is an increasingly popular approach to align these models to target rewards at in…
CABTO: Context-Aware Behavior Tree Grounding for Robot Manipulation
Behavior Trees (BTs) offer a powerful paradigm for designing modular and reactive robot controllers. BT planning, an emerging field, provid…
Towards Reliable Local Security Agents: Verifiable Post-Training for Linux Privilege Escalation
LLM agents are becoming increasingly important in the security domain, but leading systems are often closed-source, cloud-based, hard to re…
A Two-stage Transformer Framework for Temporal Localization of Distracted Driver Behaviors
The identification of hazardous driving behaviors from in-cabin video streams is essential for enhancing road safety and supporting the det…
From Arithmetic to Logic: The Resilience of Logic and Lookup-Based Neural Networks Under Parameter Bit-Flips
The deployment of deep neural networks (DNNs) in safety-critical edge environments necessitates robustness against hardware-induced bit-fli…
Estimating Individual Tree Height and Species from UAV Imagery
Accurate estimation of forest biomass, a major carbon sink, relies heavily on tree-level traits such as height and species. Unoccupied Aeri…
MCLMR: A Model-Agnostic Causal Learning Framework for Multi-Behavior Recommendation
Multi-Behavior Recommendation (MBR) leverages multiple user interaction types (e.g., views, clicks, purchases) to enrich preference modelin…
SutureFormer: Learning Surgical Trajectories via Goal-conditioned Offline RL in Pixel Space
Predicting surgical needle trajectories from endoscopic video is critical for robot-assisted suturing, enabling anticipatory planning, real…
Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching
Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality…
Reachability Across the NL/PL Boundary: A Taxonomy-Driven Dataflow Model for LLM-Integrated Applications
LLM API calls have become a standard programming primitive, but they create a program boundary that disrupts traditional dataflow analysis.…
Beyond Task Completion: A Verification-vs.-Conformance Gap in Tool-Evolving Agents
Agents that synthesize their own tools ship a second artifact alongside each answer: a software library that future tasks reuse, compose, a…
Streaming Model Cascades for Semantic SQL
Modern data warehouses extend SQL with semantic operators that invoke large language models on each qualifying row, making per-row inferenc…
Council Mode: A Heterogeneous Multi-Agent Consensus Framework for Reducing LLM Hallucination and Bias
Large Language Models (LLMs) have demonstrated advanced capabilities but often suffer from factual inaccuracies (hallucinations) and system…
AViS-Mamba: Adaptive Visual Steering of Audio State-Space Dynamics for Violence Detection
Automatic violence detection from video is challenging because violent interactions may be distant, occluded, or only partially visible. Au…
Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types
Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks…
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been…
SeaAlert: Robust Severity Classification and LLM-Based Information Extraction for Noisy Maritime Distress Communications
Maritime distress communications transmitted over very high frequency (VHF) radio are safety-critical voice messages used to report emergen…
CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas
It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the…
Don't Make Models Guess Security and Safety: Symbolic Guardrails for Domain-Specific AI Agents
There is increasing interest in integrating AI agents that invoke tools into domain-specific commercial software, where unintended tool cal…
Unveiling Stochasticity: Universal Multi-modal Probabilistic Modeling for Traffic Forecasting
Traffic forecasting is a challenging spatio-temporal modeling task and a critical component of urban transportation management. Current stu…
Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives
AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the…
Test-Time Adaptation for EEG Foundation Models: A Systematic Study under Real-World Distribution Shifts
Electroencephalography (EEG) foundation models have shown strong potential for learning generalizable representations from large-scale neur…
Towards Generalizable Deepfake Image Detection with Vision Transformers
In today's day and age, we face a challenge in detecting deepfake images because of the fast evolution of modern generative models and the…
The Rise of Verbal Tics in Large Language Models: A Systematic Analysis Across Frontier Models
As Large Language Models (LLMs) continue to evolve through alignment techniques such as Reinforcement Learning from Human Feedback (RLHF) a…
StarTSE: Towards Streaming Target Speaker Extraction via Chunk-wise Interleaved Splicing of Autoregressive Language Model
While generative models have set new benchmarks for Target Speaker Extraction (TSE), their inherent reliance on global context precludes de…
CorridorVLA: Explicit Spatial Constraints for Generative Action Heads via Sparse Anchors
Vision--Language--Action (VLA) models often use intermediate representations to connect multimodal inputs with continuous control, yet spat…
Query2Diagram: Answering Developer Queries with UML Diagrams
Software documentation frequently becomes outdated or fails to exist entirely, yet developers need focused views of their codebase to under…
Seeing Is No Longer Believing: Frontier Image Generation Models, Synthetic Visual Evidence, and Real-World Risk
Frontier image generation has moved from artistic synthesis toward synthetic visual evidence. Systems such as GPT Image 2, Nano Banana Pro,…
Kwai Summary Attention Technical Report
Long-context ability, has become one of the most important iteration direction of next-generation Large Language Models, particularly in se…
Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity
Closed-source frontier labs do not disclose parameter counts. Storing F facts requires at least F/(bits per parameter) weights, so factual…
Fitting Horn DL Ontologies to ABox and Query Examples: A Tale of Simulation Quantifiers and Finite Models
We study the problem of fitting a description logic (DL) ontology to a given set of positive and negative examples that take the form of an…
Shao: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation
A common design pattern in high-quality music generation is to handle structure and fidelity in different representation spaces: a generato…
ARISE: A Repository-level Graph Representation and Toolset for Agentic Program Repair and Fault Localization
Automated program repair at repository scale requires an agent to locate a fault among thousands of files and synthesize a correct patch. E…
Undetectable Backdoors in Model Parameters: Hiding Sparse Secrets in High Dimensions
We present Sparse Backdoor, a supply-chain attack that plants a provably undetectable backdoor in pre-trained image classifiers, including…
IRC-Bench: Recognizing Entities from Contextual Cues in First-Person Reminiscences
When people recount personal memories, they often refer to people, places, and events indirectly, relying on con-textual cues rather than e…
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
Biological laboratory automation can reduce repetitive manual work and improve reproducibility, but reliable embodied execution in wet-lab…
Defense effectiveness across architectural layers: a mechanistic evaluation of persistent memory attacks on stateful LLM agents
Persistent memory in LLM agents creates an attack surface that production safety classifiers do not observe: the payload enters via RAG ret…
Evolutionary Ensemble of Agents
We introduce Evolutionary Ensemble (EvE), a decentralized framework that organizes existing, highly capable coding agents into a live, co-e…
Clin-JEPA: A Multi-Phase Co-Training Framework for Joint-Embedding Predictive Pretraining on EHR Patient Trajectories
We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories. JE…
Active Sensing with Meta-Reinforcement Learning for Emitter Localization from RF Observations
Global navigation satellite system (GNSS) interference poses a serious threat to reliable positioning, especially in indoor and multipath-r…
CogAdapt: Adapting Clinical ECG Foundation Models for Wearable Cognitive Load Assessment
Assessing cognitive load continuously and at low latency would help adaptive human-computer interaction, but it remains hard because labele…
High-Risk AI Systems and the Problem of Identity in the European AI Act
The EU Artificial Intelligence Act (AIA) establishes a lifecycle governance regime for high-risk AI systems built around ex-ante conformity…
Hidden-State Privacy Has an Empty Middle
Of $1{,}536$ Gaussian release covariances we tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate…
IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM Inference
Large Language Models (LLMs) are increasingly expected to operate over long contexts, yet standard softmax attention incurs a KV cache that…
Learning in Low-Dimensional Subspaces: Orthogonal Bottlenecks for Reinforcement Learning
Deep reinforcement learning (RL) agents commonly rely on high-dimensional neural representations, despite growing evidence that task-releva…
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
Memory is essential for enabling large language models to support long-horizon reasoning, yet existing memory systems remain unreliable and…
Emergent Semantic Representations in World Models through Physical Interaction without Linguistic Supervision
What does a world model learn from physical exploration, without any linguistic supervision? We argue the answer is organized by a single p…
Evaluating Skill and Stability of ArchesWeather and ArchesWeatherGen under Multi-Decadal Climate Simulations
We evaluate the climate simulation capabilities of ArchesWeather and ArchesWeatherGen, two machine learning models originally trained for w…
Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms
As autonomous language model agents proliferate, forming an emerging agentic web with real-world consequences, what credibility signals can…
De-attribute to Forget for LLM Unlearning
The rapid development of large language models (LLMs) has raised concerns on the use of inappropriate data for training, which has led to a…
Automatically Differentiable Nonlinear Tensor Networks (ADNTNs) for Exponential Parameter Compression of Deep Neural Networks
Large deep neural networks are costly to store and deploy because inference must move and evaluate many parameters. This paper studies \emp…
ParetoPilot: Zero-Surrogate Offline Multi-Objective Optimization via Infer-Perturb-Guide Diffusion
Offline multi-objective optimization (Offline MOO) seeks Pareto-optimal designs from static datasets without additional environment interac…
FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (June 13th version)
Conventional HPC holds that native hardware FP64 is the irreducible foundation of scientific computing. On AI-optimized GPUs of the NVIDIA…
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify…
Quickest Detection of Hallucination Onset: Delay Bounds and Learned CUSUM Statistics
Token-level hallucination detectors are evaluated as classifiers, by AUC over all tokens, yet a streaming monitor is judged by its reaction…
A Mathematical Theory of Value: a synthesis on goal-directed agency under resource constraints
We propose that value -- the quantity goal-directed agents create, destroy, and exchange -- is a lawful structural quantity in the same cat…
Iterative Visual Thinking and the Self-Correction Mirage in VLM Grounding
Letting a vision-language model (VLM) think longer at test time has driven much recent progress. A natural way to bring this to spatial gro…
Lect\=uraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching
Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but…
Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering
Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by…
Attention is Just Another Name for Coupling? A Fast-Slow ODE Perspective on Hierarchical Pretraining
We re-interpret Transformer pretraining as a fast-slow, singularly perturbed flow along depth, with untied weights as its non-autonomous fe…
A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning
Reinforcement learning (RL) systems often degrade when operating conditions differ from those previously encountered, reflecting distributi…
TransitNet: A Compact Attention-Augmented Deep Learning Framework for Low-SNR Transit Blind Searches
Motivated by the observational incompleteness of intermediate-to-long-period Earth-size planets, we present TransitNet, a compact attention…
OrthoReg: Orthogonal Regularization for Hybrid Symbolic-Neural Dynamical Systems
Dynamical systems are fundamental to modeling the natural world, yet modeling them involves a persistent trade-off: manually prescribed mec…
Hard or Just Unreached? Diagnosing the Sampling Blind Spot in Math-Reasoning Difficulty Estimation
Math and science reasoning benchmarks rely on pass@k, the fraction of sampled chains that reach gold, as the canonical per-example difficul…
Triangular Consistency as a Universal Constraint for Learning Optical Flow
We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision…
ELVA: Exploring Ranking-Driven Universal Multimodal Retrieval
Leveraging Multimodal Large Language Models (MLLMs) via contrastive learning has become a mainstream paradigm for improving the performance…
NRT-Bench: Benchmarking Multi-Turn Red-Teaming of LLM Operator Agents in Safety-Critical Control Rooms
Large language model (LLM) agents are increasingly proposed as supervisory components for safety-critical systems, yet their robustness und…
A Digital Twin Framework for Traffic-Aware UAV Pavement Monitoring in Open-Traffic Conditions
UAV-based pavement inspection can reduce the cost and risk of road-surface monitoring, but real-world deployment remains difficult when tra…
The Power of Light: Improving Synthetic-to-Real Domain Adaptation through Physically-Based Indirect Illumination
While synthetic data generation resolves the manual labeling bottleneck in computer vision, minimizing the syn-to-real domain gap requires…
VideoAgent: All-in-One Framework for Video Understanding and Editing
Video editing has become essential in digital media creation, yet existing automated systems are restricted to short segment processing and…
DramaDirector: Geometry-Guided Short Drama Generation
Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that pr…
UC-Search: Risk-Aware Test-Time Search for Delayed Constrained Time-Series Control
Time-series deployments often need delayed feasible decisions, not only accurate forecasts. UC-Search is a trace-only retained-search layer…
Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
A language model's memory can be worse than no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behin…
Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model
Dense embeddings power semantic search and Retrieval-Augmented Generation, yet a leaked vector database leaks the text behind it, since emb…
Multipath Adaptive Gated Bottleneck Latent ODE with Raman Data Fusion for Cell Culture Process Forecasting
Mammalian cell-culture processes underpin the manufacture of many biopharmaceuticals, yet keeping a run on track is hard: critical process…
Heavy-Ball Q-Learning with Residual Weighting Correction
This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes convergence of its determinist…
The Unverifiability of Artificial General Intelligence (AGI) Alignment, Static and Dynamic: From Trakhtenbrot's Wall to the Safety-Generality Tension
We establish the mathematical limits of AGI safety in two forms: verifying a fixed system, and verifying that a certified safety property p…
Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation
Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards.…
A Deep Multiscale Neural Network for Accurate Neurological Disorder Detection from MRI Scans and Real-Time Web Deployment
Neurological disorders involve diverse pathologies of the brain and nervous system, making early and accurate detection essential. While ma…
Automating the Design of Embodied Agent Architectures
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity expo…
Curvature-Guided Sheaf Diffusion for Unsupervised Community Detection on Heterophilic Graphs
Detecting communities in heterophilic graphs -- where connected nodes often belong to different classes -- is hard for unsupervised methods…
A Stochastic--Geometric Theory of Scaling Laws in Grokking
Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only…
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech
With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly impor…
Modeling Cell-Cycle-Aware Single-Cell Drug Perturbation Responses
Single-cell drug perturbation models should capture transcriptional response magnitude and whether a treatment changes the proliferative st…
From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation
Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unpr…
CVE-TTP KG: Knowledge Graph Linking Software Vulnerabilities to Attack Behaviors
In the evolving threat landscape, adversaries exploit software vulnerabilities to launch sophisticated attacks, challenging traditional def…
WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignor…
GR2 Technical Report
Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking --…
TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands…
Spectral Geometry and Bosonic-Bloch Probes: Explorations in Quantum Learning
This paper studies how spectral geometry emerges in quantum learning models and how it can be diagnosed with physically grounded probes. In…
Hate Speech Detection in Turkish and Arabic: A Comprehensive Study
Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings,…
Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval
The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage. Industry standards fo…
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival…
Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences
A single panorama captures the full visual sphere from one camera center, yet confines users to looking around in place without enabling tr…
From World Models to World Action Models: A Concise Tutorial for Robotics
World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities…
Post-Training Pruning for Diffusion Transformers
Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhe…
Learning Cardiac Motion Priors for Implicit Neural Representations
Implicit neural representations (INRs) are well suited to cardiac motion estimation, providing continuous, compact representations of motio…
Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around ab…
Diffusion-GR2: Diffusion Generative Reasoning Re-ranker
Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list,…
KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression
Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and in…
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluatio…
Spin-Weighted Spherical Harmonics Enable Complete and Scalable $\mathrm{E}(3)$-Equivariant Networks
$\mathrm{E}(3)$-equivariant networks are promising for 3D atomistic system modeling, yet their scalability is limited by the $O(L^6)$ compl…
MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for use…
DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents
Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns requir…
Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander
We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone. Choo…
AI Virtue: What is "Good" Knowledge in the Age of Artificial Intelligence?
In the age of AI, what will be good knowledge? This article, which is accepted and forthcoming in a special issue of Modern Fiction Studies…
kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail
Large language models (LLMs) are increasingly deployed in domains requiring guardrails to detect unsafe, off-topic, or adversarial prompts.…
Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies
Flow-matching vision-language-action policies generate robot action chunks through an iterative transport process, creating an opportunity…
ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning
We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uni…
GAP-GDRNet: Geometry-Aware Monocular Visual Pose Sensing on a Single-Target Synthetic Spacecraft Dataset
Monocular relative pose sensing is a central perception problem in non-cooperative rendezvous and on-orbit servicing. In spacecraft images,…
Understanding Agent-Based Patching of Compiler Missed Optimizations
Compiler missed optimizations refer to cases in which compilers failed to optimize certain code. It takes many compiler developers' efforts…
DemoPSD: Disagreement-Modulated Policy Self-Distillation
On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single mo…
【Fable 5に聞いてみた】10年後、ITエンジニアの仕事はどうなる? 他のAIモデルの回答とも比較
注目を集めている米AnthropicのAIモデル「Claude Fable 5」の性能はどれほどなのか。ITmedia AI+編集部で試した結果を紹介する。今回は、10年後のITエンジニアの仕事について聞いてみた。
「恐縮ですが」「それな」も自然に英訳 Sakana AI“温度感”伝える翻訳サービス公開 添削機能も
チャットサービス「Sakana Chat」の新機能で、ビジネス敬語のニュアンスや日本固有の文化概念、ネットスラングなど、既存の翻訳エンジンが苦手としてきた「日本語らしさ」の難所をクリアしたのが特徴だ。
Anthropic、AIの“内なる思考”が宿る「J-space」を発見──新手法「J-lens」で可視化、安全性監視に応用
Anthropicは、LLMの内部に人間の意識研究における「グローバルワークスペース」に類する構造「J-space」が自然発生していることを発見し、新手法「Jacobian lens」を公開した。モデルが言語化しない隠れた思考を一部観測でき、評価への気づきや不正コードの意図を検…
The ‘first’ AI-run ransomware attack still needed a human
An AI agent carried out the technical execution of a real-world ransomware attack for the first known time, but new details show a human st…
US investors will soon get access to SK Hynix, another memory maker riding the AI boom
SK Hynix is experiencing a boom credited to AI. It will ride that to a multibillion-dollar U.S. IPO, expected to take place on Friday.
AI誤決裁の責任は誰が取る? 自動化していい業務/ダメな業務の境界線
エイトレッドが「ワークフローのAI代替可能性に関する実態調査」の結果を公表。7割超の担当者が「ワークフローの承認・決裁をAIに任せるべきではない」と回答したことが明らかとなった。
1年でAfter AIの組織に生まれ変わったソラコム、「トークン資本」の安全な器へ
ソラコムが、2025年7月から始めた「After AIの組織」に移行するための取り組みを説明。この新たな組織体制から生み出したマネージドAIエージェントサービス「SORACOM Agent」も紹介した。
経営層の66%が「パスワード使い回し」 「管理が面倒」なのにツール未導入が約6割
企業のパスワード管理において、本来セキュリティ対策を主導すべき意思決定層ほど、対策の遅れが目立つ――。国内のビジネスパーソン1219人を対象に実施した「企業のパスワード管理に関する実態調査2026」で分かった。
AIコーディングの「ループ」4種類を完全入門 Anthropic公式が分かりやすく整理して解説
Anthropicが、Claude Codeにおける「ループ」を4種類に整理して解説した。AIコーディングで何をAIに任せ、どこで止めるべきかを、初心者にも分かるようにかみ砕き、筆者なりの視点も添えて紹介する。
「250万超AIエージェント作成」の裏で同時多発する課題を回し切ったIT部門の運用術
ソフトバンクは、生成AIサービスを数万人規模で全社展開しましたが、その裏側では従来のSaaS導入とは異なる課題が同時多発的に発生しました。それを支えたのは、切り分け・情報整理・関係者調整・粘り強い説明といったIT部門の基礎力と現場の経験です。この事例は、生成AI時代の企業ITが…
Vercel CEO Guillermo Rauch on the fight to split off models from agents
"The reality is, when you're optimizing for production, you start looking at a price/performance," Guillermo Rauch tells TechCrunch.
You can now customize Siri’s pace and expressivity in the latest iOS 27 beta
The update is part of Apple's broader effort to make Siri feel more natural and personal, as it rebuilds the assistant around generative AI.
Every major tech layoff in 2026 that has name-checked AI
A running look — in reverse chronological order — at the bigger tech companies that have announced significant layoffs this year with AI as…
If you use Google, you’re training its AI. Here’s how to opt out.
Consider this a belated PSA: A recent change to Google’s privacy settings is allowing the company to store more of your data, including med…
Microsoft lays off nearly 5,000 employees across Xbox, commercial sales
Microsoft cut around 4,800 roles, or 2.1% of its global workforce, on Monday — the latest in a series of layoffs that’s stoking fears of AI…
Reddit is using LLMs to solve a problem LLMs largely created
In the AI era, platforms have no choice but to fight fire with fire to cull spam.
2026-07-06(13件)
Station F ramps up as a launchpad for Europe’s hottest AI startups
Station F, a Paris-based startup hub founded by French billionaire Xavier Niel, is gearing up for a new edition of its F/ai accelerator pro…
パランティアCEOがOpenAIやAnthropicを批判 「AIモデルを強化するためのデータをなぜ顧客が渡すのか」
Palantirのアレックス・カープCEOがCNBCのインタビューで、OpenAIのサム・アルトマンCEOやAnthropicのダリオ・アモデイCEOの名前を挙げつつ、フロンティアモデル提供企業が顧客のデータを使って競争優位性を確保する業界の現状を批判した。対抗して打ち出したの…
「AIはもう十分賢い」 活用のボトルネックは「モデル性能」から「評価」へ 米データ基盤開発企業のAI研究者に聞く
AIの進化は学習データや計算資源の不足で止まるのか。米DatabricksのチーフAIサイエンティスト、ジョナサン・フランクル氏は「モデルはもう十分賢い」と語り、真のボトルネックは「AIを評価すること」にあると指摘する。
目が不自由な人の歩行をAIが音声で支援、歩行補助デバイス開発へ
サフラテクノは、AI画像認識技術で障害物を検知する歩行支援デバイスの試作機を開発している。独自の電力最適化技術により、6時間以上の連続駆動と軽量設計の両立を目指す。
AI推進にブレーキ? AWS「コスト抑制の動きある」 “トークン消費問題”への有効策は
AI利用を巡って、処理コストを抑えようとする動きが出てきている。AWSも把握しているという“トークンコスト問題”について、対策を聞いた。
「AIは依然として古い性能法則に従っている」 Tenstorrent Jim Keller氏
Tenstorrentは2026年4月に、同社のチップを搭載したサーバの性能を披露するイベントを開催した。Tenstorrent CEOのJim Keller氏は米国EE Timesのインタビューに答え、同社のサーバ「Galaxy」の性能や、競争環境について語った。
FDEとリコーの新コンサルサービス、どこが違う? AXのパートナー選びを考察
企業はAXに向けてどのようなベンダーをパートナーに選べばよいのか。リコーの新たなAIコンサル会社と最近よく聞く「FDE」との共通点と相違点を明らかにしつつ、考察する。
AIサーバで稼ぐシャープ 2030年度の新規事業売上高の8割強目指す 鴻海と連携
シャープの河村哲治社長は7月3日、大阪市内で報道陣の取材に応じ、目標とする2030年度の新規事業の年間売上高2000億~3000億円のうち、8割強はAI向けサーバ関連を想定していると明らかにした。親会社の鴻海(ホンハイ)精密工業の調達・製造力を活用し、シャープが国内販売や保守を…
キオクシア、新型メモリのサンプル出荷開始 岩手の工場の最先端設備活用
半導体大手キオクシアホールディングスが、大容量で低消費電力の新型メモリのサンプル品の出荷を開始したと発表した。北上工場(岩手県北上市)第2製造棟の最先端設備を活用して生産する。
アシックス、AIでシューズ設計を高度化
アシックスは、RebuilderAIと共同で、AIを活用した次世代シューズ設計/製造シミュレーション技術を公開した。デザイナーが描いた2Dスケッチやコンセプトを基に、AIが3Dデータを生成し、CAE/FEA解析へとつなげることで、デザインから検証までの期間短縮を目指す。
カクヤス「30年物の泥沼システム」をAIでどう解読? 現場が思いついた“ある考え”
動いてはいるものの、誰もその中身を把握できない。多くの企業が直面する基幹システムの“老い”に、カクヤスは生成AIを使った。鍵となったのは、AIによる解析と現場の業務知見を組み合わせ、「AIを制御する技術」を確立したことだ。
富裕層にいかに金を使わせる? ダイナースとニューオータニ「18万円超カード」の真意
内閣府のデータが示す「85歳でも資産が減らない」というデフレマインドに縛られた日本型富裕層。インフレ時代を迎え「資産を体験に変える」ためのコンパスとして、ダイナースクラブとホテルニューオータニが年会費18万円超のプレミアムカードを解禁した。「全員平等」を守ってきた老舗ホテルの方…
Amazon will stop accepting new customers for Mechanical Turk
These may be the last days of Amazon’s Mechanical Turk.