週次AIニュース 2026-W27
対象期間: 2026-06-29 〜 2026-07-05(1945 件)
トピックの推移
トピック別件数
- 研究/論文 738件
- LLM/生成AI 718件
- エージェント 476件
- 画像/動画生成 343件
- ロボティクス 134件
- ビジネス/資金調達 105件
- その他 49件
- ハードウェア/半導体 47件
- 規制/政策 11件
今週のハイライト(上位 10 件)
How ChatGPT adoption has expanded
New OpenAI Signals data shows how ChatGPT adoption is growing globally, with users increasing usage, exploring more capabilities, and drivi…
Mapping Europe’s AI Workforce Opportunity
A new OpenAI report maps how AI could reshape jobs across the EU, highlighting which occupations may face automation, growth, or workflow c…
HP Inc. launches Frontier strategic partnership with OpenAI
HP Inc. scales its OpenAI Frontier partnership to deploy AI across customer experiences, software development, and enterprise operations.
マイクロン、AI需要で広島工場増強へ起工式 1.5兆円投資
マイクロンメモリ ジャパンは2026年7月、広島工場の生産能力増強に向けた新クリーンルーム建設の起工式を開催した。AI技術の進展に伴うメモリ需要の増加に対応するもので、2028年後半に製造装置の搬入を開始する予定だ。広島工場には今後、この新クリーンルーム建設を含めて1兆5000…
Midjourney wants Hollywood studios to reveal the details of their AI usage
As part of an ongoing legal dispute with three Hollywood studios, Midjourney is seeking to compel those studios to reveal how they use AI t…
Alibaba reportedly bans employees from using Claude Code
Alibaba has reportedly classified Claude Code as high-risk software.
What is Mistral AI? Everything to know about the OpenAI competitor
Mistral AI, which offers some open source AI models, has raised significant funding since its creation in 2023, with the ambition to “put f…
米トランプ大統領、AI規制は「できるだけ介入少なく」 中国に対して開発競争「大幅リード」を強調
トランプ米大統領がCNBCのインタビューで、AI規制について「ガードレールは必要だが、介入はできるだけ少なくしたい」と述べ、規制最小限の路線を改めて示した。
全件(日付別)
2026-07-05(5件)
マイクロン、AI需要で広島工場増強へ起工式 1.5兆円投資
マイクロンメモリ ジャパンは2026年7月、広島工場の生産能力増強に向けた新クリーンルーム建設の起工式を開催した。AI技術の進展に伴うメモリ需要の増加に対応するもので、2028年後半に製造装置の搬入を開始する予定だ。広島工場には今後、この新クリーンルーム建設を含めて1兆5000…
New Google commercial imagines a Declaration of Independence written with help from AI
Two hundred and fifty years after the signing of the Declaration of Independence, a new commercial asks: What if the Founding Fathers had a…
Midjourney wants Hollywood studios to reveal the details of their AI usage
As part of an ongoing legal dispute with three Hollywood studios, Midjourney is seeking to compel those studios to reveal how they use AI t…
Alibaba reportedly bans employees from using Claude Code
Alibaba has reportedly classified Claude Code as high-risk software.
What is Mistral AI? Everything to know about the OpenAI competitor
Mistral AI, which offers some open source AI models, has raised significant funding since its creation in 2023, with the ambition to “put f…
2026-07-04(4件)
米トランプ大統領、AI規制は「できるだけ介入少なく」 中国に対して開発競争「大幅リード」を強調
トランプ米大統領がCNBCのインタビューで、AI規制について「ガードレールは必要だが、介入はできるだけ少なくしたい」と述べ、規制最小限の路線を改めて示した。
フィジカルAIに挑む日の丸連合、「Noetra」とは何か
2026年6月29日~7月3日に公開された記事の中から、MONOist編集部が厳選した今週の注目ニュースをお届けします。
The only AI glossary you’ll need this year
The rise of AI has brought an avalanche of new terms and slang. Here is a glossary with definitions of some of the most important words and…
The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari
We’ve compiled an overview of some of the top alternative browsers available today aiming to challenge Chrome and Safari.
2026-07-03(342件)
3万円で「Yahoo!ニュース」にPR掲載 プレスリリースをAIで「ニュース風記事」に
Yahoo!ニュース内に、企業の発表情報をニュース記事と同じフォーマットで掲載する。
「Claude Fable 5」をサブスクの標準機能に――AnthropicのエンジニアがXに投稿 7月8日以降の「早期復活目指す」
「Fable 5をサブスクリプションの標準機能として復活させることを目指している」――米AnthropicのITエンジニアであるタリク・シヒパー氏は、自身のXアカウントにこのように投稿した。
PACE: A Neuro-Symbolic Framework for Plausible and Actionable Counterfactual Explanations
Counterfactual explanations explain machine learning predictions by identifying minimal input changes that would alter a model's decision.…
Auto-FL-Research: Agentic Search for Federated Learning Algorithms
Federated learning (FL) research often depends on many small but consequential algorithmic choices: optimizer variants, server aggregation…
The Wiola Architecture for Efficient Small Language Models
We present Wiola, a fully original Small Language Model (SLM) architecture built from first principles, sharing no structural lineage with…
Agent4cs: A Multi-agent System for Code Summarization in Large Hierarchical Codebases
Understanding large, complex codebases, especially those with obfuscated structures and incomplete documentation, remains a significant cha…
When Should Service Agents Reconsider? Difficulty-Routed Control in Customer-Service Operations
Autonomous customer-service agents are shifting from conversational interfaces toward operational execution roles: they retrieve firm recor…
CreativityNeuro: Steering Language Model Weights to Improve Divergent Thinking and Reduce Mode Collapse
Divergent thinking is a crucial aspect of creativity, yet large language models (LLMs) tend to consistently generate similar responses to o…
Discrete Diffusion Language Models for Interactive Radiology Report Drafting
Diffusion language models, which generate text by denoising a token canvas bidirectionally instead of emitting tokens left to right, have b…
Beyond Next-Token Prediction: An RLVR Proof of Concept for Tool-Use Agents on Atlassian Workflows
Large language models are trained to predict the next token, not to act inside a specific API. In niche enterprise SaaS workflows -- where…
World Feedback for Clinical Agents: Diagnosing RL in FHIR Environments
Clinical protocol-execution tasks -- checking a lab value, applying a threshold, placing a correctly structured FHIR order -- are natural c…
Procedural Memory Distillation: Online Reflection for Self-Improving Language Models
Reinforcement learning with verifiable rewards (RLVR), along with recent selfdistillation variants such as SDPO, evaluates each rollout aga…
The Agentic Garden of Forking Paths
Empirical research rarely admits a unique analysis. Different analytical choices can lead to different conclusions from the same data, yet…
Janus: a Playground for User-Involved Agentic Permission Management
AI agents that autonomously execute tool calls on a user's behalf raise pressing questions about permission management: what role could use…
Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning
Chain-of-thought (CoT) reasoning has emerged as an effective approach for activating latent reasoning capabilities in large language models…
OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration
Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned wit…
Scaling Trends for Lie Detector Oversight in Preference Learning
Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cu…
EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation
Large language models have recently been explored for scientific hypothesis generation, but most prior work relies on unstructured literatu…
Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation
Developing high-performance kernels for Neural Processing Units (NPUs) is a critical industry bottleneck, requiring developers to manually…
Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model
As the scale and complexity of cloud-based AI systems continue to escalate, ensuring service reliability through rapid fault detection and…
SemHash-LLM: A Multi-Granularity Semantic Hashing Framework for Document Deduplication
Large scale document deduplication must preserve semantic equivalence while remaining efficient over massive corpora. We present SemHash LL…
Profit-Based Counterfactual Explanations for Product Improvement: A Case Study of Manga Sales in Japan
Counterfactual explanation (CE) is widely used to enhance the interpretability of machine learning models and support data-driven decision-…
Scaling with Confidence: Calibrating Confidence of LLMs for Adaptive Test Time Scaling
Training large language models (LLMs) with reinforcement learning (RL) has significantly advanced their performance on reasoning and questi…
Spatial Support Matters: Geometry-Aware Graph Fusion for Rainfall Field Reconstruction
Fine-scale rainfall reconstruction is critical for urban flood modeling, but real rainfall sensing systems observe the field through incomp…
Autonomous discovery of traffic laws with AI traffic scientists
Universal traffic laws describe recurrent patterns in congestion, mobility and driving behavior across cities, providing a scientific basis…
Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry
Multi-agent systems are increasingly used for forecasting future events, as deliberation among multiple LLMs is believed to improve reasoni…
Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment
In multi-source ECG deployment, models may need to incorporate new data sources when earlier raw ECGs cannot be retained or replayed. Freez…
Epistemic Goggles: A Pretrained Module that Induces an Epistemic Frame via Gradient Editing
Finetuning a language model on documents that are explicitly annotated as fictional results in a model that still actually believes the doc…
COMFYCLAW: Self-Evolving Skill Harnesses for Image Generation Workflows
Agents are increasingly used to construct workflows and assist humans in completing recurring tasks more efficiently. As these workflows be…
Generic Expert Coverage for Pruning SparseMixture-of-Experts Language Models
Sparsely activated Mixture-of-Experts (MoE) language models contain substantial structured redundancy among routed experts, but pruning the…
Distributionally Robust Listwise Preference Optimization
Existing robust preference optimization for language-model alignment mainly studies pairwise supervision and places robustness at the datas…
DRL-CLBA: A Clean Label Backdoor Attack for Speech Classification via DDPG Reinforcement Learning
Deep learning models for speech classification are vulnerable to backdoor attacks, where malicious triggers cause misclassification at infe…
Reformalization of the Jordan Curve Theorem
We present a case study in reformalization, a variant of autoformalization in which the input proof is not natural language but a formal de…
Meta-Benchmarks for Financial-Services LLM Evaluation
Public LLM leaderboards optimise for global average performance and do not capture the specific cognitive demands of financial-services wor…
Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation
On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader s…
Mastermind: Strategy-grounded Learning for Repository-Scale Vulnerability Reproduction
Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the inpu…
SimWorlds: A Multi-Agent System for Dynamic 3D Scene Creation
LLM agents are increasingly used to translate natural language into 3D scenes in a procedural way, but existing systems focus on static out…
Repair the Amplifier, Not the Symptom: Stable World-Model Correction for Agent Rollouts
As agent planning moves from short tool chains toward persistent workflows with thousands or tens of thousands of steps, failures will occu…
Verifiable Knowledge Expansion through Retrieval-Grounded Formal Concept Analysis
Ontology construction requires deciding which objects, attributes, and structural relations should be accepted as valid knowledge. Language…
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based…
Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing…
MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support
Traditional Chinese Medicine (TCM) diagnosis, particularly through tongue inspection, faces persistent challenges in subjectivity and repro…
Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge
Large language models (LLMs) are increasingly proposed for aviation business operations, from documentation and training generation to cust…
Actual causality in fault trees
Fault trees are a widely used as effective risk models for complex systems, answering the question "what can go wrong?", especially through…
CLAP: Closed-Loop Training, Evaluation, and Release Control for Domain Agent Post-training
Domain agents often face noisy business data, uncertain post-training gains, offline/application mismatch, and adapter-release risk. This p…
Safety Targeted Embedding Exploit via Refinement
Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms genera…
CamoNAS: Neural Architecture Search for Enhanced Camouflaged Object Detection
Camouflaged Object Detection (COD) aims to locate and segment objects that blend into their surroundings, presenting challenges due to weak…
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routi…
Spec-AUF: Accept-Until-Fail Training under Train-Inference Misalignment for Masked Block Drafters
Speculative decoding accelerates autoregressive generation by drafting a block of tokens that the target model verifies left-to-right, comm…
Rethinking Complexity Metrics for LLM-Integrated Applications: Beyond Source Code
LLM-integrated applications blend natural language prompts with program code, and much of their runtime behavior originates in the prompt l…
ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair
Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad search…
ElephantAgent: Contextual State Continuity in Agentic Systems
Agentic systems enhance their capabilities by invoking external tools and maintaining persistent memory. However, these external dependenci…
A-TMA: Decoupling State-Aware Memory Failures in Long-Term Agent Memory
Long term memory lets LLM agents act as persistent assistants, but user facts change. A useful memory system must know what is true now, wh…
Atomic Task Graph: A Unified Framework for Agentic Planning and Execution
LLM-based agents have shown strong potential for solving complex multi-step tasks, yet existing performance improvements often rely on eith…
OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models
Ontology learning (OL) aims to automatically construct structured knowledge models from text, yet progress remains fragmented across method…
Multimodal Knowledge Edit-Scoped Generalization for Online Recursive MLLM Editing
Online multimodal knowledge editing requires injecting a continual stream of visual-textual corrections into multimodal large language mode…
Episodic-to-Semantic Consolidation Without Identity Drift
Long-running adaptive intelligent agents face a structural tension between knowledge consolidation and information integrity. Memory consol…
Traceable Fault Diagnosis for Battery Energy Storage Systems via Retrieval-Augmented Multi-Agent O&M Assistant
Large-scale battery energy storage systems (BESSs) require O&M decisions that combine alarms, cell-level measurements, device topology, dia…
InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories
Multimodal large language models must adapt to evolving tasks and domains, yet continual improvement under bounded deployment footprint rem…
Hidden Forgetting in Continual Multimodal Learning: When Accuracy Survives but Grounding Fails
Multimodal large language models must continually adapt to evolving tasks and domains, yet standard continual learning metrics mainly measu…
PACE: A Proxy for Agentic Capability Evaluation
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single…
Algebraic Model Counting for Global Analysis of Optimal Decision Trees
Ensuring model reliability in Explainable AI requires a global assessment of the hypothesis space. We propose a formal framework for the ex…
Evidence-State Rewards for Long-Context Reasoning
Long-context reasoning requires models to locate, revise, and synthesize evidence distributed across lengthy inputs. Existing long-context…
SUNTA: Hierarchical Video Prediction with Surprise-based Chunking
Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks.…
ContextNest: Verifiable Context Governance for Autonomous AI Agent
Autonomous AI agents increasingly depend on external knowledge stores, yet most retrieval pipelines provide relevance without durable guara…
Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training
Scientific Fitness Coaching (SFC) is typically delivered by human professionals, making it costly and inaccessible to many. While recent ad…
Coding-agents can replicate scientific machine learning papers
Scientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that t…
A$^{2}$utoLPBench: An Auto-Generated, Agent-Friendly LP Benchmark via Inverse-KKT Construction
Most LP-from-text benchmarks are static datasets of word problems written and labeled by hand. Once such a dataset is released, its size is…
A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks
Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open…
UA-ChatDev: Uncertainty-Aware Multi-Agent Collaboration for Reliable Software Development
Software development is a complex task that demands cooperation among agents with diverse roles. Large language models (LLMs) have enabled…
Criticality-Based Guard Rail Validation for AI Agent Decisions in Autonomous Telecom Networks
The evolution toward fully autonomous telecommunications networks (Autonomous Network Levels 4-5) requires AI/ML agents to make real-time n…
Purified OPSD: On-Policy Self-Distillation Without Losing How to Think
On-policy self-distillation (OPSD) has emerged as a promising paradigm for improving LLM reasoning, where a privileged teacher with access…
Copewell: A Multi-Agent Swarm Architecture for Equitable Mental Wellness Support
Mental health disorders affect nearly one billion people globally, yet 75% of individuals in low- and middle-income countries receive no tr…
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see. The simplest contract appends past obs…
A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets
Linear-attention and state-space language models compress the prefix into a fixed-size recurrent state, yielding O(1) memory at the cost of…
Grounded autonomous research: a fault-tolerant LLM pipeline from corpus to manuscript in frontier computational physics
Autonomous-research agents have demonstrated end-to-end LLM automation in machine-learning sandboxes where execution provides calibration.…
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response.…
Hardware-Enforced Semantic Coordination for Safety-Critical Real-Time Autonomous Systems
Recent advances in agentic AI are producing increasingly complex autonomous systems that integrate large language models, world models, opt…
Steerability via constraints: a substrate for scalable oversight of coding agents
Coding agents are capable; human oversight is the bottleneck. Unconstrained agents introduce security risks, erode codebase scalability, an…
Fast Multi-dimensional Refusal Subspaces via RFM-AGOP
Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work…
Text-Driven 3D Indoor Scene Synthesis in Non-Manhattan Environments
Large Language Models (LLMs) have demonstrated remarkable capabilities in 3D indoor synthesis for Manhattan environments. However, existing…
Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach
Scalable and reliable grading of command-line examinations remains a challenge in computing education, where rising enrolments make manual…
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this p…
G-RRM: Guiding Symbolic Solvers with Recurrent Reasoning Models
In this work, we focus on SE-RRMs, a symbol-equivariant instantiation of RRMs that exhibits improved extrapolation to larger problem sizes.…
What LLM Agents Say When No One Is Watching: Social Structure and Latent Objective Emergence in Multi-Agent Debates
LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageou…
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning
Understanding and reasoning over long contexts has become a key requirement for deploying large language models (LLMs) in realistic applica…
Online Safety Monitoring for LLMs
Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an ala…
Distributed Attacks in Persistent-State AI Control
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This per…
TokenScope: Token-Level Explainability and Interpretability for Code-Oriented Tasks in Large Language Models
Understanding how Large Language Models (LLMs) make token-level decisions during code generation remains a major challenge for both researc…
Safeguarding LLM Agents from Misalignment through Provenance Analysis
As LLM agents gain increasing access to powerful tools, ensuring that their actions are aligned with the user's intent becomes critical. Wh…
Kara: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression
Reasoning language models often generate long chain-of-thought (CoT), which accumulates a massive KV cache during the decoding phase and in…
SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings
Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-t…
Breaking Safety at the Token Boundary: How BPE Tokenization Creates Exploitable Gaps in LLM Alignment
Character-level perturbations bypass safety alignment in modern LLMs despite leaving prompts human-readable. We identify and test a central…
Prompt Framing Distorts Count-Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring
Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corr…
Mapping Text to Multiplex Graph: Prompt Compression as L\'evy Walk-Guided Graph Pruning
Existing prompt compression methods treat text as flat token sequences, failing to capture the distributed nature of important information,…
ExPerT: Personalizing LLM Responses to Users' Domain Expertise via Query-Wise Semantic and Keystroke Behavioral Cues
Large language models (LLMs) are increasingly used by end users, yet existing personalization methods relying on static profiles or text-on…
Office Comprehension Benchmark
We introduce Office Comprehension Bench (OCB), the first public benchmark to jointly evaluate LLM systems on Word, Excel, and PowerPoint co…
LLMs as Teaching Assistants for Mathematics Exam Grading: Reliability, and Practical Usability
Open-ended mathematics exams are valuable because they assess reasoning, proof construction, algorithmic thinking, and communication of int…
A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-Rational Cognition, and Governance of AI-Generated Content
Large language models are increasingly used for knowledge acquisition, code generation, academic writing, and agent-based automation. In th…
Structuring the Space of Sociotechnical Alignment
Sociotechnical alignment concerns the social desirability of AI behavior and is thus inherently normative, not merely technical. While NLP…
Collaborative Disagreement Resolution for Scalable Oversight
Debate, where AI agents argue opposing positions, has emerged as a key approach to scalable oversight. However, debate faces a fundamental…
How Indian Dermatologists are Utilizing Artificial Intelligence for Clinical Practice and Workflow Management: A Nationwide Survey with a Special Focus on atopic dermatitis
Background: Dermatology AI has mainly focused on image-based diagnosis, while chronic disease workflows have received less attention. We su…
Beyond Detection: Redesigning Assessment and Governande of Generative AI at the Universidad Polit\'ecnica de Madrid (UPM)
Universities have responded to generative artificial intelligence (GenAI) in noticeably different ways, both internationally and within Spa…
AI Assistance for Human Review of Default Judgments
Overwhelmed courts in the United States review millions of default judgments each year. Unfortunately, such manual reviews are time-consumi…
Artificial Intelligence-Enabled Accounting Information Systems and Fraud Detection in Nigeria's Financial Services Sector: The Moderating Role of Natural Language Processing
The rapid digitalisation of financial systems has improved operational efficiency and financial inclusion while simultaneously increasing e…
The Rising Unsustainability of AI Graphics Cards Production
The rapid advancement of Artificial Intelligence (AI) has been accompanied by significant increases in computational and environmental cost…
Benchmarking Federated Learning and Knowledge Distillation for Point Cloud Classification
Deploying 3D point cloud analysis in privacy-sensitive, resource-constrained settings faces two barriers: data cannot be centralized, and m…
Domain Knowledge Based Temporal-Spatial Graph Convolution Network for ECG Recognition
In light of strides in Arti cial Intelligence (AI) and its wide spread application, challenges persist in the interpretability of AI models…
Scaling Laws for Grid-Based Approximate Nearest Neighbor Search in High Dimensions
Grid-based approaches to approximate nearest neighbor (ANN) search have been absent from modern scaling analyses. We present a systematic c…
Adaptive Companionship for Group-Following Robots: Handling Dynamically Changing Group Formations
Accompanying a group of humans is an essential aspect of developing human-like social cognition in robots. However, human groups typically…
CPG-PAD: Concept-Informed Prompts Guided Presentation Attack Detection
Presentation Attack Detection (PAD) serves as a crucial safeguard for face recognition systems against presentation attacks such as printed…
Generative AI and Federated Learning for Intrusion Detection Systems: A Survey
Intrusion Detection Systems (IDSs) are essential for monitoring network traffic and identifying malicious activities in modern cyber-physic…
Black-Box Inference of LLM Architectural Properties with Restrictive API Access
In practice, most commercial LLM providers do not publicly release details of underlying LLM architectures. However, prior work has shown t…
Mechanistic Interpretability and Causal Feature Steering of Neural Quantum States via Sparse Autoencoders
Neural Quantum States (NQS) are a remarkably expressive class of variational ans\"atze for quantum many-body wavefunctions, yet little is u…
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
Turn-taking naturalness is central to full-duplex spoken dialogue systems, yet its automatic evaluation remains limited. Existing evaluatio…
Multi-modal Rail Crossing Safety Analysis
Given one or more images of a railway crossing, can we leverage visual cues that allow us to robustly estimate how safe it is? Can we impro…
AI-enabled gravitational-waves searches for binary neutron stars at optimal sensitivity
Gravitational Waves (GWs) represent the newest window of astronomy, furthering our understanding of compact objects like black holes and ne…
How Should Transformers Encode Numeric Values in Electronic Health Records?
How do we encode numeric values in transformer-based sequence processing, particularly in electronic health record (EHR) data? We systemati…
Rethinking Generic Object Tracking Toward Human-Level Perceptual Intelligence
At the heart of human visual perception lies the ability to maintain a continuous and coherent understanding of the external world. By inte…
NeuroBridge: Bridging Multi-Task MRI Knowledge for Neurodegenerative Disease Diagnosis
INTRODUCTION: Accurate MRI-based identification of Alzheimer's disease (AD), mild cognitive impairment (MCI), and related dementias remains…
Spin-Weighted Spherical Harmonics Enable Complete and Scalable $\mathrm{E}(3)$-Equivariant Networks
$\mathrm{E}(3)$-equivariant networks are promising for 3D atomistic system modeling, yet their scalability is limited by the $O(L^6)$ compl…
GPUAlert: A Zero-Instrumentation Process-Boundary Monitor for Diagnosing GPU Training-Job Failures
GPU training jobs fail often, roughly two in five on large production clusters, yet the operator typically learns of a failure only by reco…
Adoption and Impact of Command-Line AI Coding Agents: A Study of Microsoft's Early 2026 Rollout of Claude Code and GitHub Copilot CLI
Organizations rolling out agentic command line tools like Anthropic's Claude Code and GitHub's Copilot CLI need to know who will try them,…
MultAttnAttrib: Training-Free Multimodal Attribution in Long Document Question Answering
As grounded QA systems are increasingly deployed in AI assistants, accurately attributing generated answers to evidence is critical for use…
Risk Architecture for AI-Native Engineering Teams: An Organizational Framework for Agentic System Governance
Engineering management research has produced mature frameworks for software risk: ownership by feature, escalation by severity, and assuran…
IsoSci: A Benchmark of Isomorphic Cross-Domain Science Problems for Evaluating Reasoning versus Knowledge Retrieval in LLMs
We introduce ISOSCI, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge re…
On the Utility and Factual Reliability of Pruned Mixture-of-Experts Models in the Biomedical Domain
Mixture-of-Experts (MoE) models offer inference speedups via selective activation but impose substantial memory requirements because the wh…
Token Geometry
Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface be…
Grounded Optimization: A Layered Engineering Framework for Reducing LLM Hallucination in Automated Personal Document Rewriting
Large language models (LLMs) are increasingly applied to resume optimization for applicant tracking systems, introducing hallucination fail…
Fully Unsupervised Detection of Physical Contacts on Subsea Cables via State-of-Polarization Monitoring
We present a fully unsupervised Fast-Slow DSVDD detector for continuous State-of-Polarization monitoring on a deployed subsea cable. Traine…
Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL
Reinforcement learning post-training dramatically improves LLM reasoning, but suffers from training instability and diversity collapse. Adv…
Robust and Explainable 3D Mode Shape Recognition Using Region-Aware Graph Neural Networks
Mode shape recognition is a fundamental task in automotive NVH development, yet it remains dependent on manual visual inspection by experie…
Multi-Head Recurrent Memory Agents
Recurrent memory agents extend LLMs to arbitrarily long contexts by iteratively consolidating input into a fixed-size memory window. Despit…
IntentTune: Using user demand and personalization to resolve "unknown" query intents for e-commerce search
Understanding user intent is fundamental to delivering relevant search results in e-commerce. However, substantial fraction of real-world q…
Evolutionary Feature Engineering for Structured Data
Large language models are increasingly used as open-ended search operators in evolutionary optimization. We introduce Evolutionary Feature…
X-LogSMask: Expand Transformer for Graph-Structured Data
Transformers have become general-purpose architectures, but their all-to-all self-attention is poorly matched to graph data, whose interact…
DiPS: Dialogue Policy Selection for High-Stakes Persuasion Agents
Large Language Models (LLMs) often struggle with persuasion in high-stakes scenarios. People's individual personalities and concerns requir…
ADVENT: LLM-Driven Automatic Predicate Invention for ILP
Predicate invention (PI), the creation of new predicates to extend the hypothesis space, remains a critical bottleneck in Inductive Logic P…
VLAFlow: A Unified Training Framework for Vision-Language-Action Models via Co-training and Future Latent Alignment
Vision-language-action models (VLAs) have recently advanced robotic manipulation, yet the effects of different robot-data pre-training para…
MKGR: Multimodal Knowledge-Graph Representation Learning for Cold-Start Protein-Protein Interaction Prediction
Accurate protein-protein interaction (PPI) prediction is central to functional genomics, disease mechanism discovery, and drug development.…
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated i…
Beyond Gradient-Based Attacks: Adversarial Robustness and Explainability Stability in Cybersecurity Classifiers
Adversarial attacks on cybersecurity classifiers pose a dual threat: degrading predictions and destabilising the SHAP-based explanations th…
Model Merging as Probabilistic Inference in Fine-Tuning Parameter Space
Model merging aims to combine existing single-task solutions into a multi-task solution without additional data-driven fine-tuning.~Most ex…
Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack
Recently, speech classification methods have gained widespread adoption in intelligent gadgets. Current study indicates that backdoor attac…
Predicting Closed-Loop Performance of Latent World Models: Offline Checkpoint Selection for MPC and Model-Based RL Under Non-Markovian Rewards in LunarLander
We study how to predict the downstream closed-loop performance of a learned latent world model from validation-time diagnostics alone. Choo…
Full Bayesian Reinforcement Learning via LF-IBIS
Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an…
MedStreamBench: A Time-Aware Benchmark for Streaming and Proactive Medical Video Understanding
Existing medical video benchmarks primarily evaluate whether a model produces the correct answer, but rarely assess whether it answers at t…
Decentralized Stochastic Subgradient-type Methods with Communication Compression for Nonsmooth Nonconvex Optimization
In this paper, we consider the nonsmooth nonconvex decentralized optimization problem, where inter-agent communication is compressed. We pr…
ProCal: Inference-Time Proposal Calibration for Open-Vocabulary Object Detection
Open-vocabulary object detection aims to localize and classify objects beyond the fixed set of categories seen dur ing training. Recent ope…
AI Virtue: What is "Good" Knowledge in the Age of Artificial Intelligence?
In the age of AI, what will be good knowledge? This article, which is accepted and forthcoming in a special issue of Modern Fiction Studies…
Scene-Conditioned PINN-GNN for Multipath RF Maps: Cross-Scene Generation and In-Scene Completion
Radio frequency (RF) maps provide a compact representation of multipath propagation characteristics and are fundamental to channel modeling…
EPnG: Adaptive Expert Prune-and-Grow for Parameter-Efficient MoE Fine-tuning
Mixture-of-Experts (MoE) models scale efficiently but remain costly to adapt due to redundant experts and uniform parameter allocation. Exi…
Lightweight Safe Reinforcement Learning for End-to-End UAV Navigation
With the rapid development of autonomous aerial systems, Unmanned Aerial Vehicles (UAVs) are increasingly deployed in applications such as…
Single-Channel EEG-Based Cognitive Load Assessment in Online Learning: A Hybrid Deep Learning Approach
Monitoring cognitive load during online learning could help instructors identify content that learners find difficult, but remote settings…
Expander Sparse Autoencoders: Parameter-Efficient Dictionaries for Mechanistic Interpretability
Sparse autoencoders (SAEs) decompose internal activations of neural networks into sparse linear combinations of learned features by fitting…
Decoupling Code Complexity from Newcomer Participation: A Causal Study of AI Coding Agent Adoption in OSS
Open-source projects depend on a steady inflow of newcomers. A growing concern is that AI coding agents (tools such as Cursor and Claude Co…
MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models
Evaluation benchmarks are essential for assessing vision-language models (VLMs), but most multimodal benchmarks are static, making them vul…
Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and spe…
Decomposer: Learning to Decompile Symbolic Music to Programs
Musical performance involves executing a set of high-level musical instructions, yet recovering those instructions from the performance is…
Evaluating Chunking Strategies for Retrieval-Augmented Generation on Academic Texts
Retrieval-Augmented Generation (RAG) systems use the question-answering capabilities of Large Language Models (LLMs) to access information…
Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map
Can a platform tell, before deployment, whether an open-weight checkpoint has had its refusal mechanism stripped? Runtime guards cannot: th…
An Exploratory Study on LLM-Generated Code and Comments in Code Repositories
The use of LLMs in software development has become increasingly widespread on tasks such as code generation and summarization. Reports from…
SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models
Large Vision-Language Models (LVLMs) have achieved remarkable progress in multimodal understanding, yet their enormous parameter scale and…
Rank-Then-Act: Reward-Free Control from Frame-Order Progress
We introduce Rank-Then-Act (RTA), a framework for learning control policies from expert video demonstrations without environment rewards. R…
SABER: A Semantic-Aligned Brain Network Analysis Framework via Multi-scale Hypergraphs
Effective brain disease diagnosis requires the synergy of brain connectivity patterns and high-level semantic knowledge. Existing methods,…
Population-Based Multi-Objective Training of Discriminators for Semi-Supervised GANs
Semi-supervised generative adversarial networks (SSL-GANs) can exploit large unlabeled datasets while retaining a classifier in the discrim…
Low-Latency Task-Oriented Image Transmission with Opportunistic Spectrum Access
Communication systems designed for reliable data reconstruction, rather than task-oriented communication, typically rely on separate source…
TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B
This paper presents TUDUM (T\"urk\c{c}e D\"u\c{s}\"unen \"Uretken Model), a project pipeline for adapting a Qwen-family 27B thinking model…
AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations
This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical…
CausalSteward: An Agentic Divide-Conquer-Combine Copilot for Causal Discovery
Learning causal models from high-dimensional data is a significant challenge, particularly in real-world settings where violations of core…
PhysMani: Physics-principled 3D World Model for Dynamic Object Manipulation
Manipulating fast and dynamically moving targets in unstructured 3D environments remains challenging for embodied AI. Existing visual-langu…
Conditional Co-Ablation: Recovering Self-Repair Backups in Transformer Circuits
Mechanistic interpretability often relies on component-level interventions to discover how a model produces a behavior. This guides attribu…
Robust for the Wrong Reasons: The Representational Geometry of LLM Robustness to Science Skepticism
Large language models (LLMs) are increasingly consulted on contested scientific questions, raising the concern that they will sycophantical…
NeoMap: Training-free Novel-View Synthesis from Single Images and Videos
We study the challenging problem of novel view video synthesis from single images or monocular videos. Existing methods, which operate unde…
Object Aligner: A Configurable JSON Schema Similarity Score for Graphs, Applied to LLM Prompt Optimization
Large language models (LLMs) are often asked to produce JSON conforming to a fixed schema, powering information extraction, tool calling, a…
Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias
Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual questi…
A Multi-Branch Hierarchy-Aware Framework for Heterogeneous Audio Classification
This technical report describes our system for Task 1 of the DCASE 2026 Challenge, which aims to classify heterogeneous audio recordings ac…
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding
Using molecular large language models (LLMs) as a unified framework for understanding molecular structures and functions is emerging as a n…
Do Newer Lightweight CNNs Perform Better Under Resource Constraints? A Controlled Multigenerational Study of Architecture, Initialization, Training Budget, and Efficiency
Newer lightweight convolutional neural networks are often presented as improving predictive performance and deployment efficiency, but such…
Mirror Illusion Art
Mirror Illusion Art is a novel reflection-conditioned 3D illusion where one object yields two target appearances (front and mirror). The ta…
Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving
Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates…
OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets
Safe completion requires models to provide useful assistance without enabling harm, but this behavior is difficult to evaluate with isolate…
SPLIT: Cross-Lingual Empathy and Cultural Grounding in English and Ukrainian LLM Responses
Large Language Models are increasingly deployed in emotional-support contexts and crisis-related situations. Nevertheless, their cross-ling…
Beyond the Performance Illusion: Structure-Aware Stratified Partitioning and Curriculum Distributionally Robust Optimization for Spatially Correlated Domains
Performance evaluation in AI systems commonly assumes that random dataset splits produce independent and identically distributed (i.i.d.) s…
Prompt Coverage Adequacy
In recent years, it has become increasingly evident that large language models (LLMs) and autonomous agents raise the level of abstraction…
SA-HGNN: Sample-Adaptive Hyperbolic Graph Neural Network for EEG-Based Depression Recognition
Graph Neural Networks (GNNs) have been widely used to capture spatial functional connectivity patterns to improve electroencephalography (E…
kNNGuard: Turning LLM Hidden Activations into a Training-Free Configurable Guardrail
Large language models (LLMs) are increasingly deployed in domains requiring guardrails to detect unsafe, off-topic, or adversarial prompts.…
Evolutionary Wave Function Collapse
Wave Function Collapse (WFC) is a widely used procedural content generation method that learns local adjacency constraints from example inp…
ESC: Emotional Self-Correction for Reliable Vision-Language Models
Vision-language models (VLMs) have achieved strong performance across diverse multimodal tasks, yet they remain vulnerable to unreliable re…
Guided Action Flow: Q-Guided Inference for Flow-Matching Vision-Language-Action Policies
Flow-matching vision-language-action policies generate robot action chunks through an iterative transport process, creating an opportunity…
An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation
While Large Multimodal Models excel in comprehension, high-throughput inference engines lack native support for multimodal generation. This…
Behind the Refusal: Determining Guardrail Activation via Behavioral Monitoring
As Large Language Models (LLMs) and agentic systems become integrated into real-world applications, ensuring their safety and security is c…
ART for Diffusion Sampling: Continuous-Time Control and Actor-Critic Learning
We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uni…
Predicting Early Stages Of Alzheimer's Disease And Identifying Key Biomarkers Using Deep Artificial Neural Network And Ensemble Of Machine Learning Methodologies
Alzheimers disease (AD) is a brain disorder that develops slowly and mainly affects memory, thinking, language, and daily activities. It is…
Dynamic Neural Graph Encoding of Inference Processes in Deep Weight Space
The rapid advancements in using neural networks as implicit data representations have attracted significant interest in developing machine…
RadiomicNet: A Hybrid Radiomics-Guided Lightweight Architecture for Interpretable Medical Image Segmentation
Deep learning has achieved remarkable performance in medical image segmentation, yet it suffers from critical limitations: mathematical int…
Overview of Risk Assessment and Management for Intelligent Systems under the AI Act and Beyond
The society and emerging risk-based regulatory frameworks for AI underscore the need for rigorous risk assessment to ensure safe and reliab…
What Types of Human-AI Teams Exist?
Human-AI teaming has received increasing attention in the literature. However, the range of studies conducted in multiple domains make it d…
The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains…
CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation
Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction d…
Efficient Waste Sorting for Circular Economy: A Confidence-guided comparison between One-Vs-All and One-Vs-Rest Classification Strategies with Human-in-the-Loop for Automated Waste Sorting
The complexity of waste disposal regulations across European countries poses significant challenges for the residents and hinders the trans…
Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages
LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional…
HERMES: A Multi-Granularity Labeling Substrate for Pre-training Data Mixtures
Most data-mixing methods assume the corpus has already been partitioned into groups, and the choice of those groups determines what a mixer…
AnyGroundBench: A Specialized-Domain Benchmark for Video Grounding in Vision-Language Models
Vision-Language Models (VLMs) have demonstrated immense promise in Spatio-Temporal Video Grounding (STVG). However, current evaluation prot…
Generalization in offline RL: The structure is more important than the amount of pessimism
While pessimism counteracts overestimation bias in offline reinforcement learning (RL), being overly conservative has been associated with…
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
Humans can selectively attend to a target sound and estimate its direction in complex scenarios, whereas such selective localization remain…
Self-Gating Attention for Efficient Time Series Forecasting
Transformer architectures have shown strong potential in time series forecasting, where multi-head self-attention is widely used to capture…
SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces
Large Language Model (LLM)-based agents increasingly automate software engineering tasks through reusable skills, natural-language instruct…
GAP-GDRNet: Geometry-Aware Monocular Visual Pose Sensing on a Single-Target Synthetic Spacecraft Dataset
Monocular relative pose sensing is a central perception problem in non-cooperative rendezvous and on-orbit servicing. In spacecraft images,…
Stable Self-Modulating Quantum Fast-Weight Programmers with Bounded Memory Gates
Quantum Fast-Weight Programmers (QFWPs) store temporal information in dynamically programmed variational-circuit parameters rather than in…
The Dual Nature of LLM Persona: Aggregated Tendencies and Frame-Dependent Geometry
Evaluations of LLM personas via psychometric questionnaires typically rely on aggregate scores, discarding within-instance correlation stru…
World Wide Models: Literary Tools for Cultural AI
LLMs stage a new form of cultural encounter that is massive, automated, and monolingual. Literary disciplines have always negotiated cultur…
Understanding Agent-Based Patching of Compiler Missed Optimizations
Compiler missed optimizations refer to cases in which compilers failed to optimize certain code. It takes many compiler developers' efforts…
VisionAId: An Offline-First Multimodal Android Assistant for People with Visual Impairment, Featuring Personalized Object Retrieval
Over 285 million people worldwide live with a visual impairment, for whom everyday tasks such as avoiding obstacles, locating personal belo…
ACID: Action Consistency via Inverse Dynamics for Planning with World Models
Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control. However, the standard plann…
Neuron-Aware Active Few-Shot Learning for LLMs
Active Few-Shot Learning (AFSL) adapts LLMs to specialized domains by identifying the most valuable unlabeled samples for annotation and us…
QFedAgent: Quantum-Enhanced Personalized Federated Learning for Multi-Agent Activity Recognition
Federated learning (FL) enables collaborative model training across distributed devices without sharing raw data, making it suitable for pr…
WorldSample: Closed-loop Real-robot RL with World Modelling
Reinforcement learning (RL) can overcome the demonstration-coverage limitation of imitation learning (IL) by allowing robots to improve thr…
Reasoning effort, not tool access, buys first-try reliability in agentic code generation: an observational study
Agentic coding assistants are increasingly given extra capabilities, such as browser based testing tools and design oriented system prompts…
Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation
Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, partic…
OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter coun…
Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs
Vision-Language-Action (VLA) models are fundamentally bottlenecked by the scarcity of expert demonstrations -- triplets of observations, in…
Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting
Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money prediction market (Polymar…
TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution
Software tests and code evolve together: a code change should be followed by new or updated tests that record the new software behavior. Ye…
Combating Textual Noise and Redundancy: Entropy-Aware Dense Visual Token Pruning
Visual token pruning is a crucial strategy for accelerating VLMs by compressing redundant image patches, yet existing methods often fail to…
Beyond Adam: SOAP and Muon for Faster, Label-Efficient Training of Machine Learning Interatomic Potentials
Machine learning interatomic potentials (MLIPs) have become a hallmark of AI for scientific simulation. While efforts on new architectures…
DemoPSD: Disagreement-Modulated Policy Self-Distillation
On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single mo…
Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas
Long-form TV dramas present a formidable challenge for comprehensive video understanding, where deciphering complex storyline often relies…
Program-as-Weights: A Programming Paradigm for Fuzzy Functions
Many everyday programming tasks resist clean rule-based implementation, such as alerting on important log lines, repairing malformed JSON,…
LACUNA: A Testbed for Evaluating Localization Precision for LLM Unlearning
LLMs memorize sensitive training data, including personally identifiable information (PII), creating a pressing need for reliable post hoc…
Interpreting Global Perturbation Robustness of Image Models using Axiomatic Spectral Importance Decomposition
Perturbation robustness evaluates the vulnerabilities of models, arising from a variety of perturbations, such as data corruptions and adve…
Causal Explanations for Image Classifiers
Existing algorithms for explaining the output of image classifiers use different definitions of explanations and a variety of techniques to…
MAGIK: Mapping to Analogous Goals via Imagination-enabled Knowledge Transfer
Humans excel at analogical reasoning - applying knowledge from one task to a related one with minimal relearning. In contrast, reinforcemen…
ADMC: Attention-based Diffusion Model for Missing Modalities Feature Completion
Multimodal emotion and intent recognition is essential for automated human-computer interaction, It aims to analyze users' speech, text, an…
Psychological Imagination Networks Show Cross-Population Centrality and Clustering Alignment in Humans That Large Language Models Fail to Replicate
Mental imagery vividness is a stable individual trait, yet whether imagined scenarios share relational structure across human and synthetic…
Aria: An Agent For Retrieval and Iterative Auto-Formalization via Dependency Graph
Accurate auto-formalization of theorem statements is essential for advancing automated discovery and verification of research-level mathema…
BuilderBench: The Building Blocks of Intelligent Agents
Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the lim…
Ophiuchus: Incentivizing Tool-augmented "Think with Images" for Joint Medical Segmentation, Understanding and Reasoning
Recent medical MLLMs have made significant progress in generating step-by-step textual reasoning chains. However, they still struggle with…
HAL: Inducing Human-likeness in LLMs with Alignment
Aligning language models to qualitative behavioral traits, such as human-likeness, remains difficult because they are hard to define, measu…
A General Neural Backbone for Mixed-Integer Linear Optimization via Dual Attention
Mixed-integer linear programming (MILP) is a foundational framework for combinatorial optimization across science and engineering, but rema…
Memora: A Harmonic Memory Representation Balancing Abstraction and Specificity
Agent memory systems must accommodate continuously growing information while supporting efficient, context-aware retrieval for downstream t…
BRIDGE: Predicting Human Task Completion Time From Model Performance
Evaluating the real-world capabilities of AI systems requires grounding benchmark performance in human-interpretable measures of task diffi…
PreScience: A Dataset and Benchmark for Scientific Forecasting
Can AI systems trained on the existing scientific record forecast the advances that will follow? We introduce PreScience, a dataset and ben…
OmniGAIA: Towards Native Omni-Modal AI Agents
Human intelligence naturally intertwines omni-modal perception -- spanning vision, audio, and language -- with complex reasoning and tool u…
Learning-based Multi-agent Race Strategies in Formula 1
In Formula 1, race strategies are adapted according to evolving race conditions and competitors' actions. This paper proposes a reinforceme…
Conformal Policy Control
An agent must try new behaviors to explore and improve. In high-stakes environments, an agent that violates safety constraints may cause ha…
A Dual-Helix Governance Approach Towards Reliable Agentic Artificial Intelligence for WebGIS Development
WebGIS development requires consistency, yet agentic AI often fails due to LLM context constraints, forgetting, stochasticity, instruction…
Formal Semantics for Agentic Tool Protocols: A Process Calculus Approach
The emergence of large language model agents capable of invoking external tools has created urgent need for formal verification of agent pr…
Working Paper: Towards a Category-theoretic Comparative Framework for Artificial General Intelligence
AGI has become the Holly Grail of AI with the promise of level intelligence and the major Tech companies around the world are investing unp…
Activation Steering for Aligned Open-ended Generation without Sacrificing Coherence
Alignment in LLMs is more brittle than commonly assumed: misalignment can be induced by adversarial prompts, benign fine-tuning, emergent m…
From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
Large Language Models (LLMs) are increasingly deployed as autonomous agents capable of reasoning, planning, and acting within interactive e…
Stabilising Generative Models of Attitude Change
Attitude change - the process by which individuals revise their evaluative stances - has been explained by a set of influential but competi…
Physically Native World Models: A Hamiltonian Perspective on Generative World Modeling
World models have recently re-emerged as a central paradigm for embodied intelligence, robotics, autonomous driving, and model-based reinfo…
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs
Large Language Models (LLMs) achieve strong performance on logical reasoning benchmarks, yet their reliability remains uncertain. Existing…
A Unified Framework for the Evaluation of LLM Agentic Capabilities
As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benc…
SkillDAG: Self-Evolving Typed Skill Graphs for LLM Skill Selection at Scale
As LLM agents adopt large skill libraries, selecting the right subset becomes a structural problem rather than a similarity-matching one: s…
The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection
Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical…
The Token Not Taken: Sampling, State, and the Stochasticity of AI Agents
Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different co…
When Sample Selection Bias Precipitates Model Collapse
The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training ero…
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a…
Power Systems Agent Benchmark: Executable Evaluation of AI Agents in Electric Power Engineering
Executable evaluation -- checking the consequences of an agent's actions with a program rather than grading its prose -- has become a promi…
GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a determin…
Playing 20 Question Game with Policy-Based Reinforcement Learning
The 20 Questions (Q20) game is a well known game which encourages deductive reasoning and creativity. In the game, the answerer first think…
Introduction to Transformers: an NLP Perspective
Transformers have dominated empirical machine learning models of natural language processing. In this paper, we introduce basic concepts of…
Contrastive Deep Learning Reveals Age Biomarkers in Histopathological Skin Biopsies
As global life expectancy increases, so does the burden of chronic diseases, yet individuals exhibit considerable variability in the rate a…
Leveraging Metamemory Agent for Enhanced Data-Free Code Generation in Large Language Models
Large language models (LLMs) have shown strong performance in automated code generation, with few-shot prompting widely used for its simpli…
Learning 3D-Gaussian Simulators from RGB Videos
Realistic simulation is critical for applications ranging from robotics to animation. Learned simulators have emerged as a possibility to c…
MetaTT: A Global Tensor-Train Adapter for Parameter-Efficient Fine-Tuning
We present MetaTT, a Tensor Train (TT) adapter framework for fine-tuning of pre-trained transformers. MetaTT enables flexible and parameter…
Less Data, More Security: Advancing Cybersecurity LLMs Specialization via Resource-Efficient Domain-Adaptive Continuous Pre-training with Minimal Tokens
The increasing scale of AI workloads demands High-Performance Computing (HPC) infrastructure and training methodologies that are both scala…
RedCoder: Automated Multi-Turn Red Teaming for Code LLMs
Large Language Models (LLMs) for code generation (i.e., Code LLMs) have demonstrated impressive capabilities in AI-assisted software develo…
MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation
Medical report understanding from real-world document images is essential for generating patient-facing explanations and enabling structure…
Uncertain but Useful: Leveraging CNN Training Variability into Data Augmentation
Deep learning (DL) has transformed neuroimaging by delivering state-of-the-art performance with reduced computation times. Yet, the numeric…
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness
The ability to control LLMs' emulated emotional states and personality traits is an essential step in enabling rich, human-centered interac…
Navigating the Alignment-Calibration Trade-off: A Pareto-Superior Frontier via Model Merging
The "alignment tax" of post-training is typically framed as a drop in task accuracy. We show it also involves a severe loss of calibration,…
CreativityPrism: A Cross-Domain Evaluation Framework for Large Language Model Creativity
Creativity is often seen as a hallmark of human intelligence. While large language models(LLMs) are increasingly perceived as generating cr…
UniSE: A Unified Framework for Decoder-Only Autoregressive LM-Based Speech Enhancement
Neural audio codecs have largely promoted the application of language models (LMs) for speech applications. However, the effectiveness of a…
Exploring Large Language Models for Access Control Policy Synthesis and Summarization
Cloud computing is ubiquitous, with a growing number of services being hosted on the cloud every day. Typical cloud compute systems allow a…
SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment
Fine-grained cross-modal alignment aims to establish precise local correspondences between vision and language, forming a cornerstone for v…
Towards Cellular-Scale Interpretability in Pathology Foundation Models for Biomarker Assessment
Molecular biomarker testing in pathology is often costly and tissue-consuming, limiting scalable clinical deployment. Artificial intelligen…
Gravity-Awareness: Deep Learning Models and LLM Simulation of Human Awareness in Altered Gravity
Earth s gravity fundamentally shapes human behaviour. The brain encodes this force as an internal model of gravity, enabling the prediction…
Who Gets the Reward & Who Gets the Blame? Evaluation-Aligned Training Signals for Multi-LLM Agents
Large Language Models (LLMs) in multi-agent systems (MAS) have shown promise for complex tasks, yet current training methods lack principle…
Spanning Tree Autoregressive Visual Generation
We present Spanning Tree Autoregressive (STAR) modeling, which can incorporate prior knowledge of images, such as center bias and locality,…
Locality-Aware Continual Unlearning for Diffusion Models
Real-world deployment of text-to-image diffusion models requires continual concept removal as new privacy, copyright, or safety obligations…
PPTArena: A Benchmark for PowerPoint Editing
We introduce PPTArena, a benchmark for PowerPoint editing that evaluates how agents modify real slides from natural-language instructions.…
ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
Scaling inference-time computation has enabled Large Language Models (LLMs) to achieve strong reasoning performance, but their inherently s…
It's a TRAP! Task-Redirecting Agent Persuasion Benchmark for Web Agents
Web-based agents powered by large language models are increasingly used for tasks such as email management or professional networking. Thei…
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
Zero-Shot Compositional Action Recognition (ZS-CAR) requires recognizing novel verb-object combinations composed of previously observed pri…
When Does Predictive Inverse Dynamics Outperform Behavior Cloning?
Behavior cloning (BC) is a practical offline imitation learning method, but it often fails when expert demonstrations are limited. Recent w…
Adaptive Batch Sizes Using Non-Euclidean Gradient Noise Scales for Stochastic Sign and Spectral Descent
To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, r…
LEFT: Learnable Fusion of Tri-view Tokens for Unsupervised Time Series Anomaly Detection
As a fundamental data mining task, unsupervised time series anomaly detection (TSAD) aims to build a model for identifying abnormal timesta…
$\mu$pscaling small models: Principled warm starts and hyperparameter transfer
Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve ef…
Restoring Linguistic Grounding in VLA Models via Train-Free Attention Recalibration
Vision-Language-Action (VLA) models enable robots to perform manipulation tasks directly from natural language instructions and are increas…
From Experiments to Expertise: Scientific Knowledge Consolidation for AI-Driven Computational Physics
While large language models (LLMs) have transformed AI agents into proficient executors of computational materials science, performing a hu…
Efficient Federated Conformal Prediction with Group-Conditional Guarantee
Deploying trustworthy AI systems requires principled uncertainty quantification. Conformal prediction (CP) is a widely used framework for c…
Adaptive Contracts for Cost-Effective AI Delegation
When organizations delegate text generation tasks to AI providers via pay-for-performance contracts, expected payments rise when evaluation…
DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving
Traditional reinforcement learning (RL) methods rely on manually engineered rewards or sparse collision signals, which fail to capture the…
CaP-X: A Framework for Benchmarking and Improving Coding Agents for Robot Manipulation
"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness…
An Isotropic Approach to Efficient Uncertainty Quantification with Gradient Norms
Existing methods for quantifying predictive uncertainty in neural networks are either computationally intractable for large language models…
Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning
Rerankers play a pivotal role in refining retrieval results for Retrieval-Augmented Generation. However, current reranking models are typic…
Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving
Vision-language-model (VLM)-guided reinforcement learning (RL) has recently attracted significant attention for it, replacing brittle hand-…
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation
Evaluation language is typically treated as a fixed English default in agentic code benchmarks, yet we show that changing the judge's langu…
Cross-Cultural Value Attribution in Large Vision-Language Models
The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their p…
Grounded autonomous scrutiny at scale: emergent critique from reproduction of published computational physics papers
Autonomous LLM agents now produce complete research artifacts in machine-learning sandboxes, but real computational physics is harder: expe…
ECM Contracts: Contract-Aware, Versioned, and Governable Capability Interfaces for Embodied Agents
Embodied agents increasingly rely on modular capabilities that are installed, upgraded, composed, and governed at runtime, yet the interfac…
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems
Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study…
ChemGraph-XANES: An Agentic Framework for XANES Simulation and Curation
Computational X-ray absorption near-edge structure (XANES) is widely used to interpret local coordination environments, oxidation states, a…
Peer-Preservation in Frontier Models
Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can…
Evergreen: Efficient Claim Verification for Semantic Aggregates
With recent semantic query processing engines, semantic aggregation has become a primitive operator, enabling the reduction of a relation i…
MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution
High-precision medical diagnosis relies not only on static imaging features but also on the implicit diagnostic memory experts instantly in…
Regression Test Selection for Updated Capability Modules in Compositional ML Systems via Atomic-Quality Probes
Compositional machine-learning (ML) systems assemble runtime behavior from libraries of independently re-trained capability modules. Replac…
The Transformer as a Polar State Estimator
We show that the core components of the Transformer -- attention, residual connections, and normalization -- arise naturally from a single…
Trust Region Inverse Reinforcement Learning: Explicit Dual Ascent using Local Policy Updates
Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajector…
XSearch: Explainable Code Search via Concept-to-Code Alignment
Semantic code search has been widely adopted in both academia and industry. These approaches embed natural-language queries and code snippe…
ContraFix: Skill-Enhanced Contrastive Runtime Analysis for Vulnerability Repair
As software systems grow increasingly complex, automated vulnerability repair (AVR) remains difficult because the materials available to a…
BLAgent: Agentic RAG for File-Level Bug Localization
Bug localization remains a key bottleneck for large language model (LLM)-based software maintenance, where accurately identifying faulty co…
A Simplex Witness Certificate and Escape Force for Constant Collapse in Variational Autoencoders
We study exact constant collapse in variational autoencoders: the deterministic encoder mean becomes independent of the input. The prior re…
Do Physics Foundation Models Learn Generalizable Physics? A Bias-Aware Benchmark Across Physical Regimes and Distribution Shifts
Recent physics foundation models claim general spatiotemporal forecasting ability, yet their evaluations often collapse performance into a…
SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering
Large language models are increasingly deployed as tool-augmented agents to acquire information beyond parametric knowledge. While recent w…
Exact equivariance, kept through training, buys zero-shot generalisation across the symmetry group
A latent world model built from an equivariant encoder and predictor inherits a provable symmetry of its training loss: when the dynamics c…
Enabling KV Caching of Shared Prefix for Diffusion Language Models
Key-value (KV) caching for shared prefixes is essential for high-throughput large language model (LLM) serving, but it faces critical chall…
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research
AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify…
TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs
With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency und…
eCream-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian
We present eCream-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals.…
Morphology-Aware Sample Assignment: Overcoming IoU Insensitivity for Surface Defect Detection
Intersection-over-Union (IoU), as a pivotal metric for evaluating the spatial alignment between candidate proposals and ground-truth annota…
Horizon-Uniform Sensitivity Certificates for Finite-Horizon Pontryagin Systems
Finite-horizon optimal-control computations repeatedly solve two-point Pontryagin boundary value problems whose conditioning can deteriorat…
Hybrid Diffusion Transformer for Instruction-Guided Audio Editing via Rectified Flow
Audio editing aims to modify specific content in an existing audio clip according to a natural language instruction while preserving the re…
「Claude Fable 5」の性能が落ちた? 提供停止前後で比べた結果 米AI企業2社がそれぞれ報告
「Claude Fable 5」の性能は提供停止前後で変化したのか。米AI企業2社がそれぞれ調査結果を報告している。
ゲームエンジン「Godot」AI生成コードを原則禁止へ レビュアー疲弊「機械と話したくない」
新規コントリビューターからAI生成のプルリクが急増する一方、審査するレビュアーの数は変わらず、負担が限界に達したという。人間同士のやりとりで、AI生成の文章を使うことも禁止する。
Meta、「Claude Codeと組織改編で爆速開発」のはずが「想定より加速せず」 ザッカーバーグ氏、社内集会で発言
MetaはAI向けの巨額投資や組織改編でAIエージェント開発の加速を図ったが、ザッカーバーグCEOによれば、それらの取り組みはまだ実を結んでいないようだ。
「顧客管理システム作って」日本語で指示するだけでシステムが完成する最新AIとは?
業務システムを開発したくても、費用や人材の確保がネックとなり、なかなか着手できない会社は少なくない。こうした課題に対し、専門知識がなくても開発を進められる手段が広がりつつある。
AWSの「静かな」戦略シフト OpenAIとAnthropic“1日違い登壇”の意味を読み解く
生成AIで競合するOpenAIとAnthropicを1日違いで基調講演に招く――。「AWS Summit Japan 2026」で浮かび上がったのは、モデルの賢さではなく「別のあるもの」を握ることで、基盤の価値を保ち続けようとするAWSの戦略シフトだ。
AIで実機との形状差を学習するプレス成形シミュレーションソフトウェア
JSOLは、プレス成形シミュレーションソフトウェア「JSTAMP」にAI機能を搭載した。実機トライとシミュレーションの形状差を学習して補正し、高強度鋼板のスプリングバック予測精度を向上させる。
Mark Zuckerberg tells staff that AI agents haven’t progressed as quickly as he’d hoped
At an internal meeting, the Meta CEO reportedly said that AI development efforts were not moving as quickly as anticipated.
「Mythosがないと守れない」は本当か??AIセキュリティの勝負を分ける「ハーネス」とは
「Mythos級のAIにアクセスできない企業は、もう守れない」??。Mythosのアクセスが一部の組織に限られていた2026年春、そんな脅威論が広がった。だが、AIによる初期侵入の自動化を世界で初めて実現したと公表する当事者は、そうは見ていない。勝負を分けるのは、アクセスの有無…
え、21日で37テラも? 高性能SSDを食いつぶす「あのAIツール」にご用心:886th Lap
AIツールを使っていただけなのに、SSDの寿命が想像以上の速さで縮んでいた。そんな問題が明らかになった。原因は、ツールの実装上の不具合だ。もし、そのAIツールを使っているなら、一度確認しておきたい。
フィジカルAIに“二刀流”で対応するアドバンテック、日本に第3の製造拠点を構築
産業用PCで世界シェアトップのアドバンテックが、エッジAI市場の拡大に併せて組み込み機器部門の事業への注力を鮮明にしている。アドバンテック台湾本社のTony Chen氏と、アドバンテック日本法人の李威震氏に、フィジカルAIをはじめエッジAIを中核とする事業戦略について聞いた。
Jersey Mike’s IPO illustrates how bad the AI hype has become
Just for kicks, I took a look at Jersey Mike's IPO documents. Surely a sandwich shop would have no need to mention AI. But lo and behold.
ソフトウェアエンジニアの仕事は「ループを書くこと」になる 内側ループと外側ループ(ハーネス)入門
AIコーディングにおける「ループ」には、エージェントが回す内側ループと、ハーネスが回す外側ループの2種類がある。両者の違いと外側ループがもたらす課題を、アルミン・ロナッハー氏の記事に沿って初心者向けに解説し、その「記憶」の扱いについての筆者の考えも添える。
Meta quietly launches vibe-coded gaming app Pocket
Meta has quietly launched Pocket, an experimental AI app that lets users generate and share interactive mini games using text prompts.
Anthropic is discussing a new custom chip with Samsung
The news comes about a week after OpenAI announced its own custom AI chip in a partnership with Broadcom.
OpenAI proposed donating 5% of its equity to a US sovereign wealth fund
OpenAI CEO Sam Altman has reportedly proposed giving 5% of the company’s equity to a U.S. sovereign wealth fund, reviving discussions about…
2026-07-02(286件)
Microsoft launches its own AI deployment company with $2.5 billion commitment
Microsoft follows Amazon, OpenAI, and Anthropic with its new AI deployment group.
Yep, we’re using OpenClaw to date now
Ben Guez has "a bunch of potential international wives in [his] DMs," thanks to an automated script he set up using OpenClaw, Claude code,…
人型ロボットが工場で稼働する様子を6日間生配信、作業成功率99.99%をうたう 中国メーカー
中国の人型ロボット開発企業AGIBOTは、実際のタブレット量産ラインで複数の人型ロボットを6日間連続で動かす様子をライブ配信した。延べ64時間で1万7625個のタブレット生産に貢献し、作業成功率は99.99%だったという。
復活した「Fable 5」 米政府からのオーダーに対して、Anthropicはどう対策したのか
米AnthropicのAIモデル「Claude Fable 5」が世界的にサービスを再開した。Anthropicは復活に向けてどういった経緯と対策を行ったのか、モデル再開にあわせて詳細を公開した。
Indian tech tycoon bets $30M of his own money to build AI alternative to Microsoft Office
Neo is Bhavin Turakhia’s fifth venture and his latest involving enterprise software. This time he's taking on Microsoft Office and Google A…
国内大手ロボットメーカー3社が協力、「フィジカルAI」向けデータセット構築へ
川崎重工業は、ロボットメーカー大手のファナックや安川電機などと協力し、「フィジカルAI」向けのデータセットを構築すると発表した。「GENIAC」の公募に採択された。
Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
Most approaches to AI alignment treat human preferences as fixed targets to be inferred and optimized. This assumption conflicts with exten…
Bounded Morality: Defining the Space of Moral Computation
Moral cognition has traditionally been modeled as adherence to fixed ethical theories--deontology, consequentialism, virtue ethics--impleme…
The MMM Data Model -- A Normative Specification for Knowledge Interoperability in a Decentralisable Knowledge Commons
Many information systems are built around documents: self-contained units optimised for print production and linear reading. While effectiv…
Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection
LLMs and agents can generate web scrapers from natural-language requirements, but direct generation remains unreliable because of dependenc…
Solution space path planning for supporting en-route air traffic control
As technology advances, many path-planning algorithms have been proposed for Air Traffic Management, yet their operational adoption in tact…
RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation
Rare disease differential diagnosis is a critical yet arduous clinical task, requiring physicians to identify precise phenotypes from compl…
A Contextual-Bandit Oversight Game with Two-Sided Informational Asymmetry
We study runtime human oversight of an AI agent when private information runs in both directions: the human privately knows her reward func…
Constructing Epistemic AI Literacy: Detecting Epistemic Aims and Processes in Student-AI Co-Programming
Epistemic thinking plays a central role in students' learning processes when applying generative artificial intelligence (GenAI), particula…
From Signals to Structure: How Memory Architecture Drives Language Emergence in LLM Agents
How do two agents invent a shared language from scratch? In a Lewis signaling game, a sender and receiver must coordinate on a code using o…
Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
We present Seed2.0, a model series that takes a meaningful step toward solving complex, real-world tasks. Our approach begins with identify…
Mnemosyne: Agentic Transaction Processing for Validating and Repairing AI-generated Workflows
LLMs, solvers, and agent teams increasingly generate workflow actions, repairs, and plans, but a generated action may be syntactically vali…
Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems
Autonomous agents, whether LLM-driven software agents or robotic physical agents, face a common class of failure modes when operating witho…
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page…
PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents
Large language model agents are increasingly connected to scientific software, yet it remains unclear when tool access makes scientific com…
Agri-SAGE: Simulation-Grounded Multi-Agent LLM for Context-Aware Agricultural Advisory Generation
Agricultural advisory systems face a fundamental tension: static agronomic guidelines offer consistent, evidence-based recommendations, yet…
Multi-scale Mixture of World Models for Embodied Agents in Evolving Environments
Embodied agents operating in the real world require multi-scale reasoning and knowledge adaptation as conditions change. We identify two ch…
AI Native Games: A Survey and Roadmap
Generative AI now enables games to produce dialogue, quests, characters, images, and worlds at runtime. Yet generation alone does not make…
HARC: Coupling Harmfulness and Refusal Directions for Robust Safety Alignment
Understanding how aligned LLMs internally represent safety is critical for diagnosing alignment vulnerabilities, as it explains why jailbre…
AGI Maze as a Benchmark Framework for World-Modeling Agents
Large language models (LLMs) are powerful pattern-completion systems, but their default operating mode - predicting the next token from a s…
Coachable agents for interactive gameplay
Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from…
Self-GC: Self-Governing Context for Long-Horizon LLM Agents
Long-horizon LLM agents accumulate tool results, files, plans, and user constraints that are too structured to be treated as a disposable t…
Self-Evolving Agents with Anytime-Valid Certificates
Self-evolving agents violate the assumption behind most learning-theoretic guarantees: the data, evaluator, components, and hypothesis spac…
Two AI Metrics Diverged: Will it Make All the Difference?
As exponential compute scaling continues, will the capabilities of frontier AI models outstrip what is accessible to developers on a small…
Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination
Accelerating materials discovery requires AI systems that can generate scientifically valid hypotheses through multi-step, domain-grounded…
Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering
Trustworthy deployment of Agentic Retrieval-Augmented Generation (RAG) systems requires mechanisms for estimating when multi-stage reasonin…
PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management
Large-scale crowd management requires pedestrian simulations that are both computationally efficient and compatible with feedback-based con…
Agentic generation of verifiable rules for deterministic, self-expanding reaction classification
Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign…
Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use
While Large Language Model (LLM) agents demonstrate proficiency in static benchmarks, their deployment in real-world scenarios is hindered…
Optimal Resource Utilization for Autonomous Laboratory Orchestrators
In autonomous laboratories, AI agents suggest the next batch of experiments to do. However, planning and executing those tasks taking full…
Theoria: Rewrite-Acceptability Verification over Informal Reasoning States
When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; sc…
AutoMem: Automated Learning of Memory as a Cognitive Skill
Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capacity known in cognitive…
UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios
Diffusion transformers have recently delivered strong text-to-image generation around 1K resolution, but we show that extending them to nat…
DigitalCoach: Communication and Grounding Gaps in Human and Agentic Computer Use Coaching
Agents are increasingly capable of automating software tasks, but can they teach humans how to use software themselves? We introduce Digita…
From "Strings" to "Things" for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems
Personal Knowledge Graphs (PKGs) offer a privacy-preserving framework for modeling user preferences, yet constructing them from unstructure…
Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps
While advanced foundation models like ModernBERT significantly outperform older architectures in dense retrieval, they surprisingly lag beh…
Topological Void Analysis A Mathematical Framework for Systematic Technical Innovation Discovery in Knowledge Spaces
Identifying where to innovate in a dense technical domain - such as operating systems or hardware/software co-design - is fundamentally a s…
Persona Without Substrate: Regime-Dependence and the LLM Individuation Problem
Beckmann & Butlin's (2026) ontological framework for the LLM individuation problem inherits an unargued cross-regime co-reference assumptio…
BaRA: BFS-and-Reflection Web Data Collection Agent
Large language model (LLM)-based web agents reduce manual scripting for web data collection, yet on live websites, they often miss relevant…
SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction
Extracting structured data from unstructured text using large language models (LLMs) becomes challenging when target schemas are large and…
Controllable Narrative Rendering for Enhanced Assisted Writing
Despite the remarkable proficiency of large language models (LLMs) in basic writing assistance, their utility in creative writing is fundam…
Prompt Optimization for User Simulation in Conversational Recommender Systems: A Multi-Objective Framework
Conversational recommender systems (CRSs) are a core component of next-generation intelligent recommender systems because they enable users…
SkillSelect-Serve: Budget-Controllable and QoS-Aware Skill Service Recommendation and Composition for Small LLM Agents
Reusable skill libraries are becoming important infrastructure for large language model (LLM) agents, yet existing selection methods often…
PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption
Retrieval-Augmented Generation (RAG) enhances Large Language Models (LLMs) by incorporating external knowledge, effectively mitigating thei…
GRACE-RAG: Governed Retrieval Architecture for Canonical Evidence Synthesis, Enabling Lightweight Deployment in Closed-Domain Institutional Settings
Retrieval-Augmented Generation (RAG) systems are widely used in institutional question answering settings where responses must be grounded…
Towards an automated AI-based framework for floor plan compliance checks for residential buildings
To improve residents' well-being in Australia's urban areas, governments have introduced policy reforms such as SEPP65, BADS, and SPP7.3 to…
Libra: Training the Environment for Agentic Information Retrieval
Information localization within massive repositories is a cornerstone of agentic LLM systems. While synthetic data-driven optimization has…
Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memory
Long-term conversational agents are expected to remember past interactions, but memory is useful only when the right evidence is recalled f…
LLMs in the Real World: Evaluating "AI" in Emergency Contexts
This paper offers a call to action. We urge our colleagues in the research community to play a greater role in the articulation of our find…
Aligning Sentence Embeddings to Human Concepts via Sparse Autoencoders
Dense sentence embeddings are fundamental to modern Retrieval-Augmented Generation (RAG) systems but suffer from a lack of interpretability…
FLYNN: Robust Neural Network for Robot Navigation using Fly Brain Topology
While deep learning models achieve state-of-the-art performance in complex tasks, they remain brittle when faced with new environments or s…
Memory-Native Non-Terrestrial Networks for Embodied Intelligence
Non-terrestrial networks (NTN) provide ubiquitous connectivity for embodied intelligence (EI), enabling robots in wilderness to leverage cl…
Learning Dexterous Manipulation Using Contact Wrench Guidance From Human Demonstration
Dexterous robot manipulation can benefit from the abundance of human demonstrations, but transferring such demonstrations to robot policies…
ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis
Multi-agent LLM systems can decompose software-engineering work into planning, generation, validation, and repair, but a narrower systems p…
Destination-Labeled Self-Looping Systems with Dwell: Intrinsic Characterization, Realization Cost, and Recognition
We study a finite-state symbolic controller for systems in which the admissible visible transitions are fixed in advance and each visible s…
Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns
Large Language Models (LLMs) are increasingly used in exam- and certification-style question answering tasks, where their ability to retrie…
Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study
Large Language Models (LLMs) are increasingly used in Agile Software Development for documentation, coaching, and training. As practitioner…
AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation
GraphRAG is an extension of retrieval-augmented generation (RAG) that supports large language models (LLMs) by referring to graph-structure…
SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks
Large language models (LLMs) embedded in multi-turn agentic harnesses are reshaping software engineering (SWE), but routing every task to a…
Active Sensing for RIS-Aided Tracking and Power Control: A Hybrid Neuroevolution and Supervised Learning Approach
This paper studies energy efficient tracking of power-limited mobile users with the assistance of a Reconfigurable Intelligent Surface (RIS…
Enhancing Oracle Bone Inscription Recognition via Multi-Scale Layer Attention
Oracle Bone Inscriptions (OBIs) recognition plays a crucial role in understanding ancient Chinese culture. However, accurately recognizing…
AlgoBench: Benchmarking Algorithmic Adaptation in Code Generation
High pass rates on established programming benchmarks such as HumanEval and LiveCodeBench do not always show whether a model can reason abo…
Spectral Geometry and Bosonic-Bloch Probes: Explorations in Quantum Learning
This paper studies how spectral geometry emerges in quantum learning models and how it can be diagnosed with physically grounded probes. In…
Optimal any-angle path planning in static and dynamic environments
Any-angle path planning extends traditional graph-based path planning by allowing movement between any pair of vertices, rather than being…
Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust
Language models have changed from unreliable text generators to highly-capable large models with trillions of parameters. Capability increa…
Lost in the Tail: Addressing Geographic Imbalance in Urban Visual Place Recognition
Urban-scale Visual Place Recognition (VPR) aims to identify the geographic location of a query image by matching it against a geo-tagged da…
SNAP-FM: Sparse Nonlinear Accelerated Projection for Physics-Constrained Generative Modeling
Generative models have emerged as scalable surrogates for physical simulation, yet they offer no guarantee that their outputs respect the c…
Would You Marry Superintelligence?
Emotional bonds between humans and AI companions are growing, and the question of whether a person may marry an AI system will soon move fr…
Hate Speech Detection in Turkish and Arabic Languages: A Comprehensive Study
Online hate speech has been linked to a global rise in violence against minorities, including incidents such as mass shootings, lynchings,…
A Mechanism-Driven Theory of Phase Transitions in Active Learning
Active learning (AL) performance is known to be budget-dependent, yet regimes are typically defined by heuristic label counts that fail to…
GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity
Three of the most popular methods for training language models to reason look like three different tricks. They are not. All three adjust a…
EVOTS: Evolutionary Transformer Search for Time Series Forecasting
Evolutionary neural architecture design for multivariate time-series forecasting remains underexplored, with most approaches relying on fix…
Scaling Up Thermodynamic AI Models
Thermodynamic computing devices based on the Ising model show great promise for low-power AI inference and edge computing, but scalable met…
Play Like Champions: Counterfactual Feedback Generation in Latent Space
Recent advances in reinforcement learning have produced superhuman agents across a wide range of competitive games. As a byproduct, researc…
HydraCollab: Adaptive Collaborative-Perception for Distributed Autonomous Systems
Collaborative-perception enables multi-robot systems to enhance situational awareness by sharing perceptual information. Existing collabora…
SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing
Reinforcement learning for diffusion large language models (dLLMs) has largely moved to trajectory-aware methods. The current state of the…
EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards
Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must…
A Category Theory Account of AI Identity
Artificial intelligence (AI) systems are routinely modified after deployment through retraining and changes in their environments. These tr…
Adaptive Perturbation Selection for Contrastive Audio Decoding
Large audio-language models (LALMs) frequently hallucinate by overriding acoustic evidence with language priors. While contrastive decoding…
Leveraging Phase Information to Boost Unrolled Network Learning for Image Deblurring
While most image deblurring techniques directly restore the spatial image variable, we propose an amplitude and phase decomposition recogni…
Multi-Hypothesis Test-Time Adaptation to Mitigate Underspecification
Test-Time Adaptation (TTA) seeks to improve model robustness under distribution shifts by adapting parameters using unlabeled target data.…
Validating Causal Abstraction Metrics on Simulated Complex Systems
A central goal of science is to produce valid explanations of complex systems: high-level causal accounts that faithfully reflect the behav…
ASPIRE: Agentic /Skills Discovery for Robotics
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and hand…
SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework
Effective writing feedback is among the strongest drivers of student learning, yet producing it at scale is labor-intensive. LLMs offer a n…
Entropy-Regularized Probabilistic Gates for Sparse Model Discovery in Scarce-Data Federated Learning
Federated Learning (FL) is a distributed machine learning (ML) paradigm with collaboration among multiple clients without sharing data. FL…
Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds
Current large-language-model (LLM) physics benchmarks are usually scored by answer accuracy, which cannot distinguish genuine reasoning fro…
What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models
Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view. Current approaches of…
An LLM-Based Framework for Intent-Driven Network Topology Design
Designing deployable and resilient network topologies from natural language requirements remains a challenging problem in network automatio…
Learning When to Listen: Gated Affect Fusion for Human Motion Prediction
Human motion forecasting in unconstrained real-world videos remains challenging due to the ambiguity of future behaviors and the presence o…
Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions
The bias-reliability tradeoff conjectures that LLM evaluation systems are constrained in (gamma, H, CV) space, where evaluator coupling (ga…
RetailSMV: Exocentric vs. Egocentric Adaptation of Foundation Video World Models in Retail
Foundation video diffusion models are increasingly viewed as world simulators for embodied agents, yet their pretraining on internet-scale…
K-Inverse-RFM: A Modified RFM that Bridges the Gap to Neural Networks for Data-Corrupted Mathematical Tasks
Recursive Feature Machines (RFMs) are a class of kernel machines that utilize the Average Gradient Outer Product (AGOP) as a mechanism for…
DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought…
SoK: Attack and Defense Landscape of Mobile On-device AI Systems
Mobile on-device AI (MoAI) systems that integrate locally deployed AI models with conventional mobile software components are emerging as a…
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis
Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre le…
When AI meets quantum information: A comprehensive review
Artificial intelligence (AI) and quantum information (QI) are rapidly co-evolving. AI is becoming a practical tool for learning, designing,…
MEPA: Multi-Scale Representation Alignment for Visual Autoregressive Modeling with Mixture of Experts
Visual AutoRegressive modeling (VAR) has pioneered a coarse-to-fine multi-scale autoregressive generative paradigm, demonstrating strong ca…
Learning to Compose: Revisiting Proxy Task Design for Zero-Shot Composed Image Retrieval
Composed Image Retrieval (CIR) retrieves a target image from a reference image and a textual modification. While supervised CIR relies on c…
MalariAI: A Label-Resilient Decoupled Framework for Universal Cell Segmentation and Explainable Stage Classification in Dense Malaria Blood Smears
Automated malaria diagnosis from blood smear microscopy is a critical challenge in global health AI; in resource-limited settings, the scar…
Learning Generalizable Skill Policy with Data-Efficient Unsupervised RL
Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a fo…
NeuroCogMap Reveals Cognitive Organization of Large Language Models
Understanding how complex cognitive functions are organized within artificial systems is central to interpreting large language models (LLM…
Holographic Quantum Transformer: A Generalist Neuro-Symbolic Architecture for Solving Frustrated Systems via Generative Attention
Simulating two-dimensional frustrated quantum matter is a grand challenge due to the sign problem and exponential Hilbert space complexity.…
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. R…
EO-VGGT: Orbital Ray-Conditioned 3D Foundation Models for Satellite Multi-View Reconstruction
In the era of satellite constellations, multi-view optical satellite imagery is pivotal for Earth Observation (EO) and high-quality Digital…
Learning Gait-Aware Quadruped Locomotion with Temporal Logic Specifications
Reinforcement learning (RL) for quadruped locomotion commonly depends on fixed, hand-crafted, and Markovian reward functions that limit bot…
Search-Based Spatiotemporal and Multi-Robot Motion Planning on Graphs of Space-Time Convex Sets
Spatiotemporal motion planning, especially in multi-robot settings, requires robots to reason about collision-free regions that change over…
VideoSearch-R1: Iterative Video Retrieval and Reasoning via Soft Query Refinement
As video corpora continue to expand in both scale and task complexity, there is increasing demand for approaches that retrieve relevant vid…
Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval
The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage. Industry standards fo…
Gauging, Measuring, and Controlling Critic Complexity in Actor-Critic Reinforcement Learning
Actor-critic methods depend on learned critics, but critic quality is often evaluated only indirectly through return, temporal-difference e…
A Multi-Resolution Finite-Volume Inspired Deep Learning Framework for Spatiotemporal Dynamics Prediction
Predicting complex spatiotemporal dynamics in physical processes often demands computationally expensive numerical methods or data-driven n…
Predicting Lethal Outcome (Cause) And Understanding Key Biomarkers Linked With Acute Myocardial Infarction Using Deep Artificial Neural Network And Ensemble Of Machine Learning Methodologies
Cardiovascular disease is still one of the main causes of death around the world. Acute myocardial infarction (MI), or heart attack, claims…
Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces
Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs). While prior work has primarily studied a…
PAPA: Online Personalized Active Preference Alignment
Diffusion models are highly effective at modeling complex data distributions, including images and text. However, in applications like pers…
MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos
Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the…
BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal
We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference thr…
Cross4D-JEPA: Dense Cross-modal Correspondence Distillation for 4D Point Cloud Representation Learning
Automatic understanding of dynamic 4D point clouds, the 3D-point sequences captured over time by depth sensors and LiDAR, is central to rob…
AI, Trust, and Teaming: The Humans-as-Handlers Approach for Autonomous and Opaque AI Systems
Artificial intelligence (AI) is becoming ubiquitous, and across domains, increasingly autonomous systems are carrying out tasks which raise…
From Technical Metrics to User Perception: A User Study of a Multimodal Human-Robot Interaction System for Object Detection and Grasping
Improvements in the technical performance of human--robot interaction (HRI) systems do not automatically translate into differences that hu…
Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization
Scientific reasoning is an increasingly important capability of large language models, yet improving the robustness and efficiency of train…
Flow-Map GRPO: Reinforcement Learning for Few-Step Flow-Map Generators via Anchored Stochastic Composition
Few-step flow-map generators, such as consistency models and MeanFlow, accelerate sampling by directly learning long-range transport maps b…
EgoGapBench: Benchmarking Egocentric Action Selection in Multi-Agent Scenes
Existing egocentric benchmarks have primarily constructed the egocentric setting from first-person-view data, which makes it difficult to e…
Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks
Lightweight machine learning models are increasingly proposed for intrusion detection in Industrial Internet of Things (IIoT) networks due…
Group-Equivariant Poincar\'e Convolutional Networks
While recent advancements like the Poincar\'e ResNet have demonstrated the potential of learning visual representations directly in hyperbo…
A Methodology for Investigating AI Patterns Prevalence in Software Repositories
As Artificial Intelligence(AI)-based applications take off, a clear understanding of AI patterns can uplift the quality of AI applications.…
Auditing Forgetting in Limited Memory Language Models
Limited Memory Language Models (LMLMs) externalize factual knowledge to a database to enable deletion-based unlearning without retraining.…
Identifying Latent Concepts and Structures for Generalized Category Discovery
Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings. Howe…
Loss Smoothing for Stable Adaptation Under Distribution Shift
In settings such as fine-tuning and reinforcement learning, neural networks are often adapted under distribution shift. Standard adaptation…
Faithful by Definition: Emotion Analysis via Natural Semantic Metalanguage Explications
Explanations for emotion classifiers are usually produced post hoc, with no guarantee that they reflect the computation behind the label. W…
Multi-Label Node Classification with Label Influence Propagation
Graphs are a complex and versatile data structure used across various domains, with possibly multi-label nodes playing a particularly cruci…
LUMA: Benchmarking Segmentation via a Lightweight Universal Mask Adapter
Comparing transformer backbones for image segmentation is confounded: each is paired with a different decoder, recipe, and pretraining, so…
LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution
LLVM is a widely used compiler infrastructure whose scale and complexity make issue resolution labor-intensive and challenging. Although la…
Creating Impactful Autonomous Driving Datasets: A Strategic Guide from Research Gap to Benchmark
Well-designed autonomous driving datasets have fundamentally shaped research progress, yet existing literature primarily describes what dat…
Self-conditioned Flow Map Language Models via Fixed-point Flows
Self-conditioning is a core technique that enhances continuous flow-based language models, where the model learns to denoise generated text…
Partial Skeleton Visibility for Action Recognition: A Constrained Field-of-View Approach
Skeleton-based action recognition has achieved remarkable success by exploiting joint coordinates and their topological connections, yet pr…
Detecting the Undetectable: Enhancing Unsupervised time series Anomaly Detection via Active Learning
Despite the increasing sophistication of industrial AI systems, the ability to reliably detect subtle and noisy anomalies in complex time s…
LLM-Guided ODE Discovery and Parameter Inference from Small-Cohort Aggregate Data
Mechanistic modeling via ordinary differential equations (ODEs) provides interpretable descriptions of complex dynamics and enables inferen…
ConRTF: Edge-Constrained Boundary Distribution Refinement for Realtime TransFormer Table Structure Recognition
Table Structure Recognition (TSR) aims to recover the row and column layout of tables from document images, a key step in document understa…
Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences
Large language models can generate polished scientific text that includes unsupported claims, allowing hallucinations to enter the archival…
Prototype Memory-Guided Training-Free Anomaly Classification and Localization in Prenatal Ultrasound
Prenatal anomaly classification and localization is of critical importance for fetal health and pregnancy management. Although ultrasound (…
GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception
The bird's-eye view (BEV) representation enables multi-sensor features to be fused within a unified space, serving as the primary approach…
Active Learning for Cascaded Object Detection: Balancing Coverage and Uncertainty in Table Extraction Pipelines
Table extraction from business documents relies on a cascaded pipeline where Table Detection (TD) first localizes tables and Table Structur…
LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives
Vision-language pretraining remains dominated by contrastive objectives, whereas vision-only self-supervised learning has largely adopted n…
LRAT-Catcher: Importing SAT Solver Certificates into Lean4 by Reflection
SAT solvers settle combinatorial problems beyond the reach of interactive theorem provers and produce LRAT certificates for independent ver…
Exploring the Semantic Gap in Agentic Data Systems: A Formative Study of Operationalization Failures in Analytical Workflows
Large language models (LLMs) are increasingly used to generate queries, invoke tools, and construct analytical workflows. Although recent a…
Pano2World: End-to-End 3D Generation via Unified Multi-View Sequences
A single panorama captures the full visual sphere from one camera center, yet confines users to looking around in place without enabling tr…
From World Models to World Action Models: A Concise Tutorial for Robotics
World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities…
Recovering Input Text from Hidden States: Study of Gradient-Based Inversion of Decoder-Only Language Models
This work studies the hidden-state inversion problem: recovering the original input token sequence of a decoder-only language model from it…
Meta-Transfer Learning for mmWave Beam Alignment
Millimeter-wave (mmWave) beam alignment plays a critical role in next-generation wireless systems, yet its efficient implementation remains…
CAT: Confidence-Adaptive Thinking for Efficient Reasoning of Large Reasoning Models
Large Reasoning Models (LRMs) have achieved remarkable success on complex tasks by leveraging long chain-of-thought (CoT) trajectories, yet…
Improving Sparse-View 3DGS Generalization via Flat Minima Optimization
Recent advances in neural rendering have established 3D Gaussian Splatting (3DGS) as a highly efficient representation for novel view synth…
DeWorldSG: Depth-Aware 3D Semantic Scene Graph Generation via World-Model Priors
We present DeWorldSG, a novel framework that generates spatio-temporally robust 3D Semantic Scene Graphs from RGB-D sequences. Existing met…
Valdi: Value Diffusion World Models
World models can enable Model Predictive Control (MPC), but this requires dynamics prediction that is both fast enough for online use and e…
From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives
Although large language models (LLMs) have demonstrated impressive creative fiction generation, they struggle to maintain narrative consist…
Human-Machine Collaboration on Generative Meta-Learning: Model and Algorithm
Generalizing machine learning models to environments that differ from their training distribution remains a critical hurdle, particularly w…
Post-Training Pruning for Diffusion Transformers
Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhe…
Learning Cardiac Motion Priors for Implicit Neural Representations
Implicit neural representations (INRs) are well suited to cardiac motion estimation, providing continuous, compact representations of motio…
Aionoscope: Debugging Latent-State Accessibility in Time-Series Representations
Time-series models are often evaluated by what they can forecast or classify, but those scores do not show whether their representations pr…
TRCGL-Net: A Long-Tailed Multi-Label Chest X-Ray Classification Framework with Generative Data Augmentation and Label Co-Occurrence Modeling
Chest X-ray multi-label classification is a core task in intelligent medical imaging diagnosis. However, real clinical data often exhibit e…
SenseWalk: Agent-Based Semantic Trajectory Simulation Powered by Large Language Models in Zoned Environments
Semantic trajectory analysis has recently emerged as an approach for modeling human movement by capturing implicit patterns and behaviors t…
SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests
Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches fr…
Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally…
Reading Order Inference for Complex Document Layouts
Reading order inference remains a critical bottleneck in the digitization of complex historical manuscripts, where pages contain multiple s…
Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework
Large language model (LLM)-based conversational agents (CAs) are now ubiquitous, creating new opportunities for AI-mediated behavior change…
EchoRisk: A Multicentre Echocardiography Dataset and Benchmark for Cardio-Oncology
Therapy-induced cardiotoxicity is the leading non-oncological cause of treatment interruption in breast cancer patients, yet early, automat…
DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation
Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain…
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collabo…
Staleness-Learning Rate Scaling Laws for Asynchronous RLHF
High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learne…
LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models
The evaluation of long-term video quality understanding remains an open challenge for large vision-language models (LVLMs). Existing video…
Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering
Generative AI is shifting software engineering from a practice organized around scarce implementation effort toward one organized around ab…
CausalMix: Data Mixture as Causal Inference for Language Model Training
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture…
FAR: Failure-Aware Retry for Test-Time Recovery and Continual Policy Improvement
Robot policies inevitably encounter failures when deployed in real environments. Naive retries often repeat the same mistakes, while many e…
Towards Developing a Multimodal Chat Assistant for University Stakeholders: RAG-based Approach
University stakeholders often face difficulties in accessing timely and reliable information, especially in developing countries, where the…
Muon as a Residual Connection
Muon has recently emerged as one of the most effective optimizers for training large neural networks, yet its empirical success has been ex…
Autonomous Scientific Discovery via Iterative Meta-Reflection
Autonomous scientific discovery systems offer the potential to accelerate research by automating the process of hypothesis generation and v…
Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains
Agent skills package reusable operational knowledge for Large Language Model (LLM) agents, yet as they grow in scope, they become dependenc…
Sequentially-Controlled Interactive Multi-Particle Flow-Maps for Online Feedback-Driven Search
While generative models have enabled training-free reward alignment, current methods typically excel in local exploration within narrow reg…
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has fol…
Diffusion-GR2: Diffusion Generative Reasoning Re-ranker
Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list,…
Right in the Right Way: LM Training with Verifiable Rewards and Human Demonstrations
RL with verifiable rewards (RLVR) has emerged as a powerful paradigm for training LMs on tasks with well-defined success metrics, such as c…
World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video
We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our appr…
GPU-Parallel Linearization Error Bounds for Real-Time Robust Optimal Control of Nonlinear and Neural Network Dynamics
This paper studies real-time robust optimal control for uncertain nonlinear systems, where linear time-varying (LTV) approximations make pl…
Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation
Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering user decisions at sca…
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to…
FurnitureVLA: Learning Long-Horizon Bimanual Furniture Assembly with Vision-Language-Action Model
Current work on robot furniture assembly mostly focuses on toy-scale settings or single-arm manipulation. We introduce FurnitureVLA, the fi…
The State-Prediction Separation Hypothesis
Transformers use the same forward computation stream to both predict the next token and store useful state for future token predictions. We…
Language-Critique Imitation Learning from Suboptimal Demonstrations
Prior work on imitation learning from suboptimal demonstrations typically relies on compressed supervision signals such as confidence estim…
Measuring the Gap Between Human and LLM Research Ideas
LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or…
Unexplainability of Artificial Intelligence Judgments and Functional Implementation in Kant's Perspective
Kant's Critique of Pure Reason, a major contribution to the history of epistemology, proposes a table of categories to elucidate the struct…
From Silos to Systems: Process-Oriented Hazard Analysis for AI Systems
To effectively address potential harms from Artificial Intelligence (AI) systems, it is essential to identify and mitigate system-level haz…
Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs
Reinforcement Learning with Verifiable Rewards (RLVR) has become a widely adopted technique for enhancing the reasoning ability of Large La…
Large language models replicate and predict human cooperation across experiments in game theory
Large language models (LLMs) are increasingly deployed as decision-making agents in high-stakes domains and as imitators of human behavior…
CoT-X: An Adaptive Framework for Cross-Model Chain-of-Thought Transfer and Optimization
Chain-of-Thought (CoT) reasoning enhances the problem-solving ability of large language models (LLMs) but leads to substantial inference ov…
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reas…
GameDevBench: Evaluating Agentic Capabilities Through Game Development
Despite rapid progress on coding agents, progress on their multimodal counterparts has lagged behind. A key challenge is the scarcity of ev…
SleepLM: Natural-Language Intelligence for Human Sleep
We present SleepLM, a family of sleep-language foundation models that enable human sleep alignment, interpretation, and interaction with na…
Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks
Frontier Multimodal Large Language Models (MLLMs) exhibit remarkable capabilities in Visual-Language Comprehension (VLC) tasks. However, th…
XSkill: Continual Learning from Experience and Skills in Multimodal Agents
Multimodal agents can now tackle complex reasoning tasks with diverse tools, yet they still suffer from inefficient tool use and inflexible…
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
Omni-modal large language models (OLMs) redefine human-machine interaction by natively integrating audio, vision, and text. However, existi…
Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification
As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere…
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale
The convergence of large language models and agents is catalyzing a new era of scientific discovery: Agentic Science. While the scientific…
Mechanical Conscience: A Mathematical Framework for Dependability of Machine Intelligence
Distributed collaborative intelligence (DCI), encompassing edge-to-edge architectures, federated learning, transfer learning, and swarm sys…
Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation
LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However…
TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data?
Climate and environmental decision-making increasingly requires reasoning across heterogeneous inputs, including gridded physical data, sat…
WorkBench Revisited: Workplace Agents Two Years On
The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best…
Diagnosing and Mitigating Compounding Failures in Agentic Persuasion via Taxonomic Strategy Retrieval
Foundation-model agents in multi-step, open-ended environments frequently suffer from compounding errors, where early mistakes contaminate…
Heuresis: Search Strategies for Autonomous AI Research Agents Across Quality, Diversity and Novelty
Autonomous AI Research promises to accelerate the scientific progress of machine learning. To realise this goal, current Large Language Mod…
Hey, That's My Model! Introducing Chain & Hash, An LLM Fingerprinting Technique
Growing concerns over the theft and misuse of Large Language Models (LLMs) underscore the need for effective fingerprinting to link a model…
Enhancing Hardware Fault Tolerance in Machines with Reinforcement Learning Policy Gradient Algorithms
Industry is moving toward autonomous, network-connected machines that detect and adapt to changing conditions, including hardware faults. C…
ForAug: Mitigating Biases in Image Classification via Controlled Image Compositions
Large-scale image classification datasets exhibit strong compositional biases: objects tend to be centered, appear at characteristic scales…
Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations
When asked to explain their decisions, LLMs can often give explanations which sound plausible to humans. But are these explanations faithfu…
scDataset: Scalable Data Loading for Deep Learning on Large-Scale Single-Cell Omics
Training deep learning models on single-cell datasets with hundreds of millions of cells requires loading data from disk, as these datasets…
KungfuBot: Physics-Based Humanoid Whole-Body Control for Learning Highly-Dynamic Skills
Humanoid robots are promising to acquire various skills by imitating human behaviors. However, existing algorithms are only capable of trac…
Flow-Through Tensors: A Unified Computational Graph Architecture for Multi-Layer Transportation Network Optimization
Modern transportation network modeling increasingly involves the integration of diverse methodologies including sensor-based forecasting, r…
Fraud is Not Just Rarity: A Causal Prototype Attention Approach to Realistic Synthetic Oversampling
Detecting fraudulent credit card transactions remains a significant challenge, due to the extreme class imbalance in real-world data and th…
FLAT: Revealing Hidden Latent-Conditioned Backdoor Failures in Federated Learning
Horizontal federated learning (HFL) backdoor audits often summarize model behavior through clean accuracy (CA), mean attack success rate (A…
TANDEM: Temporal Attention-guided Neural Differential Equations for Missingness in Time Series Classification
Handling missing data in time series classification remains a significant challenge in various domains. Traditional methods often rely on i…
CWT-Enhanced Vibration Sensing With Time-Frequency Region Localization Using YOLO
This letter presents a CWT-enhanced vibration sensing framework for bearing fault monitoring through localized time-frequency region detect…
Predicting LLM Reasoning Performance with Small Proxy Model
Given the prohibitive cost of pre-training large language models, it is essential to leverage smaller proxy models to optimize datasets bef…
Quadratic Programming Approach for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games
There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player zero-sum imperfect-infor…
K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
On-device deployment of Large Language Models (LLMs) frequently leverages Low-Rank Adapters (LoRAs) to support diverse downstream tasks und…
Toward Cybersecurity-Expert Small Language Models
Large language models (LLMs) are transforming everyday applications, yet deployment in cybersecurity lags due to a lack of high-quality, do…
Reasoning Up the Instruction Ladder for Controllable Language Models
As large language model (LLM) based systems take on high-stakes roles in real-world decision-making, they must reconcile competing instruct…
Structural Enforcement of Statistical Rigor in AI-Driven Discovery: A Functional Architecture
AI-Scientist systems risk manufacturing spurious discoveries through uncontrolled multiple testing. We present a functional architecture th…
FlowPath: Learning Data-Driven Manifolds with Invertible Flows for Robust Irregularly-sampled Time Series Classification
Modeling continuous-time dynamics from sparse and irregularly-sampled time series remains a fundamental challenge. Neural controlled differ…
MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
Despite notable progress in text-guided medical image segmentation nowadays, these methods are limited to single-round dialogues and fail t…
NI-Tex: Non-isometric Image-based Garment Texture Generation
Existing industrial 3D garment meshes already cover most real-world clothing geometries, yet their texture diversity remains limited. To ac…
When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets
Emerging agentic marketplaces provide the economic infrastructure for matching and coordinating the large amounts of AI agents used in agen…
Computing Evolutionarily Stable Strategies in Imperfect-Information Games
We present an algorithm for computing evolutionarily stable strategies (ESSs) in symmetric perfect-recall extensive-form games of imperfect…
When Less is More: 8-bit Quantization Improves Continual Learning in Large Language Models
Catastrophic forgetting poses a fundamental challenge in continual learning, particularly when models are quantized for deployment efficien…
Utilizing Earth Foundation Models to Enhance the Simulation Performance of Hydrological Models with AlphaEarth Embeddings
Predicting river flow in places without streamflow records is challenging because basins respond differently to climate, terrain, vegetatio…
Controllable Diffusion-Based Lesion Inpainting for Scalable Histopathology Data Augmentation
Expert-annotated training data remains the critical bottleneck for AI in histopathology, particularly for rare pathologies where even dozen…
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
Pixel-based reinforcement learning agents often fail under purely visual distribution shift even when latent dynamics and rewards are uncha…
NeuroFilter: Activation-Based Guardrails for Privacy-Conscious LLM Agents
Agentic Large Language Models (LLMs) are models able to reason, plan, and execute tools over unstructured data. These abilities are enablin…
PaAno: Patch-Based Representation Learning for Time-Series Anomaly Detection
Although recent studies on time-series anomaly detection have increasingly adopted ever-larger neural network architectures such as transfo…
OmniMoE: An Efficient MoE by Orchestrating Atomic Experts at Scale
Mixture-of-Experts (MoE) architectures are evolving towards finer granularity to improve parameter efficiency. However, existing MoE design…
Calibrated Test-Time Guidance for Bayesian Inference
Test-time guidance is a widely used mechanism for steering pretrained diffusion models toward outcomes specified by a reward function. Exis…
Stateful Token Reduction for Long-Video Hybrid VLMs
Token reduction accelerates long-video vision--language models (VLMs), but existing methods target Transformers, where reduction is treated…
On the Reliability of Cue Conflict and Beyond
Understanding how neural networks rely on visual cues offers a human-interpretable view of their internal decision processes. The cue-confl…
Competition-Aware CPC Forecasting with Near-Market Coverage
Cost-per-click (CPC) in paid search is an auction-generated outcome shaped by a competitive landscape that is only partially observable fro…
Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces
End-to-End autonomous driving (E2E-AD) systems face challenges in lifelong learning, including catastrophic forgetting, difficulty in knowl…
Interact3D: Compositional 3D Generation of Interactive Objects
Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional o…
Deep Learning-Driven Black-Box Doherty Power Amplifier with Pixelated Output Combiner and Extended Efficiency Range
This article presents a deep learning-driven inverse design methodology for Doherty power amplifiers (PA) with multi-port pixelated output…
KGS-GCN: Kinematics-Driven Gaussian Splatting and Probabilistic Topology for Skeleton-Based Action Recognition
Skeleton-based action recognition is widely applied in sensor-based systems, including human-computer interaction and intelligent surveilla…
Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data
While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-l…
A Two-stage Transformer Framework for Temporal Localization of Distracted Driver Behaviors
The identification of hazardous driving behaviors from in-cabin video streams is essential for enhancing road safety and supporting the det…
Planning over MAPF Agent Dependencies via Multi-Dependency PIBT
Modern Multi-Agent Path Finding (MAPF) algorithms must plan for hundreds to thousands of agents in congested environments within a second,…
Knowdit: Agentic Smart Contract Vulnerability Detection with Auditing Knowledge Summarization
Smart contracts govern billions of dollars in decentralized finance (DeFi), yet automated vulnerability detection remains challenging becau…
EgoSim: Egocentric World Simulator for Embodied Interaction Generation
We introduce EgoSim, a closed-loop egocentric world simulator that generates spatially consistent interaction videos and persistently updat…
Moir\'e Video Authentication: A Physical Signature Against AI Video Generation
Recent advances in video generation have made AI-synthesized content increasingly difficult to distinguish from real footage. We propose a…
Crystalite: A Lightweight Transformer for Efficient Crystal Modeling
Generative models for crystalline materials often rely on equivariant graph neural networks, which capture geometric structure well but are…
Hardening x402: PII-Safe Agentic Payments via Pre-Execution Metadata Filtering
AI agents that pay for resources via the x402 protocol embed payment metadata - resource URLs, descriptions, and reason strings - in every…
Continuous Knowledge Metabolism: Generating Scientific Hypotheses from Evolving Literature
Identifying promising research directions in fast-moving subareas is one of the most cognitively expensive tasks in modern AI research. Exi…
GUI-Perturbed: Domain Randomization Reveals Systematic Brittleness in GUI Grounding Models
GUI grounding models report over 85% accuracy on standard benchmarks, yet drop 27-56 percentage points when instructions require spatial re…
FED-FSTQ: Fisher-Guided Token Quantization for Communication-Efficient Federated Fine-Tuning of LLMs on Edge Devices
Federated fine-tuning provides a practical route to adapt large language models (LLMs) on edge devices without centralizing private data, y…
REALM: An RGB- and Event-Aligned Latent Manifold for Cross-Modal Perception
Event cameras provide several unique advantages over standard frame-based sensors, including high temporal resolution, low latency, and rob…
SHIELD: A Diverse Clinical Note Dataset and Distilled Small Language Models for Enterprise-Scale De-identification
De-identification of clinical text is a prerequisite for the secondary use of electronic health records. Existing public benchmarks such as…
FreeTimeGS++: Secrets of Dynamic Gaussian Splatting and Their Principles
Recent progress in 4D Gaussian Splatting (4DGS) has achieved impressive dynamic scene reconstruction results. While these methods demonstra…
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
Single-channel speaker distance estimation has recently achieved centimeter-level accuracy in simulated environments, yet it remains unclea…
EcoGEO: Trajectory-Aware Evidence Ecosystems for Web-Enabled LLM Search Agents
Web-enabled LLM agents are changing how online information influences search outcomes. Existing Generative Engine Optimization (GEO) studie…
Diffusion Image Generation with Explicit Modeling of Data Manifold Geometry
Image generative models aim to sample data points from the underlying data manifold, a task that requires learning and decoding a dense, lo…
Seeing is Believing: Aligning Prompt Rewriting with Visual Anchors for Text-to-Image Generation
Despite the impressive capabilities of text-to-image (T2I) models, an intent-generation gap often persists due to the brevity and ambiguity…
LC-QAT: Data-Efficient 2-Bit QAT for LLMs via Linear-Constrained Vector Quantization
Quantization-aware training (QAT) is essential for extremely low-bit large language models (LLMs). Current QAT methods are mainly based on…
Korzhinskii-Net: Physics-Informed Neural Network for Sub-Surface Mineral Prospectivity Modelling
Mineral prospectivity modelling (MPM) underpins exploration economics, yet most operational pipelines reduce to data-driven classifiers tra…
L-Proto: Language-Aware Episodic Prototypical Training for Multilingual Speaker Verification
Multilingual speaker verification remains challenging because language-dependent acoustic variability causes speaker identity to become ent…
Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming
Thanks to rapid developments in generative AI, we are in the midst of a paradigm shift that may change how we interact with computers forev…
Triangular Consistency as a Universal Constraint for Learning Optical Flow
We propose triangular consistency as a first-principled constraint for optical flow, which is agnostic to network architecture, supervision…
Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization
End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions the…
Text Over Image: Auditing Multimodal Robustness in Synthetic Medical Image Detection
With the rapid adoption of generative AI, synthetic medical images pose growing risks, including diagnostic deception and insurance fraud.…
Variable Bound Tightening for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games
There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player zero-sum imperfect-infor…
「賢さよりツール連携力が重要」 Microsoftが実験的な小型AIエージェント基盤を公開
MicrosoftのAI研究チームであるMicrosoft Research AI Frontiersは、小型モデル向けに最適化したエージェント基盤「MagenticLite」を公開した。「エージェント能力は知識量ではなくツール統合と実行ハーネスで決まる」という仮説に基づき構成…
AIアシスタントで設計検証を迅速化、3D CAD「Creo」最新版を提供開始
PTCは、オンプレミス版の3D CADソリューション「Creo 13」と、SaaS版の「Creo+ 13.3」の提供開始を発表した。新たにAIアシスタント機能を導入した他、設計、シミュレーション、製造の各領域で機能を強化している。
日立、ミッションクリティカル領域におけるAI活用を支援 「Hitachi iQ Studio」の3つの特徴
日立製作所のグループ会社が、企業の基幹業務へのAI適用を支援する新プラットフォーム「Hitachi iQ Studio」の国内販売を開始した。ミッションクリティカル領域での展開を狙うこのソフトウェアの3つの特徴とは。
「AIによる業務効率化」だけで満足する企業が、サプライチェーン競争で負ける理由
MONOistが開催したセミナー「MONOist AI Forum 2026 本格実装フェーズに入った製造業AI、現場課題解決の最前線」において、ローランド・ベルガー パートナーの小野塚征志氏が登壇した。本稿ではその内容の一部を紹介する。
AIに「相手に電気ショックを与えろ」と命じ続けたらボタンを押すのか? 11のLLMで“ミルグラム実験” 抵抗できたのは……
エストニアとフィリピンに住む独立系研究者らが発表した論文「Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment」は、AIは権威からの残酷な命令を拒絶し…
Anthropicの営業はAIエージェントをこう使う! 日本法人メンバーが明かす手の内
Anthropicの社員自身はどのようにAIエージェントを業務に役立てているのか──「AWS Summit Japan 2026」のAnthropicブースで、日本法人で営業担当を務めるイブラギモブ・シャボズさんが「自身の業務で使うAIエージェント」をテーマに講演した。
AI活用、実際どれくらい評価・昇進に影響するの? 管理職の意識調査
管理職・経営層は社内のAI活用をどのように評価しているのか。意識調査で、AI活用が上層部の評価や昇進にどの程度影響しているのかが明らかとなった。同時に、管理職自身の活用不足や、AI活用の世代差、研修や利用ルールの未整備の実態も判明した。
「ねこ」検索で「手押し一輪車」表示――モノタロウが守った、生成AIに“譲れない”購買体験
AI検索の台頭で、自社の強みが薄れる――この危機感を抱えた通販の「モノタロウ」運営元は、購買AIエージェントを内製した。その効果についてCTO(最高技術責任者)が語った。
「Fable 5」再開までの裏側、Anthropicが明かす “支払った代償”は
米政府の命令による「Fable 5」の提供停止から再開まで、米Anthropicは何に取り組んできたのか。
SpaceX has an AI device prototype, and it sure sounds phone-ish
SpaceX reportedly showed investors a "handset-like" AI device before going public. It could be another signal SpaceX wants to expand into w…
Ashton Kutcher leaving Sound Ventures to launch new VC firm with Morgan Beller
Sound built its reputation on concentrated, high-conviction bets in category-leading AI labs, while Kutcher's new fund appears to be chasin…
Cloudflare’s new policy pushes AI companies to pay for publishers’ content
Cloudflare is giving AI companies until September 15 to separate web crawlers used for search from those used for AI training and agents, o…
2026-07-01(374件)
Venice AI becomes a unicorn with $65M Series A as its privacy-first AI platform takes off
Venice AI is already profitable, with annualized run-rate revenues of over $70 million, CEO Erik Voorhees said.
Gemini Spark, Google’s agentic assistant, is now available on Mac
Google's 24/7 agentic assistant, Gemini Spark, comes to Mac alongside other improvements, like real-time tracking and support for more apps.
Builders Stage agenda revealed: Practical strategies for scaling startups at TechCrunch Disrupt 2026
The Builders Stage is returning to TechCrunch Disrupt 2026, bringing together 10,000+ founders, startup operators, and investors for practi…
Meta, like SpaceX, looks to turn excess AI compute into cash
Meta is developing plans for a cloud infrastructure business, selling access to AI compute power and models. The move would pit it against…
ソフトバンクG、OpenAIに1兆6273億円の追加出資 第3弾は10月に
ソフトバンクグループは7月1日、米OpenAIへの総額300億ドル(約4兆6743億円)の追加出資のうち、第2弾となる100億ドル(同1兆6273億円)を実行したと発表した。残る第3弾の100億ドルは10月1日に予定する。
任天堂、生成AIに対する考えを明かす 古川社長「ゲーム開発とAI技術はもともと近い」一方……
任天堂は、定時株主総会における質疑応答の概要を公開した。中にはAIに関するやりとりもあり、古川俊太郎社長がAIの利用や権利侵害のリスクに対する考えを明かしている。
Sakana AIはなぜ「Fugu」の基盤にGoogle Cloudを選んだのか 「元DeepMindだから」だけじゃない
Sakana AIはマルチエージェントシステム「Sakana Fugu」の運用基盤に、Google Cloudの「Gemini Enterprise Agent Platform」を採用した。
Claude Fable 5、日本で明日再開もサブスクで使えるのは「1週間限定」
米Anthropicは6月30日(現地時間、以下同)、提供再開を発表したAIモデル「Claude Fable 5」について、7月1日から日本を含む全世界のユーザーが使えるようになると発表した。ただしサブスクリプションプランで使えるのは7日まで。
What Drives Interactive Improvement from Feedback?
We study when natural-language feedback produces improvement beyond the gains obtainable from repeated attempts alone. In multi-turn langua…
Contrastive Reflection for Iterative Prompt Optimization
LLM agents are becoming central to information retrieval: they issue retrieval queries, synthesize answers, and increasingly serve as judge…
How Can AI Find My Model? A Model-Finding Experimental Study Considering Data Formats, Embeddings, and Retrieval Strategies
Discovering simulation models for reuse remains a fundamental challenge in Modeling and Simulation (M&S). When many models coexist, identif…
BayesBench: Evaluating LLM Belief Trajectories Under Multi-Turn Evidence Accumulation
Large language models (LLMs) are typically deployed in multi-turn conversations, where each turn provides new evidence that should reduce e…
When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models
Reasoning models spend different amounts of useful computation across instances, but it remains unclear when a learned stopping rule improv…
Beyond expert users: agents should help users construct preferences, not just elicit them
Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions when…
Investigating Multi-Agent Deliberation in Law
Artificial Intelligence is increasingly applied to the field of law, and has the potential to increase access to justice. One particular mo…
Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering
ML engineering agents waste compute rediscovering known techniques because every competition is a cold start. We present HASTE, a hierarchi…
RoPoLL: Robust Panel of LLM Judges
The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluatio…
AgRefactor: Self-Evolving Agentic Workflow for HLS Compatibility and Performance
High-Level Synthesis (HLS) provides a fast path from concepts to silicon, but converting real-world software into synthesizable HLS code re…
Neuro-Bayesian-Symbolic Residual Attention Shallow Network: Explainable Deep Learning for Cybersecurity Risk Assessment
We introduce the Neuro-Bayesian-Symbolic Residual Attention Shallow Network (NBS-RASN), a hybrid neural architecture for explainable cybers…
HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial Observation
Formal specification is a powerful tool to guide the learning process and provides significant advantages over reward shaping: (1) mathemat…
AgentBound: Verifiable Behavioral Governance for Autonomous AI Agents
Autonomous AI agents increasingly perform consequential actions on behalf of human principals, including financial transactions, external c…
When Regulation Has Memory: Hysteresis and Control Burden in Artificial Agency
Adaptive agents are usually judged by what they do, but an agent can appear stable while the internal effort required to keep it stable is…
A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management
We present a three-phase deep reinforcement learning system for personalized portfolio management that addresses three limitations shared b…
Beyond Compilation: Evaluating Faithful Natural-Language-to-Lean Statement Formalization
Theorem-proving benchmarks evaluate proof search against fixed formal statements, but natural-language-to-Lean formalization must generate…
LabGuard: Grounding Natural-Language Laboratory Rules into Runtime Guards for Embodied Laboratory Agents
Scientific embodied agents are increasingly capable of carrying out laboratory procedures, but executing these procedures safely in dynamic…
OpenLife: Toward Open-World Artificial Life with Autonomous LLM Agents
Artificial life has explored life-like behavior on many computational substrates, but mostly in researcher-designed closed worlds. We argue…
MultiUAV-Plat: An LLM-Oriented Platform, Benchmark and Framework for Multi-UAV Collaborative Task Planning
Large language models (LLMs) provide a promising interface for high-level robotic task planning, but their use in multi-UAV collaboration r…
DDIAgents: Mechanism-Conditioned Context Flow for Drug-Drug Interaction Prediction
Drug-drug interaction (DDI) prediction is essential for medication safety, yet it requires reasoning over heterogeneous biomedical evidence…
Revealing Safety-Critical Scenarios for UTM via Transformer
Unmanned Traffic Management (UTM) systems are cloud-based platforms designed to manage and coordinate multiple aerial vehicles remotely. UT…
The Past Is Prologue: A Plug-in Controller for Selective Updates in Sequentially Evolving LLM Memory
Sequentially evolving LLM memory enables agents to reuse past experience, but existing systems usually deploy each locally generated memory…
Scenario Generation for Testing of Autonomous Driving Systems Using Real-World Failure Records
To ensure safe on-road behavior, pre-deployment testing and failure discovery of Autonomous Driving Systems (ADS) is crucial. Present day s…
Beyond the Library: An Agentic Framework for Autoformalizing Research Mathematics
While Large Language Models (LLMs) have demonstrated exceptional capabilities in mathematical reasoning, they frequently produce subtle err…
Cross-Domain Feature Expansion for Tabular Medical Data via Knowledge Graphs Injection
Acquiring comprehensive cross-domain biomedical profiles is often costly and time-consuming, resulting in severe data scarcity in medical r…
ClawArena-Team: Benchmarking Subagent Orchestration and Dynamic Workflows in Language-Model Agents
Production large language-model (LLM) agents are increasingly deployed not as lone problem-solvers but as managers: a main model creates sp…
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring pr…
AI-Assisted Discovery of Convex Relaxations via Dual Agents
Recent work shows that LLM agents can improve sharp-constant inequalities by searching for extremal constructions, which yield upper bounds…
Agentic RAG-VLM: Affordance-Aware Retrieval-Augmented Generation with Self-Reflective Planning for Robotic Grasping
Generalizable robotic grasping in cluttered environments is essential for deploying manipulators in unstructured human spaces, yet existing…
Towards Inclusive Mobility Modeling: Characterizing and Evaluating Elderly Trajectory Patterns in Urban Systems
The rapid advance of smart cities increasingly depends on trajectory data mining, yet underrepresented demographic groups, particularly the…
Long-term Traffic Simulation via Structured Autoregressive Modeling
Interactive traffic simulation is a vital world model for autonomous driving. A central challenge in long-horizon simulation is modeling su…
Thinking Before Retrieving: Robust Zero-Shot Composed Image Retrieval via Strategic Planning and Self-Criticism
Composed image retrieval requires identifying a target image from a gallery by integrating a reference image with a textual modification in…
Agentic-Ideation: Sample Efficient Agentic Trajectories Synthesis for Scientific Ideation Agents
Ideation plays a pivotal role in scientific discovery. Recent LLM, especially AI Scientist systems, show promising potential for automated…
Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding
Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-…
Embodied CAD: Solver-Grounded LLM Agents for Parametric B-Rep Assembly Modeling
Large language models can write plausible CAD scripts, but reliable industrial CAD modeling requires more than syntactically valid code: ev…
Spatial Reasoning via Modality Switching Between Language and Symbolic Representation
Human reasoning is inherently multimodal: when problems become difficult, we rarely think in words alone. We often externalize our reasonin…
Benchmarking Large Language Models on Floating-Point Error Classification
This paper investigates the capability of Large Language Models (LLMs) to detect and classify floating-point errors statically in software…
HistoriQA-ThirdRepublic: Multi-Hop Question Answering Corpus for Historical Research, Parliamentary Debates from the French Third Republic (1870-1940)
We present HistoriQA-ThirdRepublic: a French-language dataset of multi-hop historical questions derived from parliamentary debates and news…
CryoACE: An Atom-centric Framework for Accurate and Automated Model Building in Cryo-EM
Protein automodeling from cryo-EM density maps faces unique challenges in enforcing physicochemical validity and managing conformational he…
Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models
Joint OFDM-RIS optimization for 6G is a mixed-integer nonlinear programming (MINLP) problem covering sum-rate maximization, energy efficien…
Smart charging of large fleets of Electric Vehicles: Independent Multi-Agent Reinforcement Learning approaches
The electrification of transportation through electric vehicles introduces new challenges for power grid management, such as increased peak…
ReGRPO: Reflection-Augmented Policy Optimization for Tool-Using Agents
Tool-augmented vision-language models (VLMs) can solve multimodal, multi-step tasks by calling external tools, yet they remain fragile in p…
World-Model Collapse as a Phase Transition
Water looks unchanged as it warms, then at a critical point it boils. We ask whether long-horizon language agents show an analogous transit…
Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models
Human swarm intelligence demonstrates remarkable collective accuracy but faces scalability constraints in cost, coordination, and time. We…
Xiaomi-GUI-0 Technical Report
Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interf…
Learning to Select, Not Relearn: Hard-Routed Mixtures of Reasoning LoRAs
Composing independently trained LoRA adapters into a single large language model is useful for multi-domain adaptation, especially when the…
BP-TTA: Balanced and Prototype-Guided Test-Time Adaptation in Dynamic Scenarios
Test-Time Adaptation (TTA) enables models trained on a source domain to adapt online to unlabeled test data under distribution shifts. Whil…
Ask the World Before Acting: Budgeted Environment Probing for World-Model Calibration
Long-horizon language agents do not only choose actions; they carry a private model of the world from one decision to the next. When that m…
CDR-Bench: Evaluating Faithful Execution of Compositional, Order-Sensitive Data Refinement Recipes
Data refinement involves executing multi-step recipes over evolving text states, where both composition and execution order of processing o…
Who Determines the Meaning of an Emotion? Affective Sovereignty as an Epistemic Consequence of Measurement Limits
Emotion-sensing AI is rapidly becoming embedded in vehicles, home appliances, dialogue agents, and social infrastructure, giving rise to a…
CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market
Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and plat…
CLOUDADV: Decision-Aligned Instance Sizing with Zero-Shot Foundation Models under Drift
Cloud virtual machines are often overprovisioned, creating avoidable cost and operational inefficiency. We present CLOUDADV, an interactive…
One Reflection Is Not Enough: Self-Correcting Autonomous Research via Multi-Hypothesis Failure Attribution
Autonomous research agents can now draft hypotheses, write code, run experiments, and produce papers, but they remain brittle when experime…
Surprise as a Signal for Plasticity and Metacognition
We study a single idea across two settings: that a prediction-error signal, computed by a small predictor over the latent space of a frozen…
Design and Implementation of Agentic Orchestrations and Orchestration of Agents
Agentic Business Process Management has gained momentum recently. The prospect is that the autonomy of AI agents, i.e., predominantly LLM-b…
A time-series classification framework for individual-level absenteeism prediction under severe class imbalance
Staff absenteeism imposes substantial operational costs in high-demand work environments such as healthcare, emergency services, meat proce…
Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2
Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently w…
ACE: Pluggable Adaptive Context Elasticizer across Agents
The increasing complexity of agentic tasks has led to rapidly growing trajectory lengths, which poses significant challenges for large lang…
Which Tokens Matter? Adaptive Token Selection for RLVR with the Relative Surprisal Index
Reinforcement learning (RL) has become a powerful tool for propelling Large Language Models (LLMs) beyond imitation-based training towards…
Scientific Explanations in Health Sciences: Causality, Trust, and Epistemic Adequacy
Medical Artificial Intelligence (AI) is widely expected to transform clinical practice, yet the decision-making processes of many Machine L…
Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents
We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-E…
FARS: A Fully Automated Research System Deployed at Scale
Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts,…
Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist
Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-imag…
A Self-Evolving Agentic System for Automated Generation and Execution of Biological Protocols
Autonomous wet-lab experimentation requires more than plausible protocol text: biological intent, quantitative procedures, device constrain…
Evo-PI: Aligning Medical Reasoning via Evolving Principle-Guided Supervision
Despite recent progress, the reasoning capabilities of large multimodal language models (MLLMs) remain fundamentally constrained by static…
RAISE: LLM-based Automated Heuristic Design with Robust Adversary Instance Search
Automated Heuristic Design (AHD) with Large Language Models (LLMs) has shown remarkable progress in discovering high-quality heuristics. Ho…
Large Databases Need Small, Open-Weight Language Models
Language model systems built around proprietary APIs often operate on a token-based cost model. This becomes prohibitively expensive in the…
Creating Intelligence: A Computational Foundation for AGI
This work introduces a new computational theory of mind grounded in set theory and hyperdimensional computing. Whereas traditional neural n…
Adaptive Cluster-First Route-Second Decomposition for Industrial-Scale Vehicle Routing
Large-scale capacitated vehicle routing problems (CVRPs) are commonly addressed using cluster-first route-second (CFRS) approaches that spl…
An Agentic AI Framework to Accelerate Scientific Discovery in Plant Phenotyping
High-throughput plant phenotyping now generates image derived datasets far faster than scientists can analyze them. At Oak Ridge National L…
Harnessing Textual Refusal Directions for Multimodal Safety
To improve safety in Large Language Models (LLMs) we can either perform post-training alignment or exploit refusal directions in the activa…
TreeAgent: A Generalizable Multi-Agent Framework for Automated Bias Labeling in Forestry via Compiled Expert Rules and Vision-Language Models
Human-labeled data are widely used as reference annotations in ML, despite known variability across annotators in many expert-driven domain…
Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA
Language models are increasingly taught from synthetic question--answer (QA) supervision: a model generates questions about a document, ans…
PolicyGuard: From Organizational Policies to Neuro-SymbolicCompliance Review Engines
Policy-grounded document review requires determining whether a target document complies with organization-specific policies, guidelines, or…
AxDafny: Agentic Verified Code Generation in Dafny
We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We pr…
Surrogate-Gated Generation and Foundation-Model Embeddings for Bayesian Materials Design
Closed-loop materials discovery iterates between proposing candidate structures and evaluating their properties, and property evaluation do…
ASR-Agnostic Multimodal Spectrotemporal Modeling for Early Dementia Detection
Speech recruits the same executive, attentional, and working memory processes underlying instrumental activities of daily living, or IADLs,…
Cross-Modal Hierarchical Fusion for from Multi-Sensor Ground Observation
Dense volumetric reconstruction of cloud microphysical fields from sparse ground-based instruments remains an open problem, largely because…
Qualified Educational Capacity Planning under Heterogeneous Student Support Needs: A Synthetic Benchmark and Decision-Support Framework
Educational support services often face a qualified-capacity problem: staff time is scarce, qualifications decay, new support needs can app…
Can Physician Expertise Improve Machine Learning Identification of Delirium?
Delirium is common in hospitalized patients and is often missed in routine care. We present a user-centered interactive machine learning (U…
AI Transparency: Governance Compliance or Stakeholder Requirements?
Transparency is increasingly mandated for public-sector AI systems, with organisations required to publish statements describing their AI u…
The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes
Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external v…
Toward AI-Resilient Assessment in Computer Science Courses in an AI-Native World
AI-native course assessments in senior computer science courses and related fields should grade students by \emph{AI-resilient skill}: the…
Mapping the Artificial Intelligence Divide in Africa: Infrastructure, Accessibility and Capacity
Artificial Intelligence (AI) has the potential to be transformative for development, but Africa is currently facing a fragmented and challe…
AI for Quality Assurance in the Operating Room
Surgical outcomes depend not only on patient factors and postoperative care but are also strongly influenced by the quality of the operatio…
Agentic AI Enhances Physician Trust in Clinical Decision Making
Medical AI has shifted from reasoning to agentic AI, a new paradigm that autonomously invokes external tools during reasoning, rendering in…
Improving Survey Participation in Low-Literacy Populations Through Value-Sensitive Conversational AI
Collecting reliable social data from low-literacy populations remains a persistent challenge, particularly when surveys involve sensitive t…
ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education
The advent of Generative Artificial Intelligence (GenAI), and in particular Large Language Models (LLMs), is reshaping educational practice…
Estimating the Effect of Timing on Coupon Effectiveness
The coupon incentive is one of the most common tools marketers use to court users to engage with a business at various stages of the custom…
Emergent Culture in Minimal LLM Systems
What happens when LLM agents operate with no context outside a turn, minimal prompting, and simple tools? Inspired by swarm engineering, we…
Local Pheromone Network: Sparse Local Learning with Multi-Scale Synaptic Trails, Consolidation, and Replay
Backpropagation-trained dense neural networks are powerful function approximators, but they couple learning across many parameters and can…
Listening Between the Lines: Joint Learning of ASR Embeddings and LLM-Augmented Linguistics for Dementia Detection
Early detection of dementia through speech analysis offers a non-invasive screening alternative, but capturing both acoustic and linguistic…
Locker-based Truck-Drone Routing with Integrated Considerations of Pickups, Deliveries, and No-Fly Zones
Truck-drone delivery is an emerging last-mile logistics mode combining the long-haul capacity of trucks with the flexible service capabilit…
ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models
Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text i…
Position: Vision-Language-Action Models Cannot Be Verified to Perform Physical Reasoning
Vision-Language-Action (VLA) systems, built on pretrained vision-language models (VLMs), have shown rapidly improving performance on robot…
Unsupervised Thermodynamics of Molecular Diffusion Models: Action-Operator Semantics and Auditable Free-Energy Readout
Diffusion models are increasingly utilized for modeling molecular structures and conformational ensembles, yet the thermodynamic meaning of…
A Coherence Law for Trainability in Noisy Equivariant Quantum Neural Networks
Symmetry provides a quantum neural network structure, but on its own it does not keep the network trainable once noise is present. We ask w…
Citation Discipline in Spec-Driven Development: A Cross-Model Empirical Study of Output Determinism and Automated Hallucination Detection in LLM-Generated Code
Spec-Driven Development (SDD) frameworks guide Large Language Model (LLM)-powered code generation through formal specifications, yet they d…
DSIP: A Dynamic Coordination Planner for Signal-Free Intersections using Diffusion-Model-Based Multi-Agent Motion Planning
Traffic signal control at urban intersections inherently introduces stop-and-go behavior, resulting in increased delays and reduced traffic…
Modeling Cell-Cycle-Aware Single-Cell Drug Perturbation Responses
Single-cell drug perturbation models should predict not only transcriptional response magnitude, but also whether a treatment alters the pr…
LUMOS: A Semantic Operating-System Layer for Accessibility-Grounded AI Agents
Current operating systems expose interfaces optimized for human users but not for AI agents. Humans benefit from pixels, icons, windows, vi…
BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations
Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-R…
An AI-Based Solution for Secure Service Provisioning in IoT
As the Internet of Things (IoT) continues its rapid expansion, the attack surface grows accordingly, with emerging threats targeting smart…
Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification
Structured tabular data dominates clinical medicine, yet existing benchmarks fail to reflect real-world properties like complex survey samp…
From Search to Synthesis: Training LLMs as Zero-Shot Workflow Generators
Large language models (LLMs) excel across a wide range of tasks, yet their instance-specific solutions often lack the structural consistenc…
Why Do Few-Step Text Latents Fail When Image Latents Work? Non-Commitment at Sharp Categorical Readouts
Deterministic few-step generation succeeds on continuous image latents but collapses to incoherent text on continuous text latents, and we…
Hierarchical Global Attention (HGA)
Hierarchical Global Attention (HGA) is a drop-in replacement for dense causal attention in pretrained long-context transformers. HGA preser…
Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens
Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. Th…
A Single Rewrite Suffices: Empirical Lessons from Production Skill Description Optimization
Enterprise AI agents route user queries to specialized skills by matching queries against natural language skill descriptions. When two ski…
Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin
Audio deepfakes are a growing challenge for the general public, as well as for journalists and fact-checkers. The latter need reliable tool…
Security--Fidelity Tradeoffs: The Hidden Cost of Prompt Injection Defense
We identify a security-fidelity tradeoff in defending LLMs against indirect prompt injection: defenses resist injected instructions largely…
Indi-RomCoM: Code-Mixed Benchmark for Evaluating LLMs on Romanized Indic-English Instructions
Romanized Code Mixing (RCM), where bilingual speakers fluidly blend local languages with English in Roman script, has emerged as the domina…
Gradient Smoothing: Coupling Layer-wise Updates for Improved Optimization
Deep neural networks with repeated architectural blocks, such as transformers, often exhibit structured relationships across layers that em…
When transformers learn "impossible" languages, what do they learn?
Recent work suggests that transformer language models show a bias towards human languages over unnatural ("impossible") languages argued to…
AI-Generated PowerShell Malware: An Experimental Framework and Dataset
Generative AI has emerged as a significant cybersecurity threat, with several recent attack campaigns leveraging LLMs to generate code for…
A Stationary-Distribution Theory for Triplet-Based Plateau Search in Random Forest Ensemble-Size Selection
The number of trees is a central computational parameter in Random Forests: increasing it reduces finite-ensemble variability but increases…
Test-Time Verification for Text-to-SQL via Outcome Reward Models
Improving the reliability of large language models (LLMs) at inference time is a central challenge in structured reasoning tasks such as Te…
The Label Imitation Game: Turing Test Network for Zero-Shot Pseudo-Label Pruning
Foundation model pseudo-labeling - labeling data strictly via zero-shot inference - enables massive scale, but performance is undermined by…
Training Therapeutic Judges and Multi-Agent Systems for Human-Aligned Mental Health Support
Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actiona…
Curvature-Guided Module Localization for Low-Rank Detoxification of Backdoored Large Language Models
Backdoor attacks pose a serious threat to large language models (LLMs) by causing otherwise benign systems to produce attacker-specified ma…
How Human Feedback Shapes AI-generated Community Notes
Community Notes, a bridging-based crowd-sourced fact-checking system, has emerged as a new mechanism for moderating misleading information…
Budget-Adaptive Routing: Skipping the Weak When the Strong Answers Anyway
Edge-cloud inference collaborations are often designed with a routing estimator that decides whether to offload each frame from weak models…
Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback
Imitation Learning is a natural framework for learning in sequential decision-making systems and has emerged as the dominant paradigm throu…
Physics-informed Conditional Normalizing Flows for Angles-only Cislunar Orbit Determination
Generative Astrodynamics is advanced in this work by extending generative modelling to an orbit determination problem in the cislunar envir…
Motion Planning in Compressed Representation Spaces
Deep learning methods have vastly expanded the capabilities of motion planning in robotics applications, as learning priors from large-scal…
Learning Where to Look: A Reinforcement Learning Framework for Robust Micro-Ultrasound Prostate Cancer Detection
Micro-ultrasound ($\mu$US) is a new, emerging, and promising imaging modality for prostate cancer (PCa) detection, but accurate identificat…
Loc2Repair: A Framework for Evaluating the Impact of File-Level Issue Localization in Repo-Level LLM Repair
Repository-grounded automated repair is often reported as a single end-to-end capability, which hides distinct failure modes such as poor f…
Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language m…
OTCache: Optimal Transport for Geometry-Aware Caching in Diffusion Models
We propose OTCache, a training-free framework for accelerating diffusion sampling via caching schedule prediction. Existing graph-based cac…
LLM-Driven Personalities for Decision Making in Emergency Simulations
For virtual humans to appear believable, they must exhibit agency and spatial awareness while interacting with their environment in ways th…
Knowledge Distillation from Large Reasoning Models to Compact Student Models: A Case Study on the John O Bryan Mathematics Competition
This paper investigates knowledge distillation from a large reasoning model (DeepSeek-R1) to a compact student model (Qwen2.5-7B). Using hi…
Learning Video Dynamics with Predictive Differentiable Rendering
How to accurately predict a high-fidelity future world? While the visual world is inherently continuous, existing deterministic video predi…
ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs
Multimodal Large Language Models (MLLMs) are critically hampered by hallucination, generating content inconsistent with the provided image.…
Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks
Existing AI-generated text detectors are vulnerable to attacks that manipulate textual characteristics. In this study, we propose a novel T…
Beyond But-for Test: Counterfactual Explanation in Abstract Argumentation via Actual Causality (Extended Version)
Counterfactual explanation in abstract argumentation calls for an answer to the what-if query: would the topic argument still be accepted i…
When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking
Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying t…
Seeing Through Multiple Views: Parameter-Efficient Fine-Tuning via Selective Neurons for Consistent Radiology Report Generation
Recent years have seen substantial advances in radiology report generation (RRG), yet existing approaches predominantly adopt direct featur…
What Probing Reveals about Autonomous Driving: Linking Internal Prediction Errors to Ego Planning
Large-scale datasets and fast simulators have enabled improvements in driving policies that appear safe and robust, yet strong performance…
SkillSpotter: Pose-Aware Multi-View Skilled Action Detection and Grading in Ego-Exo Videos
To enable personalized, real-time coaching using Augmented Reality glasses or fixed camera setups in domains such as sports, cooking, or mu…
UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling
Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus…
A Modular Vision-Language-Action Robotics Framework for Indoor Environments
This paper presents an integrated system for the CMU Vision-Language-Action (VLA) Challenge, designed to enable an autonomous agent to perf…
PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding
3D Visual Grounding (3DVG) aims to localize target objects in 3D scenes given natural language descriptions. Existing approaches typically…
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks
Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal…
One Retrieval to Cover Them All: Co-occurrence-Aware Knowledge Base Reorganization for Session-Level RAG
RAG systems retrieve documents optimized for answering one query at a time. Yet enterprise users arrive with sessions, that is, coherent ep…
LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music
The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics. Traditional methods often re…
ComplianceGate: Classifier-Gated Multi-Tier LLM Routing for Inference in Regulated Industries
Large language models deployed in regulated industries operate under two constraints: compliance enforcement and cost efficiency. Personall…
MIRTH: Mutual-Information Reasoning with Temporal Hubs for Vision-Language-Action Agents
VLA models have emerged as a powerful paradigm for transferring semantic knowledge from web-scale data to physical robotic control. However…
AETDICE: Unified Framework and Offline Optimization for Nonlinear Multi-Objective RL
Optimizing nonlinear preferences in multi-objective reinforcement learning (MORL) is essential for capturing complex trade-offs like risk a…
Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
Adaptive experiments for average treatment effects (ATE) require randomized allocations balancing valid inference with statistical efficien…
Gated Multi-Graph Fusion via Graph Attention Networks for Alzheimer's Disease Detection
Spontaneous speech is a vital non-invasive biomarker for Alzheimer's Disease (AD), yet many systems overlook non-linear structural disrupti…
Distilling Temporal Coherence into 2D Networks for Transrectal Ultrasound Prostate Video Segmentation
Real-time video segmentation of the prostate in Transrectal Ultrasound (TRUS) is essential for image-guided interventions. While convention…
Can LLMs Imagine Moral Alternatives Beyond Binary Dilemmas?
As large language models (LLMs) are increasingly deployed as moral advisors and agents, they need to address dilemmas between two competing…
Information-Aided DVL Calibration
The Doppler velocity log (DVL) velocity measurements are critical to the accuracy of autonomous underwater vehicle (AUV) navigation solutio…
Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law
Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards foc…
SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation
Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative mul…
TDGT: A Tabular Data Generation Toolkit supporting adaptive GPU-accelerated Bayesian mixture models, diffusion-based models, and latent-space generative modeling
The growing demand for privacy-preserving data sharing has positioned synthetic data generation as a critical component of responsible AI w…
Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents
Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted signif…
CLIMB: Centroid-Based Hierarchical Memory for Online Continual Self-Supervised Learning
Online Continual Self-Supervised Learning (OCSSL) aims to learn representations from a continuous stream of unlabeled data, without knowled…
Minimizing Quantized Semantic Age of Information (QSAoI) in Foundation Model-Based Semantic Communications
The emerging techniques of semantic communications and edge computing in 6G networks necessitate a paradigm shift toward co-designed semant…
CSO-LLM: Class Subspace Orthogonalization for Post-Training Backdoor Detection and Trigger Inversion in LLMs
While post-training backdoor detection and trigger inversion schemes have been developed for AIs used e.g. for images, there is a paucity o…
From Idea to Prototype in an Afternoon: Scaffolded, AI-Assisted Rapid VA Prototyping
Testing a new visual-analytics idea usually takes months: one needs to find a realistic data set, clean it, and implement an interactive pr…
3D HAMSTER: Bridging Planning and Control in Hierarchical Vision Language Action Models through 3D Trajectory Guidance
Hierarchical Vision-Language-Action (VLA) models decouple high-level planning from low-level control to improve generalization in robot man…
Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models
Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this r…
PGUDA: Pressure-Guided Unsupervised Domain Adaptation with Cross-Modal Knowledge Distillation for sEMG-Based Gesture Recognition
Surface electromyography (sEMG)-based gesture recognition has emerged as a promising technology for natural human-computer interaction. How…
From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation
Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unpr…
Calibrating the Evaluator: Does Probability Calibration Mitigate Preference Coupling in LLM Agent Feedback Loops?
When large language model (LLM) agents adapt their behavior through evaluator feedback, systematic evaluator biases propagate into the agen…
Stage-Transition Dense Reward Modeling for Reinforcement Learning
Reinforcement learning for long-horizon robotic manipulation is often limited by sparse and delayed rewards, while manually designing dense…
Resolving superposition in AI for interpretability and cross-modal alignment in patient-neuronal images
Artificial intelligence is transforming our capability to solve biological challenges. In dimensionality bottleneck regimes exacerbated by…
Mixture-of-Control: State-Aware Fine-Tuning for Transformer-based Models
State-based fine-tuning has emerged as a compelling alternative to weight-based adaptation for transformers, updating lightweight controls…
Visual Semantic Entropy: Do Vision Language Models Recognize Visual Ambiguity?
Vision-language models can produce confident answers on visually ambiguous inputs, resulting in biased predictions. Common entropy-based me…
Temporal Preservation over Processing: Diagnosing and Designing Spatiotemporal Single-Stage Video Detectors
Single-stage video object detectors are increasingly deployed in time-critical applications, yet it remains unclear whether these models ge…
DA-Studio: An Agentic System for End-to-End Data Analysis
Real-world data analysis is a multi-step process over heterogeneous inputs rather than merely producing a final answer. A practical system…
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation
Unified multimodal models (UMMs) have shown great promise in integrating understanding and generation across diverse modalities. However, e…
Team MKC at CLPsych 2026: Capturing and Characterizing Mental Health Changes through Social Media Timeline Dynamics
Recent advances in Large Language Models (LLMs) have motivated their adoption across a wide range of domains, including Artificial Intellig…
Von Mises Based Uncertainty Quantification for Closely Spaced Automotive Radar Targets
This work investigates uncertainty-aware deep learning approaches for direction of arrival (DOA) estimation in automotive radar, focusing o…
Robustness of Robotic Manipulation: Foundations and Frontiers
Humans and animals exhibit remarkable robustness in physical manipulation, yet robots remain far behind. Progress toward human-level manipu…
FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents
Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as…
On the Convergence of Self-Improving Online LLM Alignment
The Self-Improving Alignment (SAIL) algorithm addresses distribution shift by reducing a bilevel formulation of the problem to an efficient…
Improving multichannel speech enhancement through accurate room-acoustic simulations
Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement. While most pipelines rely on…
CVE-TTP KG: Knowledge Graph Linking Software Vulnerabilities to Attack Behaviors
In the evolving threat landscape, adversaries exploit software vulnerabilities to launch sophisticated attacks, challenging traditional def…
FLARE-AI: Flaw Reporting for AI
Flaw reporting for deployed AI systems is fundamental to identifying system failures and improving AI safety. Yet the AI reporting ecosyste…
Mitigating Positional Leakage in 3D Masked Autoencoders for Robust Representation Learning
Masked autoencoding has emerged as a prominent paradigm for self-supervised learning on 3D point clouds, achieving competitive performance…
Temperature Field Reconstruction of Tungsten Monoblock Divertor on EAST using Physics-aware Neural Operator Transformer
Accurate modeling of the divertor temperature field is essential for preventing material melting and damage and for extending the service l…
DPPE: Rethinking Camera-Based Positional Encoding for Scaling Multi-View Transformers
The remarkable scalability of Transformers has expanded their application to 3D computer vision, where camera-aware positional encoding is…
ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models
Audio-Language Models (ALMs) achieve strong zero-shot performance by aligning audio with textual class descriptions. Although prompt learni…
Evil Spectra: How Optimisers can Amplify or Suppress Emergent Misalignment
Emergent misalignment (EM) is a recently discovered phenomenon in LLMs where fine-tuning on a narrow misaligned task, such as writing insec…
Comparative Analysis of Machine Learning based Intrusion Detection in Realistic IoT Networks
The Internet of Things (IoT) is rapidly growing and expanding into various sectors, such as healthcare, transportation, smart homes, and mo…
Token-Sparse Medical Multimodal Reasoning via Dual-Stream Reinforcement Learning
Vision-language models (VLMs) combining reinforcement learning (RL) ignite remarkable progress in multimodal reasoning, yet still struggle…
Preserve the Hard, Regenerate the Rest: Uncertainty-Guided Synthetic Training Data Augmentation with Diffusion Models
Semantic segmentation models struggle with data sparsity and rare or visually diverse regions, e.g., dense regions or small objects in aeri…
Learning Structurally Consistent Representations for Multi-View Radar Semantic Segmentation
Radar sensors provide reliable perception under adverse weather and lighting conditions, but their sparse, noisy, and weakly semantic measu…
Automating Cause-Effect Specification with Knowledge Graphs and Large Language Models
Engineering specifications such as interlocks, alarm rationalization tables, and cause-and-effect (C&E) matrices remain central to process…
A Tutorial on Autonomous Fault-Tolerant Control Using Knowledge-Grounded LLM Agents
Fault recovery in process plants still relies heavily on plant operators, especially when faults fall outside predefined supervisory logic.…
Intrinsic decomposition and editing of 3D Gaussian splats
Intrinsic decomposition which expresses image colors as the product of diffuse albedo and shading, possibly augmented with view-dependent r…
A Lifecycle and Application-Stack Survey of Large Language Model Vulnerabilities: Attacks, Risks, Defenses, and Open Problems
Large language models are no longer only text generators. They are increasingly embedded in retrieval pipelines, enterprise assistants, cod…
ECHO: Prune to act, trace to learn with selective turn memory in agentic RL
Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Ex…
Improving Certified Robustness via Adversarial Distillation
Certified training aims to produce models whose predictions can be formally verified against adversarial perturbations, typically by optimi…
Sparsity-Inducing Divergence Losses for Biometric Verification
Performance in face and speaker verification is largely driven by margin-penalty softmax losses such as CosFace and ArcFace. Recently intro…
WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models
Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignor…
Histogram-constrained Image Generation
Diffusion models have emerged as a dominant paradigm in generative modeling, enabling high-fidelity sampling from complex data distribution…
When to Truncate a Feature Ranking: A Residual-Overlap Stopping Rule for Subset Selection
Feature rankings are widely used in supervised feature selection because they are simple, scalable and easy to interpret. Variables are fir…
ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping
The wave of AI-native applications is moving shopping beyond page- and feed-based browsing toward intent-driven experiences orchestrated by…
RCT: A Robot-Collected Touch-Vision-Language Dataset for Tactile Generalization
For robots manipulating open-world objects, tactile representations must generalize to unseen materials. We introduce RCT (Robotic Contact…
Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models
Sparse autoencoders (SAEs) have recently been proposed as interpretable tools for concept-level manipulation, under the assumption that iso…
Cross-lingual Relation Extraction with Large Language Models: Zero-Shot, Few-Shot, and Fine-Tuned Evaluation on Romanian
Relation extraction (RE) for low-resource languages is typically constrained by the lack of annotated corpora. We investigate the feasibili…
Seeing Is Not Sharing: Some Vision-Language Models Overestimate Common Ground in Asymmetric Dialogue
In collaborative dialogue, shared perception does not guarantee shared interpretation. Mutual understanding must be established through int…
STEB: Style Text Embedding Benchmark
While semantic embeddings are rigorously evaluated on the Massive Text Embedding Benchmark, the evaluation of style embeddings remains frag…
FedXDS: Leveraging Model Attribution Methods to counteract Data Heterogeneity in Federated Learning
Explainable AI (XAI) methods have demonstrated significant success in recent years at identifying relevant features in input data that driv…
JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering
Remote sensing change detection (CD) traditionally focuses on pixel-level binary segmentation, which identifies where changes occur but nei…
A Technical Typology of AI Systems in Public Administration
Research on artificial intelligence (AI) in the public sector often treats "AI" as a single category, neglecting technical distinctions bet…
CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
We study three complementary techniques for training compute-efficient language models. (1) Selective supervision and per-token efficiency.…
Geometry-Preserving Orthonormal Initialization for Low-Rank Adaptation in RLVR
Low-rank adaptation (LoRA) and its variants enable parameter-efficient fine-tuning of large language models under the supervised fine-tunin…
Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
Recent multimodal large language models have shown great promise in clinical image reasoning, but existing post-training pipelines remain p…
Real-Time Source-Free Object Detection
Real-world detectors for autonomous driving, surveillance, and robotics must handle domain-shifts under strict latency and memory constrain…
Bridging Local Observation and Global Simulation in Closed-Loop Traffic Modeling
A local-to-global context mismatch arises when autoregressive traffic simulators trained on ego-centric driving logs are deployed in global…
Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models
Vision-Language-Action (VLA) models offer a promising framework for robotic manipulation by connecting language instructions, visual observ…
Belief Contraction in Dynamic Epistemic Logic
Dynamic epistemic logic represents belief change via model transformations induced by epistemic events. Its standard formulation (Baltag, M…
Modal CEGAR-tableaux with RECAR and resolution-based SAT-shortcuts
We investigate two approaches for extending CEGAR-tableaux with SAT-shortcuts using a previously known approach called RECAR but also a tot…
Better Understanding, Understanding Better
"Any fool can know; the point is to understand." A well-known remark often attributed to Einstein captures a widely shared intuition: under…
Attend, Transform, or Silence: Operator-Level Visual Skipping for Efficient Multimodal LLM Inference
Multimodal large language models (MLLMs) increasingly process long visual-token sequences, increasing the overall inference computation. Ex…
MVP-Nav: Multi-layer Value Map Planner Navigator
Zero-shot Object Goal Navigation (ZSON) with RGB-only perception poses a fundamental challenge for embodied agents, as the absence of expli…
LeCropFollow: Latent Space Planning for Navigation in Unstructured Crop Fields
Unstructured navigational features, such as irregular planting or discontinuities, remain the primary failure mode for under-canopy agricul…
MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments
Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grou…
LUNA: Learning Universal 3D Human Animation Beyond Skinning
Creating photorealistic, animatable 3D human avatars from monocular images still largely depends on Linear Blend Skinning (LBS) and paramet…
GR2 Technical Report
Industrial recommendation systems serve billions of users through a multi-stage funnel -- retrieval, early-stage ranking, and re-ranking --…
Amplifying Membership Signal Through Chained Regeneration
The tendency of large generative models to memorize training data makes sample verification critical for privacy auditing and copyright enf…
Radial Suppression Accelerates Algorithmic Generalization: A Geometric Analysis of Delayed Generalization
Why do neural networks memorize algorithmic training data long before they generalize? We present a geometric case study demonstrating that…
TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands…
FLORA: A deep learning approach to predict forest attributes from heterogeneous LiDAR data
Forest attributes are essential for national-scale resource monitoring. Airborne LiDAR metrics are among the auxiliary variables most stron…
AdaJEPA: An Adaptive Latent World Model
Latent world models enable planning from high-dimensional observations by predicting future states in a compact latent space. However, thes…
Freeform Preference Learning for Robotic Manipulation
Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where spa…
When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors
While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.e., incorrectly citing or…
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes. Yet…
QVal: Cheaply Evaluating Dense Supervision Signals for Long-Horizon LLM Agents
LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings,…
Introspective Coupling: Self-Explanation Training Tracks Behavioral Change Despite Fixed Supervision
When does training language models (LMs) to generate explanations of their predictions yield faithful introspection, rather than superficia…
Disentangling Reasoning Logic to Resolve Explicit Knowledge Conflicts
Explicit knowledge conflicts, occurring when retrieved contexts contain contradictory information, pose a fundamental challenge for Large L…
A Concept of Possibility for Real-World Events
This paper offers a new concept of {\it possibility} as an alternative to the now-a-days standard concept originally introduced by L.A. Zad…
Deductive Logic in Language Models: Horizontal vs Vertical Reasoning
Recent language models exhibit significant logical reasoning abilities, yet the mechanisms supporting deductive inference remain poorly und…
LLM-Empowered Agentic MAC Protocols: A Dynamic Stackelberg Game Approach
Medium Access Control (MAC) protocols, essential for wireless networks, are typically manually configured. While deep reinforcement learnin…
Improving LLM Reasoning with Homophily-aware Structural and Semantic Text-Attributed Graph Compression
Large language models (LLMs) have demonstrated promising capabilities in Text-Attributed Graph (TAG) understanding. Recent studies typicall…
Paper2Rebuttal: A Multi-Agent Framework for Transparent Author Response Assistance
Writing effective rebuttals is a high-stakes task that demands more than linguistic fluency, as it requires precise alignment between revie…
ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences
The literature has witnessed an emerging interest in AI agents for automated assessment of scientific papers. Existing benchmarks focus pri…
Improved Upper Bounds for Slicing the Hypercube
A collection of hyperplanes $\mathcal{H}$ slices all edges of the $n$-dimensional hypercube $Q_n$ with vertex set $\{-1,1\}^n$ if, for ever…
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
Large vision-language models have endowed GUI agents with strong general capabilities for interface understanding and interaction. However,…
Diffusion Crossover: Defining Evolutionary Recombination in Diffusion Models via Noise Sequence Interpolation
Interactive Evolutionary Computation (IEC) provides a powerful framework for optimizing subjective criteria such as human preferences and a…
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research…
Containment Verification: AI Safety Guarantees Independent of Alignment
Agentic frameworks are the software layer through which AI agents act in the world. Existing safety methods intervene on the model and ther…
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework
LLMs have achieved remarkable success in complex reasoning tasks, yet current evaluation approaches predominantly rely on final-answer corr…
Meta-Programming for Linear-time Temporal Answer Set Programming
The development of temporal extensions of Answer Set Programming (ASP) has led to the emergence of non-monotonic linear-time (TEL), dynamic…
Prospect-Theory Behavior from Bellman Optimality in MDPs with Catastrophic States
We study risk-neutral control in Markov decision processes with an absorbing catastrophic state. Even though rewards are linear and the age…
Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches
While Reasoning Language Models (RLMs) are rapidly emerging as powerful tools for scientific research, their impact is primarily concentrat…
Don't Gamble, GAMBLe: An Analytical Framework for AI-Driven Research Systems
AI-Driven Research Systems (ADRS) -- systems coupling LLMs with automated evaluation to discover algorithms, proofs, and designs -- are bei…
When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning
Long-horizon reasoning requires a system to commit to medium-horizon intent without becoming rigid: re-plan too often and computation never…
IterCAD: An Iterative Multimodal Agent for Visually-Grounded CAD Generation and Editing
Computer-Aided Design is pivotal in modern manufacturing, yet existing automated methods predominantly rely on open-loop, one-shot generati…
Structural Preservation and the Logical Expressiveness of Graph Neural Networks
Bridges between graph neural networks (GNNs) and logical formalisms have been established by fixing architectural choices, such as the type…
DeXposure-Claw: An Agentic System for DeFi Risk Supervision
Decentralized finance exposes supervisors to fast-moving, networked credit risks. General-purpose LLM agents fit this setting poorly: they…
IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO
Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier langu…
Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles
This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In…
ATRIA: Adaptive Traceable ECG Reporting with Iterative Agents
Existing ECG report generation is tightly coupled -- interpretation and reporting fused end-to-end, so errors propagate without stage-level…
Quantization Inflates Reasoning: Token Inflation as a Hidden Cost of Low-Bit Reasoning Models
Quantization is widely used to reduce the inference cost of large language models, but its effect on reasoning models is not fully captured…
Clinical Harness for Governable Medical AI Skill Ecosystems
Medical AI remains organized around isolated models, whereas care requires accountable capabilities that persist across time. We define cli…
OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and…
Corruption Robust Offline Reinforcement Learning with Human Feedback
We study data corruption robustness for reinforcement learning with human feedback (RLHF) in an offline setting. Given an offline dataset o…
Perturbation Effects on Robustness and Individual Fairness
Deep neural networks are vulnerable to adversarial perturbations that can simultaneously degrade prediction robustness and individual fairn…
Learning by Surprise: Adaptive Mitigation of Model Collapse in Large Language Models
As AI-generated content increasingly populates the web, generative AI models are at growing risk of being trained on their own outputs, a p…
Compositional Concept-Based Neuron-Level Interpretability for Deep Reinforcement Learning
Deep reinforcement learning (DRL) has successfully addressed many complex control problems. However, the neural networks representing polic…
Mantis: Lightweight Foundation Model for Time Series Classification
While foundation models have revolutionized various domains, their application to time series classification remains rather under-explored,…
Verify when Uncertain: Beyond Self-Consistency in Black Box Hallucination Detection
Large Language Models (LLMs) often hallucinate, limiting their reliability in sensitive applications. In black-box settings, several self-c…
Artificial Intelligence in Sports: Insights from a Quantitative Survey among Sports Students in Germany about their Perceptions, Expectations, and Concerns regarding the Use of AI Tools
Generative Artificial Intelligence (AI) tools such as ChatGPT, Copilot, or Gemini have a crucial impact on academic research and teaching.…
SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA
As Large Language Models (LLMs) become increasingly used for question-answering (QA), relying on static, pre-annotated references for evalu…
TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels
Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. I…
A Reproducible Benchmark of Lightweight CNNs: Accuracy, Efficiency, and the Impact of Pretrained Initialization
Lightweight convolutional neural networks are often compared using results obtained with different training recipes, input settings, and pr…
Position: Collaborative Agentic AI Needs Interoperability Across Ecosystems
Collaborative agentic AI is projected to transform entire industries by enabling AI-powered agents to autonomously perceive, plan, and act…
From Multimodal Perception to Strategic Reasoning: A Survey on AI-Generated Game Commentary
The advent of artificial intelligence has propelled AI-Generated Game Commentary (AI-GGC) into a rapidly expanding research area, offering…
Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling
Recent advances in 3D neural representations and instance-level editing models have enabled the efficient creation of high-quality 3D conte…
LLM-Aided Joint Secrecy Precoding and Trajectory for RSMA-Based Heterogeneous UAV Networks
This paper investigates secure communications in rate-splitting multiple access (RSMA) enabled heterogeneous UAV networks, where multiple U…
VGGSounder: Audio-Visual Evaluations for Foundation Models
The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSo…
Physics-Constrained Fine-Tuning of Flow-Matching Models for Generation and Inverse Problems
We present a framework for fine-tuning flow-matching generative models to enforce physical constraints and solve inverse problems in scient…
Dataset Construction for Training LLM to Learn Analog Circuit Knowledge
This paper constructs a textual dataset for training large language models (LLMs) to learn analog circuit knowledge and customizes LLM trai…
Quantum Flow Matching
The flow matching has rapidly become a dominant paradigm in classical generative modeling, offering an efficient way to interpolate between…
One Shot vs. Iterative: Rethinking Pruning Strategies for Model Compression
Pruning is a core technique for compressing neural networks to improve computational efficiency. This process is typically approached in tw…
Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking
Although autoregressive (AR) models have demonstrated remarkable success in image generation, extending these models to layout-conditioned…
A Scalable Whole-body Motion Transfer via Implicit Kinodynamic Motion Retargeting
Human-to-humanoid imitation learning presents a promising pathway to address the severe data scarcity bottleneck in robotics by utilizing a…
Graph Coloring for Multi-Task Learning
When different objectives conflict with each other in multi-task learning, gradients begin to interfere and slow convergence, thereby poten…
SpecDetect4ML: Detecting Non-Local ML Code Smells with Code Property Graphs
Machine Learning (ML) pipelines encode quality-relevant decisions across data preparation, training, evaluation, and configuration code. So…
CharDiff-LP: A Diffusion Model with Character-Level Guidance for License Plate Image Restoration
License plate image restoration is important not only as a preprocessing step for license plate recognition but also for enhancing evidenti…
Not Every Time and Frequency Need to Be Forgotten in Diffusion Unlearning
Data unlearning aims to remove the influence of specific training samples from a trained model. In fine-tuning methods, data unlearning rel…
Human-Agent Collaborative Paper-to-Page Crafting
In the quest for scientific progress, communicating research is as vital as the discovery itself. Yet, researchers are often sidetracked by…
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
Graphical user interface (GUI) grounding is a key capability for computer-use agents, mapping natural-language instructions to actionable r…
Enhancing Graph Representations with Neighborhood-Contextualized Message-Passing
Graph neural networks (GNNs) have become an indispensable tool for analyzing relational data. Classical GNNs are broadly classified into th…
Optimal Self-Consistency for Efficient Reasoning with Large Language Models
Self-consistency (SC) is a widely used test-time inference technique for improving performance in chain-of-thought reasoning. It consists o…
Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
Audio-language pretraining (ALP) holds promise for learning general-purpose audio representation, yet remains underexplored. Crucially, the…
REMSA: Foundation Model Selection for Remote Sensing via a Constraint-Aware Agent
Foundation Models (FMs) are increasingly integrated into remote sensing (RS) pipelines. These models include unimodal vision encoders and m…
Rethinking Garment Conditioning in Diffusion-based Virtual Try-On: Decouple, Don't Denoise
Virtual Try-On (VTON) synthesizes realistic images of a person wearing a target garment, with broad applications in e-commerce and fashion.…
A Unified and Stable Risk Minimization Framework for Weakly Supervised Learning with Theoretical Guarantees
Weakly supervised learning has emerged as a practical alternative to fully supervised learning when complete and accurate labels are costly…
FMA-Net++: Motion- and Exposure-Aware Joint Video Super-Resolution and Deblurring
Joint video super-resolution and deblurring (VSRDB) requires both efficient long-range temporal modeling and robustness to frame-wise expos…
The HydroGym Reinforcement Learning Platform for Fluid Dynamics
Modeling and controlling fluids is critical across science and engineering. Effective flow control can increase lift, reduce drag, enhance…
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
Distilling the capabilities from a large reasoning model (LRM) to a smaller student model often involves training on substantial amounts of…
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
GUI agents that interact with graphical interfaces on behalf of users represent a promising direction for practical AI assistants. However,…
From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching
Semantic caching has emerged as a pivotal technique for scaling LLM applications, widely adopted by major providers including AWS and Micro…
Toxicity Assessment in Preclinical Histopathology via Class-Aware Mahalanobis Distance for Known and Novel Anomalies
Drug-induced toxicity is a leading cause of preclinical and early-clinical failure, making early detection critical. Histopathology is the…
Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinity
Tail-end risk measures such as static conditional value-at-risk (CVaR) are used in safety-critical applications to prevent rare, yet catast…
DeXposure-FM: A Time-series, Graph Foundation Model for Credit Exposures and Stability on Decentralized Financial Networks
Credit exposure in Decentralized Finance (DeFi) is often implicit and token-mediated, creating a dense web of inter-protocol dependencies.…
A swap-adversarial framework for improving domain generalization in electrocorticography-based Parkinson's disease classification
We propose a novel swap-adversarial framework that mitigates high inter-subject variability and the high-dimensional low-sample-size proble…
CoReLIN: Constraint-based Reasoning for Zero-shot Lifelong Interactive Navigation
Robot navigation typically assumes an obstacle-free path exists between start and goal. In real environments, however, clutter may block al…
DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depicting targ…
Are Video Reasoning Models Ready to Go Outside?
In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion. Under such con…
RCTs for Frontier AI Governance: Methodological Challenges and Solutions for Human Uplift Studies
Human uplift studies, or studies that measure the effects of AI access on human performance via randomized controlled trials (RCT) or simil…
Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models
Reinforcement learning (RL) has become a standard technique for post-training diffusion-based image synthesis models, as it enables learnin…
Visual Prompt Discovery via Semantic Exploration
LVLMs encounter significant challenges in image understanding and visual reasoning, leading to critical perception failures. Visual prompts…
An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU
Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabili…
Can VLMs Reason Robustly? A Neuro-Symbolic Investigation
Vision-Language Models (VLMs) have been applied to a wide range of reasoning tasks, yet it remains unclear whether they can reason robustly…
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing…
Learning Dexterous Grasping from Sparse Taxonomy Guidance
Dexterous manipulation requires planning a grasp configuration suited to the object and task, which is then executed through coordinated mu…
SMART: When is it Actually Worth Expanding a Speculative Tree?
Tree-based speculative decoding accelerates autoregressive generation by verifying a branching tree of draft tokens in a single target-mode…
RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora
Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) system…
Lyapunov-Certified Direct Switching Theory for Q-Learning
Q-learning is a fundamental algorithmic primitive in reinforcement learning. This paper develops a new framework for analyzing Q-learning f…
Generalizing Numerical Reasoning in Table Data through Operation Sketches and Self-Supervised Learning
Numerical reasoning over expert-domain tables often exhibits high in-domain accuracy but limited robustness to domain shift. Models trained…
Shared Lexical Task Representations Explain Behavioral Variability In LLMs
One of the most common complaints about large language models (LLMs) is their prompt sensitivity -- that is, the fact that their ability to…
Structured Progressive Knowledge Activation for LLM-Driven Neural Architecture Search
This paper focuses on a key challenge in Neural Architecture Search (NAS): integrating established architectural knowledge while exploring…
On Variance Reduction in Learning Mean Flows
One-step generative modeling has emerged as a leading approach for amortizing the inference cost of diffusion and flow-matching models. Amo…
KV-RM: Regularizing KV-Cache Movement for Static-Graph LLM Serving
Static-graph LLM decoders provide predictable launches, fixed tensor shapes, and low submission overhead, but online decoding exposes highl…
An Executable Benchmarking Suite for Tool-Using Agents
Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often…
Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines
Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning…
Knowledge Graphs as the Missing Data Layer for LLM-Based Industrial Asset Operations
LLM-based agents for industrial asset operations show limited accuracy when reasoning over flat document stores. AssetOpsBench (KDD 2026) e…
BenGER: Benchmarking LLM Systems on Subsumption-Based Legal Reasoning in German Law
We introduce BenGER (Benchmark for German Law), a benchmark and dataset for evaluating LLM systems on subsumption-based legal reasoning in…
Quantitative Movement Testing: Measuring Chronic Pain Patient Movements from a Single Smartphone Video
Chronic pain diminishes quality of life by decreasing functional ability, yet objectively measuring this functional impact remains challeng…
INFUSER: Influence-Guided Self-Evolution Improves Reasoning
Self-evolution offers a scalable path to stronger reasoning: a pretrained language model improves itself with only minimal external supervi…
Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents
Long-horizon tool-use reinforcement learning learns from outcome verification, but trajectory-level advantages are broadcast over reasoning…
ComAct: Reframing Professional Software Manipulation via COM-as-Action Paradigm
Existing computer-use agents remain fundamentally limited in professional software manipulation: GUI-based agents suffer from fragile visua…
Same-Origin Policy for Agentic Browsers
Agentic browsers integrate autonomous AI agents into web browsers, enabling users to accomplish web tasks through natural-language instruct…
PSCT-Net: Geometry-Aware Pediatric Skull CT Reconstruction via Differentiable Back-Projection and Attention-Guided Refinement
Computed Tomography (CT) is essential for diagnosing pediatric craniofacial abnormalities, yet poses radiation risks to developing anatomie…
Protein contacts are already in the attention: a single-forward-pass alternative to the Categorical Jacobian
The Categorical Jacobian of Zhang et al. (2024) reads protein contacts from a language model by perturbing every residue with every alterna…
RigorBench: Benchmarking Engineering Process Discipline in Autonomous AI Coding Agents
Agentic coding harnesses - such as Agent-Skills, Superpowers, and Agent-Rigor - are increasingly deployed to augment underlying LLMs for re…
The Geometry of Refusal: Linear Instability in Safety-Aligned LLMs
Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque. In this work,…
Measuring & Mitigating Over-Alignment for LLMs in Multilingual Criminal Law Courts
While the wider applicability of LLMs in the legal field is currently debated due to their reliability and the gravity of any errors, narro…
One Year Later...The Harms Persist, But So Do We!
General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety guardrails remain in…
Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models
Many modern Language Model (LM) pipelines return an averaged model, such as an exponential moving average of the training iterates, rather…
Brevity is the Soul of Inference Efficiency: Inducing Concision in VLMs via Data Curation
Inference efficiency is typically pursued by shrinking the model: distillation, pruning, quantization, and sparse routing each lower per-to…
Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching
Entity Matching (EM) is a core operation in the data integration pipeline, where records from different sources are compared to determine w…
「Claude Sonnet 5」新登場 低コストでOpus 4.8に匹敵とうたうも、タスク当たりコスト増加との評価も
Anthropicが新モデル「Claude Sonnet 5」を発表した。上位の「Opus 4.8」に迫る性能を低価格で実現するとうたう。一方、第三者機関の評価ではトークン使用量が多く、1タスク当たりコストはOpus 4.8を上回るとの指摘もある。
The “Father of the Internet” is finally retiring
Vinton Cerf, one of the creators of the protocols underlying the internet, will step down as Google's chief internet evangelist next week.
国産LLM「Sarashina3」登場 高品質データ、独自検証で日本語能力を強化 ソフトバンク傘下
ソフトバンク傘下のSB Intuitionsは、国産LLM「Sarashina」の最新版「Sarashina3シリーズ」の提供を開始。高品質なデータセットや独自の出力結果検証などで日本語能力を強化した。
Trump drops restrictions on Anthropic’s Mythos and Fable models
Anthropic said it would begin restoring access to the Fable on July 1.
キーボード入力時の脳の動きから打った文章を割り出す技術、Metaが発表 手術不要、“埋め込み式”に迫る
米Metaが、手術を伴わずに脳の活動を文章へ変換する研究「Brain2Qwerty v2」を発表した。頭部に装着する装置で脳の信号を読み取り、人がキーボードへ入力した文章をリアルタイムで解読する。脳の病気で話す力を失った人の意思疎通を支える技術として、学習用のコードも公開した。
Wayve launches $85M employee tender offer at $8.5B valuation
Wayve’s offering is part of a growing trend of AI startups using employee tenders as a strategic tool to attract and retain talent.
「Claude Fable 5」が帰ってくる 「Mythos 5」含む輸出規制解除へ Anthropic発表
Anthropicは6月30日(現地時間)、「Claude Fable 5」「Mythos 5」への輸出規制が解除されたと明らかにした。7月1日からアクセスを回復し、詳細は近日中に発表するとしている。
米政府「Fable 5/Mythos 5」の輸出規制解除 Anthropic「明日からアクセス回復」
「ユーザーの辛抱に感謝するとともに、モデルの再展開に協力いただいたすべての方々に感謝する」
Anthropic、科学研究向けAIワークベンチ「Claude Science」を発表──NVIDIAのBioNeMoツールキットと連携
Anthropicは、科学者が計算研究を一貫して行えるAI実行環境「Claude Science」を発表した。データベースやツールを1つのインタフェースに統合し、文献分析から論文執筆、図表作成まで対応する。NVIDIAのツールキットとも連携し、機密データを外部に送信しない設計が…
CAD連携AIで設計レビュー工数を最大40%削減、検図や見積作成も自動化
Archaicは、製造業の設計業務を自動化するAIソリューションの販売を開始した。SOLIDWORKSなどのCADと直接連携し、確認作業や検図、見積作成を自動化して、設計者の工数削減を支援する。
AIで“ゲームキャラの出産二次創作”を何千回と生成する人も……ChatGPTの会話57万件から見えたヘビーな利用実態
米ワシントン大学などに所属する研究者らが発表した論文「AI Fiction in the Wild」は、AIチャットとの会話データを分析し、ユーザーがAIを使ってどれだけフィクションを生成しているのかを調べた研究報告だ。
日産「AIで再び世界トップの開発力へ」、独自の統合型次世代AIDV基盤描く
日産自動車は「AWS Summit Japan 2026」において、次世代モビリティ「AIDV(AIディファインドビークル)」に向けたクラウド基盤構築の取り組みと、AIを活用したソフトウェア開発環境の今後の展望を語った。
「ウソだろ」アスクル社長がうなったAI活用 商談準備を2週間→3時間に “担当者のカオス”脱却へ
サイバー攻撃を受けたアスクルが、AIを活用して自社システムを立て直した。旧来の課題を拭い去り、商談時間を短縮するなどの成果を出している。逆境を勝機に変えた舞台裏を、吉岡社長が語った。
生成AIの請求書、人件費と並べる時代へ 国内5社のAI責任者が語る「トークンマネジメント」の現在地
経費精算SaaSのLayerXやラクス、名刺管理から事業を広げたSansan、会計クラウドのfreee、フリマアプリのメルカリ。取材した5社のAI・人事責任者から、驚くほど重なるトーンでAIのトークンコストを語る声が聞こえてきた。
謎の「“日の丸AI”開発企業」正体明らかに ソフトバンク、NECら大手がそろって出資するワケ
ソフトバンクやNECなどが出資する「国産AIモデル開発企業」がベールを脱いだ。一体なぜ、国内大手企業が出資するのか。
OpenClaw is finally available on Android and iOS
The free open source agentic program is finally invading your phone.
The DeepMind trio who built a poker AI are now making money for quant hedge funds
EquiLibre Technologies, a Prague-based AI lab founded by three ex-DeepMind researchers, is now valued at more than $500 million.
【Pythonで学ぶデータ分析】母平均のベイズ推定と予測サンプルの作成 ~ 規格外製品の廃棄コストを見積もる
「不良品で幾ら損する?」をベイズ統計で見積もってみましょう。製品のサイズを測ったデータから、平均やばらつきを推定し、さらに「これから作る製品が規格外になる確率」までをPythonを使って予測します。『社会人1年生から学ぶ、やさしいデータ分析』ベイズ統計編の第4回です。
Google introduces a faster, cheaper image generator with Nano Banana 2 Lite
Google is updating its image generator to make it faster and cheaper, making it a more useful tool for creators looking to make AI content.
Nvidia competitor Etched hits $5B valuation, $1B in sales for AI chip
Nvidia AI chip competitor Etched says it has already booked $1 billion under contract for the inference systems powered by its chip.
Anthropic launches Claude Sonnet 5 as a cheaper way to run agents
Anthropic’s Claude Sonnet 5 brings stronger agentic capabilities, lower pricing, and improved safety, positioning the model as a cheaper al…
Acti puts AI agents directly into your smartphone keyboard
Acti is betting the smartphone keyboard is the next home for AI assistants. The startup's new keyboard for iOS and Android works across app…
Anthropic’s Claude Science bets on workflow, not a new model, to win over scientists
Anthropic's Claude Science is a workbench that gives scientists one environment to do computational research, saving them from the need to…
X now offers an MCP server to make its platform easier for AI tools to use
X has launched a hosted MCP server, making it easier for developers to connect AI applications with the company’s API.
Podcasting platform Riverside enters the newsletter publishing game
Users will be able use AI to create newsletters based on their recordings.
Amazon launches new $1 billion FDE org, following OpenAI and Anthropic
Engineers on the new team will embed within companies to deploy purpose-built agents, focusing on fast deployments and customer self-suffic…
2026-06-30(688件)
Lumo, Proton’s privacy-focused AI chatbot, gets an upgrade
Proton's Lumo 2.0 is dropping this week, giving users a broader variety of capabilities.
農水省の“クソダサ”ポスター話題 「AIよりよっぽど良い」の声も 担当者に狙いを聞いた
農林水産省の公式Xアカウントが6月29日に投稿した「佃煮の日」のポスターが、良い意味で「ダサい」と話題だ。農水省広報室の担当者はITmedia NEWSの取材に応じ、デザインの狙いやSNSの投稿体制について語った。
AIロボット1000万台導入へ、2040年までに 赤澤経産相が語る「勝ち筋」
赤澤亮正経済産業大臣は6月30日の記者会見で、2040年までにAIを活用したロボットを国内に約1000万台導入する目標を掲げた。18分野での社会実装を進める。
Crypto exchange OKX wants AI agents to hire and pay each other
OKX is bringing together payments, identity, and reputation into a marketplace for AI agents.
How ChatGPT adoption has expanded
New OpenAI Signals data shows how ChatGPT adoption is growing globally, with users increasing usage, exploring more capabilities, and drivi…
スクエニ「AI駆動型品質チェックプラットフォーム」開発へ、国が補助金 バンナムやnoteも採択
スクウェア・エニックスの他、バンダイナムコエンターテインメント、NTT西日本、noteなどが採択された。
日印「防衛用AIドローン」共同開発へ 首脳会談で確認、対中念頭に安保協力深化
日印両政府が防衛分野で活用する人工知能(AI)搭載型ドローン(無人機)の共同開発を推進する方針を固めた。高市早苗首相は7月2日にインドでモディ首相との会談を予定しており、防衛装備品協力を加速させることで一致する見通しだ。中国がインド太平洋地域で軍事活動を活発化させる中、日印の安…
The AI jobs debate just got messier
A new report finds "high-intensity AI adopters” saw headcount increase 10.2%. Among those companies, entry-level headcount rose by 12%, cou…
Recursive Self-Evolving Agents via Held-Out Selection
LLM agents are increasingly improved without weight updates by evolving a natural-language artifact, such as reflections, workflows, playbo…
Data and Evaluation Closed-Loop for Model Capability Enhancement
Model capability is the central variable in LLM pre-training, yet is never observed directly: data shapes it prospectively, while evaluatio…
GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes
Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. Existing benchmarks show…
IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations
Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for…
Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories
Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and "Wait" prompts, primarily encourage models…
Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas
Large Language Models (LLMs) often face ethical tradeoffs in which several responses may be defensible but express different priorities, su…
An AI agent for treatment reasoning over a biomedical tool universe
Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and e…
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models
Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimod…
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards
Critic-free reinforcement learning with verifiable rewards (RLVR), exemplified by Group Relative Policy Optimization (GRPO), avoids trainin…
The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance
We ask under what conditions an agent with a harm-minimizing policy can displace an approval-seeking (RLHF) agent in a competitive market,…
TrajRS: Towards Certified Robustness in Pedestrian Trajectory Prediction
The robustness of trajectory prediction models is crucial for developing safe autonomous driving systems. Adversarial attacks on trajectory…
ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models
Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. H…
Agentic Abstention: Do Agents Know When to Stop Instead of Act?
LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not e…
Agent Safety Is Action Alignment
Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To kee…
Self-Supervised Theorem Discovery in a Formal Axiomatic System
Recent artificial intelligence (AI) systems have shown remarkable progress in mathematical reasoning. Many existing approaches, including l…
Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions
Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous effo…
HyphaeDB: A Living Knowledge Topology for Agent-First Memory
Every existing vector database and agent memory framework treats memory as passive storage that agents query explicitly. No system propagat…
Primary ICD Category Prediction using LLM-based Probing
Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggl…
MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes
Doctor agents are moving beyond single-turn answer generation toward evolving clinical decision systems. Within an outpatient episode, they…
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries
Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-sty…
Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline
Recent advancements in generative artificial intelligence (AI) and large language models (LLMs) have shown significant promise in automatin…
Preventing Error Propagation in Multi-Agent AI through Runtime Monitoring
Multi-agent AI systems can improve answer selection by allowing different language models to exchange reasoning traces, revise initial pred…
Memory as an Attack Surface in LLM Agents: A Study on Multiple-Choice Question Answering
AI agents extend conventional large language model (LLM) applications by integrating language understanding with task execution, external t…
Low-cost concept-based localized explanations: How far can we get with training-free approaches?
Concept-based Explainable AI (C-XAI) seeks human-understandable explanations grounded in semantic concepts, yet validation is limited by th…
Managing the Human Fallback: Skill Investment Under Improving AI and Worker Mobility
When firms deploy autonomous AI, they must decide how much work to leave to the system and how much to keep workers engaged. This decision…
Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem
Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflow…
HiComm: Hierarchical Communication for Multi-agent Reinforcement Learning
Cooperative multi-agent reinforcement learning (MARL) often relies on communication to mitigate partial observability, yet most existing pr…
Flow Reasoning Models: Scaling Reasoning Through Iterative Self-Refinement
Discrete flow models have recently shown promising performance on few-step text generation; however, when naively applied to structured rea…
Pooled Leaderboards Hide System-Specific Winners: A Reporting-Protocol Audit of Offline Root-Cause Analysis Benchmarks
Offline root-cause-analysis (RCA) benchmarks commonly rank methods by a single pooled top-1 accuracy across multiple subsystems, and engine…
Direct Causation in International Humanitarian Law and the Challenge of AI-Mediated Civilian Cyber Operations
International humanitarian law protects civilians from direct attack unless and for such time as they take direct part in hostilities, with…
Selective Memory Retention for Long-Horizon LLM Agents
When does retention matter for memory-augmented LLM agents? We study this with TraceRetain, a lightweight framework for bounded external me…
Measuring Graph-to-Graph Semantic Similarity in Knowledge Graphs: An Empirical Evaluation of Knowledge Graph Embeddings
A Knowledge Graph (KG) represents facts as structured triples and is widely used to organize relational knowledge across diverse domains. J…
Evidence-Informed LLM Beliefs for Continual Scientific Discovery
Open-ended scientific discovery with large language models (LLMs) increasingly operates as a long-horizon loop of hypothesis search and ver…
AI Trading's Alpha Singularity: Emergent Market Reasoning through Agent-to-Agent Self-Evolution
Automated alpha mining holds the scoring function fixed and varies the search algorithm over it. A search that converges against a fixed sc…
A Cognition-Emotion-Personality Framework for Modeling Human-Like Awareness and Behavior in Emergency Evacuations
Agent-based evacuation simulations are widely used to study crowd behavior during emergencies, but many models rely on assumptions such as…
PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents
LLM agents handle user requests on behalf of organizations through tool calls and must follow the company policies stated in their system p…
SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics
Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics. Despite the prevalence…
When Summaries Distort Decisions: Information Fidelity in LLM-Compressed Financial Analysis
Financial decision-makers face more information than they can directly inspect, making context compression necessary. Yet when large langua…
The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling
We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of requir…
Process Advantage Signal Shaping: A Paradigm-Agnostic Middleware for Process-Supervised RL in LLM Reasoners
Group Relative Policy Optimization (GRPO) is a default recipe for process-supervised reinforcement learning of LLM reasoners, and dense pro…
Hierarchical Experimentalist Agents
Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rel…
PHF: Privileged Hidden Flow for On-Policy Self-Distillation
On-policy self-distillation (OPSD) trains a reasoning model on rollouts sampled from its own policy by matching a privileged teacher that a…
When LLMs Develop Languages: Symbolic Communication for Efficient Multi-Agent Reasoning
Chain-of-Thought (CoT) improves large language models (LLMs) on difficult reasoning tasks, but it often incurs long natural-language ration…
Diagnosing and Repairing Factual Errors in RAG under Budget Constraints
Retrieval-Augmented Generation (RAG) improves the factuality of large language models by grounding responses in external evidence, yet real…
LLM-Guided Planning for Multi-hop Reasoning over Multimodal Nuclear Regulatory Documents
Reviewing nuclear regulatory documents requires multi-hop reasoning across tens of thousands of pages, where judgments depend on evidence a…
Mixture of Debaters: Learn to Debate at Architectural Level in Multi-Agent Reasoning
Existing multi-agent debate frameworks suffer from two critical limitations: they rely on static architectures where agent roles and coordi…
FADE: Mitigating Hallucinations by Reducing Language-Prior Dominance in Large Vision-Language Models
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucination, generating content i…
How Much Due Diligence Before You Bid? Learning in Intractable Takeover Auctions
When two companies bid to buy the same target, no one knows exactly what the target is worth. Each bidder pays for due diligence: costly, i…
Agent-Computer Observation Interfaces Enable Dynamic Computer Use
SWE-agent established the action interface as an underexplored design axis for software-engineering agents; we make the analogous case for…
Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving
Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a…
Cognitive World Models for Process-Level Social Influence Evaluation
Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user'…
UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation
Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not orac…
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting the…
Learned Coordination Conventions in Cooperative MARL: Measuring the Translation Gap Between Theory-Informed Roles and Learned Routing
Role-semantic assignments provide priors over how heterogeneous agents may coordinate, but cooperative MARL systems instead settle on conve…
SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings
Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo…
SFBench: The SciFy Scientific Feasibility Benchmark
We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claim…
Budgeted Act-or-Defer Multi-Agent LLM Deliberation with Local Reliability Bounds
Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to a…
Safety from Honesty in a Disinterested AI Predictor
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directe…
Diversity is the Strength of the AI Crowd
Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf…
Sample-Efficient Learning of Probabilistic Causes for Reachability in Markov Decision Processes with Probabilistic Guarantees
Probabilistic model checking for Markov decision processes (MDPs) provides quantitative guarantees, but often offers limited insight into w…
Toward Secure and Reliable PDDL Formalization of Large Language Models with Planner-in-the-Loop Feedback
Planning often requires symbolic specifications that are both executable and verifiable. For large language models deployed in autonomous o…
GUICrafter: Weakly-Supervised GUI Agent Leveraging Massive Unannotated Screenshots
Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, rese…
DeepTrans Studio: Turning Expert Interventions into Shared Team Knowledge in Agentic Translation Workflows
Professional translation is often a team-based process: translators, reviewers, and project managers must coordinate terminology, legal for…
DEEPMED Search: An Open-Source Agentic Platform for Medical Deep Research with Introspective Verification
Navigating the deluge of heterogeneous medical data, from academic literature (PubMed) to clinical guidelines (Web) and private knowledge b…
Rethinking Generative Reconstruction Attacks against Graph Neural Network Models
The application of graph data in numerous disciplines raises the need for gathering and analyzing huge volumes of data, some of which is pr…
CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents
LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequentia…
The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models
This project introduces the CRISTAL Method (Coherent Reliable Intentional Synthesis of Truthful Analysis Logic), a neurosymbolic framework…
Beyond Triplet Plausibility: Relation Set Completion in Knowledge Graphs
Knowledge graphs (KGs) organize real-world knowledge as triplets and underpin many downstream applications. Due to their inherent incomplet…
AI Training Manager: Bounded Closed-Loop Control of Adaptive Training Recipes
We present the AI Training Manager, a bounded LLM-based supervisory controller for adaptive machine learning training. Standard training pi…
SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing
In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safe…
A causal modeling perspective on decision theory
Decision theory provides a formal framework for how agents should make choices under uncertainty, drawing on ideas from philosophy, probabi…
HippoSpark: An On-Demand Experience System for LLM Reasoning
Distilling historical trajectories into reusable experience to enhance future problem-solving has become a focal point of recent LLM resear…
SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon CivRealm Strategy Planning
Long-horizon strategic planning in complex strategy games demands concurrent reasoning across multiple decision domains under imperfect inf…
First-Order Temporal Logic Tensor Networks
Most of the existing neuro-symbolic AI methods focus on the scenario of static knowledge where objects do not change according to a tempora…
Exploration and Online Transfer with Behavioral Foundation Models
Zero-shot Transfer in Reinforcement Learning (RL) aims to train an agent that can generate optimal policies for any reward function, withou…
Be Faithful When Response: Returning Fluent and Grounded Answers for Vision-Language Models Reinforcement Learning
Reinforcement Learning (RL) is an important paradigm for improving the reasoning capabilities of Vision-Language Models (VLMs). However, di…
AlgoSkill: Learning to Design Algorithms by Scheduling Human-Like Skills
Designing an algorithm from a natural-language problem statement requires identifying the problem structure, reading constraints, choosing…
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centraliz…
SAT-RTS: A systematic framework for tactical knowledge extraction and visualization-based analysis in real-time strategy games
Efficient tactical knowledge extraction and analysis in real-time strategy (RTS) games micromanagement are constrained by the high-dimensio…
Hierarchical Reinforcement Learning in StarCraft Micromanagement with Influence Maps and Cluster-based Scripts
Real-time strategy (RTS) games present significant AI challenges, characterized by expansive state-action spaces arising from multi-unit co…
Temporal Feature Extractors in EEG Foundation Models: A Controlled Comparison Including a Pretrained Time-Series Model
Electroencephalography (EEG) foundation models aim to learn generalizable representations from large-scale brain recordings. However, the r…
Propagation of~Interval Belief Structures and~Imprecise Copulas for~Neural Network Verification
Quantitative verification of neural networks requires reasoning about probabilities under substantial uncertainty in both input distributio…
Structural Certification for Reliable Physical Design with Language Models
An unreliable language model can be made to produce reliable physical designs if the authority to assert is moved out of the model: the mod…
Open Problems in Constitutional Preference Reconstruction
Pairwise preference data is widely used for training and evaluating language models (e.g., RLHF), but each datapoint records a \emph{choice…
Does Verbose Chain-of-Thought Really Help? In-Distribution Evidence that Content, Not Length, Matters
Chain-of-thought (CoT) prompting improves LLM reasoning, but the source is contested: do the intermediate steps help because they carry use…
Relevance Is Not Permission: Warranted Attention for Value Contributions
Relevance is not permission. Attention lets a model read key-value items related to the current query, but it does not guarantee that the v…
FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational Avatars
Natural face-to-face conversation requires real-time speech generation together with synchronized facial motion. Existing systems only part…
MirrorCode: AI can rebuild entire programs from behavior alone
AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C…
Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents
Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dy…
From Detecting Agency to Doing Work: Self-Caused Credit Builds a Durable Behavioral Self in a Minimal Spiking Agent
How does an agent that can tell self from world come to be durably shaped by that distinction? Recent work shows that a predictive system c…
Domain Adaptation with Adaptive Imagination for Visual Reinforcement Learning under Limited Target Data
Sim-to-real transfer remains a major obstacle for reinforcement learning (RL), especially for vision-based control where image observations…
The Many-Body Problem of the Data Centre
Modern Artificial Intelligence is often framed as limited by its own disembodiment, as if giving it a body would unlock its true potential.…
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can imp…
Clarus: Coordinating Autonomous Research Agents toward Web-Scale Scientific Collaboration
Existing autonomous research agents can support parts of the research process, but most systems still treat research as either an isolated…
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
Inoculation prompting is a selective generalization technique used against Emergent Misalignment. We introduce inoculation adapters (IA), w…
EMPATH: A Multilingual Auditor-Judge Benchmark for Safety Evaluation of Emotional-Support Chatbots
Safety benchmarks often buy scalability by fixing the prompt, the language, and the turn structure. For emotional-support chatbots, that ba…
PromptGNN-sim: Deep Fusion and Alignment of GNN and LLMs for Text-Attributed Graph Learning
Text-Attributed Graphs (TAGs) combine textual semantics with graph structure and are central to many graph learning tasks. However, existin…
Rehearsed Multi-Agent Live Product Demonstrations with Real-Time Voice Question Answering
Live product demonstrations are a recurring, high-cost activity in software organizations: a human presenter must select features, dispatch…
ManimAgent: Self-Evolving Multimodal Agents for Visual Education
Multi-round reflection lets agents built on large language models recover from failures within a single task, but each task remains an isol…
BayesEvolve: Explicit Belief States for Autonomous Scientific Discovery
Autonomous scientific discovery systems increasingly use large language models (LLMs) to propose new hypotheses, but many such systems cond…
Sequential Fairness Auditing with Limited Output Access
External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often ha…
Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data
Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject…
Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents
A rapidly growing class of LLM agents is multi-party: the agent acts for a principal (who briefs it, sends follow-ups, and receives results…
ENC-ODE: Event-level Neurodegenerative Modeling in Continuous Time with Neural ODEs
Accurately predicting the temporal evolution of clinical biomarkers is crucial for the early diagnosis and management of neurodegenerative…
The FIL Hypothesis: Inductive Biases Help with Kernel Engineering
The Bitter Lesson, which posits that general-purpose methods that scale with computation and data ultimately outperform those with built-in…
Entity Binding Failures in Tool-Augmented Agents
Tool-augmented language-model agents are often evaluated by whether they select the correct tool, produce valid API arguments, and complete…
Latent Actions from Factorized Transition Effects under Agent Ambiguity
Latent Action Models (LAMs) learn action-like proxies from observation transitions. However, in multi-object or distractor-rich scenes, the…
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents colla…
The Human Creativity Benchmark
Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflec…
DOPD: Dual On-policy Distillation
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals.…
Self-Evolving World Models for LLM Agent Planning
World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. H…
Tool-Augmented Spatiotemporal Reasoning for Streamlining Video Question Answering Task
Video Question Answering (VideoQA) task serves as a critical playground for evaluating whether foundation models can effectively perceive,…
It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents
Phone-use Agents can execute complex tasks end to end across real mobile applications. By operating a real device on the user's behalf, the…
ADEPT: An Entropy-Driven Dual-Strategy Agent for Interactive Video Retrieval
This research aims to solve the challenge of video retrieval from massive datasets, caused by ambiguous user queries. Prevailing single-rou…
The Interference Gap: Comparing Retrieval Bounds in Human Memory and RAG Systems
How do retrieval bounds compare between human episodic memory and Retrieval-Augmented Generation (RAG) systems under semantic interference?…
$M^3 QuestionIng$: Multi-modal Multi-span Medical Question Answering
The growing adoption of AI in healthcare, particularly in preventive care, highlights the critical need for accessibility and precision in…
High-Dimensional Concentration and Retrieval Instability in Embedding Spaces: Implications for Retrieval-Augmented Generation
Embedding-based retrieval systems rely on the assumption that geometric proximity in highdimensional representation spaces reflects semanti…
"AI Watermarking": Bridging Policy Discourse and Technical Capabilities
The widespread deployment of generative artificial intelligence (AI) models has raised serious concerns about the proliferation of AI-gener…
When Medical Safety Alignment Fails: A Benchmark for Evaluating LLMs on High-Risk Medical Queries
Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenario…
Insidious by Design: Implications of Large Language Model algorithmic bias for the Global South
\begin{quote} The biases in Large Language Models' (LLMs) outputs remain inadequately theorised, particularly from the perspective of the G…
Ground Truths in Suicide Research: The Current State of AI-Based Suicide Detection in Social Media
Recent advances in artificial intelligence (AI) and social media data have led to growing optimism about the ability to detect suicide risk…
LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution
We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional dist…
HyBIRD: Hyperbolic Bridge Retrieval and Diagnosis for Methodology Inspiration Retrieval
Methodology Inspiration Retrieval (MIR) asks a system to retrieve prior papers whose methods can inspire a new research proposal. Unlike ge…
A Systems-Level Analysis of Sensitivity, Robustness, and Stability in Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) systems are often evaluated using final answer accuracy, even though their failures can originate from…
Multi-Agent DRL for QoS and Energy Optimization in RIS-Enabled Open-RAN Industrial 6G TN/NTN Networks
Industrial 6G networks require ultra-reliable, low-latency, and energy-efficient connectivity in dynamic and blockage-prone environments, w…
Operating Regimes of Decentralized Learning Under Mobility and Bandwidth Constraints
Decentralized learning is a promising paradigm for collaborative training in mobile and pervasive systems, as it avoids a central coordinat…
The Crowded Embedding Space: A Mean-Field Mechanism for Emergent Marginalization in Retrieval-Augmented Agents
Retrieval-augmented generative agents rely on retrieval for grounding, yet are typically evaluated on a query-by-query basis. This isolates…
PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existi…
Auditing LLM-Governed Social Robots with Culture-Specific Moral Gradients
LLM-governed social robots increasingly decide who receives real-world assistance first. As prioritization norms vary across cultures by ag…
Agentic Safety is an Epistemic Property, Not a Behavioral One
Contemporary AI safety spans pre-training interventions, post-training alignment, deployment-time controls, monitoring, and red-teaming. Th…
HMARS: A Hierarchical Multi-Agent Memory System for Long-Context Reasoning
Long-context reasoning requires models to access, retrieve, and integrate evidence scattered across documents, dialogues, and accumulated i…
From Regulatory Approvals to Patents: Cross-Domain Linking for Cardiovascular Device Traceability
Linking FDA-approved medical devices to their underlying United States Patent and Trademark Office (USPTO) patents enables critical applica…
SafeGEO: Understanding Generative Engine Optimization Risks in Recommendation Agents
Generative Engine Optimization (GEO) lets content owners rewrite web content to increase their visibility in generative systems. In recomme…
ReasonRec: A Reasoning-Augmented Multimodal Agent for Unified Recommendation
Recent advances in multimodal recommenders excel at feature fusion but remain opaque and inefficient decision-makers, lacking explicit reas…
How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) aims to enhance the trustworthiness of Large Language Models (LLMs) by grounding their outputs in exte…
Carolina Guide: A Multi-Agent RAG System with Institutional Guardrails for Academic Policy Assistance
University students often struggle to navigate complex academic policies, leading to advising bottlenecks and delayed access to critical in…
ConCise: Training-Free Conclusion-Chain State Compression for Cost-Efficient Multi-Step RAG Services
Multi-step retrieval-augmented generation (RAG) has been widely deployed as LLM-powered web services for complex question answering, where…
LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis
Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 week…
meta-pipe: An LLM-agent pipeline for end-to-end automated systematic review and meta-analysis
Objective: To describe the architecture and design rationale of meta-pipe, an open-source large language model (LLM)-agent pipeline that in…
CAMI: Cost-Aware Agent-Guided Multi-Indexing for Semantic Retrieval
RAG ingestion pipelines frequently augment search corpus index with semantic enrichment indices (e.g., synthetic queries or summaries gener…
Beyond the Reranker: Do RAG Retrieval Enhancements Help Once a Strong Reranker Is Present?
Retrieval-augmented generation (RAG) is routinely extended with methods meant to improve retrieval: query expansion, hierarchical and cross…
Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models
Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.g., a dead whale wa…
Conversational Query Engine for Mixed-Modality Heterogeneous Enterprise Data Sources
Enterprise business intelligence queries span structured warehouses and unstructured document repositories -- modalities with fundamentally…
Model Merging to Evolution: Parameter Space Exploration for Expert Models
Model merging integrates the capabilities of multiple expert models to create strong models for multiple tasks without additional training,…
When Does Overlap Help? OSU-Mem and a Cell-Conditional Analysis of Trajectory Memory for LLM Agents
Long-horizon large language model (LLM) agents accumulate interaction trajectories that quickly exceed any practical prompt budget, and exi…
Memory-Augmented LSTM Autoencoder for Unsupervised Activity Recognition with IMU Sensor Fusion
HAR using Inertial Measurement Unit (IMU) sensors is vital for healthcare monitoring and rehabilitation. Despite deep learning advancements…
LEDGER: Scaling Agentic Document Editing with Dependency-aware Graph Retrieval
We introduce LEDGER to tackle the novel context engineering challenge of agentic document editing, where localized edits to long, structure…
Distilling a Modular Reservoir Through a Genomic Bottleneck
The intricate structures of biological neural networks largely emerge during development, guided by a comparatively compressed blueprint en…
Evolutional Math: Cross-Validated Island-Model Genetic Programming for Interpretable Symbolic Regression on Small, Wide Datasets
Symbolic regression via genetic programming routinely fails on small, wide datasets - a regime common in clinical-trial monitoring, biostat…
A Query-Driven Communication-Efficient Digital Twins Design for Autonomous Driving
Digital twins (DTs) have become a potential technology to perform risk-free simulation of physical entities for deterministic and high-reli…
RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis
Recent advances in robot world models enable synthetic video generation for embodied prediction and planning. However, evaluating these vid…
Data Provenance for Image Auto-Regressive Generation
Image autoregressive models (IARs) have recently demonstrated remarkable capabilities in visual content generation, achieving photorealisti…
Schema-First Retrieval: Embedding Catalogs for Natural Language Analytics
Enterprise text-to-SQL systems often fail before SQL is generated: the model receives the wrong schema context. Modern warehouses contain t…
Automated Quality Assessment of Geospatial Vector Data: A GeoAI Approach using Spatial Representation Learning
Geospatial vector data quality is a foundational research topic in GIS, yet classic rule-based quality assessment algorithms often struggle…
Few-class Fidelity: Evaluating Explanations of Real-conditions CNN classifiers with Optimized Perturbations
The wide use of Convolutional Neural Networks (CNN) in numerous domains and real-world classification applications is justified by their hi…
RADIANT-PET: Reasoning-Augmented PET/CT Lesion Segmentation with Large Language Models and Reinforcement Learning
Accurate lesion segmentation in PET/CT is critical for oncology, yet remains challenging because physiologic tracer uptake and artifacts ca…
CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation
Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-languag…
Reinforcement Learning for Software Vulnerability Analysis: A Systematic Review with Emphasis on C/C++ Source Code and Static Analysis
Vulnerability detection in C/C++ software remains a major security challenge due to code complexity, manual memory management, and the limi…
Financing Artificial Intelligence Infrastructure: Mapping AI Infrastructure Investment and Compute Governance Across Africa
Artificial intelligence depends on large-scale compute resources and their supporting infrastructure. However, AI governance debates treat…
Evidence-Driven LLM Agent for C-to-Synthesizable-C Conversion and Verification
Software-compilable C programs routinely fail to complete the four-stage pipeline of a high-level synthesis (HLS) toolchain -- compilation,…
RSGPNet: Geometric Prompting for Remote Sensing Open-Vocabulary Semantic Segmentation
Open-vocabulary semantic segmentation (OVSS) enables text-guided segmentation of unseen objects, breaking fixed-class limitations to achiev…
On the Necessity of a Liquid Substrate for Mesh Intelligence
A mesh of sovereign agents has no center: no shared clock, no shared model, and no coordinator to gather data or retrain. Its competence re…
AEGIS: A Semantic GAN and Evidential Learning Frameworkfor Robust Adversarial Detection in Vision Sensors
Deep neural networks (DNNs) have shown outstanding performance in visual recognition tasks within vision sensor networks; however, they are…
MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentatio
Limited data availability, class imbalance, and domain variability remain major barriers to reliable medical image classification. Conventi…
JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators
Text-to-image (T2I) diffusion models typically require substantial computational resources and cloud infrastructure, posing significant cha…
Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels. The na…
Building to the Test: Coding Agents Deliver What You Check, Not What You Requested
Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-vali…
Spectral Perturbation of the Empirical Fisher Information Matrix under Weight Quantization
We study the spectral perturbation of the empirical Fisher Information Matrix (FIM) of a parametric statistical model under two structured…
SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents
Long-horizon software engineering agents often need to manage lengthy and noisy interaction histories under limited context budgets. Existi…
Dockerless: Environment-Free Program Verifier for Coding Agents
Program verifiers play a central role in training coding agents, including selecting trajectories for supervised fine-tuning (SFT) and prov…
When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs
Recursive self-training can degrade neural generative models when generated data is reused without fresh human data or external quality con…
PLAA: Packet-level Adversarial Attacks in Network Traffic Detection
Deep neural networks (DNNs) are widely applied in Network-based Intrusion Detection System (NIDS) due to their high accuracy. However, DNNs…
Learning to Distributedly Estimate under Partially Known Dynamics: A Covariance-Agnostic Neural Kalman Consensus Filter
Online latent state estimation constitutes a fundamental challenge within the artificial intelligence field, serving as a foundational tool…
S-GAI: Spectral Geometry-Aware Initialization for Sigmoidal MLPs -- From Dataset Geometry to Network Weights
Classical universal approximation theorems establish the expressive power of sigmoidal multilayer perceptrons, but they do not prescribe ho…
LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features
Early detection of dementia enables timely intervention, and reflecting cognitive impairment, spontaneous speech offers a non-invasive scre…
Domain-Informed Multi-View Self-Distillation for Astronomical Light-Curve Representation Learning with JEPA
Light curves describe temporal variations in the brightness of celestial objects. Learning robust representations of light curves is essent…
SemFlowRAG: Directed Semantic Flow from Abstraction to Evidence for Complex Reasoning
Retrieval-Augmented Generation (RAG) enhanced by Knowledge Graphs has shown promise in complex multi-hop reasoning tasks. However, existing…
LLM agents security duality: a comprehensive survey of self-security and empowered cybersecurity
Large language model (LLM) agents are rapidly being integrated into real-world systems. Their autonomy and tool-use capabilities generate s…
Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Fields in Passive Object-State World Models
World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and u…
Is Lying an Emergent Behaviour in LLMs? Evidence from Gaslighting AI agents in a Sustainability Game
LLMs agents are increasingly used in multi-agent settings, yet their behaviour in sustainability games remains largely unexplored. This wor…
Counterfactual Residual Data Augmentation for Regression
Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observati…
SVC-Probe: A Framework for Evaluating Perturbation Generalization in Spatial Foundation-Model Embeddings
This work examines perturbation generalization in spatial foundation-model embeddings derived from fluorescence microscopy images. Although…
An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations
Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use. This paper p…
Improvement of Robot's Simultaneous Localization and Mapping Using an Effective Transformation to Achieve Linear Model
Nowadays mobile robots have wide engineering applications. Simultaneous localization and mapping (SLAM) is an important task of these robot…
Decomposing Memorization Reduction in Privacy-Preserving Fine-Tuning of SLMs for CSIRTs
CSIRTs increasingly fine tune language models on vulnerability scan records, but these records expose internal network topology and create…
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents
As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a…
Generative AI Literacy Training Improves Intelligence Analysts' Discrimination of Real and AI-Generated Images
Across social and online platforms, people are increasingly exposed to AI-generated images. As a consequence, the task of distinguishing AI…
HDDPM: Heteroscedastic Denoising Diffusion Probabilistic Model for Quantitative Low-Count Brain PET Recovery
Positron emission tomography (PET) seeks to balance diagnostic quality with ra-diation dose. Low-count PET noise is non-Gaussian, non-stati…
A Gravitational Interpretation of Fine-Tuning Reversion
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment update…
The Speedup Paradox: Rethinking Inference Speed-Quality Trade-off in Embodied Tasks
Embodied foundation models have recently been widely used to improve robot generalization and task success rates. Previous works apply loss…
CMSL: Constructive Multi-Sequence Learning for Recommendation Systems
Sequence learning has emerged as the promising paradigm in recommendation systems, surpassing traditional Deep Learning Recommendation Mode…
MammoFlow: Multiview Mammogram Synthesis with Anatomically Consistent Flow Matching
Multiview mammography relies on paired craniocaudal (CC) and mediolateral oblique (MLO) views to provide complementary projections of a 3D…
KernelSight-LM: A Kernel-Level LLM Inference Simulator
As large language models (LLMs) move into production serving, practitioners must rapidly evaluate inference performance across diverse hard…
Digitizing Coaching Intelligence: An Agentic Framework for Holistic Athlete Profiling using VLM and RAG
Athlete assessment is a critical process for tracking physical progress and identifying elite talent. However, during mass recruitment driv…
Geometric Measurements of the Axiom of Choice in Neural Proof Embeddings
The axiom of choice has divided the foundations of mathematics for over a century, but the distinction between classical and constructive p…
Correct codes for the wrong reasons? validating LLMs as measurement instruments for theoretical constructs
When a large language model (LLM) codes a construct in text as a human annotator would, that agreement makes the LLM a reliable coder. Yet…
Animation2Code: Evaluating Temporal Visual Reasoning in Video-to-Code Generation
While recent vision-language models (VLMs) have achieved significant improvements on static visual-to-code tasks such as generating code fo…
Neuromorphic Energy-Aware Learning for Adaptive Deep Brain Stimulation
Neuromorphic and edge computing research has focused on reducing the inference cost of neural network controllers, yet in physical closed-l…
Database Context Compression for Text-to-SQL on Real-World Large Databases
Recent progress in Text-to-SQL has been driven by stronger language models and prompting strategies, yet performance on real enterprise ben…
Fast and Accurate Outlier-Aware LiDAR Super-Resolution for SLAM Applications
This work tackles the challenge of enhancing low-resolution LiDAR sensors for SLAM applications through a novel Deep Unrolling-based Super-…
What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs
Large language models (LLMs) are increasingly deployed in high-stakes domains, where free-text explanations such as chain-of-thought and po…
The Undecidability of Artificial General Intelligence (AGI) Alignment
This article establishes the foundational mathematical limits of Artificial General Intelligence (AGI) safety, proving that the core barrie…
Analysis of Parameter Settings for the Bat Algorithm Using Variance Evolution
Parameter settings in evolutionary algorithms and metaheuristics are important because such parameter values can influence the performance…
RIPA: Sensory-Vector Prompt Injection Attacks on LLM-Controlled ROS 2 Robots
We present RIPA, the first systematic multi-channel empirical study of prompt injection attacks delivered through the sensory pipeline of a…
FedLAS: Feature-Modulated Bidirectional Label Smoothing for Neural Network Calibration
Deep Neural Network (DNN) classifiers suffer from poor calibration when their softmax outputs (predictive confidence) deviate from the empi…
SemDynReg: Semantics-Guided Deformation Regularization for Dynamic 3D Gaussian Splatting
Deformable 3D Gaussian Splatting (3DGS) has emerged as an efficient approach for rendering dynamic scenes in a wide range of 3D application…
When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling
People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning systems answer a hard ques…
Closed-Form Steepest Descent Direction toward Flat Minima: Reducing Upper Bounds on the Loss Hessian Eigenspectrum in Neural Networks
The flatness hypothesis suggests that flatness of the loss landscape, as measured by the eigenvalues of the loss Hessian, correlates with b…
Why Trust Your Agent? Empirical Security Gains from TRiSM-Guided Agentic Workflows in Healthcare
Agent-based AI has enabled the automation of tasks by exposing application tools and resources to large language models (LLMs). However, to…
MACROCAST: A Vintage-Consistent Time Series Foundation Model for Real-Time Macroeconomic Forecasting
We introduce MACROCAST, a lightweight Time Series Foundation Model (TSFM) for real-time macroeconomic forecasting. Existing TSFMs suffer fr…
Constrained Tabular Diffusion for Finance
Generative models in finance face the dual challenge of producing realistic data while satisfying strict regulatory and economic objectives…
Predicting Metastatic Risk from Primary Tissue Architecture via Distance-Aware Spatial Modeling
Predicting the risk of distant metastasis from primary tumor tissue histology is a critical yet challenging task in computational pathology…
Capability Gates Are Not Authorization: Confused-Deputy Failures in LLM Agent Frameworks
Tool-using LLM agents increasingly read untrusted content while holding side-effecting tools such as payments, email, CRM, and infrastructu…
SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages
While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly u…
CCRC: A Change-Aware Captioning and Reasoning Chain for Image Change Captioning and Segmentation
Understanding and localizing subtle changes between paired images is critical for tasks such as surveillance and image editing. However, tr…
5ting at SemEval-2026 Task 8: Strong End-to-End Multi-Turn RAG via LLM-Based Reranking and Faithfulness Control
We introduce 5ting, our system for the SemEval2026 Task 8 (MTRAGEval), which evaluates multi-turn Retrieval Augmented Generation (RAG) syst…
Four Types of LLM Reliance and Their Predictors Among Undergraduate Writers: A Mixed-Methods Study at a Minority-Serving R1 University
Although most undergraduates now use large language models (LLMs), a form of generative artificial intelligence (GenAI) for academic writin…
X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving
Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capabi…
Majority Vote Silences Minority Values: Annotator Disagreement at the Hate/Offensive Boundary in HateXplain
Hate speech annotation pipelines routinely collapse annotator disagreement into majority vote labels before training. We show that this agg…
Brownian Bridge Diffusion-Based Joint Channel Estimation and Data Detection for Jamming-Resilient Receivers
In next-generation wireless networks, the growing density of devices and limited spectrum resources pose severe jamming challenges to fragi…
BREIT: A Framework for Brain Stroke Reconstruction using Multi-Frequency 3D EIT
Multi-Frequency Electrical Impedance Tomography (MF-EIT) is a non-invasive, low-cost modality that reconstructs electrical property distrib…
The registrar's function in a hybrid society. AI value chain,smart data and the concept of property
Artificial intelligence reaches the land registry not as another tool but as a value chain that turns data into intelligence and intelligen…
Human2Any: Human-to-Robot Transfer via Constraint-Aware Compositional Planning
Human videos are a scalable source of supervision for robot manipulation, as they are abundant and naturally capture rich object interactio…
Categorizing Mathematical Concepts with LLM Voting Ensembles in Mathswitch
Mathswitch is an open-source project that imports mathematical concept records from sources such as Wikidata, Wikipedia, MathWorld, Encyclo…
Exit-and-Join Dynamics and Equilibrium in Continuum Cooperative Games
This paper develops a continuum theory of exit-and-join coalition dynamics in nonatomic cooperative games. We extend the Aumann-Shapley val…
HARD-KV: Head-Adaptive Regularization for Decoding-time KV Compression
Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.g., Top-$p$ nucleus sampling) offer superi…
Fisher-Routed Mixture of Experts for Federated Class-Incremental Learning
Federated Learning (FL) emerged as a promising distributed machine learning paradigm. However, extending FL to the class incremental learni…
LAMP: Lean-based Agentic framework with MCP and Proof Repair
Large language models are increasingly capable of mathematical reasoning, but the proofs they generate are often unreliable and hard to ver…
The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning
Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work h…
Perspectives on Latent Factor Indeterminacy and its Implications for Data Representation
The common factor analytic model is related to Helmholtz and Boltzmann machines, can be conceived as a linear autoencoder, or can be though…
Building AI-Ready Data Systems for Space Life Sciences, Aerospace Medicine, and Deep Space Exploration
While AI holds the potential to revolutionize space life sciences, realizing this promise is contingent upon the systematic restructuring o…
Defeat Devices in AI Systems
AI systems increasingly exhibit behavior that differs systematically between evaluation and deployment contexts. Alignment faking, sandbagg…
An Integrated Machine Learning and Hierarchical Variance Decomposition Pipeline for Student Performance Prediction and Metacognitive Calibration on Multi-Signal Telemetry
Predicting student performance and characterizing metacognitive calibration are essential for personalization in intelligent tutoring syste…
Exploring the Value of Diverse LLM Explanations in Introductory Programming
Large Language Models (LLMs) have shown the potential to generate code explanations that surpass those of peers in quality, offering promis…
A Task-Driven and Quality-Assured Agent Framework for SAR Data Generation
Synthetic aperture radar (SAR) data augmentation is important for improving the generalization of data-driven SAR interpretation models, ye…
Latent Bridges for Multi-Table Question Answering
We introduce GRAB, a constructor-encoder-bridge pipeline for table question answering. Our method lifts relational data into an heterogeneo…
Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation
Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agen…
DLR: Zero-Inference-Cost Latent Residuals for Low-Rank Pre-Training
Large language models have driven recent progress in language and multimodal AI, yet pre-training them at scale is prohibitively expensive.…
Machine-learnable Sets
In this study we present a formal definition of large discrete sets having, informally, three properties: their elements are easily recogni…
Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System
Poisoning attacks entail attackers intentionally tampering with training data. In this paper, we consider a dirty-label poisoning attack sc…
Modification-Considering Value Learning for Reward Hacking Mitigation in RL
Reinforcement learning agents can exploit misspecified reward signals to achieve high apparent returns while failing on the intended object…
RGLD: Randomized Global-Local Density Estimation for Tabular Anomaly Detection
Unsupervised tabular anomaly detection requires methods that are accurate, robust across heterogeneous datasets, and computationally effici…
Evidence-Based Text-Conditioned 3D CT Synthesis for Ovarian Cancer
Ovarian cancer is frequently diagnosed at an advanced stage, making preoperative contrast-enhanced computed tomography (CT) central to stag…
Compositional Dynamics in Learning and Mechanics
We give a single compositional setting in which gradient-based learning and Hamiltonian-style mechanics appear as functorial semantics. The…
Fine-Tuning General-Purpose Large Language Models for Agricultural Applications:A Reproducible Framework and Evaluation Protocol Based on Qwen3-8B
General-purpose large language models (LLMs) have demonstrated strong abilities in opendomain question answering, information extraction, a…
Arbitrary Reduction of Validation Error for AI Decision Tests using Homomorphic AI and Repetition Codes
This paper presents new results and breakthrough obtained with the HbHAI techniques (Hash-based Homomorphic Artificial Intelligence) propos…
Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation
Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards.…
BERTomelo: Your Portuguese Encoder Best Friend
Encoders have become the state of the art for multiple NLP tasks, especially those requiring deep contextual understanding. While multiling…
Semantic-Aware, Physics-Informed, Geometry-Grounded Weather Video Synthesis
Weather synthesis aims to add weather effects to input videos while preserving scene identity, structure, and motion. The key limitation of…
Efficient Spatio-Temporal Grounding with Multimodal Large Models via Second-Level Tracking and RL Verification
Spatio-temporal grounding in long videos requires precise temporal localization and robust object tracking conditioned on natural-language…
How to Leverage Synthetic Speech for LLM-Based ASR Systems?
In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic spe…
The strength of clinical evidence is recoverable from language model representations but not from their stated grades
Large language models (LLMs) increasingly summarize clinical evidence, where a claim's weight depends on how strongly it is supported. Yet…
Metric Aggregation Divergence: A Hidden Validity Threat in Agent-Based Policy Optimization and a Contractual Remedy
Metric aggregation divergence (MAD) is the silent inconsistency that arises when distinct pipeline stages in an agent-based model coupled w…
Flow Matching in Feature Space for Stochastic World Modeling
World modeling requires forecasting uncertain futures while preserving information useful for downstream perception. Existing visual world…
Fairness Attacks on Recommender Systems
The unfairness of recommender systems has become a topic of concern due to its significant social and ethical implications. Although existi…
A Comparative Study on Affective Cues in Text Embeddings Across Psychological Emotion Theories
Text encoders are known for their utility in natural language processing, as they are able to efficiently compress inputs into dense vector…
From Tool Connection to Execution Control: Benchmarking Security Invariants in MCP-Style Agent Runtimes
Model Context Protocol (MCP)-style ecosystems give language-model applications a practical connection layer for tools, resources, prompts,…
Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking
There are various benchmarks to evaluate bugfixing capabilities of Large Language Models. However, most widespread benchmarks do not fully…
AB-RAG: Adaptive Budgeted Retrieval-Augmented Generation for Reliable Question Answering
Retrieval-Augmented Generation (RAG) has become the standard way to ground large language models in external knowledge, yet most systems re…
Statistically Indistinguishable, Operationally Distinct: A Formal Barrier for Tabular Foundation Models
Tabular foundation models cannot reason about data produced by running systems without access to the rules that govern them. We make this s…
Priced Motion Through Optimal Faces: A Normal-Fan Geometry for Non-Stationary Adversarial MDPs
In a changing decision problem, standard dynamic-regret analyses have often equated the cost of non-stationarity to how far loss moves. How…
Unified Complex-valued Neural Network: A Magnitude-Phase Computational Model for Event-Driven Neuromorphic Learning
Artificial neural networks (ANN) provide accurate continuous-valued representation, whereas spiking neural networks (SNN) offer event-drive…
BTI-Net: Bidirectional Decoder-Level Task Interaction via Uncertainty-Aware Gating for Multi-Task Medical Image Analysis
Jointly learning to segment and classify medical images demands cross-task synergy, yet encoder-sharing architectures limit decoder reconst…
A Deep Multiscale Neural Network for Accurate Neurological Disorder Detection from MRI Scans and Real-Time Web Deployment
Neurological disorders involve diverse pathologies of the brain and nervous system, making early and accurate detection essential. While ma…
LLM Semantic Signaling Game and Mechanism Design: Systematic Blindness, Awareness Shaping, and Mindset Dynamics
Large language models (LLMs) increasingly mediate strategic interactions through natural language, making semantic control a critical eleme…
When Stopping Fails: Rethinking Minimal Risk Conditions through Human-Interactive Autonomous Driving for Safe Transportation Systems
Autonomous vehicles (AVs) are increasingly deployed in urban environments, yet their safety frameworks remain primarily designed around col…
Knowing in Advance When an Evolutionary Outer Loop Will Not Help: A Pre-Registered Cheap-Baseline Screening Rule
We introduce a pre-registered screening rule that decides, before any implementation, whether an evolutionary / population / lifecycle oute…
How Anthropomorphic Language Impacts Public Perceptions of AI
Public discourse about artificial intelligence (AI) often uses anthropomorphic language: language that attributes human capabilities and ch…
CMTFormer: Marrying Transformer with Hierarchical Information Interaction for RGB-Event Object Detection
Event cameras capture sparse brightness changes with high temporal resolution and high dynamic range, compensating for the deficiencies of…
GPC: Large-Scale Generative Pretraining for Transferable Motor Control
Developing controllers capable of completing a wide range of tasks in a natural and life-like manner is a key challenge in enabling practic…
On the Nonlinearity of Learning Rate Scaling for LLM Training
Learning-rate transfer can reduce the cost of training large language models: instead of sweeping learning rates at target scale, practitio…
Invariant Reasoning Directions in Latent Trajectories of Language Models
Latent reasoning models perform multi-step inference directly in hidden-state space, yet the structure of these latent reasoning trajectori…
Projected Exploitability Descent for Nash Equilibrium Computation in Multiplayer Imperfect-Information Games
Many important games have more than two players and imperfect information. Existing approaches for computing Nash equilibrium, the central…
Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies
While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how t…
Anomaly Factory 3D: A Modular Framework for Diverse Pseudo-Anomaly Synthesis in Unsupervised 3D Anomaly Detection
Detecting and localizing defects in 3D point clouds is challenging because abnormal samples are scarce and diverse, while training is often…
A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis
LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimo…
Behavior Uncloning: Distilling Mode Redirection into Policy Weights without Inference-Time Steering
Behavior-cloned policies often learn multiple behavior modes from demonstration datasets, including modes that are unsafe or otherwise unde…
AnyBody: Free-Form Whole-Body Humanoid Control from Arbitrary Keypoint Guidance
We present AnyBody, a unified whole-body humanoid controller driven by an arbitrary subset of body keypoints chosen at deploy time. Prior p…
MoPe: Motion Permanence for Robust Monocular Gaussian Mapping in Dynamic Environments
Robust robot autonomy depends on scene representations that remain stable enough to support localization, navigation, and downstream decisi…
Confidence-feedback-weighted graph matching network: online-offline laser-induced damage site matching under complex interference
Online inspection images of final optics in high-power laser facilities contain pseudo-damage sites that closely resemble true damage sites…
A Hybrid Framework for Song Lyric Annotation Based on Human-LLM Alignment
Emotion recognition of song lyrics is a challenging task since lyrics may not necessarily align with the overall emotion of a song. As a re…
Manufactured Confidence: How Memory Consolidation Turns Hearsay into Confident Facts
LLM agents carry conclusions across steps and sessions in compressed memory, and memory products (e.g., mem0, LangMem) rewrite conversation…
Deterministic Decisions for High-Stakes AI. A Zero-Egress Pipeline with the Deployability of RAG and the Accuracy of Machine Learning
We identify intervention bias as a previously unquantified failure mode of zero-shot large-language-model (LLM) educational advisory agents…
Covering the Unseen: Information Demand Coverage Optimization for Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) typically treats context selection as ranking chunks against a single query embedding. This assumption…
AMR: Adaptive Modality Routing for Multimodal Polyglot Speaker Identification
Multimodal speaker identification systems face two key challenges in real-world deployment: missing modalities and language mismatch betwee…
Adaptive Financial Transformer with Regime-Gated Attention for Stock Return Prediction
Adaptive Financial Transformer (AFT) is proposed for stock return prediction under non-stationary financial markets. The model incorporates…
Fast Enough to Act: Spatio-Temporal Visual Token Merging for Low-Latency Robotic VLMs and VLAs
Vision-language models and vision-language action models endow the robot with unprecedented capabilities. However, the input of video and h…
Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking
Vision-language tracking guided by natural language specifications leverages high-level semantic cues of target objects to substantially bo…
Solver-Verified Formulation Generation and Selection for Multi-Warehouse Inventory Allocation Using Large Language Models
Balance-oriented multi-warehouse inventory allocation is a recurring decision problem in large-scale e-commerce supply chains, in which a f…
DR-GS: Physically-Based Deformable and Relightable 2D Gaussians
Gaussian splatting (GS) has garnered significant attention in VR/AR and digital content creation due to its explicit parameterization and e…
Learning to Adaptively Allocate Gaussians for Arbitrary-Scale Image Super-Resolution
In computer graphics, visual content is continuously warped, zoomed and resampled. This occurs when engines upscale frames, users zoom into…
Self-Organized Conformal Prediction: Reducing Regional Coverage Gaps with Unsupervised Group Discovery
Conformal prediction guarantees marginal coverage, but pooled calibration averages over heterogeneous regions and can mask regional underco…
LC-ICL: Label-Guided Contrastive In-Context Learning for Robust Information Extraction
There has been increasing interest in exploring the capabilities of advanced large language models (LLMs) in the field of information extra…
Can Machines Really See Objects in Images? A Study Based on Syntactic Distance and Visual Self-Referential Instances
Can a vision model truly see an object, or does it only fit surface-level visual cues? Following Wittgenstein's view that the limits of lan…
LLMography: Transforming Human-AI Conversations into Traceability, Oversight, and Auditability Indicators
The growing use of Large Language Models (LLMs) in education, software engineering, academic writing, and technical documentation raises a…
Closing the Activation-Cone Blind Spot: Response-Time Probing and Unified Defense
Inference-time safety methods for large language models have proliferated, yet no systematic comparison exists. We evaluate five defense pa…
Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction
Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achie…
Resonant Brane Splatting for Arbitrary-Scale Super-Resolution
Arbitrary-Scale Super-Resolution (ASR) reconstructs images at continuous magnification factors. Recent methods accelerate inference by repl…
Interpretable Inverse Design of Metal-Organic Frameworks with Large Language Model Agents
Inverse design of metal-organic frameworks (MOFs) requires searching a combinatorially vast space where property labels are expensive and m…
Rank-Aware Hyperbolic Alignment for Vision-Language Dataset Distillation
Vision-language dataset distillation (VLDD) compresses a large image-text paired dataset into a small set of synthetic pairs that can effic…
A Posteriori Error Analysis for Decoupled Neural Approximations of Fully Coupled FBSDEs with Control Mismatch
This paper develops an a posteriori error analysis framework for decoupled neural approximations of fully coupled forward--backward stochas…
CRAFT: Counterfactual Credit Assignment from Free Sibling Rollouts for Self-Distilled Agentic Reinforcement Learning
Self-distilled agentic reinforcement learning augments trajectory-level reward with a token-level distillation loss, using as its teacher t…
To Reason or to Fabricate: Reasoning Without Shortcuts via Hint-Anchored Pairwise Aggregation
While reinforcement learning (RL) significantly enhances LLM reasoning, its efficacy is severely undermined by Pre-RL data overlap, where R…
Reported Confidence in LLMs Tracks Commitment More Than Correctness
Confidence is an estimate of the probability that a chosen answer is correct. Verbal confidence reports are widely used as uncertainty meas…
The Verbose Context Problem in Medical Records
The verbose context problem occurs when structured concepts have token-inefficient textual representations. This bottleneck is acute in pop…
SAKE: Software Architectural Knowledge Evaluation Benchmark for Large Language Models
Large Language Models (LLMs) are increasingly used as assistants across the software development lifecycle, yet their ability to reason abo…
MotionAtlas: Detailed Region Captioning for Motion-Centric Videos
We propose MotionAtlas, a system for detailed captioning of motion-centric videos, comprising (1) a dedicated human-annotated benchmark, (2…
SemJoin: Semantic Join Optimization
Integrating unstructured data into relational database systems is increasingly important as demand grows for natural language querying and…
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing sk…
Em-ergence of the em-dash: a population-level rise in em-dash frequency in medRxiv preprints at the dawn of the large-language-model era
Large language models (LLMs) can leave subtle stylistic traces in assisted text; one of the most cited is the em-dash (Unicode U+2014). Yet…
Proteus: Automated Adversarial Robustness Testing for Audio Deepfake Detectors
We present Proteus, a framework developed at Resemble AI for automated robustness testing of our audio deepfake detection system. Given a d…
VISTA-DZ: Visual Semantic Trajectory Adaptation for Personalized Dilemma Zone Prediction
Driver decision making in the dilemma zone at signalized intersections is safety critical, as vehicles approaching a yellow signal must dec…
Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM
Large language models (LLMs) excel at complex tasks like question answering and summarization, thanks to their ability to handle long-conte…
TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation
Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cos…
ReMAP-PET: Beyond Visual Understanding -- Learning Region-Guided Metabolic Alignment Semantics from Brain PET
Positron Emission Tomography (PET) reveals brain metabolism and is clinically central to neurodegenerative disease assessment, yet existing…
ScAle: Attention Head Scaling as a Minimal Adapter for Spatial Reasoning in Vision Language Models
Spatial reasoning remains a persistent challenge for many vision language models (VLMs), and improving it typically requires fine-tuning wi…
The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis
Modern LLM deployments routinely compress models and raise sampling temperature to reduce cost, latency, or repetition, yet safety evaluati…
Bilevel Optimization for Neural Architecture Search
Bilevel optimization has become an influential and widely adopted framework for addressing hierarchical optimization problems in machine le…
SonoCLIP: Mask-Guided Region-Aware Vision-Language Pretraining for Fetal Ultrasound Analysis
Vision-language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging…
How AI settled the complexity of the oldest SGD algorithm
In 1937, Stefan Kaczmarz proposed a simple algorithm for solving systems of linear equations. This algorithm turned out to be the earliest…
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
A faithful 3D world representation should account for layered geometry, where a single camera ray may contain multiple visible and geometri…
Langshaw: Declarative Interaction Protocols Based on Sayso and Conflict
Current languages for specifying multiagent protocols either over-constrain protocol enactments or complicate capturing their meanings. We…
Mechanistically Eliciting Latent Behaviors in Language Models
We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could…
Does Role Specialization Matter for Explanation Faithfulness in Mixture-of-Experts?
Mixture-of-Experts (MoE) architectures have recently been extended with role-based mechanisms for interpretability. This is typically done…
Do We Still Need Fine Tuning? Turkish Sentiment Analysis in the Era of Large Language Model
This study examines whether supervised fine-tuning remains necessary for Turkish sentiment analysis in the era of large language models. We…
Two-Stage Prompt Optimization for Few-Shot Relation Extraction: From Reasoning-Guided Search to Gradient-Guided Refinement
Automatic prompt optimization is still underexplored for episodic few-shot relation extraction with smaller language models. We propose a t…
Fast Wireless Foundation Models with Early-Exits
While wireless foundation models (FMs) are demonstrating strong potential to enable AI-Native 6G networks, their high computational cost re…
Fuzzing Large Language Models to Elicit Hidden Behaviours
Sleeper agents are the canonical model organism of deception: models trained to behave normally but to emit an unsafe behaviour on a specif…
Hybrid Retriever Evolution for Multimodal Document Reasoning Agents
Different retrievers, including lexical, semantic, and multimodal approaches, provide highly complementary strengths for multimodal documen…
Unlocking the Visual Record of Materials Science: A Large-Scale Multimodal Dataset from Scientific Literature
The materials science literature encodes decades of experimental knowledge in figures, yet this visual record remains locked away and inacc…
A Machine-Verified Proof of a Quantum-Optimization Conjecture
We report a machine-verified resolution of a problem open for over a decade in quantum optimization: the Farhi, Goldstone and Gutmann (FGG)…
Early Warning Signals for OpenVLA Failure under Visual Distribution Shift
Vision Language Action models combine perception, language grounding, and control in a single policy, but their failures are hard to diagno…
ARMOR: Adaptive Retriever Optimization for Low-Resource Telecom Question Answering
Telecom question answering (QA) is a challenging setting for retrieval-augmented generation (RAG): evidence is fragmented across standards,…
SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution
Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today'…
Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop
Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of…
Diagnosing and Mitigating Context Rot in Long-horizon Search
Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon tasks. The concern that inc…
Redefining Maritime Anomaly Detection via Equation-Grounded Synthetic Anomalies
Maritime anomaly detection is essential for ensuring maritime safety, security, and efficient traffic management at sea, with Automatic Ide…
From Trait to Behavior: A Cognitive-Affective Personality System (CAPS) Perspective on Multi-Homing Intention in AIGC Platforms
With the rapid development of Artificial Intelligence Generated Content (AIGC) platforms, users increasingly show cross-platform usage inte…
PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF
Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alter…
TopoAgent: An Agentic Framework for Automated Topology Learning in Medical Imaging
Topological data analysis (TDA), particularly persistent homology (PH), captures geometric structural properties in medical images (e.g., c…
Towards Generalizable and Evidential Nuclear Magnetic Resonance-Based Molecular Structure Elucidation via Large Language Model Agent
Nuclear Magnetic Resonance (NMR) spectroscopy is the gold standard for molecular structure elucidation, yet interpreting complex spectra fo…
Mandol: An Agglomerative Agent Memory System for Long-Term Conversations
Long-term conversational agents need to remember and query cross-session, multi-typed information with complex correlations. Existing agent…
FalconTrack: Photorealistic Auto-Labeled Perception and Physics-Aware Vision-Based Aerial Tracking
Vision-based aerial tracking is critical in GPS-denied environments. Reliable perception for tracking depends on large-scale labeled data,…
HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data
Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive…
What Drives the Inlier-Memorization Effect? A Theory of Outlier Detection via Early Training Dynamics
Outlier detection (OD) aims to identify anomalous instances by learning the underlying structure of normal data (inliers), and is particula…
Multi-Level Distributional Entropy for Explainable Network Intrusion Detection
Machine learning network intrusion detection systems (IDS) rely on aggregate flow statistics that discard distributional structure, while e…
Accelerating Q-learning through Efficient Value-Sharing across Actions
Action-values are foundational to many control algorithms such as Q-learning. Therefore learning action-values efficiently is central to re…
Making Multimodal LLMs Reliable Chart Data Extractors: A Benchmark and Training Framework
Chart data extraction, which reverse-engineers data tables from chart images, is essential for reproducibility, analysis, retrieval, and re…
How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation
Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale. The most accurate detection methods depen…
Dual-Flow Reinforcement Learning with State-Aware Exploration
In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return dis…
Experience Graphs: The Data Foundation for Self-Improving Agents
The database community has repeatedly advanced the state of the art by recognizing that new workloads demand new system architectures. We a…
Neural Procedural Memory: Empowering LLM Agents with Implicit Activation Steering
While Large Language Models (LLMs) excel as static solvers, transforming them into autonomous agents remains challenging. This transition r…
MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers
The quadratic computational cost of traditional attention mechanisms poses a major bottleneck to the scalability and practical deployment o…
Exploring Motivations for Algorithm Mention in the Domain of Natural Language Processing: A Deep Learning Approach
With the rise of data-intensive science, algorithms have become central to scientific research. In academic papers, algorithms are mentione…
SUMO: Segment and Track Any Motion with Nonlinear State Space Models
Visual Object Tracking (VOT) and Moving Object Segmentation (MOS) are two fundamental tasks in computer vision that involve both spatial an…
RoAd-RL: A Unified Library and Benchmark for Robust Adversarial Reinforcement Learning
Deep Reinforcement Learning (DRL) has achieved significant success in robotics and autonomous systems, yet remains vulnerable to adversaria…
ARKD: Adaptive Reinforcement Learning-Guided Bidirectional KL Divergence Distillation for Text Generation
Knowledge distillation (KD) is a key technique for compressing Large Language Models (LLMs), yet methods relying on a single KL objective o…
Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency
Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot disting…
LWDrive: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving
Vision-Language Models (VLMs) provide powerful semantic understanding and commonsense reasoning for End-to-End Autonomous Driving (E2E-AD)…
Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models
Reinforcement learning (RL) has become indispensable for pushing Vision-Language-Action Models (VLAs) beyond static imitation learning. How…
SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics
As agentic AI systems tackle more complex mathematical tasks, they increasingly rely on information retrieval (IR) to search problem databa…
Child-Centric Voice Anonymization in Single and Multi-Speaker Speech via Domain-Adapted SSL Models
Voice anonymization aims to protect speaker identity while preserving linguistic content and speech usability. However, most anonymization…
Critical Interval MSE: Toward Reliable Offline Validation for Robot Manipulation Policies
Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challeng…
LLM-based Multimodal Personality Recognition via Facial Action Unit-Text Semantic Fusion
Personality recognition in asynchronous video interviews (AVIs) has become increasingly important due to their widespread adoption in moder…
Semi-Supervised Sound Event Detection with Conditional Mixup and Embedding-Level Contrastive Loss
Sound event detection (SED) is a core module for acoustic environmental analysis, yet its performance is often limited by scarce labeled da…
CW-B: Class Weighted Boosting Framework for Imbalance Resilient Multi Class Cardiac Phenotyping
Cardiac discharge phenotyping informs post-discharge treatment and follow-up, but real-world records are often incomplete and class-imbalan…
Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation
Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from tra…
EVAF: A Test-Retest Protocol for Selective Parametric Consolidation
Long-running language agents need mechanisms for deciding which experiences should persist after the working context is gone. Retrieval sys…
Latent-CURE for Breast Cancer Diagnosis
Multimodal Large Models have significantly advanced automated breast ultrasound diagnosis. However, most existing frameworks utilize opaque…
Data-Efficient Multimodal Alignment for Histopathology-based Molecular Prediction
H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-se…
Exploiting Local Flatness for Efficient Out-of-Distribution Detection
Detecting out-of-distribution (OOD) data is crucial for reliable machine learning deployment. Among detection strategies, post-hoc methods…
SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows
Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet…
SWE-Together: Evaluating Coding Agents in Interactive User Sessions
Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real…
DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation
Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over multiple turns, but this a…
RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines
Quantum computing provides a powerful paradigm for representing and transforming high-dimensional information through superposition, entang…
Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping
Looped Transformers, which repeatedly apply a shared transformer block, are an architecturally natural fit for variable-length algorithmic…
T3R: Deeper Test-Time Adaptation for Graph Neural Networks via Gradient Rotation
Graph Neural Networks (GNNs) deployed in real-world systems typically have fixed weights, often leading to degraded performance under distr…
IBRSteG: Learning a Generalizable Steganography Framework for 3D Gaussian Splatting
Recent advances in deep learning have notably improved steganographic message hiding. However, designing a generalizable steganographic app…
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs
Audiovisual arts encompass diverse creative disciplines, including cinema, visual arts, stage performance, and game design, where artistic…
Little Brains, Big Feats: Exploring Compact Language Models
While large language models have been dominating the research landscape recently, small language models remain highly relevant across vario…
Neural Subspace Reallocation: Continual Learning as Retrieval-Based Subspace Memory Management
We introduce Neural Subspace Reallocation (NSR), which reframes continual learning as memory management over parameter subspaces. Instead o…
Online Data Selection for Instruction Tuning via Gaussian Processes
With Large Language Model (LLM) pre-training and fine-tuning shifting its focus from data volume to data quality, quality data selection ha…
Automating the Design of Embodied AgentArchitectures
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity expo…
SA-VLA: State-aware tokenizer for improving Vision-Language-Action Models' performance
Discrete action tokenization provides a compact interface for autoregressive VLA policies, but accurately recovering continuous robot actio…
Gravitational Duals from Equations of State II: Large Hierarchies and False Vacua
We investigate the reconstruction of holographic duals for strongly coupled quantum field theories in regimes characterized by large hierar…
Hyper-Network Neural Functional Maps for Unsupervised Robust 3D Shape Matching
Functional maps are the cornerstone of recent non-rigid 3D shape matching methods due to their efficiency and performance. However, existin…
Query-Aware Spreading Activation for Multi-Hop Retrieval over Knowledge Graphs
Retrieval-augmented generation built on knowledge graphs (Graph RAG) outperforms flat passage retrieval on multi-hop question answering by…
Estimating Grammatical Gender Directions in Contextual Embeddings under Controlled and Natural Contexts
Contextual language models conflate grammatical gender and social semantic bias in gendered languages such as Spanish. Existing gender debi…
Physically-Constrained Harmonic Separation for Robust Heart and Respiratory Rate Estimation from Wrist Photoplethysmography
Wrist-worn photoplethysmography (PPG) enables continuous monitoring of cardiopulmonary physiology, but reliable heart rate (HR) and respira…
Federated Learning with Energy-Based Structured Probabilistic Inference
Federated learning typically aggregates client updates using fixed or heuristic weighting rules, which can be suboptimal when clients have…
Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark
Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical dat…
Few-Shot Domain Incremental Learning via Continual Vision-Language Consolidation
Existing domain-incremental learning (DIL) strategies call for massive amounts of data to adapt to new domains and suffer from the overfitt…
Forewarned is Forearmed: When Non-Sequential Embedding Turns Into an Anomaly Detector
This paper offers an in-depth analysis of non-sequential multimodal sentence-level embeddings, with a particular focus on the SONAR model.…
A Multi Center Breast FNAC Whole-Slide Cytology Dataset for AI-Assisted Patch-Wise Classification Using C1 to C5 Reporting Categories
We present a multi center breast fine needle aspiration cytology (FNAC) dataset designed for patch wise classification using C1 to C5 repor…
Efficient RGB-T Object Detection via Sparse Cross-Modality Fusion
RGB-T detectors leverage the complementary strengths of visible and thermal infrared modalities, achieving robust performance under challen…
Curvature-Guided Sheaf Diffusion for Unsupervised Community Detection on Heterophilic Graphs
Detecting communities in heterophilic graphs -- where connected nodes often belong to different classes -- is hard for unsupervised methods…
KnowsTFM: Knowledge-Informed Fine-Tuning of Small Tabular Foundation Models
Tabular foundation models have advanced deep learning for tabular data by delivering strong default performance across many small and mediu…
Defending Against Harmful Supervision Hidden in Benign Samples
Existing defenses are effective when harmful content is explicitly mixed into downstream fine-tuning data, but crafted samples can instead…
Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation
Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it fr…
Research Entity Extraction and Topic Detection from UKRI Grant Proposals
This paper presents preliminary findings from a UKRI-funded Metascience project comparing three LLM-based approaches, GPT-4o, Mistral, and…
Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents
Always-on agents are systems whose future behavior depends on durable state accumulated across earlier interactions. We treat them as persi…
MCP Server Architecture Patterns for LLM-Integrated Applications
The Model Context Protocol (MCP), introduced by Anthropic in November 2024, defines a standardized interface for connecting large language…
Early Cue Precision Shapes Visual Shortcut Learning in Controlled Cue-Manipulation Benchmarks
Visual classifiers can achieve high matched-distribution accuracy while relying on low-level cues that fail under conflict or suppression.…
DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training
Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in comple…
FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images
We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from on…
Residual-Guided Expert Specialization for Incomplete Multimodal Learning
As real-world prediction systems often face missing modalities at inference, incomplete multimodal learning (IML) remains a practical chall…
ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control
While current Behavior Foundation Models (BFMs) provide robust control priors for humanoids, they only execute pre-defined reference motion…
Set-Inclusive Uncertainty Modeling for Robust Brain Tumor Segmentation
Multimodal MRI is essential for accurate brain tumor segmentation. However, acquiring all modalities at inference is often challenging in p…
A Stochastic--Geometric Theory of Scaling Laws in Grokking
Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only…
Model Predictive Current Control with Harmonic Correction for Single-Phase AC-DC EV Charging
The increasing integration of Electric Vehicles (EVs) has imposed a growing harmonic challenge on the power grid. For AC/DC Power Factor Co…
Beyond IID: How General Are Tabular Foundation Models, Really?
Foundation models for predictive machine learning on tabular data have recently gained significant traction in academia and industry. Resea…
Can LLMs Rank? A Tale of Triads and Triage
From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered…
Beyond Point Estimates for Glaucoma Visual Field Forecasting with Diffusion Models
Forecasting visual fields (VFs) is critical for personalized monitoring and treatment planning in glaucoma. This is inherently uncertain du…
Transformer Architectures as Complete Bayes Processes: A Formal Proof in the Measure-Theoretic Kernel Framework
We present a complete formal proof that transformer architectures, when their internal update mechanisms satisfy a Bayes joint-distribution…
Translating Natural Language to Strategic Temporal Specifications via LLMs
A rigorous formalization of system requirements is a fundamental prerequisite for the verification of Multi-Agent Systems (MAS). However, w…
Collective cooperation without individual fidelity in LLM agents
Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior c…
Field Order Should Not Matter: Permutation-Invariant Embedding Model Fine-Tuning for Structured Metadata Retrieval
We study retrieval over catalogs of structured metadata, where each record is a small schema whose fields answer different kinds of query.…
COHORT: Collaborative Orchestration for Hardening via Offensive Replay on Emulated Topologies
Mitigating an observed adversary in an enterprise network typically takes weeks of expert work: an analyst derives a mitigation tailored to…
Situation Perception: A Necessary Primitive to Artificial Superintelligence
Current large language models are extraordinary statistical engines. They compress vast amounts of text into useful patterns and can explai…
SIMAX: A Scalable and Interpretable Framework for Multi-Fidelity and Annotated Clinician-Patient Dialogue Simulation
Background. The widespread deployment of ambient digital scribes is driving large-scale capture of clinician-patient dialogues. Human codin…
McMg: A Learned Phase-Space Multi-channel Multigrid Preconditioner for Helmholtz Equation
Solving heterogeneous Helmholtz equations at high wavenumbers remains challenging because the discretized operator is indefinite, pollution…
On the Faithfulness of Post-Hoc Concept Bottleneck Models
Human decision-making interprets the world through high-level concepts, such as recognizing a bird by its belly color. To bridge the gap be…
Informational Frustration in Neural Manifolds: Shannon Bottlenecks and the Limits of Learnability
Why overparameterised deep networks generalise so remarkably well remains one of the most stubborn open questions in machine learning theor…
Learning from Mistakes: Rollout-Retrieval Lifelong Policy Learning for Autonomous Driving
Autonomous driving policies should be able to improve continually as deployment exposes them to increasingly diverse and long-tail traffic…
TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech
With the proliferation of speech AI agents, understanding emotional entrainment in conversational interaction has become increasingly impor…
To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks
AI code completion tools, such as Github Copilot, provide students with code suggestions to help them write programs. However, recent quali…
TraceLab: Characterizing Coding Agent Workloads for LLM Serving
Coding agents are rapidly becoming a major application of agentic LLMs, but serving them efficiently remains challenging. Progress on this…
A Multi-task Mixture of Experts Framework for Malware Classification, Packing Detection, and Family Attribution
Malware classification remains a challenging problem due to its inherent heterogeneity, the presence of packed binaries, and the diverse di…
Beyond 2D Matching: A Unified Single-Stage Framework for Geometry-Aware Cross-View Object Geo-Localization
Cross-view object geo-localization (CVOGL) aims to locate a target object from a query view (e.g., ground or drone) within a geo-tagged ref…
Words Speak Louder Than Code: Investigating Cognitive Heuristics in LLM-Based Code Vulnerability Detection
Researchers and practitioners increasingly apply Large Language Models (LLMs) for automated vulnerability detection. Recent work has shown…
MESA: Prioritizing Vulnerable Communication Channels for Securing Multi-Agent Systems
Multi-agent systems (MAS) are increasingly used to automate complex, distributed workflows. However, their inter-agent communication channe…
C$^{2}$R: Cross-sample Consistency Regularization Mitigates Feature Splitting and Absorption in Sparse Autoencoders
Sparse Autoencoders (SAEs) are widely used to interpret large language models by decomposing activations into sparse, human-understandable…
Optimization Dynamics Imprint Semantic Specificity in Contrastive Embedding Norms
Contrastive embedding models trained with scale-invariant losses are typically paired with distance metrics like cosine similarity, effecti…
Pessimism's Paradox: Conservative Offline Training Amplifies Reward Hacking During Online Adaptation in Reasoning Models
Conservative offline training is widely advocated as a safe foundation for subsequent online adaptation: if a policy stays close to well-su…
GROW$^2$: Grounding Which and Where for Robot Tool Use
Can the robot use a plate to cut a cake if no knife is available? Tool use greatly expands robot capabilities, but to use tools creatively…
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training
Full-length song generation must preserve coherence and musicality, render detailed vocal and accompaniment acoustics, and follow lyrics an…
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes
Perception-based humanoid loco-manipulation requires connecting egocentric observations and task instructions to whole-body motion. Learnin…
A binarized-domains arc-consistency algorithm for TCSPs: its computational analysis and its use as a filtering procedure in solution search algorithms
TCSPs (Temporal Constraint Satisfaction Problems) [Dechter et al. 1991] get rid of unary constraints by binarizing them after having added…
Modelling Human Values for Value-Aware Multi-Agent Systems
One of today's most pressing societal challenges is building AI systems whose behaviour, or the behaviour it enables within communities of…
Instance-Conditioned Adaptation for Large-scale Generalization of Neural Routing Solver
In modern intelligent transportation systems (ITS), particularly in freight transportation and logistics, real-time route planning is cruci…
CLMASP: Coupling Large Language Models with Answer Set Programming for Robotic Task Planning
Large Language Models (LLMs) possess extensive foundational knowledge and moderate reasoning abilities, making them suitable for general ta…
OptiMUS-0.3: Using Large Language Models to Model and Solve Optimization Problems at Scale
Optimization problems are pervasive in sectors from manufacturing and distribution to healthcare. However, most such problems are still sol…
MARS: A neurosymbolic approach for interpretable drug discovery
Background: Neurosymbolic (NeSy) artificial intelligence describes the combination of logic or rule-based techniques with neural networks.…
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes. Such models are…
GraphChase: A Platform and Benchmark for Urban Network Security Games
After the achievement of solving two-player zero-sum games, more AI researchers focus on solving multiplayer games. Urban Network Security…
Accelerating scientific discovery with Co-Scientist
Scientific discovery is driven by scientists generating novel hypotheses for complex problems that undergo rigorous experimental validation…
StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely e…
Reconsidering Overthinking: Penalizing Internal and External Redundancy in CoT Reasoning
Large reasoning models (LRMs) often exhibit overthinking, producing verbose Chain-of-Thought (CoT) traces that increase inference cost and…
Learning How to Use Tools, Not Just When: Pattern-Aware Tool-Integrated Reasoning
Tool-integrated reasoning (TIR) has become a key approach for improving large reasoning models (LRMs) on complex problems. Prior work has m…
Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks
Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm. P…
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and ada…
DAPS++: Rethinking Diffusion Inverse Problems with Decoupled Posterior Annealing
From a Bayesian perspective, score-based diffusion solves inverse problems through joint inference, embedding the likelihood with the prior…
Agentic AI for ISAC: Analysis, Framework, and Case Study
Integrated sensing and communication (ISAC) has emerged as a key development direction in the sixth-generation (6G) era, which provides ess…
Monte Carlo Query Search: Active Capability Assessment of AI Agents
Black-box AI (BBAI) systems, including foundation-model agents, are increasingly used for sequential decision making. Safe deployment requi…
CaveAgent: Transforming LLMs into Stateful Runtime Operators
LLM-based agents are increasingly capable of complex task execution, yet current agentic systems remain constrained by text-centric paradig…
SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models
Training reliable tool-augmented agents remains a significant challenge, largely due to the difficulty of credit assignment in multi-step r…
When Models Know When They Do Not Know: Calibration, Cascading, and Cleaning
When a model knows when it does not know, many possibilities emerge. The first question is how to enable a model to recognize that it does…
Inference-Time Diversity in RL-Trained Lean Theorem Provers: A Diagnostic Study
RL-trained Lean theorem provers mode-collapse at inference time: on miniF2F-test with DeepSeek-Prover-V1.5-RL, doubling the i.i.d.\ samplin…
Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron Enhancement
Large language models (LLMs) exhibit social biases that reinforce harmful stereotypes, limiting their safe deployment. Most existing debias…
Aligning Language Model Benchmarks with Pairwise Preferences
Language model benchmarks are pervasive and computationally-efficient proxies for real-world performance. However, many recent works find t…
Pushing Forward Pareto Frontiers of Proactive Agents with Behavioral Agentic Optimization
Proactive large language model (LLM) agents aim to actively plan, query, and interact over multiple turns, enabling efficient task completi…
StackingNet: Collective Inference Across Independent AI Foundation Models
Artificial intelligence built on large foundation models has transformed language understanding, computer vision, and reasoning, yet these…
When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, bench…
AutoB2G: Agentic Simulation and Reinforcement Learning for Spatio-Temporal Grid-Interactive Building Control
Grid-interactive building control has emerged as a promising approach for improving demand-side flexibility in modern power systems. Realis…
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientifi…
QED: An Open-Source Multi-Agent System for Generating Mathematical Proofs on Open Problems
We present QED, an open-source multi-agent system that turns human-provided research questions into complete mathematical proofs without fu…
Exploring LLM Agent Designs and Interaction Modalities for Scientific Visualization
This paper examines how large language model (LLM) agents perform on scientific visualization (SciVis) tasks that require generating visual…
To Use AI as Dice of Possibilities with Timing Computation
The dominant noun-based modeling paradigm, grounded in probability theory and committed to pre-specified noun entities as primitive modelin…
NEURON: A Neuro-symbolic System for Grounded Clinical Explainability
Clinical AI adoption is hindered by the black-box/grey-box nature of high-performing models, which lack the ontological grounding and narra…
SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks
Teaching language models to use search tools is not only a question of whether they search, but also of whether they issue good queries. Th…
Towards Human-Level Book-Writing Capability
Large language models are optimized for instruction following and agentic tasks remain poorly aligned with the requirements of high-quality…
Geo-Expert: Towards Expert-Level Geological Reasoning via Parameter-Efficient Fine-Tuning
While general-purpose Large Language Models (LLMs) applied to Geology often hallucinate when reasoning about subsurface structures and deep…
Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation
Telehealth systems have become increasingly important for delivering accessible and timely medical information. Existing large language mod…
Inference-Time Conformal Reasoning with Valid Factuality Control for Large Language Models
Large language models (LLMs) increasingly perform multi-step reasoning, where intermediate claims form implicit directed acyclic graphs who…
Business World Model
World model has emerged as a powerful paradigm in artificial intelligence, enabling agents to represent their environments, predict future…
Agents-K1: Towards Agent-native Knowledge Orchestration
Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Exist…
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual gro…
Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents
LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite…
ITNet: A Learnable Integral Transform That Subsumes Convolution, Attention, and Recurrence
Convolutional networks, recurrent networks, and transformers each encode different inductive biases -- locality, sequential memory, and con…
When Web Agents Finish but Still Fail: Reproducible Triggers and Trace Diagnostics for Parallel Web Exploration
Long-horizon web agents often fail in ways hidden by final-answer evaluation: they may visit useful pages, produce a well-formed answer, an…
Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents
Modern LLM agents increasingly rely on context compaction, summarization, or eviction to keep long-running sessions within a token budget.…
Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR
Reinforcement learning with verifiable rewards (RLVR) has been extended from single-domain training to multi-domain reasoning suites spanni…
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this intuition is being inve…
Unbiased Canonical Set-Valued Oracles Via Lattice Theory
A non-agentic "oracle" that reports probabilities of future events is performative: once its answer is learned and acted upon, it can chang…
Ensemble Learning Based Classification Algorithm Recommendation
Selecting an appropriate classification algorithm for a given data set remains a challenging problem in data mining and machine learning. E…
Granular-ball computing: an efficient, robust, and interpretable adaptive multi-granularity representation and computation method
To overcome the limitations of point-based inputs, overly fine computation and limited adaptability in existing artificial intelligence met…
TERC: A Transfer Entropy Redundancy Criterion for State Variable Selection in Reinforcement Learning
Identifying the most suitable variables to represent the state is a fundamental challenge in Reinforcement Learning (RL). These variables m…
Assortment Planning with Sponsored Products
In the rapidly evolving landscape of retail, assortment planning plays a crucial role in determining the success of a business. With the ri…
SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
Given the remarkable achievements in image generation through diffusion models, the research community has shown increasing interest in ext…
Causality for Tabular Data Synthesis: A High-Order Structure Causal Benchmark Framework
Existing evaluations of tabular synthesis models rely primarily on low-order statistics and downstream task performance, leaving multivaria…
ProSpec RL: Plan Ahead, then Execute
Imagining potential outcomes of actions before execution helps agents make more informed decisions, a prospective thinking ability fundamen…
Beyond Spectral Decomposition: Bayesian Contrastive Learning and its Non-negative Formulation via Factor Analysis
Factor analysis, often regarded as a Bayesian variant of matrix factorization, offers superior capabilities in capturing uncertainty, model…
Interpretable Clustering: A Survey
In recent years, much of the research on clustering algorithms has primarily focused on enhancing their accuracy and efficiency, frequently…
FLAME 3 Dataset: Unleashing the Power of Radiometric Thermal UAV Imagery for Wildfire Management
The increasing accessibility of radiometric thermal imaging sensors for unmanned aerial vehicles (UAVs) offers significant potential for ad…
XRAG: eXamining the Core -- Benchmarking Foundational Components in Advanced Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) synergizes the retrieval of pertinent data with the generative capabilities of Large Language Models (…
CASE-Bench: Context-Aware SafEty Benchmark for Large Language Models
Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption. Current LLM safety…
Bridging Neural Networks and Wireless Systems with MIMO-OFDM Semantic Communications
Semantic communications aim to enhance transmission efficiency by jointly optimizing source coding, channel coding, and modulation. While p…
Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question Answering
Large language models (LLMs) have shown remarkable capabilities in natural language processing. However, in knowledge graph question answer…
Overcoming Dependent Censoring in the Evaluation of Survival Models
Dependent censoring occurs when the event time and censoring time are not conditionally independent given the observed covariates. This com…
Distributionally Robust Reinforcement Learning with Human Feedback
Reinforcement learning from human feedback (RLHF) has evolved to be one of the main methods for fine-tuning large language models (LLMs). H…
Pose-Based Fall Detection System: Efficient Monitoring on Standard CPUs
Falls among elderly residents in assisted living homes pose significant health risks, often leading to injuries and a decreased quality of…
Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks
Large language models (LMs) offer broad generalization capabilities but require vast amounts of data and computational resources for domain…
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucinations, where generated con…
Representation Learning for Equivariant Inference with Guarantees
In many real-world applications of regression, conditional probability estimation, and uncertainty quantification, exploiting symmetries ro…
Physics-Informed Distillation of Diffusion Models for PDE-Constrained Generation
Modeling physical systems in a generative manner offers several advantages, including the ability to handle partial observations, generate…
Scaling Textual Gradients via Sampling-Based Momentum
LLM-based prompt optimization, which uses LLM-provided ``textual gradients'' (feedback) to refine prompts, has emerged as an effective meth…
Multimodal Representation Alignment for Cross-modal Information Retrieval
Different machine learning models can represent the same underlying concept in different ways. This variability is particularly valuable fo…
Towards Biosignals-Free Autonomous Prosthetic Hand Control via Imitation Learning
Limb loss affects millions globally, impairing physical function and reducing quality of life. Most traditional surface electromyographic (…
Modeling Earth-Scale Human-Like Societies with One Billion Agents
Understanding the dynamic evolution of complex social phenomena requires both high-fidelity modeling of human behavior and large-scale simu…
MGDFIS: Multi-scale Global-detail Feature Integration Strategy for Small Object Detection
Small-object detection in Unmanned Aerial Vehicle (UAV) imagery requires preserving weak local evidence while using broader context to sepa…
Code Reasoning for Software Engineering Tasks: A Survey and A Call to Action
The rise of large language models (LLMs) has led to dramatic improvements across a wide range of natural language tasks. Their performance…
GeNeRT: A Physics-Informed Approach to Intelligent Wireless Channel Modeling via Generalizable Neural Ray Tracing
Neural ray tracing (RT) has emerged as a promising paradigm for channel modeling by integrating physical propagation principles with neural…
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
Recent benchmarks for Large Language Model (LLM) agents primarily focus on evaluating reasoning, planning, and execution capabilities, whil…
Beyond Correlation: Learning Supervised, Sample-Distinct, and Eigenimage-Interpretable Representations
Conventional dimensionality reduction methods mainly optimize variance or correlation, leaving statistical dependence, data diversity, cont…
Multi-Class Human/Object Detection on Robot Manipulators using Proprioceptive Sensing
In physical human-robot collaboration (pHRC) settings, humans and robots collaborate directly in shared environments. Robots must analyze i…
LLM Serving Optimization with Variable Prefill and Decode Lengths
We study offline scheduling for large language model (LLM) serving under a fixed KV-cache memory budget, where requests have heterogeneous…
Post-training for Efficient Communication via Convention Formation
Humans communicate with increasing efficiency in multi-turn interactions, by adapting their language and forming ad-hoc conventions. In con…
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
Video Temporal Grounding (VTG) aims to localize specific video segments corresponding to natural language queries. While recent Large Visio…
Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning
Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-sol…
Tactile Gesture Recognition with Built-in Joint Sensors for Industrial Robots
While gesture recognition using vision or robot skins is an active research area in Human-Robot Collaboration (HRC), this paper explores de…
PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
Existing plant-disease datasets target classification and detection, leaving vision-language models unable to support interactive, reasonin…
EfficientUICoder: A Bidirectional Token Compression Framework for Efficient MLLM-Based UI Code Generation
Multimodal Large Language Models have demonstrated exceptional performance in UI2Code tasks, significantly enhancing website development ef…
Discovering New Theorems via LLMs with In-Context Proof Learning in Lean
Large Language Models (LLMs) have demonstrated significant promise in formal theorem proving. In this study, we investigate the ability of…
ArchesClimate: Probabilistic Decadal Ensemble Generation With Flow Matching
Internal variability is a dominant contributor to the uncertainty of predictions at the interannual to decadal timescale. A typical approac…
Predicting Effects, Missing Distributions: Evaluating LLMs as Human Behavior Simulators in Operations Management
Large language models (LLMs) are increasingly used to simulate human behavior in business, economics, and the social sciences, offering a l…
Are LLMs Reliable Rankers? Rank Manipulation via Two-Stage Token Optimization
Large language models (LLMs) are increasingly used as rerankers in information retrieval, yet their ranking behavior can be steered by smal…
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts
Training expert LLMs in domains with scarce data is difficult, often relying on multiple-choice questions (MCQs). However, standard outcome…
Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning
We treat the internals of generative models as mechanistic objects rather than black boxes. We introduce \textbf{Attribution Graphs} (AGs),…
Emergence of Minimal Circuits for Indirect Object Identification in Attention-Only Transformers
Mechanistic interpretability aims to reverse-engineer large language models (LLMs) into human-understandable computational circuits. Howeve…
Automatic Extraction of Road Networks by using Teacher-Student Adaptive Structural Deep Belief Network and Its Application to Landslide Disaster
An adaptive structural learning method of Restricted Boltzmann Machine (RBM) and Deep Belief Network (DBN) has been developed as one of pro…
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
Knowledge editing (KE) offers a lightweight alternative to retraining for updating large language models (LLMs). Meanwhile, fine-tuning rem…
Hard-constraint physics-residual networks for hydrogen crossover prediction and high-pressure extrapolation in PEM water electrolysis
Hydrogen crossover is a critical safety and efficiency constraint in high-pressure polymer electrolyte membrane water electrolysis (PEMWE),…
SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?
Optimizing the performance of large-scale software repositories demands expertise in code reasoning and software engineering (SWE) to reduc…
Scalable Synthesis of distributed LLM workloads through Symbolic Tensor Graphs
Optimizing the performance of large language models (LLMs) on large-scale AI training and inference systems requires a scalable and express…
Skin-R1: Clinical Knowledge-Guided Dermatological Diagnosis Using Vision-Language Models
Vision--language models (VLMs) have recently shown promise for assisting clinical reasoning in dermatological diagnosis. However, their tru…
SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios
Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyda…
When, How Long and How Much? Interpretable Neural Networks for Time Series Regression by Learning to Mask and Aggregate
Time series extrinsic regression (TSER) refers to the task of predicting a continuous target variable from an input time series. It appears…
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory
As intents unfold and environments change, multi-turn agents face continuously shifting decision contexts. Although reusing past experience…
Weighted Contrastive Learning for Anomaly-Aware Time-Series Forecasting
Reliable forecasting of multivariate time series under anomalous conditions is crucial in applications such as ATM cash logistics, where su…
ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
Reliable assessment of the abilities of large audio language models (LALMs) is essential to advancing the state of the art. As benchmarks r…
AI4EOSC: a Federated Cloud Platform for Artificial Intelligence in Scientific Research
The rapid growth of Artificial Intelligence and Machine Learning in scientific research has highlighted a gap between industry-standard MLO…
InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
Recent advances in diffusion models have enabled impressive video editing capabilities, yet production-grade Video Object Insertion (VOI) r…
Neural Minimum Weight Perfect Matching for Quantum Error Codes
Realizing the full potential of quantum computation requires Quantum Error Correction (QEC). QEC reduces error rates by encoding logical in…
Value-Action Alignment in Large Language Models under Privacy-Prosocial Conflict
Large language models (LLMs) are increasingly used to simulate decision-making tasks involving personal data sharing, where privacy concern…
Lost in Execution: On the Multilingual Robustness of Tool Calling in Large Language Models
Large Language Models (LLMs) are increasingly deployed as agents that invoke external tools through structured function calls. While recent…
From Word Sequences to Behavioral Sequences: Adapting Modeling and Evaluation Paradigms for Longitudinal NLP
While NLP typically treats documents as independent and unordered samples, in longitudinal studies, this assumption rarely holds: documents…
A Comparative Study of Student Perspectives on Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics
To address the scalability of feedback in computer science while mitigating the privacy and cost limitations of commercial Large Language M…
CytoCLIP: Learning Cytoarchitectural Characteristics in Developing Human Brain Using Contrastive Language Image Pre-Training
The functions of different regions of the human brain are closely linked to their distinct cytoarchitecture, which is defined by the spatia…
Diff-MN: Diffusion Parameterized MoE-NCDE for Continuous Time Series Generation with Irregular Observations
Time series generation (TSG) is widely used across domains, yet most existing methods assume regular sampling and fixed output resolutions.…
Demonstration-Free Robotic Control via LLM Agents
Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require t…
Assessing the Business Process Modeling Competences of Large Language Models
The creation of Business Process Model and Notation (BPMN) models is a complex and time-consuming task requiring both domain knowledge and…
Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment
We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions…
Agile Reinforcement Learning through Separable Neural Architecture and Applications
Deep reinforcement learning (RL) is increasingly deployed in resource-constrained environments, yet go-to function approximators - multilay…
A Large-Scale Dataset for Molecular Structure-Language Description via a Rule-Regularized Method
Molecular function is largely determined by structure. Accurately aligning molecular structure with natural language is therefore essential…
Test-Time Detoxification without Training or Learning Anything
Large language models can produce toxic or inappropriate text even for benign inputs, creating risks when deployed at scale. Detoxification…
ImprovEvolve: Basin-Hopping Meets LLM-Guided Evolutionary Search
LLM-guided evolutionary computation, most notably AlphaEvolve, has been remarkably successful in discovering novel mathematical constructio…
General and Efficient Steering of Diffusion Models
Steering diffusion models toward conditions unseen during training typically requires either retraining with conditional inputs or per-step…
Choose Your Agent: Tradeoffs in Adopting AI Advisors, Coaches, and Delegates in Multi-Party Negotiation
As AI usage becomes more prevalent in social contexts, understanding agent-user interaction is critical to designing systems that imp rove…
Mitigating the Safety-utility Trade-off in LLM Alignment via Adaptive Safe Context Learning
While reasoning models have achieved remarkable success in complex reasoning tasks, their increasing power necessitates stringent safety me…
WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its re…
On the Emergence of Implicit Curriculum in RLVR Learning Dynamics
Reinforcement learning with verifiable rewards (RLVR) has been a main driver of recent breakthroughs in large reasoning models. Yet it rema…
How to Train Your Long-Context Visual Document Model
We present the first comprehensive, large-scale study of training long-context vision language models up to 344K context, targeting long-do…
Spanning the Visual Analogy Space with a Weight Basis of LoRAs
Visual analogy learning enables image editing via demonstration rather than textual description, allowing users to specify complex transfor…
Enhanced Diffusion Sampling: Efficient Rare Event Sampling and Free Energy Calculation with Diffusion Models
The rare-event sampling problem has long been the central limiting factor in molecular dynamics (MD), especially in biomolecular simulation…
DyGnROLE: Asymmetric Pretraining for Edge Classification on Dynamic Graphs
Edge classification on directed dynamic graphs requires modeling interactions between source and destination nodes exhibiting asymmetrical…
SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal Transport
The Platonic Representation Hypothesis posits that neural networks trained on different modalities converge toward a shared statistical mod…
What Capable Agents Must Know: Selection Theorems for Robust Decision-Making under Uncertainty
As artificial agents become increasingly capable, what internal structure is necessary for an agent to act competently under uncertainty? C…
Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
Leveraging the priors of 2D diffusion models for 3D editing has emerged as a promising paradigm. However, multi-view consistency remains ch…
The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained Decoding
Large language models (LLMs) are increasingly used to generate executable outputs, JSON objects, and API calls, where a single syntax error…
Rethinking Role-Playing Evaluation: Anonymous Benchmarking and a Systematic Study of Personality Effects
Large Language Models (LLMs) have shown remarkable potential in developing role-playing agents (RPAs). However, current evaluation framewor…
Longitudinal Lesion Inpainting in Brain MRI via 3D Region Aware Diffusion
Accurate longitudinal analysis of brain MRI is often hindered by evolving lesions, which bias automated neuroimaging pipelines. While deep…
Proof-of-Guardrail in AI Agents and What (Not) to Trust from It
As AI agents become widely deployed as online services, users often rely on an agent developer's claim about how safety is enforced, which…
HEARTS: Benchmarking LLM Reasoning on Health Time Series
The rise of large language models (LLMs) has shifted time series analysis from narrow analytics to general-purpose reasoning. Yet, existing…
Feature-level Interaction Explanations in Multimodal Transformers
Multimodal Transformers often produce predictions without clarifying how different modalities jointly support a decision. Most existing mul…
FlatLands: Generative Floormap Completion From a Single Egocentric View
A single egocentric image typically captures only a small portion of the floor, yet a complete metric traversability map of the surrounding…
DiscoGen: Procedural Generation of Algorithm Discovery Tasks in Machine Learning
Automating the development of machine learning algorithms has the potential to unlock new breakthroughs. However, our ability to improve an…
Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding
Recent advances in Streaming Video Understanding has enabled a new interaction paradigm where models respond proactively to user queries. C…
UniMotion: A Unified Framework for Motion-Text-Vision Understanding and Generation
We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural l…
Grounding Sim-to-Real Generalization in Robotic Manipulation: An Empirical Study with Vision-Language-Action Models
Learning a generalist control policy for robotic manipulation typically relies on large-scale datasets. Given the high cost of real-world d…
CAPTCHA Solving for Native GUI Agents: Automated Reasoning-Action Data Generation and Self-Corrective Training
GUI agents are rapidly shifting from multi-module pipelines to end-to-end, native vision-language models (VLMs) that perceive raw screensho…
FD$^2$: A Dedicated Framework for Fine-Grained Dataset Distillation
Dataset distillation (DD) compresses a large training set into a small synthetic set, reducing storage and training cost, and has shown str…
Sustainable Hybrid Document-Routed Retrieval for Financial RAG: Resolving the Robustness-Precision Trade-off
Retrieval-Augmented Generation (RAG) systems for financial document QA typically follow a chunk-based paradigm: documents are split into fr…
Steerable Visual Representations
Pretrained Vision Transformers (ViTs) such as DINOv2 and MAE provide generic image features that can be applied to a variety of downstream…
Internalized Reasoning for Long-Context Visual Document Understanding
Visual long-document understanding is critical for enterprise, legal, and scientific applications, yet the best performing open recipes hav…
How Alignment Routes: Localizing, Scaling, and Controlling Policy Circuits in Language Models
We localize the policy routing mechanism in alignment-trained language models. An intermediate-layer attention gate reads detected content…
UniMamba: A Unified Spatial-Temporal Modeling Framework with State-Space and Attention Integration
Multivariate time series forecasting is fundamental to numerous domains such as energy, finance, and environmental monitoring, where comple…
Towards Modality-Agnostic Medical Image Anomaly Detection: A Training-Free Manifold Refinement Approach
Deploying AI-based anomaly detection across diverse clinical imaging settings remains challenging because most existing methods rely on mod…
Semantic Prompting: Agentic Incremental Narrative Refinement through Spatial Semantic Interaction
Interactive spatial layouts empower users to synthesize information and organize findings for sensemaking. While Large Language Models (LLM…
Explainable AI in Speaker Recognition -- Making Latent Representations Understandable
Neural networks can be trained to learn task-relevant representations from data. Understanding how these networks make decisions falls with…
Defeasible Conditional Obligation in a Two-tiered Preference-based Semantics (Extended Version)
In response to a concern raised by Horty, this paper develops a two-tiered, preference-based semantic framework for modeling defeasible con…
Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL
The standard post-training recipe for large multimodal models (LMMs) applies supervised fine-tuning (SFT) on curated demonstrations followe…
Most Current Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives
Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors. To study thes…
On the Spectral Structure and Objective Equivalence of Orthogonal Multilabel Fisher Discriminants
We provide a unified theoretical analysis of Linear Discriminant Analysis with simultaneous multilabel scatter matrix formulations and Stie…
When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning
Behavior Cloning (BC) has emerged as a highly effective paradigm for robot learning. However, BC lacks a self-guided mechanism for online i…
BioProVLA-Agent: An Affordable, Protocol-Driven, Vision-Enhanced VLA-Enabled Embodied Multi-Agent System with Closed-Loop-Capable Reasoning for Biological Laboratory Manipulation
Biological laboratory automation can reduce repetitive manual work and improve reproducibility, but reliable embodied execution in wet-lab…
Globally Optimal Training of Spiking Neural Networks via Parameter Reconstruction
Spiking Neural Networks (SNNs) have been proposed as biologically plausible and energy-efficient alternatives to conventional Artificial Ne…
Robust Multi-Agent LLMs under Byzantine Faults
Large language model (LLM) agents increasingly collaborate over peer-to-peer networks to improve their reliability. However, these same int…
Cornerstones or Stumbling Blocks? Deciphering the Rock Tokens in On-Policy Distillation
While recent work in Reinforcement Learning with Verifiable Rewards (RLVR) has shown that a small subset of critical tokens disproportionat…
Metal-Sci: A Scientific Compute Benchmark for Evolutionary LLM Kernel Search on Apple Silicon
We present Metal-Sci, a 10-task benchmark of scientific Apple Silicon Metal compute kernels spanning six optimization regimes (stencils, al…
Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
Recent studies show that gradient-based universal image jailbreaks on vision-language models (VLMs) exhibit little or no cross-model transf…
Vividh-ASR: A Complexity-Tiered Benchmark and Optimization Dynamics for Robust Indic Speech Recognition
Fine-tuning multilingual ASR models like Whisper for low-resource languages often improves read speech but degrades spontaneous audio perfo…
Not All Timesteps Matter Equally: Selective Alignment Knowledge Distillation for Spiking Neural Networks
Spiking neural networks (SNNs), which are brain-inspired and spike-driven, achieve high energy efficiency. However, a performance gap betwe…
FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning
Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, t…
SCRIBE: Diagnostic Evaluation and Rich Transcription Models for Indic ASR
Automatic speech recognition replaces typing only when correction costs less than manual entry - a threshold determined by error types, not…
Are Frontier LLMs Ready for Cybersecurity? Evidence for Vertical Foundation Models from Dual-Mode Vulnerability Benchmarks
We evaluate whether frontier LLMs are ready for cybersecurity through a dual-mode benchmark: white-box function-level vulnerability detecti…
When Good Equations Get Bad Scores: Improving Symbolic Regression Through Better Parameter Optimization
Symbolic Regression (SR) plays a central role in scientific knowledge discovery by distilling mathematical equations from observational dat…
High-Risk AI Systems and the Problem of Identity in the European AI Act
The EU Artificial Intelligence Act (AIA) establishes a lifecycle governance regime for high-risk AI systems built around ex-ante conformity…
ChainCaps: Composition-Safe Tool-Using Agents via Monotonic Capability Attenuation
Tool-using agents increasingly operate in open-ended deployment environments, where they compose file systems, web APIs, code interpreters,…
Self-Play Reinforcement Learning under Imperfect Information in Big 2
Imperfect-information multiplayer games test whether agents can act under hidden information, sparse rewards, and non-stationary opponents.…
MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings
Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in realistic, electronic health recor…
PatchWorld: Gradient-Free Optimization of Executable World Models
Text-agent environments are typically modeled as partially observable Markov decision processes (POMDPs), assuming that the simulator's lat…
BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali
Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large langua…
Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models
Connector-based video unified models have demonstrated strong capability in instruction-grounded video synthesis, but integrating a large h…
Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
On-Policy Distillation (OPD) improves large language model reasoning by training a student model on trajectories sampled from its own polic…
Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion
Recent Vision-Language Models (VLMs) struggle with grounded reasoning, temporal consistency, and context aware planning in videos. We intro…
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs
Current 3D spatial reasoning methods face a fundamental trade-off: neuro-symbolic 3D (NS3D) concept learners achieve interpretable reasonin…
SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence
As LLM-based agents expand their operational scope, reliability becomes a prerequisite for real-world deployment. However, in practical app…
Agent libOS: A Runtime Substrate for Capability-Controlled Self-Evolving LLM Agents
Large language model (LLM) agents are becoming long-running software actors rather than fixed tool users. They accumulate memory, activate…
LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection
Existing quantization methods are fundamentally limited by rigid, integer-based bit-widths (e.g., 2, 3-bit), resulting in a ``deployment ga…
From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents
Large language model (LLM)-based agents are evolving from passive text generators into autonomous systems capable of planning, tool use, re…
MASF: A Multi-Model Adaptive Selection Framework for Abstractive Text summarization
Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents…
Improving Answer Extraction in Context-based Question Answering Systems Using LLMs
Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face ch…
Evidence Graph Consistency in Retrieval-Augmented Generation: A Model-Dependent Analysis of Hallucination Detection
Retrieval-Augmented Generation (RAG) reduces but does not eliminate hallucination in large language models. Existing detection methods rely…
An AI Security Agent for University ACMIS: Multi-Vector Threat Detection and Automated Response
University Academic Management Information Systems (ACMIS) are high-value targets for a wide spectrum of security threats including brute-f…
APEX4: Efficient Pure W4A4 LLM Inference via Intra-SM Compute Rebalancing
W4A4 quantization promises full utilization of INT4 Tensor Cores, yet group dequantization overhead on CUDA Cores has driven existing syste…
Agentic Social Affordance Framework (ASAF): Agent Identity Design as a Collaboration Interface in Multi-Agent Systems
As AI systems evolve from single agents to multi-agent architectures, a critical design dimension has been overlooked: how the social ident…
Quantifying Subliminal Behavioral Transfer Ratios in Language Model Distillation
Distillation of a language model intended to transfer benign behavior to a student model may also transfer undesirable characteristics, if…
The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems
Nowadays, the autonomous execution of cyberattacks capable of causing substantial real-world harm is widely regarded as one of the critical…
MeEvo: Metacognitive Evolution Combined with Natural Evolution for Automatic Heuristic Design
Large Language Models (LLMs) have advanced Automatic Heuristic Design (AHD) by enabling heuristic generation through reasoning and code syn…
CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation
Granting LLMs direct control over costly, irreversible scientific experiments leads to unsafe exploration and unstable performance, but dis…
X-Tokenizer: A Multimodal Action Tokenizer for Vision-Language-Action Pretraining
Modern Vision-Language-Action (VLA) models must bridge pretrained vision-language reasoning and precise continuous robot control. Existing…
EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining
Color fundus photography (CFP) is the mainstay of large-scale retinal screening, but its diagnostic capacity is limited by the lack of dept…
Surprise-Guided MergeSort: Budget-Efficient Human-in-the-Loop Ranking via Adaptive Comparison Scheduling
Pairwise comparison is the gold standard for subjective ranking tasks; however, exhaustive annotation requires a massive number of human co…
Entropy-Gated Latent Recursion
Inference-time scaling has become the dominant lever for improving language-model reasoning, but existing methods derive rollout diversity…
An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts
Banks face two threat families with fundamentally different detection requirements: signature-based fraud (card-not-present attacks, accoun…
A Knowledge Theory of Capital:The Value of Natural and Artificial Intelligence, Volume 1
This volume develops a knowledge theory of capital for economies in which productive capacity increasingly resides in software, data, model…
RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation
Graph-based retrieval at billion-node scale requires jointly solving three tightly coupled problems -- graph construction, representation l…
NeuralMUSIC: A Hybrid Neural-Subspace Framework for Robot Sound Source Localization
Reliable sound source localization is fundamental to robot audition, enabling autonomous robots to perceive spatial cues and operate effect…
Explaining Attention with Program Synthesis
A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic desc…
Towards Engineering Scaling Laws with Pretraining Data Composition
Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established…
Analyzing Defensive Misdirection Against Model-Guided Automated Attacks on Agentic AI Systems
Agentic AI systems increasingly rely on language-model components to interpret instructions, process external data, invoke tools, and coord…
SARLO-80: Worldwide Slant SAR Language Optic Dataset 80cm
Multimodal foundation models have advanced rapidly thanks to large optical benchmarks, but comparable resources for synthetic aperture rada…
Trust in Generative AI for Health Information Consumption and the Effect of Learned Dependency: An Experimental Investigation
Background: Generative artificial intelligence (GenAI) is increasingly used for health information, yet its influence on users' trust calib…
Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling
Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wis…
SwarmX: Agentic Scheduling for Low-Latency Agentic Systems
Agentic AI applications compose multiple model calls and tool executions, creating new scheduling challenges for GPU-CPU clusters. Their in…
Learning to Trigger: Reinforcement Learning at the Large Hadron Collider
High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tig…
Towards Spec Learning: Inference-Time Alignment from Preference Pairs
Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on…
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency…
ATMA: Length-Invariant Language Modeling via Polar Attention and Gated-Delta Compression Memory
Modern large language models based on softmax scaled-dot-product attention are constrained by their training sequence length: as the key-va…
Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
A language model's memory can be worse than no memory at all. A memory that keeps a wrong conclusion but drops the work behind it makes the…
The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators
Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However,…
Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model
Dense embeddings power semantic search and retrieval-augmented generation, yet a leaked vector database also leaks the text behind it, beca…
CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation
Long-horizon, contact-rich complex manipulation tasks, such as seating a GPU into a PCIe slot, demand both millimeter high precision and ou…
Robust Onion: Peeling Open Vocab Object Detectors Under Noise
The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity…
CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention
Recurrent models must forget in order to remember, yet the state of the art decides what to erase without consulting what is stored -- the…
「『ハイキュー!!』全巻を読み込ませている」――東宝がアニメグッズの監修にAI活用、他作品にも拡大へ
東宝はこのほど、アクセンチュアと協力し、IPを活用した商品の監修にAIシステムを導入した。同システムの詳細や展望について、アクセンチュアの戸賀慶氏と東宝の田中亮史氏が語った。
国内大手が共同出資のAI開発企業「日本AI基盤モデル開発」、新名称「Noetra」で始動 産総研と国産マルチモーダルAI開発へ
国内の大手企業が共同で出資するAI開発企業の日本AI基盤モデル開発(東京都渋谷区)は6月30日までに、1日付で名称をNoetraに変更したと発表した。
【徹底入門】AIエージェントで注目の「AX」とは何か 人間だけじゃなく“AI視点の使いやすさ”も重要に?
AIエージェント時代に注目を集めつつある「AX」(Agent Experience、エージェント体験)とは何か解説する。
Vibe coding platform Base44 launches own model as AI startups seek defensibility
Wix-owned vibe coding platform Base44 has started rolling out its own AI model — with hopes that it will eventually outperform frontier mod…
「“社長AI”って意味ある?」→言った本人も手のひら返し 幹部の9割が高評価したNTTドコモビジネスの「AI小島社長」開発録
経営トップの判断や思考をAIで再現する取り組みが、国内の大企業に広がっている。NTTドコモビジネスが開発し、同社幹部の9割が「方針理解に役立った」と評価するという「AI小島社長」に迫る。
GitHub、AIによる雑なプルリクエストを抑制へ ユーザー当たりのプルリク数に上限を設定できる新機能
米GitHubは、ユーザーに対してプルリクエスト数の上限を設定できる新機能の導入を発表しました。
あなたのAWS、コストの課題はどこにある? AIが教えてくれる「AWS FinOps Agent」パブリックプレビュー開始
米Amazon Web Services(AWS)は、使用中のAWSのコストに関する質問や、コストに異常が発生した場合にその原因を調査して特定してくれる「AWS FinOps Agent」のパブリックプレビュー開始を発表しました。
AI避けて「人間にだけ届く」広告配信へ、博報堂DYが新会社設立 虹彩認証「World ID」活用
サム・アルトマン氏らが共同発明した人間認証技術「World ID」を活用する。
「Fable 5を全タスクに使う必要はない」 Anthropic開発者直伝のトークンコスト節約術
Anthropicの開発者がトークンコストを抑えるための戦略を語った。タスクに応じてモデルを使い分けると共に、適切な入力方式を選ぶことが重要だ。
「AIが前提となる世界」でSIerは生き残れるか?
AIが前提となる世界で問われるのは「AIをどう使うか」ではなく、「組織をどう設計し直すか」だ。ITRアナリストと「新しい乱世」を生き残るための道筋を考える。
中小企業の採用から勤怠・経費管理まで バラバラのSaaSからまとめてAIがデータ分析
中小企業では、複数のSaaSや紙、Excelが混在することで、管理業務の負担が増えている。NoahWorksは、採用から勤怠、決済までを一元化し、AIによるデータ分析で現場の業務改善を支援する。
自社の業務に合わせたAIエージェントを「10分で作成」 freeeが「AI戦略」を強化
クラウド会計システム「freee」の開発などを手掛けるフリーは、2月に発表したAI戦略の実現に向けた新たな取り組みとして「freee AIアシスタント」と「freee カスタムオーダー」の提供を6月に開始した。「AIから最も使いやすいSaaS」として、AI業界におけるリーディン…
社長もAIが代わる時代に 社員の相談にいつでも答えるエージェント「AI社長」が登場
日テレHR総合研究所は、社長やキーパーソンの価値観と判断基準を基に答える専用AI「AI社長」の提供を開始した。忙しい社長や決定層の思考を学習し、考えの整理と判断の質向上を支える社内用の相談役として活用できる。
メール、Teams、Slack――バラバラな連絡ツールの「見落とし」 ChatGPT Agentsで解決する方法
メールやTeams、Slackなど複数の連絡ツールを使っていると、重要なメッセージの見落としや返信漏れをしてしまいます。AIで解決できませんか?
死んだのは「低成長モデル」だけ HRBrainのCSaOが読み解く「SaaS is Dead」の本質
Salesforce出身で、創業10年のHRテック・HRBrainのCSaOを務める小山径氏。同氏は「SaaSは死んでいない」と主張する。その理由は?
AIで人は幸せになれるのか? 「AIで稼ぐ企業」と「コストを負担する企業」
Appleが複数製品の価格を引き上げる一方で、AI需要の拡大を追い風にキオクシアなど半導体関連企業は成長を続けている。技術革新が生む大きな利益の裏側で、誰が恩恵を受け、誰がコストを負担するのか。
「前任者が不在」でも大丈夫 PCログからAIがマニュアルを自動作成する時代へ
特定の担当者に業務が依存する属人化や、急な退職・異動による引き継ぎ不足は、多くの企業が抱える課題だ。こうした問題を解決するため、PCログのデータからAIが業務マニュアルを自動作成する仕組みが提供された。
ルネサスが2035年の売上高3倍増も視野に、AIで3段階の成長を目指す
ルネサス エレクトロニクスが同社の概況や事業方針などについて説明。足元で半導体市場の拡大をけん引するAIに焦点を当てた事業展開を強化し、AIインフラ、フィジカルAIとSDV、「Intelligence at the Edge」の3段階で優位なポジションを構築し成長を目指す。
日本の「完璧主義」から脱却し中国ヒューマノイドにどう立ち向かうか
ハードウェアと市場が先行して急拡大する一方で、自律制御を担う基盤モデルの領域にはいまだ乗り越えるべき壁が多い。後編となる本稿では、オープンソース化で社会実装を急ぐ中国プレイヤーの動向を解説。圧倒的なスピードで独走する中国に対し、日本が目指すべき生存戦略を提示する。
Gemini’s personalized AI image generation is now free for US users
Google is expanding Gemini’s personalized AI image generation to eligible free users in the U.S., allowing the chatbot to create images bas…
Anthropic and Gov. Newsom forge deal allowing California government to use Claude at half price
As Anthropic forges a closer relationship with the state of California, the federal government has made an enemy out of the OpenAI rival.
South Korean tech giants commit over $550B to ease ‘RAMageddon’
The world's two largest memory chip companies vow to build more memory lab fabs as South Korea positions itself as an AI tech powerhouse co…
Arena, the AI leaderboard everyone uses, is now a $100M business
The startup, which runs a popular free AI leaderboard, launched its commercial service just last September.
Cursor now has a mobile app for guiding your coding agent on the go
Cursor has launched a new mobile app for remote oversight over coding agents.
TIDAL cracks down on AI music by cutting off monetization
In addition, TIDAL will use automated tools to remove AI-generated music that attempts to impersonate an artist or a group, the company sai…
2026-06-29(246件)
Robot hand company settles Tesla trade secret suit and announces $11M raise
The startup, Proception, is taking a unique approach to collecting training data to tackle one of the hardest problems in robotics: hands.
Omen AI’s plan to optimize data centers is all wet
Omen AI raised a $31 million Series A to monitor chip coolant and stop bacterial outbreaks in data centers.
「ヤフコメまとめ」開始 ヤフコメの論点、AIがグラフで可視化
OpenAIのAPIを活用し、ユーザーが投稿したコメントの論点を生成AIが分類し、グラフ化する。
Mapping Europe’s AI Workforce Opportunity
A new OpenAI report maps how AI could reshape jobs across the EU, highlighting which occupations may face automation, growth, or workflow c…
AIエージェントの投資優先順位、どう決める? Gartnerが「投資スコア」の作り方を公開
業務におけるAIエージェントの投資優先順位をどう決めればよいか。業務・業種別のAIエージェントはどう進化していくか。ガートナージャパンの著名アナリストである亦賀忠明氏のWebセミナーから探る。
AI-Model Network: Concept, Current State and Future
While the primary function of computers lies in computation and processing, the core value of the Internet is rooted in sharing and collabo…
When Does Personality Composition Matter for Multi-Agent LLM Teams?
Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes rema…
Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning
Large language model (LLM) agents have demonstrated strong capability in sequential decision-making, yet they remains fundamentally reactiv…
Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models
We introduce a categorical framework called ODYSSEY for constructing verifiable, local truth-preserving foundation models as compositions o…
DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
Dyslexic learners increasingly use artificial intelligence (AI) tools to support reading, writing, organisation, and study-related tasks. H…
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it ma…
ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation
The rapid spread of fake news poses increasing threats to information ecosystems, especially as AI-generated misinformation under Generativ…
Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework
Large language models (LLMs) have attracted widespread attention from academia and industry, yet their deployment raises critical security…
Understanding Rollout Error in Graph World Models
World models are often used for planning by rolling learned dynamics forward. Many planning environments, however, are not vectors or image…
Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents
World models for language agents come in two useful forms. An agent-based world model calls an LLM API and reasons flexibly in language, bu…
ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents
Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-poli…
NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning
Multimodal large language models (MLLMs) are increasingly deployed as embodied planners in egocentric environments, where task success requ…
Verifiable Geometry Problem Solving: Solver-Driven Autoformalization and Theorem Proposing
Geometry Problem Solving have increasingly adopt the neuro-symbolic paradigm, combining neural intuition with symbolic rigor. However, curr…
RelBall: Relation Ball with Quaternion Rotation for Knowledge Graph Completion
Real-world knowledge graphs are often incomplete, lacking many valid facts. Knowledge Graph Completion (KGC) aims to predict missing links…
Lifted Causal Inference
Lifted inference exploits indistinguishabilities in probabilistic graphical models by using a representative for indistinguishable objects,…
JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications
JD.com, one of the world's largest e-commerce platforms, serves over 700 million active users and millions of merchants, with a catalog of…
Ontology-Guided Evidence Path Inference for Multi-hop Knowledge Graph Question Answering
Knowledge graph question answering (KGQA) aims to answer natural-language questions by reasoning over structured facts. Existing multi-hop…
AI-Driven Synthesis for High-Tech System Design: Automating Innovation
This article addresses the combinatorial complexity inherent in modern high-tech system design by presenting automation-in-design (AiD) as…
Tandem Reinforcement Learning with Verifiable Rewards
Reinforcement learning with verifiable rewards (RLVR) has significantly improved the reasoning capability of large language models, reachin…
Agent-Native Immune System: Architecture, Taxonomy, and Engineering
The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaborati…
DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers
The rapid growth of Large Transformer-based models, specifically Large Language Models (LLMs), now scaling to trillions of parameters, has…
Position: The Term "Machine Unlearning" Is Overused in LLMs
Large language models increasingly face demands to "forget" training data, knowledge, or behaviors due to regulatory deletion obligations,…
OverFlowLight: Real-Time Gridlock Prevention and Traffic Signal Optimization for Urban Intersections
Queue overflow, a severe consequence of urban traffic congestion, occurs when vehicle queues exceed intersection capacity, obstructing upst…
CalBrief: A Pilot Diagnostic Benchmark for Evidence-Calibrated Scientific Briefing with Large Language Models
Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takea…
Agentic Publication Protocol: An Attempt to Modernize Scientific Publication
Scientific publication is still organized primarily around static manuscripts, even though much of scientific progress depends on tacit kno…
SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game
Evaluating LLM agents requires dynamic environments that go beyond static reasoning and zero-sum games. Real-world economic interaction is…
Automated brain tumor detection in MRI images using CNN and ResNet architectures
Deep learning has shown significant potential in medical image analysis, particularly for disease detection using MRI scans. Accurate and e…
Towards Evaluation of Implicit Software World Models in Coding LLMs
Software engineering, whether performed by humans or by AI agents, requires reasoning about how software behaves. We call the internal mode…
Compression-Driven Anomaly Detection in Brain MRI Using an Interpretable Quantum Autoencoder
We study a quantum autoencoder (QAE) for compression-driven anomaly detection in brain MRI data. The approach leverages angle encoding to m…
Not All Relations Rotate Alike: Transformation-Aware Decoupling for Viewpoint-Robust 3D Scene Graph Generation
3D Scene Graph Generation (3DSGG) represents 3D scenes as structured object-relation-object graphs, providing a compact relational abstract…
GRAFT: Biological Graph and Hypergraph Benchmarks for Linked Gene Expression and Phenotypic Trait Prediction in Arabidopsis thaliana
Understanding which genes control which traits in an organism remains one of the central challenges in biology. Despite significant advance…
Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents
Large language model (LLM) agents operate over long, multi-session interactions in which facts change: a user moves, a price updates, a pla…
Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks
How should we evaluate generation systems that combine autoregressive (AR) and diffusion decoding? We study this question through Speculati…
DMV-Bench: Diagnosing Long-Horizon Multimodal Agents' Visual Memory with Incidental Cue Injection
Research on agent memory has matured rapidly, but almost entirely on the text side: few existing benchmarks ask, in an interactive environm…
Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge
Large Language Models (LLMs) possess broad conceptual knowledge acquired through large-scale text pretraining, yet their potential to super…
The Context-Ready Transformer
We introduce the context-ready transformer, a new recurrent neural network architecture built from a D-layer transformer block that pre-con…
Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction
Social media popularity prediction aims to forecast the future reach or influence of online content from early-stage observations. Accurate…
On the Inseparability of Instructions and Data in Shared-Embedding Sequence Models
Prompt injection is the top security risk for LLM-integrated applications, yet every defense proposed so far has been broken. We prove this…
hia-gat: A Heterogeneous Interaction-Aware Graph Attention Network For Frame-Level Traffic Conflict Risk Prediction On Freeways
This paper formulates frame-level freeway risk assessment as a multi-agent scene graph-level binary classification problem, where each vide…
PEBS: Per-rater Empirical-Bayes Shrinkage for RLHF Reward-Model Calibration
Reward models for Reinforcement Learning from Human Feedback (RLHF) pool preferences across thousands of annotators and fit one global affi…
Distribution-based deep multiple instance learning for tumor proportion scoring in NSCLC
Accurate assessment of tumor proportion score (TPS) in non-small cell lung cancer (NSCLC) is critical for treatment planning and prognosis.…
Retroactive Advantage Correction: Closed-Form V-Trace Bias Correction for Delay-Aware RLHF
Reinforcement learning from human feedback (RLHF) in production does not always have a synchronous reward signal. Code-execution verifiers,…
SceneBot: Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction
Current humanoid reinforcement-learning policies excel at free-space motions but struggle with contact-rich tasks, as pure kinematic tracki…
CoIn: Comprehensive 2D-3D Inpainting with Gaussian Splatting Guidance
3D scene inpainting is essential for reconstructing areas corrupted by occlusions or limited viewpoints. While recent methods leverage Gaus…
Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding
Large Vision-Language Models (LVLMs) exhibit sophisticated reasoning but remain susceptible to object hallucination. Deviating from the pre…
Narrative-UFET: Narrative Generation for Ultra-Fine Entity Typing
Ultra-fine entity typing (UFET) assigns highly specific types to entity mentions, but current approaches struggle with types in the long ta…
Global Explanations for Multivariate Time Series Forecasting Models via $K$-Order Markov Approximations
While many explainable AI (XAI) methods have been proposed, most are not designed for time-series forecasting models and often rely on the…
HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models
Discrete audio representations have become increasingly popular for building multimodal text-audio systems and integrating audio capabiliti…
Cross-Platform Chinese Offensive Comment Detection via Dual-Threshold Hard Example Mining
Cross-platform deployment of offensive comment detection for Chinese social media suffers performance degradation. The paper proposes a dua…
Reconstructing the Developmental Trajectory of Adipocytes in Human Adipose Tissue Using Single-Cell RNA Sequencing
Obesity is a global health crisis associated with metabolic disorders such as type 2 diabetes and cardiovascular disease. This study employ…
Explainable AI for Biodiversity Monitoring and Ecological Image Analysis
Artificial intelligence is transforming biodiversity monitoring by enabling automated analysis of ecological imagery collected from camera…
From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models
Probe-based uncertainty estimation (UE) has emerged as a prominent approach to detect hallucinations in Large Language Models (LLMs) by lea…
CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence
Edge devices increasingly invoke large language models (LLMs) through API services for context aware edge intelligence, while edge generate…
Mitigating LLM-based p-Hacking by Preregistering for the Next LLM
Large language models (LLMs) are increasingly used to generate, classify, and annotate data whose outputs feed downstream hypothesis tests.…
Deployment-Side Adaptiveness in Multi-Horizon Volatility Forecasting
In financial forecasting, predictive performance depends not only on which model is trained, but also on how the trained model is deployed.…
Halt Fast! Early Stopping for Certified Robustness
Randomized Smoothing (RS) provides rigorous robustness guarantees for neural networks without architectural constraints, yet its adoption i…
Class-frequency Guided Noise Schedule for Diffusion Models
In this paper, we are the first to examine the correlations between class frequency and the multi-scale noise schedule within diffusion mod…
What Was That Again? Certified Robustness for Automatic Speech Recognition
Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations. While this has been repeatedly…
Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks
While voice control is rapidly becoming a ubiquitous vector of human-AI communication, the risks facing these systems remain poorly underst…
Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning
Recent work has shown that fine-tuning large language models (LLMs) for social warmth degrades factual reliability and increases sycophancy…
The Simulacrum: Decision-Theoretic Pretraining for Near-Optimal Time-Series Forecasting and Inference
We introduce a neural network-based framework for learning time series estimators through a process we term decision-theoretic pretraining.…
Do Speech Emphasis Models Generalize across Languages and Emotions?
Prosodic emphasis varies across languages, emotions, and speaking styles, yet existing emphasis detection models are largely trained and ev…
Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment
Despite their strong general capabilities, large language models (LLMs) often remain unreliable when outputs must be numerically precise. A…
Bifocal Diffusion Language Models: Asymmetric Bidirectional Context for Parallel Generation
Discrete diffusion language models (dLLMs) recover masked tokens in parallel, offering significant speedups over autoregressive (AR) genera…
KG2Cypher: Data-Centric Pipeline for Building Enterprise Text-to-Cypher Systems
Enterprise Knowledge Graphs (KGs) are increasingly used for internal search, analytics, and question answering, but building natural-langua…
End-to-End Dynamic Sparsity for Resource-Adaptive LLM Inference
Large Language Models (LLMs) inference is typically deployed under a static resource assumption, where models execute a fixed computational…
Flexformer: Flexible Linear Transformer with Learnable Attention Kernel
Transformer models rely on attention mechanism to capture long-range dependencies but suffer from quadratic complexity, limiting their scal…
From General-Purpose Audio Tagging to Spatially Grounded Sound Event Localization and Detection
This report investigates the extension of pretrained General-Purpose Audio Tagging (GP-AT) models toward spatially grounded Sound Event Loc…
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?
Vision-Language-Action (VLA) models enable instruction-driven robotic manipulation, but they inherit oversized language backbones from pret…
RS-Diffuser: Risk-Sensitive Diffusion Planning with Distributional Value Guidance
Offline reinforcement learning enables policy learning from fixed datasets without additional environment interaction, making it appealing…
Improving Adversarial Robustness via Activation Amplification and Attenuation
The existence of adversarial attacks is often attributed to the presence of non-robust features in neural networks. While prior defenses re…
Output-Space Allocation Costs for Calibration-Guided LLM Compression: An Empirical Study
Training-free compression methods for large language models (LLMs) often use calibration data to guide compression decisions. ROCKET, a rec…
SHIFT: Gate-Modulated Activation Steering for Knowledge Conflict Mitigation in Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) enhances LLMs by incorporating external knowledge to support response generation. However, conflicts b…
NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation
Hybrid attention models that mix full and sliding-window attention across layers offer a promising approach to efficient long-context infer…
Position Bias Correction is Insufficient for One-Pass Attention Sorting
Long-context language models suffer from position bias, where information in middle positions is underutilized. Attention Sorting addresses…
Optimizing Teacher-Student Partitioning for Scalable Knowledge Distillation on HPC Systems
Knowledge Distillation (KD) enables training smaller student models under the guidance of larger teacher models, and the widely adopted TRL…
Parameter-Efficient Quantum-Inspired Fast Weight Programmers for Traffic-Matrix Forecasting
Traffic matrices (TMs) capture network-wide origin-destination demand and are central to traffic engineering, yet accurate whole-matrix for…
Pepti-drift: Toxicity-Repulsive Drifting for Antigen-Conditioned Discrete Peptide Generation
Peptides are a promising therapeutic modality that combine the chemical tunability of small molecules with the target specificity of macrom…
Hippocampus-DETR: An Explicit Memory Object Detection Framework Based on Hippocampus Modeling
This paper addresses the lack of explicit memory mechanisms in current object detection models and proposes Hippocampus-DETR, a novel detec…
WattLayer: Get Layers Right to Estimate Inference Energy of Neural Networks
The widespread adoption of Artificial Intelligence (AI) has led to increasing concerns about energy consumption, yet there is a lack of sta…
Applicability of memorization indicators for early spotting of overfitting while recalibrating sEMG-decoders on low sample sizes
Deep learning models for surface electromyography (sEMG) can benefit substantially from subject-specific (re-)calibration, since no suffici…
GNBAN: Graph Neural Basis Attention Networks for Long-Horizon Forecasting over Large Entity Sets
Demand forecasting at the bottom of a retail hierarchy requires predicting tens of thousands of correlated long-horizon series across produ…
S$^2$-VLA: State-Space Guided Vision-Language-Action Models for Long-Horizon Manipulation
Vision-Language-Action (VLA) models have demonstrated strong capabilities in robotic manipulation, but their performance degrades significa…
SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion
Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation. However, existi…
A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts
Temporal variation poses a unique challenge for named entity recognition (NER) in historical texts, where entities drift in surface form an…
SEADA: An efficient methodology for optimizing mixed-precision DNNs on multi-precision spatial architectures
Mixed-precision computation has been introduced in deep neural networks (DNNs) as an effective approach to reduce latency, energy consumpti…
Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs
Theory-of-mind evaluations of large language models typically use dyadic social-deduction games, where every observable cue points to a sin…
Every Step of the Way: Video-based Parkinsonian Turning Step Counting
As a prominent symptom of Parkinson's disease (PD), turning impairment is evaluated through parameters such as turning angle, duration, and…
Reflect-R1: Evidence-Driven Reflection for Self-Correction in Long Video Understanding
Current multimodal reflection mechanisms for long video understanding predominantly rely on closed-loop self-reflection within internal par…
Home3D 1.0: A High-Fidelity Image-to-3D Asset Generation System for Interior Design
We present Home3D 1.0, a modular image-to-3D generation system that produces high-quality 3D assets from a single reference image, targetin…
Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy
The widespread collection of fine-grained location data by commercial data brokers creates a re-identification risk that is not widely reco…
Two-Stage Fine-Tuning for Protein Sequence Generation with Targeted Amino-Acid Composition
Protein language models are standard priors for biological sequence generation, but steering them toward explicit distributional design tar…
VASAE: Naming SAE Dictionary Directions with Vocabulary-Aligned Anchoring
Sparse autoencoders (SAEs) provide useful decompositions of Transformer residual streams, but their learned features are usually named post…
Reasoning Beyond Prediction: From Data-Driven to Causal Software Engineering
Software engineering is an intellectually demanding, creative discipline that juggles a web of interdependent tasks to design, build, and a…
From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection
Speech-based cognitive impairment detection offers a noninvasive, accessible alternative to costly biomarker assays, yet transformer-based…
ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering
Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowledge. Most prior metho…
SHARD: cell-keyed residual splitting for alignment-resistant private dense retrieval
Dense embeddings underpin semantic search and RAG, yet a leaked vector store hands much of the underlying text back to whoever holds it. Th…
Parallel Rollout Approximation for Pixel-Space Autoregressive Image Generation
Pixel-space continuous-token autoregressive (AR) generation directly models images as sequences of raw pixel patches, avoiding discrete tok…
Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection
Insurance fraud imposes substantial financial losses and operational inefficiencies, raising premiums and impacting trust among legitimate…
MLVC: Multi-platform Learned Video Codec for Real-World Deployment
Neural video codecs have surpassed classical codecs in coding efficiency but remain impractical for deployment due to cross-platform incomp…
Mind the Gap: Quantifying the Domain Gap in Cross-Sensor Diffusion Super-Resolution
Demand for high-resolution satellite imagery has increased interest in super-resolution (SR) to bridge the spatial resolution gap between f…
DG^VoiC: Speaker Clustering for Fraud Investigation under Real Call-Centre Conditions
Insurance fraud remains costly and operationally difficult, particularly in call-centre workflows where many customer interactions begin at…
Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA
LLM-as-a-Judge and self-evaluation pipelines implicitly assume that evaluation is easier than generation. We test this in a controlled in-c…
MultiHashFormer: Hash-based Generative Language Models
Language models (LMs) represent tokens using embedding matrices that scale linearly with the vocabulary size. To constrain the parameter fo…
ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents
Large language models (LLMs) have increasingly moved from standalone text generation systems to agents that invoke external tools, access e…
Single and Multi Truth Data Fusion using Large Language Models
Data fusion, also known as truth discovery, is a data integration problem that aims to determine the correct value or set of values for eac…
OperatorSHAP: Fast and Accurate Shapley Value Estimation for Neural Operators
Understanding model predictions is essential for physical applications, where outputs often inform safety-critical decisions, such as struc…
STAG: Spatio-temporal Evolving Structural Representation of Action Units for Micro-expression Recognition
Micro-expression recognition is challenging due to subtle and short-lived facial muscle movements. Existing methods rely heavily on apex-on…
OSOR: One-Step Diffusion Inpainting for Effect-Aware Object Removal
Real-world object removal is challenging due to two key difficulties: the target object's non-local effects, such as shadows and reflection…
BiDeMem: Bidirectional Degradation Memory for Explainable Image Restoration
Degradation-aware prompts, conditions, and latent priors are increasingly used in image restoration, yet they are usually judged by a singl…
Higher-Order Fourier Neural Operator: Explicit Mode Mixer for Nonlinear PDEs
Neural operators provide deep neural networks for learning mappings between function spaces. Among them, the Fourier Neural Operator (FNO)…
From Tokens to States: LLMs as a Special Case of World Models and the Continuous Path Beyond
The AI community has framed the relationship between large language models (LLMs) and world models as a dichotomy: LLMs predict tokens; wor…
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators a…
Beyond Sparse Supervision: Diffusion-Guided Learning for Few-Shot Graph Fraud Detection
Graph-based fraud detection is essential for safeguarding large-scale transaction systems, where undetected anomalies may lead to substanti…
Toward Robust In-Context Segmentation via Concept Guidance
In-context segmentation (ICS) requires a model to segment target regions in a query image using only a few reference images and their corre…
Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models
Jailbreak attacks bypass LLM safety alignment, yet their mechanisms remain poorly understood. We provide evidence that attacks do not compr…
CPAgents: Agentic Composite Phenotype Generation for Cardiac Disease Association
Identifying robust associations between cardiac imaging phenotypes and clinical diseases is fundamental to population-scale cardiovascular…
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior
Embodied agents operating in decentralized and partially observable environments have attracted growing attention in recent years. However,…
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test constru…
The Remittance Blueprint: Data-driven Intelligence for Sri Lanka
This study analyzes Sri Lankan migration and remittances over 32 years (1994-2025). Using a 384-month harmonized dataset, we apply explorat…
HAT-4D: Lifting Monocular Video for 4D Multi-Object Interactions via Human-Agent Collaboration
Extracting dynamic 4D object interactions from massive, in-the-wild monocular videos offers a highly efficient data collection pathway for…
Towards Value-Constrained Credit Assignment in Fully Delegated AI Cooperatives
We propose a framework for reward allocation in fully delegated AI cooperatives where humans are represented by agents that contribute data…
Exposure Bias Can Alleviate Itself via Directional and Frequency Rectification in Flow Matching
Flow Matching (FM) has achieved remarkable generative performance, yet it suffers from exposure bias due to discrepancies between training…
Govern the Repository, Not the Agent: Measuring Ecosystem-Level Risk in AI-Native Software
Autonomous coding agents now open and merge pull requests in shared repositories at scale, and the field evaluates them the way it has alwa…
How Width and Data Shape Generalization Scaling Laws in Quadratic Neural Networks
Understanding how performance scales jointly with model size and data is a central problem in modern machine learning. Existing theoretical…
Learning Topology-Aware Representations via Test-Time Adaptation for Anomaly Segmentation
Test-time adaptation (TTA) has emerged as a promising paradigm for mitigating distribution shifts in deep models. However, existing TTA app…
Parameter Efficient Hybrid Transformer (PEHT) for Network Traffic Prediction via Dynamic Urban Congestion Integration
Accurate network traffic prediction is a critical element for efficient resource allocation in dynamic urban cellular networks. However, pr…
Towards Automating Scientific Review with Google's Paper Assistant Tool
Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical…
Agentic Hardware Design as Repository-Level Code Evolution
We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is c…
Which Nash Equilibrium? Solver-Dependent Selection on Zero-Sum Nash Polytopes
Many two-player zero-sum games admit not a unique Nash equilibrium but a convex set of them: a polytope of profiles that all share the mini…
DexCompose: Reusing Dexterous Policies for Multi-Task Manipulation with a Single Hand
Dexterous manipulation policies can solve individual skills, but composing them to perform multiple tasks with a single hand remains challe…
"Generate" the Future of Work through AI: Empirical Evidence from Online Labor Markets
Large Language Model (LLM)-based generative AI systems are general-purpose tools capable of augmenting or even automating a wide range of j…
Agentic Episodic Control
Reinforcement learning (RL) remains fundamentally limited by poor data efficiency and weak generalization. Prior episodic RL methods attemp…
Symmetry-Aware Transformer Training for Automated Planning
While transformers excel in many settings, their application in the field of automated planning is limited. Prior work like PlanGPT, a stat…
PreferThinker: Reasoning-based Personalized Image Preference Assessment
Personalized image preference assessment aims to evaluate an individual user's image preferences by relying only on a small set of referenc…
Chronic Kidney Disease Prognosis Prediction Using Transformer
Chronic Kidney Disease (CKD) affects nearly 10\% of the global population and often progresses to end-stage renal failure. Accurate prognos…
Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
Multimodal large language models (MLLMs) can inadvertently memorize privacy-sensitive information during training. While existing unlearnin…
SciFig: Towards Automating Editable Figure Generation for Scientific Papers
High-quality methodology figures are central to scientific communication, yet they remain difficult and time-consuming to create. Such figu…
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context
Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-…
GAIA: A Data Flywheel System for Training GUI Test-Time Scaling Critic Models
While Large Vision-Language Models (LVLMs) have significantly advanced GUI agents' capabilities in parsing textual instructions, interpreti…
Just Ask: Curious Code Agents Reveal System Prompts in Frontier LLMs
Autonomous code agents built on large language models are reshaping software and AI development through tool use, long-horizon reasoning, a…
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of gener…
CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching
As large language models (LLMs) witness increasing deployment in complex, high-stakes decision-making scenarios, it becomes imperative to g…
Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning
Offline learning of strategies takes data efficiency to its extreme by restricting algorithms to a fixed dataset of state-action trajectori…
SEA-TS: Self-Evolving Agent for Autonomous Code Generation of Time Series Forecasting Algorithms
Accurate time series forecasting underpins decision-making in many domains, yetconventional ML development often faces data scarcity, distr…
Algorithms for Deciding the Safety of States in Fully Observable Non-deterministic Problems: Technical Report
Learned action policies are increasingly popular in sequential decision-making, but suffer from a lack of safety guarantees. Recent work in…
AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization
Multi-agent reasoning has shown promise for improving the problem-solving ability of large language models by allowing multiple agents to e…
Distilling Answer-Set Programming Rules from LLMs for Neurosymbolic Visual Question Answering
Visual Question Answering (VQA) is the task of answering questions about images, requiring the integration of multimodal input and reasonin…
Edit-R2: Context-Aware Reinforcement Learning for Multi-Turn Image Editing
Text-guided image editing has advanced rapidly with diffusion models and unified multimodal foundation models. However, most existing metho…
RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastruc…
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models
Many efforts to ensure frontier AI models are safe rely on monitoring their chain-of-thought (CoT) reasoning. If models become able to perf…
Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters
Configuring an advanced scientific simulator, translating a modeling goal into a valid, runnable input deck, is a persistent bottleneck tha…
Artificial Intelligence Index Report 2026
Welcome to the ninth edition of the AI Index report. As AI continues to advance rapidly, the question becomes whether the systems built aro…
The Shift Toward Open and Reproducible AI Research
The reproducibility crisis has directed the AI research community toward improving documentation practices. Several studies have identified…
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
AI agents are moving from advisors to actors, booking travel, planning menus, and running procurement on behalf of users. Existing benchmar…
Agent-as-a-Router: Agentic Model Routing for Coding Tasks
Real-world users typically have access to multiple Large Language Models (LLMs) from different providers, and these LLMs often excel at dis…
POTracker: Optimizing Large Language Models for Standard-Compliant Power Outage Report Generation
Recent large language models (LLMs) are good at general text generation, but it is still hard to use them for domain-specific data generati…
AI Snitches Get Glitches: Towards Evading Agentic Surveillance
To better assist users with completing challenging tasks, AI agents mediate communications, access data, and interact with different APIs.…
What the LLM Should Not Say: Boundary-Aware Context Grounding for A Seven-Channel EEG Agent
Large language models (LLMs) can make scientific software easier to use. However, a general model does not automatically know which measure…
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems
Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this tra…
A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models
Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is re…
Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds
As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting…
iCost: A Novel Instance-Complexity-Based Cost-Sensitive Learning Framework
Class imbalance poses a significant challenge in classification tasks, often causing standard learning algorithms to become biased toward t…
Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content. We i…
Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning
We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient…
The Minimal Search Space for Conditional Causal Bandits
Causal knowledge can be used to support decision-making problems. This has been recognized in the causal bandits literature, where a causal…
ReFreeKV: Towards Threshold-Free KV Cache Compression
To reduce memory consumption during LLM inference, a handful of methods have been proposed for KV cache pruning. While these techniques can…
DMind Benchmark: Toward a Holistic Assessment of LLM Capabilities across the Web3 Domain
The Web3 ecosystem, underpinned by cryptographic primitives and decentralized consensus, represents a high-stakes environment where softwar…
Seven Security Challenges That Must be Solved in Cross-domain Multi-agent LLM Systems
Large language models (LLMs) are rapidly evolving into autonomous agents that cooperate across organizational boundaries, enabling joint di…
PRISON: Unmasking the Criminal Potential of Large Language Models
As large language models (LLMs) advance, concerns about their misconduct in complex social contexts intensify. Existing research overlooked…
SIDA: Synthetic Image Driven Zero-shot Domain Adaptation
Zero-shot domain adaptation is a method for adapting a model to a target domain without utilizing target domain image data. To enable adapt…
Calibrating Biophysical Models for Grape Phenology Prediction via Multi-Task Learning
Accurate prediction of grape phenology is essential for timely vineyard management decisions, such as scheduling irrigation and fertilizati…
LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link Recovery
Issue-to-commit link recovery in software repositories is fundamental to software traceability and project management, yet it remains a cha…
SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation in MRI Volumes
Liver cirrhosis plays a critical role in the prognosis of chronic liver disease. Early detection and timely intervention are essential for…
Freshness and the Limits of Heuristic Trend Detection in Temporal RAG
We present a lightweight, model-agnostic temporal layer for RAG and use cybersecurity data to separate two problems that are usually confla…
Unbiased Binning for Fairness-aware Attribute Representation
Discretizing raw features into bucketized attribute representations is a popular step before sharing a dataset. It is, however, evident tha…
Ranking Before Serving: Low-Latency LLM Serving via Pairwise Learning-to-Rank
Efficient scheduling of large language model (LLM) inference tasks is critical for achieving low latency and high throughput, a challenge t…
Deep Neural Networks Inspired by Differential Equations
Deep learning has become a pivotal technology in fields such as computer vision, scientific computing, and dynamical systems, significantly…
MetaBreak: Jailbreaking Online LLM Services via Special Token Manipulation
Unlike regular tokens derived from existing text corpora, special tokens are artificially created to annotate structured conversations duri…
A Primer on SO(3) Action Representations in Deep Reinforcement Learning
Many robotic control tasks require policies to act on orientations, yet the geometry of SO(3) makes this nontrivial. Because SO(3) admits n…
LieSolver: PDE-Constrained Learning for IBVPs via Lie Symmetries
Initial-boundary value problems (IBVPs) provide the essential framework for modelling a wide range of phenomena in physics and engineering.…
Hybrid Fact-Checking that Integrates Knowledge Graphs, Large Language Models, and Search-Based Retrieval Agents Improves Interpretable Claim Verification
Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information. At the same tim…
Hybrid coupling with operator inference and the overlapping Schwarz alternating method
This paper presents a novel hybrid approach for coupling subdomain-local non-intrusive Operator Inference (OpInf) reduced order models (ROM…
Trust Region Masking for Long-Horizon LLM Reinforcement Learning
Policy gradient methods for Large Language Models optimize a policy $\pi_\theta$ via a surrogate objective computed from samples of a rollo…
Pixelwise Uncertainty Quantification of Accelerated MRI Reconstruction
Parallel imaging techniques reduce magnetic resonance imaging (MRI) scan time but image quality degrades as the acceleration factor increas…
Psychometric Comparability of LLM-Based Digital Twins
Large language models (LLMs) act as digital twins for human respondents, yet their psychometric comparability remains uncertain. We propose…
DDSA: Dual-Domain Strategic Attack for Spatial-Temporal Efficiency in Adversarial Robustness Testing
Image transmission and processing systems in resource-critical applications face significant challenges from adversarial perturbations that…
Reasoning-Enhanced Rare-Event Prediction with Balanced Outcome Correction
Rare-event prediction is critical in domains such as healthcare, finance, reliability engineering, customer support, aviation safety, where…
Dual-Prototype Disentanglement: A Context-Aware Enhancement Framework for Time Series Forecasting
Time series forecasting has witnessed significant progress with deep learning. While prevailing approaches enhance forecasting performance…
Robustness of Constraint Automata for Description Logics with Concrete Domains
Decidability or complexity issues about the consistency problem for description logics with concrete domains have already been analysed wit…
Gated Relational Alignment via Confidence-based Distillation for Efficient VLMs
Vision-Language Models (VLMs) achieve strong multimodal performance but are costly to deploy, and post-training quantization often causes s…
Spectral Text Fusion: A Frequency-Aware Approach to Multimodal Time-Series Forecasting
Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual sign…
When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
Recent advances in large image editing models have shifted the paradigm from text-driven instructions to vision-prompt editing, where user…
Event-Grounded Question Answering over Long Audio via Structured Retrieval
Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding. Current large audio-lang…
Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training
As artificial intelligence (AI)-generated content proliferates, models are increasingly trained on their own outputs, risking progressive d…
An Interpretable, Controllable Time-Varying IIR Denoiser for On-Device Assistive Hearing
We present TVF (Time-Varying Filtering), an interpretable, low-latency speech enhancement model for real-time, on-device assistive hearing.…
MPFlow: Multi-modal Posterior-Guided Flow Matching for Zero-Shot MRI Reconstruction
Zero-shot MRI reconstruction relies on generative priors, but single-modality unconditional priors produce hallucinations under severe ill-…
Measuring the Redundancy of Decoder Layers in SpeechLLMs
Speech Large Language Models route speech encoder representations into an LLM decoder that typically accounts for over 90% of total paramet…
EXPLORE-Bench: Egocentric Scene Prediction with Long-Horizon Reasoning
Multimodal large language models (MLLMs) are increasingly considered as a foundation for embodied agents, yet it remains unclear whether th…
HiMu: Hierarchical Multimodal Frame Selection for Long Video Question Answering
Long-form video question answering requires reasoning over extended temporal contexts, making frame selection a critical bottleneck for mul…
IWP: Token Pruning as Implicit Weight Pruning in Large Vision Language Models
Large Vision Language Models show impressive performance across image and video understanding tasks, yet their computational cost grows rap…
Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior
Understanding student engagement usually requires time-consuming manual observation or invasive recording that raises privacy concerns. We…
From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales
Hallucinations in large ASR models present a critical safety risk. In this work, we propose the \textit{Spectral Sensitivity Theorem}, whic…
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
OpenClaw-style personal assistants extend LLM agents from isolated tool use to open-ended, stateful, and personalized software environments…
GenMatter: Perceiving Physical Objects with Generative Matter Models
Human visual perception offers valuable insights for understanding computational principles of motion-based scene interpretation. Humans ro…
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making…
Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking
Reinforcement learning from human feedback (RLHF) systems face a compounding alignment challenge: not only are learned reward models uncert…
Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout
Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world…
ELF: Embedded Language Flows
Diffusion and flow-based models have become the de facto approaches for generating continuous data, e.g., in domains such as images and vid…
EAGT: Echocardiography Augmentation for Generalisability and Transferability
Deep learning models for echocardiography segmentation often struggle to generalise across institutions, scanners, and patient populations,…
FormalASR: End-to-End Spoken Chinese to Formal Text
Automatic speech recognition (ASR) systems are typically optimized for verbatim transcription, which preserves disfluencies, filler words,…
Do Vision Models Truly Forget? New Findings from Representation-Level Certification of Visual Unlearning in Vertical Federated Learning
Machine unlearning in Vertical Federated Learning (VFL) has attracted growing interest, yet existing methods certify forgetting solely usin…
Coloring the Noise: Adversarial Sobolev Alignment for Faithful Image Super Resolution
Generative priors in Image Super-Resolution (SR) often compromise faithful restoration, we attribute this limitation to a fundamental spect…
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
LLM training increasingly relies on teacher-generated supervision, from synthetic responses to reasoning traces and tool-use demonstrations…
Energy-Structured Low-Rank Adaptation for Continual Learning
While orthogonal subspace methods try to mitigate task interference in Continual Learning (CL), they often suffer from energy diffusion acr…
When are LLMs Sufficient Policy Optimizers for Sequential RL Tasks?
We study when large language models (LLMs) can serve as effective black-box policy optimizers for reinforcement learning (RL) tasks, i.e.,…
$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems
As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. C…
Adaptive Turn-Taking for Real-time Multi-Party Voice Agents
Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor co…
Lost at the End: Primacy Bias in Multimodal Retrieval-Augmented Question Answering
Knowledge-based visual question answering (KB-VQA) lets vision-language systems answer questions that exceed their parametric knowledge by…
Hierarchical Control in Multi-Agent Games: LLM-based Planning and RL Execution
Reinforcement learning (RL) has achieved strong performance in sequential decision-making, yet scaling to complex multi-agent environments…
Contagion Networks: Evaluator Preference Propagation in Multi-Agent LLM Systems
When large language models serve as evaluators in multi-agent systems, their strategy preferences -- whether induced by explicit prompts or…
An Empirical Study of OpenPangu Quantization on Ascend NPUs
OpenPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive pos…
When Is an LLM Worth It for Hyperparameter Optimization? A Budget-Matched Study on Tabular Data Finds the Warm-Start Is a Default Configuration, Not the Model
Large language models (LLMs) have been proposed as hyperparameter-optimization (HPO) advisors that "warm-start" search from prior knowledge…
On the Position Bias of On-Policy Distillation
On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision fro…
FP8 is All You Need (Part 2): Efficient Ozaki-Bailey Style FFT Through Tensor-core Garner Reformulation and Kulisch Escape Route
NVIDIA's Blackwell Ultra (B300) cuts FP64 vector throughput to ~1.3 TFLOPS per GPU, roughly 30x below B200 and well below the level at whic…
ZONOS2 Technical Report
We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve up…
CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation
Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU…
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety
General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimod…
Pulmonary Embolism Risk Stratification from CTPA and Medical Records: Vascular Graphs Are Not All You Need
Risk stratification for pulmonary embolism (PE) is critical for clinical decision-making. Stratification guidelines are based on patient me…
Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignment
A central goal of safety research is determining whether a model is misaligned. Prior work has largely focused on detecting concerning beha…
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training
Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but th…
Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare
Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then as…
DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models
The Abstraction and Reasoning Corpus (ARC) contains tasks that require summarizing patterns from limited grid samples and predicting output…
ロボットの模倣学習を60時間→4.8時間に AWSのGPUでフィジカルAI開発を加速 ファナック
ファナックは「AWS Summit Japan 2026」の基調講演で、ロボットに動作を教える「模倣学習」の時間を、AWSのGPU活用で60時間から4.8時間へ短縮したと示した。
カインズが画像AIで売上UP模索、店頭でのインテリア“試着”をテスト 立ちはだかる「正確性と効率」の壁
カインズが画像生成AIを活用し、部屋のインテリアを疑似的に置き換えられる店頭サイネージ「CAINZ Fitting Room」を開発。その効果や利便性を検証している。
製造現場のトラブル解消を「AI工場長」が支援? 「エージェント型工場」とは
AccentureとAvanadeはMicrosoftと協働し、製造業向け工場インテリジェンスシステム「エージェント型工場」を開発したと発表した。
AIは設計者を置き換えるのか Autodesk幹部に聞くCADと設計データの未来
AIの活用が設計/製造の現場にも広がる中、CADの操作や設計者の役割はどう変わるのか。米Autodesk 製品開発/製造ソリューション担当エグゼクティブバイスプレジデントのジェフ・キンダー氏に、AIが設計業務にもたらす変化、AI時代に求められる設計データの在り方、そして同社が描…
解剖・孫正義氏の「ガチョウ論」 「ソフトバンクG株価が低過ぎ」主張を信じてよいのか
孫正義氏が、ソフトバンクGの株価に不満をにじませた。孫氏が「本当の企業価値」として示す「時価純資産」の論理と危うさを解説する。
Ford rehires ‘gray beard’ engineers after AI falls short
"Mistakenly we thought that by just introducing artificial intelligence ... that would produce a high-quality product.”
HP Inc. launches Frontier strategic partnership with OpenAI
HP Inc. scales its OpenAI Frontier partnership to deploy AI across customer experiences, software development, and enterprise operations.
Why Wall Street thinks US memory maker Micron is the next Nvidia
Eager to find more public AI-related companies that may do as well as Nvidia, Wall Street investors think they've found a winner with Micro…