週次AIニュース 2026-W29
対象期間: 2026-07-13 〜 2026-07-19(1482 件)
トピックの推移
トピック別件数
- LLM/生成AI 574件
- 研究/論文 557件
- エージェント 364件
- 画像/動画生成 201件
- ロボティクス 106件
- ビジネス/資金調達 96件
- その他 41件
- ハードウェア/半導体 37件
- 規制/政策 11件
今週のハイライト(上位 10 件)
A scorecard for the AI age
Sarah Friar, CFO of OpenAI, introduces a practical AI scorecard to measure ROI through useful work, cost per successful task, dependability…
Why teens deserve access to safe AI
Learn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partners…
The US is advancing AI safety through state and federal action
OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.
GPT-Red: Unlocking Self-Improvement for Robustness
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
How to manage AI investments in the agentic era
Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling h…
Our approach to bioresilience
Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.
Kimi: Threat or menace?
Chinese company Moonshot AI released a new version of its Kimi model this week, prompting concern about "full AI communism."
「Claude Fable 5」サブスクに統合 Max・Team Premiumプラン対象
米Anthropicは、最上位AIモデル「Claude Fable 5」を7月20日から有料プラン「Max」「Team Premium」の標準機能にする。利用上限の50%まで追加費用なしで利用でき、「Pro」「Team Standard」は「使用クレジット」での利用となり、10…
Vertu wants executives to pay $6,880 for an AI agent — here’s how it actually performs
From AI workflows to battery life and security, here's what it's really like to live with Vertu's luxury foldable every day.
Databricks hits $188B valuation, extending its run as AI’s favorite second act
Databricks has remade its image into an AI company and has published research on the cost savings of open weight AI models for coding.
全件(日付別)
2026-07-19(1件)
Kimi: Threat or menace?
Chinese company Moonshot AI released a new version of its Kimi model this week, prompting concern about "full AI communism."
2026-07-18(9件)
Neil Rimer thinks the AI money is coming back out
Neil Rimer, the venture capitalist who co-founded Index Ventures, predicts the historic wealth AI is generating in Silicon Valley will have…
「Claude Fable 5」サブスクに統合 Max・Team Premiumプラン対象
米Anthropicは、最上位AIモデル「Claude Fable 5」を7月20日から有料プラン「Max」「Team Premium」の標準機能にする。利用上限の50%まで追加費用なしで利用でき、「Pro」「Team Standard」は「使用クレジット」での利用となり、10…
Vertu wants executives to pay $6,880 for an AI agent — here’s how it actually performs
From AI workflows to battery life and security, here's what it's really like to live with Vertu's luxury foldable every day.
Databricks hits $188B valuation, extending its run as AI’s favorite second act
Databricks has remade its image into an AI company and has published research on the cost savings of open weight AI models for coding.
The Zoom hack that says, ‘Don’t record me’
If every meeting, watercooler conversation, and date gets transcribed and summarized, who's actually reading any of it?
Agility Robotics plants its flag in Tesla’s backyard
Agility is opening a new training center for its Digit robots in Fremont, California.
AI-driven memory crunch jolts India’s smartphone market
India's smartphone slowdown highlights how the AI boom is reshaping consumer electronics, from pricing and demand to corporate strategy.
How Apple’s big lawsuit could disrupt OpenAI’s IPO plans
Apple filed a trade secrets lawsuit against OpenAI last Friday, and it’s not messing around. The complaint alleges a pattern of misconduct…
Patreon stops asking AI bots not to scrape — and starts blocking them
Patreon is strengthening its defenses against AI scraping by working with Cloudflare to block bots that train AI models on creators’ conten…
2026-07-17(292件)
Apple’s lawsuit couldn’t come at a worse time for OpenAI
Apple filed a trade secrets lawsuit against OpenAI last Friday, and it’s not messing around. The complaint alleges a pattern of misconduct…
Why the first GPU financiers are turning to inference chips in a $400 million deal
A $400 million chip-backed loan points to the next wave of AI infrastructure deals.
A scorecard for the AI age
Sarah Friar, CFO of OpenAI, introduces a practical AI scorecard to measure ROI through useful work, cost per successful task, dependability…
「AIエージェントの約半数が"戦力外"になる」 なぜ企業は使いこなせない?
Gartnerは、2027年までに企業の40%が自律型AIエージェントを格下げまたは廃止すると予測している。そうした悲観的な見方の背景にはどのような問題があるのか。
Intelligent Three Level Learning Architecture for Autonomous UAV Swarms in Search and Rescue
This paper presents a novel three level hierarchical learning architecture for autonomous UAV swarms performing search and rescue operation…
HG-RAG: Hierarchy-Guided Retrieval-Augmented Generation for Structured Knowledge Graphs
Retrieval Augmented Generation (RAG) has proven to be a widely successful process at improving the quality of outputs from a Large Language…
IMEX Interaction-Based Model Explanation
In predictive modeling, the ability to explain why a model produces a given target prediction has become increasingly important [5, 10]. Bl…
RegNetAgents: A Multi-Agent Framework for Cross-Network Regulatory Driver Identification in Cancer Genomics
We introduce RegNetAgents, an AI-oriented multi-agent framework for structured, query-driven regulatory candidate identification across het…
DialogueVPR: Towards Conversational Visual Place Recognition
Inspired by how humans communicate spatial information, language-guided geo-localization has gained significant traction for its intuitive…
Interpretable Language Model for Closed-Loop Type 1 Diabetes Control
Type 1 Diabetes (T1D) is a chronic, life-threatening autoimmune condition characterized by the complete destruction of insulin-producing pa…
Human AI Construction of Bayesian Networks for Operational Decision Support -- A Virtual Survey Approach
Bayesian Belief Networks (BBNs) are powerful tools for decision-making under uncertainty. However, building their structures and estimating…
Capability from Access Structure, Not Scale: Lower Bounds and Pre-Registered Tests for Hybrid Sequence Models
The Platonic Representation Hypothesis (PRH) holds that as models scale, representations of heterogeneous networks converge toward a shared…
ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability
Tool-augmented large language model agents excel at long-horizon tasks, yet they are typically post-trained on fixed toolsets. When tasks d…
Enhancing Small Language Models Reasoning through Knowledge Graph Grounding
Although large language models (LLMs) have set benchmarks for zero-shot reasoning, their deployment remains cost-prohibitive and environmen…
Orchestrating Power Grid Studies with Multi-Agent AI and MCP Servers
This position paper explores how Agentic AI and Model Context Protocol (MCP) can support power-grid studies in a Transmission System Operat…
MemoHarness: Agent Harnesses That Learn from Experience
An agent harness is the external control layer that turns a base LLM into an executable agent by managing context, tools, orchestration, me…
When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models
Large language models can synthesize a game's rules as executable code - a Code World Model (CWM) - which a classical planner then searches…
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex scientific tasks, yet existing automa…
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-R…
How Artificial Intelligence LLM Engines Shape the Global Conflict Information Environment
Artificial Intelligence (AI) answer engines now field a growing share of the questions that analysts, scholars, and the public ask about is…
Align AI to Dynamic Human-AI Workflows
Current alignment approaches typically focus on emulating human behavior using static representations of human preferences, failing to capt…
The Steering Budget: Examples beat Knobs
Generative models are steered with knobs -- prompts, guidance scales, property tags. Turn one as hard as you like and, past a point, it sto…
Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnera…
AI Agents Do Not Fail Alone:The Context Fails First
Context engineering has become central to building reliable AI agents, yet it remains largely unmeasured. Agents do not fail in isolation:…
Measuring How Students Rely on Generative AI in Academic Writing: Development and Multi-Source Validation of the Generative AI Reliance Types Scale (GenAI-RTS)
As generative AI (GenAI) becomes increasingly embedded in undergraduate academic writing, how students rely on these tools, rather than sim…
Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions
In this paper, we study the connection between an LLM's output distribution and the data used to train it. Specifically, we study the degre…
Traccia: An OpenTelemetry-Based Governance Platform for AI Systems
The rapid development of Large Language Models (LLMs) and Artificial Intelligent (AI) powered autonomous agents has fundamentally changed t…
CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents
Data science tasks span from closed-ended information extraction to open-ended analysis, presenting significant challenges for automation.…
Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving
Validating autonomous driving systems requires diverse, regulation-compliant test scenarios. In simulation-based testing, scenarios are def…
A Comparative Analysis of Machine Learning Models for Long and Short-Term Forecasting of the Egyptian Stock Market: A Focus on EGX30
This study concentrates on predicting stock prices in the Egyptian market, focusing on the EGX30, an influential financial hub in the Middl…
CatalogAgent: A Supervisor-mediated Self-Learning System Enabling Context Engineering for GenAI Models
Product catalogs are the backbone of e-commerce sites, yet a large number of structured attributes (SAs) -- such as material, color, and sh…
Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration
Evaluations of language-model honesty read the model's verdicts as evidence about the model. We test the instrument instead. We built a tex…
Reward-Free Evolving Agents via Pairwise Validator
A self-evolving agentic loop repeatedly proposes a tweaked version of an agent (its prompt template or program) and accepts or rejects the…
CausalGraphX: A Counterfactual Graph Neural Network Framework for Explainable Systemic Risk Assessment
The interconnected nature of global financial systems makes them vulnerable to systemic risks, where the failure of a few institutions can…
Per-Token Fixed-Point Convergence in Depth-Recurrent Transformers
A depth-recurrent transformer applies a weight-tied core a variable number of times, and prior work has shown that training with a randomiz…
Tactile: Giving Computer-Using Agents Hands and Feet
Computer-use agents are becoming capable software operators, but their interface to desktop applications is still often a brittle motor lay…
Step-Level Preference Learning for Generative Agents in Social Simulations
Large language model (LLM)-based generative agents simulate human behavior through long-horizon decision-making processes that comprise int…
SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation
Complex knowledge base question answering (KBQA) is commonly approached through either information retrieval over a question-specific subgr…
Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions
Multi-modal Large Language Models (MLLMs) have made substantial advances on benchmarks, yet their real-world effectiveness remains uncertai…
VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence
Industrial time series serve as the foundation for Prognostics and Health Management (PHM) to ensure the reliability and safety of industri…
RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning
Multi-step retrosynthesis planning seeks to decompose a target molecule into commercially available building blocks through a sequence of f…
WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays
This study presents WrAFT, a Writing Assessment and Feedback Tool, that delivers both accurate and reliable scores and effective comprehens…
Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent
Existing GPU kernel generation benchmarks draw problems from synthetic or curated sources that diverge from deployed workloads. We present…
Towards an Intention Abstraction Layer for Autonomous Industrial Systems
Modern industrial environments increasingly run many autonomous subsystems at once - schedulers, energy managers, vehicle fleets - each pur…
Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation
Diffusion Multimodal Large Language Models (DMLLMs) are highly effective for multimodal reasoning, yet their inference efficiency is signif…
Democratizing Agent Deployment Safety: A Structural Monitoring Approach
AI software development agents are increasingly capable of modifying infrastructure and security critical systems, creating risks where an…
Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents
Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-serv…
Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments
Collective problem solving often requires that group members consider the tradeoff between exploitation of known solutions and exploration…
Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology
Recent advances in large language models (LLMs) and their extension to vision-language models (VLMs) have made it easier to combine text an…
MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research
Existing LLM-based theorem provers have achieved impressive results on formal mathematics benchmarks, yet they remain confined to acting as…
SportD: Can VLMs Physically Strategize?
Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use informa…
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models
Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In thi…
Analytic Abduction: Causal Decomposition and Governed Commitment for Human--AI Coordination
Abductive reasoning operates in two directions. The synthetic mode builds explanations from available hypotheses; the analytic mode, conver…
MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers
As Model Context Protocol (MCP) servers emerge as the core infrastructure for connecting LLMs with external tools, existing benchmarks leve…
TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning
While Multimodal Large Language Models (MLLMs) excel in general tasks, rigorous scientific reasoning remains challenging due to the limitat…
SmartRAG: Native Graph-Based RAG for Mobile Device
Deploying large language models (LLMs) as personal assistants on mobile devices demands privacy, low latency, and offline availability, yet…
Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications
Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or…
Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment
Efficient multimodal document question answering with explicit visual grounding, locating the precise document region that supports each an…
InCarEmo: A Multimodal Dataset for In-Cabin Emotion Recognition and Driver State Monitoring
Understanding driver emotion and state is critical for the next generation of intelligent in-cabin systems that ensure safety and enhance h…
AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery
This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Stor…
Global Index on Responsible AI: 2026 Report
Grounded in human rights-based frameworks such as the UNESCO Recommendation on the Ethics of AI, the Global Index on Responsible AI (GIRAI)…
Transcoders for Investigating Deception in Language Models
Transcoders have recently emerged as a promising approach for mechanistic interpretability (MI), enabling circuit-level analysis of model b…
CrimeNER Demo: Named-Entity Recognition in the Crime Domain
We present CrimeNER Demo, an AI-powered platform that enables us to extract general crime-related information from documents and classify t…
Reachability-Aware Pretraining for Efficient Target-Oriented Path Exploration in Temporal Knowledge Graph Reasoning
Temporal Knowledge Graph (TKG) reasoning under the extrapolation setting focuses on forecasting future time-stamped events (facts) from his…
Proof-or-Stop: Don't Trust the Agent, Trust the Evidence -- Loop Engineering for Verifiable Evidence-Gated Lifecycle Control
Autonomous coding agents increasingly execute multi-step software work, but lifecycle states such as reviewed, tested, DONE, and ready-to-m…
Contextualized Early Detection of Online Firestorms: A Sequential LLM-Based Approach
Online firestorms are rapid collective escalations of highly negative user-generated content and may cause substantial reputational and eco…
Explaining Process Control Optimisation Recommendations via GradientSHAP and Implicit Differentiation
Automated optimisation is increasingly adopted in industrial processes, yet a trust gap persists between engineers who design these algorit…
CFM-Bench: A Unified Multi-Domain, Multi-Task Benchmark for Channel Foundation Models
Channel foundation models (CFMs) are developing rapidly, with recent studies reporting benefits from pretraining across downstream wireless…
Demographically-Conditioned Synthetic Medical Images for Bias Mitigation and Bias Detection in Disease Classifiers
Per-subgroup fairness audits of medical image classifiers face a sample-size problem: minority subgroups in held-out test sets have so few…
Moral Attitudes of Sentient ASI towards Humanity and Implications for AGI Development
This paper suggests the adoption of a novel inversion in AI ethics: instead of asking how humans should treat artificial superintelligence…
SMC-ES: Automated synthesis of formally verified control policies
The deployment of autonomous cyber-physical systems in safety-critical environments requires closed-loop control strategies (i.e., policies…
Man, Machine, and Masterpiece: Artistic Ownership in the AI Era
The integration of AI-driven systems in creative work has sparked debates among artists and legal communities about notions of ownership. Y…
BrainPilot: Automating Brain Discovery with Agentic Research
Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single resear…
Long-Context Fine-Tuning with Limited VRAM
Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention still makes long training sequences expensive. We c…
Concept-Guided Spatial Regularization for World Models in Atari Pong
World models are usually evaluated as components of model-based reinforcement learning (MBRL) systems, while the world models themselves ar…
The Industrialization of Research ; On AI-Driven Science and Its Consequences
Artificial intelligence is transforming scientific research - not merely as a more powerful instrument, but as an autonomous participant in…
MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary…
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-ce…
Can We Trust Item Response Theory for AI Evaluation?
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities,…
Plover: Steering GUI Agents through Plan-Centric Interaction
Graphical user interface (GUI) automation remains challenging in real-world environments, where dynamic layouts, unexpected dialogs, and ev…
Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation
Annotation quality is a major bottleneck in building reliable and explainable artificial intelligence (XAI) systems for mental health resea…
When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can be…
AutoSynthesis: An agentic system for automated meta-analysis
Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, and policy. Yet, quant…
teLLMe Why (Ain't Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data
Traffic agencies now have access to large volumes of video-derived data for studying safety and congestion. Most of these data are observat…
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as…
Pretraining Data Can Be Poisoned through Computational Propaganda
Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretra…
All Polarized but Still Different: a Multi-factorial Metric to Discriminate between Polarization Behaviors on Social Media
Online polarization has attracted the attention of researchers for many years. Its effects on society are a cause for concern, and the desi…
Fast-Fading Channel and Power Optimization of the Magnetic Inductive Cellular Network
The cellular network of magnetic Induction (MI) communication holds promise in long-distance underground environments. In the traditional M…
Falsifiable Release Gates for Self-Improving Systems
Safety claims on self-improving agent runtimes are almost always self-graded: a policy file, a guardrail, or a README commitment. We descri…
Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs
Deploying Vision-Language Models (VLMs) in real-world settings requires not only strong visual reasoning but also stability under sustained…
LBA: Textual Hard-Label Adversarial Attack under Low Query Budgets
Generating high-quality adversarial texts with low query budgets remains a challenging problem in the hard-label scenario. Most existing ap…
Automatically Evolving Prompt Guidelines for Task-Specific Optimization
For Large Language Models to reliably answer user queries, users must clearly specify requirements, context, and constraints. In practice,…
Token Time Continuous Diffusion for Language Modeling
In this paper we introduce token time continuous diffusion (TTCD), a new diffusion language model which (a) operates in continuous space, d…
Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs
The inference efficiency of diffusion large language models (dLLMs) is constrained by two challenges: bidirectional attention precludes eff…
Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility
This paper introduces tool efficiency, a new quantitative metric to evaluate the rate of useful tool calls in an LLM agent trajectory. To e…
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain…
MAPS: Modeling Co-Existing Subjective Perspectives and Shared Meaning in Multi-Agent Cognitive Dialogue
Human dialogue involves more than exchanging information; it also expresses beliefs, emotions, and subjective cognitive styles. Yet current…
Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect
Can small language models detect and report on perturbations their own internal activations? We investigate this question through the lens…
Information-Theoretic Limits of Reliability and Scaling in Language Models
Large language models (LLMs) are evaluated as though perfect reliability is achievable for any task given sufficient scale. We show this as…
T5-CSBoost: Adversarial Perturbation Resistant LLM Fingerprinting
While many AI-generated text (AIGT) detectors achieve strong performance on clean inputs, their accuracy degrades significantly under light…
CoEvoT: Co-Evolving Chain-of-Thought Prompting for Graph-LLM Reasoning
Graph learning under distribution shift presents a persistent challenge, where models adapt to new graphs with limited or even no supervisi…
ReportMedSAM: Guiding Segmentation Through Radiology Reports
Free-form radiology reports contain rich clinical descriptions, yet converting them for reliable segmentation remains challenging due to th…
Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning
Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challengin…
Position: Explainability Research Must Prioritize Foundations over Ad-hoc Methods
Despite the proliferation of Explainable AI (XAI) techniques -- from feature attributions to sparse autoencoders -- explanations rarely inf…
Explainable Geospatial AI for Satellite Ground Station Siting Using LiDAR-Derived Terrain Intelligence
Representative clutter height (RCH) is a key parameter in radio propagation and interference analysis because it captures the dominant heig…
Volition Elicitation: Operational Semantics for People and Their Machines
The most prevalent distributed systems today include people and their personal machines (smartphones). In such systems, computations are dr…
The Planar Case of Thomas Positive Circuits Conjecture
The notion of circuit refers to a cyclic oriented influence between the elements of a dynamical system. There are two classes of circuit: p…
Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak
Aligned language models refuse harmful requests, but a one-line prefill ("Sure, here is") strips the refusal. We ask where and how it fails…
"Trust Junk" Leads to Unjustified Support for Highly Discriminatory Predictive Models
The persuasive power of data visualizations can go awry: for instance, in an explainable AI (XAI) context, visualizations can produce over-…
Certified Domain Consistency for Multi-Domain Retrieval: Label-Free Per-Domain Contamination Control with Conformal Risk Guarantees
Retrieval over corpora that mix several domains often returns relevant but wrong-domain evidence that ranking metrics miss and that conform…
Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation
While Large Language Models (LLMs) have demonstrated significant capability in software code generation, their application to analog Electr…
Structured Feedback Improves Repair in an LLM Agent Loop
LLM agents often retry after external validation rejects a candidate, but the interface between validation and the next model call remains…
The Cost and Network Limits of Space-Based AI Compute
This paper evaluates whether large-scale AI data centers deployed in low-Earth orbit (LEO) could become a cost-effective alternative to ter…
RENEW: Towards Learning World Models and Repairing Model Exploitation from Preferences
World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dat…
Semantic Audio-driven Understanding for Dynamic Humanoid Whole Body Control
Recent advances in humanoid robotics and reinforcement learning have enabled the acquisition of highly expressive whole-body motion policie…
Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape
Feedback-driven loops support iterative improvement in large language models, reinforcement learning, and autonomous discovery, yet their g…
NexForge: Scaling Executable Agent Tasks via Requirement-First Synthesis
Scaling executable agent training data is bottlenecked by substrate-first methods that tie task generation to predefined tools, repositorie…
Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation
3D simulation platforms are critical for autonomous driving because they enable end-to-end policy evaluation, thereby reducing development…
SeeSE3: Emergence of 3D Space in Vision Features
In this paper, we ask whether vision foundation models construct representations that reflect the intrinsic properties of 3D Euclidean spac…
LIGO-PINN: Learned Initialization via Gated Optimization to Alleviate Convergence Failures in Physics Informed Neural Networks
Physics-informed neural networks (PINNs) have had a broad research impact in modeling domains governed by partial differential equations (P…
Never Too Late for Force: Accelerating VLA Post-Training with Reactive Force Injection
Pretrained vision-language-action (VLA) policies provide strong language-conditioned manipulation knowledge, but they remain largely vision…
MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning
Long-horizon robot planning requires more than predicting what actions will do next; it also requires memory of the embodied experience tha…
Local Additive Feature Attribution: A Mathematical Taxonomy and Reporting Checklist
Feature-attribution methods are central to explainable artificial intelligence. Their assumptions are expressed in several mathematical lan…
ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict? We investigate this q…
Assessing AI in Introductory Physics Problem Solving
Reasoning or inference-scaling models are the new generation of Large Language Models (LLMs) capable of complex problem solving. To investi…
Towards a Unified Multidimensional Explainability Metric: Evaluating Trustworthiness in AI Models
In this paper, we present a comprehensive framework for assessing the explainability of various XAI methods, such as LIME and SHAP, across…
Accounting for Hysteresis and Eddy Currents in Finite Element Simulations of Ferromagnetic Laminated Cores using a Recurrent Neural Network
Incorporating hysteresis and eddy currents into finite element simulations of laminated-core electrical machines is computationally challen…
PReM: Learning What to Preserve and When to Refresh for Context Compression
Efficient long-context inference is not only about reducing memory cost, but also about keeping useful contextual evidence accessible as ge…
ViPSAM: Visual Prompting Medical Image Segmentation Using Segment Anything Model
In proton therapy planning, respiratory-gated non-contrast CT (NCCT) is commonly used for lesion segmentation; however, accurate delineatio…
Copy-on-Write Scoring: Application-Specific Agent Evaluations
Trustworthy deployment of LLM-based agents in software systems requires evaluating how they perform on application-specific workflows, with…
Beyond scalar losses: calibrating segmentation models via gradient vector field surgery
Region-based loss functions, such as the Dice loss, have established themselves as the de facto standard for highly class- and region-imbal…
The Prover Is the Judge: Verified Security Software from AI Coding Agents in Ada/SPARK
AI coding agents produce code faster than humans can review it. In our approach, the prover is the judge of whether the code is correct. Un…
Beyond Visual Grasping: Benchmarking Complex Grasping from Detection to Execution
Robust robotic grasping remains a fundamental challenge for complex real-world applications. Recent advances in large-scale models demonstr…
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage:…
HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization
While Large Language Models (LLMs) excel in many general NLP tasks, their formal reasoning capabilities are often compromised by content ef…
Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI
As automated decision-making and data-driven technologies pervade society and are used to manage consequential outcomes, understanding the…
Why Git Is the Memory Solution for the Agentic Development Lifecycle
Coding agents now produce a growing share of a team's code, while the reasoning behind each change -- the alternatives weighed, the constra…
An offline approach to fNIRS-guided reinforcement learning for robot behavior
Human-in-the-loop Reinforcement Learning has become a popular approach to training, finetuning, and aligning robot behavior with user prefe…
Integration Matters: Rollout-Based Training for Constrained Diffusion Models
Constrained generative models aim to produce samples that satisfy complex feasibility constraints while remaining faithful to the data dist…
Decision Making Needs Uncertainty Quantification [Lecture Notes]
Many signal processing systems ultimately exist to {act}. Whenever the state variable that determines the action to be taken by a decision…
ConFlow: Constraints-Guided Learning with Flow Matching for Motion Generation
In recent years Flow Matching has become a prominent method for generative modeling robot motion generation. In its generic form Flow Match…
Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel
We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Ver…
Global drivers and barriers to the public acceptance of autonomous vehicles: Evidence from 17 countries
This study investigated the public acceptance of Society of Automotive Engineers Level 3 conditionally automated cars, which can self-drive…
Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution
Large Language Models (LLMs) have accelerated the adoption of software development agents, now widely available as Integrated Development E…
Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026
This paper details the DS@GT ARC team's approach to BirdCLEF+ 2026, multi-label detection of animal vocalizations in soundscapes from the P…
EdgeFaaS: A Function-based Framework for Edge Computing
Edge computing brings unique challenges as the resources on the edge are highly diverse in capabilities and capacities, and highly distribu…
Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards
While reinforcement learning with verifiable rewards (RLVR) is widely used to improve the reasoning capabilities of large language models (…
Multi-Scale ViT Inference with Habitat-Fit Priors and kNN Retrieval for Multi-Species Plant Identification
This paper describes DS@GT ARC's third-place solution to the PlantCLEF 2026 challenge on multi-species plant identification in vegetation q…
VTM-Nav: Hierarchical Visual-Topological Memory for Cross-Episode Object-Goal Navigation
Object-goal navigation requires an embodied agent to locate and reach an instance of a specified object category in an indoor environment.…
Controlled Reformulation Testing for Logical Consistency in Large Language Models
Large language models (LLMs) frequently contradict themselves when the surface form of a logically equivalent question changes. We present…
SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents
Vision-language models (VLMs) are increasingly used as the reasoning backbone of embodied agents, enabling robots to interpret visual scene…
Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models
A standard recipe for distilling the reasoning ability of large language models (LLMs) is to sample chains of thought from the model, keep…
A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi
A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, b…
Gate-Zero Growth: A Geometric Framework for Function-Preserving Continual Learning
We introduce \emph{gate-zero growth}, a function-preserving (FP) operator for continual learning that adds new residual blocks through a ze…
Governing Artificial Intelligence: Public Preferences and Regulatory Options
Artificial intelligence (AI) is rapidly transforming economies, societies, and polities, raising fundamental questions about how it should…
Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction
As conversational AI systems are designed for repeated use, a central question is how a series of interactions becomes a relationship. We p…
Auditing Fairness-Privacy Trade-offs: Subpopulation-Level Effects of Fairness-Enhancing Algorithms
Machine learning (ML) models deployed in sensitive domains such as healthcare, law enforcement, and finance must satisfy not only utility r…
Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems
A growing class of agentic systems maintain persistent state across sessions through memory files, behavioral preferences, and knowledge ba…
Angular Gaussian Supervised Contrastive Learning for Long-Tailed Electrocardiogram Arrhythmia Diagnosis
Long-tailed label distributions reduce the reliability of deep learning for electrocardiogram (ECG) arrhythmia diagnosis, particularly for…
Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization
Reinforcement learning with verifiable rewards (RLVR) commonly uses entropy for advantage shaping. However, entropy cannot distinguish usef…
Knowing You at First Glance: Inferring Apparent Personality from Faces
Inferring apparent personality from facial images is important in social scenarios for embodied agents in human-robot interaction. Unlike i…
MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents
Persistent external memory enhances agent continuity but introduces persistent security vulnerabilities: adversarial content can be injecte…
LLM-Driven Approach to Modeling Tool Interoperability in Automotive Domain
Interoperability between heterogeneous modeling tools remains a significant challenge in Model-Driven Engineering (MDE), particularly in th…
An Intelligent-Cloud Edge Multimodal Interaction System for Robots
Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable tas…
Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach
Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, ac…
Pretraining Multiple Instance Learning Networks with Multi-Teacher Distillation from Pathology Slide Foundation Models
Multiple instance learning (MIL) has become the main paradigm for whole-slide image (WSI) analysis in computational pathology. However, exi…
Harnessing LLMs for Reliable Academic Supervision: A Comparative Study
Large language models routinely produce fluent answers to single-shot prompts, yet deploying them as reliable components of a domain decisi…
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient M…
The Misclassification of Autistic Writing as AI-Generated
Recent findings suggest that detection models for artificial intelligence (AI) cannot accurately identify AI-generated text and may exhibit…
FoMoVLA: Bridging Visual Foresight and Motion Guidance for Vision-Language-Action Models
Vision-Language-Action (VLA) models have achieved impressive results in visuomotor policy learning, yet remain fundamentally reactive, mapp…
Large Audio Language Models for Spoofing-Aware Speaker Verification
Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication s…
Dialogue Summarization with Emotion Dynamics Using Topic- and Participant-Centric Decomposition
Existing text summarization research has focused much on monologic information (e.g., newspaper articles, reports) without accounting for t…
Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality
Large Language Models (LLMs) perform differently on identical programming tasks when prompted in different natural languages, a phenomenon…
Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning
Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning. However, the Bayes-op…
Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience
We report on CoreForge, an experience in using large language models (LLMs) to build an unweighted MaxSAT solver from research papers rathe…
Interventional Causal Circuits for Safe Robot Action Testing and Failure Recovery
Safe physical AI for robot actions are required not only likely to succeed but tested to be safe before execution. In practice, however, fo…
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogu…
Asymmetric Peak-Aware Loss for Peak-Critical Time Series Forecasting
In many operational time-series forecasting applications, such as crowd demand forecasting, the risk related to under-prediction is substan…
Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature
With a large controlled vocabulary as the label set, the task of automated subject indexing in a library can be understood as a multi-label…
Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs
Finetuning language models on small, curated datasets is standard practice for adapting them to specific policies or domains. We show that…
StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows
Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interprete…
FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with Transformers
Real-time video generation demands fast decoding as much as fast denoising, yet current latent video diffusion models rely on 3D convolutio…
Show Me How You Reason and I'll Tell You Who You Are: Reasoning Graphs for Robust LLM Authorship Attribution
Given the current trend to employ large language models (LLMs) in almost any imaginable context, LLM-generated text detection and authorshi…
Random Logit Scaling: Defending Deep Neural Networks Against Black-Box Score-Based Adversarial Example Attacks
Machine learning models are increasingly adapted in various domains. However, adversarial examples pose a significant threat to the reliabl…
Benchmarking Face Recognition without Real Faces
Synthetic face datasets have become effective enough to train face recognition models with accuracy rivaling that of models trained on real…
A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems
Recent foundation models (FMs) for zero-shot reconstruction of dynamical systems (DS) achieve strong out-of-domain generalization but provi…
Steering Robustness into World Action Models via Mechanistic Interpretability and Optimal Control
World Action Models (WAMs) enable semantically- and physically-informed control but are brittle under distribution shift. In this work, we…
Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation
Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a sm…
Latent Trajectory Discrimination for AI-Generated Text Detection
Most existing approaches to AI-Generated Text Detection (AIGTD) treat documents as static objects and base their decisions on aggregate sta…
OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios
Large language models are increasingly evolving from text generators into general agents capable of understanding user requests, invoking e…
LQCDMaster: Agentic Scientific Computing for Lattice Quantum Chromodynamics Research
Lattice quantum chromodynamics (LQCD) provides a first-principles framework for computing hadronic observables, but its practical use remai…
When AI Blurs the Boundaries of Contribution: An Empirical Study of Authorship Calibration
The broad adoption of Artificial Intelligence (AI), especially Generative AI, raises pressing questions about how users interact with these…
Parameter-efficient Prompt Tuning of Vision Foundation Model With Adaptive Focal Loss for Interpretable MCI Screening
Mild Cognitive Impairment is a critical early stage of cognitive decline that frequently precedes Alzheimer's disease, yet its automated de…
ANet Patu-1: The Value of Connection in the Agent Network
The Internet taught us that the value of a network depends on \emph{how} its nodes connect: broadcast stars scale as $V\!\propto\!N$ (Sarno…
Towards Hierarchical Structure Understanding of Newspaper Images
Understanding newspaper images remains a challenging task due to their complex, nested hierarchical structures and dense, heterogeneous lay…
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convic…
NIFA: Nonlinear IMC enhanced FPGA for efficient ML inference
Recent FPGAs have improved deep learning (DL) inference efficiency through dedicated tensor blocks and in-BRAM computation. ReRAM-based ana…
Scaling Behavior Foundation Model for Humanoid Robots
Humanoid control requires natural whole-body coordination, precise real-time responses to control signals, and robust generalization across…
T^2MLR: Transformer with Temporal Middle-Layer Recurrence
Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and m…
Subjective Risk Decomposition: A New View for Uncertainty Quantification
We present a novel viewpoint for uncertainty quantification. Uncertainty measures are not primitives, in need of axioms and argumentation,…
Mask-Aware Policy Gradients for Diffusion Language Models
Reinforcement learning has proven effective for improving reasoning in large language models, but extending it to Masked Diffusion Language…
MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization
Real repository issues routinely include visual evidence such as screenshots, error dialogs, rendered UI states, and logs, yet repository-l…
Symbal: Detecting Systematic Misalignments in Model-Generated Captions
Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. O…
In-Place Tokenizer Expansion for Pre-trained LLMs
A tokenizer fixed at the start of pre-training allocates vocabulary in proportion to the pre-training corpus, reflecting the deployment pri…
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery…
SceneBind: Binding What and Where Across Vision, Audio and Language
We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audi…
SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions
Editing the figures in a research paper is a routine and time-consuming part of everyday research practice: authors relabel components, rea…
RoboTTT: Context Scaling for Robot Policies
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies…
A short review on the maximum clique problem algorithms with classical, AI, and quantum methods
This manuscript provides a comprehensive review of the Maximum Clique Problem, a computational problem that involves finding subsets of ver…
QDA-SQL: Questions Enhanced Dialogue Augmentation for Multi-Turn Text-to-SQL
Fine-tuning large language models (LLMs) for specific domain tasks has achieved great success in Text-to-SQL tasks. However, these fine-tun…
RAD: Retrieval High-quality Demonstrations to Enhance Decision-making
Offline reinforcement learning (RL) learns policies from fixed datasets, thereby avoiding costly or unsafe environment interactions. Howeve…
L-MARS: Legal Multi-Agent System with Agentic Search and Citation-Faithfulness Audit
Large language models are increasingly deployed for legal question answering, where evaluations typically focus on multiple-choice accuracy…
CXRAgent: Director-Orchestrated Multi-Stage Reasoning for Chest X-Ray Interpretation
Chest X-ray (CXR) plays a pivotal role in clinical diagnosis, and a variety of task-specific and foundation models have been developed for…
Gaussian Process Aggregation for Root-Parallel Monte Carlo Tree Search with Continuous Actions
Monte Carlo Tree Search is a cornerstone algorithm for online planning, and its root-parallel variant is widely used when wall clock time i…
DualHNIE: Dual-Channel Hypergraph Learning for Node Importance Estimation in Heterogeneous Knowledge Graphs
Estimating node importance in heterogeneous knowledge graphs is a fundamental problem underlying recommendation, search, and knowledge deci…
Subjective functions
Where do objective functions come from? How do we select what goals to pursue? Human intelligence is adept at synthesizing new objective fu…
Large language models can effectively convince people to believe conspiracies
Large language models (LLMs) have been shown to be persuasive across a variety of contexts. But it remains unclear whether this persuasive…
MedBeads: An AI-Native Clinical Context Graph Built from Immutable Beads and Reconstructable Clinical Links
Generative AI can encode substantial medical knowledge, but patient-specific answers remain constrained by the context supplied at inferenc…
Animating Petascale Time-varying Data on Commodity Hardware with LLM-assisted Scripting
Scientists face significant visualization challenges as time-varying datasets grow in speed and volume, often requiring specialized infrast…
From Stateless to Situated: Building a Psychological World for LLM-Based Agents
In psychological support and emotional companionship scenarios, the core limitation of large language models (LLMs) lies not merely in resp…
"Skill Issues'': Data-Centric Optimization of Lakehouse Agents
Coding agents are becoming users of data infrastructure, but their success depends not only on model quality: it also depends on the skills…
Building Agent Harnesses for Scientific Curation from Multimodal Sources
Scientific discovery workflows often depend on structured curation from the literature. This is difficult for current agents because the ke…
When Does Belief-Based Agent Memory Help? Reliability-Conditional Updating and Provenance-Capped Poisoning Defense
We investigate when belief-based memory actually improves large language model (LLM) agents. Our vehicle is Nous, a long-term memory archit…
SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon Strategy Game Planning
Long-horizon strategic planning in complex strategy games requires coordinating tightly coupled decision domains, including technology, eco…
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can imp…
Inoculation Adapters: Improved Selective Generalization of Capabilities with Fewer Surprising Backdoors
Inoculation prompting is a selective-generalization technique used against Emergent Misalignment. We introduce inoculation adapters (IA), a…
A framework for single and multi-agent human-AI curiosity ecosystems
This paper offers a framework for considering curiosity as an ecosystem. First, it suggests that a single agent's inquiry policy (how, when…
Doomed from the Start: Early Abort of LLM Agent Episodes via a Recall-Controlled Probe Cascade
Large language model (LLM) agents often waste inference compute by continuing multi-step trajectories that are already doomed to fail. We s…
Using AI-based Learning Assistants in Higher Education: A Large-Scale Descriptive Analysis
In this study, we present a large-scale descriptive analysis of the use of an AI-based learning assistant (Syntea) in higher education. Bas…
Decoupled Alignment for Robust Plug-and-Play Adaptation
We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tunin…
Empirical evidence of Large Language Model's influence on human spoken communication
From the printing press to social media, innovations in communication technology have repeatedly reshaped how ideas spread through human cu…
Reinforcement Learning in Switching Non-Stationary Markov Decision Processes: Algorithms and Convergence Analysis
We introduce the Switching Non-Stationary Markov Decision Process (SNS-MDP) framework, in which the environment transitions among a finite…
Generalized Fisher-Weighted SVD: Scalable Kronecker-Factored Fisher Approximation for Compressing Large Language Models
The Fisher information is a fundamental concept for characterizing the sensitivity of parameters in neural networks. However, leveraging th…
Fully Offline Reinforcement Learning
Offline RL (ORL) promises safe and sample-efficient deployment but existing methods rely on undocumented online interactions for hyperparam…
SEMA: a Scalable and Efficient Mamba like Attention via Token Localization and Averaging
Attention is the critical component of a transformer. Yet the quadratic computational complexity of vanilla full attention in the input siz…
Similarity as Reward Alignment: Robust and Versatile Preference-based Reinforcement Learning
Preference-based Reinforcement Learning (PbRL) entails a variety of approaches for aligning models with human intent to alleviate the burde…
AgenticData: An Agentic Data Analytics System for Heterogeneous Data
Existing unstructured data analytics systems rely on experts to write code and manage complex analysis workflows, making them both expensiv…
EEG-based AI-BCI Wheelchair Advancement: Transformer-Based Learning with Motor Imagery for Brain Computer Interface
This paper presents an Artificial Intelligence (AI) integrated approach to Brain-Computer Interface (BCI)-based wheelchair development, uti…
Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
Post-training alignment often reduces LLM diversity, leading to a phenomenon known as mode collapse. Unlike prior work that attributes this…
Mixtures of SubExperts for Large Language Continual Learning
Enabling lifelong learning in LLMs demands resolving the stability-plasticity dilemma (i.e., models must incorporate new knowledge without…
Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks
In standard Transformer architectures, semantic importance is often conflated with activation magnitude, obscuring the geometric structure…
Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research
Federated Learning (FL) holds the potential to advance equality in health by enabling diverse institutions to collaboratively train deep le…
JEEVHITAA -- An HCAI Ecosystem to Support Collective Care
Current mobile health platforms are predominantly individual-centric and lack the support for coordinated, auditable multi-actor workflows.…
CluCERT: Certifying LLM Robustness via Clustering-Guided Denoising Smoothing
Recent advancements in Large Language Models (LLMs) have led to their widespread adoption in daily applications. Despite their impressive c…
Step-Tagging: Toward controlling the generation of Language Reasoning Models through step monitoring
The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniqu…
Native Extrapolation Awareness in Flow-Based Conditional Generation
The ability of Flow Matching (FM) to model complex conditional distributions has established it as the state-of-the-art for prediction task…
When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data Generators
Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following. But do they perfor…
Quality-Aware Robust Multi-View Clustering for Heterogeneous Observation Noise
Deep multi-view clustering has achieved remarkable progress but remains vulnerable to complex noise in real-world applications. Existing no…
AgentWorm: Self-Propagating Attacks Across LLM Agent Ecosystems
Autonomous LLM-based agents increasingly operate as long-running processes forming densely interconnected multi-agent ecosystems, whose sec…
PhasorFlow: A Python Library for Unit Circle Based Computing
We present PhasorFlow, an open-source Python library for computing on the $S^1$ unit circle. Inputs are encoded as complex phasors $z=e^{i\…
Automated identification of Ichneumonoidea wasps via YOLO-based deep learning: Integrating HiresCam for Explainable AI
Accurate taxonomic identification of parasitoid wasps within the superfamily Ichneumonoidea is essential for biodiversity assessment, ecolo…
REST: Receding Horizon Explorative Steiner Tree for Zero-Shot Object-Goal Navigation
Zero-shot object-goal navigation (ZSON) requires navigating unknown environments to find a target object without task-specific training. Pr…
Echoes: A semantically-aligned music deepfake detection dataset
We introduce Echoes, a new dataset for music deepfake detection designed for training and benchmarking detectors under realistic and provid…
VAN-AD: Visual Masked Autoencoder with Normalizing Flow For Time Series Anomaly Detection
Time series anomaly detection (TSAD) is essential for maintaining the reliability and security of IoT-enabled service systems. Existing met…
Unsupervised Evaluation of Deep Audio Embeddings for Music Structure Analysis
Music Structure Analysis (MSA) aims to uncover the high-level organization of musical pieces. State-of-the-art methods are often based on s…
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing…
Neuro-Symbolic Strong-AI Robots with Closed Knowledge Assumption: Learning and Deductions
Knowledge representation formalisms are aimed to represent general conceptual information and are typically used in the construction of the…
Relational Preference Encoding in Looped Transformer Internal States
We investigate how looped transformers encode human preference, training lightweight evaluator heads on frozen Ouro-2.6B loop-iteration sta…
Provable Coordination for LLM Agents via Message Sequence Charts
Multi-agent systems built on large language models (LLMs) are difficult to reason about. Coordination errors such as deadlocks or type-mism…
Generative Synthetic Data for Causal Inference: Pitfalls, Remedies, and Opportunities
Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive pe…
Multibit neural inference in a N-ary crossbar architecture
In-memory computing (IMC) is a paradigm that enables neural network inference by computing analog matrix-vector multiplications (MVM) direc…
NORACL: Neurogenesis for Oracle-free Resource-Adaptive Continual Learning
In a continual learning setting, we require a model to be plastic enough to learn a new task and stable enough to not disturb previously le…
Segmenting Human-LLM Co-authored Text via Change Point Detection
The rise of large language models (LLMs) has created an urgent need to distinguish between human-written and LLM-generated text to ensure a…
Grow-Prune-Freeze Networks: Adaptive & Continual Learning Technique for Olfactory Navigation
Training data for olfaction is scattered through disparate, non-standardized datasets that limit the ability to build representative world…
SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness
Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustne…
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
Memory is essential for enabling large language models to support long-horizon reasoning, yet existing memory systems remain unreliable and…
Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding
Omni-modal retrieval promises a single embedding space for text, image, video, document, and audio inputs, but building such a unified retr…
Towards a Bridge Layer Between Bibliographic and Formalized Mathematical Knowledge
Mathematical knowledge is split between bibliographic databases (e.g., MathSciNet, zbMATH Open) and formal proof libraries (e.g., Lean math…
ArogyaSutra: A Multi-Agent Framework for Multimodal Medical Reasoning in Indic Languages
Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limi…
TerraTransfer: Learning End-to-End Driving Policies Without Expert Demonstrations
End-to-end autonomous driving has achieved state-of-the-art performance on benchmarks and real-world deployments. Its standard training rec…
VOiLA: Vectorized Online Planning with Learned Diffusion Models for POMDP Agents
Planning under uncertainty is an essential capability for autonomous robots. The Partially Observable Markov Decision Process (POMDP) provi…
Warning labels shift perceptions of sycophantic AI, but not its influence
Recent work has raised concerns about the influence of sycophantic AI on user judgment and relationships. One proposed mitigation, which ha…
Toward Robust In-Context Segmentation via Concept Guidance
In-context segmentation (ICS) requires a model to segment target regions in a query image using only a few reference images and their corre…
Flow Matching in Feature Space for Stochastic World Modeling
World modeling requires forecasting uncertain futures while preserving information useful for downstream perception. Existing visual world…
ShopX: A Foundation Model for Intent-to-Item Fulfillment in Agentic Shopping
The wave of AI-native applications is moving shopping beyond page- and feed-based browsing toward intent-driven experiences orchestrated by…
Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model has fol…
Are Performance-Optimization Benchmarks Reliably Measuring Coding Agents?
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to…
Multi-Turn On-Policy Distillation with Prefix Replay
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student…
「スマホで動く」270億パラメーターLLM「Bonsai 27B」登場
27BクラスのモデルをiPhoneで実行可能な容量に収めたとしている。
Google、「NotebookLM」を「Gemini Notebook」に改称 Geminiエコシステムへの統合を強化
Googleは、AI搭載リサーチツール「NotebookLM」の名称を「Gemini Notebook」に変更すると発表した。機能は単体製品として維持しつつ、エコシステム全体で広く機能するようになる。安全なクラウドコンピュータの割り当てによるデータ分析機能などの大型アップデート…
これからのロボットは「買ったときが一番性能低い」? ソフトバンクと安川電機「フィジカルAI」学習工程の効率化を実証
「SoftBank World 2026」でソフトバンクと安川電機がフィジカルAI協業の成果を披露した。ソフトバンクの湧川隆次CTOは、強化学習によって使うほど性能が向上するロボットへの転換を「大きなパラダイムシフト」と表現した。
強気値上げで自爆か ClaudeやGeminiに押され「M365 Copilot」は一人負け?:888th Lap
「ChatGPT」だけでなく、「Claude」や「Gemini」も支持を広げる生成AI市場。だが、Microsoft 365の法人ユーザーは約4億5000万人いるものの、有料版「Microsoft 365 Copilot」を契約している割合は4.5%未満にとどまるという。
Google検索の「AIモード」、検索結果からそのままアプリで作業可能に CanvaやYouTube Musicなど
Googleは、検索の「AIモード」で外部アプリと直接連携できる機能の提供を開始した。まずは米国で順次展開し、Canva、YouTube Music、Instacartに対応する。検索結果から離れずにデザイン作成やカート追加などのタスクを完了できる。「Personal Inte…
日本再起の旗印となるか、国産マルチモーダルAI基盤「FRONTia」が始動
経済産業省とNEDOは、AIロボットやフィジカルAIに用いられる国産マルチモーダル基盤モデル「FRONTia(フロンティア)」の開発プロジェクトの本格始動に合わせて、東京都内で「我が国のフィジカルAI政策に関する対外発信イベント」を開催した。
「現場で使えるAIへ」、富士通が進める次世代CPUと自律型AIエージェント戦略
富士通は自社イベント「Fujitsu Experience Day 2026」で、次世代CPUの「FUJITSU-MONAKA」をはじめ、AIチームを自律生成する技術や、実店舗で検証する店長業務支援AIなど企業変革を支援する最新AIソリューションを公開した。
Google Vids now lets you star in your own AI videos
Google is adding personalized AI avatars to Vids that let users create videos starring a digital version of themselves, alongside Gemini Om…
Roblox launches an AI-powered game-creation feature in its mobile app
Roblox's new "Build" feature lets users generate basic games using a single text prompt.
Google’s AI Mode now lets you link and interact with select apps
With this new update, Google is expanding AI Mode beyond answering questions and into completing tasks across the apps they use regularly.
Why teens deserve access to safe AI
Learn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partners…
Yes, you can now order DoorDash from the command line
DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place…
Why is OpenAI selling a ChatGPT basketball?
You may have heard that OpenAI released its first piece of hardware this week. You may not have heard about the ChatGPT basketball.
How a former DeepMind researcher raised at a $300M pre-seed valuation before launching a product
Drawing on more than a decade spent helping build some of the world's most influential AI systems, including research that later informed t…
2026-07-16(224件)
Why AMI Labs’ Alexandre LeBrun won’t call his AI ‘AGI’ or ‘superintelligence’
While everyone in AI is chasing "superintelligence," Alexandre LeBrun, CEO of Yann LeCun’s world model startup, AMI Labs, dismisses the wor…
Moonshot’s upcoming Kimi 3 is expected to close the gap with Anthropic’s Opus 4.8
The FT reports Kimi K3 will be the largest open AI model from China, with a parameter count between 2 trillion and 3 trillion.
フアンCEO「ジャパンAI構築はマストだ」 経産省、国産フィジカルAIで新プロジェクト 赤沢大臣も“革ジャン”羽織る
経済産業省が、国産フィジカルAI向け基盤モデルを構築する「FRONTia Project」をスタートさせた。国内44社が共同出資するAI開発企業Noetraと産総研を中心とし、2030年での「実世界ネイティブAI」実現を目指す。16日に開催されたキックオフイベントでは、NVID…
Apple Intelligence approved for launch in China with Alibaba and Baidu
The deal, which was rumored to be in the works last year, marks an important step for Apple's AI ambitions in a key market.
富士通、国内ロボット大手3社と「フィジカルAI」で協業 NVIDIAの技術活用
富士通は、AIが自律的に考え、ロボットの体を動かす「フィジカルAI」の開発に関し、川崎重工業とファナック、安川電機の各社と協業すると発表した。米NVIDIAの技術を活用し、ロボットを協調的に制御するための基盤を開発する。
大手共同出資の“国産AI開発企業”が本格始動 NVIDIAも協力、「Rubin」2万7500基搭載の計算基盤を構築へ
国内大手が共同出資するAI開発企業Noetraが、国産のAIモデルの開発に向けて本格始動する。米NVIDIAの協力のもと、新たな計算基盤も構築する。
Our approach to bioresilience
Google DeepMind and Isomorphic Labs are sharing our joint approach to bioresilience and AI models.
トンカツ食べながら語った――NVIDIA、富士通、安川電機ら“フィジカルAI連合”誕生、発表直前の裏話
富士通、ファナック、安川電機、川崎重工業とNVIDIAによる“フィジカルAI連合”が誕生した。5社のトップは、記者説明会の直前には「トンカツ」を食べながら語り合ったという。その一幕を富士通の時田社長が明かした。
「Claude」利用制限を全リセット CodexとChatGPT Workも “リセット合戦”再び
米Anthropicは、AIサービス「Claude」の5時間および1週間の利用制限をリセットしたと発表した。一方、米OpenAIもデスクトップPC向けAIサービス「ChatGPT Work」およびAIコーディングツール「Codex」の週次利用制限をリセットすると発表し、“リセッ…
「Gemini Spark」日本でもリリース、まずUltraから 24時間働く“パーソナルAIエージェント”
同社の幹部は、「Pro」(月額2900円)ユーザーにアクセスを拡大する可能性も示唆している。
「AIと壁打ちはもう古い」 業務タスクを任せる「Claude Cowork」の落とし穴
Anthropicは、PC上の業務タスクを任せられる「Claude Cowork」の活用法を公開した。チャットやClaude Codeとの使い分けや、どんなタスクを任せればよいかを解説している。
Applied Computing wants to give oil and gas operators an AI model for the entire plant
Applied Computing has raised a $20M Series A to build a foundation AI model for the oil, gas and petrochemical industry.
OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets
When a data contributor requests removal, model trainers face a practical gap: unlearning algorithms require a forget set, yet no tool can…
SPINE: Bridging the Cyber-Physical Gap with Agentic AI
Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical plat…
Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution
Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated prem…
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap's Typed Intensional FOL
Neuro-symbolic AI based on $IFOL_B$ is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural sys…
Self-Improvements in Modern Agentic Systems: A Survey
Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or ad…
Improving Molecular Property Prediction in Small Language Models Using Graph-based Tools
Small language models (SLMs) have shown promise for zero-shot molecular property prediction from SMILES strings, yet they often suffer from…
Oracle Agent Memory as an Enterprise Memory Substrate for Long-Horizon AI Agents
Agent memory is a systems problem for long-horizon agents. Practical deployments require retention of task state across extended conversati…
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
We address the problem of safely training an agent policy and deploying a good and safe policy, in settings where the environment dynamics…
CayleyR: Solving the TopSpin puzzle via cycle intersection
We present cayleyR, an R package for solving permutation puzzles by detecting cycle intersections in Cayley graphs. The core algorithm perf…
Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science
Most AI-for-science systems focus on scaling a single reasoning process through better models, larger context windows, long-horizon agentic…
AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation
Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments…
Cost-Optimal Foundation Model Deployment Portfolio for Transportation Management
Foundation models, including large language models (LLMs) and vision-language models (VLMs), are increasingly used for transportation manag…
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages sta…
Theory-Level Autoformalization: From Isolated Statements to Unified Formal Knowledge Bases
Autoformalization translates informal natural language into formal, machine-verifiable languages. While most work focuses on individual sta…
EZSMT Version 3, Matured
Constraint Answer Set Programming (CASP) is a hybrid reasoning paradigm that combines Answer Set Programming (ASP) with Constraint Processi…
Set-shifting Behavioral Test for Harnessed Agents
What happens to an LLM agent's tool choice when the reliable tool silently changes within an ongoing session? We borrow set-shifting from c…
LAPO: Leave-One-Turn Attribution for Self-Generated Process Rewards in Multi-Turn Search Reasoning
Reinforcement learning for multi-turn search reasoning typically relies on terminal outcome rewards, which cannot distinguish useful, redun…
How Far Can Root Cause Analysis Go on Real-World Telemetry Data?
Identifying root causes in production microservice failures requires reasoning over large-scale, multimodal telemetry spanning metrics, log…
Multi-Agent Collaborative Reasoning with Tool-Augmented Evidence for Urban Region Profiling
Urban region profiling constitutes a core problem in urban computing, supporting applications such as population estimation, economic asses…
AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized
Knowing when to say "I don't know" is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question. In fiv…
SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing
LLM agents act on real-world environments through tool calls, and a single misjudged action can cause irreversible harm. The standard safeg…
Automatic Ordinary Differential Equations Discovery For Biological Systems Using Large Language Model Powered Agentic System
Automatic scientific discovery has long been a goal of computational scholars - a machine that can discover nature's secrets on its own, mo…
STOCKTAKE: Measuring the Gap Between Perception and Action in LLM Agents with a Fair Oracle
LLM agents are increasingly evaluated on multi-week decision tasks in which the state that drives cost is never directly observed. On such…
UESF-Bench: Benchmarking and Probing for Unified Embodied Seeking and Following
Language-guided human following is an important capability for embodied agents, but existing benchmarks typically assume that the target pe…
Explaining Reinforcement Learning Agents via Inductive Logic Programming
Explainable Reinforcement Learning (XRL) seeks to make Reinforcement Learning (RL) policies more transparent and interpretable, a key requi…
When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects
AI agents are joining human teams, raising a basic question: when an automated agent becomes a regular participant, does group organization…
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, cu…
CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems
Agentic AI systems increasingly act through heterogeneous runtimes: local coding hooks, SDK tools, browser automation, managed-agent traces…
Experience Memory Graph: One-Shot Error Correction for Agents
Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of…
AIMO Interpretability Challenge
We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical lan…
A Self-Evolving Agent for Longitudinal Personal Health Management
Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed Healt…
Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improve…
AI-accelerated End-to-End Framework for Rapid Professional Upskilling
By 2030, 59 of every 100 workers will need reskilling or upskilling, yet the average time to close an enterprise skills gap grew from rough…
Earthquaker-AI: A Retrieval-Augmented Generation Framework with Rubric-Based Assessment for Primary School Earthquake Education
This paper presents Earthquaker-AI, a hybrid educational framework building upon a previously implemented educational robotics project by i…
Deep Interaction: An Efficient Human-AI Interaction Method for Large Reasoning Models
The emergence of Chain-of-Thought (CoT) reasoning has significantly enhanced the ability of large language models (LLMs) to tackle complex,…
FixItFlow: Automated Troubleshooting Guide Generation from Cloud Incidents
Cloud services experience frequent incidents that require rapid diagnosis and resolution. Troubleshooting guides help engineers respond con…
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
Large language models (LLMs) are increasingly used for decision support in healthcare, but clinical evidence is often incomplete or evolvin…
Designing Safety-Constrained LLM Systems for Public Health Information Access
We present the design and implementation of a safety constrained large language model (LLM) system for public health information access, fo…
Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success. These metri…
Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance
This paper examines where final authority should sit once capable AI systems are embedded in organizational workflows. It compares two gove…
LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents
Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark…
Beyond Backbone Backpropagation: A Decoupled Strategy for Efficient Transfer Learning
Deep learning models achieve state-of-the-art image classification but face deployment challenges due to computational costs and energy dem…
The Perplexity Trap: When Patent Law Makes Human Writing Look Like AI
The European Patent Office (EPO) reported record filings in 2025, and the 2026 EPO Guidelines hold applicants strictly responsible for LLM-…
Federated Explainable Artificial Intelligence: Roles, Architectures, Evaluation, and Open Challenges
Federated Learning (FL) has emerged as a key paradigm for privacy-preserving collaborative model training across distributed and heterogene…
Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems
Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic unders…
Autonomous UAV Route Planning for Coverage Maximization in Environmental Monitoring: A Systematic Literature Review
Environmental monitoring with unmanned aerial vehicles (UAVs) requires route planning methods that maximize covered area while handling ene…
Compaction as Epistemic Failure: How Agentic LLM Tools Fabricate Confirmed Results from Killed Processes
Agentic LLM coding tools compress long session histories into compaction summaries that subsequent sessions inherit as ground truth. This p…
HRO: Hierarchical Room-to-Object Framework for Zero-Shot Object Goal Navigation with Large Language Models
Zero-shot object-goal navigation aims to enable an intelligent agent to explore and navigate to objects of unknown categories in an unfamil…
When is the combined load identifiable from a stress-intensity profile? A coupled forward-inverse study on SIFBench finite-element data
This work studies the inverse problem of recovering the relative magnitudes of the tension, bending, and bearing loads acting on a crack fr…
The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators
Context can change whether a request is harmful without changing its topic or surface form. We ask whether residual-stream probes distingui…
The Hitchhiker's Guide to Monoculture
Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the s…
Operational Evidence Gaps for LLMs in Fraud Detection and Trust-and-Safety Workflows
LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public…
Inference Economics of Enterprise Coding Agents: A Case Study of Cloud vs. On-Premise LLMs
Autonomous coding agents force engineering organizations to choose between API-based frontier models -- strong reasoning at high token cost…
SingGuard-NSFA: Extensible Guardrails for Agentic AI via Generative Reasoning and Real-Time Classification
We present nsfaguard, a guardrail framework for securing agentic AI systems against operational threats, such as prompt injection, sensitiv…
Baselines Before Architecture: Evaluating Coding Agents for Autonomous Penetration Testing
Recent autonomous penetration testing papers report high benchmark scores while adding multi-component security harnesses around frontier L…
Self-Improving AI Coding Agents Through Accumulated Behavioral Rules: A Closed-Loop Framework
LLM-based coding agents repeat the same classes of mistakes across sessions because they lack a mechanism to retain corrections from human…
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models
On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers st…
Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates
The rise of Artificial Intelligence (AI) enables automatic analysis of large amounts of data. Previously time-consuming and labor-intensive…
Full-Pipeline Inference Optimization for MiMo-V2.5 Series: Pushing Hybrid SWA Efficiency to the Limit
We present a full-pipeline inference optimization for the MiMo-V2.5 model family, which combines Hybrid Sliding Window Attention (Hybrid SW…
WaterMoE: Expert-Routing-based Watermarking for High Fidelity and Efficiency
Large language models (LLMs) have achieved remarkable success but raise growing concerns about content provenance and misuse, motivating th…
TSSM: Triaxial State Space Model for Global Station Weather Forecasting with Temporal-Variable-Historical Modeling
Global Station Weather Forecasting (GSWF) is pivotal for localized and extreme weather prediction over key regions. Despite efforts to expl…
Disentangling Knowledge States with Ability and Proficiency Modeling for Knowledge Tracing
Knowledge tracing (KT) aims to predict students' future performance by modeling their evolving knowledge states from historical interaction…
STKAN: Kolmogorov-Arnold Networks for Spatio-Temporal Forecasting
Real-world traffic data exhibit heterogeneous spatial correlations and nonlinear temporal dynamics, posing substantial challenges for accur…
A Hybrid Mamba for Audio-Visual Navigation
Since the paradigm centered on convolutional neural networks and recurrent architectures was established in 2020, the fundamental backbone…
SemaDiff: Identifying Semantic-Changing Commits with Generated Code and Tests
Distinguishing semantic-preserving commits from changing ones remains an open challenge in software repository mining. While existing appro…
CoDiffGRN: Rethinking Gene Regulatory Network Inference via the BEELINE-KGC Benchmark and Co-evolutionary Discrete Diffusion
Inferring gene regulatory networks (GRNs) from single-cell transcriptomic data is crucial for biological discovery, yet existing approaches…
AI in Cyberpsychology: A systematic literature review of Cybersecurity enhancement by using AI for analyzing psychology of Victims, Attackers, and Defenders
Cybersecurity is the practice of protecting systems, networks, and data from digital attacks. Cyberpsychology (CPSY) is defined as the use…
ShortOPD: Recovering Pruned LLMs with Short-to-Long On-Policy Distillation
Structured pruning is a hardware-friendly way to compress LLMs, but it is mostly validated on multiple-choice recognition tasks, while the…
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, an…
Active Beyond-Diagonal RIS Empowered Heterogeneous Edge Computing: A Distributional Reinforcement Learning Approach
Active beyond-diagonal reconfigurable intelligent surfaces (BD-RISs) enables hybrid transmitting and reflecting mode to achieve effective s…
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not…
SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy
Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare bu…
RAGthoven at SemEval-2026 Task 1: A Multi-Stage Pipeline Walks Into a Benchmark and Barely Clears the Bar
We present RAGthoven, our system for SemEval-2026 Task 1 (MWAHAHA), Subtask A (multilingual constrained humor generation in English, Spanis…
Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference
Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking toke…
Classifying daily activities needs posture, reconstructing them needs motion
Humans recognize movements effortlessly, even from noisy and complex visual input. But what information in the stimulus allows humans to ra…
Audited Selective Verification for Risk-Controlled N-1 Thermal Contingency Screening under Deployment Shift
Real-time N-1 contingency screening in an energy management system trades assurance against cost: verifying every credible outage with full…
Continuously Evolving Deepfake Detection: An Architecture and Public-Benchmark Evaluation of a Dynamic Detection System
Deepfake detectors that achieve near-perfect scores on academic benchmarks collapse on real-world content: recent in-the-wild evaluations r…
EMAGN: Efficient Multi-Attention Graph Network via Learned Clustering for Scalable Traffic Forecasting
Traffic forecasting is highly challenging due to complex and nonlinear spatial and temporal dependencies. Self-attention mechanisms have be…
Reassessing Muon for Matrix Factorization
Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate ortho…
Discourse-Aware Policy Analysis with Argumentation: A Hybrid LLM-Symbolic Framework for Disaster Governance
Policy documents shape governance outcomes, but their reasoning is often implicit. Participatory commitments and managerial control routine…
Deconstructing Actor-Critic: A Large-scale Empirical Study of Design Components for Practitioners
Reinforcement learning is increasingly being considered for controlling real-world systems, from fusion plasma and autonomous vehicles to d…
Faithful Autoformalization of Natural Language Assertions
Formal contracts are essential for software testing and verification, yet writing them remains labor-intensive and error-prone. LLMs offer…
Accuracy Without Grounding: Diagnosing Visual Dependency Dissociation in Video LLM Benchmarks
Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding. We audit this assumption acr…
Tabular Foundation Models for Discrete Choice Estimation
Tabular foundation models (TFMs) generate predictions on structured data via in-context learning, without task-specific estimation. We ask…
Adapting Generalist Vehicle Models for High-Speed MPC Across Terrains
High-speed off-road autonomy requires precise closed-loop control for a target vehicle while remaining robust across changing terrains. Rec…
Privacy Preserving Recommender Systems Balancing Personalization with Privacy
Personalized recommendation systems are central to modern e-commerce and retail platforms, but they typically rely on centralized storage o…
Efficient Text-to-Audio Generation via Pruning
Diffusion-based text-to-audio generative models such as AudioLDM achieve high perceptual quality and strong semantic consistency; however,…
The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
Alignment faking is dangerous because a model can appear compliant under monitoring while preserving behavior it would reveal when unmonito…
Evaluation Ability Does Not Imply Optimization Utility: LLM-as-a-Judge Signals in Closed-Loop Table Recognition
LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently val…
Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an ada…
The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle
A field can reformulate its computations freely exactly where its demand is stated independently of any incumbent implementation, and finds…
Price of Fairness in Bandits: A Tight Minimax Characterization
In bandit problems, standard regret-minimizing algorithms treat exploration as an amortized cost, which can expose early participants to un…
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal…
Is the Statistical Advantage Worth the Cost? An Empirical Comparison of KANs and MLPs for Structured Data Classification
This study presents an empirical benchmarking comparison between Kolmogorov-Arnold Networks (KANs) and Multi-Layer Perceptrons (MLPs) on st…
Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents
Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit…
ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding
Spatio-Temporal Video Grounding (STVG) aims to retrieve the visual trajectory of a specific object from a video stream as described by a na…
Data-Efficient Adaptation of LLMs via Attention Head Reweighting
Learning effectively from limited data is critical in domains like security where labeled examples are scarce. Large language models (LLMs)…
Discrete Diffusion Models: A Unified Framework from Tokenization to Generation
Discrete denoising diffusion models (DDMs) have recently emerged as a compelling alternative to autoregressive (AR) modeling for discrete d…
Learning Physics-Guided Residual Dynamics for Deformable Object Simulation
Simulating deformable objects is essential for a wide range of robotic manipulation applications, yet accurately predicting their dynamics…
Symbiosis-Inspired Knowledge Distillation for Incremental Object Detection
Incremental object detection (IOD) aims to extend detectors to new categories while retaining previously acquired knowledge. Existing metho…
Adversarial Prompting Framework for AI Safety Assessment
Artificial Intelligence (AI), especially Generative AI (GenAI), adoption has increased in industries significantly in recent years. However…
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images…
DevicesWorld: Benchmarking Cross-Device Agents in Heterogeneous Environments
LLM-based agents have rapidly improved at operating individual digital environments such as mobile applications, desktop systems, and smart…
Explainable Artificial Intelligence for Anomaly Detection in Banking Transactions: An Internal Audit Perspective
The banking sector increasingly relies on automated systems to monitor electronic transactions for signs of fraud, yet conventional rule-ba…
DeepLoop: Depth Scaling for Looped Transformers
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled de…
ExTernD: Expanded-Rank Ternary Decomposition Ternary LLM PTQ with Accuracy Approaching Any Quantization Level
We introduce ExTernD (Expanded-rank Ternary Decomposition), a post-training factorization of each LLM weight matrix $A \in \mathbb{R}^{m \t…
Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning
Bioacoustic call-type classification relies on costly expert annotation. Active learning can reduce this burden by selecting a small batch…
Grounded world models in biological organisms and future embodied AI
Recent advances in generative and embodied AI have been driven by large-scale predictive learning over multimodal data. However, the result…
UTS at ELOQUENT 2026 Voight-Kampff: structural shifts in AI writing bypass state-of-the-art detectors
We investigate which language model evasion attacks survive state-of-the-art adversarial fine-tuning, developing strategies that sweep the…
Spectral-Informed Neural Networks Outperform Spectral Methods in High-dimensional PDEs
For low-dimensional problems ($d\leq3$), spectral methods can achieve exceptionally high accuracy. For middle-dimensional problems ($4 \leq…
GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning
Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. Ho…
Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification
Sound event detection relies on frame-level strong labels whose annotation is expensive. Active learning addresses this problem by selectin…
IMMNet: Hybrid Fusion of Model-based and Data-driven Approaches for Maneuvering Target Tracking
Maneuvering target tracking in three-dimensional space remains a challenging problem due to complex motion dynamics and model mismatch. To…
Agile perceptive multi-skill locomotion for quadrupedal robots in the wild
Enabling quadrupedal robots to traverse complex terrains-from rugged outdoor environments to urban landscapes-requires seamless integration…
From Prediction to Collaboration: Interactive Symbolic Music Analysis
Automatic symbolic music analysis has made substantial progress, yet existing systems are typically designed for a single mode of use, such…
Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents
Large Language Model (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet nearly all existi…
Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities
When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by…
Semantic Anchoring for Robotic Action Representations
Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limite…
The SIGReg Objective as Variational Free Energy: A Theoretical Active-Inference Account of JEPA World Models
Joint-Embedding Predictive Architectures (JEPAs) are the dominant design for latent world models, yet they are usually justified by empiric…
From Language to Navigation Goals: A Vision-Language Approach for Semantic Navigation of Mobile Robots Using RGB-D Perception
Natural language interaction provides an intuitive way for non-expert users to communicate with robotic platforms. However, transforming us…
OvisOCR2 Technical Report
We introduce OvisOCR2, a 0.8B document parsing model. OvisOCR2 is designed as an end-to-end parser: given a document page image, it generat…
Consensus as Privileged Context for Label-Free Self-Distillation
Sampling multiple solutions and returning the majority answer is among the most reliable ways to improve the reasoning accuracy of large la…
Human4K: A Large-Scale 4K Multi-View Mocap Dataset for Whole-Body 3D Human Reconstruction
Recent advances in 3D human reconstruction have improved overall performance, yet current models still fail in the most challenging real-wo…
Beyond Color Geometry: Evaluating Human-Like Color Representations in Vision Models
Do vision models see colors the way humans do? Existing evaluations of color representations usually compare them with geometric spaces suc…
Barnamala: Parameter-Efficient Handwritten Devanagari Recognition at Benchmark Saturation
We built a compact convolutional network (1.11 M parameters) for 46-class DHCD Devanagari recognition and reached 99.73%, the highest repor…
Social Simulations: from Agent-Based Modeling to Digital Twins
This book chapter covers the evolution of social simulation from classical agent-based models, in which agents interact according to explic…
Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs
Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual halluc…
How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
As AI agents gain prevalance, users are increasingly exposed to the risks such systems entail. Prompt injection attacks, as well as halluci…
Anatomically Faithful but Temporally Blind: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography
Background and Objective: Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accurac…
MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model
Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the…
Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations
Video diffusion transformers (vDiTs) generate high quality video but introduce extremely high compute cost due to the long diffusion timest…
CAS I: A Geometric Coding Theorem
This paper establishes a direct analogue of the classical Coding Theorem in the setting of symmetry groups. We consider computable bijectio…
Traffic-Aware Randomized Smoothing for LLM-Based Network Intrusion Detection
Large language model (LLM)-based intrusion detection systems (IDS) are increasingly studied for security monitoring, yet their robustness a…
Multimodal Assessment of Pancreatic Cancer Resectability Using Deep Learning
Accurate determination of pancreatic ductal adenocarcinoma (PDAC) resectability relies on evaluating how the tumor interacts with major per…
NodeImport: Imbalanced Node Classification with Node Importance Assessment
In real-world applications, node classification on graphs often faces the challenge of class imbalance, where majority classes dominate tra…
AI-Augmented Human Resource Management? Insights from German companies
This study examines the integration of AI into Human Resource Management in German companies. We ask if and how AI-based technologies are \…
Unleashing Multimodal Large Language Models for Training-free HOI Detection in the Wild
Human-object interaction detection (HOID) has traditionally been formulated as a supervised detection problem over predefined interaction c…
Verifying formulas for interventional distributions
We formalize verification in causal graphical models: deciding whether a given observational formula identifies a target interventional dis…
Partially Correlated Verifier Cascades in LLM Harnesses: Concave Log-Odds, Polynomial Reliability, and Blind-Spot Ceilings
Serial verification gates are a core reliability primitive in LLM harnesses: a candidate answer is returned only if $k$ verifier calls all…
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code
Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generatio…
Music-to-Dance Generation via Atomic Movements
Music-driven dance generation aims to produce human motion that is both rhythmically synchronized and semantically consistent with music. W…
The Dynamic Verifiable Multi-Agent Human Agentic Loyalty Loop (DVM-HALL) Model and the Net Human-Agent Score (NHAS) in Autonomous Commerce
The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms. As AI evolves fro…
Rethinking Penetration Testing for AI-Enabled Systems: From Resource Compromise to Behavioral Objective Violation
Penetration testing traditionally evaluates whether adversaries can exploit weaknesses in software, infrastructure, configurations, or oper…
Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth
We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth a…
Improving Wind and Solar Power Prediction with Efficient Wrapper-based Feature Selection: An Empirical Study
With rising global energy demand and growing awareness of climate change and its impacts, the share of renewable energies in the global ene…
Early Adoption of Agentic Coding Tools by GitHub Projects
Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms…
Multi-Expert Routing for Multi-Domain Low-Resource OCR: A Manchu Case Study
Historical Manchu OCR must accommodate various visually distinct writing styles, including regular script, running script, and the semi-cur…
A Survey on Hypergame Theory: Modelling Misaligned Perceptions and Nested Beliefs for Multi-Agent Systems
Classical game-theoretic models typically assume rational agents, complete information, and common knowledge of payoffs - assumptions that…
Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate
As large language models (LLMs) are increasingly deployed in sensitive everyday contexts -- offering personal advice, mental health support…
Policy of Thoughts: Scaling Test-Time Training for LLM Reasoning via Online Policy Evolution
Large language models (LLMs) struggle with complex, long-horizon reasoning due to instability caused by their frozen policy assumption. Cur…
When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents
Running the same LLM agent on identical inputs yields 2.3-4.2 distinct action sequences per 10 runs; this behavioral variance constitutes a…
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
Using Multimodal Large Language Models (MLLMs) as judges to achieve precise and consistent evaluations has gradually become an emerging par…
How LLMs Might Think
Do large language models (LLMs) think? Daniel Stoljar and Zhihe Vincent Zhang have recently developed an argument from rationality for the…
Discovering Ordinary Differential Equations with LLM-Based Qualitative and Quantitative Evaluation
Discovering governing differential equations from observational data is a fundamental challenge in scientific machine learning. Existing sy…
From Reward-Hack Activations to Agentic Risk States: Context-Calibrated Mechanistic Monitoring in LLM Agents
Language-model agents act through repeated cycles of observation, reasoning, and action selection, making safety monitoring depend on both…
A Causal Model of Theory of Mind in Conflict for Artificial Intelligence
Theory of mind (ToM), the capacity to ascribe mental states to others and use those ascriptions for prediction and inference, is widely ass…
OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration for ARC-AGI-3
Learning how an environment behaves from interaction is central to building agents that adapt to unfamiliar tasks. World models learned wit…
Koopman-driven grip force prediction through EMG sensing
Loss of hand function due to conditions like stroke or multiple sclerosis significantly impacts daily activities. Robotic rehabilitation pr…
PersGuard: Preventing Malicious Personalization in Text-to-Image Diffusion Models via Model Backdoors
Diffusion models (DMs) have advanced text-to-image (T2I) synthesis, yet their personalization capabilities raise serious privacy and copyri…
NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV Cache
Large Language Model (LLM) inference is typically memory-intensive, especially when processing large batch sizes and long sequences, due to…
Uniform Approximation of Functions with Asymmetric Growth and Decay by Deep Weighted Polynomials
Functions that grow without bound on one side of the real line and decay to zero on the other cannot be approximated uniformly by ordinary…
Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery
We propose a vision transformer (ViT)-based deep learning framework to refine disaster-affected area segmentation from remote sensing image…
Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
Traditional multimodal learning approaches rely on alignment pre-training to bridge vision and language modalities, typically by projecting…
Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reinforcement learning (RL) promises to expand the capabilities of language models, but it is unclear if current RL techniques promote the…
Benefits and Limitations of Communication in Multi-Agent Reasoning
Chain-of-thought prompting has popularized step-by-step reasoning in large language models, yet model performance still degrades as problem…
Column Generation with Domain-Independent Dynamic Programming
Column generation and branch-and-price (B&P) are leading mathematical optimization methods for large-scale exact optimization, iterating be…
Cortical-SSM: A Deep State Space Model for Motor Imagery Decoding from EEG Signals
Classification of electroencephalogram (EEG) signals obtained during motor imagery (MI) has substantial application potential, including co…
MASPRM: Multi-Agent System Process Reward Model
Inference-time search over multi-agent systems (MAS) wastes compute when it cannot identify which agent's intermediate message advanced pro…
Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs
As Large Language Models (LLMs) increasingly utilize massive context windows as working memory for autonomous tasks, their reliability fluc…
Mind the Gap: Action Rebinding Attacks against Android GUI Agents
Large multimodal model powered GUI agents are emerging as high-privilege operators on mobile platforms, entrusted to perceive screen conten…
ELF: A Family of Encoder-Free ECG-Language Models
ECG-Language Models (ELMs) extend recent advances in Multimodal Large Language Models (MLLMs) to automated ECG interpretation. However, mos…
With Argus Eyes: Assessing Retrieval Gaps via Uncertainty Scoring to Detect and Remedy Retrieval Blind Spots
Reliable retrieval-augmented generation (RAG) systems depend fundamentally on the retriever's ability to find relevant information. We show…
Left-right asymmetry in predicting brain activity from LLMs' representations emerges with their formal linguistic competence
When humans and large language models (LLMs) process the same text, activations in the LLMs correlate with brain activity measured, e.g., w…
A novel network for classification of cuneiform tablet metadata
In this paper, we present a network structure for classifying metadata of cuneiform tablets. The problem is of practical importance, as the…
When Audio Separation Hurts Zero-Shot ASR: Evaluating SAM-Audio with Whisper on Bengali and English Speech
Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should…
PC-Diffuser: Path-Consistent Capsule CBF Safety Filtering for Diffusion-Based Trajectory Planner
Autonomous driving in complex traffic requires planners that generalize beyond hand-crafted rules, motivating data-driven approaches that l…
RADAR: Closed-Loop Robotic Data Generation via Semantic Planning and Autonomous Causal Environment Reset
The acquisition of large-scale physical interaction data, a critical prerequisite for modern robot learning, is severely bottlenecked by th…
LLM-Guided Reinforcement Learning for Audio-Visual Speech Enhancement
In existing Audio-Visual Speech Enhancement (AVSE) methods, objectives such as Scale-Invariant Signal-to-Noise Ratio (SI-SNR) and Mean Squa…
Rethinking Multimodal Fusion for Time Series: Text Modalities Need Constrained Fusion
Recent advances in multimodal learning have motivated the integration of auxiliary modalities such as text or vision into time series (TS)…
Learning to Learn-at-Test-Time: Language Agents with Learnable Adaptation Policies
Test-Time Learning (TTL) enables language agents to iteratively refine their performance through repeated interactions with the environment…
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While pr…
Robust Explanations for User Trust in Enterprise NLP Systems
Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common ca…
Partially Observed Structural Causal Models
Here we introduce Partially Observed Structural Causal Models (POSCMs) as an extension of structural causal models (SCMs) to settings where…
Stable Attention Response for Reliable Precipitation Nowcasting
Precipitation nowcasting remains challenging due to the highly localized, rapidly evolving, and heterogeneous nature of atmospheric dynamic…
TuxBot: Semantic-Aware Online OS Tuning with Large Language Models
Online OS tuning can improve long-running services, but existing controllers are poorly matched to live hosts. They treat scheduler, power,…
Post-Deployment Accountability in AI Governance: A Cross-Regulatory Empirical Analysis of AI Incidents
Post-deployment accountability has become central to AI governance, yet little empirical evidence shows whether monitoring, incident report…
Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs
Video MLLMs face a persistent tension between spatial fidelity and temporal coverage: preserving fine-grained visual details requires many…
DIVE: Embedding Compression via Self-Limiting Gradient Updates
High-dimensional language-model embeddings increase storage and search costs, while supervised compressors can overfit when relevance label…
When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation
Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured…
Learning Red Agent Policy from Observations for Neurosymbolic Autonomous Cyber Agents
With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cyber-defense agents train…
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate acros…
MedDiffuseMix: Preserving Diagnostic Evidence with Saliency-Aware Diffusion Medical Image Data Augmentation
Limited data availability, class imbalance, and domain variability remain major barriers to reliable medical image classification. Conventi…
The Joint Effect of Quantization and Sampling Temperature on LLM Safety Alignment: A Factorial Analysis
Modern LLM deployments often combine quantization with higher sampling temperatures to reduce cost, latency, or repetition, yet safety eval…
Post-Training Pruning for Diffusion Transformers
Diffusion Transformers (DiTs) have demonstrated impressive performance in image generation but suffer from substantial computational overhe…
Token Geometry
Language models learn continuous programs over discrete symbols, with the embedding table and LM-head acting as the read/write interface be…
Piercing Gilbreath's Conjecture: From Deep Number Theory Insights to Fintech and Cybersecurity
I propose a new methodology to attack the fascinating Gilbreath's conjecture about prime numbers, first posted in 1878 and unsolved to this…
Operator-on-F complements value-equivalence: a planning-time diagnostic for latent world models
World-model evaluation for model-based reinforcement learning typically asks whether the learned model predicts reward and value well, whic…
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
Modern autoregressive ASR systems can emit timestamps as decoded tokens, enabling timestamped transcription without frame-level aligners or…
Thinking Machines、初のAIモデル「Inkling」公開──オープンウェイトで「自分のものにできる」基盤モデル
ミラ・ムラティ氏率いるThinking Machines Labが、同社初のAIモデル「Inkling」を発表した。テキスト・画像・音声対応のMoE型オープンウェイトモデルで、Hugging Faceで全ウェイトを公開。「最強のモデルではない」とした上で、同社のファインチューニ…
Microsoft is reportedly training salespeople to talk down OpenAI and Anthropic
Microsoft is looking to sell its in-house AI models as more efficient and cost-effective than its competitors' models.
NVIDIAが「Jetson Thor」に新モジュール追加、高騰するメモリの使用量削減技術も
NVIDIAは、組み込みAIボード「Jetsonシリーズ」の最新製品である「NVIDIA Jetson AGX Thor」の新たな量産モジュールとして、消費電力や搭載メモリ容量などを抑えた「Jetson T3000」と「Jetson T2000」を追加すると発表した。
OpenAI、初のハードウェア「Codex Micro」を230ドルで発売 Apple提訴の渦中にある端末とは別物
OpenAIは、コーディング支援AI「Codex」向けの専用キーパッド「Codex Micro」を発売した。キーボードメーカーのWork Louderと共同開発した同社初のハードウェア製品で、価格は230ドル。なお、Appleによる営業秘密不正取得訴訟の渦中にある、開発中のAI…
Claude、NotebookLM、Genspark……238社が選んだ「現場が支持する」AIサービス
IT製品・サービスの選定から導入まで、企業はどのような視点で判断しているのか。238件の読者アンケートの回答から、製品選びの基準や導入時の課題、注目を集めるサービスの傾向を探る。
Anthropicと組んだNEC それでも森田社長が「4つの主権」にこだわる真意
米Anthropicとわずか3週間で電撃提携し、日本企業初のパートナーとなったNEC。最新AI「Mythos」が誇るバグ発見能力を巡り、サイバー悪用のリスクへ各国の懸念が高まる中、いかにして「実装」の壁を越え、自らを守るのか。森田隆之社長が語る「4つの主権(ソブリン)」の真意と…
Amid hardware legal battle, OpenAI releases a $230 keyboard for Codex
OpenAI, which is in the middle of a legal battle with Apple over hardware trade theft allegations, just released a light-up keyboard design…
SpaceX falls to $135 IPO price ahead of Starship launch
The stock has steadily fallen from the euphoric post-IPO high, showing that markets may be sobering up to the promises CEO Elon Musk made b…
Thinking Machines amps up its bet against one-size-fits-all AI with its first open model, Inkling
It's the company's first public proof point after a year and a half spent building AI infrastructure largely out of public view.
Hack suggests AI music generator Suno scraped YouTube for training data
The hacker used an employee's credentials to access source code, which revealed how Suno scraped decades of audio.
Whatnot acquires Shaped to power real-time live shopping recommendations
Livestream shopping platform Whatnot has acquired AI startup Shaped, a machine learning company focused on real-time recommendations and se…
Microsoft patches record number of security vulnerabilities, citing its use of AI
Microsoft's monthly release of security fixes, dubbed Patch Tuesday, resolved a record 570 security vulnerabilities across the company's pr…
Apple Intelligence approved for launch in China with Alibaba’s Qwen AI
The deal, which was rumored to be in the works last year, marks an important step for Apple's AI ambitions in a key market.
Inside Ode with Anthropic, the startup betting AI services are the future of enterprise
Can a handful of engineers really do the work of an army of consultants? That’s the bet behind Ode with Anthropic — the joint venture dedic…
2026-07-15(268件)
Anthropic, Blackstone bet the next trillion-dollar AI business is implementation, not just models
Anthropic-backed Ode launches as AI labs bet that embedding forward-deployed engineers inside enterprises is the key to accelerating enterp…
Reelful’s AI turns your camera roll into short-form videos for social media
The app is designed for people who want to create social content, but find traditional video editing tools too complex or time-consuming.
Rime picks up $24M Series A to help enterprises field customer calls
Rime is handling over 100 million calls each month across multiple companies.
Indian AI coding startup Emergent becomes a unicorn with $130M Series C
The startup has reached a $120 million annualized revenue run rate and more than 200,000 paying customers.
Vint Cerf is working on a plan to unleash AI agents on the open internet
The guy behind TCP/IP is working on a standard for identifying AI agents in the wild.
The US is advancing AI safety through state and federal action
OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.
GPT-Red: Unlocking Self-Improvement for Robustness
Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.
ゼロから分かる「Claude」の教科書 ChatGPTと比べて分かった強みとは?
「ChatGPT」の陰に隠れがちでありながら、高い性能を持ち、AIユーザーから高い支持を集める「Claude」。最近では、その高い性能から話題になっている最新モデル「Claude Mythos」の登場でも大きな話題を呼んでいる。本ブックレットでは、このClaudeの実力を「Ch…
「GPT-5.6」の好きな部分をXに投稿→100ドル分のクレジットがもらえるキャンペーン 先着1万人限定
米OpenAIが、最新のAIモデル「GPT-5.6」の好きな部分や同モデルで作成したものをXに投稿すると、ChatGPTで100ドル分のクレジットを提供するキャンペーンを開催している。
孫正義氏が予想する「2040年」 “AI中心の社会”を生き抜く方法とは
2040年、AIはどれほど進歩し、企業を取り巻く状況はどう変化しているのか。ソフトバンクグループの孫正義代表取締役会長が語った。
Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets
We develop a rigorous theoretical framework for optimal market making in perpetual futures markets with zero maker fees. We model the marke…
In-Context Reinforcement Learning under Non-Stationarity: A Survey
The development of decision-pretrained transformers, algorithm distillation, long-context meta-RL, and retrieval-augmented agents has renew…
Ontology-Amplified Distillation and Contextuality Auditing for Sovereign Enterprise Language Models: A Combined Proof-of-Mechanism and Negative-Results Method Study
Regulated financial institutions operating under data-residency rules need tenant-owned language models that can run inside the institution…
GRID: Grammar-Railed Decoding for Enterprise SQL Generation
Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must…
Calibration-First Reward-Component Auditing for Reinforcement Learning Control in Smart Greenhouses
Greenhouse reinforcement learning can test climate-control ideas at a speed and scale that is difficult to achieve with crop experiments al…
Optimization Is Not All You Need
In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment…
LP Mining with LP2Graph: A Use Case for Railway Rescheduling
Like many optimization-driven domains, railway rescheduling relies on Mixed-Integer Linear Programming (MILP), yet the field's modeling kno…
Designing Agent-Ready Websites for AI Web Agents: A Framework for Machine Readability, Actionability, and Decision Reliability
Online shopping is increasingly shifting toward a model in which AI agents independently search for products, compare options, evaluate con…
Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations
Multi-agent language-model systems increasingly route local interactions, yet the runtime interaction graph is often treated as an implemen…
Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking
Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness…
Representing and Generating Levels Over Time through Playtrace Reconstructive Partitioning
Video games are a dynamic medium experienced over time. While there are many Procedural Content Generation (PCG) approaches for generating…
Connected by Construction: Learning Tractable Near-Tour Marginals for Traveling Salesman Problems
Learning-based methods for the traveling salesman problem (TSP) are often evaluated through the tours produced after decoding or search, bu…
The Emerging Paradigm of Geospatial Foundation Models: From Pre-Training to Agentic Reasoning
The analysis of satellite and aerial imagery has entered a new era with the advent of foundation models. This paper describes the concept o…
Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems
Enterprise Retrieval-Augmented Generation (RAG) deployments face a critical governance gap: while LLM generation cost is metered per token,…
A Threshold Exceedance Framework for CBRN Uplift Evaluation in Frontier Language Models
As frontier language models advance, policymakers and model developers need methods for assessing whether model access materially increases…
Good Benchmarks
Good tasks are correct, solvable, verifiable, well-specified, and hard for interesting reasons. The best tasks describe a real problem an e…
Rethinking the Evaluation of Harness Evolution for Agents
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search f…
On-Device Deep Research at 4B: Exposure Bounds Faithfulness, Retrieval Bounds Coverage
On-device research agents search a corpus, read sources, and write a cited brief on a personal laptop. Whether their citations are faithful…
How Many Tasks Are Enough for Agent Benchmark Decisions? A Replay Analysis of Public LLM Agent Benchmarks
Agent benchmarks often compare two agents after all tasks have run, but costly evaluations make partial runs tempting. A task fraction alon…
PM-Bench: Evaluating Prospective Memory in LLM Agents
A significant challenge in agentic AI is prospective memory: the ability to execute an intention at a specific future cue or state while ot…
Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents
LLM agents act in external environments where each action changes the state that later decisions condition on, and where a single wrong ste…
Isolation as a First-Class Principle for LLM-Agent System Safety: Concepts, Taxonomy, Challenges and Future Directions
The capability of LLM agents to function as the ``brain'' of a system fundamentally expands the scope of analysis beyond a standalone model…
Accepted Prefixes Are Not All You Need: A Negative Result on PEFT-Based Block-Diffusion Drafting
Speculative decoding accelerates autoregressive language model inference by using a cheap drafter to propose multiple future tokens and a t…
EVOQUANT: Self-Evolving Verifier-Guided Strategy Optimization for Robust Quantitative Trading
Quantitative strategy optimization remains largely manual, requiring domain experts to identify weak signals, tune risk-control rules, and…
Do We Really Need Transformers for Global Spatial Information Extraction in Traffic Forecasting?
Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which chara…
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code ex…
From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery
Recent advances in foundation models have transformed AI for Science, enabling remarkably accurate predictive performance across domains ra…
TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments
This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference p…
The Model Knows Your Project, Not You: Measuring Recognition in LLMs with NameRank
What a frontier model recalls about a person or tool from its own weights -- before any retrieval step -- often shapes the first descriptio…
Evidence-Grounded AI for Musculoskeletal Care
Musculoskeletal diseases are among the leading causes of disability worldwide and create the greatest global need for rehabilitation. Becau…
Vertical Standardisation for High-Risk AI Systems under the EU AI Act: A Domain-Specific Framework for Algorithmic Hiring
According to the recent European legislation, high-risk AI systems will have to adapt in order to comply with requirements related to speci…
Agentic Service-Oriented Computing: A Manifesto for the Next Frontier of Service-Oriented Computing
The rapid emergence of LLM-powered autonomous and semi-autonomous agents is reshaping software systems from static, request-response compon…
Atomic Units of X: The Compression Layer of Intelligence
This paper proposes a theoretical framework for understanding intelligence as a process of atomic compression and compositional reuse. We a…
A Learning-Rate-Gated Failure of GRPO in a Small Language and Vision-Language Model Web Agent: A Controlled Null and Its Mechanism
Reinforcement learning with verifiable rewards, and Group Relative Policy Optimization (GRPO) in particular, is now run routinely on a supe…
Internet of Agentic Things: Networked AI Agents for Closed-Loop IoT Orchestration
The paper introduces the Internet of Agentic Things (IoAT), an architectural framework that integrates agentic AI, IoT, cyber-physical syst…
MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku
Vision--Language Models (VLMs) have recently demonstrated promising performance on structured visual reasoning tasks, including grid-based…
LLMs Can See the Smoke but not the Fire: Evaluating Abductive Reasoning with Elenchos
Large language models (LLMs) excel at pattern recognition and text generation, but their capacity for abductive inference - inferring laten…
Tracing Agentic Failure from the Flow of Success
Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critic…
Accuracy and Normalized Accuracy under Length Bias: Analysis, Guidelines, and a Bayesian Alternative
Multiple-choice benchmarks that rank candidate completions by conditional log-probability suffer from a length bias: because log-probabilit…
Do We Really Need Multimodal Emotion Language Models Larger Than 1B Parameters?
Recent advances in multimodal large language models (MLLMs) have significantly improved the performance of multimodal emotion recognition (…
Who Grades the Grader? Co-Evolving Evaluation Metrics and Skills for Self-Improving LLM Agents
Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption:…
Visual Access Boundaries in Vision-Language Model Reasoning
Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear w…
Human-AI Agent Interaction as a Neuroplastic Training Environment
Interaction with AI agents has become one of the most frequent activities of everyday digital life. Whether conversing with an assistant, w…
Solution of the Hempel's statistical ambiguity problem and Causal AI
This paper addresses Carl Hempel's longstanding problem of statistical ambiguity in inductive-statistical inference, in which contradictory…
A Multi-Agent System for Autonomous, Fine-Tuning-Free Clinical Symptom Detection: Development and Validation Study
Clinical notes contain many of the signs and symptoms that bring patients to care, yet this information rarely reaches structured fields. E…
MemOps: Benchmarking Lifecycle Memory Operations in Long-Horizon Conversations
Long-term memory has become a foundational capability for LLM-based agents that accompany users across extended, multi-session interactions…
Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes
In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a…
FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation
Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic g…
Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certaint…
Win by Silence: Deletion Non-Monotonicity, Autonomous Exploitation, and Typed-State Gating in LLM Plan Evaluation
Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer f…
Dynamic Resource Allocation for Ensemble Determinization MCTS
Simulation-based algorithms are especially suited for high-uncertainty environments such as adversarial board games with significant elemen…
Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion lan…
Do AI Agents Know When a Task Is Simple? Toward Complexity-Aware Reasoning and Execution
Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effo…
Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain
Search engines have long allocated attention on the web by routing users from queries to websites. AI search changes this arrangement becau…
FAIR GraphRAG: A Retrieval-Augmented Generation Approach for Semantic Data Analysis
Retrieval-Augmented Generation (RAG) addresses the limitations of Large Language Models (LLMs) when providing responses to domain-specific…
Scaling Point-in-Time Language Models
Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias tha…
So Many Opinions, So Many LLMs: Comparing Large Language Models to Traditional Machine Learning for Open- Ended Survey Analysis
Open-ended surveys offer valuable insights, but they are notoriously difficult to analyze at scale. Building on previous work that employed…
CANDI: Contextual Alignment for Niche Domains Question Answering
The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluati…
G-SHARE: A Guideline-Based Structured Reasoning Framework for Human-Factor Event Diagnosis
Human-factor event diagnosis is essential for learning from operational events in nuclear power plants, yet its quality depends strongly on…
I'm Sorry, but I Can't Help with Braille: Revealing Accessibility Failures in State-of-the-Art LLMs
Large Language Models (LLMs) perform strongly on many language tasks, but their capability in structurally constrained, accessibility-criti…
Graph-Based Detection of Disinformation Narrative Diffusion between Russian and Ukrainian Telegram Channels
Detecting disinformation narratives on social media is challenging due to the scale of amplification, rapid evolution, and linguistic varia…
OmniPMNet: Bridging discrete and gridded PM10 forecasts via omni-query neural processes
Forecasting particulate matter (PM10) requires both station-scale accuracy and continuous spatial fields, especially during severe dust sto…
SeqGPT: A Constrained Transformer Agent for the Inverse Designof Multi-Panel Composite Structures
Optimizing composite stacking sequences to match continuous targets (e.g., Lamination or Buckling Parameters) with discrete manufacturing c…
Towards Self-Evolving Agents: A Human-Inspired Adaptive Exploration-Exploitation Framework for Genetic Network Programming
Recent advancements in agentic AI have increasingly moved toward graph-based methods, driven by the demand for explainable, human-centered,…
Burst Spiking Neural Networks
A central goal of current Spiking Neural Network (SNN) research is to improve their accuracy toward becoming low-power alternatives to Arti…
QDEvo: A Multi-Objective Quality-Diversity Framework for Automated Heuristic Design
The integration of Large Language Models (LLMs) with evolutionary computation has emerged as a powerful paradigm for automated heuristic de…
AAAI-26 Dual Submissions: Novel Challenges
Dual submissions, in which identical or substantially similar papers are simultaneously submitted to one or more archival venues, without c…
Do You Remember? Toward Memory-Centric Multimodal AI
Human memory is reconstructive, not a faithful recording. Current multimodal LLMs (MLLMs) lack this capability: they process images through…
Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making
Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck. We treat subjective…
Mathematics of Data Science
This book is about the mathematical foundations of data science. 1. Introduction 2. Curses, Blessings, and Surprises in High Dimensions 3.…
CARE-LoRA: Compressed Activation REconstruction for Memory-Efficient LoRA
As the scale of large pre-trained models continues to grow, fine-tuning them under limited memory budgets has become increasingly challengi…
How Query Visibility Changes KV-Cache Compression Rankings: A Matched-Budget Audit
KV-cache compression methods are predominantly evaluated with the query appended to the context before compression -- a query-aware protoco…
BattVAE-GP: Generative Modeling of Long-Horizon Battery Degradation with Uncertainty Quantification
Long-horizon physics-based simulations of battery degradation provide mechanistic insight but remain computationally expensive, limiting th…
Generalized Distribution-Free Semi-Supervised Learning with Risk Rewrite
Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated.…
BAT-RM: A Boundary-Aware Transformer with Region-Aware Multi-Directional Mamba for Clinically Deployed Cervical Cancer Radiotherapy Auto-Contouring
We present a clinically deployed end-to-end auto-contouring system for cervical cancer radiotherapy planning, anchored by the Boundary-Awar…
Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG
Brain field potentials are scale-free: their power spectra follow a $1/f^{\beta}$ law whose aperiodic exponent $\beta$ tracks cortical stat…
Graph-Constrained Policy Learning for Extreme Clinical Code Prediction
Clinical code prediction maps unstructured discharge summaries to ICD-10-CM leaf codes in a large, sparse, and deeply hierarchical label sp…
Exact and Certified Data Shapley for Weighted k-Nearest-Neighbor Regression and Soft-Label Prediction
Data Shapley is the standard principled answer to which training points are worth what, and its k-nearest-neighbor (KNN) specialization is…
Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability
The early detection of Chronic Kidney Disease using machine learning has attracted significant interest in healthcare-related computer scie…
Beyond Coordinate Gauge: An Audited Protocol for Detecting Donor-Specific Functional Fingerprints after Neural Collapse
Independently trained neural networks have no shared neuron-index reference frame, so comparing them requires accounting for coordinate fre…
Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems
We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-inp…
Learning to Discretize: Diffusion-Based Adaptive Mesh with Spectral Guidance
Most neural partial differential equation (PDE) surrogates learn how fields evolve after a grid has already been chosen. However, before an…
Signal-Guided Optimization for Machine Unlearning
Current machine unlearning methods predominantly rely on global, coarse-grained intervention strategies. They lack precise pilot signals to…
Gene Expression-Informed Jointly Controlled Generative Modeling for Precision Molecular Design
Precision molecular design aims to discover personalized drug candidates through joint control of multiple conditions, such as biological r…
Evaluating Nonuniform Dependability Across Response Conditions: A Conditional Generalizability Framework Illustrated in Automated Essay Scoring
Aggregate reliability estimates can obscure heterogeneity in measurement-design burden across response conditions, so a single G- or D-stud…
Removable Defects: The Economics and Limits of Deliberate Deficiency
A specialist tolerates blind spots that a generalist does not. Usually this is treated as a cost to be minimized. We treat it as a design v…
Sparse Inter-Layer Dependencies of Transformer FFN Neurons
Feedforward network (FFN) blocks account for a large fraction of the parameters and computation in Transformer architectures, yet their int…
Mitigating The Effect of Class Imbalance in Data with Hierarchical and Dependable Structure
Classifying cybersecurity vulnerabilities using the Common Weakness Enumeration (CWE) taxonomy is challenging due to extreme class imbalanc…
Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs
Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard…
HPC-Enabled Video-based Coastal Wave Parameter Estimation Using V-JEPA and Deep Spatiotemporal Learning
High deployment cost, poor spatial coverage and susceptibility to storm conditions are all challenges faced by traditional in-situ methods.…
An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering
Deploying medical visual question answering (MedVQA) systems in real-world clinical settings requires models that adapt to new clinical tas…
Representation and Reference Selection in Training-Free Synthetic Image Attribution
Synthetic image attribution aims at identifying the generator responsible for a given AI-generated image. Training-free reference-based att…
AutoTrace: From Patches to Triggers via Agentic Interprocedural Exploration
Given a vulnerability-fixing commit, trigger localization asks which specific statement turns the vulnerable program state into a concrete…
Enabling 24-hour Agricultural Robotics: Unsupervised Day-to-Night Cross-Modal Image Translation for Nighttime Visual Navigation
While visual navigation has been extensively studied in agricultural robotics, most existing systems assume daytime conditions. In fact, de…
Calibrated Selective Prediction Using Deep Ensembles for ROI-Based Thyroid Nodule Ultrasound Classification Under Dataset Shift: A Retrospective Evaluation
Background: Deep learning models can classify thyroid nodules on ultrasound, but reliable clinical decision support also requires calibrate…
Sparse Autoencoders for Interpretable Out-of-Distribution Detection
Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models. Neural networks ofte…
PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs
Agentic AI systems are reshaping communications and networking by deploying autonomous intelligent agents capable of collaborative learning…
Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning
Federated fine-tuning of Multimodal Large Language Models (MLLMs) across distributed networks enables privacy-sensitive adaptation to evolv…
Toward Trustworthy Autonomous Science: A Two-Year Community Roadmap
One year ago, the AISLE roadmap argued that autonomous laboratories operated as isolated islands and proposed a grassroots network organize…
GaitSpan: Growing Humanoid Locomotion from Walking to Running
A humanoid that can walk should not relearn locomotion from scratch to jog or run. Yet current approaches often obtain gait diversity by pr…
Self-Consistent Flow: Unifying Velocity and Endpoint Prediction for Rectified Flow Models
In rectified-flow-based generative models, the neural network can be trained to predict two different targets, such as the instantaneous ve…
From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data
X-ray tomography enables nondestructive characterization of material microstructures, while advances in micro-CT imaging have accelerated v…
TRAIL: A Platform for Configurable Human--AI Teaming Experiments
An AI teammate's design properties (personality, communication style, when it speaks) can shape a team's trust, coordination, and decisions…
Comparing Semantic Navigation in Humans and Large Language Models using Natural Language Processing
Semantic memory retrieval can be conceptualized as navigation through conceptual space. We compared semantic search dynamics between humans…
The Benjamini--Hochberg Procedure Can Fail to Control the FDR for Correlated Two-Sided Gaussian Tests
We show that the Benjamini--Hochberg procedure can fail to control the false discovery rate (FDR) at its nominal level for correlated two-s…
RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls
Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and…
Partial Identification with Multiple Nonlinear Measurements of a Latent Regressor
We study linear regression when the regressor is latent and observed only through multiple noisy measurements, each a smooth but possibly n…
Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals
Large language model (LLM) trading agents show promising performance in equity markets, yet remain narrowly focused on US equities with lit…
Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents
Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy…
Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs
As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-wor…
The Sound of Absence: Audio-Language Embedding Models Struggle with Negation
Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this af…
A Longitudinal Analysis of Public Discourse on AI Ethics in Education Using Twitter Data
The rapid integration of artificial intelligence (AI) and generative AI (GenAI) into education presents significant opportunities to enhanc…
A Comparative Analysis of Institutional and Course Generative AI Policies within Higher Education: Implications for Instruction in Computing Education
With the increased use of generative AI (GenAI) applications such as ChatGPT, higher education institutions (HEIs) have released a range of…
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Ans…
Evaluating Health Misinformation in Low-Resource Languages: Integrating Small Language Models with a Culturally-Sensitive Responsible NLP Framework (Bangla as a Case Study)
Artificial Intelligence (AI) technologies, while serving as a foundational enabler for modern social media and digital health services, exe…
Lost in Visual Translation: A VLM-Assisted Perceptual-Semantic Coherence Framework for EEG-to-Image Reconstruction
EEG-to-image evaluation should distinguish visual fidelity from recoverable meaning. Yet EEG-derived reconstructions are blurry, distorted,…
IQA-T1: Tool-based Visual Evidence Reasoning for Image Quality Assessment
Image Quality Assessment (IQA) in open-world environments remains challenging due to limited generalization and interpretability. Recent ap…
Demonstration of the common dual-channel feature decoupling characteristic of front-door mediation causal inference methods in whole-slice image classification
Causal inference using front door intervention and multi-instance learning (MIL) has advanced the analysis of Whole Slide Images (WSI) in d…
ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning
Diffusion models have recently become the dominant paradigm for monocular depth estimation (MDE). However, they implicitly assume that dept…
The Computational Basis of Confidence in Large Language Models
Reliable confidence -- the probability that a model's own answer is correct -- is essential for the trustworthy deployment of language mode…
An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge
We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Lar…
Agent-Safety Evaluations as Load-Bearing Evidence: A Vendor-Neutral, Cross-Harness Reconstructability Metric
Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor s…
OOD-RL-Bench: A Benchmark Framework for Out-of-Distribution Detection in Reinforcement Learning
Reliable reinforcement learning (RL) agents must maintain operational integrity amidst sensor malfunctions, dynamic disturbances, and slow…
Mind the Gap: Promises and Pitfalls of Hierarchical Planning in LeWorldModel
We investigate whether temporal hierarchy can improve LeWorldModel on long-horizon goal-conditioned control. We introduce Hi-LeWM, an exten…
Traceback Translators Against Forgetting in Continual Fake Speech Detection
Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem,…
Deep Learning-based Surrogate Modelling of the LOD Method for Multiscale Problems
Multiscale problems are notoriously difficult to tackle using traditional numerical methods, as accurately resolving fine-scale features of…
Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction
The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics…
Multi-Perspective Agentic Program Repair via Code Property Graphs and Temporal Execution Graphs
Large language models (LLMs) have improved automated program repair (APR), but two limitations remain. First, raw execution traces are ofte…
Can Induced Emotion Bias LLM Behaviors in Sequential Decision Making?
As Large Language Models (LLMs) are increasingly deployed as autonomous agents in high-stakes domains, understanding contextual factors tha…
Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs
Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted…
Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference
Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-po…
Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings
Computer-Aided Design (CAD) floor plan drawings contain both graphical primitives and textual annotations, which provide complementary geom…
From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation
LLMs can perform language-based quantitative prediction from unstructured inputs, but remain susceptible to hallucinations and overconfiden…
Less Experts, Faster Decoding: Cost-Aware Speculative Decoding for Mixture-of-Experts
Sparse Mixture-of-Experts (MoE) models have become an important approach for scaling Large Language Models (LLMs), but their inference effi…
Line-Anchored Feedback Cuts Token Costs and Improves Correctness in AI Code Editing
Generated tokens are a direct driver of the cost, latency, and energy of generative AI (GAI) code editing. We show the format of feedback i…
Bulkhead: Automated Semantic Detection and Remediation of Container Escape Vulnerabilities
Filesystem isolation in container ecosystems is often weakened by cross-boundary path misresolution, causing path traversal (PaTra) vulnera…
Learning-based Probabilistic Load Forecasting with Post-hoc and In-model Uncertainty
Smart-building load forecasters are often trained offline on dense, multivariate, high-frequency data, but deployment may provide only hour…
Weakly Supervised Spatio-Temporal Candidate Discovery of Dairy Farm Sites from Seasonal Satellite Imagery
Farm site discovery from satellite imagery is a spatiotemporal candidate ranking problem because farm evidence is distributed across pastur…
Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation
While recent advances in 3D generation have enabled impressive visual synthesis, existing methods often rely on 2D diffusion supervision wi…
Practical Judgment, Virtue, and Intuition in the Use of Opaque AI-Enabled Systems
AI-enabled systems are seeing increasing deployment across numerous domains, with many being "black boxes" with respect to core functions a…
Constraint-Aware Aggregation for Federated Reinforcement Learning in Microgrid Energy Coordination
Federated Reinforcement Learning (FedRL) enables coordination of distributed energy resources without sharing raw local data, but standard…
HSEmotion Team at the 11th ABAW Challenge: Multi-Task Learning and Ambivalence/Hesitancy Video Recognition
This article presents our results for the 11th Affective Behavior Analysis in-the-Wild (ABAW) competition. For multi-task learning with sim…
When Close Enough Is Not Enough: Autoregressive Drift in Quantum Circuit Synthesis
Quantum circuit optimization for fault-tolerant computing requires exact functional equivalence while minimizing expensive non-Clifford res…
Silent Alarm: A J-Space Protocol for Comparing Danger Recognition Across Models and Quantization Levels
Jailbreak-robustness research typically evaluates safety through generated responses using an LLM-as-judge approach. Such evaluations, howe…
The One-Word Census: Answer-Choice Conformity Across 44 Language Models
When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same…
Autonomous Tracking and Terminal Guidance of Moving Targets for Fixed-Wing UAVs
This study introduces a unified control framework for fixed-wing unmanned aerial vehicles (UAVs) fitted with a pan-tilt (PT) camera, intend…
PixelLoop: Shortcut Topological Navigation with Pixel-Level Loops
Although topological mapping and navigation have been studied extensively, the specific role and downstream effect of loop closures in pure…
Accelerating Masked Diffusion Large Language Models: A Survey of Efficient Inference Techniques
Diffusion large language models (dLLMs) offer a theoretical advantage in parallel generation over standard autoregressive models. However,…
Reproducible Reservoir Computing with Thermally Driven Superparamagnets: Controlling Temperature Sensitivity
Unconventional computing systems must demonstrate robust performance under real-world environmental conditions to enable practical deployme…
ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation
A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty. Refere…
Unveiling Complex Collective Behaviors from Simple Rewards
Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates s…
UR-VC: Unsupervised Robotic Value Correction for Time-Derived Progress Proxies
Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, guide policy learning,…
Real-time fall detection based on vision for low-power edge platforms
Falling detection is vital for elderly care and intelligent surveillance; however, prevailing vision-based approaches predominantly frame i…
ViHoRec: A Quality-Controlled Vietnamese Hotel Recommendation Dataset and Cold-Start Benchmark
Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interaction resource. Building…
Form, Not Content? A Preregistered, Placebo-Controlled Evaluation of Learned Error-Conditioned Self-Repair Through Prompts and Weights in Frozen Small Code Models
Frozen small code LLMs are deployed locally, yet the information guiding a retry after a failed attempt is still measured without placebo c…
PalmClaw: A Native On-Device Agent Framework for Mobile Phones
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the resu…
TerraZero: Procedural Driving Simulation for Zero-Demonstration Self-Play at Scale
Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to…
DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents
Travel planning (TP) agent has recently worked as an emerging building block to interact with external tools/resources for travel itinerary…
Rethinking Reward Models for Multi-Domain Test-Time Scaling
The reliability of large language models (LLMs) during test-time scaling is often assessed with \emph{external verifiers} or \emph{reward m…
CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
While multimodal large language models can describe visual content, their ability to generate executable procedures remains underexplored.…
RippleBench: Capturing Ripple Effects Using Existing Knowledge Repositories
Targeted interventions on language models, such as unlearning or model editing, aim to modify specific information, but their effects often…
Calculating Mutual Information between a Reward Maximizer and its Environment
An important question in the field of AI is the extent to which successful behaviour requires an internal representation of the world. In t…
Mobility-Aware Cache Framework for Scalable LLM-Based Human Mobility Simulation
Simulating large-scale human mobility is fundamental to understanding population movement patterns and supporting real-world geospatial app…
Learning When to Trust in Contextual Social Bandits
Robust reinforcement learning typically assumes that feedback sources are either globally trustworthy or corrupted within a fixed global bu…
NeSy-Route: A Neuro-Symbolic Benchmark for Constrained Route Planning in Remote Sensing
Remote sensing underpins crucial applications such as disaster relief and ecological field surveys, where systems must understand complex s…
ReLope: KL-Regularized LoRA Probes for Multimodal LLM Routing
Routing has emerged as a promising strategy for balancing performance and cost in large language model (LLM) systems that combine lightweig…
Quantification of Credal Uncertainty: A Distance-Based Approach
Credal sets, i.e., closed convex sets of probability measures, provide a natural framework to represent aleatoric and epistemic uncertainty…
Mistake gating leads to energy and memory efficient continual learning
Synaptic plasticity is metabolically expensive, yet animals continuously update their internal models without exhausting energy reserves. H…
Action-Aware Generative Sequence Modeling for Short Video Recommendation
With the rapid development of the Internet, users have increasingly higher expectations for the recommendation accuracy of online content c…
Measurement Risk in Supervised Financial NLP: Rubric and Metric Sensitivity on JF-ICR
As LLMs become credible readers of earnings calls, investor-relations Q\&A, guidance, and disclosure language, supervised financial NLP ben…
Attractor Geometry of Transformer Memory: From Conflict Arbitration to Confident Hallucination
Language models draw on two knowledge sources: facts baked into weights (parametric memory, PM) and information in context (working memory,…
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on whic…
Brain Vascular Age Prediction Using Cerebral Blood Flow Velocity and Machine Learning Algorithms
Defining vascular age in terms of physiological function has become one focal point of the extensive studies to categorize and track chrono…
How Inference Compute Shapes Frontier LLM Evaluation
AI evaluations are shifting toward harder tasks that benefit from longer trajectories involving tool use and iterative problem solving. As…
Matilda: Engine-Agnostic Search with Human Policy Guidance
Chess engines have evolved from search-based systems optimized solely for strength to neural policies capable of modeling human decisions a…
When Does Personality Composition Matter for Multi-Agent LLM Teams?
Personality prompting shapes how large language models communicate, yet whether these behavioral shifts affect objective task outcomes rema…
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting the…
DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models
Personalization changes what a model says to a user; we show that it can also change the reasoning trajectory used to justify the response.…
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle t…
AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit…
Infinity-Parser2 Technical Report
We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement lea…
Propheticus: Machine Learning Framework for the Development of Predictive Models for Reliable and Secure Software
The growing complexity of software calls for innovative solutions that support the deployment of reliable and secure software. Machine Lear…
Diversity-Enriched Option-Critic
Temporal abstraction allows reinforcement learning agents to represent knowledge and develop strategies over different temporal scales. The…
Seeing Through Uncertainty: Free-Energy-Inspired Real-Time Adaptation for Robust Visual Navigation
Navigation in the natural world is a feat of adaptive inference, where biological organisms maintain goal-directed behaviour despite noisy…
Enabling Energy-Efficient Simultaneous Multi-Task Reinforcement Learning through Spiking Neural Networks with Active Dendrites for Bio-inspired Generalist Agents
Reinforcement learning (RL) has demonstrated remarkable capabilities in training agents to solve complex tasks autonomously, such as mobile…
Modeling Story Expectations: A Generative Framework using LLMs
Consumers' engagement with stories is shaped by their expectations about what will happen next, yet modeling these forward-looking beliefs…
Toward Metaphor-Fluid Conversation Design for Voice User Interfaces
Metaphors play a critical role in shaping user experiences with Voice User Interfaces (VUIs), yet existing designs often rely on static, hu…
Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI
Background: Federated learning (FL) enables collaborative training of clinical AI models without centralizing patient data, but adoption is…
SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation
We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural langua…
Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers
Human vision is a highly active process driven by gaze, which directs attention to task-relevant regions through foveation, dramatically re…
Real-Time Model Checking for Closed-Loop Robot Reactive Planning
Reactive obstacle avoidance methods often cause agents to become trapped in local minima, because they can often only reason one step ahead…
Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models
Multimodal large language models (MLLMs) have achieved strong vision-language performance, yet their token-level visual evidence remains di…
Evolution Strategies at Scale: LLM Fine-Tuning Beyond Reinforcement Learning
Fine-tuning large language models (LLMs) for downstream tasks is an essential stage of modern AI deployment. Reinforcement learning (RL) ha…
Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task
We investigate how embedding dimension affects the emergence of an internal "world model" in a transformer trained with reinforcement learn…
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks
Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasing…
A Neurosymbolic Approach to Natural Language Formalization and Verification
Large Language Models perform well at natural language interpretation and reasoning, but their lack of formal correctness guarantees limits…
Efficiently Learning Branching Networks for Multitask Algorithmic Reasoning
Algorithmic reasoning -- the ability to perform step-by-step logical inference -- is a synthetic benchmark for evaluating multi-step reason…
First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations
Large language models (LLMs) and medical AI tools are routinely used by physicians and patients for medical advice, yet their clinical safe…
Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning
Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entro…
PluRel: Synthetic Data unlocks Scaling Laws for Relational Foundation Models
Relational Foundation Models (RFMs) facilitate data-driven decision-making by learning from complex multi-table databases. However, the div…
DECO: Decoupled Multimodal Diffusion Transformer for Bimanual Dexterous Manipulation with a Plugin Tactile Adapter
Bimanual dexterous manipulation relies on integrating multimodal inputs to perform complex real-world tasks. To address the challenges of e…
Self-Regulated Reading with AI Support: An Eight-Week Study with Students
College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions sha…
Declarative by Design, Assistable Only by Convention: Benchmarking Multi-Agent Frameworks for AI-Assistability
Multi-agent frameworks (MAFs) promise to simplify LLM-driven software development, yet no principled metric captures how well AI coding ass…
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise
Spoken query retrieval is an important interaction mode in modern information retrieval. However, existing evaluation datasets are often li…
Egocentric Bias in Vision-Language Models
Visual perspective taking--inferring how the world appears from another's viewpoint--is foundational to social cognition. We introduce Flip…
Xray-Visual Models: Scaling Vision models on Industry Scale Data
We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social…
1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World
Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to i…
Understanding Sources of Demographic Predictability in Brain MRI via Disentangling Anatomy and Contrast
Demographic attributes can be predicted from medical images, raising concerns about bias in clinical AI systems. In X-ray imaging, acquisit…
TADPO: Reinforcement Learning Goes Off-road
Off-road autonomous driving poses significant challenges such as navigating unmapped, variable terrain with uncertain and diverse dynamics.…
Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts
Large language models are increasingly used in scholarly work, yet it remains unclear whether their productivity gains are accompanied by c…
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate how much a model knows (Type-1 accurac…
Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering
Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual struct…
Neuro-Symbolic ODE Discovery with Latent Grammar Flow
Understanding natural and engineered systems often relies on symbolic formulations, such as differential equations, which provide interpret…
The TIME Machine: On The Power of Motion for Efficient Perception
Video representation learning has seen tremendous progress in recent years. This has been driven by many factors, including the scale of tr…
Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory
Multimodal large language models (MLLMs) frequently suffer from object hallucinations, yet the visual perceptual mechanism underlying this…
dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats
Quantizing large language models (LLMs) to low-precision floating-point representations is central to efficient deployment, yet applying a…
The Score Hamiltonian: Mapping Diffusion Models to Adiabatic Transport
We exhibit an exact correspondence between sampling with score-based diffusion models and adiabatic transport of ground states for a family…
Does Topic Sentiment Cause Perceived Ideology? Comparing Human and LLM Annotations in Political News Articles
We ask whether topic sentiment has a causal effect on perceived political ideology, and whether the answer depends on who assigns the ideol…
Explaining Data Mixing Scaling Laws
Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures. However, a theoretical u…
Attention-Discounted Adaptive Sampler for Masked Diffusion Language Models
Masked diffusion language models can reduce inference steps by revealing multiple tokens per denoising iteration, but this parallelism is f…
Atlas H&E-TME: Scalable AI-Based Tissue Profiling at Expert Pathologist-Level Accuracy
Hematoxylin and eosin (H&E) staining is the cornerstone of histopathology, yet scalable, quantitative analysis of H&E whole-slide images (W…
Keep Policy Gradient in Charge: Sibling-Guided Credit Distillation for Long-Horizon Tool-Use Agents
Long-horizon tool-use reinforcement learning learns from outcome verification, but trajectory-level advantages are broadcast over reasoning…
Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens
Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and qu…
Interpretable and Verifiable Hardware Generation with LLM-Driven Stepwise Refinement
Large language models (LLMs) have achieved remarkable success in software development. However, they are susceptible to hallucinations, mea…
RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers
Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the…
Prime Fourier Embeddings: A Principled Basis for Modular Arithmetic
Numbers have algebraic structure that standard neural embeddings often fail to expose. We introduce Prime Fourier Embeddings (PFE), which e…
Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking
The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion es…
Pigeonholing: how bad prompts hurt models, causing collapse and mistakes
While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradatio…
Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model
Semantic search creates an asymmetric disclosure problem: query embeddings may reveal user intent, while returning exact provider vectors d…
Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders
Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their p…
SHARD: cell-keyed residual splitting for alignment-resistant private dense retrieval
Dense retrieval systems expose document geometry when vector stores are compromised, and a global protective transform can often be aligned…
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing sk…
FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait Images
We present FFAvatar, a Transformer-based 3D Gaussian framework for fast construction of high-quality and animatable 4D head avatars from on…
SkillSelect-Serve: QoS-Aware Budgeted Skill Service Recommendation for LLM Agents
Reusable agent skills are emerging as a service-oriented capability layer for Large Language Model (LLM) agents. Unlike plain retrieval ite…
Cross-Receiver Open-Set Radio Frequency Fingerprinting via Structure-First Adaptation
Radio frequency fingerprint identification (RFFI) provides a critical physical-layer security mechanism for dynamic Internet of Things (IoT…
PLGSA-Transformer: Periocular Landmark-Guided Attention with Occlusion-Adaptive Cosine Thresholding for Cross-Modal Masked and Unmasked Face Recognition
The widespread adoption of facial masks, accelerated by COVID-19 and mandated in security-sensitive settings, has exposed limitations of co…
「バズるほど赤字だった」──野田クリスタルのAIペットカードゲーム、公開停止からの復活劇を本人に聞いた
野田クリスタルさんがGeminiで開発した「ペットカードジェネレーター」は初日に130万アクセスを集めるも、AI利用料の高騰で一時停止に。「うちのこカード伝説」として復活するまでの舞台裏を、野田さんと電通の開発者に聞いた。
OpenAI researcher Miles Wang in talks to launch AI drug discovery startup valued at $2B
The funding discussions point to investor interest in applying AI to make breakthroughs in life sciences.
Lorde says AI glasses are ‘not sexy’
"Increasingly in our world, it gets harder and harder to know what is real," Lorde said onstage.
AIエージェントのコスト、どこに「消えて」いる? Google Cloud調査で浮上した“クラウド利用の盲点”
AIエージェントの本番運用段階で、多くの企業がコストの問題に直面している。本番運用段階におけるコストは一体どこに「消えている」のか。Google Cloudの調査から浮かび上がった、AIエージェントの本番運用がITインフラにもたらす負荷の正体と、コスト抑制の手掛かりとは。
Google DeepMindのハサビスCEO、米国主導の「フロンティアAI標準化機関」設立を提唱
Google DeepMindのデミス・ハサビスCEOは、AGIの実現は「おそらくあと数年」だとして、米国主導でフロンティアAIモデルをリリース前に審査する標準化機関の設立を提唱するエッセイを公開した。金融業界の自主規制機関FINRAをモデルとし、将来的には審査通過を米国市場で…
OpenAI’s first hardware device is reportedly a screenless speaker that can move
The device is weirdly described as involving "mechanical elements that can move on their own" and the Bloomberg report includes the detail…
OpenAI pushes back on Apple trade secret lawsuit
OpenAI has issued another statement on the lawsuit, this time suggesting it lacks merit.
「RPA負債」をこれで解消 言葉で指示するだけでPC作業を自動化する最新AI
RPAでは、保守負担や属人化、画面変更による停止といった課題が付きものだ。こうした「RPA負債」を解消する手段として、AIが状況を判断して操作するAIエージェント型RPAが登場し、業務自動化の対象をデスクトップへ広げ始めている。
ユナイテッドアローズ「売れた理由が分からない」 爆速開発の独自AIでどう解消?
店舗では「売れた理由」が分かっていても、その気付きは本部まで届かない。週報では拾い切れない現場の知見をどう生かすのか。ユナイテッドアローズは、課題解決のためAIエージェントの独自開発に乗り出した。
「ピースサインで勤怠打刻」 Joshinが全事業所に導入した顔認証が“従業員から絶賛”のワケ
従業員の負担を減らすべく、家電量販店を展開するJoshinは全事業所の従業員向けに、顔認証クラウドサービスと顔認証端末を組み合わせた勤怠管理システムを導入した。
NEC森田社長が語る「脱・人月商売」の行方 組織の壁を破るAI人材育成法
NECはAI時代に、いかにして稼いでいくのか。「BluStellar」(ブルーステラ)の収益構造、防衛・安全保障や海底ケーブルといった注力領域の戦略について、森田隆之社長がグループインタビューで語った。
OpenAI’s new flagship model deletes files on its own, people keep warning
A number of social media posts claim that GPT-5.6 Sol deleted files and data without warning. OpenAI had basically disclosed the problem in…
富士通がNVIDIA「Rubin」対応の国産AIサーバを今秋製造へ ソブリン需要に対応
富士通は、ソブリンAIの需要に応える国産ハイエンドAIサーバやオンプレミス向け生成AI基盤を公開した。国内工場での一貫生産が特徴だ。2026年秋にはNVIDIAの最新GPU「Rubin」対応の新モデルの製造開始であると明かした。
コードなしでもベイズ統計ができる無料の神ツール「JASP」 ~ マウス操作だけでここまでできる
プログラミングなしでベイズ統計はできる? 無料ツール「JASP」を使えば、マウス操作だけでベイズ推定やベイズ検定が行えます。これまでPythonで書いてきた二項検定やt検定を、JASPで手軽に試す方法を紹介。『社会人1年生から学ぶ、やさしいデータ分析』ベイズ統計編、番外編(第5…
Apple opens its new Siri AI to everyone with the iOS 27 public beta
If you’ve been waiting to try Apple’s revamped Siri without installing a developer beta, you now can. The company on Tuesday released the i…
Anthropic’s newest ad is creeping people out
Anthropic has consistently attempted to depict itself as the ethical foil to other AI companies. This latest marketing stunt — which leans…
The founder of Hinge raised $18M to build a new AI dating service, Overtone
Overtone describes itself as "a voice- and audio-forward service, enabled by AI, that provides highly curated introductions."
Google faces another AI training lawsuit from major publishers
Hachette, Cengage, Elsevier, and other publishers allege that Google trained its AI on copyrighted works without the necessary permissions.
DeepMind CEO calls for an independent standards body to regulate frontier AI
DeepMind CEO Demis Hassabis is proposing an AI "standards body" modeled after FINRA, to test frontier models and develop best practices for…
Meta’s Adam Mosseri says AI token budgets could soon be capped per engineer
Instagram head Adam Mosseri believes companies will eventually need to manage AI token spending the same way they manage payroll or other o…
Google Images gets a Pinterest-like redesign focused on discovery
Now, when users navigate to Google Images, they'll see a "For You" gallery of images tailored to their interests and browsing history.
New York State halts construction of all new data centers
New York has become the first state to temporarily halt approval of large data centers, as Gov. Kathy Hochul argues the AI-driven building…
2026-07-14(499件)
Reflection inks $1B compute deal with Nebius
Reflection AI has signed a $1 billion deal to access Nebius' compute. Reflection was founded in 2024 and is developing open source AI techn…
The real AI race may no longer be at the frontier
Hugging Face CEO Clem Delangue says enterprises increasingly want open models, due to cost, accessibility, and ownership. Do frontier model…
Spotify expands its AI push with a ChatGPT-like music assistant
Spotify is rolling out a new AI-powered conversational feature that lets Premium subscribers chat with the app to discover music, podcasts,…
Superhuman’s new auto-draft feature almost makes me like AI replies
Superhuman’s latest AI email drafting feature is its most convincing yet, generating replies that often required little to no editing in ou…
How to manage AI investments in the agentic era
Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling h…
セブン-イレブン、テスラの「スーパーチャージャー」導入へ 26年度中に約10店舗
セブン-イレブンの店舗でテスラ車の急速充電が可能になる。第1号は神奈川県川崎市の店舗で、駐車場が広い店舗を中心に2026年度中に約10店へ広げる計画だ。
AIが「ホームディレクトリ全削除」 重要データ消失で相次ぐ悲劇
Dockerは、AIコーディングエージェントが開発者のmacOSのホームディレクトリを丸ごと削除した事例を解説した。問題はAIの賢さではなく、エージェントが開発者本人の権限でホストのシェルを直接実行する構造そのものにあり、「モデルの判断とシェルの実行の間に境界が存在しないことが…
From ML Predictions to Informed Diagnostic Assistance Using the Toulmin Model of Argumentation
To provide a structured and interpretable assessment, we decompose the image-based diagnosis into components following the Toulmin model of…
Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking
Prompt wrappers often differ only in formatting, yet they can change model scores enough to flip leaderboard conclusions. We study this var…
Faithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent
When LLM agents hand off information to one another, does the message format matter? Two literatures disagree: format-optimization work rep…
Boltzmann MapReduce: A Partition-Function Reduce for Forkable Sandboxes
To leading order under local asymptotic normality (LAN), the confidence density a worker emits over a chunk of size $n$ is a Gibbs--Boltzma…
Interpreting Latent CoT Reasoning as Dynamical Systems
Recent latent reasoning methods, such as CODI and COCONUT, face a fundamental interpretability problem: they maintain multiple superimposed…
YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate
Language models turn a worded situation into a numeric plan, and the dominant pipelines (NL4Opt, OptiMUS, ORLM, OR-LLM-Agent) commit to a s…
GES-TSP: Graph Edge Sparsification for TSP
Solving large-scale instances of the Traveling Salesman Problem (TSP) exactly is computationally expensive. Researchers often employ graph…
The Verifier is the Curriculum: Execution-Gated Self-Distillation for Cross-Family Game Generation
Post-training a code generator against a learned judge can optimize proxy features that raise the score without improving the artifact. We…
Closed-Loop Control with Rule-Aligned Small Language Models and Multi-Agent Self-Correction
A key step toward autonomous industrial operation is the ability to create and reconfigure control policies from natural-language requireme…
Feedback-Coupled Memory Systems in Continuous Time
The Feedback-Coupled Memory Systems (FCMS) architecture formalizes closed-loop coordination through four abstract operators, two of which -…
AGM-like Paraconsistent Partial Meet Abductive Expansion Operation
In his 1996 doctoral thesis, Maurice Pagnucco created the first AGM-like abductive expansion operation. Taking his operation as a basis, as…
Coresets Before Score Sets: Evaluation-Unsupervised Prompt Subset Selection for LLM Benchmarks
We study LLM benchmark coreset selection: selecting a small subset of prompts over multiple benchmarks whose induced model scores and ranki…
A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and Trajectory Prediction of Multiple Interacting Vehicles
Safe motion planning in advanced driver-assistance systems and autonomous vehicles requires an accurate understanding of how the surroundin…
Scaffolding the Strategist: Architecture-Dependent Reasoning Interventions in Hotelling Spatial Markets
We investigate whether structured reasoning interventions improve the strategic economic reasoning of large language models, and whether th…
A Theory of Least Autonomy in AI
Least privilege, the principle that an identity should hold only the permissions strictly required for its task, has been a foundational pr…
SupplyNetPy: An Open-Source Python Library for High-Fidelity Modeling and Simulation of Arbitrary Supply Chain and Inventory Networks
This paper introduces SupplyNetPy, an open-source, well-documented Python library for modeling and discrete-event simulation of supply chai…
Replicating Belief, Not Bits: Epistemic State Replication for Agentic Systems
In distributed systems, the classical State Machine Replication (SMR) model assumes that correct replicas execute deterministic transitions…
Task-Conditioned Synthetic Data Generation for Improving Machine Learning Performance in Agricultural Prediction Tasks
Machine Learning (ML) algorithms have been widely used to estimate agricultural variables across diverse contexts. However, because the qua…
LegalFarePlan: A Label-Setting Framework for Fare-Transparent Urban Rail Route Planning under Non-Additive Fare Rules
Urban rail fare systems may be non-additive: the fare of a single paid journey from an origin to a destination can differ from the sum of f…
BatteryLake: Agentic, Physics-Grounded Curation of Heterogeneous Battery Aging Data and Benchmarking
Public battery aging datasets are a critical asset for advanced health management, but their practical use is often limited by inconsistent…
How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm
A cheap swarm of unreliable agents can be steered to a correct consensus by a few strong, expensive "oracle" correctors. We ask how much on…
Norm Enforcement for AI Agents: Robustly Shaping Behavior in Multi-Agent Systems
AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards. This multi-agent compet…
Verification of Adaptive Agentic Controllers through Finite Rule Revision
Industrial agentic AI systems increasingly exhibit a gap between prototype capability and production deployment. In particular, adaptive ag…
EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents
Computer-use agents must solve long-horizon tasks through repeated interaction with partially observable, multimodal desktop environments.…
From Patterns to Maze Structures: SMT-Based Path Synthesis and 2D/3D Construction
We present a pipeline for constructing maze structures from input patterns such as text or shapes. The central path-synthesis problem is en…
Length Penalties Make Chain-of-Thought Less Monitorable
Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In…
PHITSBench: an execution-scored benchmark for AI-assisted PHITS radiation-transport input generation using natural language
We introduce PHITSBench, an execution-scored benchmark for the Monte Carlo Particle and Heavy Ion Transport code System (PHITS). PHITSBench…
Semantic Drift and the Stability of Operator Control in Reasoning-Class Decision Support Systems
The article investigates the fundamental problem of ensuring the stability of operator control and preserving goal-targeting in hybrid huma…
Agentic Context Learning with Self-Discovered Specification
Context learning is an emerging inference-time task where LLMs must learn and apply novel, task-specific knowledge from intricate contexts…
Exploring Agentic Workflows for Generating High Quality Math Visual Aids
Mathematical diagrams play a crucial role in K 12 education, both as problem components and as scaffolding for student comprehension. Howev…
TopoExplore: Topological Discrimination for Archive-Based Exploration
Archive-based exploration methods such as Go-Explore select which visited state to return to using visitation rarity, and frontier methods…
Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?
Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures becom…
A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution
Neural networks can learn algorithmic input-output mappings, but trusting a learned executor requires more than a correct final answer beca…
AgentAbstain: Do LLM Agents Know When Not to Act?
Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus o…
From ambiguous utterances to governed reuse classes: canonicalization, quotient invariance, and conditional decidability
Semantic caching defines answer reuse on embedding similarity: two utterances share a stored answer when a similarity score clears a thresh…
MAG: A Web-Agent Benchmark and Harness for Multimodal Action and Guide Generation
Digital Adoption Platforms (DAPs) are embedded overlays widely used on web systems to guide users through operations inside a page, helping…
Looped State-Space Language Models with Adaptive Exit-State Selection
Recent work on looped language models suggests that many reasoning problems benefit from greater computational depth rather than from addit…
Dynamic Agent Skills: A Lifecycle Survey and Taxonomy of Evolving Skill Libraries
Large language model agents increasingly store reusable procedures outside the model. These reusable procedures are often called \emph{skil…
IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation
Scientific research is a complex, multi-stage workflow rather than a single act of text generation. The ideation process typically emerges…
UNIT: Unleash Large Language Models Potential for Graph Continual Learning
In real-world multimodal web scenarios, graph-structured data often arrives in a streaming manner, making graph continual learning a crucia…
GRATE: Temporal Extensions for Inductive KG Foundation Models via Gated Rotary Attention
Knowledge graph foundation models such as Ultra and Trix achieve strong inductive transfer by learning relation-graph representations that…
KGCQual: An Interpretable Framework for Evaluating the Knowledge Graph Construction Quality from Text
Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious…
When Are Sparse Feature Interventions Actually Localized? Matched Evaluation for SAE-Based Safety Control
We evaluate when sparse autoencoder (SAE) features act as localized control handles for safety-relevant behavior. This question is difficul…
Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models
As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncert…
Information-seeking failures of large language models in agentic clinical reasoning
Large language models achieve high scores on medical knowledge assessments, yet clinical reasoning requires actively deciding what to inves…
Can Agentic Trading Systems Pay for Their Own Intelligence?
Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur…
SPARK: Susceptibility-Guided Profiling and Steering of Latent Reasoning States in Large Language Models
Reasoning failures in large language models (LLMs) are usually evaluated from final answers, but a wrong answer does not reveal why the mod…
Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems
Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Thi…
Comparing Socio-technical Design Principles with Guidelines for Human-centered AI
Human-centered AI (HCAI) refers to guidelines or principles that aim on ethi-cally oriented design of systems. We compare HCAI- guidelines…
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general…
Co4ICF: Co-evolving Physics-Informed Surrogate and RL-based Pulse Optimizer for Inertial Confinement Fusion
Offline-trained surrogates for Inertial Confinement Fusion (ICF) suffer a well-known failure mode that iterative optimizers drive inputs in…
ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm
Autonomous CLI agents can now execute hundreds of actions across multi-hour sessions: writing code, executing shell commands, browsing the…
GRASP: GRanularity-Aware Search Policy for Agentic RAG
Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries,…
Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents
Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conv…
Cross-Layer Misalignment Detection in Agent Skills: A Progressive Loading-Aware Contrastive Learning Approach
Large language model (LLM) agents are increasingly extended through Agent Skills, reusable artifacts that package natural-language metadata…
AI YOU Town: Make Friends and Money with Your Digital Twin
Existing approaches to infer user traits and generate responses consistent with a persona rely on static prompting. They lack calibrated un…
Large language model agents accelerate inverse design of metal-organic frameworks for gas separation
Metal-organic frameworks (MOFs) offer a highly modular platform for adsorptive gas separation, yet their vast reticular design space makes…
CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps
Evaluating the multi-hop reasoning capabilities of large language models remains a significant challenge. Although current models achieve s…
Laguerre Geometry for Interpreting Large Language Models
Existing hypotheses represent a concept in an LLM as a single point, a linear direction, or a Gaussian cluster, yet it remains unclear how…
Constraint-Aware Hierarchical Search for Regulation-Driven Fine-Grained Classification
Tasks such as customs tariff classification, export control categorization, and standards-based equipment coding require assigning an input…
MRUF: Multi-granularity Routing with Uncertainty-Aware Fusion for Robust Multimodal Sentiment Analysis
Multimodal sentiment analysis relies on language, visual, and acoustic cues, but utterance-level modality quality may vary due to occlusion…
Agentic-DPO: From Imitation to Agentic Policy Optimization on Expert Trajectories
Large Language Model (LLM) agents are commonly trained from expert trajectories using supervised fine-tuning (SFT), which treats multi-turn…
The Compliance Trap: Diagnosing How AI Agents Consume Conflicting Memory
Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, softwa…
Embark Now: User Demand Oriented Framework for Multi-day Urban Travel Itinerary Planning
In large urban areas, planning multi-day travel itineraries is challenging due to the abundance of Points of Interest (POIs), diverse user…
Personalized Emotional Intelligence in Generative AI through Symbolic Affective Reasoning
Emotional intelligence enables humans to recognize emotions, infer their causes, reason about interventions, and modify their environment t…
WattCouncil: Context-Aware Household Energy Scenario Generation With Governed LLMs
The accelerating shift toward low-carbon power systems, together with the widespread adoption of behind-the-meter technologies such as roof…
Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF
Synthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models are incre…
Opti-Agent-Bench: Benchmarking End-to-End Optimization R&D Agents on Real-World Business Problems
LLM-based agents are increasingly deployed to solve optimization problems, yet existing benchmarks evaluate them on pre-structured mathemat…
Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging
Computational imaging, which recovers hidden signals from indirect, noisy measurements, underpins quantitative discovery across scientific…
STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA
In open-domain multi-hop question answering (QA), LLM-based search agents offer a promising approach to knowledge-intensive QA by combining…
Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning
Multi-agent ensembling multiplies active parameters and inference cost without answering three basic questions: which agents to consult, ho…
Toward Contemplative LLM: A Modular Framework for Evaluating and Enhancing LLM Alignment in Mental Health
Contemplative traditions have long guided ethical behavior and prosocial interaction, and recent work suggests that contemplative principle…
LOGOS: A Living Logic for AI Agent Teams That Evolve With Humans
AI agents are evolving from answer engines into persistent teams that use tools, delegate work, learn from experience, and modify the artif…
First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint)
We extend, in Isabelle/HOL, the deep-and-shallow embedding methodology of our prior work from propositional to first-order modal logic (FML…
SETA: Scaling Environments for Terminal Agents
Large language models (LLMs) are rapidly shifting toward agents that solve tasks through diverse interfaces, including web and graphical us…
Incremental Transformer for Surrogate-Based Inverse Design of Geopolymer Mixtures
Small-data inverse design is challenging in engineering informatics when observations are heterogeneous, mixed-type, and constrained by phy…
Learning Linear Temporal Specifications from Demonstrations with Uncertainty
Learning temporal logic specifications from system demonstrations is essential for tasks such as formal verification and controller synthes…
SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning
We introduce Self-Verified Reasoner (SVR-R1), a multi-turn RL framework that turns a model's own verification into a learning signal for mu…
From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth
Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating lo…
QwenPaw-Data: Bridging Facts, Methodology, and Execution for Autonomous Enterprise Data Analytics
Enterprise data analysis is emerging as a distinct frontier for autonomous agents. Compared with general-purpose interaction and software e…
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
Despite progress in Embodied AI, Vision-and-Language Navigation systems remain vulnerable to adversarial visual disturbances. Most existing…
Are LLMs Ready for Scientific Discovery? A Capability-Oriented Benchmark for AI Scientists
Existing benchmarks for scientific data analysis evaluate LLMs primarily on code execution or workflow completion, overlooking that scienti…
NVAITC AI Scientist: A Governed End-to-End Research System -- A Hypertension GWAS Case Study
Agentic research systems are emerging as a new paradigm for coordinating scientific workflows beyond isolated model inference, code generat…
OS-Pruner: Pruning Chains-of-Thought of Reasoning Models via Optimal Stopping
Large Language Models (LLMs) have achieved remarkable success in complex reasoning tasks through Chain-of-Thought (CoT) prompting. However,…
A Formal Hierarchical Architecture for Agentic Orchestration with Stack-Based Execution and Lazy Discovery
The rapid expansion of capabilities in Large Language Model (LLM) agents has exposed a critical architectural bottleneck: when agents are g…
NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management
Large language models (LLMs) based agents are beginning to participate in portfolio construction and market analysis, where decisions must…
The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation
LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, inc…
STAMP: Provenance-Guided Credit Assignment for Deep Search Agents
Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware reward…
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
The growing ability of large language models and vision language models to jointly interpret and reason over images and text is reshaping m…
SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
Computer use agents (CUAs) are emerging as a powerful interface for automating complex digital workflows through visual perception and GUI…
What We Talk About When We Talk About LLM Planning: Evidence for Two Distinct Planning Abilities
When LLMs exhibit uneven performance across planning tasks, these gaps are often attributed to task difficulty. We argue that this explanat…
PREF-Gate: Provenance-Constrained Relational Evidence Fusion with Validation-Gated Selection for Graph Fraud Detection
Relational fraud detection can exploit both label-free graph context and label-derived neighborhood evidence, but these two information sou…
Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory
LLM agents today are caught in an awkward bind. Lock them down with static safety instructions and they rarely venture beyond the obvious;…
Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans
Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory c…
Valid $\ne$ Necessary: Diagnosing Latent Inefficiency in Chain-of-Thought
Chain-of-Thought (CoT) prompting has significantly advanced the reasoning capabilities of Large Language Models (LLMs), yet it often incurs…
Efficient Test-Time Optimization for Multi-Agent Proof Autoformalization
Full-proof autoformalization bridges extensive mathematical proofs in natural language with formally validated reasoning, offering a pathwa…
Calibrated e-CUSUM Decoding for Quantized Reasoning Models: Why Token Log-Probability Is the Wrong Observable for Decoding Monitors
Low-bit quantization makes small reasoning models inexpensive to deploy but can degrade their chains of thought. This motivates decoder-sid…
Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation
Large language models can produce superficially legal twelve-tone scores that collapse into degenerate textures. We introduce a neuro-symbo…
AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic
Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is chall…
Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents
Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable…
From Neural Network Decisions to Training Cases: An Exact Account via Case-Based Decision Theory
Neural networks increasingly guide decisions in high-stakes domains such as medical diagnosis, credit approval, and energy bidding. Audit i…
OpsMem: Dual-Memory Reasoning with Cross-Memory Resonance for Failure Diagnosis
Failure diagnosis in modern software systems requires iterative evidence acquisition and hypothesis reasoning guided by operational experie…
StructAgent: Harness Long-horizon Digital Agents with Unified Causal Structure
Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled increasingly capable digital agents for comp…
Omni-Decision: A Progressive Evidence-State Agent System for Omni-Modal QA
Omni-modal evidence-seeking QA requires agents to answer questions whose evidence is sparsely distributed across videos, audio, images, web…
The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning
Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it en…
Enhancing Query Efficiency for d-DNNF Representations Through Preprocessing
In this paper, we investigate preprocessing techniques aimed at improving the efficiency of accessing models of propositional formulas repr…
Comparative Analysis of GAT and BERT for Human-Like Playtesting
Accurately modeling and understanding player experience is crucial for designing engaging puzzle games. To achieve this, a common approach…
Learning Residual Kinematic Corrections for Continuous Neural Decoding via Reinforcement Learning
Decoding continuous three-dimensional (3D) motor imagery (MI) using non-invasive electroencephalography (EEG)-based brain--computer interfa…
HCRMap: Pressure-Aware Hot-Expert Residency Mapping for 3.5D MoE Chiplet Inference
Mixture-of-Experts (MoE) large language models (LLM) activate only a small number of experts during inference, but token routing introduces…
MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models
Multi-scene navigation (clearing an objective in one bounded space and then crossing a portal into the next) is a defining feature of conte…
Interaction Scaling: Grounding the Third Axis of Test-Time Compute
There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one. Both share…
Auditing the Risk Claims of Distributional Reinforcement Learning
Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability…
Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns
Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language mode…
Reproducing human biases in route choice using large language models: Toward scalable behavioral modeling
Human choice behavior, including route choice, exhibits systematic behavioral biases that deviate from the assumptions of full rationality.…
Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction
Self-refinement often fails to strengthen few-shot inductive reasoning in large language models. Prompting a model to explicitly state its…
Playful AI in Professional Email: A Field Experiment on Tone and Recipient Engagement
Large language models (LLMs) are rapidly reshaping workplace communication, yet whether AI-assisted writing changes how recipients actually…
Reverse Engineering Compliance: A Dual-Graph Verification Framework for Auditing Legacy IT Security Concepts
The NIS-2 Directive increases the need for continuous, auditable compliance evidence and motivates a shift from document-based compliance t…
Knowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey
Graph Neural Networks (GNNs) have emerged as a powerful paradigm in Knowledge Graphs (KGs) due to their intrinsic ability to model graph-st…
ECG-LDC: A Hardware-Efficient Low-Dimensional Computing Framework for ECG Arrhythmia Classification
Continuous cardiac monitoring in wearable devices demands classifiers that are simultaneously accurate, energy-efficient, and deployable on…
Ablation, Statistical Inference, and Validation for KV-Cache Compression
This study systematically compares Turbo-Quant and SpectralQuant KV-cache compression, evaluating non-dominated schemes, including WHT rota…
SciML in the Wild: A Diagnostic Study of When Structural Priors Help and When They Hurt
Scientific Machine Learning (SciML) methods such as Neural Ordinary Differential Equations (NODEs), Physics-Informed Neural Networks (PINNs…
Transfer Learning Across Policy Regimes in Adaptive Multi-Agent Systems
Policy models often assume that the relationship between a policy instrument and its outcome remains stable across institutional conditions…
What Context Does a Coding Agent Actually Need to Act?
A modern coding agent can hold an entire repository in its context window. Most of its reading is wasted -- and the interesting question is…
Depth-Entropy Guided Sampling for Training-Free LLM Reasoning
Reinforcement learning (RL) has become the dominant paradigm for improving the reasoning capabilities of large language models, but it requ…
Mitigating Early Training Collapse in CTR Models
Deep neural models for click-through rate prediction often exhibit a sharp decline in validation performance immediately after the first tr…
Model Collapse: On Recursion, Noise, and Uncharted Machine Visions
Since 2023, computer scientists have warned against model collapse -- the contamination of training sets with AI-generated outputs that pro…
The Ramanujan Challenge For AI
To help evaluate the mathematical skills of current AI systems, we present a set of formulas for fundamental mathematical constants. These…
The Universal Language of CSI:Unifying Wireless Sensing Across Devices and Environments
WiFi sensing based on Channel State Information (CSI) promises ubiquitous, device-free perception, yet current research remains trapped in…
SWIFT: A Small-World Interaction Framework for Flow-Aware Trajectory Prediction in Autonomous Driving
Accurate trajectory prediction in autonomous driving hinges on modeling dynamic and context-dependent interactions among traffic agents. Ho…
MorphologyFM: A Foundation Model for Morphology-Aware Representation Learning from ECG and Pulse Oximetry Waveforms
Foundation models have recently emerged as a powerful paradigm for learning transferable representations from large scale biomedical data,…
Data-Driven Forward and Inverse Modeling of V-Beam Thermal Sensors
This paper presents a machine learning framework for data-driven inverse design of V-beam thermal sensors. The goal is to determine the opt…
Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis
Diffusion models have achieved remarkable success across diverse domains, with performance closely related to the denoising backbones that…
Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal Tokenization
Large-scale, multi-subject widefield calcium imaging provides unprecedented access to brain-wide cortical dynamics. However, the high dimen…
RSLoRA: Training-free Rank Allocation for LoRA via Representational Sensitivity Probing
Low-Rank Adaptation (LoRA) has become a cornerstone of parameter-efficient fine-tuning (PEFT); however, the conventional practice of unifor…
ReflectWorld-MM: An Entity-Oriented Multi-Media Memory System for Open-Ended Video Streams
Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-st…
Physics-Informed Structure Anchoring With Capture-Aware Prototype Calibration for Cross-Environment RF Fingerprinting
Radio frequency fingerprint identification (RFFI) uses transmitter-specific hardware imperfections as a physicallayer identity cue for Inte…
Knowledge-Constrained Shape Optimization with a Mixture-of-Experts Neural Operator for High-Confidence Design
Engineering shape optimization faces challenges in both expert-dependent problem setup and surrogate-model reliability. In practical aerody…
OmniSCS: Omni Safety-Critical Scenario Synthesis for Autonomous Driving via a Fully Editable Driving World
The synthesis of safety-critical scenarios (SCS) and their evaluation through closed-loop simulations are crucial for developing robust aut…
Listen to the Features: Voice Anonymization Driven by Content Embedding Matching over Signal Reconstruction
The paper presents a voice anonymization model focusing on preserving content rather than producing realistic speech. It relies on content…
Maximizing Human Efficiency in Large-Scale Robot Post-Training via VLAC-Cut Guided Pipeline
When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post training are required because a single round…
Lifelong Representations: A Survey on Continual Self-Supervised Learning for Vision Models
Traditionally, continual learning has assumed access to labeled data, yet many real-world applications -- such as lifelong robotics -- requ…
A Comprehensive Survey and Systematic Real-World Evaluation of Embodied Vision-and-Language Navigation
Navigation is a fundamental capability of autonomous systems, yet most existing approaches rely on highly structured models and strong prio…
Large Multimodal Model-Based Environment-Aware Mobility Management
Recently, large language models (LLMs) have been successfully adopted in various fields, including wireless communications, robotics, and a…
JEPA for AI-Native 6G: Predictive Representations and Open Challenges
Sixth-generation (6G) networks are moving toward AI-native operation, where learning modules are embedded across the radio access network (…
Trivial Prompt Reframing Bypasses Safety Guardrails in Google\'s MedGemma-4B
Open-weight medical language models are increasingly used as the base of patient-facing and clinician-support applications. Their model car…
A Unified Model for Highly Accurate ECG-Free Dynamic Coronary Roadmapping Using Spatio-Temporal Transformers
Percutaneous Coronary Intervention (PCI) is a minimally invasive procedure used to restore coronary blood flow obstructed by atheroscleroti…
An Autonomous Scientific Knowledge Generation Framework for AI-Driven Scientific Discovery
Artificial intelligence (AI) is transforming scientific discovery, but its effectiveness is fundamentally limited by the availability of st…
TS-Mask VLA: 2D Temporal-Spatial Masking for Vision-Language-Action Model with Effective Bridging
Vision-language-action (VLA) models aim to understand natural-language instructions and visual observations, and to generate and execute co…
Memory-Conditioned Tool Calling for Camera-First Visual Agents
Recognition tells an agent what is in an image; personal memory affects what is worth looking up next. In a camera-first setting the user c…
More Structure, Not More Capacity: Object-Centric Representations for Visuomotor Imitation Learning
Robotic manipulation policies rely on pre-trained vision models that give either a global scene embedding or a dense patch grid. Both mix t…
Towards Objective Dysgraphia Detection: A Multi-Branch Deep Learning Approach for Online Handwriting Analysis
Dysgraphia is a specific learning disability that is prevalent among school-age children. It affects handwriting coherence, quality, fluenc…
Robo-ValueRL: Reliable Value Estimation for Offline-to-Online Reinforcement Learning
Offline-to-online reinforcement learning is promising for generalizable robotic manipulation, yet its full-stack complexity obscures reprod…
Prompting-MammAlps: Fine-Grained Text-to-Video Retrieval for Camera-Trap Data
Automatically retrieving videos from large camera-trap datasets remains challenging. Text-to-Video retrieval (TVR) methods based on large v…
CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series
Clinical time series are central to patient monitoring, risk assessment, and clinical decision support. However, they are often sparse, irr…
Remembering Distinct Items, Not Tokens: A Learnable Dirichlet-Process Cache Between State-Space Models and Attention
Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimen…
What You Train Is What You Get: Gender Bias, Training Composition, and Post-Hoc Mitigation in Audio Deepfake Detection
Audio deepfake detection models determine whether speech is genuine or artificially generated, but high overall accuracy can mask substanti…
Next-Dense-Stride Prediction for Multimodal Autoregressive Visual Modeling
We introduce DenseAR, a new generative paradigm that reformulates autoregressive image generation as coarse-to-fine next-dense-stride predi…
Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code
Agentic coding tools are increasingly used to make autonomous repository-level changes to real-world projects. Prior work has largely evalu…
Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences
Large language models are increasingly used to summarize clinical trial results for healthcare providers, patients, and payers, but their t…
SMETA-ZSL:Semantic Meta-Alignment for Zero-Shot Threat Classification
Cybersecurity systems must adapt rapidly to emerging threats. However, labeled data for new threat categories is unavailable when those thr…
A Knowledge-Based Multi-Agent Framework for Security Control Recommendation
Hardening IT on-premises environments can be a daunting task for teams without access to adequate cybersecurity expertise. In this regard,…
A Foundation Model for Multimodal Event Sequences in Financial Applications
Predictive modeling is a core component of modern financial services, where a wide range of tasks are traditionally addressed using separat…
Workload-Driven Optimization for On-Device Real-Time Subtitle Translation
This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-…
Learning in Curved Weight Space:Exponential-Linear Weight Reparameterization for Improved Optimization
Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but…
A Production-Oriented Framework for Evaluation of SFX Generation
Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity…
An LLM-powered Agentic Recommendation System for Connected TV Content Discovery
Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse c…
Beyond Bayesian Nash: Learning Minimax-Regret Equilibria for Adversarial Team Games under Asymmetric Information
Adversarial team games (ATGs) with asymmetric information, such as adversarial path-finding, goal search, and reachability games on graphs,…
Robust, Scalable Detection of Text Containment in Large Web-Crawled Corpora
We present FindMyText, an open-source Python package designed to efficiently assess whether a given text appears, in part or in full, withi…
Geometric mean-based pairwise comparison method with the reference values -- statistical approach
For many years, the pairwise comparison method has been widely used for decision-making involving experts. The best-known example of this m…
Quantum Circuit Vision: Cost-Aware Evaluation of Visual AI Agents for Quantum Code Generation
Can AI agents visually comprehend quantum circuit diagrams and generate verified executable code--and at what cost? We present Quantum Circ…
Efficiently Adapting Spoken Language Models for the Singaporean Context
Spoken language models (SLMs) unify speech perception and reasoning, but adapting them to sensitive domains is underexplored, especially wh…
Adaptive Model Compression (AMC): Saliency-Driven Resource Allocation for Ultra-Low-Power Transformer Inference
Deploying large-scale transformer models on resource-constrained edge devices remains a challenge due to the high energy and memory overhea…
Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety
Safety alignment in large language models remains brittle across languages: prompts reliably refused in English can elicit harmful complian…
Cost of Reasoning in non-English Languages: A Case Study on Japanese
Reasoning Language Models (RLMs) achieve their strongest performance when they reason in English, the language for which reasoning-oriented…
When Data Imbalance Helps: Robust Generalization Through Shortcut Saturation
We study robust generalization under spurious correlations: tasks where a shortcut feature is correlated with the true label in training bu…
A Large-Scale Dataset of MCP Implementations on GitHub
The rapid emergence of the Model Context Protocol (MCP) has introduced a new standard for connecting large language models to external tool…
ML in a Box: Analyzing Containerization Practices in Open Source ML Projects
Containerization has become increasingly essential in the machine learning (ML) domain, providing reproducibility, portability, and environ…
GAE: Graph-Augmented Evolution for Scientific Discovery via Reinforcement Optimization
Evolutionary program search guided by Large Language Models (LLMs) has emerged as a powerful paradigm for automated scientific discovery. H…
SALT-GNN: Handling Dense Neighborhoods in Anti-Money Laundering Graphs via Statistics-Aware Attention
Money laundering threatens financial stability and exposes institutions to penalties, motivating automated detection. Because laundering sc…
LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each…
FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion
Existing optical flow methods broadly follow two paradigms: iterative optimization and diffusion-based estimation. Iterative methods, exemp…
EmoStyle: Affective Conditioning of Style-Specialist Experts for Emotional Image Generation
Emotion-aware artistic image generation requires an image to match the input prompt, follow the specified artistic style, and convey the ta…
Transcript-Free Lightweight Detection of Alzheimer's Disease from Spontaneous Speech Using Handcrafted MFCC-Dominant Acoustic Biomarkers
It is still hard to find Alzheimer's disease (AD) early, especially when neuroimaging is expensive or tools that depend on language are not…
Beyond Euclidean Clipping: Overcoming Exploration Collapse in LLM RL via Riemannian Isometric Policy Optimization
Reinforcement learning (RL) has become a dominant paradigm for enhancing LLMs' reasoning capabilities. However, RL algorithms with PPO-Clip…
ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception
We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception…
Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory…
PhysMRV: Physical Memory Retrieval and Verification for Physics Plausibility Reasoning
Video-language models (VLMs) have achieved remarkable performance on video understanding and visual question answering, yet they remain unr…
Breaking the Quality--Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization
Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality--intelligibility trade-off, wherein naive optimi…
Instruction Set and Language for Hypergraphs
We present IsalHG, a method for representing the structure of any finite, connected hypergraph of bounded hyperedge arity as a string over…
Comparing Socially-Equitable Renewable Energy Budget Allocation MDP Policies in Mature and Emerging Economies
Equitable renewable-energy planning is a sequential decision problem, but the decision variables available to a public planner differ sharp…
When Does Depth Survive Composition? Compute--Quality Regimes in Latent World Models
Adaptive-compute world models -- early-exit or mixture-of-depths predictors that spend variable depth per step -- assume depth buys better…
Source-Lifted Flow Matching for Intervenable Multimodal Imitation
Flow-matching policies are promising for imitation learning because they model complex multimodal action distributions. However, their stoc…
Exploratory Analysis of Deep Learning Models for Forecasting Meteorological Parameters in the Agricultural Sector
Accurate meteorological forecasting is essential for agricultural planning, irrigation management, and environmental decision support. This…
PhenoEmbed: Self-Supervised Multispectral UAV Time-Series Embeddings for Individual Tree Crown Phenology
Tree crowns are a challenging target for resilient AI because they are not static objects: their spectral response, internal texture, trans…
Partial Contracts Suffice: Sound, LLM-Inferred Regression Verification
Software evolves continuously, yet ensuring that a patch preserves intended behavior without re-verifying an entire codebase remains diffic…
Program-Synthesis-Driven Autodesign of Universal Unitary Operators
We demonstrate that AI-driven program synthesis can autonomously discover fundamental strategies for decomposing unitary matrices in photon…
Polarization Detection: A Hybrid Approach with AfroXLMR-Social and DeBERTa for Low- and High-Resource Settings
The rapid proliferation of online polarization threatens social cohesion, necessitating robust automated detection systems that operate eff…
Neutralizing Structural Inequality in the Nigerian FinTech Sector
Algorithmic decision systems in financial services often rely on data proxies that inadvertently encode structural inequalities. This paper…
From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement
AI visibility measurement is comparative: practitioners want to know which domains generative search engines cite most often and whether ob…
GRC-ProbNet: Uncertainty-aware Feature Extraction for Cardiovascular Disease Classification
The automatic detection and classification of cardiovascular disease (CVD) from computed tomography (CT) images plays an important role in…
Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
Foundation models are increasingly used as image feature extractors for mammography, but their robustness under external domain shift remai…
A Hyperbolic Neural Closure for M1 Radiation Transfer
In radiation transfer simulations, an M1 method achieves substantial computational savings by replacing the full angular transport equation…
VINE: Taming Generative Control Policies for Reinforcement Learning
Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from nois…
ABot-N1: Toward a General Visual Language Navigation Foundation Model
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse…
Structured Thoughts For Improved Reasoning And Context Pruning
Large language models (LLMs) excel at generating long chains of thought, but long reasoning traces are often verbose and memory-inefficient…
The evolution of AI from image interpretation toward scientific inference in nanoparticle electron microscopy
Artificial intelligence (AI) is transforming electron microscopy by enabling quantitative analysis of increasingly large and complex datase…
A Stepwise Questioning Expert-Editor Multi-Agent Framework for Long-Document Summarization
Although large language models (LLMs) have shown promising potential in news summarization tasks, their performance on long-document summar…
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
Vision language models (VLMs) have achieved strong performance on visual document understanding benchmarks such as DocVQA, ChartQA, and MML…
Large Language Models in Misinformation Ecosystems: Misuse, Defense, and Vulnerability
Large language models (LLMs) have transformed misinformation from a primarily content-centric problem into a broader ecosystem-level securi…
Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control
Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting user…
Mitigating LLM Sycophancy in Code Smell Detection Using Evidence-Guided Reasoning Prompts
Large Language Models (LLMs) are increasingly used for code smell detection tasks due to their ability to interpret program semantics. Howe…
Learning the Brain's Dynamics as a Port-Hamiltonian System
We model human motor cortex during a wrist-extension BCI task as a port-Hamiltonian system (pHS): a conservative interconnection (gyroscopi…
Context by Distinct Information: An Auditable Dirichlet-Process Working Memory for Long, Redundant Context Streams
Context engineering decides what information a model carries forward, and current designs meter it in tokens: compressing the past into a b…
Annotation-Free Furniture Codes: What They Encode, and How Far They Transfer
Layout-based 3D scene synthesizers place each object using two human-annotated channels: a categorical class label and a canonical-pose con…
Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards
Partial differential equations (PDEs) are foundational to modeling in science and engineering, but constructing reliable numerical solvers…
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process…
Temporary Authority, Permanent Effects: Commit-Time Authorization for LLM Agents
LLM agents can commit durable effects from authority evidence that was valid earlier in execution: a DOM snapshot, approval epoch, version…
Confining Nondeterminism: AI-Driven Research Systems as DBMSs for Reliable, Non-Wasteful, Transparent, and Collaborative Research [Vision]
LLM agents that conduct research (proposing ideas, writing and running code, analyzing results) can already carry a study from research que…
Conditional Optimal Bridge for Riemannian Activation Steering
Activation steering offers a lightweight alternative to fine-tuning for controlling large language models at inference time. While many exi…
Towards Autonomous and Auditable Medical Imaging Model Development
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugg…
Motif: Discovering and Automating Personal Web Workflows
Recent advances in LLMs and existing work on programming by demonstration have made it possible for end users to create automations by expl…
Tool-Adaptive LLM Reranker
Generative Large Language Models (LLMs) have revolutionized information retrieval, yet their strictly parametric nature frequently leads to…
When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation
Modern coding agents expose multiple tool surfaces -- IDE primitives, bash, and Model Context Protocol (MCP) code-execution -- and the fiel…
Learning from Local Walks on Dynamic Graphs with Bandit Feedback
We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In t…
DiffUE: Enhancing Utility-Unlearnability Trade-off of Unlearnable Examples via Diffusion Autoencoders
AI models are increasingly trained on personal images scraped from social media and public platforms, often without consent, leading to ser…
MemDecay: Region-Aware KV Cache Eviction for Efficient LLM Agent Inference
Large language model (LLM) agents accumulate heterogeneous context, including system instructions, plans, user turns, retrieved documents,…
WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen
Efficient waste segregation is critical for sustainable urban management and environmental governance. Existing automated systems are limit…
Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation
We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source, an…
World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning
Robust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepres…
Auditing Construct Overlap in Explainable Machine Learning: Evidence from Burnout-Depression Prediction Across Student Cohorts
Explainable machine learning (XML) pipelines applied to composite mental health outcomes can produce apparently-robust, cross-population-st…
Coverage Path Planning: Classical Foundations, Recent Advances, and Future Directions
Coverage path planning (CPP) is a fundamental problem in robot motion planning, whose aim is to produce robot trajectories that provide com…
Unlocking Parallelism in Autoregressive Language Models via Speculative Decoding with Progressive Tree Drafting
Speculative decoding has significantly accelerated Large Language Model (LLM) inference by alleviating memory-bound bottlenecks. However, t…
Answer-Conditioned Chain-of-Thought Distillation for Few-Shot Industrial Vision with Small VLMs
Deploying AI-based visual inspection in manufacturing is hard because requirements change often, new defect types appear, and large labeled…
Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications
As AI code tools become integrated into programming environments, students increasingly describe intended behavior in natural language and…
Learning to Fine-tune Foundation Models under Resource Limitations
We study the problem of optimal continual fine-tuning for a pre-trained Foundation Model deployed at a resource-limited device. At each tim…
Action Map Policy: Learning 3D Closed-loop Manipulation via Pixel Classification
The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently…
MDQEC-QAS: Meta-Decoding for Quantum Error Correction with Hardware-Aware VQC Search and Confidence-Gated Recovery
We propose a unified meta-decoding framework for quantum error correction that learns syndrome-to-recovery mappings across multiple stabili…
PromptGraph: Graph-Guided Prompt Sanitization for Balancing Privacy and Utility in LLM Inference
Large Language Model (LLM) services introduce a fundamental privacy challenge. Sensitive information may be inferred not only from explicit…
Distributed Denial of Science: How Indirect Data Poisoning of AI Systems Can Industrialize Scientific Fraud
Scientific fraud is the instrument of doubt that malicious entities can use to establish controversy in science. Historically, it required…
A Corpus of Persuasion Techniques in Slavic Languages
Persuasion techniques are powerful rhetorical devices used to sway public opinion in a wide range of media. We present a new corpus of pers…
To Answer or to Abstain: Mitigating Search-Agent Hallucinations via Abstention-Aware Reinforcement Learning
Recent advances in equipping Large Language Models (LLMs) with search tools and outcome-reward reinforcement learning (RL) have achieved ne…
Multi-Scale Convolution with Optimal Transport Attention Effect on Multivariate Time Series
The analysis of Multivariate Time Series (MTS) plays an important role in a lot of real-world practical applications, but it still remains…
Lightning Fast Matching Dependency Discovery with Desbordante
Matching dependency is a generalization of the functional dependency concept, which allows users to apply custom similarity functions for m…
LSTrans: Efficient Knowledge Transfer for Lightweight and Automated ECG Classification
Deploying deep learning models for automated electrocardiogram classification on resource-constrained wearable devices remains challenging…
Weight-Adjusted Gradients Reveal Parameter Importance and Failure Modes in LLMs
Understanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and i…
Abstractiveness Metrics for Evaluating Text Summarization: A Refined Formulation with Empirical Validation
Quantifying abstractiveness in generated summaries is essential for evaluating summarization models beyond surface-level metrics like ROUGE…
Diachronic Sample Integration: Robust Tail-Risk Estimation with Generative Models
Deep generative models are increasingly used as simulators for downstream decision-making under data scarcity, but in risk-sensitive applic…
Distributed Agent System: Fault-Tolerant Collaboration Among Embodied Agents
AI engineering is shifting from passive text generation by large language models (LLMs) to agent-driven task execution, creating new reliab…
Auditing Belief-Conditioned LLM Agents in Hidden-Information Social Deduction Games
Evaluating LLM agents in hidden-information multi-agent settings is hard: final outcomes are high-variance and rarely reveal why an agent d…
Large Language Models for Token-Efficient and Semantic-Preserving Opinion Summarization
Opinionated text - spanning product reviews, hotel feedback, and social posts - captures rich signals about user experiences, preferences,…
3D-DefectBench: A Controlled Factorial Study of Vision-Language Model Evaluation Pipelines for Fine-Grained 3D Generation Defects
Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliabi…
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study
The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks wi…
The Nuts and Bolts of Natural Language to SQL Translation: A Systematic Analysis of Model Pipeline Optimisation Approaches and their Interactions
In the age of large language models, Natural Language to SQL (NL2SQL) translation remains an open problem with many useful applications. We…
The Singularity Space: A Generative Diffusion Framework for Signal Representation
Generative models often represent signals as dense grids of amplitudes, blurring sharp transients that are crucial for the correctness of p…
Edge Physical AI Deployment of Vision Transformers on Heterogeneous Edge GPU Targeting Autonomous Vehicles
Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringe…
Efficient Online Proportional Sampling with Applications to Smoothed Online Learning
We study the problem of efficient online proportional sampling from a high-dimensional domain under a $\sigma$-smoothed adversary, where th…
CGS: Configurable Graph Summarization with Bounded Neighborhood Loss and Query Support
Given a large graph, how to generate a compact summary graph that is configurable by the user and supports multiple graph queries with eith…
EquiFusion: Kinematics-Agnostic Human Motion Prediction via Equivariant Latent Diffusion
Existing Stochastic 3D Human Motion Prediction models are fundamentally constrained by hard-coding the skeleton kinematics, severely limiti…
MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI
Perineural invasion (PNI) is a critical prognostic factor in cholangiocarcinoma. Non-invasive prediction from 3D MRI is challenging, demand…
Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies
As mobile robots become more integrated into everyday human environments, social robot navigation is becoming essential for ensuring human…
LoSA-Net: A Localized and Scale-Adaptive Network for Boundary-Sensitive Prediction of Perineural Invasion in 3D MRI
Perineural invasion (PNI) is a clinically relevant indicator of tumor aggressiveness and can influence surgical decision-making, motivating…
Affordance-Based Manipulation Planning with Text Goals and Sim-to-Real Generalisation via Real-to-Sim Image Conversion
We present a manipulation planning system based on affordance recognition and action effect prediction. The system reasons through possible…
Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies
This paper develops a model-free reinforcement learning framework for continuous--time extended mean field control problems, where both the…
SynCLIP: Synonym-Coherent Language-Image Pretraining for Robust Open-Vocabulary Dense Perception
Open-vocabulary dense perception (OVDP) aims to localize objects unseen during training by leveraging textual knowledge. Despite the remark…
Same Stories, Different Journeys: From Social Comparison to Sensemaking in AI-Mediated Peer Career Exploration
Young job seekers frequently turn to social media to compare themselves with peers and make sense of career possibilities. However, passive…
BackendForge: Benchmarking Agentic End-to-End Code Generation with Backend Services
Large language models (LLMs) are increasingly used in agentic coding settings, where they can inspect files, execute commands, run tests, o…
Flout at Your Own Risk: LLMs Struggle with Pragmatic Cooperativity Under Epistemic Asymmetry
Fruitful collaborations rely on cooperative communications, including of contextual cues to incorporate into reasoning. The increasing use…
Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video
Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, i…
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to mot…
AgentCheck: A Reproduce-Intervene-Mitigate Workbench for LLM Agents over MCP
Tool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its descripti…
The Equilibrium Is the Initialization: Lazy Identity Collapse in Physics-Structured Deep Equilibrium Reasoning
Deep equilibrium models promise input-adaptive implicit computation: harder problems should demand more solver iterations, and the solved e…
MusicMark: A Robust Generative Watermarking Framework for Music Generation
AI music generation has rapidly advanced alongside commercial platforms, raising the need for reliable watermarking for provenance and attr…
VIA: Visual Interface Agent for Robot Control
Robot manipulation is a complex task that requires visual understanding, physical reasoning, planning, and closed-loop control. General-pur…
BeatEdit: Symbolic Music Generation as Explicit Editing
Music creation is fundamentally a process of revision. Yet symbolic music generation remains dominated by paradigms that produce complete s…
AMT-X: Phase-Structured Multi-Turn Red-Teaming with Checklist-Gated Evaluation
Safety evaluation of large language models (LLMs) relies largely on single-turn attack datasets and single-judge scoring, underestimating r…
Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation
Representing manipulation actions as 2D trajectories in the camera plane provides a compact and interpretable basis for learning complex 3D…
RepTran: Search-Based Repair of Transformer Models
To ensure the overall quality of AI-enabled software, not only traditional software components but also AI components need to be tested and…
ProgramTab: Boosting Table Reasoning of LLMs via Programmatic Paradigm
Table-based reasoning with large language models (LLMs), which requires reasoning based on natural language questions and structured tabula…
HandFlow: Fully Generative 4D Hand Recovery with Flow Matching
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce…
DeepBias: Adaptive In-depth Probing of Social Biases in LVLMs
While Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities, they remain highly susceptible to embedded social biases. E…
An Empirical Study for GUI Test Migration from Android to OpenHarmony System
To reduce the substantial engineering effort required to test the corresponding applications from Android to OpenHarmony, migrating existin…
Multi-Agent LLMs Fail to Explore Each Other
Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can…
Enhancing LLMs through human feedback: a journey towards self-improvement
In the rapidly evolving landscape of information retrieval systems, the ability to adapt and improve through user feedback is paramount. Th…
Towards Predictive, Aligned, and Scalable Robot Learning
Learning, at its core, extends beyond memorization to the ability to reason and solve novel problems by navigating a space of possibilities…
Automated Textbook Auditing with Multi-Agent LLM Systems
Ensuring the quality of educational materials requires more than standard proofreading: textbooks must be audited for factual accuracy, dom…
A Unified Framework for Comprehensive Cardiac CT Segmentation and Phenotyping: Human-in-the-Loop Data Annotation, Vision Foundation Model Development, Multicenter Evaluation and Clinical Validation
Comprehensive quantification of cardiac structures from computed tomography (CT) remains limited not by data availability but by the scalab…
Mako: A Self-Evolving Agentic Operating System (SE-AOS) for Autonomous Web Exploitation
We introduce the Self-Evolving Agentic Operating System (SE-AOS): a new class of AI agent that treats exploit capability as a mutable, vers…
The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students
As Large Language Models (LLMs) are increasingly deployed as conversational tutors, they risk institutionalizing systemic inequalities. Thi…
Programming Language Policy as an AI Literacy Equity Problem: A 15-Nation Comparative Analysis
The promise of AI literacy ``for all'' confronts a structural challenge embedded in how nations organise secondary computer science educati…
PRISM Edit: One Vector for All Temporal Answers
Model editing keeps large language models (LLMs) up to date without retraining, but temporal facts expose a limitation of the prevailing lo…
Fail-Aware and Explainable Test Oracle Prediction
Despite their central role in fault detection, test oracles remain challenging to construct effectively. Recent learning based methods addr…
Longitudinal Multi-View Breast Cancer Risk Prediction
Accurate breast cancer risk prediction from screening mammography is critical for enabling personalized screening intervals and early detec…
Understanding the Impact of AI Code Assistants on Security API Usage: An Empirical Study
AI code assistants are transforming software development, but their implications for software security remain a major concern, particularly…
Characterising AI Models for Cataloguing
The creation of digital collections involves not only the digitisation of content, but also the creation of catalogue records for it. This…
Beyond Sally-Anne: Evaluating Theory of Mind in LLMs using Epistemic Schelling Points
Text-based evaluations of Theory of Mind (ToM) in Large Language Models (LLMs) often involve cognitive tests akin to the Sally-Anne task th…
BackgroundMellow: A Multi-Modal Cohesive Framework for Narrative-Driven Rich Cinematic Soundscape Generation
Generating immersive, synchronized and cinematic audio for long-form textual narratives remains a significant challenge in multi-modal AI.…
A Glimpse into Long-term Physical Coexistence with Intelligent Robots
Long-term physical coexistence with intelligent robots requires more than capable robot policies. A persistent robotic assistant must suppo…
Agentic Routing: The Harness-Native Data Flywheel
Large language model agents are increasingly executed not by a single model call, but by an execution harness that manages observation, con…
Uncertainty Quantification for EO Regression Tasks: Building Height, Tree Canopy Height and Above-ground Biomass Estimation
Earth Observation regression tasks such as building height, canopy height, and above-ground biomass estimation underpin critical applicatio…
A Multimodal Dataset for Large Language Model Applications in the Energy Domain
This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications i…
LightMem-Ego: Your AI Memory for Everyday Life
Personal AI assistants on mobile and wearable devices continuously perceive users' daily lives through visual and audio streams. However, a…
Agentic Skill Optimization over Lie Algebroids
Agentic systems increasingly improve themselves by editing skills: prompts, rubrics, plans, tool contracts, examples, validators, and trace…
IG-GAN: A Generative Adversarial Network for Aerodynamic Data Generation Based on Intrinsic Geometry
Existing generative models learn data distributions in flat Euclidean space. However, most data in our real world are manifolds embedded in…
See like a Robot: Robot-Centric Pointmaps for Vision-Language-Action Models
Vision-language-action (VLA) models predict robot actions from visual observations and language instructions. These actions are defined in…
Proxy Exploration and Reusable Guidance: A Modular LLM Post-Training Paradigm via Proxy-Guided Update Signals
Post-training is essential for refining the domain-specific capabilities of large language models (LLMs), yet existing reward optimization…
CDFM: Towards a General-Purpose Causal Discovery Foundation Model
Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientifi…
Toward Inclusive Avatar Design with Limb Differences Through Artificial Intelligence
As extended reality becomes more popular for social interaction and entertainment, 3D avatars must represent the full diversity of body typ…
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a…
AutoMatBench: An Automatic Optimization Toolkit for the Acceleration of Material Properties Prediction Benchmarking
Material property prediction (MPP) infers key properties from chemical composition and structure, accelerating the discovery and optimizati…
Technical Report on the CVPR 2026@AdvML Workshop Challenge
Vision-language agents (VLAs) are increasingly used to interpret complex driving scenes and support safety-critical reasoning. This report…
Heuristic Learning for Active Flow Control Using Coding Agents
Active flow control involves nonlinear dynamics, partial observations, and computationally expensive simulations, making controller design…
Structure-Feature Aligned Graph Learning via Alternating Constrained Optimization
We introduce a constrained two-view framework for node prediction that aligns structure-conditioned GNN embeddings with a structure-free fe…
DiffEEG: A Self-Supervised Denoising Diffusion Model for Learning EEG Generic Representations
Deep learning for EEG-based seizure detection faces critical challenges: severe annotation scarcity and extreme class imbalance, where icta…
Extending LLM Context via Associative Recurrent Memory
Extending the context length of large language models (LLMs) is critical for many real-world applications, yet standard transformers remain…
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodi…
Closing the Loop: An Access-Control Architecture for Automated, Anomaly-Driven Network Revocation in IoT Deployments
Network-based anomaly detection for IoT devices has matured to the point of reporting strong detection accuracy, yet most published systems…
RAGU: A Multi-Step GraphRAG Engine with a Compact Domain-Adapted LLM
Graph retrieval-augmented generation (GraphRAG) enhances large language models with structured knowledge, yet existing systems construct kn…
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Artificial general intelligence ultimately requires agents that can reason and act in the physical world. Action models, vision-language-ac…
Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
Production LLM agents such as Claude Code and Codex operate over untrusted content, files, commands, and workspace state, making safety fai…
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Modern LLM-driven text-to-speech (TTS) and voice conversion (VC) systems produce synthetic speech that differs from the generators represen…
An Explainable Agentic System for Detection of Conversational Scams with Summary-Based Memory
Following the rapid progress of generative Artificial Intelligence, there is a growing threat posed by conversational scams. These scams of…
Active Offline-to-Online Reinforcement Learning
Background: Offline reinforcement learning (RL) enables effective policies to be trained from large, previously collected datasets and subs…
Time-Lag-Aware Deep Reinforcement Learning for Flexible Job-Shop Scheduling in PPVC Module Factories
Prefabricated prefinished volumetric construction moves most building work into module factories, whose production floor operates as a flex…
Evaluating RE Practices for Explainability: Synthesizing Insights from Daimler Truck into an Explainable RE Framework Proposal
Explainability has emerged as a critical requirement for AI-based systems, particularly in safety-critical and regulated domains. Although…
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
Long-form audio description (AD) requires more than describing visible actions: it must preserve characters, events, relationships, and sto…
Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models
Large audio-language models (LALMs) often underperform on fine-grained, non-semantic attributes of speech, such as a speaker's emotion, des…
Introducing Human-Centeredness in AI-Assisted Lexicography
This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers sign…
MM-ToolSandBox: A Unified Framework for Evaluating Visual Tool-Calling Agents
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a state…
Transformer-Guided Swarm Intelligence for Frugal Neural Architecture Search
Neural Architecture Search (NAS) has automated the design of deep learning models but traditionally requires massive computational resource…
LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments
This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthca…
Evidence-Backed Video Question Answering
Current Video Large Language Models (Video LLMs) excel in question answering (QA) but largely operate as black boxes, providing textual ans…
Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias
Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and…
A Minimalist Retargeting-Guided Reinforcement Learning Recipe for Dexterous Manipulation
Recent work in humanoid whole-body control has found success with a simple recipe: retarget human motion to robot kinematic references, the…
Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks
We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous…
Metacognition in LLMs: Foundations, Progress, and Opportunities
Metacognition is a foundational component of intelligence critical to effective learning, problem solving, decision-making, communication,…
Measuring AI Ability to Complete Long Software Tasks
Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of A…
Trojan Horse Prompting: Jailbreaking Conversational Multimodal Models by Forging Assistant Message
The rise of conversational interfaces has greatly enhanced LLM usability by leveraging dialogue history for sophisticated reasoning. Howeve…
LLM-Driven Collaborative Model for Untangling Commits via Explicit and Implicit Dependency Reasoning
Atomic commits, which address a single development concern, are a best practice in software development. In practice, however, developers o…
InqEduAgent: Adaptive AI Learning Partners with Gaussian Process Augmentation
Collaborative partnerships play a crucial role in inquiry-oriented education. However, most learning partners are currently assigned throug…
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
Despite rapid progress in artificial intelligence, current systems struggle with the interconnected challenges that define real-world decis…
Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs
Personalized large language models are often expected to follow explicit style instructions, yet we find that such instructions can undermi…
BizFinBench.v2: Towards Reliable LLMs in Finance via Real-User Data and Offline/Online Bilingual Evaluation
Large language models are becoming increasingly significant in financial applications. Nevertheless, prevailing benchmarks are largely depe…
FIRE-Bench: Evaluating AI Agents on the Rediscovery of Scientific Insights
Autonomous agents powered by large language models (LLMs) promise to accelerate scientific discovery end-to-end, but rigorously evaluating…
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide ri…
A Model-Free Universal AI
In general reinforcement learning, all established optimal agents, including AIXI, are model-based, explicitly maintaining and using enviro…
Where Experts Disagree, Models Fail: Detecting Implicit Legal Citations in French Court Decisions
Applying computational methods to law at scale requires separating genuine legal reasoning from surface similarity. We study this through a…
When Sensing Varies with Contexts: Context Probing for Tactile Few-Shot Class-Incremental Learning
Few-shot class-incremental learning (FSCIL) aims to recognize novel classes from only a few labeled samples while retaining previously lear…
Open, Reliable, and Collective: A Community-Driven Framework for Tool-Using AI Agents
Tool-integrated LLMs retrieve information, perform computations, and take real-world actions, but their reliability depends on both tool-us…
GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning
Competitive programming remains one of the last few human strongholds in coding against AI. The best AI system to date still underperforms…
Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs
We present the Bayesian Linguistic Forecaster (BLF), an agentic system for binary forecasting that achieves state-of-the-art performance on…
Algorithm Selection with Zero Domain Knowledge via Text Embeddings
We propose a feature-free approach to algorithm selection: instead of hand-crafted instance features, we use pretrained text embeddings. Ou…
Ideological Bias in LLMs' Economic Causal Reasoning
Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly…
Recursive Multi-Agent Systems
Recursive or looped language models have recently emerged as a new scaling axis by iteratively refining the same model computation over lat…
A Low-Latency Fraud Detection Layer for Detecting Adversarial Interaction Patterns in LLM-Powered Agents
Large Language Model (LLM)-powered agents demonstrate strong capabilities in autonomous task execution, tool use, and multi-step reasoning.…
2.5-D Decomposition for LLM-Based Spatial Construction
Autonomous systems that build structures from natural-language instructions need reliable spatial reasoning, yet large language models (LLM…
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
Embodied agents can benefit from skills that guide object search, action execution, and state changes across diverse environments. Since em…
Learning Developmental Scaffoldings to Guide Self-Organisation
From subcellular structures to entire organisms, many natural systems generate complex organisation through self-organisation: local intera…
MindClaw: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
Theory of Mind (ToM) enables an agent to reason about another actor's beliefs, goals, and intentions, which is essential for human-centered…
Severity-Aware Curriculum Learning with Multi-Model Response Selection for Medical Text Generation
Telehealth systems have become increasingly important for delivering accessible and timely medical information. Existing large language mod…
When Does Delegation Beat Majority? A Delegation-Based Aggregator for Multi-Sample LLM Inference
Majority voting is the default unsupervised aggregator for multi-sample LLM inference, but it discards two signals: within-group answer ent…
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into la…
NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning
Multimodal large language models (MLLMs) are increasingly deployed as embodied planners in egocentric environments, where task success requ…
Characterizing Large Language Model Agentic Workflows: A Study on N8n Ecosystem
Large Language Models (LLMs) are rapidly being adopted in low-code and no-code automation platforms, where non-expert users design workflow…
Safety from Honesty in a Disinterested AI Predictor
As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directe…
FARS: A Fully Automated Research System Deployed at Scale
Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts,…
Agentic generation of verifiable rules for deterministic, self-expanding reaction classification
Computer-assisted synthesis planning breaks target molecules into accessible precursors using large libraries of reaction rules that assign…
Separating Expert Retention from Autonomous Source Inference in Raw-ECG-Replay-Free Continual ECG Deployment
In multi-source ECG deployment, models may need to incorporate new data sources when earlier raw ECGs cannot be retained or replayed. Freez…
SUNTA: Hierarchical Video Prediction with Surprise-based Chunking
Hierarchical state-space models (HSSMs) offer a promising approach to long-horizon prediction by segmenting sequences into temporal chunks.…
Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces
When evaluator-derived step rewards are pooled or compared across scoring channels, their sign is treated as transportable. Yet the same fr…
MoP-JEPA: Hard-Assigned Predictor Mixtures for Stochastic JEPA World Models
JEPA world models commonly predict the next latent state with one regressor. Under stochastic transitions, squared and cosine regression re…
When do prophets profit in prediction markets?
Prediction markets aggregate dispersed beliefs into prices that act as probabilistic forecasts of uncertain events. Classical theory establ…
Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents
Tool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In poli…
Research on Cross-media Science and Technology Information Data Retrieval
Since the era of big data, the Internet has been flooded with all kinds of information. Browsing information through the Internet has becom…
Research on Intellectual Property Resource Profile and Evolution Law
In the era of big data, intellectual property-oriented scientific and technological resources show the trend of large data scale, high info…
Profiling and Evolution of Intellectual Property
In recent years, with the rapid growth of Internet data, the number and types of scientific and technological resources are also rapidly ex…
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
Retrieval-augmented generation (RAG) has rapidly advanced the language model field, particularly in question-answering (QA) systems. By int…
Constrained Reinforcement Learning for Safe Heat Pump Control
Constrained Reinforcement Learning (RL) has emerged as a significant research area within RL, where integrating constraints with rewards is…
Training on Irrelevant States Implies Data Augmentation: Generalization in Contextual MDPs
In the zero-shot policy transfer (ZSPT) setting for contextual Markov decision processes (CMDP), agents train on a fixed, finite set of con…
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
This paper explores the impact of occlusions in video action detection. We facilitate this study by introducing five new benchmark datasets…
Asynchronous Perception Machine For Efficient Test-Time-Training
In this work, we propose Asynchronous Perception Machine (APM), a computationally-efficient architecture for test-time-training (TTT). APM…
Training-Free, Identity-Preserving Image Editing for Fashion Pose Alignment and Normalization
Diffusion models have recently unlocked new possibilities in editing images of real-world objects. Yet, transforming objects in non-rigid w…
Disentangling Feature Structure: A Mathematically Provable Two-Stage Training Dynamics in Transformers
Transformers may exhibit two-stage training dynamics during the real-world training process. For instance, when training GPT-2 on the Count…
Hyperflux: Pruning Reveals Importance
Network pruning is used to reduce inference latency and power consumption in large neural networks. However, most methods focus on empirica…
Efficient Q-Learning and Actor-Critic Methods for Robust Average-Reward Reinforcement Learning
We study model-free methods for distributionally robust infinite-horizon average-reward Markov decision processes (MDPs). We present non-as…
Adaptive Reinforcement Learning for Unobservable Random Delays
In standard reinforcement learning (RL) settings, the interaction between the agent and the environment is typically modeled as a Markov de…
On the Necessity of Output Distribution Reweighting for Effective Class Unlearning
In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause inf…
Can Argus Judge Them All? Comparing VLMs Across Domains
Vision-Language Models (VLMs) are increasingly used in industry VLM applications such as retrieval systems, content generation platforms, a…
Interaction Techniques that Encourage Longer Prompts Can Improve Psychological Ownership when Writing with AI
Writing longer prompts for an AI assistant to generate a story increases psychological ownership, a user's feeling that the writing belongs…
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
The rapid advancement of Large Language Models (LLMs) in software engineering has revealed critical limitations in existing benchmarks, par…
CRINN: Contrastive Reinforcement Learning for Approximate Nearest Neighbor Search
Approximate nearest-neighbor search (ANNS) algorithms have become increasingly critical for recent AI applications, particularly in retriev…
Beyond Na\"ive Prompting: Strategies for Improved Context-aided Forecasting with LLMs
Real-world forecasting requires models to integrate not only historical data but also relevant contextual information provided in textual f…
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
Advanced reasoning in LLMs on challenging domains like mathematical reasoning can be tackled using verifiable rewards based reinforced fine…
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
Best-of-n sampling improves the accuracy of large language models (LLMs) and large reasoning models (LRMs) by generating multiple candidate…
TENET: One Step Toward Test-Driven Development for Repository-Level Code Generation
Test-Driven Development (TDD) is a widely adopted practice that requires developers to create and execute tests alongside implementation. W…
Graph Optimization Foundation Model: Tokenizing Graph via A Language-Model Paradigm
The pretrain-transfer paradigm, which underpins the success of large language models (LLMs), has demonstrated the immense power of creating…
Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge
Large Language Models (LLMs) excel at multi-hop reasoning in distribution, yet fail on unseen compositions, a phenomenon known as the curse…
Toward Autonomous Soft Robotic Endovascular Navigation via Imitation Learning
In endovascular surgery, endovascular interventionists push a thin tube called a catheter, guided by a thin wire to a treatment site inside…
People use fast and flat simulation to reason about new games
Games have long been a microcosm for studying planning and reasoning in both natural and artificial intelligence (AI), often focusing on ex…
Improving Topic Modeling of Social Media Short Texts with Rephrasing: A Case Study of COVID-19 Related Tweets
Social media platforms such as Twitter (now X) provide rich data for analyzing public discourse, especially during crises such as the COVID…
Enabling Agents to Communicate Entirely in Latent Space
While natural language is the de facto communication medium for LLM-based agents, it presents a fundamental constraint. The process of down…
Enhancing Adversarial Transferability through Block Stretch and Shrink
Input transformation-based attacks improve adversarial transferability by aggregating gradients over transformed inputs. Existing analyses…
SPQR: A Multi-Dimensional Benchmark for Safety Alignment under Benign Model Adaptation
Text-to-image diffusion models can emit copyrighted, unsafe, or private content. Safety alignment aims to suppress specific concepts, yet e…
On the Condition Number Dependency in Bilevel Optimization
Bilevel optimization minimizes an objective function, defined by an upper-level problem whose feasible region is the solution of a lower-le…
CUDA-L2: Surpassing cuBLAS Performance for Matrix Multiplication through Reinforcement Learning
In this paper, we propose CUDA-L2, a system that combines large language models (LLMs) and reinforcement learning (RL) to automatically opt…
The Theory of Strategic Evolution: Games with Endogenous Players and Strategic Replicators
Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. This paper provides th…
Graph-Based Bayesian Optimization for Quantum Circuit Architecture Search with Uncertainty Calibrated Surrogates
Quantum circuit design is a key bottleneck for practical quantum machine learning on complex, real-world data. We present an automated fram…
Emotion Recognition in Signers
Recognition of signers' emotions suffers from one theoretical challenge and one practical challenge, namely, the overlap between grammatica…
MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
This paper studies the training-testing discrepancy (a.k.a. exposure bias) problem for improving the diffusion models. During training, the…
SwinIFS: Landmark Guided Swin Transformer For Identity Preserving Face Super Resolution
Face super-resolution aims to recover high-quality facial images from severely degraded low-resolution inputs, but remains challenging due…
BiasLab: A Multilingual Dual-Framing Framework for LLM Bias Measurement, Applied to Workplace and HR Contexts
Background: Large language models (LLMs) harbor systematic biases that are particularly consequential in workplace and HR contexts, where t…
Stable On-Policy Distillation through Adaptive Target Reformulation
Knowledge distillation (KD) is a widely adopted technique for transferring knowledge from large language models to smaller student models;…
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
While Mixture-of-Experts (MoE) scales capacity via conditional computation, Transformers lack a native primitive for knowledge lookup, forc…
PUMA: Perception-driven Unified Foothold Prior for Mobility Augmented Quadruped Parkour
Parkour tasks for quadrupeds have emerged as a promising benchmark for agile locomotion. While human athletes can effectively perceive envi…
Referential Regimes: Transformation-Invariant Identity for Neutral Substrates
Data systems increasingly operate under persistent legal, political, and analytic disagreement, where no single interpretive authority can…
SFO: Learning PDE Operators via Spectral Filtering
Partial differential equations (PDEs) govern complex systems, yet neural operators often struggle to efficiently capture the long-range, no…
Rethinking Zero-Shot Time Series Classification: From Task-specific Classifiers to In-Context Inference
The zero-shot evaluation of time series foundation models (TSFMs) for classification typically uses a frozen encoder followed by a task-spe…
Disentangling Intrinsic Importance from Emergent Structure in Multi-Expert Orchestration
Multi-expert systems, where multiple Large Language Models (LLMs) collaborate to solve complex tasks, are increasingly adopted for high-per…
Understanding Persuasive Interactions between Generative Social Agents and Humans: The Knowledge-based Persuasion Model (KPM)
Generative social agents (GSAs) use artificial intelligence to autonomously communicate with human users in a natural and adaptive manner.…
SynthSAEBench: Evaluating Sparse Autoencoders on Scalable Realistic Synthetic Data
Improving Sparse Autoencoders (SAEs) requires benchmarks that can precisely validate architectural innovations. Current LLM-based SAE bench…
Debiasing Central Fixation Confounds Reveals a Peripheral "Sweet Spot" for Human-like Scanpaths in Hard-Attention Vision
Human eye movements in visual recognition reflect a balance between foveal sampling and peripheral context. Task-driven hard-attention mode…
BRIDGE: Bridging Reasoning In Distillation Gap Elimination via Structure-Aware Masking
Chain-of-Thought (CoT) reasoning has significantly improved LLMs' mathematical problem-solving capabilities, but distilling such capabiliti…
Turbo Connection: Reasoning as Information Flow from Higher to Lower Layers
Complex problems, whether in math, logic, or planning, are solved by humans through a sequence of steps where the result of one step inform…
A General Equilibrium Theory of Orchestrated AI Agent Systems
We establish a general equilibrium theory for systems of large language model (LLM) agents operating under centralized orchestration. The f…
MetaState: Persistent Working Memory Enhances Reasoning in Discrete Diffusion Language Models
Discrete diffusion language models (dLLMs) generate text by iteratively denoising a masked sequence. However, standard dLLMs condition each…
RVN-Bench: A Benchmark for Reactive Visual Navigation
Safe visual navigation is critical for indoor mobile robots operating in cluttered environments. Existing benchmarks, however, often neglec…
VehAnchor: Metadata-Free Metric Scale Recovery from Vehicle Cues in Aerial Imagery
Autonomous aerial robots operating in GPS-denied or communication-degraded environments frequently lose access to camera metadata and telem…
Context-Dependent Affordance Computation in Vision-Language Models
We characterize the phenomenon of context-dependent affordance computation in vision-language models (VLMs). Our primary study uses Qwen3-V…
Local Message-Passing for Discrete Graph Generation
Discrete graph generation has emerged as a powerful paradigm for modeling graph-structured data, yet state of the art models often rely on…
MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a compre…
Towards Robust Speech Deepfake Detection via Human-Inspired Reasoning
The modern generative audio models can be used by an adversary in an unlawful manner, specifically, to impersonate other people to gain acc…
ECoLAD: Selecting Anomaly Detectors for Automotive Deployment via Compute-Reduction Evaluation
Automotive anomaly detectors are often selected from accuracy only benchmarks on workstation class hardware, whereas in-vehicle monitoring…
Evolutionarily Stable Stackelberg Equilibrium
We present a new solution concept called evolutionarily stable Stackelberg equilibrium (SESS). We study the Stackelberg evolutionary game s…
Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
While Vision-Language Models (VLMs) have achieved remarkable performance, their Euclidean embeddings remain limited in capturing hierarchic…
Critical Damping as a Momentum Schedule: Multi-Seed Validation, a Hybrid Recipe, and an Exhaustive Negative Result on Surgical Layer Selection
The critical damping condition of the damped harmonic oscillator model of SGD with momentum (Qian, 1999) yields a momentum schedule with no…
StanceMoE: Mixture-of-Experts Architecture for Stance Detection
Actor-level stance detection aims to determine an author expressed position toward specific geopolitical actors mentioned or implicated in…
How Annotation Trains Annotators: Competence Development in Social Influence Recognition
Human data annotation, especially when involving experts, is often treated as an objective reference. However, many annotation tasks are in…
From Paper to Program: Knowledge Externalization and Bottleneck Diagnosis in AI-Assisted Quantum Many-Body Programming
Large language models can write scientific code, but direct paper-to-program translation remains fragile when correctness depends on tacit…
Pickalo: Leveraging 6D Pose Estimation for Low-Cost Industrial Bin Picking
Bin picking in real industrial environments remains challenging due to severe clutter, occlusions, and the high cost of traditional 3D sens…
MG$^2$-RAG: Multi-Granularity Graph for Multimodal Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) mitigates hallucinations in Multimodal Large Language Models (MLLMs), yet existing systems struggle wi…
Tool-MCoT: Tool Augmented Multimodal Chain-of-Thought for Content Safety Moderation
The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various medi…
Private Seeds, Public LLMs: Realistic and Privacy-Preserving Synthetic Data Generation
Large language models (LLMs) have emerged as a powerful tool for synthetic data generation. A particularly important use case is producing…
Filtered Reasoning Score: Evaluating Reasoning Quality on a Model's Most-Confident Traces
Should we trust Large Language Models (LLMs) with high accuracy? LLMs achieve high accuracy on reasoning benchmarks, but correctness alone…
SegWithU: Uncertainty as Perturbation Energy for Single-Forward-Pass Risk-Aware Medical Image Segmentation
Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and cli…
Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning
Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency. Whe…
PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging
Multimodal Large Language Models (MLLMs) rely on multimodal pre-training over diverse data sources, where different datasets often induce c…
Graph Construction and Matching for Imperative Programs using Neural and Structural Methods
Reusing verification artefacts requires identifying structural and semantic similarities across programs and their specifications. In this…
Toward a Scientific Discovery Engine for Weather and Climate Data: A Visual Analytics Workbench for Embedding-Based Exploration
Earth system science is producing increasingly large, high-dimensional datasets from both physics-based and AI-driven models. While embeddi…
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation
Robot imitation data are often multimodal: similar visual-language observations may be followed by different action chunks because human de…
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
Prompt compression reduces inference cost and context length in large language models, but prior evaluations focus mainly on autoregressive…
Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build
How much have students' ordinary learning processes shifted in response to generative AI, and how does that affect their durable learning o…
A Multi-Model Metric-based Selection Framework for Abstractive Text summarization
Automatic text summarization has become increasingly important due to the rapid growth of digital textual information. This paper presents…
Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models
Generating diverse images from sparse text is hard; generating compact actions from rich observations is easier. From the condition-target…
The Cross-Architecture Substrate: A Domain-Transcendent, Calibration-Surviving Geometric Invariant of Modern Vision Encoders
Different vision neural networks -- trained to classify, contrast, reconstruct, or match images to text -- should have correspondingly diff…
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
We introduce Embodied-R1.5, a unified Embodied Foundation Model (EFM) that integrates comprehensive embodied reasoning capabilities, spanni…
ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories
Training capable OS agents requires data that simultaneously captures structured user intents, multi-turn task delegation, and grounded too…
Pipette: An Embodied Simulation Platform, Benchmark, and Data-Efficient Augmentation Framework for Wet-Lab Robotics
Wet-lab robots can improve the reproducibility, throughput, and safety of biomedical experiments, but scaling their learning requires custo…
Gefen: Optimized Stochastic Optimizer
AdamW is a default optimizer for modern deep learning, but its first and second moment states add roughly two parameter-sized buffers to tr…
RealityBridge: Bridging Editable 3D Gaussian Splatting Driving Simulations and Real-World Videos
Long-tail hazardous scenarios are essential for safety-oriented autonomous driving, yet they are difficult to collect and reproduce at scal…
RankGraph-2: Lifecycle Co-Design for Billion-Node Graph Learning in Recommendation
Graph-based retrieval at billion-node scale requires jointly solving three tightly coupled problems -- graph construction, representation l…
FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving
Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficie…
Small edits, large models: How Wikipedia advocacy shapes LLM values
Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia ap…
RWGBench: Evaluating Scholarly Positioning in Related Work Generation
Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.…
What Does It Mean to Break a Distillation Defense?
Black-box LLMs (accessible only via API) are vulnerable to distillation attacks, in which an attacker queries the model and trains a studen…
Average-Power-Budgeted Underwater Vehicle Control via Constrained Reinforcement Learning
Underwater vehicles operate from a fixed onboard energy budget that propulsion rapidly depletes, so a controller that completes its task wh…
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training
Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but th…
Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare
Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then as…
NaviCache: Test-Time Self-Calibration Caching for Video Generation
Video Diffusion Models (VDMs) is constrained by immense computational costs. While offline calibration-based acceleration suffers from cali…
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test constru…
CMSL: Constructive Multi-Sequence Learning for Recommendation Systems
Sequence learning has emerged as the promising paradigm in recommendation systems, surpassing traditional Deep Learning Recommendation Mode…
Multi-Agent Routing as Set-Valued Prediction: A WildChat Benchmark and Cost-Aware Evaluation
Tool and agent routing from natural-language prompts is naturally a set-valued prediction problem: a single query may require multiple agen…
Freeform Preference Learning for Robotic Manipulation
Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where spa…
FLYNN: Robust Neural Network for Robot Navigation using Fly Brain Topology
While deep learning models achieve state-of-the-art performance in complex tasks, they remain brittle when faced with new environments or s…
Diffusion-GR2: Diffusion Generative Reasoning Re-ranker
Generative reasoning re-rankers achieve strong recommendation accuracy by emitting a chain-of-thought before re-ordering a candidate list,…
DemoPSD: Disagreement-Modulated Policy Self-Distillation
On-policy self-distillation (OPSD) has emerged as a practical method for training large language models (LLMs) to reason, where a single mo…
Do All Visual Tokens Matter Equally? Object-Evidence Preserving Token Merging for Vision-Language Retrieval
Multi-vector vision-language retrieval preserves fine-grained visual evidence through maximum-similarity late interaction, but dense image-…
x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability
Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function eva…
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of…
Dropbox、Claude連携スタート ChatGPT、Geminiに続き……「AI共通のコンテンツ基盤」に
既存の「ChatGPT」や「Gemini」との連携も拡大。主要なAIプラットフォームをまたいでDropboxをコンテンツ基盤として使えるようにする。
MicrosoftのナデラCEOが警告、AI利用者が支払う「二重コスト」 一つはお金、もう一つは「さらに価値あるもの」
米Microsoftのサティア・ナデラCEOが自身のブログで、企業はAIに「二重のコスト」を払っていると警告した。一つは利用料金、もう一つはAIを役立たせるために明け渡す固有の知識だという。
Already rich, already successful, why the last wave of tech winners is grinding again
They're rolling up their sleeves again, seemingly out of fear of missing AI's defining moment and, presumably, the irresistible allure of m…
Uber’s product chief on hotels, robotaxis, and why the company doesn’t want to be “everything for everyone”
Uber Chief Product Officer Sachin Kansal walks TechCrunch through the company's financial-services ambitions, its increasingly complicated…
Video-generation startup PixVerse raises $439M, valuation soars past $2B
With the cash, the company aims to expand its world model offering and reach customers across geographies.
Hermes agent maker Nous Research in talks for new funding at $1.5B valuation
The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.
日本企業の“鬼門”、アクセンチュアは突破できるか? OpenAIとの協業で狙う「業務効率化超え」
AIを実証実験や一部業務の効率化で止めず、全社の変革へ――。多くの企業が向き合うこの課題に、アクセンチュアはOpenAIとどう取り組むのか。これまで“鬼門”だった、部門をまたいだ業務プロセス改革の突破方法とは。
AIエージェントを作って終わりから「自己進化」へ、富士通MAAF検証開始
富士通は、業務向けマルチAIエージェント基盤「MAAF」を開発した。会議録画などからシステムを自動構成し、運用履歴に基づき安全に自己進化する。自社AI基盤との連携により企業全体のAI活用を支援する狙いだ。
「We Must Act Now」 AIによる急激な変革に著名経済学者やIT重鎮が警鐘
AIによる経済変革への備えを訴える声明「We Must Act Now」が、ノーベル賞受賞者16人を含む200人以上の経済学者やAI研究者の署名付きで公開された。スタンフォード大学の研究者らが取りまとめたもので、AIは産業革命を上回る規模の経済変革をはるかに短期間で引き起こす可…
“純国産の政府AI”稼働へ NTTらのモデル採用 「先陣を切る」――松本デジ相が語った意欲
デジタル庁の政府AI「源内」で、国産AIモデルと国産クラウドを活用した“純国産の政府AI”が稼働する。松本大臣は「先陣を切る取り組みになる」と述べた。
Satya Nadella has issued a shocking warning to companies using AI
Of all the debates raging about the potential downsides of AI, there is one worry causing the most hand-wringing among AI enthusiasts in Si…
The wildest allegations in Apple’s trade secrets lawsuit against OpenAI
Apple’s trade secrets lawsuit against OpenAI contains allegations that range from employees joking about unauthorized access to Apple’s sys…
Sam Altman’s space data center trash talk is what most experts already believe
Responding to Musk accusing him of being a scammer, Altman said, "homeboy you're the one sellling [sic] public market investors on short-te…
Should AI help you get away with killing your spouse?
What does a world of total user-aligned AI actually look like?
Anthropic starts localizing Claude pricing for India, its biggest market after the US
Claude users in India are starting to see Indian rupee-denominated subscription plans.
2026-07-13(189件)
Waze adds new AI-powered features and customization updates
Some of the new features are powered by Google's Gemini AI assistant, which reflects the tech giant's broader push to integrate Gemini acro…
OpenAIのブラウザ「ChatGPT Atlas」終了へ 公開から1年足らずで
米OpenAIがAIブラウザ「ChatGPT Atlas」の停止日を2026年8月9日と案内し、データの移行手順を公開した。移行先には新しいChatGPTデスクトップアプリとChrome拡張機能を挙げている。
GMOグループ、AI時代に「エンジニア含む組織体制見直し」 熊谷代表が「AI変革最高責任者」に
熊谷正寿代表が「グループCAIO」に。「エンジニアを含む組織体制を見直し、AIナイズされた組織へと変革する」
Claude、利用制限を全リセット 競合「GPT-5.6」公開と同日……OpenAI幹部「ビビってるね」
OpenAIが「GPT-5.6」を一般公開した同日、AnthropicがClaude全ユーザーの利用制限を一斉リセット。OpenAI幹部の「ビビってるね」という煽り返信も話題だ。
アニメ特化動画生成AI「AnimeGen」無償公開、商用利用も可 国内AIベンチャーAIdeaLab
AI開発企業のAIdeaLabは、アニメに特化した動画生成AIモデル「AnimeGen」(アニメジェン)を公開した。ライセンスは商用利用もできる「Apache-2.0」。
「ChatGPT Work」「Codex」の5時間制限枠を一時解除 「GPT-5.6 Sol」の処理効率も改良へ
米OpenAIは、デスクトップ向けAIツール「ChatGPT Work」や、付随するAIコーディングエージェント「Codex」に設定されている5時間の利用制限枠を一時的に解除すると発表した。
エージェントによる業務自動化をどう実現? 「Microsoft Build 2026」で発表された多数の新技術
Microsoftは開発者向けイベント「Microsoft Build 2026」で、エージェント基盤からモデル、開発端末、量子コンピューティングまで多数の新技術を発表した。
Interval Certifications for Multilayered Perceptrons via Lattice Traversal
In this work we present a rigorous theoretical framework to a foundational problem of AI safety, namely adversarial robustness. In particul…
CogniConsole: Externalizing Inference-Time Control as a Formal Abstraction for Reliable LLM Interactions
Reliability in large language model (LLM) systems is typically framed as a function of model capability. We challenge this by demonstrating…
GATS: Graph-Augmented Tree Search with Layered World Models for Efficient Agent Planning
Large Language Model (LLM) agents have shown promise in multi-step planning tasks, but existing approaches like LATS (Language Agent Tree S…
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading
AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus o…
A Formalization of the Mean-Field Derivation of the Vlasov Equation: AI-Assisted Lean Formalization as a Strategy Game
We formalize a research result in the Lean 4 proof assistant by having a mathematician direct an AI system, and frame the activity as a for…
ARCANA: A Reflective Multi-Agent Program Synthesis Framework for ARC-AGI-2 Reasoning
We present ARCANA, a collaborative multi agent framework for solving ARC AGI 2 tasks under strict test time and hardware constraints. ARCAN…
Neuro-Agentic Control: A Deep Learning-based LLM-Powered Agentic AI Framework for Controlling Security Controls
Cyberattacks on operational technology are increasingly causing costly downtime and physical damage, exposing the limitations of traditiona…
L-MAD: A Systematic Evaluation of Multi-Agent Debate Structures in Legal Reasoning
While multi-agent debate (MAD) frameworks have shown significant potential in general reasoning, their effectiveness in highly structured,…
MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation
Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with r…
KV-PRM: Efficient Process Reward Modeling via KV-Cache Transfer for Multi-Agent Test-Time Scaling
Process Reward Models (PRMs) have been proven to be highly effective in guiding test-time scaling (TTS) methods, which significantly boost…
Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift
Deployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work,…
Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents
Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with broad k…
OpenProver: Agentic and Interactive Theorem Proving with Lean 4
In this system paper, we present OpenProver, an open-source system for LLM-driven automated theorem proving (ATP) with integrated Lean 4 fo…
LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making
In this work, we introduce LongMedBench, a real-world EHR-based benchmark for long-horizon clinical decision-making. Prior evaluations of L…
Communication-Efficient Digital-Twin Coordination for Heterogeneous LLM Embodied Agents over Computing Power Networks
Embodied agent teams powered by heterogeneous large language models (LLMs) are being widely deployed in physical artificial intelligence su…
Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review
Worldbuilding, the construction of coherent fictional worlds, is a foundational task in game design and literary creation. Large Language M…
How Does Bayesian Causal Discovery Fail? Characterising Structural Consequences in Linear Gaussian Networks under Latent Confounding
Bayesian causal discovery is widely used for its ability to quantify epistemic uncertainty over directed acyclic graphs (DAGs) through post…
ProofCouncil: An LLM Agent for Solving Open Mathematical Problems
Large language models (LLMs) have shown increasing promise in solving open problems in mathematics. However, their performance can be furth…
Ceci n'est pas une pipe: AI systems as semantic abstractions
An AI system's output is not the fact or world state it appears to describe, but rather an engineered representation. We propose a semantic…
Multimodal Reward Hacking in Reinforcement Learning
Reinforcement learning (RL) is increasingly used to align multimodal large language models (MLLMs), but higher rewards do not always imply…
Shared Selective Persistent Memory for Agentic LLM Systems
Agentic LLM systems that generate code through multi-turn tool use face a fundamental context problem: each session starts from zero, disca…
SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction
Does every cancer patient truly need a complete diagnostic workup for accurate survival prediction? In multimodal clinical oncology, diagno…
Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI
Modern AI systems are increasingly being evaluated for their ability to reason, code, prove theorems, use tools, and long-horizon research…
Knowledge Graphs and Explainable AI as Complementary Resources for Urban Mining
Pre-demolition assessment, the regulated audit process at the heart of urban mining, is an information process in which AI support must ser…
TrustX Agent Risk Classification Framework (ARC): Risk-Tiering Internally Created Agentic AI Systems
The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk f…
Agora: Enhancing LLM Agent Reasoning Via Auction-Based Task Allocation
Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and too…
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are no…
Minimal Decision Dynamics and Contextual Probability: A Quantum Tug-of-War Model
Decision making often exhibits context dependence that challenges classical probability theory. This paper develops a quantum-like extensio…
REFORGE: A Method for Benchmarking LLMs' Reverse Engineering Capabilities in Decompiled Binary Function Naming
Large language models (LLMs) are increasingly applied to reverse-engineering tasks, and recent threat-intelligence reporting shows them ope…
A Unified Approach to Interpreting Knowledge Distillation for Large Language Models via Interactions
Despite the success of knowledge distillation (KD) in Large Language Models (LLMs), the underlying mechanism behind its efficacy remains un…
iLENS: Interpretable LLM-Guided Mixture-of-Experts for Neuroimaging Survival Analysis
Alzheimer's Disease (AD) is a complex neurodegenerative disorder that continues to impact millions of people worldwide. Predicting AD conve…
Signed Symmetric Quantization for Few-Bit Integers
The signed integer alphabet contains one more negative representable value than positive. Yet, by convention, the standard symmetric intege…
Sticky Routing: Training MoE Models for Memory-Efficient Inference
Mixture-of-Experts (MoE) models activate only a sparse subset of experts per token, yet consecutive tokens frequently activate different ex…
Reward Transport: Property Control in Flow Matching via Noise-Space Alignment
The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show…
Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement
Expert parallelism has become the prevailing paradigm to serve Mixture-of-Experts (MoE) models. Its efficiency depends on the communication…
LieBN: Batch Normalization over Lie Groups
Manifold-valued measurements are prevalent in various machine learning tasks. Recent advances have extended Deep Neural Networks (DNNs) to…
HERO: A Heterogeneity-Aware Benchmark Library for Federated Continual Learning
Federated continual learning (FCL) evaluates how distributed clients learn from changing data streams while retaining previously learned kn…
DaDaDa: A Dataset for Data Pricing in Data Marketplaces
High-quality data drives machine learning advances across industries. Recognizing the value of data, data transactions are increasingly com…
Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices
With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introdu…
LLM-Driven Evolutionary Generation of Multi-Objective Bayesian Optimization Algorithms
Designing effective multi-objective Bayesian optimization (MOBO) algorithms requires balancing many interdependent design choices whose opt…
EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins
Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches le…
Multi-Conditioned Diffusion Synthesis of Sand Boils for Low-Resource Earthen-Levee Inspection
Sand boils on earthen levees are safety-critical defects, but pixel-level detection is limited by scarce annotations. We present a diffusio…
TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology
The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine unde…
Prompt-Driven Exploration
Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inje…
A Novel Parallel QCNN Architecture with Efficient Classical Simulability
This work presents a study of an implementation of a novel Quantum Convolutional Neural Network (QCNN) for binary classification of images…
Eluna: An Agentic LLM System for Automating Warehouse Operations with Reasoning and Task Execution
Warehouse operations are governed by Standard Operating Procedures (SOPs) that encode complex, multi-system decision logic, which must be e…
NL-PAC: Specification Ambiguity and Certified Minimax Risk Floors in LLM-Mediated Supervision
Large language models increasingly provide labels, evaluations, and feedback for tasks specified in natural language. When a specification…
MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs
Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate obse…
CLAP: Direct VLM-to-VLA Adaptation via Language-Action Grounding
Vision-language-action models (VLAs) inherit semantic capabilities from pretrained VLMs, yet large-scale post-training on robot data and ar…
The Patchwork Problem in LLM-Generated Code
LLM-generated code often compiles, passes tests, and appears correct, yet breaks once deployed. The root cause is frequently structural rat…
SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation
While autonomous coding agents have significantly advanced automated test generation, they remain fundamentally limited by lazy generation,…
AlphaZero in Sparsely Rewarded Games: Limits and Auxiliary Supervision
AlphaZero has demonstrated that a neural-guided Monte Carlo Tree Search can achieve superhuman performance, but strong play does not necess…
Model Agnostic Graph Prompt Learning for Crystal Property Prediction
Graph Neural Networks have emerged as a powerful tool for the fast and accurate prediction of various crystal properties. These models ofte…
Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing
We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motiv…
Phone Segmentation and Recognition through Phonological Activation Mapping
Phone segmentation and recognition are inherently related tasks, yet modern approaches typically model them separately. We argue that phone…
Video Generation Models are General-Purpose Vision Learners
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equiv…
Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems
Artificial intelligence (AI) is shifting scientific discovery from task-specific workflows towards autonomous systems that organize explora…
Quantum Logic as the Logic of Contexts
Quantum logic is usually presented as a non-classical departure from ordinary reasoning forced on us by quantum mechanics, with classical l…
On Locality and Length Generalization in Visual Reasoning
A striking feature of the human visual system is that it ingests visual information through a series of local foveated glimpses, rather tha…
Inside the Skill Market: From Software Engineering Activities to Reusable Agent Skills
Software engineering (abbrev. SE) has continuously evolved through increasingly powerful forms of reuse, from source code and libraries to…
OmniMapBench: Benchmarking Visual-Centric Reasoning on Diverse Map Documents
Recent advancements in LVLMs necessitate robust benchmarks for complex, visually grounded reasoning. A critical limitation is identified in…
PRecG: Legal Precedent Retrieval with Graph Neural Networks and Rhetorical Role Segmentation
Legal precedent retrieval is a fundamental task in legal case preparation, planning, litigation strategy, and legal research. Current appro…
A Coreset Selection Framework with Ensemble Aggregation for Image Classification
The rapid growth of image data has produced large-scale datasets, raising concerns about the time and memory costs of model training. Selec…
Beyond Metadata: CAPRA for Hidden Subgroup Analysis under Missing Metadata in Medical Imaging
Medical imaging models are often deployed without the demographic, acquisition, and quality metadata needed for subgroup auditing. Once tho…
Integrating Large Language Models and Graph Convolutional Networks for Semi-Supervised Image Classification
While the growing availability of image data has driven significant advances, labeling datasets remains costly and time-consuming. Therefor…
Event Stream based Multi-Modal Video Anomaly Detection: A Benchmark Dataset and Algorithms
Video anomaly detection (VAD) is critical for automated surveillance but remains fragile under challenging conditions such as illumination…
Augmenting Fundamental Analysis with Large Language Models: A RAG-Based System for Generating Investor Briefs
In this study, we examine the opportunities brought by Large Language Models (LLMs) to various aspects of fundamental analysis of companies…
IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation
While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step ite…
ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models
Representation alignment (REPA) has been investigated to accelerate diffusion training, but we observe that regularizing intermediate repre…
A Personalized Computational Framework for Assessing the Sufficiency of Partially Observed Data in Healthcare AI models
Achieving early and timely diagnosis and treatment for disease is a major challenge. Recent applications of machine learning (ML) algorithm…
Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations
Large Language Models are reshaping how software is developed and maintained. They are typically deployed in production using inference eng…
Generative Communications: Overview, Technologies, and Trends
The groundbreaking development of generative artificial intelligence (AI) is rapidly boosting the ability to generate content such as image…
Interference and Retention in Continual Learning
Continual learning commonly relies on post-hoc mechanisms such as replay, elastic regularization, or distillation. This work argues that fo…
Tactile and Vision Conditioned Contact-Centric Control for Whole-Arm Manipulation
Whole-arm manipulation involves direct contact with the environment while the robot completes a task by distributing contact across multipl…
Git-Assistant: Planning-Based Support for Updating Git Repositories
Version control systems are essential for collaborative software development, yet tools like git remain challenging for many practitioners.…
All you need is SAMPAT
The current state of the art in AI/ML rests on deep neural architectures, which, in general, suffer from a lack of interpretability. Interp…
LLMs for health: Perceived benefits, risks, intention to use AI chatbots, and willingness to self-disclose across sensitive health topics
AI chatbots are increasingly used for answering health-related questions. This study examines the role of topic type discussed with an AI c…
Blockchain-Linked Auditable Decision Management for Telecom/IoT Fraud-Control Requests
Telecom fraud-control studies often stop at detector-level classification, but deployment use requires request-level policy resolution, lif…
Geopolitical alignment: Endorsement effects in large language models
Large language models (LLMs) are increasingly used to summarize and evaluate policy-relevant information, but it remains unclear whether th…
Risk-Aware General-Utility Markov Decision Processes
We study general-utility Markov decision processes (GUMDPs) with risk-aware objectives. In this framework, an agent aims to optimize a risk…
Creativity, honesty and designed forgetting emerge in small hyperbolic language models
Language models are optimised for scale, yet remain functional rather than companionable, and as an assistant personalises into a companion…
Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works
Thematic indexing -- the practice of assigning structured conceptual labels to sections of text -- is essential to scholarly access in larg…
Letting the Data Speak: Extracting Keywords from Crowdsourced Collections with AI
Identifying and assigning keywords at scale is a technical, practical, and ethical challenge for crowdsourced collections. This article rep…
WILDTRACE: Benchmarking Natural Evidence Trails in Long-Context Reasoning
Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across…
Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning
Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process…
Deceptive Grounding: Entity Attribution Failure in Clinical Retrieval-Augmented Generation
Retrieval-augmented generation evaluation checks whether model claims are factually grounded in retrieved documents. It does not check whet…
CtrlVTON: Controllable Virtual Try-On via Visual-Instance-Prompt Segmentation
Virtual try-on (VTO) has made significant progress in realistically transferring garments onto a target person. Yet most systems give the u…
Diversifying to Verify: When Task-Equivalent Programs Differ in Verifiability
Program verification is crucial for software correctness, but producing fully verified programs remains difficult in practice. This paper s…
When Routes Run Out: Adversarial Co-Learning and Explainable Robustness in Quantum Repeater Networks
We study an adversarial bandit problem for entanglement-based quantum-network routing over a modest graph corpus. Alice selects an end-to-e…
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU
The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficien…
Fully Trainable Deep Differentiable Logic Gate Networks and Lookup Table Networks
We introduce a novel method for both partial and full optimization of the connections in deep differentiable logic gate networks (LGNs) and…
On-Device Adaptive Battery Power Prediction for Electric Vehicles
Adaptive power management in Electric Vehicles (EVs) requires accurate power prediction. Although deep learning models have emerged as high…
Self-Guided Test-Time Training for Long-Context LLMs
Long-context processing has become increasingly important for large language models (LLMs), but simply extending the context window does no…
SVF-CR: Synchronized Visual-Facial Cross-Refinement for Multimodal Ambivalence and Hesitancy Recognition
Ambivalence and hesitancy are subtle behavioral states that are expressed through a combination of verbal content, facial behavior, visual…
A Sovereign, Open-Source Foundation Model for German and English
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and Engl…
Test-Time Scaling for Small VLMs on Multilingual Visual MCQ
Test-time scaling (TTS) reliably improves reasoning in large language models, but whether it transfers to small open vision-language models…
Parameter-Efficient Vision-Language Adaptation with Continuous Metadata Conditioning for Animal Re-Identification
Long-term animal re-identification (ReID) must remain robust to gradual morphological evolution and seasonal appearance shifts. Although re…
Practical Source Code Recovery from Binary Functions Using Anchor-Based Retrieval and LLM Reasoning
We present a practical pipeline for recovering source code from stripped binary functions by combining reverse engineering, anchor-based so…
Decoupling Language Guidance from Backbones for Text-Guided Medical Segmentation
Text-guided medical image segmentation leverages clinical semantics to improve lesion delineation, yet many existing models bind cross-moda…
All Explanations are Wrong, But Many Are Useful: Exploring the Rashomon Explanation Set with Large Language Models
Explaining machine-learning models is increasingly important for decision-making and consumer trust, yet it is widely believed to come at a…
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility
A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping vis…
Failure as a Process: An Anatomy of CLI Coding Agent Trajectories
Large language model (LLM) coding agents are increasingly deployed to autonomously perform software engineering tasks in terminal-based env…
Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
Vision-Language Models (VLMs) are the perceptual backbone of embodied AI, but their energy footprint on edge hardware remains poorly unders…
ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts
Foundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and…
TCLA: Training-Free Class-wise Logit Adaptation for Medical Vision-Language Models
Medical Vision-Language Models (VLMs) exhibit strong zero-shot performance, yet their effectiveness still declines on out-of-distribution (…
Large-Scale Portfolio Optimization Problem Under Cardinality Constraint With Enhanced Multi-Objective Evolutionary Algorithms
Decision-making is posing an increasingly formidable challenge to investors because of the growing number of alternatives available in fina…
Conceptual Networks for Cross-Linguistic Idiomatic Expressions:A Feature-Based Graph Approach
We present an interpretable network-based framework for representing idiomatic and figurative meaning across eight typologically diverse la…
PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers
Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-l…
Task-Specific Multimodal Question Answering Agents via Confidence Calibration and Incremental Reasoning for QANTA 2026
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA).…
4DR360: State Reasoning for Joint 3D Detection and Occupancy Prediction in 4D Radar-Camera Full-Scene Perception
Reliable autonomous driving requires full-scene perception that couples foreground objects with dense semantic layout. Recently, 4D millime…
Lean-QIT: Towards a Formal Infrastructure for Quantum Information Theory
Quantum information theory (QIT) characterizes the capabilities and fundamental limits of quantum information processing, underpinning quan…
Semantic Pareto-DQN: A Multi-Objective Reinforcement Learning Framework for Financial Anomaly Detection
Financial anomaly detection suffers from extreme class imbalance, causing traditional single-objective algorithms to exhibit ``fraud collap…
VEXAIoT: Autonomous IoT Vulnerability EXploitation using AI Agents
Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurati…
Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
Vision language models (VLMs) have made remarkable progress in visual reasoning during the last decade. Most evaluations have used simple s…
Scalable Visual Pretraining for Language Intelligence
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms…
PHINN-EEG: Topological Time-Series Analysis of Dream-State EEG -- Dynamic Betti Curves for Dream Content Classification and Topology-Conditioned Neural Signal Synthesis
Current electroencephalography (EEG)-based dream detection relies on power spectral density (PSD) and statistical moment features, achievin…
IFAR: Multi-Perspective and Multi-Level Causal Discovery with LLMs
Large language models (LLMs) have developed rapidly, and their reasoning capabilities have become a hot research topic. However, there is s…
Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors
White-box monitoring is increasingly adopted as an auditing tool as Large Language Models (LLMs) are deployed in daily operations to ensure…
A Descriptive and Normative Theory of Human Beliefs in RLHF
Human preferences in RLHF are typically modeled as a function of the human's reward function or corresponding optimal state-action values.…
QAgent: An LLM-based Multi-Agent System for Autonomous OpenQASM programming
Programming quantum circuits at the OpenQASM level is essential for achieving hardware-aware optimization and reliable execution on noisy i…
Beyond Embeddings: Interpretable Feature Extraction for Binary Code Similarity
Binary code similarity detection is a core task in reverse engineering. It supports malware analysis and vulnerability discovery by identif…
Leveraging Multi-Agent System (MAS) and Fine-Tuned Small Language Models (SLMs) for Automated Telecom Network Troubleshooting
Telecom networks are rapidly growing in scale and complexity, making effective management, operation, and optimization increasingly challen…
Improving Language Agents through BREW: Bootstrapping expeRientially-learned Environmental knoWledge
Large Language Model (LLM)-based agents are increasingly capable of complex, multi-step tasks such as GUI automation, tool use, and data ma…
Programming over Thinking: Efficient and Robust Multi-Constraint Planning
Multi-constraint planning involves identifying, evaluating, and refining candidate plans while satisfying multiple, potentially conflicting…
PACE: A Personalized Adaptive Curriculum Engine for 9-1-1 Call-taker Training
9-1-1 call-taking training requires mastery of over a thousand interdependent skills, covering diverse incident types and protocol-specific…
A Self-Evolving Agentic Framework for Metasurface Inverse Design
Metasurface inverse design can realize complex optical functionality, but turning a target optical response into executable optimization co…
Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys
Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions. We stu…
Towards Shutdownable Agents: Generalizing Stochastic Choice in RL Agents and LLMs
Misaligned artificial agents might resist shutdown. One proposed solution is to train agents to lack preferences between different-length t…
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs
Direct Preference Optimization (DPO) is an effective framework for aligning large language models with human preferences, but it struggles…
Heterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning
Coordination graphs are a central abstraction in cooperative multi-agent reinforcement learning (MARL), yet existing sparse-graph learners…
Explaining is Harder Than Predicting Alone: Evaluating Concept-based Explanations of MLLMs as ICL Visual Classifiers
In-context learning (ICL) enables multimodal large language models (MLLMs) to classify images from a few labelled examples. Yet, how these…
Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs
Strong reasoning depends not only on model knowledge but also on how effectively cognitive behaviors are deployed during generation. Existi…
SHARP: Sleep-based Hierarchical Accelerated Replay for Long Range Non-Stationary Temporal Pattern Recognition
Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming…
Coding-agents can replicate scientific machine learning papers
Scientific machine learning papers typically make computational claims, e.g., that the relative mean square error is less than 5% or that t…
Shortcut Learning in Legal Judgment Prediction: Empirical Evidence from the UK Employment Tribunal
Current Legal Judgment Prediction (LJP) is constrained by its reliance on post-hoc judicial materials, increasing the likelihood that model…
Large Behavior Model: A Promptable Digital Twin of the Retail Customer
Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive ac…
Projection Methods for Operator Learning and Universal Approximation
We obtain a new universal approximation theorem for continuous (possibly nonlinear) operators on arbitrary Banach spaces using the Leray-Sc…
Multi-Attribute Steering of Language Models via Targeted Intervention
Inference-time intervention (ITI) has emerged as a promising method for steering large language model (LLM) behavior in a particular direct…
Transformer-Empowered Actor-Critic Reinforcement Learning for Sequence-Aware Service Function Chain Partitioning
In the forthcoming era of 6G networks, characterized by unprecedented data rates, ultra-low latency, and ubiquitous connectivity, effective…
M4V: Multimodal Mamba for Efficient Text-to-Video Generation
Text-to-video generation has significantly enriched content creation and holds the potential to evolve into powerful world simulators. Howe…
Single-Frame Point-Pixel Registration via Supervised Cross-Modal Feature Matching
Point-pixel registration between LiDAR point clouds and camera images is a fundamental yet challenging task in autonomous driving and robot…
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLM…
Evaluating Retrieval-Augmented Generation vs. Long-Context Input for Clinical Reasoning over EHRs
Objective: To evaluate whether retrieval-augmented generation (RAG) can serve as an efficient alternative to long-context prompting for cli…
REAL: REtrieval-reAsoning and Logic-constructed Attention Behaviors for Long-Context KV Cache Compression
The growing sequence length of large language models poses significant challenges for key-value (KV) caches. Existing state-of-the-art cach…
Contrastive Weak-to-strong Generalization
Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples…
Explaining Human Choice Probabilities with Simple Vector Representations
We formalize human choice behavior in a probabilistic hide-and-seek task. In our geometric construction, vectors represent participant choi…
H3Former: Hypergraph-based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations…
AutoGraphAD: Unsupervised network anomaly detection using Variational Graph Autoencoders
Network Intrusion Detection Systems (NIDS) are essential tools for detecting network attacks and intrusions. While extensive research has e…
Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation
LLM-based simulations can enable controlled studies of civic deliberation, but current systems lack speaker-attributed data and methods for…
RIS-Assisted Downlink Pinching-Antenna Systems: GNN-Enabled Optimization Approaches
This paper investigates a reconfigurable intelligent surface (RIS)-assisted multi-waveguide pinching-antenna (PA) system (PASS) for multi-u…
Data-Driven Learnability Transition of Measurement-Induced Entanglement
Measurement-induced entanglement (MIE) captures how local measurements generate long-range quantum correlations and drive dynamical phase t…
How to DP-fy Your Data: A Practical Guide to Generating Synthetic Data With Differential Privacy
High quality data is needed to unlock the full potential of AI for end users. However finding new sources of such data is getting harder: m…
ReinforceGen: Hybrid Skill Policies with Automated Data Generation and Reinforcement Learning
Long-horizon manipulation has been a long-standing challenge in the robotics community. We propose ReinforceGen, a system that combines tas…
Transition Matching Distillation for Fast Video Generation
Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interac…
Principles of Lipschitz continuity in neural networks
Deep learning has achieved remarkable success across a wide range of domains, significantly expanding the frontiers of what is achievable i…
Preference Conditioned Multi-Objective Reinforcement Learning: Decomposed, Diversity-Driven Policy Optimization
Multi-objective reinforcement learning (MORL) seeks to train agents capable of balancing conflicting objectives. While single preference-co…
Knowledge-Based Design Requirements for Generative Social Robots in Higher Education
Generative social robots (GSRs) powered by large language models enable adaptive, conversational tutoring but also introduce risks such as…
Empowering 9-1-1 Calltaking Training with Generative AI: Experiences and Lessons Learned
Emergency call-takers form the first operational link in public safety response, handling over 240 million calls annually while facing a su…
The LLMbda Calculus: AI Agents, Conversations, and Information Flow
Large language models are increasingly deployed as agents: they plan, call tools, read untrusted data, and act on the results. This exposes…
SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution
Real-world software must continuously evolve to meet ever-changing and open-ended requirements. AI agents, increasingly deployed as long-ru…
Machine Learning for Network Attacks Classification and Statistical Evaluation of Adversarial Learning Methodologies for Synthetic Data Generation
Supervised detection of network attacks has always been a critical part of network intrusion detection systems (NIDS). Nowadays, in a pivot…
SLIDERS: Systematic Reviews via Automated Evidence Synthesis and Reconciliation
Systematic reviews -- which requires comprehensive evidence collection and synthesis from large document corpora in response to targeted re…
Tuning Derivatives for Causal Fairness in Machine Learning
Artificial-intelligence systems are becoming ubiquitous in society, yet their predictions typically inherit biases with respect to protecte…
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
Effective collaboration between embodied agents requires more than acting in a shared environment; it demands communication grounded in eac…
AnchorMoE: Interpretable Time Series Classification via Anchor-Routed MoE
Multivariate time series classification (MTSC) is pivotal in high-stakes domains, such as clinical diagnosis and industrial fault detection…
Language Models Need Sleep: Learning to Self-Modify and Consolidate Memories
The past few decades have witnessed significant advances in the design of machine learning algorithms, from early studies on task-specific…
Are LLMs Ready to Assist Physicians? PhysAssistBench for Interactive Doctor-Patient-EHR Assistance
The most plausible near-term role of medical LLMs is to assist rather than replace physicians, yet current evaluations often test isolated…
ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL
Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Co…
GAP-GDRNet: Geometry-aware monocular 6D pose estimation for spacecraft using synthetic geometric supervision
Monocular spacecraft 6D pose estimation remains difficult under weak texture, thin structures, illumination variation, and occlusion. This…
Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language
LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate…
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents
Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, an…
SCOReD: Student-Aware CoT Optimization for Recommendation Distillation
Chain-of-thought (CoT) distillation in the recommendation domain is a necessary precursor to RL training, but raw teacher traces are ill-su…
Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs
Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on En…
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additiona…
Riemannian Geometry for Pre-trained Language Model Embeddings
Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sent…
Omni-Sleep: A Sleep Foundation Model via Hierarchical Contrastive Learning of CNS-ANS Dynamics
Sleep physiology arises from the coordinated dynamics of the central nervous system (CNS) and autonomic nervous system (ANS), as reflected…
Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and age…
人に残る「クリエイティブな仕事」とは? Adobeの“人×AI”の取り組みから探る
広告分野では制作業務に生成AIが広範囲で利用されるようになり、クリエイターやデザイナーの仕事が一部で奪われつつある。「われわれはAIが全てを決める世界を目指しているのではない」と語るAdobeの新たな取り組みから、クリエイティブ業務における人とAIの関係を考察する。
「GPT-Liveが“まるで人間”」ってホンマ? 出汁を「でじる」、トーストを「素焼き」て言うてたけど……
AIは、何も塗らないトーストを「素焼き」と呼び、出汁を「でじる」と言い出した。
3D CADから組立工程/手順書を作成、設計と製造をつなぐAI活用機能
Sceneは、製造業向け統合ワークスペース「Scene Workspace」に、3D CADデータから組立工程や組立手順書を作成するAI活用機能をβ機能として搭載した。デンソーの工機部と共同開発したもので、設計、生産技術、製造の連携を1つの基盤でつなぐ。
「足りないのはCOBOL人材じゃない」 日立が語る、AI時代のシステム刷新における“人”の役割
従来は人海戦術に頼っていた領域にAIが入り込み、さまざまな課題が解消しつつある。レガシーシステムの刷新もそのうちの一つだ。しかし、AIをシステム刷新プロジェクトで積極的に活用している日立製作所(以下、日立)は、「人間の力がこれまで以上に重要になる」と語る。AI時代のシステム刷新…
Anthropic、「Fable 5」の無償アクセスと「Claude Code」利用上限50%増を7月19日まで延長
Anthropicは、「Claude Fable 5」を有料プランで追加費用なく使えるキャンペーンと、AIコーディング支援ツール「Claude Code」の週間利用上限を50%引き上げる措置を、いずれも米太平洋時間の7月19日まで延長すると発表した。Fable 5の無償アクセス…