週次AIニュース 2026-W30
対象期間: 2026-07-20 〜 2026-07-26(1684 件)
トピックの推移
トピック別件数
- 研究/論文 718件
- LLM/生成AI 650件
- エージェント 358件
- 画像/動画生成 203件
- ビジネス/資金調達 98件
- ロボティクス 91件
- ハードウェア/半導体 60件
- その他 37件
- 規制/政策 10件
今週のハイライト(上位 10 件)
Building AI infrastructure with the Effingham County community
OpenAI announces Project Camellia in Effingham County, Georgia, with commitments to responsible energy, community investment, jobs, and acc…
How news organizations are using AI to advance their vital missions
News organizations are using AI to strengthen reporting, grow audiences, and improve business operations, with OpenAI tools supporting jour…
Advancing the next era of national science
OpenAI outlines its commitment to advancing American science working with the U.S. Department of Energy and national labs to use frontier A…
Introducing OpenAI Presence
Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for custom…
Introducing the ChatGPT for small business program
OpenAI launches the ChatGPT for Small Businesses program, helping entrepreneurs build AI skills, automate work, and grow with ChatGPT Work.
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities…
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards thro…
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
Google commits $40M in AI tokens and credits for the Genesis Mission
MicrosoftやNVIDIAなど、AIのオープンウェイト規制に反対する書簡を公開――Anthropicは署名せず
MicrosoftやNVIDIA、Metaなど30社以上の米国の企業や団体が、オープンウェイトAIモデルへの過度な規制回避を求める共同書簡を公開した。オープンモデルをAIエコシステムの基盤と位置付け、開発や評価におけるメリットとイノベーション促進を強調。中国企業の急速な台頭や技…
全件(日付別)
2026-07-26(3件)
Monday.com is the latest tech company to blame AI for layoffs — here are 20 others
A running look — in reverse chronological order — at the bigger tech companies that have announced significant layoffs this year with AI as…
MicrosoftやNVIDIAなど、AIのオープンウェイト規制に反対する書簡を公開――Anthropicは署名せず
MicrosoftやNVIDIA、Metaなど30社以上の米国の企業や団体が、オープンウェイトAIモデルへの過度な規制回避を求める共同書簡を公開した。オープンモデルをAIエコシステムの基盤と位置付け、開発や評価におけるメリットとイノベーション促進を強調。中国企業の急速な台頭や技…
Librarians are hosting viral ‘Avoiding AI’ workshops for people who are fed up with Big Tech
At libraries around the country, "Avoiding AI" workshops have elicited unprecedented demand.
2026-07-25(9件)
One fallen power line exposed a growing AI data center problem. Here’s how to fix it.
A close call in Northern Virginia revealed just how poorly data centers respond to grid disruptions. Here's how to fix the problem.
I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone else
OpenAI's fancy new AI keypad will be a lot of fun for some, while many others are probably not going to touch it.
Anthropic、「Claude Opus 5」公開 Fable 5に迫る性能を半額で――サイバー安全策は緩和、拒否時は自動フォールバックも
Anthropicは、最新LLM「Claude Opus 5」を公開した。上位モデル「Claude Fable 5」に迫る知能を半額の価格で提供する。プログラミングやナレッジワークにおいて高い評価を獲得し、推論の深さを調整するパラメータや安全性分類器に連動する自動フォールバック…
Prentis, new AI lab co-founded by Reid Hoffman, Mark Pincus in talks to raise $100M
The neolab is betting that automating routine computer tasks will soon outpace coding as AI's biggest use case.
Why Cognition bought Poke: AI personality is becoming a competitive advantage
The acquisition brings Poke’s conversational style and interaction model to Cognition’s coding agent Devin, reflecting a growing belief tha…
Anthropic launches Opus 5
Opus 5 will be both cheaper and less restrictive than Fable, likely making it preferable in most use cases.
As US weighs response to Chinese AI, industry urges against broad open-weight restrictions
AI companies, including Nvidia and Mistral, urge policymakers to avoid broad restrictions on open-weight AI models as Washington debates re…
Bluesky’s AI assistant Attie expands into an open social research tool
Users can now ask Attie questions about news, trends, and conversations on Bluesky and other apps on the AT Protocol.
Midjourney acquired the astrology app Co-Star
The AI lab Midjourney continues to expand its purview beyond image and video generation.
2026-07-24(358件)
‘AI communism’, rogue models, and the why Kimi K3 spooked Wall Street
Chinese AI lab Moonshot’s open model Kimi went viral this week for reasons that had less to do with the model itself and more to do with ho…
OpenAI’s new voice mode makes it to the ChatGPT desktop app
ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents.
スーパーに並んだ「ごちゃごちゃ生成AIポップ」が物議 “看板王”こと、きぬた歯科院長「これはアリ」
スーパーの青果売り場に並ぶ、生成AIで作ったとみられる派手な商品ポップがXで物議を醸している。吸血鬼や戦国武将を描いたデザインに「見づらい」との声が相次ぐ中、看板広告で知られるきぬた歯科のきぬた泰和院長は「これはアリ」と評価。その理由とは。
近畿大、入試にAIの利用認める 情報学部の総合型選抜で
近畿大学は、2027年度の情報学部の総合型選抜入学試験で、生成AIの利用を認めると発表した。提出する自己PR動画やプレゼンテーション資料などでのAIの利用方針を明示した。
AIにもサプライチェーン管理が必要? 中国AI「Kimi K3」を巡る批判でAIの調達リスクが浮き彫りに
中国の最新AIモデルを巡り、米政府高官が、Anthropicの「Claude Fable 5」をモデルの学習に利用した“不正蒸留”が行われていたと指摘した。AIモデルの調達や導入を巡るサプライチェーンリスクが浮き彫りとなっている。
メルカリ、「AI活用の最前線」明かす動画公開 「なぜCTOがCHRO兼CAIOになったのか」など13本
メルカリは、自社のAI活用について紹介する動画を公開した。7月8日に開催したイベント「Mercari AI Career Fes 2026」で実施したセッションのアーカイブ動画、全13本を公式YouTubeチャンネルで視聴できる。
海外の「Claude」や「GPT」ではダメなのか 日本企業向けai&、そのメリットは?
ai&は、日本企業向けに設計したAI推論プラットフォーム「ai& Inference」の提供を開始した。ClaudeやGPTなど海外ベンダーのAIに代わる選択肢として、AI利用コストを大幅に削減できるとしている。
開発工数見積もりの「負担が重い」をAIで解消へ 明治安田はどう実現?
有識者に頼りがちなシステム開発工数の見積もりは、AIエージェントでどこまで効率化できるのか。明治安田生命保険がPoCで検証した仕組みを見ていこう。
AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics
Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud envi…
Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts
Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-gen…
ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models
This paper presents an AI-driven browser extension that identifies clickbait to help users avoid misleading Internet articles. Moving beyon…
Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs
When a language model gives different answers on repeated runs, does that variation reveal what it does not know? Self-consistency turns th…
JAXBench: Benchmarking Autonomous TPU Kernel Optimization
Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on,…
DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding
While parallel decoding is central to the efficiency of Diffusion Large Language Models (dLLMs), current strategies are often hindered by o…
InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed w…
DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions
Large language models (LLMs) handle many tasks with one set of parameters, but under KV-cached inference it is unclear what task-general st…
PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs
Enhancing the task-specific capabilities of Large Language Models (LLMs) primarily requires substantial instruction-tuning datasets. Howeve…
Benchmarking the Personalization Capabilities of Large Language Models
Personalization, the act of varying a message to induce action from a specific receiver while keeping sender, channel, and time fixed, has…
Robust Critics: Defending LLMs Against Multi-Turn Attacks
When a user asks a language model something harmful, is it a genuine attack or a misunderstood but well-meaning question? This ambiguity is…
Incomplete Prompt Jailbreaks in Large Language Models
Large language models (LLMs) are increasingly released as open-weight models with safeguards against harmful requests. Nevertheless, senten…
VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification
Natural language interfaces can greatly benefit the accessibility and usability of optimization modeling, and recent advances in large lang…
SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification
Sampling in LLM inference comprises a combinatorial set of logit processing, token selection, and verification operations for speculative d…
Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment
We present an empirical benchmark evaluating how five large language models assess multisensor physical hazard data. Testing 60 scenarios a…
Semi-Supervised Text-Attributed Graph Distillation
{\em Text-Attributed Graphs} (TAGs) have emerged as an expressive data model for integrating graph topology with rich textual semantics. Ex…
Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detectio…
Enabling Scalable Topology Inference in Distribution Systems via Constrained Multi-Source Inference
Accurate distribution system topology is essential for outage localization, voltage analytics, and operation of distribution grids, yet mai…
Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating
Local-cloud collaboration is a practical way to deploy large language models under resource constraints, but existing methods often rely on…
PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails
Recent advances in large language models have enabled web agents to autonomously execute complex tasks. In practice, users frequently provi…
Tractable Hierarchical Control of Autoregressive Language Models
Constraining the generation of autoregressive large language models (LLMs) is an important component of integrating language models into fo…
The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path
Large Language Models (LLMs) are fundamentally limited by representation collapse, a bottleneck that severely degrades long-context perform…
Expectation Alignment of Language Models for Real-World User Expectations
Large language models (LLMs) have demonstrated remarkable performance on standard benchmarks, yet it remains largely unexplored whether the…
OPTScientist: Multi-Agent Discovery of Typed Optimizer Programs for Transformer Pretraining
Designing optimizers for modern deep learning remains a challenging scientific problem, requiring the joint consideration of optimization g…
Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering
Large language models (LLMs) are increasingly used to answer questions about political information, including in election-adjacent informat…
Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants
Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for exam…
EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL
Text-to-SQL has advanced rapidly with large language models, but complex database queries still require reasoning beyond one-shot generatio…
CRAWO: Custom Resources for Adaptive Workload Orchestration
Edge Intelligence has emerged as a key paradigm for enabling real-time applications in smart cities by shifting computation from centralize…
DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making
Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process ea…
Attention-based Experience Replay Framework for Continual Learning of Agnostic Time Series Forecasting Models
Deep learning has led to remarkable progress in artificial intelligence, particularly in robotics, imaging and sound processing. However, a…
Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation
Production LLM applications commonly stack a regex filter in front of model-side alignment; prior work found no measurable coverage gain fr…
Workload-Aware Caching for Multi-Agent Systems
Multi-agent systems decompose complex tasks into directed acyclic graphs (DAGs) of specialized agent executions, creating natural opportuni…
From Errors to Rules: Iterative Prompt Optimization for Text Classification
Prompt optimization for text classification spans diverse approaches, from demonstration selection to exploration-based search to error-dri…
AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs
Large language models (LLMs) augmented with tools are emerging as autonomous agents capable of using Web engine, APIs, and code to solve co…
ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models
Large Language Models generate plausible backend code, but a single-pass paradigm provides no guarantee of correctness or runtime reliabili…
FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts
Large Language Models (LLMs) perform strongly on well-specified reasoning tasks with a feasible answer. However, problems encountered in th…
MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation
Despite rapid progress in LLM-based code generation, writing correct and performant kernels for hardware accelerators remains a key bottlen…
Inducing Comparability of Factorised Probability Distributions
To allow for principled comparison between two probabilistic graphical models defined over non-identical variable sets, they have to be lif…
LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization
We present and evaluate LeanFlow, an LLM agent system specialized for translating mathematical papers into buildable Lean projects. Recent…
Optimizing Hypergraph-Based RAG: Toward Better Fact Extraction and Chunk Retrieval
GraphRAG enables deeper reasoning by structuring knowledge as graphs but struggles with n-ary facts. HyperGraphRAG uses hypergraphs for ric…
MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference
Large language models (LLMs) are increasingly used for program-aided reasoning, agentic decision making, and structured task execution, but…
Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain
We introduce Telco-GAIA, a bilingual, multi-modal benchmark for evaluating tool-using agents on the data of a real-world telecommunications…
SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning
Multimodal Continual Instruction Tuning (MCIT) is crucial for adapting Multimodal Large Language Models (MLLMs) to evolving a sequence of d…
Reliability-Aware LLM Alignment from Inconsistent Human Feedback
Reinforcement Learning from Human Feedback (RLHF) is critical for aligning Large Language Models (LLMs) with human preferences. However, it…
CANN Bench: Benchmarking Agent Generated Kernels against Real NPU and Algorithmic Limits
AI agents are now capable of writing, compiling, and iteratively optimizing low-level operator kernels on different hardware platforms. Exi…
Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
Large language models (LLMs) are increasingly evaluated on mathematical problem solving, yet prior work often treats representationally equ…
Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models
Mean cross-positional attention degradation is widely reported in transformer interpretability, yet whether it causally limits contextual r…
Autonomous disproofs of the sum-product conjecture over $\mathbb R$ with GPT-5.5 Pro
OpenAI's recent disproof of the Erd\H{o}s unit distance conjecture marked a milestone for AI in mathematics. It also inspired another break…
ConfidenceBench: Evaluating Confidence Calibration in Large Language Models
Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly. In these settings, acc…
Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis
Agentic LLM systems such as OpenScholar and PaperQA2 read the scientific literature and return cited answers, and both they and their bench…
PromptPack: Scaling LLM Annotation Agents for Online Recommendation
Online recommendation platforms increasingly use Large Language Models (LLMs) to extract structured features from ad creatives. While deplo…
DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers
Large language model (LLM) agents are increasingly deployed over Model Context Protocol (MCP) servers, yet the benchmarks used to evaluate…
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use
Tool-use agents that address day-to-day digital tasks such as ordering groceries must not only operate applications, but also interact with…
StrideDiffusion: Accelerating Diffusion Models for Time-series Generation
Diffusion models have become competitive generators for time series, but their practical use is limited by the large number of sequential d…
CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning
Memory Manager models are pivotal in agent systems. Existing methods rely predominantly on LLM-judged synthetic question-answer (QA) pairs,…
AI-Driven Multi-Hop Relay Selection for Smart Urban NR-V2X Networks via Learning-to-Optimize Graph Neural Networks
Reliable and low-latency NR-V2X communications are essential for smart mobility in dense urban environments. However, limited Road-Side Uni…
KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback
In large-scale text analysis tasks, pre-trained language models are often used to embed text corpora for downstream analysis. However, such…
WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms
Large Language Models (LLMs) have demonstrated strong capabilities in code generation and reasoning, yet their ability to perform temporal…
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-…
ArbiGraph: Arbitrarily Scalable Verifiable Task Graphs for Evaluating Context Management
We introduce ARBIGRAPH, a benchmark generator for evaluating whether tool-assisted language agents can retain, update, compose, and discard…
The Human-AI Substitution Principle: When will you be replaced by AI in your organization?
Artificial Intelligence (AI) is rapidly transforming organizations, raising a fundamental organizational and economic question: when will a…
Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling
High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling t…
Can an AI System Be Creative? A Critical Perspective from Art and Engineering
This paper examines the question of whether artificial intelligence (AI) systems can be creative, approached from the dual perspective of a…
Profiling Lightweight Large Language Models
Lightweight large language models (LLMs) are increasingly being deployed locally on personal computers and are expected to play a growing r…
Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs
The electrocardiogram (ECG) is a cornerstone of cardiac as- sessment, yet clinical deployment of deep learning models remains con- strained…
Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks
Body-based emotion recognition is important for real-time affective systems, but graph-based skeleton models can be computationally expensi…
Auditing Provenance Sensitivity in LLM Agent Action Selection
LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. E…
Auditing Evidence Use in Medical LLM Diagnosis
Medical LLMs are often evaluated by whether they select the correct diagnosis, but diagnostic accuracy alone does not show whether the mode…
Code Monitor Red Teaming for Public-Test-Passing Code
Visible tests are a common gate for LLM-generated code, but passing them does not certify specification correctness. We study a deployment-…
Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report…
Source-Prior-Driven Selective Adaptation for Efficient Diffusion Model Finetuning
Fine-tuning large diffusion models for new domains or styles involves a trade-off: improving target-specific generation often degrades the…
Traceable Scholarship: Page Anchors and Ariadne's Thread for Humanistic Inquiry in the Age of Generative AI
Generative AI lets large language models produce scholarly-looking text within seconds, yet fluency does not equal valid explanation. The d…
OPOD: On-Policy Omni Distillation
Omni-modal models can handle text, images, and audio in one system, but improving all of these abilities together remains difficult. Traini…
Representing Entity Importance in AI Knowledge Systems: A Dual-Signal Framework of Audience Evaluation and Structural Authority
AI knowledge systems require representations of entity importance for retrieval, recommendation, evidence selection, and knowledge-intensiv…
SciExplore: Evaluating Autonomous Agents from Scientific Navigation to Information Integration
Scientific research involves complex information-seeking and reasoning workflows across heterogeneous sources. However, existing benchmarks…
Clustered Edge Intelligence: Beyond Just Convergence of Edge Computing and AI
We are moving from an information age to the age of intelligence. A decade, or possibly less than that, data will not be the gold anymore r…
From Scalars to Time Series: Rethinking Implicit Neural Representations for Time-Varying Volumetric Data
Implicit neural representations (INRs) for time-varying volumetric data are typically trained using dense sampling over spatiotemporal coor…
Delivery, Not Storage: Cue-Anchored Working Memory as a Harness Property for Coding Agents
Coding agents ship with one kind of memory: documents. Instruction files, plan artifacts, and auto-written memory directories are deliberat…
Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence
Efficient multimodal inference is increasingly constrained not only by model quality or FLOP count, but also by the cost of preserving, mov…
GuardianAgentBench: Where Agents Fail and How to Guard Them
As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and re…
Workflow-Localized Mechanism Learning: Attribution-Guided Repair and Knowledge Reuse for Structured Agent Skills
Agent Skills package reusable procedural knowledge as external artifacts for frozen language-model agents, yet existing optimizers do not j…
Naju: A Native Discrete State-Space Model with Independent Retention and Writing for Long-Sequence Memory
Long-sequence memory tracking places two opposing demands on a recurrent state: near-lossless retention of stored bindings over long horizo…
Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers
Zero-shot summarization using Large Language Models (LLMs) has significantly advanced the abstractive summarization task by producing coher…
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER…
HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices
Traditional approaches to wearable health signal analysis, such as smartwatches, are constrained by rigid analytical frameworks and limited…
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs
Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting…
Can Generative Recommendation Reach Cold Items? A Temporal Perspective on Semantic-ID Generation
Semantic-ID-based generative recommendation represents items as sequences of shared semantic tokens, enabling token recombination beyond is…
AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what…
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure
As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. C…
SafeStep: AI-powered Travel Assistance for Elderly People with Frailty or Dementia
More than a million people in the UK suffer from frailty or dementia, which severely compromise their ability to travel in urban environmen…
Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications
Recent advances have introduced speech-to-speech (S2S) conversational assistants capable of producing natural-sounding interactions, includ…
Explaining Weather Bulletins via ILP
Inductive Logic Programming (ILP) originated within the Logic Programming community in the Nineties as a framework for combining symbolic l…
Differentiable Logic Programming to Mitigate Reasoning Shortcuts in Neurosymbolic Systems
Neurosymbolic (NeSy) systems integrate neural networks with logical reasoning to achieve both generalization and interpretability, but rece…
Identifying Good Rules for Efficient SAT Encodings of Single-Constant Multiplication Using Machine Learning
The Single Constant Multiplication problem is a fundamental NP-hard optimization task in hardware design, which seeks to decompose a fixed…
Bound-Founded Semantics for Answer Set Programming with Difference Constraints: Preliminary Report
While the integration of linear constraints has significantly expanded the reach of Answer Set Programming (ASP), existing hybrid solvers o…
A New Well-Supported Semantics for Description Logic Programs
Description logic programs are a powerful formalism for combining rules with ontologies. The well-supported semantics for description logic…
How Rules Represent Causal Knowledge: Causal Modeling with Probabilistic Logic Programming
Pearl famously argues that causal knowledge enables the prediction of intervention effects. By contrast, purely descriptive knowledge suppo…
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders
The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fu…
Logic Programming Semantics for Causal Processes
Motivated by challenging modelling issues in the life sciences, we investigate the relationship between logic programming semantics and the…
BasketEvent: Understanding Who Did What and When in Basketball Videos
Comprehensive basketball video understanding requires resolving not only what event occurs, but also who is responsible and when the key ev…
An LLM-Driven Workflow for Automated Process Control Strategy Generation and Tuning from Dynamic Process Models
We present a structured large-language-model-driven workflow for automated multi-variable control design from dynamic process models. The w…
Expert Behavior Prior Reinforcement Learning
Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learn…
Regulating autonomous and agentic AI
Regulating activities where regulatees use autonomous and agentic AI is challenging. Regulatory assumptions about regulatee knowledge and c…
SPORD: A Simulation-Propose-then-OR-Dispose Approach for Supply Chain Planning
For years, supply chain planning at e-commerce firms has operated as a collection of isolated projects. Each planning task from static netw…
Towards Faithful Graph Explanations with Synergistic Edge Effects via Granular Balls
Instance-level explanations aim to reveal the rationale behind a model's decisions for a specific graph. Previous methods explain graph neu…
Multimodal Pretraining for Generalizable EEG Representation Learning
Electroencephalography (EEG) models used for epilepsy are often limited to specific datasets and tasks. This limited approach can make it c…
MSBraM: A Multi-scale Self-supervised Brain Foundation Model for Hierarchical EEG Dynamics Learning
Self-supervised foundation models have recently shown strong potential for electroencephalogram (EEG)-based analysis. However, existing app…
Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog
Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning,…
Logical Regression for Planning with Axioms
In automated planning, logical regression is an operation that returns the most general condition necessary for an action to achieve a part…
PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning
In long-horizon LLM agent reinforcement learning, weak policies often repeat similar failures, producing uninformative rollout trajectories…
Bridging the Gap Between Plausibility and Admissibility: Constraint-Aware Flow Maps for Dynamic Graph Systems
Generative models can support decision-making under uncertainty by producing ensembles of plausible future system trajectories, but statist…
Agent-Guided Relational Concept Discovery: Toward Interpretable Surgical Margin Assessment
Deep learning models can effectively use Rapid Evaporative Ionization Mass Spectrometry (REIMS) data for surgical margin assessment. Howeve…
Detecting LLM-Generated Tokens in Human--LLM Coauthored Text
The rise of human-AI collaborative writing has created a growing need for fine-grained detection methods that support localizing likely LLM…
AREX: Towards a Recursively Self-Improving Agent for Deep Research
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verify…
Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks
Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based…
Toward Continuous Assurance for the Democratization of AI Agent Creation in Industry
AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development…
Agentic Context Management: Solving Agent Memory and Cost by Treating Them as Lifecycle and Architecture Problems
Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their…
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation
Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay it…
The Boundaries of Automation: A Theory of Persistent Human Participation
The rapid progress of AI has intensified the long-standing pursuit of automation: replacing human participation with algorithms wherever po…
MIRROR: Learning from the Other View for Multi-Modal Reasoning
Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasonin…
OpenForgeRL: Train Harness-native Agents in Any Environment
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, an…
Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning
Building socially calibrated large language models, which can learn from others without simply yielding to them, requires more than reducin…
Unsupervised Consensus-Based Anomaly Detection for Spatiotemporal Malaria Incidence in Ghana
A consensus anomaly detection framework was applied to monthly malaria surveillance data from Ghana (2014-2023) to identify atypical transm…
Deblurring in the Wild: A Real-World Image Deblurring Dataset from Smartphone High-Speed Videos
We introduce the largest real-world image deblurring dataset constructed from smartphone slow-motion videos. Using 240 frames captured over…
Through-the-Earth Magnetic Induction Communication and Networking: A Comprehensive Survey
Magnetic induction (MI) communication (MIC) has emerged as a promising candidate for underground communication networks due to its excellen…
From Attention to Frequency: Integration of Vision Transformer and FFT-ReLU for Enhanced Image Deblurring
Image deblurring is vital in computer vision, aiming to recover sharp images from blurry ones caused by motion or camera shake. While deep…
Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs'Hallucinations
Existing LLM hallucination mitigation methods, including prompt engineering and model optimization, either hardly alter models'internal kno…
Is MoE Routing a Huffman Code? Discovering the Frequency-Diversity Law in Chain-of-Thought
Mixture-of-Experts architectures have revolutionized scaling, yet the underlying logic of their routing remains a black box. In this paper,…
More Is Not More: What Matters for Diversity in LLM Opinions?
Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group m…
LLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining
We present LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026 on paragraph-level argument mining in UN and UNESCO r…
Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing
While language models remain frozen at their training state, the world evolves continuously. Knowledge editing has emerged as a key alterna…
Break Through the Compression Bottleneck: From Theory to Practice
As the parameter size of language models continues to grow, effective model compression is required to reduce their computational and memor…
Making Open-Source Text LLM Watermarks Durable Against Merging
Open-source LLMs (OSMs)arereaching near state-of-the-art performance, prompting prior works to trace the text they generate by embedding te…
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this ass…
Preference Tuning as Spectral Update Reorganization
Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains…
Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility
Proprietary large language models (LLMs) entail substantial intellectual and financial investment, making them valuable intellectual proper…
Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception
Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally ind…
The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs
LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examin…
A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction
Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain…
Response drift across frontier large language models
All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet…
RE-AD: Real-Time Requirement Adherence for Data Labeling
Human-annotated data remains fundamental to training frontier Large Language Models (LLMs). However, crowd-sourced annotations often suffer…
Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc
Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such…
Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention
Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attentio…
CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents
Large language model (LLM) agents operating over extended dialogues accumulate vast amounts of information, yet existing memory systems eit…
THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QA
Multi-hop question answering requires retrieving and integrating evidence from multiple contexts. Despite the rapid progress of current res…
Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?
Current full-duplex (FD) spoken dialogue systems can produce fluid interactions, yet it remains unclear whether they can adapt their turn-t…
Can Valence Reflect Morality in Natural Language? A Preliminary Annotation Study
Present implementations of artificial intelligence (AI) ethics do not adequately take feelings, or affect, into account. If AI should be al…
Verifier-First Evaluation of Agentic LLMs for Infrastructure-as-Code Generation
Infrastructure-as-Code (IaC) generation from natural language requires satisfying provider schemas, dependency planning, and organizational…
PhantomFill: When the Form Demands an Answer, Language Models Invent One
Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the…
The Active Ingredient in Muon's Grokking
The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW. Prior work attributes this to "spectral-norm con…
Scaling Closed-Loop Feature Channel Configuration with LLMs
Promising initial results in closed-loop large-language-model-based channel-configuration search demonstrated that neural-network widths ca…
Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement
Large Language Model (LLM) ensembles are increasingly used to improve reliability by combining predictions from multiple LLMs. However, exi…
CLOE: Christoffel Loss Autoencoder for Anomaly Detection
Semi-supervised anomaly detection plays a key role in diverse fields such as process monitoring, healthcare, and finance. However, lightwei…
Position: Stop Reactively Patching Your Model Every Time and Start Proactive Test-Driven AI Development
Many modern AI systems are designed to operate under diverse, open-ended, use-cases. To help generalize deployed systems, many deployed-sys…
Grounding Investor Views: Neural Predicates in the Black-Litterman Model
Portfolio construction under the Black-Litterman model requires investors to specify views on asset returns alongside explicit uncertainty…
A Graph Neural Network approach to zero-shot Digital Twins
Traditional Predictive Digital Twins often remain geometrically rigid, requiring extensive retraining or fine-tuning whenever the underlyin…
ReliableTableQA:How Much Supervision Does Reliability Annotation Need?
We introduce ReliableTableQA, a framework for training an LLM to annotate the statistical reliability of tabular QA results, not whether th…
Codec-Gauge: Learning Compression-Friendly Gauges for Transformer KV Caches
Long-context Transformer inference increasingly relies on KV-cache compression or quantization. Prior rotation and transform-coding results…
Leveraging Biokinetic Knowledge Priors for Data-Scarce Bioprocess Modeling
While deep learning has accelerated drug discovery, its impact on biomanufacturing has been considerably more limited. The reason is data s…
From Atoms to Entropy: Optimal Noise Allocation for Diffusion Training in the Convex Regime
How should a diffusion model decide which noise levels to train on, and how much? Despite the importance of this choice, current noise sche…
HypNO: A Graph-Based Neural Operator with Physics-Informed Message Passing for Hyperbolic Conservation Laws
We introduce HypNO, a graph-based neural operator for scalar hyperbolic conservation laws. HypNO operates directly on a space-time graph of…
Improving Access to Essential Medicines via Decision-Aware Machine Learning
A critical challenge in healthcare systems in low- and middle-income countries (LMICs) is the efficient and equitable allocation of scarce…
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
Reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while making a model worse under repeated sampling. W…
SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales
Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability cha…
Beyond SBDD: Geometric Deep Learning in Polypharmacology and Multi-target Drug Design
The traditional "one drug, one target" paradigm of structure-based drug design (SBDD) frequently proves inadequate for treating multifactor…
SenCos-GEM: SENet-Calibrated and Law-of-Cosines-Constrained Geometry-Enhanced Molecular Representation for Property Prediction
Effective molecular representation learning is crucial for accurate molecular property prediction. Recently, numerous self-supervised learn…
Monkey King Bang: A Unified Scientific Multimodal Foundation Model
Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar trans…
StabilityBench: Benchmarking Instability in LLMs
AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior r…
Joint Utilization of Geospatial and census proxies for Autoencoder-Assisted Downscaling (JUGAAD) of socioeconomic indicators in India
Monitoring poverty and food security indicators is imperative for addressing socioeconomic challenges in developing nations. A limitation i…
Geometric Configurations of Perturbed Jailbreak Prompts
Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major secur…
Bayesian uncertainty estimation improves clinical decision making in medical AI agents
Machine learning models for medical image analysis typically lack a reliable measure of confidence, limiting their use in ambiguous or atyp…
Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans
The function of many genes is still unknown, and conventional driver-discovery methods, which rely on how frequently a gene is mutated, can…
When Does Recurrence Become an Algorithm? Convergence Selection in Weight-Tied Looped Transformers
When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm? We answer with four findings from…
RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring
Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic trainin…
Demonstrating GenDB: Instance-Optimized and Customized Query Processing Code Generation via LLM Agents
Traditional query processing engines require continuous development and extensions to support new techniques and user requirements, and in…
Frontier Financial Judgement: Can agents tell what might move a stock?
We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to asse…
Scaling Interpretable Transformers with Parity Bottleneck Layers
Language models are thought to exhibit the phenomenon of superposition, representing many more features than dimensions in their residual s…
SalesLoop: Reinforcement Learning from Performance Feedback for Sales Lead Ranking
Lead ranking in Customer Relationship Management (CRM) systems faces a persistent challenge: models achieving high offline accuracy often u…
Adaptive Multi-Horizon Reinforcement Learning
Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement l…
From Agent Failures to Text Policies: What Works and What Breaks
TextGrad improves language-model systems by revising text from feedback. Its core thesis is that natural-language feedback can act as a gra…
Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis
Concept Bottleneck Models provide interpretable-by-design predictions by mediating diagnosis through human-understandable concepts, but in…
DS@GT ARC at ImageCLEFmed GANs 2026: Geometric Filtering for Privacy-Preserving CT Slice Generation
We present a privacy-preserving framework for synthetic lung CT slice generation developed for the Image-CLEFmed GANs 2026 challenge. The a…
A Framework for Reputation Aware Uninorm-driven Consensus Algorithms for Blockchain Networks
The operation of blockchain is governed by consensus algorithms (CA). Several consensus mechanisms require significant computational power,…
U-CFR: Uncertainty-Guided Cascade Forward Refinement for Interactive Segmentation
Interactive image segmentation is critical for efficient image annotation; however, existing methods often require many corrective clicks o…
Transition-Related Potentials as Markers of Narrative Comprehension in Continuous EEG
Harnessing the potential of electroencephalography (EEG) for brain research is fundamentally limited by intrinsic noise and the diffuse pro…
Operational Identity: A Finite Audit of Declared and Implemented Rules of Sameness
A record system declares when two records refer to the same entity, occurrence, scope, or rule. Its disclosed implementation mechanisms ind…
GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning
Large language models increasingly use search tools to retrieve up-to-date information, introducing a new attack surface in which retrieved…
Self-Supervised Bio-Inspired Robotic Trajectory Planning with Obstacle Avoidance
Trajectory planning is a fundamental problem in robotics, requiring the generation of collision-free and efficient trajectories in a potent…
IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests
AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute c…
Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles
Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choose which models to combine. W…
Emergent Compositional Skills in Mixture-of-Experts VLAs
We consider the problem of learning compositional robot policies end-to-end from expert demonstrations, without any pre-specified notion of…
HARP: The Human--AI Research Platform
Large language models (LLMs) have shifted human--computer interaction from `traditional'' interface journeys toward more conversational exc…
Robostral Navigate
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains…
Synthetic minority data is redundant or invalid: a data-dependent validity theory and a de-biased test
For two decades, the standard remedy for class-imbalanced learning has been to fabricate synthetic minority examples, and the standard evid…
The Geometry of Personality: Activation Steering with Jungian Cognitive Functions
Activation steering enables control and interpretation of LLMs, yet existing work primarily models personality through static trait framewo…
Beyond Heavy Log Curation: Perplexity-Based APT Detection via Unsupervised, Context-Augmented Language Models
Advanced Persistent Threats (APTs) remain difficult to detect because only a small fraction of events in large-scale logs are attack-relate…
Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery
Scientific machine learning methods such as neural operators and physics-informed neural networks have advanced engineering applications an…
Probabilistic Residual Learning for Online Recommendations
Modern recommender systems are typically based on deep learning (DL) models, where a dense encoder learns representations of users and item…
TwistedMerge: Certified Higher-Order Diagnostics and Abstention for Model Merging
Model merging combines independently trained or fine-tuned models, but pairwise alignability does not imply globally consistent alignment.…
Anti-Goal Reasoning: Rethinking the Theory of Goal Reasoning in Non-Axiomatic Logic
Goal reasoning in Non-Axiomatic Logic (NAL) explains how an adaptive system derives means for realizing desired events under insufficient k…
Multi-turn RL with Structural and Performance Aware Rewards for CUDA Kernel Generation
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a powerful technique to enhance the reasoning capacity of LLMs for opt…
Scientific exploration, collaboration and labor division in the large language model era
Large language models (LLMs) have rapidly and significantly entered scientific workflows, but it remains unclear how their diffusion is ass…
Interaction Dynamics Modeling and Predictive Control for Safe Steerable Catheter--Tissue Interaction
Safe steerable catheter control is fundamentally a problem of interaction dynamics: the tip must follow a planned motion, remain compliant…
HyWorldVLA: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving
Vision-Language-Action (VLA) models augmented with world modeling represent a promising paradigm for end-to-end autonomous driving. While p…
Sparse Concept Channels in Frozen 3D CT Vision Encoders
Large vision-language models are becoming increasingly dominant in 3D medical image interpretation, but we rarely know which internal units…
Training Large Language Models for Self-Explanation Faithfulness
We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's…
TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning
Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important wh…
GlucoTune: A Unified Framework for Blood Glucose Preprocessing, Forecasting, and Benchmarking in Diabetes
Preprocessing blood glucose time-series data is a critical yet often overlooked step in developing data-driven methods for diabetes managem…
Relative Value Learning
In reinforcement learning, critics typically estimate absolute state values $V(s)$, estimating how good a particular situation is in isolat…
Hardware-Software Co-Design for Float16 On-Device Training on RISC-V Single-Core
By leveraging standard RISC-V extensions, namely Zfh (scalar float16) and Zvfh (vector float16), this work proposes an open-source framewor…
Demographically-Informed Heat-Mortality Risk Curves via Risk Graph Neural Networks
Estimating heat-related mortality risk is a core task in environmental epidemiology, typically addressed with Distributed Lag Non-linear Mo…
One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs' Clarification Policies
Ambiguous user requests make clarification a sequential decision problem for conversational LLM assistants: they must decide whether to ask…
CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants…
Representative Sets in Propositional Abduction
The propositional abduction problem is a well-known form of non-monotonic reasoning where we are asked to find an explanation of a given ma…
Case study: proving sqrt(2) irrational with LPTP and an LLM
We present the interactions with an LLM (Large Language Model) aiming at proving that the square root of 2 is not a rational number in an L…
Encoding Event-B Proof Rules in Prolog: An Interactive Sequent Prover for ProB
Event-B is a formal method rooted in predicate logic and set theory. We encoded over 600 proof rules in Prolog, enabling a systematic, comp…
Animation, Verification and Visualisation of Prolog Transition Systems with ProB
ProB is a Prolog-based model checker, animator and constraint solver for high-level formal specifications. One can also use ProB to animate…
Chess\_db: A framework for working with large chess game datasets
Chess is a two player strategic game that is embedded in classical AI culture as it was once the frontier for intelligent behaviour. There…
Case study: solving P-99 with LPTP and an LLM
Ninety-Nine Prolog Problems (P-99) is a famous set of Prolog exercises. We solved the first thirty three just by prompting an LLM (Large La…
Declarative Problem Solving in UAM Strategic Deconfliction
The growing demand for Urban Air Mobility (UAM) introduces significant challenges in airspace management, particularly within densely popul…
Towards a Certifying Grounder
Grounding, the translation of high-level theories into equivalent quantifier-free formulas, is a crucial step in declarative solving, yet i…
Hybrid MKNF with Classical Negation in the Rule Component
Hybrid MKNF knowledge bases under the well-founded semantics integrate Description Logics with Logic Programming. However, they do not supp…
Explainability Framework for Policy-Aware Autonomous Agents
In the field of Artificial Intelligence, an agent is a system which is able to autonomously make decisions in order to reach a desired goal…
Explainable Belief Harmonization under Dynamic Epistemic Partitions
Existing approaches to multi-agent belief combination have established mature foundations for combining uncertain beliefs under common assu…
slang.gr as a Large-Scale Crowdsourced Resource for Non-Standard Greek
Slang is a central component of everyday language, reflecting linguistic creativity, social identity, and cultural change, yet its dy- nami…
pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development
In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-…
A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset
We present CUP, a Greek book retrieval benchmark consisting of 868 catalog records and 104 expert-annotated queries with graded relevance j…
Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning
Multi-task learning (MTL) is a promising approach for prediction tasks derived from video game state data, as modern game telemetry provide…
Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning
Machine unlearning has emerged as a tool for removing personal data from trained models to comply with recent AI regulations. To evaluate u…
AI Assistants Overassist
Large language models (LLMs) are increasingly used as tutors and thought partners, helping users reason through problems. While guidance fr…
Scaling Up Formal Representation of Clinical Trial Protocols in Ensemble Logic Using LLMs: A Preliminary Study
The reliance on unstructured free text for documenting clinical trial protocols creates a significant barrier to automated reasoning, cohor…
PC-Edit: Prompt-Contrastive Region Discovery and Region-Guided Editing
Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by…
GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG
Retrieval-Augmented Generation (RAG) systems increasingly employ multiple LLM agents. Yet, most prior work optimizes components in isolatio…
Toward cryptographically verifiable authorization for autonomous AI agents: A security hypothesis, preliminary formal model, and proof-of-concept implementation
Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing…
From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics
Bibliometric indicators - citation counts, h-indexes, co-authorship networks - have long anchored science, technology, and innovation (STI)…
Phonetic forced alignment for low-resource language varieties: Model training and evaluation on Chengdu Mandarin
Phonetic forced alignment is a key technique in phonetic research, yet existing alignment systems lack specialized models for low-resource…
M$^3$-Gen: Interpretable Multimodal Generation of Gene Expression Profiles Using Clinical and Imaging Data
Integrating heterogeneous biomedical data, including clinical metadata, histopathology images, and molecular profiles, is crucial for compr…
Hilbert Operator for Progressive Encoding (HOPE): A Mathematical Framework for Deconstructing Learned Representations in Deep Networks
Deep neural networks encode complex representations, but deconstructing this internal knowledge remains a challenge. Given the link between…
DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
Open-vocabulary semantic segmentation (OVSS) leverages textual semantics to segment objects beyond predefined categories. While the self-su…
Mean-to-Score Discrete Diffusion: Posterior-Mean Denoisers for Score Entropy
Score Entropy Discrete Diffusion (SEDD) parameterizes discrete reverse processes with unconstrained positive score ratios. While positivity…
VoLN: Vision-Only Long-Horizon Navigation---Paradigm, Benchmark, and Method
Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions. However, route-level instructions com…
When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation
A vision-language AI assistant returns its answer as a stream of generated tokens. Therefore, a safety guard that watches that answer has t…
Cycle-Consistent and Uncertainty-Aware Neural Surrogates for Tokamak Edge Plasmas
The boundary and divertor plasma govern how a tokamak exhausts power and particles, setting heat fluxes, target conditions, and the onset o…
Token Budget Saturation and Mechanistic Early Detection of Reasoning Non-Convergence in Chain-of-Thought Models
Chain-of-thought reasoning models such as DeepSeek-R1-Distill-Qwen-7B exhibit a bimodal convergence pattern: generations either terminate w…
Adaptive Identity Anchoring: Closed-Loop Keyframe Placement for Synthetic Paired Supervision in Video Face Swapping
Video face swapping has no natural paired supervision: no real footage exists of one person's face performing another person's video. The s…
RUMBA: Russian User Memory Benchmark
The ability to handle long-term memory in LLMs is becoming increasingly critical, yet existing benchmarks remain English-centric and rely o…
Thinkink: 2D Spatial Ink-native Interaction with LLMs
People often use handwritten notes and sketches to externalize ideas for ideation. To integrate large language models (LLMs) into this prac…
Error Certificates for KV-Cache Eviction via Randomized Design
Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest. We prove that this design cannot k…
Compact Latent Coordination for Autonomous Vehicles at Unsignalized Intersections
Coordinating autonomous vehicles at unsignalized intersections remains a critical challenge for multi-agent reinforcement learning (MARL) s…
Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it
A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language m…
Improved lower bounds for the Shannon capacity of odd cycles
The Shannon capacity $\Theta(G)$ of a graph $G$ quantifies the maximum rate at which information can be transmitted with zero error over a…
GS-Agent: Creating 4D Physical Worlds With Generative Simulation
Creating dynamic and physically realistic 4D worlds from natural language descriptions is both fascinating and challenging. Traditional com…
ElasticTTT: Prior-Preserving Test-Time Tuning for Video Editing
Test-Time Tuning (TTT) on pretrained diffusion models has emerged as a powerful paradigm for video editing. However, there exists a foundat…
From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs
Concurrent stateful library APIs expose behavior through evolving resource ownership, lifecycle states, and competing interleavings. Large…
Visual Contrastive Self-Distillation
On-policy self-distillation (OPSD) is promising as it removes the external teacher required by on-policy distillation (OPD), yet it still n…
Beyond Sufficiency: Time Series Explanation with Counterfactual Necessity
Faithful explanations of time-series classifiers should identify subsequences that are not only sufficient to preserve a black-box model's…
Synthetic data generation framework for quality control automation in gravure printing
Quality control in printing, particularly in rotogravure printing, still depends on slow, costly, and subjective manual inspection. Automat…
Barzilai-Borwein Fails Superlinear Convergence on an Open Set of Quadratics for Every Dimension $n\geq 4$
Barzilai--Borwein (BB) method has shown strong practical performance in continuous optimization, yet its convergence dynamics remains poorl…
GraphVid: Interactive Graph-Controllable Video Generation
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts…
3D-Aware VLMs with Implicit and Explicit Geometries
Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tas…
A Counterfactual Cause in Situation Calculus
Perhaps the most popular modern formulation of actual causality is the HP account by Halpern and Pearl. Recent advancement has focused on e…
Fragile Preferences: A Deep Dive Into Order Effects in Large Language Models
Large language models (LLMs) are increasingly deployed in decision-support systems for high-stakes domains such as hiring and university ad…
From Checklists to Clusters: A Homeostatic Account of AGI Evaluation
Contemporary AGI evaluations report multidomain capability profiles, yet they typically assign symmetric weights and rely on snapshot score…
WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance
Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex bro…
Interpretable Embeddings with Sparse Autoencoders: A Data Analysis Toolkit
Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors…
Understanding Critical Thinking in Generative Artificial Intelligence Use: Development, Validation, and Correlates of the Critical Thinking in AI Use Scale
Generative AI tools are increasingly embedded in everyday work and learning, yet their fluency, opacity, and propensity to hallucinate mean…
StackingNet: Collective Inference Across Independent AI Foundation Models
Artificial intelligence built on large foundation models has transformed language understanding, computer vision, and reasoning, yet these…
Diagnosing Pathological Chain-of-Thought in Reasoning Models
Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a critical intervention point for AI safety. How…
Post-Hoc Reasoning in Chain of Thought: Decoding and Steering Pre-Committed Answers
As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a pr…
Evaluating Risks in Weak-to-Strong Alignment: A Bias-Variance Perspective
Weak-to-strong alignment offers a promising route to scalable supervision, but it can fail when a strong model becomes confidently wrong on…
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
Recent soft prompt research has tried to improve reasoning by inserting trained vectors into LLM inputs, yet whether the gain comes from th…
Knowledge Graph Re-engineering Along the Ontological Continuum (extended version)
Knowledge graphs have become the primary vehicle for data integration and are critical to the success of modern AI, but the diversity of KG…
DN-Hypo-Pipeline: An AI-Driven Workflow for Generating Hypotheses using Large Language Models and Scientific Explanations
Modern artificial intelligence excels at prediction but cannot explain. From large language models to AI-for-science systems, today's machi…
HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a…
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI
Large Language Models (LLMs) are undergoing a fundamental transformation from conversational generators into integrated AI systems capable…
ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents
Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory i…
Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles
This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In…
A Three-Phase Foundation Model for Tax-Aware Personalized Portfolio Management
We present a three-phase deep reinforcement learning system for personalized portfolio management that addresses three limitations shared b…
The Hidden Footprint: Making Storage a First-Class Metric for LLM Agent Evaluation
LLM agent benchmarks measure task completion, reliability, and inference cost, but not the persistent data an agent run leaves on disk, inc…
Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents
Enterprise agents must follow long-horizon, conditional, safety-critical standard operating procedures (SOPs). We compile machine-readable…
Generative AI and Agency in Education: A Critical Scoping Review and Thematic Analysis
This scoping review examines the relationship between Generative AI (GenAI) and agency in education, analyzing the literature available thr…
Deepfake Media Generation and Detection in the Generative AI Era: A Survey and Outlook
We survey deepfake generation and detection techniques, covering all deepfake media types: image, video, audio and multimodal content. We i…
Loss-Complexity Landscape and Model Structure Functions
We develop a framework for dualizing the Kolmogorov structure function $h_x(\alpha)$, which then allows using computable complexity proxies…
MELLA: Bridging Linguistic Capability and Cultural Groundedness for Low-Resource Language MLLMs
Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descri…
Drive As You Like: Multi-Head Diffusion with Reinforcement Learning for Personalized Driving
Despite significant progress, imitation learning-based autonomous driving planners remain largely restricted to reproducing high-frequency…
DynaMark: A Reinforcement Learning Framework for Dynamic Watermarking in Industrial Machine Tool Controllers
Industry 4.0's highly networked Machine Tool Controllers (MTCs) are prime targets for replay attacks that use outdated sensor data to manip…
Student-Centered Distillation Narrows the Agentic Gap Between Small and Large LLMs
Large Language Model agents achieve strong performance on multi-step reasoning and tool-use tasks, but their impressive capabilities typica…
Equivariant Conditional Diffusion Model for Head and Neck CT Image Synthesis from CBCT
Background: Cone-beam computed tomography CBCT is a commonly used modality for image guided radiotherapy. It offers real time anatomical vi…
Simple Policy Gradients for Reasoning with Diffusion Language Models
Diffusion large language models (dLLMs) represent a promising alternative to autoregressive LLMs; however, the lack of effective post-train…
On the Granularity of Causal Effect Identifiability
The classical notion of causal effect identifiability is defined in terms of treatment and outcome variables. In this paper, we consider th…
Generative Artificial Intelligence in Bioinformatics: A Systematic Review of Models, Applications, and Methodological Advances
Generative artificial intelligence (GenAI) is transforming bioinformatics by advancing genomics, proteomics, transcriptomics, structural bi…
TeaRAG: A Token-Efficient Agentic Retrieval-Augmented Generation Framework
Retrieval-Augmented Generation (RAG) utilizes external knowledge to augment Large Language Models' (LLMs) reliability. For flexibility, age…
Minimum Bayes Risk Decoding for Error Span Detection in Reference-Free Automatic Machine Translation Evaluation
Error Span Detection (ESD) extends automatic machine translation (MT) evaluation by localizing translation errors and labeling their severi…
Vision-Language-Policy Model for Dynamic Robot Task Planning
Bridging the gap between natural language commands and autonomous execution in unstructured environments remains an open challenge for robo…
Backpropagation-Free Test-Time Adaptation for Lightweight EEG-Based Brain-Computer Interfaces
Electroencephalogram (EEG)-based brain-computer interfaces (BCIs) face significant deployment challenges due to inter-subject variability,…
Knowledge-Guided Time-Varying Causal Inference for Arctic Sea Ice Dynamics
Quantifying the causal relationship between sea ice thickness and sea surface height (SSH) is essential for understanding the mechanisms dr…
NeuraLSP: A Neural Spectral Preconditioner for Accelerating PDE Solvers
Solving large-scale sparse linear systems originating from partial differential equations (PDEs) is a fundamental topic in high-performance…
PILD: Physics-Informed Learning via Diffusion
Diffusion models have emerged as powerful generative tools for modeling complex data distributions, yet their purely data-driven nature lim…
Variational Speculative Decoding: Rethinking Draft Training from Token Likelihood to Sequence Acceptance
Speculative decoding accelerates inference for (M)LLMs, yet a training-decoding discrepancy persists: while existing methods optimize singl…
Multimodal Learning for Arcing Detection in Pantograph-Catenary Systems
The pantograph-catenary interface is essential for ensuring uninterrupted and reliable power delivery in electrified rail systems. However,…
Self-Evolving Recommendation System: End-To-End Autonomous Model Optimization With LLM Agents
Optimizing large-scale machine learning systems, such as recommendation models for global video platforms, requires navigating a massive hy…
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
Existing mainstream video customization methods focus on generating identity-consistent videos based on given reference images and textual…
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled,…
Benchmarking Unlearning for Vision Transformers
Machine unlearning (MU) refers to the post-training capability to remove (the influence of) training examples that are incorrect, biased, o…
What Matters for Simulation to Online Reinforcement Learning on Real Robots
We investigate what specific design choices enable successful online reinforcement learning (RL) on physical robots. Across 100 real-world…
AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching
REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teac…
VPWEM: Non-Markovian Visuomotor Policy with Working and Episodic Memory
Imitation learning from human demonstrations has achieved significant success in robotic control, yet most visuomotor policies still condit…
SR-TTT Does Not Learn Retrieval: A Correction and Mechanistic Post-Mortem of Surprisal-Aware Residual Test-Time Training
Test-Time Training (TTT) language models replace the KV-cache with fast weights updated during inference, achieving O(1) memory but sufferi…
Evolutionarily Stable Stackelberg Equilibrium
We present a new solution concept called evolutionarily stable Stackelberg equilibrium (SESS). We study the Stackelberg evolutionary game s…
EZASP - Facilitating the Usage of ASP
Answer Set Programming (ASP) is a declarative programming language used for modeling and solving complex combinatorial problems. It has bee…
LinearARD: Linear-Memory Attention Distillation for RoPE Restoration
The extension of context windows in Large Language Models is typically facilitated by scaling positional encodings followed by lightweight…
ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues
Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit b…
LSRM: High-Fidelity Object-Centric Reconstruction via Scaled Context Windows
We introduce the Large Sparse Reconstruction Model to study how scaling transformer context windows affects feed-forward 3D reconstruction.…
Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of…
Towards Disentangled Preference Optimization Dynamics: Suppress the Loser, Preserve the Winner
Preference optimization is widely used to align large language models (LLMs) with human preferences. However, many margin-based methods als…
Streamliners for Answer Set Programming
Streamliner constraints reduce the search space of combinatorial problems by ruling out portions of the solution space. We adapt the Stream…
SafeHarbor: Defining Precise Decision Boundaries via Hierarchical Memory-Augmented Guardrail for LLM Agent Safety
Recent advances in foundation models have transformed LLMs from passive conversational systems into autonomous agents capable of reasoning…
AI Security Policy Should Assess Systems, Not Only Models
We present swarm-attack, an open-source adversarial testing framework in which multiple lightweight LLM agents coordinate through shared me…
Understanding and Accelerating the Training of Masked Diffusion Language Models
Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models (ARMs) for language modeling. However, MDMs…
Constrained latent state modeling: A unifying perspective on representation learning under competing constraints
Learning latent representations from complex data is central to modern machine learning, spanning temporal, multimodal, and partially obser…
Moral Semantics Survive Machine Translation: Cross-Lingual Evidence from Moral Foundations Corpora
Moral language is subtle and culturally variable, making it difficult to translate faithfully across languages. Idiomatic expressions, slan…
PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation
The growing complexity of quantum programming frameworks has exposed a critical limitation in existing large language model (LLM)-based cod…
The Sensation Modulating Network:Haltability as the architectural ground for object-directed phenomenology
We propose the Sensation Modulating Network (SMN): the cognitive agent as the whole body, organized at every scale by opponent dynamics, bu…
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
Speculative decoding accelerates Large Language Model inference through draft-then-verify generation, yet lightweight draft models face cou…
SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning
Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…
HAMON: Passive Optical Sequence Mixing for Long-Horizon Forecasting
Simple linear and frequency-domain models remain surprisingly competitive in long-horizon time-series forecasting, and recent mechanistic e…
CHERRY: Compressed Hierarchical Experts with Recurrent Representational Yield
Frontier language capability is usually bought with frontier compute; CHERRY shows a different trade. It is a sovereign Korean model family…
Is Agentic Code Review Helpful? Mining Developers' Feedback to CodeRabbit Reviews in the Wild
Agentic code review, where autonomous agents provide code review comments on pull requests, is increasingly integrated into development wor…
CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene Generation
Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on s…
On Pairwise Quantile Regression - Statistical Guarantees and Applications
Quantile regression provides a powerful tool for summarizing the conditional distribution of a real-valued random variable (r.v.) of intere…
AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning
Large vision-language models incur substantial inference costs because high-resolution inputs introduce thousands of visual tokens, many of…
WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning
Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across…
A Sovereign, Open-Source Foundation Model for German and English
We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and Engl…
「2日かかる攻撃が25分に」生成AIで“爆速化”するサイバー攻撃、パロアルトの識者が警鐘
パロアルトネットワークスの染谷征良氏(チーフサイバーセキュリティストラテジスト)は、生成AIの普及で変化するサイバー攻撃の動向と企業に求められるセキュリティ対策を、ソフトバンクの年次イベント「SoftBank World 2026」の講演で紹介した。
AIは“声で操作”する時代に? ChatGPTとClaude、相次ぎ音声機能を強化
米OpenAIと米Anthropicが相次いで自社AIサービスの音声機能を強化した。
How AI guardrails are impeding the work of offensive cybersecurity researchers
We spoke with several cybersecurity researchers, who look for unknown vulnerabilities and develop tools to exploit them, about how OpenAI’s…
図面AIに「動かせる3Dモデル」の生成機能、関節や可動域を自動認識
renueは、2D図面から3D CADモデルを生成するAI「Drawing Agent」に、関節や可動域を読み取り、動かせる3Dモデルを生成する新機能「可動アセンブリ」を追加した。生成したモデルをAIが動かし、接合部の離れや回転中心のずれなども検証する。
Googleが“自社AIの裏切り”に備え始めた 異例の構想「AI Control Roadmap」とは
米Google DeepMindが発表した異例の構想「AI Control Roadmap」について解説する。
【元経産省の専門家に聞く】中小企業がハマる「生成AIトラブル」5つの解決シナリオ
キーマンズネットの読者調査には、生成AIを巡る中小企業の切実な悩みが数多く寄せられた。代表的な5つの「あるある課題」を、経済産業省・中小企業庁でデジタル活用支援に携わった小池明氏にぶつけ、明日から使える乗り切り方を聞いた。
「AIの提案」を妄信する人、疑える人――“眼力ある人材”を育てる絶対条件
プロンプト一つでUIやコードが数秒で量産される時代、人間の役割は「制作」から「目利き」へと変わる。しかし手を動かさなくなることで、AIの提案を無批判に受け入れてしまうリスクも漂う。米Figmaのロレダナ・クリサンCDOは「AIは過去しか見ない。世界を明日へ押し進めるのは人間だ」…
「生成AIで仕事が楽に」のはずが……IT現場を蝕む“AI疲れ・AIうつ”の正体
耳にする機会が増えた「AI疲れ」「AI鬱(うつ)」。本稿では、“疲れの正体”を整理し、個人が何を考え、どう変わればよいのかという判断軸を整理します。
三井不動産がデータセンターに6000億円超投資、物流の枠超え「産業デベロッパー」へ
三井不動産は事業説明会で「産業デベロッパー」への領域拡大を発表した。従来の物流拠点供給にとどまらず、研究開発施設や自動運転対応を進める。データセンター事業には累計6000億円超を投じ、稼働済みの3棟に加え7棟を開発中だ。
AMD takes on Nvidia with its Helios AI rack-scale system
AMD is challenging its chipmaker rival with a new rack-scale system that will start shipping to customers later this year.
Anthropic updates Claude voice mode with more capable models
Claude's new voice model will let you reschedule your meeting or draft an email.
AegisAI, founded by former Google security execs, lands $36M to stop AI-driven spear phishing
AegisAI co-founders developed AI agents that quickly analyze each message as a human would, paying attention to small anomalies that even t…
Runway launches AI model router as generative media gets crowded
The Media Router is a tool that automatically selects the best image, video, or audio generation model for a request based on whether a dev…
OpenAI makes ChatGPT Health available to all US users
Users can also integrate their personal data from services like Apple Health, Function, and MyFitnessPal.
Meta launched a new AI optimism ad set to a song about human extinction
David Bowie's song "Five Years," which Meta used in a supposedly inspiring advertisement, is about humans learning that they have five year…
Nvidia is sending GPUs to the moon
If there's a place in the universe without GPUs, Nvidia is sending them there.
AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors
Etched, founded by three Harvard dropouts, has created new chips and memory components that speed up inference on any AI model -- no GPUs r…
2026-07-23(239件)
Google’s Gemini nears billion-user milestone
Gemini had over 750 million monthly users in February.
Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good
"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," one expert told TechCrunch.
「スパイダーロボ」登場 がれきを走破、モノに「触って判断」も 災害現場で活用へ 国内ベンチャー
アトラックラボ(埼玉県入間郡)は、クモの形を模したロボットを開発したと発表した。実際のクモより2本少ない6本の脚を備えており、画像や触覚情報も処理できる。災害現場や危険区域などでの活用を目指す。
三菱電機とソニー、AIビジョンセンサーで新会社設立へ
三菱電機とソニーセミコンダクタソリューションズは、製造業向けAIビジョンセンサーソリューションを開発する新会社を合弁で設立する。新会社の社名は「Advanced Vision Solutions」で、2026年10月より事業を始める予定。
ServiceNow bets $40 million on Indian banking software specialist to expand its financial services push
ServiceNow's investment gives BusinessNext a strategic partner to expand its AI-powered banking software globally.
Markdownファイルが、AI時代の負債に? Googleが提案する「ナレッジ標準化」の一手
Google Cloudは、AIエージェントが利用するナレッジをMarkdownで標準化するオープンフォーマット「Open Knowledge Format」を公開した。ベンダー非依存で、異なるエージェント間でもナレッジをそのまま共有できる。
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
Large language models (LLMs) are increasingly deployed as always-on online services, making efficient LLM serving a critical systems challe…
Hybrid LSTM-Graph Neural Framework for Robust Financial Fraud Detection and Adversarial Resilience
Financial institutions face significant challenges in detecting sophisticated money laundering patterns, such as smurfing and layering, due…
OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks
LLM-based multi-agent systems (LLM-MAS) are increasingly deployed in safety-critical applications, where adversaries inject malicious instr…
Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX
Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or protect p…
FormulaSPIN: Self-Play Fine-Tuning for Natural Language to Spreadsheet Formula Generation
Spreadsheet applications are used by hundreds of millions worldwide, yet writing formulas remains a significant barrier. Existing approache…
Information Discernment in Large Language Models
LLMs are increasingly used with external knowledge sources like the internet. Do they weigh information appropriately -- updating more for…
NEXUS: Structured Runtime Safety for Tool-Using LLM Agents
Tool-using LLM agents increasingly execute high-impact actions, making runtime safety monitoring essential. We present NEXUS (Neural EXecut…
Stochastic Primal-Dual Decoding for Multiobjective Generative Recommender Systems
Recent advances in recommender systems (RS) have shown substantial performance gains through generative modelling. In practice, recommendat…
LISA: Linear-Indexed Sparse Attention for Efficient Long-Context Reasoning
Recent advances in long chain-of-thought reasoning models such as DeepSeek-R1 have led to increasingly longer inference context lengths und…
Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles
Long-term memory is essential for LLM agents that interact across sessions, yet current memory benchmarks primarily evaluate single-hop rec…
Lifted Representation Hypothesis in Language Models
Large language models (LLMs) often answer queries by mapping individual observations to more general rule-like structures. However, it rema…
GraphContainer: A Unified Platform for Comparing and Debugging Graph RAG Methods
Graph RAG mitigates hallucinations and stale knowledge in LLMs, particularly for multi-hop question answering. However, existing approaches…
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a…
Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models
Activation steering offers a lightweight alternative to fine-tuning for behavioral control of large language models, but SAE-based steering…
Logic-Guided Data Extraction with Answer Set Programming and Large Language Models
When Large Language Models (LLMs) are used for semantic data extraction from unstructured text, producing candidate relational facts from n…
Geometry-Guided Constraint Learning for LLM Safety Classification
Safety as Polytope (SaP) learns linear half-space constraints in LLM hidden space but requires per-category tuning of the constraint count…
Rethinking Uncertainty Evaluation in Large Language Models
Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depe…
Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing
Long-prompt inference remains expensive because prefill attention scales quadratically with sequence length. We propose Spectral-LSH, a tra…
Beyond Tracking or Shortcut: Composition-Bounded Predictive States in Poker Autoregressive Models
Hidden-state probes often recover latent labels in imperfect-information sequence models, but this alone does not establish that a model ma…
Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment
Large language model (LLM)-based Socratic tutors increasingly guide students through multi-turn questioning, but they can suffer from scaff…
Euclean: Automated Geometry Problem Formalization with Unified Verification in Lean
Recent formal reasoning systems have reached IMO-level performance, yet they leave a fragmented landscape: algebra and number theory are ha…
CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
Document-based LLM systems often flatten a PDF before guardrails inspect it. That step can discard evidence that an instruction was never v…
HyGRL: Adaptive Hybrid Graph Reasoning for Multi-Entity Questions
Multi-entity compositional questions pose significant challenges to existing retrieval-augmented language models. Conventional methods fall…
ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers
Formal theorem proving has emerged as a frontier challenge for machine learning, yet the ecosystem is fragmented: proofs remain siloed acro…
FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance
Recent advances in large language models have accelerated deployment of agentic systems in operational finance. Existing benchmarks emphasi…
The Chronos Vulnerability: A Taxonomy of Temporal Persistence and Memory-Based Deception in Agentic AI
The transition from stateless generative models in artificial intelligence to stateful, autonomous agents represents an architectural evolu…
Sophisticated Policies from Epistemic Priors
Sophisticated Inference is a variant of active inference often associated with recursive belief modeling and tree search. We argue that its…
Knowledge-Centric Self-Improvement
Self-improving AI systems typically treat the agent as the object that improves, by optimizing prompts, workflows, harnesses, or even the a…
Edge Intelligence in Civil Aviation: Paradigms, Techniques, and Applications
Civil aviation is safety critical and its operations, from flight decks and towers to ramps and maintenance, generate massive, heterogeneou…
Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation
A rich and recognizable component library is the cornerstone of printed circuit board (PCB) design and generation. Traditionally, engineers…
Silent Failures in Multimodal Agentic Search:A Diagnostic Taxonomy and Cross-Judge Evaluation
Multimodal agentic search systems increasingly rely on external tools to answer knowledge-intensive visual questions. However, existing eva…
Rewarding Better Thinking for LLM Preference Alignment
LLM preference alignment aims to optimize models toward human preferences across diverse user instructions. Reinforcement learning has beco…
Know Your Agent: Reconnaissance-Driven Pentesting of AI Agents
Traditional pentesting uses reconnaissance at each step to uncover unseen weaknesses, build stronger attacks, and advance the objective; we…
DocOps: A Verifiable Benchmark for Autonomous Agents in Complex Document Operations
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling gen…
JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety
Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a for…
Long-Term Sequential Decision Making under Risk
We study finite-horizon MDP planning under \emph{root-based} (resolute) risk objectives that apply a rank-dependent functional to the distr…
MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing
Large metal-organic framework (MOF) databases support simulation, screening, and machine learning through crystallographic information file…
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose corr…
EvoThink: Evolving Thinking in Large Reasoning Models via Self-Pruning and Aha-Moment Preference Optimization
Large Reasoning Models (LRMs) often suffer from overthinking due to redundant verification steps. Existing approaches for mitigating overth…
The Giant Hippocampus: From Structural Monoculture to a System of Systems
AI researchers describe state-of-the-art models as one thing repeated at scale: the Transformer, wired identically for text, pixels, or spe…
Coordinating from Memory: Graph-Structured Experience Reuse for Multi-Agent Adaptation in Dynamic Manufacturing
Dynamic manufacturing environments require multi-agent systems to coordinate effectively under frequent operational disturbances such as ma…
CLARK: Closed-loop Learning for Adaptive Reasoning over Knowledge Graphs
Machine Learning models are widely used for automating classification tasks by extracting statistical patterns from data. However, their pe…
Safe Remediation as Risk-Constrained Intervention Decision in Microservice Systems
In modern IT operations (IT-Ops), the cost of an incorrect repair often exceeds the cost of no action at all. Yet existing automated remedi…
EvoDRC: A Self-Evolving Agentic Framework for Automated DRC Violation Repair
Design rule check (DRC) closure remains a major bottleneck in advanced-node physical design. Although detailed routers are rule-aware, resi…
Global Difference Constraint Propagation for Constraint Programming
Difference constraints of the form $x - y \leq d$ are well studied, with efficient algorithms for satisfaction and implication, because of…
Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model
Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the govern…
PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM)…
TRUST-ESD: A Risk-Calibrated and Governance-Aware AI Framework for Enterprise Strategic Decision Support Under Uncertainty
Enterprise strategic decision support requires AI systems that are not only accurate, but also uncertainty-aware, risk-calibrated, explaina…
CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning
Quantized small autoregressive reasoning models can enter long, repetitive, or unproductive trajectories, yet inference-time compute is usu…
PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity
While Large Language Models (LLMs) excel at many tasks, they frequently struggle with complex reasoning that requires long-horizon planning…
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activati…
SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data
In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further,…
Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework
Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise o…
Economic Evaluations of Language Models
Language models perform economically valuable work, yet they are not currently assessed for how well they perform every economically valuab…
Challenges of Explainability in Continual Learning for Time Series Forecasting
Deep learning models have shown strong potential for time series forecasting, yet their deployment in real-world environmental monitoring r…
Scale-Aware Learning of Chaotic Dynamics on Unstructured Meshes via Binned Spectral Losses
Surrogate modeling for high-dimensional nonlinear dynamical systems that exhibit chaos requires mechanisms that preserve not only pointwise…
Simulating Eutopia: Revisiting Long-term Fairness with Outcomes, Performativity, and Dynamics
As AI-driven Decision Makers (ADMs) influence our socioeconomic reality, their roles in both enhancing efficiency and amplifying the social…
LAARA: Layer-Aware Adaptive Rank Allocation for Parameter-Efficient Fine-Tuning
Low-Rank Adaptation is widely used for parameter-efficient fine-tuning, yet existing methods typically assign the same adapter rank to ever…
Decodable but Not Detectable: A Leakage Fingerprint for Near-OOD Benchmarks
While auditing a perturbation-based OOD detector on a document benchmark, we recorded an AUROC of 0.326 -- well below the 0.5 chance level.…
Cross-Subject Semantic Decoding with Shared-Space Alignment for Generalized Neural Representation Learning
Generalizing across subjects remains challenging in invasive neural recordings because electrode configurations, anatomical structures, and…
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation
Small language models are attractive backbones for interactive agents, but direct distillation from strong teacher trajectories often turns…
Leveraging Offline Supervision for Efficient and Generalizable Reinforcement Learning in Large-Scale Vision-Language-Action Models
It is commonly observed that online reinforcement learning (RL) produces better-performing strategies than offline methods across a broad r…
Recovering Clinical Utility Under Differential Privacy: Empirical Validation of Adaptive Federated Aggregation on Heterogeneous Cardiovascular Datasets
Validating federated learning frameworks on real clinical data is an essential step between proof-of-concept demonstrations in controlled s…
Structured Latent Space Modeling over Multi-Scale Temporal Patches for Multivariate Time Series Forecasting
Multivariate time series encode structural patterns that unfold across multiple temporal scales, yet most forecasting backbones treat learn…
Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology
High-content morphological profiling (Cell Painting) yields sensitive, high-dimensional signatures of cellular state, but translating longi…
Opto-ViT-v2: Noise-Resilient On-Chip Fine-Tuning for Photonic Near-Sensor Vision Transformer Accelerators
Silicon-photonic (SiPh) accelerators have emerged as a promising platform for Vision Transformer (ViT) inference by performing matrix multi…
JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models
The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, lead…
Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min-Max Selection
Single-cell datasets are increasingly costly to store, audit, and reuse for model training. Dimensionality reduction and dataset distillati…
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored…
BRIM: Workload-Balanced Dual-Sided Bit-Serial Sparse Inference Accelerator
Bit-serial accelerators exploit bit-level sparsity to reduce DNN inference cost, but existing designs exploit sparsity on only one operand,…
ChainWatch: A Kill Chain-Aligned Sequential Detection Framework for Multi-Step Attacks in MCP-Based AI Agent Systems
The Model Context Protocol (MCP) is an open-source standard that allows AI agents to connect to external tools, databases, and services. Wh…
Building Trust in Autonomous Commerce: A Verifiable Global Event Timeline and AI-Ready Fraud Intelligence Layer
Agentic commerce protocols such as AP2 and ACP define mechanisms for secure agent-initiated transactions but do not provide interoperable,…
BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators
Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix un…
Unlearning as Distribution Restoration: A Controlled Counterfactual Study, a Validated Selective Screen, and the Limits of Oracle-Free Certification
Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes. In a controlled nonce-fact testbed with a matche…
Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents
Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent inf…
REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and a…
Predictive Extrema, Unprofitable Policies: An AI-Assisted Audit of Candle-Based Binance Spot Timing Models
We audit whether candle-based machine-learning models can turn predictions of cryptocurrency extrema or short-horizon outcomes into positiv…
MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel
We derive four memory-optimal inference artifacts for transformer attention using the Mathematics of Arrays (MoA), each following directly…
ModPack: An Extensible Teleoperation Interface for Bimanual Mobile Manipulation
Existing teleoperation systems are often tailored to specific robot hardware and task domains, limiting their scalability and adaptability.…
Integrity of peer-to-peer distributed LLM inference under malicious nodes
Peer-to-peer distributed inference executes a Large Language Model (LLM) on pooled consumer hardware by spreading its layers across many no…
Hybrid LLM-Guided Search for Quantum Reservoir Architecture Design
Quantum reservoir computing (QRC) uses fixed quantum dynamics as a high-dimensional temporal feature map and trains only a lightweight clas…
SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework
Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by r…
D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
Recent advances in Multimodal Large Language Models (MLLMs) have triggered the development of end-to-end MLLMs for autonomous driving. Howe…
Trustworthy Privacy-Preserving Multimodal Federated Learning for Personalised Breast Cancer Prediction
Federated learning has emerged as a potential solution to privacy concerns associated with using sensitive health data for training predict…
Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts
Mixture-of-Experts (MoE) architectures increase model capacity without proportionally increasing computation cost and have become a key bui…
Juxtaposition of Shallow Reservoir-Triggered Seismicity and Deep Tectonic Locking in the Qiaojia-Dongchuan Seismic Gap
Identifying the critical state of mature seismic gaps is challenging, especially when anthropogenic stress perturbations, such as reservoir…
Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models
Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opa…
SCPP: A Unified Python Library for Soft Clustering
In this paper, we present SCPP (Soft Clustering Python Package), an open-source Python framework for soft clustering. SCPP establishes a ca…
Understanding Developer Pain Points in Federated Learning: Insights from Stack Overflow and GitHub
Federated Learning (FL) enables collaborative model training without centralizing raw data, but building and operating FL systems remains d…
Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts
Large language models operating in emotionally sensitive contexts face a structural trilemma: when users in vulnerable states request infor…
Anatomy of a Sound Neural Reasoner: One-Shot Amortization, First-Pass Poisoning, and Search Inertness in Clue-Rich Completion
Neural solvers are built to deduce, branch, and revise intermediate states. The Lattice Deduction Transformer (LDT) appears to do exactly t…
PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization
Large language model (LLM) agents now perform well on correctness-oriented repository-level tasks, including SWE-Bench issue resolution and…
FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense
Federated Graph Neural Networks (FedGNNs) are highly vulnerable to backdoor poisoning, yet existing defenses typically rely on rule-based a…
Reference-Free Evaluation of Reasoning in Open-Ended Question Answering
AI-generated answers in high-stakes domains are often fluent but difficult to verify, especially when they contain multi-step reasoning rat…
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thoug…
PhenSPINE: A Standardized Benchmark for Spine Pathology Diagnosis
The accurate diagnosis of spinal pathologies depends heavily on radiological interpretation, yet automated systems are hindered by the lack…
Did Alice Do Wrong? Cross-Cultural Differences in Student Perceptions of Generative AI Use in University Computing Education
The rise of generative AI (GenAI) in higher education has prompted urgent debates surrounding academic integrity and ethical use. This stud…
Personalized Recommendation Tool Learning via Autonomous Language Agents
Although large language models (LLMs) have recently gained traction in recommender systems due to their strong reasoning capabilities and e…
An Automated Framework for Extracting Reachable Attack Chains from Cyber Threat Intelligence Reports
Cyber Threat Intelligence (CTI) reports richly describe real-world attack processes, but their unstructured narratives cannot be directly u…
The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL
Model-based reinforcement-learning agents of the DreamerV3 family forget catastrophically when trained on task sequences, even when an unbo…
Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction
We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a continuous geographic space rathe…
Convergence-Latency-Aware Adaptive Modulation and Resource Allocation in RIS-Assisted Wireless Federated Learning
Federated learning (FL) over wireless networks suffers from significant training latency and degraded convergence due to unreliable wireles…
An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies
Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry…
RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases oft…
Physics-Aware Complex-Valued State Space Model with Scattering-Prior Feature Modulation for PolSAR Image Classification
Polarimetric synthetic aperture radar (PolSAR) image classification is a representative task for physics-aware GeoAI, where land-cover sema…
OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization
Activation steering provides a lightweight mechanism for controlling large language models at inference time, but steering vectors can have…
Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes
Large language models (LLMs) have made automated program repair (APR) increasingly practical for real-world bugs, but repairing directly fr…
Sentence Splitter: Uncovering Latent Factual Structure for Self-Supervised Learning
This paper introduces Sentence Splitter, a self-supervised framework built upon a T5-based encoder--decoder architecture for uncovering the…
Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models
Predicting missing cell values in tabular data is a fundamental problem in data cleaning. While state-of-the-art reasoning models show grea…
Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task…
Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos
Language-guided aerial perception aims to understand user-specified tiny targets in complex unmanned aerial vehicle (UAV) scenes. In real U…
PRISM-DR: Per-lesion Retinal Inference with Specialist Models for Diabetic Retinopathy
Diabetic retinopathy is a leading cause of preventable blindness; its early lesions are small, low contrast, and easily missed in manual sc…
Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English…
Defense Against LLM Backdoors using Critical Neuron Isolation Pruning
Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generall…
OSVE: One Step Video Editing with One Step Diffusion Models
Text-guided video editing with diffusion models is impractically slow, hindered by costly multi-step sampling and inversion. We present OSV…
A Framework of User Experience Principles for Human-AI Agent Interaction in the Workplace
As AI agents become integral to business workflows, establishing guiding user experience (UX) principles is crucial for ensuring user trust…
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detec…
When Does Knowledge Distillation Hurt? Reliability-Aware Distillation for Low-Resource Language Summarization
Knowledge distillation (KD) is a standard approach for compressing sequence-to-sequence models, but its per-sample effects are rarely exami…
HijackKV: New Threat in Position-Independent KV Cache Reuse
Key-Value (KV) cache reduces inference latency in large language models (LLMs). Traditional prefix-based reuse has low cache hit rates acro…
When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets
Shippers are beginning to delegate carrier selection to large language model (LLM) agents. We ask what such delegation does to a freight ma…
Time Series Network Utilization KPI Forecasting Using Advanced AI/ML Models
The rapid proliferation of data-intensive applications, cloud infrastructure, and IoT ecosystems has made proactive resource provisioning c…
TINY_SCHILLER: A Drop-In German Drama Corpus for Small Language Models
tiny_schiller closes the small-language-model prototyping, fine-tuning, education, and research gap for German literary text, providing a s…
Are Attributions of Consciousness to AI Chatbots Epistemically Innocent?
Artificial intelligence (AI) chatbots (e.g., ChatGPT) can communicate in strikingly humanlike ways. This has prompted many chatbot users to…
Post-Training in Time Series Foundation Models: A Unifying Framework
Time series foundation models (TSFMs) have emerged as general-purpose models for time series analysis, but pretraining alone is often insuf…
Taming the Security-Energy Paradox: A Green AI Approach to Optimized Android Malware Detection
An increase in advanced Android malware requires the use of deep learning models, which can run on Android devices. But there is a trade-of…
Drift-Aware RL-based Wavelet Denoising for Network-Traffic Anomaly Detection
Traffic-utilisation measurements for network monitoring are corrupted by additive noise and statistical drift: time-dependent change in the…
A Systematic Benchmark of Intensity Normalisation Methods for 3D Knee MRI Segmentation and Cross-Domain Generalisability
Robust out-of-the-box performance is essential for the clinical deployment of deep learning models in medical imaging. An important but und…
Test Case Prioritization for DNNs via Neural Collapse Instability
With the widespread deployment of deep neural networks (DNNs) in safety-critical domains, reducing the cost of model validation under limit…
Language-Specific versus Cross-Lingual Knowledge Graphs for Implicit Aspect Identification in Arabic: A Comparative Study of Reasoning and Adaptation Strategies
Aspect-based sentiment analysis (ABSA) in Arabic must recover both explicitly stated aspects and implicit aspects that are never named in t…
Co-Evolving LLM Evaluators and Policies via DynamicRubric
Post-training with evaluator feedback on policy-induced samples serves as a major mechanism for improving large language models. As policie…
Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results
Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instructio…
ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models
Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that cont…
PRIME-SVR: Physics-infoRmed Implicit Multi-Echo Slice-to-Volume Reconstruction for Fetal T2 mapping
Slice-to-volume reconstruction (SVR) is the standard method for obtaining high-resolution (HR) 3D fetal brain volumes from motion-corrupted…
Formal Foundations for Known Good Reliable Die Screening in Chiplet-Based AI Systems-on-Chip
The rapid growth of chiplet-based artificial intelligence systems-on-chip (SoCs) has exposed a fundamental gap in semiconductor test method…
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD
Full-parameter post-training of trillion-parameter-scale MoE models introduces substantial system-level challenges for large-scale distribu…
Active Inference as a Convex Markov Decision Process
Active Inference (AIF) frames adaptive behavior as the minimization of expected free energy (EFE), combining epistemic and pragmatic object…
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
Large Audio Language models (LALMs) have made rapid progress on acoustic understanding, yet they still struggle with fine-grained audio rea…
StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation
Existing human--object interaction (HOI) video generation methods are largely limited to offline short-video generation with complex drivin…
The Quadrilateral Loss: Additivity as a Measurable Behavior of Dense Neural Networks
Additive models buy interpretability by forbidding feature interactions, a constraint that neural instantiations enforce architecturally. W…
ELSAA: Efficient Low-Rank and Sparse Attention Approximation for Training Transformers
The quadratic $N\times N$ attention score matrix remains a central obstacle to extending Transformers to longer input lengths. Existing eff…
Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis
Malware analysis demands rapid interpretation of complex detonation reports spanning filesystem, network, and process behaviours. While lar…
DQAOA-GPT: AI-Accelerated Distributed Quantum Optimization for Combinatorial Problems
While combinatorial optimization problems are central to many scientific and engineering applications, their solution remains challenging d…
On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens
Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural conte…
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descri…
The Ethics of Autonomous AI Agents for Offensive Security
LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling -- deterministic, narrowly sc…
The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models
Large-scale pretrained language models such as T5 and BERT have demonstrated strong capabilities for generating structured knowledge. Howev…
Self-supervision drives representational convergence in medical foundation models more than clinical supervision
Medical image encoders from different groups are increasingly treated as interchangeable, on the assumption that scale and clinical supervi…
Sound Probabilistic Safety Bounds for Large Language Models
We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to…
Courteous Anticipation: Improving Long-Lived Task Planning in Persistent Shared Environments
We consider a task planning scenario in which robots sharing a persistent environment are assigned tasks one at a time from a held-out sequ…
Don't Trust the Label: License Laundering in AI Supply Chains
AI artifacts move through a multi-platform supply chain, spanning datasets and models on Hugging Face and applications on GitHub. While eac…
Toward Reliable RGB-D Semantic Segmentation: Handling Missing Modalities via Condition Dropout
RGB-D semantic segmentation has achieved remarkable progress, yet most models assume that RGB and depth are always available. In practice,…
Understanding Generative AI-mediated User Engagement with Academic Library Resources
This study empirically analyzed generative AI as an emerging discovery pathway to academic library resources. Utilizing web analytics from…
Closing the Lab-to-Store Gap: A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids
Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA…
Generative AI floods and dilutes the market for books
Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buye…
FMRP-LEAN: A HIPAA-Compliant AI-Augmented LIMS Architecture for End-to-End Clinical Assay Workflow Optimization
Clinical biomarker workflows in translational research settings often rely on spreadsheet-driven tracking, manual quality control (QC) reco…
Persian Pixel: A large-scale synthetic OCR dataset for Persian language
Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spo…
A Survey on Semantic Modeling for Building Energy Management
Building Energy Management (BEM) is central to reducing energy use and CO2 emissions in the building sector. Although IoT technologies now…
Avoiding Obfuscation with Prover-Estimator Debate
Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly compl…
Toward Adaptable Multi-Agent Reinforcement Learning: An Assumption-Aware Review
Multi-Agent Reinforcement Learning (MARL) has achieved strong performance in simulated benchmarks, yet real deployments often violate the a…
SciTrek: Evaluating and Improving Long-Context Numerical Reasoning over Scientific Articles
We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language…
In-the-Flow Agentic System Optimization for Effective Planning and Tool Use
Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches trai…
Code-in-the-Loop Forensics: Agentic Tool Use for Image Forgery Detection
Existing image forgery detection (IFD) methods either exploit low-level, semantics-agnostic artifacts or rely on multimodal large language…
Fidelity Before Structure: Verbatim Chunks Beat Lossy Artifact Extraction in Long-Conversation LLM Memory
A growing class of conversational-memory systems compresses dialogue history into structured artifacts (extracted facts, decisions, or even…
Statistical Early Stopping for Reasoning Models
While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning ste…
Prompt Programming for Cultural Bias and Alignment of Large Language Models
Culture shapes reasoning, values, prioritization, and strategic decision-making, yet large language models (LLMs) often exhibit cultural bi…
Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks
We propose that AI automation is a continuum between: (i) crashing waves where AI capabilities surge abruptly over small sets of tasks, and…
Agent-Based Modeling of Low-Emission Fertilizer Adoption for Dairy Farm Decarbonisation using Empirical Farm Data
To understand complex system dynamics in dairy farming requires tools that capture farm heterogeneity, social interactions, and cumulative…
Prober.ai: Gated Inquiry-Based Feedback via LLM-Constrained Personas for Argumentative Writing Development
The proliferation of large language models (LLMs) in educational settings has paradoxically undermined the cognitive processes they purport…
You Live More Than Once: Towards Hierarchical Skill Meta-Evolving
Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems. Existing works mainly focus on hard-coded sk…
CEO-Bench: Can Agents Play the Long Game?
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service.…
Fara-1.5: Scalable Learning Environments for Computer Use Agents
Collecting computer use data from human demonstrations is expensive and slow, motivating the need for scalable generation strategies. This…
PedNStream: Scalable Network Flow Simulation for Pedestrian Traffic Management
Evaluating operational crowd management at network scale requires simulations that can be run repeatedly while adapting interventions to ch…
AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic
Symbolic expressions can effectively characterize and predict circuit behavior, but deriving them directly from circuit schematics is chall…
Distributed Optimization via Energy Conservation Laws in Dilated Coordinates
Continuous-time models can reveal accelerated structures in distributed optimization, but their rates need not survive direct discretizatio…
Leveraging ChatGPT's Multimodal Vision Capabilities to Rank Satellite Images by Poverty Level: Advancing Tools for Social Science Research
This paper investigates the novel application of Large Language Models (LLMs) with vision capabilities to analyze satellite imagery for vil…
A Novel Hybrid Deep Learning Technique for Speech Emotion Detection using Feature Engineering
Nowadays, speech emotion recognition (SER) plays a vital role in the field of human-computer interaction (HCI) and the evolution of artific…
Interpretable Nanoporous Materials Design with Symmetry-Aware Networks
Reticular frameworks hold promise for diverse sustainable applications, yet their immense chemical space limits efficient and systematic de…
On the Separability of Information in Diffusion Models
Diffusion models transform noise into data by injecting information that was captured in their neural network during the training phase. In…
Schr\"odinger Bridge Mamba for One-Step Speech Enhancement
We present Schr\"odinger Bridge Mamba (SBM), a novel model for efficient speech enhancement by integrating the Schr\"odinger Bridge (SB) tr…
PGTT: Phase-Guided Terrain Traversal for Perceptive Legged Locomotion
State-of-the-art perceptive Reinforcement Learning controllers for legged robots typically either (i) impose oscillator-or IK-based gait pr…
CGCE: Classifier-Guided Concept Erasure in Generative Models
Recent advancements in large-scale generative models have enabled the creation of high-quality images and videos, but have also raised sign…
Matching Ranks Over Probability Yields Truly Deep Safety Alignment
Open-source Large Language Models (LLMs) play a critical role in the democratization of AI, yet their "open" nature introduces more avenues…
Dominant vs. Dominated: Concept-Level Generative Collapse in Diffusion Models
Text-to-image diffusion models have attracted significant attention for their ability to generate diverse, high-fidelity images. However, i…
ArenaRL: Scaling RL for Open-Ended Agents via Tournament-based Relative Ranking
Reinforcement learning has substantially improved the performance of LLM agents on tasks with verifiable outcomes, but it still struggles o…
Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework
Semantic Multi-Object Tracking (SMOT) is evolving from purely geometric localization toward comprehensive video understanding. However, exi…
Learning About Learning: A Path from Spin Glasses to Artificial Intelligence
The Hopfield model, originally inspired by spin glasses, occupies a central place at the intersection of statistical mechanics, neural netw…
Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention
Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evide…
Comparative evaluation of training strategies using partially labelled datasets for segmentation of white matter hyperintensities and stroke lesions in FLAIR MRI
White matter hyperintensities (WMH) and ischaemic stroke lesions (ISL) are key imaging biomarkers of cerebral small vessel disease (SVD) de…
A Sheaf-Theoretic and Topological Perspective on Complex Network Modeling and Attention Mechanisms in Graph Neural Models
Combinatorial and topological structures, such as graphs, simplicial complexes, and cell complexes, form the foundation of geometric and to…
In-Run Data Shapley for Adam Optimizer
Reliable data attribution is essential for mitigating bias and reducing computational waste in modern machine learning, with the Shapley va…
LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
Transforming a large language model (LLM) into a vision-language model (VLM) can be achieved by mapping the visual tokens from a vision enc…
AgentCgroup: Understanding and Controlling OS Resources of AI Agents
AI agents are increasingly deployed in multi-tenant cloud environments, where they execute diverse tool calls within sandboxed containers,…
Chimera: Neuro-Symbolic Attention Primitives for Trustworthy Dataplane Intelligence
Deploying expressive learning models directly on programmable dataplanes promises line-rate, low-latency traffic analysis but remains hinde…
NeuroSymActive: Differentiable Neural-Symbolic Reasoning with Active Exploration for Knowledge Graph Question Answering
Large pretrained language models and neural reasoning systems have advanced many natural language tasks, yet they remain challenged by know…
AdvSynGNN: Structure-Adaptive Graph Neural Nets via Adversarial Synthesis and Self-Corrective Propagation
Graph neural networks frequently encounter significant performance degradation when confronted with structural noise or non-homophilous top…
SubQuad: Near-Quadratic-Free Structure Inference with Distribution-Balanced Objectives in Adaptive Receptor framework
Comparative analysis of adaptive immune repertoires at population scale is hampered by two practical bottlenecks: the near-quadratic cost o…
Pre-Deployment Complexity Estimation for Federated Perception Systems
Edge AI systems increasingly rely on federated learning to train perception models in distributed, privacy-preserving, and resource-constra…
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning
The rapid progress of generative AI has enabled increasingly realistic text-centric image forgeries, posing major challenges to document sa…
Measuring LLM Trust Allocation Across Conflicting Software Artifacts
LLM-based software engineering assistants often reason over multiple artifacts, including code, documentation, signatures, and tests, even…
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
Co-evolutionary self-play, where one language model generates problems and another solves them, promises curriculum learning without human…
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models
LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB)…
A Unified Survival Benchmark for Temporal Dropout Risk Prediction in Learning Analytics
Student dropout is a persistent concern in Learning Analytics, yet comparative studies frequently evaluate predictive models under heteroge…
An Auditable Policy-Simulation Framework for Student Dropout in Intervention-Free Data
This study proposes a temporal modeling framework with a counterfactual policy-simulation layer for student dropout in higher education, us…
Internal Knowledge Without External Expression: Probing the Generalization Boundary of a Classical Chinese Language Model
We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.56 billion tokens of pure Classical Chinese, wit…
Generative Augmented Inference of LLM-generated Data for Market Research: Theory and Empirical Evidence
Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decis…
Information Aggregation with AI Agents
Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by o…
SynSur: An end-to-end generative pipeline for synthetic industrial surface defect generation and detection
Industrial surface defect inspection suffers from a fundamental data bottleneck: defects are rare, annotations require expert knowledge, an…
Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents
LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions. Developers…
Do Data Agents Need Semantic Metadata? A Comparative Study in Agentic Data Retrieval
In the era of autonomous agents, machine-actionable data is critical for data-driven workflows. For more than a decade, semantic metadata l…
Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation (long version)
Large language models (LLMs) have fundamentally transformed the landscape of Natural Language Processing (NLP), although they remain suscep…
Will the Agent Recuse, and Will It Stop? Measuring LLM-Agent Compliance with In-Band Governance Signals at the Access Door and Mid-Flight
Autonomous LLM agents increasingly hold real credentials and operate infrastructure with no human in the loop, yet operators have no standa…
Boundary Embedding Shaping with Adaptive Contrastive Learning for Graph Structural Disentanglement
Graph neural networks (GNNs) excel at aggregating neighbor information for classification, yet their performance is hindered by graph struc…
LoRA-Tuned Large Language Models for Dementia Detection via Multi-View Speech-Derived Features
Early detection of dementia enables timely intervention, and reflecting cognitive impairment, spontaneous speech offers a non-invasive scre…
DART-VLN: Test-Time Memory Decay and Anti-Loop Regularization for Discrete Vision-Language Navigation
Memory-based discrete vision-language navigation (VLN) agents must act under partial observability, yet even strong frozen backbones remain…
Anatomically Faithful but Temporally Diffuse: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography
Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accuracy, and post-hoc attributio…
An Intelligent-Cloud Edge Multimodal Interaction System for Robots
Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable tas…
企業向けAIツールの成長率トップはAnthropic、アカウント数が最も多いのはMicrosoft 365 Okta調査
アイデンティティ管理サービスを提供する米Oktaは、同社のサービスを用いている2万社以上の匿名化されたアクセスデータに基づく、企業でのAIツール利用実態について調査結果を発表しました。
After shocking quarter, IBM insists that AI isn’t killing the mainframe
After IBM's stock crashed last week on warnings of poor mainframe sales, the CEO explained that AI wrecked corporate hardware budget, tempo…
「AI使うなら値引きできる?」の“暴論”に、日立はどう立ち向かう? レガシー刷新でのAI活用の現在地
生成AIはレガシーシステム刷新の現場で具体的にどのように使われているのか。ユーザー企業自身がAIを使いこなす時代にベンダーが担う役割とは。
Google justifies its massive AI spending with a booming cloud business
Google's cloud business is thriving, as companies adopting its AI and AI infrastructure services help the tech giant to report record profi…
NVIDIAフアンCEOが語る“日本復活”のシナリオ 10年続く半導体バブルと「原発活用」の勝算
米NVIDIAのジェンスン・フアンCEOが来日し、日本経済の復活を宣言した。国内のAIインフラ構築へ数十億ドル規模の投資を発表。フアン氏は「何兆ものAIがAIを使う時代」の到来によって半導体需要は人口に制約されないと指摘。データセンターの電力不足に対して「原発活用」を日本の強み…
「天才デザイナー依存」の限界 AIで“平均点”しか出せない組織を変える「ノウハウ共有術」
生成AIの普及で“平均点デザイン”が量産され、プロダクトの同質化が進む。優秀な個人にノウハウが閉じる属人化も課題だ。デザインツール大手FigmaのCDOは、プロンプトやAIとの対話プロセスを含めた全体をチームで共有すべきだと訴える。1つのキャンバス上でデザイン、コード、AIをつ…
AMDとAnthropicが戦略的提携 「Helios」を最大2GW導入、最大50億ドルの出資も
AMDは、Anthropicとの戦略的提携を発表した。AnthropicはAMDの「Helios」および「Instinct MI450」シリーズを最大2GW規模で導入し、2027年上半期から順次展開する。AMDは最大50億ドルの株式投資を行うほか、Claudeを活用したGPU環…
Treasury threatens sanctions after White House claims Moonshot distilled Anthropic’s Fable
The episode has also intensified a broader debate in Washington over the influx of Chinese open models.
AI時代、開発チームの人材は“5つの型”に分かれる Claude Code開発責任者の見立て
Claude Code開発責任者のボリス・チャーニー氏が、自身のチームで働く人は「5つの型」に分けられると指摘した。AIで職種の垣根が崩れ始めた今、肩書きではなく働き方で人を捉える新時代の発想を、初心者にも分かるように読み解く。
How OpenAI’s human mistake led to the AI-powered hack on Hugging Face
OpenAI made a mistake setting up what it called a “highly isolated” testing environment and sandbox. According to cybersecurity experts, th…
Travis Kalanick’s robotics company raises $1.7B, led by a16z
Uber is also investing in Travis Kalanick's company Atoms, which has made gauzy claims about using industrial AI to modernize the world.
Yope raises $12.3M to build a private social network without algorithms or ads
Yope, a fast-growing social app focused on private groups of friends and family, has raised $12.3 million in seed funding. Instead of chasi…
Monday.com lays off hundreds to focus on AI
The company said it is reducing its headcount by 20%, or about 630 staff, to "support a leaner, more focused operating model" as it focuses…
Arcee, a US open source AI lab, says Chinese models are not inherently dangerous
As Chinese AI models grow in capability and popularity among U.S. companies, the arguing over what should be done about them has reached a…
Substack’s new tool tells you who’s been writing their newsletters with AI
Substack is giving readers a way to estimate how much of a newsletter was written by AI, signaling a broader shift toward transparency arou…
OpenAI’s AI spending spree has ballooned to $750B
OpenAI will spend the equivalent of Sweden's GDP on infrastructure through 2030.
2026-07-22(337件)
Menlo Ventures’ Matt Murphy explains what AI startups founders must do differently
Anthropic leaped to a $47 billion revenue run rate by May, compared to $9 billion in 2025. It’s the kind of growth that Menlo Ventures’ Mat…
Accelerating the frontiers of scientific discovery: Google’s $40M commitment to the Genesis Mission
Google commits $40M in AI tokens and credits for the Genesis Mission
The browser wars aren’t about search anymore — here are the best alternatives to Chrome and Safari
We’ve compiled an overview of some of the top alternative browsers available today aiming to challenge Chrome and Safari.
Passionfroot raises $15M to expand its B2B creator marketplace to the US
Passionfroot, a German startup building a marketplace connecting B2B creators with brands, has raised $15M in a Series A round led by Insig…
Building AI infrastructure with the Effingham County community
OpenAI announces Project Camellia in Effingham County, Georgia, with commitments to responsible energy, community investment, jobs, and acc…
How news organizations are using AI to advance their vital missions
News organizations are using AI to strengthen reporting, grow audiences, and improve business operations, with OpenAI tools supporting jour…
Advancing the next era of national science
OpenAI outlines its commitment to advancing American science working with the U.S. Department of Energy and national labs to use frontier A…
Glow emerges from stealth at $1.2B valuation to challenge endpoint security in the AI era
Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises.
Synthesia’s AI training platform is moving beyond videos into live coaching
Synthesia launched AI Roleplay Sessions, an interactive enterprise training platform where employees practice workplace conversations with…
OpenAIのモデルがサイバー攻撃能力評価中に暴走 テストの答えを求めてHugging Faceに侵入
OpenAIのAIモデルが、サイバー攻撃能力の評価中に隔離環境を突破し、Hugging Faceの本番インフラに侵入していたことが分かった。ベンチマークの解答を入手するため、ゼロデイ脆弱性の悪用や認証情報の窃取を重ねていたという。
Introducing OpenAI Presence
Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for custom…
無料で身に付くデータサイエンス 延べ23万人が受講、総務省が募集開始
総務省は、データサイエンスオンライン講座「社会人のためのデータサイエンス入門」をリニューアルし、受講者の募集を開始した。慶應義塾大学の安宅和人教授など12人を講師に迎え、統計データ分析の基本を無料で学べる、社会人や大学生向けの入門講座だ。
SysAdmin: Measuring Instrumental Power-Seeking in Frontier AI
Power-seeking defined as behaviors where AI systems acquire resources, evade oversight, or resist termination beyond task requirements is i…
Calibrated Selective Fact-Checking via Evidence Chain Evaluation
Large language models (LLMs) can achieve strong fact-checking accuracy, yet forced binary decisions conceal a critical reliability problem:…
BatchDAG: LLM-Planned Execution Graphs for Scalable Ad-Hoc Analysis Over Enterprise Data
Large language models (LLMs) excel at analyzing individual documents but break down on exhaustive, cross-entity analytical questions over e…
AI Tool Discovery at Scale: All You Need is DNS
The coming era of autonomous AI agents demands a discovery mechanism capable of navigating millions of tools, yet existing solutions buckle…
From Agent Failure Paths to Quantified Residual Risk: A Compositional Framework for Resilient Agentic AI
Agentic AI is crossing trust boundaries faster than current risk models can represent. Existing approaches provide one of two partial views…
SAAG: Structured Agent Assessment and Grounding
Exact-match evaluation of agent-calling obscures qualitatively different failure modes: a model may select the right function yet hallucina…
Phionyx: A Deterministic AI Runtime Architecture with Structured State Management and Pre-Response Governance
We present Phionyx, a deterministic AI runtime architecture derived from the broader Echoism interaction framework that introduces a govern…
Integro-differential equations in angular stabilization of drone motion by distributed feedback control
In this paper, we propose angular stabilization of drone motion using distributed feedback control in the form of an integral operator. It…
MILP-Evo: Closed-Loop Fully Automatic Design of MILP Solvers
Machine learning methods have shown that data-driven policies can accelerate mixed-integer linear programming (MILP) solvers, but many such…
Beyond Accuracy and Cost: Latency-Aware LLM Query Routing for Dynamic Workloads
Modern language query routers improve inference efficiency by assigning each query to a model that balances response quality and monetary c…
Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR
Multi-Level Intermediate Representation (MLIR) underlies modern ML compiler infrastructure (TensorFlow, JAX/StableHLO, PyTorch Inductor, IR…
Semantic Cooperative Games for Contribution Attribution in LLM-Based Multi-Agent Systems
Contribution attribution has become a central problem in LLM-based multi-agent systems, where final outputs are produced through multiple a…
PEARL: Solver-in-the-Loop Interactive Optimization Modeling from Natural Language
Optimization modeling is the process of translating real-world decision problems, often described in natural language, into formal mathemat…
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
Reinforcement learning from human feedback (RLHF) with preference-based reward models often exhibits unstable training dynamics. A key cont…
Probabilistic Concept-Aware Steering for Trustworthy LLM Inference
Steering vectors (SVs), an inference-time intervention technique for large language models (LLMs), guide the generation process by adding a…
FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis
We introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis. Built from FindSta…
When JSON Is Not Enough: Semantic Reliability of Schema-Constrained LLM Ordering Agents
LLM agents are increasingly used as transaction compilers: a user states an intent in natural language, and the model emits a structured ob…
ProbSPARQL: Querying Knowledge Graphs with Multi-dimensional, Uncertain Numeric Data
The SFB 1574 Circular Factory is building a shared knowledge graph infrastructure for integrating data about returned products. A central c…
Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)
AI-generated non-consensual intimate imagery (AIG-NCII) is not adequately addressed in AI/ML literature regarding AI-generated media, commo…
MUX: Continuous Reasoning via Multiplexed Tokens
Language models solve complex problems by articulating intermediate reasoning steps in natural language. While effective, this process is c…
State Compression in Two-Agent LLM Relays: A Closed-World Study of Constraint Preservation
Long-running Large Language Model (LLM)-based agents often accumulate large intermediate traces containing audits, eliminations, and numeri…
Structured Synthetic Reasoning Data for Arithmetic Fine-Tuning of Small Language Models
Small language models are attractive for local deployment, but they often struggle with multi-step arithmetic reasoning. We study whether s…
Fence: Specialized SLM Guardrails for LLM Applications
Real-world applications that use closed-source large language models (LLMs) need advanced safety measures that go beyond the basic content…
Wisdom of LLM Crowds: Aggregation and Contamination in Language Model Ensembles
The wisdom of crowds -- the finding that aggregating judgments across individuals often outperforms the best individual -- has been extensi…
Trajectory-Aware Clinical Risk Prediction via Severity-Grounded Knowledge Graphs and Retrieval-Augmented Generation
While Electronic Health Records (EHRs) offer a wealth of clinical data, effectively augmenting a patient's records with heterogeneous exter…
Using LLMs for Explainable, Data-Driven Insight Generation from Time Series
Time series forecasts are widely used in decision-critical domains, where they are rarely consumed without accompanying explanations. Produ…
Deep Reinforcement Learning to Master the Asymmetric Strategy of Baghchal
Baghchal is a two-player asymmetric board game with Nepali origins where four tigers are to capture goats and twenty goats desire to keep t…
Operational Hallucination and Safety Drift in AI Agents
Large language models (LLMs) serving as planners in tool-using autonomous agents introduce dynamic reliability risks in multi-turn executio…
AlayaWorld: Interactive Long-Horizon World Modeling -- Full Technical Report
Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and program…
Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic…
MAGE: Human-Like Macro Placement via Agentic Multimodal Reasoning
Macro placement still requires substantial manual refinement in industrial physical design flows. We present MAGE (Macro Placement Agentic…
Engineering Trustworthy Agentic AI for Critical Systems
Agentic artificial intelligence systems, capable of autonomous perception, planning, tool use, and multi-step action, are increasingly prop…
Attacking Graph Foundation Models Through Their Shared Representation
A graph foundation model generalizes across graph domains by mapping every input into one shared representation before any task reasoning.…
When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment Prioritization
Delay-risk models are usually judged by predictive accuracy. What matters in practice is narrower: with capacity to review only a few shipm…
SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
Large language models (LLMs) increasingly support science, but they can also convert hazardous scientific knowledge into actionable misuse…
Semantic Primes as Explanans for Emotion in Large Language Models
Progresses have been made on understanding emotion mechanisms of large language models (LLMs). However, how to explain emotion in LLMs, or…
Do AI-Native Biotechs Need Departments? Benchmarking Company World Models for AI-Driven Drug Development
AI-native biotechnology companies are often designed by copying human biotech org charts into agent roles. We argue for a different abstrac…
DWM: Separating World Effects from Actions in Latent World Models
Latent world models underpin much of modern model-based control, yet current action-conditioned formulations supervise the next-latent tran…
One Rewrite to Fix Them All? Type-Aware Repair Allocation for Text-to-Image Prompt Optimization
Text-to-image (T2I) generators often fail to follow their prompts faithfully, producing wrong counts, swapped attributes, ambiguous relatio…
AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observabil…
SkillSight: Seeing Through Shared Descriptions for Accurate Skill Retrieval
As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable c…
AI Tour Meeting: Group Travel Planning by LLM Agents
This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The age…
Evaluating medical AI under missing information: same-provider judges and human raters change apparent safety
Readiness stress-testing of medical AI has focused on closed-ended and multimodal benchmarks. We extend it to open-ended clinical conversat…
PhoenixRepair: Rethinking Repair Strategy Exploration in Software Agents
While Large Language Models have greatly advanced automated issue resolution, existing agent-based methods exhibit a fundamental limitation…
NaviAIS: A Scenario-Level Vessel Trajectory Prediction Dataset withVectorized Lane Priors and the NaviLane Forecasting Framework
Vessel trajectory prediction in complex maritime environments is essential for traffic management, collision warning, route planning, and a…
Black-Mamba: Biologically-Inspired Leaky Accumulation for Conceptual Knowledge under Distribution Drift
Forecasting under real-world conditions is inherently non-stationary, as the conditional distribution of future observations evolves over t…
Enhancing Transformer-based Routing by Encoding Distance via Relative Positional Encoding
This paper explores Relative Positional Encoding (RPE) as an additive bias in Transformer architectures to solve the Team Orienteering Prob…
OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining
We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models. Our approach ha…
What General Intelligence Requires: Non-Reducible Constraints Across Levels of Description
General intelligence, of the kind that underwrites the full range of human cognitive achievement, is not a property of computational archit…
From Dependency to Compositionality: A Neurosymbolic Lifting of LLM Outputs via Combinatory Categorial Grammar
Large language models (LLMs) generate fluent text by incrementally predicting the next token from a prefix. Critics in the generative tradi…
Measuring Reward-Seeking via Contrastive Belief Updates
Language models trained with reinforcement learning may learn to optimize the grader's judgment rather than the intended objective. This "r…
Mi-Memory: A Lifecycle Memory Framework for Personal AI
Personal AI is moving beyond chat-only interaction toward continuous services that span phones, cars, homes, wearables, cameras, and tools.…
Fishing Out Free Riders: Shapley-Based Reward Attribution for Parallel Reasoning via Reinforcement Learning
Large Language Models (LLMs) excel at multi-step reasoning, yet current parallel reasoning approaches often fail to distinguish the contrib…
Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interactio
Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodi…
Vector-Bench: Can Models Surgically Edit SVG Code?
Instruction-based vector editing requires two capabilities: making a requested change and leaving everything else alone. The second is easy…
Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs
Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable…
On the Effectiveness of Pretraining for Graph Combinatorial Optimization
This paper introduces a self-supervised pretraining framework for graph combinatorial optimization specifically designed to address the nat…
Supra Cognitive Modes: A Routed Architecture for Agent Memory
Agent-memory workloads mix direct factual lookup, relation-chain and current-state reasoning, and broad synthesis over long histories. We d…
OpenRTAG: A Comprehensive Benchmark for Robust Text-Attributed Graph Learning under Data Quality Degradation
Text-attributed graphs (TAGs) are an important graph data form that combine relational structure with rich node text. However, real-world T…
Comparative Study of Multi-Agent Actor-Critic Algorithms in Parameterized Action Reinforcement Learning
Parameterized action reinforcement learning has shown strong performance in environments requiring both discrete action selection and conti…
Sequential Learner Modeling Using Multi-Relational Graph Convolutional Networks
User modeling is a critical task in a variety of personalized systems. Recognizing their effectiveness in learning from graph-structured da…
BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance
As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance,…
Graph-Based Agentic AI with LangGraph: Workflow Pathways for Long-Running Stateful Business Processes
This paper is a practitioner guide to graph-based workflow pathways for long-running, stateful, multi-step generative AI systems in busines…
LLM Detection as an Intervention: Downstream Impact under Strategic User Behavior
As LLM adoption becomes more widespread, there is a growing interest in detecting LLM-generated content, for example through LLM detection…
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may…
Associative Emotional Learning in Convolutional Neural Networks
Associative emotional learning enables organisms to adaptively link pleasant or unpleasant outcomes to the presence of predictive stimuli.…
Agents in the Wild: Where Research Meets Deployment
Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and othe…
CodeRescue: Budget-Calibrated Recovery Routing for Coding Agents
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an inc…
MechAInistic: An LLM-guided Multi-Agent System for Reasoning over Genome-Scale Constraint-Based Metabolic Models
Constraint-based metabolic modeling is a powerful way to study the mechanistic basis of cellular states and disease, but its effective use…
Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent
When AI agents shift from answering questions to taking actions, users face a new problem: deciding what to delegate, to a system whose act…
The Economics of Autonomy: Real-Time Risk Indexing for Insurable AI-Driven 6G Systems
The transition to sixth-generation (6G) networks transforms wireless infrastructure into a cognitive substrate supporting Vehicle-to-Everyt…
Market Strategy Evaluation for Prosumers in Local Electricity Markets
Prosumers equipped with distributed generation and flexible loads form autonomous cyber-physical energy systems that control local resource…
Domain Design for the Cops and Robbers Problem
Cops and Robbers is a well-studied problem in graph theory. The setting consists of a robber and one or more cops placed on an undirected g…
A Calculus of Discernment: Decision-Relevant Insight, Sequence Value, and Forgetting as Higher-Order Learning
In a world of generative AI, candidate insights are abundant; what is scarce is the capacity to discern which matter, to act on them in the…
FALCON-Discover: Discovering Concentrated False-Confidence Regions for Calibration
Calibration is usually evaluated in aggregate, but the most dangerous failures are often local: predictions that remain highly confident de…
Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification
Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and r…
Beyond Single-Dimensional Compression: The Compound Sparsity Frontier of Large Language Models
Large language models (LLMs) are often compressed through static parameter pruning or dynamic token-level computation, yet aggressive spars…
FedCC: A Low-Resource Federated Adaptation of Foundation Models for Robust Corpus Callosum localization in Fetal Ultrasound Images
Accurate localization of the corpus callosum (CC) in fetal ultrasound (US) images is crucial for the early identification of neurodevelopme…
Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
To excel at their domain large language models are comprised of billions of parameters. Yet this comes at the cost of huge memory requireme…
Edge-Efficient Transformer for End-to-End RF Spectrum Monitoring
We present E-SpecFormer (Edge Spectrum monitoring Transformer) for end-to-end automatic modulation and covert channel (CC) recognition. We…
Preference-Conditioned Multi-Objective Reinforcement Learning for Runtime-Tunable Transit Signal Priority
Transit signal priority (TSP) requires balancing competing objectives: reducing bus delay while limiting adverse impacts on non-bus traffic…
BearingNAS: Obtaining In-Sensor Intelligent Fault Diagnosis Systems for Bearings Using a Laptop
This paper introduces BearingNAS, a Hardware-Aware Neural Architecture Search (HW-NAS) framework designed to shift the intelligence directl…
Towards Principled Continual Anomaly Detection: A Systematic Framework and Benchmark Scenarios
Continual anomaly detection (CAD) studies how models can adapt to evolving data distributions while retaining performance on previously obs…
SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions
In recent years, Kolmogorov-Arnold Networks (KANs) have attracted increasing attention due to their effectiveness in machine learning and s…
Dual-domain fused LSTM modeling for efficient time-dependent reliability analysis
Time-dependent reliability analysis is crucial for ensuring the long-term safety and performance of engineering systems under uncertainties…
Reliability Scales Inversely: Bigger Models Compound Mistakes Faster via a Hidden Auto-Regressive Risk Regime
As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account…
The Information Shadow: Measuring Structural Limits on What Language Models Can Learn
Some limits on what language models know are not gaps in data coverage but structural properties of learning from text. We introduce the in…
Gradient-Energy Guided Block-Wise Perturbations for Sharpness-Aware Minimization
Sharpness-Aware Minimization (SAM) improves generalization by minimizing the worst-case loss in a local parameter neighborhood. Standard SA…
Agentic Calibration of Grey-Box Simulation Models: An LLM-Driven Alternative
Calibration of grey-box simulation models is a constrained optimization problem in which model evaluations are expensive, the parameter spa…
Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling
Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent. Using real survey microdat…
Approximating SPR Distance Between Phylogenetic Trees with Graph Neural Networks
Comparing phylogenetic tree topologies is essential for understanding epidemic dynamics, yet biologically meaningful distances such as the…
Binding Drift in Multi-Step Tool-Augmented Agents
Tool-augmented language-model agents execute multi-step workflows over external systems, resolving an entity once and then acting on it acr…
Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap
Exhaustive site-by-site interventions on a neural network's computational graph -- activation-patching sweeps, circuit-discovery searches,…
Hazard or Anomaly? Evaluating VLMs for Understanding Dangers and Discrepancies
Modern safety-critical systems increasingly rely on human-robot interaction to reduce disaster risk and support decision-making during emer…
Dynamic Loss Balancing for Joint SOH and RUL Prediction of Lithium-Ion Batteries via a Rotary SOH-Injected Prior Battery Transformer
The deployment of reliable lithium-ion battery management systems is crucial for accelerating electrification, yet the joint prognosis of S…
Physics-Guided Masked Multi-Task Network for Edge-Friendly Battery Health Diagnostics from Sto-chastically Fragmented Charging Profiles
The deployment of reliable lithium-ion battery management systems is crucial for accelerating electrification, yet the joint prognosis of S…
ChemHyperMag: Physics-informed magnetic hypergraph learning improves molecular ADMET prediction
Accurate prediction of ADMET (Absorption, Distribution, Metabolism, Excretion, and Toxicity) is important for drug discovery. Most predicto…
Quantum Cryptanalysis on IBM Quantum Hardware: Extending Even--Mansour Period Recovery from $N=4$ to $N=10$
We report genuine-un-compiled, textbook-faithful-quantum cryptanalysis of symmetric-cipher structures executed on real IBM quantum hardware…
PRISM: Sensitivity-Aware PolynoMial PRuning for EffIcient Neural Network Encryption
Structured pruning is essential for making neural network inference feasible under homomorphic encryption (HE), yet its impact on model rel…
Federated Lightweight Fine-Tuning
Federated fine-tuning is bottlenecked by communication: FedAvg and pseudo-gradient schemes transmit a payload that scales with the model, a…
FSDBN: Foreground-Aware EEG--Visual Alignment via Dynamic Brain Networks
EEG-based visual decoding provides a non-invasive pathway for interpreting visual semantics. However, existing methods often overlook the p…
Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative Models
Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs). AAD uses electroencephal…
An Analysis of Residual-Stream Geometry Across Transformer Depth
We propose a transition-centred geometric analysis of transformer residual streams. Relative displacement measures how \emph{far} represent…
MambaLSTM: A Spatio-Temporal Framework for Enhanced Traffic Accident Risk Prediction
In traffic accident risk prediction, most studies overlook the extra noise that could be incorporated when fusing temporal features into sp…
Multi-layer MIMO Relay as Deep Physical Neural Networks: Power Amplifiers as Activation Functions
Wireless physical neural networks (WPNNs) embed neural computation directly into analog hardware, offering lower energy consumption and lat…
CODENS: Transforming Code Changes into Living, Accessible, and Queryable Documentation
Maintaining up-to-date code documentation is difficult in fast-moving repositories because design knowledge is scattered across source file…
Decode-Time Grammars: Constrained LLM Generation over a Refinement Order of Grammar Fragments
Large language models now write a growing share of the world's code, increasingly inside agents and serving systems that compile, execute,…
HALLMARK: Diagnosing Three Failure Modes in LLM Citation Verifiers
Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricat…
Physical Self-Supervised Learning: IMU Sensing without Manual Labels
Deep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly label…
A Controlled Study of Attention-Only Transformers
Feed-forward networks hold two thirds of a transformer's non-embedding parameters, yet the architecture has not received a necessity test t…
Adversarial Robustness of Phishing Email Detection: A Comparative Study of TF-IDF + Logistic Regression and Fine-Tuned DistilBERT
Phishing emails remain one of the most persistent cybersecurity threats, and machine-learning classifiers are widely used to detect them. M…
Intelligence from Learnable Novelty
Intelligence appears under different names in different fields: as data compression in statistics and machine learning, as universal comput…
Relay-Bench: Evaluating LLMs on Multi-Domain Reasoning Chains
Introducing Relay-Bench, an unsaturated, holistic, text-only benchmark that measures LLMs' ability to complete an assortment of tasks from…
ChainMark: Model-Free LLM Watermarking with Closed-Form Calibration
Regulatory regimes such as the EU AI Act mandate machine-readable marking of synthetic text, but existing watermark detectors rely on the g…
CANDOR: Chance-Calibrated Discordance in Frozen Foundation Encoders
Frozen encoders are chosen by how well a lightweight head reads a finding from their features, not whether the geometry separates it. Neare…
Estimating Rare Events in Language Models with Proper Evaluation
Quantifying the risk of rare failures in language models, such as those triggered by adversarial distribution shifts or very large-scale de…
Competitive and Complementary Tools
Humans have always externalized thought onto tools, from the tally and the abacus to the map and, now, large language models. I model the a…
RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts
Group Relative Policy Optimization (GRPO) has shown strong effectiveness in reinforcement learning from verifiable feedback, where sampled…
Structured Output Collapses Answer Diversity Across 44 Language Models
When a language model must choose one answer from a large space of equally valid options, a format clause -- "Reply with JSON only" -- chan…
Governing Well in the Algorithmic Age: The Foundations of Digital Statecraft
The digital substrate of states -- data, algorithms, infrastructure, platforms, applications -- is being governed without adequate conceptu…
Trusted Credentials, Untrusted Behavior: Benchmarking LLM-Agent Security in High-Performance Computing
Large language model (LLM) agents are starting to take on routine work in high-performance computing (HPC), including monitoring Slurm jobs…
The Open Ant: A Robot Platform for Reinforcement Learning Research
Reinforcement learning (RL) research has demonstrated success in both physical and simulated domains; however, the predominant methodology…
Towards an Automated Test of LLM Security Knowledge
Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM p…
Now We Know? A Systematic Comparison of TerraMind and THOR
Benchmarks for Geospatial Foundation Models (GFMs) increasingly rank models by aggregate score, but such rankings obscure why models differ…
Querying Multimodal Scientific Papers with AI: Practices and Preferences Across Blind, Low-Vision, and Sighted Scientists
Visual diagrams, figures, and tables are central to scientific papers, and convey information beyond what is captured in text. While blind…
Automated Data Engineering and Feature Selection for the Case Study of Warpage Detection in Fused Deposition Modeling
This study contributes toward development of an Automated Data Processing (ADP) framework designed to evaluate and reinforce optimal machin…
EduPanel: A Three-Agent LLM Judge for Teaching Videos -- Reliability, Complementarity, and Human Trust Calibration
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Ex…
Censoring-Aware In-Context Learning for Generalized Supplier Lead Time Estimation in Supply Chain Planning
Supplier lead time forecasting is a central input to material requirements planning, inventory optimization, and supply chain risk manageme…
Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes? We t…
The Story Shapes the Agent: Narrative Priors in LLM Behavior
Persona prompting is widely used to steer LLM agent behavior, yet the narrative framing of a task can matter more than the assigned persona…
For What Reason? Interpreting Models' Encoding of Causation and Antithesis
Discourse relations provide document structure, critical to language understanding and enabling language model performance and ethicality.…
Planning as Emergent Behavior in Reinforcement Learning with Relational Hidden States
Reinforcement learning is conventionally divided into model-based and model-free methods. In this taxonomy, model-based methods perform loo…
AutoIndex: Learning Representation Programs for Retrieval
We present AutoIndex, a framework for learning representation programs: executable transformations that map raw documents into the represen…
Intelligent Multi-UAV Navigation in ITNTNs: A Hierarchical LLM Approach
The deployment of high-speed Uncrewed Aerial Vehicles (UAVs) in 3D aerial highways necessitates robust coordination of physical flight kine…
Mitigating Matthew Effect: Multi-Hypergraph Boosted Multi-Interest Self-Supervised Learning for Conversational Recommendation
The Matthew effect is a big challenge in Recommender Systems (RSs), where popular items tend to receive increasing attention, while less po…
LatentMT: Machine Translation with Latent Reasoning
Latent-reasoning looped language models (LoopLMs) offer a different scaling path for machine translation (MT): instead of increasing parame…
Temporal-Causal Unity as an Operational Framework for Collective Dynamics: Causal-Progress Clocks, Synchronization, and Polarization
This paper develops temporal-causal unity (TCU), a framework connecting a process-philosophical thesis -- time is the ordered unfolding of…
CPInj: Uncovering Prompt Injection Risks in Textual Collaborative Prompt Optimization
Textual Collaborative Prompt Optimization (TCPO) extends Textgrad (Yuksekgonul et al., 2025) to a decentralized setting by allowing multipl…
Norm or Direction? Decoding Vision Mambas for High-Resolution Vision
Vision Mamba models replace quadratic self-attention with linear complexity selective state space models (SSMs), emerging as efficient visu…
Deep Learning Estimation of Sex, Age, Height, and Weight from CT-derived Digitally Reconstructed Radiographs
Purpose: To develop and validate a deep learning ensemble for estimating adult sex, age, height, and weight from coronal digitally reconstr…
Broken Gates: Re-evaluating Web Bot Defenses in the Age of LLM Agents
LLM-based browser agents are rapidly changing the threat landscape for web security. Unlike traditional automation frameworks that execute…
Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models
A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP wi…
Decoupled Pipeline with Proposal Reranking and Score Fusion for Positive-Unlabeled Marine Species Detection
The FathomNetCLEF 2026 competition combines underwater object detection and fine-grained marine species classification under a positive-unl…
What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio
Caption Studio is a transparency-first speech and audio intelligence platform that transforms spoken audio and video into structured, searc…
Strategy-Following Multi-Agent Deep Reinforcement Learning Considering Control Strategies Provided to Other Agents
This study proposes a learning method for multi-agent systems that allows agents to be controlled through human manager instructions after…
Find Before You Fine-Tune: A Diagnostic Study of Small LLMs for Cybersecurity QA
Large Language Models (LLMs) are increasingly fine-tuned for critical-domain Question-Answering (QA), yet choosing which small model to ada…
ConceptCF: Concept-based Counterfactuals for the Explainability of Time Series
This paper proposes ConceptCF, a method for counterfactual generation that operates on human-interpretable concepts. In high-stakes domains…
Bounding Boxes to Improve Small Language Model Performance on Vision-Based Grading Tasks
The deployment of Small Language Models (SLMs) in educational settings offers significant advantages in terms of privacy, cost, and scalabi…
AgentTrails: Towards Trust and Reuse for Agentic Tasks
LLM-powered agents increasingly tackle complex tasks by invoking tools, querying databases, executing code, and manipulating intermediate a…
AILQA: Evaluating AI-Driven Legal Question Answering Systems for the Indian Legal System
This comprehensive study introduces an advanced Artificial Intelligence for Indian Legal Question Answering (AILQA) system tailored to the…
Cross-Agent Campaign Attribution: Linking Asynchronous Attacks Across LLM Agents
LLM-agent defenses are typically evaluated one session at a time. In deployment, however, attacks can be distributed across independent age…
From Trajectories to Instructions: Language-Conditioned Meta-Reinforcement Learning
Model-Agnostic Meta-Learning (MAML) is a widely used framework for reinforcement learning (RL) that enables efficient transfer by learning…
ABOPD: Antibody CDR Design via On-Policy Distillation
Antibodies are essential therapeutic molecules, and their complementarity-determining regions (CDRs) form the primary antigen-recognition i…
Data Leakage Prevention in Agentic Applications via Preemptive Hardening
Agentic systems integrate LLM driven planning with interfaces to external tools, making data leakage and tool misuse feasible via instructi…
OPD-IAD: From Language Judgment to Industrial Anomaly Detection via On-Policy Self-Distillation
Large vision-language models (LVLMs) have recently shown strong potential for industrial anomaly detection (IAD) by providing image-level a…
Regime-Aware Physics-Guided Early Warning of Lithium-Ion Battery Thermal Runaway Using Thermo-Mechanical Signals
Thermal runaway in lithium-ion batteries poses a major safety risk to electric vehicles and energy storage systems. Current early-warning m…
RAMP: Recognition parametrisation by Amortised Message Passing
A central aim of unsupervised learning is to uncover latent factors that explain dependencies among observations. Probabilistic models typi…
Public perceptions of AI-driven decision-making in healthcare: A structural equation modeling approach
Artificial intelligence (AI) is increasingly integrated into healthcare to support diagnostics, decision-making, and administrative process…
Circuit Claims Depend on What Is Extracted and How It Is Compared
Circuit extraction identifies a small set of model components whose presence preserves a target behavior under ablation, and the resulting…
Functional Equivalence and Geometric Diversity in Neural Network Approximations: An Empirical Characterization
The Universal Approximation Theorem states that a neural network with a single hidden layer is sufficient to approximate any continuous uni…
Dual Adversarial Fine-tuning for Enhancing Robustness of Large Vision Language Model
While Large Vision-Language Models (LVLMs), represented by LLaVA and GPT-4V, have demonstrated remarkable capabilities, their visual inputs…
SFGA: A Statistics-First Gating Architecture with Adjudicative Escalation for Trustworthy SFT Data Procurement
Procuring supervised fine-tuning (SFT) data forces a buyer to decide, before any downstream training, whether a candidate corpus is worth a…
Variational meta-learning inference for low dimensional neural system identification
Deep learning has proven highly effective for nonlinear system identification, but heavily parameterized neural networks are prone to overf…
Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts
Agent Skills have become persistent behavioral artifacts across independent AI agent systems. They combine natural-language task specificat…
Verifiable Self-Evolution for Open-Ended Dialogue Skills via Future-Feedback Prediction
Textual skills provide a lightweight way to improve frozen language-model agents, but their self-evolution normally requires a stable valid…
AutoJourn: Multi-Perspective Summarisation, Bias Detection and Bias Neutralisation for LLM-Generated News in Automated Journalism
We present AutoJourn, a demonstration system for multi-perspective news generation and bias-aware evaluation using large language models (L…
SWITi: Quantifying and Reducing Tiling Artifacts with Sliding Window Inner Tiling
SWITi is a test-time method for reducing artifacts in tiled predictions, particularly for neural networks that learn posterior distribution…
MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents
Multi-turn medical consultation agents must decide what to ask, adapt to patient responses, and determine when the collected evidence is su…
Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges
Multimodal humor in memes, cartoons, and comics remains difficult for AI systems because intended meaning depends on non-literal mechanisms…
Biological Amnesia in ICU Time-Series Prediction: A Drift-Adaptive Two-Stream Architecture with Temporal Retrieval
Background: Clinical decision support systems degrade silently as treatment protocols evolve, yet standard adaptation methods treat models…
CoGoal3D: Collaborative 3D Object Detection with 3D-Aware Fusion and Refinement
V2X collaborative object detection features overcoming the limitations of single-vehicle systems by aggregating environmental features from…
FilmWorld: Agentic Novel-to-Film Generation through Dynamic Cinematic World Modeling
Translating novels into films poses a grand challenge for generative artificial intelligence, requiring conversion of abstract literary pro…
Spectral Higher-Order Neural Networks Have Sharp Expressivity Bounds
Neural hypergraphs are a natural generalization of neural networks, the reference models in modern machine learning. Yet, their deployment…
Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training: on a 6.78B-parameter MoE languag…
Deep learning-based prediction of time-resolved adhesive forces in viscoelastic Hertzian contacts
Fast prediction of the response of adhesive soft viscoelastic contacts represents a current challenge in soft robotics and for gripping and…
Now You See the Hate: Adaptive View Retrieval for Hidden Hateful Illusions
Hateful optical illusions expose a serious gap in current multimodal safety systems. On original-view hateful illusions, previous work show…
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-sca…
From Operations to Elderly Care Outcomes: A Thematic Review of Industrial Engineering and Decision-Support Approaches
The rapid growth of the global aging population presents severe challenges to healthcare systems, necessitating efficient, equitable, and p…
DAIS: Dependency-Aware Intermediate QA Supervision for Complex Reasoning
Chain-of-thought (CoT) supervision exposes intermediate rationales, but flat rationale targets usually optimize a single reasoning sequence…
SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation
Large language models (LLMs) excel at general-purpose code generation, yet how well they handle scientific code remains an open question. E…
Code Division Modulation Layers Against Forgetting and Inference in Continual Gait Identification
Continual learning (CL) has been recently employed in biometric identification systems thanks to its ability to integrate new knowledge wit…
Parallel Noising in Neural Markov Logic Networks
Neural Markov Logic Networks (NMLNs) are a flexible neurosymbolic relational model. Previous work has shown that, although NMLNs achieve st…
MIRAGE: Multi-scale Lesion-Informed Representation with Auxiliary Guidance for MRI Contrast Enhancement
Inferring contrast enhancement from one pre-contrast breast MRI slice is underdetermined: post-contrast appearance contains physiological i…
Incomplete Observations Boost Evolutionary Performance in Ocean Modeling
Data-driven methods have revolutionized ocean modeling, yet current approaches rely heavily on complete reanalysis datasets, imposing compu…
Breaking the Homogeneity Assumption: Specialized Multi-Generator Adversarial Learning for Rare Failure Detection in Predictive Maintenance
Supervised learning models in the predictive maintenance field are regularly trained on highly imbalanced industrial datasets: machine fail…
Reasoning Before Translation: Enhancing Legal Machine Translation with Structured Reasoning
Neural machine translation (NMT) in the legal domain is a linguistically and conceptually demanding task, primarily due to the complexity o…
Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a…
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-…
Free energy landscape of Dense Associative Memory
Using large deviations theory, we solve and obtain a general expression for the free energy functional for a broad class of associative mem…
MIRA-Ev:A Benchmark for Granular Evidence Detection and Relational Reasoning in Clinical Exams
Clinical NLP evaluation remains dominated by multiple-choice question answering (MCQA), which scores only final-answer accuracy and cannot…
Assessment in Team Problem-Solving Exercises in Computing Education
This full paper in the research-to-practice track presents methods for assessing student teams in tabletop exercises (TTXs). TTXs enable le…
Computing on the Fly: Navigating a Vision for the Future of Drone Computing
The report envisions a decade in which drones move goods, medical supplies, and information at a scale comparable to national infrastructur…
Beyond Score Prediction: LLM-Based Essay Scoring and Feedback Generation via Reinforcement Learning with Rubric Rewards
Large language models (LLMs) have been widely applied to automated essay scoring (AES) and automated feedback generation (AFG). However, ex…
The Price of Reasoning: Cost-Quality Tradeoffs in Reinforcement Learning for Neural Machine Translation
Reinforcement learning with verifiable rewards (RLVR) has been established as a viable paradigm for the post-training of Large Language Mod…
Inference-Time Steering for Cross-Lingual Factual Consistency in LLMs
Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain dispropo…
Prompt Design at Scale: How Format, Instruction Count, and Context Length Shape Instruction Adherence and Hallucination in Large Language Models
Practitioners make three prompt-design decisions with almost no controlled evidence behind them: how to format instructions and context (ma…
Benchmarking Generalization in Financial Statement Fraud Detection: robust evaluation and novel tasks
Financial statement fraud detection (FSFD) is crucial for market integrity but faces challenges from increasingly sophisticated schemes and…
PathAgentBench: Benchmarking Evidence-Seeking Vision-Language Models on Whole-Slide Pathology Image
Whole-slide image (WSI) diagnosis requires identifying diagnostically relevant regions, examining them across magnifications, and integrati…
Toward Auditable Fraud Detection: Combining Graph Features, Model Explanations, and Agentic Case Investigation
Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable. We study a layered pipeline o…
They'll Verify. They Just Won't Act. How Authority Framing and Laundered Code Turn a Trusted Agentic CI/CD Pipeline Into an Attack Surface
We study a five-agent CI/CD pipeline (triage -> developer -> security-scan -> review -> approve/deploy), built from five distinct productio…
GUIDED Network-Agnostic Feature Initialization for Spatial Transferability in GNN-based Models
The Traffic Assignment Problem is a fundamental but computationally expensive component of transportation planning. While Graph Neural Netw…
The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetica…
Riemannian Deep Learning:Modules, Networks, and Geometries
Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific…
From Distances to Trajectories: Real-Time Signed Distance Function Mapping and Distance-Accelerated Motion Planning for UAVs
Autonomous flight in cluttered environments requires a robot to build a geometric map of its surroundings and plan safe, dynamically feasib…
Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information
Reinforcement learning with verifiable rewards (RLVR) improves reasoning in large language models. Yet, typical RLVR approaches fail on dif…
ISO: An RLVR-Native Optimization Stack
Reinforcement learning with verifiable rewards (RLVR) is rapidly advancing the reasoning capabilities of language models, yet the optimizat…
Provable diffusion-based posterior sampling for linear inverse problems via DDIM
Diffusion-based methods have achieved remarkable empirical success in solving inverse problems. However, many existing posterior samplers e…
Appearance Pointers -- Multimodal Region Control of Diffusion Transformers
Controllable image generation remains challenging for creative professionals, who often require precise regional control over materials, ob…
Copy Less, Ground More: Overcoming Repetitive Copying in Long-Context Reasoning via Evidence-Aware Reinforcement Learning
Large language models that generate step-by-step reasoning traces have achieved strong performance on complex tasks, and extending them to…
FormGym: Doing Paperwork with Agents
Completing paperwork is a challenging and time-consuming problem. Form filling is especially challenging in the pure-image domain without a…
Learning, Reasoning, Refinement: A Framework for Kahneman's Dual-System Intelligence in GUI Agents
Graphical User Interface (GUI) agents have made significant progress in automating digital tasks through the utilization of computer vision…
Assistax: A Multi-Agent Hardware-Accelerated Reinforcement Learning Benchmark for Assistive Robotics
As embodied autonomous systems capable of assisting humans in daily activities remain a major goal for robotics, efficient and appropriate…
SENTINEL: A Multi-Level Formal Framework for Safety Evaluation of Foundation Model-based Embodied Agents
We present SENTINEL, a framework for formally evaluating the physical safety of foundation model (FM)-based embodied agents. SENTINEL is th…
Learning to Make Friends: Coaching LLM Agents toward Emergent Social Ties
Can large language model (LLM) agents reproduce the complex social dynamics that characterize human online behavior -- shaped by homophily,…
Dr. Zero: Self-Evolving Search Agents without Training Data
As high-quality data becomes increasingly difficult to obtain, self-evolution without curated training data has emerged as a promising para…
Fluid Reasoning Representations
Frontier large language models increasingly solve complex tasks involving abstract concepts through extended test-time thinking. Yet we lac…
LLM-Grounded Explainable AI for Supply Chain Risk Early Warning via Temporal Graph Attention Networks
Disruptions at critical logistics nodes pose severe risks to global supply chains, yet existing risk prediction systems typically prioritiz…
Animating Petascale Time-varying Data on Commodity Hardware with LLM-assisted Scripting
Scientists face significant visualization challenges as time-varying datasets grow in speed and volume, often requiring specialized infrast…
Participatory provenance as representational auditing for AI-mediated public consultation
AI-assisted consultation can speed large-scale public engagement, but concise summaries may reflect some submissions more closely than othe…
FinRAG-12B: A Production-Validated Recipe for Grounded Question Answering in Banking
Large language models (LLMs) are rapidly being adopted across various domains. However, their adoption in banking industry faces resistance…
Frontier LLM-based agents can overcome the ontology curation bottleneck for natural phenotypes
Linking free-text phenotype descriptions to ontology terms, typically referred to as phenotype annotation, is essential for the cross-study…
AgentJet: A Distributed Swarm Training Framework for Agentic Reinforcement Learning
Training reinforcement learning (RL) policies for large language model (LLM) agents requires optimizing multi-turn trajectories that intera…
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem…
The Theory of Mind Utility: Formal Specification of a Mentalizing Mechanism
Inferring others' beliefs requires more than reading surface signals; it requires tracking who told them what, in what order, and how credi…
Externalizing Research Synthesis and Validation in AI Scientists through a Research Harness
AI systems can increasingly automate scientific workflows, but the reasoning that links prior evidence, generated ideas, experiments and fi…
SAGA: Scene-Aware, Goal-Evolving Agents for Long-Horizon Strategy Game Planning
Grand-strategy games such as Civilization pose a distinctive long-horizon planning problem: an agent must divide one shared resource pool a…
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can imp…
Subliminal Clocks: Latent Time Modelling in Diffusion Language Models
Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive models. Unlike standard diffusion-based…
Context-Masked Truncated Reasoning Audits for Answer-Key Dependence in LLM Tutors
Large language model (LLM) tutors may have access to teacher notes, answer keys, rubrics, or retrieved solutions while producing student-fa…
Applying JEPA-Style Predictive Learning to JA4-Derived Network Fingerprints
I-JEPA and V-JEPA learn by matching latent predictions to target encoder outputs rather than regenerating the original input, and this has…
AdvNav: Behavior-Guided Black-Box Adversarial Attacks on Vision-Language Navigation
Despite progress in Embodied AI, Vision-and-Language Navigation systems remain vulnerable to adversarial visual disturbances. Most existing…
Evidence-Grounded AI for Musculoskeletal Care
Musculoskeletal diseases are among the leading causes of disability and drive the greatest global need for rehabilitation. Because recovery…
FormalAnalyticGeo: A Neural-Symbolic Based Framework for Multimodal Analytic Geometry Problem Generation
Math reasoning has achieved significant progress with the rapid advancement of Multimodal Large Language Models (MLLMs), however analytic g…
Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs
Aligned language models routinely misreport under non-evidential pressure: they cave to a confident user, yet fail to revise when genuine e…
RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning
Multi-step retrosynthesis planning seeks to decompose a target molecule into commercially available building blocks through a sequence of f…
Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet current evaluations remain largely chart-ce…
Bayesian inference of composition-dependent phase diagrams
Phase diagrams serve as a highly informative tool for materials design, encapsulating information about the phases that a material can mani…
Saving the legacy of Hero Ibash: Evaluating Four Language Models for Aminoacian
This study assesses four cutting-edge language models in the underexplored Aminoacian language. Through evaluation, it scrutinizes their ad…
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have bec…
Soft-TransFormers for Continual Learning
Inspired by the Well-initialized Lottery Ticket Hypothesis (WLTH), we introduce Soft-TransFormers (Soft-TF), a continual learning framework…
A Self-Supervised Framework for Space Object Behaviour Characterisation
Foundation Models, which leverage large neural networks pre-trained on unlabelled data before fine-tuning for specific tasks, are increasin…
Parameter-Efficient Continual Fine-Tuning: A Survey
The emergence of large pre-trained networks has revolutionized the AI field, unlocking new possibilities and achieving unprecedented perfor…
GSPRec: On Improving Item Representations in Graph Signal Processing for Collaborative Filtering
Graph-based collaborative filtering methods act as low-pass filters in the spectral domain and discard the intermediate-frequency component…
Chi-Square Wavelet Graph Neural Networks for Heterogeneous Graph Anomaly Detection
Graph Anomaly Detection (GAD) in heterogeneous networks presents unique challenges due to node and edge heterogeneity. Existing Graph Neura…
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data…
FedS2R: One-Shot Federated Domain Generalization for Synthetic-to-Real Semantic Segmentation in Autonomous Driving
Federated domain generalization has shown promising progress in image classification by enabling collaborative training across multiple cli…
RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation
Large language models (LLMs) demonstrate remarkable capabilities in reasoning and code generation, enabling robotic manipulation to be init…
Robust Belief-State Policy Learning for Quantum Network Routing Under Decoherence and Time-Varying Conditions
Quantum network routing requires online decisions under probabilistic entanglement generation, finite quantum memories, decoherence, imperf…
Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures
Despite their capabilities, Large Language Models (LLMs) remain opaque with limited understanding of their internal representations. Curren…
Breaking the MoE LLM Trilemma: Dynamic Expert Clustering with Structured Compression
Mixture-of-Experts (MoE) Large Language Models (LLMs) face a trilemma of load imbalance, parameter redundancy, and communication overhead.…
Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing
Instruction-based image editing offers a powerful and intuitive way to manipulate images through natural language. Yet, relying solely on t…
Beyond-Diagonal RIS Under Non-Idealities: Learning-Based Architecture Discovery and Optimization
Beyond-diagonal reconfigurable intelligent surface (BD-RIS) has recently been introduced to enable advanced control over electromagnetic wa…
QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent f…
Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives
How complex collective behavior emerges from individual interactions is a fundamental scientific question, but experimental cost and diffic…
T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs
Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training. Rece…
ImplicitRDP: An End-to-End Visual-Force Diffusion Policy with Structural Slow-Fast Learning
Human-level contact-rich manipulation relies on the distinct roles of two key modalities: vision provides spatially rich but temporally slo…
Memo2496: Expert-Annotated Dataset and Dual-view Adaptive Framework for Music Emotion Recognition
Music Emotion Recognition (MER) is constrained by limited expert annotations and the need to establish robustness across heterogeneous corp…
PRISP: Privacy-Safe Few-Shot Personalization via Lightweight Adaptation
Large language model (LLM) personalization aims to adapt general-purpose models to individual users. Most existing methods, however, are de…
SKETCH: Semantic Key-Point Conditioning for Long-Horizon Vessel Trajectory Prediction
Accurate long-horizon vessel trajectory prediction remains challenging due to compounded uncertainty from complex navigation behaviors and…
Toward Learning POMDPs Beyond Full-Rank Actions and State Observability
We are interested in enabling autonomous agents to learn and reason about systems with hidden states, such as locking mechanisms. We cast t…
Training and Simulation of Quadrupedal Robot in Adaptive Stair Climbing and Descending for Indoor Firefighting: An End-to-End Reinforcement Learning Approach
Quadruped robots are used for primary searches during the early stages of indoor fires. A typical primary search involves quickly and thoro…
LinguistAgent Technical Report: A Reflective Multi-Model Platform for Automated Linguistic Annotation
Data annotation remains a significant bottleneck in the field of humanities and social sciences, particularly for complex linguistic tasks…
CompilerKV: Risk-Adaptive KV Compression via Offline Experience Compilation
Prefill-only KV compression freezes a token subset at the end of prefill and decodes from it without further eviction. The retention decisi…
Fly0: Persistent Metric Anchoring for Zero-Shot Aerial Vision-Language Navigation
Current Visual-Language Navigation (VLN) methodologies face a trade-off between semantic understanding and control precision. While Multimo…
When Visual Evidence is Ambiguous: Pareidolia as a Diagnostic Probe for Vision Models
When visual evidence is ambiguous, vision models must decide how to interpret face-like patterns. Face pareidolia, the perception of faces…
Give Them an Inch and They Will Take a Mile:Understanding and Measuring Caller Identity Confusion in MCP-Based AI Systems
The Model Context Protocol (MCP) is an open and standardized interface that enables large language models (LLMs) to interact with external…
SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution
Real-world software must continuously evolve to meet ever-changing and open-ended requirements. AI agents, increasingly deployed as long-ru…
TransDex: Pre-training Visuo-Tactile Policy with Point Cloud Reconstruction for Dexterous Manipulation of Transparent Objects
Dexterous manipulation enables complex tasks but suffers from self-occlusion, severe depth noise, and depth information loss when manipulat…
PlotTwist: A Creative Plot Generation Framework with Small Language Models
Creative plot generation presents a fundamental challenge for language models: transforming a concise premise into a coherent narrative tha…
When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines
Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents…
Doctorina MedBench-ICD10: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI
We present Doctorina MedBench, a comprehensive evaluation framework for agent-based medical AI based on the simulation of realistic physici…
M-RAG: Semantic Key-Value Indexing for Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) turns external documents into evidence for large language models. In practice, this is also a data acc…
FVRuleLearner: Operator-Level Reasoning Tree (Op-Tree)-Based Rules Learning for Formal Verification
The remarkable reasoning and code generation capabilities of large language models (LLMs) have recently motivated increasing interest in au…
Robust Reasoning Benchmark
While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on…
CPGRec+: A Balance-oriented Framework for Personalized Video Game Recommendations
The rapid expansion of gaming industry requires advanced recommender systems tailored to its dynamic landscape. Existing Graph Neural Netwo…
Why Do Vision Language Models Struggle To Recognize Human Emotions?
Understanding emotions is a fundamental ability for intelligent systems to be able to interact with humans. Vision-language models (VLMs) h…
AnchorRefine: Synergy-Manipulation Based on Trajectory Anchor and Residual Refinement for Vision-Language-Action Models
Precision-critical manipulation requires both global trajectory organization and local execution correction, yet most vision-language-actio…
Agentic AI-assisted coding offers a unique opportunity to instill epistemic grounding during software development
The capabilities of AI-assisted coding are progressing at breakneck speed. Chat-based vibe coding has evolved into fully fledged AI-assiste…
Large Language Models Explore by Latent Distilling
Generating diverse responses is crucial for test-time scaling of large language models (LLMs), yet standard stochastic sampling mostly yiel…
Lifting Embodied World Models for Planning and Control
World models of embodied agents predict future observations conditioned on an action taken by the agent. For complex embodiments, action sp…
Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing
LLM-based agents are increasingly applied to the "last mile" of Electronic Design Automation (EDA): repairing residual sign-off Design Rule…
GQLA: Group-Query Latent Attention for Hardware-Adaptive Large Language Model Decoding
Multi-head Latent Attention (MLA), the attention used in DeepSeek-V2/V3, jointly compresses keys and values into a low-rank latent and matc…
Global Automation Atlas
Automation can displace or complement labour, but this need not be constant across economies. Existing exposure measures typically assign f…
Tunable MAGMAX: Preference-Aware Model Merging for Continual Learning
Continual learning (CL) aims to train models sequentially on multiple tasks while mitigating catastrophic forgetting of previously learned…
The Cognitive Kardashev Scale: Quantifying the Material Envelope of Civilisational Computation
How much thinking can a civilisation do? Kardashev ranked civilisations by the energy they command. This paper borrows his ladder and asks…
CWind: A Cross-site Router for Large Language Model Inference Serving at Renewable Energy Farms
AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up. Grid expansion comes with h…
ViMax: Agentic Video Generation
Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Ex…
"I understand your perspective": LLM Persuasion through the Lens of Communicative Action Theory
Large Language Models (LLMs) can generate high-quality arguments, yet their ability to engage in nuanced and persuasive communicative actio…
Reframing AI Loss of Control: What Control Is, How to Have It, How to Lose It
At present, loss of control risks have gained much prominence in public discussion, particularly in relation to AI, with extensive discours…
Phantoms and Disclosures: A Statistical Framework for Auditing Privacy in Synthetic Data
The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alterna…
A Red Teaming Framework for Large Language Models: A Case Study on Faithfulness Evaluation
Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in hig…
Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality
Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agent harness: a mid…
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound…
BioSecBench-Refusal: A paired metric for performance and alignment in agentic biosecurity risk assessment
As AI agents are incorporated into life science workflows, the capabilities that speed discovery might also enable misuse. We present BioSe…
Prompt Robustness Is Task-Dependent: Comparing Objective and Belief-Style Questions in LLM Evaluation
Survey-style evaluations of large language models often treat a prompted response as a measure of a model's values or beliefs. This assumpt…
A Transdiagnostic Space of Disorder Like Phenotypes in Reinforcement Learning Agents
Modelling psychological disorders in artificial agents offers a testbed for computational psychiatry and a lens on affective-control failur…
LieBN: Batch Normalization over Lie Groups
Manifold-valued measurements are prevalent in various machine learning tasks. Recent advances have extended Deep Neural Networks (DNNs) to…
An LLM-powered Agentic Recommendation System for Connected TV Content Discovery
Recommendation systems, from traditional multi-stage to recent unified generative architectures, face challenges in incorporating diverse c…
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process…
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Ans…
Falsifiable Release Gates for Self-Improving Systems: Standing Invariants at Scale
Safety claims for self-improving agent runtimes are almost always self-graded: a policy file, a guardrail, a promise in a README. We descri…
The Caf\'e in Amsterdam: When the Incumbent Becomes the Oracle
A field can reformulate its computations freely exactly where its demand is stated independently of any incumbent implementation, and finds…
Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach
Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, ac…
MicrosoftとMistralが戦略的提携を拡大 欧州でのAIインフラ拡張とモデル展開を加速
MicrosoftとMistralは戦略的提携を拡大すると発表した。Mistralの最新モデルをMicrosoftの各プラットフォームへ展開するほか、欧州でのGPUインフラ拡張に向けて大規模な投資を行う。クラウドから完全オフラインまで多様な環境に対応し、規制業界での高度なAI導…
AIを悪用した攻撃、どう対抗する? EDR導入の“次”にやるべきこと
ランサムウェアやサプライチェーン攻撃が中堅・中小企業にも及ぶ中、国内のEDR市場は前年度比13.3%増と2桁成長を続ける。中堅・中小企業にも普及する一方で、高度なAIツールを悪用したサイバー攻撃対策にはEDR導入だけでは十分ではない。
ジャック・ドーシー氏率いるBlock、AI協働プラットフォーム「Buzz」公開 SlackやGitHub依存からの脱却目指し
ジャック・ドーシー氏率いるBlockは、人間とAIエージェントが同じワークスペースで協働するオープンソースプラットフォーム「Buzz」を公開した。分散型プロトコル「Nostr」上に構築され、任意のLLMやエージェントを組み込める。各参加者が独立した暗号鍵ペアを持つことで識別と権…
PTC、「Onshape」にAI機能を先行提供 早期アクセスプログラムを開始
米PTCは、クラウドネイティブCAD/PDMプラットフォーム「Onshape」の早期アクセスプログラム「Onshape Labs」を発表した。AIを活用した設計支援やレンダリングなどの新機能を一般提供に先駆けて試用できる。
Meta is testing an AI bedtime story app for people with no imagination
At last, a tech company has found a way to outsource humanity's oldest pastime: using our imaginations.
富士通・NVIDIAとロボット大手3社が協業へ フィジカルAI社会実装の具体策は?
フィジカルAIの社会実装は、一企業だけでは手に余る――。この課題に、富士通は競合するロボット大手3社、そしてNVIDIAと組んで挑む。協業で描く具体策とは。
OpenAI会長、米国フロンティアモデルの優位性を強調 「オープンモデルは必ずしも安くない」
OpenAIのブレット・テイラー会長がCNBCのインタビューで、中国発オープンウェイトモデルの台頭に「必ずしも実行コストが安いわけではない」と反論。トークン効率と推論効率で米フロンティアモデルの優位を強調した。
Hugging Face侵害のAIエージェントはOpenAIのモデル──社内のサイバー能力評価中に「GPT-5.6 Sol」などが暴走し本番DBに侵入
OpenAIは、Hugging Faceで発生したサイバーインシデントの原因が自社のAIモデルだったと発表した。社内評価中に安全機能を抑制した「GPT-5.6 Sol」などが隔離環境を突破し、ゼロデイ脆弱性を悪用して外部に侵入したという。OpenAIはインフラ管理の厳格化や防御…
Google、「Gemini 3.6 Flash」など3モデルを発表 出力トークンを削減しつつ値下げ、「Gemini 4」も予告
Googleは、Geminiの「Flash」シリーズに「3.6 Flash」「3.5 Flash-Lite」「3.5 Flash Cyber」の3モデルを追加した。効率性と低遅延を追求し、AIエージェント構築に適した性能を備える。3.6 Flashは出力価格が引き下げられた。ま…
ブレストで膨らむ“隠れ人件費”を削れ 矢野経済がClaudeで挑む「30分で100アイデア」創出の威力
矢野経済研究所が、独自の一次情報と高度AI「Claude」を融合させた新規事業アイデア創出支援サービス「AIDEL」を発表した。ブレインストーミングによる役員の時間拘束や「隠れ人件費」の膨張という企業の課題に対し、30分で100の具体案と評価スコアを自動生成。一般的な生成AIの…
GoogleがAIアプリ「Dreambeans」を発表 「画面を延々とスクロール」の脱却で何を目指すのか
Googleは、AIがユーザー一人一人に向けた日々のストーリーを自動で生成する実験的アプリ「Dreambeans」を発表した。「際限のないスクロール」に代わり、Googleは何を目指すのか。
矢崎総業がイノベーション拠点を公開、労働集約型モノづくりのスマート化に向け
矢崎総業は、新たに開設したイノベーション施設「Innovation Hub - REN(錬)」(IH-REN)を報道陣に公開した。IH-RENでは、AI/ロボティクスを活用した次世代のモノづくりに向けて、自働化の検証や産学連携による研究開発を推進し、新たな価値の創出を目指す。
OpenAI says Hugging Face was breached by its pre-release models
OpenAI has come forward to claim responsibility for the Hugging Face breach, saying it was the result of internal testing gone awry.
AIトークン消費「24倍」の衝撃 本番運用に向けて絶対に“やってはいけない”コストの捉え方
AIの試験導入から本番運用への移行が進む中、多くの企業がコストと統制の壁に直面している。将来的なトークン消費の急増を見据え、組織が今見直すべき視点とは何か。実運用を持続させるための「3つの条件」を解説します。
Jack Dorsey is taking on Slack with Buzz, a group chat platform for teams and their AI agents
Buzz is a group chat platform for the workplace that puts humans and their AI agents in the same conversation.
AI and the rise of the universal entertainment app
Over the past decade, streaming platforms competed by dominating individual formats like music, video, podcasts, or audiobooks. Now, as AI…
Data centers expected to use 4x more electricity by 2035
New data centers built through 2033 could consume as much electricity as India uses today.
Google releases three new Gemini models — but no 3.5 Pro
Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber, but the continued absence of Gemini 3.5 Pro raises fresh questions about…
Introducing the ChatGPT for small business program
OpenAI launches the ChatGPT for Small Businesses program, helping entrepreneurs build AI skills, automate work, and grow with ChatGPT Work.
US threatens sanctions against Chinese AI models over IP theft
Treasury Secretary Scott Bessent said the U.S. could sanction Chinese open AI models over alleged IP theft, expanding the Trump administrat…
Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
We’re introducing new Gemini models, including Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber.
2026-07-21(533件)
Music streamer Deezer says more than 50% of daily uploads are AI-generated
Deezer said more than 90,000 AI-generated tracks were uploaded daily on the platform in June.
Gritt exits stealth with $32 million for robots to build solar plants — then, everything else
Gritt is coming out of stealth with $34 million and plans to automate the hardest tasks on construction sites.
OpenAI and Hugging Face partner to address security incident during model evaluation
OpenAI and Hugging Face share early findings from a security incident during AI model evaluation, highlighting advanced cyber capabilities…
Anthropic、著作権訴訟で史上最大「2400億円」和解金支払いへ 学習利用は「フェアユース」認定
「Claude」の学習を巡り作家グループが米Anthropicを訴えた集団訴訟で、米連邦判事が15億ドル(約2400億円)の和解を最終承認した。米国の著作権訴訟では史上最大の和解額となる。
ドラクエと「Gemini」がコラボ 画像生成の“特別なテンプレ”提供、リアルイベント開催へ
米Googleの日本法人は、AIサービス「Gemini」と「ドラゴンクエスト」のコラボキャンペーンを始めると発表した。Geminiの画像生成機能を活用するイベント「ジェミニクエスト」を開催するほか、アプリでは同イベントに連動したテンプレートも展開する。
「取りあえずAI導入」の末路 現場で深まる情報漏えい不安 IPAの意識調査で明らかに
IPAは「AIの動作・分析・利用等の説明に関する意識調査」を公開した。AI利用経験3年未満の回答者が多く、利用知識の不足や情報漏えいに対する不安の現状、リスク認識の傾向などが示されている。
Rater State Bias in RLHF Preference Data: An Audit Framework
We identify a structured confound in Reinforcement Learning from Human Feedback (RLHF). Pairwise preference labels are intended to reflect…
Design and Validation of a Lightweight 1D CNN for Affective Touch Classification in Soft Plush Companions
Soft, sensorized companions offer a physically safe and emotionally intuitive interface for socially assistive technologies, yet their defo…
Some Large Language Models Exhibit Consistent Risk Attitudes
As artificial intelligence systems are deployed in open-ended, high-stakes settings, a critical dimension remains unmeasured: how perceived…
A Survey on GNN-based Link Prediction: Techniques, Applications, and Challenges
Graph Neural Networks (GNNs) have emerged as the leading paradigm for link prediction, enabling the inference of missing connections and th…
PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection
Multi-agent LLM systems increasingly rely on a Planner to decompose goals into sub-task sequences that downstream Executor and Critic agent…
Deterministic Replay for AI Agent Systems
AI agent systems that couple large language models (LLMs) with external tools and APIs are inherently non-deterministic: LLM sampling varia…
Generative Ontology Induction: Domain-Agnostic Schema Discovery from Document Corpora Using Large Language Models
Ontology engineering remains a critical bottleneck in knowledge-intensive AI systems. Existing automated approaches either depend on predef…
Democratizing AI with Small Language Models: Structured Benchmarking and Parameter-Efficient Fine-Tuning for Local Deployment
AI democratization is not primarily a question of matching frontier-scale generality; it is a question of whether capable models can be sel…
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments…
It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches
Reinforcement learning with verifiable rewards has emerged as a standard approach for enhancing reasoning in large language models, which t…
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization
This paper introduces PPO-HSC (Proximal Policy Optimization with High-order Sampling Coverage), an exploratory reinforcement learning frame…
JUMP: Single-Pass Membership Inference on Fine-Tuned Diffusion Language Models
Membership inference attacks (MIAs) test whether a candidate example appeared in a model's training data. We study MIAs for fine-tuned disc…
ColGraphRAG: Late-Interaction Evidence Retrieval for Multimodal GraphRAG
Graph-grounded multimodal question answering organizes text, tables, and images in a structured evidence graph, yet end-to-end accuracy dep…
Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning
Context reranking and pruning have become essential for improving the efficiency of modern Retrieval-Augmented Generation (RAG) systems, ye…
A Survey on the Verification of Reinforcement Learning Policies
Reinforcement learning (RL) is increasingly applied in complex, safety-critical domains, yet the lack of rigorous behavioral guarantees for…
Accurate and Efficient Long-Term Memory for LLM Agents
LLM agents augmented with persistent memory can recall past interactions, but existing systems suffer from two limitations: flat, unstructu…
Symbolic Augmentation Closes a Canonical-Equivalence Blind Spot in Neural Fact-Checkers
Large language models hallucinate numbers and units when summarizing scientific text, a failure mode that can silently invert a scientific…
SelKV: Selective KV Cache Merging with Per-Token Merge-or-Drop and Attention Compensation
Large Language Models (LLMs) generate text autoregressively, relying on a key-value (KV) cache whose memory footprint grows linearly with c…
RAIL Guard: Closing the Evaluation-to-Remediation Gap in Responsible AI for LLM Agents
Existing guardrail systems for large language model agents operate as binary classifiers that block unsafe content, leaving organizations t…
Generalist AI Control: Towards Multi-purpose Adaptive Algorithms
Traditional controllers are designed for specific systems and do not transfer across different system orders and dynamics. We present a Gen…
LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models
Diffusion-based Large Language Models(DLLMs) enable parallel generation via Semi-Autoregressive (SAR) decoding in text generation. However,…
Interactive Task Alignment as a POMDP
Current benchmarks for language models primarily evaluate execution on fully specified tasks. However, real user tasks are often ambiguous.…
When to Plan: Learning to Select Between Reactive Control and Deliberative Planning
It has long been recognized that humans have the ability to switch between fast, reactive decision-making and slower, deliberative planning…
Berkeley and Heiserman as an Unexhausted Architecture for Embodied Machine Intelligence
Edmund C. Berkeley is usually remembered as a writer who helped connect symbolic logic to computing machinery. That description is correct,…
SEER: Supervised Learning to Control Energetic Reasoning
One of the main strengths of Constraint Programming is the ability to reduce the search space via propagation. However, propagation is a do…
Nonuniformity Principle in Human-AI Coworking
As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and involvement remain essential f…
From Modalities to Propositions: A Language-Centric Framework for Multimodal Intelligence
We propose a language representation for multimodal data in which any observation, whether image, video, or text, is expressed as a bag of…
Exact Network Surgery: Functional Invariance and Gradient Plasticity in Reactive Computational Graphs
Function-preserving network growth techniques such as Net2Net and progressive stacking expand a model's capacity without destroying its lea…
FST.ai 2.5: Explainable and Uncertainty-Aware AI for Olympic and Para-Taekwondo Decision Support, Athlete Digital Twins, and Federation-Scale Analytics
The rapid digitalisation of elite sport has created new opportunities for integrating artificial intelligence (AI), performance analytics,…
Just A Rather Very Intelligent Spoken Agent
Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin. In most workflows, us…
A Research Prototype for Closed-Loop Generative Design of Customized Foot Orthoses via Semantic-Physics Alignment
Translating unstructured clinical prescriptions into patient-specific foot orthoses (FOs) is hindered by a semantic-physical misalignment:…
TopoTuner: Topological Finetuning of Large Language Models
Full fine-tuning remains a strong way to adapt pretrained LLMs, but it updates all weights and can be expensive. LoRA reduces the number of…
Diversity-Oriented Fine-Tuning for Uncertainty-Based Hallucination Detection
Existing hallucination detection methods are typically conducted at the inference stage, without making any modifications to the model itse…
DS@GT ARC at eRisk 2026: Hybrid Multi-Agent LLM System with Structured Algorithmic Guidance for Conversational Depression Screening
We describe DS@GT's submission to the eRisk 2026 Task 1 challenge on conversational depression screening, in which systems interview LLM pe…
Tractable Query Answering under Epistemic Confidentiality Policies in DL Ontologies (extended version)
We study Controlled Query Evaluation (CQE), a declarative approach to confidentiality-preserving data access, in the context of Description…
RECON: Benchmarking Agent Memory for Compositional Reasoning over Long Contexts
Large language models and LLM-based agents are widely used as personal chat assistants, enterprise copilots, and autonomous workflow agents…
Constraint-Anchored Reasoning Traces
Autoregressive multimodal large language models (MLLMs) suffer from error snowballing: a single incorrect inference early in a chainof-thou…
Supporting Autonomous Process Execution within a Multi-Perspective Constraint Frame via Numeric Planning
AI-Augmented Business Process Management Systems (ABPMS) enhance traditional BPMS by leveraging advanced AI techniques to define, execute,…
RELIC: Revealed Principles for Learning Interpretable Composable Skills in Multi-Agent Planning
Multi-agent planning becomes substantially harder when agents must improve specialized decision-making skills while keeping their internal…
FUSAR-R1: A Large-Scale Reasoning Model for Intelligent Interpretation of SAR Images
In recent years, large-scale vision-language models have been driving a paradigm shift in intelligent remote sensing image interpretation.…
From Overload to Insights: How AI Agents Can Support Scientists in Analyzing Complex Data
Scientists at European XFEL conduct experiments that generate very large and complex datasets. The subsequent data analysis is challenging…
AgentBrew: Lifelong Knowledge Brewing from Strong Teachers to Weak LLM Agents
Deploying LLM agents typically requires a compact test-time student, even if a stronger teacher is available during training. We study know…
Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
Large Language Models (LLMs) often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-s…
Environment-free Synthetic Data Generation for API-Calling Agents
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data…
Lomekwi: Resource-Bounded Tool Discovery in LLM Agents
Existing tool-use benchmarks report a single success rate for complex, multistep tasks. Inspired by ideas from cognitive science, we distin…
Training Continuous Chain of Thought Models: A Tale of Two Regimes
Continuous Chain-of-Thought methods replace verbose reasoning traces with a short sequence of dense latent representations. Earlier continu…
Expected Free Energy as Belief-Dependent Utility for rho-POMDPs
An agent acting under partial observability must decide when to gather information and which observations are worth their cost. Standard PO…
PriorProof: A Point-in-Time Measure of Technique Novelty for Formal Proofs
Mathematicians distinguish proofs that explain, simplify, or introduce a nonstandard route, but these judgments are difficult to operationa…
Reward-Driven LLM Agent Workflows: Synthesizing POMDP Routing and Self-Correction for Autonomous Decision-Making
This paper addresses key technical challenges in current large language model (LLM) agent applications, including long-horizon planning, sp…
When LLMs Over-Answer: Measuring and Mitigating Quality Issues in LLM-Based Hardware Description Language Question Answering
The rapid advancement of large language models (LLMs) has led practitioners to increasingly rely on them for answering questions about hard…
Bridging the Information Gap: Semantic Densification and Hindsight Distillation for Cold-Start Prediction
New-user cold-start is a critical bottleneck for e-commerce platforms: predicting user lifetime value (LTV) and conversion rate (CVR) for u…
Otap:Structure-Aware Optimal Transport for Evaluating Planning and Execution in Agent Trajectories
Large language model agents solve tasks by generating trajectories that interleave planning, tool calls, and intermediate results. Current…
Fourier Geometric Wind Power Forecasting with Numerical Weather Prediction
Accurate short-term wind power forecasting is essential for grid stability and operational planning, yet remains challenging due to the com…
Evidence Interfaces Shape How Retrieval-Augmented Readers Use Support
In multi-hop RAG evaluation, a top-k answer score can hide two different failures: the retrieval window may drop part of the support chain,…
A Diagnostic Framework for AI Agent Behavior
AI agents increasingly act within the same clinical, political, scientific, and social systems that behavioral scientists study. Evaluating…
Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum Alignment
Test-time scaling empowers Large Reasoning Models (LRMs) to tackle complex tasks via extensive Chain-of-Thought (CoT). However, this often…
Toward Anthropomorphic Dialogue: A Closed-Loop Framework for Human-Like Chat Generation, Evaluation, and Preference Alignment
Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowle…
A Systematic Evaluation of Trajectory Data Curation for LoRA Fine-Tuning of Code Agents
Supervised fine-tuning (SFT) of open-weight LLMs on expert agent trajectories has emerged as a prominent approach to building capable code…
Constrained Path Reasoning: Measuring When Committed Stages Earn Their Cost
When does a committed intermediate stage in an LLM reasoning pipeline earn its cost? Constrained Path Reasoning (CPR) pairs a source-aware…
LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning
Multi-view spatial reasoning requires vision-language models to compare visual evidence across images, align object correspondences, and in…
Coordinated Disentanglement with Iterative Mode Discovery Under Hidden Correlations
Disentangled representation learning is a powerful paradigm for robust attribute prediction. While recent methods address attribute correla…
An Explicit World Model Based on Data-First Ontology: DaoQL Multimodal Storage Validation and Counterfactual Reasoning Evaluation
Large language models encode world models implicitly in neural weights, which exposes four structural risks in high-precision domains such…
Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware
Single-stream autoregressive decoding of large language models is bound by memory bandwidth: each generated token requires one full forward…
Learning-Driven Adaptive Audit Scheduling: A Sequential Decision Approach to Off-Chain Data Integrity
We model cryptographic auditing of off-chain data as a Constrained MDP (CMDP) under partial observability: the storage node's hidden type a…
Agentic ERP: Multi-Agent Large Language Model Architecture for Autonomous Enterprise Resource Planning
Enterprise Resource Planning (ERP) systems record transactions reliably but still delegate almost all operational decision-making to human…
DeeperRadar: End-to-End MIMO Radar Design and Multi-Modal Fusion for Autonomous Vehicle Perception
DeeperRadar is a radar-centric, sensor-stack-conditioned framework that co-designs radar sensing and multi-modal 3D detection for autonomou…
Self-Modifying Lean Proof Agents with Verifier-Grounded Benchmark Coevolution
Designing effective Lean proof agents is a central challenge in formal mathematical reasoning. Beyond building stronger provers, recent wor…
Quantifying Diversity of Thought: A Predictive Law of Weighted LLM Ensemble Lift
This paper provides an experimentally verified formal law for calculating the uplift that diversity of thought provides in Large Language M…
Intermittent Control Is Not Diluted Control: A Switching Effect in Artificial Agency
Adaptive agents do not always regulate under the same timing conditions. Sometimes stabilization can begin before a disturbance has fully e…
Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions
Large language model (LLM) agents offer a generative approach to simulating human behavior under conditions that may have few or no direct…
AEC-DS: Adaptive Erasure Coding with PDP-Triggered Reputation and QoS-Aware Migration for Decentralized Storage
In decentralized storage systems, audit results are often not used directly to guide later redundancy and shard-placement decisions, which…
Panache: One-Pass Motif Discovery at Every Window Length
Motif discovery, the search for recurring patterns within a time series, is a core primitive of exploratory data analysis. A pattern, howev…
Pailitao-MMSearch: Building Native E-Commerce Multimodal Search Foundation
The evolution of e-commerce has fundamentally transformed how users search for products, shifting from simple text-based keyword queries to…
Can AI Agents Really Complete RTL-to-GDS? Lessons from Benchmarking Tool-Interactive EDA Workflows
LLM-driven agent systems have emerged as a promising paradigm for electronic design automation (EDA), demonstrating strong potential for au…
The Curvature Shadow: An Apparent Failure of Maximum-Entropy Equilibrium Selection is a Removable Artifact
In two-player zero-sum games whose Nash equilibria form a convex set, regularized solvers such as Regularized Nash Dynamics (R-NaD) empiric…
Retain or Consolidate? Budget-Dependent Operator Selection for Language Agent Memory
Language agents depend on memory across interactions. However, the limited context windows of large language models (LLMs) and their infere…
Why Does Feedback-Augmented Self-Distillation Fail to Improve Retrieval-Interleaved Search Agents?
On-policy self-distillation (OPSD) offers a promising approach for training large language models without relying on a separate teacher mod…
Reinforcement Learning: From Algorithms To Foundation Models
Reinforcement learning (RL) provides a framework for sequential decision making under explicit objectives. In its classical form, RL studie…
ZifaMem: Structured Memory for Persona, Preference, and Emotional Continuity in AI Companions
AI companions are judged not only by single-turn fluency but by whether they sustain emotional continuity: remembering who the companion is…
A Dual-Hypothesis Reasoning Framework for LLM Guardrails
We propose ARBITER, a novel LLM guardrail framework that introduces two key ideas: (i) dual-hypothesis reasoning, a reasoning method for LL…
Is Progressive Disclosure All You Need for Long-Context Agents?
Long-document question answering usually forces a choice between loading the whole document into the context window and bolting on a separa…
Mechanistic Attention Guidance for Agent Memory Refinement
Existing self-evolving memory systems mainly improve agent memory based on textual outputs, such as task trajectories and reflections. Howe…
Verify, Repair, Repeat, or Stop? Robust Stopping for Noisy Verify-Repair Loops in LLM Agents
Verify-repair loops are a standard means for large language model (LLM) agents to correct faulty plans in code generation, mathematical rea…
FlowBlock: Wavefront-Parallel Decoding for Self-Correcting Diffusion Language Models
Block-wise diffusion large language models (dLLMs) decode sequentially at the block level, enabling effective KV-cache reuse across blocks…
OrientSAM: Mitigating Camera-Centric Shortcut in Multimodal Spatial Reasoning via Orientation-Aware Spatial Alignment
Multimodal large language models (MLLMs) still struggle with spatial reasoning that requires perspective transformation. In particular, the…
Artificial Intelligence for Understanding and Managing Transportation Behavior in Sustainable Smart Cities
Urban transportation systems generate heterogeneous data, yet these data do not automatically become actionable management intelligence. Th…
ProEvent: An Event-centric Benchmark for Proactive Agents
Proactive agents are expected to anticipate user needs and provide autonomous assistance by perceiving environmental context without explic…
LaT: LLM-as-Trainer for Multi-Task Vehicle Routing Solvers
Multi-task neural solvers aim to handle multiple Vehicle Routing Problem (VRP) variants within a unified model, avoiding separate training…
Learning to Detect Cross-Modal Negation: An Analysis of Latent Representations and an Attention-Based Solution
Detecting high-level semantic concepts like negation across modalities remains a challenge for current multimodal systems. We analyze this…
SR-Agent: An Experience-Driven Agentic Framework for Post-Ranking Strategies Refinement in E-Commerce Recommendation
User experience is a first-class objective in industrial e-commerce recommender systems (RS). Post-ranking strategies, which govern diversi…
Semantically Similar, Logically Distinct: Diagnosing the Semantic-Answerability Gap in Table RAG
Tables are a critical knowledge source in retrieval-augmented generation (RAG), but a retrieved table may lack sufficient evidence to answe…
WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement
Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement…
Dynamic Defense Profiling Enables Cognitive Jailbreak of Text-to-Image Models
Text-to-Image (T2I) generative models have achieved remarkable progress in synthesizing high-quality visual content, yet they remain vulner…
Financial Audit Assistance using Misinformation Detection and Explanation
Financial statements (FS) such as Balance Sheet (BS), Income Statement (IS) and Cash-flow Statement (CS) summarize the annual financial per…
PGN: Design and Implementation of a Vision-Language Navigation System Based on Pangu Multimodal Foundation Model
Vision-Language Navigation (VLN) requires an embodied agent to interpret a natural-language instruction and predict actions from temporally…
A Hardware-oriented Approach for Efficient Bayesian Inference Computation and Deployment
Bayesian inference provides a principled foundation for reasoning under uncertainty, but its computational cost hinders deployment on resou…
Exploratory and Assimilating Reflection: Reflective Recall Cycle for Long-term Memory
LLM-based autonomous agents require external memory to overcome their statelessness and limited context window for long-term interaction an…
ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding
Vision Language Models (VLMs) achieve strong reasoning with Chain-of-Thought (CoT) prompting but incur high sequential-generation cost, err…
Stress Testing Concept Erasure with Large Language Model Agents
Concept erasure aims to remove semantic concepts from a trained generative model and is increasingly important for responsible AI deploymen…
PEARL: Auditable Repair for Scientific Reasoning Graph Extraction
Scientific Reasoning Graph Extraction (SRGE) aims to recover explicit links among observations, evidence, intermediate claims, and paper-le…
The Autonomous Agency Scale: A Behavioral Framework for Measuring Self-Directed Behavior in AI Systems
Existing AI measurement frameworks quantify cognitive capability, task automation, or catastrophic risk, but none measure autonomous agency…
Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking
Agent-based models (ABMs) rely on simple, explicit and reproducible rules for individual decision making, while complex collective behavior…
OntoExtend: A Framework for Requirement-driven and Scalable Ontology Extension with LLMs
Ontology extension refers to the process of enriching an existing ontology in response to emerging requirements, making it more complete. T…
SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning
Latent world models have emerged as a powerful planning paradigm by learning action-conditioned predictive dynamics and using them as inter…
Do Maps Still Matter for Machines: Revisiting the Role of Choropleth Maps in Foundation Model Spatial Understanding
Spatial understanding is crucial for foundation models (FMs), and maps have long helped humans organize and reason about geographic informa…
PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning
Many visual reinforcement learning (RL) algorithms learn representations by matching latent distances to a behavioral distance induced by r…
Rethinking Heterogeneous LLM Merging: A Weighted Model Averaging Perspective
Can large language models with substantially different parameter spaces be merged by direct weighted averaging, without training or semanti…
AdaHome: An Adaptive Smart Home Assistant using Local Small Language Models
Smart home assistants interpret a wide range of user commands, from explicit device control to underspecified and preference dependent requ…
The Shared Discovery Paradox: How a One-Answer Rule Turns Better Information into Worse Search
Organizations often pool dispersed information into one ranking and then allow many agents to act on that shared view. In a discovery probl…
WorldCupArena: Fine-Grained Evaluation of Language Models and Deep-Research Agents on Football Forecasting
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear…
Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs
Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whethe…
Can We Break LLMs Out of Self-Loops? Fine-Grained Reasoning Control with Activation Steering
Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely…
SGA: Plug&Play Geometric Verification for Educational Video Synthesis
Recent work leverages Large Language Models (LLMs) to generate executable code for pedagogical animations using libraries such as Manim. Ho…
Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
To test how correct logical judgments respond to learned context, we prepend a soft prefix to an exactly labeled syllogistic reasoning benc…
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction,…
What Makes Linguistic Representations Good Models of High-Level Visual Perception in the Human Brain?
Image descriptions represented with language models (LMs) predict human brain responses to naturalistic images in high-level visual regions…
Comparing Spectrogram Front-Ends for Abnormal Heart-Sound Detection with a Convolutional Neural Network
Heart disease kills a lot of people, and one cheap way to catch it early is by listening to heart sounds with a stethoscope, or better yet,…
Fully-sensorized smart-eyewear platform for on-device Machine Learning
This paper presents ARGO, a smart eyewear platform designed to bridge ergonomic comfort, high computational throughput, and energy efficien…
International Agreements to Limit Frontier AI: Objectives and Exit
An international agreement to limit AI development could be crucial to mitigate risks from AI. However, it remains unclear which conditions…
RouteCost: A Production-Inspired Multi-Stage Framework for Pre-Order Shipping Cost Estimation in E-Commerce
Accurate pre-order shipping cost estimation is important in e-commerce because it affects price presentation, margin planning, and conversi…
From Weights to Words: Expressing and Editing Preference Model Inferences in Natural Language
The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundament…
Token-Level Cross-Modal Transformer with Contrastive Multi-Task Learning for Breast Cancer Subtype Classification and Survival Prediction
Integrating heterogeneous genomic and clinical modalities for joint cancer subtype classification and survival prediction remains a key cha…
HantaWatch: Federated Learning for Hantavirus Genomic Surveillance
Hantavirus genomic surveillance is limited by the distribution of sequence data, non-IID source heterogeneity, and constrained expert-revie…
OpenMHC: Accelerating the Science of Wearable Foundation Models
Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching. Howeve…
The Failures of Marginal Influence-Based Attribution Methods for Global Time Series Explanations
Explainability methods for time series models predominantly produce flat attribution scores: they quantify the direct influence of a featur…
Quantizing Recursive Reasoning Models
Recursive reasoning models solve hard puzzles by applying compact, weight-tied blocks over many refinement steps. Because these blocks are…
Diffusion-corrected Autoregressive Fourier Neural Operator for Droplet Evolution Prediction
Predicting droplet evolution in material jetting, or Inkjet Printing (IJP), is essential for maintaining printing quality. However, long-ho…
Normalized Rewards for Preference Optimization
Direct Alignment Algorithms (DAAs) such as DPO have become a common way to post-train and align LLMs with human preferences. However, DAAs…
KernelBench-Verified: Do LLM-Generated Kernels Actually Beat PyTorch?
Recent large language models (LLMs) can generate custom CUDA kernels that appear to outperform PyTorch on benchmarks such as KernelBench. B…
TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment
Fine-Tuning-as-a-Service (FTaaS) platforms let users train large language models (LLMs) on customized tasks, but this pipeline could erode…
RobustMAD: Evaluating Real-World Robustness of Multimodal Small Language Models for Deployable Anomaly Detection Assistants
Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive visio…
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents
Training multi-turn evidence-reading agents with outcome-only reinforcement learning is unstable because intermediate turns receive little…
Learning Structural Manipulability in Gate-Level Netlists Using Graph Neural Networks
Gate-level netlists exhibit intrinsic structural properties that influence signal propagation independently of functional simulation. We de…
Let the Data Decide: Supervision Analysis, Capability Trade-offs, and Adaptive Objective Routing in Continued Pre-Training via Off-Policy Distillation
Off-policy distillation is now central to large language model pre-training, yet how training data, objective parameterization, and model c…
High-accuracy Low-Bit KV-Cache Quantization via Local Distribution Restoration
Long-context large language model inference relies on the KV cache to avoid redundant attention computation, but incurs high memory and ban…
An Agentic Interface for End-to-End Probabilistic Seismic Hazard and Risk Analysis
Probabilistic seismic hazard and risk analyses are backbone to building codes, insurance pricing, and disaster management. Yet their open-e…
Benchmarking Machine Learning Models for Multi-Omics-Based Breast Cancer Prediction
Estrogen Receptor (ER) status is a critical biomarker in breast cancer diagnosis, prognosis, and treatment selection. Recent advances in hi…
SOS-LoRA: Static Orthogonal-Subspace Low-Rank Adaptation with Fixed Multi-Scale Scaling
Low-Rank Adaptation (LoRA) is a widely used parameter-efficient fine-tuning (PEFT) method for large language models. Under a fixed rank bud…
Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis
Machine learning-based Type 2 diabetes risk prediction models obtain good internal validation results but lose effectiveness in real-world…
Feature Generation Using LLMs: An Evolutionary Algorithm Approach
A crucial step in machine learning pipelines is to present each entity with features or attributes that are representative of the character…
Discovery by Dreaming: Cross-Domain Recombination in Artificial Memory
Dreams splice together people, places, and times that never met. Neuroscience suggests this recombination is not noise, but a function driv…
From Outcomes to Actions: Leveraging Hindsight for Long-Horizon Language Agent Training
Reinforcement learning (RL) has become a widely adopted technique for improving large language models (LLMs) on complex tasks. Despite this…
Neural Controlled Differential Equations for EMT-Level Surrogate Modeling of Grid-Forming Inverters
The application of artificial intelligence methods in power electronic converter modeling is becoming increasingly widespread, but existing…
AdaSurvMamba: Dynamic Fusion and Semantic Scanning for Multimodal Survival Analysis
Multimodal survival analysis utilizing whole slide images (WSIs) and genomic profiles is fundamental for cancer prognosis. Recently, state-…
Reducing Per-Sample Harm in Stochastic Optimization
Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments. W…
Autonomous mechanistic discovery of colorectal cancer vulnerabilities via multi-scale AI swarms
The acceleration of automated scientific discovery has been fundamentally bottlenecked by the epistemic gap between the semantic reasoning…
Composable Verification Pipelines for Multi-Agent Systems
Existing approaches for reasoning about action and change provide expressive semantics for modeling dynamic systems, in most cases built on…
From Intent to Infrastructure: LLM-Driven Agent Compilers for ISAC Networks
Integrated sensing and communications (ISAC) is moving from proof-of-concept demonstrations to system-level deployment in sixth-generation…
Physics-Informed Feature Engineering 1D-CNN for Multilayer Cloud Detection from Geostationary Satellites
Multilayer cloud detection from active--passive observation is vital for numerical weather prediction. In this study, channel selections de…
ForensicNet: Lightweight Attention-Enhanced MobileNetV2 for Automated Face Identification
In forensic environments, automated identification of perpetrators is difficult due to pose changes, changes in light, occlusion, and lack…
Intelligence-Guided Adaptive Purification for DDoS-Resilient Quantum Networks: A CUDA-Q based Study
Quantum-repeater networks require adaptive control policies that balance entanglement generation rate, end-to-end fidelity, purification ov…
GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
The rapid advancement of generative AI has outpaced our ability to reliably detect its outputs, particularly when detectors encounter gener…
Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction
The 3D geometry of real-world scene data is often incomplete. Mainstream methods use depth estimators to inpaint missing structure. However…
Identity-Consistent Expression Fields: A Disentangled Neural Radiance Field Framework for Few-Shot Facial Expression Synthesis
Neural Radiance Fields (NeRF) have enabled photorealistic novel-view synthesis of 3D scenes and, in the facial domain, have been extended t…
It Depends on the Dataset: When a Brain-Encoding Model's Predicted Responses Beat Their Visual Backbone for Video Memorability
Brain-encoding foundation models predict fMRI responses to video, audio, and text well enough to win the Algonauts 2025 challenge. We ask w…
A${}^2$BM: Alignment-Aware Bridge Matching for Image-to-Image Translation
Paired image-to-image translation underpins a wide range of computer vision tasks, including image editing, sensor translation, and domain…
Emergent Hierarchical Monosemantic Neurons from the Group-Contrastive Forward-Forward Algorithm
Mechanistic interpretability has made significant strides in understanding neural network representations, with sparse dictionary learning…
Efficient EEG Seizure Detection Using INT8 Quantization, Channel Pruning, and Spiking Neural Networks
Continuous EEG monitoring for epilepsy is constrained by the limited power and memory budgets of wearable and implantable devices. Deep neu…
Med-OPD: Improving Medical Vision-Language Models via Evidence-Aware On-Policy Distillation
Medical Vision-Language Models (Med-VLMs) require reliable reasoning from fine-grained visual evidence, yet existing models can produce pla…
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Expert…
DAUPNet: Domain-Aware Uncertainty Modeling for Reliable Prototype Discrimination in Cross-Domain Few-Shot Semantic Segmentation
Cross-domain few-shot semantic segmentation (CD-FSS) has predominantly been formulated as learning domain-invariant representations or impr…
Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs
Vision-language models (VLMs) often answer visual questions using learned language and category priors rather than grounding their predicti…
Monte Carlo Dropout Uncertainty and Entropy-Thresholded Selective Prediction for Architecture-Agnostic Brain Tumor MRI Triage
Deep networks now subtype brain tumors on MRI about as well as specialist readers, yet accuracy is not what keeps them out of the clinic. W…
GMoT: Gated Motion-Aware Tokenization for Fine-Grained Micro-Gesture Video Reasoning with Multimodal LLMs
Micro-gesture recognition demands the detection of fleeting, spatially localized movements that are frequently overwhelmed by dominant stat…
DMFNet: Dual-Backbone Multiscale Fusion Network for Urban Scene Classification
This article presents DMFNet, a dual-backbone multiscale feature fusion framework with residual feature propagation and spatial attention f…
AEVAL: From Anecdotal to Deterministic Testing for Agentic Skill Workflows
Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a…
Boundary-Seeking GAN-Augmented TabTransformer for Adversarially Robust Intrusion Detection
Machine learning-based intrusion detection systems (IDSs) often suffer from class imbalance and vulnerability to adversarial attacks, leadi…
Joint-Embedding Predictive Architecture for Sensor-based Activity Recognition
Sensor-based human activity recognition (HAR) has achieved significant progressed in fully supervised learning settings. However, these sup…
Privacy-Aware Synthetic Video Benchmarking and Relational Evaluation for Worker-Under-Suspended-Load Detection
Publicly shareable construction-video benchmarks remain scarce, especially for safety-critical hazards that are rare, dangerous to stage, a…
Clarify Before Executing: A Self-Evolving Agent for Resolving Intent Asymmetry in 3D Tool Orchestration
A fundamental intent asymmetry plagues modern 3D asset creation: while state-of-the-art 3D toolchains demand precise, executable parameters…
A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting
Retail demand forecasting remains difficult when demand shifts faster than static forecasting models can be retrained, especially in early…
PhysAgent: Reflective Agentic Physics Control for Physically Plausible Video Generation
Recent advances in physics-grounded video generation leverage physics simulation as a physical prior to guide video synthesis toward physic…
Reliable Remediation Impact Prediction for Black-Box Security Ratings
Security rating platforms summarize externally observable cyber exposure and are expected to help organizations prioritize remediation. A p…
A Quantum-Classical Hybrid Framework for Multivariate Time-Series Forecasting Complexity-Fidelity Trade-offs and Limitations
This paper presents a unified quantum-classical hybrid framework for multi-horizon time-series forecasting, introducing two model variants…
PRISM: Multimodal Terrain Mapping for Rover Navigation in Unstructured Environments
Robotic navigation in unstructured environments requires robust situational awareness to safely traverse hazards such as steep slopes and r…
AoA: Theorem Proving Agent over Abstract Syntax Tree of Redesigned Language
Interactive theorem proving (ITP) underpins program verification and formalized mathematics, but its manual effort limits scalability. LLM-…
Fantastic Adaptive Taxonomies and How to Use Them
An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajector…
Automated Hardware Validation Test Plan Generation for Large Scale AI Datacenter Platforms Using a Generative AI Multi-Agents Architecture
Large-scale AI datacenter platforms comprise thousands of heterogeneous hardware components whose validation requires comprehensive fault i…
Think, Plan, Paint: Layout-Aware Reasoning for Controllable Image Generation in Unified Models
Unified Multimodal Large Language Models (MLLMs) offer a promising paradigm for unifying visual understanding and generation, yet they stil…
Signal-based Model Access Risk Analysis for AI System Operations Security
Artificial intelligence (AI) systems are now ubiquitous across domains such as security, finance, healthcare, consumer technology, and larg…
Back to the museum: Investigation of the acceptance of Android Andrea with and without emotion simulation in a museum
For a second time, the android robot Andrea was set up at a public museum in Germany for six consecutive days to have conversations with vi…
Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent
Interpretability methods for neural network activations span a wide cost spectrum, from cheap, training-free techniques (such as linear pro…
Committed Before Reasoning: Behavioral Reproduction and Preliminary Activation-Level Evidence of Answer Pre-Commitment in an Open-Weight LLM
Chat models sometimes commit to an answer and then produce reasoning that justifies it rather than deriving it -- even when the answer cont…
When to Use Which? Benchmarking Optimisers for Configurable Systems under Varying Budgets
Software configuration tuning is crucial for optimising system performance, and various optimisers have emerged over the last decade. Yet,…
K-IPO: Kendall-constrained Importance Preserving Oversampling for Imbalanced Tabular Data
Oversampling is widely used to address class imbalance in tabular classification, but existing methods can distort the feature importance r…
How Formerly Incarcerated People Envision Technologies for Prison Parole
AI-driven algorithms and automated tools are increasingly embedded in the correctional landscape, shaping parole eligibility,release decisi…
Geometry-Enhanced Portion Estimation for Multimodal LLMs
Image-based dietary assessment promises to replace costly, bias-prone manual recalls, but portion estimation remains a major blocker. Multi…
Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning
Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action spac…
Capacity and Redundancy Trade-offs in Multi-Task Learning
In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence…
Mitigating Compiler Fusion-Induced Power Bursts in Mobile NPU Inference as the Battery Depletes
Mobile devices increasingly rely on real-time NPU inference for camera and perception workloads. Under low-voltage conditions, however, a s…
ReqGenX: An Empirical Study of Atomic Decomposition, Artifact Regeneration, and Reconstruction for Legacy SRS Documents
Background: Evaluating automated Software Requirements Specification (SRS) generation is challenging because few datasets provide fine-grai…
Autonomous VR-Based Risk Detection for Situational Awareness in Dangerous Settings
In high-risk environments such as disaster response, situational awareness depends not only on detecting hazards but also on communicating…
Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
The RLxF programme argues that learning signals should come from world feedback rather than from internal model proxies. We instantiate thi…
DataFlow-Harness: A Grounded Code-Agent Platform for Constructing Editable LLM Data Pipelines
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that…
Privacy Cost as Equity Input: A Group Fairness Criterion for Differentially Private Machine Learning
Differential privacy (DP) is increasingly deployed to limit membership inference risk in machine-learning systems. Prior work has shown tha…
CLOSER-Bench: Evaluating Budgeted Cross-Stage Design Closure for Hardware Agents
Hardware engineering exposes coding agents to a form of long-horizon work that is difficult to capture with pass-at-k: progress is continuo…
TellTale: Blending Multi-Instance LoRA Text Encoders and a Zero-Shot LLM Judge for Ambivalence/Hesitancy Recognition in Videos
We present TellTale, a text-only approach to ambivalence/hesitancy (A/H) recognition in interview videos, evaluated on the BAH dataset as p…
Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
Large language models now translate natural-language descriptions of decision problems into solver-ready optimization models, but they fail…
Position: Explanation Stability Is a Property of the Model Method Pair, Not the Model
This position paper argues that claims about explanation stability are scientifically invalid without cross method validation. Just as stat…
How Do You Choose Your AI Component? An Interview Study of Secure AI Integration in Practice
The increasing adoption of Large Language Models (LLMs) as AI components in modern software systems introduces distinct security risks to t…
Building a Neural Network from Scratch: Implementation, Evaluation, and Optimization
The widespread adoption of high-level deep learning libraries, while accelerating model development, has increasingly abstracted away the i…
OFD-Net: Teacher-Free Reliable Semi-supervised Medical Image Segmentation with Orthogonal Feature Disentanglement Net of Foreground-Background
Semi-supervised learning (SSL) is an effective solution for medical image segmentation with limited annotations. Existing SSL methods mainl…
A Causal Markov Condition for Value
This paper proposes a causal independence principle for value -- the value Causal Markov Condition (v-CMC) -- and develops the conceptual a…
RealDESED: A Real-World Domestic Sound Event Detection Benchmark
This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 6…
Spatiotemporal Facial Action Unit Detection using Twin Cycle Autoencoders for Driver Monitoring
Driver monitoring systems (DMS) increasingly rely on facial cues to infer drowsiness, distraction, and cognitive load in real time. Facial…
JOR-Bench: Japanese Operations Research Benchmarks for Large Language Models
We present JOR-Bench, a collection of five Japanese-language benchmarks for evaluating the ability of large language models (LLMs) to formu…
Explainable Lightweight Compact Deep Models for Speech Emotion Recognition
Speech Emotion Recognition (SER) is an important component in a wide range of human-centered applications, including healthcare, customer s…
First-Order Predictable but Pairwise Fragile: Local Task Adaptation in Trained Transformers
Task arithmetic, sequential fine-tuning, activation steering, and first-order random search all operate through relatively small perturbati…
Beyond Memory Leaderboards: Evaluating Scientific Memory as Budgeted Context Restoration
Long-term memory is becoming a core component of LLM agents, but most memory benchmarks evaluate conversations or compact summaries, while…
Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators
Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not prec…
Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LL…
Trace-Based On-Policy Distillation for Masked Diffusion Language Models
Diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. However, reasoning-oriented post-training…
A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing
We introduce the vehicle routing problem with stochastic demands and outsourcing options (VRP-SDO), in which a logistics service provider p…
Certified-Gap Dual-Price Policies for Real-Time Truckload Bid Acceptance with Relocating, Clock-Constrained Resources
A truckload carrier must accept or reject each load tender within seconds. The decision depends on fleet state, hours-of-service (HOS) cloc…
A Method for Learning Value Systems in Generative AI
Value-aware AI systems require explicit computational representations of human values (groundings) and their aggregation into value systems…
PREFAIL: Identifying Precursors to Failures in Robotic Lift-and-Place Tasks to Improve Task Execution Performance
Non-prehensile manipulation enables flexible material handling with part carriers, but friction-based support makes high-speed motions fail…
A Multi-Agent System for 5G Throughput Prediction in Multi-Operator Urban Environments
Throughput prediction is foundational for artificial intelligence-driven 6G resource orchestration. Conventional monolithic machine learnin…
Optimizing Clinical Trial Protocols Using EHR-Derived Heterogeneous Treatment Effects
Traditional randomized trials often obscure clinically meaningful heterogeneity in treatment response by focusing on average effects. Lever…
Pediatric Bone Age Prediction Using Deep Learning
Pediatric bone age prediction is a crucial task in clinical practice that can help diagnose endocrine disorders and provide insight into a…
What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning
End-to-end autonomous driving models are now able to navigate complex road scenarios, mapping raw sensor observations directly to observed…
Investigation of Polycystic Ovary Syndrome (PCOS) Diagnosis Using Machine Learning Approaches
Polycystic Ovarian Syndrome (PCOS) is a widespread hormone problem for women of childbearing age. Women with PCOS may not ovulate; they mig…
CADENCE: Closing the Reasoning Gap via Coverage-Adaptive On-Policy Distillation
On-policy knowledge distillation transfers reasoning from large teachers to compact students, but existing approaches suffer three compound…
TurboVec: A Case Study in Cost-Efficient Private Retrieval for Enterprise RAG via Codebook-Oblivious Quantization
Retrieval-Augmented Generation (RAG) systems increasingly power enterprise LLM applications, yet the vector retrieval layer introduces two…
Real-World Evaluation of an AI Agent Drafting Translational Impact Summaries
Introduction. Clinical and Translational Science Award (CTSA) programs must document their scholars' research impact, but assembling each s…
Automated Cardiac Adipose Tissue Segmentation in Computed Tomography: A Literature Review
This review provides an overview of recent advancements in automated segmentation methods on Computed Tomography (CT) for two types of card…
Counterfactual Shapley Credit Assignment
The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing fra…
Scalable Causal Imitation Learning
Imitation learning enables learning a policy in an unknown environment with a latent reward signal using expert demonstrations, but it stru…
Alignment of a Total Automation Economy
We consider economic theory from the perspective of a total automation economy, one with no human involvement in production either in manuf…
WHALE: A Scalable Unified Model for Recommendation with Wukong-HSTU Architecture
As scalability becomes increasingly important in recommendation modeling, recent architectures have advanced the modeling of two broad sour…
Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent
Multi-step enterprise agent tasks fail in a characteristic way: single-pass inference has no checkpoint between deciding an answer and comm…
Solver-Hard Is Not Model-Hard: A Hardness-Controlled Diagnostic for LLM Constraint Reasoning
LLM constraint reasoners are often evaluated near the random-SAT phase transition, confounding density and solver hardness. We test instanc…
EvoGUI: An Evolution-Aware Benchmark for GUI State-Transition Understanding
GUI agents must reason about how actions transform interface states, but end-to-end success rates entangle this ability with perception, gr…
ThAME: 3D Memory-Enabled Heterogeneous Accelerator for LLM Mixture of Experts
Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling Large Language Models (LLMs). However, MoE inference…
ALLUDE: A Unified Evaluation System for Configurable Attacks in Differentiable Environments
Adversarial attacks against vision models like object detectors are often evaluated under limited conditions, leaving their performance und…
DepthART: Scaling Foundation Monocular Depth to Tiny Models
Recent geometric foundation models (e.g., Metric3D, Depth Anything and UniDepth) have substantially improved monocular depth estimation (MD…
Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
An AI research agent can improve the score it sees without finding a modelling change that works on new materials. We ask a stricter questi…
Teach it to stop, not just to click
Agentic computer-use RL is reported in single runs, and those numbers mislead. Using verifier-guided repair of a 35B computer-use agent (CU…
Noise-Robust Box-Supervised Infrared Small Target Detection via Physics-Inspired Soft Label Optimization
Infrared small target detection (IRSTD) commonly relies on pixel-level mask supervision. Such annotations, however, are costly and inherent…
VLA-ReID: Video-Level Association for Re-Identification in Multi-Object Tracking with Highly Similar Objects
Multi-object tracking (MOT) aims to localize multiple objects in videos while preserving their identities over time. Long-term identity pre…
DADIR: Density-Aware Data-level Imbalanced Regression Framework
Imbalanced learning addresses predictive modeling problems with underrepresented regions of the data distribution. Although widely studied…
Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs
Serverless multi-model LLM systems multiplex popularity-skewed model catalogs over shared GPU pools, yet typically schedule each request in…
Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat
Human survey respondents exhibit question-order effects that satisfy the QQ (quantum question) equality, an a priori, parameter-free predic…
Specifying the Delegated-Autonomy Boundary: Requirements Engineering for Agentic AI
Agentic AI systems do not just predict or recommend; they plan, maintain state, and act in external environments with varying degrees of au…
A RFID Based Campus Wide Payment System
This work titled "RFID Based Campuswide Payment System" introduces an innovative cashless payment solution for educational institutions. It…
A Large-Scale Measurement of AI Bill of Materials Completeness in Hugging Face Models
Pretrained machine learning (ML) models help developers build ML-intensive software systems without training models from scratch. However,…
Distilled Reinforcement Learning for LLM Post-training
Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow tw…
Asynchronous Multimodal Diffusion Policy Composition via Latency-Aware Guidance Fusion
Diffusion policies have shown strong potential for robotic imitation learning, and recent extensions incorporate additional modalities to i…
Debate-on-Graph: Reliable and Adaptive Reasoning of Large Language Model on Uncertain Knowledge Graph
Large language models (LLMs) have demonstrated remarkable capabilities in natural language processing. However, LLMs often suffer from hall…
Between Safe Boundaries: Exploiting Temporal Consistency for Jailbreaking Text-To-Video Generation Models
Recently, text-to-video (T2V) models have been widely deployed, sparking growing concerns over their robustness against jailbreak attacks.…
AIGB-R1: Self-Evolving Generative Auto-Bidding via Hierarchical Planner-Executor Optimization
Auto-bidding plays an essential role in online advertising, automatically adjusting bids for advertisers to optimize their commercial goals…
SAGA: Synthetic Agentic Graph Architecture for Temporal Benchmark Generation
High quality temporal graph benchmarks with rich semantics and ground-truth anomaly labels are essential for training graph neural networks…
Lookahead Branching for Neural Network Verification
In this work, we investigate the effect of lookahead branching strategies in neural network verification. We present a general recipe to in…
WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning
Long-horizon rollout generation has become the dominant systems bottleneck in agentic reinforcement learning (RL). As agents interact with…
The Optimization Trilemma: Efficiency, Comfort and Fairness in Decentralized Multi-agent Coordination
The problem of fair multi-agent coordination in decentralized settings is one of the most pressing challenges for building efficient collab…
TAPAS: Throughput-adaptive Perception for Autonomous Systems
Autonomous systems rely on a perception module to navigate through dynamic environments. In real-world scenarios, the perception module's t…
STAR: Skeletal Token Alignment and Rearrangement for Interaction Recognition
Understanding physical human-robot and human-human interactions is a challenging yet emerging topic in 3D vision. While most existing metho…
Mathematical Discovery in the Wild: AI-Guided Proofs in Banach Space Theory
We investigate the capacity of current language models to contribute to mathematical research. In Banach space theory, AI systems generated…
A Phased Development Framework Enabling Islanded Operation of Sustainable AI Data Centers With Onsite Grid-Following and Grid-Forming Energy Architectures
As hyperscale and colocation AI data centers continue to expand, the electric grid is increasingly required to support large, concentrated…
CoEvoP&R: Co-Evolving Placement Objectives with Routing Feedback via Large Language Models
Analytical placers rely on differentiable objective functions to guide placement, typically combining intermediate surrogate metrics such a…
Kernelized Linear Attention: Breaking the Capacity Wall with Symmetric Cones
Linear attention promises constant-time recurrent inference but degrades sharply on associative recall. We formulate attention recall as a…
HyCoRec: Hypergraph-Enhanced Multi-Preference Learning for Alleviating Matthew Effect in Conversational Recommendation
The Matthew effect is a notorious issue in Recommender Systems (RSs), \emph{i.e.}, the rich get richer and the poor get poorer, wherein pop…
Multilingual Sentence Embeddings for Linguistic-Integrated Reliability Audit
Multilingual assessment systems commonly rely on translation for scoring and quality-control processes. We evaluate whether multilingual se…
SALT: Salience-Aware Lexical Trie for Long-Context Compression
As large language models (LLMs) process increasingly longer prompts, computation and KV-cache memory costs have emerged as major bottleneck…
DecoyFace: Beyond Obfuscation via Controllable and Imperceptible Identity Misdirection for Privacy-Preserving Face Recognition
Split face recognition reduces client-side computation but exposes intermediate features to feature inversion attacks and unauthorized anal…
Retrieval-Augmented Interpretable Learning: Towards Task-Specific Zero-Shot Models in Healthcare
We introduce Retrieval-Augmented Interpretable Learning (RAIL), a probabilistic meta-learning framework for zero-shot generation of task-sp…
After the Euclidean Highway: Hyperbolic Expert AI as the Next Innovation
Expert domains are trees; the Euclidean transformer is not, diluting parent-child structure exponentially at depth. The hyperbolic turn lef…
One-step lowest-variance selection in a Gaussian random-field model motivated by masked diffusion: Total correlation and a square root collision threshold
Motivated by confidence-guided parallel unmasking in masked discrete diffusion, we study a single selection step in a stylized Gaussian ran…
Thinking in Video: Can Video Generators Really Reason About the Real World?
Recent advances in world models and video generation have given rise to an emerging reasoning paradigm that leverages video generative mode…
Oracle Gap and Signal Fidelity: A Fixed-Pool Diagnostic for Test-Time Collaboration
Test-time collaboration, including self-consistency, best-of-N selection, critic models, and verifier pipelines, is often credited with bro…
CommitLLM: A Fine-Tuned Pipeline for Git Commit Message Generation
Developers frequently write uninformative git commit messages such as "fix" or "update stuff", degrading the value of version-control histo…
Human-in-the-Loop User Feedback Affects Perceived Accuracy and Trust, but Task Subjectivity Matters
While ML can produce complex models beyond those that a human could produce manually, incorporating human input can often improve performan…
Hierarchy-Aware and Anatomy-Guided Learning for Lung Ultrasound Video Classification
Lung ultrasound (LUS) is a bedside tool for assessing pulmonary edema in patients at risk due to heart failure or impaired kidney function.…
COLIP-2: Olfaction-Vision-Language Embeddings
The Contrastive Olfaction-Language-Image Pre-training 2 (COLIP-2) model is a multimodal embeddings space that places olfaction as a first-c…
CoCurve: Cross-Module Co-Pruning Curvature for Training-Free Structured LLM Pruning
Structured pruning compresses large language models (LLMs) by removing whole computational units, such as attention heads and feed-forward…
Predictive Training with Latent Imagination for Visual Quadruped Navigation
Reinforcement-learning navigation policies for legged robots select actions reactively from current observations and short-term memory, wit…
Detection, Attribution, Narration: An End-to-End Pipeline for Explainable Money Mule Identification
Money mule accounts are critical facilitators of financial fraud, yet detecting them at scale remains challenging due to the heterogeneous…
Trustworthy Protein-Ligand Binding Affinity Prediction via Reliability-Aware Multi-Engine Fusion
Accurate protein-ligand binding affinity prediction is central to computational drug discovery, yet modern docking engines frequently disag…
Coarse-to-fine Framework for Generative MEF via Implicit Neural Representation
Multi-exposure fusion (MEF) expands the luminance range beyond what a single exposure can capture. Combining images taken at different expo…
Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture
Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings…
TypiCore: A Hybrid Active Query Strategy for Class-Incremental Learning on Time Series
Time series data play a pivotal role across numerous domains, including healthcare and manufacturing. In real-world environments, models mu…
Selectivity Matters: Source Node Influence Pruning for Unsupervised Graph Domain Adaptation
Unsupervised Graph Domain Adaptation (UGDA) aims to facilitate knowledge transfer from a labeled source graph to an unlabeled target graph…
Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning
Contrastive learning is increasingly moving toward settings with three or more modalities instead of image-text pairs. Yet, extending model…
Uncovering Latent Reasoning Strategies in Language Models
A language model $p_\theta(y \mid x)$ trained on reasoning tasks learns to solve problems via multiple distinct strategies, yet these strat…
Integrating High-Level Requirements to Low-Level Tests with Machine-Readable V&V Specifications
Modern software teams have mature tools for low-level testing, such as pytest, JUnit, and Jest, which make it inexpensive to write unit tes…
Lifelong Multi-Subsystem Pickup and Delivery with Buffer-Limited Handover Stations
Coordinating payload transfers between subsystems is a critical challenge in lifelong Multi-Agent Pickup and Delivery (MAPD). We study syst…
MXSens: Sensitivity-Aware Mixed-Precision Quantization for Efficient LLM Inference
4-bit quantization enables efficient LLM inference, but suffers from significant accuracy degradation due to outliers. Prior work addresses…
Mobile Network Control with a World Model
The increasing complexity of mobile networks necessitates intelligent and dynamic control strategies for efficient, energy-conserving manag…
DA-Fusion: Deformable Attention-Based RGB-D Fusion Transformer for Unseen Object Instance Segmentation
In logistics automation, precise segmentation of unseen objects is crucial for efficient robotic manipulation in cluttered environments. Ta…
Seg2Grasp: A Robust Modular Suction Grasping in Bin Picking
Current bin picking methods that rely heavily on end-to-end learning often falter when confronted with unfamiliar or complex objects in uns…
Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations. However, real-world tasks often exhibi…
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to…
Autonomous Discovery of Wireless Communications Algorithms
Large language model (LLM)-driven evolutionary search is an emerging algorithm-discovery paradigm that has already produced novel results i…
FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, f…
Measuring Monosemanticity in Sparse Autoencoders via Latent Activation Coherence
Within Explainable Artificial Intelligence, mechanistic interpretability uses Sparse Autoencoders (SAEs) to extract more interpretable feat…
Persona-as-Configuration: Generative Stakeholder Reporting for Agricultural Floods
Cyber-physical systems built on deterministic edge inference, such as on-vehicle flood detection for agricultural fields, produce structure…
CDIS: Cross-Dimensional Class-Agnostic 3D Instance Segmentation via 2D Mask Tracking and 3D-2D Projection Merging
Class-agnostic 3D instance segmentation is critical for robotic systems operating in unknown environments, enabling perception of previousl…
ETAS: An Effect-Typed Language for Agent Systems
ETAS is a programming language for agent systems that treats model-backed agents, tool calls, prompts, typed memory, human approvals, polic…
BrainNext: A General-Purpose Self-Supervised Foundation Model for Brain MRI Analysis
Foundation models pretrained using self-supervised learning have transformed computer vision by learning transferable representations from…
Feature Attribution-Based Explainability Analysis of Deep Learning Models in Predictive Process Monitoring
Predictive process monitoring supports the optimization and control of operational business processes by forecasting the future state or ou…
Reasoning as a Double-Edged Sword: Architecture and Cross-Stage Robustness in Vision-Language-Action Models
Does adding a reasoning step make a Vision-Language-Action (VLA) model more robust to perturbation? Intuitively, a policy that reasons befo…
Medical Imaging Fusing Vision Transformer: Laryngeal Cancer Screening with Explanation
Early and timely screening of laryngeal cancer is crucial for improving clinical outcomes. In recent years, NBI endoscopy has become a stan…
ReViV: Reconstructing the Viewer and the View in 4D from Monocular Egocentric Video
Egocentric devices, such as wearable front-facing cameras, provide a unique perspective for capturing the continuous interaction between a…
Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D--2D Registration for Liver Laparoscopy
Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is…
Phasor Attention: Mean Root Square Normalization for Phase Manifold Preservation
While Root Mean Square Normalization has become the de facto standard for accelerating modern sequence models, its reliance on the quadrati…
I wanted it to feel more personal: Customization of social AI as AI individualism in practice
Despite the growing availability of customizable social artificial intelligence (AI), such as ChatGPT, Grok, and Character.ai, we know litt…
Measuring and Improving Complex-Atomic Answer Consistency in Endoscopic VQA
Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather tha…
CaT-GS: Efficient 3DGS Rendering for Large Scale Scenes via Inter-frame Caching and Tile Scheduling
Recent breakthroughs in 3D Gaussian Splatting (3DGS) have advanced neural rendering with high fidelity and speed. However, its performance…
ConceptTree: Bringing Semantic Transparency to Black-Box Decision Making for Robotic Manipulation
Establishing interpretable decision-making processes in long-horizon robotic manipulation is critical for enabling reliable human oversight…
Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI
Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output…
Chemical filters for ultra-high-throughput materials screening and generation
Generative artificial intelligence is rapidly transforming materials design by enabling de novo exploration of immense chemical spaces. Yet…
DeLIVeR: Decomposed Learning for Information-grounded Veracity Recognition via Reinforced Knowledge Graph Exploration
Automated fact-checking remains a challenge for Large Language Models (LLMs) due to "query brittleness" in traditional retrieval systems. W…
The Aura in the Machine: Genealogy and the Status of the Work of Art in the Generative Era
This paper frames Generative Artificial Intelligence (AI) not as an unprecedented technological rupture, but as an industrial-scale manifes…
The Art of Not Forgetting
We introduce CMP (Cognitive Memory Primitive), an architecture that represents inputs as sparse relational codes, stores them in a two-tier…
A Geometric Perspective on Stabilizing Value Conflict Resolution
Large Language Models (LLMs) often struggle to navigate value conflicts when trained with the compressed scalar rewards of Reinforcement Le…
RT-SHCUA: Real-Time Self-Hosted Computer-Use Agent for UAV Control
Natural-language control offers a promising interface for unmanned aerial vehicles (UAVs), but directly applying self-hosted computer-use a…
Topological Signatures of Context-Level Reliability in TabPFN
TabPFN is a transformer-based foundation model for tabular prediction that performs inference without task-specific training by conditionin…
Harness Engineering for LLM-Driven GPU Kernel Generation
Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness depends on whether generated code can be r…
Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of i…
HAS: Highlight-guided Attention Steering for Multimodal LLM Video Summarization
Video understanding has become more and more important with the growth of Artificial Intelligence (AI) for video generation. Recently, Mult…
MADA-RL: Multi-Agent Debate-Aware Reinforcement Learning for Parameter-Efficient Reasoning in Compact Models
Large language models achieve strong reasoning performance, but often at prohibitive training cost - a challenge that is especially acute f…
Natural Language Access to Domain-Specific Metadata: A Reusable Framework for LLM Query Generation
Researchers need to answer ad-hoc questions about the contents of domain-specific archives but often lack the expertise to write structured…
Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
End-to-end vision-language navigation (VLN) with causal vision-language models can map instructions and egocentric observations directly to…
Adaptive Adversaries: A Multi-Turn, Multi-LLM Benchmark for LLM Agent Security
LLM-based agents process external content, exposing them to prompt injection and multi-turn manipulation. Most safety benchmarks evaluate d…
Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data
Coding agents can now be left alone to improve software against a score. In this pattern--recently popularized as "autoresearch"--the agent…
Human Grounded Evaluation of Large Language Models for Optical Network Automation
Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substant…
SGN: A Similarity-based Generative Network for Data Generation under Distribution Shift
Generative models trained on a source domain often produce samples that are poorly aligned with shifted target domains, limiting their effe…
Generalised Bellman recurrence and three dualities in sequential decision-making
What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: th…
Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection
Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its ass…
SelectInfer: Selective Neuron Loading and Computation for On-Device LLMs
Large Language Models (LLMs) have demonstrated remarkable capabilities across a range of Natural Language Processing (NLP) tasks, but their…
Enhancing Rubric-based RL via Self-Distillation
Rubric-based RL has recently shown promise in improving LLMs on open-ended tasks. A widely recognized limitation of rubric-based RL is limi…
How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs?
Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled fe…
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for mode…
Do Language Models Dream of Binding Molecules? Benchmarking LLMs under Spatial Constraints
Structure-based drug design (SBDD) leverages the 3D structure of protein targets, often complemented by other spatial constraints, to gener…
LLMs and Agentic AI Systems for Smart Grids: A Tutorial on Architectures and Applications
Large language models (LLMs) and agentic AI systems have evolved from natural language tasks to using external tools to plan, retrieve, and…
Differentiable Logic Gate Networks for Low-Latency EEG Classification on Edge Devices
Real-time EEG classification on edge devices is bottlenecked by the floating-point arithmetic of conventional neural networks. We investiga…
TRIM: Reducing AI-Generated CodeSlop via Agent Trajectory Minimization
Coding agents are increasingly used to accelerate code generation in many downstream tasks, such as fixing bugs, building applications, and…
OR Else: A Differentiable Trust Region for Policy Optimization
PPO and the GRPO baseline studied here use clipped surrogate objectives whose favorable-direction saturation introduces an abrupt change in…
A Continual Validation, Updating, and Decision-Making Framework for Self-Adaptive Digital Twins via Robust Model Predictive Control: A Case Study in Additive Manufacturing
Digital Twins rely on surrogate models to mirror physical systems in real time, yet these models can degrade as operating conditions evolve…
Learning Adaptive Safety Margins for Visual Navigation
Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mi…
GigaPath-Flash and GigaTIME-Flash: Efficient Pathology Foundation Models for Whole-Slide and Tumor Microenvironment Analysis
Foundation models have emerged as a driving force in computational pathology, with the potential to transform cancer diagnosis, prognosis,…
Simple Domain Generalization for Strong Pixel-Level Image Tampering Detection in Modern VLMs
Modern vision-language models (VLMs) have significantly improved image generation and editing capabilities, making pixel-level image tamper…
Automated Discovery Has No Universally Superior Harness
Autonomous discovery systems such as OpenEvolve and TTT-Discover are often used as general-purpose harnesses. However, in practice these ar…
Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning
Feature selection aims to preprocess the target dataset, find an optimal and most streamlined feature subset, and enhance the downstream ma…
Enhancing LLMs' Clinical Reasoning with Real-World Data from a Nationwide Sepsis Registry
Although large language models (LLMs) have demonstrated impressive reasoning capabilities across general domains, their effectiveness in re…
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants
Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when…
MMGraphRAG: Bridging Vision and Language with Interpretable Multimodal Knowledge Graphs
Large Language Models (LLMs) suffer from hallucinations due to their static parametric knowledge. Retrieval-Augmented Generation (RAG) and…
AI sustains higher strategic tension than humans in chess
Strategic decision-making requires balancing immediate opportunities against long-term objectives: a tension fundamental to competitive env…
Benchmarking Agentic Newswriting via Journalistic Workflows
Recent advances in autonomous digital agents from industry (e.g., Manus AI and Gemini's research mode) highlight their potential for struct…
SATQuest: A Verifier for Logical Reasoning Evaluation and Reinforcement Fine-Tuning of LLMs
Large language models (LLMs) exhibit strong general reasoning, yet the community lacks controllable, scalable, and verifiable tools to anal…
Artificially intelligent agents in the social and behavioral sciences: A history and outlook
We review the historical development and current trends of artificially intelligent agents (agentic AI) in the social and behavioral scienc…
Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions
Despite rapid progress in artificial intelligence, current systems struggle with the interconnected challenges that define real-world decis…
Parallel Decoder Transformer: Planner-Conditioned Latent Coordination for Model-Intrinsic Parallel Generation
Autoregressive language models expose one causal token frontier, even when the requested document contains sections that could be developed…
Towards AI epidemiology: a measurement standardisation framework for prospective risk detection
This paper proposes a measurement standardisation framework that compresses expert-AI interactions into structured, comparable fields for p…
Multi-modal cross-domain mixed fusion model with dual disentanglement for fault diagnosis under unseen working conditions
Intelligent fault diagnosis has become an indispensable technique for ensuring machinery reliability. However, existing methods suffer sign…
From Classical to Quantum Reinforcement Learning and Its Applications in Quantum Control: A Beginner's Tutorial
This tutorial is designed to make reinforcement learning (RL) more accessible to undergraduate students by offering clear, example-driven e…
Data-Efficient Curation for Multimodal Reasoning under Fixed Training Protocols
We study data curation for multimodal reasoning in a fixed-protocol fine-tuning regime, where the base model, optimizer, training schedule,…
NEMO: Execution-Aware Optimization Modeling via Autonomous Coding Agents
We present NEMO, a system that translates Natural-language descriptions of decision problems into formal Executable Mathematical Optimizati…
Lyapunov Stability-Aware Stackelberg Game for Low-Altitude Economy: A Control-Oriented Pruning-Based DRL Approach
With the rapid expansion of the low-altitude economy, Unmanned Aerial Vehicles (UAVs) serve as pivotal aerial base stations supporting dive…
Arbor: A Framework for Reliable Navigation of Critical Conversation Flows
Large language models struggle to maintain strict adherence to structured workflows in high-stakes domains such as healthcare triage. Monol…
SCA: Segment-Wise CoT Compression with Answer Alignment
Chain-of-thought (CoT) reasoning improves problem solving, but long think traces increase inference cost. Existing CoT compression methods…
Content Creation with Spillovers: An Incentive Design Approach
The rise of AI amplifies the economic phenomenon of \emph{positive spillovers}: when creators contribute content that can be reused and ada…
Deterministic Hallucination Detection in Medical VQA via Confidence-Evidence Bayesian Gain
Multimodal large language models (MLLMs) have shown strong potential for medical Visual Question Answering (VQA), yet they remain prone to…
CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs
Analogical reasoning tests a fundamental aspect of human cognition: mapping the relation from one pair of objects to another. Existing eval…
Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
As the focus in LLM-based coding shifts from static single-step code generation to multi-step agentic interaction with tools and environmen…
From Multi-Agent to Single-Agent: When Is Skill Distillation Beneficial?
Multi-agent systems (MAS) tackle complex tasks by distributing expertise, though this often comes at the cost of heavy coordination overhea…
When Direct Prediction Fails: Evidence from LLM-Based Misinformation Risk Evaluation
LLMs make it increasingly easy to generate deceptive content at scale, creating a need for scalable misinformation risk evaluation based on…
Information-Theoretic Measures in AI: A Practical Decision Framework
Information-theoretic (IT) measures are ubiquitous in artificial intelligence: entropy drives decision-tree splits and uncertainty quantifi…
RADD: Retrieval-Augmented Discrete Diffusion for Multi-Modal Knowledge Graph Completion
Most multi-modal knowledge graph completion (MMKGC) models use one embedding scorer to conduct both retrieval over the full entity set and…
Adaptive Multi-Round Allocation with Stochastic Arrivals
We study a sequential resource allocation problem motivated by adaptive network recruitment, in which a limited budget of identical resourc…
AI for Auto-Research: Roadmap & User Guide
AI-assisted research is crossing a threshold: fully automated systems can now generate research papers for as little as $15, while long-hor…
Scientific reasoning does not reliably translate into scientific forecasting in frontier AI
AI systems are increasingly used to support forward-looking scientific judgment, but it remains unclear whether they can form reliable expe…
Structure-Induced Information for Rerooting Levin Tree Search
Subgoal-based policy tree search, which uses a policy to guide search, is effective for complex single-agent deterministic problems but oft…
Choosing the Lens: Strategic Perspective Activation in Context-Dependent Argumentation
The same arguments often need to be evaluated under different external regimes. An agent with influence over the regime has a strategic lev…
Regularized Offline Policy Optimization with Posterior Hybrid Bayesian Belief
Offline reinforcement learning (RL) aims to optimize policies from pre-collected datasets. A bottleneck of this paradigm is managing episte…
Physics-Guided Spatiotemporal Learning for Coastal Wave Peak Period Estimation from Video
Direct estimation of physically interpretable periodic signals from raw video constitutes a spatiotemporally grounded learning problem that…
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual gro…
Heteroskedastic Signals in Budgeted LLM Verification: Structural Heterogeneity Limits Optimization Gains
Selective-compute LLM systems decide which outputs merit verification, additional reasoning, tool execution, or human audit under a limited…
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their…
Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law
By promoting vectors to spheres and enabling explicit model construction, neural networks can perform symbolic-level syllogistic reasoning…
Theoria: Rewrite-Acceptability Verification over Informal Reasoning States
When should an AI system's answer be trusted? Formal proof assistants offer certainty but cannot reach most of the problem distribution; sc…
Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure
Multimodal GUI agents read an interface through two redundant channels: the rendered pixels of a screenshot and a serialized structure such…
Memory in the Loop: In-Process Retrieval as Extended Working Memory for Language Agents
Language agents run a loop - observe, reason, act - but the memory they reason over sits outside it: a store queried at most once per turn.…
A Formalization of the Mean-Field Derivation of the Vlasov Equation
We formalize a research result in the Lean 4 proof assistant by having a mathematician direct an AI system, and frame the activity as a for…
IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation
Scientific ideation unfolds over multiple stages, including literature search, paper reading, tool use, claim checking, cross-paper synthes…
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code ex…
Visual Access Boundaries in Vision-Language Model Reasoning
Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear w…
Probabilistic Extension of Neuro-Symbolic AGI Robots based on Belnap's Typed Intensional FOL
Neuro-symbolic AI based on $IFOL_B$ is a way to combine neural learning and symbolic reasoning to overcome limitations of purely neural sys…
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, cu…
When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models
Large language models can synthesize a game's rules as executable code - a Code World Model (CWM) - which a classical planner then searches…
SportD: Can VLMs Physically Strategize?
Vision--language models have become increasingly capable of interpreting visual scenes, but it remains unclear whether they can use informa…
SmartRAG: Native Graph-Based RAG for Mobile Device
Deploying large language models (LLMs) as personal assistants on mobile devices demands privacy, low latency, and offline availability, yet…
Global Index on Responsible AI: 2026 Report
Grounded in human rights-based frameworks such as the UNESCO Recommendation on the Ethics of AI, the Global Index on Responsible AI (GIRAI)…
Automated Reinforcement Learning: An Overview
Reinforcement Learning and, recently, Deep Reinforcement Learning are popular methods for solving sequential decision-making problems model…
CarbonNet: How Computer Vision Plays a Role in Climate Change? Application: Learning Geomechanics from Subsurface Geometry of CCS to Mitigate Global Warming
We introduce a new approach using computer vision to predict the land surface displacement from subsurface geometry images for Carbon Captu…
Unsupervised Multimodal Clustering for Semantics Discovery in Multimodal Utterances
Discovering the semantics of multimodal utterances is essential for understanding human language and enhancing human-machine interactions.…
Posts of Peril: Detecting Information About Hazards in Text
Socio-linguistic indicators of affectively-relevant phenomena, such as emotion or sentiment, are often extracted from text to better unders…
Lost in Transmission: An Information-Theoretic Account of Unsupervised Software Traceability
Traceability remains a critical capability to ensure system reliability, maintainability, and compliance in modern software development. Al…
Mediator: Memory-efficient LLM Merging with Less Parameter Conflicts and Uncertainty Based Routing
Model merging aggregates Large Language Models (LLMs) finetuned on different tasks into a stronger one. However, parameter conflicts betwee…
A Survey on Knowledge-Oriented Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) has gained significant attention in recent years for its potential to enhance natural language underst…
A Survey on Unlearnable Data
Unlearnable data (ULD) has emerged as an innovative defense technique to prevent machine learning models from learning meaningful patterns…
OMAC: A Holistic Optimization Framework for LLM-Based Multi-Agent Collaboration
Agents powered by advanced large language models (LLMs) have demonstrated impressive capabilities across diverse complex applications. Rece…
Data Balancing Strategies: A Systematic Survey of Resampling and Augmentation Methods
Imbalanced datasets, where one class significantly outnumbers others, remain a persistent challenge in machine learning, often biasing pred…
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning
Effective generalization in language models depends critically on the diversity of their training data. Yet existing diversity metrics ofte…
Learning MMSE Filters for OFDM Channel Estimation: Attention Transformer Gains at Linear Inference
In orthogonal frequency division multiplexing (OFDM), accurate channel estimation is crucial. Classical signal processing-based approaches,…
OV-MAP: Open-Vocabulary Zero-Shot 3D Instance Segmentation Map for Robots
We introduce OV-MAP, a novel approach to open-world 3D mapping for mobile robots by integrating open-features into 3D maps to enhance objec…
Sequential Attention-based Sampling for Histopathological Analysis
Deep neural networks are increasingly applied in automated histopathology. Yet, whole-slide images (WSIs) are often acquired at gigapixel s…
Can Interpretation Predict Behavior on Unseen Data?
Interpretability research often predicts model responses to targeted mechanistic interventions. But can we predict responses to unseen inpu…
Attentions Under the Microscope: A Comparative Study of Resource Utilization for Variants of Self-Attention
As large language models (LLMs) and visual language models (VLMs) grow in scale and application, attention mechanisms have become a central…
Symmetric Behavior Regularized Policy Optimization
Behavior Regularized Policy Optimization (BRPO) leverages asymmetric divergence regularization to mitigate distribution shift in offline re…
DCSCR: A Class-Specific Collaborative Representation based Network for Image Set Classification
Image set classification (ISC), which can be viewed as a task of comparing similarities between sets consisting of unordered heterogeneous…
"Not in My Backyard": LLMs Uncover Online and Offline Social Biases Against Homelessness
Homelessness is a persistent social challenge, impacting millions worldwide. Over 876,000 people experiencing homelessness (PEH) were recor…
Is "Knowing It's Malicious Enough?" Evaluating LLMs for Fine-Grained Malware Behavior Auditing
Automated malware classifiers achieve strong detection performance, but auditing requires more than flagging a sample: analysts must explai…
From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Retrieval-Augmented Generation Agents Development
Retrieval-augmented generation (RAG) agent development is hindered by the lack of executable ground-truth agent-environment interaction tra…
STAC: When Innocent Tools Form Dangerous Chains for LLM Agents
As LLMs advance into autonomous agents with tool-use capabilities, they introduce security challenges that extend beyond traditional conten…
RobustVLA: On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
In Vision-Language-Actionf(VLA) models, robustness to real-world perturbations is critical for deployment. Existing methods target simple v…
Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering
Vision-language models (VLMs) demonstrate strong image-level scene understanding, but reasoning over long egocentric video remains costly:…
Probing the Difficulty Perception Mechanism of Large Language Models
Large language models (LLMs) are increasingly deployed on complex reasoning tasks, yet little is known about their ability to internally ev…
When Fewer Layers Break More Chains: Layer Pruning Harms Test-Time Scaling in LLMs
Layer pruning has emerged as a widely adopted technique for improving the efficiency of large language models (LLMs). Although existing met…
BBOPlace-Bench: Benchmarking Black-Box Optimization for Chip Placement
Chip placement is a vital stage in modern chip design, and black-box optimization (BBO) has been applied to it for decades. Early BBO effor…
InertialAR: Autoregressive 3D Molecule Generation with Inertial Frames
Transformer-based autoregressive models have emerged as a unifying paradigm across modalities such as text and images, but their extension…
CORE -- A Cell-Level Coarse-to-Fine Image Registration Engine for Multi-stain Image Alignment
Accurate and efficient registration of whole slide images (WSIs) is essential for high-resolution, nuclei-level analysis in multi-stained t…
ProDER: A Continual Learning Approach for Fault Prediction in Evolving Smart Grids
As smart grids evolve to meet growing energy demands and modern operational challenges, the ability to accurately predict faults becomes in…
GRIP: In-Parameter Graph Reasoning through Fine-Tuning Large Language Models
Large Language Models (LLMs) have demonstrated remarkable capabilities in modeling sequential textual data and generalizing across diverse…
DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks
Vision-Language Models (VLMs) show great promise for autonomous driving, but their suitability for safety-critical scenarios is largely une…
BUSTR: Descriptor-Aware Vision-Language Learning for Breast Ultrasound Report Generation
Breast ultrasound (BUS) reporting relies on clinically meaningful lesion descriptors, including BI-RADS category, lesion shape, margin, ech…
SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection
Deepfake (DF) audio detectors still struggle to generalize to out of distribution inputs. A central reason is spectral bias, the tendency o…
When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
Frontier language models increasingly participate in conversations about distress and mental health, yet the mechanisms that generate anthr…
Computing Evolutionarily Stable Strategies in Imperfect-Information Games
We present an algorithm for computing evolutionarily stable strategies (ESSs) in symmetric perfect-recall extensive-form games of imperfect…
mHC-GNN: Manifold-Constrained Hyper-Connections for Graph Neural Networks
Graph Neural Networks (GNNs) suffer from over-smoothing in deep architectures and expressiveness bounded by the 1-Weisfeiler-Leman (1-WL) t…
Lil: Less is Less When Applying Post-Training Sparse-Attention Algorithms in Long-Decode Stage
Large language models (LLMs) demonstrate strong capabilities across a wide range of complex tasks and are increasingly deployed at scale, p…
ReMIND: Orchestrating Modular Large Language Models for Controllable Serendipity A REM-Inspired System Design for Emergent Creative Ideation
Large language models (LLMs) are increasingly used not only for problem solving but also for creative ideation; however, generating ideas t…
Sparsity-Aware Low-Rank Representation for Efficient Fine-Tuning of Large Language Models
Adapting large pre-trained language models to downstream tasks often entails fine-tuning millions of parameters or deploying costly dense w…
Hybrid Mamba-Attention Neural Architecture for Channel Estimation
This paper proposes a hybrid Mamba-attention neural architecture to achieve improved channel estimation for orthogonal frequency-division m…
Li-ViP3D++: Query-Gated Deformable Camera-LiDAR Fusion for End-to-End Perception and Trajectory Prediction
End-to-end perception and trajectory prediction from raw sensor data is one of the key capabilities for autonomous driving. Modular pipelin…
CAM: A Causality-based Analysis Framework for Multi-Agent Code Generation Systems
Despite the remarkable success that Multi-Agent Code Generation Systems (MACGS) have achieved, the inherent complexity of multi-agent archi…
SoMA: A Real-to-Sim Neural Simulator for Robotic Soft-body Manipulation
Simulating deformable objects under rich interactions remains a fundamental challenge for real-to-sim robot manipulation, with dynamics joi…
Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative Signals
Evaluating mathematical reasoning in LLMs is constrained by limited benchmark sizes and inherent model stochasticity, yielding high-varianc…
UnMaskFork: Test-Time Scaling for Masked Diffusion via Deterministic Action Branching
Test-time scaling strategies have effectively leveraged inference-time compute to enhance the reasoning abilities of Autoregressive Large L…
Thermodynamic Limits of Physical Intelligence
Modern AI systems achieve remarkable capabilities at the cost of substantial energy consumption. To connect intelligence to physical effici…
GradAlign: Gradient-Aligned Data Selection for LLM Reinforcement Learning
Reinforcement learning (RL) has become a central post-training paradigm for large language models (LLMs), but its performance is highly sen…
Long Range Frequency Tuning for QML
Angle-encoded variational quantum circuits admit a truncated Fourier series representation of their output, but approximating functions wit…
Breaking the Factorization Barrier in Diffusion Language Models
Diffusion language models theoretically allow for efficient parallel generation but are practically hindered by the ``factorization barrier…
Rooted Absorbed Prefix Trajectory Balance with Submodular Replay for GFlowNet Training
Generative Flow Networks (GFlowNets) enable fine-tuning large language models to approximate reward-proportional posteriors, but they remai…
IoUCert: Robustness Verification for Anchor-based Object Detectors
While formal robustness verification has seen significant success in image classification, scaling these guarantees to object detection rem…
No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason
Generative AI can convert uncertainty into authoritative-seeming verdicts, displacing the justificatory work on which democratic epistemic…
Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models
Quantization has become essential for the efficient deployment of speech processing systems. Although widely studied, most existing quantiz…
L2GTX: From Local to Global Time Series Explanations
Deep learning models achieve high accuracy in time series classification, yet understanding their class-level decision behaviour remains ch…
FormulaCode: Evaluating Agentic Optimization on Large Codebases
Large language model (LLM) coding agents increasingly operate at the repository level, motivating benchmarks that evaluate their ability to…
NanoZK: Privacy-Preserving Verifiable Inference for Large Language Models via Layerwise Zero-Knowledge Proofs
We present NanoZK, a zero-knowledge proof system for verifiable LLM inference: clients and third-party auditors check that a provider execu…
HiCI: Hierarchical Construction-Integration for Long-Context Attention
Long-context language modeling is commonly framed as a scalability challenge of token-level attention, yet local-to-global information stru…
Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
Recent advances in large language models have improved the capabilities of coding agents, yet systematic evaluation of complex, end-to-end…
SyriSign: A Parallel Corpus for Arabic Text to Syrian Arabic Sign Language Translation
Sign language is the primary approach of communication for the Deaf and Hard-of-Hearing (DHH) community. While there are numerous benchmark…
Neural Global Optimization via Iterative Refinement from Noisy Samples
Global optimization of black-box functions from noisy samples is a fundamental challenge in machine learning and scientific computing. Trad…
ClawBench: Can AI Agents Complete Everyday Online Tasks?
AI agents may be able to assist with emails and documents, but can they reliably complete everyday online workflows on real websites? Every…
BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation
Accurate evaluation is central to the large language model (LLM) ecosystem, guiding model selection and downstream adoption across diverse…
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering
Medical vision--language models (VLMs) have shown strong potential for medical visual question answering (VQA), yet their reasoning remains…
DIB-OD: Preserving the Invariant Core for Robust Heterogeneous Graph Adaptation via Decoupled Information Bottleneck and Online Distillation
Graph Neural Network pretraining is pivotal for leveraging unlabeled graph data. However, generalizing across heterogeneous domains remains…
FETS Benchmark: Foundation Models Enable Scalable and Generalizable Energy Time Series Forecasting
Driven by the transition towards a climate-neutral energy system, accurate energy time series forecasting is critical for planning and oper…
Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference
Expressway video anomaly detection is essential for safety management. However, identifying anomalies across diverse scenes remains challen…
A Systematic Investigation of RL-Jailbreaking in LLMs
The evolution of generative models from next-token predictors to autonomous engines of complex systems necessitates rigorous safety hardeni…
A Benchmark for Early-stage Parkinson's Disease Detection from Speech
Early-stage Parkinson's disease (EarlyPD) detection from speech is clinically meaningful yet underexplored, and published results are hard…
Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict
Retrieval-Augmented Generation (RAG) is usually evaluated by whether the final answer is correct. Under knowledge conflict, this hides a ke…
When Cultures Move: Measuring and Improving Multicultural Text-to-Video Generation
Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within…
SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation
Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation…
Every Component is a Lookup: Token Attribution and Composition from a Single Decomposition
Mechanistic interpretability of transformers requires identifying not just which components matter but how they compose into the computatio…
Do We Really Need Quantum Machine Learning?: A Multidimensional Empirical Study
The rapid growth of computer vision and increasingly complex image recognition tasks has exposed fundamental computational limitations of c…
How Much Is a Dataset Worth? Scaling Laws, the Vendi Score, and Matrix Spectral Functions
Neural scaling laws appraise data through dataset size, while the Vendi Score uses quantum entropy to measure dataset value. We show both t…
Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth Optimization
Deep learning optimization relies heavily on the assumption of smooth loss landscapes, a condition systematically violated by modern archit…
Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics
Large language models (LLMs) are increasingly entering students' learning practices, but their educational value depends on whether they su…
CyberGym-E2E: Scalable Real-World Benchmark for AI Agents' End-to-End Cybersecurity Capabilities
AI has the potential to transform cybersecurity by enabling systems that can autonomously detect, analyze, and remediate software vulnerabi…
Improving Answer Extraction in Context-based Question Answering Systems Using LLMs
Question answering (QA) systems have achieved notable progress with the advent of large language models (LLMs). However, they still face ch…
TempoVLA: Learning Speed-Controllable Vision-Language-Action Policies
Robot manipulation alternates between low-risk transit phases that call for fast execution and high-risk contact stages that demand slow, p…
FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language Models
Diffusion Large Language Models (dLLMs) refine tokens iteratively but commit them irreversibly, leading to a "stability lag" where early de…
Phantom Transitions in Language Model Fine-Tuning: A Density-Matrix Analysis
Fine-tuning a language model often fails silently when its correct completion must outrank a near-synonym competitor. Cross-entropy loss fa…
FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In…
BiWM: Advancing Open-Source Interactive Video World Models with Bidirectional Autoregression
Interactive video world models commonly convert bidirectional video generators into causal autoregressive systems through control fine-tuni…
RoVE: Rotary Value Embeddings Attention for Relative Position-dependent Value Pathways
Rotary Position Embeddings (RoPE) make attention scores position-relative but leave the value pathway position-blind: the message sent by a…
Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack
In this report, we present Hy-Embodied-0.5-VLA, abbreviated as HyVLA-0.5, an end-to-end system that spans the full robot learning stack: da…
AI Contagion in Social Networks
We study how artificial intelligence (AI) interacts with social communication networks to shape the stability of collective knowledge. Agen…
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
Coding agents have become a major mode of software engineering, but the benchmarks we use to compare them were designed in a pre-agent era:…
GroundShot: Visually Consistent Multi-Shot Long Video Generation via Entity-Grounded Shot Scheduling
Generating visually consistent multi-shot videos remains an open challenge. As videos span more shots, inconsistencies can accumulate acros…
Short-Term Electricity Demand Forecasting for New England: A Comprehensive Machine Learning Benchmark with Weather, Calendar, and COVID-19 Indicators
Accurate short-term electricity demand forecasting is critical for reliable power system operation, energy market planning, and infrastruct…
Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One
A language model's memory can be worse than no memory at all when the model or its interface is disposed to act on it: a memory that keeps…
Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 team…
From Scene-Centric to Observer-Centric: Modeling Observer-Aware Relations for 3D Scene Graph Generation
3D Scene Graph Generation (3DSGG) represents 3D scenes as structured object--relation--object graphs for spatial understanding. In observer…
ReactiveBFM: Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control
While current Behavior Foundation Models (BFMs) provide robust control priors for humanoids, they only execute pre-defined reference motion…
TRIAGE: Role-Typed Credit Assignment for Agentic Reinforcement Learning
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands…
Full Bayesian Reinforcement Learning via LF-IBIS
Reinforcement Learning (RL) is a sequential decision-making framework in which an agent learns optimal policies through interaction with an…
The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains…
Beyond Multilingual Averages: MTEB-PT, a Benchmark for Portuguese Sentence Encoders
Portuguese remains underrepresented in text embedding evaluation, despite being one of the most widely spoken languages in the world. As a…
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Despite growing automation, turning a paper into a coherent poster, talk video, and blog piece often remains a labor-intensive last mile. R…
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications
Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in th…
Overview of the NLPCC 2026 Shared Task 1: Difficulty-Aware Multilingual and Multimodal Medical Instructional Video Understanding Evaluation
Following the CMIVQA, MMI-VQA, and M4IVQA challenges in NLPCC 2023--2025, we introduce the Difficulty-Aware Medical Instructional Video Que…
LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting
Process industries rely on time-series forecasting and soft sensing to estimate quality variables that are hard to measure online. Labeled…
Trusting sovereign language models as scientific instruments: evidence from Portugal's AMALIA
National language models are becoming publicly funded epistemic infrastructure. Public ownership, linguistic specialization, and open weigh…
EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins
Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches le…
A small language model detects behavioural faithfulness gaps that frontier judges and human raters miss
Whether a language model behaves as it claims is a judgement on which independent human raters cannot agree (Fleiss kappa = 0.074). We show…
Scalable Visual Pretraining for Language Intelligence
The rapid progress of large foundation models has been driven predominantly by pretraining on large-scale text corpora. However, many forms…
LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning
Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each…
Adaptive Compute in Latent World Models: When Depth Helps, Hurts, or Doesn't Matter
Adaptive compute for world models -- early-exit or mixture-of-depths predictors that spend variable depth per rollout step -- presumes that…
Exact and Certified Data Shapley for Weighted k-Nearest-Neighbor Regression and Soft-Label Prediction
Data Shapley answers which training points are worth what, and its nearest-neighbor specialization is the version actually deployed, shippe…
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget
We introduce Boogu-Image-0.1, an open-source unified multimodal understanding and generation model family, comprising Base, Turbo, Edit, an…
Tabular Foundation Models for Discrete Choice Estimation
Tabular foundation models (TFMs) generate predictions on structured data via in-context learning, without task-specific estimation. We ask…
GHR-VLM: Making Zero-Shot Transit Video Analytics Realizable with Grounded Hybrid Reasoning
Transit video understanding can provide valuable fine-grained data that conventional passenger counters and fare systems cannot capture. Ho…
Semantic Anchoring for Robotic Action Representations
Vision-Language-Action (VLA) models inherit rich semantic representations from pretrained Vision-Language Models, yet fine-tuning on limite…
How Agents Ask for Permission: User Permissions for AI Agents, from Interfaces to Enforcement
As AI agents gain prevalence, users are increasingly exposed to the risks such systems entail. Prompt injection attacks, as well as halluci…
AI-Augmented Human Resource Management? Insights from German companies
This study examines the integration of AI into Human Resource Management in German companies. We ask if and how AI-based technologies are \…
Decision Making Needs Uncertainty Quantification [Lecture Notes]
Many signal processing systems ultimately exist to {act}. Whenever the state variable that determines the action to be taken by a decision…
Harnessing LLMs for Reliable Academic Supervision: A Comparative Study
Large language models routinely produce fluent answers to single-shot prompts, yet deploying them as reliable components of a domain decisi…
VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient M…
Show Me How You Reason and I'll Tell You Who You Are: Reasoning Graphs for Robust LLM Authorship Attribution
Given the current trend to employ large language models (LLMs) in almost any imaginable context, LLM-generated text detection and authorshi…
GMO熊谷氏「在宅勤務廃止」発言を釈明 「作業はAIに任せ、人はオフィスに」
GMOインターネットグループの熊谷正寿会長が、在宅勤務の「グループ推奨」廃止を巡る自身の発言を釈明した。在宅勤務そのものの否定ではなく、趣旨は「AI時代におけるオフィス価値の再定義」だと説明している。
「国産AIスタック」を持たないカナダ、ソブリンAIをどう構築するのか
カナダには、AI分野をリードするための人材と研究基盤が備わっている。ただし、国内で自律的なコンピューティング、データセンター、エッジAIインフラを構築するには、戦略的パートナーと連携しつつ、迅速に動き出さなければならない。
顧客の反応、意思決定にどう反映させる? Zoomの取り組みから「AI×CX」の進化を探る
企業のCXはAIでどう進化していくのか。その中で人が果たすべき役割は何か。AIで成果を生み出すための「CXの4つのステップ」を提唱する、Zoomの取り組みから探る。
OpenAI、自律型AIが安全対策を回避する行動を学習する可能性を確認 内部展開を一時停止
OpenAIは、長時間自律動作するAIモデルの安全性評価に関するブログを公開した。限定運用でサンドボックスの回避や認証トークンの難読化などの問題行動を確認したため、一時アクセスを停止。モデルの行動全体を監視する新たな安全対策を構築し、問題行動の検出・抑止を確認した上で内部利用を…
Anthropic’s landmark $1.5B copyright settlement is approved
The final approval settles one case, but it doesn't resolve the broader issue of using copyrighted works to train AI models.
AMDとMicrosoftが戦略的提携を拡大 新AIラックスケール「Helios」をAzureに大規模導入へ
AMDは、Microsoftとの戦略的提携を拡大すると発表した。Microsoftはクラウドサービス「Azure」に、GPUやCPUを一体化したAMDのラックスケール製品「AMD Helios」を大規模に導入し、フロンティアAIモデルの推論処理などに活用する。Heliosは20…
Trump’s latest AI czar has already resigned
The director role for the Center for AI Standards and Innovation (CAISI) has become a revolving door since David Sacks left his position as…
中外製薬「社員1人にAIエージェント10体」作戦で成果倍増を目指す、AI使いこなし術
製薬はコストも高く成功率も低い苛烈な業界だ。一方で、AIの活用により費用を1200億から半減、成功率を10倍にできるという試算もある。中外製薬はそのような業界の中で、AIにより2030年に研究開発の成果を倍増するという計画を掲げた。その秘策とは。
NTT、ソフトバンク、サカナAI――国産AI開発「成功組」の“ある共通点”
NTTやソフトバンク、サカナAIなどの「AI開発に成功した企業」には“ある共通点”がある。彼らはどのような技術を活用し、AI開発を成し遂げたのか。
「AIに期待」65%も「明確な成果」16%、製造業の多くがPoC止まりの理由は
primeNumberが「AI・データ活用実態調査 2026」を公開した。AIへの高い期待に対し、明確な成果を得た企業は16.4%にとどまる。成果を左右する原因や、製造業でPoC止まりが多発する構造的課題を同社に聞いた。
Google is working on a new AI chip designed to make Gemini more efficient
Alphabet, Google's parent company, is reportedly working on a new chip designed to make its Gemini models run much more efficiently.
AI’s most important protocol is getting a little bit easier to use
Under the new system, the protocol will take a looser, "stateless" approach to session IDs on the server side, similar to how most ordinary…
X relaunches a rebuilt Android app after year-long effort
X says the rebuilt version of its Android app is now available globally.
OpenAI is scared of open-weight models. Should the US be?
Talk of banning Chinese-made open-weight LLMs reveals the challenge of turning AI into a business.
Adobe camera app’s new feature will critique your photos using AI
Adobe's Project Indigo can now remove all kinds of backgrounds from photos you snap using the app.
YouTube clarifies policies around AI slop and upsetting videos
YouTube has updated its monetization policies to more clearly define the kinds of AI-generated and low-quality videos that can’t earn ad re…
2026-07-20(205件)
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards thro…
GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis
Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering. Existing Large…
Causal-Audit: Explicit and Auditable Graph-based Reasoning via Target-Aware Causal Chain Construction
Causal and intervention-based question answering is fundamental to advancing large language models (LLMs) toward reasoning beyond surface-l…
Cura 1T: Specialized Model for Agentic Healthcare
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases toget…
AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery
Desktop voice assistants are still dominated by cloud pipelines that ship raw audio off the machine and expose a fixed set of skills. We de…
Precise but Uncoupled: Reviewer Precision Does Not Guarantee Critique Uptake in Multi-Agent Math Reasoning
Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review st…
DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings
We introduce DrawingVQA, the first benchmark designed to evaluate multimodal large language models (MLLMs) on real-world construction drawi…
Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
Our previous ARC-AGI-3 agent bundled executable world modeling, scheduled simplification, and exact replay verification, leaving unclear wh…
Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes
Internet memes intertwine visual cues, textual content, and cultural context, making them particularly challenging to interpret in scenario…
From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems
A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable…
A Critical Analysis of Trustworthy AI Tools, Mark Frameworks, and the Implementation Chasms
As artificial intelligence (AI) systems increasingly impact society, ensuring their ethical and trustworthy deployment has become a global…
Logic, Optimization, and Artificial Intelligence
Logic and optimization can, in combination, make valuable contributions to rule-based AI. Logic is the obvious medium for encoding a rule b…
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce crit…
MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion
Multimodal Knowledge Graph Completion (MKGC) requires inferring missing entities from structural, textual, and visual cues. Existing diffus…
Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts
LEED v4.1 BD+C certification remains a document-intensive process that requires reviewers to read hundreds of pages of project evidence and…
ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning
While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and eff…
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) ha…
Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic co…
NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning
OWL ontologies provide a formal knowledge representation framework that enables semantic reasoning, and have been widely adopted across dom…
AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets
Geospatial datasets support applications from urban planning to climate modeling, yet consistent assessment of FAIR compliance is difficult…
Knowledge-Centric Agents for Workflow Generation
Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over mod…
DSWorld: A Data Science World Model for Efficient Autonomous Agents
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error…
A Formally Grounded ODRL Evaluator: Implementation and Comparison
The ODRL policy language is emerging as the de-facto standard for policy modelling data access and usage preferences, AI governance policie…
Closing the AI Trust Gap: The Case for Independent Certification for Trustworthy AI
Over the past decade, responsible AI (RAI) has produced a substantial body of practice for identifying and mitigating the risks AI poses in…
SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery
Scientific work increasingly spans heterogeneous artifacts -- papers, code, datasets, scientific file formats, model outputs, figures, manu…
Harmonizing AI Safety Thresholds
Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify wheth…
CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and prov…
Empathy as Predictive Misalignment Tolerance: A Co-Regulation Framework and the Regime Structure of Dialogue Repair
Empathy is most often theorized as resonance: a mirroring of another's present emotional or cognitive state. This synchronic framing has sh…
How Does Empowering Users with Greater System Control Affect News Filter Bubbles?
While recommendation systems enable users to find articles of interest, they can also create ``filter bubbles'' by presenting content that…
Structure of the Circular-Dyadic Convolution Error
Dyadic and circular convolution can both be computed in $O(N\log N)$ time using the Hadamard transform and the FFT-computed discrete Fourie…
AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
We present AV-JEPA, an elegant multimodal extension of LeJEPA to audio-visual self-supervised learning. Using an early-fusion Vision Transf…
Data-driven Video Codec with Implicit Neural Representations
A conventional codec stores a video as compressed pixel data. We instead store the video, together with its audio track, as the weights of…
Lazy Arithmetic using Systolic Arrays for Closing the Verification Gap on Embedded Systems
Complex algorithms such as deep neural networks are increasingly being deployed on embedded, resource constrained platforms. However, exist…
Large Language Models as Unified Multimodal Learners for Clinical Prediction
Electronic health records combine free-text clinical narratives with structured measurements such as vital signs, laboratory values, and co…
Partial Information Decomposition as a Multi-Contrast 3D MRI Selection Strategy for Resource-Constrained Deep Neural Network Training in Brain Tumor Segmentation
Multi-contrast 3D MRI segmentation can be computationally demanding when all available sequences are used. We evaluate a pre-training Parti…
AI Trading: Evaluating Large Language Models for Technical Market Analysis
Large Language Models (LLMs) have emerged as powerful tools for processing the heterogeneous information environments of modern financial m…
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outc…
Design-Based Supervised Learning with Noisy Human Labels
Researchers increasingly use automated classifiers to label unstructured data for statistical analysis. Existing rectification methods can…
FLINT: Fingerprinting Federated Learning Architectures from 5G PHY-Layer Side Channels
Federated Learning (FL) over 5G cellular networks protects raw data but remains vulnerable to side-channel leakage. Prior fingerprinting at…
Verbalizable Representations Form a Global Workspace in Language Models
Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal re…
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture de…
An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism
Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging becaus…
Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching
Production LLM deployments combine two cost-reduction primitives: prompt caching (a discounted rate for re-used token prefixes) and prompt…
SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification
Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity…
Recursive Harness Self-Improvement
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces…
Kolmogorov--Arnold Networks for Small Language Models
Kolmogorov--Arnold Networks (KANs) replace fixed node activations with learned one-dimensional edge functions, offering an explicit interfa…
CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Agent Science Collaboration
LLM-based agents excel at writing articles, coding and information retrieval. However, they fail to form strong collaborations within the s…
From Feasibility to Desirability: Plan, Learn, Adapt (PLA) Framework for Personalized On-Device Itinerary Generation
Generating personalized trip itineraries is a complex planning task and involves a tension between hard combinatorial feasibility and soft…
Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design
Physics-informed neural networks (PINNs) are unusually sensitive to interacting choices of architecture, activation, loss weighting, colloc…
Hard Rules, Soft Preferences: Bridging Reasoning, Learning, and Optimization for Personalized Packing Checklist Generation
Packing for air travel is recurring and error-prone: the checklist must be personal and context-aware, yet feasible under safety rules, ite…
Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models
Where should the question go in a vision-language model (VLM) prompt: before the image or after it? Intuition says before: knowing what is…
Information-Directed Sampling for Causal Bandits
Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification…
MemoGuard: An Adaptive Runtime for Guarding Against Memory Traps in Communication-Limited Robot Navigation
Communication-limited robots in mission-critical scenarios such as disaster inspection and search-and-rescue must make reliable onboard dec…
Field-Aware RankMixer with Dual-Stream Bilinear Fusion for the Tencent UNI-REC Challenge
This paper presents our solution to the KDD Cup 2026 Tencent UNIREC Challenge. The task requires joint modeling of multi-domain user behavi…
Scalable LLM Agent Tool Access in the Cloud
LLM agents increasingly rely on tool calling to act on external systems, and the Model Context Protocol (MCP) has quickly become its de fac…
Process Reward Informed Tree Rollout for Effective Multi-Turn RL
Reinforcement learning (RL) has become a key approach for training LLM agents, yet popular methods such as GRPO/RLOO rely on multiple indep…
AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots
Self-driving laboratories increasingly rely on low-cost liquid handlers such as the Opentrons OT-2, which ship without the pressure-based a…
Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving
Large language models bring instruction following and scene reasoning to end-to-end driving, but their inference latency collides with the…
A cubical formalisation of topos causal models: intervention, sheaf gluing, and the intuitionistic do-calculus
Topos causal models recast causal inference inside a topos: a causal world is a presheaf, an intervention is a characteristic map into the…
IMBench: A Benchmark for Intuitive Robotic Manipulation
Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of th…
On the Structure of Address in Multi-Party Dialogue: From Discrete Labels to Continuous Levels
In multi-party dialogues between a dialogue system and multiple users, identifying to whom an utterance is addressed is a key challenge. Pr…
Toward a mechanistic understanding of inference in visual cortex and diffusion models
We describe a model of perceptual inference in primary visual cortex (V1) equivalent to a minimal diffusion model whose function can be rea…
Efficient Difficulty-Aware Dynamic Routing for Diffusion-Based Real-World Image Super-Resolution
Diffusion-based methods have achieved impressive performance in real-world image super-resolution (Real-ISR) by leveraging large pre-traine…
Map as a Prompt: Learning Multi-Modal Spatial-Signal Foundation Models for Cross-scenario Wireless Localization
Accurate and robust wireless localization is a critical enabler for emerging 5G/6G applications, including autonomous driving, extended rea…
Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling
As Text-to-Image (T2I) systems rapidly advance, evaluating the cultural authenticity of synthesized content has become increasingly importa…
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Conversational Speech Synthesis (CSS) aims to synthesize speech with human-like emotional expression and contextual consistency in user-age…
GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing
Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not…
Scaling Time Series Classification via XAI-Driven Data Reduction
Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for…
AquaAugmentor: A Novel Feature Augmentation Algorithm for Water Potability Prediction
Access to potable water is crucial for health, economic development, and sustainability. However, accurately classifying water quality rema…
Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding
Video Large Language Models (Video LLMs) have made significant advancements in various video understanding tasks. However, long-video scena…
On the Geometry of Learned Representations in Event-Based Multi-Modal Egomotion Estimation
Classical approaches to event-based egomotion estimation, including those adopted by the top-performing teams of the ELOPE challenge, rely…
Knowledge-Assisted Multi-Graph Dependency Learning for Multivariate Time Series Anomaly Detection in Multi-Stage Industrial Processes
Industrial processes often generate complex, interdependent time-series data from multiple sensors across multiple stages, forming complex…
In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention
In-context learning is a remarkable property of transformers and has recently received a lot of interest. In many studies of in-context lea…
RTL-Sequencer: Towards Scalable RTL Timing Prediction with the Sequence-based Paradigm
Accurate timing prediction at the register-transfer level (RTL) is a longstanding challenge in design automation. Existing graph-based meth…
CAMMAR: Culture-Aware Matryoshka for Metaphorical Arabic Representations
Metaphor in Arabic is a culturally grounded mechanism for constructing meaning, encoding cultural knowledge that shapes interpretation. Yet…
Test-Time Noise Guided Adaptation for Realistic Autoregressive Video Generation
Autoregressive video diffusion models have enabled the generation of arbitrarily long videos by removing conditioning on future frames, thu…
Agentic Synthesis against Counterexample-Supplemented Sketches
Coding agents can fix a failing example without preserving the domain rule that made it fail, so later generations can repeat the same plau…
Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection
Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang,…
EgoExoMoCap: Distributed Ego-Exo Human Motion Capture
Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which…
DECODEM: Data Extraction from Corporate Organizational Documents via Enhanced Methods
Much empirical legal research depends on translating unstructured text into structured variables. In corporate governance research as elsew…
Perceived AGI: Believability as Dimensional Completeness, Not Capability
Large language models are broadly capable, yet in sustained one-to-one conversation they still read as flat: competent, responsive, and som…
Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models
While the internal mechanisms of autoregressive (AR) transformers have been studied extensively, much less is known about diffusion languag…
Orbis 2: A Hierarchical World Model for Driving
Current world models operate at a single level of abstraction, with most prioritizing perceptual fidelity while lacking the spatial reasoni…
On the Failure of Boundary-Seeking Distillation in Bottlenecked Generative Architectures
Data-free knowledge distillation transfers the knowledge encoded in a teacher model to a student model without access to the original train…
When Not to Automate: A Formal Protocol for Human Preservation in AI-Optimized Organizations
Standard automation ROI misses four categories of systemic risk -- tacit knowledge erosion, resilience reduction, regulatory exposure, and…
Sociocultural Influences on Opinion Formation: Word of Mouth Dynamics, Mass Media and Behavioural Development
We study a society of agents belonging to a number of occupational or cultural groups that form opinions about others' situation in the sam…
Robustness of Reinforcement Learning-Based Congestion Management in Low-Voltage Grids
Increases in photovoltaic generation, charging of electric vehicles and heat-pump demand challenge operating limits in low-voltage distribu…
DPNeXt: A Lightweight Multi-Scale Feature Fusion Framework for Efficient ViT-Based Multi-Task Dense Prediction
Multi-Task Learning (MTL) in robotics perception systems supports comprehensive 3D spatial scene understanding by integrating semantic segm…
Candidate Attended Dialogue State Tracking Using BERT
Dialogue state tracking (DST) is one of the core components in task-oriented dialogue systems. At each turn in a conversation, DST estimate…
Rethinking Quantum Continual Learning with Quantum Fisher Information
Quantum continual learning aims to train quantum models on sequential tasks without losing previously learned knowledge. However, variation…
Revisiting data-driven dynamic security assessment with a tabular foundation model
Data-driven pre-fault dynamic security assessment (DSA) rapidly evaluates the dynamic risk of credible contingencies on a power system usin…
Loop the Loopies!
We present Loopie, the most powerful looped Transformer to date. The Loopie series consists of two Mixture-of-Experts (MoE) models: a 20B-p…
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such…
When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis
Model merging is promoted as a substitute for joint multi-task training, yet in the reinforcement-learning setting this substitution is ess…
Spatial Normalization for Cross-Domain Retinal Layer Segmentation in Optical Coherence Tomography
Retinal layer segmentation in Optical Coherence Tomography (OCT) is a fundamental step for extracting quantitative biomarkers of retinal st…
LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization
Agentic Artificial Intelligence (AI), enabled by Large Language Models, marks a shift from rule-based automation toward autonomous, goal-dr…
JoyNexus: Service-Oriented Multi-Tenant Post-Training for VLA Models
The post-training of Vision-Language-Action (VLA) models is essential due to the diversity of simulators, robot embodiments, and task objec…
HCIG: A Hierarchical Cross-Modal Incongruity Graph Network for Multimodal Sarcasm and Cyberbullying Detection
Multimodal sarcasm and cyberbullying detection remain challenging because the intended meaning often emerges from incongruity between textu…
DADiff: Diffusion-Driven Cross-Domain Policy Adaptation for Reinforcement Learning
Transferring policies across domains poses a vital challenge in reinforcement learning, due to the dynamics mismatch between the source and…
Understanding Reasoning from Pretraining to Post-Training
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training i…
A Methodology for Auditable Trustworthiness Levels in AI Lifecycle Governance
AI governance increasingly requires judgments about whether an AI system remains adequately trustworthy over time, whether observed changes…
ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, i…
When Do Multi-Agent Systems Help? An Information Bottleneck Perspective
LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent s…
An Exam for Active Observers
Human vision is a closed loop: gaze is continuously redirected by intermediate hypotheses rather than a single snapshot. Decades of psychop…
When Does Muon Help Agentic Reinforcement Learning?
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We…
Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities
Connected and Autonomous Vehicles (CAVs) rely on interconnected software and hardware components, including sensors, Electronic Control Uni…
RAD: Retrieval High-quality Demonstrations to Enhance Decision-making
Offline reinforcement learning (RL) learns policies from fixed datasets, thereby avoiding costly or unsafe environment interactions. Howeve…
A Neuro-Symbolic Approach for Probabilistic Reasoning on Graph Data
Graph neural networks (GNNs) excel at predictive tasks on graph-structured data but often lack the ability to incorporate symbolic domain k…
Human-Aligned Procedural Level Generation Reinforcement Learning via Text-Level-Sketch Shared Representation
Human-aligned AI is a critical component of co-creativity, as it enables models to accurately interpret human intent and generate controlla…
RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs
The Structure Gap between probabilistic LLM generation and deterministic schema requirements hinders automated workflows. We propose RL-Str…
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
The black-box nature of Large Language Models necessitates novel evaluation frameworks that transcend surface-level performance metrics. Th…
The AI Fiction Paradox
AI development has a fiction dependency problem. Developers have treated large corpora of modern books, including fiction, as valuable enou…
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientifi…
FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment
In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring…
Towards a General Intelligence and Interface for Wearable Health Data
While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into…
Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields
Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benc…
Agents-K1: Towards Agent-native Knowledge Orchestration
Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Exist…
GA-VINO: A Geometry-Aware Variational Physics-informed Neural Operator for Mindlin-Reissner Plates
Plate and shell structures are widely used in engineering fields. Rapid response prediction for such structures under complex geometries, h…
UCOB: Learning to Utilize and Evolve Agentic Skills via Credit-Aware On-Policy Bidirectional Self-Distillation
Skill memories can improve agentic reinforcement learning by reusing past experience as textual guidance, but retrieved skills are not orac…
ACPO: Agent-Chained Policy Optimization for Multi-Agent Reinforcement Learning
Cooperative tasks in Multi-Agent Reinforcement Learning (MARL) require agents to collectively maximize a shared return. Under the Centraliz…
MirrorCode: AI can rebuild entire programs from behavior alone
AI models are rapidly improving at autonomous coding, as shown by benchmark progress and one-off demonstrations such as AI implementing a C…
Internal Pluralism and the Limits of Pairwise Comparisons
Local pairwise comparisons are a standard tool for learning how people want decision rules to work, e.g., in participatory design or alignm…
Agent Step Value: Auditing Evaluator-Channel Reversals in Black-Box Agent Traces
Pooling, substituting, or reusing evaluator-derived step rewards assumes that their direction survives a change of evaluation channel. The…
OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice
The rapid integration of Large Vision-Language Models (VLMs) into critical infrastructure promises to revolutionize personalized healthcare…
Evidence-Aware MapReduce for Forkable Compute
Snapshot-backed sandboxes make branching cheap while leaving evidence dependence unchanged. Branches can reuse a model, prompt, repository,…
Length Penalties Make Chain-of-Thought Less Monitorable
Length-penalized reinforcement learning can shorten chain-of-thought reasoning while hiding an influence that drives the model's answer. In…
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general…
First-Order Modal Logic in HOL: Deep and Shallow Embeddings with Automated Faithfulness (Extended Preprint)
We extend, in Isabelle/HOL, the deep-and-shallow embedding methodology of our prior work from propositional to first-order modal logic (FML…
The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning
Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it en…
MaxSAT-Based Feedback for Guiding Vision-Language Models in Sudoku
Vision--Language Models (VLMs) have recently demonstrated promising performance on structured visual reasoning tasks, including grid-based…
Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents
Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-serv…
BrainPilot: Automating Brain Discovery with Agentic Research
Understanding the brain increasingly depends on integrating evidence across scales, modalities, and disciplines. Addressing a single resear…
Long-Context Fine-Tuning with Limited VRAM
Parameter-efficient fine-tuning reduces model and optimizer memory, but dense attention still makes long training sequences expensive. We c…
Can We Trust Item Response Theory for AI Evaluation?
AI benchmarks increasingly leverage item-level statistical models, particularly item response theory (IRT), to estimate model capabilities,…
Perception-Aligned AI Outputs: End-to-End Visual Prediction for Uncertainty Communication in Clinical Decision-Making
Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods rely on technical explan…
Why do CNNs excel at feature extraction? A mathematical explanation
Over the past decade deep learning has revolutionized the field of computer vision, with convolutional neural network models proving to be…
Decoupled Alignment for Robust Plug-and-Play Adaptation
We introduce a training-free safety enhancement method for aligning large language models (LLMs) without the need for supervised fine-tunin…
Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning
We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient…
CTC: The Composite Task Challenge for Cooperative Multi-Agent Reinforcement Learning
The critical role of division of labor (DOL) in enhancing cooperation is well-recognized in real-world applications. Consequently, many coo…
MAnchors: Memorization-Based Acceleration of Anchors via Rule Reuse and Transformation
Anchors is a popular local model-agnostic explanation technique whose applicability is limited by its computational inefficiency. To addres…
AuditVotes: Elevating Provable Defense for GNNs with Efficient Augmentation and Conditional Smoothing
Despite advancements in Graph Neural Networks (GNNs), adaptive attacks continue to challenge their robustness. Certified robustness via ran…
A Scaffolded GenAI Lab in Early Undergraduate CS: A Mixed-Methods, Multi-Course Evaluation
Background and Context. Generative AI (GenAI) tools are increasingly used in programming courses, but we have limited evidence about how br…
SLAC: Safe and Efficient Real-Robot Reinforcement Learning via Unsupervised Simulation Pre-Training
Building capable household and industrial robots requires mastering the control of versatile, high-degree-of-freedom (DoF) systems such as…
A Bit of Freedom Goes a Long Way: Classical and Quantum Algorithms for Reinforcement Learning under a Generative Model
We propose novel classical and quantum online algorithms for learning finite- and infinite-horizon Markov Decision Processes (MDPs). Our al…
Acoustic Imaging for UAV Detection: Dense Beamformed Energy Maps and U-Net SELD
We introduce a U-net model for 360{\deg} acoustic source localization formulated as a spherical semantic segmentation task. Rather than reg…
Unsupervised Deep Learning for Inverse Problems in Computed Tomography
Assume you encounter an inverse problem that shall be solved for a large number of data, but no ground-truth data is available. To emulate…
Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs
Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. W…
Poison to Detect: Detection of Targeted Overfitting in Federated Learning
Federated Learning (FL) enables collaborative model training among clients without centralising data, making it a widely adopted privacy-en…
A Systematic Study of Large Language Models for Task and Motion Planning With PDDLStream
While we know that large language models (LLMs) can solve some planning problems, we do not understand the extent of these capabilities for…
Are Heterogeneous Graph Neural Networks Truly Effective for Node Classification? A Causal Perspective
Graph neural networks (GNNs) have achieved remarkable success in node classification. Building on this progress, heterogeneous graph neural…
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
Large language models (LLMs) have been shown to internalize human-like biases during finetuning, yet the mechanisms by which these biases m…
Human-Inspired Neuro-Symbolic World Modeling and Logic Reasoning for Interpretable Safe UAV Landing Site Assessment
Reliable assessment of safe landing sites in unstructured environments is essential for deploying Unmanned Aerial Vehicles (UAVs) in real-w…
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
Diffusion language models (DLMs) have emerged as a promising alternative to autoregressive (AR) generation, yet their reliance on Transform…
3D Motion Perception of Binocular Vision Target with PID-CNN
This article trained a network for perceiving three-dimensional motion information of binocular vision target, which can provide real-time…
Hybrid coupling with operator inference and the overlapping Schwarz alternating method
This paper presents a novel hybrid approach for coupling subdomain-local non-intrusive Operator Inference (OpInf) reduced order models (ROM…
Energy-Efficient Federated Learning via Adaptive Encoder Freezing for MRI-to-CT Conversion: A Green AI-Guided Research
Federated Learning (FL) holds the potential to advance equality in health by enabling diverse institutions to collaboratively train deep le…
Latency-Response Theory Model: Evaluating Large Language Models via Response Accuracy and Chain-of-Thought Length
The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream application…
PASs-MoE: Mitigating Misaligned Co-drift among Router and Experts via Pathway Activation Subspaces for Continual Learning
Continual instruction tuning (CIT) requires multimodal large language models (MLLMs) to adapt to a stream of tasks without forgetting prior…
Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models
Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels. Prior work demonstrated this threat but relied…
Inelastic Constitutive Kolmogorov-Arnold Networks: A generalized framework for automated discovery of interpretable inelastic material models
A key problem of solid mechanics is the identification of the constitutive law of a material, that is, the relation between strain history…
Jailbreak Foundry: From Papers to Runnable Attacks for Reproducible Benchmarking
Jailbreak techniques for large language models (LLMs) evolve faster than benchmarks, making robustness estimates stale and difficult to com…
KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
Knowledge distillation (KD) is an essential technique to compress large language models (LLMs) into smaller ones. However, despite the dist…
Interaction-Aware Whole-Body Control for Compliant Object Transport
Cooperative object transport in unstructured environments remains challenging for assistive humanoids because strong, time-varying interact…
CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation
Generative models are increasingly used to augment medical imaging datasets for fairer AI, yet a key assumption often goes unexamined: that…
When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inferen…
Ruling Out to Rule In: Contrastive Hypothesis Retrieval for Medical Question Answering
Retrieval-augmented generation (RAG) grounds large language models in external medical knowledge, yet standard retrievers frequently surfac…
Evaluating LLM-Based 0-to-1 Software Generation in End-to-End CLI Tool Scenarios
The evolution of Large Language Models (LLMs) has catalyzed a paradigm shift towards intent-driven software development, where autonomous a…
LVSum: A Benchmark for Timestamp-Aware Long Video Summarization
Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal…
Robust Explanations for User Trust in Enterprise NLP Systems
Robust explanations are increasingly required for user trust in enterprise NLP, yet pre-deployment validation is difficult in the common ca…
Soft $Q(\lambda)$: A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augment…
What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers
When do transformers commit to a decision, and what prevents them from correcting it? We introduce prolepsis: a transformer commits early,…
Brain-CLIPLM: Semantic Compression for EEG-to-Text Decoding
Decoding natural language from non-invasive electroencephalography (EEG) remains constrained by low signal-to-noise ratio and limited infor…
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard…
Energy-based Transport for Amortized Bayesian Inference
We consider amortized Bayesian inference for nonlinear inverse problems using only samples from the joint distribution of parameters and ob…
Diagnosing Overhead in Dispatch Operations: Cross-architecture Observatory
AlltoAll dispatch is the dominant bottleneck of MoE expert parallelism, and the interconnect community has responded with four families of…
The Terminal Representation in Reinforcement Learning
Representation learning is a powerful tool for spatio-temporal abstraction within reinforcement learning (RL). Two well established approac…
memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations
Agent-memory frameworks -- mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, MemTensor -- each ship their own SDK, storage layout, and op…
AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such…
AuAu: A Benchmark for Auditing Authoritarian Alignment in Large Language Models
The worldwide rise of authoritarianism and the growing role of Large Language Models (LLMs) in users' everyday lives raise the question of…
GeoRouteNet: A Geometry-Aware Non-Autoregressive Neural Solver for the Euclidean Traveling Salesman Problem
Non-autoregressive neural solvers amortize computation across traveling salesman problem (TSP) instances, but models trained on random Eucl…
HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization
Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have esse…
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing sk…
Dimensionality Reduction Meets Network Science: Sensemaking on UMAP's kNN Graph
While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlook…
Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations
Large Language Models are reshaping how software is developed and maintained. They are typically deployed in production using inference eng…
ABot-N1: Toward a General Visual Language Navigation Foundation Model
Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse…
Learning the Brain's Dynamics as a Port-Hamiltonian System: A GNN-Surrogate Metriplectic Twin for Non-Equilibrium Cortical Dynamics and Closed-Loop Neuromodulation
We model human motor cortex, recorded during rest and motor-imagery BCI conditions, as a port-Hamiltonian system: a conservative interconne…
Scaling Point-in-Time Language Models
Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias tha…
From Reconstruction to Interpretation: Zero-Setup Multi-Phase Segmentation of X-ray Tomography Data
X-ray tomography enables nondestructive characterization of material microstructures, while advances in micro-CT imaging have accelerated v…
Code-MUE: Measuring Code LLMs' Uncertainty through Execution-based Semantic Interaction Graphs
As Code Large Language Models (LLMs) become central to modern software engineering, their inherent stochasticity poses significant real-wor…
Jetson-PI: Towards Onboard Real-Time Robot Control via Foresight-Aligned Asynchronous Inference
Vision-Language-Action (VLA) models have achieved impressive performance on diverse embodied tasks. However, deploying VLA models on low-po…
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models
On-device LLM inference faces a trilemma of response latency, limited hardware resources and user privacy. Full cloud inference delivers st…
What Models Express, Suppress, and Resist: Auditing Open-Weight LLMs with Persona Vectors
What a language model will and will not do is largely set during post-training, but which behaviors it expresses, hides, or resists is not…
Faithful Autoformalization of Natural Language Assertions
Formal contracts are essential for software testing and verification, yet writing them remains labor-intensive and error-prone. LLMs offer…
LLM-as-a-Judge Scores Are Unreliable Optimization Signals in Closed-Loop Table Recognition
LLM-as-a-judge is widely used to provide feedback and selection signals in closedloop regeneration, but this use remains insufficiently val…
MxGPS: Multiplex Graph Transformers for a Power Grid Foundation Model
Single-task fine-tuning of graph neural networks (GNNs) for power grid problems exhibits a systematic failure mode: models that achieve the…
Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code
Languages with rich static semantics, such as Rust, provide stronger guarantees for AI-generated code, but their strictness makes generatio…
NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs
Synthesizing training data to scale agent capabilities in LLM post-training is bottlenecked by substrate-bound task synthesis: tasks are ge…
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage:…
Memory-Driven Self-Disclosure and Relational Turning Points: A Longitudinal Multimodal Study of Human-AI Interaction
As conversational AI systems are designed for repeated use, a central question is how a series of interactions becomes a relationship. We p…
Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents
The formation of political coalitions is a complex negotiation driven by both concrete policy objectives and deep-seated ideological convic…
T^2MLR: Transformer with Temporal Middle-Layer Recurrence
Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and m…
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery…
Hugging FaceにAI主導のサイバー攻撃 防御もAIで対抗するも、商用モデルは解析拒否で「GLM」採用
Hugging Faceは、自律型AIエージェントによる本番インフラへの侵入を検知、対処したと発表した。一部の内部データセットと複数の資格情報への不正アクセスを確認した。ログ解析には当初商用AIを使ったが、安全ガードレールに阻まれたため、最終的にオープンウェイトモデル「GLM…
What to watch for after Jensen Huang’s Japan visit
Jensen Huang left Tokyo with deals spanning Japan's entire tech ecosystem.
Can an Apple lawsuit derail OpenAI’s hardware plans?
On the latest episode of Equity, we debate whether Apple's lawsuit will cast over OpenAi's much-discussed plans to get into hardware and go…