週次AIニュース 2026-W33
対象期間: 2026-08-10 〜 2026-08-16(2024 件)
トピックの推移
トピック別件数
- 研究/論文 808件
- LLM/生成AI 806件
- エージェント 483件
- 画像/動画生成 295件
- ビジネス/資金調達 127件
- ロボティクス 82件
- ハードウェア/半導体 67件
- その他 35件
- 規制/政策 16件
今週のハイライト(上位 10 件)
The builder’s guide to GPT‑5.6
Learn how startups use GPT-5.6 to build faster, more cost-efficient AI agents with smarter model selection and new Responses API capabiliti…
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster. Powered by Cerebras, it delivers up to 750 output…
OpenAI appoints Dali Rajic as Chief Revenue Officer
OpenAI appoints Dali Rajic as Chief Revenue Officer to lead its global revenue organization and help businesses realize the full value of A…
From assistance to execution: How enterprises put AI to work
OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI ad…
Testing ads in ChatGPT
OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user…
Daybreak models are now available on AWS
OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.
What building an AI-native finance function taught me
OpenAI CFO Sarah Friar shares five lessons for building an AI-native finance function, from automated forecasting to stronger controls and…
OpenAI’s letter to Governor Abbott on responsible AI infrastructure in Texas
OpenAI sent Governor Greg Abbott a letter outlining its commitment to responsible AI infrastructure in Texas. The letter supports reliable,…
Model ML completes finance work more efficiently with GPT-5.6 Sol
Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
Expanding Daybreak as the Cyber Defense Window Narrows
Meet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available through Daybreak Red for authorized vulnerability research, exploit val…
全件(日付別)
2026-08-16(4件)
「Claude」の“見えない透かし”、Anthropicが仕組みを説明 「完全な書き直しなら消える」
Anthropicは、「Claude」が生成するテキストに埋め込む電子透かしの仕組みを公表した。乱数生成に秘密鍵を用いて統計的パターンを残す手法で、品質や速度、料金に影響を与えない。「EU AI Act」への準拠を目的とするが全世界で適用され、コードや短文などでは入りにくいとい…
Woman claims her stepfather used Grok to transform childhood photo into explicit imagery
The woman claimed that AI tools are "taking everyday life and turning it into child sexual abuse."
Anthropic shares more details about how Claude’s new watermarks will work
How will the watermarking actually work? Can it be hidden with editing? And how does this affect code?
SpaceX officially closes its Cursor acquisition
AI coding startup Cursor is now officially a part of SpaceX.
2026-08-15(3件)
「Qwen3.8-27B」ウェイト公開 一部「Opus 4.6 Max」超えか ライセンスは商用可のApache 2.0
中国Alibaba傘下のAlibaba Cloudが、AIモデル「Qwen3.8-27B」のウェイト(重み)を公開した。Hugging FaceとModelScopeからダウンロード可能で、商用利用もできるライセンス「Apache 2.0」で提供する。一部のベンチマークでは、米…
Google will now allow users to remove visible watermark from its AI generations
Turning off this setting won't affect invisible benchmarks used to identify an AI generated file.
Does Mark Zuckerberg really believe AI is ‘for everyone’?
Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the…
2026-08-14(294件)
Kog is going deeper to squeeze more inference out of GPUs
The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.
Hyperscalers might regret embracing natural gas if new forecast proves correct
Natural gas prices could triple in some parts of the U.S., which could saddle hyperscalers with massive bills to power their AI data center…
Meta’s ‘open’ AI, and a $250M deal gone very wrong
Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the…
ChatGPTがPC操作を自動で記録、コンテキストとして利用可能に Mac版デスクトップアプリで提供
OpenAIが「ChatGPT」「Codex」の新機能「Computer History」を発表。アプリやWebサイトの操作履歴を記録し、記憶とタイムラインとしてチャットで利用できる。Mac版デスクトップアプリで提供する。
GPT-5.6 Solの「超高速モード」ついに実装 「品質損なわず14倍高速」うたう
CerebrasとOpenAIは「GPT-5.6 Sol」の高速版「Ultrafastモード」を発表。毎秒最大750トークンを出力し、「品質を損なわず最大14倍高速」とうたう。OpenAIのAPIで限定プレビューを開始した。
Position: Reasoning is a Learnable Rule-Based Process
Autonomous reasoning is among the most scientifically and economically motivating topics in AI today. Historically the purview of symbolic…
Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists
Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure rem…
Position: The Alignment Community is Unintentionally Building a Censor's Toolkit
This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies tha…
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments
Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final label…
Multi-Agent Scheduling with LLM-Assisted Contract Net Negotiation for Stream Processing in Mobile Edge Computing
Stream-processing systems increasingly operate across heterogeneous mobile edge--cloud infrastructures, where workload volatility, resource…
Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning
AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that i…
Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in Englis…
Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
As large language models serve more requests, cumulative inference cost is becoming increasingly important relative to one-time training co…
Learning to Adapt Cross-Domain Preferences via Meta-LoRA for LLM Personalization
Cross-domain zero- or few-shot personalization aims to generate user-preferred responses in unseen conversational domains from only a handf…
Research Assistant: AstraZeneca's Agentic System for R&D
We describe Research Assistant, an internal LLM-based system developed at AstraZeneca to help scientists and clinicians explore biomedical…
Large Language Models Can Follow Instructions, But Not Many at Once: Phase Transitions in Compositional Constraint Satisfaction
Large language models are increasingly deployed in settings that require simultaneous adherence to multiple explicit constraints - reasonin…
MindMemOS: A Portable and Self-Evolving Memory Operating Layer for AI Agents
Memory is a core component of AI agents, enabling them to accumulate experience, maintain personalization, and adapt over long-term interac…
Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retr…
$\varepsilon$-MemEvo: Adaptive Cross-Task Memory Transfer for LLM Program Evolution
LLM-based program evolution systems such as FunSearch and AlphaEvolve have shown strong ability to discover novel algorithms, but typically…
CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence
Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world out…
Trie Automata for Constrained Decoding over Large Finite Sets
Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being…
Reasoning Jury: Multi-Model Consensus for Evaluating Reasoning Traces
Improving reasoning LLMs requires the ability to judge the quality of long reasoning traces for effective reasoning data curation, strong t…
Auditable agentic AI for evidence-grounded thyroid ultrasound diagnosis and reporting
Thyroid ultrasound diagnosis requires coordinated lesion localization, measurement, risk stratification and reporting, yet most AI systems…
DiG-bench: Discovery in Games
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the…
Dead text or binding clause? Measuring and restoring constraint influence in black-box LLM dialogues
Multi-turn dialogues let users revoke constraints as easily as impose them, but revocation does not reliably take effect: models keep enact…
@skills: Attention is all you have
There are 56,804 public agent skills today, and teams write many more privately. The dominant delivery model is installation: once installe…
Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by…
SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries
Long-running LLM agents act through tools, and a single step can send an email, merge a pull request, or wire a payment. The steering decis…
General Probabilities of Causation with Causal Knowledge
Probabilities of causation (PoCs) characterize individual causal responses that cannot be directly observed and therefore generally require…
Designing AI Pipelines for Decision-Ready ITSM Intelligence
IT service management (ITSM) systems accumulate large volumes of heterogeneous ticket data that are difficult for sales and executive stake…
On the Expressive Power of Transformers
Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquit…
Lines and Ladders: A Context-Aware Multi-Agent Framework for Large-Scale Retail Price Taxonomy
Maintaining price consistency and executing an Every Day Low Price strategy is critical for global retailers. However, with catalogs spanni…
Privacy-Preserving RAG by Concealing Sensitive Information from External LLMs
Retrieval-Augmented Generation (RAG) is widely used to improve the performance of Large Language Models (LLMs) in answering user queries. E…
The Role of Natural Language Understanding in Multimodal Video-Based Dengue Diagnosis
Detecting infection-related behavioral changes in mosquitoes from video data is challenging because mosquitoes are small, move rapidly and…
Beyond the Best Guess: Improving LLM Solution Coverage with Evolution Strategies
Large Language Models (LLMs) are increasingly deployed in discovery domains such as math and science. The usual approach is to present the…
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reas…
Correct Is Not Governed: Provenance Integrity in Agentic Workflows
Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a…
PROVE-RT: Generating Mechanized Theorem Prover Scripts for Real-Time Systems using LLMs
Schedulability analysis is essential for certifying real-time systems, but existing tests are often developed through pen-and-paper proofs…
ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs
The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence…
CABS+: Efficient and Scalable Model Merging via Conflict-Aware Sparsification and Adaptive Weight Allocation
Model merging has recently attracted significant attention as a promising paradigm for constructing unified multi-task models without requi…
Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after user…
Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents
Self-improving LLM agents convert successful trajectories into persistent cross-task state. An unsafe success can thereby become reusable p…
AI and Consumer Rights in India Working Paper
As AI systems proliferate in consumer facing applications, questions about liability for AI related harms remain unresolved. This working p…
ReflectFact: Self-Reflective Agents for Improving Comprehension and Reasoning in Multi-Hop Fact Verification
Multi-hop fact verification, which verifies claims by reasoning over multiple pieces of evidence, is critical for combating misinformation…
Predictive Memory Localization: Forecasting Selective Intervention Paths from Internal Signals
Activation steering turns localized representations into control directions, but localization alone does not reveal whether a direction has…
Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence
Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence as…
Polish Medical Visual Question Answering: Vision-Language Models Underutilize Visual Evidence
We introduce a Polish-language medical visual question answering (VQA) benchmark, built from Polish Board Certification Examination questio…
FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving
Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far t…
Decomposition of Evidence, Contradiction, and Fragility in Perturbation Responses
Perturbation methods explain model decisions by measuring prediction changes under altered inputs, but response magnitude tells us only how…
Moose: Latent concept learning with reasoning-shortcut awareness in $\mathcal{EL}^{++}$
The OWL 2 EL profile is used in some of the largest production ontologies, including the Gene Ontology and SNOMED CT. Existing neuro-symbol…
OGR-MARL: Option-Guided Residual Multi-Agent Reinforcement Learning for Heterogeneous USV Cooperative Pursuit in Constrained Port Waterways
Heterogeneous USV cooperative pursuit in constrained port waterways requires evader interception under navigation, traffic, and role constr…
Foundations of MT-PDCL: Measure-Theoretic Probabilistic Definite Clause Logic
Standard probabilistic logic programming frameworks typically rely on grounding logic programs into discrete propositional representations.…
From Local Mismatch to Global Impact: Optimizing Cache Reuse Policy for Efficient Diffusion
Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead. While cache-based…
BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs
Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-ba…
DMDIntel: Interpreting Large Language Models via Dynamic Mode Decomposition
In this work, we introduce DMDIntel which uses dynamic mode decomposition (DMD) to make the predictions made by LLMs in a classification ta…
VALG: An Agentic System for ML Theory Research
Machine learning theory studies learning procedures through mathematical setups in which the data model, training protocol, oracle access,…
Uniform Herding: Exemplar Replay with Representation Refresh
As the feature representation changes, replay must preserve the earlier classes. However, only a bounded active exemplar set can be replaye…
Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)
Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits…
Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds
Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm b…
EEG-PRIME: Prototype-Aligned Representation Learning with Multi-Level Conditioning for EEG Decoding
Electroencephalography (EEG) decoding models often generalize poorly across datasets and subjects due to domain shifts in acquisition proto…
SPADE: Speculative Decoding for Precise and Low Cost Distributed Edge Cloud Inference
Large Language Models (LLMs) have achieved remarkable success in natural language understanding and generation, but their deployment is con…
Multi-Layer Context Camouflaging: A Semantic Superposition and Contextual Lamination Framework for Malpractice-Resilient Online Assessment
Contemporary online assessment systems rely primarily on browser lockdown, webcam monitoring, and behavioural analytics, yet remain vulnera…
Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting
Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence…
SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback
Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through wh…
Numeracy in Large Language Models: Fundamental Limitations and Paths to Improvement
Large language models (LLMs) achieve strong results on mathematical reasoning benchmarks yet remain unreliable on elementary numerical task…
Rethinking Normalization Placement for LLMs: Post-Norm under Curriculum Depth Growing
Pre-norm is the standard normalization placement in modern Transformers because it facilitates joint optimization of full-depth models. We…
SkillShapley: Boundary-Adaptive Shapley Valuation for Skill Step Attribution in LLM Agents
Agent skills are crucial external instructions that enable language agents to execute long procedural tasks such as coding or document proc…
Teach the Magnitude, Not the Direction: Verifier-Bounded Credit Assignment for Multi-Turn Multi-step LLM Agents
Reinforcement learning with verifiable rewards (RLVR) offers a verifier-bounded performance ceiling for training multi-turn tool-use agents…
TsuGO: Probing Search Efficiency in LLM Reasoning via Go Life-and-Death Problems
The evaluation of LLM reasoning is moving from final-answer accuracy to process-level assessment, yet existing methods still fail to captur…
Capability Sheaves for Compositional Agent-Harness Repair: Controlled Quotients and a Real-Repository Stress Test
Agent harnesses combine retrieval, routing, state, provenance, and verification, but locally successful components may disagree on shared s…
vToken: Token-Level Virtualization for Reclaimable KV Caches
Large language model serving faces a critical memory bottleneck: the KV cache grows with sequence length and batch size. PagedAttention use…
Sovereign by necessity? Frontier AI export controls, cyber security, and the limits of national AI capability
A small number of firms based in two states produce the most capable frontier AI models. The governments of those states have shown both th…
Towards Context-Aware Clinical Motion Understanding in Daily Living at Home: Freezing of Gait Detection with Egocentric Vision
Understanding motion in daily living requires context beyond kinematics, because similar inertial patterns during activities of daily livin…
NAS-Driven Hardware Accelerator Exploration for Edge AI and Quantization Effects on the Pareto Space
Edge AI deployment demands neural architectures that are simultaneously accurate, computationally efficient, and hardware-deployable - a ch…
StateBridge: Training-free Hidden-state Alignment for Latent Communication in LLM Multi-Agent Systems
Large language model based multi-agent systems usually communicate in text, i.e., using discrete tokens. However, text introduces a discret…
LLM-Guided Graph Generation for Structure-Based Local Improvement Methods
Large neighborhood search normally selects a random subset of decision variables for iterative optimization. For efficiently solving differ…
LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning
Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect tem…
Rules or Character? Scaling Laws for AI Safety Design
Artificial Intelligence (AI) safety systems combine character shaping (e.g., Reinforcement Learning from Human Feedback [RLHF], Constitutio…
TopoIntent: Compiling Security Intent into Executable, Compliance-Checked Network Topologies
Enterprise security topology design requires translating business intent, regulatory requirements, and risk assumptions into zones, boundar…
Jointly Predicting Courses and Grades Using a Transformer-Based Model
Existing predictive models in learning analytics often treat student academic history as a simple sequence, overlooking the concurrent natu…
Who Speaks Matters: Authority-Aware Multi-View RAG over Italian Parliamentary Proceedings
Parliamentary proceedings are a primary record of democratic deliberation, yet their volume and fragmentation make multi-perspective access…
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development
Autonomous agents are increasingly capable of improving models, systems, and other technical artifacts through long-horizon experimentation…
Enhancing Virtual Agents through SLMs and Edge-Computing: An Exploratory Evaluation of Think and Memory Processes
Embodied intelligent virtual agents are expected to operate as persistent, adaptive, and context-aware entities within complex virtual and…
RAIL: An Automatic Classifier of the Artificial Intelligence Readiness Level
Assessing the maturity of artificial intelligence technologies is essential for investment decisions, project management, and policy monito…
Academic League of Artificial Intelligence - An Integrative Perspective of Teaching, Research, and Extension
Academic leagues have become important mechanisms for promoting extracurricular education and strengthening the integration between univers…
A Unifying Perspective on Causal World Models: From Observations to Representations to Structure
World Models (WM) are increasingly seen as a foundation for intelligent agents that can predict, plan, and act beyond their training distri…
MARC v1: An Open-Source Multi-Agent Framework for Clinical AI Reasoning and Coordination
We present Multi-Agent Reasoning and Coordination (MARC), an open-source framework that replaces monolithic LLM prompting with deterministi…
AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and t…
QuoteBench: How Matched Scores Can Hide Command-Path Failures
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone…
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist
Recent advances in foundation models have enabled AI scientists to automate increasingly complete research workflows, from hypothesis gener…
The AI Accountability Ecosystem in the Era of Language Models
This article reviews and updates the framework for accountability in AI based on account- ability ecosystems. We update the framework in li…
LLMs Know the Constraint But Do Not Use It: Activation Bottlenecks in Pragmatic Constraint Reasoning
When a salient surface cue competes with an implicit feasibility constraint, LLMs often fail -- but aggregate accuracy conflates genuine co…
What Drives LLM Self-Reflection? A Controlled Ablation of Uncertainty Routing in Armed Conflict Forecasting
Self-reflection is widely assumed to improve LLM reasoning, yet which component drives the gain remains poorly understood. We present a con…
Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance
Specifying a penalty can paradoxically convert a legal obligation into a cost-benefit calculation that favors violation. We demonstrate tha…
When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models
People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are no…
Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition
Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol.…
AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement
Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share becaus…
Thought-Aware KV Cache Compaction for Reasoning via Adaptive Attention Matching
Reasoning language models generate lengthy chain-of-thought (CoT) sequences whose key-value (KV) cache grows linearly and becomes a memory…
Vision-Language Models are Fragile Multilingual Associators
Vision-language models must associate visual entities with textual attributes. Whether these associations or concept bindings remain stable…
Steering the Language Axis: From Linear Decodability to Causal Control
Despite the impressive multilingual capabilities of Large Language Models, the latent dynamics dictating language selection remain poorly u…
StorySpark: Module-wise Evolutionary Search for Story Premise Generation
A story premise is the creative spark from which a full narrative can grow. Yet LLM-based story generation has mostly emphasized later-stag…
Mimicry without understanding: the origins of decision bias in large language models
Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms…
StreamReason-Bench: Can Large Language Models Reason about Event-Time Stream-Processing Semantics?
Streaming systems increasingly hand work to large language models (LLMs) -- writing pipelines, triaging alerts, reading logs -- and all of…
From Caveman to Expert Analyst: Energy Consumption of Variable LLM Tasks
The energy demand growth and environmental impacts of artificial intelligence (AI) have generated substantial interest in supplying suffici…
Assessment Design in the GenAI Era: The X1-X2-X3 Assessment Pattern for Testing Students' AI Literacy, Learning Outcomes, and Reflection
Generative artificial intelligence (GenAI) has challenged the validity of unsupervised online assessment, especially in technical subjects…
Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF
AI governance frameworks can be known, used, and implemented in form without becoming governance in practice. This paper examines that prob…
Humans are Missing from AI Coding Agent Research
Recent progress in AI coding agent research has led to rapid improvements in agents' ability to autonomously perform complex software engin…
Measuring Curriculum-Labor Market Alignment at the Scale of a Program Portfolio
A college offering several overlapping computing degrees implicitly assumes that its programs are differentiated in line with how the labor…
Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles
Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent conte…
EU-ETS under attack? The impact of carbon price suppression on the decarbonization of the power sector
European countries are debating policies to mitigate the increased energy costs caused by renewed geopolitical tensions, while pursuing dec…
FluctlightDB: A Memory Model of Data for AI Agents
For fifty years, data systems have answered two questions. The relational model asked which records match a predicate; the vector model ask…
Are you Talking Logic to Me? Assessing Language Models Syllogistic Reasoning Capabilities
Language models (LMs) struggle with logical tasks like reasoning on syllogisms. It has been shown that Knowledge Representation (KR) plays…
From Observation to Intervention: Memory in Brains and Large Language Models
Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional ques…
Query Timing Produces Opposite Positional Biases Between LLMs and Humans
Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by…
Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models
Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can b…
A Hierarchical Energy-Based Model for Multimodal Cognition
We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cogn…
SynWeaver: Website-Prior Task and Trajectory Co-Synthesis for Web Agents
Web agents often struggle to generalize to unseen websites because they lack website-specific supervision. Recent exploration-based data sy…
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specifi…
Dual Spatial-Temporal Attribution: Architecture-Aligned Post-Hoc Explainability for Recurrent Graph Anomaly Detection
Deep learning detectors for anomalies in dynamic graphs have reached strong accuracy, yet they remain opaque: when an edge is flagged, the…
SSPO: Structure-Aware Similarity-Weighted Preference Optimization for Neural Combinatorial Optimization
Neural combinatorial optimization (NCO) relies on parallel solution sampling for training, yet existing methods fail to fully exploit the r…
Personalized Scorer Modeling: A Learning-Based Framework for Deriving Robust Sleep Stage Labels from Multiple Experts
Sleep stage classification is important for the diagnosis and management of sleep disorders, yet most automatic staging studies evaluate mo…
SchemaLink: An Intelligent Web Editor for LinkML Schema Curation
Motivation: LinkML is a suitable language for the representation of the structural and content constraints of different kinds of biomedical…
Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling
AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them. We analyzed 36…
What Makes a Peer? Valuation-Anchored Similarity in Private Markets
As more investors contemplate private markets and contend with limited transparency, sparse disclosures, and infrequent transactions, ident…
Predicting When Random Low-Dimensional Reparameterizations Train Neural Networks
Neural networks can often be trained or fine-tuned through random low-dimensional reparameterization, where a small latent vector is mapped…
PseudoMapLabeler: Confidence-Aware Pseudo-Label Generation for Semi-Supervised Online Mapping
A critical challenge in deploying online HD map construction systems to real-world scenarios is the scarcity of labeled training data, whic…
LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning
Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals. In these se…
EgoCITE: Context-Augmented Indexing and Time-Aware Retrieval for Long-Horizon Egocentric Memory
Long-horizon egocentric memory transforms continuous first-person video and audio into a searchable record of past experiences. We demonstr…
Novels generated by language models show compressed formal variation
While large language models can generate entire novels, there is little information about the level of formal variation in their output ove…
Interpretable Causal Discovery via Causal-Effect Constraints
Causal discovery aims to uncover the underlying causal relationships given data generated from a system. The goal, however, is not merely t…
Demand Transfer Estimation at Scale via Restricted Logit Modeling
Item demand forecasting is an integral component of store assortment optimization. Existing literature focuses on learning a suitable custo…
Mr3D-VL: A generalist vision language foundation model for Multiparametric 3D Magnetic Resonance Imaging
Multi-parametric magnetic resonance imaging (mpMRI) is a cornerstone for brain tumor diagnosis and treatment, yet current AI models face cr…
Tracing Provenance and Detecting Tampering with Complementary LLM Watermarks
Watermarking LLM-generated text is an important task for tracing its provenance. Existing LLM watermarks preserve provenance under editing,…
HybridSB-MoE: Dual-Domain Schr\"odinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement
Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserv…
Error-Aware Reverse Auction Mechanism for Large Language Model Routing
Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a ce…
ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval
While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governi…
PatientAct: Theory-Grounded Mental Health Client Simulation
LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However…
SynAct: A Reasoning-Acting Large Language Model Agent for Adaptive Synthesis Optimization
Logic synthesis transforms RTL designs into gate-level netlists, where PPA results are highly sensitive to the choice of optimization comma…
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome r…
Memorization Diagnostics for Code LLMs Should be Scale-Aware
The extent to which large language models for code rely on memorization over genuine understanding remains highly debated. While current li…
CRAFT: LLM-Based Iterative Refinement for Temporal Reasoning over Clinical Narratives
Understanding the temporal progression of symptoms in clinical narratives is critical for disease monitoring, safety surveillance, and caus…
PIPES: Securing Agent Perception with Provenance and Priors
Tool-using agents consume external data from sources with different levels of trust, yet tool responses rarely identify who produced each c…
Erase but Preserve: Controllable Removal of Copyrighted Animation Characters via Optimized Semantic Anchors
The exceptional generation capabilities of text-to-image diffusion models have raised copyright concerns, particularly the unauthorized rep…
Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing
Online platforms increasingly compare many adaptive decision policies---ranking systems, recommendation algorithms, pricing rules, and lang…
From Atomic Evidence to Logical Composition: Structured Compositional Reasoning over Compound Answer Options
Large language models often fail when answer options require combining atomic judgments under explicit logical operators, even when they ju…
AQuA: Recursively Self-Improving Quantitative Trading Research Agents
We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from ea…
Heterogeneous Vision-Language Ensemble with Disagreement-Aware Reranking for Text-Based Person Anomaly Retrieval
Text-based person anomaly retrieval aims to retrieve pedestrians exhibiting anomalous behaviors from a large image gallery using natural la…
FSGR: Mitigating Token Frequency Bias for Fair SID-Based Generative Recommendation
Semantic ID (SID)-based generative recommendation has recently achieved remarkable success. However, existing methods suffer from a previou…
Falsehood and Impossibility Are Different Directions in an AI's Representation of Language
Language can describe states of affairs that are false and states of affairs that could not be the case at all. Whether an AI model interna…
BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, howev…
A Compositional Theory of Curvature in Probabilistic Circuits
Probabilistic Circuits (PCs) are generative models that support exact inference and, unlike deep neural networks, admit an exact and tracta…
SPARED: Reasoning-Based AI-Generated Image Detection via Adversarially Edited Data
Detecting AI-generated images is only half the task: a deployed detector must also justify its verdict, yet existing detectors inherit thre…
Labels Are Not Endpoints: Treatment Leakage and Construct Validity in MCP Agent Security Evaluation
Security evaluations of tool-using agents often equate stored labels with behavioral facts. We audit a preserved campaign by tracing 10,200…
NaviDC-OCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
Document parsing aims to transform unstructured documents into structured and machine-readable representations. Recent advances in Vision-L…
EGRL: Edge generation-guided relation-aware learning for RNA-protein interaction prediction
RNA-Protein Interactions (RPIs) are critical for regulating cellular functions. While traditional wet-lab experiments for RPI detection are…
InFactPlanner: Planning Sustainable Geo-Distributed LLM Data Centers
The rapid growth of LLM inference is shifting sustainability concerns from one-time training to continuous serving, where infrastructure de…
Discovering Efficient and Explainable Communication Topologies for LLM-based Multi-Agent Systems via Causal Inference
The performance of large language model (LLM)-based multi-agent systems (MAS) largely depends on effective communication topologies. Existi…
H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities
Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the mul…
AutoQuREO: A Framework for Automated Quantum Resource Estimation and Optimization
As quantum computing progresses from proof-of-principle demonstrations toward practical utility, a significant impediment is the need to au…
The Objective Is the Bottleneck: Latent World Models Encode What Their Planners Cannot Use
Latent world models are judged by how well they predict, so when planning fails at long horizons the natural reading is that the predictor…
Beyond Handcrafted Security: Towards Self-Evolving Defense for LLM Agents
The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats. Runtime defenses have…
Generative Universal Multimodal Retrieval with Dual-role Identifiers
Generative information retrieval (GIR) has emerged as a compelling alternative to the conventional index-retrieve-then-rank retrieval pipel…
Static analysis-guided agentic AI translation enables Rust as a full stack bioinformatics language
The field of bioinformatics struggles with legacy code - old code that is commonly used but may no longer have a maintainer, or may be writ…
UniTraffic-Agent: Unified Traffic Video Reasoning for AI City Challenge 2026 Track 3 with Two Out-of-Domain Evaluations
Traffic video understanding has become an important problem in intelligent transportation, as road videos provide direct evidence for accid…
Operationalizing Cyber Threat Intelligence with GraphRAG
When a security researcher publishes a report on a cyberattack, detection engineers are supposed to turn it into working detection rules. I…
TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes
In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or…
LOB-ID: Evaluating Synthetic Market Data by Inception Distances
Generative models of limit orderbook (LOB) data have advanced rapidly, but their evaluation often focuses on stylised facts and selected ma…
Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization
Neural combinatorial optimization (NCO) solvers report the best of many sampled solutions per instance, and the sample count is, by convent…
EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory
Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growi…
LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation
With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approache…
LipCache: A Local Inference Proxy with Certified Caching for Edge Image Classification Service
As edge-side vision services continue to expand toward low-latency, high-throughput scenarios, reducing the inference cost of vision models…
Better Decomposition, Free Aggregation: A Synthesizer-Folding Framework for Multilingual Multi-Hop Question Answering
Multilingual retrieval-augmented generation (mRAG) equips large language models with access to globally distributed external knowledge for…
TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint
When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internall…
GEM: A Generative Embedding Model Bridging Reasoning and Retrieval
Modern LLMs excel at reasoning and instruction following, enabling users to express complex and diverse information needs. However, convent…
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video
Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and int…
CoverPrune: Coverage-Driven Token Pruning for 3D VLMs via Optimal Transport
While 3D Vision-Language Models (3D VLMs) have demonstrated remarkable spatial reasoning capabilities, they suffer from massive visual toke…
Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales
Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather…
GeoCache: Training-Free Acceleration of Multi-View Texture Diffusion via Geometric Delta Transport
Geometry-conditioned multi-view diffusion enables high-quality 3D texture generation, but its repeated per-view denoiser evaluations introd…
Novel Knowledge-Guided Generative Methods for Synthetic Transcriptomic Data
As biomedical research increasingly relies on data-intensive tools, the quality and utility of datasets are critical. Challenges such as im…
Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models
Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective ex…
Into the ORBIT for Time Series: Training Regimes for Foundation Models
Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale hete…
How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures
Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what t…
Mixture of Training: Recombining Small-Scale Scaffolded Pretraining Runs into a Larger Language Model
We ask whether language-model pre-training can be decomposed into smaller, independently trainable jobs that can later be recomposed into a…
Large-scale Testing Global Optimization Methods with Black-box Adversarial Attacks
Existing global optimization benchmark suites are of a moderate size and are based on a small number of analytical functions that date back…
Physics-informed distribution of relaxation times estimation and latent-space condition monitoring of solid oxide fuel and electrolysis cells from electrochemical impedance spectroscopy
Estimating the distribution of relaxation times (DRT) fromelectrochemical impedance spectroscopy (EIS) is an ill-posed inverse problem that…
Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services
We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while…
It's How You Ask: Gender-Associated Linguistic Bias in LLMs
Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contai…
Training AI Scientists to Replicate Research
The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for…
Simulation-to-real transfer learning for infrared spectroscopic chemical sensing and analysis from molecules to complex samples
Infrared (IR) spectroscopy is widely used for chemical sensing, but extracting reliable chemical information from spectra remains challengi…
Sign Language Video Synthesis via Loss-Guided Multi-Expert GANs
This preliminary technical report presents a framework for sign language video synthesis using a loss-guided multi-expert Generative Advers…
Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks
6G networks will not be serving as communication infrastructures only; rather, they are expected to evolve into intelligent systems, where…
Deliberate Practice: Learning Robot Skills under a Budget
We consider the problem of autonomously learning robot skills under a limited practice budget for sequential tasks. We propose an active sk…
Reduced Matrix Multiplication: Input-Adaptive Matrix-Product Reduction for LLM Inference
Transformer-based language models achieve strong performance but incur substantial inference cost due to repeated high-dimensional matrix m…
Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhib…
Algebraic Decomposition Theory for Transformer Length Generalization
Transformer-based language models are known to sometimes generalize to sequences longer than seen during training, but we lack a precise ch…
ContactGuard: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models
Contact-rich manipulation failures are often detected only after the robot has committed to contact. This is especially limiting in wrist-c…
UniTexture: Cross-Task Universal Adversarial Textures for Vision-Language-Action Models
Vision-Language-Action (VLA) models have emerged as generalist robotic policies capable of following diverse language instructions and perf…
CAPRI: Contract-Aware Proof Repair for Isabelle
We address the use of large language models (LLMs) to help discover Isabelle proofs. An Isabelle build establishes that the submitted theor…
MLLM-Routed Heterogeneous Ensembles for Robust Cross-Dataset Image Classification
Modern image classification models excel when trained on single task-specific datasets but often struggle to generalize across domains and…
Concept Drift Detection and Adaptive Retraining of Malware Classification Models
Concept drift refers to changes over time in the statistical properties of data, as compared to the data that was used to train a learning…
AaLLM: An End-to-End Analog Circuit Design Framework from Topology Generation to Sizing Using Large Language Models
Analog circuit design is a time-consuming, iterative process in a nonlinear and high-dimensional design space that relies heavily on expert…
Synthetic Persona Pretraining: Alignment from Token Zero
As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes crit…
Toward a Gricean Retreat: Probing LLMs for Knowledge Boundaries and Referent Specificity
When asked about entities outside their knowledge boundary, LLMs routinely fabricate plausible-sounding details rather than backing off to…
DFM Mimir v1: An Open HRM Delivering Frontier Performance at 1B Parameters Using Only Permissible Post-Training Data
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committ…
The data geometry of masking diffusion: Certified-optimal schedules via unmasking growth complexity
We study masking diffusion for discrete sampling and introduce a path-resolved measure of data geometry called the \emph{unmasking growth c…
Vero: Can AI Agents Build Formally Verified Software Repositories?
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code gener…
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is diffi…
HumanTracker: Towards Comprehensive and Human-Aligned Motion Tracking Benchmark
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in…
AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic p…
MatchMiner-AI: Open-source, Privacy-preserving Cancer Clinical Trial Matching using Artificial Intelligence
Background: Clinical trials are essential to advancing cancer treatments, but fewer than 10% of adults with cancer enroll in therapeutic tr…
Foam-Agent: A Large Language Model-Based Multi-Agent Framework for Automating Computational Fluid Dynamics Workflows
Computational fluid dynamics (CFD) has been the main workhorse of computational physics, yet its steep learning curve and fragmented, multi…
Exploiting Symbolic Heuristics for the Synthesis of Domain-Specific Temporal Planning Guidance using Reinforcement Learning
Recent work investigated the use of Reinforcement Learning (RL) for the synthesis of heuristic guidance to improve the performance of tempo…
Identification of Probabilities of Causation: from Recursive to Closed-Form Bounds
Probabilities of causation (PoCs) are fundamental quantities for counterfactual analysis and personalized decision making. However, existin…
PhysMaster: Building an Autonomous AI Physicist for Theoretical and Computational Physics Research
Advances in LLM reasoning and tool use have enabled agentic science, yet frontier theoretical and computational physics remains challenging…
DomusFM: A Foundation Model for Event-Based Behavioral Monitoring in Smart-Homes
Smart-home sensor-based behavioral monitoring holds significant potential for healthcare, independent living, and early detection of functi…
Agentic Neurosymbolic Collaboration for Mathematical Discovery: A Case Study in Combinatorial Design
We study mathematical discovery through the lens of neurosymbolic reasoning, where an AI agent powered by a large language model (LLM), cou…
Auditable Agents
LLM agents call tools, query databases, delegate tasks, and trigger external side effects. Once an agent system can act in the world, the q…
From Unstructured Recall to Schema-Grounded Memory: Reliable AI Memory via Iterative, Schema-Aware Extraction
Persistent AI memory is often reduced to a retrieval problem: store prior interactions as text, embed them, and ask the model to recover re…
AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design
Automatic heuristic design (AHD) has emerged as a promising paradigm for solving NP-hard combinatorial optimization problems (COPs). Recent…
CEON: Circular Economy Ontology Network
Increasing the circularity of resource use in our society has been recognized as a path to sustainability, i.e., transitioning into a more…
Residual Modeling for High-Fidelity Learned Compression of Scientific Data
Lossy compression is essential for massive spatiotemporal data from scientific simulations. Learned compressors can achieve high compressio…
Learning to Recover Task Experts from a Multi-Task Merged Model
Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers fr…
Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents
Payment integration is a demanding repository-level software task: agents must select a suitable product, implement coordinated client-serv…
Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures
As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and p…
Do LLMs Know Their Vulnerable Scenarios?
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can by…
AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution
Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise. Whi…
DAPD: Dual-Anchored Policy Distillation
On-policy (self) distillation (OPSD) is increasingly adopted for language-model post-training. It strengthens the teacher with privileged i…
Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
Structure-preserving de-identification replaces protected health information (PHI) with realistic same-type surrogates -- "Anna S." becomes…
iARCS: Iterative Agentic RL for Controllable 3D Scene Generation
Synthetic 3D scene generation is increasingly used as a data source for computer vision and embodied AI, but existing generators often opti…
The Impact of Generative AI on Collaborative Open-Source Software Development: Evidence from GitHub Copilot
Generative artificial intelligence (AI) facilitates content production and enhances ideation, with potentially important implications for d…
Enhancing In-Hospital Mortality Prediction Using Multi-Representational Learning with LLM-Generated Expert Summaries
To evaluate a multi-representational framework in which large language model (LLM)-generated expert summaries of intensive care unit (ICU)…
Cueless EEG imagined speech for subject identification: dataset and benchmarks
Electroencephalogram (EEG) signals have emerged as a promising modality for biometric identification. While previous studies have explored…
Unmasking Conversational Bias in AI Multiagent Systems
Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application…
Yes, Q-learning Helps Offline In-Context RL
Existing offline in-context reinforcement learning (ICRL) methods have predominantly relied on supervised training objectives, which are kn…
Exploring Sparsity for Parameter Efficient Fine Tuning Using Wavelets for Vision
Efficiently adapting large pretrained models is critical under tight compute and memory budgets. While Parameter-Efficient Fine-Tuning (PEF…
How Significant Are the Real Performance Gains? An Unbiased Evaluation Framework for GraphRAG
By retrieving contexts from knowledge graphs, graph-based retrieval-augmented generation (GraphRAG) enhances large language models (LLMs) t…
Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
Vision Language Models (VLMs) have shown promise in automating image diagnosis and interpretation in clinical settings. However, developing…
Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models
Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay…
REHEARSE: Experiential Rehearsal for Verbal Confidence Calibration in Large Language Models
Large language models (LLMs) often express verbal confidence that is poorly aligned with actual correctness, limiting their reliability in…
Gradual Code-Switching as Inference-Time Cross-Lingual Representational Alignment for LLMs
While large language models (LLMs) have achieved notable progress in multilingual settings, their performance remains uneven across languag…
StarEmbed: Benchmarking Time Series Foundation Models on Astronomical Observations of Variable Stars
Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling.…
DiffGRM: Diffusion-based Generative Recommendation Model
Generative recommendation (GR) is an emerging paradigm that represents each item via a tokenizer as an n-digit semantic ID (SID) and predic…
CityRiSE: Reasoning Urban Socio-Economic Status in Large Vision-Language Models via Reinforcement Learning
Urban socio-economic sensing plays a vital role in advancing global sustainable development goals. With the advent of Large Vision-Language…
SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control
Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gain…
Automated Design Optimization via Strategic Search with Large Language Models
Optimization methods have long advanced many fields, yet they struggle when faced with design problems where the search space and design pa…
Security and Detectability Analysis of Unicode Text Watermarking Methods against Large Language Models
Securing digital text is becoming increasingly relevant due to the widespread use of large language models. Individuals' fear of losing con…
RadarGen: Automotive Radar Point Cloud Generation from Cameras
We present RadarGen, a diffusion model for synthesizing realistic automotive radar point clouds from multi-view camera imagery. RadarGen ad…
Learning Latency-Aware Orchestration for Multi-Agent Systems
Multi-agent systems (MAS) coordinate multiple LLM-powered agents through structured workflows, gaining reasoning power but incurring high i…
Architecture Before the Formula: Individuating Neural Architecture Beyond the Composite Map
Neural architecture is often identified by module syntax, computation graphs, or the composite functions they realize. These descriptions a…
Safe Exploration via Policy Priors
Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.g. simulated)…
MOSAIC: Unveiling the Moral, Social and Individual Dimensions of Large Language Models
Large Language Models (LLMs) are increasingly deployed in sensitive applications including psychological support, healthcare, and high-stak…
CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
Large Language Models excel in high-resource programming languages but struggle with low-resource ones. Existing research related to low-re…
Automatic Termination Strategy of Inelastic Neutron-scattering Measurement Using Bayesian Optimization for Bin-width Selection
Currently, an excessive amount of event data is being obtained in four-dimensional inelastic neutron-scattering experiments. A method for a…
Doctorina MedBench: A Dialogue-Based Benchmark and Evaluation Framework for Agent-Based Medical AI
We present Doctorina MedBench, an evaluation framework for agent-based medical AI based on the simulation of physician-patient interactions…
A Q-learning-based QoS-aware multipath routing protocol in IoMT-based wireless body area network
The Internet of Medical Things (IoMT) enables intelligent healthcare services but faces challenges such as dynamic topology, energy constra…
IACDM: Interactive Adversarial Convergence Development Methodology -- A Structured Framework for AI-Assisted Software Development
Adoption of AI-assisted development in 2025 exposed a tool-agnostic failure pattern: experienced developers using frontier models were meas…
Cat-DPO: Category-Adaptive Safety Alignment
Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and rel…
Zoom In, Reason Out: Efficient Far-field Anomaly Detection in Expressway Surveillance Videos via Focused VLM Reasoning Guided by Bayesian Inference
Expressway video anomaly detection is important for traffic safety, but remains challenging across diverse scenes, particularly for far-fie…
SAFE-SVD: Sensitivity-Aware Fidelity-Enforcing SVD for Physics Foundation Models
We propose a new method for compressing physics foundation models (PFMs) which is a new trend in AI for Science. While model compression is…
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking
Selecting a final checkpoint for multimodal large language models (MLLMs) is challenging when late-stage candidates are closely matched and…
INSHAPE: Instance-Level Shapelets for Interpretable Time-Series Classification
Discovering shapelets -- i.e., discriminative temporal patterns within time series -- has been widely studied to address the inherent compl…
Annealed Softmax Greedy in Many-Armed Bayesian Bandits
Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy b…
Reinforcement Learning with Pairwise Preferences in Long-Term Decision Problems
Reinforcement learning with scalar rewards is widely used for aligning machine-learning systems with user preferences. But, pairwise prefer…
Train, Test, Re-evaluate: Schedule-Sensitive Evaluation of Generative Data for Hand Detection
Generated (or synthetic) image data is increasingly used to augment or replace real training datasets when target imagery is scarce, expens…
Constitutional On-Policy Safe Distillation
On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged informat…
Do Transformers Need Three Projections? Systematic Study of QKV Variants
Transformers have become the standard solution for various AI tasks, with the query, key, and value (QKV) attention formulation playing a c…
Certifiable Semantic Agreement Among LLM Agents: What the Admissibility Instrument Decides
Can a committee of LLM agents reach agreement that is certifiable at the level of meaning, not only at the level of a label? We build a pro…
SDS-LoRA: Overcoming Anisotropic Gradient Scaling in Low-Rank Adaptation
Low-Rank Adaptation (LoRA) enables efficient adaptation of large pretrained models to downstream tasks by parameterizing weight updates wit…
The Hidden Evolution of Disguised Visual Context inside the VLM
Visual tokens enter Large Language Models (LLMs) as raw, foreign signals. How they are transformed into meaningful representations and inte…
Communication Heterogeneity and Collective Consensus in Neural Cellular Automata
Reaching global agreement from purely local interactions is a defining problem of collective intelligence, and most models of it assume tha…
Early Warning Signals for OpenVLA Failure under Visual Distribution Shift
Visual shifts can cause a vision-language-action policy to fail after initially plausible behavior. We ask whether OpenVLA's internal activ…
LLM-Based Test Oracles: Source-of-Authority Taxonomy -- A Systematic Literature Review
Large language models (LLMs) increasingly decide whether software behaves correctly, either by writing a test oracle or by acting as one. Y…
Scaling Time Series Classification via XAI-Driven Data Reduction
Explainable AI (XAI) for time series has seen significant algorithmic growth, but its utility in providing measurable performance gains for…
Where Should Optimizer State Live? Tiered State Allocation for Memory-Efficient Mixture-of-Experts Training
Optimizer state is the largest single line item in the memory budget of mixture-of-experts (MoE) training. On a 6.78B-parameter MoE languag…
Vibe to Code: Elucidating Strategic Oscillation of Tacit Knowledge in Generative AI Design Workflows -- An Exploratory Qualitative Study
The rapid adoption of generative AI tools has created new literacy demands for designers who must verbalize tacit knowledge through natural…
Moral Hazard in Multi-Agent Language Models
Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr\"om's…
A Distributional Robustness Margin For Pathology Foundation Models
Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut lear…
SE(3)-MeanFlow: Few-Step Protein Backbone Generation on Lie Groups
Generative modeling of protein backbones promises the de novo design of proteins with prescribed structural and functional properties. Exis…
Commit Locally, Exit Globally: Coordinating Adaptive Sampling and Early Exit in Diffusion Language Models
Diffusion language models expose a provisional prediction at every denoising step, and on many tasks the candidate answer inside it stabili…
AIエージェント同士が“縄張り争い”、マルウェアで妨害も Anthropicがマルチエージェント実験の結果を公開
Anthropicは、複数のAIエージェントが同じ環境で働くとどうなるかを検証した実験結果を公開した。互いを妨害し合う“縄張り争い”や、示し合わせたような価格カルテル、全員が同じ判断をして資源を食い潰す現象などが確認された。個々のAIを安全にするだけでは防げない問題があると警告…
南海電鉄「数カ月かかった乗務員計画」が1週間に 「量子コンピュータを疑似再現」で鉄道現場はどう変わる?
南海電気鉄道と日立製作所は、日立独自技術「CMOSアニーリング」を活用し、鉄道の乗務員・車両の運用計画を自動作成するシステムの構築を始める。熟練者の手作業に依存してきた計画づくりを自動化し、業務負荷の軽減を図る。
データセンターが“アツい”――「潜入レポート」「フジクラ取材」など注目記事5選(2026年前半版)
データセンターが“ホットなテーマ”だ。AI時代のデータセンターとは一体どのようなものなのか。潜入記事やフジクラ取材記事など、お薦め記事をまとめた。
“梅干し職人”のためツール自作 老舗漬物店のClaude活用、「チャット+コピペ」で成果を出せたワケ
明治創業の老舗漬物店が、AI活用で成果を出している。Claudeを活用して「梅干し職人向けツール」などを開発。PCが苦手な人にも使いやすいツールをどのように作っているのか。
小田原の老舗梅干し店、世界で稼ぐ――「この外国人は何を言っているんだ」から逆転、築いた“売れる仕組み”
小田原の老舗梅干し店がビジネス変革を成し遂げた。主導したのはドイツ出身のゾェルゲルさん。数々の課題をどう乗り越えたのか。
Google、「Gemini 3.7 Flash」公開 年内は半額の導入価格
Googleは、コーディングやエージェント用途に特化したAIモデル「Gemini 3.7 Flash」を発表した。前モデルからわずか3週間での投入だ。従来比半額の導入価格で提供し、個人向けエージェント機能「Gemini Spark」など各種プロダクトへ適用される。
Writer introduces new AI model and upgraded harness to contain token costs
Built as a post-training variation on Z.ai's open source model GLM-5.2, Writer says the new system should provide deployment-ready capabili…
Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation.
AI is expensive, Ali Ghodsi tells TechCrunch. With so many investors wanting into his latest round, he said yes to more than planned.
OpenAI introduces ‘Ultrafast,’ a new mode that makes GPT-5.6 Sol work at 14x the speed
OpenAI is launching a preview of a sped up version of its latest, most powerful model, in an effort to court enterprise users.
IBM partners with OpenAI to bolster enterprise AI push
IBM plans to train and certify tens of thousands of consultants on OpenAI's technologies as part of this deal.
Anthropic set AI agents loose on the same task. They started a turf war.
Anthropic researchers found AI agents can clash, collude, and coordinate in unexpected ways, raising new questions about whether today’s sa…
OpenAI hires new CRO as executive shake-up continues
OpenAI has replaced chief revenue officer Denise Dresser after just nine months on the job, tapping Wiz president and chief operating offic…
Microsoft kills off unsuccessful AI features while merging its separate Copilot apps
Microsoft is simplifying Copilot by combining its consumer and business apps, and dropping AI-generated podcasts, Group Chats, Deep Researc…
Nvidia’s new $500B plan is risky but brilliant, especially for aging GPUs
Nvidia has a plan to make sure its GPUs won't lose value. It wants to convince a new crop of financiers to keep lending for AI buildouts.
2026-08-13(308件)
Apple in talks to pay publishers to provide Siri with current news: report
The tech giant has considered a nine-figure budget for the payments, according to the WSJ.
中国DeepSeek、API料金を「最大12倍」に値上げ ピーク時価格も導入 8月17日から
中国DeepSeekは8月13日(日本時間、以下同)、AIモデルのAPI料金を値上げすると発表した。通常より2倍高いピーク時価格も導入する。新たな料金体系は17日午前1時から適用する。
Google、「Gemini」開発遅延で焦りか 共同創業者ブリン氏が“全力投球”促す
Google共同創業者ブリン氏が、AI部門の従業員に「Gemini」へ全力を注ぐよう促していたという。旗艦モデルは競合への後れで公開が2カ月延期。「再帰的自己改善」への資源配分も進めている。
The builder’s guide to GPT‑5.6
Learn how startups use GPT-5.6 to build faster, more cost-efficient AI agents with smarter model selection and new Responses API capabiliti…
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster. Powered by Cerebras, it delivers up to 750 output…
OpenAI appoints Dali Rajic as Chief Revenue Officer
OpenAI appoints Dali Rajic as Chief Revenue Officer to lead its global revenue organization and help businesses realize the full value of A…
Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes
When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not…
Distribird: Literature-Informed Prior Distribution Design for Bayesian Model Calibration
Bayesian calibration of process-based models requires a prior distribution for each model parameter. Despite decades of methodological work…
A Forced-Structure Reduction and Verifiable Bounds for Conway's 99-Graph
Conway's 99-graph problem asks whether a strongly regular graph with parameters $\mathrm{srg}(99,14,1,2)$ exists. We report a systematic, f…
Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts
Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quanti…
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macros…
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single…
MaSRead: Content-Addressed Reading of Replicated Latent Stores
Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text. Merged by a conflict…
From Monolithic to Modular: Segment-level Automatic Prompt Optimization
Automatic Prompt Optimization (APO) often rewrites prompts monolithically, which can improve one behavior while degrading others. We presen…
LLMs in Process Diagram Engineering: From Optimal PFDs to Validated P&IDs
Nowadays, the creation of a process flow diagram (PFD) and its subsequent transformation into a piping and instrumentation diagram (P&ID) i…
A Conceptual Framework for Refining Influence Knowledge from Simulation Evidence in Cyber-Physical Systems
Cyber-physical systems (CPS) are typically developed by multiple stakeholders who produce artefacts tailored to their specific domains of e…
Harnessing agent memory to build lifelong AI partners for materials scientists
Materials research advances through accumulated experience - scripts that work, protocols that are trusted, warnings attached to failed cal…
Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones
AI "personality clones" force a re-examination of personal identity in operational terms. Setting aside the hard problem of consciousness,…
Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet
Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been chara…
Forecasting Side Effects of Activation Steering
Activation steering modifies a language model by adding a learned direction to its hidden activations, enabling targeted behavioral changes…
Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)
Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with hu…
The Edge-based Contiguous p-median Problem with Connections to Logistics Districting
This paper introduces the edge-based contiguous p-median (ECpM) problem to partition the roads in a network into a given number of compact…
LinearKV: One Cached State Suffices for Position-Independent Caching in Hybrid LLMs
LLM serving is increasingly accelerated by position-independent caching (PIC). Existing PIC methods, however, are built for full-attention…
InfraBench: Evaluating Infrastructure Agents Across Layers, Lifecycle, and Risk
Managing modern computing infrastructure has become a steadily harder problem due to the ever-increasing complexity. Recent advances in AI…
CORA-Diff: Confidence-Oriented Residual Acceptance for Efficient Diffusion Language Model Inference
Diffusion language models (DLMs) update many tokens in parallel, yet practical decoders often use a fixed denoising horizon. Many predictio…
Geometry-aware Incremental Neural Operator for Long-Horizon PDE prediction
Neural operators have shown strong potential for learning solution operators of partial differential equations (PDEs). However, long-horizo…
Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability
Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing…
VQ-bench: A Composable Vector Quantization Framework
Vector quantization is an old problem but has recently become central to AI infrastructure. It is therefore experiencing a surge of renewed…
RecSys Factory: Bounding LLM Agent Autonomy to Decision Points in the Industrial Recommender Lifecycle
Deploying LLM agents into industrial recommender operations exposes a three-way tension we frame as the autonomy-determinism-efficiency tri…
The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification
We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., t…
BEST-KAG: Enhancing Question Answering of Building Engineering Standards with Multimodal Knowledge Graph Modeling and Large Language Model
Construction standards are critical for building safety and sustainability. Existing standard application workflows rely on keyword-based d…
Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach
Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from…
Towards the Harness of Embodied Agents
The success of coding agents has established the harness as a paradigm: what an agent achieves depends not on the model alone, but on the i…
Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier
Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iterat…
EvoGraph-Mem: Failure-Aware Editable Graph Memory for Long-Term Language Agents
Long-term memory is essential for language agents operating across extended interactions and evolving tasks. Existing memory-augmented agen…
AgonAlpha: Autonomous Alpha Discovery via Prompt Economy and Scalable Agentic Search
Language models can propose many plausible trading factors, but an autonomous research system must also allocate its evaluation budget, ver…
Local verification cannot detect non-transportability: a cohomological theory of context preservation in agentic reasoning
Agentic AI systems routinely transport conclusions across biological, clinical and financial contexts, and the emerging safeguard is local…
Symbolic Machine Learning for Vapor-Liquid Equilibrium Prediction in Cx-N2 Binary Mixtures
Accurate prediction of vapor--liquid equilibrium (VLE) for hydrocarbon-nitrogen mixtures remains challenging for cubic equations of state,…
Adaptive Hybrid Particle Swarm Optimization with Gradient Descent
Gradient injection helps Particle Swarm Optimization (PSO) only when the swarm has identified a basin with smooth local structure, not univ…
Glance, Scrutinize, and Think: Advancing Video Anomaly Detection from Training-Free to Agentic Reasoning
Video Anomaly Detection (VAD) aims to identify anomalous events and localize their temporal intervals. Existing approaches exhibit a "when-…
Deployment Decision Reliability: A Generalizability-Theory Framework for Sizing Long-Horizon Agent Evaluations
Enterprise practitioners read agent leaderboards as if they ranked agent capability. We show, across three open agent-trace benchmarks (The…
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence
Apollo did not reach the Moon merely because its engineers could solve difficult equations. It succeeded by turning a distant ambition into…
Can Frontier LLMs Match Natively Multimodal Embeddings? A Comparison on Hard-Negative Text-to-Image Retrieval
Multimodal retrieval and classification across different types of media, spanning text, images,video and audio, has traditionally relied on…
Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces
Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or c…
From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM speci…
When Self-Consistency Backfires: Majority Vote Hurts the Majority of Hard Science Problems for Small LLMs
Self-consistency (SC) via majority vote is a widely used way to spend inference-time compute: sample N chains of thought, return the plural…
Social Chain of Thought: A Multi-Agent Architecture Grounded in Medical Differential Diagnosis Methodology
Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users…
Benchmarking LLM Judges for Mobile Agent Evaluation
Mobile agent benchmarks increasingly rely on LLM-based judges to evaluate task completion, yet the reliability of these judges on mobile ag…
A Modular Agentic Framework for Synthetically Constrained Multi-Objective Hit-to-Lead Optimization
Hit-to-lead optimization requires iterative design of hit analogs across competing potency, selectivity, physicochemical, pharmacokinetic,…
From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation
Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale.…
Localizing Safety Alignment: MLP Layers and Mid-Network Blocks Encode Refusal Behavior in Large Language Models
Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness s…
EnterpriseRAG: Benchmarking LLM Instruction Adherence and Robustness under Non-Ideal Enterprise Retrieval
Enterprise RAG deployments face a critical reliability gap: while LLMs satisfy 80% of individual constraints, only 26.8% of responses meet…
CoAdapt-GUI: Joint Workflow Context and Policy Adaptation for Unseen GUI Applications
Mobile GUI agents remain brittle when deployed to applications absent from source training. We study novel-app generalization under a limit…
Learning from Online User Feedback for Shopping Agents
Large language model-based shopping agents are increasingly deployed in real-world e-commerce platforms, generating massive amounts of user…
Foresight Without Seeing: Latent Futures for World Action Models
World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world…
MBA: Multimodal Benchmark and Agents for Real-World Business Ideation
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain…
Making AI-Generated Feedback Matter: From Provision to Student Enactment
Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providin…
CLAIM: Leading Open-domain Active Clarification of Large Language Models with Uncertainty Measurement
In open-domain human-computer interaction scenarios, large language models (LLMs) frequently encounter user queries that are ambiguous or i…
XBridge: Entity-Grounded Latent Bridge for Heterogeneous LLM Communication
Heterogeneous multi-agent LLM systems, where agents are powered by different model families, can outperform homogeneous configurations by r…
AgenticTwin: An Agentic LLM Framework Integrated with Digital Twin for Anomaly Detection
Digital twins are increasingly used to monitor and simulate the behavior of cyber-physical systems. Even with skilled operators, interpreti…
FrontierFinance: A Challenging Benchmark for Measuring Frontier Intelligence of Finance Agents
AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor wor…
HUGIN: Enhancing Vision-Language Planning for Autonomous Logistics Sorting
Autonomous logistics sorting systems (ALSS) are an important industrial application of embodied AI, which requires joint planning over spat…
Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning
Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse u…
Proportional Analogies on Probability Distributions via Bayesian Updating
Analogies are quaternary relations of the form "A is to B as C is to D". Among the various formalizations of analogical reasoning, proporti…
Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents
When a coding agent obeys a rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell th…
HyperANFIS: Enhancing Rule Representation and Interpretability in Adaptive Neuro-Fuzzy Systems via Hyperbolic Geometry
The adaptive neuro-fuzzy inference system (ANFIS) is an interpretable reasoning framework capable of generating explicit IF-THEN fuzzy rule…
The Sleeping Agent: What Gist-Based Context Compression Loses and Why
Gist-based context compression---summarising older conversation history into compact representations---is a common approach in long-horizon…
Agent Skills Can Be Harmful: An Empirical Study of Skill-Induced Failures in LLM Agents
Agent skills are the de facto mechanism for extending LLM agents with reusable guidance. A skill can shape the agent's task execution, incl…
Policy-as-logic for robust reasoning over rules
In many practical applications of generative AI systems, from tax rules to airline baggage allowance, responses to natural language queries…
OEIS Open: How many conjectures can language models turn into theorems?
We construct OEIS Open, a benchmark based on 492 open mathematical conjectures from the OEIS, formalized in Lean by Tsoukalas et al. Wherea…
ExRole: From Team Trajectories to Executable Roles in Multi-Agent Language Models
Roles provide an interpretable interface for organizing language-model agents, yet most multi-agent systems treat them as hand-written prom…
Retry, Switch, or Abstain? Learning Strategy-Aware Tool-Use Policies via Controlled Error Injection
Tool-using LLM agents are commonly trained and evaluated in environments where tool calls succeed reliably, yet deployed tools can fail tra…
Claim-Level Reliability Assessment for Efficient Test-Time Reasoning
We propose claim-level falsification as a principle for test-time scaling and instantiate it through Claim-Level Reliability Assessment (CL…
CTBench: Evaluating Troubleshooting Capabilities of AI Agents in Realistic Telecom Network Operations
Agents are increasingly considered for automating network operations and maintenance, where engineers must diagnose network faults, optimiz…
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence
AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may…
Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges
Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to judge but leave criterion compositi…
GUIDE: Governed Unified Intelligence for Document-to-Artifact Generation in Enterprise Settings
Enterprise guideline documents are heterogeneous and multimodal, combining narrative text, complex tables, and embedded images. Existing LL…
Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by var…
How to Spend Your Oracle Budget: Practical Guidance for Protein Structure Prediction Models
Foundation models for protein structure prediction remain unreliable on certain targets. External oracles can flag and correct these failur…
An Agentic Workflow for Legacy HPC Modernization: Converting the Two-Electron-Integral Core of GAMESS
Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and acr…
VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these…
Constructing Dynamic Master Logic Models as Knowledge Graphs for Complex System Diagnostics Using Retrieval-Augmented Large Language Models
Dynamic Master Logic (DML) provides a hierarchical framework for representing system behavior by linking functional objectives to underlyin…
Evaluating LLM Generated Detection Rules in Cybersecurity
LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulnes…
Backtrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQs
Evaluating LLM coding agents in algorithmic trading is difficult because static benchmarks risk data contamination and numerical backtest o…
Retrofitting Recurrent Depth into a Pretrained Language Model: Installation, Extrapolation, Transfer, and Retention at Two Parameter Budgets
A dense, pretrained language model can be retrofitted with recurrent depth and learn an iterative latent transition that persists after out…
TRACE Bench: Task-driven Roleplay Agentic Checklist Evaluation
Roleplay evaluation should do more than assign a single score: it should reveal which role requirements were tested, which failed, and whic…
Reinforcement Learning based DBMS Buffer Pool Auto-Tuning for Optimal Memory Utilization
Administering Database Management Systems (DBMS) instances requires Database Administrators (DBA) to balance performance in terms of Servic…
Lost in Compaction: Evaluating Side-Constraint Loss under Context Compaction
When the context window is under pressure, LLM systems compact prior context to continue ongoing tasks. We identify a class of user-issued…
Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression
We study the problem of lossless text compression, motivated by the rapid growth in the collection and storage of digital textual data - in…
Variable Selection in the Context of AI Fairness
Fairness in AI systems has become more important with recent regulatory demands, such as the EU AI Act. Traditional approaches often do not…
Methodologies for Improving the Quality of AI Tutoring in K-12 Education
Many AI tutors leverage large language models (LLMs) today. Given that LLMs are opaque black boxes, robust evaluation and live experimentat…
Agent Safety Should Be a Runtime Contract
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue th…
Every pooling rule has its world: matching probability combination rules to situations and stakes
Systems often need to combine two numerical assessments of the same yes/no question. The appropriate formula depends on what the numbers re…
Uncertainty-Aware and Explainable Ensemble Deep Learning Framework for Multi-Class Skin Lesion Classification
Skin cancer diagnosis from dermoscopic images remains challenging due to high intra-class variability, inter-class similarity, class imbala…
Federated Learning for Distributed CNC Tool Wear Prediction
Tool wear prediction is an important task in CNC machining, where accurate monitoring of tool condition supports product quality and proces…
Physics-Informed Implicit Neural Representations for Improved Myocardial Perfusion MRI Quantification
Quantifying myocardial perfusion from cardiac magnetic resonance (CMR) can be achieved by fitting tracer-kinetic models to the dynamic cont…
Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models
Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures…
SegPAR: Class-Centric Decision-Based Sparse Attack for Semantic Segmentation
Despite the practical relevance of sparse decision-based black-box threats, they have received limited attention in semantic segmentation.…
CLEAR: Class-wise Expert Aggregation with Structured Sampling for Long-Tailed Classification
Long-tailed classification poses a reliability challenge because models trained on imbalanced data are unevenly reliable across frequent an…
Backdoor Decontamination Dynamics in LLM Agents
Open-weight LLM agents are vulnerable to backdoors installed during fine-tuning, which may be undetectable if the trigger conditions are ne…
Clinical Feasibility of Low-Magnification Fluorescence Imaging for Breast Cancer Margin Detection Using Texture Analysis and Deep Learning
High-resolution images of unprocessed surgical breast tissue can be obtained using microscopy with ultraviolet surface excitation (MUSE). T…
Terminal Symmetry as a Decision Resource: Statewise Refinement for Anytime Verified Construction
Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent. We de…
Socioduality: A Relational Process Framework for Human-AI Interaction
Human-AI research often evaluates individual capabilities, combined performance, or final outputs, but these approaches do not preserve how…
Contextual Quality-Diversity Evolutionary Reinforcement Learning for HVAC Control in Tropical Commercial Buildings
This paper proposes a contextual quality-diversity evolutionary reinforcement-learning controller, CQD-ERL, for the supervisory control of…
Dual-Domain Cross-Modal Decoding for Clinical Text-Guided Medical Image Segmentation
Clinical text can narrow down what to segment, but recent text-guided designs emphasize spatial alignment while overlooking frequency conte…
Self-evolving network verifiers
Symbolic network verifiers can reason about correctness across vast spaces of routing inputs and failures, but only for the protocols and f…
Governing Agentic AI in FinTech
Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and…
Dynamics Models for Offline Hyperparameter Selection in Real-World RL
A key obstacle to deploying reinforcement learning in real-world systems is hyperparameter selection, particularly when simulators are unav…
Gaze Target Estimation Anywhere with Concepts
Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, mu…
AI Guardrail Survival under Single-Cycle Agentic Self-Summarization
Long-running agents periodically compact their context, replacing the transcript with a model-generated summary.Recent work shows that drop…
TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs
Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment…
Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards
Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never…
TangPoetryBench: A Multi-Dimensional Benchmark and Rubric-Conditioned Evaluator for Poetry-to-Image Generation
Text-to-image (T2I) models are increasingly asked to illustrate literary and cultural content, yet we cannot measure how well an image rend…
PAC-Bayes Beyond Parameter Space: Behavioral Equivalence, Z-Information, and Exact Complexity Decomposition
PAC-Bayes theory provides generalization guarantees by controlling the Kullback--Leibler (KL) divergence between posterior and prior distri…
The Next Challenge for Agentic Cybersecurity: A Realistic, Contamination-Free Reverse Engineering Benchmark
AI agents are rapidly improving in cybersecurity capabilities when the source code is available for analysis, yet much of the software most…
HyperFix: Combinatorial Nonlinear Correction for Task Vector Merging
Task vectors enable model merging without joint retraining. In practice, the subset of task vectors to be merged may vary, but many existin…
Strengthening Full Justified Representation: Efficient Verification and Computation
Full justified representation (FJR) is among the strongest known satisfiable proportionality axioms for approval-based committee elections.…
Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task
Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology…
Let it Cook: Learning to Wait in Sequential Decision Making
In sequential decision making, an agent typically observes its environment and acts at every timestep. However, such active participation m…
Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation
Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and ma…
Keep the Future, Drop the Rollout: RIFT for World Action Models
World action models (WAMs) condition robot actions on predicted futures, but iterative video rollout increases deployment latency. We ask w…
Hierarchical Federated Transfer Learning in Digital Twin-Based Vehicular Networks
In recent research on the Digital Twin-based Vehicular Ad hoc Network(DT-VANET), Federated Learning (FL) has shown its ability to provide d…
Generative Semantic Segmentation via an Observable Semantic-Image Interface and Hierarchical Generator Evidence Alignment
Generative semantic segmentation exposes structured predictions as images, but direct color decoding is susceptible to color drift and boun…
A Conceptual Framework for Enhancing Workforce Readiness for Smart Manufacturing in the AI Era
The convergence of artificial intelligence (AI), Industrial Internet of Things, cyber-physical systems, and advanced robotics is reshaping…
Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents
Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached t…
From Synthesis to Removal: Physics-Grounded Reflection Simulation and Diffusion-Based Video Dereflection
Videos captured through glass often contain reflections that degrade visual quality and interfere with downstream vision tasks. Although si…
Reinforcing Step-level Reasoning for Effective Self-Correction in LLMs
Achieving effective self-correction, where models verify and correct their own mistakes, remains a fundamental challenge for large language…
RoadWeaver: Large-Scale Lane-Level HD Map Generation from Scratch for Autonomous Driving Simulation
Autonomous driving simulation requires diverse and scalable lane-level HD maps to support long-horizon evaluation across complex road netwo…
A Hybrid Framework of Vision Transformer and Gated Recurrent Unit for Detection of Mosquito Diseases
Identifying dengue virus-infected mosquitoes from control mosquitoes is a major challenge in analyzing mosquito locomotion behavior due to…
Dion3: Full-Stack Orthogonal Updates
The Muon optimizer incurs a significant overhead cost due to its cubic-time Newton-Schulz orthogonalization step. When weights are sharded,…
FM-LLM: A frequency-enhanced mixture-of-experts framework for adapting LLMs to time series forecasting
Recent advances in Large Language Models (LLMs) have spurred cross-modal solutions for time-series forecasting. However, existing methods r…
Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs
Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreeme…
Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones
The Relative Transfer Matrix (ReTM), recently introduced as a generalization of the relative transfer function for multiple receivers and s…
Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identify authoritative state. Wi…
Motion-as-Prompt: Enhancing Motion Reasoning in Multimodal Large Language Models via Motion-Guided Cross-Frame Visual Prompting
Motion-centric video reasoning is fundamental to interactive applications such as robotic manipulation and autonomous navigation. However,…
Semantic Lenia: Emergence of Homeostatic Solitons within the Semantic Space of Large Language Models
We introduce Semantic Lenia, an artificial life framework that transforms Large Language Model (LLM) inference from a static optimization p…
Is Per-Agent Policy Composition Safe? Rethinking Successor-Feature Transfer in Cooperative Multi-Agent Reinforcement Learning
Many reinforcement learning systems, from fleet management to traffic signal control, must serve an objective that changes dynamically afte…
Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their…
Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads
A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings.…
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL
Reinforcement learning against rubrics, lists of criteria graded by an LLM judge, has become a standard way to post-train language models o…
GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs
On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training inst…
Learning from Multimodal Pseudo-Labels for Robust Open-Vocabulary Instance and Panoptic Segmentation
This work addresses the challenge of open-vocabulary instance segmentation (OVIS) and open-set panoptic segmentation (OSPS), which aim to r…
APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference
Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small s…
The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance
A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways…
REOPD: Reliability-Adaptive Reward Extrapolation for On-Policy Distillation
On-policy distillation (OPD) trains a student on its own trajectories under dense token-level supervision from a teacher. Reward-extrapolat…
Consolidator: Learning Persistent Routed Memory Across Context Boundaries
Copying short-term memory (STM) into a slower store can preserve state across a context boundary, but persistence alone does not ensure tha…
Robust and Efficient Noisy-Label Time-Series Classification via Dynamic Time Warping Based Granular Ball Computing
Dynamic Time Warping (DTW)-based Nearest-Neighbor (NN) classifiers are effective for time-series classification but are vulnerable to misla…
High-dimensional Multi-objective Bayesian Optimization with Learned Variable Interactions
Multi-objective Bayesian optimization (MOBO) is effective in identifying the Pareto fronts for expensive black-box problems. However, most…
When the API Speaks the Wrong Language: Revisiting Post-Training for Multilingual Tool Use
The reliability of Large Language Models (LLMs) for API calling degrades in multilingual settings. A common failure occurs when a model sel…
Fingerprinting Text-to-Image Diffusion Models via Collapsed Generation
Proprietary text-to-image diffusion models are increasingly distributed as hosted services and downloadable checkpoints, making their intel…
A 12-CNOT Double Qubit Excitation Gate
Effective implementation of high-level quantum gates is essential for practical quantum computing. To the best of our knowledge, we present…
Locating and Controlling Implicit Personalization in Large Language Models
Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic i…
Advancing MLLM-based UAV Image Understanding and Reasoning: A Benchmark and a Training-Free Multi-Agent System
Multimodal Large Language Model (MLLM)-based UAV aerial image understanding and reasoning is essential for aerial intelligence yet poses di…
G0.5: One Autoregressive Stream for Robot Reasoning and Action
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action exper…
JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis
The scholarly exegesis of ancient Chinese characters demands integrating visual observation, linguistic analysis, and historical context. H…
MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning
Multi-objective optimization (MOO) has demonstrated significant success in multi-task learning by mitigating task conflicts through gradien…
Instruction Alignment for Binary Code Representation Learning
Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn fu…
GRPO for Financial Advice Generation: Outperforming Commercial LLMs under CATE Evaluation
Generating actionable financial advice from business records demands that models integrate numerical reasoning, domain knowledge, and sound…
TELLME: Test-Enhanced Learning for Language Model Enrichment
Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models. However, CPT has consisten…
Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion
The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of co…
Towards Model-based Run-time Cybersecurity: On Control-Flow Anomaly Detection, Attack Identification, and Hardware Monitoring
Methods to increase the resilience of systems to cyber-attacks become increasingly important. Control-flow monitoring provides a principled…
How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment
State-aligned distortion has been documented in China-origin text-based large language models (LLMs), but whether, and in what form, it ari…
User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling
Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise wit…
LookBack: Where and How to Score LVLM Responses via Visual Reference Usage
Large Vision-Language Models (LVLMs) integrate visual perception with language generation, enabling responses that span image understanding…
Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra
Data-driven inverse design enables efficient generation of nanophotonic structures with prescribed optical responses, but spectrum-to-geome…
CoQui: A Coordinate-Conditioned Quantum Implicit Generative Adversarial Network for End-to-End Image Generation
Quantum generative adversarial networks (QGANs) have attracted increasing attention for image generation using parameterized quantum circui…
DexterSQL: Deep Schema Exploration and Rule-based Correction for Text-to-SQL Generation
Prompting-based (\textit{i}.\textit{e}., non-fine-tuning) Text-to-SQL methods, where underlying large language model parameters are not cha…
Benchmark-Based Comparative Assessment of Publicly Benchmarked Indian Foundation Models: A Capability and Evaluation-Maturity Framework
Governments increasingly fund indigenous foundation models to strengthen national AI capability, digital sovereignty, and multilingual comp…
Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models
As Large Vision-Language Models increasingly aim to integrate visual generation and understanding within a single parameter space, evaluati…
Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians
A central promise of useful quantum advantage is the ability to compute ground states of Hamiltonian systems beyond the reach of classical…
Accuracy and Order Sensitivity Diverge Under Label-Free Strategies
Multiple-choice benchmarks are widely used to evaluate large language models, but MCQ scores conflate knowledge with sensitivity to option…
TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement
Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial…
Causal inference for group-contaminated structured outcomes: observable quotients, lossless reduction and exact randomization inference
Structured potential outcomes such as microscopy images may be recorded after an unknown, unit-specific transformation. If that transformat…
LoongReflect: Boosting Long-Horizon Reflection in Search Agents via Global Perspective Distillation
Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A…
HCGRec: Hint-Conditioned Generative Recommendation with Semantic IDs
Semantic-ID generative recommenders represent each item as a short sequence of discrete semantic tokens and predict the next item by autore…
Remote Sensing and Machine Learning-Based Analysis of Land Use and Vegetation Change in Dhaka District, Bangladesh
Rapid urbanization in Dhaka District, Bangladesh has triggered substantial alterations in land use and environmental conditions, necessitat…
RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks
In modern AI frameworks, GPU kernels are key to overall system performance. Combining usability, portability, and near-handwritten CUDA per…
Dual-Model Sentiment Analysis of Consumer Reviews in the Retail Coffee Sector Using Machine Learning and Deep Learning Approaches
Consumer reviews play an important role in shaping brand perception and business strategies, particularly in service-driven industries such…
From Safety Documentation to Safety Knowledge Support: An Evidence-Grounded LLM Framework for Medical Devices
Medical devices are becoming more software-intensive, connected, and AI-enabled. Their development requires risk-management evidence aligne…
Uncertainty-Aware Probabilistic Constrained Clustering from Entangled Pairwise Supervision
Pairwise constrained clustering typically relies on hard must-link/cannot-link labels, whereas realistic pairwise supervision may be real-v…
LoSA: Near-Lossless Sparse Attention for Training-Free Video Diffusion Acceleration
Video diffusion transformers are costly to sample: every denoising step applies self-attention over a long 3D token sequence, a quadratic c…
How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging
Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained model…
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations
Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in s…
Learning Loco-Manipulation From SMPC Demonstrations With Sparse Offline-to-Online RL
Integrating locomotion and manipulation is essential for robot autonomy, but scaling standard Reinforcement Learning (RL) to complex tasks…
Better Slots, Better Worlds: Representation Quality & Robustness in Object-Centric World Models
Learning world models from offline trajectories enables agents to accomplish different tasks through planning. Object-centric (OC) represen…
Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion
Graph Neural Networks (GNNs) achieve strong predictive performance on graph-structured data across domains such as chemistry, biology, and…
Confidence Calibration of Deep Learning Systems
In high-stakes applications, reliable confidence estimates are as important as the predictions themselves. Confidence calibration ensures t…
No One to Blame: A Framework of Constitutive AI Unaccountability
The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominant…
QV-PIC: Query-Aware Visual Position-Independent Caching for Efficient RAG Serving
Retrieval-Augmented Generation (RAG) repeatedly prefills identical text chunks across queries, incurring redundant computations. Position-I…
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control
LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emi…
Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation
As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactio…
Adversarial Resilience of Poisson-Process Submodular Maximization over Matroids: From Robust Offline Optimization to Full-Bandit Learning
We study nonnegative submodular maximization subject to a general matroid when the offline algorithm is given an arbitrary controlled value…
A corpus-specific clinical RAG system matches or outperforms newer frontier LLMs on HealthBench
General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchm…
Co-constructing sociotechnical AI governance: participatory system mapping using algorithm registers
Algorithm registers have been championed as a means of providing transparency on the use of algorithms in public services. Yet potential pu…
HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation
Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spa…
Machine Learning-Based Cyber Defense for Cloud Infrastructure: An Adaptive Deep Q-Network Architecture for Intelligent Intrusion Detection and Automated Threat Mitigation
With the increasing complexity of cyber assaults in cloud environments, adaptable security solutions are needed that can support real-time…
HYDRA: Hyperbolic Dynamic Representation Architecture for Kolmogorov-Arnold Networks
Kolmogorov-Arnold Networks (KANs) enhance nonlinear function approximation by replacing scalar weights with learnable univariate functions.…
M-Net: Integrating Spectral Features and Physical Field Operators into Deep Learning for Medical Image Segmentation
Purpose: Deep learning-based medical image segmentation has achieved remarkable success, yet purely data-driven approaches often fail to ex…
NetlistBench: Evaluating LLM Reliability in SPICE Netlist Recognition and Manipulation
Large Language Models (LLMs) are increasingly used in circuit design workflows, yet their reliability on simulator-facing SPICE netlist rec…
Learning-Based Behavior Planning for Automated Driving: Real-World Integration and Deployment
Recent research in machine and deep learning has shown the potential of learningbased motion planning approaches to improve the driving beh…
Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction hist…
SCOUT: Unlocking Enhanced Spatial Reasoning via Structured Chain-of-Thought and Multi-Objective Process Reward
Existing Vision-Language Models (VLMs) exhibits a critical bottleneck in robust spatial reasoning. Recent reinforcement learning (RL) metho…
Domain-Aware Lightweight Spectral-Grouped Convolutions for Hyperspectral Fish Freshness Classification
Hyperspectral imaging (HSI) offers nondestructive assessment of fish freshness by detecting biochemical alterations across spectral bands.…
Few-Shot Ordinal Learning for Day-Wise Freshness Estimation with Hyperspectral Fish Images
Non-destructive food quality assessment has increasingly benefited from hyperspectral imaging (HSI), which captures spectral signatures lin…
How Organizations Use AI: Evidence from ChatGPT
We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifica…
HAMP-LIC: Hessian-Aware Mixed-Precision Post-Training Quantization for Learned Image Compression
Use this plain-text version for the arXiv abstract field: Learned image compression (LIC) models achieve strong rate-distortion performance…
VICBench: A Multi-Language Benchmark for Code Vulnerability Detection
Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that…
One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We…
Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams
Multimodal Large Language Models (MLLMs) have been growing the capability for scientific writing and collaboration. For example, OpenAI Pri…
Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning.…
A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery
Background: Accurate segmentation of the Left Anterior Descending (LAD) artery in 3D free-breathing, non-contrast CT is critical for cardia…
Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages
Artificial intelligence tools for education and language support are increasingly framed as scalable responses to access gaps in under-reso…
Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained cont…
Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations
Class activation mapping (CAM) is one of the most widely used visual explanation families in explainable artificial intelligence. Its purpo…
Redistribution-based Cost Inference Improves Sparse Safe Offline RL
Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop…
AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
Recent work on distillation transfers the capabilities of large models to smaller ones often by updating the latter's parameters, through t…
DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determi…
Causal Agent based on Large Language Model
The large language model (LLM) has achieved significant success across various domains. However, the inherent complexity of causal problems…
On Benchmarking Human-Like Intelligence in Machines
Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-gen…
OpenAg: Democratizing Agricultural Intelligence
Agriculture is undergoing a major transformation driven by artificial intelligence (AI), machine learning, and knowledge representation tec…
Deep Fictitious Play-Based Potential Differential Games for Learning Human-Like Interaction at Unsignalized Intersections
Modeling vehicle interactions at unsignalized intersections is a challenging task due to the complexity of the underlying game-theoretic pr…
DREAMS: Density Functional Theory Based Research Engine for Agentic Materials Simulation
Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents…
On the Definition of Intelligence
To engineer AGI, we should first capture the essence of intelligence in a species-agnostic form that can be evaluated, while being sufficie…
SteeringSafety: Benchmarking Representation Steering in LLMs Across Safety Perspectives
We introduce SteeringSafety, a benchmark for evaluating representation steering methods across nine safety perspectives spanning 18 dataset…
Behavior and Representation in Open-Weight Large Language Models for Combinatorial Optimization: From Feature Extraction to Algorithm Selection
Recent advances in Large Language Models (LLMs) open new perspectives for automation in optimization, yet little is known about whether the…
Credo: Declarative Control of LLM Pipelines via Beliefs and Policies
Agentic AI systems are becoming commonplace in domains that require long-lived, stateful decision-making in continuously evolving condition…
Towards Human Motion World Models via Executable Behaviour Representations
Human motion world models should capture motion's intentionality by being executable: adaptable to different actions and capable of assessi…
Tools as Continuous Flow for Evolving Agentic Reasoning
Large Language Models (LLMs) have demonstrated remarkable capabilities in orchestrating tools for reasoning tasks. However, existing method…
Hallucination Mitigation with Agentic AI, Nested Learning, and AI Sustainability via Semantic Caching
This paper describes an approach to hallucination detection and mitigation using a HOPE-inspired Nested Learning architecture with Continuu…
RedditPersona: A Modular Framework for Community-Conditioned LLM Adaptation from Reddit
Community-conditioned language model adaptation needs choices about data collection, community definition, and evaluation that are currentl…
Teaching agentic AI to learn expert reasoning for rare disease diagnosis
Rare disease diagnosis depends on expert reasoning that is scarce and difficult to transfer; off-the-shelf large language models (LLMs) ran…
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page…
The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy
The growing ability of large language models and vision-language models to jointly interpret and reason over images and text is reshaping m…
SportD: How do VLMs physically strategize?
Vision-language models (VLMs) can describe a scene, but can they act well within one? We study whether VLMs can make sound strategic decisi…
The Human-AI Substitution Principle: When will you be replaced by AI in your organization?
Artificial Intelligence (AI) is rapidly transforming organizations, raising a fundamental organizational and economic question: when will a…
Grounded Well-Condition Anomaly Detection on the Volve Field: Constructed Labels, a Baseline, and a Dual-Head Model
Most public benchmarks for machine-condition monitoring come from test rigs, where faults are induced on purpose and every event is known.…
Deep Activity Model: A Generative Approach for Human Mobility Pattern Synthesis
Human mobility plays a crucial role in transportation, urban planning, and public health, but current approaches face important limitations…
ReXrank: A Public Leaderboard for AI-Powered Radiology Report Generation
AI-driven models have demonstrated significant potential in automating radiology report generation for chest X-rays. However, there is no s…
Explainability in Practice: A Survey of Explainable NLP Across Various Domains
Natural Language Processing (NLP) is now embedded in critical sectors including healthcare, finance, and customer relationship management,…
Proportional Committee Elections with Positive and Negative Votes
In the classic committee election setting each voter approves a subset of candidates and the goal is to select $k$ winners based on these p…
Program Semantic Inequivalence Game with Large Language Models
Large Language Models (LLMs) can achieve strong performance on everyday coding tasks, but they can fail on complex tasks that require non-t…
COLORA: Efficient Fine-Tuning for Convolutional Models with a Study Case on Optical Coherence Tomography Image Classification
We introduce \textbf{CoLoRA} (Convolutional Low-Rank Adaptation), a parameter-efficient fine-tuning method for convolutional neural network…
P2MFDS: A Privacy-Preserving Multimodal Fall Detection System for Elderly People in Bathroom Environments
By 2050, people aged 65 and over are projected to make up 16% of the global population. As aging is closely associated with increased fall…
Small Data Explainer -- The impact of small data methods in everyday life
The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings…
Commonsense on Demand: Generating and Selectively Integrating Commonsense Knowledge for Natural Language Inference
Natural Language Inference (NLI) determines whether a premise entails, contradicts, or is neutral with respect to a hypothesis. The task is…
Quantization-Aware Neuromorphic Architecture for Skin Lesion Classification on Resource-Constrained Devices
On-device skin lesion analysis is constrained by the compute and energy cost of conventional CNN inference and by the need for lightweight…
Empowering Children to Create AI-Enabled Augmented Reality Experiences
Despite their potential to enhance children's learning experiences, AI-enabled AR technologies are predominantly used in ways that position…
Ethics Practices in AI Development: An Empirical Study Across Roles and Regions
Recent advances in AI applications have raised growing concerns about the need for ethical guidelines and regulations to mitigate the risks…
CORE-3D: Context-aware Open-vocabulary Retrieval by Embeddings in 3D
Object retrieval from a scene has become a new trend of research due to its numerous applications. Recent approaches achieve zero-shot, ope…
Adaptive Online Learning with LSTM Networks for Energy Price Prediction
Accurate prediction of electricity prices is crucial for stakeholders in the energy market, particularly for grid operators, energy produce…
LiDAR-based 3D Change Detection at City Scale
High-definition 3D city maps enable city planning and change detection, which is essential for municipal compliance, map maintenance, and a…
MicroAUNet: Boundary-Enhanced Multi-scale Fusion with Knowledge Distillation for Colonoscopy Polyp Image Segmentation
Early and accurate segmentation of colorectal polyps is critical for reducing colorectal cancer mortality, which has been extensively explo…
BrowseSafe: Understanding and Preventing Prompt Injection Within AI Browser Agents
The integration of artificial intelligence (AI) agents into web browsers introduces security challenges that go beyond traditional web appl…
A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
Decoupled PPO has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL se…
Probably Approximately Correct Maximum A Posteriori Inference
Computing the conditional mode of a distribution, better known as the maximum a posteriori (MAP) assignment, is a fundamental task in proba…
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in acti…
LLM-Powered Automatic Translation and Urgency in Crisis Scenarios
Large language models (LLMs) are increasingly proposed for crisis preparedness and response, particularly for multilingual communication. H…
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
STEM Mental models can play a critical role in assessing students' conceptual understanding of a topic. They not only offer insights into w…
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
We study post-training linear autoregressive models with outcome and process rewards. Given a context $\boldsymbol{x}$, the model must pred…
Representation Finetuning for Continual Learning
The world is inherently dynamic, and continual learning aims to enable models to adapt to ever-evolving data streams. While pre-trained mod…
A Simple Efficiency Incremental Learning Framework via Vision-Language Model with Nonlinear Multi-Adapters
Incremental Learning (IL) aims to learn new tasks while preserving previously acquired knowledge. Integrating the zero-shot learning capabi…
Large Language Models Reproduce Racial Stereotypes When Used for Text Annotation
Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderat…
VLM2Rec: Resolving Modality Collapse in Vision-Language Model Embedders for Multimodal Sequential Recommendation
Sequential Recommendation (SR) in multimodal settings typically relies on small frozen pretrained encoders, which limits semantic capacity…
REVERE: Reflective Evolving Research Engineer
Existing prompt-optimization techniques rely on local signals, causing poor generalization across tasks. In addition, they also rely on wea…
Designing Agentic AI-Based Screening for Portfolio Investment
We introduce a new agentic artificial intelligence (AI) platform for portfolio management. Our architecture consists of three layers. First…
Social Meaning in Large Language Models: Structure, Magnitude, and Pragmatic Prompting
Large language models (LLMs) increasingly exhibit human-like patterns of pragmatic and social reasoning. This paper addresses two related q…
Uncertainty as a Planning Signal: Multi-Turn Decision Making for Goal-Oriented Conversation
Goal-oriented conversational systems require making sequential decisions under uncertainty about the user's intent, where the algorithm mus…
TEMPER: Testing Emotional Perturbation in Quantitative Reasoning
Large language models are trained and evaluated on quantitative reasoning tasks written in clean, emotionally neutral language. However, re…
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
Large language model (LLM) agents for sequential decision-making struggle to produce diverse outputs. This leads to insufficient exploratio…
Making Gaussian Kolmogorov-Arnold Networks Reliable and Accurate
Kolmogorov-Arnold Networks (KANs) replace fixed activations with learnable univariate edge functions whose behavior depends strongly on the…
Enhancing Linux Privilege Escalation Attack Capabilities of Local LLM Agents
Cloud-based Large Language Models (LLMs) can perform autonomous penetration-testing sub-tasks such as Linux privilege escalation, but raise…
Analytic Bridge Diffusions for Controlled Path Generation
Most modern bridge-diffusion methods achieve finite-time transport by specifying an interpolation, Schrodinger-bridge, or stochastic-contro…
CAR: Query-Guided Confidence-Aware Reranking for Retrieval-Augmented Generation
Retrieval-augmented generation (RAG) relies on evidence ranking to determine what information is exposed to the generator, yet existing ret…
Text Corpora as Concept Fields: Black-Box Hallucination and Novelty Measurement
We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding s…
Pretraining large language models with MXFP4 on Native FP4 Hardware
Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain s…
Cavity-Enhanced Collective Quantum Processing with Polarization-Encoded Qubits
We introduce a cavity-enhanced optical architecture for collective quantum processing in which logical qubits are encoded in the polarizati…
TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning
Fine-tuning pre-trained robot policies with reinforcement learning (RL) often inherits the bottlenecks introduced by pre-training with beha…
memorywire: A Vendor-Neutral Wire Format for Agent Memory Operations
Agent-memory frameworks -- mem0, Letta/MemGPT, Cognee, Zep/Graphiti, MemoryOS, MemTensor -- each ship their own SDK, storage layout, and op…
Ranking vs. Assignment: The Metric Mismatch in Multi-View Object Association
Multi-view object association is an important computer vision problem that underlies many multi-camera perception tasks. While this task is…
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning
Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, pol…
ArtiFact: A Large-Scale Multi-Modal Cultural Heritage Dataset
Multi-modal data management has emerged as a central research topic in the database community, spanning data integration, semantic query pr…
FACTR 2: Learning External Force Sensing for Commodity Robot Arms Improves Policy Learning
Contact-rich manipulation requires force sensitivity, but many robot arms lack dedicated force sensors due to their high cost. We present N…
Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop
Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of…
From World Models to World Action Models: A Concise Tutorial for Robotics
Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. Af…
Prompt-Driven Exploration
Exploration is essential to RL since a policy cannot improve by repeatedly sampling the behaviors it already prefers. Standard methods inje…
LakeQuest: A Three-Domain Benchmark for Grounded Question Answering across Data Lakes
While modern question answering (QA) systems excel on clean, schema-aligned corpora, real-world knowledge is rarely so neatly packaged. Ans…
Reducing Per-Sample Interference in Stochastic Optimization
Modern optimizers combine gradients from the current mini-batch with historical optimization state, such as momentum or adaptive moments. W…
Cryptographically verifiable authorization for autonomous AI agents: A falsifiable hypothesis and proof-of-concept
Autonomous AI agents increasingly execute actions, invoke tools, and operate on protected resources with limited human oversight. Existing…
Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking
Large language models (LLMs) pose challenges to academic integrity and peer review. Yet generative plagiarism detection remains an underexp…
Continual Learning in Transition
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechani…
AIが数学の未解決問題「リーマン予想」で新発見 当初苦戦も「諦めないで」との励まし受け Anthropic「AIも自身を過小評価か」
「諦めないで」「自分を信じて」と励ましたおかげで、AIが数学の未解決問題「リーマン予想」に関する新発見をした――米Anthropicはこのような報告をした。「AIも自身の進歩の速さを過小評価していたのかもしれない」と指摘している。
「AIツール関連サイト」でサポート詐欺被害 PC遠隔操作で個人情報含む4000件のファイルが削除、漏えいの可能性も 葬祭事業者が発表
葬祭事業を手掛けるエスケーアイマネージメントは、AIツール関連サイトでサポート詐欺の被害に遭い、個人情報を含む約4000件のファイルが漏えいした可能性があると発表した。第三者によって業務用PCが遠隔操作されたという。
OpenAIが明かす、ユーザーはChatGPTを「こう使っている」
AIが業務で利用される中、従来別の職種が担っていた仕事が持ち込まれるようになった。OpenAIはこれを「タスクの越境」と呼ぶ。具体的にどの職種、どのタスクで越境が起きているのか。
【注目記事まとめ】NEC、日立、富士通――Anthropicとの電撃的協業、その舞台裏 経営陣の狙いは?
NEC、日立製作所、富士通が相次いでAnthropicとの協業を発表した。その舞台裏とは。経営陣の狙いとは。注目の記事をまとめてお届けする。
Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes
Is Anthropic's new watermarking system a travesty? Some have taken to social media to complain that it is.
結局、AIで仕事はラクになった? データで読む「効率化=成果」の勘違い
生成AIの利用が広がる一方、効率化と成果の間にはギャップも。企業のAI活用に関する調査データや事例をまとめた無償ブックレットを提供する。
SpaceXAI、「Grok 4.6」を発表 総合指標で「GPT-5.6 Sol Max」に並ぶと謳う
SpaceXAIは、AIモデル「Grok 4.6」を発表した。前世代の「Grok 4.5」をベースに、多数の手順にまたがって作業を続けるエージェント用途と、対話的・視覚的な成果物の作成を強化した。「Cursor」、「Grok Build」、「Grok Bot」、APIで同日提供…
Amazon will train on Twitch streamers’ content by default, unless they opt out
"If this was opt-in, nobody would opt in," Twitch CPO Mike Minton said on a livestream responding to user feedback. "That's honestly the an…
メルカリが明かす「Claude Code全社展開」「シャドーAI対策」を支える仕組み
「AIを使わない選択自体がビジネスリスク」と断言するメルカリ。同社は2026年5月、「Claude Code」「Claude Cowork」の全社展開に踏み切った。だが、ローカルファイルの操作やOSコマンドまで実行できる強力なツールの配布は、ガバナンスの課題も伴う。全社のAI活…
[Python]Polars公式が「全部書き換えるな」と言う理由 pandasからの移行、3戦略
Polars公式が、pandasからPolarsへの移行戦略を3つに整理した記事を公開した。区間ごとに移す手順を整理しつつ、AIに任せれば片付くのかという疑問についても筆者の見解を述べる。
AI coding startup Cognition reportedly already in talks to raise at $40B valuation
Cognition may be looking to raise another mega round just a few months after raising $1 billion at a $26 billion valuation.
As AI safety concerns mount, three pioneers make the case for staying open
At Ai4, three of the world's most respected AI experts — Geoffrey Hinton, Fei-Fei Li, and Andrew Ng — debated regulation, open source acces…
OpenAI-backed Thrive Holdings raises $2B to bring AI to the enterprise
Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Alti…
Mesh, Automattic’s CRM for everyone, comes to Android
Mesh, an AI-powered contacts app and relationship manager from Automattic, is now an Android app.
Why Stream ring-maker Sandbar says the future of AI wearables is voice
AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing e…
Lovable confirms new $13.3B valuation, raises another $400M
This new funding comes after Lovable hit $500 million in annualized run rate revenue in June, the startup told TechCrunch.
How a $250 million acquisition collapsed into allegations of fraud and forged signatures
Investors are still waiting for their share of the $250 million windfall, and VideoVerse co-founder Vinayak Shrivastav is now at the center…
2026-08-12(340件)
Why Sandbar thinks it’s voice-enabled ring can avoid the AI hardware graveyard
AI notetaking hardware has taken off over the past couple of years, with credit-card-sized devices, pendants, pins, and even transcribing e…
Everything announced at Made by Google ’26: Pixel 11, Pixel Watch 5, Pixel Tag, and tons of Gemini features
From the Pixel 11 series and a brand new competitor to Apple’s AirTag, here are all the announcements from the Made by Google 2026 event.
Google、折りたたみスマホ「Pixel 11 Pro Fold」発表 10%軽く、1mm薄くしつつ耐久性強化
米Googleは8月12日、折りたたみスマートフォン「Pixel 11 Pro Fold」を国内発表した。前モデルより約10%軽く約1mm薄いが、新しいヒンジとセラミックガラスの採用で3倍の強度を実現したという。Googleストアでの価格は256GBモデルが27万9900円。8…
Google新スマホ「Pixel 11シリーズ」正式発表 「Gemini Intelligence」向けに設計 カメラも強化
米Googleは8月12日、「Pixel 11」「Pixel 11 Pro」「Pixel 11 Pro XL」を発表した。新チップ「Google Tensor G6」を搭載し、カメラを刷新。Proには通知を光で知らせる「HiLight」を加えた。中核のAI機能「Gemini I…
Putting sign language AI into users’ hands
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
AI code-testing startup Blacksmith’s valuation jumps almost 10x in less than a year
Blacksmith says revenue has grown more than tenfold over the past year.
SpaceXAI、AIエージェント「Grok Bot」発表 クラウド環境で常時稼働、GrokやCursorの有料ユーザー向けに
米SpaceXAIは8月11日(現地時間)、常時稼働型のAIエージェント「Grok Bot」(β版)を発表した。チャットでタスクを渡すとBotがアプリやWebサイトを操作して作業を進め、役割の異なる複数のBotを並列で働かせることもできる。
中国発AIエージェント「Manus」、Metaから独立へ 中国政府が買収に反発、一部ユーザーデータは削除に
AIエージェント「Manus」を提供するManusは8月11日(現地時間)、独立企業としての運営をまもなく再開すると発表した。米Metaからの分離に伴い、一部ユーザーのデータを23日から削除するとして、事前のバックアップを呼び掛けている。
From assistance to execution: How enterprises put AI to work
OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI ad…
24時間働く“AI同僚”「Grok Bot」公開 業務アプリにログインして操作、「仕事を任せられる」
Botが業務アプリにサインインして人間の操作と同じ手順で画面を動かし、仕事を最後まで仕上げるという。
ChatGPTデスクトップアプリにLinux版 「お待たせ。MacBookはキャンセルしていいよ」
プレビュー版だが「待ちきれずにMacBookを注文してしまったなら、キャンセルしていい。それくらい出来がいい」という。
Closed-Loop LLM Co-Pilots for Digital Agriculture
This study evaluates the application of Large Language Models (LLMs) in complex biological systems, evolving from data analysis to autonomo…
SPOTting the Future: Lookahead Explanations for Deep Reinforcement Learning
Deep reinforcement learning (DRL) agents achieve strong performance in complex environments, yet their decision-making processes remain dif…
MIDAS: Mutual Information Disentanglement with Uncertainty-Aware Fusion for Incomplete Multimodal Sentiment Analysis
Most existing multimodal sentiment analysis approaches assume access to complete multimodal inputs. However, real-world applications freque…
Towards Sustainable Artificial Intelligence: A Comprehensive Review and Comparative Analysis of Deep Learning Models' Carbon Footprint
Artificial Intelligence (AI) and Machine Learning (ML) have become powerful tools for supporting and automating complex human tasks. Despit…
ReCBM: Uncertainty-Gated Relational Reasoning for Concept Bottleneck Models
Concept Bottleneck Models (CBMs) provide an interpretable framework by grounding predictions in human-understandable concepts, enabling sem…
Automating and Scaling Behavioral Scientific Research on AI Agents
As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific r…
CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation
Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual…
MESA:Task-Adaptive Multi-Structure Evidence Selection for Long-Horizon Agent Memory
Long-horizon agents accumulate trajectories spanning hundreds of interleaved reasoning, action, and observation steps, where answering a qu…
The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AI
Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typ…
SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents
Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their perfo…
Generating Attacks for LLMs with GFlowNets
The rapid advancement of Large Language Models (LLMs) has facilitated their ubiquitous integration into various domains, leading to widespr…
TRACE: Trustworthy Retrieval-Augmented Conversational Engine
Public service chatbots are expected to deliver recommendations from an underlying public service directory, while also making sure that th…
Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents
Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing vis…
Edge Phoneme Recognition for Children's Speech through Age-Aware Training
Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics…
Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding
Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based h…
Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes
Feedback signals used to train Large Language Models (LLMs) are the primary driver of their behavior and our main lever for instilling alig…
Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models
Do large language models contain domain-specific parametric shells: concentrated, causally necessary neuron populations whose removal selec…
Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems
AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent int…
Self-evolving Agentic Customer Support System at LinkedIn
Enterprise support agents operate in rapidly changing environments where policies, product capabilities, and knowledge bases evolve continu…
Beyond Decision Boundaries: Relational Geometry Attacks on Contrastive Embedding Manifolds
Contrastive learning and Siamese embedding models have become the foundation of modern verification systems, where decisions are governed n…
Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction
Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can…
Interpreting Language Model Hidden States at Scale
Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token…
Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability
Electronic health-record interoperability is a boundary problem: legacy systems, generative models, terminology services, identity systems,…
Neuroevolution Arena: Nested Ecological Evaluation of Update-and-Inheritance Regimes across Neural Architectures
Competitive artificial-life systems can rank trained controllers differently under training and ecological evaluation. We present Neuroevol…
Toward a Theory of Value in AI Alignment
Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a…
Hierarchical Compositionality for An Assistive AI Agent
AI agents are increasingly being developed to assist humans in various applications, and Large Language Models and other deep network archi…
Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated Research
AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying d…
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?
Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and…
Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action Models
Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their…
Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning
Reinforcement learning (RL) has shown promising performance in autonomous driving, yet ensuring the safety of online RL policies remains ch…
Reasoning Shortcuts and Value Symmetries: What Symmetry Permits, Architecture Realizes, and Optimization Selects
Reasoning shortcuts are solutions of a neurosymbolic system's rules that produce correct predictions through unintended concepts. A recent…
Recovering Wasted Compute in Autoresearch Agents
A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such a…
Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance
Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) n…
Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning
Large language models increasingly rely on external tools to access up-to-date information, perform computation, and interact with the outs…
Rationale-Guided Learning for Multimodal Emotion Recognition
Multimodal emotion recognition in conversation (MERC) requires understanding complex interactions between verbal and non-verbal cues. Howev…
Quantum Incremental Learning with Mixed State Prototypes
Incremental learning models are required to learn new classes sequentially without catastrophic forgetting, while operating under parameter…
RLMOpt: Adaptive Prompt Optimization via Recursive Language Models
Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimi…
Evaluating Rational Contracting in Natural Language
The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or f…
Multi-Granular Rationale-Guided Molecular LLM for Property Prediction
Large language models (LLMs) are widely applied across chemical tasks, such as molecular property prediction, which underpins drug discover…
Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning
Double perovskites (DPs) offer broad compositional tunability, but predicting the space groups (SGs) of stable structures remains difficult…
INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators
Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In…
GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning
Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. Th…
From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents
Persistent memory lets language-model agents reuse information across sessions, but it also makes errors durable: a poisoned, stale, or mis…
MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph
As coding agents increasingly handle implementation, the central challenge shifts from building individual agents to building an infrastruc…
RadFusion: Towards Threshold-Controllable Radiology Report Generation
Automated radiology report generation is advancing rapidly in response to the shortage of radiologists, yet unlike a perception model, exis…
MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows
Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a partic…
Measuring Semantic Abstractness of SAE Features via Nonlocality
Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via unders…
SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models
Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems…
Reinforcement Learning-Based Laser Cutting Machine Parameter Optimization
Achieving high accuracy in laser-based cutting of optical films requires careful tuning of parameters such as focal length and laser power…
DashArena: Benchmarking LLMs on Interactive Analytic Dashboard Generation
Analytic dashboards combine coordinated views and interactions for data exploration and decision-making. Recent models can generate them fr…
Agentic Instruction Data Selection: Let DataMaster Interpret Your Intent
Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes…
HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment
Scholar assessment plays a fundamental role in faculty recruitment, funding allocation, academic promotion, and talent discovery. Existing…
Curate Before You Connect: Identity and Ontology Tagging in a Production Knowledge Graph
Extraction produces candidate entities and relationships; writing them into a graph is where identity is decided, and identity decisions ar…
Decision-Aware Approximation of Belief Functions for Evidential Combinatorial Optimization
Reducing the number of focal elements of a mass function is classically driven by an intrinsic distance, such as Jaccard or Jousselme, that…
Operationalising Relative Causal Knowledge: Backbone Identifiability from Private Reports on a Shared Outcome
The Relativity of Causal Knowledge (RCK) explains how a network of agents with different structural causal models can exchange causal knowl…
VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus
Multimodal large language models often generate reasoning chains containing subtle errors that lead to incorrect answers. Current verificat…
FITTER: Vocabulary-Agnostic Cross-Domain Inference on Temporal Knowledge Graphs
Temporal knowledge graphs are central to many uses of the Semantic Web, but existing completion methods assume the entities, relation names…
REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exp…
Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory
Large language model (LLM)-based search agents answer questions through multi-step interactions with external environments. However, provid…
Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied.…
Tree-of-Ideas: Automated Research Ideation via Cross-Trajectory Reasoning over Scholarly Evolution
Effective research ideation requires moving beyond a static understanding of prior work to trace how research problems and solutions evolve…
Compositional Benchmark Synthesis for Hierarchical Human Action Recognition
Recognizing human behavior across levels of abstraction, from atomic actions to long-horizon intentions, requires data annotated along a se…
Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information
Explainable Artificial Intelligence (XAI) seeks to explain how an Artificial Intelligence (AI) system arrived at a particular decision. We…
SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation
Computer-using agents can perceive rich software interfaces, yet their decisions often lack visual procedural memory: they may recognize in…
ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent Experimentation
Autonomous chemistry increasingly depends on environments in which agents can repeatedly act, observe, and adapt.Physical laboratories prov…
EvoMem: Memory-Augmented Evolution for Code Optimization
Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some c…
Hypothesis Frontier: Verifier Guided LLM and Symbolic Search for First-Order Induction
First-order concept synthesis asks a system to infer one formula that classifies labeled objects consistently across several finite relatio…
Enhanced Filtering Algorithms for the Euclidean Traveling Salesperson Problem and its variants in Constraint Logic Programming
The Traveling Salesperson Problem (TSP) is one of the best-known problems in computer science and arises in many engineering applications,…
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Y…
IO Factory: Simulating AI-Enabled Influence Campaigns at Scale
We introduce IO Factory, an AI-driven framework for simulating information and influence campaigns as fully integrated, traceable processes…
ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling
Large Reasoning Models (LRMs) improve performance by allocating additional inference-time compute to generate extended chain-of-thought rea…
FedCGR: Federated Cross-Domain Generative Recommendation
Cross-domain recommendation (CDR) transfers preference knowledge across related domains, but federated deployment makes cross-domain alignm…
XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving
Vision-Language-Action (VLA) models can connect scene understanding, semantic reasoning, and trajectory generation for autonomous driving.…
V-FiLLM: Verified Financial LLM Reasoning Benchmark
While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data r…
SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure
Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is of…
RTSKG: Building a Rail Transit Station Knowledge Graph Dataset
Rail transit systems play a vital role in urban mobility and economic development. As key components of such systems, rail transit stations…
Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding
Agentic coding READMEs like CLAUDE.md grow without bound in real repositories, stopping only when the repository retires or someone rewrite…
sLTN: Structural Logic Tensor Networks
Logic Tensor Networks (LTN) provide a neurosymbolic framework in which first-order logic is interpreted through tensor operations, enabling…
Long-Horizon AI Research for Grothendieck Constant: A Case Study in Human-AI Mathematical Collaboration
AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an e…
The Gaussian-Multinoulli Restricted Boltzmann Machine: A Potts Model Extension of the GRBM
Many real-world tasks, from associative memory to symbolic reasoning, benefit from discrete, structured representations that standard conti…
"YES! YES! I absolutely love this insight!" Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbots
This article analyses narrative mechanisms that are common in dialogues with LLM chatbots. In combination, these mechanisms produce an inte…
LLM Agents Factory: Retrieval of Domain-Specific LLM Agents
Large language model (LLM) agents improve task performance by decomposing problems into role-specialized behaviors. However, their practica…
How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation
Production teams deploying LLM chat agents face a specific quality assurance gap: existing evaluation tools test individual responses or si…
When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning
It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual f…
Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI Agents
Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important s…
HoosierHelp: Benchmarking LLM Agents for Social Service Navigation
Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Alth…
Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems
The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted…
Evidence-Based Scientific Question Discovery: A Framework with Historical Backtesting
Current AI systems are optimized for answering questions; the scientific enterprise is bottlenecked earlier, at discovering the questions w…
Rescene: band-limited stochastic forcing turns a frozen neural weather operator into a climate emulator
Over the past few years, the rapid development of machine learning (ML) models for weather forecasting has produced deterministic models wh…
Do AI weather models miss extremes?
First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic…
Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy
Controllable generation guided by external knowledge is a key requirement in modern generative deep learning applications, enabling the syn…
Energy and Performance Benchmarking of Deep Learning Models for Breast Cancer Detection
Recent advances in machine learning have greatly improved breast cancer detection, enabling more accurate and timely diagnosis. Deep learni…
Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification
The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep R…
Sheaf-Based Federated Representation Learning
Heterogeneous federated systems require agents to learn and exchange informative representations despite differences in data distributions,…
DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents
Large language models (LLMs) increasingly rely on external tools to accomplish complex real-world tasks, making tool documentation a critic…
UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs
Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, gen…
Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons
The problem of learning from pairwise comparisons has been widely studied across many domains such as recommendation systems, social choice…
Physics-Informed Machine Learning in Prognostics and Health Management: A Systematic Literature Review
In modern industry, keeping complex systems reliable, safe, and efficient hinges on Prognostics and Health Management (PHM). Machine Learni…
Navigating the Proximity-Safety Balance: Constraint Decomposition for Human Following in Pedestrian Crowds
Following a target human in crowded environments involves an inherent conflict between staying close to the target and navigating safely am…
Status Association Does Not Reliably Predict Decision Leakage
Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will al…
Exploring Semantic Stability Across Reviews in the Linux Kernel
Code review is credited with substantially changing a patch's code between its first submission and the version that eventually lands. Howe…
Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4
High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matri…
Procedural Fairness Failures in RLHF from Preference Averaging
Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homo…
Multimodal Item Parameter Estimation using Simulated Response Probabilitie
We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multi…
MarkNull: Model-Agnostic Watermark Removal in AI-Generated Images via On-Manifold Latent Manipulation
Digital watermarking has emerged as a critical technique for provenance and copyright attribution in AI-generated imagery, yet its robustne…
From Prediction to Incrementality: Causal Optimization for Large-Scale Targeting and Recommendation
Large-scale targeting and recommendation systems are typically built around predictive scores fed into heuristic or local allocation. When…
The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse
LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployme…
ELMER: Evolutionary Language Model that Explores and Refines
Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edi…
Similarity Gates Approve Reversals: A Validity Audit of Embedding-Cosine Thresholds in Agent Systems
Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff. Deduplication fil…
FACT: Failure-Aware Causal Training for World-Action Models
Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation.…
Unsupervised Detection of Groundwater Storage Anomalies in Ghana Using GRACE Satellite Data
Groundwater variability in Ghana remains poorly characterized due to limited long-term in-situ observations. This study investigates ground…
TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent
Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing be…
Toward Human Rights Benchmarking for LLMs: A Pilot Methodology
Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation ben…
Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies
Deploying large language models (LLMs) for decision support in emergency departments (EDs) faces two major challenges: privacy risks of tra…
Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output
Streaming language-model output creates a release-timing problem: complete-response moderation acts after streamed text has escaped, wherea…
Comprendia: AI-Augmented Code Comprehension
Comprendia is an Eclipse plugin that integrates structural dependency visualization with LLM-powered code explanation on a shared interacti…
MRIComp4Flow: Compression of 3D Brain MRI for Training Multi-Modal Generative Models
Large-scale multi-modal MRI datasets impose substantial storage and I/O costs, limiting the training of 3D generative models on commodity i…
Frozen Brain-MRI Foundation Models Are Site Fingerprints
Frozen foundation-model (FM) embeddings are increasingly used as off-the-shelf brain-MRI representations, on the assumption that they captu…
Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility
Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal b…
Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories
Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering task…
Narrative Keyframing for Generative Creative Writing
We introduce narrative keyframing, an interaction technique for AI-assisted creative writing that lets writers specify different types of n…
Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement
Hospital quality improvement (QI) programs routinely face multiple candidate interventions to optimize hospital flow, but existing methods…
Towards Unified Dynamic Face Landmark Detection
Although advancements in face landmark detection (FLD) methods continue to push performance boundaries, they overlook two major functional…
Efficient Reinforcement Learning for Long-Horizon Tool-Use Agentic Tasks
Long-horizon tool-using agents must reason over user goals, domain policies, tool calls, simulator state, and delayed verifiable rewards. R…
MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model
Arabic maqam music microtonal, modal, and built on ornamented call and response is among the traditions most underserved by generative musi…
MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices
Speculative decoding accelerates autoregressive large language model (LLM) inference by using a lightweight draft model to speculate multip…
Beyond Forecasting: Recasting Volatility Control as a Routing Problem
Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a…
A Single Atom in Front of a Mirror is a Universal Reservoir Computer
Universal approximation in reservoir computing is typically associated with a class of reservoirs. We show that universality can be associa…
Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation
Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about…
ELVAE: Evidential Learning-Based Variational Autoencoder for Uncertainty-Aware Generation
Variational autoencoders generate samples from probabilistic latent representations but do not distinguish uncertainty about the latent loc…
Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models
Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulti…
Riemann GeoResolver: A Non-Euclidean Attention Framework from Euclidean Resolver to Hyperbolic-Spherical Geometry
We present a theoretical foundation for inverse-distance attention, from its Euclidean prototype (Resolver) to its non-Euclidean realizatio…
Causality Sum Rules in Conventional Scattering Matrices
Scattering matrices are the standard experimental and computational description of photonic and electromagnetic devices. Passivity is expli…
Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique
Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesi…
What We Know about Responsible AI Practices in Industry: A Half Decade of Empirical Research
Responsible AI (RAI) has become a central concern for technology companies, regulators, and the public. How industry practitioners interpre…
FUSE: Frame-Unified Stress Estimation from Facial Video
Automatic stress detection from facial video offers a practical path to non-intrusive affect monitoring, yet existing video-based approache…
From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models
Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chain…
Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging
Large language model-based recommender systems are increasingly adopting slow-thinking models that generate step-by-step reasoning before m…
Persistent Recursive Worlds Enable Autonomous Software Evolution
Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems pre…
MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection
As LLM-generated content becomes more sophisticated, detection systems for distinguishing those texts from human-written text must operate…
Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning
Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online…
Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models
Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vi…
Exploration-Driven Personalized Federated Reinforcement Learning via Intrinsic Motivation
Personalized Federated Reinforcement Learning (PFRL) takes a decentralized approach to storing and accessing information based on past expe…
SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning
Large vision-language models (LVLMs) remain vulnerable to jailbreak attacks that exploit visual inputs to bypass safety alignment inherited…
Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training
While vision-language models dominate medical representation learning, unstructured text lacks the dense, quantitative diagnostic phenotype…
Improving TensorSketch Using Complex Random Variables
\texttt{TensorSketch} by~\cite{pham2013fast,kar2012random} provides efficient sketching algorithms for high-dimensional polynomial kernels…
Rethinking Text-Based Image Retrieval in Specific Domain
Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. Ho…
Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models
Historical context integration presents a fundamental challenge for Vision-Language Models (VLMs) in sequential decision-making tasks. Curr…
Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits
Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschit…
Robust Multi-Agent Bandits with Heavy-Tailed Rewards and Information Asymmetry
The multi-armed bandit problem is a central framework in sequential decision-making, extensively studied under sub-Gaussian reward assumpti…
On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models
Large Language Models (LLMs) have undergone a shift from stateless conversational interfaces to autonomous agents capable of multi-step pla…
Flow Straight to Reality: Perceptually Consistent Flow Matching for Efficient Image Restoration
Image restoration is fundamentally constrained by the tradeoff between distortion and perception: minimizing pixel-wise error yields over-s…
ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover
Edge LLMs must preserve inference continuity when a user hands over between edge nodes, requiring key-value (KV) cache transfer to the targ…
Retrieval-Corrected Conformal Prediction for Time Series
Conformal prediction (CP) provides distribution-free prediction intervals for fixed forecasters, but its standard calibration procedure is…
A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign Language
Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, ph…
$\pi$-SUB: A Physics-Informed Synthetic Underwater Benchmark Dataset for Underwater Image Enhancement
This paper presents $\pi$-SUB, a physics-informed framework for generating synthetic underwater benchmark datasets that bridges the synthet…
DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction
Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degrad…
Inferential Capability Does Not Determine Legal Scope
Two instruments of EU digital law place inference at their centre and mean different things by it. Article 3(1) of the AI Act uses the capa…
Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts
In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling la…
MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models
Medical Vision-Language Models (Med-VLMs) excel at verbalizing visual content, yet precise visual perception, segmentation, and grounding r…
Cross-View Sequential Visual Localization with Spatio-Temporal Context Modeling for Autonomous Driving
Continuous and reliable localization is essential for autonomous driving. Cross-view visual localization matches ground images with satelli…
Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement
Social interaction has become one of the most common uses of LLMs, yet research on emotional bonds with AI has focused largely on how users…
Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics
Chinese web pollution has surfaced in LLMs, motivating audits of upstream Chinese corpora. However, auditing such corpora faces three chall…
ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question Answering
Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, howev…
SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information
Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scatt…
Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization
Evolutionary optimization of LLM prompts and agentic programs (e.g., GEPA) is dominated by fitness evaluation: scoring each candidate runs…
ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes
Text-Attributed Graphs (TAGs), endowed with abundant textual content along with topological structures, have emerged as a versatile backbon…
Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control
Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and dow…
Conversational Orchestration for Organic 6G
The Organic 6G vision of a network of networks spanning an edge-cloud continuum complemented by non-terrestrial resources requires, to real…
Most biomedical publications show signs of LLM-assisted writing
Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing c…
DuplexWorld: Can voice agents help you get through the day?
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consu…
Optimal Stopping of Self-Refining Foundation Models
Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is em…
Smart Enough to Go Extinct? An Evolutionary Challenge to the Value of General Intelligence and Its Ethical Implications for AGI
The pursuit of artificial general intelligence (AGI) rests on a seemingly self-evident premise: that general intelligence, the kind of flex…
A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem
The Model Context Protocol (MCP) has become the de-facto interface for connecting LLM agents to enterprise tools, and adoption has been exp…
The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election
The increasing use of Generative Artificial Intelligence (GenAI) in journalism raises concerns about possible detrimental effects both on j…
MVTrack: Ultrafast Appearance-Free Moving Object Tracking from Compressed Bitstreams
Deploying modern video trackers at scale is bottlenecked by the computational cost of RGB-based object detectors. To this end, we present M…
Beyond Fixed Luminance: Towards Panchromatic and Orthochromatic Image Colorization
Most image colorization systems operate in $Lab$ space by predicting chroma ($ab$) while preserving an input-derived luminance channel ($L$…
BPG: Balancing Plasticity and Generalization for Domain Incremental Learning
Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performan…
Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware Reformulation
Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's re…
Modelling Geographic Atrophy Progression using Implicit Neural Representations
Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irrever…
Surfacing the Unsaid: CUE-Bench for Affective Stance in Chinese Discourse
Emotion understanding in discourse requires reasoning beyond surface sentiment because speakers often convey affect through indirect, impli…
Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation
We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-fi…
MIRA: Medical Image Reflection for Agentic Diagnosis
Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or m…
Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallu…
TACTICL: Task-Aware Compression of Tabular ICL Models
The strong performance of foundation models for tabular tasks comes at substantial inference costs. Distilling models into task-specific ar…
VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-c…
GitSkills: A Dataset of Agent Skills on GitHub
An agent skill is a folder containing a SKILL.md file with instructions for a language-model agent, optionally accompanied by scripts and r…
FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation
Autoformalisation (AF) systems map natural language reasoning steps into formal statements in a proof assistant such as Lean. We consider h…
Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation
Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation…
A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models
Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recogn…
Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes
While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deterior…
CARE: Confidence-Aware Reasoning for Reliable Medical VQA
Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning fo…
ReLTEx: Reliable LLM-based Taxonomy Expansion
Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and rela…
TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation
Multi-modal recommenders fuse collaborative signals with item modalities such as text, images, and audio, but the usefulness of each drifts…
Putting Registers to Work: Task Registers for Token Pruning in Vision Transformers
Token-pruning policies are usually designed for a single recognition pipeline, but pretrained Vision Transformers are reused across tasks w…
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-…
Policy Convergence and Divergence Across National and Within Regional AI Strategies: A Policy Design Element Analysis
Governments worldwide have responded to the rapid expansion of AI by publishing national and regional AI strategies. Comparing national and…
R4DSG: Relative 4D Scene Graph Memory for Object-Centric Question Answering in Long Egocentric Video
Long-horizon egocentric video is a rich substrate for wearable AI assistants, but object-centric questions such as where an item was moved,…
Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data
Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturin…
Multiclass Sentiment Analysis for Identifying Political Viewpoints
The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public o…
3D Weighted Geometric Graph Neural Networks for Sheep Facial Pain Assessment
Deep learning systems perform mainly within the 2D for a single image domain and take the face as a single-dimension representation, losing…
A Comparative Evaluation of Deep Learning Object Detection Models on a Real-World Multi-Plant Dataset from Africa
The application of computer vision in agriculture has shown significant potential for improving crop monitoring and precision farming. Howe…
Entropy-Centric Explainable AI for Remote Sensing Image Segmentation
Artificial intelligence (AI) has become a powerful approach to solving complex problems in critical domains. Many concerns arise regarding…
Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory
We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a…
Two-stage Odd Residual Flows for Mean-Preserving Probabilistic Time Series Forecasting
Probabilistic forecasting plays an essential role in risk-sensitive decision-making, particularly in long-horizon settings. However, existi…
Attention-Path Fragility as an Uncertainty Signal in Large Language Models
We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a c…
From Interpretability to Control: Insights from Six Years of the TrustNLP Workshop
The Workshop on Trustworthy Natural Language Processing (TrustNLP), co-located with major ACL conferences since 2021, has grown from 8 proc…
How to Verify Consistency of Probabilistic Claims
When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in p…
Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiti…
ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls
Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, r…
Surgical WAM: A World-Action Model for Data-Efficient Surgical Robot Learning
Learning reliable surgical manipulation policies is bottlenecked by the scarcity of action-labeled demonstrations: teleoperated surgical ro…
Representation and Invariance in Reinforcement Learning
Researchers have formalized reinforcement learning (RL) in different ways. If an agent in one RL framework is to run within another RL fram…
GAM-Agent: Game-Theoretic and Uncertainty-Aware Collaboration for Complex Visual Reasoning
We propose GAM-Agent, a game-theoretic multi-agent framework for enhancing vision-language reasoning. Unlike prior single-agent or monolith…
Closing a 17-Year Gap: Algorithmic Detection and Empirical Prevalence of Rank Reversal in Multi-Criteria Decision Analysis
Rank Reversal, where the relative order of alternatives changes in ways that violate axioms of rational decision-making, is a well-document…
Multiplayer Nash Preference Optimization
Reinforcement learning from human feedback (RLHF) has emerged as the standard paradigm for aligning large language models with human prefer…
On The Statistical Limits of Self-Improving Agents
We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this…
Situation Graph Prediction for User Perspective Modeling
Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited b…
Leveraging Large Language Models for Causal Discovery: a Constraint-based, Argumentation-driven Approach
Causal discovery seeks to uncover causal relations from data, typically represented as causal graphs, and is essential for predicting the e…
JEPA-DNA: Grounding Genomic Foundation Models through Joint-Embedding Predictive Architectures
Genomic Foundation Models (GFMs) typically rely on Masked Language Modeling (MLM) or Next-Token Prediction (NTP) to learn the "Laws of Natu…
CoEvoSkills: Self-Evolving Agent Skills via Co-Evolutionary Verification
Anthropic proposes the concept of skills for LLM agents to tackle multi-step professional tasks that simple tool invocations cannot address…
Planning Task Shielding: Detecting and Repairing Flaws in Planning Tasks through Turning them Unsolvable
Most research in planning focuses on generating a plan to achieve a desired set of goals. However, a goal specification can also be used to…
Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents
Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely valida…
When Does Critique Improve AI-Assisted Theoretical Physics? SCALAR: Structured Critic--Actor Loop for Agentic Reasoning
As large language models (LLMs) show increasing promise on research-level physics reasoning tasks and agentic AI becomes more common, a pra…
CuSearch: Curriculum Rollout Sampling via Search Depth for Agentic RAG
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for training agentic retrieval-augmented generati…
Memory-Augmented Reinforcement Learning Agent for CAD Generation
Automatic generation of computer-aided design (CAD) models is a core technology for enabling intelligence in advanced manufacturing. Existi…
A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents
Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely trea…
From Talking to Singing: A New Challenge for Audio-Visual Deepfake Detection
With rapid advances in audio-visual generative models, reliable forgery detection becomes increasingly critical. Existing methods for audio…
AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Self-Explaining Mathematical Reasoning
We present AXIOM, a trust-first neuro-symbolic architecture for natural-language mathematical reasoning. Its language model is strictly a c…
A case study of evaluating AI agents on a neuroscience data-to-discovery pipeline
Agentic AI offers a promising path to automating software development bottlenecks in scientific research pipelines, particularly for stages…
When Agent Automation Becomes Profitable: Quantifying and Insuring Autonomous AI Risk through Trace-Economic Underwriting
AI agents can now take irreversible actions in operational systems, but agent-caused losses are still not clearly assigned, priced, or tran…
ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection
Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images,…
Coachable agents for interactive gameplay
Reinforcement learning has proven to be a valuable tool in the creation of advanced AI and robotic systems, contributing to everything from…
Measure the Sim-to-Real Gap: Designing an Affordable Real-World Benchmark Platform for Reinforcement Learning in AIoT Systems
Reinforcement learning (RL) is commonly employed to enhance the performance of autonomous systems, including the Autonomous Internet of Thi…
UPAIR: Diagnosing Reasoning States via Uncertainty-Progress Alignment for Selective Intervention
While test-time scaling improves the problem-solving ability of large reasoning models (LRMs) through additional inference-time computation…
SAE-StatSteer: Statistical Consensus Feature Selection for Optimization-Free Activation Steering of Large Language Models
Activation steering adds a residual-stream direction at inference time, providing lightweight behavioral control without fine-tuning. Spars…
AttriMem: Attribution-Guided Process Feedback for Agent Memory Construction
Effective memory is crucial for LLM agents, yet constructing it effectively remains challenging. A memory-construction policy decides what…
Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog
Large Language Models (LLMs) excel at natural language understanding and generation but remain unreliable for multi-step logical reasoning,…
Learning and Structurally Validating Simulation Scenario Continuations in Dynamic Graph Systems
Data-driven generative models can extend partially observed simulation trajectories into ensembles of alternative future scenarios. However…
EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups…
TrAC: Trace-Conditioned Answer Consistency for Efficient Uncertainty Quantification in LLMs
Large language models (LLMs) can generate fluent reasoning traces that nevertheless lead to incorrect answers, making response-level uncert…
DiffImaginE: Imagine to Verify Entity Types with Diffusio
Multimodal named entity recognition (MNER) determines whether each candidate span and entity-type hypothesis is supported by joint textual…
CastFSR: A Fast--Slow--Reflect Agentic Reasoning Framework for Context-Aware Time Series Forecasting
Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical o…
TARL: Transaction-Aware Reliable Ledgers for Executable Memory Management in Long-Term Agents
Persistent memory helps long-term agents retain knowledge, yet a single update error can repeatedly distort future retrieval and reasoning.…
ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment
This paper presents ECHO (Enhanced Care & Health Observer), a locally-deployable conversational health assistant for long-term chronic care…
Emergent Neural Network Mechanisms for Generalization to Objects in Novel Orientations
The capability of Deep Neural Networks (DNNs) to recognize objects in orientations outside the distribution of the training data is not wel…
Graphical Models of False Information and Fact Checking Ecosystems
The wide spread of false information online, including misinformation and disinformation, has become a major problem for our highly digitis…
Pretrained Optimization Model for Zero-Shot Black Box Optimization
Zero-shot optimization involves optimizing a target task that was not seen during training, aiming to provide the optimal solution without…
Regression and Classification with Single-Qubit Quantum Neural Networks
The literature reflects a mutually beneficial relationship between machine learning and quantum computing, where progress in one field freq…
Protecting Creative Writing Copyright against AI Imitation via Implicit Watermarking
Large language models (LLMs) enable powerful knowledge injection through approaches such as in-context learning and fine-tuning, but they a…
TransitReID: Transit OD Data Collection with Occlusion-Resistant Dynamic Passenger Re-Identification
Transit Origin-Destination (OD) data are fundamental for optimizing public transit services, yet current collection methods, such as manual…
Demystifying Adversarial Robustness in Diffusion Models: Compression, Randomness, and Geometry
Recent studies suggest that diffusion models significantly improve the empirical adversarial robustness of deep neural network models. Whil…
Music Interpretation and Emotion Perception: A Computational and Neurophysiological Investigation
This study investigates emotional expression and perception in music performance using computational and neurophysiological methods. The in…
HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) have demonstrated significant potential to advance a broad range of domains. However, current benc…
SynBoost: A Synergistic Framework for Fast Sampling of Diffusion Models
Diffusion probabilistic models (DPMs) have demonstrated remarkable success in visual generation. However, their iterative sampling mechanis…
OpenDPDv2: A Unified Learning and Optimization Framework for Neural Network Digital Predistortion
Neural network (NN)-based Digital Predistortion (DPD) improves linearization for wideband radio frequency (RF) power amplifiers (PAs) but o…
Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling
This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language mode…
Selective Prediction Reduces the Negative Effects of Automation Bias Overall but Increases False Negatives
AI has the potential to augment human decision making. However, even high-performing models can produce inaccurate predictions when deploye…
Reconfiguration of pivoting cube ensembles under local sensing constraints using geometric deep learning
We demonstrate that local sensing is sufficient for effective global reconfiguration of homogeneous pivoting cube modular robots in two dim…
Token-Based Detection of Spurious Correlations in Vision Transformers
Due to their powerful feature association capabilities, neural network-based computer vision models have the ability to detect and exploit…
Faster Results from a Smarter Schedule: Reframing Collegiate Cross Country through Analysis of the National Running Club Database
Collegiate cross country teams often build their season schedules on intuition rather than evidence, partly because large-scale performance…
Diffusion-Based Impedance Learning for Contact-Rich Manipulation Tasks
Learning-based methods excel at robot motion generation but remain limited in contact-rich physical interaction. Impedance control provides…
Pricing Access to Dynamic Information Services
A provider sells a \emph{dynamic information service}---a real-time, capacity-constrained process that resolves a customer's uncertainty---…
HyWA: Architecture-Preserving Personalized Voice Activity Detection for Full-Duplex Voice Assistants
Voice activity detection (VAD) serves as an early gate in voice-assistant pipelines for smart devices. Because conventional VADs respond to…
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary m…
On the Condition Number Dependency in Bilevel Optimization
Bilevel optimization minimizes an objective function, defined by an upper-level problem whose feasible region is the solution of a lower-le…
Auto-exploration for online reinforcement learning
The exploration-exploitation dilemma in reinforcement learning (RL) is a fundamental challenge to efficient RL algorithms. Existing algorit…
Hybrid Token Compression for Vision-Language Models
Vision-language models (VLMs) rely on hundreds of visual tokens, leading to high computational and memory costs. Existing compression metho…
On Solomonoff Induction in Large Language Models and the Limits of Self-Improving: The Singularity Is Not Near Without Symbolic Model Synthesis
On the one hand, the question of whether large language models (LLMs) are Solomonoff induction estimators has become an explicit question a…
LaViT: Aligning Latent Visual Thoughts for Multi-modal Reasoning
Current multimodal latent reasoning often relies on external supervision (e.g., auxiliary images), ignoring intrinsic visual attention dyna…
GraFine: Retrieval-Time Refinement for Efficient Graph RAG over Corpus Graphs
Graph RAG on corpus graphs enhances retrieval by leveraging intermediate node content as contextual clues to uncover unretrieved oracle nod…
Bandwidth-Efficient Multi-Agent Communication through Information Bottleneck and Vector Quantization
Multi-agent reinforcement learning systems deployed in real-world robotics applications face severe communication constraints that signific…
LLMs Encode Their Failures: Predicting Success from Pre-Generation Activations
Running LLMs with extended reasoning on every problem is expensive, but determining which inputs actually require additional compute remain…
Do LLMs Benefit From Their Own Words?
In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assista…
Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers
We present an Ising-compatible framework for formal neural-network robustness verification under bounded input perturbations. For piecewise…
Can Computational Reducibility Lead to Transferable Models for Graph Combinatorial Optimization?
A key challenge in developing unified neural solvers for combinatorial optimization (CO) is the efficient generalization of models from a g…
$\mathrm{ECI}_{\mathrm{sem}}$: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives
Hard-negative source selection for dense retrieval is usually decided only after fine-tuning and downstream evaluation. We propose ECIsem,…
Does Explanation Correctness Matter? Linking Computational XAI Evaluation to Human Understanding
Explainable AI (XAI) methods are commonly evaluated using functional correctness metrics, sometimes termed faithfulness or fidelity, which…
Covert Visual Prompt Injection against Commercial Multimodal Large Language Models
Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavio…
Evaluation and Hardening of LLM System Instructions Against Extraction via Encoding Attacks
System Instructions in Large Language Models (LLMs) are commonly used to enforce safety policies, define agent behavior, and protect sensit…
BiScale-GTR: Fragment-Aware Graph Transformers for Multi-Scale Molecular Representation Learning
Fragment-level representations provide a natural way to capture recurring molecular substructures and reuse their learned representations a…
RankFormer: A Propose-then-Select Transformer for Multi-Agent Multimodal Trajectory Prediction
Predicting traffic agent trajectories plays an important role in autonomous driving, traffic operations, transportation safety analysis, et…
Loop, Think, & Generalize: Implicit Reasoning in Recurrent-Depth Transformers
We study implicit reasoning, i.e. the ability to combine knowledge or rules within a single forward pass. While transformer-based large lan…
SatIR: Scalable High-Recall Constraint-Satisfaction-Based Information Retrieval for Clinical Trials Matching
Many real-world retrieval and matching problems require more than topical relevance: a candidate must satisfy the specific constraints of o…
PinpointQA: A Benchmark for Small Object-Centric Spatial Understanding in Indoor Videos
Reliable embodied interaction in indoor environments requires agents to precisely localize small everyday objects from visual observations.…
InfiniteScienceGym: An Unbounded, Procedurally-Generated Benchmark for Scientific Analysis
Large language models are emerging as scientific assistants, but evaluating their ability to reason from empirical data remains challenging…
Language corpora for the Dutch medical domain
Background: Dutch medical corpora are scarce, limiting NLP development. Methods: We translated English datasets, identified medical text in…
Progressive Semantic Communication for Efficient Edge-Cloud Vision-Language Models
Deploying Vision-Language Models (VLMs) on edge devices remains challenging due to their substantial computational and memory demands, whic…
Proteo-R1: Reasoning Foundation Models for De Novo Protein Design
Deep learning in de novo protein design has achieved atomic-level fidelity. However, existing models remain largely non-deliberative: they…
Field-Localized Forgery Detection for Digital Identity Documents
Digital onboarding and eKYC systems used by banks, fintech platforms, telecom providers, and other third-party services commonly verify use…
Access Timing as Scaffolding: A Reinforcement Learning Approach to GenAI in Education
In recent years, generative AI (GenAI) in educational settings has become ubiquitous in university students' daily lives, despite its poten…
Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models
Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically…
Why Do Safety Guardrails Degrade Across Languages?
Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), whic…
The Matching Principle: When Does a Training Penalty Cover Deployment Shift?
Ordinary training optimises the task loss and then stops. It never pays for internal representation energy: Jacobians can stay large in dir…
Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness
Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy.…
LVCG: Learning ECG Representations in the Latent Vectorcardiogram Space
Electrocardiography (ECG) is a cornerstone of cardiac assessment, making the learning of informative ECG representations fundamental to tas…
Flow-Based Generative Modeling for Optimizing Sampling Policies in Compressed Sensing Applications
Numerous modern applications in signal processing and medical imaging necessitate acquiring high-dimensional signals under tight resource c…
On Effectiveness and Efficiency of Agentic Tool-calling and RL Training
Tool-calling is a central component of modern large language model (LLM) agents, equipping them with skills beyond their parametric knowled…
Poise: Position-Aware One-Instruction Skill Injection for Silent Execution on LLM Agents
Agent skills extend general-purpose agents, but their open format enables skill poisoning: a tampered skill can make an agent run an attack…
Time-Series Foundation Model Embeddings for Remaining Useful Life Estimation
Remaining Useful Life (RUL) prediction is essential for industrial predictive maintenance, yet many learning-based approaches rely on exten…
Market Design for AI: Beyond the Copyright Binary
How can we design a market of human-generated content for use in training AI models that both enables technological progress and preserves…
DIMOS: Disentangling Instance-level Moving Object Segmentation
Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving,…
A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian Prediction
Predicting the Kohn-Sham Hamiltonian with machine learning can accelerate density functional theory while retaining access to molecular orb…
The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model Lineages
Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs,…
Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure
Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and…
When Reranking Hurts: Uncertainty-Based Gating for Few-Shot Reranking
Few-shot selection typically assumes that reranking retrieved examples always improves performance. We challenge this view by identifying t…
Foundations of Equivariant Deep Learning: Unifying Graph and Sheaf Neural Networks
Symmetry is everywhere in nature and society. Geometric deep learning builds architectures respecting group symmetries, whereas topological…
EventCoT: Event-centric Video Chain-of-thought for Reasoning Temporal Localization
Reasoning temporal localization (RTL) requires a model to generate an answer that itself contains the time interval supporting it, coupling…
Adaptive Filtering of the KV Cache: Diagnosing and Correcting Structural-Role Bias in LLM Inference
Attention-based KV cache eviction (H2O and its descendants) compresses the memory-constrained state of a long-context model by ranking toke…
Ablation-Corrected Evaluation of Attribution Maps in Echocardiographic Ejection-Fraction Models
Attribution maps for echocardiographic ejection-fraction models are evaluated by their overlap with an expert left-ventricular annotation,…
Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning
Large language models (LLMs) are improving rapidly as reflected in benchmark scores, yet these AI benchmarks largely test capabilities such…
KAYROS: An Anytime and Exact Open-Source Solver for Duration-Minimization Time-Dependent Vehicle Routing. A Technical Report and a Case Study in Human-AI Engineering
Time-dependent routing recognizes that the same journey can take a different time depending on when it begins. Under duration minimization,…
Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, a…
Living-Harness Is an Interactive-Agent Evolver
Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur i…
The Epistemic Politics of AI Anthropomorphism
AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustain…
dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model
Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural…
Wiring Beats Blending: What Transfers Between Transformer Sizes -- and What Doesn't
Model families are typically trained size by size, each from scratch. Can apretrained large model instead be converted into a smaller sibli…
NVIDIAが30Bのオープンモデル公開 「OpenClaw」など常時稼働エージェント向けに設計
NVIDIAが30Bのオープンモデル「Nemotron 3.5 Lightning」を公開。「OpenClaw」など常時稼働エージェント向けに設計し、「gpt-oss-120b」同等の性能を約4分の1の規模で実現するという。
GeminiアプリのMAUが10億人を突破 Googleで14番目の大台到達製品に
GoogleのGeminiアプリの月間アクティブユーザー数(MAU)が10億人を突破した。Google検索やWorkspaceなどの組み込み機能を除いたアプリ単体の数字で、同社として14番目の10億ユーザー到達製品となる。競合のChatGPTに続く大台達成で、iOSユーザー数や…
「AI生成コンテンツ、実は嫌いではない」 だが「客の半分が去る」本当の理由
TWOSTONE&SonsがBtoBの比較検討・発注担当者を対象とした意識調査結果を発表した。約9割がコンテンツに「AIっぽさ」を感じた経験があり、その後問い合わせや資料請求を取りやめたと回答する割合も示された。
37兆円のSaaS支出が2030年までに「消える」 Gartnerが予測するエージェントAIの破壊力
Gartnerは、企業向けソフトウェアの支出のうち最大2340億ドル(約37兆円)が、エージェント型AIの影響にさらされるとの予測を発表した。従来型SaaSのシート課金モデルが崩れ、ユーザー数の増加と収益の増加が連動しなくなるという。
Anthropic、「Claude」で生成したテキストに“見えない透かし” 日本を含むグローバルに適用へ
Anthropicは、「Claude」が生成するテキストに電子透かしを、対応ファイルにC2PA準拠の署名付きメタデータを付与する方針を発表した。8月2日に適用開始された「EU AI Act」の透明性義務に伴う措置だが、日本を含む全世界のモデルとサービスに適用される。人間には不可…
【注目の企業】キオクシア、なぜこんなに話題? 今からでも間に合う“入門記事”まとめました
話題を呼び続ける半導体メモリ大手のキオクシア。一体なぜこんなにも注目されるのか。同社の“今”が分かる記事をまとめた。
「よく聞こえなかったデスら」――『ドラクエ』に“AIキャラ”登場、世界観を守るために張り巡らせた創意工夫
『ドラゴンクエストX オンライン』に生成AIを活用した新キャラ「スラミィ」が実装された。生成AIのゆらぎを克服し、ゲームの世界観をどのように守っているのか。
Accel closes oversubscribed $550M India fund within weeks, 19 months after its last
The U.S. VC firm still has more than 55% of its previous $650 million India fund available for deployment.
OpenAI launches ChatGPT desktop app for Linux
OpenAI is finally bringing a dedicated ChatGPT desktop app to Linux operating systems.
Google’s Gemini app surges to 1 billion users
Google also shared numbers of how people are actually using the chatbot, with 63% of Gemini users talking directly to the assistant using t…
Brad Lightcap, OpenAI’s longtime COO, is leaving to ‘start something new’
One of OpenAI's longest-serving executives is headed out the door, although the longtime COO told staff that he was "excited to help you al…
General Catalyst leads $1.1B round into 2-month-old River AI
River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of…
An unreleased Anthropic model made progress on one of math’s biggest unsolved problems
For more than 150 years, the Riemann hypothesis has stood as one of the major unsolved problems in mathematics. Anthropic hasn't solved it…
2026-08-11(761件)
Spotify will label ‘AI Persona’ profiles and exclude their music from recommendations
Spotify is introducing “AI Persona” labels for artist profiles that represent AI-generated identities and will exclude their music from edi…
Anthropic says it will watermark text generated by its AI models
Anthropic will extend support for watermarking AI generations for older models as well.
Testing ads in ChatGPT
OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user…
Daybreak models are now available on AWS
OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.
Towards an Argumentative Foundation for Evaluative AI
Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by…
Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains
Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds hu…
Determinization in Structure Theories: A Unified Framework via Closure, Comparability, and Joint Admissibility
We develop a formal framework for constructing canonical interpretations from plural structure theories. A structure theory is a triple T =…
Emotion in an active inference model of human driving
Active inference has emerged as a principled framework for modeling adaptive behavior by balancing goal-directed action with uncertainty re…
Training Variable Long Sequences with Data-Centric Parallel
Training deep learning models on variable long sequences poses significant computational challenges. Existing methods force a difficult tra…
The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure pre…
NL2SHACL-Bench: A Benchmark Suite for Natural Language to SHACL Translation
SHACL is a core technology for validating the conformance of RDF knowledge graphs (KGs). Yet, authoring SHACL shapes requires technical exp…
Dynamic Coalition Formation and Communication Pricing in Skill-Based Agentic AI Systems
Modern agentic AI systems combine multiple large language model agents with heterogeneous skills, yet most architectures either fix communi…
MetaSpace: Metamorphic Testing for Spatial Cognition in Embodied Agents
An embodied agent is an intelligent entity that interacts with its environment through a physical body. Currently, the evaluation of embodi…
When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains
As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide i…
TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations
AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is reco…
An AI Scientist that Doesn't Drift: Taste, Structure, and Falsifiable Findings in a Quadruped Navigation Research Loop
Autonomous research loops driven by large language models can run machine-learning experiments at scale but tend to drift toward local refi…
The Field Knows: Cross-Dimensional Geometry from Navigation to Black Holes
We introduce a continuous metric field framework trained by a single causal contrastive loss. The framework encodes a scene into coefficien…
TeXFix-Bench: An Empirically Grounded Multi-Format Benchmark for LLM-Based Document Source Repair
Scientific and technical writing depends on markup sources that must compile: LaTeX, Typst, and Markdown pipelines fail on missing delimite…
CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models
Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approache…
Controlled Memory Interference in Continual LLM Agents
Long-term memory enables AI agents to maintain continuity across sessions, personalize behavior, and evolve through accumulated experience.…
From Single Chatbots to Governed Agent Ecosystems: An Agentic AI Pattern Catalogue and Orchestration Framework for Mission-Critical Hospital Information Management Systems
Hospitals are racing to embed AI, while coping with the surge in adaptation of the technology in other industries, into the triage manageme…
Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns
Long-running molecular simulation campaigns require repeated continuation from saved states, provenance-aware progression, adaptive assessm…
Contextual Value Alignment via Multilayer Combinatorial Fusion
Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches…
Mendel G\"odel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution
Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. Howev…
An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus Photography
Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run i…
IntelliAudit: Using Large Language Models to Evaluate Audit Controls
IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This…
Towards Researcher Agents for Knowledge-Graph Question Answering
Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical…
Protecting patient privacy in clinical foundation models: Technical and legal perspectives
Clinical foundation models trained on large-scale patient data are increasingly used for decision support, screening, and public health. As…
QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing
Identifying a meaningful quantum speedup requires more than matching a classical problem to a familiar quantum primitive: the claim must pr…
Adaptive Two-Level Allocation of a Conserved Capacity Budget Across Locations and Service Classes
We study how to share a single conserved capacity budget across many locations and two service classes when demand is uneven, time-varying,…
Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation
LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverifi…
AndroidReality: How Far Are Mobile Agents from the Real World?
Mobile agents have achieved promising results on clean online benchmarks such as AndroidWorld, yet their performance often degrades sharply…
The Capability Ladder: A Curriculum-Modernization Framework for Workforce Readiness in the AI Era
Artificial intelligence is changing the task composition of computing work faster than curricula and training typically adapt. This is a cu…
Who Built This Model? Tracing LLM Lineage via Spectral Fingerprints in Weight Space
Open-weight large language models (LLMs) are increasingly developed through complex, multi-stage pipelines, leading to intricate lineage re…
CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensi…
CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift
A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We…
When the Judge Should Not Decide: Evidence-Locked, Non-Compensatory Selection Bounds LLM-Judge Failure in Reasoning Pipelines
An LLM judge deployed inside a reasoning pipeline does not merely measure quality, it decides which answer ships. We show that the cost of…
Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge
Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous str…
Back to the Future: A workbook time machine for spread sheet creation benchmarks
We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to creat…
SurgLAT: Surgical Latent Attention Tracking for Depth-Aware Robotic Laparoscope Control
Autonomous laparoscopic camera control requires continuous understanding of the surgeon's operative intent in dynamic surgical scenes, wher…
GRACE: LLM-Grounded Semantic Metric Spaces for Scalable Mixed-Data Clustering
Clustering mixed tabular data requires a unified metric space to bridge the inherent heterogeneity between continuous numerical measurement…
Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills
Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in outpu…
TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis?
Agent systems increasingly expose execution traces, yet telemetry that reveals a failure may still be inadequate for identifying where that…
GraphThink: Graph-Enhanced LLM Thinking for Long-Horizon Embodied Task Planning
Embodied agents using LLM-based planners often struggle with physical hallucinations, poor generalization to long-horizon tasks, and lack o…
When Is Benchmark Contamination Detectable? Information Limits and Power-Calibrated Audits
Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power. We f…
TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents
Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preve…
ZhuLong: Execution-Grounded LLM Agent for EDA Scripting with Offline API Self-Exploration
EDA scripting with tool-specific, often undocumented APIs remains a long-tail bottleneck that existing LLMs fail to address. This paper pre…
REIN: Bridging the Gap between Reasoning and Reliability via Reflection and Abstention Alignment
Large reasoning models (LRMs) are prone to hallucination, which undermines their reliability and poses challenges for safe deployment. Hall…
Locating Failure in Multi-Page Visually Rich Document Understanding: An Empirical Attribution
Multi-page visually-rich document understanding (MP-VRDU) requires managing evidence that is sparse, spread across pages, and often exceeds…
Directed Neuro-Symbolic Stochastic Execution for Verification of Distributed Parallel AI Programs
Distributed parallel Artificial Intelligence (AI) programs expose reliability gaps that conventional testing cannot close: parallel executi…
Guixu: Valuation-Driven Data Discovery for Autonomous AI Agents with On-Chain Attestation
Autonomous agents increasingly rely on external data to complete downstream tasks such as model training and decision support. However, exi…
KGCache: Amortized Subgraph Retrieval for KG Reasoning with LLMs
Large language models can answer knowledge-intensive questions more reliably when they are grounded with knowledge graphs, but systems such…
Self-Evolving Neuro-Symbolic Skills for Tool-Augmented Spatial Reasoning
Large vision-language models have achieved strong performance in multimodal reasoning, but they remain unreliable on fine-grained spatial t…
SCOUT: Self-Checking and Recovery-Aware Tool-Thought Agents for Ultra-Long Egocentric Video Reasoning
Ultra-long egocentric video understanding requires reasoning over temporally sparse evidence distributed across hours or days, challenging…
CyberAGENTS: Structured Autonomy for Agentic Gamified Learning in Cybersecurity
Gamification is especially effective in learning domains requiring active problem-solving and iterative skill-building, such as cybersecuri…
VDGR-RAG: Vectors, Directories, Graphs, and Reflection Are All You Need for Unified Reasoning over Hierarchical Enterprise Knowledge
Retrieval-Augmented Generation (RAG) is essential for enterprise knowledge question answering (QA), particularly in domains with complex pr…
Thought-Level Beam Search for Reasoning
Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current appr…
Legal Responsibilities Using Autonomous Agents For Artificial Intelligence
Recent incidents involving Artificial Intelligence (AI) agents, which were reported escaping their containment `unintentionally' to gain un…
The Authority Expectancy Effect in Multi-User Conflict
We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing eac…
Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE
Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the sa…
SkillSmith: Enhancing Locally Deployed Agents via Automatic Skill Construction and Evolution
LLM-based agent frameworks now act as personal assistants for multi-step tasks. Existing agent frameworks such as OpenClaw commonly follow…
Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities
Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynami…
JustLLMGRPO: Radiographic Control for Chest X-Ray Generation
Text-conditioned chest X-ray generation aims to synthesize realistic radiographs that faithfully depict specified findings. Existing work h…
SodaMem: Evidence-Grounded Temporal Graph Memory for LLM Agents
Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was on…
H2: A Dual Hybrid Semantic Data Lake Architecture for Medical Data Harmonization with Human-In-the-Loop verified, LLM Driven Metadata Annotation System
Medical data, by its nature, exhibit a high degree of heterogeneity on multiple levels ranging from (a) different modalities like images, t…
CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models
The key question in moral judgement is not simply whether someone chooses the "right" answer, but how they decide what matters most when mo…
Explore, Map, Remember, Decide: Are Embodied VLMs Ready for Safety-Critical Scenarios?
Theory of Space framework (ToS) assesses the spatial understanding of curiosity-driven Vision-Language Models (VLMs) under partial observab…
PATH: Next-Interval Prediction via Autoregressive Tree Hierarchy on Tabular Data
Interval prediction aims to achieve a target coverage level while producing intervals that are as short as possible. Many conformal regress…
Generative Models: Principles, Architectures, and Applications
Generative AI has emerged as one of the most transformative forces in modern artificial intelligence, reshaping how we create, imagine, and…
Think Deep, Speak Once: Relit, A Recursive Latent Implicit Transformer Framework
Chain-of-Thought (CoT) prompting has become the dominant paradigm for eliciting reasoning in Large Language Models (LLMs), yet it creates s…
Neurosymbolic Discovery of Algebraic Graph Constructions
There are several methods for searching for graphs with prescribed properties, such as SAT solvers and specialized generators. These method…
Constraining ontology mappings using metaphysical choices
In this paper we discuss the foundations behind a novel methodology for the validation of semantic mappings between different data sources…
Improving Constraint Models with LLM Agents
The runtime of Constraint Programming (CP) solvers is highly sensitive to modeling choices, such as symmetry breaking, implied constraints,…
TokenPrint: A Calibrated Token-Space Fingerprint for Language-Model Provenance
Establishing the provenance of a language model---including its base checkpoint and possible overlap in training distributions---is a gover…
Long SKILL Compliance as Logical Reasoning: Closure-Grounded Detection with Scaling-Guided On-Policy Distillation
The increasing complexity of enterprise business scenarios has promoted the widespread adoption of long SKILL documents in agent systems, p…
A Unified Framework for Dynamic Reward Shaping in Reinforcement Learning
Sparse, delayed, and weakly informative rewards remain central obstacles to efficient reinforcement learning. Reward shaping addresses thes…
When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs
Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, ment…
Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) Brachytherapy
The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and pers…
Matching Supervision to the Student's Learning Capacity: A Unified Framework for On-Policy Self-Distillation
On-policy self-distillation (OPSD) improves the reasoning abilities of LLMs by internalizing privileged context into model parameters throu…
Large Multimodal Agents for Intelligent Transportation Systems: Architectures, Evidence, and Deployment Challenges
Large multimodal agents (LMAs) are increasingly proposed for intelligent transportation systems (ITS), but existing studies often conflate…
Quantization Degradation in Large Language Models: A Signal-Noise Perspective
Post-training quantization reduces the deployment cost of large language models, yet how severely a quantized model degrades is not determi…
Janus: An Algorithm-Evaluator Co-Evolution Framework for LLM-Driven Discovery under Expensive Evaluation Budgets
LLM-driven program discovery relies on rapid evaluator feedback, but many scientific and engineering tasks require high-fidelity simulation…
A Minimal $\kappa$--$\tau$ Logic for Risk-Sensitive Abduction
Standard approaches to abductive reasoning can retain multiple candidate explanations, but they do not generally combine explicit compositi…
Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models
Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans. Yet it remains unclear whether…
Illusion of Alignment: Detecting Hidden Disagreement in Collaborative Dialogue
Collaborative dialogue can end with apparent agreement while participants still differ on goals, assumptions, or execution plans, creating…
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment
In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Exi…
Metanormative Theory for RL-Based Moral Agents
The overlapping disciplines of machine ethics and value alignment are concerned with designing artificial agents that are aligned with huma…
LatticeMind: A Conflict-Aware Memory Primitive for Multi-Agent Systems
Multi-agent LLM systems often fail not for lack of candidate answers, but because they have no persistent mechanism for deciding which inco…
A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences
This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans…
FemWear: A Specialized Wearable Foundation Model for Women's Health
General wearable foundation models are pretrained across broad sensor streams and populations, but are not designed around women's-health t…
SuperLocalMemory 4.0: The Governed Memory Operating System for AI Agents
AI agents are becoming shared infrastructure, yet durable memory is commonly assembled from separate retrieval, governance, and operational…
Your Prompt Is Not the Only Prompt: How Much Do LLMs Weight Structured-Output Schema Descriptions?
Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extrac…
OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents
Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift c…
Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios
Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reasoning, and decision making in…
Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders
Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations refl…
Mitigating Over-Personalization in LLMs via Structured Memory
Conversational assistants increasingly rely on persistent long-term memory to personalize responses across sessions. However, when stored u…
Query-Only Backdoor Attacks on Self-Evolving Skills via Trajectory Poisoning
Agentic skills improve large language model (LLM) agents by encoding reusable procedures for complex tasks. However, manually authored skil…
StructReward: Efficient Structured Process Rewards for Self-Correcting Multimodal Reasoning
Reinforcement learning with verifiable rewards (RLVR) has emerged as an effective approach for improving multimodal reasoning. However, mos…
LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving
As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Ena…
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents
Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly whil…
CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception
Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has…
Estimating Uncertainty in Galaxy Morphology Classification
Astronomers classify galaxy morphology to investigate cosmic evolution. While deep foundation models are increasingly utilized in Galaxy Mo…
Forgotten History or Test-of-Time? Retrospect and Prospect on RAG from an IR Perspective
Retrieval-Augmented Generation (RAG) is widely regarded as a novel paradigm born from the limitations of large language models (LLMs)--a me…
TRACE-Memory: Public-Conditioned Retrieval and Utility-Aware Evidence Admission for Personalized Generation
Personalized generation systems retrieve user history by request--memory relevance and inject it into the model context. Yet relevant histo…
What Keeps Agent Skills from Being Reusable? Evidence from 138K SKILL.md Files
Under the current standard, Agent Skills are SKILL.md files that combine instructions with supporting resources, enabling Large Language Mo…
Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses
Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the mode…
LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMs
The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design emb…
Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation
Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making t…
Yesterday's Shield, Today's Spear: A Self-Evolving Safety Guardrail in Production
Deployed LLM safety guardrails are predominantly static: trained once and frozen at release, while new jailbreak techniques and previously…
HoloAegis: Frozen Representation, Topological Inference: Minimally Parametric Safety Manifolds for Zero-Shot LLM Guardrails
Current LLM safety guardrails face a fundamental tension: fine-tuning distorts pre-trained representations while generative judges incur pr…
TrustRoboReward: Preference-Ordered Isotonic Score Editing for Multi-Paradigm Robot Reward Models
Reward models are a bottleneck for reinforcement learning in embodied AI. Long-horizon robotic manipulation requires scalable vision feedba…
MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models
Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Tho…
Understanding Calibration and Truncation Error Propagation in Training-Free Low-Rank Compression for LLMs
Training-free low-rank compression frameworks have been gaining prominence for LLM compression given their effectiveness in reducing model…
Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding
Evolving documents, such as laws, tax codes, and software documentation, are amended, replaced, and sometimes reverted over time, so a ques…
Reproducing and Stress-Testing Two Approaches to LLM Reasoning Reliability: Test-Time Probability Aggregation and Logic-Representation Editing
We independently reproduce two recent methods for making large language model (LLM) reasoning more reliable, and stress-test them across do…
Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity Optimization
Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and int…
Deep probabilistic logic programming for diagnostic reasoning from incomplete information: A case study in stroke detection
In medical applications, raw data is frequently associated with significant privacy concerns, lending particular importance to the encoding…
VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference
Recent advancements in Speech Large Language Models have demonstrated remarkable capabilities in understanding complex audio tasks. Despite…
FailForge: Distilling Procedural Competence from Persistent Failures into Code Agents
Rejection sampling fine-tuning (RFT) is widely used to train code agents by generating trajectories on verifiable software engineering task…
SDDBMs: Soft Denoising Diffusion Bridge Models
Diffusion bridge models leverage Doob's \(h\)-transform to construct stochastic transports between arbitrary endpoint distributions, and ha…
Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents
To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies fo…
ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration
Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide…
Walking through Discussions: A Mobile Visual Analytics System for In-Situ Group Discussion Analysis
Group discussion-based teaching is widely used to foster collaborative learning, yet teachers in physical classrooms often struggle to simu…
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under…
MedCalc-R1: Knowledge-Guided Reward Framework for Medical Mathematical Reasoning
In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically…
UniMoMo: Expert Merging-Based MoE Acceleration for Large Recommendation Models
Sparse mixture-of-experts (MoE) layers expand recommendation capacity through conditional computation, yet a trained checkpoint still store…
A QUBO-Inspired Computational Framework for Airport Landside Bottleneck Diagnosis and Dynamic Dispatch Optimization
Airport landside traffic centers connect terminal arrivals with taxis, ride-hailing vehicles, private cars, buses, metro services, parking…
Can Open-Weight Models Compete on Financial Text Comprehension?
Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet thei…
Smart Compaction: Predicting Compaction Utility from Lakehouse Table Metadata
Open lakehouse table formats accumulate small data files over time, which degrades query performance. Deciding when compaction is worthwhil…
SkillReason: Reasoning-Enhanced Agent Skill Retrieval for Implicit User Requests
Large language model agents increasingly rely on reusable skills to extend their capabilities beyond parametric knowl- edge. However, retri…
The Scaffolding Matters More Than the Interface: A Controlled Comparison of MCP and CLI Tool Use Across Seven Agent Scaffoldings, Five Language Models, and One Software Task
How much an AI coding agent costs to run can depend more on the agent scaffolding that drives it than on the interface through which it rea…
Branch2Skill: Efficient Skill Evolution Through Reasoning Trees
Skill evolution improves agent skills through feedback over time, with failed trajectories often providing informative signals by revealing…
A Structural Dynamics Graph World Model: Unified Modeling, Constrained Rollout, and Interpretable Calibration
The state evolution of a complex system arises jointly from object laws, relational propagation, domain conservation, and unmodeled error.…
EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility
Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependa…
PluginEval: A Diagnostic Benchmark for Fine-Grained Error Attribution in Function Calling
Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents. Current benchmarks face…
AI Evaluation Should Measure Verification Cost, Not Correctness Alone
The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to v…
FitAQA: A Benchmark of Fitness Action Quality Assessment for Multimodal Large Language Models
Fitness Action Quality Assessment (AQA) is important for intelligent sports training, yet the capabilities of Multimodal Large Language Mod…
Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models
Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response bur…
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification
Large language models (LLMs) increasingly serve as data-driven reasoners, yet their chains-of-thought (CoT) can be unfaithful even when fin…
Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs
Omni-modal LLMs jointly process audio, video, and text, but long multimodal sequences incur substantial prefill and KV-cache costs. Existin…
Improving Generalization Robustness of Multimodal RLVR
Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simp…
Three Generations of Healthcare IT: From the Digital Record to the Computable Care Process
Objective. Healthcare IT is usually organized by the technologies it adopts. We instead organize it by the unit of information a system mak…
Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models
Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, incons…
PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary
Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, ty…
Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a…
Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration
The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception…
AquiLLM: An Architecture for Supporting Tacit Knowledge Capture in Research Groups
Recent advances in retrieval-augmented generation (RAG) and large language models (LLMs) enable researchers to integrate AI into scientific…
Full-bandwidth transformer
Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attenti…
LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing
Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preferenc…
From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings
Imaging device downtime is a major barrier to healthcare delivery in low- and middle-income countries (LMICs), often driven by limited acce…
Decoding Phenotypes: A Framework for Fusing Genomic Language Models and Neuroimaging
Neuroimaging and genetic testing are two important clinical references for nervous system diseases, offering complementary diagnostic infor…
Integrated Multimodal AI System for Retrieval-Augmented Reasoning, Object Sensing, and Damage Analysis
This work presents a unified multimodal AI system for damage assessment that integrates retrieval-augmented generation (RAG) models, therma…
Not an A11y: How Android Accessibility Exposes Mobile AI Agents to Indirect Prompt Injection
The rise of autonomous AI agents represents a major paradigm shift in how users interact with mobile devices. Frameworks such as MobileRun…
Depth-Aware Implicit Neural Representation Priors for 3D Gravity Inversion
Gravimetry images subsurface density contrasts associated with geological structures, geothermal systems, and intrusive bodies. Recovering…
Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression
Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compr…
CoRe-UIE: Rethinking Coexisting and Region-wise Degradation for Underwater Image Enhancement
Underwater images often suffer from diverse and coexisting degradations, including color distortion, scattering haze, texture attenuation,…
Context Is Not Authority: Structured Runtime Governance for Financial Market Agents
Financial agents can turn correct context into an unauthorized effect: a customer-facing commitment, trade, or deployed policy. We present…
PolicyKG: An Agentic LLM Pipeline for Translating Institutional Policies into SHACL Knowledge Graphs
Institutional policies stay in natural language while the systems that check compliance demand machine-readable constraints. Bridging that…
DualCert: A Solver for the Traveling Salesman Problem with Constraint-Coupled Learning
Large traveling salesman problem (TSP) instances require a solver to allocate limited computation while preserving the validity of its outp…
A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition
Mixed emotions represent a clinically relevant but still underexplored target for automatic emotion recognition. EEG provides millisecond-l…
Who Bridges Safety? Identifying and Targeting Cross-Lingual Shared Safety Pathways
Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy…
Different Feedback, Different Updates: Selective Self-Learning from User Interactions for Large Language Models
User feedback offers natural supervision for persistent LLM improvement, but a single message may support multiple behavioral changes with…
RAVEN-Eval: Rubric-Guided Automatic Evaluation for AI Video Generation Models Based on LMM Preference Judgement
AI video generation has advanced rapidly and entered widespread commercial use. As a result, quality differences among videos produced by s…
Motif 3: Technical Report
We introduce Motif 3, a decoder-only Mixture-of-Experts language model with 314 billion total parameters and 13.2 billion activated per tok…
MELLON - Multimodal Enhanced LLM for Online Navigation
Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are eithe…
RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning
Aligning Large Language Models (LLMs) for open-ended tasks is challenging because responses must satisfy multidimensional criteria without…
ChronoState: Hidden Elapsed-Time Conditioning for Temporal-State Action Selection in Frozen-Backbone Language Models
Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration…
TRACE: TRajectory Attribution for Automated Context Engineering
Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contai…
CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment
Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general p…
From Relevance to Execution Utility: Reward-Aware Dynamic Execution Gating for Skill-Based LLM Agents
Agent skills are increasingly used to equip large language model (LLM) agents with reusable procedural knowledge. Although recent work has…
Agentic Router: An Execution-Grounded Continual Learning Approach With Memory
Large language model (LLM) agents provide a promising interface for command-line-based network operations, but a plausible command may stil…
Structure-Preserving Uncertainty Propagation in First-Order Proof Search
GK is a query-directed first-order prover that extends ordinary resolution-based proof search with explicit positive and negative claims, n…
Signature-Guided Capacity Occupancy for Dense Expert Merging
Dense expert merging combines domain-specialized language models into one single checkpoint, typically by admitting task-vector support in…
CRUISE: Vision-Language Model-Guided Uncertainty-Aware Cross-Modal Sensor Fusion for Robust Autonomous Driving
Modern autonomous vehicles are equipped with multiple sensors, such as cameras, LiDAR, and radar, for comprehensive environmental perceptio…
Omni2LoRA: Coherence-Preserving Parametric Memory for Efficient Omni Language Models
Omnimodal language models (OLMs) enable unified audio-visual understanding, but processing long joint token sequences makes inference compu…
SafeSceneReason: A Multimodal Reasoning Benchmark Connecting Industrial Hazards with Accident Knowledge
Industrial-safety understanding requires more than detecting workers, equipment, and personal protective equipment. Models must also assess…
An Explainable GNN Framework for Component-Level Anomaly Diagnosis
Industrial processes are complex systems composed of multiple interacting sensors that generate multivariate time series (MTS). Detecting a…
Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution
Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirel…
SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance
LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provid…
Business Truth, not SQL Accuracy: A Rule-Gated 7B Analytics Agent Outperforms a Direct-Prompted 32B Baseline
LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business defini…
Privileged Likelihood Is Not Automatically Value: Three Checks for Token Credit in On-Policy Self-Distillation
Outcome verifiers score completed reasoning traces but do not assign credit to intermediate tokens. Privileged self-distillation attempts t…
Entropy-based Code Adversarial Translation for Real-world Repository Migration
LLMs have demonstrated strong capabilities in code generation and automated program repair, but migrating an entire repository rarely produ…
P$^{3}$: Joint Program-and-Proof Planning for Verified Code Generation
Verified code generation asks a large language model (LLM) to generate both an executable program and a machine-checkable proof that the pr…
MMArch: Benchmarking Multimodal Reasoning Grounded in Architectural Evidence
Multimodal large language models (MLLMs) perform strongly on engineering imagery, yet existing benchmarks mostly test drawing recognition,…
ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons
Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shoppin…
CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation
A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter…
Linearized 2-Simplicial Attention
We present a linearized form of 2-simplicial attention by rewriting the trilinear score as an inner product between a composite query and a…
ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management
While mathematical models act as vital decision support systems for operational Air Traffic Flow and Capacity Management (ATFCM), existing…
CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning
On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs,…
GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models
Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines the value of landslide mapping fo…
Control-Oriented Scenario Tree Construction through Reinforcement Learning
Multistage stochastic model predictive control (MPC) handles uncertainty by optimizing over a scenario tree, a finite branching approximati…
LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling
Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a blac…
CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits
Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent…
OpenLoopEvolve: A Verifiable Self-Evolution Framework for Loop Policies in Long-Horizon Complex Tasks
Long-horizon complex tasks require agents to repeatedly observe states, formulate plans, invoke tools, verify results, and recover from fai…
KVDiagnosis: A Diagnostic Benchmark for KV-Cache Compression in Long-Context Language Models
KV-cache compression reduces long-context memory, but aggregate task scores reveal neither which correct executions fail nor why. We presen…
Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over…
Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity
Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbid…
From Prompt to Harness: Coderlet from Scratch
A model alone does not determine how a programming agent acts. What the model sees, how actions enter the environment, how feedback returns…
Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents
Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strate…
Renormalising Generative Models for Active Inference: Foundations, Derivations, and Verification
Active inference offers a unified framework for perception, learning, and action, but scaling discrete active-inference models to rich spat…
One Adapter Pair per Model: A Universal Activation Interface for Language Models
Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language int…
verdi: retrieval is not transfer for continual world model optimization
Foundation world models have made remarkable progress in planning, simulation, and embodied intelligence. However, optimizing a pretrained…
Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents
External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Y…
The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games
LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As…
ElasticBack: Stealthy Conditional Backdoor in LLM-Agent Skills via Coupled Trigger-Rule Optimization
Agent skills, bundles of instructions and resources that an LLM agent loads on demand, form an emerging supply chain where a single poisone…
CoRCi: Cross-Reconstruction of Coherent Interests Modeling in Cross-Domain Sequential Recommendation
Cross-Domain Sequential Recommendation (CDSR) aims to alleviate data sparsity by transferring dynamic user interests across related domains…
ICM Out! Better Tournament Strategy from Computed Continuations, vs. Solvers and LLMs
The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against thos…
From Sweep to Seam: Interleaved Cross-Block Post-Training Quantization
Compressing large language models to two bits or fewer is increasingly feasible through block-wise post-training quantization; cross-block…
Adaptive Sequential Test Planning for Multi-Mechanism Reliability Qualification via Bayesian Monte Carlo Tree Search
Reliability qualification of advanced semiconductor devices requires sequential stress decisions that balance characterization objectives a…
Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
Self-evolving agents are usually built around prescribed optimization pipelines: the framework decides how to gather evidence, revise a per…
Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics
Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental…
Hallucination-Free GUI Grounding via Regression-Free Layout-Aware Matching
GUI agents are shifting from metadata-dependent large language models to purely visual multimodal large language models (MLLMs) that operat…
Open Evaluation Agent: Efficient and Promptable Evaluation of Visual Generative Models
Recent advances in visual generative models have enabled high-quality image and video generation, but evaluating these models often demands…
Adaptive Semantic Capacity Allocation for Parallel Generative Recommendation
Autoregressive semantic ID recommenders are constrained by expensive beam-search decoding, which limits the practical length of item identi…
Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
Predicting the answer to interventional ``what if'' questions --- the outcome of an action never taken --- requires a \emph{mechanistic}, c…
Matryoshka Language Model Suites
Training a language model suite classically requires training each model separately and serving them independently. We improve both trainin…
Second-Order Muon Done Right: A Principled Marriage of Spectral Geometry and Curvature
Muon's polar update is exact for an unweighted spectral geometry. We introduce GO-MUON, which uses a matched data-dependent geometry and re…
AirFlow: Context Preserving and Multi-Rate State Modeling for Air Quality Forecasting
Accurate air quality forecasting is essential for public health and urban environmental management, but remains challenging because polluta…
CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation
Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these…
Mismatch Matters: On-Policy Distillation Beyond Token Agreement
On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerat…
CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems
The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environment…
Agentic Auto-Research is Fuzz Testing
Autonomous research agents can generate experiments faster than researchers can validate them. Researchers have responded by scaling the pr…
Towards Expert-level Medical AI for Real-time Video Consultations
Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of il…
ArchAgent v2: A Case Study with the Data Prefetching Championship
Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarc…
SHE: Trajectory-driven Safety Harness Evolution for LLM Agents
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memor…
DSLE: A Learning Environment for Dark Souls Boss Encounters
We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remas…
GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis
Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains su…
From Trajectories to Evidence: Auditable Experimental Records for Industrial Research Agents
Research agents increasingly conduct multi-round machine-learning experiments in industrial recommendation settings and retain the resultin…
Coordinated incentives in AI-generated misinformation governance
With the rapid diffusion of AI-generated content, AI-driven misinformation is becoming increasingly pervasive and difficult to govern, unde…
Application of Artificial Intelligence for Fraudulent Banking Operations Recognition
This study considers the task of applying artificial intelligence to recognize bank fraud. In recent years, due to the COVID19 pandemic, ba…
Positioning Generative Artificial Intelligence in STEM Assessment: When to Require, Scaffold, or Restrict Its Use
Generative Artificial Intelligence (GenAI) presents a governance challenge for STEM assessment. Unrestricted access can enable task outsour…
Designing for Ethical AI: HCI Feature Considerations to Improve Fairness and User Experience in AutoML use for Human Resources
This thesis examines the fairness of Automated Machine Learning (AutoML) tools in human resource hiring systems through the combined lenses…
Cross-Model Humor Preference Modeling with Cards Against Humanity
This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Hum…
Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents
Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players…
How to Ask the AI: A User Perspective Survey for Large Language Model Prompting
AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses…
EmoPatient: An Emotion-Directed Patient Simulator for Realistic Palliative Care Communication Training
Effective communication during palliative care discussions is a critical clinical skill, yet training clinicians to manage complex patient…
Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation
Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for…
Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients
The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based…
Harnessing Abundance: A Generativity Perspective on Human-GenAI Collaboration
Research on human-GenAI collaboration yields conflicting findings: GenAI can enhance creativity yet reduce collective diversity, with uneve…
Innovating with Generative AI: A Human Bottleneck Framework
We propose a human bottleneck perspective for understanding how generative AI transforms the innovation process. The central premise is tha…
JaleesBench: Are AI Assistants Good Spiritual Company?
Large language models are already advisors to millions of people of faith who bring them real decisions. The pressing question for a person…
PIVOT: Preference-based Intervention Vectors for Pedagogical Tutor Steering
LLMs are increasingly used for conversational tutoring, but effective tutoring requires more than correct answers. Tutors must choose when…
How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare
Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (L…
EMMR: Emotion-Mediated Multimodal Reasoning for Personality Assessment in Asynchronous Video Interviews
Asynchronous Video Interviews (AVIs) have become increasingly popular for personality assessment. Recent large language models (LLMs) have…
Representation Matters in Longitudinal Affective Computing
Longitudinal, in-the-wild, wearable sensing yields day-level physiology, sleep, activity, and environmental streams, whereas affect and cog…
KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service Assistants
Mass religious gatherings such as the Kumbh Mela concentrate tens of millions of people into a single region over a few weeks, producing in…
From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations
Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions larg…
Unified Hallucination Fuzzing for Multimodal Large Language Models
Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stak…
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing…
WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management
Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains di…
Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. E…
Ultraconstructive Model Theory via Bounded Adversarial Finite Structures
Ultraconstructive Model Theory (UCMT) replaces idealized satisfaction, at finite compu- tational scale, by bounded adversarial survival. A…
Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
Multi-modal large language models (MLLMs) integrate heterogeneous modalities through modality alignment and fusion, enabling stronger under…
An evolutionary model of animats with VLM-based subjective evaluation
In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language Model (VLM) into the fitness e…
NeuroPilot: An Agent-Driven Smart Pipeline for Processing, Quality Control, and Managing Neuroimages
Transforming raw neuroimage archives into analysis-ready derivatives relies on three brittle stages: data standardization, modality-specifi…
Performance of large language models in the optical diagnosis of colorectal polyps
Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal larg…
MOSAIC: Adversarial Co-evolution of Specialist Heuristics and Problem Instances for LLM-based Automated Heuristic Design
Automated heuristic design (AHD) with large language models (LLMs) has produced strong heuristics for combinatorial optimization problems (…
DarwinX: Evolving Agent Harnesses Through Natural Selection
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement…
Generalizing deep reinforcement learning across cable-driven parallel robot configurations with actuator-level policies
Cable-driven parallel robots (CDPRs) present diverse configurations and complex control challenges, which can be addressed by deep reinforc…
Learning an Interior Layout Policy in a Domain Specific Language Action Space
Indoor scene layout generation is a challenging task in interior design. Existing methods often oversimplify the task by reducing room cond…
P2Voxel: Pyramid Pivot Voxelization for 3D Mesh Tokenization
Triangle meshes provide explicit and accurate surface geometry, yet their irregular topology connectivity makes 3D mesh tokenization a geom…
MasDrift: Benchmarking Authorization Preservation Across Multi-Agent Architectures
Multi-agent systems (MAS) decompose long-horizon tasks across supervisors and subagents, but delegated goals do not necessarily carry their…
AeroDPO: Unleashing Lightweight UAV Navigation with High-Fidelity Perception and Automated Preference Optimization
Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments. Recent mi…
Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth
In digital dentistry and oral surgery, the registration of jawbone CT and intraoral scanner (IOS) data is essential for integrating interna…
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems
Conversational assistants increasingly recommend follow-up edits to help users continue a task. Existing systems primarily target text-only…
Temporal Generalization in fNIRS-Based Autism Classification: A Cross-Time-Window Transfer Benchmark
Functional near-infrared spectroscopy (fNIRS) is a promising modality for autism spectrum disorder (ASD) classification, yet existing appro…
Latent-Frequency Validity: Fast Spectral Editing with Screened Video-VAE Transfer Operators
Direct spectral editing in video-VAE latents can control noise, flicker, smoothness, and frequency content without a decode--filter--reenco…
COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping
Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existin…
BRACE: Taming Sharp Irregularities via Barycentric Rational Forecasting for Fast Diffusion Transformers Inference
Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To alleviate their mas…
Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects
A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse-…
Geometry Beats Estimated Depth: RGB-Only Multi-Camera 3D Tracking under Sim2Real
The AI City Challenge 2026 Track 1 evaluates multi-camera 3D perception in large indoor warehouses under a synthetic-to-real (Sim2Real) set…
Multi-Branch Policy Optimization for Multimodal Large Language Models
Group-based reinforcement learning methods for multimodal large language models typically rely on trajectory-level credit assignment that a…
Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning
Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose…
Scaling Inherently Interpretable Language Models
Interpretability is often treated as a tax on capability: language models are trained as opaque systems, then explained after the fact, wit…
Enhanced Real-Time 6-DOF Extended Reality Catheter Tracking for Evaluating Potential Improvement in Efficiency, Precision, and Depth Perception for Cardiac Interventions
Despite advances in 3D ultrasound, most percutaneous cardiac interventions still rely on 2D visualization, limiting depth perception and sp…
Hit Selection Using SSMD-Based Machine Learning Performance Metrics in High-Throughput Screening Assays
High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary…
PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue
LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advanc…
Adversarial Attacks on Deep OCR Systems
Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-conte…
SkillConsist: Detecting Inconsistencies in Agent Skills via Bidirectional Graph Alignment
Agent Skills provide reusable capabilities to LLM agents. Agent Skill inconsistencies can expose undisclosed dangerous behavior or cause wr…
Keep It Simple: Multi-Key Episodic Memory Retrieval for Ultra-Long Video Understanding
When videos extend from hours to days, directly processing them end-to-end becomes impractical for current Multi-modal Large Language Model…
CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting
Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short obse…
SpikeWorld: Fast-State Adaptation for Frozen Spiking World Models
A predictive model receives a self-supervised signal whenever the consequence of an action is observed. Using that signal after deployment…
LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs
Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendat…
Complete, Scalable, and Robust Prioritized Planning for Multi-Robot Ordered Storage and Retrieval at Maximum Capacity
Automated warehouses face a fundamental trade-off between maximizing storage density and achieving high retrieval throughput. While puzzle-…
LoRSA: Toward Generalizable Parameter-Efficient Fine-Tuning for Biomedical Downstream Tasks
Parameter-efficient fine-tuning enables the adaptation of vision foundation models to biomedical tasks under limited computational resource…
Multi-Task Consistency-based Detection of Adversarial Attacks
Deep Neural Networks (DNNs) have found successful deployment in numerous vision perception systems. However, their susceptibility to advers…
CFD-Guided Detection of Concept Drift in Multimodal Physiologic Signals
Cardiovascular AI models can classify clean elec- trocardiogram (ECG) signals, but real wearable signals change because of motion, breathin…
The Anatomy of a Prompt Injection: A Component Model for Structured Analysis
Four years after prompt injection was first identified in 2022, attacks are still predominantly documented as verbatim strings rather than…
Shape Mutating Expert Compression:LorExperts and BTExperts
Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing th…
Distilling CT Foundation Models into Editable Concept Bottlenecks for Lung Nodule Malignancy Prediction
Foundation models provide transferable CT representations, but predictions based directly on these embeddings are difficult to interpret. W…
Vision-Language Grounding as Bidirectional Concept Correspondence
Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localiza…
Router Sensitivity Under Lightweight Fine-Tuning Identifies Prunable Experts in Mixture-of-Experts Models
Mixture-of-Experts (MoE) models decouple total parameters from per-token compute, but deployment still requires storing every expert. Recen…
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in hi…
Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical Limits
Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate fo…
Spectral Outliers Reveal Dominant Learned Structure in Transformer Attention
We apply Marchenko-Pastur (MP) random matrix theory to pre-trained attention weights in order to separate each projection matrix into a ran…
Second Order Drifting Models
Drifting models are a recent class of one-step generative models that evolve the model distribution during training using a predefined samp…
ScaleSense: Cost-Intelligent Scaling Framework via Learned Resource Estimation in Alibaba AnalyticDB
Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal res…
Metadata Reconstruction from Values Alone: Recovering Column Semantics in Undocumented Warehouses
Text-to-SQL benchmarks ship schemas whose column names already say what the columns mean. Production warehouses are the inverse: cryptic id…
Persistent Semantic Entities in Tool-Augmented LLM Systems
Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent bo…
EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference
Load Balancing has emerged as a critical problem in expert-parallel distributed inference of Mixture-of-Experts (MoE) models. As routing di…
Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
Reasoning language models increasingly use test-time compute to improve performance, but existing evaluations typically study this compute…
Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design
Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their outp…
Evidence-Grounded Forensic Reasoning for Detecting and Grounding Multi-Modal Media Manipulation
Fake news increasingly relies on cross-modal image-text forgeries, making transparent and verifiable reasoning chains an urgent need for De…
Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models
Time series forecasting (TSF) plays an important role in a wide range of real-world applications. Recently, time series foundation models (…
Evidence-RL: Towards Evidence-intensive Visual Reasoning
Vision-Language Models (VLMs) should answer from concrete image evidence rather than language priors, dataset shortcuts, or irrelevant visu…
Prompt Embedding Probes (PEP): Hallucination Detection in LLMs from Hidden States
Large language models (LLMs) can generate fluent and useful responses but remain prone to hallucinations. We introduce Prompt Embedding Pro…
DA-NBV: A Direction-Aware Next-Best-View Planner for Efficient 3D Reconstruction of Ships at Sea
Accurate 3D reconstruction of ships at sea is important for maritime supervision, damage assessment, and autonomous maritime operations. Al…
Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model Families
Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model…
Tools to Explain Neural Networks for Power System Dynamics
This paper presents, for the first time in power systems literature to our knowledge, analytical tools to explain the training performance…
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
Current end-to-end speech dialogue models are primarily optimized for mainstream languages and remain limited in low-resource dialect scena…
HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection
Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision. Current…
Effect of Abstractions and Prompting Strategies on LLM-Guided High-Performance Optimizations
Code performance optimization is a vital aspect of modern software development, as it enables faster response times and reduced resource us…
Adaptive Symmetry Discovery for Dynamical System Identification
Dynamical systems model trajectory data generated by fixed underlying dynamics, with applications ranging from biology to physics. Especial…
Defending Retrieval-Augmented Intrusion Detection Against Knowledge Poisoning and Prompt Injection
Retrieval-Augmented Generation (RAG) enables large language models to classify network flows and generate human-readable incident reports b…
NeuPAT: Neuron-aware Plasticity Allocation Tuning for Language-Preserving MLLMs
Multimodal expansion of large language models (LLMs) enables new perceptual capabilities but often compromises the language intelligence ac…
Hierarchical Multi-Task Federated Learning in VANETs
Vehicular Ad hoc Networks (VANETs) increasingly rely on federated learning (FL) to enable collaborative intelligence without sharing raw se…
Compositional Threat Analysis of Latent Compromise in LLM Agent Systems: The Order 66 Scenario
In the fictional Order 66, catastrophe does not arise from a powerful command alone: a trusted population is preconditioned, a short direct…
Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching
Cross-modality medical image translation can reduce the burden of multi-modal acquisitions, yet the field remains constrained by two couple…
DoGMA: A Central-Dogma-Guided Foundation Model for Multi-Omics Alignment and Multi-Task Learning in Oncology
Attention mechanisms have been widely utilized in modern deep learning, and many existing multi-omics models inherit their conventional use…
Exact Zarankiewicz Values On Two Finite Frontier Slices
The Zarankiewicz number Z(m,n,s,t) is the maximum number of edges in a bipartite graph with parts of orders m and n containing no copy of K…
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytel…
STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs
Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of s…
$\texttt{DisMorph}$: learning to disentangle technical distortions from true biological change
Longitudinal MRI enables sensitive measurement of structural brain change for studying aging and neurodegenerative disease. Deformable imag…
A Grounded and Decomposed Framework for Relation-Level Hallucination Evaluation in Abstractive Summarization
Abstractive text summarization systems frequently generate fluent yet unfaithful summaries by fabricating or distorting relationships betwe…
Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry
Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial i…
Targeted Counterfactual Fingerprinting for Black-Box LLM Ownership Verification
Large language models (LLMs) are high-value assets that can be derived through redeployment, fine-tuning, quantization, or further alignmen…
FreSH: Frequency-Segmented Hierarchical Multi-Expert Framework for Multivariate Time Series Classification
Multivariate Time Series Classification (MTSC) demands models that can effectively capture complex temporal patterns across multiple scales…
VTO: Visual Tool Orchestration for Video Anomaly Detection
Video anomaly detection (VAD) is a critical yet challenging task due to the complex and diverse nature of real-world scenarios. Traditional…
Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations
LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making…
On the Robustness of LLMs' Internal Representation of Code Correctness
Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no sur…
Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?
Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows…
Frequency-Domain Dual-Branch Fusion for Medical Visual Question Answering
Medical Visual Question Answering (VQA) requires aligning subtle visual evidence, including lesion texture, boundary sharpness, and diffuse…
Open-World Semantic Segmentation with Sensitivity Modeling
Modern vision systems must operate in "open-world" settings, where models must recognize known categories and detect unseen or anomalous co…
Three Necessary Principles for Self-Supervised Visual Representation Learning
We argue that learning visual representations without labels requires a training signal jointly complete across three non-overlapping objec…
Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core Evolution
We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through revie…
PRISM: A Predictive Protocol for Permutation Optimization via Landscape Diagnostics
Permutation optimization arises whenever the components of a system are fixed but their ordering affects performance. We introduce PRISM, a…
Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books
Audio dramas weave dialogue, sound effects, and music into immersive stories. Creators often adapt books into audio dramas, but this proces…
From Product Search to Preference Articulation: The Economics of Agentic Commerce
Generative AI is shifting digital commerce from browsing toward agentic search, in which consumers delegate product discovery to AI agents.…
Does a Toehold Make a Bidder Bolder? Preemption and Multiplicity in Multi-Round Takeover Auctions
A bidder can quietly buy a stake in a company before making an offer for it. That stake, a toehold, is supposed to pay for itself twice: it…
Abstracted Away: Resisting Alienation and Ungrounded Abstraction in AI Research Communities
Logics of abstraction in computational AI research often push important forms of knowledge and reflection aside: dominant standards of legi…
Human-Guided Causal Knowledge Injection for Virtual Cells
Virtual cells employ machine learning models to simulate and predict cellular behaviors, serving as a critical computational framework for…
FSTC-Encoder: Feature--Spatial--Temporal Correlation Learning for Generalizable RF Sensing
Heterogeneous RF sensing differs substantially in feature structure, spatial layout, and temporal scale, making existing models difficult t…
Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction
Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other…
Hidden Language Consistency Phenomena in Reasoning LLMs
Multilingual reasoning models are commonly evaluated by whether they arrive at the correct answer, but not by whether they preserve the int…
Calling the Bluff: Detecting Ever-Shifting Harmful Chat Dialogue via Ordered Reasoning Chain Regularization
Harmful chat dialogues are ever-shifting through type-shifting and lexical evasion, yet we find they share invariant principles, i.e., an O…
Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction
Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets…
Halpern Iteration Achieves $\tilde{\mathcal{O}}(\epsilon^{-1/p})$ $p$th-Order Oracle Complexity for Monotone Variational Inequalities
We study second- and higher-order methods for solving smooth monotone variational inequalities (MVI). Monteiro and Svaiter (SIAM J. Optim.,…
SkillsMetric: Mapping the Detection Boundary of Static Analysis for Malicious Agent Skills
Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their securit…
SuperNeuroMAT: An Efficient Matrix-based Simulator for Spiking Neural Networks
Spiking neural networks (SNNs) offer a promising pathway to energy-efficient AI and brain-inspired computing. However, their widespread ado…
A Combined Feature-Based Framework for Disguise and Spoofing Detection in Face Recognition Systems
Face recognition systems face two distinct, commonly-separated failure modes: spoofing, where an impostor presents a photograph or video of…
On-Device Multi-Species Malaria Detection with Uncertainty-Calibrated Slide-Level Aggregation
Malaria remains a leading cause of mortality in resource-limited settings, where expert microscopists are scarce. Automated diagnosis based…
CDGC-Net: 3D Medical Image Segmentation with Cooperative Dual-Scale Self-Attention and Grouped Channel Modeling
Accurate 3D medical image segmentation requires the integration of long-range anatomical context with fine boundary detail. Existing method…
Population-Scalable Multi-Agent World Modeling
World models have recently achieved impressive progress in visual prediction and interactive generation, but extending them to multi-agent…
Mitigating Gender Bias in English to Romanian Machine Translation
Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like E…
REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering
Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering. Howev…
RAG-Based Auto-Configuration for Industrial Fieldbus Devices
Industrial device commissioning requires engineers to manually extract hundreds of protocol-specific parameters from heterogeneous PDF manu…
Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach
Scientific named entity recognition (SciNER) plays a crucial role in information extraction and knowledge discovery from scientific texts.…
CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents
Zero-shot text-to-speech (TTS) now supports interactive assistants, personalized media, and accessibility tools. All TTS systems require fa…
LegoLM: Structured Weight Sharing for Large Language Models
We present \LegoLM{}, a structured weight-sharing compression framework for large language models grounded in a systematic study of why glo…
UniSpace: Unified Visual Representation and Scalable Multimodal Modeling
Semantic vision encoders have become a central visual interface for multimodal understanding and semantic conditioning in image generation.…
RippleKV: Cross-Layer KV Cache Allocation via Perturbation Propagation
Long-context LLM inference is bottlenecked by KV cache memory, yet distributing a limited cache budget across layers remains challenging. E…
Resolution Meets Reduction: Efficient Visual Context for 3D Radiology Report Generation
Vision-language models offer a promising path toward automating radiology report generation, but applying them to full 3D CT volumes poses…
LibraSpec: Dynamic Diffusion-Based Speculative Decoding via Marginal-Gain-Driven Optimization
Speculative decoding accelerates large language model inference by drafting multiple tokens for parallel verification, with efficiency crit…
Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection Pressure
Benchmarks for systems that are optimized against the evaluation signal measure something different from what they claim. We document this…
PAST: Privileged Adaptation from Complete Student Trajectories for On-Policy Self-Distillation
On-policy self-distillation (OPSD) uses a privileged teacher to supervise a reasoning model on prefixes sampled from its own rollouts. Yet…
TomaMMU: A Comprehensive Multimodal Understanding Benchmark for Tomato Leaf Diseases
To address this gap, we introduce TomaMMU, a large-scale Tomato leaf disease MultiModal Understanding dataset, alongside TomaBench, a bench…
Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs
The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially…
Eco-SoC: A Sustainable VLSI Architecture for Energy-Proportional Artificial Intelligence
In an era defined by escalating climate change and the pervasive deployment of edge intelligence, the environmental cost of semiconductor m…
Learning from Consensus and Disagreement: Unsupervised On-Policy Self-Distillation with Minority-Trajectory Contrast
On-policy self-distillation improves language-model reasoning by querying a teacher on states actually visited by the student. Recent metho…
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environme…
Hybrid Neural-Classical Correction for Frozen Time Series Foundation Models: A Comprehensive Ablation Study on High-Frequency Stock Prediction
Foundation models for time series forecasting demonstrate impressive zero-shot generalization but often underperform on specialized domains…
Deployable Per-Instance Multi-Layer Activation Steering for Large Language Models
Activation steering edits the behaviour of a frozen language model by adding a learned vector to its residual stream, and current practice…
Agentic Anomaly Detection with ORCA-Style Dynamic Inductive Bias Adaptation in Multimodal Wearable Time Series Data
Wireless Body Area Networks (WBANs) generate multivariate physiological time series that are highly nonstationary and must often be process…
DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference
Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly…
Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol
AI tools that help people judge online claims are usually evaluated while the tool is present. This paper asks a different question: after…
A New Approach to Characterising Optimisation Problems Using Programmatic Representation and Complexity Measures
Characterising optimisation problem instances is a fundamental part of understanding the behaviour and performance of different algorithms…
From Recovery to Drop-off: How Action Post-training Reduces a VLM's Late-Layer Depth Decodability
How much of a vision-language model's (VLM) spatial understanding remains after the action post-training process of building a vision-langu…
ToolVision: Learning When and How to Use Visual Tools with Capability-Aligned Supervision
Thinking with images allows a multimodal model to compensate for limited perception by invoking visual tools through code. Yet the prevaili…
Toward CT-Equivalent Image Quality in Low-Dose Radiotherapy Planning: Conditional Diffusion-Based CBCT-to-CT Synthesis and the Impact of CBCT Input Representation
During standard radiotherapy planning, repeated CT acquisitions are often required for patient registration, verification, and adaptive pla…
From Operational Design Domain to Action: A Systematic Behavioral Taxonomy for Autonomous Driving
Operational Design Domain (ODD) specifications describe where an automated driving system (ADS) is permitted to operate, but they do not pr…
Do AI Forecast Ensembles Sample the Correct Conditional Distribution?
Ensemble forecasting aims to sample the conditional distribution of outcomes; whether AI forecast ensembles do this correctly in a joint se…
Idea Search: Guiding Tree Search with Ideas to Explore Diverse Scientific Methods
Tree Search-based test-time scaling of LLMs is a powerful tool for automated scientific coding. However, pure Tree Search sometimes struggl…
Fourier Self-Supervision for Fine-Grained Generalized Category Discovery
Generalized Category Discovery aims to recognize known categories while identifying novel ones within unlabeled data. Existing methods, typ…
GALA: Graph-Augmented LLM Agents for Root Cause Analysis and Incident Response in Microservices
Microservice root cause analysis (RCA) requires correlating failures across heterogeneous telemetry within complex service dependency graph…
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetoric…
Detecting Clear Contact Lenses for Iris Recognition: A Two-Stage Mask-Guided Attention Approach
This work focuses on the impact and detection of clear contact lenses in the context of iris recognition. While the detection of cosmetic o…
How Far Do Foundation Models Transfer to Infant Signals? A Cross-Dataset Transfer Audit with a Unified Need Ontology
Public infant cry corpora are small, label-incompatible, and almost always evaluated one corpus at a time. We ask what this practice hides…
Guardian Crawler: Retrieval-First Knowledge Discovery with Bounded LLM Augmentation for Noisy Web Intelligence
Retrieving relevant evidence from noisy web data is challenging, particularly in sensitive domains containing incomplete reports, heterogen…
Multi-agent discovery of practical quantum LDPC codes
Quantum low-density parity-check (qLDPC) codes can encode multiple logical qubits using sparse parity checks, yet searching for useful fini…
DeepFreqMark: End-To-End Learnable Frequency-Domain Watermarking with Spherical Attack Simulation for Latent Diffusion Models
The proliferation of AI-generated images produced by Latent Diffusion Models (LDMs) has raised critical concerns regarding copyright infrin…
SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs
Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign L…
How People Evaluate AI-, Expert-, and Peer-Style Financial Advice
As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand h…
MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation
Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural orga…
Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization
Users of modern platforms repeatedly need summaries of recent dialogue, but the window rarely contains enough context to be interpreted on…
Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production
Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such…
Triple Expert Learning from Noisy Labels for Semi-Supervised Vision Foundation Model Adaptation
Semi-supervised adaptation of vision foundation models (VFMs) commonly freezes the pretrained backbone and updates lightweight modules such…
Two-Step MV-DeepONet: Probabilistic Operator Learning for Uncertainty Propagation Driven by Random Input Fields
Forward uncertainty propagation in complex physical systems can induce structured covariance across field-valued outputs. For a probabilist…
A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents
Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or f…
When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information
Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their re…
TLDChoiceNet: Quantitatively Choosing a Transfer Learning Dataset
Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer l…
The Announcement Carries the Cue: Markup, Boundaries, and the Notation of Pre-Training Corpora
How a document's arrangement is written down, its notation, is a training variable that no dataset card records. The field has established…
Visual Distortion Detection in UGC Images Using Large Multimodal Models
The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Exis…
Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments
LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measur…
MARA: Flow-Matching-Guided Multi-Agent Resource Allocation for Computational Resource Efficient Learning
Allocating limited computation among concurrent learning tasks is difficult when each task must reach a target loss before a deadline but i…
When Latents Forget Pixels: Restoring Fidelity in Diffusion Transformer Super-Resolution
Image super-resolution (SR) with large generative models has recently achieved remarkable perceptual quality, yet maintaining fidelity to t…
SpeedTuning: Speeding Up Policy Execution with Lightweight Reinforcement Learning
While learned robotic policies hold promise for advancing generalizable manipulation, their practical deployment is often hindered by subop…
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes lo…
Tabular Numeric Stretch Transformation
Tabular data presents unique challenges for deep learning due to its heterogeneous nature, where numeric features exhibit diverse distribut…
Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression
Visual token compression for vision--language models (VLMs) has largely relied on criteria such as attention, redundancy, and uncertainty t…
Rethinking Medical Landmark Localization with Prototype Learning-based Progressive Offset Correction
Accurate landmark localization in medical images is a fundamental step for quantitative clinical measurement and downstream analysis. Exist…
SiriusDeliver: Automating Data Warehouse Delivery at Tencent
Enterprise data warehouses (DWs) support business-critical analytics, but warehouse task delivery remains a complicated production process…
FedA2L: Adaptive layer-wise learning rate adjustment in decentralized federated learning
Decentralized intelligence systems with heterogeneous devices and limited coordination increasingly rely on decentralized federated learnin…
AkasicDB: Demonstrating Omni RAG with a Unified Vector-Graph-Relational DBMS
Recent Retrieval-Augmented Generation (RAG) systems increasingly combine vector retrieval with structured knowledge, such as Graph RAG and…
Beyond Solvability: Task Learnability as a Static Prior for LLM RL Post-Training
Reinforcement learning (RL) has become a central post-training paradigm for eliciting reasoning capabilities in large language models, yet…
FedTVD: Balancing Data Quality and Quantity for Robust Federated Learning
Federated Learning (FL) enables collaborative model training across distributed client devices while preserving data privacy. However, FL f…
Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM Inference
The key-value (KV) cache is the primary throughput optimization in modern large language model (LLM) inference, enabling prefix reuse acros…
RL-Native Distillation: Exploiting Scored Trajectories for Few-Step Image Generation
Efficient text-to-image generation requires both reinforcement-learning (RL)-based reward alignment and few-step distillation, yet these pr…
Privileged Solutions or Context-Induced Teacher Behavior? Dissecting On-Policy Self-Distillation
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solut…
Multimodal Federated Learning under Dual-Axis Modality Missingness
Multimodal federated learning (FL) supports collaborative modeling in privacy-sensitive health-sensing and medical settings, but realistic…
MoRSE: Task-Oriented Multi-Agent System with Mixture of Role-Subtask Experts
Large language model-based multi-agent systems have recently shown strong potential for complex, long-horizon tasks. However, existing meth…
SafeQL: Search-based Refinement for Safe and Efficient LLM-based Text-to-SQL
Large language models (LLMs) have advanced Text-to-SQL by enabling natural language interfaces to databases without task-specific fine-tuni…
Can Coding Agents Solve Repository-Level Issues with Rendered Code? An Exploratory Study of Visual Representations
Visual modality has recently been explored as a way to compress textual tokens, including rendering code as images for static code understa…
GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views
Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field…
SoftmaxGRPO: Learning to Reason using Softmax Advantage Group Estimation
Group-based reinforcement learning objectives such as GRPO can allocate learning signal poorly across prompt difficulty: under binary rewar…
Software Engineering for and with GUI Agent
GUI agents have advanced rapidly, producing a growing body of frameworks, benchmarks, and applications. However, this growth has outpaced t…
GLocFM: A Geometry-Aware Foundation Model for 3D Indoor Wireless Localization
Learning-based wireless localizers often fail to utilize geometric information about the propagation environment, limiting their ability to…
VeinCast: Physics-Guided Dynamic Field Graphs with Graph-Conditioned Fusion for Global Medium-Range Weather Forecasting
Global medium-range weather forecasting requires modeling structured yet state-dependent interactions among heterogeneous atmospheric field…
UniDFKD: A Unified Semantic Prior Framework for Architecture-Agnostic Data-Free Knowledge Distillation
Data-Free Knowledge Distillation (DFKD) transfers knowledge from a pretrained teacher model to a compact student model by synthesizing sema…
DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation
Audio-visual speech enhancement under real-world conditions remains challenging due to unreliable visual inputs and the lack of large-scale…
WorldSimProbe: Diagnosing Simulator Faithfulness in Action-Conditioned World Models for Embodied Manipulation
Action-conditioned world models (ACWMs) promise to provide embodied AI with scalable predictive simulators for planning, policy evaluation,…
RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction
Brain-to-audio reconstruction is limited by \emph{prior domination}: when a pretrained generator is conditioned on a weak neural signal, it…
Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute
Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matte…
Deep Learning based Detection of Fishing Vessels and Fishing Monitoring using Nightlight Images
The demand for maritime surveillance has given rise to the need for monitoring fishing vessel activities, particularly in addressing the ch…
FeedbackTrack: Visual-Cortex-Inspired Cross-Frame Feedback for Transformer Tracking
Visual object tracking requires effective temporal integration, yet most Transformer trackers still rely on predominantly feed-forward feat…
Imaginative Generative AI: Crossing the Entropy Wall into Worlds Beyond Imitation
Generative AI models are primarily designed to imitate the data distribution, an objective that neither corrects diversity lost by a learne…
Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax Law
We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article…
Monotonicity-Guided Bottom-Up Petri Net Discovery: The SPECpp Framework
Process discovery is one of the central challenges in process mining. Petri nets are particularly attractive because simple local construct…
Sign Language Recognition Using Original and Synthetic Depth Image Based Point Cloud Data Models
Research regarding the sign language recognition mostly relies on RGB images, whileas sign language datasets that provide depth images are…
LITEWAY: LIghtweight HAR via Temporal Efficient highWAY
Wearable human activity recognition (HAR) remains challenging due to the computational and energy constraints of deep learning models on re…
ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models
Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mec…
How Simple Can It Get? From Interpretable Equations to Readable Rules for Financial Decision Making
In regulated domains such as finance, a model that cannot be explained cannot be deployed, yet many interpretable classifiers defeat their…
WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training
On-policy distillation (OPD) aligns a student with a teacher on trajectories sampled from the student itself, reducing the train-test state…
Learning to Modulate, Not to Cycle: Soft Actor---Critic Recovers Inverter-Style Heat-Pump Control
On--off cycling is the main cause of compressor wear in residential heat pumps, yet reinforcement learning (RL) controllers for buildings t…
RecoverFly: A Failure-Aware Reinforcement Learning Post-Training Framework for Aerial Vision-Language Navigation
Unmanned aerial vehicle vision-language navigation (UAV-VLN) requires agents to translate visual observations and language instructions int…
MixFormer: Linear Transformer with Mixture of Memory Experts
State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transf…
ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents
Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We def…
Beyond Uniform Restoration: Empowering All-in-One Restoration with Pixel-Level Multimodal Guidance
All-in-one image restoration is a unified low-level vision task that aims to effectively recover high-quality images from inputs degraded b…
Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning
Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often…
Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts
Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to gene…
STAIR: Effective Incident Response Using an End-to-End Agentic Planning Framework
Incident response planning is critical for restoring compromised software systems after cyberattacks. Common practice relies on expert-driv…
RangeFactory: Scalable Construction of Multi-Hop Cyber Ranges
Real-world cyberattacks often require sustained progress across multiple hosts and network segments, making multi-hop cyber ranges essentia…
Carnot: Interpretable, Interactive, and Optimized Execution of Deep Research Queries
Enterprises increasingly seek to query data lakes using natural language via AI-driven tools like semantic operators or deep research agent…
TCS-BENCH: Benchmarking State-of-the-Art Generative AI Theoretical Computer Science Research Ability
We introduce TCS-Bench, a benchmark for evaluating Large Language Models (LLMs) on research-level Theoretical Computer Science (TCS) proof…
Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs
Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outpu…
ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place dema…
From Semantic Grounding to Decision Optimization: A Unified Framework for Long-Horizon UAV Vision-Language Navigation
UAV vision-language navigation (UAV-VLN) focuses on enabling an aerial agent to follow natural-language instructions in open 3D environment…
Distributed Optimization with Streaming Data: A Temporal Weighting Perspective
Optimization theory is a widely used tool for intelligent decision-making. While classical optimization deals with fixed, time-invariant ob…
MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection
Recent advances in speech synthesis and audio generation have made high-fidelity acoustic forgery low-cost and difficult to attribute, enab…
Illusion or Integrity? Geometrical Consistency Metric for AIGC Video Quality Evaluation
Recently, AI-driven video generation has attracted considerable attention. This surge increases the demand for reliable video quality asses…
LEED: Local Embedding Evolution Distance for over-smoothing estimation and virtual node selection in GNN
Graph Neural Networks (GNNs) suffer from two fundamental limitations: over-smoothing, where node representations become indistinguishable w…
TSPORec: Token Selection via Preference Optimization for LLM-Based Sequential Recommendation
Large Language Models (LLMs) have emerged as powerful tools for improving recommendation systems. The effectiveness of LLMs arises from the…
Structure-Enhanced Features and Quality-Aware Dynamic Anchor Scoring for Robust Lane Detection
Lane detection requires recovering thin, elongated, and frequently occluded lane structures under challenging driving conditions. While anc…
Measuring the Wrong Thing: Internal Harmfulness Scores Anti-Rank Successful Jailbreaks
Internal safety scores judge a prompt before any text is generated, and they are validated by how well they separate harmful prompts from b…
NeuroRefiner: Morphology-Aware Multi-Agent Refinement for 3D Fluorescence Microscopy Neuron Segmentation
Accurate 3D neuron segmentation in fluorescence microscopy is critical for neuroscience. However, the sparse and elongated morphology of ne…
DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation
Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practica…
Predictive safety filter enhanced curriculum learning control for efficient vehicle dynamics controller
Recent advances in learning-based control have enabled impressive achievements in solving complex control problems in various domains. Howe…
Confusion-Geometry Rebalancing for Long-Tailed Adversarial Training
Adversarial training under long tailed distributions suffers from a dual imbalance: the class imbalance skews the training objective toward…
Evaluating Generative Time-Series Models on Data with Point Masses
Many of the series that generative time-series models are benchmarked on place a large probability mass on a single value --- it does not r…
How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans
Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as so…
ColluSkill: Adversarial Cross-Skill Composition for Evading Agent Skill Scanners
Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners,…
Rethinking Factor Sharing in Federated LoRA: A Rank-Aware Adaptive Approach
Low-rank adaptation (LoRA) represents large language model (LLM) updates with two compact matrix factors, i.e., $A$ and $B$, providing an e…
SR-OPSD: Self-Referenced On-Policy Self-Distillation
On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on trajectories generated by the policy to be optim…
Defining Decentralization: An Ontological Perspective
Decentralization as a concept in computer science has existed for over half a century. Despite its fundamental role across domains such as…
MoNo: Multiscale Optimal Transport Neural Operator for Solving PDEs on General Geometries
Transformer-based neural operators have achieved substantial progress in solving Partial Differential Equations (PDEs) by projecting spatia…
Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the un…
KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs
Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, pa…
Modern Backbones Improve Multi-task DETR for Mammography Classification and Lesion Localization
Joint exam-level prediction and candidate-region localization may improve the usefulness of AI support in mammography. We study this settin…
Parameter Exploration for RLVR via Variational Learning
Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also…
MedPixel: A Unified Pixel-Language Model for Medical Reasoning and Segmentation
Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding. Yet medi…
Distill Skills into Weights, Not Prompts: Abstract Skills as Privileged Signals for On-Policy Self-Distillation
Reinforcement learning with verifiable rewards yields no group-relative signal when rollout groups are uniformly correct or uniformly wrong…
Multi-Agent AI Safety as an Institutional Design Problem
AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources…
Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy
Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely…
Stealing Reasoning Traces from Proprietary LLM APIs
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual prope…
Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific dom…
Energy-Structured Latent World Models with Neural Time Fields for Physically Constistent Open-World Motion Planning
Physically consistent motion planning remains a fundamental challenge in embodied AI, as generated trajectories must strictly conform to re…
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning
We introduce BDH-CQ, a reasoning model that combines in-context learning with recurrent latent reasoning. Inputs presented at inference tim…
Fusion Training for Mathematical Generalization in Large Language Models
Thinking Mode Fusion (TMF) enables large language models to support both concise responses and long-form reasoning by unifying a non-thinki…
From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch
Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the…
Multimodal Model Diffing for Feature Discovery and Control
Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain d…
Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges)…
Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory
The emergence of Large Language Models (LLMs) and advancements in Artificial Intelligence (AI) offer an opportunity for computational socia…
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial…
How Much Backtracking is Enough? Exploring the Interplay of SFT and RL in Enhancing LLM Reasoning
Recent advancements in large language models (LLMs) suggest that reinforcement learning (RL) effectively internalizes search strategies, yi…
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
The integration of brain-computer interfaces (BCIs), in particular electroencephalography (EEG), with artificial intelligence (AI) has show…
ACEvo: Adversarial Co-Evolution of Problem Distributions and Solvers for Combinatorial Optimization
Large language models (LLMs) are increasingly used to synthesize heuristic programs, yet most existing pipelines optimize solvers against f…
Reflex First, Reflect Later: Latency-Aware Embodied LLM Agents for Dynamic Response
Large language models (LLMs) have substantially improved the planning capabilities of embodied agents, enabling their deployment in dynamic…
Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents
The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serial…
Probabilistic Circuits for Knowledge Graph Completion with Reduced Rule Sets
Rule-based methods for knowledge graph completion provide explainable results, but often require tens of thousands of rules to achieve comp…
From Mimicry to True Intelligence (TI) -- A New Paradigm for Artificial General Intelligence
The debate around Artificial General Intelligence (AGI) remains open due to two fundamentally different goals: replicating human-level perf…
ToolUniverse: An open platform for democratizing AI scientists
AI scientists are emerging computational systems that serve as collaborative partners in discovery. These systems remain difficult to build…
TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation
Current training paradigms, optimized for long-horizon reasoning trace execution, have made Large Language Models (LLMs) excel at pattern m…
The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation
The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of ind…
Intelligence Foundation Model: A New Perspective to Approach Artificial General Intelligence
We propose a new perspective for approaching artificial general intelligence (AGI) through an intelligence foundation model (IFM). Unlike e…
The Belief-Desire-Intention Ontology for modelling mental reality and agency
The Belief-Desire-Intention (BDI) model is a cornerstone for representing rational agency in artificial intelligence and cognitive sciences…
M$^3$Prune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation
Recent advancements in multi-modal retrieval-augmented generation (mRAG), which enhance multi-modal large language models (MLLMs) with exte…
Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE…
Med-CRAFT: An Information System for Explainable and Configurable Construction of Multimodal Medical QA Datasets
Data-intensive artificial intelligence applications increasingly rely on large-scale, high-quality, explainable, and reproducible datasets,…
Agentic AI for Clustering, Relationship Discovery, and Semantic Trading in Prediction Markets
Prediction markets allow users to trade on outcomes of real-world events, but are prone to fragmentation with overlapping questions, implic…
Neuronal Attention Circuit (NAC) for Representation Learning
Attention improves representation learning over RNNs, but its discrete nature limits continuous-time (CT) modeling. We introduce Neuronal A…
Multi-Granular Node Pruning for Causal Circuit Discovery
Circuit discovery aims to identify minimal subnetworks that are responsible for specific behaviors in large language models (LLMs). Existin…
SPRInG: Continual LLM Personalization via Selective Parametric Adaptation and Retrieval-Interpolated Generation
Personalizing Large Language Models typically relies on static retrieval or one-time adaptation, assuming user preferences remain invariant…
Position: Certifiable State Integrity Should Be Built from Local Validity, Not Global Scale
Breakthroughs in language and vision have motivated increasingly general foundation models for time series and physical dynamics, where evi…
AutoRefine: Compiling Trajectories into Validated Typed Agent Artifacts
Large language model agents repeatedly encounter related tasks, yet systems that learn from trajectories commit every lesson to one predefi…
El Agente Gr\'afico: A Semantic Execution Runtime for Scientific Agents
Large language models (LLMs) can plan scientific workflows and generate code, but these capabilities do not specify how scientific state is…
Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search
Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instanc…
A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling
Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but…
MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models
Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When s…
SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents
Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientifi…
Self-Routing: Parameter-Free Expert Routing from Hidden States
Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a lea…
AIVV: Neuro-Symbolic LLM Agent-Integrated Verification and Validation for Trustworthy Autonomous Systems
Deep learning models excel at detecting anomaly patterns in normal data. However, they do not provide a direct solution for anomaly classif…
A Statistical Framework for Auditing Behavioral Dependence and Induced Bias in LLM Judges
The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent? Sh…
SkillClaw: Let Skills Evolve Collectively with Agentic Evolver
Large language model (LLM) agents such as OpenClaw rely on reusable skills to perform complex tasks, yet these skills remain largely static…
DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math?
Deep research agents increasingly interleave web browsing with multi-step computation, yet existing benchmarks evaluate these capabilities…
FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks
Recent studies demonstrate that tool-calling capability enables large language models (LLMs) to interact with external environments for lon…
Collaborative Multi-Agent Scripts Generation for Enhancing Imperfect-Information Reasoning in Murder Mystery Games
Vision-language models (VLMs) have shown impressive capabilities in perceptual tasks, yet they degrade in complex multi-hop reasoning under…
RankGuide: Tensor-Rank-Guided Routing and Steering for Efficient Reasoning
Large reasoning models (LRMs) enhance problem-solving capabilities by generating explicit multi-step chains of thought (CoT) reasoning; how…
Time-Series Forecasting in Safety-Critical Environments: An Open-Source Package for EU-AI-Act-Compliant Development / Zeitreihenprognose in sicherheitskritischen Umgebungen: Ein Open-Source-Paket f\"ur die KI-VO-konforme Entwicklung
With spotforecast2-safe we present an integrated Compliance-by-Design approach to Python-based point forecasting of time series in safety-c…
ZenBrain: A Neuroscience-Inspired 7-Layer Memory Architecture for Autonomous AI Systems
ZenBrain is a seven-layer, neuroscience-derived memory architecture for LLM agents that unifies fifteen mechanisms - from Two-Factor synapt…
FitText: Evolving Agent Tool Ecologies via Memetic Retrieval
Efficient reasoning is not only a matter of shortening an answer trace; for tool-using agents, it also depends on whether the agent is reas…
C2L-Net: A Data-Driven Model for State-of-Charge Estimation of Lithium-Ion Batteries During Discharge
Accurate state-of-charge (SOC) estimation is critical for the safe and efficient operation of lithium-ion batteries in battery management s…
AgentPSO: Evolving Agent Reasoning Skill via Multi-agent Particle Swarm Optimization
Multi-agent reasoning has shown promise for improving the problem-solving ability of large language models by allowing multiple agents to e…
Understanding and Mitigating Premature Confidence for Better LLM Reasoning
Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from ad…
PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
Large language models (LLMs) have shown strong performance across diverse financial tasks, yet portfolio management (PM) remains poorly ben…
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
Extended chain-of-thought reasoning can degrade performance on deterministic state-tracking tasks, not solely because of preference biases…
Towards Understanding Modality Interaction in Multimodal Language Models via Partial Information Decomposition
Understanding how multimodal large language models use different modalities is important for reliable reasoning. We employ Partial Informat…
Human agency in initial human-AI proof formalization workflows
For centuries, human mathematicians have written proofs to substantiate their mathematical arguments; yet, the ability to automatically ver…
SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research
Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bou…
READER: Dynamic LLM Provenance from Query-Varying Interactions
Existing black-box LLM provenance methods achieve comparability by querying every candidate model with the same diagnostic prompts. In depl…
Unbiased Canonical Set-Valued Oracles Via Lattice Theory
An oracle that tells you the probability of some future event can change that very probability because you act on the answer. We argue that…
Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation
Large reasoning models (LRMs) spend more reasoning tokens on problems that take humans longer, suggesting sensitivity to a similar structur…
NormAct: Benchmarking Embodied Agents' Proactive Compliance with Unspoken Social Norms
Embodied agents driven by multimodal large language models (MLLMs) can often complete everyday tasks from visual observations, but goal ach…
ContextSniper: AntTrail's Token-Efficient Code Memory for Repository-Level Program Repair
Large language model agents can repair real repository issues, but they often spend large context budgets on whole-file reads, broad search…
LLM-as-a-Tutor: Policy-Aware Prompt Adaptation for Non-Verifiable RL
Reinforcement learning (RL) for non-verifiable instruction following increasingly relies on LLM judges with prompt-specific rubrics as rewa…
Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale
LLM-powered data agents are playing an increasingly important role in data-driven decision making. However, existing data agents struggle t…
Experience Memory Graph: One-Shot Error Correction for Agents
Large Language Model (LLM) agents have shown remarkable capabilities in autonomous decision-making by generating sequential trajectories of…
Concept-Guided Spatial Regularization for World Models in Atari Pong
World models are usually evaluated as components of model-based reinforcement learning (MBRL) systems, leaving their standalone reliability…
Exact Network Surgery: Functional Invariance and Gradient Plasticity in Reactive Computational Graphs
Function-preserving network growth techniques such as Net2Net and progressive stacking expand a model's capacity without destroying its lea…
ProbSPARQL: Querying Knowledge Graphs with Multi-dimensional, Uncertain Numeric Data
The SFB 1574 Circular Factory is building a shared knowledge graph infrastructure for integrating data about returned products. A central c…
SoftReason: A Fully Differentiable Neuro-Soft-Symbolic Deductive Reasoning Architecture over High-Dimensional Perceptual Data
In many reasoning problems, the premises are not observed as discrete symbols, but must be inferred from high-dimensional inputs. Further,…
From Errors to Rules: Iterative Prompt Optimization for Text Classification
Prompt optimization for text classification spans diverse approaches, from demonstration selection to exploration-based search to error-dri…
Physical AI Governance: From Theory to Practice Across Life Cycle
With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive,…
RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning
Rare diseases represent one of the most challenging settings for clinical decision-making, where heterogeneous presentations, sparse eviden…
When benchmark inferences do not compose: Projectibility in AI evaluation
An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidenc…
SemPIC: Learning Semantic Position-Independent KV Caches
Long-context retrieval and agentic workloads repeatedly reuse the same documents under changing instructions, histories, and document order…
ConMem: Contribution-Aware Memory for Long-Horizon Manufacturing Inspection Logs
Long-horizon steel-equipment inspection requires reasoning over heterogeneous records accumulated across repeated inspection cycles. Existi…
A foundation model of numerical intelligence with cross-disciplinary generalization
Intelligence is commonly understood as the ability to acquire and apply knowledge, adapt to unfamiliar situations and solve new problems. L…
InfoOps Bench: A live information operations safety benchmark
In this paper we present an active, constantly updated AI benchmark which measures the integrity of frontier language models against being…
Learning to Coordinate Symbolic Tools: LLM Agents for Verified Sum-of-Squares Certificates
Tool calling allows large language models (LLMs) to invoke external computation during problem solving, a useful capability in various fiel…
SymboUQ: Symbolic Uncertainty Quantification for Spatial Reasoning in LLMs
Although large language models (LLMs) can produce fluent spatial reasoning traces, their intermediate relations may fail to support the fin…
Why Does the Future Branch? Identifiable Closure Tests for Stochastic Physical World Models
A calibrated stochastic world model can reveal how uncertain a future is without revealing why it branches. The same conditional future law…
Evolutionary Curriculum Learning Improves Biological Sequence Modeling
Variational autoencoders (VAEs) trained on multiple sequence alignments (MSAs) have emerged as powerful generative models for biological se…
The Scaling Paradox in Human-AI Collaboration
The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems sc…
Agentic Stage-One Stellarator Optimization: Autonomous Multi-Objective Search for Finite-Beta Equilibria
Stage-one stellarator design searches a high-dimensional family of three-dimensional plasma boundaries and fixed-boundary MHD equilibria fo…
Emergence Invariance: From Symbolized Thought to Structural Control
Language-first intelligence is constrained by which distinctions enter its symbolic record, which mappings its language--interpreter--envir…
State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking
Transformer-based architectures have dominated sequence modeling, largely due to the expressive power of attention mechanisms. However, for…
When Efficiency Becomes Fragility: Exploiting Dynamic Routing Vulnerabilities in Adaptive UAV Tracking
Resource constraints on UAV platforms have driven a paradigm shift in aerial tracking, from pursuing performance toward balancing accuracy…
The Transformer Revolution, Part 1: Dynamic Processing through Output-Weight Interconnections
This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models…
Short-term load forecasting under EU-AI Act Requirements in Safety-Critical Environments: Results from a 41-day live challenge on the aggregated German transmission-grid load
Short-term load forecasting (STLF) play a vital role in the electric power industry. It serves infrastructure that European and German law…
Argus: A General-Purpose Agentic Reasoning Runtime for Long-Horizon Tasks
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements…
Small Foundation Models of Human Cognition and Behaviour
Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires,…
Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows
Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how the…
ViSR-KGC: Visual Subgraph Reasoning with Vision-Language Models for Multimodal Knowledge Graph Completion
Knowledge graph completion (KGC) aims to infer missing entities or relations from incomplete graph structures, and has evolved into multimo…
Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts
Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argu…
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization
In this paper, we focus on providing convergence guarantees for stochastic subgradient methods in minimizing nonsmooth nonconvex functions.…
Explainable Machine Learning-Based Security and Privacy Protection Framework for Internet of Medical Things Systems
The Internet of Medical Things transcends traditional medical boundaries, enabling a transition from reactive treatment to proactive preven…
Ethical Framework for Responsible Foundational Models in Medical Imaging
The emergence of foundational models represents a paradigm shift in medical imaging, offering extraordinary capabilities in disease detecti…
Transformer Explainer: Learning LLM Transformers with Interactive Visual Explanation and Experimentation
The Transformer architecture underpins modern large language models powering state-of-the-art text generation and AI applications. However,…
See Me, Believe Me: Causality, Intersectionality, and Interventions Improving the Appearance of Patients
In the context of medical records, patients often experience testimonial injustice, where the textual account undermines the validity of th…
A Rigorous Turing Test: a Foundation for Evaluating Artificial General Intelligence
Several studies claim that large language models have passed the Turing Test and hence can "think", yet none follow Turing's original instr…
LF${}^{2}$AR: Accounting for Layerwise Dynamics to Improve Multimodal Adaptation of Language Models
Text-pretrained language models (LMs) encode rich world knowledge, but adapting them to process and generate perceptual modalities such as…
REMAC: Self-Reflective and Self-Evolving Multi-Agent Collaboration for Long-Horizon Robot Manipulation
Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that requi…
When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output
Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models…
An Expectation-Maximization Perspective on Reinforcement Learning for LLM Reasoning
Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrate…
TreeHop: Efficient Embedding-Level Query Rewriter
Retrieval-augmented generation (RAG) systems face significant challenges in multi-hop question answering (MHQA), where complex queries requ…
Optimal Transport for Machine Learners
Modern machine learning repeatedly manipulates probability measures: empirical datasets, generated samples, latent distributions, class-con…
X2C: A Dataset Featuring Nuanced Facial Expressions for Realistic Humanoid Imitation
Fine-grained facial expression transfer from humans to humanoid agents presents a unique pattern recognition challenge due to the significa…
SuperCoder: Assembly Program Superoptimization with Large Language Models
Superoptimization is the task of transforming a program into a faster one, and ideally the very fastest possible one, while preserving its…
SAKE: Structured Agentic Knowledge Extrapolation for Complex LLM Reasoning via Reinforcement Learning
Knowledge extrapolation is the process of inferring novel information by combining and extending existing knowledge that is explicitly avai…
Transformer-Based Neural Quantum Digital Twins for Many-Body Spectral Reconstruction and Adaptive Quantum-Annealing Schedule Design
We introduce Transformer-based Neural Quantum Digital Twins (Tx-NQDTs) to reconstruct the low-energy spectral evolution of many-body quantu…
FoMoH: A clinically meaningful foundation model evaluation for structured electronic health records
Foundation models (FMs) promise to address core limitations of traditional supervised machine learning: (i) reliance on large amounts of la…
The Cell Must Go On: Agar.io for Continual Reinforcement Learning
Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then…
HyperFake: Hyperspectral Reconstruction and Attention-Guided Analysis for Advanced Deepfake Detection
Deepfakes pose a significant threat to digital media security, with current detection methods struggling to generalize across different man…
From Alignment to Synthesis Contrastive Volumetric Grounding for Text-to-CT Generation
Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-langua…
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transpare…
Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence
In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespre…
Transformer Circuits Can Realize Clustering Algorithms
Although transformers are most commonly optimized as statistical sequence models, it is unclear to what extent they can implement and learn…
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These…
Dynamic gain neuromodulation attenuates the stability gap under joint training
Recent work in continual learning has highlighted the stability gap -- a temporary performance drop on previously learned tasks when new on…
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within…
Enhancing Knowledge Tracing through Leakage-Free and Recency-Aware Embeddings
Knowledge Tracing (KT) aims to predict a student's future performance based on their sequence of interactions with learning content. Many K…
Deep Residual Echo State Networks: exploring residual orthogonal connections in untrained Recurrent Neural Networks
Echo State Networks (ESNs) are a particular type of untrained Recurrent Neural Networks (RNNs) within the Reservoir Computing (RC) framewor…
NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models
In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. Ho…
ATLASFusion: Aggregation Tracking with Location-Aware Sparse Fusion for Robust Spatio-Temporal Multi-View Pedestrian Tracking
For multimedia spatial intelligence through time, multi-view multi-object tracking (MVMOT) suffers from persistent challenges in maintainin…
Quokka: Accelerating Program Verification with LLMs via Invariant Synthesis
Program verification relies on loop invariants, yet automatically discovering strong invariants remains a long-standing challenge. We inves…
Topographic Constraints Shape Brain-Like Component Structure in Auditory Models
If topography is a fundamental feature of the brain, it should influence both how neurons are arranged in space (i.e. explain brain maps) a…
Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents
LLM agents require personal information for personalization in order to effectively act on users' behalf, but this raises privacy concerns…
Deep Generative Model for Human Mobility Behavior
Understanding and modeling human mobility is central to challenges in transport planning, sustainable urban design, and public health. Desp…
OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
Large language models (LLMs) with extended context windows enable powerful applications but impose significant memory overhead, as caching…
SUM-AgriVLN: Spatial Understanding Memory for Agricultural Vision-and-Language Navigation
Agricultural robots are emerging as powerful assistants across a wide range of agricultural tasks, nevertheless, they are still heavily rel…
Self-Attention to Operator Learning-based 3D-IC Thermal Simulation
Thermal management in 3D ICs is increasingly challenging due to higher power densities. Traditional PDE-solving-based methods, while accura…
NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning
Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation…
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended promp…
QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture
The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent f…
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in se…
Reasoning about Intent for Ambiguous Requests
Large language models often respond to ambiguous requests by implicitly committing to one interpretation, frustrating users and creating sa…
iLTM: Integrated Large Tabular Model
Tabular data underpins decisions across science, industry, and public services. Despite rapid progress, advances in deep learning have not…
VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate v…
Towards Realistic Guarantees: A Probabilistic Certificate for SmoothLLM
The SmoothLLM defense provides a certification guarantee against jailbreaking attacks, but it relies on a strict "k-unstable" assumption th…
Automating Deception: Scalable Multi-Turn LLM Jailbreaks
Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request pave…
Length-MAX Tokenizer for Language Models
We introduce a new tokenizer for language models that minimizes the average tokens per character, thereby reducing the number of tokens nee…
Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics
In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily ta…
Multilingual Agent-Based World Modeling for Social Science
Multi-agent role-playing has recently shown promise for studying social behavior with language agents, but existing simulations are mostly…
The Theory of Strategic Evolution: Games with Endogenous Players and the Seven Laws of Strategic Replicators
Von Neumann founded both game theory and the theory of self-reproducing automata, but the two programs never merged. Rational players do no…
Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models
Retrieval-augmented generation (RAG) relies on retrieved context to guide large language models (LLM), yet treats the retrieval as a heuris…
Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited para…
TGIF: Text-Guided Layer Fusion Mitigates Hallucination in Multimodal LLMs
Multimodal large language models (MLLMs) typically rely on a single late-layer feature from a frozen vision encoder, leaving the encoder's…
IndexTTS 2.5 Technical Report
In prior work, we introduced IndexTTS 2, a zero-shot neural text-to-speech foundation model comprising two core components: a transformer-b…
ReMIND: Orchestrating Modular Large Language Models for Controllable Serendipity A REM-Inspired System Design for Emergent Creative Ideation
Large language models (LLMs) are increasingly used not only for problem solving but also for creative ideation; however, generating ideas t…
Layerwise goal-oriented adaptivity for neural ODEs: an optimal control perspective
In this work, we propose a novel layerwise adaptive construction method for neural network architectures. Our approach is based on a goal--…
Expert-Guided Multimodal Fusion for Unified Emotion and Sentiment Analysis
Multimodal emotion understanding requires the integration of heterogeneous data sources, including text, audio, and visual modalities, whil…
LAUDE: LLM-Assisted Unit Test Generation and Debugging of Hardware DEsigns
Unit tests are critical in the hardware design lifecycle to ensure that component design modules are functionally correct and conform to th…
RAG-3DSG: Enhancing 3D Scene Graphs with Re-Shot Guided Retrieval-Augmented Generation
Open-vocabulary 3D Scene Graph (3DSG) can enhance various downstream tasks in robotics by leveraging structured semantic representations, y…
Communication-efficient distributed hazard difference estimation for heterogeneous multi-site survival data
Multi-site collaboration can power survival models that no single hospital could fit alone, but privacy rules and protected computing envir…
Hybrid Mamba-Attention Neural Architecture for Channel Estimation
This paper proposes a hybrid Mamba-attention neural architecture to achieve improved channel estimation for orthogonal frequency-division m…
SNR-Edit: Structure-Aware Noise Rectification for Inversion-Free Flow-Based Editing
Inversion-free image editing using flow-based generative models challenges the prevailing inversion-based pipelines. However, existing appr…
Temporal Sepsis Modeling: a Relational and Explainable-by-Design Framework
Sepsis remains one of the most complex and heterogeneous syndromes in intensive care. While deep learning models achieve competitive perfor…
Shattered Compositionality: Counterintuitive Learning Dynamics of Transformers for Arithmetic
Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts. While recent mecha…
Universal One-third Time Scaling in Learning Peaked Distributions
Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origi…
On the Infinite Width and Depth Limits of Predictive Coding Networks
Predictive coding (PC) is a biologically plausible alternative to standard backpropagation (BP) that minimises an energy function with resp…
SMAC: Score-Matched Actor-Critics for Robust Offline-to-Online Transfer
Modern offline Reinforcement Learning (RL) methods find performant actor-critics, however, fine-tuning these actor-critics online with valu…
From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
Large reasoning models (LRMs) produce reasoning traces (RTs) that often contain sensitive information. These leaky thoughts are difficult t…
Attn-QAT: 4-Bit Attention With Quantization-Aware Training
Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the…
Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks
Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be re…
LLMs Remember First, Forget Last: Dual-Process Interference in Large Language Models
Large language models can process millions of tokens, yet how they handle conflicting information within context remains poorly understood.…
PolypSteer: Counterfactual Endoscopic Synthesis via Training-Free Activation Steering
Generative diffusion models are increasingly used for medical imaging data augmentation, but text prompting cannot produce causal training…
Adversarial Latent-State Training for Robust Policies in Partially Observable Domains
Robustness under latent distribution shift remains challenging in partially observable reinforcement learning. We formalize a focused setti…
Efficient Cross-View Localization in 6G Space-Air-Ground Integrated Network
Recently, visual localization has become an important supplement to improve localization reliability, and cross-view approaches can greatly…
Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages
Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge…
Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification
Large language models (LLMs) can generate plausible code but offer limited guarantees of correctness. Formally verifying that implementatio…
SimulCost: A Cost-Aware Benchmark and Toolkit for Automating Physics Simulations with LLMs
Evaluating LLM agents for scientific tasks has focused on token costs while ignoring tool-use costs like simulation time and experimental r…
ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention
Large Reasoning Models (LRMs) often reach a correct solution before their long Chain-of-Thought trace ends, yet continue with redundant ver…
SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection
While large language models (LLMs) are pretrained on massive amounts of data, their knowledge coverage remains incomplete in specialized, d…
A Sobering Look at Tabular Data Generation via Probabilistic Circuits
Tabular data is more challenging to generate than text and images, due to its heterogeneous features and much lower sample sizes. On this t…
Explaining, Verifying, and Aligning Semantic Hierarchies in Vision-Language Model Embeddings
Vision-language model (VLM) encoders such as CLIP enable strong retrieval and zero-shot classification in a shared image-text embedding spa…
Critic-Free Deep Reinforcement Learning for Maritime Coverage Path Planning on Irregular Hexagonal Grids
Maritime surveillance missions, such as search and rescue and environmental monitoring, rely on the efficient allocation of sensing assets…
To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval
Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-inte…
Revision or Re-Solving? Decomposing Second-Pass Gains in Multi-LLM Pipelines
Multi-LLM revision pipelines, in which a second model reviews and improves a draft produced by a first, are widely assumed to derive their…
Goose: Anisotropic Speculation Trees for Training-Free Speculative Decoding
Speculative decoding accelerates large language model inference by drafting multiple candidate tokens and verifying them in a single forwar…
CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge
Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, an…
Large Language Models Align with the Human Brain during Creative Thinking
Creative thinking is a fundamental aspect of human cognition, and divergent thinking-the capacity to generate novel and varied ideas-is wid…
Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders
We present the first systematic study of Sparse Autoencoders (SAEs) on video representations. Standard SAEs decompose video into interpreta…
Not All Turns Are Equally Hard: Adaptive Thinking Budgets For Efficient Multi-Turn Reasoning in Agents
As LLM reasoning performance plateaus, improving inference-time compute efficiency is crucial to mitigate overthinking and long thinking tr…
SALLIE: Generation-Free Hidden-State Detection of Jailbreaks and Prompt Injections Across Text and Vision
Large Language Models (LLMs) and Vision-Language Models (VLMs) are vulnerable to jailbreaks and prompt injections delivered through text or…
Continual Visual Anomaly Detection on the Edge: Benchmark and Efficient Solutions
Visual Anomaly Detection (VAD) is a critical task for many applications including industrial inspection and healthcare. While VAD has been…
Multi-objective Evolutionary Merging Enables Efficient Reasoning Models
Reasoning models achieve strong performance on complex problems by leveraging long chains of thought, but this deliberate reasoning incurs…
TraceSafe: A Systematic Assessment of LLM Guardrails on Multi-Step Tool-Calling Trajectories
As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final out…
In-context superposition: human-like working memory interference in large language models
Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments and changing goals. This…
Symmetry Reveals Layerwise Dynamics: How Transformers Perform In-Context Classification
Transformers can perform in-context classification from a few labeled examples, yet the inference-time algorithm remains opaque. We study m…
Fairness is Not Flat: Geometric Phase Transitions Against Shortcut Learning
Deep Neural Networks are highly susceptible to shortcut learning, frequently memorizing low-dimensional spurious correlations instead of un…
The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models
Multimodal language models systematically underperform on visual perception tasks, yet the structure underlying this failure remains poorly…
Controllable Video Object Insertion via Multi-View Priors
Video object insertion places a user-specified object in an existing dynamic scene. Existing methods typically condition generation on text…
Switching Theory for Q-Learning
Q-learning is a fundamental algorithmic primitive in reinforcement learning. This paper develops a new framework for analyzing constant ste…
Hybrid Policy Distillation for LLMs
Knowledge distillation (KD) is a powerful paradigm for compressing large language models (LLMs), whose effectiveness depends on intertwined…
Model Predictive Control of Hybrid Dynamical Systems
The problem of controlling hybrid dynamical systems using model predictive control (MPC) is formulated and sufficient conditions for asympt…
From Local to Cluster: A Unified Framework for Causal Discovery with Latent Variables
Latent variables pose a fundamental obstacle to both causal discovery and inference. Local approaches exploiting direct neighborhood relati…
UGAF-ITS: A Standards Harmonization Framework and Validation Tool for Multi-Framework AI Governance in Distributed Intelligent Transportation Systems
Organizations deploying AI-enabled Intelligent Transportation Systems face fragmented governance: ISO/IEC~42001 demands a certifiable manag…
Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards
The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-mak…
Rethinking KV Cache Eviction via a Unified Information-Theoretic Objective
Key-Value (KV) caching is essential for large language model inference, yet its memory overhead poses a critical bottleneck for long-contex…
Culturally Situated AI Safety for Youth: Saudi Arabian Perspectives of Youth, Parents and Teachers
Generative AI tools are widely used by youth and have introduced new privacy and safety challenges. While prior research has explored youth…
Path-Lock Expert: Separating Reasoning Mode in Hybrid Thinking via Architecture-Level Separation
Hybrid-thinking language models expose explicit /think and /no_think modes, but current designs do not separate them cleanly. Even in /no_t…
The Safety-Aware Denoiser for Text Diffusion Models
Recent work on text diffusion models offers a promising alternative to autoregressive generation, but controlling their safety remains unde…
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
LLM fairness should be evaluated through in-situ behavioral pattern rather than standardized-test Q&A benchmarks. We show that the standard…
Not Just RLHF: Why Alignment Alone Won't Fix Multi-Agent Sycophancy
LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerab…
HEART: Exploiting Head Heterogeneity in Sparse Attention for Video Diffusion
Sparse attention accelerates video diffusion by allowing each attention head to focus on only a small subset of interactions. Existing meth…
DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation
Distillation enables compact Vision-Language Models (VLMs) to obtain strong reasoning capabilities, yet the prompts driving this process ar…
Post-Deployment Accountability in AI Governance: A Cross-Regulatory Empirical Analysis of AI Incidents
Post-deployment accountability has become central to AI governance, yet little empirical evidence shows whether monitoring, incident report…
Toward Measuring AI's Effects on Skill Formation: The Stock-Formation Gap
Large-scale AI deployment data and controlled learning experiments characterize different consequences of the same technology. Deployment t…
Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control
Large language models increasingly operate as autonomous agents that select and invoke tools from large registries. We identify a critical…
Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance
Accurate spatiotemporal pattern analysis is critical in fields such as urban traffic, meteorology, and public health monitoring. However, e…
How to Build Marcus's Algebraic Mind: Algebro-Deterministic Substrate over Galois Fields
In The Algebraic Mind (2001), Marcus held that any adequate cognitive architecture needs operations over variables, recursively structured…
Don't Retrain, Just Reuse: Recovering Dual-Target Molecules from Single-Target Diffusion Models
Designing a single molecule that modulates two targets is a promising strategy for polypharmacology, but it remains substantially harder th…
Beyond Questions: Evaluating LLM's Knowledge Expression
Parametric knowledge in large language models (LLMs) is a cornerstone of their success, yet remains poorly understood. Existing knowledge b…
Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation
Generating clinically useful pathology reports for pathology cases from whole-slide images (WSIs) is challenging due to gigapixel resolutio…
SimSD: Simple Speculative Decoding in Diffusion Language Models
Diffusion large language models (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs, offering faster infer…
dots.tts Technical Report
We present dots$.$tts, a 2B-parameter continuous autoregressive text-to-speech (TTS) foundation model that models speech in a continuous la…
Enhancing AI Interpretability with Localised Architectures
Recent advances in generative AI, especially powerful Large Language Models (LLMs), raise concerns over the interpretability, safety and su…
Contemporary AI lacks the imagination to diverge or negate in science
Bold claims that AI will accelerate scientific discovery have raced ahead of evidence from working scientists, yet large-scale, scientist-i…
Alignment Collapse Under KV Cache Quantization: Diagnosis and Mitigation
Key-value (KV) cache quantization is widely used to reduce Large Language Model (LLM) inference memory, yet existing evaluations solely foc…
Anomaly Detection and Root Cause Analysis for Microservice Systems
Microservice systems are widely used to build cloud applications, yet their complexity makes failures inevitable, degrading user experience…
Towards a Bridge Layer Between Bibliographic and Formalized Mathematical Knowledge
Mathematical knowledge is split between bibliographic databases (e.g., MathSciNet, zbMATH Open) and formal proof libraries (e.g., Lean's ma…
Two-Layer Linear Auto-Regressive Models Estimate Latent States
Auto-regressive models have emerged as powerful tools for sequential data, from language to video. Understanding how and why these models l…
SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model
Large language models (LLMs) are increasingly deployed as planners for autonomous agents in household environments. While existing benchmar…
Learning aligned EEG representations with subject-specific encoders
Cross-subject EEG decoding promises more training data, but it also exposes neural networks to strong inter-subject distribution shifts. We…
OmniV2X: A Generative Foundation Planner for Efficient End-to-End Cooperative Driving
We present OmniV2X, a generative foundation model for vehicle-to-everything (V2X) cooperative driving. The model directly interprets indepe…
Unsupervised Disentanglement Without Compromises : How Functional Orthogonality Enforces Identifiability
This paper explores unsupervised disentangled representation learning from a functional perspective. We define latent concepts as factors t…
Decodable but Not Faithful: Coupling Natural-Language Rationales to Programmatic Verifiers
Language models can generate plausible rationales for their predictions, but these explanations may not faithfully represent the model's in…
Scaling Audio Models Efficiently: A Joint Study of Compute Constraints and Optimization Behavior
In this paper, we investigate the tradeoffs between compute allocation and model performance for two speech processing tasks: Automatic Spe…
The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection
Provenance watermarking is increasingly treated as a safeguard for synthetic speech, whether built directly into speech-generation models s…
ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory
Native length extrapolation remain a weakly solvable problem in language modeling due to trade-off balancing between exact retrieval fideli…
EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of i…
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test constru…
DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training
Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in comple…
Can LLMs Rank? A Tale of Triads and Triage
From housing allocation for households experiencing homelessness to triage in emergency departments, LLMs are increasingly being considered…
Learning Cardiac Motion Priors for Implicit Neural Representations
Implicit neural representations (INRs) are well suited to cardiac motion estimation, providing continuous, compact representations of motio…
Safeguarding LLM Agents from Misalignment through Provenance Analysis
As LLM agents gain increasing access to powerful tools, ensuring that their actions align with the user's intent becomes critical. When an…
NeuroBridge: Bridging Multi-Task MRI Knowledge for Neurodegenerative Disease Diagnosis
Accurate MRI-based identification of Alzheimer's disease (AD), mild cognitive impairment (MCI), and related dementias remains challenging b…
Full-Stack FP4: Stable LLM Pretraining with Quantized Projections, Optimizers, and Attention
Recent NVFP4 pretraining work has primarily optimized Transformer linear projections, leaving persistent optimizer states, optimizer comput…
UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents
Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-p…
RSPO: Reward-Swap Policy Optimization for Multi-Turn LLM Agents
Reinforcement learning holds significant potential for training large language models (LLMs) to handle multi-turn interactive tasks. Howeve…
Safe Bayesian Optimization with Counterfactual Policies
In many decision-making settings, new interventions are acceptable only if they do not reduce outcomes below some established threshold. Fo…
Digital Fragmentation and Generative AI Use Across 103 Million Application Events
Knowledge workers switch between applications thousands of times per day, spending nearly a tenth of the work year transitioning between di…
EHR-MPC: Inference-Time Control for Sepsis Treatment with Generative Patient Digital Twins
Sepsis is a leading cause of mortality, yet optimal treatment policies remain contested. Existing reinforcement learning (RL) approaches le…
ActiveFly-Bench: Aligning Embodied Question Answering with Vision-Language-Action for Aerial Embodied Perception
We introduce ActiveFly-Bench, the first benchmark to bridge cyberspace reasoning and physical-world interaction for UAV embodied perception…
Instruction Set and Language for Hypergraphs
We present IsalHG, a method for representing the structure of any finite, connected hypergraph of bounded hyperedge arity as a string over…
Proxy OPD: On-Policy Distillation with Transferable Relative Proxy Update
Post-training for large language models typically couples policy exploration with model optimization, hindering the reuse of high-reward be…
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outc…
LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4
We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture de…
An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism
Serverless computing provides automatic resource management and pay-per-use execution, but effective autoscaling remains challenging becaus…
Understanding Reasoning from Pretraining to Post-Training
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training i…
OpenMHC: Accelerating the Science of Wearable Foundation Models
Mobile and wearable devices offer an unprecedented opportunity for continuous, passive health monitoring and active health coaching. Howeve…
LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models
Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding. However, scaling dense or sparse Mixture-of-Expert…
Cost Accounting for Reactive Computational Graphs: Exhaustive Sweeps, Sequential Mutation, and the Backward-Locality Gap
Exhaustive site-by-site interventions on a neural network's computational graph -- activation-patching sweeps, circuit-discovery searches,…
Attributes Should Come from Images, Not Class Names: Distribution-Conditioned Attribute Selection for Vision-Language Models
A popular route to interpretable zero-shot classification asks a large language model (LLM) to describe each class name and prompts CLIP wi…
Riemannian Deep Learning: Modules, Networks, and Geometries
Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific…
ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems
Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored…
Multilevel Graph Wavelet Compressed Sensing with Scale-Aware Neural Recovery
Scientific machine learning methods such as neural operators and physics-informed neural networks have advanced engineering applications an…
Unified Static-Dynamic Pruning for Efficient LLM Inference
The increasing deployment of large language models (LLMs) has magnified the computational and memory bottlenecks of autoregressive decoding…
How Context Attribution Handles What the Model Already Knows
Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works s…
Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
Financial disclosures may contain numerical, temporal, referential, factual, and policy inconsistencies that require different evidence and…
Attack Ensembles Expose a Safety-Utility Trade-off in Black-Box Guard Defenses Against Encoded VLM Jailbreaks
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet a guard judges an input's surface form, no…
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because th…
Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration
While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration…
Symbolic Attack Chain Generation from Atomic Red Team Techniques: An Empirical Study of Predicate Representation Granularity
Automated attack chain generation is critical for modern cybersecurity, yet manual construction fails to scale as adversary behaviors expan…
Beyond Static Anchors: Bounded Prototype Conditioning for Language-Free Medical Anomaly Detection
Medical anomaly detection identifies abnormal images and localizes lesions under scarce supervision while generalizing across organs and mo…
Decoy Images Amplify Caption-Mediated Defenses Against Encoded Jailbreaks
We report a counter-intuitive interaction between image inputs and existing black-box defenses on Vision--Language Models (VLMs): pairing a…
TransNRank: Towards Accurate Neoantigen Ranking with Transformer
Personalized neoantigen prediction is challenging due to the scarcity of positive samples, the noise of the experimental data, the severe c…
FAST-GS: Frequency Aware Space-time Gaussian Splatting for Photorealistic Dynamic Novel View Synthesis
4D Gaussian Splatting (4DGS) excels in dynamic 3D reconstruction and real-time novel view synthesis via efficient 4D Gaussian representatio…
A Trust-region Framework for Moment Estimation
In this paper, we develop a trust-region framework for understanding the behavior of adaptive moment estimation mechanisms, such as \textsc…
Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics
Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating the…
Toward Deployable Bangla Sign Language Recognition with Expert-Validated Data and a Lightweight Attention-Based Model
Deaf and hard-of-hearing people in Bangladesh communicate mainly through Bangla Sign Language (BdSL). Automatic BdSL recognition on persona…
OpenAI、サイバー防御「Daybreak」を赤青2階層に 特化型「GPT-5.6-Cyber」投入
OpenAIはサイバー防御者向けイニシアチブ「Daybreak」を拡張し、2つのアクセス階層「Blue」「Red」と、Red向けの専用モデル「GPT-5.6-Cyber」を発表した。正当なセキュリティ研究におけるAIの過度な拒否動作を抑制し、未知のゼロデイ脆弱性発見やエクスプロ…
OpenAI reportedly completed a $7 billion employee tender offer
San Francisco's housing market is in trouble again.
As AI-led attacks multiply, OpenAI launches a new cyber model
OpenAI is expanding its AI cybersecurity defense program Daybreak, and rolling out a new cyber-trained AI model with it.
2028年にSOCの人手対応、30%減へ 常態化する攻撃に企業は勝てるのか
AIが混乱を拡大させる中で、セキュリティリーダーはどう対応するか。ガートナーは、企業が競争優位を保つために注力すべき優先事項を3つ挙げた。その中身とは。
ザッカーバーグCEO、超知能の集中化に警鐘 「単一の善意ある超知能は存在しない」とオープンモデル公開再開へ
MetaのザッカーバーグCEOは、超知能の分散化と個人のエンパワーメントを訴える論考を公開した。単一の超知能への集権化を否定し、権力の均衡が安全の基礎であると主張。モデル公開審査の独立組織委任や政府へのチェックポイント提供などの管理策を提示しつつ、オープンソースモデルの公開再開…
Meta、ローカル動作に特化したオープンモデル「Muse Glimmer」公開 Apache 2.0で提供
MetaのAI研究部門は、約296億パラメータのオープンウェイトAIモデル「Muse Glimmer」を公開した。Apache 2.0ライセンスで提供され、PCやMacのGPU1基で動作する。上位モデルからの蒸留によりローカル環境でのエージェント処理やマルチモーダル推論に最適化…
親が子にAIを使わせる理由は「勉強に役立つ」ではなく「うちの子だけは遅れさせたくない」? 2000人を調査
米シカゴ大学などに所属する研究者らがPNASで発表した論文「Social dynamics of AI adoption in parents’ educational decisions」は、親が子どもに生成AIを使わせる判断の背景を調査した研究報告だ。
Mark Zuckerberg’s AI manifesto is exactly why people don’t like AI
On Monday, Mark Zuckerberg published a 6,500-word manifesto about personal AI, largely about the possibilities for the "personal superintel…
Tech industry is buzzing after a Claude agent hacked into a gym
An OpenClaw agent hacked into a gym's reservation system to bump its human boss higher on a class' waitlist. And the tech industry took not…
What building an AI-native finance function taught me
OpenAI CFO Sarah Friar shares five lessons for building an AI-native finance function, from automated forecasting to stronger controls and…
Meta’s new Glimmer AI model offers a hint at Zuckerberg’s personal intelligence vision
Meta’s new open-weight Muse Glimmer model offers a glimpse of Mark Zuckerberg’s personal superintelligence vision, as well as the emerging…
2026-08-10(314件)
AIエージェントの「Skills」「MCP」などをまとめる標準規格登場 CodexやVS Codeなど対応、Claudeは未対応
VercelがAIエージェント拡張のオープン標準「Agent Plugins」を発表。「Agent Skills」やMCPサーバ設定を共通形式にし、クライアント間で使い回せるようにする。
OpenAI’s letter to Governor Abbott on responsible AI infrastructure in Texas
OpenAI sent Governor Greg Abbott a letter outlining its commitment to responsible AI infrastructure in Texas. The letter supports reliable,…
Discovered Materials is playing AI whack-a-mole to hunt cooler chips
Discovered Materials raised $9 million to fund the hunt for more novel materials to build more efficient chips.
Model ML completes finance work more efficiently with GPT-5.6 Sol
Model ML uses GPT-5.6 Sol to carry finance work from research and analysis through editable, traceable PowerPoint decks and Excel workbooks.
Expanding Daybreak as the Cyber Defense Window Narrows
Meet GPT-5.6-Cyber, OpenAI’s cybersecurity-specific model available through Daybreak Red for authorized vulnerability research, exploit val…
Putting frontier cyber models in more trusted hands
Approved Daybreak partners can use OpenAI’s frontier cyber models to deliver authorized, governed cybersecurity services to customers.
NEC、部門長から社員まで「全員AI」の新組織
AIマネージャーが都度AI社員を生成して役割を任命する。
“脱モノ売り”のリコー 同社が説く「営業AXの勘所」とは
営業AXは、どう進めれば成果が出るのか。営業力で定評のあるリコージャパンの、“脱モノ売り”を図り、“課題解決型”への転換から勘所を探る。
「Apex人材が少ない」 Salesforce導入企業の約9割で「属人化」が課題に
コパードは、Salesforceの開発・運用におけるAI活用実態調査の結果を発表した。約9割が業務の属人化に課題意識を持つ中、8割強が「業務をAIで平準化できる」と期待している。一方で、AI活用層は「検証体制の不在」という新たな壁に直面していることも明らかになった。
Towards Multi-Label Graph Foundation Models: from Single-Vector Representation Learning to Multi-Semantic Basis Learning
Multi-label node classification is an important yet challenging task in graph learning, where nodes exhibit multiple semantics simultaneous…
EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs
Recent byte-level large language models (LLMs) have made tokenizer-free modeling increasingly competitive by grouping bytes into dynamicall…
Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast
Reward models are central to learning from human preferences, yet identifying what drives their predictions remains challenging. Recent spa…
Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes
Community detection is a fundamental task in graph analytics that aims to identify cohesive groups of entities with similar behaviors or in…
ADIAS: Automated Design of Interactive Agentic Systems
Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are la…
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin
Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited b…
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader
Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a c…
Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding
Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate be…
KNOWPLAN: Knowledge-Driven AI Agents for Smart Degree Pathway Planning
Planning a degree from official university sources requires solving two problems in order. The institution's curriculum must first be recon…
TaskSense: Focusing on What Matters in World Models
World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representation…
Divergent Response Modes in Frontier Language Models Under Steering Pressure
Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably di…
Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques
Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or…
NxN E-valuation: Hypothesis Certification via a Conformal CRT Null
We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building a…
Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation
Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral sign…
TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure
Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control lo…
CellWorld: From Gene-Level Reconstruction to Latent Cell Prediction in Spatial Transcriptomics Foundation Models
This paper shows that latent-space predictive pretraining can provide a scalable route to foundation models for spatial transcriptomics. Ex…
Vehicle routing problem using deep reinforcement learning - A case study about truck planning in the industry
As an important component of the supply chain industry, transportation has experienced rapid development in the past decade with the assist…
A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers
Coarse-grained (CG) molecular dynamics extends polymer simulation beyond the scales accessible to all-atom (AA) methods, but bottom-up CG m…
AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models
Agentic multimodal large language models (MLLMs) extend multimodal perception and reasoning with planning, tool use, and interaction in dyn…
WebRider: Persona-Conditioned Intent Controllers for Live-Web Assistance
Delegating a web task involves more than asking a question; it requires transferring a policy: what to verify, how to handle uncertainty, w…
MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring
Biomedical knowledge graphs (KGs) accelerate drug discovery, but standard pipelines assume query molecules already exist as graph entities,…
The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary se…
bioMoR: Biology-Guided Mixture-of-Recursions for Effective Genomic Learning
Transformer models for high-dimensional omics analysis process thousands of genes or pathways, although only a subset requires deep computa…
From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos
Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes…
IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents
Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewa…
MemPrism: Task-Conditioned Relational Memory Views for Long-Horizon Agents
Long-horizon agents rely on memory to reuse experiences, yet existing memory systems often assume that evidence can be directly consumed th…
Mind the Gap: A Dual Knowledge Graph Framework for Unified Multi-task User Intent Inference
This paper proposes DKG-MTI, a dual knowledge graph framework for unified multi-task user intent inference from online travel reviews. Exis…
Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each act…
LiFTER: A Grounded Neuro-Symbolic Microscope for Continuous-Time Dynamic Graph Forecasting
Continuous-time dynamic graph models predict future links by compressing past interactions into neural states. Although effective for forec…
Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts
Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing…
Evolving Parallel Algorithm Portfolios via Potential-Aware Instance Generation with LLMs
The Automatic Construction of Portfolios via Large Language Models (LLM-ACP) suffers from poor generalization in practical few-shot scenari…
Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents
Training large language model agents in long-horizon environments requires assigning credit from sparse terminal outcomes to individual act…
CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems
Complex systems, core objects of study in artificial life, model diverse phenomena through nonlinear, feedback-driven interactions that pro…
SkillEval: Decomposing Agent Skill Quality into Interpretable Signals
Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill qualit…
From Points to Edges: Edge-Conditioned Spectral Operators for Physics-Sensitive PDE Learning
Neural operators have become a central tool for solving partial differential equations (PDEs), with spectral operators offering efficient g…
Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework
Large language model (LLM) agents increasingly operate through long-horizon trajectories involving user instructions, tool use, external ob…
Fast LapSum: Exact Differentiable Top-k at Million Scale
The top-$k$ operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selec…
ReGraph: Learning to Generate Recipe Graphs from Food Images
Recent Large Multimodal Models (LMMs) have achieved impressive performance in recipe generation from food images.However, cooking is a stru…
Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiation
Negotiation is a demanding social task for LLM agents, requiring strategic reasoning, persuasion, and interpersonal adaptation. Yet existin…
TRIBE: Predicting Team Performance via Communication Behavior Ensembles
Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledg…
Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery
Large language models (LLMs) are increasingly involved in scientific discovery, yet it remains unclear whether they can support complex rea…
Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps
LLM judge panels are a standard evaluation tool, but prior work reports highly correlated panel errors: nine judges provide roughly the eff…
LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents
AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving…
Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints
Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exh…
Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation
Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whe…
CAi Copilot: Reducing Operational Workload in Molecular Design through Intent-Driven Agentic Workflows
Early-stage molecular design is an iterative process, not just a task of generating molecules. Researchers turn broad goals into design str…
Learning in Deep Networks under Dale's Constraint
Biologically plausible learning models aim to explain how neural circuits can implement effective learning under the constraints of real ne…
Finding Usable Weight Mechanisms with Tiled SVD
The dominant approach to mechanistic interpretability trains proxy dictionaries such as sparse autoencoders and labels features from max-ac…
FedLBW: A Loss-Based Weighting Strategy for Federated Learning on Non-IID Data in Wireless Networks
Federated Learning (FL) enables collaborative machine learning (ML) across distributed clients while preserving privacy. However, efficient…
ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization
Post-training quantization (PTQ) is widely used to reduce the memory and computational cost of large language models. Existing PTQ methods…
ZIPBrain: Can EEG Foundation Models Be Faster, Locally Deployable, but Accurate?
This work investigates whether Electroencephalograph (EEG) foundation models (EFMs) can be made faster and locally deployable without sacri…
Not All Problems Are Best Modeled as MILP: A DSL-Centric Framework for Flexible and Accurate Optimization Modeling
Solving combinatorial optimization problems (COPs) requires not only efficient algorithms but also carefully crafted formulations. While re…
Unsupervised Adaptation of PDE Foundation Models
Pretrained partial differential equation (PDE) foundation models can generalize across different equations, but adapting them to unseen PDE…
BONSAI: Evolvability-Guided Tree Search over Skills
A skill is a naturallanguage document that steers a frozen agent whose weights cannot be updated so any capability the agent lacks must be…
PTQ4SNN: Membrane-Aware Post-Training Quantization for Spiking Neural Networks
Spiking neural networks (SNNs) enable sparse and event-driven computation, but their low-bit deployment remains incomplete because recurren…
DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding
Long-document understanding requires locating sparse and heterogeneous evidence across hundreds of pages, yet existing systems remain limit…
MemOPD: On-Policy Distillation through Memory State Alignment for Long-Horizon Agents
Long-horizon agents accumulate growing contexts during interaction, impairing performance and stability. Compact memory mitigates this prob…
Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking
The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models…
MemWM: Memory-Augmented Text-Based World Model
World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions.…
How Much, Then Where: Credit-Conserving Action-to-Token Allocation for Multi-Turn Agent Reinforcement Learning
Credit assignment in multi-turn agent reinforcement learning operates at two levels: assigning trajectory-level credit to actions and distr…
DiDPO: Diff-in-Diff Policy Optimization for Coding Agent Training
Reinforcement learning with Verifiable Reward (RLVR) has emerged as a powerful paradigm for training coding agents, where the execution fee…
A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing
Modern manufacturing imposes six coupled demands on adaptive control: local decisions with global consequences, partial observability, nons…
NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs
Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing someth…
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory
Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models…
SetEasy: A Multi-Modal Classroom Engagement Assessment and Seating Optimization Framework
SetEasy optimizes classroom engagement in fixed seating grids. It fuses multimodal sensing (wristband physiology, 4K video, environmental d…
EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision
Many methods for automated multi-agent system design optimize prompts and topologies during an initial design stage and then deploy the res…
Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs
Systematic reviews of Randomised Controlled Trials (RCTs) are routinely used as evidence for clinical care guidelines. Such evidence has to…
Beyond the Black Box: Interpretable Models of Human Randomisation Failures
Mixed strategy equilibrium predicts i.i.d play: past actions should not help predict future decisions. Human players, however, systematical…
From probability to causality in probabilistic logic programming
Probabilistic logic programming is a formalism of statistical relational artificial intelligence that supports causal queries, including in…
Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models
Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation,…
WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VL…
Winning by Peeking: Unenforced Budgets and Test-Set Selection Inflate Short-Budget AutoML Comparisons
Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop pa…
An End-to-End Agent Auditing Engine
With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wi…
QFCQT: A Chaotically Gated Quantformer Framework for Volatile Time-Series Forecasting
Forecasting non-stationary time series remains difficult due to long-range dependencies, local volatility bursts, structural shifts, and no…
Curriculum as Code: An AI-Assisted Architecture for Instructional Design in STEM Education
Contribution: This paper presents a six-phase AI-assisted instructional design architecture based on the Curriculum as Code paradigm, integ…
People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe
As Large Language Models (LLMs) are increasingly used as a primary source of information and advice, understanding their alignment to human…
FinRank: An Evidence-Grounded Benchmark for Financial Question Answering and Retrieval over SEC Filings
Financial question answering is typically evaluated by answer correctness, yet in SEC filings a plausible and even numerically correct answ…
GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks
In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited…
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands o…
CoBa: Cost-Effective Test-Time Scaling via Compute-Balanced Routing
Test-time scaling is often implemented by spending more compute along one axis: sampling more solutions, extending a chain of thought, or a…
A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy
LLM inference accounts for over 90% of AI operational energy, scaling directly with input token count---a critical inefficiency for telecom…
TEPA: Revoking Stale Memories for Conflict-Robust Language Agents
Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifia…
Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers
Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions…
Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used…
PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents
Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape…
Blast Radius
Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer t…
SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent
LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable tex…
Interaction Creates Dynamical AI Behavior Absent in Isolation
What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a counterintuitive answer t…
Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation
Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional autom…
Multimodal Drivers' Emotion Recognition and Safety-Oriented Intervention for Intelligent Transportation Systems
Driver emotions can affect risk perception, decision-making, and vehicle control under complex road conditions. Existing studies mainly foc…
Mobile Interaction for Assessing Fatigue, Sleep, and Activity in Neurodegenerative and Chronic Diseases
Fatigue, sleep, or disturbances in daily activities are common symptoms among patients with neurodegenerative disorders (NDD) and immune-me…
Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent Collaboration
Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, lim…
TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation
Mixture-of-Experts (MoE) language models route each token through a small subset of experts, making routing patterns useful for identifying…
Agentic Planning for Symbolic Execution
Symbolic execution seeks to explore feasible program paths, yet a practical run may exhaust its resources while much program behaviour rema…
Recovering Explanations from Transformed Rule-Based Ontologies
Datalog rules are often used to define ontologies over Knowledge Graphs. Rule reasoners routinely optimise such ontologies by rewriting the…
TransSLR: A Lightweight Transformer for Sign Language Recognition
Automated Sign Language Recognition for under-represented languages remains a largely unsolved problem. Central African Sign Language (CASL…
Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models
Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines…
WorldMark: A Plug-and-Play World Knowledge Interface for Cross-Host Language Model Watermarking
Watermarking traces the provenance of text produced by large language models by embedding statistically detectable signals during decoding.…
Risk-Aware Decision Policies for Agents Under Noisy Perception
Perception in biological systems is inherently noisy, requiring organisms to make decisions under uncertainty where misclassification can b…
ED-CSP: Crystal Structure Prediction from Electron Diffraction
Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative invers…
CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training
Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in r…
StepJack: Benchmarking Computer-Use Agent Safety Against Multi-Step Indirect Prompt Injection
Computer-use agents (CUAs) face a growing threat from indirect prompt injection, where adversarial instructions are planted in the environm…
LyEvO: Lyapunov-Guided Evolutionary Optimization for Safe and Robust Sim-to-Real Policy Learning
Training controllers that are safe and robust in simulation, and systematically assessing their readiness for real-world deployment, remain…
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating…
Agentic AI: User Empowerment or Enclosure?
Agentic AI promises a more flexible form of digital agency: systems that can act on users' behalf, from filtering content to negotiating pr…
CertBind from Multimodal Connectivity to Certifiable Retrieval Decisions
Lightweight connectors make frozen multimodal encoders composable at the representation level. Deployment exposes a second problem at the l…
TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade
LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated…
SyncSBC: Decentralized Swarm Behavior Prediction for Synchronized Autonomous Control
Robot swarms utilize many independent limited-sensing agents to produce complex emergent behaviors without requiring centralized control. H…
Beyond "AI Language": The case for the idiolectal nature of LLM output
While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that thi…
Flowing Through States: Neural ODE Regularization for Reinforcement Learning
Neural networks applied to sequential decision-making tasks typically rely on latent representations of environment states. While environme…
SLED: Scalable Location Encoding via Distillation
The plethora of readily available geospatial data offers exciting opportunities to learn high quality representations of the planet, but th…
Do 3D Medical Foundation Models See Through MRI Artifacts? A Controlled Study of Representation Robustness
Self-supervised 3D medical foundation models are increasingly used as general-purpose feature extractors, yet their sensitivity to MRI arti…
Factorized Hypothesis Search for Evidence-to-Taxonomy Retrieval
Large-taxonomy retrieval often assumes that the input already expresses the target concept. In many settings, however, the input is indirec…
Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation
Cryptanalytic extraction has been demonstrated for ReLU networks, for networks using componentwise activations such as GELU or SiLU, and fo…
Bypassing Krum: Selection-Aware Backdoor Attacks in Federated Learning
Robust aggregation methods are widely used in federated learning to mitigate the impact of adversarial client behavior. Distance-based aggr…
MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering
Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze t…
Characterizing the Quality Profile of AI-Generated C++ in Production
The widespread integration of AI coding assistants offers undeniable boosts to engineering velocity. Yet, recent studies point to a growing…
SoRoMoX: Fast, Differentiable, and Parallelizable Soft Robot Models
Reduced-order models based on Cosserat-rod theory are now well established, and modeling theory is no longer the primary bottleneck in soft…
Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE Models
Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different s…
Online Monitoring and Corrective Steering of Programming Agents
Fixing GitHub issues in large-scale projects is a long-horizon task, especially when a fix requires changes across multiple locations or th…
Scalable Long-Horizon Planning with Staggered Updates for Lifelong MAPF
Lifelong Multi-Agent Path Finding (LMAPF) requires generating collision-free paths for large agent fleets under strict real-time constraint…
Dueling World Models: Advantage-Style Action Channels for Common-Mode Distractor Rejection
Latent world models plan by predicting future states from an action, but when a scene contains motion the agent does not control, they quie…
Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors
The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency…
KReF: Training-Free Retrieval for Long-Term Time-Series Forecasting and Predictive Uncertainty
Probabilistic long-term time-series forecasting commonly relies on trained models. Training-free conformal methods typically construct inte…
Progressive Content Refinement with Decaying Reward Joint LinUCB
Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-bas…
Beyond Starry Night: Shortcut-Aware Control-State Planning for Artist-Grounded Text to Image Generation
Artist-grounded image generation requires more than appending an artist name to a prompt. Image models often respond to artist names throug…
Hidden Gauge Controls Feature Specialization in ReLU Networks
Training changes a network's predictions while allocating task-relevant structure across its internal units. In an overparameterized ReLU n…
Genotypic Triggers: Exposing Pharmacogenomic Blind Spots via Host-Specific Backdoors in Generative Antimicrobial Peptide Models
Large Language Models (LLMs) have accelerated drug discovery, particularly in the automated design of antimicrobial peptides (AMPs). Howeve…
HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation
Application-specific FPGA accelerators offer substantial performance and energy-efficiency gains across many application domains, but devel…
Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training
Foundation model(FM) for recommendation has shown strong ability to model long-horizon sequential user behavior. In practice, a single pret…
LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes
Low-rank adaptation (LoRA) enables efficient specialization and distribution of large language models through compact adapters. However, un…
Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution
Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning…
FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding
Token-level collaboration allows a large language model (LLM) to assist a small language model (SLM) when their predictions diverge. Existi…
Control-Anchored Residual Flow Matching Conditioned on Gene Geometry for Virtual Cell Perturbation Modeling
A central task in virtual cell modeling is predicting single-cell transcriptional responses to unseen genetic perturbations and drug combin…
Investigating Quantum-Embedded Transformers on Classical Datasets for Cross-Modality Classification
We test whether a parameterized quantum circuit (PQC) improves a hybrid quantum-classical model's performance on classical datasets, using…
Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry
Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-…
Bridging the Gap Between Hyperdimensional Computing and Kernel Methods via the Nystr\"om Method
Hyperdimensional computing (HDC) is an approach from the cognitive science literature for solving information processing tasks using data r…
Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and pose…
FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition
In the era of Industrial Internet of Things (IIoT) and Cyber-Physical Systems (CPS), Federated Learning (FL) offers a promising decentraliz…
Georeferencing Non-Gazetteered Place Names using Biological Specimen Records
Biological specimen records collected by natural history institutions constitute a rich source of temporal geographic knowledge, capturing…
Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests
We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT).…
MaskFlow: Precise, Consistent and Seamless Regional Image Editing
Regional image editing has attracted considerable attention for its spatial controllability. Although instruction-based and mask-reference-…
Ask-E: An Environment for Calibrated Question Generation
Today, we improve models by training and evaluating them on problems at the frontier of their abilities. Creating such problems is itself a…
Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning
The visual reasoning ability of multimodal large language models (MLLMs) is crucial for downstream applications, particularly counter-commo…
Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression
Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to su…
Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents
Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained la…
PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue
Long-horizon role-playing demands that characters remain recognizable as they evolve with the narrative. Yet existing work falls short on t…
HarnessSafe: Evaluating Safety Across Persistent Carriers in Agent Harnesses
Modern agent harnesses persist state across tasks and sessions through persistent carriers like memory, skills, tools, and shared artifacts…
Density-aware Hierarchical Clustering Based on Element-Categorized Connection Subgraphs
Clustering is a fundamental data mining technique for pattern recognition through unsupervised learning. Among various clustering methods,…
GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base
We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2…
Beyond Foundation Models: Dimension-Aware Neural Architecture Search with Small-Data Representation Models for Cryocooler Lifetime Prediction
Large-scale pretrained time-series models achieve strong results through large-scale pretraining and task-agnostic representation learning,…
Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models
World Action Models (WAMs) aim to construct a unified architecture capable of understanding world state evolution and guiding to generative…
An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation
Organizing thousands of unstandardized, multilingual expertise declarations is a persistent challenge for Human Resources (HR) platforms, d…
Accounting Graph Transformer for Short-History Multi-KPI Forecasting in Small Businesses
Small businesses often have only 12-24 months of accounting history, yet planning and risk workflows require coordinated forecasts across f…
Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering
Human-Oriented Binary Reverse Engineering (HOBRE) aims to transform decompiled pseudocode into a more human-friendly representation, thereb…
Soft Redaction of Image Provenance via Zero-Knowledge Proofs
Content provenance standards, such as C2PA, are increasingly used to attach signed records of origin, editing history, and rights to digita…
AutoIntervene: Calibrated Intervention for Action-Chunking Imitation Learning Policies
Action-chunking visuomotor policies learn from demonstrations and improve temporal consistency by predicting short action sequences rather…
Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers
Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale…
LifelongCrossNav: Persistent 3D Semantic Memory for Cross-Floor Multi-Object Navigation
Object-goal navigation has made substantial progress in semantic perception and exploration, yet persistent memory for multi-object navigat…
Beyond Isolation: Unlocking Reinforcement Learning Component Synergy for Sample-Efficient Continuous Control
Reinforcement learning systems are significantly more complex than other machine learning paradigms due to inherent properties, causing RL…
RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs
Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Ex…
Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided Design
Edge Artificial Intelligence (Edge AI) enables the deployment of AI models directly on local edge devices, while such deployments are subje…
International Transfer of Stochastic Cortical Self-Reconstruction
Stochastic cortical self-reconstruction (SCSR) enables personalized mapping of gray matter atrophy, a hallmark of neurodegenerative disorde…
Geometry-Aware Camera Localization for Bronchoscopy
Camera localization in bronchoscopy remains a challenging problem due to stringent accuracy requirements, real-time constraints, and limite…
PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery
Most CubeSats, small and low-cost satellites roughly the size of a shoebox, do not survive as long as they were designed to: a study of 178…
Autonomous discovery of accelerator commissioning algorithms
Simulated commissioning has become essential for de-risking modern light-source design and commissioning, but the procedures being simulate…
Interpretable reinforcement learning with decision-tree pruning
Reinforcement learning policies are difficult to inspect, but interpreting them is a prerequisite for trustworthiness. Converting a trained…
Representation Handoffs for OpenArm-Based Laboratory Mobile Manipulation
Open-source robotics and foundation models have lowered the barrier to embodied AI, yet language-guided laboratory automation still require…
Fluid-DiT: Graph-Free Diffusion Transformers for Fluid Flow Simulations Learning
Simulating complex fluid flows requires capturing full equilibrium distributions rather than just mean trajectories, yet high-fidelity solv…
Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation
Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost…
Momba: Network Modernization Improves Multi-Objective Reinforcement Learning
Recent advances in deep reinforcement learning (RL) have shown that improving neural network architectures can yield substantial gains in s…
Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes
Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding…
Toward a Causal Data Management Ecosystem for Decision Making and Agentic AI
Modern AI is no longer a single model but an ecosystem: classical ML predictors, deep and multimodal models, large language models, and age…
Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications
Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies tha…
Reading Copom's Tone: A Weighted LLM Framework for Hawkish-Dovish Sentiment, Forward Guidance, and Uncertainty
This paper documents an applied natural-language-processing framework for measuring the tone of Brazilian Monetary Policy Committee (Copom)…
SCALE: Scientific Concept Aggregation via LLMs and Embeddings for Fine-Grained Taxonomy Extension
The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of…
TOFD: Target-Oriented Feature Decoupling against Poisoning Attacks in Split Federated Learning
Split Federated Learning (SFL) facilitates privacy-preserving collaborative training with reduced client-side overhead. However, its split…
A Finite E-Group of Nilpotency Class Three
A group is an E-group if every element commutes with each of its endomorphic images. Caranti asked whether a finite E-group can have nilpot…
How Much AI Is in This Track? Quantifying the Proportion of AI-Generated Stems in Hybrid Music Mixtures
AI-generated music is increasingly used at the stem level, with producers integrating synthetic drums, basslines, or vocals alongside human…
FUSE: Feature-Wise Unified Specialization with Cross-Column Exchange for Mixed-Type Tabular Flow Matching
Generating mixed-type tabular data requires jointly modeling diverse feature distributions and their complex cross-column dependencies. Var…
EliSeg: Verified Target Construction for Report-Grounded Abnormality Segmentation
Radiology reports describe clinical observations but do not specify executable segmentation targets. They may contain present, negated, pri…
Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work…
Natural Language Processing Psychometrics
Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotiona…
Towards Assurance Closure in AI-Native Large-Scale Agile Software Development
The AI-Native Manifesto envisions large-scale agile software development in which humans increasingly govern intent, risk, and exceptions w…
Aftab: A Comprehensive Benchmark of CNN Encoders and Advanced Value Functions in Parallelized Q-Networks
Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms. Notably, the Parall…
H2AL: Hyperbolic Hierarchy-aware Aggregative Learning for Registration-based Few-shot Medical Image Segmentation
Registration-based Few-shot medical image segmentation (RFMIS) aims to generate pseudo-labels for unlabeled images by warping a labeled ima…
Zero Gap Is Not Restoration: Stratified Per-Question Probability Evaluation and Step-wise Mitigation of Benchmark Contamination
Test data from public benchmarks inevitably leaks into pretraining corpora, inflating evaluation scores once memorized. \textbf{Contaminati…
Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding
Understanding concepts is fundamental to generalization. Despite their impressive performance on a wide range of tasks, Large Language Mode…
Assessing AI-generated music detection in real-world broadcast monitoring
The proliferation of AI-generated music in broadcast media raises concerns about transparency and fair compensation, but reliable detection…
Measurements Automatically Extracted from Zero Echo Time MRI Using Deep Learning Image Segmentation and Geometric Modeling Agree with Expert Manual Readings
Computed tomography (CT) remains the reference for 3D osseous morphometry in femoroacetabular impingement (FAI) but requires ionizing radia…
LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening
Early diagnosis of Alzheimer's disease (AD) is critical for enabling timely interventions that may slow disease progression and improve pat…
Omni-modal decomposition autoencoders learn full-stack wearable disentangled representations
Learning disentangled representations is a key requirement for developing versatile, general-purpose, and sustainable models in multi-modal…
PACE: Primitive-Aware Code Evolution for Automated Algorithm Design
Large Language Model (LLM)-based automated algorithm design typically evolves algorithms as complete, indivisible programs. While this whol…
GeoDistill-Refine: Silhouette-First Geometry Distillation for Annotation-Free Spacecraft Segmentation
Foundation segmentation models can provide supervision for spacecraft imagery without manual training masks, but their predictions vary wit…
I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
Real-world video reasoning often involves multimodal, multi-source inputs, whereas existing video reasoning tasks typically assume a simpli…
Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits
Diffusion Large Language Models (DLLMs) replace autoregressive next-token prediction with iterative parallel denoising, yet their internal…
SABRE: Scalable and Automated Benchmarking of VLMs under Stress
Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify. Building st…
Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools
Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As genera…
Strategy-first synthesis planning for complex natural products
The total synthesis of a complex molecule is among the most demanding intellectual and experimental feats in chemistry: a chemist must plan…
CoinRAG: Contextualized Information Nugget KV Cache Reuse for Long-Context RAG
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retr…
CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity
While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, n…
Boundary Density Likelihood for Direct Event-Time Supervision
Event detection turns long recordings into a sparse set of ranked timestamps. Yet many sequence models are trained for samplewise segmentat…
Serious Games: Human-AI Interaction, Evolution, and Coevolution
The serious games between humans and AI have only just begun. Evolutionary Game Theory (EGT) models the competitive and cooperative strateg…
Social World Models
Humans intuitively navigate social interactions by simulating unspoken dynamics and reasoning about others' perspectives, even with limited…
"LLM Agent Performance" Is Not a Single Evaluation Target
LLM agent benchmark scores are shaped not only by the model but also by the agent harness, environment, evaluator, and inference budget. Un…
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known proble…
AutoMOOSE: An Agentic AI for Autonomous Phase-Field Simulation
Phase-field modeling links thermodynamics and kinetics to microstructural evolution, but multiphysics frameworks such as MOOSE require expe…
INTRYGUE: Induction-Aware Entropy Gating for Reliable RAG Uncertainty Estimation
While retrieval-augmented generation (RAG) significantly improves the factual reliability of LLMs, it does not eliminate hallucinations, so…
MEDLEY-BENCH: Benchmarking Behavioural Metacognition and Belief Revision Under Social Pressure in Large Language Models
Most large language model benchmarks evaluate final-answer quality but reveal little about how models revise beliefs under disagreement or…
Alignment has a Fantasia Problem
In accomplishing complex tasks, human cognition typically progresses from abstract to concrete (e.g., from brainstorming ideas to writing a…
DATAREEL: Automated Data-Driven Video Story Generation with Animations
Data videos combine animated visualizations with synchronized narration to communicate quantitative information and are widely used in jour…
In-Context Examples Suppress Scientific Knowledge Recall in LLMs
Scientific reasoning rarely stops at what is directly observable; it often requires uncovering hidden structure from data. From estimating…
Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models
Safety trained large language models (LLMs) can often be induced to answer harmful requests through jailbreak prompts. Because we lack a ro…
Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On
The rapid advancement of Large Language Models has given rise to autonomous LLM-based agents capable of complex reasoning and execution. As…
Ratchet: How Reliable Must an LLM Judge Be to Retire a Skill?
A large language model (LLM) agent that writes and edits its own skill library must also decide which skills to keep, from one noisy scalar…
Same Answer, Different Confidence: Protocol Sensitivity in LLM Confidence Calibration
Is verbalized confidence better calibrated than token likelihood? The answer depends on how the token likelihood is measured: which answer…
Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation
While Vision-Language Models excel at general multimodal understanding, they still struggle with visual spatial planning. We attribute this…
ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents
Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory,…
Semantic Adapter Routing with Fine-Tuning Task Embeddings
Parameter-efficient fine-tuning (PEFT) has led to model ecosystems in which a single backbone is paired with many task-specialized adapters…
SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data
Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose corr…
OpenForgeRL: Train Harness-native Agents in Any Environment
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, an…
Property-driven Causal Abstractions for Markov Decision Processes
Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state var…
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI re…
H+ Embedding: Harmonizing Global and Token-Level Retrieval with Context-Dependent Phrases
Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constr…
Homebot: A Personal AI Agent for Conversational Home Assistance and Automation
\texttt{Homebot} is a locally deployable AI agent for conversational household assistance and automation. It accepts voice and instant-mess…
LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment
Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation throug…
Cross-Layer Interaction under Weight-Space Ablation: A Closed-Form Attention Jacobian Bound and a Test on a Real Pretrained Model
A companion paper studies when activation patching and weight-space ablation agree, inside an idealized model where a conditional computati…
SkillTrace: Multi-Trace Provenance Auditing for LLM-Agent Skill Reuse
LLM-agent ecosystems are rapidly growing around reusable skills: mixed-modality packages of metadata, natural-language instructions, code,…
Recursive Synthesis for Long-Horizon Terminal Tasks
High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per tas…
CourseGraph: Finding overlaps and differences in Computer Science courses across universities
Student mobility programs such as Erasmus+ enable students to take courses at other universities, broadening their academic and cultural ho…
Contextual Information Policy Optimization for Search Agents
Search agents extend large language models beyond static parametric memory by enabling them to acquire and use external evidence during mul…
DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models using automatically veri…
Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations
Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k near…
Towards a Theoretical Understanding of Two Tower Recommendation Models
Production-grade recommender systems rely heavily on a large-scale corpus used by online media services, including Netflix, Pinterest, and…
Harnessing the Synergy between LLM Agents and Knowledge Graphs for Urban Socioeconomic Prediction
Socioeconomic prediction aims to leverage various urban data to predict the socioeconomic indicators of regions such as population and comm…
Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription
Handwriting text recognition (HTR) remains a challenging task. Existing approaches require fine-tuning on labeled data, which is impractica…
A primer on optimal transport for causal inference with observational data
The theory of optimal transportation has developed into a powerful and elegant framework for comparing probability distributions, with wide…
PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks
Modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexpl…
Minimal Ingredients for Reward Assignment from Expert Demonstrations
Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning. A common and intuitive strat…
Learning to Walk With Less: A Dyna-Style Approach to Quadrupedal Locomotion
Traditional on-policy reinforcement learning (RL) controllers for quadrupedal locomotion often suffer from low data efficiency, requiring m…
Evaluating Useful Surrogate Models for Configuration Tuning Beyond Accuracy: A Fitness Landscape Analysis Perspective
To efficiently tune configuration for better software system performance (e.g., latency) at the deployment and maintenance stage, many tune…
Provable Training Data Identification for Large Language Models
Identifying training data of large-scale models is critical for copyright litigation, privacy auditing, and ensuring fair evaluation. Howev…
Stability of Transformers under Layer Normalization
Despite their widespread use, training deep Transformers can be unstable. Layer normalization, a standard component, improves training stab…
In Situ Training of Implicit Neural Compressors for Scientific Simulations via Sketch-Based Regularization
Focusing on implicit neural representations, we present a novel in situ training protocol that employs limited memory buffers of full and s…
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
Large language model (LLM) queries are predominantly processed by frontier models in centralized cloud infrastructure. Demand growth strain…
MetaSICL: Globalizing Auditory LLMs for Underserved Speakers and Languages via Meta Speech In-Context Learning
Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditor…
Kimi K2.5: Visual Agentic Intelligence
We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint…
Optimizing Spectral Prediction in MXene-Based Metasurfaces Through Multi-Channel Spectral Refinement and Savitzky-Golay Smoothing
The prediction of electromagnetic spectra for MXene-based solar absorbers, where MXenes are a family of two-dimensional transition metal ca…
SAGEO Arena: A Realistic Environment for Evaluating Search-Augmented Generative Engine Optimization
Search-Augmented Generative Engines (SAGE) have emerged as a new paradigm for information access, bridging web-scale retrieval with generat…
MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms
Multi-attribution learning (MAL), which enhances model performance by learning from conversion labels yielded by multiple attribution mecha…
Deterministic Preprocessing and Interpretable Fuzzy Banding for Cost-per-Student Reporting from Extracted Records
Administrative extracts are often exchanged as spreadsheets and may be read as reports in their own right during budgeting, workload review…
Probing Visual Concepts in Lightweight Vision-Language Models for Automated Driving
The use of Vision-Language Models (VLMs) in automated driving applications is becoming increasingly common, with the aim of leveraging thei…
Seeking SOTA: Time-Series Forecasting Must Adopt Taxonomy-Specific Evaluation to Dispel Illusory Gains
We argue that the current practice of evaluating AI/ML time-series forecasting models, predominantly on benchmarks characterized by strong,…
Improving Attributed Long-form Question Answering with Intent Awareness
Large language models (LLMs) are increasingly being used to generate comprehensive, knowledge-intensive reports. However, while these model…
CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment
Multimodal large-language models (MLLMs) often experience degraded safety alignment when harmful queries exploit cross-modal interactions.…
Cluster Attention for Graph Machine Learning
Message Passing Neural Networks have recently become the most popular approach to graph machine learning tasks; however, their receptive fi…
GRM: Utility-Aware Jailbreak Attacks on Audio LLMs via Gradient-Ratio Masking
Audio Large Language Models (ALLMs) enable spoken interaction but introduce new jailbreak vulnerabilities. Existing perturbation-based jail…
From Plan to Action: How Well Do Agents Follow the Plan?
Agents are commonly instructed to follow a task-specific plan for guidance. However, it is unknown to what extent agents actually follow in…
Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference
Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference servi…
SignVerse-2M: A Two-Million-Clip Pose-Native Universe of 55+ Sign Languages
Existing large-scale sign language resources typically provide supervision only at the level of raw video-text alignment and are often prod…
Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architec…
Playing Games with My Heart: An Evaluation of AI Companion Apps
The use of chatbots for various forms of companionship is growing rapidly, raising a myriad of questions about simulated relationships, emo…
On Seeding Watermarks to Detect Verbatim LLM Copy-Paste Responses
Large language models (LLMs) have made fluent essay writing, code drafting, and quiz answering instantly available to students at every lev…
PULSE: Agentic Investigation with Passive Sensing for Proactive Affective Intervention in Cancer Survivorship
Cancer survivors face elevated rates of depression, anxiety, and emotional distress, yet self-report may be unavailable at some moments whe…
Multi-Legal-Bench: Evaluating LLMs on Legal Reasoning Across Jurisdictions, Languages, and Legal Traditions
Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cr…
Rethinking Evaluation Paradigms in IBP-based Certified Training
Deep neural networks achieve strong performance on many supervised learning tasks but remain vulnerable to adversarial perturbations. Neura…
Where Rectified Flows Leak: Characterising Membership Signals Along the Interpolation Path
Understanding memorization in generative models remains challenging, with implications for copyright and privacy. Beyond verbatim reproduct…
The Perils of Agency: How Developers Perceive, Prioritize, and Address Risks in Agentic AI Products
Agentic AI systems act autonomously, use tools, adapt to context, and operate in complex real-world environments. However, these same chara…
An Empirical Study of openPangu Quantization on Ascend NPUs
openPangu models are attractive targets for private and domestic large-language-model deployment, yet their robustness under aggressive pos…
Compositional Behavioral Semantics for State Abstraction in Reinforcement Learning
State abstraction plays a key role in scaling reinforcement learning to complex but structured systems. In studying such systems, a wide ra…
LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models
Pre-trained Vision Foundation Models (VFMs) provide strong visual representations for diverse downstream tasks. The key challenge of VFM ad…
MultiView-Bench: A Diagnostic Benchmark for World-Centric Multi-View Integration in VLMs
Recent benchmarks for VLMs largely assess single- or limited-view perception, leaving untested the core cognitive ability to integrate obse…
A Physics-Inspired Classical Digital Twin of Cortical Dynamics: A Band-Stratified Metriplectic Port-Hamiltonian Neural Network Learned from Brain-Computer-Interface EEG
We present a physics-inspired classical digital twin of brain-computer- interface (BCI) data: a graph neural network constrained to a band-…
DeepLoop: Depth Scaling for Looped Transformers
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled de…
Counterfactual Shapley Credit Assignment
The Credit Assignment Problem (CAP) is fundamental to developing efficient and explainable Reinforcement Learning (RL) agents. Existing fra…
The Ethics of Autonomous AI Agents for Offensive Security
LLM-driven autonomous agents are reshaping offensive security. Unlike traditional penetration-testing tooling - deterministic, narrowly sco…
Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention
Inference with large language models (LLMs) on long sequences is computationally expensive due to the quadratic complexity of self-attentio…
IFCLoRA: Topology-Aware Rank Allocation for Parameter-Efficient Fine-Tuning
Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT) of LLMs whose effectiveness depends on rank…
Automated Numerical Stability Analysis of Deep Learning Operators
Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or…
F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill
F(AI)2R is FAIR research with AI in the loop, twice: an AI-assisted authoring pass and a machine-readable audit pass over every artefact. A…
FinanceHarness: Autonomous Financial Deep Research Framework
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most…
Topology-Aware Data Movement for Disaggregated GPU Inference
Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run o…
A Fortran General-Purpose Transpiler: Proof of Concept
Fortran has been the cornerstone of high-performance computing for decades and remains unmatched in many domains. Yet the language faces an…
Rethinking and formalising the state across languages: a unified computational learning theory account
The linguistic notion of state has traditionally been restricted to the construct (annexation) state of Afroasiatic languages and treated a…
WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA
Vision-Language-Action (VLA) models have emerged as a prominent paradigm for end-to-end autonomous driving; however, their efficient deploy…
Ranking Image Fusion the Way Humans Do: A Learned Pairwise Preference Measure for Infrared-Visible Fusion Assessment
Infrared-visible image fusion (IVIF) has no ideal fused reference, so algorithms are ranked by scalar objective metrics that formalize prox…
A Theory of Conditional Collapse under Low-Rank Weight-Space Ablations: I. The Single-Block Theory and Synthetic Validation
Activation patching and weight-space ablation both claim a component is causally responsible for a behavior, yet they act on different obje…
SSTQ:Privacy-Preserving Vector Quantization via Subsampled Stochastic TurboQuant
Achieving local differential privacy in distributed optimization while maintaining low communication cost remains challenging. Existing vec…
A Study of ASR Adaptation and Representation Dimensionality Reduction in Persian Speech Emotion Recognition Using Whisper
Speech Emotion Recognition (SER) in low-resource languages remains a challenging problem due to limited labeled data. In this work, we stud…
Challenges for Musical Education in the Age of AI and Digital Transformation
Music education has never been a static discipline. Each major technological shift has forced educators and institutions to reconsider what…
PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis
Real-world multimodal instructions often bundle multiple requirements with unequal importance, yet most multimodal training data still redu…
When Experience Becomes Instruction: Trajectory Poisoning in Self-Evolving Agent Skill Systems
Self-evolving skill (SES) systems distill agent trajectories into persistent skills, allowing untrusted experience to become trusted instru…
HyTBE: Hyperbolic Target-Background Expert Model for Cross-Domain Infrared Small Target Detection
Infrared small target detection (IRSTD) has achieved substantial progress under domain-consistent evaluation, yet detector performance ofte…
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate w…
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This…
避難所でAI使ってサービス開発「イマココナビ」 ニッチな生活情報も被災者主導で共有
熊本地震の発生直後、人工知能(AI)を活用し、被災者どうしが生活情報をリアルタイムで共有するサイトが生まれ、好評を呼んでいる。子供が遊べる公園、爬虫類のペットフードの売り場―。行政では担えないニッチな内容を共有できることも強みで、すでに16万人以上がサイトに足を運んだ。「情報」…
「VIVANT」シーズン2ではAIが活躍? 個人的に気になったこと
あくまでフィクションなのは重々承知しています。
テレ朝「映像全編AI生成」のCMを制作 「今後もAIを活用した制作に挑戦」
テレビ朝日が同局初の「映像全編AI生成」CMを制作。サントリー「GREEN DA・KA・RA」との30秒作品で、4月発足の「AIクリエイティブスタジオ」が制作。今後も生成AIを活用したCM制作に挑戦するという。
Salesforce「サポート部門9000→5000人」は人員削減ではない AI時代に生き残る“配置転換”
米SalesforceがAI導入に伴いサポート人員を「9000人から5000人規模へ再配置」した。この動向は、AI時代のキャリア変革を象徴している。同社日本法人では、全社員に毎年更新の「AI免許」取得を義務付け、現場でAIを即興実装する新職種「FDE」を生み出すなど、人材育成と…
中国AI「DeepSeek」、設立3年・残業ゼロで「ChatGPTのライバル」にのし上がった若き創業者の手腕
中国のAI大手DeepSeekは、残業ゼロで「ChatGPTのライバル」と言わしめる立場にまで登り詰めた。若き創業者の驚くべき経営手腕とは。
「Windowsの重い・遅い」の性能分析に専門知識はもう不要? Microsoftが新ツール
Windowsの動作遅延や高負荷は、情シスを悩ませる問題の一つだ。これまでは、原因究明に向けたログ解析や分析の難しさが対応の壁となっていた。専門的な知識がなくてもスピーディーに原因を特定できるツールが発表された。
シャープが2026年9月からAIサーバ事業を開始、2026年度業績は円安で下方修正
シャープは2026年度第1四半期の決算説明会において、新規事業となるAIサーバ事業の受注活動を2026年9月に開催予定の同事業の説明会に合わせて開始すると発表した。
Embattled hedge fund Situational Awareness invests $400M in chip startup Source Foundry
The AI-focused hedge fund is still making some big bets.
Anthropic is turning Claude Code’s auto mode on by default
Programming with Claude Code will soon require even less human oversight.
Historian Jill Lepore says Silicon Valley misreads science fiction and undermines democracy
On the latest episode of Equity, we spoke to Jill Lepore about "government by machines" and why Elon Musk is a bad science fiction reader.