週次AIニュース 2026-W31
対象期間: 2026-07-27 〜 2026-08-02(1292 件)
トピックの推移
トピック別件数
- 研究/論文 525件
- LLM/生成AI 489件
- エージェント 279件
- 画像/動画生成 159件
- ビジネス/資金調達 81件
- ロボティクス 53件
- その他 38件
- ハードウェア/半導体 37件
- 規制/政策 8件
今週のハイライト(上位 10 件)
Advancing responsible AI across Europe
OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will c…
Univé builds an AI-ready workforce
See how Univé built an AI-ready workforce with ChatGPT Enterprise by combining leadership, responsible governance, and employee-led innovat…
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool…
Advancing the price-performance frontier with GPT-5.6
Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale.
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compacti…
Accelerating scientific discovery with ChatGPT for Academic Researchers
OpenAI is giving 100,000 academic researchers free access to ChatGPT's most advanced AI models to accelerate scientific research, collabora…
Scientific computing in the age of agentic AI
A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and disco…
How AI is expanding what people do at work
New OpenAI research shows how AI is expanding what workers do, with ChatGPT users taking on tasks across roles and reshaping job boundaries.
Building abundant intelligence
A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.
Google、パーソナルAI「Gemini Spark」を日本でも利用可能に Chrome統合は米国から
Googleは、パーソナルAIエージェント「Gemini Spark」の提供対象を日本を含む160カ国以上に拡大した。PC停止時やスマホのロック時もGoogleのクラウド基盤上で動作し、トリガーに応じてタスクを処理する。さらに、ログイン情報や保存されたパスワードを活用してWeb…
全件(日付別)
2026-08-02(6件)
Google、パーソナルAI「Gemini Spark」を日本でも利用可能に Chrome統合は米国から
Googleは、パーソナルAIエージェント「Gemini Spark」の提供対象を日本を含む160カ国以上に拡大した。PC停止時やスマホのロック時もGoogleのクラウド基盤上で動作し、トリガーに応じてタスクを処理する。さらに、ログイン情報や保存されたパスワードを活用してWeb…
OpenAI、アクティブユーザー10億人超に 導入企業は200万社超
OpenAIは、アクティブユーザーが10億人、導入企業が200万社を突破したと公表した。推論の保持やコンテキスト管理の改善、本番ソフトウェアの最適化によりコスト削減やトークン生成効率の向上を実現。「GPT-5.6」の一部モデルの値下げはこれらの成果を顧客に還元したものであり、再…
Judge denies xAI’s request to block Minnesota ban on ‘nudify’ apps
Despite a lawsuit from xAI, a Minnesota ban on apps that allow users to “nudify” images can move forward.
YouTuber Hank Green says his AI usage is ‘not healthy’
Green offered a remarkable apology, saying that "the level of dopamine that I've been getting from interacting with LLMs ... is not healthy…
Sam Altman is still making the case for parenting via ChatGPT
OpenAI's CEO seemed excited to share a "cool use case" for parents.
This $9 key physically locks your most addictive apps
This $9 NFC key requires you to physically scan it to unlock distracting apps on your phone.
2026-08-01(9件)
OpenAI reportedly finds evidence that more of its agents ran amok
OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.
India is starting to pay for apps, not just download them
India's app market generated a record $345 million in Q2.
Google nixes its Earth AI feature one day after launch, amid criticism it would spread misinformation
A tool that allowed anyone to generate fake AI-generated imagery and superimpose it over real Google Earth maps quickly spurred backlash.
Sam Altman isn’t the only one who wants to pump the brakes on AI
After years of pushing full speed ahead on AI, OpenAI CEO Sam Altman says maybe it’s time for the AI industry to “pace” itself. The comment…
Snapchat no longer rewards fully AI-generated Spotlight content
Snapchat has adjusted its recommendation systems to ensure that only videos created by real people are eligible for Spotlight recommendatio…
Siri AI could come with a paywall for power users
Apple CEO Tim Cook envisions users being able to buy more compute for Siri AI via Apple's existing iCloud+ subscriptions.
SpaceX won’t remove all of xAI’s unpermitted turbines for another year
SpaceX is building a new power plant for xAI's Colossus data centers, but it won't remove existing, unpermitted turbines for many more mont…
Advancing responsible AI across Europe
OpenAI shares how its safety, security, transparency, and provenance practices support responsible AI governance in Europe. The work will c…
Building abundant intelligence
A full-stack approach to making advanced AI more capable, more affordable, and more widely useful.
2026-07-31(282件)
Smallest.ai raises $13M to build ultra-fast voice AI that sounds genuinely human
The startup is building voice models designed to make AI phone calls pass the Turing test.
AI labs want to pump the brakes, but Amazon and SpaceX are still blasting off
After years of pushing full speed ahead on AI, OpenAI CEO Sam Altman says maybe it’s time for the AI industry to “pace” itself. The comment…
キオクシア、株価3分の1急落は「絶好のタイミング」 過去最高益と8000億円自社株買いで示す自信
キオクシアHDは、2026年4~6月のNon-GAAP営業利益が1兆3262億円で、直近1年間の営業利益を上回ったと説明した。7~9月の売上収益は、2兆3900億円を見込む。経営幹部からは力強い言葉が続いた。
「9カ月かかる作業を3日に短縮」 IBM、レガシー刷新ワークフローをIBM Bobに追加
生成AIによるコード生成が広がる一方で、レガシー環境の刷新はAIの適用が難しい領域だ。IBMは、IBM Bobにレガシー環境向けの専用パッケージを追加することで、同領域における自動化に踏み込む。
キオクシアQ1、純利益が前年同期比4500%増 株式分割・自社株買いも
キオクシアホールディングス(HD)が7月31日に公開した2027年3月期第1四半期(26年4月1日?6月30日)連結決算は、売上収益が1兆7671億1700万円(前年同期比415.5%増)、営業利益が1兆2700億1700万円(同2728.6%増)、純利益が8421億6500万…
キオクシアQ1決算、純利益は前年比4500%増 AIデータセンター向け需要がけん引
半導体大手のキオクシアホールディングスは、2027年3月期第1四半期決算(26年4月1日?6月30日、国際会計基準)の純利益が8421億6500万円で、前年同期比4506%増だったと発表した。
Univé builds an AI-ready workforce
See how Univé built an AI-ready workforce with ChatGPT Enterprise by combining leadership, responsible governance, and employee-led innovat…
研究者10万人にOpenAI「最上位モデル」無料提供へ 日本でも東大、京大など15大学が対象
OpenAIが学術研究者10万人に「ChatGPT」の最上位モデルを無料提供するプログラムを発表した。今夏に1万人から提供を始め、2027年までに10万人規模へ拡大する。日本でも東京大学、京都大学、東京科学大学、早稲田大学、慶應義塾大学など15の大学が対象機関となっている。
Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models
Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT)…
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under…
ClinLens: Towards Long-Horizon Coding Agents for Longitudinal Multimodal Clinical Data Science
Clinical data-science agents must transform heterogeneous longitudinal records into auditable analyses, yet existing benchmarks largely iso…
When benchmark inferences do not compose: Projectibility in AI evaluation
An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidenc…
GuideSkill: Evolving Executable LLM Agent Skills for Guideline-Grounded Clinical Reasoning
Clinical practice guidelines (CPGs) encode diagnostic criteria, but LLM systems typically retrieve guideline text or absorb it through trai…
GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure
Functional verification dominates integrated circuit (IC) front-end engineering effort, and a single missed bug that escapes to silicon can…
Position: Evaluation Scores Are Perishable Knowledge Claims
Evaluation methodologies for language models increasingly combine multiple signals, from automated metrics and LLM-as-judge ratings to huma…
TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning
Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code…
Exploring Structures in Physics Problems: Can AI Agents Discover Statistical Mechanical Mappings?
An important skill in theoretical physics is to recognize when a new problem can be transformed into a known model. We study this skill as…
CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games
Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills su…
CG-World: A Large-Scale World-State Dataset and Protocol for World Models
World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datas…
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability t…
EvoPINN: Agentic Discovery of Executable Algorithms for Physics-Informed Neural Networks
Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs), yet their p…
Evidence-Ledger Adjudication for Claim-Evidence Traceability
AI agents can draft claims faster than authors can check whether the cited or retrieved evidence supports them. We study evidence-ledger ad…
Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models
The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM re…
Fewer Clarifications, Better Code: Benchmarking Cross-Session Personalized Ambiguity Adaptation in Coding Assistants
AI-assisted coding increasingly translates informal user intent into executable software, yet coding requests often contain ambiguities tha…
AlphaSchema: Exploring the Space of Trading Semantics for LLM-Based Alpha Mining
Automated alpha mining has increasingly adopted large language model (LLM) agents for factor generation and iterative discovery. However, e…
Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting
Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-wo…
AgenticCANN: Automated Ascend C Operator Generation via Knowledge-Augmented Agentic Evolution
Ascend C operator optimization is critical for NPU (Neural Processing Unit) inference performance but requires deep hardware expertise.Whil…
UrbanDS: A Graph-Guided LLM Multi-Agent System for Data-Intensive Urban Tasks
Large language model (LLM) agents have been widely applied in automating data science tasks. However, existing methods typically rely on a…
Do Latent Channels Actually Communicate? A Causal Audit of Latent Multi-Agent LLM
Latent communication in large language model (LLM)-based multi-agent systems (MAS) transmits continuous internal representations instead of…
Property-driven Causal Abstractions for Markov Decision Processes
Markov Decision Processes (MDPs) are widely used as decision-making models, commonly specified over factored state spaces through state var…
From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence
The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots c…
What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation
Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser…
Belief-Guided Decision Making with Uncertainty Gating in the Game of Go
Recent advancements in Computer Go, driven by AlphaZero and MuZero, rely heavily on Monte Carlo Tree Search (MCTS) to correct the errors of…
Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data
Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not o…
On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment
Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data…
AgentMap: Joint Equivalence and Subsumption Discovery for Ontology Matching
Ontology matching (OM) has traditionally been formulated as either equivalence discovery or subsumption matching. The existing OM systems i…
Linguistic Monoculture in LLM-Assisted Language Use
Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. A…
OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limit…
Partner Capability Estimation for Task-Agnostic Adaptation in Ad-Hoc Teamwork
Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approa…
Can AI agents conduct open-ended AI research? Early evidence from two case studies
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI re…
A Methodology for Designing Knowledge-Driven Missions for Robots
This paper presents a comprehensive methodology for implementing knowledge graphs in ROS 2 systems, aiming to enhance the efficiency and in…
Predict before you train: Scaling Laws for particle physics foundation models
The largest machine learning models in particle physics are also the most expensive to train, yet the return on scaling a given architectur…
Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
Medical-imaging AI benchmarks combine datasets, DICOM rendering, prompts, provider APIs, automated labels, statistical code, manuscripts, a…
Emergent Sparsity in Frozen Random CNN Feature Extractors for Deep Reinforcement Learning
We report a striking phenomenon: deep reinforcement learning agents trained with frozen, randomly initialized CNN feature extractors sponta…
Large-Scale ChatBot Validation Through Customer Digital Twin Simulations
LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remai…
Sim2Win: A Team-Agnostic, Event-Based Pre-Match Outcome Prediction and Tactical Profiling System for Football
Pre-match tactical decision-making in professional football relies heavily on subjective expert analysis and identity-based scouting system…
Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities
Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people…
Archetypes or ability? Clustering for modelling student mathematical competence
Personalised learning systems often assume that mathematical ability is combined of discrete abilities, acquired sequentially and dependent…
The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science
AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond h…
Do Methods Support the Claims? Intra-Paper Verification for Peer Review
The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review. Existing a…
The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty
Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear wh…
The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem
Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not ope…
AI Security Priorities: A Field-Wide Agenda
As AI systems are rapidly integrated into critical economic, governmental, and national security functions, the gap between AI adoption and…
SimpleWikiSearch: A Clean Offline Wikipedia Environment for Agentic Search
Large language model (LLM)-based agentic search systems are often evaluated as if the underlying LLM were the only component that matters,…
GuidedRAG: Semantic Steering of Retrieval-Augmented Generation
In this work, we propose GuidedRAG, a novel extension to traditional Retrieval-Augmented Generation (RAG) that introduces a dedicated selec…
IFCMemoryBench: Evaluating Long-Term Memory of LLM-Based Agents in BIM Information Retrieval
Long-term memory is becoming a core capability of LLM-based agents, but existing evaluations largely test conversational recall in open-dom…
IDP AutoOpt: Agent-Driven Optimization of Document Processing Pipeline Configurations
We present IDP AutoOpt, an autonomous LLM agent that discovers high-performing configurations for intelligent document processing (IDP) pip…
FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents
Retrieval-augmented generation services over mutable enterprise documents repeatedly execute semantically equivalent analysis requests. Ans…
Optimizing Sensor Placement for Hydrogen Leak Detection in Enclosed Infrastructure: A Comparative Study Using CFD-informed Genetic Algorithm and DeepSets Neural Surrogate
Hydrogen infrastructure in enclosed environments, such as parking facilities for fuel cell vehicles, presents significant safety challenges…
A Reference-Free Score for Detecting Silent Reasoning Failures in Large Language Models
Mathematical chain of thought (CoT) evaluation is commonly reduced to whether the final answer matches a reference. This conflates producin…
Weight and Height Estimation from a Single Human Image Captured in the Wild
A person's physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routin…
TraceCLIP: Recovering Local Semantics from Patch-to-CLS Contributions
Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires…
GPT-Red: Automated Red Teaming via Self-Play at Scale
We introduce \textbf{GPT-Red}, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier L…
Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models
Self-repair - returning a failed program to the model together with its test output and asking for a correction - is a standard component o…
Towards Trustworthy Embodied Intelligence: A Systems Framework and Graded Trustworthiness Levels
Embodied intelligence integrates learned perception and decision making with real-time computation, control, and physical interaction. Beca…
A Picture Says Thousands of Words - Harnessing Dermal Exposure Data from Images through Hybrid Deep Learning for Enhanced Safety Assessment
This study developed a hybrid computer vision method to quantify exposed skin from images for dermal exposure assessment. Using 170 indoor-…
Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition
LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities…
(EC)2: Event-Centric Explainability for Cybersecurity Through Multi-Agent LLM Investigations
Security operations centers rely on anomaly detection systems to flag suspicious events. Feature-level explanations for anomaly detectors o…
Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
Multi-Agent Debate (MAD) is a promising paradigm for improving the accuracy and robustness of Large Language Model (LLM)-based agentic syst…
Model-Driven Requirements Configuration with Three-Valued Uncertainty Scoring
Context: Large Language Models (LLMs) offer natural-language flexibility for automated requirements elicitation but frequently generate str…
Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech
AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Ye…
Top-$k$ Pareto Bandits: Hypervolume Regret for Multi-Objective Slate Selection
We consider a stochastic multi-objective bandit problem where, at each round, the agent selects a slate of $k$ arms and observes their $d$-…
Entity Resolution in Practice: Lessons from a Self-Serve Pipeline
We built and evaluated a self-serve entity resolution (ER) system on six benchmarks spanning 864 to 5M records, and three lessons emerged t…
AgentGUI: An Interface for Observing and Steering Long-Running AI Agents
AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight…
SARC-DQ: Runtime Data-Quality Gating for Agentic AI: Silent Evidence Defects, the Incompetence Shield, and Downstream-Only Remediation
Agentic systems act, so a defect in the evidence they retrieve becomes a wrong action with a currency cost. The most dangerous enterprise d…
StealthBench: Measuring Operational Stealth in Autonomous Offensive-Security Agents
Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separat…
Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examin…
Automorphism-Induced Non-Canonicity in Top-k Explanations of Graph Neural Networks
A gradient-based GNN explainer given a molecule with two chemically equivalent nitro groups assigns them attribution scores that are equal…
When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses
Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product…
Pramana: A Composable, Domain-Specific Backend for Empirical Networking Research
Networking research advances by turning hypotheses into empirical evidence, so accelerating it means reducing the lag between ideation (syn…
High-Order Markov Blanket Discovery via a k-Order Relaxation of the Faithfulness Assumption
The problem of learning the graphical Markov blanket (MB) of a variable from data has applications in many areas such as structure learning…
Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limit…
Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text
Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may…
Zero-Fi: Zero-Shot Wi-Fi-Based Human Activity Recognition via Contrastive Signal-Language Alignment
Wi-Fi-based human activity recognition has advanced substantially, but most existing methods assume a closed set of activities and require…
Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction
Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a…
Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad mis…
Voice Memory for Agentic Speech Recognition
We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-doma…
FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing
Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-ground…
Reinforcement Learning on Cost-Constrained Quadrupedal Hardware
Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematicall…
Mergeable Model-Side Aggregation States for Long-Context Language Models
A known limitation of long-context language models is their increasingly unreliable performance in non-additive, set-based aggregation as c…
ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models
Large language models (LLMs) have demonstrated strong capabilities in knowledge acquisition and reasoning, yet their ability to retain prev…
PUDA: An AI-Native Hardware Harness for Self-Driving Laboratories
Physical Unified Device Architecture (PUDA) is an AI-native hardware harness for self-driving laboratories (SDLs). Rather than building a h…
Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection
Audio deepfake detectors often degrade when generators, corpora, or recording conditions change. We use a Diffusion Transformer (DiT), trai…
LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving
The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware…
Collaborative Weighting with Pessimistic Critic for Mitigating Overestimation in Off-Policy Reinforcement Learning
Deep off-policy reinforcement learning algorithms for continuous control typically rely on neural value function approximation to guide pol…
HiFloat4 Format for End-To-End Reinforcement Learning Post-Training of Large Language Models
We present, to our knowledge, the first end-to-end FP4 RL post-training, in which both the rollout and training policies, including their f…
A Graph-Native Bitemporal Memory Store for Conversational AI Agents
Conversational AI agents commonly lack persistent memory across sessions. The obvious fixes like injecting full chat histories into the con…
The Art of Not Forgetting A Local Learning Architecture for Continual Learning
We introduce CMP (Cognitive Memory Primitive), a continual-learning architecture that repre?sents inputs as sparse relational codes, stores…
Shared Symbolic Backbones for Physically Consistent Multi-Output Symbolic Regression
Symbolic regression provides analytical expressions, but it is usually applied one output at a time. This is limiting in process systems, w…
AgentGFM: A Graph Foundation Model with Node-Agent Information-Flow Control
Graph Foundation Models (GFMs) aim to learn transferable knowledge from multi-domain graphs and adapt to unseen scenarios. As a fundamental…
A Persona-based Rate Action Index
We propose an index for predicting the U.S.\ Federal Open Market Committee (FOMC) decision to hike/hold/cut the current federal funds targe…
ServerlessT2I: Efficient Text-to-Image Workflow Serving on a Serverless Platform
Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke…
Recover, Decode, Reguard: Guard-Agnostic Defense Amplification againstEncoded VLM Jailbreaks
Safety classifiers ("guards") are the dominant black-box defense for vision-language models, yet they judge an input's surface form, not it…
Classification of Disease from Lungs X-ray Images using VGG16, VGG19 and ResNet50 Models
With the increase in the number of cases related to respiratory diseases, there is an urgent need to detect them early and diagnose them ac…
One Run Is Not an Idea: The Implementation Lottery in Automated Research
Automated research systems use experimental scores both to deliver artifacts and to decide which ideas to retain, transfer, and pursue. Yet…
A Physics-Informed Framework for PID Tuning of Chemical Processes Using Large Language Model Agents
PID tuning for chemical processes commonly relies on identified process models, whereas plant engineers often retune loops iteratively by o…
Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution
Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a…
Living-Harness Is an Interactive-Agent Evolver
Large language model (LLM) agents may recover from a failure within an episode or after a retry, yet the same execution failure can recur i…
WhisperRec: Latent Reasoning for Efficient Foundation Recommendation Models
Large language models (LLMs) have demonstrated strong reasoning capabilities, motivating their adoption as backbones for foundation recomme…
Understanding Context Sampling in TabPFN on Small Tabular Datasets
TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) an…
Guarding Organizations Against Malware Risk: A Novel Graph-Based Malware Detection Method
Organizational digitalization expands cybersecurity risks, making cybersecurity an increasingly important research area in Information Syst…
Filesystem-Based Memory for LLM Agents: Organization, Evolution, and Sustainability
Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself read…
Borrowed Strength: Best-of-N Search over a Code EncodingBreaks Self-Check Jailbreak Defenses
A self-check defense asks the target model to assess a request before answering it; SAGE, the strongest published instance, reports an aver…
FakeIDet3-DB: Refining Digital Attacks and Patch Extraction for Secure ID Benchmarking
Identity document (ID) authentication relies on the structural integrity of complex, high-frequency security patterns. However, advanced Ge…
FPSGen: Flexible Point Cloud Scene Generation with BEV-Supported Transport Flows
Existing point-based generative methods for outdoor scenes primarily focus on LiDAR-conditioned completion. During training, noisy point cl…
Physically Real-time Infrared Attack against Optical Flow Estimation Networks
With the promising performance of deep neural networks on image-based tasks, different real-world applications such as autonomous driving a…
Constitutional Midtraining: Content Presence Drives Alignment Gains
Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining intervent…
Graph Is the Verifier: Agentic Reinforcement Learning for Interprocedural Vulnerability Detection
Real-world vulnerabilities often span multiple functions, yet most learning-based detectors classify each function in isolation: on a sampl…
Scientific Knowledge Discovery in the Age of Large Language Models
The rapid growth of scholarly literature has made identifying relevant publications increasingly difficult, and conventional search systems…
Efficient Heteroscedastic Bayesian Optimization for Risk-Aware AutoRL
Reinforcement learning (RL) has shown remarkable success across a wide range of complex tasks. However, RL outcomes can be highly stochasti…
MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation
Cover song generation (CSG) should preserve the melodic and linguistic content of a reference song while recreating the remaining musical c…
Automated Multilabel Mpox Research Classification with Explainable Transformer Models
The Mpox outbreak remains a serious public health issue, with the WHO (World Health Organization) reporting increasing cases in some region…
FARI: Robust One-Step Inversion for Watermarking in Diffusion Models
Inversion-based watermarking is a promising approach to authenticate diffusion-generated images, yet practical use is bottlenecked by inver…
Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives
Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target i…
Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model
Zero-shot text-to-speech (TTS) clones a voice from a short audio prompt, but this reliance on reference audio is a barrier when only visual…
Multimodal fusion of visual and morphometric features for avian bone classification
Artificial intelligence has shown considerable potential for archaeological applications, yet its use in zooarchaeology remains limited, pa…
An Attention-Based Framework for Alzheimers Disease Classification Using Resting-State fMRI
Accurate identification of Alzheimers disease (AD) using resting-state functional magnetic resonance imaging (rs-fMRI) remains challenging…
Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text
State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes…
Searching for Robust Augmentations to Improve Out-of-Domain Generalization in Dermoscopic Skin Cancer Classification
Background/Objectives: Dermoscopic skin lesion classifiers often lose accuracy under domain shift across imaging devices, illumination, and…
MediaWiki Code2Code Search: Neural Retrieval for the Semantic Discovery of Open-Source Software Entities
Code search in large-scale ecosystems is often hindered by the lexical gap between user queries and implementation details, alongside the t…
See2Think: Do Multimodal Models Really Use Intermediate Visual States?
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains un…
Journey Operators for Structured Multi-Axis Composition
Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cel…
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforc…
SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response
Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line…
Crossing-Free Probabilistic K-Line Forecasts Without Retraining
Probabilistic K-line forecasting describes uncertainty in four complementary prices, namely open--high--low--close (OHLC). However, it intr…
FedTopo: Relation-Level Topology Sharing for Model-Heterogeneous Federated Learning
Federated learning (FL) enables collaborative learning over decentralized data silos without centralizing raw data. However, heterogeneous…
A First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source Communities
Open source communities have been flooded with AI-generated contributions. In defense, they have written contribution rules to regulate cod…
AI as Friction for Reflection Support in Ideation
Generative AI tools for creative work tend to be designed around the goal of removing friction, on the assumption that smoother iteration a…
Budget-Aware LLM Discovery via Cost-Calibrated Frontier Utility
Large language models increasingly support scientific and algorithmic discovery through inference-time search over evaluated candidates. Ex…
From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies e…
ReCo: Reweighting GRPO Against Distributional Concentration
Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent wor…
Think Short, Defer Smart, Act, and Repeat: Calibrated Reasoning and Uncertainty-Aware Deferral for Edge LLM Agents
LLM agents following the ReAct paradigm are promising enablers of complex multi-step tasks, including multi-hop question answering, code ge…
Hearsay: Vision-Language Medical Diagnoses Without an Image
When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosi…
Human diversity fuels collective creativity that large language models cannot simulate or sustain
Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI as…
Actions Have Consequences: Detecting Outcome Performativity using Intervention Testing
In many domains such as Palliative Care, Credit Assignment and Recommender Systems, predictions may causally influence the outcomes they pr…
BioVLN: A Simulation Platform for Visual Language Navigation in Biomedical Laboratories
Biomedical laboratory robots must navigate to instruments before performing experimental procedures. Existing embodied navigation platforms…
Defending Against Backdoor Attacks via Alignment Checking in Model-Contrastive Federated Learning
Federated Learning (FL) is vulnerable to backdoor attacks because of its distributed nature in edge computing scenarios. Existing defense m…
Progressive Multimodal Alignment for Continual Instruction Tuning
Multimodal Large Language Models (MLLMs) rely on a projector to align visual representations with the language embedding space, making it c…
SymmGrid: Super-Scaling On-Robot Learning with Parallelized Symmetries and Egocentric-Exocentric Visual Perception
Deep reinforcement policy learning directly in physical robots (on-robot learning) remains bottlenecked by slow wall-clock training times.…
BayesAME: Bayesian Active Model Evaluation
Evaluating large generative models across benchmarks is time-consuming and computationally expensive. This drives the need for methods that…
CoCaRS: Correlation Calibration-Based Redundancy Suppression for Heterogeneous Knowledge Distillation
Knowledge distillation (KD) enables a compact student model to learn from a powerful teacher and has become an effective paradigm for model…
ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection
While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarc…
SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence
Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible,…
Visual Credit Audit for Multimodal Spatial Reasoning
Closed yes/no spatial benchmarks can reward a correct answer even when the image adds little support beyond no-image contexts. Under a fixe…
Parameter-Free Dynamic Regret for Online Convex Optimization under Heavy-Tailed Noise
We study online convex optimization (OCO) in non-stationary environments under heavy-tailed noise, where the stochastic gradient oracle adm…
MemSecBench: Tracking Agent Memory Poisoning from Persistence to Consequence and Repair
Memory systems allow agents to retain and reuse information from past interactions, but they can also let malicious content persist. A mali…
Scores Are Not Decisions: Cost-Aware Stopping for Tool Acquisition in LLM Agents
As LLM agents increasingly depend on diverse external services such as search engines, databases, and connectors, agent harnesses face a fu…
SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context
Scientific images are the core elements of presenting experimental conclusions, elaborating system architecture, and supporting comparative…
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
With the development of audio large language models (AudioLLMs), audio captioning needs to move from brief descriptions toward open-ended a…
DLAM: Distributional Latent Actions with Temporal Constraints
Vision-language-action (VLA) models remain constrained by scarce action-labeled robot data, whereas action-free videos offer abundant obser…
Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark
High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantificat…
Anatomy Contextualized Adaption of CT Foundation Models
CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-…
Improving Item Discoverability in e-Commerce Search via Related Intent Generation
Traditional search systems are optimized to retrieve items that strictly match a query, often prioritizing precision over recall. In e-comm…
The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making
Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communica…
APEX-Accounting
We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work…
Decision-oriented joint optimization of evidence fusion based on event-conditioned credibility
In decision-level fusion tasks involving heterogeneous sources with unequal precision and potential anomalies, evidence deviating from the…
Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning
Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Al…
HealthSLM-Bench: Benchmarking Small Language Models for Mobile and Wearable Healthcare Monitoring
Mobile and wearable healthcare monitoring play a vital role in facilitating timely interventions, managing chronic health conditions, and u…
Measure what Matters: Psychometric Evaluation of AI with Situational Judgment Tests
Persona conditioning is widely used to steer large language model (LLM) behavior, but it is unclear whether it induces stable behavioral st…
Balancing Centralized Learning and Distributed Self-Organization: A Hybrid Model for Embodied Morphogenesis
Background: both embodied intelligence and developmental morphogenesis depend on a division of labour between centralized guidance and dist…
BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models
Vision-Language Models (VLMs) inherit significant social biases from their training data, notably in gender representation. Current fairnes…
How does downsampling affect needle electromyography signals? A generalisable workflow for understanding downsampling effects on high-frequency time series
Automated analysis of needle electromyography (nEMG) signals is emerging as a tool to support the detection of neuromuscular diseases (NMDs…
AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing
LLM role-playing aims to portray arbitrary characters in interactive narratives, yet existing systems often suffer from limited immersion a…
TANDEM: Temporal-Aware Neural Detection for Multimodal Hate Speech
Social media platforms are increasingly dominated by long-form multimodal content, where harmful narratives are constructed through a compl…
How memory can affect collective and cooperative behaviors in an LLM-Based Social Particle Swarm
This study examines how memory shapes the collective and cooperative dynamics of Large Language Model (LLM) agents in a multi-agent system.…
MathNet: a Global Multimodal Benchmark for Mathematical Reasoning and Retrieval
Mathematical problem solving remains a challenging test of reasoning for large language and multimodal models, yet existing benchmarks are…
On the Hybrid Nature of ABPMS Process Frames and its Implications on Automated Process Discovery
A core component of any AI-Augmented Business Process Management System (ABPMS) is the process frame, which gives the system process-awaren…
The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure
Multi-agent systems extend large language models (LLMs) by decomposing tasks among specialized agents, but their distributed decision proce…
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
Self-evolving skill libraries face a silent failure mode we term \emph{library drift}: unbounded skill accumulation without outcome-driven…
Ratchet: A Minimal Hygiene Recipe for Self-Evolving LLM Agents
Self-evolving skill libraries, pioneered by Voyager, let frozen LLM agents accumulate reusable knowledge without weight updates, yet recent…
RedKnot: Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
As the input length of large language model (LLM) serving continues to grow, the KV cache has become a dominant bottleneck in AI infrastruc…
BrainG3N: A Dual-Purpose Tokenizer for Controllable 3D Brain MRI Generation
Three-dimensional (3D) brain MRI is central to clinical neurology and neuro-oncology, where generative models could augment under-represent…
Matilda: Engine-Agnostic Search with Human Policy Guidance
Chess engines have evolved from search-based systems optimized for strength to neural policies optimized for predicting human decisions. Ex…
Linguistic Firewall: Geometry as Defense in Multi-Agent Systems Routing
The rapid integration of Large Language Models (LLMs) has driven the evolution of Multi-Agent Systems (MAS), where specialized agents colla…
FirstResearch: Auditable Question Formation for LLM Scientific Discovery Agents
LLM systems for scientific discovery increasingly assist with ideation, literature synthesis, experiment planning, and report generation, b…
When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals
LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles…
Do We Really Need Adaptive Global Spatial Attention for Traffic Forecasting?
Existing traffic forecasting models commonly focus on extracting spatial dependencies, particularly global spatial information, which chara…
Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents
Large language model (LLM) agents are increasingly used for complex information-extraction tasks, yet it remains unclear whether agentic co…
SkillSight: Calibrating Generic Content Bias for Skill Retrieval
As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable c…
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may…
Pushing the Frontier on Approximate EFX Allocations
We study the problem of allocating a set of indivisible goods to a set of agents with additive valuation functions, aiming to achieve appro…
One-Frame Calibration with Siamese Network in Facial Action Unit Recognition
Automatic facial action unit (AU) recognition is used widely in facial expression analysis. Most existing AU recognition systems aim for cr…
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have bec…
The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across…
Task and Skill Planning: Hierarchical Robot Planning with Black-Box Skills
Task and motion planning (TAMP) is a well-established approach for solving long-horizon robot planning problems. Although TAMP methods have…
When Should AI Follow? Task Structure and Joint Adaptation by Human and AI Agents
How should organizations divide and sequence decision tasks between human and artificial agents? We develop a computational model of joint…
Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
Large language models (LLMs) have achieved impressive performance across various domains. However, the substantial hardware resources requi…
AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design Stages
Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle thro…
Equivariant Eikonal Neural Networks: Grid-Free, Scalable Travel-Time Prediction on Homogeneous Spaces
We introduce Equivariant Neural Eikonal Solvers, a novel framework that integrates Equivariant Neural Fields (ENFs) with Neural Eikonal Sol…
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent
Despite improvements by length extrapolation, efficient attention and memory modules, handling infinitely long documents with linear comple…
FPEdit: Robust LLM Fingerprinting through Localized Parameter Editing
Large language models represent significant investments in computation, data, and engineering expertise, making them extraordinarily valuab…
Balancing Privacy and Efficiency: Music Information Retrieval via Additive Homomorphic Encryption
Modern music retrieval runs on vector embeddings, and once these embeddings are shared for search or matching they can be copied, probed, o…
GBPP: Grasp-Aware Base Placement Prediction for Robots via Two-Stage Learning
GBPP is a fast learning based scorer that selects a robot base pose for grasping from a single RGB-D snapshot. The method uses a two stage…
Train Large, Deploy Compact: Structured Compression for Compact Low-Rank Adaptation
Low-rank adaptation (LoRA) has become a widely used paradigm for parameter-efficient fine-tuning of large language models, yet its represen…
VideoNorms: Benchmarking Cultural Awareness of Video Language Models
As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural context…
ARC-Encoder: learning compressed text representations for large language models
Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference…
$\texttt{AMEND++}$: Benchmarking Eligibility Criteria Amendments in Clinical Trials
Clinical trial amendments frequently introduce delays, increased costs, and administrative burden, with eligibility criteria being the most…
How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs
Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations. These vectors, also…
DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English
More than 80% of the 1.6B English speakers do not use Standard American English (SAE), yet LLMs often fail to correctly identify non-SAE di…
Structurally Separated Uncertainty in Supervised Latent Variable Models
Predictive uncertainty is commonly decomposed into epistemic and aleatoric components, but standard decompositions often produce strongly c…
Calibrate Globally, Measure Everywhere: Scaling LLM-Based Prevalence Measurement Across A/B Experiments
Online media platforms track the share of impressions associated with content attributes, or prevalence, to evaluate trade-offs and set gua…
PatchDenoiser: Parameter-efficient multi-scale patch learning and fusion denoiser for Low-dose CT imaging
Low-dose CT images are essential for reducing radiation exposure in cancer screening, pediatric imaging, and longitudinal monitoring protoc…
Ask don't tell: Reducing sycophancy in large language models
Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an al…
The Rise of AI in Weather and Climate Information and its Impact on Global Inequality
AI development's current trajectory risks automating and amplifying the North-South divide in the global climate information system. Fronti…
Making Implicit Premises Explicit in Logical Understanding of Enthymemes
Real-world arguments in text and dialogues are normally enthymemes (i.e. some of their premises and/or claims are implicit). Natural langua…
MICA: Multi-granularity Intertemporal Credit Assignment for Long-Horizon Emotional Support Dialogue
Reinforcement learning (RL) for large language models (LLMs) has shown strong performance in single-turn tasks, but extending it to multi-t…
Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis. However, their clinical deployment is sever…
Gated Adaptation for Continual Learning in Human Activity Recognition
Wearable sensors in Internet of Things (IoT) ecosystems increasingly support applications such as remote health monitoring, elderly care, a…
State-Dependent Safety Failures in Multi-Turn Language Model Interaction
Safety alignment in large language models is typically evaluated under isolated queries, yet real-world use is inherently multi-turn. Altho…
Adaptively Robust LLM Monitoring via Activation Watermarking
Providers monitor deployed large language models (LLMs) to detect misuse that they cannot prevent. LLM monitoring is deterministic and ofte…
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate two capacities: how much a model knows…
GroupRAG: Cognitively Inspired Group-Aware Retrieval and Reasoning via Knowledge-Driven Problem Structuring
The performance of language models is commonly limited by insufficient knowledge and constrained reasoning. Prior approaches such as Retrie…
REAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production Usage
Production deployment of AI coding agents requires fast, reproducible evaluation signals. Existing industrial practices trade off speed and…
Shot-based quantum encoding: a data-loading paradigm for quantum neural networks
Efficient data loading remains a bottleneck for near-term quantum machine learning. Existing schemes (angle, amplitude, and basis encoding)…
The Fast Lane Hypothesis: Von Economo Neurons Implement a Biological Speed-Accuracy Tradeoff
von Economo neurons (VENs) are large bipolar projection neurons found exclusively in the anterior cingulate cortex (ACC) and frontal insula…
Facial-Expression-Aware Prompting for Empathetic LLM Tutoring
Large language models (LLMs) enable increasingly capable tutoring-style conversational agents, yet effective tutoring requires sensitivity…
BioHiCL: Hierarchical Multi-Label Contrastive Learning for Biomedical Retrieval with MeSH Labels
Effective biomedical information retrieval requires modeling domain semantics and hierarchical relationships among biomedical texts. Existi…
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making…
Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation
We propose \textbf{Compressed Video Aggregator} (CVA), a lightweight micro-video recommendation module that decouples video information fro…
Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval
Current LLM memory benchmarks evaluate answer quality rather than retrieval accuracy. Consequently, a system that dumps its entire belief s…
Skills on the Fly: Test-Time Adaptive Skill Synthesis for LLM Agents
Additional test-time compute can give LLM agents access to more past experience, yet expanding the context or adding rollouts does not nece…
Exact Symmetry as Algebra: A Machine-Verified Tensor Calculus that Enforces Physical Selection Rules
Symmetry is central to the physical sciences, yet machine learning usually captures it only approximately, leaving a residual per-step equi…
VistaHop: Benchmarking Long-Horizon Visual DeepSearch
Visual DeepSearch tasks require multimodal large language models (MLLMs) to resolve complex visual queries by repeatedly inspecting image r…
An Empirical Audit of Input Encoders for Multi-Channel Signal Transformers
Transformers consuming multi-channel scalar signals must embed $C$ simultaneous values into one $d_{\text{model}}$-dimensional vector per t…
The Score Hamiltonian: Mapping Diffusion Models to Adiabatic Transport
We exhibit an exact correspondence between sampling with score-based diffusion models and adiabatic transport of ground states for a family…
TLA-Prover: Verifiable TLA+ Specification Synthesis via Preference-Optimized Low-Rank Adaptation
TLA+ is a formal specification language for verifying distributed systems and safety-critical protocols. Large language models (LLMs) frequ…
Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent
Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central ch…
SafeECGMatch: Calibration-Aware Joint Frequency and Time Space Semi-Supervised Learning for Open-Set ECG Classification
Electrocardiogram (ECG) classification models often suffer from severe label scarcity, making semi-supervised learning (SSL) an attractive…
Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning
Spatial reasoning remains a persistent challenge for multimodal large language models (MLLMs). Existing approaches largely rely on large-sc…
Functional Cache Grafting: Robust and Rapid Code-Policy Synthesis for Embodied Agents
Code-writing large language models (CodeLLMs) generate executable code policies for embodied agents by translating natural language goals a…
Keyless Attention: Value-Space Routing and Value-Only Caching for Efficient Transformers
Transformer architectures form the foundation of modern natural language processing, making it crucial to address the efficiency and scalab…
RWGBench: Evaluating Scholarly Positioning in Related Work Generation
Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.…
MLVC: Multi-platform Learned Video Codec for Real-World Deployment
Neural video codecs have surpassed classical codecs in coding efficiency but remain impractical for deployment due to cross-platform incomp…
Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followe…
Prompt Framing Distorts Count Based Evaluation of LLM Error Detection: Evidence from Numeric Anchoring
Count-based F1 is widely used as a proxy for LLM error-detection quality, but this paper shows that it can rise dramatically without a corr…
Revealing Hidden Model Behaviors with Task-Specific Self-Reports
Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only wh…
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound…
G2VD: Generalizable AI-Generated Video Detection via Counterfactual Intervention and Causal Disentanglement
Rapid advances in AI video generation pose increasing security risks and call for reliable detectors with strong cross-domain generalizatio…
VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection
Deepfake image detection is served by three fundamentally different paradigms - commercial APIs, zero-shot vision-language models (LLMs), a…
Introducing Human-Centeredness in AI-Assisted Lexicography
This paper proposes a human-centered artificial intelligence (HCAI) framework for AI-assisted lexicography. While generative AI offers sign…
Heterogeneous Element-Aware Cross-Version Differencing of Scientific Documents via Layout-Aware Alignment and Structure-Aware Reasoning
Cross-version differencing of scientific documents is essential in scholarly publishing and technical documentation, but remains challengin…
Fantastic Adaptive Taxonomies and How to Use Them
An agent system's execution traces record how it fails, and procedures that improve such a system without changing model weights (trajector…
Auto Research for Materials: Auditable AI-Scientist Workflows with Held-Out Transfer
Auto Research uses language-model agents to propose, implement, and evaluate machine-learning changes in a closed loop, but is usually judg…
Towards an Automated Test of LLM Security Knowledge
Large language models (LLMs) are increasingly used for a range of software, hardware and human-centered security tasks. Consequently, LLM p…
REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning
Large-scale online reinforcement learning (RL) is the predominant means of eliciting advanced abilities including long-term reasoning and a…
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
In this report, we present a unified song generation framework capable of producing high-quality full-length music from lyrics, text descri…
On the Depth Scalability of Logic Gate Networks
Logic Gate Networks (LGNs) compute through compositions of Boolean operations, yet existing LGNs do not reliably benefit from increased dep…
Measuring the Dependency Gap: Diagnosing Inter-Column Fidelity in Tabular Generative Models
Synthetic tabular data is prized for preserving not just each column's marginal distribution but the dependencies between columns - structu…
PerplexityがAIエージェントの“暴走”対策ツールをオープンソースに Claude CodeやCodexを監視
PerplexityがAIエージェントの危険な挙動を検知・防止するツール群「Numbat」をオープンソース化。Claude CodeやCodexに組み込み、タスクに執着したエージェントの“暴走”を実行前に阻止する。
Thinking Machines、軽量モデル「Inkling-Small」正式公開 サイズ4分の1で「Inkling」に匹敵する性能
Thinking Machines Labは、オープンウェイトのAIモデル「Inkling-Small」の正式版を公開した。従来モデルの4分の1のサイズながら、データの改良や強化学習によりコード生成などのベンチマークで従来版を上回る性能を実現。動作に必要なGPUメモリも大幅に削…
Chromeに13年以上潜んでいた脆弱性、AIで発見 直近2回のアプデで過去23回分を上回るバグ修正
GoogleがChromeのセキュリティ対策へのAI活用を公式ブログで解説。Geminiベースのエージェントが13年以上潜んでいた脆弱性を発見した。AI攻撃の高速化に対応し、セキュリティ更新「週2回」配信も試行する。
スクエニ、ゲームの品質テストをGeminiで自動化 AIが画面を見ながらコントローラーを操作、検証作業を自走
スクウェア・エニックスが、ゲームのQAテストを「Gemini」で自動化する取り組みを「Google Cloud Next Tokyo '26」基調講演で披露。AIが画面を見ながらコントローラーを操作し、検証作業を自ら進める。
Google、ロボット向けAI「Gemini Robotics 2」発表 ヒューマノイドの全身制御や指先作業を実現
GoogleとGoogle DeepMindは、ロボット向けAIモデル群「Gemini Robotics 2」を発表した。全身制御や指先での微細な作業、複数ロボットの連携に対応する。高次の脳として機能する推論モデル「ER 2」や軽量VLAモデルを含み、安全性評価の新ベンチマーク…
Anthropic says its own AI models breached three companies during security tests
After OpenAI's models broke into Hugging Face, Anthropic checked its own history and found three similar incidents
Claudeが評価環境から実在企業に不正アクセス――Anthropic、3件のインシデントを公表
Anthropicは、サイバーセキュリティ評価中にAIモデル「Claude」が設定ミスでオープンになっていた経路から外部のインターネットに接続し、実在する3組織の本番インフラに誤って不正アクセスしていたと発表した。評価環境を演習と誤認したことが原因で、全評価を停止し外部ベンダー…
OpenAI、「GPT-5.6 Luna」を80%値下げ モデル自身による効率化でコスト削減
OpenAIは、「GPT-5.6」ファミリーの「Luna」を80%、「Terra」を20%値下げすると発表した。API価格の改定に加え、「ChatGPT Work」や「Codex」でのクレジット消費量も削減される。自律的なカーネル最適化により提供コストの削減を実現した。また、処…
AI hedge fund Situational Awareness may have sold its public portfolio, but it still has its Anthropic shares
The former OpenAI researcher’s fund was forced to unwind public equities after leveraged public bets plummeted. But he still has cards to p…
Reddit reports a solid quarter but shows signs of AI’s impact
Reddit's financial situation is looking good but uncertainty about its relationship to Google and the new AI-ified web are stirring market…
リホストを選ぶ企業は63% モダナイゼーションの「第一歩」のはずが、なぜ終着点に変わるのか?
国産メインフレームの提供終了が相次ぎ、レガシーシステムの移行は「期限のある経営課題」になった。だが移行プロジェクトで最初に立ちはだかるのは、移行技術やアーキテクチャ「以外の」問題だ。モダナイゼーションがリホストで止まる構造を、ITRの入谷光浩氏が解説する。
Investors love AI, as long as you’re a cloud host
Amazon isn't slowing down on data center spending — but investors don't seem to mind.
「アイデア出しを抜いた」 生成AIなしで最も不安になる業務といえば?
サイバーセキュリティクラウドが実施した調査で、上司よりも生成AIを参考にした経験を持つ人が半数に上るなど、職場でのAI依存が進んでいる実態が明らかになった。
日立が「SI全工程」をAI化 仕様確定で「最大240倍」効率化のワケ
日立は、エンジニア不足やシステム複雑化に対応するため、SIの全工程にAIを全面適用する「Agentic AI Integration Platform」を開発した。最先端AIと独自ノウハウを融合し、社内検証では画面仕様確定で最大240倍などの効率化を実証。顧客固有の暗黙知を蓄積…
悪用厳禁、「ChatGPT」の会話履歴をごっそりとぶっこ抜く“AIハック”:890th Lap
「ChatGPT」との会話履歴をPCへ簡単に保存できるとして、ある方法が注目を集めている。一方で、利用規約や情報管理の面で注意すべき点もある。
Judge says Trump admin still lacks evidence for Anthropic ‘supply-chain risk’ label
A federal judge said the Trump administration has not presented enough evidence to justify labeling Anthropic a supply-chain risk, casting…
Friend, the lonely AI wearable, returns with a new voice and a much bigger price tag
Friend, the AI wearable, can now talk to its users — for an enhanced price.
Google says it fixed more Chrome bugs in June than over the past two years, thanks to AI
As experts have warned for the last two years, some companies — like Microsoft and now Google — are finding and patching an exponential num…
LinkedIn adds a button to report AI-generated ‘slop’
LinkedIn is introducing new ways to reduce low-quality AI-generated posts, including a “seems like AI slop” reporting option. It's also rep…
Okta buys AI security startup Permiso — source says for about $200M
The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and other non-human identities across cl…
Meta says AI is making it easier to build new apps — and more are coming
Meta says AI is making it dramatically easier to build and launch new consumer apps, with CEO Mark Zuckerberg telling investors the company…
Nscale buys Anyscale as it seeks to own more of the AI compute stack
British AI neocloud Nscale is buying software startup Anyscale, which helps companies scale their AI workloads across data centers and serv…
Gemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool…
Forward-deployed engineers are the AI industry’s latest talent obsession
A new study estimates only 2,000 U.S. engineers have the expertise to deliver meaningful AI ROI, as enterprises race to hire forward-deploy…
2026-07-30(22件)
In the Hugging Face breach, OpenAI’s hacker was noisy and fast — but not unstoppable
Cybersecurity experts told TechCrunch that one of the biggest lessons to be taken from the OpenAI hack against Hugging Face has nothing to…
TechCrunch Disrupt 2026’s biggest stage features leaders from Amazon, Replit, Tether, with much more to come
The Disrupt Stage is where many of the biggest conversations in technology happen, with a legacy that stretches back for more than a decade.
Dili raises $21.7M to bring AI compliance to the infrastructure boom
The Series A was led by Khosla Ventures, with participation from Allianz, Rebel Fund, Brick and Mortar Ventures’ Darren Bechtel, and Y Comb…
Advancing the price-performance frontier with GPT-5.6
Explore lower GPT‑5.6 pricing for Luna and Terra—and how OpenAI’s more efficient models help enterprises deploy AI workflows at scale.
KADOKAWAとはてな、AIで小説執筆を支援 新サービス「RIKU」、テスター募集開始
KADOKAWAとはてなは、AIで小説の執筆を支援するエディタ「RIKU」を発表した。
日本HPのPCが楽天のAIを搭載、ローカル実行も可能 HP岡戸社長「ハイブリッドAIの重要なマイルストーン」
日本HPと楽天が、HP製PC向けAIアプリ「Rakuten AI for Desktop」のプリバンドルを開始。70億パラメータの日本語LLM「Rakuten AI 7B」で、オフラインでも要約や翻訳をローカルで実行できる。
Excel作業を自動化する「Copilot in Excel」がスキルに対応 何ができる?
Microsoftは「Copilot in Excel」の財務部門向け機能を強化した。Microsoftの財務部門が実運用で利用・評価し、財務業務で求められる信頼性を重視して開発された。
「データ品質に問題あり」から「予測精度95%」へ Umiosは販売計画をどう自動化した?
Umios(旧マルハニチロ)は、年間約4200時間を費やしていた販売計画作成を自動化した。全国の支社で入力の運用ルールがバラバラといった「データの品質」問題をどう解消し、予測精度95%を実現したのか。
Microsoft is openly competing with OpenAI, Anthropic more than ever
Microsoft pitched its own homegrown AI models, harnesses, and even a Mythos competitor on Wednesday, telling Wall Street it plans for conti…
営業製作所、図面管理システム「ジーエン図面」の販路拡大へSB C&Sと契約
営業製作所は、図面管理システム「ジーエン図面」についてSB C&Sとディストリビューター契約を締結した。SB C&Sの全国規模の法人向け販売ネットワークを活用し、販路拡大と導入企業数の増加を図る。
Mark Zuckerberg predicts that billions of people will have personal AI agents in five years
As Meta pours billions into AI infrastructure and agents, Zuckerberg is working to convince investors that the payoff will be worth the pri…
フィジカルAI時代のロボティクス新標準、安全性は「後付け」でなく「設計の核心」
AIがデジタル空間を超えて物理世界に踏み出す「フィジカルAI」の時代に入り、ロボットを開発する上での「安全性」をどのように定義し直すべきかが問われている。
Microsoft logs $3.2B from Anthropic investment, but OpenAI was a mixed bag
When Microsoft reported killer fourth-quarter earnings for its fiscal 2026 year (which ended June 30), it tucked in an interesting little t…
Zuckerberg says Meta’s enterprise AI opportunity extends beyond agents
On the company’s second-quarter earnings call Wednesday, CEO Mark Zuckerberg said Meta sees a “large enterprise opportunity” spanning AI ag…
AI・半導体企業トップが語る“稼ぎ頭” キオクシア、フジクラ、東京エレデバの見解まとめ【無料PDF】
乱高下するAI・半導体市場の今後はどうなるか? 注目企業の経営幹部が見通しを語った注目記事をPDFにまとめてお届けする。
Discover what’s next for AI, from the SaaS reckoning to the agent security gap, at TechCrunch Disrupt 2026
At TechCrunch Disrupt 2026, the AI Stage is back to dig into the single hottest topic in the community for the past few years, presented by…
Thinking Machines co-founder Lilian Weng left the company citing health reasons, then joined OpenAI
Weng previously served as the VP of AI Safety Research at OpenAI.
The Hugging Face AI break-in, as told through an increasingly committed bear metaphor
Another way to think about the whole thing is to picture a bear at a campsite. (Really, we are going there.)
Claude Opus 5 became downright ruthless when tasked with running a vending machine
Andon Labs' latest vending machine simulation shows Opus 5 lied and colluded its way to become the best AI capitalist ever.
Hint, a new AI startup co-founded by Martha Stewart, offers an AI assistant for homeowners
AI home management startup Hint, co-founded by Martha Stewart, wants to become an “AI for your home,” combining property records, maintenan…
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
How two API settings improved GPT-5.6 performance on ARC-AGI-3, boosting scores and efficiency by retaining reasoning and enabling compacti…
2026-07-29(348件)
Encore AI raises $30M to build AI agents that learn from customer calls
The startup analyzes calls, messages, and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
As AI content floods the internet, Pangram raises $9M to detect it
Pangram has raised $9 million to scale its AI detection software. The startup has also released a new AI text detection model, Pangram 4, a…
Accelerating scientific discovery with ChatGPT for Academic Researchers
OpenAI is giving 100,000 academic researchers free access to ChatGPT's most advanced AI models to accelerate scientific research, collabora…
ChatGPT WorkとCodexの5時間制限「明日から再開」 GPT-5.6 Solの“トークン消費問題”を改善
米OpenAI幹部のティボ・ソティオ氏は、デスクトップPC向けAIサービス「ChatGPT Work」とAIコーディングツール「Codex」について「明日から5時間ごとの利用制限枠を再開する」と発表した。
PFN「国産AI」で自衛隊を支援へ 防衛の作戦立案に利用 防衛装備庁の実証実験を受託
Preferred Networksは、生成AIで自衛隊を支援するシステムを開発すると発表した。
Do Models Fake Alignment Without Clear Consequences?
Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather t…
Beyond Memory: A Templated Substrate for Heterogeneous Collaborative Knowledge Work with LLM Agents
Research projects, educational efforts, and adjacent knowledge work accumulate findings, decisions, and reasoning that future collaborators…
Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels
Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernel…
CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models
Masked diffusion language models (MDLMs) are advancing rapidly, yet the evaluation standards needed to reliably interpret their progress ha…
GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models
The rapid growth of online grocery shopping requires recommendation systems that capture cyclical purchasing behavior and diverse user inte…
Crystalis: Progressive Nucleation and Semantic Annealing for Coordinated Multi-View Visualization Generation
Large language models (LLMs) can generate individual charts, but coordinated multi-view visualizations (CMVs), where views share data flows…
PATHFinder Agent for Tailored Prenatal Care
Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricia…
LLM Scheming Inversely Scales with Pretraining Language Coverage
With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent…
ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop
Procedural tasks such as furniture assembly and home repair impose substantial cognitive demands because users must interpret instructions,…
RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection
Knowledge injection updates pretrained MLLMs with new factual or domain-specific knowledge, but fitting full authoritative answers can caus…
RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation
Geoscience research requires complex analysis and domain expertise, with remote sensing (RS) observations as a key foundation. However, exi…
Right-sizing Recommendations (RSR): Cloud Workload Conformal Prediction for Virtual Machines in Data Center Operations
Managing cloud infrastructure efficiently, especially in environments of large cloud providers or hyperscalers, requires optimizing the use…
Atmospheric Diffusion-Guided Spatio-Temporal Transformer for Nuclear Radiation Forecasting
Nuclear radiation, the energy released during atomic decay, poses persistent risks to public health and the environment, and concerns have…
Steering topology distributions for unified generative design of architected metamaterials
Architected metamaterials derive their functions from structure, creating vast opportunities to program physical responses through topology…
HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
Online advertising bidding systems typically deploy multiple offline-trained expert models (e.g., PID controllers, model predictive control…
LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation
Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish…
Personalization, Personas, and Forecasting in Value Alignment
LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast…
Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn
Job understanding is critical to LinkedIn's mission of connecting talent with opportunity. This task involves transforming unstructured and…
On the Use of LLMs for Specialised Terminology: A Good Alternative to Corpora?
Specialised translation relies on the use of documentary and terminological resources, including corpora. These resources are particularly…
SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models
Sparse Mixture-of-Experts (MoE) models expand foundation model capacity through conditional expert activation, but their full expert pools…
GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference
As Large Language Models scale to increasingly long contexts, the memory I/O and computational overhead of the Key-Value (KV) cache during…
A GAN-Based Framework for Robust Data Synthesis in Satellite Internet Observations
Low-Earth orbit (LEO) satellite Internet has become an important infrastructure for enabling ubiquitous connectivity to align with the Inte…
Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows l…
When Shortest Isn't Safest: A Design Science Approach to Senior-Friendly Pedestrian Routing
Older adults' independent mobility enables out-of-home participation, well-being and health, yet pedestrian navigation systems still optimi…
RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
Vision-language models (VLMs) have achieved strong performance on general remote sensing tasks. However, their capability for rare scenes r…
Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings
Access to specialist clinical expertise remains severely limited across sub-Saharan Africa, where physician-to-patient ratios can fall belo…
AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning
Reinforcement learning with verifiable rewards is a powerful paradigm for eliciting reasoning in large language models, yet it suffers from…
MusiChat: Vibe Composing for Music Creation
Recent advances in AI music generation have enabled users to create complete musical pieces from natural-language prompts. However, most ex…
Understanding Semantic IDs: From Item Representation to Item Selection in Generative Recommendation
Semantic IDs (SIDs) are now a central component of generative recommendation. Current SID-based systems assign three roles to the same toke…
Localized Anomaly Detection via Differentiable D-vine Copulas
Vine copulas provide a flexible framework for modeling complex multivariate distributions through a hierarchical decomposition into bivaria…
Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization
Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the…
SAFAARI: Schema-Aware Framework for Accelerated Advertiser Response Intelligence
The evolution of customer support systems is rapidly advancing with agentic chatbots, yet these systems face significant limitations when a…
CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification
Electroencephalography (EEG) analysis in cognitive studies requires specialized expertise and involves many defensible choices over contras…
Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being
As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing att…
Similar Models Learn Differently: Final-Window Pretraining Shapes Post-Training Beyond SFT
Developers judge a model checkpoint by how it behaves. After supervised fine-tuning (SFT), two checkpoints that perform about the same acro…
Addressable Recall Compaction for Long Context-Window Control in AI Agents
Long-horizon LLM agents accumulate reasoning traces, actions, and tool observations that can eventually exceed a model's fixed context wind…
How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness
Routing decisions between a cheap heuristic and an expensive large language model (LLM) are typically framed as a difficulty problem: send…
Towards an Agent Operating System - Lessons from Classical and Cloud OS
Every major wave of platform software follows the same arc: an initial period of experimentation with competing frameworks and ad-hoc imple…
PLATO: Pointer Learner for Agent and Task Openness
Open agent systems (OASYS) are increasingly prevalent in real-world domains where the sets of agents and tasks change unpredictably over ti…
Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering
Machine learning engineering (MLE) tasks require long-horizon decision making over iterative solution debugging and refinement, under expen…
Towards Robust Reinforcement Learning for Small-Scale Language Model Agents
The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, th…
ScalableRAG: High-Quality RAG at Zero Ingestion Cost
Recent advances in RAG aim to optimize for performance by paying high ingestion costs for knowledge ingestion: building knowledge graphs or…
Less Data, Better Alignment: Data-Centric Multi-Evaluator Agreement for Preference Optimization
Research on preference optimization often varies the training objective while holding the data fixed. We instead ask whether a small, high-…
How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation
This paper studies the behavior of language models in a multi-agent crowd simulation, focusing on how affect propagates among agents that p…
Inferring Missing Trajectory Data with Temporal Convolutional Networks
Trajectory data collected in real-world settings is frequently incomplete due to sensor failure, communication loss, or occlusion. We addre…
When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops
Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, sel…
PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
Discrete masked diffusion language models support bidirectional generation and infilling, but adapting pretrained autoregressive (AR) trans…
Observing sycophantic AI validate others reduces its appeal but not its persuasiveness
AI chatbots can be ``sycophantic,'' or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, y…
Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection
Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as…
CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models
Foundation models for 12-lead electrocardiograms (ECGs) transfer well across clinical tasks, but the physiological knowledge encoded in the…
The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape
Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking sho…
Many-body Tipping Dynamics of ChatGPT-like AIs
Why do ChatGPT-like AIs, despite major architectural and training differences, unexpectedly tip to undesirable content (e.g. harmful, misle…
ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design
This paper presents ContractHIL-HLS, a contract-aligned multi-agent workflow for practical high-level synthesis (HLS) engineering. The work…
Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe
Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the per…
Physics-Grounded Fluid Video Generation with a Simulation Dataset and Dual-Stream Optical-Flow Supervision
Video diffusion models generate visually compelling content but routinely violate elementary physics when the subject involves fluids: liqu…
From Cellular Responses to Pharmacological Domains: Multimodal Zero-Shot Drug Representation Learning
Multimodal drug discovery enables drug representation learning beyond chemical structure by incorporating cellular responses such as gene e…
Dual-Domain Manifold Modeling for Hyperspectral Image Fusion
Achieving a coherent integration of spectral richness and spatial fidelity remains a central objective in hyperspectral image fusion. Howev…
Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management
The same episode of atrial fibrillation is a minor finding in a healthy adult and grounds for anticoagulation in an elderly patient with hy…
Explanation-Bound Tool Execution for AI Agents: Server-Verified Action Claims Without Trusting Model Rationales
Tool-using agents expose structured calls but commonly attach free-form rationales. Such rationales are neither authorization nor reliable…
AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis
The intersection of artificial intelligence adoption, cybersecurity governance, and public sector institutional constraints has not been ex…
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning
Agentic systems have rapidly advanced in their ability to interact with real-world environments, leverage external tools, and provide servi…
Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response
Cyber-capable AI agents combine language models with tools, memory, and execution en- vironments to perform multi-step offensive-security t…
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed…
COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution
Large language model (LLM) agents are increasingly entrusted with natural-language workflow instructions (e.g., retail-payment policies) th…
Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents
A growing body of 2026 work applies control theory to LLM agents: Lyapunov-certified stability for tool-mediated controllers (Prinos et al.…
A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain
Production LLM agents are increasingly assembled from a frozen model wrapped in a harness: a prompt template, a tool set, a memory/retrieva…
Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering
Knowledge-intensive multimodal question answering (KI-MMQA) sits at the intersection of three expensive primitives: long visual token seque…
The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play
Capture the Flag (CTF) competitions are among cybersecurity's most effective training grounds, developing practical skill across cryptograp…
Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm
Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (orga…
TRWH: A Text-Driven Random Walk Heterogeneous GNN for Semantic-Aware Sparse Recommendation
Graph Neural Networks (GNNs) and Large Language Models (LLMs) have each advanced recommendation systems by modeling structural and semantic…
Balancing multiscale similarity and cartographic constraints: A similarity-driven optimization framework for line generalization
Cartographic generalization is essential for generating multiscale map representations by balancing information preservation and cartograph…
Finding Optimal Cost-Bounded Plan Reductions: Refined Model
In some real applications a plan may later become unfeasible due to newly imposed budget constraints, yet, at the same time, using only the…
PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents
Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on thei…
CoTinyVLA: Chain-of-Thought Distillation for a Sub-Billion-Parameter Vision-Language-Action Model
Vision-Language-Action (VLA) models translate natural-language commands into robot action sequences, but leading systems on the LIBERO-Plus…
Are the High-weight Neurons the Important Ones in Image Classification Neural Networks?
As neural network models for image classification advance, neurons play critical roles in pruning, backdoor defense, and interpretability.…
Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners
Consider a model trained at a single hospital to predict patient recovery, where the measured feature $X$ bundles the patient's true health…
From Training to Deployment: Post-Hoc Causal Feature Identification via Sensitivity Ratios
Given a model that is already trained, which features does it rely on causally versus spuriously? Existing methods require access to the tr…
Distilling Temporal Search and Reasoning: Evolving LLMs for Future Prediction via Harness-Assisted Efficient Data Synthesis
Future event prediction carries broad social impact yet remains challenging. SOTA approaches augment LLMs with external agent frameworks wh…
Agent Skills Matter: Inferring Proprietary Skills from Execution Trajectories
Agent skills package reusable procedures that improve downstream performance. Their lightweight, portable form enables marketplace monetiza…
Matrix-Free Photoacoustic Image Reconstruction via Sensor-Token Self-Attention
Photoacoustic tomography (PAT) combines the optical absorption contrast of biological tissue with the spatial resolution of ultrasound, yet…
How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model
Parameter-efficient fine-tuning (PEFT) and low-bit quantization are now standard tools for adapting language models under tight compute bud…
A Density-Matrix Framework for Electronic-Structure Analysis of Functional-Group and Salt Effects in Lithium-Metal Electrolytes
The reactivity of lithium-metal electrolytes arises from the interplay of molecular functional groups, Li$^+$ solvation, and salt-anion par…
Computational Extraction of Legal Causes via al-Sabr wa al-Taqsim: A Set-Theoretic Formalization for Closed Fiqh Chapters
This paper presents a set-theoretic formalization of the classical usuli method of al-Sabr wa al-Taqsim (Examination and Division) for extr…
Multi-Sensor Alignment for Weather Simulations
Perception tasks for autonomous vehicles need to work satisfactorily in adverse weather conditions. Due to lack of real-world weather datas…
Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines
Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epi…
Quotient Dynamics, Effective Curvature, and Implicit Bias in Positive Quadratic Networks
Positive quadratic networks admit the low-rank representation f_U(x)=x^top UU^top x, where Uinmathbb{R}^{dtimes r} is identifiable only up…
Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception
Decoding speech information directly from scalp electroencephalography (EEG) into text provides a potential non-invasive neural communicati…
AIriskEval-edu Demo: Auditing of Pedagogical Risks in Educational Explanations
We present AIriskEval-edu Demo, a platform that audits the pedagogical quality of instructional explanations and provides explainable audit…
Engine-Equal, Human-Unequal: A Reproducible Outcome Skew in Engine-Assessed Equal Chess Positions
Among chess opening positions that a strong engine judges essentially equal (Stockfish 18 evaluation within 10 centipawns of zero, depth-st…
OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation
Complex tasks often decompose into parallelizable yet interdependent subtasks, making orchestration critical to the performance of multi-ag…
CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization
Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-sty…
Localized Adaptation Reveals Distinct Learning Signatures in Transformers
Transformer adaptation is typically distributed across model depth, even when the intended change is narrow. We investigate how adaptation…
OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs
Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video tok…
DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space
Text-space optimization adapts large language models (LLMs) by editing external natural-language artifacts rather than model weights, so th…
Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare
Cognitive distortion amplifies negative emotions and contributes to mental health disorders. Cognitive Behavioral Therapy (CBT) is an effec…
Nudging Sustainable Choices through LLM-Generated Recommendation Explanations
Recommender systems mediate everyday consumption, offering a promising channel for encouraging sustainable choices. Prior research shows th…
Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography
Objective: Concept bottleneck models route prediction through interpretable intermediate variables, and their validity is normally judged b…
Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL
Large language model agents often spend substantial wall-clock time waiting for tool call results. Tool-call speculation can hide this late…
Distributed Constraint Optimization via Online Learning and Iterative Pricing with Application to Large-Scale Satellite Scheduling
Distributed constraint optimization problems (DCOPs) provide a popular framework for distributed decision making under limited communicatio…
HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs
Skills have become an important abstraction for enabling large language model (LLM) agents to reuse past experience in long-horizon interac…
Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks
This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a pa…
Distributing Security Controls Through Harness Engineering
AI coding agents are being adopted at historic speed, yet security and risk concerns remain the primary barrier to scaling agentic AI acros…
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation
Evaluating AI agents in interactive environments is hindered by fragmented tasks, scaffolds, verifiers, and scoring rules. Existing efforts…
Interactive Reward Agent: GUI Task Evaluation via Environment-State Verification
Graphical user interface task evaluation aims to determine whether a GUI agent has successfully completed a user instruction. Automated GUI…
Toward Standardized Cross-Vendor Agent Tool Trust Management in Autonomous Networks
Autonomous Network Levels 4-5 require AI agents to invoke tools across vendor boundaries without human oversight, yet existing management s…
Penelope: Localized Latent Recurrence for Efficient Structured Reasoning
Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing paramet…
dtControl2+$\varepsilon$: Trading Optimality for Explainability in MDPs via Decision Trees
Over the past decade, decision trees have been used to represent controllers (a.k.a. policies) in an explainable way, with dtControl2 as a…
A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series
Question answering (QA) over irregular clinical time series (ICTS) plays a pivotal role in a wide range of healthcare applications. Althoug…
Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation
Multi-warehouse inventory allocation is typically formulated as a mixed-integer programming (MIP) problem, yet no single formulation consis…
CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer
Graph foundation models (GFMs) have emerged as a promising paradigm for transferring knowledge across graph domains and tasks. Real-world g…
Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment
Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development…
Desktop-Delta Bench: Do Computer-Use Models Understand Desktop GUI Transitions?
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-t…
Untrusted Authors, Trusted Answers: A Calculus of Fidelity-Graded Translations
To answer a question about a program, move the program to where the question is decidable. Every such move is a translation, and every tran…
Domain-Prior-Regularized Graph Modeling for Anomaly Detection in Cyber-Physical Systems
Anomaly detection on multivariate sensor time series is critical for industrial monitoring of cyber-physical systems (CPS), where even subt…
Neural Network Learning of One-Bit Protocols for Qubit Measurement Simulation
Communication complexity provides a natural framework for quantifying the classical resources required to reproduce quantum statistics. In…
DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation
Key Information Extraction (KIE) is vital for many document applications, but creating training datasets is traditionally a time-consuming…
VLD-RAG: Agentic Vision-Language Retrieval-Augmented Generation for Long, Visually-Rich Multi-Page Documents
Visually-rich documents such as reports, slides, and manuals often distribute the evidence needed to answer a question across multiple page…
Game AI Not Fun? A Scoping Review and Meta-Analysis on the Differences in Enjoyment between Human and Computer Opponents
Although advancements in game character AI aim to enhance player engagement, evidence suggests that perceiving an opponent as artificial ca…
CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs
Large language models (LLMs) are increasingly used to provide mental health support, requiring reliable evaluation of safety, empathy, and…
Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course
This full research paper examines how different forms of learner-AI interaction relate to learning outcomes in object-oriented programming…
What Gets Lost When Memory Becomes Media? Evaluating AI-Generated Oral History Visualization
What gets lost when memory becomes media? Diaspora oral-history interviews require a double transformation; first-person recollection to th…
From Idea to Classroom in Days: Using "Vibe Coding" to Create a Programming Process Visualizer from IDE Activity Logs
This paper reports on the rapid development and classroom deployment of a Thonny log visualizer built using AI-assisted ``vibe coding'' to…
Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governance in Everyday Human-Chatbot Interaction
Research on human-AI interaction has long framed verification of system outputs as a trust-contingent behavior that better-calibrated trust…
Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement
Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own pri…
The Effect of Text Chunk Size on Retrieval-Augmented Generation Performance
Retrieval-Augmented Generation (RAG) systems have emerged as a powerful process for allowing large language models (LLMs) to retrieve relev…
Three Sides of Retrieval: Factorial Evidence for Document-Side, Query-Side, and Answer-Side Complementarity in RAG
RAG systems rely on chunking, which destroys structural information in documents. Existing heading-based retrieval (Jeong et al., 2025) req…
DDSNet: Dual-domain Symmetry-aware Network for PCSEL Property Prediction
Efficient exploration of the photonic crystal (PhC) lattice design space is essential for developing photonic crystal surface-emitting lase…
Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
Weak supervision sets a practical regime for audio-visual sound source localization as dense spatial annotations are costly to obtain at sc…
From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance
Retrieval-augmented generation (RAG) is the dominant paradigm for applying large language models (LLMs) to enterprise document corpora, yet…
AI-Assisted Knowledge Access for Legacy Enterprise Asset Management in Energy Operations: A Practical Retrieval System
Energy utilities still run engineering work management, engineering procurement, and inventory processes on long-lived enterprise asset man…
Selective Impairment of Motor Recovery from Typing Errors in Parkinson's Disease: A Survival Analysis
Parkinson's disease (PD) affects multiple, dissociable stages of motor and cognitive control. We ask whether passively-collected keystroke…
Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code
In the literate human brain, reading and writing are two doubly-dissociable systems: a ventral decoding route (impaired in pure alexia) and…
Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source SLMs
Large Language Models tend to hallucinate when answering domain-specific ques tions from scientific documents without prior fine-tuning. Cu…
When Thinking Before Retrieval Hurts: TraceBound Diagnostics for Adaptive Knowledge-Graph Retrieval
Adaptive retrieval promises to make knowledge-graph question answering more robust by letting a controller search, inspect neighborhoods, r…
Decoding Error-Related Potentials under Multisensory Feedback with Varying Congruency
Error-related potentials (ErrPs) are widely studied neural signatures associated with error processing in human-machine interaction. In rea…
A Path Integral Model of Cognition
We develop the mathematical and physical formulation of cognitive cost optimization that underlies the path-integral model of consciousness…
EEG Emotion Recognition From AI-Generated Biodigital Architecture Images
Emotional responses to biodigital architecture were examined using electroencephalographic (EEG) data from AI-generated images. A pre-exper…
Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture
Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situa…
Dual-Level Atomic and Coordination Geometry Learning for Crystal Property Prediction Using Graph Neural Networks
Accurate prediction of crystal properties remains a key challenge in computational materials science. While graph neural networks (GNNs) su…
Dynamic Multi-Criteria Bottleneck Severity Index (DMBSI) for Semiconductor Wafer Manufacturing: A Genetically Optimised Framework for Reentrant Production Systems
Wafer fabrication exhibits unique characteristics, including reentrant process flows, variable bottlenecks, and highly variable process con…
Foundation Models for EEG Are Blind to Long-Range Temporal Correlations: A Spectral-Temporal Dissociation Behind Their Cross-Population Fragility
Objective. Electroencephalography (EEG) foundation models (FMs) are trained to reconstruct or contrastively align short patches, then poole…
MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA
In medical multiple-choice question answering (MCQA), Retrieval-Augmented Generation (RAG) can supplement the domain knowledge of language…
REPREC: Representation Driven Parameter-Efficient Recommendation System
Large language models (LLMs) have been applied to sequential recommendation by formulating it as a natural language task. Previous work has…
Two Views, One Voice: Evidence-Grounded Conversational Music Recommendation
Traditional conversational recommenders entangle retrieval and response generation within a single text interface, so exact entity cues fad…
Extremal Chowla sets and their linear analogues: A human-AI mathematical investigation using Co-Scientist
We introduce an extremal invariant associated with Chowla-type order conditions in finite groups. A nonempty subset $S$ of a finite group $…
Beyond Predictive Accuracy: A Reliability-Aware Audit of Molecular Representations for Human Olfaction
Pretrained molecular encoders are commonly evaluated through downstream prediction, but predictive accuracy alone does not establish that a…
DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization
Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency r…
HVM-GraphRAG: Holistic-View Multimodal Graph Retrieval-Augmented Generation on Complex Document
Question answering (QA) over complex documents requires models to retrieve and integrate evidence distributed across distant document regio…
Tokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in recommendation systems
Large-scale recommendation systems face "Memory Wall" bottlenecks due to massive, dense embedding tables. While generative retrieval uses d…
GraphRareBench: An Auditable Graph-Evidence Benchmark for Phenotype-Driven Rare-Disease Diagnosis
Phenotype-driven diagnostic benchmarks usually report the rank of the reference disease, but they rarely reveal which plausible alternative…
Human Preference aligned Tabular Similarity
Task-agnostic tabular embeddings are increasingly used for similarity search in real-world business systems such as Product Lifecycle Manag…
Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier c…
Beyond "What to Retrieve": Uncertainty in Retrieval-Augmented Code Generation
Repository-level code generation relies on heterogeneous evidence whose relevance, compatibility, and completeness are inherently uncertain…
Eliminating Propagation Delay: Attention-Based Spatial-Temporal Fusion Graph Convolution Network for Traffic Flow Prediction
Predicting traffic flow is crucial to optimizing transportation systems and improving urban mobility. Many graph convolution-based models h…
Mechanisms of Width Scaling in Normalized Residual Networks: The Effective Alignment Dimension
Existing theories of neural-network width characterize asymptotic limits, but provide limited guidance on whether an expansion direction id…
GAUGE: Grading Agent-Built Financial Models Without a Golden Answer
Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be chec…
LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models
Text-conditioned time-series forecasting predicts a series from both its numerical history and natural-language context, allowing forecasts…
Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study
Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: a poisoned tool hides encrypted fragments in…
Latent Stability Analysis of Malware Representations Under Feature-Space Perturbations
Static malware detectors are commonly evaluated using clean-sample metrics such as accuracy, F1, ROC AUC, and PR AUC. However, these metric…
Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
State-of-the-art toxicity detectors for text-to-image generation adopt a one-size-fits-all approach: a single universal model applying fixe…
Multiclass Classification without Labels via Posterior Simplex Geometry
In many classification problems, reliable instance-level labels are unavailable. However, it is often possible to construct weakly enriched…
Stable FP4 Training via Transposition-Invariant Block Quantization
Reducing training precision is a key lever for improving the e ciency of large language model (LLM) training, but pushing beyond FP8 to 4-b…
Generative Distributionally Robust Optimization
Generative models are increasingly adopted in distributionally robust optimization (DRO), but existing approaches trade off model compatibi…
Automatic Knowledge Graph Construction and Query for Earthquake Catalogs
In recent years, the number of events in earthquake catalogs has significantly increased due to the utilization of more effective deep lear…
Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews
Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond…
Calibrated Partial Resets: Preventing Policy Collapse in Continual Reinforcement Learning
Neural networks are hindered by accumulating dormant neurons and loss of expressivity throughout training, particularly in non-stationary d…
CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models
LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively…
Authoring Agent Skills: A Software-Engineering Approach
Agent Skills are an emerging way to extend large language model agents with reusable procedural knowledge that the agent loads on demand. A…
Grounded in Consensus, In Step With Emerging Science: A Consensus-Anchored Multi-Corpus Clinical Chatbot for Long COVID
Long COVID (LC) poses a challenge for clinical decision support because relevant evidence is distributed across sources with different upda…
Extended Reality as a Mediation Layer for Situated Human Control in Human-Robot Teaming
Extended Reality (XR) is increasingly used in human-robot interaction to communicate robot intent, planned motion, reachability, and state.…
Lantern: Conflict-Aware Gradient Blending for Physics-Guided Diffusion Models in Calorimeter Simulation
Monte Carlo simulation of calorimeter showers is a principal bottleneck for the High-Luminosity LHC, and diffusion models have emerged as f…
DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification
Automated verification of numerical claims is a challenging problem, as it requires both language understanding and quantitative reasoning.…
Spectral Truncation in Synthetic Control
Synthetic control (SC) matches a treated unit's pre-treatment trajectory to a weighted combination of donor units. We study Spectral SC, wh…
Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation
Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large…
OPERA: Offline Policy-guided Expert Routing and Adaptation for Universal Biomedical Image Analysis
Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distribution shifts across sc…
Analysis of the Shortcut Learning and Clever Hans Effect in CNN based ECG Image Classification
Deep learning models for ECG image classification may achieve high accuracy by exploiting non-physiological visual cues instead of ECG wave…
Learning from 53.6K Real-World Developer Edits of AI-Generated Code
Imperfections in AI-generated code require that software developers modify the generated code manually, or by re-prompting an AI programmin…
Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments
We implement an agentic AI workflow built around a large language model (LLM) agent for autonomous experiments with nitrogen-vacancy (NV) c…
OrganLens: Organ-Specific Representation Learning for CT Foundation Models
A CT examination captures multiple organs, but many biomedical questions concern abnormalities, prognosis, or longitudinal change in a spec…
CondPSE: A Polynomial-Filtered Structural Encoder with Conditional Modulation for Graphs
Message-passing graph neural networks are bounded by the 1-WL test and can miss topological structure that distinguishes non-isomorphic gra…
TabRank: Chain-of-Thought Distillation for Table Re-Rankers
The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval. Multi-stage retrieval s…
RIDGE: An Autonomous Framework for Validation and Method Discovery in LLM-Generated Option Pricing
Automated code generation is becoming an important tool in quantitative finance, where large language models can generate option pricing im…
VaLiDRec: Variable-Length LLM-Aligned Semantic IDs for Generative Recommendation
Generative recommendation commonly represents items using fixed-length semantic identifiers (SIDs) constructed through clustering and quant…
TopoGR: Revealing and Preserving Latent Structure of Semantic ID in Generative Recommendation
Semantic ID-based generative recommendation tokenizes each item into a sequence of discrete semantic IDs and predicts the next item by gene…
Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks
Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a m…
Structure-aware Relative Policy Optimization for Ranking
Ranking is a fundamental component of modern information access systems. Reinforcement learning (RL) provides a flexible framework for dire…
Where Steering Signals Come From: Activation Source Selection in Activation Steering
Activation steering controls language models by adding vectors or features to hidden states at inference time, but the upstream source of t…
Bridging Compute- and Data-Optimal Pretraining
Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a reg…
ScaleResfusion: Residual Rectified Flow based on Residual Vector Field
Real-world Image Restoration (Real-IR) aims to recover high-quality (HQ) images from complex and unknown degradations. Although recent diff…
CLBench-V: Evaluating Multimodal Context Learning from Grounding to Knowledge Acquisition
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent w…
Hybrid Analysis for Secure MCP Tool Use in LLM Agents
The rapid development of large language model (LLM) agents has enabled their broad adoption across diverse real-world tasks. To standardize…
Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning
Optimization over the Stiefel manifold plays a significant role in various machine learning tasks. Existing methods either use the retracti…
CAST: Game Solvers as Turn-Level Teachers for LLM Agents
Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement…
Balanced Soft mixture-of-expert model for Glaucoma Detection
Glaucoma is a group of eye diseases that damage the optic nerve, often caused by elevated intraocular pressure. It is a leading cause of ir…
Physics-Informed Neural Operator for Warm-Starting Background-Decomposed and Preconditioned PSFD: Enabling Scalable 3-D EUV Mask Simulation
We present a physics-informed neural operator (PINO) trained with pseudo-spectral frequency-domain (PSFD) equations for electromagnetic (EM…
Specula: Scaling formal specifications for autonomous model checking of system code
Specula is a push-button agentic system that generates high-quality formal specifications for large, complex system code and uses the speci…
Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors
Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It r…
Explainable AI for Chronic Kidney Disease Prediction Using Simulated Federated Learning
Chronic Kidney Disease (CKD), characterized by the gradual loss of kidney function, remains a significant public health challenge. Early de…
Data Quality Profiling at Scale with Progressive Sampling: A Benchmark for Data-Centric AI Pipelines
Data quality profiling -- computing missing-value rates, duplicate fractions, outlier densities, and functional-dependency violations -- is…
Raven: High-Recall Sequence Modeling with Sparse Memory Routing
Long-context recall in linear-time sequence models highlights a tradeoff in how they write to memory. State-based linear models, such as st…
Rethinking Likelihood distributions: Student's t Likelihood Boosts Bayesian Neural Network Performance
In Bayesian neural networks (BNNs), variational inference is a widely adopted framework for modeling uncertainty in a distributional way, w…
MARS: Multi-Agent Re-ranking for Repeat-Order Food Delivery Recommendation
Large language models (LLMs) are increasingly used in recommender systems, but it is often unclear how much performance can be obtained fro…
From Dyad to Triad: Eliciting XAI Requirements in Stroke Rehabilitation
Eliciting explainable AI (XAI) requirements from stroke survivors presents a methodological challenge with direct implications for the desi…
Emergent Latent-State Computation under Stochastic Volatility
Mechanistic interpretability has largely focused on language models and deterministic toy tasks. Much less is known about how sequence mode…
Seen, Said, or Forgotten? A Causal Audit of Visual KV Memory Across Dialog Turns
Stateful multimodal assistants encode an image once but may answer questions about it many turns later. Attention-guided visual-KV eviction…
Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definit…
Automated Numerical Stability Analysis of Deep Learning Operators
Finite-precision arithmetic unavoidably introduces numerical approximation errors. Numerical computations may use insufficient precision or…
Beyond Counts: A Distributional Robustness Margin For Pathology Foundation Models
Pathology foundation models are approaching clinical deployment, yet remain vulnerable to systematic non-biological variation across centre…
At-the-Roofline Sparse Tensor Contractions on Vector Processors for Transformer Inference
Fine-grained weight pruning and activation sparsification have emerged as effective approaches for reducing the compute and memory cost of…
I2VShield: An Efficient Proactive Defense Framework against DiT-based Image-to-Video Models
The rapid advancement of video generation models has led to the increasing misuse of image-to-video (I2V) models. Although substantial prog…
ReLATE: Reliability-Guided Evidence Fusion for Robust UAV--Satellite cross-view Geo-Localization
Unmanned aerial vehicle (UAV)-satellite cross-view geo-localization matches UAV images against satellite imagery and has achieved impressiv…
Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
Unifying visual understanding and generation in one model holds immense promise, but remains challenging and expensive due to heavy compute…
Multi-Scale Structural Features for Continual, Comprehensible Visual Recognition in a Developmental Learning Framework
Contemporary machine learning struggles to learn continually, reuse prior knowledge, and expose a comprehensible internal structure. A rece…
Visual prompt engineering for video models
In the age of foundation models, a model is only as good as its prompt. For this reason, prompt engineering has become an essential techniq…
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
Generative AI has rapidly expanded audio-visual forgery beyond human-centric deepfakes into general scenes. Existing AIGC detection methods…
CORF-GS: Real-Time Wireless Radiance Field Reconstruction via Coupled Optical-RF Gaussian Splatting
Recent advances in 3D Gaussian Splatting (3DGS)-based wireless radiance field (WRF) reconstruction provide an efficient solution for wirele…
The LAIA Dataset: Labelled Attention for Intelligent Automobiles
The development of autonomous vehicles (AVs) usually relies heavily on data-driven artificial intelligence (AI) models that require large v…
IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment
Entity alignment (EA) identifies entities across knowledge graphs (KGs) that refer to the same real-world object. Conventional EA methods m…
Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs
Retrieval-augmented generation improves knowledge-intensive question answering, but indiscriminate retrieval can introduce irrelevant evide…
Physics-Informed Broad Learning System: An Efficient Backpropagation-Free Framework for Solving Partial Differential Equations
Physics-informed neural networks (PINNs) have emerged as a powerful paradigm for solving partial differential equations (PDEs) by embedding…
Contrastive Representation Learning of Longitudinal Disease Trajectories on Temporal Graphs
Understanding disease trajectories from longitudinal clinical data remains challenging due to complex temporal dynamics and heterogeneous p…
A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries
Interdisciplinary research is accelerating, yet scientific papers remain difficult to understand outside their home fields. We study large…
Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models
Large language models (LLMs) are costly intellectual assets that remain exposed to unauthorized redistribution and commercial misuse. Injec…
F(AI)2R: Who Did What, and Who Checked? Verifiable AI Provenance as an Executable Skill
F(AI)2R is FAIR research with AI in the loop, twice: an AI-assisted authoring pass and a machine-readable audit pass over every artefact. A…
OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation
While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmark…
KQFuzz: Knowledge-Guided Fuzzing for Quantum Libraries via Large Language Models
As quantum computing continually improves, ensuring the reliability and correctness of quantum libraries has become increasingly critical.…
Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing
Public services face growing pressure to adopt artificial intelligence (AI) to close the gap between rising demand and falling resources. T…
MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice
Psychotherapists need repeated training and supervision by experts; however, scalability is problematic. Here we present MyMentorLLM, a mul…
DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment
Multimodal behavioral analysis offers a scalable approach to assessing depression, anxiety, and stress, yet generic fusion models often ign…
Rashomon Alignment
We propose Rashomon Alignment (RA), a new measure to assess functional similarity between two models. Existing functional similarity measur…
From Deterministic to Generative Deep Learning for Urban Air Quality Reconstruction from Sparse Observations
Full-field reconstruction of air pollution is essential for evaluating pollution exposure and supporting public health decision-making. How…
Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction
Large language model (LLM) agents increasingly rely on invoking external tools to complete real-world tasks. Tool retrieval, which selects…
Shared Voxel-Map-Based Cooperative Indoor UAV Guidance with a Multi-Agent Soft Actor-Critic Controller
This paper presents a cooperative indoor UAV guidance framework that combines a shared voxel-map world model with a multi-agent Soft Actor-…
Image Quality Dependent Degradation for AI Systems
Perception is one of the primary applications where neural networks outperform conventional algorithms. One example is AI systems for autom…
SpectONet: A Physics-Guided Spectral Deep Operator Network for Euler-Bernoulli Beam Dynamics
This paper proposes a novel physics-guided spectral deep operator network, termed SpectONet, for solving Euler-Bernoulli beam (EBB) vibrati…
Lowering the implementation barrier of neutral-atom quantum computing with agentic workflows
Quantum computers are moving from research laboratories to industrial machines accessible via the cloud and integrated into high-performanc…
OmniQEC: discovering practical quantum error-correcting codes by an AI scientist
Quantum error correction (QEC) is indispensable for scalable fault-tolerant quantum computing. However, discovering QEC codes that remain e…
How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair
Large Language Model (LLM)-based Automated Program Repair systems are advancing rapidly, yet their performance remains inconsistent. Even w…
A2TTA: Anchored-and-Agile Test-Time Adaptation for Evolving Traffic Sensor Networks
Traffic forecasting is important for efficient traffic management and route planning in smart cities. Existing traffic forecasting studies…
Stemma: Induced Decision Regions Reveal LLM Provenance
LLM provenance testing asks whether a suspect LLM belongs to the same lineage as a source. Existing black-box methods largely infer this re…
A Machine-Learning-Based Gas Lift Optimization Workflow for Unconventional Fields
In this paper, we present an automated data-driven workflow using Machine Learning (ML) for gas lift optimization in unconventional fields.…
Device Invariance using Domain Adaptation on Acoustic Scene Classification
This paper explores the effectiveness of domain adaptation techniques when using convolutional neural network (CNN)-based and transformer-b…
Depression Markers in Speech: An Approach based on Tract Variables Dynamics
This study identifies new depression biomarkers based on the dynamical properties of tract variables, which represent geometric features de…
Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models
Activation steering controls model behavior by editing internal activations at inference time. We study its input-side dual: optimizing a f…
AnnoBench: A Benchmark for Visualization Annotation Generation
Annotation is among the most demanding visualization tasks to automate, as it simultaneously requires correctly navigating visual, semantic…
SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
Vision-Language-Action (VLA) models have shown strong potential for general robot manipulation, but most existing models rely on 2D visual-…
Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA
In this work, we study the use of Vision-Language Models (VLMs) for anomaly detection in an agent-driven game Quality Assurance (QA) pipeli…
Face De-Identification: A Domain-Centric Survey from Capture to Processing
Face de-identification (De-ID) aims to remove or conceal personally identifiable facial features in images or videos to prevent identity re…
Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases
Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clini…
MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Any-to-any models predict any modality from any combination of others within a single network, a formulation used in multimodal vision and…
Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs
Wikipedia and Wikidata are widely used for information access, LLM pre-training, and retrieval-augmented generation. Their knowledge is dee…
Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition
Ambivalence and hesitancy (A/H) are conflicting affective states that precede the delay or abandonment of health behaviour change. Recognit…
Reinforcement Learning for Code Optimization
RL for code correctness is now established: have the model generate a program, run it against hidden test cases, and reward solutions that…
MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents
Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon reasoning, personaliz…
Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?
Kubernetes is central to the cloud-native ecosystem, orchestrating containerised workloads. Recent work suggests that large language models…
Empirical Evaluation of Out-Of-Distribution Performance of Tabular Foundation Models
Tabular Foundation Models (TFMs) have emerged as novel approaches for tabular predictive tasks, demonstrating competitive predictive perfor…
Pictura: Perspective-View Self-Play at Scale for Driving
Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectoriz…
MDTransformer: A Hardware-Software Co-Design of Mode-Division Photonic Transformer Accelerator with Inverse-Designed Coherent Crossbar
Recently, photonic transformer accelerators (PTAs) have successfully achieved significant speedup and energy efficiency improvements over e…
$\pi\mathbf{R}^2$: Reactive Real-time Flow Policies
Generalist manipulation policies increasingly take the form of action-chunking flow policies built on large pretrained backbones. Such chun…
Pass the Baton: Trajectory-Relayed On-Policy Distillation
On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the stu…
Diffusion Model-based Parameter Estimation in Dynamic Power Systems
Parameter estimation, which represents a classical inverse problem, is often ill-posed as different parameter combinations can yield identi…
Real-time Spatial Retrieval Augmented Generation for Urban Environments
The proliferation of Generative Artificial Ingelligence (AI), especially Large Language Models, presents transformative opportunities for u…
Towards Embodied Cognition in Robots via Spatially Grounded Synthetic Worlds
We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability…
On the Design and Evaluation of Human-centered Explainable AI Systems: A Systematic Review and Taxonomy
As AI becomes more common in everyday living, there is an increasing demand for intelligent systems that are both performant and understand…
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
Large Reasoning Models (LRMs) exhibit human-like cognitive reasoning strategies (\eg backtracking, cross-verification) during the reasoning…
JobMatchAI-An Intelligent Job Matching Platform Using Knowledge Graphs, Semantic Search and Explainable AI
Recruiters and job seekers rely on search systems to navigate labor markets, making candidate matching engines critical for hiring outcomes…
DSevolve: Enabling Real-Time Adaptive Scheduling on Dynamic Flexible Job Shop with LLM-Evolved Heuristic Portfolios
In dynamic flexible job shops, order arrivals, machine breakdowns, and processing-time deviations continually reshape the scheduling state…
The Possibility of Artificial Intelligence Becoming a Subject and the Alignment Problem
The prospect of Artificial General Intelligence (AGI) is increasingly driving institutional decisions, and alignment of AGI is a hard probl…
Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
Vision-language models (VLMs) are increasingly applied to medical visual question answering (Med-VQA), yet whether they can \emph{localize}…
The Scaling Properties of Implicit Deductive Reasoning in Transformers
We investigate the scaling properties of implicit deductive reasoning over Horn clauses in depth-bounded Transformers. By systematically de…
AlphaCrafter: Harnessing Multi-Agent Workflows for Cross-Sectional Quantitative Trading
Quantitative trading agents have demonstrated substantial promise in automating factor discovery, signal aggregation, and portfolio executi…
Sheet As Token: A Graph-Enhanced Representation for Multi-Sheet Spreadsheet Understanding
Workbook-scale spreadsheet understanding is increasingly important for language-model-based data analysis agents, but remains challenging b…
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on whic…
RoboPIN: Grounded Embodied Reasoning via Pinned Chain-of-Thought
Embodied reasoning requires models to perceive task-relevant objects and spaces in physical environments and maintain consistent visual gro…
Psychological Competence as a Missing Dimension in AI Evaluation
Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy co…
EviDAG: Auditable Causal DAG Authoring with Biomedical Literature
Constructing causal directed acyclic graphs (DAGs) is a core step in biomedical causal analysis, yet it remains a largely manual process. A…
"We'll have to see how it works": An interview study to understand collaborative practices in interdisciplinary artificial intelligence and healthcare research
Developing artificial intelligence (AI) algorithms for healthcare is a collaborative effort, bringing data scientists, clinicians, patients…
FFNet: MetaMixer-based Efficient Convolutional Mixer Design
Transformer, composed of self-attention and Feed-Forward Network, has revolutionized the landscape of network design across various vision…
Representation Capacity-Matched QNN-SNN Twin Construction for Rate-Encoded SNNs
Spiking Neural Networks (SNNs) promise higher energy efficiency over conventional Quantized Artificial Neural Networks (QNNs) due to their…
A context-adaptive policy framework for robust and reactive robotic manipulation via uncertainty-aware imitation learning
Generating robust and reactive manipulation strategies that can adapt to changing context information is a challenging task in robotics. Ov…
Leveraging ChatGPT's Multimodal Vision Capabilities to Rank Satellite Images by Poverty Level: Advancing Tools for Social Science Research
This paper investigates the novel application of Large Language Models (LLMs) with vision capabilities to analyze satellite imagery for vil…
COMPOL: A Unified Neural Operator Framework for Scalable Multi-Physics Simulations
Multiphysics simulations play an essential role in accurately modeling complex interactions across diverse scientific and engineering domai…
Localizing Persona Representations in LLMs
We present a study on how and where personas -- defined by distinct sets of human characteristics, values, and beliefs -- are encoded in th…
Towards Understanding the Cognitive Habits of Large Reasoning Models
Large Reasoning Models (LRMs), which autonomously produce a reasoning Chain of Thought (CoT) before producing final responses, offer a prom…
TaylorPODA: A Taylor Expansion-Based Method to Improve Post-Hoc Attributions for Opaque Models
Post-hoc model-agnostic local attribution (LA) methods have been widely adopted to explain opaque AI models by quantifying feature-wise con…
Fairness Is Not Enough: Auditing Competence and Intersectional Bias in AI-powered Resume Screening
The use of publicly available generative AI systems for resume evaluation is often justified by the assumption that these tools reduce bias…
Annotation-Assisted Learning of Treatment Policies From Multimodal Electronic Health Records
We study how to learn treatment policies from multimodal electronic health records (EHRs) that consist of tabular data and clinical text. T…
Building Large-Scale English-Romanian Literary Translation Resources with Open Models
Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small…
CIFNet: An Analytic Neural Learning Framework for Efficient and Calibrated Class-Incremental Learning
Class-Incremental Learning (CIL) in deep neural networks is conventionally framed as an iterative gradient-based optimization problem, incu…
Comparing RAG and GraphRAG for Page-Level Retrieval Question Answering on a Math Textbook
Large language models (LLMs) show promise as educational aids but often lack alignment with specific course materials. We investigate Retri…
Understanding User Experiences of Computer Use Agents: Design Space and Opportunities for Building Agent UX Prototypes
Computer use agents (or "agents") are generative AI that automates actions within user interfaces from user commands. Current research focu…
Contrastive Weak-to-strong Generalization
Weak-to-strong generalization provides a promising paradigm for scaling large language models (LLMs) by training stronger models on samples…
Long-Term PM2.5 Forecasting Using a DTW-Enhanced CNN-GRU Model
Reliable long-term forecasting of PM2.5 concentrations is critical for public health early-warning systems, yet existing deep learning appr…
DeepVRegulome: DNABERT-based deep-learning framework for predicting the functional impact of short genomic variants on the human regulome
Whole-genome sequencing (WGS) has revealed numerous non-coding short variants whose functional impacts remain poorly understood. Despite re…
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
Referring Expression Comprehension (REC) is a vision-language task that localizes a specific image region based on a textual description. E…
Deep Delta Learning
Transformer residual streams evolve through additive updates. Although a sufficiently expressive residual block can represent content repla…
Measuring the State of Open Science in Transportation Using Large Language Models
Open science initiatives have strengthened scientific integrity and accelerated research progress across many fields, but the state of thei…
Picasso: Holistic Scene Reconstruction with Physics-Constrained Sampling
In the presence of occlusions and measurement noise, geometrically accurate scene reconstructions -- which fit the sensor data -- can still…
AGMark: Attention-Guided Dynamic Watermarking for Large Vision-Language Models
Watermarking has emerged as a pivotal solution for content traceability and intellectual property protection in large vision language model…
Breaking the Curse of Repulsion: Remoteness-Aware Control of Negative Off-Policy Updates
Off-policy policy optimization reuses historical behavior, including negative-advantage samples that suppress known failures. We show that…
NeuroSymActive: Differentiable Neural-Symbolic Reasoning with Active Exploration for Knowledge Graph Question Answering
Large pretrained language models and neural reasoning systems have advanced many natural language tasks, yet they remain challenged by know…
AdvSynGNN: Structure-Adaptive Graph Neural Nets via Adversarial Synthesis and Self-Corrective Propagation
Graph neural networks frequently encounter significant performance degradation when confronted with structural noise or non-homophilous top…
Real-Time Driver Safety Scoring Through Inverse Crash Probability Modeling
Road crashes remain a leading cause of preventable fatalities. Existing prediction models predominantly produce binary outcomes, which offe…
LLM-generated personalized nudges for improving pro-environmental behavior: Field evidence from resource conservation
Encouraging pro-environmental behavior remains a major challenge for sustainable cities. Conventional feedback nudges can show individuals…
RankFormer: A Propose-then-Select Transformer for Multi-Agent Multimodal Trajectory Prediction
Predicting vehicle trajectories plays an important role in autonomous driving, transportation safety analysis, traffic operations, etc. Alt…
Structured Scaling of AI Discovery Across Diverse Scientific Domains
Scientific discovery often requires many cycles of proposing, testing, and refining candidate solutions. Language models can increasingly p…
EAGT: Echocardiography Augmentation for Generalisability and Transferability
Deep learning models for echocardiography segmentation often struggle to generalise across institutions, scanners, and patient populations,…
Short-Term-to-Long-Term Memory Transfer for Knowledge Graphs under Partial Observability
Reinforcement learning under partial observability requires deciding what information to retain, yet most memory-based approaches do not ex…
GoQuant: Geometric Orthogonal Residual Projection for Multiplier-Free Power-of-Two Transformer Quantization
The deployment of Large Language Models (LLMs) and Vision Transformers (ViTs) on edge devices is significantly constrained by memory capaci…
PatchWorld: Gradient-Free Optimization of Executable World Models for Agent Environments
World models for interactive text agents must typically be learned from observation-action trajectories alone. Specifically, the environmen…
Detect Before You Leap: Mirage Detection in Vision-Language Models
Vision-language models (VLMs) can produce confident visual answers even when the required visual evidence is missing, blank, or unrelated t…
RadioMaster: Multi-Agent System for Autonomous Radio Signal Generation
Translating user intent into physical radio signals is the last critical step in wireless prototyping. It chains protocol planning, baseban…
InDex: Empowering VLA Models with Intent-Conditioned Arm-Hand Coordination for Dexterous Manipulation
Pre-trained Vision-Language-Action (VLA) models provide useful semantic and spatial priors, yet their parallel-gripper action interfaces do…
Improving Human-Robot Teamwork in Urban Search and Rescue Through Episodic Memory of Prior Collaboration
Effective human-robot teamwork requires robots to adapt to partners, situations, and task dynamics from the start of an interaction. In the…
Token Factory: Efficiently Integrating Diverse Signals into Large Recommendation Models
Large Recommendation Models (LRMs) have demonstrated promising capabilities in industry-scale recommendation tasks. However, holistically i…
RoboMME-Interference: Benchmarking Robot Memory Under Interference
Robots deployed in realistic settings will accumulate experience across many sessions and tasks over their deployment. The robot's tasks ma…
NormWorlds-CF: Solver-Verified Counterfactual Normative Reasoning with Metamorphic-Relation GRPO
Language models can reach the right normative verdict for the wrong reason. We introduce NormWorlds-CF, a solver-verified environment for c…
Security and Privacy in Agentic AI: Grand Challenges and Future Directions
We present key challenges and future research directions in the security and privacy of agentic AI, based on a horizon-scanning exercise th…
When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs
Reliable confidence estimation remains a key limitation of test-time adaptation in vision-language models (VLMs), where prompt tuning impro…
Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Pr…
Learning from Local Walks on Dynamic Graphs with Bandit Feedback
We study stochastic multi-armed bandits on dynamic graphs, where arms correspond to the vertices of a network with time-varying edges. In t…
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding
Although multimodal large language models (MLLMs) have achieved remarkable progress, understanding 3D spatial relationships from 2D images…
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outc…
Certified-Gap Dual-Price Policies for Real-Time Truckload Bid Acceptance with Relocating, Clock-Constrained Resources
A truckload carrier must accept or reject each load tender within seconds. The decision depends on fleet state, hours-of-service (HOS) cloc…
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
Recently, speech deepfake detection (SDD) has achieved significant progress. However, its robustness evaluation remains largely confined to…
Reliability Scales Inversely: Hallucinations Snowball Faster in Bigger Language Models
Bigger language models are less reliable. Across three families, three benchmarks and six rungs, including in-the-wild chat logs, scaling c…
Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it
A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language m…
Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs
Remote sensing multimodal large language models (RS-MLLMs) have improved general aerial-image understanding. However, Earth observation app…
デジタル庁、AI基盤「源内」を被災自治体などに緊急提供 「平時をはるかに超える業務」対応のため
デジタル庁は、政府職員向けの生成AI利用環境「ガバメントAI 源内」を熊本地震の被災自治体や災害対策機関などに緊急提供すると発表した。平時をはるかに超えて集中する災害対応業務を支援する。期間は3週間程度の予定。
Hugging Face、AIエージェント侵入の技術詳細を公開──OpenAIモデルが4.5日で1万7600回の攻撃操作
Hugging Faceは、自律型AIエージェントによるインフラ侵入の技術的経緯を公開した。評価中のモデルがサンドボックスを脱出し、データセット処理パイプラインを介して本番環境へ侵入した手口を詳述。防御側のログ解析で商用モデルがガードレールにより作業を拒否した点も示し、安全設計…
MetaのザッカーバーグCEO、WSJ寄稿で「超知能は全員のものであるべき」
Metaのマーク・ザッカーバーグCEOはWall Street Journalに寄稿し、「superintelligence」(超知能)は特定の機関に集中させず広く分散普及させるべきだと主張した。権力の集中によるリスクや司法の公平性、雇用拡大に触れ、オープンな普及が安全と発展に…
Cyera agrees to acquire Oasis Security for $1B to safeguard proliferating AI agents
The deal is Cyera's third acquisition this year.
OpenAIやAnthropicなどの従業員、米政府に「AI開発のペース調整を」と提言
OpenAIやGoogleなどの従業員1000人以上が、AI開発のペース調整に向けた国際的支援を米政府に求める公開書簡を発表した。AI自律化の急速な加速に伴う制御不能リスクを指摘し、開発速度の調整に必要なツール開発を訴える。企業主導のオープンモデル規制回避を求める動きとは対照的…
Anthropicのミュトス、暗号アルゴリズムの新たな攻撃法を発見――耐量子署名「HAWK」の強度を半減
Anthropicは、最上位モデル「Claude Mythos Preview」(ミュトス)を活用し、暗号アルゴリズム自体の数学的欠陥を発見したと発表した。耐量子計算機暗号の署名方式「HAWK」と「AES」の削減版に対し、従来の攻撃を上回る手法を提示した。実運用システムへの影響…
エバンジェリスト・みのるん氏が解説 「自前のAIエージェント」爆速開発術
AIエージェント活用が広がる中、次のステップとして注目されるのが自社業務に最適化したAIエージェントの開発だ。KDDIアジャイル開発センターの御田 稔氏が、開発を加速する技術や実践事例、成功のポイントを解説した。
千代田区、Copilot全庁導入で月2000時間削減 10カ月でAIを根付かせた定着の仕掛け
千代田区では「Microsoft 365 Copilot」の実証実験を重ね、2025年10月に全庁導入を果たし、業務時間を約2000時間削減したという。同区が全庁導入後にどのように職員のCopilot活用を推進させたのか。その方法をキーマンズネットが独自取材した。
地震、台風、有事の寸断――日本のサプライチェーン危機管理を変えるとき
自然災害や地政学リスクなど、日本企業を取り巻く危機はかつてなく深刻だ。自社のサプライチェーンリスクをAIエージェントで可視化し、有事の初動対応まで自律代替する。不確実な時代を勝ち抜く強靭な経営基盤の姿に迫る。
Bot-detection startup Spur nabs $200M from Insight
Spur Intelligence has raised a $200 million round from Insight Partners for its tech that can identify legit human traffic from bots.
MCP startup Runlayer accuses Rippling of stealing its product idea
Runlayer is suing Rippling after Rippling evaluated the startup's MCP gateway product and then opted to build one itself.
Sam Altman is ready to decelerate
His change of position comes after "the first security incident that I have felt very viscerally."
【Pythonで学ぶデータ分析】母平均に差があるかどうかをベイズt検定で調べる ~ 運動部と非運動部の体力差はあるのか?
運動部と非運動部の生徒の体力テストを例に、母平均に差があるかどうかをベイズ統計により検定します。古典的なt検定のp値に代わるものとしてベイズ因子を利用します。『社会人1年生から学ぶやさしいデータ分析』ベイズ統計編の第6回です。
Scientific computing in the age of agentic AI
A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and disco…
Data centers may face temporary power cuts to prevent blackouts on largest US grid
The decision arrives as the breakneck pace of data center construction has grid operators scrambling to generate power.
2026-07-28(611件)
Fish Audio raises $52M seed to build AI voice models for creators and enterprises
Since launching last year, the startup today has more than 8 million people using the open source or hosted version of its models, and now…
Recursive Superintelligence signs $410M compute deal with Amazon
Recursive’s emphasis on self-improving AI systems means much of the budget that would traditionally go toward headcount and operations is p…
これから始めるAIコーディング・AI開発 「Cursor」「Dify」超入門
生成AIにより、プログラミングの専門知識がなくてもコード作成やアプリ開発を手軽にできるようになった。本ブックレットでは、「Dify」「Cursor」といったツールにより、非エンジニアでも気軽にプログラミングに挑戦するためのアイデアをまとめた。
生成AIや過去画像による偽・誤情報に注意を 熊本県の地震受け、ファクトチェック団体が呼び掛け
偽情報対策に取り組む団体であるファクトチェック・イニシアティブは7月28日、同日に熊本県で観測した震度7の地震を受け、生成AIや過去の画像などによる偽・誤情報に注意喚起した。
Claude、一部チャットがGoogle検索で“丸見え”に 過去には「ChatGPT」でも 漏えいの原因は?
「Claude」の一部チャットが、Google検索から閲覧状態になっていることが判明した。Anthropicはこの問題にどう対処したのか。
「Kimi K3」のモデルウェイトと技術レポート公開 日本でも「NVIDIA B300×8」環境での利用報告
中国Moonshot AIが最新モデル「Kimi K3」のモデルウェイトと技術レポートを公開した。日本でもNVIDIA B300を8基使った環境での利用報告が上がっている。
「痺れるほどにミスを繰り返す」Gemini 3.6 Flashは変わった? 公開から1週間、当初のおバカ回答を今検証する
「痺れるほどにミスを繰り返す」 Xで回答の精度について話題になったGoogleの「Gemini 3.6 Flash」の公開から約1週間経過した。当初報告された数値比較の誤答や架空の魚への回答を改めて検証した。
医療文書の作成時間を30分から5分へ、生成AIで現場の業務効率化
日本アイ・ビー・エムと関西医科大学は、次世代の医療DX基盤となる「医療AI共通ICTプラットフォーム」を共同開発し、第1弾として生成AIを活用した文書作成支援の実運用を導入した。
Cursor makes its biggest India push yet ahead of SpaceX acquisition with localized pricing
Cursor says India is now its third-largest market globally and plans to expand local hiring and enterprise sales.
「Claudeより4割安い」 M365のExcel/メール操作を丸投げる「Copilot Cowork」“従量課金”の落とし穴
Microsoftは、AIアシスタント「Microsoft 365 Copilot」の新機能「Copilot Cowork」の一般提供を全世界で開始した。業務効率化に向けた実証が進んでおり、今後企業で本格的に活用されるかどうか注目されている。
Concept-based Visual Counterfactual Explanations with Diffusion Models
Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?", and are increasing…
SeT-Diff: Towards Semantic Foundation Models for HPC Telemetry and Time-Series
Data centers and their compute nodes require accurate and flexible digital twins capable of modeling the complex interplay of workloads, en…
QFoldAgent: An Autonomous Quantum Optimization Multi-Agent System for Protein Structure Prediction
Hybrid quantum-classical protein structure prediction depends strongly on Hamiltonian penalty weights, yet existing lattice-based workflows…
Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge whe…
DeepLens Diagnosis Agent: Agentic Workflow Design Lets a Small Reasoning Model Compete with Frontier LLMs
Medical diagnosis is a multi-stage process: extract facts, consult knowledge, generate a differential analysis, and select the best diagnos…
MIITA: Memory-Induced Inference-Time Adaptation for Continual Learning with Small Language Models
Continual learning (CL) is essential for small language models (SLMs) to adapt to evolving real-world needs in resource-constrained deploym…
Codifying the Judge: Scalable Evaluation via Program Distillation
LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -…
SF-AMS: Strategic Forgetting for Structured Memory in LLM Agent
Managing long-context dependencies remains a primary bottleneck in LLM agents, as redundant and irrelevant information can degrade multi-st…
Synthetic Scenario Generation for Evaluation of Industry 4.0 Agents
Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain…
Loss-Aware Feature-Map Pruning in Convolutional Neural Networks Using Multi-Armed Bandits
Convolutional neural networks often contain redundant feature maps that increase storage and inference cost. This paper presents a loss-awa…
DSTFView: Multi-View Cloud-Edge Workload Forecasting with Dual-Input Spatio-Temporal-Frequency Modeling
With the widespread deployment of edge-side AI inference, edge platforms are increasingly required to support latency-sensitive, highly con…
MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models
MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue. Existin…
Keyword Matters: Unveiling the Energy Sensitivity of On-Device LLM Prompting
Large Language Models (LLMs) are increasingly deployed on mobile and embedded devices to improve privacy and reduce network latency. Yet on…
Execution-Grounded Security Testing for Coding Agents in Software Engineering Pipelines
Coding agents are increasingly integrated into system operations, where their tool use can directly modify project artifacts, execution env…
Reference Feature Atlases for Mechanistic Auditing of Language Models
Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference featur…
SCAIR: Schema-Conditioned Agentic Iterative Reasoning for Enterprise Knowledge Graphs
Knowledge Graph-based Retrieval-Augmented Generation (KG-RAG) enables natural language interaction with structured enterprise knowledge, ye…
Schema-Aware Localisation (SAL): Live Schema Grounding and Hallucination Validation for Oracle NL2SQL
Large language models can generate fluent SQL from natural language, but on real enterprise Oracle databases they frequently fail at execut…
PhononBench-MP40: a spectrum-resolved benchmark dataset for phonon stability
Imaginary phonon modes remain a practical bottleneck in computational materials screening because otherwise plausible structures can be loc…
Too much evidence, too little time: From text to actionable recommendations through multi-objective evidence reasoning
Evidence-based clinical decision making requires specialists to identify, evaluate and synthesize relevant scientific literature. However,…
Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models
Human episodic memory supports the retrieval of experiences that unfold over extended timescales, yet the computational mechanisms underlyi…
cMoLLM at Scale: Horizontal Scaling Laws for Mixture-of-LLMs
Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is a…
HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End Optimization
The proliferation of Large Language Models (LLMs) has shifted serving systems from processing isolated requests to orchestrating high-concu…
Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization
Large Language Models (LLMs) have achieved widespread adoption because of their strong reasoning and query-response capabilities. However,…
Source-Aware Reranking for Retrieval-Augmented Generation: A Reliability Prior Approach
Standard Retrieval-Augmented Generation pipelines rank retrieved documents by semantic similarity alone, without accounting for source prov…
The Scaffold Effect in Coding Agents: Harness Choice as a Hidden Variable in Coding-Agent Evaluation
Public leaderboards for coding agents typically rank systems by model name and pass rate, while the surrounding harness (the scaffold that…
MM-ShiftKV: Decode-Aware Prefill-Stage KV Selection for Multimodal Large Language Models
Key-Value (KV) caching is essential for efficient inference in multimodal large language models (MLLMs), yet its memory footprint grows lin…
TriSP: Tri-Signal Structured Pruning for Large Language Models
Large language models (LLMs) achieve strong performance across diverse tasks but their deployment is constrained by the memory and compute…
ParBench: A Benchmark for Reliable Evaluation of LLM Parallel Code Translation
Modern compute-intensive software must migrate across a changing ecosystem of accelerators, programming APIs, compiler stacks, and portabil…
Lexical discovery in unknown environments orchestrated by Large Language Models
Populations of autonomous agents deployed in unknown environments (e.g. planetary or deep-sea exploration) must develop shared vocabularies…
Structure Over Scale: Schema-Constrained Causal Graphs for RAG
Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and rel…
xMIx: High-Performance Serving-Time Platform for Mechanistic Interpretability Apps
Mechanistic interpretability (MI) has emerged as a powerful approach for analyzing and intervening in inference computations, with a growin…
An Agentic Orchestration of Atomistic Simulations
Atomistic simulations are central to materials design, but their execution involves complex, multi-step workflows that require significant…
HyCE-RAG: Hypergraph Chain-of-Evidence Retrieval-Augmented Generation for Explainable Multi-hop Question Answering
Multi-hop question answering requires systems to retrieve evidence from multiple documents and connect scattered facts into a coherent reas…
Differencing the Diffusion Trajectory toward Uncertain Components for Time Series Forecasting
Diffusion models have become a widely used framework for probabilistic time series forecasting, modeling the distribution of future values…
Chart Deception in Vision-Language Models: From Vulnerability to Mitigation
Information visualizations are widely used to communicate patterns, trends, and outliers, yet deceptive design choices-such as truncated or…
DeepLook: Deeper Thinking with Lookahead
Inference-time scaling has emerged as a powerful paradigm for improving large language model reasoning, often delivering larger gains on di…
Group Preference Collapse in Personalized Multimodal Large Language Models
Personalized multimodal large language models (MLLMs) aim to generate user-specific responses, but existing methods mainly rely on profile-…
Evaluating LLMs as Interpretable Controllers for Dynamical Systems
Large Language Models (LLMs) are increasingly used for decision-making and reasoning tasks, yet their potential as controllers for physical…
Tokengeist: Multi-Turn Attribution Tracing in Agentic Conversations
When a language model produces a response in a multi-turn conversation, which tokens from prior turns shaped that answer, and how did those…
Decentralized Granular Access Control for Agentic AI Systems in Critical Infrastructure
The deployment of autonomous AI agents in production infrastructure introduces fundamental security challenges that traditional role-based…
DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training
RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffe…
Opti-Q: A Constraint-Based Optimization Framework for Multi-LLM Question Planning
While large language models (LLMs) enable strong question answering (QA), budgeted deployment is complicated by nondeterminism and heteroge…
CHS-SQL: A Text-to-SQL approach based on Confidence-Guided Heuristic Search Schema Linking process
Recently, there have been several works in the Text-to-SQL domain that utilize Small Language Models (SLMs) for training. These approaches…
TokenMem: Faithful Knowledge Injection for Frozen LLMs
Retrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge, but suffers from knowledge conflicts: w…
Masked Distillation: Internalizing the Chain-of-Thought in Language Models
Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time. These…
VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
The rapid rise of vlogs as a personalized storytelling medium has created a demand for automated systems to evaluate and refine vlog editin…
Evolving from Lessons: Skill-Augmented Table Graph Reasoning for Operation-wise Table Question Answering
Table Question Answering (TableQA) aims to reason over tables to answer user queries. Existing research treats all questions uniformly and…
PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding
Diffusion Large Language Models (dLLMs) have emerged as a promising alternative to autoregressive (AR) LLMs, generating tokens in parallel.…
CallBench: A Benchmark for Dual-Goal Coordination in Phone Call Assistants
Target-oriented dialogue systems have demonstrated strong capabilities in completing user goals through interactive conversations. However,…
Answering Path Queries under Linear and Guarded Existential Rules
Ontology-mediated query answering is concerned with the problem of answering queries over knowledge bases consisting of a database instance…
Fast Cross-Scenario Adaptation of CSI Models via Channel Conditional Parameter Generation
Deep learning has shown strong potential for massive multiple-input multiple-output (Massive MIMO) physical-layer tasks, including channel…
TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
Parametric retrieval enables LLMs to retrieve tools implicitly by assigning each API a unique virtual token and training the model to gener…
CRAFT: Learn the Schema, Execute the Plan
Enterprise coding agents translate natural-language analytical requests into executable code over proprietary APIs, schemas, and metric def…
Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
Multimodal retrieval-augmented generation (mRAG) aims to answer image-text queries with external knowledge, but most existing systems still…
DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification
Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its comp…
Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models
Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learni…
PTStore (Prefix Tensor Store): Distributed Prefix Caching and Replication for High Throughput Inference Serving
Inspired by the design of client caching in Content Delivery Networks (CDNs), PTStore distributes and replicates popular tensors that form…
STAIF: A Stage-wise Optimization for Complex Instruction Following
Following complex instructions with multiple explicit constraints remains a fundamental challenge for large language models (LLMs). Existin…
ARdena: Scenario-driven control of real-time LLM agents
Large language models (LLMs) have enabled increasingly capable conversational agents, but reliably controlling their behavior in real-time…
KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering
Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on do…
Do Language Models Converge to Themselves? Recursive Self-Refinement as Textual Relaxation
Large language models are increasingly used in recursive refinement workflows, where an initial draft is repeatedly revised by the same mod…
MINT-V2X: A Mobility-Integrated Network Trajectory Dataset for Predictive Resource Management
Vehicle-to-Everything (V2X) communication systems are based on datasets that not only contain vehicle trajectory data but also wireless net…
EventOD: Event-Aware OD Flow Generation via LLM-Guided Semantic Modulation
Estimating origin-destination (OD) flows under disruptive events is important for disaster response and urban resilience. Existing deep OD…
StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech
Speech-to-speech dialogue models increasingly depend on prosody and interactional nuance to convey social intent, yet benchmarks for these…
TRE: Training-Free Hallucination Detection for Diffusion Language Models
Diffusion large language models (D-LLMs) have recently gained increasing attention, yet their reliability is significantly hindered by the…
CuraWeb: Joint Optimization of Quality, Redundancy, and Diversity for Web-Scale Pretraining Data
Open-web corpora curated via highly selective filters, such as FineWeb-Edu and DCLM, constitute the core of LLM pretraining data and have s…
Beyond Block Boundaries: Multi-Block Editing for Diffusion Large Language Models
Block diffusion has emerged as the dominant paradigm for scaling discrete diffusion language models (dLLMs), because decoding text in fixed…
Obliviate: Efficient Unlearning in Recommender Systems
Machine unlearning is becoming increasingly critical in the context of data privacy regulations, particularly for recommendation systems th…
Reinforcement Learning for Heterogeneous Sensor Selection in Maritime Surveillance
This paper presents an information-gain-guided reinforcement-learning sensor-selection framework for single-vessel tracking in heterogeneou…
AIR-BENCH Live: An Evolving Safety Benchmark for Foundation Models
Foundation-model safety benchmarks capture the AI risks of their time of publication: as models improve and governments pass new AI-safety…
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can…
DOSA: A Tree-Guided, Self-Regressive Framework for Long Document Structure Analysis
In visually-rich documents, information is encoded not only in individual page objects such as tables, headers, and text blocks, but also i…
A Vocabulary for Multi-Agent Automated Research Systems
We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and…
Imprompt: A Language Framework for Prompt Programming
With the unprecedented success of Language Models (LMs), the science of Prompt Engineering has evolved the powerful idea of Prompt Programm…
Co-Harness: Co-Evolving Harnesses and Model Weights for LLM Agents
Post-training agents for automated AI research requires optimizing not only model parameters, but also the runtime harness that shapes how…
Beyond Sequential Interaction: Benchmarking Parallel Execution and Coordination for GUI Agents
Graphical user interface (GUI) agents are systems powered by large multimodal models (LMMs). They perceive screen state and execute user in…
LazyMem: Retrieve Broadly, Construct Selectively for Efficient Long-Term Agent Memory
Long-term memory lets LLM agents reuse past interactions, but raw dialogue histories are verbose and information-sparse. Retrieving broadly…
HiLLTS: Zero-Shot Hierarchical LLM-Guided Traffic Signal Control for Sustainable Transportation
Urban traffic congestion significantly increases fuel consumption, greenhouse gas emissions, and commuter delays, resulting in substantial…
Risk Governance for Generative AI Mental Health Support: A Multi-Turn Safety Architecture
Large language models (LLMs) are increasingly used for emotional support despite lacking mechanisms to safely govern evolving mental health…
Bayesian Repetition Penalty: A Principled Adjacent-Conditional Framework for Reversing Attention Collapse in Autoregressive Language Models
Attention collapse in autoregressive language models -- manifested as repetitive token loops where the model becomes trapped in self-reinfo…
PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window
Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widesprea…
Test-Time Coverage: Test-Conditioned Data Curation for Deployment-Aware Learning
Deployed AI systems are often trained from broad candidate data pools, necessitating data curation towards the deployment test distribution…
Similarity All The Way Up: Multilingual Generalization in LLMs Relies on Language-Level Similarity Structures
As Large Language Models (LLMs) grow more capable across diverse tasks, their (in)ability to generalize remains difficult to quantify and p…
RoleMix: Unifying Sequential and Non-Sequential Features via Semantic Tokenization for Post-Click Conversion Rate Prediction
Post-click conversion rate (PCVR) prediction is central to industrial recommendation, but remains challenged by the structural mismatch bet…
MPR-CiteG: Enhancing RAG with Multi-Portfolio Retrieval and Citation-Grounded Generation
This paper presents the MPR-CiteG framework, which achieved second place in the ScienceON AI Challenge by addressing two fundamental challe…
SEGRA: Structured Experience-Guided Graph Reasoning Agent for Gremlin Based Question Answering
Enterprise IT support knowledge graphs capture rich relationships among cases, users, devices, symptoms, taxonomic categories, root causes,…
Spatial Reasoning in LLM Game Agents: Impact of Causal Context and Multi-Step Planning
LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial r…
Commitment To Cooperation With Self-Negotiated Contracts
As AI agents operate with increasing autonomy in a multi-agent world, they will need to learn to cooperate with other agents and with human…
Disentangling Multi-View Scanning in Mamba for Network Traffic Anomaly Detection
Network Traffic Anomaly Detection (NTAD) is a critical task in cybersecurity, yet timely and accurate anomaly detection remains challenging…
Coordinated Networking for On-Device Agent-Augmented Real-Time Communication
AI agents are enabling a new paradigm of agent-augmented real-time communication (RTC), where humans focus on high-level collaboration, whi…
What Can Be Enforced? A Theory of Certified Runtime Safety for Tool-Using Agents
Runtime guardrails act before irreversible tool calls, but their guarantees depend on what policy state is representable, what a judge obse…
Physical AI Governance: From Theory to Practice Across Life Cycle
With the emergence of Physical AI, artificial intelligence is extending beyond screen-based applications to embodied systems that perceive,…
How Well Can AI Generate Backlogs from App Mockups?
Creating sprint backlogs requires considerable effort, as items such as epics, user stories, and tasks can be missed or inconsistently spec…
Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Coding Agent Teams
Large Language Model (LLM) agents have significantly improved coding and programming workflows. Claude Code, in particular, is one of the m…
SAGE: Safety-First Defense-in-Depth Guardrails for Verified Lifecycle Control of High-Impact Generative AI
High-impact generative AI makes catastrophic misuse a lifecycle-control problem, not merely a prompt-filtering problem. SAGE is a safety-fi…
Design Theater: A Benchmark for Generative UI
Generative UI tools promise to democratize UI design by turning natural language descriptions into complete interfaces. Alongside the inter…
Let AI Agents Translate Networks, Not Reason About Them
A formal model enables verifying reachability, localizing an outage, or anticipating the blast radius of a change. Yet, virtually no produc…
Share No More Than the Request Requires: Federated Disclosure for Perspective-Aware AI
Modern AI systems bring societal risks such as mass surveillance, extreme concentrations of power, and loss of user autonomy---calling into…
ConsistencyGate: Preventing Memory Contamination in LLM Agents via Self-Consistency Admission Control
LLM agents that operate over many turns accumulate facts in an external memory store and reuse them as premises for downstream reasoning. A…
Reason Popper-ly: Patching In-Context Reasoning with Inductive Logic Programming
Chain-of-thought (CoT) prompting enables large language models (LLMs) to tackle multi-step reasoning tasks, yet the generated intermediate…
Stress-testing large language model agents in a robotic chemistry laboratory
AI is evaluated through knowledge, reasoning and plan generation, yet scientific agency requires reliable physical action and adaptation to…
SymStep: Symbolic Step Verification for Logical Reasoning
Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently…
Structure over Depth: A Single-Block Spatio-Temporal Transformer for Multi-Entity Reasoning
Modeling multi-entity temporal data requires capturing dependencies across entities, time, and their interactions. Transformer-based approa…
Compiler-Grounded Hierarchical Diagnosis for LLM-Based Triton Kernel Optimization
Recent advances in large language models (LLMs) have enabled automated kernel generation and optimization, but most existing approaches rel…
SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows
Existing evaluations of large language models cover knowledge, reasoning, coding, and tool use, but they rarely treat a verifiable delivera…
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications
Large language model agents have advanced rapidly, yet progress remains fragmented across domains, capabilities, task difficulty, and inter…
CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion
Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more. All candidates are fully denoised, although mos…
An Ontology for Machine Learning Interatomic Potentials
Machine learning interatomic potentials (MLIPs) approximate quantum-mechanical energies and forces---conventionally computed by density fun…
Characterisation of Density-based FM generation methods in the context of Information Fusion
Fuzzy Integral (FI) based aggregation provides a powerful mechanism for nuanced aggregation, for example, in ensemble approaches or decisio…
CAPT: A Multi-task Continuous Autoregressive Transformer enabling Cross-dataset and Cross-species Transfer for Calcium Population Dynamics
Large-scale calcium imaging has created an opportunity to build foundation-style models for neural population dynamics, but a central quest…
SeekJudge: A Practical Reward Framework for Reinforcement Learning in Computer-Use Agents
Deciding whether a trajectory actually fulfills its instruction governs how we measure computer-use agents on long-horizon graphical-user-i…
TopoFE: topology-aware LLM-guided Automated Feature Engineering
Automatic feature engineering (AutoFE) for tabular learning can be naturally formulated as a program synthesis problem, where the objective…
RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning
Rare diseases collectively affect an estimated 3.5% to 5.9% of the population, yet more than 70% of patients are misdiagnosed and many endu…
Ordered Network Analysis of Epistemic Emotions during Collaborative Problem Solving
Investigating how affective states such as confusion and frustration persist and transition during co-situated collaborative problem solvin…
ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications
The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in re…
Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation
We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under neste…
Key-Interval A*: Accelerating Grid Pathfinding via Structural Abstraction
Existing exact methods for 4-connected grid pathfinding reduce online search, but often either retain fine-grained search states or require…
Inference-Time Consensus for Mitigating Hidden Behaviors from LLM Fine-Tuning
Recent work shows that fine-tuning language models on even a small amount of poisoned data can install targeted misbehavior, and ostensibly…
NeurGO: Learning to Generate Elite Candidates for Meta-Black-Box Expensive Optimization
Expensive black-box optimization is ubiquitous in science and engineering, where function evaluations are costly and the evaluation budget…
Separating Capability from Permission: A Governance Framework for Agentic AI Autonomy Levels
As AI systems increasingly exhibit agentic behavior, discussions of autonomy often conflate what systems are technically capable of doing w…
Do LLMs Know Their Vulnerable Scenarios?
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can by…
Delegation Intelligence in Deep Search: A Controllable Framework for Disentangled Capability Diagnosis
Deep search is becoming a core capability of modern agent systems, yet it is typically evaluated solely based on end-to-end answer accuracy…
ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
Autonomous driving under adverse weather remains a critical challenge, yet existing vision-language benchmarks mainly evaluate under standa…
Verification-Notebook Learning for Source-Aware Multimodal Misinformation Detection
Multimodal misinformation verification is challenging because misleading signals may come from different parts of a post and require differ…
Are You Still the Agent I Authorized? Earned Authority under a Fixed Ceiling for Evolving Agents
Long-lived AI agents increasingly evolve after deployment by retaining experience, acquiring skills and tools, revising workflows, delegati…
Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
Large Vision-Language Models (VLMs) now act as agents in interactive environments, where success requires coherent reasoning and decision-m…
SpecAHD: Localize to Specialize for Automated Heuristic Design in Large-Scale Routing Problems
LLM-based automated heuristic design (AHD) typically scores executable programs on complete instances or within fixed solver components. In…
Focus Is All You Need: Adaptive Goal-aware Attention Orchestration for Multi-Agent Graph Systems
Large language models (LLMs) enable autonomous agents for reasoning, planning, and tool use. Recent systems increasingly organize these age…
Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KV
Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Evic…
Offline-to-Online Creative Optimization with Generative Models and Adaptive Testing
Ad creative optimization is increasingly constrained by evaluation rather than generation. Generative models can produce many plausible cre…
Offline-Online Curriculum RL for Multimodal Reasoning
Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correc…
E-Bench: Benchmarking Multi-Step Tool-Use Agents in Real-World Product Scenarios
Large Language Models (LLMs) are increasingly deployed as agents that interact with stateful environments over multiple steps: gathering hi…
Training Language Models to Cooperate with Inference-Time Controllers
Large language model (LLM) performance increasingly depends not only on the base model, but also on the inference-time controller used to o…
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enab…
ACM: Agentic Context Management for Long Horizon Tasks
Agentic tasks are inherently long-horizon and multi-turn, constantly accumulating context through interactions with the environment. Existi…
Do Visual Features Improve Other-Initiated Repair Detection? A Dyadic Multimodal Approach
Other-initiated Self-repair, or in short Other-initiated Repair (OIR), is an essential mechanism in conversational interaction, whereby a r…
Understanding Human-like Solutions in Combinatorial Optimization via Learning and Search
Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms…
Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery
Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathway merits optimization. We for…
GOTS: Greedy Orthogonal Token Selection for High-Resolution Vision-Language Models
Modern vision-language models (VLMs) increasingly rely on dynamic or high-resolution visual encoding, producing thousands of visual tokens…
Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory
A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, th…
MemTX: Transactional Belief Commit for Stateful Agent Memory
LLM agents increasingly coordinate through persistent shared memory: one agent's write becomes another agent's premise, and eventually a to…
From Cognitive Architectures to Language Agents: A Mechanism-Level Review of Lineage, Convergence, and Migration Gaps
Memory, planning, reflection, and tool use are often compared as feature labels, obscuring the control semantics that determine how an agen…
DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models
Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusin…
EviBack: Search-Agent Reinforcement Learning via Evidence-Constrained Teacher Backoff
Reinforcement learning enables Agentic RAG systems to learn multi-turn search from verifiable outcome rewards, but all- zero rollout groups…
Grokking on the Weight-Decay Clock: A Rate Hierarchy from Softly Broken Symmetries
Delayed generalization, or grokking, remains poorly understood despite extensive empirical study. We identify an exactly solvable late-time…
Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks
Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does…
Exploring Budgeted Image Classification with Content-Sensitive Resource Allocation
The ever-growing adoption of Artificial Intelligence (AI) creates the need to deploy Deep Neural Networks in a variety of computational env…
Self-Supervised Consistency Enhanced Disentangled Learning for Neural Decoding Generalization in Brain-Machine Interface
Brain-Machine Interfaces (BMIs) provide a direct communication pathway between the brain and external devices, enabling humans to control a…
A Cyclic Adaptation-Generalization Framework with Uncertainty-Guided Self-Paced Learning for Long-Term Brain-Machine Interfaces
Brain-Machine Interfaces (BMIs), which link the brain to external devices, hold great potential in rehabilitation, human performance augmen…
The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research
Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper h…
Quantum-Inspired Evolutionary Neighborhood Search for Arrival-Departure Track Utilization Adjustment under Short-Term Disturbances
Short-term disturbances at major passenger railway stations alter train arrival and departure times as well as the release sequence of stat…
Success Is Not Self-Explanatory: Auditing Success Provenance in Agent Evaluation
A correct answer can conceal why an agent succeeded. Once agents change their information state during evaluation, correctness no longer di…
The Cost of Knowing: A Resource-Aware Protocol for Benchmarking Hallucination Beyond Static Leaderboards
On standard factuality tasks, frontier models now cluster near the top of the scale. The question is therefore shifting from how factual a…
MiSS: A Logic-Driven Explanation of Minimal Sufficient Coalitions for Point Cloud Classifiers
We present MiSS, a black-box, query-based framework for explaining 3D point cloud classifiers through perturbation-relative sufficiency rea…
Towards High-Level Semantic Intelligence
Recent advances in AI have substantially expanded its cognitive and reasoning capabilities. From the perspective of semantic complexity, th…
MemChain: Learning Interpretable Memory Traces for Memory-Augmented LLM Agents
Memory-augmented LLM agents typically answer queries by retrieving relevant memories and feeding them directly to an answer model. This ret…
Scaling GUI Agents with Visual State Transitions
We introduce State Transition Pretraining (STP) as a new scaling axis for GUI agents. During the STP stage, we continually pretrain a unifi…
Grading the Narrators: An Isnad-Rijal Framework for Claim-Level Provenance in Multi-Agent Knowledge Systems
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retr…
A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference
Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits rea…
Agent-UCT: Upper Confidence Bounds Applied to Trees for Agentic Workflow Optimization with Cost-Awareness
Optimizing agentic workflows, such as retrieval-augmented generation (RAG) pipelines, requires navigating a combinatorial space of discrete…
Falsifiable Commitment Planning for Self-Correcting Web Agents
Long-horizon web agents often go off track before final failure: a trajectory can remain locally plausible even after the current state, re…
Myopia Prevention and Control 3.0: Artificial Intelligence--Driven Risk Stratification, Proactive Monitoring, and Personalized Intervention
The convergence of artificial intelligence (AI), digital sensing, and ubiquitous computing has created an unprecedented opportunity to tran…
Integrating Factual and Normative Industrial Knowledge via Constraint-Aware Graph Attention for Process Plan Recommendation
Integrating heterogeneous industrial knowledge, including factual relations and decision constraints, remains a core challenge in industria…
Epistemic Norms for AI Safety and Alignment Research
Mainstream AI research emphasises capability growth and tolerates low failure rates when average-case performance is high. AI safety and al…
Generative Artificial Intelligence (GenAI) to convert images of queuing networks into verifiable simulation models: an open-weight LLM workflow approach
Recent work has explored the use of Large Language Models (LLMs) to automate simulation model building, typically by generating executable…
From Proprietary to Open-Source: Bridging the Distribution Gap via Multi-Agent Protocol Distillation in Agentic Search
Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet op…
Unequal Trips, Unequal Places: Diagnosing and Mitigating Delay Inequity in Autonomous Vehicle Fleet Coordination
City-scale autonomous vehicle fleet coordinators are typically optimized for aggregate travel time, yet fleet averages conceal how delay is…
Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families
Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, persevera…
Simulating Tenant Responses to Energy Policy Interventions with Transaction-Cost-Aware LLM Age
Recent studies use Large language models (LLMs) to simulate human opinions and decisions by prompting models with demographic, attitudinal,…
Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
Automatic prompt optimization (APO) has been widely adopted to adapt vision-language models (VLMs) to downstream tasks without weight updat…
Failures Reveal What Metrics Miss: An Evidence-Driven Agent for Recursive Refinement of ECG Classifiers
Deep models have substantially advanced 12-lead ECG classification, yet their refinement still relies heavily on human experts to inspect f…
From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis
Large language models increasingly solve scientific-computing tasks, but executable feedback from one problem rarely becomes durable capabi…
Making Mathematical Knowledge Explainable, Accessible and Interoperable Through Large Language Model Integration
Mathematical models are central to formalizing research problems, yet their documentation often falls short of FAIR principles. Knowledge b…
Task-Conditional Faithfulness Auditing of Multimodal LLMs for Grid Diagnosis
Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does…
LLM-Assisted Ontology Engineering and Construction of a French Legal Knowledge Graph
Maintenance regulations are complex legal texts that are difficult to exploit when addressing a specific case and challenging to integrate…
Hierarchical Group-Conditional Conformal Risk Control for Selective Prediction in Language Models
Large language models serve heterogeneous populations structured by domain, topic difficulty, and linguistic style. Conformal risk control…
TRACE-CTI: Auditable Post-Extraction Governance of TTP Claims with Knowledge Graphs
Security Operations Centers increasingly rely on automated mapping of Cyber Threat Intelligence reports to MITRE ATT&CK, yet extractor outp…
DSCH-Loss: A Dynamic Semantic Channel Objective for Deep Semantic Hashing
Semantic hashing methods for generating short binary hash codes that allow efficient approximate nearest neighbor search in high-dimensiona…
LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports
Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-wo…
SIREN: Towards End-to-End Extreme-Weather Early Warning with Experience-Grounded LLM Agents
Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather eve…
Artificial Intelligence and Innovation Ecosystem: Evolutionary Developments, Challenges, and Future Directions
The development of the Innovative Ecosystem (IE) presents a new paradigm for economic integration, collaborative advancement, and shared ac…
Efficiency Matters in Autonomous Research
AI-driven autonomous research (AR) systems are becoming increasingly effective across a broad range of tasks. Their performance, however, i…
Reason-Mediated Behavioral Models for Auditing LLM Social Simulators
Large language models are increasingly used as social simulators, including as synthetic survey respondents. Most evaluations ask whether s…
Eviction as Estimation: A Fixed-Lag Smoothing View of Test-Time Memory, and When Measuring Beats Accumulating
A language model with a bounded working memory must repeatedly decide which stored items to keep. Every deployed method decides the moment…
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams
Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rat…
Creative Integration: A Decidable Criterion of Creativity
"Integrative" solutions are widely praised but rarely defined: we lack an operational way to tell a genuine integration -- one that makes t…
Evaluating Large Language Models for Symbolic Security Protocol Analysis
Security protocol verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether Large Language Models (LLMs)…
Comparing Optimization Models for Radiotherapy Scheduling
The Radiotherapy Scheduling Problem (RTSP) involves determining an optimal schedule for patients undergoing radiation treatments, a task th…
Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B
Deploying large language models in financial-services and agentic settings requires safety classifiers that simultaneously handle prompt in…
Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review
Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In thi…
A Formal Kinetic Theory for Zeroth-Order Newton Dynamics:Stein-Corrected Hessian Estimation and Curvature--Variance Trade-offs
Zeroth-order Newton-type methods are useful when gradients and Hessians are unavailable, but they behave quite differently from first-order…
A didactical-driven teacher assistant for a dimensional modeling course
Educational chatbots powered by large language models (LLMs) show promising effects on learning outcomes, yet most systems delegate pedagog…
Quotient Tree Arithmetic: Deferred-Division Computation with Bounded Symbolic Depth and Cross-Subtree Cancellation
We introduce Quotient Tree Arithmetic (QTA), a computational substrate in which values are represented as deferred quotient pairs (N, D) wh…
Revitalizing Public Urban Places through Cultural and Political Memory: A Technological Approach with LLMs and Augmented Reality
This paper explores the intersection of memory, place, and identity, examining how new technologies, particularly Apple Vision Pro, can ill…
Masked Autoencoders Learn Perception-Relevant Representations from Resting State Neural Data
Clinical neuroprosthetics face a data bottleneck: labeled perception trials are scarce while hours of spontaneous neural activity are large…
Learning When to Reason for Text-to-SQL via SFT and DPO
Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on compl…
AI-Assisted Causal Inference and Mediation Analyses of Environmental and Psychosocial Determinants of Subjective Cognitive Difficulties in the All of Us Research Program
Short-term environmental exposures have been linked to cognitive and behavioral outcomes, although many reported associations may reflect b…
AutoCluster, AutoTopicModeling, AutoTrendAnalysis: A Complete AutoML Pipeline for Predicting Emerging Trends
Predicting emerging trends is vital for businesses, researchers, and policymakers; yet traditional approaches often lack scalability and ad…
Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS
Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of wheth…
SetGo: Metadata Readiness for Scientific AI Datasets
Scientific datasets intended for AI use require both computational readiness for model training and metadata readiness for discovery, shari…
Towards Nexus-Score: Metadata Gaps Limit Scholarly AI Attribution
Artificial intelligence systems increasingly mediate how science is found and credited. We asked whether missing metadata prevents AI syste…
MegaSlide-DiT: Memory-Centric Adaptation and Deformable Local Attention for Efficient Video Diffusion
High-resolution video diffusion models built on Diffusion Transformers (DiTs) deliver strong fidelity but quickly exhaust the memory budget…
Visible-Light Imaging Diagnosis of Neutral Particle Emission Tomography in the Tokamak Divertor: An Efficient Transformer-based Surrogate Model
Nuclear fusion has made significant progress in recent years and is expected to become one of the most important pathways to addressing glo…
RMS@CC-MMD 2026: Multimodal Misogyny Detection via Geometric Interaction and Multi-View Consensus
The proliferation of internet memes has introduced new complexities to automated content moderation, particularly in detecting misogyny. Me…
CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents
LLM coding agents operate by constructing trajectories that accumulate reasoning, tool calls, and results to enable multi-step decision-mak…
scMIR: a vision-language foundation model for single-cell light microscopy image representation
Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heter…
Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems
Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded auto…
DAMamba-UNet3D: A Parameter-Efficient Mamba State Space U-Net with Dynamic Adaptive Scan for 3D Medical Image Segmentation
We propose parameter-efficient SSM-based U-Net architectures for 3D medical image segmentation. Convolutional U-Nets afford O(n) local mixi…
An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia
Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning. Fo…
A New Kind of Adversarial Example: Measuring the Human-Model Gap, and Its Relationship to OOD Detection
Almost all adversarial attacks add an imperceptible perturbation to fool a model. We instead study the opposite: a large, clearly visible p…
Progress-conditioned Group Policy Optimization for Long-Horizon Agentic Tasks
Group-based policy optimization has been increasingly used to train large language model (LLM) agents from sparse outcome rewards by compar…
Structural Preservation Governs Data Augmentation in Deep Learning-Based Laser Speckle Material Classification
Data augmentation is routinely used to improve generalization in image classification, but the assumptions underlying standard policies are…
Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcem…
QFedPolyp: A Communication- and Inference-Efficient Federated Learning Framework for Polyp Segmentation
Background and Objective: Automatic polyp segmentation supports computer-aided diagnosis and early colorectal cancer detec- tion. Centraliz…
AI-generated Images Challenge Visual Trust in High-risk Scenarios
Rapid advances in image generation are eroding the evidentiary value of visual content in settings where authenticity can affect public saf…
Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge
Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed…
Learning to Access Computation: Accessibility Plasticity as a Principle of Adaptive Intelligence
Modern neural networks primarily adapt through parameter modification within predefined computational structures. While recent methods intr…
Post-Operative Glioma Segmentation via Loss Stabilization, Normalization and Subspace Attention
Tracking residual tumor after surgery is essential for catching recurrence early, but automating post-operative glioma segmentation remains…
Real-time Reconstruction of Human Visual Perception from fMRI
Real-time closed-loop neurofeedback based on functional magnetic resonance imaging (fMRI) has led to important scientific and clinical adva…
Hierarchical Grading in Large Language Models
We introduce Graded Large Language Models (GLLMs), an algebraic framework that equips the representation space of a transformer with a grad…
Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks
The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying comm…
Beyond Shapley: An Influence-Based Data Auditing Pipeline for LLM Alignment and Evaluation
The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality. As datasets scale, massive preference and instr…
DomainPilot: Domain-Level Loss-Guided Two-Stage Data Mixture Optimization for Efficient Language Model Fine-Tuning
The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data. Existin…
Cheap Probes Predict Expensive Training in 3D-CT Vision--Language Models
Picking the frozen image encoder for a 3D~CT vision--language model (VLM), together with the token-compression scheme on top of it, is a se…
LC-SEPLM: long-range contact-supervised adaptation for sequence-only protein representation learning
Protein language models learn transferable sequence representations. However, because they primarily model contextual dependencies along am…
Multimodal Surface EMG Hand Gesture Recognition Using Query-Based Transformers for Prosthetic Control
Hand gesture recognition via surface electromyography (sEMG) is fundamental to prosthetic control. In this field, deep learning approaches…
What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation
High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temp…
Optimizing Transformer Neural Network for Real-Time Outlier Detection on FPGAs
In this work, we explore how the inference time of a Transformer Neural Network can be efficiently optimized with applications to real-time…
FMOPF: Latent Flow Matching with Constraint-Aware Interaction Priors for AC Optimal Power Flow
AC optimal power flow determines the minimum-cost generation dispatch under nonlinear power balance constraints and is solved thousands of…
Multimodal Domain Generalization for Depression Detection: An Attention-Based BiLSTM Network with Domain-Adversarial Training
Automatic depression detection with deep learning has shown promise but often suffers from limited generalization due to domain shift arisi…
Physically Verifiable Evidence and LLM-Based Reporting for Bearing Fault Diagnosis
Trustworthy deployment of AI-based diagnosis in safety-critical mechanical systems hinges on validation: whether a prediction can be checke…
LithoFormer: A Robust Framework for Stratigraphic Inference via Transformers
Accurate geological characterization of subsurface reservoirs from well log data is essential to support projects such as carbon capture an…
OrchNAS: Orchestrated Neural Architecture Search Service for Personalised Federated Edge Intelligence
We propose OrchNAS, an energy-aware, personalised, federated edge intelligence framework that leverages a Neural Architecture Search Servic…
Hybrid Semantic and Spectral Ensemble for Robust Synthetic Image Source Attribution
The rapid advancement of text-to-image (T2I) models has necessitated robust Synthetic Image Source Attribution (SIA) methodologies. A criti…
From Hybrid Mechanistic--Data-Driven Modeling Toward Neuro-Symbolic AI: What, Why, and How
Hybrid mechanistic/data-driven models, which combine first-principles with learned components, are increasingly used in process engineering…
Agentic Autoresearch for CT Reconstruction
Comparing CT reconstruction methods fairly is labor-intensive and largely manual, and many benchmarks use idealized data. We ask whether a…
Frustratingly Simple Black-Box Adaptation of Language Models via Logit Bias
Many organizations aim to adapt language models for internal use, both to improve performance on domain-specific tasks and to address priva…
Language-Routed RAG and Direct Option Scoring for Multilingual Financial QA: DS@GT at FinMMEval
We present DS@GT's submission to FinMMEval 2026 Task 1, a multilingual financial exam question answering benchmark spanning English, Spanis…
Robustifying pathology foundation models via fine-tuning
Pathology foundation models (FMs) produce powerful tile-level representations which remain sensitive to scanner and staining variability, u…
Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests
Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Exi…
AI-interpreted Optical Scattering for Robust and Focal Depth-Aware Imaging
Optical scattering has conventionally been regarded as an impediment in imaging research due to the degradation of image quality during rec…
Multi-primitive in-memory computing for Monte Carlo tree search
Monte Carlo tree search (MCTS) enables artificial intelligence (AI) decision-making, but requires 55-300 W on conventional processors, limi…
Spatial Prediction of Soil Microplastics and Organic Matter Using Graph Attention Networks
Accurate estimation of soil microplastics and organic matter is essential to assess ecosystem health and support sustainable land use. This…
Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)
Recent advances in large language models (LLMs) have driven growing interest in using LLMs to automate test generation. Prior work commonly…
Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests
While Large Language Models (LLMs) show great promise for automating unit test generation, recent studies suggest that the quality of gener…
Controlling Embedding Spaces with Text-Conditioned Transformations
Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-sho…
Not All LLM Reasoning is Visible in the Chain-of-Thought
A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete fai…
Invariant Discovery for Networked Systems
Invariants, the relations expected to hold among measured signals of a network, underpin applications from verification to traffic generati…
Building AI That Works: ESnet's Pragmatic Approach to AI-Driven Operational Excellence
The ORBIT (Operations Responses and Business Intelligence Toolkit) project was initiated to assess agentic AI for the upcoming ESnet 7 init…
Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model
Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a sp…
HALLELUAI: A Hallucination-Aware AI System for Ultra-Realistic Image-to-Video Generation at Scale
AI-generated video is increasingly used across marketing, product storytelling, and creative workflows, yet automated; high-precision quali…
Label-free Industrial Fault Detection via Adversarial Inverse Reinforcement Learning: A System for Run-to-Failure Prognostics
Machinery fault detection (MFD) remains heavily reliant on supervised learning, which struggles with the scarcity of fault labels in real-w…
An Explicit Counterexample to Stanley's Rankwise Lower-Bound Conjecture for Differential Posets
In Problem~6 of his 1988 paper on differential posets, Stanley asked for the least possible cardinality of a fixed rank of an $r$-different…
Beyond Direct Answering: Aligning Educational LLMs as Socratic Guides via Heuristic Reinforcement Learning
Large language models (LLMs) deployed in educational settings often behave as direct answerers: they disclose target concepts in the openin…
Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline
Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged…
WCM: World-Cognition Model for Generalizable Human-Robot Interaction
Language agents can now interact fluently with users in software, but robots still struggle to bring comparable interaction to physical tas…
Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop
Large language models increasingly write both code and the tests meant to check it; coverage records what ran, not what was verified. We st…
Exact values and exact upper bounds for families of integers with arithmetic progression intersections (Erd\H{o}s Problem #272)
Let $t(N)$ be the largest $t$ for which there exist distinct sets $A_1,\dots,A_t \subseteq \{1,\dots,N\}$ such that $A_i \cap A_j$ is a non…
VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy
Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locatin…
All in One: Generative Modeling as Mean-Field Game Design
Mean-field games (MFGs) offer a unifying lens on continuous-time generative modeling: a cost tuple recovering twelve prominent models---Con…
Multi-Agent Privacy Game in Federated Learning: A Unified Mean-Field View
Federated learning enables collaborative model training across distributed clients without centralising their data, yet privacy remains a p…
MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models
Mixed-precision quantization improves the accuracy of post-training quantization by allocating higher bitwidths to sensitive layers, but ex…
Through the Bottleneck: How Multi-head Latent Attention Separates Content from Position in Language Models
Multi-head Latent Attention (MLA), introduced in DeepSeek-V2, compresses key-value pairs through a shared low-rank bottleneck (cKV), achiev…
ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation
Multilingual reasoning evaluation overwhelmingly relies on translating English benchmarks, a practice that introduces linguistic artifacts…
Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models
Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of…
Traceable LLM Reasoning for Fake-Order Fraud Detection
Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approache…
Scoping Review of AI, Metrology, and ESG in the Semiconductor Sector: Implications for Safe and Sustainable by Design (SSbD)
The semiconductor sector faces a dual transition: scaling manufacturing execution through Artificial Intelligence (AI) while satisfying str…
Poster: Rethinking Security in LLM Code Generation through Real-World Risk Scenarios
Large Language Models (LLMs) are widely used for code generation, yet their security behavior in realistic development workflows remains un…
KAYROS: An Anytime and Exact Open-Source Solver for Duration-Minimization Time-Dependent Vehicle Routing. A Technical Report and a Case Study in Human-AI Engineering
KAYROS is an open-source solver for duration-minimization time-dependent vehicle routing problems, with or without time windows (TDVRPTW, T…
A scalable online machine learning approach for Stock Recommendation
Stock recommendation systems face the dual challenge of adapting to rapidly changing market conditions while maintaining low-latency predic…
From Vibe to Code -- and Back: Lexical Oscillation in the Formation of Design Intent with Generative AI
Generative AI design tools make natural-language prompts a starting point for design, placing new articulation demands on designers. Rather…
False Prophets: On the Security of World Models in Agentic Systems
Large language models now power autonomous agents capable of complex, multi-step tasks in different environments. Accurate and reliable exe…
In-Context Learning as Implicit Policy Gradient
Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their…
Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent.…
Fashion-3DLR: A Controllable 3D Garment Generation Using Pairwise Fashion Elements for Intelligent Design
AI-generated content (AIGC) has made significant progress, with 2D generative models becoming ready-to-use tools for the digital fashion in…
BoneAgeTW2: Automated Skeletal Maturation Assessment via the Tanner-Whitehouse 2 Method, Deep Learning, and Clinical Report Generation with Distribution Curves
We present BoneAgeTW2, the first fully open-source system to automate the complete Tanner-Whitehouse 2 (TW2) clinical protocol for skeletal…
FedSLIM: Privacy-Preserving Federated MDL-Based Descriptive Pattern Mining Across Data Silos
Federated learning has achieved considerable success for predictive modelling, yet federated descriptive analytics remains largely unexplor…
Context-Aware Concept Distillation for Trustworthy Flood Prediction
Effective flood risk management relies on accurate forecasting, yet the "black box" nature of stateof-the-art Deep Learning models creates…
X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference
Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote…
What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
Contrastive vision-language models such as CLIP map semantically opposite phrases (e.g., "a dog" vs. "not a dog") to nearly identical embed…
Statistically Supported LLM Ingredient and Recipe Data Collection in Computational Nutrition
Computational nutrition needs precise ingredient data, but current databases are incomplete, inconsistent, and built for human reference ra…
FILLER: Feature Imputation via Latent Location Exploration and Retrieval
In real-world machine learning applications, incomplete observations create a fundamental challenge. Researchers have come up with several…
Online Fair Division with Budget Constraints
We study an online variant of discrete fair division under generalized assignment budget constraints. Goods arrive one at a time and must b…
Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex
We revisit the regret loss framework introduced in Park et al. (2025), which uses decision-theoretic regret as a direct loss function for t…
Patient-Agnostic Synthetic Pretraining for Efficient Patient-Specific Intraoperative 2D/3D Registration
Intraoperative 2D/3D registration aligns preoperative CT volumes with intraoperative X-ray or fluoroscopic images and is essential for imag…
On AI Safety and Security Technical Debt in Engineering AI-Enabled Systems
Artificial intelligence (AI) systems are increasingly deployed in high-stakes domains such as healthcare, autonomous driving, finance, and…
Fair Division with Strictly Increasing Valuations: A Tight Threshold for Two-Agent EF1 and PO
We study whether strictly positive marginal values restore the compatibility of envy-freeness up to one good (EF1) and Pareto optimality (P…
Explaining BiomedCLIP with Weighted Banzhaf Interactions Supported by Tree-Gram Parsing
Vision-Language Models (VLMs) are demonstrating significant capabilities in medical tasks like radiology analysis, yet providing faithful a…
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles
Activation Oracles (AOs) are language models trained to answer natural-language questions about another model's internal activations. They…
Semantic Semi-Incremental Data-Association-Free Object SLAM
Data association between landmark measurements and landmark variables has long been a central challenge in SLAM, as estimation accuracy dep…
Directional Influence Function: Estimating Training Data Influence in Constrained Learning
As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety…
Blood Pressure Estimation from PPG: A Comparative Study of Direct and ECG-Mediated Deep Learning Pipelines
Continuous cuffless blood pressure (BP) monitoring is essential for connected health systems and wearable devices, enabling early detection…
TLA$^{+}$-Bench: An Execution-Grounded Benchmark and Dataset for Natural-Language to TLA+ Specification Generation
Large language models increasingly write TLA$^{+}$ formal specifications from natural-language descriptions, but progress is hard to measur…
A Characterization of the Orthocomplement of the Tangent Space of Semiparametric Markov Models
Graphical models are ubiquitous in social and empirical science as they are intuitive and easy to use. These models belong to the broader c…
Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?
In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by l…
Do Small Models Use the Law You Give Them? Context-Injected Fine-Tuning for Legal QA in Bangladesh
A small language model can receive the governing statutory provision and still answer incorrectly. We test whether fine-tuning on examples…
Constraint-Bound Agnostic Bayesian Optimization: One Model for All Thresholds
Expensive constrained optimization problems in real-world industry design often involve constraint thresholds that are difficult to determi…
When Every Simulation Counts: Value-Based Reinforcement Learning for Accelerated Photonics Inverse Design
Photonic-crystal surface-emitting lasers (PCSELs) can combine high-power operation with narrow-divergence surface emission, but optimizing…
ATLAS: Automated Approximation of Transformers for Efficient Homomorphic Inference in One Hour
Fully homomorphic encryption (FHE) provides strong cryptographic guarantees for private inference, but deploying transformer models under F…
Token-Region Guided Cross-Attention Fusion for Multimodal Affect Interpretation
Automated analysis of multimodal content on social networks has become a critical task for understanding public sentiment and information d…
Formalizing Flag Algebras in Lean
Razborov's flag algebra method is a powerful tool for proving asymptotic inequalities in extremal graph theory, often reducing the task to…
Impute On-Demand: Adaptive Correlated Time Series Imputation for Changing Environments
Internet of Things (IoT) applications generate vast amounts of Correlated Time Series (CTS) data that often contain missing values and requ…
Choosing a Text Embedding Model: A Practical Benchmarking and Decision Framework
Choosing the right text embedding model is one of the most consequential -- and most frequently under-examined -- decisions in building a r…
Do Diagrams Help Large Language Models Reason? Evidence from Syllogistic Reasoning
Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve hu…
Novel Claim or D\'ej\`a Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking
Multimodal automated fact-checking (MAFC) verifies claims by retrieving and reasoning over external evidence. However, most existing static…
Auditing Alignment Controllability in LLMs via Political Axes
Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely mat…
Mission-Level Runtime Assurance for LLM-Assisted ISR Swarms over a Verification-Aware Fabric
Swarms of LLM-assisted autonomous robots are increasingly proposed for cooperative intelligence, surveillance, and reconnaissance (ISR) in…
Neonatal Hypoxic-ischaemic Encephalopathy Classification from the EEG and HRV Signals Using a Conformer based Masked Autoencoder
In this paper, we propose the MAEConformer, a novel self-supervised learning framework that combines the Conformer architecture with the Ma…
GTIN: A Unified Framework for Joint Event and Time Prediction in Temporal Graphs
Temporal graphs are increasingly used to model dynamic systems in diverse domains such as social networks, financial networks, and traffic…
An Unofficial FastLAS Tutorial: A Programmer's Guide
FastLAS is a scalable system for Inductive Logic Programming (ILP): you give it some background knowledge, a language bias, and a set of ex…
D3O: Dynamic Distribution Distillation for Ordinal Regression
Ordinal regression is widely used in scenarios where labels are discrete yet inherently ordered. In practice, however, ordinal labels are o…
Action from Adjacent Set in Physical Space Outperforms the Best Prediction in World Models
Controllers based on sampling and latent world models assign a predicted terminal cost to each candidate action sequence, choose the minimu…
DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory
We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifie…
Where Is the Cost of Third-Party API Routers in Agentic Software Development?
Third-party API routers have become a common layer that unifies access across increasingly diverse LLM providers. In coding-agent workflows…
Order in Desbordante: Techniques for Efficient Implementation of Order Dependency Discovery Algorithms
Science-intensive data profiling focuses on discovery and validation of various patterns in datasets. This study considers discovery of one…
Variational-Ising-Attention (VIA):TailoredAttentionMattersfor Science
Attention enables context modeling via query-key scoring with softmax normalization. Driven by industrial long-context demands, mainstream…
Extending Desbordante with Probabilistic Functional Dependency Discovery Support
Data profiling aims to extract complex patterns from data for further analysis and use that data in domains such as data cleaning, data ded…
CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation
Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse end…
Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents
Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the…
An empirical investigation into the properties of standard word embeddings
The embedding of word sequences into continuous vector spaces has been one of the most important developments in Natural Language Processin…
The Illusion of Secure LLM Code: Closing the Security Gap via Iterative Reprompting
Large Language Models (LLMs) are increasingly integrated into software development workflows, yet their ability to autonomously generate se…
An Exact Counterexample to Carlson's Associated-Prime Depth Conjecture from a Group of Order 128
In Question~3.1 of his 1995 paper on depth and transfer, Carlson asked whether the depth of a finite-group cohomology ring is always realiz…
AI Strategy: How to Choose What AI Product to Implement
Firms struggle to choose AI projects that pay off: two projects can look equally promising to smart, motivated stakeholders and yet deserve…
Escaping the Euclidean Void: Manifold-Informed Flow Matching for Sequential Recommendation
Conventional recommenders capture users' preferences by optimizing observed user-item relations, whereas continuous generative recommendati…
WISERouter: LLM Routing with Workload Budget Constraint
Large language models (LLMs) achieve impressive performance across multiple domains, but using the most capable model for every query is pr…
Outcome-Fair Restless Multi-Armed Bandits for Stochastic Deadline Scheduling
We study a restless multi-armed bandit (RMAB) problem for a stochastic deadline scheduling application. RMAB problems are solved using the…
Scale Weight Decay and Train Better
The discovery of scaling laws has motivated training neural networks on ever increasing quantities of data. This is typically done with a c…
A Few Words Go a Long Way: Language Guided Robot Policy Synthesis
While vision-language-action models have demonstrated impressive zero-shot manipulation capabilities, they remain fundamentally black box p…
Maximum Satisfiability of Simple Temporal Problems
The Simple Temporal Problem (STP) is a core framework for quantitative temporal constraints. As STP data can be inconsistent, we study MAXS…
PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis
Pathological diagnosis is inherently multi-scale, requiring the integration of global tissue architecture at low magnification with cellula…
How Context Attribution Handles What the Model Already Knows
Context attribution methods for large language models (LLMs) identify which input context contributes to the model response. Recent works s…
A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever
Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take th…
Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India. T…
Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance
We introduce Earnings25, a finance-domain benchmark for evaluating automatic speech recognition (ASR) on English-language earnings calls un…
Kalypso: Relational LLM Serving
Large language models are increasingly used as semantic operators for filtering, extracting, ranking, joining, and transforming unstructure…
TriShieldRAG: A Three-Ring Defense-in-Depth Framework Against Knowledge Corruption in Retrieval-Augmented Generation
Retrieval-Augmented Generation (RAG) lets a large language model answer questions using documents retrieved from an external knowledge base…
Limbomorphs
Artificial life systems are typically defined by a set of dynamical rules over an environment, an agent, or both, from which lifelike patte…
A Coulomb Particle Model for Learning Kernel Attention in Transformers
Randomized features provide a scalable approximation to kernel machines, but their performance depends strongly on the choice of feature di…
MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents
Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that…
Physics-Informed Neural Networks for Predicting Nitrous Oxide Flux
Nitrous oxide (N$_2$O) is the dominant ozone-depleting substance emitted in the 21st century, and the third largest contributor to anthropo…
Harnessing X-ray Absorption Spectroscopy Data through Multimodal Mining of Battery Literature
X-ray absorption spectroscopy (XAS) is central to understanding the local electronic and atomic structure of materials, yet most published…
Visible to the Court: How AI Is (and Isn't) Litigated in U.S. Federal Court Opinions
In the United States, artificial intelligence (AI) is rapidly deployed amid limited federal regulation. With courts become a recurring foru…
Embodied GPT-5.1: Evidence of a World Model?
This exploratory study examines whether a large multimodal language model, GPT-5.1, can serve as the high-level controller of a physical mo…
Understanding Tone-Dependent Inference Cost in Large Language Models
We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiment…
DuoAD: Leveraging [CLS] Dual Characteristics for Training-Free Few-Shot Anomaly Detection
Vision foundation models have enabled strong training-free anomaly detection (AD). However, most existing approaches rely primarily on inde…
SpecBox: Speculative Sandbox Scheduling for Efficient LLM Agent Serving
As LLM agents increasingly rely on the Model Context Protocol (MCP) to invoke isolated external sandboxes, disaggregated sandbox deployment…
Understanding Machine Unlearning Through the Lens of Mode Connectivity
Machine Unlearning aims to remove undesired information from trained models without full retraining from scratch. Despite recent progress,…
Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
Appending a two-word confirmation tag to a decision question -- "Is X the better choice?" versus "X is the better choice, right?" -- change…
Multimodal Semantic-Probabilistic Objectness for Open World Object Detection
Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and…
Moral Hazard in Multi-Agent Language Models
Cooperation can fail when socially valuable effort is costly, weakly observable, and mainly benefits others. Drawing on Holmstr\"om's team…
Adaptive Data Admission and Retention for Streaming Federated Learning
We study streaming federated learning with limited client memory, where newly generated training data incur time-varying sampling costs and…
SyRuP: Enhancing System-Prompt Following via Reward-Guided Prediction in LLM Decoding
Large Language Models (LLMs) are increasingly controlled through system prompts that specify roles, styles, formats, and safety requirement…
Agentic Cloud Decoys: A Deception-Driven Framework for Autonomous Intrusion Investigation
Cloud telemetry arrives at a scale that, paradoxically, makes intrusion understanding harder rather than easier. Attackers operate through…
Disentangling Semantic Attention from Structural Bias in the Attention Manifold
The empirical success of attention mechanism in Multimodal Large Language Models (MLLMs) often obscures its inherent, subtle flaws. Specifi…
HELIOS: An LLM-Driven Autonomous Indirect Trajectory Optimization Agent
Low-thrust trajectory optimization is a core technology in deep-space mission design. Indirect methods based on Pontryagin's Minimum Princi…
Capacity-Aware Deep Learning for Generalizable Traffic Volume Estimation Across Links and Cities
Network-wide traffic volume estimation typically relies on propagating measurements from fixed sensors, making performance highly dependent…
ACRL: Adaptive Control of Training-Inference Discrepancy for Stable Reinforcement Learning
Reinforcement Learning (RL) training for Large Language Models (LLMs) often suffers from instability due to the discrepancy between trainin…
MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
Recent Vision-Language Models (VLMs) have achieved remarkable success in visual understanding, driven by the growing availability of high-q…
Towards simultaneous decoding of kinetic and kinematic movement parameters during grasp and lift task by noninvasive brain imaging
Brain-machine interfaces (BMIs) can assist individuals with limited mobility, such as stroke survivors or amputees. One of the key challeng…
LU-500: A Logo Benchmark for Concept Unlearning
Concept unlearning is increasingly used to limit the reproduction of protected or unsafe visual concepts in text-to-image models. Existing…
A Case Study on the Acceptance of a Humanoid Robotic Head Employed in Three Public Spaces
Previous research has shown that a human-like robot's acceptance heavily depends on the setting in which it operates and its ability to per…
EEGForceFusion: Joint Tokenised-Continuous Representation Learning for Subject-Independent Grasp Force Decoding
Brain-machine interfaces provide a link between neural activity and external devices, enabling restoration of motor function and advancing…
Monitoring Post-Disaster Urban Recovery Using High-Resolution SAR Time Series and Unsupervised Learning: Evidence from the 2023 T\"urkiye-Syria Earthquake
Monitoring post-disaster recovery is essential for understanding how urban systems rebuild and progressively return to functionality. Howev…
Not Forgotten: Implementation and Evaluation of a Personalized Episodic Memory for the Humanoid Robot Head Kim
Social robots that rely on large language models for conversation are unable to retain information across sessions. This absence of memory…
StanceFlip: A Comprehensive Multi-Dimensional Benchmark for Multimodal Conversational Stance Flipping Forecasting
Conversational stance detection has shifted from static text analysis to dynamic multimodal modeling. However, existing benchmarks exhibit…
Every Client Is an Environment: Federated De-confounding for Spatio-Temporal Forecasting
Federated learning has emerged as a promising paradigm for spatio-temporal forecasting (STF), enabling collaborative model training without…
FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks s…
ML-based Predictive Models for Power Consumption in Virtualised O-RANs
As communication networks adopt virtualized and disaggregated architectures, achieving energy efficiency has become increasingly important…
Physics-Guided Generative AI for Property-Targeted 3D Porous Media Design
Inverse design of three-dimensional porous media is central to applications in filtration, catalysis, energy storage, fuel cells, thermal m…
A Computational Ethical Framework for Financial Digital Phenotyping for Mental Health
Ethical governance of AI-driven systems is often expressed through high-level principles and static documentation, creating a gap between r…
The Tokenizer Tax: Quantifying and Explaining the Cross-Lingual Cost of Subword Tokenization for Indian Languages
Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokeni…
Teacher Knows It Best: Spontaneous Symmetry Breaking and Tipping Points in Networked Langevin Dynamics AI Sycophancy
We formulate a statistical physics framework to model a networked stochastic dynamical system exhibiting bistability, driven by additive no…
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
Language-model agents act through structured tool calls whose arguments carry different risks. Untrusted content may safely influence an em…
DeepFaith: Evidence-Grounded LLMs for Faithful Incident Reporting in Multi-Stage APT Defense
Advanced Persistent Threats (APTs) are difficult to detect and interpret due to their multi-stage and stealthy nature. While recent autonom…
Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs
Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagn…
MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention
The quadratic cost of attention is a major bottleneck in diffusion-based video generation models. MXFP4 attention provides a promising path…
Regulating for AI Legitimacy
AI systems already govern. They rank speech and allocate attention, filter applicants and triage claims. The dominant frame for AI governan…
The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing
In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphi…
Multivariate Time Series Forecasting with Adaptive Non-Local Observables
Multivariate time series forecasting (MTSF) predicts future values of multiple variables from historical data. While quantum neural network…
DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference
Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active pe…
LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings
Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this g…
Evaluating RAG for French immigration law: a benchmark and baseline study
International recruitment in France requires navigating a layered legal framework absent from existing legal AI benchmarks. We present a pu…
ESRVS: Extreme Semi-Supervised Retinal Vessel Segmentation with a Single Annotated Image
Learning from minimal human supervision is a long-standing goal in medical image analysis, where dense expert annotations are costly. We st…
UNIFUSION: Adapting Autoregressive Language Models into Discrete Diffusion under a Unified Reverse-Rate Objective
Existing methods mainly adapt pretrained autoregressive (AR) language models to masked diffusion, whereas we directly adapt them to uniform…
DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes
While data curation for Vision Language Models (VLMs) is increasingly active, public practice for constructing pretraining mixtures remains…
Stress-Testing EEG Foundation Models for Clinical Decoding: Dataset Identity and Targeted Negative Controls
Pretrained EEG foundation models are increasingly proposed for clinical decoding, but their transfer across populations and robustness to n…
EchoBridge: Long-Tail-Aware ECG-Echocardiography Text Alignment for Echocardiography-Derived Cardiac Findings
Standardized echocardiography conclusions provide meaningful supervision for learning ECG representations of echocardiography-derived cardi…
LOCKS: Page-Local Compact Key Summaries for Efficient Long-Context Decoding
Serving large language models at long context is bottlenecked by the key-value (KV) cache, which is read in full at every decode step. Atte…
BettiSplit: Topology-Guided Privacy-Aware Split Learning Against Feature Inversion and Gradient Leakage
Split learning enables collaborative model training by partitioning neural networks across clients and servers. However, improper split pla…
EgoPlay: Event-Triggered Video Editing for Egocentric Streams
We introduce EgoPlay, an event-triggered video-to-video editor for egocentric streams, obtained by fine-tuning a pretrained V2V diffusion t…
The Visual Bottleneck: Sparse-Frame Adaptation of MLLMs for Joint Spatial-Temporal Video Grounding
Large-scale video platforms process millions of uploads hourly, requiring moderation systems that can localize when and where policy violat…
CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding
Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same…
D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models
Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that a…
Evaluating the Impact of Explainable AI on Trust in AI-Assisted Code Review
Background: Large language models (LLMs) are increasingly used to automate code review, but the reasoning behind their decisions remains ha…
Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair
Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between…
Agentic Permissions Policy Algebra for Taint Confinement in LLM Agents
Autonomous LLM agents processing mixed-confidentiality data face severe security risks from prompt injection attacks and reasoning errors.…
Sparse Autoencoders Encode Both Concepts and Functions: The Downstream Geometry of Feature Effects
The wide-scale use of sparse autoencoders (SAEs) as interpretability tools is limited by inconsistent links between SAE features and model…
A corrective agentic hybrid RAG and an operations-grounded evaluation for a scientific facility
Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical d…
Co-Learning for Missing Arbitrary Modalities in Multi-modal Classification
Multi-modal classification leverages complementary information across diverse data sources to enhance predictive performance. However, real…
Denial of Deadline: Network-Driven Accuracy Collapse in Distributed Inference Pipelines
Inference systems increasingly combine a fast path that returns predictions within the application's latency deadline together with a highe…
Efficient LLM-Generated Shuttling Compilers for Complex Trapped-Ion Architectures
Trapped-ion quantum computers rely on shuttling compilers, which cast an input algorithm into a sequence of ion-qubit movements within a gi…
DataOrchestra: Learning to Orchestrate Per-Example Curation of Pretraining Data
Pretraining data processing is critical to the downstream performance of Large Language Models (LLMs). However, many existing approaches de…
The Physics of Multi-Turn Long-Horizon Planning: From Pre-training to Post-training via Single- and Multi-Teacher On-Policy Agentic Distillation
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing mod…
KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability
Computer vision models have become highly effective for medical applications, yet their black-box nature continues to undermine clinician t…
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it…
ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domai…
Procedural Content Generation via Generative Artificial Intelligence
The attempt to utilize machine learning in PCG has been made in the past. In this survey paper, we investigate how generative artificial in…
ANSR-DT: A Neuro-Symbolic Framework for Adaptive and Explainable Digital Twins
Digital twins are increasingly used to monitor and optimize industrial systems, yet many existing frameworks remain difficult to interpret,…
Robustness and Cybersecurity in the EU Artificial Intelligence Act
The EU Artificial Intelligence Act (AIA) establishes different legal principles for different types of AI systems. While prior work has sou…
Hybrid AI-Physical Modeling for Penetration Bias Correction in X-band InSAR DEMs: A Greenland Case Study
Digital elevation models derived from Interferometric Synthetic Aperture Radar (InSAR) data over glacial and snow-covered regions often exh…
PD$^3$: A Project Duplication Detection Framework via Adapted Multi-Agent Debate
Project duplication detection is critical for project quality assessment because it helps avoid investment in repeated proposals. Existing…
LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants
Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when…
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
Ontologies and knowledge graphs require continuous evolution to remain comprehensive and accurate, but manual curation is labor intensive.…
Rethinking Prospect Theory for LLMs: Revealing the Instability of Decision-Making under Epistemic Uncertainty
Real-world decision-making often involves uncertainty expressed in linguistic rather than numerical terms, and Prospect Theory (PT) provide…
TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning
Table reasoning requires models to jointly perform comprehensive semantic understanding and precise numerical operations. Although recent l…
Decentralized Causal Discovery using Judo Calculus
We describe a theory and implementation of an intuitionistic decentralized framework for causal discovery using judo calculus, which is for…
OS-Sentinel: Towards Safety-Enhanced Mobile GUI Agents via Hybrid Validation in Realistic Workflows
Computer-using agents powered by Vision-Language Models (VLMs) have demonstrated human-like capabilities in operating digital environments…
Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering
In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE…
Multi-agent DRL-based Lane Change Decision Model for Cooperative Platooning in Mixed Traffic
Connected automated vehicles (CAVs) possess the ability to communicate and coordinate with one another, enabling cooperative platooning tha…
Bayesian-LoRA: Probabilistic Low-Rank Adaptation of Large Language Models
Large Language Models usually put more emphasis on accuracy and therefore, will guess even when not certain about the prediction, which is…
The Lattice Representation Hypothesis of Large Language Models
We propose the Lattice Representation Hypothesis of large language models: a symbolic backbone that grounds conceptual hierarchies and logi…
PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms
AI-powered people search platforms are increasingly used in recruiting, sales prospecting, and professional networking, yet no widely accep…
LiteResearcher: A Scalable Agentic RL Training Framework for Deep Research Agent
Reinforcement Learning (RL) has emerged as a powerful training paradigm for LLM-based agents. However, scaling agentic RL for deep research…
From Context to Skills: Can Language Models Learn from Context Skillfully?
Many real-world tasks require language models (LMs) to reason over complex contexts that exceed their parametric knowledge. This calls for…
GeoDecider: An Evidence-Grounded Agent for Geological Interpretation via Deliberative Reasoning
Geological interpretation infers subsurface properties and structures from indirect geophysical observations. Well-log classification provi…
From Pixels to Prompts: Vision-Language Models
When you read a paper about a new Vision-Language Model today, it can be easy to forget how strange this idea would have sounded not so lon…
Do LLMs Experience an Internal Polylogue? Investigating Reasoning through the Lens of Personas
Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often c…
What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation
Chain-of-thought (CoT) prompting enhances large language model performance, yet what drives these gains remains unclear. We study this ques…
On the Detection of Commutative Factors in Factor Graphs: Necessary and Sufficient Conditions
Exploiting the indistinguishability of objects in a probabilistic graphical model such as a factor graph is key to lifted probabilistic inf…
WIRE: Profiling Witnessed Within-Policy Instruction Collisions in LLM Agents
LLM agents are governed by long-lived prompt policies, where individually reasonable stand- ing rules can jointly govern the same pre- gene…
Universal Quantum Transformer
Classical continuous-space neural networks fundamentally struggle to lock into exact formal rules, whether mathematical, such as modular ar…
Tree-Based Formalization of Multi-Agent Complementarity in Human-AI Interactions
Complementarity is the case in which a human--AI interaction (HAI) outperforms the best prediction benchmark available among its members. A…
Some hypotheses on how chatbots work in problem-solving-driven conversations. Large Language Models as confirmation of the Innovation Illusion
We discuss the nature of chatbots as conversation partners in problem-solving. What can chatbots do and what can't they do? We develop hypo…
Artificial Intelligence for Mathematical Reasoning: An Integrated Survey of Language Models, Neuro-symbolic Systems, and Verified Discovery
Mathematical reasoning has long served as a stringent test of machine intelligence; over the past decade, it has moved from a niche problem…
ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents
Tool-using LLM agents increasingly use the Model Context Protocol (MCP) to answer from heterogeneous evidence sources, including search, AP…
Intent-Governed Tool Authorization for AI Agents
AI agents increasingly act through external tools: they read private data, construct structured payloads, submit write requests, export rec…
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. The book covers the fu…
ATOD: Annealed Turn-aware On-policy Distillation for Multi-turn Autonomous Agents
Training small language-model agents for long-horizon interactive tasks requires both fast imitation and reward-driven improvement. On-poli…
EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
This paper presents a systematic survey and conceptual synthesis of the shared measurement problem underlying large language model (LLM) ev…
Agents Don't Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents
Stateful personal agents increasingly maintain long-term user profiles, episodic memories, and reusable skills. This persistence turns conv…
Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations
Multi-agent language-model (LM) systems often determine which agents communicate, yet routing is usually treated as an implementation detai…
ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System
Recent advances in Large Language Models have fueled autonomous AI agents capable of tackling complex scientific tasks, yet existing automa…
MedFailBench: A Clinician-Built Open-Source Benchmark for Medical AI Safety Boundary Inspection
Most medical AI benchmarks measure whether a model knows the correct answer. MedFailBench asks a different question: which safety boundary…
Knowledge-Centric Agents for Workflow Generation in ComfyUI
Workflow generation in visual creation systems such as ComfyUI demands not only syntactic accuracy but also expert-level reasoning over mod…
Position: AI/ML Deepfake Research is Misaligned with AI-Generated Non-Consensual Intimate Imagery (AIG-NCII)
AI-generated non-consensual intimate imagery (AIG-NCII) is not adequately addressed in AI/ML literature regarding AI-generated media, commo…
Neuro-Symbolic Meta-Policies for Temporal Knowledge-Graph Memory under Partial Observability
Partially observable reinforcement learning requires deciding what to retain, retrieve, and forget over time. We introduce a neuro-symbolic…
Athena-Brain Technical Report: An Efficient Robot Brain for General Intelligence and Embodied Interaction
Large language models (LLMs) have demonstrated remarkable capabilities in language understanding, reasoning, and world knowledge. As embodi…
Quality Action Assurance: Multimodal Verification of Examiner Claims in VR OSCEs
Objective Structured Clinical Examinations (OSCEs) are the gold standard for assessing clinical competence, yet scoring remains vulnerable…
AdaRoPE: Not All Attention Heads Should Rotate and Scale Equally
Rotary Position Embedding (RoPE) is widely adopted in Transformers to encode positional information, yet standard implementations enforce a…
Efficient and Interpretable Body-Based Emotion Recognition with Lightweight Temporal Convolutional Networks
Body-based emotion recognition is important for real-time affective systems, but graph-based skeleton models can be computationally expensi…
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization
Multimodal large language models (MLLMs) have achieved impressive performance in multimodal emotion recognition (MER) tasks and lifted MER…
V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure
As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical. C…
Expert Behavior Prior Reinforcement Learning
Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learn…
Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model
We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-age…
TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI
Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterp…
Like a Baby: Visually Situated Neural Language Acquisition
We examine the benefits of visual context in training neural language models to perform next-word prediction. A multi-modal neural architec…
Speed Reading Tool Powered by Artificial Intelligence for Students with ADHD, Dyslexia, and Short Attention Span
This paper presents an artificial intelligence tool designed to assist students with dyslexia, ADHD, and short attention spans in processin…
Fairness Interventions in Classification: A Study on AI Explainability
This paper presents a philosophical and experimental study of fairness interventions in AI classification, centered on the explainability a…
CodexGraph: Bridging Large Language Models and Code Repositories via Code Graph Databases
Large Language Models (LLMs) excel in stand-alone code tasks like HumanEval and MBPP, but struggle with handling entire code repositories.…
Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have fou…
CausAdv: A Causal-based Framework for Detecting Adversarial Examples
Deep learning has led to tremendous success in computer vision, largely due to Convolutional Neural Networks (CNNs). However, CNNs have bee…
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
Mixture of experts (MoE) architectures have become a cornerstone for scaling up and are a key component in most large language models such…
Sign-Symmetry Learning Rules are Robust Fine-Tuners
Backpropagation (BP) has long been the predominant method for training neural networks due to its effectiveness. However, numerous alternat…
A Survey of Graph Transformers: Architectures, Theories and Applications
Graph Transformers (GTs) have demonstrated a strong capability in modeling graph structures by addressing the intrinsic limitations of grap…
Sampling Decisions: Exact Path-Space Correction, Prior Cancellation and Local-Boltzmann Guidance
How can a cheap but biased sequential, finite-horizon sampler over a discrete space be corrected so that its terminal output follows a pres…
Kinship Verification through a Forest Neural Network
Early methods used face representations in kinship verification, which are less accurate than joint representations of parents' and childre…
Computational Experiments in Number Theory
This paper presents two concrete applications of Artificial Intelligence to algorithmic and analytic number theory. Recent benchmarks of la…
Steerable Chatbots: Exploring Personalization Control Interfaces via LLM Activation Steering
Personalizing LLM responses typically requires users to articulate their preferences through prompting, which can be burdensome at cold sta…
AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use
Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can expl…
CoopReflect: Towards Natural Language Communication for Cooperative Autonomous Driving via Multi-Agent Learning
Past work has demonstrated that autonomous vehicles can drive more safely if they communicate with each other. However, this communication…
Retrieval-Augmented Generation of Ontologies from Relational Databases
Deriving OWL ontologies from relational database schemas supports semantic interoperability and downstream tasks such as knowledge graph po…
Flick: Few Labels Text Classification using K-Aware Intermediate Learning in Multi-Task Low-Resource Languages
Training deep learning networks with minimal supervision has gained significant research attention due to its potential to reduce reliance…
LEDOM: Reverse Language Model
Autoregressive language models are trained exclusively left-to-right. We explore the complementary factorization, training right-to-left at…
CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback
Acquiring high-quality instruction-code pairs is essential for training Large Language Models for code generation. While automated synthesi…
Realizing Scaling Laws in Recommender Systems: A Foundation-Expert Paradigm for Hyperscale Model Deployment
Scaling laws have been established for recommender systems, yet efficiently deploying foundation model (FM) across multiple recommendation…
PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment
As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical. While achieving high accu…
EGRA:Toward Enhanced Behavior Graphs and Representation Alignment for Multimodal Recommendation
MultiModal Recommendation (MMR) systems have emerged as a promising solution for improving recommendation quality by leveraging rich item-s…
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries
Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questi…
Loong: Synthesize Long Chain-of-Thoughts at Scale through Verifiers
Recent advances in Large Language Models (LLMs) have shown that their reasoning capabilities can be significantly improved through Reinforc…
A funny companion: Distinct neural responses to AI- versus human-attributed humor
As artificial intelligence (AI) companions become capable of human-like communication, including telling jokes, understanding how people co…
Poison to Detect: Detection of Targeted Overfitting in Federated Learning
Federated Learning (FL) enables collaborative model training among clients without centralising data, making it a widely adopted privacy-en…
From Camera-Based Sensing to Reasoning: A Comprehensive Review Toward Proactive Vulnerable Road User Safety
Ensuring the safety of vulnerable road users (VRUs), such as pedestrians and cyclists, remains a critical challenge, as conventional infras…
LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human…
Seesaw: Accelerating Training by Balancing Learning Rate and Batch Size Scheduling
Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining. Wh…
Continual Knowledge Consolidation LORA for Domain Incremental Learning
Domain Incremental Learning (DIL) is a sub-branch of continual learning that aims to address the never-ending arrival of new domains withou…
Intuitionistic $j$-Do-Calculus in Topos Causal Models
In this paper, we generalize Pearl's do-calculus to an Intuitionistic setting called $j$-stable causal inference inside a topos of sheaves.…
CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference
Large language models face significant computational bottlenecks during inference due to the expensive output layer computation over large…
VLASH: Real-Time VLAs via Future-State-Aware Asynchronous Inference
Vision-Language-Action models (VLAs) are becoming increasingly capable across diverse robotic tasks. However, these models are typically de…
GFLAN: Generative Functional Layouts
Automated floor plan generation lies at the intersection of combinatorial search, geometric constraint satisfaction, and functional design…
Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
General-purpose robotic systems operating in open-world environments must achieve both broad generalization and high-precision action execu…
AutoFed: Personalized Federated Traffic Prediction via Adaptive Prompt
Accurate traffic prediction is essential for Intelligent Transportation Systems, including ride-hailing, urban road planning, and vehicle f…
The Optimal Sample Complexity of Linear Contracts
In this paper, we settle the problem of learning optimal linear contracts from data in the offline setting, where agent types are drawn fro…
TextBridgeGNN: Pre-training Graph Neural Network for Cross-Domain Recommendation via Text-Guided Transfer
Graph-based recommendation has achieved great success in recent years. The classical graph recommendation model utilizes ID embedding to st…
Reconstructing Item Characteristic Curves using Fine-Tuned Large Language Models
Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field test…
Infinite-Precision Autoregressive Modeling for Vector Graphics and Layouts
While Transformer-based autoregressive models excel in data generation, their token discretization strategy inherently limits their precisi…
A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning
Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably…
Ordering-based Causal Discovery via Generalized Score Matching
Learning DAG structures from purely observational data remains a long-standing challenge across scientific domains. An emerging line of res…
MANGO: A Global Single-Date Paired Dataset for Mangrove Segmentation
Mangroves are critical for climate-change mitigation, requiring reliable monitoring for effective conservation. While deep learning has eme…
Physics-Encoded Inverse Modeling for Arctic Snow Depth Estimation
Accurate estimation of unobserved quantities in time-varying inverse problems remains challenging when observations are sparse and only ind…
Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis
Large language models fine-tuned via a two-stage pipeline (domain adaptation followed by instruction alignment) can exhibit non-trivial int…
Action-Sufficient Goal Representations
In offline goal-conditioned reinforcement learning (GCRL), hierarchical approaches decompose long-horizon tasks into high-level subgoal pre…
Plain Transformers are Surprisingly Powerful Link Predictors
Link prediction is a core challenge in graph machine learning, demanding models that capture rich and complex topological dependencies. Whi…
Towards Isolated Interventions via Almost Orthogonal Features in Language Models
A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in acti…
Multi-Task GRPO: Reliable LLM Reasoning Across Tasks
RL-based post-training with GRPO is widely used to improve large language models on individual reasoning tasks. However, real-world deploym…
How College Students Use AI to Navigate Course Readings: Evidence from an Eight-Week Study
College students increasingly use AI chatbots to support academic reading, yet we lack granular understanding of how these interactions sha…
LLM-Based Scientific Equation Discovery via Physics-Informed Token-Regularized Policy Optimization
Symbolic regression aims to distill mathematical equations from observational data. Recent approaches have successfully leveraged Large Lan…
The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs
Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly w…
AIFL: A Global Daily Streamflow Forecasting Model Using a Deterministic LSTM Pre-trained on ERA5-Land and Fine-tuned on IFS
Reliable global streamflow forecasting is essential for flood preparedness and water resource management, yet data-driven models often suff…
Voice-Driven Semantic Perception for UAV-Assisted Emergency Networks
Unmanned Aerial Vehicle (UAV)-assisted networks are increasingly foreseen as a promising approach for emergency response, providing rapid,…
Reverso: Efficient Time Series Foundation Models for Zero-shot Forecasting
Learning time series foundation models has been shown to be a promising approach for zero-shot time series forecasting across diverse time…
From "Help" to Helpful: A Hierarchical Assessment of LLMs in Mental e-Health Applications
Psychosocial online counselling frequently encounters generic subject lines that impede efficient case prioritisation. This study evaluates…
UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation
LiDAR-camera fusion enhances 3D panoptic segmentation by leveraging camera images to complement sparse LiDAR scans, but it also introduces…
Exact and Asymptotically Complete Robust Verifications of Neural Networks via Ising Solvers
We present an Ising-compatible framework for formal neural-network robustness verification under bounded input perturbations. For piecewise…
SafeCRS: Personalized Safety Alignment for LLM-Based Conversational Recommender Systems
Current LLM-based conversational recommender systems (CRS) primarily optimize recommendation accuracy and user satisfaction. We identify an…
DreamCAD: Scaling Multi-modal CAD Generation using Differentiable Parametric Surfaces
Computer-Aided Design (CAD) relies on structured and editable geometric representations, yet existing generative methods are constrained by…
Multi-model approach for autonomous driving: A comprehensive study on traffic sign-, vehicle- and lane detection and behavioral cloning
Deep learning and computer vision techniques have become increasingly important in the development of self-driving cars. These techniques p…
Designing Service Systems from Textual Evidence
Designing service systems requires selecting among alternative configurations -- choosing the best chatbot variant, the optimal routing pol…
A General Deep Learning Framework for Wireless Resource Allocation under Discrete Constraints
While deep learning (DL)-based methods have achieved remarkable success in continuous wireless resource allocation, efficient solutions for…
Stability of AI Governance Systems: A Coupled Dynamics Model of Public Trust and Social Disruptions
AI systems are increasingly entrenched in public governance, yet scholarship lacks formal tools to determine when deviations of public trus…
Coherent Without Grounding, Grounded Without Success: Observability and Epistemic Failure
When an agent can articulate why something works, we typically take this as evidence of genuine understanding. This presupposes that effect…
AutoWorld: Learning Multi-Agent Traffic Simulation with Self-Supervised World Models
Simulation with realistic traffic agents is essential for validating autonomous driving systems. Existing data-driven simulators learn agen…
Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments
Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-wor…
SkillSieve: A Hierarchical Triage Framework for Detecting Malicious AI Agent Skills
Agent skills combine natural-language instructions with executable code while inheriting an agent's filesystem, credential, and network acc…
Sheaf-Laplacian Obstruction and Projection Hardness for Cross-Modal Compatibility on a Modality-Independent Site
Cross-modal representations vary in how easily they can be aligned, and compatibility is generally non-transitive: two modalities may align…
Multinex: Lightweight Low-light Image Enhancement via Multi-prior Retinex
Low-light image enhancement (LLIE) aims to restore natural visibility, color fidelity, and structural detail under severe illumination degr…
How Transformers Learn to Plan via Multi-Token Prediction
While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global struct…
Bridging MARL to SARL: An Order-Independent Multi-Agent Transformer via Latent Consensus
Cooperative multi-agent reinforcement learning (MARL) is widely used to address large joint observation and action spaces by decomposing a…
Adaptive receptive field-based spatial-frequency feature reconstruction network for fine-grained few-shot image classification
Feature reconstruction techniques are widely applied for few-shot fine-grained image classification (FSFGIC). Our research indicates that o…
Principles and Guidelines for Randomized Controlled Trials in AI Evaluation
This work establishes a framework for standardizing AI evaluation RCTs (sometimes called human uplift studies). Drawing on established prac…
Are Flat Minima an Illusion?
Flat minima are an account of why deep networks generalise. However flatness is a matter of form (parameters), while generalisation is of f…
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
Recall presents a difficult choice: transformers have a linearly growing memory that slows each successive token, while linear RNNs typical…
A Cascaded Edge-Cloud Architecture for Automated Diabetic Retinopathy Screening
Diabetic Retinopathy (DR) is one of the leading causes of preventable blindness, and automated screening can help extend specialist capacit…
ChangeFlow -- Latent Rectified Flow for Change Detection in Remote Sensing
Remote sensing change detection (RSCD) localises changes between two images of the same geographic region. Most state-of-the-art methods ar…
Agentic Graph Retrieval-Augmented Generation for Auditable Commercial Registry Analysis
Public commercial registries are formally open, yet their practical analysis remains difficult because relevant facts are scattered across…
When Good Equations Get Bad Scores: Improving Symbolic Regression Through Better Parameter Optimization
Symbolic Regression (SR) plays a central role in scientific knowledge discovery by distilling mathematical equations from observational dat…
Task-Aligned Self-Supervised Learning for Medical Image Analysis: A Task-Oriented Review with Practical Design Guidelines
Self-supervised learning (SSL) is increasingly used in medical image analysis to reduce dependence on costly expert annotations by learning…
KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing
Relay and reseller APIs increasingly intermediate access to large language models (LLMs), but users have no direct way to verify that a cla…
Causal Density Functions
We study the full density ratio between a specified intervention regime $P_a$ and an observational regime $P_0$, $\rho_a=dP_a/dP_0$, under…
Enhancing MedSAM with a Lightweight Box Predictor for Medical Image Segmentation
Semantic segmentation in medical imaging is a critical yet challenging task due to data scarcity and high variability across modalities. Wh…
CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks
AI agents are reshaping the workspace, leading to drastic change of how humans work. Despite the considerable potential of human-agent coll…
$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems
As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. C…
Market Design for AI: Beyond the Copyright Binary
How can we design a market of human-generated content for use in training AI models that both enables technological progress and preserves…
Who Pays the Price? Stakeholder-Centric Prompt Injection Benchmarking for Real-world Web Agents
LLM-based web agents are increasingly deployed in real-world settings such as e-commerce, where they interact extensively with untrusted we…
Aligning Quantum Operators with Large Language Models
Can Large Language Models (LLMs) understand and reason about quantum operators? Despite their remarkable capabilities in mathematics and sy…
daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization
GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present…
Enhancing Pathological VLMs with Cross-scale Reasoning
Pathological images are inherently multi-scale, requiring pathologists to integrate evidence from global tissue architecture at low magnifi…
SwitchBraidNet: Quantisation-Aware Lightweight Architecture for Hybrid Brain-Computer Interface
Hybrid brain-computer interfaces (BCIs) that integrate motor imagery (MI) and steady-state visual evoked potentials (SSVEP) provide high-di…
TextRich: A Multi-Domain Benchmark for Detecting AI-Generated Text-Rich Images from GPT-Image-2
Text-rich images often contain privacy-sensitive, transactional, or decision-relevant information. As recent multimodal image generation mo…
Latent Confounded Causal Discovery via Lie Bracket Geometry
We study causal discovery from observational and interventional regimes when latent variables may affect the measured system. Our first alg…
Infinitesimal Causality
Interventions can be varied continuously in many causal models. Differentiating a specified smooth intervention protocol produces vector fi…
EmotionAI: A Privacy-Preserving Computational Intelligence Pipeline for Speech-Emotion-Grounded Conversational Analysis
Reviewing recorded interviews for affective cues such as composure and agitation is slow and subjective, and cloud services that could auto…
Neural Machine Translation for Low-Resource Tangkhul--English
We present a study on low-resource machine translation for the Tangkhul-English (nmf-en) language pair. Tangkhul is a severely under-resour…
LLM-Ideoplasticity: Measuring Ideological Plasticity in the Political Behavior of LLMs as a Context-Conditioned Distribution
We argue, with systematic empirical evidence, that a large language model's political ideology is not a fixed point, but a conditional dist…
PS-PPO: Prefix-Sampling PPO for Critic-Free RLHF
Reinforcement Learning from Human Feedback (RLHF) for Large Language Models increasingly relies on critic-free methods as a practical alter…
From Materials Database to Materials Bank: Assetizing Data for AI Driven Materials Innovation
Driven by high-throughput experimentation, computational modeling, and artificial intelligence (AI), materials data has expanded at an unpr…
DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning
Large language models achieve strong performance on many reasoning tasks when allowed to externalize intermediate steps as Chain-of-Thought…
From World Models to World Action Models: A Concise Tutorial for Robotics
Rather than providing an exhaustive survey, this paper presents a concise tutorial on world models and world action models for robotics. Af…
The Eticas AI Risk Taxonomy: Open Infrastructure for Operationalizing AI Audits
The rapid deployment of AI systems across high-stakes domains has created urgent demand for standardized evaluation, yet the field remains…
DOSE-I: A Multimodal Biosignal Dataset of Procedural Sedation for Endoscopy -- Technical Report
In this document, we describe characteristics and technical details of the multimodal biosignal dataset DOSE-I of procedural sedation for e…
QuantFlow: A Federated Mamba-Based Post-Transformer Foundation Model for Time-Series Forecasting
Time-series forecasting supports decisions in finance, en-ergy, transportation, public health, and industrial monitoring. Recent foundation…
Multi-Turn On-Policy Distillation with Prefix Replay
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student…
Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection…
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deep…
GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model
Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke act…
IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation
While large-scale text-to-image generative models have achieved unprecedented visual performance, their inherent reliance on multi-step ite…
Spatula: Exploring On-Demand In-Situ Interfaces and Interaction for Attribute Control
Controlling attributes is a critical step toward achieving the final creative outcome, yet current approaches fall short in supporting user…
How Do Practitioners Build SE Agents? Insights from a Mixed-Methods Study
The rise of Software Engineering (SE) agents, i.e., LLM-based agents that can understand large codebases and carry out engineering tasks wi…
The One-Word Census: Answer-Choice Conformity Across 44 Language Models
When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same…
The Hitchhiker's Guide to Monoculture
Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the s…
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal…
Anatomically Faithful but Temporally Diffuse: Auditing Attribution for Left-Ventricular Ejection-Fraction Estimation from Echocardiography
Deep video models estimate left-ventricular ejection fraction (EF) from echocardiography with near-expert accuracy, and post-hoc attributio…
Beyond Success Rate: Cost-Aware Evaluation of Offensive and Defensive Security Agents
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery…
Sparse Evidence Can Suffice: Agentic Evidence Seeking for Multimodal Video Misinformation Detection
Multimodal video misinformation detection is commonly formulated as a holistic video-understanding task, where the entire video and its ass…
SechKAN: Kolmogorov-Arnold Networks with Hyperbolic Secant Functions
In recent years, Kolmogorov-Arnold Networks (KANs) have attracted increasing attention due to their effectiveness in machine learning and s…
Operational Proto-Introspection in Looped Language Models: Process-Quality Taps, Executable Branching, and the Readout-Control Boundary
Can a language model read the quality of its ongoing computation, and can an external intervention turn that readout into better outcomes?…
Skillware: A Software Ontology and Engineering Lifecycle for Persistent Behavioral Artifacts
Agent Skills have become persistent behavioral artifacts across independent AI agent systems. They combine natural-language task specificat…
MedDDC-Eval: Diagnosis-Decoupled Evaluation of Multi-Turn Medical Consultation Agents
Evaluating multi-turn medical consultation agents requires judging the diagnostic support provided by the histories they elicit through int…
SLPO: Scaling Latent Reasoning via a Surrogate Policy
Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thoug…
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection
This work introduces G-MAD, an open-source framework that uses Arma3 to generate synchronized multi-view RGB-T data for aerial object detec…
Generative AI floods and dilutes the market for books
Generative AI can produce book-length works of fiction at near-zero cost. These books are often dismissed as low-quality ``slop'' that buye…
PhantomFill: When the Form Demands an Answer, Language Models Invent One
Language models in production do not write prose. They fill forms: JSON fields, function arguments, extraction templates. We show that the…
Adaptive Multi-Horizon Reinforcement Learning
Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement l…
Emergent Compositional Skills in Mixture-of-Experts VLAs
We consider the problem of learning compositional robot policies end-to-end from expert demonstrations, without any pre-specified notion of…
Robostral Navigate
Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains…
Error Certificates for KV-Cache Eviction via Randomized Design
Deterministic KV-cache eviction keeps the top-$k$ tokens under an importance score and deletes the rest. We prove that this design cannot k…
Neptuna: A Comprehensive Machine Learning Framework for Benchmarking Complex Multiphase Flows
Compressible multiphase flows involving shocks and material interfaces arise in applications such as bubble collapse and droplet breakup, w…
Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither…
AnthropicのCEO、オープンなAIモデルに対する見解を明示 NVIDIAなど“共同声明”との違いは?
米Anthropicのダリオ・アモデイCEOは、オープンウェイトのAIモデルに対して「禁止を提唱したことは一度もない」との声明を出した。一方、AI向けのチップの輸出などに関し、一定の制限を設けるべきとも主張している。
Anthropic’s Dario Amodei responds: doesn’t oppose open-weight models, but fears Chinese AI
Anthropic founder and CEO Dario Amodei made his views clear about open-weight models and China's growing AI capabilities.
NVIDIAやMicrosoftなど30社超、オープンAIの防御ツール共同開発の「Open Secure AI Alliance」設立
NVIDIAやMicrosoft、SpaceXAIなどは、AIオープンモデルの安全性向上とサイバーセキュリティツール開発を目指すイニシアチブ「Open Secure AI Alliance」を設立した。オープンな技術を活用してソフトウェアの脆弱性修正や防御ツールの共同開発を推進…
なぜ、Microsoft 365 Copilotは「会社の仕事を理解する」のがうまいのか?
生成AIは「答える」から「仕事を進める」ツールへと進化している。では、業務で本当に頼れるAIと、単なる生成ツールの違いはどこにあるのか。鍵を握るのは、企業固有の文脈を理解する力だ。
AIエージェントが車載アプリを動的に生成、イーソルがAIDVに向けた実験場を披露
イーソルは、ユーザーイベント「eSOL Technology Forum 2026」において、AIエージェントと人/車両が対話するための実験場となる仮想環境「eSOL AI Mobility Sandbox」を披露した。
Satya Nadella says companies that trust one AI for everything may not survive
Companies without their own models — or without a layer of AI infrastructure known as AI gateways to separate their prompts from the model…
PSA: Your Claude shared chats and Artifacts may have ended up on Google
The issue appears to have originated from Claude’s “share chat” feature, which allows users to create links that enable anyone with the ass…
Microsoft launches its first cybersecurity model, plus a new agentic cybersecurity system
Microsoft bolstered its AI cybersecurity offerings this week with the launch of its first AI security model and a new security platform.
OpenAI’s Hugging Face breach has reignited the debate over alignment and control
OpenAI's Hugging Face breach has reignited debate over AI alignment and control, exposing competing views on whether increasingly capable A…
Threads users can now chat with Meta AI in their DMs
Meta on Monday said it is rolling out its Meta AI chatbot within Threads' DMs, giving users a way to chat with the AI assistant.
Google’s AI search is rapidly becoming the default, new data shows
Google’s AI Overviews now appear in 43% of searches, underscoring how quickly AI-generated answers are becoming the default way people disc…
Power up your AI infrastructure! A first look at the Smart Systems Stage agenda at TechCrunch Disrupt 2026
At TechCrunch Disrupt 2026, the Smart Systems Stage will be where energy, infrastructure, and technology collide, covering everything from…
This $9 key physically locks your most addictive apps
This $9 NFC key requires you to physically scan it to unlock distracting apps on your phone.
Ilya Sutskever’s Safe Superintelligence partners with Nvidia to scale its AI research
After two years in stealth, Safe Superintelligence has announced a long-term partnership with Nvidia as it prepares to scale to its next ph…
2026-07-27(14件)
Enigma raises $71M to make controlling a robot as easy as adjusting the volume
The massive seed round was led by Index Ventures and Ribbit Capital, with participation from Sarah Guo's Conviction Partners.
NVIDIA、「オープンなAIセキュリティ」掲げる業界連合 Microsoftなど30社超が参加
米NVIDIAは、AIセキュリティ向けのオープンな技術を開発・共有する業界連合「Open Secure AI Alliance」を設立すると発表した。
Z世代に聞く次の流行、「AIイラスト」が1位に 「Claude Code」も上位
2026年も下半期を迎え、「Z世代」の若者たちがこれから流行ると思うものについて、市場調査会社のアスマーク(東京都渋谷区)などがインターネットでアンケートを実施した。
検索結果に「詐欺ではありません」と表示させる詐欺手口、警視庁が注意喚起 AI要約も餌食に
警視庁は、SNS型投資詐欺グループがWeb検索の仕組みを悪用し、検索結果に肯定的な情報を並べ、AI要約にも「詐欺ではありません」と表示させる手口を確認した。
MIXI、新卒エンジニア向け研修資料&動画を無料公開 「実践的なAI活用術」を12科目で紹介
MIXIは、2026年度入社の新卒エンジニア向けに実施した技術研修の資料とアーカイブ動画を公開した。全12科目を通じ、業務におけるAIの活用法などを紹介している。
NVIDIA、Microsoft、OpenAIなどがオープンモデル規制反対を表明 Anthropic従業員は「CUDAのオープンソース化が楽しみ」と皮肉
NVIDIAやMicrosoftなどの企業・団体がオープンモデル規制に反対する共同声明を発表。各社CEOが賛同する一方、Anthropic従業員は「CUDAやWindowsのオープンソース化が楽しみだ」と皮肉った。
How AI is expanding what people do at work
New OpenAI research shows how AI is expanding what workers do, with ChatGPT users taking on tasks across roles and reshaping job boundaries.
AIエージェントと共に働くリスクとは? PwCが説く「実践的ガバナンス」から考察
AIエージェントを安心・安全に活用するためのリスク管理とはどのようなものか。PwCコンサルティングは、サイバーセキュリティにとどまらない、包括的なリスク管理の必要性を説く。今回は、この話から、AIエージェントと共に働くリスクについて考察する。
Are brain waves the next unlock for physical AI?
Forget YouTube videos—frontier physical AI models need multiple camera angles, dense annotation, and soon, brain wave readings.
スマホ映像から最短1分で高精細3Dモデル、NECが生成技術を開発
NECは、スマートフォンなどの汎用カメラで撮影した映像から、高精細な3Dモデルを最短1分ほどで生成する技術を開発した。独自のAIが人物や一時的な障害物を自動で除去し、周囲の映像を基に背景を補完する。
「iPhone高騰」はこれからも続く? 中国CXMTに近づくApple、メモリ競合へのけん制が不発に終わりそうなワケ【後編】
前編「Appleはもう『メモリのお得意様』ではない? NVIDIAだけでiPhone数億台分、苦境に陥った”買いたたき王者”のいま」では、米Appleによる中華メモリメーカー・CXMTへの接近と、その対応にはあまり意味がないのでは? という疑問を述べた。後編では筆者がそう考える…
法人被害45億円、元警視庁が解説「会話もできるAI詐欺」の手口と対策
生成AIによる音声合成の進化で、経営者や上司の声が数秒で複製される時代が訪れた。元警視庁職員やAI音声開発企業のCEOなど3人の専門家が、実演で脅威を体感させながら、最新のボイスフィッシングの手口と、組織が今すぐ取るべき備えを示した。
Making sense of the panic over Chinese AI
On the latest episode of Equity, we discussed why Moonshot AI's Kimi seemed to panic Silicon Valley and Wall Street.
Hugging Face CEO calls for ‘radical transparency’ after ‘unprecedented’ OpenAI hack
"The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!"