週次AIニュース 2026-W26
対象期間: 2026-06-22 〜 2026-06-28(707 件)
トピックの推移
トピック別件数
- 研究/論文 260件
- LLM/生成AI 247件
- エージェント 132件
- 画像/動画生成 88件
- その他 50件
- ビジネス/資金調達 46件
- ロボティクス 39件
- ハードウェア/半導体 25件
- 規制/政策 6件
今週のハイライト(上位 10 件)
Previewing GPT-5.6 Sol: a next-generation model
OpenAI previews GPT-5.6 Sol, a next-generation model with stronger capabilities in coding, science, and cybersecurity, paired with its most…
OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI sy…
How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery
GPT-5 Pro helped solve a 3-year-old immunology mystery, offering insights into T cell behavior. The breakthrough could support cancer and a…
Helping build shared standards for advanced AI
OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the…
Patch the Planet: a Daybreak initiative to support open source maintainers
OpenAI introduces Patch the Planet, a Daybreak initiative helping open-source maintainers find, validate, and fix vulnerabilities with AI a…
Daybreak: Tools for securing every organization in the world
OpenAI introduces new Daybreak tools, including Codex Security and GPT-5.5-Cyber, to help organizations find, validate, and patch vulnerabi…
Samsung Electronics brings ChatGPT and Codex to employees
Samsung Electronics deploys ChatGPT Enterprise and Codex to employees worldwide, marking one of OpenAI’s largest enterprise AI rollouts.
Apple Vision Pro exec is reportedly leaving for OpenAI
Paul Meade, the Apple vice president in charge of the Vision Pro headset, is reportedly leaving the company to join OpenAI’s hardware team.
The fittest founder in the room got cancer. Here’s how he used AI to fight back.
When confronted with cancer, Connor Christou fed everything tied tied to his regime — blood results, scan data, wearable output, journal en…
全件(日付別)
2026-06-28(2件)
SoftBank’s CEO isn’t the only one with questions about Elon Musk’s orbital data center hype
Not everyone is buying Elon Musk’s vision for orbital data centers.
Apple Vision Pro exec is reportedly leaving for OpenAI
Paul Meade, the Apple vice president in charge of the Vision Pro headset, is reportedly leaving the company to join OpenAI’s hardware team.
2026-06-27(11件)
The fittest founder in the room got cancer. Here’s how he used AI to fight back.
When confronted with cancer, Connor Christou fed everything tied tied to his regime — blood results, scan data, wearable output, journal en…
Asian AI startups launch Mythos-like models as Anthropic’s export ban drags on
New models are launching in Asia that promise Mythos-like capabilities without fear of an export ban. U.S. AI labs may never recover this e…
AIモデル「ミュトス」、米国の一部組織に再提供へ 米政府が許可
米Anthropicは6月26日(現地時間)、12日から提供を一時停止していたAIモデル「Claude Mythos 5」について、米国の一部組織に限定して再提供を始めると発表した。米政府から許可を得たという。
Trump Admin releases Anthropic Mythos to be used by more than 100 US companies, agencies
Over 100 companies and government agencies are reportedly authorized to use Mythos 5, including their non-American employees.
東電出資に意欲 孫正義氏が「国内データセンター誘致」で狙うインフラ戦略
ソフトバンクグループ株主総会で、会長兼社長の孫正義氏が、将来的な目標として「純資産価値1000兆円」の展望を語った。AIインフラの最大のボトルネックである「電力確保」を巡り、子会社のソフトバンクが東京電力の次期オーナー候補に名乗りを上げている事実にも言及した。最先端データセンタ…
官民投資フィジカルAIに10.5兆円示す、「実証から実装へ」動き出す現場
2026年6月22日~26日に公開された記事の中から、MONOist編集部が厳選した今週の注目ニュースをお届けします。
OpenAI、次世代「GPT-5.6」シリーズを限定プレビュー 米政府と調整、命名は「Sol/Terra/Luna」に刷新
米OpenAIは6月26日(現地時間)、次世代AIモデル「GPT-5.6」シリーズの限定プレビューを始めた。フラッグシップの「Sol」、日常業務向けでバランス型の「Terra」、高速・低価格の「Luna」の3モデルで構成する。コーディングや科学、サイバーセキュリティの能力を高め…
OpenAI limits GPT-5.6 rollout after government request, says restrictions shouldn’t be the norm
“We don’t believe this kind of government access process should become the long-term default,” says OpenAI. “It keeps the best tools from u…
OpenAI poaches Uber India chief to lead its biggest market outside the US
The hire marks OpenAI's latest push into India, expanding offices, partnerships and hiring.
Why everyone from OpenAI to SpaceX is building their own chips (and turning up the heat on Nvidia)
Nvidia has dominated the AI chip market for years, but the era of total dependence might be ending. OpenAI just shared its plans to spice t…
It’s not about Anthropic vs. OpenAI anymore
AI models have progressed to the point where their capabilities have real political consequences. Dealing with those consequences will requ…
2026-06-26(294件)
OpenAI’s Jalapeño chip is Big Tech’s spiciest move away from Nvidia
Nvidia has dominated the AI chip market for years, but the era of total dependence might be ending. OpenAI just shared its plans to spice t…
Early Bird pricing ends tonight for TechCrunch Founder Summit
Save up to $190 on your pass to TechCrunch Founder Summit 2026. Early Bird pricing ends today, at 11:59 p.m. PT, after which rates increase…
Previewing GPT-5.6 Sol: a next-generation model
OpenAI previews GPT-5.6 Sol, a next-generation model with stronger capabilities in coding, science, and cybersecurity, paired with its most…
「AIを使うと他社と似てしまう」課題をどう乗り越える? 「プロダクトの差別化」の要点
生成AIの普及は製品の没個性化や、個人の生産性向上によるチームの分断という課題を生んでいる。米Figmaはカンファレンスで、AI出力を人間が微調整する「素材」として扱う手法を提示。自社ルールを組織全体で共有する仕組みを実装した。個人の暗黙知を資産化する取り組みは、現場の属人化を…
防衛省は“認知戦”にどう挑む ウクライナ脅かすAIフェイク、偽アカウントへの対応は 分析資料を公開
防衛省は6月26日、「防衛力変革推進本部」での議論に関する資料を公表した。偽情報で相手の判断をゆさぶる「認知戦」への対応方針として、戦略的な情報発信機能やAI活用、情報関連機能の強化を打ち出した。
Detecting and Controlling Sycophancy with Cascading Linear Features
Interpreting and controlling model behaviors through activation steering methods requires many pairs of contrastive samples that clearly ex…
Life After Benchmark Saturation: A Case Study of CORE-Bench
When a benchmark's accuracy saturates, it is often retired and replaced with a more challenging version. We show that this approach privile…
Refusal Lives Downstream of Persona in Chat Models
Linear directions in activation space have been identified for both refusal and persona traits in instruction-tuned chat models, but the tw…
AlgoEvolve: LLM-driven Meta-evolution of Algorithmic Trading Programs
Recent work shows that Large Language Models (LLMs) can act as semantic mutation operators for the evolutionary discovery of programs and p…
Agentic Analysis for Agentic Infrastructure: An LLM-Powered Pipeline for Comparative Governance of DAO and Corporate AI Protocols
As AI agent protocols proliferate, the governance structures shaping their interoperability standards remain empirically underexamined. We…
Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking
Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-ev…
Accelerating Skill Assessment in Chess: A Drift-Diffusion-Enhanced Elo Rating System
Rating systems such as Elo serve as the gold standard for matchmaking in competitive chess. However, they inherently suffer from response l…
Governing Actions, Not Agents: Institutional Attestation as a Governance Model for Autonomous AI Systems
Autonomous AI agents may begin to perform consequential, irreversible actions such as clinical prescribing and production software deployme…
COrigami: An AI Pipeline for Co-Designing Flat-Foldable Visually Recognisable Origami
While generative AI has achieved remarkable success in solving problems with verifiable solutions, generating physical art that satisfies b…
The Verification Horizon: No Silver Bullet for Coding Agent Rewards
A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this intuition is being inve…
How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks?
Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, ye…
What We are Missing in Multimodal LLM Evaluation?
Multimodal large language models (MLLMs) can process diverse inputs, e.g., text, images, audio, and video, and generate textual responses.…
OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents
Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across…
Instruction Bleed: Cross-Module Interference in Prompt-Composed Agentic Systems
Practitioners of prompt-composed agentic systems report a recurring failure mode: editing one prompt module silently shifts the behavior of…
Accelerating Returns and the Qualitative Engine for Science
Ray Kurzweil described a thesis of accelerating returns, which is the most influential narratives in discussions of technological progress.…
Narration-of-Thought: Inference-Time Scaffolding for Defeasible Ethical Reasoning in Large Language Models
Standard chain-of-thought on moral dilemmas exhibits two failure modes: stakeholder collapse (the trace names at most one party with a stak…
Geometry-Aware MCTS for Extremal Problems in Combinatorial Geometry
We study certain extremal problems in combinatorial geometry that ask about configurations of points in an $n \times n$ grid that satisfy s…
When Agents Meet Electric Bus Fleet Operations: Pricing Behavior, Trade-offs, and Policy Implications in an Aggregator Framework
Agentic systems are changing how complex operational tasks are coordinated, introducing a new paradigm for connecting heterogeneous data so…
Unbiased Canonical Set-Valued Oracles Via Lattice Theory
A non-agentic "oracle" AI that estimates probabilities of future events faces a self-reference problem: once its answer is learned and acte…
Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data
Researchers increasingly use text classification--supervised models or large language models--to measure constructs from natural language,…
Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law
Sphere neural networks have achieved symbolic level syllogistic reasoning without training data, raising the question of where the limit of…
MKG-RAG-Bench: Benchmarking Retrieval in Multimodal Knowledge Graph-Augmented Generation
Retrieval-augmented generation (RAG) over knowledge graphs has emerged as a promising approach for grounding large language models, yet exi…
auto-psych: Automating the science of mind using agent-driven theory discovery and experimentation
AI-based scientific automation is increasingly possible by using agents to generate hypotheses, design experiments, and analyze data. Data…
Clinical Harness for Governable Medical AI Skill Ecosystems
Medical AI remains organized around isolated models, whereas clinical care requires accountable capabilities that persist across time. We p…
Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation
Large reasoning models (LRMs) take longer on harder problems, just as humans do. This surface similarity hides an opposite pattern within i…
NeuraDock Visual Cognitive Load Agent Tutorial: A Quality-Gated Open-Source EEG Workflow for Alpha Dynamics and Real-Time Applications
This tutorial paper provides a step-by-step, reproducible walkthrough of NeuraDock Agent, an open-source EEG agent focused on Alpha dynamic…
Boundary-Aware Context Grounding for A Low-Channel EEG Agent
Large language models (LLMs) can make scientific software easier to use. However, a general model does not automatically know which measure…
Radical AI Interpretability
We develop a framework for interpreting AI systems as agents, drawing on the philosophical tradition of radical interpretation and the tool…
PMDformer: Patch-Mean Decoupling Information Transformer for Long-term Forecasting
Long-term time series forecasting (LTSF) plays a crucial role in fields such as energy management, finance, and traffic prediction. Transfo…
Explainable Ensemble-Based Machine Learning Models for Detecting the Presence of Cirrhosis in Hepatitis C Patients
Hepatitis C is a liver infection caused by a virus, which results in mild to severe inflammation of the liver. Over many years, hepatitis C…
EvoOptiGraph: Weakness-Driven Coevolution via Graph-Based Structural Generation for Optimization Modeling
Automating optimization modeling from natural language with large language models (LLMs) faces two key challenges. First, training corpora…
A Multi-Level Validation and Traceability Framework for AI-Generated Telescope Scheduling Decisions
With the gradual introduction of AI into telescope scheduling, AI-based decision-making has shown advantages in handling complex multi-cons…
Content-Based Smart E-Mail Dispatcher Using Large Language Models
Email communication has become an integral part of personal and professional life, but handling its vast volume is still a significant issu…
LLM-based Models for Detecting Emerging Topics in Service Feedback
Enhancing the analysis of service feedback is essential for public sector organizations, particularly tax administrations, where trust and…
Autoformalization of Agent Instructions into Policy-as-Code
Agent safety in high-stakes domains requires formal policy enforcement, but most existing approaches either rely on probabilistic guardrail…
SKILL-DISCO: Distilling and Compiling Agent Traces into Reusable Procedural Skills
Agents often repeatedly solve similar task instances from scratch, leading to unnecessary reasoning cost and long execution traces. Prior w…
NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research
Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing wo…
Do Safety Guardrails Need to Reason? LeanGuard: A Fast and Light Approach for Robust Moderation
In order to screen a prompt or a response, the recent guardrail methods generate a chain-of-thought (CoT) before they issue a verdict. This…
Kalman Prototypical Networks for Few-shot Fault Detection in Combined Cycle Gas Turbines
Combined-cycle gas turbines (CCGTs) play a key role in modern power generation, offering both high efficiency and reduced environmental imp…
LithoDreamer: A Physics-Informed World Model for Multi-Stage Computational Lithography
As semiconductor technology nodes scale, computational lithography is essential for ensuring yield and performance. However, lithography is…
A Latent ODE Approach to Spatiotemporal Modeling of Cine Cardiac MRI
Cardiac magnetic resonance imaging (CMR) captures rich spatiotemporal information about ventricular structure and motion, but conventional…
Socratic agents for autonomous scientific discovery in high-dimensional physical systems
The automation of scientific discovery has reached an inflection point. While AI systems now operate instruments, optimize parameters and g…
Scientific discovery as meta-optimization: a combinatorial optimization case study
Scientific discovery is fundamentally an optimization problem, defined by a vast "state space" of theories and experiments, and an evaluati…
EGG: An Expert-Guided Agent Framework for Kernel Generation
High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but t…
ResilPhase: Plug-and-Play Phase Mapping and Noise-Resilient Macro-Trajectory Extrapolation for Diffusion Acceleration
The adoption of powerful diffusion models is hindered by their significant inference latency. Recent ``cache-then-forecast'' schemes allevi…
Memory Depth, Not Memory Access: Selective Parametric Consolidation for Long-Running Language Agents
Long-running language agents need more than memory access. Retrieval systems can fetch past facts at query time, but they do not decide whi…
KARLA: Knowledge-base Augmented Retrieval for Language Models
We propose a new method that allows an LLM to automatically pull in factual knowledge from a knowledge base during token generation. This m…
Computational Analysis of Heart Rate Variability in Healthy Adults
Heart Rate Variability (HRV) analysis is a key indicator of cardiac physiological state and aids in disease diagnosis. However, research on…
The Capability Frontier: Benchmarks Miss 82% of Model Performance
Existing benchmarks typically report accuracy for a single model on a single run. This systematically understates real-world LLM capabiliti…
Context-Aware Synthesis of Optimization Pipelines for Warehouse Optimization
Order fulfillment in manual picker-to-goods warehouses involves interconnected decisions such as item assignment, order batching, and picke…
LCAi: Life Cycle Assessment with big data fusion and retrieval-augmented generation-assisted interpretation
The interpretation phase of life cycle assessment often lacks structured mechanisms for translating quantified improvement opportunities ad…
AgentX: Towards Agent-Driven Self-Iteration of Industrial Recommender Systems
Recommendation algorithm iteration is moving from an artisanal, engineer-bound process toward an industrialized research loop, but this tra…
TAVR-VLM: Risk-Conditioned Causal Grounding for Hallucination-Resistant Report Generation
Transcatheter Aortic Valve Replacement (TAVR) planning requires meticulous multimodal reasoning. However, adapting Multimodal Large Languag…
A Pipeline for Generating Longitudinal Synthetic Clinical Notes Using Large Language Models
Synthetic data is increasingly used to enable the development and evaluation of AI systems in domains where access to real-world data is re…
Generative Retrieval via Diffusion Transformer with Metric-Ordered Sequence Training and Hybrid-Policy Preference Optimization
Embedding-based retrieval ranks items by their similarity to a query in a shared vector space and usually aims to return the highest-scorin…
Learning to Recover Task Experts from a Multi-Task Merged Model
Multi-task model merging aims to consolidate several task-specific experts into a unified model, yet static merging consistently suffers fr…
Diagnosing Task Insensitivity in Language Agents
Large language models can serve as capable long-horizon agents, but their out-of-distribution (OOD) generalization remains weak. We identif…
Where Do CoT Training Gains Land in LLM based Agents?
Chain-of-thought (CoT) reasoning is widely used in language-model agents, but prior work has shown that verbalized CoT is not always faithf…
Look-Before-Move: Narrative-Grounded World Visual Attention in Dynamic 3D Story Worlds
As embodied AI and world models increasingly operate in dynamic 3D environments, visual perception must move beyond passively interpreting…
Einstein World Models
Does intelligence require the ability to reason about phenomena beyond direct experience? It is natural to suspect that some complex though…
Adaptive Utility driven Resource Orchestration for Resilient AI (AURORA-AI)
Modern AI systems are increasingly deployed under non-stationary computational, demographic, and operational conditions in which static res…
Semantic Early-Stopping for Iterative LLM Agent Loops
Multi-agent large language model (LLM) loops, for example a Writer that drafts and a Critic that revises, are almost always terminated by a…
How to evaluate clustering with ground truth?
External indexes can be used for cluster evaluation when ground truth is available. We review the most common external validity indexes foc…
Joint Learning of Experiential Rules and Policies for Large Language Model Agents
For LLM agents in multi-step interactive environments, a key challenge is to make effective use of accumulated interaction experience. Exis…
OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
Root cause analysis (RCA) poses a holistic test of LLM agentic capabilities, such as long-context understanding, multi-step reasoning, and…
TOPS: First-Principles Visual Token Pruning via Constructing Token Optimal Preservation Sets for Efficient MLLM Inference
Multimodal large language models (MLLMs) have achieved strong multimodal reasoning capabilities, but their efficiency is limited by the lar…
A Process Harness for Uplifting Legacy Workflows to Agentic BPM: Design and Realization in CUGA FLO
We introduce the process harness, a new mechanism for uplifting legacy workflows into Agentic Business Process Management (Agentic BPM) wit…
Vulnerability of Natural Language Classifiers to Evolutionary Generated Adversarial Text
Deep learning models have achieved impressive performance across various fields but remain vulnerable to adversarial inputs, particularly i…
Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with hum…
EO-WM: A Physically Informed World Model for Probabilistic Earth Observation Forecasting
Earth Observation (EO) forecasting aims to predict future Earth surface dynamics from satellite observations under changing meteorological…
Simulation-based inference for rapid Bayesian parameter estimation in epidemiological models: a comparison with MCMC
Mechanistic epidemiological models are widely used to support infectious disease forecasting and public-health decision making. Bayesian ca…
Prompt Injection in Automated R\'esum\'e Screening with Large Language Models: Single and Multi-Injection Settings
Large language models (LLMs) are increasingly used to screen and rank job applicants, creating incentives for candidates to strategically m…
When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models
Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show th…
Language-Based Digital Twins for Elderly Cognitive Assistance
Digital twins have emerged as a promising paradigm for personalized healthcare, enabling modeling of individual behavior and health traject…
Benchmarking Open-Weight Foundation Models for Global AI Technical Governance
Large language models (LLMs) are increasingly deployed in artificial intelligence (AI) governance analysis across national and internationa…
Know2Guess: A Contamination-Aware Multi-Zone Benchmark for Knowledge-Boundary Evaluation in Large Language Models
Reliable evaluation of large language models should separate supported answering from unsupported guessing without conflating either with d…
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training
Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but th…
Investigating LLM's Problem Solving Capability -- a Study on Statics Questions
Large Language Models (LLMs) have rapidly influenced many aspects of society, particularly education, due to their demonstrated ability to…
Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare
Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then as…
Context Recycling for Long-Horizon LLM Inference
Large language models (LLMs) exhibit strong capabilities in short-context reasoning but degrade in performance over long conversational hor…
Reducing Conversational Escalation in Large Language Model Dialogue with Nonviolent Communication Constraints
Large language models (LLMs) are increasingly used in emotionally charged situations involving interpersonal conflict, frustration, and dis…
Low Resource Multimodal Translation of Nepali Spoken Words into Emotion-Conditioned Sign Language Avatars
Sign language communication systems, that integrate emotional expression remain underexplored, particularly for low-resource languages. Thi…
Generative AI and Copyright Infringement: A Legal-Technical Analysis of AI Music Generation Systems Under 17 U.S.C. Title 17
Generative artificial intelligence (GenAI) has enabled users to synthesize music with text prompts, combining copyrighted lyrics, AI-compos…
From Lexicon to AI: A Structured-Data Pipeline for Specialized Conversational Systems in Low-Resource Languages
Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive t…
Dream machine -- the next creative economy
We examine the structural transformation of creative industries under generative artificial intelligence, drawing on 374 primary sources sp…
A Multi-Layer AI Framework for Information Landscape Analysis
This paper proposes a multi-layer AI framework for information landscape analysis in the context of information disorder. Rather than treat…
Divergent Recommendations, Convergent Diagnoses: Cross-Provider Failure-Mode Convergence in AI Commercial Recommendation
A brand whose customers use both ChatGPT and Claude for product recommendations faces a strategic choice: a single optimization playbook, o…
The Governance Inversion Hypothesis: Why More AI Regulation May Produce Less Organisational Control
This paper introduces the Governance Inversion Hypothesis (GIH) to explain a growing paradox in artificial intelligence (AI) governance: un…
The Open Source Economic Index of AI Adoption and Capability
We work towards measuring both AI adoption and the capability of AI to perform discrete labor tasks across various occupations. To measure…
Dot-Flik: A Scalable Edge AI Architecture for Distributed Insect Monitoring
Global insect population declines necessitate scalable, continuous monitoring systems, yet existing vision-based solutions remain constrain…
Privacy-Aware Agent Collaboration for Dynamic VR Slice Management in 6G SD-RAN
Ultra-low latency and high throughput are required for Virtual Reality (VR) services in 6G networks, which presents critical challenges for…
Geometric Fairness-Aware Routing for Federated Edge Networks
Emerging 6G and edge-intelligent networks require effective and balanced routing algorithms among varied and spatially distributed devices.…
Thinking Like a Scientist? A Structural Study of LLM-Generated Research Methods
Large Language Models (LLMs) are increasingly used to guide research methodology, yet their default methodological tendencies under minimal…
Multiscale Exit-Join Dynamics: Tactical Consensus and Strategic Coalition Formation
This paper develops a multiscale model of coalition formation in which strategic exit-and-join decisions are coupled with tactical consensu…
Unsupervised Memory-Enhanced Video Transformers: Obstacle Detection for Autonomous Agricultural Rover
While autonomous rovers have become indispensable to precision farming, achieving consistent operational safety remains a critical challeng…
Reducing Redundancy in Whole-Slide Image Patching for Scalable Indexing and Retrieval
The rapid growth of digital pathology has created an urgent need for efficient indexing and retrieval of whole slide images (WSIs). This ne…
Neural Architecture Search for Generative Adversarial Networks: A Comprehensive Review and Critical Analysis
Neural Architecture Search (NAS) has emerged as a pivotal technique in optimizing the design of Generative Adversarial Networks (GANs), aut…
LCG: Long-Context Consistent Image Generation with Sparse Relational Attention
Recent image generation models achieve impressive quality in single-image synthesis, but often fail to maintain consistency across sequenti…
KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction
While WGS-based AMR prediction has reached high accuracy, existing models lack a mechanism to ground neural attributions in established bio…
LiMoDE: Rethinking Lifelong Robot Manipulation from a Mixture-of-Dynamic-Experts Perspective
Building a generalist robot that can leverage prior knowledge for continuous task adaptation remains a significant challenge. Previous work…
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models
Multimodal Large Language Models (MLLMs) have recently made remarkable progress in unifying vision-language understanding and reasoning, es…
Statistical and Structural Approaches to Algorithmic Fairness
Modern machine learning systems have outgrown their origins as isolated predictive constructs, evolving into complex socio-technical archit…
CyberChainBench: Can AI Agents Secure Smart Contracts Against Real-World On-Chain Vulnerabilities?
We present CyberChainBench, a benchmark for evaluating LLM-based agents on smart contract security across three complementary tasks: vulner…
Lacuna: A Research Map for Machine Learning
Lacuna is a research map for machine learning that uses LLMs to turn papers and scholarly metadata into markdown summaries, concept element…
A multi-task spatiotemporal deep neural network for predicting penetration depth and morphology in laser welding
In laser penetration welding, the assessment of penetration state and weld seam morphology plays a crucial role in determining the weld qua…
From Clicks to Intent: Cross-Platform Session Embeddings with LLM-Distilled Taxonomy for Financial Services Recommendations
Sequential user behavior modeling is widely adopted in industrial recommender systems; however, significant gaps remain in financial servic…
TEMPO-Diffusion: Temporally Exposed Malicious Poisoning of Diffusion Models
Noise-based backdoor attacks on diffusion models typically rely on input-time trigger injection, untargeted activation, and out-of-distribu…
SSM Adapters via Hankel Reduced-order Modeling: Injection Site Determines Task Suitability in Long-Context Fine-Tuning
While parameter-efficient fine-tuning (PEFT) typically targets attention projectors, its efficacy for tasks requiring sequential state accu…
The Red Queen G\"odel Machine: Co-Evolving Agents and Their Evaluators
Self-improving agents are state-of-the-art (SOTA) on agentic coding benchmarks and have recently been extended to general domains. However,…
Parametric Generalized Adaptive Moment Features (PG-AMF) for Bearing Fault Diagnosis and Machine Health Monitoring
Accurate fault diagnosis of rolling element bearings in rotating machinery is considered essential for ensuring industrial safety and enabl…
EVOM: Agentic Meta-Evolution of Actor-Critic Architectures for Reinforcement Learning
In actor-critic reinforcement learning, network architectures are typically manually designed. Automating this design is challenging becaus…
Hybrid privacy-aware semantic search: SVD-truncated document geometry and CKKS-encrypted query reranking under a restricted threat model
Dense embeddings power semantic search and retrieval-augmented generation, but embedding-inversion attacks can reconstruct source text from…
Charting the Growth of Social-Physical HRI (spHRI): A Systematic Review Pipeline Augmented by Small Language Models
Social-physical human-robot interaction (spHRI) has grown rapidly across robotics, human-computer interaction, human-robot interaction, and…
SOLAR: AI-Powered Speed-of-Light Performance Analysis
How fast could a deep-learning model run on target hardware, and how far is today's implementation from that limit? These questions are cen…
Sampling sea state using a diffusion model
Sea state prediction is essential for operational maritime applications and coupled earth system modeling, yet current spectral wave models…
Deterministic Pareto-Optimal Policy Synthesis for Multi-Objective Reinforcement Learning
Real-world decision-making often requires balancing multiple conflicting objectives, a challenge that standard Reinforcement Learning (RL)…
Beyond Feedforward Networks: Reentry Neural Systems as the Fundamental Basis of Subjecthood and Intrinsic Safety of Next-Generation AGI
We propose a complete architectural blueprint for safe artificial general intelligence based on a closed reentry loop (D I cycle). In contr…
CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation
Long-horizon, contact-rich complex manipulation tasks, such as seating a GPU into a PCIe slot, demand both millimeter high precision and ou…
Play2Perfect: What Matters in Dexterous Play Pretraining for Precise Assembly?
Multi-fingered robots promise the speed and dexterity of human hands, yet challenging problems such as precise assembly have remained out o…
ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence
Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but th…
AXLE: A Cloud Infrastructure for Lean 4 Theorem Proving Utilities
We present AXLE (Axiom Lean Engine), a cloud service for Lean 4 proof manipulation, extraction, and verification. Recent progress in AI for…
WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation
A robot working alongside people must reason about what they have done, in what order, and with what intent. Video carries the spatial layo…
Closing the Loop to Discover Psychological Theories with an Automated Cognitive Scientist
Across the sciences, autonomous systems are increasingly being used in closed-loop discovery, proposing new theories and designing and runn…
ProvenAI: Provenance-Native Traces of Evidence in Generated Answers
Retrieval-augmented systems routinely present citations alongside generated answers, yet a citation does not confirm that the corresponding…
Active Adversarial Perturbation-driven Associative Memory Retrieval for RGB-Event Visual Object Tracking
RGB-Event tracking improves localization robustness by fusing RGB appearance textures and dense temporal motion cues from event sensors. Wh…
3D Spatial Pattern Matching
Spatial pattern matching is the process of matching query entities and constraints with database entities and relations. It has many applic…
Localizing RL-Induced Tool Use to a Single Crosscoder Feature
Fine-tuning through RL reshapes the internal representations of language models to enable agentic behaviors such as tool use, yet the mecha…
Retrieval-Warmed Energy-Based Reasoning: A Five-Arm Ablation Methodology for Diffusion-as-Inference on Structured Reasoning Tasks
Warm-started diffusion samplers accelerate iterative inference, but it is rarely clear which part of the pipeline carries the gain. We stud…
Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents
Recent work (2024 to 2026) has converged on a strategy for defending tool-using LLM agents against indirect prompt injection: rather than t…
Speaking Numbers to LLMs: Multi-Wavelet Number Embeddings for Time Series Forecasting
Large language models (LLMs) are attractive for context-aware time series forecasting because they can integrate heterogeneous textual sign…
An Empirical Study of LLM-Generated Specifications for VeriFast
Static verification tools can assure industrial scale software, but require significant human labor to write specifications. This is partic…
Evaluation-Strategy Gap in Fault Diagnosis of Deep Learning Programs
Deep Learning (DL) programs can fail during training for many reasons, and diagnosing the cause is a costly and time-consuming maintenance…
Temporal Validity in Retrieval Memory: Eliminating Stale-Fact Errors for AI Agents over Evolving Knowledge
Retrieval-augmented generation (RAG) gives agents access to accumulated knowledge, but has no model of time. When a fact changes (e.g., a f…
Multipath Adaptive Gated Bottleneck Latent ODE with Raman Data Fusion for Cell Culture Process Forecasting
Mammalian cell-culture processes underpin the manufacture of many biopharmaceuticals, yet keeping a run on track is hard: critical process…
The Inattentional Gap: Task-Conditioned Language and Vision Models Omit the Safety-Critical Signals They Can Otherwise Report
AI safety is evaluated by how reliably a model detects the hazards it is told to find, yet accidents often arise from the hazard no one spe…
\textsc{DiARC}: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models
The Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) contains tasks that require summarizing patterns from limited grid…
VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation
Recently, zero-shot text-to-speech (TTS) has enabled high-fidelity and expressive speech synthesis, but it often fails to imitate unseen sp…
From Hallucination to Grounding: Diagnosing Visual Spatial Intelligence via CRISP
Current VLM evaluations often conflate language priors with genuine spatial reasoning. To address this, we introduce CRISP, a novel structu…
CascadeFormer: Depth-Tapered Transformers Motivated by Gradient Fan-in Asymmetry
Deep Transformers are composed of uniformly stacked residual blocks, yet their deepest layers often add little value. We present two effici…
Perception, Verdict, and Evolution: Hindsight-Driven Self-Refining Forensics Agent for AI-Generated Image Detection
The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the…
SpaceRipple: Lightweight Semantic Delivery for Mission-Oriented LEO Earth Observation Satellite Networks
Earth observation satellite networks generate massive volumes of high-resolution imagery, whereas inter-satellite and downlink resources re…
scBench-Long: Verifiable Benchmarking of Long-Horizon Single-Cell Biology
Single-cell studies require analysts to convert raw measurements into specific biological claims through multi-step workflows and integrati…
IDEA: Insensitive to Dynamics Mismatch via Effect Alignment for Sim-to-Real Transfer in Multi-Agent Control
Complex multi-agent control tasks remain challenging for traditional rule-based and model-based approaches, motivating the adoption of lear…
SharQ: Bridging Activation Sparsity and FP4 Quantization for LLM Inference
Low-bit floating-point formats and semi-structured sparsity are increasingly supported by modern accelerators, yet combining them for LLM a…
HiLSVA: Design and Evaluation of a Human-in-the-Loop Agentic System for Scientific Visualization
Large language model (LLM) agents enable natural language interaction for scientific visualization (SciVis). Still, prior systems have esse…
Discovering Millions of Interpretable Features with Sparse Autoencoders
Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interp…
Agents That Know Too Much: A Data-Centric Survey of Privacy in LLM Agents
Large language model agents increasingly query databases, search document collections, call external APIs, remember past interactions, and…
CAT-Q: Cost-efficient and Accurate Ternary Quantization for LLMs
In this paper, we present CAT-Q, Cost-efficient and Accurate Ternary Quantization, for compressing and accelerating LLMs. Unlike existing s…
LAMP: Lane-Aligned Motion Primitives for Feasible Trajectory Prediction
Motion forecasting is essential for autonomous driving systems to enable safe decision-making and planning in complex driving scenarios. Wh…
Zero-Shot Size Transfer for Neural ODEs on Sparse Random Graphs: Graphon Limits and Adjoint Convergence
Graph Neural Differential Equations (GNDEs) model continuous-time graph dynamics by parameterizing Neural ODE velocity fields with Graph Ne…
TGHE: Template-based Graph Homomorphic Encryption for Privacy-Preserving GNN Inference in Edge-Cloud Systems
Existing homomorphic encryption (HE)-based GNN systems adopt a graph-centric paradigm that couples per-query cost to global graph size, lim…
Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization
Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos…
Beyond Logical Forms: LLM-Extracted Patterns for Fallacy Classification
In today's fast-paced information era, logical fallacies, defined as defective patterns of reasoning, inevitably contribute to the growth o…
Learning Motion Feasibility from Point Clouds in Cluttered Environments
Motion feasibility prediction plays a central role in robotics, particularly in task and motion planning and manipulation. A major bottlene…
Algorithmic Foundations of Deep Learning: Complexity-Theoretic Rates and a Characterization of Universal Approximation
Feedforward neural network (NN) expressivity is typically studied by emulating optimal basis-expansion schemes. While powerful, this perspe…
MLFFM-SegDiff: A Multi-Level Feature Fusion Diffusion Model for Skin Lesion Segmentation
Skin lesion segmentation is a key task in computer-aided dermatological diagnosis, where accuracy directly impacts downstream analysis and…
Robust Onion: Peeling Open Vocab Object Detectors Under Noise
The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity…
Anatomy-Guided Residual Motion Diffusion for Controllable 4D Cardiac MRI Synthesis
Developing robust artificial intelligence models for 4D (3D + time) medical imaging is constrained by limited annotated data, inter-device…
AIGP: An LLM-Based Framework for Long-Term Value Alignment in E-Commerce Pricing
Traditional dynamic pricing models in large-scale e-commerce suffer from limited interpretability, poor utilization of unstructured informa…
MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG
Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning…
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
CLIP and its variants are widely adopted visual backbones in multimodal systems, but their pretraining remains dominated by descriptive ima…
NaviCache: Test-Time Self-Calibration Caching for Video Generation
Video Diffusion Models (VDMs) is constrained by immense computational costs. While offline calibration-based acceleration suffers from cali…
Fortress and Gatekeeper: Theorizing Transitive Trust in Third-Party Cybersecurity Risk Governance
Third-party vendors, such as analytics platforms, cloud services, identity providers, and software suppliers, are increasingly embedded in…
Information-Aware KV Cache Compression for Long Reasoning
Reasoning capability has advanced rapidly in large language models (LLMs), leading to an increasing size of key-value (KV) cache in both pr…
Bridging Vision and Language Concepts through Optimal Transport Semantic Flow
Concept Bottleneck Models (CBMs) promise transparent reasoning by predicting through human-interpretable concepts, yet their effectiveness…
SamaVaani: Auditing and Debiasing Multilingual Clinical ASR for Indian Languages
Automatic Speech Recognition (ASR) is increasingly used to document clinical encounters, yet its reliability in multilingual and demographi…
Confidence-Aware Tool Orchestration for Robust Video Understanding
Video reasoning language models implicitly assume that every input frame is equally reliable. This leads to what we term the Blind Trust Pr…
GEOALIGN: Geometric Rollout Curation for Robust LLM Reinforcement Learning
Online reinforcement learning is widely used to align large language models (LLMs) with reward signals, yet training can be unstable under…
Risk-Aware Selective Multimodal Driver Monitoring with Driver-State World Modeling
Continuous driver monitoring in automated vehicles requires low-latency inference while avoiding unsafe decisions under uncertain driver st…
A Deterministic Control Plane for LLM Coding Agents
LLM coding harnesses grant agents broad file and shell access, yet the configuration layer that steers them -- rules files, agent definitio…
Chai: Agentic Discovery of Cryptographic Misuse Vulnerabilities
AI-assisted vulnerability discovery has proven effective for bug classes like memory safety, where instrumentation confirms memory violatio…
Scaling Multi-Reference Image Generation with Dynamic Reward Optimization
While personalized image generation has achieved remarkable progress, multi-reference image generation (MRIG) remains a challenging task. M…
XMSE-Aware Adaptive Empirical Bayes Estimation
Empirical Bayes (EB) estimators can match the first-order asymptotic risk of maximum likelihood (ML) while behaving very differently at sec…
In-Context Model Predictive Generation: Open-Vocabulary Motion Synthesis from Language Models to Physics
Synthesizing human motion from textual descriptions is essential for immersive digital applications, yet existing methods face a persistent…
Auditing Framing-Sensitive Behavioral Instability in Large Language Models for Mental Health Interactions
Large language models (LLMs) are increasingly being integrated into mental health support tools and other psychologically sensitive convers…
ReaORE: Reasoning-Guided Progressive Open Relation Extraction Empowered by Large Reasoning Models
Open Relation Extraction (OpenRE) requires a model to extract unseen relations between head and tail entities from unstructured text for re…
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Recent work identified emotion vectors in Claude Sonnet 4.5, which are internal representations that encode emotion concepts, causally infl…
Decision-Aligned Evaluation of Uncertainty Quantification
Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected ca…
Event-Aware Instructed Assistant for Referring Video Segmentation
Existing referring video segmentation methods often treat a video as a single event consisting of multiple images, overlooking the fact tha…
Inverse Design of Compact and Wideband Inverted Doherty Power Amplifiers Using Deep Learning
This paper presents a deep learning-assisted methodology for the inverse synthesis of a compact, wideband inverted Doherty power amplifier…
On-board Remote-Sensing Foundation Models for Unsupervised Change Detection of Disaster Events
Remote Sensing Foundation Models (RSFMs) have emerged as a powerful alternative to supervised models for Earth Observation, allowing satell…
ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP
With the rapid evolution of LLM-driven agents, Model Context Protocol (MCP), an open protocol bridging LLMs with external tools, has quickl…
State Representation Matters in Deep Reinforcement Learning: Application to Energy Trading
Energy trading decisions depend not only on current market prices, but also on expected future market conditions, and operational constrain…
The Spec Growth Engine: Spec-Anchored, Code-Coupled, Drift-Enforced Architecture for AI-Assisted Software Development
AI coding agents dramatically accelerate implementation speed but introduce two structural failure modes that existing spec-driven approach…
NuclearQAv2: A Structured Benchmark for Evaluating Domain-Science Competence in Large Language Models
Large language models (LLMs) have demonstrated strong performance across a wide range of tasks, but ensuring their reliability in highly te…
Parametric Open Source Games
Open-source game theory studies agents whose behavior may depend on one another's decision procedures, but most existing models use discret…
Beyond Global Divergences: A Local-Mass Perspective on Bayesian Inference
Global objectives, such as KL divergence and ELBO, are widely used in Bayesian inference for measuring distributional discrepancy. This pap…
Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation
LLMs fine-tuned for security classification are usually evaluated on held-out examples from the same distribution as their training data. W…
Data-Free Reservoir Features for Efficient Long-Horizon Cold-Start Continual Learning
Cold-start exemplar-free class-incremental learning requires learning a growing set of classes without replay, external pretraining, or a l…
Application of LLMs to Threat Assessment of Foreign Peacekeeping Missions
We present a novel approach for applying Large Language Models (LLMs) to threat assessment in the context of foreign peacekeeping missions.…
Heavy-Ball Q-Learning with Residual Weighting Correction
This paper proposes a corrected heavy-ball Q-learning method for reinforcement learning (RL) and establishes its convergence. It also ident…
Efficient foundation decoders for fault-tolerant quantum computing
Foundation decoders, a class of high-capacity neural decoders, are leading candidates for fault-tolerant quantum computing, with accurate a…
Safe Autoregressive Image Generation with Iterative Self-Improving Codebooks
Unlike diffusion-based models that operate in continuous latent spaces, autoregressive unified multimodal models produce images by sequenti…
Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
I describe my solution to the LeHome Challenge 2026, an ICRA 2026 competition on bimanual garment folding. The system placed 1st of 62 team…
Automating Potential-based Reward Shaping with Vision Language Model Guidance
Sparse rewards are inherently challenging for reinforcement learning agents as they lack intermediate feedback to guide exploration and to…
CARVE: Content-Aware Recurrent with Value Efficiency for Chunk-Parallel Linear Attention
Recurrent models must forget in order to remember, yet the state of the art decides what to erase without consulting what is stored -- the…
Bridging Talk and Thought: Understanding Dialogue Dynamics Across Collaborative Problem-Solving Contexts
We present a conceptual framework for analyzing dialogue in collaborative problem-solving contexts, with an emphasis on the emerging dynami…
From Celebrities to Anyone: Characterizing AI Nudification Content, Technology, and Community Dynamics on 4chan
AI nudification uses generative models to create synthetic non-consensual sexually explicit imagery (SNEACI) of real individuals. Prior wor…
Advancing Omnimodal Embodied Agents from Isolated Skills to Everyday Physical Autonomy
Building persistent embodied agents in unstructured environments demands unified orchestration of heterogeneous tools spanning both cyber (…
E-TTS: A New Embodied Test-Time Scaling Framework for Robotic Manipulation
Recently, a few works have made early attempts to study test-time scaling for embodied tasks. However, two major challenges remain unsolved…
AI Healthcare Chatbots as Information Infrastructure: A Large-Scale Study of User-Reported Breakdowns
AI healthcare chatbots are increasingly used to support health information seeking and self-management, yet their performance and impact on…
Beyond the Hard Budget: Sparsity Regularizers for More Interpretable Top-k Sparse Autoencoders
Sparse autoencoders (SAEs) have become a leading tool for interpreting the representations of vision foundation models, decomposing their p…
Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task Planning
Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing compl…
Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching
Entity Matching (EM) is a core operation in the data integration pipeline, where records from different sources are compared to determine w…
Error-Conditioned Neural Solvers
Neural surrogate models offer fast approximate mappings from PDE parameters to solutions, but they typically treat solving as a purely stat…
Autoregressive Boltzmann Generators
Efficient sampling of molecular systems at thermodynamic equilibrium is a hallmark challenge in statistical physics. This challenge has dri…
A Concept of Possibility for Real-World Events
This paper offers a new concept of {\it possibility} as an alternative to the now-a-days standard concept originally introduced by L.A. Zad…
Human-AI Complementarity: A Goal for Amplified Oversight
Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging ta…
SciFig: Towards Automating Editable Figure Generation for Scientific Papers
High-quality methodology figures are central to scientific communication, yet they remain difficult and time-consuming to create. Such figu…
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models
Reward models are critical for reinforcement learning from human feedback, as they determine the alignment quality and reliability of gener…
CLEF HIPE-2026: Evaluating Accurate and Efficient Person-Place Relation Extraction from Multilingual Historical Texts
HIPE-2026 is a CLEF evaluation lab dedicated to person-place relation extraction from noisy, multilingual historical texts. Building on the…
Experience Compression Spectrum: Unifying Memory, Skills, and Rules in LLM Agents
As LLM agents scale to long-horizon, multi-session deployments, efficiently managing accumulated experience becomes a critical bottleneck.…
To Use AI as Dice of Possibilities with Timing Computation
The dominant noun-based modeling paradigm has fundamentally constrained AI development, precluding any adequate representation of the futur…
Evaluating Deep Research Agents on Expert Consulting Work: A Benchmark with Verifiers, Rubrics, and Cognitive Traps
Frontier deep research agents (DRAs) are being deployed in enterprise workflows faster than they are being evaluated. Existing benchmarks m…
Library Drift: Diagnosing and Fixing a Silent Failure Mode in Self-Evolving LLM Skill Libraries
Self-evolving skill libraries face a silent failure mode we term \emph{library drift}: unbounded skill accumulation without outcome-driven…
Augmentation techniques for video surveillance in the visible and thermal spectral range
In intelligent video surveillance, cameras record image sequences during day and night. Commonly, this demands different sensors. To achiev…
Automated reproducibility assessments in the social and behavioral sciences using large language models
Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to…
R2D-RL: A RoboCup 2D Soccer Environment for Multi-Agent Reinforcement Learning
Robot soccer is a challenging testbed for multi-agent reinforcement learning because it combines partial observability, cooperative and adv…
A-Evolve-Training: Autonomous Post-Training of a 30B Model
Post-training a frontier model is normally weeks of human work: proposing data and recipe changes, launching runs, reading evals, deciding…
Cliff Tokens: Identifying Single-Token Failure Triggers in LLM Mathematical Reasoning
Large language models (LLMs) reach high accuracy in mathematical reasoning, but individual traces on the same problem diverge; some arrive…
Autodata: An agentic data scientist to create high quality synthetic data
We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation dat…
Wearable Device-Based Real-Time Monitoring of Physiological Signals: Evaluating Cognitive Load Across Different Tasks
This study employs cutting-edge wearable monitoring technology to conduct high-precision, high-temporal-resolution (1-second interval) cogn…
Byzantine-Robust Aggregation for Securing Decentralized Federated Learning
Federated Learning (FL) emerges as a distributed machine learning approach that addresses privacy concerns by training AI models locally on…
Tuning Language Models by Mixture-of-Depths Ensemble
Transformer-based Large Language Models (LLMs) traditionally rely on final-layer loss for finetuning and final-layer representations for pr…
Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
Despite the impressive capabilities of Large Vision-Language Models (LVLMs), they remain susceptible to hallucinations, where generated con…
HauntAttack: When Attack Follows Reasoning as a Shadow
Emerging Large Reasoning Models (LRMs) consistently excel in mathematical and reasoning tasks, showcasing remarkable capabilities. However,…
DMSC: Dynamic Multi-Scale Coordination Framework for Time Series Forecasting
Time Series Forecasting (TSF) faces persistent challenges in modeling intricate temporal dependencies across different scales. Despite rece…
Learning to Select Maximum Clique Algorithms: From Traditional Machine Learning to a Dual-Channel Hybrid Neural Architecture
The Maximum Clique Problem (MCP) is an NP-hard problem with wide-ranging applications in fields such as bioinformatics, network science, an…
Through the Looking Glass: A Dual Perspective on Weakly-Supervised Few-Shot Segmentation
Meta-learning aims to uniformly sample homogeneous support-query pairs, characterized by the same categories and similar attributes, and ex…
Reconstruction Alignment Improves Unified Multimodal Models
Unified multimodal models (UMMs) unify visual understanding and generation within a single architecture. However, conventional training rel…
Limited Reference, Reliable Generation: A Two-Component Framework for Tabular Data Generation in Low-Data Regimes
Synthetic tabular data generation is increasingly essential in machine learning, supporting downstream applications when real-world, high-q…
Rotary Position Encodings for Graphs
We study the extent to which rotary position encodings (RoPE), a recent transformer position encoding algorithm broadly adopted in large la…
The Journal of Prompt-Engineered (Moral) Philosophy Or: Why AI-Assisted Ethics Research Requires Process Transparency
Existing AI disclosure mandates in scholarship require that AI assistance be reported but leave transparency philosophically unspecified: t…
Patent Representation Learning via Self-supervision
We study self-supervised patent representation learning with contrastive objectives. A standard baseline constructs positives by encoding t…
Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
Existing methods for expressive music performance rendering, a conditional generation task that aims to generate a human-like performance f…
The Best of the Two Worlds: Harmonizing Semantic and Hash IDs for Sequential Recommendation
Conventional Sequential Recommender Systems (SRS) typically assign unique hash IDs (HID) to construct item embeddings, which mainly capture…
Improved Bounds for Private and Robust Alignment
In this paper, we study the private and robust alignment of language models from a theoretical perspective by establishing upper bounds on…
Digital Twin-Driven Communication-Efficient Federated Anomaly Detection for Industrial IoT
Anomaly detection is increasingly becoming crucial for maintaining the safety, reliability, and efficiency of industrial systems. Recently,…
Metaphors are a Source of Cross-Domain Misalignment of Large Reasoning Models
Earlier research has shown that metaphors influence human decision-making, raising the question of whether metaphors also influence large l…
Dual-Prototype Disentanglement: A Context-Aware Enhancement Framework for Time Series Forecasting
Time series forecasting has witnessed significant progress with deep learning. While prevailing approaches enhance forecasting performance…
VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image
3D editing has emerged as a critical research area to provide users with flexible control over 3D assets. While current editing approaches…
Revisiting the Platonic Representation Hypothesis: An Aristotelian View
The Platonic Representation Hypothesis suggests that representations from neural networks are converging to a common statistical model of r…
ReportLogic: Evaluating Logical Quality in Deep Research Reports
Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports…
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
The effectiveness of multimodal instruction tuning depends not only on dataset scale, but critically on whether training samples genuinely…
Delegation and Verification Under AI
As AI systems enter institutional workflows, workers must decide whether to delegate task execution to AI and how much effort to invest in…
Latent-Mark: An Audio Watermark Robust to Neural Codec Compression
While existing audio watermarking techniques have achieved strong robustness against traditional digital signal processing (DSP) attacks, t…
Residual RL-MPC for Robust Microrobotic Cell Pushing Under Time-Varying Flow
Contact-rich micromanipulation in microfluidic flow is challenging because small disturbances can break pushing contact and induce large la…
A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation
Delineating the clinical target volume (CTV) in radiotherapy involves complex margins constrained by tumor location and anatomical barriers…
MedPruner: Training-Free Hierarchical Token Pruning for Efficient 3D Medical Image Understanding in Vision-Language Models
While specialized Medical Vision-Language Models (VLMs) have achieved remarkable success in interpreting 2D and 3D medical modalities, thei…
Power Couple? AI Growth and Renewable Energy Investment
AI and renewable energy are increasingly framed as a "power couple," on the premise that surging AI demand will accelerate clean-energy inv…
Scalable AI-assisted Workflow Management for Detector Design Optimization Using Distributed Computing
The Production and Distributed Analysis (PanDA) system, originally developed for the ATLAS experiment at the CERN Large Hadron Collider (LH…
The Augmentation Trap: AI Productivity and the Cost of Cognitive Offloading
Experimental evidence suggests that AI tools raise worker productivity, but also that sustained use can erode the expertise on which those…
Statistical Properties of the King Wen Sequence: An Anti-Habituation Structure That Does Not Improve Neural Network Training
The King Wen sequence of the I-Ching (c. 1000 BC) orders 64 hexagrams -- states of a six-dimensional binary space -- in a pattern that has…
Finetuning-Free Diffusion Model with Adaptive Constraint Guidance for Inorganic Crystal Structure Generation
Generative diffusion models have emerged as powerful tools for the discovery of inorganic crystal structures, yet steering their sampling p…
TransXion: A High-Fidelity Graph Benchmark for Realistic Anti-Money Laundering
Money laundering poses severe risks to global financial systems, driving the widespread adoption of machine learning for transaction monito…
Peer-Preservation in Frontier Models
Recent work has found that frontier AI models can exhibit misaligned behaviors in pursuit of assigned goals. We demonstrate that models can…
Mapping License Plate Recoverability Under Extreme Viewing Angles for Opportunistic Urban Sensing
Urban environments contain many imaging sensors built for specific purposes, including ATM, body-worn, CCTV, and dashboard cameras. Under t…
Hierarchical Fault Detection and Diagnosis for Transformer Architectures
Transformers now underpin critical AI systems across industry and research. Yet their faults can silently alter model behavior without runt…
S2P-Net: A Spectral-Spatial Polar Network for Rotation-Invariant Object Recognition in Low-Data Regimes
We present S2P-Net (Spectral-Spatial Polar Network), a compact deep learning architecture that achieves mathematically guaranteed rotation…
Weak-to-Strong Elicitation via Mismatched Wrong Drafts
We consider whether off-policy experience from a smaller, weaker model can elicit capability in a stronger learner that on-policy RL fine-t…
Semantic Generative Tuning for Unified Multimodal Models
Unified multimodal models (UMMs) strive to consolidate visual understanding and visual generation within a single architecture. However, pr…
Sutra: Tensor-Op RNNs as a Compilation Target for Vector Symbolic Architectures
Sutra is a typed, purely functional programming language whose compiled forward pass is a PyTorch neural network. The compiler beta-reduces…
Beyond Independent Manipulation: Individual Fairness-aware Strategic Classification with Peer Imitation
Strategic classification (SC) investigates scenarios where agents manipulate their features to obtain favorable decisions from predictive m…
Symbolic Reasoning Frameworks Trigger Memory-Mediated Ecosystem Dynamics in Multi-Agent LLM Systems
Large language models exhibit a risk-averse "turtle" bias as strategic agents. We show that injecting a symbolic reasoning framework as a p…
An LLM-Native Psychometric Instrument Does Not Predict LLM Behavior: Evidence Across 25 Models
Large language models (LLMs) give stable answers to personality questionnaires, yet these self-reports fail to predict how the models actua…
When Role-playing, Do Models Believe What They Say?
Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite. Recent work argues that pe…
SymQNet: Amortized Acquisition for Low-Latency Adaptive Hamiltonian Learning
Adaptive Hamiltonian learning is central to calibrating and characterizing quantum devices. In an adaptive controller, choosing the next ex…
Position: Align AI to Our Aspirations, Not Our Flaws
We argue that aligning AI to aggregated human preferences is the wrong target. With current technology, one can train AIs to share the valu…
Trust in Generative AI for Health Information Consumption and the Effect of Learned Dependency: An Experimental Investigation
Background: Generative artificial intelligence (GenAI) is increasingly used for health information, yet its influence on users' trust calib…
Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior
Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on t…
A3C3: AI Algorithm and Accelerator Co-design, Co-search, and Co-generation
We present a holistic methodology for artificial intelligence algorithm and accelerator co-design, co-search, and co-generation (A3C3), whi…
MMGist: A Comprehensive Multimodal Benchmark for 2027
We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on vi…
Small edits, large models: How Wikipedia advocacy shapes LLM values
Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia ap…
Noise-Aware Boundary-Enhanced Generative Learning for Ultrasound Speckle Reduction
Ultrasound is a non-invasive, real-time, and cost-effective imaging technique widely used in clinical diagnosis. However, its diagnostic ef…
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
We present Wan-Streamer, a native-streaming, end-to-end interactive foundation model designed from the ground up for real-time, low-latency…
MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources
Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a chal…
「OpenAIはAzureだけ」の時代が終了 「GPT-5.5」「Codex」をAWSで利用するメリットは何か
OpenAIとAmazon Web Services(AWS)が戦略的パートナーシップを拡大した。OpenAIのモデル、コーディングエージェント「Codex」、マネージドエージェントを、企業がAWS環境で利用できるようにする。
Amazon Bedrockのトークン処理量、26年1Qだけで過去累計超え AWSが目指す、AIのための「信頼できるインフラ」
Amazon Bedrockのトークン処理量が2026年第1四半期だけで過去の累計を上回ったという。「AWS Summit Japan 2026」の基調講演では、AI推論が最大の負荷となる時代の「信頼できるインフラ」の必要性が訴えられた。
企業のAI支出そろそろ“様子見”は終わり? Gartner予測、本格投資の行方は
Gartnerは2026年の世界AI支出が前年比47%増の2兆5956億ドルに達するとの予測を発表。2026年は企業によるAI支出が拡大局面へ移行する転換点になるとしている。
The White House is asking OpenAI to slow roll the release of its new model over safety concerns
penAI reportedly plans to share its newest model, GPT 5.6, with a select group of partners instead of to the broader public. The reason: th…
日本の“眠れるデータ”を競争力へ、日本発のプラットフォーム「xIPF」が始動
欧州を中心に進むデータ共有圏の動向やその日本へのインパクトについて解説してきた本連載だが、第9回は日本独自のデータ連携エコシステムを創出することを目指す「xIPFコンソーシアム」を取り上げる。
ITインフラの要件に“大変化” なぜ「電力」と「冷却」にAI活用が縛られるのか?
AI時代の到来により、ITインフラの要件に大きな変化が訪れています。データセンターを「どこに置くか」「どうやって冷やすか」が、AI活用の制約になるのはなぜか。新しいITインフラの要件に、IT部門がどのように対応すべきでしょうか。
“AIの回答が薄い問題”をどう解決? 日本ハム、「AIが食べやすいデータ」作戦の全貌
データ×AIによる意思決定を推進するには、非構造化データを「AIが食べやすいデータ」に変える必要がある。日本ハムが選んだアプローチとは。
“流れていく会話”をAIが構造化 チャット履歴を会社の資産に
業務上の会話には、後から参照すべき貴重な情報が多く含まれる。しかし、従来型のチャットでは、こうした情報が整理されないまま流れ、必要な場面で見つけにくい。そうした課題を解決するために、AIチャットエージェントが役立つかもしれない。
リコーが多能工ヒューマノイドを披露、工場ではPoCから導入に向けた実証段階へ
リコーは、「AWS Summit Japan 2026」において、フィジカルAI搭載の多能工ヒューマノイドのデモンストレーションを披露した。既に工場内でPoCを始めており、今夏までをめどに多能工ヒューマノイドが一部の工程を担うより実用的な実証を始めたい考えだ。
Patronus AI lands $50M to build ‘digital worlds’ that stress-test AI agents
Agent-testing startup Patronus AI, founded by former Meta AI researchers, is experiencing nearly insatiable demand, its investor says.
Anthropic’s Claude is winning over paid consumers, a market owned by ChatGPT
Despite ChatGPT's commanding market lead, consumers who pay for AI have been increasingly choosing Anthropic's Claude, data shows.
General Intuition’s $2.3B bet that video games can train AI agents for the real world
General Intuition has raised $320 million to scale AI trained on millions of hours of gameplay, betting action data can help AI develop som…
Databricks’ former AI chief thinks he can cut AI’s power bill by 1,000x
Un-0 is an image-generation system tool that shows for the first time how the company's technology can replicate conventional AI systems.
2026-06-25(32件)
Netris raises $15M Series A from a16z to help AI neoclouds go live faster
Netris provides software that runs on network switches, and offers a platform that helps neocloud operators reduce the time it takes to go…
2 days left to save up to $190: Join 1,000+ founders and investors at TechCrunch Founder Summit
Two days left to lock in your spot at TechCrunch Founder Summit 2026 and save up to $190 before Early Bird rates expire on June 26 at 11:59…
Adobe acquires image and video enhancement tool maker Topaz Labs
Adobe said that it will integrate Topaz Labs' tools across its apps.
Flashの再来? Figmaの新機能「Figma Motion」に懐かしいとの声 アニメーション生成するAI機能も
Figmaが発表した新機能「Figma Motion」が、かつてのAdobe Flashを思わせるとSNSで話題だ。タイムラインでキーフレームを打つ操作感が懐かしさを呼んだが、実際に触ると別物との声もみられる。
Amazon ups India bet with fresh $13B AI infrastructure investment
Amazon’s latest India investment comes as global tech companies race to expand AI infrastructure in the country.
男性に美人局容疑で3人逮捕 ChatGPTの示談相場示し脅迫か 警視庁
美人局の手口で、少女とホテルに入った男性から現金を脅し取ったとして、警視庁少年事件課は恐喝の疑いで、東京都練馬区上石神井の職業不詳、斎藤蓮容疑者(22)ら男3人を逮捕、男女2人を書類送検した。斎藤容疑者は黙秘し、4人は容疑を認めている。
中国が人型ロボット開発で急成長しているワケ 日本が学ぶべきポイントは? 専門家が解説
なぜ人型ロボットの開発で中国が急成長しているのか。日本が学ぶべきポイントを野村総合研究所の李智慧氏が解説した。
「教員を生成AIに置き換える考えはない」東京外大が声明 SNSで拡散した懸念にコメント
東京外大が、「2人体制で運営する小規模語科を1人体制に縮小し、AIに代替させようとしている」との懸念がSNSで広まったことについて、「教員を生成AIに置き換える考えはない」との声明を出した。
白血病など16疾患の診断をAIが支援、日立がAUC0.9以上の新技術
日立製作所と九州大学病院は、血液悪性腫瘍の診断に用いるフローサイトメトリー検査において、医師の鑑別診断を支援する機械学習型のAI技術を開発した。複数疾患の同時分類において、識別性能を示す指標AUCで0.9以上の性能を確認した。
シャープブランドのAIサーバ展開も検討 シャープと鴻海、5分野で戦略的協業へ
シャープはこれまで、鴻海が製造するAIサーバの国内販売方針を示していたが、新たに、製品をシャープブランドとして展開を検討する方針が示された。
「今日言うつもりはなかったが……」 孫正義氏が明かした「ロボット自動量産工場」の実態
「今日ここで言うつもりはなかったんですが」──。ソフトバンクグループが6月24日に開催した株主総会の質疑応答で、会長兼社長の孫正義氏が、投資先の現場で起きている現場実態を明かす一幕があった。
Google、「Gemini 3.5 Flash」に「Computer Use」を標準搭載──AIが画面を見てブラウザやアプリを操作
Googleは、AIモデル「Gemini 3.5 Flash」に、AIがコンピュータの画面を認識してマウス操作やキーボード入力を自動で実行する「Computer Use」機能を標準ツールとして搭載したと発表した。これまで専用モデルでのみ提供していた機能を主力モデルに統合したもの…
OpenAI、「GPT-5.5 Instant」をアップデート 会話の文脈維持や箇条書き減など「読みやすさ」を改善
OpenAIは、ChatGPTで最も広く利用されているモデル「GPT-5.5 Instant」のアップデートを発表した。新機能の追加ではなく、日常的な会話の品質向上が中心。質問の意図を的確に捉えて文脈を維持する能力が向上するほか、テンプレート的な回答が減り、位置情報を活用した地…
Europe is pushing back on Washington’s chip war
As ASML CEO Christophe Fouquet told TechCrunch in May, what China can currently buy are older-generation deep ultraviolet tools — gear firs…
Former Infosys chief has a new startup that wants to challenge the IT services world
Backed by Mayfield and Aramco Ventures, Vishal Sikka’s new venture brings together veterans from SAP, Infosys, and VianAI.
富士通と日本IBMの協業、ついに始動 COBOL刷新における「役割分担」は?
レガシーシステムをどうモダナイズするかは、多くの企業における課題だ。富士通と日本IBMがこの領域での協業を発表した。ついに始動する、両社の協業における役割分担とは。
【役に立つの?】「Google公式」の初心者向けAI講座、受けてみたら想像以上にすごかった
1日で1万人以上が登録したGoogleの初心者向けAI講座を実際に体験。想像以上の学びとは?
慶応大がNotionを選んだ「3つの理由」 “何ができるか”以外の決め手は?
慶應義塾大学が全教職員にNotionを導入し、「AIキャンパス構想」を本格始動した。数あるツールからNotionを選んだ理由は何か。また、同大が目指す、生成AI時代におけるナレッジ管理の形とは。
Cerebras stock plunges after earnings as CEO says margin outlook was misunderstood
In its first earnings report since going public, the AI chipmaker forecast a narrower gross margin in its core business, scaring investors.
OpenAI、初の独自AIチップ「Jalapeno」発表──Broadcomと共同開発の推論用アクセラレータ
OpenAIは、Broadcomと共同開発したLLMの推論に特化した独自AIチップ「Jalapeno」を発表した。同社初の「インテリジェンスプロセッサ」と位置付けており、設計から製造用のテープアウトまでを9カ月で完了したという。数世代にわたる計算基盤の第1弾として、年内に展開を…
味の素、“万能DX人材”増員へ 育成のきっかけは新規プロジェクトの苦い経験
味の素が、AIを活用して「フルスタック人財」の育成を目指している。DX推進のキーパーソンにする狙いだ。既に300時間の工数削減例も生まれている。一体どう育成したのか。
AI was supposed to kill engineering jobs, but new data suggests they’re the most resilient
While AI dominates the layoff narrative, engineers are actually making up a larger share of new hires, according to SignalFire data.
AI researchers continue to leave Google for its rivals
Top AI researchers Jonas Adler and Alexander Pritzel are leaving Google for Anthropic, following departures from top scientists Noam Shazee…
The memory chip crunch is paying off for this US company
Revenue quadrupled to $41.45 billion compared with the same period a year ago. The company's profit, meanwhile, rose from $1.88 billion to…
Companies are scrambling to stop employees from maxing out AI budgets with small tasks
The tokenmaxxing era was brief. We now appear to be entering the era of token rationing.
「AI教育、どこから手を付ける?」 全社導入のカギは“生成AIリテラシー向上研修”(前編)
生成AIは業務の現場に急速に浸透し、「使って当たり前」の時代が到来しています。その活用範囲は広がる一方、情報漏洩や誤情報のリスクが企業の大きな課題になっています。今求められるのは、誰もが“安全かつ賢く”生成AIを使いこなすリテラシーです。本稿は、社内の誰もが生成AIを安全に、自…
「AI教育、どこから手を付ける?」 全社導入のカギは“生成AIリテラシー向上研修”(後編)
生成AIは業務の現場に急速に浸透し、「使って当たり前」の時代が到来しています。その活用範囲は広がる一方、情報漏洩や誤情報のリスクが企業の大きな課題になっています。今求められるのは、誰もが“安全かつ賢く”生成AIを使いこなすリテラシーです。本稿は、社内の誰もが生成AIを安全に、自…
「最強モデル」はもう無意味 ナデラCEOが語る、企業の生き残り新戦略「学習ループ」
AIの進化で、自社システムの模倣やコモディティ化への不安が広がっている。MicrosoftのナデラCEOが示す「学習ループ」戦略とは何か。日本のソフトウェア企業の生き残りにも通じる筆者の視点を交えて解説する。
Facebook rolls out an AI companion app for creators
The new app, which is currently being tested with select creators, will have Facebook's recently launched AI creator assistant built into i…
Agility Robotics plans to go public via SPAC in a $2.5B deal
Agility Robotics, the humanoid robotics startup that spun out of Oregon State University in 2015, expects to generate $620 million in proce…
Figma adds code layers, support for animations, more AI features in new update
Figma's update adds a new code layer, support for motion and shaders, and the ability to create custom plug-ins for various tasks using AI.
2026-06-24(322件)
OpenAI unveils its first custom chip, built by Broadcom
Named Jalapeño, the new processor was designed specifically for the unique needs of OpenAI's inference systems.
3 days left to save up to $190 on your TechCrunch Founder Summit 2026 pass
You have just 3 days left to save up to $190 on your pass to TechCrunch Founder Summit 2026 before Early Bird rates end on June 26 at 11:59…
「Transformerの最大475倍」 富士通、GPUを効率的に使うLLMアーキテクチャ「PHOTON」開発
富士通が、大規模言語モデル(LLM)を少ないGPUで動かせる新アーキテクチャ「PHOTON」(フォトン)を開発した。GPU当たりの処理性能(スループット)が、現在のLLMで主流のアーキテクチャ「Transformer」の最大475倍に達するという。LLMの運用に必要なGPUを抑…
陸自駐屯地で四足歩行型の警備用ロボットが見回り GMOインターネットグループが開発
GMOインターネットグループ4社は、国産ロボット開発を担う未来ロボットと組み、国産の四足歩行型警備用ロボットを開発し、陸上自衛隊の駐屯地での導入検証を始めると発表した。警備の省人化を図り、24時間警備体制の実現を目指すという。
民生VRグローブにロボット業界が注目 日本発ベンチャーがB2B加速
Diver-Xは2026年6月23日、グローブ型仮想現実(VR)コントローラー新製品「ContactGlove3」を発表した。電磁場トラッキング方式の採用により推奨環境下で中央値0.5mm、最大値1.5mmの誤差という高精度を実現した。民生用と業務用を用意していて、ロボティクス…
【解説】キオクシアなぜ急成長? 半導体メモリって何? AIブームを見通すための基礎知識
注目を集める半導体メモリ大手のキオクシア。同社はなぜAI需要を取り込めたのか。いま押さえたい基礎知識を解説する。
ClaudeをSlackチャンネルに召喚、“チームの一員”として直接指示 新機能「Claude Tag」登場
Anthropicは、「Slack」上で「Claude」を利用できる新機能「Claude Tag」のベータ版提供を23日から開始した。指定したチャンネルにClaudeを招待し、メンションをつけてタスクを依頼すると、Claudeが関連情報をもとにタスク計画を自動構築する。
国内ユーザー数「前年比1582%増」――AI開発支援「Devin」は競合と何が違う? 日本法人代表が語る事業戦略
ソフトウェア開発向けAIエージェント「Devin」をどのように日本市場で展開するのか。Devinを手掛ける米Cognition AI日本法人の正井拓己代表が語った。
OpenAI and Broadcom unveil LLM-optimized inference chip
OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI sy…
RIFT-Bench: Dynamic Red-teaming For Agentic AI Systems
Agentic AI systems powered by large language models (LLMs) are rapidly evolving into autonomous decision-making systems, exposing attack ve…
Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs
Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and e…
Critique of Agent Model
What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding agents'', ``AI co-scien…
Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control
Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints. Exis…
Reinforcement Learning Towards Broadly and Persistently Beneficial Models
As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domai…
Can Language Model Agents be Helpful Circuit Explainers in Mechanistic Interpretability?
Mechanistic interpretability has made substantial progress in automatically localizing circuits, but explaining what localized components d…
Breaking the Filter Bubble: A Semantic Pareto-DQN Framework for Multi-Objective Recommendation
Recommender systems often induce filter bubbles and semantic homogenization by monolithically optimizing for immediate user engagement. Sta…
Ensemble Feature Selection and Harris Hawks Optimization for Explainable Mental Health Risk Prediction in Female Sex Workers
One of the significant mental health issues affecting female sex workers (FSWs) is mental disorders, especially depression. Exposure to vio…
Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning
Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectiv…
Exploring Academic Influence of Algorithms by Co-occurrence Network Based on Full-text of Academic Papers
Algorithms have become central to scientific research in the era of artificial intelligence (AI). Although algorithm mentions in papers are…
ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection
Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images,…
VeryTrace: Verifying Reasoning Traces through Compilable Formalism and Structured Verification
Multi-step reasoning with Chain-of-Thought (CoT) prompting remains fragile: logical errors or hallucinations in early steps silently propag…
OmniPath: A Multi-Modal Agentic Framework for Auditing Wheelchair Accessibility
For a wheelchair user, a standard blue line on a map is often a broken promise. While platforms like OpenStreetMap (OSM) successfully captu…
T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph
Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constrain…
The Geometry Behind Diffusion and Flow Matching: Gradient Flows and Geodesics in Wasserstein Space
The space $\mathcal{P}_2(\mathbb{R}^d$) of probability measures with finite second moment carries a natural geometry: the quadratic Wassers…
An Introduction to Causal Reinforcement Learning
Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with qu…
Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR
Adapting a streaming speech recognition model to a new language requires choosing between two plausible warm starts: a multilingual (ML) en…
Navigating User Behavior toward Personalized Multimodal Generation
Modern AIGC pipelines deliver high-fidelity images and videos but presuppose a well-formed creation instruction, while end users rarely art…
Exploring the relationship between human-centric AI and firm idiosyncratic risks
Despite the extensive discussions of human-centric AI (HCAI) in Industry 5.0, its effects on firms' idiosyncratic risks (IR) remains undere…
FlowR2A: Learning Reward-to-Action Distribution for Multimodal Driving Planning
Multimodal driving planning faces a long-standing tension between two paradigms: scoring-based methods benefit from dense reward supervisio…
SP-Mind: An Autonomous Reasoning Agent for Spatial Proteomics Analysis
Spatial proteomics enables single-cell-resolution characterization of protein expression within tissue architecture, playing a critical rol…
Towards Federated Long-Tailed Graph Learning: An Energy-Guided Dual Decoupling Approach
Federated Graph Learning facilitates collaborative graph modeling across distributed clients while preserving data privacy. However, real-w…
Probing the Misaligned Thinking Process of Language Models
Large language models exhibit a growing range of misaligned behaviors such as strategic deception, sandbagging, and self-preservation. As t…
Tractable Reasoning and Conjunctive Query Answering for Defeasible DL-Lite under Rational Closure
In Description Logics (DLs), reasoning under Rational Closure (RC) is a well-known and widely accepted non-monotonic formalism to handle de…
LemonHarness Technical Report
As large language model (LLM) agents are applied to longer tasks, they increasingly modify workspace state across multiple rounds of iterat…
Prob-BBDM: a Probabilistic Brownian Bridge Diffusion Model for MRI sequence image-to-image translation
AI-driven image-to-image synthesis is rapidly advancing, with growing applications in medical imaging. Multi-modal image analysis plays a c…
MVG-KAN: Multi-View Geo-Wind Guided KAN for PM$_{2.5}$ Forecasting
Accurate short-term PM$_{2.5}$ forecasting is important for public health protection, air-quality early warning, and urban environmental ma…
Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation
Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL…
When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs
Large language models (LLMs) are increasingly integrated into decision-support roles in business and policy contexts. While prior benchmark…
PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models
We introduce a large-scale, open-source dataset of pre-generated adversarial attacks for vision-language models (VLMs). The dataset is desi…
Age of LLM: A Strategic 1v1 Benchmark for Reasoning, Diplomacy and Reliability of Large Language Models under Fog of War
We introduce Age of LLM, a turn-based 1v1 benchmark in which two LLMs face off on a 13x7 grid to destroy the enemy base. Three stressors ar…
ATRIA: Adaptive Traceable ECG Reporting with Iterative Agents
Existing ECG report generation is tightly coupled -- interpretation and reporting fused end-to-end, so errors propagate without stage-level…
Cycle-Consistent Neural Explanation of Formal Verification Certificates
Formal verification produces machine-checkable certificates that attest to the satisfaction or violation of temporal properties, yet these…
Agentic AI for Bilevel Long-Term Optimization of Policy-Driven Physical Layer Systems
Network operators' changing policies, service requirements, and stringent real-time constraints render existing methods designed with fixed…
Can Aggregate Invariants Accelerate Continuous Subgraph Matching? Limits, Laws, and a Dynamic Spectral Index
Spectral filtering recently delivered substantial pruning for \emph{static} subgraph matching: Laplacian interlacing rejects candidates who…
ReM-MoA: Reasoning Memory Sustains Mixture-of-Agents Scaling
Mixture-of-Agents (MoA) architectures improve inference-time scaling by organizing multiple LLM agents into layered reasoning pipelines. Ho…
Bayesian control for coding agents
Modern coding agents pair LLM generators with various tools, including cheap diagnostics and expensive verifiers. The tool-use decisions ar…
CompressKV: Semantic-Retrieval-Guided KV-Cache Compression for Resource-Efficient Long-Context LLM Inference
Long-context large language model (LLM) inference is increasingly constrained by the memory footprint and decoding cost of key-value (KV) c…
The Latent Bridge: A Continuous Slow-Fast Channel for Real-Time Game Agents
A real-time agent for general computer use - with games as the most demanding case - must act within tens of milliseconds while still plann…
On the Smallness of the Large Language Models Scaling Exponents
We discuss reasons why the scaling exponents of current Large Language Models (LLMs) applications are indicating an unsustainable regime in…
A specialized reasoning large language model for accelerating rare disease diagnosis: a randomized AI physician assistance trial
Rare diseases affect millions of individuals worldwide, yet timely diagnosis remains a major public health challenge due to scarcity of spe…
Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation
Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinf…
Governed Shared Memory for Multi-Agent LLM Systems
Multi-agent LLM environments require robust mechanisms for shared knowledge management. This paper formalizes the fleet-memory problem and…
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents
Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluat…
Quant Convergence: Bridging Classical Value Investing and Modern Factor Models for Systematic Equity Selection
Modern finance relies heavily on complex machine learning models to find patterns in the stock market. However, as these AI models get more…
LLMs Prompted for Legal Context Object More: Overrefusal from Small On-Premises LLMs in Criminal Legal Context
While the validity of LLMs' use in the legal context remains subject to ethical and legal debate, legal professionals are already experimen…
AdversaBench: Automated LLM Red-Teaming with Multi-Judge Confirmation and Cross-Model Transferability
Scaling adversarial evaluation of large language models requires both a method for generating hard inputs and a reliable way to confirm tha…
ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning
Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixe…
Uncertainty-Aware Longitudinal Forecasting of Alzheimer's Disease Progression Using Deep Learning
Longitudinal modelling of Alzheimer's disease progression is clinically useful only if it can describe not just the most likely next diagno…
ScaleToT: Generalizing Structured LLM Reasoning for Billion-Scale Low-Activity User Modeling
Accurate user modeling often depends on rich interaction histories, which are unavailable for billions of low-activity users. Large Languag…
AI Tokenomics: The Economics of Tokens, Computation, and Pricing in Foundation Models
Tokens have become the practical accounting unit for modern foundation model services, linking information processing, computation, memory…
Abstractions of Queries in Ontology-Based Data Access
In ontology-based data access (OBDA), multiple data sources are integrated via mappings to an ontology. We consider an OBDA setting based o…
When CQs Go Wrong: Challenges in CQ Verification with OE-Assist
Competency Questions (CQs) are the central component of CQ-verification, an established process in which an ontology is evaluated against a…
Themis: An explainable AI-enabled framework for Reinforcement Learning with Human Feedback
Training safe Reinforcement Learning (RL) systems is inherently challenging, with no guarantee of avoiding unwanted behaviors. The most eff…
SAFARI: Scaling Long Horizon Agentic Fault Attribution via Active Investigation
As autonomous agents tackle increasingly complex multi-step, multi-agent tasks, their execution trajectories have scaled beyond the constra…
CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning
Cinematographic captioning aims to describe how a video is filmed using professional film-language concepts such as camera movement, shot s…
LaGO: Latent Action Guidance for Online Reinforcement Learning
Large language models (LLMs) have shown strong potential for planning and sequential decision-making, but prior work often relies on using…
Cost-Optimal Decision Diagrams for Stochastic Boolean Function Evaluation
In many decision-making scenarios, acquiring information incurs different costs. We consider the problem of constructing a deterministic ev…
Decentralised AI Training and Inference with BlockTrain
Frontier AI training is increasingly shaped by access to dense, centrally controlled accelerator clusters. This creates a structural advant…
Scaling Laws for Task-Specific LLM Distillation
Large Language Models (LLMs) achieve strong performance across a growing range of domains, yet their scale poses deployment challenges in a…
Can Scale Save Us From Plasticity Loss in Large Language Models?
The loss of plasticity - the ability of a network to learn new information after having already learned older information - is a fundamenta…
BluTrain: A C++/CUDA Framework for AI Systems
Progress in deep learning is, at scale, more a matter of systems engineering than of modelling: the behaviour of a model in training (its t…
Assessing Distribution Shift in Human Activity Recognition for Domain Generalization
While the field of Human Activity Recognition (HAR) continues to draw interest from researchers and advance in important ways, some key cha…
Solving Inverse Problems of Chaotic Systems with Bidirectional Conditional Flow Matching
Modeling chaotic systems is crucial yet challenging. Inverse problems in chaotic dynamics, namely inferring initial conditions from final s…
Difference-Making without Making a Difference
Over a series of seven papers, Andreas & G\"unther have introduced seven definitions of actual causation and have classified them as belong…
Accuracy and Satisfaction in Multi-Turn LLM Dialogues for NFR Assessment
LLM-based dialogue assistants have become mainstream tools for software developers, yet current evaluation benchmarks focus exclusively on…
Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System
Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challen…
Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models
Prompt-based learning has emerged as a dominant paradigm in natural language processing. This study explores the impact of diverse pre-trai…
World Models in Pieces: Structural Certification for General Agents
In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. C…
OpenThoughts-Agent: Data Recipes for Agentic Models
Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broad…
Reentrant value fields as delayed coupled reaction-diffusion systems on finite graphs
We describe a dynamical system in which a symbolic field is coupled to a geometric field via a bipartite Hilbert-Schmidt kernel. The system…
Beyond the Autoregressive Horizon: A Comprehensive Survey of Diffusion Models, World Modelling, and State Space Models for Code
Autoregressive (AR) language models have driven significant progress in automated software engineering, enabling powerful code generation a…
Quantifying Prior Dominance in RAG Systems
Retrieval-Augmented Generation (RAG) grounds Large Language Models in external knowledge, yet current evaluations rely on discrete heuristi…
SemChunk-C: Semantic Segmentation for C Code
Semantic segmentation of code written in a C-family language remains a challenging problem, due to the language's complex syntax, macro exp…
FP8 is All You Need (Part 2): Efficient Ozaki-Bailey Style FFT Through Tensor-core Garner Reformulation and Kulisch Escape Route
NVIDIA's Blackwell Ultra (B300) cuts FP64 vector throughput to ~1.3 TFLOPS per GPU, roughly 30x below B200 and well below the level at whic…
Self-Recognition Finetuning can Prevent and Reverse Emergent Misalignment
Emergent misalignment (EM) has been linked to the activation of misaligned persona vectors and evil character traits, suggesting that EM op…
Evaluating LLM Usage for Efficient and Explainable Numerical and Classified Implicit Sentiment Analysis of Product Desirability
Qualitative product feedback can reveal nuanced user experiences, but its implicit sentiment is difficult to measure. This paper presents a…
Heterogeneous 2D/1D Signal Representation Fusion for Underwater Acoustic Modulation Recognition Under Distribution Shift
Modulation recognition systems rely on heterogeneous signal representations. 2D signal-image modalities such as time-frequency and cyclosta…
Event-Aligned Analysis of Multi-Rater Pain Assessments Using Continuous Wearable Physiology
Pain is assessed differently by patients, nurses, and clinicians, yet most computational approaches assume a single ground-truth label - ef…
Coordinate-Queryable Neural Field Reconstruction for EEG Spatial Super-Resolution with Unseen-Electrode Generation
EEG spatial super-resolution (EEGSR) in real deployments is challenged by random channel missingness, unstable electrode quality, and chang…
Audio-visual Contrastive Alignment for Diffusion-based Visual-conditioned Speech Enhancement
Audio-visual speech enhancement (AVSE) exploits visual cues such as lip movements to recover speech in noisy environments. Recent work intr…
Random coloured digraphs defined by a Markov logic network
A Markov Logic Network (MLN) is a probabilistic relational model used in Statistical Relational Artificial Intelligence for defining a prob…
Legal Reasoning Is Not Lawyering: Rethinking Legal Benchmarks for Pro Se Access to Justice
Legal AI benchmark research frequently invokes the assumption that large language models can improve access to justice, including for peopl…
A Unified Framework for Runtime Verification and Model-Based Diagnosis in LOLA
We present an integrated framework that unifies runtime verification and model-based diagnosis within the stream specification language LOL…
Weight-Space Geometry of Offline Reasoning Training
Offline reinforcement-learning losses (RFT, RIFT, DFT, Offline GRPO, DPO) are widely used to distill reasoning from large teachers into sma…
A Survey on Federated Causal Discovery and Inference
Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven d…
Low-power analogue neural networks with trainable nonlinear connections for continuous control
Physical neural networks promise low-power machine learning by computing directly with analogue device physics, but most architectures forc…
Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation
Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acce…
JEDEL: Zero-Shot DNA-Encoded Library Design for Early-Stage Drug Discovery
We present JEDEL, a framework for generating synthesis-ready DNA-encoded libraries (DELs) directly from three-dimensional pharmacophore rep…
Synergizing Physically Constrained MCMC and Chemical-Informed Gaussian Processes for Reaction Network Discovery
Extracting interpretable governing equations from sparse, noisy chemical time-series data remains difficult because discrete reaction topol…
Exploring Dualistic Meta-Learning to Enhance Domain Generalization in Open Set Scenarios
Domain generalization learns from multiple source domains to generalize to unseen target domains. However, it often neglects the realistic…
VeriPilot: An LLM-Powered Verilog Debugging Framework
Verilog debugging remains one of the most time-consuming stages in digital circuit design. Recent advances in Large Language Models (LLMs)…
Engineering Reliable Autonomous Systems: Challenges and Solutions
Engineering reliable autonomous systems is an important and growing topic in computer science. As autonomous systems become more prevalent,…
Neuromorphic Speech Enhancement with Dual-Branch Spiking Neural Networks
Spiking neural network (SNN)-based neuromorphic speech enhancement has emerged as a promising paradigm due to its energy efficiency, yet it…
Listening makes Vision Clear for VLMs
Recent work typically assesses vision--language consistency using attention distributions of answer-side tokens. However, we observe that h…
Emergent Relational Order in LLM Agent Societies: From Collective Affect to Authority Stratification
Fei Xiaotong's Differential Order Pattern characterizes rural society as egocentric and relationally graded, with cooperation attenuating o…
Cryptographic certificates of validity for trustworthy AI
We propose cryptographic certificates of validity for agentic AI systems. The core idea is to formally specify a correctness or policy cond…
Integrated Sensing and Communications for Real-time Avatar Control in XR over 5G
Extended Reality (XR) presents a challenging use case for 5G and 6G networks, requiring high data-rates and lowlatency communication to del…
From Task-Guided Conversational Graphs to Goal-Oriented Dialogue Runtimes
Graph and multi-agent orchestration frameworks make production large language model (LLM) workflows practical, but they do not by themselve…
Ten Digits on a Train: AI-Assisted Verification of Two Eigenvalue Problems
Accurate numerical eigenvalues are often difficult to certify, especially in singular or non-normal settings. This article reports a human-…
From Spatial to Spectral: An Efficient, Frequency-Guided Feature Representation Learner for Small Object Detection
Efficient small object detection is bottlenecked by the inherent feature scarcity of tiny targets, which is further aggravated by operation…
Deciphering Fingerprints of 3D Molecular Surfaces for Accurate Epitope Prediction
Molecular surfaces encode the geometric and physicochemical patterns that determine antibody-antigen recognition, central to epitope predic…
Decentralized Coordination of Autonomous Traffic Through Advanced Air Mobility Corridors
The use of dedicated corridors for Advanced Air Mobility (AAM) traffic is one of the most commonly proposed pathways to integrating them in…
The Measurable Majority
This paper studies strict majority reasoning in finite electorates using so-called $\textit{social decision frames}$: finite sets of voters…
Are Safety Guarantees in Neural Networks Safe? How to Compute Trustworthy Robustness Certifications
A primary challenge in AI safety is the existence of adversarial examples -- slightly distorted inputs that cause a neural network (NN) to…
MGI: Member vs Generated Inference
As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine…
JupOtter: Cell-Level Bug Detection in Jupyter Notebooks
Jupyter Notebooks are an increasingly popular coding environment used across many domains, especially in Python-based data science and scie…
Promise and challenges of heart chamber segmentation from non-contrast CT scans using contrastive unpaired image translation: a feasibility study
Purpose: To evaluate the feasibility and challenges of heart chamber segmentation from non-contrast CT scans using contrastive unpaired ima…
One Year Later...The Harms Persist, But So Do We!
General-purpose large language models (LLMs) are increasingly used for mental health-related conversations, yet safety safeguards remain in…
Mind the Heads: Topological Representation Alignment for Multimodal LLMs
Representation alignment has emerged as an effective approach to improve Multimodal Large Language Models (MLLMs) by regularizing their int…
E-MRL: Cross-view Aligned Evidence-driven Multimodal Reinforcement Learning for Reliable 3D Tumor Analysis
While Vision-Language Models (VLMs) show great promise in volumetric medical report generation, they frequently suffer from visual hallucin…
The Professor: Multi-Teacher Unsupervised Prompt Distillation for Vision-Language Models
Prompt distillation compresses large vision-language models (VLMs) such as CLIP into lightweight student models by matching teacher predict…
ARIA: Adaptive Region-Based Importance Allocation for Conditional Diffusion Distillation
Distilling conditional diffusion models aims to transfer the behavior of a large teacher to a smaller student while preserving alignment ac…
Catastrophic Compositional Generation: Why Vanilla Diffusion Models Fail to Extrapolate
The task of compositional generation involves using a conditional generative model, trained only on a subset of the possible conditions, to…
When Retrieval Metrics Mislead: Measuring Policy Signal in Long-Horizon Tool-Use Agents
Exact-match retrieval recall is often used as a proxy for whether a retriever supplies useful policy context to a downstream decision model…
Offline Reinforcement Learning for Warehouse SLAM Throughput Control
We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment.…
Maestro Order: A Model-Agnostic Orchestration Harness
A single forward pass of a capable model is a fast, fluent, and unreliable problem-solver: it is right often enough to be useful and wrong…
Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization
End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions the…
RASC+: Retrieval-Constrained LLM Adjudication for Clinical Value Set Authoring
Clinical value sets define the standardized terminology codes used in quality measurement, phenotyping, cohort construction, and clinical d…
Learning to Trigger: Reinforcement Learning at the Large Hadron Collider
High-throughput scientific facilities such as the Large Hadron Collider depend on real-time event filtering (\textit{triggering}) under tig…
EMAgnet: Parameter-Space EMA Regularization for Policy Gradient Self-Play in Large Games
Recent work has established that regularized policy gradient methods such as PPO, when used in self-play, can match or exceed specialized g…
Towards Spec Learning: Inference-Time Alignment from Preference Pairs
Steering a large language model (LLM) toward a desired behavior typically relies on an iterative process of hand-crafting a prompt based on…
Fast and Slow Variational Continual Learning
Continual learning remains a major challenge for modern deep networks, partly because commonly used optimizers lack inherent mechanisms for…
Towards Version-aware Operations and Transaction Memories for Multi-layer MeMo
MeMo proposes language models with explicit multi-layer correlation matrix memories (CMMs), where memorization, retrieval, and forgetting a…
Rapid FinFET Modelling Using an Autoencoder
This work presents a machine learning framework that leverages an autoencoder (AE) for the efficient modeling of FinFET. We first calibrate…
RAVEN: A Regime-Aware Variable-context Expert Network for Financial Time Series Forecasting
Financial time series forecasting presents structural challenges absent from standard benchmarks. Log-returns are non-stationary, exhibit e…
Selective Capability Unlearning in End-to-End Spoken Language Understanding
Modern spoken language understanding (SLU) systems are increasingly deployed in real-world settings, where specific functionalities may nee…
Token Complexity of Certifying Stochastic-Oracle Reliability
Wang~\cite{Wang2026} introduced the Stochastic-Oracle Turing Machine (SOTM) framework and defined token complexity as the minimum expected…
End-to-End Radar and Communication Modulation Recognition with Neuromorphic Computing
Although deep learning-based methods can achieve high accuracy in automatic modulation recognition (AMR) tasks, their high computational co…
PixJail: Self-Evolving Paper-to-Pipeline Reproduction for Text-to-Image Jailbreak Evaluation
As Text-to-Image (T2I) jailbreak techniques evolve rapidly, existing benchmarks and reproduction workflows often struggle to keep pace. Mor…
CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression
"Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves…
Blockwise Policy-Drift Gating for On-Policy Distillation
On-policy distillation (OPD) trains a student policy using teacher signals computed on trajectories sampled by the student itself. Recent w…
DynaWM: Dynamics-Aware Distillation with World Model and Momentum Targets for Smooth Locomotion over Continuous Stairs
Recent advances in control have enabled bipedal-wheeled robots to traverse slopes and single-step obstacles, yet long staircase traversal r…
Predicting Poets' Origins from Verse: A Computational Analysis of Regional Linguistic Fingerprints in the Complete Tang Poems
We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work. Aggregating every poem attri…
Beyond Bayer: Task-Optimal Sensor Co-Design for Robust Autonomous-Driving Segmentation
Robust perception underpins autonomous driving, and most recent progress comes from scaling the model-larger backbones, foundation models,…
The impact of generative artificial intelligence on academic development of Chinese students in humanities and social sciences
Generative artificial intelligence(GenAI) is reshaping learning in higher education, with particularly pronounced implications for the huma…
DramaDirector: Geometry-Guided Short Drama Generation
Short dramas, with their rapid shot rhythms, dialogue-driven focus shifts, and demanding cinematographic grounding, pose challenges that pr…
A Benchmark for Hallucination Detection in VLMs for Gastrointestinal Endoscopy
Vision-language models (VLMs) are prone to hallucination, which remains a major barrier to their safe deployment in clinical practice. To d…
DTT-BSR+: A Generative-Regression Cascade for Music Source Restoration
Music source restoration (MSR) requires jointly addressing source unmixing and the inversion of non-linear production effects. Current meth…
Metis: Bridging Text and Code Memory for Self-Evolving Agents
Self-evolving agents improve over time by distilling experience from past executions and reusing it in future tasks. Existing systems repre…
Breaking Shortcut Learning for Cross-Trial EEG-Guided Target Speech Extraction via Two-Stage Training
Recent end-to-end models for EEG-guided target speech extraction report impressive results, underscoring potential for neuro-steered hearin…
A P\={a}ninian Foundation for Indic Language Processing
More than a billion people communicate in Indic languages, yet the natural language processing infrastructure serving them remains fragment…
Lightweight Transformer Models for On-Device Fault Detection: A Benchmark Study on Resource-Constrained Deployment
On-device fault detection enables real-time diagnostics without cloud dependency, but deploying machine learning models on resource-constra…
Agon: An Autonomous Large-Scale Omnidisciplinary Research System Built on Prompt Economy
Large language models are making research production scalable, shifting the bottleneck from producing artifacts to judging claims. We prese…
Zero-Shot Test-Time Canonicalization using Out-of-Distribution Scoring
Pretrained vision models often misclassify inputs that are rotated, scaled, or sheared, even though these affine transformations leave the…
Deep Learning Approaches for 3D Medical Scene Completion: From Geometric Modeling to Generative Paradigms
Three-dimensional scene completion has evolved as a major problem in computer vision and robotics, and its applications are diverse, includ…
Co-occurring associated retained concepts in Diffusion Unlearning
Unlearning has emerged as a key technique to mitigate harmful content generation in diffusion models. However, existing methods often remov…
MMed-Bench-IR: A Heterogeneous Benchmark for Multilingual Medical Information Retrieval
Retrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English eviden…
Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation
Recent breakthroughs in 3D generation have advanced notably with the development of text-to-image diffusion model. However, existing method…
AutoSpec: Safety Rule Evolution for LLM Agents via Inductive Logic Programming
Large language model (LLM) agents increasingly automate complex tasks by integrating language models with external tools and environments.…
Social Structure Matters in 3D Human-Human Interaction Generation
Although text-to-motion generation has achieved strong progress in synthesizing realistic single-person motions from language, extending it…
SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization
Fine-tuned encoders deployed across heterogeneous NLP tasks face three compounding problems: mismatched inductive biases, class-imbalance c…
Neural Network-Based Parametric Model Reduction for Predicting Turbulent Flow for Different Vehicle Geometries
Numerical simulations in industrial applications often require performing numerous high-precision computations parameterized by specific ex…
Pigeonholing: Bad prompts hurt models to collapse and make mistakes
While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradatio…
CALIBER: Calibrating Confidence Before and After Reasoning in Language Models
Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success.…
Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation
Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from thi…
ZONOS2 Technical Report
We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve up…
What Does ODRL Mean? A Cross-Level Ontological Grounding of Permissions, Prohibitions, and Duties in UFO-L
ODRL policy evaluators produce verdicts, but say nothing about the normative positions a policy brings into existence, the authority struct…
Structural Kolmogorov-Arnold Convolutions: Learnable Function on the Values or the Filter Shape as Parameter-Efficient Alternative to Per-Edge Convolutional KANs
Convolutional Kolmogorov--Arnold Networks (KANs) replace the fixed weights of a convolutional kernel with learnable univariate functions. T…
On the Stability of Prompt Ranking in Large Language Model Evaluation
Prompt-based interaction has become a dominant paradigm for using large language models (LLMs), where multiple candidate prompts are evalua…
Female-RHINO: A Real-Time Scanner-Integrated Framework for Automated Quantitative Uterine MRI Analysis and Structured Reporting
Standardized assessment of uterine MRI remains challenging due to anatomical variability, observer dependence, and the lack of workflow-int…
Average Rankings Mask Per-Subject Optimality: A Friedman-Nemenyi Benchmark of EEG Motor-Imagery BCI Decoders
Electroencephalography (EEG) is the dominant non-invasive modality for brain-computer interfaces (BCIs), yet reliable decoding of motor ima…
Entity Resolution via Batched Oracle Queries
We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity. We s…
Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories
Generative AI coding agents are entering the open-source supply chain, yet their diverse and often invisible traces leave their prevalence…
Transformation Behavior of Images in Latent Space
Training of neural networks for histopathology classification tasks typically relies on data encoding into latent space, which reduces comp…
MedPCFM: Improving Medical Point Cloud Completion by Integrating Point Transformers and Flow Matching
Medical point cloud completion is important for anatomical reconstruction and downstream clinical workflows, yet generative modeling in thi…
NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation
Perceiving physical contact is fundamental to dexterous manipulation. While robots often rely on dedicated hardware tactile sensors, humans…
The African Language Tax: Quantifying the Cost, Latency, and Context Penalty of Tokenizing African Languages in Frontier LLMs
Commercial large language models bill, scale latency, and budget context per token. Yet tokenizers assign more subword tokens to the same m…
G$^3$VLA: Geometric inductive bias for Vision-Language-Action Models
Vision-language-action (VLA) models have made rapid progress in generalist robot manipulation by harnessing semantic knowledge from pretrai…
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding
Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spa…
Adaptive Machine Learning Framework for UAV Trajectory Optimization in O-RAN
The deployment of unmanned aerial vehicles (UAV) as open radio units (O-RUs) in 6G cellular systems presents a promising opportunity to ach…
RetiSEM: Generalising Causal Models for Fragmented Biomedical Data
Learning causal models from fragmented biomedical data is challenging because clinical, molecular, and imaging variables are often incomple…
Red-Teaming the Agentic Red-Team
The use of agentic systems to perform offensive security operations has moved from a theoretical possibility to a commoditized capability.…
CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation
Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU…
A Fair Evaluation of Graph Foundation Models for Node Property Prediction
Due to the wide use of graph-structured data in different fields of industry and science, the development of Graph Foundation Models (GFMs)…
Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams
Scam phone calls exploit vulnerable communities worldwide, yet research on detection has focused almost exclusively on English and other hi…
Toward Self-Evolution-Ready Workflow Harnesses: A Reversible Migration Path and Convertibility Taxonomy for Expert LLM Pipelines
While expert-validated "LLM + script" workflows deliver significant value, they remain static: they encode hard-won domain knowledge yet fa…
Infinitesimal Causality
This paper introduces a categorical account of infinitesimal causality in Frobenius Markov categories equipped with tangent-bundle semantic…
Privacy-Preserving RAG via Multi-Agent Semantic Rewriting: Achieving Confidentiality Without Compromising Contextual Fidelity
Retrieval-Augmented Generation enhances large language models by incorporating external knowledge, but deploying it in sensitive scenarios…
Visualizing "We the People": Bridging the Perception Gap through Pluralistic Data Storytelling
Traditional visual data storytelling relies on binary graphics that depict two simplified groups in conflict. This can increase political p…
AI-PAVE-Br: Leveraging Large Language Models for Enhanced Product Attribute Value Extraction through a Golden Set Approach
The explosive growth and complexity of product data within the dynamic Brazilian e-commerce landscape demand robust and specialized methods…
FlowPipe: LLM-Enhanced Conditional Generative Flow Networks for Data Preparation Pipeline Construction
Data preparation pipelines improve data quality in machine learning by transforming raw tables into learning-ready data through sequential…
TACTFUL: Tactile-Driven Exploration For Object Localization and Identification in Confined Environments
Humans effortlessly locate and identify objects by touch alone, even without vision. In contrast, robotic systems rely heavily on vision an…
Evaluating the Interpretability of Sparse Autoencoders with Concept Annotations
Sparse autoencoders (SAEs) are increasingly used to extract interpretable concepts from vision and vision language models, yet existing eva…
Task Decomposition for Efficient Annotation
High-quality annotations of structured representations are expensive to collect over large corpora. Manual annotation of structure is labor…
Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement
Generative models, particularly diffusion and score-based approaches, have recently achieved strong performance in speech enhancement, but…
UniDrive: A Unified Vision-Language and Grounding Framework for Interpretable Risk Understanding in Autonomous Driving
Recent multimodal large language models (MLLMs) have shown strong potential for autonomous driving scene understanding, yet existing method…
Context-Aware Prediction of Student Quiz Performance with Multimodal Textbook Features
Educational platforms often predict student performance from prior interactions, but the assessment content itself also varies in linguisti…
DeepBD: A Grounded Agentic Workflow for Variant Prioritization and Diagnosis of Genetic Birth Defects
Birth defects are a major cause of fetal loss, neonatal morbidity and long-term disability. In the subset with suspected genetic etiologies…
Paying to Know: Micro-Transaction Markets for Verified Product Information in Agentic E-Commerce
Commercial NLP treats the shopping chatbot as a recommender or a conversion tool: its job is to match a user to a catalogue entry and close…
Grad Detect: Gradient-Based Hallucination Detection in LLMs
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet they remain prone to generating hallucinat…
EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence
Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Never…
OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis
Generic text-to-video models can be used as rich open-world scene priors. Despite the high quality of today's generated videos, they do not…
Large-Language-Model Discovery of Quantum LDPC Codes through Structured Concept Evolution
Quantum computers could outperform classical machines on important problems, but only if the errors that pervade quantum hardware can be co…
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation
Unified multi-modal large language models (MLLMs) have achieved strong text-to-image generation quality, but still struggle with structure-…
It's Complicated: On the Design and Evaluation of AI-Powered AAC Interfaces
Artificial intelligence (AI) can enhance what people who use augmentative and alternative communication (AAC) are able to do with their sys…
FLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse Representation
Sparse voxel representation has emerged as a scalable foundation for image-to-3D Gaussian Splatting (3DGS) generation, yet current methods…
InSight: Self-Guided Skill Acquisition via Steerable VLAs
Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in…
Random Rule Forest (RRF): Interpretable and Manageable Ensembles of LLM-Generated Questions for Predicting Success from Unstructured Data
Many high-stakes screening tasks require predicting rare outcomes from unstructured text, where errors are costly and decisions must be aud…
TIP-Search: Time-Predictable Inference Scheduling for Market Prediction under Uncertain Load
Real-time market prediction services need correct predictions before a decision deadline; a correct prediction delivered late is not usable…
From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control
Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated,…
A global log for medical AI
Modern computer systems rely on syslog, a universal protocol that records critical events across heterogeneous infrastructure. Medicine's r…
Representation Interventions Enable Lifelong Knowledge Memory Control in LLMs
Large language models (LLMs) often produce incorrect or outdated content after being employed. Efficient and accurate knowledge updates wit…
Evolving Programmatic Skill Networks
We study continual skill acquisition in open-ended embodied environments where an agent must construct, refine, and reuse an expanding libr…
BioPIE: A Biomedical Protocol Information Extraction Dataset for Experiment Understanding
Understanding biomedical experiments provides a foundation for downstream tasks, e.g., laboratory automation, and facilitates effective cro…
LLM-MINE: Large Language Model based Alzheimer's Disease and Related Dementias Phenotypes Mining from Clinical Notes
Accurate extraction of Alzheimer's Disease and Related Dementias (ADRD) phenotypes from electronic health records (EHR) is critical for ear…
Grounded Chess Reasoning in Language Models via Master Distillation
Language models often lack grounded reasoning capabilities in specialized domains where training data is scarce but bespoke systems excel.…
Subjective-Graph LLM Agents for Simulating Uncertainty in Classroom Social Perception
Social actors do not observe a common social world: each individual forms judgments from a partial and potentially distorted view of the su…
Riemann-Bench: A Benchmark for Moonshot Mathematics
Recent AI systems have achieved gold-medal-level performance on the International Mathematical Olympiad, demonstrating remarkable proficien…
Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization
Multi-Hop Fact Verification requires complex reasoning across disparate evidence, posing significant challenges for Large Language Models ,…
BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accura…
2.5-D Decomposition for LLM-Based Spatial Construction
Autonomous systems that build structures from natural-language instructions need reliable spatial reasoning, yet large language models (LLM…
Efficient Test-time Inference for Generative Planning Models with OCL Search
Generative models have emerged as a powerful paradigm for AI planning, yet their performance remains constrained by the training data distr…
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into la…
EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent
As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly i…
BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling
Large language models (LLMs) are increasingly applied to computer-aided design (CAD) to generate design artifacts from textual instructions…
Repeated Shared Access Enables Grokking, but Edit Propagation Depends on an Addressable Memory
We study factual edit propagation in a controlled synthetic knowledge-graph QA setting using a 2x2 grid that crosses loop recurrence with s…
When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models
Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two out…
IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO
Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier langu…
HOLMES: Evaluating Higher-Order Logical Reasoning in LLMs
Logical reasoning is essential for reliable AI, yet existing benchmarks are largely first-order-logic-centric, focusing on object-level ded…
Invariant Graph Representations for Continuous-Time Dynamic Graphs Under Distribution Shifts
Continuous-Time Dynamic Graphs (CTDGs) enable fine-grained modeling of evolving relational systems. However, most existing CTDG representat…
When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments
Can AI Agents simulate real-world trading environments to investigate the impact of external factors on stock trading activities (e.g., mac…
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
AI agents have the potential to aid users on a variety of consequential tasks, including conducting scientific research. To spur the develo…
Variational Model Merging for Pareto Front Estimation in Multitask Finetuning
Pareto fronts are useful to find good task-mixing strategies for multitask finetuning, but they are also costly to compute. To reduce costs…
Impatient Bandits: Optimizing for the Long-Term Without Delay
Increasingly, recommender systems are tasked with improving users' long-term satisfaction. In this context, we study a content exploration…
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
Recent studies have raised significant concerns regarding the reliability of current mathematics benchmarks, highlighting issues such as si…
Societal Alignment Frameworks Can Improve LLM Alignment
Recent progress in large language models (LLMs) has focused on producing responses that meet human expectations and align with shared value…
Reward-Centered ReST-MCTS: A Robust Decision-Making Framework for Robotic Manipulation in High Uncertainty Environments
Monte Carlo tree search is attractive for robotic manipulation because it can improve action selection through simulation without requiring…
Ensemble Learning for Large Language Models in Text and Code Generation: A Survey
Generative Pretrained Transformers (GPTs) are foundational Large Language Models (LLMs) for text generation. However, individual LLMs often…
Multimedia and Visual Analytics in the Agentic Era
Professional users need tools to help them gain actionable insights from large multimedia collections. Foundation models and AI agents have…
MuTRAP: Multi-trigger Trojans Attacking Robot Task Planning Systems
Robots need task planning methods to achieve goals that require more than one action. Recently, large pretrained models have demonstrated i…
Minimisation of Quasar-Convex Functions Using Random Zeroth-Order Oracles
This paper explores the performance of a random Gaussian smoothing zeroth-order (ZO) scheme for minimising quasar-convex (QC) and strongly…
SEAL: Searching Expandable Architectures for Incremental Learning
Incremental learning is a machine learning paradigm where a model learns from a sequential stream of tasks. This setting poses a key challe…
Graph Alignment for Benchmarking Graph Neural Networks and Learning Positional Encodings
We propose a novel benchmarking methodology for graph neural networks (GNNs) based on the graph alignment problem, a combinatorial optimiza…
Render-FM: Feedforward Model for Real-time Photorealistic Volumetric Rendering
Photorealistic volumetric rendering of CT scans greatly benefits clinical workflows, yet neural approaches such as Neural Radiance Fields (…
Tuning without Peeking: Provable Generalization Bounds and Robust LLM Post-Training
Gradient-based optimization is the workhorse of deep learning, offering efficient and scalable training via backpropagation. However, expos…
FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation
Industrial signal analysis is hindered by severe data heterogeneity, which we characterize as the M5 problem. Existing solutions rely on sp…
Rule2Text: A Framework for Generating and Evaluating Natural Language Explanations of Knowledge Graph Rules
Knowledge graphs (KGs) can be enhanced through rule mining; however, the resulting logical rules are often difficult for humans to interpre…
FALCON: Transforming Cyber Threat Intelligence into Deployable IDS Rules with Self-Reflection
Signature-based Intrusion Detection Systems (IDS) detect malicious activity by matching network or host events against predefined rules. Se…
Breaking the Mirror: Activation-Based Mitigation of Self-Preference in LLM Evaluators
Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor t…
VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing
The energy cost of Large Language Model (LLM) inference is rapidly becoming a barrier to sustainable and scalable deployment. Although mode…
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
Personalized object detection aims to adapt a general-purpose detector to recognize user-specific instances from only a few examples. Light…
ATHENA: Agentic Team for Hierarchical Evolutionary Numerical Algorithms
Progress in computational science depends on complex numerical workflows that must faithfully encode physical laws, yet translating concept…
Computing Evolutionarily Stable Strategies in Imperfect-Information Games
We present an algorithm for computing evolutionarily stable strategies (ESSs) in symmetric perfect-recall extensive-form games of imperfect…
EMFusion: Uncertainty-Aware Conditional Diffusion Model for Multivariate Narrow-band Exposure Forecasting
The rapid growth in wireless infrastructure has increased the need to accurately estimate and forecast electromagnetic field (EMF) levels t…
Attention in Motion: Secure Platooning via Transformer-based Misbehavior Detection
Vehicular platooning promises transformative improvements in transportation efficiency and safety through the coordination of multi-vehicle…
Disentangling Aleatoric and Epistemic Uncertainty in Physics-Informed Neural Networks. Application to Insulation Material Degradation Prognostics
Physics-Informed Neural Networks (PINNs) provide a framework for integrating physical laws with data. However, their application to Prognos…
The $\mathbf{P}$-Completeness of Inverted Index Traversal: On the Complexity of Evaluating Boolean Query DAGs
Modern AI agents increasingly rely on search infrastructure to execute complex, neuro-symbolic reasoning workflows. These workflows often c…
Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations
Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of aut…
Toward Autonomous O-RAN: A Multi-Scale Agentic AI Framework for Real-Time Network Control and Management
Open Radio Access Networks (O-RAN) promise flexible 6G network access through disaggregated, software-driven components and open interfaces…
Event-Grounded Question Answering over Long Audio via Structured Retrieval
Answering natural-language questions over multi-hour audio requires both event recognition and temporal grounding. Current large audio-lang…
MyoInteract: A Framework for Fast Prototyping of Biomechanical HCI Tasks using Reinforcement Learning
Reinforcement learning (RL)-based biomechanical simulations have the potential to revolutionise HCI research and interaction design, but cu…
Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators
Neural network accelerators have been widely applied to edge devices for complex tasks like object tracking, image recognition, etc. Previo…
No Certificate, No Categorical Speech Act: A Brouwerian Assertibility Constraint for Public Reason
Generative AI can convert uncertainty into authoritative-seeming verdicts, intensifying the hypersuasive force of automated speech and disp…
An Approach to Simultaneous Acquisition of Real-Time MRI Video, EEG, and Surface EMG for Articulatory, Brain, and Muscle Activity During Speech Production
Speech production is a complex process spanning neural planning, motor control, muscle activation, and articulatory kinematics. While the a…
CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance
We introduce CRAFT hand, a tendon-driven anthropomorphic hand with hybrid hard-soft compliance for contact-rich manipulation. The design is…
AI-Driven Predictive Maintenance with Environmental Context Integration for Connected Vehicles: Simulation, Benchmarking, and Field Validation
Predictive maintenance for connected vehicles offers the potential to reduce unexpected breakdowns and improve fleet reliability, but most…
HiPath: Hierarchical Vision-Language Alignment for Structured Pathology Report Prediction
Pathology reports are structured, multi-granular documents encoding diagnostic conclusions, histological grades, and ancillary test results…
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polish…
WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models
Recent decoder-only autoregressive text-to-speech (AR-TTS) models produce high-fidelity speech, but their memory and compute costs scale qu…
THEIA: Learning Complete Kleene Three-Valued Logic in a Pure-Neural Modular Architecture
We present THEIA, a 2.75M-parameter modular neural architecture that learns the complete Kleene three-valued logic (K3) truth table from ta…
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While rec…
Fix Initial Programs and Iteratively Refine Repair Instructions Toward Non-Elimination Multi-Turn Program Correction
Recent work on large language models (LLMs) has emphasized the importance of scaling inference compute. From this perspective, the state-of…
DynamicPO: Dynamic Preference Optimization for Recommendation
In large language model (LLM)-based recommendation systems, direct preference optimization (DPO) effectively aligns recommendations with us…
Ensemble Distributionally Robust Bayesian Optimisation with Continuous Context
We study Bayesian Optimisation (BO) in settings where the objective function is influenced by uncontrollable environmental contexts governe…
When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely h…
A Simplex Witness Certificate and Escape Force for Constant Collapse in Variational Autoencoders
We study exact constant collapse in variational autoencoders: the deterministic encoder mean becomes independent of the input. The prior re…
Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment
Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in…
Sensing Intelligence as a Trainable Metamaterial Property
In biological systems, sensing is not performed by the brain alone: the body deforms, vibrates, and filters external stimuli before they ar…
More Skills, Worse Agents? Skill Shadowing Degrades Performance When Expanding Skill Libraries
Skill libraries allow LLM agents to load task-specific instructions on demand, letting non-expert users solve domain-specific tasks through…
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
We present VISTA (VIsual Spec-To-App Benchmark), a benchmark for evaluating the end-to-end web-app generation capabilities of LLM-based age…
QSignAI: Quantum-Randomness-Seeded Identity Signatures at the Intersection of AI for Science and Science for AI
The 2024-2025 Nobel and Turing awards recognised AI and quantum science simultaneously. Yet no deployed system has brought these streams to…
Cosmos 3: Omnimodal World Models for Physical AI
We introduce Cosmos 3, a family of omnimodal world models designed to jointly process and generate language, image, video, audio, and actio…
ASymPO: Asymmetric-Scale Policy Optimization for Asynchronous LLM Post-Training Without Behavior Information
Asynchronous reinforcement learning can improve language-model post-training throughput by decoupling response generation from policy optim…
A Training-Free Mixture-of-Agents Framework for Multi-Document Summarization using LLMs and Knowledge Graphs
Multi-Document Summarization (MDS) plays a critical role in distilling essential information from collections of textual data. Existing app…
Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing
Purpose: Digitization projects in the humanities produce vast, heterogeneous archives of historical documents, making manual sorting imprac…
AI-Driven Analytics of Team-Teaching Talk: Acoustic Patterns across Experience, Cohorts and the Learning Design
As classroom cohorts expand, team teaching is increasingly used to integrate the expertise and pedagogical perspectives of multiple teacher…
FedSteer: Taming Extreme Gradient Staleness in Federated Learning with Corrective Projections and Caching
Federated learning (FL) is often subject to aggregation variance if clients do not consistently participate in training rounds. While reusi…
Acquisition state behaves as a structured, measurable variable governing lung-nodule AI: kernel-driven measurement instability and noise-driven detection fragility, invisible to DICOM metadata
AI governance for medical imaging is formalizing: the 2026 ACR-SIIM Practice Parameter recommends local acceptance testing and ongoing drif…
AgentRivet: an automated system for producing Rivet routines from journal publications
Particle physics collider experiments provide Rivet routines as part of the analysis preservation strategy for model-independent measuremen…
Surprise-Guided MergeSort: Budget-Efficient Human-in-the-Loop Ranking via Adaptive Comparison Scheduling
Pairwise comparison is the gold standard for subjective ranking tasks; however, exhaustive annotation requires a massive number of human co…
Lect\=uraAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching
Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but…
Robust Dual-Signal Fusion: Hybrid Neuro-Symbolic Gating with Compressed Chain-of-Thought Refinement for Irony Detection in Social Media Texts
Small-scale Large Language Models (LLMs) natively default to literal semantic interpretations, making few-shot irony detection a persistent…
Quantum Cinema: An Interactive Cinematic Exploration of Quantum Computing Hardware via Generative World Models
Quantum computing promises transformative advances across science and industry, yet the physical hardware that enables these computations r…
Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference
Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, i…
KANLib -- A Modular, Extensible and Fast Kolmogorov-Arnold Network Implementation
Kolmogorov-Arnold Networks (KANs) have recently emerged as a promising alternative to traditional multilayer perceptrons by replacing linea…
Essential Subspace Merging for Multi-Task Learning
Model merging aims to enable multi-task learning by integrating the capabilities of multiple models fine-tuned from the same pre-trained ch…
Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR
The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired ar…
HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training
Leveraging Vision Foundation Models (VFMs) for camera-to-LiDAR knowledge distillation offers a promising solution to the scarcity of annota…
Topological Neural Dynamics: A Neuron-wise Framework for Sequence Modeling
Existing sequence models, including RNNs, LSTMs, continuous-time networks, and Transformers, share a common structural principle: layer-wis…
Sexualised synthetic personas encode and amplify gendered power asymmetries through voice
This work examines sexualised AI-generated English-speaking voices offered by a popular commercial platform. New technologies may enable se…
Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study
Mixture-of-Experts (MoE) language models are often described as ideal for resource-constrained inference. Each token activates only a small…
Skills for the future software profession: beyond agentic AI!
As coding agents are rapidly changing software engineering, a natural question is: what are the core skills needed by future software engin…
Alternate loss functions and regression models that achieve robustness to outliers by modulating the learning rate
Most real-world datasets used for training supervised learning models are contaminated with noisy data and outliers leading to large predic…
MultiMem: Measuring and Mitigating Memorization in Multi-Modal Contrastive Learning
Memorization in machine learning models enables high performance on rare in-distribution samples by capturing their atypical patterns. Howe…
Diffusion Integrated Gradients: Controllable Path Generation for Flexible Feature Attribution
Path-based attribution methods such as Integrated Gradients (IG) are widely adopted for their strong axiomatic properties and effectiveness…
On the Position Bias of On-Policy Distillation
On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision fro…
AI Fiction in the Wild
Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, to…
Polycepta: Object-Centric Appearance Estimation for Multi-Object Tracking
The tracking-by-detection paradigm in multi-object tracking (MOT) typically relies on static appearance descriptors to complement motion es…
主婦がAIアニメでYouTube登録者100万人を突破――1億ユーザーのAI動画生成サービス「PixVerse」の実態
1億人超のユーザーを持つAI動画生成プラットフォーム「PixVerse」とは何か。運営企業の担当者が活用事例などを語った。
「Apps in ChatGPT」にメルカリ登場 自社MCP基盤を活用、会話で商品検索など
1月に公開したAI接続基盤「Mercari MCP」(Model Context Protocol)を活用。ChatGPTでの会話を通じて商品を検索したり、出品時の説明文の下書きを作成したりできる。
「夏場は50度以上のコンテナで作業」に対処 サンワサプライが西日本で荷降ろしロボット活用
サンワサプライが、物流倉庫における荷降ろし作業の自動化と労働環境の改善を目的に、AI搭載のコンテナ向け荷降ろしロボット「RockyOne」を採用した。5月から同社の西日本物流センターで運用を開始している。
AmazonはNVIDIAに挑戦状を突きつけるのか
世界最大のハイパースケーラーであるAWSは、AIアクセラレーターを大規模に販売することで、半導体市場の好機を捉えようとしているのだろうか。
Mistral、文書解析OCRの新版「OCR 4」公開 文字の位置や信頼度スコアを出力、日本語を含む170言語に対応
Mistral AIは、文書のテキストや構造を抽出するOCRモデルの最新版「Mistral OCR 4」を公開した。日本語を含む170言語に対応し、要望の多かったバウンディングボックスや信頼度スコアの出力に対応した。単一コンテナによる自己ホスティングも可能で、厳格なデータ主権や…
Anthropic、Slackで「@Claude」を呼べる「Claude Tag」提供──チームの一員として非同期でタスク遂行
Anthropicは、チーム向け新機能「Claude Tag」を発表し、「Slack」での提供を開始した。チャンネル内で「@Claude」とメンションすることで、各種ツールやコードベースと連携し、非同期かつ自律的にタスクを遂行する。基盤モデルには「Opus 4.8」が採用され、…
India’s MoEngage bets that the future of marketing is millions of AI agents
The all-cash deal gives MoEngage access to technology that assigns AI agents to individual customers.
再生師、テクスチャー翻訳家……オカムラがAIで導く2045年“未来の職業”
オカムラは、自社の保有特許とAIを掛け合わせて導き出した「まだ存在しない未来の職業展 2045」を開催。手の動きで環境音を奏でる「ジェスチャーオーケストラ」や、対話から人生を再仕立てする「エンディングエディター」といった体験型展示を披露した。
「最初は壊れ過ぎてビビった」──1220億円投じたソフトバンク「AIスパコン」、それでもNVIDIAのGPUを選ぶワケ
AIブームの波に乗って時価総額世界1位に躍り出たNVIDIA。一体なぜ、AIインフラにNVIDIA製GPUが採用されるのか。その理由をAIスパコン開発者に聞いた。
Geminiで「AIを使いたい現場」と「ダメと言う会社」のギャップを埋める方法
生成AIやAIエージェントの取り組みが活発化している。ただ、その取り組みが検証止まりになるケースも多い。グーグル・クラウド・ジャパンの北瀬公彦氏は、そうした「AI活用のカベ」を乗り越えるために「スピードと守りを両立が重要」と語った。
「賞金1000万円」コンテストに「AI番付」 サイバーエージェントが明かす、AIを使い倒させる仕組み作り
AI活用を一部の意欲的な社員にとどめず、組織全体の競争力へとつなげるには何が必要なのか。全社的にAI活用を進めているサイバーエージェント AIオペレーション室長の上野千紘氏と、フリー 常務執行役員の前村菜緒氏が語った。
“中国ヒューマノイド革命”はなぜ起きた、異業種や大手テックが動かす市場の今
中国のヒューマノイドロボット市場は、劇的なパラダイムシフトの渦中にある。出荷台数は前年比約7倍、世界シェアは8割に達し、異業種企業の参入で本体企業数は倍増した。野村総合研究所の李智慧氏による、量産化フェーズへ突入した中国市場の急成長を支えるマクロ動向の解説を紹介する。
Anthropic’s Claude Tag is learning your company, one Slack message at a time
Anthropic’s new Claude Tag brings an always-on AI teammate to Slack. But beyond productivity, the feature is a strategic play to capture or…
How GPT-5 helped immunologist Derya Unutmaz solve a 3-year-old mystery
GPT-5 Pro helped solve a 3-year-old immunology mystery, offering insights into T cell behavior. The breakthrough could support cancer and a…
2026-06-23(24件)
4 days left to save up to $190 on TechCrunch Founder Summit 2026
Four days left to save up to $190 on your pass to TechCrunch Founder Summit 2026 — the ultimate founder bootcamp — before Early Bird rates…
Fika Jobs raises $4M to build a video-first hiring platform where AI agents interview candidates
Stockholm-based startup Fika Jobs is building a video-first hiring platform that combines AI interview agents with short-form video profile…
Helping build shared standards for advanced AI
OpenAI helps build shared standards for advanced AI, supporting evaluation frameworks, safety practices, and global cooperation through the…
業務でAIを使う人の約38%「禁止されても利用継続」 セキュリティ企業が調査
業務でAIを使っている人の37.8%が勤務先に禁止されても利用を継続する意向を示した――Webセキュリティサービスなどを手掛けるサイバーセキュリティクラウドは、このような調査結果を発表した。
国産AI「Sakana Fugu」なぜドル建て? 円建てニーズ「受け止める」とSakana AI
円建てプランへのニーズは「日本のユーザーの皆様からいただくご意見・ご要望として、引き続きしっかりと受け止める」という。
NRIセキュア、未公表の脆弱性を「Mythosと同等のレベルで」検出する診断サービス提供
「米AnthropicのClaude Mythos Previewと同等のレベルで未公表の脆弱性を検出できる」のが売り。Anthropicの日本法人代表も新サービスにコメントを寄せている。
AIに自然言語で3Dモデル作成を頼んでみたら?
生成AI×3D CADの現在地。
日立、メインフレーム事業から撤退へ ハード製造終了から9年後の決断
日立製作所が、自社のメインフレーム環境提供から撤退する。同社のメインフレーム事業撤退の経緯を整理する。
The running list: major tech layoffs in 2026 where employers cited AI
A running look — in reverse chronological order — at the bigger tech companies that have announced significant layoffs this year with AI as…
ダイハツ、自動車部品のキズ検査をAIで自動化 “人の目と感性”を代替
アルミ加工ラインで生産されているトランスミッション用の部品について、AIで加工穴内部のキズなどを検査。人の目と感性に頼っていた検査工程を自動化した。
OpenAI launches new initiative to help find and patch open-source bugs
OpenAI is attempting to tackle the security issues of the open source software community.
Copilotの“元”は取れるのか問題、ついに決着? 住友商事、京都市が掴んだ「AI活用の勝ち筋」
生成AIの導入効果が問われる時期に入っている。Microsoftが試算した削減見込み額を独自検証した住友商事が出した結論とは。住友商事と京都市の生成AI活用例を紹介し、生成AI投資の勝ち筋に迫る。
AIに詳しくなくても大丈夫、月額制で中小企業のAI活用をプロが支える新サービス
ソルバは、AIアークスと提携し、地域企業向けのAI顧問サービスを開始した。AI・DXの専門人材による伴走支援を月額制で提供し、企業ごとの課題に応じたAI活用を後押しする。生産性向上や原価改善、人手不足対策を通じて、地域企業の競争力強化を目指す。
経営者は「何を捨てるか考えて」――中東紛争“サイバー戦”のリアルが突き付ける、無視できない脅威
サイバー戦やAIの悪用など、サイバー空間は“危険と隣り合わせの場所”になった。企業の経営者は、この課題にどう向き合えばよいのだろうか。
Excelの10万行データを3分でAIに処理させる、M365 Copilotの使い方
上司から突然「このデータを見ておいて」と任され、Excelを前に手が止まった経験がある人も多いはずだ。大量データの集計は負担が大きく、分析業務は特定の担当者に偏ってしまう。Copilotのチャットや関数機能は、その前提を変えつつある。
「取引先経由」のサイバー攻撃が増加 狙われる中小企業が見直すべき「セキュリティ対策」の3ステップ
企業を狙ったサイバー攻撃が後を絶たない。2025年後半に、アスクルやアサヒグループホールディングスが被害を受け、商品の受注や出荷が停止したことは記憶に新しい。ローカーや交渉人、実行組織などが絡み合う犯罪集団に、人手やリソースが限られる中堅・中小企業が単独で立ち向かうのは難しい。…
ダイハツがAI品質検査システムを共同開発、アルミ加工穴内部の目視検査を自動化
ダイハツ工業は、滋賀(竜王)工場 第1地区にAIを用いた自動車部品の品質検査システムを導入したと発表した。製造業向けAIソリューションを提供するスタートアップのVRAIN Solutionと共同開発したもので、現場主導のDX推進の取り組みに位置付けられる。
The AI world is getting ‘loopy’
The loop takes agentic AI a step further by authorizing a swarm of agents to work continuously in the background, endlessly.
AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia’s $20B not-acqui-hire deal
What does an AI company do after one of those not-acqui-hire deals? Groq raised money, is leaning into its neocloud business, and is hiring…
Nvidia wants to cut data center water use, but that’s not the same as fixing AI’s water problem
Nvidia announced a new cooling system that cuts water use inside the data center. But it does nothing to address AI's biggest water use — f…
ループエンジニアリングとは? チャットとAIコーディングの往復から卒業する新しい開発スタイル
AI開発で人間がプロンプトを書く時代は終わるのか。ループを回してAIエージェントを動かし続ける新概念「ループエンジニアリング」の基本を、提唱者の記事に沿って解説。チャットAIとAIコーディングの往復から卒業したい筆者の考えも示す。
Google DeepMind bets $75M on AI’s future in Hollywood with A24 deal
Google DeepMind and A24 are teaming up to build AI filmmaking tools.
Amazon is testing Alexa+ in India with Hindi support
Amazon is planning to increase the footprint of its new conversational AI assistant Alexa+ to India and is inviting users in the country to…
SpaceX inks compute deal with Reflection AI, an open source AI lab
Reflection AI will pay $150 million a month beginning July 1, 2026 through 2029 for immediate access to Nvidia's latest GB300 AI chips and…
2026-06-22(22件)
The founder conference built for growth: TechCrunch Founder Summit pass rates increase June 26
Save up to $190 on your pass to TechCrunch Founder Summit 2026 by June 26, 11:59 p.m. PT. Designed for founders first on November 4 in Bost…
Claude Codeに指示を出す「7つの方法」と使い分け 公式が解説
米Anthropicは、AIコーディング支援ツール「Claude Code」に指示を出す7つの方法を解説する公式ブログを公開した。各方法の概要や使い分け方などを紹介している。
Patch the Planet: a Daybreak initiative to support open source maintainers
OpenAI introduces Patch the Planet, a Daybreak initiative helping open-source maintainers find, validate, and fix vulnerabilities with AI a…
Daybreak: Tools for securing every organization in the world
OpenAI introduces new Daybreak tools, including Codex Security and GPT-5.5-Cyber, to help organizations find, validate, and patch vulnerabi…
踏切に取り残された人をAIで検知→列車を自動停止 小田急が実運用
遮断機が降りた後に取り残しを検知すると、信号設備と連動し、接近する列車を停止させるための信号を発するとともに、乗務員に危険を知らせてブレーキを操作させる。
日立はAX事業にどう臨む? 徳永CEOの話から「成長につなげるための勘所」を探る
企業はAIトランスフォーメーション(AX)にどう取り組めばよいのか。日立のAX事業に臨む姿勢から、その勘所を探る。
国産フルスクラッチAI「PLaMo 3.0 Prime」提供開始、“高コスパ”うたう 無料のAPIプランも PFN
Preferred Networksは、AIモデル「PLaMo 3.0 Prime」の提供を始めると発表した。
OpenAIが明かす、新職種「FDE」の実態 半年で様変わり、「仕事の7割が消滅」したことも
新職種「FDE」の実態はどのようなものなのか。米OpenAIのColin Jarvis氏(Global Head of Forward Deployed Engineering)が語った。
「ChatGPTにうちの会社が出てこない」──採用担当を悩ます“AI就活時代”の容赦なき実態
採用やHRの界隈で、じわじわと語られ始めている問題意識がある。就活生、それも優秀な層ほど、キャリア相談をLLMにするようになった。その結果、Web上での露出が乏しくLLMO対策をしていない企業は、学生から“見つけてもらえなく”なってきている――というものだ。
Sakana AI、一部「ミュトス越えの性能」うたうAIを提供 複数モデルの“集合知”を活用
Sakana AIは、複数のAIモデルを組み合わせてタスクをこなすAIシステム「Sakana Fugu」の提供を始めると発表した。
日経のAI、NTTデータが販売契約 出典明示でハルシネーションに強い法人向け「NIKKEI KAI」
日本経済新聞社とNTTデータが6月19日、日経の法人向け生成AIサービス「NIKKEI KAI」の販売契約を結んだと発表した。
Anthropicへの500万ドル間接出資を解消、広告事業のイオレ 軸足移すAIデータセンター事業に資金投入
イオレは6月19日、3月に決めた米Anthropicへの間接出資を解消し、出資金500万ドル(約7億9355万円)全額の返還を受けると発表した。返還資金は自社で開発を進めるAIDCへの投資に充てる。
7割超の企業はシャドーAIを管理できていない ガートナーがガバナンスの現実解を提唱
AIの能力向上に伴って、シャドーAIのリスクも増している。ガートナーの調査によると、国内企業の73%はシャドーAIを管理できていないという。同社が推奨する、事業部門を巻き込んだガバナンスの仕組みとは。
シャドーAI対策「7割が未着手」 「AIは全て禁止」は限界 IT部門が採るべき一手とは? Gartner提言
生成AIの爆発的な普及に伴い、企業のITガバナンスは新たな局面に直面している。情報システム部門が抱えてきた旧来のシャドーSaaSといった問題に、個人契約のAIツールやローカルLLMなど幾つものリスクが積み重なった「難局」を迎えているためだ。限られたリソースで推進と統制をどう両立…
Samsung Electronics brings ChatGPT and Codex to employees
Samsung Electronics deploys ChatGPT Enterprise and Codex to employees worldwide, marking one of OpenAI’s largest enterprise AI rollouts.
AIに頼ると技術が落ちる? 医師・エンジニアたちの懸念、検証結果は……Natureも警鐘
AIツールが職場に普及するにつれて、専門家たちが長年培ってきたスキルが衰えてしまうのではないかという懸念が広がっている。
定番データベースを捨て、あのコーディングAIにエンジニアたちが群がる理由
技術の世界では、長く使われてきた“定番”が必ずしも選ばれ続けるとは限らない。データベースとAI開発ツールという異なる領域で起きている変化は、エンジニアが重視する基準の変化を映し出しているのかもしれない。
情シスが「日本1位のAIスパコン」作るまで 猶予は4カ月、ソフトバンク“社長プロジェクト”の舞台裏
ソフトバンクのスパコンが、AI計算性能で国内1位を獲得した。実は、これを開発したのは「情シス」のメンバーだった。“社長プロジェクト”に挑んだメンバーの奮闘を追う。
エッジAIと無線リモコンを融合、DMG MORI Digitalが建機操作支援システム
双葉電子工業とDMG MORI Digitalは、「第8回 国際 建設・測量展(CSPI 2026)」において、エッジAIの画像認識技術と無線リモコンを組み合わせた、建設重機の遠隔操作支援システムを展示した。
インテル「シリーズ3」はフィジカルAIでも力を発揮、“ヤマネコ”の実力は?
インテルがメインストリームPC向け製品「インテル Core シリーズ3 プロセッサー」と、ハンドヘルドゲーミングPC向け製品「インテル Arc G3 プロセッサー」について説明。また、「COMPUTEX TAIPEI 2026」に併せて発表したエッジAI/フィジカルAI向けソリ…
千葉県印西市はなぜ「データセンターの聖地」になったのか Google、Microsoftを呼び込んだ半世紀前の“読み違い”
東京都心から北東へおよそ40キロメートルにある千葉県印西市。かつて「千葉ニュータウン」として開発されたこの郊外の街が、いま世界の巨大IT企業にとって“聖地”になっている。
When the Trump administration cracks down on Anthropic, who benefits?
On the new episode of Equity, we discussed what actually prompted the administration's latest moves against Anthropic, and what this might…