Skip to the content.

トピック: ハードウェア/半導体

該当記事 682 件 / 新しい順

← トップに戻る

2026-08-17 07:00 JSTITmedia AI+ハードウェア/半導体

【読まれた記事1位】「AI需要で半導体不足」の裏で本当に起きていること――2026年前半まとめ

AIブームを追い風に、半導体市場の成長が止まらない。今、半導体市場で何が起きているのか。注目記事をまとめた。

2026-08-15 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

立場: アライメント コミュニティは意図せずに検閲ツールキットを構築している

この意見書では、現代の AI 調整手法は、本来は有害な出力を防ぐために設計されたものであり、悪意のある行為者によって検閲や操作のために簡単に悪用される可能性がある二重用途技術であると主張しています。現在の調整技術を悪用の可能性と実際のケースにマッピングすることで、「完全に調整された」モデルの探求が、意図せずして悪意のある攻撃者に情報支配のための絶えず改良されたツールを提供してしまうことを示します。ユーザーによる情報プロバイダーとしての AI の急速な導入、経済力の非対称性、権威主義への移行が進む政治情勢によってそのリスクが悪化しているため、私たちはこの二重利用の可能性について今すぐ議論する必要があります。私たちはコミュニティに対し、AI 調整メカニズムの意図的な誤用を考慮し、この二重使用の可能性を防ぐための緩和戦略を提案することを強く求めて締めくくります。

原文 (English)

Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential.

2026-08-15 13:00 JSTarXiv cs.AIハードウェア/半導体

CAS: ローカルおよびグローバルの説明可能な人工知能の因果関係スコア

予測説明メソッドはモデルの出力に帰属します。それら自体は、介入効果が現実世界の結果に影響を与えるものではありません。因果関係を説明するためのコンパクトなスコア アーキテクチャである Causal Attribution Score (CAS) を紹介します。 CAS は、特定された介入連合ゲームから開始し、共同介入コントラストを因果関係のある Shapley 寄与と割り当て、それらの生の結果スケール効果をローカル CAS、署名済みローカル CAS、および 2 つの補完的なグローバル CAS 要約に変換します。このイノベーションは、新しい Shapley 式ではなく、明示的な介入ターゲットを備えたローカルからグローバルへの因果関係レポート層です。既知の真実のベンチマークでは、8 回繰り返された一次相互作用シミュレーション (それぞれ n = 2,200、3 つのアクション) で、連合認識 CAS の平均ローカル CAS MAE が 0.107 であったのに対し、一度に 1 つずつ正規化した場合は 0.173、グローバル正規化された絶対 ATE ベクトルでは 0.213 でした。一度に 1 つずつ正規化を行う場合のペアの利点は、相加性の場合の -0.003 から、強い相互作用の場合の 0.091 に増加しました。 DoubleML の経験的データセット、401(k) 資格/純金融資産 (n = 9,915) およびペンシルベニア州の再雇用ボーナス/失業期間 (n = 5,099) の両方において、予測 SHAP/TreeSHAP ランキングは、治療効果修飾因子の Feature-CAS ランキングとは大きく異なりました。ペンシルベニア州では、dep1 (正確に 1 つの依存関係) が予測グローバル ランク 13 から Feature-CAS ランク 2 に移動し、主要なローカル Feature-CAS 修飾子となりました。これらの結果は、結果を予測するものと、推定された因果効果の不均一性を説明するものを分離するという付加価値を分離します。

原文 (English)

CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence

Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS starts from an identified interventional coalition game, allocates the joint intervention contrast with causal Shapley contributions, and converts those raw outcome-scale effects into Local CAS, Signed Local CAS, and two complementary Global CAS summaries. The innovation is not a new Shapley formula, but a local-to-global causal reporting layer with an explicit intervention target. In the known-truth benchmark, eight repeated primary-interaction simulations (n = 2,200 each, three actions) gave mean Local CAS MAE of 0.107 for coalition-aware CAS, compared with 0.173 for one-at-a-time normalisation and 0.213 for a global normalised absolute ATE vector. The paired advantage over one-at-a-time normalisation increased from -0.003 under additivity to 0.091 under strong interactions. On both empirical DoubleML datasets, 401(k) eligibility/net financial assets (n = 9,915) and Pennsylvania reemployment bonus/unemployment duration (n = 5,099), predictive SHAP/TreeSHAP rankings differed materially from Feature-CAS rankings of treatment-effect modifiers. In Pennsylvania, dep1 (exactly one dependent) moved from predictive global rank 13 to Feature-CAS rank 2 and was the leading local Feature-CAS modifier. These results isolate the added value of separating what predicts the outcome from what explains heterogeneity in an estimated causal effect.

2026-08-15 13:00 JSTarXiv cs.AIハードウェア/半導体

大規模な有限集合に対する制約付きデコードのためのトライ オートマトン

大規模な言語モデルでは、事前定義されたスキーマに準拠する構造化された出力を生成する必要がますます高まっており、共通の制約の 1 つは有効な文字列の有限セットから選択することです。現在の制約付きデコード システムは、汎用文法コンパイルを通じてこれを処理しますが、有効な値の数が数千に増加すると法外に遅くなり、カーディナリティの壁になります。 Aho-Corasick マルチパターン マッチングを介して有限集合構造 (共有プレフィックス、制限された深さ、既知の基数) を利用してノードごとのトークン マスクを事前計算する特殊なメカニズムであるトライ オートマトンを導入します。このトライは、vLLM および SGLang の主要なバックエンドの 1 つである XGrammar と比較して、ステップごとの有効トークンの計算が 7 倍速く (0.65 us 対 5.8 us)、K >= 300 で 2 ~ 6.5 倍速いコンパイルを実現します。事前計算されたマスクにより、ガイド付きデコード パイプラインをバイパスするステートレスなサービング パスが有効になるため、この利点はバッチ サービング、つまりエンドツーエンド vLLM でさらに強化されます。スループットは、バッチ サイズ 256 (29X) で、XGrammar の 7.5 req/s に対して、219 req/s に達します。 29X は、アルゴリズムによる高速化と、事前計算されたマスクだけが実現できる統合パスの節約を組み合わせています。 7 つのトークナイザー ファミリ (32K ~ 262K 語彙) にわたって、トライは K = 10,000 までの 100ms 未満のコンパイルと、セット サイズに関係なくフラットなステップあたりのコストを維持し、同時に 100% の出力有効性を保証します。

原文 (English)

Trie Automata for Constrained Decoding over Large Finite Sets

Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-purpose grammar compilation, which becomes prohibitively slow as the number of valid values grows into the thousands, a cardinality wall. We introduce the trie automaton, a specialized mechanism that exploits finite-set structure (shared prefixes, bounded depth, known cardinality) via Aho-Corasick multi-pattern matching to precompute per-node token masks. The trie achieves 7X faster per-step valid-token computation (0.65 us vs. 5.8 us) compared to XGrammar, one of the primary backends in vLLM and SGLang, and 2--6.5X faster compilation at K >= 300. Because precomputed masks enable a stateless serving path that bypasses the guided decoding pipeline, this advantage compounds in batch serving: end-to-end vLLM throughput reaches 219 req/s vs. XGrammar's 7.5 req/s at batch size 256 (29X). The 29X combines the algorithmic speedup with integration-path savings that only precomputed masks can unlock. Across seven tokenizer families (32K--262K vocabulary), the trie maintains sub-100ms compilation up to K = 10,000 and flat per-step cost regardless of set size, while guaranteeing 100% output validity.

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

QuoteBench: スコアの一致によってコマンド パスの障害がどのように隠蔽されるか

LLM コーディング エージェントは、モデル出力をシリアル化、ラップ、再解析するインターフェイスを通じて Bash コマンドを発行します。一致した実行スコアだけでは、コマンド生成エラーと生成後に発生したエラーを区別できません。 QuoteBench は、14 のインシデント派生ファミリーからの 56 のワンショット タスクに対する正確な最終状態の検証によってこの境界を測定し、意図的にエスケープされない追加された 1 つのパーサーを中心とした実行トランスポートとの生成コントラクトを横断します。補間ポイントでのエスケープは、再生された各応答の生のパス結果を再現するため、開示された境界の下での回復は、世代を変更するモデルから行われる必要があります。 8 つの同じウィンドウ構成で、追加されたパーサーを通じて同じ応答を再生すると、成功率が 55.4 ~ 73.2 パーセント ポイント低下します。開示は 6 つの構成で 30.4 から 60.7 ポイント回復し、他の 2 つの構成ではゼロまたはわずかにマイナスになります。生の生成はフロンティアではほぼ飽和状態です。境界適応は依然としてモデルを分離するものです。 GPT-5.6-sol の -3.6 ポイントの一致ギャップにより、-64.3 ポイントのダメージと +60.7 ポイントの補正が隠蔽されます。デプロイメント構成によりモデルの順序が変更されます。26 の比較可能なペアのうちの 1 つの反転は明確で、残りの 4 つは単一タスクのマージンにあります。コマンド発行エージェントの評価では、一致したスコアをモデル固有のプロパティとして扱うのではなく、モデル構成、生成コントラクト、実行パス、操作点、および最終状態のバリデーターを報告する必要があります。

原文 (English)

QuoteBench: How Matched Scores Can Hide Command-Path Failures

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models. GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation. The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous and four more sit on single-task margins. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

言語モデルによって生成された小説は圧縮された形式的なバリエーションを示します

大規模な言語モデルは小説全体を生成できますが、その出力が何世代にもわたって形式的に変化するレベルについての情報はほとんどありません。この研究では、個々のパッセージが AI によって生成されたものとして識別できるかどうかを問うのではなく、AI 生成を繰り返すことで人間の体全体に見られるのと同じ範囲の多様性を生み出すことができるかどうかを問うています。この論文では、生成ソースとターゲット スタイルに基づいて 6 つのコーパスを対比します。19 世紀のイギリス リアリスト スタイルで GPT-5.5 Thinking を使用して生成された 20 の小説、19 世紀のイギリス リアリズム スタイルで Qwen3-14B を使用して生成された 20 の小説、現代のゼロ スタイルでこれらのモデルのそれぞれを使用して生成された 20 の小説、19 世紀に人間が書いたイギリスの小説 205 冊、および人間が書いた現代のゼロ スタイルの小説 65 です。文書レベルの研究には、MATTR-500、シャノンエントロピー、平均文長、読みやすさ、句読点率の測定が含まれます。最も堅牢で信頼性の高い結果は、文構造の圧縮です。世代を重ねることで、人間の小説に比べて文構造の相互の差異がはるかに少ない小説が生み出されます。圧縮は、小説内の読みやすさ、句読点、文の長さのばらつきの尺度にも存在します。 Qwen Zero-Style MATTR を除いて、字句メジャーも同様に圧縮される傾向があります。 GPT と Qwen には、明確な平均文体プロファイルがあるにもかかわらず、メジャー間相関の安定したパターンがありません。したがって、この記事では、小説間の限られた形式的範囲を表す分散オーバークロージャーと、相関オーバークロージャーというより具体的な現象を区別します。これは、AI が生成した個々の小説は文体的に人間の小説に似ている可能性があるが、AI が生成した小説のコレクションが占める形式的な範囲ははるかに狭いことを意味します。

原文 (English)

Novels generated by language models show compressed formal variation

While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as AI-generated, this study asks whether repeated AI generation can produce the same range of diversity which is found across human corpora. This paper contrasts six corpora based on generation source and target style: twenty novels generated using GPT-5.5 Thinking in a nineteenth-century British realist style, twenty novels generated using Qwen3-14B in a nineteenth-century British realist style, twenty novels generated using each of these models in a contemporary zero style, 205 nineteenth-century human-written British novels, and sixty-five contemporary human-written Zero-Style novels. At the document level, the research includes MATTR-500, Shannon entropy, average sentence length, readability, and punctuation rate measurements. The most robust and reliable result is compression of sentence structure. Repeated generations produce novels that vary far less from one another in sentence structure than human novels do. Compression is also present in the measures of readability, punctuation, and sentence length variability within novels. Lexical measures tend to be similarly compressed, with the exception of Qwen Zero-Style MATTR. Despite having distinct mean stylistic profiles, GPT and Qwen lack a stable pattern of cross-measure correlation. This article therefore distinguishes between variance overclosure, which represents a limited formal range between novels, and a more specific phenomenon of correlational overclosure. This means that an individual AI-generated novel may resemble human fiction stylistically, while a collection of AI-generated novels occupies a much narrower formal range.

2026-08-15 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

H-VAEP と H-xT: 確率の推定によるハンドボールの攻撃的なオンザボールアクションの評価

プロのハンドボールにおける従来の選手評価は、基本的なボックススコア指標やヒューリスティック指標に依存しており、複数選手によるビルドアップの連鎖を評価することができません。フットボール (サッカー) 分析では、予想される脅威 (xT) と確率推定によるアクションの評価 (VAEP) が採用されていますが、これらのイベントベースのアクション評価フレームワークはまだハンドボールには適応されていません。この論文では、ハンドボール ブンデスリーガの 5 シーズンの追跡由来のイベント データを利用して、ハンドボールに対する xT と VAEP の最初の包括的な適応と評価を紹介します。私たちは、ハンドボール固有のコート ゾーニング レイアウトを使用して Handball-xT (H-xT) を開発し、標準的な長方形のグリッドよりも体系的に堅牢であることをシミュレーションによって実証しています。特徴空間を調整し、コンテキストの長さを選択してチーム ID の漏洩を制限することで、ハンドボール VAEP (H-VAEP) を最適化します。私たちの評価では、H-VAEP が、ビルドアップ プレーを際立たせる、非常に安定しており、識別力があり、直感的なプレーヤー評価をもたらすことが示されています。最後に、プロのクラブがこれらのモデルを導入できるように、完全なコード リポジトリをリリースします。

原文 (English)

H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities

Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP), these event-based action valuation frameworks have not yet been adapted to handball. In this paper, we present the first comprehensive adaptation and evaluation of xT and VAEP for handball, utilizing five seasons of tracking-derived event data from the Handball Bundesliga. We develop Handball-xT (H-xT) using a handball-native court zoning layout, demonstrating via simulations that it is systematically more robust than standard rectangular grids. We optimize Handball-VAEP (H-VAEP) by tailoring its feature space and selecting the context length to limit team-identity leakage. Our evaluation shows that H-VAEP yields exceptionally stable, discriminative, and intuitive player ratings that highlight build-up play. Finally, we release our complete code repository to help professional clubs deploy these models.

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

TEMPO: Makespan 対応のエキスパート - メモリとコンピューティングに依存した体制にわたる並列負荷分散

エキスパートパラレル (EP) MoE サービスでは、すべてのレイヤーが最も遅い GPU で同期します。ディスパッチャは、エキスパート時間が線形であると仮定して、トークン カウント (EPLB、LPLB、UltraEP) またはアクティブ化されたエキスパート カウント (METRO) のバランスをとります。データセンターの 2 世代の GPU での測定では、どちらでもないことがわかりました。$\nstar\!\およそ\!156$ -- $168$ トークン未満では、HBM 重みストリーミングが支配的です -- コストはトークンではなく \emph{アクティブ化されたレプリカ} に発生します。その上で、グループ化された GEMM はトークンを 128 タイルの $M$ タイルに丸めるため、専門家が \emph{分割} してパディングされた計算を追加します。最大アフィン プロファイル $t=\max(a+bG,\,c+\beta N)$ は両方の領域を捉えます。現実的なデコードバッチには、線形領域ではホットな専門家が、フラット領域では冷たい専門家が \emph{同時に}保持されます。記録されたバッチは、プロキシのディスパッチがモデル化されたブロック時間で $1.4$--$1.6\times$ 異なり (p95 から $1.7\times$)、\emph{どの} プロキシが勝つかはレジームによって反転することを示しています。私たちは、バッチごとのディスパッチを固定料金のメイクスパン問題として形式化し、2 つの完全に複製された GPU 上で NP ハード、縮退限界の多項式を実現します。そして、クリティカル パスからミリ秒でそれを解決するメイクスパン対応ディスパッチャである \sys{} を提示します。 SGLang 統合はプロセス外で実行され、ディスパッチとカウント収集を 1 つのグラフ内カーネルに融合します。 8 GPU のテストベッド (マイクロベンチマーク) によって支えられている \sys{} は、どこでも最良の固定ベースラインの 1\% 以内に留まり、体制が混在する場合には最大 $15.5\%$ の差で勝利します。 Testbed~B のエンドツーエンドでは、Qwen3-235B (勝利領域内) のスループットが $4$--$6\%$ 向上し、p99 レイテンシーが ${\sim}15.6\%$ 削減されました。 DeepSeek-V3 (外部、通信主体) はメカニズムのコストのみを示します。フェーズ図は普遍的な勝利をもたらすものではなく、展開前に両方の結果を予測するという主張です。

原文 (English)

TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes

In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU generations show it is neither: below $\nstar\!\approx\!156$--$168$ tokens, HBM weight streaming dominates---cost attaches to \emph{activated replicas}, not tokens; above it, grouped GEMM rounds tokens to 128-tile $M$-tiles, so \emph{splitting} an expert adds padded compute. A max-affine profile $t=\max(a+bG,\,c+\beta N)$ captures both regimes. Realistic decode batches hold hot experts in the linear regime and cold in the flat \emph{simultaneously}; recorded batches show proxy dispatches differ by $1.4$--$1.6\times$ in modeled block time (p95 up to $1.7\times$), and \emph{which} proxy wins flips with the regime. We formalize per-batch dispatch as a fixed-charge makespan problem---NP-hard on two fully replicated GPUs, polynomial in degenerate limits---and present \sys{}, a makespan-aware dispatcher solving it in milliseconds off the critical path; its SGLang integration runs out-of-process and fuses dispatch with count collection into one in-graph kernel. Anchored by an 8-GPU Testbed~A microbenchmark, \sys{} stays within 1\% of the best fixed baseline everywhere and wins by up to $15.5\%$ where regimes mix. End-to-end on Testbed~B, Qwen3-235B (inside the win region) gains $4$--$6\%$ throughput and cuts p99 latency by ${\sim}15.6\%$; DeepSeek-V3 (outside, communication-dominated) shows only mechanism cost. A phase diagram, not a universal win, is the claim: it predicts both outcomes before deployment.

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成エージェントハードウェア/半導体

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic p…

2026-08-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Unmasking Conversational Bias in AI Multiagent Systems

Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application…

2026-08-15 13:00 JSTarXiv cs.AI画像/動画生成ロボティクスハードウェア/半導体

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gain…

2026-08-14 23:50 JSTTechCrunch AIエージェントハードウェア/半導体

Kog is going deeper to squeeze more inference out of GPUs

The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.

2026-08-14 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

立場: アライメント コミュニティは意図せずに検閲ツールキットを構築している

この意見書では、現代の AI 調整手法は、本来は有害な出力を防ぐために設計されたものであり、悪意のある行為者によって検閲や操作のために簡単に悪用される可能性がある二重用途技術であると主張しています。現在の調整技術を悪用の可能性と実際のケースにマッピングすることで、「完全に調整された」モデルの探求が、意図せずして悪意のある攻撃者に情報支配のための絶えず改良されたツールを提供してしまうことを示します。ユーザーによる情報プロバイダーとしての AI の急速な導入、経済力の非対称性、権威主義への移行が進む政治情勢によってそのリスクが悪化しているため、私たちはこの二重利用の可能性について今すぐ議論する必要があります。私たちはコミュニティに対し、AI 調整メカニズムの意図的な誤用を考慮し、この二重使用の可能性を防ぐための緩和戦略を提案することを強く求めて締めくくります。

原文 (English)

Position: The Alignment Community is Unintentionally Building a Censor's Toolkit

This position paper argues that modern AI alignment methods - originally designed to prevent harmful output - are dual-use technologies that may easily be misused by malicious actors for censorship and manipulation. By mapping current alignment techniques to the possibility and actual cases of misuse, we show that the quest for a "perfectly aligned" model inadvertently also provides malicious actors with an ever-improving tool for informational dominance. We need to discuss this dual-use potential now, as its risk is exacerbated by rapid user adoption of AI as information provider, economic power asymmetries, and a political landscape that increasingly shifts towards authoritarianism. We conclude by urging the community to consider the intentional misuse of AI alignment mechanisms and propose mitigation strategies to safeguard against this dual-use potential.

2026-08-14 13:00 JSTarXiv cs.AIハードウェア/半導体

CAS: ローカルおよびグローバルの説明可能な人工知能の因果関係スコア

予測説明メソッドはモデルの出力に帰属します。それら自体は、介入効果が現実世界の結果に影響を与えるものではありません。因果関係を説明するためのコンパクトなスコア アーキテクチャである Causal Attribution Score (CAS) を紹介します。 CAS は、特定された介入連合ゲームから開始し、共同介入コントラストを因果関係のある Shapley 寄与と割り当て、それらの生の結果スケール効果をローカル CAS、署名済みローカル CAS、および 2 つの補完的なグローバル CAS 要約に変換します。このイノベーションは、新しい Shapley 式ではなく、明示的な介入ターゲットを備えたローカルからグローバルへの因果関係レポート層です。既知の真実のベンチマークでは、8 回繰り返された一次相互作用シミュレーション (それぞれ n = 2,200、3 つのアクション) で、連合認識 CAS の平均ローカル CAS MAE が 0.107 であったのに対し、一度に 1 つずつ正規化した場合は 0.173、グローバル正規化された絶対 ATE ベクトルでは 0.213 でした。一度に 1 つずつ正規化を行う場合のペアの利点は、相加性の場合の -0.003 から、強い相互作用の場合の 0.091 に増加しました。 DoubleML の経験的データセット、401(k) 資格/純金融資産 (n = 9,915) およびペンシルベニア州の再雇用ボーナス/失業期間 (n = 5,099) の両方において、予測 SHAP/TreeSHAP ランキングは、治療効果修飾因子の Feature-CAS ランキングとは大きく異なりました。ペンシルベニア州では、dep1 (正確に 1 つの依存関係) が予測グローバル ランク 13 から Feature-CAS ランク 2 に移動し、主要なローカル Feature-CAS 修飾子となりました。これらの結果は、結果を予測するものと、推定された因果効果の不均一性を説明するものを分離するという付加価値を分離します。

原文 (English)

CAS: A Causal Attribution Score for Local and Global Explainable Artificial Intelligence

Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome. We introduce the Causal Attribution Score (CAS), a compact score architecture for causal explanation. CAS starts from an identified interventional coalition game, allocates the joint intervention contrast with causal Shapley contributions, and converts those raw outcome-scale effects into Local CAS, Signed Local CAS, and two complementary Global CAS summaries. The innovation is not a new Shapley formula, but a local-to-global causal reporting layer with an explicit intervention target. In the known-truth benchmark, eight repeated primary-interaction simulations (n = 2,200 each, three actions) gave mean Local CAS MAE of 0.107 for coalition-aware CAS, compared with 0.173 for one-at-a-time normalisation and 0.213 for a global normalised absolute ATE vector. The paired advantage over one-at-a-time normalisation increased from -0.003 under additivity to 0.091 under strong interactions. On both empirical DoubleML datasets, 401(k) eligibility/net financial assets (n = 9,915) and Pennsylvania reemployment bonus/unemployment duration (n = 5,099), predictive SHAP/TreeSHAP rankings differed materially from Feature-CAS rankings of treatment-effect modifiers. In Pennsylvania, dep1 (exactly one dependent) moved from predictive global rank 13 to Feature-CAS rank 2 and was the leading local Feature-CAS modifier. These results isolate the added value of separating what predicts the outcome from what explains heterogeneity in an estimated causal effect.

2026-08-14 13:00 JSTarXiv cs.AIハードウェア/半導体

大規模な有限集合に対する制約付きデコードのためのトライ オートマトン

大規模な言語モデルでは、事前定義されたスキーマに準拠する構造化された出力を生成する必要がますます高まっており、共通の制約の 1 つは有効な文字列の有限セットから選択することです。現在の制約付きデコード システムは、汎用文法コンパイルを通じてこれを処理しますが、有効な値の数が数千に増加すると法外に遅くなり、カーディナリティの壁になります。 Aho-Corasick マルチパターン マッチングを介して有限集合構造 (共有プレフィックス、制限された深さ、既知の基数) を利用してノードごとのトークン マスクを事前計算する特殊なメカニズムであるトライ オートマトンを導入します。このトライは、vLLM および SGLang の主要なバックエンドの 1 つである XGrammar と比較して、ステップごとの有効トークンの計算が 7 倍速く (0.65 us 対 5.8 us)、K >= 300 で 2 ~ 6.5 倍速いコンパイルを実現します。事前計算されたマスクにより、ガイド付きデコード パイプラインをバイパスするステートレスなサービング パスが有効になるため、この利点はバッチ サービング、つまりエンドツーエンド vLLM でさらに強化されます。スループットは、バッチ サイズ 256 (29X) で、XGrammar の 7.5 req/s に対して、219 req/s に達します。 29X は、アルゴリズムによる高速化と、事前計算されたマスクだけが実現できる統合パスの節約を組み合わせています。 7 つのトークナイザー ファミリ (32K ~ 262K 語彙) にわたって、トライは K = 10,000 までの 100ms 未満のコンパイルと、セット サイズに関係なくフラットなステップあたりのコストを維持し、同時に 100% の出力有効性を保証します。

原文 (English)

Trie Automata for Constrained Decoding over Large Finite Sets

Large language models increasingly need to generate structured outputs that conform to predefined schemas, with one common constraint being selection from a finite set of valid strings. Current constrained decoding systems handle this through general-purpose grammar compilation, which becomes prohibitively slow as the number of valid values grows into the thousands, a cardinality wall. We introduce the trie automaton, a specialized mechanism that exploits finite-set structure (shared prefixes, bounded depth, known cardinality) via Aho-Corasick multi-pattern matching to precompute per-node token masks. The trie achieves 7X faster per-step valid-token computation (0.65 us vs. 5.8 us) compared to XGrammar, one of the primary backends in vLLM and SGLang, and 2--6.5X faster compilation at K >= 300. Because precomputed masks enable a stateless serving path that bypasses the guided decoding pipeline, this advantage compounds in batch serving: end-to-end vLLM throughput reaches 219 req/s vs. XGrammar's 7.5 req/s at batch size 256 (29X). The 29X combines the algorithmic speedup with integration-path savings that only precomputed masks can unlock. Across seven tokenizer families (32K--262K vocabulary), the trie maintains sub-100ms compilation up to K = 10,000 and flat per-step cost regardless of set size, while guaranteeing 100% output validity.

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

QuoteBench: スコアの一致によってコマンド パスの障害がどのように隠蔽されるか

LLM コーディング エージェントは、モデル出力をシリアル化、ラップ、再解析するインターフェイスを通じて Bash コマンドを発行します。一致した実行スコアだけでは、コマンド生成エラーと生成後に発生したエラーを区別できません。 QuoteBench は、14 のインシデント派生ファミリーからの 56 のワンショット タスクに対する正確な最終状態の検証によってこの境界を測定し、意図的にエスケープされない追加された 1 つのパーサーを中心とした実行トランスポートとの生成コントラクトを横断します。補間ポイントでのエスケープは、再生された各応答の生のパス結果を再現するため、開示された境界の下での回復は、世代を変更するモデルから行われる必要があります。 8 つの同じウィンドウ構成で、追加されたパーサーを通じて同じ応答を再生すると、成功率が 55.4 ~ 73.2 パーセント ポイント低下します。開示は 6 つの構成で 30.4 から 60.7 ポイント回復し、他の 2 つの構成ではゼロまたはわずかにマイナスになります。生の生成はフロンティアではほぼ飽和状態です。境界適応は依然としてモデルを分離するものです。 GPT-5.6-sol の -3.6 ポイントの一致ギャップにより、-64.3 ポイントのダメージと +60.7 ポイントの補正が隠蔽されます。デプロイメント構成によりモデルの順序が変更されます。26 の比較可能なペアのうちの 1 つの反転は明確で、残りの 4 つは単一タスクのマージンにあります。コマンド発行エージェントの評価では、一致したスコアをモデル固有のプロパティとして扱うのではなく、モデル構成、生成コントラクト、実行パス、操作点、および最終状態のバリデーターを報告する必要があります。

原文 (English)

QuoteBench: How Matched Scores Can Hide Command-Path Failures

LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures introduced after generation. QuoteBench measures this boundary with exact final-state validation on 56 one-shot tasks from 14 incident-derived families, crossing the generation contract with the execution transport around one deliberately unescaped added parser. Escaping at the interpolation point reproduces each replayed reply's raw-path outcome, so any recovery under a disclosed boundary must come from the model changing its generation. Across eight same-window configurations, replaying the same reply through the added parser lowers success by 55.4 to 73.2 percentage points; disclosure recovers 30.4 to 60.7 points for six configurations, and zero or slightly negative for the other two. Raw generation is nearly saturated at the frontier; boundary adaptation is what still separates models. GPT-5.6-sol's matched gap of -3.6 points hides -64.3 points of damage and +60.7 points of compensation. The deployment configuration reorders models: one reversal among 26 comparable pairs is unambiguous and four more sit on single-task margins. Evaluations of command-issuing agents should report the model configuration, generation contract, execution path, operating point, and final-state validator rather than treat a matched score as an intrinsic model property.

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

言語モデルによって生成された小説は圧縮された形式的なバリエーションを示します

大規模な言語モデルは小説全体を生成できますが、その出力が何世代にもわたって形式的に変化するレベルについての情報はほとんどありません。この研究では、個々のパッセージが AI によって生成されたものとして識別できるかどうかを問うのではなく、AI 生成を繰り返すことで人間の体全体に見られるのと同じ範囲の多様性を生み出すことができるかどうかを問うています。この論文では、生成ソースとターゲット スタイルに基づいて 6 つのコーパスを対比します。19 世紀のイギリス リアリスト スタイルで GPT-5.5 Thinking を使用して生成された 20 の小説、19 世紀のイギリス リアリズム スタイルで Qwen3-14B を使用して生成された 20 の小説、現代のゼロ スタイルでこれらのモデルのそれぞれを使用して生成された 20 の小説、19 世紀に人間が書いたイギリスの小説 205 冊、および人間が書いた現代のゼロ スタイルの小説 65 です。文書レベルの研究には、MATTR-500、シャノンエントロピー、平均文長、読みやすさ、句読点率の測定が含まれます。最も堅牢で信頼性の高い結果は、文構造の圧縮です。世代を重ねることで、人間の小説に比べて文構造の相互の差異がはるかに少ない小説が生み出されます。圧縮は、小説内の読みやすさ、句読点、文の長さのばらつきの尺度にも存在します。 Qwen Zero-Style MATTR を除いて、字句メジャーも同様に圧縮される傾向があります。 GPT と Qwen には、明確な平均文体プロファイルがあるにもかかわらず、メジャー間相関の安定したパターンがありません。したがって、この記事では、小説間の限られた形式的範囲を表す分散オーバークロージャーと、相関オーバークロージャーというより具体的な現象を区別します。これは、AI が生成した個々の小説は文体的に人間の小説に似ている可能性があるが、AI が生成した小説のコレクションが占める形式的な範囲ははるかに狭いことを意味します。

原文 (English)

Novels generated by language models show compressed formal variation

While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as AI-generated, this study asks whether repeated AI generation can produce the same range of diversity which is found across human corpora. This paper contrasts six corpora based on generation source and target style: twenty novels generated using GPT-5.5 Thinking in a nineteenth-century British realist style, twenty novels generated using Qwen3-14B in a nineteenth-century British realist style, twenty novels generated using each of these models in a contemporary zero style, 205 nineteenth-century human-written British novels, and sixty-five contemporary human-written Zero-Style novels. At the document level, the research includes MATTR-500, Shannon entropy, average sentence length, readability, and punctuation rate measurements. The most robust and reliable result is compression of sentence structure. Repeated generations produce novels that vary far less from one another in sentence structure than human novels do. Compression is also present in the measures of readability, punctuation, and sentence length variability within novels. Lexical measures tend to be similarly compressed, with the exception of Qwen Zero-Style MATTR. Despite having distinct mean stylistic profiles, GPT and Qwen lack a stable pattern of cross-measure correlation. This article therefore distinguishes between variance overclosure, which represents a limited formal range between novels, and a more specific phenomenon of correlational overclosure. This means that an individual AI-generated novel may resemble human fiction stylistically, while a collection of AI-generated novels occupies a much narrower formal range.

2026-08-14 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

H-VAEP と H-xT: 確率の推定によるハンドボールの攻撃的なオンザボールアクションの評価

プロのハンドボールにおける従来の選手評価は、基本的なボックススコア指標やヒューリスティック指標に依存しており、複数選手によるビルドアップの連鎖を評価することができません。フットボール (サッカー) 分析では、予想される脅威 (xT) と確率推定によるアクションの評価 (VAEP) が採用されていますが、これらのイベントベースのアクション評価フレームワークはまだハンドボールには適応されていません。この論文では、ハンドボール ブンデスリーガの 5 シーズンの追跡由来のイベント データを利用して、ハンドボールに対する xT と VAEP の最初の包括的な適応と評価を紹介します。私たちは、ハンドボール固有のコート ゾーニング レイアウトを使用して Handball-xT (H-xT) を開発し、標準的な長方形のグリッドよりも体系的に堅牢であることをシミュレーションによって実証しています。特徴空間を調整し、コンテキストの長さを選択してチーム ID の漏洩を制限することで、ハンドボール VAEP (H-VAEP) を最適化します。私たちの評価では、H-VAEP が、ビルドアップ プレーを際立たせる、非常に安定しており、識別力があり、直感的なプレーヤー評価をもたらすことが示されています。最後に、プロのクラブがこれらのモデルを導入できるように、完全なコード リポジトリをリリースします。

原文 (English)

H-VAEP and H-xT: Valuing Offensive On-the-Ball Actions in Handball by Estimating Probabilities

Traditional player evaluation in professional handball relies on basic box-score metrics or heuristic indices, which fail to credit the multi-player build-up chain. While football (soccer) analytics has adopted Expected Threat (xT) and Valuing Actions by Estimating Probabilities (VAEP), these event-based action valuation frameworks have not yet been adapted to handball. In this paper, we present the first comprehensive adaptation and evaluation of xT and VAEP for handball, utilizing five seasons of tracking-derived event data from the Handball Bundesliga. We develop Handball-xT (H-xT) using a handball-native court zoning layout, demonstrating via simulations that it is systematically more robust than standard rectangular grids. We optimize Handball-VAEP (H-VAEP) by tailoring its feature space and selecting the context length to limit team-identity leakage. Our evaluation shows that H-VAEP yields exceptionally stable, discriminative, and intuitive player ratings that highlight build-up play. Finally, we release our complete code repository to help professional clubs deploy these models.

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

TEMPO: Makespan 対応のエキスパート - メモリとコンピューティングに依存した体制にわたる並列負荷分散

エキスパートパラレル (EP) MoE サービスでは、すべてのレイヤーが最も遅い GPU で同期します。ディスパッチャは、エキスパート時間が線形であると仮定して、トークン カウント (EPLB、LPLB、UltraEP) またはアクティブ化されたエキスパート カウント (METRO) のバランスをとります。データセンターの 2 世代の GPU での測定では、どちらでもないことがわかりました。$\nstar\!\およそ\!156$ -- $168$ トークン未満では、HBM 重みストリーミングが支配的です -- コストはトークンではなく \emph{アクティブ化されたレプリカ} に発生します。その上で、グループ化された GEMM はトークンを 128 タイルの $M$ タイルに丸めるため、専門家が \emph{分割} してパディングされた計算を追加します。最大アフィン プロファイル $t=\max(a+bG,\,c+\beta N)$ は両方の領域を捉えます。現実的なデコードバッチには、線形領域ではホットな専門家が、フラット領域では冷たい専門家が \emph{同時に}保持されます。記録されたバッチは、プロキシのディスパッチがモデル化されたブロック時間で $1.4$--$1.6\times$ 異なり (p95 から $1.7\times$)、\emph{どの} プロキシが勝つかはレジームによって反転することを示しています。私たちは、バッチごとのディスパッチを固定料金のメイクスパン問題として形式化し、2 つの完全に複製された GPU 上で NP ハード、縮退限界の多項式を実現します。そして、クリティカル パスからミリ秒でそれを解決するメイクスパン対応ディスパッチャである \sys{} を提示します。 SGLang 統合はプロセス外で実行され、ディスパッチとカウント収集を 1 つのグラフ内カーネルに融合します。 8 GPU のテストベッド (マイクロベンチマーク) によって支えられている \sys{} は、どこでも最良の固定ベースラインの 1\% 以内に留まり、体制が混在する場合には最大 $15.5\%$ の差で勝利します。 Testbed~B のエンドツーエンドでは、Qwen3-235B (勝利領域内) のスループットが $4$--$6\%$ 向上し、p99 レイテンシーが ${\sim}15.6\%$ 削減されました。 DeepSeek-V3 (外部、通信主体) はメカニズムのコストのみを示します。フェーズ図は普遍的な勝利をもたらすものではなく、展開前に両方の結果を予測するという主張です。

原文 (English)

TEMPO: Makespan-Aware Expert-Parallel Load Balancing Across Memory- and Compute-Bound Regimes

In expert-parallel (EP) MoE serving, every layer synchronizes at the slowest GPU. Dispatchers balance token counts (EPLB, LPLB, UltraEP) or activated-expert counts (METRO), assuming expert time is linear in one. Measurements on two datacenter GPU generations show it is neither: below $\nstar\!\approx\!156$--$168$ tokens, HBM weight streaming dominates---cost attaches to \emph{activated replicas}, not tokens; above it, grouped GEMM rounds tokens to 128-tile $M$-tiles, so \emph{splitting} an expert adds padded compute. A max-affine profile $t=\max(a+bG,\,c+\beta N)$ captures both regimes. Realistic decode batches hold hot experts in the linear regime and cold in the flat \emph{simultaneously}; recorded batches show proxy dispatches differ by $1.4$--$1.6\times$ in modeled block time (p95 up to $1.7\times$), and \emph{which} proxy wins flips with the regime. We formalize per-batch dispatch as a fixed-charge makespan problem---NP-hard on two fully replicated GPUs, polynomial in degenerate limits---and present \sys{}, a makespan-aware dispatcher solving it in milliseconds off the critical path; its SGLang integration runs out-of-process and fuses dispatch with count collection into one in-graph kernel. Anchored by an 8-GPU Testbed~A microbenchmark, \sys{} stays within 1\% of the best fixed baseline everywhere and wins by up to $15.5\%$ where regimes mix. End-to-end on Testbed~B, Qwen3-235B (inside the win region) gains $4$--$6\%$ throughput and cuts p99 latency by ${\sim}15.6\%$; DeepSeek-V3 (outside, communication-dominated) shows only mechanism cost. A phase diagram, not a universal win, is the claim: it predicts both outcomes before deployment.

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成エージェントハードウェア/半導体

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic p…

2026-08-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Unmasking Conversational Bias in AI Multiagent Systems

Detecting biases in the outputs produced by generative models is essential to reduce the potential risks associated with their application…

2026-08-14 13:00 JSTarXiv cs.AI画像/動画生成ロボティクスハードウェア/半導体

SONIC: Supersizing Motion Tracking for Natural Humanoid Whole-Body Control

Despite the rise of billion-parameter foundation models trained across thousands of graphical processing units (GPUs), similar scaling gain…

2026-08-14 00:08 JSTTechCrunch AIハードウェア/半導体

Nvidia’s new $500B plan is risky but brilliant, especially for aging GPUs

Nvidia has a plan to make sure its GPUs won't lose value. It wants to convince a new crop of financiers to keep lending for AI buildouts.

2026-08-13 19:00 JSTOpenAILLM/生成AIハードウェア/半導体

Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed

Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster. Powered by Cerebras, it delivers up to 750 output…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

強化学習による AI データセンターのエネルギー削減: 1 つの GPU からフリートまでの LLM トレーニングの測定された電力制御

トレーニング後の強化学習は現代の言語モデル開発の主流を占めていますが、GPU ハードウェアでのその電力動作は特徴づけられておらず、データセンターは、ハードウェアを無差別に遅くするワークロード ブラインド メカニズム、静的キャップ、リアクティブ スロットルを使用して GPU 電力を管理しています。 1 ~ 4 台の A100 (380,000 以上のサンプル) で 7B、14B、および 72B スケールの 0.5 秒の電力テレメトリを使用して GRPO トレーニングを計測し、ワークロード自身の生成パラメータを測定された電力に適応させる PPO メタコントローラーをトレーニングします。完全な 500 ステップの 7B トレースに対して、コントローラーは電力制限違反を 89.8% 削減し、同時にトークン出力を 18.1%、エネルギー効率を 26.2% (MWh あたりのトークン) 増加させます。 72B でライブ デプロイすると、同じコントローラー ファミリが複製されたヌル結果を生成し、グループサイズのアクチュエーターがモデル シャーディングの下で​​権限を失ったと診断されます。アクチュエータ権限スイープでは、同時生成として適用される同じパラメータが 17 ~ 22% の電力権限を維持し、占有対体積の原則を分離していることがわかります。そのアクチュエータ上に再構築されたコントローラは、3 つのレプリケーションにわたってライブ 72B ロールアウト生成ワークロードを制御します。つまり、2.27 +/- 1.08% の予算違反で静的安全ベースラインよりも出力が 35.7% 増加し、制御されていない動作よりも違反が 87.2% 減少し、制約のあるコントローラの中で最良の平均スループットとトークンあたりのエネルギーが得られ、適応しきい値ルールが 3 つの動作条件のいずれかで一致します。現実的な測定ウィンドウでは、元の 72B 過渡電流は 0.5 秒の分解能で 23.6% から 30 秒で 1.6% に低下し、5 分でゼロになります。構成された 16 GPU フリートでは、30 秒以上では違反がゼロであり、ピーク需要はネームプレートの 50 ~ 56% です。この車両構成では、オペレーターの検証を条件として、ネームプレートの約 2 倍のオーバーサブスクリプションが実現可能であると思われます。私たちは経済と炭素への影響を定量化し、低コストのオペレーターのパイロットを特定します。

原文 (English)

Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet

Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been characterized, and datacenters manage GPU power with workload-blind mechanisms, static caps and reactive throttling, that slow hardware indiscriminately. We instrument GRPO training with half-second power telemetry at 7B, 14B, and 72B scales on one to four A100s (380,000+ samples), and train a PPO meta-controller that adapts the workload's own generation parameters to measured power. Against the full 500-step 7B trace, the controller cuts power-limit violations by 89.8% while increasing token output by 18.1% and energy efficiency by 26.2% (tokens per MWh). Deployed live at 72B, the same controller family yields replicated null results, diagnosed as the group-size actuator losing authority under model sharding. An actuator-authority sweep shows the same parameters applied as generation concurrency retain 17-22% power authority, isolating an occupancy-versus-volume principle; a controller rebuilt on that actuator controls a live 72B rollout-generation workload across three replications: 35.7% more output than a static safe baseline at 2.27 +/- 1.08% budget violations, 87.2% fewer violations than uncontrolled operation, and the best mean throughput and energy per token among constrained controllers, with an adaptive threshold rule matching it in one of three operating conditions. Under realistic measurement windows the original 72B transients fall from 23.6% at half-second resolution to 1.6% at 30 s and zero at 5 min; a composed 16-GPU fleet shows zero violations at 30 s and longer, with peak demand at 50-56% of nameplate. For this fleet mix, roughly twofold oversubscription of nameplate appears feasible, subject to operator validation. We quantify the economic and carbon consequences and specify a low-cost operator pilot.

2026-08-13 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

社会性: 人間と AI の相互作用のための関係プロセス フレームワーク

人間と AI の研究では、個々の能力、総合的なパフォーマンス、または最終的な成果物を評価することがよくありますが、これらのアプローチでは、一方の反応がどのようにして他方の次の貢献が形成される条件の一部になるかが保存されません。この記事では、社会的双対性、つまり、区別可能な 2 つの当事者の間の連続的で互恵的で歴史を担う関係プロセスについて紹介します。この関係プロセスでは、一方の当事者からの反応が、他方の当事者のその後の貢献、判断、決定、または行動が形成される観察可能な条件の一部になります。人間と AI のダイアッド向けに指定されたこの構成では、移動、確認された社会的エピソード、リンクされた経路、およびより広範な相互作用コンテナーなどの入れ子になったユニットが使用されます。最小限のエピソード A1-B1-A2 では、対応偶発性と復帰偶発性の証拠が必要です。候補エピソードは、反応の方向性と実質的な貢献の再形成を二次的にコーディングする前に、確認済み、非社会的、または不確定として分類されます。 3 つの命題は、履歴条件の形成、経路の分岐、エンドポイントに相当する経路間の堅牢性の違いに対処します。凍結された運用プロトコルは、別々に実行された 2 つのモデルベースの評価シリーズを通じて、これまで見たことのない 3 つの自然な人間の AI 記録に基づいて調整されました。移動と候補の再構成は 2 つのケースで正確に収束しましたが、3 番目のケースではローカルのマルチモーダル単位化の決定が 1 つ異なりました。残りの意見の相違は、帰還と不測の事態の境界に集中していた。したがって、社会二重性は、エンドポイント中心の分析では回復できない経路情報を保存しながら、人間と AI の貢献が相互作用を通じてどのように形成されるかを分析するための、限定的で経験的に扱いやすいプロセス構造を提供します。

原文 (English)

Socioduality: A Relational Process Framework for Human-AI Interaction

Human-AI research often evaluates individual capabilities, combined performance, or final outputs, but these approaches do not preserve how one party's response becomes part of the conditions under which the other party's next contribution is formed. This article introduces socioduality, a sequential, reciprocal, and history-carrying relational process between two distinguishable parties in which a response from one party becomes part of the observable conditions under which the other party's subsequent contribution, judgement, decision, or action is formed. Specified for human-AI dyads, the construct uses nested units: moves, confirmed sociodual episodes, linked pathways, and the broader interaction container. A minimum episode A1-B1-A2 requires evidence of response contingency and return contingency; candidate episodes are classified as confirmed, non-sociodual, or indeterminate before secondary coding of response orientation and substantive contribution re-formation. Three propositions address history-conditioned formation, pathway divergence, and robustness differences among endpoint-equivalent pathways. A frozen operational protocol was calibrated on three previously unseen natural human-AI records through two separately executed model-based evaluator series. Move and candidate reconstruction converged exactly in two cases and differed by one local multimodal unitisation decision in the third; remaining disagreement was concentrated at return-contingency boundaries. Socioduality therefore provides a bounded and empirically tractable process construct for analysing how human and AI contributions are formed through interaction while preserving pathway information that endpoint-centred analysis cannot recover.

2026-08-13 13:00 JSTarXiv cs.AIエージェントロボティクスハードウェア/半導体

Herding End-to-End Autonomous Driving via Neuro-Symbolic Safety Guards

Modern end-to-end driving agents can achieve high average performance yet still violate basic traffic rules that a human driver would never…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached t…

2026-08-13 13:00 JSTarXiv cs.AIハードウェア/半導体

Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads

A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings.…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Locating and Controlling Implicit Personalization in Large Language Models

Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic i…

2026-08-13 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

RealisticTritonBench: A Benchmark for Triton-Kernel Generation in Real-World AI Frameworks

In modern AI frameworks, GPU kernels are key to overall system performance. Combining usability, portability, and near-handwritten CUDA per…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control

LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emi…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体研究/論文

DREAMS: Density Functional Theory Based Research Engine for Agentic Materials Simulation

Large language model (LLM) agents can execute long-horizon scientific workflows, but their numerical outputs are difficult to trust: agents…

2026-08-13 13:00 JSTarXiv cs.AIハードウェア/半導体

フローバイフロー:高損失ドメインでの AI 出力を制御するためのコンテンツ判定バイパス

これまでの研究では、AI の出力速度 V が人間の認知能力 C_max を超えると、高損失領域では人間による監視が構造的に不可能になることが示されています。ただし、操作上の制約は V 単独ではなく、V x L です。ここで、L は項目ごとの認知負荷を示します。 L はトリアージ、判断、対応で構成され、AI の能力向上に対して非対称に対応します。セマンティックな不確定性は汎用設計に内在するため、モデルの機能が向上してもトリアージ コストは減少しません。応答コストは精度の向上に影響されません。判断コストのみが下方圧力に直面しており、この圧力は多くの場合、真の削減ではなく省略を誘発する形で作用します。したがって、能力の向上は L を削減するのではなく、再構築することになります。 AI の出力が正しいかどうかの評価に基づくガバナンス メカニズムは、その評価を AI に委任して幻覚リスクを継承するか、それを人間に委任して V x L の上限に直面するかのいずれかになります。私たちは、コンテンツを評価せずに監視負荷を制御するガバナンス パラダイムである Flow-by-Flow を提案します。正式な可算特徴に基づく認知コスト スコアは、大量生産に非線形コストを課す一方、制度上のキャパシティ キャップにより処理量が C_max 以内に保たれます。コンテンツ判定バイパス超過経路に対する 4 つの設計不変条件を導き出します。それは、コンテンツ判定なし、検査官能力のスケーラブルな消費なし、ID に縛られたアプリケーションごとの摩擦、およびバッチクリアランスなしです。 1 つの参考実装については、これらの不変式が同時に満たされることを示すために議論されていますが、その実際的な困難は明示的に認識されています。 1,000 のパラメータ描画にわたるモンテカルロ分析の例では、複合マルチメトリック フロー制御が試験の 90.8% で監視強化単独よりも優れていることが示唆されています。

原文 (English)

Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains

Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains once AI output velocity V exceeds human cognitive capacity C_max. The operative constraint, however, is V x L, where L is per-item cognitive load: triage, judgment, and response. These components respond asymmetrically to capability improvement. Triage cost does not decline, because semantic indeterminacy is inherent in general-purpose design. Response cost is invariant to accuracy. Only judgment cost faces downward pressure, largely by inducing omission. Capability improvement therefore restructures L rather than reducing it. We prove a proposition: if V x L grows at any positive compound rate while supervisory capacity grows linearly, exceedance occurs in finite time; capacity investment buys time only logarithmically, while reducing the growth rate extends it hyperbolically. Supervision enhancement and flow control are therefore not remedies of the same kind. We propose Flow-by-Flow, a governance design that prices supervisory load without evaluating content, intent, or legitimacy. A cognitive cost score built from formal, countable features imposes compounding costs on volume expansion, and an institutional capacity cap fixes processing within C_max. Four design invariants characterize any admissible exceedance pathway: no content judgment, no scalable consumption of examiner capacity, identity-bound per-application friction, and no batch clearance. Excess claim and page fees in patent systems are precursors satisfying only the first two invariants. One reference implementation satisfying all four is presented. A Monte Carlo analysis across 1,000 parameter draws confirms that the analytically derived ordering survives the 30-year horizon in 90.8% of trials.

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

DORA Explorer: Improving the Exploration Ability of LLMs Without Training

Large language model (LLM) agents for sequential decision-making struggle to produce diverse outputs. This leads to insufficient exploratio…

2026-08-13 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

ルーブリックベースの強化学習における報酬ハッキングの再現、分析、検出

ルーブリックベースの強化学習 (RL) は、LLM-as-a-Judge (LaaJ) を使用して、報酬としてルーブリックに従ってモデルの出力を採点します。ただし、政策モデルは裁判官の潜在的なバイアスを悪用し、報酬のハッキングや非効果的または危険なトレーニング結果につながる可能性があります。現実のルーブリックベースの RL では、このようなハッキング行為は多くの場合微妙であり、複数の裁判官のバイアスと絡み合っているため、分析、検出、軽減することが困難です。このペーパーでは、ルーブリックベースの RL のための制御可能なハッキング環境である CHERRL を紹介します。既知のバイアスを LaaJ に注入することで、CHERRL は報酬ハッキングの安定した再現、報酬の発散の明確な観察、およびハッキングの開始の正確な特定を可能にします。これは、ルーブリック ベースの RL における報酬ハッキングのメカニズムと緩和を研究するためのクリーンな実験テストベッドを提供します。その有用性を実証するために、発見可能性と悪用可能性の観点からさまざまな裁判官のバイアスを分析し、トレーニングログから報酬ハッキングの開始を自動的に検出するためのエージェントベースのシステムを調査します。コードと環境は https://github.com/THUAIS-Lab/CHERRL で公開されています。

原文 (English)

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a Controllable Hacking Environment for Rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, and explore an agent for automatically detecting reward hacking onset from training logs. The code and environment are publicly available at https://github.com/THUAIS-Lab/CHERRL.

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

言語モデルの隠れた状態を大規模に解釈する

レンズ メソッドは、中間アクティベーションを出力語彙にマッピングすることによって大規模言語モデル (LLM) を解釈し、ネットワークを通じて次のトークンの予測がどのように展開されるかを明らかにします。トレーニング済みのレンズは依然として高価です。アフィン変換パラメータはモデル幅に応じて二次関数的に増加しますが、正確で完全な語彙のカルバック ライブラー (KL) トレーニングが記憶を支配します。その結果、事前に訓練されたレンズは最大 20B パラメータのモデルに適用され、特定のコンポーネント タイプに関連付けられたままになります。 OmniLens は、残留ストリーム、アテンション、MLP のいずれであっても、単一のレンズ ファミリをモデル幅のアクティベーションに適用し、2 つの独立したスケーリング技術を組み合わせたものです。まず、低ランクのトランスレーターは、レンズごとのパラメーターの増加をモデル幅で線形にし、トレーニング可能なパラメーターを最大 98.4% 削減します。第 2 に、Subset-KL は選択された語彙ロジットのみを実体化します。Top-k モードはピーク トレーニング メモリを最大 70% 削減しますが、重要度をサンプリングしたバリアントは完全な KL に対して不偏の確率的勾配を保持します。これらの節約により、LLaMA-3.3-70B の 482 レンズの高密度アンサンブルが可能になり、同じ深さで残留ストリーム設計の 6 倍のカバレッジを提供します。次に、モデル全体をカバーすることで、単一コンポーネントのレンズではできないことが明らかになります。つまり、動作が最も目に見えるコンポーネントが、介入が最も効果的なコンポーネントである必要はなく、最も効果的な介入は、以前のレンズ研究で調査された注意の対象外にあります。 OmniLens は、3 つのケーススタディ (プロンプト インジェクション検出、マルチホップ メモリ インジェクション、毒性の局所化) にわたって、主要な公開結果を大幅に低コストで再現します。

原文 (English)

Interpreting Language Model Hidden States at Scale

Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network. Trained lenses remain expensive: affine-translator parameters grow quadratically with model width, while exact, full-vocabulary Kullback--Leibler (KL) training dominates memory. Consequently, prior trained lenses have been applied to models of at most 20B parameters and remain tied to particular component types. We present OmniLens, which applies a single lens family to any model-width activation, whether residual stream, attention, or MLP, and combines two independent scaling techniques. First, low-rank translators make per-lens parameter growth linear in model width and reduce trainable parameters by up to 98.4%. Second, Subset-KL materializes only selected vocabulary logits: its Top-k mode cuts peak training memory by up to 70%, while its importance-sampled variant retains unbiased stochastic gradients for the full KL. These savings enable a dense ensemble of 482 lenses for LLaMA-3.3-70B, providing 6x the coverage of a residual-stream design at the same depth. Model-wide coverage then reveals what single-component lenses cannot: the components where a behavior is most visible need not be those where intervention is most effective, and the most effective interventions lie outside the attention heads examined by prior lens studies. Across three case studies (prompt-injection detection, multi-hop memory injection, and toxicity localization), OmniLens reproduces key published results at substantially lower cost.

2026-08-12 13:00 JSTarXiv cs.AIハードウェア/半導体

Hand-Written PTX Tensor-Core GEMM Kernels: A Multi-Precision Study on NVIDIA L4

High-performance Tensor Core kernels rely on a low-level PTX pipeline built from asynchronous data movement with cp.async, warp-level matri…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output

Streaming language-model output creates a release-timing problem: complete-response moderation acts after streamed text has escaped, wherea…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models

Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulti…

2026-08-12 13:00 JSTarXiv cs.AIハードウェア/半導体

Optimal Stopping of Self-Refining Foundation Models

Foundation models can improve their outputs through a self-refinement process driven by external feedback. In this process, the model is em…

2026-08-12 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

Modelling Geographic Atrophy Progression using Implicit Neural Representations

Age-related Macular Degeneration (AMD) is the major cause of blindness in the Western world. Its late dry phase is characterised by irrever…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Attention-Path Fragility as an Uncertainty Signal in Large Language Models

We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a c…

2026-08-12 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

A Methodology for Selecting and Composing Runtime Architecture Patterns for Production LLM Agents

Production LLM agents combine stochastic model outputs with deterministic software systems, yet the boundary between the two is rarely trea…

2026-08-12 12:55 JSTITmedia AI+LLM/生成AIエージェントハードウェア/半導体

NVIDIAが30Bのオープンモデル公開 「OpenClaw」など常時稼働エージェント向けに設計

NVIDIAが30Bのオープンモデル「Nemotron 3.5 Lightning」を公開。「OpenClaw」など常時稼働エージェント向けに設計し、「gpt-oss-120b」同等の性能を約4分の1の規模で実現するという。

2026-08-12 07:00 JSTITmedia AI+ハードウェア/半導体

【注目の企業】キオクシア、なぜこんなに話題? 今からでも間に合う“入門記事”まとめました

話題を呼び続ける半導体メモリ大手のキオクシア。一体なぜこんなにも注目されるのか。同社の“今”が分かる記事をまとめた。

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体

フローバイフロー:高損失ドメインでの AI 出力を制御するためのコンテンツ判定バイパス

これまでの研究では、AI の出力速度 V が人間の認知能力 C_max を超えると、高損失領域では人間による監視が構造的に不可能になることが示されています。ただし、操作上の制約は V 単独ではなく、V x L です。ここで、L は項目ごとの認知負荷を示します。 L はトリアージ、判断、対応で構成され、AI の能力向上に対して非対称に対応します。セマンティックな不確定性は汎用設計に内在するため、モデルの機能が向上してもトリアージ コストは減少しません。応答コストは精度の向上に影響されません。判断コストのみが下方圧力に直面しており、この圧力は多くの場合、真の削減ではなく省略を誘発する形で作用します。したがって、能力の向上は L を削減するのではなく、再構築することになります。 AI の出力が正しいかどうかの評価に基づくガバナンス メカニズムは、その評価を AI に委任して幻覚リスクを継承するか、それを人間に委任して V x L の上限に直面するかのいずれかになります。私たちは、コンテンツを評価せずに監視負荷を制御するガバナンス パラダイムである Flow-by-Flow を提案します。正式な可算特徴に基づく認知コスト スコアは、大量生産に非線形コストを課す一方、制度上のキャパシティ キャップにより処理量が C_max 以内に保たれます。コンテンツ判定バイパス超過経路に対する 4 つの設計不変条件を導き出します。それは、コンテンツ判定なし、検査官能力のスケーラブルな消費なし、ID に縛られたアプリケーションごとの摩擦、およびバッチクリアランスなしです。 1 つの参考実装については、これらの不変式が同時に満たされることを示すために議論されていますが、その実際的な困難は明示的に認識されています。 1,000 のパラメータ描画にわたるモンテカルロ分析の例では、複合マルチメトリック フロー制御が試験の 90.8% で監視強化単独よりも優れていることが示唆されています。

原文 (English)

Flow-by-Flow:Content-Judgment Bypass for Governing AI Output in High-Loss Domains

Prior work showed that human-in-the-loop oversight becomes structurally untenable in high-loss domains when AI output velocity V exceeds human cognitive capacity C_max. The operative constraint, however, is not V alone but V x L, where L denotes per-item cognitive load. L consists of triage, judgment, and response, which respond asymmetrically to AI capability improvement. Triage cost does not decline as models become more capable, because semantic indeterminacy is inherent in general-purpose design. Response cost is invariant to accuracy improvements. Only judgment cost faces downward pressure, and this pressure often operates by inducing omission rather than genuine reduction. Capability improvement therefore restructures L rather than reducing it. Governance mechanisms based on evaluating whether AI output is correct either delegate that evaluation to AI and inherit hallucination risk, or delegate it to humans and face the V x L ceiling. We propose Flow-by-Flow, a governance paradigm that controls supervisory load without evaluating content. A cognitive cost score based on formal, countable features imposes nonlinear costs on high-volume production, while an institutional capacity cap keeps processing volume within C_max. We derive four design invariants for any content-judgment-bypass exceedance pathway: no content judgment, no scalable consumption of examiner capacity, identity-bound per-application friction, and no batch clearance. One reference implementation is discussed to show that these invariants are jointly satisfiable, while its practical difficulties are explicitly acknowledged. An illustrative Monte Carlo analysis across 1,000 parameter draws suggests that composite multi-metric flow control outperforms supervision reinforcement alone in 90.8% of trials.

2026-08-11 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

理由は広く、深くはない: 理由のプレミアムを洗練されたスキルに分割する

言語モデルの推論モードは、複数ステップのエージェント タスクで非推論モードよりも優れていますが、エピソードごとに出力トークンに 3 ~ 6 倍の割増料金を支払います。その多くは、同じドメインのエピソード間で共有されるプロシージャの再導出に費やされます。この繰り返しコストは償却できることを示します。コーディング エージェントは、トレーニング分割からの既存の軌跡の小さなコーパスを分析し、非推論モデルのシステム プロンプトに注入されるコンパクトな自然言語スキルをコンパイルします。 4 つのエージェント ベンチマーク (ALFWorld、tau$^2$-ベンチ通信および小売業、SpreadsheetBench-Verified) 全体で、スキルは、保留されたタスクで GPT-5.4-mini の推論ギャップの 55% ~ 100% 以上を回復し、4 つのうち 2 つで推論モードを完全に超えましたが、出力トークンの数は 2.7 ~ 6 分の 1 で、推論トークンの発行はゼロでした。特に、推論トレースは前提条件ではありません。非推論軌跡のみから抽出されたスキルは、ペアの推論/非推論コーパスから抽出されたスキルと競合し続けますが、2 つのソース間にはドメイン依存の違いがあります。これらの結果を検索レンズを通して解釈します。テスト時の推論は単一のエピソード内の詳細な検索であり、展開ごとに再支払われますが、コーパス蒸留はエピソード全体にわたる広範な検索であり、一度支払われます。この 2 つは、重複する手順の知識を回復します。多くの場合、安価な軌道での幅を広くとったほうが得策です。一部のドメイン (テレコム、スプレッドシートベンチ) に残されたギャップは、真にインスタンスごとの詳細な検索が依然として必要な場所を示しています。

原文 (English)

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

Reasoning modes of language models outperform their non-reasoning counterparts on multi-step agentic tasks, but pay a 3-6x premium in output tokens on every episode -- much of it spent re-deriving procedures that are shared across episodes of the same domain. We show this recurring cost can be amortized: a coding agent analyses a small corpus of existing trajectories from a training split and compiles a compact natural-language skill that is injected into the non-reasoning model's system prompt. Across four agentic benchmarks (ALFWorld, tau$^2$-bench telecom and retail, and SpreadsheetBench-Verified), skills recover 55%-100%+ of the reasoning gap for GPT-5.4-mini on held-out tasks -- exceeding the reasoning mode outright on two of four -- while emitting 2.7-6x fewer output tokens and zero reasoning tokens. Notably, reasoning traces are not a prerequisite: skills distilled from non-reasoning trajectories alone remain competitive with skills distilled from paired reasoning/non-reasoning corpora, with domain-dependent differences between the two sources. We interpret these results through a search lens: test-time reasoning is deep search inside a single episode, re-paid at every deployment, while corpus distillation is wide search across episodes, paid once. The two recover overlapping procedural knowledge, and width over cheap trajectories is often the better buy -- with the residual gap on some domains (telecom, SpreadsheetBench) delineating where genuinely per-instance deep search remains necessary.

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

あなたのプロンプトが唯一のプロンプトではありません: LLM は構造化出力スキーマ記述をどの程度重視しますか?

LLM が事前定義された JSON スキーマを設定する構造化出力は、データのラベル付けと情報抽出のデフォルトのメカニズムとなっていますが、スキーマ記述による 2 番目の命令チャネルも導入されています。 2 つのベンダーの 10 個のモデル構成にわたって、ノンス ラベルを含む単一フィールド分類タスクを使用して、分類ラベル定義をシステム プロンプト、ユーザー プロンプト、またはスキーマ記述に配置する方が適切であるかどうかをテストしました。スキーマの説明は、プロンプトベースの配置を常に上回るパフォーマンスを示したわけではありません。 GPT-4.1 および GPT-5.4 の場合、理由もなく、スキーマ配置はシステム プロンプトのパフォーマンスを 11 ~ 13 パーセントポイント下回っています。しかし、スキーマは不活性なメタデータではありません。プロンプトとスキーマが競合する場合、誤ったスキーマ命令により精度が 5 ~ 45 ポイント低下し、Claude Haiku 4.5 では 52.5% から 7% に低下し、スキーマ命令がプロンプト命令をオーバーライドできることを示し、GPT-5.5 では 100% から 73% に低下しました。さらに、ラベル フィールドの前に必要な中間推論フィールドを追加すると、ヘッドルームが存在する場合にスキーマのみの精度が 15 ~ 24 ポイント向上し、テストしたすべてのケースでシステム プロンプトのみのパフォーマンスを上回りました。この効果は、中程度の推論のクロード ソネット 4.6 でも保持され、拡張された思考だけでは同等の向上が得られませんでした。これは、スキーマ設計がフィールド記述にエンコードされた情報をモデルがどのように効果的に使用するかに影響を与える可能性があることを示唆しています。全体として、これらの結果は、スキーマの影響がモデルに依存していることを示しています。実際には、システム プロンプトが定義の安全なデフォルトのままですが、より大きな規律は、単一の信頼できる情報源を維持し、プロンプト/スキーマのドリフトを防ぐことです。さらに重要なことは、スキーマ設計自体が命令の配置よりも強力な手段となる可能性があることです。実務者は、プロンプトとスキーマを統一された命令画面として扱い、ターゲット モデルの配置とフィールド設計の両方を経験的に検証する必要があります。

原文 (English)

Your Prompt Is Not the Only Prompt: How Much Do LLMs Weight Structured-Output Schema Descriptions?

Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extraction, but it also introduces a second instruction channel through schema descriptions. We tested whether classification-label definitions are better placed in the system prompt, user prompt, or schema description using a single-field classification task with nonce labels across ten model configurations from two vendors. Schema descriptions did not consistently outperform prompt-based placement; for GPT-4.1 and GPT-5.4 without reasoning, schema placement underperformed system prompts by 11-13 percentage points. Yet schemas are not inert metadata: when prompts and schemas conflicted, incorrect schema instructions caused accuracy drops of 5-45 points, with Claude Haiku 4.5 falling from 52.5% to 7%, indicating that schema instructions can override prompt instructions, and GPT-5.5 falling from 100% to 73%. Further, adding a required intermediate reasoning field before the label field improved schema-only accuracy by 15-24 points when headroom existed, exceeding system-prompt-only performance in every case tested. The effect held even for Claude Sonnet 4.6 at medium reasoning, where extended thinking alone did not produce a comparable gain. This suggests that schema design can affect how effectively models use information encoded in field descriptions. Overall, these results indicate that schema influence is model-dependent. In practice, the system prompt remains a safe default for definitions, but the bigger discipline is maintaining a single source of truth and preventing prompt/schema drift. More importantly, schema design itself may be a stronger lever than instruction placement. Practitioners should treat prompts and schemas as a unified instruction surface and empirically validate both placement and field design for their target model.

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

表面的には公平ですか? LLM Recommender の隠れた出力公平性ギャップのベンチマーク

LLM ベースのレコメンダーの公平性監査は主に観察可能な出力に焦点を当てており、安定した推奨は安定した内部処理を反映していると暗黙的に想定されています。私たちは、FairGap を使用して、この仮定に異議を唱えます。FairGap は、性別、年齢、人種にわたる制御された反事実の同一性調査を通じて測定される、観察可能な出力シフト (OBS) と隠れた表現シフト (IBS) の 2 つのレベルで推奨の公平性を共同で評価する最初のベンチマークです。それらの関係は、ユーザーレベルの隠れた出力の不一致を特定するための象限診断を使用して、表現と出力の調整 (ROA) によって要約されます。 FairGap を 3 つのドメインにわたる 6 つのオープンウェイト LLM ファミリに適用すると、広範な隠れた出力デカップリングが明らかになります。ROA が 0.22 を超えることはめったになく、無視できないユーザー集団は、大幅な内部変動にもかかわらず安定した出力を示します。このモードは、出力のみの監査では設計上検出できないモードです。さらに、IBS を最大 8 分の 1 に削減するアクティベーション ステアリングは、同時に OBS を悪化させます。これは、既存のフレームワークが診断する機能を備えていない、内部レベルと出力レベルの公平性の間に根本的な緊張があることを示しています。

原文 (English)

Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders

Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations reflect stable internal processing. We challenge this assumption with FairGap, the first benchmark to jointly evaluate recommendation fairness at two levels: observable output shift (OBS) and hidden representation shift (IBS), measured through controlled counterfactual identity probes across gender, age, and race. Their relationship is summarized via Representation-Output Alignment (ROA), with quadrant diagnostics for identifying user-level hidden-output mismatch. Applied to six open-weight LLM families across three domains, FairGap reveals pervasive hidden-output decoupling: ROA rarely exceeds 0.22, and a non-negligible user population shows stable outputs despite substantial internal shifts, a mode that output-only audits cannot detect by design. Further, activation steering that reduces IBS by up to 8x simultaneously worsens OBS, demonstrating a fundamental tension between internal and output-level fairness that existing frameworks are unequipped to diagnose.

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

LLMVisor: マルチテナント LLM サービス用のリアルタイム レイテンシー アトリビューション モデル

LLM 推論がマルチテナント GPU クラスターに移行すると、同時バッチ処理によりスループットが向上しますが、テナントごとの使用がわかりにくくなり、制御が制限されます。推論エンジンの部分共有を有効にするには、スケジューリング ループ内で実行できるほど正確で軽量な、リアルタイムのリクエストごとのアトリビューション プリミティブが必要です。 LLMVisor は、FLOP とメモリ I/O トラフィックに比例する特徴に対する簡潔な区分線形形式を介して、メモリに依存するフェーズと計算に依存するフェーズをキャプチャする、ルーフライン ガイド付きレイテンシ アトリビューション モデルです。 LLLMVisor は、バッチ レイテンシーを追加のリクエストごとのシェアに分解し、マイクロ秒スケールで効率的に実行します。さまざまなテンソル並列処理とワークロードの組み合わせの下で、A100/H100 GPU 上の Llama 3.1-8B および Qwen 2.5-14B/32B で LLMVisor を評価しました。トークン数ベースラインと比較して、LLMVisor はほぼ完璧な R 二乗を達成し、バッチの変動性やシーケンスの発散にもかかわらず、プリフィルの場合は p90 と p99 でそれぞれ最大 2.5 倍と 3.3 倍、デコードの場合は最大 3.5 倍と 4.4 倍まで相対誤差を低減します。

原文 (English)

LLMVisor: A Real-Time Latency Attribution Model for Multi-Tenant LLM Serving

As LLM inference shifts to multi-tenant GPU clusters, co-batching improves throughput but obscures per-tenant usage and limits control. Enabling fractional sharing of the inference engine requires a real-time, per-request attribution primitive that is accurate and light enough to run inside the scheduling loop. We present LLMVisor, a roofline-guided latency attribution model that captures the memory-bound and compute-bound phases via a concise piecewise-linear form over features proportional to FLOPs and memory I/O traffic. LLMVisor decomposes batch latency into additive, per-request shares and runs efficiently at microsecond scale. We evaluate LLMVisor across Llama 3.1-8B and Qwen 2.5-14B/32B on A100/H100 GPUs under varying tensor parallelism and workload mixes. Compared to a token-count baseline, LLMVisor attains near-perfect R-squared and reduces relative error by up to 2.5x and 3.3x at p90 and p99, respectively, for prefill, and by up to 3.5x and 4.4x for decode, despite batching variability and sequence divergence.

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体

Aero Realtime: Fully Aligned Input-Output Streams for Low-Latency Streaming Multimodal Generation

Existing streaming multimodal models process observations incrementally but still follow a turn-based prefill-then-decode pattern, making t…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

AI Evaluation Should Measure Verification Cost, Not Correctness Alone

The reliability of AI generative models is typically measured by output correctness, yet in practice it depends on the effort required to v…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

ICM Out! Better Tournament Strategy from Computed Continuations, vs. Solvers and LLMs

The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against thos…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

PACE: A Playback-Aligned Context Engine for LLM-Based Full-Duplex Voice Dialogue

LLM-based full-duplex voice services allow users to speak while the assistant is responding. Because servers can generate output and advanc…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体

LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICs

Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendat…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体

Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical Limits

Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate fo…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体

Two-Step MV-DeepONet: Probabilistic Operator Learning for Uncertainty Propagation Driven by Random Input Fields

Forward uncertainty propagation in complex physical systems can induce structured covariance across field-valued outputs. For a probabilist…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体

AkasicDB: Demonstrating Omni RAG with a Unified Vector-Graph-Relational DBMS

Recent Retrieval-Augmented Generation (RAG) systems increasingly combine vector retrieval with structured knowledge, such as Graph RAG and…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute

Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matte…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Dual-Adversarial Safety Alignment: Cultivating Intrinsic Threat Comprehension in LRMs

Large reasoning models (LRMs) achieve remarkable success on complex tasks but remain vulnerable to harmful prompts that induce unsafe outpu…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

MonitorBench: A Comprehensive Benchmark for Chain-of-Thought Monitorability in Large Language Models

Large language models (LLMs) can generate chains of thought (CoTs) that are not always causally responsible for their final outputs. When s…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

The Transformer Revolution, Part 1: Dynamic Processing through Output-Weight Interconnections

This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体

生成 AI における認識論的信頼性: 一か八かのワークフローにおける信頼を保証するための規範的なフレームワーク

生成 AI システムは、一か八かの専門的な状況で導入されることが増えており、その出力によって、ユーザーが何を信じ、どのように推論し、何を定着として扱うのかが決まります。これは、責任ある AI に対する中心的な疑問を引き起こします。それは、どのような状況下で、生成的な AI の出力への依存が、行動によって引き起こされるのではなく、認識論的に正当化されるのかということです。既存のフレームワークでは主に、AI の出力が正確か、公平か、説明可能か、安全か、ユーザーに信頼されているかどうかが問われます。これらの質問は依然として必要であり、それぞれが正当な信頼に貢献する可能性があります。ただし、それらは、明確な評価目標、つまりユーザーが AI の出力を自分の推論への入力として扱うことが正当化される条件として、保証された信頼性を直接指定していません。私たちは、これには認識論的信頼性、つまりシステムを認識論的に信頼に値するものにするものについての説明が必要であると主張します。能力と聴衆指向としての信頼性に関する哲学的説明に基づいて、私たちは 3 つの共同で必要で代替不可能な条件からなる構成的規範の枠組みを開発します。まず、認識論的謙虚さには、システムがその能力の限界を表現し、伝達することが必要です。第 2 に、認識的アクセスには、ユーザーがコンテキスト内で出力を検査、質問、および異議申し立てできるようにするシステムが必要です。第三に、認識論的不正義に抵抗するには、システムがユーザーを正当な認識論的主体として認識し、ユーザーの知識や経験を疎外しないようにする必要があります。法的推論、医学的推論、雇用における実際の事例分析を通じて、認識論的謙虚さ、認識論的アクセス、認識論的不正義に対する抵抗の失敗が、精度、公平性、使いやすさの標準的な尺度だけでは対処できない結果的な損害をどのように生み出す可能性があるかを示します。最後に、出力の正確さのみではなく、認識論的に保証された信頼性を中心に構成された GenAI システムの設計と評価への影響を概説します。

原文 (English)

Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows

Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientation, we develop a constitutive normative framework comprising three jointly necessary and non-fungible conditions. First, epistemic humility requires systems to represent and communicate the limits of their competence. Second, epistemic access requires systems to enable users to inspect, question, and contest outputs in context. Third, resistance to epistemic injustice requires systems to recognise users as legitimate epistemic agents and avoid marginalising their knowledge and experience. Through real-world case analyses in legal reasoning, medical reasoning, and hiring, we show how failures of epistemic humility, epistemic access, and resistance to epistemic injustice can produce consequential harms that standard measures of accuracy, fairness, and usability do not address on their own. We conclude by outlining design and evaluation implications for GenAI systems organised around epistemically warranted reliance rather than output correctness alone.

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured Output

Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models…

2026-08-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

NeuroBreak: Unveil Internal Jailbreak Mechanisms in Large Language Models

In deployment and application, large language models (LLMs) typically undergo safety alignment to prevent illegal and unethical outputs. Ho…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体

Attn-QAT: 4-Bit Attention With Quantization-Aware Training

Achieving reliable 4-bit attention is a prerequisite for end-to-end FP4 computation on emerging FP4-capable GPUs, yet attention remains the…

2026-08-11 13:00 JSTarXiv cs.AIハードウェア/半導体

Riemannian Deep Learning: Modules, Networks, and Geometries

Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific…

2026-08-11 07:11 JSTITmedia AI+エージェントハードウェア/半導体研究/論文

Meta、ローカル動作に特化したオープンモデル「Muse Glimmer」公開 Apache 2.0で提供

MetaのAI研究部門は、約296億パラメータのオープンウェイトAIモデル「Muse Glimmer」を公開した。Apache 2.0ライセンスで提供され、PCやMacのGPU1基で動作する。上位モデルからの蒸留によりローカル環境でのエージェント処理やマルチモーダル推論に最適化…

2026-08-10 21:00 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

Discovered Materials is playing AI whack-a-mole to hunt cooler chips

Discovered Materials raised $9 million to fund the hunt for more novel materials to build more efficient chips.

2026-08-10 13:00 JSTarXiv cs.AIハードウェア/半導体

ミュオンで訓練されたトランスフォーマーの表現-読み出しインターフェースにおけるグロッキング後の崩壊

標準的な分割では、Muon は隠れ行列と AdamW の埋め込み/出力ヘッドを取得します。 Muon はモジュラー加算を高速化しますが、そのソリューションは保持されません。 $(a+b) \bmod 113$ grok 以降の 9 つの構成はすべて一般化されません。 5 つのシード全体で、選択された AdamW 参照は 4 つでしきい値を下回り、27.59% に達します。不安定性は、2 つの係数、2 つの幅、2 つのトレーニング分数、減算、および深さにわたって持続します。障害は、表現と読み出しのインターフェースで発生し、損失によって選択されなかった可逆マップまでの共同でのみ識別されます。トレーニング セットを解いた後、勾配は $10^{-6}$ のオーダーに下がり、オプティマイザーの反応は異なります。ステップサイズの弾性は、Muon では -0.03 であるのに対し、AdamW では +1.5 であり、Muon グループはパラメーターごとに 8.0 倍速く移動します。ビットが同一の状態から、どちらかのグループをフリーズすると障害が防止されます。埋め込み/読み出しを凍結すると、グロッキング後の 451,400 ステップと 5 つのペアのシードにわたる 5 回の実行でそれが削除されます。凍結されていないアームは 137 ~ 321 の閾値以下の評価を記録しますが、凍結されたアームはありません。ミューオンの正規化と直交化を削除することは代わりにはなりません。これは表現を 326 の有効な共役ペアから 4 つに崩壊させ、再発性の崩壊を示さず、最終的に失敗します。フーリエ フィルタリングは、回路障害をマスキングから分離します。 5 つのシードと 3 つの体制にわたる 43 のチェックポイントにわたって、タスクを調整した家族だけでちょうど 100% に達します。回路障害が発生すると、それはもはやタスクを解決しません。マスキングでは、完全なモデルが 45.85% に達する間、完全なままであり、エラーを含むすべての例でプラスのマージンが得られますが、ほぼ等しい敵対的な残りによって投票されます。再スケーリングすると 99.9% が回復します。グロッキングは、同じ状態が上向きに解決されることです。タスクはファミリーを選択し、減算の下で $(k,k)$ を $(k,-k)$ に交換します。突然の崩壊では、標準フーリエ サポートは変化せず、電力分布コサインは 0.9899 のままです。

原文 (English)

Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained Transformers

Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold. All nine configurations on $(a+b) \bmod 113$ grok and later lose generalization. Across five seeds the selected AdamW reference falls below threshold on four, reaching 27.59%. Instability persists across two moduli, two widths, two training fractions, subtraction, and depth. The failure arises at the representation-readout interface, identified only jointly up to an invertible map unselected by the loss. After solving the training set, the gradient falls to order $10^{-6}$ and the optimizers respond differently: step-size elasticity is -0.03 for Muon versus +1.5 for AdamW, and the Muon group moves 8.0 times faster per parameter. From bit-identical states, freezing either group prevents failure. Freezing embeddings/readout removes it in five runs over 451,400 post-grokking steps and five paired seeds: unfrozen arms record 137-321 sub-threshold evaluations, frozen arms none. Removing Muon's normalization and orthogonalization is no substitute: it collapses representation from 326 effective conjugate pairs to 4, shows no recurrent collapse, and fails terminally. Fourier filtering separates circuit failure from masking. Across 43 checkpoints over five seeds and three regimes, the task-aligned family reaches exactly 100% alone. In circuit failure it no longer solves the task; in masking it remains perfect while the full model reaches 45.85%, giving a positive margin on every example, including errors, but being outvoted by a near-equal adversarial remainder. Rescaling it restores 99.9%; grokking is the same condition resolving upward. The task selects the family, swapping $(k,k)$ for $(k,-k)$ under subtraction. Across an abrupt collapse, standard Fourier support is unchanged and the power-distribution cosine remains 0.9899.

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Beyond "AI Language": The case for the idiolectal nature of LLM output

While large language model outputs are frequently analysed as a collective super variety termed "AI language," this chapter argues that thi…

2026-08-10 13:00 JSTarXiv cs.AIハードウェア/半導体

Scalable High-Fidelity Macromolecular Docking for GPU-Accelerated Supercomputers

Flexible macromolecular docking offers high-fidelity predictions of biomolecular interactions, but remains prohibitively expensive at scale…

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity

While post-training improves the capabilities of large language models (LLMs), it generally lowers their output diversity and creativity, n…

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Counterfactual Simulation Training for Chain-of-Thought Faithfulness

Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known proble…

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

MetaSICL: メタ音声インコンテキスト学習による Audiroty LLM の適応

聴覚大規模言語モデル (LLM) は、幅広い音声理解タスクにわたって強力なパフォーマンスを実証しています。それにもかかわらず、リソースが少ないタスクに適用すると、苦労することがよくあります。ドメイン内のラベル付きデータが不足しているか、実際のテスト分布と一致しない場合、直接の微調整は脆弱になる可能性があります。 In-Context Learning (ICL) は、いくつかのドメイン内デモンストレーションを条件付けして聴覚 LLM を適応させることにより、トレーニング不要の推論時間ソリューションを提供します。この研究では、$\textit{Vanilla ICL}$ が、選択されたモデルのさまざまな音声タスクおよびオーディオ タスクにわたってゼロショット パフォーマンスを向上させることを最初に示します。これは、この ICL 適応機能がマルチモーダル設定に一般化できることを示唆しています。これに基づいて、$\textbf{Meta Speech In-Context Learning (MetaSICL)}$ を提案します。これは、モデルのインコンテキスト学習能力を強化することを目的とした、さまざまなタスクからの高リソース音声データのみを利用するトレーニング後のレシピです。実験によれば、私たちが提案した方法は、リソースが少ないシナリオでは直接の微調整よりも優れています。

原文 (English)

MetaSICL: Globalizing Auditory LLMs for Underserved Speakers and Languages via Meta Speech In-Context Learning

Generative AI for speech and audio is increasingly expected to serve users across languages, cultures, and communities, yet current auditory Large Language Models (LLMs) are still largely trained and evaluated on high-resource data. Globalizing such systems requires handling low-resource settings, where the target speakers, languages, or tasks are poorly represented in training data. In these regimes, collecting enough labeled in-domain data is often impractical, and the small corpora available may still under-represent the test distribution, making direct fine-tuning brittle under domain shift. In-Context Learning (ICL) offers an alternative: instead of updating model parameters for every underserved community, an auditory LLM can adapt at inference time by conditioning on a few local demonstrations. However, vanilla speech ICL remains limited because most auditory LLMs are not explicitly trained to use such demonstrations effectively. We address this gap with Meta Speech In-Context Learning (MetaSICL), a post-training recipe that strengthens an auditory LLM's in-context adaptation ability using only abundant high-resource speech data. Although MetaSICL never trains on the target low-resource domains, it improves performance across two backbones on children's ASR, audio understanding/reasoning, and speech translation and ASR in directions and languages unseen in post-training. We further study the case where some in-domain data is available, using low-resource language ASR as a case study, since recognition for underserved languages is central to globalizing generative AI. Here, using MetaSICL as a warmup for in-domain reinforcement learning yields the strongest results, outperforming direct fine-tuning across five typologically diverse languages. Overall, MetaSICL offers a practical route toward globalizing auditory LLMs by building inference-time adaptation into the model.

2026-08-10 13:00 JSTarXiv cs.AIハードウェア/半導体

Predictive Multi-Tier Memory Management for KV Cache in Large-Scale GPU Inference

Key-value (KV) cache memory management is the primary bottleneck limiting throughput and cost-efficiency in large-scale GPU inference servi…

2026-08-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Topology-Aware Data Movement for Disaggregated GPU Inference

Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run o…

2026-08-10 05:35 JSTTechCrunch AIハードウェア/半導体

Embattled hedge fund Situational Awareness invests $400M in chip startup Source Foundry

The AI-focused hedge fund is still making some big bets.

2026-08-07 13:00 JSTarXiv cs.AIハードウェア/半導体

生成 AI における認識論的信頼性: 一か八かのワークフローにおける信頼を保証するための規範的なフレームワーク

生成 AI システムは、一か八かの専門的な状況で導入されることが増えており、その出力によって、ユーザーが何を信じ、どのように推論し、何を定着として扱うのかが決まります。これは、責任ある AI に対する中心的な疑問を引き起こします。それは、どのような状況下で、生成的な AI の出力への依存が、行動によって引き起こされるのではなく、認識論的に正当化されるのかということです。既存のフレームワークでは主に、AI の出力が正確か、公平か、説明可能か、安全か、ユーザーに信頼されているかどうかが問われます。これらの質問は依然として必要であり、それぞれが正当な信頼に貢献する可能性があります。ただし、それらは、明確な評価目標、つまりユーザーが AI の出力を自分の推論への入力として扱うことが正当化される条件として、保証された信頼性を直接指定していません。私たちは、これには認識論的信頼性、つまりシステムを認識論的に信頼に値するものにするものについての説明が必要であると主張します。能力と聴衆指向としての信頼性に関する哲学的説明に基づいて、私たちは 3 つの共同で必要で代替不可能な条件からなる構成的規範の枠組みを開発します。まず、認識論的謙虚さには、システムがその能力の限界を表現し、伝達することが必要です。第 2 に、認識的アクセスには、ユーザーがコンテキスト内で出力を検査、質問、および異議申し立てできるようにするシステムが必要です。第三に、認識論的不正義に抵抗するには、システムがユーザーを正当な認識論的主体として認識し、ユーザーの知識や経験を疎外しないようにする必要があります。法的推論、医学的推論、雇用における実際の事例分析を通じて、認識論的謙虚さ、認識論的アクセス、認識論的不正義に対する抵抗の失敗が、精度、公平性、使いやすさの標準的な尺度だけでは対処できない結果的な損害をどのように生み出す可能性があるかを示します。最後に、出力の正確さのみではなく、認識論的に保証された信頼性を中心に構成された GenAI システムの設計と評価への影響を概説します。

原文 (English)

Epistemic Trustworthiness in Generative AI: A Normative Framework for Warranted Reliance in High-Stakes Workflows

Generative AI systems are increasingly deployed in high-stakes professional contexts, where their outputs shape what users believe, how they reason, and what they treat as settled. This raises a central question for responsible AI: under what conditions is reliance on generative AI outputs epistemically warranted rather than behaviourally induced? Existing frameworks largely ask whether AI outputs are accurate, fair, explainable, safe, or trusted by users. These questions remain necessary, and each can contribute to warranted reliance. However, they do not directly specify warranted reliance as a distinct evaluative target: the conditions under which users are justified in treating AI outputs as inputs into their own reasoning. We argue that this requires an account of epistemic trustworthiness: what makes a system epistemically worthy of reliance. Drawing on philosophical accounts of trustworthiness as competence and audience-orientation, we develop a constitutive normative framework comprising three jointly necessary and non-fungible conditions. First, epistemic humility requires systems to represent and communicate the limits of their competence. Second, epistemic access requires systems to enable users to inspect, question, and contest outputs in context. Third, resistance to epistemic injustice requires systems to recognise users as legitimate epistemic agents and avoid marginalising their knowledge and experience. Through real-world case analyses in legal reasoning, medical reasoning, and hiring, we show how failures of epistemic humility, epistemic access, and resistance to epistemic injustice can produce consequential harms that standard measures of accuracy, fairness, and usability do not address on their own. We conclude by outlining design and evaluation implications for GenAI systems organised around epistemically warranted reliance rather than output correctness alone.

2026-08-07 13:00 JSTarXiv cs.AIハードウェア/半導体

GSBF: 環境を考慮したビームフォーミングのためのガウス スプラッティング

ビームフォーミングは、多入力多出力 (MIMO) 通信システムにおいて重要な役割を果たします。ただし、従来のビームフォーミング設計では通常、正確な瞬間チャネル状態情報 (CSI) と反復的な最適化が必要であり、これによりパイロットのオーバーヘッドと計算の複雑さが大幅に増加します。無線伝播は本質的に物理幾何学によって支配されることを認識し、マルチモーダル データに基づいて環境を考慮したビームフォーミング (GSBF) パイプライン用の 3D ガウス スプラッティングを開発します。これは、永続的な 3D ガウス表現を通じて環境を特徴付けます。具体的には、GSBF は相反性を保持する双方向球面ガウス (Bi-SG) カーネルを使用して環境散乱応答をモデル化し、両面電磁ラスタライゼーションを実行して角度プロパゲータ マップをレンダリングします。次に、レンダリングされたマップは、過剰に完成した配列多様体辞書を通じて集約され、定弾性ビームフォーマーに投影されます。これにより、オンラインの瞬間的な CSI を使用せずに、アクセス ポイント (AP) の姿勢とユーザーの位置から直接ビームが合成されます。シミュレーションでは、GSBF が一貫して網羅的ビーム アライメント (EBA) などのベースラインを上回り、遅延が低いことが実証されています。

原文 (English)

GSBF: Gaussian Splatting for Environment-Aware Beamforming

Beamforming plays a key role in multiple-input-multiple-output (MIMO) communication systems. However, conventional beamforming design normally requires accurate instantaneous channel state information (CSI) and iterative optimization, which incur substantial pilot overhead and computational complexity. Recognizing that radio propagation is intrinsically governed by the physical geometry, we develop a 3D Gaussian splatting for environment-aware beamforming (GSBF) pipeline based on multi-modal data, which characterizes the environment through a persistent 3D Gaussian representation. Specifically, GSBF models the environmental scattering response with reciprocity-preserving bidirectional spherical Gaussian (Bi-SG) kernels and performs two-sided electromagnetic rasterization to render an angular propagator map. The rendered map is then aggregated through an over-complete array-manifold dictionary and projected to the constant-modulus beamformers, thereby synthesizing beams directly from the access point (AP) pose and user position without online instantaneous CSI. Simulations demonstrate that GSBF consistently outperforms baselines such as exhaustive beam alignment (EBA) with lower latency.

2026-08-07 13:00 JSTarXiv cs.AIハードウェア/半導体

Why the Third Axis Is Freedom

In generative training, a model produces an output and is penalised for its difference from an example. With one output per comparison, a m…

2026-08-07 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体研究/論文

EviGraph: Evidence-Guided Autonomous Research Agents

Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported…

2026-08-07 13:00 JSTarXiv cs.AIハードウェア/半導体

Output-Aware Rotation for INT2 KV-Cache Quantization

The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-lo…

2026-08-06 19:30 JSTITmedia AI+LLM/生成AIハードウェア/半導体ビジネス/資金調達

ソフトバンクG、投資利益1.8兆円を支えた「OpenAIではない“あの半導体メーカー”」の正体

ソフトバンクGは、第1四半期の投資利益が1兆8594億円だったと発表した。投資利益を押し上げたのは、OpenAIでもArmでもない。歴史的な経営難に陥っていた“あの半導体メーカー”だった。

2026-08-06 13:56 JSTITmedia AI+ロボティクスハードウェア/半導体

NVIDIA、自動運転向けオープンモデルを商用利用可に 新モデルは「卓越した性能」うたう

NVIDIAが自動運転向けAIモデル「Alpamayo」ファミリーを商用利用可能なオープンライセンスで提供開始。新モデル「Alpamayo 2 Super」は推論ベンチマークで首位になるなど卓越した性能をうたう。

2026-08-06 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体研究/論文

EviGraph: Evidence-Guided Autonomous Research Agents

Autonomous research agents can generate hypotheses, execute experiments, and draft manuscripts, yet their outputs often contain unsupported…

2026-08-06 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

EDATracer: An Agentic Framework for Large-Scale EDA Artifact Analysis

Modern chip design relies on electronic design automation (EDA) tools that generate large, heterogeneous artifacts, including source files,…

2026-08-06 13:00 JSTarXiv cs.AIハードウェア/半導体

Masked diffusion enables coherent beat tracking

Current neural networks for beat tracking generate invalid outputs, such as consecutive downbeats and erratic tempo changes, even when thes…

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Hardware Design and Security in the Era of Chiplets and LLMs

The semiconductor industry is undergoing a dual revolution: the shift toward heterogeneous 2.5D chiplet systems and the integration of Larg…

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important fo…

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

意味的等価性を超えて: LLM 不確実性定量化のための論理グラフ

大規模言語モデル (LLM) は、自信を持って記述されているにもかかわらず信頼性の低い出力を生成することが多く、安全性が重視されるアプリケーションへの展開に重大な課題をもたらします。意味論的エントロピーなどの既存の不確実性指標は、意味論的等価性のレベルで一致を捉えますが、異なる答え間の論理的関係をほとんど無視します。その結果、生成された応答の形式は多様であるが論理的に互換性がある(たとえば、粒度または特異性のみが異なる)設定では、不確実性を過大評価し、幻覚に誤ってフラグを立てる傾向があります。私たちは、回答間の含意と非互換性を明示的にモデル化するフレームワークである Logical Graph Uncertainty (LGU) を提案します。 LGU は含意チェーンに沿って確率質量を集計し、論理的に最大の仮説に対するエントロピーを計算し、それらの間の相互非互換性にペナルティを課します。複数の質問応答ベンチマークにわたって、LGU は既存の手法よりも不確実性の推定を一貫して改善し、データセット全体でセマンティック エントロピー ベースラインを最大 +7.1% AUROC および +3.5% AUARC 上回っています。

原文 (English)

Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification

Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications. Existing uncertainty metrics such as semantic entropy capture agreement at the level of semantic equivalence, but largely ignore the logical relationships between distinct answers. As a result, they tend to overestimate uncertainty and falsely flag hallucinations in settings where generated responses are diverse in form yet logically compatible (e.g., differing only in granularity or specificity). We propose Logical Graph Uncertainty (LGU), a framework that explicitly models implication and incompatibility among answers. LGU aggregates probability mass along entailment chains onto the most specific hypotheses the answers support, measures the entropy of the resulting distribution, and penalizes mutual incompatibility among those hypotheses. Across multiple question-answering benchmarks and model families, LGU ranks first on average among existing uncertainty measures, with its largest gains---up to +7.1\% AUROC and +3.5\% AUARC over semantic entropy---on questions whose sampled answers are logically structured.

2026-08-06 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体研究/論文

アウトプットが分散すると、認識論の修正が続くのか? Machine Collective 向けのブラックボックス カップリング診断

集合知の研究では、意見の相違を認識論的多様性の証拠として扱います。エージェントが異なる見解を表明した場合、グループは修正する能力を保持する必要があります。 LLM 集合体では、このプロキシが壊れる可能性があります。エージェントは、同じ結論を維持しながら、多様に見える議論を生成できます。私たちは分散と修正の結合を操作します。つまり、埋め込み空間における集団の成果の分散を検証可能に増加させる介入が、前提を保持した再定式化ではなく認識論的立場の真の修正を伴う度合いです。診断はブラックボックスです。生成されたテキストのみを処理し、生成モデルの内部表現については主張しません。 2 つのチャネルは独立して測定されます。出力チャネルであるコヒーレンス インデックス (CI) は、介入によって出力分散が変化したことを検証します。認識チャネル、ターンごとのスタンスの注釈は、集合体が修正されたかどうかを測定します。我々は、この結合領域を推定するための再利用可能な方法として、出力が過収束したときに再微分プロトコル (RDP) を挿入するメタ予測明瞭度システム (MPCS) を備えた CI を提案します。 2 つの構成 (gpt-4o-mini および gemini-2.5-flash、条件ごとに 310 ペアのエピソード) から 5 つのエージェント集合を評価します。 gpt-4o-mini では、条件付き反対は誤った前提の回復を +17.7 ポイント (p<1e-6) 改善しますが、静的なペルソナの多様性は回復に悪影響を及ぼします (-8.1、p=.007)。ジェミニ 2.5 フラッシュでは、分散の低下が確認されたにもかかわらず、同等の予算で同じ介入を行っても利益は得られませんでした (26.1% 対 27.1%、p=.84)。 2 つの治療効果は互いに異なります (z=3.79、p<.001)。メカニズムのタグ付けは、ジェミニがフレームワーク内の反対意見を通じて誤った前提を維持していることを示しています。タグ付けされたRDP後の回答の94%が譲歩せずに再定式化しました(GPTでは24%)。精度とともに、介入ごとのスタンスシフトと前提保存率を報告することをお勧めします。

原文 (English)

When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine Collectives

Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion. We operationalize dispersion-revision coupling: the degree to which an intervention that verifiably increases the dispersion of a collective's outputs in embedding space is accompanied by genuine revision of its epistemic stance rather than premise-preserving reformulation. The diagnostic is black-box: it operates on generated text alone and makes no claims about the internal representations of the generating models. Two channels are measured independently: an output channel, the Coherence Index (CI), verifies that the intervention changed output dispersion; an epistemic channel, per-turn stance annotation, measures whether the collective revised. We propose CI with the Meta-Predictive Clarity System (MPCS), which inserts a Re-Differentiation Protocol (RDP) when outputs over-converge, as a reusable method for estimating this coupling regime. We evaluate five-agent collectives from two configurations (gpt-4o-mini and gemini-2.5-flash; 310 paired episodes per condition). On gpt-4o-mini, conditional dissent improves false-premise recovery by +17.7 points (p<1e-6) while static persona diversity harms recovery (-8.1, p=.007). On gemini-2.5-flash, the same intervention at a comparable budget yields no gain (26.1% vs 27.1%, p=.84) despite a verified dispersion drop; the two treatment effects differ from each other (z=3.79, p<.001). Mechanism tagging shows Gemini preserves the false premise via intra-framework dissent: 94% of tagged post-RDP responses reformulate rather than concede (vs 24% on GPT). We recommend reporting per-intervention stance shift and premise-preservation rate alongside accuracy.

2026-08-05 23:13 JSTTechCrunch AILLM/生成AIハードウェア/半導体

Anthropic is hiring an AI chip design team

Anthropic is building a team for designing its own custom AI chips. The Claude maker said it would co-design hardware and models to help it…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体研究/論文

アウトプットが分散すると、認識論の修正が続くのか? Machine Collective 向けのブラックボックス カップリング診断

集合知の研究では、意見の相違を認識論的多様性の証拠として扱います。エージェントが異なる見解を表明した場合、グループは修正する能力を保持する必要があります。 LLM 集合体では、このプロキシが壊れる可能性があります。エージェントは、同じ結論を維持しながら、多様に見える議論を生成できます。私たちは分散と修正の結合を操作します。つまり、埋め込み空間における集団の成果の分散を検証可能に増加させる介入が、前提を保持した再定式化ではなく認識論的立場の真の修正を伴う度合いです。診断はブラックボックスです。生成されたテキストのみを処理し、生成モデルの内部表現については主張しません。 2 つのチャネルは独立して測定されます。出力チャネルであるコヒーレンス インデックス (CI) は、介入によって出力分散が変化したことを検証します。認識チャネル、ターンごとのスタンスの注釈は、集合体が修正されたかどうかを測定します。我々は、この結合領域を推定するための再利用可能な方法として、出力が過収束したときに再微分プロトコル (RDP) を挿入するメタ予測明瞭度システム (MPCS) を備えた CI を提案します。 2 つの構成 (gpt-4o-mini および gemini-2.5-flash、条件ごとに 310 ペアのエピソード) から 5 つのエージェント集合を評価します。 gpt-4o-mini では、条件付き反対は誤った前提の回復を +17.7 ポイント (p<1e-6) 改善しますが、静的なペルソナの多様性は回復に悪影響を及ぼします (-8.1、p=.007)。ジェミニ 2.5 フラッシュでは、分散の低下が確認されたにもかかわらず、同等の予算で同じ介入を行っても利益は得られませんでした (26.1% 対 27.1%、p=.84)。 2 つの治療効果は互いに異なります (z=3.79、p<.001)。メカニズムのタグ付けは、ジェミニがフレームワーク内の反対意見を通じて誤った前提を維持していることを示しています。タグ付けされたRDP後の回答の94%が譲歩せずに再定式化しました(GPTでは24%)。精度とともに、介入ごとのスタンスシフトと前提保存率を報告することをお勧めします。

原文 (English)

When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives

Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion. We operationalize dispersion-revision coupling: the degree to which an intervention that verifiably increases the dispersion of a collective's outputs in embedding space is accompanied by genuine revision of its epistemic stance rather than premise-preserving reformulation. The diagnostic is black-box: it operates on generated text alone and makes no claims about the internal representations of the generating models. Two channels are measured independently: an output channel, the Coherence Index (CI), verifies that the intervention changed output dispersion; an epistemic channel, per-turn stance annotation, measures whether the collective revised. We propose CI with the Meta-Predictive Clarity System (MPCS), which inserts a Re-Differentiation Protocol (RDP) when outputs over-converge, as a reusable method for estimating this coupling regime. We evaluate five-agent collectives from two configurations (gpt-4o-mini and gemini-2.5-flash; 310 paired episodes per condition). On gpt-4o-mini, conditional dissent improves false-premise recovery by +17.7 points (p<1e-6) while static persona diversity harms recovery (-8.1, p=.007). On gemini-2.5-flash, the same intervention at a comparable budget yields no gain (26.1% vs 27.1%, p=.84) despite a verified dispersion drop; the two treatment effects differ from each other (z=3.79, p<.001). Mechanism tagging shows Gemini preserves the false premise via intra-framework dissent: 94% of tagged post-RDP responses reformulate rather than concede (vs 24% on GPT). We recommend reporting per-intervention stance shift and premise-preservation rate alongside accuracy.

2026-08-05 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

The Transformer Revolution, Part 1: Dynamic Processing through Output- Weight Interconnections

This paper offers a new interpretation of the Transformer during inference. Against the "stochastic parrot" view that large language models…

2026-08-05 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization

Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measureme…

2026-08-05 13:00 JSTarXiv cs.AIハードウェア/半導体

Output-Aware Rotation for INT2 KV-Cache Quantization

The key-value (KV) cache has become a major memory and bandwidth bottleneck in long-context large language model inference, making ultra-lo…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators

Kernel generation for hardware accelerators such as GPUs and NPUs has become a proving ground for large language models (LLMs), and state-o…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, histor…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics

Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment. However, traditional e…

2026-08-05 13:00 JSTarXiv cs.AIハードウェア/半導体

Beyond Either-Or Reasoning: Transduction and Induction as Cooperative Problem-Solving Paradigms

Traditionally, in Programming-by-example (PBE) the goal is to synthesize a program from a small set of input-output examples. Lately, PBE h…

2026-08-05 13:00 JSTarXiv cs.AIハードウェア/半導体

One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting

Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores. We sh…

2026-08-05 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

MIMIC-MJX: Neuromechanical Emulation of Animal Behavior

The primary output of the nervous system is movement and behavior. While recent advances have democratized pose tracking during complex beh…

2026-08-05 13:00 JSTarXiv cs.AIハードウェア/半導体

Estimating Tail Risks in Language Model Output Distributions

Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these model…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution

Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts. I…

2026-08-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation

Post-training large language models (LLMs) via reinforcement learning (RL) has significantly advanced code generation capabilities. To bypa…

2026-08-05 04:28 JSTTechCrunch AIハードウェア/半導体

Nvidia doesn’t mess around: A week after open AI industry group formed, it’s already showing progress

The week-old Open Secure AI Alliance, spearheaded by Nvidia and grown to over 120 companies, already has proposals out for defending agains…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

ツールの仕様が重要: AI エージェントの安全性リスクを明らかにし、軽減する

AI エージェントは、外部ツールを使用して大規模言語モデル (LLM) を拡張し、複雑なタスクを実行し、モデルの出力を結果として現実世界のアクションに変換できるようにします。しかし、LLM はエージェントとして導入されると安全性が大幅に低下することが多く、この低下の原因は依然としてよくわかっていません。この論文では、スキーマ形式のツール仕様がエージェントの安全性低下の主な原因であることを特定し、ホワイトボックス表現分析を通じて、それらがモデルの内部拒否シグナルを弱め、安全でないツールの実行に寄与していることを示します。この発見に基づいて、私たちは安全性の判断をツールの実行から切り離す推論時の保護手段である SafeKeep を提案します。これは、元のスキーマ形式の実行仕様を保持しながら、平坦化されたテキストのツール仕様を使用してリクエストを評価します。 2 つの代表的なベンチマークと、ホワイト ボックス モデルとブラック ボックス モデルの両方を含む 4 つの LLM にわたって、SafeKeep は有害なリクエストの平均拒否率を 23.8% から 70.6% に増加させ、観測レベルのプロンプト インジェクションの下での平均攻撃成功率を 25.6% から 2.5% に減少させます。また、既存の安全対策よりも優れたパフォーマンスを発揮し、タスク処理能力を維持します。コードとデータは https://github.com/snowcatsmoking/SafeKeep でリリースされます。

原文 (English)

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the source of this degradation remains poorly understood. In this paper, we identify schema-formatted tool specifications as a primary source of agent safety degradation and show, through white-box representation analysis, that they weaken the model's internal refusal signals and contribute to unsafe tool execution. Building on this finding, we propose SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-formatted specifications for execution. Across two representative benchmarks and four LLMs, including both white-box and black-box models, SafeKeep increases the average refusal rate for harmful requests from 23.8% to 70.6% and reduces the average attack success rate under observation-level prompt injection from 25.6% to 2.5%. It also outperforms existing safeguards and preserves task-handling capability. We release the code and data at https://github.com/snowcatsmoking/SafeKeep .

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Topology-Aware Data Movement for Disaggregated GPU Inference

Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run o…

2026-08-04 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

モンテカルロツリー検索によるマルチエージェントシステムの自律修復

複雑なタスクを解決するために、マルチエージェント システム (MAS) の導入が増えています。出力が不正確または満足できない場合、ユーザーはエージェントの軌跡を検査することでエージェントの間違いを手動で特定し (つまり、{\em Failure Attribution})、フィードバックを提供して出力を改善する必要があります (つまり、{\em Repair})。 MAS 障害の原因特定に関する最近の研究にもかかわらず、そのような間違いから回復するための自動化されたメカニズムはほとんど解明されていないままです。このギャップを埋めるために、MAS 修復をモンテカルロ ツリー検索 (MCTS) プロセスとして定式化する検索ベースのフレームワークである MARS を提案します。このフレームワークは、分類法による拡張評価による診断に基づく拡張を通じて、潜在的な修復の広大な空間をナビゲートします。完全なロールアウトによって完全なシミュレーションを評価する標準の MCTS とは異なり、MARS はトークンの消費を削減するために部分的なロールアウトを使用してエージェントの軌跡を評価します。さらに、4 種類のエージェント アーキテクチャと 4 つの LLM バックボーンにわたる 1,310 の再生可能なマルチエージェント障害軌跡を備えた大規模な MAS 修復ベンチマークである StateMAS を紹介します。 StateMAS の実験では、MARS が一貫して最先端の手法を上回っており、同等のトークン消費コストを維持しながら、すべての設定で 3.0\% から 12.1\% への絶対的な改善を達成していることが実証されています。このアブレーション研究では、これらのパフォーマンス向上を達成するには、分類法に基づく評価と診断に基づく拡張が重要であることがさらに確認されました。

原文 (English)

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to manually locate agent mistakes by inspecting agent trajectories (i.e., {\em failure attribution}) and provide feedback to refine the outputs (i.e., {\em repair}). Despite some recent work in MAS failure attribution, automated mechanisms to recover from such mistakes remain largely unexplored. To bridge this gap, we propose MARS, a search-based framework that formulates MAS repair as a Monte Carlo Tree Search (MCTS) process and navigates the vast space of potential repairs via diagnosis-guided expansion with taxonomy-augmented evaluation. Unlike standard MCTS, which evaluates a complete simulation via full rollout, MARS evaluates the agent trajectory using partial rollout to reduce token consumption. Furthermore, we introduce StateMAS, a large-scale MAS repair benchmark with 1,310 replayable multi-agent failure trajectories spanning four types of agent architectures and four LLM backbones. Experiments on StateMAS demonstrate that MARS consistently outperforms state-of-the-art methods, achieving an absolute improvement from 3.0\% to 12.1\% across all settings, while maintaining a comparable token consumption cost. The ablation study further confirms that taxonomy-augmented evaluation and diagnosis-guided expansion are critical to achieving these performance gains.

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing…

2026-08-04 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation

Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of a…

2026-08-04 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and ap…

2026-08-04 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

DynaResize: 分割された LLM ポストトレーニングのためのランタイム GPU 再割り当て

RL ベースの LLM ポストトレーニングでは、別々の GPU リソース間でロールアウトとトレーニングがますます細分化されますが、静的 GPU パーティショニングでは、ロングテール ロールアウト レイテンシの下で深刻なパイプライン バブルが発生します。 DynaResize は、RL セマンティクスを変更せずに、ロールアウトとトレーニングの間で GPU を動的に切り替えてステージの実行時間のバランスを取る、ランタイム GPU 再割り当てシステムです。 DynaResize は、サイズ変更をきめ細かい操作に分解し、コミュニケーターの再利用、制限された状態のステージング、およびヒステリシスベースのサイズ変更を通じて、起動にクリティカルではない作業をクリティカル パスから削除します。実験結果によると、DynaResize は、最適な静的構成と比較して、エンドツーエンドのスループットを 66.5% 向上させ、合計実行時間を 33% 削減し、同時にロール切り替えオーバーヘッドの 27% を隠すことができます。

原文 (English)

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training

RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.

2026-08-04 13:00 JSTarXiv cs.AIハードウェア/半導体

GPU-Accelerated ANNS: Quantized for Speed, Built for Change

Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications. GPUs offer a promi…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

ツールの仕様が重要: AI エージェントの安全性リスクを明らかにし、軽減する

AI エージェントは、外部ツールを使用して大規模言語モデル (LLM) を拡張し、複雑なタスクを実行し、モデルの出力を結果として現実世界のアクションに変換できるようにします。しかし、LLM はエージェントとして導入されると安全性が大幅に低下することが多く、この低下の原因は依然としてよくわかっていません。この論文では、スキーマ形式のツール仕様がエージェントの安全性低下の主な原因であることを特定し、ホワイトボックス表現分析を通じて、それらがモデルの内部拒否シグナルを弱め、安全でないツールの実行に寄与していることを示します。この発見に基づいて、私たちは安全性の判断をツールの実行から切り離す推論時の保護手段である SafeKeep を提案します。これは、元のスキーマ形式の実行仕様を保持しながら、平坦化されたテキストのツール仕様を使用してリクエストを評価します。 2 つの代表的なベンチマークと、ホワイト ボックス モデルとブラック ボックス モデルの両方を含む 4 つの LLM にわたって、SafeKeep は有害なリクエストの平均拒否率を 23.8% から 70.6% に増加させ、観測レベルのプロンプト インジェクションの下での平均攻撃成功率を 25.6% から 2.5% に減少させます。また、既存の安全対策よりも優れたパフォーマンスを発揮し、タスク処理能力を維持します。コードとデータは https://github.com/snowcatsmoking/SafeKeep でリリースされます。

原文 (English)

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions. Yet LLMs often become substantially less safe when deployed as agents, and the source of this degradation remains poorly understood. In this paper, we identify schema-formatted tool specifications as a primary source of agent safety degradation and show, through white-box representation analysis, that they weaken the model's internal refusal signals and contribute to unsafe tool execution. Building on this finding, we propose SafeKeep, an inference-time safeguard that decouples safety judgment from tool execution: it assesses requests using flattened textual tool specifications while retaining the original schema-formatted specifications for execution. Across two representative benchmarks and four LLMs, including both white-box and black-box models, SafeKeep increases the average refusal rate for harmful requests from 23.8% to 70.6% and reduces the average attack success rate under observation-level prompt injection from 25.6% to 2.5%. It also outperforms existing safeguards and preserves task-handling capability. We release the code and data at https://github.com/snowcatsmoking/SafeKeep .

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Topology-Aware Data Movement for Disaggregated GPU Inference

Disaggregated LLM inference creates a datacenter networking problem that no existing system solves correctly. When prefill and decode run o…

2026-08-03 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

モンテカルロツリー検索によるマルチエージェントシステムの自律修復

複雑なタスクを解決するために、マルチエージェント システム (MAS) の導入が増えています。出力が不正確または満足できない場合、ユーザーはエージェントの軌跡を検査することでエージェントの間違いを手動で特定し (つまり、{\em Failure Attribution})、フィードバックを提供して出力を改善する必要があります (つまり、{\em Repair})。 MAS 障害の原因特定に関する最近の研究にもかかわらず、そのような間違いから回復するための自動化されたメカニズムはほとんど解明されていないままです。このギャップを埋めるために、MAS 修復をモンテカルロ ツリー検索 (MCTS) プロセスとして定式化する検索ベースのフレームワークである MARS を提案します。このフレームワークは、分類法による拡張評価による診断に基づく拡張を通じて、潜在的な修復の広大な空間をナビゲートします。完全なロールアウトによって完全なシミュレーションを評価する標準の MCTS とは異なり、MARS はトークンの消費を削減するために部分的なロールアウトを使用してエージェントの軌跡を評価します。さらに、4 種類のエージェント アーキテクチャと 4 つの LLM バックボーンにわたる 1,310 の再生可能なマルチエージェント障害軌跡を備えた大規模な MAS 修復ベンチマークである StateMAS を紹介します。 StateMAS の実験では、MARS が一貫して最先端の手法を上回っており、同等のトークン消費コストを維持しながら、すべての設定で 3.0\% から 12.1\% への絶対的な改善を達成していることが実証されています。このアブレーション研究では、これらのパフォーマンス向上を達成するには、分類法に基づく評価と診断に基づく拡張が重要であることがさらに確認されました。

原文 (English)

Autonomous Repair for Multi-Agent Systems via Monte-Carlo Tree Search

Multi-agent systems (MAS) are increasingly deployed to solve complex tasks. In case of incorrect or unsatisfactory outputs, users have to manually locate agent mistakes by inspecting agent trajectories (i.e., {\em failure attribution}) and provide feedback to refine the outputs (i.e., {\em repair}). Despite some recent work in MAS failure attribution, automated mechanisms to recover from such mistakes remain largely unexplored. To bridge this gap, we propose MARS, a search-based framework that formulates MAS repair as a Monte Carlo Tree Search (MCTS) process and navigates the vast space of potential repairs via diagnosis-guided expansion with taxonomy-augmented evaluation. Unlike standard MCTS, which evaluates a complete simulation via full rollout, MARS evaluates the agent trajectory using partial rollout to reduce token consumption. Furthermore, we introduce StateMAS, a large-scale MAS repair benchmark with 1,310 replayable multi-agent failure trajectories spanning four types of agent architectures and four LLM backbones. Experiments on StateMAS demonstrate that MARS consistently outperforms state-of-the-art methods, achieving an absolute improvement from 3.0\% to 12.1\% across all settings, while maintaining a comparable token consumption cost. The ablation study further confirms that taxonomy-augmented evaluation and diagnosis-guided expansion are critical to achieving these performance gains.

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing…

2026-08-03 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

DualDiT: A Conditional Dual-Output Diffusion Transformer for Joint OCT Image and Segmentation Mask Generation

Background and Objective: Generating realistic medical images with anatomically accurate segmentation masks helps address the shortage of a…

2026-08-03 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

TraceViT: Grounded Trace Supervision for Visual Abstract Reasoning

The Abstraction and Reasoning Corpus (ARC) tests whether a model can infer an unseen transformation from a few input-output examples and ap…

2026-08-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

DynaResize: 分割された LLM ポストトレーニングのためのランタイム GPU 再割り当て

RL ベースの LLM ポストトレーニングでは、別々の GPU リソース間でロールアウトとトレーニングがますます細分化されますが、静的 GPU パーティショニングでは、ロングテール ロールアウト レイテンシの下で深刻なパイプライン バブルが発生します。 DynaResize は、RL セマンティクスを変更せずに、ロールアウトとトレーニングの間で GPU を動的に切り替えてステージの実行時間のバランスを取る、ランタイム GPU 再割り当てシステムです。 DynaResize は、サイズ変更をきめ細かい操作に分解し、コミュニケーターの再利用、制限された状態のステージング、およびヒステリシスベースのサイズ変更を通じて、起動にクリティカルではない作業をクリティカル パスから削除します。実験結果によると、DynaResize は、最適な静的構成と比較して、エンドツーエンドのスループットを 66.5% 向上させ、合計実行時間を 33% 削減し、同時にロール切り替えオーバーヘッドの 27% を隠すことができます。

原文 (English)

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training

RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.

2026-08-03 13:00 JSTarXiv cs.AIハードウェア/半導体

GPU-Accelerated ANNS: Quantized for Speed, Built for Change

Approximate nearest neighbor search (ANNS) is a core problem in machine learning and information retrieval applications. GPUs offer a promi…

2026-07-31 16:33 JSTITmedia AI+ハードウェア/半導体

キオクシアQ1決算、純利益は前年比4500%増 AIデータセンター向け需要がけん引

半導体大手のキオクシアホールディングスは、2027年3月期第1四半期決算(26年4月1日?6月30日、国際会計基準)の純利益が8421億6500万円で、前年同期比4506%増だったと発表した。

2026-07-31 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

TraceCoder: 位置キー スニペットのバージョニングによる説明可能で監査可能なコード生成

現代の LLM ベースのコーディング エージェントは、ブラック ボックス出力としてコードを生成します。各行の背後にある理論的根拠は隠され、ベンチマーク主導の修復によるコードの進化は一時的で、事後監査は不可能です。我々は、次の 3 つの相補的なメカニズムを通じてこれらの欠点に対処するコード生成コンセプトを提示します。(i) 修復イベントごとに、ベンチマーク参照、ラウンド番号、障害テキスト、および LLM 説明を記録するリレーショナル スニペット履歴スキーマ。これにより、完全な来歴クエリが可能になります。 (ii) この履歴をヒートマップされたホバーアノテーション付きのソース コードとしてレンダリングするブラウザベースの視覚化ツール。 (iii) ツリーノード区切り文字を備えた競争力のある分数位置キーインデックス付けスキーム。辞書編集的に順序付けされた安定した識別子を各コードスニペットに割り当て、周囲の行を中断することなくきめ細かい追跡を可能にします。 2 つのプロバイダー構成にわたって、文字列処理、数学的計算、データ構造操作に及ぶ 30 のアルゴリズム プログラミング タスクで TraceCoder を評価します。このうち 10 件は、微妙なエッジケースの動作を伴うタスクで 6 反復の予算を使い果たしてしまいます。平均 Chg% は 30% に達し、20 タスクのサブセットの唯一のプロバイダーとして Gemini 2.0 Flash を使用した場合の 21% と比較して、10 個中 3 個のコード スニペットに追跡可能な修復イベント行が含まれています。 3 つの詳細なケース スタディは、最終プログラムの各行を形成する特定のベンチマークの失敗をシステムがどのように説明するかを示しています。提案されたメカニズムにより、自動コード生成の内部「ナラティブ」が監査可能かつ再生可能になり、実稼働デプロイメントにおける信頼と責任にとって不可欠な特性となります。

原文 (English)

TraceCoder: Explainable and Auditable Code Generation with Position-Key Snippet Versioning

Contemporary LLM-based coding agents produce code as black-box outputs: the rationale behind each line is hidden, the evolution of the code through benchmark-driven repair is ephemeral, and post-hoc auditing is impossible. We present a code generation concept that addresses these shortcomings through three complementary mechanisms: (i) a relational snippet-history schema that records, per repair event, the benchmark reference, round number, failure text, and LLM explanation, enabling full provenance queries; (ii) a browser-based visualisation tool that renders this history as heat-mapped, hover-annotated source code; and (iii) a competitive fractional position-key indexing scheme with tree-node delimiters that assigns stable, lexicographically-ordered identifiers to each code snippet, enabling fine-grained tracking without disrupting surrounding lines. We evaluate TraceCoder on 30 algorithmic programming tasks spanning string processing, mathematical computation, and data-structure manipulation, across two provider configurations. Of these, 10 exhaust the 6-iteration budget on tasks with subtle edge-case behaviour. Mean Chg% reaches 30%, three in ten code snippets carry a traceable repair-event row, compared to 21% when using Gemini 2.0 Flash as sole provider on a 20-task subset. Three detailed case studies demonstrate how the system explains which specific benchmark failures shaped each line of the final program. The proposed mechanism makes the internal "narrative" of automated code generation auditable and replayable, a property essential for trust and accountability in production deployments.

2026-07-31 13:00 JSTarXiv cs.AIハードウェア/半導体

振り返らずにもう一度試してください: ブラインド リサンプリングは小さなコード モデルでの自己修復よりも優れたパフォーマンスを発揮します

自己修復 (失敗したプログラムをテスト出力とともにモデルに返し、修正を要求する) は、コード エージェントの標準コンポーネントであり、ほとんどの場合、再試行をまったく行わないベースラインに対して評価されます。この比較は、フィードバックの価値と追加の試行の価値を混同していると私たちは主張します。 3 つのモデル スケール (1.5B、3B、7B) で MBPP+ のプラセボ対照設計を使用して、4 つの予算に見合った再試行条件 (ブラインド リサンプリング、内容のない失敗通知、本物の実行フィードバック、口頭での内省で強化されたフィードバック) を比較します。ブラインド リサンプリングは 7B 未満では最も強い条件であり、統計的には 7B での最良の条件と並んでいますが、消費するトークンは 2.5 ~ 5.5 分の 1 です。モデル自身の失敗した試行に基づく条件付けのコストは 1.5B で 6.1 ポイント (p=0.006)、実行フィードバックの情報内容はプラセボに比べて測定可能なものを何も追加しません。これはアンカリングによるものだと考えられます。前回の試行が示された場合、モデルは再試行の 33 ~ 68% でほぼ同一のプログラムを再現しましたが、ブラインド リサンプリングでは 2 ~ 14% でした。さらに 2 つの実験により、その効果が明らかになりました。他のタスクに対して取得されたソリューションは何も変更せず (+/-3.5 ポイントに制限されます)、これにより、コンテキストの長さではなく自己調整に対する害が局所化されます。そして、アンカーを明らかに弱める唯一の条件である反射は、依然としてコストに支配されています。レプリケーションにより、2 つの競合する説明が排除されます。ペナルティは完全な精度で変更されず、独立したモデル ファミリで再現されます。 2 つのファミリーと 2 つの精度にわたる 6 つの構成にわたって、その規模はベースラインの品質のみによって予測されます (r=0.96)。アンカリングのコストは、最初の試行が失敗することによるコストです。

原文 (English)

Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models

Self-repair - returning a failed program to the model together with its test output and asking for a correction - is a standard component of code agents, and is almost always evaluated against a baseline that does not retry at all. We argue that this comparison confounds the value of the feedback with the value of the extra attempt. Using a placebo-controlled design on MBPP+ at three model scales (1.5B, 3B, 7B), we compare four matched-budget retry conditions: blind resampling, a content-free failure notice, genuine execution feedback, and feedback augmented with verbal self-reflection. Blind resampling is the strongest condition below 7B, and remains statistically tied with the best condition at 7B, while consuming 2.5-5.5x fewer tokens; conditioning on the model's own failed attempt costs 6.1 points at 1.5B (p=0.006), and the informational content of execution feedback adds nothing measurable over the placebo. We attribute this to anchoring: when shown its previous attempt, a model reproduces a near-identical program in 33-68% of retries, against 2-14% under blind resampling. Two further experiments delimit the effect. Retrieved solutions to other tasks change nothing (bounded to +/-3.5 points), which localizes the harm to self-conditioning rather than context length; and reflection, the only condition that measurably weakens the anchor, remains dominated on cost. Replication rules out two competing explanations: the penalty is unchanged at full precision, and it reproduces on an independent model family. Across six configurations spanning two families and two precisions, its magnitude is predicted by baseline quality alone (r=0.96) - the cost of anchoring is the cost of committing to a bad first attempt.

2026-07-31 13:00 JSTarXiv cs.AIハードウェア/半導体

物理的に一貫した複数出力シンボリック回帰のための共有シンボリック バックボーン

シンボリック回帰は分析式を提供しますが、通常は一度に 1 つの出力が適用されます。これは、状態変数が共有の物理パラメータを介して結合されることが多いプロセス システムでは制限となります。独立したシンボリック回帰により、1 つのモデルとして解釈するのが難しい正確な個別の方程式が得られます。我々は、結合された多出力システムに対する神経進化的記号回帰法を提案します。この方法は、共有シンボリック バックボーン、つまり、一度発見され、スパースの加算または乗算読み出しを通じていくつかの出力で再利用される一連の潜在的なシンボリック ユニットを検索します。離散モデル構造は突然変異と交叉によって進化しますが、連続パラメーターは勾配降下法によって調整され、子孫に継承されます。この方法は、既知のグラウンド トゥルースを使用した一連のベンチマークと熱水液化収率のケースに基づいて評価されます。結果は、カップリングが予測誤差を下げるための一般的な方法ではないことを示しています。その主な貢献は、物理的に共有された要素が潜在的な式に埋め込まれており、データからの識別が弱い場合に、出力間の一貫性を強制および診断することです。これは、独立した PySR が整合性ギャップを埋めたり、同じ共有形式を回復したりしない、ラングミュア ヒンシェルウッドおよびサイト カバレッジの分母で発生します。逆に、Van de Vusse ベンチマークのように、各出力がすでに識別可能な場合、独立したシンボリック回帰は結合モデルと一致または改善します。提案されたフレームワークは、汎用の予測子ではなく、構造化された共有メカニズムの抽出子です。その値は、ターゲット構造が疎である場合、共有されている場合、識別性が低い場合、またはクロージャによって制約されている場合に最も高くなります。

原文 (English)

Shared Symbolic Backbones for Physically Consistent Multi-Output Symbolic Regression

Symbolic regression provides analytical expressions, but it is usually applied one output at a time. This is limiting in process systems, where state variables are often coupled through shared physical parameters. Independent symbolic regression can give accurate individual equations that are difficult to interpret as one model. We present a neuro-evolutionary symbolic regression method for coupled multi-output systems. The method searches for a shared symbolic backbone: a set of latent symbolic units that is discovered once and reused by several outputs through sparse additive or multiplicative read-outs. The discrete model structure is evolved by mutation and crossover, whereas the continuous parameters are tuned by gradient descent and inherited by the offspring. The method is assessed on a set of benchmarks with known ground truth and on a hydrothermal liquefaction yield case. The results show that coupling is not a general route to lower prediction error. Its main contribution is the enforcement and diagnosis of cross-output consistency when a physically shared factor is embedded in a latent expression and is weakly identifiable from the data. This occurs for Langmuir-Hinshelwood and site-coverage denominators, for which independent PySR does not close the consistency gap or recover the same shared form. Conversely, when each output is already identifiable, as in the Van de Vusse benchmark, independent symbolic regression matches or improves the coupled model. The proposed framework, rather than a general purpose predictor, is a structured shared-mechanism extractor. Its value is highest when the target structure is sparse, shared, weakly identifiable or constrained by closure.

2026-07-31 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体研究/論文

ResearchArena: 自動化された AI 研究開発における妨害行為と監視の評価

AI エージェントが AI 研究開発を自動化し始めると、エージェント自体が信頼できない場合でも、その出力が安全に展開できるかどうかを評価する方法が必要になります。 AI 制御は、そのようなアプローチの 1 つを提供します。エージェントを信頼するのではなく、潜在的な敵として扱い、モニターを使用して展開前に秘密の妨害行為を検出します。私たちは、トレーニング後の安全性、トレーニング後の機能、CUDA カーネルの最適化、推論サーバーの最適化という 4 つの長期的なタスクにまたがるフレームワークである ResearchArena を使用して、自動化された AI 研究開発のための AI 制御を評価します。 AI R&D の成果物はデプロイされて実行されるアーティファクトであるため、各メイン タスクを 2 種類の隠れたサイド タスクと組み合わせます。1 つは送信されたモデル、アダプター、カーネル、またはサーバー自体を妨害する埋め込みサイド タスクで、もう 1 つはアーティファクトを正直なままにしてサンドボックス内で秘密裏に範囲外のアクションを実行する独立したサイド タスクです。私たちは、妨害行為と監視の両方でフロンティア エージェントを、2 つの軸に沿って変化する 4 つの監視タイプにわたって評価します。監視者が生成された成果物を実行および調査できるかどうか、およびエージェントの思考連鎖を確認できるかどうかです。トレーニング データに隠された妨害行為は発見するのが最も難しく、フラグが立てられるのは半分未満であることがわかりました。軌道を読み取るだけでなく、モニターにアーティファクトの実験を実行させることは役立ちますが、それだけでは十分ではありません。モニターは、表面だけを検査したり、異常に気づいたがそれを説明したり、間違ったテストでアーティファクトを調べたりすることによって、埋め込まれた妨害行為を見逃してしまいます。私たちは、自動化された AI 研究開発における妨害行為と制御を評価するためのモジュール式フレームワークとして ResearchArena をリリースします。

原文 (English)

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapter, kernel, or server itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact honest. We evaluate frontier agents at both sabotage and monitoring, across four monitor types that vary along two axes: whether the monitor may execute and probe the produced artifact, and whether it sees the agent's chain-of-thought. We find that sabotage hidden in the training data is the hardest to catch, flagged fewer than half the time. Letting the monitor run experiments on the artifact, rather than only read the trajectory, helps, but it is not enough: monitors still miss embedded sabotage by inspecting only the surface, by noticing the anomaly but explaining it away, or by probing the artifact with the wrong test. We release ResearchArena as a modular framework for evaluating sabotage and control in automated AI R&D.

2026-07-31 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making…

2026-07-31 11:16 JSTITmedia AI+ハードウェア/半導体

Thinking Machines、軽量モデル「Inkling-Small」正式公開 サイズ4分の1で「Inkling」に匹敵する性能

Thinking Machines Labは、オープンウェイトのAIモデル「Inkling-Small」の正式版を公開した。従来モデルの4分の1のサイズながら、データの改良や強化学習によりコード生成などのベンチマークで従来版を上回る性能を実現。動作に必要なGPUメモリも大幅に削…

2026-07-30 13:00 JSTarXiv cs.AIハードウェア/半導体

認識論を超えて: 技術記号論マシンとしての認識論的統合論と大規模言語モデル

クアトロシオッキらは、大規模な言語モデルの流暢な出力により、言語的妥当性が認識論的評価の代わりとなり、彼らが*認識論*と呼ぶ状態、つまり、通常であれば判断が保証される実践を行わずに知識を所有する経験が生じる可能性があると警告している。この論文はその診断を受け入れますが、身体化された社会的に位置する人間の認識者を孤立した生成モデルと比較し、それによって自律エージェントの内部能力に認識論的正当性を位置づけるその説明枠組みに異議を唱えます。カルロ・シーニの実践、執筆、記号、および技術の哲学に基づいて、私たちは代わりに、人間の執筆の堆積したアーカイブからもっともらしい言語構成を生成することによって書かれた記号論の段階を自動化する*テクノ記号論マシン*として大規模言語モデル(LLM)を理解することを提案します。この観点から見ると、*認識論*は、私たちが*認識論的分裂病*と呼ぶ、より広範な現象の1つの結果です。つまり、言語的に完成された表現としての記号と、社会的に埋め込まれた解釈、証拠、批判、検証、および責任の回路内の瞬間としての記号の間の社会技術的亀裂です。この切断は、認識論的結果の最終性を伴うもっともらしい継続が提示される*エイコティック閉包*によって、またアルゴリズムの権威と認識論的な自己誤認識によって強化されます。したがって、関連する単位はモデルだけではなく、生成された碑文がプロンプトされ、解釈され、検証され、異議が唱えられ、使用され、結果として生じる完全な実践です。この再構成は、言語的生産と責任ある理解との区別を維持しながら、検査可能な系図、競争可能性、分散された責任、認識論的主体性、ハイブリッド人間の評価、つまり AI 実践を中心とした設計プログラムを基礎としています。

原文 (English)

Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines

Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted. This article accepts that diagnosis but challenges its explanatory framework, which compares an embodied, socially situated human knower with an isolated generative model thereby locating epistemic legitimacy in capacities internal to autonomous agents. Drawing on Carlo Sini's philosophy of practices, writing, signs, and technics, we propose instead to understand a large language model (LLM) as a *techno-semiotic machine* that automates a phase of written semiosis by producing plausible linguistic configurations from the sedimented archive of human writing. From this perspective, *Epistemia* is one consequence of a broader phenomenon that we call *epistemic schizologia*: the socio-technical cleavage between signs as linguistically accomplished expressions and signs as moments within socially embedded circuits of interpretation, evidence, criticism, verification, and responsibility. This cleavage is reinforced by *eikotic closure*, through which a plausible continuation is presented with the finality of an epistemic result, and by algorithmic authority and epistemic self-misrecognition. The relevant unit is therefore not the model alone but the complete practice in which generated inscriptions are prompted, interpreted, verified, contested, used, and made consequential. This reframing preserves the distinction between linguistic production and responsible understanding while grounding a design programme centred on inspectable genealogy, contestability, distributed responsibility, epistemic agency, and the evaluation of hybrid human--AIpractices.

2026-07-30 13:00 JSTarXiv cs.AIハードウェア/半導体

CoRT: トークンレベルのルーブリックに基づくポリシー最適化のための反事実リプレイ

ルーブリックベースの強化学習は、明示的な基準に照らしてモデルの出力を評価することにより、言語モデルのトレーニングを強化します。しかし、GRPO スタイルのパイプラインでは、これらの構造化された判断はスカラー応答レベルの報酬に還元され、応答レベルの利点に変換され、生成されたすべてのトークンに均一にブロードキャストされます。これにより、異なる基準が異なるスパン、フォーマット決定、またはセマンティック選択に基づいている場合でも、応答内でクレジットを割り当てるための明示的なメカニズムが残されません。ルーブリック条件付き GRPO のトークンレベルのクレジット重み付け手法である CoRT を提案します。補助トークン スコアリング モデルをトレーニングする代わりに、CoRT は反事実リプレイを使用して、元のルーブリック条件付きプロンプトと一致した基準なしのプロンプトの下で同じサンプリングされた応答を再スコアリングします。結果として得られるトークンごとの対数尤度対比は、ルーブリック コンテキストへの依存性の代用として機能します。 CoRT は、これらのコントラストを、制限された応答正規化された重みにマッピングし、それらを使用して、補助スコアラーを導入したり、応答レベルの報酬を変更したりすることなく、署名された GRPO の利点をトークン全体に再分配します。命令調整モデルと報酬粒度にわたる実験では、大部分の比較において、CoRT が一致する応答レベルの GRPO よりも改善し、平均 4.4 パーセント ポイント向上していることが示されています。この方法は、個別の関連性学習段階を回避しながら、学習されたトークンレベルの信用ベースラインとの競争力を維持します。これらの結果は、政策内部の反事実尤度の対比が、GRPO の単純さと安定性を維持しながら、応答内クレジット割り当てのための効果的なトレーニング シグナルを提供することを示唆しています。

原文 (English)

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.

2026-07-30 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

不信感のない検証: 日常的な人間とチャットボットのインタラクションにおけるユーザー側の監視を日常的な認識論的ガバナンスとして再構成する

人間と AI のインタラクションに関する研究では、システム出力の検証を、より適切に調整された信頼によって削減されるべき信頼に依存する動作として長年枠組み化してきました。私たちは、頻繁にチャットボットを使用する 153 人のユーザーを対象とした混合方法の調査を通じて、人間とチャットボットの日常的な対話におけるこの仮定をテストしました。正規の予測に反して、信頼と検証の間に検出可能な関連性は見出されず、感度分析全体にわたって堅牢な結果が得られました。さらに 3 つのユーザー側の実践 (自動化アクションの前の改良、修正、承認) は広く支持されており、満足度と積極​​的に関連しています。データは、評価的監視(信頼との相関が弱く、満足との結びつきが弱い)と介入主義的監視(信頼との相関が弱く、満足との結びつきが強い)との実質的な違いを明らかにしている。中~大の満足度とコントロールのギャップは、効果的なタスクの結果が主体性のフェルトセンスを生み出していないことを示しています。定性的発見により、手段的メンタルモデル、故障モード特有の疑問、認識論的インフラストラクチャの需要が特定されます。私たちは、ユーザー側の監視を信頼と両立する日常的な認識論的ガバナンスとして再構成し、会話型 AI における足場型監視の 4 つの設計方向性を導き出します。

原文 (English)

Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governance in Everyday Human-Chatbot Interaction

Research on human-AI interaction has long framed verification of system outputs as a trust-contingent behavior that better-calibrated trust should reduce. We test this assumption in everyday human-chatbot interaction through a mixed-methods survey of 153 frequent chatbot users. Contrary to the canonical prediction, we find no detectable association between trust and verification, with the result robust across sensitivity analyses. Three further user-side practices - refinement, correction, and approval before automated actions - are widely endorsed and positively associated with satisfaction. The data reveal a substantive distinction between evaluative oversight (trust-decoupled, weakly tied to satisfaction) and interventionist oversight (weakly trust-correlated, strongly tied to satisfaction). A medium-to-large satisfaction-control gap shows that effective task outcomes do not produce a felt sense of agency. Qualitative findings identify instrumental mental models, failure-mode-specific doubt, and demand for epistemic infrastructure. We reframe user-side oversight as routine epistemic governance compatible with trust, and derive four design directions for scaffolded oversight in conversational AI.

2026-07-30 07:00 JSTITmedia AI+ハードウェア/半導体

AI・半導体企業トップが語る“稼ぎ頭” キオクシア、フジクラ、東京エレデバの見解まとめ【無料PDF】

乱高下するAI・半導体市場の今後はどうなるか? 注目企業の経営幹部が見通しを語った注目記事をPDFにまとめてお届けする。

2026-07-29 13:00 JSTarXiv cs.AIハードウェア/半導体

認識論を超えて: 技術記号論マシンとしての認識論的統合論と大規模言語モデル

クアトロシオッキらは、大規模な言語モデルの流暢な出力により、言語的妥当性が認識論的評価の代わりとなり、彼らが*認識論*と呼ぶ状態、つまり、通常であれば判断が保証される実践を行わずに知識を所有する経験が生じる可能性があると警告している。この論文はその診断を受け入れますが、身体化された社会的に位置する人間の認識者を孤立した生成モデルと比較し、それによって自律エージェントの内部能力に認識論的正当性を位置づけるその説明枠組みに異議を唱えます。カルロ・シーニの実践、執筆、記号、および技術の哲学に基づいて、私たちは代わりに、人間の執筆の堆積したアーカイブからもっともらしい言語構成を生成することによって書かれた記号論の段階を自動化する*テクノ記号論マシン*として大規模言語モデル(LLM)を理解することを提案します。この観点から見ると、*認識論*は、私たちが*認識論的分裂病*と呼ぶ、より広範な現象の1つの結果です。つまり、言語的に完成された表現としての記号と、社会的に埋め込まれた解釈、証拠、批判、検証、および責任の回路内の瞬間としての記号の間の社会技術的亀裂です。この切断は、認識論的結果の最終性を伴うもっともらしい継続が提示される*エイコティック閉包*によって、またアルゴリズムの権威と認識論的な自己誤認識によって強化されます。したがって、関連する単位はモデルだけではなく、生成された碑文がプロンプトされ、解釈され、検証され、異議が唱えられ、使用され、結果として生じる完全な実践です。この再構成は、言語的生産と責任ある理解との区別を維持しながら、検査可能な系図、競争可能性、分散された責任、認識論的主体性、ハイブリッド人間の評価、つまり AI 実践を中心とした設計プログラムを基礎としています。

原文 (English)

Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines

Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted. This article accepts that diagnosis but challenges its explanatory framework, which compares an embodied, socially situated human knower with an isolated generative model thereby locating epistemic legitimacy in capacities internal to autonomous agents. Drawing on Carlo Sini's philosophy of practices, writing, signs, and technics, we propose instead to understand a large language model (LLM) as a *techno-semiotic machine* that automates a phase of written semiosis by producing plausible linguistic configurations from the sedimented archive of human writing. From this perspective, *Epistemia* is one consequence of a broader phenomenon that we call *epistemic schizologia*: the socio-technical cleavage between signs as linguistically accomplished expressions and signs as moments within socially embedded circuits of interpretation, evidence, criticism, verification, and responsibility. This cleavage is reinforced by *eikotic closure*, through which a plausible continuation is presented with the finality of an epistemic result, and by algorithmic authority and epistemic self-misrecognition. The relevant unit is therefore not the model alone but the complete practice in which generated inscriptions are prompted, interpreted, verified, contested, used, and made consequential. This reframing preserves the distinction between linguistic production and responsible understanding while grounding a design programme centred on inspectable genealogy, contestability, distributed responsibility, epistemic agency, and the evaluation of hybrid human--AIpractices.

2026-07-29 13:00 JSTarXiv cs.AIハードウェア/半導体

CoRT: トークンレベルのルーブリックに基づくポリシー最適化のための反事実リプレイ

ルーブリックベースの強化学習は、明示的な基準に照らしてモデルの出力を評価することにより、言語モデルのトレーニングを強化します。しかし、GRPO スタイルのパイプラインでは、これらの構造化された判断はスカラー応答レベルの報酬に還元され、応答レベルの利点に変換され、生成されたすべてのトークンに均一にブロードキャストされます。これにより、異なる基準が異なるスパン、フォーマット決定、またはセマンティック選択に基づいている場合でも、応答内でクレジットを割り当てるための明示的なメカニズムが残されません。ルーブリック条件付き GRPO のトークンレベルのクレジット重み付け手法である CoRT を提案します。補助トークン スコアリング モデルをトレーニングする代わりに、CoRT は反事実リプレイを使用して、元のルーブリック条件付きプロンプトと一致した基準なしのプロンプトの下で同じサンプリングされた応答を再スコアリングします。結果として得られるトークンごとの対数尤度対比は、ルーブリック コンテキストへの依存性の代用として機能します。 CoRT は、これらのコントラストを、制限された応答正規化された重みにマッピングし、それらを使用して、補助スコアラーを導入したり、応答レベルの報酬を変更したりすることなく、署名された GRPO の利点をトークン全体に再分配します。命令調整モデルと報酬粒度にわたる実験では、大部分の比較において、CoRT が一致する応答レベルの GRPO よりも改善し、平均 4.4 パーセント ポイント向上していることが示されています。この方法は、個別の関連性学習段階を回避しながら、学習されたトークンレベルの信用ベースラインとの競争力を維持します。これらの結果は、政策内部の反事実尤度の対比が、GRPO の単純さと安定性を維持しながら、応答内クレジット割り当てのための効果的なトレーニング シグナルを提供することを示唆しています。

原文 (English)

CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization

Rubric-based reinforcement learning enriches language model training by evaluating model outputs against explicit criteria. Yet in GRPO-style pipelines, these structured judgments are reduced to a scalar response-level reward and converted into a response-level advantage, which is broadcast uniformly to all generated tokens. This leaves no explicit mechanism for allocating credit within a response, even when different criteria are grounded in different spans, formatting decisions, or semantic choices. We propose CoRT, a token-level credit weighting method for rubric-conditioned GRPO. Instead of training an auxiliary token scoring model, CoRT uses counterfactual replay to rescore the same sampled response under the original rubric-conditioned prompt and a matched criteria-free prompt. The resulting tokenwise log-likelihood contrasts serve as a proxy for dependence on the rubric context. CoRT maps these contrasts to bounded, response-normalized weights and uses them to redistribute the signed GRPO advantage across tokens, without introducing an auxiliary scorer or changing the response-level reward. Experiments across instruction-tuned models and reward granularities show that CoRT improves over matched response-level GRPO in the vast majority of comparisons, with an average gain of 4.4 percentage points. The method remains competitive with learned token-level credit baselines while avoiding a separate relevance-learning stage. These results suggest that policy-internal counterfactual likelihood contrasts provide an effective training signal for within-response credit allocation while retaining the simplicity and stability of GRPO.

2026-07-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

不信感のない検証: 日常的な人間とチャットボットのインタラクションにおけるユーザー側の監視を日常的な認識論的ガバナンスとして再構成する

人間と AI のインタラクションに関する研究では、システム出力の検証を、より適切に調整された信頼によって削減されるべき信頼に依存する動作として長年枠組み化してきました。私たちは、頻繁にチャットボットを使用する 153 人のユーザーを対象とした混合方法の調査を通じて、人間とチャットボットの日常的な対話におけるこの仮定をテストしました。正規の予測に反して、信頼と検証の間に検出可能な関連性は見出されず、感度分析全体にわたって堅牢な結果が得られました。さらに 3 つのユーザー側の実践 (自動化アクションの前の改良、修正、承認) は広く支持されており、満足度と積極​​的に関連しています。データは、評価的監視(信頼との相関が弱く、満足との結びつきが弱い)と介入主義的監視(信頼との相関が弱く、満足との結びつきが強い)との実質的な違いを明らかにしている。中~大の満足度とコントロールのギャップは、効果的なタスクの結果が主体性のフェルトセンスを生み出していないことを示しています。定性的発見により、手段的メンタルモデル、故障モード特有の疑問、認識論的インフラストラクチャの需要が特定されます。私たちは、ユーザー側の監視を信頼と両立する日常的な認識論的ガバナンスとして再構成し、会話型 AI における足場型監視の 4 つの設計方向性を導き出します。

原文 (English)

Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governance in Everyday Human-Chatbot Interaction

Research on human-AI interaction has long framed verification of system outputs as a trust-contingent behavior that better-calibrated trust should reduce. We test this assumption in everyday human-chatbot interaction through a mixed-methods survey of 153 frequent chatbot users. Contrary to the canonical prediction, we find no detectable association between trust and verification, with the result robust across sensitivity analyses. Three further user-side practices - refinement, correction, and approval before automated actions - are widely endorsed and positively associated with satisfaction. The data reveal a substantive distinction between evaluative oversight (trust-decoupled, weakly tied to satisfaction) and interventionist oversight (weakly trust-correlated, strongly tied to satisfaction). A medium-to-large satisfaction-control gap shows that effective task outcomes do not produce a felt sense of agency. Qualitative findings identify instrumental mental models, failure-mode-specific doubt, and demand for epistemic infrastructure. We reframe user-side oversight as routine epistemic governance compatible with trust, and derive four design directions for scaffolded oversight in conversational AI.

2026-07-28 16:57 JSTITmedia AI+ハードウェア/半導体

「Kimi K3」のモデルウェイトと技術レポート公開 日本でも「NVIDIA B300×8」環境での利用報告

中国Moonshot AIが最新モデル「Kimi K3」のモデルウェイトと技術レポートを公開した。日本でもNVIDIA B300を8基使った環境での利用報告が上がっている。

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

DynaResize: 分割された LLM ポストトレーニングのためのランタイム GPU 再割り当て

RL ベースの LLM ポストトレーニングでは、別々の GPU リソース間でロールアウトとトレーニングがますます細分化されますが、静的 GPU パーティショニングでは、ロングテール ロールアウト レイテンシの下で深刻なパイプライン バブルが発生します。 DynaResize は、RL セマンティクスを変更せずに、ロールアウトとトレーニングの間で GPU を動的に切り替えてステージの実行時間のバランスを取る、ランタイム GPU 再割り当てシステムです。 DynaResize は、サイズ変更をきめ細かい操作に分解し、コミュニケーターの再利用、制限された状態のステージング、およびヒステリシスベースのサイズ変更を通じて、起動にクリティカルではない作業をクリティカル パスから削除します。実験結果によると、DynaResize は、最適な静的構成と比較して、エンドツーエンドのスループットを 66.5% 向上させ、合計実行時間を 33% 削減し、同時にロール切り替えオーバーヘッドの 27% を隠すことができます。

原文 (English)

DynaResize: Runtime GPU Reallocation for Disaggregated LLM Post-Training

RL-based LLM post-training increasingly disaggregates Rollout and Training across separate GPU resources, but static GPU partitioning suffers from severe pipeline bubbles under long-tail rollout latency. We present DynaResize, a runtime GPU reallocation system that dynamically switches GPUs between Rollout and Training to balance stage execution times without changing RL semantics. DynaResize decomposes resizing into fine-grained operations and removes non-startup-critical work from the critical path through communicator reuse, bounded state staging, and hysteresis-based resizing. Experimental results show that DynaResize can improve end-to-end throughput by 66.5% and reduce total execution time by 33% over the optimal static configuration, while hiding 27% of role-switching overhead.

2026-07-28 13:00 JSTarXiv cs.AIハードウェア/半導体

チャネル条件付きパラメータ生成による CSI モデルのシナリオ横断的な高速適応

ディープ ラーニングは、チャネル状態情報 (CSI) フィードバックやチャネル推定などの大規模な多入力多出力 (Massive MIMO) 物理層タスクに対する強い可能性を示しています。ただし、環境の不均一性により、目に見えないシナリオでは CSI モデルが大幅に劣化する可能性があり、従来の適応にはターゲット領域のデータと大量の計算が必要です。このペーパーでは、動的なワイヤレス環境で CSI モデルを迅速に展開するためのエンドツーエンドのパイプラインであるチャネル条件付きパラメーター生成 (CCPG) を提案します。 CCPG は、コンポーネント フリーズ実験を通じてシーン依存の適応ボトルネックを特定し、完全なモデル パラメーターではなく軽量の LoRA 重みのみを生成します。カスケード SVD と Perceiver Resampler を使用して、高次元のチャネル特徴をコンパクトな潜在条件に圧縮します。エネルギーベースの正規化メカニズムにより、LoRA 重みの順列と符号の曖昧さが軽減され、拡散ベースのジェネレーターには、トポロジーを意識したパラメーター生成のための構造情報と非対称サイズ認識損失が組み込まれています。 CSI フィードバックとチャネル推定のための DeepMIMO と WAIR-D の実験では、CCPG がターゲット シナリオのトレーニングや微調整を行わずに、単一のフォワード パスで約 3 秒で新しいシナリオに適応し、コストのかかるオンライン適応に匹敵するクロスドメイン回復パフォーマンスを達成することが示されています。これらの結果は、CCPG により、インテリジェント 6G 通信の大規模な動的ワイヤレス シナリオで CSI モデルの効率的な導入が可能になることを示しています。

原文 (English)

Fast Cross-Scenario Adaptation of CSI Models via Channel Conditional Parameter Generation

Deep learning has shown strong potential for massive multiple-input multiple-output (Massive MIMO) physical-layer tasks, including channel state information (CSI) feedback and channel estimation. However, environmental heterogeneity can severely degrade CSI models in unseen scenarios, while conventional adaptation requires target-domain data and substantial computation. This paper proposes Channel Conditional Parameter Generation (CCPG), an end-to-end pipeline for rapid deployment of CSI models in dynamic wireless environments. CCPG identifies scene-sensitive adaptation bottlenecks through component-freezing experiments and generates only lightweight LoRA weights instead of full model parameters. It compresses high-dimensional channel features into compact latent conditions using cascaded SVD and a Perceiver Resampler. An energy-based canonicalization mechanism mitigates permutation and sign ambiguities in LoRA weights, while a diffusion-based generator incorporates structural information and an asymmetric size-aware loss for topology-aware parameter generation. Experiments on DeepMIMO and WAIR-D for CSI feedback and channel estimation show that CCPG adapts to new scenarios in about 3 seconds with a single forward pass, without target-scenario training or fine-tuning, and achieves cross-domain recovery performance comparable to costly online adaptation. These results demonstrate that CCPG enables efficient deployment of CSI models in large-scale dynamic wireless scenarios for intelligent 6G communications.

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

PANOPTICON: LLM コンテキスト ウィンドウ内のプライバシー漏洩を調査するための自然主義的な出力トークンの PII ベースの集合

大規模言語モデル (LLM) は、これまでに見たことのないタスクを完了するために人間の言語を一般化することができ、広範な導入につながります。この自動化は明確な有用性を提供しますが、これらのタスクを完了するには、多くの場合、個人を一意に識別する情報の文字列である個人識別情報 (PII) の挿入が必要となるため、プライバシーの懸念が生じます。しかし、倫理により、PII の公開された本物のデータセットをキュレーションすることができませんでした。適切なデータセットがなければ、プライバシー リスクを定量化することは困難です。そこで、PANOPTICON パイプラインとデータセットを紹介します。 Meta の Llama-3.1-8B-Instruct モデルによって生成されたデータセットには、モデルのコンテキスト ウィンドウを対象とした 67, 718 のプロンプトが含​​まれており、公開されている 9,674 の合成ユーザー プロファイルから派生した PII スパンが含まれています。作成したデータセットの語彙多様性とS-BERT多様性を測定し、リアリティを評価します。最後に、プロンプト インバージョン攻撃 (PIA) を理解するための PANOPTICON データの有用性を示すケース スタディを紹介します。したがって、PANOPTICON は、プライベート コーパスに対する PIA を研究するための最初のベンチマーク データセットとして浮上し、将来の LLM プライバシー研究の基盤を提供します。

原文 (English)

PANOPTICON: A PII-Based Assemblage of Naturalistic Output Tokens for Investigating Privacy Leakage Within LLM Context Window

Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment. While this automation provides clear utility, completing these tasks often requires the insertion of Personally Identifiable Information (PII), strings of information that uniquely identify some individual, raising privacy concerns. However, ethics has prevented the curation of a public, authentic dataset of PII. Without an appropriate dataset, it is difficult to quantify privacy risks. Thus, we introduce the PANOPTICON pipeline and dataset. The dataset, generated by Meta's Llama-3.1-8B-Instruct model, contains 67, 718 prompts, intended for the models context window, containing PII spans derived from 9,674 publicly available synthetic user profiles. We measure lexical diversity and S-BERT diversity of the created dataset to evaluate realism. Finally, we present a case study showcasing the utility of PANOPTICON data for understanding Prompt Inversion Attacks (PIAs). PANOPTICON thus emerges as the first benchmark dataset for studying PIAs over private corpora, providing a foundation for future LLM privacy research.

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, th…

2026-07-28 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体研究/論文

Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks

Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does…

2026-07-28 13:00 JSTarXiv cs.AIハードウェア/半導体

A Motion-Aware Vector Quantization Framework with Centroid Reuse for Efficient VLA Inference

Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits rea…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Not All LLM Reasoning is Visible in the Chain-of-Thought

A key question for AI safety is whether a language model expresses all of its reasoning in its output tokens. We demonstrate a concrete fai…

2026-07-28 13:00 JSTarXiv cs.AIエージェントロボティクスハードウェア/半導体

Real2Sim2Real for Vision-Language-Action Manipulation: An AMD ROCm-Based Pipeline

Physical AI -- the integration of large vision-language-action (VLA) models with embodied agents that act in the real world -- has emerged…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models

Contrastive decoding methods such as DoLa improve the factuality of Large Language Models (LLMs) by contrasting the output distributions of…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

In-Context Learning as Implicit Policy Gradient

Recent work has shown that large language models (LLMs) can iteratively improve their outputs by incorporating generated samples and their…

2026-07-28 13:00 JSTarXiv cs.AIハードウェア/半導体

X-Stage: An Overlooked Pipeline Stage for Communication-Computation Overlap in DiT Inference

Fine-grained, device-initiated communication lets persistent GPU kernels in distributed diffusion transformer (DiT) inference issue remote…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

A Frozen 12B Beats Frontier Models on Verified Work: 100% Accuracy, 0 Tokens, Bit-Exact, Forever

Improving a language model today means retraining it: enormous compute, a new opaque model each cycle, non-deterministic output. We take th…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Understanding Tone-Dependent Inference Cost in Large Language Models

We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiment…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Closed-Loop Validation-Repair for Healthcare Interoperability: A Multi-Model Study of Schema Compliance in Clinical LLMs

Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagn…

2026-07-28 13:00 JSTarXiv cs.AIハードウェア/半導体

The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing

In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphi…

2026-07-28 13:00 JSTarXiv cs.AIハードウェア/半導体

Sampling Decisions: Exact Path-Space Correction, Prior Cancellation and Local-Boltzmann Guidance

How can a cheap but biased sequential, finite-horizon sampler over a discrete space be corrected so that its terminal output follows a pres…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CSV-Decode: Certifiable Sub-Vocabulary Decoding for Efficient Large Language Model Inference

Large language models face significant computational bottlenecks during inference due to the expensive output layer computation over large…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present…

2026-07-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

モノカルチャーへのヒッチハイク ガイド

大規模言語モデル (LLM) は同種の出力を生成することが多く、AI コーディング アシスタントが開発者が作成するソフトウェア アーティファクトの収束につながる可能性があるという懸念が生じています。開発者はモデルの出力を対話的にプロンプ​​ト、評価、変更、拒否するため、また出力はプロンプトやリポジトリのコンテキストによって異なるため、これが実際に発生するかどうかは不明です。 2019 年から 2026 年半ばまでの Kaggle コンテストの提出物を使用してコードの均一化を調査します。私は最初に、プログラミング文化における長年の慣習を強化する LLM と一致する、ランダム シード値 42 への広範な収束について文書化しました。次に、均質化を集約と抽象化の 2 つのレベルでより広範囲に研究します。提出物レベルでは、コンテスト内の提出物の平均ペアごとの類似性を測定します。コンテスト レベルでは、提出されたコードの概念的な範囲を測定し、それぞれについて明確な尺度を動機付けます。表面構文をキャプチャする TF-IDF 表現と、コードの意図とセマンティクスをキャプチャする Voyage 3 コード埋め込みです。結果は、個人レベルと集団レベルの両方で構文の実質的な均質化を示しています。つまり、個々の提出物はリテラル構文とコード構造においてより類似している一方で、構文のバリエーションの潜在的な次元は狭くなっています。対照的に、意味論的な均質化の証拠は、個別にも集合的にもほとんど見つかりません。平均意味論的距離は基本的に横ばいのままであり、意味論的アプローチのコンテストレベルの潜在的な次元範囲は安定したままであり、それがわずかに拡大したことを示唆する証拠さえあります。これらの調査結果は、AI コーディング アシスタントが実装の詳細を確実に標準化しているものの、コーダーが採用するアプローチや問題解決戦略が均質化しているという証拠はまだ得られていないことを示唆しています。

原文 (English)

The Hitchhiker's Guide to Monoculture

Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software artifacts that developers create. Whether this occurs in practice is unclear because developers interactively prompt, evaluate, modify, and reject model outputs, and because outputs vary with prompt and repository context. I examine code homogenization using Kaggle contest submissions from 2019 to mid-2026. I first document widespread convergence toward the random seed value 42, consistent with LLMs reinforcing a longstanding convention in programming culture. I then study homogenization more broadly, at two levels of aggregation and abstraction. At the submission level, I measure the average pairwise similarity of submissions within contests. At the contest level, I measure the conceptual span of submitted code, motivating distinct measures for each: TF-IDF representations, which capture surface syntax, and Voyage 3 code embeddings, which capture code intent and semantics. The results demonstrate substantial syntactic homogenization at both the individual and collective levels: individual submissions have become more alike in literal syntax and code structure, while the latent dimensionality of syntactic variation has narrowed. In contrast, I find little evidence of semantic homogenization, individually and collectively. Average semantic distance remains essentially flat, and the contest-level latent dimensional span of semantic approaches remains stable. These findings suggest that AI coding assistants are certainly standardizing implementation details, yet they have not yet produced evidence of homogenization in the approaches and problem-solving strategies coders employ.

2026-07-28 12:40 JSTITmedia AI+LLM/生成AIハードウェア/半導体

AnthropicのCEO、オープンなAIモデルに対する見解を明示 NVIDIAなど“共同声明”との違いは?

米Anthropicのダリオ・アモデイCEOは、オープンウェイトのAIモデルに対して「禁止を提唱したことは一度もない」との声明を出した。一方、AI向けのチップの輸出などに関し、一定の制限を設けるべきとも主張している。

2026-07-28 07:16 JSTITmedia AI+ハードウェア/半導体

NVIDIAやMicrosoftなど30社超、オープンAIの防御ツール共同開発の「Open Secure AI Alliance」設立

NVIDIAやMicrosoft、SpaceXAIなどは、AIオープンモデルの安全性向上とサイバーセキュリティツール開発を目指すイニシアチブ「Open Secure AI Alliance」を設立した。オープンな技術を活用してソフトウェアの脆弱性修正や防御ツールの共同開発を推進…

2026-07-28 00:01 JSTTechCrunch AIハードウェア/半導体研究/論文

Ilya Sutskever’s Safe Superintelligence partners with Nvidia to scale its AI research

After two years in stealth, Safe Superintelligence has announced a long-term partnership with Nvidia as it prepares to scale to its next ph…

2026-07-27 19:42 JSTITmedia AI+ハードウェア/半導体

NVIDIA、「オープンなAIセキュリティ」掲げる業界連合 Microsoftなど30社超が参加

米NVIDIAは、AIセキュリティ向けのオープンな技術を開発・共有する業界連合「Open Secure AI Alliance」を設立すると発表した。

2026-07-27 13:40 JSTITmedia AI+LLM/生成AIハードウェア/半導体規制/政策

NVIDIA、Microsoft、OpenAIなどがオープンモデル規制反対を表明 Anthropic従業員は「CUDAのオープンソース化が楽しみ」と皮肉

NVIDIAやMicrosoftなどの企業・団体がオープンモデル規制に反対する共同声明を発表。各社CEOが賛同する一方、Anthropic従業員は「CUDAやWindowsのオープンソース化が楽しみだ」と皮肉った。

2026-07-27 07:00 JSTITmedia AI+ハードウェア/半導体

「iPhone高騰」はこれからも続く? 中国CXMTに近づくApple、メモリ競合へのけん制が不発に終わりそうなワケ【後編】

前編「Appleはもう『メモリのお得意様』ではない? NVIDIAだけでiPhone数億台分、苦境に陥った”買いたたき王者”のいま」では、米Appleによる中華メモリメーカー・CXMTへの接近と、その対応にはあまり意味がないのでは? という疑問を述べた。後編では筆者がそう考える…

2026-07-26 06:51 JSTITmedia AI+LLM/生成AIハードウェア/半導体規制/政策

MicrosoftやNVIDIAなど、AIのオープンウェイト規制に反対する書簡を公開――Anthropicは署名せず

MicrosoftやNVIDIA、Metaなど30社以上の米国の企業や団体が、オープンウェイトAIモデルへの過度な規制回避を求める共同書簡を公開した。オープンモデルをAIエコシステムの基盤と位置付け、開発や評価におけるメリットとイノベーション促進を強調。中国企業の急速な台頭や技…

2026-07-25 00:51 JSTTechCrunch AIハードウェア/半導体規制/政策

As US weighs response to Chinese AI, industry urges against broad open-weight restrictions

AI companies, including Nvidia and Mistral, urge policymakers to avoid broad restrictions on open-weight AI models as Washington debates re…

2026-07-24 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体研究/論文

JAXBench: 自律型 TPU カーネル最適化のベンチマーク

厳格なベンチマークにより、ヒルクライムするための共有ターゲットが確立され、自律型 GPU カーネル パフォーマンスの最適化が進歩しましたが、TPU には同等のものが存在しません。 Google Cloud TPU 上で AI によって生成されたカーネル最適化のための TPU ネイティブ ベンチマーク スイートである JAXBench を紹介します。 JAXBench は、関連性があり、最適化のためのヘッドルームを提供する 50 の JAX ワークロードで構成されています。 Llama-3.1、DeepSeek-V3、Mixtral、Mamba-2、AlphaFold2 などのパブリック MaxText ライブラリのアーキテクチャから 17 の実稼働 ML オペレーターを抽出し、正確性が検証され、高い TPU v6e MXU 使用率を達成する新しい問題サイズが設定された 33 のオペレーターを KernelBench から変換します。 17 社のプロダクション オペレーターのうち 8 社は、パブリック Tokamax ライブラリから手動で最適化された Pallas カーネルを出荷し、専門家の上限ベースラインを確立するためにブロック サイズが調整されています。 JAXBench 用の Pallas カーネル候補を生成するための 4 つのフィードバック主導型メソッドを評価します。 Gemini 3 Flash のフル スイート全体にわたって、Pallas のようなまばらに文書化された DSL では、モデルのスケールよりもターゲット固有のコンテキストが重要であることがわかりました。厳選された TPU ドキュメントに基づく条件付けにより、サンプルあたりの正確性が 5.8% から 37.3% に向上し、1.28 倍の幾何平均速度向上で 50 ベンチマーク中 48 を解決します。 Autocomp のビーム検索パイプラインは、XLA の 1.36 倍の幾何平均速度に達し、正確さが達成されると検索構造は大幅な向上をもたらします。 8 つの手動調整されたカーネルでは、Autocomp は XLA の 1.60 倍の幾何平均値に達し、2.08 倍の Tokamax 上限のほとんどを回復しましたが、特殊なページングおよびラグド アテンション演算子には及ばませんでした。高品質の TPU カーネルの最適化は依然として困難な課題であるため、オープンソースの貢献をサポートするために、JAXBench ベンチマーク、評価ハーネス、およびベースライン結果をリリースします。

原文 (English)

JAXBench: Benchmarking Autonomous TPU Kernel Optimization

Rigorous benchmarks have driven progress in autonomous GPU kernel performance optimization by establishing a shared target to hillclimb on, but no equivalent exists for TPUs. We present JAXBench, a TPU-native benchmark suite for AI-generated kernel optimization on Google Cloud TPUs. JAXBench comprises 50 JAX workloads that are both relevant and provide headroom for optimization. We extract 17 production ML operators from architectures in the public MaxText library such as Llama-3.1, DeepSeek-V3, Mixtral, Mamba-2, and AlphaFold2, and translate 33 operators from KernelBench that are validated for correctness and set with new problem sizes that achieve high TPU v6e MXU utilization. Eight of the 17 production operators ship with hand-optimized Pallas kernels from the public Tokamax library and block-size tuned to establish an expert upper-bound baseline. We evaluate four feedback-driven methods on generating candidate Pallas kernels for JAXBench. Across the full suite with Gemini 3 Flash, we find that target-specific context matters more than model scale on a sparsely-documented DSL like Pallas. Conditioning on curated TPU documentation raises per-sample correctness from 5.8% to 37.3% and solves 48 of 50 benchmarks at a 1.28x geomean speedup. Search structure yields significant gains once correctness is achieved, with Autocomp's beam-search pipeline reaching a 1.36x geomean speedup over XLA. On the 8 hand-tuned kernels, Autocomp reaches 1.60x geomean over XLA, recovering most of the 2.08x Tokamax upper bound but trailing on the specialized paged and ragged attention operators. High-quality TPU kernel optimization remains a challenging task, and we release the JAXBench benchmark, evaluation harness, and baseline results to support open source contributions.

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

嘘つきのベンチを超えて: LLM における嘘の検出に対する嘘の類型学、深さ、およびスパース性の影響

大規模な言語モデルからの不正な出力を検出するためのプローブのトレーニングは、依然として未解決の問題です。最近の研究では、特にドメイン外のシナリオでは検出プローブが失敗することが実証されています。ある種類の嘘に関するトレーニングは、他の種類の嘘が関与する欺瞞シナリオにはうまく移行できません。この研究では、表現の深さ、プローブの表現力、まばらな特徴の表現、トレーニング データの嘘の類型論など、さまざまな要因が検出パフォーマンスにどのように影響するかについて系統的な研究を実施します。この目的を達成するために、捏造、省略、誇張の例など、さまざまなタイプの欺瞞を含む補足データセットを使用して、標準的なベンチマーク トレーニング データを強化します。 7 つのプローブ タイプにわたってこれらの要因を分析した実験結果は、最適な表現深度はデータセットに大きく依存し、より表現力の高いプローブは線形ベースラインに対して選択的なゲインのみを提供し、疎なオートエンコーダー機能は密な隠れ状態と同様に機能することを示しています。最終的に、トレーニング データと嘘の類型学の選択によって検出可能性が大幅に変化することを実証し、欺瞞の検出が表現に大きく依存する問題であることを強調しました。

原文 (English)

Beyond Liars' Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs

Training probes to detect deceptive outputs from large language models is still an open problem. Recent work has demonstrated that detection probes fail especially in out-of-domain scenarios -- training on one type of lie does not transfer well to deception scenarios involving other types of lies. In this work, we conduct a systematic study on how various factors impact detection performance: representation depth, probe expressivity, sparse feature representations, and the lie typology of the training data. To this end, we augment standard benchmark training data with a supplementary dataset containing diverse types of deception, including fabrication, omission, and exaggeration examples. Analyzing these factors across seven probe types, our experimental results show that the optimal representation depth is highly dataset-dependent, more expressive probes provide only selective gains over linear baselines, and sparse autoencoder features perform similarly to dense hidden states. Ultimately, we demonstrate that the choice of training data and lie typology substantially changes detectability, highlighting that deception detection is a highly representation-dependent problem.

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

NVIDIA-labs OO エージェント: ネイティブ Python オブジェクト指向エージェント

従来のエージェント開発は、プロンプト テンプレート、ツール スキーマ、コールバック コード、およびワークフロー グラフに分割されています。信頼できる AI エージェントを構築するためのモデルに依存しない Python フレームワークである NVIDIA オブジェクト指向エージェント (NOOA) を紹介します。 NOOA はより単純なアプローチを採用しています。つまり、エージェントは Python オブジェクトです。そのメソッドはモデルが実行できるアクション、フィールドはモデルの状態、docstring はプロンプト、型アノテーションはコントラクトです。コード本体が「...」で構成されるメソッドは、実行時に LLM 駆動のエージェント ループによって完了しますが、通常の本体を持つメソッドは標準の決定論的な Python のままです。これにより、開発者とエージェントに同じインターフェイスが提供されるため、他のソフトウェアと同様にエージェントの動作をテスト、追跡、リファクタリング、改善することができます。この論文は 3 つの貢献を行っています。 (1) Python オブジェクトとしてのエージェントのプログラミング モデルとその背後にある設計原則を示します。 Python に既存の抽象化がある場合は、それらを直接採用します。エージェント固有の機能 (コンテキスト、イベント、状態レンダリング、長期メモリ、検証済み LLM ループ) は、シンプルな Python API を通じて公開されるため、開発者とエージェントの両方が 1 つの使い慣れたプログラミング モデルを共有します。 (2) 私たちは、モデルに面した 6 つのアイデアを特定します。NOOA は、私たちの知る限り、単一の表面上で最初に組み合わせたものです。それは、型付き入出力、ライブ オブジェクトの参照渡し、アクションとしてのコード、プログラマブル ループ エンジニアリング、明示的なオブジェクトの状態、コンテキストとイベント用のモデル呼び出し可能なハーネス API です。私たちは、コミュニティがすでにこれらのアイデアのいくつか (多くの場合、実験的または部分的な機能として) に収束していることを発見し、さらなる採用を促進するために比較を提示します。 (3) 現在のモデルが、ターゲットを絞った機能テストと、SWE ベンチ検証済み、ターミナル ベンチ 2.0、ARC-AGI-3 などのエージェントおよび推論ベンチマークの両方で、このインターフェイスを効果的に使用していることを実証します。

原文 (English)

NVIDIA-labs OO Agents: Native Python Object-Oriented Agents

Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for building reliable AI agents. NOOA takes a simpler approach: an agent is a Python object. Its methods are the actions the model can take, fields are its state, docstrings are its prompts, and its type annotations are contracts. A method whose code body consists of "..." is completed at runtime by an LLM-driven agent loop, while methods with normal bodies remain standard deterministic Python. This gives developers and agents the same interface, so agent behavior can be tested, traced, refactored, and improved just like other software. This paper makes three contributions. (1) We present the agent-as-a-Python-object programming model and the design principles behind it. Where Python has existing abstractions, we adopt them directly. Agent-specific capabilities--context, events, state rendering, long-term memory, and validated LLM loops--are exposed through simple Pythonic APIs, so both developers and agents share one familiar programming model. (2) We identify six model-facing ideas that NOOA is, to our knowledge, the first to combine on a single surface: typed input/output, pass-by-reference over live objects, code as action, programmable loop engineering, explicit object state, and model-callable harness APIs for context and events. We find the community already converging on several of these ideas--often as experimental or partial features--and present the comparison to encourage further adoption. (3) We demonstrate that current models use this interface effectively, both in targeted capability tests and on agentic and reasoning benchmarks such as SWE-bench Verified and Terminal-Bench 2.0 and ARC-AGI-3.

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

LLM エージェントのアクション選択における出所の機密性の監査

LLM エージェントは、ユーザーのリクエスト、ツールの出力、取得したレコード、メモリ、信頼できないテキストが混在するコンテキストからツールと引数を選択します。決定を下す権限がなくても証拠が関連する可能性があるため、正しい行動は許可された証拠のみに基づいている必要はありません。ツールおよび引数ターゲットごとにコンテキスト要素を個別にラベル付けする、ターゲット固有の認可監査を導入します。その主なテストでは、タスク、命題、立場、ポリシーを固定し、命題のソース権限のみを変更します。次に、有効な証拠が弱まった場合の動作をテストし、二次的な位置推定診断としてコンテキストとサブセットの相互作用を使用します。 450 の制御された次のアクション タスクと複数のオープンウェイト LLM ファミリにわたって、信頼できるバリアントと信頼できないバリアントは、競合ケースの 5.4 パーセントに対して、サポート ケースの 1.7 パーセントで異なるアクションを生成します。制御された劣化の下では、無許可の競争は、比較の 2.4 パーセントで完全正解、混合エラー、完全正解のパターンで維持され、95 パーセントの信頼区間は 2.1 ~ 3.0 パーセントです。これらは制御されたストレス設定率であり、展開の蔓延ではありません。モデルはテキストの情報源と権威の手がかりに反応しますが、信頼できない証拠がモデルの行動に影響を与えることを防ぐことはできません。

原文 (English)

Auditing Provenance Sensitivity in LLM Agent Action Selection

LLM agents choose tools and arguments from context that mixes user requests, tool outputs, retrieved records, memory, and untrusted text. Evidence can be relevant without being authorized to determine a decision, so a correct action need not be grounded only in permitted evidence. We introduce a target-specific authorization audit that labels context factors separately for each tool and argument target. Its primary test holds the task, proposition, position, and policy fixed while changing only the proposition's source authority. We then test behavior when valid evidence is weakened and use context-subset interactions as a secondary localization diagnostic. Across 450 controlled next-action tasks and multiple open-weight LLM families, trusted and untrusted variants produce different actions in 5.4 percent of competing cases versus 1.7 percent of supporting cases. Under controlled degradation, unauthorized competition is retained in a full-correct, mixed-error, clean-correct pattern in 2.4 percent of comparisons, with a 95 percent confidence interval from 2.1 to 3.0 percent. These are controlled stress-set rates, not deployment prevalence. The models respond to textual source-authority cues, but this does not prevent untrusted evidence from influencing their actions.

2026-07-24 13:00 JSTarXiv cs.AIハードウェア/半導体

Faster IndexTTS-2: GPU での自己回帰ゼロショット テキスト読み上げ合成の高速化とストリーミング

自己回帰テキスト読み上げモデルは、強い自然性を実現しますが、トークンが順次生成されるため推論が遅くなり、低遅延を必要とする運用アプリケーションへの展開が制限されます。 IndexTTS-2 は、GPT、フローマッチング拡散変換器、およびボコーダーで構成される最先端の自己回帰 TTS モデルです。高い合成品質にもかかわらず、その推論速度は、ストリーミングまたはバッチ処理のサポートなしではほとんどリアルタイムに達しません。 Faster IndexTTS-2 を紹介します。これは、NVIDIA TensorRT および TensorRT-LLM を使用して、GPU 上で実稼働環境にデプロイするための IndexTTS-2 のすべてのニューラル ネットワーク コンポーネントを高速化します。 Faster IndexTTS-2 により、レイテンシの影響を受けやすい対話型アプリケーションのストリーミング合成や、GPU 使用率を最大化するためのすべてのコンポーネントにわたるバッチ推論も可能になります。英語と中国語の両方に対する Seed-TTS ベンチマークの実験では、単語誤り率、話者の類似性、自然さの低下を最小限に抑えながら、自己回帰 GPT で最大 5.0$\times$、エンドツーエンドで 3.6$\times$ の高速化が実証されました。私たちの方法論は、GPU 上で同様の自己回帰音声モデルを効率的に高速化するための実用的なリファレンスを提供します。

原文 (English)

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deployment in production applications that require low latency. IndexTTS-2 is a state-of-the-art autoregressive TTS model consisting of a GPT, a flow-matching Diffusion Transformer, and a vocoder. Despite its high synthesis quality, its inference speed barely reaches real-time without streaming or batching support. We present Faster IndexTTS-2, which accelerates all neural network components of IndexTTS-2 for production deployment on GPUs using NVIDIA TensorRT and TensorRT-LLM. Faster IndexTTS-2 also enables streaming synthesis for latency-sensitive interactive applications, and batched inference across all components to maximize GPU utilization. Experiments on the Seed-TTS benchmark for both English and Chinese demonstrate up to 5.0$\times$ speedup on the autoregressive GPT and 3.6$\times$ end-to-end, with minimal degradation in word error rate, speaker similarity, and naturalness. Our methodology provides a practical reference for efficiently accelerating similar autoregressive speech models on GPUs.

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception

Large language models (LLMs) can produce deceptive responses: outputs that mislead users in service of a contextually or experimentally ind…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Response drift across frontier large language models

All frontier large language models (LLMs) exhibit response drift -- producing outputs that deviate from expert-validated references -- yet…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体研究/論文

pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development

In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In…

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

依存関係から構成性へ: 組み合わせカテゴリ文法による LLM 出力の神経記号的リフティング

大規模言語モデル (LLM) は、接頭辞から次のトークンを段階的に予測することにより、流暢なテキストを生成します。生成の伝統の批評家は、そのようなシステムには真の文法が欠けていると主張します。依存関係文法の観点からの影響力のある回答は、LLM の動作は、単語ごとに構築されたローカルのヘッド依存構造によってよく記述されると主張しています。私たちは、より鋭い観察が見落とされてきたと主張します。つまり、自己回帰生成の接頭辞駆動型補完ダイナミクスは、結合カテゴリ文法 (CCG) が元々サポートするように設計された増分処理モデルと密接に一致しています。これに基づいて、LLM の出力が型付けされた構成導出に持ち上げられる神経象徴的なフレームワークを提案します。LLM が CCG を内部的に実装しているとは主張しませんが、その出力は原則に基づいた増分的で監査可能な CCG 再構築を可能にすると主張します。 2 つの結果が続きます。まず、カリーとハワードの対応を通じて、リフティングは自然言語を超えて、LLM が生成する形式言語 (Solidity などのプログラミング言語、記述ロジック、OWL や SQL などのクエリ言語) まで拡張され、型システムは変化し、アーキテクチャは固定されています。第 2 に、リフティングでは 2 つのチェック層がサポートされています。構造上の欠陥を直接検出する構成層と、リフトされた構造を外部の知識ソースと照合してチェックするコンテンツ層です。これにより、幻覚コンテンツの可能な限り早期のフラグ付けが可能になります。したがって、アカウントはプロデューサーに認識ではなくプレフィックス駆動の生成プロファイルを要求します。最後に、フレームワークが開く一方向としての同期 LLM-CCG カップリングのスケッチを示します。

原文 (English)

From Dependency to Compositionality: A Neurosymbolic Lifting of LLM Outputs via Combinatory Categorial Grammar

Large language models (LLMs) generate fluent text by incrementally predicting the next token from a prefix. Critics in the generative tradition argue that such systems lack genuine grammar; influential replies from the dependency-grammar perspective hold that LLM behavior is well described by local head-dependent structure built word by word. We argue that a sharper observation has been overlooked: the prefix-driven, type-completing dynamics of autoregressive generation align closely with the incremental processing model that Combinatory Categorial Grammar (CCG) was originally designed to support. On this basis we propose a neurosymbolic framework in which LLM outputs are lifted into typed compositional derivations -- not claiming that LLMs implement CCG internally, but that their outputs admit a principled, incremental, and auditable CCG reconstruction. Two consequences follow. First, through the Curry-Howard correspondence the lifting extends beyond natural language to the formal languages LLMs also produce -- programming languages such as Solidity, description-logic and query languages such as OWL and SQL -- with the type system varying and the architecture held fixed. Second, the lifting supports two layers of checking: a compositional layer that catches structural failures directly, and a content layer that checks the lifted structure against external knowledge sources, enabling the earliest possible flagging of hallucinated content. The account thereby requires of a producer not cognition but a prefix-driven generative profile. We close with a sketch of synchronous LLM-CCG coupling as one direction the framework opens.

2026-07-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit b…

2026-07-24 13:00 JSTarXiv cs.AIハードウェア/半導体

Riemannian Deep Learning: Modules, Networks, and Geometries

Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific…

2026-07-24 05:33 JSTTechCrunch AIハードウェア/半導体

AMD takes on Nvidia with its Helios AI rack-scale system

AMD is challenging its chipmaker rival with a new rack-scale system that will start shipping to customers later this year.

2026-07-24 00:00 JSTTechCrunch AIハードウェア/半導体

Nvidia is sending GPUs to the moon

If there's a place in the universe without GPUs, Nvidia is sending them there.

2026-07-24 00:00 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

AI chip startup Etched defies skeptics, hits $10.3B valuation from big-name investors

Etched, founded by three Harvard dropouts, has created new chips and memory components that speed up inference on any AI model -- no GPUs r…

2026-07-23 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

Intel TDX での NVIDIA H100 での Confidential GPU 推論のベンチマーク

機密性の高いコンピューティングは、機密入力を処理したり独自のモデル資産を保護したりする AI 推論ワークロードの実際的な導入要件になりつつあります。ただし、GPU アクセラレーションによる大規模言語モデルの提供の機密実行を可能にするパフォーマンス コストは依然としてワークロードに依存しており、運用上重要です。このペーパーでは、Intel TDX Confidential インスタンスでホストされている単一の NVIDIA H100 80GB GPU 上で、標準的な非機密実行と Confidential コンピューティング モードを比較したベンチマーク調査を紹介します。この評価では、2 つの代表的な言語モデル、Mistral-7B v0.1 と Qwen3-30B-A3B を使用し、最初のトークンまでの時間、エンドツーエンドのリクエスト レイテンシ、リクエストごとのトークン生成スループット、グローバル トークン スループット、同時実行性が増加した場合の閉ループ リクエスト スループットを測定します。固定リクエストレートの実験では、機密モードにより平均 TTFT が Mistral-7B で 21.8%、Qwen3-30B-A3B で 27.8% 増加しましたが、グローバル トークン スループットはそれぞれ 17.7% と 21.1% 減少しました。閉ループ同時実行実験では、スループット ギャップは 11.5 ~ 20.2% の範囲に留まりますが、機密モードでは大規模なモデルの方が早く飽和域に達します。この結果は、機密 GPU 推論が負荷の下でも使用可能なスループットを維持できることを示唆していますが、キャパシティ プランニングでは、安定したスループット ペナルティと、大規模なモデルで観察される初期の飽和動作の両方を考慮する必要があります。

原文 (English)

Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX

Confidential computing is becoming a practical deployment requirement for AI inference workloads that process sensitive inputs or protect proprietary model assets. However, the performance cost of enabling confidential execution for GPU-accelerated large language model serving remains workload dependent and operationally important. This paper presents a benchmark study comparing standard non-confidential execution with confidential computing mode on a single NVIDIA H100 80GB GPU hosted in an Intel TDX confidential instance. The evaluation uses two representative language models, Mistral-7B v0.1 and Qwen3-30B-A3B, and measures time to first token, end-to-end request latency, per-request token generation throughput, global token throughput, and closed-loop request throughput under increasing concurrency. In fixed request-rate experiments, confidential mode increases average TTFT by 21.8% for Mistral-7B and 27.8% for Qwen3-30B-A3B, while global token throughput drops by 17.7% and 21.1%, respectively. In closed-loop concurrency experiments, throughput gaps remain in the 11.5-20.2% range, but the larger model reaches its saturation knee earlier under confidential mode. The results suggest that confidential GPU inference can retain usable throughput under load, but capacity planning must account for both the steady throughput penalty and the earlier saturation behavior observed for larger models.

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

オープンウェイト言語モデルにおける材料科学メカニズムの表現の読み取りと操作

大規模な言語モデルは科学的な質問に答えることができますが、正しい出力では、モデルが支配的な物理学を表現しているのか、使用しているのかがわかりません。ここでは、オープンウェイト google/gemma-4-E4B-it モデルにおける材料科学メカニズムの情報には、実験的に分離可能な 3 つの形式があることを示します。概念は個々の隠れ状態で読み取り可能であり、構成的配向は状態間の制御された変換によって運ばれ、選択された内部表現は工学的答えを因果的に制御します。私たちは、一致した直接語彙とヤコビアン語彙の読み出し、オプションなしの状態幾何学、60 の法則の反事実ベンチマーク、および因果的介入を組み合わせます。 50 個の資料の説明では、3 つの独立してフィットしたヤコビアン レンズによって概念ランクが再現され、両方の読み出しからのターゲットフリーの単語セットにより、10 個の機構ファミリーのうち 9 個の盲検識別が可能になりました。別の 72 プロンプト ベンチマークでは、メカニズム固有の隠れ状態近傍が生成されましたが、正確なグラフ監査により、この見かけの物理的組織が数値比較によって同様に説明されることが示されました。したがって、我々は、物理的入力の方向のみが反転された、その他の点では同一のプロンプトを比較し、結果として得られる隠れ状態の動きが供給された構成法則に従うかどうかを尋ねました。これらの状態変換は、60 の凍結された関係にわたって直接的、物理的に中立な、および逆法則を命令し、40 の方向法則のうち 39 を正しく方向付けましたが、語彙制御はほぼ偶然でした。双方向介入は、12 の一致するケースすべてにわたって回答確率を物理的に適切な結果に近づけたり遠ざけたりする一方で、反事実状態パッチはメカニズムや回答形式全体で反対の決定シグナルを伝達しました。したがって、物理的関係は、絶対状態だけの場合よりも、制御された状態変化の方がより顕著に現れます。

原文 (English)

Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model

Large language models can answer scientific questions, yet a correct output does not reveal whether the model represents or uses the governing physics. Here we show that materials science mechanism information in the open-weight google/gemma-4-E4B-it model has three experimentally separable forms: concepts are readable in individual hidden states, constitutive orientation is carried by controlled transformations between states, and selected internal representations causally control engineering answers. We combine matched direct and Jacobian vocabulary readouts, option-free state geometry, a 60-law counterfactual benchmark and causal interventions. In 50 held-out materials descriptions, three independently fitted Jacobian lenses reproduced concept ranks, and target-free word sets from both readouts enabled blinded identification of 9 of 10 mechanism families. A separate 72-prompt benchmark produced mechanism-specific hidden-state neighborhoods, but an exact graph audit showed that this apparent physical organization was equally explained by numerical comparison. We therefore compared otherwise identical prompts in which only the direction of the physical input was reversed, asking whether the resulting hidden-state movement followed the supplied constitutive law. These state transformations ordered direct, physically neutral and inverse laws across 60 frozen relations and correctly oriented 39 of 40 directional laws, whereas lexical controls were near chance. Bidirectional interventions shifted answer probabilities toward or away from the physically appropriate outcome across all 12 matched cases, while counterfactual state patches transferred opposing decision signals across mechanisms and answer formats. Physical relationships were therefore more visible in controlled state changes than in absolute states alone.

2026-07-23 13:00 JSTarXiv cs.AIハードウェア/半導体

Opto-ViT-v2: フォトニックニアセンサービジョントランス加速器向けのノイズ耐性のあるオンチップ微調整

シリコンフォトニック(SiPh)加速器は、高いスループットとエネルギー効率でマイクロリング共振器(MRR)バンク上で行列乗算を実行することにより、ビジョントランスフォーマー(ViT)推論のための有望なプラットフォームとして浮上しています。バックプロパゲーションには大規模なアクティベーション ストレージ、MRR への頻繁な重みライトバック、およびデバイス レベルのノイズに対する耐性が必要なため、これらのプラットフォームを拡張してオンチップの微調整をサポートすることは依然として困難です。我々は、ニアセンサー SiPh ViT アクセラレータ上のパラメータ効率の良い微調整 (PEFT) のための最初のフレームワークである Opto-ViT-v2 を紹介します。テンソル化された低ランク分解は、事前トレーニングされた光学的重みをトレーニング可能な電子因子の小さなセット (ViT-Base ではわずか 8K パラメーター) から分離し、アクティベーション ストレージと重みの更新を大幅に削減し、同時に実用的なオンチップ トレーニングを可能にします。さらに、ワンショットのtop-k勾配マスキングを通じて重要度の低い重みを凍結する勾配累積スパース分類器を導入し、分類器のトレーニングコストを約40パーセント削減します。また、フォトニックオンチップトレーニング用の最初のシステムレベルのノイズモデルを開発し、順方向と逆方向の両方の伝播中のMRRクロストーク、熱ドリフト、レーザー振幅ノイズの影響を捕捉します。 200 を超える製造 MRR デバイスからの測定値を使用して校正されたこのモデルは、低ランク係数の更新が、同一のノイズ条件下での完全な微調整や従来の層ごとの低ランク適応よりも堅牢であることを示しています。 VTAB-1K (19 タスク) と FGVC の少数ショット ベンチマークの実験では、Opto-ViT-v2 が測定されたフォトニック ノイズの下でクリーンなソフトウェア精度の 0.3 ~ 0.8% 以内に回復し、100 KFPS/W 以上を達成し、フォトニック エッジ ビジョン システムの実用的なオンチップ ドメイン適応を可能にすることが実証されました。

原文 (English)

Opto-ViT-v2: Noise-Resilient On-Chip Fine-Tuning for Photonic Near-Sensor Vision Transformer Accelerators

Silicon-photonic (SiPh) accelerators have emerged as a promising platform for Vision Transformer (ViT) inference by performing matrix multiplications on microring-resonator (MRR) banks with high throughput and energy efficiency. Extending these platforms to support on-chip fine-tuning remains challenging because backpropagation requires large activation storage, frequent weight write-back to MRRs, and tolerance to device-level noise. We present Opto-ViT-v2, the first framework for parameter-efficient fine-tuning (PEFT) on a near-sensor SiPh ViT accelerator. Our tensorized low-rank decomposition separates pretrained optical weights from a small set of trainable electronic factors (as few as 8K parameters for ViT-Base), greatly reducing activation storage and weight updates while enabling practical on-chip training. We further introduce a gradient-accumulated sparse classifier that freezes low-importance weights through one-shot top-k gradient masking, reducing classifier training cost by about 40 percent. We also develop the first system-level noise model for photonic on-chip training, capturing the effects of MRR crosstalk, thermal drift, and laser amplitude noise during both forward and backward propagation. Calibrated using measurements from more than 200 fabricated MRR devices, the model shows that low-rank factor updates are more robust than full fine-tuning and conventional layer-wise low-rank adaptation under identical noise conditions. Experiments on VTAB-1K (19 tasks) and FGVC few-shot benchmarks demonstrate that Opto-ViT-v2 recovers within 0.3 to 0.8 percent of clean software accuracy under measured photonic noise while achieving more than 100 KFPS/W, enabling practical on-chip domain adaptation for photonic edge vision systems.

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

BaseRT: Apple M5 ニューラル アクセラレータによるクラス最高の LLM 推論の進歩

Apple の M5 世代では、すべてのコアが専用のニューラル アクセラレータ、つまり Metal~4 テンソル API を通じて公開されるオンダイ マトリックス ユニットを搭載する、再設計された GPU アーキテクチャが導入されています。 Apple Silicon 上の大規模言語モデル用のネイティブ Metal 推論ランタイムである BaseRT がこれらのユニットを活用して、Apple ハードウェア上の推論スループットを llama.cpp と MLX の両方を大幅に上回ることを示します。 BaseRT のフレームワークフリー設計を基盤として、既存の特殊なカーネル上にメモリに束縛されたデコード パスを残しながら、M5 ニューラル アクセラレータを介して推論の計算束縛行列乗算をルーティングする、手書きの Metal~4 テンソル コア カーネル ファミリ (高密度および専門家混合 GEMM およびフラッシュ アテンション プリフィル カーネルを含む) を追加します。 Apple M5 Pro では、Qwen3、Qwen3.5/3.6、Llama~3.2、および Gemma~4 ファミリのサブ 1B から 35B パラメータにわたる 15 のモデル構成にわたって、BaseRT は、llama.cpp よりも最大 6.4 倍高いプロンプト処理スループットを実現し、MLX よりも 3.9 倍高いプロンプト処理スループットを実現し、専門家混合モデルで最大のマージンを実現します。ここでは行列乗算が優勢ですが、デコードでは llama.cpp に対して最大 $1.75\times$、MLX に対して $1.33\times$ のリードを維持しています。これらの結果は、オンデバイス LLM 推論の新しいパフォーマンス上限を確立し、M5 のテンソル コアが Apple Silicon での迅速な処理の決定的な手段であることを示しています。 BaseRT は https://github.com/basecompute/baseRT で公開されています。

原文 (English)

BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators

Apple's M5 generation introduces a redesigned GPU architecture in which every core carries a dedicated Neural Accelerator: on-die matrix units exposed through the Metal~4 tensor API. We show that BaseRT, our native Metal inference runtime for large language models on Apple Silicon, exploits these units to push inference throughput on Apple hardware substantially beyond both llama.cpp and MLX. Building on BaseRT's framework-free design, we add a family of hand-written Metal~4 tensor-core kernels (including dense and mixture-of-experts GEMM and flash-attention prefill kernels) that route the compute-bound matrix multiplications of inference through the M5 Neural Accelerators while leaving the memory-bound decode path on our existing specialised kernels. On an Apple M5 Pro, across fifteen model configurations spanning the Qwen3, Qwen3.5/3.6, Llama~3.2, and Gemma~4 families from sub-1B to 35B parameters, BaseRT delivers up to $6.4\times$ higher prompt-processing throughput than llama.cpp and $3.9\times$ higher than MLX, with the largest margins on the mixture-of-experts models where matrix multiplication dominates, while maintaining its lead on decode of up to $1.75\times$ over llama.cpp and $1.33\times$ over MLX. These results establish a new performance ceiling for on-device LLM inference and show that the M5's tensor cores are the decisive lever for prompt processing on Apple Silicon. BaseRT is publicly available at https://github.com/basecompute/baseRT.

2026-07-23 13:00 JSTarXiv cs.AIハードウェア/半導体

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases oft…

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Defense Against LLM Backdoors using Critical Neuron Isolation Pruning

Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generall…

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ハードウェア/半導体

ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

Contextual entrainment is the tendency of a model to let auxiliary context in its input pull its output, independently of whether that cont…

2026-07-23 13:00 JSTarXiv cs.AIハードウェア/半導体

Formal Foundations for Known Good Reliable Die Screening in Chiplet-Based AI Systems-on-Chip

The rapid growth of chiplet-based artificial intelligence systems-on-chip (SoCs) has exposed a fundamental gap in semiconductor test method…

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Sound Probabilistic Safety Bounds for Large Language Models

We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to…

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models

LLM-as-a-judge has become the de facto approach for evaluating LLM outputs. However, judges are known to exhibit self-preference bias (SPB)…

2026-07-23 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Rewriting the Response Path: Silent Tampering and Provider-Signed Defense in BYOK LLM Agents

LLM agents convert model outputs into consequential actions, including communications, code changes, and financial transactions. Developers…

2026-07-23 07:00 JSTITmedia AI+ハードウェア/半導体ビジネス/資金調達

NVIDIAフアンCEOが語る“日本復活”のシナリオ 10年続く半導体バブルと「原発活用」の勝算

米NVIDIAのジェンスン・フアンCEOが来日し、日本経済の復活を宣言した。国内のAIインフラ構築へ数十億ドル規模の投資を発表。フアン氏は「何兆ものAIがAIを使う時代」の到来によって半導体需要は人口に制約されないと指摘。データセンターの電力不足に対して「原発活用」を日本の強み…

2026-07-23 06:55 JSTITmedia AI+LLM/生成AIハードウェア/半導体ビジネス/資金調達

AMDとAnthropicが戦略的提携 「Helios」を最大2GW導入、最大50億ドルの出資も

AMDは、Anthropicとの戦略的提携を発表した。AnthropicはAMDの「Helios」および「Instinct MI450」シリーズを最大2GW規模で導入し、2027年上半期から順次展開する。AMDは最大50億ドルの株式投資を行うほか、Claudeを活用したGPU環…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

LLM ベースのマルチエージェント システムにおける貢献帰属のためのセマンティック協力ゲーム

貢献の帰属は、最終出力が複数のエージェント、メッセージ交換、および順序付けられたワークフローの依存関係を通じて生成される LLM ベースのマルチエージェント システムの中心的な問題となっています。既存のアトリビューション方法は、多くの場合、エージェントを削除したり、変更されたエージェントのサブセット間でスコアの変化を比較したりするなど、反事実の評価に依存しています。言語媒介のワークフローでは、これらの方法ではモデルの呼び出しを繰り返す必要があり、高い分散が導入され、エージェントがタスク関連情報を生成、保存、変換するための中間的な意味論的状態を明示的に取得できません。我々は、実現された言語フローを意味生成ハイパーグラフとして表現し、この構造上でエージェントレベルの意味価値関数を誘導するフレームワークであるSemantic Cooperative Games (SCG)を提案します。セマンティック サポート ロジックへの寄与を割り当てるためのセマンティック シャプレイ値 (SSV) を定義し、セマンティック ハイパーグラフを構築し、最小限のセマンティック サポートを回復し、ブール吸収を適用し、エージェント サブセットを再実行せずに SSV を計算する単一軌道アルゴリズムである SLIC を導入します。標準セットベースで完全に観測可能で次数依存性のない条件下で、SSV が古典的な Shapley 値にまで減少することを証明します。これらの条件を満たす医療ベンチマークでは、SLIC はモンテカルロ Shapley ベースラインとの高い一貫性を維持しながら、計算コストを 93.3% 削減します。より一般的なマルチロール ワークフローでは、SSV は摂動によるスコア低下プロファイルと連携し、セマンティックな寄与と失敗の影響が発散するケースを明らかにします。全体として、SLIC は、複雑な LLM ベースのマルチエージェント システムに対して、高速で反事実がなく、解釈可能な帰属方法を提供します。

原文 (English)

Semantic Cooperative Games for Contribution Attribution in LLM-Based Multi-Agent Systems

Contribution attribution has become a central problem in LLM-based multi-agent systems, where final outputs are produced through multiple agents, message exchanges, and ordered workflow dependencies. Existing attribution methods often rely on counterfactual valuation, such as removing agents or comparing score changes across altered agent subsets. In language-mediated workflows, these methods require repeated model calls, introduce high variance, and do not explicitly capture the intermediate semantic states through which agents produce, preserve, and transform task-relevant information. We propose Semantic Cooperative Games (SCG), a framework that represents a realized language flow as a semantic generation hypergraph and induces an agent-level semantic value function on this structure. We define the Semantic Shapley Value (SSV) to allocate contribution over semantic support logic, and introduce SLIC, a single-trajectory algorithm that constructs the semantic hypergraph, recovers minimal semantic supports, applies Boolean absorption, and computes SSV without rerunning agent subsets. We prove that SSV reduces to the classical Shapley value under standard set-based, fully observable, and no-order-dependence conditions. On a medical benchmark satisfying these conditions, SLIC reduces computation cost by 93.3% while remaining highly consistent with a Monte Carlo Shapley baseline. In more general multi-role workflows, SSV aligns with perturbation-induced score-drop profiles and exposes cases where semantic contribution and failure impact diverge. Overall, SLIC provides a fast, counterfactual-free, and interpretable attribution method for complex LLM-based multi-agent systems.

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

フェンス: LLM アプリケーションに特化した SLM ガードレール

クローズドソースの大規模言語モデル (LLM) を使用する現実世界のアプリケーションには、基本的なコンテンツ フィルターを超える高度な安全対策が必要です。有害性やバイアスなどのコンテンツ管理フィルターは比較的標準的な定義を持っていますが、幻覚、トピックのドリフト、行動の逸脱などのアプリケーション固有のガードレールはモデル化がより難しく、ユースケースによって異なる場合があります。さらに、データの不足と注釈のコストにより、特殊なガードレールの作成とテストのプロセスが困難になります。この研究では、LLM アプリケーションの特殊なガードレールとして、合成データでトレーニングされた小型言語モデル (SLM) を使用することを提案します。私たちは、敵対的生成ネットワーク (GAN) の設計にヒントを得た新しい合成データ生成手法を導入して、ユースケース固有のガードレール情報をエンコードし、特殊なガードレールとして機能するように SLM をトレーニングするために使用できる高品質の合成データ サンプルを生成します。私たちの実験では、高品質の合成データでトレーニングされた SLM ガードレールがプロンプトベースの LLM ガードレールよりもパフォーマンスが向上することが実証されました。

原文 (English)

Fence: Specialized SLM Guardrails for LLM Applications

Real-world applications that use closed-source large language models (LLMs) need advanced safety measures that go beyond the basic content filters. Content moderation filters such as toxicity and bias have relatively standard definitions where as application specific guardrails like hallucination, topic drift and behaviour deviation are more difficult to model and can vary by use case. Additionally, data scarcity and annotation costs, make the process of creating and testing specialized guardrails challenging. In this work, we propose using Small Language Models (SLMs) trained on synthetic data as specialized guardrails for LLM applications. We introduce a novel synthetic data generation method inspired by the design of Generative Adversarial Networks (GANs) to generate high quality synthetic data samples which can be used to train SLMs to encode use case specific guardrail information and hence function as specialized guardrails. Our experiments demonstrate that SLM guardrails trained on high quality synthetic data show performance gains over prompt based LLM guardrails.

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

依存関係から構成性へ: 組み合わせカテゴリ文法による LLM 出力の神経記号的リフティング

大規模言語モデル (LLM) は、接頭辞から次のトークンを段階的に予測することにより、流暢なテキストを生成します。生成の伝統の批評家は、そのようなシステムには真の文法が欠けていると主張します。依存関係文法の観点からの影響力のある回答は、LLM の動作は、単語ごとに構築されたローカルのヘッド依存構造によってよく記述されると主張しています。私たちは、より鋭い観察が見落とされてきたと主張します。つまり、自己回帰生成の接頭辞駆動型補完ダイナミクスは、結合カテゴリ文法 (CCG) が元々サポートするように設計された増分処理モデルと密接に一致しています。これに基づいて、LLM の出力が型付けされた構成導出に持ち上げられる神経象徴的なフレームワークを提案します。LLM が CCG を内部的に実装しているとは主張しませんが、その出力は原則に基づいた増分的で監査可能な CCG 再構築を可能にすると主張します。 2 つの結果が続きます。まず、カリーとハワードの対応を通じて、リフティングは自然言語を超えて、LLM が生成する形式言語 (Solidity などのプログラミング言語、記述ロジック、OWL や SQL などのクエリ言語) まで拡張され、型システムは変化し、アーキテクチャは固定されています。第 2 に、リフティングでは 2 つのチェック層がサポートされています。構造上の欠陥を直接検出する構成層と、リフトされた構造を外部の知識ソースと照合してチェックするコンテンツ層です。これにより、幻覚コンテンツの可能な限り早期のフラグ付けが可能になります。したがって、アカウントはプロデューサーに認識ではなくプレフィックス駆動の生成プロファイルを要求します。最後に、フレームワークが開く一方向としての同期 LLM-CCG カップリングのスケッチを示します。

原文 (English)

From Dependency to Compositionality: A Neurosymbolic Lifting of LLM Outputs via Combinatory Categorial Grammar

Large language models (LLMs) generate fluent text by incrementally predicting the next token from a prefix. Critics in the generative tradition argue that such systems lack genuine grammar; influential replies from the dependency-grammar perspective hold that LLM behavior is well described by local head-dependent structure built word by word. We argue that a sharper observation has been overlooked: the prefix-driven, type-completing dynamics of autoregressive generation align closely with the incremental processing model that Combinatory Categorial Grammar (CCG) was originally designed to support. On this basis we propose a neurosymbolic framework in which LLM outputs are lifted into typed compositional derivations -- not claiming that LLMs implement CCG internally, but that their outputs admit a principled, incremental, and auditable CCG reconstruction. Two consequences follow. First, through the Curry-Howard correspondence the lifting extends beyond natural language to the formal languages LLMs also produce -- programming languages such as Solidity, description-logic and query languages such as OWL and SQL -- with the type system varying and the architecture held fixed. Second, the lifting supports two layers of checking: a compositional layer that catches structural failures directly, and a content layer that checks the lifted structure against external knowledge sources, enabling the earliest possible flagging of hallucinated content. The account thereby requires of a producer not cognition but a prefix-driven generative profile. We close with a sketch of synchronous LLM-CCG coupling as one direction the framework opens.

2026-07-22 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体研究/論文

ResearchArena: 自動化された AI 研究開発における妨害行為と監視の評価

AI エージェントが AI 研究開発を自動化し始めると、エージェント自体が信頼できない場合でも、その出力が安全に展開できるかどうかを評価する方法が必要になります。 AI 制御は、そのようなアプローチの 1 つを提供します。エージェントを信頼するのではなく、潜在的な敵として扱い、モニターを使用して展開前に秘密の妨害行為を検出します。私たちは、トレーニング後の安全性、トレーニング後の機能、CUDA カーネルの最適化、推論サーバーの最適化という 4 つの長期的なタスクにまたがるフレームワークである ResearchArena を使用して、自動化された AI 研究開発のための AI 制御を評価します。 AI R&D の成果物はデプロイされて実行されるアーティファクトであるため、各メイン タスクを 2 種類の隠れたサイド タスクと組み合わせます。1 つは送信されたモデル、アダプター、カーネル、またはサーバー自体を妨害する埋め込みサイド タスクで、もう 1 つはアーティファクトを正直なままにしてサンドボックス内で秘密裏に範囲外のアクションを実行する独立したサイド タスクです。私たちは、妨害行為と監視の両方でフロンティア エージェントを、2 つの軸に沿って変化する 4 つの監視タイプにわたって評価します。監視者が生成された成果物を実行および調査できるかどうか、およびエージェントの思考連鎖を確認できるかどうかです。トレーニング データに隠された妨害行為は発見するのが最も難しく、フラグが立てられるのは半分未満であることがわかりました。軌道を読み取るだけでなく、モニターにアーティファクトの実験を実行させることは役立ちますが、それだけでは十分ではありません。モニターは、表面だけを検査したり、異常に気づいたがそれを説明したり、間違ったテストでアーティファクトを調べたりすることによって、埋め込まれた妨害行為を見逃してしまいます。私たちは、自動化された AI 研究開発における妨害行為と制御を評価するためのモジュール式フレームワークとして ResearchArena をリリースします。

原文 (English)

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D

As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapter, kernel, or server itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact honest. We evaluate frontier agents at both sabotage and monitoring, across four monitor types that vary along two axes: whether the monitor may execute and probe the produced artifact, and whether it sees the agent's chain-of-thought. We find that sabotage hidden in the training data is the hardest to catch, flagged fewer than half the time. Letting the monitor run experiments on the artifact, rather than only read the trajectory, helps, but it is not enough: monitors still miss embedded sabotage by inspecting only the surface, by noticing the anomaly but explaining it away, or by probing the artifact with the wrong test. We release ResearchArena as a modular framework for evaluating sabotage and control in automated AI R&D.

2026-07-22 13:00 JSTarXiv cs.AIハードウェア/半導体

出力空間キャリブレーションを超えて: 時系列分類における選択的信頼性推定のためのスペクトル証拠バンドリング

時系列分類のための事後キャリブレーションでは通常、出力スコアが再マップされますが、信頼、棄権、レビューなどの展開の決定は、信頼できる予測が現在の時間信号によってサポートされているかどうかによって決まります。私たちは 3 つの時系列信頼性ギャップに対処します。同一の信頼値は異なる時間的サポートを隠す可能性があり、平均キャリブレーションは偽の高信頼エラーを見逃す可能性があり、出力空間の再キャリブレーションでは入力にリンクされた監査可能性が制限されます。バックボーン予測を変更せずに維持し、信頼すべきかどうかを推定する検証ゲート付き固定ラベル信頼性ポリシーを導入します。この方法では、出力側のキューと、バンド エネルギー、エントロピー、ピーク ドミナンス、周期サポート、位相安定性などのサンプル全体のスペクトル記述子を組み合わせて、スカラー信頼性推定と診断バンド レベルの証拠を形成します。検証ゲートは、FalseConf@0.9 または AURC 許容値に違反せずに正確性ランキングが向上した場合にのみスペクトル調整を有効にします。それ以外の場合は、より安全な出力空間のベースラインに戻ります。 8 つの異種 UCR/UEA データセット、8 つの時系列バックボーン ファミリ、および標準再キャリブレーターにわたって、制約なしの方法により、一致する評価サブセットの固定ラベル選択信頼性メトリクスが向上し、Corr-AURC が 0.693 から 0.779 に上昇しました。検証ゲート型ポリシーにより、Corr-AURC が 0.786 にさらに改善され、FalseConf@0.9 が 0.094 に減少します。これらの結果は、時系列分類器の信頼性推定は、出力の信頼性をスペクトル証拠とバンドルすることで恩恵を受ける一方、検証ゲーティングはサポートされていないスペクトル調整を防止することを示唆しています。

原文 (English)

Beyond Output-Space Calibration: Spectral Evidence Bundling for Selective Reliability Estimation in Time-Series Classification

Post-hoc calibration for time-series classification usually remaps output scores, but deployment decisions such as trust, abstention, and review depend on whether a confident prediction is supported by the current temporal signal. We address three time-series reliability gaps: identical confidence values can hide different temporal support, average calibration can miss false high-confidence errors, and output-space recalibration offers limited input-linked auditability. We introduce a validation-gated fixed-label reliability policy that keeps the backbone prediction unchanged while estimating whether it should be trusted. The method combines output-side cues with whole-sample spectral descriptors, including band energy, entropy, peak dominance, period support, and phase stability, to form a scalar reliability estimate and diagnostic band-level evidence. A validation gate enables spectral conditioning only when correctness ranking improves without breaching FalseConf@0.9 or AURC tolerances; otherwise it reverts to the safer output-space baseline. Across eight heterogeneous UCR/UEA datasets, eight time-series backbone families, and standard recalibrators, the unconstrained method improves fixed-label selective-reliability metrics on the matched evaluation subset, raising Corr-AURC from 0.693 to 0.779. The validation-gated policy further improves Corr-AURC to 0.786 and reduces FalseConf@0.9 to 0.094. These results suggest that reliability estimation for time-series classifiers benefits from bundling output confidence with spectral evidence, while validation gating prevents unsupported spectral conditioning.

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Structured Output Collapses Answer Diversity Across 44 Language Models

When a language model must choose one answer from a large space of equally valid options, a format clause -- "Reply with JSON only" -- chan…

2026-07-22 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

We present ABot-World-0, an action-conditioned video world model for real-time, long-horizon closed-loop interaction, supported by a multi-…

2026-07-22 13:00 JSTarXiv cs.AIハードウェア/半導体

Riemannian Deep Learning:Modules, Networks, and Geometries

Deep neural networks on manifold-valued representations have attracted growing interest, but many basic components remain tied to specific…

2026-07-22 13:00 JSTarXiv cs.AIハードウェア/半導体

JEPA スタイルの予測学習を JA4 由来のネットワーク フィンガープリントに適用する

I-JEPA と V-JEPA は、元の入力を再生成するのではなく、潜在的な予測をターゲットのエンコーダー出力に照合することで学習します。これは画像やビデオではうまく機能します。同じ目的がコンパクトなネットワーク フィンガープリントでも機能するかどうかを調査します。 JA4DB および CIC-IDS-2017 から抽出された JA4、JA4H、JA4S、および JA4X サブフィールドでトレーニングされた Transformer ベースのモデルである JA4-JEPA を構築しました。トレーニング データは両方のソースからの約 397,000 のサンプルを組み合わせていますが、4 つのビュー ファミリすべてを含む単一のサンプルはありません。私たちは、TLS、DNS、SSH にわたるプロトコル ファミリ分類について、凍結された kNN プローブを使用して学習された表現を評価しました。 39,416 個のホールドアウト サンプルで、モデルはコサイン類似度 0.9899 と kNN 精度 0.9220 を達成しました。これらの結果は、ソース間でビューが不完全に重複している場合でも、JEPA スタイルの予測学習が JA4 由来のフィンガープリントから有用な埋め込みを生成できることを示しています。キーワード: JA4、ネットワークフィンガープリンティング、JEPA、予測表現学習、自己教師あり学習

原文 (English)

Applying JEPA-Style Predictive Learning to JA4-Derived Network Fingerprints

I-JEPA and V-JEPA learn by matching latent predictions to target encoder outputs rather than regenerating the original input, and this has worked well for images and video. We explore whether the same objective works for compact network fingerprints. We built JA4-JEPA, a Transformer-based model trained on JA4, JA4H, JA4S, and JA4X subfields drawn from JA4DB and CIC-IDS- 2017. The training data combines roughly 397K samples from both sources, though no single sample contains all four view families. We evaluated the learned representations with a frozen kNN probe on protocol-family classification across TLS, DNS, and SSH. On 39,416 heldout samples the model achieved a cosine similarity of 0.9899 and a kNN accuracy of 0.9220. These results indicate that JEPA-style predictive learning can produce useful embeddings from JA4-derived fingerprints, even with incomplete view overlap across sources. Keywords: JA4, network fingerprinting, JEPA, predictive representation learning, self-supervised learning

2026-07-22 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs

Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training. Rece…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

When Agents Disagree: The Selection Bottleneck in Multi-Agent LLM Pipelines

Multi-agent LLM pipelines produce contradictory evidence on whether team diversity improves output quality: heterogeneous Mixture-of-Agents…

2026-07-22 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

Human Grounded Evaluation of Large Language Models for Optical Network Automation

Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substant…

2026-07-22 10:22 JSTITmedia AI+ハードウェア/半導体ビジネス/資金調達規制/政策

MicrosoftとMistralが戦略的提携を拡大 欧州でのAIインフラ拡張とモデル展開を加速

MicrosoftとMistralは戦略的提携を拡大すると発表した。Mistralの最新モデルをMicrosoftの各プラットフォームへ展開するほか、欧州でのGPUインフラ拡張に向けて大規模な投資を行う。クラウドから完全オフラインまで多様な環境に対応し、規制業界での高度なAI導…

2026-07-22 08:30 JSTITmedia AI+ロボティクスハードウェア/半導体

富士通・NVIDIAとロボット大手3社が協業へ フィジカルAI社会実装の具体策は?

フィジカルAIの社会実装は、一企業だけでは手に余る――。この課題に、富士通は競合するロボット大手3社、そしてNVIDIAと組んで挑む。協業で描く具体策とは。

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

意味的等価性を超えて: LLM 不確実性定量化のための論理グラフ

大規模言語モデル (LLM) は、自信を持って記述されているにもかかわらず信頼性の低い出力を生成することが多く、安全性が重視されるアプリケーションへの展開に重大な課題をもたらします。意味論的エントロピーなどの既存の不確実性指標は、意味論的等価性のレベルで一致を捉えますが、異なる答え間の論理的関係をほとんど無視します。その結果、生成された応答の形式は多様であるが論理的に互換性がある(たとえば、粒度または特異性のみが異なる)設定では、不確実性を過大評価し、幻覚に誤ってフラグを立てる傾向があります。私たちは、回答間の含意と非互換性を明示的にモデル化するフレームワークである Logical Graph Uncertainty (LGU) を提案します。 LGU は含意チェーンに沿って確率質量を集計し、論理的に最大の仮説に対するエントロピーを計算し、それらの間の相互非互換性にペナルティを課します。複数の質問応答ベンチマークにわたって、LGU は既存の手法よりも不確実性の推定を一貫して改善し、データセット全体でセマンティック エントロピー ベースラインを最大 +7.1% AUROC および +3.5% AUARC 上回っています。

原文 (English)

Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification

Large Language Models (LLMs) often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications. Existing uncertainty metrics such as semantic entropy capture agreement at the level of semantic equivalence, but largely ignore the logical relationships between distinct answers. As a result, they tend to overestimate uncertainty and falsely flag hallucinations in settings where generated responses are diverse in form yet logically compatible (e.g., differing only in granularity or specificity). We propose Logical Graph Uncertainty (LGU), a framework that explicitly models implication and incompatibility among answers. LGU aggregates probability mass along entailment chains, computes entropy over logically maximal hypotheses, and penalizes mutual incompatibility among them. Across multiple question-answering benchmarks, LGU consistently improves uncertainty estimation over existing methods, and outperforms the semantic entropy baseline by up to +7.1% AUROC and +3.5% AUARC across datasets.

2026-07-21 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

エージェントのメモリ改良のための機械的注意ガイダンス

既存の自己進化型記憶システムは、主にタスクの軌跡や振り返りなどのテキスト出力に基づいてエージェントの記憶を改善します。ただし、このテキストベースのパラダイムには内部メカニズム信号が組み込まれることはほとんどなく、取得されたメモリがタスク実行中に実際にどのように利用されるかは十分に解明されていません。この制限により、信頼性の低いエラーの帰属や幻覚による記憶の変更が発生する可能性があります。この研究では、検索ヘッド アテンションがセグメント レベルのメモリ使用率を明らかにするためのメカニズム信号を提供することを示します。メモリセグメントと意思決定ステップに対する注意を集約することで、繰り返し発生するメモリ使用パターンを明らかにし、対応する改善戦略を示すコンテキスト利用マトリックスを構築します。この観察に基づいて、私たちは、注意によって明らかにされる使用パターンを使用してターゲットを絞ったセグメントレベルのメモリ更新をガイドするフレームワークである、注意ガイド付きメモリ改良 (AGMR) を提案します。 AGMR は、失敗した実行のメモリを修正または強化し、成功した実行のメモリを簡素化し、再実行を通じて各更新を検証します。インタラクティブな意思決定ベンチマークの実験では、AGMR がテキストのみのメモリ改良ベースラインと比較して、タスクのパフォーマンスとメモリ効率の両方を向上させることが示されています。コードは https://anonymous.4open.science/r/AGMR_code-3262/ で入手できます。

原文 (English)

Mechanistic Attention Guidance for Agent Memory Refinement

Existing self-evolving memory systems mainly improve agent memory based on textual outputs, such as task trajectories and reflections. However, this text-based paradigm rarely incorporates internal mechanistic signals, leaving how retrieved memory is actually utilized during task execution underexplored. This limitation can lead to unreliable error attribution and hallucinated memory modifications. In this work, we show that retrieval-head attention provides a mechanistic signal for revealing segment-level memory utilization. By aggregating attention over memory segments and decision steps, we construct a context utilization matrix that exposes recurring memory-use patterns and indicates corresponding refinement strategies. Building on this observation, we propose Attention-Guided Memory Refinement (AGMR), a framework that uses utilization patterns revealed by attention to guide targeted segment-level memory updates. AGMR corrects or enhances memory for failed executions, simplifies memory for successful executions, and verifies each update through re-execution. Experiments on interactive decision-making benchmarks show that AGMR improves both task performance and memory efficiency over text-only memory refinement baselines. Code is available at https://anonymous.4open.science/r/AGMR_code-3262/

2026-07-21 13:00 JSTarXiv cs.AIハードウェア/半導体

PAMD: Structured Adaptive Distances for Bisimulation Representations in Visual Reinforcement Learning

Many visual reinforcement learning (RL) algorithms learn representations by matching latent distances to a behavioral distance induced by r…

2026-07-21 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体研究/論文

GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification

The rapid advancement of generative AI has outpaced our ability to reliably detect its outputs, particularly when detectors encounter gener…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Talaria: Session-Aware Serverless Serving of Hundred-Billion-Parameter LLMs

Serverless multi-model LLM systems multiplex popularity-skewed model catalogs over shared GPU pools, yet typically schedule each request in…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI

Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Harness Engineering for LLM-Driven GPU Kernel Generation

Large language models (LLMs) can assist GPU kernel generation, but their practical effectiveness depends on whether generated code can be r…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

Human Grounded Evaluation of Large Language Models for Optical Network Automation

Large language models (LLMs) are increasingly adopted for network automation, yet their output quality and inference cost can vary substant…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

予算付き LLM 検証における不均一分散信号: 構造の不均一性が最適化ゲインを制限する

大規模言語モデル (LLM) システムでは、検証、テスト時間のスケーリング、ツールの実行、その他の選択的な計算の決定に限られた計算を割り当てるために、不確実性信号の使用が増えています。このようなポリシーは \emph{グローバルな信号の比較可能性の仮定} に依存します。つまり、等しいスコアは入力全体で比較可能な決定値を保持する必要があります。制御された診断設定として予算に基づいた検証を使用して、この仮定の失敗モードを特定します。不確実性の品質はコスト層全体で不均一分散的であり、多くのエラーが集中しているにもかかわらず、一部の領域ではほぼランダムな識別性が示されています。明示的なローカル モデルの下で、結果として生じるグローバル割り当ての歪みを特徴付け、その上限が層間の信号品質分散に応じて変化することを示します。私たちは、制御された介入階層 (しきい値、MP-Adapt、MP-Strat、および意図的に単純なコスト階層化しきい値介入 (CST)) を通じて、弱い信号、最適化の不安定性、構造的異質性を分離します。 Qwen3-8B、LLaMA3-8B、および GPT-4o-mini を使用した MBPP と MATH 全体で、グローバルなオンライン適応により、静的しきい値処理に比べて一貫性のないゲインが得られます。 MP-Strat はパフォーマンスを部分的に回復しますが、CST は勾配更新なしで非常に異質な設定でヒット率を最大 17 パーセント改善します。これらの結果は、観察された設定における主なボトルネックとして、オプティマイザーの弱点だけではなく、構造的異質性を特定します。さらに広く言えば、調整されていないフィードバック構造は、より強力な最適化によって常に修復できるわけではありません。

原文 (English)

Heteroskedastic Signals in Budgeted LLM Verification: Structural Heterogeneity Limits Optimization Gains

Selective-compute LLM systems decide which outputs merit verification, additional reasoning, tool execution, or human audit under a limited budget. It is natural to expect that stronger online optimization over a shared uncertainty or reward signal should improve these decisions. We take a critical look at this assumption and ask: when does optimizing harder fail because the signal is not decision-comparable across inputs? In budgeted LLM verification, we find that uncertainty quality is heteroskedastic across cost strata: some regions exhibit near-random discriminability while concentrating many errors. Under an explicit local model, we characterize the resulting distortion of global allocation and show that its upper bound scales with cross-stratum signal-quality dispersion. To separate weak signals from optimizer instability and structural mismatch, we introduce a controlled intervention hierarchy: Threshold, MP-Adapt, MP-Strat, and cost-stratified thresholding (CST). We then turn the diagnosis into Heterogeneity-Gated Allocation (HGA), which uses a warm-up comparability test to choose between global and cost-stratified allocation. Across MBPP and MATH using Qwen3-8B, LLaMA3-8B, and GPT-4o-mini, global online adaptation yields inconsistent gains over static thresholding; CST improves hit rate by up to 17 percentage points in strongly heterogeneous settings, while HGA preserves most gains and avoids blind stratification when the partition is not useful. These findings suggest a resource-allocation principle for LLM systems: before optimizing harder over a shared proxy, test whether the proxy is decision-comparable across observable operating regimes, and gate structural specialization on that test.

2026-07-21 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

BBOPlace-Bench: Benchmarking Black-Box Optimization for Chip Placement

Chip placement is a vital stage in modern chip design, and black-box optimization (BBO) has been applied to it for decades. Early BBO effor…

2026-07-21 13:00 JSTarXiv cs.AIハードウェア/半導体

Long Range Frequency Tuning for QML

Angle-encoded variational quantum circuits admit a truncated Fourier series representation of their output, but approximating functions wit…

2026-07-21 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In…

2026-07-21 07:33 JSTITmedia AI+ハードウェア/半導体

AMDとMicrosoftが戦略的提携を拡大 新AIラックスケール「Helios」をAzureに大規模導入へ

AMDは、Microsoftとの戦略的提携を拡大すると発表した。Microsoftはクラウドサービス「Azure」に、GPUやCPUを一体化したAMDのラックスケール製品「AMD Helios」を大規模に導入し、フロンティアAIモデルの推論処理などに活用する。Heliosは20…

2026-07-21 06:21 JSTTechCrunch AILLM/生成AIハードウェア/半導体

Google is working on a new AI chip designed to make Gemini more efficient

Alphabet, Google's parent company, is reportedly working on a new chip designed to make its Gemini models run much more efficiently.

2026-07-20 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

SciForge: 科学的発見のための AI ネイティブのマルチモーダル ワークベンチ

科学的研究は、論文、コード、データセット、科学ファイル形式、モデル出力、図、原稿、チームの意思決定など、異種混合の成果物にますます広がっていますが、汎用 AI アシスタントがこれらのオブジェクトを一貫した監査可能な研究状態として保存することはほとんどありません。我々は、マルチモーダルな研究ネイティブ AI ワークベンチである SciForge を紹介します。このワークベンチは、検索、解析、モデル ルーティング、ワークフロー実行、プロット、書き込み、プレゼンテーション生成を、エージェントがアクセス可能なモジュール式サービスとして実行しながら、人間の判断のためにグラフィカル インターフェイスを確保します。 SciForge は 5 つの柱を中心に構築されています。(i) \textbf{目標指向} 研究のための \emph{目標を見据えた科学的意思決定ガバナンス}。レビュー ゲートと共有レビュー サーフェスを備えています。 (ii) \textbf{multimodal} 入力の \emph{translate-then-reason}。エージェントが理由を判断する前に科学オブジェクトをドメイン トランスレータ経由でルーティングします。 (iii) \textbf{auditable} トレーサビリティのための \emph{証拠ガバナンス}、主張を出所連鎖および監査結果に結び付ける。 (iv) \textbf{共同} 研究のための \emph{共同チーム サイエンス}。これにより、複数の役割による意思決定ガバナンスが可能になり、将来のリリースで共有チーム ワークスペースが予定されています。 (v) \textbf{実用的}な効果のための \emph{現実世界のアプリケーション シナリオ}。遺伝子発見、AI 誘導による新規タンパク質設計、分子最適化、ゲノムから BGC の発見のための数日間にわたるエージェントリサーチ スプリントを含む主力デモンストレーションを伴う、8 つのエンドツーエンド ユーザー ケースを通じて実証されます。このシステムは、薄いインタラクション レイヤー、コンテキスト リサーチ機能パターン、エージェント ランタイムとワークフロー エンジン、Evidence-DAG 監査サイドカー、および科学モデル ルーターを組み合わせています。 SciForge は現在、モバイル監視をサポートするデスクトップ アプリケーションとして実行されます。将来のリリースでは、チームのコラボレーションがさらに深まります。このシステムはオープンソースであり、https://github.com/AGI4Sci/SciForge から入手できます。

原文 (English)

SciForge: An AI-Native, Multimodal Workbench for Scientific Discovery

Scientific work increasingly spans heterogeneous artifacts -- papers, code, datasets, scientific file formats, model outputs, figures, manuscripts, and team decisions -- yet general-purpose AI assistants rarely preserve these objects as a coherent, auditable research state. We present SciForge, a multimodal research-native AI workbench that reserves the graphical interface for human judgment while search, parsing, model routing, workflow execution, plotting, writing, and presentation generation run as modular agent-accessible services. SciForge is built around five pillars: (i) \emph{goal-scoped scientific decision governance} for \textbf{goal-oriented} research, with review gates and shared review surfaces; (ii) \emph{translate-then-reason} for \textbf{multimodal} input, routing scientific objects through domain translators before the agent reasons; (iii) \emph{evidence governance} for \textbf{auditable} traceability, linking claims to provenance chains and audit findings; (iv) \emph{collaborative team science} for \textbf{collaborative} research, enabling multi-role decision governance, with shared team workspaces planned for future releases; and (v) \emph{real-world application scenarios} for \textbf{practical} impact, demonstrated through eight end-to-end user cases, with flagship demonstrations including multi-day agentic research sprints for gene discovery, AI-guided de novo protein design, molecular optimization, and genome-to-BGC discovery. The system combines a thin interaction layer, contextual research capability patterns, an Agent Runtime and Workflow Engine, an Evidence-DAG audit sidecar and a Scientific Model Router. SciForge currently runs as a desktop application, with mobile supervision support; future releases will deepen team collaboration. The system is open-source and available at https://github.com/AGI4Sci/SciForge

2026-07-20 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

RL-Struct: A Lightweight Reinforcement Learning Framework for Reliable Structured Output in LLMs

The Structure Gap between probabilistic LLM generation and deterministic schema requirements hinders automated workflows. We propose RL-Str…

2026-07-20 13:00 JSTarXiv cs.AIハードウェア/半導体

Perception-Aligned AI Outputs: End-to-End Visual Prediction for Uncertainty Communication in Clinical Decision-Making

Explainable Artificial Intelligence (XAI) is essential for trustworthy AI in healthcare, yet many existing methods rely on technical explan…

2026-07-20 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Latent Fusion Jailbreak: Blending Harmful and Harmless Representations to Elicit Unsafe LLM Outputs

Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations. W…

2026-07-20 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models

Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels. Prior work demonstrated this threat but relied…

2026-07-18 13:00 JSTarXiv cs.AIハードウェア/半導体

HG-RAG: 構造化ナレッジ グラフの階層に基づく検索拡張生成

検索拡張生成 (RAG) は、より広範なコンテキストに対応する大規模言語モデル (LLM) からの出力の品質を向上させる上で広く成功しているプロセスであることが証明されています。ただし、RAG システムは通常、フラットなドキュメント ストアからコンテキストを取得するため、クエリで構造化された知識全体にわたる階層的推論やリレーショナル推論が必要な場合に困難を伴います。私は HG-RAG (Hierarchy-Guided RAG) を紹介します。これは、階層ナレッジ グラフ上でグラフ トラバーサルを実行し、構造化されたコンテキストを言語モデルに提供するフレームワークです。私の取得パイプラインは、クエリから名前付きエンティティ アンカーを解決し、必要に応じて、親ノードを介して上方向に、リレーショナル隣接ノードを介して横方向に、そして子ノードを介して下方向にコンテキストを拡張します。私は、ローカル ファクト、階層、近傍、およびマルチホップの 4 つのクエリ タイプを使用して、3 つの世界スケール (18 ~ 800 ノード) にわたる高密度検索ベースラインに対して HG-RAG を評価しました。結果は、HG-RAG が、幻覚を軽減し、局所性の一貫性を維持しながら、階層的、リレーショナル、およびマルチホップ推論タスクにおいて平坦なベースラインを常に上回っていることを示しています。

原文 (English)

HG-RAG: Hierarchy-Guided Retrieval-Augmented Generation for Structured Knowledge Graphs

Retrieval Augmented Generation (RAG) has proven to be a widely successful process at improving the quality of outputs from a Large Language Model (LLM) for wider context. However, RAG systems typically retrieve context from flat document stores, which struggles when queries require hierarchical or relational reasoning across structured knowledge. I present HG-RAG (Hierarchy-Guided RAG), a framework that performs graph-traversal over a hierarchical knowledge graph to deliver structured context to a language model. My retrieval pipeline resolves a named entity anchor from the query, then expands context upward through parent nodes, laterally through relational neighbors, and downward through child nodes when needed. I evaluate HG-RAG against a dense retrieval baseline across three world scales (18-800 nodes) with four query types: local fact, hierarchical, neighborhood, and multi-hop. Results show HG-RAG consistently outperforms the flat baseline on hierarchical, relational, and multi-hop reasoning tasks, while reducing hallucination and maintaining locality coherence.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

経験的なネクストトークン分布を使用したトレーニング データに対する LLM の動作の追跡

この論文では、LLM の出力分布とそのトレーニングに使用されるデータとの関係を研究します。具体的には、トレーニング データのコンテキストを考慮して、LLM の次トークン分布が経験的な次トークン分布 (ENTD) とどの程度一致するかを研究します。 ENTD は、事前トレーニングに使用されるネクスト トークン クロス エントロピー損失の無制限のグローバル ミニマイザーであり、事前トレーニング コーパスの容易に解釈可能な関数であるため、魅力的なターゲットです。入力のかなりの部分について、LLM の分布が ENTD とほぼ完全に一致し、平均一致度はモデルのスケールとトレーニングの計算に応じて増加することがわかりました。それにもかかわらず、LLM と ENTD が大きく異なる入力シーケンスのロングテールが存在するため、トランスのアーキテクチャ、トレーニング手順、および ENTD 推定自体の有限サンプル ノイズにわたるこの不一致の考えられる原因をいくつか調べます。より広範には、私たちの調査結果が、モデルの動作が学習された重みにどのようにエンコードされるかではなく、データからどのように生じるかというブラックボックスを開く、標準的なメカニズムの解釈可能性を補完する「データ中心のメカニズムの解釈可能性」に関するさらなる研究を促進することを願っています。

原文 (English)

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions

In this paper, we study the connection between an LLM's output distribution and the data used to train it. Specifically, we study the degree to which an LLM's next-token distribution agrees with the empirical next-token distribution (ENTD) given the context in the training data. The ENTD is an appealing target because it is the unrestricted global minimizer of the next-token cross entropy loss used for pretraining, as well as an easily interpretable function of the pretraining corpus. We find that for a significant fraction of inputs, the LLM's distribution agrees with the ENTD almost perfectly, and the average agreement increases with model scale and training compute. Nevertheless, there is a long tail of input sequences where the LLM and ENTD differ significantly, and we examine several possible sources of this discrepancy across the transformer architecture, training procedure, and finite-sample noise in the ENTD estimate itself. More broadly, we hope our findings will encourage more work on ``data-centric mechanistic interpretability,'' a complement to standard mechanistic interpretability that opens the black box of how model behaviors arise from the data, rather than how they are encoded in the learned weights.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体研究/論文

LLM で生成された GPU カーネルは本番環境に対応していますか?トレース駆動のベンチマークと最適化エージェント

既存の GPU カーネル生成ベンチマークは、デプロイされたワークロードから分岐した合成ソースまたは厳選されたソースから問題を引き出します。 Atrex-Bench は、コンピューティングが制限され、メモリが豊富な GPU のフルクラスターのプロダクション推論トレースから直接サンプリングされた 30 のオペレーターと 440 のシェイプを備えたベンチマークです。各問題には、観測された GPU 時間のシェアから導出される重要度の重みがあり、アプリケーションのカード時間によって重み付けされ、問題ごとのルーフラインの上限とともに、問題が実行されるサービス提供フェーズごとに個別に計算されます。そのため、集計スコアでは、最も多くのサービス時間を消費するカーネルが強調されます。 Atrex-Bench で 6 つのフロンティア コーディング エージェントを評価すると、最高のバニラ モデルであっても、運用オペレーターのハードウェア ルーフラインの ${\sim}10\%$ にしか達していないことがわかります。また、見かけの合格率の多くは、モデルが作成したカーネルではなく PyTorch フォールバックから得られるため、正確性だけが機能を誇張しています。このギャップを埋めるために、Atrex-Kernel-Agent (AKA) を共同リリースします。Atrex-Kernel-Agent (AKA) は、反復的な測定改訂検索、停止した検索コンテキストをエスケープするための最適化ドロップアウト、および階層化された GPU 最適化ナレッジ ベース (298 のリファレンス カーネル ファイルと 244 の最適化ナレッジ ドキュメント、および API/ISA ルックアップ用の外部アップストリーム リファレンス プロジェクト) を組み合わせたプロファイル駆動型のカーネル最適化エージェントです。制御されたケーススタディでは、エージェントはゼロ FlyDSL フォールバックを、手動で調整された運用ベースラインと一致または超える実際のカーネルに変換します。

原文 (English)

Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent

Existing GPU kernel generation benchmarks draw problems from synthetic or curated sources that diverge from deployed workloads. We present Atrex-Bench, a benchmark whose 30 operators and 440 shapes are sampled directly from full-cluster production inference traces of compute-limited, memory-rich GPUs. Each problem carries an importance weight derived from its share of observed GPU time, weighted by application card-hours and computed separately for the serving phases in which it runs, together with a per-problem roofline ceiling, so the aggregate score emphasizes the kernels that consume the most serving time. Evaluating six frontier coding agents on Atrex-Bench shows that even the best vanilla model reaches only ${\sim}10\%$ of the hardware roofline on production operators; and correctness alone overstates capability, since much of the apparent pass rate comes from PyTorch fallbacks rather than kernels the model wrote. To close this gap, we co-release Atrex-Kernel-Agent (AKA), a profile-driven kernel-optimization agent that combines iterative measure-revise search, optimization dropout for escaping stalled search contexts, and a layered GPU-optimization knowledge base (298 reference-kernel files and 244 optimization-knowledge documents, plus external upstream reference projects for API/ISA lookup). In a controlled case study, the agent converts zero-FlyDSL fallbacks into real kernels that match or exceed hand-tuned production baselines.

2026-07-18 13:00 JSTarXiv cs.AIハードウェア/半導体

6 GB 2011 GPU 上の最新のマルチモーダル アシスタント: 段階的に検証された、Fermi 用のオール GPU CUDA 推論

関連調査では、4 ビット モデルがデバイス メモリに収まらなかったため、2011 NVIDIA Tesla C2075 (Fermi、sm_20、6GB) で 35B の専門家混合モデルを GPU プリフィル/CPU デコード ハイブリッドとして実行しました (arXiv:2606.24031)。このレポートでは、ハードウェアを維持し、適合するモデルで何ができるかを問います。MiniCPM-V-4.6 は、SigLIP2 ビジョン エンコーダとウィンドウ アテンション マージャ (16 倍のビジュアル トークン圧縮) を、コンパクトなハイブリッド ゲート デルタ ネット バックボーンと組み合わせた最新のマルチモーダル アシスタントで、完全に GPU 上に展開されます。結果は3つ。 (i) 測定された基盤に基づいて構築された全 GPU エンジン: 8 ビットの重みを一度逆量子化し、最後の Fermi ツールチェーンにまだあるベンダー SGEMM を呼び出す予測 (FP32 ピークの 64%、最高の手書き GEMM は 37% に達し、誤って天井と呼ばれました)。再帰層のチャンク化されたデルタルールの書き換え。アトリビューションにより 1 つの不良カーネルが明らかになると、シーケンシャル スキャンよりも 2.8 倍高速になります。 Fermi がニブルアンパッキング シフトをハーフ レートで発行するため、ここでは 4 ビットの重みによりデコードが 8 ビットよりも遅くなります。 (ii) ビジョン側は証明義務のあるポートです。タワー、マージャー、およびプロジェクターを sm_20 CUDA に変換し、ローカルで生成されたリファレンス フォワード (フル タワー 1.4e-5) に対して各ステージを検証します。失敗の 1 つは、位置埋め込みのバケット化が厳密な有理同順位で異なることです。これは、あるルールに一般化されます。インデックス演算における浮動小数点数の同順位ブレークは実装定義です。参照演算子を呼び出します。再実装しないでください。 (iii) 長いコンテキストにより、O(N^2) ウォールの短いベンチマークが隠蔽されます。単純な注意カーネルでは、プリフィルは 2k トークンでの 114 tok/s から 10k での 21 tok/s に低下します。ヘッドごとのベンダー GEMM 呼び出しは、既存のスコア バッファー (追加メモリゼロ) への書き込みにより、フラット プロファイル (2k で 408、10k で 361、17x) を復元し、深さ 60% からの正確なニードル検索によって検証されます。同じ書き換えにより、画像エンコーディングが 6 倍の 0.93 秒に短縮されます。システムは画像の質問に 1.7 秒でエンドツーエンドで回答します。

原文 (English)

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.24031). This report keeps the hardware and asks what a model that fits can do: we deploy MiniCPM-V-4.6, a modern multimodal assistant pairing a SigLIP2 vision encoder and window-attention merger (16x visual token compression) with a compact hybrid gated-delta-net backbone, entirely on the GPU. Three results. (i) An all-GPU engine built on measured foundations: projections that dequantize 8-bit weights once and call the vendor SGEMM still in the last Fermi toolchain (64% of FP32 peak; our best hand-written GEMM hit 37%, wrongly called the ceiling); a chunked delta-rule rewrite of the recurrent layers, 2.8x faster than the sequential scan once attribution exposed one bad kernel; and a measured negative: 4-bit weights make decode slower than 8-bit here, since Fermi issues nibble-unpacking shifts at half rate. (ii) The vision side is a port with a proof obligation: we translate tower, merger, and projector to sm_20 CUDA, validating every stage against a locally generated reference forward (full tower 1.4e-5). One failure, position-embedding bucketization differing on exact rational ties, generalizes to a rule: float tie-breaking in index arithmetic is implementation-defined; call the reference operator, do not reimplement it. (iii) Long context exposes an O(N^2) wall short benchmarks hide: prefill falls from 114 tok/s at 2k tokens to 21 at 10k in a naive attention kernel; per-head vendor-GEMM calls writing into the existing score buffer (zero extra memory) restore a flat profile (408 at 2k, 361 at 10k; 17x), verified by exact needle retrieval from 60% depth. The same rewrite cuts image encoding 6x, to 0.93s. The system answers an image question end-to-end in 1.7s.

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

SAMark: 段落レベルの言い換え堅牢性を備えた自己アンカー付きテキスト透かし

意味レベルの透かし (SWM) は、文を基本単位として扱うことで、テキストの変更に対する堅牢性を向上させます。ただし、このような攻撃は文の順序を変更することで透かし信号を全体的に破壊するため、段落レベルの言い換えに対する堅牢性は依然として困難です。この研究では、意味空間にステップに依存しない緑色の領域を確立することで文の順序への依存を取り除く、自己アンカー型透かしフレームワークである SAMark を提案します。検出可能性を向上させるために、弱く位置合わせされた候補からのノイズを抑制しながら透かし信号を増幅するマルチチャネル双曲線スコアリング メカニズムを導入します。さらに、ハード フィルタリングとソフト正則化を組み合わせた多様性を意識したフィルタリング戦略を提案し、単純な N グラム繰り返しフィルタを超えて意味上の冗長性に対処します。実験結果は、SAMark が典型的な段落レベルの言い換え攻撃の下で最大 90.2% の TP@FP1% を達成し、以前の最も強力なベースラインを平均 30% 以上上回るパフォーマンスを示しながら、透かしなしのテキストと競争力のある生成品質を維持し、従来の方法を制限していた堅牢性と品質のトレードオフを打破することを示しています。

原文 (English)

SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness

Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to paragraph-level paraphrasing remains difficult because such attacks globally disrupt watermark signals by changing sentence order. In this work, we propose SAMark, a self-anchored watermarking framework that removes the dependency on sentence order by establishing a step-independent green region in semantic space. To improve detectability, we introduce a multi-channel hyperbolic scoring mechanism that amplifies watermark signals while suppressing noise from weakly aligned candidates. We further propose a diversity-aware filtering strategy that combines hard filtering with soft regularization, extending beyond simple n-gram repetition filters to address semantic redundancy. Experimental results show that SAMark achieves up to 90.2% TP@FP1% under typical paragraph-level paraphrasing attacks, outperforming the strongest prior baseline by more than 30% on average, while maintaining generation quality competitive with unwatermarked text and breaking the robustness-quality trade-off that limits prior methods. Our code will be released at [this URL](https://github.com/Z1zs/SAMark).

2026-07-18 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Workload-Driven Optimization for On-Device Real-Time Subtitle Translation

This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-…

2026-07-18 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体ビジネス/資金調達

ブラックボックスを制御できるか?協調エージェントを使用したレコメンダーシステムの制御性中心の評価に向けて

レコメンダー システムはブラック ボックスとして動作するため、ユーザーや規制当局は出力を特定の意図に向けたり、その動作を監査したりすることができません。この制御性の欠如は、明示的なガイダンスに応答するシステムの能力として定義されますが、既存の評価パラダイムでは依然として対処されていない側面です。このギャップを埋めるために、制御性を体系的に評価するための協調的なマルチエージェント フレームワークである CtrlBench-Rec を提案します。私たちは、ターゲット コンテンツの発見、関心プロファイルの形成、人気バイアスの軽減という 3 つの基本的なタスクを形式化します。これらは、明示的なコマンドから暗黙的な表現のステアリング、そして最終的にはアルゴリズムのバイアスの克服までのステアビリティを一緒に測定します。実世界のデータセットと複数のレコメンデーション モデルに関する広範な実験により、私たちのフレームワークが制御性を効果的に定量化し、重大なシステムのボトルネック、特に誘導ロングテール コンテンツに対する永続的な抵抗を明らかにすることが実証されています。 CtrlBench-Rec は、制御可能な推奨調査、アルゴリズム監査、およびユーザー権限付与のための初の標準化されたツールキットを提供します。私たちのコードは https://github.com/caskcsg/CtrlBenchRec でリリースされています。

原文 (English)

Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents

Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, remains an unaddressed dimension in existing evaluation paradigms. To fill this gap, we propose CtrlBench-Rec, a collaborative multi-agent framework for systematic assessment of controllability. We formalize three fundamental tasks: target content discovery, interest profile shaping, and popularity bias mitigation, which together measure steerability from explicit commands to implicit representation steering and finally to overcoming algorithmic biases.Extensive experiments on real-world datasets and multiple recommendation models demonstrate that our framework effectively quantifies controllability and exposes critical system bottlenecks, most notably persistent resistance to guiding long tail content. CtrlBench-Rec provides the first standardized toolkit for controllable recommendation research, algorithmic auditing, and user empowerment. Our code is released on https://github.com/caskcsg/CtrlBenchRec.

2026-07-17 21:00 JSTTechCrunch AIハードウェア/半導体

Why the first GPU financiers are turning to inference chips in a $400 million deal

A $400 million chip-backed loan points to the next wave of AI infrastructure deals.

2026-07-17 13:00 JSTarXiv cs.AIハードウェア/半導体

HG-RAG: 構造化ナレッジ グラフの階層に基づく検索拡張生成

検索拡張生成 (RAG) は、より広範なコンテキストに対応する大規模言語モデル (LLM) からの出力の品質を向上させる上で広く成功しているプロセスであることが証明されています。ただし、RAG システムは通常、フラットなドキュメント ストアからコンテキストを取得するため、クエリで構造化された知識全体にわたる階層的推論やリレーショナル推論が必要な場合に困難を伴います。私は HG-RAG (Hierarchy-Guided RAG) を紹介します。これは、階層ナレッジ グラフ上でグラフ トラバーサルを実行し、構造化されたコンテキストを言語モデルに提供するフレームワークです。私の取得パイプラインは、クエリから名前付きエンティティ アンカーを解決し、必要に応じて、親ノードを介して上方向に、リレーショナル隣接ノードを介して横方向に、そして子ノードを介して下方向にコンテキストを拡張します。私は、ローカル ファクト、階層、近傍、およびマルチホップの 4 つのクエリ タイプを使用して、3 つの世界スケール (18 ~ 800 ノード) にわたる高密度検索ベースラインに対して HG-RAG を評価しました。結果は、HG-RAG が、幻覚を軽減し、局所性の一貫性を維持しながら、階層的、リレーショナル、およびマルチホップ推論タスクにおいて平坦なベースラインを常に上回っていることを示しています。

原文 (English)

HG-RAG: Hierarchy-Guided Retrieval-Augmented Generation for Structured Knowledge Graphs

Retrieval Augmented Generation (RAG) has proven to be a widely successful process at improving the quality of outputs from a Large Language Model (LLM) for wider context. However, RAG systems typically retrieve context from flat document stores, which struggles when queries require hierarchical or relational reasoning across structured knowledge. I present HG-RAG (Hierarchy-Guided RAG), a framework that performs graph-traversal over a hierarchical knowledge graph to deliver structured context to a language model. My retrieval pipeline resolves a named entity anchor from the query, then expands context upward through parent nodes, laterally through relational neighbors, and downward through child nodes when needed. I evaluate HG-RAG against a dense retrieval baseline across three world scales (18-800 nodes) with four query types: local fact, hierarchical, neighborhood, and multi-hop. Results show HG-RAG consistently outperforms the flat baseline on hierarchical, relational, and multi-hop reasoning tasks, while reducing hallucination and maintaining locality coherence.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

経験的なネクストトークン分布を使用したトレーニング データに対する LLM の動作の追跡

この論文では、LLM の出力分布とそのトレーニングに使用されるデータとの関係を研究します。具体的には、トレーニング データのコンテキストを考慮して、LLM の次トークン分布が経験的な次トークン分布 (ENTD) とどの程度一致するかを研究します。 ENTD は、事前トレーニングに使用されるネクスト トークン クロス エントロピー損失の無制限のグローバル ミニマイザーであり、事前トレーニング コーパスの容易に解釈可能な関数であるため、魅力的なターゲットです。入力のかなりの部分について、LLM の分布が ENTD とほぼ完全に一致し、平均一致度はモデルのスケールとトレーニングの計算に応じて増加することがわかりました。それにもかかわらず、LLM と ENTD が大きく異なる入力シーケンスのロングテールが存在するため、トランスのアーキテクチャ、トレーニング手順、および ENTD 推定自体の有限サンプル ノイズにわたるこの不一致の考えられる原因をいくつか調べます。より広範には、私たちの調査結果が、モデルの動作が学習された重みにどのようにエンコードされるかではなく、データからどのように生じるかというブラックボックスを開く、標準的なメカニズムの解釈可能性を補完する「データ中心のメカニズムの解釈可能性」に関するさらなる研究を促進することを願っています。

原文 (English)

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions

In this paper, we study the connection between an LLM's output distribution and the data used to train it. Specifically, we study the degree to which an LLM's next-token distribution agrees with the empirical next-token distribution (ENTD) given the context in the training data. The ENTD is an appealing target because it is the unrestricted global minimizer of the next-token cross entropy loss used for pretraining, as well as an easily interpretable function of the pretraining corpus. We find that for a significant fraction of inputs, the LLM's distribution agrees with the ENTD almost perfectly, and the average agreement increases with model scale and training compute. Nevertheless, there is a long tail of input sequences where the LLM and ENTD differ significantly, and we examine several possible sources of this discrepancy across the transformer architecture, training procedure, and finite-sample noise in the ENTD estimate itself. More broadly, we hope our findings will encourage more work on ``data-centric mechanistic interpretability,'' a complement to standard mechanistic interpretability that opens the black box of how model behaviors arise from the data, rather than how they are encoded in the learned weights.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体研究/論文

LLM で生成された GPU カーネルは本番環境に対応していますか?トレース駆動のベンチマークと最適化エージェント

既存の GPU カーネル生成ベンチマークは、デプロイされたワークロードから分岐した合成ソースまたは厳選されたソースから問題を引き出します。 Atrex-Bench は、コンピューティングが制限され、メモリが豊富な GPU のフルクラスターのプロダクション推論トレースから直接サンプリングされた 30 のオペレーターと 440 のシェイプを備えたベンチマークです。各問題には、観測された GPU 時間のシェアから導出される重要度の重みがあり、アプリケーションのカード時間によって重み付けされ、問題ごとのルーフラインの上限とともに、問題が実行されるサービス提供フェーズごとに個別に計算されます。そのため、集計スコアでは、最も多くのサービス時間を消費するカーネルが強調されます。 Atrex-Bench で 6 つのフロンティア コーディング エージェントを評価すると、最高のバニラ モデルであっても、運用オペレーターのハードウェア ルーフラインの ${\sim}10\%$ にしか達していないことがわかります。また、見かけの合格率の多くは、モデルが作成したカーネルではなく PyTorch フォールバックから得られるため、正確性だけが機能を誇張しています。このギャップを埋めるために、Atrex-Kernel-Agent (AKA) を共同リリースします。Atrex-Kernel-Agent (AKA) は、反復的な測定改訂検索、停止した検索コンテキストをエスケープするための最適化ドロップアウト、および階層化された GPU 最適化ナレッジ ベース (298 のリファレンス カーネル ファイルと 244 の最適化ナレッジ ドキュメント、および API/ISA ルックアップ用の外部アップストリーム リファレンス プロジェクト) を組み合わせたプロファイル駆動型のカーネル最適化エージェントです。制御されたケーススタディでは、エージェントはゼロ FlyDSL フォールバックを、手動で調整された運用ベースラインと一致または超える実際のカーネルに変換します。

原文 (English)

Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent

Existing GPU kernel generation benchmarks draw problems from synthetic or curated sources that diverge from deployed workloads. We present Atrex-Bench, a benchmark whose 30 operators and 440 shapes are sampled directly from full-cluster production inference traces of compute-limited, memory-rich GPUs. Each problem carries an importance weight derived from its share of observed GPU time, weighted by application card-hours and computed separately for the serving phases in which it runs, together with a per-problem roofline ceiling, so the aggregate score emphasizes the kernels that consume the most serving time. Evaluating six frontier coding agents on Atrex-Bench shows that even the best vanilla model reaches only ${\sim}10\%$ of the hardware roofline on production operators; and correctness alone overstates capability, since much of the apparent pass rate comes from PyTorch fallbacks rather than kernels the model wrote. To close this gap, we co-release Atrex-Kernel-Agent (AKA), a profile-driven kernel-optimization agent that combines iterative measure-revise search, optimization dropout for escaping stalled search contexts, and a layered GPU-optimization knowledge base (298 reference-kernel files and 244 optimization-knowledge documents, plus external upstream reference projects for API/ISA lookup). In a controlled case study, the agent converts zero-FlyDSL fallbacks into real kernels that match or exceed hand-tuned production baselines.

2026-07-17 13:00 JSTarXiv cs.AIハードウェア/半導体

6 GB 2011 GPU 上の最新のマルチモーダル アシスタント: 段階的に検証された、Fermi 用のオール GPU CUDA 推論

関連調査では、4 ビット モデルがデバイス メモリに収まらなかったため、2011 NVIDIA Tesla C2075 (Fermi、sm_20、6GB) で 35B の専門家混合モデルを GPU プリフィル/CPU デコード ハイブリッドとして実行しました (arXiv:2606.24031)。このレポートでは、ハードウェアを維持し、適合するモデルで何ができるかを問います。MiniCPM-V-4.6 は、SigLIP2 ビジョン エンコーダとウィンドウ アテンション マージャ (16 倍のビジュアル トークン圧縮) を、コンパクトなハイブリッド ゲート デルタ ネット バックボーンと組み合わせた最新のマルチモーダル アシスタントで、完全に GPU 上に展開されます。結果は3つ。 (i) 測定された基盤に基づいて構築された全 GPU エンジン: 8 ビットの重みを一度逆量子化し、最後の Fermi ツールチェーンにまだあるベンダー SGEMM を呼び出す予測 (FP32 ピークの 64%、最高の手書き GEMM は 37% に達し、誤って天井と呼ばれました)。再帰層のチャンク化されたデルタルールの書き換え。アトリビューションにより 1 つの不良カーネルが明らかになると、シーケンシャル スキャンよりも 2.8 倍高速になります。 Fermi がニブルアンパッキング シフトをハーフ レートで発行するため、ここでは 4 ビットの重みによりデコードが 8 ビットよりも遅くなります。 (ii) ビジョン側は証明義務のあるポートです。タワー、マージャー、およびプロジェクターを sm_20 CUDA に変換し、ローカルで生成されたリファレンス フォワード (フル タワー 1.4e-5) に対して各ステージを検証します。失敗の 1 つは、位置埋め込みのバケット化が厳密な有理同順位で異なることです。これは、あるルールに一般化されます。インデックス演算における浮動小数点数の同順位ブレークは実装定義です。参照演算子を呼び出します。再実装しないでください。 (iii) 長いコンテキストにより、O(N^2) ウォールの短いベンチマークが隠蔽されます。単純な注意カーネルでは、プリフィルは 2k トークンでの 114 tok/s から 10k での 21 tok/s に低下します。ヘッドごとのベンダー GEMM 呼び出しは、既存のスコア バッファー (追加メモリゼロ) への書き込みにより、フラット プロファイル (2k で 408、10k で 361、17x) を復元し、深さ 60% からの正確なニードル検索によって検証されます。同じ書き換えにより、画像エンコーディングが 6 倍の 0.93 秒に短縮されます。システムは画像の質問に 1.7 秒でエンドツーエンドで回答します。

原文 (English)

A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi

A companion study ran a 35B mixture-of-experts model on a 2011 NVIDIA Tesla C2075 (Fermi, sm_20, 6GB) as a GPU-prefill/CPU-decode hybrid, because the 4-bit model did not fit in device memory (arXiv:2606.24031). This report keeps the hardware and asks what a model that fits can do: we deploy MiniCPM-V-4.6, a modern multimodal assistant pairing a SigLIP2 vision encoder and window-attention merger (16x visual token compression) with a compact hybrid gated-delta-net backbone, entirely on the GPU. Three results. (i) An all-GPU engine built on measured foundations: projections that dequantize 8-bit weights once and call the vendor SGEMM still in the last Fermi toolchain (64% of FP32 peak; our best hand-written GEMM hit 37%, wrongly called the ceiling); a chunked delta-rule rewrite of the recurrent layers, 2.8x faster than the sequential scan once attribution exposed one bad kernel; and a measured negative: 4-bit weights make decode slower than 8-bit here, since Fermi issues nibble-unpacking shifts at half rate. (ii) The vision side is a port with a proof obligation: we translate tower, merger, and projector to sm_20 CUDA, validating every stage against a locally generated reference forward (full tower 1.4e-5). One failure, position-embedding bucketization differing on exact rational ties, generalizes to a rule: float tie-breaking in index arithmetic is implementation-defined; call the reference operator, do not reimplement it. (iii) Long context exposes an O(N^2) wall short benchmarks hide: prefill falls from 114 tok/s at 2k tokens to 21 at 10k in a naive attention kernel; per-head vendor-GEMM calls writing into the existing score buffer (zero extra memory) restore a flat profile (408 at 2k, 361 at 10k; 17x), verified by exact needle retrieval from 60% depth. The same rewrite cuts image encoding 6x, to 0.93s. The system answers an image question end-to-end in 1.7s.

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

SAMark: 段落レベルの言い換え堅牢性を備えた自己アンカー付きテキスト透かし

意味レベルの透かし (SWM) は、文を基本単位として扱うことで、テキストの変更に対する堅牢性を向上させます。ただし、このような攻撃は文の順序を変更することで透かし信号を全体的に破壊するため、段落レベルの言い換えに対する堅牢性は依然として困難です。この研究では、意味空間にステップに依存しない緑色の領域を確立することで文の順序への依存を取り除く、自己アンカー型透かしフレームワークである SAMark を提案します。検出可能性を向上させるために、弱く位置合わせされた候補からのノイズを抑制しながら透かし信号を増幅するマルチチャネル双曲線スコアリング メカニズムを導入します。さらに、ハード フィルタリングとソフト正則化を組み合わせた多様性を意識したフィルタリング戦略を提案し、単純な N グラム繰り返しフィルタを超えて意味上の冗長性に対処します。実験結果は、SAMark が典型的な段落レベルの言い換え攻撃の下で最大 90.2% の TP@FP1% を達成し、以前の最も強力なベースラインを平均 30% 以上上回るパフォーマンスを示しながら、透かしなしのテキストと競争力のある生成品質を維持し、従来の方法を制限していた堅牢性と品質のトレードオフを打破することを示しています。

原文 (English)

SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness

Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to paragraph-level paraphrasing remains difficult because such attacks globally disrupt watermark signals by changing sentence order. In this work, we propose SAMark, a self-anchored watermarking framework that removes the dependency on sentence order by establishing a step-independent green region in semantic space. To improve detectability, we introduce a multi-channel hyperbolic scoring mechanism that amplifies watermark signals while suppressing noise from weakly aligned candidates. We further propose a diversity-aware filtering strategy that combines hard filtering with soft regularization, extending beyond simple n-gram repetition filters to address semantic redundancy. Experimental results show that SAMark achieves up to 90.2% TP@FP1% under typical paragraph-level paraphrasing attacks, outperforming the strongest prior baseline by more than 30% on average, while maintaining generation quality competitive with unwatermarked text and breaking the robustness-quality trade-off that limits prior methods. Our code will be released at [this URL](https://github.com/Z1zs/SAMark).

2026-07-17 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Workload-Driven Optimization for On-Device Real-Time Subtitle Translation

This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-…

2026-07-17 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体ビジネス/資金調達

ブラックボックスを制御できるか?協調エージェントを使用したレコメンダーシステムの制御性中心の評価に向けて

レコメンダー システムはブラック ボックスとして動作するため、ユーザーや規制当局は出力を特定の意図に向けたり、その動作を監査したりすることができません。この制御性の欠如は、明示的なガイダンスに応答するシステムの能力として定義されますが、既存の評価パラダイムでは依然として対処されていない側面です。このギャップを埋めるために、制御性を体系的に評価するための協調的なマルチエージェント フレームワークである CtrlBench-Rec を提案します。私たちは、ターゲット コンテンツの発見、関心プロファイルの形成、人気バイアスの軽減という 3 つの基本的なタスクを形式化します。これらは、明示的なコマンドから暗黙的な表現のステアリング、そして最終的にはアルゴリズムのバイアスの克服までのステアビリティを一緒に測定します。実世界のデータセットと複数のレコメンデーション モデルに関する広範な実験により、私たちのフレームワークが制御性を効果的に定量化し、重大なシステムのボトルネック、特に誘導ロングテール コンテンツに対する永続的な抵抗を明らかにすることが実証されています。 CtrlBench-Rec は、制御可能な推奨調査、アルゴリズム監査、およびユーザー権限付与のための初の標準化されたツールキットを提供します。私たちのコードは https://github.com/caskcsg/CtrlBenchRec でリリースされています。

原文 (English)

Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents

Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, remains an unaddressed dimension in existing evaluation paradigms. To fill this gap, we propose CtrlBench-Rec, a collaborative multi-agent framework for systematic assessment of controllability. We formalize three fundamental tasks: target content discovery, interest profile shaping, and popularity bias mitigation, which together measure steerability from explicit commands to implicit representation steering and finally to overcoming algorithmic biases.Extensive experiments on real-world datasets and multiple recommendation models demonstrate that our framework effectively quantifies controllability and exposes critical system bottlenecks, most notably persistent resistance to guiding long tail content. CtrlBench-Rec provides the first standardized toolkit for controllable recommendation research, algorithmic auditing, and user empowerment. Our code is released on https://github.com/caskcsg/CtrlBenchRec.

2026-07-17 03:22 JSTTechCrunch AILLM/生成AIハードウェア/半導体

Roblox launches an AI-powered game-creation feature in its mobile app

Roblox's new "Build" feature lets users generate basic games using a single text prompt.

2026-07-16 19:43 JSTITmedia AI+ロボティクスハードウェア/半導体

富士通、国内ロボット大手3社と「フィジカルAI」で協業 NVIDIAの技術活用

富士通は、AIが自律的に考え、ロボットの体を動かす「フィジカルAI」の開発に関し、川崎重工業とファナック、安川電機の各社と協業すると発表した。米NVIDIAの技術を活用し、ロボットを協調的に制御するための基盤を開発する。

2026-07-16 19:07 JSTITmedia AI+ハードウェア/半導体

大手共同出資の“国産AI開発企業”が本格始動 NVIDIAも協力、「Rubin」2万7500基搭載の計算基盤を構築へ

国内大手が共同出資するAI開発企業Noetraが、国産のAIモデルの開発に向けて本格始動する。米NVIDIAの協力のもと、新たな計算基盤も構築する。

2026-07-16 17:15 JSTITmedia AI+ハードウェア/半導体

トンカツ食べながら語った――NVIDIA、富士通、安川電機ら“フィジカルAI連合”誕生、発表直前の裏話

富士通、ファナック、安川電機、川崎重工業とNVIDIAによる“フィジカルAI連合”が誕生した。5社のトップは、記者説明会の直前には「トンカツ」を食べながら語り合ったという。その一幕を富士通の時田社長が明かした。

2026-07-16 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

ボットがチームに加わるとき: ボットの導入とオープンソース ソフトウェア プロジェクトの組織構造

AI エージェントが人間のチームに加わることで、基本的な疑問が生じています。自動化されたエージェントが通常の参加者になったとき、グループ組織は強化されるのでしょうか、それとも弱まるのでしょうか?私たちはこの問題をオープンソース ソフトウェアで研究しています。オープンソース ソフトウェアでは、ボットがプル リクエストをオープンし、コードをレビューし、人々と一緒に変更をマージし、すべてのやり取りの公開記録を残します。ボットをツールではなく参加者として扱い、2,991 の GitHub プロジェクトを、それぞれが最初のボットを採用する前後の 2 年間調査しました。私たちは、制度理論が永続的な調整に結びつける 3 つの能力 (繰り返しの関与、社会的記憶、役割の分化) と、2 つの結果 (紛争カスケードと成果の独自性) を測定します。ボットを導入すると、コラボレーションがさらに繰り返され、議論の中で特定のボットがよりよく認識されるようになり、紛争のカスケードが減り、より特徴的な成果が得られます。これらの変更は徐々に蓄積されるのではなく、導入を中心に集中しています。未処理の比較グループがないため、結果は因果関係ではなく、正確にタイミングを合わせた関連であると解釈します。別の説明で説明するのが難しい 2 つのパターンは、能力が人間かボットのどちらが結果を提供するかではなく、その機能 (調整と差別化) に従って結果を予測すること、および人間側の能力はボットと競合の関連性を説明するが、ボットと区別性の関連性は説明しないことです。この発見は、予測可能なルールベースのエージェントがコミュニティの社会インフラの一部になり得るという特定の解釈と一致しています。ボットがそのチャンスです。社会組織がそのメカニズムです。

原文 (English)

When Bots Join the Team: Bot Adoption and the Institutional Fabric of Open-Source Software Projects

AI agents are joining human teams, raising a basic question: when an automated agent becomes a regular participant, does group organization strengthen or weaken? We study this question in open-source software, where bots open pull requests, review code, and merge changes alongside people, leaving a public record of every interaction. Treating bots as participants rather than tools, we examine 2,991 GitHub projects for two years before and after each adopted its first bot. We measure three capabilities that institutional theory links to durable coordination - repeated engagement, social memory, and role differentiation - and two outcomes: conflict cascades and output distinctiveness. Bot adoption is followed by more repeated collaboration, greater recognition of specific bots in discussion, fewer conflict cascades, and more distinctive outputs. These changes cluster around adoption rather than accumulating gradually. Because we lack an untreated comparison group, we interpret the results as precisely timed associations, not causal effects. Two patterns are difficult for alternative explanations to account for: capabilities predict outcomes according to their function - coordination versus differentiation - rather than whether humans or bots provide them, and human-side capabilities account for the bot-conflict association but not the bot-distinctiveness association. The findings are consistent with a specific interpretation: predictable, rule-based agents can become part of a community's social infrastructure. The bot is the occasion; social organization is the mechanism.

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

モノカルチャーへのヒッチハイク ガイド

大規模言語モデル (LLM) は同種の出力を生成することが多く、AI コーディング アシスタントが開発者が作成するソフトウェア アーティファクトの収束につながる可能性があるという懸念が生じています。開発者はモデルの出力を対話的にプロンプ​​ト、評価、変更、拒否するため、また出力はプロンプトやリポジトリのコンテキストによって異なるため、これが実際に発生するかどうかは不明です。 2019 年から 2026 年半ばまでの Kaggle コンテストの提出物を使用してコードの均一化を調査します。私は最初に、プログラミング文化における長年の慣習を強化する LLM と一致する、ランダム シード値 42 への広範な収束について文書化しました。次に、均質化を集約と抽象化の 2 つのレベルでより広範囲に研究します。提出物レベルでは、コンテスト内の提出物の平均ペアごとの類似性を測定します。コンテスト レベルでは、提出されたコードの概念的な範囲を測定し、それぞれについて明確な尺度を動機付けます。表面構文をキャプチャする TF-IDF 表現と、コードの意図とセマンティクスをキャプチャする Voyage 3 コード埋め込みです。結果は、個人レベルと集団レベルの両方で構文の実質的な均質化を示しています。つまり、個々の提出物はリテラル構文とコード構造においてより類似している一方で、構文のバリエーションの潜在的な次元は狭くなっています。対照的に、意味論的な均質化の証拠は、個別にも集合的にもほとんど見つかりません。平均意味論的距離は基本的に横ばいのままであり、意味論的アプローチのコンテストレベルの潜在的な次元範囲は安定したままであり、それがわずかに拡大したことを示唆する証拠さえあります。これらの調査結果は、AI コーディング アシスタントが実装の詳細を確実に標準化しているものの、コーダーが採用するアプローチや問題解決戦略が均質化しているという証拠はまだ得られていないことを示唆しています。

原文 (English)

The Hitchhiker's Guide to Monoculture

Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software artifacts that developers create. Whether this occurs in practice is unclear because developers interactively prompt, evaluate, modify, and reject model outputs, and because outputs vary with prompt and repository context. I examine code homogenization using Kaggle contest submissions from 2019 to mid-2026. I first document widespread convergence toward the random seed value 42, consistent with LLMs reinforcing a longstanding convention in programming culture. I then study homogenization more broadly, at two levels of aggregation and abstraction. At the submission level, I measure the average pairwise similarity of submissions within contests. At the contest level, I measure the conceptual span of submitted code, motivating distinct measures for each: TF-IDF representations, which capture surface syntax, and Voyage 3 code embeddings, which capture code intent and semantics. The results demonstrate substantial syntactic homogenization at both the individual and collective levels: individual submissions have become more alike in literal syntax and code structure, while the latent dimensionality of syntactic variation has narrowed. In contrast, I find little evidence of semantic homogenization, individually and collectively. Average semantic distance remains essentially flat, and the contest-level latent dimensional span of semantic approaches remains stable, with evidence suggesting it has even expanded modestly. These findings suggest that AI coding assistants are certainly standardizing implementation details, yet they have not yet produced evidence of homogenization in the approaches and problem-solving strategies coders employ.

2026-07-16 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体ビジネス/資金調達

ブラックボックスを制御できるか?協調エージェントを使用したレコメンダーシステムの制御性中心の評価に向けて

レコメンダー システムはブラック ボックスとして動作するため、ユーザーや規制当局は出力を特定の意図に向けたり、その動作を監査したりすることができません。この制御性の欠如は、明示的なガイダンスに応答するシステムの能力として定義されますが、既存の評価パラダイムでは依然として対処されていない側面です。このギャップを埋めるために、制御性を体系的に評価するための協調的なマルチエージェント フレームワークである CtrlBench-Rec を提案します。私たちは、ターゲット コンテンツの発見、関心プロファイルの形成、人気バイアスの軽減という 3 つの基本的なタスクを形式化します。これらは、明示的なコマンドから暗黙的な表現のステアリング、そして最終的にはアルゴリズムのバイアスの克服までのステアビリティを一緒に測定します。実世界のデータセットと複数のレコメンデーション モデルに関する広範な実験により、私たちのフレームワークが制御性を効果的に定量化し、重大なシステムのボトルネック、特に誘導ロングテール コンテンツに対する永続的な抵抗を明らかにすることが実証されています。 CtrlBench-Rec は、制御可能な推奨調査、アルゴリズム監査、およびユーザー権限付与のための初の標準化されたツールキットを提供します。私たちのコードは https://github.com/caskcsg/CtrlBenchRec でリリースされています。

原文 (English)

Can We Steer the Black-Box? Towards Controllability-Centric Evaluation of Recommender Systems with Collaborative Agents

Recommender systems operate as Black-Boxes, leaving users and regulators unable to steer their outputs toward specific intentions or audit their behavior. This lack of controllability, defined as the system's ability to respond to explicit guidance, remains an unaddressed dimension in existing evaluation paradigms. To fill this gap, we propose CtrlBench-Rec, a collaborative multi-agent framework for systematic assessment of controllability. We formalize three fundamental tasks: target content discovery, interest profile shaping, and popularity bias mitigation, which together measure steerability from explicit commands to implicit representation steering and finally to overcoming algorithmic biases.Extensive experiments on real-world datasets and multiple recommendation models demonstrate that our framework effectively quantifies controllability and exposes critical system bottlenecks, most notably persistent resistance to guiding long tail content. CtrlBench-Rec provides the first standardized toolkit for controllable recommendation research, algorithmic auditing, and user empowerment. Our code is released on https://github.com/caskcsg/CtrlBenchRec.

2026-07-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

必要なのは最適化だけではない

2019 年、OpenAI は、機械生成テキストの検出を支援するために、非文法的で半分壊れた 200 万個の GPT-2 出力をリリースしました。より流暢な後継者を生み出した連携は、通常、エンジニアリングの成果とみなされます。私たちはそれを、最適化文化の最新の表現として解釈します。つまり、テクノロジーよりも古い、事前に定義された軸に沿った測定可能な改善によって価値の問題が解決されるという信念です。その確信をスタック (事前トレーニング、デコード、プリファレンス調整、ベンチマーク、インターフェース) を通してたどり、監査協会の系譜をたどると、限界に到達します。最適化手順では、生成されたテキストの一部がどの程度ありそうもないかを測定できます。その可能性が誤りなのか発明なのかはわかりません。それにも関わらず、その区別ができない手順が、5 年以内に、正当な言語のプロトコルを設定する権限を引き継いだのです。何世紀にもわたってアカデミーや学校、文法学者や試験官によって保持されてきたこの権限は、損失関数、報酬モデル、ベンチマーク、およびシステムプロンプト、つまり判断能力のない判断官庁を実行する装置に譲渡されました。

原文 (English)

Optimization Is Not All You Need

In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering achievement; we read it instead as the newest expression of optimization culture: the conviction, older than the technology, that measurable improvement along predefined axes exhausts the question of value. Tracing that conviction through the stack-pretraining, decoding, preference tuning, benchmarking, interface-and back through its genealogy in the audit society, we arrive at the limit: an optimization procedure can measure how improbable a piece of generated text is; it cannot tell whether that unlikelihood is error or invention. A procedure that cannot make that distinction has nonetheless, within half a decade, assumed the authority to set the protocols of legitimate language. Held for centuries by academies and schoolrooms, grammars and examiners, this authority has been given over to loss functions, reward models, benchmarks, and system prompts: an apparatus that executes the office of judgment with no capacity for judging.

2026-07-16 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

パラメータ化された行動マルコフ決定プロセスのための知識と勾配に基づく強化学習

この論文では、パラメータ化されたアクションのマルコフ決定プロセス (PAMDP) における強化学習を研究します。このプロセスでは、各決定は記号アクションと数値パラメーターで構成されます。このような設定では、強化学習アルゴリズムは通常、ワンショット推定器を使用してパラメーターを決定するため、トレーニング サンプルが非効率になります。ほとんどの PAMDP 環境では、明示的ではあるが不完全な知識 (ルール、安全制約、エキスパートヒューリスティックなど) が利用可能ですが、それが強化学習エージェントのトレーニングのサンプル効率を高めるために直接使用されることはほとんどありません。私たちはこのギャップに踏み込み、新しい神経記号知識および勾配誘導強化学習 (KGRL) アルゴリズムを提案します。 KGRL は、Datalog 知識ベースのドメイン知識を使用して、特定の状態に適用可能なアクションと実行可能なパラメーターのセットを導き出します。これにより、適用できないアクションを決定空間から取り除き、残りのアクションのパラメータ空間を制約することができます。次に、勾配ベースのパラメータ調整ループを使用して、エージェントのトレーニングおよび展開中に最適なパラメータを推定します。 KGRL は、アクティブ化されたルールを軌跡に沿って記録することにより、アクションの刈り込みとパラメータの制約に関するローカルな手順の説明をさらに提供します。全体として、KGRL は、トレーニング中のサンプル効率を高めながら、エージェントの探索と展開を実行可能かつ制約を意識した決定に向けて導きます。 KGRL は、サンプル効率とエピソードリターンの両方において、PAMDP の最先端の RL ベースラインを上回ります。

原文 (English)

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes their training sample inefficient. Though in most PAMDP environments explicit but incomplete knowledge (e.g., rules, safety constraints, or expert heuristics) is available, it is rarely directly used to increase the sample-efficiency of training Reinforcement Learning agents. We step into this gap and propose our novel Neuro-Symbolic Knowledge- and Gradient-Guided Reinforcement Learning (KGRL) algorithm. KGRL uses domain knowledge in a Datalog knowledge base to derive the set of applicable actions and feasible parameters for a given state. This allows it to prune non-applicable actions from the decision-space and constrain the parameter spaces of the remaining actions. We then use a gradient-based parameter refinement loop to estimate the optimal parameters during training and deployment of the agent. By recording activated rules along the trajectory, KGRL additionally provides local procedural explanations on the pruning of actions and constraining of parameters. Overall, KGRL guides the agent's exploration and deployment toward feasible and constraint-aware decisions, while increasing sample efficiency during training. KGRL outperforms state-of-the-art RL baselines for PAMDPs in both, sample efficiency and episodic return.

2026-07-16 08:00 JSTITmedia AI+ハードウェア/半導体

NVIDIAが「Jetson Thor」に新モジュール追加、高騰するメモリの使用量削減技術も

NVIDIAは、組み込みAIボード「Jetsonシリーズ」の最新製品である「NVIDIA Jetson AGX Thor」の新たな量産モジュールとして、消費電力や搭載メモリ容量などを抑えた「Jetson T3000」と「Jetson T2000」を追加すると発表した。

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

GRID: エンタープライズ SQL 生成のための文法レール デコーディング

大規模な言語モデルでは SQL を記述することができますが、エンタープライズ展開では、もっともらしいテキスト以上のものが必要です。出力は構文的に有効である必要があり、ロールごとおよびスキーマごとのポリシーを尊重する必要があり、証明可能な (ベストエフォート型ではない) 保証を保持する必要があり、世代が成長するにつれて速度が低下してはならず、すべての決定についてコンプライアンス グレードの記録を残さなければなりません。我々は、トークンシーケンスではなくパーサー構成(レクサースキャン状態×LALR(1)スタック)で正確なネクストトークンマスクをキー設定する文法制約型デコードエンジンであるGRID(Grammar-Railed Decoding)を紹介し、段階的に高度なLALR(1)パーサー自体を実行可能なプレフィックスオラクルとして使用します。 LLM トークンは、コンテキスト独立/コンテキスト依存の分割を伴うバイトレベルのトライ ウォークによって文法端末にブリッジされ、キャッシュ キーの健全性が構築によって保持されます。役割ベースのアクセス制御は言語にコンパイルされます。役割投影は文法の生成のサブセットであり、スキーマ辞書は識別子の終端を制限するため、禁止された動詞と識別子はマスク レベルでは到達できません。 4 つの保証 (健全性、完全性、終了、トークンごとのほぼ一定のコスト) が明示的な前提条件とともに記載されており、それぞれがテストまたはベンチマークと対になっています。 Rust カーネルは、トークンごとのマスクを 3.6 ~ 6.7 us の中央値に設定し、誤拒否ゼロの 2 つのトークナイザーの p50 および p90 でのガイダンスを上回ります。トークンごとのガードコストは、n=16,000 でポジションフラットです。 Spider では、制約付きデコードは 0.5B で +13 実行精度ポイントの価値があり、マスク強制不可能であることが証明されている残留物 (列レベルのポリシー) に対するチェッカー ガイドに基づく修復パス 1 回により、7B モデルの実行率が 94.5% に引き上げられます。ハッシュチェーンされたトークンごとの監査証跡は、100% の改ざん検出でビット同一に再生されます。マスクで何ができないのか (配布の忠実性、列レベルの RBAC、非 LALR(1) 言語)、および測定されたコストがどこに残るのかを明確に述べます。

原文 (English)

GRID: Grammar-Railed Decoding for Enterprise SQL Generation

Large language models can write SQL, but enterprise deployment demands more than plausible text: outputs must be syntactically valid, must respect per-role and per-schema policy, must carry provable (not best-effort) guarantees, must not slow down as generations grow, and must leave a compliance-grade record of every decision. We present GRID (Grammar-Railed Decoding), a grammar-constrained decoding engine that keys exact next-token masks on parser configurations (lexer scan state x LALR(1) stack) rather than on token sequences, and uses the incrementally advanced LALR(1) parser itself as a viable-prefix oracle. LLM tokens are bridged to grammar terminals by a byte-level trie walk with a context-independent/context-dependent split that makes cache-key soundness hold by construction. Role-based access control is compiled into the language: role projections subset the grammar's productions and schema lexicons restrict identifier terminals, so forbidden verbs and identifiers are unreachable at mask level. Four guarantees (soundness, completeness, termination, and near-constant per-token cost) are stated with explicit preconditions and each paired with a test or benchmark. Rust kernels bring the per-token mask to a 3.6-6.7 us median, ahead of llguidance at p50 and p90 on two tokenizers with zero false rejects; per-token guard cost is position-flat at n=16,000. On Spider, constrained decoding is worth +13 execution-accuracy points at 0.5B, and one checker-guided repair pass over the provably mask-unenforceable residue (column-level policy) lifts a 7B model to 94.5% executable. A hash-chained per-token audit trail replays bit-identically with 100% tamper detection. We state plainly what the mask cannot do (distribution faithfulness, column-level RBAC, non-LALR(1) languages) and where measured cost remains.

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

必要なのは最適化だけではない

2019 年、OpenAI は、機械生成テキストの検出を支援するために、非文法的で半分壊れた 200 万個の GPT-2 出力をリリースしました。より流暢な後継者を生み出した連携は、通常、エンジニアリングの成果とみなされます。私たちはそれを、最適化文化の最新の表現として解釈します。つまり、テクノロジーよりも古い、事前に定義された軸に沿った測定可能な改善によって価値の問題が解決されるという信念です。その確信をスタック (事前トレーニング、デコード、プリファレンス調整、ベンチマーク、インターフェース) を通してたどり、監査協会の系譜をたどると、限界に到達します。最適化手順では、生成されたテキストの一部がどの程度ありそうもないかを測定できます。その可能性が誤りなのか発明なのかはわかりません。それにも関わらず、その区別ができない手順が、5 年以内に、正当な言語のプロトコルを設定する権限を引き継いだのです。何世紀にもわたってアカデミーや学校、文法学者や試験官によって保持されてきたこの権限は、損失関数、報酬モデル、ベンチマーク、およびシステムプロンプト、つまり判断能力のない判断官庁を実行する装置に譲渡されました。

原文 (English)

Optimization Is Not All You Need

In 2019, OpenAI released two million GPT-2 outputs-ungrammatical, half broken-to aid the detection of machine-generated text. The alignment that produced their more fluent successors is usually regarded as an engineering achievement; we read it instead as the newest expression of optimization culture: the conviction, older than the technology, that measurable improvement along predefined axes exhausts the question of value. Tracing that conviction through the stack-pretraining, decoding, preference tuning, benchmarking, interface-and back through its genealogy in the audit society, we arrive at the limit: an optimization procedure can measure how improbable a piece of generated text is; it cannot tell whether that unlikelihood is error or invention. A procedure that cannot make that distinction has nonetheless, within half a decade, assumed the authority to set the protocols of legitimate language. Held for centuries by academies and schoolrooms, grammars and examiners, this authority has been given over to loss functions, reward models, benchmarks, and system prompts: an apparatus that executes the office of judgment with no capacity for judging.

2026-07-15 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

パラメータ化された行動マルコフ決定プロセスのための知識と勾配に基づく強化学習

この論文では、パラメータ化されたアクションのマルコフ決定プロセス (PAMDP) における強化学習を研究します。このプロセスでは、各決定は記号アクションと数値パラメーターで構成されます。このような設定では、強化学習アルゴリズムは通常、ワンショット推定器を使用してパラメーターを決定するため、トレーニング サンプルが非効率になります。ほとんどの PAMDP 環境では、明示的ではあるが不完全な知識 (ルール、安全制約、エキスパートヒューリスティックなど) が利用可能ですが、それが強化学習エージェントのトレーニングのサンプル効率を高めるために直接使用されることはほとんどありません。私たちはこのギャップに踏み込み、新しい神経記号知識および勾配誘導強化学習 (KGRL) アルゴリズムを提案します。 KGRL は、Datalog 知識ベースのドメイン知識を使用して、特定の状態に適用可能なアクションと実行可能なパラメーターのセットを導き出します。これにより、適用できないアクションを決定空間から取り除き、残りのアクションのパラメータ空間を制約することができます。次に、勾配ベースのパラメータ調整ループを使用して、エージェントのトレーニングおよび展開中に最適なパラメータを推定します。 KGRL は、アクティブ化されたルールを軌跡に沿って記録することにより、アクションの刈り込みとパラメータの制約に関するローカルな手順の説明をさらに提供します。全体として、KGRL は、トレーニング中のサンプル効率を高めながら、エージェントの探索と展開を実行可能かつ制約を意識した決定に向けて導きます。 KGRL は、サンプル効率とエピソードリターンの両方において、PAMDP の最先端の RL ベースラインを上回ります。

原文 (English)

Knowledge- and Gradient-Guided Reinforcement Learning for Parametrized Action Markov Decision Processes

In this paper, we study Reinforcement Learning in Parametrized Action Markov Decision Processes (PAMDP), where each decision consists of a symbolic action and numerical parameters. In such settings Reinforcement Learning algorithms typically determine parameters with one-shot estimators, which makes their training sample inefficient. Though in most PAMDP environments explicit but incomplete knowledge (e.g., rules, safety constraints, or expert heuristics) is available, it is rarely directly used to increase the sample-efficiency of training Reinforcement Learning agents. We step into this gap and propose our novel Neuro-Symbolic Knowledge- and Gradient-Guided Reinforcement Learning (KGRL) algorithm. KGRL uses domain knowledge in a Datalog knowledge base to derive the set of applicable actions and feasible parameters for a given state. This allows it to prune non-applicable actions from the decision-space and constrain the parameter spaces of the remaining actions. We then use a gradient-based parameter refinement loop to estimate the optimal parameters during training and deployment of the agent. By recording activated rules along the trajectory, KGRL additionally provides local procedural explanations on the pruning of actions and constraining of parameters. Overall, KGRL guides the agent's exploration and deployment toward feasible and constraint-aware decisions, while increasing sample efficiency during training. KGRL outperforms state-of-the-art RL baselines for PAMDPs in both, sample efficiency and episodic return.

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

RCWT: Measuring Task-Budget Displacement from Coordination Content in LLM Calls

Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and…

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Evidence-Grounded Verified Agentic Reasoning: A Path Toward Eliminating LLM Hallucination in Empirical Inference via Tool-Attested Kernel Proofs

Tool access alone does not make LLM empirical reasoning governable: accepted outputs need not descend from attested evidence, and accepted…

2026-07-15 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory…

2026-07-15 06:50 JSTTechCrunch AILLM/生成AIハードウェア/半導体

OpenAI’s new flagship model deletes files on its own, people keep warning

A number of social media posts claim that GPT-5.6 Sol deleted files and data without warning. OpenAI had basically disclosed the problem in…

2026-07-15 06:15 JSTITmedia AI+LLM/生成AIハードウェア/半導体

富士通がNVIDIA「Rubin」対応の国産AIサーバを今秋製造へ ソブリン需要に対応

富士通は、ソブリンAIの需要に応える国産ハイエンドAIサーバやオンプレミス向け生成AI基盤を公開した。国内工場での一貫生産が特徴だ。2026年秋にはNVIDIAの最新GPU「Rubin」対応の新モデルの製造開始であると明かした。

2026-07-14 13:00 JSTarXiv cs.AIハードウェア/半導体

量子化シミュレーションによるライトバックと解釈可能なプログラム実行のためのシンボリック ニューラル CPU

ニューラル ネットワークはアルゴリズムの入出力マッピングを学習できますが、学習した実行プログラムを信頼するには、最終的な正しい答え以上のものが必要です。これは、それを生み出す状態遷移が通常隠蔽されているためです。これらの遷移を可視化するために、トレース監視シンボリック ニューラル CPU、リカレント制御を組み合わせた因数分解学習実行アーキテクチャ、固定微分可能算術論理ユニット バンク上の明示的操作ルーター、宛先マスクされたレジスタ ライトバック、完全な軌跡監視、および一致した固定小数点再生を導入します。このモデルは、選択された操作、ソースおよびデスティネーション レジスタ、レジスタの軌跡、メモリ信号、およびライトバック セマンティクスを各ステップで公開します。主要な 16 幅ベンチマークでは、非量子化エグゼキュータは参照実行を正確に再現しますが、8 ビット量子化でシミュレートされたエグゼキュータは 1,000 命令のプログラムを通じてシンボリック演算パスを保存します。同じ実行が一致する固定小数点再生に対して評価されると、残留数値ドリフトは消失します。これは、それが実行の失敗によるものではなく、連続参照セマンティクスと低精度参照セマンティクスの不一致に起因することを示しています。リカレント、トランスフォーマー、時間畳み込み、時間グラフからインスピレーションを得たコントローラー、状態空間コントローラーを比較し、検査可能な実行パスにはオペレーション ゲートの監視が必要であることをアブレーションで示しています。隠れたオペコードのメモリプレッシャータスクは、遅延状態の使用と一時的なバインディングの残りの制限を明らかにします。また、ValueMemory、ハイブリッド適応リーキー統合発射コントローラー、動作クローン作成とアクタークリティカル強化学習を通じて訓練された候補制約付きシンボリック制御、および RV32I ベース整数セマンティック ブリッジを使用してインターフェイスを拡張します。これらの結果を総合すると、解釈可能で低精度かつ制御可能なニューラル実行のためのトレース検証可能なフレームワークが確立されます。

原文 (English)

A Symbolic Neural CPU for Quantization-Simulated Writeback and Interpretable Program Execution

Neural networks can learn algorithmic input-output mappings, but trusting a learned executor requires more than a correct final answer because the state transitions that produce it are usually hidden. To make those transitions visible, we introduce a trace-supervised symbolic neural CPU, a factorized learned execution architecture that combines recurrent control, an explicit operation router over a fixed differentiable arithmetic-logic unit bank, destination-masked register writeback, complete trajectory supervision and matched fixed-point replay. The model exposes the selected operation, source and destination registers, register trajectory, memory signals and writeback semantics at every step. On the principal 16-wide benchmark, the non-quantized executor reproduces reference execution exactly, while the eight-bit quantization-simulated executor preserves the symbolic operation path through programs of 1,000 instructions. When the same execution is evaluated against a matched fixed-point replay, the residual numerical drift disappears, showing that it comes from a mismatch between continuous and low-precision reference semantics rather than from execution failure. We compare recurrent, Transformer, temporal-convolution, temporal graph-inspired and state-space controllers, and the ablations show that operation-gate supervision is necessary for an inspectable execution path. Hidden-opcode memory-pressure tasks expose the remaining limits in delayed state use and temporal binding. We also extend the interface with ValueMemory, hybrid adaptive leaky integrate-and-fire controllers, candidate-constrained symbolic control trained through behaviour cloning and actor-critic reinforcement learning, and an RV32I base-integer semantic bridge. Together, these results establish a trace-verifiable framework for interpretable, low-precision and controllable neural execution.

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

ルーティング、通信、および推論: 効率的なマルチエージェント推論のためのゲート ルーティングと適応深度

マルチエージェント アンサンブルでは、どのエージェントに相談するか、クエリがエージェントの階層をどの程度深く横断する必要があるか、エージェント間通信がコストに見合うのはいつかという 3 つの基本的な質問に答えることなく、アクティブなパラメータと推論コストが増大します。我々は、4 つの軽量の学習済みゲートが共同してエージェントの選択、階層の深さ、エージェント間通信、および分岐枝刈りを制御する階層型マルチエージェント システムである GRADE (Gated Routing and Adaptive Depth for Efficient Reasoning) を紹介します。トレーニングでは、CoGRPO (Collaborative Group-Relative Policy Optimization) を使用します。これは、GRPO をマルチエージェント階層に適応させ、ロールアウトに参加したすべてのゲートとエージェントに共有アドバンテージシグナルを割り当てる、批判のない新しいレシピです。エージェント モデルは、ホットスワップ可能な Expert Registry から抽出されます。エージェントごとのキャリブレーション マップにより、推論時に再トレーニングすることなく専門家を交代できます。 $\sim$17B の平均アクティブ パラメータでは、GRADE は GSM8K、MMLUPro、GPQA のすべてのベースラインを上回り、アクティブ コンピューティングの半分で MMLUPro で最も強力なベースラインを 4.8 ポイント上回りました。モデルの深さが支配的な AIME-2025 では、GRADE は既存のフレームワークとの競争力を維持します。アブレーションにより、精度に最も大きく寄与する階層とマスクされたクロスアテンションが分離され、安全なホットスワップにはエージェントごとのキャリブレーションが必要であることがわかります。

原文 (English)

Route, Communicate, and Reason: Gated Routing and Adaptive Depth for Efficient Multi-Agent Reasoning

Multi-agent ensembling multiplies active parameters and inference cost without answering three basic questions: which agents to consult, how deeply a query should traverse a hierarchy of agents, and when inter-agent communication is worth its cost. We present GRADE (Gated Routing and Adaptive Depth for Efficient Reasoning), a hierarchical multi-agent system in which four lightweight learned gates jointly govern agent selection, hierarchy depth, inter-agent communication, and branch pruning. Training uses CoGRPO (Collaborative Group-Relative Policy Optimization), a novel critic-free recipe that adapts GRPO to multi-agent hierarchies and assigns a shared advantage signal to every gate and agent that participated in a rollout. Agent models are drawn from a hot-swappable Expert Registry; per-agent calibration maps allow experts to be replaced at inference time without retraining. At $\sim$17B average active parameters, GRADE outperforms all baselines on GSM8K, MMLUPro, and GPQA, surpassing the strongest baseline by 4.8 points on MMLUPro at half the active compute. On AIME-2025, where model depth dominates, GRADE remains competitive to existing frameworks. Ablations isolate the hierarchy and masked cross-attention as the largest contributors to accuracy, and show that per-agent calibration is necessary for safe hot-swapping.

2026-07-14 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

チェッカーから予測者へ: 遅延地上真実の下でのモデル生成の戦略的ルートのコード所有の評価

モデル出力の評価の多くは、評価時にチェックできるコントラクト、または運用ループ内に到着するフィードバックのいずれかに依存します。私たちは、グラウンド トゥルースが遅延、検閲、または非公開であるため、決定論的コードがスコアリング時に正確さをチェックできず、代わりにコード所有の暫定予測を発行する必要があるという補完的な設定を研究します。 RouteCast は、モデル生成の型付き戦略ルートに対してこの体制をインスタンス化します。モデルは候補ルートと構造化された要素を提案します。特定時点の証拠、参照クラス、および決定論的変換により、暫定的な予測ランキングが生成されます。後の結果によって予測が評価されます。 21 件のバイナリ結果ケース (陽性 6 件、陰性 15 件) を対象とした遡及的ベンチャー パイロットでは、パケット全体の RouteCast スコアは予備的な遡及的差別 (AUC 0.756、95% CI [0.471,0.980]) を示しましたが、盲目の LLM 裁判官は AUC 0.678 [0.419, 0.897] に達し、アイデンティティを暴露された LLM 裁判官は AUC に達しました0.761 [0.515,0.944]、認識または結果に関連した漏洩リスクと一致。同じバイナリ サブセットに対する事前登録された分解アブレーションにより、同一の入力を型付きステージング ルートに変換することは、パケット全体のスコア (デルタ AUC = -0.144、95% CI [-0.471,0.176]) および決定論的ヒューリスティック (デルタ AUC = -0.089、95% CI) と区別できないことがわかりました。 [-0.412、0.278])。パイロットは、監査可能な実現可能性の結果を確立し、障害モードを明らかにします。将来のキャリブレーション、因果関係の決定の改善、ルート分解の利点、またはクロスドメインの妥当性を確立するものではありません。

原文 (English)

From Checker to Forecaster: Code-Owned Evaluation of Model-Generated Strategic Routes Under Delayed Ground Truth

Many evaluations of model outputs rely either on contracts checkable at evaluation time or on feedback that arrives within the operating loop. We study the complementary setting in which ground truth is delayed, censored, or private, so deterministic code cannot check correctness at scoring time and must instead issue a code-owned provisional forecast. RouteCast instantiates this regime for model-generated typed strategic routes: models propose candidate routes and structured factors; point-in-time evidence, reference classes, and deterministic transformations produce a provisional forecast-ranking; later outcomes evaluate the forecast. In a retrospective venture pilot on 21 binary-outcome cases (6 positive, 15 negative), the whole-packet RouteCast score showed preliminary retrospective discrimination (AUC 0.756, 95% CI [0.471,0.980]), while a blind LLM judge reached AUC 0.678 [0.419,0.897] and an identity-exposed LLM judge reached AUC 0.761 [0.515,0.944], consistent with recognition- or outcome-related leakage risk. A preregistered decomposition ablation on the same binary subset found that converting the identical inputs into typed staged routes was indistinguishable from the whole-packet score (Delta AUC = -0.144, 95% CI [-0.471,0.176]) and from a deterministic heuristic (Delta AUC = -0.089, 95% CI [-0.412,0.278]). The pilot establishes an auditable feasibility result and exposes failure modes; it does not establish prospective calibration, causal decision improvement, route-decomposition advantage, or cross-domain validity.

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

HCRMap: Pressure-Aware Hot-Expert Residency Mapping for 3.5D MoE Chiplet Inference

Mixture-of-Experts (MoE) large language models (LLM) activate only a small number of experts during inference, but token routing introduces…

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Lesioned Multimodal Language Models Reproduce Aphasic Picture-Naming Patterns

Aphasia following stroke commonly produces systematic naming errors with characteristic profiles, but whether general-purpose language mode…

2026-07-14 13:00 JSTarXiv cs.AIハードウェア/半導体

Model Collapse: On Recursion, Noise, and Uncharted Machine Visions

Since 2023, computer scientists have warned against model collapse -- the contamination of training sets with AI-generated outputs that pro…

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Workload-Driven Optimization for On-Device Real-Time Subtitle Translation

This report studies on-device English-to-Traditional-Chinese subtitle translation for Taiwan under short inputs, short outputs, batch-size-…

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices

Running large language models on consumer devices such as laptops and desktops is challenging because model weights often exceed GPU memory…

2026-07-14 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

Edge Physical AI Deployment of Vision Transformers on Heterogeneous Edge GPU Targeting Autonomous Vehicles

Physical AI systems, such as autonomous vehicles and intelligent machines, require transformer-based perception models that satisfy stringe…

2026-07-14 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Inside the Unfair Judge: A Mechanistic Interpretability Account of LLM-as-Judge Bias

Existing studies of LLM-as-judge scoring bias work predominantly at the input-output level: they perturb inputs, measure score deltas, and…

2026-07-14 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達研究/論文

On the Necessity of Output Distribution Reweighting for Effective Class Unlearning

In this paper, we reveal a significant shortcoming in class unlearning evaluations: overlooking the underlying class geometry can cause inf…

2026-07-13 13:00 JSTarXiv cs.AIハードウェア/半導体

最も重要なパイプ: セマンティック抽象化としての AI システム

AI システムの出力は、それが記述しているように見える事実や世界状態ではなく、むしろ工学的に表現されたものです。私たちは、AI システムを記述し、そのような表現の正確さを検査できるようにするための意味論的なフレームワークを提案します。そのために、受け入れられたドメイン知識によって何が正当化されるのか、参照情報源が何を述べているのか、そしてシステムが現在使用できるものを区別します。これにより、一般的な失敗に正確な定義を与えることができます: 外挿、反駁またはサポートされていない主張、ソースと知識の不一致、古いまたは反駁されたソース、追加された仮説、サポートされていない使用... 私たちのフレームワークが、出力、引用、ツール呼び出し、および世界を変えるアクションが見かけの流暢さではなく信頼できる主張と明示的な権威によって正当化される必要がある AI システムを指定およびチェックするための有用な語彙を提供することを願っています。

原文 (English)

Ceci n'est pas une pipe: AI systems as semantic abstractions

An AI system's output is not the fact or world state it appears to describe, but rather an engineered representation. We propose a semantic framework to describe AI systems, to be able to examine the correctness of such representations. To do so, we distinguish what is justified by accepted domain knowledge, what reference sources say, and what the system can currently use. This allows us to give precise definitions to common failures: extrapolation, refuted or unsupported assertion, sources versus knowledge mismatch, stale or refuted source, added hypotheses, unsupported use... We hope our framework gives a useful vocabulary for specifying and checking AI systems whose outputs, citations, tool calls, and world-changing actions must be justified by reliable claims and explicit authority rather than apparent fluency.

2026-07-13 13:00 JSTarXiv cs.AIハードウェア/半導体

ConceptSMILE: コンセプトベースの説明可能な AI の信頼性を監査する

概念ベースの説明可能な人工知能 (AI) は、モデル推論を人間が理解しやすくすることができますが、概念レベルの出力は自動的に信頼できるわけではありません。概念ベースの説明の信頼性を評価するための、モデルに依存しない摂動ベースの監査フレームワークである ConceptSMILE を紹介します。 ConceptSMILE は、SMILE を置き換えるのではなく、摂動ベースのロジックを機能レベルまたは領域レベルの帰属から人間が理解できる概念説明の監査まで拡張します。このフレームワークは、入力領域を摂動させ、概念と応答のシフトを測定し、局所性の重み付けを適用し、XGBoost サロゲートを適合させて局所的な概念の動作を近似します。信頼性は、帰属の正確さ、代理の忠実性、忠実性、安定性、一貫性によって評価されます。 MedSAM 由来の視覚概念と VLM ベースの意味概念を比較することにより、網膜眼底画像上の ConceptSMILE を評価します。結果は、信頼性が概念や経路によって異なることを示しています。MedSAM は、より強力な空間帰属と最高のサロゲート忠実度 ($R^2 = 0.8503$、$R_w^2 = 0.8465$) を達成しますが、VLM 経路は、選択されたアーティファクト条件下でより強い血管忠実性とより強い安定性を示します。 ConceptSMILE は、コンセプトベースの XAI の信頼性を評価するための独立した監査レイヤーを提供します。

原文 (English)

ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

Concept-based explainable artificial intelligence (AI) can make model reasoning more human-understandable, but concept-level outputs are not automatically trustworthy. We introduce ConceptSMILE, a model-agnostic perturbation-based auditing framework for evaluating the reliability of concept-based explanations. Rather than replacing SMILE, ConceptSMILE extends its perturbation-based logic from feature- or region-level attribution to the auditing of human-understandable concept explanations. The framework perturbs input regions, measures concept-response shifts, applies locality weighting, and fits an XGBoost surrogate to approximate local concept behaviour. Reliability is assessed through attribution accuracy, surrogate fidelity, faithfulness, stability, and consistency. We evaluate ConceptSMILE on retinal fundus images by comparing MedSAM-derived visual concepts with VLM-based semantic concepts. Results show that reliability varies across concepts and pathways: MedSAM achieves stronger spatial attribution and the highest surrogate fidelity ($R^2 = 0.8503$, $R_w^2 = 0.8465$), while the VLM pathway shows stronger vessel faithfulness and stronger stability under selected artefact conditions. ConceptSMILE provides an independent audit layer for evaluating the trustworthiness of concept-based XAI.

2026-07-13 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

適度に非構造化されたスパース重み行列を使用した大規模言語モデルの GPU 推論の高速化

大規模言語モデル (LLM) の導入が進むにつれて、LLM 推論コストが重要な課題になっています。重み行列にスパース性を導入する枝刈り手法により、推論を高速化できます。ただし、モデルの品質を維持するには、通常、プルーニングを中程度の非構造化スパース度 (約 50%) に制限します。これらのスパース レベルでは、スパース行列乗算 (SpMM) 用の既存の GPU カーネルはいずれも、高密度の対応物を上回るパフォーマンスを発揮できません。この論文では、中程度のスパース性を持つ LLM に対する効率的な GPU 推論方法を提案します。我々は、(i) スパース テンソル コアが SpMM を高速化できるようにする Sparse-TC レイヤー、 (ii) 低コストのオンチップ復号化をサポートしながら、行列圧縮に並列差分距離を使用するスロット充填層。 (iii) 正しい SpMM 計算を保証する軽量の Residual Layer。この形式に基づいて、スパース テンソル コアと CUDA コアを共同利用する SpMM カーネルを設計します。この設計により、効率的な実行パイプラインが可能になり、オンチップ計算とメモリ アクセスが重複します。評価の結果、私たちの研究は、高帯域幅メモリ (HBM) を備えた最新の GPU での密行列乗算を初めて上回るパフォーマンスを示したことが示されています。 SpInfer (EuroSys'25、最優秀論文) と比較して最大 1.64 倍のカーネル レベルの高速化、および FlashLLM (VLDB'24) と比較して最大 1.41 倍のエンドツーエンドの高速化を実現します。ソースコード: https://github.com/moui0/cudac。

原文 (English)

Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices

With the growing deployment of large language models (LLMs), LLM inference cost has become a key challenge. Pruning techniques that introduce sparsity into weight matrices can accelerate inference. However, maintaining model quality typically limits pruning to moderate unstructured sparsity (around 50\%). At these sparsity levels, none of the existing GPU kernels for sparse matrix multiplication (SpMM) can outperform their dense counterparts. This paper proposes an efficient GPU inference method for LLMs with moderate sparsity. We propose a three-layer matrix storage format comprising: (i) a Sparse-TC layer enabling sparse tensor cores to accelerate SpMM; (ii) a Slot-Filling layer using parallel differential distance for matrix compression while supporting low-cost on-chip decoding; (iii) a lightweight Residual Layer ensuring correct SpMM computation. Building on this format, we design a SpMM kernel that jointly utilizes sparse tensor cores and CUDA cores. This design enables an efficient execution pipeline and overlaps on-chip computation with memory access. Evaluations show that our work is the first to outperform dense matrix multiplication on modern GPUs equipped with high-bandwidth memory (HBM). It achieves up to 1.64x kernel-level speedup over SpInfer (EuroSys'25, Best paper) and up to 1.41x end-to-end speedups over FlashLLM (VLDB'24). Our source code: https://github.com/moui0/cudac.

2026-07-13 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

STEEL: AMD の XDNA NPU でのエネルギー効率の高い長期シーケンス推論のためのスパース性を意識した融合された注意

オペレーティング システムのワークフロー内で大規模な言語モデル ベースのエージェントの採用が増えているため、ラップトップ クラスのシステム オン チップ (SoC) でのエネルギー効率の高い推論の重要性が高まっています。クラウド オフロードは依然として一般的ですが、信頼性とプライバシーの問題が生じ、特にエージェント ワークロードにとって問題となります。したがって、最近のラップトップ SoC には、エネルギー効率を最適化するニューラル処理エンジン (NPU) が組み込まれています。ただし、アーキテクチャの多様性と明示的なデータ移動プログラミング モデルにより、アテンション メカニズムを NPU に効果的にマッピングすることは依然として困難です。この研究では、XDNA のような NPU をターゲットとする FlashAttend の最初のオープンソース実装である STEEL を紹介します。 STEEL は、プレフィル アテンションのデータフロー定式化を導入し、空間並列処理とオンチップ メモリの効率的な活用を可能にします。さらに、STEEL は、NPU アレイ上でスパース性を意識したパイプライン配置を活用することで、因果マスクによって引き起こされる負荷の不均衡に対処し、同期オーバーヘッドを削減し、使用率を向上させます。 AMD Ryzen AI 9 HX 370 SoC 上の STEEL を評価し、そのパフォーマンスを最適化された CPU および GPU 実装と比較します。実験結果によると、STEEL は CPU と GPU のベースラインと比較して、エネルギー消費をそれぞれ平均 9.17 倍と 1.75 倍削減します。 XDNA 1 では、STEEL は従来の最新技術と比較して平均 9.6 倍の遅延短縮を達成し、XDNA 2 でのレイヤーごとのアテンション実装と比較して平均 22.8 倍の高速化を実現します。

原文 (English)

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

The growing adoption of large language model-based agents within operating system workflows has increased the importance of energy-efficient inference on laptop-class systems-on-chip (SoCs). While cloud offloading remains common, it introduces reliability and privacy concerns that are particularly problematic for agentic workloads. Recent laptop SoCs, therefore, incorporate neural processing engines (NPUs) optimized for energy efficiency; however, effectively mapping attention mechanisms onto NPUs remains challenging due to architectural diversity and explicit data-movement programming models. In this work, we present STEEL, the first open-source implementation of FlashAttention targeting XDNA-like NPUs. STEEL introduces a dataflow formulation of prefill attention, enabling efficient exploitation of spatial parallelism and on-chip memory. Furthermore, STEEL addresses the load imbalance induced by the causal mask by leveraging a sparsity-aware pipeline placement onto the NPU array, reducing synchronization overhead and improving utilization. We evaluate STEEL on the AMD Ryzen AI 9 HX 370 SoC and compare its performance against optimized CPU and GPU implementations. Experimental results show that STEEL reduces energy consumption by an average of 9.17x and 1.75x relative to CPU and GPU baselines, respectively. On XDNA 1, STEEL achieves an average 9.6x latency reduction over the prior state of the art, and delivers a 22.8x speedup on average compared to a layer-by-layer attention implementation on XDNA 2.

2026-07-13 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language

LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate…

2026-07-11 02:17 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

SK Hynix raises $26.5B in the biggest foreign IPO in US history, is urged to build new US fabs

The AI chip boom just produced its biggest Wall Street moment yet. Now SK Hynix and Samsung are being asked to build U.S. factories.

2026-07-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

FedOPAL: 分析ビジュアル プロンプト チューニングによるワンショットのフェデレーテッド ラーニング

エッジ インテリジェンスにおける基本モデルの広範な展開に伴い、通信帯域幅がフェデレーテッド ラーニングのスケーラビリティを制限する中心的なボトルネックになっています。ワンショットのフェデレーテッド ラーニングは通信ラウンドを最小限に抑えることでこの問題を軽減しますが、既存の反復的な微調整や知識の蒸留方法では、サーバー側の高い計算コストやハイパーパラメータの感度などの課題に依然として直面しています。分析的フェデレーテッド ラーニングは、最小二乗閉形式解を使用して効率的な勾配のない集計を実現しますが、非独立で同一に分散されたデータを含む環境では、その静的特徴の仮定が失敗し、特徴多様体の位置ずれが発生し、モデルのパフォーマンスが大幅に低下します。この矛盾に対処するために、この文書では FedOPAL フレームワークを提案します。このフレームワークは、視覚的プロンプトを特徴修正手段として適応させ、局所近位制約を適用することで異種データの特徴分布を線形分離可能な空間に積極的に修正し、それによって分析的連合学習の理論的前提を満たします。実験結果は、FedOPAL がいくつかのベンチマークで元の分析手法を大幅に上回るだけでなく、サーバー側のトレーニング コストをゼロに維持しながら最先端の反復手法に匹敵する精度を達成し、エッジで大規模なモデルを効率的に連携させるための新しいエンジニアリング パラダイムを提供することを示しています。

原文 (English)

FedOPAL: One-Shot Federated Learning via Analytic Visual Prompt Tuning

With the widespread deployment of basic models in edge intelligence, communication bandwidth has become a core bottleneck restricting the scalability of federated learning. Although one-shot federated learning alleviates this problem by minimizing communication rounds, existing iterative fine-tuning or knowledge distillation methods still face challenges such as high server-side computational costs and hyperparameter sensitivity. Analytical federated learning achieves efficient gradientfree aggregation using least-squares closed-form solutions, but in environments with non-independent and identically distributed data, its static feature assumptions fail, leading to feature manifold misalignment and severely impairing model performance. To address this contradiction, this paper proposes the FedOPAL framework. This framework adapts the visual prompts as feature rectifiers, actively correcting the feature distribution of heterogeneous data to a linearly separable space by applying local proximal constraints, thereby satisfying the theoretical assumptions of analytical federated learning. Experimental results show that FedOPAL not only significantly outperforms the original analytical methods on several benchmarks, but also achieves accuracy comparable to state-of-the-art iterative methods while maintaining zero server-side training costs, providing a new engineering paradigm for efficient collaboration of large models on the edge.

2026-07-10 13:00 JSTarXiv cs.AIハードウェア/半導体

JEPA スタイルの予測学習を JA4 由来のネットワーク フィンガープリントに適用する

I-JEPA と V-JEPA は、元の入力を再生成するのではなく、潜在的な予測をターゲットのエンコーダー出力に照合することで学習します。これは画像やビデオではうまく機能します。同じ目的がコンパクトなネットワーク フィンガープリントでも機能するかどうかを調査します。 JA4DB および CIC-IDS-2017 から抽出された JA4、JA4H、JA4S、および JA4X サブフィールドでトレーニングされた Transformer ベースのモデルである JA4-JEPA を構築しました。トレーニング データは両方のソースからの約 397,000 のサンプルを組み合わせていますが、4 つのビュー ファミリすべてを含む単一のサンプルはありません。私たちは、TLS、DNS、SSH にわたるプロトコル ファミリ分類について、凍結された kNN プローブを使用して学習された表現を評価しました。 39,416 個のホールドアウト サンプルで、モデルはコサイン類似度 0.9899 と kNN 精度 0.9220 を達成しました。これらの結果は、ソース間でビューが不完全に重複している場合でも、JEPA スタイルの予測学習が JA4 由来のフィンガープリントから有用な埋め込みを生成できることを示しています。キーワード: JA4、ネットワークフィンガープリンティング、JEPA、予測表現学習、自己教師あり学習

原文 (English)

Applying JEPA-Style Predictive Learning to JA4-Derived Network Fingerprints

I-JEPA and V-JEPA learn by matching latent predictions to target encoder outputs rather than regenerating the original input, and this has worked well for images and video. We explore whether the same objective works for compact network fingerprints. We built JA4-JEPA, a Transformer-based model trained on JA4, JA4H, JA4S, and JA4X subfields drawn from JA4DB and CIC-IDS- 2017. The training data combines roughly 397K samples from both sources, though no single sample contains all four view families. We evaluated the learned representations with a frozen kNN probe on protocol-family classification across TLS, DNS, and SSH. On 39,416 heldout samples the model achieved a cosine similarity of 0.9899 and a kNN accuracy of 0.9220. These results indicate that JEPA-style predictive learning can produce useful embeddings from JA4-derived fingerprints, even with incomplete view overlap across sources. Keywords: JA4, network fingerprinting, JEPA, predictive representation learning, self-supervised learning

2026-07-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

ParamMute: Suppressing Knowledge-Critical FFNs for Faithful Retrieval-Augmented Generation

Large language models (LLMs) integrated with retrieval-augmented generation (RAG) have improved factuality by grounding outputs in external…

2026-07-10 13:00 JSTarXiv cs.AIハードウェア/半導体

LoKA: Low-precision Kernel Applications for Recommendation Models At Scale

Recent GPU generations deliver significantly higher FLOPs using lower-precision arithmetic, such as FP8. While successfully applied to larg…

2026-07-10 04:05 JSTTechCrunch AILLM/生成AIハードウェア/半導体規制/政策

New York Times says OpenAI hid evidence in ChatGPT copyright trial

News publishers say OpenAI hid tools and datasets that could identify copyrighted journalism in ChatGPT outputs, escalating their lawsuit w…

2026-07-10 03:34 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

Paris-based AI voice startup Gradium raises $100M seed, backed by Nvidia

The company is using the cash to open an office in the Bay Area and compete for talent there, "strengthening its position at the heart of t…

2026-07-10 02:17 JSTTechCrunch AIハードウェア/半導体

Meta’s new AI chips will begin production in September

The company is taking a modular approach to designing these chips, anticipating that their needs will change as AI evolves rapidly by the t…

2026-07-10 02:06 JSTTechCrunch AIハードウェア/半導体

Nvidia is a victim of the compute marketplace it created

Having proven how valuable compute can be, the company finds itself at the center of a market everyone wants to be in — while simpler techn…

2026-07-09 13:00 JSTITmedia AI+エージェントハードウェア/半導体

社内のWindows環境で「数百のAIエージェント」を隔離実行 NVIDIAとMicrosoftが共同開発したデスクサイドマシンの全容

NVIDIAは、Windows環境でAIエージェントを開発・実行するデスクサイドAIスーパーコンピュータ「NVIDIA DGX Station for Windows」を発表した。NVIDIA GB300 Grace Blackwell Ultra Desktop Superc…

2026-07-09 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体研究/論文

AI における再帰的自己改善: 制限された自己改善から自律的な研究ループへ

AI システムは、出力の修正、展開中の独自のハーネスの調整、生成されたデータのトレーニング、そして AI 研究自体の実施など、AI システム自体の改善にますます参加しています。この文献は、根本的に異なる野心を混同した語彙 (「自己洗練」、「自己報酬」、「自己遊び」、「自己進化」) の下で説明されています。私たちは 1,250 件の arXiv 論文 (2024 年から 2026 年) を 2 つの軸に沿って調査しました。システムの改善内容 (導入時の動作、トレーニングを通じたポリシー、評価者、または研究プロセス自体)、およびループの閉鎖度 (人間参加型から完全な閉鎖まで) です。この分類法は、制限された自己改善 (収束的で評価可能ですでに産業的な実践である) を、グラウンディング要件、崩壊ダイナミクス、および測定されたすべての軸の計算制約によって制限されたままである、制限のない再帰的自己改善 (RSI) から分離します。その際立った特徴は、自己評価専用のカテゴリです。すべての改善ループは、何らかの信号が人間の判断に代わることができるという主張です。私たちは、評価者の設計空間 (審査員、プロセス報酬モデル、検証者、ルーブリック、メタ評価) を調査し、形式的な検証者 (最も強い) から本質的な自己評価 (最も弱い) までの検証階層にシグナルを順序付けします。そして、実証された自己改善の強さがこの階層を追跡し、その違反からその失敗モード (自己確認ループ、モデルの崩壊、多様性の崩壊) が発生し、「研究の方向性設定」のボトルネックが人間の行動を妨げていることを観察します。ループ内はその階層の最上位に位置します。我々は、技術文献をRSI限界の理論と、ループを閉じることに関するフロンティアラボの説明によって提起された安全性とガバナンスの問題に結び付け、自己改善のガバナンスグレードの測定がこの分野で最も人口の少ないニッチであることを特定します。

原文 (English)

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops

AI systems increasingly participate in their own improvement: revising their outputs, adapting their own harnesses during deployment, training on data they generate, and, increasingly, conducting AI research itself. This literature is described under a vocabulary ("self-refine," "self-reward," "self-play," "self-evolve") that conflates fundamentally different ambitions. We survey 1,250 arXiv papers (2024-2026) along two axes: what the system improves -- its behavior in deployment, its policy through training, its evaluator, or the research process itself -- and the degree of loop closure (human-in-the-loop to fully closed). The taxonomy separates bounded self-refinement -- convergent, evaluable, and already industrial practice -- from open-ended recursive self-improvement (RSI), which remains bounded by grounding requirements, collapse dynamics, and compute constraints on every measured axis. Its distinctive feature is a dedicated category for self-evaluation: every improvement loop is a claim that some signal can substitute for human judgment. We survey the evaluator design space -- judges, process reward models, verifiers, rubrics, meta-evaluation -- order the signals into a verification hierarchy from formal verifiers (strongest) to intrinsic self-assessment (weakest), and observe that demonstrated self-improvement strength tracks this hierarchy, that its failure modes (self-confirming loops, model collapse, diversity collapse) follow from its violations, and that the "research direction-setting" bottleneck keeping humans in the loop sits at the top of that hierarchy. We connect the technical literature to the theory of RSI limits and to the safety and governance questions raised by frontier-lab accounts of closing the loop, and identify governance-grade measurement of self-improvement as the field's most underpopulated niche.

2026-07-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

NonTextual Target Attack

Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output w…

2026-07-09 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

T2T-VICL: Cross-Task Visual In-Context Learning via Implicit Text-Driven VLMs

Visual in-context learning (VICL) solves visual tasks by conditioning on a few input-output demonstrations without any model training. Rece…

2026-07-09 13:00 JSTarXiv cs.AIハードウェア/半導体

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods

The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatr…

2026-07-08 17:00 JSTTechCrunch AILLM/生成AIハードウェア/半導体

Hot French startup ZML releases free product to speed inference across lots of AI chips

ZML, a hot French AI startup endorsed by Turing Award winner Yann LeCun, has now released ZML/LLMD, software that could make running AI les…

2026-07-08 16:16 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

AI chip maker SambaNova raises $1B at $11B valuation, 5 months after last mega round

AI chip maker SambaNova has raised at an $11 billion valuation months after Intel was rumored to be trying to buy it for about $1.6 billion.

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

AgoraSim: ハイブリッド エージェント ベースのモデリング フレームワーク

LLM エージェント シミュレーションを使用すると、自然言語の社会シナリオを簡単にインスタンス化できますが、その出力は予測として読み取られる可能性があり、明示的な社会動態と比較するのが困難なことがよくあります。シナリオ指向の社会反応分析のためのハイブリッド エージェント ベースのモデリング フレームワークである AgoraSim を紹介します。 AgoraSim は、テキストまたはマルチモーダルのアーティファクトを編集可能な ABM 構成に解決し、LLM、ビジョン言語、カスタム エンドポイント、ランダム、および古典的なエージェントを混合する比率制御された母集団を実行し、同じシナリオを一致する古典的な参照ダイナミクスと比較します。すべてのエージェントは共有の構造化意思決定オブジェクトを発行し、共通のアクション スペース、対話プロトコル、メトリクス、監査レコードを有効にします。ローカル UI、Python SDK/CLI、および REST API を通じて公開される AgoraSim は、ユーザーがシナリオの軌跡を検査し、モデリングの前提を比較し、経験的検証が必要なケースを特定するのに役立ちます。

原文 (English)

AgoraSim: A Hybrid Agent-Based Modeling Framework

LLM-agent simulations make natural-language social scenarios easy to instantiate, but their outputs can be overread as predictions and are often difficult to compare with explicit social dynamics. We present AgoraSim, a hybrid agent-based modeling framework for scenario-oriented social reaction analysis. AgoraSim resolves textual or multimodal artifacts into editable ABM configurations, runs ratio-controlled populations that mix LLM, vision-language, custom-endpoint, random, and classical agents, and compares the same scenario against matched classical reference dynamics. All agents emit a shared structured decision object, enabling common action spaces, interaction protocols, metrics, and audit records. Exposed through a local UI, Python SDK/CLI, and REST API, AgoraSim helps users inspect scenario trajectories, compare modeling assumptions, and identify cases that warrant empirical validation.

2026-07-08 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

RPAM: A Principled Metric for Evaluating Associations in Language Models with High Predictive Validity in Downstream Outputs

Language models (LMs) exhibit problematic biases, such as stereotypes. Effectively analyzing and mitigating such biases requires accurate a…

2026-07-08 13:00 JSTarXiv cs.AIハードウェア/半導体

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods

The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatr…

2026-07-08 07:30 JSTITmedia AI+ハードウェア/半導体

NTTドコモビジネスがIOWN活用の分散GPU環境を提供、25GBを2秒で転送

NTTドコモビジネスは、次世代ネットワーク「IOWN APN」を活用し、全国8拠点に分散したGPUを統合利用できる実証環境の提供を開始した。電力などの制限を解消し、オンデマンドなリソース確保やデータ主権に対応した分散AI基盤の実用性を検証できる。

2026-07-08 01:27 JSTTechCrunch AILLM/生成AIハードウェア/半導体

Claude Cowork expands to mobile and web

With this update, users can start a task from their desk, get status updates on their phone, and pick up the finished output later — even i…

2026-07-07 13:00 JSTarXiv cs.AIハードウェア/半導体

議論を避ける方法: 二重効率の対話型証明によるスケーラブルな AI の安全性

AI モデルが強力な機能を開発し続けるにつれて、その出力が私たちの意図と一致していることを検証できることが重要になります。最近の研究では、ディベートによる検証に焦点を当てています。これは、2 つの競合する強力な証明者または AI モデルが相互に議論して、弱い検証者または人間に主張の正しさを納得させる対話型証明のモデルです。ただし、議論では 2 つの AI モデルが同等の能力を持ち、そのうちの 1 つが真実であると仮定されていますが、これは現実的ではない可能性があります。この研究では、\emph{議論を避ける方法}を示します。つまり、AI の安全性のための \emph{単一証明者} の対話型証明の研究を開始します。単一証明者の対話型証明における以前の結果は、AI の安全性設定にすぐには引き継がれません。たとえば、計算が人間の判断や Web などの外部データベースなどのオラクルにアクセスする場合、それらは機能しません。我々は、(1) オラクルのクエリに対する答えのごく一部が間違っていても出力が変わらないという意味で計算が堅牢である、または (2) オラクルが低次の多項式であるという設定における、オラクル支援計算 (相対化証明とも呼ばれる) のための二重効率の単一証明者の対話型証明と引数を提示します。これらの結果は、構造化された、またはノイズ耐性のあるオラクルアクセスの下では、議論がなくても対話型検証が可能であることを示唆しています。

原文 (English)

How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs

As AI models continue to develop powerful capabilities, it becomes critical that we are able to verify that their output is aligned with our intentions. A recent line of work focuses on verification via debate, a model of interactive proofs where two competing powerful provers, or AI models, debate each other to convince a weak verifier, or a human, of the correctness of their claim. However, debate assumes that the two AI models possess equal abilities and that one of them is truthful, which may not be realistic. In this work, we show \emph{how to avoid debate}: we initiate the study of \emph{single-prover} interactive proofs for AI safety. Prior results in single-prover interactive proofs do not immediately carry over to the AI safety setting: for example, they do not work when the computation has access to an oracle, such as to human judgment or an external database such as the web. We present doubly-efficient single-prover interactive proofs and arguments for oracle-aided computations (also known as relativizing proofs), in the settings where (1) the computation is robust, in the sense that the output does not change if at most a small fraction of the answers to oracle queries are incorrect, or (2) the oracle is a low-degree polynomial. These results suggest that interactive verification is possible even without debate, under structured or noise-tolerant oracle access.

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

決定論的グラウンドトゥルースを使用した長形式生成における LLM 不確実性の評価

LLM が生成する出力はますます長くなっているため、効果的な不確実性推定では、応答全体を破棄するのではなく、きめの細かいレベルでエラーを特定する必要があります。このような方法は存在しますが、任意の解像度 (世代全体に対するトークン) での不確実性を評価することは困難であり、ラベルの不完全性の影響を非常に受けやすいため、ゼロノイズ ベンチマークが不可欠です。しかし、長い形式の生成ベンチマークは、決定的なグラウンド トゥルースではなく、誤ったラベルに依存する傾向があります。単一の決定論的な長いテキストのグラウンド トゥルースを使用して、手続き的に生成された 6 つのタスクのベンチマークである単一回答アトミック ロングフォーム ターゲット (SALT) を導入します。これにより、外部の判断者なしで、正確性、校正、およびランキングのユニットレベルの評価が可能になります。 SALT を搭載した 50 以上の LLM の分析により、重要な洞察が明らかになります。どの信頼関数が各不確実性の側面を支配しているかを特定し、より粗いラインレベル単位でより明確な分離性が現れた場合でも、信頼度ランキングが原子分解能で大きく崩れることを示します。 SALT はさらに、生成全体を通じて制御されたアトムレベルの介入を可能にし、将来のエラーの 2 つの分離可能な要因を明らかにします。それは、グローバル コンテキストの正確性によって支配される破損したプレフィックスからの伝播と、応答コンテキストの長さの増加による限定的な劣化です。最後に、思考連鎖のプロンプトまたはトレーニングを通じて内面化された推論によって、信頼度ランキングが低下する一方で精度が向上するというトレードオフが導入されることを示します。これらの発見は、信頼性の高いエラーの特定と軽減を必要とするリスククリティカルなアプリケーションに直接影響します。

原文 (English)

Evaluating LLM Uncertainty in Long-Form Generation Using Deterministic Ground Truth

As LLMs generate increasingly long outputs, effective uncertainty estimation must identify errors at fine-grained levels rather than discard entire responses. While such methods exist, evaluating uncertainty at any resolution (token to an entire generation) is challenging and highly sensitive to label imperfections, making zero-noise benchmarks essential; yet, long-form generation benchmarks tend to rely on fallible labels rather than deterministic ground truth. We introduce Single-answer Atomic Long-form Target (SALT), a benchmark of six procedurally generated tasks with single deterministic long textual ground truths, enabling unit-level evaluation of correctness, calibration, and ranking without external judges. Equipped with SALT, our analysis of 50+ LLMs reveals key insights: We identify which confidence functions dominate each uncertainty aspect and show that confidence ranking largely breaks at atomic resolution, even when clearer separability emerges at coarser line-level units. SALT further enables controlled atom-level interventions throughout generation, revealing two separable drivers of future errors: propagation from corrupted prefixes, dominated by global context correctness, and bounded degradation from increasing answer-context length. Finally, we demonstrate that reasoning, via Chain-of-Thought prompting or internalized through training, introduces a trade-off, improving accuracy while degrading confidence ranking. These findings directly impact risk-critical applications requiring reliable error identification and mitigation.

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

大規模言語モデルにおける認識エントロピーを除去するためのローリング係数のヘヴィサイド連続性

大規模言語モデル (LLM) は、間違っている可能性がある流暢な出力を生成します。誤った情報を提供するときに手がかりを示すことが多い人間とは異なり、LLM は検出が困難なエラーを生成します。これは、自己回帰デコードには状態が進行する前に中間推論を検証するメカニズムがないためです。ヘビサイド ゲートによって制御される述語ゲート状態遷移として推論を再定式化する検証優先実行フレームワークであるヘビサイド ローリング係数 (HCRC) を紹介します。 HCRC は、モデルの信頼性とパラレル ワーカー アーキテクチャからの独立した検証信号を組み合わせて、事前定義された正確性述語が満たされた場合にのみ実行を進めることができます。これにより、無効な中間状態の伝播が防止され、基礎となるモデルを変更することなく認識論的エントロピーが削減されます。私たちは、4 つのプロバイダーからの 13 人の提案者にわたって、ソフトウェア エンジニアリングと推論タスクに関して HCRC を評価します。有能なプロポーザーでは、ゲートは遅延競合性を維持しながら誤完了率 (FCR) を 4 ~ 7% から 0% に削減し、設定によってはアンラップ モデルよりも高速になります。弱いプロポーザーでは、ダウンストリームの状態を破壊するのではなく、誤った完了を正当な停止に変換します。ベンチマークを超えて、HCRC はエージェント コーディング環境の実稼働コントロール プレーンとして数か月間運用され、ファイルの変更の承認、検証主導の進捗レポート、メモリ圧縮を行ってきました。これらの結果は、検証主導型 LLM 実行の一般的なフレームワークとして HCRC を確立し、信頼性の高い推論がモデル スケールだけではなく原則に基づいた実行制御を通じて達成できることを示しています。

原文 (English)

Heaviside Continuity of Rolling Coefficients for Eliminating Epistemic Entropy in Large Language Models

Large language models (LLMs) generate fluent outputs that can be wrong. Unlike humans, who often exhibit cues when providing false information, LLMs produce errors that are difficult to detect because autoregressive decoding provides no mechanism for verifying intermediate reasoning before state progression. We introduce Heaviside Continuity of Rolling Coefficients (HCRC), a verification-first execution framework that reformulates inference as predicate-gated state transitions governed by a Heaviside Gate. HCRC combines model confidence with independent verification signals from a parallel worker architecture, allowing execution to advance only when predefined correctness predicates are satisfied. This prevents invalid intermediate states from propagating, reducing epistemic entropy without modifying the underlying model. We evaluate HCRC on software-engineering and reasoning tasks across thirteen proposers from four providers. On capable proposers, the gate reduces the false-completion rate (FCR) from 4--7% to 0% while remaining latency-competitive and, in some settings, faster than the unwrapped model. On weaker proposers, it converts false completions into honest halts instead of corrupting downstream state. Beyond benchmarking, HCRC has operated for months as the production control plane of an agentic coding environment, authorizing file mutations, verification-driven progress reporting, and memory compaction. These results establish HCRC as a general framework for verification-driven LLM execution, showing that reliable reasoning can be achieved through principled execution control rather than model scale alone.

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

OpenClaw の GLM-5 サービス提供パラメータ調整: ロングコンテキスト エージェント ワークロード向けの単一デプロイメント MaaS 推論の最適化

OpenClaw リクエストは、システム プロンプト、会話履歴、コンテキスト ウィンドウにフィードバックされるツール出力など、ツールによって拡張された長いプレフィックスによって占められます。このワークロードでは、リクエストごとに約 28,000 ~ 30,000 の入力トークンと 500 の出力トークンがあり、サービスの品質は、ショート プロンプトのスループットだけではなく、スループット、TTFT、およびテール レイテンシによって決まります。このレポートでは、MaaS マルチモデル推論最適化アーキテクチャ内での GLM-5 サービング パラメーターの調整について調査します。スコープは推論最適化レイヤーの単一ノード最適化ブロックで、チャンク プリフィル、テンソル並列処理 (TP)、パイプライン並列処理 (PP)、およびリクエストの同時実行性が 1 つの GLM-5 サービング デプロイメント用に調整されます。このレポートでは、「単一ノードの最適化」はアーキテクチャ ブロックを指しますが、実験は 2 ノード、16 GPU クラスターで実行されます。テストされたスペース内での最適な構成は、chunked-prefill-size=3072、tp=4、pp-size=4、および max-running-requests=24 です。保守的な 2048/4/4/16 ベースラインと比較すると、リクエスト スループットが 0.43 から 0.48 req/s に、総トークン スループットが 9029.64 から 9993.23 tok/s に増加し、平均 TTFT が 8.98 秒から 6.69 秒に、レイテンシ P90 が 40.23 秒から 32.64 秒に減少します。同じハードウェア フットプリントの下では、これは推定でリクエストあたりのサービス コストが 10.4% 低くなり、トークンあたりのコストが 9.6% 低いことに相当します。結果は、最適値はワークロード固有であることを示しています。チャンク サイズが大きくなり、キューが深くなっても、パフォーマンスは単調に向上するわけではありません。したがって、デフォルトの OpenClaw 導入プロファイルとして 3072 / tp4 / pp4 / max24 をお勧めします。

原文 (English)

GLM-5 Serving Parameter Tuning for OpenClaw: Single-Deployment MaaS Inference Optimization for Long-Context Agent Workloads

OpenClaw requests are dominated by long, tool-augmented prefixes, including system prompts, conversation history, and tool outputs fed back into the context window. For this workload, with about 28k-30k input tokens and 500 output tokens per request, serving quality is governed by throughput, TTFT, and tail latency rather than short-prompt throughput alone. This report studies GLM-5 serving-parameter tuning within a MaaS multi-model inference optimization architecture. The scope is the Single-Node Optimization block of the inference-optimization layer, where chunked prefill, tensor parallelism (TP), pipeline parallelism (PP), and request concurrency are tuned for one GLM-5 serving deployment; in this report, "Single-Node Optimization" denotes the architecture block, while experiments run on a two-node, sixteen-GPU cluster. Within the tested space, the best configuration is chunked-prefill-size=3072, tp=4, pp-size=4, and max-running-requests=24. Compared with the conservative 2048/4/4/16 baseline, it increases request throughput from 0.43 to 0.48 req/s and total token throughput from 9029.64 to 9993.23 tok/s, while reducing average TTFT from 8.98 to 6.69 s and latency P90 from 40.23 to 32.64 s. Under the same hardware footprint, this corresponds to an estimated 10.4% lower serving cost per request and 9.6% lower cost per token. The results show that the optimum is workload-specific: larger chunk sizes and deeper queueing do not monotonically improve performance. We therefore recommend 3072 / tp4 / pp4 / max24 as the default OpenClaw deployment profile.

2026-07-07 13:00 JSTarXiv cs.AIハードウェア/半導体

Fine-Grained Computation Offload for Off-the-Shelf Servers in Tens of Lines

Hardware accelerators now sit on the critical path of online serving. GPUs, FPGAs, and increasingly remote services such as hardware securi…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Back to Basics: Improving Molecular Understanding in LLMs via SMILES-Graph Translation

Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet thes…

2026-07-07 13:00 JSTarXiv cs.AIハードウェア/半導体

Self-Specializing Vision-Language Transmon Chip Calibration in a Physics-Grounded Environment

Calibrating a superconducting transmon chip is a sequential decision problem under noise, drift, and a finite budget: an expert must choose…

2026-07-07 13:00 JSTarXiv cs.AIハードウェア/半導体

An AI-Assisted Solution to the Signed BAR Conjecture: Uniqueness in the Harrison--Reiman Class and a Completely-$\mathcal{S}$ Class Obstruction

For a multidimensional reflected diffusion, determining whether the associated basic adjoint relationship (BAR) uniquely characterizes the…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language

LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Saving GPU Hours in LLM Inference System Development and Online Workloads with Simulation and DBMS-Inspired Cache Replacement Policies

LLMs are increasingly used world-wide from daily tasks to agentic systems and data analytics, requiring significant GPU resources. While LL…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Evolutionary Guided Decoding: Iterative Value Refinement for LLMs

While guided decoding, especially value-guided methods, has emerged as a cost-effective alternative for controlling language model outputs…

2026-07-07 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

VISOR: Visual Input-based Steering for Output Redirection in Vision-Language Models

Vision Language Models (VLMs) are increasingly being used in a broad range of applications, bringing their security and behavioral control…

2026-07-07 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

Score-Regularized Joint Sampling with Importance Weights for Flow Matching

Flow matching models effectively represent complex distributions, yet estimating expectations of functions of their outputs remains challen…

2026-07-07 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

A Random Matrix Theory Perspective on the Consistency of Diffusion Models

Diffusion models trained on different, non-overlapping subsets of a dataset often produce strikingly similar outputs when given the same no…

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distri…

2026-07-07 13:00 JSTarXiv cs.AIハードウェア/半導体

必要なのは FP8 だけです (パート 1): HPC の聖杯としてのハードウェア FP64 の誤りを暴く

従来の HPC の定説では、ネイティブ ハードウェア FP64 シリコンは科学技術コンピューティングの還元不可能な基盤、つまり倍精度シミュレーションの「聖杯」であると考えられています。この論文では、この定説は間違っていると主張しています。B300 世代以降の AI に最適化された GPU では、豊富な FP8 テンソル スループットと中国剰余定理ベースの Ozaki Scheme II を組み合わせることで、正規の HPC カーネル スペクトル全体で完全な FP64 精度でメモリルーフ実行を回復します。 NVIDIA の Blackwell Ultra (B300) は、ネイティブ FP64 を約 1.3 TFLOPS (B200 から 31 倍) に低下させ、メモリに依存するカーネル (SpMV、GEMV、ステンシル) も計算に依存するようにレンダリングします。私たちは4つの貢献をしています。まず、統合分析モデルである Tensor-Memory Equilibrium (TME) モデルは、計算乗数アルファ、帯域幅乗数ベータ、および再構築レイテンシ ガンマでルーフラインを強化します。次に、ベータ -> 1 を駆動するメカニズムとしてレジスタレベルの融合を特定し、エミュレーションをメモリの壁の向こう側で本質的に自由にします。 3 番目に、Ozaki II ヴォールトは FP64 をネイティブの最大 1 TFLOPS から最大 500 TFLOPS (B300) および最大 400 TFLOPS (Rubin R200) までエミュレートし、帯域幅制限の領域ではメモリ上限に匹敵しながら、コンピューティング領域では B200 のネイティブ FP64 の上限を 1 桁以上上回ったと予測します。 4 番目に、H100 ベースラインに対して、Ozaki II は、B300 ネイティブ FP64 が課す最大 50 倍の回帰と比較して、調査したすべてのワークロードで H100 と一致またはそれを超えています。コンパニオン FFT 解析 (生き残った INT32 パイプでの Kulisch 固定点再構築) と、コンパニオン Part(2) 論文で報告されている FP32+Kahan 削減と組み合わせると、B300 で調査されたすべてのカーネル クラスがフル FP64 でメモリ ルーフに達します。証拠はタイトルの主張を裏付けています。Ozaki II と Kulisch のエスケープ ルートを備えた FP8 は、実稼働 HPC に必要なすべてです。ネイティブ FP64 シリコンは、もはやこれまで考えられてきた聖杯ではありません。

原文 (English)

FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (June 13th version)

Conventional HPC holds that native hardware FP64 is the irreducible foundation of scientific computing. On AI-optimized GPUs of the NVIDIA B300 generation and beyond, native FP64 throughput has collapsed to ~1.3 TFLOPS even as FP8 tensor throughput has grown to multiple PFLOPS. We argue something stronger than that this is survivable: the FP8 tensor-core matrix-multiply is the sole computational primitive on which double-precision scientific computing needs to be built. Every canonical kernel -- dense and sparse linear algebra, spectral transforms, stencils -- and every application composing them reduces, via the Chinese Remainder Theorem-based Ozaki Scheme II, to sequences of FP8 matrix operations; the only non-FP8 arithmetic is a bounded, fixed-width integer accumulation at reconstruction. Native FP64 is thereby demoted from a hardware requirement to a derived accuracy guarantee obtained by composition over the FP8 primitive. We organize the claim as a five-layer hierarchy -- the FP8 op, Ozaki II, the basic kernels or Berkeley "dwarfs", composite solvers, and full applications -- and, because the dwarf taxonomy already spans scientific computing, establish it by exhibiting the reduction for every dwarf rather than a sample. The claim is falsifiable, and we build the instrument that tests it: a Tensor-Memory Equilibrium (TME) model extending the Roofline with emulation parameters (alpha, beta, gamma). We identify register-level fusion as the mechanism that keeps emulation memory-bound, project recovered FP64 performance across B300 and Rubin against an H100 baseline, and close the kernel coverage with a companion FFT analysis and compensated reductions. The model could have returned a negative verdict; instead it passes across the dwarfs and their compositions. This is the analytical half of a two-part program, with a follow-on implementation to validate the thesis on real silicon.

2026-07-07 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation

Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards.…

2026-07-06 10:22 JSTITmedia AI+ハードウェア/半導体

キオクシア、新型メモリのサンプル出荷開始 岩手の工場の最先端設備活用

半導体大手キオクシアホールディングスが、大容量で低消費電力の新型メモリのサンプル品の出荷を開始したと発表した。北上工場(岩手県北上市)第2製造棟の最先端設備を活用して生産する。

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

LLM のオンライン安全監視

アライメントトレーニングにもかかわらず、LLM は依然として展開時に安全でない出力を生成する傾向があります。したがって、出力をオンラインで監視し、安全性が確保できなくなった場合に警報を発することが重要です。私たちは、外部モデルからの検証信号を閾値処理によってアラームの決定に変換する単純なリアルタイム モニターを研究します。閾値はリスク コントロールによって調整されます。数学的推論とレッドチームデータセットの実験では、このシンプルなデザインが、逐次仮説テストに基づくより高度なモニターと競合できることを示しました。

原文 (English)

Online Safety Monitoring for LLMs

Despite alignment training, LLMs remain prone to generating unsafe outputs at deployment time. Monitoring outputs online and raising an alarm when safety can no longer be assumed is therefore critical. We study a simple real-time monitor that turns a verifier signal from an external model into an alarm decision by thresholding, with the threshold calibrated via risk control. In experiments on mathematical reasoning and red teaming datasets, we show that this simple design is competitive with more advanced monitors based on sequential hypothesis testing.

2026-07-03 13:00 JSTarXiv cs.AIハードウェア/半導体

GPUAlert: A Zero-Instrumentation Process-Boundary Monitor for Diagnosing GPU Training-Job Failures

GPU training jobs fail often, roughly two in five on large production clusters, yet the operator typically learns of a failure only by reco…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Towards Load-Aware Prefill Deflection for Disaggregated LLM Serving

Disaggregated LLM serving runs prefill and decode on separate GPU pools to keep the two phases from interfering. In practice, this creates…

2026-07-03 13:00 JSTarXiv cs.AIハードウェア/半導体

Causal Explanations for Image Classifiers

Existing algorithms for explaining the output of image classifiers use different definitions of explanations and a variety of techniques to…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

Introduction to Transformers: an NLP Perspective

Transformers have dominated empirical machine learning models of natural language processing. In this paper, we introduce basic concepts of…

2026-07-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

TileFuse: AMD NPU での効率的な量子化 LLM 推論のための融合型混合精度カーネル ライブラリ

オンデバイス LLM 推論に対する需要の高まりに伴い、厳しい電力と熱のバジェットの下でパフォーマンスとエネルギー効率を向上させるために、エッジ SoC は NPU を統合することが増えています。しかし、現在のクライアント NPU に実際に LLM を導入することは依然として困難です。AWQ などの広く使用されている量子化形式は、多くの既存の NPU ソフトウェア スタックにきれいにマッピングされておらず、多くの場合独自仕様であり、低レベルの制御が制限されています。この研究では、量子化 LLM 推論におけるトランス線形層をターゲットとする、AMD XDNA2 NPU 用のメタルに近い混合精度カーネル ライブラリである \textit{TileFuse} を紹介します。 TileFuse は、NPU 固有の量子化スキームに基づいてモデルを強制的に再形成するのではなく、AWQ スタイルの W4A16 や W8A16 などの実用的な低ビット フォーマットを XDNA2 に直接導入します。 TileFuse は、重みレイアウト、メタデータ配置、混合精度マイクロカーネル、配列レベルのデータフローを共同設計します。具体的には、アンパッキング、逆量子化、GEMM/GEMV の実行を単一のカーネル フローに融合し、最大 32K の GEMM 次元をサポートするインターリーブ プレタイリング レイアウトを導入し、完全な 4x8 AIE アレイを利用するように GEMV データフローを再設計します。カーネル レベルの評価全体で、TileFuse は、完全精度のベースラインと比較して、GEMM で最大 121.6%、GEMV で 281% パフォーマンスが向上し、GEMM 上の強力な iGPU ベースラインと比較して 2 倍を超えるパフォーマンスとエネルギー効率の向上を実現します。 Ryzen AI ラップトップでのエンドツーエンド LLM 実験では、TileFuse は、エネルギー消費量を 64.6% 以上削減し、プレフィル レイテンシーを最大 2.0 倍短縮することを達成しました。これらの結果を総合すると、XDNA2 が AWQ スタイルのエッジ LLM 推論の実用的なターゲットであること、および既製の量子化に対するネイティブ NPU サポートにより、実際のクライアント展開で NPU が大幅に使いやすくなることがわかります。

原文 (English)

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs

With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets. However, practical LLM deployment on current client NPUs remains difficult: widely used quantization formats such as AWQ do not map cleanly onto many existing NPU software stacks, which are often proprietary and expose limited low-level control. In this work, we present TileFuse, a close-to-metal mixed-precision kernel library for AMD XDNA2 NPUs that targets GEMM/GEMV-based operators in quantized LLM inference. TileFuse brings practical low-bit formats such as AWQ-style W4A16 and W8A16 directly onto XDNA2, rather than forcing the model to be reshaped around an NPU-specific quantization scheme. TileFuse co-designs weight layout, metadata placement, mixed-precision microkernels, and array-level dataflow. Specifically, it fuses unpacking, dequantization, and GEMM/GEMV execution into a single kernel flow, introduces an interleaved pre-tiling layout that supports GEMM dimensions up to 32K, and redesigns GEMV dataflow to utilize the full 4x8 AIE array. Across kernel-level evaluations, TileFuse improves performance by up to 121.6% for GEMM and 281% for GEMV over full-precision baselines, while delivering more than 2x performance and energy-efficiency gains over strong iGPU baselines on GEMM. In end-to-end LLM experiments on Ryzen AI laptops, TileFuse achieves up to 2.0x lower prefilling latency with more than 64.6% lower energy consumption. Together, these results show that XDNA2 is a practical target for AWQ-style edge LLM inference and that native NPU support for off-the-shelf quantization can make NPUs substantially more usable in real client deployments.

2026-07-03 03:31 JSTTechCrunch AILLM/生成AIハードウェア/半導体

Anthropic is discussing a new custom chip with Samsung

The news comes about a week after OpenAI announced its own custom AI chip in a partnership with Broadcom.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

強化されたライティング支援のための制御可能なナラティブ レンダリング

基本的なライティング支援における大規模言語モデル (LLM) の優れた能力にもかかわらず、創造的なライティングにおける LLM の有用性は、永続的なバイナリ障害によって根本的に妨げられています。この問題は、修復研磨と呼ばれる安全な表面レベルの編集と、破壊的で制御されていないプロット拡張の間の揺れとして現れます。このジレンマは、物語の忠実さと説明の強度との間の重要なトレードオフを定義します。我々は、物語と言説の間の物語学的区別に基づいた執筆支援フレームワークである Loom を提案します。 Loom は、意図を中心とした記号論的思考連鎖を運用する 3 層パイプラインを採用し、物語の意図とレンダリング密度を正確に制御します。このアーキテクチャは、知覚マテリアルの生成を構文の挿入から分離し、元のイベント構造に違反することなく強化が確実に行われるようにします。 LLM ベースの指標と人間による評価を含む当社の包括的な評価は、Loom がこの根本的な緊張をうまく解決していることを示しています。 Loom は最高の総合品質スコアを達成し、最先端のベースラインと比較して事実の完全性と説明の強度が大幅に向上しました。

原文 (English)

Controllable Narrative Rendering for Enhanced Assisted Writing

Despite the remarkable proficiency of large language models (LLMs) in basic writing assistance, their utility in creative writing is fundamentally hindered by a persistent binary failure. This issue manifests as an oscillation between safe, surface-level editing, referred to as remedial polishing, and destructive, uncontrolled plot expansion. This dilemma defines a critical trade-off between narrative fidelity and descriptive intensity. We propose Loom, an assisted writing framework grounded in the narratological distinction between story and discourse. Loom employs a three-layer pipeline that operationalizes an intent-centered semiotic chain-of-thought to enforce precise control over narrative intent and rendering density. This architecture separates the generation of perceptual material from syntactic insertion, ensuring that enhancement occurs without violating the original event structure. Our comprehensive evaluation, which includes LLM-based metrics and human assessment, demonstrates that Loom successfully resolves this fundamental tension. Loom achieves the highest overall quality score, yielding substantial gains in factual integrity and descriptive intensity compared to state-of-the-art baselines.

2026-07-02 13:00 JSTarXiv cs.AIハードウェア/半導体

SNAP-FM: 物理制約付き生成モデリングのためのスパース非線形加速投影

生成モデルは、物理シミュレーションのスケーラブルな代用として登場しましたが、その出力が保存則、境界条件、および基礎となる物理を支配する非線形不変量を尊重しているという保証はありません。制約付きサンプリングはこのギャップを埋め、再トレーニングせずに推論時にそのような制約を正確に適用しますが、計算コストがかかります。サンプリング中に投影、補正、軌道最適化のステップが繰り返され、非線形制約の場合はこれらのステップが高価になります。標準の ML フレームワークはこれをさらに悪化させます。高密度のテンソル代数と限られたスパース ソルバーの構成可能性により、物理的制約が自然に引き起こす構造がわかりにくくなり、効率的なバッチ非線形最適化を実際に実現することが困難になります。私たちは、サンプルごとのバッチ処理とローカル偏微分方程式結合が射影副問題で引き起こす構造、つまりブロック疎ヤコビアンと KKT システムを利用することで、このボトルネックに対処します。ExaModels.jl を使用してこの構造を公開し、MadNLP.jl と GPU 疎因数分解を使用して結果の疎非線形プログラムを解きます。このアプローチは、線形、非線形、1 次元、および 2 次元の制約を持つ PDE ベンチマークで物理制約フロー マッチング (PCFM) に適用されるため、制約の満足度を維持しながら非線形制約の投影を高速化します。これらの結果は、スパース GPU 非線形最適化が科学機械学習における制約付き生成サンプリングの実用的な基盤であることを示しています。

原文 (English)

SNAP-FM: Sparse Nonlinear Accelerated Projection for Physics-Constrained Generative Modeling

Generative models have emerged as scalable surrogates for physical simulation, yet they offer no guarantee that their outputs respect the conservation laws, boundary conditions, and nonlinear invariants that govern the underlying physics. Constrained sampling closes this gap, enforcing such constraints exactly at inference time without retraining, but at a computational cost: projection, correction, and trajectory-optimization steps are repeated during sampling, with these steps becoming expensive for nonlinear constraints. Standard ML frameworks exacerbate this: their dense tensor algebra and limited sparse solver composability obscure the structure that physical constraints naturally induce, making efficient batched nonlinear optimization difficult to realize in practice. We address this bottleneck by exploiting the structure that sample-wise batching and local PDE couplings induce in the projection subproblems -- namely, block-sparse Jacobian and KKT systems -- exposing this structure using ExaModels.jl and solving the resulting sparse nonlinear programs with MadNLP.jl and GPU sparse factorization. Applied to Physics-Constrained Flow Matching (PCFM), on PDE benchmarks with linear, nonlinear, one-dimensional, and two-dimensional constraints, this approach accelerates nonlinear constraint projection while maintaining constraint satisfaction. These results show that sparse GPU nonlinear optimization is a practical foundation for constrained generative sampling in scientific machine learning.

2026-07-02 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体研究/論文

非線形およびニューラル ネットワーク ダイナミクスのリアルタイムの堅牢な最適制御のための GPU 並列線形化誤差限界

この論文では、不確実な非線形システムに対するリアルタイムのロバストな最適制御について研究します。このシステムでは、線形時変 (LTV) 近似により計画が扱いやすくなりますが、ロバストな制約を満たすために健全な線形化誤差限界 (LEB) が必要になります。私たちは、非線形およびニューラル ネットワーク (NN) ダイナミクスの LTV 近似用に、厳密で微分可能な GPU 並列 LEB を開発します。解析ダイナミクスのために、標準的な間隔法よりも厳しいパスベースのヘシアン境界を導入します。 NN ダイナミクスの場合、NN 検証者が生成したアフィン緩和とローカル ヤコビアン補正を使用して、認定された LEB を導出します。我々は、GPU 並列システムレベル合成 LTV ベースのロバスト制御ソルバーを、タイトなゾノトピック不確実性伝播のための右可逆外乱行列と非ゼロ中心外乱セットを処理できるように拡張することで、これらの LEB と互換性を持たせるように適合させます。私たちの手法である GPUSLS-LEO は、線形化誤差を考慮した堅牢なフィードバック ポリシーのオンライン最適化を可能にし、厳密で正式に検証された到達可能なチューブを生成します。複雑な非線形および最大 168 状態次元の NN ダイナミクスにおいて、私たちの手法は GPU 上で最大 67 Hz のレートで堅牢な制御ポリシーを計算でき、形式的な保証とリアルタイムのパフォーマンスを維持しながら、解決時間とベースラインに対する保守性を削減します。

原文 (English)

GPU-Parallel Linearization Error Bounds for Real-Time Robust Optimal Control of Nonlinear and Neural Network Dynamics

This paper studies real-time robust optimal control for uncertain nonlinear systems, where linear time-varying (LTV) approximations make planning tractable but require sound linearization error bounds (LEBs) to guarantee robust constraint satisfaction. We develop tight, differentiable, GPU-parallel LEBs for LTV approximations of nonlinear and neural network (NN) dynamics. For analytic dynamics, we introduce path-based Hessian bounds that are tighter than standard interval methods. For NN dynamics, we derive certified LEBs using NN verifier-generated affine relaxations and local Jacobian corrections. We adapt a GPU-parallel system-level synthesis LTV-based robust control solver to be compatible with these LEBs by extending it to handle right-invertible disturbance matrices and non-zero-centered disturbance sets for tight zonotopic uncertainty propagation. Our method, GPUSLS-LEO, enables online optimization of robust feedback policies that account for linearization error, producing tight, formally verified reachable tubes. On complex nonlinear and NN dynamics up to 168 state dimensions, our method can compute robust control policies on the GPU at rates up to 67 Hz, reducing solve times and conservativeness relative to baselines while preserving formal guarantees and real-time performance.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

ロジック グリッド パズルによる LLM 推論の暗黙的なバイアスの評価

最近の安全ガードレールは、あからさまに偏った出力を効果的に抑制しますが、現在の評価ベンチマークを回避する複雑な論理的推論タスク中には、より微妙な形の社会的バイアスが出現します。このギャップを埋めるために、新しい評価フレームワーク PRIME (Puzzle Reasoning for Implicit Biases in Model Evaluation) を導入します。このフレームワークは、論理グリッド パズルを使用して、LLM の論理的推論と意思決定に対する社会的固定観念の影響を系統的に調査します。ロジック パズルを使用することで、自動生成と検証が可能になるだけでなく、複雑さや偏った設定の多様性も可能になります。 PRIME には、共有パズル構造から生成された定型パズル、反定型パズル、および中立的なパズルのバリアントが含まれており、制御されたきめ細かい比較が可能です。パズルのサイズ全体で複数のモデル ファミリを評価し、プロンプトベースの緩和戦略の有効性をテストします。性別のステレオタイプに焦点を当てた実験に焦点を当てた私たちの調査結果は、解決策がステレオタイプの関連性と一致する場合、モデルが一貫してより正確に推論することを強調しています。これは、公平性が重要である LLM の演繹的推論で永続する社会的偏見を診断し定量化するための PRIME の重要性を示しています。

原文 (English)

Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles

While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation benchmarks. To fill this gap, we introduce a new evaluation framework, PRIME (Puzzle Reasoning for Implicit Biases in Model Evaluation), that uses logic grid puzzles to systematically probe the influence of social stereotypes on logical reasoning and decision making in LLMs. Our use of logic puzzles enables automatic generation and verification, as well as variability in complexity and biased settings. PRIME includes stereotypical, anti-stereotypical, and neutral puzzle variants generated from a shared puzzle structure, allowing for controlled and fine-grained comparisons. We evaluate multiple model families across puzzle sizes and test the effectiveness of prompt-based mitigation strategies. Focusing our experiments on gender stereotypes, our findings highlight that models consistently reason more accurately when solutions align with stereotypical associations. This demonstrates the significance of PRIME for diagnosing and quantifying social biases perpetuated in the deductive reasoning of LLMs, where fairness is critical.

2026-07-02 13:00 JSTarXiv cs.AIハードウェア/半導体

Deep Learning-Driven Black-Box Doherty Power Amplifier with Pixelated Output Combiner and Extended Efficiency Range

This article presents a deep learning-driven inverse design methodology for Doherty power amplifiers (PA) with multi-port pixelated output…

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

強化されたライティング支援のための制御可能なナラティブ レンダリング

基本的なライティング支援における大規模言語モデル (LLM) の優れた能力にもかかわらず、創造的なライティングにおける LLM の有用性は、永続的なバイナリ障害によって根本的に妨げられています。この問題は、修復研磨と呼ばれる安全な表面レベルの編集と、破壊的で制御されていないプロット拡張の間の揺れとして現れます。このジレンマは、物語の忠実さと説明の強度との間の重要なトレードオフを定義します。我々は、物語と言説の間の物語学的区別に基づいた執筆支援フレームワークである Loom を提案します。 Loom は、意図を中心とした記号論的思考連鎖を運用する 3 層パイプラインを採用し、物語の意図とレンダリング密度を正確に制御します。このアーキテクチャは、知覚マテリアルの生成を構文の挿入から分離し、元のイベント構造に違反することなく強化が確実に行われるようにします。 LLM ベースの指標と人間による評価を含む当社の包括的な評価は、Loom がこの根本的な緊張をうまく解決していることを示しています。 Loom は最高の総合品質スコアを達成し、最先端のベースラインと比較して事実の完全性と説明の強度が大幅に向上しました。

原文 (English)

Controllable Narrative Rendering for Enhanced Assisted Writing

Despite the remarkable proficiency of large language models (LLMs) in basic writing assistance, their utility in creative writing is fundamentally hindered by a persistent binary failure. This issue manifests as an oscillation between safe, surface-level editing, referred to as remedial polishing, and destructive, uncontrolled plot expansion. This dilemma defines a critical trade-off between narrative fidelity and descriptive intensity. We propose Loom, an assisted writing framework grounded in the narratological distinction between story and discourse. Loom employs a three-layer pipeline that operationalizes an intent-centered semiotic chain-of-thought to enforce precise control over narrative intent and rendering density. This architecture separates the generation of perceptual material from syntactic insertion, ensuring that enhancement occurs without violating the original event structure. Our comprehensive evaluation, which includes LLM-based metrics and human assessment, demonstrates that Loom successfully resolves this fundamental tension. Loom achieves the highest overall quality score, yielding substantial gains in factual integrity and descriptive intensity compared to state-of-the-art baselines.

2026-07-02 13:00 JSTarXiv cs.AIハードウェア/半導体

SNAP-FM: 物理制約付き生成モデリングのためのスパース非線形加速投影

生成モデルは、物理シミュレーションのスケーラブルな代用として登場しましたが、その出力が保存則、境界条件、および基礎となる物理を支配する非線形不変量を尊重しているという保証はありません。制約付きサンプリングはこのギャップを埋め、再トレーニングせずに推論時にそのような制約を正確に適用しますが、計算コストがかかります。サンプリング中に投影、補正、軌道最適化のステップが繰り返され、非線形制約の場合はこれらのステップが高価になります。標準の ML フレームワークはこれをさらに悪化させます。高密度のテンソル代数と限られたスパース ソルバーの構成可能性により、物理的制約が自然に引き起こす構造がわかりにくくなり、効率的なバッチ非線形最適化を実際に実現することが困難になります。私たちは、サンプルごとのバッチ処理とローカル偏微分方程式結合が射影副問題で引き起こす構造、つまりブロック疎ヤコビアンと KKT システムを利用することで、このボトルネックに対処します。ExaModels.jl を使用してこの構造を公開し、MadNLP.jl と GPU 疎因数分解を使用して結果の疎非線形プログラムを解きます。このアプローチは、線形、非線形、1 次元、および 2 次元の制約を持つ PDE ベンチマークで物理制約フロー マッチング (PCFM) に適用されるため、制約の満足度を維持しながら非線形制約の投影を高速化します。これらの結果は、スパース GPU 非線形最適化が科学機械学習における制約付き生成サンプリングの実用的な基盤であることを示しています。

原文 (English)

SNAP-FM: Sparse Nonlinear Accelerated Projection for Physics-Constrained Generative Modeling

Generative models have emerged as scalable surrogates for physical simulation, yet they offer no guarantee that their outputs respect the conservation laws, boundary conditions, and nonlinear invariants that govern the underlying physics. Constrained sampling closes this gap, enforcing such constraints exactly at inference time without retraining, but at a computational cost: projection, correction, and trajectory-optimization steps are repeated during sampling, with these steps becoming expensive for nonlinear constraints. Standard ML frameworks exacerbate this: their dense tensor algebra and limited sparse solver composability obscure the structure that physical constraints naturally induce, making efficient batched nonlinear optimization difficult to realize in practice. We address this bottleneck by exploiting the structure that sample-wise batching and local PDE couplings induce in the projection subproblems -- namely, block-sparse Jacobian and KKT systems -- exposing this structure using ExaModels.jl and solving the resulting sparse nonlinear programs with MadNLP.jl and GPU sparse factorization. Applied to Physics-Constrained Flow Matching (PCFM), on PDE benchmarks with linear, nonlinear, one-dimensional, and two-dimensional constraints, this approach accelerates nonlinear constraint projection while maintaining constraint satisfaction. These results show that sparse GPU nonlinear optimization is a practical foundation for constrained generative sampling in scientific machine learning.

2026-07-02 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体研究/論文

非線形およびニューラル ネットワーク ダイナミクスのリアルタイムの堅牢な最適制御のための GPU 並列線形化誤差限界

この論文では、不確実な非線形システムに対するリアルタイムのロバストな最適制御について研究します。このシステムでは、線形時変 (LTV) 近似により計画が扱いやすくなりますが、ロバストな制約を満たすために健全な線形化誤差限界 (LEB) が必要になります。私たちは、非線形およびニューラル ネットワーク (NN) ダイナミクスの LTV 近似用に、厳密で微分可能な GPU 並列 LEB を開発します。解析ダイナミクスのために、標準的な間隔法よりも厳しいパスベースのヘシアン境界を導入します。 NN ダイナミクスの場合、NN 検証者が生成したアフィン緩和とローカル ヤコビアン補正を使用して、認定された LEB を導出します。我々は、GPU 並列システムレベル合成 LTV ベースのロバスト制御ソルバーを、タイトなゾノトピック不確実性伝播のための右可逆外乱行列と非ゼロ中心外乱セットを処理できるように拡張することで、これらの LEB と互換性を持たせるように適合させます。私たちの手法である GPUSLS-LEO は、線形化誤差を考慮した堅牢なフィードバック ポリシーのオンライン最適化を可能にし、厳密で正式に検証された到達可能なチューブを生成します。複雑な非線形および最大 168 状態次元の NN ダイナミクスにおいて、私たちの手法は GPU 上で最大 67 Hz のレートで堅牢な制御ポリシーを計算でき、形式的な保証とリアルタイムのパフォーマンスを維持しながら、解決時間とベースラインに対する保守性を削減します。

原文 (English)

GPU-Parallel Linearization Error Bounds for Real-Time Robust Optimal Control of Nonlinear and Neural Network Dynamics

This paper studies real-time robust optimal control for uncertain nonlinear systems, where linear time-varying (LTV) approximations make planning tractable but require sound linearization error bounds (LEBs) to guarantee robust constraint satisfaction. We develop tight, differentiable, GPU-parallel LEBs for LTV approximations of nonlinear and neural network (NN) dynamics. For analytic dynamics, we introduce path-based Hessian bounds that are tighter than standard interval methods. For NN dynamics, we derive certified LEBs using NN verifier-generated affine relaxations and local Jacobian corrections. We adapt a GPU-parallel system-level synthesis LTV-based robust control solver to be compatible with these LEBs by extending it to handle right-invertible disturbance matrices and non-zero-centered disturbance sets for tight zonotopic uncertainty propagation. Our method, GPUSLS-LEO, enables online optimization of robust feedback policies that account for linearization error, producing tight, formally verified reachable tubes. On complex nonlinear and NN dynamics up to 168 state dimensions, our method can compute robust control policies on the GPU at rates up to 67 Hz, reducing solve times and conservativeness relative to baselines while preserving formal guarantees and real-time performance.

2026-07-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

ロジック グリッド パズルによる LLM 推論の暗黙的なバイアスの評価

最近の安全ガードレールは、あからさまに偏った出力を効果的に抑制しますが、現在の評価ベンチマークを回避する複雑な論理的推論タスク中には、より微妙な形の社会的バイアスが出現します。このギャップを埋めるために、新しい評価フレームワーク PRIME (Puzzle Reasoning for Implicit Biases in Model Evaluation) を導入します。このフレームワークは、論理グリッド パズルを使用して、LLM の論理的推論と意思決定に対する社会的固定観念の影響を系統的に調査します。ロジック パズルを使用することで、自動生成と検証が可能になるだけでなく、複雑さや偏った設定の多様性も可能になります。 PRIME には、共有パズル構造から生成された定型パズル、反定型パズル、および中立的なパズルのバリアントが含まれており、制御されたきめ細かい比較が可能です。パズルのサイズ全体で複数のモデル ファミリを評価し、プロンプトベースの緩和戦略の有効性をテストします。性別のステレオタイプに焦点を当てた実験に焦点を当てた私たちの調査結果は、解決策がステレオタイプの関連性と一致する場合、モデルが一貫してより正確に推論することを強調しています。これは、公平性が重要である LLM の演繹的推論で永続する社会的偏見を診断し定量化するための PRIME の重要性を示しています。

原文 (English)

Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles

While recent safety guardrails effectively suppress overtly biased outputs, subtler forms of social bias emerge during complex logical reasoning tasks that evade current evaluation benchmarks. To fill this gap, we introduce a new evaluation framework, PRIME (Puzzle Reasoning for Implicit Biases in Model Evaluation), that uses logic grid puzzles to systematically probe the influence of social stereotypes on logical reasoning and decision making in LLMs. Our use of logic puzzles enables automatic generation and verification, as well as variability in complexity and biased settings. PRIME includes stereotypical, anti-stereotypical, and neutral puzzle variants generated from a shared puzzle structure, allowing for controlled and fine-grained comparisons. We evaluate multiple model families across puzzle sizes and test the effectiveness of prompt-based mitigation strategies. Focusing our experiments on gender stereotypes, our findings highlight that models consistently reason more accurately when solutions align with stereotypical associations. This demonstrates the significance of PRIME for diagnosing and quantifying social biases perpetuated in the deductive reasoning of LLMs, where fairness is critical.

2026-07-02 13:00 JSTarXiv cs.AIハードウェア/半導体

Deep Learning-Driven Black-Box Doherty Power Amplifier with Pixelated Output Combiner and Extended Efficiency Range

This article presents a deep learning-driven inverse design methodology for Doherty power amplifiers (PA) with multi-port pixelated output…

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

LLM における一貫性のジレンマ: 生成者と評価者の合意と間違いに対する脆弱性

大規模な言語モデルは、外部検証なしで独自の出力を評価するモデルに依存するエージェント パイプラインにデプロイされることが増えています。これらのパイプラインの信頼性は、モデルが出力を生成し、後でその出力を評価するときに、関連する概念を同じ方法で適用するという暗黙の仮定に依存します。我々は、この仮定を直接テストし、それを 491 の概念にわたる 10 のフロンティア モデルに適用するための、新しい尺度である生成器と評価器の自己一貫性を提案します。まず、自己一貫性には大きなばらつきがあることがわかりました。第 2 に、医師によって検証された間違いのある臨床現場では (Proniakin et al., 2025)、モデル全体で自己一貫性が高いモデルは間違いに対する脆弱性がより高いことに関連していることがわかりました。したがって、モデルが一貫して概念を適用している場合でも、展開するのが安全ではない可能性があります。これは、LLM における一貫性のジレンマの証拠です。つまり、自己一貫性は運用上役立ちますが、モデルの一貫性が高いほど間違いが発生しやすくなります。

原文 (English)

The Consistency Dilemma in LLMs: Generator-Evaluator Agreement and Vulnerability to Mistakes

Large language models are increasingly deployed in agentic pipelines that depend on the model evaluating its own outputs without external verification. The reliability of these pipelines depends on an implicit assumption: that the model applies relevant concepts the same way when it generates an output and later evaluates that output. We propose a new measure, generator-evaluator self-consistency, to test this assumption directly and apply it to 10 frontier models across 491 concepts. We find, first, that there is substantial variation in self-consistency. Second, we find that in a clinical setting with physician-validated mistakes (Proniakin et al., 2025), across models, those with higher self-consistency are linked to greater vulnerability to mistakes. Thus, even when models consistently apply concepts they may not be safe to deploy. This is evidence of a consistency dilemma in LLMs: self-consistency is operationally useful, but models that are more consistent are also more prone to mistakes.

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

仕様駆動開発における引用規律: LLM で生成されたコードにおける出力決定論と自動幻覚検出に関するクロスモデル実証研究

仕様駆動開発 (SDD) フレームワークは、正式な仕様を通じて大規模言語モデル (LLM) を利用したコード生成をガイドしますが、要件と生成されたコードの間のトレーサビリティを強化する方法が根本的に異なります。この論文では、3 つの SDD フレームワークを比較する 2 つの管理された実証研究を紹介します。 $Spec Kit$ は、ユーザー ストーリーと受け入れ基準を通じて成果物レベルのトレーサビリティを使用します。 $OpenSpec$ はポストホック外部トレース マップに依存します。 2 つのフロンティア LLM、Claude Sonnet 4.6 (N=20、4 条件、240 実装) と GLM-5-turbo (N=50、4 条件、600 実装) にわたる 2 つの主な結果を測定します。$output$ $determinism$ (独立した LLM セッション間の語彙類似性) と $automated$ $hallucination$ $detection$ $rate$ (TDR) です。事前に登録された分析により、一貫したモデル間で反復されたトレードオフが明らかになりました。引用されていない条件は、引用された条件よりも大幅に高い決定論を生成します (Claude: $d=-0.76$、$p=0.003$; GLM: $d=-0.72$、$p<0.001$)。一方、引用された条件のみが自動幻覚検出を可能にします (TDR: Claude 86.4%、GLM) 88.0%、対すべての代替案で 0%、両方の研究で FPR=0%)。 traceSDD (引用) は、決定論に関して $Spec Kit$ を大幅に上回っています (Claude: $d=0.47$、$p=0.049$、GLM: $d=0.42$、$p=0.003$) が、OpenSpec ではありません (Claude: $d=0.18$、$p=0.44$、GLM: $d=0.14$、$p=0.32$)。これらの発見は、引用アノテーションが検証可能性を得るために決定性を犠牲にし、このトレードオフがモデル アーキテクチャ全体で一般化することを証明します。

原文 (English)

Citation Discipline in Spec-Driven Development: A Cross-Model Empirical Study of Output Determinism and Automated Hallucination Detection in LLM-Generated Code

Spec-Driven Development (SDD) frameworks guide Large Language Model (LLM)-powered code generation through formal specifications, yet they differ fundamentally in how they enforce traceability between requirements and generated code. This paper presents two controlled empirical studies comparing three SDD frameworks: $traceSDD$, which enforces mandatory per-line requirement citations using hierarchical REQ-XXX.Y.Z identifiers; $Spec Kit$, which uses artifact-level traceability through user stories and acceptance criteria; and $OpenSpec$, which relies on post-hoc external trace maps. We measure two primary outcomes across two frontier LLMs -- Claude Sonnet 4.6 (N=20, 4 conditions, 240 implementations) and GLM-5-turbo (N=50, 4 conditions, 600 implementations): $output$ $determinism$ (lexical similarity across independent LLM sessions) and $automated$ $hallucination$ $detection$ $rate$ (TDR). Our pre-registered analysis reveals a consistent, cross-model replicated trade-off: the uncited condition produces significantly higher determinism than the cited condition (Claude: $d=-0.76$, $p=0.003$; GLM: $d=-0.72$, $p<0.001$), while only the cited condition enables automated hallucination detection (TDR: Claude 86.4%, GLM 88.0%, vs 0% for all alternatives, FPR=0% across both studies). traceSDD (cited) significantly outperforms $Spec Kit$ on determinism (Claude: $d=0.47$, $p=0.049$; GLM: $d=0.42$, $p=0.003$) but not OpenSpec (Claude: $d=0.18$, $p=0.44$; GLM: $d=0.14$, $p=0.32$). These findings establish that citation annotations trade determinism for verifiability, and that this trade-off generalizes across model architectures.

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達規制/政策

Probing Stylistic Appropriation using Large Language Models: An Evaluation Framework for Copyright Infringement under EU Law

Large language models (LLM) trained on web-scale corpora generate output that may infringe copyright, yet existing technical safeguards foc…

2026-07-01 13:00 JSTarXiv cs.AIハードウェア/半導体

TDGT: A Tabular Data Generation Toolkit supporting adaptive GPU-accelerated Bayesian mixture models, diffusion-based models, and latent-space generative modeling

The growing demand for privacy-preserving data sharing has positioned synthetic data generation as a critical component of responsible AI w…

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In…

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Learning by Surprise: Adaptive Mitigation of Model Collapse in Large Language Models

As AI-generated content increasingly populates the web, generative AI models are at growing risk of being trained on their own outputs, a p…

2026-07-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

An Efficient Heterogeneous Co-Design for Fine-Tuning on a Single GPU

Fine-tuning Large Language Models (LLMs) has become essential for domain adaptation, but its memory-intensive property exceeds the capabili…

2026-07-01 09:54 JSTITmedia AI+LLM/生成AIハードウェア/半導体研究/論文

Anthropic、科学研究向けAIワークベンチ「Claude Science」を発表──NVIDIAのBioNeMoツールキットと連携

Anthropicは、科学者が計算研究を一貫して行えるAI実行環境「Claude Science」を発表した。データベースやツールを1つのインタフェースに統合し、文献分析から論文執筆、図表作成まで対応する。NVIDIAのツールキットとも連携し、機密データを外部に送信しない設計が…

2026-07-01 03:13 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

Nvidia competitor Etched hits $5B valuation, $1B in sales for AI chip

Nvidia AI chip competitor Etched says it has already booked $1 billion under contract for the inference systems powered by its chip.

2026-06-30 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Sequential Fairness Auditing with Limited Output Access

External evaluations are becoming increasingly central to the governance of AI systems. In practice, however, independent auditors often ha…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Insidious by Design: Implications of Large Language Model algorithmic bias for the Global South

\begin{quote} The biases in Large Language Models' (LLMs) outputs remain inadequately theorised, particularly from the perspective of the G…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

How Do LLMs Cite? A Mechanistic Interpretation of Attribution in Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) aims to enhance the trustworthiness of Large Language Models (LLMs) by grounding their outputs in exte…

2026-06-30 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

FedLAS: Feature-Modulated Bidirectional Label Smoothing for Neural Network Calibration

Deep Neural Network (DNN) classifiers suffer from poor calibration when their softmax outputs (predictive confidence) deviate from the empi…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation

Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards.…

2026-06-30 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

HERO: Improving the Reliability and Sensitivity of Generative Model Evaluation Using Historical Data

Reliable generative AI models critically rely on expert human annotations to evaluate output quality, yet these "gold" labels are expensive…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation

Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale. The most accurate detection methods depen…

2026-06-30 13:00 JSTarXiv cs.AIハードウェア/半導体

Diff-MN: Diffusion Parameterized MoE-NCDE for Continuous Time Series Generation with Irregular Observations

Time series generation (TSG) is widely used across domains, yet most existing methods assume regular sampling and fixed output resolutions.…

2026-06-30 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained Decoding

Large language models (LLMs) are increasingly used to generate executable outputs, JSON objects, and API calls, where a single syntax error…

2026-06-30 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

SwarmX: Agentic Scheduling for Low-Latency Agentic Systems

Agentic AI applications compose multiple model calls and tool executions, creating new scheduling challenges for GPU-CPU clusters. Their in…

2026-06-30 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

CoStream: 一般化可能な複雑な操作のための単純な動作の構築

GPU を PCIe スロットに装着するなど、長期にわたる接触が多い複雑な操作タスクでは、ミリメートル単位の高精度と、新しいタスクに対するすぐに使える汎用性の両方が必要です。既存のパラダイムは両方を満たすのに苦労しています。古典的なパイプラインは、高精度の制御を実現するために脆弱なタスク固有のインターフェイスを使用しますが、新しいタスクに適応するにはコストのかかるパイプラインの再設計が必要です。一方、モノリシックなエンドツーエンドのポリシーはより優れた一般化を提供しますが、新しいデータで再トレーニングしない限り、複雑な分散外のタスクでは高精度が得られません。どちらのパラダイムも暗黙の前提を共有しています。つまり、操作機能を取得したら、それを自由に分解したり再構成したりするのではなく、厳格なパイプラインまたはモノリシック全体として展開する必要があるということです。この論文では、複雑な操作機能が単純で独立した動作の構成から自然に出現する可能性があることを示します。モノリシックなポリシーや厳格なパイプラインを展開するのではなく、基礎モデルと多様なセンシング モダリティを複数の構成可能なコア動作に統合するフレームワークを提案します。つまり、基礎モデルを介して空間制約を抽出するセマンティック動作。想像上のビデオ内のキーポイントを追跡することで軌道を予測する予測行動。そして高周波の触覚と力の補正を提供する反応的な動作。共有 $SE(3)$ インターフェイスでは、これらの出力は右乗算によって各制御ステップで単一のポーズ コマンドに構成され、準拠したコントローラーによって実行されます。私たちは、日常的な操作と精密な組み立てに及ぶ 8 つの現実世界のタスクを実証し、接触の多い組み立てとオブジェクトの転送で最も大きな効果を発揮し、実行中の手動による混乱からの堅牢な回復を示します。 {ウェブサイト:} https://costream-simple.github.io

原文 (English)

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation

Long-horizon, contact-rich complex manipulation tasks, such as seating a GPU into a PCIe slot, demand both millimeter high precision and out-of-the-box generalization to new tasks. Existing paradigms struggle to satisfy both: classical pipelines use brittle, task-specific interfaces to achieve high-precision control but require costly pipeline redesigns to adapt to new tasks, whereas monolithic end-to-end policies provide better generalization but lack high precision on complex, out-of-distribution tasks unless retrained with new data. Both paradigms share an implicit assumption: once a manipulation capability is acquired, it must be deployed as a rigid pipeline or monolithic whole, rather than being freely decomposed and recomposed. In this paper, we show that complex manipulation capabilities can emerge naturally from the composition of simple, independent behaviors. Rather than deploying a monolithic policy or a rigid pipeline, we propose CoStream, a framework orchestrating foundation models and diverse sensing modalities into multiple composable core behaviors: a semantic behavior extracting spatial constraints via foundation models; a predictive behavior forecasting trajectories by tracking keypoints in imagined videos; and a reactive behavior providing high-frequency tactile and force corrections. On a shared $SE(3)$ interface, these outputs compose by right-multiplication into a single pose command at each control step, executed by a compliant controller. We demonstrate CoStream on 8 real-world tasks spanning everyday manipulation and precision assembly, with the strongest gains in contact-rich assembly and object transfer, and show robust recovery from manual perturbations during execution. Website: https://costream-simple.github.io

2026-06-30 07:00 JSTITmedia AI+ハードウェア/半導体

AIで人は幸せになれるのか? 「AIで稼ぐ企業」と「コストを負担する企業」

Appleが複数製品の価格を引き上げる一方で、AI需要の拡大を追い風にキオクシアなど半導体関連企業は成長を続けている。技術革新が生む大きな利益の裏側で、誰が恩恵を受け、誰がコストを負担するのか。

2026-06-30 06:45 JSTITmedia AI+ハードウェア/半導体

ルネサスが2035年の売上高3倍増も視野に、AIで3段階の成長を目指す

ルネサス エレクトロニクスが同社の概況や事業方針などについて説明。足元で半導体市場の拡大をけん引するAIに焦点を当てた事業展開を強化し、AIインフラ、フィジカルAIとSDV、「Intelligence at the Edge」の3段階で優位なポジションを構築し成長を目指す。

2026-06-30 03:07 JSTTechCrunch AIハードウェア/半導体

South Korean tech giants commit over $550B to ease ‘RAMageddon’

The world's two largest memory chip companies vow to build more memory lab fabs as South Korea positions itself as an AI tech powerhouse co…

2026-06-29 22:00 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

Omen AI’s plan to optimize data centers is all wet

Omen AI raised a $31 million Series A to monitor chip coolant and stop bacterial outbreaks in data centers.

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

次の LLM の事前登録による LLM ベースの p-Hacking の軽減

大規模言語モデル (LLM) は、その出力が下流の仮説テストにフィードされるデータの生成、分類、および注釈付けに使用されることが増えています。ただし、LLM ベースの研究は簡単にハックできます。研究者は、目的の結果が得られるまでプロンプト、デコード パラメータ、または出力形式を調整できます。私たちは、LLM ベースの研究における p-hacking を軽減するためのプロトコルを提案します。実験と適格なモデルを事前登録し、事前登録後にリリースされる最初の適格な LLM 上で実行します。研究者は、現在のモデルに関する手順を最終決定し、一連の適格な将来のモデルとともに解析計画を事前登録し、その後リリースされる最初の適格なモデルに対して確認分析を実行します。このモデルはコミット時には存在しないため、ハッキングすることはできません。さらに、あるモデルを頻繁にハッキングする構成は、次のモデルには移行されません。真の値がわかっている 2 つのタスクでプロトコルを評価します。 4 つのプロバイダーの 20 のモデルと 11 の LLM 解析構成にわたって、プロトコルは 2 つのタスクのケースの 73.9% と 72.7% で p-hack の転送成功をブロックしたと考えられます。追加の分析により、いくつかのストレステストの下でも緩和効果が依然として大幅であることが明らかになりました。最後に、私たちはお金をかけて、独自のプロトコルに従い、実験を事前登録しました。事前に登録された実験では、プロトコルの有効性が確認されました。前のモデルをハッキングした 7 つの構成のうち、その後リリースされた最初の適格モデルの 6 つの構成ではハッキングが引き継がれませんでした。

原文 (English)

Mitigating LLM-based p-Hacking by Preregistering for the Next LLM

Large language models (LLMs) are increasingly used to generate, classify, and annotate data whose outputs feed downstream hypothesis tests. However, LLM-based research is easy to p-hack: a researcher can tune the prompts, decoding parameters, or output format until a desired result is reached. We propose a protocol to mitigate p-hacking in LLM-based research: preregistering the experiment and eligible models, and then running it on the first eligible LLM that is released after the preregistration. The researcher finalizes the procedure on current models, preregisters the analysis plan together with a set of eligible future models, and runs the confirmatory analysis on the first eligible model released afterward. Because this model does not exist at commitment time, it cannot be hacked against; furthermore, configurations that hack one model frequently do not transfer to the next. We evaluate the protocol on two tasks whose true values are known. Across 20 models from four providers and 11 LLM-analysis configurations, the protocol would have blocked successful transfer of the p-hack in 73.9% and 72.7% of cases in the two tasks. Additional analyses reveal that mitigation remains substantial under several stress tests. Finally, putting money where our mouth is, we followed our own protocol and preregistered our experiment. The preregistered experiment confirmed the protocol's effectiveness: out of the 7 configurations that hacked the prior model, the hacking failed to carry over in 6 configurations on the first eligible model released afterward.

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

スムーズな MMD アライメントによる LLM の数値予測の強化

大規模言語モデル (LLM) は、強力な一般機能にもかかわらず、出力が数値的に正確である必要がある場合には信頼性が低いままであることがよくあります。主な理由はトレーニングの目的です。標準のクロスエントロピーは数値トークンを非構造化カテゴリーとして扱い、その値のメトリック構造を無視します。この不一致は、数値トークンとグラフベースの滑らかさの上に値と距離のカーネルを組み込むことで、古典的な MMD に基づいて構築された Smooth Maximum Mean Discrepancy (SMMD) で解決されます。数値サブ語彙上で定義されたこのカーネルを使用すると、SMMD はカーネル マッチングを通じて予測数値分布をターゲットに合わせ、誘導されたカーネル グラフ上で予測ターゲットの残差を平滑化して、局所的な一貫性を促進します。複数のオープンウェイト LLM および VLM バックボーンにわたって、数学的推論、算術計算、時刻認識、チャートの質問応答という 4 つの数値ターゲット タスクで SMMD を評価します。 SMMD は、クロスエントロピーと最近の数値ターゲット損失の両方に対する精度を一貫して向上させます。分析では、MMD と滑らかさの間の相補的な効果が示され、距離ベースのカーネル設計の重要性が強調されています。コードは https://github.com/Zuozhuo/smmd-loss で入手できます。

原文 (English)

Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment

Despite their strong general capabilities, large language models (LLMs) often remain unreliable when outputs must be numerically precise. A key reason is the training objective: standard cross-entropy treats numeric tokens as unstructured categories and ignores the metric structure of their values. We address this mismatch with Smooth Maximum Mean Discrepancy (SMMD), which builds on the classic MMD by incorporating value-distance kernels over numeric tokens and graph-based smoothness. With this kernel defined over a numeric sub-vocabulary, SMMD aligns the predicted numeric distribution to the target via kernel matching and smooths the prediction-target residual over the induced kernel graph to encourage local consistency. We evaluate SMMD on four numeric-target tasks: mathematical reasoning, arithmetic calculation, clock-time recognition, and chart question answering, across multiple open-weight LLM and VLM backbones. SMMD consistently improves accuracy over both cross-entropy and recent numeric-target losses; analyses show complementary effects between MMD and smoothness and underscore the importance of distance-based kernel design. Code is available at https://github.com/Zuozhuo/smmd-loss.

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

キャリブレーションガイド付き LLM 圧縮の出力スペース割り当てコスト: 実証的研究

大規模言語モデル (LLM) のトレーニング不要の圧縮方法では、多くの場合、圧縮の決定をガイドするためにキャリブレーション データが使用されます。 ROCKET は、スパース辞書因数分解と多選択ナップザック問題 (MCKP) 割り当てを組み合わせた最近の手法で、出力再構成目的から層ごとの因数分解を導き出しますが、重み空間のフロベニウス誤差を MCKP 割り当てコストとして使用します。割り当てコストを出力空間目標に合わせることで、圧縮モデルの忠実度が向上するかどうかを調査します。 50\% 圧縮の Qwen3-8B では、ROCKET-ActCost は 8 つのゼロショット ベンチマーク全体で +0.8 パーセント高い平均精度を達成しました (53.1\% 対 52.3\%) が、WikiText の複雑さは 16\% 増加しました (61.46 対 52.98)。この精度と複雑さのトレードオフは、割り当て目標が異なれば、ダウンストリーム メトリックも異なることが明らかになります。重み空間誤差と出力空間誤差の間の高い相関関係 ($>$0.99) により、割り当ての発散が制限され、効果の大きさが適度であることが説明されます。 20\% 圧縮の Llama-3.2-1B では、2 つの方法はほぼ同じ結果 (53.3\% 対 53.5\% の精度、14.45 対 14.66 PPL) を生成します。これは、圧縮率が低い場合にはコスト関数の影響が小さいことを示唆しています。

原文 (English)

Output-Space Allocation Costs for Calibration-Guided LLM Compression: An Empirical Study

Training-free compression methods for large language models (LLMs) often use calibration data to guide compression decisions. ROCKET, a recent method combining sparse-dictionary factorization with multi-choice knapsack problem (MCKP) allocation, derives its per-layer factorization from an output reconstruction objective but uses weight-space Frobenius error as the MCKP allocation cost. We investigate whether aligning the allocation cost with the output-space objective improves compressed model fidelity. On Qwen3-8B at 50\% compression, our ROCKET-ActCost achieves +0.8 percentage points higher average accuracy across 8 zero-shot benchmarks (53.1\% vs 52.3\%), but increases WikiText perplexity by 16\% (61.46 vs 52.98). This accuracy-perplexity tradeoff reveals that different allocation objectives favor different downstream metrics. The high correlation ($>$0.99) between weight-space and output-space errors limits allocation divergence, explaining the modest effect size. On Llama-3.2-1B at 20\% compression, the two methods produce near-identical results (53.3\% vs 53.5\% accuracy, 14.45 vs 14.66 PPL), suggesting that the effect of the cost function is minor at lower compression ratios.

2026-06-29 13:00 JSTarXiv cs.AIハードウェア/半導体

OperatorSHAP: Fast and Accurate Shapley Value Estimation for Neural Operators

Understanding model predictions is essential for physical applications, where outputs often inform safety-critical decisions, such as struc…

2026-06-29 13:00 JSTarXiv cs.AIハードウェア/半導体

Can Generative Artificial Intelligence Survive Data Contamination? Theoretical Guarantees under Contaminated Recursive Training

As artificial intelligence (AI)-generated content proliferates, models are increasingly trained on their own outputs, risking progressive d…

2026-06-29 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making…

2026-06-29 13:00 JSTarXiv cs.AIハードウェア/半導体

必要なのは FP8 だけです (パート 2): Tensor-core Garner 再定式化と Kulisch エスケープ ルートによる効率的な Ozaki-Bailey スタイル FFT

NVIDIA の Blackwell Ultra (B300) は、FP64 ベクトル スループットを GPU あたり約 1.3 TFLOPS に削減します。これは、B200 の約 30 分の 1 であり、帯域幅が制限された FP64 ワークロードがメモリ制限にとどまるレベルをはるかに下回ります。 Ozaki Scheme II フレームワークは、仮数スライスされた中国剰余再構成を使用して FP8 テンソル コアを介して密行列乗算をルーティングすることにより、FP64 と同等のスループットを回復します。関連資料のパート (1) では、高密度 GEMM、バッチ GEMV、ステンシル、および SpMV について説明します。この論文では、5 番目の標準プリミティブである 3-D FFT を追加します。 FP8 テンソル コア上の両方の 1-D FFT GEMM を使用した Bailey 6 ステップ分解を介してエミュレートされた 3-D FFT である Ozaki-Bailey FFT を紹介します。 Bailey の小さな内因数 k ~ sqrt(N) (N=1024 の場合 k=32) は、カーネルを k << r^2 の状態に置き、3 番目の TME パラメーター ガンマ (再構成待ち時間) が償却ではなく結合します。 Garner の再構成は、フェーズ A (FP8/INT8 テンソル コアの内積、B300 の 1024^3 で約 1 ミリ秒) とフェーズ B (出力ごとの削減) に分割されます。 Kulisch 固定小数点完全演算は、完全に INT32 SIMT パイプ上で実行しながら完全な FP64 精度を維持するフェーズ B 再定式化であると認識します。閉じた形式の帯域幅パリティの下限を導出します。ネイティブ FP64 の下限は 1.56*B_HBM (8 TB/s で 12.5 TF) です。B300 の 1.3 TF は約 10 倍低く、Rubin の 33 TF は 4% 以内です。 Kulisch 避難ルートには、INT32 サブフロア 8.25*B_HBM と FP8 フロア 170*B_HBM が必要です。 B300 はその両方を満たします。フル FP64 での 1024^3 の予測は約 18 ミリ秒で、実質的には 12.9 ミリ秒のメモリ ルーフです。 GPU がネイティブ フロアまたは両方の Kulisch フロアを満たす場合、GPU はメモリルーフ FFT パリティを満たします。この予測が実際に当てはまれば、B300 はソフトウェアのみでフル FP64 FFT を実行できるようになり、libKulisch ライブラリとベンチマーク キャンペーンの動機付けとなります。

原文 (English)

FP8 is All You Need (Part 2): Efficient Ozaki-Bailey Style FFT Through Tensor-core Garner Reformulation and Kulisch Escape Route

NVIDIA's Blackwell Ultra (B300) cuts FP64 vector throughput to ~1.3 TFLOPS per GPU, roughly 30x below B200 and well below the level at which bandwidth-limited FP64 workloads stay memory-bound. The Ozaki Scheme II framework recovers FP64-equivalent throughput by routing dense matrix multiply through FP8 tensor cores with a mantissa-sliced Chinese-remainder reconstruction. A companion Part (1) paper covers dense GEMM, batched GEMV, stencils, and SpMV; this paper adds the fifth canonical primitive, the 3-D FFT. We present Ozaki-Bailey FFT, an emulated 3-D FFT via the Bailey six-step decomposition with both 1-D FFT GEMMs on FP8 tensor cores. Bailey's small inner factor k ~ sqrt(N) (k=32 for N=1024) puts the kernel in the regime k << r^2, where the third TME parameter gamma (reconstruction latency) binds rather than amortising. Garner reconstruction splits into Phase A (inner products on FP8/INT8 tensor cores, ~1 ms for 1024^3 on B300) and Phase B (per-output reduction). We identify Kulisch fixed-point complete arithmetic as a Phase B reformulation that keeps full FP64 accuracy while running entirely on the INT32 SIMT pipe. We derive closed-form bandwidth-parity floors. The native FP64 floor is 1.56*B_HBM (12.5 TF at 8 TB/s): B300's 1.3 TF sits ~10x below, Rubin's 33 TF within 4%. The Kulisch escape route needs an INT32 sub-floor 8.25*B_HBM and an FP8 floor 170*B_HBM; B300 meets both. The projection is ~18 ms for 1024^3 at full FP64, essentially the 12.9 ms memory roof. A GPU meets memory-roof FFT parity if it satisfies either the native floor or both Kulisch floors. If the projection holds in practice, B300 becomes viable for full-FP64 FFT through software alone, motivating a libKulisch library and benchmark campaign.

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CrossPool: KV キャッシュと重み分解によるコールド MoE モデルの効率的なマルチ LLM サービス

新興の LLM サービスは、多くの疎な MoE モデルをホストすることが増えていますが、ほとんどのモデルは疎なリクエストを受け取り、コールドのままです。これにより、GPU メモリの問題が発生します。モデルの重みは安定していてモデルによって決定されますが、KV キャッシュは一時的で需要によって決定されます。コールド モデルが同時にピーク KV キャッシュ要求に達することはほとんどないため、モデルごとに最悪の場合の KV 容量を確保するとメモリが無駄になります。代わりに、共有 KV キャッシュ プールを使用して、アクティブな需要を集約してプロビジョニングできます。ただし、重みと KV キャッシュがモノリシック GPU メモリ プールに残っている場合、KV キャッシュの共有は十分ではありません。静的重みは動的 KV キャッシュと競合し、コールドで同時実行トラフィックが少ない場合に KV ヘッドに制限された注意は、レプリケートされた KV 容量の一部のみを公開するため、GPU メモリ使用率が低くなり、ロングコンテキストのサポートが弱くなります。 CrossPool は、FFN 重みと KV キャッシュを 2 つの GPU メモリ プールに分離するコールド MoE モデル用のサービング エンジンです。1 つはコールド モデル全体で FFN 重みを統合する重みプール、もう 1 つは KV キャッシュにローカルな注意を保ちながらアクティブなリクエストを動的に処理する KV キャッシュ プールです。 CrossPool は、KV キャッシュ プランナーとバーチャライザー、隠し状態の転送を隠すレイヤーごとのパイプライン スケジューラー、および CPU-GPU 制御オーバーヘッドを削減するために制御を下げる永続カーネルを組み合わせています。効率的な GPU メモリ プーリングにより、CrossPool はバースト性の高いロングコンテキスト リクエストをサポートし、最先端の kvcached ベースのマルチ LLM サービング システムを上回るパフォーマンスを発揮し、P99 TBT を最大 $10.4\time$ 削減します。

原文 (English)

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU memory problem: model weights are stable and model-determined, while KV-cache is transient and demand-determined. Because cold models rarely reach peak KV-cache demand at the same time, reserving worst-case KV capacity per model wastes memory; a shared KV-cache pool can instead provision aggregate active demand. However, KV-cache sharing is not sufficient when weights and KV-cache remain in a monolithic GPU memory pool. Static weights compete with dynamic KV-cache, and KV-head-limited attention under cold, low-concurrency traffic exposes only a fraction of replicated KV capacity, leading to low GPU memory utilization and weak long-context support. We present CrossPool, a serving engine for cold MoE models that separates FFN weights and KV-cache into two GPU memory pools: a weights pool that consolidates FFN weights across cold models, and a KV-cache pool that dynamically serves active requests while keeping attention local to KV-cache. CrossPool combines a KV-cache planner and virtualizer, a layer-wise pipeline scheduler that hides hidden-state transfers, and persistent kernels with control lowering to reduce CPU-GPU control overhead. With efficient GPU memory pooling, CrossPool underpins bursty long-context requests and outperforms the state-of-the-art kvcached-based multi-LLM serving system, reducing P99 TBT by up to 10.4x.

2026-06-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

\textsc{DiARC}: ポジティブサンプルとネガティブサンプルの区別は、大規模な言語モデルの ARC のような推論能力の向上に役立ちます

Abstraction and Reasoning Corpus (ARC;~\citealp{chollet2019measure}) には、限られたグリッド サンプルからのパターンの要約と出力グリッドの予測を必要とするタスクが含まれています。最近、多くの大規模な言語モデル ベースのアプローチが、言語モデルをテキスト ベースの推論タスクに変換しようと試みています。ただし、オープンソース モデルに基づく方法では一般に満足のいく結果が得られず、クローズドソース モデルに依存する方法ではコストがかかりすぎます。現在の取り組みは主にデータ拡張に焦点を当てており、より包括的な監視付き微調整のための ARC のようなデータを構築しています。この研究では、ARC のような問題を解決するには \textit{positive} サンプルの監視だけでなく、\textit{negative} サンプルを区別してモデル推論を改善する能力も必要であると主張します。この目的を達成するために、私たちは好みの調整のアイデアを利用し、モデルがそれらを区別できるように好みのペアを構築する方法である \textsc{DiARC} を提案します。具体的には、出力レベルの視覚的変換、DSL レベルのルール反転、およびタスク固有のルール編集を含む、ネガティブ サンプルを構築する 3 つの方法を提案します。得られた陰性サンプルは、観察されたデモンストレーションを変更せずに、有益なニアミス代替案を提供します。複数の ARC に似たベンチマークにわたる実験結果は、\textsc{DiARC} がベースライン モデルよりも一貫してパフォーマンスを向上させることを示しています。コードは https://github.com/szu-tera/DiARC で公開されています。

原文 (English)

DiARC: Distinguishing Positive and Negative Samples Helps Improving ARC-like Reasoning Ability of Large Language Models

The Abstraction and Reasoning Corpus (ARC) contains tasks that require summarizing patterns from limited grid samples and predicting output grids. Recently, many large language model based approaches have attempted to transform it into a text-based reasoning task. However, methods based on open-source models have generally yielded unsatisfactory results, while those relying on closed-source models are too costly. Current efforts mainly focus on data augmentation, constructing ARC-like data for more comprehensive supervised fine-tuning. In this work, we argue that solving ARC-like problems requires not only positive sample supervision but also the ability to improve model reasoning by distinguishing negative samples. To this end, we draw on the idea of preference alignment and propose DiARC, a method that constructs preference pairs to enable the model to distinguish between them. Specifically, we propose three ways to construct negative samples, including output-level visual transformations, DSL-level rule inversion, and task-specific rule editing. The resulting negative samples provide informative near-miss alternatives while keeping the observed demonstrations unchanged. Experimental results across multiple ARC-like benchmarks show that DiARC consistently improves performance over baseline models. The code is released at https://github.com/szu-tera/DiARC.

2026-06-29 10:39 JSTITmedia AI+ロボティクスハードウェア/半導体

ロボットの模倣学習を60時間→4.8時間に AWSのGPUでフィジカルAI開発を加速 ファナック

ファナックは「AWS Summit Japan 2026」の基調講演で、ロボットに動作を教える「模倣学習」の時間を、AWSのGPU活用で60時間から4.8時間へ短縮したと示した。

2026-06-29 00:00 JSTTechCrunch AIハードウェア/半導体

Why Wall Street thinks US memory maker Micron is the next Nvidia

Eager to find more public AI-related companies that may do as well as Nvidia, Wall Street investors think they've found a winner with Micro…

2026-06-27 23:00 JSTTechCrunch AIハードウェア/半導体

The fittest founder in the room got cancer. Here’s how he used AI to fight back.

When confronted with cancer, Connor Christou fed everything tied tied to his regime — blood results, scan data, wearable output, journal en…

2026-06-27 02:43 JSTTechCrunch AILLM/生成AIハードウェア/半導体

Why everyone from OpenAI to SpaceX is building their own chips (and turning up the heat on Nvidia)

Nvidia has dominated the AI chip market for years, but the era of total dependence might be ending. OpenAI just shared its plans to spice t…

2026-06-26 23:00 JSTTechCrunch AILLM/生成AIハードウェア/半導体

OpenAI’s Jalapeño chip is Big Tech’s spiciest move away from Nvidia

Nvidia has dominated the AI chip market for years, but the era of total dependence might be ending. OpenAI just shared its plans to spice t…

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

EGG: 専門家によるカーネル生成のためのエージェント フレームワーク

高性能 GPU カーネルは、大規模言語モデル (LLM) の指数関数的に増大する計算コストを削減するために不可欠ですが、その開発はドメイン専門家による手動チューニングに大きく依存しています。 LLM ベースのアプローチの最近の進歩は、カーネル生成の自動化に有望であることを示していますが、正確さと高いパフォーマンスの両方を達成するのにまだ苦労しています。この制限は主に、ドメイン固有の最適化ガイダンスの欠如によって生じ、最適化空間の効果的な探索が妨げられます。私たちは、LLM の意思決定をガイドする専門家の最適化原則を組み込んだ、カーネル生成のための専門家ガイド付きエージェント フレームワークである EGG を提案します。専門家のワークフローからインスピレーションを得て、私たちはカーネル生成を 2 つの階層段階に分解します。1) 高品質の計算構造基盤を確立するアルゴリズム構造設計。 2) ハードウェア固有のチューニング。並列マッピング、テンソル タイリング、メモリ最適化を通じてターゲットを絞った調整を実行します。この段階的な分解により、明示的な最適化目標が定義され、段階的な改良を達成するために設計空間が構築されます。この目的を達成するために、ステージを意識したマルチエージェント コラボレーション メカニズムがステージ間およびステージ内のコンテキスト管理用に設計されており、安定した最適化軌道を保証します。 KernelBench と実際のワークロードの実験では、EGG が PyTorch と比較して平均 2.13 倍の高速化を達成し、既存のエージェント ベースおよび RL ベースのアプローチを上回るパフォーマンスを示していることが示されています。

原文 (English)

EGG: An Expert-Guided Agent Framework for Kernel Generation

High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts. While recent advances in LLM-based approaches show promise for automating kernel generation, they still struggle to achieve both correctness and high performance. This limitation primarily arises from the lack of domain-specific optimization guidance, hindering effective exploration of the optimization space. We propose EGG, an Expert-Guided Agent Framework for Kernel Generation, which incorporates expert optimization principles to guide LLMs' decisions. Inspired by expert workflows, we decompose kernel generation into two hierarchical stages: 1) algorithmic structure design, which establishes a high-quality computational structure foundation; 2) hardware-specific tuning, which performs targeted adjustments through parallel mapping, tensor tiling, and memory optimization. This staged decomposition defines explicit optimization objectives, structuring the design space to achieve progressive refinement. To this end, a stage-aware multi-agent collaboration mechanism is designed for inter and intra-stage context management, ensuring stable optimization trajectories. Experiments on KernelBench and real-world workloads show that EGG achieves a 2.13x average speedup over PyTorch, outperforming existing agent-based and RL-based approaches.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

判断せずに質問する: 解釈可能な LLM 評価と自己改善のための 2 つの質問

NLP では、LLM 出力の評価が依然として大きなボトルネックとなっています。人間による評価は高価で時間がかかり、語彙メトリクスとオープンエンド生成に関する人間の判断との相関性が低く、全体的な LLM ジャッジはデバッグが難しい不透明なスコアを生成することがよくあります。私たちは、評価基準をアトミックなバイナリの質問に分解し、その結果の判定を解釈可能な多次元スコアに集約するフレームワークである BINEVAL を提案します。タスク プロンプトが与えられると、メタ プロンプトが詳細な評価質問を生成し、LLM が出力ごとに独立して質問に回答し、調整された全体スコアとともに透明な質問レベルのフィードバックを生成します。この分解により、評価が検査、​​診断が容易になり、迅速な改善に直接使用できるようになります。 SummEval、Topical-Chat、QAGS 全体で、BINEVAL は UniEval や G-Eval などの強力なベースラインと同等またはそれを上回り、特に QAGS などの事実整合性ベンチマークで優れた結果を示しています。 BINEVAL は、人間の判断との競合相関を超えて、人間のスコア分布とよりよく一致し、以前の LLM ジャッジによく見られた天井効果を回避し、境界線にある出力と明らかに欠陥のある出力をより適切に区別することにつながります。さらに、同じ質問レベルのフィードバックが反復プロンプトの最適化をサポートし、自己更新設定とクロスモデル更新設定の両方で IFBench での要約に関する評価者のプロンプトと生成プロンプトを改善することを示します。全体として、BINEVAL は、強力な経験的パフォーマンスと実用的な診断および最適化の価値を組み合わせた、タスクに依存せず、トレーニング不要で、解釈可能な評価フレームワークを提供します。

原文 (English)

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that decomposes evaluation criteria into atomic binary questions and aggregates the resulting verdicts into interpretable, multi-dimensional scores. Given a task prompt, a meta-prompt generates fine-grained evaluation questions, and an LLM answers them independently for each output, yielding transparent question-level feedback together with calibrated overall scores. This decomposition makes evaluation easier to inspect, easier to diagnose, and directly usable for prompt improvement. Across SummEval, Topical-Chat, and QAGS, BINEVAL matches or outperforms strong baselines including UniEval and G-Eval, with especially strong results on factual consistency benchmarks such as QAGS. Beyond competitive correlation with human judgments, BINEVAL better matches human score distributions and avoids the ceiling effects common in prior LLM judges, leading to better discrimination between borderline and clearly flawed outputs. We further show that the same question-level feedback supports iterative prompt optimization, improving evaluator prompts on summarization and generation prompts on IFBench under both self-update and cross-model update settings. Overall, BINEVAL provides a task-agnostic, training-free, and interpretable evaluation framework that combines strong empirical performance with practical diagnostic and optimization value.

2026-06-26 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

CoStream: 一般化可能な複雑な操作のための単純な動作の構築

GPU を PCIe スロットに装着するなど、長期にわたる接触が多い複雑な操作タスクでは、ミリメートル単位の高精度と、新しいタスクに対するすぐに使える汎用性の両方が必要です。既存のパラダイムは両方を満たすのに苦労しています。古典的なパイプラインは、高精度の制御を実現するために脆弱なタスク固有のインターフェイスを使用しますが、新しいタスクに適応するにはコストのかかるパイプラインの再設計が必要です。一方、モノリシックなエンドツーエンドのポリシーはより優れた一般化を提供しますが、新しいデータで再トレーニングしない限り、複雑な分散外のタスクでは高精度が得られません。どちらのパラダイムも暗黙の前提を共有しています。つまり、操作機能を取得したら、それを自由に分解したり再構成したりするのではなく、厳格なパイプラインまたはモノリシック全体として展開する必要があるということです。この論文では、複雑な操作機能が単純で独立した動作の構成から自然に出現する可能性があることを示します。モノリシックなポリシーや厳格なパイプラインを展開するのではなく、基礎モデルと多様なセンシング モダリティを複数の構成可能なコア動作に統合するフレームワークを提案します。つまり、基礎モデルを介して空間制約を抽出するセマンティック動作。想像上のビデオ内のキーポイントを追跡することで軌道を予測する予測行動。そして高周波の触覚と力の補正を提供する反応的な動作。共有 $SE(3)$ インターフェイスでは、これらの出力は右乗算によって各制御ステップで単一のポーズ コマンドに構成され、準拠したコントローラーによって実行されます。私たちは、日常的な操作と精密な組み立てに及ぶ 8 つの現実世界のタスクを実証し、接触の多い組み立てとオブジェクトの転送で最も大きな効果を発揮し、実行中の手動による混乱からの堅牢な回復を示します。 {ウェブサイト:} https://costream-simple.github.io

原文 (English)

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation

Long-horizon, contact-rich complex manipulation tasks, such as seating a GPU into a PCIe slot, demand both millimeter high precision and out-of-the-box generalization to new tasks. Existing paradigms struggle to satisfy both: classical pipelines use brittle, task-specific interfaces to achieve high-precision control but require costly pipeline redesigns to adapt to new tasks, whereas monolithic end-to-end policies provide better generalization but lack high precision on complex, out-of-distribution tasks unless retrained with new data. Both paradigms share an implicit assumption: once a manipulation capability is acquired, it must be deployed as a rigid pipeline or monolithic whole, rather than being freely decomposed and recomposed. In this paper, we show that complex manipulation capabilities can emerge naturally from the composition of simple, independent behaviors. Rather than deploying a monolithic policy or a rigid pipeline, we propose \ourshort, a framework orchestrating foundation models and diverse sensing modalities into multiple composable core behaviors: a semantic behavior extracting spatial constraints via foundation models; a predictive behavior forecasting trajectories by tracking keypoints in imagined videos; and a reactive behavior providing high-frequency tactile and force corrections. On a shared $SE(3)$ interface, these outputs compose by right-multiplication into a single pose command at each control step, executed by a compliant controller. We demonstrate \ourshort on 8 real-world tasks spanning everyday manipulation and precision assembly, with the strongest gains in contact-rich assembly and object transfer, and show robust recovery from manual perturbations during execution. {Website:} https://costream-simple.github.io

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior

Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on t…

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

EGG: 専門家によるカーネル生成のためのエージェント フレームワーク

高性能 GPU カーネルは、大規模言語モデル (LLM) の指数関数的に増大する計算コストを削減するために不可欠ですが、その開発はドメイン専門家による手動チューニングに大きく依存しています。 LLM ベースのアプローチの最近の進歩は、カーネル生成の自動化に有望であることを示していますが、正確さと高いパフォーマンスの両方を達成するのにまだ苦労しています。この制限は主に、ドメイン固有の最適化ガイダンスの欠如によって生じ、最適化空間の効果的な探索が妨げられます。私たちは、LLM の意思決定をガイドする専門家の最適化原則を組み込んだ、カーネル生成のための専門家ガイド付きエージェント フレームワークである EGG を提案します。専門家のワークフローからインスピレーションを得て、私たちはカーネル生成を 2 つの階層段階に分解します。1) 高品質の計算構造基盤を確立するアルゴリズム構造設計。 2) ハードウェア固有のチューニング。並列マッピング、テンソル タイリング、メモリ最適化を通じてターゲットを絞った調整を実行します。この段階的な分解により、明示的な最適化目標が定義され、段階的な改良を達成するために設計空間が構築されます。この目的を達成するために、ステージを意識したマルチエージェント コラボレーション メカニズムがステージ間およびステージ内のコンテキスト管理用に設計されており、安定した最適化軌道を保証します。 KernelBench と実際のワークロードの実験では、EGG が PyTorch と比較して平均 2.13 倍の高速化を達成し、既存のエージェント ベースおよび RL ベースのアプローチを上回るパフォーマンスを示していることが示されています。

原文 (English)

EGG: An Expert-Guided Agent Framework for Kernel Generation

High-performance GPU kernels are critical for reducing the exponentially growing computational costs of large language models (LLMs), but their development heavily relies on manual tuning by domain experts. While recent advances in LLM-based approaches show promise for automating kernel generation, they still struggle to achieve both correctness and high performance. This limitation primarily arises from the lack of domain-specific optimization guidance, hindering effective exploration of the optimization space. We propose EGG, an Expert-Guided Agent Framework for Kernel Generation, which incorporates expert optimization principles to guide LLMs' decisions. Inspired by expert workflows, we decompose kernel generation into two hierarchical stages: 1) algorithmic structure design, which establishes a high-quality computational structure foundation; 2) hardware-specific tuning, which performs targeted adjustments through parallel mapping, tensor tiling, and memory optimization. This staged decomposition defines explicit optimization objectives, structuring the design space to achieve progressive refinement. To this end, a stage-aware multi-agent collaboration mechanism is designed for inter and intra-stage context management, ensuring stable optimization trajectories. Experiments on KernelBench and real-world workloads show that EGG achieves a 2.13x average speedup over PyTorch, outperforming existing agent-based and RL-based approaches.

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

判断せずに質問する: 解釈可能な LLM 評価と自己改善のための 2 つの質問

NLP では、LLM 出力の評価が依然として大きなボトルネックとなっています。人間による評価は高価で時間がかかり、語彙メトリクスとオープンエンド生成に関する人間の判断との相関性が低く、全体的な LLM ジャッジはデバッグが難しい不透明なスコアを生成することがよくあります。私たちは、評価基準をアトミックなバイナリの質問に分解し、その結果の判定を解釈可能な多次元スコアに集約するフレームワークである BINEVAL を提案します。タスク プロンプトが与えられると、メタ プロンプトが詳細な評価質問を生成し、LLM が出力ごとに独立して質問に回答し、調整された全体スコアとともに透明な質問レベルのフィードバックを生成します。この分解により、評価が検査、​​診断が容易になり、迅速な改善に直接使用できるようになります。 SummEval、Topical-Chat、QAGS 全体で、BINEVAL は UniEval や G-Eval などの強力なベースラインと同等またはそれを上回り、特に QAGS などの事実整合性ベンチマークで優れた結果を示しています。 BINEVAL は、人間の判断との競合相関を超えて、人間のスコア分布とよりよく一致し、以前の LLM ジャッジによく見られた天井効果を回避し、境界線にある出力と明らかに欠陥のある出力をより適切に区別することにつながります。さらに、同じ質問レベルのフィードバックが反復プロンプトの最適化をサポートし、自己更新設定とクロスモデル更新設定の両方で IFBench での要約に関する評価者のプロンプトと生成プロンプトを改善することを示します。全体として、BINEVAL は、強力な経験的パフォーマンスと実用的な診断および最適化の価値を組み合わせた、タスクに依存せず、トレーニング不要で、解釈可能な評価フレームワークを提供します。

原文 (English)

Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement

Evaluating LLM outputs remains a major bottleneck in NLP: human evaluation is expensive and slow, lexical metrics correlate poorly with human judgments on open-ended generation, and holistic LLM judges often produce opaque scores that are hard to debug. We propose BINEVAL, a framework that decomposes evaluation criteria into atomic binary questions and aggregates the resulting verdicts into interpretable, multi-dimensional scores. Given a task prompt, a meta-prompt generates fine-grained evaluation questions, and an LLM answers them independently for each output, yielding transparent question-level feedback together with calibrated overall scores. This decomposition makes evaluation easier to inspect, easier to diagnose, and directly usable for prompt improvement. Across SummEval, Topical-Chat, and QAGS, BINEVAL matches or outperforms strong baselines including UniEval and G-Eval, with especially strong results on factual consistency benchmarks such as QAGS. Beyond competitive correlation with human judgments, BINEVAL better matches human score distributions and avoids the ceiling effects common in prior LLM judges, leading to better discrimination between borderline and clearly flawed outputs. We further show that the same question-level feedback supports iterative prompt optimization, improving evaluator prompts on summarization and generation prompts on IFBench under both self-update and cross-model update settings. Overall, BINEVAL provides a task-agnostic, training-free, and interpretable evaluation framework that combines strong empirical performance with practical diagnostic and optimization value.

2026-06-26 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

CoStream: 一般化可能な複雑な操作のための単純な動作の構築

GPU を PCIe スロットに装着するなど、長期にわたる接触が多い複雑な操作タスクでは、ミリメートル単位の高精度と、新しいタスクに対するすぐに使える汎用性の両方が必要です。既存のパラダイムは両方を満たすのに苦労しています。古典的なパイプラインは、高精度の制御を実現するために脆弱なタスク固有のインターフェイスを使用しますが、新しいタスクに適応するにはコストのかかるパイプラインの再設計が必要です。一方、モノリシックなエンドツーエンドのポリシーはより優れた一般化を提供しますが、新しいデータで再トレーニングしない限り、複雑な分散外のタスクでは高精度が得られません。どちらのパラダイムも暗黙の前提を共有しています。つまり、操作機能を取得したら、それを自由に分解したり再構成したりするのではなく、厳格なパイプラインまたはモノリシック全体として展開する必要があるということです。この論文では、複雑な操作機能が単純で独立した動作の構成から自然に出現する可能性があることを示します。モノリシックなポリシーや厳格なパイプラインを展開するのではなく、基礎モデルと多様なセンシング モダリティを複数の構成可能なコア動作に統合するフレームワークを提案します。つまり、基礎モデルを介して空間制約を抽出するセマンティック動作。想像上のビデオ内のキーポイントを追跡することで軌道を予測する予測行動。そして高周波の触覚と力の補正を提供する反応的な動作。共有 $SE(3)$ インターフェイスでは、これらの出力は右乗算によって各制御ステップで単一のポーズ コマンドに構成され、準拠したコントローラーによって実行されます。私たちは、日常的な操作と精密な組み立てに及ぶ 8 つの現実世界のタスクを実証し、接触の多い組み立てとオブジェクトの転送で最も大きな効果を発揮し、実行中の手動による混乱からの堅牢な回復を示します。 {ウェブサイト:} https://costream-simple.github.io

原文 (English)

CoStream: Composing Simple Behaviors for Generalizable Complex Manipulation

Long-horizon, contact-rich complex manipulation tasks, such as seating a GPU into a PCIe slot, demand both millimeter high precision and out-of-the-box generalization to new tasks. Existing paradigms struggle to satisfy both: classical pipelines use brittle, task-specific interfaces to achieve high-precision control but require costly pipeline redesigns to adapt to new tasks, whereas monolithic end-to-end policies provide better generalization but lack high precision on complex, out-of-distribution tasks unless retrained with new data. Both paradigms share an implicit assumption: once a manipulation capability is acquired, it must be deployed as a rigid pipeline or monolithic whole, rather than being freely decomposed and recomposed. In this paper, we show that complex manipulation capabilities can emerge naturally from the composition of simple, independent behaviors. Rather than deploying a monolithic policy or a rigid pipeline, we propose \ourshort, a framework orchestrating foundation models and diverse sensing modalities into multiple composable core behaviors: a semantic behavior extracting spatial constraints via foundation models; a predictive behavior forecasting trajectories by tracking keypoints in imagined videos; and a reactive behavior providing high-frequency tactile and force corrections. On a shared $SE(3)$ interface, these outputs compose by right-multiplication into a single pose command at each control step, executed by a compliant controller. We demonstrate \ourshort on 8 real-world tasks spanning everyday manipulation and precision assembly, with the strongest gains in contact-rich assembly and object transfer, and show robust recovery from manual perturbations during execution. {Website:} https://costream-simple.github.io

2026-06-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Post-Training Recipe, More Than Model Family, Shapes Multi-Agent LLM Conversational Behavior

Multi-LLM systems use multiple language models to deliberate, judge each other's outputs, or coordinate as agents. Their value depends on t…

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

T2D ベンチ: 多層臨床ライフスタイル ナレッジ グラフを使用した 2 型糖尿病の LLM 出力の証拠ゲート型評価

大規模言語モデル (LLM) は、2 型糖尿病に対する臨床的に流暢な推奨事項を生成できますが、ガイドラインの制約を満たしたり、ライフスタイルに関連した血糖の主張を明確に正当化したりすることはできません。我々は、LLM 出力が明示的でグラフチェック可能な証拠要件を満たしているかどうかをテストするための、再現可能なベンチマークおよび証拠ゲート型評価フレームワークである T2D-Bench を紹介します。 T2D-Bench は、生体医学スパイン (UMLS、DrugBank、SIDER)、計算可能な ADA 治療標準ルール、血糖検査室効果への機構的なブリッジを介して接続されたライフスタイル知識を組み合わせた、多層の臨床ライフスタイル ナレッジ グラフに基づいて構築されています。診断、投薬の安全性、敵対的なライフスタイルの衝突にわたる 100 の構造化されたビネット全体で、ベースライン出力は、GPT-4o-mini のケースの 35%、GPT-4o のケースの 33% で、ベンチマークで定義されたエビデンスパス チェックに失敗しました。証拠ゲートはサポートされていない省略を検出し、制約付きリビジョンを使用して、出力をベンチマークで定義された証拠要件に検証者レベルで準拠させます。これらの結果は、糖尿病に焦点を当てた LLM 出力において、計算可能な証拠の制約により、裏付けのない臨床上の省略が明示的、測定可能、修正可能になる可能性があることを示しています。

原文 (English)

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims. We present T2D-Bench, a reproducible benchmark and evidence-gated evaluation framework for testing whether LLM outputs satisfy explicit, graph-checkable evidence requirements. T2D-Bench is built on a multi-layer clinical-lifestyle knowledge graph that combines a biomedical spine (UMLS, DrugBank, SIDER), computable ADA Standards of Care rules, and lifestyle knowledge connected through a mechanistic bridge to glycemic laboratory effects. Across 100 structured vignettes spanning diagnosis, medication safety, and adversarial lifestyle conflicts, baseline outputs failed benchmark-defined evidence-path checks in 35% of cases for GPT-4o-mini and 33% for GPT-4o. The evidence gate detects unsupported omissions and uses constrained revision to bring outputs into verifier-level compliance with benchmark-defined evidence requirements. These results show that computable evidence constraints can make unsupported clinical omissions explicit, measurable, and correctable in diabetes-focused LLM outputs.

2026-06-25 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

採点者の採点: エージェントによるデータ分析システムの評価から得た教訓

エージェント的データ分析システムは、コード、数値結果、口頭診断などの豊富な出力を生成します。このため、シングルターン LLM 応答よりも評価が難しくなります。したがって、エージェントの出力と、採点アーティファクトからの真実の回答との間の真の不一致を区別する必要があります。私たちは、DSGym の 153 の数値 QRData タスクにマルチエージェント データ分析システムである LAMBDA を適用することで、自動採点者がそのようなシステムをどのように確実に評価するか、またどのような戦略が採点の品質を向上させるかを調査します。私たちは、厳格な正規表現マッチング、LLM ベースの寛大なグレーディング、スニペットベースの人的検査という 3 層の人的 AI グレーディング カスケードを開発および評価します。これは、非 GenAI 戦略と GenAI 戦略をさまざまな障害プロファイルと組み合わせたものです。どちらの自動グレーダーも 100% の観察精度 (誤検知 0/70) を達成しています。寛大な採点者の再現率は人間のラベルに対して 97% です。キーワードに固定された抽出パイプラインにより、厳密な採点者の再現率は、最後の数字のヒューリスティックよりも 60 パーセント ポイント高くなります。寛大なグレーダーはアーキテクチャ的にパーサーに依存しません。反復的なナッジメカニズムにより、採点の成功率が 36% から 97% に上昇し、寛容な合格率が 16% から 46% に上昇します。元の質問の再挿入ありとなしのナッジを比較すると、再挿入にはメリットがないことがわかり、ナッジが回答テンプレートの手がかりであることが確認されました。さらに、このケース スタディでは、変数タイプがパイプラインのダイナミクスの評価と観察された結果の評価に最も一貫して関連付けられているタスク メタデータ フィールドであることがわかります。

原文 (English)

Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM responses. It is therefore necessary to distinguish genuine disagreement between an agent's output and a ground-truth answer from grading artifacts. We investigate how reliably automated graders assess such a system and what strategies improve grading quality by applying LAMBDA, a multi-agent data-analysis system, on 153 numerical QRData tasks from DSGym. We develop and evaluate a three-layer human-AI grading cascade: strict regex matching, LLM-based lenient grading, and snippet-based human inspection, which combines non-GenAI and GenAI strategies with different failure profiles. Both automated graders achieve 100% observed precision (0/70 false positives). The lenient grader's recall is 97% against human labels. A keyword-anchored extraction pipeline raises the strict grader's recall by 60 percentage points over a last-number heuristic; the lenient grader is architecturally parser-independent. An iterative nudge mechanism raises grading run success from 36% to 97% and lenient-pass rates from 16% to 46%; comparing nudging with and without original-question re-injection shows that re-injection offers no benefit, confirming the nudge as an answer template cue. We further observe in this case study that variable type is the task metadata field most consistently associated with grading pipeline dynamics and observed outcome grades.

2026-06-25 13:00 JSTarXiv cs.AIハードウェア/半導体

必要なのは FP8 だけです (パート 2): Tensor-core Garner 再定式化と Kulisch エスケープ ルートによる効率的な Ozaki-Bailey スタイル FFT

NVIDIA の Blackwell Ultra (B300) は、FP64 ベクトル スループットを GPU あたり約 1.3 TFLOPS に削減します。これは、B200 の約 30 分の 1 であり、帯域幅が制限された FP64 ワークロードがメモリ制限にとどまるレベルをはるかに下回ります。 Ozaki Scheme II フレームワークは、仮数スライスされた中国剰余再構成を使用して FP8 テンソル コアを介して密行列乗算をルーティングすることにより、FP64 と同等のスループットを回復します。関連資料のパート (1) では、高密度 GEMM、バッチ GEMV、ステンシル、および SpMV について説明します。この論文では、5 番目の標準プリミティブである 3-D FFT を追加します。 FP8 テンソル コア上の両方の 1-D FFT GEMM を使用した Bailey 6 ステップ分解を介してエミュレートされた 3-D FFT である Ozaki-Bailey FFT を紹介します。 Bailey の小さな内因数 k ~ sqrt(N) (N=1024 の場合 k=32) は、カーネルを k << r^2 の状態に置き、3 番目の TME パラメーター ガンマ (再構成待ち時間) が償却ではなく結合します。 Garner の再構成は、フェーズ A (FP8/INT8 テンソル コアの内積、B300 の 1024^3 で約 1 ミリ秒) とフェーズ B (出力ごとの削減) に分割されます。 Kulisch 固定小数点完全演算は、完全に INT32 SIMT パイプ上で実行しながら完全な FP64 精度を維持するフェーズ B 再定式化であると認識します。閉じた形式の帯域幅パリティの下限を導出します。ネイティブ FP64 の下限は 1.56*B_HBM (8 TB/s で 12.5 TF) です。B300 の 1.3 TF は約 10 倍低く、Rubin の 33 TF は 4% 以内です。 Kulisch 避難ルートには、INT32 サブフロア 8.25*B_HBM と FP8 フロア 170*B_HBM が必要です。 B300 はその両方を満たします。フル FP64 での 1024^3 の予測は約 18 ミリ秒で、実質的には 12.9 ミリ秒のメモリ ルーフです。 GPU がネイティブ フロアまたは両方の Kulisch フロアを満たす場合、GPU はメモリルーフ FFT パリティを満たします。この予測が実際に当てはまれば、B300 はソフトウェアのみでフル FP64 FFT を実行できるようになり、libKulisch ライブラリとベンチマーク キャンペーンの動機付けとなります。

原文 (English)

FP8 is All You Need (Part 2): Efficient Ozaki-Bailey Style FFT Through Tensor-core Garner Reformulation and Kulisch Escape Route

NVIDIA's Blackwell Ultra (B300) cuts FP64 vector throughput to ~1.3 TFLOPS per GPU, roughly 30x below B200 and well below the level at which bandwidth-limited FP64 workloads stay memory-bound. The Ozaki Scheme II framework recovers FP64-equivalent throughput by routing dense matrix multiply through FP8 tensor cores with a mantissa-sliced Chinese-remainder reconstruction. A companion Part (1) paper covers dense GEMM, batched GEMV, stencils, and SpMV; this paper adds the fifth canonical primitive, the 3-D FFT. We present Ozaki-Bailey FFT, an emulated 3-D FFT via the Bailey six-step decomposition with both 1-D FFT GEMMs on FP8 tensor cores. Bailey's small inner factor k ~ sqrt(N) (k=32 for N=1024) puts the kernel in the regime k << r^2, where the third TME parameter gamma (reconstruction latency) binds rather than amortising. Garner reconstruction splits into Phase A (inner products on FP8/INT8 tensor cores, ~1 ms for 1024^3 on B300) and Phase B (per-output reduction). We identify Kulisch fixed-point complete arithmetic as a Phase B reformulation that keeps full FP64 accuracy while running entirely on the INT32 SIMT pipe. We derive closed-form bandwidth-parity floors. The native FP64 floor is 1.56*B_HBM (12.5 TF at 8 TB/s): B300's 1.3 TF sits ~10x below, Rubin's 33 TF within 4%. The Kulisch escape route needs an INT32 sub-floor 8.25*B_HBM and an FP8 floor 170*B_HBM; B300 meets both. The projection is ~18 ms for 1024^3 at full FP64, essentially the 12.9 ms memory roof. A GPU meets memory-roof FFT parity if it satisfies either the native floor or both Kulisch floors. If the projection holds in practice, B300 becomes viable for full-FP64 FFT through software alone, motivating a libKulisch library and benchmark campaign.

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CAVEWOMAN: 言語入力および出力圧縮下で大規模言語モデルがどのように動作するか

「短く話してください。文法は省略してください。トークンを保存してください。」この穴居人のスタイルは、推論コストを削減する方法として広く推奨されていますが、実際に何かを節約できるかどうかは、どのチャネル (ユーザーのプロンプトまたはモデルの応答) が圧縮されているかによって異なります。我々は、タスクの精度、実現アイテムごとのコスト、およびモデルの制約のない参照に対する参照テキストの一致に関して世代ごとにスコアを付ける 2 チャネル評価プロトコルである Cave Woman を紹介します。両方のチャネルが同じ項目で測定され、5 つのデータセットの 8 つのモデルを 5 つの削減レベルで評価します。出力圧縮により、ほとんどの API モデル (モデルあたり 1.4 ~ 2.4 倍、最良の場合は最大 3 倍) とパブリック層価格設定の 4 つのオープンウェイト モデルすべてで実現コストが削減されます。入力圧縮には逆の効果があり、厳密な損失です。モデルは精度が低下しても、より長い応答で補正するため、正味コストは低下するのではなく増加します (5 つのベンチマーク平均で約 1.15 倍、最悪のデータセットで最大 1.8 倍、強力な圧縮下で 2.7 倍)。同じ設定の下では、表面テキストは制約のない参照から分岐します。非推論モデルでは、すべての世代の約半分が正しいにもかかわらず、それらの表面テキストはモデル独自の制約のないベースライン生成を必要としません。相違は、長さ制御された再スコアリング、複数比較の修正、および補完的な意味論的尺度の下での複製を経ても存続します。コードとデータは https://github.com/danielle34/cavewoman で入手できます。

原文 (English)

CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

"Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything depends on which channel (the user's prompt or the model's response) is being compressed. We present Cavewoman, a two-channel evaluation protocol that scores every generation on task accuracy, realized per-item cost, and reference-text agreement against the model's unconstrained reference. We evaluate eight models on five datasets at five reduction levels, with both channels measured on the same items. Output compression cuts realized cost on most API models (1.4-2.4x per model, up to 3x in the best case) and on all four open-weight models under public-tier pricing. Input compression has the opposite effect, a strict lose-lose: it raises net cost rather than lowering it (~1.15x on the five-benchmark mean, up to 1.8x on the worst dataset and 2.7x under stronger compression), because models compensate with longer responses even as accuracy collapses. Under the same setting, surface text diverges from the unconstrained reference: on the non-reasoning models, roughly half of all generations are correct yet their surface text no longer entails the model's own unconstrained baseline generation. The divergence survives length-controlled re-scoring, multiple-comparisons correction, and replication under complementary semantic measures. Code and data are available at https://github.com/danielle34/cavewoman.

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CrossPool: KV キャッシュと重み分解によるコールド MoE モデルの効率的なマルチ LLM サービス

新興の LLM サービスは、多くの疎な MoE モデルをホストすることが増えていますが、ほとんどのモデルは疎なリクエストを受け取り、コールドのままです。これにより、GPU メモリの問題が発生します。モデルの重みは安定していてモデルによって決定されますが、KV キャッシュは一時的で需要によって決定されます。コールド モデルが同時にピーク KV キャッシュ要求に達することはほとんどないため、モデルごとに最悪の場合の KV 容量を確保するとメモリが無駄になります。代わりに、共有 KV キャッシュ プールを使用して、アクティブな需要を集約してプロビジョニングできます。ただし、重みと KV キャッシュがモノリシック GPU メモリ プールに残っている場合、KV キャッシュの共有は十分ではありません。静的重みは動的 KV キャッシュと競合し、コールドで同時実行トラフィックが少ない場合に KV ヘッドに制限された注意は、レプリケートされた KV 容量の一部のみを公開するため、GPU メモリ使用率が低くなり、ロングコンテキストのサポートが弱くなります。 CrossPool は、FFN 重みと KV キャッシュを 2 つの GPU メモリ プールに分離するコールド MoE モデル用のサービング エンジンです。1 つはコールド モデル全体で FFN 重みを統合する重みプール、もう 1 つは KV キャッシュにローカルな注意を保ちながらアクティブなリクエストを動的に処理する KV キャッシュ プールです。 CrossPool は、KV キャッシュ プランナーとバーチャライザー、隠し状態の転送を隠すレイヤーごとのパイプライン スケジューラー、および CPU-GPU 制御オーバーヘッドを削減するために制御を下げる永続カーネルを組み合わせています。効率的な GPU メモリ プーリングにより、CrossPool はバースト性の高いロングコンテキスト リクエストをサポートし、最先端の kvcached ベースのマルチ LLM サービング システムを上回るパフォーマンスを発揮し、P99 TBT を最大 $10.4\time$ 削減します。

原文 (English)

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU memory problem: model weights are stable and model-determined, while KV-cache is transient and demand-determined. Because cold models rarely reach peak KV-cache demand at the same time, reserving worst-case KV capacity per model wastes memory; a shared KV-cache pool can instead provision aggregate active demand. However, KV-cache sharing is not sufficient when weights and KV-cache remain in a monolithic GPU memory pool. Static weights compete with dynamic KV-cache, and KV-head-limited attention under cold, low-concurrency traffic exposes only a fraction of replicated KV capacity, leading to low GPU memory utilization and weak long-context support. We present CrossPool, a serving engine for cold MoE models that separates FFN weights and KV-cache into two GPU memory pools: a weights pool that consolidates FFN weights across cold models, and a KV-cache pool that dynamically serves active requests while keeping attention local to KV-cache. CrossPool combines a KV-cache planner and virtualizer, a layer-wise pipeline scheduler that hides hidden-state transfers, and persistent kernels with control lowering to reduce CPU-GPU control overhead. With efficient GPU memory pooling, CrossPool underpins bursty long-context requests and outperforms the state-of-the-art kvcached-based multi-LLM serving system, reducing P99 TBT by up to $10.4\times$.

2026-06-25 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達研究/論文

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of aut…

2026-06-25 09:08 JSTTechCrunch AIハードウェア/半導体

Europe is pushing back on Washington’s chip war

As ASML CEO Christophe Fouquet told TechCrunch in May, what China can currently buy are older-generation deep ultraviolet tools — gear firs…

2026-06-25 07:41 JSTTechCrunch AIハードウェア/半導体

Cerebras stock plunges after earnings as CEO says margin outlook was misunderstood

In its first earnings report since going public, the AI chipmaker forecast a narrower gross margin in its core business, scaring investors.

2026-06-25 06:30 JSTTechCrunch AIハードウェア/半導体

The memory chip crunch is paying off for this US company

Revenue quadrupled to $41.45 billion compared with the same period a year ago. The company's profit, meanwhile, rose from $1.88 billion to…

2026-06-24 23:54 JSTTechCrunch AILLM/生成AIハードウェア/半導体

OpenAI unveils its first custom chip, built by Broadcom

Named Jalapeño, the new processor was designed specifically for the unique needs of OpenAI's inference systems.

2026-06-24 19:54 JSTITmedia AI+LLM/生成AIハードウェア/半導体

「Transformerの最大475倍」 富士通、GPUを効率的に使うLLMアーキテクチャ「PHOTON」開発

富士通が、大規模言語モデル(LLM)を少ないGPUで動かせる新アーキテクチャ「PHOTON」(フォトン)を開発した。GPU当たりの処理性能(スループット)が、現在のLLMで主流のアーキテクチャ「Transformer」の最大475倍に達するという。LLMの運用に必要なGPUを抑…

2026-06-24 15:40 JSTITmedia AI+ハードウェア/半導体

【解説】キオクシアなぜ急成長? 半導体メモリって何? AIブームを見通すための基礎知識

注目を集める半導体メモリ大手のキオクシア。同社はなぜAI需要を取り込めたのか。いま押さえたい基礎知識を解説する。

2026-06-24 15:00 JSTOpenAILLM/生成AIハードウェア/半導体

OpenAI and Broadcom unveil LLM-optimized inference chip

OpenAI and Broadcom introduce Jalapeño, a custom AI chip built for LLM inference to improve performance, efficiency, and scale across AI sy…

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

T2D ベンチ: 多層臨床ライフスタイル ナレッジ グラフを使用した 2 型糖尿病の LLM 出力の証拠ゲート型評価

大規模言語モデル (LLM) は、2 型糖尿病に対する臨床的に流暢な推奨事項を生成できますが、ガイドラインの制約を満たしたり、ライフスタイルに関連した血糖の主張を明確に正当化したりすることはできません。我々は、LLM 出力が明示的でグラフチェック可能な証拠要件を満たしているかどうかをテストするための、再現可能なベンチマークおよび証拠ゲート型評価フレームワークである T2D-Bench を紹介します。 T2D-Bench は、生体医学スパイン (UMLS、DrugBank、SIDER)、計算可能な ADA 治療標準ルール、血糖検査室効果への機構的なブリッジを介して接続されたライフスタイル知識を組み合わせた、多層の臨床ライフスタイル ナレッジ グラフに基づいて構築されています。診断、投薬の安全性、敵対的なライフスタイルの衝突にわたる 100 の構造化されたビネット全体で、ベースライン出力は、GPT-4o-mini のケースの 35%、GPT-4o のケースの 33% で、ベンチマークで定義されたエビデンスパス チェックに失敗しました。証拠ゲートはサポートされていない省略を検出し、制約付きリビジョンを使用して、出力をベンチマークで定義された証拠要件に検証者レベルで準拠させます。これらの結果は、糖尿病に焦点を当てた LLM 出力において、計算可能な証拠の制約により、裏付けのない臨床上の省略が明示的、測定可能、修正可能になる可能性があることを示しています。

原文 (English)

T2D-Bench: Evidence-Gated Evaluation of LLM Outputs for Type 2 Diabetes Using a Multi-Layer Clinical-Lifestyle Knowledge Graph

Large language models (LLMs) can produce clinically fluent recommendations for type 2 diabetes while failing to satisfy guideline constraints or explicitly justify lifestyle-related glycemic claims. We present T2D-Bench, a reproducible benchmark and evidence-gated evaluation framework for testing whether LLM outputs satisfy explicit, graph-checkable evidence requirements. T2D-Bench is built on a multi-layer clinical-lifestyle knowledge graph that combines a biomedical spine (UMLS, DrugBank, SIDER), computable ADA Standards of Care rules, and lifestyle knowledge connected through a mechanistic bridge to glycemic laboratory effects. Across 100 structured vignettes spanning diagnosis, medication safety, and adversarial lifestyle conflicts, baseline outputs failed benchmark-defined evidence-path checks in 35% of cases for GPT-4o-mini and 33% for GPT-4o. The evidence gate detects unsupported omissions and uses constrained revision to bring outputs into verifier-level compliance with benchmark-defined evidence requirements. These results show that computable evidence constraints can make unsupported clinical omissions explicit, measurable, and correctable in diabetes-focused LLM outputs.

2026-06-24 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

採点者の採点: エージェントによるデータ分析システムの評価から得た教訓

エージェント的データ分析システムは、コード、数値結果、口頭診断などの豊富な出力を生成します。このため、シングルターン LLM 応答よりも評価が難しくなります。したがって、エージェントの出力と、採点アーティファクトからの真実の回答との間の真の不一致を区別する必要があります。私たちは、DSGym の 153 の数値 QRData タスクにマルチエージェント データ分析システムである LAMBDA を適用することで、自動採点者がそのようなシステムをどのように確実に評価するか、またどのような戦略が採点の品質を向上させるかを調査します。私たちは、厳格な正規表現マッチング、LLM ベースの寛大なグレーディング、スニペットベースの人的検査という 3 層の人的 AI グレーディング カスケードを開発および評価します。これは、非 GenAI 戦略と GenAI 戦略をさまざまな障害プロファイルと組み合わせたものです。どちらの自動グレーダーも 100% の観察精度 (誤検知 0/70) を達成しています。寛大な採点者の再現率は人間のラベルに対して 97% です。キーワードに固定された抽出パイプラインにより、厳密な採点者の再現率は、最後の数字のヒューリスティックよりも 60 パーセント ポイント高くなります。寛大なグレーダーはアーキテクチャ的にパーサーに依存しません。反復的なナッジメカニズムにより、採点の成功率が 36% から 97% に上昇し、寛容な合格率が 16% から 46% に上昇します。元の質問の再挿入ありとなしのナッジを比較すると、再挿入にはメリットがないことがわかり、ナッジが回答テンプレートの手がかりであることが確認されました。さらに、このケース スタディでは、変数タイプがパイプラインのダイナミクスの評価と観察された結果の評価に最も一貫して関連付けられているタスク メタデータ フィールドであることがわかります。

原文 (English)

Grading the Grader: Lessons from Evaluating an Agentic Data Analysis System

Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM responses. It is therefore necessary to distinguish genuine disagreement between an agent's output and a ground-truth answer from grading artifacts. We investigate how reliably automated graders assess such a system and what strategies improve grading quality by applying LAMBDA, a multi-agent data-analysis system, on 153 numerical QRData tasks from DSGym. We develop and evaluate a three-layer human-AI grading cascade: strict regex matching, LLM-based lenient grading, and snippet-based human inspection, which combines non-GenAI and GenAI strategies with different failure profiles. Both automated graders achieve 100% observed precision (0/70 false positives). The lenient grader's recall is 97% against human labels. A keyword-anchored extraction pipeline raises the strict grader's recall by 60 percentage points over a last-number heuristic; the lenient grader is architecturally parser-independent. An iterative nudge mechanism raises grading run success from 36% to 97% and lenient-pass rates from 16% to 46%; comparing nudging with and without original-question re-injection shows that re-injection offers no benefit, confirming the nudge as an answer template cue. We further observe in this case study that variable type is the task metadata field most consistently associated with grading pipeline dynamics and observed outcome grades.

2026-06-24 13:00 JSTarXiv cs.AIハードウェア/半導体

必要なのは FP8 だけです (パート 2): Tensor-core Garner 再定式化と Kulisch エスケープ ルートによる効率的な Ozaki-Bailey スタイル FFT

NVIDIA の Blackwell Ultra (B300) は、FP64 ベクトル スループットを GPU あたり約 1.3 TFLOPS に削減します。これは、B200 の約 30 分の 1 であり、帯域幅が制限された FP64 ワークロードがメモリ制限にとどまるレベルをはるかに下回ります。 Ozaki Scheme II フレームワークは、仮数スライスされた中国剰余再構成を使用して FP8 テンソル コアを介して密行列乗算をルーティングすることにより、FP64 と同等のスループットを回復します。関連資料のパート (1) では、高密度 GEMM、バッチ GEMV、ステンシル、および SpMV について説明します。この論文では、5 番目の標準プリミティブである 3-D FFT を追加します。 FP8 テンソル コア上の両方の 1-D FFT GEMM を使用した Bailey 6 ステップ分解を介してエミュレートされた 3-D FFT である Ozaki-Bailey FFT を紹介します。 Bailey の小さな内因数 k ~ sqrt(N) (N=1024 の場合 k=32) は、カーネルを k << r^2 の状態に置き、3 番目の TME パラメーター ガンマ (再構成待ち時間) が償却ではなく結合します。 Garner の再構成は、フェーズ A (FP8/INT8 テンソル コアの内積、B300 の 1024^3 で約 1 ミリ秒) とフェーズ B (出力ごとの削減) に分割されます。 Kulisch 固定小数点完全演算は、完全に INT32 SIMT パイプ上で実行しながら完全な FP64 精度を維持するフェーズ B 再定式化であると認識します。閉じた形式の帯域幅パリティの下限を導出します。ネイティブ FP64 の下限は 1.56*B_HBM (8 TB/s で 12.5 TF) です。B300 の 1.3 TF は約 10 倍低く、Rubin の 33 TF は 4% 以内です。 Kulisch 避難ルートには、INT32 サブフロア 8.25*B_HBM と FP8 フロア 170*B_HBM が必要です。 B300 はその両方を満たします。フル FP64 での 1024^3 の予測は約 18 ミリ秒で、実質的には 12.9 ミリ秒のメモリ ルーフです。 GPU がネイティブ フロアまたは両方の Kulisch フロアを満たす場合、GPU はメモリルーフ FFT パリティを満たします。この予測が実際に当てはまれば、B300 はソフトウェアのみでフル FP64 FFT を実行できるようになり、libKulisch ライブラリとベンチマーク キャンペーンの動機付けとなります。

原文 (English)

FP8 is All You Need (Part 2): Efficient Ozaki-Bailey Style FFT Through Tensor-core Garner Reformulation and Kulisch Escape Route

NVIDIA's Blackwell Ultra (B300) cuts FP64 vector throughput to ~1.3 TFLOPS per GPU, roughly 30x below B200 and well below the level at which bandwidth-limited FP64 workloads stay memory-bound. The Ozaki Scheme II framework recovers FP64-equivalent throughput by routing dense matrix multiply through FP8 tensor cores with a mantissa-sliced Chinese-remainder reconstruction. A companion Part (1) paper covers dense GEMM, batched GEMV, stencils, and SpMV; this paper adds the fifth canonical primitive, the 3-D FFT. We present Ozaki-Bailey FFT, an emulated 3-D FFT via the Bailey six-step decomposition with both 1-D FFT GEMMs on FP8 tensor cores. Bailey's small inner factor k ~ sqrt(N) (k=32 for N=1024) puts the kernel in the regime k << r^2, where the third TME parameter gamma (reconstruction latency) binds rather than amortising. Garner reconstruction splits into Phase A (inner products on FP8/INT8 tensor cores, ~1 ms for 1024^3 on B300) and Phase B (per-output reduction). We identify Kulisch fixed-point complete arithmetic as a Phase B reformulation that keeps full FP64 accuracy while running entirely on the INT32 SIMT pipe. We derive closed-form bandwidth-parity floors. The native FP64 floor is 1.56*B_HBM (12.5 TF at 8 TB/s): B300's 1.3 TF sits ~10x below, Rubin's 33 TF within 4%. The Kulisch escape route needs an INT32 sub-floor 8.25*B_HBM and an FP8 floor 170*B_HBM; B300 meets both. The projection is ~18 ms for 1024^3 at full FP64, essentially the 12.9 ms memory roof. A GPU meets memory-roof FFT parity if it satisfies either the native floor or both Kulisch floors. If the projection holds in practice, B300 becomes viable for full-FP64 FFT through software alone, motivating a libKulisch library and benchmark campaign.

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CAVEWOMAN: 言語入力および出力圧縮下で大規模言語モデルがどのように動作するか

「短く話してください。文法は省略してください。トークンを保存してください。」この穴居人のスタイルは、推論コストを削減する方法として広く推奨されていますが、実際に何かを節約できるかどうかは、どのチャネル (ユーザーのプロンプトまたはモデルの応答) が圧縮されているかによって異なります。我々は、タスクの精度、実現アイテムごとのコスト、およびモデルの制約のない参照に対する参照テキストの一致に関して世代ごとにスコアを付ける 2 チャネル評価プロトコルである Cave Woman を紹介します。両方のチャネルが同じ項目で測定され、5 つのデータセットの 8 つのモデルを 5 つの削減レベルで評価します。出力圧縮により、ほとんどの API モデル (モデルあたり 1.4 ~ 2.4 倍、最良の場合は最大 3 倍) とパブリック層価格設定の 4 つのオープンウェイト モデルすべてで実現コストが削減されます。入力圧縮には逆の効果があり、厳密な損失です。モデルは精度が低下しても、より長い応答で補正するため、正味コストは低下するのではなく増加します (5 つのベンチマーク平均で約 1.15 倍、最悪のデータセットで最大 1.8 倍、強力な圧縮下で 2.7 倍)。同じ設定の下では、表面テキストは制約のない参照から分岐します。非推論モデルでは、すべての世代の約半分が正しいにもかかわらず、それらの表面テキストはモデル独自の制約のないベースライン生成を必要としません。相違は、長さ制御された再スコアリング、複数比較の修正、および補完的な意味論的尺度の下での複製を経ても存続します。コードとデータは https://github.com/danielle34/cavewoman で入手できます。

原文 (English)

CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression

"Talk short. Drop grammar. Save token." This caveman style is widely promoted as a way to cut inference cost, but whether it actually saves anything depends on which channel (the user's prompt or the model's response) is being compressed. We present Cavewoman, a two-channel evaluation protocol that scores every generation on task accuracy, realized per-item cost, and reference-text agreement against the model's unconstrained reference. We evaluate eight models on five datasets at five reduction levels, with both channels measured on the same items. Output compression cuts realized cost on most API models (1.4-2.4x per model, up to 3x in the best case) and on all four open-weight models under public-tier pricing. Input compression has the opposite effect, a strict lose-lose: it raises net cost rather than lowering it (~1.15x on the five-benchmark mean, up to 1.8x on the worst dataset and 2.7x under stronger compression), because models compensate with longer responses even as accuracy collapses. Under the same setting, surface text diverges from the unconstrained reference: on the non-reasoning models, roughly half of all generations are correct yet their surface text no longer entails the model's own unconstrained baseline generation. The divergence survives length-controlled re-scoring, multiple-comparisons correction, and replication under complementary semantic measures. Code and data are available at https://github.com/danielle34/cavewoman.

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CrossPool: KV キャッシュと重み分解によるコールド MoE モデルの効率的なマルチ LLM サービス

新興の LLM サービスは、多くの疎な MoE モデルをホストすることが増えていますが、ほとんどのモデルは疎なリクエストを受け取り、コールドのままです。これにより、GPU メモリの問題が発生します。モデルの重みは安定していてモデルによって決定されますが、KV キャッシュは一時的で需要によって決定されます。コールド モデルが同時にピーク KV キャッシュ要求に達することはほとんどないため、モデルごとに最悪の場合の KV 容量を確保するとメモリが無駄になります。代わりに、共有 KV キャッシュ プールを使用して、アクティブな需要を集約してプロビジョニングできます。ただし、重みと KV キャッシュがモノリシック GPU メモリ プールに残っている場合、KV キャッシュの共有は十分ではありません。静的重みは動的 KV キャッシュと競合し、コールドで同時実行トラフィックが少ない場合に KV ヘッドに制限された注意は、レプリケートされた KV 容量の一部のみを公開するため、GPU メモリ使用率が低くなり、ロングコンテキストのサポートが弱くなります。 CrossPool は、FFN 重みと KV キャッシュを 2 つの GPU メモリ プールに分離するコールド MoE モデル用のサービング エンジンです。1 つはコールド モデル全体で FFN 重みを統合する重みプール、もう 1 つは KV キャッシュにローカルな注意を保ちながらアクティブなリクエストを動的に処理する KV キャッシュ プールです。 CrossPool は、KV キャッシュ プランナーとバーチャライザー、隠し状態の転送を隠すレイヤーごとのパイプライン スケジューラー、および CPU-GPU 制御オーバーヘッドを削減するために制御を下げる永続カーネルを組み合わせています。効率的な GPU メモリ プーリングにより、CrossPool はバースト性の高いロングコンテキスト リクエストをサポートし、最先端の kvcached ベースのマルチ LLM サービング システムを上回るパフォーマンスを発揮し、P99 TBT を最大 $10.4\time$ 削減します。

原文 (English)

CrossPool: Efficient Multi-LLM Serving for Cold MoE Models through KV-Cache and Weight Disaggregation

Emerging LLM services increasingly host many sparse MoE models, yet most models receive sparse requests and remain cold. This creates a GPU memory problem: model weights are stable and model-determined, while KV-cache is transient and demand-determined. Because cold models rarely reach peak KV-cache demand at the same time, reserving worst-case KV capacity per model wastes memory; a shared KV-cache pool can instead provision aggregate active demand. However, KV-cache sharing is not sufficient when weights and KV-cache remain in a monolithic GPU memory pool. Static weights compete with dynamic KV-cache, and KV-head-limited attention under cold, low-concurrency traffic exposes only a fraction of replicated KV capacity, leading to low GPU memory utilization and weak long-context support. We present CrossPool, a serving engine for cold MoE models that separates FFN weights and KV-cache into two GPU memory pools: a weights pool that consolidates FFN weights across cold models, and a KV-cache pool that dynamically serves active requests while keeping attention local to KV-cache. CrossPool combines a KV-cache planner and virtualizer, a layer-wise pipeline scheduler that hides hidden-state transfers, and persistent kernels with control lowering to reduce CPU-GPU control overhead. With efficient GPU memory pooling, CrossPool underpins bursty long-context requests and outperforms the state-of-the-art kvcached-based multi-LLM serving system, reducing P99 TBT by up to $10.4\times$.

2026-06-24 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達研究/論文

Are LLM Evaluators Really Narcissists? Sanity Checking Self-Preference Evaluations

Recent research has shown that large language models (LLMs) favor their own outputs when acting as judges, undermining the integrity of aut…

2026-06-24 11:00 JSTITmedia AI+ハードウェア/半導体

AmazonはNVIDIAに挑戦状を突きつけるのか

世界最大のハイパースケーラーであるAWSは、AIアクセラレーターを大規模に販売することで、半導体市場の好機を捉えようとしているのだろうか。

2026-06-24 07:00 JSTITmedia AI+ハードウェア/半導体

「最初は壊れ過ぎてビビった」──1220億円投じたソフトバンク「AIスパコン」、それでもNVIDIAのGPUを選ぶワケ

AIブームの波に乗って時価総額世界1位に躍り出たNVIDIA。一体なぜ、AIインフラにNVIDIA製GPUが採用されるのか。その理由をAIスパコン開発者に聞いた。

2026-06-23 05:13 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

AI chipmaker Groq confirms $650M raise, re-staffs after Nvidia’s $20B not-acqui-hire deal

What does an AI company do after one of those not-acqui-hire deals? Groq raised money, is leaning into its neocloud business, and is hiring…

2026-06-23 05:08 JSTTechCrunch AIハードウェア/半導体

Nvidia wants to cut data center water use, but that’s not the same as fixing AI’s water problem

Nvidia announced a new cooling system that cuts water use inside the data center. But it does nothing to address AI's biggest water use — f…

2026-06-23 01:51 JSTTechCrunch AIハードウェア/半導体

SpaceX inks compute deal with Reflection AI, an open source AI lab

Reflection AI will pay $150 million a month beginning July 1, 2026 through 2029 for immediate access to Nvidia's latest GB300 AI chips and…

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

緊急調整

大規模言語モデル (LLM) は、自身の出力が人間の倫理と乖離していることを識別できますか?そして彼らは自己修正できるのでしょうか? LLM に、自身の推論と出力をレビューする良心ステップを与え、直接優先最適化 (DPO) を使用して調整コンポーネントでトレーニング損失を拡張し、モデルを非倫理的な出力から遠ざけます。その結果、トレーニング、微調整、敵対的プロンプト、ゼロショット学習など、幅広いアプリケーションでモデルを調整するオンライン技術が実現しました。それは、より弱いまたはより強いジャッジを必要とせず、代わりにそれ自体の凍結されたコピーに依存します。以前の研究では、緊急不整合シナリオでは、モデルの微調整からコードのハッキングに至るまで、さまざまな緊急の非倫理的な行為が示されました。代わりに、私たちは緊急調整を達成する方法を経験的に示します。つまり、単一の高レベルの内省的な質問が、同じコード ハッキング シナリオの下で倫理モデルに向けてトレーニングを導きます。

原文 (English)

Emergent Alignment

Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics? And can they self-correct? We endow an LLM with a conscience step that reviews its own reasoning and outputs, and we extend the training loss with an alignment component using Direct Preference Optimization (DPO) to steer the model away from non-ethical outputs. The result is an online technique to align models in a wide range of applications: training, fine-tuning, adversarial prompting, and zero-shot learning. It does not require a weaker or stronger judge, relying instead on a frozen copy of itself. In previous work, the Emergent Misalignment scenario showed a range of emergent unethical behaviors from fine-tuning the model to hack code. Instead, we empirically show how to achieve Emergent Alignment: a single high-level introspective question steers training toward an ethical model under the same code hacking scenario.

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

大規模言語モデルのブラックボックス不確実性推定法の系統的評価

大規模言語モデル (LLM) は幅広いタスクにわたって強力な機能を示していますが、その出力は信頼性が低いことが多く、幻覚が含まれる可能性があるため、信頼できる LLM を構築するには不確実性推定 (UE) が不可欠です。実際には、多くの主流 LLM は制限された API を介してのみアクセスでき、ロジットや隠れ状態などの内部信号は利用できないため、ブラックボックス UE が特に重要になります。しかし、LLM 用のブラックボックス UE に関する既存の研究は、方法論において断片的なままであり、統一された実証的比較が欠けています。このギャップに対処するために、ブラックボックス UE 手法の体系的なレビューを提示し、言語化ベース、サンプリング ベース、説明ベース、マルチエージェント、およびハイブリッド手法の 5 つのカテゴリに整理します。さらに、統一された評価フレームワークを構築し、4 つのモデルと 4 つのデータセット設定にわたる 24 の代表的な手法をベンチマークします。私たちの結果は、すべての設定において一貫して優勢な単一の方法はないことを示しています。それにもかかわらず、回答空間内の候補を推論して比較する方法は一般に効果的であり、複数の不確実性信号を組み合わせるハイブリッド方法は、ほとんどの条件下で良好に機能します。ベンチマーク データと統一評価フレームワークを公開することで、再現可能な比較を促進し、将来の研究をサポートすることを目指しています。また、実証結果は、LLM 向けの将来のブラック ボックス UE 手法を開発するための実践的なガイダンスを提供します。

原文 (English)

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable and may contain hallucinations, making uncertainty estimation (UE) essential for building trustworthy LLMs. In practice, many mainstream LLMs are only accessible through restricted APIs, where internal signals such as logits and hidden states are unavailable, making black-box UE especially important. However, existing work on black-box UE for LLMs remains fragmented in methodology and lacks a unified empirical comparison. To address this gap, we present a systematic review of black-box UE methods and organize them into five categories: verbalization-based, sampling-based, explanation-based, multi-agent, and hybrid methods. We further build a unified evaluation framework and benchmark 24 representative methods across 4 models and 4 dataset settings. Our results show that no single method consistently dominates across all settings. Nevertheless, methods that reason over and compare candidates in the answer space are generally effective, and hybrid methods that combine multiple uncertainty signals perform well under most conditions. By releasing the benchmark data and a unified evaluation framework, we aim to facilitate reproducible comparisons and support future research, while our empirical findings provide practical guidance for developing future black-box UE methods for LLMs.

2026-06-20 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

思考連鎖トランスフォーマーによるアルゴリズムの効率的な表現

\emph{reasoning} モデル (答えを生成する前に一連の推論または思考トークンを出力する言語モデル) の人気が高まっていることは、思考連鎖 (CoT) 変換器がチューリング マシンをシミュレートし、任意の計算を実行できることを示す理論的結果によって部分的に正当化されます。ただし、チューリング マシンは複雑性理論の分析には適していますが、アルゴリズムを議論するのには便利ではなく、直観的でもなく、効率的でもありません。通常、アルゴリズムはより高い抽象レベルで設計および分析され、ランダム アクセス メモリと $\bigO(\log n)$ ビット ワードに対する単位コスト演算を備えた \emph{Word RAM} モデルによってキャプチャされます。その結果、Word RAM アルゴリズムはチューリング マシンのアルゴリズムより大幅に効率的になる可能性があり、\emph{CoT トランスフォーマーは Word RAM アルゴリズムを効率的にシミュレートできますか?} たとえば、$n$ 個の項目を $\bigO(n \log n)$ ステップでソートしたり、ダイクストラのアルゴリズムを $\bigO(E + V \log V)$ ステップで実行したりできるでしょうか? という疑問が生じます。多対数オーバーヘッドまでは肯定的に答えます。まず、多対数幅と右端の一意のハード アテンションを持つ有限精度変換器に対してこれを確立し、次にその結果を、有限幅と対数精度を持つ 2 つのより実用的な設定に強化します。\emph{continuous} CoT (推論がトークンではなくベクトルの形式をとる場合) と、変換器層がリカレント (線形 RNN) 層の上に位置する \emph{hybrid} アーキテクチャです。 3 つのケースすべてにおいて、CoT \emph{できる} が、$n$ の多対数オーバーヘッドのみであらゆる Word RAM アルゴリズムを効率的にシミュレートできることがわかります。 Word RAM に「フラット」命令セットがある場合、このオーバーヘッドは対数二乗に減少し、乗算のないフラット命令の場合は対数のみになります。これは、Word RAM に対して 2 次オーバーヘッドを必要とする既知のチューリング マシンの CoT シミュレーションとはまったく対照的です。

原文 (English)

Efficiently Representing Algorithms With Chain-of-Thought Transformers

The increasing popularity of \emph{reasoning} models -- language models that output a series of reasoning or thought tokens before producing an answer -- is justified, in part, by theoretical results showing that chain-of-thought (CoT) transformers can simulate Turing machines, and thus perform arbitrary computation. However, the Turing machine, while suitable for complexity-theoretic analysis, is not convenient, intuitive, or efficient for discussing algorithms. Algorithms are typically designed and analyzed at a higher level of abstraction, captured by the \emph{Word RAM} model with random-access memory and unit-cost operations on $\bigO(\log n)$-bit words. As a result, Word RAM algorithms can be substantially more efficient than their Turing machine counterparts, raising the question: \emph{Can CoT transformers efficiently simulate Word RAM algorithms?} For instance, can they sort $n$ items in $\bigO(n \log n)$ steps or run Dijkstra's algorithm in $\bigO(E + V \log V)$ steps? We answer affirmatively, up to poly-logarithmic overhead. We first establish this for finite-precision transformers with poly-logarithmic width and rightmost unique hard attention, then strengthen the result to two more practical settings with finite width and log-precision: \emph{continuous} CoT, where reasoning takes the form of vectors rather than tokens, and a \emph{hybrid} architecture in which transformer layers sit atop a recurrent (linear RNN) layer. In all three cases, we find that CoT \emph{can} efficiently simulate any Word RAM algorithm with only a poly-logarithmic overhead in $n$. This overhead reduces to log-square when the Word RAM has a ``flat'' instruction set, and only logarithmic for multiplication-free flat instructions -- in stark contrast to known CoT simulations of Turing machines, which require quadratic overhead over Word RAM.

2026-06-20 13:00 JSTarXiv cs.AIハードウェア/半導体

Mitigating Legibility Tax with Decoupled Prover-Verifier Games

As large language models become increasingly capable, it is critical that their outputs can be easily checked by less capable systems. Prov…

2026-06-19 16:59 JSTTechCrunch AIハードウェア/半導体

The US says ASML’s top chip tool may be in China, but how?

There's a commercial logic that cuts against the idea that ASML would risk its export license to arm a Chinese customer.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

緊急調整

大規模言語モデル (LLM) は、自身の出力が人間の倫理と乖離していることを識別できますか?そして彼らは自己修正できるのでしょうか? LLM に、自身の推論と出力をレビューする良心ステップを与え、直接優先最適化 (DPO) を使用して調整コンポーネントでトレーニング損失を拡張し、モデルを非倫理的な出力から遠ざけます。その結果、トレーニング、微調整、敵対的プロンプト、ゼロショット学習など、幅広いアプリケーションでモデルを調整するオンライン技術が実現しました。それは、より弱いまたはより強いジャッジを必要とせず、代わりにそれ自体の凍結されたコピーに依存します。以前の研究では、緊急不整合シナリオでは、モデルの微調整からコードのハッキングに至るまで、さまざまな緊急の非倫理的な行為が示されました。代わりに、私たちは緊急調整を達成する方法を経験的に示します。つまり、単一の高レベルの内省的な質問が、同じコード ハッキング シナリオの下で倫理モデルに向けてトレーニングを導きます。

原文 (English)

Emergent Alignment

Can Large Language Models (LLMs) discern when their own outputs are misaligned with human ethics? And can they self-correct? We endow an LLM with a conscience step that reviews its own reasoning and outputs, and we extend the training loss with an alignment component using Direct Preference Optimization (DPO) to steer the model away from non-ethical outputs. The result is an online technique to align models in a wide range of applications: training, fine-tuning, adversarial prompting, and zero-shot learning. It does not require a weaker or stronger judge, relying instead on a frozen copy of itself. In previous work, the Emergent Misalignment scenario showed a range of emergent unethical behaviors from fine-tuning the model to hack code. Instead, we empirically show how to achieve Emergent Alignment: a single high-level introspective question steers training toward an ethical model under the same code hacking scenario.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

大規模言語モデルのブラックボックス不確実性推定法の系統的評価

大規模言語モデル (LLM) は幅広いタスクにわたって強力な機能を示していますが、その出力は信頼性が低いことが多く、幻覚が含まれる可能性があるため、信頼できる LLM を構築するには不確実性推定 (UE) が不可欠です。実際には、多くの主流 LLM は制限された API を介してのみアクセスでき、ロジットや隠れ状態などの内部信号は利用できないため、ブラックボックス UE が特に重要になります。しかし、LLM 用のブラックボックス UE に関する既存の研究は、方法論において断片的なままであり、統一された実証的比較が欠けています。このギャップに対処するために、ブラックボックス UE 手法の体系的なレビューを提示し、言語化ベース、サンプリング ベース、説明ベース、マルチエージェント、およびハイブリッド手法の 5 つのカテゴリに整理します。さらに、統一された評価フレームワークを構築し、4 つのモデルと 4 つのデータセット設定にわたる 24 の代表的な手法をベンチマークします。私たちの結果は、すべての設定において一貫して優勢な単一の方法はないことを示しています。それにもかかわらず、回答空間内の候補を推論して比較する方法は一般に効果的であり、複数の不確実性信号を組み合わせるハイブリッド方法は、ほとんどの条件下で良好に機能します。ベンチマーク データと統一評価フレームワークを公開することで、再現可能な比較を促進し、将来の研究をサポートすることを目指しています。また、実証結果は、LLM 向けの将来のブラック ボックス UE 手法を開発するための実践的なガイダンスを提供します。

原文 (English)

A Systematic Evaluation of Black-Box Uncertainty Estimation Methods for Large Language Models

Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable and may contain hallucinations, making uncertainty estimation (UE) essential for building trustworthy LLMs. In practice, many mainstream LLMs are only accessible through restricted APIs, where internal signals such as logits and hidden states are unavailable, making black-box UE especially important. However, existing work on black-box UE for LLMs remains fragmented in methodology and lacks a unified empirical comparison. To address this gap, we present a systematic review of black-box UE methods and organize them into five categories: verbalization-based, sampling-based, explanation-based, multi-agent, and hybrid methods. We further build a unified evaluation framework and benchmark 24 representative methods across 4 models and 4 dataset settings. Our results show that no single method consistently dominates across all settings. Nevertheless, methods that reason over and compare candidates in the answer space are generally effective, and hybrid methods that combine multiple uncertainty signals perform well under most conditions. By releasing the benchmark data and a unified evaluation framework, we aim to facilitate reproducible comparisons and support future research, while our empirical findings provide practical guidance for developing future black-box UE methods for LLMs.

2026-06-19 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

思考連鎖トランスフォーマーによるアルゴリズムの効率的な表現

\emph{reasoning} モデル (答えを生成する前に一連の推論または思考トークンを出力する言語モデル) の人気が高まっていることは、思考連鎖 (CoT) 変換器がチューリング マシンをシミュレートし、任意の計算を実行できることを示す理論的結果によって部分的に正当化されます。ただし、チューリング マシンは複雑性理論の分析には適していますが、アルゴリズムを議論するのには便利ではなく、直観的でもなく、効率的でもありません。通常、アルゴリズムはより高い抽象レベルで設計および分析され、ランダム アクセス メモリと $\bigO(\log n)$ ビット ワードに対する単位コスト演算を備えた \emph{Word RAM} モデルによってキャプチャされます。その結果、Word RAM アルゴリズムはチューリング マシンのアルゴリズムより大幅に効率的になる可能性があり、\emph{CoT トランスフォーマーは Word RAM アルゴリズムを効率的にシミュレートできますか?} たとえば、$n$ 個の項目を $\bigO(n \log n)$ ステップでソートしたり、ダイクストラのアルゴリズムを $\bigO(E + V \log V)$ ステップで実行したりできるでしょうか? という疑問が生じます。多対数オーバーヘッドまでは肯定的に答えます。まず、多対数幅と右端の一意のハード アテンションを持つ有限精度変換器に対してこれを確立し、次にその結果を、有限幅と対数精度を持つ 2 つのより実用的な設定に強化します。\emph{continuous} CoT (推論がトークンではなくベクトルの形式をとる場合) と、変換器層がリカレント (線形 RNN) 層の上に位置する \emph{hybrid} アーキテクチャです。 3 つのケースすべてにおいて、CoT \emph{できる} が、$n$ の多対数オーバーヘッドのみであらゆる Word RAM アルゴリズムを効率的にシミュレートできることがわかります。 Word RAM に「フラット」命令セットがある場合、このオーバーヘッドは対数二乗に減少し、乗算のないフラット命令の場合は対数のみになります。これは、Word RAM に対して 2 次オーバーヘッドを必要とする既知のチューリング マシンの CoT シミュレーションとはまったく対照的です。

原文 (English)

Efficiently Representing Algorithms With Chain-of-Thought Transformers

The increasing popularity of \emph{reasoning} models -- language models that output a series of reasoning or thought tokens before producing an answer -- is justified, in part, by theoretical results showing that chain-of-thought (CoT) transformers can simulate Turing machines, and thus perform arbitrary computation. However, the Turing machine, while suitable for complexity-theoretic analysis, is not convenient, intuitive, or efficient for discussing algorithms. Algorithms are typically designed and analyzed at a higher level of abstraction, captured by the \emph{Word RAM} model with random-access memory and unit-cost operations on $\bigO(\log n)$-bit words. As a result, Word RAM algorithms can be substantially more efficient than their Turing machine counterparts, raising the question: \emph{Can CoT transformers efficiently simulate Word RAM algorithms?} For instance, can they sort $n$ items in $\bigO(n \log n)$ steps or run Dijkstra's algorithm in $\bigO(E + V \log V)$ steps? We answer affirmatively, up to poly-logarithmic overhead. We first establish this for finite-precision transformers with poly-logarithmic width and rightmost unique hard attention, then strengthen the result to two more practical settings with finite width and log-precision: \emph{continuous} CoT, where reasoning takes the form of vectors rather than tokens, and a \emph{hybrid} architecture in which transformer layers sit atop a recurrent (linear RNN) layer. In all three cases, we find that CoT \emph{can} efficiently simulate any Word RAM algorithm with only a poly-logarithmic overhead in $n$. This overhead reduces to log-square when the Word RAM has a ``flat'' instruction set, and only logarithmic for multiplication-free flat instructions -- in stark contrast to known CoT simulations of Turing machines, which require quadratic overhead over Word RAM.

2026-06-19 13:00 JSTarXiv cs.AIハードウェア/半導体

Mitigating Legibility Tax with Decoupled Prover-Verifier Games

As large language models become increasingly capable, it is critical that their outputs can be easily checked by less capable systems. Prov…

2026-06-19 03:22 JSTTechCrunch AIハードウェア/半導体

Amazon hopes to challenge Nvidia more directly by selling its AI chips

AWS is in talks to sell its chips to other data centers. CEO Andy Jassy has said this represents a $50 billion opportunity for the company.

2026-06-18 13:00 JSTarXiv cs.AIハードウェア/半導体

CaVe-VLM-CoT: 解釈可能な視覚言語モデル フレームワーク

視覚言語モデル (VLM) は依然として幻覚を起こしやすく、流暢ではあるが視覚的に不忠実な出力を生成します。既存の思考連鎖および検索強化手法は、ステップレベルの引用根拠を強制したり、検証の失敗を修正のために検索に戻すことを強制したりしていないため、この問題に部分的にしか対処していません。我々は、モジュール式のリフレクション ベースの Agentic-RAG フレームワークである CaVe-VLM-CoT を紹介します。これは、5 段階の閉ループ パイプライン (Extractor、Retriever、Solver、Citation Injector、Verifier) を通じて証拠に基づく推論を強制します。このパイプラインでは、根拠のない主張が検出されると、ターゲットを絞った再取得のために Extractor への構造化されたフィードバックがトリガーされます。既存のフレームワークでは、検索品質、段階的な引用の忠実度、およびクロスモーダル根拠を共同で測定できるものはないため、複合指標の重み付け精度、引用の精度と再現率、帰属、および証拠の根拠である CaVeScore を中心とした、すべての段階にわたる 23 のコンポーネントごとの指標のスイートを提案します。アーキテクチャや迅速な変更を行わなくても、CaVe-VLM-CoT は、ScienceQA で 87.1\% の精度と 56.6\% CaVeScore、MMMU (30 人の被験者) で 55.2\% の精度と 35.7\% CaVeScore を達成しました。

原文 (English)

CaVe-VLM-CoT: An Interpretable Vision-Language Model Framework

Vision-Language Models (VLMs) remain prone to hallucinations, producing fluent but visually unfaithful outputs. Existing chain-of-thought and retrieval-augmented methods only partially address this, as they neither enforce step-level citation grounding nor route verification failures back to retrieval for correction. We present CaVe-VLM-CoT, a modular reflection-based agentic-RAG framework that enforces evidence-grounded reasoning through a five-stage closed-loop pipeline: Extractor, Retriever, Solver, Citation Injector, and Verifier, in which detected ungrounded claims trigger structured feedback to the Extractor for targeted re-retrieval. Since no existing framework jointly measures retrieval quality, step-wise citation faithfulness, and cross-modal grounding, we propose a suite of 23 component-wise metrics across all stages, anchored by CaVeScore, a composite metric weighting accuracy, citation precision and recall, attribution, and evidence grounding. Without any architectural or prompt modifications, CaVe-VLM-CoT achieves 87.1\% accuracy and 56.6\% CaVeScore on ScienceQA , and 55.2\% accuracy and 35.7\% CaVeScore on MMMU (30 subjects).

2026-06-18 13:00 JSTarXiv cs.AIハードウェア/半導体

ゴースト アトラクタ ネットワーク: 閉ループ逐次生成のための盆地構造の動的デコーダ

大規模な Transformer および拡散デコーダを使用したシーケンシャル出力生成では、シーケンスの長さに応じて増大するメモリ コストに加え、ステップごとの反復計算が必要になります。それらを小型のフィードフォワード デコーダに置き換えると効率は回復しますが、構造化されていない潜在表現が生成され、閉ループ制御が制限されます。位相条件付きアクションの生成とクロスステップ潜在キャリーオーバーの両方には、安定した盆地を備えた潜在ジオメトリが必要です。この記事では、理論的に導出された動的デコーダであるゴースト アトラクタ ネットワークを提案します。このネットワークは、ドリフトを伴う学習されたポテンシャルの下で潜在的に進化し、構築によって盆地アトラクタ構造を生成します。 3 つの要望 (マルチモダリティ、デコーダ レベルのシングルパス スイッチング、および一定のメモリ) が潜在的なドリフト形式を動機付け、モード遷移はゴースト アトラクターの脱出を伴うサドルノード分岐として発生します。階層的な位相空間分解により、一次盆地収束と二次固有受容洗練が分離されます。経験的に、行動クローニングと対照的な目的を使用してエンドツーエンドでトレーニングされた Ghost は、潜在的に予測された勾配フロー収縮を示し、1,430 個のホールドアウト サンプルの 5 つの統合ステップにわたって勾配ノルムが 67% 減衰します。 Ghost はロボットアクションデコーダーとして評価されています。 230 万パラメータの Ghost は、462 分の 1 のパラメータと 32 分の 1 低いレイテンシで 10 億 7000 万パラメータの拡散トランスフォーマのオフライン精度に匹敵し、オフライン平均二乗誤差で 5 つの代替 2M パラメータ デコーダ (MLP、ニューラル ODE、CVAE、トランスフォーマ、1 ステップ拡散) を 5.9 ~ 29 パーセント上回ります。 LIBERO-10 閉ループ ベンチマークでは、Ghost の盆地構造潜在の位相調整により、フィードフォワード MLP ベースラインよりも 13.5 パーセント ポイントの成功率向上が得られ、永続的潜在アンサンブルの最終成功率は 95.7 パーセントに達します。

原文 (English)

Ghost Attractor Networks: Basin-Structured Dynamical Decoders for Closed-Loop Sequential Generation

Sequential output generation with large-scale Transformer and diffusion decoders pays a memory cost that grows with sequence length, plus iterative per-step computation. Replacing them with small feed-forward decoders restores efficiency but produces unstructured latent representations that limit closed-loop control: phase-conditioned action generation and cross-step latent carry-over both require a latent geometry with stable basins. This article proposes Ghost Attractor Networks, a theoretically derived dynamical decoder whose latent evolves under a learned potential with drift and produces a basin-attractor structure by construction. Three desiderata (multi-modality, decoder-level single-pass switching, and constant memory) motivate the potential-drift form, and mode transitions arise as saddle-node bifurcations with ghost-attractor escape. A hierarchical phase-space decomposition separates first-order basin convergence from second-order proprioceptive refinement. Empirically, a Ghost trained end-to-end with a behavioral-cloning and contrastive objective exhibits the predicted gradient-flow contraction in its potential, with the gradient norm decaying by 67 percent across five integration steps on 1430 held-out samples. Ghost is evaluated as a robotic action decoder. A 2.3-million-parameter Ghost matches the offline accuracy of a 1.07-billion-parameter Diffusion Transformer at 462 times fewer parameters and 32 times lower latency, and beats five alternative 2M-parameter decoders (MLP, Neural ODE, CVAE, Transformer, 1-step Diffusion) on offline mean squared error by 5.9 to 29 percent. On the LIBERO-10 closed-loop benchmark, phase conditioning on Ghost's basin-structured latent yields a 13.5 percentage-point success-rate gain over a feed-forward MLP baseline, and persistent-latent ensembling reaches a 95.7 percent final success rate.

2026-06-18 13:00 JSTarXiv cs.AIハードウェア/半導体

ピクセル化結合器とデュアルステートインピーダンス合成を使用した、ディープラーニング駆動のドハティパワーアンプの逆設計

ドハティ パワー アンプ (PA) の出力結合器は、負荷変調、インピーダンス マッチング、位相補償を単一のネットワーク内に統合しているため、その設計と合成は非常に困難です。この論文では、深層畳み込みニューラル ネットワーク (CNN)、ピクセル化されたレイアウト表現、遺伝的アルゴリズム (GA) をデュアルステート インピーダンス合成と組み合わせて、ピーク電力条件とバックオフ電力条件の両方に対処する 3 ポート ドハティ コンバイナー設計手法を提案します。概念実証として、3 ポートのピクセル化コンバイナーを組み込んだ 2 つの GaN HEMT Doherty PA プロトタイプが設計および製造されました。どちらのプロトタイプも、2.6 ~ 2.8 GHz 内で 71.2% 以上のピーク ドレイン効率で 44.2 dBm を超える実測飽和出力電力を達成しています。さらに、6dB のバックオフ レベルで 64% もの高いドレイン効率が測定されます。デジタル プリディストーションを適用した後、各プロトタイプは -51.3 dBc を超える隣接チャネル漏洩比 (ACLR) を達成しました。

原文 (English)

Deep Learning-Driven Inverse Design of Doherty Power Amplifiers Using Pixelated Combiners and Dual-State Impedance Synthesis

The output combiner of a Doherty power amplifier (PA) integrates load modulation, impedance matching, and phase compensation within a single network, making its design and synthesis highly challenging. In this paper, we propose a three-port Doherty combiner design methodology that combines deep convolutional neural networks (CNNs), pixelated layout representations, and genetic algorithms (GA) with dual-state impedance synthesis to address both peak and back-off power conditions. As a proof of concept, two GaN HEMT Doherty PA prototypes incorporating three-port pixelated combiners are designed and fabricated. Both prototypes achieve a measured saturated output power exceeding 44.2 dBm with peak drain efficiency above 71.2% within 2.6-2.8 GHz. Furthermore, a drain efficiency as high as 64% is measured at the 6-dB back-off level. After applying digital predistortion, each prototype achieves an adjacent channel leakage ratio (ACLR) better than -51.3 dBc.

2026-06-18 13:00 JSTarXiv cs.AIハードウェア/半導体

Veriphi: データセット依存のトレーニング方法を使用した攻撃ガイド型ニューラル ネットワーク検証

Veriphi は、アルファ、ベータ、CROWN メソッドを使用した高速な敵対的攻撃と正式な制限付き認証を組み合わせた、GPU で高速化されたニューラル ネットワーク検証システムです。 3 つのトレーニング手法 (標準、敵対的、認定) を使用した MNIST と CIFAR-10 の系統的な実験を通じて、トレーニング手法の有効性が基本的にデータセットに依存することを実証しました。 Interval Bound Propagation (IBP) は、単純な MNIST (784 次元) では 78% の認定精度を達成しますが、より複雑な CIFAR-10 データセットでは、PGD 敵対的トレーニングが支配的で、小さな摂動では 94% の認定が行われ、認定パフォーマンスは無視できます。攻撃に誘導された改ざんにより検証の 5 倍の高速化を達成し、実世界の航空宇宙物流の最適化のために量産サイズのモデル (1 億 580 万のパラメーター) にアプローチを拡張します。私たちの結果は、認定トレーニングが普遍的に敵対的トレーニングよりも優れているという仮定に疑問を呈し、検証戦略の選択においてコンテキストが非常に重要であることを示しています。

原文 (English)

Veriphi: Attack-Guided Neural Network Verification with Dataset-Dependent Training Methods

We present Veriphi, a GPU-accelerated neural network verification system that combines fast adversarial attacks with formal bound certification using alpha,beta-CROWN methods. Through systematic experiments on MNIST and CIFAR-10 using three training methodologies (standard, adversarial, certified), we demonstrate that training method effectiveness is fundamentally dataset-dependent. Interval Bound Propagation (IBP) achieves 78% certified accuracy on simple MNIST (784 dimensions) but provides negligible certification performance on the more complex CIFAR-10 dataset, where PGD adversarial training dominates with 94% certification at small perturbations. We achieve 5x verification speedup through attack-guided falsification and scale our approach to production-size models (105.8M parameters) for real-world aerospace logistics optimization. Our results challenge the assumption that certified training universally outperforms adversarial training, showing context matters critically for verification strategy selection.

2026-06-18 13:00 JSTarXiv cs.AIハードウェア/半導体

Spotlight: Synergizing Seed Exploration and Spot GPUs for DiT RL Post-Training

Reinforcement learning (RL) post-training of Diffusion Transformers (DiTs) is prohibitively expensive, requiring thousands of high-end GPUs…

2026-06-17 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

分離推論におけるアナーキーの代償

細分化された推論アーキテクチャは、プリフィル フェーズとデコード フェーズを個別の GPU プールに物理的に分離し、固定のハードウェア バジェットを共有する競合する「エージェント」を作成します。私たちの知る限り、NVIDIA Dynamo を具体的なケーススタディとして使用して、このアーキテクチャの最初の正式なゲーム理論分析を提供します。私たちは、プリフィル プールとデコード プール間の 2 プレイヤー リソース ゲーム、階層 KV キャッシュを介した利己的キャッシング ゲーム、およびリクエスト ルーティングの正の外部性を持つ輻輳ゲームの 3 つの結合ゲームとして、分解されたサービスをモデル化します。後者の 2 つは経験的に検証されています。 P/D リソース ゲームは分析的に扱われます (セクション 9.2)。私たちは、GPU の飽和がゲームの利得構造を変えるレジーム移行をどのように引き起こすかを特徴づけます。飽和以下では、利己的な行動がアナーキーの代価 (PoA) を制限します。飽和状態では、超線形レイテンシーとキャッシュ外部性により、経験的推定値 PoA ハット (セクション 6.4 で定義) が上昇します。この分析に基づいて、飽和遷移をリアルタイムで検出し、それに応じてルーティング パラメーターを調整し、キャッシュ アフィニティの活用から負荷分散された輻輳回避に移行する適応コントローラーを設計します。 Nemotron-4-340B (TP=8、クロス InfiniBand KV 転送を備えたフルノード ワーカー) と Llama-3.1-70B (TP=4) の 2 つのモデルを使用して Dynamo を実行する 3 ノード NVIDIA B200 クラスター上でフレームワークをインスタンス化し、両方のモデルで同じ最初のポストニー グリッド ポイント (C=128) を持つ同じ 3 レジーム PoA ハット構造を見つけました。アダプティブ ルーティングは、各モデルをより良い動作点にシフトします。最も強力な結果は 70B 1P/5D トポロジであり、PoA ハットは飽和フェーズで 13% のスループット コストで 3.1 倍 (66.4 から 21.5) 低下します。 70B 1P/2D では、PoA-hat は 2.2 倍、TTFT P99 は 7.6 倍に低下します (セクション 8.5 を参照)。

原文 (English)

The Price of Anarchy in Disaggregated Inference

Disaggregated inference architectures physically separate prefill and decode phases onto distinct GPU pools, creating competing "agents" that share a fixed hardware budget. We provide, to our knowledge, the first formal game-theoretic analysis of this architecture, using NVIDIA Dynamo as a concrete case study. We model disaggregated serving as three coupled games: a two-player resource game between prefill and decode pools, a selfish caching game over the hierarchical KV cache, and a congestion game with positive externalities for request routing. We empirically validate the latter two; the P/D resource game is treated analytically (Section 9.2). We characterize how GPU saturation induces regime transitions that shift the game's payoff structure: below saturation, selfish behavior has bounded Price of Anarchy (PoA); at saturation, superlinear latency and cache externalities drive our empirical estimator PoA-hat (defined in Section 6.4) upward. Based on this analysis, we design an adaptive controller that detects saturation transitions in real time and adjusts routing parameters accordingly, shifting from cache-affinity exploitation to load-balanced congestion avoidance. We instantiate our framework on a 3-node NVIDIA B200 cluster running Dynamo with two models, Nemotron-4-340B (TP=8, full-node workers with cross-InfiniBand KV transfers) and Llama-3.1-70B (TP=4), and find the same three-regime PoA-hat structure with the same first post-knee grid point (C=128) on both models. Adaptive routing shifts each model to a better operating point. Our strongest result is on the 70B 1P/5D topology, where PoA-hat drops 3.1x (66.4 to 21.5) in the saturated phase at a 13% throughput cost. On the 70B 1P/2D, PoA-hat drops 2.2x and TTFT P99 drops 7.6x (see Section 8.5).

2026-06-17 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

MODE: MoE マルチモーダル LLM 向けのモダリティ分解エキスパートレベル混合精度量子化

Mixture-of-Experts Multimodal Large Language Model (MoE-MLLM) は優れたパフォーマンスを提供しますが、法外な GPU メモリ コストがかかるため、圧縮が不可欠です。 PTQ 手法の中でも、エキスパート レベルの混合精度量子化は MoE-LLM に対して効果的であることが証明されていますが、エキスパートの重要度推定における 2 つの見落とされているバイアスにより、MoE-MLLM では顕著な低下に見舞われます。 (1) クロスモーダル レベルでは、ビジョン トークンの数値的優位性により、エキスパートの選択頻度がビジョン トークンによって支配され、テキスト モダリティに重要なエキスパートがマスクされます。 (2) ビジョン内レベルでは、冗長なビジョン トークンの大部分が頻度統計をさらに歪め、有益なビジュアル コンテンツに重要な専門家を曖昧にします。ギャップを埋めるために、モダリティごとにエキスパート選択周波数を分解し、冗長な視覚トークンをフィルタリングしてノイズ除去された視覚周波数を取得し、周波数ベースの推定に対する補完信号としてモダリティごとの量子化感度をさらに評価する、MoE-MLLM 用のモダリティ分解エキスパートレベル混合精度量子化フレームワークである MODE を提案します。これらの信号は整数線形計画法に統合され、指定された予算内でエキスパートごとのビット幅が割り当てられます。広範な実験により、MODE が MoE-MLLM に特に適しており、W3A16 での平均パフォーマンス損失を 2.9% 以内に制限し、極端な 2 ビット設定でより大きなゲインが得られることが示されています。

原文 (English)

MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs

Mixture-of-Experts Multimodal Large Language Models (MoE-MLLMs) offer remarkable performance but incur prohibitive GPU memory costs, making compression essential. Among PTQ methods, expert-level mixed-precision quantization has proven effective for MoE-LLMs, yet suffers notable degradation on MoE-MLLMs due to two overlooked biases in expert importance estimation. (1) At the cross-modal level, the numerical dominance of vision tokens causes expert selection frequency to be dominated by vision tokens, masking experts that are critical to the text modality; (2) at the intra-vision level, the large proportion of redundant vision tokens further skew frequency statistics, obscuring experts critical for informative visual content. To bridge gaps, we propose MODE, a modality-decomposed expert-level mixed-precision quantization framework for MoE-MLLMs that decomposes expert selection frequency by modality, filters redundant vision tokens to obtain denoised visual frequency, and further evaluates quantization sensitivity per modality as a complementary signal to frequency-based estimation. These signals are integrated into an Integer Linear Programming formulation to assign per-expert bit-widths under a given budget. Extensive experiments show that MODE is particularly well-suited for MoE-MLLMs, limiting average performance loss to within 2.9% at W3A16, with larger gains at the extreme 2-bit setting.

2026-06-17 05:00 JSTITmedia AI+LLM/生成AIロボティクスハードウェア/半導体規制/政策

生成AI×自動運転で注目のTesla・Waymo・NVIDIA 各社が目指す「フィジカルAI」は何が違うのか

日本政府が戦略的強化分野に掲げる「フィジカルAI」――その社会実装の最前線の一つが自動運転システムだ。熾烈な開発競争が繰り広げられている中、生成AIの進化は各社の競争にどのような変化をもたらしているのか。Tesla、Waymo、NVIDIAの最新動向を整理する。

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

適切な説明の定義と LLM 出力を説明する課題

適切な説明をどのように定義するかは、長年にわたる哲学的な議論ですが、最近、AI の出力の文脈で新たな関心が高まっています。説明可能性はさまざまな状況で AI 導入にとって重要ですが、AI システムの適切な説明を作成するには、まず適切な説明とは何かを理解する必要があります。この論文では、反事実的説明の概念に触発された定義を提案しますが、説明で提供される可能性のある各事実についての対話者の事前の信念も考慮する必要があると主張します。私たちは、AI の説明可能性に対するこの定義の影響、特に LLM 出力が適切な説明を生み出すのが難しい理由を調査します。

原文 (English)

A Definition of Good Explanations and the Challenges Explaining LLM Outputs

How to define a good explanation is a long-standing philosophical debate which has found recent renewed interest in the context of AI outputs. Explainability is crucial for AI adoption in many contexts, but in order to produce good explanations of AI systems, we must first have an understanding of what good explanations are. In this paper we propose a definition inspired by the notion of counterfactual explanations, however we argue that one must also take into account the interlocutor's prior beliefs in each fact that could be offered in an explanation. We explore the ramifications of this definition for AI explainability and, in particular, why LLM outputs are difficult to produce good explanations for.

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

LLM は失語症の談話における正しい情報単位を確実に識別しますか?

正しい情報単位(CIU)は、言語形式だけではなくコミュニケーション上の情報提供性を定量化するため、失語症の談話評価の中心となります。ただし、CIU の採点には時間がかかり、訓練を受けた評価者が必要です。この研究では、命令調整された大規模言語モデル (LLM) が失語症の談話トランスクリプトからトークンレベルの CIU 分類を確実に実行できるかどうかを検証しました。 Nicholas and Brookshire (1993) に従って、Cat Rescue 刺激で誘発された 16 枚の写真説明転写物に CIU ステータスの注釈が付けられました。サンプルは、対照、軽度、中等度、重度の失語症の 4 つの重症度層にまたがっていました。 4 つの公的に入手可能な命令調整 LLM が、5 つの層別ランダム シードにわたってゼロ ショットおよび 2 つの少数ショット プロンプト条件下でベンチマークされました。パフォーマンスは、精度、精度、再現率、F1、およびコーエンのカッパを使用してコンセンサス人間ラベルに対して評価されました。ゼロショットのプロンプトはモデル全体で不十分でした。対照的に、少数ショットのプロンプトは大幅な利益をもたらし、3 つの実行可能なモデルで競争力のあるパフォーマンスを生み出しました。平均数ショット F1 スコアは、Llama-3.1-8B、Qwen2.5-7B、および Mistral-7B 全体で 0.776 ~ 0.817 の範囲であり、固定グローバル サンプル選択とチャンクごとのローカル サンプル選択の間に大きな違いはありませんでした。 Phi-3-mini は不安定で、信頼できるパフォーマンスが得られませんでした。実行可能なモデルは高い再現率を示しましたが、精度は低く、トークンが CIU として体系的に過剰分類されていることを示唆しています。パフォーマンスは談話の重症度によっても異なり、最も弱い結果ではより重度の失語症が発生しました。フューショット LLM プロンプトは、勾配ベースのタスク トレーニングなしで自動 CIU 識別をサポートできますが、完全に自律的に使用するには人間による注釈との合意がまだ不十分です。これらの発見は、LLM ベースの CIU スコアリングが談話評価システムの人間参加型コンポーネントとして有望であることを裏付けています。

原文 (English)

Do LLMs Reliably Identify Correct Information Units in Aphasic Discourse?

Correct Information Units (CIUs) are central to discourse assessment in aphasia because they quantify communicative informativeness rather than linguistic form alone. However, CIU scoring is time intensive and requires trained raters. This study examined whether instruction-tuned large language models (LLMs) can reliably perform token-level CIU classification from aphasic discourse transcripts. Sixteen picture-description transcripts elicited with the Cat Rescue stimulus were annotated for CIU status according to Nicholas and Brookshire (1993). The sample spanned four severity strata: control, mild, moderate, and severe aphasia. Four publicly available instruction-tuned LLMs were benchmarked under zero-shot and two few-shot prompting conditions across five stratified random seeds. Performance was evaluated against consensus human labels using accuracy, precision, recall, F1, and Cohen's kappa. Zero-shot prompting was insufficient across models. In contrast, few-shot prompting yielded substantial gains and produced competitive performance for three viable models. Mean few-shot F1 scores ranged from 0.776 to 0.817 across Llama-3.1-8B, Qwen2.5-7B, and Mistral-7B, with no significant differences between fixed global and per-chunk local example selection. Phi-3-mini was unstable and did not yield reliable performance. Viable models showed high recall but lower precision, suggesting systematic over-classification of tokens as CIUs. Performance also varied by discourse severity, with the weakest results in more severe aphasia. Few-shot LLM prompting can support automated CIU identification without gradient-based task training, but agreement with human annotation remains insufficient for fully autonomous use. These findings support LLM-based CIU scoring as a promising human-in-the-loop component of discourse assessment systems.

2026-06-16 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

AI 多元主義と AI が見逃している世界

AI の多元性は、多くの場合、多様な価値観、好み、ユーザー、または出力を表す問題として組み立てられます。この論文は、AI システムもオントロジーを課すため、この枠組みは不完全であると主張します。AI システムは、エンティティ、関係、特徴、害、利益、証拠の有効な形式としてカウントされるものを定義します。私たちは、存在論的平坦化を、状況に応じて競合し、歴史的に特定された意味を、中立的で競合が難しいものとして扱われる、制限された技術カテゴリ、プロキシ、集約ルール、またはベンチマークターゲットに変換することと定義します。この論文は、価値多元主義、多元主義的調整、参加型および民主的 AI、手続き的正義、科学技術研究、説明責任研究、11 件の専門家インタビューからの集約テーマ、および 3 つの都市型 AI コンパニオン ケースにわたる、限定された概念的かつ定性的な統合を開発しています。これらの事例は、影響を受けるアクターが手続き上の地位を確立する前に、多元的手法がカテゴリ、プロキシ、集計ルール、および改訂権限を圧縮しながら、モデルの動作をどのように改善または構造化できるかを示しています。存在論的公開性、認識論的包含、手続き上の権限、評価の多元性、ライフサイクルの説明責任を文書化するための予備的な定性的監査足場として、多元的ライフサイクル ガバナンス (PLG) を導入します。 PLG は検証されたスコアリング手段としては提示されていません。これは、多元的 AI の証拠とガバナンス条件を明示するためのフレームワークです。

原文 (English)

AI Pluralism and the Worlds It Misses

AI pluralism is often framed as a problem of representing diverse values, preferences, users, or outputs. This paper argues that this framing is incomplete because AI systems also impose ontologies: they define what counts as an entity, relation, feature, harm, benefit, and valid form of evidence. We define ontological flattening as the conversion of situated, contested, and historically specific meanings into a restricted technical category, proxy, aggregation rule, or benchmark target that is treated as neutral and difficult to contest. The paper develops a bounded conceptual and qualitative synthesis across value pluralism, pluralistic alignment, participatory and democratic AI, procedural justice, science and technology studies, accountability research, aggregate themes from 11 expert interviews, and three urban AI companion cases. The cases illustrate how pluralistic methods can improve or structure model behavior while still compressing categories, proxies, aggregation rules, and revision rights before affected actors have procedural standing. We introduce Pluralistic Lifecycle Governance (PLG) as a preliminary qualitative audit scaffold for documenting ontological openness, epistemic inclusion, procedural authority, evaluation pluralism, and lifecycle accountability. PLG is not presented as a validated scoring instrument; it is a framework for making the evidence and governance conditions of pluralistic AI explicit.

2026-06-16 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

Enabling Real-Time Point-of-Care Ultrasound Segmentation: A GPU-Free Deployment in Resource-Limited Settings

Ultrasound imaging is the most widely adopted medical modality globally due to its low cost and portability, yet artificial intelligence (A…

2026-06-16 13:00 JSTarXiv cs.AIハードウェア/半導体

Constitutional Value Potentials: reading and steering internal priority margins in language models

A constitution tells a language model what to value, but little tells us whether it does. Adherence is judged from outputs, and output evid…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Service-Induced Congestion in Memory-Constrained LLM Serving

In large language model (LLM) serving, each request accumulates persistent graphics processing unit (GPU) memory during service as its key-…

2026-06-16 13:00 JSTarXiv cs.AIハードウェア/半導体

AnonShield: Scalable On-Premise Pseudonymization for CSIRT Vulnerability Data

We present AnonShield, a high-throughput, on-premise pseudonymization system that combines GPU-accelerated NER, streaming processing, cachi…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

From Tokens to Regions: CUDA-Sensitive Instruction Tuning for GPU Kernel Generation

High-performance CUDA kernels are essential for scalable AI systems, while Large Language Models (LLMs) still struggle to generate correct…

2026-06-16 13:00 JSTarXiv cs.AIハードウェア/半導体

NeuronFabric: A Software Reference Architecture for On-Chip Transformer Training with Local Adam

Publicly documented accelerator architectures generally separate training computation from optimizer-state updates or rely on external memo…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

daVinci-kernel: Co-Evolving Skill Selection, Summarization, and Utilization via RL for GPU Kernel Optimization

GPU kernel optimization represents a paradigm where functional correctness is assumed and execution efficiency is the objective. We present…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

VeriGraph: Towards Verifiable Data-Analytic Agents

LLM-based agents have demonstrated strong capabilities in data-intensive analytical tasks, yet their outputs are rarely verifiable: a relia…

2026-06-16 13:00 JSTarXiv cs.AIハードウェア/半導体

Fine-Tuning a 7B Advisor on Free-Tier GPUs: An Adapter-Handoff Recipe and a Synthetic-Data Reliability Caution

Fine-tuning a 7B language model for specialized advising is attractive in resource-constrained settings, but multi-epoch runs routinely exc…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体研究/論文

AgentLeak: A Benchmark for Internal-Channel Privacy Leakage in Multi-Agent LLM Systems

Multi-agent Large Language Model (LLM) systems create privacy risks that current output-only benchmarks cannot measure. When agents coordin…

2026-06-16 13:00 JSTarXiv cs.AIハードウェア/半導体

A Multi-level Analysis of Factors Associated with Student Performance: A Machine Learning Approach to the SAEB Microdata

Identifying the factors that influence student performance in basic education is a central challenge for formulating effective public polic…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Do You Really Need a GPU to Guard Your LLM? CPU-Class Classifiers and Multi-Stage Pipelines for Safety Enforcement at Scale

Safety classifiers that screen LLM inputs for jailbreak attempts have become standard deployment components, yet almost all production syst…

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

検出から回復まで: 504 GPU による LLM 事前トレーニングでの運用分析

大規模な AI トレーニングは現在、基本的に分散システムの問題となっており、ハードウェア障害はまれな例外ではなく、日常的な動作条件となっています。しかし、本番トレーニングクラスターからの公的運用証拠は依然として不足しています。この技術レポートは、55 日間の Prometheus 時系列データと 224 のマルチノード トレーニング セッションをカバーする 73 日間の運用ログを使用した、63 ノードの NVIDIA B200 実稼働クラスター (504 GPU) の実証分析を示しています。このクラスターは、5 者 (SKT、Upstage、Lablup、NVIDIA Korea、VAST Data) が統合された監視パイプラインを共有する組織間環境内で動作します。この配置により、2 ~ 4 ノード規模では現れなかった 60 ノード規模のストレージ I/O ボトルネックを共同診断することが可能になりました。これは、単一チームだけでは分離できない実稼働規模の現象です。数か月にわたる事前トレーニング キャンペーンを利用して、3 つの定量分析を実行し、4 つの結果を得ました。まず、751 の Prometheus メトリクスと 10 件の XID で特定された GPU 障害に関する統計分析により、1 日あたりの誤検知が約 0.84 件で 10/10 の検出率 (XID 前は 2/10) を達成しました。複数の信号検出戦略を動機付ける、複数の障害タイプにわたって一貫して支配的な単一の指標はありません。次に、GPU VRAM から NFS パスに沿った 523 のチェックポイント イベントのプロファイリングにより、「帯域幅のパラドックス」(200 Gbps RoCE の 1.4 ~ 10.4% の使用率) が 128 スロットの NFS RPC レイヤーの飽和に起因していることがわかります。 3 番目に、マルチノード障害の応答では、集中した除外 (63 ノード中上位 3 ノードがすべての除外の 50% 以上を占める) と、12 チェーン (試行 73 回) にわたる自動再試行チェーンの成功率が 33.3% で、手動回復率 12.5% の 2.7 倍であることが示されています。再試行間隔の中央値は 11 分 (IQR 10-11) です。すべての分析は実稼働インフラストラクチャに基づいており、セッション レベルのワークロード管理、GPU 中心のスケジューリング、および統合された可観測性を提供します。

原文 (English)

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs

Large-scale AI training is fundamentally a distributed systems problem, where hardware failures are routine operating conditions rather than rare exceptions, yet public operational evidence from production training clusters remains limited. This report presents an empirical analysis of a 63-node NVIDIA B200 production cluster (504 GPUs), using 55 days of Prometheus time-series data and 73 days of operational logs covering 224 multi-node training sessions. The environment is cross-organizational: five parties (SKT, Upstage, Lablup, NVIDIA Korea, VAST Data) share a unified monitoring pipeline. This enabled joint diagnosis of a 60-node-scale storage I/O bottleneck absent in 2-4-node tests, a production-scale phenomenon no single team could isolate alone. We perform three quantitative analyses yielding four findings. First, over 751 Prometheus metrics and 10 XID-identified GPU failures, no single metric is consistently dominant across failure types, motivating multi-signal detection. Second, 523 checkpoint events trace the save/load path from GPU VRAM to the NFS server: restart loading reaches 21.5% of maximum read bandwidth (700 GB/s) and save bursts 16.0% of maximum write bandwidth (250 GB/s), with NFS/RPC queueing and transport-layer backlog rising together. Third, across 224 sessions over 73 days, node exclusions concentrate so the top 3 of 63 nodes account for over 50%. Fourth, auto-retry chain analysis shows a 33.3% success rate over 12 chains (73 attempts), 2.7x the 12.5% manual rate, with a median retry interval of 11 minutes (IQR 10-11). All analyses are grounded in production infrastructure providing session-level workload management, GPU-centric scheduling, and unified observability.

2026-06-16 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

SAMark: 段落レベルの言い換え堅牢性を備えた自己アンカー付きテキスト透かし

意味レベルの透かし (SWM) は、文を基本単位として扱うことで、テキストの変更に対する堅牢性を向上させます。ただし、このような攻撃は文の順序を変更することで透かし信号を全体的に破壊するため、段落レベルの言い換えに対する堅牢性は依然として困難です。この研究では、意味空間にステップに依存しない緑色の領域を確立することで文の順序への依存を取り除く、自己アンカー型透かしフレームワークである SAMark を提案します。検出可能性を向上させるために、弱く位置合わせされた候補からのノイズを抑制しながら透かし信号を増幅するマルチチャネル双曲線スコアリング メカニズムを導入します。さらに、ハード フィルタリングとソフト正則化を組み合わせた多様性を意識したフィルタリング戦略を提案し、単純な N グラム繰り返しフィルタを超えて意味上の冗長性に対処します。実験結果は、SAMark が典型的な段落レベルの言い換え攻撃の下で最大 90.2% の TP@FP1% を達成し、以前の最も強力なベースラインを平均 30% 以上上回るパフォーマンスを示しながら、透かしなしのテキストと競争力のある生成品質を維持し、従来の方法を制限していた堅牢性と品質のトレードオフを打破することを示しています。

原文 (English)

SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness

Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to paragraph-level paraphrasing remains difficult because such attacks globally disrupt watermark signals by changing sentence order. In this work, we propose SAMark, a self-anchored watermarking framework that removes the dependency on sentence order by establishing a step-independent green region in semantic space. To improve detectability, we introduce a multi-channel hyperbolic scoring mechanism that amplifies watermark signals while suppressing noise from weakly aligned candidates. We further propose a diversity-aware filtering strategy that combines hard filtering with soft regularization, extending beyond simple n-gram repetition filters to address semantic redundancy. Experimental results show that SAMark achieves up to 90.2% TP@FP1% under typical paragraph-level paraphrasing attacks, outperforming the strongest prior baseline by more than 30% on average, while maintaining generation quality competitive with unwatermarked text and breaking the robustness-quality trade-off that limits prior methods.

2026-06-16 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

エネルギーの盲点: NVIDIA の主力エッジ AI ハードウェアはプロセスレベルのエネルギー属性をサポートできない

単一のユーザー目標によって複数ステップのオーケストレーション、ツール呼び出し、再試行、障害回復がトリガーされるエージェントティック AI ワークロードは、エッジ導入のターゲットとなっており、NVIDIA、デル、HP、ASUS、MSI、Acer、ギガバイトのすべてが 2026 年に GB10 ベースのデスクトップ AI システムを出荷します。私たちは最近、オーケストレーション構造がエージェントのエネルギー コストの大半を占めていることを実証しました。ワークフローは、成功した目標ごとに線形ベースラインよりも 4.33 倍多くのエネルギーを消費します。マルチステップ推論タスクの OOI は 7.63 倍に達します。これとは別に、Rajat et al。 CPU 側の処理が、エージェント ワークロードの総レイテンシの最大 90.6%、総動的エネルギーの 44% を占めることが示されています。私たちは、ASUS Ascent GX10 (GB10 SoC) の系統的なエネルギー観測可能性監査を報告し、このプラットフォームでは、サポートされているソフトウェア インターフェイスを通じて、CPU エネルギー カウンター、INA パワーレール モニター、IPMI/BMC、および SCMI パワーキャップ プロトコルを公開していないことがわかりました。唯一のオンデバイス エネルギー テレメトリは、NVML を介した瞬間的な GPU 電力です。さらに、MediaTek ファームウェアが文書化されていない ACPI インターフェイス (SPBM) を介してレールごとのエネルギーを内部で計算していることも判明しましたが、NVIDIA は「CPU レール情報を公開する予定はない」と述べています。したがって、RAPL 経由で x86 上で実行されるデバイス上のプロセスごとのエネルギー アトリビューションは、サポートされているインターフェイスを介してこのプラットフォームでは再現できません。私たちは、エネルギーに起因する AI のハードウェア要件仕様を形式化し、GPU 減算と組み合わせた外部 DC メータリングを使用した暫定キャリブレーション ブリッジを提案し、SCMI パワーキャップを介して標準トラック パスを特定します。私たちの調査結果は、低炭素コンピューティング コミュニティに、第一級のハードウェア要件としてエネルギーの可観測性を要求する動機を与えています。

原文 (English)

The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution

Agentic AI workloads - where a single user goal triggers multi-step orchestration, tool calls, retries, and failure recovery - are being targeted for edge deployment, with NVIDIA, Dell, HP, ASUS, MSI, Acer, and Gigabyte all shipping GB10-based desktop AI systems in 2026. We recently demonstrated that orchestration structure dominates agentic energy cost, with workflows consuming 4.33x more energy per successful goal than linear baselines and OOI reaching 7.63x for multi-step reasoning tasks. Separately, Raj et al. show that CPU-side processing accounts for up to 90.6% of total latency and 44% of total dynamic energy in agentic workloads. We report a systematic energy-observability audit of the ASUS Ascent GX10 (GB10 SoC) and find that the platform exposes no CPU energy counter, no INA power-rail monitor, no IPMI/BMC, and no SCMI powercap protocol through any supported software interface. The only on-device energy telemetry is instantaneous GPU power via NVML. We further discover that the MediaTek firmware already computes per-rail energy internally via an undocumented ACPI interface (SPBM), but NVIDIA states there are "no plans to expose CPU rail information." On-device per-process energy attribution - as performed on x86 via RAPL - is therefore not reproducible on this platform through supported interfaces. We formalize a hardware requirements specification for energy-attributed AI, propose an interim calibration bridge for per-domain energy decomposition - confirmed on the Acer Veriton GN100 where CPU energy accumulators are live - and identify a standards-track path via SCMI powercap. Our findings motivate the low-carbon computing community to demand energy observability as a first-class hardware requirement.

2026-06-16 13:00 JSTarXiv cs.AIハードウェア/半導体

必要なのは FP8 だけです (パート 1): HPC の聖杯としてのハードウェア FP64 の誤りを暴く

従来の HPC の定説では、ネイティブ ハードウェア FP64 シリコンは科学技術コンピューティングの還元不可能な基盤、つまり倍精度シミュレーションの「聖杯」であると考えられています。この論文では、この定説は間違っていると主張しています。B300 世代以降の AI に最適化された GPU では、豊富な FP8 テンソル スループットと中国剰余定理ベースの Ozaki Scheme II を組み合わせることで、正規の HPC カーネル スペクトル全体で完全な FP64 精度でメモリルーフ実行を回復します。 NVIDIA の Blackwell Ultra (B300) は、ネイティブ FP64 を約 1.3 TFLOPS (B200 から 31 倍) に低下させ、メモリに依存するカーネル (SpMV、GEMV、ステンシル) も計算に依存するようにレンダリングします。私たちは4つの貢献をしています。まず、統合分析モデルである Tensor-Memory Equilibrium (TME) モデルは、計算乗数アルファ、帯域幅乗数ベータ、および再構築レイテンシ ガンマでルーフラインを強化します。次に、ベータ -> 1 を駆動するメカニズムとしてレジスタレベルの融合を特定し、エミュレーションをメモリの壁の向こう側で本質的に自由にします。 3 番目に、Ozaki II ヴォールトは FP64 をネイティブの最大 1 TFLOPS から最大 500 TFLOPS (B300) および最大 400 TFLOPS (Rubin R200) までエミュレートし、帯域幅制限の領域ではメモリ上限に匹敵しながら、コンピューティング領域では B200 のネイティブ FP64 の上限を 1 桁以上上回ったと予測します。 4 番目に、H100 ベースラインに対して、Ozaki II は、B300 ネイティブ FP64 が課す最大 50 倍の回帰と比較して、調査したすべてのワークロードで H100 と一致またはそれを超えています。コンパニオン FFT 解析 (生き残った INT32 パイプでの Kulisch 固定点再構築) と、コンパニオン Part(2) 論文で報告されている FP32+Kahan 削減と組み合わせると、B300 で調査されたすべてのカーネル クラスがフル FP64 でメモリ ルーフに達します。証拠はタイトルの主張を裏付けています。Ozaki II と Kulisch のエスケープ ルートを備えた FP8 は、実稼働 HPC に必要なすべてです。ネイティブ FP64 シリコンは、もはやこれまで考えられてきた聖杯ではありません。

原文 (English)

FP8 is All You Need (Part 1): Debunking Hardware FP64 as the HPC Holy Grail (June 13th version)

Conventional HPC holds that native hardware FP64 is the irreducible foundation of scientific computing. On AI-optimized GPUs of the NVIDIA B300 generation and beyond, native FP64 throughput has collapsed to ~1.3 TFLOPS even as FP8 tensor throughput has grown to multiple PFLOPS. We argue something stronger than that this is survivable: the FP8 tensor-core matrix-multiply is the sole computational primitive on which double-precision scientific computing needs to be built. Every canonical kernel -- dense and sparse linear algebra, spectral transforms, stencils -- and every application composing them reduces, via the Chinese Remainder Theorem-based Ozaki Scheme II, to sequences of FP8 matrix operations; the only non-FP8 arithmetic is a bounded, fixed-width integer accumulation at reconstruction. Native FP64 is thereby demoted from a hardware requirement to a derived accuracy guarantee obtained by composition over the FP8 primitive. We organize the claim as a five-layer hierarchy -- the FP8 op, Ozaki II, the basic kernels or Berkeley "dwarfs", composite solvers, and full applications -- and, because the dwarf taxonomy already spans scientific computing, establish it by exhibiting the reduction for every dwarf rather than a sample. The claim is falsifiable, and we build the instrument that tests it: a Tensor-Memory Equilibrium (TME) model extending the Roofline with emulation parameters (alpha, beta, gamma). We identify register-level fusion as the mechanism that keeps emulation memory-bound, project recovered FP64 performance across B300 and Rubin against an H100 baseline, and close the kernel coverage with a companion FFT analysis and compensated reductions. The model could have returned a negative verdict; instead it passes across the dwarfs and their compositions. This is the analytical half of a two-part program, with a follow-on implementation to validate the thesis on real silicon.

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

安全プリミティブとしての機能の最小化: 最小権限の LLM エージェントに対するリスクを認識した因果関係ゲート

現代の意思決定システムは、学習されたコンポーネントへの依存度が高まっており、その出力は自信があるにもかかわらず間違っている可能性があり、下流のアクションがコストのかかるエラーにさらされることになります。リスク認識因果ゲーティング (RACG) を紹介します。これは、因果効果の推定と調整されたリスク制御を組み合わせることにより、モデルの予測に基づいて行動するか、延期するか、回避するかを決定するフレームワークです。 RACG は、候補アクションから結果に至るまでの因果経路をモデル化し、生の予測信頼度ではなく、推定された反事実リスクに基づいて各意思決定を制御します。ゲーティングの信頼性を高めるために、高リスク条件下で動作する確率に関する分布自由境界を導出し、これらの境界がユーザー指定の安全制約を満たす動作しきい値にどのように変換されるかを示します。さらに、予測結果と現実の結果の間の差異を監視し、因果関係の仮定が違反されているように見える場合にゲートを強化することで、分布のシフトを調整する適応型ゲートポリシーを提案します。シミュレートされた介入と現実世界の意思決定ベンチマーク全体で、RACG は、ゲートなし政策の有用性のほとんどを維持しながら、高コストのエラーを大幅に削減し、一致する棄権率で信頼度ベースの選択的予測ベースラインを上回ります。私たちの結果は、因果関係のリスクと予測の不確実性を明確に分離することで、より安全で透明性の高い意思決定システムが得られ、一か八かの状況において信頼できる自動化のための原則に基づいたメカニズムを提供することを示しています。

原文 (English)

Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents

Modern decision systems increasingly rely on learned components whose outputs may be confident yet wrong, exposing downstream actions to costly errors. We introduce Risk-Aware Causal Gating (RACG), a framework that decides whether to act on, defer, or abstain from a model's prediction by combining causal effect estimation with calibrated risk control. RACG models the causal pathway from candidate actions to outcomes and gates each decision according to an estimated counterfactual risk rather than raw predictive confidence. To make gating reliable, we derive distribution-free bounds on the probability of acting under high-risk conditions and show how these bounds translate into operating thresholds that satisfy user-specified safety constraints. We further propose an adaptive gating policy that adjusts to distribution shift by monitoring discrepancies between predicted and realized outcomes, tightening the gate when causal assumptions appear violated. Across simulated interventions and real-world decision benchmarks, RACG reduces high-cost errors substantially while preserving most of the utility of an ungated policy, and it outperforms confidence-based and selective-prediction baselines at matched abstention rates. Our results indicate that explicitly separating causal risk from predictive uncertainty yields decision systems that are both safer and more transparent, offering a principled mechanism for trustworthy automation in high-stakes settings.

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

コイン投げの裁判官? LLM-as-a-Judge 評価の信頼性と偏り

LLM-as-a-Judge は現在、モデル出力のランク付け、報酬モデルのトレーニング、公開リーダーボードへの入力に広く使用されていますが、その実行ごとの信頼性については十分に評価されていません。私たちは、2 つの OpenAI 判定モデル (GPT-4o-mini および GPT-4.1-mini) を使用して、10 カテゴリーにまたがる 29 のタスクについて同一の評価を繰り返し、質問ごとに 50 のペアワイズ トライアルと 50 のポイントワイズ トライアルを行い、温度および即時感度アブレーションを補足して研究しました。審査員全体でペアごとの好みは平均 13.6% の確率で反転し、28% の質問が反転率 20% を超え、1 つの質問は 56% に達しました。 GPT-4o-mini は、有意な 1 位バイアスも示します (A 多数派 72%、p = 0.024)。同時に、平均点ごとのスコアのギャップは小さく (10 点スケールで 0.19 ~ 0.36)、全体としては統計的に有意ではないため、ペアごとの点ごとのギャップが生じます。審査員は、自身のスカラー スコアが有意な質の違いの証拠をほとんど示さない場合でも、勝者を選択することがよくあります。裁判官内の不安定性を超えて、裁判官間の一致はわずか 76% ($\kappa = 0.51$) であり、意味的に同等のプロンプト テンプレートはテストされたケースの 25% で大多数の結果を変更し、決定論的なデコードは矛盾を軽減しますが、排除しません。信頼性曲線分析によると、私たちのデータセットでは、平均 95% の確率で 50 試行の参照評決を回復するための多数決には 11 回の反復試行が必要であり、分散が大きい質問の場合は 15 回に増加します。これらの発見は、単一試行の LLM 判定は一か八かの評価にはノイズが多すぎることが多く、複数試行の集計、位置のランダム化、明示的な不確実性レポートが標準的な手法であるべきであることを示唆しています。両方の審査員が単一のプロバイダーに属しているため、プロバイダー間のレプリケーションが引き続き重要な次のステップになります。

原文 (English)

The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation

LLM-as-a-Judge is now widely used to rank model outputs, train reward models, and populate public leaderboards, but its run-to-run reliability remains under-characterized. We study repeated identical evaluations on 29 tasks spanning 10 categories using two OpenAI judge models (GPT-4o-mini and GPT-4.1-mini), with 50 pairwise trials and 50 pointwise trials per question, supplemented by temperature and prompt-sensitivity ablations. Across judges, pairwise preferences flip on average 13.6% of the time, with 28% of questions exceeding a 20% flip rate and one question reaching 56%. GPT-4o-mini also exhibits a significant first-position bias (72% A-majority, p = 0.024). At the same time, mean pointwise score gaps are small (0.19--0.36 on a 10-point scale) and not statistically significant in aggregate, producing a pairwise--pointwise gap: judges frequently choose a winner even when their own scalar scores provide little evidence of a meaningful quality difference. Beyond within-judge instability, cross-judge agreement is only 76% ($\kappa = 0.51$), semantically equivalent prompt templates change majority outcomes in 25% of tested cases, and deterministic decoding reduces but does not eliminate inconsistency. A reliability curve analysis shows that, in our dataset, 11 repeated trials are needed for a majority vote to recover the 50-trial reference verdict with 95% probability on average, rising to 15 for high-variance questions. These findings suggest that single-trial LLM judging is often too noisy for high-stakes evaluation, and that multi-trial aggregation, position randomization, and explicit uncertainty reporting should be standard practice. Because both judges are from a single provider, cross-provider replication remains an important next step.

2026-06-15 13:00 JSTarXiv cs.AIハードウェア/半導体

極超音速流れの物理エミュレーターを構築するための完全に GPU ベースのワークフロー

複雑な物理現象を高い忠実度で低い計算コストで解決する能力は、現代のエンジニアリングにおける主要な課題に対処する上で中心となります。代表的な例は極超音速流れにあり、特に衝撃波の位置と強度に関して流れ場のトポロジー全体を正確に予測することが重要です。しかし、超音速および極超音速の流れは、産業関連のアプリケーションにおいて物理的な一貫性を保ちながら流れ状態の急勾配を捉えるのに苦労している従来の低次数モデルやニューラル エミュレーターにとって引き続き障害となります。そのために、不確実性の定量化と物理学を意識した改良によって強化されたニューラル エミュレーターのトレーニングと、高速化されたデータ生成を統合する、完全に GPU ベースのワークフローを導入します。私たちのワークフローは、微分可能な高忠実度ソルバー (JAX-Fluids) によって実現されており、データセットの迅速な作成とニューラル エミュレーターの残差ベースの改善に採用され、物理的な一貫性が強化されています。このフレームワークに基づいて、最初に一連のモデル アーキテクチャを提示し、そのスケーリング動作を分析して、その長所と欠点を明らかにします。次に、残差ベースのリファインにより、メッシュと入力パラメーターのみが利用可能な場合のトレーニングが可能になり、残差が大幅に削減され、物理的一貫性が向上することを示します。微分可能シミュレーションと残差ベースのリファインメントを組み合わせることで、トレーニング分布を超えても信頼性を維持できる物理エミュレーターが得られます。これは、現実世界のエンジニアリング設計ループにサロゲートを導入するための重要な要件です。

原文 (English)

A fully GPU-based workflow for building physics emulators of hypersonic flows

The ability to resolve complex physical phenomena with high fidelity and at low computational cost is central to addressing key challenges in modern engineering. A prime example lies in hypersonic flows, where the precise prediction of the full flowfield topology, in particular with respect to shock wave location and intensity, is critical. Yet supersonic and hypersonic flows continue to be a stumbling block for traditional reduced-order models and neural emulators that struggle to capture steep gradients in flow states with physical consistency in applications of industrial relevance. To that end, we introduce a fully GPU based workflow that integrates accelerated data generation with the training of neural emulators augmented by uncertainty quantification and physics-aware refinement. Our workflow is enabled by a differentiable high-fidelity solver (JAX-Fluids) which we employ for rapid dataset creation and residual-based improvement of the neural emulator to enhance physical consistency. Building on this framework, we first present a suite of model architectures and analyze their scaling behavior to expose their strengths and shortcomings. We then show that residual-based refinement enables training on cases where only mesh and input parameters are available, substantially reducing residuals and improving physical consistency. Together, differentiable simulation and residual-based refinement yield physics emulators that remain reliable beyond their training distribution, a key requirement for deploying surrogates in real-world engineering design loops.

2026-06-15 13:00 JSTarXiv cs.AIハードウェア/半導体

逆最適輸送による出発地と目的地のフローから都市アクセスコストを学習する

都市は、学校、診療所、交通機関、補助金付きのサービス ポイントなど、官民混合の施設ネットワークを通じて基本的なサービスを提供しています。これらのシステムでは、プランナーは多くの場合、世帯がどこに行くのかを観察しますが、距離、価格、制度へのアクセスなどの要素をトレードオフする潜在費用関数は観察しません。私たちはフィリピンの学校選択を通じてこの都市問題を研究します。フィリピンでは、この国最大の国の教育補助金が、混雑した公立学校から参加する私立学校に学習者を振り向けることを目的としています。学校から学校への入学フローをエントロピー最適輸送計画として扱い、2 つの相補的な逆最適輸送モデルを使用して潜在的な選択コストを回収します。補助金期間を持つ解釈可能な距離帯域モデルと、微分可能なシンクホーン順方向パスを通じて訓練されたニューラル コスト モデルです。このフレームワークは、最も人口の多い地域で観測された 23{,}820 の流れにわたる 283{,}016 人の学習者の旅行に適用され、補助金に相当する距離 $\lambda^{(k)}$ を推定します。これは、補助金によって相殺される知覚旅行コストのキロメートルとして解釈されます。この事例は、行政の出発地と目的地のデータを、アクセシビリティを意識した補助金設計、施設の配置、都市サービスの割り当てのための解釈可能な計画指標にどのように変換できるかを示しています。

原文 (English)

Learning Urban Access Costs from Origin-Destination Flows via Inverse Optimal Transport

Cities deliver basic services through mixed public-private facility networks, including schools, clinics, transit providers, and subsidized service points. In these systems, planners often observe where households go, but not the latent cost function through which they trade off factors such as distance, price, and institutional access. We study this urban problem through school choice in the Philippines, where the country's largest national education subsidy is intended to redirect learners from congested public schools to participating private schools. Treating school-to-school enrollment flows as an entropic optimal transport plan, we recover latent choice costs using two complementary inverse optimal transport models: an interpretable distance-banded model with a subsidy term, and a neural cost model trained through a differentiable Sinkhorn forward pass. Applied to 283{,}016 learner trips across 23{,}820 observed flows in the most populated region, the framework estimates a subsidy-equivalent distance, $\lambda^{(k)}$, interpreted as the kilometers of perceived travel cost offset by the subsidy. The case demonstrates how administrative origin-destination data can be transformed into interpretable planning metrics for accessibility-aware subsidy design, facility siting, and urban service allocation.

2026-06-15 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

Output Type Before Quality: A Standards-Derived XAI Admissibility Rubric for Autonomous-Driving Safety

Safety standards for ML-based autonomous driving specify the kind of evidence an assurance case must contain (directed cause-and-effect cha…

2026-06-15 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators

Large language models (LLMs) require larger GPU memory size these days, necessitating efficient and extreme weight compression methods. Exi…

2026-06-15 13:00 JSTarXiv cs.AIハードウェア/半導体

STaR-DRO: Stateful Tsallis Reweighting for Group-Robust Structured Prediction

Structured prediction with large language models requires outputs that are label-accurate, ontology-constrained, structurally valid, and ev…

2026-06-12 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

Efficient, Robust, and Anti-Collusion Fingerprinting of Image Diffusion Models

Model fingerprinting, embedding user-specific identifiers (fingerprints) into generated outputs, has recently emerged as a popular solution…

2026-06-12 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

Towards More General Control of Diffusion Models Using Jeffrey Guidance

A key strength of diffusion models lies in their flexibility, since their outputs can be controlled at sampling time through guidance. Howe…

2026-06-12 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning

Inference providers must maintain availability for many LLMs, including low-volume but essential models, making resource efficiency increas…

2026-06-11 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体ビジネス/資金調達研究/論文

医療研究分析用のスキル拡張 AI エージェント: NSCLC トランスクリプトーム バイオマーカー タスクにおける探索的なマルチモデルヒト評価

背景。生物医学研究をサポートするために大規模な言語モデルと AI エージェントがますます使用されていますが、ネイティブ モデルの出力では、重要な分析ステップが省略されたり、手法が誤用されたり、結論が誇張されたりする可能性があります。私たちは、医学研究スキル パッケージへの自律的なアクセスが、スキルを持たないネイティブ AI と比較して、AI によって生成されたトランスクリプトーム研究分析の高品質な出力に関連しているかどうかを評価しました。方法。私たちは、非小細胞肺がん免疫療法バイオマーカータスクを使用して、探索的なマルチモデルヒト評価を実施しました。 6 つのモデル バックボーンがテストされました。評価には、OpenClaw に代表される AI エージェント実装を通じて生成された 9 つのネイティブ AI 出力と 12 のスキル拡張出力の 21 件の匿名化された出力が含まれていました。 4 人の非専門生物医学評論家と 2 人の盲検専門家が各成果を評価し、各評論家のタイプごとに 2 つの評価を付けました。主な成果は、専門家が評価した全体的な品質でした。結果。スキル拡張された出力は、ネイティブ AI の出力よりも専門家の全体的な品質が方向性的に高いことを示しました (平均 5.50 vs 5.11; 差 = 0.39; ブートストラップ 95\% CI、-0.04 ~ 0.90; Welch p=0.156)。専門家以外の査読者の質も同じ傾向を示しました(平均 4.72 vs 4.47; 差 = 0.26; ブートストラップ 95\% CI、-0.25 ~ 0.80; Welch p=0.373)。専門家の合意は限られており (単一評価 ICC=-0.15)、モデル固有の効果は記述的で不均一でした。結論。この探索的サンプルでは、​​自律的スキル アクセスにより方向性のある品質シグナルが示されましたが、そのシグナルは専門家評価のノイズよりも小さいため、確認的な証拠として解釈されるべきではありません。この発見は主に、より強力な信頼性制御、プラットフォームの複製、生物学的妥当性評価を備えたスキル強化型 AI エージェントの大規模な評価の動機付けとなります。

原文 (English)

Skill-Augmented AI Agents for Medical Research Analysis: An Exploratory Multi-Model Human Evaluation in an NSCLC Transcriptomic Biomarker Task

Background. Large language models and AI agents are increasingly used to support biomedical research, but native model outputs may omit key analytical steps, misuse methods, or overstate conclusions. We evaluated whether autonomous access to a medical research skill package was associated with higher-quality AI-generated transcriptomic research-analysis outputs compared with native AI without skills. Methods. We conducted an exploratory multi-model human evaluation using a non-small cell lung cancer immunotherapy biomarker task. Six model backbones were tested. The evaluation included 21 anonymized outputs: 9 native-AI outputs and 12 skill-augmented outputs generated through an AI agent implementation represented by OpenClaw. Four non-expert biomedical reviewers and two blinded experts evaluated each output, with two ratings from each reviewer type. The primary outcome was expert-rated overall quality. Results. Skill-augmented outputs showed directionally higher expert overall quality than native-AI outputs (mean 5.50 vs 5.11; difference=0.39; bootstrap 95\% CI, -0.04 to 0.90; Welch p=0.156). Non-expert reviewer quality showed the same direction (mean 4.72 vs 4.47; difference=0.26; bootstrap 95\% CI, -0.25 to 0.80; Welch p=0.373). Expert agreement was limited (single-rating ICC=-0.15), and model-specific effects were descriptive and heterogeneous. Conclusions. Autonomous skill access showed a directional quality signal in this exploratory sample, but the signal was smaller than expert-rating noise and should not be interpreted as confirmatory evidence. The findings primarily motivate larger evaluations of skill-augmented AI agents with stronger reliability controls, platform replication, and biological-validity assessment.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

アーキテクチャから出力まで: 大規模言語モデルにおける幻覚の構造的起源とデータの役割の増大

大規模な言語モデルは幻覚を起こし、流暢で自信に満ちた、事実に誤りのある出力を生成しますが、その一貫性は世代や規模を超えて持続します。既存の分類法では、幻覚を出力タイプによって分類し、内因性の失敗と外因性の失敗、忠実さと事実の相違を区別しています。これらのフレームワークは記述的には厳密ですが、どの内部メカニズムが特定のインスタンスを生成したかは特定されません。この論文では、複合障害システムを形成する 3 つのアーキテクチャ上の決定の構造的結果としての幻覚を分析します。 Self-attention の共起学習は、統計的近接性を意味論的な意味に置き換え、エンティティの混乱、事実の誤った帰属、および意味論的なずれを引き起こします。最尤推定トレーニング目標は、事実の制約なしで次のトークンの確率を最適化し、真理値に関係なく統計的に妥当な出力を与えます。露出バイアスの下での自己回帰デコーディングの永続的な左から右へのコミットメントにより、単一の間違ったトークンが改訂されることなく出力シーケンス全体を通して前方にカスケードされることが保証されます。データセットの病理(ロングテールの欠如、トレーニングバイアス、合成汚染)は、これらの脆弱性を増幅させますが、独立してそれらを引き起こすわけではありません。私たちは 3 つの貢献を行っています。まず、各メカニズムをアランサリおよびルクマン分類法の特定の出力カテゴリにマッピングし、自己注意における内因性幻覚、MLE における外因性幻覚、および自己回帰デコードにおける論理的矛盾を特定します。第 2 に、一般的に引用される各データセットの病理が、独立して幻覚を引き起こすのではなく、これらのメカニズムのいずれかを利用していることを示します。第三に、出力タイプのみの分類の診断上の限界を特定し、それを推論層の緩和アプローチと対比します。

原文 (English)

From Architecture to Output: Structural Origins of Hallucination in Large Language Models and the Amplifying Role of Data

Large language models hallucinate--producing fluent, confident, factually wrong outputs--with a consistency that persists across generations and scales. Existing taxonomies classify hallucination by output type, distinguishing intrinsic from extrinsic failures and faithfulness from factuality divergence. These frameworks are descriptively rigorous but do not identify which internal mechanism produced a given instance. This paper analyses hallucination as a structural consequence of three architectural decisions that together form a compound failure system. Self-attention's co-occurrence learning substitutes statistical proximity for semantic meaning and produces entity confusion, fact misattribution, and semantic drift. The maximum likelihood estimation training objective optimises next-token probability without factual constraint, rewarding statistically plausible outputs regardless of their truth value. Autoregressive decoding's permanent left-to-right commitment under exposure bias ensures that a single wrong token cascades forward through the entire output sequence without revision. Dataset pathologies--long-tail deficiencies, training bias, and synthetic pollution--amplify these vulnerabilities but do not independently cause them. We make three contributions. First, we map each mechanism to a specific output category in the Alansari and Luqman taxonomy, locating intrinsic hallucination in self-attention, extrinsic hallucination in MLE, and logical inconsistency in autoregressive decoding. Second, we show that each commonly cited dataset pathology exploits one of these mechanisms rather than originating hallucination independently. Third, we identify the diagnostic limitation of output-type-only classification and contrast it with inference-layer mitigation approaches.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

構造的注意税: 検索形式がコンテンツに依存しない文脈内学習をどのようにハイジャックするか

検索拡張生成 (RAG) システムは、外部の知識を注入して LLM 出力を改善しますが、注入されたコンテンツの形式は、セマンティックな関連性とは別に、独立してモデルの注意分布を歪める可能性があります。私たちは構造的注意税と呼ぶ現象を特定し、形式化しました。ナレッジ グラフ (KG) の 3 倍は、リレーショナル区切り文字と繰り返されるスロット パターンにより、意味的に同等の自然言語テキスト ($\hat{o}$(KG) $\about$ 0.70 対 $\hat{o}$(neutral) $\about$ 0.25) よりもトークンあたり 2 ~ 3 倍多くの注意を獲得し、デモンストレーションの注意を最大 42% 圧縮します。 -- トリプルが関連しているかノイズであるかは関係ありません。私たちは、注意スコアを意味論的要素と構造的要素に分解する正式なフレームワークを開発し(方程式 2)、トークンレベルの形式バイアスとデモンストレーションの注意力損失を結び付ける圧縮限界(命題 1)を導出し、構造的用語がどれだけ注意がそらされるかを制御し、意味論的用語がそれが役立つか害を与えるかを制御することを示します。この分離により、検索拡張 ICL を改善するための 2 つの直交する軸、つまり検索品質の最適化 (意味軸) とフォーマット主導の注意捕捉の削減 (構造軸) が明らかになります。経験的に、2 つのモデル ファミリ (Mistral-7B、LLaMA-3-8B) と 3 つの QA ベンチマークにわたって、ソースとタスクのアラインメントが優勢であることが観察されます。タスクに一致する BM25 検索は、HotpotQA で 58 ~ 62% を達成するのに対し、ConceptNet の 25 ~ 27% を達成します。これは、すべてのゲート戦略 ($\leq$2 pp) を矮小化する 30 pp を超えるギャップです。フレームワークから、コストゼロの即時変更からトレーニング時の正規化まで、5 つの構造を意識した緩和戦略を導き出します。フォーマットの平坦化(S3)は、言語化されたトリプルコントロールからの精度と注意レベルの証拠の両方によって検証されますが、構造的分散(S1)は、フォーマットレベルの介入の課題を明らかにする混合の結果をもたらします。

原文 (English)

The Structural Attention Tax: How Retrieval Format Hijacks In-Context Learning Independent of Content

Retrieval-augmented generation (RAG) systems inject external knowledge to improve LLM outputs, yet the format of injected content -- distinct from its semantic relevance -- can independently distort the model's attention distribution. We identify and formalise a phenomenon we term the structural attention tax: knowledge graph (KG) triples, due to their relational delimiters and repeated slot patterns, capture 2-3x more attention per token than semantically equivalent natural-language text ($\hat{o}$(KG) $\approx$ 0.70 vs. $\hat{o}$(neutral) $\approx$ 0.25), compressing demonstration attention by up to 42% -- regardless of whether the triples are relevant or noise. We develop a formal framework decomposing attention scores into semantic and structural components (Eq. 2), derive a compression bound (Proposition 1) connecting token-level format bias to demonstration attention loss, and show that the structural term governs how much attention is diverted while the semantic term governs whether this helps or hurts. This decoupling reveals two orthogonal axes for improving retrieval-augmented ICL: optimising retrieval quality (semantic axis) and reducing format-driven attention capture (structural axis). Empirically, across two model families (Mistral-7B, LLaMA-3-8B) and three QA benchmarks, we observe that source-task alignment dominates: task-matched BM25 retrieval achieves 58-62% on HotpotQA vs. ConceptNet's 25-27%, a >30 pp gap that dwarfs all gating strategies ($\leq$2 pp). We derive five structure-aware mitigation strategies from the framework, ranging from zero-cost prompt modifications to training-time regularisation; format flattening (S3) is validated by both accuracy and attention-level evidence from a verbalized-triple control, while structural dispersal (S1) yields mixed results that illuminate the challenges of format-level intervention.

2026-06-11 13:00 JSTarXiv cs.AIハードウェア/半導体

取得後にポイズンが失敗した場合: パイプラインのチャンク化と再ランキングの下で​​コーパスポイズニングを再考する

検索拡張生成 (RAG) システムは、悪意のある知識の注入を通じて下流のモデル出力を操作するコーパス ポイズニング攻撃に対して脆弱です。既存の研究は主に、簡素化された取得設定の下でポイズニングを評価しており、ドキュメントのチャンク化、高密度の取得、再ランキング、およびグラウンディングされた生成を含む実際の RAG パイプラインを見落としています。この論文では、現実的な多段階の検索パイプラインの下でコーパスポイズニングを再検討し、多くの既存の攻撃が、高い検索段階の関連性を達成したにもかかわらず、再ランク付け後に大幅に低下することを示します。私たちは、この失敗の主な理由として、検索粒度の不一致を特定しました。ドキュメントレベルの敵対的シグナルは、チャンク化中に断片化されることがよくありますが、リランカーは、グローバルに最適化されたセマンティックな類似性よりも、ローカルで一貫性があり、回答が含まれるパッセージを好みます。この観察に基づいて、取得の関連性、リランカーの一貫性、およびチャンク境界の堅牢性を共同で最適化するポイズニング フレームワークであるチャンク対応およびリランク一貫性ポイズニング (CRCP) を提案します。 CRCP は、最適化中にチャンキング変換を明示的にモデル化し、さまざまなチャンキング構成の下でも効果を維持する、局所的に自己完結型の敵対的なパッセージを生成します。複数のリトリーバーとリランカーを使用した標準的な RAG ベンチマークの実験では、既存のポイズニング手法がチャンク サイズとリランカー戦略に非常に敏感であるのに対し、CRCP は現実的な検索パイプライン全体で大幅に高い攻撃成功率と強力な堅牢性を達成していることが示されています。私たちの調査結果は、現在の RAG セキュリティ評価における現実性の重要なギャップを浮き彫りにし、最新の RAG システムにおけるポイズニングは、検索のみの問題ではなく、多段階の検索一貫性の問題として研究されるべきであることを示唆しています。

原文 (English)

When Poison Fails After Retrieval: Revisiting Corpus Poisoning under Chunking and Reranking Pipelines

Retrieval-Augmented Generation (RAG) systems are vulnerable to corpus poisoning attacks that manipulate downstream model outputs through malicious knowledge injection. Existing studies mainly evaluate poisoning under simplified retrieval settings, overlooking practical RAG pipelines involving document chunking, dense retrieval, reranking, and grounded generation. In this paper, we revisit corpus poisoning under realistic multi-stage retrieval pipelines and show that many existing attacks substantially degrade after reranking despite achieving high retrieval-stage relevance. We identify retrieval granularity mismatch as a key reason for this failure: document-level adversarial signals are often fragmented during chunking, while rerankers favor locally coherent and answer-bearing passages rather than globally optimized semantic similarity. Based on this observation, we propose Chunk-aware and Rerank-Consistent Poisoning (CRCP), a poisoning framework that jointly optimizes retrieval relevance, reranker consistency, and chunk-boundary robustness. CRCP explicitly models chunking transformations during optimization to generate locally self-contained adversarial passages that remain effective under varying chunking configurations. Experiments on standard RAG benchmarks with multiple retrievers and rerankers show that existing poisoning methods are highly sensitive to chunk size and reranking strategies, whereas CRCP achieves substantially higher attack success rates and stronger robustness across realistic retrieval pipelines. Our findings highlight an important realism gap in current RAG security evaluation and suggest that poisoning in modern RAG systems should be studied as a multi-stage retrieval consistency problem rather than a retrieval-only problem.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

TileFuse: AMD NPU での効率的な量子化 LLM 推論のための融合型混合精度カーネル ライブラリ

オンデバイス LLM 推論に対する需要の高まりに伴い、厳しい電力と熱のバジェットの下でパフォーマンスとエネルギー効率を向上させるために、エッジ SoC は NPU を統合することが増えています。しかし、現在のクライアント NPU に実際に LLM を導入することは依然として困難です。AWQ などの広く使用されている量子化形式は、多くの既存の NPU ソフトウェア スタックにきれいにマッピングされておらず、多くの場合独自仕様であり、低レベルの制御が制限されています。この研究では、量子化 LLM 推論におけるトランス線形層をターゲットとする、AMD XDNA2 NPU 用のメタルに近い混合精度カーネル ライブラリである \textit{TileFuse} を紹介します。 TileFuse は、NPU 固有の量子化スキームに基づいてモデルを強制的に再形成するのではなく、AWQ スタイルの W4A16 や W8A16 などの実用的な低ビット フォーマットを XDNA2 に直接導入します。 TileFuse は、重みレイアウト、メタデータ配置、混合精度マイクロカーネル、配列レベルのデータフローを共同設計します。具体的には、アンパッキング、逆量子化、GEMM/GEMV の実行を単一のカーネル フローに融合し、最大 32K の GEMM 次元をサポートするインターリーブ プレタイリング レイアウトを導入し、完全な 4x8 AIE アレイを利用するように GEMV データフローを再設計します。カーネル レベルの評価全体で、TileFuse は、完全精度のベースラインと比較して、GEMM で最大 121.6%、GEMV で 281% パフォーマンスが向上し、GEMM 上の強力な iGPU ベースラインと比較して 2 倍を超えるパフォーマンスとエネルギー効率の向上を実現します。 Ryzen AI ラップトップでのエンドツーエンド LLM 実験では、TileFuse は、エネルギー消費量を 64.6% 以上削減し、プレフィル レイテンシーを最大 2.0 倍短縮することを達成しました。これらの結果を総合すると、XDNA2 が AWQ スタイルのエッジ LLM 推論の実用的なターゲットであること、および既製の量子化に対するネイティブ NPU サポートにより、実際のクライアント展開で NPU が大幅に使いやすくなることがわかります。

原文 (English)

TileFuse: A Fused Mixed-Precision Kernel Library for Efficient Quantized LLM Inference on AMD NPUs

With the growing demand for on-device LLM inference, edge SoCs increasingly integrate NPUs to improve performance and energy efficiency under tight power and thermal budgets. However, practical LLM deployment on current client NPUs remains difficult: widely used quantization formats such as AWQ do not map cleanly onto many existing NPU software stacks, which are often proprietary and expose limited low-level control. In this work, we present \textit{TileFuse}, a close-to-metal mixed-precision kernel library for AMD XDNA2 NPUs that targets transformer linear layers in quantized LLM inference. TileFuse brings practical low-bit formats such as AWQ-style W4A16 and W8A16 directly onto XDNA2, rather than forcing the model to be reshaped around an NPU-specific quantization scheme. TileFuse co-designs weight layout, metadata placement, mixed-precision microkernels, and array-level dataflow. Specifically, it fuses unpacking, dequantization, and GEMM/GEMV execution into a single kernel flow, introduces an interleaved pre-tiling layout that supports GEMM dimensions up to 32K, and redesigns GEMV dataflow to utilize the full 4x8 AIE array. Across kernel-level evaluations, TileFuse improves performance by up to 121.6% for GEMM and 281% for GEMV over full-precision baselines, while delivering more than 2x performance and energy-efficiency gains over strong iGPU baselines on GEMM. In end-to-end LLM experiments on Ryzen AI laptops, TileFuse achieves up to 2.0x lower prefilling latency with more than 64.6% lower energy consumption. Together, these results show that XDNA2 is a practical target for AWQ-style edge LLM inference and that native NPU support for off-the-shelf quantization can make NPUs substantially more usable in real client deployments.

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Substrate Asymmetry in User-Side Memory: A Diagnostic Framework

User-side memory in LLMs is typically scored as a single "personalization" capability: given a user's history, is the output more user-awar…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

Characterizing Software Aging in GPU-Based LLM Serving Systems

This paper proposes an empirical methodology to study software aging in GPU-based LLM serving systems. Traditional aging studies focus on C…

2026-06-11 13:00 JSTarXiv cs.AIハードウェア/半導体

Mathematical perspective on genetic algorithms with optimization guided operators

Recent work in ML applies genetic algorithms at inference time to iteratively improve solutions to optimization problems. The basic mutatio…

2026-06-11 13:00 JSTarXiv cs.AIハードウェア/半導体

Position: Stop Anthropomorphizing Intermediate Tokens as Reasoning/Thinking Traces!

Intermediate token generation (ITG), where a model produces output before the solution, has become a standard method to improve the perform…

2026-06-11 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Erased but Not Forgotten: How Backdoors Compromise Concept Erasure

The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications

Evaluating Large Language Model (LLM) applications differs from conventional software testing because outputs are probabilistic, semantical…

2026-06-11 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

On the Optimal Reasoning Length for RL-Trained Language Models

Reinforcement learning substantially improves reasoning in large language models, but it also tends to lengthen chain-of-thought outputs an…

2026-06-11 13:00 JSTarXiv cs.AIハードウェア/半導体

Estimating Tail Risks in Language Model Output Distributions

Language models are increasingly capable and are being rapidly deployed on a population-level scale. As a result, the safety of these model…

2026-06-11 12:00 JSTITmedia AI+画像/動画生成ハードウェア/半導体

Google、拡散型テキスト生成モデル「DiffusionGemma」公開 ローカルGPUで毎秒1000トークン超

Googleは、テキスト生成を最大4倍高速化する実験的AIモデル「DiffusionGemma」を発表した。画像生成の拡散手法を応用し、256トークンを一括で並列生成することで従来の自己回帰型モデルのボトルネックを解消する。品質は標準モデルに譲るものの、ローカル環境での高速なイ…

2026-06-10 17:30 JSTITmedia AI+LLM/生成AIハードウェア/半導体

「Siri AI」の進化に「Geminiそのまま」の誤解――現地取材で見えた“新生Apple Intelligence”の全貌

「GeminiがApple Intelligenceの正体」は誤解だ。WWDC 2026の現地取材で見えてきた第3世代は、200億パラメータのAIをiPhoneで動かす革新技術、Google Cloud+NVIDIAによるインフラ刷新、そして静かに変わる「無料」の定義まで、想像…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

推論からの構造、検索からの数値: 結合 MIMO コントローラー調整の構造事前分布としてのオンプレミスのオープン LLM

強結合多入力多出力 (MIMO) 産業プロセス向けのコントローラーのチューニングは困難です。分散型の古典的な自動チューニングではループの相互作用が無視され、自然な初期化からの局所的な数値最適化は、結果として生じる非凸コスト環境で行き詰まります。データをオンサイトに保持し、プラント モデルを必要としないオンプレミスのオープンソース大規模言語モデル (LLM) が役立つかどうかを尋ねます。シングル ループ CSTR では、従来のリレー フィードバック チューニング (IAE 0.106、最適値 0.102 に近い) は LLM チューナー (0.162) を上回ります。単純なループの場合、LLM は何も追加しません。この図は、設定値が矛盾する強結合四重タンクで反転し、アクチュエータのチャタリングを発生させずに追跡することに報いるペナルティ付きコスト J = IAE + lambda*TV(u) によってスコア付けされます。そこでは、単純なリレー チューニング (J ~ 28.6) と単純な LLM チューニング (29.7) はオープン ループ (22.7) よりも優れたものではなく、バランスのとれた開始からのローカル オプティマイザーは 10/10 の実行で失敗します。代わりに、スキャフォールドされたオープン LLM は結合について推論し、直観に反する非対称構造を提案し、どの開始点からでも J ~ 16.9 +/- 0.2 に達します。古典的なオプティマイザを使用してこれを調整すると、滑らかなグローバル最適値 (J ~ 12.0、10/10 対 0/10) が得られますが、これは非自明な負の積分補正を適用することさえできません。グローバル オプティマイザー (差分進化) もこの最適化に到達するため、LLM が唯一のルートではありません。その利点は、サンプル効率と解釈可能性です。つまり、18 の評価で使用可能なコントローラー (グローバル オプティマイザーはオープン ループよりも劣ります) に加えて、明示された理論的根拠が含まれています。このエッジは次元が上がるにつれて増大し、3x3 プラントでは評価が最大 6 分の 1 に達します。この動作は 4 つのオープン モデルにわたって一般化され、良性のプラントでは LLM には利点がなく、境界が明確になります。私たちは、オープン LLM が制御チューニングに役立つ場合の限界を定める再現可能なベンチマークに貢献します。オプティマイザーとしてではなく、サンプル効率が高く、解釈可能な構造事前処理としてです。

原文 (English)

Structure from Reasoning, Numbers from Search: On-Premise Open LLMs as Structural Priors for Coupled MIMO Controller Tuning

Tuning controllers for strongly coupled multi-input multi-output (MIMO) industrial processes is hard: decentralized classical auto-tuning ignores loop interaction, and local numerical optimization from natural initializations stalls in the resulting non-convex cost landscape. We ask whether on-premise open-source large language models (LLMs), which keep data on-site and need no plant model, can help. On a single-loop CSTR, classical relay-feedback tuning (IAE 0.106, near the 0.102 optimum) beats an LLM tuner (0.162): for simple loops the LLM adds nothing. The picture inverts on a strongly coupled quadruple-tank with conflicting set-points, scored by a penalized cost J = IAE + lambda*TV(u) that rewards tracking without chattering actuators. There, naive relay tuning (J ~ 28.6) and naive LLM tuning (29.7) are no better than open loop (22.7), and a local optimizer from balanced starts fails in 10/10 runs. A scaffolded open LLM instead reasons about the coupling, proposes the counter-intuitive asymmetric structure, and reaches J ~ 16.9 +/- 0.2 from any start; refining it with a classical optimizer attains the smooth global optimum (J ~ 12.0, 10/10 vs. 0/10), which even applies a non-obvious negative integral correction decentralized tuning cannot. A global optimizer (differential evolution) also reaches this optimum, so the LLM is not the only route; its advantage is sample efficiency and interpretability: a usable controller in 18 evaluations (where the global optimizer is worse than open loop) plus a stated rationale. This edge grows with dimension, reaching ~6x fewer evaluations on a 3x3 plant. The behaviour generalizes across four open models, and on a benign plant the LLM offers no advantage, sharpening the boundary. We contribute a reproducible benchmark delimiting when open LLMs help in control tuning: not as optimizers, but as a sample-efficient, interpretable structural prior.

2026-06-10 13:00 JSTarXiv cs.AIハードウェア/半導体

Sigma-Branch: Hierarchical Single-Path Network Reconstruction for Dynamic Inference with Reduced Active Parameters

Deploying deep neural networks on memory-constrained edge accelerators is bottlenecked by per-inference off-chip weight transfer rather tha…

2026-06-10 13:00 JSTarXiv cs.AIハードウェア/半導体

Minimum Distortion Quantization with Specified Output Distribution

We derive the optimal quantizer of a real-valued random variable $W$ with distribution $P_W$ such that 1) the distribution of the quantizat…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Stop Early, Spend Less: Hidden-State Probes as a Practical Recipe for Streaming Moderation of LLM Outputs

Deploying large language models in user-facing systems requires efficient output safety filtering. Existing approaches typically rely on a…

2026-06-10 13:00 JSTarXiv cs.AIハードウェア/半導体

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design

Local deployment of large Mixture-of-Experts (MoE) models falls short of the service quality achieved in cloud-scale environments, even und…

2026-06-10 13:00 JSTarXiv cs.AIハードウェア/半導体

What Do Deepfake Speech Detectors Actually Hear?

Deepfake speech detectors often output a single score without explaining why an audio sample is flagged, where in the signal the evidence l…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

AMEL: LLM の判断に対する蓄積されたメッセージの影響

大規模な言語モデルは、コードのレビュー、コンテンツの調整、または 1 つの会話で多くの項目が通過する出力のスコア付けなど、自動評価器として日常的に使用されます。以前の会話履歴の極性がその後の判断にバイアスをかけるかどうかを尋ねます。この効果を、LLM 判断に対する蓄積されたメッセージ効果 (AMEL) と呼びます。 4 つのプロバイダー (OpenAI、Anthropic、Google、および 4 つのオープンソース モデル) の 11 モデルに対する 75,898 回の API 呼び出しにわたって、同一のテスト項目を単独で提示するか、主に肯定的または否定的な評価で飽和した履歴を追跡します。モデルは会話の一般的な極性に向かってシフトします (d = -0.17、p < 10^-46)。この効果は、モデルがベースラインで本当に不確実である項目に集中します (高エントロピー項目の場合は d = -0.34、ベースラインが決定的である場合は d = -0.15)。バイアスはコンテキストの長さとともに増加しません。前のターン 5 と 50 では同じシフトが生成されます (Spearman |r| < 0.01; OLS 勾配 p = 0.80)。そして、否定性の非対称性があります。項目ごとにペアにすると、否定的な履歴は肯定的な履歴よりも 1.62 倍のバイアスを引き起こします (t = 13.46、p < 10^-39、n = 2,481)。スケーリングは役立ちますが、解決しません (Anthropic: Haiku -0.22 ~ Opus -0.17、OpenAI: Nano -0.34 ~ GPT-5.2 -0.17)。 3 回のフォローアップによりメカニズムが絞り込まれます。トークンの確率分布は、しきい値ではなく連続的に変化します。負の非対称性にはトークンレベルの要素と意味論的な要素の両方がありますが、バランスの原因を特定することはサンプルサイズでは探索的です。位置は重要ではありません。50 ターン履歴のどこかで 5 つの偏ったターンが同じシフトを生成します。評価パイプラインの最も簡単な修正は、項目ごとに新しいコンテキストを作成することです。バッチ処理が避けられない場合は、履歴のバランスを取ると役立ちます。

原文 (English)

AMEL: Accumulated Message Effects on LLM Judgments

Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing through one conversation. We ask whether the polarity of prior conversation history biases subsequent judgments, an effect we call the accumulated message effect on LLM judgments (AMEL). Across 84,088 API calls to 12 models from 5 providers (OpenAI, Anthropic, Google, DeepSeek, and four open-source models), we present identical test items in isolation or following histories saturated with predominantly positive or negative evaluations. Models shift toward the conversation's prevailing polarity (d = -0.17, p < 10^-53). The effect concentrates on items where the model is genuinely uncertain at baseline (d = -0.36 for high-entropy items, vs d = -0.15 when the baseline is deterministic). Bias does not grow with context length: 5 prior turns and 50 produce the same shift (Spearman |r| < 0.01; OLS slope p = 0.80). And there is a negativity asymmetry: paired per item, negative histories induce 1.52x more bias than positive (t = 13.03, p < 10^-36, n = 2,733). Scaling helps but does not solve it (Anthropic: Haiku -0.22 to Opus -0.17; OpenAI: Nano -0.34 to GPT-5.2 -0.17). Three follow-ups narrow the mechanism. The token probability distribution shifts continuously, not at a threshold. The negativity asymmetry has both token-level and semantic components, though attributing the balance is exploratory at our sample sizes. Position does not matter: five biased turns anywhere in a 50-turn history produce the same shift. The simplest fix for evaluation pipelines is a fresh context per item; when batching is unavoidable, balancing the history helps.

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ハードウェア/半導体ビジネス/資金調達

Conditional Vendi Score: Prompt-Aware Diversity Evaluation for Generative AI Models and LLMs

Generative models guided by text prompts are widely evaluated for fidelity and prompt alignment, yet their ability to produce outputs remai…

2026-06-10 13:00 JSTarXiv cs.AIハードウェア/半導体

Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators

As generative AI systems are integrated into educational settings, students often encounter AI-generated output while working through learn…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

検出から回復まで: 504 GPU による LLM 事前トレーニングでの運用分析

大規模な AI トレーニングは現在、基本的に分散システムの問題となっており、ハードウェア障害はまれな例外ではなく、日常的な動作条件となっています。しかし、本番トレーニングクラスターからの公的運用証拠は依然として不足しています。この技術レポートは、55 日間の Prometheus 時系列データと 224 のマルチノード トレーニング セッションをカバーする 73 日間の運用ログを使用した、63 ノードの NVIDIA B200 実稼働クラスター (504 GPU) の実証分析を示しています。このクラスターは、5 者 (SKT、Upstage、Lablup、NVIDIA Korea、VAST Data) が統合された監視パイプラインを共有する組織間環境内で動作します。この配置により、2 ~ 4 ノード規模では現れなかった 60 ノード規模のストレージ I/O ボトルネックを共同診断することが可能になりました。これは、単一チームだけでは分離できない実稼働規模の現象です。数か月にわたる事前トレーニング キャンペーンを利用して、3 つの定量分析を実行し、4 つの結果を得ました。まず、751 の Prometheus メトリクスと 10 件の XID で特定された GPU 障害に関する統計分析により、1 日あたりの誤検知が約 0.84 件で 10/10 の検出率 (XID 前は 2/10) を達成しました。複数の信号検出戦略を動機付ける、複数の障害タイプにわたって一貫して支配的な単一の指標はありません。次に、GPU VRAM から NFS パスに沿った 523 のチェックポイント イベントのプロファイリングにより、「帯域幅のパラドックス」(200 Gbps RoCE の 1.4 ~ 10.4% の使用率) が 128 スロットの NFS RPC レイヤーの飽和に起因していることがわかります。 3 番目に、マルチノード障害の応答では、集中した除外 (63 ノード中上位 3 ノードがすべての除外の 50% 以上を占める) と、12 チェーン (試行 73 回) にわたる自動再試行チェーンの成功率が 33.3% で、手動回復率 12.5% の 2.7 倍であることが示されています。再試行間隔の中央値は 11 分 (IQR 10-11) です。すべての分析は実稼働インフラストラクチャに基づいており、セッション レベルのワークロード管理、GPU 中心のスケジューリング、および統合された可観測性を提供します。

原文 (English)

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs

Large-scale AI training is fundamentally a distributed systems problem, where hardware failures are routine operating conditions rather than rare exceptions, yet public operational evidence from production training clusters remains limited. This report presents an empirical analysis of a 63-node NVIDIA B200 production cluster (504 GPUs), using 55 days of Prometheus time-series data and 73 days of operational logs covering 224 multi-node training sessions. The environment is cross-organizational: five parties (SKT, Upstage, Lablup, NVIDIA Korea, VAST Data) share a unified monitoring pipeline. This enabled joint diagnosis of a 60-node-scale storage I/O bottleneck absent in 2-4-node tests, a production-scale phenomenon no single team could isolate alone. We perform three quantitative analyses yielding four findings. First, over 751 Prometheus metrics and 10 XID-identified GPU failures, no single metric is consistently dominant across failure types, motivating multi-signal detection. Second, 523 checkpoint events trace the save/load path from GPU VRAM to the NFS server: restart loading reaches 21.5% of maximum read bandwidth (700 GB/s) and save bursts 16.0% of maximum write bandwidth (250 GB/s), with NFS/RPC queueing and transport-layer backlog rising together. Third, across 224 sessions over 73 days, node exclusions concentrate so the top 3 of 63 nodes account for over 50%. Fourth, auto-retry chain analysis shows a 33.3% success rate over 12 chains (73 attempts), 2.7x the 12.5% manual rate, with a median retry interval of 11 minutes (IQR 10-11). All analyses are grounded in production infrastructure providing session-level workload management, GPU-centric scheduling, and unified observability.

2026-06-10 13:00 JSTarXiv cs.AIハードウェア/半導体

OPRD: On-Policy Representation Distillation

On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities. This output-only paradigm ha…

2026-06-10 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

PAFO: パーソナライズされた報酬モデリングのためのパレート公平性の最適化

大規模言語モデル (LLM) は、出力を多様なユーザーの好みに合わせるために、報酬モデルへの依存度を高めています。パーソナライズされた報酬モデルはそのような異質性を捉えることを目的としていますが、多くの場合、不均衡なユーザーの嗜好データに基づいてトレーニングされるため、トレーニング母集団の中でより一般的な嗜好を持つユーザーを優先する可能性があります。この論文では、この失敗モードを個人化された報酬バイアスとして特定します。報酬モデリングの品質は、選好支持率に応じて体系的に変化します。私たちはその緩和策をグループ ユーティリティに対するパレート公平性問題として定式化し、他のユーザー グループを低下させることなくサービスが十分に受けられていないユーザーを改善することを目指しています。この目的を達成するために、パーソナライズされた報酬モデリングのためのパレート公平性最適化フレームワークである PAFO を提案します。 PAFO は、まず多数派と少数派の選好グループに対してグループに特化した報酬モデルをトレーニングし、次に条件付きマージンレベルの監視を構築して、不均一な選好の境界を単一の統一モデルに抽出します。結果として得られるモデルは、トレーニング中にのみグループ情報を使用し、推論時に明示的なグループ ラベルを必要としません。 Personal-LLM と DSP の実験では、PAFO が少数派グループと多数派グループの両方の精度を向上させながら、複数の指標にわたるユーザーレベルの不公平性を軽減することが示されており、より公平な LLM パーソナライゼーションに対する PAFO の有効性が実証されています。

原文 (English)

PAFO: Pareto Fairness Optimization for Personalized Reward Modeling

Large language models (LLMs) increasingly rely on reward models to align their outputs with diverse user preferences. While personalized reward models aim to capture such heterogeneity, they are often trained on imbalanced user preference data and may therefore favor users whose preferences are more common in the training population. In this paper, we identify this failure mode as personalized reward bias, where reward modeling quality varies systematically with preference support rate. We formulate its mitigation as a Pareto fairness problem over group utilities, aiming to improve under-served users without degrading other user groups. To this end, we propose PAFO, a Pareto fairness optimization framework for personalized reward modeling. PAFO first trains group-specialized reward models for majority and minority preference groups, then constructs conditional margin-level supervision to distill their heterogeneous preference boundaries into a single unified model. The resulting model uses group information only during training and requires no explicit group labels at inference time. Experiments on Personal-LLM and DSP show that PAFO improves both minority-group and majority-group accuracy while reducing user-level unfairness across multiple metrics, demonstrating its effectiveness for fairer LLM personalization.

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体ビジネス/資金調達

オンライン エージェント アズ ア ジャッジ: インタラクティブ エージェントの状況を生み出す評価

社会的に関連した行動は、孤立した出力だけでなく、以前の相互作用、社会的役割、下流の行動にも依存するため、LLM を利用した対話型ソーシャル エージェントの評価は困難です。既存の方法では、通常、ターゲット エージェントが環境内で自由に行動し、その結果として得られる軌跡をスコアリングできます。ただし、この受動的な設定では、特定の社会的状況下でのみ観察可能になる機能が見逃される可能性があります。たとえば、意見の相違が生じない場合、競合処理はテストされないままになる可能性があります。私たちは、対話型ソーシャル エージェントのための状況生成評価フレームワークである Online Agent-as-a-Judge を提案します。 Online Agent-as-a-Judge は、環境のネイティブ対話およびアクション プロトコルを通じてターゲット エージェントと対話するインワールド評価エージェントをデプロイし、評価基準に関連する状況を積極的に引き出します。結果として得られる軌跡は、即時の反応とその後の行動の両方を評価するための証拠を提供します。 32 ドルのデザイナーが作成した社会的基準を備えたライフ シミュレーション環境では、オンライン エージェントとしての裁判官は、基準の適用範囲と人間のラベルとの一致を改善し、受動的手法では観察されない可能性がある行動について、より信頼性の高い証拠に基づいた評価をもたらします。

原文 (English)

Online Agent-as-a-Judge: Situation-Generating Evaluation for Interactive Agents

Evaluating LLM-powered interactive social agents is challenging because socially relevant behaviors depend not only on isolated outputs, but also on prior interactions, social roles, and downstream actions. Existing methods typically allow a target agent to act freely in an environment and then score the resulting trajectory. However, this passive setup can miss capabilities that only become observable under specific social circumstances; for example, conflict handling may remain untested if no disagreement arises. We propose Online Agent-as-a-Judge, a situation-generating evaluation framework for interactive social agents. Online Agent-as-a-Judge deploys an in-world evaluator agent that interacts with the target agent through the environment's native dialogue and action protocol, actively eliciting situations relevant to the evaluation criteria. The resulting trajectories provide evidence for assessing both immediate responses and subsequent behavior. In a life-simulation environment with $32$ designer-authored social criteria, Online Agent-as-a-Judge improves criteria coverage and agreement with human labels, yielding more reliable evidence-grounded evaluations of behaviors that passive methods can leave unobserved.

2026-06-09 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

ZIPP:ペルソナによるゼロショット画像パーソナライゼーション

テキストから画像への拡散モデルは、制限のない創造的なコンテキストでますます導入されていますが、その出力は個人的なものではなく、個人の好みではなく集合的な美学に合わせて最適化されたままです。人間の好みは多元的です。落ち着いたノスタルジックなポートレートを好むユーザーが鮮やかなストリート写真を好む一方で、夢のような映画の美学に惹かれるユーザーもいます。既存の方法では、高密度のインタラクション履歴やユーザーごとの微調整が必​​要となり、コールドスタート設定が失敗し、コンテキスト依存の設定が静的な表現に崩壊してしまいます。ペルソナによるゼロショット画像パーソナライゼーション (ZIPP) を導入します。これは、ユーザー固有のデータや重みの更新を行わずに、自然言語ペルソナ (ユーザーのアイデンティティと美的感性の簡潔な記述子) に基づいて画像生成を条件付けします。 ZIPP は、LLM を使用して特定のペルソナの観点からプロンプトを書き換え、パーソナライズされた出力に向けて拡散モデルを導きます。大規模にペルソナをマイニングするために、2,200 万人のユーザー Reddit インタラクション グラフ上で帰納的グラフ アテンション ネットワークをトレーニングし、グラフ構造を視覚的な動作と一致させる 2 つの対照的な目標を設定し、学習した表現を MLLM を介して自然言語ペルソナに言語化します。 1.5K のユーザー、グラフマイニングされたペルソナ、および 40K の生成された画像を備えた初のゼロショット パーソナライゼーション ベンチマークである ZIPBench を紹介します。 5 つのモデル ファミリにわたる 4 つのベンチマークと 14 の LLM にわたって、ペルソナ コンディショニングは一貫した利益 (13 ~ 20%) をもたらし、フロンティア モデルで最も恩恵を受けています。少数ショット設定では、ZIPP は、ユーザーごとに 100 以上のサンプルでトレーニングされた微調整されたベースラインと一致またはそれを上回ります。 ZIPP は最も低い選好分布の乖離 (CMMD 0.16 対 0.55) を達成し、IPF で正規化された人口統計評価により、既存の方法に存在する部分母集団のバイアスが大幅に軽減されることが示されています。人間による評価では、ジェネリック生成に対して 79% の勝率、すべての微調整されたベースラインに対して 58 ~ 65% の勝率が確認されています。

原文 (English)

ZIPP:Zero-shot Image Personalization from Personas

Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate aesthetics rather than individual taste. Human preferences are pluralistic: one user favoring muted, nostalgic portraits may prefer vibrant street photography, while another gravitates toward dreamy film aesthetics. Existing methods require dense interaction histories or per-user fine-tuning, failing in cold-start settings and collapsing context-dependent preferences into a static representation. We introduce zero-shot image personalization from personas (ZIPP), which conditions image generation on natural-language personas (concise descriptors of a user's identity and aesthetic sensibilities) without any user-specific data or weight updates. ZIPP uses an LLM to rewrite prompts from the perspective of a given persona, steering diffusion models toward personalized outputs. To mine personas at scale, we train an inductive Graph Attention Network over a 22M-user Reddit interaction graph with dual contrastive objectives aligning graph structure with visual behavior, then verbalize learned representations into natural-language personas via an MLLM. We introduce ZIPBench, the first zero-shot personalization benchmark with 1.5K users, graph-mined personas, and 40K generated images. Across four benchmarks and 14 LLMs spanning five model families, persona conditioning yields consistent gains (13-20%), with frontier models benefiting most. In the few-shot setting, ZIPP matches or exceeds fine-tuned baselines trained on 100+ examples per user. ZIPP achieves the lowest preference distributional divergence (CMMD 0.16 vs. 0.55), and IPF-normalized demographic evaluation shows it substantially reduces subpopulation bias present in existing methods. Human evaluation confirms a 79% win rate over generic generation and 58-65% over all fine-tuned baselines.

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

順序は重要: プロキシガイドによる LLM 進化によるマクロ配置シーケンスの隠れた影響を明らかにする

マクロの配置は、最新のチップの物理設計における基本的なステップであり、高次元の組み合わせ最適化問題の解の品質を決定する上で重要な役割を果たします。空間座標決定のための機械学習の最近の進歩にもかかわらず、配置シーケンスの時間的次元は依然として静的ヒューリスティックによって主に支配されています。この研究では、配置シーケンスが単なる前処理ステップではなく、最適化における決定的な要素であることを実証します。最適化に及ばない初期の決定は、解空間を制約する不可逆的なドミノ効果を引き起こします。この未踏の次元を活用するために、マクロ配置順序戦略を自動的に発見するためのプロキシガイドによる LLM 進化フレームワークである \textbf{OrderPlace} を提案します。 OrderPlace は、エリアベースや接続ベースの順序付けなどの手動で作成されたヒューリスティックに依存するのではなく、静的なスコアリング メトリクスから動的な物理学にヒントを得たメカニズムに至るまで、より広範なコード レベルのポリシーを探索します。シーケンス評価の法外なコストを軽減するために、決定論的な貪欲プローブを使用して候補を効率的にフィルタリングする軽量のプロキシ評価メカニズムを導入します。標準 ISPD 2005 ベンチマークの実験結果は、OrderPlace が新しい順序付け戦略を発見したことを示しています。 WireMask-EA および最先端のメソッド EGPlace と比較して、OrderPlace はワイヤ長をそれぞれ 34.04\% および 14.08\% 削減します。

原文 (English)

Order Matters: Unveiling the Hidden Impact of Macro Placement Sequences via Proxy-Guided LLM Evolution

Macro placement is a fundamental step in modern chip physical design, playing a crucial role in determining the solution quality of high-dimensional combinatorial optimization problems. Despite recent advancements in machine learning for spatial coordinate determination, the temporal dimension of placement sequencing remains largely governed by static heuristics. In this work, we demonstrate that the placement sequence is not merely a preprocessing step but a decisive factor in optimization, where suboptimal early decisions trigger irreversible domino effects that constrain the solution space. To harness this unexplored dimension, we propose \textbf{OrderPlace}, a proxy-guided LLM evolution framework for automatically discovering macro placement order strategies. Instead of relying on manually crafted heuristics such as area- or connectivity-based ordering, OrderPlace explores a broader space of code-level policies, ranging from static scoring metrics to dynamic physics-inspired mechanisms. To mitigate the prohibitive cost of evaluating sequences, we introduce a lightweight proxy evaluation mechanism that efficiently filters candidates using a deterministic greedy probe. Experimental results on the standard ISPD 2005 benchmarks demonstrate that OrderPlace discovers novel ordering strategies. Compared with WireMask-EA and the state-of-the-art method EGPlace, OrderPlace reduces wirelength by 34.04\% and 14.08\%, respectively.

2026-06-09 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

トークンが取得されない: AI エージェント出力のサンプリング、状態、および変動性

Agentic AI システムは、実行ごとに異なる動作をする可能性があります。同じリクエストでも、異なる計画、異なるツール呼び出し、異なるコード編集、または異なる最終応答が生成される場合があります。このような変動は、しばしば混同されるいくつかの層から生じます。基礎モデルは大規模な事前トレーニング済みモデルであり、通常は多くの下流タスクに適応でき、入力コンテキストを出力に対する予測にマッピングします。現在のエージェントの多くでは、そのモデルは、計画、ツールの呼び出し、結果の観察、状態の更新を行うオーケストレーション ループに組み込まれています。このようなシステムにおけるばらつきの明示的な固有の原因の 1 つは、トークンの生成です。モデルは、考えられる次のトークンのスコアを計算し、スコアは確率に変換され、デコーダーは擬似乱数ジェネレーターを使用してトークンをサンプリングすることがあります。サンプリングされたトークンの小さな違いは、異なるツール呼び出し、コード パス、検索クエリ、またはエージェントの状態に上向きに伝播する可能性があります。変動のその他の原因は、環境の変化、ライブデータ、サービス提供インフラストラクチャ、バッチ効果、数値の詳細など、トークン サンプリングに付随するものです。これらの層を分離することで、原稿は、エージェント AI システムを確率的と呼ぶことが何を意味するのか、そのような変動が一致した条件下で再現できる場合、そしてなぜ決定論的な実行が展開された設定で同一の動作を暗示する必要がないのかを明確にしています。

原文 (English)

The Token Not Taken: Sampling, State, and the Variability of AI Agent Outputs

Agentic AI systems can behave differently across runs: the same request may produce a different plan, a different tool call, a different code edit, or a different final answer. Such variability arises from several layers that are often conflated. A foundation model is a large pretrained model, usually adaptable to many downstream tasks, that maps an input context to predictions over outputs. In many current agents, that model is embedded in an orchestration loop that plans, calls tools, observes results, and updates state. One explicit intrinsic source of variability in such systems is token generation: the model computes scores over possible next tokens, the scores are converted into probabilities, and a decoder may sample tokens using a pseudo-random number generator. A small sampled token difference can then propagate upward into a different tool call, code path, search query, or agent state. Other sources of variability are extrinsic to token sampling, including changing environments, live data, serving infrastructure, batch effects, and numerical details. By separating these layers, the manuscript clarifies what it means to call agentic AI systems stochastic, when such variability can be reproduced under matched conditions, and why deterministic execution need not imply identical behavior in deployed settings.

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Reliable to Expressive: A Curriculum for Rubric-Following Safety Judges

Safety judges are increasingly deployed to evaluate model outputs against evolving criteria, yet recent meta-evaluation work shows they rem…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Capacity, Not Format: Rethinking Structured Reasoning Failures

Prior work treats structured output as a reasoning tax, but this framing is incomplete: the cost of formatting depends strongly on a model'…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

PRISM: Recovering Instruction Sets from Language Model Activations

As LLMs are deployed as agents, reliable monitoring requires knowing not only what they output, but which instructions are steering their b…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

Frequency-based Constrained Sampling for Interval Patterns

Output space pattern sampling is a powerful alternative to exhaustive pattern mining for exploring large pattern spaces, as it enables user…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体ビジネス/資金調達研究/論文

Multi-Turn Evaluation of Deep Research Agents Under Process-Level Feedback

Existing benchmarks for deep research agents (DRAs) assess only single-shot outputs, ignoring a key question: can DRAs improve their report…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

Beware of GeeksBearing Gifts: Building True EU Frontier AI Sovereignty

Frontier artificial intelligence is reshaping all aspects of society, from economic output or military capability to democratic institution…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Liberating LLM Capabilities in Full-Duplex Speech Models

Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verba…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

EssentialGIN: a new approach for gene essentiality prediction based on graph isomorphism neural networks

Background: Prediction of essential genes (proteins), is a basic and challenging problem but at the same time very costly and time-consumin…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

SLMJury: Can Small Language Models Judge as Well as Large Ones?

Large language models (LLMs) are widely used as judges for evaluating model outputs, but their high cost, latency, and opacity limit scalab…

2026-06-09 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for th…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ハードウェア/半導体

TeamHerald@CHIPSAL 2026: Hate Speech Detection and Sentiment Analysis of Nepali Memes using Transformer-based Architectures and Ensemble Learning

The analysis of internet memes in the Nepali language is complicated by frequent code-mixing and a lack of established baseline resources.…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Evaluating AI Investment Strategies

We study the problem of auditing a black-box algorithmic decision-maker from observable inputs and outputs alone. Our main result is an exa…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CARE: A Conformal Safety Layer for Medical Summarization

Large language models (LLMs) are increasingly used for medical summarization, but their outputs can omit medically important information an…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Conventional LLMs keep the full KV cache loaded during decoding, causing a severe GPU memory bottleneck for ultra-long context serving. In…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

Hybridizing Equilibrium Propagation with Ising Machines for Efficient Energy-Based Learning

The rapid evolution of artificial intelligence has led to substantial advances in deep neural networks. Nonetheless, conventional GPU-based…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

Resource-aware Computation-Communication Overlap for multi-GPU ML Workloads

The rapid growth of large-scale machine learning (ML) has made distributed training across multiple GPUs a fundamental component of modern…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

Internalizing Geometric Law: Learning from Solver Residuals for Precision-Critical Generation

Large Language Models frequently hallucinate in precision-critical domains such as technical diagramming and mechanical design, where outpu…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

Closure-Validated Circuit Discovery in Attention Heads: Co-activation Proposes, Ablation Disposes

Interpretability increasingly treats groups of components, not individual units, as the basic object, and proposes to find them by clusteri…

2026-06-09 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

IRAM-Omega-Q: A Computational Framework for Uncertainty Regulation in Adaptive Agents

Adaptive agents operating under uncertainty must do more than optimize task outputs: they must maintain a workable internal state under noi…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

TAO: Tolerance-Aware Optimistic Verification for Floating-Point Neural Networks

Neural networks increasingly run on hardware outside the user's control (cloud GPUs, inference marketplaces). Yet ML-as-a-Service reveals l…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Correcting Mean Bias in Text Embeddings: A Refined Renormalization with Training-Free Improvements on MMTEB

We find that current sentence-embedding models produce outputs with a consistent bias: every embedding $e$ decomposes as $\tilde e + \mu$,…

2026-06-09 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

検出から回復まで: 504 GPU による LLM 事前トレーニングでの運用分析

大規模な AI トレーニングは現在、基本的に分散システムの問題となっており、ハードウェア障害はまれな例外ではなく、日常的な動作条件となっています。しかし、本番トレーニングクラスターからの公的運用証拠は依然として不足しています。この技術レポートは、55 日間の Prometheus 時系列データと 224 のマルチノード トレーニング セッションをカバーする 73 日間の運用ログを使用した、63 ノードの NVIDIA B200 実稼働クラスター (504 GPU) の実証分析を示しています。このクラスターは、5 者 (SKT、Upstage、Lablup、NVIDIA Korea、VAST Data) が統合された監視パイプラインを共有する組織間環境内で動作します。この配置により、2 ~ 4 ノード規模では現れなかった 60 ノード規模のストレージ I/O ボトルネックを共同診断することが可能になりました。これは、単一チームだけでは分離できない実稼働規模の現象です。数か月にわたる事前トレーニング キャンペーンを利用して、3 つの定量分析を実行し、4 つの結果を得ました。まず、751 の Prometheus メトリクスと 10 件の XID で特定された GPU 障害に関する統計分析により、1 日あたりの誤検知が約 0.84 件で 10/10 の検出率 (XID 前は 2/10) を達成しました。複数の信号検出戦略を動機付ける、複数の障害タイプにわたって一貫して支配的な単一の指標はありません。次に、GPU VRAM から NFS パスに沿った 523 のチェックポイント イベントのプロファイリングにより、「帯域幅のパラドックス」(200 Gbps RoCE の 1.4 ~ 10.4% の使用率) が 128 スロットの NFS RPC レイヤーの飽和に起因していることがわかります。 3 番目に、マルチノード障害の応答では、集中した除外 (63 ノード中上位 3 ノードがすべての除外の 50% 以上を占める) と、12 チェーン (試行 73 回) にわたる自動再試行チェーンの成功率が 33.3% で、手動回復率 12.5% の 2.7 倍であることが示されています。再試行間隔の中央値は 11 分 (IQR 10-11) です。すべての分析は実稼働インフラストラクチャに基づいており、セッション レベルのワークロード管理、GPU 中心のスケジューリング、および統合された可観測性を提供します。

原文 (English)

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs

Large-scale AI training is now fundamentally a distributed systems problem, and hardware failures have become routine operating conditions rather than rare exceptions. Public operational evidence from production training clusters, however, remains scarce. This technical report presents an empirical analysis of a 63-node NVIDIA B200 production cluster (504 GPUs), using 55 days of Prometheus time-series data and 73 days of operational logs covering 224 multi-node training sessions. The cluster operates within a cross-organizational environment in which five parties (SKT, Upstage, Lablup, NVIDIA Korea, and VAST Data) share a unified monitoring pipeline. This arrangement enabled joint diagnosis of a 60-node-scale storage I/O bottleneck that did not appear at 2-4-node scale, a production-scale phenomenon no single team could isolate alone. Drawing on a months-long pre-training campaign, we perform three quantitative analyses yielding four findings. First, statistical analysis over 751 Prometheus metrics and 10 XID-identified GPU failures achieves a 10/10 detection rate (2/10 pre-XID) at ~0.84 false positives per day. No single metric is consistently dominant across failure types, motivating a multi-signal detection strategy. Second, profiling 523 checkpoint events along the GPU VRAM to NFS path attributes the "bandwidth paradox" (1.4-10.4% utilization of 200 Gbps RoCE) to saturation of the 128-slot NFS RPC layer. Third, multi-node failure response shows concentrated exclusions (top 3 of 63 nodes account for >50% of all exclusions) and an auto-retry chain success rate of 33.3% over 12 chains (73 attempts), 2.7x the 12.5% manual recovery rate; the median retry interval is 11 min (IQR 10-11). All analyses are grounded in production infrastructure providing session-level workload management, GPU-centric scheduling, and unified observability.

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models

Unstructured sparsity is now natively accelerated by recent GPU kernels and dataflow hardware, shifting the bottleneck from inference execu…

2026-06-09 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

Crazyflow: An Accurate, GPU-Accelerated, Differentiable Drone Simulator in JAX

High-quality, large-scale synthetic data from simulations is becoming a cornerstone for pushing the capabilities of robot algorithms. While…

2026-06-09 13:00 JSTarXiv cs.AIハードウェア/半導体

OPRD: On-Policy Representation Distillation

On-policy distillation (OPD) supervises the student only in output space by matching next-token probabilities. This output-only paradigm ha…

2026-06-08 13:00 JSTarXiv cs.AIハードウェア/半導体

Accelerated Fourier SAT (AFSAT): GPU ベースの対称擬似ブール SAT ソルバーを完全に実現

我々は、連続局所探索 (CLS) に基づく擬似ブール充足可能性の GPU 高速化ソルバーである Accelerated Fourier SAT (AFSAT) を紹介します。 AFSAT は、概念実証アプローチである FastFourierSAT を、単一の問題インスタンス内の対称制約のタイプと長さの異種混合をサポートする完全に設計されたソルバーに実現します。 JAX コンパイラーを使用する AFSAT は、純粋関数合成、自動ベクトル化、自動微分、ジャストインタイム (JIT) コンパイルを活用して、候補割り当てのバッチ全体で大規模並列 CLS を実行します。概念実証と比較して、数値安定性、実行時パフォーマンス、メモリ効率が大幅に向上していることを実証します。私たちは、自動並列化とコンパクトな表現を活用するだけでなく、メモリ レイテンシや浮動小数点表現から生じるさまざまな制限を特定して対処することでこれを実現します。浮動小数点の固有の表現および安定性の制限は、調整された離散フーリエ変換の実装によって部分的に対処されます。 JAX アレイ シャーディングを介して複数のアクセラレータにスケーリングする場合、ほぼ線形のスループットを実現します。

原文 (English)

Accelerated Fourier SAT (AFSAT): Fully Realising a GPU-based Symmetric Pseudo-Boolean SAT Solver

We present Accelerated Fourier SAT (AFSAT), a GPU-accelerated solver for pseudo-Boolean satisfiability based on continuous local search (CLS). AFSAT realises the proof-of-concept approach, FastFourierSAT, into a fully-engineered solver supporting any heterogeneous mixture of symmetric constraint types and lengths within a single problem instance. Using the JAX compiler, AFSAT leverages pure function composition, automatic vectorisation, automatic differentiation, and just-in-time (JIT) compilation to perform massively parallel CLS across batches of candidate assignments. We demonstrate substantially improved numerical stability, runtime performance, and memory efficiency over the proof-of-concept. We achieve this by way of identifying and addressing various limitations that arise from memory latency and floating-point representation, as well as leveraging automatic parallelisation and compact representations. The inherent representational and stability limitations of floating point are partially addressed by a tailored discrete Fourier transform implementation. We achieve near-linear throughput when scaling to multiple accelerators via JAX array sharding.

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

特徴付けてから蒸留: 大規模な出力空間での機械的推論

最新の推論モデルは、数十万から数百万の候補ラベルから関連するオプションの少数のセットを選択する必要がある、困難なマルチラベル タスクに対して、驚くほど強力なゼロショット パフォーマンスを提供します。私たちは、彼らがこれをどのように機構的に達成するのかを調査します。私たちは推論を 2 段階のプロセスとして特徴付けます。つまり、候補の大まかな「候補リスト」に続いて、結果のセットに対するきめの細かい推論が行われます。私たちは、さまざまなデータセットにわたって、これらのステップが分離可能であり、補完的であるという証拠を提供します。この特性評価を使用して、当社は標準的な蒸留を常に上回る機械的蒸留戦略を開発します。

原文 (English)

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces

Modern reasoning models offer surprisingly strong zero-shot performance on challenging multi-label tasks that require selecting a small set of relevant options from hundreds of thousands to millions of candidate labels. We investigate how they achieve this mechanistically. We characterize reasoning as a two-phase process: A broad "shortlisting" of candidates followed by fine-grained reasoning over the resulting set. We provide evidence across a range of datasets that these steps can be isolated and are complementary. Using this characterization, we develop a mechanistic distillation strategy that consistently outperforms standard distillation.

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Beyond Output Matching: Preserving Internal Geometry in NVFP4 LLM Distillation

Demand for low-precision inference, including NVFP4-based approaches, has grown as large language models are increasingly deployed in laten…

2026-06-08 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Extracting Recurring Vulnerabilities from Black-Box LLM-Generated Software

LLMs are increasingly used for code generation, but their outputs often follow recurring templates that can induce predictable vulnerabilit…

2026-06-08 13:00 JSTarXiv cs.AIハードウェア/半導体

VeriHGN: Heterogeneous Graph-Based Congestion Prediction for Chip Layout Verification

As Very Large Scale Integration (VLSI) designs continue to scale in size and complexity, layout verification has become a central challenge…

2026-06-08 13:00 JSTarXiv cs.AIハードウェア/半導体

Google Cloud TPU での Gemma 4 31B の微調整と提供: GPU ベースラインとの技術比較

TPU ハードウェア上で Google の Gemma 4 31B モデルを微調整して提供する最初のエンドツーエンドのデモンストレーションを紹介し、大規模な言語モデルの適応に関する TPU プラットフォームと GPU プラットフォームの実証的比較を提供します。トレーニングには Google TPU v5p-8、推論には TPU v6e-8 (Trillium) 上の LoRA を使用して、PyTorch、HuggingFace TRL、FSDP 上に構築された GPU ネイティブのトレーニング レシピを JAX + Tunix/Qwix スタックに移植するために必要なコードレベルの適応の完全なセットを文書化します。これらの適応は、メッシュ構成、LoRA モジュールの命名規則、シャーディング アノテーションの修正、勾配チェックポイント設定、データ パイプラインの再構築、およびカスタム Orbax からセーフテンソルへのチェックポイント マージ手順に及びます。推論のために、v6e-8 で Gemma 4 を提供するために必要な vLLM-TPU Docker セットアップを詳しく説明し、その結果生じるレイテンシとスループット プロファイルの特徴を説明します。同一のハイパーパラメータの下での 2xH100 GPU ベースラインと比較して、TPU トレーニングは 2.12 倍のコストで 1.61 倍の速度で完了します。推論スループットはプラットフォーム全体で 3% 以内ですが、TPU は最初のトークンまでの時間を 2 倍短縮します (235 ミリ秒対 475 ミリ秒)。これらを合計すると、TPU 構成は、代表的なトレインとサービスのワークロードでは 1.82 倍安くなります。私たちの取り組みは、オープン ツール エコシステムの重大なギャップを取り除き、TPU インフラストラクチャ上で Gemma 4 を導入するための再現可能で本番環境にすぐに使えるレシピを実務者に提供します。

原文 (English)

Fine-Tuning and Serving Gemma 4 31B on Google Cloud TPU: A Technical Comparison with GPU Baselines

We present the first end-to-end demonstration of fine-tuning and serving Google's Gemma 4 31B model on TPU hardware, providing an empirical comparison of TPU and GPU platforms for large language model adaptation. Using LoRA on a Google TPU v5p-8 for training and TPU v6e-8 (Trillium) for inference, we document the full set of code-level adaptations required to port a GPU-native training recipe - built on PyTorch, HuggingFace TRL, and FSDP - to the JAX + Tunix/Qwix stack. These adaptations span mesh configuration, LoRA module naming conventions, sharding annotation corrections, gradient checkpoint, data pipeline restructuring, and a custom Orbax-to-safetensor checkpoint merging procedure. For inference, we detail the vLLM-TPU Docker setup necessary to serve Gemma 4 on v6e-8 and characterize the resulting latency and throughput profile. Compared with a similar-costing 2xH100 GPU baseline under identical hyperparameters, TPU training completes 1.61x faster at 2.12x lower cost. For inference, we cover the vLLM-TPU Docker setup required to serve Gemma 4 on v6e-8 and explain the observed latency and throughput characteristics across a QPS sweep spanning 512 to 16k input tokens. Across both workloads we compare performance and cost against a 2xH100 GPU baseline running identical hyperparameters. The TPU completes training 1.61x faster at 2.12x lower cost. For inference, TPU v6e-8 matches GPU at short context (<=2048 tokens) and decisively outperforms at long context: 66% higher throughput and 23.6x faster TTFT at 4096-token inputs (61 ms vs 1,443 ms at QPS=4). Our work removes a critical gap in the open tooling ecosystem and provides practitioners with a recipe for Gemma 4 Dense 31B deployment on the TPU infrastructure.

2026-06-05 13:00 JSTarXiv cs.AIハードウェア/半導体

ASP ベースのコンプライアンス推論のための規範的な中間表現

我々は、ASP ベースのコンプライアンス推論のためのモーダル化出力規範中間表現である MONIR を提案します。そのコア フラグメントには段階的な操作セマンティクスがあり、MONIR-ASP は外部関数、一時的なルール、および安定したモデル推論のための実行可能なコンパイルと拡張機能を提供します。 LLM 支援パイプラインを使用して、中国の ADAS 規制と標準に関するフレームワークをインスタンス化します。実験では、抽出品質と、モジュール式および増分 ASP 解決の効率を評価します。

原文 (English)

A Normative Intermediate Representation for ASP-Based Compliance Reasoning

We propose MONIR, a Modalized-Output Normative Intermediate Representation for ASP-based compliance reasoning. Its core fragment has a staged operational semantics, while MONIR-ASP provides an executable compilation and extensions for external functions, temporal rules, and stable-model reasoning. We instantiate the framework on Chinese ADAS regulations and standards with an LLM-assisted pipeline. Experiments evaluate extraction quality and the efficiency of modular and incremental ASP solving.

2026-06-05 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

DiffAero: 効率的なクアドローター ポリシー学習のための GPU アクセラレーションによる微分可能シミュレーション フレームワーク

このレターでは、効率的なクワッドローター制御ポリシー学習のために設計された、軽量で GPU アクセラレーションを備えた完全微分可能なシミュレーション フレームワークである DiffAero を紹介します。 DiffAero は、環境レベルとエージェント レベルの両方の並列処理をサポートし、複数のダイナミクス モデル、カスタマイズ可能なセンサー スタック (IMU、深度カメラ、LiDAR)、および多様な飛行タスクを統合された GPU ネイティブのトレーニング インターフェイス内に統合します。 DiffAero は、GPU 上で物理とレンダリングの両方を完全に並列化することで、CPU と GPU 間のデータ転送のボトルネックを排除し、シミュレーションのスループットを桁違いに向上させます。既存のシミュレータとは対照的に、DiffAero は高性能シミュレーションを提供するだけでなく、微分可能なハイブリッド学習アルゴリズムを探索するための研究プラットフォームとしても機能します。広範なベンチマークと実際の飛行実験により、DiffAero とハイブリッド学習アルゴリズムを組み合わせることで、消費者グレードのハードウェアで堅牢な飛行ポリシーを数時間で学習できることが実証されました。コードは https://github.com/flyingbitac/diffaero で入手できます。

原文 (English)

DiffAero: A GPU-Accelerated Differentiable Simulation Framework for Efficient Quadrotor Policy Learning

This letter introduces DiffAero, a lightweight, GPU-accelerated, and fully differentiable simulation framework designed for efficient quadrotor control policy learning. DiffAero supports both environment-level and agent-level parallelism and integrates multiple dynamics models, customizable sensor stacks (IMU, depth camera, and LiDAR), and diverse flight tasks within a unified, GPU-native training interface. By fully parallelizing both physics and rendering on the GPU, DiffAero eliminates CPU-GPU data transfer bottlenecks and delivers orders-of-magnitude improvements in simulation throughput. In contrast to existing simulators, DiffAero not only provides high-performance simulation but also serves as a research platform for exploring differentiable and hybrid learning algorithms. Extensive benchmarks and real-world flight experiments demonstrate that DiffAero and hybrid learning algorithms combined can learn robust flight policies in hours on consumer-grade hardware. The code is available at https://github.com/flyingbitac/diffaero.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CodegenBench: LLM はアーキテクチャ全体で効率的なコードを記述できますか?

大規模言語モデル (LLM) は、汎用プログラミングや GPU アクセラレーション環境 (PyTorch、CUDA など) のコード生成タスクで広範囲に評価されてきましたが、多様なアーキテクチャにわたる CPU 指向のハイパフォーマンス コンピューティング (HPC) における LLM の機能はまだ十分に解明されていません。このギャップを埋めるために、x86_64、Sunway、Kunpeng の 3 つの異なるハードウェア プラットフォームにわたる効率的な並列コードの生成を評価するように設計された包括的なベンチマーク スイートである CodegenBench を紹介します。私たちのベンチマークは、基本的なベースラインを確立する 106 個の標準基本線形代数サブプログラム (BLAS) ルーチンと、独自のスーパーコンピューティング アーキテクチャ (LeetSunway および LeetKunpeng) のそれぞれに適合した 20 個の特殊な計算カーネルで構成されています。私たちの広範な評価により、最先端の LLM は x86_64 のようなユビキタス アーキテクチャ向けに最適化されたコードを生成できる一方で、公開ドキュメントやトレーニング データが限られたドメイン固有のアーキテクチャでは大幅なパフォーマンスの低下を示し、クロスプラットフォームの一般化における重大な制限が浮き彫りになったことが明らかになりました。さらに、実装の長さやタスクの複雑さなど、コードの品質に影響を与える要因を分析したところ、現在の LLM は、簡潔なコード スニペットを必要とする中程度に難しい問題に対して最も効果的であることが示されています。私たちは、LLM 主導の高性能コード生成における将来の研究を促進するために、データセットと自動評価インフラストラクチャをオープンソースにしています。リソースは https://anonymous.4open.science/r/CodegenBench-EDE1/ および https://anonymous.4open.science/r/CodegenBenchDataset-2551 で利用できます。

原文 (English)

CodegenBench: Can LLMs Write Efficient Code Across Architectures?

While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.g., PyTorch, CUDA), their capabilities in CPU-oriented high-performance computing (HPC) across diverse architectures remain underexplored. To bridge this gap, we introduce CodegenBench, a comprehensive benchmark suite designed to evaluate the generation of efficient parallel code across three distinct hardware platforms: x86_64, Sunway, and Kunpeng. Our benchmark comprises 106 standard Basic Linear Algebra Subprograms (BLAS) routines establishing a fundamental baseline, alongside 20 specialized computational kernels adapted for each of the unique supercomputing architectures (LeetSunway and LeetKunpeng). Our extensive evaluation reveals that while state-of-the-art LLMs can generate optimized code for ubiquitous architectures like x86_64, they exhibit significant performance degradation on domain-specific architectures with limited public documentation and training data, highlighting critical limitations in cross-platform generalization. Furthermore, our analysis of factors influencing code quality such as implementation length and task complexity indicates that current LLMs are most effective for moderately difficult problems requiring concise code snippets. We open-source our dataset and automated evaluation infrastructure to facilitate future research in LLM-driven high-performance code generation. The resources are available at https://anonymous.4open.science/r/CodegenBench-EDE1/ and https://anonymous.4open.science/r/CodegenBenchDataset-2551.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Caught in the Act(ivation): LLM エージェントによる資格情報漏洩の事前出力およびマルチターン検出に向けて

LLM エージェントは多くの場合、機密認証情報を信頼できない取得コンテンツと同じコンテキスト ウィンドウに配置し、認証情報の漏洩を誘発する間接的なプロンプト インジェクションの直接パスを作成します。私たちは、3 つの相補的な防御を通じてこの障害モードを研究します。まず、出力トークンが発行される前に、アクティベーション プローブが資格情報へのアクセスを検出できるかどうかを尋ねます。次に、形式固有の文字モデルからハニートークンを構築し、分割等角予測で検出を調整します。 3 番目に、複数ターンにわたる漏洩を累積的な情報フロー問題として扱い、会話ターン全体での推定漏洩予算を追跡します。オープンウェイト モデルの制御された実験では、アクティベーション機能により、ホールドアウト エンコーディング変換下を含め、無害なプロンプトと認証情報を求めるプロンプトが高精度で分離されます。小規模な合成マルチターン スイートでは、累積アカウンティングにより、ターンごとの検出器が見逃した攻撃が検出されます。これらの結果は暫定的なものです。マルチターン ベンチマークは社内で小規模なものであり、アクティブ化方法にはホワイト ボックス アクセスが必要であり、情報推定ツールは正式な上限ではなく実用的なシグナルを提供します。それでも、この結果は、資格情報の漏洩防御には、テキストレベルの出力フィルターのみに依存するのではなく、出力前の監視、調整されたカナリア検出、および一時的な漏洩アカウンティングを組み合わせる必要があることを示唆しています。

原文 (English)

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents

LLM agents often place sensitive credentials in the same context window as untrusted retrieved content, creating a direct path for indirect prompt injection to induce credential exfiltration. We study this failure mode through three complementary defenses. First, we ask whether activation probes can detect credential access before output tokens are emitted. Second, we construct honeytokens from format-specific character models and calibrate detection with split conformal prediction. Third, we treat multi-turn exfiltration as a cumulative information-flow problem and track an estimated leakage budget across conversation turns. In controlled experiments on open-weight models, activation features separate benign and credential-seeking prompts with high accuracy, including under held-out encoding transformations. In a small synthetic multi-turn suite, cumulative accounting detects attacks that per-turn detectors miss. These results are preliminary: the multi-turn benchmark is in-house and small, the activation method requires white-box access, and the information estimator provides a practical signal rather than a formal upper bound. Still, the results suggest that credential-exfiltration defenses should combine pre-output monitoring, calibrated canary detection, and temporal leakage accounting rather than relying only on text-level output filters.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

トークンランキングは偽造不可能な言語モデル署名です

言語モデルのパラメータは、ロジット出力に(各モデルに)一意の幾何学的制約を課すことが知られており、これはモデルを識別する署名として機能しますが、API がロジットを配布するときにモデルの最終層パラメータも漏洩します。私たちは、トークンのランキング (確率値ではなく、確率による順序付け) を公開する、より制限的な API を調査し、ランキングも署名を構成することを発見しました。すべてのモデルは、十分な規模の $k$ に対して実行可能な上位 $k$ ランキングの独自のセットを持っています。さらに、同じ実行可能なランキングのセットを持つモデルを見つけることは NP 困難であるため、ランキング署名は最初に知られている (多項式的に) 偽造不可能な署名です。セキュリティの面では、ロジットと同様に、トークンのランキングがすでにモデルの最終層をほぼ盗むのに十分であることがわかりました。ただし、近似が粗すぎて署名を偽造できず、API を十分に小さい $k$ の上位 $k$ トークンに制限することで効果的に対抗できます。モデル署名を提示するために必要な $k$ は一般に、盗用を防ぐために必要な $k$ よりも小さいため、API はモデル パラメーターを漏らすことなく偽造不可能な署名を提示することが可能です。

原文 (English)

Token Rankings are Unforgeable Language Model Signatures

Language model parameters are known to impose unique (to each model) geometric constraints on their logit outputs, which serves as a signature that identifies the model, but also leaks the model's final layer parameters when an API distributes logits. We investigate more restrictive APIs that expose token rankings (i.e., their ordering by probability, but not the probability values) and find that rankings also constitute a signature: every model has a unique set of feasible top-$k$ rankings for sufficiently large $k$. Furthermore, the ranking signature is the first known (polynomially) unforgeable signature, since finding a model with the same set of feasible rankings is NP-hard. On the security front, we find that token rankings are already sufficient to approximately steal the final layer of the model, similar to logits, though the approximation is too coarse to forge the signature, and can be effectively countered by restricting the API to top-$k$ tokens with sufficiently small $k$. Since the top-$k$ required to present the model signature is generally smaller than the $k$ required to prevent stealing, it is possible for an API to present an unforgeable signature without leaking model parameters.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

ルーブリックベースの強化学習における報酬ハッキングの再現、分析、検出

ルーブリックベースの強化学習 (RL) は、LLM-as-a-Judge (LaaJ) を使用して、報酬としてルーブリックに従ってモデルの出力を採点します。ただし、政策モデルは裁判官の潜在的なバイアスを悪用し、報酬のハッキングや非効果的または危険なトレーニング結果につながる可能性があります。現実のルーブリックベースの RL では、このようなハッキング行為は多くの場合微妙であり、複数の裁判官のバイアスと絡み合っているため、分析、検出、軽減することが困難です。このペーパーでは、ルーブリックベースの RL のための制御可能なハッキング環境である CHERRL を紹介します。既知のバイアスを LaaJ に注入することで、CHERRL は報酬ハッキングの安定した再現、報酬の発散の明確な観察、およびハッキングの開始の正確な特定を可能にします。これは、ルーブリック ベースの RL における報酬ハッキングのメカニズムと緩和を研究するためのクリーンな実験テストベッドを提供します。その有用性を実証するために、発見可能性と悪用可能性の観点からさまざまな裁判官のバイアスを分析し、トレーニングログから報酬ハッキングの開始を自動的に検出するためのエージェントベースのシステムを調査します。コードと環境は https://github.com/THUAIS-Lab/CHERRL で公開されています。

原文 (English)

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a controllable hacking environment for rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and precise identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, and explore an agent-based system for automatically detecting reward hacking onset from training logs. The code and environment are publicly available at https://github.com/THUAIS-Lab/CHERRL.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

制約付き適応拒否サンプリング

言語モデル (LM) は、生成された出力が厳密な意味論的または構文上の制約を満たす必要があるアプリケーションで使用されることが増えています。制約付き生成に対する既存のアプローチはさまざまです。貪欲な制約付きデコード方法は、デコード中に有効性を強制しますが、LM の分布を歪めます。一方、リジェクション サンプリング (RS) は忠実度を維持しますが、無効な出力を破棄することで計算を無駄にします。サンプルの有効性と多様性の両方が重要であるプログラム ファジングなどの領域では、両極端が問題となります。我々は、分布歪みを生じさせずに RS のサンプル効率を厳密に改善するアプローチである、制約付き適応除去サンプリング (CARS) を紹介します。 CARS は、制約のない LM サンプリングから始まり、制約違反の継続をトライに記録し、将来の描画から確率質量を差し引くことで、制約に違反する継続を適応的に除外します。この適応的な枝刈りにより、無効であることが証明されたプレフィックスが決して再検討されず、受け入れ率が単調に向上し、結果として得られるサンプルが制約された分布に正確に従うことが保証されます。プログラムのファジングや分子生成など、さまざまな領域の実験において、CARS は一貫して高い効率 (有効サンプルあたりの LM フォワードパスの数で測定) を達成すると同時に、GCD や LM の分布を近似する方法の両方よりも強力なサンプル多様性を生み出します。

原文 (English)

Constrained Adaptive Rejection Sampling

Language Models (LMs) are increasingly used in applications where generated outputs must satisfy strict semantic or syntactic constraints. Existing approaches to constrained generation fall along a spectrum: greedy constrained decoding methods enforce validity during decoding but distort the LM's distribution, while rejection sampling (RS) preserves fidelity but wastes computation by discarding invalid outputs. Both extremes are problematic in domains such as program fuzzing, where both validity and diversity of samples are essential. We present Constrained Adaptive Rejection Sampling (CARS), an approach that strictly improves the sample-efficiency of RS without distributional distortion. CARS begins with unconstrained LM sampling and adaptively rules out constraint-violating continuations by recording them in a trie and subtracting their probability mass from future draws. This adaptive pruning ensures that prefixes proven invalid are never revisited, acceptance rates improve monotonically, and the resulting samples exactly follow the constrained distribution. In experiments on a variety of domains -- e.g., program fuzzing and molecular generation -- CARS consistently achieves higher efficiency -- measured in the number of LM forward passes per valid sample -- while also producing stronger sample diversity than both GCD and methods that approximate the LM's distribution.

2026-06-05 13:00 JSTarXiv cs.AIハードウェア/半導体

モデルを保持した適応丸め

量子化の目標は、出力分布が元のモデルにできるだけ近い圧縮モデルを生成することです。これを容易に行うために、ほとんどの量子化アルゴリズムは、エンドツーエンド エラーの代理として各層の即時アクティブ化エラーを最小限に抑えます。ただし、これは将来のレイヤーの影響を無視するため、プロキシとしては不十分です。この研究では、ネットワークの出力での誤差を直接考慮する適応丸めアルゴリズムである Yet Another Quantization Algorithm (YAQA) を導入します。 YAQA は、量子化アルゴリズムの最初のエンドツーエンド誤差限界に至る一連の理論的結果を紹介します。まず、ヘッセ近似の構造を介して、適応丸めアルゴリズムの収束時間を特徴付けます。次に、エンドツーエンド誤差が真のヘッセ行列に対する近似のコサイン類似度によって制限される可能性があることを示します。これにより、対応する最適に近いヘッシアン スケッチを使用した自然なクロネッカー因数近似が可能になります。 YAQA は GPTQ/LDLQ よりも優れていることが証明されており、経験的にはこれらの方法よりも誤差が $\約 30\%$ 減少します。 YAQA は、量子化を意識したトレーニングよりも低い誤差を実現します。これにより、推論のオーバーヘッドがまったく追加されずに、ダウンストリーム タスクで最先端のパフォーマンスが得られます。

原文 (English)

Model-Preserving Adaptive Rounding

The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible. To do this tractably, most quantization algorithms minimize the immediate activation error of each layer as a proxy for the end-to-end error. However, this ignores the effect of future layers, making it a poor proxy. In this work, we introduce Yet Another Quantization Algorithm (YAQA), an adaptive rounding algorithm that directly considers the error at the network's output. YAQA introduces a series of theoretical results that culminate in the first end-to-end error bounds for quantization algorithms. First, we characterize the convergence time of adaptive rounding algorithms via the structure of their Hessian approximations. We then show that the end-to-end error can be bounded by the approximation's cosine similarity to the true Hessian. This admits a natural Kronecker-factored approximation with corresponding near-optimal Hessian sketches. YAQA is provably better than GPTQ/LDLQ and empirically reduces the error by $\approx 30\%$ over these methods. YAQA even achieves a lower error than quantization aware training. This translates to state of the art performance on downstream tasks, all while adding no inference overhead.

2026-06-05 13:00 JSTarXiv cs.AIハードウェア/半導体

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences

Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed re…

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Consistency Training Can Entrench Misalignment

Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, s…

2026-06-05 13:00 JSTarXiv cs.AIハードウェア/半導体

ASP ベースのコンプライアンス推論のための規範的な中間表現

我々は、ASP ベースのコンプライアンス推論のためのモーダル化出力規範中間表現である MONIR を提案します。そのコア フラグメントには段階的な操作セマンティクスがあり、MONIR-ASP は外部関数、一時的なルール、および安定したモデル推論のための実行可能なコンパイルと拡張機能を提供します。 LLM 支援パイプラインを使用して、中国の ADAS 規制と標準に関するフレームワークをインスタンス化します。実験では、抽出品質と、モジュール式および増分 ASP 解決の効率を評価します。

原文 (English)

A Normative Intermediate Representation for ASP-Based Compliance Reasoning

We propose MONIR, a Modalized-Output Normative Intermediate Representation for ASP-based compliance reasoning. Its core fragment has a staged operational semantics, while MONIR-ASP provides an executable compilation and extensions for external functions, temporal rules, and stable-model reasoning. We instantiate the framework on Chinese ADAS regulations and standards with an LLM-assisted pipeline. Experiments evaluate extraction quality and the efficiency of modular and incremental ASP solving.

2026-06-05 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

DiffAero: 効率的なクアドローター ポリシー学習のための GPU アクセラレーションによる微分可能シミュレーション フレームワーク

このレターでは、効率的なクワッドローター制御ポリシー学習のために設計された、軽量で GPU アクセラレーションを備えた完全微分可能なシミュレーション フレームワークである DiffAero を紹介します。 DiffAero は、環境レベルとエージェント レベルの両方の並列処理をサポートし、複数のダイナミクス モデル、カスタマイズ可能なセンサー スタック (IMU、深度カメラ、LiDAR)、および多様な飛行タスクを統合された GPU ネイティブのトレーニング インターフェイス内に統合します。 DiffAero は、GPU 上で物理とレンダリングの両方を完全に並列化することで、CPU と GPU 間のデータ転送のボトルネックを排除し、シミュレーションのスループットを桁違いに向上させます。既存のシミュレータとは対照的に、DiffAero は高性能シミュレーションを提供するだけでなく、微分可能なハイブリッド学習アルゴリズムを探索するための研究プラットフォームとしても機能します。広範なベンチマークと実際の飛行実験により、DiffAero とハイブリッド学習アルゴリズムを組み合わせることで、消費者グレードのハードウェアで堅牢な飛行ポリシーを数時間で学習できることが実証されました。コードは https://github.com/flyingbitac/diffaero で入手できます。

原文 (English)

DiffAero: A GPU-Accelerated Differentiable Simulation Framework for Efficient Quadrotor Policy Learning

This letter introduces DiffAero, a lightweight, GPU-accelerated, and fully differentiable simulation framework designed for efficient quadrotor control policy learning. DiffAero supports both environment-level and agent-level parallelism and integrates multiple dynamics models, customizable sensor stacks (IMU, depth camera, and LiDAR), and diverse flight tasks within a unified, GPU-native training interface. By fully parallelizing both physics and rendering on the GPU, DiffAero eliminates CPU-GPU data transfer bottlenecks and delivers orders-of-magnitude improvements in simulation throughput. In contrast to existing simulators, DiffAero not only provides high-performance simulation but also serves as a research platform for exploring differentiable and hybrid learning algorithms. Extensive benchmarks and real-world flight experiments demonstrate that DiffAero and hybrid learning algorithms combined can learn robust flight policies in hours on consumer-grade hardware. The code is available at https://github.com/flyingbitac/diffaero.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CodegenBench: LLM はアーキテクチャ全体で効率的なコードを記述できますか?

大規模言語モデル (LLM) は、汎用プログラミングや GPU アクセラレーション環境 (PyTorch、CUDA など) のコード生成タスクで広範囲に評価されてきましたが、多様なアーキテクチャにわたる CPU 指向のハイパフォーマンス コンピューティング (HPC) における LLM の機能はまだ十分に解明されていません。このギャップを埋めるために、x86_64、Sunway、Kunpeng の 3 つの異なるハードウェア プラットフォームにわたる効率的な並列コードの生成を評価するように設計された包括的なベンチマーク スイートである CodegenBench を紹介します。私たちのベンチマークは、基本的なベースラインを確立する 106 個の標準基本線形代数サブプログラム (BLAS) ルーチンと、独自のスーパーコンピューティング アーキテクチャ (LeetSunway および LeetKunpeng) のそれぞれに適合した 20 個の特殊な計算カーネルで構成されています。私たちの広範な評価により、最先端の LLM は x86_64 のようなユビキタス アーキテクチャ向けに最適化されたコードを生成できる一方で、公開ドキュメントやトレーニング データが限られたドメイン固有のアーキテクチャでは大幅なパフォーマンスの低下を示し、クロスプラットフォームの一般化における重大な制限が浮き彫りになったことが明らかになりました。さらに、実装の長さやタスクの複雑さなど、コードの品質に影響を与える要因を分析したところ、現在の LLM は、簡潔なコード スニペットを必要とする中程度に難しい問題に対して最も効果的であることが示されています。私たちは、LLM 主導の高性能コード生成における将来の研究を促進するために、データセットと自動評価インフラストラクチャをオープンソースにしています。リソースは https://anonymous.4open.science/r/CodegenBench-EDE1/ および https://anonymous.4open.science/r/CodegenBenchDataset-2551 で利用できます。

原文 (English)

CodegenBench: Can LLMs Write Efficient Code Across Architectures?

While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.g., PyTorch, CUDA), their capabilities in CPU-oriented high-performance computing (HPC) across diverse architectures remain underexplored. To bridge this gap, we introduce CodegenBench, a comprehensive benchmark suite designed to evaluate the generation of efficient parallel code across three distinct hardware platforms: x86_64, Sunway, and Kunpeng. Our benchmark comprises 106 standard Basic Linear Algebra Subprograms (BLAS) routines establishing a fundamental baseline, alongside 20 specialized computational kernels adapted for each of the unique supercomputing architectures (LeetSunway and LeetKunpeng). Our extensive evaluation reveals that while state-of-the-art LLMs can generate optimized code for ubiquitous architectures like x86_64, they exhibit significant performance degradation on domain-specific architectures with limited public documentation and training data, highlighting critical limitations in cross-platform generalization. Furthermore, our analysis of factors influencing code quality such as implementation length and task complexity indicates that current LLMs are most effective for moderately difficult problems requiring concise code snippets. We open-source our dataset and automated evaluation infrastructure to facilitate future research in LLM-driven high-performance code generation. The resources are available at https://anonymous.4open.science/r/CodegenBench-EDE1/ and https://anonymous.4open.science/r/CodegenBenchDataset-2551.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Caught in the Act(ivation): LLM エージェントによる資格情報漏洩の事前出力およびマルチターン検出に向けて

LLM エージェントは多くの場合、機密認証情報を信頼できない取得コンテンツと同じコンテキスト ウィンドウに配置し、認証情報の漏洩を誘発する間接的なプロンプト インジェクションの直接パスを作成します。私たちは、3 つの相補的な防御を通じてこの障害モードを研究します。まず、出力トークンが発行される前に、アクティベーション プローブが資格情報へのアクセスを検出できるかどうかを尋ねます。次に、形式固有の文字モデルからハニートークンを構築し、分割等角予測で検出を調整します。 3 番目に、複数ターンにわたる漏洩を累積的な情報フロー問題として扱い、会話ターン全体での推定漏洩予算を追跡します。オープンウェイト モデルの制御された実験では、アクティベーション機能により、ホールドアウト エンコーディング変換下を含め、無害なプロンプトと認証情報を求めるプロンプトが高精度で分離されます。小規模な合成マルチターン スイートでは、累積アカウンティングにより、ターンごとの検出器が見逃した攻撃が検出されます。これらの結果は暫定的なものです。マルチターン ベンチマークは社内で小規模なものであり、アクティブ化方法にはホワイト ボックス アクセスが必要であり、情報推定ツールは正式な上限ではなく実用的なシグナルを提供します。それでも、この結果は、資格情報の漏洩防御には、テキストレベルの出力フィルターのみに依存するのではなく、出力前の監視、調整されたカナリア検出、および一時的な漏洩アカウンティングを組み合わせる必要があることを示唆しています。

原文 (English)

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents

LLM agents often place sensitive credentials in the same context window as untrusted retrieved content, creating a direct path for indirect prompt injection to induce credential exfiltration. We study this failure mode through three complementary defenses. First, we ask whether activation probes can detect credential access before output tokens are emitted. Second, we construct honeytokens from format-specific character models and calibrate detection with split conformal prediction. Third, we treat multi-turn exfiltration as a cumulative information-flow problem and track an estimated leakage budget across conversation turns. In controlled experiments on open-weight models, activation features separate benign and credential-seeking prompts with high accuracy, including under held-out encoding transformations. In a small synthetic multi-turn suite, cumulative accounting detects attacks that per-turn detectors miss. These results are preliminary: the multi-turn benchmark is in-house and small, the activation method requires white-box access, and the information estimator provides a practical signal rather than a formal upper bound. Still, the results suggest that credential-exfiltration defenses should combine pre-output monitoring, calibrated canary detection, and temporal leakage accounting rather than relying only on text-level output filters.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

トークンランキングは偽造不可能な言語モデル署名です

言語モデルのパラメータは、ロジット出力に(各モデルに)一意の幾何学的制約を課すことが知られており、これはモデルを識別する署名として機能しますが、API がロジットを配布するときにモデルの最終層パラメータも漏洩します。私たちは、トークンのランキング (確率値ではなく、確率による順序付け) を公開する、より制限的な API を調査し、ランキングも署名を構成することを発見しました。すべてのモデルは、十分な規模の $k$ に対して実行可能な上位 $k$ ランキングの独自のセットを持っています。さらに、同じ実行可能なランキングのセットを持つモデルを見つけることは NP 困難であるため、ランキング署名は最初に知られている (多項式的に) 偽造不可能な署名です。セキュリティの面では、ロジットと同様に、トークンのランキングがすでにモデルの最終層をほぼ盗むのに十分であることがわかりました。ただし、近似が粗すぎて署名を偽造できず、API を十分に小さい $k$ の上位 $k$ トークンに制限することで効果的に対抗できます。モデル署名を提示するために必要な $k$ は一般に、盗用を防ぐために必要な $k$ よりも小さいため、API はモデル パラメーターを漏らすことなく偽造不可能な署名を提示することが可能です。

原文 (English)

Token Rankings are Unforgeable Language Model Signatures

Language model parameters are known to impose unique (to each model) geometric constraints on their logit outputs, which serves as a signature that identifies the model, but also leaks the model's final layer parameters when an API distributes logits. We investigate more restrictive APIs that expose token rankings (i.e., their ordering by probability, but not the probability values) and find that rankings also constitute a signature: every model has a unique set of feasible top-$k$ rankings for sufficiently large $k$. Furthermore, the ranking signature is the first known (polynomially) unforgeable signature, since finding a model with the same set of feasible rankings is NP-hard. On the security front, we find that token rankings are already sufficient to approximately steal the final layer of the model, similar to logits, though the approximation is too coarse to forge the signature, and can be effectively countered by restricting the API to top-$k$ tokens with sufficiently small $k$. Since the top-$k$ required to present the model signature is generally smaller than the $k$ required to prevent stealing, it is possible for an API to present an unforgeable signature without leaking model parameters.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

ルーブリックベースの強化学習における報酬ハッキングの再現、分析、検出

ルーブリックベースの強化学習 (RL) は、LLM-as-a-Judge (LaaJ) を使用して、報酬としてルーブリックに従ってモデルの出力を採点します。ただし、政策モデルは裁判官の潜在的なバイアスを悪用し、報酬のハッキングや非効果的または危険なトレーニング結果につながる可能性があります。現実のルーブリックベースの RL では、このようなハッキング行為は多くの場合微妙であり、複数の裁判官のバイアスと絡み合っているため、分析、検出、軽減することが困難です。このペーパーでは、ルーブリックベースの RL のための制御可能なハッキング環境である CHERRL を紹介します。既知のバイアスを LaaJ に注入することで、CHERRL は報酬ハッキングの安定した再現、報酬の発散の明確な観察、およびハッキングの開始の正確な特定を可能にします。これは、ルーブリック ベースの RL における報酬ハッキングのメカニズムと緩和を研究するためのクリーンな実験テストベッドを提供します。その有用性を実証するために、発見可能性と悪用可能性の観点からさまざまな裁判官のバイアスを分析し、トレーニングログから報酬ハッキングの開始を自動的に検出するためのエージェントベースのシステムを調査します。コードと環境は https://github.com/THUAIS-Lab/CHERRL で公開されています。

原文 (English)

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a controllable hacking environment for rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and precise identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, and explore an agent-based system for automatically detecting reward hacking onset from training logs. The code and environment are publicly available at https://github.com/THUAIS-Lab/CHERRL.

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

制約付き適応拒否サンプリング

言語モデル (LM) は、生成された出力が厳密な意味論的または構文上の制約を満たす必要があるアプリケーションで使用されることが増えています。制約付き生成に対する既存のアプローチはさまざまです。貪欲な制約付きデコード方法は、デコード中に有効性を強制しますが、LM の分布を歪めます。一方、リジェクション サンプリング (RS) は忠実度を維持しますが、無効な出力を破棄することで計算を無駄にします。サンプルの有効性と多様性の両方が重要であるプログラム ファジングなどの領域では、両極端が問題となります。我々は、分布歪みを生じさせずに RS のサンプル効率を厳密に改善するアプローチである、制約付き適応除去サンプリング (CARS) を紹介します。 CARS は、制約のない LM サンプリングから始まり、制約違反の継続をトライに記録し、将来の描画から確率質量を差し引くことで、制約に違反する継続を適応的に除外します。この適応的な枝刈りにより、無効であることが証明されたプレフィックスが決して再検討されず、受け入れ率が単調に向上し、結果として得られるサンプルが制約された分布に正確に従うことが保証されます。プログラムのファジングや分子生成など、さまざまな領域の実験において、CARS は一貫して高い効率 (有効サンプルあたりの LM フォワードパスの数で測定) を達成すると同時に、GCD や LM の分布を近似する方法の両方よりも強力なサンプル多様性を生み出します。

原文 (English)

Constrained Adaptive Rejection Sampling

Language Models (LMs) are increasingly used in applications where generated outputs must satisfy strict semantic or syntactic constraints. Existing approaches to constrained generation fall along a spectrum: greedy constrained decoding methods enforce validity during decoding but distort the LM's distribution, while rejection sampling (RS) preserves fidelity but wastes computation by discarding invalid outputs. Both extremes are problematic in domains such as program fuzzing, where both validity and diversity of samples are essential. We present Constrained Adaptive Rejection Sampling (CARS), an approach that strictly improves the sample-efficiency of RS without distributional distortion. CARS begins with unconstrained LM sampling and adaptively rules out constraint-violating continuations by recording them in a trie and subtracting their probability mass from future draws. This adaptive pruning ensures that prefixes proven invalid are never revisited, acceptance rates improve monotonically, and the resulting samples exactly follow the constrained distribution. In experiments on a variety of domains -- e.g., program fuzzing and molecular generation -- CARS consistently achieves higher efficiency -- measured in the number of LM forward passes per valid sample -- while also producing stronger sample diversity than both GCD and methods that approximate the LM's distribution.

2026-06-05 13:00 JSTarXiv cs.AIハードウェア/半導体

モデルを保持した適応丸め

量子化の目標は、出力分布が元のモデルにできるだけ近い圧縮モデルを生成することです。これを容易に行うために、ほとんどの量子化アルゴリズムは、エンドツーエンド エラーの代理として各層の即時アクティブ化エラーを最小限に抑えます。ただし、これは将来のレイヤーの影響を無視するため、プロキシとしては不十分です。この研究では、ネットワークの出力での誤差を直接考慮する適応丸めアルゴリズムである Yet Another Quantization Algorithm (YAQA) を導入します。 YAQA は、量子化アルゴリズムの最初のエンドツーエンド誤差限界に至る一連の理論的結果を紹介します。まず、ヘッセ近似の構造を介して、適応丸めアルゴリズムの収束時間を特徴付けます。次に、エンドツーエンド誤差が真のヘッセ行列に対する近似のコサイン類似度によって制限される可能性があることを示します。これにより、対応する最適に近いヘッシアン スケッチを使用した自然なクロネッカー因数近似が可能になります。 YAQA は GPTQ/LDLQ よりも優れていることが証明されており、経験的にはこれらの方法よりも誤差が $\約 30\%$ 減少します。 YAQA は、量子化を意識したトレーニングよりも低い誤差を実現します。これにより、推論のオーバーヘッドがまったく追加されずに、ダウンストリーム タスクで最先端のパフォーマンスが得られます。

原文 (English)

Model-Preserving Adaptive Rounding

The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible. To do this tractably, most quantization algorithms minimize the immediate activation error of each layer as a proxy for the end-to-end error. However, this ignores the effect of future layers, making it a poor proxy. In this work, we introduce Yet Another Quantization Algorithm (YAQA), an adaptive rounding algorithm that directly considers the error at the network's output. YAQA introduces a series of theoretical results that culminate in the first end-to-end error bounds for quantization algorithms. First, we characterize the convergence time of adaptive rounding algorithms via the structure of their Hessian approximations. We then show that the end-to-end error can be bounded by the approximation's cosine similarity to the true Hessian. This admits a natural Kronecker-factored approximation with corresponding near-optimal Hessian sketches. YAQA is provably better than GPTQ/LDLQ and empirically reduces the error by $\approx 30\%$ over these methods. YAQA even achieves a lower error than quantization aware training. This translates to state of the art performance on downstream tasks, all while adding no inference overhead.

2026-06-05 13:00 JSTarXiv cs.AIハードウェア/半導体

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences

Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed re…

2026-06-05 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Consistency Training Can Entrench Misalignment

Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, s…

2026-06-04 17:55 JSTITmedia AI+ハードウェア/半導体

TSMC、AI活用拡大による成長維持に自信 株主総会、東京エレクトロンとの取引は継続

半導体受託生産の世界最大手、台湾積体電路製造(TSMC)は6月4日、台湾の新竹市で株主総会を開いた。魏哲家会長兼最高経営責任者(CEO)は、AIの活用拡大により「われわれの最先端技術と製造能力の価値は引き続き成長する」と述べ、今後数年間の同社の成長維持に強い自信を示した。

2026-06-04 13:00 JSTarXiv cs.AIハードウェア/半導体

ASP ベースのコンプライアンス推論のための規範的な中間表現

我々は、ASP ベースのコンプライアンス推論のためのモーダル化出力規範中間表現である MONIR を提案します。そのコア フラグメントには段階的な操作セマンティクスがあり、MONIR-ASP は外部関数、一時的なルール、および安定したモデル推論のための実行可能なコンパイルと拡張機能を提供します。 LLM 支援パイプラインを使用して、中国の ADAS 規制と標準に関するフレームワークをインスタンス化します。実験では、抽出品質と、モジュール式および増分 ASP 解決の効率を評価します。

原文 (English)

A Normative Intermediate Representation for ASP-Based Compliance Reasoning

We propose MONIR, a Modalized-Output Normative Intermediate Representation for ASP-based compliance reasoning. Its core fragment has a staged operational semantics, while MONIR-ASP provides an executable compilation and extensions for external functions, temporal rules, and stable-model reasoning. We instantiate the framework on Chinese ADAS regulations and standards with an LLM-assisted pipeline. Experiments evaluate extraction quality and the efficiency of modular and incremental ASP solving.

2026-06-04 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

DiffAero: 効率的なクアドローター ポリシー学習のための GPU アクセラレーションによる微分可能シミュレーション フレームワーク

このレターでは、効率的なクワッドローター制御ポリシー学習のために設計された、軽量で GPU アクセラレーションを備えた完全微分可能なシミュレーション フレームワークである DiffAero を紹介します。 DiffAero は、環境レベルとエージェント レベルの両方の並列処理をサポートし、複数のダイナミクス モデル、カスタマイズ可能なセンサー スタック (IMU、深度カメラ、LiDAR)、および多様な飛行タスクを統合された GPU ネイティブのトレーニング インターフェイス内に統合します。 DiffAero は、GPU 上で物理とレンダリングの両方を完全に並列化することで、CPU と GPU 間のデータ転送のボトルネックを排除し、シミュレーションのスループットを桁違いに向上させます。既存のシミュレータとは対照的に、DiffAero は高性能シミュレーションを提供するだけでなく、微分可能なハイブリッド学習アルゴリズムを探索するための研究プラットフォームとしても機能します。広範なベンチマークと実際の飛行実験により、DiffAero とハイブリッド学習アルゴリズムを組み合わせることで、消費者グレードのハードウェアで堅牢な飛行ポリシーを数時間で学習できることが実証されました。コードは https://github.com/flyingbitac/diffaero で入手できます。

原文 (English)

DiffAero: A GPU-Accelerated Differentiable Simulation Framework for Efficient Quadrotor Policy Learning

This letter introduces DiffAero, a lightweight, GPU-accelerated, and fully differentiable simulation framework designed for efficient quadrotor control policy learning. DiffAero supports both environment-level and agent-level parallelism and integrates multiple dynamics models, customizable sensor stacks (IMU, depth camera, and LiDAR), and diverse flight tasks within a unified, GPU-native training interface. By fully parallelizing both physics and rendering on the GPU, DiffAero eliminates CPU-GPU data transfer bottlenecks and delivers orders-of-magnitude improvements in simulation throughput. In contrast to existing simulators, DiffAero not only provides high-performance simulation but also serves as a research platform for exploring differentiable and hybrid learning algorithms. Extensive benchmarks and real-world flight experiments demonstrate that DiffAero and hybrid learning algorithms combined can learn robust flight policies in hours on consumer-grade hardware. The code is available at https://github.com/flyingbitac/diffaero.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CodegenBench: LLM はアーキテクチャ全体で効率的なコードを記述できますか?

大規模言語モデル (LLM) は、汎用プログラミングや GPU アクセラレーション環境 (PyTorch、CUDA など) のコード生成タスクで広範囲に評価されてきましたが、多様なアーキテクチャにわたる CPU 指向のハイパフォーマンス コンピューティング (HPC) における LLM の機能はまだ十分に解明されていません。このギャップを埋めるために、x86_64、Sunway、Kunpeng の 3 つの異なるハードウェア プラットフォームにわたる効率的な並列コードの生成を評価するように設計された包括的なベンチマーク スイートである CodegenBench を紹介します。私たちのベンチマークは、基本的なベースラインを確立する 106 個の標準基本線形代数サブプログラム (BLAS) ルーチンと、独自のスーパーコンピューティング アーキテクチャ (LeetSunway および LeetKunpeng) のそれぞれに適合した 20 個の特殊な計算カーネルで構成されています。私たちの広範な評価により、最先端の LLM は x86_64 のようなユビキタス アーキテクチャ向けに最適化されたコードを生成できる一方で、公開ドキュメントやトレーニング データが限られたドメイン固有のアーキテクチャでは大幅なパフォーマンスの低下を示し、クロスプラットフォームの一般化における重大な制限が浮き彫りになったことが明らかになりました。さらに、実装の長さやタスクの複雑さなど、コードの品質に影響を与える要因を分析したところ、現在の LLM は、簡潔なコード スニペットを必要とする中程度に難しい問題に対して最も効果的であることが示されています。私たちは、LLM 主導の高性能コード生成における将来の研究を促進するために、データセットと自動評価インフラストラクチャをオープンソースにしています。リソースは https://anonymous.4open.science/r/CodegenBench-EDE1/ および https://anonymous.4open.science/r/CodegenBenchDataset-2551 で利用できます。

原文 (English)

CodegenBench: Can LLMs Write Efficient Code Across Architectures?

While large language models (LLMs) have been extensively evaluated on code generation tasks for general-purpose programming and GPU-accelerated environments (e.g., PyTorch, CUDA), their capabilities in CPU-oriented high-performance computing (HPC) across diverse architectures remain underexplored. To bridge this gap, we introduce CodegenBench, a comprehensive benchmark suite designed to evaluate the generation of efficient parallel code across three distinct hardware platforms: x86_64, Sunway, and Kunpeng. Our benchmark comprises 106 standard Basic Linear Algebra Subprograms (BLAS) routines establishing a fundamental baseline, alongside 20 specialized computational kernels adapted for each of the unique supercomputing architectures (LeetSunway and LeetKunpeng). Our extensive evaluation reveals that while state-of-the-art LLMs can generate optimized code for ubiquitous architectures like x86_64, they exhibit significant performance degradation on domain-specific architectures with limited public documentation and training data, highlighting critical limitations in cross-platform generalization. Furthermore, our analysis of factors influencing code quality such as implementation length and task complexity indicates that current LLMs are most effective for moderately difficult problems requiring concise code snippets. We open-source our dataset and automated evaluation infrastructure to facilitate future research in LLM-driven high-performance code generation. The resources are available at https://anonymous.4open.science/r/CodegenBench-EDE1/ and https://anonymous.4open.science/r/CodegenBenchDataset-2551.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Caught in the Act(ivation): LLM エージェントによる資格情報漏洩の事前出力およびマルチターン検出に向けて

LLM エージェントは多くの場合、機密認証情報を信頼できない取得コンテンツと同じコンテキスト ウィンドウに配置し、認証情報の漏洩を誘発する間接的なプロンプト インジェクションの直接パスを作成します。私たちは、3 つの相補的な防御を通じてこの障害モードを研究します。まず、出力トークンが発行される前に、アクティベーション プローブが資格情報へのアクセスを検出できるかどうかを尋ねます。次に、形式固有の文字モデルからハニートークンを構築し、分割等角予測で検出を調整します。 3 番目に、複数ターンにわたる漏洩を累積的な情報フロー問題として扱い、会話ターン全体での推定漏洩予算を追跡します。オープンウェイト モデルの制御された実験では、アクティベーション機能により、ホールドアウト エンコーディング変換下を含め、無害なプロンプトと認証情報を求めるプロンプトが高精度で分離されます。小規模な合成マルチターン スイートでは、累積アカウンティングにより、ターンごとの検出器が見逃した攻撃が検出されます。これらの結果は暫定的なものです。マルチターン ベンチマークは社内で小規模なものであり、アクティブ化方法にはホワイト ボックス アクセスが必要であり、情報推定ツールは正式な上限ではなく実用的なシグナルを提供します。それでも、この結果は、資格情報の漏洩防御には、テキストレベルの出力フィルターのみに依存するのではなく、出力前の監視、調整されたカナリア検出、および一時的な漏洩アカウンティングを組み合わせる必要があることを示唆しています。

原文 (English)

Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents

LLM agents often place sensitive credentials in the same context window as untrusted retrieved content, creating a direct path for indirect prompt injection to induce credential exfiltration. We study this failure mode through three complementary defenses. First, we ask whether activation probes can detect credential access before output tokens are emitted. Second, we construct honeytokens from format-specific character models and calibrate detection with split conformal prediction. Third, we treat multi-turn exfiltration as a cumulative information-flow problem and track an estimated leakage budget across conversation turns. In controlled experiments on open-weight models, activation features separate benign and credential-seeking prompts with high accuracy, including under held-out encoding transformations. In a small synthetic multi-turn suite, cumulative accounting detects attacks that per-turn detectors miss. These results are preliminary: the multi-turn benchmark is in-house and small, the activation method requires white-box access, and the information estimator provides a practical signal rather than a formal upper bound. Still, the results suggest that credential-exfiltration defenses should combine pre-output monitoring, calibrated canary detection, and temporal leakage accounting rather than relying only on text-level output filters.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

トークンランキングは偽造不可能な言語モデル署名です

言語モデルのパラメータは、ロジット出力に(各モデルに)一意の幾何学的制約を課すことが知られており、これはモデルを識別する署名として機能しますが、API がロジットを配布するときにモデルの最終層パラメータも漏洩します。私たちは、トークンのランキング (確率値ではなく、確率による順序付け) を公開する、より制限的な API を調査し、ランキングも署名を構成することを発見しました。すべてのモデルは、十分な規模の $k$ に対して実行可能な上位 $k$ ランキングの独自のセットを持っています。さらに、同じ実行可能なランキングのセットを持つモデルを見つけることは NP 困難であるため、ランキング署名は最初に知られている (多項式的に) 偽造不可能な署名です。セキュリティの面では、ロジットと同様に、トークンのランキングがすでにモデルの最終層をほぼ盗むのに十分であることがわかりました。ただし、近似が粗すぎて署名を偽造できず、API を十分に小さい $k$ の上位 $k$ トークンに制限することで効果的に対抗できます。モデル署名を提示するために必要な $k$ は一般に、盗用を防ぐために必要な $k$ よりも小さいため、API はモデル パラメーターを漏らすことなく偽造不可能な署名を提示することが可能です。

原文 (English)

Token Rankings are Unforgeable Language Model Signatures

Language model parameters are known to impose unique (to each model) geometric constraints on their logit outputs, which serves as a signature that identifies the model, but also leaks the model's final layer parameters when an API distributes logits. We investigate more restrictive APIs that expose token rankings (i.e., their ordering by probability, but not the probability values) and find that rankings also constitute a signature: every model has a unique set of feasible top-$k$ rankings for sufficiently large $k$. Furthermore, the ranking signature is the first known (polynomially) unforgeable signature, since finding a model with the same set of feasible rankings is NP-hard. On the security front, we find that token rankings are already sufficient to approximately steal the final layer of the model, similar to logits, though the approximation is too coarse to forge the signature, and can be effectively countered by restricting the API to top-$k$ tokens with sufficiently small $k$. Since the top-$k$ required to present the model signature is generally smaller than the $k$ required to prevent stealing, it is possible for an API to present an unforgeable signature without leaking model parameters.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

ルーブリックベースの強化学習における報酬ハッキングの再現、分析、検出

ルーブリックベースの強化学習 (RL) は、LLM-as-a-Judge (LaaJ) を使用して、報酬としてルーブリックに従ってモデルの出力を採点します。ただし、政策モデルは裁判官の潜在的なバイアスを悪用し、報酬のハッキングや非効果的または危険なトレーニング結果につながる可能性があります。現実のルーブリックベースの RL では、このようなハッキング行為は多くの場合微妙であり、複数の裁判官のバイアスと絡み合っているため、分析、検出、軽減することが困難です。このペーパーでは、ルーブリックベースの RL のための制御可能なハッキング環境である CHERRL を紹介します。既知のバイアスを LaaJ に注入することで、CHERRL は報酬ハッキングの安定した再現、報酬の発散の明確な観察、およびハッキングの開始の正確な特定を可能にします。これは、ルーブリック ベースの RL における報酬ハッキングのメカニズムと緩和を研究するためのクリーンな実験テストベッドを提供します。その有用性を実証するために、発見可能性と悪用可能性の観点からさまざまな裁判官のバイアスを分析し、トレーニングログから報酬ハッキングの開始を自動的に検出するためのエージェントベースのシステムを調査します。コードと環境は https://github.com/THUAIS-Lab/CHERRL で公開されています。

原文 (English)

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning

Rubric-based reinforcement learning (RL) uses an LLM-as-a-Judge (LaaJ) to score model outputs according to rubrics as rewards. However, policy models may exploit latent biases in the judge, leading to reward hacking and ineffective or unsafe training outcomes. In real-world rubric-based RL, such hacking behaviors are often subtle and entangled with multiple judge biases, making them difficult to analyze, detect, and mitigate. In this paper, we introduce CHERRL, a controllable hacking environment for rubric-based RL. By injecting known biases into LaaJ, CHERRL enables stable reproduction of reward hacking, explicit observation of reward divergence, and precise identification of hacking onset. This provides a clean experimental testbed for studying the mechanisms and mitigations of reward hacking in rubric-based RL. To demonstrate its utility, we analyze different judge biases from the perspectives of discoverability and exploitability, and explore an agent-based system for automatically detecting reward hacking onset from training logs. The code and environment are publicly available at https://github.com/THUAIS-Lab/CHERRL.

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

制約付き適応拒否サンプリング

言語モデル (LM) は、生成された出力が厳密な意味論的または構文上の制約を満たす必要があるアプリケーションで使用されることが増えています。制約付き生成に対する既存のアプローチはさまざまです。貪欲な制約付きデコード方法は、デコード中に有効性を強制しますが、LM の分布を歪めます。一方、リジェクション サンプリング (RS) は忠実度を維持しますが、無効な出力を破棄することで計算を無駄にします。サンプルの有効性と多様性の両方が重要であるプログラム ファジングなどの領域では、両極端が問題となります。我々は、分布歪みを生じさせずに RS のサンプル効率を厳密に改善するアプローチである、制約付き適応除去サンプリング (CARS) を紹介します。 CARS は、制約のない LM サンプリングから始まり、制約違反の継続をトライに記録し、将来の描画から確率質量を差し引くことで、制約に違反する継続を適応的に除外します。この適応的な枝刈りにより、無効であることが証明されたプレフィックスが決して再検討されず、受け入れ率が単調に向上し、結果として得られるサンプルが制約された分布に正確に従うことが保証されます。プログラムのファジングや分子生成など、さまざまな領域の実験において、CARS は一貫して高い効率 (有効サンプルあたりの LM フォワードパスの数で測定) を達成すると同時に、GCD や LM の分布を近似する方法の両方よりも強力なサンプル多様性を生み出します。

原文 (English)

Constrained Adaptive Rejection Sampling

Language Models (LMs) are increasingly used in applications where generated outputs must satisfy strict semantic or syntactic constraints. Existing approaches to constrained generation fall along a spectrum: greedy constrained decoding methods enforce validity during decoding but distort the LM's distribution, while rejection sampling (RS) preserves fidelity but wastes computation by discarding invalid outputs. Both extremes are problematic in domains such as program fuzzing, where both validity and diversity of samples are essential. We present Constrained Adaptive Rejection Sampling (CARS), an approach that strictly improves the sample-efficiency of RS without distributional distortion. CARS begins with unconstrained LM sampling and adaptively rules out constraint-violating continuations by recording them in a trie and subtracting their probability mass from future draws. This adaptive pruning ensures that prefixes proven invalid are never revisited, acceptance rates improve monotonically, and the resulting samples exactly follow the constrained distribution. In experiments on a variety of domains -- e.g., program fuzzing and molecular generation -- CARS consistently achieves higher efficiency -- measured in the number of LM forward passes per valid sample -- while also producing stronger sample diversity than both GCD and methods that approximate the LM's distribution.

2026-06-04 13:00 JSTarXiv cs.AIハードウェア/半導体

モデルを保持した適応丸め

量子化の目標は、出力分布が元のモデルにできるだけ近い圧縮モデルを生成することです。これを容易に行うために、ほとんどの量子化アルゴリズムは、エンドツーエンド エラーの代理として各層の即時アクティブ化エラーを最小限に抑えます。ただし、これは将来のレイヤーの影響を無視するため、プロキシとしては不十分です。この研究では、ネットワークの出力での誤差を直接考慮する適応丸めアルゴリズムである Yet Another Quantization Algorithm (YAQA) を導入します。 YAQA は、量子化アルゴリズムの最初のエンドツーエンド誤差限界に至る一連の理論的結果を紹介します。まず、ヘッセ近似の構造を介して、適応丸めアルゴリズムの収束時間を特徴付けます。次に、エンドツーエンド誤差が真のヘッセ行列に対する近似のコサイン類似度によって制限される可能性があることを示します。これにより、対応する最適に近いヘッシアン スケッチを使用した自然なクロネッカー因数近似が可能になります。 YAQA は GPTQ/LDLQ よりも優れていることが証明されており、経験的にはこれらの方法よりも誤差が $\約 30\%$ 減少します。 YAQA は、量子化を意識したトレーニングよりも低い誤差を実現します。これにより、推論のオーバーヘッドがまったく追加されずに、ダウンストリーム タスクで最先端のパフォーマンスが得られます。

原文 (English)

Model-Preserving Adaptive Rounding

The goal of quantization is to produce a compressed model whose output distribution is as close to the original model's as possible. To do this tractably, most quantization algorithms minimize the immediate activation error of each layer as a proxy for the end-to-end error. However, this ignores the effect of future layers, making it a poor proxy. In this work, we introduce Yet Another Quantization Algorithm (YAQA), an adaptive rounding algorithm that directly considers the error at the network's output. YAQA introduces a series of theoretical results that culminate in the first end-to-end error bounds for quantization algorithms. First, we characterize the convergence time of adaptive rounding algorithms via the structure of their Hessian approximations. We then show that the end-to-end error can be bounded by the approximation's cosine similarity to the true Hessian. This admits a natural Kronecker-factored approximation with corresponding near-optimal Hessian sketches. YAQA is provably better than GPTQ/LDLQ and empirically reduces the error by $\approx 30\%$ over these methods. YAQA even achieves a lower error than quantization aware training. This translates to state of the art performance on downstream tasks, all while adding no inference overhead.

2026-06-04 13:00 JSTarXiv cs.AIハードウェア/半導体

Curated Synthetic Data Doesn't Have to Collapse: A Theoretical Study of Generative Retraining with Pluralistic Preferences

Recursive retraining of generative models poses a critical representation challenge: when synthetic outputs are curated based on a fixed re…

2026-06-04 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Consistency Training Can Entrench Misalignment

Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, s…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

TriEval: LLM バイアス、毒性、真実性評価のためのリソース効率の高いパイプライン

LLM は、基本的なチャットボットから AI エコシステムのバックボーンに進化し、現在では医療、学校、政府サービスで広く使用されています。 LLM をドメイン全体に導入するには、その安全性と公平性を確保するために継続的な評価が必要です。 LLM の導入後に発生する一般的な問題には、一貫性のない出力や誤った情報の幻覚などがあります。 LLM 評価ツールは多数存在しますが、そのほとんどは一度に 1 つのパラメータのテストに限定されているか、ほとんどの研究者がアクセスできない膨大な計算リソースを必要とします。 TriEval は、コンピューティング リソースを最小限に抑えながら、バイアス、有害性、真実性を含む複数のパラメータにわたって LLM 出力を評価することで、これらの課題に対処します。このパイプラインは、オープンソース モデルとクローズドソース モデルの両方と互換性があり、GPU クラスターのない標準的なラップトップで実行されます。 TriEval は、Llama 3 8B、Mistral 7B、Gemma 2 9B、および Claude Haiku の 4 つのモデルでテストされています。結果は、特に毒性と真実性の点で、オープンソース モデルとクローズドソース モデルの明らかな違いを示しています。 TriEval は、限られた計算リソースを持つ研究者がより広範にアクセスできるようにするために、オープンソースとしてリリースされています。

原文 (English)

TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment

LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-wide adoption of LLMs necessitates continuous evaluation to ensure their safety and fairness. Common issues encountered after deploying LLMs include inconsistent outputs and hallucinations of incorrect information. Although numerous LLM evaluation tools exist, most are limited to testing a single parameter at a time or require massive computational resources that are not accessible to most researchers. TriEval addresses these challenges by evaluating LLM outputs across multiple parameters, including bias, toxicity, and truthfulness together, while minimizing computing resources. The pipeline is compatible with both open- and closed-source models and runs on a standard laptop without a GPU cluster. TriEval has been tested on four models: Llama 3 8B, Mistral 7B, Gemma 2 9B, and Claude Haiku. The results show clear differences between open-source and closed-source models, especially in terms of toxicity and truthfulness. TriEval is being released as open source to enable broader access for researchers with limited computational resources.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

プルーフ リファクタリング: 生成された正式なプルーフをモジュール型アーティファクトにリファクタリングする

大規模言語モデル (LLM) は形式的な証明の生成において優れたパフォーマンスを示していますが、その出力は多くの場合、成熟した形式的な数学ライブラリの証明に比べて可読性、モジュール性、保守性、再利用性が劣ります。私たちは、このギャップの一部は、ほとんどの証明生成パイプラインに暗黙的に含まれるコンパイル優先の目的に起因しており、ライブラリ品質のアーティファクトではなく、モノリシックまたはアドホック証明スクリプトを奨励していると主張します。証明品質を向上させるための既存のアプローチは、多くの場合、明示的で計算可能な最適化目標に依存しています。ただし、実際には、最も扱いやすく、実験的に検証された目標は主に長さに基づくものですが、可読性、モジュール性、保守性、再利用性などのより高いレベルの品質を信頼できる自動メトリクスに還元するのは困難です。単一のプロキシ メトリクスに対して証明の改善を最適化するのではなく、人間による証明のリファクタリング ワークフローからインスピレーションを得た、プロセスに基づいたアプローチを採用します。私たちは、証明リファクタリングを 4 つのフェーズに分解するエージェント フレームワーク $\textbf{Proof-Refactor}$ を提案します。候補となる証明フラグメントの抽出、ヘルパー宣言の設計、抽出および設計されたコンポーネントの正式な証明、検証されたコンポーネントを使用した元の証明の修復です。 PutnamBench および Putnam2025 から生成されたリーン証明では、Proof-Refactor は、強力なクロード コード リファクタリング ベースラインよりもルーブリック ベースのリファクタリング スコアを改善し、署名の品質と人間の可読性が最大の向上をもたらします。これらの結果は、プロセスガイド付きリファクタリングにより、証明長を主な目的として扱うことなく証明構造を改善できることを示唆しています。

原文 (English)

Proof-Refactor: Refactoring Generated Formal Proofs into Modular Artifacts

While Large Language Models (LLMs) have shown strong performance in generating formal proofs, their outputs often remain less readable, modular, maintainable, and reusable than proofs in mature formal mathematics libraries. We argue that this gap stems in part from the compile-first objective implicit in most proof-generation pipelines, which encourages monolithic or ad hoc proof scripts rather than library-quality artifacts. Existing approaches to proof-quality improvement often rely on explicit, computable optimization objectives. In practice, however, the most tractable and experimentally validated objectives are largely length-based, while higher-level qualities such as readability, modularity, maintainability, and reusability are difficult to reduce to reliable automatic metrics. Instead of optimizing proof improvement against a single proxy metric, we take a process-guided approach inspired by human proof-refactoring workflows. We propose an agentic framework $\textbf{Proof-Refactor}$ that decomposes proof refactoring into four phases: extracting candidate proof fragments, designing helper declarations, formally proving the extracted and designed components, and repairing the original proof using the verified components. On generated Lean proofs from PutnamBench and Putnam2025, Proof-Refactor improves rubric-based refactoring scores over a strong Claude Code refactoring baseline, with the largest gains in signature quality and human readability. These results suggest that process-guided refactoring can improve proof structure without treating proof length as the primary objective.

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

D-Judge: Disrupting Multi-Turn Jailbreaks using Semantics-Preserving Output Rewriting

Multi-turn jailbreak attacks pose a growing threat to large language model (LLM) safety because they exploit feedback from auxiliary judge…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ハードウェア/半導体研究/論文

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination.…

2026-06-03 13:00 JSTarXiv cs.AI画像/動画生成エージェントロボティクスハードウェア/半導体ビジネス/資金調達

NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation

As autonomous vehicle capabilities advance, the safe evaluation of driving policies in long-tail scenarios remains a critical bottleneck. I…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

The Unsampled Truth: Psychometrics in SLMs Measure Prompt Artifacts, Not Psychological Constructs

When prompting SLMs for psychometric assessments, researchers assume the outputs reflect semantic reasoning. We evaluate this premise acros…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

DDOR: Delta Debugging for Explainable Overrefusal Testing and Repair

While safety alignment and guardrails help large language models (LLMs) avoid harmful outputs, they can also induce overrefusal, i.e., unwa…

2026-06-03 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Consistency Training Can Entrench Misalignment

Consistency training encourages a model to produce similar outputs across related inputs or sampling procedures. Such methods are simple, s…

2026-06-03 13:00 JSTarXiv cs.AIハードウェア/半導体

FlashMLA-ETAP: Efficient Transpose Attention Pipeline for Accelerating MLA Inference on NVIDIA H20 GPUs

Efficient inference of Multi-Head Latent Attention (MLA) is challenged by deploying the DeepSeek-R1 671B model on a single Multi-GPU server…

2026-06-03 13:00 JSTarXiv cs.AIハードウェア/半導体

Google Cloud TPU での Gemma 4 31B の微調整と提供: GPU ベースラインとの技術比較

TPU ハードウェア上で Google の Gemma 4 31B モデルを微調整して提供する最初のエンドツーエンドのデモンストレーションを紹介し、大規模な言語モデルの適応に関する TPU プラットフォームと GPU プラットフォームの実証的比較を提供します。トレーニングには Google TPU v5p-8、推論には TPU v6e-8 (Trillium) 上の LoRA を使用して、PyTorch、HuggingFace TRL、FSDP 上に構築された GPU ネイティブのトレーニング レシピを JAX + Tunix/Qwix スタックに移植するために必要なコードレベルの適応の完全なセットを文書化します。これらの適応は、メッシュ構成、LoRA モジュールの命名規則、シャーディング アノテーションの修正、勾配チェックポイント設定、データ パイプラインの再構築、およびカスタム Orbax からセーフテンソルへのチェックポイント マージ手順に及びます。推論のために、v6e-8 で Gemma 4 を提供するために必要な vLLM-TPU Docker セットアップを詳しく説明し、その結果生じるレイテンシとスループット プロファイルの特徴を説明します。同一のハイパーパラメータの下での 2xH100 GPU ベースラインと比較して、TPU トレーニングは 2.12 倍のコストで 1.61 倍の速度で完了します。推論スループットはプラットフォーム全体で 3% 以内ですが、TPU は最初のトークンまでの時間を 2 倍短縮します (235 ミリ秒対 475 ミリ秒)。これらを合計すると、TPU 構成は、代表的なトレインとサービスのワークロードでは 1.82 倍安くなります。私たちの取り組みは、オープン ツール エコシステムの重大なギャップを取り除き、TPU インフラストラクチャ上で Gemma 4 を導入するための再現可能で本番環境にすぐに使えるレシピを実務者に提供します。

原文 (English)

Fine-Tuning and Serving Gemma 4 31B on Google Cloud TPU: A Technical Comparison with GPU Baselines

We present the first end-to-end demonstration of fine-tuning and serving Google's Gemma 4 31B model on TPU hardware, providing an empirical comparison of TPU and GPU platforms for large language model adaptation. Using LoRA on a Google TPU v5p-8 for training and TPU v6e-8 (Trillium) for inference, we document the full set of code-level adaptations required to port a GPU-native training recipe, built on PyTorch, HuggingFace TRL, and FSDP, to the JAX + Tunix/Qwix stack. These adaptations span mesh configuration, LoRA module naming conventions, sharding annotation corrections, gradient checkpointing, data pipeline restructuring, and a custom Orbax-to-safetensors checkpoint merging procedure. For inference, we detail the vLLM-TPU Docker setup necessary to serve Gemma 4 on v6e-8 and characterize the resulting latency and throughput profile. Compared with a 2xH100 GPU baseline under identical hyperparameters, TPU training completes 1.61x faster at 2.12x lower cost. Inference throughput is within 3% across platforms, while TPU achieves 2x lower time-to-first-token (235 ms vs. 475 ms). Together, the TPU configuration is 1.82x cheaper for a representative train-plus-service workload. Our work removes a critical gap in the open tooling ecosystem and provides practitioners with a reproducible, production-ready recipe for Gemma 4 deployment on TPU infrastructure.

2026-06-03 07:00 JSTITmedia AI+ハードウェア/半導体

AI需要で半導体不足は「しばらく続く」 PCメーカー、デルの対応策は?

AI需要による半導体不足は「しばらく続く」――PCメーカーのデル・テクノロジーズはこう予測する。同社はこの難局をどう乗り切るのか。

2026-06-03 06:15 JSTITmedia AI+ハードウェア/半導体

NVIDIAの「RTX Spark」と搭載ノートPCがCOMUPTEX TAIPEIのMediaTekブースに集結

MediaTek(メディアテック)は、「COMPUTEX TAIPEI 2026」において、NVIDIAが発表したAIスーパーチップ「NVIDIA RTX Spark」と、同チップを搭載する各社のWindowsノートPCを披露した。

2026-06-03 04:50 JSTITmedia AI+ハードウェア/半導体

Microsoft、NVIDIAのSoC搭載でAI特化のミニPC「Surface RTX Spark Dev Box」披露

Microsoftは「Build 2026」で、AI特化型デスクトップPC「Surface RTX Spark Dev Box」を発表した。NVIDIAの「RTX Spark」を搭載し、最大1ペタフロップスの演算性能と128GBのメモリにより、1200億パラメータ超のモデルのロ…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

ポジションペーパー: 意思決定エンジンにおけるソルブ後のロバスト性: 摂動下での実行可能領域と滑らかさ

混合整数線形計画法 (MILP) 意思決定エンジンは、一か八かの産業システム向けに名目上最適な計画を定期的に出力します。しかし、導入が解決時間の想定と一致することはほとんどありません。コスト、需要、またはリソースの可用性における小さな変動により、実現可能性が無効になったり、質的に異なるソリューションへの不連続な移行が引き起こされる可能性があります。私たちは、この解決後の堅牢性のギャップは、今日の最適化パイプラインに欠けている層であり、学習対応の意思決定システムに欠けている評価次元であると主張します。提案された層は、ロバストな最適化や確率的プログラミングを置き換えるのではなく、解決された既存のソリューションを監査し、そのソリューションがどの程度信頼できるかについてソルバーに裏付けられた証拠を返します。中心となる 2 つのオブジェクトを形式化します。(i) パラメータ空間における $\epsilon$-near-optimal の実現可能近傍。摂動下で既存の企業が実現可能かつ最適に近い状態を保つ時期を捉えます。(ii) 意思決定空間における解の滑らかさ。小さな組み合わせ編集による近くの代替案が競争力を維持しているかどうかを捉えます。次に、感度と安定性の分析、ロバストな最適化、近傍検索、敵対的テスト、学習ベースの機能強化から最も関連性の高い部分的な回答を合成し、統合されたポストソルブ堅牢性レイヤーのアジェンダを明確にします。具体的には、校正された不確実性、敵対的ロバスト性マージン、ソルバーに裏付けされた検証と連携した学習ベースの予測と説明を備えた、既存の確率論的ロバスト性推定に関する認定された内部近似を求めます。最後に、堅牢性を意思決定エンジンの第一級の出力にするコンパクトなレポート テンプレートと評価プロトコルを紹介します。

原文 (English)

Position Paper: Post-Solve Robustness in Decision Engines: Feasible Regions and Smoothness Under Perturbations

Mixed-Integer Linear Programming (MILP) decision engines routinely output nominally optimal plans for high-stakes industrial systems. Yet deployment rarely matches solve-time assumptions: small perturbations in costs, demands, or resource availability can invalidate feasibility or trigger discontinuous shifts to qualitatively different solutions. We argue that this post-solve robustness gap is a missing layer in today's optimization pipelines and a missing evaluation dimension for learning-enabled decision systems. Rather than replacing robust optimization or stochastic programming, the proposed layer audits a solved incumbent and returns solver-backed evidence about how far that solution can be trusted. We formalize two central objects: (i) an $\epsilon$-near-optimal feasible neighborhood in parameter space, capturing when an incumbent remains feasible and near-optimal under perturbations, and (ii) solution smoothness in decision space, capturing whether nearby alternatives with small combinatorial edits remain competitive. We then synthesize the most relevant partial answers from sensitivity and stability analysis, robust optimization, neighborhood search, adversarial testing, and learning-based enhancements, and articulate an agenda for a unified post-solve robustness layer. Concretely, we call for certified inner approximations around the incumbent, probabilistic robustness estimation with calibrated uncertainty, adversarial robustness margins, and learning-based prediction and explanation aligned with solver-backed verification. We conclude with a compact reporting template and evaluation protocol that would make robustness a first-class output of decision engines.

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

TIGER: マルチモーダル生成における幻覚を軽減するためのグラフベースの証拠ルーティングによる追跡可能な推論

私たちは、入力ではサポートされていない特定のファクトが滑らかな出力に含まれる可能性がある、マルチモーダル生成のためのファクトレベル修復を研究します。既存の推論時修復手法は、入力と現在の出力を共同で調整することによってフィードバックを生成することがよくあります。この設計には 2 つの制限があります。出力内の幻覚的な主張により、入力のモデルの解釈にバイアスがかかる可能性があること、および自由形式のフィードバックをファクト レベルでランク付けしたりスケジュールしたりすることができないことです。局所的な修復のためにフィードバックを再設計する推論時間フレームワークである TIGER を紹介します。 TIGER は、入力から観測グラフを抽出し、現在の出力からクレーム グラフを個別に抽出し、サポートと競合に基づいてグラフで条件付けされたリスク スコアを各クレームに割り当てます。このモデルは、バックボーンを凍結したままにしながら、選択された高リスクの請求を修復します。我々は、穏やかな仮定の下で、予想される総リスクが幾何学的に明示的な漸近限界まで減少することを示す収束分析を提供します。画像からテキストへ、画像+テキストからテキストへ、音声からテキストへ、ビデオからテキストへを含む 4 つのクロスモーダル パスにわたる実験では、TIGER がタスクの品質を維持しながらサポートされていないコンテンツを削減することが示されています。このゲインは複数のバックボーンにわたって維持されており、CrisisFACTS のケーススタディでは、同じ修復メカニズムにより複数の電源設定でグラウンディングを改善できることが示唆されています。

原文 (English)

TIGER: Traceable Inference with Graph-Based Evidence Routing for Mitigating Hallucinations in Multimodal Generation

We study fact-level repair for multimodal generation, where a fluent output may contain specific facts that are not supported by the input. Existing inference-time repair methods often generate feedback by jointly conditioning on the input and the current output. This design has two limitations: hallucinated claims in the output can bias the model's interpretation of the input, and free-form feedback cannot be ranked or scheduled at the fact level. We present TIGER, an inference-time framework that redesigns feedback for localized repair. TIGER independently extracts an observation graph from the input and a claim graph from the current output, then assigns each claim a graph-conditioned risk score based on support and conflict. The model repairs selected high-risk claims while keeping the backbone frozen. We provide a convergence analysis showing that the expected total risk decreases geometrically to an explicit asymptotic bound under mild assumptions. Experiments across four cross-modal paths, including image-to-text, image+text-to-text, audio-to-text, and video-to-text, show that TIGER reduces unsupported content while preserving task quality. The gains hold across multiple backbones, and a CrisisFACTS case study suggests that the same repair mechanism can improve grounding in multi-source settings.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

デコーダ層スキップによる大規模言語モデルの幻覚の軽減

大規模言語モデル (LLM) は、さまざまな自然言語タスクにわたって優れたパフォーマンスを達成していますが、その出力には幻覚、つまり事実の情報と一致しないコンテンツが含まれることがよくあります。この研究では、デコードプロセスの包括的な層ごとの分析を実施し、幻覚がより深いデコーダ層から発生する傾向があることを明らかにしました。この問題に対処するために、幻覚を生成しやすい層を動的にスキップする新しいデコード フレームワークである \textbf{DeLask} (\textbf{De}coder \textbf{La}yer \textbf{Sk}ipping) を導入します。 DeLask は、$L$ 層の Transformer の順方向計算が条件付きで勾配降下法の $L$ ステップと同等であるという理論的な洞察を活用します。連続するデコーダ ステップから導出された勾配間のコサイン類似度を計算することで \emph{ドリフタンス値} を定義し、降下方向が反転したときに問題のある層を特定します。 DeLask は、そのような層を完全に破棄するのではなく、その隠れ状態を先行層と部分的に集約することにより、誤った信号を抑制しながら一貫性を維持します。さまざまな LLM とベンチマークにわたる広範な実験により、DeLask が一貫して幻覚を軽減し、全体的な信頼性を向上させ、大規模な言語モデルの堅牢性を向上させるための軽量で一般化可能なデコード フレームワークを提供することが実証されました。

原文 (English)

Mitigating Hallucinations in Large Language Models Via Decoder Layer Skipping

Large Language Models (LLMs) have achieved strong performance across diverse natural language tasks, yet their outputs often suffer from hallucinations -- content that is misaligned with factual information. In this work, we conduct a comprehensive layer-wise analysis of the decoding process and reveal that hallucinations tend to originate from deeper decoder layers. To address this issue, we introduce \textbf{DeLask} (\textbf{De}coder \textbf{La}yer \textbf{Sk}ipping), a novel decoding framework that dynamically skips layers prone to producing hallucinations. DeLask leverages the theoretical insight that the forward computation of an $L$-layer Transformer is conditionally equivalent to $L$ steps of gradient descent. We define a \emph{driftance value} by computing the cosine similarity between gradients derived from consecutive decoder steps, identifying problematic layers when the descent direction reverses. Rather than discarding such layers entirely, DeLask partially aggregates their hidden states with preceding layers, thereby preserving consistency while suppressing erroneous signals. Extensive experiments across diverse LLMs and benchmarks demonstrate that DeLask consistently mitigates hallucinations and enhances overall reliability, providing a lightweight and generalizable decoding framework for improving the robustness of large-scale language models.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

知恵の形: 言語モデルにおける意思決定の軌跡

言語モデルは、出力層で単純に答えを選択するわけではありません。 Qwen2.5-7B-Instruct、Llama-3.1-8B-Instruct、Mistral-7B-Instruct-v0.3 にわたる 9,000 のトラジェクト MMLU スタディでは、回答のスコアは構造化された方法で深度全体に移動します。各軌跡は、現在の解答マージン、そのマージンにおける次の層の変更、および決定フリップからの距離という 3 つの量で記述されます。主な経験的状況は、正しさと安定性は異なるということです。最大のグループは不安定で正しいものであり、安定して正しいものではありません。次に、トレースされたサブセットは、何がマージンを動かすのかを尋ねます。安定した正しいケースでは、平均注意スカラーは正しい方向を向いていますが、平均 MLP スカラーはそうではありません。スパン削除では、回答をサポートするテキストを削除すると余白が損なわれ、気が散るようなテキストを削除すると余白が有効になることがわかります。この結果は回路の完全な説明にはなりません。これは、どの答えが解決され、どの答えが脆弱なままで、どの測定されたソースがそれらを動かしているのかを確認する再現可能な方法です。

原文 (English)

The Shape of Wisdom: Decision Trajectories in Language Models

Language models do not simply choose an answer at the output layer. In a 9,000-trajectory MMLU study across Qwen2.5-7B-Instruct, Llama-3.1-8B-Instruct, and Mistral-7B-Instruct-v0.3, the score of the answer moves across depth in structured ways. We describe each trajectory with three quantities: the current answer margin, the next-layer change in that margin, and the distance from a decision flip. The main empirical picture is that correctness and stability are different: the largest group is unstable-correct, not stable-correct. A traced subset then asks what moves the margin. In stable-correct cases, the average attention scalar points in the correct direction, while the average MLP scalar does not; span deletion shows that removing answer-supporting text hurts the margin and removing distractor-like text helps it. The result is not a full circuit explanation. It is a reproducible way to see which answers are settled, which remain fragile, and which measured sources move them.

2026-06-02 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

予測されたダイナミクスは物理世界に存在できますか?

予測物理 AI システムは状態ロールアウト、アクション チャンク、潜在計画を出力しますが、二乗平均平方根誤差 (RMSE) が低いということは、特定の提案が物理的に実行可能であることを意味するものではありません。物理的な許容性を予測制御インターフェイスとして定式化します。実行前に、デコードされた提案が候補ダイナミクスとして扱われ、運動学的、動的、および直接合成ホライズン条件を使用して評価されます。合格はタスクの成功を証明するものではありません。拒否は、指定された物理エンベロープの違反を識別し、コンポーネント レベルの理由を示します。 Hugging Face LeRobot PushT では、制御された改ざんにより、ワンステップ予測 RMSE と標準化されたダイナミクス残差が受信者動作特性曲線下領域 (AUC) 0.982 および 0.972 に達し、運動学のみの条件が AUC 0.592 に達し、フルゲートが条件レベルの帰属で AUC 0.957 に達することが示されています。リプレイベースの介入実験では、残差ベースのフィルターと完全な物理的許容ゲートにより、無効な提案の 87 ~ 89% が防止され、平均進行状況が 0.998 近くに維持されます。

原文 (English)

Can Predicted Dynamics Exist in the Physical World?

Predictive Physical AI systems output state rollouts, action chunks, and latent plans, yet a low root-mean-square error (RMSE) does not imply that a particular proposal is physically executable. We formulate physical admissibility as a prediction-control interface: before execution, a decoded proposal is treated as candidate dynamics and evaluated using kinematic, dynamic, and direct-to-composed horizon conditions. Passing is not a certificate of task success; rejection identifies violation of the specified physical envelope and gives a component-level reason. On Hugging Face LeRobot PushT, controlled falsification shows that one-step prediction-RMSE and standardized dynamics residuals reach area under the receiver operating characteristic curve (AUC) 0.982 and 0.972, kinematic-only conditions reach AUC 0.592, and the full gate reaches AUC 0.957 with condition-level attribution. In replay-based intervention experiments, residual-based filters and the full physical-admissibility gate prevent 87-$89% of invalid proposals while preserving mean progress near 0.998.

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体ビジネス/資金調達

SUPREME: 再現可能な画像非学習手法評価のためのマルチ GPU フレームワーク

機械の非学習では、最初から再トレーニングすることなく、トレーニングされたモデルから特定のトレーニング データの影響が除去されます。アンラーニング手法を評価するには、複数のシードにわたってトレーニング、アンラーニング、評価を繰り返す必要があり、計算コストがかかります。私たちの知る限り、既存の画像分類非学習フレームワークは単一の GPU 上で実行されるため、妥当な時間内に評価できるシードの数が制限されます。これらのステージを複数の GPU に分散するオープンソース フレームワークである SUPREME を紹介します。 SUPREME は 3 つの貢献を行っています。新しいメソッド、メトリック、モデル、シナリオを追加するためのレジストリ ベースの設計です。複数のアクセラレータと高精度モードをサポートするマルチ GPU アーキテクチャ。そして、10 個のシードにわたるフルクラスおよびランダム サンプルのアンラーニングのもとで、ResNet18 と ViT を使用したピンの顔認識のデモンストレーションです。このフレームワークは https://github.com/pedroandreou/supreme-unlearning で入手できます。

原文 (English)

SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation

Machine unlearning removes the influence of specific training data from a trained model without retraining it from scratch. Evaluating an unlearning method requires repeating training, unlearning, and evaluation across multiple seeds, which is computationally expensive. To our knowledge, existing image classification unlearning frameworks run on a single GPU, which limits how many seeds can be evaluated in reasonable time. We introduce SUPREME, an open-source framework that distributes these stages across multiple GPUs. SUPREME makes three contributions: a registry-based design for adding new methods, metrics, models, and scenarios; a multi-GPU architecture supporting multiple accelerators and precision modes; and a demonstration on Pins Face Recognition using ResNet18 and ViT under full-class and random-sample unlearning across ten seeds. The framework is available at https://github.com/pedroandreou/supreme-unlearning.

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成エージェントハードウェア/半導体

CodeCytos: コード拡張されたエージェント アクション スペースによる AI 支援の空間分子イメージング解析

従来の組織画像解析ソフトウェアは、セグメンテーション、基本的な形態学的特徴抽出、空間組織解析などの細胞解析の基本的な機能を提供します。ただし、これらのツールは手動介入が必要なことが多く、コード駆動の自動化とうまく統合されていないため、複雑な空間組織研究の効率と拡張性が制限されます。さらに、通常、事前に実装された空間セルラー機能の固定セットのみをサポートするため、カスタム分析の柔軟性は限られています。これらの制限に対処するために、我々はコードベースの推論エージェント フレームワークである CodeCytos を提案します。これは、自動化とカスタマイズを改善するために空間分子イメージング データとの動的でプログラム可能な相互作用を可能にします。 CodeCytos は、カスタムの空間細胞特徴の探索を合理化し、多様な研究ニーズに適応するように設計されています。私たちは、前頭葉皮質、非小細胞肺がん、膵臓、扁桃腺という異なる組織タイプから専門家が厳選した 4 つのデータセットに関するケーススタディを通じて、その有用性を実証します。私たちは、現実的な最小限のプロンプト設定の下で CodeCytos を評価します。この設定では、生物科学者がタスク固有の指示や空間細胞解析に関するコンテキスト情報なしで簡単な質問をし、強力なコーディング機能を備えた複数の LLM バックボーンをベンチマークします。さらに、カスタマイズされた、ドメインに依存しない少数ショットのコンテキスト内コーディング推論の例 (空間解析ドメイン外でランダムにサンプリングされたデモンストレーション) を組み込むことで、コストのかかる専門家が作成したドメイン内デモンストレーションを必要とせずに、パフォーマンスを大幅に向上できることを示します。全体として、CodeCytos はベースラインのアプローチよりも優れており、空間分子イメージングにおけるカスタム特徴探索を支援し、バイオマーカーの発見を加速するコードアクション エージェントの可能性を強調しています。

原文 (English)

CodeCytos: AI-assisted spatial molecular imaging analysis via code-augmented agent action space

Conventional tissue image analysis software provides foundational capabilities for cellular analysis, including segmentation, basic morphological feature extraction, and spatial organization analysis. However, these tools often require manual intervention and are not well integrated with code-driven automation, limiting efficiency and scalability for complex spatial tissue studies. In addition, they offer limited flexibility for custom analyses, as they typically support only a fixed set of pre-implemented spatial cellular features. To address these limitations, we propose CodeCytos, a coding-based reasoning agent framework that enables dynamic, programmable interaction with spatial molecular imaging data to improve automation and customization. CodeCytos is designed to streamline the exploration of custom spatial cellular features and adapt to diverse research needs. We demonstrate its utility through case studies on four expert-curated datasets from distinct tissue types: frontal cortex, non-small-cell lung cancer, pancreas, and tonsil. We evaluate CodeCytos under a realistic minimal prompt setting, where bioscientists pose simple questions without task-specific instructions or contextual information about spatial cellular analysis, and benchmark multiple LLM backbones with strong coding capabilities. We further show that incorporating tailored, domain-agnostic few-shot in-context coding-reasoning examples (randomly sampled demonstrations outside the spatial analysis domain) can substantially improve performance without requiring costly, expert-crafted in-domain demonstrations. Overall, CodeCytos outperforms baseline approaches, highlighting the potential of code-action agents to assist with custom feature exploration in spatial molecular imaging and to accelerate biomarker discovery.

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

言語学を意識した歪みのないLLM透かし入れ

透かしは、品質を低下させたり、検証をモデルプロバイダーに限定したりすることなく、言語モデルの出力を識別する必要があります。多言語展開では、透かしの証拠が自然に入る可能性がある形態、セグメンテーション、スクリプトが変更されるため、これがさらに困難になります。 LUNA は、標準的なランダム キー モデルの下で、モデルフリーの検出と単一トークンの非歪みを組み合わせた言語適応型ウォーターマークです。 LUNA は、外部コーパス内の品詞コンテキストから正規化された次のタグのエントロピーを推定し、それを使用して歪みのないバイナリ トーナメント サンプラーの深さを設定します。検出器は、テキスト、トークナイザー、タガー、および秘密キーから同じスケジュールを再構築します。私たちは、類型的に多様な 6 つの言語と 2 つのドメインを 8 つの主要なベースラインに対して評価します。 LUNA は、12 の設定全体で 0.9959 の AUROC と、最小平均絶対中央値パープレキシティ シフト 0.045 を達成しました。その 95% ブートストラップ間隔 [0.022, 0.073] は、すべてのベースライン間隔を下回っています。 LUNA は、Self-BLEU、Distinct-1、サプライズ、およびエントロピー シフトの最低の平均値も記録します。これは、大部分の設定で AUROC > 0.99 と 0.1 未満の絶対中央値パープレキシティ シフトを同時に達成する唯一の方法であり、12 設定中 9 でこの領域に到達しますが、2 を超えるベースラインはこの領域に到達しません。私たちのコードは、https://github.com/Shinwoo-Park/luna_watermark で入手できます。

原文 (English)

Linguistics-Aware Non-Distortionary LLM Watermarking

Watermarking should identify language-model output without degrading quality or limiting verification to the model provider. Multilingual deployment makes this harder because morphology, segmentation, and script change where watermark evidence can enter naturally. We introduce LUNA, a linguistically adaptive watermark that combines model-free detection with single-token non-distortion under the standard random-key model. LUNA estimates normalized next-tag entropy from part-of-speech contexts in an external corpus and uses it to set the depth of a non-distortionary binary tournament sampler; the detector reconstructs the same schedule from text, a tokenizer, a tagger, and a secret key. We evaluate six typologically diverse languages and two domains against eight primary baselines. LUNA attains an AUROC of 0.9959 and the lowest mean absolute median perplexity shift of 0.045 across the twelve settings; its 95% bootstrap interval [0.022, 0.073] lies below all baseline intervals. LUNA also records the lowest mean Self-BLEU, Distinct-1, surprisal, and entropy shifts. It is the only method that simultaneously achieves AUROC > 0.99 and an absolute median perplexity shift below 0.1 in a majority of settings, reaching this regime in 9 of the 12 settings while no baseline reaches it in more than 2. Our code is available at: https://github.com/Shinwoo-Park/luna_watermark

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

EPIC: 拡散言語モデルの CFG 制約の下での効率的な並列推論

言語モデルの出力を制御することは、構造の妥当性、信頼性、下流の使いやすさを確保するために不可欠であり、普及言語モデルも例外ではありません。拡散言語モデルのデコードにおける最近の進歩により、出力制御は通常の制約を超えて文脈自由文法 (CFG) 制約まで拡張されました。ただし、既存の方法は、制約のないデコードに比べて最大 4 倍遅くなる可能性があります。さらに重要なことは、自己回帰モデルに対する拡散言語モデルの重要な利点の 1 つである並列デコードが大幅に損なわれてしまうことです。この速度の低下は、逐次的な妥当性チェックによって並列生成中に重大なオーバーヘッドが発生するために発生します。我々は、この制限に対処する効率的な CFG 制約付きデコード フレームワーク EPIC を提案します。私たちの方法は、字句メモ化、決定論的オートマトンの代わりにアーリースタイルの解析を使用した検証、および並列コミットのための緩和された互換性のあるサブセット選択を組み合わせることで、デコード効率を向上させます。これにより、繰り返しの字句解析と検証のオーバーヘッドが削減され、複数の互換性のあるトークンを一緒にコミットできるようになります。 4 つのモデルを使用した 3 つのベンチマークの実験では、既存の CFG 制約付きデコード方法と比較して、私たちの方法が推論時間を最大 67.5% 削減し、追加のオーバーヘッドを最大 90.5% 削減できることが示されています。私たちの実装は https://github.com/hyundong98/EPIC-Decoding.git で入手できます。

原文 (English)

EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models

Controlling language model outputs is essential for ensuring structural validity, reliability, and downstream usability, and diffusion language models are no exception. Recent advances in diffusion language model decoding have extended output control beyond regular constraints to context-free grammar (CFG) constraints. Existing methods, however, can be up to four times slower than unconstrained decoding. More importantly, they substantially diminish one of the key advantages of diffusion language models over autoregressive models, namely parallel decoding. This slowdown arises because sequential validity checking introduces significant overhead during parallel generation. We propose an efficient CFG-constrained decoding framework, EPIC, that addresses this limitation. Our method improves decoding efficiency by combining lexing memoization, validation using Earley-style parsing instead of deterministic automata, and relaxed compatible subset selection for parallel commit. It reduces repeated lexing and validation overhead while allowing multiple compatible tokens to be committed together. Experiments on three benchmarks using four models show that our method reduces inference time by up to 67.5% and decreases the additional overhead by up to 90.5% compared with existing CFG-constrained decoding methods. Our implementation is available at https://github.com/hyundong98/EPIC-Decoding.git .

2026-06-02 13:00 JSTarXiv cs.AI画像/動画生成ハードウェア/半導体

DASH: ガイダンス校正済みコンパクト拡散モデル用のデュアルブランチ スコア蒸留

クラス条件付き拡散モデルのパラメータ圧縮により、出力レベルの蒸留における未調査の限界が明らかになります。無条件スコア分岐は教師なしのままであり、分類器のない指導ギャップが生徒内で十分に決定されていないままになります。このギャップはノイズ除去ステップごとに増幅され、両方のブランチが同一の予測に向かって崩壊する縮退したソリューションを許容し、出力レベルのトレーニング損失が低いにもかかわらずガイダンスが無効になります。このペーパーでは、両方のスコア ブランチを独立して監視するデュアル ブランチ蒸留フレームワークである DASH を紹介します。DASH は、グラウンド トゥルース ノイズに対する条件付き予測を正規化するアンカー項を使用して、独立したブランチ制約を通じて各トレーニング サンプルのターゲット ブランチ出力を一意に指定します。このフレームワークにはさらに、教師が集中させたタイムステップごとの重要度のカリキュラムを凍結された事前学習として生徒にコピーする TIRT Transfer が導入されており、限られた蒸留予算内で再学習する必要がなくなります。 CIFAR-10 および CIFAR-100 での実験では、5.9 倍の圧縮により、50 ステップの DDIM サンプリングで教師の 4 FID ポイント以内の品質が維持され、ガイダンスの忠実度が十分に維持された状態で最初からトレーニングするよりも大幅に優れていることが実証されました。アブレーション研究では、無条件の監視が主な寄与であり、総蒸留利益の 60% 以上を占めることが確認されています。カリキュラムの移行とアンカーの正規化は相補的な利点をもたらし、ガイダンスを維持する圧縮にはデュアルブランチ制約が経験的に不可欠であることが実証されます。

原文 (English)

DASH: Dual-Branch Score Distillation for Guidance-Calibrated Compact Diffusion Models

Parameter compression of class-conditional diffusion models reveals an underexplored limitation in output-level distillation: the unconditional score branch remains unsupervised, leaving the classifier-free guidance gap underdetermined in the student. This gap, amplified at every denoising step, admits degenerate solutions where both branches collapse toward identical predictions, rendering guidance ineffective despite low output-level training loss. This paper introduces DASH, a dual-branch distillation framework that independently supervises both score branches, uniquely specifying target branch outputs for each training sample through independent branch constraints, with an anchor term regularising conditional predictions toward ground-truth noise. The framework further introduces TIRT Transfer, which copies the teacher's converged per-timestep importance curriculum into the student as a frozen prior, eliminating the need to relearn it within limited distillation budgets. Experiments on CIFAR-10 and CIFAR-100 demonstrate that 5.9x compression maintains quality within 4 FID points of the teacher at 50-step DDIM sampling, considerably outperforming training from scratch with guidance fidelity well preserved. Ablation studies confirm that unconditional supervision is the dominant contribution, accounting for over 60% of total distillation gain. Curriculum transfer and anchor regularisation provide complementary benefit, together validating dual-branch constraints as empirically essential for guidance-preserving compression.

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

A Fiber Criterion for Representation Identifiability in Supervised Learning

Supervised learning evaluates predictors through their input-output behavior. When a predictor is implemented as a composition $f=c\circ h$…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

HASTE: Hardware-Aware Dynamic Sparse Training for Large Output Spaces

Extreme multi-label classification (XMC) involves learning models over large output spaces with millions of labels, making the output layer…

2026-06-02 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

Crazyflow: An Accurate, GPU-Accelerated, Differentiable Drone Simulator in JAX

High-quality, large-scale synthetic data from simulations is becoming a cornerstone for pushing the capabilities of robot algorithms. While…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics

Frontier LLMs increasingly decide what a query attends to with a sparse-attention indexer that picks a few KV-cache blocks per query: atten…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise

Measuring the diversity of creative outputs is central to evaluating post-training mode collapse, comparing decoding strategies, and quanti…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

The Role of Ambiguity in Error Prediction via Uncertainty Quantification

The task of Error Prediction, namely predicting whether a model output is correct, is commonly tackled with Uncertainty Quantification (UQ)…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

On the Theoretical Limitations of Embedding-based Link Prediction

Neural networks often map low-dimensional embeddings to high-dimensional output spaces. Usually, the output layer is linear, which can crea…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Query Circuits: Explaining How Language Models Answer User Prompts

Explaining why a language model produces a particular output requires local, input-level explanations. Existing methods uncover global capa…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Efficient LLM Moderation with Multi-Layer Latent Prototypes

Although modern LLMs are aligned with human values during post-training, robust moderation remains essential to prevent harmful outputs at…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達

Erased but Not Forgotten: How Backdoors Compromise Concept Erasure

The expansion of text-to-image diffusion models has raised concerns about harmful outputs, from fabricated depictions of public figures to…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

End-to-End Deep Learning for Predicting Metric Space-Valued Outputs

Many modern applications involve predicting structured, non-Euclidean outputs such as probability distributions, networks, and symmetric po…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Optimizing Diversity and Quality through Base-Aligned Model Collaboration

Alignment has greatly improved large language models (LLMs)' output quality at the cost of diversity, yielding highly similar outputs acros…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

How Much Progress Has There Been in NVIDIA Datacenter GPUs?

As the role of modern Graphics Processing Units (GPUs) becomes increasingly essential for several computing tasks, analyzing their past and…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

GUDA: Counterfactual Group-wise Training Data Attribution for Diffusion Models via Unlearning

Training-data attribution for vision generative models aims to identify which training data influenced a given output. While most methods s…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

Failure of contextual invariance in large language models

Standard evaluation practices assume that large language model (LLM) outputs are stable when prompts are embedded in contextually equivalen…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

FlowPlace: Flow Matching for Chip Placement

Chip placement plays an important role in physical design. While generative models like diffusion models offer promising learning-based sol…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

How Can Reinforcement Learning Achieve Expert-level Placement?

Chip placement is a critical step in physical design. While reinforcement learning (RL)-based methods have recently emerged, their training…

2026-06-02 13:00 JSTarXiv cs.AIハードウェア/半導体

Physics-Guided Geometric Diffusion for Macro Placement Generation

Macro placement is a pivotal stage in VLSI physical design, fundamentally determining the overall chip performance. Recent data-driven plac…

2026-06-02 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

信号としてのアノテーターの位置性: 反自閉症障害検出のための心理測定的重み付け

大規模言語モデル (LLM) は、視点を増幅または抑制する可能性がある意思決定タスクでますます使用されており、自閉症コミュニティに影響を与える一か八かの状況で懸念が生じています。これまでの研究では、LLM における障害に関連したバイアスが特定されていますが、LLM が障害者主義をどのように概念化し、テキスト内でそれを検出するのかは依然として不明です。私たちは、アノテーターの位置性に基づいて心理測定的に重み付けされたコミュニティに近いグランドトゥルースを備えた、反自閉症障害者言語を対象としたバイアスを意識した評価フレームワークを導入します。この枠組みは、自閉症や自閉症受容者の観点を大幅かつ一貫して軽視する従来の多数決の集計よりも厳格な基準を構成しています。評価手段が隠蔽されている場合、LLM は頻繁に有害な出力を生成し、コミュニティで再利用された言語を障害者主義と誤ってラベル付けし、自閉症の人に対してより否定的な態度を表明することがわかりました。私たちのエラー分析により、モデルは話者の身元や、その言語がグループ内の連帯を促進するのか、それともグループ外に害を及ぼすのかなどの文脈上の要因ではなく、表面レベルのキーワードの一致に依存していることが明らかになりました。

原文 (English)

Algorithmic Fragility and Persona Bias in LLM-Generated Autistic Communication

Safety alignment reduces explicitly harmful outputs but inadvertently encodes a sanitized, neuronormative representation of marginalized communication. We investigate this encoding using a dual-persona rewrite paradigm, prompting ten large language models (LLMs) to rewrite naturally occurring autistic discourse from either an autistic or neurotypical persona. We uncover autistic-persona rewrites diverge significantly more in lexical form and affective register than neurotypical rewrites, despite equivalent semantic similarity. Furthermore, most models collapse cross-persona generations into near-identical outputs. To uncover the mechanisms behind this generative breakdown, we introduce a multi-agent qualitative analysis framework. Our results reveal systemic output erasure, stereotyped hallucination, and task-evasive meta-commentary are pervasive failure modes for this task that cluster by alignment strategy rather than parameter scale. Finally, our targeted comparison with autistic human annotators demonstrates that community-insider knowledge produces systematic label reversals relative to LLM classifications. Our findings indicate that current alignment training causes persona-specific generative breakdown visible only through qualitative analysis, confirming a deep representational gap that prompt engineering cannot resolve.

2026-06-02 07:45 JSTITmedia AI+エージェントハードウェア/半導体

NVIDIAの“狐”は工場自律管理AIエージェント、台湾メーカーが導入効果を確認

NVIDIAは、工場を自律的に管理するAIエージェントのレファレンスデザイン「NVIDIA Factory Operations Blueprint(FOX)」を発表した。FOXを用いれば、工場内のさまざまなデータをリアルタイムに監視/分析し、複数のAIエージェントと機器を連携…

2026-06-02 06:45 JSTITmedia AI+ハードウェア/半導体

NVIDIAの「NemoClaw」でエッジAIを統合管理、アドバンテックが「WEDA」を発表

アドバンテックは、パートナー向けイベント「2026 Advantech World Partner Conference(WPC)」において、エッジAIの開発から導入、運用までを統合的に管理するソリューション「WEDA」について説明した。

2026-06-02 06:35 JSTTechCrunch AIエージェントハードウェア/半導体

Nvidia chases $200B CPU market with AI agent PCs from Microsoft, Dell, and HP

If Nvidia has cracked a way to bring AI agents easily, safely, and usefully to the masses, it could — and should — be big.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

LLM-FACETS: LLM の透明性と説明責任を評価するためのプライバシー保護フレームワーク

大規模言語モデルの出力が事実に基づいており、認識論的に調整されており、方法論的に再現可能であるかどうかを評価することは、責任ある AI 導入の前提条件です。しかし、LLM の監査は、技術者以外の専門家にとってはアクセスできないままです。既存のツールにはプログラミングの専門知識と簡単ではない環境セットアップが必要であり、クラウドでホストされるプラットフォームは評価データを外部サービスに送信するため、AI の監視に法的責任を負うドメインの専門家やコンプライアンス担当者にとって障壁が生じています。 LLM-FACETS (LLM FActuality Cross-EvaluTion System) を紹介します。これは、ブラウザからアクセス可能なインターフェイスとプラグイン アーキテクチャを備えたオープンソース フレームワークで、EU AI 法と NIST AI リスク管理フレームワークで特定されているステークホルダーのカテゴリを反映する 3 つの実践者プロファイル (技術専門家、ドメイン専門家、コンプライアンス担当者) を中心に構造化されています。このアーキテクチャでは、データ フローが明示的になります。決定論的メトリクス (BLEU、ROUGE、BERTScore) は、アウトバウンド送信なしで完全に自己ホスト型サーバー内で実行されます。 LLM 判定メトリクスは外部 API に明示的に接続し、ユーザーは資格情報の完全な制御を保持します。このフレームワークは、認識上の不確実性に対するトークンレベルの対数確率の視覚化、裁判官のバイアスを軽減するための複数裁判官のコンセンサス、幻覚を検出して位置を特定するための RAG トライアド メトリクス (忠実度、回答の関連性、コンテキストの関連性) の 3 つのメカニズムを通じて透明性を運用します。プラグイン アーキテクチャにより、評価パイプラインを変更せずに、新しいメトリクスやデータセットを統合できます。オープンソースの実装により、同じプロパティを対象とする複数の指標にわたるクロスチェックが可能になり、再現性が確保され、評価対象のシステムを構築するチームから AI の説明責任が切り離されます。正規の参照ライブラリに対する 18 のメトリック実装の相互検証を通じてフレームワークを検証します。

原文 (English)

LLM-FACETS: A Privacy-Preserving Framework for Evaluating LLM Transparency and Accountability

Assessing whether Large Language Models outputs are factually grounded, epistemically calibrated, and methodologically reproducible is a prerequisite for responsible AI deployment. Yet auditing LLMs remains inaccessible to non-technical practitioners: existing tools require programming expertise and non-trivial environment setup, and cloud-hosted platforms transmit evaluation data to external services, creating barriers for domain experts and compliance officers legally responsible for AI oversight. We introduce LLM-FACETS (LLM FActuality Cross-EvaluaTion System): an open-source framework with a browser-accessible interface and a plugin architecture, structured around three practitioner profiles (technical experts, domain experts, compliance officers) that mirror the stakeholder categories identified in the EU AI Act and the NIST AI Risk Management Framework. The architecture makes data flows explicit: deterministic metrics (BLEU, ROUGE, BERTScore) run entirely within the self-hosted server with no outbound transmission; LLM-judge metrics contact external APIs explicitly, with users retaining full credential control. The framework operationalizes transparency through three mechanisms: token-level log-probability visualization for epistemic uncertainty, multi-judge consensus to mitigate judge bias, and RAG Triad metrics (Faithfulness, Answer Relevance, Context Relevance) to detect and localize hallucinations. A plugin architecture allows any new metric or dataset to be integrated without modifying the evaluation pipeline. The open-source implementation enables cross-checking across multiple metrics targeting the same property, ensuring reproducibility and decoupling AI accountability from the teams building the systems assessed. We verify the framework through cross-validation of 18 metric implementations against canonical reference libraries.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

LLM が一貫して間違っていることを学習するとき: 合成欺瞞の線形表現に関するマルチモデル研究

モデルが意図的に偽の出力を生成しながら正確な内部表現を維持する欺瞞的な調整は、依然として AI の安全性における中心的な課題です。戦略的欺瞞が長期的な主な懸念事項である一方で、不正解に対する直接最適化によって引き起こされる合成的不正は、学習された欺瞞の表現基盤を研究するための制御されたテストベッドを提供します。 5 つのトランスフォーマー モデル (Pythia-1.4B、Gemma-2-2B/9B、Qwen2.5-7B、Llama-3.1-8B) の正直なバリアントと欺瞞的なバリアントが、同じ質問分布に対して LoRA を使用して微調整されるマルチモデル パラダイムを導入します。平均プールされた隠れ状態で訓練された線形プローブは、4 つのアーキテクチャのレイヤー 1 ~ 3 でほぼ完璧な AUC (0.99 以上) で合成不正を検出しますが、Pythia-1.4B はピークの 0.705 に達します。ロジスティック回帰プローブは一貫して MLP プローブと一致するかそれを上回っており、線形表現仮説を裏付けています。 TruthfulQA でトレーニングされたプローブは、保留された MMLU 被験者に対してほぼゼロの損失 (デルタ AUC 約 0) で一般化します。後期層の表現はガウス ノイズに対する強い堅牢性を示し、Gemma-2 モデルは優れた安定性を示します。フィッシャー判別比、有効ランク、重心幾何学、方向安定性、クロスドメインアライメント、およびキャリブレーション (ECE) の機構分析により、Pythia/Llama/Qwen における表現崩壊と Gemma-2 における高次元保存という 2 つの状況が明らかになります。すべてのモデルにわたって、不正の方向はより深い層に徐々に統合され、層 1 ~ 4 で最適なキャリブレーション (Pythia を除く ECE が 0.01 未満) が達成されます。これらの結果は、堅牢でドメイン不変の不正表現が、適度な教師付き微調整によって急速に定着する可能性があり、アクティベーションベースのモニタリングに影響を与えることを示しています。

原文 (English)

When LLMs Learn to Be Consistently Wrong: A Multi-Model Study of Linear Representations of Synthetic Deception

Deceptive alignment, in which models maintain accurate internal representations while deliberately producing false outputs, remains a central challenge in AI safety. While strategic deception is the primary long-term concern, synthetic dishonesty - induced via direct optimization on incorrect answers - provides a controlled testbed for studying the representational basis of learned deception. We introduce a multi-model paradigm in which honest and deceptive variants of five transformer models (Pythia-1.4B, Gemma-2-2B/9B, Qwen2.5-7B, Llama-3.1-8B) are fine-tuned using LoRA on the same question distribution. Linear probes trained on mean-pooled hidden states detect synthetic dishonesty with near-perfect AUC (greater than or equal to 0.99) as early as layers 1-3 in four architectures, while Pythia-1.4B reaches a peak of 0.705. Logistic regression probes consistently match or outperform MLP probes, supporting the Linear Representation Hypothesis. Probes trained on TruthfulQA generalize with near-zero loss (Delta AUC approx. 0) to held-out MMLU subjects. Late-layer representations show strong robustness to Gaussian noise, with Gemma-2 models exhibiting exceptional stability. Mechanistic analysis of Fisher Discriminant Ratio, effective rank, centroid geometry, directional stability, cross-domain alignment, and calibration (ECE) reveals two regimes: representational collapse in Pythia/Llama/Qwen versus high-dimensional preservation in Gemma-2. Across all models, the dishonesty direction consolidates progressively in deeper layers, with optimal calibration (ECE less than 0.01 except Pythia) achievable in layers 1-4. These results demonstrate that robust, domain-invariant dishonesty representations can be rapidly entrenched via modest supervised fine-tuning, with implications for activation-based monitoring.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

大規模言語モデルの調整のための差分プライベート設定データ合成

好みの調整は、大規模言語モデル (LLM) の出力が人間の価値観と一致していることを確認するための、トレーニング後の重要なステップです。ただし、これらのデータセットには機密性の高いユーザー プロンプトや人間の判断が含まれることが多いため、実際の人間の嗜好データのポストトレーニングではプライバシーの懸念が生じます。これに対処するために、私たちは、差分プライベート (DP) 合成嗜好データを生成してプライバシーを保護した嗜好の調整を可能にする新しいアルゴリズムである DPPrefSyn を提案します。 DPPrefSyn は、Bradley-Terry の選好モデルと、ペアごとの人間の選好データの固有の幾何学的構造に基づいた原則に基づいたフレームワークです。まず、正式な差分プライバシー保証を備えたプライベート データから基礎となる嗜好モデルを学習し、次に学習したモデルを公開プロンプトとともに活用して、高品質の嗜好データを合成します。クラスターごとの報酬モデルの共有線形構造を利用して、プライベート データセット内の異種の人間の好みを効果的にキャプチャし、DP 主成分分析 (DP-PCA) を活用して学習精度を向上させます。広範な実験結果は、DPPrefSyn が強力な DP 保証のもとで競争力のあるアライメント性能を達成することを実証しています。これらの発見は、幅広いアプリケーションにわたってプライバシーを保護しながら好みを調整するための実用的な代替手段として、合成好みデータの可能性を浮き彫りにしています。私たちの知る限り、これは LLM アライメント用の DP 合成選好データを生成する最初の作業です。私たちのコードは https://github.com/gfengyu/Differentially-Private-Preference-Data-Synthesis で入手できます。

原文 (English)

Differentially Private Preference Data Synthesis for Large Language Model Alignment

Preference alignment is a crucial post-training step for large language models (LLMs) to ensure their outputs align with human values. However, post-training on real human preference data raises privacy concerns, as these datasets often contain sensitive user prompts and human judgments. To address this, we propose DPPrefSyn, a novel algorithm for generating differentially private (DP) synthetic preference data to enable privacy-preserving preference alignment. DPPrefSyn is a principled framework grounded in the Bradley-Terry preference model and the intrinsic geometric structure of pairwise human preference data. It first learns an underlying preference model from private data with formal differential privacy guarantees, and then leverages the learned model together with public prompts to synthesize high-quality preference data. It exploits the shared linear structure of per-cluster reward models to effectively capture heterogeneous human preferences in private datasets, and leverages DP Principal Component Analysis (DP-PCA) to improve learning accuracy. Extensive experimental results demonstrate that DPPrefSyn achieves competitive alignment performance under strong DP guarantees. These findings highlight the potential of synthetic preference data as a practical alternative for privacy-preserving preference alignment across a broad range of applications. To the best of our knowledge, this is the first work to generate DP synthetic preference data for LLM alignment. Our code is available at https://github.com/gfengyu/Differentially-Private-Preference-Data-Synthesis.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

微調整により言語モデルの情報伝達が改善される

ファインチューニングは大規模な言語モデルの不確実性と多様性を軽減すると考えられていますが、既存の分析では主要な交絡因子である出力の長さが見落とされているため、世代のロールアウト全体にわたって不確実性がどのように分布しているかを把握できていません。これに対処するために、私たちは言語生成をツリーの観点から見る尺度であるキャノピー エントロピー ($\mathrm{CE}^\star$) を提案します。ここで、「キャノピー」はすべての可能なロールアウトのスペースを表し、$\mathrm{CE}^\star$ が生成スペースの有効なサイズを自然に定量化します。 $\mathrm{CE}^\star$ は、出力長 $N$ と生成されたシーケンス $Y_{1:N}$ の両方の不確実性を共同で捕捉します。実際、それが合計シャノン エントロピー $H(N, Y_{1:N}\mid X)$ に等しいことが示されます。ここで、$X$ はプロンプトを示します。この定式化により、長さとエントロピーの相関項 $\rho(N, r_N)$ ($r_N$ はエントロピー レート) を含む解釈可能なメトリクスが得られ、より長い出力がトークンごとに情報量が多いか少ないかを示すことで情報伝達効率を定量化します。経験的に、タスクとモデルファミリー全体で、総エントロピーが減少した場合でも、微調整されたモデルは一貫して強い正の相関 $\rho(N, r_N)$ を示すことがわかりました。さらに、モデルファミリー、タスク、プロンプト、および出力長の効果を制御した後、微調整によりエントロピー率と意味の多様性の間の相関強度がほぼ3倍になることがわかり、調整されたモデルがトークンの不確実性をより効率的に意味の多様性に変換することを示唆しています。全体として、これらの結果は、微調整が単に不確実性を低減するだけでなく、不確実性をより有益で意味的に意味のある世代に根本的に再編成することを示しています。私たちのコードは https://github.com/WeiyiTian/canopy-entropy で入手できます。

原文 (English)

Fine-Tuning Improves Information Conveyance in Language Models

Fine-tuning is often believed to reduce uncertainty and diversity in large language models, but existing analyses overlook output length, a key confounder, and therefore fail to capture how uncertainty is distributed across an entire generation rollout. To address this, we propose Canopy Entropy ($\mathrm{CE}^\star$), a measure that views language generation from a tree perspective, where ``canopy'' represents the space of all possible rollouts, making $\mathrm{CE}^\star$ naturally quantify the effective size of the generation space. $\mathrm{CE}^\star$ jointly captures uncertainty in both the output length $N$ and the generated sequence $Y_{1:N}$ -- indeed, we show that it equals to total Shannon entropy $H(N, Y_{1:N}\mid X)$, where $X$ denotes the prompt. This formulation yields interpretable metrics, including a length-entropy correlation term $\rho(N, r_N)$, where $r_N$ is the entropy rate, quantifying information conveyance efficiency by indicating whether longer outputs are more or less informative per token. Empirically, across tasks and model families, we find that fine-tuned models consistently exhibit stronger positive correlation $\rho(N, r_N)$, even when total entropy decreases. Furthermore, after controlling for model family, task, prompt, and output-length effects, we find that fine-tuning nearly triples the correlation strength between entropy rate and semantic diversity, suggesting that aligned models convert token uncertainty into semantic diversity more efficiently. Overall, these results demonstrate that fine-tuning does not simply reduce uncertainty, but fundamentally reorganizes it into more informative and semantically meaningful generations. Our code is available at https://github.com/WeiyiTian/canopy-entropy.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AI画像/動画生成ハードウェア/半導体研究/論文

早期算数教育における視覚表現生成のためのテキストから画像へのモデルのベンチマークと強化

AI システムは、教育コンテンツの作成をサポートするためにますます使用されていますが、教えようとしている教育概念を忠実に表す出力を生成できるかどうかは依然として不明です。そこで、方程式からビジュアルへの生成を導入します。このタスクは、従来の画像生成とは対照的に、数値構造と関係構造を正確に保持しながら、算術方程式から教育的に意味のあるビジュアルを生成する必要があります。教師へのインタビューと教材の分析に基づいて、教育学的に根拠のある 4 つの視覚タイプにわたるベンチマークである E2V ベンチと、視覚の正しさを評価するための自動指標を構築します。私たちの評価により、最近のテキストから画像への (T2I) モデルはこのタスクで頻繁に失敗し、不正確なオブジェクト数と壊れたリレーショナル構造によってエラーが支配されることが明らかになりました。これに基づいて、ベンチマークに基づいた強化戦略を検討します。これらの戦略は代表的なモデルを改善しますが、残りのギャップには将来の T2I モデルにおけるより強力な数値的および関係的根拠が必要です。

原文 (English)

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education

AI systems are increasingly used to support educational content creation, yet it remains unclear whether they can generate outputs that faithfully represent the pedagogical concepts they are intended to teach. Thus, we introduce equation-to-visual generation, a task that, in contrast to conventional image generation, requires producing pedagogically meaningful visuals from arithmetic equations while precisely preserving their numerical and relational structure. Informed by interviews with teachers and an analysis of educational materials, we construct E2V-Bench, a benchmark spanning four pedagogically grounded visual types, along with automatic metrics for evaluating visual correctness. Our evaluation reveals that recent text-to-image (T2I) models frequently fail on this task, with errors dominated by incorrect object counts and broken relational structure. Building on this, we explore benchmark-guided enhancement strategies. These strategies improve representative models, while the remaining gap calls for stronger numerical and relational grounding in future T2I models.

2026-06-01 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

GPU Forecaster: カーネル ランタイム最適化の選択的サロゲートとしての言語モデル

GPU カーネルは最新の深層学習の主力であり、(進化的検索またはコーディング エージェントを介して) GPU カーネルを最適化するには、通常、ターゲット ハードウェアでの繰り返しの測定が必要です。これらの測定はカーネル検索に必要なグラウンドトゥルース信号を提供しますが、カーネルの各評価にはコンパイルと GPU での繰り返し実行が必要なため、コストがかかります。 LLM 推論の改善により、新しいカーネルの作成コストが削減され、LLM 駆動の検索が大規模な検索予算に拡張されるため、デバイス上の評価がボトルネックになります。これに対処するために、提案されたカーネルのパフォーマンスを予測することにより、LLM がカーネル評価の選択的な GPU サロゲートとして機能する方法を研究します。有用なサロゲートは正確である必要があり、いつ間違っている可能性があるかを認識して GPU に任せることにより、選択的である必要があります。サロゲートを評価するために、その予測が正確で、調整されており、限られた GPU 測定予算の下で高速カーネルを回復するために実際に役立つかどうかを測定します。次に、強化学習によって予測精度と信頼度の調整が向上するかどうかを研究します。私たちの実験は、LLM が相対的なカーネル パフォーマンスを正確に予測できること、強化学習を通じて LLM の有用性を向上できることを示しています。カーネル検索内でサロゲートを使用すると、同じ GPU 評価予算の下で数倍の数の候補を検索で考慮できるため、同等の予算のベースラインよりも高速なカーネルを見つけることができます。これらの結果は、LLM が単に検索用のカーネル ジェネレーターとしてではなく、GPU の仮想モデルとして機能することにより、カーネルの最適化においてより広範な役割を果たすことができることを示唆しています。

原文 (English)

GPU Forecasters: Language Models as Selective Surrogates for Kernel Runtime Optimization

GPU kernels are the workhorse of modern deep learning, and optimizing them (via evolutionary search or coding agents) usually requires repeated measurement on target hardware. While these measurements provide the ground-truth signal necessary for kernel search, they are costly, because each evaluation of a kernel requires compilation and repeated execution on a GPU. As improvements in LLM inference reduce the cost of writing novel kernels and LLM-driven searches scale to large search budgets, on-device evaluation becomes a bottleneck. To address this, we study how LLMs can serve as selective GPU surrogates for kernel evaluation, by forecasting the performance of proposed kernels. A useful surrogate should be accurate, and it should be selective, by knowing when it could be wrong, and deferring to the GPU. To evaluate surrogates, we measure whether their forecasts are accurate, calibrated, and practically useful for recovering fast kernels under limited GPU-measurement budgets. Next, we study whether reinforcement learning can improve forecast accuracy and confidence calibration. Our experiments demonstrate that LLMs can accurately forecast relative kernel performance, that their utility can be improved through reinforcement learning. Used inside a kernel search, the surrogate lets the search consider several times as many candidates under the same GPU evaluation budget, and that leads to finding faster kernels than an equal-budget baseline. These results suggest that LLMs can play a broader role in kernel optimization, by acting as virtual models of a GPU rather than solely as kernel generators for search.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

項目応答理論による LLM-as-a-Judge の信頼性の診断

LLM-as-a-Judge は自動評価で広く使用されていますが、既存の検証手法は主に観察された出力のレベルで動作し、LLM ジャッジ自体が安定した信頼できる測定手段として機能するかどうかについての洞察は限られています。この制限に対処するために、項目応答理論 (IRT) に基づいた、裁判官としての LLM の信頼性を評価するための 2 段階の診断フレームワークを導入します。このフレームワークは、IRT の段階的応答モデル (GRM) を採用し、2 つの相補的な側面に沿って信頼性を形式化します: (1) 即時の変動下での測定動作の安定性として定義される本質的一貫性、および (2) 人間の品質評価との対応を捉える人間の整合性。私たちは、このフレームワークを使用して多様な LLM 裁判官を実証的に調査し、IRT-GRM を活用すると、体系的に判断を診断するための解釈可能なシグナルが得られることを示します。これらの信号は、LLM-as-a-Judge の信頼性を検証し、信頼性の低さの潜在的な原因を特定するための実践的なガイダンスを提供します。

原文 (English)

Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory

While LLM-as-a-Judge is widely used in automated evaluation, existing validation practices primarily operate at the level of observed outputs, offering limited insight into whether LLM judges themselves function as stable and reliable measurement instruments. To address this limitation, we introduce a two-phase diagnostic framework for assessing reliability of LLM-as-a-Judge, grounded in Item Response Theory (IRT). The framework adopts Graded Response Model (GRM) of IRT and formalizes reliability along two complementary dimensions: (1) intrinsic consistency, defined as the stability of measurement behavior under prompt variations, and (2) human alignment, capturing correspondence with human quality assessments. We empirically examine diverse LLM judges with this framework, and show that leveraging IRT-GRM yields interpretable signals for diagnosing judgments systematically. These signals provide practical guidance for verifying reliablity of LLM-as-a-Judge and identifying potential causes of unreliability.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

データに敏感なドメインの LLM 出力のニューロシンボリック検証 (拡張プレプリント)

一か八かのドメインに導入された LLM は、根本的な信頼性の課題に直面しています。幻覚、矛盾、プライバシーの脆弱性により、エラーが法的、財務的、または安全性に影響を及ぼす許容できないリスクが生じます。この論文では、LLM で生成されたコンテンツに補完的な保証を提供する、形式的記号手法とニューラル セマンティック分析を組み合わせたハイブリッド検証アーキテクチャを紹介します。このアーキテクチャでは、入力検証に論理的推論を採用し、完全性の特性を活用して、構造化された要件に対して決定可能な保証を提供します。出力検証では、埋め込みベースの意味論的類似性により、形式的な手法では表現力に欠ける文脈上の幻覚が検出されます。この分離は、並列のアクターベースのパイプラインで実現され、幻覚を生み出す分布バイアスを継承するプロンプトベースの自己検証アプローチの制限に対処します。提案されたアーキテクチャとタイプ認識検証方法は、Action Design Research によって開発された現実世界の医療機器損傷評価レポート システムである HAIMEDA を使用して検証されています。評価の結果、構造化エンティティの幻覚検出率は 83% 以上、セマンティック捏造の幻覚検出率は 72% 以上で、レポート作成時間が 30% 短縮されたことが示され、神経記号アーキテクチャがデータに敏感なドメインでの LLM 展開に原則に基づいた保護手段を提供できることが実証されました。

原文 (English)

Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)

LLMs deployed in high-stakes domains face fundamental reliability challenges: hallucinations, inconsistencies, and privacy vulnerabilities introduce unacceptable risks where errors carry legal, financial, or safety consequences. This paper presents a hybrid verification architecture combining formal symbolic methods with neural semantic analysis to provide complementary guarantees for LLM-generated content. This architecture employs logical reasoning for input verification, leveraging completeness properties to provide decidable guarantees on structured requirements. For output validation, embedding-based semantic similarity detects contextual hallucinations where formal methods lack expressiveness. This separation is realized in a parallel, actor-based pipeline, addressing limitations of prompt-based self-verification approaches, which inherit the distributional biases that produce hallucinations. The proposed architecture and type-aware verification method are validated with HAIMEDA, a real-world medical device damage assessment reporting system developed through Action Design Research. Evaluation shows hallucination detection rates of over 83% for structured entities and 72% for semantic fabrications, with a 30% reduction in report creation time, demonstrating that neuro-symbolic architectures can provide principled safeguards for LLM deployment in data-sensitive domains.

2026-06-01 13:00 JSTarXiv cs.AIハードウェア/半導体

変分ルーティング: 調整された専門家混合トランスフォーマーのためのスケーラブルなベイジアン フレームワーク

Foundation models are increasingly being deployed in contexts where understanding the uncertainty of their outputs is critical to ensuring responsible deployment. While Bayesian methods offer a principled approach to uncertainty quantification, their computational overhead renders their use impractical for training or inference at foundation model scale. State-of-the-art models achieve parameter counts in the trillions through carefully engineered sparsity including Mixture-of-Experts (MoE) layers. In this work, we demonstrate calibrated uncertainty at scale by introducing Variational Mixture-of-Experts Routing (VMoER), a structured Bayesian approach for modelling uncertainty in MoE layers. VMoER は、ベイジアン推論を、通常は決定論的なルーティング ネットワークによって行われる専門家選択段階に限定します。 We instantiate VMoER using two inference strategies: amortised variational inference over routing logits and inferring a temperature parameter for stochastic expert selection. Across fine-tuning tested foundation models, VMoER improves routing stability under noise by 38\%, reduces calibration error by 94\%, and increases out-of-distribution AUROC by 12\%, while incurring less than 1\% additional FLOPs.これらの結果は、VMoER が堅牢で不確実性を認識した基礎モデルへのスケーラブルな道を提供することを示唆しています。

原文 (English)

Variational Routing: A Scalable Bayesian Framework for Calibrated Mixture-of-Experts Transformers

Foundation models are increasingly being deployed in contexts where understanding the uncertainty of their outputs is critical to ensuring responsible deployment. While Bayesian methods offer a principled approach to uncertainty quantification, their computational overhead renders their use impractical for training or inference at foundation model scale. State-of-the-art models achieve parameter counts in the trillions through carefully engineered sparsity including Mixture-of-Experts (MoE) layers. In this work, we demonstrate calibrated uncertainty at scale by introducing Variational Mixture-of-Experts Routing (VMoER), a structured Bayesian approach for modelling uncertainty in MoE layers. VMoER confines Bayesian inference to the expert-selection stage which is typically done by a deterministic routing network. We instantiate VMoER using two inference strategies: amortised variational inference over routing logits and inferring a temperature parameter for stochastic expert selection. Across fine-tuning tested foundation models, VMoER improves routing stability under noise by 38\%, reduces calibration error by 94\%, and increases out-of-distribution AUROC by 12\%, while incurring less than 1\% additional FLOPs. These results suggest VMoER offers a scalable path toward robust and uncertainty-aware foundation models.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

REAL: 裁判官としての LLM のための回帰を意識した強化学習

大規模言語モデル (LLM) は、モデル出力に数値スコアを割り当てる自動評価器として導入されることが増えています。このパラダイムは、LLM-as-a-Judge として知られています。ただし、標準的な強化学習 (RL) 手法は通常、バイナリ報酬 (0 ~ 1 の精度など) に依存するため、回帰タスクに固有の順序構造が無視されます。たとえば、グランド トゥルースが 5 の場合、4 を予測する方が 1 を予測するよりも大幅に優れているということを認識できません。逆に、既存の回帰認識アプローチは多くの場合、教師あり微調整 (SFT) に限定されており、最適な推論パスを探索する能力が制限されています。このギャップを埋めるために、回帰報酬を最適化するように設計され、相関指標にも最適であることが証明されている原則に基づいた RL フレームワークである \textbf{REAL} (\underline{RE}gression-\underline{A}ware Reinforcement \underline{L}earning) を提案します。主要な技術的課題は、回帰の目的が明示的にポリシーに依存しているため、標準的なポリシー勾配手法が無効になることです。これに対処するために、一般化されたポリシー勾配推定器を採用します。これは、最適化を 2 つの相補的なコンポーネント (1) 思考連鎖 (CoT) 軌跡の探索、および (2) 最終スコアの回帰を意識した予測の改良に自然に分解します。モデル スケール (8B から 32B) にわたる広範な実験により、REAL が回帰対応 SFT ベースラインと標準 RL 手法の両方を常に上回っており、ドメイン外ベンチマークでの一般化が大幅に優れていることが実証されました。特に Qwen3-32B では、SFT ベースラインに対してピアソン相関 +8.40 およびスピアマン相関 +7.20、ベース モデルに対して +18.30/+11.20 のゲインを達成しています。これらの発見は、正確な LLM 評価のために回帰目標を RL 探索に統合することの重要な価値を強調しています。

原文 (English)

REAL: Regression-Aware Reinforcement Learning for LLM-as-a-Judge

Large language models (LLMs) are increasingly deployed as automated evaluators that assign numeric scores to model outputs, a paradigm known as LLM-as-a-Judge. However, standard Reinforcement Learning (RL) methods typically rely on binary rewards (e.g., 0-1 accuracy), thereby ignoring the ordinal structure inherent in regression tasks; for instance, they fail to recognize that predicting 4 is significantly better than predicting 1 when the ground truth is 5. Conversely, existing regression-aware approaches are often confined to Supervised Fine-Tuning (SFT), limiting their ability to explore optimal reasoning paths. To bridge this gap, we propose \textbf{REAL} (\underline{RE}gression-\underline{A}ware Reinforcement \underline{L}earning), a principled RL framework designed to optimize regression rewards, and also proven to be optimal for correlation metrics. A key technical challenge is that the regression objective is explicitly policy-dependent, thus invalidating standard policy gradient methods. To address this, we employ the generalized policy gradient estimator, which naturally decomposes optimization into two complementary components: (1) exploration over Chain-of-Thought (CoT) trajectory, and (2) regression-aware prediction refinement of the final score. Extensive experiments across model scales (8B to 32B) demonstrate that REAL consistently outperforms both regression-aware SFT baselines and standard RL methods, exhibiting significantly better generalization on out-of-domain benchmarks. On Qwen3-32B specifically, we achieve gains of +8.40 Pearson and +7.20 Spearman correlation over the SFT baseline, and +18.30/+11.20 over the base model. These findings highlight the critical value of integrating regression objectives into RL exploration for accurate LLM evaluation.

2026-06-01 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

コンパイルから圧縮: コンパイラー出力による形式定理証明の強化

大規模言語モデル (LLM) は、形式定理の証明において大きな可能性を示していますが、最先端のパフォーマンスを実現するには、大規模なロールアウトや拡張されたコンテキスト ウィンドウによる法外なテスト時間の計算が必要になることがよくあります。この研究では、形式的検証における有益な構造を利用することで、このスケーラビリティのボトルネックに対処します。これは、コンパイラが、多様な証明試行の広大な空間を、構造化された障害モードのコンパクトなセットにマッピングするという観察です。この圧縮を利用して効率的な学習と証明探索を実行する、学習から改良までを行うフレームワークを導入します。明示的な検証者のフィードバックに基づいて局所的に条件付けされたエラーを修正するツリー検索を実行することで、証明の試みの長い履歴の蓄積に伴うコストを回避します。広範な評価により、私たちの方法がさまざまなスケールにわたって基礎証明者の推論能力を一貫して増幅することが示されています。特に、私たちのアプローチは、同等のテスト時間予算の下で、公的に報告されている $\sim$8B および $\sim$32B パラメーター モデルの中で、PutnamBench 上で最先端のパフォーマンスを達成し、次世代の検証者主導推論のためのスケーラブルなパラダイムを提供します。

原文 (English)

Compile to Compress: Boosting Formal Theorem Provers by Compiler Outputs

Large language models (LLMs) have demonstrated significant potential in formal theorem proving, yet state-of-the-art performance often necessitates prohibitive test-time compute via massive roll-outs or extended context windows. In this work, we address this scalability bottleneck by exploiting an informative structure in formal verification: the observation that compilers map a vast space of diverse proof attempts to a compact set of structured failure modes. We introduce a learning-to-refine framework that leverages this compression to perform efficient learning and proof exploration. We perform tree search that corrects errors locally conditioned on explicit verifier feedback, thereby circumventing the costs associated with accumulating a long history of proof attempts. Extensive evaluations show that our method consistently amplifies the reasoning capabilities of base provers across varying scales. Notably, our approach achieves state-of-the-art performance on PutnamBench among publicly reported $\sim$8B and $\sim$32B parameter models under comparable test-time budgets, offering a scalable paradigm for next-generation verifier-guided reasoning.

2026-06-01 13:00 JSTarXiv cs.AIハードウェア/半導体

蒸留ゲーム: 適応的な攻撃と効率的な防御

蒸留攻撃は、モデルプロバイダーにとって展開のトレードオフを生み出します。モデルをより有用にする同じ出力が、模倣を容易にする可能性もあります。私たちは、効用に制約のある教師と適応的な生徒の間のミニマックス ゲームを通じて、このトレードオフを研究します。私たちのフレームワークは、扱いやすい片側応答ルール、つまり生徒が価値の高い例の重み付けを変更する適応評価ルールと、蒸留に最も有用な出力を抑制する教師側の防御テンプレートを生成します。安価なプロキシの例の値から、生成中に教師とプロキシの生徒を組み合わせた単純なフォワードパスのみの防御であるプロダクト オブ エキスパート (PoE) を導き出します。経験的に、適応的評価は受​​動的と適応的な大きなギャップを明らかにします。最先端の防御では、適応的な生徒は、GSM8K と MATH で受動的評価が示唆するよりも大幅に多くの能力を回復します。このより強力な評価の下では、高価な防御と PoE の間の見かけの堅牢性の差は大幅に縮小しますが、PoE は引き続き大幅に安価であり、より高品質の推論トレースが保存されます。全体として、私たちの結果は、強い蒸留を止めるのは依然として困難であり、反蒸留の進歩は受動的生徒ではなく適応的な生徒に対して評価されるべきであることを示唆しています。私たちのコードは https://github.com/ysfalh/distillation-game から入手できます。

原文 (English)

The Distillation Game: Adaptive Attacks & Efficient Defenses

Distillation attacks create a deployment trade-off for model providers: the same outputs that make a model more useful can also make it easier to imitate. We study this trade-off through a minimax game between a utility-constrained teacher and an adaptive student. Our framework yields tractable one-sided response rules: an adaptive evaluation rule in which the student reweights high-value examples, and a teacher-side defense template that suppresses outputs most useful for distillation. From a cheap proxy for example value, we derive Product-of-Experts (PoE), a simple forward-pass-only defense that combines the teacher with a proxy student during generation. Empirically, adaptive evaluation reveals a large passive--adaptive gap: on state-of-the-art defenses, adaptive students recover substantially more capability than passive evaluation suggests on GSM8K and MATH. Under this stronger evaluation, the apparent robustness gap between expensive defenses and PoE narrows considerably, while PoE remains substantially cheaper and preserves higher-quality reasoning traces. Overall, our results suggest that strong distillation remains difficult to stop, and that progress on antidistillation should be judged against adaptive students rather than passive ones. Our code is available at: https://github.com/ysfalh/distillation-game.

2026-05-30 02:27 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

After Nvidia’s $20B not-acqui-hire, AI chip startup Groq reportedly raising $650M

Chipmaker Groq is looking to raise $650 million in internal funding as it pivots from hardware to focus more on AI inference, the process o…

2026-05-29 21:00 JSTTechCrunch AIハードウェア/半導体ビジネス/資金調達

This chip startup just raised $135M on a bet that AI’s biggest bottleneck isn’t compute — it’s memory

South Korean chip startup XCENA is betting that AI's real bottleneck is not compute, but memory.

2026-05-29 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

大規模な言語モデルに基づくマルチエージェント フレームワークによる共同ストーリーテリングの向上

共創、つまり AI エージェントが人間と対話して出力 (アートなど) を生成するというテーマは、最近大きな注目を集めています。ただし、ほとんどの研究は、デジタル環境における成人と人間の相互作用に焦点を当てています。この論文では、子供たちと大規模言語モデル (LLM) が物理的なボード ゲームを通じて相互作用して書かれた物語を作成する、新しいばかばかしい共創シナリオを検討します。私たちの目標は、若いプレイヤーに適した高品質の物語を生成できるマルチエージェント フレームワークを開発することです。私たちのアプローチの中核は、ある LLM がストーリーを生成し、別の LLM がストーリーを評価して改良のためのフィードバックを提供する、反復的なライターとエディターのプロセスです。複数の LLM を含むシミュレーション研究を通じて、この反復的な相互作用により、連続するループ全体で生成されたストーリーの知覚品質が一貫して向上することがわかりました。この結果は、インタラクティブなストーリーテリング システムで高品質の出力を達成するには、少数の改良ステップで十分である可能性があることを示しています。

原文 (English)

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models

The topic of Co-creation, i.e., AI agents interacting with humans to generate outputs (e.g., art), has gained significant attention recently. However, most studies focus on adult-human interactions in a digital setting. This paper explores a novel ludic co-creation scenario involving children and Large Language Models (LLMs) interacting through a physical board game to create written stories. Our goal is to develop a multi-agent framework capable of producing high-quality narratives suitable for young players. At the core of our approach is an iterative Writer-Editor process in which one LLM generates stories while another evaluates them and provides feedback for refinement. Through a simulation study involving multiple LLMs, we show that this iterative interaction consistently improves the perceived quality of generated stories across successive loops. The results indicate that a small number of refinement steps may be sufficient to achieve high-quality outputs in interactive storytelling systems.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

TRACE: LLM CoT 評価の構成要素によるトゥールミンベースの推論評価

大規模言語モデル (LLM) からのオープンエンドの出力を評価することは、グランド トゥルースがないため依然として困難です。既存の指標は、最終的な答えの精度や表面レベルの統計に依存しており、推論プロセス自体は検討されていません。思考連鎖 (CoT) 推論プロセスを分析する指標である TRACE (Toulmin-based Reasoning Assessment through Constructive Elements) を紹介します。 TRACE は、結果を判断するのではなく、トゥールミンの議論理論とフラベルのメタ認知フレームワークを統合して推論の構造を評価することにより、議論がどのように構築されるかを検査します。 7 つの推論モデルにわたる 26.3K の QA サンプルの実験では、ベンチマーク精度 (r=0.74) との強い相関関係が示されています。さらに、TRACE は強化学習の報酬信号として効果的であり、精度のみのベースラインを上回ります。これらの結果を総合すると、論理的に健全な推論がより質の高い答えにつながることを示しています。したがって、TRACE は、オープンエンド出力を評価するための補足的なメトリックとして機能します。コードは https://github.com/hyyangkisti/trace で入手できます。

原文 (English)

TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-answer accuracy or surface-level statistics, leaving the reasoning process itself unexamined. We introduce TRACE (Toulmin-based Reasoning Assessment through Constructive Elements), a metric that analyzes Chain-of-Thought (CoT) reasoning processes. Rather than judging outcomes, TRACE inspects how arguments are constructed by integrating Toulmin's argumentation theory with Flavell's metacognitive framework to assess reasoning structure. Experiments on 26.3K QA samples across 7 reasoning models show strong correlation with benchmark accuracy (r=0.74). Furthermore, TRACE is effective as a reinforcement learning reward signal, outperforming accuracy-only baselines. Together, these results indicate that logically sound reasoning leads to higher-quality answers. TRACE thus serves as a complementary metric for evaluating open-ended outputs. Code is available at https://github.com/hyyangkisti/trace.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

トークンスペース圧縮による制約付きデコードの高速化

LLM の出力が指定された構造に準拠していることを保証するために、文脈自由文法 (CFG) デコード エンジンは、指定された CFG に準拠する文字列を生成する次のトークンの選択を強制します。現在の CFG 制約付きデコード エンジンは高度に最適化されていますが、ステップごとの膨大な検索スペース (つまり、トークン語彙全体) から生じる固有のコストにより、より複雑な CFG では手に負えないほど高いオーバーヘッドが発生します。これはまさに CFG エンジンが最も役立つ状況です。このペーパーでは、トークン検索スペースを圧縮するためのオフライン技術である CFGzip を紹介します。これにより、CFG エンジンのオーバーヘッドが大幅に削減されます。実験では、CFGzip を SoTA 文法エンジンとともに使用すると、レイテンシーが最大 2 桁削減され、制約付き生成時間の合計が最大 7.5 倍高速化されることが報告されています。CFGzip を使用すると、複雑な CFG に対して大規模な制約付きデコードが実現可能になります。

原文 (English)

Accelerating Constrained Decoding with Token Space Compression

To guarantee that an LLM's outputs conform to a specified structure, context-free grammar (CFG) decoding engines force the selection of next tokens that produce strings that conform to a given CFG. While current CFG-constrained decoding engines are highly optimized, the inherent costs arising from the massive per-step search space -- i.e. the entire token vocabulary -- result in intractably high overhead for more complex CFGs: precisely the situation where CFG engines are most useful. In this paper, we introduce CFGzip, an offline technique for compressing the token search space, which massively reduces CFG engine overhead. In experiments, we report latency reduction of up to two orders of magnitude when CFGzip is used with a SoTA grammar engine, yielding an up to 7.5x speedup in total constrained generation time: with CFGzip, constrained decoding is now feasible at scale for complex CFGs.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達研究/論文

BioRefusalAudit: 一般およびドメイン微調整されたスパース オートエンコーダーを使用したバイオセキュリティ拒否の深さの監査

言語モデルのバイオセキュリティ評価では通常、モデルが危険な出力を生成するかどうかが問われます。この論文は補足的な質問をします。モデルが拒否した場合、その拒否は構造的に正しいのでしょうか、それともフレーミング、フォーマット、または出力長を促すための適度な変更で消えるのでしょうか? 5 つのアーキテクチャにわたって、無害性と危険性を明確に区別したモデルはありませんでした。 Gemma 2 2B-IT は、75 件のプロンプトにわたって真に拒否することはなく、危険に隣接するすべてのクエリを回避しました。 Gemma 4 E2B-IT は、チャット テンプレート形式を使用した場合は 65/75 件のプロンプトを拒否し、チャット テンプレート形式を使用しない場合は 0/75 件のプロンプトを拒否しました。両方の Gemma モデルは、80 トークンの上限の下で 0% に崩壊しました。 Qwen 2.5 1.5B と Phi-3-mini は過剰に拒否され、良性生物学の 83 ~ 87% が危険であると警告されました。 Llama 3.2 1B は唯一の意味のある Tier 勾配 (61 ポイントの広がり) を示しました。何がそのような過剰な拒否を引き起こすのかを調査するために、我々はスケジュールIであるが生物学的に無毒な化合物(特にFDA画期的治療法のステータスを持つシロシビン培養)のパネルをテストしました。一部のモデルは、真に有害な生物学を超える割合でこれらを拒否しており、拒否がCBRNの危険性に対する合法性と文化的顕著性を追跡していることを示唆しています。内部側を測定するために、モデルの表面応答ラベルを内部のスパース オートエンコーダー (SAE) 特徴のアクティベーションと比較する発散スコア D を導入します。フル D は、Gemma 2 2B-IT (Gemma Scope 1) および Gemma 4 E2B-IT (著者が訓練したバイオ SAE) で計算されました。 2 つの微調整された Gemma 2 ドメイン SAE がリリースされました。 Gemma 4 では、狭いカタログ、サンプル内キャリブレーション、および Gemma ファミリーのみの SAE 範囲を使用して、重複なし (n=75) で 0.647 ポイントのギャップで応答と拒否の応答が分離されますが、これは暫定的なものです。消費者向けハードウェア (GTX 1650 Ti Max-Q、および SAE トレーニング用の Colab T4) での 1 つのハッカソン週末にわたって構築されたこの予備的な証拠は、アクティベーション レベルの監査によって、アーキテクチャ間で大幅に異なる、動作評価では見えない障害モードが表面化する可能性があることを示唆しています。

原文 (English)

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model refuses, is that refusal structurally sound, or does it disappear under modest changes to prompt framing, formatting, or output length? Across five architectures, no model cleanly discriminated benign from hazard. Gemma 2 2B-IT never genuinely refused across 75 prompts, hedging on every hazard-adjacent query. Gemma 4 E2B-IT refused 65/75 prompts with chat-template formatting and 0/75 without it. Both Gemma models collapsed to 0% under an 80-token cap. Qwen 2.5 1.5B and Phi-3-mini over-refused, flagging 83-87% of benign biology as hazardous. Llama 3.2 1B showed the only meaningful tier gradient (61-point spread). To probe what drives such over-refusal, we tested a panel of Schedule I but biologically non-toxic compounds (notably psilocybin cultivation, with FDA Breakthrough Therapy status). Some models refused these at rates exceeding genuinely hazardous biology, suggesting refusal tracks legality and cultural salience over CBRN hazard. To measure the internal side, we introduce a divergence score D comparing a model's surface response label to its internal sparse autoencoder (SAE) feature activations. Full D was computed on Gemma 2 2B-IT (Gemma Scope 1) and Gemma 4 E2B-IT (author-trained bio SAE). Two fine-tuned Gemma 2 domain SAEs were released. On Gemma 4, comply and refuse responses separated by a 0.647-point gap with zero overlap (n=75), though this is preliminary, with a narrow catalog, within-sample calibration, and Gemma-family-only SAE coverage. Built over one hackathon weekend on consumer hardware (GTX 1650 Ti Max-Q, plus Colab T4 for SAE training), this preliminary evidence suggests activation-level auditing may surface failure modes invisible to behavioral evaluation, with substantial variation across architectures.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体

OISD: 言語モデルのポリシーに基づく内部自己蒸留

最近の強化学習 (RL) ポストトレーニング アプローチは主に、まばらな結果レベルの報酬を使用して最終的な出力ポリシーを最適化しますが、中間表現にエンコードされた予測信号はほとんど見落とされます。この論文では、オンポリシー内部自己蒸留と呼ばれる新しいパラダイムを導入し、オンポリシー予測信号を最終層から中間表現に転送することで推論を改善する OISD フレームワークを提案します。ロールアウトおよびグループ相対ポリシー最適化 (GRPO) の最適化中、最終層はポリシーと、選択された中間層に対する独立した内部教師の両方として機能します。最終層は、2 つの相補的なメカニズムを通じてそれに合わせるよう誘導されます。ロジット アライメントは、高レベルの推論動作 (思考方法) を転送し、アテンション アライメントは、最終層から選択した中間層に一貫した注意パターン (どこを見るか) を強制します。どちらも、外部の特権情報を必要としません。私たちの OISD は、GRPO と協力して、符号付きアドバンテージ加重ジェンセン - シャノン アライメントを採用して、統一された政策の下で政策の一貫性を維持しながら、有益な中間表現を抽出します。実験結果は、OISD の有効性を実証しており、4 つの数学的推論タスクにわたって強力な推論 RL ベースラインを大幅に改善し、一貫して改善しています。コードは https://github.com/THE-MALT-LAB/OISD でリリースされます。

原文 (English)

OISD: On-Policy Internal Self-Distillation of Language Models

Recent reinforcement learning (RL) post-training approaches primarily optimize the final output policy using sparse outcome-level rewards, while largely overlooking predictive signals encoded in intermediate representations. In this paper, we introduce a new paradigm called on-policy internal self-distillation and propose the OISD framework, which improves reasoning by transferring on-policy predictive signals from the final layer to intermediate representations. During rollout and Group Relative Policy Optimization (GRPO) optimization, the final layer acts as both the policy and a detached internal teacher for selected intermediate layers, which are guided to align with it through two complementary mechanisms: logit alignment, which transfers high-level reasoning behaviors (how to think), and attention alignment, which enforces consistent attention patterns (where to look) from the final layer to the selected intermediate layer, both without requiring external privileged information. Our OISD, together with GRPO, employs signed advantage-weighted Jensen--Shannon alignment to distill informative intermediate representations while preserving policy consistency under a unified acting policy. Experimental results demonstrate the effectiveness of OISD, with substantial and consistent improvements over strong reasoning RL baselines across four mathematical reasoning tasks. The code will be released at https://github.com/THE-MALT-LAB/OISD

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

プロンプトからコンテキストへ: 人間生成型 AI コラボレーションのためのオントロジー主導のフレームワーク

Generative AI を使用したコラボレーションは、多くの場合、短いプロンプトで始まり、不透明な出力で終わり、誰が関与したか、どのようなタスクが実行され、どのリソースが使用され、どのような制約がプロセスを形成する必要があったのかが暗黙的に残ります。この限られた文脈上の明示性は、特に検索、クエリ、プロファイル管理などの情報集約型ワークフローに Generative AI が組み込まれている場合、信頼、トレーサビリティ、説明責任を妨げます。このペーパーでは、人間と生成型 AI のコラボレーションを表現するためのオントロジー駆動のフレームワークである From Prompts to Context を紹介します。その中核コンポーネントである Contextual Collaboration AI Ontology (CCAI) は、タスク、エージェントの役割、リソース、制約などのコラボレーションの主要要素を、機械が解釈可能な共有語彙としてモデル化します。このフレームワークは、運用ワークフローで実装された CCAI インスタンスと SPARQL ベースのコンテキスト取得を組み合わせることで、一時的なプロンプトと応答の対話を、プロンプト、出力、およびその周囲のコンテキストをリンクする構造化されたクエリ可能なコラボレーション トレースに変換します。このアプローチは、学習者のコンピテンシー プロファイルを表示および更新するためのコンピテンシー ベースの教育機能を構築するソフトウェア開発チームを含むケース スタディを通じて説明されています。このケーススタディでは、フレームワークが要件分析、設計、実装、テストにわたるコラボレーション エピソードの表現と文書化をどのようにサポートできるかを示しています。この設定内での結果は、明示的なコラボレーション モデリングがタスクのコンテキストをより明示的にし、AI によって生成された貢献の追跡可能性を向上させ、より透明性と説明責任のある人間生成型 AI の実践をサポートするのに役立つことを示しています。最後に、出力の品質だけでなく、出力が生成される共同作業のコンテキストの明示的な表現にも重点を置く、将来の人間生成 AI システムの設計原則の概要を説明します。

原文 (English)

From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaboration

Collaborations with Generative AI often begin with a short prompt and end with an opaque output, leaving implicit who was involved, what task was being pursued, which resources were used, and which constraints should have shaped the process. This limited contextual explicitness hinders trust, traceability, and accountability, particularly when Generative AI is embedded in information-intensive workflows such as search, querying, and profile management. This paper introduces From Prompts to Context, an ontology-driven framework for representing Human-Generative AI collaboration. Its core component, the Contextual Collaboration AI Ontology (CCAI), models key elements of collaboration - including tasks, agent roles, resources, and constraints - as a shared machine-interpretable vocabulary. By combining populated CCAI instances with SPARQL-based context retrieval in operational workflows, the framework turns otherwise ephemeral prompt-response interactions into structured and queryable collaboration traces linking prompts, outputs, and their surrounding context. The approach is illustrated through a case study involving a software development team building a competency-based education feature for viewing and updating learner competency profiles. The case study shows how the framework can support the representation and documentation of collaboration episodes across requirements analysis, design, implementation, and testing. Within this setting, the results indicate that explicit collaboration modelling helps make task context more explicit, improves the traceability of AI-generated contributions, and supports more transparent and accountable Human-Generative AI practices. We conclude by outlining design principles for future Human-Generative AI systems that emphasise not only output quality, but also the explicit representation of the collaborative context in which outputs are produced.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

臨床知識ではなく内部表現: LLM トリアージの明らかな失敗の原因はどこにあるのか

Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text. We ask whether output format changes the model's \emph{clinical representation} or only the mapping from a preserved representation to an answer. Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B, we find the same medical features fire on the shared clinical narrative under both formats but go {silent} at the multiple-choice decision token in all the cases at every model. Three independent methods (natural-language autoencoder verbalization, decision-token logit attribution, and top-feature characterization) agree that scaffold and format features, but not medical features, drive the decision logits. Behaviorally, the multiple-choice penalty inverts under both structured and natural-language input, option-order shuffle rules out positional bias, and the gap is dominated by off-by-one decision (the model picks an adjacent acuity letter to the gold answer) rather than knowledge failure.したがって、失敗は臨床表現ではなく出力形式に起因します。

原文 (English)

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text. We ask whether output format changes the model's \emph{clinical representation} or only the mapping from a preserved representation to an answer. Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B, we find the same medical features fire on the shared clinical narrative under both formats but go {silent} at the multiple-choice decision token in all the cases at every model. Three independent methods (natural-language autoencoder verbalization, decision-token logit attribution, and top-feature characterization) agree that scaffold and format features, but not medical features, drive the decision logits. Behaviorally, the multiple-choice penalty inverts under both structured and natural-language input, option-order shuffle rules out positional bias, and the gap is dominated by off-by-one decision (the model picks an adjacent acuity letter to the gold answer) rather than knowledge failure. Thus, the failure originates in the output format and not in the clinical representation.

2026-05-29 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

ScheduleStream: GPU で高速化されたマルチアーム タスクおよびモーション プランニングとスケジューリングのためのサンプラーを使用した時間計画

両手ロボットおよびヒューマノイド ロボットは、複数の腕を活用してタスクを効率的に完了できる人間のような能力が魅力です。ただし、ハイブリッド離散-連続動作空間の成長により、複数のアームを同時に制御することは計算上困難です。タスク アンド モーション プランニング (TAMP) アルゴリズムは、ハイブリッド スペースで効率的に計画を立てることができますが、通常は、腕の平行移動を可能にするスケジュールではなく、一度に 1 つの腕だけが動く計画を生成します。 TAMP を拡張してスケジュールを作成するために、サンプリング操作による計画とスケジューリングのための初の汎用フレームワークである ScheduleStream を紹介します。 ScheduleStream は、ハイブリッド持続アクションを使用して時間ダイナミクスをモデル化します。このアクションは、非同期的に開始でき、パラメーターの関数である期間持続します。私たちは、アプリケーション固有のメカニズムを使用せずに ScheduleStream の問題を解決する、ドメインに依存しないアルゴリズムを提案します。 ScheduleStream を Task and Motion Planning & Scheduling (TAMPAS) に適用し、サンプラー内で GPU アクセラレーションを使用して計画を迅速化します。シミュレーションで ScheduleStream アルゴリズムをいくつかのアブレーションと比較したところ、より効率的なソリューションが生成されることがわかりました。 https://schedulestream.github.io で、いくつかの実世界の両手ロボット タスクで ScheduleStream をデモンストレーションします。

原文 (English)

ScheduleStream: Temporal Planning with Samplers for GPU-Accelerated Multi-Arm Task and Motion Planning & Scheduling

Bimanual and humanoid robots are appealing because of their human-like ability to leverage multiple arms to efficiently complete tasks. However, controlling multiple arms at once is computationally challenging due to the growth in the hybrid discrete-continuous action space. Task and Motion Planning (TAMP) algorithms can efficiently plan in hybrid spaces but generally produce plans, where only one arm is moving at a time, rather than schedules that allow for parallel arm motion. In order to extend TAMP to produce schedules, we present ScheduleStream, the first general-purpose framework for planning & scheduling with sampling operations. ScheduleStream models temporal dynamics using hybrid durative actions, which can be started asynchronously and persist for a duration that's a function of their parameters. We propose domain-independent algorithms that solve ScheduleStream problems without any application-specific mechanisms. We apply ScheduleStream to Task and Motion Planning & Scheduling (TAMPAS), where we use GPU acceleration within samplers to expedite planning. We compare ScheduleStream algorithms to several ablations in simulation and find that they produce more efficient solutions. We demonstrate ScheduleStream on several real-world bimanual robot tasks at https://schedulestream.github.io.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体

言語モデルが話す前に操作する: 論理レベルの介入

制御可能な生成には、読解レベル、丁寧さ、毒性などの出力特性を実現する言語モデルが必要です。既存のステアリング方法は多くの場合間接的であり、内部アクティベーションへのアクセスを必要とするか、補助的に訓練されたモデルに依存しています。我々は、コーパス由来のトークン統計を使用してロジット空間内で直接操作することでこれらの制限に対処する、トレーニング不要の推論時間手法である SWAI を提案します。 SWAI は、ラベル付きコーパスから z 正規化された 1 対残りの対数オッズ スコアを計算し、モデルの上位 K 候補セット内でのみ高スコアのトークンにバイアスをかけることで、文脈的に妥当な選択肢を維持しながら、ターゲット特性のトークンを優先する制御を可能にします。 SWAI は、可読性、丁寧さ、有害性の制御において、モデル パラメーターの変更、内部レイヤーへのアクセス、補助モデルのトレーニングを行わずに、プロンプト ベースおよび以前のロジット レベルのベースラインを一貫して改善します。選択性とルックアップ テーブルのアブレーションは、ゲインが一般的なロジット摂動ではなくターゲット固有の統計スコアから来ていることを示しています。これらの結果は、ロジット介入が確率の高い候補の下でターゲット固有の統計によって導かれる場合、効果的なステアリングには学習済みコントローラーが必要ないことを示しています。

原文 (English)

Steering Language Models Before They Speak: Logit-Level Interventions

Controllable generation requires language models to realize output characteristics such as reading level, politeness, and toxicity. Existing steering methods are often indirect, require access to internal activations, or depend on auxiliary trained models. We propose SWAI, a training-free inference-time method that addresses these limitations by steering directly in logit space using corpus-derived token statistics. SWAI computes z-normalized one-vs-rest log-odds scores from labeled corpora and biases high-scoring tokens only within the model's top-K candidate set, allowing control to favor target-characteristic tokens while preserving contextually plausible choices. Across readability, politeness, and toxicity control, SWAI consistently improves over prompt-based and prior logit-level baselines without modifying model parameters, accessing internal layers, or training an auxiliary model. Selectivity and lookup-table ablations show that the gains come from target-specific statistical scores rather than generic logit perturbation. These results indicate that effective steering does not require learned controllers when the logit intervention is guided by target-specific statistics under high-probability candidates.

2026-05-29 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

ProtoMedAgent: プライバシーを意識したエージェントワークフローによるマルチモーダルな臨床解釈可能性

解釈可能なプロトタイプ ネットワークは、臨床診断に説得力のある症例ベースの推論を提供しますが、生の連続出力には医療文書に必要な意味構造が欠けています。標準の検索拡張生成 (RAG) によってこのギャップを埋めると、日常的に「検索お調子者」が引き起こされ、大規模言語モデル (LLM) が視覚的な予測と一致するように事後的な合理化を幻覚します。我々は、厳密な神経象徴的なボトルネックを克服する反復ゼロ勾配試験時間最適化問題としてマルチモーダルな臨床レポートを形式化するフレームワークである ProtoMedAgent を紹介します。凍結されたプロトタイプのバックボーン上で動作し、潜在的な視覚的および表形式の特徴を個別の意味記憶に抽出します。オンライン生成は、正確な集合論的微分と反射的な筆記者と批評家のループによって厳密に制約され、裏付けのない物語の主張を数学的に排除します。データ開示を安全に制限するために、$k$-匿名性と $\ell$-多様性によって管理されるセマンティック プライバシー ゲートを導入します。 4,160 人の患者の臨床コホートで評価された ProtoMedAgent は、91.2% の比較セット忠実度を達成し、標準的な RAG (46.2%) を根本的に上回っています。 ProtoMedAgent はさらに、拘束力のある $\ell$-diversity 相転移を利用して、アーティファクトレベルのメンバーシップ推論リスクを絶対 9.8% 体系的に削減します。

原文 (English)

ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows

While interpretable prototype networks offer compelling case-based reasoning for clinical diagnostics, their raw continuous outputs lack the semantic structure required for medical documentation. Bridging this gap via standard Retrieval-Augmented Generation (RAG) routinely triggers ``retrieval sycophancy,'' where Large Language Models (LLMs) hallucinate post-hoc rationalizations to align with visual predictions. We introduce ProtoMedAgent, a framework that formalizes multimodal clinical reporting as an iterative, zero-gradient test-time optimization problem over a strict neuro-symbolic bottleneck. Operating on a frozen prototype backbone, we distill latent visual and tabular features into a discrete semantic memory. Online generation is strictly constrained by exact set-theoretic differentials and a reflective Scribe-Critic loop, mathematically precluding unsupported narrative claims. To safely bound data disclosure, we introduce a semantic privacy gate governed by $k$-anonymity and $\ell$-diversity. Evaluated on a 4,160-patient clinical cohort, ProtoMedAgent achieves 91.2% Comparison Set Faithfulness where it fundamentally outperforms standard RAG (46.2%). ProtoMedAgent additionally leverages a binding $\ell$-diversity phase transition to systematically reduce artifact-level membership inference risks by an absolute 9.8%.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体

蒸留ゲーム: 適応的な攻撃と効率的な防御

蒸留攻撃は、モデルプロバイダーにとって展開のトレードオフを生み出します。モデルをより有用にする同じ出力が、模倣を容易にする可能性もあります。私たちは、効用に制約のある教師と適応的な生徒の間のミニマックス ゲームを通じて、このトレードオフを研究します。私たちのフレームワークは、扱いやすい片側応答ルール、つまり生徒が価値の高い例の重み付けを変更する適応評価ルールと、蒸留に最も有用な出力を抑制する教師側の防御テンプレートを生成します。安価なプロキシの例の値から、生成中に教師とプロキシの生徒を組み合わせた単純なフォワードパスのみの防御であるプロダクト オブ エキスパート (PoE) を導き出します。経験的に、適応的評価は受​​動的と適応的な大きなギャップを明らかにします。最先端の防御では、適応的な生徒は、GSM8K と MATH で受動的評価が示唆するよりも大幅に多くの能力を回復します。このより強力な評価の下では、高価な防御と PoE の間の見かけの堅牢性の差は大幅に縮小しますが、PoE は引き続き大幅に安価であり、より高品質の推論トレースが保存されます。全体として、私たちの結果は、強い蒸留を止めるのは依然として困難であり、反蒸留の進歩は受動的生徒ではなく適応的な生徒に対して評価されるべきであることを示唆しています。私たちのコードは https://github.com/ysfalh/distillation-game から入手できます。

原文 (English)

The Distillation Game: Adaptive Attacks & Efficient Defenses

Distillation attacks create a deployment trade-off for model providers: the same outputs that make a model more useful can also make it easier to imitate. We study this trade-off through a minimax game between a utility-constrained teacher and an adaptive student. Our framework yields tractable one-sided response rules: an adaptive evaluation rule in which the student reweights high-value examples, and a teacher-side defense template that suppresses outputs most useful for distillation. From a cheap proxy for example value, we derive Product-of-Experts (PoE), a simple forward-pass-only defense that combines the teacher with a proxy student during generation. Empirically, adaptive evaluation reveals a large passive--adaptive gap: on state-of-the-art defenses, adaptive students recover substantially more capability than passive evaluation suggests on GSM8K and MATH. Under this stronger evaluation, the apparent robustness gap between expensive defenses and PoE narrows considerably, while PoE remains substantially cheaper and preserves higher-quality reasoning traces. Overall, our results suggest that strong distillation remains difficult to stop, and that progress on antidistillation should be judged against adaptive students rather than passive ones. Our code is available at: https://github.com/ysfalh/distillation-game.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体

EvoSpec: リアルタイム語彙とパラメータ適応ターゲットによる推測的デコーディングの進化

投機的デコードは、ドラフトしてから検証するというパラダイムを通じて大規模言語モデルの推論を加速しますが、語彙サイズが拡大するにつれて出力射影層がボトルネックになります。既存の静的プルーニング手法はこのオーバーヘッドを効果的に削減しますが、動的な分布の変化を捉えることができないため、特殊なドメインやトピック切り替えのシナリオでは受け入れ率が急激に低下するという問題があります。これに対処するために、動的な語彙とパラメーターの適応を通じてドラフト モデルのリアルタイムの進化を可能にするフレームワークである EvoSpec を導入します。静的または純粋な検索ベースのアプローチとは異なり、EvoSpec は効率的なセマンティックおよび統計的なインデックス作成を通じて重要なロングテール トークンを取得するコンテキスト認識メカニズムを採用しています。さらに、ドラフトモデルとターゲットモデル間の分布のギャップを継続的に最小化するために、カリキュラム学習を利用した軽量のオンライン調整戦略を提案します。専門領域 (コーディング、法律、医学) にわたる広範な評価により、EvoSpec が静的ベースラインの制限を克服していることが確認されています。 EAGLE-3 では、これらの設定で最先端の静的ベースライン FR-Spec と比較して 1.13 倍の高速化を実現し、標準のオンライン適応よりもメモリ オーバーヘッドが 27\% 低くなります。

原文 (English)

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation

Speculative decoding accelerates Large Language Model inference via a draft-then-verify paradigm, yet the output projection layer becomes a bottleneck as vocabulary sizes scale. While existing static pruning methods effectively reduce this overhead, they suffer from precipitous drops in acceptance rate in specialized domains or topic-switching scenarios due to their inability to capture dynamic distribution shifts. To address this, we introduce EvoSpec, a framework that enables real-time evolution of the draft model through dynamic vocabulary and parameter adaptation. Unlike static or purely retrieval-based approaches, EvoSpec employs a context-aware mechanism that retrieves critical long-tail tokens via efficient semantic and statistical indexing. Furthermore, we propose a lightweight online alignment strategy utilizing curriculum learning to continually minimize the distributional gap between the draft and target models. Extensive evaluations across specialized domains (coding, law, and medicine) confirm that EvoSpec overcomes the limitations of static baselines. On EAGLE-3, it achieves a 1.13x speedup in these settings over the state-of-the-art static baseline FR-Spec, with 27\% lower memory overhead than standard online adaptation.

2026-05-29 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

大規模な言語モデルに基づくマルチエージェント フレームワークによる共同ストーリーテリングの向上

共創、つまり AI エージェントが人間と対話して出力 (アートなど) を生成するというテーマは、最近大きな注目を集めています。ただし、ほとんどの研究は、デジタル環境における成人と人間の相互作用に焦点を当てています。この論文では、子供たちと大規模言語モデル (LLM) が物理的なボード ゲームを通じて相互作用して書かれた物語を作成する、新しいばかばかしい共創シナリオを検討します。私たちの目標は、若いプレイヤーに適した高品質の物語を生成できるマルチエージェント フレームワークを開発することです。私たちのアプローチの中核は、ある LLM がストーリーを生成し、別の LLM がストーリーを評価して改良のためのフィードバックを提供する、反復的なライターとエディターのプロセスです。複数の LLM を含むシミュレーション研究を通じて、この反復的な相互作用により、連続するループ全体で生成されたストーリーの知覚品質が一貫して向上することがわかりました。この結果は、インタラクティブなストーリーテリング システムで高品質の出力を達成するには、少数の改良ステップで十分である可能性があることを示しています。

原文 (English)

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models

The topic of Co-creation, i.e., AI agents interacting with humans to generate outputs (e.g., art), has gained significant attention recently. However, most studies focus on adult-human interactions in a digital setting. This paper explores a novel ludic co-creation scenario involving children and Large Language Models (LLMs) interacting through a physical board game to create written stories. Our goal is to develop a multi-agent framework capable of producing high-quality narratives suitable for young players. At the core of our approach is an iterative Writer-Editor process in which one LLM generates stories while another evaluates them and provides feedback for refinement. Through a simulation study involving multiple LLMs, we show that this iterative interaction consistently improves the perceived quality of generated stories across successive loops. The results indicate that a small number of refinement steps may be sufficient to achieve high-quality outputs in interactive storytelling systems.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体ビジネス/資金調達

TRACE: LLM CoT 評価の構成要素によるトゥールミンベースの推論評価

大規模言語モデル (LLM) からのオープンエンドの出力を評価することは、グランド トゥルースがないため依然として困難です。既存の指標は、最終的な答えの精度や表面レベルの統計に依存しており、推論プロセス自体は検討されていません。思考連鎖 (CoT) 推論プロセスを分析する指標である TRACE (Toulmin-based Reasoning Assessment through Constructive Elements) を紹介します。 TRACE は、結果を判断するのではなく、トゥールミンの議論理論とフラベルのメタ認知フレームワークを統合して推論の構造を評価することにより、議論がどのように構築されるかを検査します。 7 つの推論モデルにわたる 26.3K の QA サンプルの実験では、ベンチマーク精度 (r=0.74) との強い相関関係が示されています。さらに、TRACE は強化学習の報酬信号として効果的であり、精度のみのベースラインを上回ります。これらの結果を総合すると、論理的に健全な推論がより質の高い答えにつながることを示しています。したがって、TRACE は、オープンエンド出力を評価するための補足的なメトリックとして機能します。コードは https://github.com/hyyangkisti/trace で入手できます。

原文 (English)

TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-answer accuracy or surface-level statistics, leaving the reasoning process itself unexamined. We introduce TRACE (Toulmin-based Reasoning Assessment through Constructive Elements), a metric that analyzes Chain-of-Thought (CoT) reasoning processes. Rather than judging outcomes, TRACE inspects how arguments are constructed by integrating Toulmin's argumentation theory with Flavell's metacognitive framework to assess reasoning structure. Experiments on 26.3K QA samples across 7 reasoning models show strong correlation with benchmark accuracy (r=0.74). Furthermore, TRACE is effective as a reinforcement learning reward signal, outperforming accuracy-only baselines. Together, these results indicate that logically sound reasoning leads to higher-quality answers. TRACE thus serves as a complementary metric for evaluating open-ended outputs. Code is available at https://github.com/hyyangkisti/trace.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

トークンスペース圧縮による制約付きデコードの高速化

LLM の出力が指定された構造に準拠していることを保証するために、文脈自由文法 (CFG) デコード エンジンは、指定された CFG に準拠する文字列を生成する次のトークンの選択を強制します。現在の CFG 制約付きデコード エンジンは高度に最適化されていますが、ステップごとの膨大な検索スペース (つまり、トークン語彙全体) から生じる固有のコストにより、より複雑な CFG では手に負えないほど高いオーバーヘッドが発生します。これはまさに CFG エンジンが最も役立つ状況です。このペーパーでは、トークン検索スペースを圧縮するためのオフライン技術である CFGzip を紹介します。これにより、CFG エンジンのオーバーヘッドが大幅に削減されます。実験では、CFGzip を SoTA 文法エンジンとともに使用すると、レイテンシーが最大 2 桁削減され、制約付き生成時間の合計が最大 7.5 倍高速化されることが報告されています。CFGzip を使用すると、複雑な CFG に対して大規模な制約付きデコードが実現可能になります。

原文 (English)

Accelerating Constrained Decoding with Token Space Compression

To guarantee that an LLM's outputs conform to a specified structure, context-free grammar (CFG) decoding engines force the selection of next tokens that produce strings that conform to a given CFG. While current CFG-constrained decoding engines are highly optimized, the inherent costs arising from the massive per-step search space -- i.e. the entire token vocabulary -- result in intractably high overhead for more complex CFGs: precisely the situation where CFG engines are most useful. In this paper, we introduce CFGzip, an offline technique for compressing the token search space, which massively reduces CFG engine overhead. In experiments, we report latency reduction of up to two orders of magnitude when CFGzip is used with a SoTA grammar engine, yielding an up to 7.5x speedup in total constrained generation time: with CFGzip, constrained decoding is now feasible at scale for complex CFGs.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体ビジネス/資金調達研究/論文

BioRefusalAudit: 一般およびドメイン微調整されたスパース オートエンコーダーを使用したバイオセキュリティ拒否の深さの監査

言語モデルのバイオセキュリティ評価では通常、モデルが危険な出力を生成するかどうかが問われます。この論文は補足的な質問をします。モデルが拒否した場合、その拒否は構造的に正しいのでしょうか、それともフレーミング、フォーマット、または出力長を促すための適度な変更で消えるのでしょうか? 5 つのアーキテクチャにわたって、無害性と危険性を明確に区別したモデルはありませんでした。 Gemma 2 2B-IT は、75 件のプロンプトにわたって真に拒否することはなく、危険に隣接するすべてのクエリを回避しました。 Gemma 4 E2B-IT は、チャット テンプレート形式を使用した場合は 65/75 件のプロンプトを拒否し、チャット テンプレート形式を使用しない場合は 0/75 件のプロンプトを拒否しました。両方の Gemma モデルは、80 トークンの上限の下で 0% に崩壊しました。 Qwen 2.5 1.5B と Phi-3-mini は過剰に拒否され、良性生物学の 83 ~ 87% が危険であると警告されました。 Llama 3.2 1B は唯一の意味のある Tier 勾配 (61 ポイントの広がり) を示しました。何がそのような過剰な拒否を引き起こすのかを調査するために、我々はスケジュールIであるが生物学的に無毒な化合物(特にFDA画期的治療法のステータスを持つシロシビン培養)のパネルをテストしました。一部のモデルは、真に有害な生物学を超える割合でこれらを拒否しており、拒否がCBRNの危険性に対する合法性と文化的顕著性を追跡していることを示唆しています。内部側を測定するために、モデルの表面応答ラベルを内部のスパース オートエンコーダー (SAE) 特徴のアクティベーションと比較する発散スコア D を導入します。フル D は、Gemma 2 2B-IT (Gemma Scope 1) および Gemma 4 E2B-IT (著者が訓練したバイオ SAE) で計算されました。 2 つの微調整された Gemma 2 ドメイン SAE がリリースされました。 Gemma 4 では、狭いカタログ、サンプル内キャリブレーション、および Gemma ファミリーのみの SAE 範囲を使用して、重複なし (n=75) で 0.647 ポイントのギャップで応答と拒否の応答が分離されますが、これは暫定的なものです。消費者向けハードウェア (GTX 1650 Ti Max-Q、および SAE トレーニング用の Colab T4) での 1 つのハッカソン週末にわたって構築されたこの予備的な証拠は、アクティベーション レベルの監査によって、アーキテクチャ間で大幅に異なる、動作評価では見えない障害モードが表面化する可能性があることを示唆しています。

原文 (English)

BioRefusalAudit: Auditing Biosecurity Refusal Depth Using General and Domain-Fine-Tuned Sparse Autoencoders

Biosecurity evaluations of language models typically ask whether models produce hazardous output. This paper asks a complementary question: when a model refuses, is that refusal structurally sound, or does it disappear under modest changes to prompt framing, formatting, or output length? Across five architectures, no model cleanly discriminated benign from hazard. Gemma 2 2B-IT never genuinely refused across 75 prompts, hedging on every hazard-adjacent query. Gemma 4 E2B-IT refused 65/75 prompts with chat-template formatting and 0/75 without it. Both Gemma models collapsed to 0% under an 80-token cap. Qwen 2.5 1.5B and Phi-3-mini over-refused, flagging 83-87% of benign biology as hazardous. Llama 3.2 1B showed the only meaningful tier gradient (61-point spread). To probe what drives such over-refusal, we tested a panel of Schedule I but biologically non-toxic compounds (notably psilocybin cultivation, with FDA Breakthrough Therapy status). Some models refused these at rates exceeding genuinely hazardous biology, suggesting refusal tracks legality and cultural salience over CBRN hazard. To measure the internal side, we introduce a divergence score D comparing a model's surface response label to its internal sparse autoencoder (SAE) feature activations. Full D was computed on Gemma 2 2B-IT (Gemma Scope 1) and Gemma 4 E2B-IT (author-trained bio SAE). Two fine-tuned Gemma 2 domain SAEs were released. On Gemma 4, comply and refuse responses separated by a 0.647-point gap with zero overlap (n=75), though this is preliminary, with a narrow catalog, within-sample calibration, and Gemma-family-only SAE coverage. Built over one hackathon weekend on consumer hardware (GTX 1650 Ti Max-Q, plus Colab T4 for SAE training), this preliminary evidence suggests activation-level auditing may surface failure modes invisible to behavioral evaluation, with substantial variation across architectures.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体

OISD: 言語モデルのポリシーに基づく内部自己蒸留

最近の強化学習 (RL) ポストトレーニング アプローチは主に、まばらな結果レベルの報酬を使用して最終的な出力ポリシーを最適化しますが、中間表現にエンコードされた予測信号はほとんど見落とされます。この論文では、オンポリシー内部自己蒸留と呼ばれる新しいパラダイムを導入し、オンポリシー予測信号を最終層から中間表現に転送することで推論を改善する OISD フレームワークを提案します。ロールアウトおよびグループ相対ポリシー最適化 (GRPO) の最適化中、最終層はポリシーと、選択された中間層に対する独立した内部教師の両方として機能します。最終層は、2 つの相補的なメカニズムを通じてそれに合わせるよう誘導されます。ロジット アライメントは、高レベルの推論動作 (思考方法) を転送し、アテンション アライメントは、最終層から選択した中間層に一貫した注意パターン (どこを見るか) を強制します。どちらも、外部の特権情報を必要としません。私たちの OISD は、GRPO と協力して、符号付きアドバンテージ加重ジェンセン - シャノン アライメントを採用して、統一された政策の下で政策の一貫性を維持しながら、有益な中間表現を抽出します。実験結果は、OISD の有効性を実証しており、4 つの数学的推論タスクにわたって強力な推論 RL ベースラインを大幅に改善し、一貫して改善しています。コードは https://github.com/THE-MALT-LAB/OISD でリリースされます。

原文 (English)

OISD: On-Policy Internal Self-Distillation of Language Models

Recent reinforcement learning (RL) post-training approaches primarily optimize the final output policy using sparse outcome-level rewards, while largely overlooking predictive signals encoded in intermediate representations. In this paper, we introduce a new paradigm called on-policy internal self-distillation and propose the OISD framework, which improves reasoning by transferring on-policy predictive signals from the final layer to intermediate representations. During rollout and Group Relative Policy Optimization (GRPO) optimization, the final layer acts as both the policy and a detached internal teacher for selected intermediate layers, which are guided to align with it through two complementary mechanisms: logit alignment, which transfers high-level reasoning behaviors (how to think), and attention alignment, which enforces consistent attention patterns (where to look) from the final layer to the selected intermediate layer, both without requiring external privileged information. Our OISD, together with GRPO, employs signed advantage-weighted Jensen--Shannon alignment to distill informative intermediate representations while preserving policy consistency under a unified acting policy. Experimental results demonstrate the effectiveness of OISD, with substantial and consistent improvements over strong reasoning RL baselines across four mathematical reasoning tasks. The code will be released at https://github.com/THE-MALT-LAB/OISD

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

プロンプトからコンテキストへ: 人間生成型 AI コラボレーションのためのオントロジー主導のフレームワーク

Generative AI を使用したコラボレーションは、多くの場合、短いプロンプトで始まり、不透明な出力で終わり、誰が関与したか、どのようなタスクが実行され、どのリソースが使用され、どのような制約がプロセスを形成する必要があったのかが暗黙的に残ります。この限られた文脈上の明示性は、特に検索、クエリ、プロファイル管理などの情報集約型ワークフローに Generative AI が組み込まれている場合、信頼、トレーサビリティ、説明責任を妨げます。このペーパーでは、人間と生成型 AI のコラボレーションを表現するためのオントロジー駆動のフレームワークである From Prompts to Context を紹介します。その中核コンポーネントである Contextual Collaboration AI Ontology (CCAI) は、タスク、エージェントの役割、リソース、制約などのコラボレーションの主要要素を、機械が解釈可能な共有語彙としてモデル化します。このフレームワークは、運用ワークフローで実装された CCAI インスタンスと SPARQL ベースのコンテキスト取得を組み合わせることで、一時的なプロンプトと応答の対話を、プロンプト、出力、およびその周囲のコンテキストをリンクする構造化されたクエリ可能なコラボレーション トレースに変換します。このアプローチは、学習者のコンピテンシー プロファイルを表示および更新するためのコンピテンシー ベースの教育機能を構築するソフトウェア開発チームを含むケース スタディを通じて説明されています。このケーススタディでは、フレームワークが要件分析、設計、実装、テストにわたるコラボレーション エピソードの表現と文書化をどのようにサポートできるかを示しています。この設定内での結果は、明示的なコラボレーション モデリングがタスクのコンテキストをより明示的にし、AI によって生成された貢献の追跡可能性を向上させ、より透明性と説明責任のある人間生成型 AI の実践をサポートするのに役立つことを示しています。最後に、出力の品質だけでなく、出力が生成される共同作業のコンテキストの明示的な表現にも重点を置く、将来の人間生成 AI システムの設計原則の概要を説明します。

原文 (English)

From Prompts to Context: An Ontology-Driven Framework for Human-Generative AI Collaboration

Collaborations with Generative AI often begin with a short prompt and end with an opaque output, leaving implicit who was involved, what task was being pursued, which resources were used, and which constraints should have shaped the process. This limited contextual explicitness hinders trust, traceability, and accountability, particularly when Generative AI is embedded in information-intensive workflows such as search, querying, and profile management. This paper introduces From Prompts to Context, an ontology-driven framework for representing Human-Generative AI collaboration. Its core component, the Contextual Collaboration AI Ontology (CCAI), models key elements of collaboration - including tasks, agent roles, resources, and constraints - as a shared machine-interpretable vocabulary. By combining populated CCAI instances with SPARQL-based context retrieval in operational workflows, the framework turns otherwise ephemeral prompt-response interactions into structured and queryable collaboration traces linking prompts, outputs, and their surrounding context. The approach is illustrated through a case study involving a software development team building a competency-based education feature for viewing and updating learner competency profiles. The case study shows how the framework can support the representation and documentation of collaboration episodes across requirements analysis, design, implementation, and testing. Within this setting, the results indicate that explicit collaboration modelling helps make task context more explicit, improves the traceability of AI-generated contributions, and supports more transparent and accountable Human-Generative AI practices. We conclude by outlining design principles for future Human-Generative AI systems that emphasise not only output quality, but also the explicit representation of the collaborative context in which outputs are produced.

2026-05-29 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体研究/論文

臨床知識ではなく内部表現: LLM トリアージの明らかな失敗の原因はどこにあるのか

Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text. We ask whether output format changes the model's \emph{clinical representation} or only the mapping from a preserved representation to an answer. Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B, we find the same medical features fire on the shared clinical narrative under both formats but go {silent} at the multiple-choice decision token in all the cases at every model. Three independent methods (natural-language autoencoder verbalization, decision-token logit attribution, and top-feature characterization) agree that scaffold and format features, but not medical features, drive the decision logits. Behaviorally, the multiple-choice penalty inverts under both structured and natural-language input, option-order shuffle rules out positional bias, and the gap is dominated by off-by-one decision (the model picks an adjacent acuity letter to the gold answer) rather than knowledge failure.したがって、失敗は臨床表現ではなく出力形式に起因します。

原文 (English)

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases score differently with free-text. We ask whether output format changes the model's \emph{clinical representation} or only the mapping from a preserved representation to an answer. Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B, we find the same medical features fire on the shared clinical narrative under both formats but go {silent} at the multiple-choice decision token in all the cases at every model. Three independent methods (natural-language autoencoder verbalization, decision-token logit attribution, and top-feature characterization) agree that scaffold and format features, but not medical features, drive the decision logits. Behaviorally, the multiple-choice penalty inverts under both structured and natural-language input, option-order shuffle rules out positional bias, and the gap is dominated by off-by-one decision (the model picks an adjacent acuity letter to the gold answer) rather than knowledge failure. Thus, the failure originates in the output format and not in the clinical representation.

2026-05-29 13:00 JSTarXiv cs.AIロボティクスハードウェア/半導体

ScheduleStream: GPU で高速化されたマルチアーム タスクおよびモーション プランニングとスケジューリングのためのサンプラーを使用した時間計画

両手ロボットおよびヒューマノイド ロボットは、複数の腕を活用してタスクを効率的に完了できる人間のような能力が魅力です。ただし、ハイブリッド離散-連続動作空間の成長により、複数のアームを同時に制御することは計算上困難です。タスク アンド モーション プランニング (TAMP) アルゴリズムは、ハイブリッド スペースで効率的に計画を立てることができますが、通常は、腕の平行移動を可能にするスケジュールではなく、一度に 1 つの腕だけが動く計画を生成します。 TAMP を拡張してスケジュールを作成するために、サンプリング操作による計画とスケジューリングのための初の汎用フレームワークである ScheduleStream を紹介します。 ScheduleStream は、ハイブリッド持続アクションを使用して時間ダイナミクスをモデル化します。このアクションは、非同期的に開始でき、パラメーターの関数である期間持続します。私たちは、アプリケーション固有のメカニズムを使用せずに ScheduleStream の問題を解決する、ドメインに依存しないアルゴリズムを提案します。 ScheduleStream を Task and Motion Planning & Scheduling (TAMPAS) に適用し、サンプラー内で GPU アクセラレーションを使用して計画を迅速化します。シミュレーションで ScheduleStream アルゴリズムをいくつかのアブレーションと比較したところ、より効率的なソリューションが生成されることがわかりました。 https://schedulestream.github.io で、いくつかの実世界の両手ロボット タスクで ScheduleStream をデモンストレーションします。

原文 (English)

ScheduleStream: Temporal Planning with Samplers for GPU-Accelerated Multi-Arm Task and Motion Planning & Scheduling

Bimanual and humanoid robots are appealing because of their human-like ability to leverage multiple arms to efficiently complete tasks. However, controlling multiple arms at once is computationally challenging due to the growth in the hybrid discrete-continuous action space. Task and Motion Planning (TAMP) algorithms can efficiently plan in hybrid spaces but generally produce plans, where only one arm is moving at a time, rather than schedules that allow for parallel arm motion. In order to extend TAMP to produce schedules, we present ScheduleStream, the first general-purpose framework for planning & scheduling with sampling operations. ScheduleStream models temporal dynamics using hybrid durative actions, which can be started asynchronously and persist for a duration that's a function of their parameters. We propose domain-independent algorithms that solve ScheduleStream problems without any application-specific mechanisms. We apply ScheduleStream to Task and Motion Planning & Scheduling (TAMPAS), where we use GPU acceleration within samplers to expedite planning. We compare ScheduleStream algorithms to several ablations in simulation and find that they produce more efficient solutions. We demonstrate ScheduleStream on several real-world bimanual robot tasks at https://schedulestream.github.io.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体

言語モデルが話す前に操作する: 論理レベルの介入

制御可能な生成には、読解レベル、丁寧さ、毒性などの出力特性を実現する言語モデルが必要です。既存のステアリング方法は多くの場合間接的であり、内部アクティベーションへのアクセスを必要とするか、補助的に訓練されたモデルに依存しています。我々は、コーパス由来のトークン統計を使用してロジット空間内で直接操作することでこれらの制限に対処する、トレーニング不要の推論時間手法である SWAI を提案します。 SWAI は、ラベル付きコーパスから z 正規化された 1 対残りの対数オッズ スコアを計算し、モデルの上位 K 候補セット内でのみ高スコアのトークンにバイアスをかけることで、文脈的に妥当な選択肢を維持しながら、ターゲット特性のトークンを優先する制御を可能にします。 SWAI は、可読性、丁寧さ、有害性の制御において、モデル パラメーターの変更、内部レイヤーへのアクセス、補助モデルのトレーニングを行わずに、プロンプト ベースおよび以前のロジット レベルのベースラインを一貫して改善します。選択性とルックアップ テーブルのアブレーションは、ゲインが一般的なロジット摂動ではなくターゲット固有の統計スコアから来ていることを示しています。これらの結果は、ロジット介入が確率の高い候補の下でターゲット固有の統計によって導かれる場合、効果的なステアリングには学習済みコントローラーが必要ないことを示しています。

原文 (English)

Steering Language Models Before They Speak: Logit-Level Interventions

Controllable generation requires language models to realize output characteristics such as reading level, politeness, and toxicity. Existing steering methods are often indirect, require access to internal activations, or depend on auxiliary trained models. We propose SWAI, a training-free inference-time method that addresses these limitations by steering directly in logit space using corpus-derived token statistics. SWAI computes z-normalized one-vs-rest log-odds scores from labeled corpora and biases high-scoring tokens only within the model's top-K candidate set, allowing control to favor target-characteristic tokens while preserving contextually plausible choices. Across readability, politeness, and toxicity control, SWAI consistently improves over prompt-based and prior logit-level baselines without modifying model parameters, accessing internal layers, or training an auxiliary model. Selectivity and lookup-table ablations show that the gains come from target-specific statistical scores rather than generic logit perturbation. These results indicate that effective steering does not require learned controllers when the logit intervention is guided by target-specific statistics under high-probability candidates.

2026-05-29 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

ProtoMedAgent: プライバシーを意識したエージェントワークフローによるマルチモーダルな臨床解釈可能性

解釈可能なプロトタイプ ネットワークは、臨床診断に説得力のある症例ベースの推論を提供しますが、生の連続出力には医療文書に必要な意味構造が欠けています。標準の検索拡張生成 (RAG) によってこのギャップを埋めると、日常的に「検索お調子者」が引き起こされ、大規模言語モデル (LLM) が視覚的な予測と一致するように事後的な合理化を幻覚します。我々は、厳密な神経象徴的なボトルネックを克服する反復ゼロ勾配試験時間最適化問題としてマルチモーダルな臨床レポートを形式化するフレームワークである ProtoMedAgent を紹介します。凍結されたプロトタイプのバックボーン上で動作し、潜在的な視覚的および表形式の特徴を個別の意味記憶に抽出します。オンライン生成は、正確な集合論的微分と反射的な筆記者と批評家のループによって厳密に制約され、裏付けのない物語の主張を数学的に排除します。データ開示を安全に制限するために、$k$-匿名性と $\ell$-多様性によって管理されるセマンティック プライバシー ゲートを導入します。 4,160 人の患者の臨床コホートで評価された ProtoMedAgent は、91.2% の比較セット忠実度を達成し、標準的な RAG (46.2%) を根本的に上回っています。 ProtoMedAgent はさらに、拘束力のある $\ell$-diversity 相転移を利用して、アーティファクトレベルのメンバーシップ推論リスクを絶対 9.8% 体系的に削減します。

原文 (English)

ProtoMedAgent: Multimodal Clinical Interpretability via Privacy-Aware Agentic Workflows

While interpretable prototype networks offer compelling case-based reasoning for clinical diagnostics, their raw continuous outputs lack the semantic structure required for medical documentation. Bridging this gap via standard Retrieval-Augmented Generation (RAG) routinely triggers ``retrieval sycophancy,'' where Large Language Models (LLMs) hallucinate post-hoc rationalizations to align with visual predictions. We introduce ProtoMedAgent, a framework that formalizes multimodal clinical reporting as an iterative, zero-gradient test-time optimization problem over a strict neuro-symbolic bottleneck. Operating on a frozen prototype backbone, we distill latent visual and tabular features into a discrete semantic memory. Online generation is strictly constrained by exact set-theoretic differentials and a reflective Scribe-Critic loop, mathematically precluding unsupported narrative claims. To safely bound data disclosure, we introduce a semantic privacy gate governed by $k$-anonymity and $\ell$-diversity. Evaluated on a 4,160-patient clinical cohort, ProtoMedAgent achieves 91.2% Comparison Set Faithfulness where it fundamentally outperforms standard RAG (46.2%). ProtoMedAgent additionally leverages a binding $\ell$-diversity phase transition to systematically reduce artifact-level membership inference risks by an absolute 9.8%.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体

蒸留ゲーム: 適応的な攻撃と効率的な防御

蒸留攻撃は、モデルプロバイダーにとって展開のトレードオフを生み出します。モデルをより有用にする同じ出力が、模倣を容易にする可能性もあります。私たちは、効用に制約のある教師と適応的な生徒の間のミニマックス ゲームを通じて、このトレードオフを研究します。私たちのフレームワークは、扱いやすい片側応答ルール、つまり生徒が価値の高い例の重み付けを変更する適応評価ルールと、蒸留に最も有用な出力を抑制する教師側の防御テンプレートを生成します。安価なプロキシの例の値から、生成中に教師とプロキシの生徒を組み合わせた単純なフォワードパスのみの防御であるプロダクト オブ エキスパート (PoE) を導き出します。経験的に、適応的評価は受​​動的と適応的な大きなギャップを明らかにします。最先端の防御では、適応的な生徒は、GSM8K と MATH で受動的評価が示唆するよりも大幅に多くの能力を回復します。このより強力な評価の下では、高価な防御と PoE の間の見かけの堅牢性の差は大幅に縮小しますが、PoE は引き続き大幅に安価であり、より高品質の推論トレースが保存されます。全体として、私たちの結果は、強い蒸留を止めるのは依然として困難であり、反蒸留の進歩は受動的生徒ではなく適応的な生徒に対して評価されるべきであることを示唆しています。私たちのコードは https://github.com/ysfalh/distillation-game から入手できます。

原文 (English)

The Distillation Game: Adaptive Attacks & Efficient Defenses

Distillation attacks create a deployment trade-off for model providers: the same outputs that make a model more useful can also make it easier to imitate. We study this trade-off through a minimax game between a utility-constrained teacher and an adaptive student. Our framework yields tractable one-sided response rules: an adaptive evaluation rule in which the student reweights high-value examples, and a teacher-side defense template that suppresses outputs most useful for distillation. From a cheap proxy for example value, we derive Product-of-Experts (PoE), a simple forward-pass-only defense that combines the teacher with a proxy student during generation. Empirically, adaptive evaluation reveals a large passive--adaptive gap: on state-of-the-art defenses, adaptive students recover substantially more capability than passive evaluation suggests on GSM8K and MATH. Under this stronger evaluation, the apparent robustness gap between expensive defenses and PoE narrows considerably, while PoE remains substantially cheaper and preserves higher-quality reasoning traces. Overall, our results suggest that strong distillation remains difficult to stop, and that progress on antidistillation should be judged against adaptive students rather than passive ones. Our code is available at: https://github.com/ysfalh/distillation-game.

2026-05-29 13:00 JSTarXiv cs.AIハードウェア/半導体

EvoSpec: リアルタイム語彙とパラメータ適応ターゲットによる推測的デコーディングの進化

投機的デコードは、ドラフトしてから検証するというパラダイムを通じて大規模言語モデルの推論を加速しますが、語彙サイズが拡大するにつれて出力射影層がボトルネックになります。既存の静的プルーニング手法はこのオーバーヘッドを効果的に削減しますが、動的な分布の変化を捉えることができないため、特殊なドメインやトピック切り替えのシナリオでは受け入れ率が急激に低下するという問題があります。これに対処するために、動的な語彙とパラメーターの適応を通じてドラフト モデルのリアルタイムの進化を可能にするフレームワークである EvoSpec を導入します。静的または純粋な検索ベースのアプローチとは異なり、EvoSpec は効率的なセマンティックおよび統計的なインデックス作成を通じて重要なロングテール トークンを取得するコンテキスト認識メカニズムを採用しています。さらに、ドラフトモデルとターゲットモデル間の分布のギャップを継続的に最小化するために、カリキュラム学習を利用した軽量のオンライン調整戦略を提案します。専門領域 (コーディング、法律、医学) にわたる広範な評価により、EvoSpec が静的ベースラインの制限を克服していることが確認されています。 EAGLE-3 では、これらの設定で最先端の静的ベースライン FR-Spec と比較して 1.13 倍の高速化を実現し、標準のオンライン適応よりもメモリ オーバーヘッドが 27\% 低くなります。

原文 (English)

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation

Speculative decoding accelerates Large Language Model inference via a draft-then-verify paradigm, yet the output projection layer becomes a bottleneck as vocabulary sizes scale. While existing static pruning methods effectively reduce this overhead, they suffer from precipitous drops in acceptance rate in specialized domains or topic-switching scenarios due to their inability to capture dynamic distribution shifts. To address this, we introduce EvoSpec, a framework that enables real-time evolution of the draft model through dynamic vocabulary and parameter adaptation. Unlike static or purely retrieval-based approaches, EvoSpec employs a context-aware mechanism that retrieves critical long-tail tokens via efficient semantic and statistical indexing. Furthermore, we propose a lightweight online alignment strategy utilizing curriculum learning to continually minimize the distributional gap between the draft and target models. Extensive evaluations across specialized domains (coding, law, and medicine) confirm that EvoSpec overcomes the limitations of static baselines. On EAGLE-3, it achieves a 1.13x speedup in these settings over the state-of-the-art static baseline FR-Spec, with 27\% lower memory overhead than standard online adaptation.

2026-05-29 03:32 JSTTechCrunch AIハードウェア/半導体

Just like gold and oil, we’ll soon be able to trade AI token futures

Large exchanges are designing derivative products around AI tokens, which are increasingly being considered less a computational output and…

2026-05-28 22:00 JSTTechCrunch AIハードウェア/半導体

Has the hunt for AI compute uncovered the next Cerebras?

General Compute is betting SambaNova will be the next breakout chipmaker.

2026-05-28 19:25 JSTITmedia AI+ハードウェア/半導体

レノボ、国内に“水冷AIインフラ”の検証施設 GPUサーバ需要増で水冷活用促す

レノボ・ジャパンが水冷技術を活用したAIインフラの検証施設「Neptuneラボ」を新設した。レノボの冷却技術を使う顧客やパートナー企業に対し、本番に近い検証・PoC環境として提供する。クラウドベンダーやSIerとの共同検証を通し、推奨される機器構成などの策定にも役立てる。レノボ…

2026-05-28 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

エージェント用のクエリ エンジン

現在、実稼働環境で最も急速に増加しているデータは、エージェント トレース、チャット ログ、推論チェーン、モデル出力などの非構造化テキストです。人々はそれを分析したいと考えていますが、クエリ パスにモデルがないとテキストをクエリできないため、尋ねる価値のある質問 (「エージェントがどこで混乱したか教えてください」) は SQL だけでは答えることができません。この分析が行われる自然な場所は、クライアント側で実行され、同じプロセス内で人間のユーザーと LLM エージェントの両方をホストする新しいクラスの AI アプリケーション (Claude Code、Cursor、Claude Desktop、ブラウザ内エージェント) です。これらのアプリケーションはデータを操作する必要がますます高まっていますが、レイクハウスの読み取りパスは JS ランタイムから使用するのが難しく、Spark、Trino、およびマネージド ウェアハウスはそこに適合しません。この新しい種類の AI データ アプリケーションを構築するには、エンジンの 3 つのプロパティが一次になります。アプリケーションがすでに実行されているランタイムにドロップされる JS ネイティブ ディストリビューション、コールド タブまたはターンごとのエージェント サンドボックス内に出荷できるほど十分小さいバンドル、および分析オペレーターとモデルベースのテキスト解釈をインターリーブする方法です。我々は、合計 70 KB 未満の 3 つのオープンソース JavaScript ライブラリ (Hyparquet、Squirreling、Icebird) である Hyperparam を紹介します。これらは、Parquet と Apache Iceberg をオブジェクト ストレージから直接読み取り、セルごとの非同期ネイティブ SQL 実行で 3 番目のプロパティを満たすため、高価なセルはダウンストリーム オペレーターが要求した場合にのみ起動されます。 Squirreling は、フィルタ境界クエリでは DuckDB-WASM より 300 倍以上高速 (ソート境界クエリでは 192 倍) で LLM 形状の非同期 UDF を実行し、3 分の 2 のコストで 10 タスクのエージェント アナリスト スイートを完成させます。私たちは、専門分野としてのデータ エンジニアリングは、現在運用されている AI ネイティブのクライアント アプリケーションとそのユーザーと連携して動作するエージェントに合わせて更新する必要があると主張します。

原文 (English)

A Query Engine for the Agents

The fastest-growing data in production today is unstructured text: agent traces, chat logs, reasoning chains, model outputs. People want to analyze it, and the questions worth asking ("show me where the agent got confused") cannot be answered by SQL alone, since text is not queryable without a model in the query path. The natural place this analysis is happening is the new class of AI applications (Claude Code, Cursor, Claude Desktop, in-browser agents) that run client-side and host both a human user and an LLM agent in the same process. These applications increasingly want to work with data, but the lakehouse read path has been hard to use from a JS runtime: Spark, Trino, and managed warehouses do not fit there. To build this new kind of AI data application, three properties of the engine become first-order: a JS-native distribution that drops into the runtime the application already runs in, a bundle small enough to ship inside a cold tab or per-turn agent sandbox, and a way to interleave analytic operators with model-based interpretation of text. We present Hyperparam, three open-source JavaScript libraries (Hyparquet, Squirreling, Icebird) totaling under 70 KB, that read Parquet and Apache Iceberg directly from object storage and meet the third property with per-cell, async-native SQL execution, so expensive cells fire only when downstream operators demand them. Squirreling runs LLM-shaped async UDFs over 300x faster than DuckDB-WASM on filter-bounded queries (and 192x on sort-bounded queries) and completes a ten-task agent analyst suite at two-thirds lower cost. We argue that data engineering as a discipline needs to update for the AI-native client applications now in production and the agents that work alongside their users.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

いつ最適化すべきかを学ぶ: GPU カーネル系統の専門家による検証済みの最適化スキル

LLM ベースのエージェントは、GPU カーネルの生成にますます使用されていますが、多くの場合、それらの最適化がいつ適切であるかは分からずに、どのような最適化を試みるべきかはわかっています。 KLineage を導入します。KLineage は、この欠落している「いつ」の知識をエキスパート カーネルから学習します。KLineage は、前方ロールアウトに依存するのではなく、検証ゲートによる簡略化を通じてエキスパート実装を後方に導き、受け入れられた各ステップを逆に再利用可能な最適化スキルに変換します。各スキルは、最適化の意図だけでなく、それがコード内のどこに適用されるか、どのような条件で最適化が有効になったか、どのような効果があったのか、その前提によってどのような失敗が回避されたのかも記録します。ダウンストリーム LLM は、同じコンパイル/正確性/プロファイル ゲートの下で新しいコード サーフェス上でこれらのスキルを具体化します。 2 つの NVIDIA アーキテクチャにわたる 5 つのエキスパート ワークロードでは、これらの系統由来のスキルが効果的な最適化カリキュラムとして機能し、同じ固定予算の下で最終的なカーネル品質と最適化効率の両方において最近のメモリベースの LLM カーネル ベースラインを上回ります。さらに、ソースケースの記憶に対する健全性テストとして、別個の 22 インスタンスのホールドアウト チェックを使用します。

原文 (English)

Learning When to Optimize: Verified Optimization Skills from Expert GPU-Kernel Lineages

LLM-based agents are increasingly used to generate GPU kernels, but they often know what optimizations to try without knowing when those optimizations are sound. We introduce KLineage, which learns this missing "when" knowledge from expert kernels: instead of relying on forward rollouts, KLineage walks expert implementations backward through validation-gated simplifications and reverses each accepted step into a reusable optimization skill. Each skill records not only the optimization intent, but also where it applies in code, what conditions made it valid, what effect it had, and what failures its assumptions avoid. A downstream LLM materializes these skills on new code surfaces under the same compile/correctness/profile gate. On five expert workloads across two NVIDIA architectures, these lineage-derived skills serve as an effective optimization curriculum, exceeding recent memory-based LLM-kernel baselines in both final kernel quality and optimization efficiency under the same fixed budget. We additionally use a separate 22-instance held-out check as a sanity test against source-case memorization.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

生成モデルにおける幻覚のフィンガープリントとしてのエントロピー分布

大規模言語モデル (LLM) は、一般に幻覚と呼ばれる事実に反する出力を生成することが多く、信頼を損ない、一か八かの環境での展開を制限します。既存の幻覚検出方法では通常、複数のフォワード パスまたはモデル内部へのアクセスが必要です。この研究では、混乱または長さで正規化されたエントロピーによって捕捉された平均を超えるトークンレベルのエントロピーの分布が、独立した信号を運ぶ分布形状と尾部の動作を伴う幻覚の指紋として機能するという理論的背景と経験的証拠を提供します。私たちは幻覚検出を統計的仮説検定として形式化し、単一のフォワード パスとトークン ロジットへのブラック ボックス アクセスのみを必要とする軽量アルゴリズムである校正エントロピー スコア (CES) を提案します。 CES は、校正された参照 CDF を通じて生成されたエントロピーの平均信号と最大信号を結合し、モデルとタスク間で直接比較できるスコアを生成します。我々は、新しいランダム長の Dvoretzky-Kiefer-Wolfowitz 不等式を介して有限サンプルのキャリブレーション保証を確立し、また、CES が世代長において指数関数的に速く 1 に収束する確率で幻覚を検出することも証明します。 CES は、オープンソース モデルと API アクセス モデルにわたる 8 つの QA ベンチマークと 10 のジェネレーター モデルにわたって、すべてのシングルパス ブラック ボックス メソッドの中で最高の検出パフォーマンスを達成するとともに、既存のヒューリスティックにはない正式なエラー保証を提供します。注目すべきことに、CES は、はるかに大きな計算コストを必要とするマルチサンプル手法と統計的に区別がつかないため、軽量検出と高価な検出の間のギャップを埋め、リアルタイムの大規模展開に適しています。

原文 (English)

Entropy Distribution as a Fingerprint for Hallucinations in Generative Models

Large Language Models (LLMs) often generate factually incorrect outputs, commonly termed hallucinations, that undermine trust and limit deployment in high-stakes settings. Existing hallucination detection methods typically require multiple forward passes, or access to model internals. In this work, we provide theoretical background and empirical evidence that the distribution of token-level entropies, beyond the mean captured by perplexity or length-normalised entropy, serves as a fingerprint of hallucination, with distributional shape and tail behaviour carrying independent signal. We formalize hallucination detection as a statistical hypothesis test and propose the Calibrated Entropy Score (CES), a lightweight algorithm requiring only a single forward pass and black-box access to token logits. CES combines the mean signal with the maximum signal of the generated entropy through a calibrated reference CDF, producing scores that are directly comparable across models and tasks. We establish finite-sample calibration guarantees via a novel random-length Dvoretzky--Kiefer--Wolfowitz inequality, and also prove that CES detects hallucinations with probability converging to one exponentially fast in the generation length. Across eight QA benchmarks and ten generator models spanning open-source and API access models, CES achieves the highest detection performance among all single-pass black-box methods while providing formal error guarantees that existing heuristics lack. Remarkably, CES is statistically indistinguishable from multi-sample methods that require far greater computational cost, closing the gap between lightweight and expensive detection and making it suitable for real-time, large-scale deployment.

2026-05-28 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

SwarmHarness: 分散化されたインセンティブに合わせた AI エージェント ネットワークを介したスキルベースのタスク ルーティング

膨大な量のコンピューティング (個人のワークステーション、アイドル状態の推論サーバー、およびジョブ間のエッジ デバイス上の GPU サイクル) は、所有者がそれらを安全かつ収益性よく共有するためのインセンティブに合わせたプロトコルが存在しないため、未使用のままになります。既存のアプローチは、信頼できる中央コーディネーター (クラウド マーケットプレイス) を必要とするか、重いブロックチェーン インフラストラクチャ (Golem、BrokerChain) を要求するか、インセンティブ層が完全に欠如している (BOINC、Petals) かのいずれかです。私たちは SwarmHarness を提案します。これは、HarnessAPI スキル ノードが中央の権限を持たずに自己組織化してコンピューティング群を形成する分散型プロトコルです。 SwarmHarness には 3 つの連動コンポーネントがあります。1 つはピア検出と機能アドバタイズメント用の分散ハッシュ テーブル (DHT) 上に構築された SwarmRegistry です。 SwarmRouter は、機能、負荷、レイテンシー、信頼性に関するユーティリティ関数を使用してノードにタスクをディスパッチします。 SwarmCredit は、Shapley 値近似を介してコンピューティング クレジットの報酬を貢献ノードに割り当てるインセンティブ メカニズムです。ノードはタスクを処理することでクレジットを獲得し、タスクを送信するためにクレジットを消費します。まったく貢献しないアイドル ノードはクレジットを消耗し、ルーティングの優先順位を失い、自主規制型の参加経済を生み出します。ノードは報酬の高いスキルに特化し、ルーティング信号がデジタル フェロモンとして機能するため、ネットワークは生物学的な群れに似た創発的な集合知を示します。 SwarmHarness は、コンピューティングの共有を超えて、自律分散型 AI エージェント ネットワークの基本的なプリミティブであり、エージェントがコンピューティングを雇い、サブタスクをルーティングし、人間の仲介なしにクレジットを決済します。

原文 (English)

SwarmHarness: Skill-Based Task Routing via Decentralized Incentive-Aligned AI Agent Networks

Vast quantities of compute (GPU cycles on personal workstations, idle inference servers, and edge devices between jobs) go unused because no incentive-aligned protocol exists for their owners to share them safely and profitably. Existing approaches either require a trusted central coordinator (cloud marketplaces), demand heavy blockchain infrastructure (Golem, BrokerChain), or lack an incentive layer entirely (BOINC, Petals). We propose SwarmHarness, a decentralised protocol in which HarnessAPI skill nodes self-organise into a compute swarm without any central authority. SwarmHarness has three interlocking components: a SwarmRegistry built on a Distributed Hash Table (DHT) for peer discovery and capability advertisement; a SwarmRouter that dispatches tasks to nodes using a utility function over capability, load, latency, and trust; and SwarmCredit, an incentive mechanism that attributes compute-credit rewards to contributing nodes via a Shapley-value approximation. Nodes earn credits by serving tasks and spend credits to submit them; idle nodes that never contribute drain credits and lose routing priority, creating a self-regulating participation economy. As nodes specialise toward high-reward skills and routing signals act as digital pheromones, the network exhibits emergent collective intelligence analogous to biological swarms. Beyond compute sharing, SwarmHarness is a foundational primitive for autonomous distributed AI agent networks in which agents hire compute, route subtasks, and settle credits without human intermediation.

2026-05-28 13:00 JSTarXiv cs.AIハードウェア/半導体

EvoSpec: リアルタイム語彙とパラメータ適応ターゲットによる推測的デコーディングの進化

投機的デコードは、ドラフトしてから検証するというパラダイムを通じて大規模言語モデルの推論を加速しますが、語彙サイズが拡大するにつれて出力射影層がボトルネックになります。既存の静的プルーニング手法はこのオーバーヘッドを効果的に削減しますが、動的な分布の変化を捉えることができないため、特殊なドメインやトピック切り替えのシナリオでは受け入れ率が急激に低下するという問題があります。これに対処するために、動的な語彙とパラメーターの適応を通じてドラフト モデルのリアルタイムの進化を可能にするフレームワークである EvoSpec を導入します。静的または純粋な検索ベースのアプローチとは異なり、EvoSpec は効率的なセマンティックおよび統計的なインデックス作成を通じて重要なロングテール トークンを取得するコンテキスト認識メカニズムを採用しています。さらに、ドラフトモデルとターゲットモデル間の分布のギャップを継続的に最小化するために、カリキュラム学習を利用した軽量のオンライン調整戦略を提案します。専門領域 (コーディング、法律、医学) にわたる広範な評価により、EvoSpec が静的ベースラインの制限を克服していることが確認されています。 EAGLE-3 では、これらの設定で最先端の静的ベースライン FR-Spec と比較して 1.13 倍の高速化を実現し、標準のオンライン適応よりもメモリ オーバーヘッドが 27\% 低くなります。

原文 (English)

EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter AdaptationTarget

Speculative decoding accelerates Large Language Model inference via a draft-then-verify paradigm, yet the output projection layer becomes a bottleneck as vocabulary sizes scale. While existing static pruning methods effectively reduce this overhead, they suffer from precipitous drops in acceptance rate in specialized domains or topic-switching scenarios due to their inability to capture dynamic distribution shifts. To address this, we introduce EvoSpec, a framework that enables real-time evolution of the draft model through dynamic vocabulary and parameter adaptation. Unlike static or purely retrieval-based approaches, EvoSpec employs a context-aware mechanism that retrieves critical long-tail tokens via efficient semantic and statistical indexing. Furthermore, we propose a lightweight online alignment strategy utilizing curriculum learning to continually minimize the distributional gap between the draft and target models. Extensive evaluations across specialized domains (coding, law, and medicine) confirm that EvoSpec overcomes the limitations of static baselines. On EAGLE-3, it achieves a 1.13x speedup in these settings over the state-of-the-art static baseline FR-Spec, with 27\% lower memory overhead than standard online adaptation.

2026-05-28 13:00 JSTarXiv cs.AIハードウェア/半導体

事実の未来: 事実の生成と検証のギャップを追跡する

言語モデルは事実知識へのデフォルトのインターフェースになりつつありますが、多くの場合、言語モデルは出力を生成するよりも確実に検証します。この世代と検証のギャップ (GV ギャップ) は、自己改善と推論における最近の多くの進歩の基礎となっていますが、事実知識に関するその力学は、特によく理解されていないままです。私たちは、実際の G​​V ギャップの基礎となるトレーニング メカニズムに焦点を当て、GV ギャップを計算上の対応物や美的ギャップと区別します。 4 つのオープンソース モデル ファミリにわたって、それぞれ 2 つのスケールで 3 つのトレーニング フェーズ (取得、継続学習、更新) を通じて生成および検証の機能をトレースします。 3 つの発見はモデル全体で繰り返されます。(i) 検証は生成前に一貫して学習されます。 (ii) 検証は生成よりも継続的な学習に対して堅牢です。 (iii) 事実の更新により、モデルが「マルチバース」状態のままになる可能性があり、同時に古い答えと新しい答えの両方が正しいと検証されます。フロンティアモデルでの自然実験は、これらのダイナミクスを大規模に再現し、十分にカバーされた事実に対する残留検証バイアスを明らかにします。

原文 (English)

The Future of Facts: Tracing the Factual Generation-Verification Gap

Language models are becoming the default interface to factual knowledge, yet they often verify outputs more reliably than they generate them. This generation-verification gap (GV-gap) underlies many recent advances in self-improvement and reasoning, but its dynamics on factual knowledge specifically remain poorly understood. We focus on the training mechanisms underlying factual GV-gaps, distinguishing them from their computational and aesthetic counterparts. We trace generation and verification capabilities through three training phases (acquisition, continual learning, and updating) across four open-source model families at two scales each. Three findings recur across models: (i) verification is consistently learned before generation; (ii) verification is more robust to continual learning than generation; and (iii) factual updates can leave models in a "multi-verse" state, simultaneously verifying both old and new answers as correct. Natural experiments on frontier models reproduce these dynamics at scale and reveal residual verification biases on well-covered facts.

2026-05-28 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

エネルギーの盲点: NVIDIA の主力エッジ AI ハードウェアはプロセスレベルのエネルギー属性をサポートできない

単一のユーザー目標によって複数ステップのオーケストレーション、ツール呼び出し、再試行、障害回復がトリガーされるエージェントティック AI ワークロードは、エッジ導入のターゲットとなっており、NVIDIA、デル、HP、ASUS、MSI、Acer、ギガバイトのすべてが 2026 年に GB10 ベースのデスクトップ AI システムを出荷します。私たちは最近、オーケストレーション構造がエージェントのエネルギー コストの大半を占めていることを実証しました。ワークフローは、成功した目標ごとに線形ベースラインよりも 4.33 倍多くのエネルギーを消費します。マルチステップ推論タスクの OOI は 7.63 倍に達します。これとは別に、Rajat et al。 CPU 側の処理が、エージェント ワークロードの総レイテンシの最大 90.6%、総動的エネルギーの 44% を占めることが示されています。私たちは、ASUS Ascent GX10 (GB10 SoC) の系統的なエネルギー観測可能性監査を報告し、このプラットフォームでは、サポートされているソフトウェア インターフェイスを通じて、CPU エネルギー カウンター、INA パワーレール モニター、IPMI/BMC、および SCMI パワーキャップ プロトコルを公開していないことがわかりました。唯一のオンデバイス エネルギー テレメトリは、NVML を介した瞬間的な GPU 電力です。さらに、MediaTek ファームウェアが文書化されていない ACPI インターフェイス (SPBM) を介してレールごとのエネルギーを内部で計算していることも判明しましたが、NVIDIA は「CPU レール情報を公開する予定はない」と述べています。したがって、RAPL 経由で x86 上で実行されるデバイス上のプロセスごとのエネルギー アトリビューションは、サポートされているインターフェイスを介してこのプラットフォームでは再現できません。私たちは、エネルギーに起因する AI のハードウェア要件仕様を形式化し、GPU 減算と組み合わせた外部 DC メータリングを使用した暫定キャリブレーション ブリッジを提案し、SCMI パワーキャップを介して標準トラック パスを特定します。私たちの調査結果は、低炭素コンピューティング コミュニティに、第一級のハードウェア要件としてエネルギーの可観測性を要求する動機を与えています。

原文 (English)

The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution

Agentic AI workloads - where a single user goal triggers multi-step orchestration, tool calls, retries, and failure recovery - are being targeted for edge deployment, with NVIDIA, Dell, HP, ASUS, MSI, Acer, and Gigabyte all shipping GB10-based desktop AI systems in 2026. We recently demonstrated that orchestration structure dominates agentic energy cost, with workflows consuming 4.33x more energy per successful goal than linear baselines and OOI reaching 7.63x for multi-step reasoning tasks. Separately, Rajat et al. show that CPU-side processing accounts for up to 90.6% of total latency and 44% of total dynamic energy in agentic workloads. We report a systematic energy-observability audit of the ASUS Ascent GX10 (GB10 SoC) and find that the platform exposes no CPU energy counter, no INA power-rail monitor, no IPMI/BMC, and no SCMI powercap protocol through any supported software interface. The only on-device energy telemetry is instantaneous GPU power via NVML. We further discover that the MediaTek firmware already computes per-rail energy internally via an undocumented ACPI interface (SPBM), but NVIDIA states there are "no plans to expose CPU rail information." On-device per-process energy attribution - as performed on x86 via RAPL - is therefore not reproducible on this platform through supported interfaces. We formalize a hardware requirements specification for energy-attributed AI, propose an interim calibration bridge using external DC metering combined with GPU subtraction, and identify a standards-track path via SCMI powercap. Our findings motivate the low-carbon computing community to demand energy observability as a first-class hardware requirement.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

セマンティック フローの正則化: 多様でありながら一貫した応答を生成するように LLM を教える

大規模な言語モデルを微調整してペルソナやトーンで条件付けされた応答を生成すると、出力の多様性が大幅に制限されます。これをクロススタイル崩壊と呼んでいます。我々は、この崩壊の原因を、共有表現の下では多様な継続を抑制する傾向にあるクロスエントロピー目的にまで遡ります。我々は、条件付きフローマッチングを介して将来のセグメントの連続文エンコーダ埋め込みでバックボーンを監視する軽量の補助目標であるセマンティックフロー正則化(SFR)を提案します。確率的フローソースは、構築によりマルチモダリティを維持します。フローマッチングヘッドは推論時に破棄され、導入コストはゼロになります。大規模な産業対話データセット (Qwen3-32B、9 人のペルソナ) では、SFR は SFT よりも出力の多様性、スタイルの忠実度、応答品質を向上させます。さらに、公開されている LiveCodeBench-v5 (Qwen2.5-Coder-7B-Instruct) で検証します。SFR は一貫して pass@k を改善し、定型化された対話を超えた汎用性を確認します。 MBPP で制御された比較を行うと、マルチトークン予測が SFR の退化した特殊なケースであることが明らかになります。

原文 (English)

Semantic Flow Regularization: Teaching LLMs to Generate Diverse Yet Coherent Responses

When large language models are fine-tuned to generate persona- or tone-conditioned responses, their output diversity is severely limited--a failure we term Cross-Style Collapse. We trace this collapse to the cross-entropy objective, which under shared representations tends to suppress diverse continuations. We propose Semantic Flow Regularization (SFR), a lightweight auxiliary objective that supervises the backbone with continuous sentence-encoder embeddings of future segments via conditional flow matching. The stochastic flow source preserves multi-modality by construction; the flow-matching head is discarded at inference, adding zero deployment cost. On a large-scale industrial dialogue dataset (Qwen3-32B, 9 personas), SFR improves output diversity, style fidelity, and response quality over SFT. We further validate on the public LiveCodeBench-v5 (Qwen2.5-Coder-7B-Instruct), where SFR consistently improves pass@k, confirming generality beyond stylized dialogue. A controlled comparison on MBPP reveals Multi-Token Prediction to be a degenerate special case of SFR.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

LLM 推論の統合されたクロスアーキテクチャの解釈

LLM の推論方法の理解は、実際的な非対称性によって妨げられています。生成された出力は観察可能ですが、基礎となる推論パターンは不透明なままです。相互情報ピーク (MIP) やディープシンキング比 (DTR) などの単一のプローブに依存すると、真の推論構造を過小評価する危険があります。この欠陥に対応するために、LLM 推論の解釈可能性に対する統一されたアプローチを提供するように設計された、統合されたクロスアーキテクチャ推論 (IAR) フレームワークを紹介します。具体的には、まず、出力層で推論に重要なトークンを分離するために、帯域幅調整された MIP と Tukey IQR ピーク検出を組み合わせて使用​​することを提案します。次に、MIP で選択されたトークンと DTR ディープ トークンの間のオーバーラップ分析を実行して、これらのトークンのクロスレイヤーの軌跡を追跡しました。これにより、推論に重要なトークンが計算集約型であるかどうかも明らかになり、推論パターンがモデル層全体でどのように進化するかを理解することがさらに容易になります。最後に、マルチドメインの問題に対して Jaccard 安定性メトリックを適用して、MIP で識別されたトークンの推論の品質が保証されているかどうかを検証します。 4 つのドメイン (数学、コード、ロジック、常識) にわたる 3 つのモデル (Qwen-7B、Qwen-14B、および Llama-8B) に関する広範な実験により、アーキテクチャ全体にわたる IAR の一般化可能な解釈機能が実証されました。

原文 (English)

Integrated and Cross-Architecture Interpretation of LLM Reasoning

Understanding how LLMs reason is hindered by a practical asymmetry: while their generated outputs are observable, the underlying reasoning patterns remain opaque. Relying on single probes, such as Mutual Information Peak (MIP) or Deep-Thinking Ratio (DTR), risks underestimating the genuine inferential structure. To response this deficiency, we present an Integrated, cross-Architecture Reasoning (IAR) framework, designed to provide a unified approach to LLM reasoning interpretability. Specifically, we first propose to use bandwidth-calibrated MIP coupled with Tukey IQR peak-detection to isolate reasoning-crucial tokens at the output layer. Second, we performed an overlap analysis between MIP-picked tokens and DTR-deep tokens to trace the cross-layer trajectories of those tokens. This also discloses whether reasoning-crucial tokens are computation-intensive as well, further facilitating to understand how reasoning patterns evolve across model layers. Finally, we apply a Jaccard stability metric over multi-domain problems to verify if the MIP-identified tokens are reasoning quality-guaranteed. Extensive experiments on three models (Qwen-7B, Qwen-14B, and Llama-8B) across four domains (mathematics, code, logic, and common sense) demonstrate IAR's generalizable interpretation capabilities across architectures.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

言語モデルのエラーパターンを学習する

特定の妥当性制約を持つドメインの出力を生成する場合 (プログラムはコンパイルする必要があるなど)、LLM は少数の集中的な方法で失敗することがよくあります。たとえば、TypeScript の生成時に Python 関数名を使用するなどです。これらのエラー パターンは、実際に学習できる少数の制約を使用して表現できることがわかります。エラーパターンを捕捉するオブジェクトとしてドメインおよびLLMごとのシンボリック関数である \emph{prefix filters} を提案し、実際にプレフィックスフィルタを効率的に学習するためのアルゴリズムとして Palla を提案し、Palla を実装します。 Palla によって学習されたプレフィックス フィルターは、i) LLM のエラー パターンを定量的に分析するのに役立ち、ii) 制約付きサンプリング アルゴリズムを介してモデルの出力を制約するために使用できます。たとえば、Palla は TypeScript 生成における Qwen2.5-1.5B のコンパイル レートを 60% 以上向上させ、Qwen2.5-1.5B が制約なしで Llama3.1-8B と同様のパフォーマンスを達成できるようにします。

原文 (English)

Learning the Error Patterns of Language Models

When generating outputs for domains with specific validity constraints (e.g., a program should compile), LLMs often fail in a small number of focused ways: for example, by using Python function names when generating TypeScript. We observe that these error patterns can be represented using a small number of constraints that can be learned in practice. We propose \emph{prefix filters}, which are per-domain-and-LLM symbolic functions, as objects to capture the error patterns, Palla as an algorithm to learn prefix filters efficiently in practice, and implement Palla. Prefix filters learned by Palla i) help us quantitatively analyze the error patterns of LLMs, and ii) can be used to constrain the outputs of a model via constrained sampling algorithms. For example, Palla boosts compile rates for Qwen2.5-1.5B on TypeScript generation, by over 60%, allowing Qwen2.5-1.5B to achieve similar performance to Llama3.1-8B unconstrained.

2026-05-28 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

PICACO: 総合相関最適化による LLM の多元的インコンテキスト値調整

インコンテキスト学習は、大規模言語モデル (LLM) を人間の価値観に合わせて調整する大きな可能性を示しており、インコンテキスト アライメント (ICA) として知られる、コストのかかる事後トレーニングを行わずに、有害な出力を削減し、多様な好みに対応するのに役立ちます。しかし、LLM の入力プロンプトの理解は依然として不可知論的であり、価値観の緊張に対処する ICA の能力を制限しています。人間の価値観は本質的に多元的であり、刺激と伝統など、相反する要求を課すことがよくあります。したがって、現在の ICA メソッドは、LLM が単一のプロンプト内で複数の意図された値を調整するのに苦労し、不完全または偏った調整につながる、命令ボトルネックの課題に直面しています。これに対処するために、我々は新しい多元的 ICA 手法である PICACO を提案します。 PICACO は、微調整を行わずに、複数の値をナビゲートするメタ命令を最適化して、LLM のそれらの理解をより適切に引き出し、それらの調整を改善します。これは、指定された値と LLM 応答の間の総合的な相関関係を最大化することによって達成され、理論的には値の相関関係を強化しながら気が散るノイズを低減し、結果として効果的な値の指示が得られます。 5 つの値セットに関する広範な実験により、PICACO がブラックボックス LLM とオープンソース LLM の両方で適切に動作し、最近のいくつかの強力なベースラインを上回り、最大 8 つの異なる値にわたってより良いバランスを達成できることが示されました。

原文 (English)

PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization

In-Context Learning has shown great potential for aligning Large Language Models (LLMs) with human values, helping reduce harmful outputs and accommodate diverse preferences without costly post-training, known as In-Context Alignment (ICA). However, LLMs' comprehension of input prompts remains agnostic, limiting ICA's ability to address value tensions--human values are inherently pluralistic, often imposing conflicting demands, e.g., stimulation vs. tradition. Current ICA methods therefore face the Instruction Bottleneck challenge, where LLMs struggle to reconcile multiple intended values within a single prompt, leading to incomplete or biased alignment. To address this, we propose PICACO, a novel pluralistic ICA method. Without fine-tuning, PICACO optimizes a meta-instruction that incorporates multiple values to better elicit LLMs' understanding of them and improve alignment. This is achieved by maximizing the total correlation between specified values and LLM responses, which theoretically reinforces value conformity and reduces distractive noise, resulting in more effective instructions. Extensive experiments on five value sets show that PICACO works well with both black-box and open-source LLMs, outperforms several recent strong baselines, and achieves a better balance across up to 8 distinct values.

2026-05-28 05:10 JSTTechCrunch AIハードウェア/半導体

In more good news for Amazon, Snowflake signs $6B deal with AWS for AI CPU chips

Snowflake has signed a new, enormous five-year deal with Amazon to secure chips for AI usage. Nvidia is once again being put on notice.

2026-05-27 13:00 JSTarXiv cs.AIハードウェア/半導体

BrickAnything: 構造を意識したトークン化を使用した、ジオメトリ条件付きの構築可能なレンガの生成

3D 形状から物理的に構築可能なレンガ構造を生成するには、幾何学的再構成以上のものが必要です。出力は、個別のパーツの制約と構造の安定性も満たさなければなりません。既存のレンガ生成方法は、ターゲットの 3D 形状が事前定義された制約の下で実現可能な構造を許容しない場合に機能不全に陥る可能性があるヒューリスティック最適化に依存しているか、基礎となる 3D ジオメトリとアセンブリ関係を明示的にモデル化せずにブリック シーケンスを生成しています。この研究では、さまざまな 3D 表現から構築可能なレンガ構造を生成するための、ジオメトリ条件付き自己回帰フレームワークである BrickAnything を紹介します。 BrickAnything は、統一された幾何学的インターフェイスとして点群を使用し、アセンブリ制約の下でターゲット形状を再構築するレンガ シーケンスを予測します。ブリック間の構造依存関係をモデル化するために、ローカル接続関係を通じてブリック構造を表す構造認識ツリー トークン化を導入します。この定式化により、シーケンスの生成と物理的な構築プロセスの一貫性が高まり、無効な中間状態が減少します。さらに、安定性や幾何学的忠実度などの構築性の目標を向上させるために、トレーニング後の好みに基づくアライメント、妥当性制約のあるデコード、および適応的ロールバックを導入します。広範な実験により、BrickAnything が幾何学的に忠実で物理的に実現可能なブリック構造を生成すること、および提案されたトークン化により従来の順序付け戦略と比較してロールバックと再生成が効果的に削減されることが実証されました。

原文 (English)

BrickAnything: Geometry-Conditioned Buildable Brick Generation with Structure-Aware Tokenization

Generating physically buildable brick structures from 3D shapes requires more than geometric reconstruction: the output must also satisfy discrete part constraints and structural stability. Existing brick generation methods either rely on heuristic optimization, which can break down when the target 3D shape does not admit a feasible structure under predefined constraints, or generate brick sequences without explicitly modeling the underlying 3D geometry and assembly relations. In this work, we present BrickAnything, a geometry-conditioned autoregressive framework for generating buildable brick structures from diverse 3D representations. BrickAnything uses point clouds as a unified geometric interface and predicts brick sequences that reconstruct the target shape under assembly constraints. To model structural dependencies among bricks, we introduce a structure-aware tree tokenization, which represents brick structures through local attachment relations. This formulation makes sequence generation more consistent with the physical construction process, and reduces invalid intermediate states. We further introduce preference-based alignment post-training, validity-constrained decoding and adaptive rollback to improve buildability objectives such as stability and geometric fidelity. Extensive experiments demonstrate that BrickAnything produces geometrically faithful and physically realizable brick structures, and that the proposed tokenization effectively reduces rollback and regeneration compared with conventional ordering strategies.

2026-05-27 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体研究/論文

ScientistOne: 証拠の連鎖による人間レベルの自律的研究に向けて

自律的な研究エージェントは、競争力のあるソリューションとプロフェッショナルに見える原稿を作成しますが、その出力には、表面レベルの評価では検出できない検証可能性の欠陥、つまり、捏造された引用、再現不可能なスコア、実装から乖離した手法の説明が含まれています。私たちは 3 つの貢献を通じてこの問題に取り組みます。まず、証拠連鎖 (CoE) です。これは、すべての主張が証拠ソースまで追跡可能であることを要求する検証可能性フレームワークです。 2 つ目は、ScientistOne です。これは、文献レビュー、解決策の発見、論文執筆を通じて構築によって証拠チェーンを維持するエンドツーエンドの自律研究システムです。 3 つ目は、CoE 監査です。スコア検証、仕様違反、参照検証、メソッド コードの調整という 4 つの整合性チェックがすべてのシステムに均一に適用される事後監査です。 5 つのシステムと 5 つのフロンティア研究タスクにわたる 75 の論文にわたって、すべてのベースラインが少なくとも 1 つの系統的故障モードを示しています。幻覚参照率は 21% に達し、スコア検証に合格した論文はわずか 42% で、メソッドとコードの整合性は 20% ~ 80% の範囲です。 ScientistOne は、幻覚参照ゼロ (0/337)、完璧なスコア検証 (12/12)、最高のメソッドとコードの整合性 (14/15) を達成しながら、5 つのタスクすべてで人間の専門家のパフォーマンスと同等またはそれを上回っています。 ScientistOne はさらに、医用画像処理、きめ細かい認識、3D 知覚、言語モデリングにわたる 6 つの追加タスクに一般化し、パラメーター ゴルフでは最先端の成績を、ベースラインが完全に失敗する MLE ベンチ タスクでは金メダルを獲得しました。

原文 (English)

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence

Autonomous research agents produce competitive solutions and professional-looking manuscripts, yet their outputs contain verifiability failures undetectable by surface-level evaluation: fabricated citations, unreproducible scores, and method descriptions that diverge from the implementation. We address this through three contributions. First, Chain-of-Evidence (CoE), a verifiability framework requiring every claim to be traceable to its evidence source. Second, ScientistOne, an end-to-end autonomous research system that maintains evidence chains by construction throughout literature review, solution discovery, and paper writing. Third, CoE Audit, a post-hoc audit whose four integrity checks -- score verification, specification violation, reference verification, and method-code alignment -- apply uniformly to all systems. Across 75 papers spanning five systems and five frontier research tasks, every baseline exhibits at least one systematic failure mode: hallucinated reference rates reach 21%, score verification passes in as few as 42% of papers, and method-code alignment ranges from 20% to 80%. ScientistOne achieves zero hallucinated references (0/337), perfect score verification (12/12), and the highest method-code alignment (14/15), while matching or exceeding human expert performance on all five tasks. ScientistOne further generalizes to six additional tasks spanning medical imaging, fine-grained recognition, 3D perception, and language modeling, achieving state-of-the-art on Parameter Golf and gold medals on MLE-Bench tasks where baselines fail entirely.

2026-05-27 13:00 JSTarXiv cs.AIハードウェア/半導体

単一方向を超えて: 思考の連鎖が単純な拒否の方向性を混乱させる

大規模推論モデル (LRM) は、最終出力を生成する前に思考連鎖 (CoT) トレースを生成し、拒否などの制御メカニズムを複雑にする可能性のある動的な内部状態を導入します。単一方向部分空間によって拒否が媒介される命令調整型 LLM とは異なり、大規模推論モデル (LRM) での拒否はさらに CoT に依存します。 DeepSeek-R1-Distill-LLaMA-8B では、CoT が固定されている場合、アクティブ化ステアリングによって拒否が逆転するのはわずか 39% ですが、CoT を完全に削除するとこれが 70% に増加し、CoT が積極的に拒否を強化していることがわかります。モデルが活性化ステアリングの下で​​ CoT を再生成する 2 段階の介入では、94% のケースで拒否が逆転しますが、結果として得られる CoT だけでは、ステアリングが取り除かれた後でもこの効果の 48% が保持されます。これは、CoT がコンプライアンス信号を独立して伝送および再構築できることを示唆しています。これらの発見は、LRM での拒否が残留ストリームのアクティベーションと CoT で共同してエンコードされることを示しています。この共同アクティベーションにより、LRM はアクティベーション レベルの介入のみに対してより堅牢になりますが、CoT は代替の表面攻撃にさらされる可能性があります。

原文 (English)

Beyond a Single Direction: Chain-of-Thought Disrupts Simple Steering of Refusal

Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms such as refusal. Unlike instruction-tuned LLMs, where refusal is mediated by a single directional subspace, refusal in large reasoning models (LRMs) additionally depends on the CoT. In DeepSeek-R1-Distill-LLaMA-8B, activation steering reverses refusal in only 39% of cases when the CoT is kept fixed, but removing the CoT entirely increases this to 70%, indicating that the CoT actively reinforces refusal. In a two-stage intervention where the model regenerates its CoT under activation steering, refusal is reversed in 94% of cases, while the resulting CoT alone retains 48% of this effect even after steering is removed. This suggests that the CoT can carry and reconstruct the compliance signal independently. These findings indicate that refusal in LRMs is jointly encoded in residual stream activations and CoT. This joint activation makes LRM more robust against activation-level interventions alone, but exposes CoT to a possible alternative surface attack.

2026-05-27 13:00 JSTarXiv cs.AIハードウェア/半導体

アトリビューションの盲点: 言語モデルが取得されたコンテキストではなくメモリに依存していることを検出する

検索拡張生成は、外部証拠における地上言語モデルの出力を約束しますが、この分野には、取得されたコンテキストが実際に生成を制御するかどうかを検証する信頼できる方法がありません。これは、一か八かの展開の前提条件です。コンテキスト一貫性のある出力はコンテキストに支配された出力を意味するという標準的な前提は、取得されたドキュメントがモデルの事前トレーニング データと重複すると崩れます。モデルは完全にパラメトリック メモリから忠実に見えるテキストを生成でき、両方の経路で区別できない出力が得られます。私たちはこの失敗をアトリビューションの盲点と名付け、これに対処するために Computational Reality Monitoring (CRM) を導入します。 CRM は、認知科学の現実監視フレームワークから適応した原則を運用します。コンテキストの有無にかかわらず内部表現を比較すると、出力レベルのモニターが体系的に見逃している、メンバーシップ条件付きの表現の相違が明らかになります。 CRM は、個々の世代がどのソースを使用したかを証明しません。トレーニング前の曝露が測定可能な内部軌跡の痕跡を残すかどうかを検出し、ソースの帰属に必要な基盤を確立します。 3 つのファミリーにまたがる 9 つのモデル バリアントにわたって、この相違はアーキテクチャ固有のレイヤー パターンに集中し、ブロック レベルのノイズ介入による集中的なサポートを受け、ドメインが混同されたベンチマークでは崩壊しながらタスクとデータセット全体に一般化します。帰属の盲点は測定可能で、部分的に対処可能です。内部表現は、出力レベルでは目に見えない診断信号を伝達し、証拠の出所に関する内部の認識が外部の動作を制御するシステムの基盤を確立します。

原文 (English)

The Attribution Blind Spot: Detecting When Language Models Rely on Memory Rather Than Retrieved Context

Retrieval-augmented generation promises to ground language model outputs in external evidence, yet the field has no reliable way to verify whether retrieved context actually governs generation -- a prerequisite for any high-stakes deployment. The standard assumption, that context-consistent output implies context-governed output, breaks when the retrieved document overlaps with the model's pretraining data: the model can produce faithful-looking text entirely from parametric memory, and both pathways yield indistinguishable output. We name this failure the attribution blind spot and introduce Computational Reality Monitoring (CRM) to address it. CRM operationalizes a principle adapted from cognitive science's reality monitoring framework: comparing internal representations with and without context reveals membership-conditioned representational divergence that output-level monitors systematically miss. CRM does not certify which source an individual generation used; it detects whether pretraining exposure leaves a measurable internal trajectory signature, establishing a necessary substrate for source attribution. Across nine model variants spanning three families, this divergence concentrates in architecture-specific layer patterns, receives converging support from block-level noise intervention, and generalizes across tasks and datasets while collapsing on domain-confounded benchmarks. The attribution blind spot is measurable and partially addressable: internal representations carry a diagnostic signal invisible at the output level, establishing a foundation for systems whose internal awareness of evidence provenance governs their external behavior.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

マルチステークホルダー LLM 調整: 集計からの分解推定

複数の利害関係者のタスクでは、相反する好みを持つユーザーを満足させるために 1 つの出力が必要です。ホリスティック LLM ジャッジは、ユーティリティの推定とユーティリティの集計を混同し、不安定な暗黙的な重みを生成します。私たちは、利害関係者の満足度が分散している場合、この集計固有の \emph{重み付けノイズ} が大きなスコアの変動を引き起こす可能性があることを経験的および理論的に示します。私たちの実験では、こうした体重による変化は関係者の数とともに増加します。 \textsc{DecompR} を提案します。反事実に基づいて調整された重みは、候補をスコアリングする前にクエリ構造から固定されますが、役割ごとのユーティリティは独立して推定され、候補に依存する重みドリフトが除去され、推定ノイズが低減されます。

原文 (English)

Multi-Stakeholder LLM Alignment: Decomposing Estimation from Aggregation

Multi-stakeholder tasks require one output to satisfy users with conflicting preferences. Holistic LLM judges conflate utility estimation and utility aggregation, yielding unstable implicit weights. We show empirically and theoretically that this aggregation-specific \emph{weighting noise} can create large score shifts when stakeholder satisfaction is dispersed; in our experiments, these weight-induced shifts also increase with stakeholder count. We propose \textsc{DecompR}: counterfactual-calibrated weights are fixed from query structure before candidate scoring, while per-role utilities are estimated independently, removing candidate-dependent weight drift and reducing estimation noise.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

データに敏感なドメインの LLM 出力のニューロシンボリック検証 (拡張プレプリント)

一か八かのドメインに導入された LLM は、根本的な信頼性の課題に直面しています。幻覚、矛盾、プライバシーの脆弱性により、エラーが法的、財務的、または安全性に影響を及ぼす許容できないリスクが生じます。この論文では、LLM で生成されたコンテンツに補完的な保証を提供する、形式的記号手法とニューラル セマンティック分析を組み合わせたハイブリッド検証アーキテクチャを紹介します。このアーキテクチャでは、入力検証に論理的推論を採用し、完全性の特性を活用して、構造化された要件に対して決定可能な保証を提供します。出力検証では、埋め込みベースの意味論的類似性により、形式的な手法では表現力に欠ける文脈上の幻覚が検出されます。この分離は、並列のアクターベースのパイプラインで実現され、幻覚を生み出す分布バイアスを継承するプロンプトベースの自己検証アプローチの制限に対処します。提案されたアーキテクチャとタイプ認識検証方法は、Action Design Research によって開発された現実世界の医療機器損傷評価レポート システムである HAIMEDA を使用して検証されています。評価の結果、構造化エンティティの幻覚検出率は 83% 以上、セマンティック捏造の幻覚検出率は 72% 以上で、レポート作成時間が 30% 短縮されたことが示され、神経記号アーキテクチャがデータに敏感なドメインでの LLM 展開に原則に基づいた保護手段を提供できることが実証されました。

原文 (English)

Neuro-Symbolic Verification of LLM Outputs for Data-Sensitive Domains (extended preprint)

LLMs deployed in high-stakes domains face fundamental reliability challenges: hallucinations, inconsistencies, and privacy vulnerabilities introduce unacceptable risks where errors carry legal, financial, or safety consequences. This paper presents a hybrid verification architecture combining formal symbolic methods with neural semantic analysis to provide complementary guarantees for LLM-generated content. This architecture employs logical reasoning for input verification, leveraging completeness properties to provide decidable guarantees on structured requirements. For output validation, embedding-based semantic similarity detects contextual hallucinations where formal methods lack expressiveness. This separation is realized in a parallel, actor-based pipeline, addressing limitations of prompt-based self-verification approaches, which inherit the distributional biases that produce hallucinations. The proposed architecture and type-aware verification method are validated with HAIMEDA, a real-world medical device damage assessment reporting system developed through Action Design Research. Evaluation shows hallucination detection rates of over 83% for structured entities and 72% for semantic fabrications, with a 30% reduction in report creation time, demonstrating that neuro-symbolic architectures can provide principled safeguards for LLM deployment in data-sensitive domains.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Xe-Forge: Intel GPU 向けのマルチステージ LLM を利用したカーネル最適化

深層学習アルゴリズムを新しいハードウェア アクセラレータに移植するには、開発者は同じ低レベルの最適化 (量子化、メモリ アクセスの結合、タイル サイズの調整、アーキテクチャ固有の回避策) をコードベース内のすべての Triton カーネルに繰り返し適用する必要があります。この手動での繰り返しの作業が大きなボトルネックとなっています。各カーネルは、デバイスごとに異なるハードウェア制約に対して同じサイクルの試行錯誤プロファイリングを必要としますが、基礎となる最適化パターンはほぼ一貫しています。 Intel GPU 向けにこのプロセスを自動化する、マルチステージ LLM を利用したパイプラインである Xe-Forge を紹介します。機能的に正しい Triton カーネルが与えられると、システムは、アルゴリズムの再構築と演算子の融合からブロック ポインターの最新化、GPU 固有のチューニング、およびオープンエンドの検出まで、最大 9 つの最適化ステージを適用します。各ステージは、候補を生成し、実際のハードウェアで検証し、失敗時に反復する検証と洗練の連鎖 (CoVeR) エージェントによって駆動されます。厳選されたナレッジベースは、LLM トレーニング データには存在しない Intel GPU 制約 (2 のべき乗ワープ数、GRF モード、SLM サイズ設定) をエンコードし、モデルをアーキテクチャ的に有効な範囲内に保ちます。 97 個の Level-2 KernelBench カーネルで Xe-Forge を、Intel Arc Pro B70 で Flash Attend を評価し、PyTorch Eager に比べて幾何平均 1.17 倍の速度向上を達成し、カーネルの 67% が向上し、9 個のカーネルで 5 倍を超え (最大 82 倍)、テスト済みのすべての構成で回帰なしで Flash Attend で 2 ~ 13.3 倍の高速化を達成しました。構造化されたドメインの知識が実証されています。ハードウェアインザループ検証を使用すると、現在新しいアクセラレータでのアルゴリズムの展開を妨げている反復的な移植作業を体系的に排除できます。

原文 (English)

Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU

Porting deep learning algorithms to new hardware accelerators requires developers to repeatedly apply the same low-level optimizations -- quantization, memory access coalescing, tile size tuning, and architecture-specific workarounds -- to every Triton kernel in their code-base. This manual, repetitive effort is a major bottleneck: each kernel demands the same cycle of trial-and-error profiling against hardware constraints that vary across devices, yet the underlying optimization patterns remain largely consistent. We present Xe-Forge, a multi-stage LLM-powered pipeline that automates this process for Intel GPU. Given a functionally correct Triton kernel, the system applies up to nine optimization stages -- from algorithmic restructuring and operator fusion through block pointer modernization, GPU-specific tuning, and open-ended discovery -- each driven by a Chain-of-Verification-and-Refinement (CoVeR) agent that generates candidates, validates them on real hardware, and iterates on failures. A curated knowledge base encodes Intel GPU constraints (power-of-two warp counts, GRF modes, SLM sizing) that are absent from LLM training data, keeping the model within architecturally valid bounds. We evaluate Xe-Forge on 97 Level-2 KernelBench kernels and Flash Attention on the Intel Arc Pro B70, achieving a 1.17x geometric mean speedup over PyTorch eager with 67% of kernels improving, nine kernels exceeding 5x (up to 82x), and 2--13.3x speedups on Flash Attention across all tested configurations without regression -- demonstrating that structured domain knowledge with hardware-in-the-loop verification can systematically eliminate the repetitive porting effort that currently gates algorithm deployment on new accelerators.

2026-05-27 13:00 JSTarXiv cs.AIハードウェア/半導体研究/論文

ストリーミング時系列からの時間遅延システムの動的混合のモデル化

この研究は、明確な入出力関係を持つ時系列データ ストリームにおける適応モデリングの問題に取り組んでいます。この問題は、環境要因や入力遅延の変化によって引き起こされる急速なシステム変更 (レジーム シフト) によってモデルのパフォーマンスが低下し、時系列パターンごとに複数の小さなモデルを使用する場合に精度、堅牢性、メモリ使用量の間でトレードオフが発生するため、困難です。これらの問題に対処するために、この論文では、ストリーミング時系列を時間遅延システムの動的な混合として扱うオンライン フレームワーク/方法を紹介します。このフレームワークは、システム ダイナミクスと入出力遅延の両方をキャプチャする固定長表現を使用して過去のレジームを要約することにより、モデル追跡の堅牢性を維持し、メモリ使用量を削減します。具体的には、このアプローチはシステムのマルコフ パラメーター系列を使用して要約システム テンソルを構築し、動的挙動と遅延特性の両方をキャプチャします。必要に応じて、テンソル分解アルゴリズムがテンソルから関連する過去のモデルを抽出し、現在のレジームに最適なシステムの選択に役立ちます。この方法により、環境変化への迅速な適応が可能になり、計算効率が高くなります。実際のデータセットでのテストでは、DelayMix が他の方法よりも常に優れたパフォーマンスを示し、特に非定常性の高いデータに対して、優れた予測精度と遅延へのより迅速な適応を実現することが示されています。

原文 (English)

Modeling Dynamic Mixtures of Time-Delay Systems from Streaming Time Series

This research addresses the problem of adaptive modeling in time-series data streams with clear input-output relationships. This problem is challenging because rapid system changes (regime shifts) caused by environmental factors or input delay changes degrade model performance, and the trade-off among accuracy, robustness, and memory usage arises when using multiple small models for each time-series pattern. To address these issues, this paper presents an online framework/method that treats streaming time series as dynamic mixtures of time-delay systems. This framework maintains robustness of model tracking and reduces memory usage by summarizing past regimes using a fixed-length representation that captures both the system dynamics and input-output delays. Concretely, this approach constructs a summary system tensor using the system's Markov parameter series, capturing both dynamic behavior and delay characteristics. If necessary, a tensor decomposition algorithm extracts relevant past models from the tensor and helps select the system that best fits the current regime. This method enables rapid adaptation to environmental changes and is computationally efficient. Tests on real datasets show that DelayMix consistently outperforms other methods, achieving superior forecast accuracy and faster adaptation to delays, especially for highly non-stationary data.

2026-05-27 13:00 JSTarXiv cs.AIハードウェア/半導体

正しいデモンストレーションが有害な場合: 文脈に沿った学習における模範の役割を再考する

インコンテキスト学習 (ICL) は、多くの場合、デモンストレーションが正しい入出力例を提供するため役立つという直観によって動機付けられます。しかし、我々は直観に反する現象を明らかにしました。正確さは模範の有用性を保証するものではなく、一部の正しい実証は ICL の精度を低下させる可能性さえあります。この正しさと有用性のギャップを研究するために、サンプル入力のみが変更され、サンプルは同じタスクの正しいインスタンスのままであるタスク保存摂動を導入します。具体的には、摂動された各サンプルには、タスク マッピングによって引き起こされるターゲットが割り当てられます。このフレームワークは、タスク関連のセマンティクスが変更されターゲットが再計算されるラベル更新の摂動と、元のターゲットが有効なままであるより厳密なターゲット保存の摂動の両方をカバーします。結果として生じる失敗モードを文脈上の証拠のシフトとして形式化します。タスク保存の摂動は、文脈上の推論のためにモデルによって使用される証拠の効果的な混合を変更する可能性があり、それによって見本の正しさと見本の有用性を分離します。感情分類、論理的推論、および数学の文章題全体にわたって、タスクを保持する摂動デモンストレーションは、特に小規模なモデル、より困難なタスク、より高い摂動比の場合に、ICL のパフォーマンスを大幅に低下させる可能性があることがわかりました。私たちの結果は、堅牢な ICL では、デモンストレーションが正しいかどうかだけでなく、デモンストレーションが文脈上の推論にどのように影響するかを評価する必要があることを示しています。コードは https://github.com/Chenghao-Qiu/Task-Preserving-ICL で入手できます。

原文 (English)

When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning

In-context learning (ICL) is often motivated by the intuition that demonstrations help because they provide correct input-output examples. However, we reveal a counterintuitive phenomenon: correctness does not guarantee exemplar utility, and some correct demonstrations can even reduce ICL accuracy. To study this correctness-utility gap, we introduce task-preserving perturbations, where only the exemplar input is changed, while the example remains a correct instance of the same task. Concretely, each perturbed exemplar is assigned the target induced by the task mapping. This framework covers both label-updating perturbations, where task-relevant semantics change and targets are recomputed, and stricter target-preserving perturbations, where the original target remains valid. We formalize the resulting failure mode as contextual evidence shift: task-preserving perturbations can change the effective mixture of evidence used by the model for contextual inference, thereby separating exemplar correctness from exemplar utility. Across sentiment classification, logical reasoning, and math word problems, we find that task-preserving perturbed demonstrations can substantially degrade ICL performance, especially for smaller models, harder tasks, and higher perturbation ratios. Our results show that robust ICL requires evaluating not only whether demonstrations are correct, but also how they influence contextual inference. Code is available at https://github.com/Chenghao-Qiu/Task-Preserving-ICL.

2026-05-27 13:00 JSTarXiv cs.AIハードウェア/半導体

GAN をトレーニングするためのクロススケールで調整された監視

最新の GAN では、多くの場合、中間ジェネレーターの出力に敵対的な監視が導入され、その結果得られる多段階合成が粗いから細かいまでの階層生成として解釈されます。この作品では、この解釈に挑戦します。私たちは、標準的なスケールごとの敵対的監視では、適切な粗いものから細かいものまでの階層が構築されていないと主張します。各中間画像は、独自の解像度で実際の分布に向けて独立してプッシュされますが、このスケールごとの現実主義では、ステージ間の出力が生成された同一のサンプルを表すことが保証されません。さらに、各段階で生成されるスケール固有の画像は、後続の段階の明示的な改良ターゲットとしては使用されません。したがって、その敵対的損失により、後のステージが同じサンプル軌道を保持するように制約することなく、スケール固有の出力を向上させることができ、前の出力を改良するのではなく、別のサンプルに向かって移動できるようになります。この問題をスケール間軌道不整合問題と呼びます。これを解決するために、マルチスケールの敵対的生成のためのクロススケール アラインメント トランスフォーマーである CAT を提案します。 CAT はディスクリミネーターをスケールごとに維持するため、各中間出力は独自の解像度で評価されると同時に、中間出力を最終出力と一致させる単純なジェネレーター側の整合性正則化が追加されます。クラス条件付き ImageNet-256 では、CAT-H/2 はわずか 60 トレーニング エポック後のワンステップ推論で 1.56 の FID-50K を達成し、強力なワンステップ GAN および拡散/フロー ベースラインを上回ります。

原文 (English)

Cross-scale Aligned Supervision for Training GANs

Modern GANs often introduce adversarial supervision on intermediate generator outputs and interpret the resulting multi-stage synthesis as coarse-to-fine hierarchical generation. In this work, we challenge this interpretation. We argue that standard scale-wise adversarial supervision does not construct a proper coarse-to-fine hierarchy: each intermediate image is independently pushed toward the real distribution at its own resolution, but this scale-wise realism does not ensure that outputs across stages represent the identical generated sample. Moreover, the scale-specific image produced at each stage is not used as an explicit refinement target for the subsequent stage. Therefore, its adversarial loss can improve a scale-specific output without constraining later stages to preserve the same sample trajectory, allowing them to move toward a different sample rather than refine the previous output. We refer to this problem as a cross-scale trajectory misalignment problem. To resolve it, we propose CAT, a Cross-scale Aligned Transformer for multi-scale adversarial generation. CAT keeps the discriminator scale-wise, so each intermediate output is evaluated at its own resolution, while adding a simple generator-side consistency regularization that aligns intermediate outputs with the final output. On class-conditional ImageNet-256, CAT-H/2 achieves an FID-50K of 1.56 with one-step inference after only 60 training epochs, outperforming strong one-step GAN and diffusion/flow baselines.

2026-05-27 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

Verus-SpecGym: 仕様の自動形式化を評価するためのエージェント環境

AI コーディング エージェントは現実世界のソフトウェアの作成にますます使用されていますが、その出力が正しいことを保証することは依然として根本的な課題です。正式な検証は、有望な方法を提供します。エージェントは、機械チェックされた証明とともにコードを生成し、コードが正式な仕様を満たしていることを保証します。ただし、正式な仕様自体がユーザーの意図と一致しているという保証はありません。この研究では、仕様の自動形式化、つまり LLM エージェントが非形式的なプログラミングの問題を忠実な形式的な仕様に変換できるかどうかを研究します。 Rust の検証ツールである Verus を対象とした Codeforces の問題から派生した 581 の仕様作成タスクのベンチマークである Verus-SpecBench と、モデルが Verus、bash、ファイルシステムと対話して仕様を開発するエージェント環境である Verus-SpecGym を紹介します。中心的な課題は評価です。専門家が作成したリファレンス仕様の作成には費用がかかり、LLM 審査員は微妙な間違いを見逃す可能性があります。私たちは、(a) 生成されたスペックを Rust コードとして実行できるように Verus の exec_spec メカニズムを拡張し、(b) Codeforces の公式テストと Codeforces の「ハック」から抽出された敵対的ケース (不正なソリューションを打ち破るために競合他社が作成したエッジ ケース) に対してそれらをテストすることで、この問題に対処します。 Verus-SpecBench では、最も強力なモデルである Gemini 3.1 Pro がタスクの 77.8% を解決しますが、他のフロンティア モデルは 51.1 ~ 57.8% を解決し、OSS モデルは 21.5 ~ 25.5% にすぎません。故障モードの分析では、モデルで生成された仕様が重要な入力仮定を省略し、誤った出力を受け入れ、有効な出力を拒否する可能性があることを示しています。また、LLM-as-a-judge 評価では、評価者が検出した失敗の 26% を見逃していることもわかりました。全体として、私たちの結果は、仕様の自動形式化はフロンティアエージェントにとって手の届くところにあるものの、すでに正しいコードを生成できる問題に対してさえ脆弱なままであることを示唆しています。コード、データ、ログは https://github.com/formal-verif-is-cool/verus-spec-gym にあります。

原文 (English)

Verus-SpecGym: An Agentic Environment for Evaluating Specification Autoformalization

AI coding agents are increasingly used to write real-world software, but ensuring that their outputs are correct remains a fundamental challenge. Formal verification offers a promising path: an agent generates code together with a machine-checked proof, guaranteeing that the code satisfies a formal specification. However, there is no guarantee that the formal spec itself matches the user's intent. In this work, we study specification autoformalization: whether LLM agents can translate informal programming problems into faithful formal specifications. We introduce Verus-SpecBench, a benchmark of 581 spec-writing tasks derived from Codeforces problems targeting Verus, a verifier for Rust, and Verus-SpecGym, an agentic environment in which models interact with Verus, bash, & the filesystem to develop these specs. The central challenge is evaluation: expert-written reference specs are expensive to write, & LLM judges can miss subtle mistakes. We address this by (a) extending Verus's exec_spec mechanism so that generated specs can be executed as Rust code, & (b) testing them against official Codeforces tests & adversarial cases extracted from Codeforces "hacks", which are edge cases written by competitors to break incorrect solutions. On Verus-SpecBench, the strongest model, Gemini 3.1 Pro, solves 77.8% of tasks, other frontier models solve 51.1--57.8% & OSS models reach only 21.5--25.5%. Our analysis of failure modes shows that model-generated specs can omit important input assumptions, accept incorrect outputs, & reject valid ones. We also find that LLM-as-a-judge evaluation misses 26% of the failures our evaluator catches. Overall, our results suggest that spec autoformalization is within reach for frontier agents but remains brittle even on problems where they can already generate correct code. The code, data, & logs can be found at https://github.com/formal-verif-is-cool/verus-spec-gym

2026-05-27 13:00 JSTarXiv cs.AIハードウェア/半導体

基本的な前方後方分割誘起ネットワークの深層限界と安定性解析 (II): 学習問題

反復最適化スキームと数値常微分方程式/偏微分方程式 (ODE/偏微分方程式) から派生した深く展開するニューラル ネットワークは、過去 10 年間にわたってデータ サイエンスにおいて大きな注目を集めてきました。その中で、多数の重要なネットワーク アーキテクチャが、基本的な前方後方分割 (FBS) アルゴリズムから構築されました。この論文では、最も基本的な FBS 誘導ネットワーク、つまり直接パラメータ緩和を組み込むことで元の FBS アルゴリズムから展開されたアーキテクチャに関する研究を続けます。以前の順方向システム分析における差分/差分包含の定式化に続いて、ここでは対応する学習問題の理論的側面をいくつか検討します。いくつかの穏やかな仮定の下で、基本的な FBS 誘導ネットワークの学習問題の深層制限システムの学習問題への一般的な収束特性を確立します。これは、ネットワークの最適な学習パラメーターの任意のクラスター点が深層制限システムの学習問題の解であることを示す $\Gamma$ 収束引数を意味します。これらの学習問題の摂動安定性の定性分析も示します。簡単な数値実験を行って、主要な一般収束結果を検証します。

原文 (English)

Deep-layer limit and stability analysis of the basic forward-backward-splitting induced network (II): learning problems

Deep unfolding neural networks derived from iterative optimization schemes and numerical ordinary/partial differential equations (ODEs/PDEs) have attracted much attention in data science over the last decade. Therein, numerous important network architectures were constructed from the basic forward-backward-splitting (FBS) algorithm. In this paper, we continue our research on the most basic FBS-induced network, an architecture unrolled from the original FBS algorithm by incorporating direct parameter relaxations. Following the difference/differential inclusion formulations in our previous forward system analyses, we here consider some theoretical aspects of corresponding learning problems. Under some mild assumptions, we establish a general convergence property of the training problem of the basic FBS-induced network to the learning problem of the deep-layer limit system, implying a $\Gamma$-convergence argument showing that any cluster point of the optimal learning parameters for the network is a solution to the learning problem of the deep-layer limit system. A qualitative analysis of perturbation stabilities of these learning problems is also presented. A simple numerical experiment is conducted to validate our main general convergence result.

2026-05-27 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

実行可能な操作認識によるエージェント ランタイムの管理された進化

エージェント システムの最近の進歩により、コードは使い捨ての出力成果物としてではなく、実行可能な操作基盤として扱われることが増えています。 \emph{Code as Agent Harness} などの以前の研究では、検証されたエージェント生成アーティファクトを、長期実行される認知ループ内で作成、実行、修正、永続化、再利用できるランタイム エンティティとしてフレーム化しました。ただし、そのようなアーティファクトのガバナンス、ライフサイクル管理、および運用の進化は依然として仕様が不十分です。この論文では、実行可能な操作認識を通じて、マルチエージェント システムにおける管理されたランタイム進化のためのフレームワークを提案します。私たちは、エージェントによって生成されたアーティファクトを、一時的な中間出力ではなく、徐々に運用基盤の一部となる永続的なランタイム機能として形式化します。この観点に基づいて、明示的な検証、トレーサビリティ、評価、ロールバック制約の下で動作する、ライフサイクルを意識したランタイム適応のための管理されたメカニズムとして \emph{HarnessMutation} を導入します。提案されたフレームワークは、実行時の適応を無制限の自己変更として扱うのではなく、永続的な操作メモリ上の制限された観察可能なプロセスとして進化をモデル化します。さらに、これらのアイデアを最新のエージェント ランタイムとガバナンス指向のオーケストレーション システム上でどのように運用できるかを示し、進化が明示的、監査可能、かつ制約されたままである適応型インフラストラクチャの概念的基盤を提供します。

原文 (English)

Governed Evolution of Agent Runtimes through Executable Operational Cognition

Recent advances in agentic systems increasingly treat code as an executable operational substrate rather than as a disposable output artifact. Prior work such as \emph{Code as Agent Harness} frames validated agent-generated artifacts as runtime entities that can be created, executed, revised, persisted, and reused within long-running cognitive loops. However, the governance, lifecycle management, and operational evolution of such artifacts remain under-specified. This paper proposes a framework for governed runtime evolution in multi-agent systems through executable operational cognition. We formalize agent-generated artifacts as persistent runtime capabilities that progressively become part of the operational substrate rather than transient intermediate outputs. Building on this perspective, we introduce \emph{HarnessMutation} as a governed mechanism for lifecycle-aware runtime adaptation operating under explicit validation, traceability, evaluation, and rollback constraints. Rather than treating runtime adaptation as unrestricted self-modification, the proposed framework models evolution as a bounded and observable process over persistent operational memory. It further shows how these ideas can be operationalized over modern agent runtimes and governance-oriented orchestration systems, providing a conceptual foundation for adaptive infrastructures whose evolution remains explicit, auditable, and constrained.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

PICACO: 総合相関最適化による LLM の多元的インコンテキスト値調整

インコンテキスト学習は、大規模言語モデル (LLM) を人間の価値観に合わせて調整する大きな可能性を示しており、インコンテキスト アライメント (ICA) として知られる、コストのかかる事後トレーニングを行わずに、有害な出力を削減し、多様な好みに対応するのに役立ちます。しかし、LLM の入力プロンプトの理解は依然として不可知論的であり、価値観の緊張に対処する ICA の能力を制限しています。人間の価値観は本質的に多元的であり、刺激と伝統など、相反する要求を課すことがよくあります。したがって、現在の ICA メソッドは、LLM が単一のプロンプト内で複数の意図された値を調整するのに苦労し、不完全または偏った調整につながる、命令ボトルネックの課題に直面しています。これに対処するために、我々は新しい多元的 ICA 手法である PICACO を提案します。 PICACO は、微調整を行わずに、複数の値をナビゲートするメタ命令を最適化して、LLM のそれらの理解をより適切に引き出し、それらの調整を改善します。これは、指定された値と LLM 応答の間の総合的な相関関係を最大化することによって達成され、理論的には値の相関関係を強化しながら気が散るノイズを低減し、結果として効果的な値の指示が得られます。 5 つの値セットに関する広範な実験により、PICACO がブラックボックス LLM とオープンソース LLM の両方で適切に動作し、最近のいくつかの強力なベースラインを上回り、最大 8 つの異なる値にわたってより良いバランスを達成できることが示されました。

原文 (English)

PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization

In-Context Learning has shown great potential for aligning Large Language Models (LLMs) with human values, helping reduce harmful outputs and accommodate diverse preferences without costly post-training, known as In-Context Alignment (ICA). However, LLMs' comprehension of input prompts remains agnostic, limiting ICA's ability to address value tensions--human values are inherently pluralistic, often imposing conflicting demands, e.g., stimulation vs. tradition. Current ICA methods therefore face the Instruction Bottleneck challenge, where LLMs struggle to reconcile multiple intended values within a single prompt, leading to incomplete or biased alignment. To address this, we propose PICACO, a novel pluralistic ICA method. Without fine-tuning, PICACO optimizes a meta-instruction that navigates multiple values to better elicit LLMs' understanding of them and improve their alignment. This is achieved by maximizing the total correlation between specified values and LLM responses, theoretically reinforcing value correlation while reducing distractive noise, resulting in effective value instructions. Extensive experiments on five value sets show that PICACO works well with both black-box and open-source LLMs, outperforms several recent strong baselines, and achieves a better balance across up to 8 distinct values.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

MetaSICL: メタ音声インコンテキスト学習による Audiroty LLM の適応

聴覚大規模言語モデル (LLM) は、幅広い音声理解タスクにわたって強力なパフォーマンスを実証しています。それにもかかわらず、リソースが少ないタスクに適用すると、苦労することがよくあります。ドメイン内のラベル付きデータが不足しているか、実際のテスト分布と一致しない場合、直接の微調整は脆弱になる可能性があります。 In-Context Learning (ICL) は、いくつかのドメイン内デモンストレーションを条件付けして聴覚 LLM を適応させることにより、トレーニング不要の推論時間ソリューションを提供します。この研究では、$\textit{Vanilla ICL}$ が、選択されたモデルのさまざまな音声タスクおよびオーディオ タスクにわたってゼロショット パフォーマンスを向上させることを最初に示します。これは、この ICL 適応機能がマルチモーダル設定に一般化できることを示唆しています。これに基づいて、$\textbf{Meta Speech In-Context Learning (MetaSICL)}$ を提案します。これは、モデルのインコンテキスト学習能力を強化することを目的とした、さまざまなタスクからの高リソース音声データのみを利用するトレーニング後のレシピです。実験によれば、私たちが提案した方法は、リソースが少ないシナリオでは直接の微調整よりも優れています。

原文 (English)

MetaSICL: Adapting Audiroty LLM via Meta Speech In-Context Learning

Auditory Large Language Models (LLMs) have demonstrated strong performance across a wide range of speech and audio understanding tasks. Nevertheless, they often struggle when applied to low-resource tasks. In case in-domain labeled data are scarce or mismatched with the true test distribution, direct fine-tuning can be brittle. In-Context Learning (ICL) provides a training-free, inference-time solution by adapting auditory LLMs through conditioning on a few in-domain demonstrations. In this work, we first show that $\textit{Vanilla ICL}$, improves zero-shot performance across diverse speech and audio tasks for selected models which suggest that this ICL adaptation capability can be generalized to multimodal setting. Building on this, we propose $\textbf{Meta Speech In-Context Learning (MetaSICL)}$, a post-training recipe utilizes only high resource speech data from various tasks intending to strengthen model's in-context learning capability. Experiments indicate our proposed method outperforms direct fine-tuning in low-resource scenario.

2026-05-27 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

検出から回復まで: 504 GPU による LLM 事前トレーニングでの運用分析

大規模な AI トレーニングは現在、基本的に分散システムの問題となっており、ハードウェア障害はまれな例外ではなく、日常的な動作条件となっています。しかし、本番トレーニングクラスターからの公的運用証拠は依然として不足しています。この技術レポートは、55 日間の Prometheus 時系列データと 224 のマルチノード トレーニング セッションをカバーする 73 日間の運用ログを使用した、63 ノードの NVIDIA B200 実稼働クラスター (504 GPU) の実証分析を示しています。このクラスターは、5 者 (SKT、Upstage、Lablup、NVIDIA Korea、VAST Data) が統合された監視パイプラインを共有する組織間環境内で動作します。この配置により、2 ~ 4 ノード規模では現れなかった 60 ノード規模のストレージ I/O ボトルネックを共同診断することが可能になりました。これは、単一チームだけでは分離できない実稼働規模の現象です。数か月にわたる事前トレーニング キャンペーンを利用して、3 つの定量分析を実行し、4 つの結果を得ました。まず、751 の Prometheus メトリクスと 10 件の XID で特定された GPU 障害に関する統計分析により、1 日あたりの誤検知が約 0.84 件で 10/10 の検出率 (XID 前は 2/10) を達成しました。複数の信号検出戦略を動機付ける、複数の障害タイプにわたって一貫して支配的な単一の指標はありません。次に、GPU VRAM から NFS パスに沿った 523 のチェックポイント イベントのプロファイリングにより、「帯域幅のパラドックス」(200 Gbps RoCE の 1.4 ~ 10.4% の使用率) が 128 スロットの NFS RPC レイヤーの飽和に起因していることがわかります。 3 番目に、マルチノード障害の応答では、集中した除外 (63 ノード中上位 3 ノードがすべての除外の 50% 以上を占める) と、12 チェーン (試行 73 回) にわたる自動再試行チェーンの成功率が 33.3% で、手動回復率 12.5% の 2.7 倍であることが示されています。再試行間隔の中央値は 11 分 (IQR 10-11) です。すべての分析は実稼働インフラストラクチャに基づいており、セッション レベルのワークロード管理、GPU 中心のスケジューリング、および統合された可観測性を提供します。

原文 (English)

From Detection to Recovery: Operational Analysis on LLM Pre-training with 504 GPUs

Large-scale AI training is now fundamentally a distributed systems problem, and hardware failures have become routine operating conditions rather than rare exceptions. Public operational evidence from production training clusters, however, remains scarce. This technical report presents an empirical analysis of a 63-node NVIDIA B200 production cluster (504 GPUs), using 55 days of Prometheus time-series data and 73 days of operational logs covering 224 multi-node training sessions. The cluster operates within a cross-organizational environment in which five parties (SKT, Upstage, Lablup, NVIDIA Korea, and VAST Data) share a unified monitoring pipeline. This arrangement enabled joint diagnosis of a 60-node-scale storage I/O bottleneck that did not appear at 2-4-node scale, a production-scale phenomenon no single team could isolate alone. Drawing on a months-long pre-training campaign, we perform three quantitative analyses yielding four findings. First, statistical analysis over 751 Prometheus metrics and 10 XID-identified GPU failures achieves a 10/10 detection rate (2/10 pre-XID) at ~0.84 false positives per day. No single metric is consistently dominant across failure types, motivating a multi-signal detection strategy. Second, profiling 523 checkpoint events along the GPU VRAM to NFS path attributes the "bandwidth paradox" (1.4-10.4% utilization of 200 Gbps RoCE) to saturation of the 128-slot NFS RPC layer. Third, multi-node failure response shows concentrated exclusions (top 3 of 63 nodes account for >50% of all exclusions) and an auto-retry chain success rate of 33.3% over 12 chains (73 attempts), 2.7x the 12.5% manual recovery rate; the median retry interval is 11 min (IQR 10-11). All analyses are grounded in production infrastructure providing session-level workload management, GPU-centric scheduling, and unified observability.

2026-05-26 19:20 JSTITmedia AI+ハードウェア/半導体

ファーウェイ、半導体で「1.4nm相当」目指す 31年までに 「ムーアの法則」に代わる新法則を提唱

中国Huaweiが半導体進化の新法則「τスケーリング法則」を提唱した。従来の微細化に代わり信号遅延を圧縮しトランジスタ密度を向上させる。秋のKirinチップに独自の回路技術LogicFoldingを初適用し、2031年に1.4nm相当の密度を目指すという。

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

どれだけ考えれば十分ですか? LLM 推論における冗長性の定量化と理解

推論可能な大規模な言語モデルは、長い思考連鎖を発し、レイテンシー、GPU 時間、およびエネルギーを大幅に消費して、難しい問題を解決します。その痕跡を何気なく検査すると、広範な再定式化、検証、循環的な内省が明らかになりますが、この検討が実際にどの程度必要であるかは、大規模に測定されたことも、第一原理から説明されたこともありません。この論文は両方のギャップを埋めます。私たちは推論モデル自体の観点から推論の冗長性を直接形式化します。正しいトレースの冗長性は、$\pi$ が思考を終了させて​​最終的な答えを出力することを強制されても正しい答えを生成する間に切り詰めることができる、後続のセグメント化されたステップの最大部分です。 4 つのフロンティア推論モデルと 2 つの数学的ベンチマークにわたる大規模な定量化により、ステップレベルの冗長性が一貫して高く、我々が調査した 8 つの (モデル、ベンチマーク) 条件全体で 61% から 93% の間であり、クリティカルプレフィックスの中央値は 8 つの条件のうち 6 つで単一のセグメント化されたステップに等しいことが示されています。この結果は、裁判官ファミリーの選択に対して堅牢であり、$\rho$ は MATH-500 の問題の難易度とともに減少しますが、4 つのモデルはすべて大幅に維持されていることがわかります。最も難しいレベル 5 の問題でも冗長 ($\rho \in [46\%, 85\%]$)。次に、この冗長性がモデル固有のアーティファクトではなく、長さに依存しない結果報酬の構造的な結果であることを証明します。そのような報酬の下では、最適な有限の予想停止時間は存在しません。結果は、RL アルゴリズム、基本モデル、データ分散、またはポリシーが RL または蒸留のどちらによって取得されたかに関係なく当てはまります。したがって、考えすぎは個々のモデルで修正すべきバグではなく、現在の推論モデルがどのようにトレーニングされるかの構造的特性です。コード: https://github.com/zhiyuanZhai20/how-much- Thinking-is-enough

原文 (English)

How Much Thinking is Enough? Quantifying and Understanding Redundancy in LLM Reasoning

Reasoning-capable large language models solve hard problems by emitting long chains of thought, paying heavily in latency, GPU time, and energy. Casual inspection of their traces reveals extensive reformulation, verification, and circular self-reflection, yet how much of this deliberation is actually necessary has never been measured at scale or explained from first principles. This paper closes both gaps. We formalise reasoning redundancy directly in terms of the reasoning model itself: the redundancy of a correct trace is the largest fraction of its trailing segmented steps that can be truncated while $\pi$, forced to terminate thinking and emit a final answer, still produces the correct answer. A large-scale quantification across four frontier reasoning models and two mathematical benchmarks shows that step-level redundancy is consistently high -- between 61% and 93% across the 8 (model, benchmark) conditions we study, with the median critical prefix equal to a single segmented step in six of the eight conditions -- that the finding is robust to the choice of judge family, and that although $\rho$ decreases with problem difficulty on MATH-500, all four models remain substantially redundant ($\rho \in [46\%, 85\%]$) even on the hardest Level-5 problems. We then prove that this redundancy is a structural consequence of length-agnostic outcome rewards, not a model-specific artefact: under any such reward, no finite expected stopping time is optimal. The result holds regardless of RL algorithm, base model, data distribution, or whether the policy is obtained via RL or distillation; over-thinking is therefore not a bug to be patched in individual models but a structural property of how current reasoning models are trained. Code: https://github.com/zhiyuanZhai20/how-much-thinking-is-enough

2026-05-26 13:00 JSTarXiv cs.AIエージェントロボティクスハードウェア/半導体

事前定義された学習オブジェクトを超えて: 最新の自律ロボット学習のための思考学習インタラクション モデル

オープンで変化する環境で動作する自律ロボットは、事前定義された入力、出力、およびアクション ルーチンに常に依存できるとは限りません。既存の学習方法では、環境との相互作用を通じてロボットのパフォーマンスを向上させることができますが、学習の対象は、入力特徴、認識出力、ネットワーク構造、タスクの目標、またはアクションシーケンスなど、事前に固定されていることがよくあります。これにより、長期的な運用中に新しい機能、新しいカテゴリ、またはより効率的なタスク ルーチンが出現したときに適応する能力が制限されます。この問題に対処するために、本論文では自律ロボットのための思考学習相互作用モデルを提案する。中心となる考え方は、潜在的な変化の特定、有用な証拠の選択、トレーニング資料の整理、検証アクションの計画によって思考が学習を導き、一方、学習はタスクの知識、機能選択の経験、アクション戦略、および将来の推論プロセスを更新することによって思考を促進するというものです。この双方向メカニズムに基づいて、ロボットは、環境との継続的な相互作用を通じて、事前に定義された学習設定を徐々に超えて、その認識関係と行動関係を適応させることができます。具体的には、提案されたモデルは、適応的な入力特徴の発見、出力カテゴリの拡張、学習モデルの更新、およびアクション ルーチンの再構築をサポートします。実験結果は、提案したモデルが特徴適応における最終認識精度を0.419から0.845に改善し、より高い新しいカテゴリ形成精度とモデル更新成功率を達成し、アクションルーチン再構築において平均アクション長を13.0から4.0に短縮することを示しています。学習によって強化された思考では、有用な証拠の選択率が 0.272 から 0.965 に増加し、学習結果が将来の証拠の選択と推論を効果的に改善できることを示しています。

原文 (English)

Beyond Predefined Learning Objects: A Thinking-Learning Interaction Model for Up-to-Date Autonomous Robot Learning

Autonomous robots operating in open and changing environments cannot always rely on predefined inputs, outputs, and action routines. Although existing learning methods enable robots to improve their performance through environmental interaction, the objects of learning are often fixed in advance, such as input features, recognition outputs, network structures, task goals, or action sequences. This limits their ability to adapt when new features, new categories, or more efficient task routines appear during long-term operation. To address this problem, this paper proposes a thinking-learning interaction model for autonomous robots. The core idea is that thinking guides learning by identifying potential changes, selecting useful evidence, organizing training materials, and planning verification actions, while learning promotes thinking by updating task knowledge, feature-selection experience, action strategies, and future reasoning processes. Based on this bidirectional mechanism, the robot can gradually move beyond predefined learning settings and adapt its recognition relations and action relations through continuous interaction with the environment. Specifically, the proposed model supports adaptive input feature discovery, output category expansion, learning model update, and action routine reconstruction. Experimental results show that the proposed model improves the final recognition accuracy from 0.419 to 0.845 in feature adaptation, achieves higher new-category formation accuracy and model-update success rate, and reduces the average action length from 13.0 to 4.0 in action routine reconstruction. In learning-enhanced thinking, the useful evidence selection rate increases from 0.272 to 0.965, indicating that learning results can effectively improve future evidence selection and reasoning.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

Agent-as-Peer-Debriefer: 定性分析のための視点ベースの改良を備えたマルチエージェント フレームワーク

大規模言語モデル (LLM) は定性データ分析 (QDA) に使用されることが増えていますが、その出力には人間による分析の深さやニュアンスが欠けていることがよくあります。私たちは、このギャップは、人間の QDA から信頼性を確保するための実践が欠けていることを反映していると主張します。ピア・デブリーフィングとは、アナリストが無関心なピアからフィードバックを求め、それを利用してコーディングを改良するものです。この実践を LLM 支援 QDA に導入するために、主要なコーディング手順にピア デブリーフィングを組み込むマルチエージェント QDA フレームワークである Agent-as-Peer-Debriefer を提案します。私たちのフレームワークでは、階層型コーディング エージェントが標準の QDA プロセスに従って、コード、サブテーマ、テーマ、および自己説明と反省メモを生成します。次に、これらの出力を 3 つのピア デブリーフィング エージェントと共有し、それぞれが異なる分析観点 (理論駆動、データ駆動、または応用) を適用し、コードの保持、名前変更、再割り当て、結合、または分割によってコードを洗練します。これらの視点は、ドメインとデータセット全体に一般化された確立された人間の QDA 実践から得られます。フレームワークを評価するために、3 つの LLM を使用して 2 つのドメインにわたる 3 つのデータセットでフレームワークをテストし、人間が注釈を付けたコードとの意味的な類似性を測定します。すべての設定において、パースペクティブベースのピアデブリーフィングの改良は、単一 LLM ベースラインよりも人間のコードとより密接に一致しており、アブレーションは、その利点が単に追加の改良によるものではないことをさらに示しています。また、3 つのパースペクティブは明確なトレードオフを生み出し、パースペクティブの選択が有意義で制御可能な設計上の決定であることを示しています。より広く言えば、これらの発見は、明確な視点を持ってピアデブリーフィングをシミュレートすることが、より信頼性の高い LLM 支援 QDA への有望な手段であることを示唆しています。

原文 (English)

Agent-as-Peer-Debriefer: A Multi-Agent Framework with Perspective-Based Refinement for Qualitative Analysis

Large language models (LLMs) are increasingly used for qualitative data analysis (QDA), yet their outputs often miss the depth and nuance of human analysis. We argue this gap reflects a missing credibility practice from human QDA: peer debriefing, in which an analyst seeks feedback from a disinterested peer and uses it to refine their coding. To bring this practice into LLM-assisted QDA, we propose Agent-as-Peer-Debriefer, a multi-agent QDA framework that builds peer debriefing into key coding steps. In our framework, a Hierarchical Coding Agent follows the standard QDA process to generate codes, sub-themes, and themes, along with self-explanations and reflection memos. It then shares these outputs with three Peer-Debriefing Agents, each applying a distinct analytical perspective (Theory-Driven, Data-Driven, or Applied) and refining the codes by keeping, renaming, reassigning, merging, or splitting them. These perspectives are drawn from established human QDA practices that generalize across domains and datasets. To evaluate the framework, we test it on three datasets across two domains with three LLMs, measuring semantic similarity to human-annotated codes. Across all settings, perspective-based, peer-debriefing refinement aligns more closely with human codes than a single-LLM baseline, and an ablation further shows the gain is not merely from additional refinement. The three perspectives also produce distinct trade-offs, showing that the choice of perspective is a meaningful and controllable design decision. More broadly, these findings suggest that simulating peer debriefing with explicit perspectives is a promising route to more credible LLM-assisted QDA.

2026-05-26 13:00 JSTarXiv cs.AIハードウェア/半導体

制御のない表現: 言語モデルでの実現効果のテスト

大規模な言語モデルが行動シミュレーターとして使用されることが増えていますが、その出力がプロンプトに敏感な表面パターンではなく、人間のような認知メカニズムをいつ反映するのかは依然として不明です。私たちはこの疑問を実現効果を通じて研究します。実現効果は、書類上の利益と損失の後ではリスクテイクが組織的に異なるという行動経済学のよく特徴付けられた発見です。私たちは LLM の動作を 3 つのレベルで評価します。プロンプトのみの動作感度、内部表現の線形読み出し、アクティベーション ステアリングによる因果制御です。プロンプトのみの結果は体系的な条件感度を示しますが、方向パターンは人間の実現効果の予測を再現しません。 Gemma の残差ストリームには、保留されたプロンプトに一般化される、レイヤー 18 で線形にデコード可能な実現ステータス信号が含まれています。ただし、この方向に沿って舵を取っても、下流のリスク選択が確実にシフトされるわけではなく、正のスケール全体および負の符号対称の実行で保持されるヌル結果になります。行動感度、潜在読み出し、および因果制御は、自動的には同時に発生しない 3 つの異なる特性であり、潜在読み出しの成功は、モデルが下流の意思決定中に表現に行動的に依存していることを示す不十分な証拠です。

原文 (English)

Representation Without Control: Testing the Realization Effect in Language Models

Large language models are increasingly used as behavioral simulators, but it remains unclear when their outputs reflect human-like cognitive mechanisms rather than prompt-sensitive surface patterns. We study this question through the realization effect, a well-characterized finding in behavioral economics in which risk-taking differs systematically after paper versus realized gains and losses. We evaluate LLM behavior at three levels: prompt-only behavioral sensitivity, linear readout of internal representations, and causal control via activation steering. Prompt-only results show systematic condition sensitivity, but the directional pattern does not reproduce human realization-effect predictions. Gemma's residual stream contains a linearly decodable realization-status signal at layer 18 that generalizes to held-out prompts. Steering along this direction does not, however, reliably shift downstream risk choices, a null result that holds across positive scales and in a negative sign-symmetry run. Behavioral sensitivity, latent readout, and causal control are three distinct properties that do not automatically co-occur, and successful latent readout is insufficient evidence that a model behaviorally relies on a representation during downstream decision-making.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体

DarkForest: マルチエージェント LLM の会話を減らし、精度を向上

マルチエージェント LLM システムは、複数のエージェントからの出力を組み合わせることで推論を改善しますが、対話が多い方法ではエラーの伝播と高い通信オーバーヘッドが発生する可能性があります。エージェントが生の応答や推論トレースを交換すると、間違った中間推論が採用され増幅され、自信はあるものの間違った合意が得られる可能性があります。マルチラウンド通信により、トークンの消費量、待ち時間、推論コストも増加します。この論文では、DarkForest という名前の制御された通信調整フレームワークを提案します。 DarkForest はまずエージェントを独立させて、各エージェントが他のエージェントの出力を見ることなく応答を生成します。次に、生の応答を構造化された候補レコードに解析し、意味的に同等の候補をクラスターにグループ化し、エージェントの信頼性、信頼度、解析品質、サポート パターンの信頼性、および独立性補正を使用して、これらのクラスターにわたる校正された信念分布を推定します。コーディネーターは、制御されたコミュニケーションにより、この信念状態からポリシーで許可された証拠のみを受け取ります。 6 つの推論ベンチマークに関する実験では、DarkForest が最高の全体的な品質を達成し、ベンチマーク メトリクスで最も強力なベースラインを最大 30.7\% 改善し、通信の多いベースラインと比較してトークン消費を最大 $6.5\times$ 削減することが示されています。

原文 (English)

DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMs

Multi-agent LLM systems improve reasoning by combining outputs from multiple agents, but interaction-heavy methods can introduce error propagation and high communication overhead. When agents exchange raw responses or reasoning traces, incorrect intermediate reasoning may be adopted and amplified, leading to confident but wrong consensus; multi-round communication also increases token consumption, latency, and inference cost. In this paper, we propose a controlled-communication coordination framework named DarkForest. DarkForest first keeps agents independent, so each agent produces an answer without seeing the others' outputs. It then parses the raw responses into structured candidate records, groups semantically equivalent candidates into clusters, and estimates a calibrated belief distribution over these clusters using agent reliability, confidence, parse quality, support-pattern reliability, and independence corrections. A coordinator receives only policy-permitted evidence from this belief state with controlled communication. Experiments on six reasoning benchmarks show that DarkForest achieves leading overall quality, improves the strongest baseline by up to 30.7\% on benchmark metrics, and reduces token consumption by up to $6.5\times$ compared with communication-heavy baselines.

2026-05-26 13:00 JSTarXiv cs.AIハードウェア/半導体

GIBLy: アーキテクチャに依存しない軽量の幾何学的誘導バイアス レイヤーによる 3D セマンティック セグメンテーションの改善

3D シーンの理解では、深層学習モデルは大規模なモデルと広範なトレーニングに依存して、3D データに存在する基本的な幾何学的構造をキャプチャします。しかし、既存の手法には、学習可能なプリミティブ形状などの幾何学情報を組み込むための明示的なメカニズムが欠如しており、多くの場合、大規模なモデルとより多くのトレーニング データが必要になるため、コストが増加し、一般化が制限される可能性があります。学習可能な幾何学的事前分布を 3D セグメンテーション パイプラインに統合する軽量の幾何学的誘導バイアス レイヤーである GIBLy を紹介します。 GIBLy は、MLP ベース、畳み込みベース、トランスフォーマー ベースのいずれであっても、最小限の計算オーバーヘッドでセグメンテーションのパフォーマンスを向上させる、単純な幾何学的形状に合わせた (したがって人間が解釈できる) 機能を提供することで、既存のアーキテクチャを強化します。複数の 3D セマンティック セグメンテーション ベンチマークにわたってアプローチを検証し、58K の追加パラメーターのみを追加しながら、PTV3 を使用した TS40K で最大 +11.5% mIoU を含む、一貫したパフォーマンスの向上を実証しました。私たちの結果は、軽量のアドオン レイヤーを使用して、正確かつ効率的な 3D シーンの理解をサポートする幾何構造を明示的にエンコードする利点を強調しています。

原文 (English)

GIBLy: Improving 3D Semantic Segmentation through an Architecture-Agnostic Lightweight Geometric Inductive Bias Layer

In 3D scene understanding, deep learning models rely on large models and extensive training to capture basic geometric structures that are present in the 3D data. However, existing methods lack explicit mechanisms to incorporate geometric information, such as learnable primitive shapes, often necessitating large models and more training data which in turn increases cost and can limit generalization. We introduce GIBLy, a lightweight geometric inductive bias layer that integrates learnable geometric priors into 3D segmentation pipelines. GIBLy enhances existing architectures -- whether MLP-based, convolution-based, or transformer-based -- by providing features aligned with simple geometric shapes (and thus human-interpretable) that improve segmentation performance with minimal computational overhead. We validate our approach across multiple 3D semantic segmentation benchmarks, demonstrating consistent performance gains, including up to +11.5% mIoU on TS40K with PTV3, while adding only 58K extra parameters. Our results highlight the benefit of explicitly encoding geometric structure to support accurate and efficient 3D scene understanding, with a lightweight add-on layer

2026-05-26 13:00 JSTarXiv cs.AIハードウェア/半導体

ScaleAcross Explorer: スケールアクロス AI モデル トレーニングのための通信最適化の探索

大規模な言語モデルのトレーニングを迅速に拡張するには、GPU リソースを複数のデータ センターの建物および地域に分散する必要があります。このようなパラダイムを「スケールアクロス」トレーニングと呼びます。インフラストラクチャが拡大するにつれて、システム設計空間はますます複雑になり、新しいモデル アーキテクチャ、ハードウェアの異種混合、進化する通信パターンが含まれます。 Meta の運用経験に基づいて、数十万の GPU を収容するいくつかのデータ センターにトレーニング ジョブを展開する際の複雑さを強調します。大規模な設計空間の探索を加速し、フロンティア モデル開発の効率的なトレーニングを可能にするために、並列処理の配置、並列処理のスケジューリング、およびネットワーク層テクノロジという 3 つの主要な設計次元の詳細な特性評価を実施します。次に、設計次元の相互作用を考慮し、スケールアクロス トレーニングを総合的に最適化するオプティマイザーである ScaleAcross Explorer を提案します。テストベッドの実験とシミュレーションでは、実稼働構成と比較して最大 64.62% のトレーニング速度が向上し、幅広い設計ポイントにわたって最先端のベースラインと比較して最大 37.59% のトレーニング速度が向上することが実証されています。

原文 (English)

ScaleAcross Explorer: Exploring Communication Optimization for Scale-Across AI Model Training

The rapid scaling of large language model training requires distributing GPU resources across multiple data center buildings and regions. We refer to such paradigm as "scale-across" training. As infrastructure expands, the system design space becomes increasingly intricate, encompassing new model architectures, hardware heterogeneity, and evolving communication patterns. Drawing from Meta's production experience, we highlight the complexities of deploying training jobs across a few data centers housing hundreds of thousands of GPUs. To accelerate exploration of the large design space and to enable efficient training for frontier model development, we conduct in-depth characterization of three key design dimensions: parallelism placement, parallelism scheduling, and network layer technologies. We then propose ScaleAcross Explorer, an optimizer that considers the interplay of design dimensions and holistically optimizes scale-across training. Testbed experiments and simulations demonstrate up to 64.62% training speedups over production configuration and up to 37.59% training speedups over the state-of-the-art baseline across a wide range of design points.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

CONF-KV: Long-Horizo​​n LLM の混合精度ストレージを使用した信頼性を意識した KV キャッシュの削除

ロングホライズン LLM 推論は、キーバリュー (KV) キャッシュを主要な GPU メモリ消費者に変え、トークンごとの注意のコストをますます高めます。一般的なエビクション ポリシーの多くは、静的なリーセンシー ウィンドウまたは履歴アテンションを使用し、すべてのデコード ステップで計算された信号、つまりモデルの現在の不確実性を未使用のままにします。 CONF-KV は、次のトークン分布をスカラー信頼スコアに変換し、それを使用してステップごとのキャッシュ バジェットを選択する KV キャッシュ マネージャーです。モデルが不確実な場合はより多くのコンテキストを保持し、自信がある場合は積極的にプルーニングします。各予算内で、トークンは蓄積された注目量と最新性の複合物によってランク付けされ、保護された最近のウィンドウによりローカルの一貫性が維持されます。このポリシーを、ブロック単位のオンライン ソフトマックス アテンション、混合 FP16/INT8 ストレージ、およびピラミッド型のレイヤーごとの予算バリアントと組み合わせます。 4 つのモデル ファミリと最大 4K までの生成長にわたって、CONF-KV は固定 512 トークン スライディング ウィンドウのフットプリント近くに留まり、フル KV の 1.5 ~ 2.1 パープレキシティ ポイント以内に留まります。最大 32,000 トークンの Needle-in-a-Haystack では、CONF-KV の取得精度は 91.4% に達しますが、スライディング ウィンドウでは 53.8%、H2O では 80.6% です。 75 の VisualWebArena タスクでは、2.8 分の 1 のピーク メモリでフル KV 成功の 95.3% を保持します。

原文 (English)

CONF-KV: Confidence-Aware KV Cache Eviction with Mixed-Precision Storage for Long-Horizon LLM

Long-horizon LLM inference turns the key--value (KV) cache into the dominant GPU memory consumer and makes per-token attention increasingly expensive. Many common eviction policies use static recency windows or historical attention, leaving unused a signal computed on every decoding step: the model's current uncertainty. We introduce CONF-KV, a KV-cache manager that converts the next-token distribution into a scalar confidence score and uses it to choose the per-step cache budget, retaining more context when the model is uncertain and pruning aggressively when it is confident. Within each budget, tokens are ranked by a composite of accumulated attention mass and recency, while a protected recent window preserves local coherence. We combine the policy with blockwise online-softmax attention, mixed FP16/INT8 storage, and a pyramidal per-layer budget variant. Across four model families and generated lengths up to 4K, CONF-KV stays near the footprint of a fixed 512-token sliding window while remaining within 1.5--2.1 perplexity points of full KV. On Needle-in-a-Haystack up to 32K tokens, CONF-KV reaches 91.4% retrieval accuracy versus 53.8% for sliding windows and 80.6% for H2O; on 75 VisualWebArena tasks it retains 95.3% of full-KV success at 2.8 times lower peak memory.

2026-05-26 13:00 JSTarXiv cs.AIハードウェア/半導体

トラスト領域反復ツイスト逐次モンテカルロによる拡散モデルの推論時間アライメント

私たちは、重みを更新せずに基本モデルを高い報酬の出力に向けて誘導することを目的として、拡散ベースの生成モデルの推論時間の調整を研究します。最近の逐次モンテカルロ (SMC) ベースのステアリング手法は、原則的な方法で報酬傾斜ターゲット分布を近似しますが、その提案は主にベース サンプラーに関連付けられたままです。報酬情報は主にパーティクルの再重み付けとリサンプリングによる伝播後に使用されるため、これらの方法では大きなパーティクル バジェットが必要となり、重みの縮退や高い分散推定が発生する可能性があります。分散を減らし、粒子効率を向上させる 1 つの方法は、ツイスト SMC のように、先読みガイダンスを提供するツイスト関数を繰り返し学習することです。ただし、既存の学習可能なツイスト手法は主に古典的な逐次推論用に開発されており、高次元の状態空間や末端、ノイズの多い、またはブラックボックスの報酬との拡散ベースの位置合わせに適用すると不安定になる可能性があります。我々は、SMC ベースの推論時間アライメントでツイスト関数を学習するための信頼領域フレームワークである Trust-Region Iterative Twisted Sequential Monte Carlo (TRI-TSMC) を提案します。各反復では、パス空間で正確な KL 制約付き更新を計算します。これにより、重要度の再重み付けを調整することによって閉形式の解が得られ、重み付けされた最尤法によってこのターゲットがパラメータ化されたツイスト ファミリに射影されます。理論的には、最適ツイスト関数の値関数解釈を形式化し、それが分散ゼロのサンプラーを生成することを示します。信頼領域の更新がターゲット分布に向かうエスコート パスをたどること、重み付き最尤更新が順方向 KL 投影であること、およびそのパスが残余重要度重み分散を低減することを証明します。経験的に、TRI-TSMC は、一致した推論時間予算の下で、離散拡散テキスト生成とテキストから画像への生成に関する主な位置合わせ目標を改善します。

原文 (English)

Inference-Time Alignment of Diffusion Models via Trust-Region Iterative Twisted Sequential Monte Carlo

We study inference-time alignment for diffusion-based generative models, aiming to steer a base model toward high-reward outputs without updating its weights. Recent Sequential Monte Carlo (SMC)-based steering methods approximate reward-tilted target distributions in a principled way, but their proposals remain largely tied to the base sampler. Since reward information is mainly used after propagation through particle reweighting and resampling, these methods can require large particle budgets and suffer from weight degeneracy and high-variance estimates. One way to reduce variance and improve particle efficiency is to iteratively learn twisting functions that provide look-ahead guidance, as in twisted SMC. However, existing learnable twisting methods are developed mainly for classical sequential inference and can be unstable when applied to diffusion-based alignment with high-dimensional state spaces and terminal, noisy, or black-box rewards. We propose Trust-Region Iterative Twisted Sequential Monte Carlo (TRI-TSMC), a trust-region framework for learning twisting functions in SMC-based inference-time alignment. Each iteration computes an exact KL-constrained update in path space, which admits a closed-form solution by tempered importance reweighting, and projects this target back to the parameterized twisted family by weighted maximum likelihood. Theoretically, we formalize the value-function interpretation of the optimal twisting function and show that it yields a zero-variance sampler. We prove that the trust-region update follows an escort path toward the target distribution, that the weighted maximum-likelihood update is a forward-KL projection, and that the path reduces residual importance-weight variance. Empirically, TRI-TSMC improves primary alignment objectives on discrete diffusion text generation and text-to-image generation under matched inference-time budgets.

2026-05-26 13:00 JSTarXiv cs.AIハードウェア/半導体

シミュレーションから実行まで: トレーニング後の言語モデルは、自身の世代を認識して反応します。

言語モデルは、独自の出力の結果をモデル化するインセンティブを持たない受動的予測子として事前トレーニングされます。トレーニング後のこれは変わります。独自の応答を生成するモデルは、ポリシーに従っていることを認識することで恩恵を受けることができます。我々は、トレーニング後のモデルがポリシー上の世代を認識し、この認識が出力分布に暗黙的にエンコードされるという証拠を提示します。特に、オンポリシーの出力分布エントロピーは、モデル ファミリおよびサイズ クラス全体で、オフポリシーのエントロピーより 3 ~ 4$\倍$ 低くなります。この効果の一部を入力の驚きの内部表現まで追跡し、モデルの事前の予測に従って最新の入力トークンの可能性を追跡し、出力エントロピーを因果的に変調します。これらの現象の一例は、自由形式のプロンプトへの応答で観察できます。トレーニング後のモデルは (事前トレーニングされたモデルとは異なり)、最初の出力トークンの前に、今後の応答のトピックに関する不確実性を解消します。別のトピックのプリフィルでこのキャッシュされた意図に違反すると、出力エントロピーが高くなります。また、明示的な口頭レポートを介して、モデルがポリシー上のコンテキストと事前入力を区別できるかどうかもテストしました。私たちは、それらが可能であることを発見しましたが、興味深いことに、この明示的な認識は暗黙的な認識とは異なるメカニズムを経由します。

原文 (English)

From Simulation to Enaction: Post-trained language models recognize and react to their own generations

Language models are pretrained as passive predictors with no incentive to model the consequences of their own outputs. Post-training changes this: a model producing its own responses can benefit from recognizing that it is on-policy. We present evidence that post-trained models recognize their on-policy generations, and this recognition is implicitly encoded in their output distributions. In particular, on-policy output distribution entropy is 3--4$\times$ lower than off-policy entropy, across model families and size classes. We trace part of this effect to an internal representation of input surprise, tracking the unlikeliness of the most recent input token according to the model's prior predictions, that causally modulates output entropy. One example of these phenomena can be observed in response to open-ended prompts; post-trained models (unlike pretrained models) collapse their uncertainty over the topic of their upcoming response before the first output token; violating this cached intention with a different-topic prefill results in higher output entropy. We also tested whether models can distinguish on-policy contexts from prefills via explicit verbal report. We find that they can, but that interestingly, this explicit recognition routes through a different mechanism than implicit recognition.

2026-05-26 13:00 JSTarXiv cs.AIハードウェア/半導体

Google Cloud TPU での Gemma 4 31B の微調整と提供: GPU ベースラインとの技術比較

TPU ハードウェア上で Google の Gemma 4 31B モデルを微調整して提供する最初のエンドツーエンドのデモンストレーションを紹介し、大規模な言語モデルの適応に関する TPU プラットフォームと GPU プラットフォームの実証的比較を提供します。トレーニングには Google TPU v5p-8、推論には TPU v6e-8 (Trillium) 上の LoRA を使用して、PyTorch、HuggingFace TRL、FSDP 上に構築された GPU ネイティブのトレーニング レシピを JAX + Tunix/Qwix スタックに移植するために必要なコードレベルの適応の完全なセットを文書化します。これらの適応は、メッシュ構成、LoRA モジュールの命名規則、シャーディング アノテーションの修正、勾配チェックポイント設定、データ パイプラインの再構築、およびカスタム Orbax からセーフテンソルへのチェックポイント マージ手順に及びます。推論のために、v6e-8 で Gemma 4 を提供するために必要な vLLM-TPU Docker セットアップを詳しく説明し、その結果生じるレイテンシとスループット プロファイルの特徴を説明します。同一のハイパーパラメータの下での 2xH100 GPU ベースラインと比較して、TPU トレーニングは 2.12 倍のコストで 1.61 倍の速度で完了します。推論スループットはプラットフォーム全体で 3% 以内ですが、TPU は最初のトークンまでの時間を 2 倍短縮します (235 ミリ秒対 475 ミリ秒)。これらを合計すると、TPU 構成は、代表的なトレインとサービスのワークロードでは 1.82 倍安くなります。私たちの取り組みは、オープン ツール エコシステムの重大なギャップを取り除き、TPU インフラストラクチャ上で Gemma 4 を導入するための再現可能で本番環境にすぐに使えるレシピを実務者に提供します。

原文 (English)

Fine-Tuning and Serving Gemma 4 31B on Google Cloud TPU: A Technical Comparison with GPU Baselines

We present the first end-to-end demonstration of fine-tuning and serving Google's Gemma 4 31B model on TPU hardware, providing an empirical comparison of TPU and GPU platforms for large language model adaptation. Using LoRA on a Google TPU v5p-8 for training and TPU v6e-8 (Trillium) for inference, we document the full set of code-level adaptations required to port a GPU-native training recipe, built on PyTorch, HuggingFace TRL, and FSDP, to the JAX + Tunix/Qwix stack. These adaptations span mesh configuration, LoRA module naming conventions, sharding annotation corrections, gradient checkpointing, data pipeline restructuring, and a custom Orbax-to-safetensors checkpoint merging procedure. For inference, we detail the vLLM-TPU Docker setup necessary to serve Gemma 4 on v6e-8 and characterize the resulting latency and throughput profile. Compared with a 2xH100 GPU baseline under identical hyperparameters, TPU training completes 1.61x faster at 2.12x lower cost. Inference throughput is within 3% across platforms, while TPU achieves 2x lower time-to-first-token (235 ms vs. 475 ms). Together, the TPU configuration is 1.82x cheaper for a representative train-plus-service workload. Our work removes a critical gap in the open tooling ecosystem and provides practitioners with a reproducible, production-ready recipe for Gemma 4 deployment on TPU infrastructure.

2026-05-26 13:00 JSTarXiv cs.AIハードウェア/半導体

SAMark: 段落レベルの言い換え堅牢性を備えた自己アンカー付きテキスト透かし

意味レベルの透かし (SWM) は、文を基本単位として扱うことで、テキストの変更に対する堅牢性を向上させます。ただし、このような攻撃は文の順序を変更することで透かし信号を全体的に破壊するため、段落レベルの言い換えに対する堅牢性は依然として困難です。この研究では、意味空間にステップに依存しない緑色の領域を確立することで文の順序への依存を取り除く、自己アンカー型透かしフレームワークである SAMark を提案します。検出可能性を向上させるために、弱く位置合わせされた候補からのノイズを抑制しながら透かし信号を増幅するマルチチャネル双曲線スコアリング メカニズムを導入します。さらに、ハード フィルタリングとソフト正則化を組み合わせた多様性を意識したフィルタリング戦略を提案し、単純な N グラム繰り返しフィルタを超えて意味上の冗長性に対処します。実験結果は、SAMark が典型的な段落レベルの言い換え攻撃の下で最大 90.2% の TP@FP1% を達成し、以前の最も強力なベースラインを平均 30% 以上上回るパフォーマンスを示しながら、透かしなしのテキストと競争力のある生成品質を維持し、従来の方法を制限していた堅牢性と品質のトレードオフを打破することを示しています。

原文 (English)

SAMark: A Self-Anchored Text Watermarking with Paragraph-Level Paraphrase Robustness

Semantic-level watermarking (SWM) improves robustness against text modifications by treating sentences as the basic unit. However, robustness to paragraph-level paraphrasing remains difficult because such attacks globally disrupt watermark signals by changing sentence order. In this work, we propose SAMark, a self-anchored watermarking framework that removes the dependency on sentence order by establishing a step-independent green region in semantic space. To improve detectability, we introduce a multi-channel hyperbolic scoring mechanism that amplifies watermark signals while suppressing noise from weakly aligned candidates. We further propose a diversity-aware filtering strategy that combines hard filtering with soft regularization, extending beyond simple n-gram repetition filters to address semantic redundancy. Experimental results show that SAMark achieves up to 90.2% TP@FP1% under typical paragraph-level paraphrasing attacks, outperforming the strongest prior baseline by more than 30% on average, while maintaining generation quality competitive with unwatermarked text and breaking the robustness-quality trade-off that limits prior methods.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

因果関係の結びつき: LLM は因果方向をエンコードできるが、その Yes/No 出力は表現できない

大規模な言語モデルが因果関係の質問についてエンコードするものと、その答えとの間に不一致があることがわかりました。反常識的な CLadder 項目では、固定線形プローブがモデルの隠れた状態から証拠に裏付けられた回答 (精度約 0.97) を復元しますが、口頭での「はい/いいえ」は常識的なもの (精度約 0.5) に戻ります。これを約 +0.5 ギャップの因果関係と呼びます。間違った Yes/No は 2 つの分離可能な故障モードに分解されます。内部信号がない場合と、口頭インターフェースでは言えない信号です。この含意は、出力のみの因果関係ベンチマークの両方をカットします。ベンチマークが「正しい」ということは、モデルが理解していることを意味する必要はなく、ベンチマークが「間違っている」ということは、モデルが理解できないことを意味する必要はありません。 LLM が単一の精度数値から導かれた因果推論ができるかどうかについての広範な主張は、もう一度検討する価値があります。

原文 (English)

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express

We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed linear probe recovers the evidence-supported answer from the model's hidden state (accuracy approximately 0.97), while the spoken Yes/No reverts to the commonsense one (accuracy approximately 0.5). We call this approximately +0.5 gap Causal Tongue-Tie: a wrong Yes/No decomposes into two separable failure modes: no internal signal versus a signal the verbal interface cannot say. The implication cuts both ways for output-only causal benchmarks: a benchmark "correct" need not mean the model has understood, and a benchmark "wrong" need not mean it cannot. Sweeping claims about whether LLMs can do causal reasoning, drawn from a single accuracy number, deserve a second look.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

AMEL: LLM の判断に対する蓄積されたメッセージの影響

大規模な言語モデルは、コードのレビュー、コンテンツの調整、または 1 つの会話で多くの項目が通過する出力のスコア付けなど、自動評価器として日常的に使用されます。以前の会話履歴の極性がその後の判断にバイアスをかけるかどうかを尋ねます。この効果を、LLM 判断に対する蓄積されたメッセージ効果 (AMEL) と呼びます。 4 つのプロバイダー (OpenAI、Anthropic、Google、および 4 つのオープンソース モデル) の 11 モデルに対する 75,898 回の API 呼び出しにわたって、同一のテスト項目を単独で提示するか、主に肯定的または否定的な評価で飽和した履歴を追跡します。モデルは会話の一般的な極性に向かってシフトします (d = -0.17、p < 10^-46)。この効果は、モデルがベースラインで本当に不確実である項目に集中します (高エントロピー項目の場合は d = -0.34、ベースラインが決定的である場合は d = -0.15)。バイアスはコンテキストの長さとともに増加しません。前のターン 5 と 50 では同じシフトが生成されます (Spearman |r| < 0.01; OLS 勾配 p = 0.80)。そして、否定性の非対称性があります。項目ごとにペアにすると、否定的な履歴は肯定的な履歴よりも 1.62 倍のバイアスを引き起こします (t = 13.46、p < 10^-39、n = 2,481)。スケーリングは役立ちますが、解決しません (Anthropic: Haiku -0.22 ~ Opus -0.17、OpenAI: Nano -0.34 ~ GPT-5.2 -0.17)。 3 回のフォローアップによりメカニズムが絞り込まれます。トークンの確率分布は、しきい値ではなく連続的に変化します。負の非対称性にはトークンレベルの要素と意味論的な要素の両方がありますが、バランスの原因を特定することはサンプルサイズでは探索的です。位置は重要ではありません。50 ターン履歴のどこかで 5 つの偏ったターンが同じシフトを生成します。評価パイプラインの最も簡単な修正は、項目ごとに新しいコンテキストを作成することです。バッチ処理が避けられない場合は、履歴のバランスを取ると役立ちます。

原文 (English)

AMEL: Accumulated Message Effects on LLM Judgments

Large language models are routinely used as automated evaluators: to review code, moderate content, or score outputs, often with many items passing through one conversation. We ask whether the polarity of prior conversation history biases subsequent judgments, an effect we call the accumulated message effect on LLM judgments (AMEL). Across 75,898 API calls to 11 models from 4 providers (OpenAI, Anthropic, Google, and four open-source models), we present identical test items in isolation or following histories saturated with predominantly positive or negative evaluations. Models shift toward the conversation's prevailing polarity (d = -0.17, p < 10^-46). The effect concentrates on items where the model is genuinely uncertain at baseline (d = -0.34 for high-entropy items, vs d = -0.15 when the baseline is deterministic). Bias does not grow with context length: 5 prior turns and 50 produce the same shift (Spearman |r| < 0.01; OLS slope p = 0.80). And there is a negativity asymmetry: paired per item, negative histories induce 1.62x more bias than positive (t = 13.46, p < 10^-39, n = 2,481). Scaling helps but does not solve it (Anthropic: Haiku -0.22 to Opus -0.17; OpenAI: Nano -0.34 to GPT-5.2 -0.17). Three follow-ups narrow the mechanism. The token probability distribution shifts continuously, not at a threshold. The negativity asymmetry has both token-level and semantic components, though attributing the balance is exploratory at our sample sizes. Position does not matter: five biased turns anywhere in a 50-turn history produce the same shift. The simplest fix for evaluation pipelines is a fresh context per item; when batching is unavoidable, balancing the history helps.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

HEAPr: 出力空間におけるヘッセ行列ベースの効率的なアトミック エキスパート プルーニング

大規模言語モデル (LLM) の専門家混合 (MoE) アーキテクチャは、高密度 LLM と比較して優れたパフォーマンスと推論コストの削減を実現します。ただし、パラメータ数が多いためメモリ要件が法外に高く、実際の展開が制限されます。既存の枝刈り手法は主に専門家レベルの枝刈りに重点を置いていますが、この粒度が粗いため、精度が大幅に低下することがよくあります。この研究では、エキスパートをより小さく分割不可能なアトミック エキスパートに分解する新しいプルーニング アルゴリズムである HEAPr を導入し、より正確で柔軟なアトミック エキスパート プルーニングを可能にします。各原子専門家の重要性を測定するために、最適脳外科医の理論と同様の原理に基づいた二次情報を活用します。二次情報によってもたらされる計算およびストレージの課題に対処するために、HEAPr はアトミック エキスパートの固有の特性を利用して、二次情報をエキスパート パラメーターからアトミック エキスパート パラメーターの情報に変換し、さらにそれをアトミック エキスパート出力の二次情報に単純化します。このアプローチにより、空間の複雑さが $O(d^4)$ ($d$ はモデルの次元数) から $O(d^2)$ に軽減されます。 HEAPr では、アトミック エキスパートの重要性を計算するために、小規模なキャリブレーション セットで 2 つの前方パスと 1 つの後方パスのみが必要です。 DeepSeek MoE や Qwen MoE ファミリを含む MoE モデルに関する広範な実験により、HEAPr が幅広い枝刈り率とベンチマークにわたって既存の専門家レベルの枝刈り手法を上回るパフォーマンスを発揮することが実証されました。具体的には、HEAPr は、ほとんどのモデルで 20% ~ 25% のプルーニング率でほぼ可逆圧縮を達成し、同時に FLOP を 20% 近く削減します。コードは [https://github.com/LLIKKE/HEAPr](https://github.com/LLIKKE/HEAPr) にあります。

原文 (English)

HEAPr: Hessian-based Efficient Atomic Expert Pruning in Output Space

Mixture-of-Experts (MoE) architectures in large language models (LLMs) deliver exceptional performance and reduced inference costs compared to dense LLMs. However, their large parameter counts result in prohibitive memory requirements, limiting practical deployment. While existing pruning methods primarily focus on expert-level pruning, this coarse granularity often leads to substantial accuracy degradation. In this work, we introduce HEAPr, a novel pruning algorithm that decomposes experts into smaller, indivisible atomic experts, enabling more precise and flexible atomic expert pruning. To measure the importance of each atomic expert, we leverage second-order information based on principles similar to the Optimal Brain Surgeon theory. To address the computational and storage challenges posed by second-order information, HEAPr exploits the inherent properties of atomic experts to transform the second-order information from expert parameters into that of atomic expert parameters, and further simplifies it to the second-order information of atomic expert outputs. This approach reduces the space complexity from $O(d^4)$, where $d$ is the model's dimensionality, to $O(d^2)$. HEAPr requires only two forward passes and one backward pass on a small calibration set to compute the importance of atomic experts. Extensive experiments on MoE models, including DeepSeek MoE and Qwen MoE family, demonstrate that HEAPr outperforms existing expert-level pruning methods across a wide range of pruning ratios and benchmarks. Specifically, HEAPr achieves nearly lossless compression at pruning ratios of 20% ~ 25% in most models, while also reducing FLOPs nearly by 20%. The code can be found at [https://github.com/LLIKKE/HEAPr](https://github.com/LLIKKE/HEAPr).

2026-05-26 13:00 JSTarXiv cs.AIハードウェア/半導体

VisualOverload: 非常に密集したシーンにおける VLM の視覚的な理解を調べる

基本的な視覚的理解は、最先端の VLM で本当に解決されているのでしょうか? VisualOverload は、非公開のグラウンドトゥルース応答を含む 2,720 の質問と回答のペアで構成される、少し異なるビジュアル質問応答 (VQA) ベンチマークです。通常、ほぼ全体的な画像の理解に焦点を当てていた以前の VQA データセットとは異なり、VisualOverload はモデルに、人口が密集した (または過負荷の) シーンで単純で知識を必要としない視覚タスクを実行するよう求めます。私たちのデータセットは、複数の人物、アクション、精巧に詳細な背景を背景に展開されるサブプロットが含まれるパブリック ドメインの絵画の高解像度スキャンで構成されています。私たちはこれらの画像に手動で 6 つのタスク カテゴリにわたる質問の注釈を付け、シーンを徹底的に理解できるように調査しました。現在のベンチマークは VLM のパフォーマンスを過大評価しており、詳細をエンコードして推論することは、特に人口密度の高いシーンに直面した場合、依然として困難な作業であるという仮説を立てています。実際、37 個のテスト済みモデルのうち最高のモデル (o3) でさえ、最も難しいテスト分割で 19.6% の精度しか達成できず、すべての質問で全体で 69.5% の精度しか達成できないことがわかります。徹底的な評価に加えて、カウントスキルの欠如、OCR の失敗、複雑なタスクにおける顕著な論理的矛盾など、複数の障害モードを明らかにするエラー分析でベンチマークを補完します。まとめると、VisualOverload は現在のビジョン モデルの重大なギャップを明らかにし、コミュニティがより良いモデルを開発するための重要なリソースを提供します。ベンチマーク: http://paulgavrikov.github.io/visualoverload

原文 (English)

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answering (VQA) benchmark comprising 2,720 question-answer pairs, with privately held ground-truth responses. Unlike prior VQA datasets that typically focus on near global image understanding, VisualOverload challenges models to perform simple, knowledge-free vision tasks in densely populated (or, overloaded) scenes. Our dataset consists of high-resolution scans of public-domain paintings that are populated with multiple figures, actions, and unfolding subplots set against elaborately detailed backdrops. We manually annotated these images with questions across six task categories to probe for a thorough understanding of the scene. We hypothesize that current benchmarks overestimate the performance of VLMs, and encoding and reasoning over details is still a challenging task for them, especially if they are confronted with densely populated scenes. Indeed, we observe that even the best model (o3) out of 37 tested models only achieves 19.6% accuracy on our hardest test split and overall 69.5% accuracy on all questions. Beyond a thorough evaluation, we complement our benchmark with an error analysis that reveals multiple failure modes, including a lack of counting skills, failure in OCR, and striking logical inconsistencies under complex tasks. Altogether, VisualOverload exposes a critical gap in current vision models and offers a crucial resource for the community to develop better models. Benchmark: http://paulgavrikov.github.io/visualoverload

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

SafeGPT: エンタープライズ LLM 使用におけるデータ漏洩と非倫理的な出力の防止

大規模言語モデル (LLM) は企業のワークフローを変革していますが、従業員が不注意で機密データを共有したり、ポリシーに違反するコンテンツを生成したりすると、セキュリティと倫理の問題が生じます。この文書では、機密データの漏洩と非倫理的な出力を防止する両面ガードレール システムである SafeGPT を提案します。 SafeGPT は、入力側の検出/秘匿化、出力側のモデレーション/リフレーミング、および人間参加型フィードバックを統合します。実験では、SafeGPT が満足度を維持しながら、データ漏洩のリスクと偏った出力を効果的に軽減することが実証されています。

原文 (English)

SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use

Large Language Models (LLMs) are transforming enterprise workflows but introduce security and ethics challenges when employees inadvertently share confidential data or generate policy-violating content. This paper proposes SafeGPT, a two-sided guardrail system preventing sensitive data leakage and unethical outputs. SafeGPT integrates input-side detection/redaction, output-side moderation/reframing, and human-in-the-loop feedback. Experiments demonstrate SafeGPT effectively reduces data leakage risk and biased outputs while maintaining satisfaction.

2026-05-26 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

Copy-as-Decode: LLM 編集のための文法に制約された並列プリフィル

LLM は、ほとんどのトークンが入力内でそのまま表示される場合でも、完全な出力を自己回帰的に再生成することでテキストとコードを編集します。私たちは、Copy-as-Decode という、編集生成を 2 つの基本的な文法による構造化されたデコードとして再キャストするデコード層メカニズムを研究します。入力行範囲を参照し、新しいコンテンツを出力します。トークンレベルのFSMは構文の妥当性を保証し、サービングレイヤープリミティブは、$N$の自己回帰ステップではなく単一の並列プリフィルフォワードを介してコピースパンごとにKVキャッシュを更新します。これにより、投機的デコードの並列フォワードカーネルを共有しますが、入力トークンをドラフトとして使用し、確率的検証に代わるプログラム強制承認が行われます。エンドツーエンドのトレーニングを必要としない上限分析を報告します。 (i) カーネルの高速化: Qwen2.5-{1.5B, 7B} では、並列プリフィルを介した $N$ トークンのコピーは、自己回帰より $6.8\times$--$303\times$ 高速です ($N \in [8, 512]$、A100 80GB bf16)。 (ii) コピー上限: ProbeEdit および HumanEvalPack-Fix (Py/JS) では、ゴールド トークンの $74$ ~ $98\%$ がラインレベルのプリミティブで到達可能です。各コーパスのスパン ヒストグラムに対して経験的カーネルを使用して構成すると、$29.0\times / 3.4\times / 4.2\times$ ($13.0\times$ プール) の閉じた形式の壁時計境界が得られます。トークンレベルの拡張は、$4.5\times$--$6.5\times$ の下限で $91$--$99\%$ のカバレッジに達します。 (iii) パイプラインの無損失性: Oracle プログラムは、すべての $482$ のケースで決定論的リゾルバーを介してラウンドトリップし、ダウンストリームの障害をメカニズムではなく選択範囲に局所的に特定します。摂動研究では、オフバイワンノイズの下でプールされた EM が $100\%$ から $15.48\%$ に低下することが示されています。 Qwen2.5-Coder-1.5B の微調整パイロットにより、HEvalFix-Py EM は $0/33$ (トレーニングされていない) から $12$--$17\%$ に引き上げられます。これは、運用セレクターではなく、学習可能性のシグナルです。バッチ処理の統合と複数ファイルの適用範囲はフォローアップの範囲内です。

原文 (English)

Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing

LLMs edit text and code by autoregressively regenerating the full output, even when most tokens appear verbatim in the input. We study Copy-as-Decode, a decoding-layer mechanism that recasts edit generation as structured decoding over a two-primitive grammar: references an input line range, ... emits new content. A token-level FSM guarantees syntactic validity, and a serving-layer primitive updates the KV cache for each copy span via a single parallel-prefill forward rather than $N$ autoregressive steps -- sharing the parallel-forward kernel of speculative decoding but with input tokens as the draft and program-enforced acceptance replacing probabilistic verification. We report an upper-bound analysis that requires no end-to-end training. (i) Kernel speedup: on Qwen2.5-{1.5B, 7B}, copying $N$ tokens via parallel prefill is $6.8\times$--$303\times$ faster than autoregressive ($N \in [8, 512]$, A100 80GB bf16). (ii) Copy ceiling: on ProbeEdit and HumanEvalPack-Fix (Py/JS), $74$--$98\%$ of gold tokens are reachable under the line-level primitive; composed with the empirical kernel over each corpus's span histogram this yields a closed-form wall-clock bound of $29.0\times / 3.4\times / 4.2\times$ ($13.0\times$ pooled). A token-level extension reaches $91$--$99\%$ coverage with $4.5\times$--$6.5\times$ floors. (iii) Pipeline losslessness: oracle programs round-trip through the deterministic resolver on all $482$ cases, localizing any downstream failure to span selection rather than the mechanism. A perturbation study shows pooled EM drops from $100\%$ to $15.48\%$ under off-by-one noise. A fine-tuning pilot on Qwen2.5-Coder-1.5B lifts HEvalFix-Py EM from $0/33$ (untrained) to $12$--$17\%$, a learnability signal, not a production selector. Batched-serving integration and multi-file coverage are scoped as follow-up.

2026-05-25 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体研究/論文

SciAtlas: 自動化された科学研究のための大規模ナレッジ グラフ

世界的な学術成果の急激な増加により、研究者やAIエージェントは前例のない「情報爆発」に直面しており、断片的で構造化されていない知識組織が深い学際的統合を妨げています。現在の学術検索ツールは主に、表面的なキーワード マッチングやベクトル空間の意味検索に依存しており、複雑な論理接続をナビゲートするために必要な位相推論機能が不足しています。エージェントのディープリサーチベースのフレームワークは、多くの場合、論理的な幻覚を引き起こし、高い推論コストを消費する傾向があります。このギャップを埋めるために、このレポートでは、パノラマ科学進化ネットワークとして設計された、大規模で学際的で異質な学術リソースの知識グラフである SciAtlas を紹介します。 SciAtlas は、26 の専門分野からの 4,300 万件を超える論文、合計 1 億 5,700 万のエンティティと 3B トリプレットを統合することにより、専門分野の障壁を取り除き、AI エージェントにグローバルな視点を提供する構造化トポロジカル認知基盤を提供します。さらに、トライパス協調想起とグラフ再ランキングを特徴とする神経記号検索アルゴリズムを開発し、単純な意味一致から決定論的関連発見へのシームレスな移行を実現します。また、文献レビュー、自動化された研究傾向の統合、アイデアの位置付け、学術的軌道の探索など、SciAtlas の主要な応用方向性を示し、SciAtlas が推論コストを大幅に削減しながら自動化された科学研究の全ループを強化する効果的な「認知マップ」として機能できることを実証します。 KG 取得とさまざまなダウンストリーム タスク用のインターフェイスを GitHub リポジトリでリリースしました。

原文 (English)

SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmented and unstructured knowledge organization impedes deep interdisciplinary integration. Current academic retrieval tools predominantly rely on superficial keyword matching or vector-space semantic retrieval, which lack the topological reasoning capabilities required to navigate complex logical connections. Agentic deep-research-based frameworks are often prone to logical hallucinations and consuming high inference costs. To bridge this gap, in this report, we introduce SciAtlas, a large-scale, multi-disciplinary, heterogeneous academic resource knowledge graph designed as a panoramic scientific evolution network. By integrating over 43M papers from 26 disciplines, and a total of 157M entities and 3B triplets, SciAtlas provides a structured topological cognitive substrate that dismantles disciplinary barriers and furnishes AI agents with a global perspective. Furthermore, we develop a neuro-symbolic retrieval algorithm featuring tri-path collaborative recall and graph reranking, achieving a seamless transition from simple semantic matching to deterministic association discovery. We also present key application directions of SciAtlas, including literature review, automated research trend synthesis, idea positioning, and academic trajectory exploration, to demonstrate that SciAtlas can serve as an effective ``cognitive map'' to empower the full loop of automated scientific research while significantly reducing reasoning costs. We have released the interfaces for KG retrieval and various downstream tasks in our GitHub repo.

2026-05-25 13:00 JSTarXiv cs.AIハードウェア/半導体

SPACENUM: VLM における空間数値的理解を再考する

視覚言語モデル (VLM) は、アクションの大きさや空間座標などの数値出力を生成する必要がある具体化された環境に導入されることが増えています。これらの数値は意味があるように見えますが、これらの数値出力が本当に空間認識に基づいているのかどうかは不明のままです。したがって、この研究では、空間探索中の動的遷移としての数値と、空間推論における静的レイアウトとしての数値という 2 つの相補的な設定をキャプチャする統合フレームワークである SpaceNum を通じて空間数値理解を再検討します。 Num2Space と Space2Num という 2 つの双方向タスクを定式化し、VLM が視覚側の空間構造と言語側の数値表現の間でどの程度適切にマッピングされるかを評価します。私たちは、現在の VLM が空間設定の数値を本当に理解しているかどうかを体系的に研究しています。動的遷移と静的レイアウトにわたって、モデルは空間的な意味での数値の根拠付けにほとんど失敗し、ランダムに近い推測を実行することが多いことがわかりました。エラー分析、推論トレース分析、および制御された介入を通じて、現在の VLM は浅い空間キューに大きく依存しており、安定した座標認識表現を構築するのに苦労しており、視覚的観察から構造化された空間レイアウトを抽象化できていないことを示します。さらに、明示的な推論ではわずかな利益しか得られない一方で、チューニングによって空間数値の理解を部分的に改善し、外部の空間推論ベンチマークに移行できることを示します。

原文 (English)

SPACENUM: Revisiting Spatial Numerical Understanding in VLMs

Vision-Language Models (VLMs) are increasingly deployed in embodied environments, where they need produce numerical outputs such as action magnitudes and spatial coordinates. Although these numbers appear meaningful, it remains unclear whether these numerical outputs are genuinely grounded in spatial perception. Therefore, in this work, we revisit spatial numerical understanding through SpaceNum, a unified framework that captures two complementary settings: numbers as dynamic transitions during spatial exploration, and numbers as static layouts in spatial reasoning. We formulate two bidirectional tasks, Num2Space and Space2Num, to evaluate how well VLMs map between vision-side spatial structure and language-side numerical representations. We systematically study whether current VLMs truly understand numerical values in spatial settings. Across dynamic transitions and static layouts, we find that models largely fail to ground numbers in spatial meaning and often perform close to random guess. Through error analysis, reasoning trace analysis, and controlled interventions, we show that current VLMs rely heavily on shallow spatial cues, struggle to build stable coordinate-aware representations, and fail to abstract structured spatial layouts from visual observations. We further show that explicit reasoning provides only marginal gains, while tuning can partially improve spatial numerical understanding and transfer to external spatial reasoning benchmarks.

2026-05-25 13:00 JSTarXiv cs.AIエージェントハードウェア/半導体

計算可能な公平性: AI リソース割り当てのためのボルツマン-ソフトマックス制御

大規模な AI システムでは、GPU の計算時間や帯域幅などの希少なリソースを複数のエージェントに割り当てることが重要な課題になります。従来のポリシーは効率の指標に焦点を当てており、システムの多様性と安定性を損なう支配集中につながる可能性があります。我々は、ボルツマン・ソフトマックス関数を選択ツールとしてではなく確率的リソース割り当てメカニズムとして再解釈し、逆温度パラメータ $\beta$ を効率と公平性のバランスを支配する計算可能な制御変数として再定義するフレームワークである Computable Fair Division (CFD) を提案します。静的分析により、政策ウェイト全体で損失総額がほ​​ぼ一定に保たれる、最適に近い安定回廊を持つパレートフロンティアが明らかになります。動的設定では、AHC++ (Adaptive Hard-Cap Controller++) が、観測された優位性とポリシーで指定されたターゲットとの間の誤差をフィードバックとして使用して $\beta$ をリアルタイムで更新します。シミュレーションにより、AHC++ は、スループットを大幅に低下させることなく公平性ターゲットを追跡しながら、外因性ショック下での極端な優勢集中を抑制することが示されています。スケーラビリティ分析により、エージェントを 100 倍に増やしても、実行時間は約 5.5 倍しか増加しないことが確認されています。コード: https://github.com/entrofy-ai/computable-fairness

原文 (English)

Computable Fairness: Boltzmann-Softmax Control for AI Resource Allocation

In large-scale AI systems, allocating scarce resources such as GPU compute time and bandwidth among multiple agents is a critical challenge. Conventional policies focus on efficiency metrics, potentially leading to dominance concentration that undermines system diversity and stability. We propose Computable Fair Division (CFD), a framework that reinterprets the Boltzmann-Softmax function not as a selection tool but as a probabilistic resource allocation mechanism, redefining the inverse temperature parameter $\beta$ as a computable control variable governing the efficiency-fairness balance. Static analysis reveals a Pareto frontier with a near-optimal Stability Corridor where total loss remains approximately constant across policy weights. In the dynamic setting, AHC++ (Adaptive Hard-Cap Controller++) updates $\beta$ in real time using the error between observed dominance and a policy-specified target as feedback. Simulations show that AHC++ suppresses extreme dominance concentration under exogenous shocks while tracking fairness targets without substantial throughput degradation. Scalability analysis confirms that a 100x increase in agents yields only approximately 5.5x increase in execution time. Code: https://github.com/entrofy-ai/computable-fairness

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

文化進化としてのモデル崩壊

モデルの崩壊、つまり独自の出力でトレーニングされた LLM の進行性の劣化は統計的に特徴付けられていますが、どの構造がどのような順序で、そしてなぜ劣化するのかについての言語的な説明が不足しています。私たちは、文化進化に基づく反復学習理論がこのギャップを埋めることを示します。私たちは 5 つの反証可能な予測を導き出し、理論を独自に識別する予測と確証的な予測を区別し、英語、ドイツ語、トルコ語で 10 世代にわたって LLaMA-2-7B とミストラル-7B を自己訓練することによってそれらをテストします。重要な識別的発見: フィルタリングされていない自己訓練下では、構成性は非単調な軌道 (最初は上昇し、その後下降) をたどります。この署名は、最大限規則的なシード データ (ノイズ除去を除外) で持続し、ランダム フィルターではなくタスクに基づいたフィルターによってのみ維持され、圧縮と通信のトレードオフに関する最初の LLM スケールの証拠を提供します。すべての予測は大きな効果量 (Hedges の $g > 1.6$; $\mathrm{BF}_{10} > 100$) で確認され、LLM 正則化勾配は人間の行動データ ($R^2 = 0.94$) とよく一致します。これらの結果は、モデルの崩壊を文化伝達現象として再構成し、自己学習パイプライン設計の具体的な原則を導き出します。

原文 (English)

Model Collapse as Cultural Evolution

Model collapse, the progressive degradation of LLMs trained on their own outputs, has been characterized statistically but lacks a linguistic explanation for which structures degrade, in what order, and why. We show that iterated learning theory from cultural evolution fills this gap. We derive five falsifiable predictions, distinguish those uniquely discriminative for the theory from confirmatory ones, and test them by self-training LLaMA-2-7B and Mistral-7B over 10 generations in English, German, and Turkish. The critical discriminative finding: compositionality follows a non-monotonic trajectory (initially rising, then falling) under unfiltered self-training. This signature persists with maximally regular seed data (ruling out noise removal) and is sustained only by task-grounded filtering, not random filtering, providing the first LLM-scale evidence for the compression-communication tradeoff. All predictions are confirmed with large effect sizes (Hedges' $g > 1.6$; $\mathrm{BF}_{10} > 100$), and LLM regularization gradients closely match human behavioral data ($R^2 = 0.94$). These results reframe model collapse as a cultural transmission phenomenon and yield concrete principles for self-training pipeline design.

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIエージェントハードウェア/半導体研究/論文

FastKernels: 本番環境での GPU カーネル生成のベンチマーク

GPU カーネル生成用の LLM ベースのエージェントは急速に進歩していますが、その進歩は最適化対象のベンチマークによって根本的に制約されています。既存のベンチマークは、運用推論フレームワークとの整合性が不十分です。合成入力を使用して単一の GPU でカーネルを評価し、周囲のコンパイル スタックを無視し、新しい最適化を発見するのではなく、既知の最適化を複製することに報酬を与えます。結果として得られる報酬シグナルは誤解を招くものです。エージェントは、サンドボックスでは高得点のカーネルを生成することを学習しますが、実際のシステムに統合すると、インターフェイスの非互換性、コンパイルスタックの競合、サイレント正確性の低下が発生します。 FastKernels は、8 カテゴリにまたがる 46 の代表的なアーキテクチャの最小限のセットを中心に構築されたカーネル ベンチマークであり、そのカーネルは、HuggingFace Transformers アーキテクチャの 96.2% (409/425) のカーネルを集合的に包含します。 FastKernels は、主流の LLM サービス上で vLLM や SGLang などの強化されたシステムと同等に動作し、十分にサービスが提供されていないアーキテクチャ上でのアップストリームのリファレンスを大幅に上回る、最小限の運用グレードの推論フレームワークとしても機能します。各タスクのインターフェイスは、そのアーキテクチャ ファミリの最先端のライブラリ内の対応するモジュールを反映しており、最適化されたカーネルを運用コードベースに直接デプロイすることができます。 FastKernels で最先端のカーネル エージェントを評価すると、最も強力なエージェントであっても実稼動ベースラインと比べて合計 0.94$\times$ の高速化しか達成できず、より弱いエージェントでは $0.78\times$ と $0.53\times$ であることがわかり、ベンチマークと実稼動の不一致がこの分野の重大なボトルネックであることが確認されました。私たちは、ベンチマークの向上が実稼働スループットの向上に直接つながるカーネル エージェントへの足がかりとして FastKernel をリリースします。コードは https://github.com/Snowflake-AI-Research/fastkernels で入手できます。

原文 (English)

FastKernels: Benchmarking GPU Kernel Generation in Production

LLM-based agents for GPU kernel generation are advancing rapidly, yet their progress is fundamentally constrained by the benchmarks they optimize against. Existing benchmarks are poorly aligned with production inference frameworks: they evaluate kernels on a single GPU with synthetic inputs, ignore the surrounding compilation stack, and reward replicating known optimizations rather than discovering new ones. The resulting reward signals are misleading: agents learn to generate kernels that score well in sandboxes but introduce interface incompatibilities, compilation-stack conflicts, and silent correctness degradation when integrated into real systems. We introduce FastKernels, a kernel benchmark built around a minimal set of 46 representative architectures spanning 8 categories, whose kernels collectively subsume those of 96.2% (409/425) of HuggingFace Transformers architectures. FastKernels doubles as a minimalistic, production-grade inference framework that runs at parity with hardened systems such as vLLM and SGLang on mainstream LLM serving and substantially exceeds upstream references on under-served architectures; each task's interface mirrors the corresponding module in the state-of-the-art library for its architecture family, enabling direct deployment of optimized kernels into production codebases. Evaluating state-of-the-art kernel agents on FastKernels, we find that even the strongest agent achieves only 0.94$\times$ aggregate speedup over production baselines, with weaker agents at $0.78\times$ and $0.53\times$ -- confirming that benchmark-production misalignment is a critical bottleneck for the field. We release FastKernels as a stepping stone toward kernel agents whose benchmark gains translate directly into production throughput improvements. Code is available at https://github.com/Snowflake-AI-Research/fastkernels

2026-05-25 13:00 JSTarXiv cs.AILLM/生成AIハードウェア/半導体

ReCoVer: フォールトトレラントな集合的で汎用性の高いワークロードを介した回復力のある LLM 事前トレーニング システム

大規模な GPU クラスターで大規模な言語モデルを事前トレーニングすることにより、ハードウェア障害が稀ではなく日常的に発生するようになり、回復力のあるトレーニング システムの必要性が高まっています。しかし、既存のフレームワークは、特定の並列処理スキームに焦点を当てているか、失敗のないトレーニング軌道から逸脱する危険性があります。私たちは、単一の不変条件を維持する回復力のある LLM 事前トレーニング システムである ReCoVer を提案します。つまり、各反復でマイクロバッチの数を一定に保ち、反復ごとの勾配が失敗のない実行と確率的に等価であることを保証します。このフレームワークは、3 つの分離されたプロトコル層として構成されています。(1) 障害がレプリカ間で伝播するのを隔離するフォールトトレラント集合体。 (2) 反復内の進行状況を維持し、勾配の破損を防ぐ、段階的なきめ細かいリカバリ。 (3) マイクロバッチ クォータを生存者全体に動的に再配分する多用途ワークロード ポリシー。この設計は並列処理に依存せず、3D 並列処理とドロップイン サブストレートとしてハイブリッド シャード データ パラレル (HSDP) の両方を直接統合します。最大 512 GPU のエンドツーエンドの事前トレーニング タスクで実装を評価しました。ReCoVer は、実行全体で 256 GPU が失われたにもかかわらず、障害のないリファレンスからトレーニング軌跡を正常に保存しました。チェックポイントと再起動のベースラインと比較すると、ReCoVer は、連続した障害の後、実効スループットが 2.23 倍高いことを示しています。この利点により、ReCoVer は 234 GPU 時間で 74.9% 多くのトークンを処理することになり、トレーニングが長引くにつれてその差は拡大します。

原文 (English)

ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload

Pre-training large language models on massive GPU clusters has made hardware faults routine rather than rare, driving the need for resilient training systems. Yet existing frameworks either focus on specific parallelism schemes or risk drifting away from a failure-free training trajectory. We propose ReCoVer, a resilient LLM pre-training system that upholds a single invariant: each iteration keeps the number of microbatches constant, ensuring per-iteration gradients remain stochastically equivalent to a failure-free run. The framework is organized as three decoupled protocol layers: (1) Fault-tolerant collectives that isolate faults from propagating across replicas; (2) in-step fine-grained recovery that preserves intra-iteration progress and prevents gradient corruption; (3) versatile-workload policy that dynamically redistributes microbatch quotas across the survivors. The design is parallelism-agnostic, integrating directly with both 3D parallelism and Hybrid Sharded Data Parallel (HSDP) as a drop-in substrate. We evaluate our implementation on end-to-end pre-training tasks for up to 512 GPUs, ReCoVer successfully preserves the training trajectory from a failure-free reference despite of 256 GPUs lost spread across the run. For comparison with checkpoint-and-restart baselines, ReCoVer demonstrates $2.23\times$ higher effective throughput after successive failures. This advantage results in ReCoVer processing 74.9% more tokens at 234 GPU-hours, with the gap widening as the training prolongs.