Source ledger

90 sources·FOLLOW 18 · ADAPT 28 · REFERENCE 34 · DEPRECATED 10

Every citation token like [EXT-STATS-001] anywhere on this site links to its entry here. Rendered from references/sources.yaml, the record of record.

Adopt this recipe as written; validate only your project-specific parameters.

Metrics, ISL/OSL sweeps, operating-point selection, sizing, $/M-token TCO; AIPerf-based

official-docsmaturity: maturestrength: strong-evidencefreshness: activeverified 2026-08-21published 2026-07-20cited in 00, 04, 06, 11

Verified live (W1): AIPerf-based; TTFT/ITL/TPS definitions, concurrency sweeps, ISL/OSL pairs, operating-point selection all present verbatim. 'latest' resolves to guide v2.0.0. 2025 blog series' GenAI-Perf commands deprecated, unbannered.

McNemar paired testing, power/MDE tables, GO/NO-GO gate w/ 95% CIs

official-docsmaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-21published 2026-06-03cited in 03, 04, 14

Verified live (W1): compare (McNemar exact on discordant pairs, power table 10~28%/1000~2.8%, MDE reported, INCONCLUSIVE when underpowered, identical-prompt-template requirement) and quality-gate (tiers, 95% CI on paired delta, INSUFFICIENT_EVIDENCE <10 items, GO/NO-GO/INCONCLUSIVE exit codes) confirmed verbatim. SDK v0.3.0 (2026-06-03) still latest. CITATION PATH: /latest/ 308-redirects to /nightly/ — no pinned doc tree.

Source method nel compare productized; usable standalone on lm-eval output

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21cited in 03, 04

Verified (W1): stable citation is arXiv:2602.10144 (ICLR 2026). Code repo amazon-science/LLM-Accuracy-Stats archived read-only 2026-05-08 — usable, unmaintained.

EXT-STATS-001Miller 2024 'Adding Error Bars to Evals' + McNemar exact power (NCSS PASS ch.150) (arXiv:2411.00640)Miller; NCSS

clustered/paired SEs, power/MDE formulas; largest pure gap in the source project practice

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2024-11cited in 00, 04

arXiv 2411.00640 verified resolving, title 'Adding Error Bars to Evals' (W1). Drives clustered/paired SE and power/MDE machinery in ch.04.

objective→dataset→metrics→run→continuous; validate model-graders against humans

official-docsmaturity: currentstrength: consensusfreshness: activeverified 2026-08-21cited in 00, 01, 03

Verified live (W1) WITH PLATFORM SUNSET: the hosted OpenAI Evals platform 'will become read-only for existing users on October 31, 2026' and 'is scheduled to shut down on November 30, 2026'. The methodology content (objective→dataset→metrics→run→iterate; validate model graders against humans) remains citable; do NOT present the platform as current tooling.

EXT-EVAL-004EvalGen / Who Validates the Validators (criteria drift) (arXiv:2404.12272)Shankar et al.

criteria and ground truth co-evolve; most load-bearing validation of the source project's design

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2024cited in 00, 03

n=9 practitioner study; validates frozen-suite+versioned-release discipline

EXT-EVAL-006GSM1k contamination study (arXiv:2405.00332)Zhang et al.

up to 8pp accuracy drop on fresh benchmark; contamination r²=0.36

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2024-05cited in 00, 03

validates the source project disjoint seed ranges while corpus stays private

EXT-JUDGE-002JudgeBench: LLM judges on objectively verifiable tasks (arXiv:2410.12784)Tan et al.

GPT-4o-judge ~56.6% near chance; vindicates the source project's deterministic grading

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2024-10cited in 00, 03

numbers verified from paper's tables, not summarizer

EXT-FT-001RAG-vs-fine-tuning comparison trio (EMNLP 2024, 2403.01432, 2401.08406) (arXiv:2403.01432)Ovadia; Soudani; Balaguer

three studies agree RAG/prompt-first for factual/citation tasks

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-20published 2024cited in 00, 07, 09

validates the source project's evidence-gated ladder ordering

EXT-FT-003QLoRA quantized fine-tuning (arXiv:2305.14314)Dettmers et al.

VRAM floors: 9B~6.5GB, 14B~8.5GB, 27B~22GB, 70B~41GB

papermaturity: maturestrength: strong-evidencefreshness: stableverified 2026-08-21published 2023cited in 09, 11

cross-checked against Unsloth docs

EXT-FT-005LoRA quality package (LoRA Learns Less/Forgets Less + LoRA Without Regret) (arXiv:2405.09673)Biderman; Thinking Machines

all-layer LoRA at post-training scale ~= full FT at ~67% FLOPs

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2024; 2025-09-29cited in 09

Thinking Machines paper reconciles with Biderman's single-model scope caveat

28-test rubric, score=minimum across categories; adopt as annual self-audit

papermaturity: maturestrength: consensusfreshness: stableverified 2026-08-21published 2017; 2015cited in 01, 10, 12, 13

Verified from the PDF itself (W1): 28 tests; final score = MINIMUM across the 4 sections; Monitor-3 = training/serving compute the same values; Monitor-7 = prediction-quality regression via held-out production examples. NOTE: a first-pass WebFetch summarization hallucinated wrong authors/venue — verified against page images; cite Breck, Cai, Nielsen, Salib, Sculley, IEEE Big Data 2017.

correctness-only sanity gates before scaling up

third-partymaturity: maturestrength: consensusfreshness: stableverified 2026-08-20published 2019cited in 00, 03, 06, 09

the source project's make reachability/gold/determinism probes formalize this

'Use the simplest model that meets your technical and business objectives'; run ONE canary at a time; canary metrics must be causally attributable to the change

official-docsmaturity: maturestrength: consensusfreshness: stableverified 2026-08-21published 2018cited in 10, 12

Verified live (W1), phrasing near-verbatim.

Endpoint-agnostic benchmarking vs any OpenAI-compatible endpoint; v0.12.0 adds agent workloads, Anthropic Messages endpoint, accuracy benchmarking; drop-in GenAI-Perf replacement

official-repomaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-21published v0.12.0 2026-08-06cited in 06

Verified active (W1). Pin one tool + one metric-definition set across tiers.

ADAPT Sound core, but adapt it — parts are project-specific, stale, or need validation.

Weak/strong escalation router; 4 mechanisms; confirmations=2, fail-open, latch

official-repomaturity: pre-alphastrength: strong-evidencefreshness: volatileverified 2026-08-21published 2026-08-11tool v0.2.0 (2026-08-10)cited in 08, 14

Verified live (W1): still pre-alpha 'Not for production use'; escalation semantics confirmed (weak-first, judge on actual output, confirmations=2, per-session latch, fail-open, judge failure HOLDS streak). PRECISION: capable_target/efficient_target two-tier schema belongs to the stage_router route type. v0.2.0 was a substantial redesign (native Rust server) — do not cite v0.1.0-era CLI. Prefill router still unshipped (zero 'prefill' paths on main).

Break-even rule: min offload=judge_cost/(strong-weak); 74% cost cut, 93% accuracy retained

third-partymaturity: currentstrength: strong-evidencefreshness: snapshotverified 2026-08-21published 2026-08-11cited in 08, 11

Verified live (W1): break-even rule and all figures verbatim (74% cost cut, 93% accuracy retention 80.0/86.0, escalation mean 6.9% over 5 runs, range 4.1–9.1%). Scope: 145-task agent suite, one weak/strong pairing — numbers are conditioned on that setup.

Evaluate-first ordering; verifiability/resources/maturity framework; TSR over accuracy

official-blogmaturity: currentstrength: consensusfreshness: stableverified 2026-08-21published 2026-05-20cited in 00, 01, 03, 07, 09, 14

Verified live (W1): 'The most successful teams start with lightweight methods, invest early in evaluation, and layer in training-based techniques where measurement shows they're needed.' Companion evaluation post (2026-05-19): developer.nvidia.com/blog/mastering-agentic-techniques-ai-agent-evaluation/ — TSR over accuracy, instrument from day one.

Demonstrated >99% median-recovery PTQ bar; 95-99% deliberate QAD target; worked example

official-blogmaturity: currentstrength: case-studyfreshness: snapshotverified 2026-08-21published 2026-08-17cited in 03, 07

Verified live (W1): '>99% median accuracy recovery is targeted when performing PTQ only'; '95–99% ... when combined with QAD'; PTQ 96.33% → QAD 99.72% worked example — all verbatim. Demonstrated practice, not documented doctrine.

On-device 35B-MoE to NVFP4 in ~2h; W4A16 vs W4A4 recipe rule

official-repomaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-21published 2026-07-22cited in 02, 07

Verified live (W1, file updated 2026-07-27): W4A16 interactive/low-concurrency vs W4A4 high-concurrency recipe rule verbatim; 'we recommend running evaluations' with no procedure/threshold. Format is Blackwell-only — quantize for the deployment target.

Publish the complete evaluation recipe alongside results; methodological consistency with clear provenance rather than bit-wise identical outputs

official-blogmaturity: maturestrength: strong-evidencefreshness: stableverified 2026-08-21published 2025-12-17cited in 02, 04, 13

Verified live (W1) WITH CORRECTION: two claims previously attributed here (pin containers/sampling-params/judges; smoke-test discipline) are NOT on the live page — do not cite this source for them. Confirmed: full-recipe publication; 'the purpose of open evaluation is not to force bit-wise identical outputs, but to deliver methodological consistency with clear provenance'. A --dry-run flag is mentioned in passing only.

Prometheus /v1/metrics (vLLM passthrough), OTel traces, JSONL logs = production minimum

official-docsmaturity: maturestrength: heuristicfreshness: activeverified 2026-08-21published 2026-08-19cited in 12

Verified live (W1, page updated 2026-08-19): Prometheus /v1/metrics with unmodified vLLM-native metric passthrough, OTel spans via OTEL_EXPORTER_OTLP_TRACES_ENDPOINT, JSONL logs via NIM_JSONL_LOGGING. Classified ADAPT: production mechanics documented and adoptable; the experimentation telemetry minimum is missing and is supplied by chapter 12.

NV-FINETUNESTACK-001ADAPTNeMo AutoModel + NeMo RL + NeMo Customizer (fine-tuning stack)NVIDIA

LoRA-vs-SFT selection guidance (LoRA for 1–2 GPUs / fast iteration / multiple specialized versions); GPU sizing guidance

official-docsmaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-21cited in 07, 09

Verified (W1) WITH CORRECTION: Customizer docs carry LoRA-vs-SFT guidance, but the early-stopping defaults (val_loss, patience 10, min-delta 0.001) are NOT in the Customizer docs — they exist in NeMo's OSS training-framework code. Cite as code-level defaults, not published vendor guidance. K8s footprint.

Only source for LoRA-vs-QLoRA criteria; NVIDIA delegates methodology here

partnermaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-21cited in 07, 09, 11

Verified live (W1): VRAM floors exact match (9B 6.5GB QLoRA/24GB LoRA; 27B 22GB/64GB; 70B 41GB/164GB — stated as absolute minimums); all-major-linear-layers LoRA targeting, rank 8–128 guidance, alpha=2r. CHANGE: multi-GPU now 'works but a much better version is coming' (no longer just 'coming soon'). Unsloth publishes no dates — staleness undatable from the page.

Dataset thresholds: 100-1,000 pairs PEFT / 1,000+ full FT

official-blogmaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2025-12-15cited in 09, 14

Verified live (W1) WITH CORRECTIONS: correct host is blogs.nvidia.com (not developer.nvidia.com/blog); published 2025-12-15. Thresholds confirmed exact: PEFT 100–1,000 prompt-sample pairs; full fine-tuning 1,000+.

NV-TOOLCALLTUTORIAL-001ADAPTNeMo tool-calling evaluation tutorialNVIDIA

85/15 split + independent xlam golden set; 500+ synthetic samples to ~93%

official-docsmaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-20cited in 03, 09
NV-GARAK-001ADAPTgarak (LLM security probing tool)NVIDIA

Z-score-vs-calibration-bag pattern for security/regression baselines

official-repomaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-20cited in 03, 13

Pattern worth copying for internal regression baselines

NV-ATIF-001ADAPTATIF (agent trajectory interchange format)NVIDIA / Harbor community

Trajectory interchange format across NAT/Relay/Platform

official-repomaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-20cited in 12, 13

Adopt if the source project stores trajectories beyond its event-bus records

EXT-STOPPING-002ADAPTClinical-trial adaptive design & DSMB pre-specification (Lan-DeMets spending, PMC3248853) (PMC3248853)Lan & DeMets; PMC/NIH

pre-register adaptation rule, DSMB executes it not investigator judgment

papermaturity: maturestrength: consensusfreshness: stableverified 2026-08-20cited in 00, 03, 04

CP~0.20 convention & Lan-DeMets spending explicitly dropped, unverifiable

EXT-STOPPING-003ADAPTRacing and successive-halving screening (Hoeffding/Bernstein racing, Hyperband, ASHA)Loh & Nowozin; Jamieson & Talwalkar; Li et al.

ASHA-style rungs adopted for TRAIN screening, must be class-balanced

papermaturity: maturestrength: strong-evidencefreshness: stableverified 2026-08-20published 2013-2020cited in 04, 05, 07

corrected: rungs must be class-balanced, Bernstein racing at class level

multidimensional explicit criteria, grader by task shape, edge-case checklists

official-docsmaturity: currentstrength: consensusfreshness: activeverified 2026-08-21published fetched 2026-08-20cited in 00, 01, 03

Verified live (W1): multidimensional explicit criteria, grader-by-task-shape, edge-case checklists; volume heuristic ('prefer higher volume with slightly lower signal') present. The volume-vs-curation tension stays visible in ch.03.

read traces, open-ended failure notes, synthesize taxonomy

third-partymaturity: currentstrength: heuristicfreshness: activeverified 2026-08-21published 2025-03-24cited in 03

Verified live (W1): the error-analysis-first / bottom-up taxonomy process is the field-guide post (2025-03-24) — cite it, not the older evals post.

EXT-EVAL-007ADAPTBIG-bench canary strings convention (arXiv:2206.04615)BIG-bench authors

cooperative-scraper exclusion convention for published task text

papermaturity: maturestrength: heuristicfreshness: stableverified 2026-08-21published 2022cited in 03, 13

new action: add canary GUID to TEST files and published excerpts

EXT-AGENT-001ADAPTpass@k estimator (Chen 2021) + τ-bench pass^k reliability (Yao 2406.12045) (arXiv:2406.12045)Chen et al.; Yao et al.

gpt-4o agent pass^1>60% collapses <25% at pass^8, identical tasks

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2021; 2024-06cited in 03, 04

new action: measure frontier pass^k (k=3-5) on TEST subset

EXT-DETERM-001ADAPTNondeterminism/batch-invariance package (Defeating Nondeterminism, arXiv:2506.09501, vLLM & SGLang docs, llama.cpp #4130) (arXiv:2506.09501)Thinking Machines; SGLang; vLLM; llama.cpp

reduction-order/GPU-state variance dominant; batch-invariance costs 34-110% throughput

papermaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-21published 2025-09cited in 00, 02, 03, 05, 06, 10

validates the source project's single-slot same-session reproducibility rule

EXT-FT-006ADAPTCatastrophic forgetting in fine-tuning (2308.08747, 2510.17776) (arXiv:2308.08747)Luo et al.

forgetting real, scale-trends non-obvious, merging doesn't reliably fix it

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2023; 2025cited in 09

mandates full TEST suite as FT gate, never just target-behavior slice

EXT-PERF-001ADAPTvLLM benchmarks + DistServe goodput conceptvLLM project; DistServe (OSDI'24)

inference benchmarking method + goodput concept for serving

papermaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-20published 2024cited in 06

DistServe headline figures blocked PDF, not independently verified

smoke/load/stress/spike/soak test types; thresholds-as-SLOs

official-docsmaturity: maturestrength: consensusfreshness: stableverified 2026-08-21cited in 06, 10

Verified live (W1): smoke / average-load / stress / spike / soak / breakpoint all present with definitions.

Mirrors a copy of live inference traffic to a shadow variant in real time; comparison/interpretation left to the user

official-docsmaturity: maturestrength: consensusfreshness: stableverified 2026-08-21cited in 10

Verified live (W1).

Staged manual traffic shifting (small % → all) plus mirrored/shadow traffic (≤50%, one deployment, documented limits)

official-docsmaturity: maturestrength: consensusfreshness: activeverified 2026-08-21published 2026-02-08 (rev. 2026-04-24)cited in 10

Verified live (W1).

'QAD is Model Optimizer's recommended strategy for accuracy recovery after quantization' (verbatim); references NVFP4 (which the choosing-quant-methods decision page omits)

official-repomaturity: currentstrength: strong-evidencefreshness: activeverified 2026-08-21cited in 07

Added during W1: the QAD-recommendation line lives here, not on the decision page — cite this for the recovery-strategy recommendation.

De-facto standard open serving engine; deterministic batch-invariant mode is OPT-IN (VLLM_BATCH_INVARIANT=1, CC>=8.0, reduced performance)

official-repomaturity: maturestrength: strong-evidencefreshness: activeverified 2026-08-21cited in 02, 05, 06

Verified active (W1). Throughput-regime default candidate in ch.05 regime table.

REFERENCE Background evidence and context; not an executable recipe.

Offline/online benchmark procedure, 4 engines, ttft/tpot/itl/e2el, concurrency=1

official-repomaturity: currentstrength: heuristicfreshness: snapshotverified 2026-08-20cited in 06

Buried asset in the connect-two-sparks playbook; no thresholds; pins stale containers. Not re-verified in authoring pass (REFERENCE stance).

Integrated evaluate/secure/tune/build; Switchyard under Tune; Studio alpha

official-repomaturity: newstrength: heuristicfreshness: volatileverified 2026-08-21published created 2026-05-14cited in 02, 05, 14

Verified live (W1): active; pillars Secure/Evaluate/Tune/Build + synthetic data; Switchyard under Tune; driven from coding agents. Experiment-registry/run-history capability STILL ABSENT — the watch item stands.

$/M-token formula chain, Pareto precision comparison, datacenter-scale assumptions

official-blogmaturity: currentstrength: heuristicfreshness: stableverified 2026-08-20published 2025-06-18cited in 11

No local-vs-API framework; assumes $320K 8-GPU servers

Git-versioned envs across RTX/Spark/Brev; agent sandboxing; no run/eval tracking by design

official-docsmaturity: currentstrength: heuristicfreshness: activeverified 2026-08-20published 2026-08-07cited in 02, 13

Marketing-quiet but actively developed, not deprecated

Batch-size heuristics; 'FP8 first'; omits NVFP4 entirely

official-docsmaturity: currentstrength: heuristicfreshness: stableverified 2026-08-21cited in 07, 14

Verified live (W1): 'prioritizing using FP8 first' verbatim; NVFP4 STILL ABSENT from this decision page (present in the same repo's llm_qat guide, NV-MODELOPTQAD-001) — the decision-page-lags-flagship trap stands as of 2026-08-21.

NV-CURATORDESIGNER-001REFERENCENeMo Curator + Data Designer (data curation & synthetic data tooling)NVIDIA

Pretraining-scale cleaning/dedup/decontamination; beta schema-driven synthetic data

official-docsmaturity: currentstrength: heuristicfreshness: activeverified 2026-08-20cited in 09

No small-dataset fine-tuning methodology; Data Designer beta API

NV-DYNAMOAICONFIG-001REFERENCEDynamo profiler + AIConfiguratorNVIDIA

Datacenter operating-point automation and sizing; no RTX/GB10 data

official-docsmaturity: currentstrength: heuristicfreshness: activeverified 2026-08-20cited in 06, 11

Dynamo router = KV-cache worker placement, not model selection

Knative-revision traffic splitting; rollback mechanics delegated to Knative

official-docsmaturity: currentstrength: heuristicfreshness: stableverified 2026-08-20cited in 10

Native NIM Operator path = plain rolling update, rollback manual/undocumented

NV-NAT-001REFERENCENeMo Agent Toolkit (NAT: nat eval, profiler, sizing calculator)NVIDIA

RAGAS/trajectory judges, token/latency profiler, whole-workflow sizing

official-repomaturity: currentstrength: heuristicfreshness: activeverified 2026-08-20cited in 03, 05, 06

No harness-selection guidance; no variance/CI for stochastic agents

NV-NEMOCLAW-001REFERENCENemoClaw / OpenShell agent sandboxNVIDIA

Credential-custody wrapper; host-side cost/quality Model Router; demo-grade

official-repomaturity: pre-alphastrength: heuristicfreshness: volatileverified 2026-08-20published 2026-03-16cited in 08, 13

Wraps community harnesses OpenClaw (Steinberger)/Hermes (Nous); 'privacy router' is positioning

NV-NEMOTRONCC-LICENSE-001REFERENCENemotron-CC quality-routing rule + Open Model LicenseNVIDIA

Score >11 vs <=11 quality routing; license permits Nemotron outputs for training

official-docsmaturity: maturestrength: heuristicfreshness: stableverified 2026-08-20published 2025-10-24cited in 09, 13

Pretraining-scale corpus tooling, tangential to small-lab curation

NV-LEPTONBREV-001REFERENCEDGX Cloud Lepton / BrevNVIDIA

Rentable multi-provider GPU marketplace/console; no local hardware on-ramp documented

official-docsmaturity: currentstrength: heuristicfreshness: activeverified 2026-08-21cited in 10, 11

Flagged NOT-RELEVANT-YET; register-to-brev does not migrate workloads out

EXT-PYTORCH-SPARKFT-001REFERENCEPyTorch.org DGX Spark fine-tune case study (synthetic CoT -> TorchTune -> BFCL eval)PyTorch (Meta/PyTorch Foundation)

Only end-to-end Spark fine-tune experiment with real before/after evaluation

third-partymaturity: maturestrength: heuristicfreshness: snapshotverified 2026-08-20published 2026-02-02cited in 03, 06, 09

Not an NVIDIA property; fills the Spark eval gap NVIDIA leaves open

NV-TAO-HPO-001REFERENCETAO Toolkit AutoML HPO guidanceNVIDIA

Only full HPO methodology NVIDIA publishes: Hyperband/ASHA/BOHB/PBT selection

official-docsmaturity: maturestrength: heuristicfreshness: stableverified 2026-08-20cited in 04

Computer-vision-only; ASHA/Hyperband selection logic reads across to LLM sweeps

EXT-STOPPING-001REFERENCEWald SPRT + always-valid sequential testing (incl. confidence sequences, Waudby-Smith & Ramdas 2103.06476) (arXiv:2410.16076)Wald; Fischer & Ramdas

SPRT sequential test; dominated by exact counting on the source project's clustered data

papermaturity: maturestrength: contestedfreshness: stableverified 2026-08-21published 2024-10cited in 04

SPRT deleted from contract; curtailed exact counting exact counting adopted instead

EXT-EVAL-005REFERENCESWE-bench Verified curationOpenAI / swebench.com

human review 1699→500 tasks; solve rate 16%→33.2% from decontamination

official-blogmaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2024cited in 00, 03

W1 WITH DOWNGRADE (FOLLOW→REFERENCE): the historical curation figures (1,699 reviewed → 500 kept; 16%→33.2%) are corroborated via the SWE-bench repo + swebench.com and remain valid as the eval-curation case study. But the publisher's own ~Feb-2026 follow-up (EXT-EVAL-008) declares the benchmark saturated for frontier claims and reports flawed test cases in a hard-problem sample — cite this source ONLY for the historical curation lesson, never as a currently-recommended live benchmark. openai.com pages 403-blocked; secondary corroboration.

EXT-JUDGE-001REFERENCEMT-Bench LLM-as-judge agreement study (arXiv:2306.05685)Zheng et al. (LMSYS)

GPT-4-judge ~80% human agreement on open-ended chat preference

papermaturity: maturestrength: strong-evidencefreshness: stableverified 2026-08-21published 2023cited in 03

strong evidence but wrong domain for the source project's verifiable tasks

EXT-JUDGE-003REFERENCEReliability without Validity (judge κ-deflation study) (arXiv:2606.19544)arXiv:2606.19544 authors

chance-corrected κ deflates judge quality 33-41pp; stable judge bias

papermaturity: newstrength: contestedfreshness: activeverified 2026-08-21published 2026cited in 03

21 judges, 541K judgments; sets bar if the source project ever adds a judge

EXT-FT-002REFERENCELIMA (style-alignment sample size) (arXiv:2305.11206)Zhou et al.

n=1000 sizing is for broad style alignment, not narrow fixes

papermaturity: maturestrength: research-onlyfreshness: stableverified 2026-08-21published 2023cited in 09

no primary source exists for minimum-n on narrow behavioral fix; field gap

EXT-FT-007REFERENCEError-driven data curation research cluster (LoopTool, CurateEvo, AlpaGasus) (arXiv:2511.09148)LoopTool; CurateEvo; Chen et al.

promising unreviewed preprints; pattern matches the source project's design, pilot-grade only

papermaturity: research-onlystrength: research-onlyfreshness: activeverified 2026-08-20published 2025/26cited in 09

AlpaGasus filtering grouped in gap matrix alongside these

EXT-ROUTE-001REFERENCECascade/learned-router literature (FrugalGPT, RouteLLM, AutoMix, Hybrid LLM) (arXiv:2305.05176)FrugalGPT/RouteLLM/AutoMix/HybridLLM authors

all four use learned gates; the source project's deterministic the deterministic-gate experiment gate is novel

papermaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published 2023-2024cited in 08

RouteLLM's OOD near-random result is the key caution for learned routers

EXT-OPS-003REFERENCEExperiment lineage tooling docs (MLflow/W&B/DVC)MLflow; Weights & Biases; DVC

provide versioning/lineage-by-reference; none has tamper-evidence

official-docsmaturity: maturestrength: heuristicfreshness: stableverified 2026-08-20cited in 00, 13

adopt only as metrics dashboard, never as the record of record

EXT-PERF-003REFERENCEServing-sensitivity research cluster (LLM-42, MarginGate, Silent Hyperparameter, LLM Output Drift) (arXiv:2605.19537)multiple (arXiv authors; ICAIF)

backend choice alone moves scores up to 16.6pp; fintech-domain LLM output-drift study (ICAIF)

papermaturity: research-onlystrength: research-onlyfreshness: activeverified 2026-08-20published 2025/26cited in 02, 05, 06

cited in source register, not independently surfaced in main narrative

EXT-TESTBED-002REFERENCESmall-to-large transfer literature (muP, scaling laws) (arXiv:2001.08361)Microsoft; Kaplan et al.

transfer validated only for smooth continuous metrics, not discrete pass/fail

papermaturity: maturestrength: strong-evidencefreshness: stableverified 2026-08-21published 2020cited in 00, 14

rebuts the minimal-testbed steelman for the source project's discrete decisions

EXT-TESTBED-003REFERENCEEfficient-eval/small-model-set reliability literature (tinyBenchmarks, Anchor Points, IRT-reliability, test-pyramid) (arXiv:2607.15190)Maia Polo et al.; Vivek et al.; multiple

scalable estimators can produce unreliable item/ranking inferences at small N

papermaturity: currentstrength: contestedfreshness: activeverified 2026-08-21published 2024-2025cited in 03, 04

quote resolves Opus adversary's bare-citation critique

EXT-HW-001REFERENCE2026-08-20 hardware/cloud price-verification snapshotNewegg; NVIDIA; Apple; RunPod; hardware-corner; community threads

GPU/DRAM/cloud prices verified live amid AI memory crunch, weeks shelf-life

third-partymaturity: currentstrength: heuristicfreshness: snapshotverified 2026-08-20published 2026-08-20cited in 11

2026-08-20 price snapshot — stale by design; the durable content is the re-verify-at-order-time rule. Every price cited from here MUST print its as-of date.

Deterministic mode via Thinking Machines batch-invariant kernels; documented 25–45% slowdown, recommended for debugging/reproducibility

official-repomaturity: maturestrength: strong-evidencefreshness: activeverified 2026-08-21cited in 02, 05

Verified active (W1); deterministic blog lmsys.org/blog/2025-09-22-sglang-deterministic/ live.

CPU/GPU GGUF engine; very active (multiple releases/day at verification); wide quantization-format support; single-slot serving suits determinism/provenance regimes

official-repomaturity: maturestrength: strong-evidencefreshness: activeverified 2026-08-21cited in 02, 05

Verified active (W1): release b10545 dated 2026-08-21.

NVIDIA-native engine; active (1.3.0rc-series on NGC 2026-08); NIM has moved its default backend to vLLM

official-repomaturity: currentstrength: heuristicfreshness: activeverified 2026-08-21cited in 05

Verified active (W1).

The publisher of the curated benchmark later declared it saturated for frontier claims and reported that a large share of a sampled hard-problem subset had flawed test cases rejecting functionally correct submissions; recommends a successor benchmark

official-blogmaturity: currentstrength: strong-evidencefreshness: stableverified 2026-08-21published ~2026-02 (per third-party citation)cited in 03

Discovered during W1 (2026-08-21); direct fetch 403-blocked, content corroborated via secondary citation — flagged as not first-party-verified. Doubly instructive for ch.03: even a human-curated suite carried instrument defects found only later — instrument validation never ends; and benchmark verdicts age (freshness discipline).

EXT-OPS-004REFERENCEKayenta automated canary analysisNetflix/Google (Kayenta)

Automated statistical canary judgment — the CEILING of canary automation, explicitly NOT the small-team default (SRE restraint doctrine applies)

official-repomaturity: maturestrength: heuristicfreshness: stableverified 2026-08-20published 2018cited in 10

Exact statistical test not independently verified (inherited flag from research pass) — cite as existence proof only.

DEPRECATED Superseded or no longer current — kept for the record, do not follow.

log-curate-tune-judge-human-gated promotion loop; 'flashlight, not autopilot'

official-repomaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21published Apr 2026cited in 12, 14

Verified live (W1): Apr-2026 deprecation notice verbatim, 'reference only', no successor named. Loop design (log→stratify→tune→judge→human-gated promotion; 'flashlight, not autopilot') remains citable as a design reference.

Superseded by NeMo Switchyard

official-repomaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21published Jul 2026cited in 08

Verified (W1): deprecation notice (added 2026-07-08) lives on the 'experimental' branch README — which IS the default branch, so visitors land on it; 'main' carries no notice. Points to Switchyard.

Superseded by AIPerf

official-repomaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21published 2026cited in 06

Verified live (W1): perf_analyzer README carries the deprecation warning verbatim, pointing to AIPerf (github.com/ai-dynamo/aiperf, drop-in replacement per its migrating.md).

NV-RTXAITOOLKIT-DEP-001DEPRECATEDRTX AI ToolkitNVIDIA

Only official RTX fine-tune-to-deploy workflow; died with no successor

official-repomaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21published 2025-11-21cited in 07, 10

Verified (W1): archived:true; README deprecation line present; no successor named.

NV-NEMOALIGNER-DEP-001DEPRECATEDNeMo-AlignerNVIDIA

Superseded by NeMo RL

official-repomaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21published 2025-05-15cited in 09

Verified (W1): deprecated in favor of NeMo RL (github.com/NVIDIA-NeMo/RL).

NV-NIMTRTLLM-DEP-001DEPRECATEDNIM TensorRT-LLM backendNVIDIA

Superseded by NIM vLLM backend

official-docsmaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21published 2026cited in 10

Verified (W1, PARTIAL — release-notes level): NIM vLLM backend is the going-forward path; Spark still ships a TRT-LLM playbook alongside.

NV-MODELOPTRENAME-DEP-001DEPRECATEDTensorRT Model Optimizer (old name/docs domain)NVIDIA

Renamed to NVIDIA Model Optimizer at 0.40, new docs domain

official-docsmaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21published 2025-12-12cited in 07

Verified (W1): rebrand announced 2025-12-08 in README; old GitHub-Pages docs subsite hard-404s; old repo slug transparently redirects (GitHub rename). Update docs links, repo links survive.

NV-EVALDOCTREE-DEP-001DEPRECATEDNeMo Evaluator 2025-era doc tree (adapters/interceptors/params)NVIDIA

Superseded by SDK 0.3.x docs

official-docsmaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21published 2026cited in 03

Verified (W1): /latest/ redirects to /nightly/; versioned static doc paths do not exist; 2025-era adapters/interceptors URLs dead. Cite /nightly/ paths with as-of dates.

NV-LLMPTQ-DEP-001DEPRECATEDllm_ptq exampleNVIDIA

Superseded by hf_ptq example (ModelOpt 0.46)

official-repomaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21cited in 07

Verified (W1): examples/llm_ptq 404s; hf_ptq live. Supersession is inferred from the 404 + directory listing — no in-repo prose migration note exists.

NV-DGXCLOUD-DEP-001DEPRECATED'DGX Cloud' as rentable productNVIDIA

Repositioned as NVIDIA-internal; rentable destination now Lepton/Brev

official-docsmaturity: deprecatedstrength: snapshotfreshness: snapshotverified 2026-08-21published 2026cited in 10, 11

Verified (W1): rentable destinations present as DGX Cloud Lepton and Brev; 'DGX Cloud' proper no longer presents as a rentable product.