Source ledger
FOLLOW Adopt this recipe as written; validate only your project-specific parameters.
Metrics, ISL/OSL sweeps, operating-point selection, sizing, $/M-token TCO; AIPerf-based
Verified live (W1): AIPerf-based; TTFT/ITL/TPS definitions, concurrency sweeps, ISL/OSL pairs, operating-point selection all present verbatim. 'latest' resolves to guide v2.0.0. 2025 blog series' GenAI-Perf commands deprecated, unbannered.
McNemar paired testing, power/MDE tables, GO/NO-GO gate w/ 95% CIs
Verified live (W1): compare (McNemar exact on discordant pairs, power table 10~28%/1000~2.8%, MDE reported, INCONCLUSIVE when underpowered, identical-prompt-template requirement) and quality-gate (tiers, 95% CI on paired delta, INSUFFICIENT_EVIDENCE <10 items, GO/NO-GO/INCONCLUSIVE exit codes) confirmed verbatim. SDK v0.3.0 (2026-06-03) still latest. CITATION PATH: /latest/ 308-redirects to /nightly/ — no pinned doc tree.
Source method nel compare productized; usable standalone on lm-eval output
Verified (W1): stable citation is arXiv:2602.10144 (ICLR 2026). Code repo amazon-science/LLM-Accuracy-Stats archived read-only 2026-05-08 — usable, unmaintained.
arXiv:2411.00640)Miller; NCSSclustered/paired SEs, power/MDE formulas; largest pure gap in the source project practice
arXiv 2411.00640 verified resolving, title 'Adding Error Bars to Evals' (W1). Drives clustered/paired SE and power/MDE machinery in ch.04.
objective→dataset→metrics→run→continuous; validate model-graders against humans
Verified live (W1) WITH PLATFORM SUNSET: the hosted OpenAI Evals platform 'will become read-only for existing users on October 31, 2026' and 'is scheduled to shut down on November 30, 2026'. The methodology content (objective→dataset→metrics→run→iterate; validate model graders against humans) remains citable; do NOT present the platform as current tooling.
arXiv:2404.12272)Shankar et al.criteria and ground truth co-evolve; most load-bearing validation of the source project's design
n=9 practitioner study; validates frozen-suite+versioned-release discipline
up to 8pp accuracy drop on fresh benchmark; contamination r²=0.36
validates the source project disjoint seed ranges while corpus stays private
arXiv:2410.12784)Tan et al.GPT-4o-judge ~56.6% near chance; vindicates the source project's deterministic grading
numbers verified from paper's tables, not summarizer
arXiv:2403.01432)Ovadia; Soudani; Balaguerthree studies agree RAG/prompt-first for factual/citation tasks
validates the source project's evidence-gated ladder ordering
VRAM floors: 9B~6.5GB, 14B~8.5GB, 27B~22GB, 70B~41GB
cross-checked against Unsloth docs
arXiv:2405.09673)Biderman; Thinking Machinesall-layer LoRA at post-training scale ~= full FT at ~67% FLOPs
Thinking Machines paper reconciles with Biderman's single-model scope caveat
prohibits training other models on Claude I/O without authorization
Verified live (W1), both pages directly fetchable. §D.4 verbatim: customers 'may not ... access the Services to build a competing product or service, including to train competing AI models...'. AUP separately prohibits 'utilization of inputs and outputs to train an AI model' (model scraping/distillation) absent authorization. Re-verify every authoring pass and per project.
28-test rubric, score=minimum across categories; adopt as annual self-audit
Verified from the PDF itself (W1): 28 tests; final score = MINIMUM across the 4 sections; Monitor-3 = training/serving compute the same values; Monitor-7 = prediction-quality regression via held-out production examples. NOTE: a first-pass WebFetch summarization hallucinated wrong authors/venue — verified against page images; cite Breck, Cai, Nielsen, Salib, Sculley, IEEE Big Data 2017.
correctness-only sanity gates before scaling up
the source project's make reachability/gold/determinism probes formalize this
'Use the simplest model that meets your technical and business objectives'; run ONE canary at a time; canary metrics must be causally attributable to the change
Verified live (W1), phrasing near-verbatim.
Consumer ToU: 'Use Output to develop models that compete with OpenAI' is prohibited; Business Terms §3.3(e): may not 'use Output to develop artificial intelligence models that compete' except Permitted Exceptions (§17)
First verification 2026-08-21 — via text-extraction proxy only (direct fetch 403-blocked; archive.org unreachable). Content confirmed; treat as verified-with-caveat and re-verify per project.
'You may not use the Services to develop models that compete with the Services' (Use Restrictions)
Verified live 2026-08-21, directly fetchable. Materially newer effective date than the Anthropic/OpenAI documents — vendor terms move on different clocks; check each provider at training time.
Endpoint-agnostic benchmarking vs any OpenAI-compatible endpoint; v0.12.0 adds agent workloads, Anthropic Messages endpoint, accuracy benchmarking; drop-in GenAI-Perf replacement
Verified active (W1). Pin one tool + one metric-definition set across tiers.
ADAPT Sound core, but adapt it — parts are project-specific, stale, or need validation.
Weak/strong escalation router; 4 mechanisms; confirmations=2, fail-open, latch
Verified live (W1): still pre-alpha 'Not for production use'; escalation semantics confirmed (weak-first, judge on actual output, confirmations=2, per-session latch, fail-open, judge failure HOLDS streak). PRECISION: capable_target/efficient_target two-tier schema belongs to the stage_router route type. v0.2.0 was a substantial redesign (native Rust server) — do not cite v0.1.0-era CLI. Prefill router still unshipped (zero 'prefill' paths on main).
Break-even rule: min offload=judge_cost/(strong-weak); 74% cost cut, 93% accuracy retained
Verified live (W1): break-even rule and all figures verbatim (74% cost cut, 93% accuracy retention 80.0/86.0, escalation mean 6.9% over 5 runs, range 4.1–9.1%). Scope: 145-task agent suite, one weak/strong pairing — numbers are conditioned on that setup.
Evaluate-first ordering; verifiability/resources/maturity framework; TSR over accuracy
Verified live (W1): 'The most successful teams start with lightweight methods, invest early in evaluation, and layer in training-based techniques where measurement shows they're needed.' Companion evaluation post (2026-05-19): developer.nvidia.com/blog/mastering-agentic-techniques-ai-agent-evaluation/ — TSR over accuracy, instrument from day one.
Demonstrated >99% median-recovery PTQ bar; 95-99% deliberate QAD target; worked example
Verified live (W1): '>99% median accuracy recovery is targeted when performing PTQ only'; '95–99% ... when combined with QAD'; PTQ 96.33% → QAD 99.72% worked example — all verbatim. Demonstrated practice, not documented doctrine.
Binomial margin-of-error table by sample size; first-N ordering-bias warning
Verified live (W1): margin-of-error table (±8.3pp @140 → ±0.9pp full) and first-N ordering-bias warning confirmed verbatim.
On-device 35B-MoE to NVFP4 in ~2h; W4A16 vs W4A4 recipe rule
Verified live (W1, file updated 2026-07-27): W4A16 interactive/low-concurrency vs W4A4 high-concurrency recipe rule verbatim; 'we recommend running evaluations' with no procedure/threshold. Format is Blackwell-only — quantize for the deployment target.
Publish the complete evaluation recipe alongside results; methodological consistency with clear provenance rather than bit-wise identical outputs
Verified live (W1) WITH CORRECTION: two claims previously attributed here (pin containers/sampling-params/judges; smoke-test discipline) are NOT on the live page — do not cite this source for them. Confirmed: full-recipe publication; 'the purpose of open evaluation is not to force bit-wise identical outputs, but to deliver methodological consistency with clear provenance'. A --dry-run flag is mentioned in passing only.
Prometheus /v1/metrics (vLLM passthrough), OTel traces, JSONL logs = production minimum
Verified live (W1, page updated 2026-08-19): Prometheus /v1/metrics with unmodified vLLM-native metric passthrough, OTel spans via OTEL_EXPORTER_OTLP_TRACES_ENDPOINT, JSONL logs via NIM_JSONL_LOGGING. Classified ADAPT: production mechanics documented and adoptable; the experimentation telemetry minimum is missing and is supplied by chapter 12.
LoRA-vs-SFT selection guidance (LoRA for 1–2 GPUs / fast iteration / multiple specialized versions); GPU sizing guidance
Verified (W1) WITH CORRECTION: Customizer docs carry LoRA-vs-SFT guidance, but the early-stopping defaults (val_loss, patience 10, min-delta 0.001) are NOT in the Customizer docs — they exist in NeMo's OSS training-framework code. Cite as code-level defaults, not published vendor guidance. K8s footprint.
Only source for LoRA-vs-QLoRA criteria; NVIDIA delegates methodology here
Verified live (W1): VRAM floors exact match (9B 6.5GB QLoRA/24GB LoRA; 27B 22GB/64GB; 70B 41GB/164GB — stated as absolute minimums); all-major-linear-layers LoRA targeting, rank 8–128 guidance, alpha=2r. CHANGE: multi-GPU now 'works but a much better version is coming' (no longer just 'coming soon'). Unsloth publishes no dates — staleness undatable from the page.
Dataset thresholds: 100-1,000 pairs PEFT / 1,000+ full FT
Verified live (W1) WITH CORRECTIONS: correct host is blogs.nvidia.com (not developer.nvidia.com/blog); published 2025-12-15. Thresholds confirmed exact: PEFT 100–1,000 prompt-sample pairs; full fine-tuning 1,000+.
85/15 split + independent xlam golden set; 500+ synthetic samples to ~93%
Z-score-vs-calibration-bag pattern for security/regression baselines
Pattern worth copying for internal regression baselines
Trajectory interchange format across NAT/Relay/Platform
Adopt if the source project stores trajectories beyond its event-bus records
PMC3248853)Lan & DeMets; PMC/NIHpre-register adaptation rule, DSMB executes it not investigator judgment
CP~0.20 convention & Lan-DeMets spending explicitly dropped, unverifiable
ASHA-style rungs adopted for TRAIN screening, must be class-balanced
corrected: rungs must be class-balanced, Bernstein racing at class level
multidimensional explicit criteria, grader by task shape, edge-case checklists
Verified live (W1): multidimensional explicit criteria, grader-by-task-shape, edge-case checklists; volume heuristic ('prefer higher volume with slightly lower signal') present. The volume-vs-curation tension stays visible in ch.03.
read traces, open-ended failure notes, synthesize taxonomy
Verified live (W1): the error-analysis-first / bottom-up taxonomy process is the field-guide post (2025-03-24) — cite it, not the older evals post.
cooperative-scraper exclusion convention for published task text
new action: add canary GUID to TEST files and published excerpts
arXiv:2406.12045)Chen et al.; Yao et al.gpt-4o agent pass^1>60% collapses <25% at pass^8, identical tasks
new action: measure frontier pass^k (k=3-5) on TEST subset
arXiv:2506.09501)Thinking Machines; SGLang; vLLM; llama.cppreduction-order/GPU-state variance dominant; batch-invariance costs 34-110% throughput
validates the source project's single-slot same-session reproducibility rule
arXiv:2308.08747)Luo et al.forgetting real, scale-trends non-obvious, merging doesn't reliably fix it
mandates full TEST suite as FT gate, never just target-behavior slice
inference benchmarking method + goodput concept for serving
DistServe headline figures blocked PDF, not independently verified
smoke/load/stress/spike/soak test types; thresholds-as-SLOs
Verified live (W1): smoke / average-load / stress / spike / soak / breakpoint all present with definitions.
Mirrors a copy of live inference traffic to a shadow variant in real time; comparison/interpretation left to the user
Verified live (W1).
Staged manual traffic shifting (small % → all) plus mirrored/shadow traffic (≤50%, one deployment, documented limits)
Verified live (W1).
'QAD is Model Optimizer's recommended strategy for accuracy recovery after quantization' (verbatim); references NVFP4 (which the choosing-quant-methods decision page omits)
Added during W1: the QAD-recommendation line lives here, not on the decision page — cite this for the recovery-strategy recommendation.
De-facto standard open serving engine; deterministic batch-invariant mode is OPT-IN (VLLM_BATCH_INVARIANT=1, CC>=8.0, reduced performance)
Verified active (W1). Throughput-regime default candidate in ch.05 regime table.
REFERENCE Background evidence and context; not an executable recipe.
Offline/online benchmark procedure, 4 engines, ttft/tpot/itl/e2el, concurrency=1
Buried asset in the connect-two-sparks playbook; no thresholds; pins stale containers. Not re-verified in authoring pass (REFERENCE stance).
46 install-verify-cleanup tutorials, no eval playbook, no sequencing
build.nvidia.com/spark JS pages unverified; categories/ordering not confirmed
Integrated evaluate/secure/tune/build; Switchyard under Tune; Studio alpha
Verified live (W1): active; pillars Secure/Evaluate/Tune/Build + synthetic data; Switchyard under Tune; driven from coding agents. Experiment-registry/run-history capability STILL ABSENT — the watch item stands.
$/M-token formula chain, Pareto precision comparison, datacenter-scale assumptions
No local-vs-API framework; assumes $320K 8-GPU servers
Git-versioned envs across RTX/Spark/Brev; agent sandboxing; no run/eval tracking by design
Marketing-quiet but actively developed, not deprecated
Batch-size heuristics; 'FP8 first'; omits NVFP4 entirely
Verified live (W1): 'prioritizing using FP8 first' verbatim; NVFP4 STILL ABSENT from this decision page (present in the same repo's llm_qat guide, NV-MODELOPTQAD-001) — the decision-page-lags-flagship trap stands as of 2026-08-21.
Intellectual backbone of SLM-default/LLM-fallback cascade thesis
Rationale, not procedure; cite as external validation when writing up
Pretraining-scale cleaning/dedup/decontamination; beta schema-driven synthetic data
No small-dataset fine-tuning methodology; Data Designer beta API
Datacenter operating-point automation and sizing; no RTX/GB10 data
Dynamo router = KV-cache worker placement, not model selection
Knative-revision traffic splitting; rollback mechanics delegated to Knative
Native NIM Operator path = plain rolling update, rollback manual/undocumented
Semver + per-version metrics; HF-compatible one-hop base_model/peft lineage fields
No end-to-end dataset-to-deployment lineage prescription; LoRA adapters versionless
RAGAS/trajectory judges, token/latency profiler, whole-workflow sizing
No harness-selection guidance; no variance/CI for stochastic agents
Credential-custody wrapper; host-side cost/quality Model Router; demo-grade
Wraps community harnesses OpenClaw (Steinberger)/Hermes (Nous); 'privacy router' is positioning
Score >11 vs <=11 quality routing; license permits Nemotron outputs for training
Pretraining-scale corpus tooling, tangential to small-lab curation
Rentable multi-provider GPU marketplace/console; no local hardware on-ramp documented
Flagged NOT-RELEVANT-YET; register-to-brev does not migrate workloads out
Only end-to-end Spark fine-tune experiment with real before/after evaluation
Not an NVIDIA property; fills the Spark eval gap NVIDIA leaves open
Only full HPO methodology NVIDIA publishes: Hyperband/ASHA/BOHB/PBT selection
Computer-vision-only; ASHA/Hyperband selection logic reads across to LLM sweeps
arXiv:2410.16076)Wald; Fischer & RamdasSPRT sequential test; dominated by exact counting on the source project's clustered data
SPRT deleted from contract; curtailed exact counting exact counting adopted instead
human review 1699→500 tasks; solve rate 16%→33.2% from decontamination
W1 WITH DOWNGRADE (FOLLOW→REFERENCE): the historical curation figures (1,699 reviewed → 500 kept; 16%→33.2%) are corroborated via the SWE-bench repo + swebench.com and remain valid as the eval-curation case study. But the publisher's own ~Feb-2026 follow-up (EXT-EVAL-008) declares the benchmark saturated for frontier claims and reports flawed test cases in a hard-problem sample — cite this source ONLY for the historical curation lesson, never as a currently-recommended live benchmark. openai.com pages 403-blocked; secondary corroboration.
GPT-4-judge ~80% human agreement on open-ended chat preference
strong evidence but wrong domain for the source project's verifiable tasks
arXiv:2606.19544)arXiv:2606.19544 authorschance-corrected κ deflates judge quality 33-41pp; stable judge bias
21 judges, 541K judgments; sets bar if the source project ever adds a judge
n=1000 sizing is for broad style alignment, not narrow fixes
no primary source exists for minimum-n on narrow behavioral fix; field gap
arXiv:2511.09148)LoopTool; CurateEvo; Chen et al.promising unreviewed preprints; pattern matches the source project's design, pilot-grade only
AlpaGasus filtering grouped in gap matrix alongside these
arXiv:2305.05176)FrugalGPT/RouteLLM/AutoMix/HybridLLM authorsall four use learned gates; the source project's deterministic the deterministic-gate experiment gate is novel
RouteLLM's OOD near-random result is the key caution for learned routers
provide versioning/lineage-by-reference; none has tamper-evidence
adopt only as metrics dashboard, never as the record of record
arXiv:2605.19537)multiple (arXiv authors; ICAIF)backend choice alone moves scores up to 16.6pp; fintech-domain LLM output-drift study (ICAIF)
cited in source register, not independently surfaced in main narrative
arXiv:2001.08361)Microsoft; Kaplan et al.transfer validated only for smooth continuous metrics, not discrete pass/fail
rebuts the minimal-testbed steelman for the source project's discrete decisions
arXiv:2607.15190)Maia Polo et al.; Vivek et al.; multiplescalable estimators can produce unreliable item/ranking inferences at small N
quote resolves Opus adversary's bare-citation critique
GPU/DRAM/cloud prices verified live amid AI memory crunch, weeks shelf-life
2026-08-20 price snapshot — stale by design; the durable content is the re-verify-at-order-time rule. Every price cited from here MUST print its as-of date.
Deterministic mode via Thinking Machines batch-invariant kernels; documented 25–45% slowdown, recommended for debugging/reproducibility
Verified active (W1); deterministic blog lmsys.org/blog/2025-09-22-sglang-deterministic/ live.
CPU/GPU GGUF engine; very active (multiple releases/day at verification); wide quantization-format support; single-slot serving suits determinism/provenance regimes
Verified active (W1): release b10545 dated 2026-08-21.
NVIDIA-native engine; active (1.3.0rc-series on NGC 2026-08); NIM has moved its default backend to vLLM
Verified active (W1).
The publisher of the curated benchmark later declared it saturated for frontier claims and reported that a large share of a sampled hard-problem subset had flawed test cases rejecting functionally correct submissions; recommends a successor benchmark
Discovered during W1 (2026-08-21); direct fetch 403-blocked, content corroborated via secondary citation — flagged as not first-party-verified. Doubly instructive for ch.03: even a human-curated suite carried instrument defects found only later — instrument validation never ends; and benchmark verdicts age (freshness discipline).
Automated statistical canary judgment — the CEILING of canary automation, explicitly NOT the small-team default (SRE restraint doctrine applies)
Exact statistical test not independently verified (inherited flag from research pass) — cite as existence proof only.
DEPRECATED Superseded or no longer current — kept for the record, do not follow.
log-curate-tune-judge-human-gated promotion loop; 'flashlight, not autopilot'
Verified live (W1): Apr-2026 deprecation notice verbatim, 'reference only', no successor named. Loop design (log→stratify→tune→judge→human-gated promotion; 'flashlight, not autopilot') remains citable as a design reference.
Superseded by NeMo Switchyard
Verified (W1): deprecation notice (added 2026-07-08) lives on the 'experimental' branch README — which IS the default branch, so visitors land on it; 'main' carries no notice. Points to Switchyard.
Superseded by AIPerf
Verified live (W1): perf_analyzer README carries the deprecation warning verbatim, pointing to AIPerf (github.com/ai-dynamo/aiperf, drop-in replacement per its migrating.md).
Only official RTX fine-tune-to-deploy workflow; died with no successor
Verified (W1): archived:true; README deprecation line present; no successor named.
Superseded by NeMo RL
Verified (W1): deprecated in favor of NeMo RL (github.com/NVIDIA-NeMo/RL).
Superseded by NIM vLLM backend
Verified (W1, PARTIAL — release-notes level): NIM vLLM backend is the going-forward path; Spark still ships a TRT-LLM playbook alongside.
Renamed to NVIDIA Model Optimizer at 0.40, new docs domain
Verified (W1): rebrand announced 2025-12-08 in README; old GitHub-Pages docs subsite hard-404s; old repo slug transparently redirects (GitHub rename). Update docs links, repo links survive.
Superseded by SDK 0.3.x docs
Verified (W1): /latest/ redirects to /nightly/; versioned static doc paths do not exist; 2025-era adapters/interceptors URLs dead. Cite /nightly/ paths with as-of dates.
Superseded by hf_ptq example (ModelOpt 0.46)
Verified (W1): examples/llm_ptq 404s; hf_ptq live. Supersession is inferred from the 404 + directory listing — no in-repo prose migration note exists.
Repositioned as NVIDIA-internal; rentable destination now Lepton/Brev
Verified (W1): rentable destinations present as DGX Cloud Lepton and Brev; 'DGX Cloud' proper no longer presents as a rentable product.
No sources match.