Glossary¶
Expansions for the acronyms and short names used across these docs and the
codebase. Terms are referenceable from other pages with MyST {term}`NCCL`
syntax. Widely-known terms (GPU, HTTP, JSON, YAML, CLI, SSH, VM) are omitted.
Algorithms and losses¶
- GRPO¶
Group Relative Policy Optimization — replaces PPO’s value network with a group-relative baseline: several rollouts per prompt, each advantage is its reward minus the group mean. See GRPO and the loss.
- GSPO¶
Group Sequence Policy Optimization — a GRPO variant using sequence-level importance ratios instead of token-level (the
sequence_level_importance_ratiosknob on the shared loss).- CISPO¶
Clipped Importance-Sampling-weight Policy Optimization (MiniMax-M1) — a clip variant that keeps every token in the gradient but freezes its importance weight at the clipped value via a stop-gradient, instead of dropping clipped tokens as the
min/maxsurrogate does (theuse_cispoknob on the shared loss). See GRPO and the loss.- DAPO¶
Decoupled Clip and Dynamic Sampling Policy Optimization — GRPO with an asymmetric “clip-higher” range and dynamic resampling of degenerate groups.
- GDPO¶
A multi-reward (multi-objective) advantage estimator — a GRPO generalization that computes a per-component leave-one-out baseline for each reward axis, sums them, then renormalizes. The name is not an expanded acronym in this codebase; read it as “multi-reward GRPO”. See GRPO and the loss.
- OPD¶
On-Policy Distillation — the student generates on-policy and a frozen teacher scores those rollouts via log-probs;
OPDAdvantageEstimatorforms the advantagesg[log π_teacher − log π_student](selected byadv_estimator='opd'). MOPD is the multi-teacher form (routed per sample). See Distillation.- xtoken¶
Cross-tokenizer distillation — distillation from a teacher whose tokenizer differs from the student’s: a token aligner maps teacher to student tokens and the teacher’s logits are projected onto the student vocabulary for a chunk-averaged KL. Generalizes to several teachers, each with its own projection (a null projection marks a same-tokenizer teacher), combined per
kd_loss_mode. See Distillation.- RLOO¶
REINFORCE Leave-One-Out — the leave-one-out baseline (each sample’s baseline excludes itself), used by the GRPO estimator.
- PPO¶
Proximal Policy Optimization — the clipped-surrogate policy-gradient method that GRPO/DAPO/GSPO are expressed as variants of.
- PG¶
Policy Gradient. “Clipped-PG” is the single configurable clipped policy-gradient loss (
ClippedPGLossFn) that expresses PPO, GRPO, RLOO, DAPO, GSPO, and CISPO.- DPO¶
Direct Preference Optimization — trains directly on preference pairs, no reward model or sampling loop. See Preference optimization.
- DPOP¶
DPO-Positive — adds a positive-penalty margin term to the DPO loss (
extra_margin).- R-DPO¶
Regularized DPO — adds a response-length penalty (
pre_sigmoid_offset).- cDPO¶
conservative DPO — label-smoothed DPO that mixes in the flipped target.
- SimPO¶
Simple Preference Optimization — reference-free preference loss (drops the reference-model term).
- ORPO¶
Odds Ratio Preference Optimization — a reference-free preference method.
- RLAIF¶
Reinforcement Learning from AI Feedback — preference/critique signal from an LLM judge rather than human labels.
- SFT¶
Supervised Fine-Tuning — standard next-token training on demonstrations; also the auxiliary term blended into DPO to keep the policy on-distribution.
- RM¶
Reward Model / Reward Modeling — training a model to score responses.
- KL¶
Kullback–Leibler divergence — the reference-policy penalty that bounds how far the policy drifts (
k1/k2/k3estimators).- NLL¶
Negative Log-Likelihood.
- IS¶
Importance Sampling — the off-policy correction that reweights rollouts generated by a slightly older policy (mandatory under async GRPO).
Models, parallelism, precision¶
- DTensor¶
Distributed Tensor — PyTorch’s sharded-tensor abstraction underpinning the default FSDP2 trainer backend.
- FSDP2¶
Fully Sharded Data Parallel, version 2 — PyTorch-native parameter/gradient/ optimizer-state sharding.
- TP¶
Tensor Parallelism — splits individual layers (matmuls) across devices.
- CP¶
Context Parallelism — splits the sequence dimension across devices.
- EP¶
Expert Parallelism — splits Mixture-of-Experts experts across devices.
- DP¶
Data Parallelism — replicates the model and splits the batch.
- MoE¶
Mixture of Experts — a sparse layer that routes each token to a few of many expert MLPs. See Mixture of Experts.
- Router replay (R3)¶
Reusing generation’s recorded per-token expert selection in the trainer’s MoE forward, so train-time routing matches generation and their log-probs line up (removing the router-nondeterminism component of the log-prob error). Gated by
policy.router_replay.enabled. See Mixture of Experts.- GEMM¶
General Matrix Multiply. “Grouped-GEMM” batches the per-expert matmuls of an MoE layer into one call.
- LoRA¶
Low-Rank Adaptation — parameter-efficient fine-tuning via low-rank adapter matrices.
- FP8¶
8-bit floating point — the optional quantized format for vLLM inference serving (training stays bf16).
- NVFP4¶
NVIDIA 4-bit floating point — a ModelOpt quantized format optionally served during rollout (
real_quant) through vLLM’s FP4 kernel and re-applied across refits. See Quantized generation.- bf16¶
bfloat16 — 16-bit “brain” floating point; the default training precision.
- KV cache¶
Key–Value cache — the attention cache the inference engine reuses across decode steps.
JAX backend¶
- NNX¶
Flax NNX — the Flax neural-network module system the JAX trainer is built on (a name, not an acronym). See the JAX trainer.
- GSPMD¶
General/Global Single-Program Multiple-Data — XLA’s automatic-sharding parallelization that
jax.shardingdrives.- XLA¶
Accelerated Linear Algebra — the compiler backend JAX lowers to.
- HF¶
Hugging Face — the source of model weights/configs; the trainer keeps HF-layout parity and an HF-weight name map.
- J0–J11¶
The JAX trainer re-platform build phases. J0–J9/J11 are CPU-validated; J10 (GPU bring-up) and J7b are hardware-gated. Tracked in the HV ledger (
hardware-deferred-validation.md).
Systems, networking, deployment¶
- NCCL¶
NVIDIA Collective Communications Library — the GPU collective transport used for the trainer→inference weight sync. See weight sync.
- IPC¶
Inter-Process Communication — a colocated weight-sync transport.
- ZMQ¶
ZeroMQ — the messaging library behind one of the weight-sync/data-plane transports.
- RDMA¶
Remote Direct Memory Access — NIC-to-NIC memory transfer used by high-speed cluster fabrics.
- RoCE¶
RDMA over Converged Ethernet — RDMA on an Ethernet fabric; auto-provisioned as an opt-in DRA resource on supported nodes.
- CNI¶
Container Network Interface — the Kubernetes networking plugin; NetworkPolicy is only enforced by a CNI that supports it.
- PDB¶
PodDisruptionBudget — a Kubernetes object that bounds voluntary pod eviction (node drain/upgrade). (Not the Python debugger.)
- DRA¶
Dynamic Resource Allocation — the Kubernetes mechanism for claiming specialized devices (here, NVLink ComputeDomain and RoCE NICs).
- KAI¶
The KAI Scheduler — gang scheduling + hierarchical fair-share queues on Kubernetes (a product name).
- KVM¶
Kernel-based Virtual Machine — a Linux hypervisor used as one of the computer-use agent’s guest backends.
- ABC¶
Abstract Base Class — the Python interface classes the trainer backends implement (
models/policy/interfaces.py).- HV¶
Hardware-deferred Validation — the ledger of checks that are written and CPU-validated but unverifiable without GPU/cluster hardware (
hardware-deferred-validation.md), each taggedHV-<n>.
Benchmarks and environments¶
- SWE-bench¶
Software Engineering benchmark — real GitHub issues scored by running the repository’s held-out tests. See SWE environments.
- CUA¶
Computer-Use Agent — the multimodal GUI-control agent (mouse/keyboard on a desktop). See OSWorld / CUA.
- HLE¶
Humanity’s Last Exam — a hard knowledge/reasoning benchmark.
- GDPval¶
A benchmark of 220 long-horizon, file-producing professional tasks (the name is the benchmark’s own). See GDPval.