CHAPTER 01

Chapter 1: Chapter 1: Running and Phenomena: Observing External Behavior Starting from One AllReduce

Official Source: NVIDIA/nccl · Version: Commit @12df1a11 · Book Progress: Chapter 1 / 25

Chapter 1: Running and Phenomena: Observing External Behavior Starting from One AllReduce

Before diving into any kernel code, let's first get NCCL running and observe its externally exposed behavior. This chapter does not read the kernel; it does only one thing: establish a verifiable reference frame—any subsequent internal mechanism analysis must ultimately be able to explain the external behavior seen here.

1.1 Understanding NCCL's Engineering Structure from the Build Entry Point

Intuitive Model

The build system is like the construction blueprint of a building: it does not determine who lives in the building, but it determines what rooms exist and which way the doors open. If the build entry point is chaotic, you cannot even take the first step of "getting it running." NCCL provides two build entry points, Makefile and CMake. Understanding their differences is the first step to understanding this project's engineering organization.

Structure of the Two Build Entry Points

Top-levelMakefileis an extremely thin scheduling layer. It does not compile any source files itself, but forwards work to the Makefile in each subdirectory.

📎 Makefile:44-45definessrc.%pattern rules, forwarding targets such assrc.build、src.installtosrc/Makefile:

code
src.%:
	${MAKE} -C src $* BUILDDIR=${ABSBUILDDIR}

📎 Makefile:47-48definesexamplestarget, which depends onsrc.buildand then entersdocs/examplesdirectory to build examples:

code
examples: src.build
	${MAKE} -C docs/examples NCCL_HOME=${ABSBUILDDIR}

Note the dependency relationship here: the build of the examples depends onsrc.buildcompleting first, because the examples need to link against the NCCL library, and theNCCL_HOMEenvironment variable passes the build output directory to the examples' Makefile. This is the build order constraint of "library first, examples second."

📎 Makefile:29Lists all the target sets that can be cleaned:

code
TARGETS := src pkg nccl4py ir

📎 Makefile:30Using GNU Make's substitution reference syntax${TARGETS:%=%.clean}to expandsrc pkg nccl4py irintosrc.clean pkg.clean nccl4py.clean ir.clean, defining all cleanup targets at once. This is a common "data-driven rules" technique in Makefiles—adding a new module only requires adding one word toTARGETS.

CMake entry point: where the version number comes from

The CMake entry point is much more complex than the Makefile, because it has to handle cross-platform support, CUDA version detection, architecture selection, and more. We only focus on the parts directly related to "getting it running."

📎 CMakeLists.txt:5-11shows the source of the version number—it is not hardcoded in CMakeLists.txt, but read frommakefiles/version.mkand then extracted with a regex:

cmake
file(READ ${CMAKE_SOURCE_DIR}/makefiles/version.mk VERSION_CONTENT)
string(REGEX REPLACE ".*NCCL_MAJOR[ ]*:=[ ]*([0-9]+).*" "\\1" NCCL_MAJOR "${VERSION_CONTENT}")
...
math(EXPR NCCL_VERSION_CODE "(${NCCL_MAJOR} * 10000) + (${NCCL_MINOR} * 100) + ${NCCL_PATCH}")
[Design inference and architectural trade-offs]

Centralizing the version number inversion.mkallows both build systems, Makefile and CMake, to share the same version source, avoiding the classic engineering pitfall of "inconsistent version numbers across two build systems."NCCL_VERSION_CODEThe calculation formula forMAJOR*10000 + MINOR*100 + PATCHis consistent with theNCCL_VERSIONmacro in the header file.

📎 CMakeLists.txt:14-20Inject these version numbers into all C++ source files viaadd_compile_definitions:

cmake
add_compile_definitions(
    NCCL_USE_CMAKE
    NCCL_MAJOR=${NCCL_MAJOR}
    NCCL_MINOR=${NCCL_MINOR}
    NCCL_PATCH=${NCCL_PATCH}
    NCCL_VERSION_CODE=${NCCL_VERSION_CODE}
)

📎 CMakeLists.txt:24-25declares the project languages as CUDA, CXX, and C:

cmake
project(NCCL VERSION ${NCCL_MAJOR}.${NCCL_MINOR}.${NCCL_PATCH}
        LANGUAGES CUDA CXX C)

CUDA architecture selection: why the default value is so complex

📎 CMakeLists.txt:140-171is a large block of logic that determinesCMAKE_CUDA_ARCHITECTURESbased on the CUDA version. Take CUDA 12.8 and above as an example:

cmake
elseif(${CUDA_MAJOR} EQUAL 12)
    if(${CUDA_MINOR} LESS 8)
        set(CMAKE_CUDA_ARCHITECTURES "50;60;61;70;80;90")
    else()
        set(CMAKE_CUDA_ARCHITECTURES "50;60;61;70;80;90;100;120")
    endif()
[Design inference and architectural trade-offs]

The design motivation behind this logic is that the PTX of new architectures (such as 100 and 120) is only recognized by newer CUDA toolchains. If new architectures are forcibly specified for older CUDA versions, compilation will fail outright. Therefore, the default architecture list must be adjusted dynamically with the CUDA version. For readers, this means:If you do not explicitly setCMAKE_CUDA_ARCHITECTURES, the compiled artifact will contain a fatbin with a long list of architectures, and compilation time will increase significantly. Production environments usually specify the target architecture explicitly to speed up builds.

Build process decision diagram

The diagram below shows the complete decision path from executingmaketo producing a runnable example:

mermaid
flowchart TD
    start["执行 make 或 make examples"] --> check_ir{"EMIT_LLVM_IR 或<br/>NCCL_EMIT_LTO_IR 非 0?"}
    check_ir -->|是| add_ir["IR_GOALS 加入 llvm_ir/ltoir<br/>default 依赖 ir-emit"]
    check_ir -->|否| only_src["default 仅依赖 src.build"]
    add_ir --> src_build["make -C src build<br/>BUILDDIR=build"]
    only_src --> src_build
    src_build --> build_ok{"src.build 成功?"}
    build_ok -->|否| fail["构建失败,终止"]
    build_ok -->|是| is_examples{"目标是 examples?"}
    is_examples -->|是| ex_build["make -C docs/examples<br/>NCCL_HOME=build"]
    is_examples -->|否| done["产出 libnccl.so"]
    ex_build --> ex_ok{"示例链接成功?"}
    ex_ok -->|否| fail
    ex_ok -->|是| runnable["产出可执行示例"]

The key branch in this diagram is whetherIR_GOALSis non-empty—it determines whether the default build additionally triggers LLVM IR generation. For readers who just want to "get it running," keepingEMIT_LLVM_IR=0is enough to take the shortest path.

1.2 Prerequisites for a minimal runnable program

Intuitive model

Writing an NCCL program is like organizing a multi-party conference call. You first need to confirm: how many people are participating (number of devices), who each person is (rank), and what line is used for the call (stream). If any one of these is missing, the meeting cannot start. In this section, through the01_communicatorsexample, we will see clearly what these three prerequisites look like in code.

Data structures: three arrays carry all the state

📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:88-92defines the core variables of the example:

c
int num_gpus;                 // Number of available CUDA devices
ncclComm_t *comms = NULL;     // Array of NCCL communicators (one per GPU)
cudaStream_t *streams = NULL; // Array of CUDA streams (one per GPU)
int *devices = NULL;          // Array of device IDs to use

This reflects the core of NCCL's single-process multi-GPU programming model:one communication domain, one stream, and one device ID per GPU. The length of all three arrays isnum_gpus, and the indexicorresponds to theith GPU.

ncclComm_tis defined in the header file as an opaque pointer.📎 src/nccl.h.in:36gives its real type:

c
typedef struct ncclComm* ncclComm_t;
[Design inference and architectural trade-offs]

"Opaque pointer" is a classic technique in C for achieving information hiding: the header file only exposesstruct ncclComm*as a pointer type, user code cannot access the internal fields of the struct, and all operations must be performed through API functions. In this way, NCCL can freely modify the internal layout ofncclCommwithout breaking the ABI. For beginner readers, this can be understood as "what you get is a black-box handle, and you can only operate it through the official interface."

Step-by-Step: from device detection to communication domain creation

Step 1: Detect the number of devices. 📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:96-104callscudaGetDeviceCountand checks whether it is 0:

c
CUDACHECK(cudaGetDeviceCount(&num_gpus));

if (num_gpus == 0) {
    fprintf(stderr, "ERROR: No CUDA devices found on this system\n");
    ...
    return 1;
}

What this step is doing: asking the CUDA runtime, "How many GPUs are on this machine?" If it returns 0, it means there are no available devices, and the program exits directly—this is the earliest guard condition.

Step 2: Allocate host memory and fill in the device list. 📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:114-121allocates three arrays and checks whether allocation succeeded:

c
devices = (int *)malloc(num_gpus * sizeof(int));
comms = (ncclComm_t *)malloc(num_gpus * sizeof(ncclComm_t));
streams = (cudaStream_t *)malloc(num_gpus * sizeof(cudaStream_t));

if (!devices || !comms || !streams) {
    fprintf(stderr, "ERROR: Failed to allocate memory for device arrays\n");
    return 1;
}

📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:126-136fillsdevices[i] = iwith a loop and prints the properties of each device:

c
for (int i = 0; i < num_gpus; i++) {
    devices[i] = i; // Use device i for communicator i
    cudaDeviceProp prop;
    CUDACHECK(cudaGetDeviceProperties(&prop, devices[i]));
    printf("  GPU %d: %s (CUDA Device %d)\n", i, prop.name, devices[i]);
    ...
}

Step 3: Create a stream for each GPU. 📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:140-145is the key:

c
for (int i = 0; i < num_gpus; i++) {
    CUDACHECK(cudaSetDevice(devices[i]));
    CUDACHECK(cudaStreamCreate(&streams[i]));
}

Note thatcudaSetDevicemust be called beforecudaStreamCreate. This is a basic rule of CUDA programming:a stream belongs to the currently active device. If you do not switch devices first, the stream will be created on the wrong GPU. This is one of the pitfalls that beginners most easily fall into.

Step 4: Create the communication domain. 📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:169is the core call of the entire example:

c
NCCLCHECK(ncclCommInitAll(comms, num_gpus, devices));

ncclCommInitAllis a convenient entry point for the single-process multi-GPU scenario. The header file📎 src/nccl.h.in:301-301gives its contract:

c
/* Creates a clique of communicators (single process version).
 * This is a convenience function to create a single-process communicator clique.
 * Returns an array of ndev newly initialized communicators in comm.
 * comm should be pre-allocated with size at least ndev*sizeof(ncclComm_t).
 * If devlist is NULL, the first ndev CUDA devices are used.
 * Order of devlist defines user-order of processors within the communicator. */
ncclResult_t  ncclCommInitAll(ncclComm_t* comm, int ndev, const int* devlist);

The meanings of the three parameters:commis the preallocated communication domain array,ndevis the number of devices,devlistis the list of device IDs (if NULL is passed, the firstndevdevices are used). After the call returns,comms[i]is the communication domain of theith device, and its rank isi。

Step 5: Verify the communication domain properties. 📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:185-189verifies with three query APIs:

c
NCCLCHECK(ncclCommUserRank(comms[i], &rank));
NCCLCHECK(ncclCommCount(comms[i], &size));
NCCLCHECK(ncclCommCuDevice(comms[i], &device));

The definitions of these three APIs in the header file are📎 src/nccl.h.in:396、📎 src/nccl.h.in:400、📎 src/nccl.h.in:404. They answer three questions respectively: who am I (rank), how many people are there in total (size), and which GPU am I on (device).

Sequence diagram of the communication domain creation process

mermaid
sequenceDiagram
    participant App as 应用主线程
    participant CUDA as CUDA Runtime
    participant NCCL as NCCL 库
    App->>CUDA: cudaGetDeviceCount(&num_gpus)
    CUDA-->>App: num_gpus = N
    loop i in 0..N-1
        App->>CUDA: cudaSetDevice(devices[i])
        App->>CUDA: cudaStreamCreate(&streams[i])
        CUDA-->>App: streams[i]
    end
    App->>NCCL: ncclCommInitAll(comms, N, devices)
    Note over NCCL: 内部为每个设备建立通信域<br/>分配 rank 0..N-1
    NCCL-->>App: comms[0..N-1]
    loop i in 0..N-1
        App->>NCCL: ncclCommUserRank(comms[i], &rank)
        NCCL-->>App: rank = i
        App->>NCCL: ncclCommCount(comms[i], &size)
        NCCL-->>App: size = N
    end

This sequence diagram reveals the key point:ncclCommInitAllis asynchronous blocking call, which internally completes coordination among all devices, and when it returns, all communication domains are ready.

Design thinking: Why is ncclCommInitAll needed

[Design inference and architectural trade-offs]

In multi-process scenarios, each process manages only one GPU, so usingncclCommInitRankto initialize each separately is sufficient. But in single-process multi-GPU scenarios, if the user is asked to manually callncclCommInitRankfor each GPU, they must handle "synchronization among multiple ranks"—yet in a single process there is only one thread, which cannot advance the initialization of multiple ranks simultaneously, causing deadlock.ncclCommInitAllencapsulates this coordination inside the library, using internal mechanisms (usually multithreading or a state machine) to complete synchronized initialization of all ranks, exposing it to the user as a simple synchronous call. This is the fundamental reason the "convenience function" exists.

1.3 The complete external behavior of one AllReduce

Intuitive model

AllReduce is the most commonly used operation in collective communication: each participant contributes a piece of data, and everyone gets the sum of all data. It is like calculating the total score for a group assignment—everyone reports their own score, and in the end everyone has a copy of the class total. In this section we trace the03_collectives/01_allreduceexample to see the complete external behavior of one AllReduce from call to result verification.

Data structures: data buffers and initialization

📎 docs/examples/03_collectives/01_allreduce/c/main.cc:59-63defines the core variables:

c
int num_gpus = 0;
ncclComm_t *comms;
cudaStream_t *streams;
float **sendbuff;
float **recvbuff;

Note thatsendbuffandrecvbuffarefloat**—pointers to arrays of pointers. Eachsendbuff[i]is the device memory address on theith GPU.

📎 docs/examples/03_collectives/01_allreduce/c/main.cc:99defines the data size:

c
const size_t size = 32 * 1024 * 1024; // 32M floats for demonstration

32M floats, 4 bytes each, that is, a 128 MB send buffer and a 128 MB receive buffer, one copy per GPU.

📎 docs/examples/03_collectives/01_allreduce/c/main.cc:101-120is the initialization loop for each device:

c
for (int i = 0; i < num_gpus; i++) {
    CUDACHECK(cudaSetDevice(i));
    CUDACHECK(cudaStreamCreate(&streams[i]));
    CUDACHECK(cudaMalloc((void **)&sendbuff[i], size * sizeof(float)));
    CUDACHECK(cudaMalloc((void **)&recvbuff[i], size * sizeof(float)));
    CUDACHECK(cudaMemset(sendbuff[i], 0, size * sizeof(float)));
    float rank_value = (float)i;
    CUDACHECK(cudaMemcpy(sendbuff[i], &rank_value, sizeof(float),
                         cudaMemcpyHostToDevice));
    printf("  Device %d initialized with data value %d\n", i, i);
}

The clever part of this code: first clear the entire send buffer to zero, then set only thefirst elementtoi(the rank value of that device). In this way, after AllReduce summation, the result of the first element is0 + 1 + 2 + ... + (num_gpus-1), while all the remaining elements are 0. During verification, checking only the first element is enough to confirm whether AllReduce is correct.

Step-by-Step: AllReduce call and verification

Step 1: Group wrapping. 📎 docs/examples/03_collectives/01_allreduce/c/main.cc:130-136is the core call:

c
NCCLCHECK(ncclGroupStart());
for (int i = 0; i < num_gpus; i++) {
    NCCLCHECK(ncclAllReduce(sendbuff[i], recvbuff[i], size, ncclFloat, ncclSum,
                            comms[i], streams[i]));
}
NCCLCHECK(ncclGroupEnd());

Here there is anextremely important detail: the comment📎 docs/examples/03_collectives/01_allreduce/c/main.cc:128-129clearly states:

c
// NOTE: ncclGroupStart and ncclGroupEnd are essential to avoid
// deadlock when using ncclCommInitAll and multiple communication calls.

Why must Group be used? The header file📎 src/nccl.h.in:844-864gives the explanation:

c
/* Group semantics
 *
 * When managing multiple GPUs from a single thread, and since NCCL collective
 * calls may perform inter-CPU synchronization, we need to "group" calls for
 * different ranks/devices into a single call.
 * ...
 * Both collective communication and ncclCommInitRank can be used in conjunction
 * of ncclGroupStart/ncclGroupEnd, but not together.
 */
[Design inference and architectural trade-offs]

The core contradiction is: collective communication requires all ranks to participate at the same time, but in a single thread you can only callncclAllReduceone by one. If the firstncclAllReducecall blocks waiting for other ranks while the calls for the other ranks have not yet been issued, deadlock occurs. The role of the Group mechanism is:ncclGroupStartall calls afterncclGroupEndonly "register" and do not actually start;

only when 📎 docs/examples/03_collectives/01_allreduce/c/main.cc:139-142:

c
for (int i = 0; i < num_gpus; i++) {
    CUDACHECK(cudaSetDevice(i));
    CUDACHECK(cudaStreamSynchronize(streams[i]));
}

Step 2: Synchronize the stream.📎 src/nccl.h.in:854-856CopyncclGroupEndThe header fileemphasizes:only guarantees that the operation isenqueued to the stream, not that the operation

has completed 📎 docs/examples/03_collectives/01_allreduce/c/main.cc:152-169:

c
float expected = (float)(num_gpus * (num_gpus - 1) / 2);
...
for (int i = 0; i < num_gpus; i++) {
    float result;
    CUDACHECK(cudaSetDevice(i));
    CUDACHECK(cudaMemcpy(&result, recvbuff[i], sizeof(float),
                         cudaMemcpyDeviceToHost));
    if (result != expected) {
        printf("  Device %d received incorrect result: %.0f (expected %.0f)\n", i,
               result, expected);
        success = false;
    } else {
        printf("  Device %d correctly received sum: %.0f\n", i, result);
    }
}

Step 3: Verify the result.0 + 1 + ... + (N-1) = N*(N-1)/2Copy

The expected value is the sum of an arithmetic progression

mermaid
flowchart LR
    subgraph dev0["GPU 0 (rank 0)"]
        s0["sendbuff[0]<br/>首元素=0"]
        r0["recvbuff[0]"]
    end
    subgraph dev1["GPU 1 (rank 1)"]
        s1["sendbuff[1]<br/>首元素=1"]
        r1["recvbuff[1]"]
    end
    subgraph dev2["GPU 2 (rank 2)"]
        s2["sendbuff[2]<br/>首元素=2"]
        r2["recvbuff[2]"]
    end
    s0 -->|ncclAllReduce<br/>ncclFloat ncclSum| reduce["归约求和<br/>0+1+2=3"]
    s1 -->|ncclAllReduce<br/>ncclFloat ncclSum| reduce
    s2 -->|ncclAllReduce<br/>ncclFloat ncclSum| reduce
    reduce -->|广播结果| r0
    reduce -->|广播结果| r1
    reduce -->|广播结果| r2

AllReduce data flow diagramrecvbuffCopy

This diagram shows the two phases of AllReduce: first reduce, then broadcast. The

of each rank ultimately obtains the same result.

Design thinking: Why use Group instead of calling one by onencclGroupStart/ncclGroupEnd[Design inference and architectural trade-offs]

c
for (int i = 0; i < num_gpus; i++) {
    ncclAllReduce(sendbuff[i], recvbuff[i], size, ncclFloat, ncclSum,
                  comms[i], streams[i]);
}

is removed, the code becomes:ncclAllReduceCopy

In a single thread, when the first iteration calls

, NCCL needs to wait for all ranks to initiate AllReduce before it can proceed. But the calls for the other ranks have not yet been reached in the loop, so the first call can never wait for the other ranks, resulting in deadlock. The Group mechanism separates "initiation" and "execution," allowing all rank calls to be registered first and then executed together, fundamentally avoiding single-thread deadlock.

1.4 The lifecycle of a communicator and resource cleanup

Intuitive model

📎 docs/examples/03_collectives/01_allreduce/c/main.cc:176-183A communicator is like a meeting. Before the meeting you sign in (initialization), and after the meeting you adjourn (destruction). If the adjournment order is wrong—for example, locking the meeting room before people have left—problems arise. In this section we look at the destruction order of an NCCL communicator and why this order cannot be reversed.

c
NCCLCHECK(ncclGroupStart());
for (int i = 0; i < num_gpus; i++) {
    NCCLCHECK(ncclCommFinalize(comms[i]));
}
NCCLCHECK(ncclGroupEnd());
for (int i = 0; i < num_gpus; i++) {
    NCCLCHECK(ncclCommDestroy(comms[i]));
}

shows the standard destruction flow:📎 src/nccl.h.in:309-309CopyncclCommFinalizeThe header file

c
/* Finalize a communicator. ncclCommFinalize flushes all issued communications,
 * and marks communicator state as ncclInProgress. The state will change to ncclSuccess
 * when the communicator is globally quiescent and related resources are freed; then,
 * calling ncclCommDestroy can locally free the rest of the resources (e.g. communicator
 * itself) without blocking. */
ncclResult_t  ncclCommFinalize(ncclComm_t comm);

📎 src/nccl.h.in:313-313:ncclCommDestroy:

c
/* Frees local resources associated with communicator object. */
ncclResult_t  ncclCommDestroy(ncclComm_t comm);
explains

CopyncclCommFinalize[Design inference and architectural trade-offs]Why should destruction be split into two steps?is ancclCommDestroyglobal operation—it requires all ranks to participate, ensuring there is no in-flight communication.is ancclCommDestroylocal operation

—it only releases the resources of this process and does not block. This design decouples "waiting for all ranks to become quiet" from "releasing local resources": the former may take a long time (waiting for network peers), while the latter is a purely local operation. If there were only one

📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:221-249, it would have to assume both responsibilities at once, either blocking too long or failing to guarantee global quiescence.📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:218-219The complete chain of the destruction order

c
// IMPORTANT: Proper cleanup is critical for NCCL applications
// Resources must be cleaned up in the correct order to avoid issues

emphasizes:

Copy📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:224-227)

2. Finalize + Destroy communication domain (📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:233-240)

3. Destroy CUDA stream (📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:246-249)

4. Free host memory (📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:253-255)

Communication domain state machine

ncclCommFinalizeThe documentation explicitly mentions state transitions, which meets the admission criteria for a state machine:

mermaid
stateDiagram-v2
    [*] --> Active : ncclCommInitAll() 成功
    Active --> InProgress : ncclCommFinalize()<br/>刷新在途通信
    InProgress --> Quiescent : 全局静默<br/>相关资源释放
    Quiescent --> Destroyed : ncclCommDestroy()<br/>释放本地资源
    Destroyed --> [*]
    Active --> Aborted : ncclCommAbort()<br/>中止在途操作
    Aborted --> [*]

The key transition of this state machine isInProgress -> Quiescent: it is triggered by the "global quiescence" event, rather than directly triggered by a function call. This means that afterncclCommFinalizereturns, the communication domain may still be in theInProgressstate, and you need to pollncclCommGetAsyncErrorto know when it entersQuiescent。

Design consideration: why the destruction order cannot be reversed

[Design inference and architectural trade-offs]

If the CUDA stream is destroyed before the communication domain, what problems would occur? The communication domain may internally hold a reference to the stream (for example, for completion notification of asynchronous operations). If the stream is destroyed first, the communication domain accessing the already-destroyed stream during Finalize will cause undefined behavior. Similarly, if the host memory (commsarray) is freed before the communication domain is destroyed,ncclCommDestroythen a dangling pointer is obtained. This is why the order must be "synchronize first, then destroy the communication domain, then destroy the stream, and finally free host memory" —the dependency relationship determines that the destruction order must be the reverse of the creation order。

1.5 Production pitfall avoidance guide

Pitfall 1: Forgetting Group causes deadlock

This is the pitfall that beginners most often encounter. In a single-process multi-GPU scenario, if you directly callncclAllReducein a loop without adding Group, the program will deadlock on the first call. The symptom is: the program hangs and does not move, CPU usage is close to 0, and there is no output.

Troubleshooting method: usegdbto attach to the process and see whether the stack is stuck in NCCL's internal waiting logic. If so, check whetherncclGroupStart/ncclGroupEnd。

was omitted.

📎 src/nccl.h.in:854-856Pitfall 2: Reading results without synchronizing the streamncclGroupEndexplicitly states that📎 docs/examples/03_collectives/01_allreduce/c/main.cc:139-142only guarantees enqueueing, not completion. If you omitrecvbuffstream synchronization and directly read

, you will read incomplete data.cudaMemcpyThe symptom is: results are sometimes correct and sometimes wrong, or all zeros are read. This is becauseis synchronous by default, but what it synchronizes isthe current streamcudaStreamSynchronize, while AllReduce may execute on another stream. Troubleshooting method: add

before reading the result. If the problem disappears, this is the pitfall.

Pitfall 3: Incorrect destruction order causes segmentation faultncclCommDestroyIfcudaFreeis done beforesendbuff/recvbuff, the communication domain may still be accessing these buffers during Finalize, causing a segmentation fault or data corruption.

The symptom is: the program crashes during the exit phase, or occasionally reads garbage data. Troubleshooting method: check the order of the cleanup code and ensure that the communication domain is destroyed before all CUDA resources are released.

Pitfall 4: Confusing device number with rank

📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:198-200There is a validation:

c
if (device != devices[i]) {
    printf(" [WARNING: Expected device %d]", devices[i]);
}
[Design inference and architectural trade-offs]

rank and device are two different concepts. rank is the logical number within the communication domain (0 to nRanks-1), and device is the physical GPU number. InncclCommInitAll's default usage,devices[i] = i, so rank and device happen to be equal. But if a customdevlistis passed in (for example,{2, 0, 1}), rank 0 corresponds to device 2. Confusing these two concepts will cause data to be sent to the wrong GPU.

Chapter summary

In this chapter we completed three things:

1. Build entry point: understood the Makefile forwarding mechanism, the source of the CMake version number, and the CUDA architecture selection logic. The key conclusion is thatmake exampleswill build the library first and then build the examples,NCCL_HOMEand passes the build artifact directory to the examples.

2. The three elements of a minimal runnable program: number of devices (cudaGetDeviceCount), rank (automatically assigned byncclCommInitAll), stream (one per GPU).ncclCommInitAllis a convenient entry point for single-process multi-GPU, encapsulating multi-rank synchronous initialization inside the library.

3. The complete external behavior of one AllReduce: fromncclGroupStartwrapping multiplencclAllReducecalls, toncclGroupEndsubmission, tocudaStreamSynchronizewaiting for completion, and finally verifying the result. The Group mechanism is the key to avoiding deadlock in single-threaded multi-GPU scenarios.

4. Communication domain lifecycle:ncclCommFinalize(global quiescence) +ncclCommDestroy(local release) two-phase destruction, and the ordering constraint of "synchronize first, then destroy the communication domain, then destroy the stream, and finally free host memory."

Chapter reflection and self-test

Q1: If the ncclGroupStart/ncclGroupEnd of📎 docs/examples/03_collectives/01_allreduce/c/main.cc:130-136are removed and changed to directly calling ncclAllReduce in a loop, what will happen in a single-process multi-GPU scenario? Why?

Reference analysis: Deadlock will occur. The header file📎 src/nccl.h.in:844-864explains the reason: collective communication calls may perform inter-CPU synchronization and require all ranks to participate at the same time. In a single thread, when the first loop iteration callsncclAllReduce(comms[0], ...), NCCL needs to wait for other ranks to also initiate AllReduce before it can proceed. But the calls for other ranks have not yet been reached in the loop (because the current thread is blocked on the first call), so the first call will never wait for the other ranks, resulting in deadlock.

The role of the Group mechanism is to separate "initiation" and "execution":ncclGroupStartafterncclGroupEnd, all calls only register,

and only atgdbWhen you attach and look at the stack, it will be stuck in NCCL's internal wait logic, with CPU usage close to 0.

Q2: 📎 docs/examples/03_collectives/01_allreduce/c/main.cc:139-142Can the cudaStreamSynchronize be replaced with cudaDeviceSynchronize? What is the semantic difference between the two? In what scenarios would this replacement cause problems?

Reference analysis: It can be replaced withcudaDeviceSynchronize, but the semantics differ.cudaStreamSynchronize(streams[i])only waits for operations on the specified stream to complete;cudaDeviceSynchronizewaits for operations onallstreams on the current device to complete.

In a single-process multi-GPU scenario,cudaDeviceSynchronizeonly synchronizes the current device (determined bycudaSetDevice), so it needs to be used in conjunction with acudaSetDevice(i)loop. IfcudaSetDevice,cudaDeviceSynchronizeis omitted, only the default device (usually device 0) will be synchronized, and the AllReduce on other devices may not have completed yet.

The header file📎 src/nccl.h.in:854-856emphasizes thatncclGroupEndonly guarantees enqueueing, not completion, so synchronization is necessary. UsingcudaStreamSynchronizeis more precise, because it only waits for the relevant stream and will not mistakenly wait for unrelated operations. The problem with usingcudaDeviceSynchronizeis that if there are other unrelated long-running kernels on the device, they will be mistakenly waited on, reducing performance.

Q3: 📎 docs/examples/01_communicators/01_multiple_devices_single_process/c/main.cc:233-240The destruction order of is "first Finalize all communication domains, then Destroy all communication domains." If it were changed to "for each communication domain, first Finalize then Destroy" (that is, completing both operations in one loop), what problems would arise?

Reference analysis: It would break the Group semantics. The current form is:

c
ncclGroupStart();
for (i) ncclCommFinalize(comms[i]);
ncclGroupEnd();
for (i) ncclCommDestroy(comms[i]);

ncclCommFinalizeis wrapped by Group, meaning that the Finalize of all communication domains will be submitted together and can progress concurrently. If it were changed to:

c
for (i) {
    ncclCommFinalize(comms[i]);
    ncclCommDestroy(comms[i]);
}

The first iteration'sncclCommFinalize(comms[0])will block waiting for all ranks to become silent, but the Finalize of other communication domains has not yet been initiated, causing a deadlock - this is the same type of problem as the deadlock in Q1.

In addition, the header file📎 src/nccl.h.in:309-309states thatncclCommFinalizewhen returns, the communication domain may still be in thencclInProgressstate, and it needs to wait for global silence before enteringncclSuccess. IfncclCommDestroyis called immediately afterward, local resources may be released before the communication domain has fully become silent, leading to undefined behavior. The correct approach is to pollncclCommGetAsyncErrorafter Finalize to confirm the state, and then Destroy.

These external behaviors form the reference frame for all subsequent source code analysis. In Chapter 2, we will establish the core mental model: the five-piece set of communication domain, channel, algorithm, protocol, and transport layer, and see how NCCL internally organizes these concepts.

CHAPTER 02

Chapter 2: Chapter 2: Core Abstract Model: Communication Operators, Topology, Algorithms, Protocols, and Transport Layer

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 2 / 25

Chapter 2: Core Abstract Model: Communication Operators, Topology, Algorithms, Protocols, and Transport Layer

In the previous chapter, we got NCCL running and observed the external behavior of three APIs: ncclCommInitRank, ncclAllReduce, and ncclCommDestroy. But external behavior is only the tip of the iceberg - when ncclAllReduce returns, what exactly happened on the GPU? Which path did the data take? Why does the same AllReduce show huge performance differences on different machines? To answer these questions, we must first establish NCCL's common vocabulary. This chapter will break down five core abstractions one by one: communication domain (ncclComm), channel, algorithm, protocol, and transport layer. These five concepts run through the entire book, and every subsequent chapter's analysis will use them. Once you understand the relationships among them, you understand NCCL's skeleton.

2.1 Communication Domain ncclComm: A Process's Communication Context

Intuitive model

Think ofncclCommas a "group chat": after each process joins the group chat, it gets a group ID, and afterward all messages are sent in this group. How many people are in the group (nRanks), who I am (rank), which route to take (channels), and which rules to use (config) are all recorded in this group chat object.

WithoutncclComm, NCCL would not know "who communicates with whom" or "where the data is sent" - every API call would have to renegotiate the rank list and rebuild connections, and the overhead would be unbearable.

Data structure and memory layout

ncclCommis the most core struct in all of NCCL, defined insrc/include/comm.h. It is extremely large (nearly 300 lines), so let us look at the key fields grouped by function.

Identity markers and lifecycle sentinels

📎 src/include/comm.h:576-580definesstartMagic,📎 src/include/comm.h:879-881definesendMagic. These two fields are not security keys, but memory out-of-bounds detection sentinels. At📎 src/include/comm.h:883-885there are twostatic_assert:

c
static_assert(offsetof(struct ncclComm, startMagic) == 0, "startMagic must be the first field of ncclComm");
static_assert(offsetof(struct ncclComm, endMagic) == sizeof(struct ncclComm) - sizeof(uint64_t),
              "endMagic must be the last field of ncclComm");
{Design inference and architectural trade-offs}

These two assertions enforce at compile time thatstartMagicis located at the first address of the struct andendMagicis located at the end. At runtime, by checking whether these two magic numbers have been tampered with, one can quickly determine whether thencclCommpointer is valid - this is very useful when troubleshooting bugs such as "wild pointer accessing a destroyed communication domain" in a multithreaded environment.

Rank and topology information

📎 src/include/comm.h:628-629definesrankandnRanks- my number in the communication domain and the total number of participants.📎 src/include/comm.h:644-652defines node-related fields:node(the node number where I am located),nNodes(total number of nodes),localRank(number within the node),localRanks(number of GPUs within the node), and three mapping tablesrankToNode、rankToLocalRank、localRankToRank。

{Design inference and architectural trade-offs}

These three mapping tables are the foundation of topology-aware algorithms. For example, the Ring algorithm needs to know "whether my next rank is within the same node" to decide whether to use NVLink or the network. Without these mapping tables, every algorithm selection would require re-querying the topology graph, resulting in enormous overhead.

Channels and Buffers

📎 src/include/comm.h:593-593defineschannels[MAXCHANNELS]—this is the array of all channels within the communicator.📎 src/include/comm.h:674-676defines the number of channels:nChannels(number of connection channels),collChannels(number of collective communication enqueue channels),nvlsChannels(number of NVLS channels).

📎 src/include/comm.h:691-693defines buffer sizes:buffSizes[NCCL_NUM_PROTOCOLS](buffer size for each protocol),p2pChunkSize(P2P chunk size),nvlsChunkSize(NVLS chunk size).

[Design Inference and Architectural Trade-offs]

buffSizesThe index of the array is the protocol enum value (LL/LL128/Simple), which means each protocol has its own independent buffer size configuration. The LL protocol needs small buffers to reduce latency, while the Simple protocol needs large buffers to improve bandwidth—this array allows both requirements to coexist.

Work Queue and FIFO

📎 src/include/comm.h:719-728defines work FIFO related fields:workFifoBytes(FIFO size, power of 2),workFifoBuf(host-side FIFO buffer),workFifoBufDev(device-side FIFO buffer),workFifoProduced(bytes produced),workFifoConsumed(bytes consumed).

[Design Inference and Architectural Trade-offs]

This is a typical producer-consumer ring buffer. The host side (producer) writes work descriptors into the FIFO, and the GPU kernel (consumer) reads and executes them.workFifoBytesmust be a power of 2, so that bitmasking can replace modulo operations, accelerating index computation.

Intra-process Synchronization Barrier

📎 src/include/comm.h:731-731defines the intra-process multi-communicator synchronization mechanism:

c
struct ncclComm* intraComm0; // leader of intra-process comms (self possible)
struct ncclComm* intraNext; // next of intra-process comms, intraComm0 is head
int intraRank;
int intraRanks;
uint32_t intraBarrierPhase;
char intraPad1[64 - sizeof(uint64_t)];
uint64_t intraBarrierCounter; // only used if this is intraComm0
char intraPad2[64 - sizeof(uint64_t)];
uint64_t intraBarrierGate; // only used if this is intraComm0

NoteintraPad1andintraPad2have a size of64 - sizeof(uint64_t), which is 56 bytes. Adding the precedinguint64_tfield, each field group occupies exactly 64 bytes—this is one cache line.

[Design Inference and Architectural Trade-offs]

This is a typicalcache line paddingtechnique.intraBarrierCounterandintraBarrierGateare frequently read and written by multiple threads. If they share the same cache line, it causesfalse sharing: one thread modifyingintraBarrierCounterinvalidates another thread'sintraBarrierGatecache, causing a sharp performance degradation. Using 56 bytes of padding to separate them into different cache lines is a standard technique in high-performance concurrent programming.

Asynchronous Error State

📎 src/include/comm.h:705-705definesasyncResult—this field records the asynchronous operation state of the communicator. In the previous chapter, we mentioned that whenncclCommFinalizereturns, the communicator may still be in thencclInProgressstate, which is tracked through this field.

Scenario-Driven Walkthrough: From ncclCommInitRank to Struct Population

When the user callsncclCommInitRank(&comm, nranks, commId, rank), NCCL internally allocates ancclCommstruct and populates it field by field. Let us follow this process to see how key fields are set:

Step 1: Allocation and Zeroing

NCCL usesncclCallocto allocatencclComm, ensuring all fields are initialized to 0. At this point,startMagicandendMagicare set toNCCL_MAGIC(📎 src/include/comm.h:563-569defined as0x0280028002800280, with the comment saying "Nickel atomic number is 28").

Step 2: Populating Identity Information

rank、nRanks、cudaDevobtained from parameters and CUDA APIs.commHashis derived by hashingncclCommId, used for consistency verification in subsequent network communication.

Step 3: Building the Topology Graph

NCCL calls the topology detection module to enumerate all GPUs, NICs, and PCI switches, building thetopofield (📎 src/include/comm.h:595-595). This topology graph determines subsequent algorithm selection and path planning.

Step 4: Initializing Channels

channels[MAXCHANNELS]Theidarray is initialized one by one. Each channel'speersis set to the array index,devPeersand

pointers are allocated.

Step 5: Establishing Transport ConnectionssetupBased on the topology graph, NCCL selects the transport layer (P2P/SHM/NET) for each pair of ranks, calling the correspondingconnectandchannels[i].peers[j]callbacks. Connection information is stored in

.

Step 6: Setting the Magic NumberendMagicFinally,NCCL_MAGICis set to

, marking the struct initialization as complete.

Design Reflections and Production PitfallsncclCommWhy is

so large?

ncclComm[Design Inference and Architectural Trade-offs]

contains nearly 300 fields because it carries the entire state of a communicator. NCCL's design philosophy is "initialize once, reuse many times"—during initialization, all potentially useful information is computed and stored, and at runtime, tables are looked up directly to avoid redundant computation. The cost is higher memory usage (a few KB per communicator), but compared to GPU memory and network bandwidth, this memory is negligible.

Pitfall Scenario 1: Multi-threaded Sharing of a Communicator

ncclComm[Design Inference and Architectural Trade-offs]ncclCommis not thread-safe. If two threads simultaneously callncclAllReduce,workFifoProducedon the same

, fields such as

ncclCommDestroywill race, causing data corruption. The correct approach is for each thread to use an independent communicator, or to serialize calls with an external lock.startMagicPitfall Scenario 2: Access After DestructionendMagicAfter

frees the struct memory, if a thread still holds a pointer and accesses it, it will read freed memory.

andintraBarrierCountercan help detect this situation—if the magic number does not match, the pointer is invalid.intraBarrierGatePitfall Scenario 3: Cache Line False Sharing

In multi-process scenarios (one rank per process),

and

padding is particularly important. If padding is omitted, barrier operations across multiple processes will interfere with each other, causing synchronization latency to rise from nanoseconds to microseconds.channelIt is NCCL's "conveyor belt" — splitting the data of a single collective communication into multiple parts, with each channel independently carrying one part, advancing in parallel to improve bandwidth utilization.

Without channels, all data can only travel along a single path, and the multiple physical links between GPUs (multiple NICs, multiple NVLink groups) cannot be utilized simultaneously, causing bandwidth utilization to drop significantly.

Data Structures and Memory Layout

ncclChannelDefined in📎 src/include/comm.h:169-191:

c
struct ncclChannel {
  struct ncclChannelPeer** peers;
  struct ncclDevChannelPeer** devPeers;
  /* devPeer pointer array used for host side access */
  struct ncclDevChannelPeer** devPeersHostPtr;
  struct ncclRing ring;
  int* devRingUserRanks;
  struct ncclTree tree;

  struct ncclTree collnetChain;
  struct ncclDirect collnetDirect;

  struct ncclNvls nvls;

  int id; // index of this channel
  uint32_t workFifoProduced; // +1 successor of last used work fifo byte

  /* comm split sharable resources */
  struct ncclChannelPeer* collnetPeers;
  struct ncclDevChannelPeer* collnetDevPeers;
  struct ncclChannelPeer* nvlsPeers;
  struct ncclDevChannelPeer* nvlsDevPeers;
};

Key Field Analysis

  • peers / devPeers: Points to the connection information of all ranks within that channel.peersis the host-side view,devPeersis the device-side view (directly accessed by GPU kernel).
  • ring: Topology description for the Ring algorithm — the predecessor and successor of each rank.
  • tree: Topology description for the Tree algorithm — parent node and child node list.
  • collnetChain / collnetDirect: Two variant topologies for the CollNet algorithm.
  • nvls: Topology description for NVLink SHARP.
  • id: Channel index, from 0 tonChannels-1。
  • workFifoProduced: The work FIFO production pointer for that channel.
[Design Inference and Architectural Trade-offs]

Note thatring、tree、collnetChain、collnetDirect、nvlsthese five fields areparallel— the same channel can simultaneously hold topology descriptions for multiple algorithms. At runtime, the algorithm selection determines which field to use. This design allows algorithm switching without rebuilding channels — only the field being read needs to be switched.

Channel Count Calculation

The channel count is defined inncclComm(in📎 src/include/comm.h:674-676):

c
int nChannels; // connection nChannels
int collChannels; // enqueue nChannels
int nvlsChannels; // enqueue nChannels
[Design Inference and Architectural Trade-offs]

nChannelsis the actual number of connections established,collChannelsis the number of channels used when enqueuing collective communication,nvlsChannelsis the number of NVLS-dedicated channels. These three may differ — for example, some channels are used only for P2P and not for collective communication.

P2P Channel Scheduling

📎 src/include/channel.h:21-33defines thencclP2pChannelBaseForRoundfunction, used to calculate the channel base address used by each round in P2P communication:

c
inline uint8_t ncclP2pChannelBaseForRound(struct ncclComm* comm, int p2pRound) {
  int base;
  if (comm->nNodes > 1) {
    int localSize = comm->p2pSchedGroupSize;
    int groupDelta = p2pRound / localSize;
    int localDelta = p2pRound % localSize;
    base = groupDelta * divUp(localSize, NCCL_MAX_DEV_WORK_P2P_PER_BATCH);
    base += localDelta / NCCL_MAX_DEV_WORK_P2P_PER_BATCH;
  } else {
    base = p2pRound;
  }
  return reverseBits(base, log2Up(comm->p2pnChannels));
}
[Design Inference and Architectural Trade-offs]

The logic of this function is: in multi-node scenarios, P2P communication is scheduled by "groups," with ranks within each group using adjacent channels; in single-node scenarios, each round maps directly to one channel.reverseBitsis a bit-reversal operation, used to scatter channel assignments and avoid hotspot concentration.

Scenario-Driven Walkthrough: How an AllReduce Allocates Channels

Assume 8 ranks and 4 channels, executing one AllReduce. The data is split into 4 parts, each handled by one channel.

Step 1: Algorithm Selection

NCCL's tuning module selects the algorithm (e.g., Ring) and protocol (e.g., Simple) based on message size and topology.

Step 2: Channel Allocation

ncclTaskCollThe struct (📎 src/include/comm.h:212-273) is created, where thenChannelsfield is set to 4 (📎 src/include/comm.h:254-254)。channelLoandchannelHifields (📎 src/include/comm.h:256-257) mark the channel range used by that task.

Step 3: Data Splitting

Each channel is responsible forcount / nChannelselements. Channel 0 handles elements 0 to count/4-1, channel 1 handles elements count/4 to count/2-1, and so on.

Step 4: Parallel Execution

The GPU kernels of the 4 channels are launched simultaneously, each executing Ring AllReduce on its own data slice. Since there is no data dependency between channels, they can run fully in parallel.

Step 5: Result Merging

After all channels complete, the recv buffer of each rank contains the complete AllReduce result.

Concurrency Control and Hardware Interaction

Mapping of Channels to GPU Resources

[Design Inference and Architectural Trade-offs]

Each channel is typically bound to an independent CUDA stream or GPU hardware queue. This allows kernels of different channels to execute concurrently on the GPU, fully utilizing SM (Streaming Multiprocessor) resources.

Mapping of Channels to Network Devices

In multi-NIC scenarios, different channels can be bound to different NICs. For example, with 4 channels and 2 NICs, channels 0 and 1 go through NIC A, and channels 2 and 3 go through NIC B. This way, the bandwidth of both NICs can be utilized.

Choosing the Number of Channels

[Design Inference and Architectural Trade-offs]

More channels is not always better. Increasing the number of channels brings:

  • More kernel launch overhead
  • More connection establishment overhead
  • More complex synchronization

NCCL's tuning module automatically selects the optimal number of channels based on message size. Small messages use few channels (reducing overhead), while large messages use multiple channels (improving bandwidth).

Production Pitfall Guide

Pitfall Scenario 1: Improper Channel Count Configuration

[Design Inference and Architectural Trade-offs]

If manually settingNCCL_NCHANNELStoo large, the kernel launch overhead in small message scenarios will exceed the benefit, causing performance to degrade instead. It is recommended to let NCCL choose automatically, unless there is a clear tuning requirement.

Pitfall Scenario 2: Channel-Topology Mismatch

[Design Inference and Architectural Trade-offs]

If the number of channels exceeds the number of physical links, some channels will share links and cannot achieve true parallelism. For example, with 2 NICs and 8 channels, only 2 channels can actually transmit simultaneously, while the other 6 are queued.

Pitfall Scenario 3: P2P Channel Conflict

ncclP2pChannelBaseForRoundIf thereverseBitsoperation is implemented incorrectly, multiple rounds may map to the same channel, causing serialization.📎 src/include/channel.h:32-32ThereverseBits(base, log2Up(comm->p2pnChannels))ensures even channel distribution.

2.3 Algorithm: Topology Organization of Tree/Ring/CollNet/NVLS/PAT

Intuitive Model

From Beijing to Shanghai, you can take the high-speed rail, fly, or drive yourself, and each mode suits different distances and group sizes. NCCL's algorithms are these "travel modes"—Ring is suited for stable bandwidth with large messages, Tree is suited for low latency with small messages, CollNet leverages NIC offloading, NVLS leverages NVLink SHARP hardware acceleration, and PAT is a parallelized variant of NVLS.

Without algorithm selection, NCCL could only communicate in one fixed mode, unable to adapt to different message sizes and topologies, and performance would suffer greatly.

Data Structures and Memory Layout

Ring Algorithm

The core of the Ring algorithm is thencclRingstruct (insrc/include/comm.hreferenced viachannels[i].ring).📎 src/include/collectives.h:81-116defines theRingAlgorithmbase class:

c
class RingAlgorithm {
protected:
  int refCount;
  int nRanks;
  int nStepsPerLoop;
  int chunkSteps;
  int sliceSteps;
  ssize_t sliceSize;
  ssize_t loopSize;
  ssize_t channelSize;
  uint8_t* sendbuff;
  uint8_t* recvbuff;
  void* sendMhandle;
  void* recvMhandle;
  void* srecvMhandle;

public:
  virtual void getNextSendAddr(int curStep, uint8_t** sendbuffOut, size_t* sizeOut, void** mhandleOut) = 0;
  virtual void getNextRecvAddr(int curStep, uint8_t** recvbuffOut, size_t* sizeOut, void** mhandleOut) = 0;
  int incRefCount() {
    return (int)COMPILER_ATOMIC_ADD_FETCH(&refCount, 1, std::memory_order_relaxed);
  }
  int decRefCount() {
    return (int)COMPILER_ATOMIC_SUB_FETCH(&refCount, 1, std::memory_order_release);
  }
  RingAlgorithm() {
    refCount = 0;
  }
  virtual ~RingAlgorithm() {};
};

Key Field Analysis

  • refCount: reference count, used for sharing algorithm objects between proxy threads and GPU kernels.
  • nRanks: number of nodes in the ring.
  • nStepsPerLoop: number of steps per loop iteration. AllReduce is2*(nRanks-1)*chunkSteps(📎 src/include/collectives.h:218-218)。
  • chunkSteps / sliceSteps: chunk steps and slice steps, controlling pipeline granularity.
  • sliceSize / loopSize / channelSize: slice size, loop size, channel size.
  • sendbuff / recvbuff: send and receive buffer pointers.
  • sendMhandle / recvMhandle / srecvMhandle: memory handle, used for network registration.

Atomic Operations for Reference Counting

📎 src/include/collectives.h:106-108demonstratesincRefCountanddecRefCount:

c
int incRefCount() {
  return (int)COMPILER_ATOMIC_ADD_FETCH(&refCount, 1, std::memory_order_relaxed);
}
int decRefCount() {
  return (int)COMPILER_ATOMIC_SUB_FETCH(&refCount, 1, std::memory_order_release);
}
[Design Inference and Architectural Trade-offs]

incRefCountusesmemory_order_relaxed—incrementing the reference count does not require synchronization, only atomicity needs to be guaranteed.decRefCountusesmemory_order_release—when decrementing the reference count, it is necessary to ensure that prior writes are visible to other threads (because it may trigger object destruction).

RingARAlgorithm: Ring Implementation of AllReduce

📎 src/include/collectives.h:118-234definesRingARAlgorithm, inheriting fromRingAlgorithm. The core methods aregetNextSendAddrandgetNextRecvAddr。

📎 src/include/collectives.h:126-167'sgetNextSendAddrlogic:

c
void getNextSendAddr(int curStep, uint8_t** sendbuffOut, size_t* sizeOut, void** mhandleOut) {
  int curLoop = curStep / nStepsPerLoop;
  int curLoopStage = (curStep % nStepsPerLoop) / chunkSteps;
  int chunkStage = curLoopStage % nRanks;
  int sliceStage = (curStep % chunkSteps) / sliceSteps;
  ssize_t elemOffset = curLoop * loopSize;
  ssize_t remSize = channelSize - elemOffset;
  // ... 计算 chunkOffset, sliceOffset, curSliceSize ...
  if (remSize < loopSize) {
    curChunkSize = alignUp(divUp(remSize / elemSize, nRanks), 16 / elemSize) * elemSize;
  } else {
    curChunkSize = chunkSize;
  }
  chunkId = (ringIndex + nRanks - 1 - chunkStage) % nRanks;
  chunkOffset = chunkId * curChunkSize;
  nelem = std::min(remSize - chunkOffset, curChunkSize);
  curSliceSize = std::max(divUp(nelem / elemSize, 16 * slicePerChunk) * 16, sliceSize / elemSize / 32) * elemSize;
  sliceOffset = sliceStage * curSliceSize;
  // ... 设置 sendbuffOut, sizeOut, mhandleOut ...
}
[Design Inference and Architectural Trade-offs]

The core of this code isaddress calculation: given the current stepcurStep, calculate which slice of which data chunk should be sent.chunkId's calculation(ringIndex + nRanks - 1 - chunkStage) % nRanksimplements backpropagation along the ring—each rank receives data from its predecessor, processes it, and sends it to its successor.

PAT Algorithm

PAT (Parallel Aggregated Tree) is a parallelized variant of NVLS.📎 src/include/collectives.h:416-423definesncclPatStep:

c
struct ncclPatStep {
  int recvDim, sendDim, recvOffset, sendOffset, stepOffset, postRecv, postSend, nelem, last, flags;
  // PAT algo computation thread step number; -1 while the slot is free.
  int step;
  // This PAT group's offset within the shared NVLS slot.
  int nvlsOffset;
  size_t inpIx, outIx;
};

📎 src/include/collectives.h:425-435definesncclPatPeer:

c
struct ncclPatPeer {
  uint64_t step;
  struct ncclConnInfo* conn;
  struct ncclConnFifo* connFifo;
  void* buff;
  uint64_t* headPtr;
  uint64_t* tailPtr;
  uint64_t stepCache;
  long long int accSize;
  int connStepSize;
};
[Design Inference and Architectural Trade-offs]

The core idea of the PAT algorithm is toaggregate multiple small steps into one large step, reducing synchronization overhead.ncclPatStepdescribes the send/receive dimensions, offsets, element counts, and other information of an aggregation step.ncclPatPeerdescribes the connection state and buffer pointers of a peer node.

Scenario-Driven Walkthrough: Step Evolution of Ring AllReduce

Assume 4 ranks (0, 1, 2, 3), each with 4 elements, executing Ring AllReduce.

Reduce-Scatter Phase

  • Step 0: rank 0 sends element 0 to rank 1, rank 1 sends element 1 to rank 2, rank 2 sends element 2 to rank 3, rank 3 sends element 3 to rank 0.
  • Step 1: each rank adds the received element to the corresponding local element, then sends it to the next rank.
  • Step 2: continue accumulating and passing.
  • Step 3: at this point each rank has a complete reduction result (rank 0 has the result for element 3, rank 1 has the result for element 0, etc.).

AllGather Phase

  • Steps 4-6: each rank propagates the reduction result it holds along the ring, and finally all ranks have the complete result.

📎 src/include/collectives.h:218-218'snStepsPerLoop = 2 * (nRanks - 1) * chunkStepsexactly corresponds to this flow: Reduce-Scatter requires(nRanks-1)*chunkStepssteps, AllGather also requires(nRanks-1)*chunkStepssteps, for a total of2*(nRanks-1)*chunkStepssteps.

Design Reflections and Production Pitfalls

Why do Ring and Tree coexist?

[Design Inference and Architectural Trade-offs]

The Ring algorithm has high bandwidth utilization (every link is transmitting), but latency grows linearly with the number of ranks. The Tree algorithm has logarithmic latency, but low bandwidth utilization (only some links are working). NCCL automatically selects based on message size: small messages use Tree (latency-sensitive), large messages use Ring (bandwidth-sensitive).

Pitfall Scenario 1: Wrong Algorithm Selection

[Design Inference and Architectural Trade-offs]

If Ring is manually forced for small messages, latency will increase significantly. It is recommended to let the tuning module select automatically, unless there is clear profiling data supporting manual intervention.

Pitfall Scenario 2: NVLS Hardware Not Supported

NVLS requires specific hardware support (NVLink SHARP). If the hardware does not support it but the code forces NVLS, it will fall back to Ring or Tree, but may be accompanied by performance jitter.📎 src/include/comm.h:755-755'snvlsSupportfield marks whether the hardware supports NVLS.

Pitfall Scenario 3: Aggregation Factor Configuration of the PAT Algorithm

The PAT algorithm'saggFactordetermines how many steps to aggregate.📎 src/include/collectives.h:537-560demonstratesaggFactor's calculation logic:

c
aggFactor = 1;
size_t channelSize = end - offset;
while (stepSize / (channelSize * sizeof(T) * aggFactor) >= 2 && aggFactor < nranks / 2) {
  aggFactor *= 2;
  aggDelta /= 2;
}
postFreq = aggFactor;
if (postFreq < parallelFactor) parallelFactor = postFreq;
int d = stepDepth;
while (d > 1 && aggFactor < nranks / 2) {
  d /= 2;
  aggFactor *= 2;
  aggDelta /= 2;
}
[Design Inference and Architectural Trade-offs]

aggFactorIf too small, synchronization overhead will be large; if too large, it will cause pipeline bubbles. NCCL automatically calculates the optimal value based onstepSize、channelSize、nranks.

2.4 Protocol: LL/LL128/Simple Three Data Movement Strategies

Intuitive Model

Sending a package can be done via "same-city instant delivery," "next-day delivery," or "standard courier," each with different speed and cost. NCCL's protocols are these "shipping methods" — LL (Low Latency) is suited for low-latency transmission of small messages, LL128 is suited for 128-byte aligned transmission of medium messages, and Simple is suited for high-bandwidth transmission of large messages.

Without protocol selection, NCCL could only use a single fixed strategy to move data, unable to balance between latency and bandwidth.

Data Structures and Memory Layout

Protocol Enum

📎 src/include/comm.h:55-57Defines protocol-related thread thresholds:

c
#define NCCL_LL_THREAD_THRESHOLD 8
#define NCCL_LL128_THREAD_THRESHOLD 8
#define NCCL_SIMPLE_THREAD_THRESHOLD 64
[Design Inference and Architectural Trade-offs]

These thresholds determine how many threads each protocol uses. LL and LL128 use 8 threads (low latency, few threads suffice), while Simple uses 64 threads (high bandwidth, requiring more threads for parallel data movement).

Protocol Buffers

📎 src/include/comm.h:691-691DefinesbuffSizes[NCCL_NUM_PROTOCOLS]— each protocol has independent buffer sizes.

Protocol-related FIFO Structures

📎 src/include/comm.h:59-83DefinesncclSendMemandncclRecvMem:

c
struct ncclSendMem {
  union {
    struct {
      uint64_t head;
      char pad1[CACHE_LINE_SIZE - sizeof(uint64_t)];
      void* ptrExchange;
      uint64_t redOpArgExchange[2];
      char pad2[CACHE_LINE_SIZE - sizeof(void*) - 2 * sizeof(uint64_t)];
      int offsFifo[NCCL_STEPS];
    };
    char pad3[MEM_ALIGN];
  };
};

struct ncclRecvMem {
  union {
    struct {
      uint64_t tail;
      char pad1[CACHE_LINE_SIZE - sizeof(uint64_t)];
      struct ncclConnFifo connFifo[NCCL_STEPS];
      int flush; // For GDRCopy-based flush
    };
    char pad4[MEM_ALIGN];
  };
};
[Design Inference and Architectural Trade-offs]

ncclSendMemandncclRecvMemare shared memory structures for sending and receiving.headandtailare the read and write pointers of the ring buffer,pad1ensuring they are on different cache lines.connFifoThe array stores connection information for each step (mode, offset, size, pointer), defined in📎 src/include/collectives.h:72-77:

c
struct ncclConnFifo {
  int mode;
  ssize_t offset;
  ssize_t size;
  void* ptr;
};

Protocol Selection Logic

[Design Inference and Architectural Trade-offs]

Protocol selection is handled by the tuning module, considering factors including:

  • Message size: small messages use LL, medium use LL128, large use Simple.
  • Topology: NVLink connections suit LL128, network connections suit Simple.
  • Hardware capabilities: certain GPU architectures have optimizations for specific protocols.

Scenario-Driven Walkthrough: LL Protocol Data Movement

Assume using the LL protocol to transmit 1KB of data.

Step 1: Data written to send buffer

The host side writes data tosendbuff, then updates thencclSendMem.headpointer, notifying the GPU kernel of new data.

Step 2: GPU kernel reads data

The GPU kernel polls theheadpointer, and upon detecting new data, reads fromsendbuff.

Step 3: Data transmission

The GPU kernel sends data to the target rank via NVLink or network.

Step 4: Target rank receives data

The target rank's GPU kernel writes data torecvbuff, then updates thencclRecvMem.tailpointer.

Step 5: Host side reads data

The host side polls thetailpointer, and upon detecting new data, reads fromrecvbuff.

Concurrency Control and Hardware Interaction

LL Protocol's Low-Latency Mechanism

[Design Inference and Architectural Trade-offs]

The LL protocol usesPollingrather than interrupts to detect data arrival. The GPU kernel continuously reads theheadpointer and processes immediately upon detecting a change. This has lower latency than interrupts but occupies GPU compute resources.

LL128 Protocol's 128-Byte Alignment

[Design Inference and Architectural Trade-offs]

The LL128 protocol requires data to be 128-byte aligned, so each transmission exactly fills one cache line. The benefits of alignment are:

  • Reduces partial cache line writes
  • Improves memory bandwidth utilization
  • Simplifies hardware processing logic

Simple Protocol's Batch Transmission

[Design Inference and Architectural Trade-offs]

The Simple protocol usesBatch transmissionmode: accumulating a certain amount of data before sending at once, reducing synchronization frequency. This suits large message scenarios because synchronization overhead is amortized over large amounts of data.

Production Pitfall Guide

Pitfall Scenario 1: Protocol and Message Size Mismatch

[Design Inference and Architectural Trade-offs]

If the LL protocol is forced to transmit large messages, performance drops sharply. This is because the LL protocol's design goal is low latency, not high bandwidth. Large messages should use the Simple protocol.

Pitfall Scenario 2: LL128 Alignment Issues

[Design Inference and Architectural Trade-offs]

If data is not 128-byte aligned, the LL128 protocol falls back to LL or Simple, causing unstable performance. It is recommended to ensure both send and receive buffers are 128-byte aligned.

Pitfall Scenario 3: Protocol Switching Overhead

[Design Inference and Architectural Trade-offs]

Dynamically switching protocols at runtime incurs additional overhead. NCCL determines the protocol at initialization and does not switch at runtime. If switching is needed, the communication domain must be reinitialized.

2.5 Transport Layer: P2P/SHM/NET/CollNet Underlying Data Movement Channels

Intuitive Model

Getting from point A to point B can be done by walking, cycling, taking the subway, or taking a taxi. NCCL's transport layer is these different "travel methods." The upper layers don't care how it gets there, only whether it can be delivered. P2P is "walking" (same-machine GPU direct connection), SHM is "cycling" (shared memory), NET is "taking the subway" (network), and CollNet is "taking a taxi" (NIC offload).

Without transport layer abstraction, upper-layer algorithms would need to write different code for each physical link, unable to reuse.

Data Structures and Memory Layout

Transport Layer Enum

📎 src/include/transport.h:18-23Defines transport layer types:

c
#define NTRANSPORTS 4
#define TRANSPORT_UNDEFINED -1
#define TRANSPORT_P2P 0
#define TRANSPORT_SHM 1
#define TRANSPORT_NET 2
#define TRANSPORT_COLLNET 3

Transport Layer Interface

📎 src/include/transport.h:129-146DefinesncclTransportComm— the transport layer's communication interface:

c
struct ncclTransportComm {
  ncclResult_t (*setup)(struct ncclComm* comm, struct ncclTopoGraph* graph, struct ncclPeerInfo*, struct ncclPeerInfo*,
                        struct ncclConnect*, struct ncclConnector*, int channelId, int connIndex);
  ncclResult_t (*connect)(struct ncclComm* comm, struct ncclConnect*, int nranks, int rank, struct ncclConnector*);
  ncclResult_t (*free)(struct ncclComm* comm, struct ncclConnector*);
  ncclResult_t (*proxySharedInit)(struct ncclProxyConnection* connection, struct ncclProxyState* proxyState,
                                  int nChannels);
  ncclResult_t (*proxySetup)(struct ncclProxyConnection* connection, struct ncclProxyState* proxyState, void* reqBuff,
                             int reqSize, void* respBuff, int respSize, int* done);
  ncclResult_t (*proxyConnect)(struct ncclProxyConnection* connection, struct ncclProxyState* proxyState, void* reqBuff,
                               int reqSize, void* respBuff, int respSize, int* done);
  ncclResult_t (*proxyFree)(struct ncclProxyConnection* connection, struct ncclProxyState* proxyState);
  ncclResult_t (*proxyProgress)(struct ncclProxyState* proxyState, struct ncclProxyArgs*);
  ncclResult_t (*proxyRegister)(struct ncclProxyConnection* connection, struct ncclProxyState* proxyState,
                                void* reqBuff, int reqSize, void* respBuff, int respSize, int* done);
  ncclResult_t (*proxyDeregister)(struct ncclProxyConnection* connection, struct ncclProxyState* proxyState,
                                  void* reqBuff, int reqSize, int* done);
};

Key Callback Analysis

  • setup: Preparation work before establishing a connection, exchanging connection parameters.
  • connect: Actually establishing the connection.
  • free: Releasing connection resources.
  • proxySharedInit: Initialize proxy thread shared resources.
  • proxySetup / proxyConnect: Connection establishment on the proxy thread side.
  • proxyProgress: Proxy thread advances data transfer.
  • proxyRegister / proxyDeregister: Memory registration and deregistration.

Transport layer struct

📎 src/include/transport.h:148-154definesncclTransport:

c
struct ncclTransport {
  const char name[8];
  ncclResult_t (*canConnect)(int*, struct ncclComm* comm, struct ncclTopoGraph* graph, struct ncclPeerInfo*,
                             struct ncclPeerInfo*);
  struct ncclTransportComm send;
  struct ncclTransportComm recv;
};
[Design Inference and Architectural Trade-offs]

nameis the transport layer name (e.g., "P2P", "SHM", "NET"),canConnectdetermines whether this transport layer can be used between two ranks,sendandrecvare the communication interfaces for send and receive directions respectively.

Transport layer instances

📎 src/include/transport.h:36-36declares four transport layer instances:

c
extern struct ncclTransport p2pTransport;
extern struct ncclTransport shmTransport;
extern struct ncclTransport netTransport;
extern struct ncclTransport collNetTransport;

📎 src/include/transport.h:36-36defines the transport layer array:

c
extern struct ncclTransport* ncclTransports[];

Peer node information

📎 src/include/transport.h:46-74definesncclPeerInfo— metadata exchanged between ranks:

c
struct ncclPeerInfo {
  int rank;
  int cudaDev;
  int nvmlDev;
  int gdrSupport;
  uint64_t hostHash;
  uint64_t pidHash;
  dev_t shmDev;
  int64_t busId;
  cudaUUID_t gpuUuid;
  struct ncclComm* comm;
  int cudaCompCap;
  int gpuCftSupport;
  size_t totalGlobalMem;
  // MNNVL support
  nvmlGpuFabricInfoV_t fabricInfo;
  int fabricHandleSupport;
  int cuMemSupport;
  int version;
  uint64_t supportedGinTypeBitMask;
  bool crossNicSupport;
  bool rmaPluginAvailable;
  bool cuMemGdrSupport;
  int mloPart; // MLOPart partition index, or -1 if not an MLOPart GPU
  int cudaDriverVersion;
  bool gpuCftMulticastSupport;
  bool gpuCftCountedSupport;
  uint32_t gitVersionHash;
};
[Design Inference and Architectural Trade-offs]

These fields are used to determine which transport layer can be used between two ranks:

  • hostHashSame → same host → P2P or SHM available
  • hostHashDifferent → different hosts → must use NET
  • gdrSupport→ whether GPUDirect RDMA is supported
  • cudaCompCap→ GPU compute capability, affects protocol selection

Scenario-Driven Walkthrough: Establishing a P2P Connection

Assume two ranks are on the same host, and NCCL selects the P2P transport layer.

Step 1: Exchange PeerInfo

The two ranks exchangencclPeerInfothrough the bootstrap channel, confirming they are on the same host and the GPUs support P2P.

Step 2: Call canConnect

📎 src/include/transport.h:148-154'scanConnectcallback is invoked, checking the topology graph to confirm there is an NVLink or PCIe connection between the two GPUs.

Step 3: Call setup

p2pTransport.send.setupandp2pTransport.recv.setupare invoked, preparing connection parameters (such as IPC handles).

Step 4: Call connect

p2pTransport.send.connectandp2pTransport.recv.connectare invoked, actually establishing the connection.

Step 5: Register memory

If RDMA is needed, callproxyRegisterto register send and receive buffers.

Concurrency Control and Hardware Interaction

P2P Transport Layer

[Design Inference and Architectural Trade-offs]

P2P uses the CUDA IPC (Inter-Process Communication) mechanism, allowing one GPU to directly access another GPU's memory. This requires:

  • Both GPUs in the same PCIe domain or NVLink domain
  • OS support for CUDA IPC
  • Sufficient permissions

SHM Transport Layer

[Design Inference and Architectural Trade-offs]

SHM uses host shared memory as an intermediary. When there is no direct connection between two GPUs, data is first copied to host memory, then copied to the target GPU. This is slower than P2P but has better compatibility.

NET Transport Layer

[Design Inference and Architectural Trade-offs]

NET uses network devices (InfiniBand or RoCE) to transfer data. This requires:

  • Network devices support GPUDirect RDMA (optional, but recommended)
  • Correct network configuration (IP address, subnet mask, etc.)
  • Sufficient network bandwidth

CollNet Transport Layer

[Design Inference and Architectural Trade-offs]

CollNet leverages the collective communication offload capability of network cards (such as NVIDIA SHARP). The network card directly performs reduction operations, reducing the GPU's computational burden. This requires:

  • Network cards that support SHARP
  • Correct SHARP configuration

Production Pitfall Guide

Pitfall Scenario 1: P2P Unavailable

[Design Inference and Architectural Trade-offs]

If there is no NVLink between two GPUs and the PCIe topology does not support P2P, NCCL falls back to SHM. This causes performance degradation. You can useNCCL_P2P_DISABLE=1to force-disable P2P and observe performance changes.

Pitfall Scenario 2: Network Configuration Error

[Design Inference and Architectural Trade-offs]

If the network device's IP address is misconfigured, the NET transport layer cannot establish a connection. Common errors include: wrong subnet mask, missing routing table entries, firewall blocking. It is recommended to useibstatandibpingto check the InfiniBand connection.

Pitfall Scenario 3: GPUDirect RDMA Not Enabled

[Design Inference and Architectural Trade-offs]

IfgdrSupportis 0, the NET transport layer falls back to the "copy to host memory first, then send" mode, significantly increasing latency. Check whether thenvidia-peermemmodule is loaded, and whether the network card driver supports GPUDirect.

2.6 How the Five Components Combine: The Complete Lifecycle of a Single Communication

Combination Relationship Diagram

mermaid
flowchart TD
    api["ncclAllReduce(sendbuff, recvbuff, count, ...)"] --> comm_lookup["查找 ncclComm"]
    comm_lookup --> task_create["创建 ncclTaskColl"]
    task_create --> tuning{"tuning 模块选择算法和协议"}
    tuning -->|"小消息"| tree_ll["Tree + LL"]
    tuning -->|"中等消息"| ring_ll128["Ring + LL128"]
    tuning -->|"大消息"| ring_simple["Ring + Simple"]
    tuning -->|"NVLS 可用"| nvls["NVLS + Simple"]
    tree_ll --> channel_assign["分配通道"]
    ring_ll128 --> channel_assign
    ring_simple --> channel_assign
    nvls --> channel_assign
    channel_assign --> transport_select{"选择传输层"}
    transport_select -->|"同机 GPU 直连"| p2p["P2P"]
    transport_select -->|"同机无直连"| shm["SHM"]
    transport_select -->|"跨机"| net["NET"]
    transport_select -->|"CollNet 可用"| collnet["CollNet"]
    p2p --> kernel_launch["启动 GPU kernel"]
    shm --> kernel_launch
    net --> kernel_launch
    collnet --> kernel_launch
    kernel_launch --> execute["执行通信"]
    execute --> complete["完成,更新 asyncResult"]

Complete Lifecycle

Phase 1: API Call

The user callsncclAllReduce, passing in the send buffer, receive buffer, element count, data type, reduction operation, communication domain, and CUDA stream.

Phase 2: Task Creation

NCCL creates thencclTaskCollstruct (📎 src/include/comm.h:212-273), filling in fields such asfunc(AllReduce)、sendbuff、recvbuff、count、datatype、opHost.

Phase 3: Algorithm and Protocol Selection

The Tuning module selects the algorithm (Ring/Tree/NVLS) and protocol (LL/LL128/Simple) based on message size, topology, and hardware capabilities. The selection results are written intoncclTaskColl'salgorithmandprotocolfields (📎 src/include/comm.h:227-227)。

Phase 4: Channel Allocation

Based on the algorithm and protocol, determine the number of channels and channel range to use.nChannels、channelLo、channelHiThe field is set (📎 src/include/comm.h:254-257)。

Phase 5: Transport Layer Selection

Based on the topology graph, select the transport layer (P2P/SHM/NET/CollNet) for each pair of ranks. Connection information is stored inchannels[i].peers[j].

Phase 6: Kernel Launch

NCCL buildsncclKernelPlan(📎 src/include/comm.h:357-410), containing work queues, cleanup queues, task queues, etc. Then launches the GPU kernel.

Phase 7: Execute Communication

The GPU kernel reads the work FIFO and performs data transfer and reduction operations. Proxy threads asynchronously advance network I/O.

Phase 8: Completion

After all channels complete,asyncResultis set toncclSuccess. Users can query the status viancclCommGetAsyncError.

Design Reflections

Why are the five components needed?

[Design Inference and Architectural Trade-offs]

These five abstractions each address problems in different dimensions:

  • ncclComm: Solves the "who communicates with whom" problem.
  • channel: Solves the "how to parallelize" problem.
  • algorithm: Solves the "what topology to use" problem.
  • protocol: Solves the "what strategy to use" problem.
  • transport: Solves the "what physical link to use" problem.

They combine orthogonally, allowing NCCL to adapt to various hardware configurations and message sizes without writing specialized code for each combination.

Flexibility of Combination

[Design Inference and Architectural Trade-offs]

The number of combinations for the five components is:

  • Algorithms: 5 types (Tree/Ring/CollNet/NVLS/PAT)
  • Protocols: 3 types (LL/LL128/Simple)
  • Transport layers: 4 types (P2P/SHM/NET/CollNet)

Chapter Reflections and Self-Assessment

Q1: If the📎 src/include/comm.h:731-731inintraPad1[64 - sizeof(uint64_t)]is changed tointraPad1[0](i.e., removing the cache line padding), what performance issues would arise in multi-process scenarios? Why?

Reference Analysis:

After removing the padding,intraBarrierPhase、intraBarrierCounter、intraBarrierGatethe three fields would be tightly packed in memory, likely sharing the same cache line (typically 64 bytes).

In multi-process scenarios, each process has its own copy ofncclComm, but theintraComm0andintraBarrierCounterof the leader communicator pointed to byintraBarrierGateare read and written by all processes. When process A callsncclCommIntraBarrierInto updateintraBarrierCounter(📎 src/include/comm.h:943-959), it causes process B'sintraBarrierGatecache line to be invalidated. When process B pollsncclCommIntraBarrierOutinintraBarrierGate(📎 src/include/comm.h:962-977), each cache invalidation requires reloading from memory, with latency rising from nanoseconds to microseconds.

This is theFalse Sharingproblem. Padding 56 bytes ensures each field occupies its own cache line, eliminating false sharing.

Q2: If the📎 src/include/collectives.h:106-108ofincRefCountis changed frommemory_order_relaxedtomemory_order_seq_cst, what impact would it have? Why did the author chooserelaxed?

Reference Analysis:

memory_order_seq_cstwould enforce global sequential consistency, requiring a memory barrier to be inserted on every reference count increment, causing performance degradation.

incRefCountonly needs to guarantee atomicity, without synchronizing other memory operations. Because incrementing the reference count does not trigger object destruction, nor does it depend on other threads' write operations.memory_order_relaxedexactly satisfies this requirement—guaranteeing only atomicity without inserting barriers.

In contrast,decRefCount(📎 src/include/collectives.h:109-111) usesmemory_order_release, because decrementing the reference count may trigger object destruction, requiring that prior write operations be visible to other threads.

This is a classic application of the C++ memory model: choosing the weakest memory order based on operation semantics, maximizing performance while ensuring correctness.

Q3: If the📎 src/include/channel.h:32-32ofreverseBits(base, log2Up(comm->p2pnChannels))is changed to directly returnbase % comm->p2pnChannels, in what scenarios would performance degrade? Why?

Reference Analysis:

reverseBitsis a bit-reversal operation used to scatter channel assignments. Direct modulo would cause channel assignments to exhibit regularity: round 0 uses channel 0, round 1 uses channel 1, ..., round N uses channel N%p2pnChannels.

In multi-node scenarios, if multiple ranks' P2P communications occur simultaneously, regular channel assignments would cause hotspot concentration—certain channels used by multiple ranks simultaneously while others are idle. This causes link congestion and reduces overall bandwidth utilization.

reverseBitsscatters channel assignments, making different rounds use seemingly random channels and distributing load evenly. This is a classic technique forload balancing.

Additionally,reverseBitsis a pure bit operation, faster than modulo (modulo requires a division instruction, while bit operations need only a few instructions).

---

In the next chapter, we will dive deep into the internal implementation ofncclCommInitRank, seeing how NCCL starts from an emptyncclCommstruct, progressively builds the topology graph, initializes channels, establishes transport connections, and ultimately constructs a usable communicator. The mental model of the five components established in this chapter will be put into practice one by one in the next chapter.

These five abstractions do not exist in isolation: the communicator is the container, channels are the units of parallel execution, algorithms determine how data is reduced, protocols specify how data is encoded, and transport layers handle how data moves. Their combination—5 dimensions, each with 3 to 4 choices—constitutes the search space for NCCL performance tuning. So, how exactly is this communicator object built from scratch? In the next chapter, we will dive into the ncclCommInitRank call chain, examining how NCCL completes device probing, topology discovery, and channel allocation during initialization, and revealing the assignment timing of key fields such as comm->rank, comm->nRanks, and comm->channels.

CHAPTER 03

Chapter 3: Chapter 3: Initialization: How ncclCommInitRank Builds a Group of Isolated Processes into a Communicator

Official Source: NVIDIA/nccl · Version: Commit @12df1a11 · Book Progress: Chapter 3 / 25

Chapter 3: Initialization: How ncclCommInitRank Builds a Group of Isolated Processes into a Communicator

In the previous chapter, we established five core abstractions that run through the entire book: ncclComm, channel, algorithm, protocol, and transport. Together they form the common vocabulary of "one communication = several channels × one algorithm × one protocol × several transports." Now we need to answer a more fundamental question: how exactly is this ncclComm object constructed from nothing? When you call ncclCommInitRank, NCCL needs to complete a series of complex operations within a few hundred milliseconds: confirm that all ranks have arrived, exchange device information, probe machine topology, compute data paths, allocate GPU memory and host memory, and finally package all of this into an ncclComm object. This chapter will follow this call chain, drilling down from the API entry point all the way to the last capillary of initTransportsRank.

3.1 API Entry: The Synchronous Shell and Asynchronous Core of ncclCommInitRank

Intuitive Model

ncclCommInitRankOn the surface, it is "creating a communication domain," but in reality what it does is "launch a background task, then (by default) wait for it to complete." This is like ordering food at a restaurant: the act of ordering (the API call) returns instantly, but the kitchen preparing the food (the actual initialization) happens in the background. The default "blocking mode" simply makes you wait at the counter until the food is ready, while "non-blocking mode" gives you a pickup number so you can go do something else first.

Without this asynchronous design layer, NCCL would not be able to cooperate with scenarios such as CUDA Graph capture and parallel initialization of multiple communication domains during initialization—all initialization would become serialized blocking operations that cannot overlap with user code.

Data Structures and Memory Layout

Let us first look at the API entry point itself.ncclCommInitRankIt is an extremely thin synchronous shell:

📎 src/init.cc:2946-2970

It does four things: callncclInitEnv()load the environment variable plugin, turn on NVTX performance markers, read the current CUDA device number, and then callncclGroupStartInternal()enter group semantics, and finally delegate the actual work toncclCommInitRankDev。

NotencclGroupStartInternal() / ncclGroupEndInternal()This pair of calls—even if you are initializing only one communication domain, NCCL still wraps it in group semantics. This is to uniformly handle the scenario where "the user initializes multiple communication domains within one group," avoiding the need to write two code paths for single-domain and multi-domain cases.

The real parameter validation and object allocation happen inncclCommInitRankDevinside:

📎 src/init.cc:2851-2943

This function is the "central dispatch desk" of the entire chain. It first performs parameter validation (nIdrange,nranks/myrankvalidity), then allocatesncclCommthe structure itself, as well as three fields related to the abort mechanism:abortFlag(host-side atomic flag),abortFlagDev(device-visible pinned memory copy),abortFlagRefCount(reference count, because child communication domains created by split may share the parent communication domain's abortFlag).

There is a detail worth noting here—comm->startMagic = comm->endMagic = NCCL_MAGIC:

📎 src/init.cc:2886-2886

This pair of magic values acts like a "seal" clamped at the beginning and end of thencclCommstructure. Any out-of-bounds write or structure corruption will destroy this pair of magic values, and subsequent operations can detect memory trampling by validating them. This is a cheap but effective memory integrity protection.

Step-by-Step Walkthrough

WhenncclCommInitRankDevreaches the end, it constructs ancclCommInitRankAsyncJoband starts an asynchronous task:

📎 src/init.cc:2896-2929

jobThe structure carries all the parameters needed for initialization. Note thatjob->commIdiscopiedout, rather than directly referencing the user-passedcommId:

📎 src/init.cc:2903-2910

Why copy? The source code comments give the answer:ncclUniqueIdandncclBootstrapHandlehave different alignment requirements, and the array passed in by the user may not be correctly aligned to the boundary required byncclBootstrapHandleCopying to newly allocated memory can guarantee alignment. This is a typical "ABI compatibility trap"—what the user sees isncclUniqueIdbut internally it must be used asncclBootstrapHandleThe two have the same size but different alignment.

Finally, according to the value ofncclParamEnqueueRearchEnable()the task either enters the management queue or is started directly throughncclAsyncLaunch:

📎 src/init.cc:2922-2929

ncclAsyncLaunchcreates a new thread to executencclCommInitRankFuncIf it is blocking mode (the default), the caller waits inncclGroupEndInternal()for this thread to complete; if it is non-blocking mode, the caller returns immediately, and the user later polls the status throughncclCommGetAsyncError.

Design Considerations

The core of the design here is "synchronous API + asynchronous implementation." Why not letncclCommInitRankdirectly execute all initialization synchronously? Because NCCL needs to supportncclCommInitRankConfignon-blocking mode, and non-blocking mode requires initialization to run in a background thread. If the synchronous path and the asynchronous path were two separate pieces of code, the maintenance cost would double. By uniformly going through the asynchronous path, the synchronous path is just "start and immediately wait," and there is only one copy of the code.

mermaid
flowchart TD
    api["ncclCommInitRank(newcomm, nranks, commId, myrank)"]
    env["ncclInitEnv() 加载环境变量插件"]
    group["ncclGroupStartInternal()"]
    dev["ncclCommInitRankDev(...)"]
    check{"nId/nranks/myrank 合法?"}
    alloc["ncclCalloc 分配 comm + abortFlag"]
    parse["parseCommConfig() 解析配置"]
    job["构造 ncclCommInitRankAsyncJob"]
    copyid["拷贝 commId 保证对齐"]
    enq{"ncclParamEnqueueRearchEnable()?"}
    mgmt["ncclMgmtTaskEnqueue()"]
    async["ncclAsyncLaunch() 启动后台线程"]
    func["ncclCommInitRankFunc() 执行初始化"]
    fail["返回 ncclInvalidArgument"]

    api --> env --> group --> dev --> check
    check -->|否| fail
    check -->|是| alloc --> parse --> job --> copyid --> enq
    enq -->|是| mgmt --> func
    enq -->|否| async --> func

3.2 Bootstrap: The First Control Channel Between Ranks

Intuitive Model

Bootstrap is NCCL's "pre-meeting WeChat group." Before formal communication begins, all ranks need to first establish a control channel to exchange metadata such as "who I am, which machine I am on, what model my GPU is, and what my NIC address is." Without bootstrap, the ranks are just a group of strangers who do not know each other and cannot coordinate any communication.

If bootstrap fails or times out, the entire communication domain initialization will hang—this is one of the most common causes of NCCL hangs in production environments.

Data Structures and Memory Layout

The core state of Bootstrap is stored in thebootstrapStatestructure:

📎 src/bootstrap.cc:527-546

Several key fields in this struct are worth elaborating on:

  • ring: a union, either a network device handle (net.sendComm/net.recvComm), or a pair of sockets (socket.send/socket.recv). This corresponds to two bootstrap modes: the default socket-based mode and the network-device-basedNCCL_OOB_NET_ENABLEmode.
  • listen: listener-side information, which likewise has two forms: network and socket.
  • peerP2pAddresses / peerProxyAddresses: arrays of P2P addresses and proxy addresses for all ranks, populated via ring allgather.
  • unexpectedConnections: a linked list that caches connections that have been "received but not yet matched." This is a key design of the bootstrap protocol—because the receiver cannot predict who will connect first, unmatched connections must be stored first.
  • asyncSendQueue + asyncSendLock + asyncSendCond: the asynchronous send queue and its synchronization primitives, used for concurrent sends in TLS encryption mode.

bootstrapStateThe allocation ofbootstrapInitoccurs at the beginning of

📎 src/bootstrap.cc:769-776

Note thecomm->bootstrap = stateline—the bootstrap state is attached to the communicator, and all subsequent bootstrap operations access it throughcomm->bootstrap.

Step-by-Step Walkthrough

bootstrapInitis the main function of bootstrap. Let's break it down in execution order:

Step 1: Determine the magic value.magic is the "secret code" for bootstrap communication; only ranks holding the same magic can connect to each other.

📎 src/bootstrap.cc:778-788

If it is normal initialization (handles != NULL), magic comes from the first handle; if it is split/grow (parent != NULL), magic is derived viahashCombine(parent->magic, parent->childCount). This ensures each sub-communicator has a unique magic.

Step 2: Create listening sockets.Each rank needs two listening endpoints: one for ring neighbor connections (STATE_LISTEN(state, socket)), and one for root connections (listenSockRoot):

📎 src/bootstrap.cc:797-831

There is a key division of labor here: the ring listening socket usescomm->magic, while the root listening socket usesBOOTSTRAP_HANDLE(handles, curr_root)->magic. Why? Because root is the global coordinator, and all ranks need to connect to it, so it uses a unified magic; whereas ring neighbors are point-to-point, so the communicator's own magic is sufficient.

Step 3: Staggered connections.When the number of ranks is very large, all ranks connecting to root simultaneously causes a connection storm. NCCL usesNCCL_UID_STAGGER_RATEandNCCL_UID_STAGGER_THRESHOLDto control staggering:

📎 src/bootstrap.cc:833-843

When the number of ranks handled by a root exceeds a threshold (default 256), each rank calculates a delay in microseconds based on its local ID under that root, and then sleeps. This is a simple but effective "token bucket"-style rate limiting.

Step 4: Send your own connection information to root.Each rank sends its listening address to root:

📎 src/bootstrap.cc:845-867

After root receives the information of all ranks, it performs a "ring pairing"—sending rank i's address to rank i-1, and rank i+1's address to rank i. In this way, each rank learns its predecessor and successor neighbors on the ring.

Step 5: Establish ring connections.Each rank connects to its "next" neighbor while accepting the connection from its "previous" neighbor:

📎 src/bootstrap.cc:885-894

HeresocketRingConnectinternally usesbootstrapConcurrent—in TLS encryption mode, connect and accept must be executed concurrently, otherwise it will deadlock (because the TLS handshake requires both parties to participate simultaneously). In non-encrypted mode, connect is executed serially followed by accept.

Step 6: AllGather all addresses.After the ring is established, perform an allgather of all ranks' P2P addresses, proxy addresses, and UDS addresses viaringAllInfo:

📎 src/bootstrap.cc:934-938

ringAllInfointernally callsbootstrapAllGather, which in socket mode usessocketRingAllGather—a bidirectional ring allgather algorithm, requiring only N/2 steps for N ranks:

📎 src/bootstrap.cc:1363-1412

This bidirectional algorithm is a key optimization for bootstrap performance. The traditional unidirectional ring allgather requires N-1 steps, while the bidirectional version halves the number of steps. Each step simultaneously sends and receives data in both directions, usingsocketDoubleSendRecvto package 4 operations (2 sends and 2 receives) into a single system call.

Concurrency control and low-level interaction

Bootstrap's concurrency control has several layers:

Layer 1: abort checking.All blocking loops periodically check abortFlag:

📎 src/bootstrap.cc:150-159

BOOTSTRAP_N_CHECK_ABORTSet to 10000, meaning the abort flag is checked once every 10000 loop iterations. This number is a tradeoff between performance and responsiveness—checking too frequently hurts performance, while checking too infrequently delays abort response.

Layer 2: asynchronous send queue.In TLS encryption mode,bootstrapSendcannot be executed synchronously (because the TLS handshake requires the receiver to participate as well), so NCCL places send operations on a separate thread:

📎 src/bootstrap.cc:1161-1217

There is an ingenious ordering guarantee mechanism here.bootstrapAsyncSendMainBefore sending, it checks whether there is an "earlier send to the same (peer, tag)" in the queue:

📎 src/bootstrap.cc:1124-1152

Why must the send order for the same (peer, tag) be guaranteed? The source code comments explain this clearly: the receiver matches connections by (peer, tag). If two messages destined for the same (peer, tag) arrive out of order, the receiver will match them incorrectly. During NVLS initialization, broadcasts are sent multiple times to the same peer with the same tag, so this ordering guarantee is essential.

Third layer: the unexpected connection queue.The receiver cannot predict who will connect first, sosocketAcceptit stores unmatched connections in aunexpectedConnectionslinked list:

📎 src/bootstrap.cc:1276-1300

This design solves a classic distributed problem: multiple ranks may initiate connections to you simultaneously, but yourbootstrapRecvcall order is fixed. If unmatched connections were simply dropped, the sender would time out; if it blocked and waited, it could deadlock. Storing them in a queue is the safest approach.

Production Pitfall Guide

Pitfall 1: bootstrap timeout causing initialization to hang.If a rank cannot connect to the root due to network issues, all other ranks will wait indefinitely onncclSocketAcceptorncclSocketRecv. NCCL has no built-in bootstrap timeout mechanism; the only escape route is abortFlag. In production environments, it is recommended to setNCCL_UID_STAGGER_RATEto mitigate connection storms in large-scale clusters.

Pitfall 2:NCCL_COMM_IDconflicts with multiple handles.When the user sets theNCCL_COMM_IDenvironment variable, NCCL forcibly downgradesnIdto 1:

📎 src/init.cc:2912-2921

This means thatncclCommInitRankScalable's multi-handle feature is silently disabled. If you are using scalable initialization and also setNCCL_COMM_ID, the behavior will differ from what you expect.

Pitfall 3: deadlock in TLS mode.In TLS encryption mode, if connect and accept are not executed concurrently, both sides will get stuck in the TLS handshake.bootstrapConcurrentThis is precisely to solve this problem:

📎 src/bootstrap.cc:648-669

In non-encrypted mode, execution is serial (send first, then recv); in encrypted mode, a thread is started to handle send, while the main thread handles recv.

mermaid
sequenceDiagram
    participant R0 as Rank 0
    participant Root as Bootstrap Root
    participant R1 as Rank 1
    participant R2 as Rank 2

    R0->>Root: sendToRoot(extInfo{rank=0, listenAddr})
    R1->>Root: sendToRoot(extInfo{rank=1, listenAddr})
    R2->>Root: sendToRoot(extInfo{rank=2, listenAddr})
    Note over Root: 收集所有 rank 的监听地址
    Root-->>R0: rootSend(rank2.addr) 下一个邻居
    Root-->>R1: rootSend(rank0.addr) 下一个邻居
    Root-->>R2: rootSend(rank1.addr) 下一个邻居
    R0->>R1: socketRingConnect(connect to next)
    R1->>R2: socketRingConnect(connect to next)
    R2->>R0: socketRingConnect(connect to next)
    Note over R0,R2: Ring 建立完成
    R0->>R1: socketRingAllGather 双向交换
    R1->>R2: socketRingAllGather 双向交换
    R2->>R0: socketRingAllGather 双向交换
    Note over R0,R2: 所有地址交换完成

3.3 commAlloc: the memory skeleton of the communicator object

Intuitive model

commAllocis the "roughcast delivery" of a communicator—it allocates the struct memory, initializes all fields to safe defaults, and creates the necessary CUDA objects and synchronization primitives, but has not yet filled in the "fine decoration" content such as topology information, channel configuration, and transport connections. IfncclCommis compared to a building,commAllocis laying the foundation and pouring the frame,initTransportsRankis the interior decoration.

Without the initialization performed bycommAlloc, subsequent code accessing uninitialized fields will lead to unpredictable behavior—for example, ifcomm->channels[c].idis a random value, the channel initialization logic will misjudge the channel state.

Data structures and memory layout

commAllocThe signature and initial validation of

📎 src/init.cc:512-526

It first validates the legality ofndevandrank, then constructs two memory stacks (memPermanentandmemScoped), and setsrankandnRanks. These two memory stacks are NCCL's memory management infrastructure—memPermanentis used for allocations whose lifetime is the same as the communicator,memScopedis used for temporary allocations.

Next is CUDA device probing:

📎 src/init.cc:528-531

cudaGetDeviceobtains the current device number,ncclCudaCompCapobtains the compute capability. The source code comment says it plainly: "Try to create a CUDA object right away. If there is something wrong with the device we're on, better know it early."—expose device problems as early as possible to avoid discovering them late in initialization.

Then comes the allocation or inheritance of shared resources:

📎 src/init.cc:533-555

There is an important branch here: ifparent == NULL || !parent->shareResources, create a newncclSharedResources; otherwise inherit the parent communicator's shared resources and increment the reference count.ncclSharedResourcescontains device streams, host streams, launch events, scratch events, etc.—these resources can be reused by sub-communicators in split scenarios, avoiding repeated creation.

Note thesharedRes->refCount = 1line—the initial reference count is 1, incremented each time it is shared by split, and only truly destroyed when the last reference is released.

Next is the initialization of network, RMA, and GIN:

📎 src/init.cc:547-549

These three subsystems are respectively responsible for network transport, remote memory access, and GPU-initiated network communication. Their initialization order matters—ncclNetInitmust precedencclRmaInit, because RMA depends on the network plugin.

Initialization of the memory manager:

📎 src/init.cc:567-576

There are likewise two paths: shared/new.ncclMemManageris responsible for managing the CUDA memory pool and registration cache.

Channel initialization marker:

📎 src/init.cc:607-608

This line sets all channels'idto -1, indicating "uninitialized". Later,setupChannelwill check this value to decide whether initialization is needed.

Construction of interrupt queues:

📎 src/init.cc:619-632

NCCL uses intrusive queues to manage various tasks. These queues are all constructed as empty during thecommAllocphase, and are used directly when subsequent tasks are enqueued.

Creation of the CUDA memory pool:

📎 src/init.cc:636-652

If the device supports memory pools (cudaDevAttrMemoryPoolsSupported), create a pinned-type memory pool and set the release threshold to the maximum value (~uint64_t(0)), meaning "never automatically release". This is to prevent the CUDA runtime from reclaiming memory without NCCL's knowledge.

Step-by-Step Walkthrough

Let us trace a specific initialization scenario: a single machine with 8 GPUs, one rank per process, normal initialization.

1. commAlloc(comm, NULL, 8, rank)is called,parent == NULL。

2. Validation passes,comm->rank = rank,comm->nRanks = 8。

3. cudaGetDevicereturns the current device number,comm->compCapis set.

4. Create a newncclSharedResources, with a reference count of 1.

5. ncclNetInitInitialize the network plugin (possibly Socket or IB).

6. ncclMemManagerInitCreate the memory manager.

7. getBusIdGet the PCI bus ID,ncclNvmlDeviceGetHandleByPciBusIdGet the NVML handle.

8. dmaBufSupportedDetect DMA-BUF support.

9. AllocateconnectSend / connectRecvbitmap array.

10. All channelsidset to -1.

11. Construct all interrupt queues.

12. Create CUDA memory pool.

Design considerations

commAllocThe most intriguing design is the "fail fast" principle. It callscudaGetDeviceat the beginning of the function, rather than waiting until device information is needed later. The benefit is that if there's a problem with the device (e.g., it's exclusively held by another process), the error is exposed early in initialization, rather than after allocating a large amount of memory.

Another design is the initialization ofpreconnectNext:

📎 src/init.cc:598-598

reinterpret_cast<struct ncclComm*>(0x1)is a sentinel value used to mark the state of "the next pre-connection." This technique of using an invalid pointer value as a state marker is common in systems programming—it saves memory compared to an extra boolean field, but care must be taken not to dereference it.

3.4 initTransportsRank: Topology Discovery and Channel Allocation

Intuitive model

initTransportsRankis the "heart" of initialization. It does three major things: exchange device information and topology information of all ranks through two AllGathers; compute the graph structures for algorithms such as ring/tree/collnet/nvls based on this information; and finally establish all transport connections. If the communication domain is likened to a city's transportation system,initTransportsRankis the process of planning all roads, overpasses, and bus routes.

Without this step, NCCL wouldn't know which path data should take—it might route data on a detour, or fail to find any reachable path at all.

Data structures and memory layout

initTransportsRankhas a large number of local variables; let's look at the key ones:

📎 src/init.cc:1163-1179

Here, the various graph structures in thecomm->graphsarray are extracted and aliases are created.graphsThe array is indexed by algorithm; note thatnvlsGraphis used twice (NVLS and NVLSTree share the same graph structure).

Two key temporary structures:

📎 src/init.cc:1181-1206

graphInfoholds the graph information of a single rank for a certain algorithm (number of channels, bandwidth, type, etc.),allGatherInfois the data unit for AllGather, containing graph information for all algorithms plus topology rank information.

Step-by-Step Walkthrough

Phase one: AllGather1—exchange device information.

📎 src/init.cc:1234-1239

Each rank callsfillInfoto fill its ownncclPeerInfo, then exchanges viabootstrapAllGather.fillInfoThe information filled in includes: rank number, CUDA device number, NVML device number, NCCL version, git hash, host hash, process hash, GPU UUID, bus ID, memory size, driver version, etc.

📎 src/init.cc:888-982

Noteinfo->hostHash = getHostHash() + commHashandinfo->pidHash = getPidHash() + commHash—both host hash and pid hash have commHash added. This is to distinguish different communication domains on the same machine.

After AllGather completes, each rank iterates through all peers' information and computes global attributes:

📎 src/init.cc:1250-1303

This loop does many things: detects version mismatches, counts the number of nodes, computes the intersection ofcuMemSupport, detects whether multiple ranks use the same GPU, computes the intersection of GIN type masks, etc. Note the counting method fornNodes—it increments each time a different hostHash is encountered, which assumes ranks are arranged contiguously by node.

Phase two: Topology discovery.

📎 src/init.cc:1390-1403

These six steps are the core process of topology discovery:ncclTopoGetSystemenumerates system devices to build the topology graph,ncclTopoComputePathscomputes GPU-to-NIC paths,ncclTopoTrimSystemremoves unreachable devices and computes paths again,ncclTopoSearchInitinitializes search state, and finally prints the topology.

Phase three: Graph computation.

📎 src/init.cc:1421-1468

Sequentially compute five graphs: ring, tree, collnet chain, collnet direct, and nvls. Each graph has different pattern and channel count constraints. NotetreeGraph->minChannels = ringGraph->nChannels—the tree's channel count is constrained to be the same as ring's, to ensure channel alignment between different algorithms.

Phase four: AllGather3—exchange graph information.

📎 src/init.cc:1490-1533

Each rank fills its graph information intoallGather3Data[rank], thenbootstrapAllGatheragain. The information exchanged this time includes: pattern/nChannels/bwIntra/bwInter/typeIntra/typeInter/crossNic for each algorithm, CPU architecture, P2P channel count, number of network devices, number of CollNet devices, etc.

After AllGather3 completes, each rank iterates through all peers' graph information and takes the minimum/maximum values to align:

📎 src/init.cc:1687-1703

Note the alignment strategy here:nChannels、sameChannels、bwIntra、bwIntertakes the minimum,typeIntra、typeInter、crossNictakes the maximum. Why? Because channel count and bandwidth are limited by the weakest link, while type and crossNic need to take the union to ensure compatibility.

Phase five: Establish transport connections.

📎 src/init.cc:1811-1892

There are two branches here:runtimeConnWhen true, only channel setup is done without connections (deferred to runtime), otherwise all connections are established immediately. The connection order is: ring → tree → NVLS → PAT → NVLS tree → CollNet.

Concurrency control and hardware interaction

initTransportsRankThere are several noteworthy concurrency/hardware interaction points in

CPU affinity setting:

📎 src/init.cc:1406-1412

NCCL binds the current thread to a CPU core near the GPU, ensuring that host memory allocations are on the local NUMA node. This reduces the latency of cross-NUMA accesses.

NVLS initialization:

📎 src/init.cc:1419-1419

ncclNvlsInitDetect NVLink SHARP support. NVLS allows the switch to directly perform reduce operations, greatly reducing AllReduce latency.

Proxy thread creation:

📎 src/init.cc:1780-1786

The proxy thread is responsible for asynchronously driving network I/O. It is created ininitTransportsRank, and all subsequent network operations go through the proxy.

Production Pitfall Guide

Pitfall 1: Mismatched number of network devices.If different ranks have different numbers of local NICs, NCCL will report an error:

📎 src/init.cc:1576-1596

UnlessNCCL_IGNORE_NET_MISMATCH=1is set. This is common in heterogeneous clusters—some nodes have 8 NICs, others only 4. Ignoring the mismatch can lead to performance degradation, because the number of channels will be limited by the weakest node.

Pitfall 2: Multiple ranks sharing the same GPU.If two ranks have the same GPU UUID, NCCL will refuse to initialize:

📎 src/init.cc:1291-1296

UnlessNCCL_MULTI_RANK_GPU_ENABLE=1is set. This check prevents performance issues caused by user misconfiguration.

Pitfall 3: Insufficient number of CollNet nodes.CollNet requires at leastNCCL_COLLNET_NODE_THRESHOLDnodes to be enabled:

📎 src/init.cc:1720-1728

The default threshold is 2. In a single-node environment, CollNet is automatically disabled.

mermaid
flowchart TD
    start["initTransportsRank(comm, parent, timers)"]
    ag1["AllGather1: fillInfo + bootstrapAllGather"]
    check_ver{"版本匹配?"}
    fail_ver["返回 ncclInvalidUsage"]
    topo["ncclTopoGetSystem + ComputePaths + TrimSystem"]
    graphs["计算 ring/tree/collnet/nvls 图"]
    ag3["AllGather3: 交换图信息"]
    align["对齐 nChannels/bwIntra/bwInter"]
    setup["setupChannel 初始化所有通道"]
    conn_ring["ncclTransportRingConnect"]
    conn_tree["ncclTransportTreeConnect"]
    conn_nvls["ncclNvlsSetup + ncclNvlsBufferSetup"]
    conn_collnet{"collnetEnable?"}
    conn_collnet_yes["ncclCollNetSetup + BufferSetup"]
    devcomm["devCommSetup 映射到设备"]
    barrier["bootstrapIntraNodeBarrier"]
    done["初始化完成"]

    start --> ag1 --> check_ver
    check_ver -->|否| fail_ver
    check_ver -->|是| topo --> graphs --> ag3 --> align --> setup
    setup --> conn_ring --> conn_tree --> conn_nvls --> conn_collnet
    conn_collnet -->|是| conn_collnet_yes --> devcomm
    conn_collnet -->|否| devcomm
    devcomm --> barrier --> done

3.5 NCCL_PARAM: The compile-time magic of the environment variable system

Intuitive model

NCCL_PARAMIt is NCCL's "configuration switch factory." It uses a macro to generate a function at compile time, and on the first runtime call it reads the environment variable and caches the result. This is like a light switch at home—you flip it (call the function), the light turns on (returns the configuration value), and afterward the switch state is remembered, so you don't need to flip it again every time.

Without this mechanism, NCCL would need to manually callgetenvand parse strings everywhere configuration is used, making the code extremely verbose and error-prone.

Data structures and memory layout

NCCL_PARAMThe definition of the macro:

📎 src/include/param.h:22-31

This macro expands to generate a functionncclParam##name(), with three static variables inside:

  • uninitialized = INT64_MIN: sentinel value, indicating "not yet initialized."
  • noCache: tri-state flag, -1 means uninitialized, 0 means cache, 1 means do not cache.
  • cache: the cached value, initiallyuninitialized。

The function logic is: ifcacheis stilluninitialized, callncclLoadParamto load; otherwise directly returncache。COMPILER_EXPECT(..., false)tells the compiler that this branch is rarely taken, optimizing the hot path.

ncclLoadParamThe implementation of:

📎 src/misc/param.cc:78-108

It uses a mutex to protect the entire loading process, first checking thenoCachepolicy, then checking whether the cache is valid, and then reading and parsing the environment variable. On parse failure, it uses the default value and prints a warning.

Step-by-Step Walkthrough

TakingNCCL_PARAM(BuffSize, "BUFFSIZE", -2)as an example:

📎 src/init.cc:1007-1007

After macro expansion it generates:

cpp
int64_t ncclParamBuffSize() {
  constexpr int64_t uninitialized = INT64_MIN;
  static int8_t noCache = -1;
  static_assert(-2 != uninitialized, "...");
  static int64_t cache = uninitialized;
  if (COMPILER_EXPECT(COMPILER_ATOMIC_LOAD(&cache, std::memory_order_relaxed) == uninitialized, false)) {
    return ncclLoadParam("NCCL_BUFFSIZE", -2, uninitialized, &cache, &noCache);
  }
  return cache;
}

On the first call,cache == uninitialized, entersncclLoadParam. It reads theNCCL_BUFFSIZEenvironment variable, and if it is not set, returns the default value -2. Then, based on thenoCachepolicy, it decides whether to cache.

noCacheThe policy is determined byncclParamIsCacheDisabled:

📎 src/misc/param.cc:74-76

If the environment variable name matches a certain pattern (for example, ends with_), it is not cached and is re-read every time. This allows users to dynamically modify certain configurations at runtime.

Design considerations

The brilliance of this design lies in "zero-cost abstraction": on the hot path there is only one atomic load and comparison, with no locks and no string parsing. Only the cold path (first load) pays the full cost.COMPILER_EXPECTIt hints to the compiler to place the hot path earlier in the instruction cache, further improving performance.

Another design is the tri-state design ofnoCache. -1 means "not yet decided," 0 means "cache," and 1 means "do not cache." This decision is made only once on the first load and does not change afterward.

Production Pitfall Guide

Pitfall 1: Misspelled environment variable.If the user writesNCCL_BUFSIZEinstead ofNCCL_BUFFSIZE, NCCL will not report an error and will only use the default value. It is recommended to useNCCL_DEBUG=ENVto view all recognized environment variables.

Pitfall 2: The loading order ofNCCL_CONF_FILE.NCCL loads$NCCL_CONF_FILE(or~/.nccl.conf) and/etc/nccl.conf:

📎 src/misc/param.cc:52-67

in sequence. Files loaded later override those loaded earlier. If both files set the same variable,/etc/nccl.conf's value takes effect.

Pitfall 3: Thread safety of thenoCachevariable.The source comment says "noCache is only load/stored within the mutex, no need for atomic":

📎 src/misc/param.cc:74-76

This means that reads and writes ofnoCacheare both protected by the mutex and do not require atomic operations. But reads ofcacheare lock-free (hot path), so atomic loads are used.

3.6 devCommSetup: Mapping the communication domain to the device

Intuitive model

devCommSetupIt is the "device-side projection" of the communication domain. GPU kernels run on the device and cannot directly access thencclCommstructure in host memory. Therefore, NCCL needs to copy the key fields of the communication domain into device-accessible memory, formingncclDevComm. This is like making a copy of the company directory and placing it at every employee's workstation—employees don't have to run to the front desk every time to ask for a colleague's phone number.

WithoutdevCommSetup, the GPU kernel cannot know its own rank, channel configuration, buffer size, and other information, and the collective communication kernel cannot start at all.

Data structures and memory layout

devCommSetupIt uses a temporary structurencclKernelCommAndChannelsto package the data to be copied to the device:

📎 src/init.cc:712-746

This structure containsncclDevComm(the device-side communication domain) and the channel array. The function first fills the host-side data into the temporary structure, then performs a singlecudaMemcpyAsyncto the device.

Filling of key fields:

📎 src/init.cc:734-746

Notecomm->devComm = &devCommAndChans->comm—the host-sidecomm->devCommpoints to thencclDevCommin device memory. When the kernel is subsequently launched,comm->devCommwill be passed in as a parameter.

Filling of channel information:

📎 src/init.cc:829-843

The peers, ring, tree, collnetChain, collnetDirect, and nvls pointers of each channel are copied to the device side. Note thatring.userRanksrequires an additionalcudaMemcpyAsyncbecause it is an array.

Step-by-Step Walkthrough

1. Get the device stream:ncclStrongStreamAcquireGet a strong stream to ensure that subsequent asynchronous copies execute in order.

2. Allocate device memory:ncclCudaCallocAsyncAllocatedevCommAndChans。

3. Fill the host-side temporary struct: set rank, nRanks, node, nNodes, abortFlag, buffSizes, etc.

4. Allocate and copy therankToLocalRankarray.

5. ComputeworkFifoBytes: determined based on the CC (Confidential Computing) state.

6. Allocate the workFifo buffer: usencclGdrCudaCallocin GDR mode, otherwise usencclCudaHostCalloc。

7. Allocate profiler counters.

8. Allocate progress counters (if enabled).

9. Fill channel information.

10. Copy to the device in one go:ncclCudaMemcpyAsync(devCommAndChans, &tmpCommAndChans, 1, deviceStream)。

11. Release the strong stream and synchronize.

Design considerations

devCommSetupThe most noteworthy design in this is "batch copy". NCCL does not callcudaMemcpyseparately for each field; instead, it packs all fields into a temporary struct and uses a singlecudaMemcpyAsyncto complete it. This greatly reduces the number of CUDA API calls and synchronization overhead.

Another design isworkFifoBytes's CC handling:

📎 src/init.cc:750-763

In CC (Confidential Computing) mode,workFifoBytesis set to 0 because GDR copy is unavailable in CC mode. This is an elegant degradation due to a hardware limitation.

Production pitfall guide

Pitfall 1:devCommSetupmust be called before the barrier.The source code comments explain the reason:

📎 src/init.cc:1950-1952

If it is called after the barrier, some threads may have already started launching the NCCL kernel, while device memory has not yet been fully allocated, which can cause a deadlock.

Pitfall 2:workFifoBytesmust be a power of 2.If it is not, NCCL will warn and use the default value:

📎 src/init.cc:757-762

Chapter review and self-test

Q1: If the logic in📎 src/init.cc:1291-1296that detects "multiple ranks using the same GPU" is removed, in what scenarios would it cause problems? Why does NCCL reject this configuration by default?

Reference analysis:

This code checks whether the GPU UUIDs of two ranks on the same host are the same. If they are the same andNCCL_MULTI_RANK_GPU_ENABLE=0(default), it returnsncclInvalidUsage。

After removing this check, multiple ranks will share the same GPU. This will cause:

1. P2P transfer conflicts: NCCL's P2P transfer assumes that each rank exclusively occupies one GPU. If two ranks share a GPU, they will simultaneously write data to the same buffer of the same GPU, causing data races and incorrect results.

2. Channel allocation conflicts:comm->channelsThe channel resources (buffers, FIFO) in are allocated per rank. Ranks sharing a GPU will contend for the same resources.

3. Performance disaster: Even if there are no correctness issues, two ranks sharing one GPU's compute power and memory bandwidth will cause performance to drop sharply.

NCCL rejects this configuration by default in order to "fail fast" - rather than letting users waste hours debugging a misconfiguration, it is better to report an error clearly during initialization.NCCL_MULTI_RANK_GPU_ENABLE=1is an escape hatch prepared for users who clearly know what they are doing (such as in MPS scenarios).

Q2: If the logic in📎 src/bootstrap.cc:1129-1134that waits for "an earlier send to the same (peer, tag)" is removed, in what scenarios would it cause the receiver to match incorrectly?

Reference analysis:

This code waits in the asynchronous send thread until there is no earlier send to the same (peer, tag) in the queue.

After removing this wait, two sends to the same (peer, tag) may execute concurrently, and the order in which they arrive at the receiver is uncertain. The receiver'ssocketAcceptmatches connections by (peer, tag):

📎 src/bootstrap.cc:1291-1292

If sender A callsbootstrapSendfirst but arrives later, and sender B calls later but arrives first, the receiver will treat B's message as A's response. This causes data misalignment - the receiver thinks it received the response to the first request, but it is actually the response to the second request.

The source code comments explicitly point out this scenario: "NVLS setup broadcasts to the same peers with the same tag several times during init". During NVLS initialization, broadcasts are sent multiple times to the same peer with the same tag. If the order is reversed, the NVLS configuration will be completely disrupted.

The cost of this ordering guarantee is that sends to the same (peer, tag) are serialized. However, sends to different (peer, tag) are still concurrent, so overall throughput is not affected.

Q3: If the alignment strategy in📎 src/init.cc:1691-1697is changed from "take min for nChannels and max for typeIntra" to "take min for all" or "take max for all", what problems would each cause?

Reference analysis:

The current strategy is:nChannels、sameChannels、bwIntra、bwIntertake min,typeIntra、typeInter、crossNictake max.

If all take min:typeIntraandtypeInterTaking min would cause the transport type of some ranks to be downgraded. For example, if rank A supports P2P (typeIntra=P2P) and rank B only supports SHM (typeIntra=SHM), taking min would make all ranks use SHM. But the enum value of SHM may be smaller than that of P2P, so taking min would select the wrong type. In realitytypeIntrais a bitmask or enum, and taking max is to select the "most capable" type.

If max is taken for everything:nChannelsTaking max would cause some ranks to be allocated more channels than they can support. For example, if rank A can only support 4 channels and rank B supports 8, taking max would make all ranks try to use 8 channels, and rank A would fail or suffer performance degradation.bwIntraTaking max would make the bandwidth estimate overly optimistic, and the tuning module might select an unsuitable algorithm.

The essence of this alignment strategy is:Resource constraints take the intersection (min), capability enums take the union (max). Channel count and bandwidth are "upper bound" constraints and must take the most conservative value; transport type is a "capability" enum, and taking the maximum ensures that all ranks can find a compatible transport method.

In the next chapter, we will dive into topology discovery and graph search, and see how NCCL enumerates the GPUs, NICs, and PCI switches in a machine, builds a complete topology graph, and searches this graph for the optimal ring and tree structures. The bootstrap communication, commAlloc memory skeleton, and initTransportsRank main flow established in this chapter will have their topology details unfolded one by one in the next chapter.

At this point, we have fully walked through the call chain of ncclCommInitRank and seen the entire process of building the ncclComm object from scratch. But there is one key step in the initialization process that we only briefly skimmed over: how does NCCL detect the GPUs and NICs inside a machine and use that to decide which path the data should take? This is exactly the topic to be explored in depth in the next chapter—topology discovery and graph search. We will break down how src/graph/topo.cc enumerates PCI/NVLink/NIC devices and builds the topology graph, how src/graph/search.cc searches for the optimal path on that graph, and how src/graph/rings.cc and trees.cc materialize the search results into Ring and Tree algorithm topologies. Once you understand this mechanism, you will understand why NCCL can automatically select suitable algorithms on different machines.

CHAPTER 04

Chapter 4: Chapter 4: Topology Discovery and Graph Search: How NCCL "Sees" the Physical Interconnect of a Multi-GPU System

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 4 / 25

Chapter 4: Topology Discovery and Graph Search: How NCCL "Sees" the Physical Interconnect of a Multi-GPU System

In the previous chapter, we drilled down layer by layer along the call chain of ncclCommInitRank and saw when the comm->topo field is populated, but we did not expand its internal structure. So how exactly does NCCL "see" the GPUs and NICs in a machine and organize them into usable topology information? This chapter will break down the three key steps of this process: topo.cc is responsible for enumerating physical devices into a graph, search.cc searches for the optimal path on this graph, and rings.cc and trees.cc materialize the search results into the two algorithm topologies, Ring and Tree. Only by understanding how these three work together can you understand why NCCL can automatically select suitable algorithms on different machines.

Topology Graph: Drawing a Machine as a "Subway Map"

Intuitive Model

Imagine you are a courier who has just arrived in an unfamiliar city. You need to deliver a package from point A to point B, but you do not know which route is fastest. You need a map—one marked with all stations (GPUs, NICs, CPUs, PCI switches) and the connections between stations (NVLink, PCIe, network). NCCL's topology graph is this map.

Without this map, NCCL can only blindly assume that "all GPUs have the same bandwidth." On a machine with 8 GPUs fully interconnected by NVLink, this might still barely work, but once it encounters a complex topology with cross-NUMA, cross-PCI-switch, and mixed NVLink + PCIe, it will choose the wrong path, stuffing data that should go over NVLink into slow PCIe, and performance will be cut in half.

Data Structures and Memory Layout

The core of the topology graph isncclTopoSystem, which stores all devices grouped by node type. The node types are defined in thetopoNodeTypeStrarray:

📎 src/graph/topo.cc:33-35

c
const char* topoNodeTypeStr[] = {"GPU", "PCI", "NVS", "CPU", "NIC", "NET", "GIN", "RMA", "DEV", "CXB"};
const char* topoLinkTypeStr[] = {"LOC", "NVL", "", "C2C", "PCI", "", "", "", "", "SYS", "NET"};
const char* topoPathTypeStr[] = {"LOC", "NVL", "NVB", "C2C", "PIX", "PXB", "P2C", "PXN", "PHB", "SYS", "NET", "DIS"};

These three arrays define the string representations of node types, link types, and path types, respectively. Note thetopoPathTypeStrorder—it also serves as the ranking of path quality: the smaller the index, the faster the path.LOC(local) is the fastest,DIS(disconnected) is the slowest. This order will be used repeatedly in subsequent searches to compare the quality of paths.

Each node is represented byncclTopoNodeand different fields are initialized according to the type at creation time. Taking a GPU node as an example:

📎 src/graph/topo.cc:105-141

c
ncclResult_t ncclTopoCreateNode(struct ncclTopoSystem* system, struct ncclTopoNode** node, int type, uint64_t id) {
  if (system->nodes[type].count == NCCL_TOPO_MAX_NODES) {
    WARN("Error : tried to create too many nodes of type %d", type);
    return ncclInternalError;
  }
  struct ncclTopoNode* n = system->nodes[type].nodes + system->nodes[type].count;
  system->nodes[type].count++;
  n->type = type;
  n->id = id;
  if (type == GPU) {
    n->gpu.dev = NCCL_TOPO_UNDEF;
    n->gpu.rank = NCCL_TOPO_UNDEF;
    n->gpu.cudaCompCap = NCCL_TOPO_UNDEF;
    n->gpu.mloPart = NCCL_TOPO_UNDEF;
  } else if (type == CPU) {
    ...

There are several key design points here. First, nodes are stored in a preallocated array (system->nodes[type].nodes), rather than a linked list. This means that nodes are arranged contiguously in memory, making traversal cache-friendly. Second,NCCL_TOPO_MAX_NODESis a hard upper limit, and exceeding it raises an error—this is to prevent unbounded growth when the topology is abnormal. Third, each node has aidfield, which is a 64-bit integer. The upper 32 bits are systemId (identifying which host), and the lower 32 bits are localId (the device number within the host).

The connections between nodes are represented byncclTopoLink.ncclTopoConnectNodesis responsible for establishing bidirectional connections:

📎 src/graph/topo.cc:179-204

c
ncclResult_t ncclTopoConnectNodes(struct ncclTopoNode* node, struct ncclTopoNode* remNode, int type, float bw) {
  // Aggregate links into higher bw for NVLink
  struct ncclTopoLink* link;
  for (link = node->links; link - node->links != NCCL_TOPO_MAX_LINKS && link->remNode; link++) {
    if (link->remNode == remNode && link->type == type) break;
  }
  if (link - node->links == NCCL_TOPO_MAX_LINKS) {
    WARN("Error : too many Topo links (max %d)", NCCL_TOPO_MAX_LINKS);
    return ncclInternalError;
  }
  if (link->remNode == NULL) node->nlinks++;
  link->type = type;
  link->remNode = remNode;
  link->bw += bw;

  // Sort links in BW descending order
  struct ncclTopoLink linkSave;
  memcpy(&linkSave, link, sizeof(struct ncclTopoLink));
  while (link != node->links) {
    if ((link - 1)->bw >= linkSave.bw) break;
    memcpy(link, link - 1, sizeof(struct ncclTopoLink));
    link--;
  }
  memcpy(link, &linkSave, sizeof(struct ncclTopoLink));
  return ncclSuccess;
}

This function does three things. First, it checks whether a link to the same target and of the same type already exists—if so, it accumulates the bandwidth (link->bw += bw). This handles the case where multiple NVLinks connect to the same GPU: 4 NVLinks at 25 GB/s each aggregate to 100 GB/s. Second, if none is found, it adds a new link. Third, after insertion, it sorts in descending order by bandwidth, so that subsequent traversals see high-bandwidth links first.

[Design Inference and Architectural Trade-offs]

The design motivation for sorting in descending order by bandwidth is to let the search algorithm discover high-bandwidth paths as early as possible, thereby converging faster to a better solution. The search has a timeout limit (as will be seen later inNCCL_SEARCH_TIMEOUT), and sorting allows the limited time budget to be spent on more promising paths.

Scenario-Driven Step-by-Step Walkthrough

Now let us plug in a concrete scenario: an 8-GPU A100 server, where each GPU is fully interconnected via NVLink, and there are also 4 Mellanox ConnectX-6 NICs plugged into PCIe slots. During NCCL initialization,ncclTopoGetSystemis called. It reads device information from an XML file (generated bynvidia-topologydor by NCCL itself), and then builds the topology graph.

Step 1: Parse the CPU node.ncclTopoAddCpureads the CPU architecture, vendor, and model from the XML, and creates the CPU node:

📎 src/graph/topo.cc:806-875

c
ncclResult_t ncclTopoAddCpu(struct ncclXmlNode* xmlCpu, struct ncclTopoSystem* system) {
  int numaId;
  NCCLCHECK(xmlGetAttrInt(xmlCpu, "numaid", &numaId));
  int systemId;
  NCCLCHECK(ncclGetSystemId(system, xmlCpu, &systemId));
  struct ncclTopoNode* cpu;
  NCCLCHECK(ncclTopoCreateNode(system, &cpu, CPU, NCCL_TOPO_ID(systemId, numaId)));
  ...
  for (int s = 0; s < xmlCpu->nSubs; s++) {
    struct ncclXmlNode* node = xmlCpu->subs[s];
    if (strcmp(node->name, "pci") == 0) NCCLCHECK(ncclTopoAddPci(node, system, cpu, systemId, numaId));
    if (strcmp(node->name, "nic") == 0) {
      ...
    }
  }
  return ncclSuccess;
}

The CPU node is the root of the topology tree. Under each CPU hang the PCI subtree and NIC nodes.ncclTopoAddPcirecursively processes the PCI tree, creating a GPU node when it encounters a GPU and a NIC node when it encounters a NIC.

Step 2: Add NVLink connections. Note thatncclTopoAddGpuonly reads the basic attributes of the GPU, and the comment explicitly says "Do not go any further, nvlinks will be added in a second pass":

📎 src/graph/topo.cc:590-598

c
ncclResult_t ncclTopoAddGpu(struct ncclXmlNode* xmlGpu, struct ncclTopoSystem* system, struct ncclTopoNode* gpu) {
  NCCLCHECK(xmlGetAttrInt(xmlGpu, "rank", &gpu->gpu.rank));
  NCCLCHECK(xmlGetAttrInt(xmlGpu, "sm", &gpu->gpu.cudaCompCap));
  NCCLCHECK(xmlGetAttrInt(xmlGpu, "dev", &gpu->gpu.dev));
  NCCLCHECK(xmlGetAttrInt(xmlGpu, "gdr", &gpu->gpu.gdrSupport));
  NCCLCHECK(xmlGetAttrIntDefault(xmlGpu, "mlopart", &gpu->gpu.mlopart, NCCL_TOPO_UNDEF));
  // Do not go any further, nvlinks will be added in a second pass
  return ncclSuccess;
}

Why split it into two passes? Because NVLink is a connection between GPUs, and links can only be established after the GPU nodes at both ends already exist. The first pass creates all nodes, and the second passncclTopoAddNvLinksthen connects them.

Step 3: Handle network devices.ncclTopoAddNictraverses the net/gin/rma child nodes under the NIC and calls the corresponding add functions respectively. TakingncclTopoAddNetas an example:

📎 src/graph/topo.cc:461-503

c
static ncclResult_t ncclTopoAddNet(struct ncclXmlNode* xmlNet, struct ncclXmlNode* parent,
                                   struct ncclTopoSystem* system, struct ncclTopoNode* nic, int systemId) {
  int dev;
  NCCLCHECK(xmlGetAttrInt(xmlNet, "dev", &dev));
  int64_t netId = NCCL_TOPO_ID(systemId, dev);
  struct ncclTopoNode* net;
  NCCLCHECK(ncclTopoCreateNode(system, &net, NET, netId));
  net->net.dev = dev;
  int mbps;
  NCCLCHECKNOWARN(xmlGetAttrIntDefault(xmlNet, "speed", &mbps, 0), NCCL_GRAPH);
  if (mbps <= 0) mbps = 10000; // Some NICs define speed = -1
  net->net.bw = mbps / 8000.0;
  ...
  NCCLCHECK(ncclTopoConnectNodes(nic, net, LINK_NET, net->net.bw));
  NCCLCHECK(ncclTopoConnectNodes(net, nic, LINK_NET, net->net.bw));
  return ncclSuccess;
}

Note the conversion inmbps / 8000.0: mbps is megabits per second, and dividing by 8000 gives GB/s (because 1 GB/s = 8000 Mbps). If the NIC reports speed = -1 (as some virtual NICs do), it defaults to 10000 Mbps = 1.25 GB/s.

Step 4: Finalization.ncclTopoGetSystemFromXmlAfter all nodes and links have been added,

📎 src/graph/topo.cc:1080-1088

c
  NCCLCHECK(ncclTopoAddNvLinks(topNode, *topoSystem, NULL, 0));
  NCCLCHECK(ncclTopoAddC2c(topNode, *topoSystem, NULL, 0));
  NCCLCHECK(ncclTopoAddPciLinks(topNode, *topoSystem, NULL, 0));

  NCCLCHECK(ncclTopoFlattenBcmSwitches(*topoSystem));
  NCCLCHECK(ncclTopoConnectCpus(*topoSystem));
  NCCLCHECK(ncclTopoSortSystem(*topoSystem));

ncclTopoFlattenBcmSwitchesCopyncclTopoConnectCpushandles the special case of Broadcom Gen4 PCIe switches—they present themselves as two-layer switches, but are actually full-bandwidth, and need to be "flattened" to avoid misleading the search algorithm.ncclTopoSortSystemconnects all CPU nodes to one another (cross-NUMA access goes over SYS links).

sorts the links so that PCI downstream links come first, making traversal easier.

Design Reflections and Production Pitfalls

[Design Inference and Architectural Trade-offs]Why use XML as the intermediate format?

Because topology discovery needs to be shared across processes—each rank only probes the GPUs it manages, then exchanges XML through bootstrap, and finally merges it into the complete topology. XML is a self-describing text format, which is convenient for debugging (it can be dumped and inspected) and version compatibility.ncclTopoGetNodePitfall 1:does not report an error when it cannot find a node.

📎 src/graph/topo.cc:95-103

c
ncclResult_t ncclTopoGetNode(struct ncclTopoSystem* system, struct ncclTopoNode** node, int type, uint64_t id) {
  for (int i = 0; i < system->nodes[type].count; i++) {
    if (system->nodes[type].nodes[i].id == id) {
      *node = system->nodes[type].nodes + i;
      return ncclSuccess;
    }
  }
  return ncclSuccess;
}

CopyncclSuccessIf it is not found, it returns*nodebut*node == NULLremains unchanged (callers usually initialize it to NULL). The caller must check

itself. This design is prone to missed checks—if the caller forgets to check, a later dereference will crash.ncclTopoConnectNodesPitfall 2:bandwidth accumulation may cause overflow.link->bw += bwIf there are a large number of links between the same pair of nodes (such as in an NVSwitch scenario),

may accumulate to a very large value. Although float precision is sufficient, if the number of links is abnormally large, the sorting logic may go wrong.ncclTopoRemoveNodePitfall 3:pointer fixup.

📎 src/graph/topo.cc:143-177

c
ncclResult_t ncclTopoRemoveNode(struct ncclTopoSystem* system, int type, int index) {
  struct ncclTopoNode* delNode = system->nodes[type].nodes + index;
  for (int t = 0; t < NCCL_TOPO_NODE_TYPES; t++) {
    if (delNode->paths[t] != nullptr) {
      WARN("Cannot remove topology node %d/%lx while paths are computed", type, delNode->id);
      return ncclInternalError;
    }
    for (int n = 0; n < system->nodes[t].count; n++) {
      struct ncclTopoNode* node = system->nodes[t].nodes + n;
      if (node == delNode) continue;
      for (int l = 0; l < node->nlinks; l++) {
        while (l < node->nlinks && node->links[l].remNode == delNode) {
          memmove(node->links + l, node->links + l + 1, (node->nlinks - l - 1) * sizeof(struct ncclTopoLink));
          node->nlinks--;
        }
        if (l < node->nlinks && node->links[l].remNode->type == type && node->links[l].remNode >= delNode) {
          node->links[l].remNode--;
        }
      }
    }
  }
  ...

Copynode->links[l].remNode--There is a subtle point here:sizeof(struct ncclTopoNode)is fixing pointers. Because nodes are stored in a contiguous array, after deleting one node, the addresses of all subsequent nodes shift forward by onememmoveExecuted before, order is critical.

Path search: finding the "optimal route" on the graph

Intuitive model

Having a map isn't enough—you also need a navigation algorithm. NCCL's path search has two layers: the first layer is preprocessing, computing the shortest path between all pairs of nodes (BFS); the second layer is graph search, trying different Ring/Tree structures on the preprocessing results to find the one with the highest bandwidth.

Without path search, NCCL could only hardcode fixed orders like "GPU 0 connects to GPU 1 connects to GPU 2...", which would select slow paths on non-uniform topologies.

Data structures and memory layout

The core data structure of path search isncclTopoLinkList, which stores the complete path from a source node to a target node:

c
struct ncclTopoLinkList {
  struct ncclTopoLink* list[NCCL_TOPO_MAX_HOPS];  // 路径上的链路
  int count;      // 跳数
  float bw;       // 瓶颈带宽
  int type;       // 路径类型(PATH_LOC, PATH_NVL, ...)
  int capacity;   // list 数组的容量
};

Each node has apaths[type]array, storing paths to all nodes of that type. For example, a GPU node'spaths[NET]stores paths to all NICs.

Path computation is done byncclTopoSetPaths, which is a BFS:

📎 src/graph/paths.cc:52-147

c
static ncclResult_t ncclTopoSetPaths(struct ncclTopoNode* baseNode, struct ncclTopoSystem* system) {
  if (baseNode->paths[baseNode->type] == NULL) {
    NCCLCHECK(ncclCalloc(baseNode->paths + baseNode->type, system->nodes[baseNode->type].count));
    for (int i = 0; i < system->nodes[baseNode->type].count; i++) baseNode->paths[baseNode->type][i].type = PATH_DIS;
  }

  // breadth-first search to set all paths to that node in the system
  struct ncclTopoNodeList nodeList;
  struct ncclTopoNodeList nextNodeList = {{0}, 0};
  nodeList.count = 1;
  nodeList.list[0] = baseNode;
  ...
  while (nodeList.count) {
    nextNodeList.count = 0;
    for (int n = 0; n < nodeList.count; n++) {
      struct ncclTopoNode* node = nodeList.list[n];
      struct ncclTopoLinkList* path;
      NCCLCHECK(getPath(system, node, baseNode->type, baseNode->id, &path));
      for (int l = 0; l < node->nlinks; l++) {
        struct ncclTopoLink* link = node->links + l;
        struct ncclTopoNode* remNode = link->remNode;
        ...
        float bw = std::min(path->bw, link->bw);
        ...
        // Update if better path type, OR same type with higher bw, OR same type/bw with strickly fewer hops.
        if (newType < remPath->type || (newType == remPath->type && remPath->bw < bw) ||
            (newType == remPath->type && remPath->bw == bw && remPath->count > (path->count + 1))) {
          ...
          remPath->bw = bw;
          remPath->type = newType;
          ...
        }
      }
    }
    memcpy(&nodeList, &nextNodeList, sizeof(nodeList));
  }
  return ncclSuccess;
}

BFS starts frombaseNodeand expands layer by layer. Each time a new node is reached, the path's bottleneck bandwidth (std::min(path->bw, link->bw)) and path type are computed. There are several special rules for computing path type:

  • If it passes through two PCI switches, the type is upgraded toPATH_PXB
  • If it passes through the CPU, the type is upgraded toPATH_PHB
  • If it passes through a DEV node and is NVLink, the type is upgraded toPATH_NVB

The update condition is "better path": better type, or same type but higher bandwidth, or same type and bandwidth but fewer hops.

Scenario-driven Step-by-Step Walkthrough

Now let's look at the second layer of search.ncclTopoComputeis the entry point, which tries different parameter combinations and callsncclTopoSearchRecto perform the search.

The core of the search is the recursive functionncclTopoSearchRecGpu. It starts from a certain GPU and tries to reach the next GPU until all GPUs are traversed, forming a path:

📎 src/graph/search.cc:639-756

c
ncclResult_t ncclTopoSearchRecGpu(struct ncclTopoSystem* system, struct ncclTopoGraph* graph,
                                  struct ncclTopoGraph* saveGraph, struct ncclTopoNode* gpu, int step, int backToNet,
                                  int backToFirstRank, int forcedOrder, int* time) {
  if ((*time) <= 0) return ncclSuccess;
  (*time)--;
  ...
  if (step == ngpus) {
    // Determine whether we found a better solution or not
    int copy = 0;
    graph->nChannels++;
    NCCLCHECKGOTO(ncclTopoCompareGraphs(system, graph, saveGraph, &copy), ret, exit);
    if (copy) {
      memcpy(saveGraph, graph, sizeof(struct ncclTopoGraph));
      if (graph->nChannels == graph->maxChannels) *time = -1;
    }
    if (graph->nChannels < graph->maxChannels) {
      NCCLCHECKGOTO(ncclTopoSearchRec(system, graph, saveGraph, time), ret, exit);
    }
    graph->nChannels--;
    ret = ncclSuccess;
    goto exit;
  }
  graph->intra[graph->nChannels * ngpus + step] = gpu->gpu.rank;
  g = gpu - system->nodes[GPU].nodes;
  if (step == backToNet) {
    // first get back to NIC
    ...
  } else if (graph->pattern == NCCL_TOPO_PATTERN_NVLS) {
    ...
  } else if (step < system->nodes[GPU].count - 1) {
    // Go to next GPU
    ...
  } else if (step == backToFirstRank) {
    // Find first GPU and loop back to it
    ...
  } else {
    // Next path
    NCCLCHECKGOTO(ncclTopoSearchRecGpu(system, graph, saveGraph, gpu, ngpus, -1, -1, forcedOrder, time), ret, exit);
  }
  ...
}

This function has several key branches:

1. step == ngpus: all GPUs have been traversed, forming a complete path. At this point incrementnChannels, compare the current graph with the saved optimal graph, and save if better. Then recursively callncclTopoSearchRecto try searching for the next channel.

2. step == backToNet: need to return to the NIC. This happens in Ring mode (the last GPU must connect back to the starting NIC) or Tree mode (the first GPU must connect to the NIC).

3. step < ngpus - 1: continue to the next GPU. HerencclTopoSearchNextGpuSortis called to sort candidate GPUs.

4. step == backToFirstRank: in Ring mode, the last GPU must connect back to the first GPU.

5. else: the path ends, proceed to the next round.

ncclTopoSearchNextGpuSortdetermines the order in which the next GPU is tried:

📎 src/graph/search.cc:254-327

c
ncclResult_t ncclTopoSearchNextGpuSort(struct ncclTopoSystem* system, struct ncclTopoGraph* graph,
                                       struct ncclTopoNode* gpu, int* next, int* countPtr, int sortNet) {
  const uint64_t flag = 1ULL << (graph->nChannels);
  int ngpus = system->nodes[GPU].count;
  struct ncclTopoLinkList* paths = gpu->paths[GPU];
  ...
  for (int i = 1; i < ngpus; i++) {
    int g = (start + i) % ngpus;
    if (paths[g].count == 0) continue; // There is no path to that GPU
    if (system->nodes[GPU].nodes[g].used & flag) continue;
    scores[count].g = g;
    scores[count].startIndex = i;
    scores[count].intraNhops = paths[g].count;
    scores[count].intraBw = paths[g].bw;
    if (netPaths) {
      scores[count].interNhops = netPaths[g].count;
      scores[count].interPciBw = gpuPciBw(system->nodes[GPU].nodes + g);
      scores[count].interBw = netPaths[g].bw;
    }
    count++;
  }

  // Sort GPUs
  qsort(scores, count, sizeof(struct ncclGpuScore), cmpScore);
  ...
}

It scores each candidate GPU, with sorting rules: first compare interBw (bandwidth to NIC), then interPciBw, then interNhops, then intraBw, and finally intraNhops. This priority reflects NCCL's optimization goal: cross-machine communication is the bottleneck, so GPUs with high NIC bandwidth are preferred.

Design considerations and production pitfalls

Why does the search have a timeout?Look at these constants:

📎 src/graph/search.cc:329-330

c
#define NCCL_SEARCH_GLOBAL_TIMEOUT (1ULL << 19)
#define NCCL_SEARCH_TIMEOUT (1 << 14)
#define NCCL_SEARCH_TIMEOUT_TREE (1 << 14)
#define NCCL_SEARCH_TIMEOUT_SAMECHANNELS (1 << 8)

The search space is exponential—each channel has O(ngpus!) permutations. An 8-GPU machine has 40320, and 16 GPUs has 2 trillion. The search time must be limited.NCCL_SEARCH_TIMEOUTis 16384 iterations,NCCL_SEARCH_GLOBAL_TIMEOUTis 524288. After timeout, the current optimal solution is returned.

Pitfall one:ncclTopoFollowPath's bandwidth deduction is a global side effect.Look at this function:

📎 src/graph/search.cc:127-173

c
static ncclResult_t ncclTopoFollowPath(struct ncclTopoSystem* system, struct ncclTopoGraph* graph, int type1,
                                       int index1, int type2, int index2, float mult, struct ncclTopoNode** node) {
  ...
  bw *= mult;
  // Check there is enough bandwidth on paths.
  int step = 0;
  NCCLCHECK(followPath(path, node1, path->count, bw, &step));
  if (step < path->count) goto rewind;
  // Enough bandwidth : return destination node.
  graph->nHops += mult * path->count;
  *node = system->nodes[type2].nodes + index2;
  return ncclSuccess;
rewind:
  // Not enough bandwidth : rewind and exit.
  NCCLCHECK(followPath(path, node1, step, -bw, &step));
  return ncclSuccess;
}

followPathmodifies each link'sbwon the path (deducting used bandwidth). If the search fails,followPathmust be called to restore using-bw. This "deduct-restore" pattern is error-prone in recursive search—if a branch forgets to restore, subsequent searches will see incorrect bandwidth.

Pitfall two:ncclTopoCompareGraphs's comparison logic is very subtle.It prioritizes comparingnChannels * bwIntra, but there are also a bunch of special cases:

📎 src/graph/search.cc:446-477

c
ncclResult_t ncclTopoCompareGraphs(struct ncclTopoSystem* system, struct ncclTopoGraph* graph,
                                   struct ncclTopoGraph* refGraph, int* copy) {
  // 1. Try to get the same nChannels between Rings and Trees
  if (graph->nChannels < graph->minChannels) return ncclSuccess;
  const bool evenReference = refGraph->nChannels > 0 && !(refGraph->nChannels & 1);
  const bool evenReferenceIsBetter = refGraph->nChannels * refGraph->bwIntra >= graph->nChannels * graph->bwIntra;
  // Favor an even number of channels when aggregate bandwidth is equal or better.
  if (graph->pattern != NCCL_TOPO_PATTERN_NVLS && evenReference && (graph->nChannels & 1) &&
      graph->nChannels < system->nodes[NET].count && evenReferenceIsBetter)
    return ncclSuccess;
  ...
[Design inference and architectural trade-offs]

Why prefer even channels? Because the Ring algorithm can pair better with even channels—each channel can be split into two halves, one clockwise and one counterclockwise, reducing network congestion.

Ring and Tree: turning search results into algorithm topology

Intuitive model

The search algorithm finds a set of paths, but the algorithm needs an explicit "who sends to whom" order. Ring strings all ranks into a loop, where each rank receives from the previous and sends to the next. Tree is a tree, where data flows down from the root or converges up from the leaves.

Without these two modules, the search algorithm would just find a bunch of paths and couldn't tell the GPU kernel exactly how to send data.

Data structures and memory layout

Ring construction is done byncclBuildRings:

📎 src/graph/rings.cc:29-74

c
ncclResult_t ncclBuildRings(int nrings, int* rings, int rank, int nranks, int* prev, int* next) {
  ncclResult_t ret = ncclSuccess;
  uint64_t* rankFound;
  int rankFoundSize = DIVUP(nranks, 64);
  NCCLCHECK(ncclCalloc(&rankFound, rankFoundSize));

  for (int r = 0; r < nrings; r++) {
    int current = rank;
    for (int i = 0; i < nranks; i++) {
      rankFound[current / 64] |= (1ULL << (current % 64));
      rings[r * nranks + i] = current;
      current = next[r * nranks + current];
    }
    ...
    if (current != rank) {
      WARN("Error : ring %d does not loop back to start (%d != %d)", r, current, rank);
      ret = ncclInternalError;
      goto end;
    }
    // Check that all ranks are there
    for (int i = 0; i < nranks; i++) {
      uint64_t bits = rankFound[i / 64], mask = 1ULL << (i % 64);
      // Fast check 64 ranks at a time
      if (mask == 1 && bits == 0xffffffffffffffff) {
        i += 63;
        continue;
      }
      if ((bits & mask) == 0) {
        WARN("Error : ring %d does not contain rank %d", r, i);
        ret = ncclInternalError;
        goto end;
      }
    }
    memset(rankFound, 0, rankFoundSize * sizeof(uint64_t));
  }
end:
  free(rankFound);
  return ret;
}

The inputs areprevandnextarrays (each rank's predecessor and successor), and the output is theringsarray (the complete rank order for each channel). It starts from the current rank, follows thenextpointers around a full loop, verifies whether it returns to the starting point, and checks that all ranks are visited.

Tree construction is done byncclGetBtree:

📎 src/graph/trees.cc:32-67

c
ncclResult_t ncclGetBtree(int nranks, int rank, int* u, int* d0, int* d1, int* parentChildType) {
  int up, down0, down1;
  int bit;
  for (bit = 1; bit < nranks; bit <<= 1) {
    if (bit & rank) break;
  }

  if (rank == 0) {
    *u = -1;
    *d0 = -1;
    // Child rank is > 0 so it has to be our child 1, not 0.
    *d1 = nranks > 1 ? bit >> 1 : -1;
    return ncclSuccess;
  }

  up = (rank ^ bit) | (bit << 1);
  // if smaller than the parent, we are his first child, otherwise we're his second
  if (up >= nranks) up = (rank ^ bit);
  *parentChildType = (rank < up) ? 0 : 1;
  *u = up;

  int lowbit = bit >> 1;
  // down0 is always within bounds
  down0 = lowbit == 0 ? -1 : rank - lowbit;

  down1 = lowbit == 0 ? -1 : rank + lowbit;
  // Make sure down1 is within bounds
  while (down1 >= nranks) {
    down1 = lowbit == 0 ? -1 : rank + lowbit;
    lowbit >>= 1;
  }
  *d0 = down0;
  *d1 = down1;

  return ncclSuccess;
}

This function uses bit operations to build a binary tree. The core idea is: find the rank's lowest non-zero bitbit, the parent is(rank ^ bit) | (bit << 1), the left child isrank - (bit >> 1), and the right child isrank + (bit >> 1). The ASCII diagram in the comments clearly shows this structure.

Scenario-driven Step-by-Step Walkthrough

Take 8-GPU Ring as an example. Assume the search results give each rank'snextpointer:

code
rank 0 -> rank 1
rank 1 -> rank 2
...
rank 7 -> rank 0

ncclBuildRingsStarting from rank 0, visit 1, 2, ..., 7 in sequence, and finally return to 0. The generatedrings[0..7] = {0, 1, 2, 3, 4, 5, 6, 7}。

For Tree,ncclGetBtreecompute the parent node and child nodes for each rank. Taking rank 1 as an example:

  • bit= 1 (the lowest non-zero bit is bit 0)
  • up = (1 ^ 1) | (1 << 1) = 0 | 2 = 2
  • up >= nranks? 2 < 8, soup = 2
  • parentChildType = (1 < 2) ? 0 : 1 = 0(is the first child of the parent node)
  • lowbit = 0, sodown0 = -1
  • down1 = -1

Therefore, rank 1's parent node is rank 2, and it has no child nodes. This matches the tree structure in the comment: rank 1 is a leaf.

Design considerations and production pitfalls

[Design inferences and architectural trade-offs]

Why does Tree use bit operations instead of explicitly building a tree?Because each rank only needs to know its own parent node and child nodes, and does not need the global tree structure. Bit operations can compute this information in O(1) time, avoiding the overhead of storing and synchronizing the entire tree.

Pitfall one:ncclBuildRingsvalidation may be skipped.If thenextarray has a cycle (for example, rank 0 -> rank 1 -> rank 0), the loop will exit afternranksiterations, but thecurrent != rankcheck will catch this problem. However, if the cycle length happens to benranksa factor of and does not include all ranks,rankFoundthe check will catch it.

Pitfall two:ncclGetDtreehandling of odd ranks.For an odd number of ranks, the second tree is "shifted" rather than "mirrored":

📎 src/graph/trees.cc:90-112

c
ncclResult_t ncclGetDtree(int nranks, int rank, int* s0, int* d0_0, int* d0_1, int* parentChildType0, int* s1,
                          int* d1_0, int* d1_1, int* parentChildType1) {
  // First tree ... use a btree
  ncclGetBtree(nranks, rank, s0, d0_0, d0_1, parentChildType0);
  // Second tree ... mirror or shift
  if (nranks % 2 == 1) {
    // shift
    int shiftrank = (rank - 1 + nranks) % nranks;
    ...
  } else {
    // mirror
    int u, d0, d1;
    ncclGetBtree(nranks, nranks - 1 - rank, &u, &d0, &d1, parentChildType1);
    *s1 = u == -1 ? -1 : nranks - 1 - u;
    ...
  }
  return ncclSuccess;
}

Double Tree is NCCL's Tree algorithm implementation - two trees work simultaneously, one responsible for the first half of the data and one for the second half, improving bandwidth utilization. With an odd number of ranks, mirroring would cause an incomplete rank mapping, so shifting is used instead.

The coordination of the three: from topology to algorithm

Now connect the three modules together. The entire process can be represented by a diagram:

mermaid
flowchart TD
    A["ncclTopoGetSystem()"] --> B["解析 XML,创建节点"]
    B --> C["ncclTopoConnectNodes() 建立链路"]
    C --> D["ncclTopoComputePaths() 计算所有路径"]
    D --> E{"ncclTopoCompute() 搜索"}
    E -->|"Ring 模式"| F["ncclTopoSearchRecNet()"]
    E -->|"Tree 模式"| G["ncclTopoSearchRecNet()"]
    F --> H["ncclTopoSearchRecGpu() 递归搜索"]
    G --> H
    H --> I{"找到更优解?"}
    I -->|"是"| J["memcpy 保存到 saveGraph"]
    I -->|"否"| K["继续尝试其他路径"]
    J --> L["ncclBuildRings() 或 ncclGetDtree()"]
    K --> H
    L --> M["生成最终算法拓扑"]

This diagram shows the complete process from topology discovery to algorithm generation. Note thatncclTopoSearchRecGpuis a recursive function that continuously tries different GPU orders until it times out or finds the optimal solution.

Now look at a more fine-grained sequence diagram, showing the interaction of the modules during the search process:

mermaid
sequenceDiagram
    participant Init as ncclTopoCompute
    participant Search as ncclTopoSearchRec
    participant Net as ncclTopoSearchRecNet
    participant Gpu as ncclTopoSearchRecGpu
    participant Follow as ncclTopoFollowPath
    participant Compare as ncclTopoCompareGraphs

    Init->>Search: ncclTopoSearchRec(system, tmpGraph, graph, &time)
    Search->>Net: ncclTopoSearchRecNet(system, graph, saveGraph, backToNet, backToFirstRank, time)
    Net->>Net: ncclTopoSelectNets() 选择候选网卡
    Net->>Gpu: ncclTopoSearchTryGpu(..., NET, n, gpu)
    Gpu->>Follow: ncclTopoFollowPath(system, graph, NET, n, GPU, g, 1, &gpu)
    Follow-->>Gpu: 返回目标 GPU 节点
    Gpu->>Gpu: 递归 ncclTopoSearchRecGpu(step+1)
    Gpu->>Compare: ncclTopoCompareGraphs(system, graph, saveGraph, &copy)
    Compare-->>Gpu: copy=1 表示更优
    Gpu->>Gpu: memcpy(saveGraph, graph)
    Gpu->>Follow: ncclTopoFollowPath(..., -1, &gpu) 恢复带宽

This sequence diagram shows the core loop of the search: select NIC -> try GPU -> recursive search -> compare results -> restore bandwidth.

Chapter summary

This chapter breaks down the three stages of NCCL topology awareness:

1. Topology discovery(topo.cc): read device information from XML, create GPU/CPU/PCI/NIC nodes, establish NVLink/PCIe/network links, and form a complete topology graph.

2. Path search(search.cc + paths.cc): first use BFS to precompute the shortest paths between all node pairs, then use recursive search to try different Ring/Tree structures and find the solution with the highest bandwidth.

3. Algorithm topology generation(rings.cc + trees.cc): convert the search results into a specific rank order. Ring usesncclBuildRingsto generate the ring, and Tree usesncclGetBtreeto generate the binary tree.

Chapter review questions

Q1: If the bandwidth accumulation inncclTopoConnectNodeslink->bw += bwis changed tolink->bw = std::max(link->bw, bw), in what scenarios would this cause performance degradation? Why?

Reference analysis: Bandwidth accumulation handles the case of multiple parallel links. Taking 4 NVLinks at 25 GB/s each as an example, the accumulated value is 100 GB/s, while the max value is only 25 GB/s. InncclTopoSetPaths, the path bandwidth isstd::min(path->bw, link->bw). If the link bandwidth is underestimated, the bandwidth of the entire path will be underestimated. This will causencclTopoCompareGraphsto choose the wrong graph - it may choose a solution with more channels but lower bandwidth per channel, resulting in worse actual performance. Specific scenario: 8-GPU A100 fully interconnected with NVLink, with 4 NVLinks between each pair of GPUs. Accumulation gives 100 GB/s, while max gives 25 GB/s. The search algorithm will consider NVLink and PCIe Gen4 x16 (about 25 GB/s) to have the same bandwidth, and may choose a path through PCIe.

Q2: ncclTopoSearchRecGpuIn(*time)--is executed at the function entry. If the search times out (*time <= 0), the function returns directly. Under what circumstances would this design cause the search to fall into an infinite loop? How to fix it?

Reference analysis:(*time)--decrements at the entry. If*timethe initial value is 0 or negative, the function returns directly and will not decrement. But if*timeis a very large positive number, each recursion will decrement it, and it will eventually reach 0. The problem is: if a branch has a very large recursion depth, but after each decrement*timeis still greater than 0, the search will continue. The real risk isncclTopoSearchRecthegoto searchloop in - iftimeis not correctly reset in the loop, it may loop infinitely. Look atncclTopoComputetheglobalTimeoutlogic in:globalTimeout -= timeis executed at eachsearchlabel. IfglobalTimeoutbecomes negative, it willgoto done. But iftimeis reset toNCCL_SEARCH_TIMEOUT,globalTimeoutit may never become negative. The fix is to ensure thatglobalTimeoutis decremented after each search and has a hard upper limit.

Q3: ncclTopoFollowPathWhen the search fails,followPath(path, node1, step, -bw, &step)is called to restore bandwidth. If a recursive branch returns before restoration (for example,NCCLCHECKGOTOjumps toexit), what happens? How to detect this problem?

Reference analysis: If restoration is skipped, the link bandwidths along the path will remain in the deducted state. Subsequent searches will see incorrect bandwidth and may miss the optimal solution. Detection method: inncclTopoComputeAfter completion, traverse all links and check whether the bandwidth matches the initial value. If a mismatch is found, it indicates a recovery omission. Fix: use an RAII-style guard object to automatically restore bandwidth during destruction. Alternatively, save a bandwidth snapshot of all links before each search and restore it after the search. NCCL's current approach is to manually pair forward and reverse calls at eachncclTopoFollowPathcall site, which is error-prone. A more robust design is to encapsulate bandwidth deduction and restoration into a function, ensuring they occur in pairs.

In the next chapter, we will dive into the tuning module to see how NCCL makes the final choice among Ring, Tree, CollNet, and other algorithms based on topology search results and message size. The topology graph, path search results, and algorithm templates established in this chapter will become the input to the tuning module.

Through graph construction in topo.cc, path search in search.cc, and topology generation in rings.cc and trees.cc, NCCL realizes the design philosophy of describing arbitrary topologies with a general graph structure, finding the optimal solution with configurable search algorithms, and generating the final algorithm with simple templates. This mechanism allows NCCL to automatically select suitable algorithms on machines ranging from 2-GPU workstations to 10,000-GPU clusters. However, the topology graph only provides candidate paths for algorithms. Deciding which path to take and which protocol to use for a specific communication still requires more fine-grained decisions. In the next chapter, we will focus on the src/tuning directory to see how the tuning module combines cost models and algorithm estimates to make the final choice among Ring/Tree/NVLS/PAT and LL/LL128/Simple.

CHAPTER 05

Chapter 5: Chapter 5: Algorithm and Protocol Selection: How the tuning Module Determines the Communication Path

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 5 / 25

Chapter 5: Algorithm and Protocol Selection: How the tuning Module Determines the Communication Path

In the previous chapter, we broke down NCCL's topology-awareness capability: from enumerating devices in src/graph/topo.cc to build the topology graph, to searching for the optimal path in src/graph/search.cc, and then to rings.cc and trees.cc materializing the search results into Ring and Tree algorithm topologies. But the topology graph only answers "which paths data can take"; it does not answer "which path this communication should take." On the same machine, a 4KB AllReduce and a 400MB AllReduce may have completely different optimal solutions: the former competes on latency, while the latter competes on bandwidth; the former may choose Tree/LL, while the latter may choose Ring/Simple or NVLS. The tuning module is the one that "makes the call." Its inputs are message size, number of ranks, topology graph (the product of the previous chapter), and user environment variables; its output is an ncclTuningResult_t, which specifies which algorithm (algo) to use, which protocol (proto) to use, how many channels to open, and how many warps to use. In this chapter, we will break open the src/tuning directory in the order of "overall scheduling → cost model → estimates for each algorithm → final decision." There is only one core question: how does NCCL choose the fastest one among dozens of (algorithm, protocol) combinations using a purely CPU-based mathematical model within microseconds?

1. tuning.cc: Overall Scheduling and Decision Backbone

Intuitive Model

Imagine the tuning module as amoving company. A customer (one collective communication) arrives and says, "I want to move 100MB of goods from 8 warehouses to 8 warehouses." The dispatcher (ncclTuningCompute) will not actually try moving it once, but instead takes out aprice list(cost model), estimates an "expected time" for each option (Ring/LL, Tree/Simple, NVLS/Simple, ...), and then picks the shortest quote for the customer.

Without this dispatcher, NCCL could only hard-code "AllReduce always uses Ring," which would be crushed by Tree in small-message scenarios and by NVLS in large-scale NVLink scenarios.The cost is that performance is halved or even worse in specific scenarios.

Data Structures and Memory Layout

The carrier of the decision isncclTuningResult_t, and the candidate set isncclTuningResultList_t(a singly linked list). The linked list node is defined intuning_int.h, but the push logic is intuning.cc:

📎 src/tuning/tuning.cc:32-39

c
ncclResult_t ncclTuningResultListPushFront(struct ncclTuningResultList_t* list, struct ncclTuningResult_t result) {
  struct ncclTuningResultListNode* node = nullptr;
  NCCLCHECK(ncclCalloc(&node, 1));
  node->result = result;
  node->next = list->head;
  list->head = node;
  return ncclSuccess;
}
[Design Inference and Architectural Trade-offs]

Note that this ishead insertion: each time a valid candidate is computed, it is inserted at the head of the linked list. This means the linked list order and the id order arereversed. Why use a linked list instead of an array? Because the number of candidates is determined at compile time byNCCL_TUNING_COUNT, but the actually valid candidates are dynamic (affected bytuningMask, platform capabilities, and user environment variables). A linked list allows "only attaching the valid ones," avoiding repeated checks ofvalidduring traversal. The cost is that each decision requiresncclCalloconce, but tuning happens on the enqueue path and at low frequency, so this allocation overhead is acceptable.

ncclTuningResult_tThe two most critical fields intimeUs(estimated time, microseconds) andselectionTimeUs(the time used for selection, which may be overridden by a tuner plugin). The selection logic only looks at the latter:

📎 src/tuning/tuning.cc:155-173

c
static ncclResult_t ncclTuningSelectBestTuning(struct ncclTuningResultList_t* tunings,
                                               struct ncclTuningResult_t* const bestTuning) {
  bestTuning->timeUs = FLT_MAX;
  float bestSelectionTimeUs = FLT_MAX;
  struct ncclTuningResultListNode* node = tunings->head;
  while (node != nullptr) {
    const struct ncclTuningResult_t& tuning = node->result;
    float selectionTimeUs = tuning.selectionTimeUs > 0.0f ? tuning.selectionTimeUs : tuning.timeUs;
    ...
    if (selectionTimeUs < bestSelectionTimeUs) {
      *bestTuning = tuning;
      bestSelectionTimeUs = selectionTimeUs;
    }
    node = node->next;
  }
  return ncclSuccess;
}

There is a detail here:bestTuning->timeUsis first set toFLT_MAX, then iterated. If the linked list is empty (all candidates are invalid),bestTuningwill keepNCCL_TUNING_RESULT_INIT's initial value, and both algo/proto areUNDEF. This "empty result" is specially handled by the caller—see the error branch later.

Step-by-Step Walkthrough: The Decision Flow of a Single AllReduce

Suppose the application callsncclAllReduce, with a 1MB message and 8 ranks on a single-node NVLink. We followncclTuningComputethrough it once.

Step 0: Single-rank short circuit.IfnRanks <= 1, communication is not needed at all, and it directly returns Ring/Simple with the channel count set to 0:

📎 src/tuning/tuning.cc:191-200

c
  // Set tuning to Ring/Simple for single rank case
  if (input->comm->nRanks <= 1) {
    bestTuning.algo = NCCL_ALGO_RING;
    bestTuning.proto = NCCL_PROTO_SIMPLE;
    bestTuning.symKernelId = ncclSymkKernelId_Count;
    bestTuning.ceMethodId = ncclCeMethodId_Count;
    bestTuning.nChannels = 0;
    bestTuning.maxChannels = 0;
    bestTuning.nWarps = 0;
    bestTuning.forced = 0;
  } else {

This short circuit is important: with a single rank, any algorithm estimate will divide by quantities likenRanks-1, which can easily produce NaN or division by zero.Fallback first, then do the accounting, is a typical example of defensive programming.

Step 1: Enumerate all candidates.entersncclTuningComputeAllTunings, which iterates overNCCL_TUNING_COUNTids:

📎 src/tuning/tuning.cc:128-149

c
ncclResult_t ncclTuningComputeAllTunings(struct ncclTuningInput_t* const input,
                                         struct ncclTuningResultList_t* const tunings) {
  ncclResult_t ret = ncclSuccess;

  for (int i = 0; i < NCCL_TUNING_COUNT; i++) {
    struct ncclTuningResult_t tuning = NCCL_TUNING_RESULT_INIT;
    tuning.id = i;
    tuning.valid = 1;

    if (!(input->tuningMask & (1ULL << i))) {
      tuning.valid = 0;
      continue;
    }
    NCCLCHECK(ncclTuningExpandId(i, &tuning.algo, &tuning.proto, &tuning.symKernelId, &tuning.ceMethodId));
    NCCLCHECKGOTO(ncclTuningComputeTuning(i, input, &tuning), ret, fail);
    if (tuning.valid) NCCLCHECKGOTO(ncclTuningResultListPushFront(tunings, tuning), ret, fail);
  }
...
}

Note thattuningMaskis a 64-bit mask, where bit i indicates "whether the i-th (algo, proto) combination is allowed." This mask is computed at a higher layer based on platform capabilities, user environment variables, and function type.The mask is a "coarse filter," and the cost model is the "fine calculation"—first exclude what is fundamentally impossible (for example, NVLS cannot exist on a PCI machine), then compute time for the rest.

ncclTuningExpandIdexpands the one-dimensional id into (algo, proto, symKernelId, ceMethodId). This mapping must strictly match thecost_model.ccinmodelMaparray, otherwise the model will be computed incorrectly.

Step 2: Compute the cost one by one. ncclTuningComputeTuninghas only one line, delegating to the cost model:

📎 src/tuning/tuning.cc:339-343

c
ncclResult_t ncclTuningComputeTuning(int id, struct ncclTuningInput_t* const input,
                                     struct ncclTuningResult_t* const result) {
  NCCLCHECK(ncclTuningCostModelSimModel(id, input, result));
  return ncclSuccess;
}

Step 3: Tuner plugin intervention (optional).If the user has installed a tuner plugin (such as a custom tuner from some cloud vendors), NCCL packs thetimeUsof all candidates into a two-dimensional tablegeneralTable[algo][proto]and hands it to the plugin, letting the plugin override:

📎 src/tuning/tuning.cc:203-230

c
    if (input->comm->tuner != NULL) {
      float generalTable[NCCL_NUM_ALGORITHMS][NCCL_NUM_PROTOCOLS];
      for (int i = 0; i < NCCL_NUM_ALGORITHMS; i++) {
        for (int j = 0; j < NCCL_NUM_PROTOCOLS; j++) {
          generalTable[i][j] = NCCL_TUNING_IGNORE;
        }
      }
      struct ncclTuningResultListNode* node = tunings.head;
      while (node != nullptr) {
        const struct ncclTuningResult_t& tuning = node->result;
        node = node->next;
        if (tuning.algo == NCCL_ALGO_UNDEF || tuning.proto == NCCL_PROTO_UNDEF) continue;
        generalTable[tuning.algo][tuning.proto] = tuning.timeUs;
      }
      node = tunings.head;
      int nMaxChannels = 0;
      NCCLCHECKGOTO(input->comm->tuner->getCollInfo(input->comm->tunerContext, input->func, input->nBytes,
                                                    input->numPipeOps, (float**)generalTable, NCCL_NUM_ALGORITHMS,
                                                    NCCL_NUM_PROTOCOLS, input->regBuff, &nMaxChannels),
                    ret, exit);
      while (node != nullptr) {
        struct ncclTuningResult_t& tuning = node->result;
        node = node->next;
        if (tuning.algo == NCCL_ALGO_UNDEF || tuning.proto == NCCL_PROTO_UNDEF) continue;
        tuning.maxChannels = nMaxChannels;
        tuning.timeUs = generalTable[tuning.algo][tuning.proto];
      }
    }

HereNCCL_TUNING_IGNOREis a sentinel value indicating "this combination has not been computed/is not applicable." The plugin can modify only the cells it cares about, leaving other cells as IGNORE, and NCCL will skip them.

Step 4: Select the best.callsncclTuningSelectBestTuning, traversing the linked list to take the one with the smallestselectionTimeUs.

Step 5: Compute the channel count.After selecting the algorithm, it still needs to decide how many channels to open:

📎 src/tuning/tuning.cc:233-235

c
  if (bestTuning.algo != NCCL_ALGO_UNDEF && bestTuning.proto != NCCL_PROTO_UNDEF) {
    NCCLCHECKGOTO(ncclTuningGetChannels(input, &bestTuning), ret, exit);
  }

ncclTuningGetChannelsIntuning_int.h, the logic is to interpolate betweenminChannelsandmaxChannelsbased on message size and algorithm type. The channel count directly affects bandwidth: more channels mean higher parallelism, but the startup overhead of each channel is also greater.

Step 6: CTA Policy bias (NVLS priority).If the user has setNCCL_CTA_POLICY_EFFICIENCY, and the current operation is AllGather/ReduceScatter and the buffer is already registered, NCCL will try to change the result to NVLS:

📎 src/tuning/tuning.cc:240-257

c
  if (input->comm->tuner == NULL && (input->CTAPolicy & NCCL_CTA_POLICY_EFFICIENCY) &&
      ncclGetEnv("NCCL_ALGO") == NULL && ncclGetEnv("NCCL_PROTO") == NULL && !input->comm->MNNVL &&
      (input->tuningMask & (1ull << (NCCL_ALGO_NVLS * NCCL_NUM_PROTOCOLS + NCCL_PROTO_SIMPLE)))) {
    if (input->regBuff && (input->func == ncclFuncAllGather || input->func == ncclFuncReduceScatter)) {
      if ((input->comm->nNodes > 1 && input->collNetSupport && input->nvlsSupport) ||
          (input->comm->nNodes == 1 && input->nvlsSupport)) {
        int recChannels;
        NCCLCHECKGOTO(ncclNvlsRegResourcesQuery(input->comm, input->func, &recChannels), ret, exit);
        if (recChannels <= bestTuning.nChannels) {
          bestTuning.algo = NCCL_ALGO_NVLS;
          ...

The comment in this code is crucial:The EFFICIENCY bias must run afterGetChannels, because it needs to usebestTuning.nChannels; and it must check whether the NVLS bit intuningMaskis allowed, otherwise it will "revive" an algorithm that was excluded by the upper layer. This is a typicalstate-dependent ordering trap。

Step 7: Symmetric kernel fallback.If the selected one is a symmetric kernel (symKernelId), but the buffer is not registered or the platform does not support it, it needs to fall back to a normal kernel. This logic is intuning.cc:258-298, and it is the most convoluted part of the whole chapter, so we will discuss it specifically in Section 5.

Step 8: No solution error.If all candidates are invalid and both algo/proto are UNDEF, NCCL will print a WARN and return different error codes depending on whether the user has set environment variables:

📎 src/tuning/tuning.cc:308-329

c
  if ((bestTuning.algo == NCCL_ALGO_UNDEF || bestTuning.proto == NCCL_PROTO_UNDEF) &&
      bestTuning.symKernelId == ncclSymkKernelId_Count && bestTuning.ceMethodId == ncclCeMethodId_Count) {
    ...
    WARN("No algorithm/protocol nor symKernelId available for function %s with datatype %s.%s%s%s",
         ncclFuncToString(input->func), ncclDatatypeToString(input->datatype), ncclAlgoEnvStr, ncclProtoEnvStr,
         ncclSymKernelIdEnvStr);
    ret = (algoEnv || protoEnv || symKernelIdEnv) ? ncclInvalidUsage : ncclInternalError;
  }

Why distinguish the error codes?If the user has setNCCL_ALGO=ringbut the current platform does not support ring (for example, some special topologies), then that isa user configuration error(ncclInvalidUsage); if the user has not set any environment variables but still cannot select an algorithm, then that isan NCCL internal bug(ncclInternalError). This distinction is crucial for troubleshooting.

Main Decision Flowchart

mermaid
flowchart TD
    start["ncclTuningCompute(input)"] --> check_rank{"comm->nRanks <= 1?"}
    check_rank -->|是| single["bestTuning = Ring/Simple<br/>nChannels = 0"]
    check_rank -->|否| enum["ncclTuningComputeAllTunings<br/>遍历 NCCL_TUNING_COUNT"]
    enum --> mask{"tuningMask & (1<<i)?"}
    mask -->|否| skip["tuning.valid = 0<br/>continue"]
    mask -->|是| expand["ncclTuningExpandId(i)"]
    expand --> sim["ncclTuningComputeTuning<br/>-> ncclTuningCostModelSimModel"]
    sim --> valid{"result.valid?"}
    valid -->|是| push["ncclTuningResultListPushFront"]
    valid -->|否| skip
    push --> tuner{"comm->tuner != NULL?"}
    tuner -->|是| plugin["tuner->getCollInfo<br/>覆盖 generalTable"]
    tuner -->|否| select
    plugin --> select["ncclTuningSelectBestTuning<br/>取 selectionTimeUs 最小"]
    select --> getch["ncclTuningGetChannels"]
    getch --> cta{"CTA_POLICY_EFFICIENCY<br/>且 NVLS 在 mask 内?"}
    cta -->|是| nvls["ncclNvlsRegResourcesQuery<br/>可能改写为 NVLS"]
    cta -->|否| symk
    nvls --> symk{"symKernelId 需要回退?"}
    symk -->|是| fallback["ncclTuningCompute(generalInput)<br/>回退普通 kernel"]
    symk -->|否| done
    fallback --> done["*result = bestTuning"]
    single --> done
    done --> undef{"algo/proto 仍 UNDEF?"}
    undef -->|是| warn["WARN + 返回<br/>InvalidUsage 或 InternalError"]
    undef -->|否| ret_ok["返回 ncclSuccess"]

---

2. cost_model.cc: Model Registry and Switch Matrix

Intuitive model

cost_model.ccis tuning'sgeneral ledger. It maintains amodelMaptable, where each row corresponds to an (algo, proto) combination and records "who this combination's initialization function is, who its simulation function is, and for which functions it is enabled." At the same time, it is responsible for parsing the user environment variableNCCL_ALGO/NCCL_PROTO/NCCL_SYM_KERNEL, translating the user's intent into aenabled[i][f]switch matrix.

Without this table, every time a new algorithm is added, the main tuning flow would have to be modified again, and the code would rot into a mess.Table-drivenmakes "adding an algorithm" become "adding a row."

Data structure: modelMap and switch matrix

modelMapis a static array, and each element isncclTuningModelEntry_t:

📎 src/tuning/cost_model.cc:230-277

c
static struct ncclTuningModelEntry_t modelMap[] = {
  {ncclTuningTreeModelInit, ncclTuningTreeModelSim, nullptr, {0, 0, 0, 0, 1}},       // Tree/LL
  {ncclTuningTreeModelInit, ncclTuningTreeModelSim, nullptr, {0, 0, 0, 0, 1}},       // Tree/LL128
  {ncclTuningTreeModelInit, ncclTuningTreeModelSim, nullptr, {0, 0, 0, 0, 1}},       // Tree/Simple
  {ncclTuningRingModelInit, ncclTuningRingModelSim, nullptr, {1, 1, 1, 1, 1}},       // Ring/LL
  ...
  {nullptr, nullptr, nullptr, {0}}, // CollNetDirect/LL, disabled as there is no implementation
  ...
};

Each entry has four fields:init(initialization, computing latency/bandwidth and storing them in comm),model(simulation, computing the final timeUs based on message size),finalize(cleanup),enabled[5](whether the five functions Broadcast/Reduce/AllGather/ReduceScatter/AllReduce are enabled).

NoteenabledThe order of the array is annotated at L234:Enable order: Broadcast, Reduce, AllGather, ReduceScatter, AllReduce. This order must match thencclFunc_tenum, otherwise things will get mixed up.

[Design Inference and Architectural Trade-offs]

Why should init and sim be separated?Because the things computed in init (latency, bandwidth)only depend on the static properties of comm(topology, number of ranks, compCap), and are independent of the specific message size. Within a single communication, tuning may be called multiple times in succession (for example, when a group has multiple ops), init runs only once, while sim runs every time. This is a typical "precompute + fast lookup" optimization.

Step-by-Step: Environment Variable Parsing and Switch Matrix Construction

Step 1: All enabled by default, LL128 is special. ncclTuningCostModelInitInitially all protos are set to 1 (enabled), but LL128 is set to 2:

📎 src/tuning/cost_model.cc:313-323

c
  for (int f = 0; f < NCCL_NUM_FUNCTIONS; f++) {
    for (int p = 0; p < NCCL_NUM_PROTOCOLS; p++) {
      protoEnable[f * NCCL_NUM_PROTOCOLS + p] = p == NCCL_PROTO_LL128 ? 2 : 1;
    }
    for (int a = 0; a < NCCL_NUM_ALGORITHMS; a++) {
      algoEnable[f * NCCL_NUM_ALGORITHMS + a] = 1;
    }
    for (int k = 0; k < ncclSymkKernelId_Count; k++) {
      symKernelIdEnable[f * ncclSymkKernelId_Count + k] = 1;
    }
  }

Why is LL128 2 instead of 1?Because LL128 is not "enabled by default", but "conditionally enabled". 2 is a special marker indicating "the user did not explicitly request it, and it will later be determined byisLL128Enabledbased on platform capabilities". 1 means "unconditionally enabled", 0 means "disabled". This tri-state design is reflected in the check at L366:

📎 src/tuning/cost_model.cc:364-370

c
      // Disable LL128 when 1) it is not supported on the platform, and 2) user did not explicitly request it.
      // protoEnable[..] == 2 indicates that user did not set NCCL_PROTO=LL128 explicitly.
      if (proto == NCCL_PROTO_LL128 && protoEnable[f * NCCL_NUM_PROTOCOLS + proto] == 2 &&
          !isLL128Enabled(comm->minCompCap, comm->maxCompCap, comm->graphs[algo].typeInter,
                          comm->graphs[algo].typeIntra, comm->nRanks, f, algo, comm->minDriverVersion)) {
        comm->tuningContext.enabled[i][f] = 0;
      }

Step 2: Parse user environment variables.If the user setsNCCL_ALGOorNCCL_SYM_KERNEL, first clear algo and symKernel entirely (because the user has specified a whitelist):

📎 src/tuning/cost_model.cc:327-345

c
  if ((algoStr && strlen(algoStr) > 0) || (symKernelIdStr && strlen(symKernelIdStr) > 0)) {
    std::fill_n(algoEnable, NCCL_NUM_FUNCTIONS * NCCL_NUM_ALGORITHMS, 0);
    std::fill_n(symKernelIdEnable, NCCL_NUM_FUNCTIONS * ncclSymkKernelId_Count, 0);
  }
  if (protoStr) {
    INFO(NCCL_ENV, "NCCL_PROTO set by environment to %s", protoStr);
    NCCLCHECK(parseList(protoStr, ncclFuncStr, NCCL_NUM_FUNCTIONS, ncclProtoStr, NCCL_NUM_PROTOCOLS, protoEnable,
                        comm->tuningContext.forced));
  }

Note that proto is not cleared—because proto's default value is 1/2, when the user setsNCCL_PROTO=LL,parseListwill set LL to 1 and others to 0 (because of theunsetlogic). This asymmetry is intentional: algo is fully enabled by default but must be narrowed after the user specifies it, while proto's narrowing is handled internally byparseList.

Step 3: The syntax of parseList.This function supports fairly complex syntax, and the comments give examples:

📎 src/tuning/cost_model.cc:14-32

c
// Parse a map of prefixes to a list of elements. The first prefix is
// optional and, if not present, the list of elements will be applied
// to all prefixes. Only the first list of elements can lack a
// prefix. Prefixes (if present) are followed by a colon. Lists of
// elements are comma delimited. Mappings of prefix to the lists of
// elements are semi-colon delimited.
//
// For example:
//
//     NCCL_ALGO="ring,collnetdirect;allreduce:tree,collnetdirect;broadcast:ring"
// Enable ring and collnetdirect for all functions, then select tree
// and collnetdirect for allreduce and ring for broadcast.

^The prefix means "negation":

📎 src/tuning/cost_model.cc:59-67

c
    int unset, set;
    if (elemList[0] == '^') {
      unset = 1;
      set = 0;
      elemList++;
    } else {
      unset = 0;
      set = 1;
    }

SoNCCL_PROTO="^LL128;allreduce:LL128"means: globally disable LL128, but enable LL128 as an exception for AllReduce.

Step 4: Merge the enabled matrix.Finally, iterate over all models and ANDmodel->enabled[f]with the user switches:

📎 src/tuning/cost_model.cc:371-383

c
      //  Check the user env vars only for functions that have a forced configuration and not already disabled.
      if (comm->tuningContext.forced[f] == 0 || comm->tuningContext.enabled[i][f] == 0) continue;
      comm->tuningContext.enabled[i][f] = 0;
      ...
      if (((algo != NCCL_ALGO_UNDEF && algoEnable[f * NCCL_NUM_ALGORITHMS + algo] != 0) &&
           (proto != NCCL_PROTO_UNDEF && protoEnable[f * NCCL_NUM_PROTOCOLS + proto] != 0)) ||
          (symKernelId != ncclSymkKernelId_Count && symKernelIdEnable[f * ncclSymkKernelId_Count + symKernelId] != 0)) {
        comm->tuningContext.enabled[i][f] = 1;
      }

The logic is:Only when the user has set a forced configuration for some function is the user configuration used to override the model default value. If the user has not set it,forced[f] == 0, directlycontinue, retaining the model's ownenabled. This is the priority of "user explicit specification > model default".

Unified entry point for model simulation

All models are ultimately invoked throughncclTuningCostModelSimModel:

📎 src/tuning/cost_model.cc:470-497

c
ncclResult_t ncclTuningCostModelSimModel(int id, struct ncclTuningInput_t* const input,
                                         struct ncclTuningResult_t* const result) {
  struct ncclTuningModelEntry_t* model = nullptr;
  ncclResult_t ret = ncclSuccess;
  result->forced = input->comm->tuningContext.forced[input->func];
  NCCLCHECKGOTO(getModelEntry(id, &model), ret, not_valid);
  if (model == nullptr) {
    ret = ncclInternalError;
    goto not_valid;
  }
  if (input->comm->tuningContext.enabled[id][input->func] == 0) {
    goto not_valid;
  }
  if (model->model != nullptr) {
    NCCLCHECKGOTO(model->model(input, result), ret, not_valid);
    if (result->timeUs <= 0.0) {
      goto not_valid;
    }
  } else {
    goto not_valid;
  }
exit:
  return ret;
not_valid:
  result->timeUs = NCCL_TUNING_IGNORE;
  result->valid = 0;
  goto exit;
}

Three layers of filtering:id out of bounds → model disabled → model returns non-positive time, if any layer fails, go tonot_valid, settimeUstoNCCL_TUNING_IGNORE(a negative sentinel),valid = 0. When the caller seesvalid == 0, it will not attach it to the candidate linked list.

Design considerations

modelMapThere is a key warning in the comments of

📎 src/tuning/cost_model.cc:229

c
// IMPORTANT: this table need must be consistent with the algRegistry in src/config/algorithm_registry.cc
[Design Inference and Architectural Trade-offs]

This means that themodelMapofindex ordermust strictly match the algorithm registration order inalgorithm_registry.cc. If someone inserts a new algorithm into the registry but forgets to changemodelMap, all ids will be misaligned, and tuning will select a completely wrong algorithm.This is the classic trap of table-driven design: implicit contracts.A more robust approach would be to use enum names as keys instead of indices, but that would sacrifice a bit of compile-time optimization.

---

III. ring.cc: Cost Estimation for the Ring Algorithm

Intuitive model

The Ring algorithm arranges N ranks into a ring, and data is passed around the ring circle by circle. Its cost model must answer two questions:How much data is passed per step (bandwidth)、How many steps are needed in total (latency)。

The intuition of Ring is "pipeline": imagine N people standing in a circle passing a bucket, and each person pours a little water into it before passing it to the next person. After the bucket goes around once, everyone's water is mixed evenly. The faster the bucket goes around (higher bandwidth) and the smaller the circle (fewer steps), the faster the whole thing is.

Data structure: latency/bandwidth table

The Ring model does not introduce new structures; it writes the estimation results intocomm->tuningContext.generalLatencies[c][algo][proto]andgeneralBandwidths[c][algo][proto]. These two are three-dimensional arrays: function × algorithm × protocol.

During initialization, all are first set to -1.0 (sentinel, meaning "not computed yet"):

📎 src/tuning/ring.cc:31-33

c
  for (int c = 0; c < NCCL_NUM_FUNCTIONS; c++) {
    comm->tuningContext.generalLatencies[c][algo][proto] = -1.0;
    comm->tuningContext.generalBandwidths[c][algo][proto] = -1.0;

The -1.0 sentinel is checked during the sim stage:

📎 src/tuning/ring.cc:94-97

c
  if (inputs->comm->tuningContext.generalBandwidths[inputs->func][tuning->algo][tuning->proto] == -1.0f) {
    tuning->valid = 0;
    return ncclSuccess;
  }

Why use -1.0 instead of 0?Because 0 is a legal bandwidth value (though physically impossible), while -1.0 clearly indicates "uninitialized". Using==for floating-point comparison is safe here, because -1.0 is exactly representable.

Step-by-Step: Ring Bandwidth Estimation

Step 1: Determine whether to use intra or inter bandwidth.Single node (nNodes==1) uses intra, multi-node uses inter:

📎 src/tuning/ring.cc:34-37

c
    int nSteps = ncclTuningGetNsteps(c, comm->nRanks);
    float bw = (comm->nNodes == 1 || (comm->nNodes <= 2 && comm->minCompCap < 100)) ? comm->graphs[algo].bwIntra :
                                                                                      comm->graphs[algo].bwInter;
    float busBw = bw * comm->graphs[algo].nChannels;

nStepsis the number of steps required by the algorithm; for Ring, AllReduce is2*(nRanks-1), others arenRanks-1。busBwis the "bus bandwidth" = single-link bandwidth × number of channels.

Step 2: Apply protocol discount.The LL protocol only uses half the bandwidth (because of LL's flag overhead), while LL128 uses 92% (120/128):

📎 src/tuning/ring.cc:38-42

c
    if (proto == NCCL_PROTO_LL) {
      busBw = std::min(llMaxBw, busBw * .5);
    }
    if (proto == NCCL_PROTO_LL128)
      busBw = std::min(busBw * (0.92 /*120.0/128.0*/), comm->graphs[algo].nChannels * perChMaxRingLL128Bw);

0.92 = 120/128This is because in LL128, 8 out of every 128 bytes are flags, leaving only 120 bytes of payload. This number comes directly from the protocol design.

Step 3: Calculate effective bandwidth.Note that here it is multiplied bynRanks / nSteps:

📎 src/tuning/ring.cc:44-46

c
    comm->tuningContext.generalLatencies[c][algo][proto] =
      comm->tuningContext.tuningConstants.baseLatencies[algo][proto];
    comm->tuningContext.generalBandwidths[c][algo][proto] = busBw * comm->nRanks / nSteps;

Why multiply bynRanks / nSteps?This is the core characteristic of the Ring algorithm: the amount of data each rank actually moves isnBytes * nSteps / nRanks(because the data must go around the ring multiple times). So "effective bandwidth" = bus bandwidth × nRanks / nSteps. For AllReduce, nSteps = 2(nRanks-1), so effective bandwidth ≈ busBw/2.

Step 4: Calculate latency.Latency is split into two parts: intra and inter:

📎 src/tuning/ring.cc:48-63

c
    int intraHw, interHw;
    ncclTuningGetHwIndexes(comm, algo, &intraHw, &interHw);
    int hwLevel = comm->nNodes == 1 ? intraHw : interHw;

    float intraLat = comm->tuningContext.tuningConstants.hwLatencies[intraHw][algo][proto];
    // Preserve the pre-refactor model: with one rank per node, Ring inter-node steps use the exposed Tree NET latency.
    float interLat;
    if (comm->nNodes == 1) {
      interLat = intraLat;
    } else if (comm->maxLocalRanks == 1) {
      interLat = comm->tuningContext.tuningConstants.hwLatencies[NCCL_HW_NET][NCCL_ALGO_TREE][proto];
    } else {
      interLat = comm->tuningContext.tuningConstants.hwLatencies[interHw][algo][proto];
    }
    interLat += comm->graphs[algo].latencyInter;
    if (proto == NCCL_PROTO_SIMPLE) interLat += comm->graphs[algo].latencyInter;

Note the special handling at L57-58: whenmaxLocalRanks == 1(each node has only 1 rank), Ring's inter-node latency usesTree's NET latency. The comment says this is to "preserve the pre-refactor model" — that is, a deliberate "quirk" kept to maintain consistency with pre-refactor behavior.This kind of historical baggage is very common in mature systems. When reading source code and seeing the word "preserve," be especially careful; it often means there is a compatibility constraint here that cannot be changed.

Step 5: Accumulate by function type.The latency models for Reduce/Broadcast and AllReduce/AllGather/ReduceScatter are different:

📎 src/tuning/ring.cc:65-87

c
    if ((c == ncclFuncReduce || c == ncclFuncBroadcast)) {
      float lat = comm->tuningContext.tuningConstants.hwLatencies[hwLevel][algo][proto];
      if (comm->graphs[algo].sameChannels) {
        comm->tuningContext.generalLatencies[c][algo][proto] += lat;
      } else {
        if (proto == NCCL_PROTO_SIMPLE)
          lat =
            comm->tuningContext.tuningConstants
              .hwLatencies[hwLevel][NCCL_ALGO_TREE][proto]; // Add some chunk latency, waiting for proper chunk modeling
        comm->tuningContext.generalLatencies[c][algo][proto] += nSteps * lat;
      }
    } else {
      // Inter-node rings still have to launch nsteps * net overhead.
      float netOverhead = 0.0;
      if (comm->nNodes > 1) {
        netOverhead = getNetOverhead(comm);
        if (proto == NCCL_PROTO_SIMPLE) netOverhead *= 3;
      }
      intraLat = std::max(intraLat, netOverhead);
      int nInterSteps = comm->nNodes == 1 ? 0 : c == ncclFuncAllReduce ? 2 * (comm->nNodes - 1) : comm->nNodes - 1;
      comm->tuningContext.generalLatencies[c][algo][proto] +=
        (nSteps - nInterSteps) * intraLat + nInterSteps * interLat;
    }

sameChannelsis a topology property indicating "whether the intra and inter steps on the ring use the same set of channels." If they are different, latency must be multiplied bynSteps(each step must wait).netOverheadis the network post overhead. For the Simple protocol, multiply by 3 (because Simple has three network round trips: send, recv, ack).

Production pitfall avoidance: the plateau effect of Ring/Simple

ncclTuningRingModelSimThere is a section of code specifically handling "plateau":

📎 src/tuning/ring.cc:105-137

c
  // Update Ring/Simple latency for multi-node AllReduce and
  // single NVL Domain AllReduce/AllGather/ReduceScatter for Blackwell
  bool isBlackwellNvLink =
    inputs->comm->minCompCap >= 100 && inputs->comm->graphs[NCCL_ALGO_RING].typeIntra == PATH_NVL;
  bool ringSimplePlateau =
    (inputs->comm->nNodes > 1 && inputs->func == ncclFuncAllReduce) ||
    (inputs->comm->nNodes == 1 && isBlackwellNvLink &&
     (inputs->func == ncclFuncAllReduce || inputs->func == ncclFuncAllGather || inputs->func == ncclFuncReduceScatter));
  size_t bytesPerRankPerChannel = inputs->nBytes / (inputs->comm->nChannels * inputs->comm->nRanks);

  if (tuning->algo == NCCL_ALGO_RING && tuning->proto == NCCL_PROTO_SIMPLE && ringSimplePlateau &&
      bytesPerRankPerChannel >= 64) {
    float plateauFactor = inputs->comm->minCompCap < 80 ? 1.9 : 1.4;
    ...
    lat *= plateauFactor; // Plateau effect of ring
  }
[Design inference and architectural trade-offs]

What is plateau?In Ring/Simple, when the message becomes large enough, latency no longer grows linearly with message size but instead "gets stuck" on a plateau — because at this point the bottleneck shifts from "startup overhead" to "bandwidth," and bandwidth is already saturated. This phenomenon is especially obvious on Blackwell NVLink (because NVLink bandwidth is so high that latency accounts for a larger proportion). The code usesplateauFactor(1.4 or 1.9) multiplied onto the latency to simulate this "latency amplification" effect.

bytesPerRankPerChannel >= 64is the trigger condition: each rank must transfer at least 64 bytes per channel, otherwise the plateau does not hold. This 64 bytes comes from the LL protocol's flag size.

Pitfall scenario: If you run a 1MB AllReduce on Blackwell and find that the actual latency is 40% higher than the model predicts, do not assume it is a bug — this is the plateau effect, and the model has already accounted for it. If you manually reduceplateauFactor, the model will underestimate latency, leading to the wrong algorithm being selected.

---

IV. tree.cc and nvls.cc: Cost estimation for Tree and NVLS

Intuitive model

The Tree algorithmis "tree broadcast": the root node distributes data to child nodes, and child nodes then distribute it to grandchild nodes. Its advantage isfewer steps(log N instead of N), making it suitable for small messages; its disadvantage islow bandwidth utilization(each non-leaf node must forward, so the actual effective bandwidth is only half).

NVLS(NVLink SHARP) is "hardware multicast": the switch directly copies data to multiple GPUs without software forwarding. Its advantages arehigh bandwidth and low latency, but it requires specific hardware (Hopper or above) and specific configuration.

Tree model: only serves AllReduce

The Tree model has a hard restriction —it is only enabled for AllReduce:

📎 src/tuning/tree.cc:21-27

c
  for (int c = 0; c < NCCL_NUM_FUNCTIONS; c++) {
    if (c != ncclFuncAllReduce) {
      comm->tuningContext.generalLatencies[c][algo][proto] = -1.0;
      comm->tuningContext.generalBandwidths[c][algo][proto] = -1.0;
      enabled[c] = 0; // Hard disable
      continue;
    }
[Design inference and architectural trade-offs]

Why?Because NCCL's Tree implementation only supports AllReduce (other collective operations do not have Tree versions). This is an implementation constraint, not a theoretical limitation.enabled[c] = 0is a "hard disable," more thorough thangeneralBandwidths = -1— the former directly makesncclTuningCostModelSimModelreturn at L480not_valid, while the latter only checks inside the sim function.

Tree bandwidth estimation:

📎 src/tuning/tree.cc:28-43

c
    float bw = (comm->minCompCap < 100) ?
                 ((comm->nNodes <= 2) ? comm->graphs[algo].bwIntra : comm->graphs[algo].bwInter) :
                 std::min(comm->graphs[algo].bwInter, comm->graphs[algo].bwIntra);
    float busBw = bw * comm->graphs[algo].nChannels;
    if (c == ncclFuncAllReduce) busBw = std::min(busBw * .92, comm->graphs[algo].nChannels * perChMaxTreeBw);
    if (proto == NCCL_PROTO_LL) {
      busBw = std::min(busBw * 1.0 / 3.8, llMaxBw);
    }
    if (proto == NCCL_PROTO_LL128)
      busBw = std::min(busBw * (comm->nNodes == 1 ? 7.0 / 9.0 : 120.0 / 128.0),
                       comm->graphs[algo].nChannels * perChMaxTreeLL128Bw);
    if (comm->maxTreePattern == NCCL_TOPO_PATTERN_TREE) busBw *= .85;
[Design inference and architectural trade-offs]

Note that the LL protocol's discount factor is1/3.8, which is even more aggressive than Ring's0.5.Why is Tree's LL efficiency lower?Because each intermediate node in Tree must both receive and send, and LL's flag overhead is amplified under bidirectional traffic.1/3.8This number comes from measurement.

Tree latency estimation:

📎 src/tuning/tree.cc:55-58

c
    if (c == ncclFuncAllReduce) {
      comm->tuningContext.generalLatencies[c][algo][proto] +=
        2 * ((comm->nRanks / comm->nNodes - 1) * intraLat + log2i(comm->nNodes) * interLat);
    }

2 *is because AllReduce = ReduceScatter + AllGather, two passes.(nRanks/nNodes - 1)is the intra-node step count (the number of ranks within each node minus one),log2i(nNodes)is the inter-node step count (the height of the tree).

Tree's correction factor: The Tree model is multiplied by a factor during the sim stagetreeCorrectionFactor:

📎 src/tuning/tree.cc:75-79

c
  int logSize = log2i(inputs->nBytes >> 6);
  float bw = inputs->comm->tuningContext.generalBandwidths[inputs->func][tuning->algo][tuning->proto];
  float lat = inputs->comm->tuningContext.generalLatencies[inputs->func][tuning->algo][tuning->proto];
  if (inputs->func == ncclFuncAllReduce && logSize >= 0 && logSize < 23)
    bw *= treeCorrectionFactor[tuning->proto][logSize];

treeCorrectionFactoris a 3×24 table:

📎 src/tuning/cost_model.cc:223-227

c
float treeCorrectionFactor[NCCL_NUM_PROTOCOLS][24] = {
  {1.0, 1.0, 1.0, 1.0, .9, .8, .7, .7, .7, .7, .6, .5, .4, .4, .5, .6, .7, .8, .9, 1.0, 1.0, 1.0, 1.0, 1.0},
  {1.0, 1.0, 1.0, 1.0, 1.0, .9, .8, .8, .8, .7, .6, .6, .6, .6, .6, .6, .8, .9, .9, .9, .9, 1.0, 1.0, 1.0},
  {.9, .9, .9, .9, .9, .9, .9, .8, .7, .6, .6, .5, .5, .5, .5, .6, .7, .8, .7, .7, .8, .9, .9, .9}
};

logSize = log2(nBytes >> 6), i.e., the message size is taken as log2 in units of 64 bytes. The table indices 0-23 correspond to 64B to 64B×2^23 ≈ 512MB.This table is a measured "Tree efficiency curve": For small messages, efficiency is 1.0 (latency-dominated); for medium messages, efficiency drops to 0.4-0.5 (bandwidth not saturated); for large messages, it returns to 1.0 (bandwidth saturated). This "mid-range dip" is an inherent characteristic of the Tree algorithm.

NVLS model: The cost of hardware multicast

The NVLS model first checks whether the hardware supports it:

📎 src/tuning/nvls.cc:19-24

c
ncclResult_t ncclTuningNvlsModelInit(struct ncclComm* comm, int id, int enabled[NCCL_NUM_FUNCTIONS]) {
  ncclResult_t ret = ncclSuccess;
  if (!ncclNvlsTransportEnabled(comm)) {
    memset(enabled, 0, NCCL_NUM_FUNCTIONS * sizeof(int));
    return ncclSuccess;
  }

Then there is a series of hard constraints: only the Simple protocol is supported, NVLSTree is not supported on a single node, and multi-node NVLS requires CollNet:

📎 src/tuning/nvls.cc:28-41

c
  if ((algo == NCCL_ALGO_NVLS || algo == NCCL_ALGO_NVLS_TREE) && (proto != NCCL_PROTO_SIMPLE)) {
    memset(enabled, 0, NCCL_NUM_FUNCTIONS * sizeof(int));
    return ncclSuccess;
  }

  if (comm->nNodes == 1 && algo == NCCL_ALGO_NVLS_TREE) {
    memset(enabled, 0, NCCL_NUM_FUNCTIONS * sizeof(int));
    return ncclSuccess;
  }

  if (comm->config.collnetEnable == 0 && algo == NCCL_ALGO_NVLS && comm->nNodes > 1) {
    memset(enabled, 0, NCCL_NUM_FUNCTIONS * sizeof(int));
    return ncclSuccess;
  }

NVLS bandwidth estimationuses an efficiency factor:

📎 src/tuning/nvls.cc:12-17

c
static const float nvlsEfficiency[NCCL_NUM_COMPCAPS] = {
  0.0f, // Volta
  0.0f, // Ampere
  0.85f, // Hopper
  0.74f, // Blackwell
};
〔Design inference and architectural trade-offs〕

Hopper is 0.85, while Blackwell actually drops to 0.74.Why is the efficiency lower on the newer generation of hardware?Because Blackwell's NVLink bandwidth is higher, but the NVLS switch processing capability has not increased proportionally, resulting in a relative efficiency decrease. This number is measured, not theoretical.

In the bandwidth calculation there is a(nChannels - 1) / nChannelsfactor:

📎 src/tuning/nvls.cc:62-74

c
    int nSteps = ncclTuningGetNsteps(c, comm->nRanks);
    float intraBw = comm->graphs[algo].bwIntra * nvlsEfficiency[compCapIndex] * (comm->graphs[algo].nChannels - 1) /
                    comm->graphs[algo].nChannels;
    if (c == ncclFuncAllReduce) {
      intraBw *= 2.0f;
    } else {
      float ppn = comm->minLocalRanks;
      intraBw *= (ppn - 1) / ppn;
    }
    float interBw = comm->graphs[algo].bwInter * ((comm->nNodes <= 2 && algo == NCCL_ALGO_NVLS_TREE) ? 2 : 1);
    bw = std::min({intraBw, interBw,
                   algo == NCCL_ALGO_NVLS_TREE ? (float)perChMaxNVLSTreeBw : std::numeric_limits<float>::max()});
    bw = bw * comm->graphs[algo].nChannels;

(nChannels - 1) / nChannelsbecause NVLS needs to reserve one channel for synchronization.(ppn - 1) / ppnis the additional overhead of AllGather/ReduceScatter (each rank has to wait for the previous rank's data).

Production pitfalls: NVLS hard constraints

The NVLS model also has a runtime check during the sim stage:

📎 src/tuning/nvls.cc:136-156

c
  int nvlsSupport = inputs->nvlsSupport;
  if (!nvlsSupport) {
    tuning->valid = 0;
    tuning->timeUs = -1.0;
    return ret;
  }
  if (inputs->func != ncclFuncAllReduce && inputs->comm->graphs[tuning->algo].nChannels > NCCL_MAX_NVLS_ARITY) {
    tuning->valid = 0;
    tuning->timeUs = -1.0;
    return ret;
  }
  if (inputs->func != ncclFuncAllReduce && inputs->comm->localRanks > NCCL_MAX_NVLS_ARITY) {
    tuning->valid = 0;
    tuning->timeUs = -1.0;
    return ret;
  }

NCCL_MAX_NVLS_ARITYis the maximum number of GPUs that an NVLS multicast group can accommodate. If this number is exceeded, NVLS is unavailable.Pitfall scenario: Running AllGather in a 16-GPU NVLink domain, ifNCCL_MAX_NVLS_ARITYis 8, NVLS will be disabled and tuning will fall back to Ring. If you don't know this limitation, you'll wonder, "NVLS is clearly supported by the hardware, why isn't it being used?"

---

V. Symmetric kernel fallback and error recovery chain

Intuitive model

Symmetric kernel is a new NCCL feature: when all ranks' buffers are registered to symmetric memory, the kernel can use more efficient instructions to access peer memory. Butif the buffer is not registered, or the platform does not support it, it must fall back to the normal kernel. This fallback logic is the most convoluted part of tuning.

Step-by-Step: Fallback decision

The fallback logic is intuning.cc:258-298. Let's break it down.

Step 1: Determine whether fallback is needed.Entry conditions:

📎 src/tuning/tuning.cc:258-263

At this point, the decision chain of the tuning module is clear: it receives the topology graph and communication parameters, and through cost models and algorithm estimation, outputs the optimal (algorithm, protocol, channel, warp) combination within microseconds. But selection is only the beginning—how is this decision result used downstream? In the next chapter, we will enter the main trunk of src/enqueue/enqueue.cc and see how a single ncclAllReduce call goes through parameter validation, algorithm/protocol determination, and channel partitioning, ultimately generating the ncclInfo and ncclTaskColl structures. This is the key chapter in the book where we switch from the "user perspective" to the "engine perspective." You will discover what a collective communication call is translated into on the host side, and the boundary between it and subsequent kernel launch.

CHAPTER 06

Chapter 6: Chapter 6: Operator dispatch panorama: How ncclAllReduce becomes an executable kernel task

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 6 / 25

Chapter 6: Operator dispatch panorama: How ncclAllReduce becomes an executable kernel task

In the previous chapter, we walked through the tuning module and learned that NCCL selects an (algorithm, protocol, channel, warp) combination for a collective communication within microseconds. But the selection result itself is just a bunch of numbers—it needs to be "translated" into a task description object that the GPU kernel can understand before it can actually be executed. This chapter enters the main body of src/enqueue/enqueue.cc and answers a core question: when the user calls ncclAllReduce, what exactly happens on the host side? From ncclAllReduce to ncclEnqueueCheck, it goes through parameter validation, algorithm/protocol determination, and channel partitioning, ultimately generating the ncclInfo and ncclTaskColl structures. This is the key chapter where the book switches from the "user perspective" to the "engine perspective." If NCCL is compared to a restaurant, then the enqueue module is the "front desk ordering system": the user (application layer) says "I want an AllReduce," and the front desk translates it into a work order that the kitchen (GPU kernel) can execute—which stove, what pan to use, and how many batches to make it in. Without this translation layer, the kitchen wouldn't know what dish to make at all.

I. Entry Point: How ncclAllReduce Constructs ncclInfo

Intuitive Model

ncclAllReduceIt is the API function directly called by the user. Its responsibility is extremely singular:Package the raw parameters passed in by the user into ancclInfostructure, then hand it off toncclEnqueueCheck. This is like going to a bank counter to handle business—the teller first fills your request into a standard form, then forwards it to the backend system.

Without this layer, every collective communication API would have to handle parameter validation, group semantics, and profiler instrumentation on its own—the code would become so repetitive it would be unmaintainable.

Data Structure: Memory Layout of ncclInfo

ncclInfoIt is the core carrier that runs through the entire enqueue process. Its definition is insrc/include/info.h:

📎 src/include/info.h:17-44

This structure has 20+ fields, which we can divide into four groups by function:

Field GroupFieldPurpose
Collective Communication Parameterscoll, sendbuff, recvbuff, count, datatype, op, rootDescribes "what to do"
Communication Domain and Streamcomm, streamDescribes "where to do it"
Algorithm DetailschunkSteps, sliceStepsDescribes "how to partition"
One-sided OperationspeerWinOffset, peerWin, sigIdx, ctx, flags, nDesc, signalDescsRMA-specific
User ConfigurationcollConfigA private copy copied from the user config

Note thecollConfigcomment:"A config copied from config passed by user so older user config can be safely accessed during synchronous host scheduling (never at launch/replay)" 📎 src/include/info.h:41-43. This is a key design—the config pointer passed in by the user may be destroyed beforencclGroupEnd, so NCCL makes a copy inncclInfo.

Step-by-Step: The Call Chain of ncclAllReduce

We takencclAllReduceas an example, tracing the complete path from the user call to the construction ofncclInfo.

Step 1: The user calls ncclAllReduce.The entry point is insrc/collectives.cc:

📎 src/collectives.cc:206-211

Three things are done here:

1. NVTX3_FUNC_WITH_PARAMSAdd an NVTX marker (for visualization in tools like Nsight)

2. CallncclAllReduceConfigImpl, passing inconfig = nullptr

3. Return the result

Step 2: ncclAllReduceConfigImpl constructs ncclInfo.This is the key step:

📎 src/collectives.cc:192-202

Note that C-style aggregate initialization is used here:

c
struct ncclInfo info = {ncclFuncAllReduce, "AllReduce",
                        sendbuff, recvbuff, count, datatype, op, 0, comm, stream,
                        ALLREDUCE_CHUNKSTEPS, ALLREDUCE_SLICESTEPS};

The fields correspond one-to-one in the declaration order ofncclInfo.ALLREDUCE_CHUNKSTEPSandALLREDUCE_SLICESTEPSare defined insrc/include/collectives.h:

📎 src/include/collectives.h:19-20

NCCL_STEPSis the number of steps in the ring buffer (usually 8 or 16), so AllReduce's chunkSteps isNCCL_STEPS/2, and sliceSteps isNCCL_STEPS/4. This means one chunk contains 2 slices.

Step 3: Parse the user config. ncclParseCollConfigParses the user-passedncclCollConfig_t*intoinfo.collConfig. Ifconfig == nullptr, this field remains zero-initialized.

Step 4: Hand off to ncclEnqueueCheck.This is the true entry point of the enqueue module.

Design Thinking: Why Use Aggregate Initialization Instead of Field-by-Field Assignment?

[Design Inference and Architectural Trade-offs]

Aggregate initialization has two benefits: first, the compiler checks whether the number of fields matches (a missing field triggers a warning); second, the code is more compact. But the drawback is thatthe field order must strictly match the struct declaration—if someone inserts a field in the middle ofncclInfo, all aggregate initialization sites will silently misalign. This is an implicit maintenance risk in the NCCL codebase.

Production Pitfall: Config Lifetime

A real pitfall scenario: the user writes code like this:

c
ncclCollConfig_t config = {...};
ncclAllReduceConfig(..., &config);
// config 在这里被销毁(比如是栈变量,函数返回了)

If NCCL did not copy the config inncclInfo, then accessingncclGroupEndatinfo.collConfigwould read already-freed memory.src/include/info.h:41-43The comment inis precisely to explain this design—。

---

the config is parsed and copied during the task append phase, and afterward no longer depends on the user pointer

II. ncclEnqueueCheck: Parameter Validation and Group Semantics

ncclEnqueueCheckIntuitive ModelIt is the "main gate" of the enqueue module. All collective communication APIs ultimately converge here. Its responsibilities are:Validate parameter legality, handle group semantics, and call taskAppend to generate tasksncclEnqueueCheck。

. If compared to airport security, then each API function is a check-in counter—check-in only takes luggage; the real security check is at

Step-by-Step: The execution flow of ncclEnqueueCheck

📎 src/enqueue/enqueue.cc:3478-3527

Let's break it down step by step:

Step 1: CommCheck validates the communicator. CommCheck(info->comm, info->opName, "comm")Check whether the comm pointer is non-null and whether it has been initialized. If the comm has been revoked (for example, if some rank encounters an error), return an error directly:

📎 src/enqueue/enqueue.cc:3480-3485

Step 2: Handle profiler depth.If already inside a group (profilerGroupDepth > 0), increment the depth counter. This is to correctly handle implicitncclGroupStartInternal/ncclGroupEndInternalcalls.

Step 3: Enter the internal group. ncclGroupStartInternal()This is NCCL's internal group mechanism.Key point: Even if the user does not explicitly callncclGroupStart, NCCL will create an implicit group for each API call. This guarantees the atomicity of a single call.

Step 4: Ensure comm is ready. ncclCommEnsureReady(info->comm)Wait for communicator initialization to complete (for example, bootstrap completion and connection establishment).

Step 5: ArgsCheck parameter validation.This is the most complex validation step:

📎 src/enqueue/enqueue.cc:3497-3503

Note the handling ofcheckMode: If it isncclCheckModeDebugGlobal,ArgsCheck, info will be enqueued, and global validation will be performed atncclGroupEnd(for example, checking whether the count is consistent across all ranks).

Step 6: Call taskAppend.This is the core conversion step:

📎 src/enqueue/enqueue.cc:3513

Step 7: Increment opCount.After each successful enqueue,comm->opCount++. This counter is used to match send/recv operations and is also the basis for the profiler timeline.

Step 8: Exit the group. ncclGroupEndInternal()If depth drops to 0, the actual group operation is triggered (scheduling and kernel launch).

Concurrency control: group semantics and thread safety

[Design inference and architectural trade-offs]

ncclGroupStartInternal/ncclGroupEndInternalThread-local storage (TLS) is used to maintain group state. This means thatmultiple API calls within the same thread will be merged into one group, but calls from different threads are independent. This is the foundation of NCCL's support for multithreaded calls.

An easy pitfall: If the user calls a non-NCCL CUDA API betweenncclGroupStartandncclGroupEnd(for example,cudaMemcpy), it may cause stream ordering issues. NCCL's group mechanism assumes that operations within a group are all on the same set of streams.

Error recovery chain

ncclEnqueueCheckThe error handling of

📎 src/enqueue/enqueue.cc:3524-3526

has an ingenious design:taskAppendIfncclCommSetAsyncErrorfails and comm is in non-blocking mode, it will call

---

to record the error. In this way, subsequent API calls will immediately return an error instead of continuing to try. This is the asynchronous error propagation mechanism.

III. taskAppend: The crossroads of task dispatch

taskAppendIntuitive modelinfo->collIt is the "transport hub" of the enqueue module. Based on the value of

, it dispatches tasks to different processing paths: P2P, RMA, CE, or ordinary collective communication. This is like a post office sorting center - based on the address on the envelope, it delivers letters to different mailboxes.

Without this dispatch layer, all types of operations would have to be crammed into one huge if-else, making the code difficult to maintain.

📎 src/enqueue/enqueue.cc:3337-3476

Step-by-Step: The dispatch logic of taskAppend ncclParamEnqueueRearchEnable()Step 1: Determine whether the new architecture is enabled.rawTaskAppendIt is an environment variable switch (default 0). If enabled, take the

path - this is the new task model that NCCL is developing.Step 2: P2P dispatch.p2pTaskAppend:

📎 src/enqueue/enqueue.cc:3343-3345

If it is Send/Recv, callStep 3: RMA dispatch.rmaTaskAppend:

📎 src/enqueue/enqueue.cc:3346-3347

If it is PutSignal/Signal/WaitSignal, call if (info->count == 0) return ncclSuccess;Step 4: Early return for empty collective communication.

- Collective communication with count 0 is discarded directly. ncclCollConfigGetAlgMaskStep 5: Algorithm selection validation.

📎 src/enqueue/enqueue.cc:3357-3358

Validate whether the algorithm selection passed in by the user is legal:Step 6: FP8 type check.

📎 src/enqueue/enqueue.cc:3360-3366

FP8 reduction requires sm90+: hostToDevRedOpStep 7: Reduction operation conversion.ncclRedOp_tConvert the host-sidencclDevRedOpFull:

📎 src/enqueue/enqueue.cc:3370-3371

to the device-sideStep 8: Early return for single rank.comm->nRanks == 1IfncclLaunchOneRank, directly call

📎 src/enqueue/enqueue.cc:3373-3377

to perform local reduction without generating a task:Step 9: Multi-rank path.

📎 src/enqueue/enqueue.cc:3378-3470

This is the most complex branch, including CE routing, AllToAll/Gather/Scatter fallback, and ordinary collective communication:

collTaskAppendData structure: fields of ncclTaskCollncclTaskCollThis is where

📎 src/enqueue/enqueue.cc:2757-2851

is generated. Let's look at its core logic:

Key field assignments:FieldSource
funcinfo->collMeaning
sendbuff/recvbuffinfo->sendbuff/recvbuffCollective communication type
countinfo->countBuffer pointer
datatypeinfo->datatypeElement count
trafficBytescount * elementSize * ncclFuncTrafficPerByteData type
opHost/opDevinfo->op/opDevTraffic estimation
chunkSteps/sliceStepsinfo->chunkSteps/sliceStepsReduction operation
minCTAs/maxCTAs/nvlsCTAsNumber of split stepsConfiguration parsing
algMaskncclCollConfigGetAlgMaskResource limit

Algorithm selection masktrafficBytesNote the calculation of

📎 src/enqueue/enqueue.cc:2813

ncclFuncTrafficPerByte:

📎 src/enqueue/enqueue.cc:123-134

It returns the traffic multiplier for each collective communication type:

AllReduce returns 2 (because it needs reduce + broadcast), AllGather/ReduceScatter returns nRanks, and others return 1.

📎 src/enqueue/enqueue.cc:2808-2812

Design consideration: Why should AllGather/Broadcast be converted to int8?ncclInt8. This is an optimization:These two operations do not involve reduction, so there is no need to care about data types. Handling them uniformly as bytes can simplify the kernel logic.。

Production pitfall: the parsing order of CTAPolicy

📎 src/enqueue/enqueue.cc:3390-3397

CTAPolicy parsing has a subtle priority:env > per-call > comm. AndNCCL_CTA_POLICY_ZEROtakes precedence overNCCL_CTA_POLICY_EFFICIENCY. If the user sets both flags at the same time, ZERO will take effect.

A real pitfall scenario: the user setNCCL_CTA_POLICY=EFFICIENCY, but found that the CE path was not used. The reason is that CE routing requiresCTAPolicy & NCCL_CTA_POLICY_ZEROto be true, and EFFICIENCY does not satisfy this condition.

---

4. ncclPrepareTasks: from task list to scheduling queue

Intuitive model

ncclPrepareTasksIt is the "preprocessor" of the enqueue module. It buckets the scattered task list by (func, op, datatype), and then computes the algorithm and protocol for each bucket. This is like a librarian - first sorting returned books by category, then deciding which shelf each category of books goes on.

Without this step, the subsequentscheduleCollTasksToPlanwould have to compute the algorithm separately for each task, which is extremely inefficient.

Step-by-Step: the bucketing logic of ncclPrepareTasks

📎 src/enqueue/enqueue.cc:423-642

Step 1: Broadcast task conversion.If there is only one broadcast peer, convert the broadcast task into a coll task:

📎 src/enqueue/enqueue.cc:430-461

Note that here the fields ofbcastTaskare copied to the newncclTaskColl, andtrafficBytesis computed. Then the original task is released frommemPool_ncclTaskBcast.

Step 2: Bucket by (func, op, datatype).Tasks come out of the sorter in descending order of size, and are then assigned to thetasksByFnOpTyarray:

📎 src/enqueue/enqueue.cc:464-487

Index calculation:((int)task->func * ncclNumDevRedOps + (int)task->opDev.op) * ncclNumTypes + (int)task->datatype. This is the linearization of a three-dimensional array.

Step 3: Aggregation and algorithm selection.For each bucket, aggregate tasks with similar sizes (within 4x), and then callncclGetAlgoInfo:

📎 src/enqueue/enqueue.cc:503-547

Step 4: Bucket by (collnet, nvls).According to the algorithm type, assign tasks tocollBins[2][2]:

📎 src/enqueue/enqueue.cc:517-544

Step 5: Concatenate the final queue.Concatenate the four buckets intoplanner->collTaskQueue:

📎 src/enqueue/enqueue.cc:553-557

Data structure: ncclTaskCollSorter

ncclTaskCollSorteris an insertion sorter ordered bytrafficBytes.ncclTaskCollSorterInsertinserts the task into the correct position,ncclTaskCollSorterDequeueAllretrieves all tasks in order.

[Design inference and architectural trade-offs]

The design motivation of this sorter is:Large tasks are scheduled first. Because large tasks have long transfer times, starting them first allows better overlap of computation and communication.

Concurrency control: runtimeConn and connection establishment

📎 src/enqueue/enqueue.cc:572-583

Ifcomm->runtimeConnis true (runtime connection mode), and the channel of some algorithm has not yet been initialized, markalgoNeedConnect. This will trigger connection establishment later.

Production pitfall: boundary conditions of aggregation

📎 src/enqueue/enqueue.cc:507-508

The aggregation condition isaggEnd->trafficBytes < 4 * aggBeg->trafficBytes, and neither task setsaggIsolate. If the user sets a per-call config (for example,maxCTAs),aggIsolatewill be set to true, this task will not be aggregated.

A real pitfall scenario: the user setmaxCTAs=4for a certain AllReduce, expecting it to use only 4 CTAs. However, due to the aggregation logic, this task may be merged with adjacent tasks, causing the actual number of CTAs used to not match expectations. The solution is to setaggIsolate- NCCL has already handled this incollTaskAppend:

📎 src/enqueue/enqueue.cc:2821-2822

---

5. scheduleCollTasksToPlan: channel splitting and budget control

Intuitive model

scheduleCollTasksToPlanIt is the "scheduler" of the enqueue module. It assigns tasks to specific channels and computes the data split for each channel. This is like a factory's production scheduling system - deciding what each production line does and how much it does.

Without this step, the GPU kernel would not know which part of the data it needs to process.

Step-by-Step: channel splitting algorithm

📎 src/enqueue/enqueue.cc:644-947

Step 1: Budget estimation.First estimate the number of tasks that can fit into this plan:

📎 src/enqueue/enqueue.cc:648-689

ncclTestBudgetCheck whether the work byte count exceeds the budget:

📎 src/enqueue/enqueue.cc:343-349

Step 2: Compute the traffic for each channel.According to kind (collnet/nvls), computetrafficPerChannel:

📎 src/enqueue/enqueue.cc:701-707

Step 3: Collnet path.If it is a collnet algorithm, channel assignment is relatively simple:

📎 src/enqueue/enqueue.cc:709-739

Step 4: Cell splitting for the normal path.This is the most complex part. NCCL splits data into "cells", and each cell is a minimum transfer unit:

📎 src/enqueue/enqueue.cc:740-845

Key variables:

  • cellSize: the number of bytes per cell, at leastMinTrafficPerChannel(32KB)
  • cells: total number of cells
  • cellsPerChannel: number of cells processed by each channel
  • cellsLo/cellsHi: number of cells for the first and last channels (may be less than full)

Step 5: Compute chunkGrains.Call for each channel segmentcalcCollChunking:

📎 src/enqueue/enqueue.cc:811-825

Step 6: Generate proxyOp.Generate a proxy operation for each channel:

📎 src/enqueue/enqueue.cc:844-894

Data structure: ncclDevWorkColl

ncclDevWorkCollis the device-side work descriptor. Its key fields:

FieldMeaning
sendbuff/recvbuffBuffer pointer
channelLo/channelHiChannel range
cbd.countLo/countMid/countHiNumber of elements in each segment
cbd.chunkGrainsLo/Mid/HiChunk granularity of each segment
directDirect flag

Concurrency control: bit operations of channelMask

📎 src/enqueue/enqueue.cc:897

This line of code uses bit operations to set channelMask:(2ull << channelHi) - (1ull << channelLo)For example, channelLo=2, channelHi=5, the result is(2<<5) - (1<<2) = 64 - 4 = 60 = 0b111100, meaning bits 2-5 are set.

Production pitfall: budget overflow

📎 src/enqueue/enqueue.cc:792-794

If the budget is insufficient, directly returnncclSuccess, letting the outer loop create a new plan. This is an elegant degradation strategy—no error, just batch processing。

A real pitfall scenario: ifNCCL_WORK_FIFO_BYTESis set too small, each plan can only hold very few tasks, increasing the number of kernel launches and reducing performance.

---

6. finishPlan: From tasks to kernel parameters

Intuitive model

finishPlanis the "packer" of the enqueue module. It packs tasks, batches, and proxyOps into a parameter structure that the kernel can directly read. This is like express packaging—putting loose items into boxes, attaching waybills, and waiting for shipment.

Step-by-Step: The packing logic of finishPlan

📎 src/enqueue/enqueue.cc:236-330

Step 1: Decide the storage type.If all work can fit into kernel args, usencclDevWorkStorageTypeArgs:

📎 src/enqueue/enqueue.cc:244-250

Step 2: Allocate kernelArgs.Allocate from the memory stack:

📎 src/enqueue/enqueue.cc:251-255

Step 3: Round-robin placement of batches.The first batch of each channel must be placed atbatchZero[blockIdx.x]:

📎 src/enqueue/enqueue.cc:257-280

Step 4: Merge proxyOp queues.Merge sort by opCount:

📎 src/enqueue/enqueue.cc:282-329

Data structure: ncclDevKernelArgs

ncclDevKernelArgsis the parameter structure passed to the kernel. It contains:

  • comm: device-side communicator
  • channelMask: channel bitmask
  • workStorageType: work storage type
  • workBuf: work buffer pointer
  • workMask: work buffer mask

Production pitfall: batch order

📎 src/enqueue/enqueue.cc:257-259

The comment states it clearly: "The first batch for each channel must be located at batchZero[blockIdx.x]". If this order is wrong, the kernel will read the wrong batch, causing data corruption.

---

Chapter summary

In this chapter, we traced the complete path fromncclAllReducetoncclTaskColl:

1. ncclAllReduceconstructsncclInfo, packing user parameters

2. ncclEnqueueCheckvalidates parameters, handles group semantics

3. taskAppenddispatches to different paths based on operation type

4. collTaskAppendgeneratesncclTaskColl, parses configuration

5. ncclPrepareTasksbuckets by (func, op, datatype), computes algorithms

6. scheduleCollTasksToPlansplits channels, generatesncclDevWorkColl

7. finishPlanpacks into kernel parameters

Key design principles:

  • Layered decoupling: each function does only one thing, passing state throughncclInfoandncclTaskColl
  • Budget control: controlling the size of each plan throughncclTestBudget
  • Aggregation optimization: tasks of similar size are aggregated, reducing the number of kernel launches
  • Configuration priority:env > per-call > comm

In the next chapter, we will entertask_sched, to see how NCCL orchestrates the execution order of multiple channels and multiple kernels.

Chapter review and self-test

Q1: If thecollTaskAppendinaggIsolateis removed (i.e.,src/enqueue/enqueue.cc:2821-2822always returns false), in what scenarios would the user-setmaxCTAsbecome ineffective? Why?

Reference analysis:aggIsolate's purpose is to mark "this task cannot be aggregated". If this check is removed, tasks with per-call config set will be merged with adjacent tasks. InncclPrepareTasks's aggregation loop (src/enqueue/enqueue.cc:507-508), the aggregation condition isaggEnd->trafficBytes < 4 * aggBeg->trafficBytes && !aggBeg->aggIsolate && !aggEnd->aggIsolate. IfaggIsolatealways returns false, then even if a task hasmaxCTAs=4set, it may be merged with a task that hasmaxCTAs=32. The mergedaggwill take some combination of the two (depending on the implementation ofncclGetAlgoInfo), causing the actual number of CTAs used to not match user expectations.

More seriously, inscheduleCollTasksToPlan(src/enqueue/enqueue.cc:665-666),taskAggIsolateis used to ensure that tasks with per-call resources configured occupy a separate plan. If this check fails, multiple tasks will share the plan's channel budget, causing resource allocation to not match expectations.

Q2: InncclEnqueueCheck, ifncclGroupEndInternal()returns an error (e.g., ArgsCheck fails for some rank), buttaskAppendhas already executed successfully, what happens? How does NCCL ensure state consistency?

Reference analysis: Look atsrc/enqueue/enqueue.cc:3513-3519's control flow:

c
NCCLCHECKGOTO(taskAppend(info->comm, info), ret, fail);
info->comm->opCount++;
exit:
  if (devOld != -1) CUDACHECK(cudaSetDevice(devOld));
  ncclGroupErrCheck(ret);
  NCCLCHECK(ncclGroupEndInternal());

IftaskAppendsucceeds butncclGroupEndInternalfails,opCounthas already been incremented. This will cause the opCount of subsequent operations to mismatch with the peer, potentially triggering a hang.

NCCL's approach is:ncclGroupErrCheck(ret)will check for errors, and if there are any, will set the comm's error state. Subsequent API calls will detect this error throughncclCommGetAsyncErrorand return immediately. This is a "fail-fast" strategy—once an error occurs, the entire comm enters an error state and no longer attempts recovery.

In a production environment, this means that once a group error occurs, the user needs to destroy and rebuild the communicator.

Q3: scheduleCollTasksToPlanThe cell splitting algorithm insrc/enqueue/enqueue.cc:740-845) has a boundary condition: whencellsLo == 0, it skips the minimum number of channels. If this skip logic has a bug (e.g.,channelIdis not correctly incremented), what consequences would it cause?

Reference analysis: Look atsrc/enqueue/enqueue.cc:770-780:

c
if (cellsLo == 0) {
  // Least channel skipped. Make the next channel the new least.
  channelId += 1;
  if (nMidChannels == 0) {
    cellsLo = cellsHi;
    cellsHi = 0;
  } else {
    cellsLo = cellsPerChannel;
    nMidChannels -= 1;
  }
}

IfchannelIdis not correctly incremented, then the next task will start allocating from the wrong channel. This will cause:

1. Channel overlap: two tasks may be allocated to the same segment of data in the same channel

2. Data corruption: the kernel will repeatedly process or miss data

3. Performance degradation: channel load imbalance

More insidiously, this kind of bug may only be triggered under specific message sizes (whencellsLo == 0), making it hard to reproduce. NCCL usesplan->channelMask |= (2ull << devWork->channelHi) - (1ull << devWork->channelLo)to track the channels already in use, but this is only a record; it cannot prevent overlap.

At this point, we have clearly seen how ncclAllReduce goes from a user call to a series of executable kernel tasks: parameter validation, algorithm/protocol determination, channel partitioning, and finally generating ncclInfo and ncclTaskColl. But creating the tasks is only the first step—they still need to be scheduled onto multiple channels, generate kernel launch parameters, and handle batch submission and dependency ordering under group semantics. The next chapter will dive into src/enqueue/task_sched and src/enqueue/task_prep to answer "why a single AllReduce launches multiple kernels, and how their order and dependencies are guaranteed," while also revealing how ncclGroupStart/ncclGroupEnd in src/group.cc merge multiple API calls into a single submission.

CHAPTER 07

Chapter 7: Chapter 7: Task Scheduler: How task_sched orchestrates multi-channel and kernel execution order

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 7 / 25

Chapter 7: Task Scheduler: How task_sched orchestrates multi-channel and kernel execution order

In the previous chapter, we traced ncclAllReduce all the way to ncclTaskColl—the task description object is already sitting in comm->planner. But a task description is only a "work order"; it has not yet become a kernel actually running on the GPU. This chapter answers three questions: How are multiple API calls accumulated and submitted together? How are the accumulated tasks split across multiple channels? What guarantees the order and dependencies among multiple kernels? First, here is an overall mental model. Think of NCCL as a restaurant: ncclGroupStart/ncclGroupEnd is the "shopping cart," where the user puts several dishes (multiple collective communication calls) into the cart; ncclGroupEnd is "placing the order," and only then does the kitchen start cooking according to the order. And doLaunches is the "dish dispatch coordinator," deciding which dishes go out first and which can be prepared in parallel. Without group semantics, each dish is ordered separately, and the kitchen has to relight the fire (launch a kernel) for every dish, which is extremely expensive; without doLaunches' round-based scheduling, multi-channel kernels would launch out of order, breaking data dependencies.

1. Global state of group semantics: thread_local variables and the "shopping cart" model

Intuitive model

ncclGroupStartandncclGroupEndAll communication calls between them do not immediately launch kernels, but are "accumulated." Where are they accumulated? They are accumulated inthread-local (thread_local)global variables. Why thread_local? Because NCCL assumes that group calls within the same thread are serial, and different threads each have independent shopping carts that do not interfere with each other. If these states were global variables rather than thread_local, two threads callingncclGroupStartat the same time would step on each other, causing one thread's tasks to be submitted by another thread'sncclGroupEnd—this would be catastrophic.

Data structures and memory layout

First look at the global state definition of group.

📎 src/group.cc:34-34

cpp
thread_local int ncclGroupDepth = 0; // depth of ncclGroupStart nesting
thread_local ncclResult_t ncclGroupError = ncclSuccess;
thread_local struct ncclComm* ncclGroupCommHead[ncclGroupTaskTypeNum] = {nullptr};
thread_local struct ncclComm* ncclGroupCommPreconnectHead = nullptr;
thread_local struct ncclIntruQueue<struct ncclAsyncJob, &ncclAsyncJob::next> ncclAsyncJobs;
thread_local int ncclGroupBlocking = -1; /* default mode */

Breaking down each field one by one:

  • ncclGroupDepth: nesting depth.ncclGroupStartcan be nested (though uncommon); each timencclGroupStartincrements by one,ncclGroupEnddecrements by one. Only when it reaches 0 is the submission actually performed. This is like a shopping cart being nestable—you open a sub-cart inside a cart, and only the outermost checkout actually places the order.
  • ncclGroupError: if any call within the group errors, the error is recorded here,ncclGroupEndand handled uniformly at that time. This avoids the inconsistent state where "after one call fails, subsequent calls are still adding things to the shopping cart."
  • ncclGroupCommHead[ncclGroupTaskTypeNum]: head of the communication domain linked list grouped by task type.ncclGroupTaskTypeNumis the number of task types (collective communication, raw tasks, management tasks, symmetric registration, etc.). Each type has a linked list, and the list nodes arencclComm, connected throughcomm->groupNext[type]. Why group by type? Because different types of tasks have different submission timing and dependency relationships—collective communication tasks need preconnect first, and management tasks (such as destroy) need to execute last.
  • ncclGroupCommPreconnectHead: linked list of communication domains that need preconnection. Preconnection means "establishing network connections in advance" to avoid latency caused by establishing connections only at kernel launch time.
  • ncclAsyncJobs: asynchronous task queue. Some tasks (such asncclCommInitRank) are asynchronous; they are placed into this queue and uniformly started atncclGroupEnd.
  • ncclGroupBlocking: blocking mode flag.-1means not yet determined,0means non-blocking,1indicates blocking. Mixing blocking and non-blocking communication domains within the same group is not allowed; otherwise, an error will be reported.

There is a key design here:ncclGroupCommHeadisarray, and each element is a linked list. The linked list nodes are connected throughcomm->groupNext[type]instead of using a separate linked list node structure. This means thatncclCommthe struct must reservegroupNextan array field. This "intrusive linked list" design avoids additional memory allocation, but the cost is thatncclCommthe struct becomes larger.

Scenario-driven Step-by-Step Walkthrough

Scenario: The user callsncclGroupStart(), then calls it twice in successionncclAllReduce(for two different communication domains commA and commB respectively), and finally callsncclGroupEnd()。

Step 1:ncclGroupStartWhat was done?

📎 src/include/group.h:63-66

cpp
inline ncclResult_t ncclGroupStartInternal() {
  ncclGroupDepth++;
  return ncclSuccess;
}

Extremely simple: increment the depth by one. No memory allocation, no locks, no system calls. This is whyncclGroupStarthas almost zero overhead.

Step 2:ncclAllReduceWhat happens when it is called within a group?

ncclAllReduceInternally it will callncclGroupCommJoin(comm, ncclGroupTaskTypeCollective), adding the communication domain to the group linked list.

📎 src/include/group.h:80-116

cpp
inline void ncclGroupCommJoin(struct ncclComm* comm, int type) {
  if (comm->groupNext[type] == reinterpret_cast<struct ncclComm*>(NCCL_COMM_GROUP_INVALID)) {
    // Insert comm into ncclGroupCommHead adjacent to sibling comms. This preserves
    // the users program order yet insures siblings occur consecutively. This
    // is required by doLaunches() in "group.cc".
    struct ncclComm** pp = &ncclGroupCommHead[type];
    while (*pp != nullptr && comm->intraComm0 != (*pp)->intraComm0) pp = &(*pp)->groupNext[type];

    // didn't find its clique, we need to insert it with ascending order based on commHash
    if (*pp == nullptr) {
      pp = &ncclGroupCommHead[type];
      while (*pp != nullptr && (*pp)->commHash < comm->commHash) pp = &(*pp)->groupNext[type];
    }
    comm->groupNext[type] = *pp;
    *pp = comm;
    // Comms gets a new memory stack scope upon joining. Each task batched for
    // this comm is allocated there.
    if (type == ncclGroupTaskTypeCollective || type == ncclGroupTaskTypeRawTask) {
      // Initialize planner
      ncclMemoryStackPush(&comm->memScoped);
      ncclKernelPlanner::Peer* tmp = comm->planner.peers;
      ncclIntruQueue<ncclTaskRma, &ncclTaskRma::next>* tmpRmaQueues = comm->planner.rmaTaskQueues;
      int numRmaCtx = comm->config.numRmaCtx;
      memset(&comm->planner, 0, sizeof(comm->planner));
      comm->planner.peers = tmp;
      comm->planner.bcast_info.minBcastPeer = INT_MAX;
      comm->planner.bcast_info.maxBcastPeer = INT_MIN;
      comm->planner.rmaTaskQueues = tmpRmaQueues;
      if (comm->planner.rmaTaskQueues != NULL) {
        for (int i = 0; i < numRmaCtx; i++) {
          ncclIntruQueueConstruct(&comm->planner.rmaTaskQueues[i]);
        }
      }
    }
  }
  ncclGroupBlocking = comm->config.blocking;
}

This code has several ingenious aspects:

1. Idempotency check:if (comm->groupNext[type] == NCCL_COMM_GROUP_INVALID)ensures that the same communication domain is added only once within the same group. If the user calls it twice for the same commncclAllReduce, the second time it will not be added to the linked list again, but the task will be appended tocomm->planner.

2. Clique ordering:intraComm0is the identifier of a "global entity." If multiple communication domains belong to the same global entity (for example, split throughncclCommSplit), theirintraComm0are the same, and they are called a clique. The code first finds the clique byintraComm0, and inserts the comm next to its sibling nodes in the same clique. If no clique is found, it inserts in ascending order bycommHash. This ordering is so thatdoLaunchescan correctly handle barrier synchronization within the clique.

3. Memory stack scope:ncclMemoryStackPush(&comm->memScoped)allocates a new memory stack scope for this comm within the group. All tasks allocated for this comm (ncclTaskColl, etc.) are allocated from this stack.ncclGroupCommLeavewillncclMemoryStackPoprelease all task memory at once - this is the classic optimization of "batch allocation, batch release," avoiding the overhead of separatemalloc/freefor each task.

4. Planner reset:memset(&comm->planner, 0, sizeof(comm->planner))clears the planner, but retains thepeersandrmaTaskQueuespointers (first stored in temporary variables, then restored after memset). Why retain them? Because these two are preallocated arrays and do not need to be reallocated each time.bcast_infoThe min/max of are reset toINT_MAX/INT_MIN, used for the merge optimization of subsequent broadcast tasks.

Step 3:ncclGroupEndWhat was done?

📎 src/group.cc:1039-1164

ncclGroupEndInternalis the core. Parse it section by section:

📎 src/group.cc:1048-1061

cpp
if (ncclGroupDepth == 0) {
  WARN("ncclGroupEnd: not in a group call.");
  ret = ncclInvalidUsage;
  goto exit;
}
// ...
if ((--ncclGroupDepth) > 0) goto exit;

First check the depth, then decrement by one. If after decrementing it is still greater than 0, it means it is still inside a nested inner group, so return directly without submitting. Only when it reaches 0 does it continue.

📎 src/group.cc:1063

cpp
if ((ret = ncclGroupError) != ncclSuccess) goto fail;

If any call within the group has errored, jump directly to fail cleanup.

📎 src/group.cc:1084-1093

cpp
NEW_NOTHROW_GOTO(groupJob, ncclGroupJob, ret, fail);
ncclIntruQueueConstruct(&groupJob->asyncJobs);
groupJob->groupRefCount = 0;
groupJob->nonBlockingInit = false;
memcpy(groupJob->groupCommHead, ncclGroupCommHead, sizeof(ncclGroupCommHead));
groupJob->groupCommPreconnectHead = ncclGroupCommPreconnectHead;
groupJob->groupError = ncclSuccess;
groupJob->abortFlag = false;
groupJob->joined = false;
ncclIntruQueueTransfer(&groupJob->asyncJobs, &ncclAsyncJobs);

Create ancclGroupJob, and "transfer" the thread_local group state into the job object.ncclIntruQueueTransfertransfers the entirencclAsyncJobsqueue togroupJob->asyncJobs. This step is crucial: the thread_local state is "temporary," while the job object is "persistent" and can be held by an asynchronous thread.

📎 src/group.cc:1095-1147

cpp
if (hasCommHead || !ncclIntruQueueEmpty(&groupJob->asyncJobs) || ncclGroupCommPreconnectHead != nullptr) {
  /* make sure ncclGroupBlocking has been set. */
  if (ncclGroupBlocking != 0 && ncclGroupBlocking != 1) {
    WARN("Invalid group blocking state %d", ncclGroupBlocking);
    ret = ncclInternalError;
    goto fail;
  }
  if (ncclGroupBlocking == 0) {
    /* nonblocking group */
    // ... 设置 async error 为 ncclInProgress,创建线程执行 groupLaunchNonBlocking
    groupJob->base.func = groupLaunchNonBlocking;
    STDTHREADCREATE_GOTO(groupJob->base.thread, ncclAsyncJobMain, ret, fail, &groupJob->base);
    groupJob->nonBlockingInit = true;
    ret = ncclInProgress;
  } else {
    /* blocking group */
    int savedDev;
    CUDACHECKGOTO(cudaGetDevice(&savedDev), ret, fail);
    NCCLCHECKGOTO(groupLaunch(&groupJob->base, internalSimInfoPtr), ret, fail);
    CUDACHECKGOTO(cudaSetDevice(savedDev), ret, fail);
    if (simInfo) memcpy((void*)simInfo, (void*)internalSimInfoPtr, realSize);
    delete groupJob;
  }
} else {
  // Free when not needed (single rank case)
  delete groupJob;
}

Blocking mode: directly callgroupLaunchon the current thread and complete synchronously. Non-blocking mode: create a thread to executegroupLaunchNonBlocking, and immediately returnncclInProgress. The user subsequently queries progress throughncclCommGetAsyncError.

Note the saving and restoring ofcudaGetDevice/cudaSetDevice:groupLaunchInternally it will switch the CUDA device (because different comms may be on different GPUs), and after execution restores the user's original device. This is to prevent "NCCL internally switching devices and not switching back" from causing the user's subsequent CUDA calls to run on the wrong device.

Design Thinking and Production Pitfalls

Pitfall 1: Mixing blocking and non-blocking communication domains。ncclAsyncLaunchThere is a check in:

📎 src/group.cc:55-64

cpp
/* check if there are blocking and nonblocking comms at the same time in group. */
if (comm->destroyFlag) {
  ncclGroupBlocking = 1;
} else if (ncclGroupBlocking == -1) {
  /* first met communicator */
  ncclGroupBlocking = comm->config.blocking;
} else if (ncclGroupBlocking != comm->config.blocking) {
  WARN("Blocking and nonblocking communicators are not allowed in the same group.");
  ret = ncclInvalidArgument;
}

Why is mixing not allowed? Because a blocking group executes synchronously on the current thread, while a non-blocking group executes asynchronously on a separate thread. If mixed, it is impossible to determine whetherncclGroupEndshould return synchronously or returnncclInProgress. In a production environment, if the user accidentally puts blocking and non-blocking comms into the same group, they will receivencclInvalidArgument, but at this point the group state has already been polluted, and it must be re-ncclGroupStart。

Pitfall 2:ncclGroupErrorpropagation.. If a call within the group fails,ncclGroupErroris set,ncclGroupEndwill jump to the fail branch to executegroupCleanup。groupCleanupwill traverse all comms, release the plan memory in the planner, reset the planner, and clean up rawTaskQueue. If this step is not done cleanly, the next timencclGroupStartthe planner will still contain old data, causing duplicate task submission or memory leaks.

📎 src/group.cc:514-607

cpp
static void groupCleanup(struct ncclComm** groupCommHeadPtr,
                         struct ncclIntruQueue<struct ncclAsyncJob, &ncclAsyncJob::next>* asyncJobsPtr,
                         ncclResult_t error) {
  struct ncclComm* comm;
  for (int type = 0; type < ncclGroupTaskTypeNum; ++type) {
    comm = groupCommHeadPtr[type];
    groupCommHeadPtr[type] = nullptr;
    while (comm != nullptr) {
      struct ncclComm* next = comm->groupNext[type];
      (void)ncclGroupCommLeave(comm, type);
      // We don't know if preconnect succeeded or happened at all, so clear
      // the flags that let `taskAppend()` skip over checking if preconnect
      // is needed.
      if (type == ncclGroupTaskTypeCollective || type == ncclGroupTaskTypeRawTask) {
        comm->preconnectNext = reinterpret_cast<struct ncclComm*>(0x1);
        for (int i = 0; i < comm->nRanks; i++) {
          comm->connectSend[i] = 0UL;
          comm->connectRecv[i] = 0UL;
        }
        // Reclaim abandoned kernel plan memory.
        while (!ncclIntruQueueEmpty(&comm->planner.planQueue)) {
          struct ncclKernelPlan* plan = ncclIntruQueueDequeue(&comm->planner.planQueue);
          if (!plan->persistent) {
            while (!ncclIntruQueueEmpty(&plan->proxyOpQueue)) {
              struct ncclProxyOp* pxop = ncclIntruQueueDequeue(&plan->proxyOpQueue);
              ncclMemoryPoolFree(&comm->memPool_ncclProxyOp, pxop);
            }
            ncclMemoryPoolFree(&comm->memPool_ncclKernelPlan, plan);
          }
        }
        // Reset comm->planner to empty.
        // ...
      }
      // ...
    }
  }
  // ...
}

Notecomm->preconnectNext = reinterpret_cast<struct ncclComm*>(0x1)this line. This is a "sentinel value," indicating "this comm needs to be preconnected again." Why? Because during cleanup it is unknown whether preconnect succeeded, so the next check is forced to be repeated.0x1This value is very clever - it is not a valid pointer, but it can be used as an "uninitialized" marker.ncclGroupCommPreconnectCheck insideif (comm->preconnectNext == reinterpret_cast<struct ncclComm*>(0x1))to determine whether it needs to be added to the preconnect linked list.

---

2. Task Preparation:ncclPrepareTasksHow to Turn Task Descriptions into Schedulable Units

Intuitive Model

ncclPrepareTasksThis is the "prep work" phase. The ingredients in the shopping cart (task descriptions) are still raw and need to be washed, cut, and prepared (determining algorithms, protocols, channel partitioning) before they can go into the pot (launching kernels). If you skip this step and launch kernels directly, the kernel won't know how to partition the data or which path to take, and will crash immediately.

Scenario-Driven Step-by-Step Walkthrough

ncclPrepareTasksIngroupLaunchLegacyis called:

📎 src/group.cc:705-746

cpp
static ncclResult_t ncclPrepareTasksAndCollPreconnect(
  struct ncclComm* comm, ncclSimInfo_t* simInfo,
  struct ncclIntruQueue<struct ncclAsyncJob, &ncclAsyncJob::next>* asyncCollJobs) {
  if (ncclParamSingleProcMemRegEnable()) {
    // 单进程内存注册模式:把 prepare 和 preconnect 合并成一个异步 job
    struct ncclPrepareTasksAndCollPreconnectJob* job;
    NEW_NOTHROW(job, ncclPrepareTasksAndCollPreconnectJob);
    job->base.func = ncclPrepareTasksAndCollPreconnectFunc;
    // ...
    ncclIntruQueueEnqueue(asyncCollJobs, &job->base);
  } else {
    bool needConnect = false;
    bool algoNeedConnect[NCCL_NUM_ALGORITHMS];
    memset(algoNeedConnect, 0, sizeof(bool) * NCCL_NUM_ALGORITHMS);

    CUDACHECK(cudaSetDevice(comm->cudaDev));
    NCCLCHECK(ncclPrepareTasks(comm, algoNeedConnect, &needConnect, simInfo));

    if (comm->cuMemSupport && needConnect) {
      // 创建 preconnect job
      struct ncclPreconnectJob* job;
      NEW_NOTHROW(job, ncclPreconnectJob);
      job->base.func = ncclCollPreconnectFunc;
      // ...
      ncclIntruQueueEnqueue(asyncCollJobs, &job->base);
    }
  }
  return ncclSuccess;
}

ncclPrepareTasksThe output is two things:algoNeedConnectarray (which algorithms need to establish connections) andneedConnectflag (whether a connection is needed). IfneedConnectis true and cuMem is supported, a preconnect job is created and executed asynchronously.

ncclPrepareTasksWhat does it do internally? It iterates overcomm->plannertasks, determines the algorithm and protocol for each task, then callstaskAppendto append the task to the planner's plan. This logic was covered in the previous chapter and won't be repeated here.

Key points:ncclPrepareTasksiscalled per comm individuallybut preconnect isexecuted in batches per cliqueWhy? See the comments ingroupLaunchLegacy:

📎 src/group.cc:818-834

cpp
do {
  // We need to preconnect connections for collectives clique by clique to avoid
  // race condition for split shared comms which can connect the same connections
  // at the same time.
  comm = cliqueHead;
  do {
    NCCLCHECKGOTO(ncclPrepareTasksAndCollPreconnect(comm, simInfo, &asyncCollJobs), ret, fail);
    comm = comm->groupNext[ncclGroupTaskTypeCollective];
  } while (comm != nullptr && comm->intraComm0 == cliqueHead->intraComm0);
  // connect
  NCCLCHECKGOTO(asyncJobLaunch(&asyncCollJobs, groupAbortFlag), ret, fail);
  // ...
  cliqueHead = comm;
} while (cliqueHead != nullptr);

The comment explains it clearly:Preconnect per clique one at a time to avoid split shared comms simultaneously connecting the same set of connections causing races. If two comms are split from the same parent comm, they may share some connections. If preconnected in parallel, two threads might simultaneously try to establish the same connection, causing duplicate connections or inconsistent connection state. Executing serially per clique ensures only one clique is establishing connections at any given time.

Concurrency Control and Low-Level Interaction

asyncJobLaunchis the core of asynchronous task launching:

📎 src/group.cc:609-678

cpp
static ncclResult_t asyncJobLaunch(struct ncclIntruQueue<struct ncclAsyncJob, &ncclAsyncJob::next>* asyncJobsMain,
                                   volatile bool* groupAbortFlag) {
  ncclResult_t ret = ncclSuccess;
  bool jobsDone = false;
  bool errorJobAbortFlag = false;

  if (!ncclIntruQueueEmpty(asyncJobsMain)) {
    struct ncclAsyncJob* job = ncclIntruQueueHead(asyncJobsMain);
    if (job->next == nullptr) {
      // 只有一个 job,直接在当前线程执行,避免线程创建开销
      job->isThreadMain = true;
      ncclAsyncJobMain(job);
      job->state = ncclGroupJobJoined;
      return job->result;
    }
    // 多个 job,每个创建一个线程
    do {
      STDTHREADCREATE(job->thread, ncclAsyncJobMain, job);
      job = job->next;
    } while (job != nullptr);

    do {
      jobsDone = true;
      job = ncclIntruQueueHead(asyncJobsMain);
      do {
        ncclGroupJobState_t state = COMPILER_ATOMIC_LOAD(&job->state, std::memory_order_acquire);
        if (state == ncclGroupJobRunning) {
          jobsDone = false;
        } else if (state == ncclGroupJobDone) {
          int err;
          if ((err = ncclThreadJoin(job->thread)) != ncclSuccess) {
            WARN("asyncJobLaunch: failed to join thread for job");
            ret = ncclSystemError;
          }
          job->state = ncclGroupJobJoined;
          if (job->result != ncclSuccess && ret == ncclSuccess) {
            ret = job->result;
            errorJobAbortFlag = true;
          }
        } else {
          // safety check
          if (state != ncclGroupJobJoined) {
            WARN("Async job state is %d, expected %d", state, ncclGroupJobJoined);
            if (ret == ncclSuccess) ret = ncclInternalError;
            errorJobAbortFlag = true;
          }
        }

        if (!job->destroyFlag &&
            (COMPILER_ATOMIC_LOAD(groupAbortFlag, std::memory_order_acquire) || errorJobAbortFlag == true)) {
          COMPILER_ATOMIC_STORE(job->abortFlag, uint32_t(1), std::memory_order_release);
          COMPILER_ATOMIC_STORE(job->abortFlagDev, uint32_t(1), std::memory_order_release);
          if (job->childAbortFlag) {
            COMPILER_ATOMIC_STORE(job->childAbortFlag, uint32_t(1), std::memory_order_release);
            COMPILER_ATOMIC_STORE(job->childAbortFlagDev, uint32_t(1), std::memory_order_release);
          }
        }

        job = job->next;
      } while (job != nullptr);
      // Let preconnect threads progress.
      if (jobsDone == false) std::this_thread::sleep_for(std::chrono::microseconds(1));
    } while (jobsDone == false);

    if (ret != ncclSuccess) goto fail;
  }

exit:
  return ret;
fail:
  goto exit;
}

This code has several key design points:

1. Single job optimization: If there's only one job in the queue, no thread is created and it executes directly on the current thread. This avoids the overhead of thread creation and join. For single-comm groups, this is the common case.

2. Atomic state machine:job->stateis an atomic variable with three states:ncclGroupJobRunning、ncclGroupJobDone、ncclGroupJobJoined. After the worker thread finishes execution, it usesCOMPILER_ATOMIC_STORE(..., std::memory_order_release)to set it toDone; the main thread usesCOMPILER_ATOMIC_LOAD(..., std::memory_order_acquire)to read. The release/acquire pairing guarantees that all memory writes by the worker thread are visible to the main thread.

3. Busy-wait + micro-sleep: The main thread polls the status of all jobs. If any job is still running,sleep_for(1us)then continues polling. Why use 1 microsecond instead of a condition variable? Because preconnect is a short task (typically tens of microseconds to a few milliseconds), and the wake-up overhead of a condition variable may be greater than busy-waiting. A 1-microsecond sleep avoids CPU waste from pure spinning.

4. Error propagation and abort: If any job fails,errorJobAbortFlagis set, and all subsequent jobs'abortFlagare atomically set to 1. The worker thread checksabortFlagduring execution, and if aborted, exits early. This is a "fail-fast" mechanism, preventing other jobs from continuing to run foolishly after one job fails.

Mermaid diagram: control flow of group submission

mermaid
flowchart TD
    gs["ncclGroupStart()"] --> depth_inc["ncclGroupDepth++"]
    depth_inc --> api_calls["用户调用 ncclAllReduce 等"]
    api_calls --> join["ncclGroupCommJoin(comm, type)"]
    join --> check_dup{"comm->groupNext[type]<br/>== NCCL_COMM_GROUP_INVALID?"}
    check_dup -->|是| insert["插入 clique 链表<br/>ncclMemoryStackPush"]
    check_dup -->|否| skip["跳过(已加入)"]
    insert --> ge["ncclGroupEnd()"]
    skip --> ge
    ge --> depth_dec["--ncclGroupDepth"]
    depth_dec --> depth_zero{"depth == 0?"}
    depth_zero -->|否| ret_early["返回(嵌套内层)"]
    depth_zero -->|是| check_err{"ncclGroupError<br/>== ncclSuccess?"}
    check_err -->|否| fail_cleanup["groupCleanup()"]
    check_err -->|是| create_job["创建 ncclGroupJob<br/>转移 thread_local 状态"]
    create_job --> blocking{"ncclGroupBlocking?"}
    blocking -->|0 非阻塞| spawn_thread["STDTHREADCREATE<br/>groupLaunchNonBlocking"]
    blocking -->|1 阻塞| sync_launch["groupLaunch() 同步执行"]
    spawn_thread --> ret_progress["返回 ncclInProgress"]
    sync_launch --> ret_ok["返回 ncclSuccess"]
    fail_cleanup --> reset["groupLocalResetJobState()"]
    ret_progress --> reset
    ret_ok --> reset

---

Three,doLaunches: round scheduling for multi-channel multi-kernel

Intuitive model

doLaunchesis the "dish delivery dispatcher." The kitchen (GPU) has multiple stoves (channels), and each dish (kernel plan) needs to be served in order. But dishes from different comms may be served in parallel, while dishes from the same comm must be served in order. The dispatcher must ensure: comms within the same clique advance synchronously (using a barrier), while different cliques can advance independently.

Data structures and memory layout

doLaunchesThe core data structures ofncclKernelPlanarecomm->planner.unlaunchedPlansHead。

📎 src/group.cc:427-503

cpp
ncclResult_t doLaunches(struct ncclComm* head, int taskType) {
  ncclResult_t result = ncclSuccess;
  struct ncclComm* cliqueHead = head;
  struct ncclComm* cliqueNextHead;
  bool useBarrier = ncclParamLaunchMode == ncclLaunchModeGroup;
  // This outer loop iterates over cliques of comms which are siblings of the
  // same global entity. We calculate a clique as all comms which have the same
  // `intraComm0` value.
  do {
    struct ncclComm* comm = cliqueHead;
    bool capturingYes = false, capturingNo = false;
    do {
      (ncclCudaGraphValid(comm->planner.capturingGraph) ? capturingYes : capturingNo) = true;
      CUDACHECKGOTO(cudaSetDevice(comm->cudaDev), result, failure);
      NCCLCHECKGOTO(ncclLaunchPrepare(comm), result, failure);
      if (useBarrier) ncclCommIntraBarrierIn(comm, 1);
      comm = comm->groupNext[taskType];
    } while (comm != nullptr && comm != reinterpret_cast<struct ncclComm*>(NCCL_COMM_GROUP_INVALID) &&
             comm->intraComm0 == cliqueHead->intraComm0);
    cliqueNextHead = comm;

    if (capturingYes && capturingNo) {
      // We have entered barriers but are aborting without leaving them. Thus
      // these comms are permanently trashed. We need a good mechanism for
      // tracking and reporting that.
      WARN("Either none or all communicators in a ncclGroup() can be CUDA graph captured.");
      result = ncclInvalidUsage;
      goto failure;
    }

    while (true) {
      // Iterate rounds of launches for clique.
      bool moreRounds = false;
      comm = cliqueHead;
      do {
        // Iterate clique members.
        struct ncclComm* next = comm->groupNext[taskType];
        if (useBarrier) {
          // Barrier reduction result tells us if this was the final round.
          moreRounds = 0 != ncclCommIntraBarrierOut(comm);
        } else {
          moreRounds |= comm->planner.unlaunchedPlansHead != nullptr;
        }
        if (moreRounds) {
          // Pop next unlaunched kernel
          struct ncclKernelPlan* plan = comm->planner.unlaunchedPlansHead;
          if (plan != nullptr) {
            comm->planner.unlaunchedPlansHead = plan->next;
            CUDACHECKGOTO(cudaSetDevice(comm->cudaDev), result, failure);
            NCCLCHECKGOTO(ncclLaunchKernelBefore_NoUncapturedCuda(comm, plan), result, failure);
            if (plan->isCeColl) {
              NCCLCHECKGOTO(ncclLaunchCeColl(comm, plan), result, failure);
            } else if (plan->isRma) {
              NCCLCHECKGOTO(ncclLaunchRma(comm, plan), result, failure);
            } else {
              NCCLCHECKGOTO(ncclLaunchKernel(comm, plan), result, failure);
            }
          }
          // Barrier reduction input indicates if we require further rounds.
          if (useBarrier) ncclCommIntraBarrierIn(comm, comm->planner.unlaunchedPlansHead != nullptr ? 1 : 0);
          if (plan != nullptr) {
            NCCLCHECKGOTO(ncclLaunchKernelAfter_NoCuda(comm, plan), result, failure);
          }
        } else {
          // Final round.
          CUDACHECKGOTO(cudaSetDevice(comm->cudaDev), result, failure);
          NCCLCHECKGOTO(ncclLaunchFinish(comm), result, failure);
        }
        comm = next;
      } while (comm != reinterpret_cast<struct ncclComm*>(NCCL_COMM_GROUP_INVALID) && comm != cliqueNextHead);
      if (!moreRounds) break;
    }
    cliqueHead = cliqueNextHead;
  } while (cliqueHead != nullptr && cliqueHead != reinterpret_cast<struct ncclComm*>(NCCL_COMM_GROUP_INVALID));
failure:
  return result;
}

Copy

Scenario-Driven Step-by-Step WalkthroughScenariointraComm0: Two comms (commA and commB) belong to the same clique (

are the same), and each comm has 3 kernel plans pending launch.

First-level loop: iterate over cliquesdo-whileThe outercliqueHeaditerates over all cliques.do-whileis the first comm of the current clique. The innercomm->intraComm0 == cliqueHead->intraComm0)。

iterates over all comms in the clique (

  • cudaSetDevice(comm->cudaDev)For each comm:
  • ncclLaunchPrepare(comm): Switch to the GPU corresponding to that comm.
  • ncclCommIntraBarrierIn(comm, 1): Prepare for launch, including setting up the CUDA stream, checking resources, etc.

: Enter the barrier, with an initial value of 1.

while (true)Second-level loop: round scheduling

The loop executes "rounds." In each round, each comm in the clique launches one kernel plan.moreRoundsThe key is in the computation of

  • :(useBarrier == true):moreRounds = 0 != ncclCommIntraBarrierOut(comm)。ncclCommIntraBarrierOutWith barrier modeis across-comm barrier reduction operationncclCommIntraBarrierIn. It waits for all comms in the clique to callmoreRounds, then returns the reduction result of all input values (here, logical OR). If any comm still has unlaunched plans, the reduction result is 1,moreRoundsis true, and the next round continues. If all comms have no unlaunched plans, the reduction result is 0,
  • is false, and it enters the final round.:moreRounds |= comm->planner.unlaunchedPlansHead != nullptr. Directly check whether each comm still has an unstarted plan. Note that here it uses|=, as long as one comm still has a plan,moreRoundsis true.

Why is a barrier needed? Because the comms within a clique are "siblings"; they may share GPU resources or network connections. If one comm launches 3 kernels and another launches only 1, the comm that finishes launching first will enterncclLaunchFinish, release resources, while the other comm is still using these resources, causing a use-after-free. The barrier ensures that all comms within the clique advance synchronously: either they all launch round N, or they all enter the final round.

Kernel launch branch

📎 src/group.cc:477-483

cpp
if (plan->isCeColl) {
  NCCLCHECKGOTO(ncclLaunchCeColl(comm, plan), result, failure);
} else if (plan->isRma) {
  NCCLCHECKGOTO(ncclLaunchRma(comm, plan), result, failure);
} else {
  NCCLCHECKGOTO(ncclLaunchKernel(comm, plan), result, failure);
}

Three plan types:

  • isCeColl: CollNet collective communication (using NIC offload for collective communication).
  • isRma: RMA (Remote Memory Access) tasks.
  • Default: normal GPU kernel.

Each type has a different launch function, but all follow the "Before -> Launch -> After" pattern:

  • ncclLaunchKernelBefore_NoUncapturedCuda: preparation before launch (setting kernel parameters, uploading to device, etc.).
  • ncclLaunchKernel: actually launch the kernel (cudaLaunchKernel)。
  • ncclLaunchKernelAfter_NoCuda: cleanup after launch (updating state, releasing temporary resources).

Final round

WhenmoreRoundsis false, executencclLaunchFinish(comm). This step performs final cleanup: freeing plan memory, updating comm state, notifying the proxy thread, etc.

Concurrency control and hardware interaction

ncclCommIntraBarrierIn/Outis the synchronization primitive for comms within a clique. Its implementation involves atomic operations and spin-waiting.Inwrites the value to shared memory,Outwaits for all comms to write before reading the reduction result. This barrier iscross-process(if the comms are in different processes), and the underlying implementation may use shared memory or the network.

Why use a barrier instead of simply "checking whether all comms still have plans"? Because "checking" is non-atomic: when commA checks, commB still has a plan, so commA decides to continue; but commB immediately finishes launching its last plan after commA's check and enters the final round. commA is still launching kernels, while commB has already released shared resources. The barrier turns "checking" and "deciding" into a single atomic operation, eliminating this race.

Production Pitfall Guide

Pitfall 1: Mixing CUDA graph capture。

📎 src/group.cc:448-455

cpp
if (capturingYes && capturingNo) {
  // We have entered barriers but are aborting without leaving them. Thus
  // these comms are permanently trashed. We need a good mechanism for
  // tracking and reporting that.
  WARN("Either none or all communicators in a ncclGroup() can be CUDA graph captured.");
  result = ncclInvalidUsage;
  goto failure;
}

If some comms within a clique are in CUDA graph capture mode and others are not, it directly errors out. The comment says "these comms are permanently trashed" — because they have entered the barrier but not exited, the barrier states of these comms will forever be inconsistent, and they can no longer be used afterward. This is anunrecoverable error, and the user must rebuild the communication domain. In production, if a user mixes graph-capture and non-capture comms, they will receivencclInvalidUsage, but more seriously, the comm is already corrupted.

Pitfall 2:useBarrierconfiguration dependency。useBarrier = ncclParamLaunchMode == ncclLaunchModeGroup. If the user setsNCCL_LAUNCH_MODE=GROUP, the barrier path is taken; otherwise, the non-barrier path is taken. Under the non-barrier path,moreRoundsuses|=to accumulate, but each comm decides independently. If commA still has a plan while commB does not, commB will enter the final round and executencclLaunchFinish, while commA is still launching kernels. This is safe in some scenarios (there are no shared resources between comms), but if proxy threads or network connections are shared, it may cause problems. Therefore, barrier mode is recommended by default.

---

IV.groupLaunchLegacy's complete execution chain

Scenario-driven Step-by-Step Walkthrough

groupLaunchLegacyis the complete submission process in blocking mode. Execute in order:

Phase 1: P2P preconnect

📎 src/group.cc:756-774

cpp
if (!simInfo && groupCommPreconnectHeadMain != nullptr) {
  struct ncclComm* comm = groupCommPreconnectHeadMain;
  do {
    struct ncclPreconnectJob* job;
    NEW_NOTHROW_GOTO(job, ncclPreconnectJob, ret, fail);
    job->base.func = ncclP2PPreconnectFunc;
    // ...
    ncclIntruQueueEnqueue(asyncJobsMain, (struct ncclAsyncJob*)job);
    struct ncclComm* next = comm->preconnectNext;
    comm->preconnectNext = reinterpret_cast<struct ncclComm*>(0x1);
    comm = next;
  } while (comm != nullptr);
}
NCCLCHECKGOTO(asyncJobLaunch(asyncJobsMain, groupAbortFlag), ret, fail);

For each comm that needs preconnect, create ancclP2PPreconnectFuncjob, then launch them in batches.ncclP2PPreconnectFuncinternally callsncclTransportP2pSetupto establish the P2P connection.

Phase 2: Symmetric memory registration

📎 src/group.cc:778-808

cpp
// only loop through sym alloc and register tasks
for (int type = ncclGroupTaskTypeSymRegister; type <= ncclGroupTaskTypeSymRegister; ++type) {
  if (groupCommHeadMain[type]) {
    // 按 clique 批量执行 ncclCommGroupRegisterSymmetric
  }
}

Symmetric memory registration (ncclCommWindowRegister, etc.) is executed in batches by clique.

Phase 3: Collective communication preconnect

📎 src/group.cc:810-870

cpp
if (groupCommHeadMain[ncclGroupTaskTypeCollective] != nullptr) {
  // 按 clique 逐个 prepare + preconnect
  // 然后 ncclTasksRegAndEnqueue
  // 然后 debug check
}

This is the core phase. CallncclPrepareTasksAndCollPreconnectclique by clique, thenasyncJobLaunchexecutes preconnect. After preconnect completes, callncclTasksRegAndEnqueueto register the task into the plan and generate kernel launch parameters.

Phase 4:doLaunches

📎 src/group.cc:872-874

cpp
if ((!simInfo) && (groupCommHeadMain[ncclGroupTaskTypeCollective] != nullptr)) {
  NCCLCHECKGOTO(doLaunches(groupCommHeadMain[ncclGroupTaskTypeCollective], ncclGroupTaskTypeCollective), ret, fail);
}

Launch all kernel plans.

Phase 5: Cleanup

📎 src/group.cc:876-903

cpp
while (!ncclIntruQueueEmpty(asyncJobsMain)) {
  struct ncclAsyncJob* job = ncclIntruQueueDequeue(asyncJobsMain);
  if (!job->destroyFlag && job->comm && !job->comm->config.blocking &&
      groupCommHeadMain[ncclGroupTaskTypeCollective] == nullptr) {
    (void)ncclCommSetAsyncError(job->comm, ret);
  }
  if (job->destructor) job->destructor((void*)job);
}

for (int type = 0; type < ncclGroupTaskTypeNum; ++type) {
  while (groupCommHeadMain[type] != nullptr) {
    struct ncclComm* comm = groupCommHeadMain[type];
    struct ncclComm* next = comm->groupNext[type];
    // Poll for callbacks sent to us from other threads.
    if (comm->reclaimSteps == GROUP_MAX_RECLAIM_STEPS) {
      NCCLCHECKGOTO(ncclCommPollCallbacks(comm, /*waitSome=*/false), ret, fail);
      comm->reclaimSteps = 0;
    } else {
      comm->reclaimSteps++;
    }
    (void)ncclGroupCommLeave(comm, type);
    if (!comm->config.blocking) {
      (void)ncclCommSetAsyncError(comm, ret);
    }
    groupCommHeadMain[type] = next;
  }
}

Clean up asynchronous jobs, then iterate over all comms and callncclGroupCommLeave. Note the count ofreclaimSteps: everyGROUP_MAX_RECLAIM_STEPS(10) group calls, poll callbacks once. This is to avoid the overhead of polling callbacks on every group, while ensuring callbacks do not accumulate indefinitely.

Mermaid diagram:groupLaunchLegacydata flow of

mermaid
flowchart LR
    subgraph input["输入"]
        preconnect["ncclGroupCommPreconnectHead"]
        coll["ncclGroupCommHead[Collective]"]
        sym["ncclGroupCommHead[SymRegister]"]
    end

    subgraph phase1["阶段1: P2P preconnect"]
        p2p_job["ncclPreconnectJob<br/>func=ncclP2PPreconnectFunc"]
        p2p_launch["asyncJobLaunch"]
    end

    subgraph phase2["阶段2: 对称内存注册"]
        sym_job["ncclGroupSymmetricJob<br/>func=ncclCommGroupRegisterSymmetric"]
    end

    subgraph phase3["阶段3: 集合通信 prepare+preconnect"]
        prep["ncclPrepareTasksAndCollPreconnect"]
        coll_job["ncclPreconnectJob<br/>func=ncclCollPreconnectFunc"]
        reg_enq["ncclTasksRegAndEnqueue"]
    end

    subgraph phase4["阶段4: kernel 启动"]
        do_launch["doLaunches<br/>轮次调度"]
        plan["ncclKernelPlan"]
        kernel["ncclLaunchKernel"]
    end

    preconnect --> p2p_job --> p2p_launch
    sym --> sym_job
    coll --> prep --> coll_job --> reg_enq
    reg_enq --> plan --> do_launch --> kernel

---

V.groupLaunchEnqueueRearch: The new architecture scheduler

Intuitive model

groupLaunchEnqueueRearchis the new scheduling architecture being developed by NCCL. It divides task preparation, scheduling, and launch into finer stages, managed with an asynchronous job queue. Currently, the scheduler and launcher modules are "not yet implemented," falling back to legacydoLaunches。

📎 src/group.cc:991-996

cpp
// Schedule and launch tasks. Scheduler and launcher module of the enqueue framework
// is not yet implemented and falls back to the legacy launcher: a single phased
// doLaunches over the clique, run here on the user's thread.
if (!simInfo && groupCommHeadMain[ncclGroupTaskTypeRawTask] != nullptr) {
  NCCLCHECKGOTO(doLaunches(groupCommHeadMain[ncclGroupTaskTypeRawTask], ncclGroupTaskTypeRawTask), ret, fail);
}

Execution flow of the new architecture:

1. Manage tasks:ncclMgmtTaskJobFuncHandlemgmtTaskQueuetasks in (such as destroy).

2. Task preparation:ncclTaskPrepareJobFuncCallncclTaskPrepare。

3. Scheduling and launch: Fall back todoLaunches。

The new architecture usesncclGroupJobLaunchinstead ofasyncJobLaunch, adding stricter state checks:

📎 src/group.cc:113-116

cpp
} else {
  /* safety check */
  assert(state == ncclGroupJobJoined);
}

The legacy version usesWARNinstead ofassert, while the new architecture usesassert. This indicates that the new architecture has higher requirements for state machine correctness.

Design considerations

The motivation for the new architecture isdecoupling: The legacygroupLaunchLegacycrams all stages into one function, making it difficult to maintain and extend. The new architecture splits each stage into independent job types, connected through a queue. However, the scheduler and launcher are not yet implemented, so it is currently "framework first."

ncclParamEnqueueRearchEnable()controls whether to use the new architecture or legacy:

📎 src/group.cc:1031-1033

cpp
static ncclResult_t groupLaunch(struct ncclAsyncJob* job_, ncclSimInfo_t* simInfo = NULL) {
  return ncclParamEnqueueRearchEnable() ? groupLaunchEnqueueRearch(job_, simInfo) : groupLaunchLegacy(job_, simInfo);
}

Users can switch via the environment variableNCCL_ENQUEUE_REARCH_ENABLE. For production environments, it is recommended to keep the default (legacy), because the new architecture is still under development.

---

VI. Non-blocking group and asynchronous error handling

Scenario-driven Step-by-Step Walkthrough

The core of non-blocking group isncclGroupJobCompleteandncclGroupJobAbort:

📎 src/group.cc:1166-1190

cpp
ncclResult_t ncclGroupJobComplete(struct ncclGroupJob* groupJob) {
  ncclResult_t ret = ncclSuccess;
  if (groupJob && groupJob->nonBlockingInit) {
    if (!COMPILER_ATOMIC_EXCHANGE(&groupJob->joined, true, std::memory_order_acq_rel)) {
      ret = ncclAsyncJobComplete(&groupJob->base);
    }
    if (ncclAtomicRefCountDecrement(&groupJob->groupRefCount) == 0) {
      delete groupJob;
    }
  }
  return ret;
}

ncclResult_t ncclGroupJobAbort(struct ncclGroupJob* groupJob) {
  if (groupJob && groupJob->nonBlockingInit) {
    if (!COMPILER_ATOMIC_EXCHANGE(&groupJob->joined, true, std::memory_order_acq_rel)) {
      COMPILER_ATOMIC_STORE(&groupJob->abortFlag, true, std::memory_order_relaxed);
      ncclAsyncJobComplete(&groupJob->base);
    }
    if (ncclAtomicRefCountDecrement(&groupJob->groupRefCount) == 0) {
      delete groupJob;
    }
  }
  return ncclSuccess;
}

Key design:

1. joinedAtomic flag: UseCOMPILER_ATOMIC_EXCHANGEto ensure only one thread can execute the join logic. If two threads callncclGroupJobCompleteat the same time, only one will actually join, and the other will skip directly. This prevents double-join.

2. Reference counting:groupRefCountrecords how many comms are associated with this group job. Each comm increments the reference count inncclGroupEndInternal:

📎 src/group.cc:1108-1111

cpp
if (job->comm->groupJob == NULL) {
  job->comm->groupJob = groupJob;
  groupJob->groupRefCount++;
}

Only when all comms have calledncclGroupJobCompleteorncclGroupJobAbort, and the reference count drops to 0, is the group job deleted. This ensures that the group job's lifetime covers all associated comms.

3. abort semantics:ncclGroupJobAbortfirst setsabortFlag, then joins. The worker thread checksabortFlagduring execution, and if it finds it has been aborted, exits early. This is "cooperative cancellation" — not forcibly killing the thread, but letting the thread check the flag and exit on its own.

Production Pitfall Guide

Pitfall 3: Error querying for non-blocking group. Non-blocking group returnsncclInProgress, and the user needs to query progress throughncclCommGetAsyncError. If the user forgets to query and directly calls the next communication, they may encounterncclInProgresserrors. More seriously, if the group job is still running and the user callsncclCommDestroy, it will cause use-after-free. NCCL prevents this throughcomm->groupJobpointers and reference counting:ncclCommDestroywill first checkcomm->groupJob, and if there is an unfinished group job, it will wait or report an error.

Pitfall 4:ncclGroupJobCompletereturn value of. If the group job execution fails,ncclAsyncJobCompletereturns an error code. ButncclGroupJobCompleteonly returns this error code on the first call, and subsequent calls returnncclSuccess(becausejoinedis already true). The user must check the return value on the first call, otherwise the error information will be lost.

---

Chapter Summary

In this chapter, we broke down NCCL's complete scheduling chain from "task description" to "kernel launch":

1. Group semantics:ncclGroupStart/ncclGroupEndaccumulate tasks through thread_local variables,ncclGroupEndand submit them uniformly at . Blocking mode executes synchronously, while non-blocking mode creates a thread for asynchronous execution.

2. Task preparation:ncclPrepareTasksdetermines the algorithm/protocol,ncclPrepareTasksAndCollPreconnectand preconnects one by one by clique to avoid races in split comms.

3. Round scheduling:doLaunchesgroups by clique, uses a barrier to synchronize comms within the clique, and launches one kernel plan per round until all plans have been launched.

4. Asynchronous tasks:asyncJobLaunchuse an atomic state machine and busy waiting to manage asynchronous jobs, supporting fast failure and abort.

5. New architecture:groupLaunchEnqueueRearchis a new scheduling framework under development, currently falling back to legacydoLaunches。

The next chapter will enter the last mile of kernel launch:ncclLaunchKernelhow to turnncclKernelPlaninto a kernel actually executed on the GPU, and how the device side readsDevCommmetadata.

Chapter Reflection and Self-Test

Q1: IfncclGroupCommJoininncclMemoryStackPush(&comm->memScoped)is removed, what will happen? In what scenarios will it cause memory leaks or data corruption?

Reference analysis:ncclMemoryStackPushfor comm in group

At this point, the task description has become an executable launch plan: group semantics merge multiple API calls into a single submission, channel partitioning distributes tasks across multiple execution streams, and doLaunches' round scheduling ensures ordering and dependencies between kernels. But a plan is still just a plan—how does the host-side task description become a grid on the GPU? In the next chapter, we will dive into ncclLaunchKernel to see parameter preparation, kernel variant selection, and the cudaLaunchKernel call, completing the final leap from host to device.

CHAPTER 08

Chapter 8: Chapter 8: Kernel Launch and Device-Side Execution: From Host-Side Invocation to GPU Thread Block Startup

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 8 / 25

Chapter 8: Kernel Launch and Device-Side Execution: From Host-Side Invocation to GPU Thread Block Startup

In the previous chapter, we broke down how tasks are partitioned across multiple channels, how kernel launch parameters are generated, and the mechanisms for batch submission and dependency ordering under group semantics. Now, the launch plan is ready, but it is still only a host-side data structure. The core question this chapter answers is:ncclKernelPlanHow does it become a grid actually running on the GPU? We will follow thencclLaunchKernelcall chain to see how parameters are packed into kernel args, how kernel variants are selected,cuLaunchKernelExhow it is invoked, and how the device-sidencclKernelMainreads the work description from shared memory and dispatches it to the concrete implementation.

From Plan to Grid: A Panorama of the Launch Path

Before diving into details, let us first build an overall mental model. Think ofncclKernelPlanas a "construction blueprint": it records how many channels (how many blocks) to launch this time, how many threads per block, which work items to execute, and which kernel function to use. AndncclLaunchKernelis the action of "the construction crew entering the site"—it translates the information on the blueprint intoCUlaunchConfigthat the CUDA driver can understand, and then callscuLaunchKernelExto actually launch the grid onto the GPU.

Without this layer, all host-side scheduling (the previous chapter's channel partitioning, batch organization, and proxy op ordering) would be nothing but talk on paper; no kernel would run on the GPU, and communication would never happen. This is the final link in the end-to-end backbone, and also the boundary between host and device.

The entire launch path can be summarized in three stages:

1. Parameter preparation(finishPlan + uploadWork): organize the work structs, batch descriptors, and kernel args into a contiguous block of memory, deciding whether to place them in kernel parameters, in the FIFO, or in a persistent buffer.

2. Kernel launch(ncclLaunchKernel): compute grid/block dimensions, assemble launch attributes (CGA cluster, mem sync domain, launch completion event), and callcuLaunchKernelEx。

3. Device-side entry(ncclKernelMain): each block determines its own channelId based onblockIdx.xloads the work batch from args or the FIFO into shared memory, and then dispatches throughncclDevFuncTableto the concrete algorithm/protocol implementation.

The figure below shows the complete control flow from plan to grid, including the key branch decisions:

mermaid
flowchart TD
    plan["ncclKernelPlan<br/>channelMask / workBytes / kernelFn"]
    finish["finishPlan()<br/>决定 workStorageType"]
    check_budget{"sizeof(args)+batchBytes<br/>+workBytes <= workArgsBytes?"}
    args_type["workStorageType = Args<br/>work 直接放 kernel 参数"]
    fifo_type["workStorageType = Fifo/Persistent<br/>work 放外部缓冲区"]
    upload["uploadWork()<br/>拷贝 work 到目标缓冲区"]
    launch["ncclLaunchKernel()<br/>组装 CUlaunchConfig"]
    check_cluster{"compCap >= 90<br/>且 clusterSize > 0?"}
    add_cluster["添加 CLUSTER_DIMENSION<br/>+ SPREAD 调度策略"]
    no_cluster["不添加 cluster 属性"]
    check_event{"userKernelEvent<br/>且 driver >= 12030?"}
    add_event["添加 LAUNCH_COMPLETION_EVENT"]
    no_event["无 completion event"]
    cu_launch["cuLaunchKernelEx()<br/>发射 grid 到 GPU"]

    plan --> finish --> check_budget
    check_budget -->|是| args_type
    check_budget -->|否| fifo_type
    args_type --> upload
    fifo_type --> upload
    upload --> launch --> check_cluster
    check_cluster -->|是| add_cluster
    check_cluster -->|否| no_cluster
    add_cluster --> check_event
    no_cluster --> check_event
    check_event -->|是| add_event
    check_event -->|否| no_event
    add_event --> cu_launch
    no_event --> cu_launch

This figure anchors the three core functions of this chapter:finishPlan、uploadWork、ncclLaunchKernel. Next, we will break them down one by one.

Parameter Preparation: How the Work Struct Finds Its Place

Intuitive model

finishPlan's role is similar to the "packer" at a courier sorting center. It faces a pile of scattered work structs (one for each collective or p2p operation) and needs to decide: should these work items be stuffed into the "carry-on backpack" of kernel parameters, placed on the "conveyor belt" of the FIFO, or put into the "warehouse" of a persistent buffer?

If this decision is made incorrectly—for example, if the work is too large to fit into kernel parameters but is forced in anyway—the kernel launch will fail outright. If the work is placed in the wrong location, the device side will read garbage data, and the communication result will be completely wrong.

Data Structures and Memory Layout

First look atncclDevKernelArgs's structure; it is the "envelope" between host and device:

📎 src/include/device.h:514-522

c
struct alignas(16) ncclDevKernelArgs {
  struct ncclKernelComm* comm;      // 指向设备侧通信器元数据
  uint64_t channelMask;             // 哪些 channel 有工作
  enum ncclDevWorkStorageType workStorageType;  // work 存在哪里
  uint32_t workMask;                // FIFO 环形缓冲区的掩码
  void* workBuf;                    // work 缓冲区指针
  // struct ncclDevWorkBatch batches[];  // 紧随其后的是 batch 数组
};

This struct has only 5 fields, but each field carries critical information.channelMaskis a 64-bit mask, with each bit corresponding to a channel; the device side computes__popcllto determineblockIdx.x's corresponding channelId.workStorageTypedetermines where the device side reads work from:Argsmeans the work is in the kernel parameters,Fifomeans it is in the ring buffer,Persistentmeans it is in the persistent buffer.

ncclDevWorkBatchis the batch descriptor, which tells the device side "where the work for this channel is and how many there are":

📎 src/include/device.h:400-421

c
struct alignas(16) ncclDevWorkBatch {
  union {
    struct {
      uint32_t nextJump:14, nextExtends:1;
      uint32_t workType:2, funcId : NCCL_DEV_WORK_BATCH_FUNC_ID_BITS, func : NCCL_DEV_WORK_BATCH_FUNC_BITS;
    };
    uint32_t flags;
  };
  uint32_t offsetBase;    // work 在 FIFO 中的起始偏移
  uint64_t offsetBitset;  // 哪些 work 属于这个 channel
};

offsetBitsetis a 64-bit mask, with each bit corresponding to a work struct. The device side uses__popcandfns(find n-th set) instructions to locate the offset of each work.nextJumpandnextExtendsare used to chain multiple batches together—when there is too much work to fit in one batch, an "extended batch" is created.

Step-by-Step Walkthrough

Now consider a concrete scenario: one AllReduce is split across 4 channels, each channel has 2 work structs, for a total of 8 work items.

Step 1:finishPlandetermines the storage type.

📎 src/enqueue/enqueue.cc:245-255

c
if (sizeof(ncclDevKernelArgs) + batchBytes + workBytes <= comm->workArgsBytes) {
  plan->workStorageType = ncclDevWorkStorageTypeArgs;
}
plan->kernelArgsSize = sizeof(struct ncclDevKernelArgs) + batchBytes;
plan->kernelArgsSize += (plan->workStorageType == ncclDevWorkStorageTypeArgs) ? workBytes : 0;
plan->kernelArgsSize = alignUp(plan->kernelArgsSize, 16);
plan->kernelArgs =
  (struct ncclDevKernelArgs*)ncclMemoryStackAlloc(&comm->memScoped, plan->kernelArgsSize, /*align=*/16);
plan->kernelArgs->comm = comm->devComm;
plan->kernelArgs->channelMask = plan->channelMask;
plan->kernelArgs->workStorageType = plan->workStorageType;

The key judgment here is: ifsizeof(ncclDevKernelArgs) + batchBytes + workBytescan fit intocomm->workArgsBytes(usually 4KB), then put the work directly into the kernel parameters. Otherwise, the work is placed into the FIFO or a persistent buffer, and only the batch descriptor is placed in the kernel parameters.

[Design inference and architectural trade-offs]

Why prefer putting it in kernel parameters? Because kernel parameters are passed through constant memory in the CUDA driver, and when the device side reads them it uses theld.paraminstruction, which is much faster than reading the FIFO from global memory. For small messages (small total work), this can significantly reduce latency.

Step 2: Place batches into kernel args by rotating across channels.

📎 src/enqueue/enqueue.cc:257-280

c
uint64_t hasBatchMask = plan->channelMask;
struct ncclDevWorkBatch* batchPrev[MAXCHANNELS] = {};
struct ncclDevWorkBatch* batchZero = (struct ncclDevWorkBatch*)(plan->kernelArgs + 1);
int batchIx = 0;
while (hasBatchMask != 0) {
  uint64_t tmpMask = hasBatchMask;
  do {
    int c = popFirstOneBit(&tmpMask);
    if (!ncclIntruQueueEmpty(&wipChannels[c].workBatchQueue)) {
      struct ncclWorkBatchList* batchNode = ncclIntruQueueDequeue(&wipChannels[c].workBatchQueue);
      if (batchPrev[c] != nullptr) {
        batchPrev[c]->nextJump = int(&batchZero[batchIx] - batchPrev[c]);
      }
      batchPrev[c] = &batchZero[batchIx];
      batchZero[batchIx++] = batchNode->batch;
    }
    if (ncclIntruQueueEmpty(&wipChannels[c].workBatchQueue)) {
      hasBatchMask ^= 1ull << c;
    }
  } while (tmpMask != 0);
}

The logic of this code is "round-robin": in each round, take one batch from each channel that still has batches, and place them into thebatchZeroarray in ascending channel number order. The purpose of this is to ensure that "the first batch of each channel is located atbatchZero[blockIdx.x]"—each block on the device side directly indexes to its own first batch throughblockIdx.xwithout needing to search.

nextJumpThe field records the offset of the next batch in the same channel relative to the current batch. The device side can jump to the next batch throughbatchIx += batch.nextJump, forming a linked list.

Step 3:uploadWorkcopies the work to the target buffer.

📎 src/enqueue/enqueue.cc:1365-1430

c
static ncclResult_t uploadWork(struct ncclComm* comm, struct ncclKernelPlan* plan) {
  if (plan->isSymColl || plan->isCeColl || plan->isRma) return ncclSuccess;
  size_t workBytes = plan->workBytes;
  size_t batchBytes = plan->nWorkBatches * sizeof(struct ncclDevWorkBatch);
  void* fifoBufHost;
  uint32_t fifoCursor, fifoMask;
  switch (plan->workStorageType) {
  case ncclDevWorkStorageTypeArgs:
    plan->kernelArgs->workBuf = nullptr;
    fifoBufHost = (void*)plan->kernelArgs;
    fifoCursor = sizeof(ncclDevKernelArgs) + batchBytes;
    fifoMask = ~0u;
    break;
  case ncclDevWorkStorageTypeFifo:
    fifoBufHost = comm->workFifoBuf;
    fifoCursor = comm->workFifoProduced;
    fifoMask = comm->workFifoBytes - 1;
    NCCLCHECK(waitWorkFifoAvailable(comm, fifoCursor + workBytes));
    plan->kernelArgs->workBuf = comm->workFifoBufDev;
    break;
  // ...
  }
  plan->kernelArgs->workMask = fifoMask;
  // 修正 batch 的 offsetBase
  struct ncclDevWorkBatch* batchZero = (struct ncclDevWorkBatch*)(plan->kernelArgs + 1);
  for (int b = 0; b < plan->nWorkBatches; b++) {
    batchZero[b].offsetBase += fifoCursor;
  }
  // 拷贝 work 结构体
  struct ncclWorkList* workNode = ncclIntruQueueHead(&plan->workQueue);
  while (workNode != nullptr) {
    char* dst = (char*)fifoBufHost;
    char* src = (char*)(workNode + 1);
    for (int n = workNode->size; n != 0; n -= 16) {
      memcpy(COMPILER_ASSUME_ALIGNED(dst + (fifoCursor & fifoMask), 16), COMPILER_ASSUME_ALIGNED(src, 16), 16);
      fifoCursor += 16;
      src += 16;
    }
    workNode = workNode->next;
  }
  // ...
}

There are several key points here:

1. fifoCursorThe semantics of: for theArgstype, it is an offset relative to thekernelArgsstart address; for theFifotype, it is an offset relative to the FIFO base address; for thePersistenttype, it starts from 0.

2. offsetBaseCorrection of:finishPlanThe batch'soffsetBaseinuploadWorkis relative to the work start position of the plan (starting from 0).ArgsIt needs to be converted into an offset relative to the actual storage location. For thesizeof(ncclDevKernelArgs) + batchBytestype, addFifo; for thecomm->workFifoProduced。

3. type, add16-byte aligned copyalignas(16): work structs are all 16-byte aligned (COMPILER_ASSUME_ALIGNED), so copying is done in units of 16 bytes.

4. tells the compiler that this address is 16-byte aligned, allowing the compiler to generate more efficient vectorized instructions.FIFO waitFifo: for thewaitWorkFifoAvailabletype,comm->abortFlagspins waiting for the FIFO to have enough space. This wait checks

to avoid deadlock on abort.

Design considerations and production pitfalls

[Design inference and architectural trade-offs]Why are there three storage types?

  • ArgsThis is a trade-off between space and latency:
  • Fifo: fastest (constant memory), but limited capacity (4KB). Suitable for small messages and a small amount of work.
  • Persistent: large capacity (ring buffer), but device-side reads must go through global memory. Suitable for medium messages.cudaMemcpy: used for CUDA Graph capture scenarios. Because graph capture cannot perform

, a persistent buffer must be preallocated, the work copied into it, and then the kernel reads from there.Pitfall 1: FIFO overflow causing deadlock.waitWorkFifoAvailableIfabortFlagdoes not check📎 src/enqueue/enqueue.cc:1333-1349, then when the FIFO is full and the consumer (GPU kernel) stops consuming for some reason, the host will spin forever. In the source code,

c
if (COMPILER_ATOMIC_LOAD(comm->abortFlag, std::memory_order_acquire)) {
  return ncclInternalError;
}

CopyoffsetBitsetPitfall 2: offsetBitsetoverflow.1ull << (offset / workSize)is 64-bit and supports at most 64 work items in one batch. If there are more than 64,NCCL_MAX_DEV_WORK_BATCH_BYTESwill overflow. In the source code,ncclDevWorkColllimits the batch size (1024 bytes), and the smallest work struct is

(about 80 bytes), so there are at most 12 work items, and it will not overflow.Pitfall 3: Memory leak in Persistent mode.uploadWorkInPersistent'sfifoBufHostbranch,ncclOsAlignedAllocis allocated throughuploadWork_cleanup_fnand needs to be freed incudaMemcpyAsync. Iffailfails, thecleanuplabel checks whetherfifoBufHostis null, and if it is null, directly frees📎 src/enqueue/enqueue.cc:1483-1485. This error recovery chain can be seen in

Kernel launch: from CUlaunchConfig to cuLaunchKernelEx

Intuitive model

ncclLaunchKernelrole is similar to a "rocket launch console." It receives a plan already loaded with fuel (work data), calculates the rocket's flight parameters (grid/block dimensions), sets various launch options (cluster, mem sync domain, completion event), and then presses the launch button (cuLaunchKernelEx)。

If something goes wrong at this stage—for example, if the grid dimensions are calculated incorrectly—the wrong number of blocks will be launched on the GPU, causing the work of some channels to never be executed and communication to hang.

Data Structures and Memory Layout

CUlaunchConfigis the launch configuration struct of the CUDA driver API, and NCCL constructs it on the stack:

📎 src/enqueue/enqueue.cc:1916-1917

c
CUlaunchConfig launchConfig = {0};
CUlaunchAttribute launchAttrs[6] = {};
int attrs = 0;

launchAttrsis an array of at most 6 elements, each element being aCUlaunchAttribute. NCCL conditionally adds different attributes based on hardware capabilities and driver version:

  • CU_LAUNCH_ATTRIBUTE_CLUSTER_DIMENSION: CGA cluster dimension (sm90+)
  • CU_LAUNCH_ATTRIBUTE_CLUSTER_SCHEDULING_POLICY_PREFERENCE: cluster scheduling policy
  • CU_LAUNCH_ATTRIBUTE_MEM_SYNC_DOMAIN: memory sync domain (CUDA 12.0+)
  • CU_LAUNCH_ATTRIBUTE_LAUNCH_COMPLETION_EVENT: launch completion event (CUDA 12.3+)
  • CU_LAUNCH_ATTRIBUTE_PROGRAMMATIC_STREAM_SERIALIZATION: programmatic stream serialization (sym kernel)
  • CU_LAUNCH_ATTRIBUTE_NVLINK_UTIL_CENTRIC_SCHEDULING: NVLink utilization-centric scheduling (CUDA 13.0+)

Step-by-Step Walkthrough

Step 1: Calculate grid and block dimensions.

📎 src/enqueue/enqueue.cc:1889-1893

c
int nChannels = countOneBits(plan->channelMask);
void* sym = plan->kernelFn;
dim3 grid = {(unsigned)nChannels, 1, 1};
dim3 block = {(unsigned)plan->threadPerBlock, 1, 1};
int smem = plan->isSymColl ? plan->kernelDynSmem : ncclShmemDynamicSize(comm->cudaArch);

nChannelsischannelMaskthe number of set bits in , that is, how many blocks this plan needs to launch. Each block is responsible for one channel.threadPerBlockis inscheduleCollTasksToPlancomputed viaplan->threadPerBlock = std::max(plan->threadPerBlock, task->nWarps * WARP_SIZE), taking the maximum among all tasksnWarps * 32。

smemis the dynamic shared memory size. For a normal kernel, it isncclShmemDynamicSize(comm->cudaArch), which is a compile-time constant depending on the architecture (sm70+ isncclShmemScratchWarpSize * (NCCL_MAX_NTHREADS / WARP_SIZE)). For a sym kernel, it isplan->kernelDynSmem, because the shared memory requirements of a sym kernel may differ.

Step 2: Assemble kernel arguments.

📎 src/enqueue/enqueue.cc:1902-1903

c
void* extra[] = {CU_LAUNCH_PARAM_BUFFER_POINTER, plan->kernelArgs, CU_LAUNCH_PARAM_BUFFER_SIZE, &plan->kernelArgsSize,
                 CU_LAUNCH_PARAM_END};

This is a way of passing arguments in the CUDA driver API:CU_LAUNCH_PARAM_BUFFER_POINTERtells the driver "the arguments are not passed one by one, but as one contiguous memory block,"CU_LAUNCH_PARAM_BUFFER_SIZEtells the driver the size of this block. The advantage of doing this is that NCCL can passncclDevKernelArgsand the following batch array all at once, without needing to pack each argument individually.

Step 3: Add launch attributes.

📎 src/enqueue/enqueue.cc:1929-1936

c
if (clusterSize) {
  if (grid.x % clusterSize) clusterSize = 1;
  launchAttrs[attrs].id = CU_LAUNCH_ATTRIBUTE_CLUSTER_DIMENSION;
  launchAttrs[attrs++].value.clusterDim = {clusterSize, 1, 1};
  launchAttrs[attrs].id = CU_LAUNCH_ATTRIBUTE_CLUSTER_SCHEDULING_POLICY_PREFERENCE;
  launchAttrs[attrs++].value.clusterSchedulingPolicyPreference = CU_CLUSTER_SCHEDULING_POLICY_SPREAD;
}

CGA (Cooperative Group Array) is a hardware feature introduced in sm90 that allows multiple blocks to be grouped into a cluster. Blocks within a cluster are guaranteed to be scheduled simultaneously onto a set of SMs and can access each other's shared memory. NCCL uses this feature to implement algorithms such as NVLS that require cross-block synchronization.

Note theif (grid.x % clusterSize) clusterSize = 1;protection: the cluster dimension must evenly divide the grid dimension, otherwise the driver will report an error. Ifgrid.xis not divisible byclusterSize, it degrades to not using a cluster.

Step 4: Add launch completion event.

📎 src/enqueue/enqueue.cc:1944-1964

c
#if CUDART_VERSION >= 12030
enum ncclImplicitOrder implicitOrder;
NCCLCHECKGOTO(getImplicitOrder(&implicitOrder, comm, plan->persistent, driverVersion), ret, do_return);
if (implicitOrder == ncclImplicitOrderLaunch) {
  launchAttrs[attrs].id = CU_LAUNCH_ATTRIBUTE_LAUNCH_COMPLETION_EVENT;
  launchAttrs[attrs].value.launchCompletionEvent.event = comm->sharedRes->launchEvent;
  launchAttrs[attrs].value.launchCompletionEvent.flags = 0;
  attrs++;
  if (userKernelEvent) {
    NCCLCHECKGOTO(ncclUncapturedStreamPoolAcquire(&comm->sharedRes->uncapturedStreamPool, &relayStream), ret, do_return);
    relayUserLaunchCompletionEvent = true;
    userKernelEventArmed = true;
  }
} else if (userKernelEvent && driverVersion >= 12030) {
  launchAttrs[attrs].id = CU_LAUNCH_ATTRIBUTE_LAUNCH_COMPLETION_EVENT;
  launchAttrs[attrs].value.launchCompletionEvent.event = plan->launchCompletionEvent;
  launchAttrs[attrs].value.launchCompletionEvent.flags = 0;
  attrs++;
  userKernelEventArmed = true;
}
#endif

CU_LAUNCH_ATTRIBUTE_LAUNCH_COMPLETION_EVENTis a feature introduced in CUDA 12.3: the driver records an event when the kernel actually starts executing (rather than when the host-side call returns). This is crucial for implementing "implicit order"—NCCL needs to ensure that multiple kernels execute in order, but does not want the host side to block and wait.

getImplicitOrderThe logic of is: if the user has setlaunchOrderImplicit, and the driver version is new enough, usencclImplicitOrderLaunch(ordering via launch event); otherwise usencclImplicitOrderSerial(ordering via completion event, i.e., serial execution).

Step 5: CallcuLaunchKernelEx。

📎 src/enqueue/enqueue.cc:1978-1996

c
launchConfig.gridDimX = grid.x;
launchConfig.gridDimY = grid.y;
launchConfig.gridDimZ = grid.z;
launchConfig.blockDimX = block.x;
launchConfig.blockDimY = block.y;
launchConfig.blockDimZ = block.z;
launchConfig.sharedMemBytes = smem;
launchConfig.attrs = launchAttrs;
launchConfig.numAttrs = attrs;
launchConfig.hStream = launchStream;
if (userKernelEvent && !userKernelEventArmed) {
  WARN("CUDA launch-completion events require CUDA 12.3 or newer; recording the user event before launch");
  CUDACHECKGOTO(cudaEventRecord(plan->launchCompletionEvent, launchStream), ret, do_return);
}
CUCHECKGOTO(cuLaunchKernelEx(&launchConfig, fn, nullptr, extra), ret, do_return);
if (relayUserLaunchCompletionEvent) {
  CUDACHECKGOTO(cudaStreamWaitEvent(relayStream, comm->sharedRes->launchEvent, 0), ret, do_return);
  CUDACHECKGOTO(cudaEventRecord(plan->launchCompletionEvent, relayStream), ret, do_return);
}

cuLaunchKernelExis a new API introduced in CUDA 12.0 that supports launch attributes. For older drivers (< 11.8), NCCL falls back tocuLaunchKernel:

📎 src/enqueue/enqueue.cc:1998-2007

c
} else {
  // Standard kernel launch
  if (userKernelEvent) {
    WARN("CUDA launch-completion events require CUDA 12.3 or newer; recording the user event before launch");
    CUDACHECKGOTO(cudaEventRecord(plan->launchCompletionEvent, launchStream), ret, do_return);
  }
  CUCHECKGOTO(cuLaunchKernel(fn, grid.x, grid.y, grid.z, block.x, block.y, block.z, smem, launchStream, nullptr,
                             extra),
              ret, do_return);
}

Concurrency Control and Hardware Interaction

The relay mechanism of the Launch completion event.When usingncclImplicitOrderLaunchand the user has providedlaunchCompletionEvent, NCCL cannot pass the user's event directly to the driver, because the driver supports only one launch completion event. NCCL's approach is:

1. Passcomm->sharedRes->launchEventto the driver.

2. Wait onrelayStreamforlaunchEvent。

3. Record the user's event onrelayStreamThis way, the user's event is triggered after the kernel actually starts executing, rather than when the host-side call returns.

On sm90+, NCCL sets

Mem Sync Domain。 📎 src/enqueue/enqueue.cc:1938-1942toCU_LAUNCH_ATTRIBUTE_MEM_SYNC_DOMAIN. This is the memory sync domain mechanism introduced by the Hopper architecture, used to isolate memory barriers of different kernels and reduce unnecessary synchronization overhead.cudaLaunchMemSyncDomainRemoteProduction Pitfall Guide

Pitfall 1: Cluster dimension not evenly dividing causes launch failure.

Ifis not divisible bygrid.x, the driver returnsclusterSize. The source code protects against this viaCUDA_ERROR_INVALID_VALUE, but this also means the cluster feature is silently disabled. If the user expects the performance improvement brought by clusters, they need to checkif (grid.x % clusterSize) clusterSize = 1;andcgaClusterSizethe relationship betweennChannelsPitfall 2: Driver version not meeting requirements causes the kernel to be unavailable.

踩坑点 2:驱动版本不满足导致 kernel 不可用。 ncclInitKernelsForDevicechecks the driver requirements of each kernel during initialization:

📎 src/enqueue/enqueue.cc:71-76

c
for (int k = 0; k < kcount; k++) {
  if (kptrs[k] != nullptr && driverVersion < krequires[k]) {
    INFO(NCCL_INIT, "Skipping %skernel %d which requires driver %d", sym ? "symmetric " : "", k, krequires[k]);
    kptrs[k] = nullptr;
    if (kptrsProfile != nullptr) kptrsProfile[k] = nullptr;
  }

If the driver version is insufficient, the kernel pointer will be set to null. Later, if the scheduler selects this kernel,cuLaunchKernelExwill fail. NCCL's tuner should avoid selecting unavailable kernels, but if the user forcibly specifies an algorithm (NCCL_ALGO), this problem may be triggered.

Pitfall 3:launchCompletionEventBehavior on old drivers.If the driver version < 12.3, NCCL will record an event before the kernel starts, which means the event will be triggered before the kernel begins execution, rather than when the kernel actually starts executing. This may invalidate timing assumptions in user code.

Device-side entry: from blockIdx to the concrete implementation

Intuitive model

ncclKernelMainis the "entrance hall" for each block on the GPU. When a block is scheduled onto an SM and starts executing, it first enters this hall and completes three things: determine its own identity (which channel am I), claim its own task (load the work batch), and then go to the corresponding window to do the work (call the concrete algorithm implementation).

Without this entry, every kernel variant would need to handle the questions "who am I and what am I supposed to do" on its own, and the code would be heavily duplicated.ncclKernelMainThrough the template parametersSpecializedFnIdandSpecializedRunWorkBatchimplements the pattern of "generic entry + specialized execution."

Data structures and memory layout

The device-side shared memory layout is the key to understandingncclKernelMain.ncclShmemDatais the "workbench" shared by all blocks:

📎 src/device/common.h:48-72

c
struct ncclShmemData {
  struct ncclDevKernelArgs args;
  int channelId;
  int aborted;
  alignas(16) struct ncclKernelComm comm;
  alignas(16) struct ncclDevChannel channel;

  int batchIx, nextBatchIx;
  enum ncclDevWorkType workType;
  uint8_t directMode;
  uint16_t funcId;
  int nWorks;
  int workSize;
  uint64_t workCounter;
  bool profilerEnabled;
  uint8_t func;
  struct ncclShmemGroup groups[NCCL_MAX_GROUPS];

  alignas(16) char workStorage[ncclMaxDevWorkBatchBytes()];

  alignas(16) union {
    unpackShmem unpack;
  } devicePlugin;
};

The layout of this struct has been carefully designed:

  • argsis placed first because it is copied from the kernel parameters and needs 16-byte alignment.
  • commandchannelare also 16-byte aligned because they are copied viacopyToShmem16using vectorized instructions.
  • workStorageis the temporary storage area for the work struct, and its size isncclMaxDevWorkBatchBytes()(16KB on sm90+).
  • groupsThe array is used to store the connection information of each group,NCCL_MAX_GROUPSis 16.

Step-by-Step Walkthrough

Step 1: Copy kernel args to shared memory.

📎 src/device/common.h:426-428

c
if (tid < sizeof(ncclDevKernelArgs) / sizeof(uint32_t)) {
  ((uint32_t*)&ncclShmem.args)[tid] = ((uint32_t*)args)[tid];
}

Here the firstsizeof(ncclDevKernelArgs) / 4threads are used, and each thread copies one 32-bit word. Why copy to shared memory? Because kernel parameters are in constant memory. Although access is fast, there is broadcast overhead when every thread needs to access them. After copying to shared memory, all threads access the same block of shared memory, which is more efficient.

Step 2: Determine channelId.

📎 src/device/common.h:430-437

c
if (tid < MAXCHANNELS && (args->channelMask & (1ull << tid))) {
  int n = __popcll(args->channelMask & ((1ull << tid) - 1));
  if (blockIdx.x == n) ncclShmem.channelId = tid;
}
__syncthreads();

The logic of this code is: for each set channel (args->channelMask & (1ull << tid)), calculate how many set channels precede it (__popcll). If this count equalsblockIdx.x, then the current block is responsible for this channel.

For example:channelMask = 0b1011(channels 0, 1, and 3 have work).blockIdx.x = 0The block withblockIdx.x = 1is responsible for channel 0 (0 set bits before it),blockIdx.x = 2the block with

is responsible for channel 1 (1 set bit before it),

📎 src/device/common.h:446-478

c
switch (tid / WARP_SIZE) {
case 0:
  {
    void* dst = &ncclShmem.comm;
    void* src = ncclShmem.args.comm;
    int bytes = sizeof(ncclKernelComm);
    static_assert(sizeof(ncclKernelComm) <= 16 * WARP_SIZE,
                  "ncclKernelComm cannot be loaded by a single warp in one insn.");
    copyToShmem16(tid, dst, src, bytes);
  }
  break;
case 1:
  {
    void* dst = &ncclShmem.channel;
    void* src = &((ncclKernelCommAndChannels*)ncclShmem.args.comm)->channels[ncclShmem.channelId];
    int bytes = sizeof(ncclDevChannel);
    static_assert(sizeof(ncclDevChannel) <= 16 * WARP_SIZE,
                  "ncclDevChannel cannot be loaded by a single warp in one insn.");
    copyToShmem16(tid - WARP_SIZE, dst, src, bytes);
  }
  break;
default:
  {
    int subtid = tid - 2 * WARP_SIZE;
    int subtn = tn - 2 * WARP_SIZE;
    loadWorkBatchToShmem(subtid, subtn, args, /*batchIx=*/blockIdx.x);
  }
  break;
}
__syncthreads();

is responsible for channel 3 (2 set bits before it).

  • Step 3: Load comm and channel into shared memory.CopyncclKernelCommHere the threads are divided into three groups:
  • warp 0: loadncclDevChannel(communicator metadata) into shared memory.
  • warp 1: load the current channel's

copyToShmem16(channel metadata) into shared memory.

📎 src/device/common.h:131-139

c
inline __device__ void copyToShmem16(int tid, void* dst, void const* src, int bytes) {
  int offset = 16 * tid;
  if (offset < bytes) {
    uint64_t a = 0, b = 0;
    asm volatile("ld.v2.u64 {%0,%1},[%2];" : "=l"(a), "=l"(b) : "l"((char const*)src + offset) : "memory");
    uint32_t udst = (uint32_t)__cvta_generic_to_shared(dst);
    asm volatile("st.shared.v2.u64 [%0],{%1,%2};" ::"r"(udst + offset), "l"(a), "l"(b) : "memory");
  }
}

: load the work batch into shared memory.ld.v2.u64is a 16-byte copy function implemented with inline PTX:st.shared.v2.u64Copy__cvta_generic_to_sharedIt uses

to load 16 bytes from global memory and

loadWorkBatchToShmemto store them to shared memory.workStorageconverts a generic address into a shared memory address (the shared memory address space is 32-bit).

📎 src/device/common.h:142-260

c
__device__ __forceinline__ void loadWorkBatchToShmem(int tid, int tn, struct ncclDevKernelArgs const* args,
                                                     int batchIx) {
  int lane = tid % WARP_SIZE;
  int workCursor = 0;
  while (true) {
    struct ncclDevWorkBatch batch = ((struct ncclDevWorkBatch*)(args + 1))[batchIx];

    uint8_t* fnsOfBitset = (uint8_t*)ncclScratchForWarp(threadIdx.x / WARP_SIZE);
    __syncwarp();
    if (uint32_t(batch.offsetBitset) & (1u << lane)) {
      int nWorksBelow = __popc(uint32_t(batch.offsetBitset) & ((1u << lane) - 1));
      fnsOfBitset[nWorksBelow] = lane;
    }
    int nWorksLow32 = __popc(uint32_t(batch.offsetBitset));
    if (uint32_t(batch.offsetBitset >> 32) & (1u << lane)) {
      int nWorksBelow = nWorksLow32;
      nWorksBelow += __popc(uint32_t(batch.offsetBitset >> 32) & ((1u << lane) - 1));
      fnsOfBitset[nWorksBelow] = 32 + lane;
    }
    int nWorks = nWorksLow32 + __popc(uint32_t(batch.offsetBitset >> 32));
    __syncwarp();
    // ...
  }
}

is the most complex part. Its task is to copy the work struct pointed to by the batch descriptor from global memory (or kernel parameters) into the shared memoryfnsOfBitset.offsetBitsetCopyfnsThe core of this code is computingfnsOfBitset[nWorksBelow]。

: for the nth set bit in

📎 src/device/common.h:209-241

c
if (tid < nPacks) {
  int srcWork = fnsOfBitset[dstWork];
  ulonglong2 tmp;
  if (ncclShmem.args.workStorageType == ncclDevWorkStorageTypeArgs) {
    char* src = (char*)args + (batch.offsetBase + srcWork * workSize + packInWork * 16);
    tmp = *(ulonglong2*)src; // becomes ld.param.v2.u64
  } else {
    char* src = (char*)ncclShmem.args.workBuf +
                ((batch.offsetBase + srcWork * workSize + packInWork * 16) & ncclShmem.args.workMask);
    tmp = *(ulonglong2*)src; // becomes ld.v2.u64
  }
  char* dst = ncclShmem.workStorage;
  dst += (workCursor + dstWork) * workSize + packInWork * 16;
  *(ulonglong2*)dst = tmp;
}

instruction to do this, but it expands into many SASS instructions. NCCL's approach is to use shared memory: each lane checks whether its own bit is set, and if so, calculates how many set bits precede it, then writes its own lane number toArgsNext is the actual copy:(char*)args + offsetCopyld.param.v2.u64There is a key optimization here: for theFifotype, the source code directly writes(char*)ncclShmem.args.workBuf + (offset & workMask), and the compiler will recognize that this is a read from kernel parameters and generate theld.v2.u64instruction. For the

type, the source code writes

📎 src/device/common.h:212-229

c
// The loads done in these two cases must be kept separate since we are
// relying on the compiler to use "ld.param" in the first one. The parameter
// space is not generically addressable, so any attempt to load through
// a pointer that *might* be parameter space backed will cause the
// compiler to spill the parameter struct (4K!) to each thread's local space
// before creating a pointer (to the spill) and decimate perf.

instruction.

The comments specifically emphasize that these two cases must not be merged:

📎 src/device/common.h:481-497

c
while (ncclShmem.aborted == 0) {
  profiler(START);
  if (0 <= SpecializedFnId && ncclShmem.funcId == (unsigned)SpecializedFnId) {
    SpecializedRunWorkBatch().run();
  } else {
    ncclDevFuncTable[ncclShmem.funcId]();
  }

  if (ncclShmem.nextBatchIx == -1) break;
  int batchIx = ncclShmem.nextBatchIx;
  __syncthreads();
  profiler(STOP);
  if (ncclShmem.comm.progressCounters != nullptr) __syncthreads();
  loadWorkBatchToShmem(tid, tn, args, batchIx);
  __syncthreads();
}

If the compiler cannot determine whether the pointer points to parameter space or global space, it will spill the entire parameter struct (4KB) to each thread's local memory, and performance will drop sharply.SpecializedFnIdStep 5: Execute the work.funcIdCopySpecializedRunWorkBatch().run()There is an important optimization here: ifncclDevFuncTable[ncclShmem.funcId]()matches the current batch's

ncclDevFuncTable, directly callgenerate.pyGenerate:

📎 src/device/generate.py:261-270

python
out("__device__ ncclDevFuncPtr_t const ncclDevFuncTable[] = {\n")
index = 0
for fn in primary_funcs:
  sym = paste("_", "ncclDevFunc", *fn)
  cudart, arch = required_cuda(*fn)
  if (cudart, arch) != (0, 0):
    out("#if CUDART_VERSION >= %d && __CUDA_ARCH__ >= %d\n" % (cudart ,arch))
  out("/*%4d*/ %s,\n" % (index, sym))
  if (cudart, arch) != (0, 0):
    out("#else\n" "/*%4d*/ nullptr,\n" "#endif\n" % index)
  index += 1
out("nullptr};\n")

Design Thinking and Production Pitfalls

Why use__grid_constant__? 📎 src/device/common.h:19-24

c
#if __CUDA_ARCH__ >= 700
// __grid_constant__ appears to break cuda-gdb
#define NCCL_GRID_CONSTANT __grid_constant__
#else
#define NCCL_GRID_CONSTANT
#endif

__grid_constant__Tells the compiler this parameter is read-only and can be placed in constant memory. This way, device-side reads go throughld.paraminstructions, which is faster than reading from global memory. The comment mentions it breaks cuda-gdb, so it's only enabled on sm70+.

Pitfall 1:workStorageoverflow. workStorageThe size of isncclMaxDevWorkBatchBytes(), sm90+ is 16KB. IfnWorks * workSizeexceeds this value, it will write out of bounds. In the source code,NCCL_MAX_DEV_WORK_BATCH_BYTESlimits the batch size on the host side, but there is no additional check on the device side. If the host-side constraint is bypassed (e.g., by modifying environment variables), it will cause shared memory out-of-bounds.

Pitfall 2:__syncthreads()The absence of causes data races.AfterloadWorkBatchToShmem, there must be a__syncthreads()for all threads to see the completeworkStorage. In the source code, at📎 src/device/common.h:479there is__syncthreads(); // publish ncclShmem. If this synchronization is removed, some threads may start reading beforeworkStoragehas finished writing, resulting in reading garbage data.

Pitfall 3: Timing of abort checks. while (ncclShmem.aborted == 0)Only checks abort at the start of each batch. If a batch takes a long time to execute, the abort signal may take a long time to take effect. This is a design trade-off: more frequent checks add overhead, but respond faster.

Kernel Variant Selection: How generate.py Generates the Kernel List

Intuitive Model

generate.py's role is similar to a "production line planner in a car factory." It faces a huge combinatorial space (7 set operations × 5 reduction operations × 12 data types × 7 algorithms × 3 protocols) and needs to decide: which combinations need dedicated kernels generated? Which can share a generic kernel?

If a kernel is generated for every combination, compilation time and binary size will explode. If only one generic kernel is generated, runtime will be slow due to function pointer calls and branch judgments.generate.py's solution is "representative kernels": generate one kernel for each equivalence class, and dispatch at runtime through a function pointer table.

Data Structures and Memory Layout

generate.pyGenerates three key files:

1. device_table.cu: device-sidencclDevFuncTable, mapping funcId to specific device functions.

2. host_table.cc: host-sidencclDevKernelList、ncclDevKernelForFunc、ncclDevFuncRowToIdand other tables.

3. Each<coll>_<op>_<ty>.cu: specific kernel implementation.

Step-by-Step Walkthrough

Step 1: Enumerate all function rows.

📎 src/device/generate.py:186-199

python
def enumerate_func_rows():
  yield ("SendRecv", None, None, None, None)
  for coll in ("AllGather", "Broadcast", "AllGatherV"):
    algos = algos_of_coll[coll]
    for algo in algos:
      for proto in all_protos:
        yield (coll, None, None, algo, proto)
  for coll in ("AllReduce", "Reduce", "ReduceScatter"):
    algos = algos_of_coll[coll]
    for redop in all_redops:
      for ty in all_tys:
        for algo in algos:
          for proto in all_protos:
            yield (coll, redop, ty, algo, proto)

This enumeration order must matchncclDevFuncId()'s calculation formula:

📎 src/include/device.h:646-706

c
inline int ncclDevFuncId(int coll, int devRedOp, int type, int algo, int proto) {
  constexpr int NumTypes = ncclNumTypes;
  int row;
  do {
    row = 0; // ncclDevFuncIndex_P2p
    if (coll == ncclFuncSendRecv) break;
    row += 1;
    // ...
  } while (false);
  return ncclDevFuncRowToId[row];
}

ncclDevFuncIdWhat is calculated is the "row number," which is then mapped throughncclDevFuncRowToIdto the "main function ID." The reason for this mapping is: many rows may map to the same main function (e.g., allAllReduce Sum i32rows map toAllReduce Sum u32's main function).

Step 2: Compute main functions and kernel functions.

📎 src/device/generate.py:211-225

python
func_rows = [validate(*fn) for fn in enumerate_func_rows()]
primary_funcs = sorted(set(equivalent_primary(*fn) for fn in func_rows if fn is not None))
primary_to_index = {fn: i for (i,fn) in zip(range(len(primary_funcs)), primary_funcs)}
kernel_funcs = sorted(set(best_kernel(*fn) for fn in primary_funcs))

equivalent_primaryMaps signed integers to unsigned integers (because addition/multiplication are the same for both):

📎 src/device/generate.py:158-166

python
def equivalent_primary(coll, redop, ty, algo, proto):
  if coll in ("AllReduce", "Reduce", "ReduceScatter"):
    if redop in ("Sum","Prod","PreMulSum","SumPostDiv") and ty[0]=="i":
      return (coll, redop, "u"+ty[1:], algo, proto)
    if redop=="MinMax" and ty[0]=="i" and ("NVLS" not in algo):
      return (coll, redop, "u"+ty[1:], algo, proto)
  return (coll, redop, ty, algo, proto)

best_kernelMaps multiple main functions to the same kernel (e.g., allAllGatheralgorithms map toAllGather RING LL):

📎 src/device/generate.py:171-183

python
def best_kernel(coll, redop, ty, algo, proto):
  def best(coll, redop, ty, algo, proto):
    if coll=="Nop": return ("Generic", None, None, None, None)
    if coll=="SendRecv": return ("SendRecv", None, None, None, None)
    if exact_kernel_names: return (coll, redop, ty, algo, proto)
    if coll in ("AllGather","Broadcast","AllGatherV"): return (coll, None, None, "RING", "LL")
    return (coll, "Sum", ty, ("TREE" if algo=="TREE" else "RING"), "LL")
  kfn = equivalent_primary(*best(coll, redop, ty, algo, proto))
  if not func_filter(*kfn): return ("Generic", None, None, None, None)
  return kfn

Step 3: Generate kernel definitions.

📎 src/device/generate.py:458-480

python
(_, kfns) = name_to_kernels.get(name) or (None, [])
for kfn in kfns:
  (coll, redop, ty, algo, proto) = kfn
  sym = kernel_suffix(kfn)
  fn_id = primary_to_index[kfn]
  cudart, arch = required_cuda(*kfn)
  s = "DEFINE_ncclDevKernel({sym}, ncclFunc{coll}, {redop_cxx}, {ty_cxx}, NCCL_ALGO_{algo}, NCCL_PROTO_{proto}, {fn_id})\n"
  # ...
  out(s.format(...))

DEFINE_ncclDevKernelAfter macro expansion:

📎 src/device/common.h:507-509

c
#define DEFINE_ncclDevKernel(suffix, coll, redop, ty, algo, proto, specializedFnId) \
  __global__ void ncclDevKernel_##suffix(ncclDevKernelArgs4K NCCL_GRID_CONSTANT const args4K) { \
    ncclKernelMain<specializedFnId, RunWorkBatch<coll, ty, redop<ty>, algo, proto>>(&args4K.args); \
  }

So each kernel is a__global__function, callingncclKernelMain, with template parametersspecializedFnIdandRunWorkBatch<coll, ty, redop<ty>, algo, proto>。

Design Thinking and Production Pitfalls

[Design Inference and Architectural Trade-offs]

Why use "representative kernels" instead of one kernel per combination?Trade-off between compilation time and binary size. The full combinatorial space is 7 × 5 × 12 × 7 × 3 ≈ 8820 kernels, each kernel takes a few seconds to compile, totaling several hours. And the binary size would reach hundreds of MB. By mapping to representative kernels, the actual number of generated kernels is reduced to a few dozen.

Pitfall 1:NCCL_EXACT_KERNEL_NAMEScauses compilation explosion.If this environment variable is set,best_kernelreturns the original function, and a kernel is generated for every combination. This is useful during development (you can precisely control which kernel is compiled), but in production it causes excessively long compilation times.

Pitfall 2:required_cuda's version check.Some kernels require a specific CUDA version or architecture:

📎 src/device/generate.py:130-154

At this point, the kernel has been launched on the GPU, and the device side has obtained the work descriptor. But what really determines performance is how data is moved inside the device. The next chapter will dive into the three protocol primitives under src/device: LL, LL128, and Simple, to see why the same AllReduce logic requires three sets of transport primitives, and their differences in synchronization methods, buffer layouts, and flag semantics.

CHAPTER 09

Chapter 9: Chapter 9: Device-Side Communication Primitives: Data Transport Implementation of the Three Protocols LL, LL128, and Simple

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 9 / 25

Chapter 9: Device-Side Communication Primitives: Data Transport Implementation of the Three Protocols LL, LL128, and Simple

In the previous chapter, we traced how the host side translates an AllReduce into a __global__ kernel, and saw that the device-side entry point ncclKernelMain dispatches based on algorithm and protocol. But dispatch only selects the tool; what truly determines performance is how these tools execute data movement. This chapter dives deep into the three sets of data movement primitives under src/device: LL, LL128, and Simple, dissecting their data movement implementations one by one to understand the trade-offs different protocols make between latency and bandwidth.

Why the same AllReduce needs three sets of data movement primitives

Let's first build an intuitive model. Imagine a pipeline factory: raw materials (user data) enter from one end, finished products exit from the other, and in between there are several workstations (ranks) that need to exchange semi-finished products with each other. There are three ways to move semi-finished products:

  • LL(Low Latency): Like two people passing a note face-to-face; the moment it's handed over, the other person knows "this is for you," with almost zero handshake overhead. But the note is very small—only 8 bytes of effective data can be passed at a time. Suitable for small messages.
  • LL128: Replace the note with a 128-byte sticky note, passing 120 bytes of effective data at a time, but the sticky note must be placed 16-byte aligned, otherwise it must first be "re-laid out" in shared memory. Suitable for medium messages.
  • Simple: Like a parcel locker—first put the package into the locker (FIFO buffer), then send a notification "locker number N has goods." Handshake overhead is high, but a lot can be moved at once. Suitable for large messages.
[Design inference and architectural trade-offs]

What if there were only one set of primitives? With only LL, large messages would crush bandwidth because "every message must wait for the other side to confirm the flag"; with only Simple, small messages would explode in latency due to the fixed overhead of "write FIFO + send notification + wait for notification." This is the root cause of why NCCL's performance curve has obvious inflection points around 8KB and 128KB.

The three sets of primitives share the same template skeletonPrimitives<T, RedOp, Fan, Direct, Proto, P2p, isNetOffload>, and throughProtothis template parameter, three versions are specialized📎 src/device/primitives.h:117-117。ProtoLL、ProtoLL128、ProtoSimpleThe three structs each carry protocol-related constants and computation methods📎 src/device/primitives.h:25-75, and the algorithm code only callsprims.send()、prims.recvReduceSend()unified interfaces like these, without caring which protocol underlies them.

mermaid
flowchart TD
    algo["算法层 all_reduce.h<br/>调用 prims.recvReduceSend()"] --> dispatch{"Proto 模板参数?"}
    dispatch -->|ProtoLL| ll["Primitives&lt;..., ProtoLL, ...&gt;<br/>prims_ll.h"]
    dispatch -->|ProtoLL128| ll128["Primitives&lt;..., ProtoLL128, ...&gt;<br/>prims_ll128.h"]
    dispatch -->|ProtoSimple| simple["Primitives&lt;..., ProtoSimple&lt;...&gt;, ...&gt;<br/>prims_simple.h"]
    ll --> llop["LLGenericOp&lt;RECV,SEND,SrcBuf,DstBuf&gt;"]
    ll128 --> ll128op["GenericOp -&gt; recvReduceSendCopy"]
    simple --> simpleop["genericOp -&gt; waitPeer / reduceCopy / postPeer"]

This diagram explains "why the same AllReduce logic needs three sets of data movement primitives": the algorithm layer is protocol-agnostic, and protocol differences are encapsulated inPrimitivesthe three specializations.

LL: zero-handshake data movement with flags embedded in data lines

Intuitive model

The core idea of LL is:pack "data" and the marker for "whether the data is ready" into the same 16-byte read/write unit. The receiver does not need an extra "notification message"; it only needs to poll the flag field in the data line, and a flag match means the data has arrived. This is like printing the "recipient signature" directly on the envelope when sending a letter—the mail carrier can tell at a glance whether it should be delivered, without sending a separate receipt.

Without this design, the receiver would have to first wait for a "data has been written" notification, then go back and read the data—two memory round trips, doubling latency.

Data structures and memory layout

LL's data movement unit isunion ncclLLFifoLine, and fromstoreLL's assembly we can see its layout📎 src/device/prims_ll.h:154-158:

code
st.volatile.global.v4.u32 [%0], {%1,%2,%3,%4};
// 写入 4 个 u32:data1, flag, data2, flag

OnencclLLFifoLineis 16 bytes, arranged as[data1(4B) | flag(4B) | data2(4B) | flag(4B)]. Only 8 bytes are effective data (data1 + data2); the other 8 bytes are all flag. This is whyProtoLL::calcBytePerGrain()returnssizeof(uint64_t)—"One 16-byte line has 8-bytes of data"📎 src/device/primitives.h:55-57。

Key fields (Primitives's LL specialization)📎 src/device/prims_ll.h:20-42:

FieldTypePurpose
recvStep[i] / sendStep[i]uint64_t[MaxRecv/MaxSend]Per-peer step counter, determining buffer offset and flag value
recvBuff[i] / sendBuff[i]ncclLLFifoLine*Points to each peer's FIFO buffer base address
recvConnHeadPtrvolatile uint64_t*Global pointer on the receive side for "how many steps have been consumed"
sendConnHeadPtrvolatile uint64_t*Global pointer on the send side for "how many steps the peer has consumed"
sendConnHeadCacheuint64_tCaches the last read head value to avoid reading global memory every time

The buffer offset is computed byrecvOffset(i) = (recvStep[i] % NCCL_STEPS) * stepLinesis the number of slots in the ring buffer,📎 src/device/prims_ll.h:44-46,NCCL_STEPSis the number of lines per slot. The flag value is computed bystepLines, noterecvFlag(i) = NCCL_LL_FLAG(recvStep[i] + 1)—because the initial flag value is 0, the first step's flag must be 1 to distinguish it from "not written."📎 src/device/prims_ll.h:56-58Scenario-driven Walkthrough: a recvReduceSend+1Suppose rank 0 executes

in Ring AllReduce: receive data from the previous rank, reduce it with local data, then send it to the next rank. The call chain is

Step 1: Wait for the send buffer to become available.recvReduceSendchecksrecvReduceSend(inpIx, eltN) → LLGenericOp<1, 1, Input, -1>(inpIx, -1, eltN, false) 📎 src/device/prims_ll.h:403-405。

. The meaning is: if the peer's consumption progress (head) lags too far behind me, the ring buffer is almost full, and I must wait. waitSendis the total number of buffer slots,sendConnHeadCache + NCCL_STEPS < sendConnHead + 1 📎 src/device/prims_ll.h:73-89is the slot I am about to occupy. While waiting, pollNCCL_STEPSto update the cache, and periodically callsendConnHead + 1to check whether it has been aborted*sendConnHeadPtrStep 2: Load local data.checkAborthandles the alignment problem📎 src/device/prims_ll.h:73-89。

. When DataLoader::loadBegin(such as half or int8), the source address may not be 4-byte aligned, so first read it into📎 src/device/prims_ll.h:200-216with 4-byte alignment, recordsizeof(T) <= 2, and then inu4[0..2]usemisalignto perform byte-level shifts and assemble the correct 64-bit valueloadFinish. This is a typical "aligned read + shift reassembly" technique, avoiding the performance penalty of unaligned access.__funnelshift_rStep 3: Read peer data and wait for the flag.📎 src/device/prims_ll.h:218-225is the core

Copy readLLIt uses📎 src/device/prims_ll.h:108-122:

cpp
do {
  asm volatile("ld.volatile.global.v4.u32 {%0,%1,%2,%3}, [%4];" ...);
  if (checkAbort(abort, 1, spins)) break;
} while ((flag1 != flag) || (flag2 != flag));

它用 ld.volatile.global.v4.u32Read 16 bytes at once (4 u32s), then check whether both flag fields equal the expected values.volatileThe keyword ensures the compiler will not optimize away this read or cache it in a register—because the peer may write new data at any time. Both flags must match because the writerstoreLLwrites 4 u32s at once, which in theory may be split into two 8-byte writes. Only when both flags match can the 16 bytes be guaranteed complete.

Step 4: reduce and send.After receiving peerData,applyReduce(redOp, peerData, data)perform the reduction📎 src/device/prims_ll.h:279. ThenstoreLL(sendPtr(i) + offset, data, sendFlag(i))write the result to the send buffer📎 src/device/prims_ll.h:295-296. Note the send order: first sendi=1..MaxSend(usually the network peer), and finally sendi=0(usually the local peer)📎 src/device/prims_ll.h:291-297. The comment is very clear: "Send : inter-node, then intra-node, then local"—send the slow one (network) first so it can fly in the background, then send the fast one (local), so the local peer does not wait for the network.

Step 5: advance step and post. incRecv(i)Increment the receive step📎 src/device/prims_ll.h:91-93,postRecv()and writerecvConnHeadback to the global pointer📎 src/device/prims_ll.h:94-97, notifying the peer "I have consumed this step." On the send side,incSendthere is special logic📎 src/device/prims_ll.h:99-106:

cpp
if ((sendStep[i] & NCCL_LL_CLEAN_MASK) == NCCL_LL_CLEAN_MASK) {
  for (int o = offset; o < stepLines; o += nthreads) storeLL(sendPtr(i) + o, 0, sendFlag(i));
}

When step reaches theNCCL_LL_CLEAN_MASKboundary, all rows of the entire slice must be written once with the current flag (data filled with 0). Why? Because flags are reused cyclically; if the previous flag of a row happens to equal the expected value this time, the receiver will mistakenly think the data is ready. This "cleanup" operation uniformly flushes the flags of all rows to the new value, eliminating ambiguity.

Concurrency control and hardware interaction

LL synchronization relies entirely onvolatilereads and writes + flag polling, with no locks.barrier()Use__syncwarp()(for a single warp) orbarrier_sync(15 - group, nthreads)(for multiple warps)📎 src/device/prims_ll.h:63-69。15 - groupis the barrier number. NCCL uses different barrier numbers to isolate different groups and avoid mutual interference.

checkAbortis the key to preventing infinite loops📎 src/device/primitives.h:154-164: only everyNCCL_SPINS_BEFORE_CHECK_ABORT(10000) spins isabortFlagread once, avoiding frequent global memory reads that slow down the hot path. Once abort is detected, setncclShmem.abortedand cache it; all subsequent wait loops will exit quickly.

Production pitfalls

Pitfall 1: false readiness caused by flag wraparound.IfNCCL_LL_CLEAN_MASK's cleanup logic is removed, after long runtime (step exceeds the mask period), the receiver may read a residual flag from the previous round, mistakenly judge the data as ready, and read stale data. This kind of bug is extremely hard to reproduce because it depends on step happening to wrap around to a specific value.

Pitfall 2:MaxRecv == 0compilation trap.In the code,MaxRecv = Fan::MaxRecv > 1 ? Fan::MaxRecv : 1 📎 src/device/prims_ll.h:13, because even if only sending and not receiving, a receive buffer of length MaxRecv will still be allocated. If MaxRecv is 0, it will cause a zero-length array compilation failure. On Windows,MaxSendhas the same handling📎 src/device/prims_ll.h:14-19。

LL128: trading 128-byte alignment for higher payload

Intuitive model

LL's pain point is that the payload is only 50% (8 of 16 bytes are flags). LL128's idea is:concentrate the flags into the last 8 bytes of every 128 bytes, with the first 120 bytes all data. This increases the payload from 50% to 93.75%. The cost is that 128-byte alignment must be guaranteed, otherwise "shared memory repacking" is required.

Data structures and memory layout

LL128's transfer unit isuint64_t(8 bytes), but organized into 128-byte "lines".NCCL_LL128_LINEELEMSis the number of 64-bit elements per line (16),NCCL_LL128_DATAELEMSis the number of data elements among them (15), and the last element holds the flag.

Key constants📎 src/device/prims_ll128.h:292-294:

cpp
static constexpr int WireWordPerSlice = WARP_SIZE * NCCL_LL128_SHMEM_ELEMS_PER_THREAD;
static constexpr int DataEltPerSlice =
  (WireWordPerSlice - WireWordPerSlice / NCCL_LL128_LINEELEMS) * (sizeof(uint64_t) / sizeof(T));

WireWordPerSliceis the number of 64-bit words transferred by one warp at a time,DataEltPerSliceis the number of valid data elements among them (minus one flag element per line).

LL128's flag mechanism differs from LL:only the 7th of every 8 threads (flagThread) is responsible for checking the flag 📎 src/device/prims_ll128.h:373。flagThread = ((tid % 8) == 7). Why? Because there is one flag per 128 bytes, and a warp has 32 threads. Every 8 threads handle 128 bytes (8 threads × 16 bytes = 128 bytes), so only 1 of every 8 threads needs to read the flag.

Scenario-driven walkthrough: one recvReduceSendCopy

Call chain:recvReduceSend(inpIx, eltN) → GenericOp<1, 1, Input, -1> → recvReduceSendCopy<NCCL_LL128_SHMEM_ELEMS_PER_THREAD, RECV, SEND, SrcBuf, DstBuf> 📎 src/device/prims_ll128.h:422-423, 296-333。

Step 1: load local data into registers. loadRegsBeginThere are two cases📎 src/device/prims_ll128.h:99-142:

  • 16-byte aligned: directlyload128to registers, with no shared memory staging. Note thatflagThreadonly loads half the data (g % 2 == 0), because its other half of registers must be reserved for the flag📎 src/device/prims_ll128.h:109-114。
  • unaligned: first load the aligned region into shared memoryncclScratchForWarp(warpInBlock),__syncwarp(), then read it back from shared memory into registers at the correct offset📎 src/device/prims_ll128.h:115-141。

Step 2: wait for and read peer data. recvReduceSendCopyThe wait loop in📎 src/device/prims_ll128.h:190-207:

cpp
do {
  needReload = false;
  for (int u = 0; u < ELEMS_PER_THREAD; u += 2) {
    load128(ptr + u * WARP_SIZE, vr[u], vr[u + 1]);
    needReload |= flagThread && (vr[u + 1] != flag);
  }
  needReload &= (0 == checkAbort(abort, 1, spins));
} while (__any_sync(WARP_MASK, needReload));

Key point: onlyflagThreadchecks the flag, then uses__any_syncfor warp-level voting—as long as one flagThread finds that the flag does not match, the entire warp continues spinning. This saves more instructions than having every thread check the flag.

Step 3: register rearrangement. loadRegsFinishMove the flagThread's flag register to an idle register📎 src/device/prims_ll128.h:145-151. The comment explains this design: "By deferring register shuffle here we've overlapped spinning on first peer's data with memory loads of src data" — deferring the register shuffle until after the wait allows the wait time to overlap with local data loading.

Step 4: reduce and send.After receiving data, doapplyReduce 📎 src/device/prims_ll128.h:227-230, thenstore128write to the send buffer📎 src/device/prims_ll128.h:274-287. Note that when sending,flagThread ? flag : v[u+1]—flagThread writes the flag, while other threads write data.

Step 5: advance step.Unlike LL, LL128's step advancement is done uniformly at the end ofGenericOp📎 src/device/prims_ll128.h:324-332, rather than inrecvReduceSendCopy. Moreover,postSenduses__threadfence_system()(SM90+) or__threadfence() 📎 src/device/prims_ll128.h:87-96, ensuring data is visible to other GPUs/NICs before updating the tail pointer.

Concurrency Control and Hardware Interaction

LL128'sbarrier()always usesbarrier_sync(15 - group, nthreads) 📎 src/device/prims_ll128.h:64-66, unlike LL which has a single-warp optimization. This is because LL128's data movement is warp-level and requires cross-warp synchronization.

loadRegsBeginThe shared memory repacking in__syncwarp()uses📎 src/device/prims_ll128.h:129to synchronize

, ensuring all threads finish writing to shared memory before reading.

Production PitfallsPitfall 1: The performance cliff of unaligned access.

If the user buffer is not 16-byte aligned, every transfer must go through shared memory as an intermediary, and performance may drop by more than 30%. In production environments, ensure input and output buffers are allocated with 16-byte alignment.flagThreadPitfall 2:register pressure.

flagThread only loads half the data, meaning its register utilization differs from other threads. If the compiler does not allocate registers correctly, it may cause register spilling to local memory and a sharp performance drop.

Simple: Achieving high throughput for large messages with FIFO + notification

Intuitive Model

The Simple protocol is like a parcel locker: the sender puts data into a FIFO buffer (a locker), then updates a step pointer indicating "slot N has been filled" (sends a notification); the receiver polls the step pointer, and upon seeing a new value, goes to the corresponding locker to pick up the goods. The handshake overhead is high (must write pointer + read pointer), but it can move a lot of data at once, making it suitable for large messages.

Data Structures and Memory Layout📎 src/device/prims_simple.h:28-46:

Simple's fields are much more complex than LL/LL128FieldType
flagsintPurpose
stepuint64_tBit flags, encoding role (WaitRecv/WaitSend/PostRecv/PostSend), Direct mode, NetReg, etc.
connStepPtruint64_t*Current step
connStepCacheuint64_tPoints to the peer's step pointer for the connection
connEltsFifoT*Caches the last read step value
connStepSizeintFIFO buffer base address
directBuffT*Bytes per step

flagsDirect buffer pointer in Direct mode📎 src/device/prims_simple.h:23-27:

cpp
RoleInput = 0x01, RoleOutput = 0x02, RoleWaitRecv = 0x04, RoleWaitSend = 0x08,
RolePostSend = 0x10, RolePostRecv = 0x20, Aborted = 0x40, NetRegMode = 0x80,
ConnFifoEnabled = 0x100, DirectWrite = 0x200, DirectRead = 0x400, PatMode = 0x800,
NvlsMinPolling = 0x1000, NetDeviceUnpack = 0x2000, AnyNetDeviceUnpack = 0x4000,
RoleWaitPatNvls = 0x8000, RolePostPatNvls = 0x10000;

CopytidThis is a typical design of "using bit operations instead of multiple bool fields" to save registers. Each thread is assigned a role based on its📎 src/device/prims_simple.h:651-666:nrecvthe firstnsendthreads are WaitRecv, the nextnrecvare WaitSend, the lastnsendare PostRecv, and the last

are PostSend.

Scenario-Driven Walkthrough: A recvReduceSendrecvReduceSend(inpIx, eltN) → genericOp<0, 0, 1, 1, Input, -1> 📎 src/device/prims_simple.h:994-996。

Call chain: sliceSize = max(divUp(nelem, 16 * SlicePerChunk) * 16, sliceSize / 32) 📎 src/device/prims_simple.h:185-186Step 1: Compute slice size.

. This formula ensures the slice is at least 16-byte aligned and not too small.Step 2: Worker loop.tid < nworkersOnly📎 src/device/prims_simple.h:190。nworkers = nthreads - (MaxSend > 0 && nthreads >= NCCL_SIMPLE_EXTRA_GROUP_IF_NTHREADS_GE ? WARP_SIZE : 0) 📎 src/device/prims_simple.h:626threads enter the main loop

—reserving one warp to overlap threadfence and copy. waitPeerStep 3: Wait for the peer.📎 src/device/prims_simple.h:103-164:

cpp
while (connStepCache + (isSendNotRecv ? NCCL_STEPS : 0) < step + StepPerSlice) {
  connStepCache = loadStepValue(connStepPtr);
  if (checkAbort(flags, Aborted, spins)) break;
}

isSendNotRecvCopyNCCL_STEPSDistinguishes send and receive: when sending, it waits for "the peer has consumed" (head); when receiving, it waits for "the peer has produced" (tail).StepPerSliceis the number of buffer slots,

is the number of steps per slice.ptrs[index] 📎 src/device/prims_simple.h:123-158After the wait completes, set

according to Direct mode. Direct mode allows directly reading and writing the peer's buffer, bypassing the FIFO and reducing one copy.Step 4: reduceCopy.reduceCopySelect different📎 src/device/prims_simple.h:241-277calls based on the Direct combinationsrcs[0] && dsts[0]. The most complex branch is when📎 src/device/prims_simple.h:258-271both existreduceCopy<Unroll, RedOp, T, MultimemSrcs, Recv+Src, Recv*MaxRecv+Src, MultimemDsts, Send+Dst, Send*MaxSend+Dst, PreOpSrcs>, callingRecv*MaxRecv+Src, where the parameters mean: read fromSend*MaxSend+Dstsources, reduce, and write to

destinations. postPeerStep 5: postPeer.📎 src/device/prims_simple.h:167-175:

cpp
if (Send && (flags & RolePostSend) && (dataStored || (flags & ConnFifoEnabled))) {
  fence_acq_rel_sys();
}
st_relaxed_sys_global(connStepPtr, step);

Copyfence_acq_rel_sys()The send side must

before updating step, ensuring data writes are visible to other GPUs/NICs. The receive side does not need a fence, because the receiver is only notifying "I have consumed" and does not involve data visibility.

Concurrency Control and Hardware Interactionst_relaxed_sys_globalSimple's synchronization uses📎 src/device/prims_simple.h:167-175to write the step pointerloadStepValue, and uses📎 src/device/prims_simple.h:86-100。loadStepValueto readNvlsMinPollingOn SM90+ withmultimem.ld_reduce.acquire.sys.global.min.u64enabled, it uses the📎 src/device/prims_simple.h:86-100instruction

barrier(), which is NVLink SHARP's hardware-accelerated polling.subBarrier()and📎 src/device/prims_simple.h:49-55:barrier()differencenthreadssynchronizes allsubBarrier()threads,nworkersonly synchronizessubBarrierworker threads.15 - group - (nworkers != nthreads ? 1 : 0)'s barrier number isbarrier(), and when the number of workers is not equal to the total number of threads, a different barrier is used to avoid conflict with

.

Production PitfallsPitfall 1: Destructor wait under NetRegMode.📎 src/device/prims_simple.h:794-804:

cpp
if ((flags & NetRegMode) && (flags & RoleWaitSend)) {
  uint64_t prevStep = step - StepPerSlice;
  volatile ssize_t* ptr = &(connFifo[prevStep % NCCL_STEPS].size);
  while (*ptr != -1) { ... }
}

In NetRegMode, the send buffer is directly accessed by the NIC, and it must wait for the proxy thread to confirm that the send is complete (size is set to -1) before returning; otherwise, the next kernel may overwrite data that is being read by the NIC.

Pitfall 2: DirectRead sendrecv deadlock.There is also a section in the destructor📎 src/device/prims_simple.h:814-824:

cpp
if ((flags & DirectRead) && (flags & RoleWaitSend) && P2p) {
  while (*tail > *head) { ... }
}

In DirectRead mode of sendrecv, the sender must wait for the receiver to finish reading the data before returning. If the receiver does not advance tail for some reason, the sender will deadlock. This wait must be done afterbarrier()otherwise it may race with the post thread.

Pitfall 3:roundUpThe resulting step jump. loadRecvConnandloadSendConnboth containstep = roundUp(step, SlicePerChunk * StepPerSlice) 📎 src/device/prims_simple.h:486, 533. This aligns step to the slice boundary, but if the previous step is not aligned, it causes skipped slots to not be initialized correctly. The code adds a line inloadRecvConnto return credit*connStepPtr = stepComparison and selection of the three primitive sets📎 src/device/prims_simple.h:489。

Copy

mermaid
flowchart LR
    subgraph LL["LL 协议"]
        ll_data["ncclLLFifoLine 16B<br/>data1(4B)+flag(4B)+data2(4B)+flag(4B)"]
        ll_sync["flag 内嵌数据行<br/>轮询 flag 匹配"]
    end
    subgraph LL128["LL128 协议"]
        ll128_data["128B line<br/>15×8B data + 1×8B flag"]
        ll128_sync["flagThread 每8线程1个<br/>__any_sync 投票"]
    end
    subgraph Simple["Simple 协议"]
        simple_data["FIFO 缓冲区<br/>connEltsFifo + step*connStepSize"]
        simple_sync["step 指针 + fence<br/>loadStepValue 轮询"]
    end
    ll_data --> ll_sync
    ll128_data --> ll128_sync
    simple_data --> simple_sync
Effective payload rateLLLL128Simple
Synchronization method50%93.75%~100%
flag embedded, pollingflagThread + warp votestep pointer + fenceAlignment requirement
None (with shift and reassembly)16 bytesNoneApplicable message size
Small (< 8KB)Medium (8KB ~ 128KB)Large (> 128KB)Buffer layout
By 128B linencclLLFifoLine[]uint64_t[]Direct supportT[] FIFO
None (downgrade)PrimitivesWithoutDirectNone (same as left)Fully supportedBoth LL and LL128 inherit

because their buffer layouts do not support directly reading and writing peer memory. Simple fully implements Direct mode, supporting P2P direct connections and NVLS.PrimitivesWithoutDirect 📎 src/device/prims_ll.h:9-10, src/device/prims_ll128.h:13-14Design considerations

[Design inference and architectural tradeoffs]

Why does the LL flag need to be duplicated twice?

Because GPU global memory writes are not guaranteed to be atomic.For a 16-byte write, the hardware may split it into two 8-byte writes. If only one flag is placed, the receiver may consider the data ready when only half of it has been written. The two flags are located in the first half and second half of the 16 bytes respectively, so only when both writes are complete will both flags match.storeLLWhy does Simple reserve one warp?

The comment says, "For send operations, we need an extra warp to overlap the threadfence and the copy." 📎 src/device/prims_simple.h:625-626It is an expensive operation. If all threads wait for the fence to complete before continuing, a large amount of time will be wasted. Reserving one warp specifically for the fence allows the other warps to continue moving the next batch of data.fence_acq_rel_sys()[Design inference and architectural tradeoffs]

Why is LL128's step advancement at the end of GenericOp instead of in recvReduceSendCopy?

Because LL128's transfer is warp-level, and multiple warps may process different slices in parallel. If step is advanced ineach warp will advance it once, causing step to be advanced multiple times. Placing the unified advancement at the end ofrecvReduceSendCopyensures that each slice advances only once.GenericOpChapter summary

This chapter took a deep dive into the implementation of the three transfer primitive sets:

: uses 16-byte

1. LLto embed the flag in the data row, and the receiver only needs to poll for a flag match to confirm that the data is ready. Payload is 50%, suitable for small messages. The core isncclLLFifoLine'sreadLLandld.volatile.global.v4.u32'sstoreLL: concentrates the flag into the last 8 bytes of every 128 bytes, increasing the payload to 93.75%. Usesst.volatile.global.v4.u32。

2. LL128(1 per 8 threads) to check the flag,flagThreadand performs a warp vote. When unaligned, it goes through shared memory repacking.__any_sync: uses a FIFO buffer + step pointer notification to achieve high throughput for large messages.

3. Simplebit flags encode the role,flagspolls step,waitPeerupdates step and fences. Fully supports Direct mode.postPeerThe three primitive sets share the same template skeleton, specialized through

template parameters. The algorithm layer only calls the unified interface and does not care about the underlying protocol. This is the answer to "why the same AllReduce logic needs three transfer primitive sets": different message sizes require different synchronization strategies and buffer layouts, and the three primitive sets are optimized for small, medium, and large messages respectively.ProtoChapter review questions

Q1: If the cleanup logic in

(incSend) is removed, in what scenario will data corruption be triggered? Why?📎 src/device/prims_ll.h:99-106Reference analysis

: The cleanup logic writes all rows of the entire slice once with the current flag (data filled with 0) when. If removed, when step wraps around to thesendStep[i] & NCCL_LL_CLEAN_MASK == NCCL_LL_CLEAN_MASKboundary, the flags of some rows may still be the values from the previous round. If the previous round's flag happens to equal the flag expected by the receiver in this round, the receiver will mistakenly believe the data is ready and read the residual data from the previous round. This is a typical ABA problem. The trigger condition is long-running operation (step exceedsNCCL_LL_CLEAN_MASKcycles) and the flag happens to wrap around to the same value. This kind of bug is extremely difficult to reproduce because it requires precise step alignment.NCCL_LL_CLEAN_MASK 周期)且 flag 恰好回绕到相同值。这类 bug 极难复现,因为需要精确的 step 对齐。

Q2: In the destructor of the Simple protocol, what do the wait under NetRegMode (📎 src/device/prims_simple.h:794-804) and the wait under DirectRead (📎 src/device/prims_simple.h:814-824) each prevent? If one of them is removed, what happens in a high-concurrency scenario?

Reference analysis: NetRegMode waits for the proxy thread to setconnFifo[prevStep].sizeto -1, indicating that the NIC has finished sending. If this is removed, the next kernel may overwrite the send buffer that is currently being read by the NIC via DMA, causing the NIC to read dirty data. DirectRead waits for the receiver to advance the tail (*tail > *head), indicating that the receiver has finished reading the direct buffer. If this is removed, the sender may overwrite the buffer before the receiver has finished reading, causing the receiver to read new data instead of old data. In high-concurrency scenarios, both waits are necessary; removing either one will cause a data race. The difference is that NetRegMode prevents "NIC reads," while DirectRead prevents "peer GPU reads."

Q3: LL128'sloadRegsBegintakes the shared memory repacking path (📎 src/device/prims_ll128.h:115-141) when unaligned. How much slower is this path than the aligned path? Why doesn't NCCL directly require user buffers to be 16-byte aligned?

Reference analysis: The unaligned path has three extra steps: write to shared memory,__syncwarp(), read from shared memory. Although shared memory bandwidth is high,__syncwarp()is a synchronization point that blocks the warp until all threads finish writing. A rough estimate is that the unaligned path is 20-40% slower than the aligned path, depending on shared memory bank conflicts. NCCL does not enforce alignment because users may pass buffers with arbitrary offsets (such as tensor slices), and enforcing alignment would limit API flexibility. NCCL's strategy is "use the fast path when aligned, and use the slow path when unaligned while guaranteeing correctness." In production environments, users are advised to allocate buffers aligned to 16 bytes as much as possible in order to use the fast path.

At this point, we have mastered the data movement mechanisms of the three primitives LL, LL128, and Simple, which provide flexible performance tuning means for upper-layer algorithms. The next chapter will go deep into the collective communication algorithm kernels to see how AllReduce, AllGather, ReduceScatter, etc. call these primitives, and how algorithms such as Ring, Tree, and CollNet organize data flow, ultimately completing end-to-end collective communication.

CHAPTER 10

Chapter 10: Chapter 10: Collective Communication Algorithm Kernels: Device-Side Implementation of AllReduce, AllGather, ReduceScatter

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 10 / 25

Chapter 10: Collective Communication Algorithm Kernels: Device-Side Implementation of AllReduce, AllGather, ReduceScatter

The previous chapter broke down the three protocol primitives LL, LL128, and Simple. They are the "engines" of data movement, but the engines themselves do not know what to move, where to move it, or in what order. The set of algorithm kernel files under src/device that this chapter examines is the "gearbox" - they translate collective communication semantics such as AllReduce, AllGather, and ReduceScatter into a series of primitive calls such as prims.directSend and prims.directRecvReduceDirectSend. To summarize the core contradiction of this chapter in one sentence: for the same AllReduce, why are four completely different device-side implementations needed: Ring, Tree, CollNet, and NVLS? The answer lies in matching "data flow topology" with "hardware capabilities." Ring uses the least network bandwidth to perform two-stage pipelining, Tree uses tree-shaped reduction to reduce latency to log(n), and CollNet/NVLS offload reduction to the NIC or NVLink switch. This chapter examines them one by one.

10.1 Ring AllReduce: How two-stage pipelining is implemented inside the kernel

Intuitive model: a "relay race" on a ring-shaped pipeline

Imagine n workers standing in a circle, each holding a box of raw materials. The goal of AllReduce is for everyone to ultimately receive the "finished product after all raw materials are mixed." The Ring algorithm works in two stages: in the first stage (reduce-scatter), each person passes the box along the ring, and at each stop mixes in their own raw materials. After n-1 stops, each person has exactly one "fully mixed" finished product, but only a 1/n share; in the second stage (all-gather), these finished-product shares are passed around the ring again, and each person completes all shares.

Without Ring, the most naive approach is for each rank to send data to the root, and after the root reduces it, broadcast it - the root's network bandwidth becomes the bottleneck, and the larger n is, the slower it gets. The brilliance of Ring is:Each rank's send and receive volume is 2(n-1)/n times the data volume, flattened across all links regardless of n。

Data Structures and Memory Layout

The core state of the Ring algorithm is in thencclRingstructure (defined in device.h, not covered in this chapter),runRingonly two fields are taken from it:

  • ring->index: this rank's logical position in the ring, used to compute "which chunk to process at step j".
  • ring->prev / ring->next: the predecessor and successor rank numbers, used as thePrimitivesconstructor's recv/send peer parameters.

The key chunking parameters are computed byncclCollCbdPart(📎 src/device/all_reduce.h:21-22):

code
ncclCollCbdPart(work, ncclShmem.channelId, Proto::Id, sizeof(T), (ssize_t*)nullptr, &gridOffset, &channelCount, &chunkCount);

This function splits the entire communication domain's data by channel and outputs three values:gridOffset(the starting offset of the data this channel is responsible for within the entire buffer),channelCount(the total number of elements this channel is responsible for),chunkCount(the number of chunk elements each rank gets).chunkCountis the granularity of the Ring algorithm—one chunk is moved per step.

loopCount = nranks * chunkCount(📎 src/device/all_reduce.h:23) represents the amount of data processed in "one full lap". The outer loopfor (elemOffset = 0; elemOffset < channelCount; elemOffset += loopCount)(📎 src/device/all_reduce.h:34) means: if the channel's data volume exceeds what one lap can process, it runs in multiple laps.

Step-by-Step Walkthrough: The Complete Call Flow of a Ring AllReduce

Scenario: 4 ranks (nranks=4), this rank'sringIx=0,chunkCount=100,channelCount=400(exactly one lap).

Step 0: Push "your own chunk" to the next GPU(📎 src/device/all_reduce.h:42-47)

code
chunk = modRanks(ringIx + nranks - 1);   // = 3
chunkOffset = chunk * chunkCount;         // = 300
offset = gridOffset + elemOffset + chunkOffset;
nelem = min(chunkCount, remCount - chunkOffset);
prims.directSend(offset, offset, nelem);

modRanksis a lambda that performs modulo-nranks subtraction (📎 src/device/all_reduce.h:40)。ringIx + nranks - 1represents "this rank's previous chunk number". Why is chunk 3 sent at step 0? Because in the Ring's reduce-scatter phase, each rank first sends out the portion of data it "should not keep" (i.e., the predecessor rank's chunk).directSendonly sends without receiving, because no data has been received yet at this point.

Steps 1 to nranks-2: Receive, reduce, and forward simultaneously(📎 src/device/all_reduce.h:50-56)

code
for (int j = 2; j < nranks; ++j) {
  chunk = modRanks(ringIx + nranks - j);
  ...
  prims.directRecvReduceDirectSend(offset, offset, nelem);
}

directRecvReduceDirectSendis the core primitive of Ring: receive a chunk fromprev, reduce it with local data (e.g., addition), then send the result tonext. Note thatoffsetandnelemare recomputed on each iteration—because the chunk processed differs each step. j goes from 2 to nranks-1, for a total of nranks-2 steps.

Step nranks-1: Receive the last chunk and reduce, producing the final result(📎 src/device/all_reduce.h:58-64)

code
chunk = ringIx + 0;
...
prims.directRecvReduceCopyDirectSend(offset, offset, nelem, /*postOp=*/true);

ThepostOp=trueof this step is key: after the reduction completes, a post-operation must be performed (e.g., division when computing the average).directRecvReduceCopyDirectSendhas one moreCopythan the previous step—writing the reduction result simultaneously to the local recvbuff and to next. At this point the reduce-scatter phase ends, and each rank holds a "fully reduced" chunk.

all-gather phase: nranks-2 steps of pure forwarding(📎 src/device/all_reduce.h:66-73)

code
for (int j = 1; j < nranks - 1; ++j) {
  chunk = modRanks(ringIx + nranks - j);
  ...
  prims.directRecvCopyDirectSend(offset, offset, nelem);
}

Note thatdirectRecvCopyDirectSendis used here, withoutReduce—because the data has already been reduced, only copy-forwarding is needed.

Final step: Receive the last chunk(📎 src/device/all_reduce.h:75-81)

code
chunk = modRanks(ringIx + 1);
...
prims.directRecv(offset, nelem);

Only receives without sending, completing the last piece.

The entire flow can be summarized by the following control flow diagram:

mermaid
flowchart TD
    start["runRing 入口<br/>计算 chunkCount/loopCount"] --> loop{"elemOffset < channelCount?"}
    loop -->|否| done["返回"]
    loop -->|是| s0["step 0: directSend<br/>chunk = ringIx-1"]
    s0 --> mid{"j 从 2 到 nranks-1?"}
    mid -->|是| s1["directRecvReduceDirectSend<br/>chunk = ringIx-j"]
    s1 --> mid
    mid -->|否| s2["step nranks-1<br/>directRecvReduceCopyDirectSend<br/>postOp=true"]
    s2 --> ag{"j 从 1 到 nranks-2?"}
    ag -->|是| s3["directRecvCopyDirectSend<br/>纯转发"]
    s3 --> ag
    ag -->|否| s4["directRecv<br/>收最后一块"]
    s4 --> loop

Design Thinking: Why Ring's chunk order goes "backwards"

Note the pattern of chunk numbering: step 0 sendsringIx-1, step j processesringIx-j, the final step processesringIx+0. This iscounterclockwiseprogression. Why? Because each rank in Ring only keeps "the chunk it is responsible for reducing" (i.e.,ringIx+0), and all other chunks just pass through. Counterclockwise progression guarantees: when a chunk completes a full lap back to its starting point, it has exactly completed nranks reductions, producing the final result. If it progressed clockwise, the chunk would complete its reduction on the wrong rank.

Production Pitfall:remCount < loopCountalignment trap when

📎 src/device/all_reduce.h:38There is a line of code that is easy to overlook:

code
if (remCount < loopCount) chunkCount = alignUp(divUp(remCount, nranks), 16 / sizeof(T));

When the remaining data is less than one lap, chunkCount must be recomputed, andalignUp(..., 16/sizeof(T))is forced to 16-byte alignment. Why? Because the LL128 protocol requires 128-byte alignment, and the Simple protocol also has alignment requirements for vectorized access. If this alignment is removed, unaligned chunks take the slow path, degrading performance by 20-40%. In production, if you find Ring AllReduce performance jitter at the tail of small messages, it is often because this alignment is not in effect—check whetherchannelCountis an integer multiple ofnranks * 16/sizeof(T).

10.2 Tree AllReduce: Compressing Latency to log(n) with Tree Reduction

Intuitive Model: "Level-by-Level Reporting" in a Company

Ring's latency is O(n)—data must go around a full lap. When n is very large (e.g., 1024 GPUs), even if bandwidth is flattened, the latency becomes unbearable. The Tree algorithm takes a different approach: like a company's organizational structure, each rank only communicates with its "parent node" and "child nodes". In the reduction phase, leaf nodes report data upward, and parent nodes merge their child nodes' data; in the broadcast phase it reverses, with the root node sending the result downward. Latency drops from O(n) to O(log n).

Without Tree, AllReduce latency in large-scale clusters grows linearly with the number of ranks, and training iteration time is dragged down by communication.

Data Structures and Memory Layout

Tree's state is inncclTree:

  • tree->up: parent node rank (-1 means this rank is the root).
  • tree->down[]: child node array, at mostNCCL_MAX_TREE_ARITY(typically 3, i.e., binary + local).

runTreeUpDownandrunTreeSplitare two variants. The former uses a two-phase mode of "reduce all first, then broadcast all," while the latter splits threads into two halves, with one half doing reduction and the other half doing broadcast, achieving pipeline overlap.

Step-by-Step Walkthrough: The Three Branches of runTreeUpDown

runTreeUpDownThe first code block is the reduction phase (📎 src/device/all_reduce.h:96-118), which has three cases depending on this rank's position in the tree:

Case A: This rank is the root (tree->up == -1)(📎 src/device/all_reduce.h:99-104)

code
prims.directRecvReduceCopy(offset, offset, nelem, /*postOp=*/true);

The root node only receives and does not send; it receives data from all child nodes, reduces, and writes to recvbuff.postOp=trueExecute the post-operation.

Case B: This rank is a leaf (tree->down[0] == -1)(📎 src/device/all_reduce.h:105-110)

code
prims.directSend(offset, offset, nelem);

A leaf node only sends and does not receive; it sends its own data to the parent node.

Case C: Intermediate node(📎 src/device/all_reduce.h:111-117)

code
prims.directRecvReduceDirectSend(offset, offset, nelem);

Receive from child nodes, reduce, and send to the parent node.

Broadcast phase (📎 src/device/all_reduce.h:120-142) logic is symmetric: root nodedirectSendFromOutput(sends from recvbuff), leaf nodedirectRecv, intermediate nodedirectRecvCopyDirectSend。

runTreeSplit: Using Thread Splitting to Implement Reduce-Broadcast Pipeline

runTreeUpDownThe problem withrunTreeSplitis that the reduction phase and broadcast phase are serial, with a global synchronization point in between.📎 src/device/all_reduce.h:155-164):

code
if (Proto::Id == NCCL_PROTO_SIMPLE) {
  nthreadsSplit = nthreads / 2;
  if (nthreadsSplit >= 256) nthreadsSplit += 64;
} else {
  nthreadsSplit = (nthreads * 7 / (10 * WARP_SIZE)) * WARP_SIZE;
}

Copy

The Simple protocol splits them evenly; the LL/LL128 protocols split them 7:3, because "receiving data from 3 sources for reduction" is more compute-intensive than "sending to 3 targets," so the reduction group gets more threads.tid < nthreadsSplitThen📎 src/device/all_reduce.h:175-202threads perform reduction push-up (📎 src/device/all_reduce.h:203-224), and the remaining threads perform broadcast push-down (Proto::MaxGroupWidth). The two groups distinguish their respective communication groups via the📎 src/device/all_reduce.h:189offset (0 * Proto::MaxGroupWidth's📎 src/device/all_reduce.h:210and1 * Proto::MaxGroupWidth)。

's

Design Thinking: Why Tree's Root Node Needs Special HandlingdirectRecvReduceDirectSendThe root node of tree reduction is the "convergence point"; its receive volume is a multiple of the number of child nodes, and its send volume is zero (during the reduction phase). If the root node also went through the generictree->up, it would try to send toif (tree->up == -1)(-1), causing an out-of-bounds error. Therefore, it must be handled separately with thetree->down[0] == -1branch. Similarly, the leaf node's

check.

Production Pitfall: The "Hot Root" Problem of the Tree AlgorithmTree's root node bears all reduction traffic. If the GPU where the root node resides happens to be a slow node (e.g., limited PCIe bandwidth), the entire AllReduce will be slowed down. NCCL's response is:Each channel selects a different rootrunTreeSplit, distributing the root node's load across multiple ranks. This is whyFanSymmetric<NCCL_MAX_TREE_ARITY_TOP>(📎 src/device/all_reduce.h:168the root node branch uses

)—it must handle reduction from multiple child nodes simultaneously. In production, if Tree AllReduce performance is uneven, check whether the channel's root node distribution is balanced.

10.3 AllGather and ReduceScatter: Ring's "Half-Journey" Variants

Intuitive Model: AllReduce Split into Two Halves

AllGather and ReduceScatter are essentially the two phases of AllReduce made into independent APIs. AllGather only does "gather"—each rank contributes a piece of data, and ultimately everyone gets all the data. ReduceScatter only does "reduce + scatter"—everyone contributes data, and after reduction each person gets one piece.

Without these two independent APIs, when users want to do "reduce first, then gather" or "gather first, then reduce," they can only call AllReduce and then manually slice, wasting half the bandwidth.

all_gather.hAllGather's Ring ImplementationrunRing(📎 src/device/all_gather.h:14-88's

) is simpler than AllReduce: no reduction, only copy-and-forward.(📎 src/device/all_gather.h:51-60)

code
rankDest = ringRanks[0];
offset = dataOffset + rankDest * count;
if ((inputBuf + dataOffset == outputBuf + offset) || isNetOffload) {
  prims.directSend(dataOffset, offset, nelem);
} else {
  prims.directCopySend(dataOffset, offset, nelem);
}

CopyinputBuf + dataOffset == outputBuf + offsetThere is an in-place check here: ifdirectSend, it means input and output are the same memory block (in-place AllGather), directlydirectCopySend; otherwise

(copy to output first, then send).(📎 src/device/all_gather.h:62-67)

code
prims.directRecvCopyDirectSend(offset, offset, nelem);

Copy(📎 src/device/all_gather.h:69-74)

code
prims.directRecv(offset, nelem);

Copy

📎 src/device/all_gather.h:28-36isNetOffload: Single Warp Drives Network + Multiple Warps Copy in Parallel

code
if (isNetOffload) {
  workNthreads = WARP_SIZE;
  chunkCount = NCCL_MAX_NET_SIZE;
} else {
  workNthreads = nthreads;
}

CopyisNetOffload=trueWhen📎 src/device/all_gather.h:76-82(single RPN + network registration mode), only 1 warp drives Ring communication, and the remaining warps perform "copy source data to target buffer" in parallel (

). This is to overlap copy overhead with communication overhead during non-in-place AllGather.barrier_sync(14, nthreads)(📎 src/device/all_gather.h:87Finally there is a__syncthreads()。

), and the comment explains it clearly: must wait for all warps to complete, otherwise the next work may reuse outputBuf and cause a race. Barrier 14 is used to avoid prims' own barrier and

reduce_scatter.hReduceScatter's Ring ImplementationrunRing(📎 src/device/reduce_scatter.h:14-56's

) is the reduce-scatter phase of AllReduce extracted separately:(📎 src/device/reduce_scatter.h:39-42)

code
rankDest = ringRanks[nranks - 1];
offset = dataOffset + rankDest * count;
prims.send(offset, nelem);

Copy(📎 src/device/reduce_scatter.h:44-49)

code
prims.recvReduceSend(offset, nelem);

Copy(📎 src/device/reduce_scatter.h:61-64)

code
prims.recvReduceCopy(offset, dataOffset, nelem, /*postOp=*/true);

Note the last step'srecvReduceCopyhas two offsets:offset(receive source) anddataOffset(local input), the reduction result is written todataOffset。

Data flow comparison diagram

mermaid
flowchart LR
    subgraph AllReduce["AllReduce (两阶段)"]
        A1["reduce-scatter<br/>n-1 步"] --> A2["all-gather<br/>n-1 步"]
    end
    subgraph AG["AllGather (单阶段)"]
        B1["directSend<br/>step 0"] --> B2["directRecvCopyDirectSend<br/>n-2 步"] --> B3["directRecv<br/>step n-1"]
    end
    subgraph RS["ReduceScatter (单阶段)"]
        C1["send<br/>step 0"] --> C2["recvReduceSend<br/>n-2 步"] --> C3["recvReduceCopy<br/>step n-1"]
    end
    AllReduce -.->|"拆解"| AG
    AllReduce -.->|"拆解"| RS

Production pitfall: the boundary of in-place detection

📎 src/device/all_gather.h:55's in-place detectioninputBuf + dataOffset == outputBuf + offsetrelies on exact pointer equality. If the sendbuff and recvbuff passed by the user have an offset but are logically the same memory block, this check will fail, causing it to take thedirectCopySendpath—correct but with an extra copy. In production, when doing in-place AllGather, ensure sendbuff and recvbuff are exactly identical.

10.4 CollNet and NVLS: Offloading reduction to hardware

Intuitive model: let the "switch" help with the computation

Ring and Tree both have "the GPU compute the reduction itself." CollNet and NVLS take a different approach: offload the reduction operation to the NIC (CollNet) or the NVLink switch (NVLS). The GPU is only responsible for sending data out, and the hardware completes the reduction and then broadcasts it back. This is like changing from "each worker mixing the ingredients themselves" to "sending the ingredients to a central blender, which mixes them and then distributes them."

Without hardware offload, the reduction operation occupies the GPU's SM resources, and the reduction latency cannot be hidden.

Thread division of labor in CollNet Direct

RunWorkColl<ncclFuncAllReduce, ..., NCCL_ALGO_COLLNET_DIRECT, ...>'srun(📎 src/device/all_reduce.h:249-386) divides threads into four groups:

code
const int nThreadsScatter = WARP_SIZE + ((hasUp && hasDn) ? COLLNET_COPY_THREADS : ...);
const int nThreadsGather = ((hasUp && hasDn) ? COLLNET_COPY_THREADS : ...);
const int nThreadsBcast = WARP_SIZE + ((hasUp && hasDn) ? COLLNET_COPY_THREADS : ...);
const int nThreadsReduce = work->nWarps * WARP_SIZE - nThreadsScatter - nThreadsGather - nThreadsBcast;

The four thread groups are respectively responsible for: Scatter (scattering data to each rail), Reduce (sending to the network after reduction), Gather (collecting from each rail), Bcast (broadcasting after receiving from the network).COLLNET_COPY_THREADS = 96(📎 src/device/all_reduce.h:250) is the fixed number of copy threads.

netRegUsed: buffer layout in network registration mode

📎 src/device/all_reduce.h:280-288has a key branch:

code
if (work->netRegUsed) {
  offsetBase = bid * chunkSize;
  maxNelems = size;
  peerOffset = nChannels * chunkSize;
} else {
  offsetBase = bid * direct->nHeads * chunkSize;
  maxNelems = direct->nHeads * chunkSize;
  peerOffset = chunkSize;
}

netRegUsedIn mode, buffers are arranged contiguously by channel (bid * chunkSize), and the peer offset isnChannels * chunkSize; in non-registered mode, they are arranged by head (bid * nHeads * chunkSize), and the peer offset ischunkSize. This difference stems from the fact that network registration mode requires buffers to be contiguous so that the NIC can perform DMA.

NVLS warp allocation

RunWorkColl<ncclFuncAllReduce, ..., NCCL_ALGO_NVLS, ...>'srun(📎 src/device/all_reduce.h:391-523) uses finer warp allocation:

code
const int bcastWarps = hasOut ? (work->regUsed ? ((totalWarps - 2) >> 1) - 1 : 2) : 0;
const int reduceWarps = work->regUsed ? (totalWarps - bcastWarps - 2) : (hasOut ? 3 : nranks <= 6 ? 7 : 5);
const int scatterWarps = work->regUsed ? 1 : (totalWarps - reduceWarps - bcastWarps + 1) >> 1;
const int gatherWarps = work->regUsed ? 1 : (totalWarps - reduceWarps - bcastWarps) >> 1;

regUsedIn mode, scatter/gather each occupy only 1 warp (because NVLS hardware directly operates on registered memory), and reduce takes the majority; in non-registered mode, scatter/gather each take about half, and reduce is adjusted according to the number of ranks (≤6 uses 7 warps, otherwise 5 warps).

Timing interaction diagram

mermaid
sequenceDiagram
    participant App as 应用层
    participant Scatter as Scatter Warps
    participant NVLS as NVLS 硬件
    participant Reduce as Reduce Warps
    participant Bcast as Bcast Warps

    App->>Scatter: prims.scatter(offset, nelem, chunkSize)
    Scatter->>NVLS: 写入 NVLink SHARP 缓冲区
    NVLS->>NVLS: 硬件归约 (multimem)
    NVLS->>Reduce: prims.directRecvDirectSend(offset, nelem)
    Reduce->>NVLS: 归约结果写回
    NVLS->>Bcast: prims.directRecvDirectSend(offset, nelem)
    Bcast->>App: 广播到所有 rank

Production pitfall: CollNet'sdirect->out == -1trap

📎 src/device/reduce_scatter.h:521has a line:

code
if (direct->out == -1) __trap();

If CollNet's out connection is not established (-1), directly__trap()causes the kernel to crash. This is defensive programming—CollNet depends on the NIC. If NIC initialization fails, out will be -1, and continuing execution at this point will lead to undefined behavior. In production, if you see a kernel trap, check whether the CollNet NIC is initialized properly.

10.5 Broadcast and Reduce: the two simplest collective operations

Broadcast: fan-out from root

broadcast.h'srunRing(📎 src/device/broadcast.h:14-64) logic is straightforward: the root node sends data, other nodes forward it, and the last node only receives.

code
if (rank == root) {
  if (inputBuf == outputBuf || isNetOffload) {
    prims.directSend(offset, offset, nelem);
  } else {
    prims.directCopySend(offset, offset, nelem);
  }
} else if (nextRank == root) {
  prims.directRecv(offset, nelem);
} else {
  prims.directRecvCopyDirectSend(offset, offset, nelem);
}

Three branches: root sends, root's predecessor receives, intermediate nodes forward. Note thatnextRank == rootchecks whether "this node's next is root," that is, this node is the last on the ring—it only receives and does not send.

Reduce: converge toward root

reduce.h'srunRing(📎 src/device/reduce.h:14-53) is the inverse operation of Broadcast:

code
if (prevRank == root) {
  prims.send(offset, nelem);
} else if (rank == root) {
  prims.recvReduceCopy(offset, offset, nelem, /*postOp=*/true);
} else {
  prims.recvReduceSend(offset, nelem);
}

prevRank == rootThe node at

Design consideration: why Broadcast/Reduce also use Ring

Broadcast and Reduce could theoretically use Tree to achieve lower latency, but NCCL chooses Ring because:the data volume of these two operations is usually small, Ring's implementation is simpler, and it can reuse AllReduce's Ring code path. Tree's complexity (root selection, thread splitting) does not bring obvious benefits in small-message scenarios.

Production pitfall: Broadcast's root node bandwidth bottleneck

Broadcast's root node must send all data. If root is a slow node, the entire Broadcast is slowed down. NCCL's response is:Broadcast also supports multiple channels, and each channel's root can be different. But note thatwork->rootis global, and all channels share the same root—this is determined by Broadcast's semantics (there is only one source). In production, if Broadcast is slow, check the root node's network bandwidth.

10.6 Algorithm selection matrix: RunWorkColl template specialization

All algorithm kernels are registered throughRunWorkColltemplate specialization (📎 src/device/all_reduce.h:228-788). Each specialization corresponds to a combination of "function × algorithm × protocol":

FunctionAlgorithmProtocolSpecialization location
AllReduceRINGSIMPLE📎 src/device/all_reduce.h:230-233
AllReduceTREESIMPLE📎 src/device/all_reduce.h:238-244
AllReduceCOLLNET_DIRECTSIMPLE📎 src/device/all_reduce.h:249-386
AllReduceNVLSSIMPLE📎 src/device/all_reduce.h:391-523
AllReduceNVLS_TREESIMPLE📎 src/device/all_reduce.h:528-634
AllReduceCOLLNET_CHAINSIMPLE📎 src/device/all_reduce.h:639-759
AllReduceRINGLL📎 src/device/all_reduce.h:764-766
AllReduceTREELL📎 src/device/all_reduce.h:771-773
AllReduceRINGLL128📎 src/device/all_reduce.h:778-780
AllReduceTREELL128📎 src/device/all_reduce.h:785-787

Note:CollNet and NVLS only support the SIMPLE protocol. This is because these two algorithms rely on hardware offload, and the low-latency synchronization mechanisms of LL/LL128 are incompatible with hardware offload—the latency of hardware reduction is far greater than LL's flag polling, so using LL instead increases overhead.

The inherent logic of protocol selection

  • LL: Small messages (< 8KB), low latency prioritized. Both Ring and Tree support it.
  • LL128: Medium messages (8KB - 1MB), 128-byte aligned. Both Ring and Tree support it.
  • SIMPLE: Large messages (> 1MB), bandwidth prioritized. All algorithms support it.

Production pitfalls: Combination constraints of protocols and algorithms

If the user forcibly specifiesNCCL_PROTO=LLbut the algorithm is CollNet, NCCL will fall back to SIMPLE during the tuning phase. In production, if you find that the protocol setting is not taking effect, check whether the algorithm supports that protocol.

Design reflection: Why the same AllReduce logic needs so many implementations

Reviewing this chapter, AllReduce has six algorithm implementations: Ring, Tree, CollNet Direct, CollNet Chain, NVLS, and NVLS Tree. This is not redundancy, but ratheroptimal solutions targeting different hardware topologies and message sizes:

  • Ring: General-purpose, suitable for large messages, highest bandwidth utilization.
  • Tree: Suitable for large-scale clusters, latency O(log n).
  • CollNet: Suitable for clusters with NICs that support reduction, offloading GPU computation.
  • NVLS: Suitable for single-node NVLink full connectivity, hardware multicast reduction.

NCCL's tuning module (Chapter 5) automatically selects based on message size, number of ranks, and topology. The device-side implementation only needs to ensure "every combination is correct"; the selection logic is on the host side.

Chapter Summary

This chapter dissectedsrc/devicesix algorithm kernel files under

1. Ring AllReduce(📎 src/device/all_reduce.h:14-83): Two-phase pipeline, reduce-scatter + all-gather, n-1 steps per phase.

2. Tree AllReduce(📎 src/device/all_reduce.h:86-225): Tree reduction, latency O(log n),runTreeSplituses thread splitting to implement reduction-broadcast pipeline.

3. AllGather(📎 src/device/all_gather.h:14-88): Ring single-phase, supports in-place and netOffload.

4. ReduceScatter(📎 src/device/reduce_scatter.h:14-56): Ring single-phase, is the reduce-scatter phase of AllReduce.

5. Broadcast/Reduce(📎 src/device/broadcast.h:14-64、📎 src/device/reduce.h:14-53): The simplest Ring variant.

6. CollNet/NVLS(📎 src/device/all_reduce.h:247-635): Hardware offload, only supports SIMPLE protocol.

Chapter Reflection and Self-Test

Q1: In the reduce-scatter phase of Ring AllReduce, step 0 usesdirectSend, intermediate steps usedirectRecvReduceDirectSend, and the last step usesdirectRecvReduceCopyDirectSend. If the last step'spostOp=trueis removed, in what scenarios would incorrect results occur?

Reference Analysis:postOp=truetriggers post-operations (such as division when computing the average). TakingncclAvgas an example, the reduction is summation, and postOp is division by nranks. IfpostOpis removed, the last step only performs reduction without division, and recvbuff stores the "sum" rather than the "average". In the reduce-scatter phase, each rank only retains the final result of one chunk, and this chunk happens to beringIx+0(📎 src/device/all_reduce.h:60). If postOp is missing, this chunk's sum is not divided by nranks, and the subsequent all-gather phase will propagate this incorrect "sum" to all ranks. Note: Only the last step needs postOp, because only this step produces a "complete reduction" result; intermediate steps' reductions are partial sums and do not need postOp. In production, if you find that AllReduce results are larger by a factor of nranks, check whether postOp is correctly passed.

Q2: runTreeSplitUnder the LL/LL128 protocol, threads are split 7:3 (📎 src/device/all_reduce.h:163), while under the Simple protocol they are split 1:1 (📎 src/device/all_reduce.h:157). If the LL protocol is forcibly changed to 1:1 as well, what would happen?

Reference Analysis: The reduction group of LL/LL128 needs to receive data from up to 3 child nodes and perform reduction (📎 src/device/all_reduce.h:187'sFanAsymmetric<NCCL_MAX_TREE_ARITY, 1>), which is computation-intensive; the broadcast group only performs copy-and-forward (📎 src/device/all_reduce.h:208'sFanAsymmetric<1, NCCL_MAX_TREE_ARITY>), which is computation-light. The 7:3 split gives the reduction group enough threads to handle 3-way reduction, while the broadcast group has fewer threads but sufficient. If changed to 1:1, the reduction group would have insufficient threads, making reduction the bottleneck; the broadcast group would have excess threads, wasting resources. More seriously, LL protocol's flag polling is busy-waiting, and more threads increase flag contention. In production, if you find that Tree AllReduce performs abnormally under the LL protocol, check whethernthreadsSplit's computation has been modified.

Q3: In AllGather'sisNetOffloadmode, only 1 warp drives Ring communication (📎 src/device/all_gather.h:32), while the remaining warps copy in parallel (📎 src/device/all_gather.h:76-82). If the finalbarrier_sync(14, nthreads)(📎 src/device/all_gather.h:87) is removed, in what scenarios would data races occur?

Reference Analysis:barrier_syncEnsure all warps (including communication warps and copy warps) complete this work before proceeding to the next work. If removed, the communication warp might start the next work's communication while the copy warp hasn't finished writing outputBuf, and the next work might reuse the same outputBuf. Specific scenario: two consecutive AllGathers — the first copy warp is still writing the tail of outputBuf while the second communication warp has already started writing new data to outputBuf, causing the first data to be overwritten. The comment states it clearly: "otherwise, we can have contention if next work will use the outputBuf in this work". Using barrier 14 instead of the default barrier is to avoid the barrier inside prims and__syncthreads(), preventing deadlock. In production, if AllGather results show intermittent errors, check whether the barrier in theisNetOffloadpath has been optimized away.

At this point, we have seen how the device-side algorithm kernels organize data flow. Each algorithm calls the primitives from the previous chapter throughPrimitives, and the algorithm layer only cares about "who sends to whom, which chunk to send, reduce or copy". The next chapter will dive into the transport layer abstraction, examining how P2P, SHM, NET, and NVLS are unified into a single interface, and how host-side proxy threads cooperate with device-side kernels to accomplish cross-machine communication.

Core pattern: all algorithms call primitives through the Primitives template class; algorithms are only responsible for "data flow topology", while primitives handle "data movement". This layering allows new algorithms to implement only topology logic without worrying about low-level synchronization. But regardless of how the topology changes, data must ultimately be transmitted over physical links. The next chapter will dive into the src/transport directory to see how NCCL uses a unified transport interface to mask the differences between P2P, SHM, NET, and NVLS, and the setup/connect/send/recv semantics of each transport. This is the foundation for understanding cross-machine communication.

CHAPTER 11

Chapter 11: Chapter 11: Transport Layer Abstraction: How P2P, SHM, NET, and NVLS Are Unified Under a Single Interface

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 11 / 25

Chapter 11: Transport Layer Abstraction: How P2P, SHM, NET, and NVLS Are Unified Under a Single Interface

In the previous chapter, we delved into the algorithm kernels and saw how Ring AllReduce splits data and performs two-phase reduction, and how Tree AllReduce leverages a tree structure to lower latency — but these algorithms only define the logical view of "who sends to whom, which chunk to send". Data must ultimately traverse real physical links: NVLink, PCIe, shared memory, or network cards. This chapter dissects the src/transport directory to see how NCCL uses a unified ncclTransport interface to mask the four physical channels — P2P, SHM, NET, and NVLS — behind a single facade, completing the last mile from algorithm topology to physical transmission.

I. Unified Interface: How ncclTransport Masks Four Physical Channels

Intuitive Model

Imagine a logistics company: whether a customer is sending a same-city express (P2P), intra-building delivery (SHM), cross-province transport (NET), or dedicated line direct delivery (NVLS), the front desk only fills out one "waybill". This waybill is thencclTransportstruct — it specifies the fixed actions each transport method must provide, such ascanConnect、setup、connect、free. Without this layer of abstraction, upper-layer algorithms would have to write four sets ofif-elseto determine which link to use, and adding a new hardware type would require modifying all algorithms.

Data Structures and Memory Layout

NCCL uses a global array to register all transports, with order determining priority:

📎 src/transport.cc:15-20

c
struct ncclTransport* ncclTransports[NTRANSPORTS] = {
  &p2pTransport,
  &shmTransport,
  &netTransport,
  &collNetTransport,
};

The array order determines selection order: P2P first, then SHM, then NET, and finally CollNet. Each transport is described by ancclTransportstruct, which contains acanConnectfunction pointer and twoncclTransportComm(one each for send/recv). Taking P2P as an example:

📎 src/transport/p2p.cc:1493-1498

c
struct ncclTransport p2pTransport = {"P2P",
                                     p2pCanConnect,
                                     {p2pSendSetup, p2pSendConnect, p2pSendFree, NULL, p2pSendProxySetup, NULL,
                                      p2pSendProxyFree, NULL, p2pProxyRegister, p2pProxyDeregister},
                                     {p2pRecvSetup, p2pRecvConnect, p2pRecvFree, NULL, p2pRecvProxySetup, NULL,
                                      p2pRecvProxyFree, NULL, p2pProxyRegister, p2pProxyDeregister}};

ncclTransportCommThe field order ofsetupis a fixed "lifecycle slot" sequence:connect(prepare resources),free(exchange connection info),proxySharedInit(release),proxySetup、proxyConnect、proxyFree、proxyProgress、proxyRegister、proxyDeregister(proxy shared initialization),proxyProgress. Note that P2P'sNULLslot isproxyProgress— because P2P uses the GPU to directly read/write the peer's memory, without needing host proxy threads to move data; while NET'ssendProxyProgress/recvProxyProgressis

, because network card I/O must be driven by host threads.

Scenario-Driven Walkthrough: How a Connection Selects a TransportselectTransport:

📎 src/transport.cc:23-44

c
template <int type>
static ncclResult_t selectTransport(struct ncclComm* comm, struct ncclTopoGraph* graph, struct ncclConnect* connect,
                                    int channelId, int peer, int connIndex, int* transportType) {
  struct ncclPeerInfo* myInfo = comm->peerInfo + comm->rank;
  struct ncclPeerInfo* peerInfo = comm->peerInfo + peer;
  struct ncclConnector* connector = (type == 1) ? comm->channels[channelId].peers[peer]->send + connIndex :
                                                  comm->channels[channelId].peers[peer]->recv + connIndex;
  for (int t = 0; t < NTRANSPORTS; t++) {
    struct ncclTransport* transport = ncclTransports[t];
    struct ncclTransportComm* transportComm = type == 1 ? &transport->send : &transport->recv;
    int ret = 0;
    NCCLCHECK(transport->canConnect(&ret, comm, graph, myInfo, peerInfo));
    if (ret) {
      connector->transportComm = transportComm;
      NCCLCHECK(transportComm->setup(comm, graph, myInfo, peerInfo, connect, connector, channelId, connIndex));
      if (transportType) *transportType = t;
      return ncclSuccess;
    }
  }
  WARN("No transport found for rank %d[%lx] -> rank %d[%lx]", myInfo->rank, myInfo->busId, peerInfo->rank,
       peerInfo->busId);
  return ncclSystemError;
}

type==1indicates the send direction,type==0indicates the recv direction. The loop queries each transport'scanConnectin turn: returningret=1means "I can do this job", immediately pointingconnector->transportCommto the corresponding direction of that transport, and calling itssetup. If all transports return 0, print a warning and returnncclSystemError。

canConnect's decision logic embodies each transport's "territory boundaries". Taking P2P as an example:

📎 src/transport/p2p.cc:129-157

c
ncclResult_t p2pCanConnect(int* ret, struct ncclComm* comm, struct ncclTopoGraph* graph, struct ncclPeerInfo* info1,
                           struct ncclPeerInfo* info2) {
  initCeOperation();
  int intermediateRank;
  int isCrossClique;
  NCCLCHECK(ncclTopoCheckP2p(comm, comm->topo, info1->rank, info2->rank, ret, NULL, &intermediateRank, NULL,
                             &isCrossClique));
  if (*ret == 0) return ncclSuccess;
  if (intermediateRank != -1) {
    if (useMemcpy) *ret = 0;
    return ncclSuccess;
  }
  if (!isCrossClique) {
    int useNet = 0;
    NCCLCHECK(ncclTopoCheckNet(comm->topo, info1->rank, info2->rank, &useNet));
    if (useNet) {
      *ret = 0;
      return ncclSuccess;
    }
  }
  if (info1->hostHash != comm->peerInfo[comm->rank].hostHash || info1->hostHash != info2->hostHash) {
    return ncclSuccess;
  }
  ...

P2P's decision chain: first ask the topology "is there a P2P path between the two ranks"; if there are intermediate hops (intermediateRank != -1) and CE memcpy is enabled, give up P2P and yield to SHM/NET; if the topology suggests going over the network (useNet), also give up; finally check whether they are on the same host. SHM's decision is simpler:

📎 src/transport/shm.cc:61-83

c
static ncclResult_t shmCanConnect(int* ret, struct ncclComm* comm, struct ncclTopoGraph* graph,
                                  struct ncclPeerInfo* info1, struct ncclPeerInfo* info2) {
  *ret = 0;
  initShmLocality();
  if (ncclParamShmDisable() == 1) return ncclSuccess;
  int useNet = 0;
  NCCLCHECK(ncclTopoCheckNet(comm->topo, info1->rank, info2->rank, &useNet));
  if (useNet) return ncclSuccess;
  if (info1->hostHash != info2->hostHash) return ncclSuccess;
  if (info1->shmDev != info2->shmDev) return ncclSuccess;
  *ret = 1;
  return ncclSuccess;
}

SHM requires the same host (hostHashidentical) and sharing the same/dev/shm(shmDevidentical, used for inter-container communication). NET almost always returns 1, only checking whether intra-node net is disabled when on the same host:

📎 src/transport/net.cc:160-168

c
static ncclResult_t canConnect(int* ret, struct ncclComm* comm, struct ncclTopoGraph* graph, struct ncclPeerInfo* info1,
                               struct ncclPeerInfo* info2) {
  *ret = 1;
  if (info1->hostHash == info2->hostHash) {
    NCCLCHECK(ncclTopoCheckNet(comm->topo, info1->rank, info2->rank, ret));
  }
  return ncclSuccess;
}

NET is the "fallback" — as long as no one ahead takes it, it takes it. NVLS'scanConnectdirectly returns 0:

📎 src/transport/nvls.cc:21-26

c
ncclResult_t nvlsCanConnect(int* ret, struct ncclComm* comm, struct ncclTopoGraph* graph, struct ncclPeerInfo* info1,
                            struct ncclPeerInfo* info2) {
  // This transport cannot be used for p2p
  *ret = 0;
  return ncclSuccess;
}

NVLS does not go through the regular peer-to-peer connection path; it establishes multicast groups separately viancclNvlsSetup, socanConnectalways returns 0.

mermaid
flowchart TD
    start["selectTransport(comm, peer, connIndex)"] --> loop{"遍历 ncclTransports[t]"}
    loop -->|t=0| p2p["p2pCanConnect()"]
    p2p --> p2p_chk{"拓扑有P2P路径<br/>且非中间跳<br/>且同主机?"}
    p2p_chk -->|是| use_p2p["connector->transportComm = p2pTransport<br/>调用 p2pSendSetup/p2pRecvSetup"]
    p2p_chk -->|否| shm["shmCanConnect()"]
    shm --> shm_chk{"同hostHash<br/>且同shmDev?"}
    shm_chk -->|是| use_shm["connector->transportComm = shmTransport<br/>调用 shmSendSetup/shmRecvSetup"]
    shm_chk -->|否| net["canConnect() (NET)"]
    net --> net_chk{"同主机时<br/>intra-node net 启用?"}
    net_chk -->|是/跨机| use_net["connector->transportComm = netTransport<br/>调用 sendSetup/recvSetup"]
    net_chk -->|否| collnet["collNetTransport"]
    collnet --> fail["WARN: No transport found<br/>return ncclSystemError"]
    use_p2p --> done["return ncclSuccess"]
    use_shm --> done
    use_net --> done

Design Thinking

[Design Inference and Architectural Trade-offs]

Why use "array order + canConnect voting" instead of an explicit routing table? Because the topology is dynamic: the same machine may have P2P unavailable due to factors such asNCCL_P2P_DISABLE, container isolation, CUDA IPC availability, etc., in which case it automatically degrades to SHM or NET. The voting mechanism lets each transport judge for itself "can I do this", and adding a new transport only requires adding an entry to the array, without changing the selection logic. This is precisely the embodiment of the open-closed principle in systems programming.

II. P2P: Four Forms of Same-Machine GPU Direct Connection

Intuitive Model

P2P is "handing things directly between neighbors" — GPU 0 directly reads and writes GPU 1's memory, without going through the CPU or NIC. Without P2P, same-machine multi-GPU communication would have to detour through host memory, doubling latency and halving bandwidth.

Data Structures and Memory Layout

P2P has four internal forms, distinguished byenum p2pType:

📎 src/transport/p2p.cc:19-24

c
enum p2pType {
  P2P_DIRECT,
  P2P_INTERMEDIATE,
  P2P_IPC,
  P2P_CUMEM
};
  • P2P_DIRECT: different GPUs within the same process, accessed directly via pointers (fastest).
  • P2P_INTERMEDIATE: no direct connection between the two GPUs, requiring forwarding through an intermediate GPU.
  • P2P_IPC: cross-process, using the traditionalcudaIpcOpenMemHandleto import the peer's memory.
  • P2P_CUMEM: cross-process, using the cuMem API (cuMemExportToShareableHandle) to import, supporting finer-grained memory management.

Core resource struct:

📎 src/transport/p2p.cc:79-94

c
struct p2pResources {
  enum p2pType type;
  union {
    struct ncclSendMem* sendDevMem;
    struct ncclRecvMem* recvDevMem;
  };
  void* sendMemIpc;
  int sendMemSameProc;
  void* recvMemIpc;
  int recvMemSameProc;
  // CE memcpy support
  struct p2pShmProxyInfo proxyInfo;
  struct p2pShm* shm;
  struct p2pShm* devShm;
  ncclShmIpcDesc_t desc;
};

sendDevMem/recvDevMemis a union — the sender only cares aboutsendDevMem, the receiver only cares aboutrecvDevMem, sharing one block of memory.sendMemIpc/recvMemIpcstores the imported peer memory handle,sendMemSameProc/recvMemSameProcmarks whether it is the same process (determining whether to usencclCuMemFreeAddrorcudaIpcCloseMemHandle)。

when releasing).p2pConnectInfoThe connection info struct

📎 src/transport/p2p.cc:38-44

c
struct p2pConnectInfo {
  int rank;
  int read;
  struct ncclP2pBuff p2pBuff;
  // Used by CE memcpy
  ncclShmIpcDesc_t desc;
};
static_assert(sizeof(struct p2pConnectInfo) <= CONNECT_SIZE, "p2pConnectInfo is too large");

static_assertCopyCONNECT_SIZEensures the connection info does not exceedread(the fixed buffer size for a single bootstrap exchange).read=1The field determines the data flow direction:read=0means the receiver actively reads the sender's memory (P2P Read),

means the sender actively writes the receiver's memory (P2P Write).

Scenario-Driven Walkthrough: Establishing P2P SendselectTransportWhenp2pSendSetup:

📎 src/transport/p2p.cc:393-471

c
ncclResult_t p2pSendSetup(struct ncclComm* comm, struct ncclTopoGraph* graph, struct ncclPeerInfo* myInfo,
                          struct ncclPeerInfo* peerInfo, struct ncclConnect* connectInfo, struct ncclConnector* send,
                          int channelId, int connIndex) {
  struct p2pResources* resources;
  struct ncclP2pRequest req;
  NCCLCHECK(ncclCalloc(&resources, 1));
  send->transportResources = resources;
  int useRead, intermediateRank;
  NCCLCHECK(p2pGetInfo(comm, myInfo, peerInfo, &useRead, &intermediateRank));
  if (useMemcpy) useRead = 0;
  ...
  int sendSize = sizeof(struct ncclSendMem);
  if (info->read) sendSize += comm->buffSizes[NCCL_PROTO_SIMPLE];
  ALIGN_SIZE(sendSize, CUDA_IPC_MIN);
  ...

CopysendSizeKey points:ncclSendMemIn P2P Read mode, the SIMPLE protocol buffer size must be added extra — because in read mode the sender's SIMPLE buffer is directly read by the receiver, it must be allocated together withALIGN_SIZE(sendSize, CUDA_IPC_MIN)in the same shareable memory block.

ensures the size is aligned to the CUDA IPC minimum granularity.intermediateRankThen, based on

📎 src/transport/p2p.cc:416-437

c
  if (intermediateRank == -1) {
    info->rank = myInfo->rank;
    if (P2P_SAME_PID(myInfo, peerInfo) && ncclParamP2pDirectDisable() == 0 && useMemcpy == 0) {
      resources->type = P2P_DIRECT;
      ...
    } else {
      if (ncclCuMemEnable()) {
        resources->type = P2P_CUMEM;
        ...
      } else {
        resources->type = P2P_IPC;
        ...
      }
    }
    send->conn.flags |= info->read ? NCCL_P2P_READ : NCCL_P2P_WRITE;
  } else {
    resources->type = P2P_INTERMEDIATE;
    info->rank = intermediateRank;
    ...
  }

P2P_SAME_PIDCopy

📎 src/transport/p2p.cc:334-335

c
#define P2P_SAME_PID(MYINFO, PEERINFO) \
  ((MYINFO->hostHash == PEERINFO->hostHash) && (MYINFO->pidHash == PEERINFO->pidHash))

CopyP2P_DIRECTSame process and direct not disabled and memcpy not enabled means the fastest

— directly taking the peer pointer. Otherwise go through IPC/CUMEM.

📎 src/transport/p2p.cc:457-468

c
  NCCLCHECK(ncclProxyConnect(comm, TRANSPORT_P2P, 1, info->rank, &send->proxyConn));
  if (useMemcpy) {
    NCCLCHECK(ncclProxyCallBlocking(comm, &send->proxyConn, ncclProxyMsgSetup, NULL, 0, &resources->proxyInfo,
                                    sizeof(struct p2pShmProxyInfo)));
    memcpy(&info->desc, &resources->proxyInfo.desc, sizeof(ncclShmIpcDesc_t));
  } else {
    NCCLCHECK(ncclProxyCallBlocking(comm, &send->proxyConn, ncclProxyMsgSetup, &req, sizeof(struct ncclP2pRequest),
                                    &info->p2pBuff, sizeof(struct ncclP2pBuff)));
    NCCLCHECK(p2pMap(comm, &send->proxyConn, myInfo, comm->peerInfo + info->rank, &info->p2pBuff,
                     (void**)&resources->sendDevMem, &resources->sendMemIpc));
    resources->sendMemSameProc = P2P_SAME_PID(myInfo, (comm->peerInfo + info->rank));
  }

ncclProxyCallBlockingCopyp2pSendProxySetupis a synchronous RPC: the host thread sends a message to the proxy thread, the proxy thread callsncclP2pBuffto allocate a shareable buffer, and sends backp2pMap(including the IPC handle). Then

p2pMapmaps the peer buffer into the local address space.

📎 src/transport/p2p.cc:349-390

c
static ncclResult_t p2pMap(struct ncclComm* comm, struct ncclProxyConnector* proxyConn, struct ncclPeerInfo* myInfo,
                           struct ncclPeerInfo* peerInfo, struct ncclP2pBuff* p2pBuff, void** devMem, void** ipcPtr) {
  if (P2P_SAME_PID(myInfo, peerInfo)) {
    if (peerInfo->cudaDev != myInfo->cudaDev) {
      cudaError_t err = cudaDeviceEnablePeerAccess(peerInfo->cudaDev, 0);
      ...
      if (ncclCuMemEnable()) {
        NCCLCHECK(ncclCuMemAllocAddr(devMem, &p2pBuff->ipcDesc.memHandle, p2pBuff->size));
        CUCHECK(cuMemRelease(p2pBuff->ipcDesc.memHandle));
        *ipcPtr = *devMem;
        ...
      } else {
        *devMem = p2pBuff->directPtr;
        *ipcPtr = NULL;
      }
    } else {
      *devMem = p2pBuff->directPtr;
      *ipcPtr = NULL;
    }
  } else {
    NCCLCHECK(ncclP2pImportShareableBuffer(comm, peerInfo->rank, p2pBuff->size, &p2pBuff->ipcDesc, devMem,
                                           p2pBuff->directPtr, ncclMemOffload));
    *ipcPtr = *devMem;
  }
  return ncclSuccess;
}

CopycudaDeviceEnablePeerAccessSame process, different GPUs: firstdirectPtropens the P2P channel, then directly usesncclP2pImportShareableBuffer(because the same process shares the address space). Cross-process: calls

to import the peer memory handle.

Concurrency Control and Hardware InteractionncclSendMem/ncclRecvMemP2P synchronization relies on thehead/tailpointer inhead. The sender writestailto tell the receiver "how far I've written", and the receiver writes

📎 src/transport/p2p.cc:571-576

c
  } else {
    send->conn.tail = &remDevMem->tail;
    send->conn.head = &resources->sendDevMem->head;
    send->conn.ptrExchange = &resources->sendDevMem->ptrExchange;
    send->conn.redOpArgExchange = resources->sendDevMem->redOpArgExchange;
  }

headCopysendDevMem,tailpoints to the localremDevMempoints to the peer

. The GPU kernel achieves cross-GPU synchronization by reading and writing these two pointers, without CPU intervention.

Production Pitfall GuidePitfall 1: P2P Read and memcpy are mutually exclusive.p2pSendConnect:

📎 src/transport/p2p.cc:551-559

c
  for (int p = 0; p < NCCL_NUM_PROTOCOLS; p++) {
    if (info->read && p == NCCL_PROTO_SIMPLE) {
      /* For P2P Read the SIMPLE buffer is local (ncclSendMem) */
      if (resources->sendDevMem == NULL) return ncclInternalError; // We should not use read + memcpy
      send->conn.buffs[p] = (char*)(resources->sendDevMem + 1);
    } else {
      send->conn.buffs[p] = buff;
      buff += comm->buffSizes[p];
    }
  }

Copyread=1IfsendDevMem==NULLbutncclInternalError, directly returnNCCL_P2P_READ_ENABLE=1. If you see this error in production, check whether bothNCCL_P2P_USE_CUDA_MEMCPY=1and

are set — these two have conflicting semantics. p2pSendFreePitfall 2: Cross-process release order.sendMemSameProcDetermines the release method based on

📎 src/transport/p2p.cc:624-651

c
ncclResult_t p2pSendFree(struct ncclComm* comm, struct ncclConnector* send) {
  struct p2pResources* resources = (struct p2pResources*)send->transportResources;
  if (resources) {
    if (ncclCuMemEnable()) {
      if (resources->sendMemIpc) {
        if (resources->sendMemSameProc) {
          NCCLCHECK(ncclCuMemFreeAddr(resources->sendMemIpc, comm->memManager));
        } else {
          NCCLCHECK(ncclCudaFree(resources->sendMemIpc, comm->memManager));
        }
      }
      ...

CopyncclCuMemFreeAddrSame process usesncclCudaFree(release physical memory). Getting this backwards causes memory leaks or use-after-free.

3. SHM: The "Who Hosts the Memory" Debate in Shared Memory

Intuitive Model

SHM is "two processes sharing a whiteboard"—the sender writes, the receiver reads. But whose house is the whiteboard in? At the sender's house (sender-side), with the receiver coming over to read? Or at the receiver's house (receiver-side), with the sender going over to write? This is the problem that theNCCL_SHM_LOCALITYparameter is meant to solve.

Data Structures and Memory Layout

📎 src/transport/shm.cc:28-34

c
struct shmSendResources {
  struct ncclRecvMem* remHostMem;
  struct ncclRecvMem* devRemHostMem;
  ncclShmIpcDesc_t remDesc;
  struct ncclSendMem* hostMem;
  struct ncclSendMem* devHostMem;
};

struct shmRecvResources {
  struct ncclSendMem* remHostMem;
  struct ncclSendMem* devRemHostMem;
  ncclShmIpcDesc_t remDesc;
  struct ncclRecvMem* hostMem;
  struct ncclRecvMem* devHostMem;
};

Note thathostMemanddevHostMemappear in pairs:hostMemis the host-side pointer,devHostMemis the device-side pointer (mapped via UVA or cuMem).remHostMem/devRemHostMemis the local mapping of the peer's shared memory.

Scenario-Driven Walkthrough: SHM's locality Selection

shmSendSetupdetermines how much memory to allocate based on locality:

📎 src/transport/shm.cc:88-119

c
static ncclResult_t shmSendSetup(struct ncclComm* comm, struct ncclTopoGraph* graph, struct ncclPeerInfo* myInfo,
                                 struct ncclPeerInfo* peerInfo, struct ncclConnect* connectInfo,
                                 struct ncclConnector* send, int channelId, int connIndex) {
  struct shmSendResources* resources;
  struct shmConnectInfo* info = (struct shmConnectInfo*)connectInfo;
  size_t shmSize = sizeof(struct ncclSendMem);
  struct shmRequest req;

  NCCLCHECK(ncclCalloc(&resources, 1));
  send->transportResources = resources;

  if (shmLocality == SHM_SEND_SIDE) {
    for (int p = 0; p < NCCL_NUM_PROTOCOLS; p++) shmSize += comm->buffSizes[p];
  }
  req.size = shmSize;
  if (myInfo->hostHash == peerInfo->hostHash && myInfo->pidHash == peerInfo->pidHash) req.legacy = true;
  else req.legacy = false;

  NCCLCHECK(ncclProxyConnect(comm, TRANSPORT_SHM, 1, myInfo->rank, &send->proxyConn));
  NCCLCHECK(ncclProxyCallBlocking(comm, &send->proxyConn, ncclProxyMsgSetup, (void*)&req, sizeof(struct shmRequest),
                                  (void*)info, sizeof(struct shmConnectInfo)));

  info->rank = comm->rank;
  resources->hostMem = (struct ncclSendMem*)info->buf.hptr;
  resources->devHostMem = (struct ncclSendMem*)info->buf.dptr;
  ...

shmLocality == SHM_SEND_SIDEWhen , the sender allocates the data buffer (shmSizeplus all protocol buffers); otherwise it only allocates thencclSendMemcontrol structure.req.legacymarks whether it's the same process—same process can use traditionalmmap, cross-process requires cuMem or/dev/shmfiles.

shmSendConnectdetermines based on locality whetherbuffspoints to local or peer:

📎 src/transport/shm.cc:153-176

c
static ncclResult_t shmSendConnect(struct ncclComm* comm, struct ncclConnect* connectInfo, int nranks, int rank,
                                   struct ncclConnector* send) {
  struct shmConnectInfo* info = (struct shmConnectInfo*)connectInfo;
  struct shmSendResources* resources = (struct shmSendResources*)send->transportResources;
  char* buff;

  NCCLCHECK(ncclShmImportShareableBuffer(comm, info->rank, &info->desc, (void**)&resources->remHostMem,
                                         (void**)&resources->devRemHostMem, &resources->remDesc));

  buff = shmLocality == SHM_SEND_SIDE ? (char*)(resources->devHostMem + 1) : (char*)(resources->devRemHostMem + 1);
  for (int p = 0; p < NCCL_NUM_PROTOCOLS; p++) {
    send->conn.buffs[p] = buff;
    buff += comm->buffSizes[p];
  }
  send->conn.tail = &resources->devRemHostMem->tail;
  send->conn.head = &resources->devHostMem->head;
  send->conn.stepSize = comm->buffSizes[NCCL_PROTO_SIMPLE] / NCCL_STEPS;
  ...

SHM_SEND_SIDE:buffspoints to localdevHostMem(the sender writes its own memory);SHM_RECV_SIDE:buffspoints to peerdevRemHostMem(the sender writes the receiver's memory).headalways points to local,tailalways points to peer—because the sender updateshead, and the receiver updatestail。

Design Considerations

[Design Inference and Architectural Trade-offs]

Why default toSHM_RECV_SIDE? Because the receiver typically needs to copy data from shared memory to its own GPU memory. If the shared memory is local to the receiver, the copy path is shorter (local memory → local GPU), avoiding cross-NUMA access. Although the sender writing to remote memory incurs an extra cross-node write, the sender is usually a compute-intensive GPU, and write operations can proceed asynchronously.

Production Pitfall Guide

Pitfall: Between containers,/dev/shmis not shared. shmCanConnectCheckinfo1->shmDev != info2->shmDev:

📎 src/transport/shm.cc:76-78

c
  TRACE(NCCL_INIT | NCCL_SHM, "peer1 shmDev %lx peer2 shmDev %lx", info1->shmDev, info2->shmDev);
  if (info1->shmDev != info2->shmDev) return ncclSuccess;

If two containers mount different/dev/shm,shmDevdiffer, SHM automatically degrades to NET. In production, if you find same-host communication going over the network, check whether the containers'/dev/shmmounts are consistent.

4. NET: Network Transport's Mapping Table and Proxy Progress

Intuitive Model

NET is "intercity express delivery"—data is packaged and handed to the NIC, which sends it to the peer over fiber. But the NIC doesn't understand GPU memory addresses; it needs an "address mapping table" to translate GPU virtual addresses into physical addresses the NIC can understand. This table isconnectMap。

Data Structures and Memory Layout

📎 src/transport/net.cc:73-86

c
struct connectMapMem {
  char* gpuPtr;
  char* cpuPtr;
  ssize_t size;
  ncclIpcDesc ipcDesc;
  ncclShmIpcDesc_t attachDesc;
  ncclShmIpcDesc_t createDesc;
};

struct connectMap {
  int sameProcess;
  int shared;
  int cudaDev;
  // First 3 bits of offsets determine the mem bank. 001 is host mem, 011 is dev mem, 101 is shared host mem and 111
  // is shared dev mem.
  struct connectMapMem mems[NCCL_NET_MAP_MEMS];
  // Offsets. 3 MSBs indicate mem bank, 111 indicates NULL.
  struct {
    uint32_t sendMem;
    uint32_t recvMem;
    uint32_t buffs[NCCL_NUM_PROTOCOLS];
  } offsets;
};

connectMapis a "memory bank" system:memsThe array has 5 slots (NCCL_NET_MAP_MEMS=5), corresponding to host mem, dev mem, shared host mem, shared dev mem, and GDC mem.offsetsEach field in is a 32-bit integer; the upper 3 bits encode "which bank," and the lower 29 bits encode "offset within the bank."

Decoding macro:

📎 src/transport/net.cc:36-46

c
#define NCCL_NET_MAP_OFFSET_BANK(mapStruct, offsetName) ((mapStruct)->offsets.offsetName >> 30)

#define NCCL_NET_MAP_OFFSET_NULL(mapStruct, offsetName) (((mapStruct)->offsets.offsetName >> 29) == 0)

#define NCCL_NET_MAP_GET_POINTER(mapStruct, cpuOrGpu, offsetName) \
  (NCCL_NET_MAP_OFFSET_NULL(mapStruct, offsetName) ? \
     NULL : \
     (mapStruct)->mems[NCCL_NET_MAP_OFFSET_BANK(mapStruct, offsetName)].cpuOrGpu##Ptr + \
       ((mapStruct)->offsets.offsetName & NCCL_NET_MAP_MASK_OFFSET))

#define NCCL_NET_MAP_DEV_MEM(mapStruct, offsetName) (((mapStruct)->offsets.offsetName & NCCL_NET_MAP_MASK_DEVMEM) != 0)

NCCL_NET_MAP_GET_POINTER(map, gpu, sendMem)After expansion: takeoffsets.sendMem's upper 2 bits as the bank index, add the lower 29-bit offset tomems[bank].gpuPtr, to get the actual pointer. This encoding compresses "which memory region + offset within the region" into a single 32-bit integer, savingconnectMap's transmission size.

Scenario-Driven Walkthrough: Mapping Establishment in sendProxyConnect

sendProxyConnectis NET's most complex function, responsible for establishing the NIC connection, allocating buffers, and registering memory:

📎 src/transport/net.cc:858-1041

c
static ncclResult_t sendProxyConnect(struct ncclProxyConnection* connection, struct ncclProxyState* proxyState,
                                     void* reqBuff, int reqSize, void* respBuff, int respSize, int* done) {
  struct sendNetResources* resources = (struct sendNetResources*)(connection->transportResources);
  ...
  if (resources->shared) {
    // Shared buffers
    ...
    if (resources->maxRecvs > 1 && ncclParamNetSharedComms()) {
      // Connect or reuse connection for a netdev/remote rank.
      ...
      if (comms->sendComm[resources->channelId] == NULL &&
          comms->activeConnect[resources->channelId] == (resources->tpLocalRank + 1)) {
        ret = proxyState->ncclNet->connect(proxyState->netContext, resources->netDev, req->handle,
                                           comms->sendComm + resources->channelId, &resources->netDeviceHandle);
      }
      ...

maxRecvs > 1When , "shared connection" is enabled: multiple channels reuse the same NIC connection, reducing the connection count.activeConnectThe array ensures only one local rank initiates the connection, avoiding duplication.

Next, allocate buffers and register:

📎 src/transport/net.cc:933-956

c
  if (resources->shared == 0) {
    // Only allocate dedicated buffers for ring/tree, not for p2p
    for (int p = 0; p < NCCL_NUM_PROTOCOLS; p++) {
      NCCL_NET_MAP_ADD_POINTER(map, 0, p != NCCL_PROTO_LL && resources->useGdr ? 1 : 0, proxyState->buffSizes[p],
                               buffs[p]);
      resources->buffSizes[p] = proxyState->buffSizes[p];
    }
  } else {
    // Get shared buffers
    int bank = resources->useGdr ? NCCL_NET_MAP_SHARED_DEVMEM : NCCL_NET_MAP_SHARED_HOSTMEM;
    struct connectMapMem* mapMem = map->mems + bank;
    NCCLCHECK(sharedNetBuffersInit(proxyState, resources->useGdr, resources->tpLocalRank, 0, map->sameProcess,
                                   proxyState->p2pnChannels, &mapMem->gpuPtr, &mapMem->cpuPtr, &mapMem->size,
                                   &mapMem->ipcDesc));
    resources->buffSizes[NCCL_PROTO_SIMPLE] = mapMem->size;
    ...

NCCL_NET_MAP_ADD_POINTERThe macro registers the buffer toconnectMap:

📎 src/transport/net.cc:48-62

c
#define NCCL_NET_MAP_ADD_POINTER(mapStruct, shared, dev, memSize, offsetName) \
  do { \
    int bank = NCCL_NET_MAP_MASK_USED + (dev) * NCCL_NET_MAP_MASK_DEVMEM + (shared) * NCCL_NET_MAP_MASK_SHARED; \
    if ((shared) == 0) { \
      if (dev) { \
        (mapStruct)->offsets.offsetName = bank + (mapStruct)->mems[NCCL_NET_MAP_DEVMEM].size; \
        (mapStruct)->mems[NCCL_NET_MAP_DEVMEM].size += memSize; \
      } else { \
        (mapStruct)->offsets.offsetName = bank + (mapStruct)->mems[NCCL_NET_MAP_HOSTMEM].size; \
        (mapStruct)->mems[NCCL_NET_MAP_HOSTMEM].size += memSize; \
      } \
    } else { \
      (mapStruct)->offsets.offsetName = bank; \
    } \
  } while (0);

Non-shared buffer: write the current bank'ssizeas the offset intooffsets, thensize += memSize—this is a bump allocator. Shared buffer: directly write the bank number with offset 0 (because the entire shared buffer is one bank).

Finally, register the memory with the NIC:

📎 src/transport/net.cc:1004-1035

c
  for (int p = 0; p < NCCL_NUM_PROTOCOLS; p++) {
    resources->buffers[p] = NCCL_NET_MAP_GET_POINTER(map, cpu, buffs[p]);
    if (resources->buffers[p]) {
#if CUDA_VERSION >= 11070
      int type = NCCL_NET_MAP_DEV_MEM(map, buffs[p]) ? NCCL_PTR_CUDA : NCCL_PTR_HOST;
      if (type == NCCL_PTR_CUDA && resources->useDmaBuf) {
        int dmabuf_fd;
        size_t dmaBufSize = resources->buffSizes[p];
        ALIGN_SIZE(dmaBufSize, ncclOsGetPageSize());
        CUCHECK(cuMemGetHandleForAddressRange((void*)&dmabuf_fd, (CUdeviceptr)resources->buffers[p], dmaBufSize,
                                              CU_MEM_RANGE_HANDLE_TYPE_DMA_BUF_FD,
                                              getHandleForAddressRangeFlags(resources->useGdr)));
        NCCLCHECK(proxyState->ncclNet->regMrDmaBuf(resources->netSendComm, resources->buffers[p],
                                                   resources->buffSizes[p], type, 0ULL, dmabuf_fd,
                                                   &resources->mhandles[p]));
        (void)close(dmabuf_fd);
      } else
#endif
      {
        NCCLCHECK(proxyState->ncclNet->regMr(resources->netSendComm, resources->buffers[p], resources->buffSizes[p],
                                             NCCL_NET_MAP_DEV_MEM(map, buffs[p]) ? NCCL_PTR_CUDA : NCCL_PTR_HOST,
                                             &resources->mhandles[p]));
      }
      ...

Prefer the DMA-BUF path (cuMemGetHandleForAddressRangeobtains the fd and passes it to the NIC plugin); on failure, fall back toregMr(traditional nv_peermem GDR).

Concurrency Control and Hardware Interaction: sendProxyProgress's Three-Stage Pipeline

sendProxyProgressis NET's data movement engine, using a "post → transmit → done" three-stage approach:

📎 src/transport/net.cc:1324-1491

c
static ncclResult_t sendProxyProgress(struct ncclProxyState* proxyState, struct ncclProxyArgs* args) {
  ...
  if (args->state == ncclProxyOpProgress) {
    int p = args->protocol;
    int maxDepth = std::min(NCCL_STEPS, NCCL_SHARED_STEPS / args->nsubs);
    for (int s = 0; s < args->nsubs; s++) {
      struct ncclProxySubArgs* sub = args->subs + s;
      ...
      // Post buffers to the GPU
      if (sub->posted < sub->nsteps && sub->posted < sub->done + maxDepth) {
        ...
        if (resources->shared) {
          ...
          volatile uint64_t* sendHead = resources->gdcSync ? resources->gdcSync : &resources->sendMem->head;
          sub->posted += args->sliceSteps;
          *sendHead = sub->base + sub->posted - NCCL_STEPS;
          if (resources->gdcSync) wc_store_fence(); // Flush out WC write
        } else {
          sub->posted += args->sliceSteps;
        }
        ...
        continue;
      }
      // Check whether we received data from the GPU and send it to the network
      if (sub->transmitted < sub->posted && sub->transmitted < sub->done + NCCL_STEPS) {
        ...
        if (connFifo[buffSlot].size != -1 && (*recvTail > tail || p == NCCL_PROTO_LL)) {
          ...
          if (ready) {
            ...
            NCCLCHECK(proxyState->ncclNet->isend(resources->netSendComm, buff, size, resources->tpRank,
                                                 sub->sendMhandle, phandle, sub->requests + buffSlot));
            ...
  • post: The proxy thread updatessendMem->head, telling the GPU "the buffer is ready, you can write data."
  • transmit: Check whetherrecvMem->tailhas advanced (GPU has finished writing), checkconnFifo[buffSlot].size != -1(data size has been filled), then callncclNet->isendto initiate an asynchronous send.
  • done: CallncclNet->testto check whether the send is complete, updatesendMem->headto return the buffer.

wc_store_fence()is a write-combining barrier—in the GDRCopy scenario, after the CPU writesgdcSync, the write-combining buffer must be flushed, otherwise the GPU won't see the update.

Production Pitfall Guide

Pitfall 1: LL128 protocol's flag validation.When data is in sysmem (non-GDR), the proxy thread must check the LL128 flag line by line:

📎 src/transport/net.cc:1388-1403

c
          if (p == NCCL_PROTO_LL128) {
            ready = resources->useGdr;
            if (!ready) {
              uint64_t flag = sub->base + sub->transmitted + 1;
              int nFifoLines = DIVUP(connFifo[buffSlot].size, sizeof(uint64_t) * NCCL_LL128_LINEELEMS);
              volatile uint64_t* lines = (volatile uint64_t*)buff;
              ready = 1;
              for (int i = 0; i < nFifoLines; i++) {
                if (lines[i * NCCL_LL128_LINEELEMS + NCCL_LL128_DATAELEMS] != flag) {
                  ready = 0;
                  break;
                }
              }
            }
          }

Because the GPU only calledthreadfence(), the data may still be in the L2 cache and not yet landed in sysmem. The proxy thread must confirm that each line's flag is correct before sending. In production, if you find LL128 data corruption, checkuseGdrIs it correct—under the GDR path, data lands directly in VRAM and does not require line-by-line verification.

Pitfall 2: Memory ordering of GDRCopy flush.On the receiving side, inrecvProxyProgressthere is a clever piece of inline assembly:

📎 src/transport/net.cc:1664-1682

c
          if (totalSize > 0 && p == NCCL_PROTO_SIMPLE && needFlush) {
            struct recvNetResources* resources = (struct recvNetResources*)(subGroup->connection->transportResources);
            if (resources->gdcFlush) {
#if defined(__x86_64__)
              asm volatile("mfence" ::: "memory");
              asm volatile("mov (%0), %%eax" ::"l"(resources->gdcFlush) : "%eax", "memory");
#else
              std::atomic_thread_fence(std::memory_order_seq_cst);
              uint64_t dummy;
              NCCLCHECK(ncclGdrCudaRead(resources->gdrDesc, &dummy, resources->gdcFlush, sizeof(dummy)));
#endif
            }

mfenceEnsure that the read of the CQE poll is not reordered before the flush read;mov (%0), %%eaxForce a PCIe read to make the CPU stall until all prior PCIe posted writes (including NIC DMA) are committed. This is the key to preventing "the NIC says the write is done but the data is still in the PCIe buffer" in the GDRCopy scenario. Remove this section, and the receiver may read stale data.

mermaid
sequenceDiagram
    participant GPU as GPU Kernel
    participant SM as ncclSendMem
    participant Proxy as sendProxyProgress
    participant NIC as ncclNet->isend
    participant Peer as 对端网卡

    GPU->>SM: 写数据到 buffs[p]
    GPU->>SM: 更新 recvMem->tail
    Proxy->>SM: 读 recvTail, connFifo[buffSlot].size
    Proxy->>Proxy: 检查 ready (LL128 flag / GDR)
    Proxy->>NIC: isend(comm, buff, size, mhandle)
    NIC->>Peer: DMA 发送
    Proxy->>NIC: test(request, &done)
    NIC-->>Proxy: done=1
    Proxy->>SM: 更新 sendMem->head (归还缓冲区)
    Proxy->>GPU: 下一轮 post

V. NVLS: Multicast groups and UC/MC memory binding

Intuitive model

NVLS is a "broadcast station"—one rank writes data to a multicast group, and the hardware automatically replicates it to all subscribers. Traditional AllReduce requires N-1 point-to-point transfers, while NVLS needs only 1 multicast write + 1 multicast read. Without NVLS, the latency of large-scale AllReduce grows linearly with the number of ranks.

Data structures and memory layout

The core of NVLS is the binding of "UC (unicast) memory" and "MC (multicast) memory."nvlsAllocBindUcAllocate UC memory and bind it to an MC group:

📎 src/transport/nvls.cc:225-277

c
static ncclResult_t nvlsAllocBindUc(struct ncclComm* comm, const struct ncclMcPartition* partition, size_t size,
                                    struct ncclNvlsUcSegment* outUc) {
  CUmemAllocationProp ucprop;
  ...
  ucprop.type = CU_MEM_ALLOCATION_TYPE_PINNED;
  ucprop.location.type = CU_MEM_LOCATION_TYPE_DEVICE;
  ucprop.location.id = comm->cudaDev;
  ucprop.requestedHandleTypes = ncclCuMemHandleType;
  CUCHECKGOTO(cuMemGetAllocationGranularity(&ucgran, &ucprop, CU_MEM_ALLOC_GRANULARITY_RECOMMENDED), ret, fail);
  ALIGN_SIZE(ucsize, ucgran);
  CUCHECKGOTO(cuMemAddressReserve((CUdeviceptr*)&ucptr, ucsize, ucgran, 0U, 0), ret, fail);
  CUCHECKGOTO(cuMemCreate(&ucHandle, ucsize, &ucprop, 0), ret, fail1);
  CUCHECKGOTO(cuMemMap((CUdeviceptr)ucptr, ucsize, 0, ucHandle, 0), ret, fail2);
  CUCHECKGOTO(cuMemSetAccess((CUdeviceptr)ucptr, ucsize, &comm->nvlsResources->accessDesc, 1), ret, fail3);
  CUDACHECKGOTO(cudaMemset(ucptr, 0, ucsize), ret, fail3);
  NCCLCHECKGOTO(ncclMemTrack(comm->memManager, ucptr, ucsize, ucHandle, ncclCuMemHandleType, ncclMemPersist), ret,
                fail3);
  NCCLCHECKGOTO(bootstrapIntraNodeBarrier(comm->bootstrap, comm->localRankToRank, comm->localRank, comm->localRanks,
                                          comm->localRankToRank[0]),
                ret, fail3);
  NCCLCHECKGOTO(ncclMcPartitionBindMem(partition, 0 /*offsetInPartition*/, ucHandle, 0 /*memOffset*/, ucsize), ret,
                fail3);
  ...

Flow:cuMemCreateAllocate physical memory →cuMemMapmap to a virtual address →cuMemSetAccessset GPU access permissions →ncclMcPartitionBindMembind the UC physical memory to the specified offset of the MC group. After binding, any rank writing to the MC address causes the hardware to copy the data to all bound UC memory.

NotebootstrapIntraNodeBarrierbeforecuMulticastBindMem—the comment says this is to "mitigate the possible hang in cuMulticastBindMem during abort." This is a hardware-level defense: if a rank aborts during the binding process, other ranks may hang incuMulticastBindMem.

Scenario-driven Walkthrough: Buffer layout of ncclNvlsBufferSetup

📎 src/transport/nvls.cc:279-368

c
ncclResult_t ncclNvlsBufferSetup(struct ncclComm* comm) {
  ...
  nvlsStepSize = comm->nvlsChunkSize;
  buffSize = nvlsStepSize * NCCL_STEPS;
  nvlsPerRankSize = nChannels * 2 * buffSize;
  nvlsTotalSize = nvlsPerRankSize * nHeads;
  ...
  if (resources->dataUc.ptr == NULL) {
    NCCLCHECKGOTO(nvlsAllocBindUc(comm, &resources->dataPartition, nvlsTotalSize, &resources->dataUc), res, fail);
  }
  ...
  for (int h = 0; h < nHeads; h++) {
    int nvlsPeer = comm->nRanks + 1 + h;
    for (int c = 0; c < nChannels; c++) {
      struct ncclChannel* channel = comm->channels + c;
      struct ncclChannelPeer* peer = channel->peers[nvlsPeer];

      // Reduce UC -> MC
      peer->send[1].conn.buffs[NCCL_PROTO_SIMPLE] = (char*)resources->dataUc.ptr + (h * 2 * nChannels + c) * buffSize;
      peer->recv[0].conn.buffs[NCCL_PROTO_SIMPLE] =
        (char*)resources->dataPartition.ptr + (h * 2 * nChannels + c) * buffSize;

      // Broadcast MC -> UC
      peer->recv[1].conn.buffs[NCCL_PROTO_SIMPLE] =
        (char*)resources->dataUc.ptr + ((h * 2 + 1) * nChannels + c) * buffSize;
      peer->send[0].conn.buffs[NCCL_PROTO_SIMPLE] =
        (char*)resources->dataPartition.ptr + ((h * 2 + 1) * nChannels + c) * buffSize;
      ...

Buffer layout: each head has2 * nChannelsbuffers (half for reduce, half for broadcast).send[1]andrecv[0]are the reduce direction (UC → MC),recv[1]andsend[0]are the broadcast direction (MC → UC).dataUc.ptris the local UC memory,dataPartition.ptris the MC group mapped address.

Design considerations

[Design inference and architectural trade-offs]

Why does NVLS'scanConnectreturn 0? Because NVLS is not point-to-point transfer—it is a "one-to-many" multicast model.selectTransportThe loop in is designed for point-to-point connections; NVLS connection establishment goes throughncclNvlsSetupan independent path. Putting NVLS intoncclTransportsthe array is only to unifyfreethe interface (nvlsSendFree/nvlsRecvFree), while the actual connection logic is completely independent.

Production pitfall avoidance guide

Pitfall: MNNVL does not support NVLS buffer registration.SeencclNvlsSetup:

At this point, through the ncclTransport abstraction layer, NCCL has successfully unified the four heterogeneous channels P2P, SHM, NET, and NVLS into a consistent interface, and the algorithm kernel does not need to care whether the underlying layer is NVLink or a NIC. But the transport layer only solves "how channels are abstracted"; it has not yet answered "how data is driven asynchronously." In the next chapter, we will focus on src/proxy.cc and src/include/proxy.h to see how the proxy thread asynchronously advances network send/receive on the host side and forms a producer-consumer relationship with the GPU kernel, revealing the key mechanism of NCCL asynchrony.

CHAPTER 12

Chapter 12: Chapter 12: Proxy thread asynchronous scheduling: how proxy.cc decouples I/O from kernel execution

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 12 / 25

Chapter 12: Proxy thread asynchronous scheduling: how proxy.cc decouples I/O from kernel execution

The previous chapter broke down the transport abstraction layer and saw how NCCL uses a unified interface to shield the differences among P2P/SHM/NET/NVLS. But the transport layer only answered "which channel the data takes"; it has not yet answered "how the data is driven asynchronously." If the GPU kernel blocks directly on network waits, the compute units will be dragged down by I/O. This chapter focuses onsrc/proxy.ccandsrc/include/proxy.hto see how NCCL uses an independent host thread to peel network I/O away from the kernel execution path and form a producer-consumer relationship with the GPU.

12.1 Why proxy threads are needed: starting from "who waits for the network"

Intuitive model

Imagine a restaurant: the kitchen (GPU kernel) is only responsible for cooking, and the food runner (proxy thread) is responsible for delivering the dishes to the customers (network peers). If the chef were made to deliver the dishes personally, he would have to stop cooking every time he makes a delivery, and the serving speed would plummet. NCCL's proxy is exactly that dedicated food runner—the kernel only writes data into and reads data from the shared buffer, while all the dirty and tiring work of network send/receive is handed off to the proxy thread on the host side.

[Design inference and architectural trade-offs]

What disaster would the system face without the proxy? The GPU kernel is SIMT massively parallel, and a single warp blocking on network polling would waste the compute power of an entire SM; even more fatally, network send/receive involves socket system calls, verbs polling, and DMA descriptor submission, and these operations simply cannot be executed in device code. Therefore NCCL must move network I/O to the host, letting the kernel and proxy exchange "data ready" signals through a FIFO in shared memory.

Division of labor between the two types of threads

NCCL starts two types of proxy threads on the host side, with completely different responsibilities:

  • Service thread(ncclProxyService): handles control-plane requests—connection establishment, memory registration, FD queries. It listens on a socket, receives RPC requests from the local rank, and asynchronously advances operations such as setup/connect.
  • Progress thread(ncclProxyProgress): handles the data plane—actually driving network send/receive. It takes proxy ops from the shared memory pool and calls the transport'sproxyProgresscallback to advance data movement.

📎 src/include/proxy.h:343-345showsncclProxyStatesimultaneously holdsthread(Service) andthreadUDS(UDS service), while the Progress thread's handle is hidden inprogressState.threadinside📎 src/include/proxy.h:261-261。

Establishment of the producer-consumer relationship

📎 src/proxy.cc:2130-2166'sncclProxyCreateis where the thread is born: whenrefCount == 1(first comm creation), it copies the comm's key fields intoproxyState, then starts the Service thread and the UDS thread. Note that the Progress thread is not started here—it is lazily started byproxyProgressInitonly when a connection that needs proxy progress is established for the first time📎 src/proxy.cc:1523-1524。

mermaid
flowchart TD
    create["ncclProxyCreate(comm)"] --> check_ref{"proxyState->refCount == 1?"}
    check_ref -->|否| skip["复用已有线程,直接返回"]
    check_ref -->|是| copy["拷贝 comm 字段到 proxyState"]
    copy --> start_svc["std::thread(ncclProxyService)"]
    start_svc --> start_uds["std::thread(ncclProxyServiceUDS)"]
    start_uds --> wait["等待连接建立请求"]
    wait --> conn_init{"proxyConnInit 发现<br/>tcomm->proxyProgress != NULL?"}
    conn_init -->|是| prog_init["proxyProgressInit()"]
    conn_init -->|否| no_prog["不启动 Progress 线程"]
    prog_init --> shm["ncclShmOpen 创建 opsPool 共享内存"]
    shm --> start_prog["std::thread(ncclProxyProgress)"]

This diagram anchors the real branch for thread startup: only whentcomm->proxyProgressis non-null (that is, the transport needs data-plane progress) is the Progress thread created.

12.2 Data structures and memory layout: shared memory pool and op pool

Overview of core structures

The proxy's concurrency model is built on two blocks of shared memory, and understanding their memory layout is the prerequisite for understanding the entire mechanism.

First block:ncclProxyOpsPool(📎 src/include/proxy.h:218-226). This is the "task delivery box" between the main thread and the Progress thread, shared across processes through/dev/shm.

FieldTypePurpose
ops[]ncclProxyOp[]Preallocated op array, sizeMAX_OPS_PER_PEER * NCCL_MAX_LOCAL_RANKS
nextOpsvolatile intHead index of the pending op linked list, -1 means empty
nextOpsEndvolatile intTail index of the pending op linked list
freeOps[]volatile int[]Head of the free op linked list for each local rank
syncObjectsInitializedintMarks whether the mutex/cond has been initialized
mutex / condstd::mutex / std::condition_variableCross-process synchronization primitive

MAX_OPS_PER_PEERdefinition of📎 src/include/proxy.h:218-226is2 * MAXCHANNELS * 2 * NCCL_MAX_DEV_WORK_P2P_PER_BATCH. The comment explains why it is 2x: each p2p work contains one send and one recv proxy op, so it must be multiplied by 2; multiplying by 2 again is to be able to store two full rounds of operations, otherwise it would be impossible to "deliver half and release half."

Second block:ncclProxyArgs(📎 src/include/proxy.h:174-209). This is the "runtime op description" used internally by the Progress thread, allocated fromncclProxyPool, and not shared across processes.

Key fields:

  • subs[NCCL_PROXY_MAX_SUBS]: sub-operation array,NCCL_PROXY_MAX_SUBS = MAXCHANNELS 📎 src/include/proxy.h:55-55. Operations of the same type from multiple channels are aggregated into multiple subs of one args.
  • progress: function pointer, pointing to the transport'sproxyProgresscallback📎 src/include/proxy.h:176-176。
  • next / nextPeer / proxyAppendPtr: three linked-list pointers, forming a complex op organization relationship.
  • state:ncclProxyOpNone / ncclProxyOpReady / ncclProxyOpProgressThree-state📎 src/include/proxy.h:48-52。

Layered design of the memory pool

ncclProxyPool 📎 src/proxy.cc:50-53is a batch allocation unit, and each pool containsPROXYARGS_ALLOCATE_SIZE(that is,NCCL_MAX_OPS) ofncclProxyArgs。allocateArgs 📎 src/proxy.cc:207-231The allocation logic is worth a closer look:

c
if (state->pool == NULL) {
    struct ncclProxyPool* newPool;
    NCCLCHECK(ncclCalloc(&newPool, 1));
    struct ncclProxyArgs* newElems = newPool->elems;
    for (int i = 0; i < PROXYARGS_ALLOCATE_SIZE; i++) {
      if (i + 1 < PROXYARGS_ALLOCATE_SIZE) newElems[i].next = newElems + i + 1;
    }
    state->pool = newElems;
    newPool->next = state->pools;
    state->pools = newPool;
}
elem = state->pool;
state->pool = state->pool->next;

📎 src/proxy.cc:207-231

[Design inference and architectural trade-offs]

The design motivation here is:ncclProxyArgsThe structure is very large (containingsubs[MAXCHANNELS]array, and each sub also hasrequests[NCCL_STEPS]). If each op were malloc'ed separately, it would cause severe memory fragmentation and allocation overhead. Batch allocation + free-list reuse amortizes the allocation cost to almost zero. The comment "Make sure we allocate the memory close to the network thread" suggests that this is for NUMA affinity—the pool is created when the Progress thread first allocates, naturally close to the CPU on which that thread runs.

False sharing and atomic variables

ncclProxyOpsPoolinnextOps、nextOpsEnd、freeOps[]are allvolatile int. They are read and written simultaneously by the main thread and the Progress thread, but NCCL does not use locks to protect all accesses—instead it uses atomic operations + memory ordering to ensure correctness.

Look atncclLocalOpAppendthe logic for taking a free op from freeOps📎 src/proxy.cc:503-513:

c
int freeOp = -1;
while (freeOp == -1) {
  freeOp = COMPILER_ATOMIC_EXCHANGE(&pool->freeOps[tpLocalRank], -1, std::memory_order_acquire);
  if (freeOp == -1) std::this_thread::yield();
}

The main thread usesatomic_exchangetofreeOps[tpLocalRank]Set to -1 and retrieve the old value—this is a "preemptive acquisition": whoever succeeds in the exchange first gets the entire free list. When the Progress thread returns an op, it uses a CAS loop📎 src/proxy.cc:898-907:

c
oldFree = COMPILER_ATOMIC_LOAD(&pool->freeOps[i], std::memory_order_acquire);
do {
  pool->ops[freeOpEnd[i]].next = oldFree;
} while (!COMPILER_ATOMIC_COMPARE_EXCHANGE(&pool->freeOps[i], &oldFree, newFree,
                                           std::memory_order_release,
                                           std::memory_order_acquire));
[Design inference and architectural trade-offs]

Acquire/release is used here instead of seq_cst because it only needs to ensure that "the write to the linked list node's next pointer" is visible to the acquiring side, and does not require global ordering.freeOps[]Each element of the array corresponds to a local rank, naturally distributed near different cache lines, reducing false sharing.

12.3 Control plane: connection establishment and RPC mechanism

Intuitive model

[Design inference and architectural trade-offs]

The Service thread is like a "front desk receptionist": when a local rank wants to establish a network connection, it does not connect directly itself, but sends an RPC request to the Service thread, which performs setup/connect on its behalf. Why do it this way? Because network connection establishment (especially verbs QP creation and memory registration) may block, and certain resources (such as the listen socket) must be held by a single thread. By centralizing the control plane in the Service thread, the main thread can continue doing other things without blocking.

Encoding of RPC requests

ncclProxyCallAsync 📎 src/proxy.cc:1369-1394It is the sender side of the RPC. It sends sequentially over the socket: type, connection pointer, reqSize, respSize, reqBuff, opId.

c
NCCLCHECKGOTO(ncclSocketSend(sock, &type, sizeof(int)), ret, error);
NCCLCHECKGOTO(ncclSocketSend(sock, &proxyConn->connection, sizeof(void*)), ret, error);
NCCLCHECKGOTO(ncclSocketSend(sock, &reqSize, sizeof(int)), ret, error);
NCCLCHECKGOTO(ncclSocketSend(sock, &respSize, sizeof(int)), ret, error);
if (reqSize) NCCLCHECKGOTO(ncclSocketSend(sock, reqBuff, reqSize), ret, error);
NCCLCHECKGOTO(ncclSocketSend(sock, &opId, sizeof(opId)), ret, error);
NCCLCHECK(expectedProxyResponseEnqueue(sharedProxyState, opId, respSize));

📎 src/proxy.cc:1369-1394

Note the last step: after sending the request, immediately register the opId with theexpectedResponsesqueue. This is the key to asynchronous RPC—the caller does not wait for a reply, but first registers "I expect a response for this opId," and afterward usesncclPollProxyResponsepolling.

Linked-list implementation of the response queue

expectedProxyResponseEnqueue 📎 src/proxy.cc:97-117It uses a singly linked list to store ops awaiting responses.expectedProxyResponseStore 📎 src/proxy.cc:67-95When a response is received, it matches by opId, memcpy's the response data into the preallocatedrespBuff, and marksdone = true。expectedProxyResponseDequeue 📎 src/proxy.cc:119-141During polling, it looks up completed responses and removes them.

There is a detail here:expectedProxyResponseStoreCheckrespSizewhether it matches📎 src/proxy.cc:72-75, and if it does not match, reportncclInternalError. This is defensive programming—if the requester and responder have inconsistent understandings of the response size, it means the protocol is corrupted, and it must fail immediately rather than silently continue.

Main loop of the Service thread

ncclProxyService 📎 src/proxy.cc:1789-2016The core is a poll loop. It usespollfdsan array to manage all connections, including the listen socket and each peer's socket.

c
while (stop == PROXY_RUNNING || npeers > 0) {
    if (COMPILER_ATOMIC_LOAD(proxyState->abortFlag, std::memory_order_acquire) != 0) stop = PROXY_ABORT;
    int ret = 0;
    const int timeout = asyncOpCount ? 0 : 500;
    ...
    ret = poll(activePollfds, nfds_to_poll, timeout);

📎 src/proxy.cc:1842-1863

timeoutThe choice of is very particular: if there is an asynchronous op in progress (asyncOpCount > 0), timeout is set to 0 (non-blocking polling), because it needs to callproxyProgressAsyncfrequently to advance them; otherwise it is set to 500ms to avoid spinning and burning CPU. The comment "never let proxy service thread blocks in poll, or it cannot receive abortFlag"📎 src/proxy.cc:1847-1847clarifies why it cannot block indefinitely—it must periodically wake up to check abortFlag.

Advancement of asynchronous ops

proxyProgressAsync 📎 src/proxy.cc:1626-1700It is the core of the Service thread advancing asynchronous operations. It dispatches to different transport callbacks according to the op type:

c
if (op->type == ncclProxyMsgSetup) {
    res = op->connection->tcomm->proxySetup(op->connection, proxyState, op->reqBuff, op->reqSize, op->respBuff,
                                            op->respSize, &done);
} else if (op->type == ncclProxyMsgConnect) {
    res = op->connection->tcomm->proxyConnect(...);
} else if (op->type == ncclProxyMsgInit) {
    res = proxyConnInit(peer, connectionPool, proxyState, ...);
}

📎 src/proxy.cc:1631-1664

Each callback carries andoneoutput parameter. Ifdone == 0, it means the operation is not yet complete (for example, the network connection is still in the three-way handshake), and it returnsncclInProgress, and the next loop continues advancing. Ifdone == 1, then it sends the response header + response body to the requester📎 src/proxy.cc:1681-1689。

mermaid
sequenceDiagram
    participant Main as 主线程 (ncclSend)
    participant Svc as Service 线程
    participant Net as 网络插件 (ncclNet)
    Main->>Svc: ncclProxyCallAsync(ncclProxyMsgConnect)
    Note over Main: expectedProxyResponseEnqueue(opId)
    Svc->>Svc: proxyServiceInitOp 读取请求
    Svc->>Net: proxyConnect() 调用 ncclNet->connect
    alt connect 未完成
        Net-->>Svc: netSendComm == NULL, done=0
        Svc->>Svc: 返回 ncclInProgress,下次 poll 重试
    else connect 完成
        Net-->>Svc: netSendComm != NULL, done=1
        Svc->>Main: ncclSocketSend(resp header + connectMap)
    end
    Main->>Main: ncclPollProxyResponse 轮询
    Main->>Main: expectedProxyResponseDequeue 取回结果

This sequence diagram anchorssendProxyConnectin*done = 0; return ncclInProgressthe real branch📎 src/transport/net.cc:913-916。

12.4 Data plane: how the Progress thread drives network send/receive

Intuitive model

The Progress thread is a "conveyor belt operator": it watches the FIFO in the shared buffer, and as soon as the GPU has written the data (size != -1 in the FIFO), it immediately callsisendto send the data out; once the network has finished receiving data, it updates recvTail to notify the GPU that it can read. Throughout the process, the GPU and proxy synchronize through the head/tail pointers in the FIFO, without needing any locks.

Op submission: from the main thread to the Progress thread

The main thread, inncclProxySaveOp 📎 src/proxy.cc:591-761, decides which proxy ops are needed according to the pattern, and then usesSaveProxy → ncclLocalOpAppendto write the op into the shared memory pool.

ncclLocalOpAppend 📎 src/proxy.cc:488-554The flow of is:

1. FromproxyOps->freeOporpool->freeOps[tpLocalRank]take a free op slot.

2. memcpy(op, proxyOp, sizeof(struct ncclProxyOp))Copy the op contents into shared memory📎 src/proxy.cc:515-515。

3. Attach the op to the tail of theproxyOps->nextOpslinked list.

4. If the accumulated number of ops reachesMAX_OPS_PER_PEER, trigger a batch submission📎 src/proxy.cc:525-551。

The logic of batch submission is very subtle: it cannot simply send all ops out, because "multiple ops with the same opCount must be submitted together, otherwise it will break the sub-aggregation of proxyArgs." So it finds the boundary of the last opCount change and submits only up to there📎 src/proxy.cc:529-548。

Submission is completed throughncclProxyPost 📎 src/proxy.cc:476-486, which locks, updatespool->nextOps、notify_oneand wakes up the Progress thread.

Main loop of the Progress thread

ncclProxyProgress 📎 src/proxy.cc:951-1011The structure of is:

c
do {
    int idle = 1;
    ncclResult_t ret = progressOps(proxyState, state, state->active, &idle);
    ...
    if (idle || !state->active || (++proxyOpAppendCounter == ncclParamProgressAppendOpFreq())) {
      int added = 0;
      proxyOpAppendCounter = 0;
      ret = ncclProxyGetPostedOps(proxyState, &added);
      ...
    }
    lastIdle = idle;
    stopv = state->stop.load(std::memory_order_acquire);
} while ((stopv == 0 || (stopv == 1 && state->active)) &&
         COMPILER_ATOMIC_LOAD(proxyState->abortFlag, std::memory_order_acquire) == 0);

📎 src/proxy.cc:976-1009

There is a performance optimization worth noting here:proxyOpAppendCountercounter📎 src/proxy.cc:974-974. The comment explains📎 src/proxy.cc:969-973: callingncclProxyGetPostedOpstoo frequently will cause performance regression in small-message communication, so every time it advancesProgressAppendOpFreq(default 8) times before fetching a new op.

Op aggregation: ProxyAppend

ProxyAppend 📎 src/proxy.cc:437-474Determines whether an op is "appended to the sub of an existing args" or "creates a new args". The criterion isconnection->shared && args->opCount == op->opCount 📎 src/proxy.cc:443-443— multiple channel operations with the same connection and same opCount are aggregated.

[Design inference and architectural trade-offs]

Value of aggregation: Similar operations from multiple channels are merged into one args, so the Progress thread can advance all channels in a single loop iteration, reducing function call overhead and cache invalidation.ncclProxyOpToArgs 📎 src/proxy.cc:368-435When appending a sub, it validatessliceSteps、chunkSteps、protocol、dtype、redOp、collwhether they are consistent📎 src/proxy.cc:401-406, and reports an error if not — this is the defense against incorrect aggregation.

sendProxyProgress: the four-stage state machine on the send side

sendProxyProgress 📎 src/transport/net.cc:1324-1491It is the core of the send side. It advances sub by sub, and each sub has four counters:posted、transmitted、done。

Stage 1: Ready initialization 📎 src/transport/net.cc:1326-1339

c
sub->base = ROUNDUP(resources->step, args->chunkSteps);
resources->step = sub->base + sub->nsteps;
sub->posted = sub->transmitted = sub->done = 0;

baseis the starting number of the step,ROUNDUPensuring alignment tochunkSteps。resources->stepaccumulation, reserving space for the next op.

Stage 2: Post the buffer to the GPU 📎 src/transport/net.cc:1355-1376

c
if (sub->posted < sub->nsteps && sub->posted < sub->done + maxDepth) {
    int buffSlot = (sub->base + sub->posted) % NCCL_STEPS;
    if (resources->shared) {
        ...
        *sendHead = sub->base + sub->posted - NCCL_STEPS;
    } else {
        sub->posted += args->sliceSteps;
    }
}

maxDepthis the pipeline depth📎 src/transport/net.cc:1343-1343, limiting the number of simultaneously in-flight steps. In shared mode, the proxy tells the GPU "this slot can be written" by updatingsendHead.

Stage 3: Check whether the GPU has finished writing, and initiate isend 📎 src/transport/net.cc:1378-1452

c
if (sub->transmitted < sub->posted && sub->transmitted < sub->done + NCCL_STEPS) {
    int buffSlot = (sub->base + sub->transmitted) % NCCL_STEPS;
    volatile uint64_t* recvTail = &resources->recvMem->tail;
    uint64_t tail = sub->base + sub->transmitted;
    if (connFifo[buffSlot].size != -1 && (*recvTail > tail || p == NCCL_PROTO_LL)) {
        int size = connFifo[buffSlot].size;
        ...
        NCCLCHECK(proxyState->ncclNet->isend(resources->netSendComm, buff, size, resources->tpRank,
                                             sub->sendMhandle, phandle, sub->requests + buffSlot));
        if (sub->requests[buffSlot] != NULL) {
            sub->transmitted += args->sliceSteps;
        }
    }
}

The key condition here isconnFifo[buffSlot].size != -1 && *recvTail > tail— after the GPU finishes writing data, it updates the FIFO size and recvTail, and the proxy initiates isend only after seeing both conditions satisfied. For the LL protocol, because it has "zero-copy" semantics, there is no need to wait for recvTail.

Stage 4: Check whether sending is complete, and update sendHead 📎 src/transport/net.cc:1455-1481

c
if (sub->done < sub->transmitted) {
    int buffSlot = (sub->base + sub->done) % NCCL_STEPS;
    NCCLCHECK(proxyState->ncclNet->test(sub->requests[buffSlot], &done, &size));
    if (done) {
        connFifo[buffSlot].size = -1;
        std::atomic_thread_fence(std::memory_order_seq_cst);
        sub->done += args->sliceSteps;
        if (resources->shared == 0) {
            volatile uint64_t* sendHead = resources->gdcSync ? resources->gdcSync : &resources->sendMem->head;
            *sendHead = sub->base + sub->done;
        }
    }
}

testAfter returns done, first reset the FIFO size to -1, insert a seq_cst fence, and then update sendHead to notify the GPU that "this slot can be reused". The purpose of the fence is to prevent reordering of the size reset and the head update — if head is updated first, the GPU may start writing while size is still the old value.

recvProxyProgress: the four stages on the receive side

recvProxyProgress 📎 src/transport/net.cc:1493-1788It is more complex because it involves sub grouping (multirecv is used when multiple subs share the same recvComm).

Stage 1: Group by recvComm during Ready 📎 src/transport/net.cc:1495-1538

c
for (int s = 0; s < args->nsubs; s++) {
    ...
    if (groupSize == maxRecvs) {
        groupSize = 0;
    } else if (s > 0) {
        int next;
        for (next = s; next < args->nsubs; next++) {
            struct recvNetResources* nextRes = ...;
            if (nextRes->netRecvComm == recvComm) break;
        }
        if (next == args->nsubs) {
            groupSize = 0;
        } else if (s != next) {
            // swap subs
        }
    }
    groupSize++;
    ...
    for (int i = 0; i < groupSize; i++) sub[-i].groupSize = groupSize;
}
[Design inference and architectural trade-offs]

This code groups subs that use the samerecvCommtogether and recordsgroupSize. Why group? Becauseirecvsupports receiving multiple buffers at once (multirecv), and merging requests for the same comm into one call can significantly reduce plugin overhead.

Stage 2: Initiate irecv 📎 src/transport/net.cc:1543-1631

c
if (subCount) {
    uint64_t step = subGroup->posted;
    void** requestPtr = subGroup->requests + (step % NCCL_STEPS);
    bool ignoreCompletion = ncclParamNetOptionalRecvCompletion() &&
                            ((args->protocol == NCCL_PROTO_LL128) || (args->protocol == NCCL_PROTO_LL)) &&
                            (subCount == 1);
    if (ignoreCompletion) *requestPtr = (void*)NCCL_NET_OPTIONAL_RECV_COMPLETION;
    NCCLCHECK(proxyState->ncclNet->irecv(resources->netRecvComm, subCount, ptrs, sizes, tags, mhandles, phandles,
                                         requestPtr));
    if (*requestPtr) {
        subGroup->recvRequestsCache[step % NCCL_STEPS] = *requestPtr;
        subGroup->recvRequestsSubCount = subCount;
        for (int i = 0; i < subGroup->groupSize; i++) {
            sub->posted += args->sliceSteps;
        }
    }
}

ignoreCompletionOptimization📎 src/transport/net.cc:1608-1610: For single-buffer receives in the LL/LL128 protocols, completion notification is optional (because the data itself carries a flag), so the completion check can be skipped.

Stage 3: Check whether receiving is complete, and update recvTail 📎 src/transport/net.cc:1634-1743

c
NCCLCHECK(proxyState->ncclNet->test(subGroup->requests[step % NCCL_STEPS], &done, sizes));
if (done) {
    for (int i = 0; i < subGroup->groupSize; i++) {
        struct ncclProxySubArgs* sub = subGroup + i;
        int buffSlot = (sub->base + sub->received) % NCCL_STEPS;
        connFifo[buffSlot].size = -1;
        sub->received += args->sliceSteps;
    }
    ...
}

After receiving is complete, reset the FIFO size, then enter the flush stage (the GDRDMA scenario requires flush to ensure data visibility).

Stage 4: Wait for the GPU to consume, and update done 📎 src/transport/net.cc:1745-1779

c
if (sub->transmitted > sub->done) {
    volatile uint64_t* sendHead = &resources->sendMem->head;
    uint64_t done = *sendHead;
    while (done > sub->base + sub->done && sub->transmitted > sub->done) {
        if (subGroup->recvRequestsCache[sub->done % NCCL_STEPS]) {
            if (proxyState->ncclNet->irecvConsumed) {
                NCCLCHECK(proxyState->ncclNet->irecvConsumed(resources->netRecvComm, subGroup->recvRequestsSubCount,
                                                             subGroup->recvRequestsCache[sub->done % NCCL_STEPS]));
            }
            subGroup->recvRequestsCache[sub->done % NCCL_STEPS] = NULL;
        }
        sub->done += args->sliceSteps;
    }
}

Here it readssendHeadto determine whether the GPU has already consumed the data.irecvConsumedIt is a callback to the plugin, telling it that "the buffer for this receive request has been consumed and can be reused".

Overview of the data flow

mermaid
flowchart LR
    subgraph GPU["GPU Kernel"]
        gpu_write["写入数据到 buff"]
        gpu_fifo["更新 connFifo.size<br/>和 recvTail"]
    end
    subgraph SHM["共享内存 FIFO"]
        fifo["ncclConnFifo<br/>size / offset"]
        head["sendMem->head"]
        tail["recvMem->tail"]
    end
    subgraph PROXY["Progress 线程"]
        check["检查 size != -1<br/>且 recvTail > tail"]
        isend["ncclNet->isend()"]
        test["ncclNet->test()"]
        update["更新 sendHead"]
    end
    gpu_write --> gpu_fifo
    gpu_fifo --> fifo
    gpu_fifo --> tail
    fifo --> check
    tail --> check
    check -->|数据就绪| isend
    isend --> test
    test -->|发送完成| update
    update --> head
    head -->|GPU 可复用 slot| gpu_write

This data flow diagram shows the closed loop formed by the GPU and proxy through the FIFO and the head/tail pointers: GPU writes data → updates tail → proxy detects it and initiates isend → test confirms completion → updates head → GPU reuses the slot.

12.5 Concurrency control, memory barriers, and hardware interaction

Memory ordering of the lock-free FIFO

Synchronization between the proxy and the GPU relies entirely onncclConnFifoand the head/tail pointers, without any locks. This requires extremely careful memory ordering control.

On the send side, aftertestreturns done, the proxy📎 src/transport/net.cc:1460-1473:

c
connFifo[buffSlot].size = -1;
std::atomic_thread_fence(std::memory_order_seq_cst);
...
*sendHead = sub->base + sub->done;

The seq_cst fence ensures that only after the size reset is visible to the GPU does the head update become visible. If the order were reversed, the GPU might see the new head but the old size, mistakenly assuming there is data in the slot.

On the receive side, before updating recvTail, the proxy📎 src/transport/net.cc:1731-1736:

c
if (step < sub->nsteps) {
    std::atomic_thread_fence(std::memory_order_seq_cst);
    volatile uint64_t* recvTail = resources->gdcSync ? resources->gdcSync : &resources->recvMem->tail;
    *recvTail = sub->base + sub->transmitted;
}

The same principle applies: first fence to ensure the data write is visible, then update tail to notify the GPU that it can read.

GDRCOPY's flush mechanism

When GDRDMA is used, the NIC writes directly to GPU memory, but the write may still be uncommitted on the PCIe bus. The proxy needs to actively flush to ensure data visibility. SeerecvProxyProgressthe flush logic in📎 src/transport/net.cc:1664-1709:

c
if (totalSize > 0 && p == NCCL_PROTO_SIMPLE && needFlush) {
    if (resources->gdcFlush) {
#if defined(__x86_64__)
        asm volatile("mfence" ::: "memory");
        asm volatile("mov (%0), %%eax" ::"l"(resources->gdcFlush) : "%eax", "memory");
#else
        std::atomic_thread_fence(std::memory_order_seq_cst);
        uint64_t dummy;
        NCCLCHECK(ncclGdrCudaRead(resources->gdrDesc, &dummy, resources->gdcFlush, sizeof(dummy)));
#endif
    } else {
        // iflush 路径
        NCCLCHECK(proxyState->ncclNet->iflush(resources->netRecvComm, subCount, ptrs, sizes, mhandles,
                                              subGroup->requests + (step % NCCL_STEPS)));
    }
}

The comments on the x86 path are excellent.📎 src/transport/net.cc:1668-1674:mfencePrevent the CQE-poll load from being reordered before the flush load;mov (%0), %%eaxForce a PCIe read, making the CPU stall until all prior PCIe posted writes (including NIC DMA) are committed to the endpoint. This is hardware-level memory ordering control, more hardcore than any software fence.

Coordination between atomic variables and stop/abort

Exit conditions of the Progress thread📎 src/proxy.cc:1007-1009:

c
stopv = state->stop.load(std::memory_order_acquire);
} while ((stopv == 0 || (stopv == 1 && state->active)) &&
         COMPILER_ATOMIC_LOAD(proxyState->abortFlag, std::memory_order_acquire) == 0);

stop == 1Butstate->active != NULLcontinue running during — this is for "graceful stop": already-posted ops must be fully advanced, otherwise the GPU will wait forever for data. Onlystop == 2(abort) orabortFlag != 0forces exit.

ncclProxyProgressDestroy 📎 src/proxy.cc:1039-1065The stop procedure of:

c
std::lock_guard<std::mutex> lock(state->opsPool->mutex);
state->stop.store(1, std::memory_order_release);
state->opsPool->cond.notify_one();
state->thread.join();

Lock first, then store stop, then notify — this is the standard pattern to prevent lost wakeup. The Progress thread holds the lock duringpool->cond.waitand checks the predicate📎 src/proxy.cc:850-851, ensuring it won't miss the wakeup.

12.6 Production Pitfall Guide and Failure Recovery Chain

Pitfall 1: Connection leak prevents the Service thread from exiting

ncclProxyServiceThe main loop condition of isstop == PROXY_RUNNING || npeers > 0 📎 src/proxy.cc:1842-1842. The comment explains📎 src/proxy.cc:1843-1845: even if the local comm aborts, as long as there are still peer connections, the proxy thread cannot exit, otherwise it may segfault.

Troubleshooting scenario: if a rank crashes without notifying the peer, the peer's Service thread will be stuck forever in the loop ofnpeers > 0. In this case, you need to rely onabortFlagor a timeout mechanism. In production, if you see a process hanging atncclProxyService, first check whether a peer rank exited abnormally.

Pitfall 2: Response queue mismatch causes memory leak

expectedProxyResponseStorereturns when opId doesn't matchncclInternalError 📎 src/proxy.cc:93-94. But if the requester has already given up by the time the response arrives (e.g., timeout), this response will remain in the queue forever,respBuffleak.

Defensive measures:expectedProxyResponseFree 📎 src/proxy.cc:55-65cleans up the entire queue atncclProxyDestroy. But this is the last resort; under normal operation there should be no residue.📎 src/proxy.cc:2226-2226Pitfall 3: head initialized to a negative value in shared mode

In

sendProxyConnectCopy📎 src/transport/net.cc:999-1000:

c
// Don't give credits yet in shared mode.
(resources->gdcSync ? *resources->gdcSync : resources->sendMem->head) = (map->shared ? -NCCL_STEPS : 0);

, meaning the GPU initially has no credit to write. The proxy needs to gradually increase head during the post phase to "grant credit." If this initialization is forgotten, the GPU will mistakenly think it has credit and write to slots that aren't ready, causing data corruption.-NCCL_STEPSPitfall 4: flag validation in the LL128 protocol

In

sendProxyProgressCopy📎 src/transport/net.cc:1388-1403:

c
if (p == NCCL_PROTO_LL128) {
    ready = resources->useGdr;
    if (!ready) {
        uint64_t flag = sub->base + sub->transmitted + 1;
        int nFifoLines = DIVUP(connFifo[buffSlot].size, sizeof(uint64_t) * NCCL_LL128_LINEELEMS);
        volatile uint64_t* lines = (volatile uint64_t*)buff;
        ready = 1;
        for (int i = 0; i < nFifoLines; i++) {
            if (lines[i * NCCL_LL128_LINEELEMS + NCCL_LL128_DATAELEMS] != flag) {
                ready = 0;
                break;
            }
        }
    }
}

, and the proxy must check the flag line by line to confirm data integrity. If this check is skipped and isend is called directly, half-baked data may be sent. This is a pitfall unique to LL128.threadfence()Failure recovery chain

When

returns non-proxyProgressAsync,ncclSuccess/ncclInProgressthe Service thread closes the connection and cleans up all async ops for that peer📎 src/proxy.cc:1929-1937. This cleanup is a "full drain" — it doesn't just clean up the failed op, but empties the entire peer's asyncOps queue, preventing residual ops from referencing an already-freed connection.📎 src/proxy.cc:1984-1995When the Progress thread encounters an error

, it writes the error code to📎 src/proxy.cc:979-983and exits the loop. The main thread can later detect the error by checking this field.proxyState->asyncResultChapter Summary

In this chapter, we dissected the complete mechanism of NCCL proxy threads:

Division of labor between two thread types

1. : the Service thread handles control-plane RPCs (connection establishment, memory registration), and the Progress thread handles the data plane (advancing network send/receive).Shared memory pool

2. passes ops across processes,:ncclProxyOpsPoolaggregates operations from multiple channels within the Progress thread.ncclProxyArgsLock-free FIFO synchronization

3. : the GPU and proxy exchange data-ready signals viaand head/tail pointers, using seq_cst fences to guarantee memory ordering.connFifoFour-phase state machine

4. : send/recv each have their own posted → transmitted → received → done counters driving the pipeline.Hardware-level flush

5. : in GDRDMA scenarios, use+ PCIe read to force commit of posted writes.mfenceChapter Review Questions

Q1: If you remove the logic in

that updatessendProxyProgresswhensub->done == sub->nsteps(i.e., not notifying the GPU that the slot has been released), in what scenario would a deadlock be triggered? Why?sendHeadReference analysis

is the sole basis for the GPU to determine "which slots can be reused." See:sendHeadCopy📎 src/transport/net.cc:1469-1473:

c
if (resources->shared == 0) {
    volatile uint64_t* sendHead = resources->gdcSync ? resources->gdcSync : &resources->sendMem->head;
    *sendHead = sub->base + sub->done;
}

in shared mode, 0 in non-shared mode). The GPU kernel checks-NCCL_STEPSatwaitSendbefore considering there is credit to write. If head doesn't advance, the GPU will block forever waiting for credit after fillinghead + NCCL_STEPS > stepslots, while the proxy is waiting for the GPU to write new data before it can isend — a classic producer-consumer deadlock. It's even worse in shared mode, because the initial head is negative, so the GPU has no credit from the start.NCCL_STEPS 个 slot 后就永远阻塞在等待 credit 上,而 proxy 又在等 GPU 写新数据才能 isend——经典的生产者-消费者死锁。在 shared 模式下更严重,因为初始 head 是负值,GPU 一开始就没有 credit。

Q2: ncclLocalOpAppendWhen the cumulative op reachesMAX_OPS_PER_PEERit triggers batch delivery, but the code deliberately "does not deliver all ops of the last opCount". If it were changed to simply deliver all ops, what mechanism would be broken?

Reference analysis: Look at📎 src/proxy.cc:525-548's comments and logic:

c
// Do not post last operations as we could have more coming with the same opCount, and posting
// them in different batches would break proxyArgs aggregation with subs.
uint64_t lastOpCount = pool->ops[proxyOps->nextOpsEnd].opCount;
int lastOp = -1;
...
for (int op = proxyOps->nextOps; op != proxyOps->nextOpsEnd; op = pool->ops[op].next) {
    ops++;
    if (pool->ops[op].opCount != lastOpCount) {
        lastOp = op;
        toSend = ops;
    }
}

ProxyAppend's aggregation logic📎 src/proxy.cc:443-443depends onargs->opCount == op->opCountto determine whether to append a sub. If multiple channel ops of the same opCount are split across two batches for delivery, the first batch creates an args, and when the second batch arrives,args->opCountis already not equal to the new op's opCount (because args may have already been advanced), causing subs that should have been aggregated to be split into independent args. This not only reduces performance, but may also breakncclProxyOpToArgs'snChannels/nPeersmin-taking logic📎 src/proxy.cc:399-400, leading to incorrect channel count calculation.

Q3: recvProxyProgress's Ready phase will regroup and reorder subs byrecvComm. If this grouping logic is removed and each sub independently callsirecv, what consequences would there be onmaxRecvs > 1's NIC?

Reference analysis: Look at📎 src/transport/net.cc:1495-1538's grouping logic and📎 src/transport/net.cc:1613-1614's multirecv call:

c
NCCLCHECK(proxyState->ncclNet->irecv(resources->netRecvComm, subCount, ptrs, sizes, tags, mhandles, phandles,
                                     requestPtr));

maxRecvsis the "maximum number of buffers a single irecv can receive" declared by the NIC plugin📎 src/transport/net.cc:1525-1525. WhenmaxRecvs > 1, plugins such as IB support receiving multiple buffers with one WQE, which can significantly reduce doorbell overhead and CQE processing cost. If grouping is removed and each sub is irecv'd independently,subCountis always 1, the plugin degrades to single-buffer mode, and throughput will drop. More critically,recvRequestsCacheandirecvConsumedmechanisms📎 src/transport/net.cc:1616-1617are designed for multirecv—under single-buffer mode these caching logics become ineffective and may cause request leaks.

At this point, we understand how the proxy thread decouples network I/O from kernel execution, allowing GPU computation and communication to truly run in parallel. But the proxy is only the driver; the concrete implementation of the underlying network transport remains to be revealed. In the next chapter we will go deep intonet_ibto see how NCCL encapsulates the verbs API to implement InfiniBand transport, and how GPUDirect RDMA allows the NIC to directly read and write GPU memory.

CHAPTER 13

Chapter 13: Chapter 13: InfiniBand Network Transport: How net_ib Encapsulates verbs and GPUDirect RDMA

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 13 / 25

Chapter 13: InfiniBand Network Transport: How net_ib Encapsulates verbs and GPUDirect RDMA

In the previous chapter we saw how the proxy thread strips network I/O out of the GPU kernel, allowing computation and communication to truly run in parallel. But the proxy is only a "driver"—it calls abstract interfaces such as ncclNet->isend/irecv, yet does not know whether the underlying transport is TCP, InfiniBand, or something else. In this chapter we lift this layer of abstraction and enter src/transport/net_ib and src/misc/ibvwrap.cc to see how NCCL encapsulates the libibverbs C library into a pluggable symbol table, how it establishes a Queue Pair (QP), and how GPUDirect RDMA allows the NIC to bypass host memory and directly read and write GPU memory.

13.1 Why NCCL Does Not Directly Call libibverbs

Intuitive model: the symbol table is a "pluggable power socket"

Imagine you bought an imported appliance, and the plug shape does not match your home socket. You have two choices: either take the appliance apart and rewire it (directly#include <infiniband/verbs.h>and link-libverbs), or buy a universal adapter plug (dynamically load symbols at runtime). NCCL chose the latter.

[Design inference and architectural trade-offs]

The core motivation for this choice isdeployment flexibility: as a library loaded by upper-layer frameworks such as PyTorch and TensorFlow, NCCL cannot assume that the runtime environment definitely haslibibverbs.soinstalled. If it were hard-linked at compile time, then on machines without an InfiniBand driver, the entire NCCL library could not be loaded—even if you only wanted to use NVLink for single-machine communication. Through runtimedlopen+ symbol resolution, NCCL can gracefully degrade on machines without IB.

If this layer of encapsulation were missing, the disaster the system would face is:a pure NVLink single-machine training job would crash directly because the machine does not have an IB driver installed. This is extremely common in cloud environments and on development machines.

Data structures and memory layout: symbol table container

The core data structure isncclIbvSymbols, defined inibvsymbols.h(this chapter's material does not include that file, but its structure can be inferred from usage). It is a pure function pointer container, with each field corresponding to a libibverbs function:

c
struct ncclIbvSymbols {
  int (*ibv_internal_fork_init)(void);
  struct ibv_device** (*ibv_internal_get_device_list)(int* num_devices);
  int (*ibv_internal_modify_qp)(struct ibv_qp*, struct ibv_qp_attr*, int);
  // ... 数十个函数指针
};

There is only one global instance, together withstd::once_flagensuring thread-safe initialization:

📎 src/misc/ibvwrap.cc:26-29

c
static std::once_flag initOnceFlag;
static ncclResult_t initResult;
struct ncclIbvSymbols ibvSymbols;

The design here is very restrained:initOnceFlagisstd::once_flag,initResultCache initialization result,ibvSymbolsis the global symbol table. All three have static storage duration, with lifetimes spanning the entire process.

[Design Inference and Architectural Trade-offs]

Why usestd::once_flaginstead ofpthread_once? Because NCCL's C++ code already depends on<mutex>and<thread>, using the standard library is more consistent.call_onceThe semantics of are: no matter how many threads callwrap_ibv_symbols()simultaneously, the lambda executes only once, the remaining threads block and wait, then all receive the sameinitResult. This is far safer than hand-written double-checked locking (DCLP)—DCLP has a well-known reordering pitfall under the C++ memory model.

Step-by-Step: The Complete Symbol Resolution Flow

When NCCL first needs IB transport, it callswrap_ibv_symbols():

📎 src/misc/ibvwrap.cc:26-29

c
ncclResult_t wrap_ibv_symbols(void) {
  std::call_once(initOnceFlag, []() { initResult = buildIbvSymbols(&ibvSymbols); });
  return initResult;
}

buildIbvSymbolsis defined inibvsymbols.cc(not included in this chapter), its job is to usedlopen("libibverbs.so")to open the library, then for each function name calldlsymto fill in pointers. If a symbol is not found, the corresponding field remains NULL.

This "allow NULL" design runs through the entire wrapper layer. Look atCHECK_NOT_NULLmacro:

📎 src/misc/ibvwrap.cc:26-29

c
#define CHECK_NOT_NULL(container, internal_name) \
  if (container.internal_name == NULL) { \
    WARN("lib wrapper not initialized."); \
    return ncclInternalError; \
  }

Each wrapper function checks whether the corresponding symbol is non-null before calling. This means:If an older version of libibverbs lacks a certain new function, NCCL won't crash at load time, but will report an error only when that function is actually used. This is the key to graceful degradation.

Design Thinking: The Triple Responsibility of Macro Wrappers

ibvwrap.ccdefines 7 macros, which are not simple syntactic sugar but carry three responsibilities:

1. Null pointer protection:CHECK_NOT_NULLintercepts uninitialized

2. Error code normalization: translating libibverbs' various error conventions (returning -1, returning errno, returning NULL pointer) uniformly intoncclResult_t

3. Logging instrumentation: on failureWARNprints the function name and errno

Look atIBV_PTR_CHECK_ERRNOthis most complex macro:

📎 src/misc/ibvwrap.cc:38-45

c
#define IBV_PTR_CHECK_ERRNO(container, internal_name, call, retval, error_retval, name) \
  CHECK_NOT_NULL(container, internal_name); \
  retval = container.call; \
  if (retval == error_retval) { \
    WARN("Call to " name " failed with error %s", strerror(errno)); \
    return ncclSystemError; \
  } \
  return ncclSuccess;

After expansion it does four things: check symbol is non-null, execute the call, write the return value intoretval(typically returned via pointer parameter such asibv_pd*etc.), determine whether it equals the error value. Notestrerror(errno)—libibverbs' pointer-returning functions (such asibv_alloc_pd) return NULL on failure and seterrno, so readingerrnohere is correct.

WhileIBV_INT_CHECKis used for functions returning int:

📎 src/misc/ibvwrap.cc:84-91

c
#define IBV_INT_CHECK(container, internal_name, call, error_retval, name) \
  CHECK_NOT_NULL(container, internal_name); \
  int ret = container.call; \
  if (ret == error_retval) { \
    WARN("Call to " name " failed"); \
    return ncclSystemError; \
  } \
  return ncclSuccess;

Here does not readerrno, because such functions (such asibv_fork_init) directly return -1 to indicate failure, and the error information is already lost.

[Design Inference and Architectural Trade-offs]

This approach of "using a different macro for each function" looks cumbersome, but it is necessary: libibverbs' API error conventions are extremely inconsistent—some return 0/-1, some return errno values, some return pointers. Forcing uniformity would instead lose error information. NCCL chooses to "translate faithfully," keeping complexity in the wrapper layer, so that the upper layernet_ib.cconly needs to checkncclSuccess。

13.2 ibvcore.h: ABI Contract Without Header Dependencies

Intuitive Model: A Translator with Its Own Dictionary

ibvcore.his a peculiar file—it redefines libibverbs' core structs, enums, and constantsfrom scratch. Why? Because NCCL needs to use these types without#include <infiniband/verbs.h>.

[Design Inference and Architectural Trade-offs]

This solves a real engineering problem:infiniband/verbs.hhas different contents across distributions and driver versions. If NCCL directly included it, it would be bound to a specific version at compile time. By defining its own "minimal necessary subset," NCCL can avoid needing IB headers at compile time and load any version of the library at runtime viadlopen.

Without this layer, the disaster would be:Cannot compile NCCL on machines withoutlibibverbs-devinstalled. Yet in reality the library file might be provided at runtime viardma-core.

Memory Layout of Key Structs

Let's dissect a few structs most critical to understanding RDMA.

ibv_gid: Global Identifier

📎 src/include/ibvcore.h:58-64

c
union ibv_gid {
	uint8_t			raw[16];
	struct {
		uint64_t	subnet_prefix;
		uint64_t	interface_id;
	} global;
};

GID is InfiniBand's "IP address," 16 bytes. It can be accessed both as a 16-byte array and as two 64-bit integers. In RoCE (RDMA over Converged Ethernet) scenarios, the GID is actually an IPv6 address—which is also whyibvGetGidStrusesinet_ntop(AF_INET6, ...)to format it:

📎 src/include/ibvwrap.h:102-108

c
static inline const char* ibvGetGidStr(union ibv_gid* gid, char* gidStr, size_t strLen) {
  static_assert(sizeof(union ibv_gid) == sizeof(struct in6_addr),
                "the sizeof struct ibv_gid must be the size of struct in6_addr");
  return inet_ntop(AF_INET6, gid->raw, gidStr, strLen);
}

static_assertguarantees at compile time thatibv_gidandin6_addrhave the same size, so thatinet_ntopcan correctly interpret these 16 bytes.

ibv_mr: Memory Registration Handle

📎 src/include/ibvcore.h:402-410

c
struct ibv_mr {
	struct ibv_context     *context;
	struct ibv_pd	       *pd;
	void		       *addr;
	size_t			length;
	uint32_t		handle;
	uint32_t		lkey;
	uint32_t		rkey;
};

This is the core of GPUDirect RDMA.addris the starting address of the registered memory (can be host memory, or GPU memory mapped to host address),lengthis the length.lkey(local key) andrkey(remote key) are the "keys" the NIC uses to verify access permissions—the sender includeslkeyin the WQE, and the receiver usesrkeyto validate.

[Design Inference and Architectural Trade-offs]

Why is registration needed? Because the NIC uses physical addresses for DMA, whileaddris a virtual address. The registration process makes the driver "pin" the page table for this virtual address range, establish IOMMU mappings, and returnlkey/rkeyas a handle for subsequent references. Registration is expensive (involving page table walks and IOMMU programming), so NCCL caches MRs to avoid registering on every transfer.

ibv_send_wr: Send Work Request

📎 src/include/ibvcore.h:704-738

c
struct ibv_send_wr {
	uint64_t		wr_id;
	struct ibv_send_wr     *next;
	struct ibv_sge	       *sg_list;
	int			num_sge;
	enum ibv_wr_opcode	opcode;
	int			send_flags;
	uint32_t		imm_data;
	union {
		struct {
			uint64_t	remote_addr;
			uint32_t	rkey;
		} rdma;
		// ...
	} wr;
};

This is the description of "what I want the NIC to do."wr_idis a user-defined tag (returned as-is upon completion),sg_listis the scatter-gather list,opcodedetermines the operation type (RDMA_WRITE, SEND, etc.),wr.rdma.remote_addrandwr.rdma.rkeySpecify the target address and access key of the peer.

ibv_sgeDescribe a segment of local memory:

📎 src/include/ibvcore.h:698-702

c
struct ibv_sge {
	uint64_t		addr;
	uint32_t		length;
	uint32_t		lkey;
};

Noteaddrisuint64_trather than a pointer—because the WQE is read by the NIC hardware, it must be in a fixed 64-bit format.

Inline functions: the fast path that bypasses the symbol table

Some functions NCCL chooses to implement inline rather than going through the symbol table. For exampleibv_post_send:

📎 src/include/ibvcore.h:1099-1101

c
static inline int ibv_post_send(struct ibv_qp *qp, struct ibv_send_wr *wr, struct ibv_send_wr **bad_wr) {
  return qp->context->ops.post_send(qp, wr, bad_wr);
}

It directly calls through theqp->context->ops.post_sendfunction pointer. This is the classic design of libibverbs:ibv_contextThere is aopsstruct that contains all operation function pointers, filled in by the specific driver.

[Design inference and architectural trade-offs]

Why doespost_sendgo throughopsinstead of the symbol table? Becausepost_sendis adata pathhot function, called on every send. If it went through thedlsymresolved global symbol table, there would be an extra level of indirection. But throughqp->context->ops, the compiler can perform better optimizations, and this pointer is fixed at QP creation time. In contrast,ibv_modify_qpis a control path function, called infrequently, so going through the symbol table doesn't matter.

NCCL's wrapperwrap_ibv_post_sendis also inline:

📎 src/include/ibvwrap.h:77-85

c
static inline ncclResult_t wrap_ibv_post_send(struct ibv_qp* qp, struct ibv_send_wr* wr, struct ibv_send_wr** bad_wr) {
  int ret = qp->context->ops.post_send(
    qp, wr, bad_wr);
  if (ret != IBV_SUCCESS) {
    WARN("ibv_post_send() failed with error %s, Bad WR %p, First WR %p", strerror(ret), wr, *bad_wr);
    return ncclSystemError;
  }
  return ncclSuccess;
}

NoteIBV_SUCCESSis defined as 0:

📎 src/include/ibvwrap.h:23-25

c
typedef enum ibv_return_enum {
  IBV_SUCCESS = 0,
} ibv_return_t;

Design thinking: ABI compatibility "version probing"

ibvcore.hThere is a clever piece of ABI version probing code in

📎 src/include/ibvcore.h:81

c
static void *__VERBS_ABI_IS_EXTENDED = ((uint8_t *)NULL) - 1;

This is a "magic pointer"—the value is(uint8_t*)0 - 1, i.e.0xFFFFFFFFFFFFFFFF. It is used as a marker value for theibv_context.abi_compatfield:

📎 src/include/ibvcore.h:1072-1081

c
static inline struct verbs_context *verbs_get_ctx(struct ibv_context *ctx)
{
	if (ctx->abi_compat != __VERBS_ABI_IS_EXTENDED)
		return NULL;
	return (struct verbs_context *)(((uintptr_t)ctx) -
					offsetof(struct verbs_context,
						 context));
}

Ifabi_compatequals this magic value, it means the underlying library supports the extended ABI, and at this point you can use thecontainer_oftrick to derive fromibv_contextthe outerverbs_context。verbs_contextwhose last field isibv_context:

📎 src/include/ibvcore.h:1068-1069

c
	size_t   sz;			/* Must be immediately before struct ibv_context */
	struct ibv_context context;	/* Must be last field in the struct */
[Design inference and architectural trade-offs]

This is the classic technique for implementing "inheritance" in C:verbs_context"inherits"ibv_context, and by placing the base class at the end, you can usecontainer_ofto derive the derived class pointer from the base class pointer.szThe field records the struct size, used for version compatibility—new versions of the library can extend the struct, and old code can checkszto determine whether a certain field exists.

verbs_get_ctx_opThe macro further encapsulates this check:

📎 src/include/ibvcore.h:1083-1086

c
#define verbs_get_ctx_op(ctx, op) ({ \
	struct verbs_context *__vctx = verbs_get_ctx(ctx); \
	(!__vctx || (__vctx->sz < sizeof(*__vctx) - offsetof(struct verbs_context, op)) || \
	 !__vctx->op) ? NULL : __vctx; })

It checks three things: whether it is the extended ABI, whether the struct is large enough to contain the field, and whether the field is non-null. Only if all are satisfied does it return a valid pointer. This is the basis foribv_query_port_exbeing able to call safely:

📎 src/include/ibvcore.h:1121-1132

c
static inline int ibv_query_port_ex(struct ibv_context *context,
				    uint8_t port_num,
				    struct ibv_port_attr *port_attr)
{
	struct verbs_context *vctx = verbs_get_ctx_op(context, query_port);
        if (vctx) {
          return vctx->query_port(context, port_num, port_attr, sizeof(*port_attr));
        }
        return -1;
}

If the underlying library does not support extendedquery_port, it returns -1, and the callerwrap_ibv_query_portfalls back to the old API:

📎 src/misc/ibvwrap.cc:156-171

c
ncclResult_t wrap_ibv_query_port(struct ibv_context* context, uint8_t port_num, struct ibv_port_attr* port_attr) {
#ifndef NCCL_BUILD_RDMA_CORE
  // First try and query the extended port attributes (e.g. active_speed_ex)
  if (ibv_query_port_ex(context, port_num, port_attr) != 0) {
    // Fall back to the original attribute API call, but zero all members first
    memset(port_attr, 0, sizeof(*port_attr));
    IBV_INT_CHECK_RET_ERRNO(ibvSymbols, ibv_internal_query_port, ibv_internal_query_port(context, port_num, port_attr),
                            0, "ibv_query_port");
  }
#else
  IBV_INT_CHECK_RET_ERRNO(ibvSymbols, ibv_internal_query_port, ibv_internal_query_port(context, port_num, port_attr), 0,
                          "ibv_query_port");
#endif
  return ncclSuccess;
}

Notememset(port_attr, 0, sizeof(*port_attr))—clear to zero before falling back, because the old API will not fill inactive_speed_exand other new fields; if not cleared, it will read garbage values from the stack.

13.3 QP state machine and the retry art of modify_qp

Intuitive model: QP is the complete process of "making a phone call"

Queue Pair (QP) is the basic unit of RDMA communication, containing a send queue (SQ) and a receive queue (RQ). Establishing a QP is like making a phone call: first dial (RESET→INIT), wait for the other party to answer (INIT→RTR), confirm both sides can hear each other (RTR→RTS), and then you can talk.

If the QP state machine goes wrong, the disaster is:the NIC cannot establish a connection, all cross-machine communication fails, and the training task hangs or crashes. And QP state transitions are exactly where problems are most likely to occur—network jitter, GID changes, and cross-rail connection errors can all causeibv_modify_qpfailures.

State enum and transitions

📎 src/include/ibvcore.h:636-645

c
enum ibv_qp_state {
	IBV_QPS_RESET,
	IBV_QPS_INIT,
	IBV_QPS_RTR,
	IBV_QPS_RTS,
	IBV_QPS_SQD,
	IBV_QPS_SQE,
	IBV_QPS_ERR,
	IBV_QPS_UNKNOWN
};

This is the standard RDMA QP state machine. NCCL'sibvQpStateNametranslates the enum into readable strings for logging:

📎 src/misc/ibvwrap.cc:263-293

c
static void ibvQpStateName(enum ibv_qp_state state, char* msg, const size_t len) {
  switch (state) {
  case (IBV_QPS_RESET):
    snprintf(msg, len, "RESET");
    break;
  case (IBV_QPS_INIT):
    snprintf(msg, len, "INIT");
    break;
  // ...
  }
}

The state diagram below precisely corresponds to the enum and transition semantics in the source code:

mermaid
stateDiagram-v2
    [*] --> RESET : ibv_create_qp()
    RESET --> INIT : modify_qp(IBV_QPS_INIT) [设置 pkey_index, port]
    INIT --> RTR : modify_qp(IBV_QPS_RTR) [设置 ah_attr, dest_qp_num, rq_psn]
    RTR --> RTS : modify_qp(IBV_QPS_RTS) [设置 sq_psn, timeout, retry_cnt]
    RTS --> SQD : modify_qp(IBV_QPS_SQD) [SQ Drain]
    SQD --> RTS : modify_qp(IBV_QPS_RTS)
    RTS --> ERR : 硬件错误 / WC 错误
    RTR --> ERR : 硬件错误
    ERR --> RESET : modify_qp(IBV_QPS_RESET) [错误恢复]
[Design inference and architectural trade-offs]

NoteIBV_QPS_SQD(SQ Drained) andIBV_QPS_SQE(SQ Error), these two states. SQD is used for graceful shutdown—transition after draining the send queue. SQE indicates a send queue error. NCCL does not actively enter these two states on the normal path, but error handling needs to recognize them.

Step-by-Step: the retry logic of modify_qp

wrap_ibv_modify_qpis the most complex function in this chapter, implementing a complete retry mechanism:

📎 src/misc/ibvwrap.cc:360-385

c
ncclResult_t wrap_ibv_modify_qp(struct ibv_qp* qp, struct ibv_qp_attr* attr, int attr_mask) {
  char qpMsg[1024];
  int ret = 0, attempts = 0;
  int maxCnt = (int)ncclParamIbMQpRetryCnt() + 1; // number of attempts = number of retry + 1
  int timeOut = (int)ncclParamIbMQpRetryTimeout();
  CHECK_NOT_NULL(ibvSymbols, ibv_internal_modify_qp);
  do {
    if (attempts > 0) {
      unsigned int sleepTime = timeOut * attempts;
      ibvModifyQpLog(qp, attr->qp_state, attr, attr_mask, qpMsg, sizeof(qpMsg));
      INFO(NCCL_NET, "Call to ibv_modify_qp failed with %d %s, %s, retrying %d/%d after %u msec of sleep", ret,
           strerror(ret), qpMsg, attempts, maxCnt, sleepTime);
      // sleep before retrying
      std::this_thread::sleep_for(std::chrono::milliseconds(sleepTime));
    }
    ret = ibvSymbols.ibv_internal_modify_qp(qp, attr, attr_mask);
    attempts++;
  } while (IBV_MQP_RETRY_ERRNO_ALL(ret) && attempts < maxCnt);
  if (ret != 0) {
    ibvModifyQpLog(qp, attr->qp_state, attr, attr_mask, qpMsg, sizeof(qpMsg));
    WARN("Call to ibv_modify_qp failed with %d %s, %s", ret, strerror(ret), qpMsg);
    printIbModifyQpHint(ret);
    return ncclSystemError;
  }
  return ncclSuccess;
}

Step-by-step breakdown:

Step 1: Read parameters。maxCnt = IbMQpRetryCnt() + 1, default retry 34 times, so at most 35 attempts.timeOutDefault 100 milliseconds.

Step 2: Enter the retry loop. The first timeattempts == 0, no sleep, call directly. After that, on each failure,sleepTime = timeOut * attempts—this islinear backoff, the 1st retry waits 100ms, the 2nd waits 200ms, the 34th waits 3400ms.

Step 3: Determine whether to retry。IBV_MQP_RETRY_ERRNO_ALL(ret)decides whether to continue:

📎 src/misc/ibvwrap.cc:107-109

c
#define IBV_ERR_EQ(e, code) (e == code || e == (-code))
#define IBV_MQP_RETRY_ERRNO(e) (IBV_ERR_EQ(e, ETIMEDOUT))
#define IBV_MQP_RETRY_ERRNO_ALL(e) (ncclParamIbMQpRetryAll() ? (e != 0) : IBV_MQP_RETRY_ERRNO(e))

By default only retriesETIMEDOUT.IBV_ERR_EQmatches both positive and negative values, because different drivers may returnETIMEDOUTor-ETIMEDOUT. IfNCCL_IB_MQP_RETRY_ALL=1is set, then retry on any non-zero error.

Step 4: Print diagnostic information on failure。ibvModifyQpLogcollects the device name, port number, current state, target state, local/remote GID:

📎 src/misc/ibvwrap.cc:297-339

c
static void ibvModifyQpLog(struct ibv_qp* qp, enum ibv_qp_state qpState, struct ibv_qp_attr* userAttr, int userFlag,
                           char* msg, size_t msgLen) {
  // ...
  char nextState[32], currState[32];
  ibvQpStateName(qp->state, currState, sizeof(currState));
  ibvQpStateName(qpState, nextState, sizeof(nextState));
  char devName[IBV_SYSFS_NAME_MAX] = "";
  snprintf(devName, sizeof(devName), "%s",
           (qp->pd->context) ? wrap_ibv_get_device_name(qp->pd->context->device) : "N/A");
  // ...
}

Note the clever design of theQP_ATTRmacro:

📎 src/misc/ibvwrap.cc:295

c
#define QP_ATTR(attr, userAttr, userFlag, mask) ((userFlag & mask) ? (userAttr) : (attr))

It prioritizes the attributes passed in by the user (if the corresponding bit is set inattr_mask), otherwise falls back to the current attributes found byquery_qp. This way, even ifquery_qpfails, some information can still be obtained from the user parameters.

Step 5: Give hints on failure。printIbModifyQpHintprovides troubleshooting suggestions for common error codes:

📎 src/misc/ibvwrap.cc:341-358

c
static void printIbModifyQpHint(int status) {
  switch (status) {
  case ETIMEDOUT:
    INFO(NCCL_NET, "HINT: In many cases this error indicates that the NICs are not cross-rail connected.");
    INFO(NCCL_NET, "HINT: To confirm, set NCCL_CROSS_NIC=0 to disable cross-rail communication ...");
    return;
  case EINVAL:
    INFO(NCCL_NET, "HINT: In many cases this error indicates that an incorrect GID index is forced by "
                   "NCCL_IB_GID_INDEX, or that a NIC's GID changed mid-run.");
    // ...
  }
}
[Design inference and architectural trade-offs]

This hint is the crystallization of production experience.ETIMEDOUTThe most common cause is cross-rail connection problems—in a multi-rail network, if rank A's NIC 0 tries to connect to rank B's NIC 1, and they are not on the same rail, it will time out.EINVALThis is usually a GID index configuration error, or the GID changes during operation (such as a NIC reset).

Concurrency Control and Hardware Interaction

wrap_ibv_modify_qpitself is not locked—it assumes the caller guarantees that the same QP will not be modified by multiple threads simultaneously. This holds in NCCL: QP establishment occurs during the initialization phase and is completed by a single thread.

[Design Inference and Architectural Trade-offs]

But in the retry loop,std::this_thread::sleep_foris worth noting. It yields the CPU but does not release any locks (since no locks are held in the first place). When this function is called in the proxy thread, the sleep will block the proxy's progress—if QP establishment gets stuck, the entire communication will stall. This is why the default retry count is 34 times, with a total time of about 60 seconds—enough to cover brief network jitter, but not to wait indefinitely.

13.4 Memory Registration: The Entry Point of GPUDirect RDMA

Intuitive Model: Issuing an "Access Card" to the NIC

For the NIC to directly read and write memory, it must first "recognize" this memory. Memory registration (ibv_reg_mr) is issuing an access card to the NIC—telling it the physical address range of this memory, and returning alkey(local key) andrkey(remote key). Afterward, when the NIC performs DMA, it accesses memory using this key.

If memory registration is missing, the disaster is:The NIC cannot access any memory, and RDMA completely fails to work. A more insidious problem is: if host memory is registered but you want to access GPU memory, the NIC will read incorrect data or trigger a protection error.

Three Registration Paths

NCCL encapsulates three memory registration functions, corresponding to different usage scenarios:

Path One: Regular Registration

📎 src/misc/ibvwrap.cc:198-201

c
ncclResult_t wrap_ibv_reg_mr(struct ibv_mr** ret, struct ibv_pd* pd, void* addr, size_t length, int access) {
  IBV_PTR_CHECK_ERRNO(ibvSymbols, ibv_internal_reg_mr, ibv_internal_reg_mr(pd, addr, length, access), *ret, NULL,
                      "ibv_reg_mr");
}

This is the standard path,addris the virtual address,accessis the access permission flag (IBV_ACCESS_LOCAL_WRITE | IBV_ACCESS_REMOTE_WRITEetc.).

Path Two: Registration with Specified IOVA

📎 src/misc/ibvwrap.cc:211-219

c
ncclResult_t wrap_ibv_reg_mr_iova2(struct ibv_mr** ret, struct ibv_pd* pd, void* addr, size_t length, uint64_t iova,
                                   int access) {
  if (ibvSymbols.ibv_internal_reg_mr_iova2 == NULL) {
    return ncclInternalError;
  }
  if (ret == NULL) return ncclSuccess; // Assume dummy call
  IBV_PTR_CHECK_ERRNO(ibvSymbols, ibv_internal_reg_mr_iova2, ibv_internal_reg_mr_iova2(pd, addr, length, iova, access),
                      *ret, NULL, "ibv_reg_mr_iova2");
}

iova(I/O Virtual Address) allows specifying the address as seen by the NIC. This is useful in scenarios requiring fixed address mapping. Note thatret == NULLdirectly returns success—this is a "probe call" that only checks whether the function exists, without actually registering.

Path Three: DMA-BUF Registration (The Key to GPUDirect RDMA)

📎 src/misc/ibvwrap.cc:222-227

c
ncclResult_t wrap_ibv_reg_dmabuf_mr(struct ibv_mr** ret, struct ibv_pd* pd, uint64_t offset, size_t length,
                                    uint64_t iova, int fd, int access) {
  IBV_PTR_CHECK_ERRNO(ibvSymbols, ibv_internal_reg_dmabuf_mr,
                      ibv_internal_reg_dmabuf_mr(pd, offset, length, iova, fd, access), *ret, NULL,
                      "ibv_reg_dmabuf_mr");
}

This is the core of GPUDirect RDMA.fdis a DMA-BUF file descriptor—it represents a block of GPU memory. NCCL obtains this fd throughcuMemGetHandleForAddressRangeor similar CUDA APIs, then passes it toibv_reg_dmabuf_mr. The NIC driver directly maps GPU memory through the DMA-BUF mechanism, without going through host memory copies.

[Design Inference and Architectural Trade-offs]

DMA-BUF is the Linux kernel's buffer sharing framework. The GPU driver (such as NVIDIA's nvidia.ko) exports GPU memory as a DMA-BUF, and the NIC driver (such as mlx5) imports it, establishing an IOMMU mapping. The entire process is completed in the kernel, with user space only passing an fd. This is the underlying mechanism for "the NIC directly reading and writing GPU memory."

Direct Registration vs. Wrapped Registration

Note that there are two "direct" versions:

📎 src/misc/ibvwrap.cc:203-209

c
struct ibv_mr* wrap_direct_ibv_reg_mr(struct ibv_pd* pd, void* addr, size_t length, int access) {
  if (ibvSymbols.ibv_internal_reg_mr == NULL) {
    WARN("lib wrapper not initialized.");
    return NULL;
  }
  return ibvSymbols.ibv_internal_reg_mr(pd, addr, length, access);
}

📎 src/misc/ibvwrap.cc:229-236

c
struct ibv_mr* wrap_direct_ibv_reg_dmabuf_mr(struct ibv_pd* pd, uint64_t offset, size_t length, uint64_t iova, int fd,
                                             int access) {
  if (ibvSymbols.ibv_internal_reg_dmabuf_mr == NULL) {
    errno = EOPNOTSUPP; // ncclIbDmaBufSupport() requires this errno being set
    return NULL;
  }
  return ibvSymbols.ibv_internal_reg_dmabuf_mr(pd, offset, length, iova, fd, access);
}

They directly returnibv_mr*rather thanncclResult_t, and do not print WARN logs. Why?

[Design Inference and Architectural Trade-offs]

Because these two functions are used forcapability probing。ncclIbDmaBufSupport()will callwrap_direct_ibv_reg_dmabuf_mrto probe whether the NIC supports DMA-BUF. If it fails, it expects to geterrno == EOPNOTSUPPto determine "not supported" rather than "error." If a WARN were printed here, it would flood the screen on machines that do not support DMA-BUF. So the direct versions delegate error handling responsibility to the caller.

Access Permission Flags

📎 src/include/ibvcore.h:365-372

c
enum ibv_access_flags {
	IBV_ACCESS_LOCAL_WRITE		= 1,
	IBV_ACCESS_REMOTE_WRITE		= (1<<1),
	IBV_ACCESS_REMOTE_READ		= (1<<2),
	IBV_ACCESS_REMOTE_ATOMIC	= (1<<3),
	IBV_ACCESS_MW_BIND		= (1<<4),
	IBV_ACCESS_RELAXED_ORDERING     = (1<<20),
};

These flags are bitmasks and can be combined.LOCAL_WRITEallows local write (needed when receiving data),REMOTE_WRITEallows remote write (needed for RDMA WRITE targets),REMOTE_READallows remote read (needed for RDMA READ targets).

IBV_ACCESS_RELAXED_ORDERINGis a performance optimization flag—it allows the NIC to access with a more relaxed memory ordering, potentially improving throughput, but requires the application layer to guarantee correctness.

Data Flow: The Complete Path from GPU Memory to the NIC

The following diagram shows the data flow of a cross-machine RDMA write, anchored to the structs covered in this chapter:

mermaid
flowchart LR
    subgraph GPU["GPU 显存"]
        buf["ncclSendBuff<br/>(device ptr)"]
    end
    subgraph Host["Host 进程"]
        dmabuf["DMA-BUF fd<br/>(cuMemGetHandleForAddressRange)"]
        mr["ibv_mr<br/>{addr, lkey, rkey}"]
        wr["ibv_send_wr<br/>{opcode=RDMA_WRITE,<br/>sg_list, wr.rdma.remote_addr, rkey}"]
    end
    subgraph NIC["网卡 mlx5"]
        qp["ibv_qp<br/>(SQ + RQ)"]
        wqe["WQE<br/>(硬件工作队列元素)"]
    end
    buf -->|导出| dmabuf
    dmabuf -->|ibv_reg_dmabuf_mr| mr
    mr -->|填充 sge.lkey| wr
    wr -->|ibv_post_send| qp
    qp -->|DMA 读取| wqe
    wqe -->|PCIe P2P| buf
    wqe -->|网络| remote["对端 GPU 显存<br/>(remote_addr + rkey)"]

Each node in the diagram corresponds to a real type in the source code:ibv_mrfrom📎 src/include/ibvcore.h:402-410,ibv_send_wrfrom📎 src/include/ibvcore.h:704-738,ibv_qpfrom📎 src/include/ibvcore.h:787-802。

13.5 Work Completion and Error Diagnosis

Intuitive Model: Delivery Receipt

RDMA is asynchronous—after youpost_send, you will not immediately know the result. After the NIC completes the operation, it places a Work Completion (WC) in the Completion Queue (CQ), just like a courier putting a delivery receipt in your mailbox. You need to activelypoll_cqto retrieve it.

If WC diagnosis is missing, the disaster is:When communication fails, you only know "it failed," not "why it failed". RDMA has more than 20 error codes, each corresponding to a different root cause.

WC Struct

📎 src/include/ibvcore.h:349-363

c
struct ibv_wc {
	uint64_t		wr_id;
	enum ibv_wc_status	status;
	enum ibv_wc_opcode	opcode;
	uint32_t		vendor_err;
	uint32_t		byte_len;
	uint32_t		imm_data;	/* in network byte order */
	uint32_t		qp_num;
	uint32_t		src_qp;
	int			wc_flags;
	uint16_t		pkey_index;
	uint16_t		slid;
	uint8_t			sl;
	uint8_t			dlid_path_bits;
};

wr_idis the tag you filled in when posting,statusis the completion status,opcodeis the operation type,byte_lenis the actual number of bytes transferred.qp_numandsrc_qpare used to identify which QP completed in multi-QP scenarios.

Status Code Translation

ibvWcStatusStrtranslates the status enum into strings:

📎 src/misc/ibvwrap.cc:415-464

c
const char* ibvWcStatusStr(enum ibv_wc_status status) {
  switch (status) {
  case IBV_WC_SUCCESS:
    return "IBV_WC_SUCCESS";
  case IBV_WC_LOC_LEN_ERR:
    return "IBV_WC_LOC_LEN_ERR";
  // ... 20 多个 case
  default:
    return "UNKNOWN_STATUS";
  }
}

The meanings of these status codes:

Status CodeMeaningCommon Root Cause
IBV_WC_SUCCESSSuccess—
IBV_WC_LOC_LEN_ERRLocal length errorSGE length exceeds MR range
IBV_WC_LOC_ACCESS_ERRLocal access errorInvalid lkey or insufficient permissions
IBV_WC_REM_ACCESS_ERRRemote access errorInvalid rkey or the peer's MR has been deregistered
IBV_WC_RETRY_EXC_ERRRetry exhaustedNetwork unreachable or the peer's QP not ready
IBV_WC_RNR_RETRY_EXC_ERRRNR retry exhaustedThe peer did not post recv
IBV_WC_RESP_TIMEOUT_ERRResponse timeoutPeer not responding
[Design Inference and Architectural Trade-offs]

IBV_WC_RNR_RETRY_EXC_ERR(Receiver Not Ready) is one of the most common issues in production environments. It means the sender sent data, but the receiver did not pre-post enough recv buffers. In NCCL, this typically occurs during the connection establishment phase—the QP states on both sides are out of sync, one side has already started sending, while the other side is not yet ready to receive.

opcode translation

ibvWcOpcodeStrandibvWrOpcodeStrrespectively translate the completion opcode and request opcode:

📎 src/misc/ibvwrap.cc:467-488

c
const char* ibvWcOpcodeStr(enum ibv_wc_opcode opcode) {
  switch (opcode) {
  case IBV_WC_SEND:
    return "IBV_WC_SEND";
  case IBV_WC_RDMA_WRITE:
    return "IBV_WC_RDMA_WRITE";
  case IBV_WC_RDMA_READ:
    return "IBV_WC_RDMA_READ";
  // ...
  }
}

Note thatIBV_WC_RECVthe value of is1 << 7:

📎 src/include/ibvcore.h:329-342

c
enum ibv_wc_opcode {
	IBV_WC_SEND,
	IBV_WC_RDMA_WRITE,
	IBV_WC_RDMA_READ,
	IBV_WC_COMP_SWAP,
	IBV_WC_FETCH_ADD,
	IBV_WC_BIND_MW,
	IBV_WC_RECV			= 1 << 7,
	IBV_WC_RECV_RDMA_WITH_IMM
};
[Design Inference and Architectural Trade-offs]

Why isIBV_WC_RECV1 << 7instead of a sequential value? Because receive completion and send completion are two different types of operations, using the high bit to distinguish them allows the code to useopcode & IBV_WC_RECVto quickly determine "is this a receive completion." This is the API design convention of libibverbs.

Poll CQ

wrap_ibv_poll_cqis inlined:

📎 src/include/ibvwrap.h:60-69

c
static inline ncclResult_t wrap_ibv_poll_cq(struct ibv_cq* cq, int num_entries, struct ibv_wc* wc, int* num_done) {
  int done = cq->context->ops.poll_cq(cq, num_entries,
                                      wc);
  if (done < 0) {
    WARN("Call to ibv_poll_cq() returned %d", done);
    return ncclSystemError;
  }
  *num_done = done;
  return ncclSuccess;
}

It is called throughcq->context->ops.poll_cqand, likepost_sendgoes through theopsfast path. The return valuedoneis the number of WCs polled this time, 0 means no new completions, negative means error.

[Design Inference and Architectural Trade-offs]

poll_cqisbusy polling—it does not block and returns immediately. NCCL's proxy thread will repeatedly call it in a loop until it gets a completion event. This is the key to low latency: compared to interrupt-driven, busy polling avoids the overhead of interrupt context switching. The cost is high CPU usage, but in high-performance computing scenarios this is acceptable.

13.6 Production Pitfall Guide

Pitfall One: Cross-rail Connection Timeout

Symptom:ibv_modify_qpreturnsETIMEDOUT, fails after 34 retries.

Root Cause: In a multi-rail network, each GPU is typically bound to a specific NIC. If rank A's GPU 0 is bound to NIC 0, and rank B's GPU 0 is bound to NIC 1, and NIC 0 and NIC 1 are not on the same rail (i.e., they connect to different switches), then QP establishment will time out.

Troubleshooting: The source code already provides a hint:

📎 src/misc/ibvwrap.cc:343-347

c
  case ETIMEDOUT:
    INFO(NCCL_NET, "HINT: In many cases this error indicates that the NICs are not cross-rail connected.");
    INFO(NCCL_NET, "HINT: To confirm, set NCCL_CROSS_NIC=0 to disable cross-rail communication ...");
    return;

SettingNCCL_CROSS_NIC=0can force same-rail communication. If this resolves the issue, it confirms it is indeed a cross-rail problem.

Recovery Chain: NCCL's retry mechanism (34 times, linear backoff) gives the network enough time to recover. But if the root cause is a topology configuration error, retries are useless, and theNCCL_IB_HCAorNCCL_CROSS_NICconfiguration must be corrected.

Pitfall Two: GID Index Error

Symptom:ibv_modify_qpreturnsEINVAL。

Root Cause:NCCL_IB_GID_INDEXforcibly specified a non-existent GID index, or the NIC's GID changed during runtime (e.g., a RoCE NIC re-acquiring an IP).

Troubleshooting:

📎 src/misc/ibvwrap.cc:341-358

c
  case EINVAL:
    INFO(NCCL_NET, "HINT: In many cases this error indicates an incorrect GID index is forced by "
                   "NCCL_IB_GID_INDEX, or that a NIC's GID changed mid-run.");
    INFO(NCCL_NET, "HINT: To confirm, set NCCL_IB_GID_INDEX=-1 to enable automatic detection and check "
                   "'dmesg | grep -i gid' for GID changes ...");
    return;

SettingNCCL_IB_GID_INDEX=-1enables automatic detection. Also checkdmesgfor GID change events.

Pitfall Three: DMA-BUF Not Supported Causing Fallback to Host Copy

Symptom: GPUDirect RDMA is not taking effect, performance is lower than expected.

Root Cause: The NIC driver or kernel does not support DMA-BUF,wrap_direct_ibv_reg_dmabuf_mrreturns NULL and setserrno = EOPNOTSUPP:

📎 src/misc/ibvwrap.cc:229-236

c
struct ibv_mr* wrap_direct_ibv_reg_dmabuf_mr(struct ibv_pd* pd, uint64_t offset, size_t length, uint64_t iova, int fd,
                                             int access) {
  if (ibvSymbols.ibv_internal_reg_dmabuf_mr == NULL) {
    errno = EOPNOTSUPP; // ncclIbDmaBufSupport() requires this errno being set
    return NULL;
  }
  return ibvSymbols.ibv_internal_reg_dmabuf_mr(pd, offset, length, iova, fd, access);
}

Note the comment:ncclIbDmaBufSupport()relies on thiserrnoto determine whether it is supported. IfEOPNOTSUPPis not set here, the upper layer will misinterpret it as "error" rather than "not supported."

Troubleshooting: Check the kernel version (requires 5.12+), NIC driver version, and whether thenvidia-peermemmodule is loaded. If it is indeed not supported, NCCL will fall back to host memory staging, performance will decrease but functionality remains normal.

Pitfall Four: MR Cache and Memory Leak

[Design Inference and Architectural Trade-offs]

Memory registration is an expensive operation (involving IOMMU programming), NCCL cachesibv_mr. But if the caching strategy is improper, it leads to two problems: first, memory leaks (MRs are never deregistered), and second, cache invalidation (memory is freed but the MR still points to the old address).

wrap_ibv_dereg_mris the deregistration entry point:

📎 src/misc/ibvwrap.cc:238-241

c
ncclResult_t wrap_ibv_dereg_mr(
  struct ibv_mr* mr) {
  IBV_INT_CHECK_RET_ERRNO(ibvSymbols, ibv_internal_dereg_mr, ibv_internal_dereg_mr(mr), 0, "ibv_dereg_mr");
}
[Design Inference and Architectural Trade-offs]

In production environments, if training jobs frequently create/destroy communication domains and MRs are not properly deregistered, it will cause the IOMMU mapping table to bloat, eventually triggeringibv_reg_mrfailure (returningENOMEM). The troubleshooting method is to monitor the number of mappings under/sys/kernel/debug/iommu.

Design Reflection: Why Is the Wrapper Layer So "Thick"

Reviewing this chapter,ibvwrap.cchas 509 lines,ibvcore.hhas 1134 lines. For a wrapper layer that "just calls libibverbs," this is quite substantial. Why?

[Design Inference and Architectural Trade-offs]

Three reasons:

First, the complexity of error handling. libibverbs' API error conventions are extremely inconsistent, and NCCL needs to write a macro for each convention and use it correctly in every function. This is not over-engineering, but the necessary cost of "faithful translation."

Second, the burden of ABI compatibility。ibvcore.hredefines all structures and also handlesverbs_contextversion detection. This is to avoid depending on IB header files at compile time and be compatible with any version at runtime.

Third, the value of diagnostic information。ibvModifyQpLog、printIbModifyQpHint、ibvWcStatusStrThese functions are not called on the normal path, but they are invaluable during troubleshooting. NCCL chose to "pre-embed" diagnostic information in the wrapper layer, rather than collecting it on the fly when errors occur.

The cost of this "thick wrapper" is a large codebase and high maintenance cost. But the benefit is: the upper layernet_ib.cccan be written using a unifiedncclResult_tinterface, without needing to care about the various quirks of libibverbs. This is a typical "complexity isolation" design.

Chapter Summary

In this chapter, we took a deep dive into NCCL's InfiniBand transport wrapper layer. The core points are:

1. Symbol table wrapper:ncclIbvSymbolsThroughdlopen + dlsymruntime loading of libibverbs, combined withstd::once_flagensuring thread-safe initialization. This allows NCCL to load even on machines without an IB driver.

2. ABI contract:ibvcore.hRedefines the core types of libibverbs, using__VERBS_ABI_IS_EXTENDEDmagic pointers andverbs_contextthecontainer_oftechnique for version detection.

3. QP state machine:wrap_ibv_modify_qpImplements 34 linear backoff retries, providing diagnostic hints forETIMEDOUTandEINVAL.

4. GPUDirect RDMA:wrap_ibv_reg_dmabuf_mrThrough the DMA-BUF mechanism, the NIC can directly map GPU memory,wrap_direct_ibv_reg_dmabuf_mrused for capability probing.

5. Error diagnostics:ibvWcStatusStr、ibvWcOpcodeStr、ibvWrOpcodeStrTranslating hardware error codes into readable strings is a key tool for production troubleshooting.

Chapter Review and Self-Test

Q1: Ifwrap_ibv_symbolsinstd::call_onceis replaced with a regularif (initResult == ncclSuccess) return initResult;double-checked lock, in what concurrency scenarios would problems arise?

Reference Analysis: See📎 src/misc/ibvwrap.cc:26-29:

c
ncclResult_t wrap_ibv_symbols(void) {
  std::call_once(initOnceFlag, []() { initResult = buildIbvSymbols(&ibvSymbols); });
  return initResult;
}

If replaced with a naive double-checked lock, the problem lies inmemory reordering。buildIbvSymbolswill fill inibvSymbolsthe various fields of, and then writeinitResult. Without a memory barrier, the CPU or compiler may reorderinitResult = ncclSuccessto `

At this point, we have seen clearly how NCCL wraps libibverbs into a pluggable transport layer through net_ib, and uses GPUDirect RDMA to enable direct NIC access to GPU memory. This mechanism addresses the latency and bandwidth bottlenecks of inter-node communication. But intra-node communication is equally critical—in the next chapter, we will enter symmetric memory and NVLS to see how NCCL leverages NVLink multicast for hardware-accelerated collective communication. You will then find that the RDMA mechanism in this chapter complements NVLS: the former handles inter-node, while the latter handles intra-node.

CHAPTER 14

Chapter 14: Chapter 14: Symmetric Memory and NVLS: Multicast Acceleration and LSA Device-Side Direct Addressing

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 14 / 25

Chapter 14: Symmetric Memory and NVLS: Multicast Acceleration and LSA Device-Side Direct Addressing

In the previous chapter, we followed an inter-node AllReduce and saw how data travels from GPU memory through the NIC to the peer GPU. That path solves communication between machines. But in modern AI clusters, the communication volume between GPUs within the same machine or even within the same NVLink domain is equally enormous—gradient synchronization in data parallel training and activation exchange in tensor parallelism mostly occur within a node. If intra-node communication still goes through the inter-node flow of GPU→memory→NIC→peer NIC→memory→GPU, it is like sending a local package by air freight, wasting latency for no reason. This chapter will dissect exactly the two powerful tools NCCL prepares for intra-node communication: symmetric memory and NVLS. The former lets each rank use the same set of virtual addresses to access all ranks' buffers, while the latter uses the multicast capability of NVSwitch hardware for reduction. Combined, they can push the latency of small-message collective communication close to the hardware limit.

14.1 Symmetric Memory: Making "Row 3, Seat 5" Point to the Same Location in Everyone's Home

Intuitive Model

Imagine a class exchanging homework notebooks. The traditional approach is: everyone numbers their own notebooks, then shouts, "Zhang San, my 5th notebook is for you; Li Si, my 8th notebook is for you"—everyone has to remember "whose notebook is where, and which number it is." This is ordinary communication: addresses arerelative and private, and to access peer data, you must first know the peer's address mapping.

Symmetric memory takes a different approach: the whole class agrees that the coordinate "Row 3, Seat 5" points to the same physical location in everyone's home. So if Zhang San wants Li Si's 5th notebook, he can just say "Li Si's home, Row 3, Seat 5," without any address translation. This is the core of symmetric memory:each rank's buffer is mapped to the same virtual address in all ranks' address spaces。

[Design Inference and Architectural Trade-offs]

Without symmetric memory, what disaster would intra-node collective communication face? Each rank accessing a peer buffer would have to go through an "address translation"—looking up tables, calculating offsets, and possibly even cross-process communication to confirm the mapping relationship. For small messages (a few KB), the overhead of this translation may be greater than the data transmission itself. Symmetric memory eliminates this overhead entirely, which is precisely the fundamental reason it "significantly reduces small-message latency."

Data Structures and Memory Layout

The registration type of symmetric memory is described byncclSymRegType_t,ncclGetSymRegTypewhich divides registration states into four categories based on whether the send/recv windows carry theNCCL_WIN_COLL_SYMMETRICflag.

📎 src/sym_kernels.cc:395-412

c
ncclResult_t ncclGetSymRegType(struct ncclDevrWindow* sendWin, struct ncclDevrWindow* recvWin,
                               ncclSymRegType_t* winRegType) {
  bool isSendSymmReg = false;
  bool isRecvSymmReg = false;
  if (sendWin && (sendWin->winFlags & NCCL_WIN_COLL_SYMMETRIC)) isSendSymmReg = true;
  if (recvWin && (recvWin->winFlags & NCCL_WIN_COLL_SYMMETRIC)) isRecvSymmReg = true;
  // determine the registration type
  if (!isSendSymmReg && !isRecvSymmReg) {
    *winRegType = ncclSymSendNonregRecvNonreg;
  } else if (isSendSymmReg && !isRecvSymmReg) {
    *winRegType = ncclSymSendRegRecvNonreg;
  } else if (!isSendSymmReg && isRecvSymmReg) {
    *winRegType = ncclSymSendNonregRecvReg;
  } else if (isSendSymmReg && is isRecvSymmReg) {
    *winRegType = ncclSymSendRegRecvReg;
  }
  return ncclSuccess;
}

These four states determine which path subsequent kernels take: fully symmetric registration (SendRegRecvReg) takes the fastest LSA path, fully non-registered (SendNonregRecvNonreg) takes the ordinary path, and mixed states require special handling.winFlagsTheNCCL_WIN_COLL_SYMMETRICbit in

is the marker for "whether this window has undergone symmetric registration."ncclSymkInitOnceThe initialization entry point for symmetric memory ishasLsaMultimem)。

📎 src/sym_kernels.cc:185-196

c
ncclResult_t ncclSymkInitOnce(struct ncclComm* comm) {
  // ncclTeamLsa() below calls this internally but drops the error code so we do it here.
  NCCLCHECK(ncclDevrInitOnce(comm));

  struct ncclSymkState* symk = &comm->symkState;
  if (!symk->initialized) {
    symk->initialized = true;
    struct ncclDevCommRequirements reqs = NCCL_DEV_COMM_REQUIREMENTS_INITIALIZER;
    // Disable LSA multicast for cross-clique since NVLS isn't available across cliques
    symk->hasLsaMultimem =
      ncclNvlsSymmetricMultimemEnabled(comm) && ncclTeamLsa(comm).nRanks > 2 && !comm->p2pCrossClique;
    reqs.lsaMultimem = symk->hasLsaMultimem;

hasLsaMultimemAll three conditions are indispensable: NVLS symmetric multicast is enabled, the LSA team rank count is greater than 2 (two ranks are faster with direct point-to-point, no multicast needed), and it does not cross cliques (NVSwitch multicast is unavailable when crossing cliques). This determination directly decides whetherreqs.lsaMultimemis set, which in turn affects the resource allocation of the device-side communicator.

Scenario-Driven Step-by-Step Walkthrough

Suppose we initiate an AllReduce with a message size of 4KB, and 8 ranks are within the same NVLink domain.ncclSymkMaskwill determine which kernels are available.

📎 src/sym_kernels.cc:304-352

c
uint32_t ncclSymkMask(struct ncclComm* comm, ncclFunc_t coll, int /*ncclDevRedOp_t*/ red, ncclDataType_t ty,
                      size_t nElts, bool symAligned16B) {
  uint32_t kmask = kernelMask_coll(coll);

  bool hasSTMC = comm->symkState.hasLsaMultimem;
  bool hasLDMC = false;
  if (comm->symkState.hasLsaMultimem) {
    switch (ty) {
    case ncclInt32:
    ...
      hasLDMC = red == ncclDevSum || red == ncclDevMinMax || red == ncclDevSumPostDiv;
      break;
    ...
    }
  }
  if (!hasSTMC) kmask &= ~kernelMask_STMC;
  if (!hasLDMC) kmask &= ~kernelMask_LDMC;

Step one:kernelMask_collBased on the collective type (AllReduce), retrieve the candidate kernel setkernelMask_AR. Step two: checkhasLsaMultimem. If multicast is supported, further determine whether the data type and reduction operation support LDMC (Load-Multicast). Step three: use a bitmask to clear unsupported features—kmask &= ~kernelMask_STMCremove all kernels that do not support STMC.

Next is the size limit:

📎 src/sym_kernels.cc:336-342

c
  size_t nBytes = alignUp(nElts * ncclTypeSize(ty), NCCL_SYM_KERNEL_CELL_SIZE);
  size_t nBusBytes = (coll == ncclFuncAllReduce ? 1 : comm->nRanks) * nBytes;
  // LL kernels use 32-bit ints to track element counts and indices.
  if (nBusBytes >= (size_t(2) << 30)) kmask &= ~kernelMask_LL;
  // Any kernel might use 32-bit int to track unrolled loop chunks (which are going
  // to be at least 32 bytes per chunk)
  if (nBusBytes >= 32 * (size_t(2) << 30)) kmask = 0;

There are two hard boundaries here: LL-series kernels use 32-bit integers to track element counts, so when the total bus bytes exceed 2GB, LL kernels are removed; when it exceeds 64GB, all kernels are removed (kmask = 0). This is a classic case of "trading bit width for performance"—32-bit indices save registers and instructions compared to 64-bit, but at the cost of a message size cap.

Finally, the availability checks for TMA and GIN:

📎 src/sym_kernels.cc:344-350

c
  if (!ncclSymkTmaAvailable(comm)) kmask &= ~kernelMask_Tma;
  if (!symAligned16B) kmask &= ~kernelMask_Tma;

  bool hasGin = ncclParamSymGinKernelsEnable() != 0;
  if (!hasGin) kmask &= ~kernelMask_Gin;
  bool needGin = ncclTeamLsa(comm).nRanks < comm->nRanks;
  kmask &= needGin ? kernelMask_Gin : ~kernelMask_Gin;
  return kmask;

TMA requires SMEM capacity to meet the threshold (ncclSymkTmaAvailablecheckmaxSharedMemOptin) and 16-byte alignment. GIN is only needed when "the LSA team rank count is less than the total rank count"—that is, GIN only makes sense when the communication domain crosses the LSA boundary (requiring network traversal). If the entire communication domain is within the LSA, GIN kernels are removed.

Concurrency Control and Hardware Interaction

Address resolution for symmetric memory ultimately lands on the device side.ncclSymkMakeDevWorktranslates the host-side task description into device-readable work items.

📎 src/sym_kernels.cc:380-393

c
ncclResult_t ncclSymkMakeDevWork(struct ncclComm* comm, struct ncclTaskColl* task, struct ncclSymkDevWork* outDevWork) {
  outDevWork->rootRank = task->root;
  outDevWork->redOpArg = task->opDev.scalarArg;
  outDevWork->nElts = task->count;
  outDevWork->inputWin = task->sendWin ? task->sendWin->vidmem : nullptr;
  outDevWork->inputOff =
    task->sendWin ? (uint8_t*)task->sendbuff - (uint8_t*)task->sendWin->userPtr : (size_t)task->sendbuff;
  outDevWork->outputWin = task->recvWin ? task->recvWin->vidmem : nullptr;
  outDevWork->outputOff =
    task->recvWin ? (uint8_t*)task->recvbuff - (uint8_t*)task->recvWin->userPtr : (size_t)task->recvbuff;
  outDevWork->sChannelId = 0xffff;
  outDevWork->nChannels = 0;
  return ncclSuccess;
}

Note the computation ofinputOff: if sendWin exists (a symmetrically registered window), the offset issendbuff - sendWin->userPtr—this is theoffset within the window, and the device side can compute the actual address by takinginputWin(the window base address) plusinputOff. If sendWin does not exist, the offset is directly the absolute address ofsendbuff. This design lets device-side kernels handle both registered and unregistered buffers with the same logic.

ncclSymkInitOncealso initializes GIN-related resource requirements, including inbox, outbox, accumulation buffer, and rail signal.

📎 src/sym_kernels.cc:208-251

c
    struct ncclDevResourceRequirements ginInboxRailReq = {};
    struct ncclDevResourceRequirements ginOutboxReq = {};
    struct ncclDevResourceRequirements rsGinAccumReq = {};
    struct ncclDevResourceRequirements railSignalReq = {};
    if (ncclParamSymGinKernelsEnable() && ncclTeamLsa(comm).nRanks < comm->nRanks) {
      int maxBlocks;
      size_t bufSize;
      getRequirements_gin(comm, &maxBlocks, &bufSize);

      maxBlocks = std::max(maxBlocks, comm->config.minCTAs);
      maxBlocks = std::min(maxBlocks, comm->config.maxCTAs);
      if (ncclParamSymCTAs() >= 1) maxBlocks = ncclParamSymCTAs();
      maxBlocks = std::min(maxBlocks, ncclSymkMaxBlocks);
      symk->maxGinInboxBlocks = maxBlocks;
      symk->kcomm.rsGinAccumBytesPerBlock = ncclSymkRsGinAccumBytesPerBlock();

      rsGinAccumReq.bufferSize = (size_t)maxBlocks * symk->kcomm.rsGinAccumBytesPerBlock;
      rsGinAccumReq.bufferAlign = 128;
      rsGinAccumReq.outBufferHandle = &symk->kcomm.rsGinAccumBuf;
      ...
      uint32_t railSignalCount = ncclTeamRail(comm).nRanks * ncclSymkMaxBlocks;
      ...
      reqs.barrierCount = ncclSymkMaxBlocks;
      reqs.ginConnectionType = NCCL_GIN_CONNECTION_RAIL;
      reqs.ginStrongSignalsRequired = true;
      reqs.ginVaSignalsRequired = true;
    }

getRequirements_ginuses the tuning model to compute the required number of blocks and buffer size, which are then clamped to the[minCTAs, maxCTAs]range.rsGinAccumBytesPerBlockis the accumulation buffer size per block, aligned to 128 bytes—the cache line size, to avoid false sharing.

mermaid
flowchart TD
    start["ncclSymkMask(comm, coll, red, ty, nElts)"] --> coll{"集合类型?"}
    coll -->|AllGather| mask_ag["kmask = kernelMask_AG"]
    coll -->|AllReduce| mask_ar["kmask = kernelMask_AR"]
    coll -->|ReduceScatter| mask_rs["kmask = kernelMask_RS"]
    mask_ag --> check_stmc{"hasLsaMultimem?"}
    mask_ar --> check_stmc
    mask_rs --> check_stmc
    check_stmc -->|否| clear_stmc["kmask &= ~kernelMask_STMC"]
    check_stmc -->|是| check_ldmc{"数据类型+归约支持LDMC?"}
    clear_stmc --> size_check
    check_ldmc -->|否| clear_ldmc["kmask &= ~kernelMask_LDMC"]
    check_ldmc -->|是| size_check
    clear_ldmc --> size_check
    size_check{"nBusBytes >= 2GB?"} -->|是| clear_ll["kmask &= ~kernelMask_LL"]
    size_check -->|否| tma_check
    clear_ll --> tma_check{"TMA可用且16B对齐?"}
    tma_check -->|否| clear_tma["kmask &= ~kernelMask_Tma"]
    tma_check -->|是| gin_check
    clear_tma --> gin_check{"需要GIN? LSA rank < 总rank"}
    gin_check -->|否| clear_gin["kmask &= ~kernelMask_Gin"]
    gin_check -->|是| done
    clear_gin --> done["返回 kmask"]

This diagram fully depicts the decision chain ofncclSymkMask: starting from the collective type, it passes through five filters in sequence—multicast support, data type, size boundary, TMA availability, and GIN requirement—and finally returns a bitmask. Each filter may eliminate a batch of kernels, which is exactly the embodiment of NCCL's "select the optimal kernel by scenario."

Production Pitfall Guide

Pitfall 1: Multicast silently fails when crossing cliques. hasLsaMultimemThe third condition of!comm->p2pCrossCliqueisncclNvlsSymmetricMultimemEnabled. If your cluster is configured with MNNVL (Multi-Node NVLink) but some ranks cross cliques, multicast will be disabled, and performance will silently degrade to the normal path. When troubleshooting, check the log output of

Pitfall 2: The implicit requirement of 16-byte alignment. ncclSymkMaskInif (!symAligned16B) kmask &= ~kernelMask_Tma;—if the user buffer is not 16-byte aligned, TMA kernels are removed. TMA is the fastest copy engine on Hopper/Blackwell, and losing it means a performance drop. In production environments, user-passed buffers often come fromcudaMalloc, which are naturally aligned; but if they come from a custom allocator or a slice, you may hit this pitfall.

Pitfall 3: The 2GB boundary.LL kernels use 32-bit indices, and are removed once the total bus bytes exceed 2GB. For large model training, the gradients of a single AllReduce may exceed this value, in which case NCCL automatically switches to the STMC or Simple protocol. This is not a bug, but if you manually specify the LL protocol, you will getncclInvalidArgument。

---

14.2 NVLS: Let NVSwitch Hardware Do the Reduction for You

Intuitive Model

Traditional AllReduce is "software reduction": each GPU sends data to its neighbor, the neighbor performs addition, and then forwards it—data is shuttled back and forth between GPUs, and the addition is executed on the SM. This is like 8 people passing notes to compute a sum, where each person has to read it, add it, and pass it on.

NVLS takes a different approach: the NVSwitch chip has built-inmulticast and reduction capabilitiesYou write the data to the multicast address, and NVSwitch automatically broadcasts it to all members and performs the addition in hardware. It's like 8 people writing numbers on the same whiteboard, and the whiteboard automatically displays the sum—the GPU only writes once and reads once, while all the intermediate movement and addition are handled entirely by the switch hardware.

Without NVLS, the bandwidth of intra-node AllReduce would be limited by the point-to-point links between GPUs, and the SM would have to spend a large number of cycles doing additions. NVLS offloads both of these to hardware, freeing the SM to do other computations.

Data Structures and Memory Layout

The core of NVLS is themulticast group (MC group)。ncclMcGroupThe struct describes the entire state of a multicast group.

📎 src/transport/multicast.cc:72-77

c
struct ncclMcGroup {
  CUmemGenericAllocationHandle handle;  // the MC object
  char* base;                          // mapped MC VA base
  size_t capacity;                      // total mapped VA size
  int dev;                           // local device, for unbind
};

Four fields:handleis the handle of the CUDA multicast object,baseis the base address of the multicast virtual address,capacityis the total mapping size,devis the local device number (used for unbinding). Note that there is no lock here—the creation and destruction of multicast groups happen during the initialization/destruction phase, not on the hot path.

The multicast group is divided into multiplepartitions, and each partition is an immutable slice.ncclMcPartitiondescribes a partition.

📎 src/transport/multicast.cc:162-170

c
  // A partition is self-sufficient for binds: it carries the group's handle, device and
  // bind granularity alongside its own extent.
  for (int i = 0; i < nRequests; i++) {
    if (outPartitions[i].size == 0) continue;
    outPartitions[i].ptr = group->base + outPartitions[i].offset;
    outPartitions[i].mcHandle = mcHandle;
    outPartitions[i].minGranularity = minGran;
    outPartitions[i].dev = comm->cudaDev;
  }

Each partition carries its ownoffset、size、ptr, as well as the owning group'smcHandle、minGranularity、dev. This "self-contained" design allows partitions to be passed independently to the bind function without needing to look up group information again.

Scenario-Driven Step-by-Step Walkthrough

Suppose 8 ranks want to establish an NVLS domain.ncclMcGroupBuildPartitionsis responsible for creating the multicast group and splitting partitions.

📎 src/transport/multicast.cc:79-121

c
ncclResult_t ncclMcGroupBuildPartitions(struct ncclComm* comm, const struct ncclMcRequest* requests, int nRequests,
                                        struct ncclMcGroup** outGroup, struct ncclMcPartition* outPartitions) {
  ...
  mcprop.numDevices = comm->localRanks;
  mcprop.handleTypes = ncclCuMemHandleType;
  mcprop.flags = 0;
  mcprop.size = 0;
  for (int i = 0; i < nRequests; i++) mcprop.size += requests[i].size;
  CUCHECKGOTO(cuMulticastGetGranularity(&recGran, &mcprop, CU_MULTICAST_GRANULARITY_RECOMMENDED), ret, fail);
  CUCHECKGOTO(cuMulticastGetGranularity(&minGran, &mcprop, CU_MULTICAST_GRANULARITY_MINIMUM), ret, fail);

  // Bump-allocate an immutable slice per request. Offsets and sizes are rounded
  // to the recommended granularity (a multiple of the MC minimum) so every slice
  // boundary is a valid bind offset.
  for (int i = 0; i < nRequests; i++) {
    outPartitions[i] = {};
    if (requests[i].size == 0) continue;
    size_t align = requests[i].alignment > recGran ? requests[i].alignment : recGran;
    ALIGN_SIZE(capacity, align);
    size_t slice = requests[i].size;
    ALIGN_SIZE(slice, recGran);
    outPartitions[i].offset = capacity;
    outPartitions[i].size = slice;
    capacity += slice;
  }

Step 1: Accumulate the sizes of all requests to get the total multicast group size. Step 2: Query CUDA's recommended granularity and minimum granularity—these are hardware constraints, and the address and size of the multicast object must be integer multiples of the granularity. Step 3: Bump allocation—carve out a block for each request, with offsets and sizes aligned to the recommended granularity.ALIGN_SIZE(capacity, align)ensures that the starting offset of each slice is a valid bind offset.

Next is cross-rank creation and import:

📎 src/transport/multicast.cc:125-146

c
  if (comm->localRank == 0) {
    NCCLCHECKGOTO(ncclMcCreate(comm, &mcprop, comm->localRank, comm->localRanks, &mcHandle, shareableHandle), ret,
                  fail);
    mcCreated = 1;
    NCCLCHECKGOTO(bootstrapIntraNodeBroadcast(comm->bootstrap, comm->localRankToRank, comm->localRank, comm->localRanks,
                                              0, shareableHandle, NVLS_HANDLE_SIZE),
                  ret, fail);
  } else {
    NCCLCHECKGOTO(bootstrapIntraNodeBroadcast(comm->bootstrap, comm->localRankToRank, comm->localRank, comm->localRanks,
                                              0, shareableHandle, NVLS_HANDLE_SIZE),
                  ret, fail);
    NCCLCHECKGOTO(ncclMcImport(comm, shareableHandle, comm->localRankToRank[0], &mcHandle), ret, fail);
    mcCreated = 1;
  }
  CUCHECKGOTO(cuMulticastAddDevice(mcHandle, comm->cudaDev), ret, fail);

  // cuMemMap of an MC object blocks until every device has been added. This
  // abort-aware barrier makes a peer failing before cuMulticastAddDevice trip the
  // abort flag here instead of stranding survivors in the blocking cuMemMap.
  NCCLCHECKGOTO(bootstrapIntraNodeBarrier(comm->bootstrap, comm->localRankToRank, comm->localRank, comm->localRanks,
                                          comm->localRankToRank[0]),
                ret, fail);

localRank 0 creates the multicast object, then broadcasts the shareable handle via bootstrap; other ranks receive the handle and import it.cuMulticastAddDeviceadds the local device to the multicast group. Note that barrier—the comment makes it very clear:cuMemMapblocks until all devices have joined, and if some peer fails beforecuMulticastAddDevice, the survivors will hang incuMemMap. This barrier allows failures to be captured by the abort flag before blocking.

Finally, mapping and access permission setup:

📎 src/transport/multicast.cc:148-155

c
  // Reserve and map the whole MC VA once; each consumer slice is a view into it.
  CUCHECKGOTO(cuMemAddressReserve(&base, capacity, recGran, 0U, 0), ret, fail);
  CUCHECKGOTO(cuMemMap(base, capacity, 0, mcHandle, 0), ret, fail);
  mapped = 1;
  desc.flags = CU_MEM_ACCESS_FLAGS_PROT_READWRITE;
  desc.location.type = CU_MEM_LOCATION_TYPE_DEVICE;
  desc.location.id = comm->cudaDev;
  CUCHECKGOTO(cuMemSetAccess(base, capacity, &desc, 1), ret, fail);

The entire multicast VA is reserved and mapped only once, and each consumer slice is a view of this VA. This is the "map once, slice many times" design—more resource-efficient than creating a separate multicast object for each consumer.

Concurrency Control and Hardware Interaction

Binding is the most critical operation in NVLS.ncclMcPartitionBindMembinds a UC (unicast) memory handle to a certain offset in the multicast group.

📎 src/transport/multicast.cc:200-225

c
ncclResult_t ncclMcPartitionBindMem(const struct ncclMcPartition* partition, size_t offsetInPartition,
                                    CUmemGenericAllocationHandle mem, size_t memOffset, size_t bindSize) {
  // A bind overrunning its partition would corrupt the next consumer's partition; fail
  // cleanly instead (possible when UC rounding exceeds the MC-rounded partition).
  if (offsetInPartition + bindSize > partition->size) {
    WARN("NVLS MC bind of size %zu at slice offset %zu exceeds slice size %zu (UC/MC granularity mismatch)", bindSize,
         offsetInPartition, partition->size);
    return ncclInternalError;
  }
  size_t mcOffset = partition->offset + offsetInPartition;
  ...
  CUresult err = CUPFN(cuMulticastBindMem(partition->mcHandle, mcOffset, mem, memOffset, bindSize, 0 /*flags*/));
  if (err != CUDA_SUCCESS) {
    ...
    WARN("Failed to bind NVLink SHARP (NVLS) Multicast memory of size %zu at MC group %llx offset %zu : CUDA error %d "
         "'%s'.\nThis is usually caused by a system or configuration error in the Fabric Manager or NVSwitches.\n"
         "Disable NVLS (NCCL_NVLS_ENABLE=0) if you wish to avoid this error in the future.",
         bindSize, partition->mcHandle, mcOffset, err, errStr);
    return ncclUnhandledCudaError;
  }
  return ncclSuccess;
}

The first line of defense is bounds checking:offsetInPartition + bindSize > partition->sizeand it errors out. The comment explains why—the granularity of UC memory may be larger than the MC partition, and if the UC alignment exceeds the boundary of the MC partition, it will step on the next consumer's partition. This is a typical "two granularities mismatch" trap.

cuMulticastBindMemis a hardware call, and the comment says it "blocks until all ranks have been added to the group"—this is where NVLS is most prone to problems. If Fabric Manager is misconfigured or there is an issue with the NVSwitch firmware, it will hang or return an error here. The error message directly suggests that the userNCCL_NVLS_ENABLE=0, which is the standard escape hatch in production environments.

There is also a "try bind" variant, used for user buffer registration:

📎 src/transport/multicast.cc:237-268

c
ncclResult_t ncclMcPartitionTryBindAddr(const struct ncclMcPartition* partition, size_t offsetInPartition,
                                        CUdeviceptr address, size_t bindSize, enum ncclMcBindStatus* outStatus) {
  const char* errStr = NULL;

  *outStatus = ncclMcBindStatusTransient;
  if (offsetInPartition + bindSize > partition->size) {
    ...
    return ncclInternalError;
  }
  size_t mcOffset = partition->offset + offsetInPartition;
  CUresult err = CUPFN(cuMulticastBindAddr(partition->mcHandle, mcOffset, address, bindSize, 0 /*flags*/));
  if (err == CUDA_SUCCESS) {
    *outStatus = ncclMcBindStatusOk;
    return ncclSuccess;
  }

  (void)pfn_cuGetErrorString(err, &errStr);
  // Only an outright rejection of the input is a property of the buffer. Anything else,
  // notably OUT_OF_MEMORY, may succeed later, so it must not be reported as permanent.
  if (err == CUDA_ERROR_INVALID_VALUE || err == CUDA_ERROR_NOT_SUPPORTED || err == CUDA_ERROR_NOT_PERMITTED) {
    *outStatus = ncclMcBindStatusNoSupport;
    ...
  } else {
    WARN("NVLS Multicast bind of size %zu at MC group %llx offset %zu dev %d failed transiently: CUDA error %d '%s'.\n"
         "The buffer is left unregistered for this operation and will be retried; repeated occurrences indicate "
         "sustained resource pressure.",
         bindSize, partition->mcHandle, mcOffset, partition->dev, err, errStr);
  }
  return ncclSuccess;
}

Here there is a subtle error classification:CUDA_ERROR_INVALID_VALUE、NOT_SUPPORTED、NOT_PERMITTEDis classified asncclMcBindStatusNoSupport—this is apermanent failure, indicating that this buffer itself does not support multicast binding. Other errors (especiallyOUT_OF_MEMORY) are classified asncclMcBindStatusTransient—this is atemporary failure, and can be retried. This distinction is crucial: if OOM is treated as a permanent failure, a registration that could have succeeded will be mistakenly abandoned; if a parameter error is treated as a temporary failure, it will be retried indefinitely.

Production Pitfall Avoidance Guide

Pitfall 1: Fabric Manager misconfiguration causescuMulticastBindMemto hang.This is the most classic production failure of NVLS. The error message explicitly points to Fabric Manager or NVSwitch. Troubleshooting steps: firstNCCL_NVLS_ENABLE=0to confirm the problem disappears, then check the Fabric Manager logs and NVSwitch firmware version.

Pitfall 2: UC/MC granularity mismatch. ncclMcPartitionBindMemThe bounds check in

will catch this problem, but if you see the "UC/MC granularity mismatch" warning, it means that the UC size of some request exceeds the MC partition after alignment. This usually happens when the request size is close to the granularity boundary. ncclMcGroupBuildPartitionsPitfall 3: Resource leak after multicast group creation failure.CUCALLThe fail path ofCUCHECK:

📎 src/transport/multicast.cc:179-184

c
fail:
  // Best-effort (CUCALL) so a failing cleanup op cannot skip releasing the MC handle.
  if (mapped) CUCALL(cuMemUnmap(base, capacity));
  if (base) CUCALL(cuMemAddressFree(base, capacity));
  if (mcCreated) CUCALL(cuMemRelease(mcHandle));
  return ret;

The comment explains the reason: if the cleanup operation itself fails, releasing the MC handle must not be skipped because of it—MC slots are a scarce resource, and a leak will cause subsequent creations to fail. This is a typical design of "the cleanup path must do its best."

mermaid
sequenceDiagram
    participant R0 as "Rank 0 (localRank=0)"
    participant R1 as "Rank 1..N-1"
    participant BS as "bootstrapIntraNode"
    participant CU as "CUDA Driver"

    R0->>CU: "cuMulticastCreate(mcHandle, prop)"
    CU-->>R0: "mcHandle"
    R0->>BS: "bootstrapIntraNodeBroadcast(shareableHandle)"
    BS-->>R1: "shareableHandle"
    R1->>CU: "cuMemImportFromShareableHandle(mcHandle)"
    CU-->>R1: "mcHandle"
    R0->>CU: "cuMulticastAddDevice(mcHandle, cudaDev)"
    R1->>CU: "cuMulticastAddDevice(mcHandle, cudaDev)"
    R0->>BS: "bootstrapIntraNodeBarrier()"
    R1->>BS: "bootstrapIntraNodeBarrier()"
    Note over R0,R1: "barrier 防止 cuMemMap 阻塞时 peer 失败"
    R0->>CU: "cuMemAddressReserve(base, capacity)"
    R0->>CU: "cuMemMap(base, capacity, mcHandle)"
    R0->>CU: "cuMemSetAccess(base, capacity, desc)"
    R0->>CU: "cuMulticastBindMem(mcHandle, mcOffset, ucHandle)"
    CU-->>R0: "绑定完成,硬件多播就绪"

This sequence diagram depicts the complete flow of a multicast group from creation to binding. The key point is that barrier—it decouples "peer failure" from "cuMemMap blocking," preventing survivors from getting stuck.

---

14.3 The Combination of Symmetric Memory and NVLS: How LSA Pointers Are Resolved on the Device Side

Intuitive Model

Symmetric memory solves the "address consistency" problem, and NVLS solves the "hardware reduction" problem. But for the two to truly work together, a key mechanism is still needed:How does the device side know that a certain address is symmetric and can take the multicast path?

The answer lies in the LSA (Load-Store Accessible) pointer. LSA is short for "Load-Store Accessible," meaning that the memory this pointer points to can be directly accessed by the GPU using ordinary load/store instructions—regardless of whether it is physically local or remote. If the address falls within a multicast group, the load/store will be intercepted and broadcast by the NVSwitch hardware.

Data Structures and Memory Layout

ncclSymkDevWorkIt is the device-side work descriptor, and it carries the key information of symmetric memory.

📎 src/sym_kernels.cc:380-393

c
ncclResult_t ncclSymkMakeDevWork(struct ncclComm* comm, struct ncclTaskColl* task, struct ncclSymkDevWork* outDevWork) {
  outDevWork->rootRank = task->root;
  outDevWork->redOpArg = task->opDev.scalarArg;
  outDevWork->nElts = task->count;
  outDevWork->inputWin = task->sendWin ? task->sendWin->vidmem : nullptr;
  outDevWork->inputOff =
    task->sendWin ? (uint8_t*)task->sendbuff - (uint8_t*)task->sendWin->userPtr : (size_t)task->sendbuff;
  outDevWork->outputWin = task->recvWin ? task->recvWin->vidmem : nullptr;
  outDevWork->outputOff =
    task->recvWin ? (uint8_t*)task->recvbuff - (uint8_t*)task->recvWin->userPtr : (size_t)task->recvbuff;
  outDevWork->sChannelId = 0xffff;
  outDevWork->nChannels = 0;
  return ncclSuccess;
}

inputWinIt is the device-side virtual address of the window (vidmem),inputOffIt is the offset of the buffer within the window. After the device-side kernel obtains these two values, it computesinputWin + inputOffto get the actual address. If this address falls within the multicast group, the hardware will automatically handle the broadcast.

ncclSymkInitOnceIt also sets up the LSA barrier and LLA2A (Low-Latency All-to-All) resources.

📎 src/sym_kernels.cc:197-206

c
    reqs.lsaBarrierCount = ncclSymkMaxBlocks;
    reqs.ginStrongSignalsRequired = false;
    reqs.ginVaSignalsRequired = false;

    struct ncclDevResourceRequirements lla2aReq;
    ncclLLA2ACreateRequirement(ncclSymkMaxBlocks,
                               ncclLLA2ACalcSlots(ncclTeamLsa(comm).nRanks * ncclSymkMaxThreads, ncclSymkLLMaxEltSize),
                               &symk->kcomm.lsaLLA2A, &lla2aReq);
    lla2aReq.next = reqs.resourceRequirementsList;
    reqs.resourceRequirementsList = &lla2aReq;

lsaBarrierCountSet toncclSymkMaxBlocks—one barrier slot per block. LLA2A is short for low-latency all-to-all, used for fast data exchange within the LSA domain.ncclLLA2ACalcSlotsThe required number of slots is calculated based on the number of ranks, the number of threads, and the maximum element size.

Scenario-Driven Step-by-Step Walkthrough

Suppose an AllReduce usesAllReduce_AGxLLMC_Rkernel (AllGather + LL + MC + Reduce). The workflow of this kernel is:

1. AllGather Phase: Each rank writes its own data into the multicast group, and the NVSwitch hardware broadcasts it to all ranks.

2. Reduce Phase: Each rank reads the data of all ranks from the multicast group and performs the reduction locally.

ncclSymkMaskIt will check whether this kernel is available.kernelMask_LLIt includesAllReduce_AGxLLMC_R, but only ifhasLsaMultimemis true (otherwisekernelMask_STMCis cleared, andAllReduce_AGxLLMC_Rbelongs to the STMC set).

Wait, there is a detail here:kernelMask_STMCDoes it includeAllReduce_AGxLLMC_R? Look at the source code:

📎 src/sym_kernels.cc:17-21

c
constexpr uint32_t kernelMask_STMC =
  1 << ncclSymkKernelId_AllGather_LLMC | 1 << ncclSymkKernelId_AllGather_STMC |
  1 << ncclSymkKernelId_AllGather_TmaSTMC | 1 << ncclSymkKernelId_AllReduce_AGxLLMC_R |
  1 << ncclSymkKernelId_AllReduce_RSxLDMC_AGxSTMC | 1 << ncclSymkKernelId_ReduceScatter_LDMC |
  1 << ncclSymkKernelId_AllGather_RailRing_LsaSTMC;

Yes,AllReduce_AGxLLMC_Ris inkernelMask_STMC. So ifhasLsaMultimemis false, this kernel will be filtered out. This explains why symmetric memory and NVLS must work together—without multicast, all MC-series kernels are unavailable.

After the device side obtainsncclSymkDevWork, it will, based oninputWinandinputOffcompute the address. If the address is within the multicast group, the load/store instruction will be intercepted by the NVSwitch. This is the resolution process of the LSA pointer:No software translation is needed; the hardware automatically determines it based on the address range.。

Concurrency Control and Hardware Interaction

The synchronization mechanism of NVLS relies oncredit。ncclNvlsSetupThe credit partition is initialized in

📎 src/transport/nvls.cc:407-447

c
    int nChannels = comm->nvlsChannels;
    size_t creditSize = nChannels * 2 * memSize * nHeads;
    int nvlsStepSize = comm->nvlsChunkSize;

    NCCLCHECKGOTO(ncclCalloc(&comm->nvlsResources, 1), res, fail);
    comm->nvlsResources->inited = false;
    comm->nvlsResources->refCount = 1;
    comm->nvlsResources->nChannels = nChannels;
    comm->nvlsResources->nHeads = nHeads;
    comm->nvlsResources->chunkSize = comm->nvlsChunkSize;
    comm->nvlsResources->treeMaxChunkSize = comm->nvlsTreeMaxChunkSize;
    resources = comm->nvlsResources;

    for (int c = 0; c < nChannels; c++) {
      NCCLCHECKGOTO(initNvlsChannel(comm, c, NULL, false), res, fail);
    }

    memset(&resources->accessDesc, 0, sizeof(resources->accessDesc));
    resources->accessDesc.flags = CU_MEM_ACCESS_FLAGS_PROT_READWRITE;
    resources->accessDesc.location.type = CU_MEM_LOCATION_TYPE_DEVICE;
    resources->accessDesc.location.id = comm->cudaDev;
    resources->dev = comm->cudaDev;

    // Build the single shared MC group for this NVLS domain. The data slice is
    // reserved here but bound later by ncclNvlsBufferSetup.
    {
      size_t buffSize = nvlsStepSize * NCCL_STEPS;
      size_t dataSize = nChannels * 2 * buffSize * nHeads;
      size_t ubSize = ncclNvlsUbSize(comm);
      struct ncclMcRequest requests[3] = {{creditSize, 0}, {dataSize, 0}, {ubSize, 0}};
      struct ncclMcPartition partitions[3];
      NCCLCHECKGOTO(ncclMcGroupBuildPartitions(comm, requests, 3, &resources->mcGroup, partitions), res, fail);
      resources->creditPartition = partitions[0];
      resources->dataPartition = partitions[1];
      if (ubSize) {
        resources->ubPartition = partitions[2];
        NCCLCHECKGOTO(ncclMcArenaInit(comm, &resources->ubArena, &resources->ubPartition), res, fail);
        resources->ubEnabled = true;
      }
      NCCLCHECKGOTO(nvlsAllocBindUc(comm, &resources->creditPartition, creditSize, &resources->creditUc), res, fail);
    }

The multicast group is divided into three partitions:creditPartition(credit),dataPartition(data),ubPartition(user buffer). The credit partition is used for synchronization—each channel has independent head/tail pointers, shared through the multicast group.

The initialization of credit is in the later loop:

📎 src/transport/nvls.cc:456-491

c
    for (int h = 0; h < nHeads; h++) {
      int nvlsPeer = comm->nRanks + 1 + h;
      for (int c = 0; c < nChannels; c++) {
        struct ncclChannel* channel = comm->channels + c;
        char* mem = NULL;
        struct ncclChannelPeer* peer = channel->peers[nvlsPeer];

        // Reduce UC -> MC
        mem = (char*)resources->creditUc.ptr + (h * 2 * nChannels + c) * memSize;
        peer->send[1].transportComm = &nvlsTransport.send;
        peer->send[1].conn.buffs[NCCL_PROTO_SIMPLE] = NULL;
        peer->send[1].conn.head = (uint64_t*)mem;
        peer->send[1].conn.tail = (uint64_t*)(mem + memSize / 2);
        peer->send[1].conn.stepSize = nvlsStepSize;
        mem = (char*)resources->creditPartition.ptr + (h * 2 * nChannels + c) * memSize;
        peer->recv[0].transportComm = &nvlsTransport.recv;
        peer->recv[0].conn.buffs[NCCL_PROTO_SIMPLE] = NULL;
        peer->recv[0].conn.head = (uint64_t*)mem;
        peer->recv[0].conn.tail = (uint64_t*)(mem + memSize / 2);
        peer->recv[0].conn.stepSize = nvlsStepSize;
        peer->recv[0].conn.flags |= NCCL_NVLS_MIN_POLL;

Each combination of head and channel has an independent credit region.headandtailare 64-bit pointers,memSizeis 64 bytes (size_t memSize = 64;), so head and tail each occupy 32 bytes—exactly half a cache line.NCCL_NVLS_MIN_POLLThe flag lets the receiver use the minimum polling mode, reducing CPU overhead.

Production Pitfall Avoidance Guide

Pitfall 1: head/tail contention in the credit partition.Multiple channels share the same multicast group, but each channel has an independent credit region. If the number of channels is configured improperly (for example,nvlsCTAsis set too large), the credit region will expand and occupy precious multicast address space.ncclNvlsChannelsThe number of channels is automatically adjusted based on the GPU architecture and the number of nodes:

📎 src/transport/nvls.cc:100-133

c
  if (comm->config.nvlsCTAs != NCCL_CONFIG_UNDEF_INT) {
    channels = comm->config.nvlsCTAs;
  } else if (channels == 0 && comm->compCap >= 100) {
    // Use a reduced number of channels for single node/MNNVL domain on Blackwell and above.
    // comm->nNodes is not yet initialized at this point so we need to use local information.
    bool multiNode = false;
    if (comm->MNNVL) {
      multiNode = (comm->clique.size < comm->nRanks);
    } else {
      int i;
      for (i = 1; i < comm->nRanks; i++) {
        if (comm->peerInfo[i].hostHash != comm->peerInfo[0].hostHash) break;
      }
      multiNode = (i < comm->nRanks);
    }
    if (multiNode) {
      channels = RUBIN_AND_LATER(comm->compCap) ? /*RUBIN=*/64 : /*SM100=*/32;
    } else {
      channels = RUBIN_AND_LATER(comm->compCap) ? /*RUBIN=*/48 : /*SM100=*/24;
    }
  } else if (channels == 0) {
    channels = /*SM90=*/16;
  }

Note thatcomm->nNodeshas not been initialized at this stage, so the code usespeerInfo[i].hostHashto manually determine whether it is multi-node. This is a classic trap in initialization order—you cannot rely on a field that has not yet been computed.

Pitfall 2: MNNVL does not support NVLS buffer registration. 📎 src/transport/nvls.cc:516-517

c
  // MNNVL does not support NVLS buffer registration
  if (!comm->MNNVL && comm->nvlsResources->nvlsShmemHandle == NULL) {

In an MNNVL (Multi-Node NVLink) environment, user buffer registration is skipped. If your cluster is MNNVL and relies on UB registration to improve performance, you will find that the registration does not take effect. This is a hardware limitation, not a bug.

Pitfall 3: Reference counting for shared resources. ncclNvlsSetupSupports parent-child communicator sharing of NVLS resources:

📎 src/transport/nvls.cc:380-392

c
  if (nvlsShare) {
    /* reuse NVLS resources */
    comm->nvlsChannels = std::min(comm->nvlsChannels, parent->nvlsResources->nChannels);
    /* Inherit chunk sizes from the shared resource since we're reusing the parent's
     * NVLS buffers, which were allocated and laid out based on these values. */
    comm->nvlsChunkSize = parent->nvlsResources->chunkSize;
    comm->nvlsTreeMaxChunkSize = parent->nvlsResources->treeMaxChunkSize;
    for (int c = 0; c < comm->nvlsChannels; c++) {
      NCCLCHECKGOTO(initNvlsChannel(comm, c, parent, true), res, fail);
    }

    comm->nvlsResources = parent->nvlsResources;
    ncclAtomicRefCountIncrement(&parent->nvlsResources->refCount);
  }

The child communicator reuses the parent communicator's resources, incrementing the reference count by one.ncclNvlsFreeThe resource is only truly released when the reference count drops to zero. If reference counting is mismanaged, it can lead to premature resource release or leaks. Note thatnvlsChunkSizeandnvlsTreeMaxChunkSizemust inherit the parent communicator's values—because buffers are laid out according to these values, and changing them would cause address calculation errors.

mermaid
flowchart LR
    subgraph host["Host 侧"]
        task["ncclTaskColl<br/>sendbuff/recvbuff"]
        devwork["ncclSymkDevWork<br/>inputWin + inputOff"]
        task -->|"ncclSymkMakeDevWork"| devwork
    end
    subgraph device["Device 侧"]
        kernel["SymKernel<br/>load/store"]
        lsa{"地址在多播组内?"}
        devwork --> kernel
        kernel --> lsa
    end
    subgraph hw["NVSwitch 硬件"]
        mc["多播组<br/>MC group"]
        reduce["硬件归约<br/>Reduction"]
        lsa -->|"是"| mc
        lsa -->|"否"| local["本地显存<br/>UC memory"]
        mc --> reduce
        reduce -->|"广播结果"| kernel
    end

This data flow diagram shows the complete chain from host-side tasks to device-side execution. The key branch islsa{"地址在多播组内?"}—if yes, it goes through NVSwitch hardware multicast and reduction; if no, it goes through local memory. This determination is made automatically by the hardware based on the address range, requiring no software intervention.

---

14.4 Design Reflection: Why Symmetric Memory Reduces Small Message Latency

Returning to the core question at the beginning of this chapter: why can symmetric memory significantly reduce small message latency?

First, it eliminates address translation overhead.In traditional communication, each rank accessing a peer's buffer must look up a table and calculate offsets. Symmetric memory lets all ranks use the same set of addresses, and the device-side kernel can directly computebase + offsetFor small messages, the overhead of this translation is proportionally very high.

Second, it eliminates control message round trips.Traditional communication requires exchanging control information such as "which buffer of yours do I want to write to." With symmetric memory, addresses are pre-agreed upon and no runtime negotiation is needed.

Third, it makes hardware multicast possible.Only when addresses are symmetric can NVSwitch use the same set of addresses for multicast. If each rank has different addresses, the hardware cannot know where to broadcast.

Fourth, it reduces the SM's reduction burden.NVLS offloads addition to NVSwitch, so the SM only needs to issue one write and one read. For small messages, the SM's instruction overhead is the main source of latency.

The combination of these four factors reduces small message latency from "microsecond-level" to "sub-microsecond-level."

[Design Inference and Architectural Trade-offs]

From an engineering perspective, the design of symmetric memory embodies a core philosophy of NCCL:Push complexity to the initialization phase, keeping the hot path as simple as possible.Address negotiation, multicast group creation, and credit allocation are all completed at initialization time, and the runtime kernel only needs to perform the simplest address calculations and load/store operations. This "heavy initialization, light runtime" design is a common pattern in high-performance communication libraries.

---

Chapter Summary

This chapter dissected the two pillars of NCCL intra-node communication:

1. Symmetric Memory: ThroughncclSymkInitOnceandncclSymkMaskestablish buffers with consistent addresses, allowing each rank to access all ranks' data using the same set of addresses.ncclSymkMakeDevWorktranslates host-side tasks into device-side work items,inputWin + inputOffis the core formula for address resolution.

2. NVLS Multicast: ThroughncclMcGroupBuildPartitionscreate multicast groups,ncclMcPartitionBindMembinds UC memory to multicast groups,cuMulticastBindMemis the hardware call. The multicast group is divided into three partitions—credit, data, and ub—used for synchronization, data transfer, and user buffer registration respectively.

3. LSA Pointer Resolution: The device side automatically determines whether to use the multicast path based on the address range, requiring no software translation.NCCL_NVLS_MIN_POLLThe flag optimizes polling overhead.

4. Error Handling:ncclMcPartitionTryBindAddrdistinguishes permanent failures from transient failures,ncclMcGroupBuildPartitionsthe fail path usesCUCALLto ensure resource release.

Chapter Review Questions

Q1: If the boundary check inncclMcPartitionBindMemis removed,if (offsetInPartition + bindSize > partition->size)under what scenarios would an out-of-bounds memory access be triggered? Why can't this check be replaced by "UC and MC have the same granularity"?

Reference Analysis: See📎 src/transport/multicast.cc:200-208:

CHAPTER 15

Chapter 15: Chapter 15: RMA and GIN: The Evolution of Remote Memory Access and GPU Direct Communication

Official Source: NVIDIA/nccl · Version: Commit @12df1a11 · Book Progress: Chapter 15 / 25

Chapter 15: RMA and GIN: The Evolution of Remote Memory Access and GPU Direct Communication

In the previous chapter, we saw how symmetric memory allows each rank to access all ranks' buffers using the same set of addresses, and how NVLS leverages NVSwitch's multicast capability to push hardware-accelerated reduction to the extreme. But collective communication is not everything—when applications need point-to-point remote memory operations, or want GPU kernels to directly initiate network requests, RMA and GIN come into play. RMA provides put/get semantics for remote memory access, while GIN allows the GPU to bypass host proxy threads and interact directly with the network. This chapter follows the order of "RMA first, then GIN," dissecting layer by layer the data structures, scheduling logic, concurrency control, and production pitfalls of these two mechanisms.

RMA's Dual-Channel Model: The Division of Labor Between CE and Proxy

Intuitive Model

Imagine a cross-border courier system: intra-city deliveries (ranks reachable via LSA) can be delivered directly by local delivery vehicles, while cross-city deliveries (ranks not reachable via LSA) must be handed off to air freight forwarders. NCCL's RMA is exactly this model—the same put operation, depending on whether the target rank is within the LSA (Load-Store Accessible) team, is routed to two completely different execution paths: the CE (Copy Engine) path and the Proxy (proxy thread) path.

Without this splitting mechanism, all RMA operations would go through the proxy thread, so even intra-node puts would have to be relayed through a host thread, needlessly adding a host-device round-trip latency. Conversely, if all operations went through CE, cross-node operations could not leverage the asynchronous capabilities of the network plugin.

Data Structures and Memory Layout

The core scheduling structure for RMA isncclRmaArgs, which records the splitting result of RMA tasks within a plan. Key fields include:

FieldMeaning
funcOperation type (PutSignal / Signal / WaitSignal)
nRmaTasksTotal task count
nRmaTasksProxyNumber of tasks going through the proxy path
nRmaTasksCeNumber of tasks going through the CE path

Each plan internally maintains two intrusive queues:rmaTaskQueueCeandrmaTaskQueueProxy, which respectively hold the tasks for the two paths.📎 src/rma/rma.cc:166-171

The logic for determining whether a rank is LSA-reachable is straightforward—iterate over thelsaRankListarray and perform a linear search.📎 src/rma/rma.cc:34-41This lookup is performed once per peer during task scheduling, with complexity O(lsaSize), and for typical small-scale LSA teams (usually 2-8 ranks) the overhead is negligible.

Step-by-Step Scheduling Flow

When the application calls an RMA put operation, the task entersplanner->rmaTaskQueues[ctx]。scheduleRmaTasksToPlan, which is responsible for distributing tasks from the queue into plans.📎 src/rma/rma.cc:141-296

Step 1: Find the first non-empty context queue. NCCL supports multiple RMA contexts (configured bynumRmaCtx), each with its own independent queue.📎 src/rma/rma.cc:148-155

Step 2: Take out the first task and determine the operation type. If it is WaitSignal, follow the special splitting logic; if it is Put/Signal, follow the batch merging logic.📎 src/rma/rma.cc:163-168

For WaitSignal tasks, the scheduler needs to split the peers list into two groups based on LSA reachability: the CE group and the Proxy group.📎 src/rma/rma.cc:187-204After splitting, two newncclTaskRmastructures are created respectively, each holding the peers array for the corresponding group.📎 src/rma/rma.cc:207-246The original task is released.📎 src/rma/rma.cc:251

For Put/Signal tasks, the logic is more complex—the scheduler iterates over the queues of all contexts, pulling all consecutive put/signal tasks into the same plan until it encounters a WaitSignal, at which point it stops.📎 src/rma/rma.cc:279-295The purpose of this design is clearly stated in the comments: let a single kernel launch cover the put/signal of all contexts, so that the proxy can issue all asynchronous requests at once before any blocking operation, while the CE path submits the copies and signals of all contexts in a batch.📎 src/rma/rma.cc:270-278

mermaid
flowchart TD
    start["scheduleRmaTasksToPlan(comm, plan)"]
    find_ctx{"找到非空 ctx 队列?"}
    no_task["返回 ncclSuccess"]
    dequeue["取出 firstTask"]
    check_func{"firstTask->func == WaitSignal?"}
    ws_split["按 isLsaAccessible 拆分 peers"]
    ws_ce{"npeersCe > 0?"}
    ws_proxy{"npeersProxy > 0?"}
    ws_ce_task["创建 CE WaitSignal 任务"]
    ws_proxy_task["创建 Proxy WaitSignal 任务"]
    ws_free["释放原始 firstTask"]
    put_check{"firstTask 的 peer LSA 可达?"}
    put_ce["入队 rmaTaskQueueCe"]
    put_proxy["入队 rmaTaskQueueProxy"]
    batch_loop["遍历所有 ctx 队列, 拉取连续 put/signal"]
    batch_check{"isRmaPutOrSignal(task->func)?"}
    batch_route{"isLsaAccessible(comm, task->peer)?"}
    batch_ce["入队 CE, nRmaTasksCe++"]
    batch_proxy["入队 Proxy, nRmaTasksProxy++"]
    done["记录 INFO 日志, 返回"]

    start --> find_ctx
    find_ctx -->|否| no_task
    find_ctx -->|是| dequeue
    dequeue --> check_func
    check_func -->|是| ws_split
    ws_split --> ws_ce
    ws_ce -->|是| ws_ce_task
    ws_ce -->|否| ws_proxy
    ws_ce_task --> ws_proxy
    ws_proxy -->|是| ws_proxy_task
    ws_proxy -->|否| ws_free
    ws_proxy_task --> ws_free
    ws_free --> done
    check_func -->|否| put_check
    put_check -->|是| put_ce
    put_check -->|否| put_proxy
    put_ce --> batch_loop
    put_proxy --> batch_loop
    batch_loop --> batch_check
    batch_check -->|否, 遇到 WaitSignal| done
    batch_check -->|是| batch_route
    batch_route -->|是| batch_ce
    batch_route -->|否| batch_proxy
    batch_ce --> batch_loop
    batch_proxy --> batch_loop

Parallel Execution and Stream Synchronization

After scheduling is complete,ncclLaunchRmadispatches tofuncorncclRmaPutbased on thencclRmaWaitSignal。📎 src/rma/rma.cc:109-131

field.ncclRmaPutTaking📎 src/rma/rma.cc:80-96as an example, when both proxy and CE tasks exist in a plan, the two paths need to execute in parallel. NCCL's approach is: record an event on the input stream, have the CE stream wait on this event, then launch operations on both streams simultaneously, and finally record another event on the CE stream, having the input stream wait on it.

This event chain ensures that: CE operations do not start before the input stream's dependencies are ready, and subsequent operations on the input stream do not start before CE completes.📎 src/rma/rma.cc:97-101

If there are only proxy tasks or only CE tasks, the corresponding operation is launched directly on the input stream, with no additional stream synchronization needed.

Design Considerations and Production Pitfalls isLsaAccessiblePitfall 1: The static nature of LSA reachability determination.comm->devrState.lsaRankListAt scheduling time,

is queried, and this list no longer changes after the communication domain is initialized. If the topology changes during runtime (for example, NVLink failure degradation), the LSA list will not update automatically, which may cause operations that should go through proxy to still take the CE path, triggering an unrecoverable error.Pitfall 2: FIFO guarantee of batch merging.📎 src/rma/rma.cc:283The batch merging logic only pulls consecutive put/signal tasks and stops when it encounters a WaitSignal.

This guarantees FIFO order within each context, but tasks across contexts may be merged into the same plan. If the application relies on operation order across contexts, it needs to explicitly use WaitSignal to establish a barrier.Pitfall 3: Memory leak paths.npeersProxy == 0In the WaitSignal branch, ifpeersProxy、nsignalsProxy、signalIdxsProxy, the code releases📎 src/rma/rma.cc:239-244three arrays.npeersCe == 0But ifnpeersProxy > 0,peersCeandncclMemoryStackAllocand other arrays are allocated via📎 src/rma/rma.cc:176-178This asymmetry can easily confuse readers, but it is actually correct—the stack-allocated memory is managed bycomm->memScopedin a unified manner.

RMA Proxy Context: Signals, Queues, and Lock-Free Ring Buffers

Intuitive Model

The proxy context is like a "post office sorting center": the GPU places packages to be sent (put requests) into the inbox (ring buffer), the proxy thread takes packages out of the inbox and hands them to the courier company (network plugin), and the courier company stamps the receipt (signal) after delivery. Throughout this process, the GPU and the proxy thread communicate through lock-free data structures, avoiding expensive lock contention.

Data Structures and Memory Layout

ncclRmaProxyCtxis the host structure of the proxy context, and its core fields include:

Signal region (signalsDev): a block of memory allocated on the GPU, with sizenRanks * numRmaSig * sizeof(uint64_t)。📎 src/rma/rma_proxy.cc:120-123Each rank hasnumRmaSigsignal slots, used to receive signals from that rank. When this block of memory is registered with the network plugin, it carriesNCCL_NET_MR_FLAG_FORCE_SO(force strong ordering) andNCCL_NET_MR_FLAG_SIGNAL_NEVER_RESET(signals are never reset) flags.📎 src/rma/rma_proxy.cc:125-127The strong ordering flag ensures the ordering relationship between put and signal—if put is issued before signal, the network must guarantee that signal is written only after the put data arrives.

Sequence number region (opSeqs/readySeqs/doneSeqs): one group per rank, allocated throughallocMemCPUAccessibleand may be GDR (GPU Direct RDMA) memory or ordinary host memory.📎 src/rma/rma_proxy.cc:132-137These three sequence numbers respectively track: the sequence number of submitted operations, the sequence number of ready operations, and the sequence number of completed operations.

Lock-free ring buffers (circularBuffers): an array of pointers with sizenRanks * queueSizewith one independent ring queue per rank.📎 src/rma/rma_proxy.cc:163-164The accompanyingpis(Producer Index) andcis(Consumer Index) arrays each havenRankselements.📎 src/rma/rma_proxy.cc:165-166The queue size must be a power of 2, so that index wraparound can use bitwise AND& (queueSize - 1)instead of modulo.📎 src/rma/rma_proxy.cc:156-160

InProgress queue: one intrusive linked list per peer, storing descriptors that have been submitted to the network plugin but have not yet completed.📎 src/rma/rma_proxy.cc:170-175This is a single-consumer queue, accessed only by the proxy thread, and requires no atomic operations.

Step-by-Step: From Context Creation to Progress Advancement

Context Creation:ncclRmaProxyCreateContextFirst, create the network context through the RMA plugin.📎 src/rma/rma_proxy.cc:229Then callncclRmaProxyCtxAllocto allocate resources such as signals, sequence numbers, and ring buffers.📎 src/rma/rma_proxy.cc:231Next, callncclRmaProxyCtxAllocGraphto allocate the resources required for graph capture mode—CPU-accessible signals, flush buffers, and persistent queues.📎 src/rma/rma_proxy.cc:232

Graph capture mode exists because CUDA Graph requires all operations to be replayable. In normal mode, signals are in GPU memory and the proxy reads them through GDR; in graph capture mode, signals are in CPU-accessible memory and the proxy can read and write them directly, avoiding the nondeterminism of GDR.📎 src/rma/rma_proxy.cc:184-190

Progress Thread:ncclRmaProxyProgressThreadis the main loop of the proxy.📎 src/rma/rma_proxy.cc:354-389It decides its behavior based on thermaProgressstate word:

  • rmaProgress == 1: normal progress mode, iterating over all proxy contexts and callingncclRmaProxyProgress。📎 src/rma/rma_proxy.cc:361-372
  • rmaProgress == 2: pause mode, used for resource reclamation. After the thread confirms the pause, it waits on a condition variable.📎 src/rma/rma_proxy.cc:373-378
  • rmaProgress == -1: exit signal, and the thread returns.📎 src/rma/rma_proxy.cc:379-380
  • rmaProgress == 0: idle wait.📎 src/rma/rma_proxy.cc:381-382

IfncclRmaProxyProgressreturns an error, the thread writes the error code intoasyncResult, setsrmaProgress = -2, and then exits.📎 src/rma/rma_proxy.cc:365-369This error code will be read by the main thread in a subsequentncclCommGetAsyncErrorcall.

Concurrency Control and Memory Ordering

The concurrency model of the RMA proxy is "single-producer-single-consumer": the GPU kernel is the producer, and the proxy thread is the consumer. The PI of the ring buffer is updated by the GPU, and the CI is updated by the proxy. Because it is single-producer-single-consumer, no CAS operation is needed, only correct memory ordering.

The strong ordering flag of the signal regionNCCL_NET_MR_FLAG_FORCE_SOis key.📎 src/rma/rma_proxy.cc:127Without this flag, the network plugin may reorder put and signal, causing the receiver to see the signal before the data arrives and read stale data.

NCCL_NET_MR_FLAG_SIGNAL_NEVER_RESETThe flag tells the network plugin: once a signal is written, it will not be reset.📎 src/rma/rma_proxy.cc:127This allows the plugin to optimize the signal write path—there is no need to clear it before each write.

Production Pitfalls

Pitfall 1: The queue size is not a power of 2.If the user sets a value that is not a power of 2 throughNCCL_RMA_PROXY_QUEUE_SIZEthe code falls back to the default value and prints an INFO log.📎 src/rma/rma_proxy.cc:156-159This fallback is silent (only INFO level), and is easily overlooked in production environments. If the user expects a larger queue to absorb burst traffic but the default value is actually used, backpressure may result.

Pitfall 2: The fallback chain when DMA-BUF registration fails. ncclRmaProxyRegMrSymThere are three layers of fallback for registering CUDA memory: first try DMA-BUF in DataDirect mode, then try non-DataDirect DMA-BUF after failure, and only fall back to ordinaryregMrSym。📎 src/rma/rma_proxy.cc:76-108The comments specifically warn: if one MR enters the non-DataDirect path, all other MRs must do the same; mixing them will break GIN's ordering guarantees.📎 src/gin/gin_host_proxy.cc:429-430This constraint is not explicitly checked in the RMA path, making it a potential hidden risk.

Pitfall 3: Delayed error propagation in the progress thread.WhenncclRmaProxyProgressreturns an error, the thread setsasyncResultand exits.📎 src/rma/rma_proxy.cc:366-369But the main thread may be executing a long-running kernel and will not immediately checkasyncResult. During this period, subsequent RMA operations will continue to be enqueued but will not be processed until the main thread discovers the error. This is the inherent delay of asynchronous error propagation, and the application needs to periodically callncclCommGetAsyncErrorto shorten this window.

GIN Architecture: GPU Directly Initiates Network Requests

Intuitive Model

In the traditional model, for the GPU to send network data, it must go through the path "GPU → host memory → proxy thread → NIC." The goal of GIN (GPU-Initiated Networking) is to let the GPU directly write to the NIC's send queue, just as the CPU directly writes to the NIC's MMIO registers. This requires the NIC to support doorbell writes initiated by the GPU, as well as a communication protocol between the GPU and the proxy thread.

Data Structures and Memory Layout

The core data structure of GIN isginProxyHostGpuCtx, which represents a GPU-host communication context:

FieldTypeMeaning
queuesncclGinProxyGfd_t*GFD queue, sizenRanks * queueSize
pisuint32_t*Producer index (written by GPU)
cisuint32_t*Consumer index (written by proxy)
cisShadowuint32_t*Shadow copy of CI (proxy-local)
sisuint32_t*Seen index (proxy-local)
statesginProxyGfdState*Status of each GFD slot
inlinesuint64_t*Inline data buffer

A GFD (GIN Forwarding Descriptor) is a request descriptor written by the GPU to the proxy. Each GFD consists of multiple qwords, containing the operation type, source address, destination address, size, signal information, and so on.📎 src/gin/gin_host_proxy.cc:158-163

queuesThere is a key detail in the memory allocation of the array: it is allocated viaallocMemCPUAccessible, but theforceHost=trueparameter is passed in.📎 src/gin/gin_host_proxy.cc:564This means the queue itself is in host memory, and the GPU writes to it via PCIe. Whereas thecisarray is allocated in GPU-accessible memory (possibly GDR), because the proxy needs to update it frequently.📎 src/gin/gin_host_proxy.cc:565-566

cisShadowandsisare local copies of the proxy thread, avoiding the need to readcis。📎 src/gin/gin_host_proxy.cc:44-47which may be located in GPU memory every time. Only whencisShadowadvances arecis。

Step-by-Step: Polling and Processing of GFDs

ncclGinProxyProgressis the main loop of the GIN proxy.📎 src/gin/gin_host_proxy.cc:648-669

Step 1: For each context, first callproxyGinPollCompletionsto check the completion status of submitted requests.📎 src/gin/gin_host_proxy.cc:653

Step 2: For each target rank, poll GFDs in batches.pollBatchcontrols the maximum number of GFDs processed each time.📎 src/gin/gin_host_proxy.cc:654-655

Step 3:proxyGinPollGfdCheck whether there is a new GFD at the head of the queue. The criterion is whether the flag bit in the GFD header is non-zero.📎 src/gin/gin_host_proxy.cc:176-182If so, first copy the first qword (the header), then wait for the remaining qwords to become ready.📎 src/gin/gin_host_proxy.cc:194-202After copying is complete, zero out the GFD in the queue to prevent duplicate processing.📎 src/gin/gin_host_proxy.cc:206-208

Step 4:proxyGinProcessGfdDispatch to different processing paths according to the operation type.📎 src/gin/gin_host_proxy.cc:246-340

mermaid
flowchart TD
    poll_start["proxyGinPollGfd(ctx, hostGpuCtx, targetRank)"]
    check_avail{"isGfdAvailable?"}
    no_gfd["返回 0, 跳出批量循环"]
    copy_header["拷贝 GFD header qword"]
    copy_rest["循环等待并拷贝其余 qword"]
    reset_gfd["清零队列中的 GFD"]
    set_state["设置 state->op, counterId, done=0"]
    inc_sis["sis[targetRank]++"]
    process["proxyGinProcessGfd(ctx, hostGpuCtx, targetRank, gfd, state, isLastInBatch)"]
    check_va{"op & ncclGinProxyOpVASignal?"}
    check_get{"op & ncclGinProxyOpGet?"}
    check_flush{"op & ncclGinProxyOpFlush?"}
    check_inline{"op & ncclGinProxyOpWithInline?"}
    va_signal["rmaBackend->iputSignal(...)"]
    get_op["rmaBackend->iget(...)"]
    flush_op["rmaBackend->iflush(...)"]
    inline_src["从 inlines 缓冲区取源地址"]
    normal_src["从 GFD 取源地址"]
    put_signal["rmaBackend->iputSignal(...)"]
    put_only["rmaBackend->iput(...)"]

    poll_start --> check_avail
    check_avail -->|否| no_gfd
    check_avail -->|是| copy_header
    copy_header --> copy_rest
    copy_rest --> reset_gfd
    reset_gfd --> set_state
    set_state --> inc_sis
    inc_sis --> process
    process --> check_va
    check_va -->|是| va_signal
    check_va -->|否| check_get
    check_get -->|是| get_op
    check_get -->|否| check_flush
    check_flush -->|是| flush_op
    check_flush -->|否| check_inline
    check_inline -->|是| inline_src
    check_inline -->|否| normal_src
    inline_src --> put_signal
    normal_src --> put_signal
    put_signal --> put_only

Complete polling and counter updates

proxyGinPollCompletionsis responsible for checking the completion status of submitted requests.📎 src/gin/gin_host_proxy.cc:113-156

For each target rank, fromcisShadowtosisiterate over all seen but unconsumed GFD states.📎 src/gin/gin_host_proxy.cc:117If the state is not complete, callrmaBackend->testto check.📎 src/gin/gin_host_proxy.cc:122If it is complete and the operation carries a counter flag, update the counter value.📎 src/gin/gin_host_proxy.cc:132-141

Counter updates use atomic loads and atomic stores, but the comment explains why atomic addition is not needed: the GPU kernel does not allow resetting the counter while there are outstanding operations, so there is no race.📎 src/gin/gin_host_proxy.cc:133-135

The update of CI has a mechanism that "allows holes": only whenstate->done && i == cisShadow[targetRank]does CI advance.📎 src/gin/gin_host_proxy.cc:145-151This ensures that CI is monotonically increasing, and even if some GFDs complete first, it will not skip incomplete GFDs.

Concurrency Control and Memory Barriers

The concurrency model of the GIN proxy is more complex than that of the RMA proxy, because there are multiple proxy threads (controlled byGIN_PROXY_NTHREADS).📎 src/gin/gin_host.cc:90

ncclGinProgressIn , each thread is responsible for a set of connections: thread t handles connections t, t+proxyNthreads, t+2*proxyNthreads, ....📎 src/gin/gin_host.cc:72This allocation method ensures that each connection is handled by only one thread, avoiding connection-level races.

Modifications to the devComms linked list require write-lock protection.ginProgressWriteLockFirst set thewritePendingflag, then acquire the write lock.📎 src/gin/gin_host.cc:43-47The progress thread checkswritePendingat the beginning of each loop, and yields the CPU if it is true.📎 src/gin/gin_host.cc:63-66This design avoids the progress thread being blocked by a write lock while holding a read lock.

writePendingusesstd::atomic<bool>, but the comment points out that this logic assumes there is only one writer.📎 src/gin/gin_host.cc:43-47In NCCL's usage scenario, only the main thread modifies the devComms linked list, so this assumption holds.

Production Pitfalls

Pitfall 1: The memory location of the GFD queue. queuesis forcibly allocated in host memory (forceHost=true),📎 src/gin/gin_host_proxy.cc:564This means that GPU writes to GFDs must go through the PCIe bus. If the GFD write frequency is very high (small-message scenarios), PCIe bandwidth may become a bottleneck. In contrast,cisis allocated in GPU-accessible memory, because the proxy needs to update it frequently.📎 src/gin/gin_host_proxy.cc:565-566

Pitfall 2: Reconstruction of inline data.When a GFD carries inline data, the proxy needs to reconstruct the inline value from multiple qwords.📎 src/gin/gin_host_proxy.cc:298-305The reconstruction logic decides which qwords to read based on size: size ≤ 4 reads only the low 32 bits, size > 4 reads the low 64 bits, and size > 6 additionally reads the high 16 bits. This segmented logic must strictly correspond to the write logic on the GPU side; any inconsistency will cause data corruption.

Pitfall 3: Multithreaded progress and connection allocation.If different ranks set differentGIN_PROXY_NTHREADS, after taking the minimum via AllGather, some threads may not be allocated any connections.📎 src/gin/gin_host.cc:181-183Comments point out that these threads will spin in the stride loop, which will not cause correctness issues but will waste CPU resources.

GIN backend selection and version compatibility

Intuitive model

GIN supports multiple backends: Proxy (software emulation based on the RMA plugin), GDAKI (GPU Direct Async Kernel Initiated), GPI (GPU-Initiated), and EFA GDA (AWS EFA's GPU Direct Async). This is like how the same API can have multiple implementations - the software emulation version has the best compatibility but average performance, while the hardware-offload version has the best performance but requires support from specific NICs.

Backend version matrix

Each backend has a version compatibility array, where the index is the backend version number and the value is the minimum NCCL version required by that version.📎 src/gin/gin_host.cc:27-33

BackendVersion 0Version 1Version 2Version 3
Proxy02.30.32.30.52.32.0
GDAKI02.30.32.30.5-
GPI02.30.5--
EFA GDA02.31.02.32.0-

Version selection logic: iterate through the version array and find the first entry whose required version is higher than the current device code version; the previous version is then the available version.📎 src/gin/gin_host.cc:300-304

Backend selection process

ncclGinDevCommSetupIterate through all active backends and try to create a DevComm with each backend.📎 src/gin/gin_host.cc:427-442Selection conditions include: the requested GIN type matches (or is unspecified), and the signal capability meets the requirements.📎 src/gin/gin_host.cc:430-435

ncclGinValidateSignalRequestCheck two capabilities: strong signal (supportsStrongSignals) and VA signal (supportsVASignals)。📎 src/gin/gin_host.cc:230-243If the request requires a strong signal but the backend does not support it, skip that backend.

Connection establishment and stride calculation

ncclGinConnectOnceEstablish the GIN connection.📎 src/gin/gin_host.cc:92-228

The connection type determines the stride: in FULL mode the stride is 1 (connect to all ranks), and in RAIL mode the stride iscontiguousRanksPerHost(only connect to ranks on the same rail).📎 src/gin/gin_host.cc:139-145

InginDevCommSetupWithBackend, the stride validation logic is very strict:

  • The requested stride cannot be 0.📎 src/gin/gin_host.cc:318-323
  • The requested stride cannot be greater than the stride of the rail team.📎 src/gin/gin_host.cc:324-330
  • The requested stride must be a multiple of the connected stride.📎 src/gin/gin_host.cc:331-337

The motivation for these constraints is that the hierarchical barrier assumes GIN is at least RAIL-connected.📎 src/gin/gin_host.cc:325If the stride does not satisfy these conditions, the communication path between some ranks may not exist.

Production pitfalls

Pitfall 1: Backend version mismatch.If the device code version is lower than the minimum version required by the backend,backendVersionwill remain at a lower value.📎 src/gin/gin_host.cc:301-303This may make some new features unavailable (such as signals never being reset), but it will not cause errors. However, if the device code version is higher than all known versions,backendVersionwill take the maximum value, which may trigger undefined behavior.

Pitfall 2: The boundary of stride validation.IfrequestedStride % connectedStride != 0, creation fails.📎 src/gin/gin_host.cc:331-337This check assumes that connectedStride is a power of 2 (1 in FULL mode, andcontiguousRanksPerHostin RAIL mode). IfcontiguousRanksPerHostis not a power of 2 (for example, 3), the multiple check may reject a legal stride.

Chapter review and self-test

Q1: InscheduleRmaTasksToPlan's WaitSignal branch, if you removeplan->rmaArgs->nRmaTasks = (npeersCe > 0 ? 1 : 0) + (npeersProxy > 0 ? 1 : 0)this line and change it to directly set 1, in what scenario would it cause a problem?

Reference analysis: Look at📎 src/rma/rma.cc:248。nRmaTasksrecords the actual number of tasks enqueued. If all peers are LSA-reachable (npeersProxy == 0), only 1 CE task is actually enqueued,nRmaTasksshould be 1. If all peers are unreachable (npeersCe == 0), only 1 Proxy task is actually enqueued,nRmaTasksshould also be 1. But if the peers are mixed, both tasks are enqueued,nRmaTasksshould be 2.

If this line is changed toplan->rmaArgs->nRmaTasks = 1, then in a mixed distribution scenario,nRmaTaskswill underestimate the actual number of tasks. The subsequent judgment inncclRmaWaitSignalplan->rmaArgs->nRmaTasksProxy > 0 && plan->rmaArgs->nRmaTasksCe > 0can still work correctly (because it usesnRmaTasksProxyandnRmaTasksCe),📎 src/rma/rma.cc:47), but any code that relies onnRmaTasksfor resource estimation or logging statistics will get incorrect results. More seriously, if subsequent code usesnRmaTasksto allocate arrays or calculate loop counts, it may cause buffer overflow or missed tasks.

Q2: InproxyGinPollGfd, ifhostGpuCtx->sis[targetRank]++is moved to after theproxyGinProcessGfdcall, in what concurrency scenario would GFD be processed repeatedly?

Reference analysis: Look at📎 src/gin/gin_host_proxy.cc:228。sisis the "seen index", indicating the number of GFDs that the proxy has seen and started processing.proxyGinPollGfdImmediately increments after copying the GFDsis, and then returns 1 to indicate success. The callerncclGinProxyProgresscallsproxyGinPollGfdin a loop, and if it returns 1, continues processing the next GFD.📎 src/gin/gin_host_proxy.cc:648-669

Ifsis++is moved to afterproxyGinProcessGfd, then during the execution ofproxyGinProcessGfd(which may involve asynchronous calls to the network plugin),sisstill points to the current GFD. If at this time the GPU writes a new GFD to the same slot (because the queue is circular,pismay have already wrapped around),proxyGinPollGfdwill see this slot again, butsishas not advanced, causing the same slot to be processed repeatedly.

Even more dangerous is that,proxyGinPollGfdAfter copying the GFD, the GFD in the queue is cleared.📎 src/gin/gin_host_proxy.cc:206-208Ifsisdoes not advance, the next poll will see the cleared GFD (flag is 0),isGfdAvailablereturns false, causing the GFD to be lost. This causes the GPU side to wait for a request that will never be processed, ultimately leading to deadlock.

Q3: InncclRmaProxyProgressThread, ifrmaProgress == 2the branch forgets to callrmaProxyState->cond.notify_one(), in what scenario will it cause the main thread to block permanently?

Reference analysis: Look at📎 src/rma/rma_proxy.cc:373-378。rmaProgress == 2is the "pause request" state, used for resource reclamation. After the main thread setsrmaProgress = 2, it waits for the progress thread to confirm the pause. The progress thread waits incond.wait(lock), and the main thread needs to callcond.notify_one()to wake it up.📎 src/rma/rma_proxy.cc:377

If the progress thread, after settingrmaProgress = 0, forgetsnotify_one(), the main thread will wait forever on the condition variable. But more critically, while the progress thread is waiting incond.wait(lock), the main thread needs to acquire the lock first before it can setrmaProgress = 2. If the progress thread does not release the lock beforewait, the main thread cannot acquire the lock, forming a deadlock.

The correct order is: the progress thread setsrmaProgress = 0, callsnotify_one()to wake the main thread, then callscond.wait(lock)to release the lock and wait. After the main thread is woken up, it acquires the lock, setsrmaProgress = 2, callsnotify_one()to wake the progress thread, and then waits for the progress thread to confirm. After the progress thread is woken up, it setsrmaProgress = 0, againnotify_one(), and thenwait. In this handshake protocol, the absence ofnotify_one()at any step will cause permanent blocking.

From RMA's put/get semantics to GIN's GPU-initiated network communication, we have completed a key step in NCCL's evolution toward a general-purpose remote memory access engine. But no matter how ingenious the mechanism is, it must ultimately interface with external network backends, tuning strategies, and performance collectors through the plugin system. The next chapter enters the plugin world to see how NCCL dynamically loads extensions such as net, tuner, profiler, and env without modifying the core code, and uses google-fastsocket and google-CoMMA as examples to reveal the key implementation points of ecosystem extensibility.

CHAPTER 16

Chapter 16: Chapter 16: Plugin ecosystem and environment variables: how net, tuner, profiler, and env extend NCCL behavior

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 16 / 25

Chapter 16: Plugin ecosystem and environment variables: how net, tuner, profiler, and env extend NCCL behavior

In the previous chapter, we saw how NCCL extends communication capabilities from collective operations to point-to-point remote access through RMA and GIN, and even lets the GPU directly initiate network requests. This evolution toward new hardware and low-latency scenarios places higher demands on the flexibility of the communication engine: if adapting to a new network, a new tuning strategy, or a new collection tool required recompiling the core code every time, NCCL would struggle to keep up with ecosystem changes. This chapter breaks down the src/plugin and plugins directories to answer a core question: how does NCCL replace network backends, tuning strategies, performance collectors, and configuration sources without recompiling the core code?

16.1 Plugin loader: how plugin_open.cc turns a .so into a usable backend

Intuitive model

Think ofplugin_open.ccas NCCL's "recruitment agency": it holds a list of positions (NET, GIN, RMA, TUNER, PROFILER, ENV), and each position corresponds to a candidate library name. When NCCL needs someone for a position, the agency goes to the talent market (dynamic linker) in a fixed order to find someone, signs a contract if found (dlopen), records "this person does not exist" if not found, and finally returns a handle. Without this intermediary layer, NCCL could only hardcode network backends into the binary, and any NIC vendor wanting to integrate would have to modify the NCCL source code—this is exactly the disaster the plugin system aims to eliminate.

Data structures and memory layout

The loader's entire state is six parallel arrays, with the index being the plugin type enum:

code
static char* libNames[NUM_LIBS];              // 已加载库的名字
char* ncclPluginLibPaths[NUM_LIBS];           // 库的绝对路径
static void* libHandles[NUM_LIBS];            // dlopen 返回的句柄
static const char* pluginNames[NUM_LIBS];     // 日志用的人类可读名
static const char* pluginPrefix[NUM_LIBS];    // 库名前缀
static const char* pluginFallback[NUM_LIBS];  // 找不到时的提示
static unsigned long subsys[NUM_LIBS];        // 日志子系统位掩码

The subscripts of these seven arrays must be strictly aligned,pluginNames[type]、pluginPrefix[type]、subsys[type]describes the same plugin type.📎 src/plugin/plugin_open.cc:18-29definesNUM_LIBS = 6, the type order is{"NET", "GIN", "RMA", "TUNER", "PROFILER", "ENV"}, and the prefix is{"libnccl-net", "libnccl-gin", "libnccl-rma", "libnccl-tuner", "libnccl-profiler", "libnccl-env"}。

[Design inference and architectural trade-offs]

Parallel arrays are used here instead of an array of structs so thatopenPluginLibthis single function can serve six kinds of plugins at the same time—the type is only used as a subscript, and the logic is fully reused. The cost is that when adding a new plugin type, six arrays must be modified in sync, and the compiler cannot help you check for omissions.

subsysThe array determines log ownership: NET/GIN/RMA all attach toNCCL_INIT | NCCL_NET, TUNER attaches toNCCL_INIT | NCCL_TUNING, PROFILER attaches only toNCCL_INIT, ENV attaches toNCCL_INIT | NCCL_ENV。📎 src/plugin/plugin_open.cc:26-29In this way,NCCL_DEBUG_SUBSYS=NETonly network plugin logs will be seen, without being drowned in tuning logs.

Step-by-Step Walkthrough: a complete journey ofncclOpenNetPluginLib("mlx5")

Suppose the user setsNCCL_NET_PLUGIN=mlx5, and during NCCL initializationncclOpenNetPluginLib("mlx5")is called, which directly forwards toopenPluginLib(ncclPluginTypeNet, "mlx5")。📎 src/plugin/plugin_open.cc:132-134

Step 1: Construct the candidate library name.Because a non-emptylibNameis passed in, it goes through thesnprintf(libName_, MAX_STR_LEN, "%s", libName)branch,libName_becomes"mlx5"。📎 src/plugin/plugin_open.cc:85-89Note that at this point it is not yet a valid library file name—it has no prefix and no.sosuffix.

Step 2: First attempt to open. tryOpenLib("mlx5", ...)is called.📎 src/plugin/plugin_open.cc:91After enteringtryOpenLib, first check whethernameis empty or has zero length, then there is a special branch: if the name starts withSTATIC_PLUGIN, setnametonullptr。📎 src/plugin/plugin_open.cc:37-39This is the sentinel for plugins statically linked into NCCL—dlopen(nullptr)On Linux, it returns the main program handle, thereby allowingdlsymto find plugin symbols in the main program's symbol table.

Then it callsncclOsDlopen(name)。📎 src/plugin/plugin_open.cc:41because"mlx5"is neither a path nor a valid library name,dlopenwill fail. After the failure, the code takesncclOsDlerror()'s error string and makes a fine-grained judgment: if the error string contains bothnameand"No such file or directory", then set*errtoENOENT。📎 src/plugin/plugin_open.cc:42-55The significance of this judgment is to distinguish "the file does not exist at all" from "the file exists but failed to load"—the former just means the candidate name is wrong, and the next candidate name should be silently tried; the latter is a real error and should be logged.

Step 3: Handling after the first failure.Return toopenPluginLib,libHandles[type]is empty, andopenErr == ENOENT, so append"mlx5"toeNoEntNameList。📎 src/plugin/plugin_open.cc:97-101This list will ultimately be assembled into a log line: "Could not find: mlx5 libnccl-net-mlx5.so".

Step 4: Second attempt—add the prefix.The code checkslibNamewhether it is neither a path (does not contain/) nor a library name (does not start withlib, does not end with.so).📎 src/plugin/plugin_open.cc:105-107 "mlx5"The condition is satisfied, so it assembles"libnccl-net-mlx5.so"and tries again.📎 src/plugin/plugin_open.cc:108This timedlopensucceeds,libHandles[type]is assigned,libNames[type]records the library name,ncclPluginLibPaths[type]obtains the absolute path viagetLibPath, and the function returns the handle.📎 src/plugin/plugin_open.cc:110-115

Step 5: Obtain the absolute path. getLibPathOn Linux, usedlinfo(handle, RTLD_DI_LINKMAP, &lm)to retrievelink_map, thenstrdup(lm->l_name)。📎 src/plugin/plugin_open.cc:65-69This path will appear in all subsequent logs, letting users see at a glance exactly which file was loaded—when troubleshooting in production why the wrong plugin was loaded, this log line is the primary scene.

The entire decision flow is as follows:

mermaid
flowchart TD
    start["openPluginLib(type, libName)"] --> build{"libName 非空?"}
    build -->|是| use_name["libName_ = libName"]
    build -->|否| use_prefix["libName_ = pluginPrefix[type] + .so"]
    use_name --> try1["tryOpenLib(libName_)"]
    use_prefix --> try1
    try1 --> ok1{"handle 非空?"}
    ok1 -->|是| success["记录 libNames/libPaths, 返回 handle"]
    ok1 -->|否| enoent{"openErr == ENOENT?"}
    enoent -->|是| append1["appendNameToList(eNoEntNameList)"]
    enoent -->|否| log1["INFO 打印 dlopen 错误"]
    append1 --> shape{"非路径且非库名?"}
    log1 --> shape
    shape -->|是| try2["tryOpenLib(prefix-libName.so)"]
    shape -->|否| report["打印 Could not find 列表"]
    try2 --> ok2{"handle 非空?"}
    ok2 -->|是| success
    ok2 -->|否| report
    report --> retnull["返回 nullptr"]

Design considerations and production pitfalls

[Design inferences and architectural trade-offs]

The order of candidate names is the priority.First try the bare name given by the user, then try the name with the prefix added. This means that if the current directory happens to contain a file namedmlx5, it will be loaded first—this is a potential security surface, and in production environments you should avoid placing an executable with the same name as the plugin inLD_LIBRARY_PATH.

STATIC_PLUGINThe semantics ofWhenNCCL_NET_PLUGIN=STATIC_PLUGIN,tryOpenLibsets the name to empty,dlopen(nullptr)opens the main program,dlsymand looks for symbols such asncclNet_v12from the main program's symbol table.📎 src/plugin/plugin_open.cc:37-39This allows the plugin to be statically linked into the NCCL binary, eliminating the hassle of deploying.so, at the cost of losing runtime replaceability.

Reference counting and unloading. ncclClosePluginLibOnly whenlibHandles[type] == handledoes it actuallydlclose, and clears the path and name.📎 src/plugin/plugin_open.cc:176-186This equality check prevents mistakenly closing a handle that has already been replaced. The GIN and RMA plugins reuse the NET library's handle throughncclGetGinPluginLib/ncclGetNetPluginLib, implemented by callingdlopenagain with the same library name to increase the reference count.📎 src/plugin/plugin_open.cc:156-164This isdlopen's reference counting semantics—the same library opened twice requiresdlclosetwice to actually unload.

16.2 net.cc: The state machine and lifecycle of network plugins

Intuitive model

net.ccis the "dispatch center" for network plugins. It maintains an array of plugin libraries, each with its own state (not loaded, load failed, pending load, pending initialization, enabled). When a new communicator is born, the dispatch center traverses all candidate plugins, trying to initialize them one by one; the first successful one is "assigned" to this communicator, and all other external plugins are disabled. Without this state machine, NCCL would be unable to handle real-world problems such as "the plugin loaded but the device is unavailable," "which one to choose when multiple plugins coexist," and "how to safely unload when the communicator is destroyed."

Data structures and memory layout

The core structure isnetPluginLib_t:

FieldTypeMeaning
namechar[255]Plugin library name
dlHandlevoid*dlopen handle
ncclNetncclNet_t*Network function table
ncclNetVerintNetwork API version number
ncclCollNetncclCollNet_t*Collective communication offload function table
ncclNetPluginStateEnumNetwork plugin state
ncclCollNetPluginStateEnumCollNet plugin state
ncclNetPluginRefCountintReference count
netPhysDevs/netVirtDevsintNumber of physical/virtual devices
collNetPhysDevs/collNetVirtDevsintNumber of CollNet devices

📎 src/plugin/net.cc:63-76defines these fields. Note thatncclNetandncclCollNetare two separate function tables, and the states are also two separate enums—a plugin can provide network functionality but not CollNet offload.

The state enum has five values:Disabled = -2(initialization failed),LoadFailed = -1(load failed),LoadReady = 0(pending load),InitReady = 1(loaded pending initialization),Enabled = 2(enabled).📎 src/plugin/net.cc:54-60uses negative numbers to represent failure states, so that comparisons like "state >= InitReady" can naturally express "at least loaded."

The global state consists of three variables:pluginCountrecords the total number of plugins,netPluginLibs[NCCL_NET_MAX_PLUGINS]is the plugin array,netPluginMutexprotects concurrent access,initPluginLibsOnceFlagensures initialization is done only once.📎 src/plugin/net.cc:78-81

Step-by-Step Walkthrough: A complete journey of onencclNetInit(comm)

Step 1: One-time initialization. std::call_once(initPluginLibsOnceFlag, initPluginLibsOnceFunc)ensures the plugin list is built only once.📎 src/plugin/net.cc:360 initPluginLibsOnceFuncreads theNCCL_NET_PLUGINenvironment variable; if not set, it adds by default"libnccl-net.so", then registers two built-in pluginsncclNetIbandncclNetSocket。📎 src/plugin/net.cc:288-340

Environment variable parsing usesstrtok_rto split by commas, supporting multiple plugin names.📎 src/plugin/net.cc:303-324has a capacity check: the number of external plugins cannot exceedNCCL_NET_MAX_PLUGINS - NCCL_NET_NUM_INTERNAL_PLUGINS; the excess is ignored and logged.📎 src/plugin/net.cc:307-311Built-in plugins are fixed at 2 (IB and Socket), so external plugins are at mostNCCL_NET_MAX_PLUGINS - 2.

Step 2: Locked traversal. std::lock_guard<std::mutex> lock(netPluginMutex)protects the entire traversal process.📎 src/plugin/net.cc:361For each plugin index, first determine whether it is an external plugin and in theLoadReadystate; if so, callncclNetPluginLoad。📎 src/plugin/net.cc:364-367

Step 3: Load the plugin. ncclNetPluginLoadcallsncclOpenNetPluginLibto get the handle, then tries from high version to low version in ordergetNcclNet_v12togetNcclNet_v6; the first version that returns non-null is adopted.📎 src/plugin/net.cc:103-112The version arrayncclNetVersionand function pointer arraygetNcclNetare arranged in descending order, ensuring the latest API is used first.📎 src/plugin/net.cc:41-43

If all versions fail to obtainncclNet, it means this library is not a valid network plugin. At this point, check whetherNCCL_NET_PLUGINis explicitly set: if set, warn atATTNlevel (the user explicitly requested it but it failed); if not set, useINFOlevel (just the default attempt failure).📎 src/plugin/net.cc:115-125This distinction is important—if the user's explicit configuration fails, they must see it.

Step 4: Initialize the plugin.Return toncclNetInit, for state>= InitReadyand name matchingcomm->config.netNameplugin callncclNetPluginInit。📎 src/plugin/net.cc:369-372 ncclNetPluginInitDo two things: call the plugin'sinitfunction to establish the communication domain context, and on first initialization calldevicesto probe the device count.📎 src/plugin/net.cc:186-236

Noteinitcall conditions:pluginLib->ncclNetPluginState >= ncclNetPluginStateInitReady。📎 src/plugin/net.cc:190The comment explicitly states "every new communication domain must call init to set the correct context."📎 src/plugin/net.cc:189But device probing is only done once at== InitReady.📎 src/plugin/net.cc:201This distinction of "init called every time, devices called only once" is a performance optimization—device probing can be slow, but the context must be independent for each communication domain.

Step 5: Allocation and disabling.After successful initialization, callncclNetPluginAssignToComm, which assigns the plugin'sncclNettocomm->ncclNet, increments the reference count, setscomm->netPluginIndex。📎 src/plugin/net.cc:238-255After successful allocation, immediately callncclNetPluginDisableOtherExternalto disable all other external plugins.📎 src/plugin/net.cc:377-380

[Design inference and architectural trade-offs]

The disable logic has a key judgment: only when the allocated plugin is an external plugin (pluginIndex >= pluginCount - NCCL_NET_NUM_INTERNAL_PLUGINS) are other external plugins disabled.📎 src/plugin/net.cc:257-259If a built-in IB plugin is allocated, external plugins remain as-is—this leaves room for choice in subsequent communication domains.

mermaid
flowchart TD
    init["ncclNetInit(comm)"] --> once["call_once(initPluginLibsOnceFunc)"]
    once --> lock["lock(netPluginMutex)"]
    lock --> loop{"遍历 pluginIndex"}
    loop -->|外部且 LoadReady| load["ncclNetPluginLoad()"]
    loop -->|状态 >= InitReady| namechk{"netName 匹配?"}
    load --> namechk
    namechk -->|否| loop
    namechk -->|是| plugininit["ncclNetPluginInit()"]
    plugininit --> enabled{"状态 == Enabled?"}
    enabled -->|否| loop
    enabled -->|是| assign["ncclNetPluginAssignToComm()"]
    assign --> assigned{"isAssigned?"}
    assigned -->|否| finalize["ncclNetPluginFinalize()"]
    finalize --> loop
    assigned -->|是| disable["ncclNetPluginDisableOtherExternal()"]
    disable --> ok["返回 ncclSuccess"]
    loop -->|遍历结束| fail["WARN 无可用插件, 返回 ncclInvalidUsage"]

Concurrency control and hardware interaction

netPluginMutexprotects all reads and writes tonetPluginLibs.ncclNetInit、ncclNetFinalizeAll are locked.📎 src/plugin/net.cc:361📎 src/plugin/net.cc:411-416ButncclNetGetDevCountand other function comments say "no lock needed, because the caller is already withinncclTopoGetSystem's lock."📎 src/plugin/net.cc:418-429This is a convention of "the lock is held by the upper layer," reducing the overhead of nested locks, at the cost that callers must follow the convention.

ncclGpuGdrSupportdemonstrates direct interaction between the plugin and hardware: it allocates a 2MB GPU buffer, establishes a loopback connection through the plugin'slisten/connect/accept, and then attemptsregMrto register GPU memory.📎 src/plugin/net.cc:464-535If registration succeeds, it indicates the NIC supports GPUDirect RDMA. This probe result is cached ingdrSupportMatrix[32], indexed by CUDA device number.📎 src/plugin/net.cc:478-480

[Design inference and architectural trade-offs]

NotegdrSupportMatrixisstatic's, shared across communication domains.📎 src/plugin/net.cc:478This means multiple communication domains within the same process will reuse the probe result, avoiding repeated expensive probing. But the array size is hardcoded to 32, and machines with more than 32 GPUs will go out of bounds—this is an implicit upper-limit assumption.

Production pitfall avoidance guide

Pitfall 1: Plugin loads successfully but device count is zero. ncclNetPluginInitCheckdevices(&ndev) != ncclSuccess || ndev <= 0and jump to the failure branch.📎 src/plugin/net.cc:202After failure, callfinalizeto clean up the established context, reset the device count toNCCL_UNDEF_DEV_COUNT, and set the state toDisabled。📎 src/plugin/net.cc:229-234If this cleanup is not done, subsequent communication domains will see a plugin that is "initialized but has no devices," causing hard-to-diagnose errors.

[Design inference and architectural trade-offs]

Pitfall 2:initsucceeds butdevicesfails.The code usesinitCompletedflag to trackinitwhether it succeeded.📎 src/plugin/net.cc:178-184📎 src/plugin/net.cc:198In the failure branch, only ifinitCompletedis true isfinalize。📎 src/plugin/net.cc:230called. This prevents callingfinalizeon an uninitialized context—many plugins'finalizedo not check for null pointers, and an erroneous call will crash.

Pitfall 3: Reference counting when destroying a communication domain. ncclNetPluginFinalizeFirst call the plugin'sfinalize, then decrement the reference count, and finally unload the library when the reference count reaches zero and it is an external plugin.📎 src/plugin/net.cc:342-355 ncclNetPluginUnloadCheckdlHandleis non-null and the reference count is zero before actuallydlclose。📎 src/plugin/net.cc:84-101After unloading, reset the fields but retainname, so it can be reused when reloaded.📎 src/plugin/net.cc:84-101

16.3 tuner.cc and profiler.cc: Different contracts for strategy plugins and observation plugins

Intuitive model

The Tuner plugin is like "route preference settings in navigation software"—it does not change how the car is driven, only which route is chosen. The Profiler plugin is like a "dashcam"—it does not intervene in driving, only records what happened. What they have in common is that both are connected through a function table. The difference is that Tuner is a lightweight strategy object with "one instance per communication domain," while Profiler requires a separate thread to asynchronously consume events generated by the GPU.

tuner.cc: A minimalist global singleton

Tuner's state is extremely simple: one mutex, one reference count, one library handle, one symbol pointer, and one state variable.📎 src/plugin/tuner.cc:24-37There is no plugin array, no coexistence of multiple plugins—there is only one global tuner.

ncclTunerPluginLoadThe logic is "load on first use, reuse afterward": if the state isLoadSuccess, directly assign the symbol tocomm->tunerand increment the reference count.📎 src/plugin/tuner.cc:53-57Otherwise read theNCCL_TUNER_PLUGINenvironment variable; if it is"none", fail directly.📎 src/plugin/tuner.cc:59-63

[Design inference and architectural trade-offs]

Version negotiation drops from v6 to v2, trying one by one.📎 src/plugin/tuner.cc:75-87Note that there is no v1 here—the tuner API only has a stable function table structure starting from v2.

[Design inference and architectural trade-offs]

An interesting detail: ifncclOpenTunerPluginLibreturns empty, the code triesncclGetNetPluginLib(ncclPluginTypeTuner)。📎 src/plugin/tuner.cc:65-70This means the tuner can be packaged in the net plugin library—this reduces deployment complexity, with one.soproviding both networking and tuning functionality.

profiler.cc: Asynchronous event consumption thread

Profiler is the most complex plugin in this chapter because it needs to handle events asynchronously generated by the GPU. The core structure isncclProfilerThread:

FieldTypePurpose
threadstd::threadConsumption thread
mutexstd::mutexProtects the queue
condcondition_variableWakes up when there is new work
condIterationInactivecondition_variableWaits for iteration to end
stopintStop flag
refCountintCommunication domain reference count
cudaDevintBound CUDA device
abortFlagvolatile uint32_t*Abort flag
iterationActiveboolWhether iterating
pending/pendingTailLinked listPending work
active/activeTailLinked listWork in progress
opStack/opPoolMemory poolWork object allocation
inflight/maxInflightSeen/maxInflightsize_tBackpressure observation
droppedOpsuint64_tAllocation failure count

📎 src/plugin/profiler.cc:38-69defines this structure. Notependingandactiveare two independent linked lists: producers append topending, and the consumption thread splicespendingintoactivewithin the lock, then traversesactive。📎 src/plugin/profiler.cc:56-59

iterationActiveoutside the lock. The flag is key to concurrency correctness: the consumption thread sets it totruewithin the lock, then releases the lock to call the plugin callback. The destruction thread must wait for this flag to return tofalseonly then can the communicator state be torn down.📎 src/plugin/profiler.cc:52-55

Step-by-Step Walkthrough: Generation and Consumption of a KernelCh Event

Step 1: Host-side enqueue.When the kernel plan is submitted,ncclProfilerPostPlanWorkiterate over the collective tasks in the plan, and for each task withncclProfileKernelChenabled, callprofilerPostWorkInternal。📎 src/plugin/profiler.cc:1315-1331

profilerPostWorkInternalby channel range. First incrementcomm->profiler.workCounter[channelId]then callprofilerEnqueueOp。📎 src/plugin/profiler.cc:1259-1266The comment emphasizes that this increment must be "exactly once per call, even if allocation fails," to stay in sync with the device kernel.📎 src/plugin/profiler.cc:1259-1266

Step 2: Allocate the work object. profilerEnqueueOpInside the lock, allocate from the memory poolncclProfilerWorkOpand fill in fields such as channel number, work counter, activation mask, task event handle, and communicator context.📎 src/plugin/profiler.cc:1199-1223On allocation failure, incrementdroppedOpsand log it, butdo notroll backworkCounter—this is the key to staying in sync with the device.📎 src/plugin/profiler.cc:1202-1207

After successful allocation, append the object to the tail of thependinglinked list, incrementinflightupdatemaxInflightSeenand wake up the consumer thread.📎 src/plugin/profiler.cc:1225-1239

Step 3: The consumer thread waits. ncclProfilerThreadFuncIt loops callingwaitForAction。📎 src/plugin/profiler.cc:1074-1077 waitForActionwaiting on the condition variable inside the lock untilpendingoractiveis non-empty, or a stop/abort signal is received.📎 src/plugin/profiler.cc:1017-1031

After being woken up, it callsappendWorkToActiveQueueto splicependingonto the tail ofactivesetiterationActive = trueand returnNCCL_PROFILER_THREAD_PROGRESS。📎 src/plugin/profiler.cc:1017-1031

Step 4: Process the work. profilerProgressOpsOutsidethe lockiterate over theactivelinked list.📎 src/plugin/profiler.cc:958-999For each work object, check whether the device has already written the start timestamp:wc <= op->workStarted[ch].data[slot].counter。📎 src/plugin/profiler.cc:972Note that<=is used rather than==because the device wraps aroundMAX_PROFILER_EVENTS_PER_CHANNELslots, and if the host falls behind, the device may have already overwritten that slot.📎 src/plugin/profiler.cc:969-971

If the start condition is satisfied, callncclProfilerStartKernelChEventto notify the plugin.📎 src/plugin/profiler.cc:973Then check the completion condition; if satisfied, first trigger the phase event, then callncclProfilerStopKernelChEvent。📎 src/plugin/profiler.cc:978-985

Completed work objects are removed from the linked list and collected into therecycledlist.📎 src/plugin/profiler.cc:987-991

Step 5: Reclaim and publish. cleanupAndStopInside the lock, reclaim therecycledlist, publish the newactiveTailcleariterationActiveand notify waiters.📎 src/plugin/profiler.cc:1036-1050

mermaid
sequenceDiagram
    participant Host as 主机线程
    participant PT as Profiler 线程
    participant Plugin as Profiler 插件
    participant Dev as GPU 内核

    Host->>Host: profilerPostWorkInternal() 递增 workCounter
    Host->>PT: profilerEnqueueOp() 追加到 pending
    Host->>PT: cond.notify_one()
    PT->>PT: waitForAction() 返回 PROGRESS
    PT->>PT: appendWorkToActiveQueue() 拼接 pending 到 active
    Dev->>Dev: 内核写入 workStarted/workCompleted 时间戳
    PT->>PT: profilerProgressOps() 检查 wc <= counter
    PT->>Plugin: startEvent(ncclProfileKernelCh)
    PT->>Plugin: recordEventState(ncclProfilerKernelChStop)
    PT->>Plugin: stopEvent()
    PT->>PT: cleanupAndStop() 回收对象, 清除 iterationActive

Concurrency Control and Backpressure

NCCL_PROFILER_DEFAULT_MAX_INFLIGHTis defined asMAXCHANNELS * MAX_PROFILER_EVENTS_PER_CHANNEL * 4。📎 src/plugin/profiler.cc:32-32This is a "soft cap"—exceeding it does not prevent enqueueing, it only logs.📎 src/plugin/profiler.cc:1233-1238The comment explains that keeping the enqueue is to pair KernelCh events with their parent task events.📎 src/plugin/profiler.cc:32-32

Logging is triggered at powers of 2:(pt->inflight & (pt->inflight - 1)) == 0。📎 src/plugin/profiler.cc:1233This ensures logging only when inflight is 1, 2, 4, 8..., avoiding log spam.

The consumer thread's backoff strategy is inupdateProgressIntervalwhen there is progress, retry immediately; when there is no progress, start at 1 microsecond and double, up to a maximum of 10 microseconds.📎 src/plugin/profiler.cc:1054-1057This design balances latency and CPU usage.

Production Pitfall Guide

Pitfall 1: Work leak during destruction. ncclProfilerThreadDestroyFirst wait foriterationActiveto become false, then callprofilerPurgeByContextto clear all pending work referencing that communicator context.📎 src/plugin/profiler.cc:1162-1169If this clearing is not done, plugin callbacks will receive a pointer to an already-destroyed context, causing a use-after-free.

Pitfall 2: Draining on stop.When a stop signal is received butactiveis non-empty, returnNCCL_PROFILER_THREAD_CLEANUP_AND_STOP,cleanupAndStopwith thedrainStuckparameter set to true, directly reclaiming all remaining work.📎 src/plugin/profiler.cc:1029📎 src/plugin/profiler.cc:1036-1050The comment says the kernels for this work will never run, so it is simply discarded.📎 src/plugin/profiler.cc:1034-1035

Pitfall 3: CUDA device binding.When the consumer thread starts, callcudaSetDevice(pt->cudaDev)。📎 src/plugin/profiler.cc:1054-1057The comment explains: the thread itself only reads host pinned memory, but plugins may make context-dependent driver calls, so binding is defensive.📎 src/plugin/profiler.cc:1054-1057Binding failure only logs and does not abort, because the thread itself does not depend on CUDA.📎 src/plugin/profiler.cc:1065-1070

16.4 Official Examples: Implementation Highlights of google-fastsocket and google-CoMMA

Intuitive Model

The official examples are the "reference implementations" of the plugin API.google-fastsocketshows how to replace kernel TCP with a userspace network stack;google-CoMMAshows how to implement a profiler plugin to collect communication performance. Their existence proves that the plugin API is expressive enough for real requirements.

google-fastsocket: Replacing the Network Backend

[Design Inference and Architectural Trade-offs]

FastSocket is Google's open-source userspace network stack that bypasses the kernel TCP/IP stack through theAF_FABRICaddress family. As an NCCL net plugin, it needs to implementncclNet_tall functions:init、devices、getProperties、listen、connect、accept、regMr、isend、irecv、test、closeSendetc.

The key implementation point is thegetPropertiesreturned byptrSupportif FastSocket supports GPUDirect RDMA, it should be set toNCCL_PTR_HOST|NCCL_PTR_CUDAotherwise it can only be set toNCCL_PTR_HOSTand NCCL will copy GPU data to host memory before sending.📎 plugins/net/README.md:245-245

connectandacceptThe "non-blocking" contract ofsendComm/recvCommis the core difficulty of plugin implementation: they must return immediately, settingNULLto📎 plugins/net/README.md:299-311and letting NCCL call repeatedly until success.

This requires the plugin to maintain a connection state machine internally, putting the time-consuming handshake in the background.

google-CoMMA: Implementing a Profiler Plugin

[Design Inference and Architectural Trade-offs]ncclProfiler_tCoMMA (Collective Memory Monitoring Agent) is Google's communication performance collector. As a profiler plugin, it implements theinit、finalize、startEvent、stopEvent、recordEventState。

initfunction table:ncclProfilerEventMaskreceives the📎 src/plugin/profiler.cc:341pointer, and the plugin selects which events to subscribe to by writing to this mask.📎 src/plugin/profiler.cc:285-307

startEventThe event types supported by NCCL include Group, Coll, P2p, ProxyOp, ProxyStep, ProxyCtrl, KernelCh, KernelPhase, NetPlugin, etc.stopEventreturns an event handle, and subsequentrecordEventStateand📎 src/plugin/profiler.cc:392📎 src/plugin/profiler.cc:400-407use this handle to associate events.

The plugin can use the handle to store its own state, implementing event pairing and duration statistics.

Design ReflectionsBecause the net API involves device-side code (ncclNetDeviceHandle), a version mismatch will cause a kernel crash; whereas tuner/profiler are purely host-side, and a version mismatch at most results in missing functionality.📎 src/plugin/net.cc:153-176showsncclNetCheckDeviceVersionhow to check the device type and version, returning when there is a mismatchncclInternalError。

Why does the profiler need a separate thread?Because profiler callbacks may block (such as writing files or making network requests), and calling them on the host thread would slow down communication.📎 src/plugin/profiler.cc:950-952The comment explicitly states "plugin callbacks may block, so they must not be called while holding the lock."

16.5 Production Pitfall Guide and Failure Recovery Chain

Pitfall 1: Plugin version mismatch causes kernel crash

ncclNetCheckDeviceVersionCheckprops.netDeviceTypeandprops.netDeviceVersion。📎 src/plugin/net.cc:153-176If the plugin reports aNCCL_NET_DEVICE_UNPACKversion that is inconsistent with theNCCL_NET_DEVICE_UNPACK_VERSIONused when NCCL was compiled, returnncclInternalErrorand raise a warning.📎 src/plugin/net.cc:153-176This check is called inncclNetPluginAssignToComm, and on failure the plugin will not be assigned to a communication domain.📎 src/plugin/net.cc:241

Recovery chain: version mismatch →ncclNetCheckDeviceVersionreturns an error →ncclNetPluginAssignToCommreturnsisAssigned = false → ncclNetInitcontinues trying the next plugin → may ultimately fall back to the built-in Socket plugin.

Pitfall 2: The profiler thread cannot exit

If the profiler plugin blocks instopEvent, the consumer thread will get stuck inprofilerProgressOps,iterationActiveis always true,ncclProfilerThreadDestroywill wait forever.📎 src/plugin/profiler.cc:1166This is a real deadlock risk.

[Design Inference and Architectural Trade-offs]

Recovery chain:comm->abortFlagis set →waitForActiondetects the abort → returnsCLEANUP_AND_STOP → cleanupAndStopdrains the queue.📎 src/plugin/profiler.cc:1017-1031But if the thread is already stuck in a plugin callback, the abort flag cannot interrupt it—this is the responsibility of the plugin implementer; callbacks must have timeouts.

Pitfall 3: Reference count leak in the tuner plugin

ncclTunerPluginLoadIncrement on successtunerPluginRefCount。📎 src/plugin/tuner.cc:98 ncclTunerPluginUnloadDecrement whencomm->tunerPluginLoadedis true.📎 src/plugin/tuner.cc:111-123If a communication domain loads a tuner buttunerPluginLoadedis accidentally cleared on destruction, the reference count will never return to zero, and the plugin library will never be unloaded.

Chapter Review and Self-Test

Q1: If the loop inncclNetPluginLoadthat "tries from higher versions to lower versions" is changed to "only try the highest version," in what scenario would a previously usable plugin fail to load?

Reference analysis: See📎 src/plugin/net.cc:108-112. The loop iterates overNCCL_NET_VERSION_COUNTversions, from v12 down to v6, and the first one that returns non-null is adopted. If only v12 is tried, then an old plugin that only implements v11 will fail to load.

[Design Inference and Architectural Trade-offs]

This design is for backward compatibility: after the NCCL core is upgraded to support v12, it can still load plugins that only provide v11. Plugin authors are encouraged to provide symbols for multiple versions (see📎 plugins/net/README.md:35-37), so that the same.socan serve multiple NCCL versions.

If the downgrade attempts were removed, old plugins would suddenly become unavailable after users upgrade NCCL, and they could only fall back to the built-in Socket plugin, causing a significant performance drop. This is exactly the purpose of version negotiation.

Q2: InprofilerProgressOps, ifwc <= op->workStarted[ch].data[slot].counteris changed towc == op->workStarted[ch].data[slot].counter, in what high-concurrency scenario would the event never trigger?

Reference analysis: See📎 src/plugin/profiler.cc:969-972. The comment explicitly states that the device wraps aroundMAX_PROFILER_EVENTS_PER_CHANNELslots. If the host consumes more slowly than the device produces, the device may have already overwritten slotwc + Nwith counterwc % MAX_PROFILER_EVENTS_PER_CHANNEL。

At this point the value ofop->workStarted[ch].data[slot].counteriswc + N, whileop->workCounteriswc. Using==for the check will fail, the event will never trigger, and the work object will remain forever in theactivelinked list,inflightonly increasing and never decreasing, eventually exhausting the memory pool.

Using<=handles this situation correctly: as long as the counter written by the device is not less than the expected value, the event is considered ready. This is a typical correctness condition for a "producer-consumer ring buffer."

Q3: If the loop inncclProfilerThreadDestroythat waits foriterationActiveto become false is removed, under what timing would the profiler plugin access an already-freed communication domain context?

Reference analysis: See📎 src/plugin/profiler.cc:1162-1166. The comment states thatncclProfilerPluginFinalizewill destroy the communication domain'sncclProfilerThreadDestroyimmediately afterprofilerContext。

returns. When the consumer thread calls the plugin callback inprofilerProgressOps, what is passed in isop->profilerContext。📎 src/plugin/profiler.cc:938If the destruction thread returns without waiting foriterationActiveto become false,ncclProfilerPluginFinalizewill free the context, while the consumer thread may be using this context to call the plugin—use-after-free.

iterationActiveThe handshake protocol is: the consumer thread sets it totrueinside the lock, then releases the lock to call the plugin, and the destruction thread waits inside the lock for it to return tofalse。📎 src/plugin/profiler.cc:1028📎 src/plugin/profiler.cc:1054-1057This protocol guarantees that the context remains valid during the plugin callback.

After removing the wait, the destruction thread may return just as the consumer thread enters the plugin callback, causing the plugin to receive a dangling pointer. This is a typical "lifetime and concurrent access" race.

The plugin system has moved NCCL from closed to open: network backends, tuning strategies, performance collectors, and configuration sources can all be replaced without modifying the core code. But plugins also introduce new failure surfaces—version mismatches, lifetime races, and reference count leaks. In the next chapter we will enter the RAS and diagnostics subsystem to see how NCCL detects failures, monitors progress, and achieves self-healing in long-running training jobs.

The plugin system draws a clear boundary between NCCL's core communication path and replaceable components. The four types of plugins—net, tuner, profiler, and env—each safely intervene in runtime behavior through registration and reference counting mechanisms. But an extensible communication engine must not only be able to flexibly replace components, but also run stably during long training sessions—when a NIC or GPU fails, how does NCCL detect it, monitor it, and trigger recovery? In the next chapter we will enter the RAS and diagnostic mechanisms to see how reliability in production environments is systematically guaranteed.

CHAPTER 17

Chapter 17: Chapter 17: RAS Mechanisms and Fault Tolerance: Link Failure Detection, Heartbeat, and Graceful Degradation

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 17 / 25

Chapter 17: RAS Mechanisms and Fault Tolerance: Link Failure Detection, Heartbeat, and Graceful Degradation

In the previous chapter, we saw how the plugin system draws a clear boundary between the core communication path and replaceable components, allowing network backends, tuning strategies, and performance collectors to be swapped without modifying core code. But extensibility is only one dimension of production readiness. Another equally hardcore question is: when an AllReduce has been running for 72 hours and a machine's NIC silently fails, how can NCCL detect it, isolate it, and continue? The RAS subsystem is precisely the watershed that takes NCCL from "it runs" to "it's production-ready." This chapter will dissect the design behind fault detection, progress monitoring, and self-healing mechanisms.

17.1 RAS Master Control: A Global Coordinator with One RAS Thread per Process

Intuitive Model

Think of RAS as the "duty room" for the entire job. Each NCCL process (each rank) opens a duty room during initialization, staffed by a dedicated thread. All communicator creation, destruction, and diagnostic requests must first be registered with the duty room; the duty rooms then communicate with each other over an independent RAS network to report "who is still alive and who is dead."

Without this duty room, NCCL could only perceive failures through timeouts on the communication path itself—and timeouts on the communication path are both slow and prone to false positives (a single network jitter could be mistaken for node death). RAS strips "fault perception" out of the data plane and into the control plane, using an independent lightweight heartbeat and diagnostic channel to determine health status.

Data Structures and Memory Layout

The core state of RAS is scattered across the global variables inras.cc. Let's break them down one by one:

VariableTypePurpose
rasInitMutexstd::mutexProtects RAS singleton initialization
rasInitializedboolWhether already initialized
rasInitRefCountintReference count, equal to the number of active comms
rasNetListeningSocketstruct ncclSocketRAS network listening socket
rasNotificationPipe[2]ncclSocketPairDescriptorNotification pipe from local thread → RAS thread
rasPfdsstruct pollfd*Poll array of the main event loop
ncclCommsstruct ncclComm**Array of all communicator pointers

📎 src/ras/ras.cc:49-61defines these global states. Note thatrasInitRefCountusesncclAtomicRefCountIncrementto increment/decrement📎 src/ras/ras.cc:129, whilerasInitializeduses a plain bool plus double-checked locking to protect📎 src/ras/ras.cc:103-105—this is the typical "initialize once, read-only thereafter" pattern.

ncclCommsThe allocation strategy for theRAS_INCREMENT * 8array is worth noting: it does not grow on demand, but expands by📎 src/ras/ras.cc:139-140each time (i.e., 32 slots)nullptr. The array allows📎 src/ras/ras.cc:135-137。

holes (set to null when a comm is destroyed), and a new comm reuses the first hole

Scenario-Driven Walkthrough: From Comm Initialization to RAS Thread StartupncclRasCommInitStep 1:is called.📎 src/ras/ras.cc:101This is the first RAS function called during each comm initializationrasInitialized. It first checks

, and if not initialized, enters the critical section:rasNetListeningSocket1. Initialize📎 src/ras/ras.cc:108-109

with the bootstrap network interface address, setting the port to 0 to let the kernel assign a random one📎 src/ras/ras.cc:113

2. Listen on that socket📎 src/ras/ras.cc:118

3. Create the local notification pipe📎 src/ras/ras.cc:120

4. Initialize the diagnostic subsystemrasThreadMain5. Start the📎 src/ras/ras.cc:121

threadatexit(rasTerminate)6. Register📎 src/ras/ras.cc:126

to ensure cleanup on process exitStep 2: Register the comm.commRegardless of whether this is the first initialization, it writes thencclCommspointer into the📎 src/ras/ras.cc:142arrayncclCommsSorted, and sets📎 src/ras/ras.cc:143to false

—because the array order has changed, the previous sorting is invalidated.Step 3: Backfill the port.rasNetListeningSocket.addrAt the end of the function,myRank->addr 📎 src/ras/ras.cc:146(including the kernel-assigned port) is copied back to

, so the caller can know which port the RAS network is listening on.

rasThreadMainMain Event Loop: Poll-Driven Multiplexing📎 src/ras/ras.cc:633is the heart of the RAS thread📎 src/ras/ras.cc:641-652. It first registers three fixed fds: the notification pipe, the RAS network listening socket, and the client listening socket

code
for (int64_t nextWakeup = 0;;) {
  // 计算超时
  timeoutMs = min(..., 1000);
  nEvents = poll(rasPfds, nRasPfds, timeoutMs);
  // 处理事件
  for (pollIdx...) { ... }
  // 处理各类超时
  rasSocksHandleTimeouts(now, &nextWakeup);
  rasConnsHandleTimeouts(now, &nextWakeup);
  rasNetHandleTimeouts(now, &nextWakeup);
  rasCollsHandleTimeouts(now, &nextWakeup);
}

📎 src/ras/ras.cc:655-728CopytimeoutMsshows this loop. Note that📎 src/ras/ras.cc:664is hard-capped at 1000msnextWakeup—even if

is far away, it must wake up once per second to ensure timely timeout checks.📎 src/ras/ras.cc:684-715Event dispatch logic uses fd values for routingrasLocalHandle: if it's the notification pipe, callrasSocketsHead; if it's a listening socket, accept; otherwise traverse therasClientsHeadand

linked lists to find the corresponding socket to handle.

Local Notification Mechanism: Pipe + Fixed-Length StructurerasNotificationThe local NCCL thread and the RAS thread communicate through a socketpair. The notification structure📎 src/ras/ras.cc:35-46is fixed-lengthstatic_assert, andPIPE_BUF 📎 src/ras/ras.cc:47is used to ensure it does not exceed

—this is to ensure write atomicity (POSIX guarantees that writes smaller than PIPE_BUF are atomic).rasLocalNotifyThe senderrasNotificationMutexuses📎 src/ras/ras.cc:224-237to serialize writes from multiple user threads📎 src/ras/ras.cc:224-237, then loops until all writes are completerasLocalHandle. The receiver📎 src/ras/ras.cc:247-256similarly loops to read the entire structurencclSystemError 📎 src/ras/ras.cc:251-253。

, returning on EOFRAS_ADD_RANKSThree notification types:RAS_RUN_DIAG(new rank joins),RAS_TERMINATE(run diagnostics),📎 src/ras/ras.cc:28-32。

(terminate)

Message Send/Receive: Length Prefix + Incremental Progress📎 src/ras/ras_internal.h:110-117The wire format of RAS messages is "4-byte length + message body"rasConnSendMsg. When sending,📎 src/ras/ras.cc:362-390sends the length first, then the message bodymeta->offset, usingrasMsgRecvto record progress, supporting continuation after a partial send. When receiving,📎 src/ras/ras.cc:393-412。

first receives the length, allocates a buffer according to the length, then receives the message bodyrasMsgAllocThere is a detail here:rasMsgMetaallocates themsgstructure,offsetoffield is at the end of the structure, calculated via📎 src/ras/ras.cc:313-319to compute the offset📎 src/ras/ras.cc:323-328. This "metadata-first" layout allows messages to carry local information such as send progress and enqueue time without occupying the wire format.

Design Considerations

[Design Inference and Architectural Trade-offs]

Why use poll instead of epoll?The O(n) complexity of poll is acceptable in RAS scenarios—the number of RAS connections is far smaller than data-plane connections, and the RAS thread itself is not on the performance-critical path. poll also has better cross-platform support (Windows compatibility).

[Design Inference and Architectural Trade-offs]

Why use a pipe instead of a condition variable for notification?A pipe can be seamlessly integrated into the poll loop, allowing the RAS thread to use a unifiedpollwait for all event sources. If a condition variable were used, an additional mechanism would be needed to wake up poll.

mermaid
flowchart TD
    start["rasThreadMain 启动"] --> reg_pipe["注册通知管道 fd"]
    reg_pipe --> reg_net["注册 RAS 网络监听 fd"]
    reg_net --> reg_client["注册客户端监听 fd"]
    reg_client --> poll["poll(rasPfds, timeout<=1000ms)"]
    poll --> check{"nEvents == -1?"}
    check -->|"是且非 EINTR"| log_err["记录 poll 错误并继续"]
    check -->|"否"| dispatch["遍历 revents 分发事件"]
    log_err --> dispatch
    dispatch --> is_pipe{"fd == 通知管道?"}
    is_pipe -->|"是"| local_handle["rasLocalHandle()"]
    is_pipe -->|"否"| is_net{"fd == RAS 监听?"}
    is_net -->|"是"| accept_net["rasNetAcceptNewSocket()"]
    is_net -->|"否"| is_client{"fd == 客户端监听?"}
    is_client -->|"是"| accept_client["rasClientAcceptNewSocket()"]
    is_client -->|"否"| find_sock["遍历 rasSocketsHead 找匹配 socket"]
    find_sock --> sock_loop["rasSockEventLoop(sock, pollIdx)"]
    local_handle --> terminate{"terminate?"}
    terminate -->|"是"| cleanup["rasThreadCleanup() 并退出"]
    terminate -->|"否"| timeouts
    sock_loop --> timeouts["rasSocksHandleTimeouts / rasConnsHandleTimeouts / rasNetHandleTimeouts / rasCollsHandleTimeouts"]
    accept_net --> timeouts
    accept_client --> timeouts
    timeouts --> poll

17.2 Progress Monitoring: Using DMA to Move GPU Counters to the Host

Intuitive Model

Progress monitoring is like the "tachometer" on a car's dashboard. It doesn't participate in driving (doesn't participate in communication), but continuously copies the GPU's internal progress counters to host memory, allowing the host to determine whether "this communication domain is stuck." Without it, when an AllReduce hangs, you can only see "the program doesn't return," but you can't tell whether the GPU is computing, waiting on the network, or completely deadlocked.

Data Structures and Memory Layout

Each CUDA device corresponds to onencclGpuProgressCounterMonitorworker thread📎 src/ras/progress_monitor.cc:35-52:

FieldTypePurpose
cudaDevintBound CUDA device number
threadstd::threadWorker thread
mutex / cvstd::mutex / condition_variableProtects mutable state and wakeups
running / shouldStopboolThread lifecycle flag
copyInFlightboolWhether a DMA copy is in flight
copyStallWarnedboolWhether an alert has already been issued for this stall
copyStartNsuint64_tStart time of this copy
sideStreamcudaStream_tDedicated non-blocking stream
copyDonecudaEvent_tCopy completion event
warningMutexstd::mutexProtects alert timestamps
lastStaleWarnNs / lastErrorWarnNsuint64_tRate-limiting timestamp
destroyRefsintDestruction reference count
registrationsIntrusive queueList of comms registered to this device

📎 src/ras/progress_monitor.cc:59-62Clarifies the lock order:gpuProgressCounterMonitorsMubeforencclGpuProgressCounterMonitor::mutex. This is a key convention for avoiding deadlocks.

Global arraygpuProgressCounterMonitors[kRasMaxCudaDevices]indexed by device number📎 src/ras/progress_monitor.cc:59-62。

Scenario-Driven Walkthrough: A Single Counter Copy

Step 1: Registration. ncclProgressCounterMonitorInitis called📎 src/ras/progress_monitor.cc:319. IfdeviceCountersBlockis empty, return directly (this comm does not participate in monitoring)📎 src/ras/progress_monitor.cc:323. Otherwise, within the global lock, find or create the worker for that device📎 src/ras/progress_monitor.cc:328-335, then enqueue the comm toregistrations 📎 src/ras/progress_monitor.cc:339。

Step 2: Worker thread startup. createGpuProgressCounterMonitorcreates the worker, setscudaSetDevice, createssideStream(cudaStreamNonBlocking) andcopyDoneevent📎 src/ras/progress_monitor.cc:280-282, and after starting the thread waits up to 2000ms to confirmrunningbecomes true📎 src/ras/progress_monitor.cc:287-303。

Step 3: Loop copy. progressCounterMonitorLoopFirst bind the device and set relaxed stream capture mode (to avoid interfering with the application's graph capture)📎 src/ras/progress_monitor.cc:97-121, then enter the main loop:

1. Wait forpollIntervalMs(default 1000ms)📎 src/ras/progress_monitor.cc:132-136

2. If the previous copy is still in flight, usecudaEventQueryto check📎 src/ras/progress_monitor.cc:140. IfcudaErrorNotReadyand the stale threshold is exceeded (default 5000ms), issue a rate-limited alert📎 src/ras/progress_monitor.cc:141-154

3. Iterate over all registered comms, and for each callcudaMemcpyAsyncto copydeviceCountersBlocktohostCountersBlock 📎 src/ras/progress_monitor.cc:170-185

4. If any copy succeeds, record thecopyDoneevent and setcopyInFlight 📎 src/ras/progress_monitor.cc:194-202

Concurrency Control and Rate Limiting

Alert rate limiting is implemented byprogressCounterMonitorShouldWarn📎 src/ras/progress_monitor.cc:78-87: under the protection ofwarningMutex, check whether more thanwarnIntervalNshas elapsed since the last alert, and only then update and return true. The defaultstaleWarnSecis 600 seconds📎 src/ras/progress_monitor.cc:27, meaning the same type of alert is emitted at most once every 10 minutes.

Parameters have lower-bound clamping: the minimum poll interval is 50ms📎 src/ras/progress_monitor.cc:29, and the minimum stale threshold is 1000ms📎 src/ras/progress_monitor.cc:30. This prevents overly aggressive user configuration from causing CPU spinning.

Destruction: Reference Counting + Stream Synchronization

ncclProgressCounterMonitorDestroyThe destruction logic of📎 src/ras/progress_monitor.cc:352-354:

is one of the most elegant concurrency designs in this chapterregistrations1. Under the global lock + worker lock, remove the comm from📎 src/ras/progress_monitor.cc:368

2. If removal succeeds,destroyRefs++and sethaveDestroyRef 📎 src/ras/progress_monitor.cc:371-372

3. If the registration list becomes empty, remove it from the global array and setshouldStop 📎 src/ras/progress_monitor.cc:373-376

4. After releasing the lock,cudaStreamSynchronize(g->sideStream)drain copies that may still reference this comm's buffer📎 src/ras/progress_monitor.cc:393

5. FinallyreleaseGpuProgressCounterMonitorDestroyRefdecrement the reference count; when it reaches zero and the queue is empty, join the thread and delete📎 src/ras/progress_monitor.cc:219-246

[Design Inference and Architectural Trade-offs]

Why isdestroyRefs?needed? BecausecudaStreamSynchronizeexecutes outside the lock, and during that time another thread may also be destroying the same worker. The reference count ensures that only the last destroyer actually joins and deletes.

mermaid
sequenceDiagram
    participant App as 应用线程
    participant Mon as 监控线程
    participant GPU as CUDA 设备
    App->>Mon: ncclProgressCounterMonitorInit(comm)
    Mon->>Mon: 查找/创建 worker
    Mon->>Mon: registrations 入队 comm
    loop 每 pollIntervalMs
        Mon->>GPU: cudaEventQuery(copyDone)
        GPU-->>Mon: cudaErrorNotReady / cudaSuccess
        Mon->>GPU: cudaMemcpyAsync(hostCounters, deviceCounters, D2H, sideStream)
        Mon->>GPU: cudaEventRecord(copyDone, sideStream)
    end
    App->>Mon: ncclProgressCounterMonitorDestroy(comm)
    Mon->>Mon: registrations 删除 comm, destroyRefs++
    Mon->>GPU: cudaStreamSynchronize(sideStream)
    GPU-->>Mon: 拷贝排空完成
    Mon->>Mon: releaseGpuProgressCounterMonitorDestroyRef
    Mon->>Mon: join 线程, delete worker

Production Pitfalls

Pitfall 1:cudaSetDevicefailure causes monitoring to silently fail.IfcudaSetDevicefails when the thread starts, the worker setsshouldStopand exits📎 src/ras/progress_monitor.cc:97-107, but the comm that registered it still believes monitoring is running. At this point the counter mirror remains stale until the failure is exposed during the Init phase. When troubleshooting, check whether theNCCL_RASlogs contain "progress-counter mirrors will remain stale".

Pitfall 2: graph capture conflict.If the application is performing stream capture when the monitoring thread calls the CUDA API, it will pollute the capture graph. The code usescudaThreadExchangeStreamCaptureMode(cudaStreamCaptureModeRelaxed)to avoid📎 src/ras/progress_monitor.cc:110-111, which is a necessary safeguard.

17.3 Diagnostic Framework: Table-Driven Check Dispatch

Intuitive Model

The diagnostic framework is like a hospital's "health check package." Each check item (GPU model, ECC status, NVLink health, XID errors, etc.) is an independent "check department," and the framework is responsible for collecting the check results from each rank and summarizing them into a report. Without it, operations can only rely onnvidia-smimanually troubleshooting machine by machine, which is completely infeasible on a thousand-GPU cluster.

Data Structure: Check Dispatch Table

At the core is a static dispatch tablerasDiagnosticsChecks 📎 src/ras/diagnostics.cc:63-77, where each entry binds a check ID and two callbacks:collectLocal(local collection) andsummarize(aggregation). 11 checks in total: GPU model, CUDA driver version, ECC, NVLink, NCCL environment, RDMA topology, IOMMU mode, ATS, XID/SXID, NVIDIA driver version, path.

rasDiagnosticsGetCheckPerforms triple validation: ID range, table entry ID match, callback non-null📎 src/ras/diagnostics.cc:104-128. This is defensive programming—preventing table entries from being incorrectly modified, which would lead to calling a null pointer.

Scenario-driven Walkthrough: The Complete Lifecycle of a Single Diagnosis

Step 1: Build the local payload. rasDiagnosticsCollectLocalPeerPayloadFirst write the peer header📎 src/ras/diagnostics.cc:226-227, then iterate over the dispatch table, calling for each entryrasDiagnosticsAppendCheckPayload 📎 src/ras/diagnostics.cc:229-231。

rasDiagnosticsAppendCheckPayloadCallcollectLocalto getrasDiagnosticsLocalData, usencclUniquePtrto take ownership of records📎 src/ras/diagnostics.cc:191-192, validate metadata📎 src/ras/diagnostics.cc:193, if the record count is 0 then skip📎 src/ras/diagnostics.cc:194, otherwise write the check header + record data📎 src/ras/diagnostics.cc:196-201。

Step 2: Initiate collective communication. rasDiagnosticsStartConstructRAS_COLL_DIAGrequest📎 src/ras/diagnostics.cc:532-537, send viarasNetSendCollReq📎 src/ras/diagnostics.cc:539, set client state toRAS_CLIENT_DIAG_FINI 📎 src/ras/diagnostics.cc:541。

Step 3: Merge responses. rasCollDiagMergeAppend each peer's payload to the collective buffer📎 src/ras/diagnostics.cc:310-337. Note that it performs extensive overflow checks: peer count upper limit📎 src/ras/diagnostics.cc:320-324, total size upper limit📎 src/ras/diagnostics.cc:325-328。

Step 4: Aggregation. rasDiagnosticsSummarizePeerPayloadsis a two-pass scan📎 src/ras/diagnostics.cc:399:

  • First pass: validate each peer header and check header, accumulate the record count and byte count for each check type📎 src/ras/diagnostics.cc:418-470
  • Allocate the merge buffer for each check type📎 src/ras/diagnostics.cc:472-476
  • Second pass: copy each peer's records into the corresponding buffer📎 src/ras/diagnostics.cc:479-497
  • Finally call for each check typesummarize 📎 src/ras/diagnostics.cc:499-506

Client State and Cancellation

Diagnostic state is stored inrasDiagnosticsClientState📎 src/ras/diagnostics.cc:242-245, attached torasClient->diagnostics.rasDiagnosticsCancelTargetWhen the client socket is closed, replaces the reporter with noop📎 src/ras/diagnostics.cc:286-293, preventing writes to an already-closed socket after asynchronous diagnosis completes📎 src/ras/diagnostics.cc:48-52。

Design Considerations

〔Design Inference and Architectural Trade-offs〕

Why use a two-pass scan?Because the payload is variable-length; only the first pass can compute how large a buffer each check type needs. A single-pass scan would either require dynamic growth (multiple reallocs) or over-allocation. The two-pass scan trades a single precise allocation for determinism.

Why include in the check headerrecordStride? 📎 src/ras/diagnostics.cc:197Because different checks have different record structure sizes, and aggregation needs to know the stride to correctly copy and validate.rasDiagnosticsAccountCheckRecordsEnforces that the stride is consistent for the same check📎 src/ras/diagnostics.cc:381-385。

mermaid
flowchart TD
    start["rasDiagnosticsStart"] --> build_req["构造 RAS_COLL_DIAG 请求"]
    build_req --> send["rasNetSendCollReq"]
    send --> all_done{"allDone?"}
    all_done -->|"是"| fini["client->status = DIAG_FINI"]
    all_done -->|"否"| in_progress["返回 ncclInProgress"]
    fini --> resume["rasDiagnosticsResume"]
    in_progress --> resume
    resume --> summarize["rasDiagnosticsSummarizePeerPayloads"]
    summarize --> pass1["第一遍: 校验头 + 累计每类记录数"]
    pass1 --> valid{"payload 合法?"}
    valid -->|"否"| err["返回 ncclInternalError"]
    valid -->|"是"| alloc["为每类检查分配合并缓冲区"]
    alloc --> pass2["第二遍: 拷贝各 peer 记录"]
    pass2 --> emit["对每类检查调用 summarize"]
    emit --> finish["reporter.finish + rasCollFree"]

17.4 Peer Management: Sorted Array + Hash Synchronization

Intuitive Model

peers.ccmaintains a "class roster." Each RAS thread keeps an identical copy of the roster, recording each NCCL process's address, PID, and managed GPUs. When a new member joins or someone "goes missing," the change is broadcast over the RAS network. The roster uses a hash value as a version number to avoid full synchronization every time.

Data Structures and Memory Layout

Two core arrays:

  • rasPeers: all known peers, sorted by address📎 src/ras/peers.cc:18-19. Includes dead peers.
  • rasDeadPeers: dead peer addresses, stored separately📎 src/ras/peers.cc:37-38。

Why store dead peers separately? 📎 src/ras/peers.cc:25-28The comments in explain it clearly:rasPeersis essentially static and very large at scale, whilerasDeadPeersis dynamic and much smaller. Storing them separately avoids transmitting the hugerasPeersarray on every synchronization.

rasPeerInfoStructure📎 src/ras/ras_internal.h:110-117:

FieldTypeDescription
addrncclSocketAddressNetwork address (sort key)
pidncclPid_tProcess ID
cudaDevsuint64_tCUDA device bitmask (affected by CUDA_VISIBLE_DEVICES)
nvmlDevsuint64_tNVML device bitmask (not affected)
hostHash / pidHashuint64_tExtracted from comm, minus commHash to make it independent of the communication domain

Two hashesrasPeersHashandrasDeadPeersHashare the core of synchronization📎 src/ras/peers.cc:21📎 src/ras/peers.cc:37-38。

Scenario-driven Walkthrough: A New Rank Joins

Step 1: Conversion. rasRanksConvertToPeersConverts therasRankInitarray intorasPeerInfo 📎 src/ras/peers.cc:104. First sort by address + cudaDev📎 src/ras/peers.cc:114, skip empty addresses📎 src/ras/peers.cc:127-130, merge multi-GPU processes at the same address (bitmask OR)📎 src/ras/peers.cc:134-139。

Step 2: Update the local array. rasPeersUpdateis the most complex merge algorithm in this chapter📎 src/ras/peers.cc:197. It first computes the new array size📎 src/ras/peers.cc:202-229, then merges the two sorted arrays📎 src/ras/peers.cc:244-361. Key point: during the merge, it transformsrankPeersinto a "diff"—keeping only the genuinely new GPU bits📎 src/ras/peers.cc:301-308, and finally clears entries with no contribution📎 src/ras/peers.cc:393-402. This minimizes the amount of broadcast data.

Step 3: Propagation. rasNetUpdatePeersPropagates alongrasNextLinkandrasPrevLinkin both directions📎 src/ras/peers.cc:430-450, then rebuilds connections📎 src/ras/peers.cc:443-444。

Step 4: Send updates. rasConnSendPeersUpdateFirst check the hash📎 src/ras/peers.cc:500-508: if the peer already knows the current hash, skip. The message carriespeersHashanddeadPeersHash 📎 src/ras/peers.cc:521-524; if after merging the receiver's hash still doesn't match, it sends back📎 src/ras/peers.cc:608-653。

Declaration and Propagation of Dead Peers

rasPeerDeclareDeadAdds the address torasDeadPeers, re-sorts and recomputes the hash📎 src/ras/peers.cc:793-812。rasMsgHandleBCDeadPeerHandles broadcast dead peer messages📎 src/ras/ras.cc:578-591: if locally unknown, disconnect and declare dead; otherwise mark*pDone = trueStop re-broadcasting.

rasDeadPeersUpdateUses merge sort to combine the old and new dead peer lists📎 src/ras/peers.cc:838-893. Note that it usesmemmoveinstead ofmemcpy 📎 src/ras/peers.cc:855, because the source and destination may overlap.

Connection Rebuilding: Avoiding Duplicate Connection Races

rasLinkReinitConnsRebuilds link connections after peer updates📎 src/ras/peers.cc:680. Core strategy: initiate the connection from the side with the smaller address📎 src/ras/peers.cc:706-711, avoiding both sides initiating simultaneously and causing duplicates.

rasLinkCalculatePeerComputes the next peer index, skipping dead peers📎 src/ras/peers.cc:743-785. There is an additional optimization for fallback: skip peers on the same node as the previous fallback📎 src/ras/peers.cc:743-785, avoiding waiting one by one when an entire node goes down.

Production Pitfalls

Pitfall 1: The byte-order trap in address comparison. ncclSocketsCompareSorts by address family → address → port📎 src/ras/peers.cc:960-990. The comment points out that you cannot simplymemcmpthe entire structure, because the memory layout order differs from the desired sort order📎 src/ras/peers.cc:957-959. IPv4 addresses and ports can be compared byte by byte under network byte order, but the address family field cannot.

Pitfall 2:myPeerIdxfails.When the array grows,myPeerIdxchanges.📎 src/ras/peers.cc:22-23。rasPeersUpdateUpdate it synchronously during the merge process.📎 src/ras/peers.cc:312📎 src/ras/peers.cc:358If the update fails, fall back to binary search.📎 src/ras/peers.cc:374-388。

[Design Inference and Architectural Trade-offs]

Pitfall 3: Hash collisions cause synchronization omissions.The hash is only used to determine "whether synchronization is needed," not for correctness. Even if a hash collision causes synchronization to be skipped, subsequent keep-alive exchanges will still carry the hash, and it will eventually converge.

mermaid
flowchart LR
    subgraph 输入
        ranks["rasRankInit[]"]
    end
    subgraph 转换
        convert["rasRanksConvertToPeers: 排序+合并同地址"]
        rankPeers["rasPeerInfo[] (rankPeers)"]
    end
    subgraph 合并
        update["rasPeersUpdate: 归并到 rasPeers"]
        diff["rankPeers 改造为差异"]
        hash["重算 rasPeersHash"]
    end
    subgraph 传播
        send["rasConnSendPeersUpdate: 带哈希"]
        recv["rasMsgHandlePeersUpdate: 合并+回发"]
        reinit["rasLinkReinitConns: 重建连接"]
    end
    ranks --> convert --> rankPeers --> update
    update --> diff --> hash
    hash --> send --> recv --> reinit

17.5 Design Considerations: The Boundary Between RAS and the Main Communication Path

The most core design decision of the RAS subsystem iscomplete decoupling from the data plane. RAS threads do not participate in any data movement for collective communication; they only do three things: maintain the peer list, detect connection health, and perform diagnostics. This decoupling brings several benefits:

1. Fault isolation: A crash of the RAS thread will not directly cause communication failure (although it will lose fault-awareness capability).

2. No performance loss: RAS heartbeat and synchronization traffic use an independent network and do not consume data plane bandwidth.

3. Observability: Diagnostics and monitoring can be performed in parallel while communication is in progress.

The cost isstate consistencychallenges: The comm state seen by RAS may lag behind the data plane.ncclRasCommInitandncclRasCommFinithroughncclCommsMutexprotect📎 src/ras/ras.cc:77-77, but when the RAS thread reads, it only takes a snapshot and does not provide strong consistency guarantees.

Another key design istimeout layering。ras_internal.hdefines a complete set of timeout constants📎 src/ras/ras_internal.h:214-249: keep-alive interval 1 second, warning threshold 5 seconds, error threshold 20 seconds, peer death threshold 60 seconds. This layering allows the system to take different actions at different severity levels—first warn, then try an alternate connection, and only finally declare death.

17.6 Chapter Summary

This chapter broke down the four core modules of the NCCL RAS subsystem:

  • ras.cc: A singleton RAS thread + poll event loop, receiving local notifications through a pipe and exchanging messages with other ranks over an independent network.
  • progress_monitor.cc: One worker thread per device, using DMA to move GPU progress counters to the host, with throttling warnings and reference-counted destruction.
  • diagnostics.cc: A table-driven check dispatch framework, with two-pass scanning to aggregate diagnostic payloads from each rank.
  • peers.cc: Peer list management with a sorted array + hash synchronization, with dead peers stored separately to save bandwidth.

Chapter Review and Self-Test

Q1:rasLocalNotifyusesrasNotificationMutexserialized writes, butrasLocalHandlereads without a corresponding lock. Why is this safe? Ifstatic_assert(sizeof(struct rasNotification) <= PIPE_BUF)is removed, in what scenarios would problems occur?

Reference Analysis: Safety comes from POSIX's guarantee of atomicity for pipe writes—writes smaller thanPIPE_BUFare atomic.📎 src/ras/ras.cc:47。rasLocalNotify's loop write📎 src/ras/ras.cc:224-237will not interleave with other writes when it can be completed in a single write.rasLocalHandle's loop read📎 src/ras/ras.cc:247-256may read partial data, but because writes are atomic, what is read must be a prefix of a complete message, and the next read can fill in the rest.

Removestatic_assert, ifrasNotificationexceedsPIPE_BUF, the write may be split into multiple non-atomic writes. When two threads write concurrently, their bytes may interleave, causing the RAS thread to read malformed data formed by concatenating two notifications.msg.typemay come from thread A whilemsg.addRanks.rankscomes from thread B, triggeringrasLocalHandle's unknown type branch📎 src/ras/ras.cc:267-269or, worse, a wild pointer dereference.

Q2:ncclProgressCounterMonitorDestroyexecutescudaStreamSynchronize 📎 src/ras/progress_monitor.cc:381-400only after releasing the lock. What happens if another thread also calls Destroy to destroy the same comm during synchronization?destroyRefsHow can the problem be prevented?

Reference Analysis:destroyRefsis a reference count that prevents the worker from being deleted too early. After the first thread deletes the comm,destroyRefs++ 📎 src/ras/progress_monitor.cc:371, at this pointhaveDestroyRef = true. When the second thread tries to delete the same comm,ncclIntruQueueDeletereturns nullptr (already deleted),haveDestroyRefremains false📎 src/ras/progress_monitor.cc:368, and synchronization and release are skipped directly.

After the first thread completescudaStreamSynchronize, it callsreleaseGpuProgressCounterMonitorDestroyRef 📎 src/ras/progress_monitor.cc:402, decrementingdestroyRefsto 0, and only when the registration queue is empty does it actually join the thread and delete📎 src/ras/progress_monitor.cc:225。

If there were nodestroyRefs, the first thread might have its worker released by the second thread'sdelete gduring synchronization, causing a use-after-free. Note thatreleaseGpuProgressCounterMonitorDestroyRefdecrements under the global lock + worker lock📎 src/ras/progress_monitor.cc:222-225, ensuring the atomicity of checkingregistrationsis empty anddestroyRefs == 0.

Q3:rasDiagnosticsSummarizePeerPayloadsDuring the first pass scan, validatecheckHeader->payloadBytes != checkHeader->nRecords * checkHeader->recordStride 📎 src/ras/diagnostics.cc:451-454. If some malicious or corrupted peer sendsrecordStride = 0andnRecords = 0, will this validation pass? What happens afterward?

Reference Analysis:recordStride <= 0will be intercepted by the first condition📎 src/ras/diagnostics.cc:451, returningncclInternalError. SorecordStride = 0will not pass.

But ifrecordStride > 0andnRecords = 0, thenpayloadBytes = 0, and validation passes.rasDiagnosticsAccountCheckRecordsFornRecords == 0directly returns success📎 src/ras/diagnostics.cc:378, without updatingcombined. During subsequent allocation,recordsBytes == 0does not allocate📎 src/ras/diagnostics.cc:473, and during copy,payloadBytes > 0is false and skips📎 src/ras/diagnostics.cc:490. Ultimatelysummarizereceivesrecords = nullptr, recordsBytes = 0, and the summarize implementation of each check needs to handle empty input.

The real risk is innRecords > INT_MAX / recordStride's check📎 src/ras/diagnostics.cc:453—this preventsnRecords * recordStrideinteger overflow from bypassing the equality validation. If this check is removed, an attacker can constructnRecords = 2^31, recordStride = 2, the product overflows to 0, equal topayloadBytes = 0, and after passing validationrasDiagnosticsAccountCheckRecordswill accumulate a hugenRecords, causing out-of-bounds allocation or copying later.

RAS gives NCCL fault awareness and self-healing capability during long training runs, but it relies on a control network independent of the data plane. In the next chapter, we will enter the memory management subsystem and see how NCCL optimizes memory allocation and RDMA registration overhead through allocators, registration caches, and user buffer registration—this is the third pillar beyond performance and reliability.

The design principles running through this chapter are: decoupling the control plane from the data plane, using hashes for state versioning, layered timeout handling, and using reference counting to protect object lifetimes under concurrency. These principles allow RAS to achieve fault detection and self-healing without dragging down communication performance. And another key pillar supporting communication performance—memory management—likewise requires careful engineering trade-offs: Why does NCCL need to register memory before communication? How does the registration cache affect performance? In the next chapter, we will dive into the allocator, the registration cache, and user buffer registration to uncover the answers to these questions.

CHAPTER 18

Chapter 18: Chapter 18: Memory Allocation and Device Memory Management: Allocator, Registration Cache, and User-Registered Memory Optimization

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 18 / 25

Chapter 18: Memory Allocation and Device Memory Management: Allocator, Registration Cache, and User-Registered Memory Optimization

In the previous chapter, we saw how the RAS subsystem runs independently of the data plane on the control plane, using hashes for versioning and reference counting to protect object lifetimes. This chapter enters NCCL's third pillar—memory management. The upper bound of communication performance often does not depend on the algorithm itself, but on "whether the data can be directly read and written by the NIC." To this end, NCCL builds a three-layer mechanism: at the bottom layer, it usesncclSpaceandncclShadowPoolto manage the address space and shadow objects; at the middle layer, it usesncclMemManagerto track the import/export and suspend/resume of dynamic memory; at the upper layer, it usesncclCommRegisterto register user buffers into the cache, avoiding repeated pinning of memory for every communication. This chapter will dismantle these three mechanisms layer by layer and answer "why NCCL needs to register memory before communication" and "how the registration cache affects performance."

18.1 ncclSpace: Slicing the Address Space into Alternating Full/Empty Segments

Intuitive Model

Imagine an infinitely long line of parking space numbers, starting at 0 and extending to the right. Some spaces have cars parked in them (allocated), while others are empty (unallocated).ncclSpaceis the "parking space status record book" for this number line—it does not record every space, but only the "boundary points where the status flips." Without it, when managing the virtual address range of symmetric memory, NCCL would have to maintain a flag bit for every byte, making memory overhead proportional to the address space, which is completely unacceptable.

Data Structure and Memory Layout

ncclSpaceThe definition of is extremely minimal📎 src/include/allocator.h:20-24:

c
struct ncclSpace {
  int count;        // cuts[] 中有效元素个数
  int capacity;     // cuts[] 已分配容量
  int64_t* cuts;    // 升序排列的边界点数组
};

The core insight is stated very clearly in the source code comments📎 src/allocator.cc:151-153:cuts[]splits the non-negative integer axis into alternating "full" and "empty" segments, with cut points arranged in ascending order. The segment after the last cut point must be empty (the unallocated frontier). From this, we can derive the formula for determining whether theith segment is full:

code
isFull(i) = (i%2 != ncuts%2)

The meaning of this formula is: the full/empty state of a segment is jointly determined by "the parity of the segment index" and "the parity of the total number of cut points." Whenncutsis even, segment 0 (beforecuts[0]) is empty; whenncutsis odd, segment 0 is full. This invariant runs through the entire module.

Step-by-Step Walkthrough: How a Single Allocation Changes cuts[]

Scenario: initiallyncclSpaceis empty (count=0), callncclSpaceTryAlloc(a, limit=1000, size=100, align=1, &outOffset)。

Step 1: Locate the first empty segment 📎 src/allocator.cc:209。i = a->count % 2, at this pointcount=0, soi=0, and scanning starts from segment 0.

Step 2: Compute the segment boundaries 📎 src/allocator.cc:212-213。i==0whenlo=0;i==a->countwhenhi=limit=1000. So the empty segment is[0, 1000)。

Step 3: Align and check capacity 📎 src/allocator.cc:214-215。off = alignUp(0, 1) = 0,0 + 100 <= 1000holds, and the allocation succeeds.

Step 4: Insert cut points 📎 src/allocator.cc:217-223. Becausei==0(insertion at the head), take the slow pathinsertSegment(a, 0, 0, 100)。insertSegmentatindex=0insert two cut pointslo=0, hi=100 📎 src/allocator.cc:172-174, then perform "adjacent duplicate value filtering"📎 src/allocator.cc:185-203. The filtering logic is very elegant: it scans with read and write cursors, and when it encounters duplicate values, it moves the write cursor back, deleting duplicate pairs—because a duplicate pair means an empty segment is sandwiched between two full segments and can be merged. But leading zeros are a special case and can be deleted separately📎 src/allocator.cc:182-184。

After allocationcuts = [0, 100],count=2. At this pointisFull(0) = (0%2 != 2%2) = false, segment 0 ([0,0), empty) is empty; segment 1 ([0,100)) is full. Correct.

Step 5: Free 📎 src/allocator.cc:239-267. CallncclSpaceFree(a, 0, 100). First check whethercuts[count-1] <= offsetholds📎 src/allocator.cc:231-237, that is,100 <= 0is false, so continue. Locate the first full segmenti = 1 - count%2 = 1 - 0 = 1 📎 src/allocator.cc:246,cuts[1]=100 > 0, soi=1。lo = cuts[0] = 0,hi = cuts[1] = 100. Checkoffset < lo || hi < offset+size 📎 src/allocator.cc:252,0<0false,100<100false, pass. Becauselo==offsetandoffset+size==hi, neither fast path is satisfied (the first requiresoffset+size != hi, the second requireslo != offset), so take the slow pathinsertSegment(a, 1, 0, 100) 📎 src/allocator.cc:264. After insertioncuts = [0, 0, 100, 100], and after filtering it becomes[],count=0. Back to the initial state.

This "insert then filter" design avoids complex segment merging logic during allocation/free, concentrating the complexity ininsertSegmentin one place.

Design Considerations and Production Pitfalls

Why use int64_t instead of size_t?BecausencclSpacemanages "offsets" rather than "pointers," offsets may be negative (although this does not happen in actual use), and it needs to match the width of CUDA'sCUdeviceptr. Using a signed type makes out-of-bounds issues easier to spot during debugging.

Performance Pitfall:ncclSpaceFreeThe comment directly states, "This could be binary search, but since allocate is linear there's no point"📎 src/allocator.cc:245. This means both allocation and free are O(n) scans. If a communication domain frequently allocates and frees a large number of small segments,cuts[]will grow, and every operation will become slower. In production environments, registered buffers should be reused as much as possible rather than repeatedly registered/unregistered.

Alignment Overflow Risk:alignUp(lo, align)whenlois close toINT64_MAXandalignis large, overflow may occur. The source code does not explicitly check this becauselimitis guaranteed by the caller to be within a reasonable range.

18.2 ncclShadowPool: Paired Management of Device Objects and Host Shadows

Intuitive Model

GPU kernels run on the device and cannot directly access C++ objects in host memory (such as the metadata inncclDevComm).ncclShadowPoolIt acts like a "translator": it allocates a block of device memory for each device-side object, while simultaneously allocating a corresponding "shadow" memory block on the host side, and maintains a "device address → host address" mapping table. When the host needs to modify the configuration of a device object, it first modifies the host shadow, then copies it to the device. Without it, every time a kernel needs to read metadata it would have to pull it from the host viacudaMemcpy, resulting in unacceptably high latency.

Data Structures and Memory Layout

Two core structs📎 src/allocator.cc:272-277:

c
struct ncclShadowPage {   // 最多 64 个对象的连续块
  struct ncclShadowPage* next;
  int objSize;
  uint64_t freeMask;      // 位图,1=空闲,0=已占用
  void* devObjs;
};
struct ncclShadowObject {
  struct ncclShadowObject* next;
  void* devObj;
  void* hostObj;
  struct ncclShadowPage* page;  // null 表示直接分配在 CUDA mempool
};

ncclShadowPoolitself📎 src/include/allocator.h:42-47:

c
struct ncclShadowPool {
  int count, hbits;                       // 对象数、哈希位数
  struct ncclShadowObject** table;        // 哈希桶数组
  cudaMemPool_t memPool;                  // 可选的 CUDA 内存池
  struct ncclShadowPage* pages;           // 页链表
};

Key design points:freeMaskis uint64_t, so each page holds at most 64 objects. This is not an arbitrary choice—64 bits is exactly the width of a cache line,popFirstOneBitand a single__builtin_ctzllinstruction can be used to find the first free slot without looping.

Hash Table Growth Strategy: Source code comment "Maintain 2:1 object:bucket ratio"📎 src/allocator.cc:368, meaning it expands when the number of objects exceeds twice the number of buckets. Initialhbits=4(16 buckets)📎 src/allocator.cc:363, doubling each time.

Step-by-Step Walkthrough: How a Single Allocation Chooses Between Page or Direct

Scenario:ncclShadowPoolAlloc(pool, size=1024, &devObj, &hostObj, stream)。

Step 1: Lazy initialization 📎 src/allocator.cc:347-366. Ifhbits==0, first query whether the device supports memory pools📎 src/allocator.cc:352, and if supported, createcudaMemPool_t, setmaxSizeto the parameterSHADOW_MEMPOOL_MAX_SIZE(default 1GB)📎 src/allocator.cc:359. Then allocate a hash table with 16 buckets.

Step 2: Check whether expansion is needed 📎 src/allocator.cc:369-386. Ifcount+1 > 2<<hbits, allocate a double-sized bucket array, traverse the old table and reinsert (hashInsertusingncclHashPointerto compute the bucket index📎 src/allocator.cc:333-337), and free the old table.

Step 3: Decide whether to take the page path or the direct path 📎 src/allocator.cc:390. The condition is(64<<10)/size >= 3, i.e., whensize <= 21845, take the page path. Forsize=1024,65536/1024=64 >= 3, take the page path.

Step 4: Compute the in-page object size 📎 src/allocator.cc:391-392。shift = max(0, log2Down(1024)+1-4) = max(0, 10+1-4) = 7。pageObjSize = ((1024 + 127) >> 7) << 7 = 1024. That is, the in-page object size is aligned to a power of 2, rounded up to a multiple of 128 bytes.

Step 5: Find or create a page 📎 src/allocator.cc:393-415. Traverse thepool->pageslinked list to find a page withobjSize == pageObjSize. If none exists, create a new page:pageSize = min(65536, 64*1024) = 65536,freeMask = uint64_t(-1) >> (64 - 65536/1024) = uint64_t(-1) >> 0 = 全 1(all 64 slots empty)📎 src/allocator.cc:400. UsecudaMallocFromPoolAsyncorcudaMallocto allocate device memory📎 src/allocator.cc:403-404, andcudaMemsetAsynczero out📎 src/allocator.cc:405。

Step 6: Take a slot from the page 📎 src/allocator.cc:408-412。popFirstOneBit(&page->freeMask)to find the first free bit,devObj = page->devObjs + slot * pageObjSize. IffreeMaskbecomes 0 (page full), remove the page from the free list📎 src/allocator.cc:411。

Step 7: Allocate the host shadow object 📎 src/allocator.cc:423-428。malloc(sizeof(ncclShadowObject) + alignof(max_align_t)-1 + size), note that here extraalignof(max_align_t)-1bytes are allocated for alignment padding.hostObj = alignUp((char*)(obj+1), alignof(max_align_t)), i.e., after the object header, align to the maximum alignment boundary. Thenmemset(hostObj, 0, size)zero out.

Step 8: Insert into the hash table and update the count 📎 src/allocator.cc:429-430。

Concurrency Control and Hardware Interaction

ncclShadowPoolitselfhas no lock. This means it can only be used in a single-threaded context, or mutual exclusion must be guaranteed by the caller. From NCCL's actual usage, it is mainly called during the communication domain initialization phase, which is single-threaded.

cudaMallocFromPoolAsyncandcudaFreeAsyncare asynchronous operations, relying on thestreamparameter to guarantee ordering📎 src/allocator.cc:403,459。ncclShadowPoolDestructis called after all resources are releasedcudaStreamSynchronize(stream) 📎 src/allocator.cc:333-337, ensuring that all asynchronous frees complete before destroying the memory pool.

Production Pitfall Guide

Pitfall 1: Memory waste caused by in-page object size alignment。pageObjSizeis aligned to a power of 2; ifsize=1000,shift = log2Down(1000)+1-4 = 9+1-4 = 6,pageObjSize = ((1000+63)>>6)<<6 = 1024. Each object wastes 24 bytes, and 64 objects in a page waste 1536 bytes. For a large number of small objects, this overhead cannot be ignored.

Pitfall 2:ncclShadowPoolFreeBehavior when an object cannot be found 📎 src/allocator.cc:442-445. It returnsncclInternalErrorand prints a warning, butdoes not release any resources. If the caller ignores the return value, it will cause a memory leak. Production code must check the return value.

Pitfall 3:ncclShadowPoolDestructInfreeMask==0, a page with 📎 src/allocator.cc:301-306is reclaimedfreeMask. Note that herepool->pagesis set to 1 (rather than all 1s), meaning only the first slot is marked as free. This is to put the "full page" back into the

linked list, but the other slots in the page are still occupied—in fact, these objects are about to be released, so this operation is safe. However, if there is concurrent access during destruction, an inconsistent state will be read.

18.3 ncclMemManager: Reference Counting and Suspend/Resume for Dynamic Memory

Intuitive ModelncclMemManagerTraining tasks may run for days, during which the GPU may be preempted by other tasks, or checkpoints may need to be taken.

It acts like a "memory steward": it records all dynamically allocated memory (scratch/offload), and when needed "suspends" GPU memory (unmaps physical pages, retains virtual addresses), backs up the data to the CPU, and upon resumption reallocates physical pages, remaps, and restores the data. Without it, after a task is preempted it can only start over from the beginning, wasting hours of training progress.

ncclMemManagerData Structures and Memory Layout📎 src/mem_manager.cc:32-60:

Core fields of(inferred from the initialization code)Field
entriesncclDynMemEntry*Type
numEntriesintMeaning
releasedintHead of the dynamic memory entry linked list
refCountintLinked list length
totalPersistsize_t0=active, 1=suspended
totalScratchsize_tReference count (multiple comms can share)
totalOffloadsize_tTotal persistent memory (atomic)
cpuBackupUsagesize_tTotal scratch memory (atomic)
lockstd::mutexTotal offload memory (atomic)
initializedintTotal CPU backup memory

Protects the entries linked list:lockAtomic flag to prevent accessing a destroyed mutexstd::mutexKey design of the memory layoutncclMemManageris ancclCalloc, but📎 src/mem_manager.cc:39is allocated with~mutex() 📎 src/mem_manager.cc:120(C style), so placement new must be used to explicitly construct

, andmust be explicitly called during destruction. This is a classic pitfall of mixed C/C++ programming.totalPersistDivision of labor between atomic variables and locksentries: Statistical fields (lock, etc.) are updated with atomic operations and do not need locks;ncclCommMemStatsthe linked list is protected by📎 src/mem_manager.cc:1117-1130. In this way, statistical queries (

Step-by-Step Walkthrough: The Complete Suspend and Resume Flow

Suspend Flow ncclCommMemSuspend 📎 src/mem_manager.cc:418-540:

Step 1: Pre-checks 📎 src/mem_manager.cc:419-430. Check whether the memory manager is disabled, whether comm is empty, and whether it is already suspended.

Step 2: Device Synchronization and Barrier 📎 src/mem_manager.cc:440-441。cudaDeviceSynchronize()Ensure all GPU operations are complete, thenbootstrapBarrierEnsure all ranks are synchronized. The barrier tag is0xBEEF。

Step 3: First Pass — Unmap All Peer-Imported Buffers 📎 src/mem_manager.cc:444-465. For eachisImportedFromPeer && state==Activeentry, callcuMemUnmapto unmap📎 src/mem_manager.cc:451, release the handle📎 src/mem_manager.cc:456, and change the state toReleased。

Step 4: Second Pass — Offload Local Memory 📎 src/mem_manager.cc:468-526. Skip peer-imported and already-released entries. ForncclMemOffloadtype, first allocate a CPU backup📎 src/mem_manager.cc:484, thencudaMemcpycopy from GPU to CPU📎 src/mem_manager.cc:492. ForncclMemScratchtype, only accumulate statistics. Then close the shareable FD📎 src/mem_manager.cc:508-513,cuMemUnmap 📎 src/mem_manager.cc:516,cuMemRelease 📎 src/mem_manager.cc:519, and change the state toReleased。

Step 5: Mark as Suspended 📎 src/mem_manager.cc:528。

Resume Flow ncclCommMemResume 📎 src/mem_manager.cc:550-942:

Step 1: Restore Local Memory 📎 src/mem_manager.cc:577-668. For each!isImportedFromPeer && state==Releasedentry, re-cuMemCreate 📎 src/mem_manager.cc:599,ncclCuMemMapAndSetAccessmap to the same virtual address📎 src/mem_manager.cc:602, restore peer access permissions📎 src/mem_manager.cc:610-626, restore data from the CPU backup for offload types📎 src/mem_manager.cc:632-643, and re-export the FABRIC handle📎 src/mem_manager.cc:646-658。

Step 2: Barrier Synchronization 📎 src/mem_manager.cc:671-679. The tag is still0xBEEF。

Step 3: Exchange New Handle Information 📎 src/mem_manager.cc:688-816. Count how many local buffers each rank needs to broadcast📎 src/mem_manager.cc:689-696, usebootstrapAllGatherto exchange counts📎 src/mem_manager.cc:710, compute offsets📎 src/mem_manager.cc:724-728, then firstbootstrapSendthenbootstrapRecv(the comment explicitly states "send first, then receive to avoid deadlock"📎 src/mem_manager.cc:783)。

Step 4: Re-import Peer Buffers 📎 src/mem_manager.cc:822-911. For eachisImportedFromPeer && state==Releasedentry, look up the matching handle information in the exchange results📎 src/mem_manager.cc:829-835. For POSIX FD type, check whether hostHash is the same📎 src/mem_manager.cc:853-859, then obtain the FD through the proxy📎 src/mem_manager.cc:866,cuMemImportFromShareableHandleimport📎 src/mem_manager.cc:873. For FABRIC type, import directly📎 src/mem_manager.cc:878. ThenncclCuMemMapAndSetAccessremap📎 src/mem_manager.cc:893。

Step 5: Final Barrier 📎 src/mem_manager.cc:916-928. The tag is0xCAFE, distinguished from the earlier0xBEEF.

Concurrency Control and Hardware Interaction

Reference Counting Protects the Lifecycle:ncclMemManagerDestroyFirst decrementrefCount 📎 src/mem_manager.cc:76, and if it is still greater than 0, only clear the pointer for the current comm📎 src/mem_manager.cc:81, without releasing resources. This allows multiple comms to share the same memory manager (such as in the split_share scenario).

Atomic initialized Flag: Check before all operationsCOMPILER_ATOMIC_LOAD(&manager->initialized, memory_order_acquire) 📎 src/mem_manager.cc:136,242,338,358, to prevent accessing a destroyed mutex. During destruction, usememory_order_releaseto store 0📎 src/mem_manager.cc:87, ensuring that previous write operations are visible to other threads.

Use of the CUDA VMM API:cuMemCreate/cuMemMap/cuMemUnmap/cuMemReleaseis the CUDA virtual memory management API, which allows physical memory and virtual addresses to be separated. This is the foundation of suspend/resume — during suspend, unmap the physical pages but retain the virtual addresses; during resume, remap to the same virtual addresses, so that all established pointer relationships do not need to be modified.

Production Pitfall Guide

Pitfall 1: The split_share communication domain does not support suspend 📎 src/mem_manager.cc:1014-1018. IfrefCount > 1, directly returnncclInvalidUsage. Because when multiple comms share a memory manager, suspending one comm will affect the memory of other comms.

Pitfall 2: POSIX FD becomes invalid across nodes 📎 src/mem_manager.cc:853-859. POSIX file descriptors are only valid within the same node, and must be skipped when resuming across nodes. The source code useshostHashcomparison to determine whether they are on the same node.

Pitfall 3: Keep the backup when offload data restoration fails 📎 src/mem_manager.cc:635. IfcudaMemcpyrestoring from CPU to GPU fails, the source code prints a warning and keepscpuBackup, without releasing it. This is to give the caller a chance to retry, but if no retry occurs, CPU memory will leak.

Pitfall 4:ncclMemUntrackDynamicuse-after-free risk in. The source code finds the entry while holding the lock, saves the necessary information, releases the entry📎 src/mem_manager.cc:302, and then updates statistics outside the lock📎 src/mem_manager.cc:311-327. This order is correct, but if theinfopointer points to the caller's stack memory and the caller reads it outside the lock, you need to ensure thatinfo's lifetime covers the entire function.

mermaid
flowchart TD
    start["ncclCommMemSuspend(comm)"] --> check{"manager->released?"}
    check -->|"是"| err1["返回 ncclInvalidUsage"]
    check -->|"否"| sync["cudaDeviceSynchronize()"]
    sync --> barrier1["bootstrapBarrier(tag=0xBEEF)"]
    barrier1 --> pass1["第一遍: 遍历 entries"]
    pass1 --> cond1{"isImportedFromPeer && Active?"}
    cond1 -->|"是"| unmap1["cuMemUnmap + cuMemRelease"]
    cond1 -->|"否"| skip1["跳过"]
    unmap1 --> pass2["第二遍: 遍历 entries"]
    skip1 --> pass2
    pass2 --> cond2{"memType == Offload?"}
    cond2 -->|"是"| backup["ncclCudaHostCalloc + cudaMemcpy D2H"]
    cond2 -->|"否"| scratch["累加 releasedScratch"]
    backup --> unmap2["cuMemUnmap + cuMemRelease"]
    scratch --> unmap2
    unmap2 --> mark["manager->released = 1"]
    mark --> done["返回 ncclSuccess"]
    err1 --> done

The figure above shows the control flow of the suspend process. Note two key branches: the first pass only handles peer-imported buffers, and the second pass only handles local buffers. The order cannot be reversed — you must first release references to peer memory, and then release local memory.

18.4 Registration Cache: How ncclRegister Avoids Repeated Pinning

Intuitive Model

For the NIC to directly read and write GPU memory (GPUDirect RDMA), the memory must first be "registered" — telling the NIC "you can directly access this address." The registration process involves pinning pages and establishing IOMMU mappings, and is very expensive (millisecond-level). If re-registration happens on every AllReduce, the latency of small-message communication will be completely overwhelmed by registration overhead.ncclRegisteris a "registration cache": it records already-registered address ranges in an ordered array, and the next time it encounters the same or a contained buffer, it directly reuses it without re-registering.

Data Structure and Memory Layout

ncclRegCacheThe core ofslotsis an ordered arrayncclReg*。ncclReg, where each element is

's key fields (inferred from usage):FieldType
begAddruintptr_tMeaning
endAddruintptr_tPage-aligned start address
localRefsintPage-aligned end address
graphRefsintLocal reference count
stateintGraph reference count
netHandleHeadncclRegNetHandles*Registration state bits (NET/NVLS/COLLNET/IPC)
ipcInfosncclIpcInfo**IPC information array

Page alignment:begAddr = (uintptr_t)data & -pageSize 📎 src/register/register.cc:31,endAddr = ((uintptr_t)data + size + pageSize - 1) & -pageSize 📎 src/register/register.cc:32。-pageSizeispageSize's two's complement, equivalent to "rounding down to a multiple of pageSize". The reason for this is: the minimum granularity of registration is a page, so even if only 1 byte is registered, the entire page must be registered.

Step-by-Step Walkthrough: How a single registration hits the cache

Scenario:ncclCommRegister(comm, buff=0x7f0000001000, size=4096, &handle)。

Step 1: Parameter check and page alignment 📎 src/register/register.cc:18-24。CommCheckValidate comm validity. AssumepageSize=4096,begAddr = 0x7f0000001000 & -4096 = 0x7f0000001000,endAddr = (0x7f0000001000 + 4096 + 4095) & -4096 = 0x7f0000002000。

Step 2: System memory check 📎 src/register/register.cc:36-64. IfncclCuMemEnable(), query the address range and memory type. IfmemType == CU_MEMORYTYPE_HOST, it indicates CPU memory, skip registration📎 src/register/register.cc:58-61. Otherwise check whether there is a Sysmem segment📎 src/register/register.cc:50-55。

Step 3: Traverse the cache to find the insertion position 📎 src/register/register.cc:66-89. Loopslotstarting from 0:

  • Ifslot == population(reached the end) orbegAddr < slots[slot]->begAddr(the current address is before the cache entry), it indicates a new entry needs to be created📎 src/register/register.cc:67。
  • Ifslots[slot]->begAddr <= begAddr && slots[slot]->endAddr >= endAddr, it indicates the current buffer is fully contained by an existing entry, directly increment the reference count📎 src/register/register.cc:83-87。

Step 4: Create a new entry 📎 src/register/register.cc:68-82. If the cache is full, expand it (initially 32, then double)📎 src/register/register.cc:70. Usememmoveatslotposition to make room📎 src/register/register.cc:73,ncclCallocallocate a new entry📎 src/register/register.cc:74, setbegAddr/endAddr, according toisGraphsetgraphRefsorlocalRefsto 1📎 src/register/register.cc:78-79,population++, return handle.

Step 5: Deregistration 📎 src/register/register.cc:172-195。commDeregisterFirst find the slot corresponding to the handle📎 src/register/register.cc:180, decrement the reference count📎 src/register/register.cc:185-186. If there are still references, return directly📎 src/register/register.cc:187. Otherwise callregCleanupto clean up all underlying registrations📎 src/register/register.cc:188, free the entry, usememmoveto fill the hole📎 src/register/register.cc:190,population--。

Design considerations and production pitfalls

Why use a sorted array instead of a hash table?Because registration queries are "range containment" queries, not exact matches. A sorted array supports binary search (although the source code uses linear scan), and has good memory locality. A hash table cannot efficiently handle queries like "is this address contained by some larger range".

regCleanupStatus bit design of 📎 src/register/register.cc:95-134。stateis a bitmask, where each bit corresponds to a registration type (NET/NVLS/COLLNET/IPC). During cleanup, check bit by bit and only clean up completed registrations. This design allows partial registration success and partial failure—for example, network registration succeeds but IPC registration fails, and cleanup only cleans up the network part.

Production pitfall: the registration cache is unaware of memory release. If the user registers a buffer and then, without deregistering,cudaFreeit, the cache still retains this entry. The next allocation may reuse the same address, causing a cache hit but the actual memory is already invalid. NCCL's convention is: registration and deregistration must be paired, and the user is responsible for ensuring the memory is not freed during registration.

ncclCommRegisterSkip conditions of 📎 src/register/register.cc:150-159. IfLocalRegister=0orP2pUsesMemcpy=1, directly returnNULLhandle. This means that under certain configurations (such as P2P using memcpy instead of RDMA), registration is completely skipped. The caller must check whether the handle is NULL.

18.5 Collective communication registration: how coll_reg chooses registration strategies for different algorithms

Intuitive model

Different collective communication algorithms take different transport paths: NVLS uses NVLink SHARP, Ring uses P2P or the network, and Tree uses a tree topology. Each path requires a different registration method: NVLS needs to be registered with NVLS hardware, the network needs to be registered with the NIC, and IPC needs to be registered with the peer GPU.coll_reg.ccis the "registration strategy router": it decides which registration functions to call based on the algorithm, protocol, and buffer type. Without it, each algorithm would have to implement registration logic itself, resulting in duplicated code and being error-prone.

Step-by-Step Walkthrough: Registration decision for the Ring algorithm

Scenario:ncclRegisterCollBuffers(comm, info, outRegBufSend, outRegBufRecv, cleanupQueue, regNeedConnect), whereinfo->algorithm == NCCL_ALGO_RING,info->protocol == NCCL_PROTO_SIMPLE。

Step 1: Pre-checks 📎 src/register/coll_reg.cc:155-157. SetregBufType = NCCL_REGULAR_BUFFER,regNeedConnect = true. IfLocalRegister=0and it is not persistent graph registration, exit directly.

Step 2: Enter the Ring branch 📎 src/register/coll_reg.cc:338. InitializerecvRegRecord/sendRegRecordto NULL, allocatesendNetConns/sendNetHandles/recvNetConns/recvNetHandles/srecvNetHandlesarray📎 src/register/coll_reg.cc:356-360。

Step 3: Find existing registration records 📎 src/register/coll_reg.cc:351-355。ncclRegFindSearch the cache for recv/send buffers. If recv is not found and it is not persistent graph registration, exit📎 src/register/coll_reg.cc:352. If cross-node and send is not found and it is not persistent graph registration, exit📎 src/register/coll_reg.cc:354。

Step 4: Traverse all channels to collect peers 📎 src/register/coll_reg.cc:362-393. For each channel, checkring.prevandring.next. If the connection flag containsNCCL_DIRECT_NIC, record it torecvNetConns/sendNetConns 📎 src/register/coll_reg.cc:370-379. If it containsNCCL_P2P_READ | NCCL_P2P_WRITE, add the peer topeerRanksarray📎 src/register/coll_reg.cc:382-391。

Step 5: IPC registration 📎 src/register/coll_reg.cc:394-407. IfnPeers > 0 && comm->isAllDirectP2p, first try graph registration📎 src/register/coll_reg.cc:395-399, and if it fails, try local registration📎 src/register/coll_reg.cc:400-403. If successful, setregBufType = NCCL_IPC_REG_BUFFER 📎 src/register/coll_reg.cc:406。

Step 6: Network registration 📎 src/register/coll_reg.cc:409-457. Check!comm->useNetPXN && comm->useGdr && netDeviceType != UNPACKand not AllReduce's PreMulSum/SumPostDiv📎 src/register/coll_reg.cc:415-418. First try graph registration📎 src/register/coll_reg.cc:419-430, and if it fails, local registration📎 src/register/coll_reg.cc:431-442. If successful, setregBufType |= NCCL_NET_REG_BUFFER, save the handle array📎 src/register/coll_reg.cc:445-452。

Step 7: Adjust the number of channels 📎 src/register/coll_reg.cc:551-554. If only IPC registration exists and it is single-node and the number of channels is between 17-24, reduce it to 16. This is to match the bandwidth characteristics after IPC registration.

Design considerations and production pitfalls

Why are the registration orders of NVLS and Ring opposite?The NVLS branch first tries graph registration and then local registration📎 src/register/coll_reg.cc:86-94, while the Ring branch first does local and then graph📎 src/register/coll_reg.cc:395-403. This is because NVLS graph registration is more likely to succeed (NVLS hardware has optimizations for persistent buffers), while Ring local registration is more lightweight.

isMloPartBufRdmaCapableGlobal decision of 📎 src/register/coll_reg.cc:14-37. The comment emphasizes "Registration decision must be global, using communicator-wide guarantees"📎 src/register/coll_reg.cc:20. This means that even if a certain rank's buffer supports RDMA, as long as one rank within the communication domain does not support it, the entire communication domain will not register. This is to avoid inconsistency caused by some ranks registering and others not registering.

Production pitfall: Silent degradation when registration fails。ncclRegisterCollBuffersWhen registration fails, no error is reported; it simply does not setregBufTypethe corresponding bit. This means communication can still work, just with degraded performance. In production environments, if performance does not meet expectations, you should checkNCCL_REGthe logs to confirm whether registration succeeded.

mermaid
flowchart LR
    subgraph input["输入"]
        task["ncclTaskColl<br/>algorithm=RING<br/>protocol=SIMPLE"]
    end
    subgraph ipc["IPC 注册路径"]
        find["ncclRegFind<br/>查找缓存"]
        collect["遍历 channel<br/>收集 peerRanks"]
        ipcReg["ncclIpcLocalRegisterBuffer<br/>或 GraphRegister"]
    end
    subgraph net["网络注册路径"]
        checkGdr{"useGdr &&<br/>!useNetPXN?"}
        netReg["ncclNetLocalRegisterBuffer<br/>或 GraphRegister"]
    end
    subgraph output["输出"]
        regType["info->regBufType<br/>NCCL_IPC_REG_BUFFER<br/>NCCL_NET_REG_BUFFER"]
        handles["info->sendNetHandles<br/>info->recvNetHandles"]
    end
    task --> find
    find --> collect
    collect --> ipcReg
    ipcReg --> regType
    find --> checkGdr
    checkGdr -->|"是"| netReg
    checkGdr -->|"否"| regType
    netReg --> regType
    netReg --> handles

The above diagram shows two parallel registration paths under the Ring algorithm: the IPC path handles same-node P2P connections, and the network path handles cross-node RDMA connections. The two paths execute independently and ultimately both converge toinfo->regBufType。

18.6 Production Pitfall Avoidance and Failure Recovery Chain

Pitfall 1: Interaction between registration cache and memory pool

When usingncclMemAllocto allocate memory, the underlying implementation goes through the CUDA VMM API📎 src/allocator.cc:38-94. The physical memory created by this allocation method carries thegpuDirectRDMACapableflag📎 src/allocator.cc:54, meaning it natively supports RDMA. But whenncclMemFreereleases it, if the memory manager has already been destroyed, it will go through thecudaFreefallback path📎 src/allocator.cc:130-132. This may cause VMM-allocated memory to be incorrectly freed withcudaFree. In production environments, you must ensure thatncclMemAlloc/ncclMemFreeare used in pairs, and do not release after the memory manager has been destroyed.

Pitfall 2: Communication requests during suspension

ncclCommMemSuspendDuring execution, what happens if new communication requests arrive? The source code callscudaDeviceSynchronize() 📎 src/mem_manager.cc:440before suspension to ensure all queued GPU operations complete. However, if host-side communication requests are being enqueued, there is no explicit protection. In production environments, you should stop all communication threads before suspension, or use group semantics to ensure suspension operations are serialized with other operations.

Pitfall 3: Compatibility of FABRIC handle

ncclMemAllocOn CUDA 12.3+, it will attempt to use FABRIC handle📎 src/allocator.cc:60-71. IfcuMemCreatereturnsCUDA_ERROR_NOT_PERMITTEDorCUDA_ERROR_NOT_SUPPORTED, it will fall back to POSIX FD📎 src/allocator.cc:63-65. But during recovery, if the handle type is FABRIC but export fails, it will directly report an error and unmap📎 src/mem_manager.cc:649-655. This means that in mixed environments (some GPUs support FABRIC, some do not), suspend/resume may fail.

Pitfall 4: Reference count leak

ncclRegisterEach cache hit increments the reference count📎 src/register/register.cc:84-85. If the caller registers N times but only deregisters M times (M < N), the reference count will never reach zero,regCleanupwill never be called, and the underlying registration resources will leak. Production code must strictly pairncclCommRegister/ncclCommDeregister。

mermaid
sequenceDiagram
    participant App as 应用层
    participant Reg as ncclRegister
    participant Cache as ncclRegCache
    participant Net as ncclNetLocalRegisterBuffer
    participant GPU as CUDA Driver

    App->>Reg: ncclCommRegister(comm, buff, size, &handle)
    Reg->>Reg: begAddr = data & -pageSize
    Reg->>Cache: 遍历 slots 查找包含范围
    alt 缓存命中
        Cache-->>Reg: 返回已有 ncclReg*
        Reg->>Reg: localRefs++
    else 缓存未命中
        Reg->>Cache: memmove 腾出插入位置
        Reg->>Cache: ncclCalloc 新条目
        Reg->>Reg: localRefs = 1
    end
    Reg-->>App: 返回 handle
    App->>Net: 首次注册时调用
    Net->>GPU: cuMemExportToShareableHandle
    GPU-->>Net: 返回 handle
    Net-->>App: 注册完成

Chapter Review and Self-Test

Q1: If thencclSpaceFreeinif (a->count == 0 || a->cuts[a->count - 1] <= offset)check📎 src/allocator.cc:231-237is removed, under what scenarios would out-of-bounds access be triggered?

Reference analysis: This check has two purposes. First,a->count == 0prevents empty array accesscuts[-1]. Second,a->cuts[a->count-1] <= offsetpreventsoffsetfrom exceeding the allocated range. If removed, whencount == 0,a->cuts[a->count - 1]will readcuts[-1], which is undefined behavior and may read heap metadata or trigger a segmentation fault. More subtly, even ifcount > 0, ifoffsetis greater than the last split point, the subsequentwhile (a->cuts[i] <= offset) i += 2loop📎 src/allocator.cc:247will keep incrementingiuntil out of bounds, becausecuts[]does not contain any element greater thanoffset. The triggering scenario in production is: the caller passes in an offset that was never allocated (for example, calling free again after the buffer has been externally released), orncclSpaceis concurrently modified causing inconsistent state. The fix is to keep this check and printoffsetandcountwhen returning an error for easier troubleshooting.

Q2: ncclMemManagerDestroyInrefCount, if📎 src/mem_manager.cc:78-83is still greater than 0 after decrementing, only the current comm's pointer is cleared without releasing resourcesncclMemTrack. If at this time another comm is calling

, what will happen?:ncclMemTrackReference analysismanager->initialized 📎 src/mem_manager.cc:136First checksrefCount > 0. Sinceinitialized = 0does not setmanager->lock, the check passes. Then it will acquireentriesand modify the📎 src/mem_manager.cc:188-192linked listrefCount > 0. This is safe becausencclMemManagerDestroymeans at least one comm still holds a reference, and the memory manager will not be destroyed. The real risk is: if the last comm callsrefCount,initialized = 0 📎 src/mem_manager.cc:87decrements to 0, it will setncclMemTrackand release all resources. If at this time another thread is ininitializedand has already passed themanager->lockcheck but has not yet acquired the lock, it will access the already-freedmemory_order_acquire/release, causing use-after-free. The source code mitigates this problem through

pairing, but strictly speaking there is still a race window. In production environments, you should ensure all communication threads have stopped before destroying the memory manager.ncclCommMemResumeQ3: In📎 src/mem_manager.cc:853-859, POSIX FD type peer buffers are skipped when crossing nodesrestoredPeerCount. If all peer buffers are skipped,manager->releasedis 0, but📎 src/mem_manager.cc:913is still set to 0

. What consequences will this cause?:manager->released = 0Reference analysisstateindicates that the memory manager considers recovery complete. But if peer buffers were skipped, theirncclDynMemStateReleased,handleis stillncclCommMemStatsis still 0. If subsequent communication accesses these buffers, it will trigger a CUDA error (accessing unmapped virtual addresses). More seriously,ncclStatGpuMemSuspendedquerying📎 src/mem_manager.cc:1130will return 0 (active)entriesIn this case, the correct approach is to mark cross-node POSIX FD entries as unrecoverable at suspend time, or to return an error at resume time rather than silently skipping them. In production, if POSIX FD is used across nodes, FABRIC handles should be used instead, or suspend/resume should be ensured to occur only within a single node.

Memory management is the invisible pillar of NCCL performance:ncclSpaceIt uses a minimalist split-point array to manage the address space,ncclShadowPooluses a 64-bit bitmap and hash table to manage device/host object pairing,ncclMemManageruses reference counting and the CUDA VMM API to implement suspend/resume,ncclRegisterand uses a sorted array to cache registration results and avoid repeated pinning. These four layers of mechanisms together support the key performance guarantee that "memory does not need to be re-registered before communication." In the next chapter, we will move on to the device-side communicator and ABI compatibility, and seedevcommhow these host-side memory layouts are mapped into structures accessible to GPU kernels.

The figure above shows the registration timing: on a cache hit, only the reference count is incremented and the underlying registration is not called; only on a cache miss is a new entry created and the underlying registration triggered. At this point, the host-side memory management mechanism is already clear. But communication ultimately happens on the GPU, and the kernel needs direct access to the peer rank's address and connection state. The next chapter will move on to the device-side communicator and ABI compatibility, to see how devcomm maps the metadata of the host-side ncclComm into structures accessible on the device side, and how the versioned ABI ensures compatibility between old and new kernels and the library.

CHAPTER 19

Chapter 19: Chapter 19: Device-side communication domain and ABI compatibility: the communication contract between devcomm and kernel

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 19 / 25

Chapter 19: Device-side communication domain and ABI compatibility: the communication contract between devcomm and kernel

In the previous chapter, we saw that the host-side ncclMemManager uses reference counting and the CUDA VMM API to manage the lifecycle of communication buffers. But the place where communication actually happens is the GPU kernel—threads in the kernel need to know: which rank am I? At which virtual address is the peer rank's buffer? Is the connection ready? This information is in the host-side ncclComm structure, but the kernel cannot directly dereference host pointers. If NCCL made the kernel obtain this metadata through parameter passing or global memory queries every time, then every communication would incur extra latency and bandwidth overhead. Worse, once kernel code is compiled, the field offsets it accesses are fixed—if the layout of ncclComm changes after a library upgrade, the old kernel will read incorrect data. This is the core problem that devcomm must solve: map the key metadata of the host-side communication domain, with a stable, versioned memory layout, into structures accessible on the device side. The files devcomm_v22902.cc, devcomm_v22907.cc, devcomm_v23000.cc, and devcomm_v23100.cc under the src/devcomm directory are the concrete implementations of this versioned ABI. Each file corresponds to a NCCL version range, defines the exact memory layout of ncclDevComm within that range, and specifies the field-copying logic between old and new versions. This chapter will break down in turn: what the core data structures of the device-side communicator look like, how the registration and matching mechanism of the versioned ABI works, how field-level conversion is performed between old and new versions, and the boundaries and pitfalls of this mechanism in production environments.

I. Core structure of the device-side communicator: the memory layout of ncclDevComm

Intuitive model

Think ofncclDevCommas a "workstation card": when each GPU kernel starts, it receives a card printed with "you are rank 3, there are 8 ranks in total, there are 4 ranks in your LSA group, and the peer buffer base address is at 0x7f...". This card must be small enough (to fit into kernel parameters), and it must contain all key information. If this card did not exist, the kernel could only rely on repeatedly passing parameters from the host side and reassembling them for every communication—high latency and error-prone.

Data structures and memory layout

TakingncclDevComm_v23000as an example, its complete definition is in📎 src/devcomm/devcomm_v23000.cc:25-62:

c
struct ncclDevComm_v23000 {
  unsigned int magic;          // 偏移 0,魔数校验
  unsigned int version;        // 偏移 4,版本号

  int rank, nRanks;            // 偏移 8, 12
  uint32_t nRanks_rcp32;       // 偏移 16,nRanks 的倒数(定点数)
  int lsaRank, lsaSize;        // 偏移 20, 24
  uint32_t lsaSize_rcp32;      // 偏移 28

  ncclDevCommWindowTable_t windowTable;  // 偏移 32
  ncclWindow_t resourceWindow;           // 偏移 40
  ncclResourceWindow_vidmem_v23000_t resourceWindow_inlined;  // 偏移 48
  ncclGinBarrierHandle_t hybridWorldGinBarrier;  // 偏移 112
  ...
};

📎 src/devcomm/devcomm_v23000.cc:64-93It uses a series ofstatic_assertto pin down the offset of each field. This is not decoration—it is a compile-time contract for ABI compatibility. If the offset of a field moves because of a change in the compiler's alignment strategy, compilation will fail, rather than producing hard-to-debug memory misalignment at runtime.

The design motivations for several key fields:

[Design inference and architectural trade-offs]

nRanks_rcp32andlsaSize_rcp32: this isnRanksandlsaSizeThe reciprocal of , represented as a 32-bit fixed-point number. When the kernel performs the division operation for rank-to-buffer offset, the GPU's integer division is very slow, and using multiplication by the reciprocal followed by a shift can significantly speed it up. This is a classic case of "trading space for time" — storing 4 extra bytes to save dozens of clock cycles per division.

resourceWindow_inlined: This is an inline window descriptor, of typencclResourceWindow_vidmem_v23000_t. Note📎 src/devcomm/devcomm_v23000.cc:11-18its definition in :

c
typedef struct ncclResourceWindow_vidmem_v23000 {
  char reserved1[8];
  char* lsaFlatBase;
  char reserved2[8];
  uint32_t stride4G;
  uint32_t mcOffset4K;
  char reserved3[32];  // NOTE: shrunk from 40 in 2.30u1 to reclaim 8 bytes
} ncclResourceWindow_vidmem_v23000_t;

Herereserved1、reserved2、reserved3is apadding field, used as a placeholder. Why is padding needed? BecausencclDevComm_v23000's layout must maintain consistent offsets with a certain "baseline version." Even if some fields are no longer used in the current version, placeholders must be retained to keep the offsets of subsequent fields unchanged.📎 src/devcomm/devcomm_v23000.cc:11-18's comment explicitly states: 2.30u1 shrinksreserved3from 40 bytes to 32 bytes, freeing up 8 bytes forhybridWorldGinBarrier. This is alayout rearrangement— by shrinking the padding area, new fields are inserted without changing the overall size.

📎 src/devcomm/devcomm_v23000.cc:11-18'sstatic_assertfurther verifies:lsaFlatBase、stride4G、mcOffset4KThe offsets of the three fields must match the "current version's"ncclWindow_vidmem, and the entire struct size is 64 bytes. This meansresourceWindow_inlinedisbinary-compatiblebetween v23000 and the current version — it can be directly memcpy'd.

The family of versioned structs

ComparingncclDevComm_v22902 📎 src/devcomm/devcomm_v22902.cc:38-62andncclDevComm_v22907 📎 src/devcomm/devcomm_v22907.cc:13-41, we can see the evolution of fields:

Fieldv22902v22907v23000
magic/versionNoneNoneYes (offset 0/4)
ginContextCountuint8_tuint32_tuint32_t
ginNetDeviceTypes[4][NCCL_GIN_MAX_CONNECTIONS][NCCL_GIN_MAX_CONNECTIONS]
ginIsRailedNoneboolSplit intoginConnectionsRailed + ginContextsRailed
hybridWorldGinBarrierNoneNoneYes (offset 112)
Struct size200224240
[Design inference and architectural trade-offs]

This evolution path reveals NCCL's versioning strategy:only add fields when necessary, and make use of padding areas as much as possible. From v22902 to v22907,ginSignalBase、ginCounterBase、ginContextBase、ginIsRailedand other GIN-related fields were added; from v22907 to v23000,magic/versionvalidation fields andhybridWorldGinBarrierwere added, whileginIsRailedwas split into two more precise flag bits.

---

II. Registration and matching of versioned ABI: the ncclDevCommCompat struct

Intuitive model

Think of the versioned ABI as a set of "translation plugins": when an application is compiled with NCCL 2.29.2 but linked at runtime against the 2.31.0 library, the library needs to know "what kind ofncclDevCommlayout the 2.29.2 kernel expects," and then translate the current version'sncclDevComminto the old layout. Each version range corresponds to a translation plugin, registered in a global table.

Core struct: ncclDevCommCompat

At the end of eachdevcomm_vXXXXX.ccfile, ancclDevCommCompatstruct is defined. Taking v23000 as an example📎 src/devcomm/devcomm_v23000.cc:192-199:

c
struct ncclDevCommCompat ncclDevCommCompat_v23000 = {
  NCCL_VERSION(2, 30, 0),               // minVersion
  NCCL_VERSION(2, 30, 7),               // maxVersion
  nullptr,                              // commPropertiesFilter
  ncclDevCommRequirementsFilter_v23000, // devCommRequirementsFilter
  ncclDevCommCopyNewToOld_v23000,       // devCommCopyNewToOld
  ncclDevCommCopyOldToNew_v23000,       // devCommCopyOldToNew
};

The meaning of the six fields:

1. minVersion / maxVersion: the version range this plugin is responsible for. v23000 covers 2.30.0 to 2.30.7.

2. commPropertiesFilter: an optional filter, used to adjust the capability flags exposed to older versions inncclCommProperties. v23000 sets it tonullptr, indicating no filtering is needed.

3. devCommRequirementsFilter: checks whether the device-side resources requested by the application are compatible with the old version. The v23000 implementation📎 src/devcomm/devcomm_v23000.cc:95-98simply copiesginTypefromcomm->sharedRestoreqs。

4. devCommCopyNewToOld: copies the current version'sncclDevCommto the old version layout.

5. devCommCopyOldToNew: copies the old version layout back to the current version.

Division of version ranges

The version ranges of the four files:

FileminVersionmaxVersionNotes
devcomm_v22902.cc2.29.22.29.3The earliest versioned implementation
devcomm_v22907.cc2.29.52.29.7Added GIN fields, but does not provide GIN backward compatibility
devcomm_v23000.cc2.30.02.30.7Added magic/version validation
devcomm_v23100.cc2.31.0Current versionAll filters are nullptr, indicating full compatibility

📎 src/devcomm/devcomm_v23100.cc:10-17All callbacks ofnullptr's v23100 plugin arencclDevComm, which means that starting from 2.31.0,

's layout has stabilized and requires no conversion.

[Design inference and architectural trade-offs]

Note that there is a "gap" in the version ranges between v22902 and v22907 (2.29.4 and 2.29.6 have no corresponding plugins). This may be because these versions were not released, or their layouts are completely identical to adjacent versions and can be reused.

Matching processncclCommGetDeviceHandleWhen an application calls

or a similar API, NCCL needs to:reqs->version)。

1. Read the NCCL version number embedded at compile time in the application (viancclDevCommCompat2. Look up the plugin covering that version in the global

table.devCommCopyNewToOld3. If found, call the plugin's

to convert the current layout to the old layout.

4. If not found, return an error or use default behavior.

mermaid
flowchart TD
    start["应用请求设备侧通信器"] --> read_ver["读取 reqs->version<br/>(应用编译时版本)"]
    read_ver --> find_compat{"在 ncclDevCommCompat 表中<br/>查找覆盖该版本的插件?"}
    find_compat -->|找到| check_filter["调用 devCommRequirementsFilter<br/>检查资源请求兼容性"]
    find_compat -->|未找到| err_unsupported["返回 ncclInvalidUsage<br/>版本不兼容"]
    check_filter --> filter_ok{"过滤器返回<br/>ncclSuccess?"}
    filter_ok -->|是| copy_new_to_old["调用 devCommCopyNewToOld<br/>把当前布局转为旧布局"]
    filter_ok -->|否| err_gin["返回 ncclInvalidUsage<br/>GIN 资源不兼容"]
    copy_new_to_old --> done["返回旧布局 ncclDevComm"]
    err_unsupported --> done_err["应用收到错误"]
    err_gin --> done_err

---

Copy

III. Field-level conversion: how old and new layouts are converted to each other

Intuitive modelncclDevCommVersion conversion is like "translation": the new version'srankis a modern Chinese article, and the old version's layout is classical Chinese. The translator needs to map field by field — some fields correspond directly (ranktoginConnectionStride > 1), some fields require "free translation" (ginConnectionsRailed = truetranslated to

), and some fields do not exist in the old version (simply discarded).

NewToOld conversion: from the current version to the old versionncclDevCommCopyNewToOld_v23000Taking📎 src/devcomm/devcomm_v23000.cc:114-152:

c
static ncclResult_t ncclDevCommCopyNewToOld_v23000(ncclComm_t comm, void* oldDevComm,
                                                   struct ncclDevComm const* newDevComm) {
  struct ncclDevComm_v23000* old = (struct ncclDevComm_v23000*)oldDevComm;

  memset(old, '\0', sizeof(*old));  // 先清零,防止未初始化字段泄露
  old->magic = newDevComm->magic;
  old->version = newDevComm->version;
  old->rank = newDevComm->rank;
  ...
  old->ginConnectionsRailed = (newDevComm->ginConnectionStride > 1);
  old->ginStrongLegacySignals = newDevComm->ginStrongLegacySignals;
  old->ginContextsRailed = (newDevComm->ginContextStride > 1);
  ...
}

Copy

1. memsetKey steps: 📎 src/devcomm/devcomm_v23000.cc:118Zeroing

2. : this is a safety measure — the old struct may contain fields that do not exist in the new version, and zeroing prevents uninitialized memory from leaking to the device side.:rank、nRanks、lsaRankDirect field copy

3. and other direct assignments.Inline window conversionncclDevCommCopyResourceWindowNewToOld_v23000 📎 src/devcomm/devcomm_v23000.cc:100-105: calllsaFlatBase、stride4G、mcOffset4K。

4. , copying field by field:ginConnectionsRailed = (newDevComm->ginConnectionStride > 1) 📎 src/devcomm/devcomm_v23000.cc:142Semantic conversionginConnectionStride. The new version uses

5. (an integer stride) to indicate whether it is railed, while the old version uses a boolean value. When the stride is greater than 1, it indicates that the connection is railed.:memcpyArray copyginNetDeviceTypescopies theginHandlesand📎 src/devcomm/devcomm_v23000.cc:135-136。

arrays

OldToNew conversion: from the old version to the current version📎 src/devcomm/devcomm_v23000.cc:154-190:

c
static ncclResult_t ncclDevCommCopyOldToNew_v23000(ncclComm_t comm, struct ncclDevComm* newDevComm,
                                                   void const* oldDevComm) {
  struct ncclDevComm_v23000 const* old = (struct ncclDevComm_v23000 const*)oldDevComm;

  newDevComm->magic = old->magic;
  ...
  newDevComm->ginConnectionStride = old->ginConnectionsRailed ? old->lsaSize : 1;
  newDevComm->ginContextStride = old->ginContextsRailed ? old->lsaSize : 1;
  ...
}
Copy

[Design inference and architectural trade-offs]📎 src/devcomm/devcomm_v23000.cc:180-181NoteginConnectionsRailed's semantic conversion: if in the old versionginConnectionStrideis true, then the new version'slsaSize; otherwise set to 1. Here we uselsaSizeas the step size because in railed mode, ranks within each LSA group share a single GIN connection, and the step size equals the size of the LSA group.

Special handling for v22902

ncclDevCommCopyOldToNew_v22902 📎 src/devcomm/devcomm_v22902.cc:149-167There is an important comment:

c
// Note: this callback will be used with v22907 as well because, prior to 2.30.0, ncclDevComm was unversioned,
// so v22902 and v22907 variants are indistinguishable.
[Design inference and architectural trade-offs]

This means that before 2.30.0,ncclDevCommdoes not have themagic/versionfield, so the library cannot distinguish whether an old struct is v22902 or v22907. Therefore, v22907'sdevCommCopyOldToNewis set tonullptr 📎 src/devcomm/devcomm_v22907.cc:128, and the v22902 version is actually used. Since neither supports GIN backward compatibility, the differences in GIN-related fields do not affect correctness.

Versioning of resource windows

ncclWindow_vidmem_v22902The definition ofdevcomm_v22902.his in📎 src/devcomm/devcomm_v22902.cc:141(the content of this file is not provided in this chapter), but from📎 src/devcomm/devcomm_v22902.cc:164andncclDevCommCopyResourceWindow_v22902we can see that v22902 usesdevcomm_v22902.hfor window conversion. This function is declared in

📎 src/devcomm/devcomm_v23000.cc:11-18, and its specific implementation is not shown in this chapter's source code.static_assert's

---

verifies that the window layout of v23000 is consistent with the current version, so the conversion function for v23000 can directly copy field by field.

IV. Capability filtering and resource checking: preventing old kernels from accessing unsupported features

Intuitive modelncclDevCommVersion conversion is not just "moving fields around"—it also needs to check whether the old version supports the features requested by the application. For example, a kernel compiled with 2.29.2 requests GIN resources, but in 2.29.2's

layout, the GIN fields are incomplete, and direct conversion would cause the kernel to read garbage data. Therefore, a "filter" is needed to intercept such requests before conversion.

ncclCommPropertiesFilter_v22907 📎 src/devcomm/devcomm_v22907.cc:69-77:

c
static ncclResult_t ncclCommPropertiesFilter_v22907(ncclComm_t comm, struct ncclCommProperties* props) {
  // We don't provide backwards compatibility for GIN with 2.29.7.  If a communicator needs it, we indicate that
  // the Device API is not available.
  props->deviceApiSupport = (props->deviceApiSupport && ncclTeamLsa(comm).nRanks == comm->nRanks);
  props->ginType = NCCL_GIN_TYPE_NONE;
  props->railedGinType = NCCL_GIN_TYPE_NONE;
  return ncclSuccess;
}

Copy

1. deviceApiSupportThree operations:Downgrade

2. ginType: if the number of ranks in the LSA group is not equal to the total number of ranks (that is, cross-node communication exists), disable the device API. This is because GIN in 2.29.7 does not support cross-node.Set to NONE

3. railedGinType: explicitly tell the application that "this version does not support GIN."Set to NONE

ncclCommPropertiesFilter_v22902 📎 src/devcomm/devcomm_v22902.cc:86-96: same as above.

c
// v22902 ncclCommProperties is _almost_ compatible with newer ones, with the exception of ginType, which in that
// version was based on uint_8, not an int.
((struct ncclCommProperties_v22902*)props)->ginType = NCCL_GIN_TYPE_NONE_v22902;

📎 src/devcomm/devcomm_v22902.cc:13-17Copy

c
typedef enum : uint8_t {
  NCCL_GIN_TYPE_NONE_v22902 = 0,
  NCCL_GIN_TYPE_PROXY_v22902 = 2,
  NCCL_GIN_TYPE_GDAKI_v22902 = 3,
} ncclGinType_t_v22902;

Copyuint8_tNote that this isginTypetype, whereas in the new versionintisprops. Therefore, the filter for v22902 needs to castncclCommProperties_v22902*touint8_t, and then write it intoginType。📎 src/devcomm/devcomm_v22902.cc:35-36type'sstatic_assert'sginTypeverifies that

is at offset 34, and the struct size is 40 bytes.

ncclDevCommRequirementsFilter_v22907 📎 src/devcomm/devcomm_v22907.cc:79-98devCommRequirementsFilter: resource request checking

c
static ncclResult_t ncclDevCommRequirementsFilter_v22907(ncclComm_t comm, ncclDevCommRequirements_t* reqs) {
  bool requestedGinResources =
    reqs->ginSignalCount > 0 || reqs->ginCounterCount > 0 || reqs->barrierCount > 0 || reqs->railGinBarrierCount > 0;
  struct ncclDevResourceRequirements* node = reqs->resourceRequirementsList;
  while (!requestedGinResources && node != nullptr) {
    requestedGinResources = node->ginSignalCount > 0 || node->ginCounterCount > 0;
    node = node->next;
  }
  if (requestedGinResources && (reqs->ginConnectionType != NCCL_GIN_CONNECTION_NONE || reqs->ginForceEnable)) {
    // 打印警告并返回错误
    return ncclInvalidUsage;
  }
  return ncclSuccess;
}

Copy

1. The logic is divided into two steps::reqs->ginSignalCount、ginCounterCount、barrierCount、railGinBarrierCountCheck top-level requests

2. If any is greater than 0, it indicates that GIN resources have been requested.Traverse the resource requirement linked listresourceRequirementsList: if there is no top-level request, continue traversing theginSignalCountlinked list and check each node'sginCounterCount。

andginConnectionTypeIf GIN resources are indeed requested, andNONEis notginForceEnableorncclInvalidUsageis true, then return

ncclDevCommRequirementsFilter_v22902 📎 src/devcomm/devcomm_v22902.cc:98-126and print a warning, indicating that the application needs to be recompiled.barrierCountis more complex. In addition to the GIN check, it also handles the semantic change of

c
// Prior to 2.29.4, a non-zero barrierCount did not imply GIN, but it does since.
if (reqs->barrierCount) {
  reqs->lsaBarrierCount = std::max(reqs->lsaBarrierCount, reqs->barrierCount);
  reqs->barrierCount = 0;
}
// Strangely, neither did railGinBarrierCount.
reqs->railGinBarrierCount = 0;
Copy

[Design inference and architectural trade-offs]barrierCountBefore 2.29.4,barrierCountonly represented LSA barrier and did not imply a GIN requirement. Starting from 2.29.4,barrierCountimplies a GIN requirement. To be compatible with older versions, the filter convertslsaBarrierCounttobarrierCount, and clearsrailGinBarrierCount。

and

mermaid
sequenceDiagram
    participant App as 应用层
    participant Host as Host 侧 NCCL 库
    participant Compat as ncclDevCommCompat 插件
    participant Dev as 设备侧 ncclDevComm

    App->>Host: ncclCommGetDeviceHandle(comm, &devComm)
    Host->>Host: 读取 reqs->version(应用编译版本)
    Host->>Compat: 查找覆盖该版本的插件
    Compat-->>Host: 返回 ncclDevCommCompat_vXXXXX
    Host->>Compat: devCommRequirementsFilter(comm, reqs)
    alt 请求了不支持的 GIN 资源
        Compat-->>Host: ncclInvalidUsage
        Host-->>App: 返回错误 + 警告日志
    else 资源兼容
        Compat-->>Host: ncclSuccess
        Host->>Compat: devCommCopyNewToOld(comm, oldDevComm, newDevComm)
        Compat->>Compat: memset(old, 0, sizeof(*old))
        Compat->>Compat: 逐字段拷贝 + 语义转换
        Compat-->>Host: ncclSuccess
        Host->>Dev: 返回旧布局 ncclDevComm
        Dev-->>App: 设备侧可访问的通信器
    end

---

Copy

V. Production pitfall guide and failure recovery chain

Pitfall 1: Conflict between GIN resource requests and old-version kernelsScenarioncclGinPut)。

: The application is compiled with NCCL 2.29.2, but at runtime links against the 2.31.0 library. The application calls GIN-related device-side APIs in the kernel (such as:ncclDevCommRequirementsFilter_v22902 📎 src/devcomm/devcomm_v22902.cc:98-126What happensginForceEnabledetectsginSignalCount > 0orncclInvalidUsage, returns

code
The application was compiled with too old version of NCCL. It was compiled with NCCL version 2.29.2, but is
running with NCCL library version 2.31.0. Because of its use of GIN device kernels, it needs to be recompiled,
preferably with the same NCCL version that it will be running with.

CopyRoot causencclDevComm_v22902: In 2.29.2'sginContextCount、ginNetDeviceTypes、ginHandleslayout, the GIN fields (

, etc.) are incompatible with the 2.31.0 layout. If forced conversion is performed, the kernel will read the wrong offsets, resulting in undefined behavior.Correct approach

: The application must be recompiled with the same (or compatible) NCCL version as the runtime library. If recompilation is not possible, avoid using GIN APIs in the kernel.

Pitfall 2: Device API silently disabled during cross-node communicationScenarioncclTeamLsa(comm).nRanks != comm->nRanks)。

: The application is compiled with 2.29.7, and the communication domain contains cross-node ranks (:ncclCommPropertiesFilter_v22907 📎 src/devcomm/devcomm_v22907.cc:69-77What happensprops->deviceApiSupportsetsfalseto

. If the application checks this flag, it will know that the device API is unavailable; but if it does not check and directly calls the device-side API, undefined behavior will result.Root cause

: GIN in 2.29.7 does not support cross-node. Only ranks within an LSA (Local SHARP Aggregation) group can use the device-side API.Correct approachncclCommProperties.deviceApiSupport: The application should checkfalseafter initialization, and if it is

, fall back to the host-side API.

Pitfall 3: memset zeroing and leakage of uninitialized fields:ncclDevCommCopyNewToOld_v23000 📎 src/devcomm/devcomm_v23000.cc:118Scenariomemset(old, '\0', sizeof(*old))。

executesbefore copyingginSignalBase、ginCounterBaseWhy it is needed

: Old structs may contain fields that do not exist in the new version (such as: If developers manually implement version conversion and forget to zero it out, the kernel may read random values, manifesting as intermittent errors—difficult to reproduce and debug.

Correct approach: Always zero out the entire target struct before conversion. All of NCCL'sCopyNewToOldimplementations follow this pattern📎 src/devcomm/devcomm_v22902.cc:132 📎 src/devcomm/devcomm_v22907.cc:104 📎 src/devcomm/devcomm_v23000.cc:118。

Trap Four: Matching failures caused by gaps in version ranges

Scenario: The application is compiled with NCCL 2.29.4. Looking at the version range table:

FileminVersionmaxVersion
v229022.29.22.29.3
v229072.29.52.29.7

2.29.4 has no corresponding plugin.

[Design inference and architectural trade-offs]

What happens: If the matching logic strictly searches by range, 2.29.4 will fail to match and return an error. But in the actual implementation, there may be a "nearest match" strategy—2.29.4 may be routed to the v22902 or v22907 plugin.

Correct approach: Applications should try to use the same major version number as the runtime library. If cross-version use is necessary, test whether the target version range has a corresponding compatible plugin.

Failure recovery chain

When version conversion fails, NCCL's error recovery chain:

1. The filter returns an error:devCommRequirementsFilterreturnsncclInvalidUsage。

2. The upper-layer API catches the error:ncclCommGetDeviceHandlechecks the return value; if it is notncclSuccess, does not populate thedevCommstruct.

3. Application handling: The application should check the return value, and if it fails, fall back to the host-side API or terminate communication.

4. Logging: NCCL printsWARN-level logs, including the compile version and runtime version, to help locate the problem.

[Design inference and architectural trade-offs]

Currently NCCL does not provide an "automatic downgrade" mechanism—if version conversion fails, it will not automatically fall back to the host-side API. The application needs to implement the fallback logic itself.

---

Design considerations

Why use versioned structs instead of a "stable ABI"?

[Design inference and architectural trade-offs]

One alternative is to design ancclDevCommlayout that "never changes," with all new fields accessed through indirect pointers. But this brings two problems: first, indirect access increases latency (the kernel needs an extra dereference); second, it cannot use padding areas to optimize layout. NCCL chooses versioned structs as a trade-off between "performance" and "compatibility"—kernels within each version range get the optimal layout, and compatibility across versions is ensured through a conversion layer.

Why is v22907'sdevCommCopyOldToNewset to nullptr?

📎 src/devcomm/devcomm_v22902.cc:153-155The comments explain the reason: before 2.30.0,ncclDevCommhad no version field, so the old layouts of v22902 and v22907 cannot be distinguished. Since neither supports GIN backward compatibility, the differences in GIN fields do not affect correctness, so the conversion function of v22902 is reused.

Why doesnRanks_rcp32use fixed-point numbers instead of floating-point numbers?

[Design inference and architectural trade-offs]

The precision of GPU floating-point division may be insufficient to accurately represent1/nRanks, especially whennRanksis not a power of 2. Fixed-point numbers (decimals represented by 32-bit integers) can provide sufficient precision, and integer multiplication is faster than floating-point multiplication.

---

Chapter summary

This chapter dismantled the versioned ABI implementation under thesrc/devcommdirectory:

1. ncclDevCommmemory layout: Each version has precise field offsets, verified at compile time withstatic_assert. Key fields includerank、nRanks、nRanks_rcp32、lsaRank、lsaSize、windowTable、resourceWindow, etc.

2. Registration of the versioned ABI: Each version range corresponds to ancclDevCommCompatstruct, containingminVersion、maxVersion, a filter function, and a conversion function.

3. Field-level conversion:CopyNewToOldandCopyOldToNewcopy field by field and handle semantic changes (such asginConnectionStride > 1converted toginConnectionsRailed = true)。

4. Capability filtering:commPropertiesFilteradjusts the capability flags exposed to older versions,devCommRequirementsFilterchecks whether resource requests are compatible with older versions.

5. Production traps: conflicts between GIN resource requests and older-version kernels, device APIs being disabled during cross-node communication, the necessity of memset zeroing, and matching failures caused by gaps in version ranges.

In the next chapter we will move into device-side APIs and kernel fusion, looking at how thenccl_deviceheader file organizes device-side functions, and how kernel fusion combines multiple collective communication operations into a single kernel for execution.

Chapter review and self-test

Q1: IfncclDevCommCopyNewToOld_v23000inmemset(old, '\0', sizeof(*old))is removed, in what scenarios would the kernel read incorrect data? Please analyze based on the field differences between v22902 and v23000.

Reference analysis:

ncclDevComm_v22902The struct size of📎 src/devcomm/devcomm_v22902.cc:84is 200 bytesncclDevComm_v23000, while📎 src/devcomm/devcomm_v23000.cc:95-98is 240 bytesginSignalBase. v22902 hasginCounterBase(offset 176),ginContextBase(offset 184),

(offset 204), and other fields, which do not exist or have different semantics in v23000.memsetIfoldis removed, when converting from v23000 to v22902,ginSignalBase、ginCounterBasefields in the struct that do not exist in v23000 (such as

  • ) will retain garbage values on the stack. If the kernel happens to read these fields (for example, the GIN code path of an old kernel), it will get random values, causing:
  • The signal base address to be wrong, and GIN operations to write to the wrong memory location.
  • The counter base address to be wrong, causing counter overflow or underflow.

memsetIn extreme cases, this may trigger illegal memory access and cause the kernel to crash.CopyNewToOldZeroing ensures that all fields not explicitly assigned are 0, which is a safe default value. All of NCCL's📎 src/devcomm/devcomm_v22902.cc:132 📎 src/devcomm/devcomm_v22907.cc:104 📎 src/devcomm/devcomm_v23000.cc:118。

implementations include this stepncclDevCommCompatplugin. Please analyze how NCCL might handle this situation, and how applications should avoid it.

Reference Analysis:

Version range table:

  • v22902:2.29.2 - 2.29.3
  • v22907:2.29.5 - 2.29.7
  • v23000:2.30.0 - 2.30.7
  • v23100: 2.31.0 - current

2.29.4 falls into the gap between v22902 and v22907. Possible handling approaches:

1. Nearest match: NCCL might choose the largest range less than or equal to the requested version, i.e., v22902. But v22902'smaxVersionis 2.29.3, which strictly speaking does not cover 2.29.4.

2. Return error: If the matching logic strictly follows ranges, 2.29.4 will fail to match and returnncclInvalidUsage。

3. Match upward: Choose the smallest range greater than or equal to the requested version, i.e., v22907. But v22907'sminVersionis 2.29.5, which also does not cover 2.29.4.

[Design Inference and Architectural Trade-offs]

In actual implementation, NCCL may have a "fault tolerance" strategy—if no exact match is found, try using a plugin from an adjacent range. But this is not a reliable guarantee.

Application avoidance methods:

  • Use the same major version number as the runtime library (e.g., 2.31.x).
  • If cross-version is necessary, test whether the target version range has a corresponding compatible plugin.
  • After initialization, checkncclCommProperties.deviceApiSupport, if it isfalse, fall back to the host-side API.
Q3: ncclDevCommRequirementsFilter_v22902There is a piece of logic in:if (reqs->barrierCount) { reqs->lsaBarrierCount = std::max(reqs->lsaBarrierCount, reqs->barrierCount); reqs->barrierCount = 0; }. Please explain why this conversion is needed, and what would happen if it were not converted.

Reference Analysis:

📎 src/devcomm/devcomm_v22902.cc:117-121The comment in explains: "Prior to 2.29.4, a non-zero barrierCount did not imply GIN, but it does since."

Before 2.29.4,barrierCountonly indicated the number of LSA barriers and did not imply GIN requirements. Starting from 2.29.4,barrierCountimplies GIN requirements (i.e., requesting a barrier means requiring GIN resources).

When an application is compiled with 2.29.2, it may setbarrierCount > 0to indicate LSA barrier requirements, but is unaware that this implies GIN requirements. If the NCCL library (2.31.0) directly processes according to the new semantics, it will consider that the application has requested GIN resources, and thenncclDevCommRequirementsFilter_v22902will detect the GIN request and returnncclInvalidUsage—this is a false positive.

The conversion logic convertsbarrierCounttolsaBarrierCount(taking the maximum of the two), and clearsbarrierCount. This way:

  • lsaBarrierCountpreserves the application's barrier requirements.
  • barrierCount = 0avoids false GIN requirement reports.
  • railGinBarrierCount = 0Similarly, because in older versions it also did not imply GIN requirements.

If not converted, when an application is compiled with 2.29.2 and setsbarrierCount > 0, it will be incorrectly rejected and unable to use the device API.

At this point, we have seen clearly how devcomm safely maps key metadata of the host-side communication domain to the device side through versioned ABI, allowing kernels to obtain rank, addresses, and connection status without host pointers. This mechanism solves the basic problem of kernel access to the communication domain, but device-side capabilities go far beyond this. When users want to directly call communication primitives in their own kernels, or even fuse communication and computation into the same kernel, higher-level device-side APIs and kernel fusion techniques are needed. The next chapter will delve into the nccl_device directory and related examples, exploring how device-side APIs such as ncclBarrier, ncclLsaBarrier, ncclGinBarrier enable user kernels to participate in communication, and how kernel fusion reduces launch overhead, thereby pushing NCCL from a library toward a programming model.

CHAPTER 20

Chapter 20: Chapter 20: Device-Side Native APIs and Operator Fusion: nccl_device and kernel fusion practices

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 20 / 25

Chapter 20: Device-Side Native APIs and Operator Fusion: nccl_device and kernel fusion practices

In the previous chapter, we saw clearly how devcomm maps the host-side ncclComm metadata in a versioned way to the device side, allowing the kernel to read rank, addresses, and connection status. But "being able to read metadata" and "being able to initiate communication" are two different things. If there is only metadata, a user kernel can at most calculate addresses on its own and write flags on its own. Once cross-rank synchronization or cross-machine signal transmission is involved, it still has to go back to the host side to call collective APIs such as ncclAllReduce—and every such call means a kernel launch and a host-device round trip. The src/nccl_device directory analyzed in this chapter is precisely the key to NCCL's transition from "a library that is called" to "a model that can be programmed." What it provides is not a new collective communication algorithm, but a set of device-side primitives: allowing the user's own kernel to internally call synchronization operations such as ncclBarrier, ncclLsaBarrier, and ncclGinBarrier, thereby packing "communication" and "computation" into the same kernel and eliminating the intermediate launch overhead. The source material in this chapter focuses on the host-side requirement declaration (CreateRequirement) and Team abstraction for this set of primitives, which is exactly the entry point of the device-side API. A key premise for understanding this chapter: the design philosophy of the device-side API is "the host side declares resource requirements, and the device side consumes resources." The host side does not directly create barriers; instead, it tells NCCL "I need nBarriers barriers, and the team has team.nRanks members." Based on this, NCCL calculates how many buffers and how many GIN signals are needed, and then instantiates these resources on the device side. This separation of "declaration-consumption" is the fundamental reason why device-side code can work without host pointers.

1. Team abstraction: the coordinate system of the device-side API

Intuitive model

Imagine the organizational structure of a multinational company. To send an email, you first need to know "who to send it to"—whether it is to the whole company (World), to colleagues in the same office (LSA), or to a cross-office team on the same business line (Rail).ncclTeam_tThis is the descriptor for the "recipient scope." Without the Team abstraction, every device-side API would have to recalculate on its own "what my position is in this communication domain and how many members there are in total," which would lead to duplicated code and be highly error-prone.

Data structure and memory layout

ncclTeam_tIt is the coordinate system of the device-side API, and its three fields define anarithmetic progression:

FieldMeaningAnalogy
nRanksTotal number of members in the teamHow many people are in the group
rankThe number of the current rank within the teamMy sequence number in the group
strideThe stride in world space between adjacent members in the teamHow much the student IDs of two adjacent people in the group differ

strideIt is the most easily overlooked but most critical field. In the World team,stride = 1, because all ranks are arranged consecutively; but in the Rail team,stride = lsaSize, because ranks on the same rail appear in world space only once everylsaSizeentries.

📎 src/nccl_device/core.cc:13-19It shows the construction of the World team: directly takecomm->nRanksandcomm->rank,stridefixed to 1. This is the only team that does not requirencclDevrInitOnce, because all its information is in the host-sidecomm.

📎 src/nccl_device/core.cc:22-33It is the LSA team. Note L26'sncclDevrInitOnce(comm)—this is the idempotent entry point for device-side resource initialization. The comments at L23-25 are very important:errors are deliberately ignored here, because if initialization fails, the returned team is a "garbage value," but the next API call that actually needs resources will triggerncclDevrInitOnceagain and report the error. This is a "delayed error reporting" strategy, avoiding throwing heavy errors on lightweight operations such as team queries.

Scenario-driven walkthrough: coordinate transformation from World to Rail

Suppose an 8-GPU machine,lsaSize = 4(one LSA domain per 4 GPUs),nRanks = 8. Let us look at howncclTeamRailis constructed:

📎 src/nccl_device/core.cc:70-79InnRanks = 8 / 4 = 2,rank = comm->rank / 4,stride = 4. If the current rank is 5, then itsrank = 5 / 4 = 1,stride = 4in the Rail team means that the members of the Rail team are rank 1 and rank 5 in world.

Now look atncclTeamRankToWorld's conversion formula:

📎 src/nccl_device/core.cc:82-84'scomm->rank + (rank - team.rank) * team.strideis arelative offsetcalculation: first calculate the offset of the target rank relative to the current rank within the team(rank - team.rank), then multiply by the stridestride, and add the world number of the current rank. This formula is universal for all teams, becausestridealready encodes the team's arrangement pattern.

ncclTeamRankToLsais different:

📎 src/nccl_device/core.cc:87-92usescomm->devrState.lsaSelf + (rank - team.rank) * team.stride. Note that here it useslsaSelfrather thancomm->rank—because the LSA number is known only after device-side resource initialization and may differ from the world rank.

mermaid
flowchart TD
    start["用户调用 ncclTeamRail(comm)"] --> init{"ncclDevrInitOnce(comm)<br/>成功?"}
    init -->|"否"| empty["返回 ncclTeam_t{}<br/>空团队"]
    init -->|"是"| calc["计算 nRanks = comm->nRanks / lsaSize<br/>rank = comm->rank / lsaSize<br/>stride = lsaSize"]
    calc --> ret["返回 ncclTeam_t"]
    empty --> caller["调用方继续<br/>下一个 API 会报错"]
    ret --> caller

This figure reveals the execution path of the "delayed error reporting" strategy: when initialization fails, an empty team is returned, but the caller is not interrupted; the error will be exposed at the next API that actually needs resources (such asncclLsaBarrierCreateRequirement).

Design considerations and pitfalls

Why doesncclTeamWorldnot callncclDevrInitOnce?Because the information of the World team comes entirely from the host sidecomm, without requiring any device-side resources. Forcing a call would make a pure host query operation depend on device-side initialization, adding unnecessary failure points.

Pitfalls:ncclTeamRankToLsareturns on initialization failure-1(📎 src/nccl_device/core.cc:87-92), whilencclTeamRankToWorldnever fails. If a caller mixes these two functions and does not check return values, it may obtain-1when LSA initialization fails and use it as a valid rank, causing out-of-bounds access. In production code,ncclTeamRankToLsa's return value should be treated as an operation that may fail.

---

II. Barrier Requirement Declaration: How the Host Side "Reserves" Device Resources

Intuitive Model

Resource allocation on the device side is likebooking a meeting room: you cannot just rush into the meeting room for a meeting; you must first submit an application to the front desk (host-sideCreateRequirement)—"I want to hold 3 meetings, each with 8 attendees." Based on this, the front desk calculates how large a venue is needed (bufferSize), how many chairs are needed (ginSignalCount), and then gives you the venue number (outBufferHandle). Without this reservation mechanism, the device-side kernel would not know where its barrier buffer is or how large it is, and could not safely read from or write to it.

Data Structures and Memory Layout

TheCreateRequirementfunctions of the three barriers share the same pattern:zero the requirement struct → fill in buffer size/alignment → fill in the output handle pointer. But their resource types differ:

Barrier TypeResource TypeSize FormulaAlignment
LSA BarrierBuffer(3*n + n*team.nRanks) * sizeof(uint32_t)alignof(uint32_t)
CFT BarrierBuffer(3*n + n*team.nRanks) * NCCL_CFT_BARRIER_GRANNCCL_CFT_BARRIER_ALIGN
GIN BarrierGIN signaln * team.nRankssignalsDoes not involve a buffer

First, look at the size formula for the LSA Barrier:

📎 src/nccl_device/lsa_barrier.cc:14-22's(3 * nBarriers + nBarriers * team.nRanks) * sizeof(uint32_t)can be broken down into two parts:

  • 3 * nBarriers: each barrier needs 3uint32_tcontrol fields ([INFERENCE] usually "arrival count," "round," and "status flag").
  • nBarriers * team.nRanks: each barrier needs to reserve oneuint32_tarrival slot for each member in the team.

So the total size of a single barrier is3 + team.nRanksofuint32_t. This formula is exactly the same in LSA and CFT, except that CFT usesNCCL_CFT_BARRIER_GRANas the granularity unit (possibly to align to a larger boundary).

GIN Barrier is completely different:

📎 src/nccl_device/gin_barrier.cc:14-20does not allocate a buffer; instead it setsginSignalCount = nBarriers * team.nRanks, and pointsoutGinSignalStartto thesignal0in the handle. This is because the GIN barrier uses the network signal path and does not need a shared memory buffer; instead, it needs signal slots that the NIC can recognize.

Scenario-Driven Walkthrough: A Complete Reservation for One LSA Barrier

Suppose the user wants to create 2 barriers on a 4-GPU LSA team:

1. Call ncclLsaBarrierCreateRequirement(team, 2, &handle, &req)。

2. zero:memset(outReq, 0, sizeof(*outReq))(📎 src/nccl_device/lsa_barrier.cc:14-22)—ensures that unset fields have deterministic values, preventing the caller from reading garbage from the stack.

3. Record the number of barriers:outHandle->nBarriers = 2(📎 src/nccl_device/lsa_barrier.cc:14-22)。

4. Calculate the buffer size:(3*2 + 2*4) * 4 = (6 + 8) * 4 = 56bytes (📎 src/nccl_device/lsa_barrier.cc:14-22)。

5. Set alignment:alignof(uint32_t) = 4(📎 src/nccl_device/lsa_barrier.cc:14-22)。

6. Write back the handle pointer:outReq->outBufferHandle = &outHandle->bufHandle(📎 src/nccl_device/lsa_barrier.cc:14-22)—lets NCCL write the address back to the handle after actually allocating the buffer.

mermaid
flowchart LR
    subgraph host["host 侧声明阶段"]
        req["ncclLsaBarrierCreateRequirement<br/>team, nBarriers=2"]
        calc["bufferSize = (3*2 + 2*4)*4 = 56<br/>bufferAlign = 4"]
        handle["outHandle->nBarriers = 2<br/>outReq->outBufferHandle = &handle->bufHandle"]
    end
    subgraph dev["device 侧消费阶段"]
        buf["缓冲区 56 字节<br/>3 控制字段 + 4 到达槽位"]
        bar["ncclLsaBarrier 实例"]
    end
    req --> calc --> handle
    handle -.->|"NCCL 分配后回填"| buf
    buf --> bar

This data flow diagram shows the separation of "declaration" and "consumption": the host side only calculates the size and pointer; the actual buffer allocation and instantiation happen inside NCCL, and the device-side kernel receives an already-filled handle.

Design Considerations and Pitfalls

Why usememsetto zero the entireoutReq?BecausencclDevResourceRequirements_tis a multi-field struct, and different barrier types fill only part of its fields. Zeroing ensures that unused fields (such asginSignalCount, which LSA barrier does not use) are 0, and NCCL internally uses this to determine "this resource is not needed." If it is not zeroed, random values on the stack may be mistaken for "GIN resources are needed," triggering the false-positive problem mentioned in the previous chapter.

Pitfalls:outReq->outBufferHandle = &outHandle->bufHandlehands the address of a field inside the handle to NCCL. This means thatoutHandlemust remain valid until NCCL completes buffer allocation (it cannot be reclaimed by the stack or moved). If the user placesoutHandlein a scope that is released early, NCCL will write to a dangling pointer when writing back.

[Design Inference and Architectural Trade-offs]

Granularity Difference of CFT Barrier:📎 src/nccl_device/cft_barrier.cc:13-21usesNCCL_CFT_BARRIER_GRANandNCCL_CFT_BARRIER_ALIGNto replace LSA'ssizeof(uint32_t)andalignof(uint32_t). This indicates that the barrier of CFT (possibly Cross-Fabric Team or a similar cross-domain team) needs a larger alignment granularity, possibly because it must cross multicast memory regions, and the hardware has stricter requirements for address alignment.

---

III. Semantic Division of Labor Among the Three Barriers: What LSA, CFT, and GIN Each Handle

Intuitive Model

The three barriers are like "assembly whistles" for three different scopes:

  • LSA Barrier: colleagues in the same office assemble, using shared memory, the fastest.
  • CFT Barrier: assembly across offices but within the same building, using multicast memory, medium speed.
  • GIN Barrier: assembly across cities or even countries, using network signals, the slowest but with the widest coverage.

Choosing the wrong barrier type will not cause errors, but it will bring huge performance losses—using a GIN barrier for same-office synchronization is like sending a document to the desk next door via international express.

Comparison of Data Structures and Memory Layout

From the host-side requirement declaration, the resource requirements of the three are completely different:

DimensionLSA BarrierCFT BarrierGIN Barrier
RequirescommparameterNoNoYes
BufferYesYesNo
GIN signalNoNoYes
Size unituint32_tNCCL_CFT_BARRIER_GRANNumber of signals
Output handle fieldbufHandlebufHandlesignal0

Note that GIN Barrier is the only one that requires thecommparameter:

📎 src/nccl_device/gin_barrier.cc:14-20's function signature includesncclComm_t comm, while the signatures of LSA and CFT only havencclTeam_t team. This is because GIN signals need to be bound to specific network connections, and the network connection information is incomm.

Scenario-Driven Walkthrough: Signal Allocation for GIN Barrier

📎 src/nccl_device/gin_barrier.cc:14-20The logic of is simpler than LSA, but the semantics are more subtle:

1. Zero out:memset(outReq, 0, sizeof(*outReq))(L16)。

2. Set signal count:outReq->ginSignalCount = nBarriers * team.nRanks(L17) — each barrier needs to allocate a signal slot for each member in the team.

3. Backfill signal start pointer:outReq->outGinSignalStart = &outHandle->signal0(L18) — note that is not set here, because GIN barrier does not use a shared memory buffer.bufferSize[Design Inference and Architectural Trade-offs]

The name suggests that the handle may contain a set of contiguous signal fields (

signal0points to the first one, and NCCL uses this to know where to start allocatingsignal0, signal1, ...),outGinSignalStartsignals.nBarriers * team.nRanksConcurrency Control and Hardware Interaction

The concurrency control mechanisms of the three barriers are completely different:

: atomic operations based on shared memory.

  • LSA Barrierout of3 + team.nRanks, the arrival slot uses atomic add or atomic write to mark "I have arrived," and the control field uses atomic read to check "whether everyone has arrived." This is pure intra-GPU synchronization and does not involve the network.uint32_t: based on multicast memory (multimem). [INFERENCE] Multicast memory allows a single write operation to simultaneously update the view of multiple ranks, so the CFT barrier may use fewer control fields to achieve broader synchronization.
  • CFT Barrier: based on network signals.
  • GIN Barriersignals are sent through the NIC, and the receiver polls the signal slots. This is the only barrier that involves cross-machine hardware.ginSignalCountCopy
mermaid
sequenceDiagram
    participant K as "用户 Kernel"
    participant LSA as "LSA 共享内存"
    participant CFT as "CFT 多播内存"
    participant NIC as "网卡 GIN 信号"
    K->>LSA: "原子写到达槽位"
    LSA-->>K: "轮询所有槽位"
    Note over K,LSA: LSA barrier 完成
    K->>CFT: "多播写控制字段"
    CFT-->>K: "读多播状态"
    Note over K,CFT: CFT barrier 完成
    K->>NIC: "发送 GIN 信号"
    NIC-->>K: "轮询信号槽位"
    Note over K,NIC: GIN barrier 完成

Design Considerations and Pitfalls

Why do LSA and CFT not need the

parameter?commBecause their resources (shared memory, multicast memory) have already been bound to the team during thephase,ncclDevrInitOnceitself implicitly contains the resource location information. However, GIN signals need to dynamically allocate network resources and must access the network connection state throughteam.commPitfall

: GIN Barrier'sisginSignalCount. If the team is very large (e.g., 1024 ranks) and there are many barriers (e.g., 100), the total number of signals will reach 102400. The NIC's signal slots are a limited resource, and excessive requests may causenBarriers * team.nRanksto fail. Production code should request based on the actual minimum number of barriers needed, rather than requesting a large number of spares all at once.ncclDevrInitOnceIV. From Requirement Declaration to Device-Side Consumption: The Complete Lifecycle

---

Intuitive Model

is only "placing an order"; the actual "shipping" and "receiving" happen inside NCCL and in the device-side kernel. The entire lifecycle is like

CreateRequirementonline shopping: you place an order (CreateRequirement) -> the merchant prepares the goods (NCCL allocates resources) -> the courier delivers (resources are bound to DevComm) -> you sign for and use it (the device-side kernel calls barrier).Data Structures and Memory Layout: Field Evolution of the Handle

Taking

as an example, it goes through three stages in its lifecycle:ncclLsaBarrierHandle_tStage

Other FieldsnBarriersbufHandleAfter CreateRequirement
Already setThe address has been backfilled, but the content has not been allocatedNot setAfter NCCL allocation
Already setPoints to the actual bufferAlready setDevice-side use
Read-onlyRead-onlyRead-onlysets

📎 src/nccl_device/lsa_barrier.cc:14-22backfills the address ofnBarriers,📎 src/nccl_device/lsa_barrier.cc:14-22. Between these two operations, NCCL internally completes the actual allocation of the buffer.bufHandleScenario-Driven Walkthrough: A Complete Barrier Usage

Host-side declaration

1. : the user callsand obtainsncclLsaBarrierCreateRequirement(team, 2, &handle, &req)Host-side submissionreq.bufferSize = 56。

2. : the user handstoreq(content from the previous chapter), NCCL allocates a 56-byte buffer and writes the address intoncclDevCommCreateDevice-side initializationhandle.bufHandle。

3. : when the user kernel starts, it retrievesfrom DevComm and useshandleto locate the buffer.bufHandleDevice-side synchronization

4. : the kernel calls, writes an arrival marker into the corresponding slot of the buffer, and polls the other slots.ncclLsaBarrier(handle, barrierIndex)Device-side completion

5. : after all ranks have arrived, the barrier returns and the kernel continues execution.Copy

mermaid
flowchart TD
    a["ncclLsaBarrierCreateRequirement<br/>算出 bufferSize=56"] --> b["ncclDevCommCreate<br/>分配 56 字节缓冲区"]
    b --> c{"分配成功?"}
    c -->|"否"| err["返回 ncclSystemError<br/>句柄无效"]
    c -->|"是"| d["回填 handle.bufHandle<br/>指向实际缓冲区"]
    d --> e["用户 kernel 启动<br/>从 DevComm 取 handle"]
    e --> f["ncclLsaBarrier(handle, idx)<br/>写到达槽位 + 轮询"]
    f --> g{"所有 rank 到达?"}
    g -->|"否"| f
    g -->|"是"| h["barrier 返回<br/>kernel 继续"]
    err --> i["用户需检查返回值<br/>不可使用无效句柄"]

itself always returnsncclLsaBarrierCreateRequirement); the real failure occurs in the subsequent resource allocation stage.ncclSuccess(📎 src/nccl_device/lsa_barrier.cc:14-22Concurrency Control and Hardware Interaction

The core of concurrency control for the device-side barrier is

atomic operations + memory barriers. Taking the LSA barrier as an example:Arrival phase

  • : each rank uses an atomic write (or atomic add) to update its own arrival slot. This step must use release semantics to ensure that all memory operations before the barrier are visible to other ranks.Polling phase
  • : each rank uses an atomic read (or volatile read) to check all slots. This step must use acquire semantics to ensure that after seeing "everyone has arrived," it can read the data written by others before the barrier.Reset phase
  • : after the barrier completes, the slots need to be reset for the next use. The concurrency control in this step is the most subtle — if the reset is too fast, it may overwrite the markers of ranks that have not yet read them.[Design Inference and Architectural Trade-offs]
〔设计推断与架构权衡〕

3 * nBarriersThese control fields are very likely used to handle this kind of "round" problem: one field records the current round, one field records the arrival count, and one field serves as a reset flag. This way, multiple barriers can reuse the same set of slots without confusing rounds.

Production Pitfall Guide

Pitfall 1: Handle Lifetime Management。outReq->outBufferHandle = &outHandle->bufHandleThe address of the handle's internal field was handed to NCCL. If the user destroysncclDevCommCreatebeforeoutHandlereturns, NCCL will write to freed memory when backfilling. The correct approach is to bind the lifetime ofoutHandleto the DevComm, rather than to the function scope that created it.

Pitfall 2: The Product of Barrier Count and Team Size。bufferSize = (3*n + n*team.nRanks) * sizeof(uint32_t)Inn*team.nRanksthe term dominates the size for large teams. 1024 ranks and 100 barriers require100*1024*4 = 409600bytes, about 400KB. If every rank requests this much, the memory pressure cannot be ignored. You should request based on the number of barriers actually used concurrently, not the total number of barriers.

Pitfall 3: GIN Barrier Signal Exhaustion. GIN signals are NIC resources and are limited in number. If multiple DevComms request a large number of GIN signals at the same time, they may exhaust the NIC slots. Production code should check whether DevComm creation failure is due to insufficient GIN signals, and consider reducingnBarriersor switching to LSA barriers.

Pitfall 4: Delayed Exposure of Initialization Failure。ncclTeamLsaFunctions such asncclDevrInitOncereturn an empty team when📎 src/nccl_device/core.cc:22-33fails, without reporting an error. If user code does not check the return values of subsequent APIs, it may continue operating on an empty team, causing hard-to-locate errors. It is recommended to explicitly check the validity of the team when using the device-side API for the first time (such asteam.nRanks > 0)。

---

V. Kernel Fusion: Why Put Communication and Computation into One Kernel

Intuitive Model

In the traditional model, one "AllReduce + activation function" requires two kernels: one for communication and one for computation. There is an implicit global synchronization between the two kernels - the communication kernel must fully finish before the computation kernel can start. This is likea relay race: the first runner must hand the baton to the second runner after finishing, and at the moment of handoff both are waiting. Kernel fusion, by contrast, lets the same kernel perform both communication and computation, likea person changing shoes while running, eliminating the waiting at the handoff.

Data Structures and Memory Layout

The key to kernel fusion is that communication primitives (such as barriers) and computation logic share the same kernel's registers and shared memory. This means:

  • Register Pressure: The atomic operations and polling loops of communication primitives consume registers, squeezing the register budget of the computation logic.
  • Shared Memory Contention: If the LSA barrier buffer is placed in shared memory, it will compete with the shared memory needs of the computation logic.
  • Occupancy Impact: The occupancy of a fused kernel is usually lower than that of a pure computation kernel, because communication primitives require additional resources.
[Design Inference and Architectural Trade-offs]

The design of the device-side API (host side declares resources, device side consumes them) is precisely intended to alleviate these pressures: resources are pre-allocated on the host side, and the device-side kernel only needs to read and write, without dynamic allocation, reducing register usage.

Scenario-Driven Walkthrough: Execution Flow of a Fused Kernel

Suppose the user wants to write a fused "AllReduce + ReLU" kernel:

1. Host-side preparation: callncclLsaBarrierCreateRequirementto request a barrier, and callncclDevCommCreateto allocate resources.

2. Kernel launch: the user kernel receives the DevComm and barrier handle as parameters.

3. Communication phase: inside the kernel, callncclLsaBarrierto synchronize all ranks, then each rank exchanges data (directly reading and writing through symmetric memory).

4. Computation phase: after synchronization is complete, the kernel directly applies ReLU to the local data, without an additional kernel launch.

5. Completion: the kernel exits, and the host side does not need to wait for an additional communication kernel.

mermaid
flowchart LR
    subgraph old["传统模式:两个 kernel"]
        k1["通信 kernel<br/>AllReduce"] --> sync["隐式全局同步<br/>kernel 边界"]
        sync --> k2["计算 kernel<br/>ReLU"]
    end
    subgraph fused["融合模式:一个 kernel"]
        f1["通信阶段<br/>ncclLsaBarrier + 数据交换"]
        f1 --> f2["计算阶段<br/>ReLU"]
    end
    old -.->|"融合后省掉"| fused

This comparison diagram shows the core benefit of fusion: eliminating the implicit global synchronization at the kernel boundary. In the traditional model, the cost of this synchronization is the latency of two kernel launches plus the draining of the GPU pipeline.

Design Reflections and Pitfalls

Why does the device-side API not directly provide "fused AllReduce"?Because the specific form of fusion depends on the user's computation logic. What NCCL provides areprimitives(barriers, signals, symmetric memory access), notfinished products(fused AllReduce+ReLU). Users need to combine these primitives themselves to implement a fused kernel that meets their own needs. This is the essential difference between a "programming model" and a "library."

Pitfalls: Debugging a fused kernel is much harder than debugging separate kernels. If the barrier logic has a bug, it can cause the kernel to hang (deadlock), and a hung GPU kernel is not as easy to diagnose as a hung host process. It is recommended to add a timeout mechanism to the fused kernel, or first validate the barrier logic with a small-scale team.

Pitfalls: The reduced occupancy of a fused kernel may cause a loss in compute performance that exceeds the gains from saved communication. Before deciding to fuse, you should measure the end-to-end time before and after fusion, rather than only looking at the reduction in communication latency.

Chapter Review and Self-Test

Q1: If inncclTeamLsaL26, thencclDevrInitOncecall is removed and it directly returnscomm->devrState.lsaSizeandlsaSelf, in what scenarios would the device-side kernel read incorrect team information?

Reference Analysis:ncclDevrInitOnceis the idempotent entry point for device-side resource initialization. If it is removed,comm->devrState.lsaSizeandlsaSelfmay still be at their initial values (usually 0 or undefined). In the scenario where the device-side API is used for the first time, when the user callsncclTeamLsa, they will getnRanks = 0's empty team. Later, if the user does not check the validity of the team and directly uses this team to callncclLsaBarrierCreateRequirement, it will computebufferSize = (3*n + n*0) * 4 = 12nbytes—smaller than actually needed, because then*team.nRanksentry becomes 0. This leads to a buffer overflow: the barrier runtime tries to writeteam.nRanksarrival slots, but the buffer has only allocated space for3nofuint32_t. More subtly, iflsaSelfis also 0,ncclTeamRankToLsawill return an incorrect rank number, causing the barrier's arrival slot to be written to the wrong location, and it may never wait for all ranks to arrive, causing the kernel to hang. This is exactly the situation that the L23-25 comment's "return garbage value, next API reports error" strategy is meant to prevent—but only if the next API actually reports an error, rather than silently using the wrong size.

Q2:ncclLsaBarrierCreateRequirementThe size formula is(3*nBarriers + nBarriers*team.nRanks) * sizeof(uint32_t). If the team has 8 ranks and the user requests 1 barrier, the buffer is 44 bytes. Assume that in the barrier implementation the "3 control fields" are "arrival count", "round", and "reset flag". Reason through: when 8 ranks arrive simultaneously, if the "arrival count" uses a non-atomic++operation, what will happen?

Reference Analysis: A non-atomic++on the GPU is a three-step "read-modify-write", not an atomic operation. When 8 ranks executecount++simultaneously, multiple ranks may read the same old value (for example, all read 0), and then all write back 1. In the end,countincreases by only 1 instead of 8, causing the barrier to always think that "not everyone has arrived yet", and all ranks spin forever in the polling phase. This is why the arrival slot of an LSA barrier must use atomic operations (such asatomicAdd) or each rank writes to its own independent slot (thenBarriers * team.nRanksentry is precisely to reserve an independent slot for each rank). If the "each rank writes its own slot" scheme is adopted, atomic add is not needed; only atomic write + memory barrier is needed, because each slot has only one writer. This also explains why the size formula includes thenBarriers * team.nRanksentry—it trades space for atomicity and avoids multi-writer contention.

Q3:ncclGinBarrierCreateRequirementrequires thecommparameter whilencclLsaBarrierCreateRequirementdoes not. If thecommparameter were forcibly added to the LSA barrier as well (assuming for interface uniformity), what design problems would be introduced? Conversely, if thecommparameter were removed from the GIN barrier, in what scenarios would it fail?

Reference Analysis: The problem with adding thecommparameter to the LSA barrier is that it introduces an unnecessary dependency. The LSA barrier's resources (shared memory) have already been bound to the team during thencclDevrInitOncephase, andteamitself already implies the resource location. Addingcommwould make a pure team operation depend on communication domain state, increasing failure points (for example, ifcommis invalid, the LSA barrier also cannot be created), and it violates the principle of "least privilege". Conversely, removing thecommparameter from the GIN barrier would fail, because GIN signals need to be bound to a specific network connection.ncclGinBarrierCreateRequirement'sginSignalCountneeds to know which NIC and which QP (Queue Pair) to send the signal to, and this information is incomm's network transport layer state. Withoutcomm, NCCL cannot determine which NIC's slot the signal should be assigned to, nor can it guarantee that the signal can be correctly routed to the target rank. This reflects a design principle of device-side APIs:resource requirement declarations depend only on the context they truly need—LSA only needs the team topology, while GIN needs the network connection.

---

Device-side APIs and kernel fusion turn NCCL from "a library you call" into "a programming model you use".ncclTeam_tprovides the coordinate system,CreateRequirementprovides the resource reservation mechanism, and the three barriers cover the entire synchronization range from shared memory to network signals. But declaring resources and writing a fused kernel does not mean the performance will be good—the number of barriers, the size of the team, and the granularity of fusion, every choice affects end-to-end performance. In the next chapter, we will enter the practice of performance tuning, looking at how tuning parameters affect algorithm selection and how to use real benchmarks to verify the tuning results.

At this point, we have completed the entire journey from devcomm metadata mapping to nccl_device device-side primitives, and seen how NCCL, through the "host declares, device consumes" model, allows user kernels to directly call barrier-type synchronization operations, fusing communication and computation into the same kernel. But after mastering these mechanisms, a more practical question naturally arises: when a real training task fails to meet performance targets, how do we determine whether it is due to an inappropriate algorithm choice, a protocol mismatch, or an unreasonable channel count configuration? The next chapter will string together the mechanisms from the previous 20 chapters into an actionable tuning methodology, combining performance reports, cost models, and environment variables to provide a troubleshooting path from symptoms to root causes.

CHAPTER 21

Chapter 21: Chapter 21: Performance Tuning in Practice: tuning hands-on, benchmark tools, and tuning methodology

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 21 / 25

Chapter 21: Performance Tuning in Practice: tuning hands-on, benchmark tools, and tuning methodology

In the previous chapter, we saw how user-defined kernels can cooperate with NCCL communication primitives through device-side APIs, and even fuse communication and computation into the same kernel. This opens up the possibility of NCCL as a programming model, but it also brings a practical question: when communication performance is not as expected, where should we start? NCCL exposes hundreds of NCCL_PARAMs, but what really determines which path a collective communication takes are actually only three knobs: algorithm (Algo), protocol (Proto), and number of channels (nChannels). This chapter strings together the mechanisms from the previous 20 chapters into an actionable troubleshooting path—first look at performance reports to locate the symptoms, then read the cost model to understand how NCCL itself chooses, and finally use environment variables and benchmarks to verify your hypothesis.

21.1 Performance Reports: First Establish a "Normal" Baseline

The first step in tuning is not to change parameters, but to know what "normal" looks like. If you do not even know what the peak bandwidth of the current system is, any parameter tuning is blind guessing.

NCCL officially publishes reference performance data underdocs/perf, and its positioning is very clear—it is not a product-level guarantee, but a reference point for aligning expectations.

📎 docs/perf/README.md:3-14

code
NCCL publishes reference performance data to:

1. Provide reference points that help users align performance expectations.
2. Help users validate their system setup.
3. Reduce repeated requests to the NCCL team for basic performance numbers.

These results are references, and NOT product-level guarantees that the same
performance is achievable on every system. Performance depends on a complex
combination of software versions, system configuration, hardware, and operating
conditions, including factors outside NCCL's control. A difference within 5% is
generally considered acceptable variance due to differences in the underlying
systems.

There are two key pieces of information here that beginners easily overlook:

First,differences within 5% are normal fluctuation. This means that when you measure 3% lower than the official result, do not rush to tune parameters—first confirm whether it is measurement noise, GPU clock jitter, or interference from neighboring jobs.

Second,the official data only publishes peak bandwidth, not latency。

📎 docs/perf/README.md:24-24

code
We publish peak bandwidth for a selection of commonly used platforms. We do not
currently publish latency because it is typically more sensitive to factors
outside NCCL's control.
[Design inference and architectural trade-offs]

Why is latency not published? Because latency is extremely sensitive to system state—CPU frequency, PCIe link status, NIC firmware version, and even the BIOS power policy can affect it. Bandwidth tends to saturate under large messages and is relatively stable; latency under small messages is formed by the superposition of countless tiny stages, and jitter in any one of them will be amplified. So when tuning,large messages look at bandwidth, small messages look at latency, and these are two different troubleshooting paths.

📎 docs/perf/README.md:24-24

code
If your workload differs significantly from the published results, open an
issue in the [NCCL repository](https://github.com/NVIDIA/nccl/issues) or contact
NVIDIA Support. We will try our best to help.

The first item in the troubleshooting order: first run a standard benchmark (such asnccl-tests'sall_reduce_perf), and compare the result with the official report. If the gap is within 5%, it means the system configuration is fine, and the performance bottleneck is in your application layer (such as communication frequency or message splitting method); if the gap is significant, then proceed to NCCL parameter tuning.

21.2 Cost Model: How NCCL Itself Chooses Algorithms and Protocols

To tune parameters, you first need to understand how NCCL chooses by default. Internally it has a "cost model," which is essentially a lookup table plus formula calculation: given the message size, topology type, and number of ranks, it estimates the time cost of each "algorithm × protocol" combination and chooses the smallest one.

Intuitive model

Think of the cost model as navigation software. You input the start and end points (message size, topology), and internally it estimates the time for each route (algorithm/protocol combination), then recommends the fastest one. Navigation estimates are based on historical data and road classes, while NCCL's estimates are based on a hardcoded table of latency/bandwidth parameters.

Without this model, NCCL could only use the same fixed algorithm for all scenarios—small messages would become slower due to excessive startup overhead, and large messages would become slower due to insufficient bandwidth utilization, so the system would perform poorly at both extremes.

Data structures: model table and tuning context

The core of the cost model is themodelMaparray, and each element corresponds to an "algorithm/protocol/symmetric kernel" combination.

📎 src/tuning/cost_model.cc:230-277

code
static struct ncclTuningModelEntry_t modelMap[] = {
    /*
Initialize default, static models here
{mod_init, mod_sim, mod_final, enabled}
Enable order: Broadcast, Reduce, AllGather, ReduceScatter, AllReduce
*/
  {ncclTuningTreeModelInit, ncclTuningTreeModelSim, nullptr, {0, 0, 0, 0, 1}},       // Tree/LL
  {ncclTuningTreeModelInit, ncclTuningTreeModelSim, nullptr, {0, 0, 0, 0, 1}},       // Tree/LL128
  {ncclTuningTreeModelInit, ncclTuningTreeModelSim, nullptr, {0, 0, 0, 0, 1}},       // Tree/Simple
  {ncclTuningRingModelInit, ncclTuningRingModelSim, nullptr, {1, 1, 1, 1, 1}},       // Ring/LL
  ...

Each entry has four fields:mod_init(initialization function),mod_sim(simulation function),mod_final(cleanup function),enabled(enable flags for each of the 5 functions).enabledThe order of the{Broadcast, Reduce, AllGather, ReduceScatter, AllReduce}array is

—note this order, because it will be used repeatedly when reading the code later.

[Design inference and architectural trade-offs]Key observation:({0,0,0,0,1}Tree is only enabled for AllReduce{1,1,1,1,1}). This is because the advantage of the Tree algorithm lies in the fact that the reduction phase of AllReduce can be parallelized, but for operations like AllGather/ReduceScatter that are essentially ring-based pipelines, Ring is more natural.

The specific parameters of the model are stored inncclTunerConstants_t, including the base latency and bandwidth under each topology.

📎 src/tuning/cost_model.cc:142-152

code
static const ncclTunerConstants_t ncclTunerConstantsDefaults = {
    // baseLatencies
  {
    {6.8, 14.0, 8.4},  // Tree
    {6.6, 14.0, 8.4},  // Ring
    {0, 0, 0},         // Collnet Direct
    {0, 0, 0},         // Collnet Chain
    {0, 0, 0},         // NVLS
    {0, 0, 0},         // NVLS Tree
    {8.0, 8.0, 8.0}    // PAT
  },

Each algorithm has three base latency values, corresponding to the three protocols LL / LL128 / Simple. For example, Ring's{6.6, 14.0, 8.4}means: LL protocol base latency 6.6 microseconds, LL128 is 14.0, Simple is 8.4. These numbers are empirical values measured by NVIDIA on real hardware.

Hardware latency is given separately by topology type (NVLink / PCI / NET).

📎 src/tuning/cost_model.cc:153-184

code
    // hwLatencies
  {
    /* NVLINK */
    {
      {0.6, 1.25, 4.0}, // Tree (LL/LL128/Simple)
      {0.6, 1.9, 3.4},  // Ring (LL/LL128/Simple)
      ...
    },
    /* PCI */
    {
      {1.0, 1.9, 4.0}, // Tree (LL/LL128/Simple)
      {1.0, 2.5, 5.7}, // Ring (LL/LL128/Simple)
      ...
    },
    /* NET */
    {
      {5.0, 8.5, 14},   // Tree (LL/LL128/Simple)
      {2.7, 4.0, 14.0}, // Ring (LL/LL128/Simple)
      ...
    },
  },

A comparison reveals the topology differences: on NVLink, the per-hop latency of Ring/Simple is 3.4 microseconds, on PCI it is 5.7, and on NET it is 14.0. This is why cross-machine communication is slow—each hop costs an extra 10 microseconds.

Bandwidth parameters are given by GPU architecture generation.

📎 src/tuning/cost_model.cc:183-183

code
    // llMaxBws
  {
    {39.0, 39.0, 20.4}, /* Volta-N1/Intel-N2/Intel-N4) */
    {87.7, 22.5 /*avg of ring & tree*/, 19.0}, /* Ampere-N1/AMD-N2/AMD-N4) */
    {141.0, 45.0 /*avg of ring & tree*/, 35.0}, /* Hopper-N1/AMD-N2/AMD-N4) */
    {2 * 141.2, 2 * 45.0 /*avg of ring & tree*/, 2 * 35.0}, /* Blackwell-N1/AMD-N2/AMD-N4) */
  },

Each row corresponds to one architecture generation, and the three values are the maximum bandwidth of the LL protocol under single-machine (N1), dual-machine (N2), and quad-machine (N4) scenarios, respectively. Hopper single-machine is 141 GB/s, and Blackwell doubles it to 282 GB/s—this explains why the same algorithm performs much better on new cards.

Tuning context: per-comm state

Each communicator holds a copy ofncclTuningContext_t, which stores the tuning state of this comm.

📎 src/include/tuning.h:81-95

code
struct ncclTuningContext_t {
  // Persistant tuning parameters tied to a communicator.
  ncclTunerConstants_t tuningConstants;
  // State of the tuning models
  // Forced function is set via env var
  int forced[NCCL_NUM_FUNCTIONS];
  // Disabled tuning models are not execute and excluded from implemetation selection.
  int enabled[NCCL_TUNING_COUNT][NCCL_NUM_FUNCTIONS];
  // Store of model contexts per communicator.
  float generalLatencies[NCCL_NUM_FUNCTIONS][NCCL_NUM_ALGORITHMS][NCCL_NUM_PROTOCOLS];
  float generalBandwidths[NCCL_NUM_FUNCTIONS][NCCL_NUM_ALGORITHMS][NCCL_NUM_PROTOCOLS];

  ssize_t threadThresholds[NCCL_NUM_ALGORITHMS][NCCL_NUM_PROTOCOLS];
  int maxThreads[NCCL_NUM_ALGORITHMS][NCCL_NUM_PROTOCOLS];
};

Four key fields:

  • forced[NCCL_NUM_FUNCTIONS]: marks which functions have their algorithm/protocol forcibly specified by environment variables. This is whereNCCL_ALGO/NCCL_PROTOtakes effect.
  • enabled[NCCL_TUNING_COUNT][NCCL_NUM_FUNCTIONS]: a two-dimensional boolean table, marking whether a certain model is enabled for a certain function. Disabled models do not participate in selection.
  • generalLatencies / generalBandwidths: a three-dimensional array, storing the estimated latency and bandwidth by "function × algorithm × protocol". This is the source of the large table printed byncclTuningInit.
  • threadThresholds / maxThreads: thresholds related to the number of threads, determining how many threads each block uses.

Scenario-driven Walkthrough: Algorithm selection for one AllReduce

Suppose you callncclAllReduce, with a message size of 1MB, 8 cards single-machine NVLink. Internally, NCCL will construct ancclTuningInput_t, and then callncclTuningCompute。

📎 src/tuning/tuning.cc:180-202

code
ncclResult_t ncclTuningCompute(struct ncclTuningInput_t* const input, struct ncclTuningResult_t* const result) {
  ncclResult_t ret = ncclSuccess;
  TRACE(NCCL_TUNING, ...);
  struct ncclTuningResultList_t tunings;
  tunings.head = nullptr;
  struct ncclTuningResult_t bestTuning = NCCL_TUNING_RESULT_INIT;
  // Set tuning to Ring/Simple for single rank case
  if (input->comm->nRanks <= 1) {
    bestTuning.algo = NCCL_ALGO_RING;
    bestTuning.proto = NCCL_PROTO_SIMPLE;
    ...
  } else {
    NCCLCHECKGOTO(ncclTuningComputeAllTunings(input, &tunings), ret, exit);

Step 1: A single rank directly returns Ring/Simple without any computation. This is a short-circuit optimization—with a single card there is no communication, so whichever algorithm is chosen makes no difference.

Step 2: When there are multiple ranks, callncclTuningComputeAllTunings, iterating over all candidate combinations.

📎 src/tuning/tuning.cc:128-149

code
ncclResult_t ncclTuningComputeAllTunings(struct ncclTuningInput_t* const input,
                                         struct ncclTuningResultList_t* const tunings) {
  ncclResult_t ret = ncclSuccess;

  for (int i = 0; i < NCCL_TUNING_COUNT; i++) {
    struct ncclTuningResult_t tuning = NCCL_TUNING_RESULT_INIT;
    tuning.id = i;
    tuning.valid = 1;

    if (!(input->tuningMask & (1ULL << i))) {
      tuning.valid = 0;
      continue;
    }
    NCCLCHECK(ncclTuningExpandId(i, &tuning.algo, &tuning.proto, &tuning.symKernelId, &tuning.ceMethodId));
    NCCLCHECKGOTO(ncclTuningComputeTuning(i, input, &tuning), ret, fail);
    if (tuning.valid) NCCLCHECKGOTO(ncclTuningResultListPushFront(tunings, tuning), ret, fail);
  }

There is an ingenious design here:tuningMaskis a 64-bit mask, with each bit corresponding to one candidate combination.NCCL_TUNING_MASK_GENERAL_KERNELS、NCCL_TUNING_MASK_SYM_KERNELS、NCCL_TUNING_MASK_CErespectively delimit different categories of candidates.

📎 src/include/tuning.h:17-25

code
#define NCCL_TUNING_SYM_KERNEL_ID_OFFSET (NCCL_NUM_ALGORITHMS * NCCL_NUM_PROTOCOLS)
#define NCCL_TUNING_CE_METHOD_ID_OFFSET (NCCL_TUNING_SYM_KERNEL_ID_OFFSET + ncclSymkKernelId_Count)
#define NCCL_TUNING_COUNT (NCCL_TUNING_CE_METHOD_ID_OFFSET + ncclCeMethodId_Count)

#define NCCL_TUNING_MASK_GENERAL_KERNELS ((1ULL << NCCL_TUNING_SYM_KERNEL_ID_OFFSET) - 1ULL)
#define NCCL_TUNING_MASK_SYM_KERNELS \
  ((1ULL << NCCL_TUNING_CE_METHOD_ID_OFFSET) - 1ULL - NCCL_TUNING_MASK_GENERAL_KERNELS)
#define NCCL_TUNING_MASK_CE ((1ULL << NCCL_TUNING_COUNT) - (1ULL << NCCL_TUNING_CE_METHOD_ID_OFFSET))
#define NCCL_TUNING_MASK_ALL ((1ULL << NCCL_TUNING_COUNT) - 1ULL)

The layout of the mask is: the lowerNCCL_NUM_ALGORITHMS × NCCL_NUM_PROTOCOLSbits are the traditional "algorithm × protocol" combinations, the middlencclSymkKernelId_Countbits are symmetric kernels, and the high bits are CE (Copy Engine) methods. Using a bitmask instead of an array is for quickly determining inncclTuningComputewhether "this candidate is within the scope of this tuning".

Step 3: For each candidate, callncclTuningComputeTuning, which forwards toncclTuningCostModelSimModel。

📎 src/tuning/cost_model.cc:470-497

code
ncclResult_t ncclTuningCostModelSimModel(int id, struct ncclTuningInput_t* const input,
                                         struct ncclTuningResult_t* const result) {
  struct ncclTuningModelEntry_t* model = nullptr;
  ncclResult_t ret = ncclSuccess;
  result->forced = input->comm->tuningContext.forced[input->func];
  NCCLCHECKGOTO(getModelEntry(id, &model), ret, not_valid);
  if (model == nullptr) {
    ret = ncclInternalError;
    goto not_valid;
  }
  if (input->comm->tuningContext.enabled[id][input->func] == 0) {
    goto not_valid;
  }
  if (model->model != nullptr) {
    NCCLCHECKGOTO(model->model(input, result), ret, not_valid);
    if (result->timeUs <= 0.0) {
      goto not_valid;
    }
  } else {
    goto not_valid;
  }
exit:
  return ret;
not_valid:
  result->timeUs = NCCL_TUNING_IGNORE;
  result->valid = 0;
  goto exit;
}

Note the handling of thenot_validlabel: if any step fails (model does not exist, is disabled, or simulation returns a non-positive time), it will settimeUstoNCCL_TUNING_IGNORE、validand set it to 0. This candidate is then excluded from subsequent selection.

Step 4: Select the one with the minimum cost from all valid candidates.

📎 src/tuning/tuning.cc:155-173

code
static ncclResult_t ncclTuningSelectBestTuning(struct ncclTuningResultList_t* tunings,
                                               struct ncclTuningResult_t* const bestTuning) {
  bestTuning->timeUs = FLT_MAX;
  float bestSelectionTimeUs = FLT_MAX;
  struct ncclTuningResultListNode* node = tunings->head;
  while (node != nullptr) {
    const struct ncclTuningResult_t& tuning = node->result;
    float selectionTimeUs = tuning.selectionTimeUs > 0.0f ? tuning.selectionTimeUs : tuning.timeUs;
    TRACE(NCCL_TUNING, "A/P/S %s/%s/%s, time: %f, selection time: %f", ...);
    if (selectionTimeUs < bestSelectionTimeUs) {
      *bestTuning = tuning;
      bestSelectionTimeUs = selectionTimeUs;
    }
    node = node->next;
  }
  return ncclSuccess;
}

There is a detail here: the selection usesselectionTimeUs, and if it is greater than 0, it is used; otherwise it falls back totimeUs。selectionTimeUsis the "selection time", which may include additional penalty terms (for example, some algorithms incur extra overhead in specific scenarios). This gives the cost model the ability to separate "estimated time" and "selection time".

Flowchart

mermaid
flowchart TD
    start["ncclTuningCompute(input, result)"] --> check_ranks{"comm->nRanks <= 1?"}
    check_ranks -->|是| single["bestTuning = Ring/Simple<br/>nChannels = 0"]
    check_ranks -->|否| all["ncclTuningComputeAllTunings()"]
    all --> loop{"遍历 i in NCCL_TUNING_COUNT"}
    loop -->|mask 未命中| skip["tuning.valid = 0<br/>continue"]
    loop -->|mask 命中| expand["ncclTuningExpandId(i, ...)"]
    expand --> sim["ncclTuningComputeTuning()<br/>→ ncclTuningCostModelSimModel()"]
    sim --> sim_check{"enabled[id][func] != 0<br/>且 model->model != nullptr?"}
    sim_check -->|否| invalid["timeUs = NCCL_TUNING_IGNORE<br/>valid = 0"]
    sim_check -->|是| push["ncclTuningResultListPushFront()"]
    skip --> loop
    invalid --> loop
    push --> loop
    loop -->|遍历结束| tuner_check{"comm->tuner != NULL?"}
    tuner_check -->|是| plugin["tuner->getCollInfo()<br/>覆盖 generalTable"]
    tuner_check -->|否| select["ncclTuningSelectBestTuning()"]
    plugin --> select
    select --> channels["ncclTuningGetChannels()"]
    channels --> eff{"CTAPolicy & EFFICIENCY<br/>且 NCCL_ALGO/NCCL_PROTO 未设置?"}
    eff -->|是| nvls["尝试 NVLS 覆盖<br/>ncclNvlsRegResourcesQuery()"]
    eff -->|否| done["*result = bestTuning"]
    nvls --> done
    single --> done

This diagram fully depicts the decision path from the entry point to the final result, including all branches such as single-rank short-circuiting, mask filtering, model disabling, tuner plugin intervention, and CTAPolicy override.

21.3 Environment variables: the three knobs that truly affect performance

Once you understand the cost model, you know how environment variables intervene.NCCL_ALGO、NCCL_PROTO、NCCL_SYM_KERNELThese three variables, after being parsed byparseList, directly modify theenabledtable, disabling all candidates that do not match the user's intent.

Parsing syntax

parseListThe syntax supported by

📎 src/tuning/cost_model.cc:14-32

code
// Parse a map of prefixes to a list of elements. The first prefix is
// optional and, if not present, the list of elements will be applied
// to all prefixes. Only the first list of elements can lack a
// prefix. Prefixes (if present) are followed by a colon. Lists of
// elements are comma delimited. Mappings of prefix to the lists of
// elements are semi-colon delimited.
//
// For example:
//
//     NCCL_ALGO="ring,collnetdirect;allreduce:tree,collnetdirect;broadcast:ring"
// Enable ring and collnetdirect for all functions, then select tree
// and collnetdirect for allreduce and ring for broadcast.
//
//     NCCL_PROTO="LL,Simple;allreduce:^LL"
// Enable LL and Simple for all functions, but everything except LL
// for allreduce.
//
//     NCCL_PROTO="^LL128;allreduce:LL128"
// Enable everything but LL128, but only LL128 for allreduce.

Copy

1. Three usages::NCCL_ALGO="ring,tree"Global list

2. — all functions use only ring and tree.:NCCL_ALGO="ring;allreduce:tree"By function prefix

3. — default is ring, but allreduce uses tree.:NCCL_PROTO="^LL128"Exclusion syntax

^— everything except LL128 is enabled.

📎 src/tuning/cost_model.cc:59-67

code
    int unset, set;
    if (elemList[0] == '^') {
      unset = 1;
      set = 0;
      elemList++;
    } else {
      unset = 0;
      set = 1;
    }

prefix is the key—it means "unset", i.e., excluding a certain option from the default of all-enabled.^Copyunset=1、set=0When parsing tounset,set。

📎 src/tuning/cost_model.cc:69-96

code
    bool foundPrefix = false;
    for (int p = 0; p < nprefixes; p++) {
      if (prefix && strcasecmp(prefix, prefixElems[p]) != 0) continue;
      foundPrefix = true;
      for (int e = 0; e < nelems; e++) list[p * nelems + e] = unset;

      tokStr = strdup(elemList);
      char* tmpStr;
      char* elem = strtok_r(tokStr, ",", &tmpStr);
      while (elem) {
        int e;
        for (e = 0; e < nelems; e++) {
          if (strcasecmp(elem, elems[e]) == 0) {
            list[p * nelems + e] = set;
            forced[p] = 1;
            break;
          }
        }
        if (e == nelems) {
          WARN("Unrecognized element token \"%s\" when parsing \"%s\"", elem, str);
          ret = ncclInvalidUsage;
          goto fail;
        }
        elem = strtok_r(NULL, ",", &tmpStr);
      }

(all excluded), and then set the listed elements toforced[p] = 1Copy

Note the

ncclTuningCostModelInitline—as long as the user explicitly lists a certain element, the corresponding function is marked as "forced". This mark will later be used to determine whether the cost model is allowed to choose freely.

📎 src/tuning/cost_model.cc:363-384

code
    for (int f = 0; f < NCCL_NUM_FUNCTIONS; f++) {
      // Disable LL128 when 1) it is not supported on the platform, and 2) user did not explicitly request it.
      // protoEnable[..] == 2 indicates that user did not set NCCL_PROTO=LL128 explicitly.
      if (proto == NCCL_PROTO_LL128 && protoEnable[f * NCCL_NUM_PROTOCOLS + proto] == 2 &&
          !isLL128Enabled(comm->minCompCap, comm->maxCompCap, comm->graphs[algo].typeInter,
                          comm->graphs[algo].typeIntra, comm->nRanks, f, algo, comm->minDriverVersion)) {
        comm->tuningContext.enabled[i][f] = 0;
      }
      //  Check the user env vars only for functions that have a forced configuration and not already disabled.
      if (comm->tuningContext.forced[f] == 0 || comm->tuningContext.enabled[i][f] == 0) continue;
      comm->tuningContext.enabled[i][f] = 0;
      TRACE(NCCL_TUNING, "a/p/s %s/%s/%s enabled %d/%d/%d", ...);
      if (((algo != NCCL_ALGO_UNDEF && algoEnable[f * NCCL_NUM_ALGORITHMS + algo] != 0) &&
           (proto != NCCL_PROTO_UNDEF && protoEnable[f * NCCL_NUM_PROTOCOLS + proto] != 0)) ||
          (symKernelId != ncclSymkKernelId_Count && symKernelIdEnable[f * ncclSymkKernelId_Count + symKernelId] != 0)) {
        comm->tuningContext.enabled[i][f] = 1;
      }
    }

There is a key piece of logic in

1. that handles the interaction between user forcing and environment variables and platform capabilities.CopyisLL128EnabledThe order of this logic is important:protoEnable == 2First handle LL128 platform capability

2. : if the platform does not support LL128 (returns 0) and the user did not explicitly request it (forced[f] != 0), disable it directly.enabled[i][f] = 0), then check whether the user allows this combination—if allowed, re-enable it.

protoEnablehas three values: 0 (user excluded), 1 (user enabled), 2 (user did not mention, enabled by default). This tri-state design allows "explicit user request" and "platform default" to be distinguished.

Caching Mechanism for Environment Variable Reads

AllNCCL_PARAMmacros ultimately go throughncclLoadParam。

📎 src/misc/param.cc:78-108

code
int64_t ncclLoadParam(char const* env, int64_t deftVal, int64_t uninitialized, int64_t* cache, int8_t* noCache) {
  static std::mutex mutex;
  std::lock_guard<std::mutex> lock(mutex);

  // noCache is only load/stored within the mutex, no need for atomic
  if (*noCache == /*uninitialized*/ -1) ncclGetCachePolicy(env, noCache);

  if (COMPILER_ATOMIC_LOAD(cache, std::memory_order_relaxed) != uninitialized) {
    return COMPILER_ATOMIC_LOAD(cache, std::memory_order_relaxed);
  }

  // Read the environment variable
  const char* str = ncclGetEnv(env);
  int64_t value = deftVal;

  if (str && strlen(str) > 0) {
    errno = 0;
    char* end = nullptr;
    value = strtoll(str, &end, 0);
    // Preserve numeric-prefix parsing while rejecting non-numeric values.
    if (errno || end == str) {
      value = deftVal;
      ATTN("Invalid value %s for %s, using default %lld.", str, env, (long long)deftVal);
    } else {
      INFO(NCCL_ENV, "%s set by environment to %lld.", env, (long long)value);
    }
  }

  if (*noCache == /*cache*/ 0) COMPILER_ATOMIC_STORE(cache, value, std::memory_order_relaxed);
  return value;
}

This code has several noteworthy design points:

Global Mutex:static std::mutex mutexprotects the entire read process. This means the first read of all parameters is serialized. Why use a lock instead of lock-free? Because parameter reads only happen during the initialization phase, not on the hot path, so the lock overhead is negligible, while correctness is more important.

Double-Check: first atomically readcache, and if already initialized, return directly. This avoids entering the lock every time a parameter is read—although the lock itself has almost no contention after initialization, atomic reads are faster.

Caching Strategy:noCacheflag determines whether to write the read value back tocache. Some parameters (such as those requiring dynamic response) may disable caching and re-read the environment variable every time.

Error Handling:strtollWhen parsing fails, use the default value and printATTNwarning. Note theend == strcheck—if the string does not start with a digit,endwill equalstr, indicating that no number was parsed at all.

Configuration File Support

Environment variables do not necessarily have to be set from the shell; NCCL supports reading from configuration files.

📎 src/misc/param.cc:52-67

code
static void initEnvFunc() {
  char confFilePath[1024];
  const char* userFile = std::getenv("NCCL_CONF_FILE");
  if (userFile && strlen(userFile) > 0) {
    snprintf(confFilePath, sizeof(confFilePath), "%s", userFile);
    setEnvFile(confFilePath);
  } else {
    const char* userDir = userHomeDir();
    if (userDir) {
      snprintf(confFilePath, sizeof(confFilePath), "%s/.nccl.conf", userDir);
      setEnvFile(confFilePath);
    }
  }
  snprintf(confFilePath, sizeof(confFilePath), "/etc/nccl.conf");
  setEnvFile(confFilePath);
}

Load order:NCCL_CONF_FILEthe specified file (if set) →~/.nccl.conf → /etc/nccl.conf. Later loads override earlier ones (becausesetEnvFilecallsncclOsSetEnv)。

📎 src/misc/param.cc:69-72

code
void initEnv() {
  static std::once_flag once;
  std::call_once(once, initEnvFunc);
}

std::call_onceensures the configuration file is loaded only once, even if multiple threads call it for the first time simultaneouslyncclGetEnv。

21.4 Number of Channels: The Underestimated Performance Knob

Algorithms and protocols determine "how to go"; the number of channels determines "how many paths to open." Many people focus only on the first two when tuning and ignore the number of channels—but in large-message scenarios, the number of channels is often the key to bandwidth utilization.

Where the Number of Channels Comes From

ncclTuningComputeAfter selecting the best algorithm/protocol, it callsncclTuningGetChannelsto calculate the number of channels.

📎 src/tuning/tuning.cc:233-235

code
  if (bestTuning.algo != NCCL_ALGO_UNDEF && bestTuning.proto != NCCL_PROTO_UNDEF) {
    NCCLCHECKGOTO(ncclTuningGetChannels(input, &bestTuning), ret, exit);
  }

The calculation logic for the number of channels is not in this chapter's source material, but its role can be seen from the fields ofncclTuningResult_t.

📎 src/include/tuning.h:42-55

code
struct ncclTuningResult_t {
  int id;
  int valid;
  float timeUs;
  float selectionTimeUs;
  int algo;
  int proto;
  int symKernelId;
  int ceMethodId;
  int nChannels;
  int maxChannels;
  int nWarps;
  int forced;
};

nChannelsis the final number of channels used,maxChannelsis the upper limit.nWarpsis the number of warps per block.

CTAPolicy Override of the Number of Channels

There is a special piece of logic handling theNCCL_CTA_POLICY_EFFICIENCYpolicy.

📎 src/tuning/tuning.cc:236-257

code
  // NCCL_CTA_POLICY_EFFICIENCY requires user (non-symmetric) buffer registration (currently unsupported with MNNVL).
  // Run after GetChannels so bestTuning.nChannels is valid. Skip when a tuner plugin owns selection
  // (same as pre-rearch). The NVLS-bit guard keeps this bias inside the candidate set: a per-call
  // algSelection may have narrowed tuningMask, so EFFICIENCY must not resurrect NVLS when excluded.
  if (input->comm->tuner == NULL && (input->CTAPolicy & NCCL_CTA_POLICY_EFFICIENCY) &&
      ncclGetEnv("NCCL_ALGO") == NULL && ncclGetEnv("NCCL_PROTO") == NULL && !input->comm->MNNVL &&
      (input->tuningMask & (1ull << (NCCL_ALGO_NVLS * NCCL_NUM_PROTOCOLS + NCCL_PROTO_SIMPLE)))) {
    if (input->regBuff && (input->func == ncclFuncAllGather || input->func == ncclFuncReduceScatter)) {
      if ((input->comm->nNodes > 1 && input->collNetSupport && input->nvlsSupport) ||
          (input->comm->nNodes == 1 && input->nvlsSupport)) {
        int recChannels;
        NCCLCHECKGOTO(ncclNvlsRegResourcesQuery(input->comm, input->func, &recChannels), ret, exit);
        if (recChannels <= bestTuning.nChannels) {
          bestTuning.algo = NCCL_ALGO_NVLS;
          bestTuning.proto = NCCL_PROTO_SIMPLE;
          bestTuning.nChannels = recChannels;
          bestTuning.maxChannels = recChannels;
          bestTuning.nWarps = input->comm->tuningContext.maxThreads[bestTuning.algo][bestTuning.proto] / WARP_SIZE;
        }
      }
    }
  }

The guard conditions in this code are very dense and worth interpreting one by one:

1. input->comm->tuner == NULL: only take this path when there is no tuner plugin. When the plugin has the right to choose, NCCL does not intervene.

2. input->CTAPolicy & NCCL_CTA_POLICY_EFFICIENCY: the user has set the efficiency-first policy.

3. ncclGetEnv("NCCL_ALGO") == NULL && ncclGetEnv("NCCL_PROTO") == NULL: the user has not forced an algorithm/protocol. If forced, respect the user's choice.

4. !input->comm->MNNVL: not supported in the MNNVL scenario.

5. input->tuningMask & (1ull << (NCCL_ALGO_NVLS * NCCL_NUM_PROTOCOLS + NCCL_PROTO_SIMPLE)): NVLS/Simple is in the candidate set. This guard prevents "reviving" excluded options.

After the conditions are met, query the number of channels supported by NVLS registered resources, and if it does not exceed the current selection, switch to the NVLS algorithm.

[Design Inference and Architectural Trade-offs]

Why does the EFFICIENCY policy favor NVLS? Because NVLS (NVLink SHARP) uses switch hardware for reduction, which can reduce GPU computation and communication overhead and is more efficient for operations such as AllGather/ReduceScatter. However, its number of channels is limited by hardware resources, soncclNvlsRegResourcesQueryis needed to query the actual available amount.

Fallback Logic for Symmetric Kernels

Symmetric kernels are a newer feature, and when they are unavailable, they need to fall back to general kernels.

📎 src/tuning/tuning.cc:258-298

code
  if ((bestTuning.symKernelId != ncclSymkKernelId_Count ||
       (input->tuningMask & NCCL_TUNING_MASK_SYM_KERNELS && bestTuning.symKernelId == ncclSymkKernelId_Count)) &&
      bestTuning.algo == NCCL_ALGO_UNDEF && bestTuning.proto == NCCL_PROTO_UNDEF) {
    bool isLLKernel = (1 << bestTuning.symKernelId) & ncclSymkLLKernelMask();
    bool isOneThreadMultiGpus = input->comm->intraRanks > 1 && !ncclParamSingleProcMemRegEnable();
    bool needFallback = bestTuning.symKernelId != ncclSymkKernelId_Count ? false : true;

    // General kernel tuning structs if fallback is needed
    struct ncclTuningResult_t generalTuning = NCCL_TUNING_RESULT_INIT;
    struct ncclTuningInput_t generalInput = *input;
    generalInput.tuningMask = NCCL_TUNING_MASK_GENERAL_KERNELS;

    // Fallback logic for symmetric LL kernels:
    // - If both src and dst are registered, we don't fall back if a symmetric kernel is available.
    // - Otherwise, we have to fall back to generl kernel if running the selected symmetric LL kernel is
    //   not possible (if the buffers are not registered and we manage multiple GPUs).
    // - If the user forced a symmetric kernel via NCCL_SYM_KERNEL or requested preference for using
    //   symmetric kernels even without symmetric buffers via NCCL_SYM_NOWIN_ENABLE, we respect that.
    // - Otherwise, we query the general cost model and if it selects a non-LL proto, we pick that.
    if (bestTuning.symKernelId != ncclSymkKernelId_Count) {
      if (input->winRegType == ncclSymSendRegRecvReg) {
        needFallback = false;
      } else if (isLLKernel) {
        needFallback = isOneThreadMultiGpus && input->winRegType == ncclSymSendNonregRecvNonreg;
        if (!needFallback && !result->forced) {
          needFallback = !ncclParamSymNoWinEnable() && input->winRegType == ncclSymSendNonregRecvNonreg;
          if (!needFallback) {
            NOWARN(ncclTuningCompute(&generalInput, &generalTuning), NCCL_TUNING);
            needFallback = (generalTuning.proto != NCCL_PROTO_LL);
          }
        }
      }
    }

Fallback decision tree:

  • If both send and receive buffers are registered (ncclSymSendRegRecvReg), do not fall back.
  • If it is an LL kernel and a single thread manages multiple GPUs and the buffers are not registered, fall back.
  • If the user has not setNCCL_SYM_NOWIN_ENABLEand the buffers are not registered, fall back.
  • Otherwise, query the general cost model, and if it selects a non-LL protocol, fall back.
[Design Inference and Architectural Trade-offs]

The core of this logic is: symmetric LL kernels need buffer registration to leverage their advantages. When unregistered, the advantage of LL kernels (low latency) may be offset by additional address translation overhead, so falling back to general kernels is more worthwhile.

Error Handling When No Combination Is Available

If all candidates are excluded, NCCL reports an error and provides diagnostic information.

📎 src/tuning/tuning.cc:308-329

code
  if ((bestTuning.algo == NCCL_ALGO_UNDEF || bestTuning.proto == NCCL_PROTO_UNDEF) &&
      bestTuning.symKernelId == ncclSymkKernelId_Count && bestTuning.ceMethodId == ncclCeMethodId_Count) {
    char ncclAlgoEnvStr[1024] = "";
    char ncclProtoEnvStr[1024] = "";
    char ncclSymKernelIdEnvStr[1024] = "";
    const char* symKernelIdEnv = ncclGetEnv("NCCL_SYM_KERNEL");
    if (symKernelIdEnv) {
      snprintf(ncclSymKernelIdEnvStr, 1023, " NCCL_SYM_KERNEL was set to %s.", symKernelIdEnv);
    }
    const char* algoEnv = ncclGetEnv("NCCL_ALGO");
    if (algoEnv) {
      snprintf(ncclAlgoEnvStr, 1023, " NCCL_ALGO was set to %s.", algoEnv);
    }
    const char* protoEnv = ncclGetEnv("NCCL_PROTO");
    if (protoEnv) {
      snprintf(ncclProtoEnvStr, 1023, " NCCL_PROTO was set to %s.", protoEnv);
    }
    WARN("No algorithm/protocol nor symKernelId available for function %s with datatype %s.%s%s%s",
         ncclFuncToString(input->func), ncclDatatypeToString(input->datatype), ncclAlgoEnvStr, ncclProtoEnvStr,
         ncclSymKernelIdEnvStr);
    ret = (algoEnv || protoEnv || symKernelIdEnv) ? ncclInvalidUsage : ncclInternalError;
  }

The choice of error code is deliberate: if the user has set an environment variable (algoEnv || protoEnv || symKernelIdEnv), returnncclInvalidUsage—this is a user configuration problem; otherwise returnncclInternalError—this is an internal NCCL problem (all candidates were unexpectedly excluded).

21.5 Production Pitfall Guide

Pitfall 1: Environment Variable Typos Causing Silent Fallback

parseListreturnsncclInvalidUsagewhen encountering an unrecognized token, but if you writeNCCL_ALGO=RING(uppercase),strcasecmpwill match correctly. What is truly dangerous is a typo, such asNCCL_ALGO=rnig。

📎 src/tuning/cost_model.cc:87-91

code
        if (e == nelems) {
          WARN("Unrecognized element token \"%s\" when parsing \"%s\"", elem, str);
          ret = ncclInvalidUsage;
          goto fail;
        }

Here it will print WARN and return an error. But if you have not enabledNCCL_DEBUG=WARN, you may not see this warning.Recommendation: always setNCCL_DEBUG=WARNorNCCL_DEBUG=INFOwhen tuning to ensure you can see the results of configuration parsing.

Pitfall 2: Interaction Between NCCL_ALGO and NCCL_PROTO

If you setNCCL_ALGO=treebut do not setNCCL_PROTO, NCCL will choose the optimal protocol under the Tree algorithm. But if you set bothNCCL_ALGO=treeandNCCL_PROTO=LL, and the Tree/LL combination is disabled for certain functions (for example, Tree is only enabled for AllReduce), it will trigger a "no available combination" error.

📎 src/tuning/cost_model.cc:379-383

code
      if (((algo != NCCL_ALGO_UNDEF && algoEnable[f * NCCL_NUM_ALGORITHMS + algo] != 0) &&
           (proto != NCCL_PROTO_UNDEF && protoEnable[f * NCCL_NUM_PROTOCOLS + proto] != 0)) ||
          (symKernelId != ncclSymkKernelId_Count && symKernelIdEnable[f * ncclSymkKernelId_Count + symKernelId] != 0)) {
        comm->tuningContext.enabled[i][f] = 1;
      }

Only when the algorithm and protocol aresimultaneouslyallowed is the combination enabled. This is AND logic, not OR.

Pitfall 3: Platform Limitations of LL128

LL128 is not supported on all platforms.isLL128EnabledChecked compute capability, driver version, and connection type.

📎 src/tuning/cost_model.cc:119-139

code
static int isLL128Enabled(int minCompCap, int maxCompCap, int interType, int intraType, int nRanks, int func, int algo,
                          int minDriverVersion) {
  int ret = 1;
  if (ncclParamLl128C2c() && minCompCap >= 90 && (!RUBIN_AND_LATER(minCompCap) || minDriverVersion >= 13030)) {
    // Rubin, Blackwell, and Hopper: Enable LL128 for all P2C and PXN if CUDA supports it.
    ret &= (interType <= PATH_PXN);
  } else {
    // Enable LL128 only up to PXB. Don't enable LL128 over PxN because PxN can encapsulate PxB or P2C links.
    ret &= (interType <= PATH_PXB);
    if (!ncclParamLl128C2c() && minCompCap >= 90)
      INFO(
        NCCL_GRAPH | NCCL_TUNING,
        "Disabling LL128 over all PxN connections (PXB and C2C). This ensures that no C2C link will be used by LL128.");
  }
  ret &= (intraType <= PATH_NVB);
  // Enable LL128 for interoperability between GPUs with different compcap (Hopper and above)
  ret &= (minCompCap == maxCompCap || minCompCap >= 90);
  ret &= !(minCompCap < 70 || (minCompCap == 90 && CUDART_VERSION == 11080 && func == ncclFuncAllReduce &&
                               algo == NCCL_ALGO_RING && nRanks == 2));
  return ret;
}

Several key limitations:

  • minCompCap < 70: GPUs before Volta do not support LL128.
  • intraType <= PATH_NVB: Intra-node connections must be at NVLink level.
  • Hopper + CUDA 11.8 + AllReduce + Ring + 2 ranks: This is a known bug scenario and is explicitly excluded.

Recommendations: If your platform does not support LL128, do not force settingNCCL_PROTO=LL128, otherwise it will trigger an error. Let NCCL choose automatically.

Pitfall Four: Number of Channels and VRAM

The more channels there are, the larger the buffers required. In scenarios with tight VRAM, too many channels may cause OOM.

📎 src/tuning/tuning.cc:246-253

code
        int recChannels;
        NCCLCHECKGOTO(ncclNvlsRegResourcesQuery(input->comm, input->func, &recChannels), ret, exit);
        if (recChannels <= bestTuning.nChannels) {
          bestTuning.algo = NCCL_ALGO_NVLS;
          bestTuning.proto = NCCL_PROTO_SIMPLE;
          bestTuning.nChannels = recChannels;
          bestTuning.maxChannels = recChannels;
          bestTuning.nWarps = input->comm->tuningContext.maxThreads[bestTuning.algo][bestTuning.proto] / WARP_SIZE;
        }

The number of channels for NVLS is determined byncclNvlsRegResourcesQueryquerying hardware resources, not set arbitrarily. If hardware resources are insufficient, the number of channels will be limited.

21.6 Tuning Decision Flow

Connect the previous content together to obtain an actionable troubleshooting process.

mermaid
flowchart TD
    start["性能不达标"] --> baseline["跑 nccl-tests 对比官方报告"]
    baseline --> diff{"差距 > 5%?"}
    diff -->|否| app["检查应用层:<br/>通信频率、消息切分"]
    diff -->|是| debug["设置 NCCL_DEBUG=INFO<br/>查看算法/协议选择"]
    debug --> check_algo{"选择的算法合理?"}
    check_algo -->|否| force_algo["尝试 NCCL_ALGO 强制<br/>对比不同算法"]
    check_algo -->|是| check_proto{"协议合理?"}
    check_proto -->|否| force_proto["尝试 NCCL_PROTO 强制<br/>小消息 LL,大消息 Simple"]
    check_proto -->|是| check_chan{"通道数合理?"}
    check_chan -->|否| tune_chan["调整 NCCL_NCHANNELS<br/>或检查显存限制"]
    check_chan -->|是| check_topo["检查拓扑:<br/>NCCL_TOPO_DUMP 确认链路"]
    force_algo --> verify["重新 benchmark 验证"]
    force_proto --> verify
    tune_chan --> verify
    check_topo --> verify
    verify --> improved{"性能提升?"}
    improved -->|是| done["固化配置"]
    improved -->|否| escalate["提交 issue 或联系支持"]

The core idea of this process is:locate first, then tune parameters, and finally verify. Do not randomly set environment variables right away.

Chapter Summary

This chapter breaks down the NCCL tuning path into four levels:

1. Baseline: Use official performance reports to establish expectations. Within 5% is normal fluctuation. For large messages, look at bandwidth; for small messages, look at latency.

2. Cost model: Internally, NCCL uses themodelMaptable plus latency/bandwidth parameters to estimate the time cost of each combination and selects the smallest one. Understanding this model is the prerequisite for tuning parameters.

3. Environment variables:NCCL_ALGO、NCCL_PROTO、NCCL_SYM_KERNELare parsed throughparseListand then modify theenabledtable to force or exclude specific combinations. The syntax supports three modes: global, per-function, and exclusion.

4. Number of channels: Calculated byncclTuningGetChannels, and affected by hardware resources and CTAPolicy.

Chapter Review Questions

Q1: If the single-rank short-circuit logic inncclTuningCompute(theinput->comm->nRanks <= 1branch) is removed, what will happen? In what scenarios will it cause problems?

Reference Analysis:

The single-rank short-circuit in📎 src/tuning/tuning.cc:191-200:

cpp
  // Set tuning to Ring/Simple for single rank case
  if (input->comm->nRanks <= 1) {
    bestTuning.algo = NCCL_ALGO_RING;
    bestTuning.proto = NCCL_PROTO_SIMPLE;
    bestTuning.symKernelId = ncclSymkKernelId_Count;
    bestTuning.ceMethodId = ncclCeMethodId_Count;
    bestTuning.nChannels = 0;
    bestTuning.maxChannels = 0;
    bestTuning.nWarps = 0;
    bestTuning.forced = 0;
  } else {
    NCCLCHECKGOTO(ncclTuningComputeAllTunings(input, &tunings), ret, exit);
    ...

If this branch is removed, the single-rank scenario will enterncclTuningComputeAllTuningsand iterate through all candidate combinations. The problem is:

1. Performance waste: A single rank has no communication. The time estimates for all algorithms are pure overhead, and choosing any of them makes no difference. Iterating through all candidates is pure waste.

2. It may fail to select a result: Some algorithms may be judged invalid by the model under a single rank (for example, Ring requires at least 2 ranks to form a ring), causing thetuningslist to be empty,ncclTuningSelectBestTuningreturning the initial value ofFLT_MAX, and ultimatelybestTuning.algostill beingNCCL_ALGO_UNDEF。

3. Triggering the error path: IfbestTuning.algo == NCCL_ALGO_UNDEF, it will enter the error handling of📎 src/tuning/tuning.cc:308-329, print the "No algorithm/protocol available" warning, and returnncclInternalError。

So this short-circuit is not just an optimization, but also a correctness guarantee - the single-rank scenario must have a definite default value.

Q2: parseListInforced[p] = 1, what is the purpose of this line of code (📎 src/tuning/cost_model.cc:83)? If it is removed,NCCL_ALGO=ringhow will the behavior of

change?:

forced[p] = 1Reference Analysis📎 src/tuning/cost_model.cc:80-85:

cpp
        for (e = 0; e < nelems; e++) {
          if (strcasecmp(elem, elems[e]) == 0) {
            list[p * nelems + e] = set;
            forced[p] = 1;
            break;
          }
        }

forcedCopyncclTuningContext_tThe array is defined in

There are only three key knobs: algorithm, protocol, and number of channels. Most other parameters are auxiliary diagnostics or optimizations for specific scenarios. Once you have mastered this tuning path, you can already make NCCL achieve performance close to the hardware in most scenarios. But beyond performance, production environments have another, more troublesome class of problems: code that appears normal may hang or fail under specific conditions. In the next chapter, we will summarize typical NCCL pitfalls in production - deadlocks, timeouts, version mismatches, and common misuse - and see how NCCL internally detects and reports these problems.

CHAPTER 22

Chapter 22: Chapter 22: Production Troubleshooting and Pitfalls: Common Deadlocks, Timeouts, Version Mismatches, and Troubleshooting Solutions

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 22 / 25

Chapter 22: Production Troubleshooting and Pitfalls: Common Deadlocks, Timeouts, Version Mismatches, and Troubleshooting Solutions

In the previous chapter, we sorted out the troubleshooting order and key knobs for performance tuning, but NCCL failures in production environments are often not due to performance falling short, but because the program hangs directly or crashes. The root cause of these failures is usually not that some function was written incorrectly, but that the call order, lifecycle, or version contract was violated. This chapter focuses on the four most typical pitfalls: deadlocks caused by misuse of group semantics, silent errors caused by missing parameter validation, ABI version mismatches, and the boundaries of timeouts and retries. We will follow four clues - src/group.cc, src/misc/argcheck.cc, src/include/checks.h, and contrib/nccl_ep/nccl_ep.cc - to see clearly how NCCL blocks errors before they occur.

Misuse of Group Semantics: Why "forgetting a GroupEnd" causes a hang

Intuitive model: Group is a "shopping cart", not an "acceleration switch"

Think ofncclGroupStart() / ncclGroupEnd()as an online shopping cart: you put multiple items (multiple communication calls) into the cart, and finally check out all at once (ncclGroupEnd). If you only add items without checking out, the cart remains suspended forever - thencclGroupDepthcounter maintained internally by NCCL will not return to zero, and all subsequent communication calls will think they are "still accumulating the order", never actually launching the kernel, so the entire process hangs.

[Design inference and architectural trade-offs]

This is the most common deadlock pattern in production: code in some exception branchreturn, skippingncclGroupEnd, andncclGroupDepthisthread_local, and will not be automatically cleaned up when the function returns.

Data structure: thread_local group state

NCCL stores all group state in thread-local storage, which is the key to understanding the deadlock.

📎 src/group.cc:34-34

cpp
thread_local int ncclGroupDepth = 0; // depth of ncclGroupStart nesting
thread_local ncclResult_t ncclGroupError = ncclSuccess;
thread_local struct ncclComm* ncclGroupCommHead[ncclGroupTaskTypeNum] = {nullptr};
thread_local struct ncclComm* ncclGroupCommPreconnectHead = nullptr;
thread_local struct ncclIntruQueue<struct ncclAsyncJob, &ncclAsyncJob::next> ncclAsyncJobs;
thread_local int ncclGroupBlocking = -1; /* default mode */

Field-by-field interpretation:

  • ncclGroupDepth: nesting depth.ncclGroupStartincrements,ncclGroupEnddecrements, and only when it reaches 0 does it actually trigger submission. Supporting nesting is a design convenience, but it also means that "missing one End" will leave the depth stuck at 1 forever.
  • ncclGroupError: the group error accumulated by this thread. Once a call fails, subsequentncclGroupEndwill directly take the failure path.
  • ncclGroupCommHead[]: the heads of the communication domain linked lists grouped by task type (collective / rawTask / mgmtTask / symRegister).
  • ncclAsyncJobs: the queue of asynchronous tasks to be executed (such as preconnect, symmetric register).
  • ncclGroupBlocking:-1means "no communication domain has been encountered yet,"0means non-blocking,1means blocking. This field is the core of the later "mixed blocking and non-blocking" detection.
[Design inference and architectural trade-offs]

Usingthread_localinstead of a global variable has a straightforward motivation: NCCL allows multiple threads to each hold independent group contexts without interfering with each other. The cost is that these states are not automatically cleaned up when a thread exits. If a thread exits in the middle of a group, the state leaks.

Step-by-Step: the complete validation chain of a GroupEnd

Scenario: the application callsncclGroupEnd(), and at this pointncclGroupDepthis 1.

Step one, check whether it is really inside a group:

📎 src/group.cc:1048-1052

cpp
  if (ncclGroupDepth == 0) {
    WARN("ncclGroupEnd: not in a group call.");
    ret = ncclInvalidUsage;
    goto exit;
  }

If the user did not callncclGroupStartand directlyncclGroupEnd, this will print "not in a group call" and returnncclInvalidUsage. This is the friendliest error - it reports immediately and will not hang.

Step two, decrement the depth and determine whether this is the outermost layer:

📎 src/group.cc:1061-1063

cpp
  if ((--ncclGroupDepth) > 0) goto exit;

  if ((ret = ncclGroupError) != ncclSuccess) goto fail;

If multiple layers are nested, the innerEndonly decrements the depth and returns without triggering submission. Only the outermost layer continues. At the same time, accumulated errors are checked.

Step three, validate consistency of the blocking mode. This is the detection point for "mixed blocking and non-blocking":

📎 src/group.cc:1095-1101

cpp
  if (hasCommHead || !ncclIntruQueueEmpty(&groupJob->asyncJobs) || ncclGroupCommPreconnectHead != nullptr) {
    /* make sure ncclGroupBlocking has been set. */
    if (ncclGroupBlocking != 0 && ncclGroupBlocking != 1) {
      WARN("Invalid group blocking state %d", ncclGroupBlocking);
      ret = ncclInternalError;
      goto fail;
    }

ncclGroupBlockingmust be between{0, 1}. If it is still-1, it means the group contains neither a communication domain nor an asynchronous task, and logically execution should not reach here.

Step four, branch according to the blocking mode. Non-blocking goes through asynchronous submission by a thread, while blocking goes through synchronous submission:

📎 src/group.cc:1102-1134

cpp
    if (ncclGroupBlocking == 0) {
      /* nonblocking group */
      if (!ncclIntruQueueEmpty(&groupJob->asyncJobs)) {
        ncclAsyncJob* job = ncclIntruQueueHead(&groupJob->asyncJobs);
        do {
          NCCLCHECKGOTO(ncclCommSetAsyncError(job->comm, ncclInProgress), ret, fail);
          if (job->comm->groupJob == NULL) {
            job->comm->groupJob = groupJob;
            groupJob->groupRefCount++;
          }
          job = job->next;
        } while (job);
      }
      ...
      groupJob->base.func = groupLaunchNonBlocking;
      STDTHREADCREATE_GOTO(groupJob->base.thread, ncclAsyncJobMain, ret, fail, &groupJob->base);
      groupJob->nonBlockingInit = true;
      ret = ncclInProgress;
    }

NotegroupRefCount++andret = ncclInProgress: in non-blocking mode,ncclGroupEndreturns immediately withncclInProgress, and the actual submission runs in a background thread. The caller must subsequently poll withncclCommGetAsyncError, or wait withncclGroupJobComplete.

Mixing blocking and non-blocking: why it is forbidden

Return toncclAsyncLaunchand look at the mixed-use detection:

📎 src/group.cc:55-64

cpp
    /* check if there are blocking and nonblocking comms at the same time in group. */
    if (comm->destroyFlag) {
      ncclGroupBlocking = 1;
    } else if (ncclGroupBlocking == -1) {
      /* first met communicator */
      ncclGroupBlocking = comm->config.blocking;
    } else if (ncclGroupBlocking != comm->config.blocking) {
      WARN("Blocking and nonblocking communicators are not allowed in the same group.");
      ret = ncclInvalidArgument;
    }
[Design inference and architectural trade-offs]

Why is mixing forbidden? Because the submission semantics of a blocking communication domain are "the kernel has been submitted when the call returns," while non-blocking means "the task has been enqueued but not submitted when the call returns." If both are in the same group,ncclGroupEndcannot provide a unified return semantics - should it wait or not? NCCL chooses to reject it outright and expose the problem at the API boundary.

Production pitfalls: three real scenarios

Scenario one: an exception branch misses GroupEnd.Code throws an exception betweenncclGroupStartandncclGroupEndor returns earlyreturn,ncclGroupDepthstays at 1. All subsequent communication calls enter a "batching" state and are never submitted. Troubleshooting method: printncclGroupEndbeforencclGroupDepth, or usegdbto observe that thread_local variable.

Scenario two: using the same comm across threads.Because group state isthread_local, after thread A callsncclGroupStart, thread B callingncclAllReducewill not enter A's group. If A and B operate on the same comm, there will be confusion where "some calls are inside the group and some are outside the group." NCCL does not detect this situation because it assumes that a comm is operated on by only one thread at any given time.

Scenario three: the interaction between CUDA graph capture and groups.Look at the detection indoLaunches:

📎 src/group.cc:448-455

cpp
    if (capturingYes && capturingNo) {
      // We have entered barriers but are aborting without leaving them. Thus
      // these comms are permanently trashed. We need a good mechanism for
      // tracking and reporting that.
      WARN("Either none or all communicators in a ncclGroup() can be CUDA graph captured.");
      result = ncclInvalidUsage;
      goto failure;
    }

The comment says it very directly: once a barrier is entered and then abandoned midway, these comms are "permanently corrupted." So the rule is - all communication domains in a group must either all be in capture or all not be. Mixing them will cause inconsistent comm state, and NCCL currently has no good recovery mechanism.

mermaid
flowchart TD
    start["ncclGroupEnd()"] --> depth_check{"ncclGroupDepth == 0?"}
    depth_check -->|是| err_usage["WARN not in a group call<br/>return ncclInvalidUsage"]
    depth_check -->|否| dec["--ncclGroupDepth"]
    dec --> nested{"depth > 0?"}
    nested -->|是| exit_ok["goto exit 返回"]
    nested -->|否| err_check{"ncclGroupError == success?"}
    err_check -->|否| fail_clean["groupCleanup 清理所有 comm 与 asyncJobs"]
    err_check -->|是| blocking_check{"ncclGroupBlocking in {0,1}?"}
    blocking_check -->|否| err_internal["WARN Invalid group blocking state<br/>return ncclInternalError"]
    blocking_check -->|是| mode_split{"ncclGroupBlocking == 0?"}
    mode_split -->|是 非阻塞| async_launch["STDTHREADCREATE groupLaunchNonBlocking<br/>ret = ncclInProgress"]
    mode_split -->|否 阻塞| sync_launch["groupLaunch 同步下发<br/>delete groupJob"]
    async_launch --> reset["groupLocalResetJobState"]
    sync_launch --> reset
    reset --> exit_ok
    fail_clean --> reset

Parameter validation and silent errors: how ArgCheck blocks calls that "look normal"

Intuitive model: ArgCheck is "airport security"

Parameter validation is like airport security: it is not responsible for making you fly faster, but it can block those things that "look like luggage but are actually dangerous goods." Without it, a pointer with the wrong device passed in will make the GPU kernel read garbage data, or worse - silently corrupt someone else's memory.

Data structure: validation modes and the global check queue

NCCL parameter validation is not "check everything every time," but is divided into modes. The core iscomm->checkMode:

📎 src/misc/argcheck.cc:227-251

cpp
  if (info->comm->checkMode != ncclCheckModeDefault) {
    if ((info->coll == ncclFuncSend || info->coll == ncclFuncRecv)) {
      if (info->count > 0) NCCLCHECK(CudaPtrCheck(info->recvbuff, info->comm, "buff", info->opName));
    } else if (info->coll == ncclFuncPutSignal || info->coll == ncclFuncSignal || info->coll == ncclFuncWaitSignal) {
      // One-sided RMA ops specify the remote destination via peerWin, not sendbuff/recvbuff,
      // so the standard CUDA pointer checks do not apply here.
      INFO(NCCL_COLL, "%s : skipping sendbuff/recvbuff pointer check (one-sided RMA uses peerWin)", info->opName);
    } else {
      // Check CUDA device pointers
      if (info->coll != ncclFuncBroadcast || info->comm->rank == info->root) {
        NCCLCHECK(CudaPtrCheck(info->sendbuff, info->comm, "sendbuff", info->opName));
      }
      if (info->coll != ncclFuncReduce || info->comm->rank == info->root) {
        NCCLCHECK(CudaPtrCheck(info->recvbuff, info->comm, "recvbuff", info->opName));
      }
    }

    if (info->comm->checkMode == ncclCheckModeDebugGlobal) {
      struct ncclArgsInfo* argsInfo;
      NCCLCHECK(ncclCalloc(&argsInfo, 1));
      argsInfo->info = *info;
      argsInfo->next = NULL;
      ncclIntruQueueEnqueue(&info->comm->argsInfoQueue, argsInfo);
    }
  }

Three modes:

  • ncclCheckModeDefault: only the cheapest checks are performed (root range, datatype range, op range), without touching the CUDA API.
  • Non-default mode: callsCudaPtrCheck, which actually callscudaPointerGetAttributes, and has performance overhead.
  • ncclCheckModeDebugGlobal: in addition to local checks, it also putsncclInfointoargsInfoQueue, perform a cross-rank global consistency check when the group ends.
[Design Inference and Architectural Trade-offs]

This design is a trade-off between performance and correctness:cudaPointerGetAttributesIt is a synchronous CUDA call, and calling it on every communication in the hot path would significantly slow down small messages. So the default mode only performs "zero-cost" checks, leaving expensive pointer validation to debug mode.

Step-by-Step: The Three Layers of Defense in CudaPtrCheck

Scenario: The user passes in asendbuff, and NCCL validates it in debug mode.

First layer, whether the pointer is valid:

📎 src/misc/argcheck.cc:12-18

cpp
ncclResult_t CudaPtrCheck(const void* pointer, struct ncclComm* comm, const char* ptrname, const char* opname) {
  cudaPointerAttributes attr;
  cudaError_t err = cudaPointerGetAttributes(&attr, pointer);
  if (err != cudaSuccess || attr.devicePointer == NULL) {
    WARN("%s : %s %p is not a valid pointer", opname, ptrname, pointer);
    return ncclInvalidArgument;
  }

cudaPointerGetAttributesIt returns an error for invalid pointers, ordevicePointeris NULL. This blocks "passed a host stack address" or "passed a freed pointer."

Second layer, whether the device matches:

📎 src/misc/argcheck.cc:19-26

cpp
#if CUDART_VERSION >= 10000
  if (attr.type == cudaMemoryTypeDevice && attr.device != comm->cudaDev) {
#else
  if (attr.memoryType == cudaMemoryTypeDevice && attr.device != comm->cudaDev) {
#endif
    WARN("%s : %s allocated on device %d mismatchs with NCCL device %d", opname, ptrname, attr.device, comm->cudaDev);
    return ncclInvalidArgument;
  }

This is the most insidious pitfall: the pointer is a valid GPU pointer, but it belongs to another GPU. On multi-GPU machines, if the user forgetscudaSetDevice, it is very easy to pass the wrong one. NCCL explicitly rejects it here.

Third layer, communication domain object integrity:

📎 src/misc/argcheck.cc:38-45

cpp
ncclResult_t CommCheck(struct ncclComm* comm, const char* opname, const char* ptrname) {
  NCCLCHECK(PtrCheck(comm, opname, ptrname));
  if (comm->startMagic != NCCL_MAGIC || comm->endMagic != NCCL_MAGIC) {
    WARN("Error: corrupted comm object detected");
    return ncclInvalidArgument;
  }
  return ncclSuccess;
}

startMagic / endMagicIt is a sentinel value placed at the beginning and end of thencclCommstruct. If the user passes a wild pointer, or comm has already been freed, the magic will not match. This is the classic technique for "memory corruption detection" — sandwiching the struct with two sentinels, so any out-of-bounds write is likely to corrupt one of them.

Global Consistency Check: registrationCheck's Cross-Rank Validation

This is the "heaviest" validation in NCCL, and is triggered only underncclCheckModeDebugGlobal. What it checks is — whether the symmetric memory registration state of all ranks is consistent.

📎 src/misc/argcheck.cc:95-111

cpp
  NCCLCHECKGOTO(bootstrapAllGather(comm->bootstrap, bufInfo, sizeof(struct symBufInfo) * 2), ret, fail);

  cmpBufInfo[0] = bufInfo[0];
  cmpBufInfo[1] = bufInfo[1];
  for (int r = 1; r < comm->nRanks; r++) {
    int infoIdx = r * 2;
    if (cmpBufInfo[0].isSymRegistered != bufInfo[infoIdx].isSymRegistered ||
        cmpBufInfo[1].isSymRegistered != bufInfo[infoIdx + 1].isSymRegistered) {
      if (comm->rank == 0) {
        WARN("Coll %s size %ld symmetric registration check failed on rank %d: sendReg %d recvReg %d mismatch with "
             "rank 0 sendReg %d recvReg %d",
             info->opName, size, r, bufInfo[infoIdx].isSymRegistered, bufInfo[infoIdx + 1].isSymRegistered,
             cmpBufInfo[0].isSymRegistered, cmpBufInfo[1].isSymRegistered);
      }
      ret = ncclInvalidArgument;
      goto fail;
    }

It uses the bootstrap'sallGatherto collect each rank's(isSymRegistered, bigOffset, userOffset), then compares them rank by rank. If rank 0's send buffer has symmetric memory registered, but rank 3 does not, an error will be reported here.

[Design Inference and Architectural Trade-offs]

Why is this check important? Symmetric memory requires all ranks to access the buffer using the same set of virtual addresses. If some rank's buffer is not registered, the address computed in the kernel will be wrong, causing reads of garbage or out-of-bounds access. This kind of error manifests at runtime as "results are occasionally wrong," and is extremely difficult to troubleshoot. NCCL chooses to block it at the API boundary at the cost of one allGather.

Production Pitfalls

Pitfall 1: Pointer errors are not reported in default mode.If the user has not enabled debug mode and passes a pointer for the wrong device, NCCL will not report an error during theArgsCheckphase, but will only discover it when the kernel executes — by which time it may have already corrupted another rank's GPU memory. It is recommended to useNCCL_DEBUG=WARNpluscheckModefor debugging during development.

Pitfall 2:ncclCheckModeDebugGlobal's allGather overhead.Every communication performs a bootstrap allGather, which becomes a bottleneck in high-frequency small-message scenarios. This mode is only suitable for debugging and cannot be used in production.

Pitfall 3: The lifecycle of userRedOp.Look at this section:

📎 src/misc/argcheck.cc:220-225

cpp
  int opIx = int(ncclUserRedOpMangle(info->comm, info->op)) - int(ncclNumOps);
  if (ncclNumOps <= info->op &&
      (info->comm->userRedOpCapacity <= opIx || info->comm->userRedOps[opIx].freeNext != -1)) {
    WARN("%s : reduction operation %d unknown to this communicator", info->opName, info->op);
    return ncclInvalidArgument;
  }

The user-defined reduction op is registered on comm. If the user passes an op that "was once registered but has already been freed,"freeNext != -1will detect that it has been reclaimed. This is a check to prevent "dangling op handles."

Error Propagation Macros: How the NCCLCHECK Family Ensures "Errors Are Not Lost"

Intuitive Model: Error propagation macros are a "relay baton"

NCCL's error handling relies on a set of macros in a relay: the lower-level function returnsncclResult_t, the upper layer usesNCCLCHECKto check, and if it is not successful, it returns immediately. This is like a relay race — the baton (error code) must be passed all the way to the end; if any leg drops it, the whole chain breaks.

Data Structures: Overview of the Macro Family

📎 src/include/checks.h:148-166

cpp
#define NCCLCHECK(call) \
  do { \
    ncclResult_t RES = call; \
    if (RES != ncclSuccess && RES != ncclInProgress) { \
      /* Print the back trace*/ \
      if (ncclDebugNoWarn == 0) INFO_LOC(NCCL_ALL, "-> %d", RES); \
      return RES; \
    } \
  } while (0)

#define NCCLCHECKGOTO(call, RES, label) \
  do { \
    RES = call; \
    if (RES != ncclSuccess && RES != ncclInProgress) { \
      /* Print the back trace*/ \
      if (ncclDebugNoWarn == 0) INFO_LOC(NCCL_ALL, "-> %d", RES); \
      goto label; \
    } \
  } while (0)

Key details:ncclInProgressis treated as "not an error." This is the core of non-blocking communication —ncclGroupEndreturningncclInProgressmeans "the task has been submitted but not yet completed," and the caller should continue polling rather than treating it as an error.

NCCLCHECKdirectlyreturn,NCCLCHECKGOTOjumps tolabel. The latter is used in scenarios that require resource cleanup.

Cleanup Path: NCCLCHECKIGNORE Preserves the First Error

📎 src/include/checks.h:168-177

cpp
// Report failure but continue - useful for cleanup paths where we want to
// attempt all cleanup steps. Preserves the first error in RES.
#define NCCLCHECKIGNORE(call, RES) \
  do { \
    ncclResult_t TMPRES = call; \
    if (TMPRES != ncclSuccess && TMPRES != ncclInProgress) { \
      if (ncclDebugNoWarn == 0) INFO_LOC(NCCL_ALL, "-> %d", TMPRES); \
      if (RES == ncclSuccess) RES = TMPRES; \
    } \
  } while (0)

The comment makes it very clear: on the cleanup path, it should "attempt all cleanup steps" and must not be interrupted by the first error. But the error code should preserve the first one — because the first error is usually the root cause with the most diagnostic value.

Waiting and Aborting: NCCLWAIT's abortFlag Check

📎 src/include/checks.h:196-205

cpp
#define NCCLWAIT(call, cond, abortFlagPtr) \
  do { \
    uint32_t* tmpAbortFlag = (abortFlagPtr); \
    ncclResult_t RES = call; \
    if (RES != ncclSuccess && RES != ncclInProgress) { \
      if (ncclDebugNoWarn == 0) INFO_LOC(NCCL_ALL, "-> %d", RES); \
      return ncclInternalError; \
    } \
    if (COMPILER_ATOMIC_LOAD(tmpAbortFlag, std::memory_order_acquire)) NEQCHECK(*tmpAbortFlag, 0); \
  } while (!(cond))

This is the template for polling waits: each loop iteration callscall(advance progress), checkscond(whether it is satisfied), and also checksabortFlag(whether it has been aborted).abortFlagusesmemory_order_acquireloading to ensure it sees the abort signal written by other threads.

[Design Inference and Architectural Trade-offs]

This design solves a classic problem: when one rank errors out, other ranks may still be waiting forever for its data.abortFlagis the mechanism for propagating the abort signal across ranks — once set, all waiting loops will exit.

Safe Macros for Thread Creation and Memory Allocation

📎 src/include/checks.h:237-256

cpp
#define STDTHREADCREATE_IMPL(var, func, error_action, ...) \
  do { \
    try { \
      (var) = std::thread(func, __VA_ARGS__); \
    } catch (const std::exception& e) { \
      WARN("Thread creation failed: %s", e.what()); \
      error_action; \
    } \
  } while (0)

#define STDTHREADCREATE(var, func, ...) STDTHREADCREATE_IMPL(var, func, return ncclSystemError, __VA_ARGS__)

#define STDTHREADCREATE_GOTO(var, func, RES, label, ...) \
  STDTHREADCREATE_IMPL( \
    var, func, \
    do { \
      RES = ncclSystemError; \
      goto label; \
    } while (0), \
    __VA_ARGS__)

std::threadConstruction failure throws an exception (for example, exceeding the thread limit). This macro converts the exception intoncclSystemError, preventing exceptions from crossing the C API boundary.

📎 src/include/checks.h:258-275

cpp
#define NEW_NOTHROW(var, x) \
  do { \
    (var) = new (std::nothrow) x{}; \
    if (!(var)) { \
      WARN("Allocation failed"); \
      return ncclSystemError; \
    } \
  } while (0)

new (std::nothrow)Returns nullptr instead of throwing an exception when allocation fails. This is the standard practice for C++ code at the C API boundary.

Production Pitfalls

Pitfall 1:ncclInProgressis mistakenly treated as success.Some user code writesif (ret == ncclSuccess)to determine success, but in non-blocking mode what is returned isncclInProgress. The correct approach isif (ret == ncclSuccess || ret == ncclInProgress), or usencclCommGetAsyncErrorto query.

Pitfall 2:NCCLCHECKUsed in destructors.If used in a destructorNCCLCHECK, the error will directlyreturn, skipping subsequent cleanup. Should useNCCLCHECKIGNORE。

ABI version mismatch: nccl_ep's size-based design

Intuitive model: ABI is a "socket standard"

ABI (Application Binary Interface) is like a power socket standard: if the library and the caller have inconsistent understanding of "what the struct looks like," it's like plugging a US-standard plug into a European-standard socket—at best it won't work, at worst it burns out.contrib/nccl_epuses a clever design: every cross-boundary struct starts with asizefield.

Data structure: size + magic dual validation

📎 contrib/nccl_ep/nccl_ep.cc:70-76

cpp
// Size-based ABI versioning: every cross-boundary struct starts with a `size`
// field set by the caller to sizeof(struct). The library checks that against
// its own known size; any mismatch means caller and library are from different
// releases. Strict equality for now — see nccl_ep.h for the planned future
// relaxation (all-zero-trailing-bytes escape hatch).
// Immediately after `size` there is a `magic` field pre-filled by NCCL_EP_*_INIT
// to catch unininitialized structures.

Design points:

  • sizeThe field is filled in by the caller withsizeof(struct), and the library checks whether it equals the size it recognizes.
  • magicThe field is pre-filled by theNCCL_EP_*_INITmacro, used to catch "uninitialized" structs.
  • Currently it's strict equality; in the future there are plans to support a lenient mode where "if the tail is all zeros, a smaller size is allowed."

Step-by-Step: EP_REQUIRE_STRUCT's validation flow

📎 contrib/nccl_ep/nccl_ep.cc:77-80

cpp
#define EP_REQUIRE_STRUCT(ptr) \
    do { \
        assert( \
            (ptr) != nullptr && (ptr)->size == sizeof(*(ptr)) && \

This macro is called at entry points such asncclEpDispatch、ncclEpCombine:

📎 contrib/nccl_ep/nccl_ep.cc:2827-2830

cpp
    EP_REQUIRE_STRUCT(inputs);
    EP_REQUIRE_STRUCT(outputs);
    EP_OPTIONAL_LAYOUT_INFO(layout_info);
    EP_OPTIONAL_STRUCT(config);

inputsandoutputsare required parameters, usingEP_REQUIRE_STRUCT;layout_infoandconfigare optional parameters, usingEP_OPTIONAL_*。

Version-safe field reading: layoutInfoRecvTopkIdxKind

This is the most ingenious part—how to safely read a field when "the caller's struct may be smaller."

📎 contrib/nccl_ep/nccl_ep.cc:139-144

cpp
// Safe field reader for ncclEpLayoutInfo_t::recv_topk_idx_kind. Returns AUTO
// when the caller's struct (size) does not cover the field, preserving the
// pre-flag default.
static inline ncclEpExpertIdKind_t layoutInfoRecvTopkIdxKind(const ncclEpLayoutInfo_t* lip) {
    if (lip == nullptr) return NCCL_EP_EXPERT_ID_AUTO;
    constexpr size_t field_end = offsetof(ncclEpLayoutInfo_t, recv_topk_idx_kind) + sizeof(ncclEpExpertIdKind_t);
    if (lip->size < field_end) return NCCL_EP_EXPERT_ID_AUTO;
    return lip->recv_topk_idx_kind;
}

The logic is: if the caller'ssizeis less than "the offset where this field ends," it means the caller is using an older version of the struct, this field doesn't exist, and the default valueAUTOis returned. Otherwise, read normally.

[Design inference and architectural trade-offs]

This is the standard technique for ABI compatibility: new fields can only be added at the end of the struct, and when reading, usesizeto determine whether the field exists. This way, old callers use the old struct, and the new library can still handle it correctly.

Version number check: soft warning rather than hard rejection

📎 contrib/nccl_ep/nccl_ep.cc:1393-1400

cpp
    if (in_config->version != NCCL_EP_API_VERSION) {
        fprintf(
            stderr,
            "NCCL EP WARN: ncclEpGroupConfig_t.version=%u, library API_VERSION=%u; "
            "behavior may differ across versions.\n",
            in_config->version,
            (unsigned)NCCL_EP_API_VERSION);
    }

Note that here it'sWARNrather thanreturn error. A version number mismatch is only a warning, because thesizecheck already guarantees memory layout safety. The version number is more of a hint that "behavior may differ."

Production pitfalls

Pitfall 1: Forgetting to initialize with the INIT macro.If the user manuallymemsetthe struct to 0,magicwill be 0,EP_REQUIRE_STRUCTwill fail. Must use theNCCL_EP_*_INITmacro.

Pitfall 2: Mixing dynamic libraries across versions.If the application links against a new version oflibnccl_ep.so, but the header file is an old version,sizeof(struct)will be inconsistent,EP_REQUIRE_STRUCTwill immediately report an error. This is by design—failing fast is better than silent errors.

Pitfall 3:EP_OPTIONAL_LAYOUT_INFOrange check.Look at this:

📎 contrib/nccl_ep/nccl_ep.cc:114-123

cpp
            if ((ptr)->size < kNcclEpLayoutInfoMinSize || (ptr)->size > sizeof(*(ptr))) { \
                fprintf( \
                    stderr, \
                    "NCCL EP: ncclEpLayoutInfo_t size out of supported range: " \
                    "got %u, expected [%zu, %zu]\n", \
                    (ptr)->size, \
                    kNcclEpLayoutInfoMinSize, \
                    sizeof(*(ptr))); \
                return ncclInvalidArgument; \
            } \

layout_infoallows size within the[min, sizeof]range, which is more lenient thanEP_REQUIRE_STRUCT's strict equality. The reason is thatlayout_infois an optional parameter, and historically fields have been added and removed.

mermaid
flowchart TD
    entry["ncclEpDispatch(inputs, outputs, layout_info, config)"] --> req_inputs{"EP_REQUIRE_STRUCT(inputs)<br/>size == sizeof?"}
    req_inputs -->|否| err_size["assert 失败 / 返回错误"]
    req_inputs -->|是| req_outputs{"EP_REQUIRE_STRUCT(outputs)"}
    req_outputs -->|否| err_size
    req_outputs -->|是| opt_layout{"layout_info != nullptr?"}
    opt_layout -->|否| skip_layout["跳过 layout 校验"]
    opt_layout -->|是| range_check{"size in [min, sizeof]?"}
    range_check -->|否| err_range["fprintf size out of range<br/>return ncclInvalidArgument"]
    range_check -->|是| magic_check{"magic == NCCL_EP_MAGIC?"}
    magic_check -->|否| err_magic["fprintf magic mismatch<br/>return ncclInvalidArgument"]
    magic_check -->|是| read_field["layoutInfoRecvTopkIdxKind<br/>size < field_end ? AUTO : 实际值"]
    skip_layout --> read_field
    read_field --> proceed["继续执行 dispatch 逻辑"]

Timeout, retry, and abort: from NCCLWAIT to nccl_ep's timeout_cycles

Intuitive model: timeout is a "fuse"

In distributed communication, one stuck rank causes all ranks to wait forever. The timeout mechanism is like a fuse: under normal conditions it doesn't act, but once the current is abnormal it blows, preventing the entire system from burning out.

Data structure: abortFlag and timeout_cycles

The NCCL core usesabortFlagto propagate the abort signal. Look at the propagation inncclAsyncLaunch:

📎 src/group.cc:49-52

cpp
    job->abortFlag = comm->abortFlag;
    job->abortFlagDev = comm->abortFlagDev;
    job->childAbortFlag = comm->childAbortFlag;
    job->childAbortFlagDev = comm->childAbortFlagDev;

Each job holds a pointer to the comm's abortFlag. When the group detects an error:

📎 src/group.cc:118-126

cpp
        if (!job->destroyFlag &&
            (COMPILER_ATOMIC_LOAD(groupAbortFlag, std::memory_order_acquire) || errorJobAbortFlag == true)) {
          COMPILER_ATOMIC_STORE(job->abortFlag, uint32_t(1), std::memory_order_release);
          COMPILER_ATOMIC_STORE(job->abortFlagDev, uint32_t(1), std::memory_order_release);
          if (job->childAbortFlag) {
            COMPILER_ATOMIC_STORE(job->childAbortFlag, uint32_t(1), std::memory_order_release);
            COMPILER_ATOMIC_STORE(job->childAbortFlagDev, uint32_t(1), std::memory_order_release);
          }
        }

OncegroupAbortFlagorerrorJobAbortFlagis true, all jobs' abortFlag are set to 1.memory_order_releaseensures that previous write operations are visible to other threads.

nccl_ep's timeout design: GPU clock cycles

nccl_epuses a more refined timeout—in units of GPU clock cycles.

📎 contrib/nccl_ep/nccl_ep.cc:1558-1591

cpp
    // Resolve timeout_cycles: env var > config field > compile-time default
    {
        int dev;
        int clock_khz_int;
        CUDA_CHECK(cudaGetDevice(&dev));
        CUDA_CHECK(cudaDeviceGetAttribute(&clock_khz_int, cudaDevAttrClockRate, dev));
        uint64_t clock_khz = static_cast<uint64_t>(clock_khz_int);

        uint64_t resolved = NUM_TIMEOUT_CYCLES;
        const char* source = "compile-time default";
        const uint64_t env_ms = static_cast<uint64_t>(ep_group->env.timeout_ms.value.ul);
        // Only a positive timeout overrides the default.
        const bool have_env_ms = ep_group->env.timeout_ms.is_set && env_ms > 0;

        if (have_env_ms) {
            resolved = clock_khz * 1000ULL * env_ms / 1000ULL;
            source = "NCCL_EP_TIMEOUT_MS env var";
            ...
        } else if (ep_group->config.timeout_ns != 0) {
            resolved = clock_khz * 1000ULL * (ep_group->config.timeout_ns / 1000000ULL) / 1000ULL;
            source = "config.timeout_ns";
        }

        ep_group->timeout_cycles = resolved;

The priority is: environment variableNCCL_EP_TIMEOUT_MS> config fieldtimeout_ns> compile-time default. The conversion formula isclock_khz * 1000 * ms / 1000, i.e., converting milliseconds to clock cycles.

[Design inference and architectural trade-offs]

Why use clock cycles instead of milliseconds? Because the wait loop inside a GPU kernel cannot call system time APIs, it can only read theclock64()register. Using clock cycles for timeout determination allows direct comparison inside the kernel, with no host involvement.

Asynchronous error flag: host-pinned memory

📎 contrib/nccl_ep/nccl_ep.cc:1767-1778

cpp
    // Allocate mask buffer and async error flag for active-mask support
    if (ep_group->config.enable_mask && ep_group->config.algorithm == NCCL_EP_ALGO_LOW_LATENCY) {
        size_t mask_bytes = ep_group->nRanks * sizeof(int);
        CUDA_CHECK(cudaMalloc(reinterpret_cast<void**>(&ep_group->mask_buffer), mask_bytes));
        // Initialize all ranks as active (1 = active, 0 = masked/failed)
        std::vector<int> all_active(ep_group->nRanks, 1);
        CUDA_CHECK(
            cudaMemcpyAsync(ep_group->mask_buffer, all_active.data(), mask_bytes, cudaMemcpyHostToDevice, stream));
        CUDA_CHECK(
            cudaHostAlloc(reinterpret_cast<void**>(&ep_group->async_error_flag), sizeof(int), cudaHostAllocMapped));
        *ep_group->async_error_flag = 0;
    }

async_error_flagusescudaHostAllocMappedfor allocation, which is host-pinned memory mapped into the device address space. The GPU kernel can write to it, and the host can read it, with no explicit copy needed.

Reading asynchronous errors: atomic load

📎 contrib/nccl_ep/nccl_ep.cc:4312-4321

cpp
ncclResult_t ncclEpGetAsyncError(ncclEpGroup_t ep_group, int* error_out) {
    EP_HOST_ASSERT(ep_group != nullptr);
    if (!ep_group->config.enable_mask) {
        return ncclInvalidUsage;
    }
    EP_HOST_ASSERT(ep_group->async_error_flag != nullptr && "ncclEpGetAsyncError: enable_mask must be true");
    EP_HOST_ASSERT(error_out != nullptr);
    *error_out = __atomic_load_n(ep_group->async_error_flag, __ATOMIC_ACQUIRE);
    return ncclSuccess;
}

uses__atomic_load_nplus__ATOMIC_ACQUIRE, ensuring that what's read is the latest value written by the GPU, not a stale cached value.

Production pitfalls

Pitfall 1: Timeout set too short causing false positives.IfNCCL_EP_TIMEOUT_MSis set too small, normal network jitter will be misjudged as a timeout. It's recommended to set it based on the actual network RTT, generally no less than 10 seconds.

Pitfall 2: abortFlag not cleared after being set.Once abortFlag is set to 1, the comm enters an "aborted" state. If the user wants to continue using this comm, abortFlag must first be cleared. NCCL'sncclCommAbortdoes this cleanup.

Pitfall 3:ncclEpMaskClean's precondition.Look at this:

📎 contrib/nccl_ep/nccl_ep.cc:4262-4266

cpp
    EP_HOST_ASSERT(ep_group->config.algorithm == NCCL_EP_ALGO_LOW_LATENCY);
    EP_HOST_ASSERT(
        ep_group->rdma_buffer != nullptr &&
        "ncclEpMaskClean: rdma_buffer not yet allocated; create at least one LL handle first");
    EP_HOST_ASSERT(ep_group->sync_buffer != nullptr && ep_group->sync_window != nullptr);

ncclEpMaskCleanrequiresrdma_bufferto already be allocated. If the user has created a group but hasn't created any LL handle yet,rdma_bufferis nullptr (because LL is lazily allocated), and this will fail the assert.

Chapter summary

This chapter ties together four types of production pitfalls:

1. Group semantics misuse:ncclGroupDepthis thread_local, missingncclGroupEndwill cause a permanent hang; blocking and non-blocking communication domains cannot be mixed; CUDA graph capture must be all-or-nothing.

2. Parameter validation:ArgsCheckMode-specific validation; default mode performs only zero-cost checks;CudaPtrCheckThree layers of defense block invalid pointers, wrong devices, and corrupted comms;registrationCheckPerform cross-rank symmetric memory consistency checks.

3. Error propagation:NCCLCHECKThe family guarantees errors are not lost;ncclInProgressis not an error;NCCLCHECKIGNOREUsed in cleanup paths to preserve the first error;NCCLWAITCheck abortFlag during polling.

4. ABI versioning:nccl_epUses a size-based design; every cross-boundary struct begins withsize, paired withmagicto catch uninitialized fields; new fields can only be appended at the end, and reads usesizeto determine whether they exist.

5. Timeout and abort: The core usesabortFlagto propagate aborts;nccl_epUses GPU clock cycles for timeouts,async_error_flagUses host-pinned memory to implement GPU→host asynchronous notification.

Chapter Review and Self-Test

Q1: If inncclGroupEndInternaltheif ((--ncclGroupDepth) > 0) goto exit;(📎 src/group.cc:1061) is changed toif (ncclGroupDepth > 0) goto exit;(no decrement), what happens? What are the consequences in nested group scenarios?

Reference Analysis:

The original code--ncclGroupDepthdecrements first, then checks. If changed to not decrement:

cpp
if (ncclGroupDepth > 0) goto exit;  // 错误版本

Then each timencclGroupEndthe depth will not decrease. Suppose the user writes:

cpp
ncclGroupStart();  // depth = 1
ncclGroupStart();  // depth = 2
ncclAllReduce(...);
ncclGroupEnd();    // 原版: depth = 1, 返回; 错误版: depth = 2, 返回
ncclGroupEnd();    // 原版: depth = 0, 触发下发; 错误版: depth = 2, 返回

In the buggy version, on the secondncclGroupEnd,ncclGroupDepthis still 2,> 0holds, directlygoto exit, and the dispatch is never triggered. All communication calls remain in the "batching" state, and the process hangs.

What is more insidious:ncclGroupDepthis thread_local and will not be reset when the function returns. Even if subsequent code no longer calls the group API, all communication on this thread will fail.

This change also breaks the pairing semantics ofncclGroupStart—ncclGroupStartincrements,ncclGroupEnddoes not decrement, so the depth only grows and eventually overflows (although int overflow requires 2 billion calls, in practice it is more likely to be a logical hang).

Q2: CudaPtrCheckInattr.type == cudaMemoryTypeDevice && attr.device != comm->cudaDev(📎 src/misc/argcheck.cc:20), if theattr.type == cudaMemoryTypeDevicecondition is removed, what problems arise? In what scenarios would it produce false positives?

Reference Answer:

cudaPointerAttributes.typehas three possible values:cudaMemoryTypeDevice(device memory),cudaMemoryTypeHost(host memory),cudaMemoryTypeManaged(unified memory).

If theattr.type == cudaMemoryTypeDevicecondition is removed, it becomes:

cpp
if (attr.device != comm->cudaDev) {  // 错误版本

Then for host memory or managed memory,attr.devicemay be -1 or 0, which does not matchcomm->cudaDev, causing a false "device mismatch" report.

Specific scenario: the user passes a pointer allocated bycudaMallocManaged. Theattr.deviceof managed memory is usually the device at allocation time, but if the memory is migrated to another device,attr.devicemay change. More commonly, for host memory (such ascudaHostAllocallocated pinned memory),attr.deviceis -1, which is unequal to anycudaDev, causing a false positive.

NCCL allows host memory as a communication buffer (relayed throughcudaMemcpy), so it is necessary to distinguish "device memory but wrong device" from "non-device memory." The former is an error; the latter is legal.

Q3: layoutInfoRecvTopkIdxKind(📎 contrib/nccl_ep/nccl_ep.cc:139-144) useslip->size < field_endto determine whether a field exists. If a new version inserts a field in the middle of the struct (rather than at the end), how does this check fail? Why does ABI design require that new fields can only be added at the end?

Reference Analysis:

Suppose the original struct is:

c
struct ncclEpLayoutInfo_t {
    unsigned int size;
    unsigned int magic;
    ncclEpExpertIdKind_t recv_topk_idx_kind;  // offset = 8
};

field_end = offsetof(recv_topk_idx_kind) + sizeof(...) = 8 + 4 = 12。

If the new version inserts a field betweenmagicandrecv_topk_idx_kind:

c
struct ncclEpLayoutInfo_t {
    unsigned int size;
    unsigned int magic;
    unsigned int new_field;                    // 新插入
    ncclEpExpertIdKind_t recv_topk_idx_kind;  // offset 变成 12
};

At this pointfield_end = 12 + 4 = 16. The old caller'ssizeis 12 (the old struct size),12 < 16holds, and the function returnsAUTO—but the old caller actually has therecv_topk_idx_kindfield, just at a different offset. This causes therecv_topk_idx_kindset by the old caller to be ignored.

Worse, if the old caller writesrecv_topk_idx_kindat the old offset (8), the new library reads at the new offset (12) and will read the value ofnew_field, causing complete corruption.

So the iron rule of ABI design is:New fields can only be added at the end of the struct. In this way, the old caller'ssizeis smaller than the new field'sfield_end, and the function correctly returns the default value; the new caller'ssizecovers the new field and reads it normally. Inserting fields in the middle breaks alloffsetof-based version checks.

This chapter analyzed four typical pitfalls in production environments and their internal defense mechanisms. These boundary conditions remind us that the stable operation of NCCL depends not only on the core implementation, but also on the adaptation and extension of the surrounding ecosystem. In the next chapter, we will turn to the ecosystem and extensions to see how peripheral projects such as nccl4py, nccl4rust, nccl_ep, and nccl_ubx bring NCCL's capabilities to a broader set of users.

CHAPTER 23

Chapter 23: Chapter 23: Ecosystem Extensions: Peripheral Projects Such as nccl4py, nccl4rust, nccl_ep, and nccl_ubx

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 23 / 25

Chapter 23: Ecosystem Extensions: Peripheral Projects Such as nccl4py, nccl4rust, nccl_ep, and nccl_ubx

In the previous chapter, we investigated typical NCCL failures in production environments—group semantics misuse, rank count mismatches, stream interactions, ABI version conflicts, and network timeouts. Most of these issues occur when using the C ABI directly, but modern large model training frameworks often do not call the C ABI directly. Instead, they reuse NCCL's capabilities through language bindings such as Python and Rust, or through extension projects targeting scenarios like MoE and ultra-high-bandwidth communication. These peripheral projects are placed under the bindings/ and contrib/ directories, positioned as experimental and community-maintained, and do not inherit the release quality assurance of the core library. This chapter examines nccl4py, nccl4rust, nccl_ep, nccl_ubx, and nccl_checkpoint one by one, looking at how they build a rich ecosystem outside the core through three paths: language bindings, device API extensions, and symbol interception.

nccl4py: Cython bindings and namespace package design

Intuitive model: translating the C ABI into something Python can understand

Imagine the NCCL core is a diplomat who only speaks C, and a Python training script is an intern who only speaks Python. nccl4py is that translator—it does not change what the diplomat says (NCCL's behavior), it only translates "ncclAllReduce(sendbuff, recvbuff, count, ...)" into "nccl.all_reduce(tensor)". Without this layer of translation, every Python framework would have to write its own ctypes bindings, which is repetitive work and error-prone.

Layered structure: Cython low-level + Python high-level

The design of nccl4py has two layers: the low level is Cython bindings (nccl/bindings/cynccl.pxd), and the high level is the Python API (nccl.core). The README explicitly states this layering.📎 bindings/nccl4py/README.md:4-4:

nccl4py provides low-level Cython bindings and a high-level Python API

The Cython bindings are distributed with the wheel as.pxdfiles, for other Cython extensions to directlycimport 📎 bindings/nccl4py/README.md:39-43:

cython
from nccl.bindings cimport cynccl
[Design inference and architectural trade-offs]

Why expose the Cython layer and not just the Python layer? Because some frameworks (such as DeepSpeed and Megatron) have their core loops in Cython, and going through the Python interpreter on every call is too expensive. Directlycimport cyncclallows Cython extensions to call NCCL functions with near-C zero overhead. This is a typical "layered exposure" design—the high level is for ordinary users, and the low level is for performance-sensitive scenarios.

Namespace package: multiple distributions share thencclprefix

This is the most ingenious design of nccl4py.ncclis a PEP 420 implicit namespace package📎 bindings/nccl4py/README.md:50-51:

nccl is a PEP 420 implicit namespace package. nccl4py provides nccl.bindings and nccl.core; other NCCL extension distributions can provide additional nccl.* subpackages.
[Design inference and architectural trade-offs]

In traditional Python packages,nccl/__init__.pywould "own" the entirencclnamespace. If both nccl4py and nccl_ep's Python bindings want to providenccl.xxx, they will conflict—whoever installs first wins. PEP 420 namespace packages solve this problem: without__init__.py, multiple distributions can each place subpackages into thenccl/directory, and the Python import system will merge them. So nccl4py providesnccl.bindingsandnccl.core, while nccl_ep providesnccl.ep, and the two can coexist📎 contrib/nccl_ep/README.md:80-82。

This design is crucial for ecosystem expansion: in the future, any third party that wants to addnccl.monitoring、nccl.profilingdoes not need to modify nccl4py's code.

CUDA version selection: the extra mechanism

During installation, usenccl4py[cu12]ornccl4py[cu13]to select the CUDA major version📎 bindings/nccl4py/README.md:13-17. The README explains the reason: extras will install the corresponding NCCL runtime and CUDA Python dependencies📎 bindings/nccl4py/README.md:19. Published wheels do not requireCUDA_HOMEor a local CUDA Toolkit, but building from source requires📎 bindings/nccl4py/README.md:20-21。

[Design inference and architectural trade-offs]

This is the standard approach in the Python ecosystem for handling CUDA version fragmentation. CUDA 12 and 13 are ABI-incompatible, so one wheel cannot cover both. Using extras lets pip choose the correct binary dependency according to the user's environment, avoiding version mismatches that are only discovered at runtime.

Production pitfalls

Pitfall 1: namespace packages and__init__.pyconflict.If some third-party package placesnccl/under__init__.py, the PEP 420 namespace package mechanism will be broken, causingnccl.coreimport failure. Troubleshooting method:python -c "import nccl; print(nccl.__path__)", if it reportsAttributeErrorit meansncclis not a namespace package.

Pitfall 2: Cython ABI version drift. cynccl.pxdis an experimental API📎 bindings/nccl4py/README.md:32-32, and when NCCL is upgraded.pxdmay change. Cython extensions that depend oncimport cyncclmust strictly match the nccl4py version, otherwise symbol resolution fails at compile time.

nccl4rust: RAII ownership and device-side boundaries

Intuitive model: let the compiler manage the lifecycle for you

In C, youncclCommInitRankget a communicator, and when you are done you mustncclCommDestroy. Forgetting to destroy it leaks; destroying it early crashes. Rust's RAII (Resource Acquisition Is Initialization) mechanism lets the compiler automatically call the destructor when a variable leaves scope—like a hotel room card, where the system automatically settles the bill when you check out, without you having to go to the front desk manually.

The core value of nccl4rust is applying this ownership semantics to NCCL's C ABI.

Layered structure: five crates each with their own role

The README's Layout table lists five crates📎 contrib/nccl4rust/README.md:20-28:

PathPurpose
crates/nccl-sysRaw host ABI generated by bindgen
crates/ncclRust-style host wrapper + RAII ownership
crates/nccl-device-sysno_stdCUDA-Oxide device declarations
crates/nccl-deviceTypedDevComm、Team、WindowWrapper
shim/Pure C-ABI shim, using only public headers
[Design inference and architectural trade-offs]

This split is deliberate. The README explains the motivation📎 contrib/nccl4rust/README.md:30-32: host applications can use onlyncclwithout needing the Rust GPU compiler; CUDA-Oxide kernels usenccl-device; consumers who need the raw ABI can choose the-syscrate. This "layer on demand" approach lets different users pay only the compilation cost they need.

Key design: pass device communicators by pointer rather than by value

This is the most instructive design decision in nccl4rust. The README's Host/device ownership boundary section📎 contrib/nccl4rust/README.md:211-219:

ncclDevCommCreate produces a versioned public structure in host memory. The host DeviceCommunicator wrapper owns that structure and destroys it before its parent communicator. CUDA-Oxide remains responsible for allocating device memory, copying those bytes, and keeping the copy alive while kernels execute. Kernels construct nccl_device::DevComm from a pointer to that device copy. Using a pointer rather than a by-value Rust mirror keeps the versioned C struct layout out of the kernel argument ABI.
[Design inference and architectural trade-offs]

Why not mirror C structs with Rust structs? BecausencclDevComm_tis versioned—different NCCL versions may have different fields. If kernel parameters pass a Rust mirror by value, then the kernel ABI is bound to a specific NCCL version's struct layout. Once NCCL upgrades the struct, all compiled kernels must be recompiled. Passing by pointer only passes an address, and the kernel accesses through the pointer, so layout changes do not affect the ABI. This is the same idea as thencclEpLayoutInfo_tsize-based ABI discussed in the previous chapter—isolating version differences behind a pointer。

Safety boundary: what is unsafe

The README's Current API contracts section lists six contracts📎 contrib/nccl4rust/README.md:230-249, with several key ones:

  • The raw-syscrate only mirrors the C ABI, without adding ownership or lifetime checks📎 contrib/nccl4rust/README.md:232-233
  • Current collective communication and point-to-point wrappers accept raw device pointers, declared asunsafe 📎 contrib/nccl4rust/README.md:42-45
  • Pointer translation methods return raw device pointers and cannot validate offset bounds, alignment, peer membership, aliasing, or window lifetime📎 contrib/nccl4rust/README.md:242-244
[Design inference and architectural trade-offs]

This is the fundamental difficulty of Rust bindings to NCCL: many of NCCL's API contracts are "the buffer must remain valid until the CUDA stream completes," but Rust's type system cannot express the asynchronous event of "stream completion." So these methods can only beunsafe, handing responsibility back to the caller. The README also points out the direction for improvement📎 contrib/nccl4rust/README.md:44-45: a stream-aware buffer abstraction could encode these requirements into a safe API. This is future work.

Device side: CUDA-Oxide and the LTOIR shim

The core challenge on the device side is that NCCL's device API is C++ templates, while Rust device code (CUDA-Oxide) needs a C ABI. The solution is a C++ shim📎 contrib/nccl4rust/README.md:26:

shim/ — CUDA C++ C-ABI shim built exclusively from public nccl.h and nccl_device.h

The shim is compiled into LTOIR (LLVM intermediate representation) and linked with Rust PTX into a cubin📎 contrib/nccl4rust/README.md:165-167. The README explains the build process📎 contrib/nccl4rust/README.md:158-163:

bash
make device \
  NCCL_INCLUDE_DIR="$NCCL_INCLUDE_DIR" \
  CUDA_HOME="$CUDA_HOME" \
  ARCH=90
[Design inference and architectural trade-offs]

LTOIR is NVIDIA's link-time optimization intermediate format. Using LTOIR rather than compiling directly to cubin is to allow the shim and Rust kernels to perform cross-language optimization at link time—for example, inlining shim functions into Rust kernels. This is the key technique for hybrid programming with "C++ templates + Rust kernels."

Production pitfalls

Pitfall one: the NCCL version must match exactly.The README explicitly requiresMatching NCCL 2.31 headers and runtime 📎 contrib/nccl4rust/README.md:80-81, because the prototype directly initializes fields that differ in earlier NCCL device API versions. If the headers andlibnccl.soversions are inconsistent, device communicator fields will be misaligned.

Pitfall two: CUDA graph and device communicators.The device communicator is a versioned structure in host memory; after being copied to the device, the kernel accesses it through a pointer. If CUDA graph capture bakes the device pointer into kernel parameters, recreating the communicator later will invalidate the pointer in the graph. This is the same root cause as the RDMA buffer reallocation problem in nccl_ep.

Pitfall three: safe initialization cannot be mixed with raw groups.The README warns📎 contrib/nccl4rust/README.md:238-239: safe initialization and managed calls that produce output cannot be mixed with rawnccl-sysgroup state, because the wrapper layer cannot observe raw group state. Mixing them will cause the wrapper layer's polling logic to conflict with raw group semantics.

nccl_ep: dispatch/combine primitives for expert parallelism

Intuitive model: MoE's "sorting center"

In MoE (Mixture of Experts) models, each token must be routed to top-k experts. The experts are distributed across different GPUs, so tokens need to be transferred across GPUs—this is dispatch. After the experts compute, the results must be sent back to the GPU where the original token resides—this is combine. nccl_ep is the communication engine for this "sorting center."

Without it, every MoE framework would have to implement the dispatch/combine communication logic itself, which is repetitive and difficult to optimize. nccl_ep turns it into a standard primitive in the NCCL ecosystem.

Two algorithms: LL and HT

The README explains two algorithms📎 contrib/nccl_ep/README.md:36-40:

  • Low-Latency (LL): small batch, latency-sensitive (LLM inference). Uses direct point-to-point all-to-all communication.
  • High-Throughput (HT): large-batch training and inference prefill. Uses hierarchical communication—intra-node NVLink aggregation, inter-node RDMA. Leverages Hopper's warp-specialized pipeline and TMA.
[Design inference and architectural trade-offs]

The split between these two algorithms reflects the different bottlenecks of MoE inference and training. During inference, the batch is small, and latency is the main contradiction, so LL uses direct point-to-point to avoid aggregation overhead. During training, the batch is large, and bandwidth is the main contradiction, so HT uses hierarchical aggregation to reduce cross-node traffic. This is a typical "choose the algorithm based on workload characteristics" design.

Core data structure: ncclEpGroupConfig_t

This is the EP configuration structure, with many fields📎 contrib/nccl_ep/README.md:339-362. Key fields:

  • sizeandversion: ABI version check, from the same origin as the size-based ABI discussed in the previous chapter📎 contrib/nccl_ep/README.md:340-341
  • algorithm: HT or LL📎 contrib/nccl_ep/README.md:342
  • max_dispatch_tokens_per_rank: maximum number of tokens dispatched by a single rank📎 contrib/nccl_ep/README.md:344
  • rdma_buffer_size: RDMA buffer size in LL mode📎 contrib/nccl_ep/README.md:356-356
  • alloc: custom device memory allocator📎 contrib/nccl_ep/README.md:359
[Design inference and architectural trade-offs]

rdma_buffer_size'sNCCL_EP_AUTOsemantics are worth digging into. The README explains📎 contrib/nccl_ep/README.md:396-406: in AUTO mode, the buffer is not allocated atncclEpCreateGroup, but instead at the firstncclEpInitHandleaccording to the actual(layout, num_topk). When a later handle needs a larger buffer, it will collectively reallocate. This "lazy allocation" design avoids requiring users to guess the buffer size, but introduces three constraints📎 contrib/nccl_ep/README.md:396-406:

1. All ranks must use the same(layout, num_topk)synchronous callncclEpInitHandle

2. Reallocation discards the old buffer contents,send_onlytemporarily stored data will be lost

3. CUDA graph capture bakes in the RDMA base address pointer, and after reallocation it must be recaptured

This is one of the most important production pitfalls in this chapter.Lazy allocation buys ease of use, but shifts the complexity of "when to reallocate" onto the user.

Tensor descriptors: static and dynamic forms

ncclEpTensor_tis a lightweight value type📎 contrib/nccl_ep/README.md:310-332. The README shows two usages:

Static descriptor(on the stack,NCCL_EP_TENSOR_INIT_INLINE)📎 contrib/nccl_ep/README.md:806-809:

c
ncclEpTensor_t expert_counters = { NCCL_EP_TENSOR_INIT_INLINE,
                                   .ndim = 1, .datatype = ncclInt32,
                                   .data = expert_counters_data,
                                   .sizes = expert_counters_dims };

Dynamic descriptor(on the heap,ncclEpTensorAlloc)📎 contrib/nccl_ep/README.md:793-798:

c
ncclEpTensor_t* topk_idx = nullptr;
{
    size_t dims[2] = { num_tokens, top_k };
    ncclEpTensorAlloc(&topk_idx, 2, ncclInt64, dims, /*config=*/NULL);
    cudaMalloc(&topk_idx->data, num_tokens * top_k * sizeof(int64_t));
}
[Design inference and architectural trade-offs]

The difference between the two forms lies insizesownership of the array. The static descriptor'ssizesis a caller-owned stack array, and must outlive the descriptor📎 contrib/nccl_ep/README.md:325-326. The dynamic descriptor'ssizesis a library-owned heap copy, freed byncclEpTensorDestroy. The public struct holds the📎 contrib/nccl_ep/README.md:514-514pointer, so the two forms can be mixed in the same callncclEpTensor_t*. This design gives zero heap allocation for simple scenarios and library-managed convenience for complex scenarios.📎 contrib/nccl_ep/README.md:514-514Execution modes: synchronous and staged

The README's Execution Modes section

explains two modes:📎 contrib/nccl_ep/README.md:701-741Synchronous mode

(default): occupies GPU resources for the entire operation, including the time spent waiting to receive dataStaged mode📎 contrib/nccl_ep/README.md:705-709。

(LL only): the operation is split into send and receive phases. Initiated with📎 contrib/nccl_ep/README.md:718-726, GPU resources are released after data transfer starts, the application can use these resources for computation, and finallysend_only = 1is used to completencclEpCompletecopy📎 contrib/nccl_ep/README.md:728-741。

mermaid
sequenceDiagram
    participant App as 应用线程
    participant EP as ncclEpDispatch
    participant GPU as GPU 内核
    participant Net as RDMA 网卡
    App->>EP: ncclEpDispatch(send_only=1)
    EP->>GPU: 启动发送内核
    GPU->>Net: GIN put/signal 发起传输
    EP-->>App: 立即返回,释放 SM
    Note over App: 应用用释放的 SM 做计算
    App->>EP: ncclEpComplete()
    EP->>GPU: 启动接收内核
    GPU->>Net: 等待数据到达
    Net-->>GPU: 数据写入
    GPU-->>EP: 完成
    EP-->>App: 返回,数据就绪

returns immediately after initiation, SM resources are released for computation, and after the application finishes other work it callssend_onlyto wait for receive completion. This is the classic "compute-communication overlap" pattern.ncclEpCompleteProduction pitfalls

Pitfall one:

conditional collectivity ofncclEpInitHandle. In AUTO mode,is a conditional collective callncclEpInitHandle. If some rank triggers reallocation due to a different layout, the other ranks must participate synchronously. Lack of synchronization will cause deadlock or data corruption.📎 contrib/nccl_ep/README.md:396-406Pitfall two: prohibited during CUDA graph capture

The README explicitly warnsncclEpInitHandle。: in AUTO mode, you must not call📎 contrib/nccl_ep/README.md:396-406betweencudaStreamBeginCaptureandcudaStreamEndCapture. Because reallocation changes the RDMA base address, while graph capture has already baked in the old pointer.ncclEpInitHandlePitfall three: guard overhead.

The README mentions: EP adds a guard to internal communication buffers by default to prevent adjacent dispatch/combine calls from corrupting each other's data. Advanced users who have already ensured that consecutive operations will not contend can use📎 contrib/nccl_ep/README.md:299-303to disable it and reclaim the overhead. But disabling it incorrectly will cause silent data corruption.NCCL_EP_DISABLE_GUARD=1nccl_ubx: fused collective communication and symmetric allocator

Intuitive model: also hand over the "packing and unpacking before and after moving" to the moving company

直觉模型:把「搬家前后的打包拆包」也交给搬家公司

Ordinary collective communication is only responsible for moving data. But in real models, residual addition often needs to be done before AllReduce, and RMSNorm after it. If these operations are done separately, the data has to make extra trips through GPU memory. The idea of nccl_ubx is: fuse residual addition, RMSNorm, and mxfp8 quantization into the collective communication kernel📎 contrib/nccl_ubx/README.md:6-9. It's like a moving company not only moving boxes, but also helping you pack and unpack, all in one trip.

Hardware prerequisite: NVLink multicast is required

The README explicitly requires SM 9.0+ (Hopper/Blackwell), and the MC kernel path requires NVLink multicast hardware📎 contrib/nccl_ubx/README.md:24-24. SM 8.0 (A100) is not supported, because Ampere does not have NVLink multicast hardware,multimem.*and inline PTX cannot be assembled for arch 8.0📎 contrib/nccl_ubx/README.md:24-24。

[Design inference and architectural trade-offs]

This explains why ubx is "experimental" — it depends on the NVLink multicast capability introduced only with Hopper.multimem.*The instruction allows one GPU to write data to the symmetric addresses of multiple GPUs with a single instruction, which is the hardware-accelerated foundation of collective communication. Without this hardware, the core optimization of ubx does not hold.

Symmetric allocator: turning PyTorch tensors into NCCL windows

The core of ubx is a custom symmetric allocator📎 contrib/nccl_ubx/README.md:11-14:

A central piece of the design is a custom symmetric allocator that provides zero-copy collective input/output buffers while remaining easy to plug into existing PyTorch code: tensors are ordinary torch.Tensor instances backed by an NCCL-managed symmetric window.
[Design inference and architectural trade-offs]

This is the most ingenious part of ubx. NCCL's symmetric memory requires all ranks to access the buffer using the same set of virtual addresses (as discussed in Chapter 14). But PyTorch users are used to usingtorch.Tensor. ubx makestorch.Tensorthe underlying storage directly be an NCCL symmetric window, so user code does not need to change, but collective communication can be zero-copy — the input and output buffers are the symmetric memory itself, with no extra copy needed.

Collective communication variants and automatic selection

The Available collectives table in the README📎 contrib/nccl_ubx/README.md:90-90:

OpVariantsAuto-select
AllReducemc, uc, lamport, autoLamport ≤ 0.25 MB, else MC
AllToAlluc, lamport, autoLamport ≤ 0.25 MB, else UC
AllGathermc—
[Design inference and architectural trade-offs]

The differences among the three variants:mcuses NVLink multicast hardware,ucuses ordinary unicast,lamportis a low-latency algorithm. Automatic selection uses a 0.25 MB threshold — small messages use Lamport low latency, large messages use MC/UC high bandwidth. This threshold is similar to the tuning logic in the NCCL core, but ubx simplifies it to a fixed threshold.

Fused operations: residual + RMSNorm

The README mentions📎 contrib/nccl_ubx/README.md:103-103:

SymmAllocator.allreduce_mc() and allreduce_lamport() accept optional gamma/residual_in parameters to fuse residual addition + RMSNorm into the same kernel.
[Design inference and architectural trade-offs]

This is the core selling point of ubx. The traditional flow is: AllReduce → residual addition → RMSNorm, with three rounds of GPU memory reads and writes. After fusion, it is completed in one kernel, saving 2/3 of GPU memory bandwidth. For bandwidth-constrained large-model training, this is a real speedup.

MoE token dispatch + mxfp8 quantization

The README describesa2av_token_bf16_mxfp8 📎 contrib/nccl_ubx/README.md:103-103:

a single GPU kernel that routes bf16 tokens to remote ranks while quantizing them to mxfp8 (E8M0 scale per 32 elements) on the fly.
[Design inference and architectural trade-offs]

This kernel fuses "routing + quantization." bf16 is 16-bit, mxfp8 is 8-bit, and after quantization the data volume is halved, so the cross-node transmission bandwidth requirement is halved. Quantizing before transmission is better than quantizing after transmission — what is saved is network bandwidth rather than GPU memory bandwidth. This is a key optimization for MoE inference.

Production pitfalls

Pitfall one:TORCH_CUDA_ARCH_LISTmust include theasuffix.The README emphasizes📎 contrib/nccl_ubx/README.md:47-56: use theasuffix to ensure access to the completemultimem.*instruction set. Some acceleration-specific variants are unavailable on ordinary9.0/10.0, and future kernels using these variants will silently degrade performance or fail to assemble.

Pitfall two:UBX_BUILD_TIMEOUTruntime overhead.The README states📎 contrib/nccl_ubx/README.md:47-56: setting it to 1 compiles a spinloop timeout into the kernel side, increasing runtime overhead (extraclock64()checks andprintfon timeout). Only enable it when troubleshooting hangs.

Pitfall three:NCCL_NVLS_ENABLE=0degradation.The README lists this environment variable📎 contrib/nccl_ubx/README.md:202: setting it to 0 allows running without NVLink multicast. But the MC kernel path becomes unavailable, leaving only the UC/Lamport variants, and performance drops significantly.

nccl_checkpoint: LD_PRELOAD interception and state replay

Intuitive model: taking a snapshot of the communication domain

A training job runs for several hours, and suddenly it needs to be migrated to another machine, or its state needs to be saved for recovery. Ordinary checkpoints only save model weights and optimizer state, but the state of the NCCL communication domain (rank IDs, connections, buffers) cannot be serialized directly. The idea of nccl_checkpoint is: intercept all NCCL calls, record the initialization steps, and replay these steps during recovery📎 contrib/nccl_checkpoint/README.md:3-7。

It's like recording every step of assembling furniture, then reassembling it according to the recording after moving, instead of trying to move the assembled furniture as a whole.

Core mechanism: LD_PRELOAD symbol interception

The Design section of the README📎 contrib/nccl_checkpoint/README.md:17-20:

The application is launched with LD_PRELOAD=/path/to/libnccl-checkpoint-shim.so in the environment. This allows the library to intercept all calls to NCCL functions to capture all resource initialization steps.
[Design inference and architectural trade-offs]

LD_PRELOADis a mechanism of the Linux dynamic linker: load the specified.sobefore the application normally loads shared libraries. If this.sodefines symbols with the same names as NCCL (such asncclCommInitRank), the dynamic linker will preferentially use the version in.so. In this way, the shim can intercept all NCCL calls, record parameters, and then replay them during recovery.

Checkpoint flow

The Python example in the README📎 contrib/nccl_checkpoint/README.md:44-58shows the complete workflow:

python
nccl_checkpoint.checkpoint_prepare()
drv.cuCheckpointProcessLock(os.getpid(), None)
drv.cuCheckpointProcessCheckpoint(os.getpid(), None)
# CRIU dump happens here.
drv.cuCheckpointProcessRestore(os.getpid(), None)
drv.cuCheckpointProcessUnlock(os.getpid(), None)
nccl_checkpoint.checkpoint_restore()
[Design Inference and Architectural Trade-offs]

The workflow has four steps:

1. checkpoint_prepare(): Destroy all communicators so that CUDA Checkpoint and CRIU can safely dump process state📎 contrib/nccl_checkpoint/README.md:25-27

2. cuCheckpointProcessLock/Checkpoint: CUDA driver locks the process and performs checkpointing

3. CRIU dump: External tools dump process memory and file descriptors to disk

4. cuCheckpointProcessRestore/Unlock + checkpoint_restore(): Restore the process and replay NCCL configuration📎 contrib/nccl_checkpoint/README.md:29-31

Redis KVS: Cross-machine rendezvous

The README explains why Redis is needed📎 contrib/nccl_checkpoint/README.md:33-38:

Because it is useful to restore on different hardware, IP addresses may have changed. There is no convenient way to directly inform the NCCL Checkpoint library of all peer addresses during the restore process, so the library depends on a temporary Redis Key-Value store to be made available.
[Design Inference and Architectural Trade-offs]

During recovery, the machine may change and the IP may change. Rebuilding the NCCL communication domain requires knowing the new addresses of all peers. But the shim cannot directly know these addresses, so a Redis KVS is used for rendezvous—all processes write their new addresses to the KVS and read other processes' addresses from the KVS. This is like after moving, everyone agrees to exchange new addresses on a public message board.

The README states that Redis is only needed during the recovery bootstrap phase📎 contrib/nccl_checkpoint/README.md:221-221,checkpoint_restore()After returning, it can be stopped.

Limitations: Three unsupported cases

The Limitations section of the README📎 contrib/nccl_checkpoint/README.md:119-129lists three limitations:

1. ncclWinGetUserPtr()The returned pointer is invalid after recovery📎 contrib/nccl_checkpoint/README.md:125-126

2. CUDA graph capture is not supported📎 contrib/nccl_checkpoint/README.md:136-136

3. Device API is not supported—ncclDevCommobjects and device-visiblencclWindow_tvalues cannot be restored📎 contrib/nccl_checkpoint/README.md:136-136

[Design Inference and Architectural Trade-offs]

The third limitation is the most serious. The device API is a new direction for NCCL (DevComm discussed in Chapter 19), but checkpointing does not support it. This means applications using the device API (such as nccl_ep, nccl_ubx) cannot be recovered with checkpointing. This reflects ecosystem fragmentation—new features move fast, but reliability tools cannot keep up.

Production Pitfalls

Pitfall One:NCCL_CHECKPOINT_KVS_PATHSet before checkpointing; cannot be changed during recovery.The README warns📎 contrib/nccl_checkpoint/README.md:221-221: This environment variable is not used during the checkpoint preparation phase, but it will be captured into the checkpoint and cannot be easily modified during recovery. Therefore, it must be set before checkpointing, and the Redis address in the recovery environment must match.

Pitfall Two:NCCL_CHECKPOINT_KVS_TIMEOUTOnly covers the shim's Redis rendezvous.The README states📎 contrib/nccl_checkpoint/README.md:221-221: Default is 300 seconds. Once communicator replay enters the NCCL transport establishment phase, the underlying NCCL transport calls use their own behavior and may require transport-specific diagnostics. In other words, the timeout only protects the Redis phase; if the transport establishment phase hangs, you need toNCCL_DEBUGtroubleshoot.

Pitfall Three: NCCL version must match.The README requires NCCL 2.31.0 or newer📎 contrib/nccl_checkpoint/README.md:158, and recommendsNCCL_SRCthat the NCCL version in the path exactly matches the runtime NCCL library version📎 contrib/nccl_checkpoint/README.md:156-158. Version mismatch will cause struct layout misalignment during replay.

Design Reflection: Three Modes of Ecosystem Extension

Reviewing these five projects, we can summarize three modes of NCCL ecosystem extension:

Mode One: Language bindings (nccl4py, nccl4rust).The core challenge is ownership and lifecycle. C's ABI has no ownership semantics, so the binding layer must add them itself. nccl4py uses Cython layering, nccl4rust uses RAII +unsafeboundaries. The common point is:isolating version differences behind pointers—nccl4rust passes DevComm via pointers, nccl4py isolates versions with namespace packages.

Mode Two: Device API extensions (nccl_ep, nccl_ubx).The core challenge is ABI version management and resource lifecycle. nccl_ep uses size-based ABI (detailed in the previous chapter), nccl_ubx uses a symmetric allocator. The common point is:lazy allocation + collective reallocation—nccl_ep's RDMA buffer and nccl_ubx's symmetric pool are both allocated on demand, but reallocation requires synchronization across all ranks.

Mode Three: Symbol interception (nccl_checkpoint).The core challenge is state capture and replay. UsingLD_PRELOADto intercept all NCCL calls, record initialization steps, and replay during recovery. This mode does not modify the NCCL core, but can transparently add checkpointing capability to existing applications.

[Design Inference and Architectural Trade-offs]

The common constraint of the three modes isNCCL version compatibility. All projects require an exact matching NCCL version because NCCL's ABI is evolving. This reflects a fundamental tension in the NCCL ecosystem: the core iterates rapidly, but surrounding projects need stability. Size-based ABI, pointer passing, and namespace packages are all technical means to alleviate this tension.

mermaid
flowchart TD
    start["用户想扩展 NCCL"] --> q1{"扩展什么?"}
    q1 -->|"语言互操作"| lang["语言绑定"]
    q1 -->|"新通信模式"| dev["设备 API 扩展"]
    q1 -->|"可靠性"| ckpt["符号拦截"]
    lang --> q2{"性能敏感?"}
    q2 -->|"是"| cython["Cython 底层 + Python 高层<br/>nccl4py"]
    q2 -->|"否"| raii["RAII 包装<br/>nccl4rust"]
    dev --> q3{"需要 MoE?"}
    q3 -->|"是"| ep["dispatch/combine<br/>nccl_ep"]
    q3 -->|"否"| ubx["融合集合通信<br/>nccl_ubx"]
    ckpt --> preload["LD_PRELOAD 拦截<br/>nccl_checkpoint"]
    cython --> abi{"ABI 版本管理"}
    raii --> abi
    ep --> abi
    ubx --> abi
    preload --> abi
    abi -->|"指针传递"| safe["版本差异隔离"]
    abi -->|"size-based"| safe
    abi -->|"命名空间包"| safe

This decision diagram shows the path for choosing NCCL extensions. Whichever path you take, you ultimately face the core issue of ABI version management, and the three technical means (pointer passing, size-based ABI, namespace packages) all isolate version differences behind stable interfaces.

Chapter Summary

This chapter analyzed five peripheral projects in the NCCL ecosystem:

  • nccl4pyUse Cython layering + PEP 420 namespace packages to let the Python ecosystem extend with zero conflictsnccl.*subpackages.
  • nccl4rustUse RAII ownership + pointer-passed device communicators to isolate versioned C struct layouts outside the kernel ABI.
  • nccl_epUse LL/HT dual algorithms + lazy RDMA buffer allocation to provide dispatch/combine primitives for MoE, but this introduces constraints of conditional collective calls and CUDA graph invalidation.
  • nccl_ubxUse symmetric allocators + kernel fusion to fold residual addition, RMSNorm, and mxfp8 quantization into collective communication kernels, but this depends on Hopper+ NVLink multicast hardware.
  • nccl_checkpointUseLD_PRELOADsymbol interception + Redis rendezvous to implement cross-machine communication domain checkpointing, but it does not support device APIs and CUDA graphs.

Chapter Review and Self-Test

Q1: In nccl_ep'srdma_buffer_size = NCCL_EP_AUTOmode, if rank 0 first callsncclEpInitHandleand triggers buffer reallocation, while rank 1 does not trigger reallocation because its layout is different, what happens? Please analyze in combination with📎 contrib/nccl_ep/README.md:396-406's constraints.

Reference Analysis: The README explicitly states📎 contrib/nccl_ep/README.md:396-406:All ranks must call ncclEpInitHandle in lockstep with the same (layout, num_topk). In AUTO mode,ncclEpInitHandleis a conditional collective call—whether reallocation is triggered depends on whether that handle's(layout, num_topk)requires more space than the current buffer.

If rank 0's layout requires a larger buffer and triggers reallocation, while rank 1's layout does not, then rank 0 will execute the collective operation sequence "deregister window → free → ncclMemAlloc → register"📎 contrib/nccl_ep/README.md:396-406, while rank 1 will not. This causes two problems:

1. Collective operation mismatch: NCCL's window deregister/register are collective operations and require all ranks to participate. Rank 0 unilaterally executing them will cause rank 1 to reference the old window handle in subsequent communication, while rank 0 has already switched to a new window, resulting in communication failure or data corruption.

2. Base address inconsistency: After reallocation, rank 0's RDMA base address changes, while rank 1's does not. Although the README says "recorded layout offsets on every live handle are pure offsets relative to the group's rdma_buffer and resolve correctly against the new base"📎 contrib/nccl_ep/README.md:396-406, this only holds under the premise that all ranks reallocate. Rank 1's base address has not changed, while rank 0's has, so cross-rank address resolution will be misaligned.

The correct approach is: all ranks use the same(layout, num_topk)to synchronously callncclEpInitHandle, ensuring consistent reallocation decisions. If this cannot be guaranteed, explicitrdma_buffer_size > 0mode should be used, allocating a sufficiently large buffer atncclEpCreateGrouptime to avoid runtime reallocation📎 contrib/nccl_ep/README.md:396-406。

Q2: Why does nccl4rust passncclDevComm_tto device kernels by pointer rather than by value? If changed to pass-by-value, what happens after NCCL upgrades the struct layout? Please analyze in combination with📎 contrib/nccl4rust/README.md:211-219.

Reference Analysis: The README explicitly states📎 contrib/nccl4rust/README.md:217-219:Kernels construct nccl_device::DevComm from a pointer to that device copy. Using a pointer rather than a by-value Rust mirror keeps the versioned C struct layout out of the kernel argument ABI.

ncclDevComm_tis a versioned public struct, and fields may differ across NCCL versions. If passed by value:

1. Kernel ABI binds to struct layout: When kernel parameters are passed by value, the compiler bakes the entire struct's byte layout into the kernel's calling convention. After NCCL upgrades the struct (adding fields, changing field order, changing alignment), already-compiled kernels still parse parameters according to the old layout, causing field misalignment.

2. All kernels must be recompiled: Every NCCL upgrade requires recompiling all kernels that use device communicators. For training jobs deployed on a large number of machines, this is a huge operational burden.

3. Cross-version incompatibility: If the host side creates a communicator with the new NCCL, while the device-side kernel is compiled with the old NCCL, pass-by-value will cause the kernel to read incorrect fields.

Passing by pointer only passes an 8-byte address, and the kernel accesses the struct through the pointer. When NCCL upgrades the struct layout, as long as the host side creates the communicator with the new version and copies it to the device, the kernel accesses the new layout through the pointer. The kernel itself does not need to be recompiled, because its parameter is just an address. This isolates version differences behind the pointer—the pointer is stable, while the content pointed to by the pointer can change。

This is the same design philosophy as nccl_ep's size-based ABI: use a layer of indirection to isolate volatile version details behind a stable interface.

Q3: nccl_checkpoint usesLD_PRELOADto intercept NCCL calls, but if an application links both nccl4py and nccl_checkpoint, nccl4py's Cython bindings directly calllibnccl.so's symbols,LD_PRELOADcan it intercept them? Please analyze the symbol resolution order.

Reference Analysis: This depends on the symbol resolution order.LD_PRELOADThe mechanism is: before loading the shared libraries that the application normally depends on, the dynamic linker first loads theLD_PRELOADspecified.so. When the application (or a library it depends on) references a symbol, the dynamic linker searches in "first loaded, first resolved" order—LD_PRELOAD's.sotakes precedence overlibnccl.so。

So in theory, when nccl4py's Cython binding callsncclCommInitRank, the dynamic linker will first find the symbol with the same name inlibnccl-checkpoint-shim.so, and the interception succeeds.

But there are several edge cases:

1. Directdlopen + dlsym: If nccl4py usesdlopen("libnccl.so")and thendlsymto get the function pointer,LD_PRELOADcannot intercept it, becausedlsymdirectly searches for the symbol in the specified.so, without going through the global symbol table. The README mentions that C applications usedlsymto resolvencclCheckpointPrepare 📎 contrib/nccl_checkpoint/README.md:109-109, but that is resolving the checkpoint's own symbols, not NCCL symbols.

2. Symbol binding timing: If nccl4py binds NCCL symbols beforeLD_PRELOADtakes effect (for example, in__attribute__((constructor))), interception may fail. But under normal circumstancesLD_PRELOADtakes effect at process startup, earlier than any user code.

3. RTLD_DEEPBIND: If nccl4py usesdlopenand specifiesRTLD_DEEPBIND, symbol lookup will preferentially resolve insidelibnccl.so, bypassingLD_PRELOAD. This is a common pitfall.

4. Static linking: If nccl4py statically links NCCL,LD_PRELOADis completely ineffective, because the symbols have already been resolved at compile time.

So the conclusion is:In normal dynamic linking scenariosLD_PRELOADcan intercept nccl4py's calls, but if nccl4py usesdlopen + RTLD_DEEPBINDor static linking, interception will fail. In production use, you should useLD_DEBUG=bindingsto verify symbol binding and confirm that NCCL calls are intercepted by the shim.

In the next chapter we will turn to architectural evolution and future directions, and look at how NCCL evolves from a collective communication library into a programmable communication engine.

These surrounding projects demonstrate, through language bindings, device API extensions, and symbol interception, how NCCL's core capabilities can be reused in different scenarios. The core constraint running through all the projects is NCCL ABI version compatibility—size-based ABI, pointer passing, and namespace packages are all technical means of isolating version differences behind a stable interface. Understanding these means is the prerequisite for safely using these surrounding projects. As these extension projects keep probing the boundaries of the core, NCCL itself is also quietly evolving: from fixed collective operations to a programmable communication engine, from host proxy to GPU direct issue, from registered buffers to symmetric memory. In the next chapter, based on the traces of evolution in the source code, we will discuss how these changes will reshape the communication methods of upper-layer frameworks.

CHAPTER 24

Chapter 24: Chapter 24: Architectural Evolution and Future Directions: From Static Communication to Programmable Communication

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 24 / 25

Chapter 24: Architectural Evolution and Future Directions: From Static Communication to Programmable Communication

In the previous chapter we saw how the community builds a surrounding ecosystem around the NCCL core: Python bindings, Rust bindings, expert-parallel communication, ultra-bandwidth primitives, and communication checkpoints. These projects all reuse NCCL's stable API, but their demands have already gone beyond the scope of traditional collective communication—expert parallelism requires fine-grained point-to-point send/receive, checkpoints require pausing/resuming communication state, and ultra-bandwidth primitives require bypassing standard collective operations to directly operate the network. These demands point to the same problem: NCCL's fixed collective operation model is being stretched to the breaking point by more flexible communication needs. In this chapter we will no longer look at a single module, but instead start from the traces of evolution that have already appeared in the source code and discuss where NCCL is heading. Specifically, we will analyze three intertwined forces of evolution: communication primitives moving from fixed collectives to programmable—the RMA task scheduling in src/rma/rma.cc allows upper layers to compose Put/Signal/WaitSignal primitives instead of only calling AllReduce; network initiation moving from host proxy to GPU direct issue—the GIN backend management in src/gin/gin_host.cc allows GPU kernels to directly drive the NIC; and the memory model moving from registered buffers to symmetric memory—the symmetric memory kernel selection in src/sym_kernels.cc allows all ranks to use the same set of virtual addresses to access each other's buffers. These three forces are not isolated; they share the same infrastructure: the team abstraction in src/nccl_device/core.cc and the versioned DevComm in src/devcomm/devcomm_v23100.cc. Understanding how they mesh together means understanding the evolution logic of NCCL from a "collective communication library" to a "programmable communication engine."

1. Programmable Communication Primitives: How RMA Turns a "Fixed Recipe" into a "Buffet"

Intuitive model

Traditional NCCL collective communication is like a fixed set meal: you order AllReduce, and the kitchen just follows the AllReduce procedure to completion. But in expert parallelism (MoE) scenarios, each token needs to be sent to a different expert, and the sending pattern is completely unknown at compile time—this is like a buffet, where you have to decide what to take, how much to take, and when to take it.

RMA is the "buffet counter" that NCCL provides to upper layers: Put (write data to the peer's memory), Signal (notify the peer), WaitSignal (wait for the peer's signal). Upper-layer frameworks can freely combine these three primitives to implement arbitrary communication patterns.

Without RMA, MoE's all-to-all can only be simulated through multiple small-scale collective operations, each requiring the full kernel launch and synchronization process, resulting in unacceptably high latency.

Data Structures and Memory Layout

The core data structures of RMA arencclTaskRma(task description) andncclRmaArgs(plan parameters). Let's first look atncclRmaArgs's fields, which are initialized inscheduleRmaTasksToPlan.

📎 src/rma/rma.cc:166-171

cpp
plan->isRma = true;
plan->rmaArgs = ncclMemoryStackAlloc<struct ncclRmaArgs>(&comm->memScoped);
plan->rmaArgs->func = firstTask->func;
plan->rmaArgs->nRmaTasks = 0;
plan->rmaArgs->nRmaTasksProxy = 0;
plan->rmaArgs->nRmaTasksCe = 0;

The key fields here arenRmaTasksProxyandnRmaTasksCe. They split RMA tasks into two execution paths:

  • CE path(Copy Engine): The target rank is within the LSA (Local Symmetric Access) range, which can be completed directly using the GPU's copy engine without needing the network.
  • Proxy path: The target rank is not within the LSA range and must go through a host proxy thread to drive the network.
[Design Inference and Architectural Trade-offs]

The motivation behind this dichotomy is straightforward: communication within the LSA range goes over NVLink or PCIe, which has high bandwidth and low latency, making asynchronous copy with CE the most cost-effective; cross-machine communication must go through the NIC and can only be driven by proxy threads. Only by scheduling the two types of tasks separately can CE and proxy execute in parallel, rather than waiting serially.

ncclTaskRmaitself containspeers、nsignals、signalIdxsthree array pointers, recording the peer rank, signal count, and signal index respectively. For WaitSignal tasks, one task can wait for multiple peers; for Put/Signal tasks, one task targets only one peer.

Step-by-Step Walkthrough: Scheduling of a WaitSignal

Let's plug in a concrete scenario: rank 0 callsncclWaitSignal, waiting for signals from rank 1 and rank 3. Assume rank 1 is within the LSA range and rank 3 is not.

Step 1: Find the first non-empty context queue.

📎 src/rma/rma.cc:148-158

cpp
int ctx = -1;
for (int i = 0; i < comm->config.numRmaCtx; i++) {
  if (!ncclIntruQueueEmpty(&planner->rmaTaskQueues[i])) {
    ctx = i;
    break;
  }
}
if (ctx == -1) return ncclSuccess;

RMA tasks are queued by context, and each context is an independent RMA channel. Here, find the first context that has tasks and take out its queue.

Step 2: Take out the first task and determine its type.

📎 src/rma/rma.cc:163-168

cpp
struct ncclTaskRma* firstTask = ncclIntruQueueDequeue(ctxQueue);
plan->isRma = true;
plan->rmaArgs = ncclMemoryStackAlloc<struct ncclRmaArgs>(&comm->memScoped);
plan->rmaArgs->func = firstTask->func;

firstTask->funcisncclFuncWaitSignal, enter the WaitSignal branch.

Step 3: Split peers by LSA reachability.

📎 src/rma/rma.cc:187-204

cpp
for (int i = 0; i < firstTask->npeers; i++) {
  int peerRank = firstTask->peers[i];
  bool lsaAccessible = isLsaAccessible(comm, peerRank);
  if (lsaAccessible) {
    peersCe[npeersCe] = peerRank;
    nsignalsCe[npeersCe] = firstTask->nsignals[i];
    signalIdxsCe[npeersCe] = firstTask->signalIdxs[i];
    npeersCe++;
  } else {
    peersProxy[npeersProxy] = peerRank;
    nsignalsProxy[npeersProxy] = firstTask->nsignals[i];
    signalIdxsProxy[npeersProxy] = firstTask->signalIdxs[i];
    npeersProxy++;
  }
}

isLsaAccessibleiterates overcomm->devrState.lsaRankList, determining whether the peer is within the LSA team. Rank 1 is within LSA, so it goes into the CE list; rank 3 is not, so it goes into the Proxy list.

Step 4: Create a new task for each of CE and Proxy.

📎 src/rma/rma.cc:206-246

cpp
if (npeersCe > 0) {
  struct ncclTaskRma* waitSignalTaskCe = ...;
  waitSignalTaskCe->peers = peersCe;
  waitSignalTaskCe->npeers = npeersCe;
  ncclIntruQueueEnqueue(&plan->rmaTaskQueueCe, waitSignalTaskCe);
  plan->rmaArgs->nRmaTasksCe = 1;
}
if (npeersProxy > 0) {
  struct ncclTaskRma* waitSignalTaskProxy = ...;
  waitSignalTaskProxy->peers = peersProxy;
  waitSignalTaskProxy->npeers = npeersProxy;
  ncclIntruQueueEnqueue(&plan->rmaTaskQueueProxy, waitSignalTaskProxy);
  plan->rmaArgs->nRmaTasksProxy = 1;
}

The original single WaitSignal task is split into two: the CE task waits for rank 1, and the Proxy task waits for rank 3. The two tasks can execute in parallel—the CE path waits on the GPU, and the Proxy path waits on the host thread.

Step 5: Release the original task.

📎 src/rma/rma.cc:249-251

cpp
planner->nTasksRma -= 1;
ncclMemoryPoolFree(&comm->memPool_ncclTaskRma, firstTask);

The original task has already been split into two new tasks, so it is released back to the memory pool.

Concurrency Control and Hardware Interaction

The parallel execution of RMA is reflected inncclRmaWaitSignal.

📎 src/rma/rma.cc:43-74

cpp
if (plan->rmaArgs->nRmaTasksProxy > 0 && plan->rmaArgs->nRmaTasksCe > 0) {
  cudaStream_t ceStream = comm->rmaState.rmaCeState.ceStream;
  cudaEvent_t ceEvent = comm->rmaState.rmaCeState.ceEvent;
  CUDACHECKGOTO(cudaEventRecord(ceEvent, stream), ret, fail);
  CUDACHECKGOTO(cudaStreamWaitEvent(ceStream, ceEvent, 0), ret, fail);
  NCCLCHECKGOTO(ncclRmaProxyWaitLaunch(comm, plan, stream), ret, fail);
  NCCLCHECKGOTO(ncclRmaCeWaitLaunch(comm, plan, ceStream), ret, fail);
  CUDACHECKGOTO(cudaEventRecord(ceEvent, ceStream), ret, fail);
  CUDACHECKGOTO(cudaStreamWaitEvent(stream, ceEvent, 0), ret, fail);
}

This code uses CUDA events for inter-stream synchronization: first record an event on the input stream, let the CE stream wait for this event, then launch the proxy and CE tasks on the two streams respectively, and finally let the input stream wait for the CE stream's event. In this way, the two paths advance in parallel, but externally it appears as a single synchronous operation.

[Design Inference and Architectural Trade-offs]

The design trade-off here is: parallel execution can reduce latency, but it introduces additional event recording and stream synchronization overhead. For small messages, this overhead may exceed the parallel benefit; for large messages, the parallel benefit is significant. NCCL does not make an adaptive judgment here, but uniformly takes the parallel path—because the typical scenario for RMA is fine-grained communication of large messages.

Production Pitfall Avoidance Guide

Pitfall 1: Incorrect LSA reachability judgment causes tasks to take the wrong path. isLsaAccessibleiterates overlsaRankList, iflsaSizeis 0 (for example, a single-rank communication domain), all peers will be judged unreachable and all will take the Proxy path. This will not be exposed during small-scale testing, but it will cause a sharp performance drop in large-scale deployments. The troubleshooting method is to look atscheduleRmaTasksToPlan's INFO logs for the ratio ofnRmaTasksProxyandnRmaTasksCe.

Pitfall 2: The lifetime of the peer array after a WaitSignal task is split.The CE path'speersCeusesncclMemoryStackAllocallocation, and its lifetime followscomm->memScoped; the Proxy path'speersProxyusesncclCallocallocation, and after the task finishes executing it needs to be manuallyfree. If Proxy task creation fails,failbranch will release these arrays.

📎 src/rma/rma.cc:302-308

cpp
exit:
  return ret;
fail:
  free(peersProxy);
  free(nsignalsProxy);
  free(signalIdxsProxy);
  goto exit;

Pitfall 3: Cross-context batching of Put/Signal tasks.In the Put/Signal branch, NCCL pulls put/signal tasks from all contexts into the same plan, but stops when it encounters a WaitSignal.

📎 src/rma/rma.cc:279-295

cpp
for (int c = 0; c < comm->config.numRmaCtx; c++) {
  struct ncclIntruQueue<struct ncclTaskRma, &ncclTaskRma::next>* q = &planner->rmaTaskQueues[c];
  while (!ncclIntruQueueEmpty(q)) {
    struct ncclTaskRma* task = ncclIntruQueueHead(q);
    if (!isRmaPutOrSignal(task->func)) break;
    ncclIntruQueueDequeue(q);
    ...
  }
}

The intent of this design is: a single kernel launch covers put/signal for all contexts, reducing launch overhead. But each context's queue only consumes up to the first WaitSignal, ensuring per-context FIFO ordering. If the upper layer alternates put and waitSignal calls within the same context, the batching effect is greatly diminished—this is a pattern to watch out for when using RMA.

---

II. GPU Direct Network: How GIN Lets Kernels Bypass the Host Proxy

Intuitive Model

Traditional NCCL network communication is like mailing a letter: the GPU kernel puts data into a buffer, the host proxy thread hands the data to the NIC, and the NIC sends it out. GIN, on the other hand, lets the GPU kernel drop the letter directly into the recipient's mailbox—the kernel writes directly to the NIC's send queue, and the NIC reads directly from GPU memory.

Without GIN, every network communication must go through host memory as an intermediary, adding at least one PCIe round-trip of latency. For fine-grained communication like MoE, this latency is fatal.

Data Structures and Memory Layout

The core state of GIN isncclGinState, which manages multiple backends and multiple DevComms. Let's first look at the backend version compatibility table.

📎 src/gin/gin_host.cc:27-33

cpp
const int proxyBackendMinVersions[] = {0, NCCL_VERSION(2, 30, 3), NCCL_VERSION(2, 30, 5), NCCL_VERSION(2, 32, 0)};
const int gdakiBackendMinVersions[] = {0, NCCL_VERSION(2, 30, 3), NCCL_VERSION(2, 30, 5)};
const int gpiBackendMinVersions[] = {0, NCCL_VERSION(2, 30, 5)};
constexpr int efaGdaBackendMinVersions[] = {0, NCCL_VERSION(2, 31, 0), NCCL_VERSION(2, 32, 0)};

The index of these arrays is the backend version number, and the value is the minimum compatible NCCL version. For example,proxyBackendMinVersions[3]corresponds to backend version 3, requiring NCCL at least 2.32.0. This design allows NCCL to select the appropriate backend version at runtime based on the device code version, rather than binding at compile time.

[Design Inference and Architectural Trade-offs]

The motivation behind this version compatibility table design is: the GIN backend (NIC driver, firmware) and the NCCL library evolve at different paces. If version requirements were hardcoded, an upgrade on either side would cause incompatibility. Using arrays for version mapping allows dynamic selection at runtime, maintaining backward compatibility with older backends.

ncclGinStateDevCommis the GIN state for each DevComm, containingcontextCount、backendIndex、ginCtx[]、devHandles[]and other fields. It is chained into a linked list attached toginState->devComms.

Step-by-Step Walkthrough: Establishing a GIN Connection

Let's walk through a scenario: rank 0 initializes the communication domain and needs to establish a GIN connection.

Step 1: Check whether GIN is enabled and supported.

📎 src/gin/gin_host.cc:96-107

cpp
if (ginState->connected) return ncclSuccess;
if (ncclParamGinEnable() == 0) {
  WARN("GIN is disabled.");
  return ncclInternalError;
}
if (!ginState->supported) {
  WARN("GIN not supported.");
  return ncclInvalidUsage;
}

ncclParamGinEnable()reads the environment variableNCCL_GIN_ENABLE, defaulting to 1. If the user explicitly disables it, return an error directly.

Step 2: Check symmetric memory support.

📎 src/gin/gin_host.cc:111-114

cpp
if (!comm->symmetricSupport) {
  WARN("Communicator does not support symmetric memory!");
  return ncclInternalError;
}

GIN relies on symmetric memory—because the GPU kernel needs to know the virtual address of the peer's buffer, and only symmetric memory can guarantee address consistency.

Step 3: Get the local GIN device list.

📎 src/gin/gin_host.cc:116-122

cpp
int nLocalGinDevs;
int localGinDevs[NCCL_TOPO_MAX_NODES];
NCCLCHECK(ncclTopoGetLocalGinDevs(comm, localGinDevs, &nLocalGinDevs));
if (nLocalGinDevs > NCCL_GIN_MAX_CONNECTIONS) {
  ATTN("Found %d local devices, but GIN supports at most %d connections. Using the first %d connections.",
       nLocalGinDevs, NCCL_GIN_MAX_CONNECTIONS, NCCL_GIN_MAX_CONNECTIONS);
}

ncclTopoGetLocalGinDevsfinds all NICs that support GIN from the topology graph. If it exceedsNCCL_GIN_MAX_CONNECTIONS, only take the first few and print a warning.

Step 4: Compute the GIN team.

📎 src/gin/gin_host.cc:138-149

cpp
ginTeam = ncclTeamWorld(comm);
if (ginState->ginConnectionType != NCCL_GIN_CONNECTION_FULL) {
  ginTeam = {
    .nRanks = comm->nRanks / comm->contiguousRanksPerHost,
    .rank = comm->rank / comm->contiguousRanksPerHost,
    .stride = comm->contiguousRanksPerHost,
  };
}
for (int r = 0; r < ginTeam.nRanks; r++) {
  int worldRank = ncclTeamRankToWorld(comm, ginTeam, r);
  handles[r] = allHandles + worldRank * NCCL_NET_HANDLE_MAXSIZE;
}

If the connection type is FULL, the GIN team is the entire world team; otherwise, only connect the first rank of each host (rail connection).ncclTeamRankToWorldconverts team ranks to world ranks.

Step 5: Establish connections backend by backend.

📎 src/gin/gin_host.cc:151-202

cpp
for (int backendIdx = 0; backendIdx < ginState->numActiveBackends; backendIdx++) {
  backend = &ginState->backends[backendIdx];
  NCCLCHECKGOTO(backend->ncclGin->devices(&ndev), ret, fail);
  ...
  for (int commIdx = 0; commIdx < backend->ginCommCount; commIdx++) {
    NCCLCHECKGOTO(backend->ncclGin->listen(...), ret, fail);
    NCCLCHECKGOTO(backend->ncclGin->getProperties(...), ret, fail);
    NCCLCHECKGOTO(bootstrapAllGather(comm->bootstrap, allHandles, NCCL_NET_HANDLE_MAXSIZE), ret, fail);
    NCCLCHECKGOTO(backend->ncclGin->connect(...), ret, fail);
    NCCLCHECKGOTO(backend->ncclGin->closeListen(...), ret, fail);
  }
}

Each backend first callsdevicesto get the device count, then performs the listen→getProperties→allGather→connect→closeListen flow for each connection.bootstrapAllGatherexchanges handles among all ranks, so that each rank knows the connection information of its peers.

Concurrency Control and Hardware Interaction

The GIN progress thread is the core concurrency mechanism.

📎 src/gin/gin_host.cc:56-87

cpp
void* ncclGinProgress(struct ncclGinState* ginState, int threadIdx) {
  if (ncclOsCpuCount(ginState->cpuAffinity)) {
    ncclOsSetAffinity(ginState->cpuAffinity);
  }
  while (1) {
    if (ginState->proxyThreadStopSignal.load()) return NULL;
    if (ginState->writePending.load()) {
      std::this_thread::yield();
      continue;
    }
    {
      std::shared_lock<std::shared_timed_mutex> rlock(ginState->devCommRwMutex);
      struct ncclGinStateDevComm* dc = ginState->devComms;
      while (dc) {
        struct ncclGinBackendState* backend = &ginState->backends[dc->backendIndex];
        for (int commIdx = threadIdx; commIdx < backend->ginCommCount; commIdx += ginState->proxyNthreads) {
          if (dc->devHandles[commIdx]->needsProxyProgress) {
            ncclResult_t ret = backend->ncclGin->ginProgress(dc->ginCtx[commIdx]);
            if (ret != ncclSuccess) {
              COMPILER_ATOMIC_STORE(&ginState->asyncResult, ret, std::memory_order_release);
              return NULL;
            }
          }
        }
        dc = dc->next;
      }
    }
    std::this_thread::yield();
  }
}

There are several key design points here:

1. CPU Affinity:ncclOsSetAffinitybinds the progress thread to a specified CPU core, avoiding cache invalidation caused by thread migration.

2. Write Lock Backoff:writePendingis an atomic flag. When the main thread needs to modify thedevCommslinked list, it sets this flag first, and the progress thread yields voluntarily upon seeing it, avoiding lock contention.

3. Read-Write Lock:devCommRwMutexisshared_timed_mutex. The progress thread holds the read lock to traverse the linked list, and the main thread holds the write lock to modify the linked list.

4. Thread Division of Labor: thread t is responsible for connections t, t+proxyNthreads, t+2*proxyNthreads, ..., achieving load balancing through a stride loop.

📎 src/gin/gin_host.cc:43-47

cpp
static void ginProgressWriteLock(struct ncclGinState* ginState) {
  ginState->writePending.store(true);
  ginState->devCommRwMutex.lock();
}
static void ginProgressWriteUnlock(struct ncclGinState* ginState) {
  ginState->devCommRwMutex.unlock();
  ginState->writePending.store(false);
}

This write lock implementation assumes there is only one writer (the main thread), so no additional mutex is needed.writePendingsets the flag first before acquiring the lock, ensuring the progress thread can see the write intent before acquiring the lock and voluntarily back off.

Production Pitfall Guide

Pitfall 1: Mismatched GIN connection counts causing AllGather deadlock.Each rank'sginCommCountmay differ (depending on the number of local NICs), and NCCL takes the minimum across all ranks viabootstrapAllGather.

📎 src/gin/gin_host.cc:176-180

cpp
ginCommCountHandles[comm->rank] = backend->ginCommCount;
NCCLCHECKGOTO(bootstrapAllGather(comm->bootstrap, ginCommCountHandles, sizeof(int)), ret, fail);
for (int r = 0; r < comm->nRanks; r++) {
  backend->ginCommCount = std::min(backend->ginCommCount, ginCommCountHandles[r]);
}

If a rank has fewer NICs than other ranks, all ranks drop to the minimum. This guarantees connection symmetry but wastes NIC resources.

Pitfall 2: proxyNthreads exceeding ginCommCount causes thread spinning.If the user setsNCCL_GIN_PROXY_NTHREADSgreater thanginCommCountthe excess threads will spin in the stride loop.

📎 src/gin/gin_host.cc:181-183

cpp
// After cross-rank min, proxyNthreads may exceed ginCommCount if ranks disagree
// on NCCL_GIN_PROXY_NTHREADS (atypical — env vars are normally uniform across a job).
// Extra threads simply idle in the stride loop; no correctness issue.

This is not a correctness issue, but it wastes CPU resources. The way to troubleshoot is to check whetherNCCL_GIN_PROXY_NTHREADSis greater than the actual number of NICs.

Pitfall 3: race condition when releasing DevComm. ncclGinDevCommFreeFirst remove DevComm from the linked list, then destroy the context.

📎 src/gin/gin_host.cc:464-475

cpp
ginProgressWriteLock(ginState);
if (prevDc) prevDc->next = dc->next;
else ginState->devComms = dc->next;
ginProgressWriteUnlock(ginState);
struct ncclGinBackendState* backend = &ginState->backends[dc->backendIndex];
for (int commIdx = 0; commIdx < backend->ginCommCount; commIdx++) {
  NCCLCHECK(backend->ncclGin->destroyContext(dc->ginCtx[commIdx]));
}

After removal, the progress thread can no longer see this DevComm, so destroying the context is safe. However, if there are in-flight network operations during destruction, it may lead to undefined behavior—this is what must be ensured when using GIN: before releasing DevComm, all operations must be confirmed complete.

---

III. Symmetric memory kernel: from "registered buffers" to "unified address space"

Intuitive model

Traditional NCCL buffers are "registration-based": each rank registers its own buffer, and addresses are exchanged via handles during communication. Symmetric memory, by contrast, is a "unified address space": all ranks agree on the same set of virtual addresses; address A on rank 0 and address A on rank 1 point to their respective physical memory, but the same address can be used in code to access them.

This is like everyone agreeing that "row 3, seat 5" refers to the same location in each person's home, so when looking for something you don't have to first ask "where is row 3, seat 5 in your home?"

Without symmetric memory, every kernel would have to resolve the peer address first, increasing instruction overhead and register pressure.

Data structures and memory layout

The core of the symmetric memory kernel is the kernel mask—a bitmap that marks which kernels are available in the current communication domain.

📎 src/sym_kernels.cc:17-63

cpp
constexpr uint32_t kernelMask_STMC =
  1 << ncclSymkKernelId_AllGather_LLMC | 1 << ncclSymkKernelId_AllGather_STMC |
  ...
constexpr uint32_t kernelMask_LDMC = ...;
constexpr uint32_t kernelMask_LL = ...;
constexpr uint32_t kernelMask_AG = ...;
constexpr uint32_t kernelMask_AR = ...;
constexpr uint32_t kernelMask_RS = ...;
constexpr uint32_t kernelMask_LSA = ...;
constexpr uint32_t kernelMask_Gin = ...;
constexpr uint32_t kernelMask_Tma = ...;

Each mask is a 32-bit integer, and bit i being 1 indicates that kernel i is available. These masks are grouped by different dimensions:

  • By protocol:STMC(Simple TMA Multimem Copy)、LDMC(Low-latency Direct Multimem Copy)、LL(Low Latency)
  • By operation:AG(AllGather)、AR(AllReduce)、RS(ReduceScatter)
  • By hardware:LSA(Local Symmetric Access)、Gin(GPU-Initiated Networking)、Tma(Tensor Memory Accelerator)
[Design inference and architectural trade-offs]

The advantage of this bitmap design is that bitwise operations can quickly filter available kernels. For example,kmask &= ~kernelMask_STMCa single line can disable all STMC kernels without traversing the list.

Step-by-Step Walkthrough: one kernel mask computation

Let's plug in a scenario: rank 0 wants to execute AllReduce, the data type is float16, the message size is 1MB, the communication domain has 8 ranks, and all are interconnected via NVLink.

Step 1: Get the base mask corresponding to the operation.

📎 src/sym_kernels.cc:304-306

cpp
uint32_t kmask = kernelMask_coll(coll);

kernelMask_coll(ncclFuncAllReduce)returnskernelMask_ARcontaining 5 AllReduce kernels.

Step 2: Check STMC and LDMC availability.

📎 src/sym_kernels.cc:308-334

cpp
bool hasSTMC = comm->symkState.hasLsaMultimem;
bool hasLDMC = false;
if (comm->symkState.hasLsaMultimem) {
  switch (ty) {
  case ncclFloat16:
  case ncclBfloat16:
    hasLDMC = red == ncclDevSum || red == ncclDevMinMax || red == ncclDevSumPostDiv;
    break;
  ...
  }
}
if (!hasSTMC) kmask &= ~kernelMask_STMC;
if (!hasLDMC) kmask &= ~kernelMask_LDMC;

hasLsaMultimemis computed inncclSymkInitOncerequiring NVLS symmetric multicast to be available and the LSA team to have more than 2 ranks. float16 supports LDMC, so ifhasLsaMultimemis true, the LDMC kernel is retained.

Step 3: Check the message size limit.

📎 src/sym_kernels.cc:336-342

cpp
size_t nBytes = alignUp(nElts * ncclTypeSize(ty), NCCL_SYM_KERNEL_CELL_SIZE);
size_t nBusBytes = (coll == ncclFuncAllReduce ? 1 : comm->nRanks) * nBytes;
if (nBusBytes >= (size_t(2) << 30)) kmask &= ~kernelMask_LL;
if (nBusBytes >= 32 * (size_t(2) << 30)) kmask = 0;

The LL kernel uses a 32-bit integer to track element counts, so it is disabled when the total bus bytes exceed 2GB. If it exceeds 64GB, all kernels are disabled (32-bit integer overflow).

Step 4: Check TMA availability.

📎 src/sym_kernels.cc:344-345

cpp
if (!ncclSymkTmaAvailable(comm)) kmask &= ~kernelMask_Tma;
if (!symAligned16B) kmask &= ~kernelMask_Tma;

TMA requires SMEM capacity and compute capability 10.0+, and the buffer must be 16-byte aligned.

Step 5: Check GIN requirements.

📎 src/sym_kernels.cc:347-350

cpp
bool hasGin = ncclParamSymGinKernelsEnable() != 0;
if (!hasGin) kmask &= ~kernelMask_Gin;
bool needGin = ncclTeamLsa(comm).nRanks < comm->nRanks;
kmask &= needGin ? kernelMask_Gin : ~kernelMask_Gin;

If the LSA team covers all ranks, GIN is not needed; otherwise, only the GIN kernel is retained.

Concurrency control and hardware interaction

Initialization of the symmetric memory kernel involves DevComm creation and resource allocation.

📎 src/sym_kernels.cc:185-264

cpp
ncclResult_t ncclSymkInitOnce(struct ncclComm* comm) {
  NCCLCHECK(ncclDevrInitOnce(comm));
  struct ncclSymkState* symk = &comm->symkState;
  if (!symk->initialized) {
    symk->initialized = true;
    struct ncclDevCommRequirements reqs = NCCL_DEV_COMM_REQUIREMENTS_INITIALIZER;
    symk->hasLsaMultimem = ncclNvlsSymmetricMultimemEnabled(comm) && ncclTeamLsa(comm).nRanks > 2 && !comm->p2pCrossClique;
    reqs.lsaMultimem = symk->hasLsaMultimem;
    reqs.lsaBarrierCount = ncclSymkMaxBlocks;
    ...
    NCCLCHECK(ncclDevrCommCreateInternal(comm, &reqs, &symk->kcomm.devComm, /*isInternal=*/true, /*deviceCodeVersion=*/NCCL_VERSION_CODE));
  }
  return ncclSuccess;
}

The key here isncclDevrCommCreateInternalwhich creates an internal DevComm containing resources such as LSA multicast, GIN inbox/outbox, and signals.reqs.ginConnectionType = NCCL_GIN_CONNECTION_RAILspecifies that GIN uses rail connection mode.

📎 src/sym_kernels.cc:257-261

cpp
symk->kcomm.workStarted = comm->profiler.symWorkStarted;
symk->kcomm.workCompleted = comm->profiler.symWorkCompleted;
symk->kcomm.workPhases = comm->profiler.symWorkPhases;

The symmetric memory kernel uses an independent profiler buffer to avoid interleaving with the workCounter of regular kernels.

Production pitfall avoidance guide

Pitfall 1: SMEM requirements of the TMA kernel.TMA requires about 8KB of SMEM scratch per warp, so 16 warps means 128KB.

📎 src/sym_kernels.cc:135-142

cpp
bool ncclSymkTmaAvailable(struct ncclComm* comm) {
  if (comm->maxSharedMemOptin < ncclTmaShmemScratchWarpSize() * 16) {
    return false;
  }
  return comm->minCompCap >= 100 && ncclParamSymTmaEnable();
}

If the GPU's SMEM capacity is insufficient (such as in a MIG instance), the TMA kernel will be disabled. The way to troubleshoot is to check whethermaxSharedMemOptinis less thanncclTmaShmemScratchWarpSize() * 16。

Pitfall 2: boundaries of the GIN chunk size.The chunk size of the ReduceScatter GIN kernel has upper and lower limits.

📎 src/sym_kernels.cc:148-153

cpp
static constexpr size_t ncclSymkRsGinDefaultChunkBytes = 128 << 10;
static constexpr size_t ncclSymkRsGinMinChunkBytes = 128;
static constexpr size_t ncclSymkRsGinMaxChunkBytes = size_t(1) << 30;
size_t ncclSymkRsGinChunkBytes() {
  int64_t param = ncclParamSymRsGinChunkSize();
  size_t chunkBytes = param > 0 ? (size_t)param : ncclSymkRsGinDefaultChunkBytes;
  chunkBytes = std::max(ncclSymkRsGinMinChunkBytes, std::min(chunkBytes, ncclSymkRsGinMaxChunkBytes));
  return pow2Down(chunkBytes);
}

If the user setsNCCL_SYM_RS_GIN_CHUNK_SIZEexceeding 1GB, it will be truncated to 1GB; if it is less than 128 bytes, it will be raised to 128 bytes. The final value will also be rounded down to a power of 2.

Pitfall 3: Symmetric memory registration type mismatch. ncclGetSymRegTypeBased on the flags of sendWin and recvWin,NCCL_WIN_COLL_SYMMETRICdetermine the registration type.

📎 src/sym_kernels.cc:395-412

cpp
if (!isSendSymmReg && !isRecvSymmReg) {
  *winRegType = ncclSymSendNonregRecvNonreg;
} else if (isSendSymmReg && !isRecvSymmReg) {
  *winRegType = ncclSymSendRegRecvNonreg;
} else if (!isSendSymmReg && isRecvSymmReg) {
  *winRegType = ncclSymSendNonregRecvReg;
} else if (isSendSymmReg && isRecvSymmReg) {
  *winRegType = ncclSymSendRegRecvReg;
}

If the registration types of send and recv are inconsistent, the kernel needs to take different code paths. This affects performance but does not cause errors.

---

IV. Team Abstraction and Versioned DevComm: Evolving Infrastructure

Intuitive Model

The Team abstraction is like "grouping": the world team is the whole class, the LSA team is deskmates, and the Rail team is seats in the same column. Different communication patterns require different grouping perspectives.

Versioned DevComm is like a "translator": different versions of device code speak different "dialects," and the DevComm compatibility layer handles translation so old and new code can understand each other.

Without the Team abstraction, every kernel would have to compute rank mappings itself; without versioned DevComm, any ABI change would force all device code to be recompiled.

Data Structures and Memory Layout

A Team is a simple triple:nRanks、rank、stride。

📎 src/nccl_device/core.cc:13-19

cpp
ncclTeam_t ncclTeamWorld(ncclComm_t comm) {
  ncclTeam_t ans;
  ans.nRanks = comm->nRanks;
  ans.rank = comm->rank;
  ans.stride = 1;
  return ans;
}

The world team's stride is 1 because all ranks are arranged consecutively.

📎 src/nccl_device/core.cc:70-79

cpp
ncclTeam_t ncclTeamRail(ncclComm_t comm) {
  if (ncclSuccess != ncclDevrInitOnce(comm)) return ncclTeam_t{};
  ncclTeam_t ans;
  ans.nRanks = comm->nRanks / comm->devrState.lsaSize;
  ans.rank = comm->rank / comm->devrState.lsaSize;
  ans.stride = comm->devrState.lsaSize;
  return ans;
}

The Rail team's stride islsaSize, because ranks on each rail are separated by the size of one LSA team.

The core of versioned DevComm is thencclDevCommCompatstructure.

📎 src/devcomm/devcomm_v23100.cc:10-17

cpp
struct ncclDevCommCompat ncclDevCommCompat_v23100 = {
  NCCL_VERSION(2, 31, 0), // minVersion
  NCCL_VERSION_CODE, // maxVersion
  nullptr,           // commPropertiesFilter
  nullptr,           // devCommRequirementsFilter
  nullptr,           // devCommCopyNewToOld
  nullptr,           // devCommCopyOldToNew
};

This structure defines the compatibility rules for version 2.31.0.minVersionandmaxVersiondefine the applicable version range, and the following four function pointers define attribute filtering and structure conversion logic. If all are nullptr, it means this version has no special compatibility requirements.

Step-by-Step Walkthrough: A Team Conversion

Let's use a scenario: rank 5 in an 8-rank communication domain, with an LSA team size of 4. We want to compute rank 5's rank in the Rail team.

Step 1: Initialize DevR state.

📎 src/nccl_device/core.cc:70-79

cpp
if (ncclSuccess != ncclDevrInitOnce(comm)) return ncclTeam_t{};

ncclDevrInitOnceComputes derived information such as the LSA team and CFT team. If it fails, returns an empty team.

Step 2: Compute Rail team parameters.

📎 src/nccl_device/core.cc:70-79

cpp
ncclTeam_t ans;
ans.nRanks = comm->nRanks / comm->devrState.lsaSize;  // 8 / 4 = 2
ans.rank = comm->rank / comm->devrState.lsaSize;       // 5 / 4 = 1
ans.stride = comm->devrState.lsaSize;                  // 4

Rank 5's rank in the Rail team is 1, the team has 2 ranks, and the stride is 4.

Step 3: Convert back to world rank.

📎 src/nccl_device/core.cc:82-84

cpp
int ncclTeamRankToWorld(ncclComm_t comm, ncclTeam_t team, int rank) {
  return comm->rank + (rank - team.rank) * team.stride;
}

If you want to convert Rail rank 0 to a world rank:5 + (0 - 1) * 4 = 1. Verification: rank 1 and rank 5 are on the same rail (separated by 4).

Concurrency Control and Hardware Interaction

The Team abstraction itself is stateless and requires no concurrency control. ButncclDevrInitOnceis lazily loaded, and all derived information is computed on the first call.

📎 src/nccl_device/core.cc:22-33

cpp
ncclTeam_t ncclTeamLsa(ncclComm_t comm) {
  if (ncclSuccess != ncclDevrInitOnce(comm)) return ncclTeam_t{};
  ncclTeam_t ans;
  ans.nRanks = comm->devrState.lsaSize;
  ans.rank = comm->devrState.lsaSelf;
  ans.stride = 1;
  return ans;
}

The comment says "Ignoring errors since if it fails ncclDevrInitOnce will try again" — if initialization fails, it returns an empty team, and the next call will retry.

Production Pitfall Guide

Pitfall 1: Stride assumptions in Team conversion. ncclTeamRankToWorldassumes that ranks within a team form an arithmetic sequence.

📎 src/nccl_device/core.cc:82-84

cpp
int ncclTeamRankToWorld(ncclComm_t comm, ncclTeam_t team, int rank) {
  return comm->rank + (rank - team.rank) * team.stride;
}

If the team is not an arithmetic sequence (for example, a custom arbitrary grouping), this function will compute incorrectly. NCCL currently only supports regular teams.

Pitfall 2: Null pointers in versioned DevComm. ncclDevCommCompat_v23100All function pointers in are nullptr, indicating there is no special compatibility logic. If a future version requires conversion, these functions must be implemented; otherwise, old and new code cannot interoperate.

Pitfall 3: Hierarchy modes of the CFT team. ncclTeamCftsupports three modes: FLAT, HIER_MULTIMEM, and HIER_LSA.

📎 src/nccl_device/core.cc:36-55

cpp
if (mode == NCCL_CFT_TEAM_FLAT) return flatTeam;
int innerSize;
if (mode == NCCL_CFT_TEAM_HIER_MULTIMEM) {
  innerSize = comm->devrState.cftMcSize;
} else if (mode == NCCL_CFT_TEAM_HIER_LSA) {
  innerSize = comm->devrState.lsaSize;
} else {
  return ncclTeam_t{};
}
return ncclTeamOuterFactor(flatTeam, innerSize);

If an invalid mode is passed in, it returns an empty team. When using a CFT team, you need to ensure the mode is correct.

---

Design Reflections

Why does NCCL support three evolution paths simultaneously: RMA, GIN, and symmetric memory?

[Design Inference and Architectural Trade-offs]

These three paths solve problems at different levels:

  • RMAsolves the problem of "fixed communication patterns" — allowing upper layers to compose primitives and implement arbitrary communication patterns.
  • GINsolves the problem of "high network latency" — allowing the GPU to directly drive the NIC, bypassing the host proxy.
  • Symmetric memorysolves the problem of "address resolution overhead" — allowing the kernel to directly access peer memory using a unified address.

They are not substitutes but complements. RMA can use GIN as the underlying transport, and GIN relies on symmetric memory to provide address consistency. Together, the three form the infrastructure of a "programmable communication engine."

What is the design philosophy of versioned DevComm?

[Design Inference and Architectural Trade-offs]

The core idea of versioned DevComm is "stable ABI, evolving API." Device code (kernels) is compiled and embedded in binaries and cannot be recompiled as the NCCL library upgrades. Therefore, NCCL must ensure that old device code can run on the new library.ncclDevCommCompatThe structure is the entry point of the compatibility layer: the new library selects the appropriate compatibility rules based on the device code version and performs structure conversion when necessary.

---

Chapter Summary

In this chapter, starting from the traces of evolution in the source code, we analyzed the three forces driving NCCL from a collective communication library toward a programmable communication engine:

1. RMA(src/rma/rma.cc): Through the combination of Put/Signal/WaitSignal primitives, upper layers can implement arbitrary communication patterns. The core design splits tasks into two parallel execution paths, CE and Proxy, based on LSA reachability.

2. GIN(src/gin/gin_host.cc): Through GPU direct network transmission, bypassing the host proxy. The core design includes multi-backend management, version compatibility tables, and a progress thread pool.

3. Symmetric memory kernel(src/sym_kernels.cc): Through a unified address space, eliminating address resolution overhead. The core design includes kernel mask bitmaps and TMA/GIN hardware acceleration.

4. Team abstraction and versioned DevComm(src/nccl_device/core.cc、src/devcomm/devcomm_v23100.cc): Providing infrastructure for evolution. Team provides a grouping perspective, and versioned DevComm provides ABI compatibility.

The impact of these changes on upper-layer frameworks is profound: PyTorch's ProcessGroup can directly call RMA primitives to implement custom communication patterns; Megatron's expert parallelism can leverage GIN to reduce all-to-all latency; symmetric memory makes kernel code more concise.

Chapter Review and Self-Test

Q1: If thescheduleRmaTasksToPlanLSA reachability check in the WaitSignal branch is removed, and all peers go through the Proxy path, what would be the consequences? In what scenarios would this trigger a performance disaster?

Reference Analysis:

The LSA reachability check is in📎 src/rma/rma.cc:187-204, which divides peers into two groups: CE and Proxy. If this check is removed, all peers go through the Proxy path,nRmaTasksCeis always 0.

The consequence is: the CE path is completely unused, and all WaitSignal operations poll the network through host proxy threads. For peers within LSA range (same-machine NVLink interconnect), which could originally use GPU copy engines for asynchronous waiting, now become host thread polling, with latency rising from microseconds to milliseconds.

Performance disaster scenario: In MoE training, each token needs to wait for signals from multiple experts. If all signals go through Proxy, the host thread becomes the bottleneck, and the GPU spends a large amount of time waiting for host polling. On an 8-GPU all-NVLink machine, this degradation is especially pronounced—all communication that could originally go through CE now crowds onto the host.

Troubleshooting method: Check thescheduleRmaTasksToPlanINFO logs. IfnRmaTasksCeis always 0 whilenRmaTasksProxyis very large, it indicates a problem with the LSA check.

Q2:ncclGinProgressInwritePendingflag anddevCommRwMutexread-write lock coordination, if thewritePendingcheck is removed and only the read-write lock is kept, what problems would arise?

Reference Analysis:

writePendingThe check is in📎 src/gin/gin_host.cc:63-66, which makes the progress thread actively yield when the main thread wants to write. If this check is removed, the progress thread will directly attempt to acquire the read lock.

The problem is:std::shared_timed_mutex's read lock is shared, and multiple progress threads can hold it simultaneously. If the main thread wants to acquire the write lock, it must wait for all read locks to be released. Under high load, progress threads frequently acquire read locks, and the main thread may be unable to acquire the write lock for a long time, causingncclGinDevCommSetuporncclGinDevCommFreeto block.

More seriously: if the main thread first setsginProgressWriteLockinwritePendingbefore acquiring the lock, and the progress thread does not checkwritePending, then the progress thread may still acquire the read lock after the main thread sets the flag, causing unpredictable wait times for the main thread.

writePendingThe purpose of is a "soft notification": telling progress threads "I'm about to write, please yield." This is more efficient than relying solely on lock fairness, because progress threads can actively yield rather than block on the lock.

Q3:ncclSymkMaskInnBusBytes >= 32 * (size_t(2) << 30), if all kernels are disabled whenkmask = 0), at this pointncclSymkAvailablereturns false, what path will NCCL fall back to? What performance impact does this fallback path have?

Reference Analysis:

kmask = 0In📎 src/sym_kernels.cc:342, at this pointncclSymkAvailablereturns false (📎 src/sym_kernels.cc:354-361)。

The fallback path is: NCCL will use traditional collective communication kernels (non-symmetric memory kernels). These kernels access peer memory through registered buffers, requiring address resolution first, with higher instruction overhead.

Performance impact: For very large messages (exceeding 64GB bus bytes), the address resolution overhead of traditional kernels is a small proportion, because data transfer itself dominates. But in boundary cases (just exceeding 64GB), traditional kernels may be 10-20% slower than symmetric memory kernels.

The root cause of this limitation is: symmetric memory kernels use 32-bit integers to track unrolled loop chunks, with each chunk being at least 32 bytes, so the maximum addressable range is 32 * 2^31 = 64GB. Exceeding this range causes integer overflow.

In actual production, scenarios where a single collective communication exceeds 64GB are rare (usually all-reduce after gradient accumulation), but not impossible. If such a scenario is encountered, consider sharded communication or using traditional kernels.

---

Chapter Transition

In this chapter, we have seen NCCL moving from "fixed collective operations" toward a "programmable communication engine": RMA provides primitive composition, GIN provides GPU direct transmission, symmetric memory provides a unified address space, and Team and versioned DevComm provide infrastructure.

These evolutions are not isolated; they collectively point toward one goal:Enable upper-layer frameworks to implement custom communication patterns with lower latency and greater flexibility. For frameworks like PyTorch and Megatron, this means they can directly build complex communication patterns such as MoE all-to-all, pipeline parallelism, and expert parallelism on top of NCCL, without needing to bypass NCCL and implement the network layer themselves.

The next chapter is the final chapter of the book. We will walk through the complete path of a single AllReduce once again—starting from thencclAllReducecall, going through task enqueue, algorithm selection, kernel launch, proxy progression, network transmission, until the result is returned. This review will connect the knowledge points from the previous 24 chapters into a complete cognitive map.

At this point, we have seen the three main lines of NCCL's evolution from fixed collective operations to a programmable communication engine: RMA primitive composition, GPU direct network transmission, symmetric memory model, and the team abstraction and versioned DevComm that support them. These mechanisms together point toward a more flexible communication future that is closer to hardware capabilities. However, no matter how the architecture evolves, the complete path of a single AllReduce remains the cornerstone of understanding NCCL. In the next chapter, we will not introduce new code, but instead re-narrate the end-to-end flow from Chapter 3 to Chapter 10—from the ncclAllReduce call, to communicator establishment, topology search, algorithm selection, task enqueue, kernel launch, device-side primitive execution, and result write-back. You will reassemble the mechanisms scattered across chapters into a complete mental model, and obtain an index of "which chapter to check when encountering a problem."

CHAPTER 25

Chapter 25: Chapter 25: Panoramic Review and Reflections: The Ultimate Journey and Design Essence of an AllReduce

Official source: NVIDIA/nccl · Version: Commit @12df1a11 · Book progress: Chapter 25 / 25

Chapter 25: Panoramic Review and Reflections: The Ultimate Journey and Design Essence of an AllReduce

In the previous chapter, based on the traces of evolution in the source code, we looked ahead at NCCL's architectural trends: from fixed collective operations to programmable ones, from host proxy to GPU direct transmission, and from registered buffers to symmetric memory. Now, it is time to put these trends back into a concrete execution flow for verification. This chapter does not introduce any new code, but instead reconnects the end-to-end path from Chapter 3 to Chapter 10—starting from the single call ncclAllReduce, all the way to writing the result back to device memory. After reading this, you should be able to clearly answer: which functions does a single AllReduce actually go through? In which file and on which line is each function? Which chapter should you consult when encountering a problem?

1. Initialization: How the communicator "grows" out

Intuitive model

Think of the communicator as a "group chat." When you callncclCommInitRankit is like "applying to join the group chat." At this point, NCCL must determine the full member list (peerInfo), who connects to whom through which route (topology graph), and how many pipelines each route opens (channel).If this step goes wrong, all subsequent communication will be wrong—just like when someone in a group chat has not been pulled in, the messages you send will always be missing one recipient.

Data structures and memory layout

The core structure of the communicator isncclComm, and its initialization is divided into two stages:commAllocis responsible for "allocating the skeleton,"initTransportsRankis responsible for "filling in the flesh and blood."

commAllocThe most noteworthy thing inis the design ofshared resource reference countingncclSharedResources. When a sub-communicator (produced by split/shrink) reuses the parent communicator's resources, it does not copy a separate set, but shares the same

📎 src/init.cc:533-555

cpp
if (parent == NULL || !parent->shareResources) {
    struct ncclSharedResources* sharedRes;
    NEW_NOTHROW(sharedRes, ncclSharedResources);
    sharedRes->owner = comm;
    ...
    comm->sharedRes = sharedRes;
    sharedRes->refCount = 1;
    NCCLCHECK(ncclNetInit(comm));
    NCCLCHECK(ncclRmaInit(comm));
    NCCLCHECK(ncclGinInit(comm));
} else {
    comm->sharedRes = parent->sharedRes;
    ncclAtomicRefCountIncrement(&parent->sharedRes->refCount);
    NCCLCHECK(ncclNetInitFromParent(comm, parent));
    NCCLCHECK(ncclRmaInitFromParent(comm, parent));
}

CopyrefCountThe intent of this code is very clear: "heavy resources" such as network plugins, RMA, and GIN are initialized only once, and sub-communicators directly borrow them.

uses atomic operations to increment, ensuring that under multithreading there will be no duplicate release.commAllocAnother key point is theinitialization ofchannels inid = -1. All channels are first marked as "uninitialized" (setupChannel), and only later will

📎 src/init.cc:607-608

cpp
// Mark channels as non initialized.
for (int c = 0; c < MAXCHANNELS; c++) comm->channels[c].id = -1;

Copy-1Thisid == -1is a sentinel value. If any code mistakenly uses an uninitialized channel,

will immediately expose the problem, rather than reading a bunch of random memory.

Step-by-Step: From ncclCommInitRank to initTransportsRankncclCommInitRankAfter the user calls

1. ncclCommInitRank, the actual execution flow is as follows:ncclInitEnvfirst callsncclGroupStartInternalto load environment plugins, then calls

to enter group semantics (this is to support "initializing multiple communicators within one group").ncclCommInitRankDev2. Next, it callscomm, which performs parameter validation, allocates thestructure, parses config, and then:

📎 src/init.cc:2923-2929

cpp
if (ncclParamEnqueueRearchEnable()) {
    NCCLCHECKGOTO(ncclMgmtTaskEnqueue((struct ncclAsyncJob*)job, ncclCommInitRankFunc, ncclCommInitJobFree, comm), res, fail);
} else {
    NCCLCHECKGOTO(ncclAsyncLaunch((struct ncclAsyncJob*)job, ncclCommInitRankFunc, NULL, ncclCommInitJobFree, comm), res, fail);
}

CopyncclParamEnqueueRearchEnable()Note thencclAsyncLaunchbranch here—this is a trace of the "enqueue refactor" currently underway in NCCL. By default it goes throughncclMgmtTaskEnqueue, and after enabling the refactor it goes throughncclCommInitRankFunc。

3. ncclCommInitRankFunc. Both paths will eventually call

📎 src/init.cc:2119-2127

cpp
timers[TIMER_INIT_TOTAL] = clockNano();
CUDACHECKGOTO(cudaSetDevice(cudaDev), res, fail);
CUDACHECKGOTO(cudaDeviceGetAttribute(&maxSharedMem, cudaDevAttrMaxSharedMemoryPerBlockOptin, cudaDev), res, fail);
CUDACHECKGOTO(cudaDeviceGetAttribute(&archMajor, cudaDevAttrComputeCapabilityMajor, cudaDev), res, fail);
CUDACHECKGOTO(cudaDeviceGetAttribute(&archMinor, cudaDevAttrComputeCapabilityMinor, cudaDev), res, fail);
cudaArch = 100 * archMajor + 10 * archMinor;

timers[TIMER_INIT_KERNELS] = clockNano();
NCCLCHECKGOTO(ncclInitKernelsForDevice(cudaArch, maxSharedMem, &maxLocalSizeBytes), res, fail);

cudaArch = 100 * archMajor + 10 * archMinorCopy

4. Then, depending on whether it is normal initialization or split/shrink/grow, take different bootstrap paths:

📎 src/init.cc:2136-2191

cpp
if (job->parent && !job->isGrow) {
    // SPLIT/SHRINK: use bootstrapSplit
    ...
    NCCLCHECKGOTO(bootstrapSplit(comm->commHash, comm, job->parent, job->color, job->key, parentRanks), res, fail);
} else {
    // GROW or NORMAL INIT: use bootstrapInit
    ...
    NCCLCHECKGOTO(bootstrapInit(job->nId, (struct ncclBootstrapHandle*)job->commId, comm, job->parent), res, fail);
}

5. Finally callinitTransportsRank, which is the heaviest function in the entire initialization (about 800 lines). Internally it performs two AllGathers:

  • AllGather1: exchangencclPeerInfo(each rank's device information, host hash, pid hash, GPU UUID, etc.):

📎 src/init.cc:1236-1239

cpp
NCCLCHECKGOTO(ncclCalloc(&comm->peerInfo, nranks + 1), ret, fail); // Extra rank to represent CollNet root
NCCLCHECKGOTO(fillInfo(comm, comm->peerInfo + rank, comm->commHash), ret, fail);
NCCLCHECKGOTO(bootstrapAllGather(comm->bootstrap, comm->peerInfo, sizeof(struct ncclPeerInfo)), ret, fail);
COMPILER_ATOMIC_STORE(&comm->peerInfoValid, true, std::memory_order_release);

Notenranks + 1this allocation—the extra slot is for the CollNet root.peerInfoValidStore with release semantics to ensure that when other threads see this flag, the contents of peerInfo are already visible.

  • AllGather3: exchange topology computation results (the ring/tree structure, bandwidth, channel count, etc. computed by each rank), then take theminimum valueacross all ranks to align:

📎 src/init.cc:1687-1703

cpp
for (int i = 0; i < nranks; i++) {
    allTopoRanks[i] = &allGather3Data[i].topoRanks;
    // Make sure we align all ranks so that the tuning is consistent across ranks
    for (int a = 0; a < NCCL_NUM_ALGORITHMS; a++) {
        graphs[a]->nChannels = std::min(allGather3Data[i].graphInfo[a].nChannels, graphs[a]->nChannels);
        graphs[a]->sameChannels = std::min(allGather3Data[i].graphInfo[a].sameChannels, graphs[a]->sameChannels);
        graphs[a]->bwIntra = std::min(allGather3Data[i].graphInfo[a].bwIntra, graphs[a]->bwIntra);
        graphs[a]->bwInter = std::min(allGather3Data[i].graphInfo[a].bwInter, graphs[a]->bwInter);
        graphs[a]->typeIntra = std::max(allGather3Data[i].graphInfo[a].typeIntra, graphs[a]->typeIntra);
        graphs[a]->typeInter = std::max(allGather3Data[i].graphInfo[a].typeInter, graphs[a]->typeInter);
        graphs[a]->crossNic = std::max(allGather3Data[i].graphInfo[a].crossNic, graphs[a]->crossNic);
    }
    ...
}

Bandwidth takes the min, type takes the max—this is the "barrel principle": the performance of the entire communication domain is determined by the slowest rank. If not aligned, different ranks may compute different algorithm choices, leading to communication deadlock.

Initialization Flowchart

mermaid
flowchart TD
    api["ncclCommInitRank()"] --> env["ncclInitEnv()"]
    env --> grp["ncclGroupStartInternal()"]
    grp --> dev["ncclCommInitRankDev()"]
    dev --> alloc["ncclCalloc(comm) + parseCommConfig()"]
    alloc --> launch{"ncclParamEnqueueRearchEnable()?"}
    launch -->|是| mgmt["ncclMgmtTaskEnqueue(ncclCommInitRankFunc)"]
    launch -->|否| async["ncclAsyncLaunch(ncclCommInitRankFunc)"]
    mgmt --> func["ncclCommInitRankFunc()"]
    async --> func
    func --> kernels["ncclInitKernelsForDevice(cudaArch)"]
    kernels --> branch{"job->parent && !job->isGrow?"}
    branch -->|是 split/shrink| split["bootstrapSplit()"]
    branch -->|否 grow/normal| init["bootstrapInit()"]
    split --> transports["initTransportsRank()"]
    init --> transports
    transports --> ag1["bootstrapAllGather(peerInfo)"]
    ag1 --> topo["ncclTopoGetSystem() + ncclTopoComputePaths()"]
    topo --> graphs["ncclTopoCompute(ringGraph/treeGraph/nvlsGraph)"]
    graphs --> ag3["bootstrapAllGather(allGather3Data)"]
    ag3 --> align["min/max 对齐所有 rank 的图参数"]
    align --> connect["setupChannel() + ncclTransportRingConnect()"]
    connect --> devcomm["devCommSetup()"]
    devcomm --> done["initState = ncclSuccess"]

Design Considerations and Pitfalls

Why does initialization need to be asynchronous?Because multi-rank initialization requires cross-process synchronization (bootstrap), and if executed synchronously it would block the calling thread. After making it asynchronous, users can initialize multiple communication domains simultaneously within a group, advancing them in parallel.

Pitfalls:initTransportsRankThere is an intra-node barrier at the end:

📎 src/init.cc:1968-1971

cpp
/* Local intra-node barrier */
NCCLCHECKGOTO(bootstrapIntraNodeBarrier(comm->bootstrap, comm->localRankToRank, comm->localRank, comm->localRanks, comm->localRankToRank[0]), ret, fail);

This barrier ensures that all ranks on the same machine have completed resource allocation before continuing. If some rank is stuck indevCommSetup(e.g., out of GPU memory), other ranks will wait here forever. When encountering "initialization hang" in production, the first thing to check is whether some rank'sdevCommSetupfailed.

II. Task Enqueueing: From API Call to Internal Task Object

Intuitive Model

When a user callsncclAllReduceit's like ordering food at a restaurant.ncclEnqueueCheckis the waiter, which translates your order into a "work order" (ncclTaskColl) that the kitchen can understand, and puts it intocomm->plannerthis "order pool".Without this layer, NCCL would not be able to merge multiple calls into a single kernel launch—lighting the stove separately for each order is extremely inefficient.

Data Structures and Memory Layout

The core of task enqueueing isncclKernelPlanner, which hangs offcomm->planner. Key fields include:

  • collSorter: a collection of collective communication tasks sorted by traffic size
  • collTaskQueue: the final sorted task queue
  • peers[]: each peer's send/recv queue (for P2P)
  • wipPlan: the kernel plan being constructed

The key fields of the task objectncclTaskCollare filled incollTaskAppend:

📎 src/enqueue/enqueue.cc:2800-2847

cpp
struct ncclTaskColl* t = ncclMemoryPoolAlloc<struct ncclTaskColl>(&comm->memPool_ncclTaskColl, &comm->memPermanent);
t->func = info->coll;
t->sendbuff = info->sendbuff;
t->recvbuff = info->recvbuff;
t->count = info->count;
t->root = info->root;
t->datatype = info->datatype;
size_t elementSize = ncclTypeSize(t->datatype);
if (t->func == ncclFuncAllGather || t->func == ncclFuncBroadcast) {
    t->count *= elementSize;
    t->datatype = ncclInt8;
    elementSize = 1;
}
t->trafficBytes = t->count * elementSize * ncclFuncTrafficPerByte(t->func, comm->nRanks);
...
t->aggIsolate = ncclCollConfigNeedAggIsolate(&info->collConfig) || info->collConfig.CTAPolicy != comm->config.CTAPolicy;
NCCL_CONFIG_SET(t, minCTAs, ncclParamMinCTAs(), info->collConfig.minCTAs, comm->config.minCTAs, 1, MAXCHANNELS);
NCCL_CONFIG_SET(t, maxCTAs, ncclParamMaxCTAs(), (std::min(info->collConfig.maxCTAs, comm->config.maxCTAs)), comm->config.maxCTAs, 1, MAXCHANNELS);
...
planner->nTasksColl += 1;
ncclTaskCollSorterInsert(&planner->collSorter, t, t->trafficBytes);

Note a few details:

1. Special handling for AllGather/Broadcast: multiply count by the element size and change datatype toncclInt8. This is because the semantics of these two operations is "moving bytes" and does not need to care about the original type.

2. trafficBytesComputation of:ncclFuncTrafficPerBytereturns how many times each byte needs to be transferred. AllReduce returns 2 (reduce + broadcast), AllGather returns nRanks:

📎 src/enqueue/enqueue.cc:123-134

cpp
static inline int ncclFuncTrafficPerByte(ncclFunc_t func, int nRanks) {
  switch (func) {
  case ncclFuncAllReduce:
    return 2;
  case ncclFuncAllGather:
    return nRanks;
  case ncclFuncReduceScatter:
    return nRanks;
  default:
    return 1;
  }
}

3. NCCL_CONFIG_SETMacro: this is "env > per-call > comm" three-level configuration resolution. Environment variables have the highest priority, followed by the per-call config, and finally the communication domain-level default value.

Step-by-Step: The Enqueue Path of ncclAllReduce

1. ncclEnqueueCheckFirst perform communication domain validation and group entry:

📎 src/enqueue/enqueue.cc:3478-3495

cpp
ncclResult_t ncclEnqueueCheck(struct ncclInfo* info) {
  ncclResult_t ret = CommCheck(info->comm, info->opName, "comm");
  if (ret != ncclSuccess) return ncclGroupErrCheck(ret);
  if (info->comm->revokedFlag) {
    WARN("%s: communicator was revoked", info->opName);
    return ncclGroupErrCheck(ncclInvalidUsage);
  }
  ...
  NCCLCHECK(ncclGroupStartInternal());
  ret = ncclSuccess;
  int devOld = -1;
  NCCLCHECKGOTO(ncclCommEnsureReady(info->comm), ret, fail);

2. Then calltaskAppend, which dispatches based on the operation type:

📎 src/enqueue/enqueue.cc:3337-3348

cpp
static ncclResult_t taskAppend(struct ncclComm* comm, struct ncclInfo* info) {
  ncclFunc_t collAPI = info->coll;
  bool hasLaunchCompletionEvent = ncclInfoHasLaunchCompletionEvent(info);

  if (ncclParamEnqueueRearchEnable()) {
    NCCLCHECK(rawTaskAppend(comm, info));
  } else if (info->coll == ncclFuncSend || info->coll == ncclFuncRecv) {
    NCCLCHECK(p2pTaskAppend(comm, info, info->coll, collAPI, (void*)info->recvbuff, info->count, info->datatype, info->root, true));
  } else if (info->coll == ncclFuncPutSignal || info->coll == ncclFuncSignal || info->coll == ncclFuncWaitSignal) {
    NCCLCHECK(rmaTaskAppend(comm, info));
  } else {
    ...
  }
}

For AllReduce, it goes through the finalelsebranch, ultimately callingcollTaskAppend。

3. collTaskAppendto insert the task intocollSorter, sorted bytrafficBytes. The purpose of sorting is to let the scheduler prioritize large tasks and avoid small tasks fragmenting channel resources.

Task Enqueueing Data Flow

mermaid
flowchart LR
    api["ncclAllReduce()"] --> info["ncclInfo 填充"]
    info --> enq["ncclEnqueueCheck()"]
    enq --> check["CommCheck + ncclCommEnsureReady()"]
    check --> append["taskAppend()"]
    append --> coll["collTaskAppend()"]
    coll --> task["ncclTaskColl 分配"]
    task --> sorter["ncclTaskCollSorterInsert(collSorter)"]
    sorter --> prepare["ncclPrepareTasks()"]
    prepare --> algo["ncclGetAlgoInfo() 选算法"]
    algo --> schedule["scheduleCollTasksToPlan()"]
    schedule --> plan["ncclKernelPlan"]

Design Considerations and Pitfalls

Why usencclMemoryPoolAllocinstead ofmalloc?Because task objects have a short lifecycle and are allocated frequently. The memory pool avoids the system call overhead ofmalloc/freeeach time. Note that the second parameter ofncclMemoryPoolAllocis&comm->memPermanent—this means task objects are released uniformly when the communication domain is destroyed, rather than each task being released individually.

Pitfalls:ncclPrepareTasksThere is an "aggregation" logic in

📎 src/enqueue/enqueue.cc:506-512

cpp
// We aggregate operations that are within 4X size of each other.
while (aggEnd != nullptr && aggEnd->trafficBytes < 4 * aggBeg->trafficBytes && !aggBeg->aggIsolate && !aggEnd->aggIsolate) {
    agg.count += aggEnd->count;
    agg.trafficBytes += aggEnd->trafficBytes;
    aggEnd = aggEnd->next;
}

CopyaggIsolateThis aggregation is to make algorithm selection more stable—if each small task selects an algorithm individually, it may select a bunch of different algorithms, causing kernel fragmentation. But the

flag prevents aggregation, used for those tasks that "must be scheduled individually" (such as those with per-call config).

III. Algorithm Selection: How the Cost Model Picks the Optimal Solution

Intuitive ModelAlgorithm selection is like navigation software choosing a route. NCCL's "cost model" (tuning module) estimates the time cost of each algorithm/protocol combination under a given message size and topology, then picks the fastest one.。

Without a cost model, NCCL could only hardcode a single set of algorithms, wasting bandwidth on small messages and wasting latency on large messages

Data Structures and Memory LayoutncclGetAlgoInfo:

📎 src/enqueue/enqueue.cc:2159-2185

cpp
ncclResult_t ncclGetAlgoInfo(struct ncclComm* comm, struct ncclTaskColl* info, int collNetSupport, int nvlsSupport,
                             int numPipeOps, ncclSimInfo_t* simInfo) {
  size_t elementSize = ncclTypeSize(info->datatype);
  size_t nBytes = elementSize * ncclFuncMaxSendRecvCount(info->func, comm->nRanks, info->count);
  info->algorithm = NCCL_ALGO_UNDEF;
  info->protocol = NCCL_PROTO_UNDEF;
  struct ncclTuningInput_t input;
  input.comm = comm;
  input.tuningMask = NCCL_TUNING_MASK_GENERAL_KERNELS;
  uint64_t effAlgMask = comm->tuningContext.forced[info->func] ? 0 : info->algMask;
  if (effAlgMask != 0) {
    input.tuningMask = effAlgMask & NCCL_TUNING_MASK_GENERAL_KERNELS;
  }
  input.CTAPolicy = info->CTAPolicy;
  input.func = info->func;
  input.redOp = info->opHost;
  input.devRedOp = info->opDev.op;
  input.datatype = info->datatype;
  input.nBytes = nBytes;
  input.numPipeOps = numPipeOps;
  input.collNetSupport = collNetSupport;
  input.nvlsSupport = nvlsSupport;
  input.count = info->count;
  NCCLCHECK(ncclGetRegBuff(comm, info, &input.regBuff));
  ...
}

CopyeffAlgMaskNote the logic ofcomm->tuningContext.forced[info->func]: if an environment variable forces a specific algorithm (algMaskis non-zero), then the user's

is ignored and the environment variable's is used. This reflects the "env > per-call" priority.ncclTuningComputeThen call

📎 src/enqueue/enqueue.cc:2213-2224

cpp
} else {
    NCCLCHECK(ncclTuningCompute(&input, &bestTuning));
}
INFO(NCCL_TUNING, "Best tuning, algorithm, %s, protocol, %s", ncclAlgoToString(bestTuning.algo), ncclProtoToString(bestTuning.proto));
info->algorithm = bestTuning.algo;
info->protocol = bestTuning.proto;
info->nWarps = bestTuning.nWarps;
if (simInfo) simInfo->estimatedTime = bestTuning.timeUs;
TRACE(NCCL_COLL, "%ld Bytes -> Algo %d proto %d time %f", nBytes, info->algorithm, info->protocol, bestTuning.timeUs);
info->nMaxChannels = bestTuning.maxChannels == 0 ? info->nMaxChannels : bestTuning.maxChannels;

Step-by-Step: Algorithm Selection for a Single AllReduce

Assume 8 GPUs on a single node, message size 1MB, AllReduce:

1. nBytes = 1MB,numPipeOpsis the number of tasks already in the current plan.

2. collNetSupportandnvlsSupportdetermined byncclGetCollNetSupportandncclNvlsTransportEnabled.

3. ncclTuningComputeIterate over all available (algo, proto) combinations and estimate time using the cost model.

4. For a 1MB single-node scenario, NVLS or Tree+LL128 typically wins.

5. Write the result back toinfo->algorithm、info->protocol、info->nWarps。

Algorithm Selection Decision Diagram

mermaid
flowchart TD
    start["ncclGetAlgoInfo()"] --> nbytes["计算 nBytes = elementSize * count"]
    nbytes --> forced{"comm->tuningContext.forced[func]?"}
    forced -->|是| envMask["effAlgMask = 0, 用环境变量强制"]
    forced -->|否| userMask{"info->algMask != 0?"}
    userMask -->|是| useUser["tuningMask = algMask"]
    userMask -->|否| full["tuningMask = GENERAL_KERNELS"]
    envMask --> compute["ncclTuningCompute(input, bestTuning)"]
    useUser --> compute
    full --> compute
    compute --> result{"bestTuning.algo == UNDEF?"}
    result -->|是| fallback["重算全量菜单"]
    fallback --> force{"forceAlgSelection?"}
    force -->|是| err["返回 ncclInvalidArgument"]
    force -->|否| auto["回退到自动选择"]
    result -->|否| assign["info->algorithm = bestTuning.algo"]
    auto --> assign
    assign --> done["返回 ncclSuccess"]

Design Considerations and Pitfalls

Why must algorithm selection be "cross-rank aligned"?Because if different ranks choose different algorithms, the communication patterns won't match, causing deadlock. SoinitTransportsRankuses min/max to align all graph parameters, ensuring every rank's cost model input is consistent.

Pitfalls:ncclGetAlgoInfoThere is a "recompute" logic — if the user specifiesalgMaskbut no algorithm matches, it first silently recomputes the full menu, then determines whether it's a hard error or soft fallback:

📎 src/enqueue/enqueue.cc:2192-2208

cpp
NOWARN(ncclTuningCompute(&input, &bestTuning), NCCL_TUNING);
if (bestTuning.algo == NCCL_ALGO_UNDEF) {
    input.tuningMask = NCCL_TUNING_MASK_GENERAL_KERNELS;
    bestTuning = NCCL_TUNING_RESULT_INIT;
    bestTuning.maxChannels = 0;
    NCCLCHECK(ncclTuningCompute(&input, &bestTuning));
    if (info->forceAlgSelection) {
        WARN("algSelection: no algorithm in the selected set is available for %s", ncclFuncToString(info->func));
        return ncclInvalidArgument;
    }
    INFO(NCCL_TUNING, "algSelection: selected set unavailable for %s; falling back to automatic selection", ncclFuncToString(info->func));
}

NOWARNThe macro temporarily suppresses warnings, because "no algorithm matches" may be a normal situation (the user-selected set is indeed unavailable). Only whenforceAlgSelectionis true does it report an error.

IV. Task Scheduling and Kernel Plan Construction

Intuitive Model

Task scheduling is like distributing a bunch of orders across several assembly lines.scheduleCollTasksToPlandetermines how many channels each task uses and how much data each channel processes, ultimately generating ancclKernelPlan— this is the "work order" to be passed to the GPU.

Data Structures and Memory Layout

ncclKernelPlanCore fields of

  • channelMask: which channels this plan uses (bitmap)
  • workBytes: total bytes of all work structures
  • nWorkBatches: number of work batches
  • kernelArgs: kernel launch parameters
  • workStorageType: where work data is stored (args/fifo/persistent)

finishPlandetermines the storage location of work data:

📎 src/enqueue/enqueue.cc:244-255

cpp
// If we can fit everything into the kernel args we do so.
if (sizeof(ncclDevKernelArgs) + batchBytes + workBytes <= comm->workArgsBytes) {
    plan->workStorageType = ncclDevWorkStorageTypeArgs;
}
plan->kernelArgsSize = sizeof(struct ncclDevKernelArgs) + batchBytes;
plan->kernelArgsSize += (plan->workStorageType == ncclDevWorkStorageTypeArgs) ? workBytes : 0;
plan->kernelArgsSize = alignUp(plan->kernelArgsSize, 16);
plan->kernelArgs = (struct ncclDevKernelArgs*)ncclMemoryStackAlloc(&comm->memScoped, plan->kernelArgsSize, /*align=*/16);
plan->kernelArgs->comm = comm->devComm;
plan->kernelArgs->channelMask = plan->channelMask;
plan->kernelArgs->workStorageType = plan->workStorageType;

Trade-offs of the three storage types:

  • Args: fastest, but kernel parameter size is limited (typically 4KB)
  • Fifo: ring buffer, suitable for medium sizes
  • Persistent: separate device memory allocation, suitable for CUDA Graph scenarios

Step-by-Step: Channel Allocation in scheduleCollTasksToPlan

1. First estimate how many tasks this plan can hold:

📎 src/enqueue/enqueue.cc:654-687

cpp
do {
    size_t workBytes = 0;
    struct ncclTaskColl* task = ncclIntruQueueHead(&planner->collTaskQueue);
    struct ncclWorkList* workNode = ncclIntruQueueHead(&planner->collWorkQueue);
    while (task != nullptr) {
        int nBatches = divUp(nPlanColls, 4); // Rough guess: 4 colls per batch.
        if (!ncclTestBudget(budget, nBatches, workBytes + workNode->size)) goto plan_full;
        bool taskAggIsolate = task->aggIsolate;
        if (taskAggIsolate && nPlanColls > 0) goto plan_full;
        nPlanColls += 1;
        workBytes += workNode->size;
        int kind = 2 * task->isCollnet + task->isNvls;
        trafficBytes[kind] += std::max(MinTrafficPerChannel, task->trafficBytes);
        ...
    }
plan_full:;
} while (0);

2. Then allocate channels to tasks by traffic. For non-CollNet tasks, split using "cell" as the unit:

📎 src/enqueue/enqueue.cc:742-759

cpp
int trafficPerByte = ncclFuncTrafficPerByte(task->func, comm->nRanks);
if (task->protocol == NCCL_PROTO_LL) trafficPerByte *= 4;
size_t cellSize = divUp(divUp(MinTrafficPerChannel, (size_t)trafficPerByte), 16) * 16;
int elementsPerCell = cellSize / elementSize;
size_t cells = divUp(task->count * elementSize, cellSize);
size_t trafficPerElement = elementSize * trafficPerByte;
size_t trafficPerCell = cellSize * trafficPerByte;
size_t cellsPerChannel = std::min(cells, divUp(trafficPerChannel, trafficPerCell));
size_t cellsLo;
if (channelId + 1 == nMaxChannels[kind]) {
    cellsLo = cells;
} else {
    cellsLo = std::min(cells, divUp((trafficPerChannel - currentTraffic), trafficPerCell));
}
int nMidChannels = (cells - cellsLo) / cellsPerChannel;
size_t cellsHi = (cells - cellsLo) % cellsPerChannel;
int nChannels = (cellsLo != 0 ? 1 : 0) + nMidChannels + (cellsHi != 0 ? 1 : 0);

This code splits data into "low/mid/high" three segments:countLo、countMid、countHi. The low and high segments are boundary channels, and the mid segment is the middle channel. This split is to make the data volume processed by each channel as even as possible.

3. Finally callcalcCollChunkingto compute the chunk size for each channel:

📎 src/enqueue/enqueue.cc:2228-2275

cpp
static ncclResult_t calcCollChunking(struct ncclComm* comm, struct ncclTaskColl* info, int nChannels, size_t nBytes,
                                     uint32_t* outChunkSize, uint32_t* outDirectFlags, struct ncclProxyOp* proxyOp) {
  ncclPattern_t pattern;
  size_t grainSize = ncclProtoGrainSize(info->protocol);
  switch (info->func) {
  case ncclFuncAllReduce:
    pattern = info->algorithm == NCCL_ALGO_NVLS           ? ncclPatternNvls :
              info->algorithm == NCCL_ALGO_NVLS_TREE      ? ncclPatternNvlsTree :
              info->algorithm == NCCL_ALGO_COLLNET_DIRECT ? ncclPatternCollnetDirect :
              info->algorithm == NCCL_ALGO_COLLNET_CHAIN  ? ncclPatternCollnetChain :
              info->algorithm == NCCL_ALGO_TREE           ? ncclPatternTreeUpDown :
                                                            ncclPatternRingTwice;
    break;
  ...
  }
  int stepSize = comm->buffSizes[info->protocol] / NCCL_STEPS;
  int chunkSteps = (info->protocol == NCCL_PROTO_SIMPLE && info->algorithm == NCCL_ALGO_RING) ? info->chunkSteps : 1;
  int sliceSteps = (info->protocol == NCCL_PROTO_SIMPLE && info->algorithm == NCCL_ALGO_RING) ? info->sliceSteps : 1;
  int chunkSize = stepSize * chunkSteps;
  if (info->protocol == NCCL_PROTO_LL) chunkSize /= 2;
  if (info->protocol == NCCL_PROTO_LL128) chunkSize = (chunkSize / NCCL_LL128_LINEELEMS) * NCCL_LL128_DATAELEMS;
  ...
}

Scheduling Flow Diagram

mermaid
flowchart TD
    prep["ncclPrepareTasks()"] --> sort["collSorter 按 trafficBytes 排序"]
    sort --> agg["按 (fn,op,ty) 聚合任务"]
    agg --> algo["ncclGetAlgoInfo() 选算法"]
    algo --> bins["按 isCollnet/isNvls 分箱"]
    bins --> sched["scheduleCollTasksToPlan()"]
    sched --> budget{"ncclTestBudget()?"}
    budget -->|否| full["plan_full: 停止添加"]
    budget -->|是| kind{"task->isCollnet?"}
    kind -->|是| collnet["calcCollChunking + 全通道分配"]
    kind -->|否| cells["cell 切分: countLo/Mid/Hi"]
    collnet --> batch["ncclAddWorkBatchToPlan()"]
    cells --> batch
    batch --> proxy["ncclAddProxyOpIfNeeded()"]
    proxy --> finish["finishPlan()"]
    finish --> storage{"workBytes 能放进 args?"}
    storage -->|是| args["ncclDevWorkStorageTypeArgs"]
    storage -->|否| fifo["ncclDevWorkStorageTypeFifo"]

Design Considerations and Pitfalls

Why are CollNet tasks handled separately?Because CollNet uses network switches for reduction, and the channel allocation logic is completely different from regular ring/tree. CollNet tasks directly occupy all available channels, while regular tasks need to be split by traffic.

Pitfalls:ncclTestBudgetThe estimation uses a rough formulanBatches = divUp(nPlanColls, 4)— assuming one batch is produced every 4 collective operations. This estimate may be inaccurate, so there's a precise check afterward:

📎 src/enqueue/enqueue.cc:711-714

cpp
// Ensure room for worst case of one new batch per channel
if (!ncclTestBudget(budget, plan->nWorkBatches + nChannels, plan->workBytes + workNode->size)) {
    return ncclSuccess;
}

If the precise check fails, return directly (without error), letting the upper layer open a new plan.

V. Kernel Launch and Device-Side Execution

Intuitive Model

Kernel launch is like handing work orders to the factory.ncclLaunchKerneltranslatesncclKernelPlaninto CUDA kernel launch parameters, then callscuLaunchKernelEx. After the device-side kernel receives the work order, it executes data movement according to the algorithm.

Data Structures and Memory Layout

ncclLaunchKernelKey steps of

📎 src/enqueue/enqueue.cc:1886-1909

cpp
ncclResult_t ncclLaunchKernel(struct ncclComm* comm, struct ncclKernelPlan* plan) {
  ncclResult_t ret = ncclSuccess;
  struct ncclKernelPlanner* planner = &comm->planner;
  int nChannels = countOneBits(plan->channelMask);
  void* sym = plan->kernelFn;
  dim3 grid = {(unsigned)nChannels, 1, 1};
  dim3 block = {(unsigned)plan->threadPerBlock, 1, 1};
  int smem = plan->isSymColl ? plan->kernelDynSmem : ncclShmemDynamicSize(comm->cudaArch);
  cudaStream_t launchStream = planner->streams->stream;
  ...
  void* extra[] = {CU_LAUNCH_PARAM_BUFFER_POINTER, plan->kernelArgs, CU_LAUNCH_PARAM_BUFFER_SIZE, &plan->kernelArgsSize, CU_LAUNCH_PARAM_END};
  ...
  CUfunction fn;
  CUDACHECKGOTO(cudaGetFuncBySymbol(&fn, sym), ret, do_return);

Notegrid.x = nChannels— one block per channel.block.x = plan->threadPerBlock— the number of threads per block is determined by the task.

Step-by-Step: From Plan to Kernel Launch

1. First calluploadWorkto write work data to the target location (args/fifo/persistent):

📎 src/enqueue/enqueue.cc:1365-1407

cpp
static ncclResult_t uploadWork(struct ncclComm* comm, struct ncclKernelPlan* plan) {
  if (plan->isSymColl || plan->isCeColl || plan->isRma) return ncclSuccess;
  size_t workBytes = plan->workBytes;
  size_t batchBytes = plan->nWorkBatches * sizeof(struct ncclDevWorkBatch);
  void* fifoBufHost;
  uint32_t fifoCursor, fifoMask;
  switch (plan->workStorageType) {
  case ncclDevWorkStorageTypeArgs:
    plan->kernelArgs->workBuf = nullptr;
    fifoBufHost = (void*)plan->kernelArgs;
    fifoCursor = sizeof(ncclDevKernelArgs) + batchBytes;
    fifoMask = ~0u;
    break;
  case ncclDevWorkStorageTypeFifo:
    fifoBufHost = comm->workFifoBuf;
    fifoCursor = comm->workFifoProduced;
    fifoMask = comm->workFifoBytes - 1;
    NCCLCHECK(waitWorkFifoAvailable(comm, fifoCursor + workBytes));
    plan->kernelArgs->workBuf = comm->workFifoBufDev;
    break;
  ...
  }
}

2. Then construct CUDA launch attributes. For sm90+, cluster dimensions are set:

📎 src/enqueue/enqueue.cc:1929-1936

cpp
if (clusterSize) {
    // Grid dimension must be divisible by clusterSize
    if (grid.x % clusterSize) clusterSize = 1;
    launchAttrs[attrs].id = CU_LAUNCH_ATTRIBUTE_CLUSTER_DIMENSION;
    launchAttrs[attrs++].value.clusterDim = {clusterSize, 1, 1};
    launchAttrs[attrs].id = CU_LAUNCH_ATTRIBUTE_CLUSTER_SCHEDULING_POLICY_PREFERENCE;
    launchAttrs[attrs++].value.clusterSchedulingPolicyPreference = CU_CLUSTER_SCHEDULING_POLICY_SPREAD;
}

3. Finally callcuLaunchKernelEx:

📎 src/enqueue/enqueue.cc:1992

cpp
CUCHECKGOTO(cuLaunchKernelEx(&launchConfig, fn, nullptr, extra), ret, do_return);

Device Side: Execution of runRing

After the device-side kernel receives the work order, it calls the correspondingRunWorkCollspecialization based on the algorithm. Taking Ring AllReduce as an example:

📎 src/device/all_reduce.h:14-83

cpp
template <typename T, typename RedOp, typename Proto>
__device__ __forceinline__ void runRing(int tid, int nthreads, struct ncclDevWorkColl* work) {
  ncclRing* ring = &ncclShmem.channel.ring;
  int ringIx = ring->index;
  const int nranks = ncclShmem.comm.nRanks;
  ssize_t gridOffset;
  ssize_t channelCount;
  ssize_t chunkCount;
  ncclCollCbdPart(work, ncclShmem.channelId, Proto::Id, sizeof(T), (ssize_t*)nullptr, &gridOffset, &channelCount, &chunkCount);
  const ssize_t loopCount = nranks * chunkCount;
  ...
  Primitives<T, RedOp, FanSymmetric<1>, 1, Proto, 0> prims(tid, nthreads, &ring->prev, &ring->next, work->sendbuff, work->recvbuff, work->redOpArg, 0, 0, 0, work);

  for (ssize_t elemOffset = 0; elemOffset < channelCount; elemOffset += loopCount) {
    ssize_t remCount = channelCount - elemOffset;
    ssize_t chunkOffset;
    if (remCount < loopCount) chunkCount = alignUp(divUp(remCount, nranks), 16 / sizeof(T));
    auto modRanks = [&] __device__(int r) -> int { return r - (r >= nranks ? nranks : 0); };

    // step 0: push data to next GPU
    chunk = modRanks(ringIx + nranks - 1);
    chunkOffset = chunk * chunkCount;
    offset = gridOffset + elemOffset + chunkOffset;
    nelem = (int)min(chunkCount, remCount - chunkOffset);
    prims.directSend(offset, offset, nelem);

    // k-2 steps: reduce and copy to next GPU
    for (int j = 2; j < nranks; ++j) {
      chunk = modRanks(ringIx + nranks - j);
      chunkOffset = chunk * chunkCount;
      offset = gridOffset + elemOffset + chunkOffset;
      nelem = (int)min(chunkCount, remCount - chunkOffset);
      prims.directRecvReduceDirectSend(offset, offset, nelem);
    }

    // step k-1: reduce this buffer and data, which will produce the final result
    chunk = ringIx + 0;
    chunkOffset = chunk * chunkCount;
    offset = gridOffset + elemOffset + chunkOffset;
    nelem = (int)min(chunkCount, remCount - chunkOffset);
    prims.directRecvReduceCopyDirectSend(offset, offset, nelem, /*postOp=*/true);

    // k-2 steps: copy to next GPU
    for (int j = 1; j < nranks - 1; ++j) {
      chunk = modRanks(ringIx + nranks - j);
      chunkOffset = chunk * chunkCount;
      offset = gridOffset + elemOffset + chunkOffset;
      nelem = (int)min(chunkCount, remCount - chunkOffset);
      prims.directRecvCopyDirectSend(offset, offset, nelem);
    }

    // Make final copy from buffer to dest.
    chunk = modRanks(ringIx + 1);
    chunkOffset = chunk * chunkCount;
    offset = gridOffset + elemOffset + chunkOffset;
    nelem = (int)min(chunkCount, remCount - chunkOffset);
    prims.directRecv(offset, nelem);
  }
}

The classic two phases of Ring AllReduce:

  • Reduce-Scatter Phase(first nranks-1 steps): each rank sends its own data to the next, while receiving the previous rank's data and reducing.
  • AllGather Phase(last nranks-1 steps): propagate the reduced result along the ring.

modRanksThis lambda handles ring index wraparound: whenr >= nranks, subtract nranks.

Kernel Launch Timing Diagram

mermaid
sequenceDiagram
    participant Host as Host 线程
    participant Plan as ncclKernelPlan
    participant CUDA as CUDA Driver
    participant Kernel as GPU Kernel
    participant Proxy as Proxy 线程

    Host->>Plan: ncclLaunchPrepare()
    Plan->>Plan: scheduleCollTasksToPlan()
    Plan->>Plan: finishPlan() 分配 kernelArgs
    Host->>Plan: ncclLaunchKernelBefore_NoUncapturedCuda()
    Plan->>Plan: uploadWork() 写 work 数据
    Host->>CUDA: cuLaunchKernelEx(fn, grid, block, smem)
    CUDA->>Kernel: 启动 nChannels 个 block
    Kernel->>Kernel: runRing() 执行 Ring AllReduce
    Host->>Plan: ncclLaunchKernelAfter_NoCuda()
    Plan->>Proxy: hostStreamPlanTask() + uploadProxyOps()
    Proxy->>Proxy: ncclProxyStart() 推进网络 I/O
    Kernel-->>Host: kernel 完成
    Host->>Plan: ncclLaunchFinish()
    Plan->>Plan: reclaimPlan() 释放资源

Design Considerations and Pitfalls

Why usecuLaunchKernelExinstead ofcudaLaunchKernel?Because launch attributes need to be set (cluster dimensions, mem sync domain, launch completion event). These attributes are only supported in CUDA 12.0+.

Pitfalls:uploadWorkThe handling of persistent mode here is very complex—it needs to allocate GPU memory, copy data, record events, and also work correctly under CUDA Graph capture mode:

📎 src/enqueue/enqueue.cc:1445-1478

cpp
CUDACHECKGOTO(cudaThreadExchangeStreamCaptureMode(&mode), result, fail);
NCCLCHECKGOTO(ncclStrongStreamAcquire(ncclCudaGraphNone(comm->config.graphUsageMode), &comm->sharedRes->deviceStream, /*concurrent=*/false, &deviceStream), result, fail);
if (comm->memPool) {
    CUDACHECKGOTO(cudaMallocAsync(&fifoBufDev, workBytes, comm->memPool, deviceStream), result, fail);
} else {
    CUDACHECKGOTO(cudaMalloc(&fifoBufDev, workBytes), result, fail);
}
plan->workBufPersistent = fifoBufDev;
plan->kernelArgs->workBuf = fifoBufDev;
CUDACHECKGOTO(cudaMemcpyAsync(fifoBufDev, fifoBufHost, workBytes, cudaMemcpyDefault, deviceStream), result, fail);
cudaEvent_t memcpyDone;
CUDACHECKGOTO(cudaEventCreateWithFlags(&memcpyDone, cudaEventDisableTiming), result, fail);
CUDACHECKGOTO(cudaEventRecord(memcpyDone, deviceStream), result, fail);

cudaThreadExchangeStreamCaptureModeis to temporarily switch to relaxed mode during capture mode, allowing GPU memory allocation. After the copy is complete, record the event, and later reclaim it throughncclCommPollEventCallbacks.

6. Production Pitfall Guide

Pitfall 1: Initialization hangs

Symptom:ncclCommInitRankgets stuck and does not return.

Troubleshooting: Check theNCCL_DEBUG=INFOlogs and find the last rank that printed. If all ranks printed "Init START" but not "Init COMPLETE", it means it is stuck ininitTransportsRank.

Common causes:

  • A certain rank'sdevCommSetupfailed (out of GPU memory, CUDA error)
  • bootstrap network is unreachable (firewall, port occupied)
  • Different ranks have inconsistent NCCL versions

Source code basis:initTransportsRankThe intra-node barrier at the end will wait for all local ranks:

📎 src/init.cc:1968-1971

cpp
/* Local intra-node barrier */
NCCLCHECKGOTO(bootstrapIntraNodeBarrier(comm->bootstrap, comm->localRankToRank, comm->localRank, comm->localRanks, comm->localRankToRank[0]), ret, fail);

Pitfall 2: work FIFO overflow

Symptom: after the kernel starts, it hangs, or reportsncclInternalError。

Cause:waitWorkFifoAvailableis waiting for FIFO space, but the consumer side (kernel) is not making progress.

📎 src/enqueue/enqueue.cc:1333-1349

cpp
static ncclResult_t waitWorkFifoAvailable(struct ncclComm* comm, uint32_t desiredProduced) {
  bool hasRoom = (desiredProduced - comm->workFifoConsumed) <= comm->workFifoBytes;
  if (!hasRoom) {
    while (true) {
      // Check abort flag to break deadlock when abort is signaled
      if (COMPILER_ATOMIC_LOAD(comm->abortFlag, std::memory_order_acquire)) {
        return ncclInternalError;
      }
      NCCLCHECK(ncclCommPollEventCallbacks(comm, /*waitSome=*/true));
      hasRoom = (desiredProduced - comm->workFifoConsumed) <= comm->workFifoBytes;
      if (hasRoom) break;
      std::this_thread::yield();
    }
  }
  return ncclSuccess;
}

Note the abort flag check—this is the only escape path. If abort is also not set, it will loop forever.

How to avoid: increaseNCCL_WORK_FIFO_BYTES, or reduce the number of operations in a single group.

Pitfall 3: CUDA Graph capture failure

Symptom: calling NCCL during CUDA Graph capture reports "operation not permitted".

Cause: certain CUDA operations cannot be performed in capture mode (such ascudaMalloc). NCCL usescudaThreadExchangeStreamCaptureModeto temporarily switch modes, but not all operations can be bypassed.

Source code basis:uploadWorkThe persistent branch of

📎 src/enqueue/enqueue.cc:1445

cpp
CUDACHECKGOTO(cudaThreadExchangeStreamCaptureMode(&mode), result, fail);

How to avoid: useNCCL_GRAPH_MIXING_SUPPORT=1to enable graph mixed mode, or preallocate the work buffer.

Chapter summary

In this chapter, we walked through the complete path of one AllReduce again:

1. Initialization:ncclCommInitRank → ncclCommInitRankFunc → initTransportsRank, establishing the communication domain, searching the topology, and aligning graph parameters.

2. Task enqueue:ncclEnqueueCheck → taskAppend → collTaskAppend, translating API calls intoncclTaskColl。

3. Algorithm selection:ncclGetAlgoInfo → ncclTuningCompute, using the cost model to choose the optimal (algo, proto).

4. Task scheduling:ncclPrepareTasks → scheduleCollTasksToPlan → finishPlan, assigning tasks to channels and generatingncclKernelPlan。

5. Kernel launch:ncclLaunchKernel → cuLaunchKernelEx, translating the plan into CUDA launch parameters.

6. Device-side execution:runRing / runTreeUpDown / runNvls, performing data movement according to the algorithm.

Chapter review and self-test

Q1: If the min/max alignment logic after AllGather3 ininitTransportsRank(L1690-L1698) is removed, in what scenarios would it cause communication deadlock? Why?

Reference analysis: This logic ensures that all ranks agree on parameters such asnChannels、bwIntra、bwInterfor each algorithm. If removed, each rank would compute the result using its own local topology. Consider a heterogeneous cluster: rank 0 is on an 8-GPU NVLink machine, and rank 8 is on a 4-GPU PCIe machine. Rank 0 computes that the ring has 8 channels, and rank 8 computes 4. When they execute Ring AllReduce, rank 0 will wait for rank 8 to send data on 8 channels, but rank

At this point, we have completed the review of the full path of one AllReduce. From initialization, topology search, algorithm selection, task enqueue, and kernel launch, to device-side execution and network transmission, each step corresponds to the in-depth analysis in the previous chapters. This path diagram is not only the skeleton for understanding NCCL, but also an index for troubleshooting: for initialization failures, check Chapters 3 and 4; for wrong algorithm selection, check Chapter 5; for task enqueue errors, check Chapters 6 and 7; for kernel launch failures, check Chapter 8; for device-side hangs, check Chapters 9 and 10; for network issues, check Chapters 12 and 13. As NCCL evolves toward programmable communication, GPU-initiated communication, and symmetric memory, this path will continue to extend—and you have already mastered the method to trace it.

To understand any complex project, all you really need is a good book

This book was automatically compiled by AiReadCode by scanning the official open-source repository, with real commit line numbers permanently anchored.

Star the GitHub repository ★ Browse more open-source books →
🇨🇳 Chinese · 🇺🇸 EN · 🇯🇵 Japanese · 🇰🇷 한국어 · 🌐 Traditional Chinese · 🇪🇸 ES · 🇩🇪 DE · 🇫🇷 FR · 🇧🇷 PT · 🇷🇺 RU