CHAPTER 01

Chapter 1: The Mental Model of Async: The Trio of Future, Waker, and Executor

Project: tokio-rs/tokio · Book Progress: Chapter 1 / 14 · Verification Status: FACT Line Numbers Genuinely Anchored

Async programming in Rust is not a library, but a language-level protocol. Tokio became a production-grade runtime not because it invented Future, but because it precisely implements the boundary conditions of every contract in this protocol. This chapter does not rush into Tokio's scheduler code, but first thoroughly explains the "trio"—Future, Waker, Executor—their responsibility boundaries and reverse control flow. Once you understand how these three interlock, the subsequent chapters on Runtime assembly, work-stealing scheduling, and I/O drivers have a foundation to stand on.

1.1 From Blocking to Pulling: Why Rust Chooses poll Over Callbacks

Intuitive Model

Imagine you order a dish at a restaurant that needs to be made fresh. Callback-style async (like early Node.js style) is equivalent to leaving your phone number, and the chef calls youproactively—control is in the chef's hands, and your code merely responds passively. Pull-style async (Rust's choice) is equivalent to getting a pickup ticket, and youdecide for yourselfwhen to go to the window and ask "is it ready?": if not, go do something else; if ready, pick it up.

This difference seems minor, but it determines the shape of the entire system. In the callback model, every async operation must carry a closure for "what to do when done," closures nest layer upon layer forming callback hell, and cancellation is extremely difficult—you cannot "withdraw" an already-registered callback. In the pull model, a Future is just a state machine,pollis a pure query action; if you don't advance it, it consumes no resources; cancellation is just drop, clean and neat.

The Core Contract of the Pull Model

TheFuturetrait defined by the Rust standard library has only two elements: apollmethod, and aOutputassociated type. Tokio does not redefine this trait, but directly reuses the standard library's implementation. This is clearly reflected in the source code:

rust
// tokio/src/future/mod.rs
cfg_not_trace! {
    cfg_rt! {
        pub(crate) use std::future::Future;
    }
}

📎 tokio/src/future/mod.rs:24-28

This code reveals an important fact: when thetracingfeature is not enabled, Tokio's internalFutureis an alias forstd::future::Future, with no wrapping whatsoever. Only whentracingis enabled is it replaced withInstrumentedFuture:

rust
cfg_trace! {
    mod trace;
    #[allow(unused_imports)]
    pub(crate) use trace::InstrumentedFuture as Future;
}

📎 tokio/src/future/mod.rs:18-22

[Design Inference and Architectural Trade-offs]

This "zero overhead by default, instrumentation on demand" design is Tokio's consistent philosophy: the core path introduces no additional abstraction layers, and observability is layered on as an optional feature.InstrumentedFutureThe existence of

shows that the Tokio team believes the instrumentation cost of tracing should not be borne by all users.

pollThree Implicit Constraints of the poll Contractfn poll(self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<Self::Output>The signature of the

method is Pin<&mut Self>. This signature hides three contracts; violating any one of them leads to undefined behavior or logical errors:

Contract One: Pin Guarantees Self-Reference Safety.means that once a Future is polled, its memory address cannot be moved again. This is because an async block compiles into a state machine containing self-references—local variables may hold references pointing to other fields within the same state machine. If movement were allowed, these references would dangle.pollContract Two: Pending Must Have Registered a Waker.Poll::PendingWhencx.waker()Obtain and save the Waker, or have already registered the Waker with some event source. Otherwise, the executor will never know when this Future can be polled again, causing the task to be permanently suspended.

Contract Three: After Ready, it should not be polled again.OncepollreturnsPoll::Ready, polling the same Future again is a logical error (although it will not cause UB, the behavior is undefined). The executor is responsible for no longer scheduling the task after receiving Ready.

Among these three contracts, Contract Two is the most error-prone place, and it is also the fundamental reason for the existence of Waker.

1.2 Waker: The Carrier of Reverse Control Flow

Intuitive Model

Waker is the "vibrating pager" the restaurant gives you. You do not need to stand at the window repeatedly asking "Is it ready?" - that would waste your time. You only need to hand the pager to the chef the first time you go to the window (register the Waker), and then go do other things with peace of mind. When the food is ready, the chef presses the button, the pager vibrates (callswake), and after you receive the signal, you go to the window to pick up the food (poll again).

Without Waker, the executor has only two choices: either busy-poll all tasks (wasting CPU), or never poll tasks that have already returned Pending (task starvation). Waker is the only mechanism that breaks this deadlock.

Waker's Memory Layout and Vtable Design

Waker is a standard library type, but its design directly influenced Tokio's task structure.Wakeris essentially a fat pointer: aRawWakerstruct containing a data pointer and a vtable pointer.

rust
// 标准库中的定义(非 Tokio 源码,此处为背景说明)
pub struct RawWaker {
    data: *const (),
    vtable: &'static RawWakerVTable,
}

pub struct RawWakerVTable {
    clone: unsafe fn(*const ()) -> RawWaker,
    wake: unsafe fn(*const ()),
    wake_by_ref: unsafe fn(*const ()),
    drop: unsafe fn(*const ()),
}
[Design Inference and Architectural Trade-offs]

The brilliance of this design lies in:Wakeritself does not care what "wake" specifically means. It is just a carrier of four function pointers. Tokio can provide a Waker whosewakefunction pushes the task back into the scheduling queue; while another runtime (such as thefuturescrate'sblock_on) can provide a completely different Waker implementation. This "data + vtable" pattern allows Waker to be passed between different runtimes without losing semantics.

wakeandwake_by_refThe difference is crucial:wakeconsumes ownership of the Waker (the Waker is dropped after the call), whilewake_by_refonly borrows. Executors usually implementwake_by_refas "mark the task as ready and enqueue it", whilewakeadditionally handles the decrement of the reference count on top of that. In Tokio's task structure, the Waker's data pointer points to the task's reference count header. Each clone increases the count, and drop decreases the count. When the count reaches zero, the task memory is released.

Complete Timing of Waking

The sequence diagram below shows the complete chain from initiation to being woken for a TCP read operation. Note how the Waker is passed all the way from the task context to the I/O driver:

mermaid
sequenceDiagram
    participant App as 应用任务
    participant Exec as 调度器 Worker
    participant Future as TcpStream::read Future
    participant Reactor as I/O 驱动 (epoll)
    participant Kernel as 操作系统内核

    App->>Future: poll(cx) 携带 Waker
    Future->>Reactor: 注册可读兴趣 + 保存 Waker
    Reactor->>Kernel: epoll_ctl(ADD, fd, EPOLLIN)
    Future-->>Exec: 返回 Poll::Pending
    Note over Exec: 任务挂起,Worker 去执行其他任务
    Kernel-->>Reactor: epoll_wait 返回 fd 就绪
    Reactor->>Reactor: 查找 fd 对应的 Waker
    Reactor->>Exec: waker.wake_by_ref()
    Note over Exec: 任务重新入队
    Exec->>Future: 再次 poll(cx)
    Future->>Kernel: read(fd, buf) 非阻塞读取
    Kernel-->>Future: 返回数据
    Future-->>App: 返回 Poll::Ready(n)

The key to this diagram is:Waker is the only channel that can reach the Executor in reverse from the Reactor. The Reactor does not hold any other information about the task; it only knows "when this fd is ready, call this Waker." This decoupling allows the I/O driver to be implemented independently of the scheduler, and the two communicate only through the narrow interface of Waker.

Spurious Wakeup: The Gray Area of the Contract

Tokio's documentation explicitly acknowledges the existence of spurious wakeups:

Normally, tasks are scheduled only if they have been woken by calling wake on their waker. However, this is not guaranteed, and Tokio may schedule tasks that have not been woken under some circumstances.

📎 tokio/src/runtime/mod.rs:306-309

[Design Inference and Architectural Trade-offs]

This means that the implementation ofpollmust be able to tolerate the situation of "being polled again without having been woken." A correct Future, after returning Pending, should still return Pending when polled again even if no event has occurred, rather than panicking or producing an erroneous result. This constraint may seem loose, but in fact it imposes requirements on the design of the state machine: it cannot assume that "an event must occur between two polls."

1.3 Executor: Encapsulation from Future to Task

Intuitive Model

The Executor is the restaurant's dispatcher. He has a stack of orders (task queue) in his hands and decides which order to make first and who makes it. When the pager vibrates, he puts the corresponding order back into the queue. Without a dispatcher, the chefs would not know which dish to make or when to switch work.

But the Executor's responsibilities go far beyond "polling Futures." It must solve three core problems:Task lifecycle management(creation, scheduling, completion, cancellation),fairness guarantees(preventing one task from starving other tasks),resource driver integration(how I/O and timer events are converted into wakeups).

Task Memory Layout: From Future to Task

When callingtokio::spawn, the passed-in Future is not directly placed into the queue. It is wrapped into aTaskstruct containing a reference count header, scheduling metadata, and the Future itself. This wrapping process has a key optimization decision:

rust
/// Boundary value to prevent stack overflow caused by a large-sized
/// Future being placed in the stack.
pub(crate) const BOX_FUTURE_THRESHOLD: usize = if cfg!(debug_assertions)  {
    2048
} else {
    16384
};

pub(crate) struct AutoBox<T>(std::marker::PhantomData<T>);

impl<T> AutoBox<T> {
    /// `true` if a value of type `T` is larger than [`BOX_FUTURE_THRESHOLD`].
    pub(crate) const SHOULD_BOX: bool = std::mem::size_of::<T>() > BOX_FUTURE_THRESHOLD;
}

📎 tokio/src/runtime/mod.rs:649-673

This code solves a very specific problem: if the Future is too large (over 16KB, 2KB in debug mode), inlining it directly into the Task struct will cause stack overflow or memory waste.AutoBoxuses the compile-time constantSHOULD_BOXto decide whether to box the Future.

[Design Inference and Architectural Trade-offs]

The comment particularly emphasizes "using associated constants rather than runtimeifThe reason: if runtime judgment is used, the compiler will, for eachTsimultaneously instantiate code for both branches (one handlingT, one handlingPin<Box<T>>), causing code bloat. With constant branching, the monomorphization collector prunes unreachable branches and generates code only for the types actually used. This is a classic optimization of "replacing runtime judgment with the type system."

Scheduling fairness: the magic numbers 31 and 61

Tokio's scheduler documentation defines a formal fairness guarantee:

If the total number of tasks does not grow without bound, and no task is blocking the thread, then it is guaranteed that tasks are scheduled fairly.

📎 tokio/src/runtime/mod.rs:279-281

The implementation of this guarantee depends on two key parameters. For the current-thread runtime:

The runtime will prefer to choose the next task to schedule from the local queue, and will only pick a task from the global queue if the local queue is empty, or if it has picked a task from the local queue 31 times in a row.

📎 tokio/src/runtime/mod.rs:328-333

The runtime will check for new IO or timer events whenever there are no tasks ready to be scheduled, or when it has scheduled 61 tasks in a row.

📎 tokio/src/runtime/mod.rs:335-337

These two numbers (31 and 61) are not chosen arbitrarily. 31 is 2 to the 5th power minus 1, which can be quickly checked with bitwise operations; 61 is chosen to ensure that I/O events are not delayed indefinitely—even if the task queue is never empty, I/O must be checked once every 61 scheduling rounds.

[Design inference and architectural trade-offs]

Why 31 and not 32? Because the counter starts at 0 and increments by 1 on each scheduling round, and when the counter reaches 31, a global queue check is triggered. Usingcounter & 31 == 31to check is more efficient thancounter % 32 == 0(although modern compilers will optimize it automatically). The choice of 61 is more subtle: it needs to be large enough to avoid the overhead of frequent epoll_wait system calls, yet small enough to keep I/O latency within an acceptable range.

LIFO slot optimization in the multi-threaded runtime

On top of fairness, the multi-threaded runtime adds a performance optimization—the LIFO slot:

The multi thread runtime uses the lifo slot optimization: Whenever a task wakes up another task, the other task is added to the worker thread's lifo slot instead of being added to a queue.

📎 tokio/src/runtime/mod.rs:373-377

The intuition behind this optimization is: when a task wakes another task, the awakened task is very likely to have a data dependency with the current task (such as in the producer-consumer pattern). By placing it in the LIFO slot, the current task can execute it immediately after finishing, taking advantage of hot data in the CPU cache.

But the LIFO slot has an anti-abuse mechanism:

if a worker thread uses the lifo slot three times in a row, it is temporarily disabled until the worker thread has scheduled a task that didn't come from the lifo slot.

📎 tokio/src/runtime/mod.rs:380-382

[Design inference and architectural trade-offs]

This rule of "disabled after three consecutive uses" is intended to prevent two tasks from waking each other and forming a livelock. If task A wakes task B, and B wakes A, without this restriction the LIFO slot would be permanently occupied by these two tasks, and other tasks would never get scheduled. The limit of three gives other tasks a chance to be inserted.

Task cancellation: the real semantics of abort

JoinHandle::abortThe behavior of

Be aware that calls to JoinHandle::abort just schedule the task for cancellation, and will return before the cancellation has completed.

📎 tokio/src/task/mod.rs:146-148

is often misunderstood. The documentation clearly states:abortThis means that.awaitis not synchronous. It only sets a flag, and the task will check this flag at the next.awaitpoint and terminate itself. If the task is executing a CPU-intensive section of code with noabort,

will not take effect immediately.

Note that aborting a task does not guarantee that it fails with a cancelled error, since it may complete normally first.

📎 tokio/src/task/mod.rs:134-138

More subtly:

[Design inference and architectural trade-offs]spawn_blockingThe design motivation for this semantics is: cancellation is a "best-effort" operation. Tokio does not forcibly kill tasks (Rust has no safe forced-termination mechanism), but instead cooperatively requests that the task exit on its own. This is consistent with the design that.awaittasks are not cancellable—blocking tasks have no

points and cannot check the cancellation flag.

1.4 Design reflections: the boundaries and costs of the trio

Why Future does not include ExecutorFutureRust's

trait deliberately does not include information about "how to schedule itself." This is a deliberate decoupling decision. If a Future knew its Executor, then:

1. The same Future could not be executed on different runtimes (for example, migrating from Tokio to async-std)block_on2. During testing, it could not be driven with a simple

select!、join!3. Combinators (such as

) could not work across runtimes

The existence of Waker is precisely to preserve this decoupling while still allowing the Future to notify the Executor. Waker is a "capability token"—the Future only knows "I can call this to request rescheduling," but does not know how scheduling actually happens.

The cost of cooperative scheduling.awaitTokio's tasks are cooperative: a task yields execution only at

code that spends a long time without reaching an .await will prevent other tasks from running.

📎 tokio/src/lib.rs:178-179

points. This means:

[Design inference and architectural trade-offs].awaitThis is the fundamental cost of cooperative scheduling. The operating system can preempt a thread at any instruction boundary, but Tokio can switch tasks only at.awaitpoints. If a task executes a 10-second CPU-intensive loop with nospawn_blockingin between, then all other tasks on the same worker thread will be blocked for 10 seconds. Tokio's response is to provideblock_in_placeand

, moving this kind of work to a dedicated thread pool. But this is the user's responsibility; the runtime cannot detect it automatically.

Boundary conditions of the fairness guarantee

  • Tokio's fairness guarantee has two preconditions: the total number of tasks is bounded, and no task blocks the thread. These two conditions are often violated in real production environments:
  • If tasks continuously spawn new tasks and do not reclaim them, the total number of tasks is unbounded and the fairness guarantee fails
If a task executes a blocking system call (such as synchronous file I/O), it blocks the entire worker thread

[Design inference and architectural trade-offs]

This is why Tokio's documentation repeatedly emphasizes "do not perform blocking operations in asynchronous tasks." The fairness guarantee is not a hard guarantee of the runtime, but a guarantee "under the premise of correct use." The runtime does not detect violations, because detection itself has overhead.

1.5 Chapter summary

Future is a pull-based state machine. pollis a pure query action, returningPendingmust have already registered a waker when returningReadyshould no longer be polled after returning. Tokio directly reusesstd::future::Futurewithout additional wrapping (unless tracing is enabled).

Waker is the only channel for reverse control flow.It achieves runtime independence through a "data pointer + vtable" design.wakeconsumes ownership,wake_by_refonly borrows. Spurious wakeups are allowed, and Future must tolerate them.

Executor is responsible for lifecycle, fairness, and resource integration.It wraps Future into a Task, throughAutoBoxdeciding at compile time whether to box, through the two magic numbers 31/61 balancing scheduling between the local queue and the global queue, and through the LIFO slot optimizing performance in data-dependent scenarios.

These three components are decoupled through narrow interfaces: Future only knowspoll, Waker only knowswake, Executor only knows "poll until Pending or Ready." It is precisely this decoupling that allows Tokio to implement advanced features such as work-stealing scheduling, I/O driver integration, and cooperative budgeting without modifying the Future definition.

Chapter Review and Self-Test

Q1: If theAutoBox::SHOULD_BOXjudgment is changed from a compile-time constant to a runtimeif size_of::<T>() > THRESHOLD, what impact will it have on the compiled artifact? Why does Tokio's comment specifically emphasize this point?

Reference Analysis: According to the📎 tokio/src/runtime/mod.rs:657-667comment, if a runtimeifis used, the compiler will instantiate code for both branches for eachT—one handling the case whereTis directly inlined, and one handling thePin<Box<T>>case. This means that each spawned Future type will generate two copies of the task harness code, causing the binary size to double. With the associated constantSHOULD_BOX, since it is a compile-time constant onceTis determined, the monomorphization collector will prune unreachable branches and generate code only for the path actually used. This is a typical optimization of "replacing runtime judgment with the type system," at the cost thatAutoBoxmust be a generic struct rather than an ordinary function.

Q2: Suppose a task returnspollinPendingbut forgets to register a Waker. What happens to this task under the current-thread runtime and the multi-thread runtime respectively? Does Tokio have a mechanism to detect this situation?

Reference Analysis: According to📎 tokio/src/runtime/mod.rs:306-309, Tokio allows spurious wakeups, which means a task may be rescheduled without being woken. But this does not mean forgetting to register a Waker is safe. Under the current-thread runtime, if both the local queue and the global queue are empty, the runtime enters theparkstate waiting for I/O or timer events. A task that forgets to register a Waker will never be re-enqueued, resulting in permanent suspension. Under the multi-thread runtime, the situation is similar, but if other tasks continue to wake, the task may be accidentally rescheduled due to spurious wakeups—but this cannot be relied upon. Tokio has no runtime detection mechanism to discover the case of "returning Pending but not registering a Waker," because this would require checking after every poll whether the Waker was used, which is too expensive. This is the responsibility of the Future implementer.

Q3: What specific scenario is the "disabled after three consecutive uses" rule of the LIFO slot intended to prevent? If this restriction were removed, under what kind of task dependency pattern would other tasks be starved?

Reference Analysis: According to📎 tokio/src/runtime/mod.rs:380-382, the LIFO slot is temporarily disabled after three consecutive uses, until a task from a non-LIFO source is scheduled. The scenario this rule prevents is: two tasks waking each other in a tight loop. For example, task A wakes task B after processing a batch of data, and task B immediately wakes task A after processing. Without the three-use limit, A and B would forever occupy the LIFO slot, the worker thread would switch infinitely between these two tasks, and other tasks in the local queue and global queue would never get a chance to execute. The three-use limit ensures that after every three rounds of "mutual waking," at least one other task is scheduled, breaking the livelock. The choice of this number is empirical: too small reduces the benefit of the LIFO optimization, too large increases the latency of other tasks.

At this point, the responsibility boundaries and collaboration mechanisms among Future, Waker, and Executor are clear: Future defines computation, Waker is responsible for waking, and Executor drives execution. But a single component cannot work independently; they must be assembled into a unified runtime environment. In the next chapter, we will trace the complete assembly chain of Runtime::new and Builder::build, see how the scheduler, I/O driver, time driver, and blocking thread pool are injected into the same Runtime instance, and reveal the fundamental differences between the current_thread and multi_thread forms during the assembly stage.

CHAPTER 02

Chapter 2: Runtime Assembly: How Builder Assembles Drivers, Scheduler, and Thread Pool

Project: tokio-rs/tokio · Book progress: Chapter 2 / 14 · Verification status: FACT line numbers are genuinely anchored

FromBuildertoRuntime: The complete journey of an assembly

In the previous chapter, we clarified the responsibility boundaries of Future, Waker, and Executor. But a truly usable runtime is far more than "an Executor" — it also needs an I/O event loop, timers, a blocking thread pool, and these components must share the same set of handles and the same lifecycle. This chapter traces the complete assembly chain ofBuilder::build, answering a core question:What components exist inside aRuntime, and how are they assembled and share handles。

Tokio's assembly entry point isBuilder. It is itself a pure configuration container; all fields are "intent declarations" and hold no runtime resources. The actual resource creation happens whenbuild()is called.

Intuitive model: Builder is the "renovation blueprint", Runtime is the "house after delivery"

Builderis like a renovation blueprint: you mark on it "how many rooms (worker_threads)", "whether to run water (enable_io)", "whether to run electricity (enable_time)", "outsourced helper limit (max_blocking_threads)". The blueprint itself produces no physical entity. Only whenbuild()is called does the construction crew build according to the blueprint, actually constructing the "rooms" — the scheduler, drivers, thread pool — and deliver aRuntimeinstance.

Without theBuilderlayer, users would have to manually new each component, manually wire them up, and manually handle failure rollback — any ordering mistake would lead to dangling handles or resource leaks.BuilderThe value oflies in: completely separating "configuration" from "construction", so that the construction process can centrally perform validation, failure cleanup, and handle sharing。

Memory layout:Builder's field partitions

Builder's fields can be divided into four groups by responsibility. The first group isform and switches:kinddetermines the scheduler form,enable_io / enable_timedetermines whether to create the corresponding driver.

📎 tokio/src/runtime/builder.rs:55-68

rust
pub struct Builder {
    kind: Kind,
    name: Option<String>,
    enable_io: bool,
    nevents: usize,
    nevents_busy: Option<usize>,
    enable_time: bool,
    start_paused: bool,
    // ...
}

The second group isthread pool parameters:worker_threadsisOption<usize>,Nonemeaning "defer to build time to auto-detect based on CPU core count";max_blocking_threadsdefaults to 512.

📎 tokio/src/runtime/builder.rs:73-79

rust
worker_threads: Option<usize>,
max_blocking_threads: usize,

The third group iscallback hooks, all of which areOption<Arc<dyn Fn ...>>. Note that they useArcrather thanBox, because these callbacks need to be cloned into each worker thread'sConfig。

📎 tokio/src/runtime/builder.rs:87-97

rust
pub(super) after_start: Option<Callback>,
pub(super) before_stop: Option<Callback>,
pub(super) before_park: Option<Callback>,
pub(super) after_unpark: Option<Callback>,

The fourth group isscheduling heuristics and random seed:global_queue_interval、event_interval、disable_lifo_slot、seed_generator。

📎 tokio/src/runtime/builder.rs:116-134

rust
pub(super) global_queue_interval: Option<u32>,
pub(super) event_interval: u32,
pub(super) disable_lifo_slot: bool,
pub(super) seed_generator: RngSeedGenerator,

There is a noteworthy design here:Kindis aCopysmall enum with only two variants.

📎 tokio/src/runtime/builder.rs:261-265

rust
#[derive(Clone, Copy)]
pub(crate) enum Kind {
    CurrentThread,
    #[cfg(feature = "rt-multi-thread")]
    MultiThread,
}

MultiThreadThert-multi-threadvariant is gated by thertfeature. This means that in a build where only theKindfeature is enabled,build()has only one variant, andmatch'swill be optimized by the compiler into a single branch —。

using the type system rather than runtime checks to eliminate the code size of the multi-threaded scheduler

Builder::newThe philosophy of defaults: why I/O and time are disabled by defaultenable_iois the common entry point for all construction. It sets bothenable_timeandfalse。

📎 tokio/src/runtime/builder.rs:309-318

rust
// I/O defaults to "off"
enable_io: false,
nevents: 1024,
nevents_busy: None,

// Time defaults to "off"
enable_time: false,

// The clock starts not-paused
start_paused: false,
Copy

〔Design inference and architectural trade-offs〕#[tokio::main]This default choice is deliberate: creating the I/O driver requires requesting epoll/kqueue handles from the operating system, and creating the time driver requires starting timer infrastructure. If the user only wants a pure computation task scheduler (e.g., running CPU-intensive async logic), forcibly creating these drivers is pure waste.enable_all()。

enable_all()The reason the

📎 tokio/src/runtime/builder.rs:398-419

rust
pub fn enable_all(&mut self) -> &mut Self {
    #[cfg(any(
        feature = "net",
        all(unix, feature = "process"),
        all(unix, feature = "signal")
    ))]
    self.enable_io();

    #[cfg(all(
        tokio_unstable,
        feature = "io-uring",
        // ...
    ))]
    self.enable_io_uring();

    #[cfg(feature = "time")]
    self.enable_time();

    self
}

's implementation reveals how feature gating affects the semantics of "all enabled".enable_io()Copynet、processNote thatsignalis only called when thetime feature,enable_all()or

feature is enabled. If the user only enabledbuild()will not enable the I/O driver — because there is simply no I/O driver code in the compiled artifact.

build()Assembly main path:kind's branching

📎 tokio/src/runtime/builder.rs:1146-1152

rust
pub fn build(&mut self) -> io::Result<Runtime> {
    match &self.kind {
        Kind::CurrentThread => self.build_current_thread_runtime(),
        #[cfg(feature = "rt-multi-thread")]
        Kind::MultiThread => self.build_threaded_runtime(),
    }
}

into two completely different paths.

Copy

build_current_thread_runtimeThe difference between these two paths is far more than "one thread vs multiple threads". Let's expand each below.build_current_thread_runtime_componentsPath one: current_thread assemblyRuntime。

📎 tokio/src/runtime/builder.rs:1725-1736

rust
fn build_current_thread_runtime(&mut self) -> io::Result<Runtime> {
    use crate::runtime::runtime::Scheduler;

    let (scheduler, handle, blocking_pool) =
        self.build_current_thread_runtime_components(None)?;

    Ok(Runtime::from_parts(
        Scheduler::CurrentThread(scheduler),
        handle,
        blocking_pool,
    ))
}

, then wraps the returned triple intobuild_current_thread_runtime_componentsCopy

📎 tokio/src/runtime/builder.rs:1760-1766

rust
let mut cfg = self.get_cfg();
cfg.timer_flavor = TimerFlavor::Traditional;
let (driver, driver_handle) = driver::Driver::new(cfg)?;

// Blocking pool
let blocking_pool = blocking::create_blocking_pool(self, self.max_blocking_threads, 0);
let blocking_spawner = blocking_pool.spawner().clone();

. Its execution order is crucial:driverCopy(driver, driver_handle)The first step creates?, returning a pair ofbuild. Note that hereErrdirectly propagates the error upward — if I/O driver initialization fails (e.g., epoll creation fails), the entire

returnsspawner, and at this point the blocking pool has not yet been created, so no cleanup is needed.spawnerThe second step creates the blocking pool and immediately takes out its

clone. This

📎 tokio/src/runtime/builder.rs:1768-1770

rust
let seed_generator_1 = self.seed_generator.next_generator();
let seed_generator_2 = self.seed_generator.next_generator();
The third step generates two independent RNG seed generators.

Copyseed_generator_1〔Design inference and architectural trade-offs〕ConfigWhy are two needed?select!is placed intoseed_generator_2, for internal use by the scheduler (e.g.,CurrentThread::new's random branch order);rng_seedis passed to

, for use on the task side. Separating the two generators prevents the scheduler's internal consumption of random numbers from affecting the user-visible random sequence, thereby guaranteeingConfig's reproducibility.CurrentThread::new。

📎 tokio/src/runtime/builder.rs:1776-1807

rust
let (scheduler, handle) = CurrentThread::new(
    driver,
    driver_handle,
    blocking_spawner,
    seed_generator_2,
    Config {
        before_park: self.before_park.clone(),
        after_unpark: self.after_unpark.clone(),
        // ...
        global_queue_interval: self.global_queue_interval,
        event_interval: self.event_interval,
        // ...
        enable_eager_driver_handoff: false,
        seed_generator: seed_generator_1,
        // ...
    },
    local_tid,
    self.name.clone(),
);

together toenable_eager_driver_handoffCopyfalse。

📎 tokio/src/runtime/builder.rs:1795-1798

rust
// This setting never makes sense for a current thread runtime,
// as it only configures how the I/O driver is stolen across
// workers.
enable_eager_driver_handoff: false,
is hardcoded to

This comment points out the essence of this option: it describes "how multiple workers compete for the I/O driver," and current_thread has only one thread, so there is no contention, and it is therefore forcibly disabled. This is a typical example of "configuration item semantics being strongly correlated with form"—the sameBuilderfield has different meanings under different forms.

Finally,CurrentThread::newthe returnedhandleis wrapped intoscheduler::Handle::CurrentThread, and then wrapped into the publicHandle。

📎 tokio/src/runtime/builder.rs:1816-1822

rust
let handle = Handle {
    inner: scheduler::Handle::CurrentThread(handle),
};

Ok((scheduler, handle, blocking_pool))

Path two: assembly of multi_thread

build_threaded_runtimeThe skeleton of is similar to current_thread, but there are three essential differences. The first is the determination of the number of worker threads:

📎 tokio/src/runtime/builder.rs:2185

rust
let worker_threads = self.worker_threads.unwrap_or_else(num_cpus);

NoneHere it is parsed asnum_cpus(). This is where "delayed automatic detection" lands—detection happens at build time rather than atBuilder::newtime, because CPU affinity may change between the two.

The second difference is in the capacity calculation of the blocking pool:

📎 tokio/src/runtime/builder.rs:2189-2192

rust
let blocking_pool =
    blocking::create_blocking_pool(self, self.max_blocking_threads + worker_threads, worker_threads);
let blocking_spawner = blocking_pool.spawner().clone();

Notemax_blocking_threads + worker_threads. In contrast, the current_thread path passes inself.max_blocking_threadsand0。

📎 tokio/src/runtime/builder.rs:1765

rust
let blocking_pool = blocking::create_blocking_pool(self, self.max_blocking_threads, 0);
[Design inference and architectural trade-offs]

This difference reveals the semantics of blocking pool capacity: under multi_thread,max_blocking_threadsis the upper limit of "additional" blocking threads, and the actual total thread limit must add the number of worker threads. The third parameter (current_thread passes 0, multi_thread passesworker_threads) is very likely a hint for "reserved thread count" or "initial thread count." This design keeps the semantics ofmax_blocking_threadsconsistent across the two forms: it describes "how many extra blocking threads can be opened beyond the core workers."

The third difference is thatMultiThread::newreturns a triple instead of a pair:

📎 tokio/src/runtime/builder.rs:2198-2226

rust
let (scheduler, handle, launch) = MultiThread::new(
    worker_threads,
    driver,
    driver_handle,
    blocking_spawner,
    seed_generator_2,
    Config {
        // ...
        enable_eager_driver_handoff: self.enable_eager_driver_handoff,
        // ...
    },
    self.timer_flavor,
    self.name.clone(),
);

The extralaunchis a "startup handle."MultiThread::newis only responsible for constructing the scheduler structure,and does not immediately start worker threads. The actual startup happens later:

📎 tokio/src/runtime/builder.rs:2228-2234

rust
let handle = Handle { inner: scheduler::Handle::MultiThread(handle) };

// Spawn the thread pool workers
let _enter = handle.enter();
launch.launch();

Ok(Runtime::from_parts(Scheduler::MultiThread(scheduler), handle, blocking_pool))

handle.enter()enters the runtime context, and thenlaunch.launch()actually spawns all worker threads. This two-phase design of "construct first, start later" is very critical.

[Design inference and architectural trade-offs]

Why can't it start while constructing? Because once worker threads start, they immediately begin polling tasks, and tasks may referencehandle. Ifhandlehas not finished being constructed, there will be a race where "workers hold a half-finished handle." The two-phase design guarantees:when all worker threads start, the completeHandleis already ready。_enterThe guard ensures that worker threads are in the correct runtime context at the moment they start.

Assembly flow diagram

The diagram below puts the assembly order, key branches, and error paths of both paths together. Note that whendriver::Driver::newfails, it directly returnsErr, and at this point the blocking pool has not yet been created.

mermaid
flowchart TD
    start["Builder::build()"] --> match_kind{"self.kind?"}

    match_kind -->|CurrentThread| ct_cfg["get_cfg() + timer_flavor=Traditional"]
    match_kind -->|MultiThread| mt_workers["worker_threads = self.worker_threads.unwrap_or_else(num_cpus)"]

    ct_cfg --> ct_driver["driver::Driver::new(cfg)?"]
    mt_workers --> mt_driver["driver::Driver::new(self.get_cfg())?"]

    ct_driver -->|Err| ret_err["return Err(io::Error)"]
    mt_driver -->|Err| ret_err

    ct_driver -->|Ok driver, driver_handle| ct_pool["create_blocking_pool(self, max_blocking_threads, 0)"]
    mt_driver -->|Ok driver, driver_handle| mt_pool["create_blocking_pool(self, max_blocking_threads + worker_threads, worker_threads)"]

    ct_pool --> ct_seed["next_generator() x2"]
    mt_pool --> mt_seed["next_generator() x2"]

    ct_seed --> ct_new["CurrentThread::new(driver, driver_handle, blocking_spawner, ...)"]
    mt_seed --> mt_new["MultiThread::new(worker_threads, driver, ...) -> (scheduler, handle, launch)"]

    ct_new --> ct_wrap["Handle { inner: CurrentThread(handle) }"]
    mt_new --> mt_wrap["Handle { inner: MultiThread(handle) }"]

    ct_wrap --> ct_rt["Runtime::from_parts(Scheduler::CurrentThread, handle, blocking_pool)"]
    mt_wrap --> mt_enter["handle.enter()"]
    mt_enter --> mt_launch["launch.launch() 启动 worker 线程"]
    mt_launch --> mt_rt["Runtime::from_parts(Scheduler::MultiThread, handle, blocking_pool)"]

Handle sharing:HandleHow becomes a "pass" across components

After assembly is complete,Runtimeholdsscheduler、handle、blocking_poolthe trio. Among them,handleis the shared core. Internally it is an enum:

📎 tokio/src/runtime/scheduler/mod.rs:29-41

rust
#[derive(Debug, Clone)]
pub(crate) enum Handle {
    #[cfg(feature = "rt")]
    CurrentThread(Arc<current_thread::Handle>),

    #[cfg(feature = "rt-multi-thread")]
    MultiThread(Arc<multi_thread::Handle>),

    #[cfg(not(feature = "rt"))]
    #[allow(dead_code)]
    Disabled,
}

Note that both variants wrapArc. This means that cloningHandleis a cheap reference-count increment and can be freely distributed to any thread.Handleprovides a unified access interface, encapsulating form differences insidematch. For exampledriver():

📎 tokio/src/runtime/scheduler/mod.rs:53-64

rust
pub(crate) fn driver(&self) -> &driver::Handle {
    match *self {
        #[cfg(feature = "rt")]
        Handle::CurrentThread(ref h) => &h.driver,

        #[cfg(feature = "rt-multi-thread")]
        Handle::MultiThread(ref h) => &h.driver,

        #[cfg(not(feature = "rt"))]
        Handle::Disabled => unreachable!(),
    }
}

blocking_spawner()uses thematch_flavor!macro to eliminate duplication:

📎 tokio/src/runtime/scheduler/mod.rs:96-98

rust
pub(crate) fn blocking_spawner(&self) -> &blocking::Spawner {
    match_flavor!(self, Handle(h) => &h.blocking_spawner)
}

After expansion, this macro is thedriver()like thematchabove. Its value is that when adding a new accessor that needs to dispatch by form, only one line ofmatch_flavor!is needed, instead of manually writing thematchbranches twice.

The publicHandleis a thin wrapper around the internalscheduler::Handle:

📎 tokio/src/runtime/handle.rs:13-15

rust
pub struct Handle {
    pub(crate) inner: scheduler::Handle,
}

TheHandleusers get can be cloned across threads, canspawn, canblock_on。spawnThe implementation shows the compile-time branching ofAutoBox:

📎 tokio/src/runtime/handle.rs:197-208

rust
pub fn spawn<F>(&self, future: F) -> JoinHandle<F::Output>
where
    F: Future + Send + 'static,
    F::Output: Send + 'static,
{
    let fut_size = mem::size_of::<F>();
    if AutoBox::<F>::SHOULD_BOX {
        self.spawn_named(Box::pin(future), SpawnMeta::new_unnamed(fut_size))
    } else {
        self.spawn_named(future, SpawnMeta::new_unnamed(fut_size))
    }
}

AutoBox::<F>::SHOULD_BOXis an associated constant, derived by comparingsize_of::<F>()with a threshold.

📎 tokio/src/runtime/mod.rs:668-673

rust
pub(crate) struct AutoBox<T>(std::marker::PhantomData<T>);

impl<T> AutoBox<T> {
    pub(crate) const SHOULD_BOX: bool = std::mem::size_of::<T>() > BOX_FUTURE_THRESHOLD;
}
[Design inference and architectural trade-offs]

The comment explains why an associated constant is used instead of a runtimeif: if runtime judgment were used,spawn_namedwould be monomorphized twice (once forF, once forPin<Box<F>>), causing every spawned future to generate two task harnesses and doubling code size. With constant branching, the monomorphization collector keeps only the branch actually taken.

Design thinking: assembly order, error recovery, and production pitfalls

Order is contract. The assembly orderdriver -> blocking_pool -> scheduleris not arbitrary. The driver is created first because it is the only step that may fail due to insufficient OS resources and, after failure, requires no cleanup of other components. blocking_pool comes after the driver and before the scheduler, because the scheduler needs blocking_spawner. If blocking_pool creation fails (in practice it is unlikely to fail), the driver will be automatically cleaned up by drop.

current_thread'slocal_tidbranch。build_localtakesbuild_current_thread_local_runtime, passing in the current thread ID:

📎 tokio/src/runtime/builder.rs:1738-1751

rust
fn build_current_thread_local_runtime(&mut self) -> io::Result<LocalRuntime> {
    use crate::runtime::local_runtime::LocalRuntimeScheduler;

    let tid = std::thread::current().id();

    let (scheduler, handle, blocking_pool) =
        self.build_current_thread_runtime_components(Some(tid))?;

    Ok(LocalRuntime::from_parts(
        LocalRuntimeScheduler::CurrentThread(scheduler),
        handle,
        blocking_pool,
    ))
}

Thistidis stored inHandle, and latercan_spawn_local_on_local_runtimeuses it to verify "whether spawn_local is called on the owner thread":

📎 tokio/src/runtime/scheduler/mod.rs:140-147

rust
pub(crate) fn can_spawn_local_on_local_runtime(&self) -> bool {
    match self {
        Handle::CurrentThread(h) => h.local_tid.is_some_and(|x| std::thread::current().id() == x),

        #[cfg(feature = "rt-multi-thread")]
        Handle::MultiThread(_) => false,
    }
}
[Design inference and architectural trade-offs]

This is the cornerstone ofLocalRuntimesafety:!Send's future can only be polled on its owner thread, andlocal_tidis the runtime checkpoint for this constraint. If this check were removed, cross-thread spawn_local would cause!Senddata to be accessed concurrently, leading to UB.

Production pitfall one:worker_threads(0)will panic。worker_threadsThe method has an assertion:

📎 tokio/src/runtime/builder.rs:582-586

rust
pub fn worker_threads(&mut self, val: usize) -> &mut Self {
    assert!(val > 0, "Worker threads cannot be set to 0");
    self.worker_threads = Some(val);
    self
}

This assertion fails at the configuration stage rather than waiting until build. The benefit is earlier error localization; the downside is that if the thread count comes from a dynamic value in a config file, users must validate it themselves before calling.

Production Pitfall Two:max_blocking_threadsSetting it too small will hang. The documentation explicitly warns:

📎 tokio/src/runtime/builder.rs:600-601

rust
/// It's recommended to not set this limit too low in order to avoid hanging on operations
/// requiring [`spawn_blocking`].
[Design Inference and Architectural Trade-offs]

Because the blocking pool's queue has no backpressure—tasks will keep accumulating until a thread becomes available. If all blocking threads are waiting on some operation that "requires a new blocking thread to complete," a deadlock occurs. The documentation's statement "the queue does not apply any backpressure, it could potentially grow unbounded" is precisely a footnote to this risk.

Production Pitfall Three:UnhandledPanic::ShutdownRuntimeOnly supports current_thread。

📎 tokio/src/runtime/builder.rs:1374-1381

rust
pub fn unhandled_panic(&mut self, behavior: UnhandledPanic) -> &mut Self {
    if !matches!(self.kind, Kind::CurrentThread) && matches!(behavior, UnhandledPanic::ShutdownRuntime) {
        panic!("UnhandledPanic::ShutdownRuntime is only supported in current thread runtime");
    }

    self.unhandled_panic = behavior;
    self
}
[Design Inference and Architectural Trade-offs]

The reason for this limitation is: under multi_thread, "immediately shutting down the runtime" requires coordinating the stopping of all worker threads, which is highly complex to implement and semantically ambiguous (what about other tasks currently being polled?). current_thread has only one thread, so the shutdown semantics are clear.

Chapter Summary

This chapter traced theBuilder::buildcomplete assembly chain. Core conclusions:

1. Builderis a pure configuration container,build()is what actually creates resources. The assembly orderdriver -> blocking_pool -> scheduleris determined by error recovery requirements.

2. The difference between current_thread and multi_thread is not just the thread count: the blocking pool capacity calculation differs (max_blocking_threads vs max_blocking_threads + worker_threads), multi_thread has an extralaunchtwo-phase startup,enable_eager_driver_handoffis forcibly disabled under current_thread.

3. Handleis the core shared across components, internally usingArcto wrap form-specific handles, accessed uniformly through thematchormatch_flavor!macros.

4. AutoBoxuses associated constants to decide at compile time whether to box the future, avoiding doubling the code size.

5. local_tidis theLocalRuntimeruntime checkpoint for safety.

In the next chapter, we will enter the task lifecycle:spawnhow a Future becomes a schedulable entity,JoinHandlehow it interacts with the task state machine, and the state transitions of tasks betweenPENDING / RUNNING / COMPLETE.

Chapter Review and Self-Test

Q1: If inbuild_threaded_runtimethe capacity parameter ofcreate_blocking_poolis changed fromself.max_blocking_threads + worker_threadstoself.max_blocking_threads, in what scenarios would blocking tasks starve? Why can the current_thread path passself.max_blocking_threads?

Reference Analysis: According to📎 tokio/src/runtime/builder.rs:2189-2192, the multi_thread path passes inself.max_blocking_threads + worker_threads, while the current_thread path📎 tokio/src/runtime/builder.rs:1765passes inself.max_blocking_threads. The root of the difference is: under multi_thread, worker threads themselves also execute blocking tasks (for example,block_in_placetemporarily converts a worker thread into a blocking thread), so the total budget for blocking threads must include the number of worker threads. If changed to pass onlyself.max_blocking_threads, whenmax_blocking_threadsis set small (say 1) and worker threads are already occupying the budget inblock_in_place, newspawn_blockingtasks will have no threads available and will pile up in the backpressure-free queue, causing async tasks that depend on these blocking tasks to hang permanently. current_thread has only one thread and does not supportblock_in_place's worker conversion semantics, so there is no need to add the worker count.

Q2: MultiThread::newreturns thelaunchhandle; what actually starts the worker threads islaunch.launch(). If thehandle.enter()line is removed andlaunch.launch()is called directly, what happens?

Reference Analysis: According to📎 tokio/src/runtime/builder.rs:2230-2232, before startup there islet _enter = handle.enter();and only thenlaunch.launch()。handle.enter(). The purpose is to set the thread-local context, making the current thread "appear" to be inside the runtime. After worker threads start, they immediately begin polling tasks, and task code may callHandle::current()、tokio::spawnand other context-dependent APIs. If_enteris removed, the context setup at the moment worker threads start may be incomplete (depending on whetherlaunchsets it internally), and in the worst case, initialization code executing on the worker thread that callsHandle::current()will panic (CONTEXT_MISSING_ERROR). Even iflaunchinternally sets the context for each worker,_enterensures that "the startup action itself" occurs in the correct context, avoiding races during startup.

Q3: AutoBox::<F>::SHOULD_BOXuses associated constants rather than runtimeif size_of::<F>() > THRESHOLD. Suppose it were changed to runtime judgment—besides doubling the code size, in what situations would it cause performance degradation?

Reference Analysis: According to📎 tokio/src/runtime/mod.rs:657-673's comments, runtimeifcausesspawn_namedto monomorphize eachTtwice (TandPin<Box<T>>each once). Besides doubling the code size, performance degradation manifests as: 1) increased instruction cache (i-cache) pressure, because both sets of harness code must reside; 2) the compiler cannot optimize for "actually only one branch is taken," and although runtime branch prediction is usually accurate, the branch itself and the register allocation differences between the two code sets accumulate; 3) more subtly,Pin<Box<T>>the path forces heap allocation, and if the runtime judgment misjudges for some reason (e.g.,size_ofis not fully constant-folded in a generic context), small futures will also be boxed, adding an extra heap allocation per spawn. Associated constants let the monomorphization collector prune the untaken branch at compile time, with zero runtime overhead.

CHAPTER 03

Chapter 3: The Life of a Task (Part 1): How spawn Turns a Future into a Schedulable Entity

Project: tokio-rs/tokio · Book Progress: Chapter 3 / 14 · Verification Status: FACT line numbers are genuinely anchored

In the previous chapter, we completed the assembly of the Runtime: the I/O driver, time driver, blocking pool, and scheduler are injected into the sameRuntimeinstance,Handlebecoming a shared handle for cross-thread access to these components. But the assembled runtime is still just an empty shell at this point—it has the engine to drive tasks, but no tasks to drive. The question this chapter aims to answer is precisely: when you typetokio::spawn(async { ... })at that moment, what exactly does thatasyncblock go through to transform from ordinary Rust code into an entity that "can be taken over by the scheduler, can be woken up, and can be joined." This is the first half of "The Life of a Task," and we focus on birth: starting fromHandle::spawngoing throughnew_task's reference-count allocation, landing onCell<T, S>'s memory layout, and finally seeing clearly how a task is delivered to some worker's local queue or the global injection queue. The second half (Chapter 4) will enter the scheduling loop and the poll/wake closed loop.

3.1 A Future Is Not a Task: What Exactly Does One spawn Create

Intuitive Model

Think ofFutureas a "recipe," and think of a task as "a dish currently being cooked in the kitchen." The recipe itself is static, copyable, and has no execution state; only when the kitchen (scheduler) decides "make this dish now," assigns it a stove (worker), an order number (TaskId), and a serving window (JoinHandle), does it become a "dish in production." Without this layer of wrapping, the scheduler would have no way to know "how far along this dish is," "who is waiting for it," or "whom to notify when it's done"—it can only see a recipe and cannot manage it.

Data Structures and Memory Layout

Tokio usesTask<S>to represent "a task reference owned by the runtime," which is a transparent wrapper aroundRawTask:

rust
#[repr(transparent)]
pub(crate) struct Task<S: 'static> {
    raw: RawTask,
    _p: PhantomData<S>,
}

📎 tokio/src/runtime/task/mod.rs:233-238

#[repr(transparent)]means thatTask<S>andRawTaskare completely identical in memory, with no additional overhead.PhantomData<S>is only a compile-time type marker, marking which scheduler type this task belongs toS。

What truly carries all the task's state isCell<T, S>, and its layout is the cornerstone of the entire task module:

rust
#[repr(C)]
pub(super) struct Cell<T: Future, S> {
    pub(super) header: Header,
    pub(super) core: Core<T, S>,
    pub(super) trailer: Trailer,
}

📎 tokio/src/runtime/task/core.rs:126-136

The three fields are arranged as "hot-warm-cold."Headeris hot data (accessed on every scheduling and every state transition),Coreis warm data (accessed during poll),Traileris cold data (accessed only during creation and destruction). The comment explicitly states:Headermust be the first field, because the task struct will be referenced by both*mut Celland*mut Header📎 tokio/src/runtime/task/core.rs:37-43。

More critical is cache line alignment.Cellhas a long list of#[cfg_attr(..., repr(align(...)))]attached to it, selecting the alignment byte count according to the target architecture: x86_64/aarch64/powerpc64 use 128 bytes, arm/mips/sparc/hexagon use 32 bytes, m68k uses 16 bytes, s390x uses 256 bytes, and the rest default to 64 bytes📎 tokio/src/runtime/task/core.rs:64-125. The comment explains why x86_64 should use 128 rather than 64: starting with Intel Sandy Bridge, the spatial prefetcher fetchespairsof 64-byte cache lines at once, so it must be aligned to 128 bytes to avoid false sharing📎 tokio/src/runtime/task/core.rs:45-53。

[Design Inference and Architectural Trade-offs]

The cost of this alignment strategy is that each task wastes at least one cache line of space. But the task state bits (state) are read and written at high frequency by multiple worker threads—one thread sets the RUNNING bit during poll, another thread reads the NOTIFIED bit during wake—if the state bits of two tasks fall on the same cache line, every state transition will trigger the cache line to bounce back and forth between cores (cache line ping-pong), and the performance loss far exceeds the memory waste. Tokio chooses to trade space for time.

Headeritself is constrained to within 8 pointer sizes:

rust
#[test]
#[cfg(not(loom))]
fn header_lte_cache_line() {
    assert!(std::mem::size_of::<Header>() <= 8 * std::mem::size_of::<*const ()>());
}

📎 tokio/src/runtime/task/core.rs:591-593

This test ensures thatHeaderdoes not exceed 64 bytes (8 × 8), so that on architectures with 64-byte cache lines it can fit entirely within one line.Header's fields include:state: State(atomic state bits),queue_next: UnsafeCell<Option<NonNull<Header>>>(linked-list pointer for the injection queue),vtable: &'static Vtable(function pointer table),owner_id: UnsafeCell<Option<NonZeroU64>>(the ID of theOwnedTaskslist it belongs to),scheduled_at: UnsafeCell<ScheduleLatencyInstant>(scheduling delay measurement)📎 tokio/src/runtime/task/core.rs:169-198。

Core<T, S>holds the scheduler handlescheduler: S, the task IDtask_id: Id, and the most centralstage: CoreStage<T> 📎 tokio/src/runtime/task/core.rs:148-165。Stageis a three-state enum:

rust
#[repr(C)]
pub(super) enum Stage<T: Future> {
    Running(T),
    Finished(super::Result<T::Output>),
    Consumed,
}

📎 tokio/src/runtime/task/core.rs:225-229

This is precisely the key to "Future and Output reusing the same memory": during task executionStage::Runningholds the future, and after completion it is replaced in place withStage::Finished(output), and after being taken byJoinHandleit becomesStage::Consumed。#[repr(C)]The comment points to a Miri issue, indicating that this layout has hard requirements for the correctness of unsafe code📎 tokio/src/runtime/task/core.rs:225-229。

Trailerstores cold data:owned: linked_list::Pointers<Header>(OwnedTaskslinked-list pointer),waker: UnsafeCell<Option<Waker>>(the consumer waker waiting for task completion),hooks: TaskHarnessScheduleHooks 📎 tokio/src/runtime/task/core.rs:205-213。

Step-by-Step: From spawn to Enqueue

Let us plug in a concrete scenario: in a multi_thread runtime, worker thread A executestokio::spawn(async { 42 })。

Step 1: Construct the task trio. new_taskis the only entry point for a task's birth:

rust
fn new_task<T, S>(
    task: T,
    scheduler: S,
    id: Id,
    spawned_at: SpawnLocation,
) -> (Task<S>, Notified<S>, JoinHandle<T::Output>)

📎 tokio/src/runtime/task/mod.rs:336-346

It callsRawTask::new::<T, S>to allocateCell, then derives three references from the samerawpointer:Task(owned reference, usually immediately placed intoOwnedTasks)、Notified(notification reference, handed to the scheduler),JoinHandle(result-reading handle)📎 tokio/src/runtime/task/mod.rs:347-363. Note that the three share the sameraw, each holding a reference count.

Step 2: AllocateCelland write the initial state. Cell::newAllocate the entire structure on the heap:

rust
let result = Box::new(Cell {
    trailer: Trailer::new(scheduler.hooks()),
    header: new_header(state, vtable, ...),
    core: Core {
        scheduler,
        stage: CoreStage {
            stage: UnsafeCell::new(Stage::Running(future)),
        },
        task_id,
        ...
    },
});

📎 tokio/src/runtime/task/core.rs:261-278

vtableGenerated byraw::vtable::<T, S>(), it is a function pointer table monomorphized for a specificTandS.📎 tokio/src/runtime/task/core.rs:260. The future is moved directly intoStage::Running, with no additional boxing.

Step 3: Debug assertions verify the layout.Underdebug_assertions,Cell::newwill call thecheckfunction, usingHeader::get_trailer、Header::get_scheduler、Header::get_id_ptrand other pointer arithmetic based on vtable offsets to assert one by one that "the field address looked up through the header" matches "the actual field address"📎 tokio/src/runtime/task/core.rs:280-321. This is a runtime self-check of the correctness of vtable offsets.

Step 4: Submit to the scheduler.After the scheduler receivesNotified<S>, it callsSchedule::schedule 📎 tokio/src/runtime/task/mod.rs:315. Under multi_thread, this goes throughpush_back_or_overflow, pushing the task into the current worker's local queue, and overflowing to the injection queue when the queue is full.

The following diagram depicts the control flow and branches fromnew_taskto enqueueing:

mermaid
flowchart TD
    spawn_call["Handle::spawn(future)"] --> new_task["new_task::<T,S>(future, scheduler, id)"]
    new_task --> raw_new["RawTask::new::<T,S>"]
    raw_new --> cell_new["Cell::new: Box::new(Cell{header, core, trailer})"]
    cell_new --> vtable["raw::vtable::<T,S>() 生成函数指针表"]
    cell_new --> stage["Stage::Running(future) 移入"]
    cell_new --> debug_check{"debug_assertions?"}
    debug_check -->|是| check_layout["check(): 断言 trailer/scheduler/id 偏移量"]
    debug_check -->|否| skip_check["跳过"]
    check_layout --> triple["派生 (Task, Notified, JoinHandle)"]
    skip_check --> triple
    triple --> owned["Task 存入 OwnedTasks"]
    triple --> sched["Notified 交给 Schedule::schedule"]
    sched --> push{"本地队列有容量?"}
    push -->|是| local_push["push_back_finish: 写入 buffer[tail & MASK]"]
    push -->|否| overflow_check{"steal == real?"}
    overflow_check -->|否, 有并发窃取| inject_only["overflow.push(task) 仅注入"]
    overflow_check -->|是| push_overflow["push_overflow: CAS 认领后半批"]
    push_overflow --> cas_ok{"CAS 成功?"}
    cas_ok -->|是| inject_batch["overflow.push_batch(后半批 + 当前 task)"]
    cas_ok -->|否| retry["返回 Err(task), 重试 push_back_or_overflow"]
    retry --> push

This diagram reveals several key branches: debug assertions only take effect in debug builds; when the local queue is full, it does not overflow directly, but first checks whether there is a concurrent stealer (steal != real). If so, it only pushes the current task into the injection queue, because the space freed up by the stealer will soon become available.

Design consideration: why three references instead of one

new_taskreturns three references, not one. This is the core of the reference counting design:Taskrepresents "the runtime owns this task",Notifiedrepresents "this task has been notified and is pending scheduling",JoinHandlerepresents "someone cares about its result". The three have independent lifetimes—JoinHandlecan be dropped (the task continues running, the result is discarded),Notifieddisappears after poll,Taskis released after the task completes and is removed fromOwnedTasks. If there were only one reference, it would be impossible to express the state "the task is still running but no one is joining".

UnownedTaskis another important branch: it holdstworeference counts, used for blocking tasks (not stored inOwnedTasks)📎 tokio/src/runtime/task/mod.rs:286-295。unowned. The function merges the two references intomem::forget(task)andmem::forget(notified)viaUnownedTask 📎 tokio/src/runtime/task/mod.rs:388-397. The design motivation for "two references" is: blocking tasks do not have aOwnedTaskslist to hold an owned reference, so an extra reference count is needed to ensure the task is not released during execution.

3.2 State bits: how a single usize encodes the entire lifecycle of a task

Intuitive model

Think of task state as a "medical report form" with several independent checkboxes: whether it is being polled, whether it is completed, whether it has been notified, whether it has been cancelled, whether someone is joining. Tokio does not use multiple boolean fields, but instead packs these check bits intoaAtomicUsize. This way, each state transition requires only one CAS instead of multiple locks. Without this design, task state transitions would become nested multiple locks, and both deadlock risk and overhead would soar.

Bitfield layout

StateThe bitfields of📎 tokio/src/runtime/task/mod.rs:32-53:

  • RUNNINGare fully defined in the module documentation: whether the task is being polled or cancelled. 📎 tokio/src/runtime/task/mod.rs:37-38。
  • COMPLETEThis bit also serves as the task's lockRUNNING: the future has fully completed and been dropped. Once set, it is never cleared, and is never set at the same time as📎 tokio/src/runtime/task/mod.rs:40-41。
  • NOTIFIED: whether aNotifiedobject currently exists📎 tokio/src/runtime/task/mod.rs:43。
  • CANCELLED: the task should be cancelled as soon as possible📎 tokio/src/runtime/task/mod.rs:45-46。
  • JOIN_INTEREST: existsJoinHandle 📎 tokio/src/runtime/task/mod.rs:48。
  • JOIN_WAKER: serves as the access control bit for the join handle waker📎 tokio/src/runtime/task/mod.rs:50-51。

The remaining bits are used for the reference count📎 tokio/src/runtime/task/mod.rs:53。

RUNNINGThe fact that the bit serves as a lock is worth elaborating on. The Safety section of the module documentation states: any mutable access to the future must occur after acquiring the lock by modifying theRUNNINGbit, thereby guaranteeing exclusive access📎 tokio/src/runtime/task/mod.rs:130-133. This means that when polling a task, the thread first CASes to setRUNNING, and on success exclusively owns the future; if it fails, it means another thread is polling, and this poll returns directly. This merges "mutual exclusion for poll" and "state transition" into a single atomic operation, avoiding a separate mutex.

JOIN_WAKER access control protocol

JOIN_WAKERThe bit is the most ingenious part of the entire state machine. The problem it solves is:wakerThe field (inTrailer) is accessed concurrently by two threads—the runtime, when the task completes,readsit to wake the joiner,JoinHandleand during pollwritesit to register the waker. The module documentation gives 7 rules📎 tokio/src/runtime/task/mod.rs:75-120:

1. JOIN_WAKERis initially 0.

2. When it is 0,JoinHandlehas exclusive (mutable) access to the waker field.

3. When it is 1,JoinHandlehas only shared (read-only) access.

4. When it is 1 andCOMPLETEis 1, the runtime has shared (read-only) access to the waker field.

5. JoinHandleTo write the waker, it must: (i) successfully setJOIN_WAKERto 0 to obtain exclusive rights, (ii) write the waker, (iii) successfully setJOIN_WAKERto 1.

6. JoinHandlecan only modifyCOMPLETEwhenJOIN_WAKERis 0; the runtime can only modify whenCOMPLETEis 1.

7. IfJOIN_INTERESTis 0 andCOMPLETEis 1, the runtime has exclusive access to the waker field (used to drop the waker).

Rule 6 implies a race: step (i) or (iii) may fail. If (i) fails, abandon writing the waker; if (iii) fails (another thread setCOMPLETEduring this period), then clear the waker field📎 tokio/src/runtime/task/mod.rs:110-120. The essence of this protocol is: use a single atomic bit to dynamically transfer ownership between "writer" and "reader", avoiding a separate lock for the waker field.

Two kinds of reference count decrements

TaskThe drop ofUnownedTask's drop decrements twice:

rust
impl<S: 'static> Drop for Task<S> {
    fn drop(&mut self) {
        if self.header().state.ref_dec() {
            self.raw.dealloc();
        }
    }
}

📎 tokio/src/runtime/task/mod.rs:580-586

rust
impl<S: 'static> Drop for UnownedTask<S> {
    fn drop(&mut self) {
        if self.raw.header().state.ref_dec_twice() {
            self.raw.dealloc();
        }
    }
}

📎 tokio/src/runtime/task/mod.rs:590-596

ref_decreturnstrueindicates this is the last reference, and only then is it truly releasedCellmemory.ref_dec_twiceisUnownedTaskA direct manifestation of holding two counts.

Design consideration: why the state bit and reference count share a single atomic

[Design inference and architectural trade-offs]

Putting the state bit and reference count in the sameAtomicUsizeis to allow the two actions of "decrementing the reference count" and "setting the state bit" to be completed ina single CAS. The module documentation explicitly mentions in the comment atSchedule::release: "The task module will batch-process ref-dec and other option settings"📎 tokio/src/runtime/task/mod.rs:302-304. If the state bit and reference count belonged to two separate atomic variables, then a window would appear between "releasing the last reference" and "marking completion," requiring additional synchronization. After merging,ref_deccan atomically complete "decrement count + check whether it reached zero," avoiding ABA-like problems.

3.3 JoinHandle: how results are passed back across task boundaries

Intuitive model

JoinHandleis like the "meal pickup ticket" a restaurant gives you. When the task (kitchen) completes, it places the dish (output) at the pickup counter (Stage::Finished), then rings your pager (waker). You come to pick it up with the ticket; the ticket itself does not hold the dish, it is just a pointer to the pickup counter. If you lose the ticket (dropJoinHandle), the dish will be thrown away directly (output is dropped), but the kitchen will not stop working because of this.

Data structure

JoinHandle<T>is likewise a transparent wrapper aroundRawTask:

rust
pub struct JoinHandle<T> {
    raw: RawTask,
    _p: PhantomData<T>,
}

📎 tokio/src/runtime/task/join.rs:163-166

PhantomData<T>Marks the output type.JoinHandle<T>Only atT: Sendis itSend/Sync 📎 tokio/src/runtime/task/join.rs:169-170, which ensures that non-Send output will not be moved across threads.

Step-by-Step: awaiting a JoinHandle

JoinHandleimplementsFuture, whosepollis the core of result passing:

rust
fn poll(self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<Self::Output> {
    ready!(crate::trace::trace_leaf());
    let mut ret = Poll::Pending;
    let coop = ready!(crate::task::coop::poll_proceed(cx));
    unsafe {
        self.raw.try_read_output(&mut ret, cx.waker());
    }
    if ret.is_ready() {
        coop.made_progress();
    }
    ret
}

📎 tokio/src/runtime/task/join.rs:327-354

Note several details:trace_leafis used for tracing instrumentation;coop::poll_proceedconsumes the cooperative budget (detailed in Chapter 12);try_read_outputerases generics through the vtable, placing the return value on the stack and passing it in with*mut ()to📎 tokio/src/runtime/task/join.rs:327-354. This "return value on the stack" trick is because vtable functions cannot genericize the return typeT, and can only write back through a raw pointer.

[Design inference and architectural trade-offs]

try_read_outputInternal logic (in raw.rs, source not provided in this chapter): first check theCOMPLETEbit; if it is already set, calltake_outputto take the result fromStage::Finished; otherwise registercx.waker()into theTrailer::wakerfield and returnPending. The registration process follows theJOIN_WAKERprotocol in Section 3.2.

Ownership transfer of the result

The "Non-Send output" section of the module documentation precisely describes the ownership rules for the result📎 tokio/src/runtime/task/mod.rs:151-170:

  • When the task completes, output is placed intoStage, then the transition that "sets COMPLETE" is executed, and theJOIN_INTERESTvalue at that moment is read.
  • IfJOIN_INTERESTis 0 (noJoinHandle), output is dropped immediately📎 tokio/src/runtime/task/mod.rs:157-158。
  • IfJOIN_INTERESTis 1,JoinHandleis responsible for cleaning up output📎 tokio/src/runtime/task/mod.rs:160-161。

For non-Send output, the documentation gives a three-step argument: output is created on the thread that polls the future;JoinHandle<Output>is also non-Send when Output is non-Send, so it is also on the spawn thread; thereforeJoinHandlewill not move output across threads when taking it or dropping it📎 tokio/src/runtime/task/mod.rs:164-170。

JoinHandle's drop: fast and slow paths

rust
impl<T> Drop for JoinHandle<T> {
    fn drop(&mut self) {
        if self.raw.state().drop_join_handle_fast().is_ok() {
            return;
        }
        self.raw.drop_join_handle_slow();
    }
}

📎 tokio/src/runtime/task/join.rs:358-364

drop_join_handle_fasttries to complete "clear theJOIN_INTERESTbit + decrement the reference count" with a single CAS. If it fails (for example, the task is completing and the state bit is occupied), it takes the slow path ofdrop_join_handle_slow. This is a typical "optimistic fast path + pessimistic slow path" pattern.

Design consideration: why JoinHandle does not directly hold output

[Design inference and architectural trade-offs]

IfJoinHandledirectly held output, then output would have to be moved to the thread whereJoinHandleresides when the task completes. ButJoinHandlemay be moved to any thread (as long asT: Send), while the thread that produces output is the poll thread. Directly holding it would cause a cross-thread move where "output is produced on the poll thread but must be dropped on the join thread," directly violating the type system for non-Send output. Tokio chooses to leave output inCell(Stage::Finished),JoinHandleonly holds aCellpointing toRawTask, and when retrieving the result, takes it in place throughtake_output. In this way, output's drop occurs on the thread whereJoinHandleresides, but the premise is that this thread is the same as the poll thread (which holds in the non-Send scenario).

3.4 Local queue: the producer-consumer structure of work-stealing

Intuitive model

Each worker has a "private to-do list" (local queue) with capacity 256. The worker itself takes tasks from thehead(LIFO, exploiting cache locality), while other workers steal tasks from thetail(FIFO, taking the oldest and most likely already-completed tasks). Without a local queue, all tasks would crowd into the global queue, and every task retrieval would contend for the global lock, causing multicore scalability to collapse.

Memory layout: separation of head and tail

rust
pub(crate) struct Inner<T: 'static> {
    head: AtomicUnsignedLong,
    tail: AtomicUnsignedShort,
    buffer: Box<[UnsafeCell<MaybeUninit<task::Notified<T>>>; LOCAL_QUEUE_CAPACITY]>,
}

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:36-57

headisAtomicUnsignedLong(64-bit, if the platform supports u64),tailisAtomicUnsignedShort(32-bit). The comment explains why the indices are wider than actually needed: for ABA mitigation, and to distinguish "full" and "empty" buffers📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:37-49。

headinternally packstwo UnsignedShortThe low bits are the "real head", and the high bits are the "steal head" (the first position the stealer is processing). When the two are equal, there is no active stealer.📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:39-49This dual-value packing is the core trick of the work-stealing queue: the stealer first CAS-updates the steal value to "claim" a batch of tasks, and after completion advances the steal value to catch up with the real value, indicating that stealing has ended.

LOCAL_QUEUE_CAPACITYUnder non-loom it is 256, and under loom it shrinks to 4 to test more boundary cases.📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:62-69。MASK = LOCAL_QUEUE_CAPACITY - 1, used for ring buffer indexing.📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:71。

Step-by-Step: the complete branches of push_back_or_overflow

This is the most complex function of the local queue; let's analyze it branch by branch:

rust
pub(crate) fn push_back_or_overflow<O: Overflow<T>>(
    &mut self,
    mut task: task::Notified<T>,
    overflow: &O,
    stats: &mut Stats,
) {
    let tail = loop {
        let head = self.inner.head.load(Acquire);
        let (steal, real) = unpack(head);
        let tail = unsafe { self.inner.tail.unsync_load() };

        if tail.wrapping_sub(steal) < LOCAL_QUEUE_CAPACITY as UnsignedShort {
            break tail;
        } else if steal != real {
            overflow.push(task);
            return;
        } else {
            match self.push_overflow(task, real, tail, overflow, stats) {
                Ok(_) => return,
                Err(v) => { task = v; }
            }
        }
    };
    self.push_back_finish(task, tail);
}

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:188-223

Three branches:

1. Has capacity(tail - steal < CAPACITY):break tail, after breaking out of the loop, callpush_back_finishto write to the buffer.

2. No capacity but there are concurrent stealers(steal != real): the stealer will free up space, so just push the current task into the injection queue and return immediately.📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:204-208。

3. No capacity and no stealers: callpush_overflowto overflow the latter half batch of tasks to the injection queue📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:209-219. If the CAS fails (loses to a concurrent stealer),push_overflowreturnsErr(task), and the loop retries.

push_back_finishWrite the task and update tail:

rust
fn push_back_finish(&self, task: task::Notified<T>, tail: UnsignedShort) {
    let idx = tail as usize & MASK;
    self.inner.buffer[idx].with_mut(|ptr| {
        unsafe { ptr::write((*ptr).as_mut_ptr(), task); }
    });
    self.inner.tail.store(tail.wrapping_add(1), Release);
}

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:226-244

ReleaseThe ordering guarantees that the written task is visible to stealers.

push_overflow: why overflow the latter half batch

rust
const NUM_TASKS_TAKEN: UnsignedShort = (LOCAL_QUEUE_CAPACITY / 2) as UnsignedShort;

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:265

When overflowing, take 128 tasks. The comment explains in detail why to takethe latter half batchrather than the former half batch📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:295-306: when taking tasks from the injection queue, they are always placed in the former half. So if a task is in the latter half, it can be determined that it was not just taken from the injection queue. This guarantees that "a task taken from the injection queue will not be immediately put back into the injection queue" (at least before it has been polled once).

CAS claims the latter half batch:

rust
if self.inner.head.compare_exchange_weak(
    pack(head, head), pack(tail, tail), Release, Relaxed
).is_err() {
    return Err(task);
}

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:283-293

Updateheadfrom(head, head)to(tail, tail), that is, advance both steal and real to tail, claiming all tasks. After success, roll back tail totail + NUM_TASKS_TAKEN, indicating that the former half batch remains in the local queue.📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:314-316。

pop and steal_into: the two paths for taking tasks

popis the worker itself taking a task (from the head, LIFO):

rust
pub(crate) fn pop(&mut self) -> Option<task::Notified<T>> {
    let mut head = self.inner.head.load(Acquire);
    let idx = loop {
        let (steal, real) = unpack(head);
        let tail = unsafe { self.inner.tail.unsync_load() };
        if real == tail { return None; }
        let next_real = real.wrapping_add(1);
        let next = if steal == real {
            pack(next_real, next_real)
        } else {
            assert_ne!(steal, next_real);
            pack(steal, next_real)
        };
        let res = self.inner.head.compare_exchange_weak(head, next, AcqRel, Acquire);
        match res {
            Ok(_) => break real as usize & MASK,
            Err(actual) => head = actual,
        }
    };
    Some(self.inner.buffer[idx].with(|ptr| unsafe { ptr::read(ptr).assume_init() }))
}

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:361-399

Key branch: ifsteal == real(no stealer), advance both at the same time; otherwise advance only real and leave steal unchanged.📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:377-384。assert_ne!(steal, next_real)ensures that real will not be advanced to steal's position, otherwise it would corrupt the stealer's claim state.

steal_intois the stealing path, first checking whether the target queue has enough space:

rust
if dst_tail.wrapping_sub(steal) > LOCAL_QUEUE_CAPACITY as UnsignedShort / 2 {
    return None;
}

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:431-435

If the target queue is more than half full, do not steal, to avoid immediately overflowing again after stealing.

steal_into2is the core of stealing, calculating the number to steal:

rust
let n = src_tail.wrapping_sub(src_head_real);
let n = n - n / 2;

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:487-488

Steal half (rounded up). Then CAS-update head's steal value to claim:

rust
let steal_to = src_head_real.wrapping_add(n);
next_packed = pack(src_head_steal, steal_to);
let res = self.0.head.compare_exchange_weak(prev_packed, next_packed, AcqRel, Acquire);

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:496-506

Note that only the real value is updated here (pack(src_head_steal, steal_to)steal remains unchanged insteal_to), advancing real to

rust
loop {
    let head = unpack(prev_packed).1;
    next_packed = pack(head, head);
    let res = self.0.head.compare_exchange_weak(prev_packed, next_packed, AcqRel, Acquire);
    match res {
        Ok(_) => return n,
        Err(actual) => prev_packed = actual,
    }
}

📎 tokio/src/runtime/scheduler/multi_thread/queue.rs:548-561

The sequence diagram below depicts the three-party concurrent interaction of "producer push, consumer pop, stealer steal":

mermaid
sequenceDiagram
    participant P as "Worker A (生产者)"
    participant Q as "Local 队列 Inner"
    participant C as "Worker A (消费者 pop)"
    participant S as "Worker B (窃取者)"

    P->>Q: "load head (Acquire)"
    P->>Q: "unsync_load tail"
    Note over P: "tail - steal < 256?"
    P->>Q: "push_back_finish: buffer[idx] = task"
    P->>Q: "store tail+1 (Release)"

    C->>Q: "load head (Acquire)"
    C->>Q: "unsync_load tail"
    Note over C: "real == tail? 空则返回 None"
    C->>Q: "CAS head: pack(real+1, real+1)"
    Q-->>C: "Ok, 读取 buffer[real & MASK]"

    S->>Q: "load head (Acquire)"
    S->>Q: "load tail (Acquire)"
    Note over S: "src_head_steal != src_head_real? 返回 0"
    S->>Q: "CAS head: pack(steal, real+n) 认领一半"
    Q-->>S: "Ok, 拷贝 n 个任务到 dst"
    S->>Q: "CAS head: pack(real+n, real+n) 完成窃取"
    Q-->>S: "返回 n"

Design thinking: why the local queue is LIFO while stealing is FIFO

{Design inference and architectural trade-offs}

The worker itself takes from the head (LIFO), because the most recently pushed task is most likely still in the CPU cache, and is most likely the task that was "just woken up and whose data is still hot." The stealer takes from the tail (FIFO), because the oldest task is most likely to have already completed most of its work, and stealing it can reduce the victim's load the fastest. This combination of "LIFO local + FIFO stealing" is the classic design of work-stealing scheduling, balancing cache locality and load balancing.

At this point, the task has completed its transformation from a Future to a schedulable entity: it has been assigned a reference count, placed intoCell's memory layout, and successfully delivered to the worker's local queue or the global injection queue. But putting a task into a queue is only the beginning; what really makes it run is the worker thread's scheduling loop. In the next chapter we will enter the second half of "the life of a task", tracing how the worker takes tasks from the queue, callsFuture::poll, and upon returningPendingregisters a wakeup throughWaker, ultimately triggeringscheduleto re-enqueue—the complete call path of the closed loop "wakeup -> enqueue -> poll again", as well as the work-stealing strategy and LIFO slot optimization, will all be revealed there.

CHAPTER 04

Chapter 4: The Life of a Task (Part 2): The Closed Loop of the Scheduling Loop, Poll, and Wakeup

Project: tokio-rs/tokio · Book progress: Chapter 4 / 14 · Verification status: FACT line numbers are truly anchored

From queue to execution: the skeleton of the worker main loop

In the previous chapter we sent tasks into theLocalqueue or the global injection queue. But the queue is only a "to-do list"; what really makes tasks run is the never-ending loop in the worker thread. In this chapter we traceContext::run—it is the heart of the entire multi-threaded scheduler.

First build intuition: a worker thread is like a chef, with a stack of their own orders in front of them (run_queue), and also a public order rack nearby (inject). The chef first looks at the nearest one at hand (lifo_slot), if not, take from their own stack, if still not, grab a handful from the public shelf, and if that still doesn't work, steal a few from other chefs' stacks. Only when everything is empty does he go rest, but even while resting his ears stay perked—as soon as an order comes in, he wakes up immediately.

Without this loop, once a task is enqueued it would forever lie in the queue,Future::pollnever to be invoked, and the entire runtime would just be a pile of dead data.

Core's memory layout and state fields

The worker's mutable state is all contained inCore, which isBoxallocated on the heap, and passed betweenAtomicCell<Core>viaWorkerand the thread-localContext.

CoreThe key fields of📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:113-167:

  • tick: u32are as follows: incremented each loop iteration, used to periodically trigger maintenance (maintenance) and global queue checks.
  • lifo_slot: Option<Notified>:LIFO slot, this is the most ingenious design in this chapter. When a worker schedules a task itself, it does not go intorun_queue, but instead places it into this slot, and the next time it fetches a task itprioritizestaking from here.
  • lifo_enabled: bool: A switch for the LIFO slot, used to prevent starvation in ping-pong scenarios.
  • run_queue: queue::Local<Arc<Handle>>: The local queue, theLocalstructure analyzed in the previous chapter.
  • is_searching: bool: Whether the worker is currently searching for stealable tasks.
  • is_shutdown: bool / is_traced: bool: Shutdown and tracing flags.
  • park: Option<Parker>: The parker, wrapped withOptionto conveniently take out/put back under the borrow checker.
  • global_queue_interval: u32: How often to check the global queue.
  • rand: FastRand: A fast random number generator, used to randomly select the stealing start point.
[Design inference and architectural trade-offs]

Note thatlifo_slotis aOption<Notified>rather than a queue—it only storesonetask. The design motivation is clearly stated in the source code comments📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:117-121: tasks scheduled by the worker itself are stored in this slot, and the worker checks itrun_queue beforechecking

, with the effect that "the last scheduled task runs next" (LIFO). This is to improve locality, is especially effective for message-passing patterns, and can reduce latency.

Why can LIFO reduce latency? Consider a typical message-passing scenario: task A finishes processing a message and wakes task B, and B finishes processing and wakes A again. If B runs immediately after A wakes it, the data B needs is very likely still in the CPU cache (because A just touched it). If B is pushed to the tail of the queue, by the time the dozens of tasks ahead of it finish, the cache will long since have been flushed.MAX_LIFO_POLLS_PER_TICK = 3But LIFO has a starvation risk. The source code uses📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:263-263to limit

: each tick prioritizes the LIFO slot at most 3 times, after which it is disabled to give other tasks a chance to execute.

Main loop walkthrough: one complete scheduling cycleparkLet us plug in a concrete scenario: worker 0 has just woken up fromrun_queue,lifo_slothas 5 tasks,

has 1 task, and the global queue has 3 tasks.Context::run 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:570-642The main loop entry islifo_enabled. It first resetsblock_in_place(because the core may have been stolen by📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:571-573, and the state needs to be restored)while !core.is_shutdown, then enters the

loop.

Each loop iteration does four things: core.tick()Step one: tick and maintenance.📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:587increments the counterself.maintenance(core). Thentick % event_interval == 0checkspark_yield, and if so calls📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:809-826。

to drive I/O and timers with a 0 timeout core.next_task(&self.worker)Step two: fetch a task.📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1090-1156is the core task-fetching logic

  • . It has two paths:tick % global_queue_interval == 0When,prioritize📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1091-1098taking from the global queue, and if that fails, take from the local
  • . This is to prevent tasks in the global queue from starving.Otherwiseprioritize📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1090-1156。

taking local tasksnext_local_taskLocal task fetching is done by📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1158-1160:

rust
fn next_local_task(&mut self) -> Option<Notified> {
    self.lifo_slot.take().or_else(|| self.run_queue.pop())
}

first take from the LIFO slot, then take from the head of the queue (LIFO pop). This is the "local LIFO" mentioned in the previous chapter.

If the local queue is empty but the global queue is non-empty, the worker willbatchpull tasks from the global queue📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1110-1154. The calculation of the batch sizenis quite particular:min(inject.len() / remotes.len() + 1, cap), wherecapis again taken asmin(remaining_slots, max_capacity / 2). The source code comments explain why it is limited to half the queue capacity📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1120-1131: to ensure that the pulled tasks land in thefirst halfof the local queue, so that even if overflow occurs later, these tasks will not be pushed back to the global queue (overflow only affects the second half).

Step three: run the task.After obtaining the task, callsrun_task 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:647-796. This is the most complex function in this chapter, and we will expand on it specifically in the next section.

Step four: steal or park.Ifnext_taskreturnsNone, it means there is no work left locally or globally, so callssteal_work 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1167-1195. If stealing fails, it entersparkorpark_yield 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:613-621。

The entire control flow is as follows:

mermaid
flowchart TD
    start["Context::run 进入循环"] --> tick["core.tick() 自增"]
    tick --> maint{"tick % event_interval == 0?"}
    maint -->|是| park_yield["park_yield 驱动 I/O 与定时器"]
    maint -->|否| next
    park_yield --> next["core.next_task()"]
    next --> has_task{"取到任务?"}
    has_task -->|是| run_task["run_task 执行 poll"]
    run_task --> cont{"core 还在?"}
    cont -->|是| tick
    cont -->|否| ret["return 退出"]
    has_task -->|否| steal["core.steal_work()"]
    steal --> steal_ok{"窃取成功?"}
    steal_ok -->|是| run_task
    steal_ok -->|否| defer_check{"defer 非空?"}
    defer_check -->|是| py["park_yield"]
    defer_check -->|否| pk["park 阻塞等待"]
    py --> tick
    pk --> tick

run_task: the closed loop of poll and the LIFO slot

run_taskis where the task is actuallypoll, and also the closing point of the "wake -> enqueue -> poll again" loop.

The first thing after entering the function isassert_owner 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:648, convertingNotifiedintoTask, while asserting that the current thread is indeed the owner of this task (debug assertion).

Next istransition_from_searching 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:652—if the worker was previously in the searching state, now that it has found a task, it must exit the searching state and may wake other parked workers.

Then comes the key budget wrapper📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:695-795:

rust
coop::budget(|| {
    task.run();
    let mut lifo_polls = 0;
    loop {
        let mut core = match self.core.borrow_mut().take() {
            Some(core) => core,
            None => return ControlFlow::Break(()),
        };
        let task = match core.lifo_slot.take() {
            Some(task) => task,
            None => {
                self.reset_lifo_enabled(&mut core);
                core.stats.end_poll();
                return ControlFlow::Continue(core);
            }
        };
        if !coop::has_budget_remaining() {
            core.run_queue.push_back_or_overflow(task, ...);
            return ControlFlow::Continue(core);
        }
        lifo_polls += 1;
        if lifo_polls >= MAX_LIFO_POLLS_PER_TICK {
            core.lifo_enabled = false;
        }
        let task = self.worker.handle.shared.owned.assert_owner(task);
        *self.core.borrow_mut() = Some(core);
        task.run();
    }
})

This code reveals the complete closed loop of the LIFO slot:task.run()executesFuture::poll, and during polling, if the task wakes itself or another task,schedule_localwill place the new task intolifo_slot 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1396-1408. After poll returns, the loop immediately checkslifo_slot, and if there is a task it continues running—without returning to the main loop, directly polling continuously within the same budget.

This is the manifestation of "wake -> enqueue -> poll again" on the LIFO path: when woken, the task is placed intolifo_slot, and immediately after poll returns it is taken out and polled again, forming a tight closed loop.

Note theself.core.borrow_mut().take()branch ofNone:📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:716-724if the core has been stolen (for example, if the task calledblock_in_place), the worker must returnControlFlow::Break(()), lettingContext::runexit. This isblock_in_placeInteraction points with the scheduling loop.

Wake path: How the Waker triggers re-enqueueing

WhenFuture::pollreturnsPendingthe task needs to register aWakerand be woken when the event is ready. Tokio'sWakerimplementation is extremely lean—it's just a raw pointer to the task'sHeaderplus a vtable.

waker_refConstructWakerRef 📎 tokio/src/runtime/task/waker.rs:11-34wrapManuallyDropwithWakerto avoid decrementing the reference count on drop. The vtable is a static📎 tokio/src/runtime/task/waker.rs:119-119:

rust
static WAKER_VTABLE: RawWakerVTable =
    RawWakerVTable::new(clone_waker, wake_by_val, wake_by_ref, drop_waker);

All four functions simply restore the raw pointer toHeaderand then callRawTask's corresponding method📎 tokio/src/runtime/task/waker.rs:70-116For example,wake_by_refultimately callsraw.wake_by_ref() 📎 tokio/src/runtime/task/waker.rs:106-116。

wake_by_refThe semantics are: transition the task state fromPENDINGtoSCHEDULEDand if the transition succeeds (i.e., it was indeed PENDING), callSchedule::scheduleto re-enqueue the task.

For the multi-threaded scheduler,schedule's implementation is inHandle::schedule_task 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1353-1376:

rust
pub(super) fn schedule_task(&self, task: Notified, is_yield: bool) {
    with_current(|maybe_cx| {
        if let Some(cx) = maybe_cx {
            if self.ptr_eq(&cx.worker.handle) {
                if let Some(core) = cx.core.borrow_mut().as_mut() {
                    self.schedule_local(core, task, is_yield);
                    return;
                }
            }
        }
        self.push_remote_task(task);
        self.notify_parked_remote();
    });
}

The logic branches into two paths:

  • If the current thread is this scheduler's worker and holds the core, go throughschedule_local 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1385-1417—place it into the LIFO slot or the local queue.
  • Otherwise (woken from an external thread, or the core was stolen), go throughpush_remote_taskpush into the global injection queue andnotify_parked_remotewake a parked worker📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1379-1383。

schedule_localInternally it branches again📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1385-1417: if it'syieldor LIFO is disabled, push into therun_queuetail; otherwise place it intolifo_slotand push the task originally in the slot to the tail of the queue.

mermaid
sequenceDiagram
    participant Future as "Future::poll"
    participant Waker as "Waker(wake_by_ref)"
    participant RawTask as "RawTask::wake_by_ref"
    participant Handle as "Handle::schedule_task"
    participant Core as "Core(schedule_local)"
    participant Inject as "InjectQueue"
    participant Parker as "Unparker"

    Future->>Waker: "返回 Pending, 注册 waker"
    Note over Future: "事件就绪(如 epoll)"
    Waker->>RawTask: "raw.wake_by_ref()"
    RawTask->>RawTask: "state: PENDING -> SCHEDULED"
    RawTask->>Handle: "schedule(Notified)"
    alt 当前线程是同一 worker 且持有 core
        Handle->>Core: "schedule_local: 放入 lifo_slot"
    else 外部线程或 core 被偷走
        Handle->>Inject: "push_remote_task"
        Handle->>Parker: "notify_parked_remote().unpark()"
    end

park and unpark: Atomicity of the state machine and wakeup

When a worker has nothing to do it must park, but park/unpark is the most race-prone area. Tokio uses aAtomicUsizestate machine plusCondvaras a fallback to solve this.

Inner's field📎 tokio/src/runtime/scheduler/multi_thread/park.rs:31-43:state: AtomicUsize、mutex: Mutex<()>、condvar: Condvar、shared: Arc<Shared>There are four state constants📎 tokio/src/runtime/scheduler/multi_thread/park.rs:36-45:

  • EMPTY = 0: not parked.
  • PARKED_CONDVAR = 1: parked on the condvar.
  • PARKED_DRIVER = 2: parked on the I/O driver.
  • NOTIFIED = 3: already woken.

This is an explicit state machine, and we use it to draw the state diagram (this is the only place in this chapter that meets thestateDiagram-v2admission criteria—the source code really does have these four state constants):

mermaid
stateDiagram-v2
    [*] --> Empty
    Empty --> ParkedCondvar : "park_condvar() CAS(EMPTY->PARKED_CONDVAR)"
    Empty --> ParkedDriver : "park_driver() CAS(EMPTY->PARKED_DRIVER)"
    Empty --> Notified : "unpark() swap(NOTIFIED)"
    ParkedCondvar --> Empty : "condvar 唤醒后 CAS(NOTIFIED->EMPTY)"
    ParkedCondvar --> Empty : "超时 swap(EMPTY)"
    ParkedDriver --> Empty : "driver 返回后 swap(EMPTY)"
    Notified --> Empty : "park() CAS(NOTIFIED->EMPTY) 消费通知"
    Notified --> Notified : "再次 unpark() swap(NOTIFIED)"

unpark's implementation📎 tokio/src/runtime/scheduler/multi_thread/park.rs:277-290usesswapinstead of CAS; the source comments explain why📎 tokio/src/runtime/scheduler/multi_thread/park.rs:277-290: a release operation must be performed so that the parked thread observes the writes before unpark, so even if state is alreadyNOTIFIEDit must still be written once.

parkfirst tries to consume an existing notification📎 tokio/src/runtime/scheduler/multi_thread/park.rs:132-149: if CASNOTIFIED -> EMPTYsucceeds, it means it was already woken, so return directly without blocking. Otherwise it tries to acquire the driver lock; if acquired, park on the driver; if not, fall back to the condvar📎 tokio/src/runtime/scheduler/multi_thread/park.rs:143-148。

park_condvarThere is a classic double-check📎 tokio/src/runtime/scheduler/multi_thread/park.rs:162-180: first CASEMPTY -> PARKED_CONDVAR, and if it fails and it'sNOTIFIED, it means it was woken before the state was set, so at this point it mustswap(EMPTY)to synchronize the unpark write📎 tokio/src/runtime/scheduler/multi_thread/park.rs:167-177The comment specifically emphasizes: even if you know it'sNOTIFIEDyou must still read once, because unpark may have been called again after we readNOTIFIED.

unpark_condvar's comment📎 tokio/src/runtime/scheduler/multi_thread/park.rs:292-307points out the classic condvar trap: there is a window between the parked thread setting thePARKEDstate and actuallywait, and if notify happens during this window it will be ignored. The solution is that the park thread holdsmutexat this time, and the unpark thread firstdrop(self.mutex.lock())acquires the lock (thereby waiting for the park thread to release), thennotify_one。

Design thinking: Why the LIFO slot is a single slot rather than a queue

[Design inference and architectural trade-offs]

The single-slot design is a deliberate trade-off. If a queue were used, every wakeup would require enqueueing and every task fetch would require dequeueing, which is more expensive; moreover, the queue would accumulate multiple tasks, breaking the locality assumption of "the most recently woken runs first." The semantics of a single slot are "remember only the most recent one," and the evicted task goes into the normal queue—this exactly matches the law of diminishing locality returns: the most recent task is the hottest, the second is next, and beyond the third the benefit becomes very small.

MAX_LIFO_POLLS_PER_TICK = 3This magic number📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:263-263is also an empirical value. The source comment says "running a few times through the LIFO slot seems enough to benefit from locality; more than 3 times may over-weight it." This prevents the ping-pong scenario where A wakes B and B wakes A from starving other tasks.

Another noteworthy design issteal_work's "half-search" strategy📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1158-1160: a new worker truly attempts to steal only when fewer than half of the workers are searching. This avoids CAS contention caused by all workers frantically stealing at the same time.transition_to_searchingcoordinates throughidle.transition_worker_to_searching()to📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1197-1203。

Stealing starts from a random starting point📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1172-1174, iterates over all remotes, skips itself📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1179-1182, and callssteal_intoto attempt stealing. After all fail, it falls back to the global queue📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1197-1203。

Chapter summary

The worker main loopContext::runis the heart of the scheduler: after each tick it first fetches a task (LIFO slot → local queue → global queue); if it gets one, itrun_taskexecutes poll; if not, it steals; if stealing fails, it parks.run_taskThe internal LIFO loop compresses "wake → enqueue → poll again" within the same budget, forming a low-latency closed loop.Wakeris a raw pointer plus a static vtable,wake_by_reftriggers through a state transitionschedule, and decides whether to go through the local queue or the global queue based on whether the current thread is the same worker.park/unparkuses a four-state atomic machine plus a condvar fallback to solve the classic race of lost wakeups.

In the next chapter we will leave the scheduler and enter the I/O world: how the Reactor translates epoll events intoWakerwakeups, turningAsyncFd'sPendingintoReady。

Chapter review and self-test

Q1: Ifnext_local_taskis changed to first fetchrun_queueThen takelifo_slot, what are the consequences in message-passing-intensive scenarios?

Reference analysis:next_local_taskThe current implementation isself.lifo_slot.take().or_else(|| self.run_queue.pop()) 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:1158-1160, taking the LIFO slot first. If instead we takerun_queuefirst, then tasks that were just woken up and whose data is still hot would be scheduled to execute after other tasks in the queue. In an A→B→A message-passing pattern, B, after being woken up, would not run immediately but would wait for other tasks in the queue to finish; by then, the data written by A may have been evicted from the CPU cache, and the locality benefit is lost. More seriously,lifo_slottasks in would wait untilrun_queueis emptied before being executed, causing a significant increase in latency. The source code comment📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:117-121explicitly states that this order is intended to "improve locality, benefit from message-passing patterns, and reduce latency."

Q2: park_condvarIn , if we removeErr(NOTIFIED)from the branchself.state.swap(EMPTY, SeqCst)and keep onlyreturn, what problems would arise?

Reference analysis: The source code executesErr(NOTIFIED)in the branchlet old = self.state.swap(EMPTY, SeqCst) 📎 tokio/src/runtime/scheduler/multi_thread/park.rs:167-177. The comment explains📎 tokio/src/runtime/scheduler/multi_thread/park.rs:168-173: unpark may be called again after we readNOTIFIED, so an acquire operation must be performed to synchronize with that unpark in order to observe all writes before it. If we onlyreturnwithout swapping, state would remain atNOTIFIED, and on the next park, CASNOTIFIED -> EMPTYwould succeed and return immediately (consuming an already-expired notification). But worse, the release write of unpark would not be synchronized, and the parking thread might not see the data written before unpark, leading to memory visibility issues. This is a classic double bug of "lost wakeup + memory ordering."

Q3: run_taskIn , whenself.core.borrow_mut().take()returnsNone, why returnControlFlow::Break(())instead ofContinue?

Reference analysis:self.core.borrow_mut().take()ReturningNonemeans the core has already been stolen📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:716-724. The only way the core can be stolen is if a task internally callsblock_in_place, which throughmaybe_move_runtimetakes the core out ofcx.coreand hands it to a new thread📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:473-497. At this point, the current thread no longer holds scheduling capability. If it returnsContinue,Context::run, it would continue looping and callcore.next_task()and other methods that require the core, but the core is no longer inself.core, causing a panic or inconsistent state. ReturningBreakletsContext::rundirectlyreturn 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:594-597, handing control back to therunfunction, which handles the rest (such ascx.defer.wake() 📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:564). The comment also states📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:719-721: at this point,reset_lifo_enabledcannot be called because the core has been stolen, and the thief will handle it at the top ofContext::run.

CHAPTER 05

Chapter 5: I/O readiness notification: How the Reactor translates epoll events into Waker wakeups

Project: tokio-rs/tokio · Book progress: Chapter 5 / 14 · Verification status: FACT line numbers truly anchored

In the previous chapter, we traced the main loop of a worker thread: a task is polled, and when it returns Pending, the Waker is stored somewhere; after the event becomes ready, the Waker is triggered and the task is re-enqueued. But where exactly is "somewhere"? How is the Waker found again when an epoll event arrives? This is exactly the question the Reactor must answer. Let's first build an intuitive model: imagine the entire I/O readiness notification mechanism as a restaurant's order pickup calling system—after customers (tasks) place their orders, they do not stand at the window waiting forever, but take a pager (Waker) back to their seats; after the kitchen (kernel epoll) finishes the meal, the front desk (Reactor) finds the corresponding pager by order number (Token) and presses the button. Without this system, each task could only poll the socket, burning up CPU; or it could use blocking threads to wait, one thread per connection, which does not scale. Tokio's Reactor consists of three files forming a three-layer structure with strictly separated responsibilities: driver.rs is the event loop itself, holding mio::Poll, responsible for calling poll() to block waiting for kernel events and translating events into reads and writes on ScheduledIo; registration.rs is the user-facing registration handle, which is what TcpStream holds internally, providing APIs such as poll_read_ready / poll_write_ready; scheduled_io.rs is the state slot for each fd, storing read/write readiness bits and the Waker list, and is the bridge between events and tasks. For the module assembly relationship, see tokio/src/runtime/io/mod.rs:5-16: driver exports Driver, Handle, ReadyEvent; registration exports Registration; scheduled_io exports ScheduledIo. The following diagram anchors the complete data flow to be traced in this chapter: TcpStream → Registration → ScheduledIo → Handle/Driver → kernel → back to ScheduledIo → Waker. Next, we will break it down layer by layer.

Driver layer:DriverandHandledivision of responsibilities

Intuitive model

Driveristhe only entity that ownsmio::Poll, and it can only beaccessed in a single thread—this is the exclusivity requirement of the event loop. While&mutisHandlea registration entry point that is cloneable and shareable across threads可克隆、可跨线程共享的注册入口, any thread that wants to register a new fd goes through it. Without this split, eithermio::Pollwould need to be locked (contending on every registration), or all registrations would have to go back to the driver thread (introducing a cross-thread message queue). Tokio chooses to letHandledirectly hold a clone ofmio::Registry, so registration operations can proceed concurrently, and only actual event waiting requires exclusivity.

Memory Layout and Fields

First look atDriver's fields📎 tokio/src/runtime/io/driver.rs:25-38:

  • signal_ready: bool: whether a Unix signal event has arrived, used for signal driving.
  • events: mio::Events: the main event buffer, reused acrossturncalls to avoid allocating each time.
  • events_busy: Option<mio::Events>:Dedicated buffer for non-blocking poll, present only whenmax_io_events_per_busy_tickis set.
  • poll: mio::Poll: a wrapper around the kernel event queue.

Next look atHandle 📎 tokio/src/runtime/io/driver.rs:41-75:

  • registry: mio::Registry:mio::Poll::registry()'s clone, used forregister/deregister。
  • registrations: RegistrationSet: the set of all active registrations, responsible for allocatingTokenandScheduledIo。
  • synced: Mutex<registration_set::Synced>: protects the synchronization state ofRegistrationSet.
  • waker: mio::Waker: used to wake up the driver blocked inturnfrom any thread.
  • metrics: IoDriverMetrics: tracks the number of fds and ready events.

There is a key design here:events_busythe existence of📎 tokio/src/runtime/io/driver.rs:25-38is meant to solvethe problem that non-blocking poll would swallow events. The comment📎 tokio/src/runtime/io/driver.rs:189-190makes it clear: if events taken by a non-blocking poll were left in the main buffer, the next poll would not see them; with a separate buffer, unprocessed events remain in the kernel queue and will be returned again on the next poll.

Step-by-Step: One Execution ofturn

turnis the core function of the driver📎 tokio/src/runtime/io/driver.rs:184-261. Suppose a worker thread finds there is no task to run and callspark → turn(handle, None)to block and wait:

Step One: assert it is not shutdown📎 tokio/src/runtime/io/driver.rs:185, and release registrations pending cleanup📎 tokio/src/runtime/io/driver.rs:187。release_pending_registrationscheckneeds_release(), and if present callregistrations.release() 📎 tokio/src/runtime/io/driver.rs:336-340。

Step Two: choose the event buffer📎 tokio/src/runtime/io/driver.rs:191-194. Ifmax_waitis zero andevents_busyexists, use the busy buffer; otherwise use the main buffer.

Step Three: callself.poll.poll(events, max_wait) 📎 tokio/src/runtime/io/driver.rs:198. This is where it truly blocks on epoll_wait. Error handling is very restrained:Interruptedis ignored directly (signal interruption is normal)📎 tokio/src/runtime/io/driver.rs:200, under WASIInvalidInputis also ignored📎 tokio/src/runtime/io/driver.rs:201-205, and other errors panic directly📎 tokio/src/runtime/io/driver.rs:206。

Step Four: iterate over events📎 tokio/src/runtime/io/driver.rs:211-233. For eachevent:

  • iftoken == TOKEN_WAKEUP(value 0)📎 tokio/src/runtime/io/driver.rs:214, do nothing—this is whatunparkuses to interrupt blocking.
  • Iftoken == TOKEN_SIGNAL(value 1)📎 tokio/src/runtime/io/driver.rs:216, setsignal_ready = true。
  • Otherwise it is a normal I/O event📎 tokio/src/runtime/io/driver.rs:218-231: convertmio::Readyinto Tokio'sReady, useEXPOSE_IO.from_exposed_addr(token.0)to turn the token back into a*const ScheduledIopointer, thenset_readiness(Tick::Set, |curr| curr | ready)accumulate the readiness bits, and thenio.wake(ready)trigger the corresponding direction'sWaker。

HereEXPOSE_IOis aPtrExposeDomain<ScheduledIo> 📎 tokio/src/runtime/io/mod.rs:21-22, which "exposes" the pointer as ausizeasmio::Token. The safety comment📎 tokio/src/runtime/io/driver.rs:222-225explains why this unsafe conversion is safe: the pointer will not be freed before it is deregistered from mioandthe driver no longer concurrently polls, and the driver holds ownership ofArc<ScheduledIo>.

Step Five: handle the io_uring completion queue (Linux + tokio_unstable only)📎 tokio/src/runtime/io/driver.rs:235-258, including the flush loop when the CQ overflows.

Step Six: accumulate metrics📎 tokio/src/runtime/io/driver.rs:265-267。

mermaid
flowchart TD
    start["turn(handle, max_wait)"] --> assert["debug_assert!(!is_shutdown)"]
    assert --> release["release_pending_registrations()"]
    release --> pick{"max_wait == 0<br/>且 events_busy 存在?"}
    pick -->|是| busy["events = events_busy"]
    pick -->|否| main["events = events"]
    busy --> poll["poll.poll(events, max_wait)"]
    main --> poll
    poll --> pollres{"poll 返回?"}
    pollres -->|"Ok / Interrupted"| iter["遍历 events.iter()"]
    pollres -->|"其他 Err"| panic["panic!(unexpected error)"]
    iter --> tok{"event.token()?"}
    tok -->|"TOKEN_WAKEUP"| skip["忽略,仅用于打断阻塞"]
    tok -->|"TOKEN_SIGNAL"| sig["signal_ready = true"]
    tok -->|"普通 fd token"| cast["EXPOSE_IO.from_exposed_addr(token.0)"]
    cast --> setr["io.set_readiness(Tick::Set, curr | ready)"]
    setr --> wake["io.wake(ready)"]
    wake --> iter
    skip --> iter
    sig --> iter
    iter --> uring["dispatch_completions() (io-uring)"]
    uring --> metrics["metrics.incr_ready_count_by(ready_count)"]

Design Reflection: WhyHandlemust holdmio::Waker

unpark 📎 tokio/src/runtime/io/driver.rs:280-283callsself.waker.wake(). Thismio::Wakeris used inDriver::newto registerTOKEN_WAKEUPwith📎 tokio/src/runtime/io/driver.rs:124. When the driver is blocked inpoll.poll(), another thread callingunparkwill stuff aTOKEN_WAKEUPevent into epoll,pollreturns immediately, and when iterating it sees this token and skips it directly📎 tokio/src/runtime/io/driver.rs:214-215。

[Design Inference and Architectural Trade-offs]

This mechanism is used inderegister_source📎 tokio/src/runtime/io/driver.rs:315-334: after deregistering a source, ifregistrations.deregisterreturns true (indicating this is the last reference), thenunpark(). Why? Because the driver may be blocked inpollwaiting for events on this fd, and the fd has already been deregistered, so the kernel will no longer produce events; the driver must be actively woken up so it can re-check the registration set and possibly exit blocking. Otherwise the driver would sleep untilmax_waittimes out, delaying shutdown.

Another detail:deregister_sourcefirst callsself.registry.deregister(source) 📎 tokio/src/runtime/io/driver.rs:322, then cleans upregistrations 📎 tokio/src/runtime/io/driver.rs:315-334. The comment📎 tokio/src/runtime/io/driver.rs:320-321says "Cleanup ALWAYS happens"—even if OS-level deregistration fails, internal state must still be cleaned up, and only then is the OS error returned📎 tokio/src/runtime/io/driver.rs:336-340. This is the typicalresource cleanup takes priority over error propagationpattern.

Registration Layer:RegistrationHowWakeris stored intoScheduledIo

Intuitive Model

Registrationisthe contract between a task and an fd. It holds two things: ascheduler::Handle(used to access the runtime when needed), and aArc<ScheduledIo>(the state slot for the fd). When a task callspoll_read_ready,RegistrationhandsWakerover toScheduledIofor safekeeping; when the driver receives an event, it takesScheduledIoout ofWakerto wake it up.

Memory Layout and Fields

Registrationhas only two fields📎 tokio/src/runtime/io/registration.rs:46-54:

  • handle: scheduler::Handle: the runtime handle, with comment📎 tokio/src/runtime/io/registration.rs:46-54saying "TODO: this can probably be moved into ScheduledIo", indicating the author thinks this field's placement can be optimized.
  • shared: Arc<ScheduledIo>: shared state,Arcensuring both the driver and the task can access it.
[Design Inference and Architectural Trade-offs]

Note thatRegistrationmanually implementsSendandSync 📎 tokio/src/runtime/io/registration.rs:57-58. Why is an unsafe impl needed? Becausescheduler::Handlemay internally contain fields that are notSend/Sync(such asRc), butRegistration's usage scenario requires it to be able to cross threads. The doc comment📎 tokio/src/runtime/io/registration.rs:28-33gives the key constraint:the caller must guarantee that at most two tasks concurrently use the sameRegistration, one reading and one writing. Violating this constraint is still memory-safe, but will cause lost notifications and task hangs.

Step-by-Step:poll_read_ready's call chain

Suppose a task is inTcpStream::poll_readdiscovers that the socket has no data and needs to register read interest. The call chain isTcpStream::poll_read_priv → PollEvented::poll_read → Registration::poll_read_io → poll_io → poll_ready。

poll_readyis the core📎 tokio/src/runtime/io/registration.rs:155-171:

Step one:trace_leaf() 📎 tokio/src/runtime/io/registration.rs:160, used for tracing instrumentation.

Step two:coop::poll_proceed(cx) 📎 tokio/src/runtime/io/registration.rs:155-171. This is the cooperative budget mechanism to be discussed in Chapter 12. If the budget is exhausted, returnPendingand register a specialWaker, causing the task to be rescheduled in the next round.

Step three:self.shared.poll_readiness(cx, direction) 📎 tokio/src/runtime/io/registration.rs:155-171. This is where it truly interacts withScheduledIo: check the current readiness bit; if already ready, return immediatelyReady; otherwise storecx.waker()intoScheduledIo's corresponding direction slot and returnPending。

Step four: checkev.is_shutdown 📎 tokio/src/runtime/io/registration.rs:155-171. If the runtime is shutting down, returnRUNTIME_SHUTTING_DOWN_ERROR。

Step five:coop.made_progress() 📎 tokio/src/runtime/io/registration.rs:169, mark the budget as consumed, and return the readiness event.

poll_ioadds a retry loop on top ofpoll_ready📎 tokio/src/runtime/io/registration.rs:173-192:

rust
loop {
    let ev = ready!(self.poll_ready(cx, direction))?;
    match f() {
        Ok(ret) => return Poll::Ready(Ok(ret)),
        Err(ref e) if e.kind() == io::ErrorKind::WouldBlock => {
            self.clear_readiness(ev);
        }
        Err(e) => return Poll::Ready(Err(e)),
    }
}

This reflectsreadiness is a hint, not a guaranteethe core idea of:poll_readysays it is readable, but when actuallyread()it may returnWouldBlock(for example, another thread read the data first). In this case, it mustclear_readiness(ev) 📎 tokio/src/runtime/io/registration.rs:187clear the readiness bit and then loop to wait again. If it is not cleared, the task will fall into a busy loop of "thinks it is readable -> read fails -> thinks it is readable again."

Design consideration:try_ioandasync_iodivision of labor

try_io 📎 tokio/src/runtime/io/registration.rs:194-213is the synchronous version: firstready_event(interest)check the readiness bit; if empty, return directlyWouldBlock 📎 tokio/src/runtime/io/registration.rs:194-213; otherwise executef(), and iff()returnsWouldBlockthen clear the readiness bit📎 tokio/src/runtime/io/registration.rs:207-210. Itdoes not register a Waker, suitable fortry_readscenarios like "try once and leave."

async_io 📎 tokio/src/runtime/io/registration.rs:225-245is the asynchronous version:readiness(interest).awaitregisters a Waker and waits, then when executingf(),WouldBlockclears the readiness bit and loops. Note that inside the loop it also callscoop::poll_proceed 📎 tokio/src/runtime/io/registration.rs:233, preventing exhaustion of the budget during large numbers ofWouldBlockretries.

Production pitfalls:DropWaker cleanup in

Registration::drop 📎 tokio/src/runtime/io/registration.rs:253-262callsself.shared.clear_wakers(). The comment📎 tokio/src/runtime/io/registration.rs:253-262explains the reason:ScheduledIotheWakerstored inArc<driver::Inner>may holddriver::Inner, andScheduledIoin turn holdsRegistration, forming a circular reference. Cleaning up the Waker is a means of breaking the cycle. But the comment also admits this is an "imperfect solution" - ifWakeritself is stored in

, the cycle still exists. This is the issue discussed in tokio-rs/tokio#3481.

[Design inference and architectural trade-offs]clear_wakersThe behavior in production is: if a large number of connections are dropped but the runtime has not exited, memory will not be reclaimed immediately until the nextScheduledIoor runtime shutdown. For long-lived connection services, this is usually not a problem; but for scenarios with high-frequency creation/destruction of short-lived connections, attention must be paid to

the reclamation timing ofTcpStream::read.WakerThe complete chain from

to

wakeupTcpStreamIntuitive model.read().awaitNow connect the three layers. The user callsAsyncRead::poll_read → PollEvented::poll_read → Registration::poll_read_ioonWaker, and what is actually executed isScheduledIo. When data has not arrived,ScheduledIois stored intoWaker; when epoll reports readable, the driver takespoll_readinessfromReady,read()and wakes it, the task is rescheduled, and on the next poll

finds that the readiness bit has been set and directly returns

success.。TcpStream::new 📎 tokio/src/net/tcp/stream.rs:166-169Step-by-Step: a complete read waitPollEvented::new(connected)Phase one: register interestRegistration::new_with_interest_and_handle 📎 tokio/src/runtime/io/registration.rs:73-81callshandle.driver().io().add_source(io, interest) 📎 tokio/src/runtime/io/registration.rs:73-81。

add_source 📎 tokio/src/runtime/io/driver.rs:288-312, which internally calls

1. registrations.allocate(&mut synced.lock()), and thenScheduledIodoes three things:token 📎 tokio/src/runtime/io/driver.rs:293-294。

2. self.registry.register(source, token, interest.to_mio())allocate a📎 tokio/src/runtime/io/driver.rs:298, obtainregisterwith the kernel. If it fails,ScheduledIomust📎 tokio/src/runtime/io/driver.rs:300-303remove the just-allocated

3. metrics.incr_fd_count()from the set📎 tokio/src/runtime/io/driver.rs:309。

, otherwise it leaks.countTcpStream::poll_read 📎 tokio/src/net/tcp/stream.rs:1492-1498 → poll_read_priv 📎 tokio/src/net/tcp/stream.rs:1451-1458 → PollEvented::poll_read → Registration::poll_read_io 📎 tokio/src/runtime/io/registration.rs:133-139 → poll_io → poll_ready → ScheduledIo::poll_readinessPhase two: wait for readinessWaker. The task pollsScheduledIo. At this point, if not ready,Pending。

is stored into's read slot and returnsturnPhase three: event arrivespoll.poll(). The driver's📎 tokio/src/runtime/io/driver.rs:198obtains eventio.set_readiness(Tick::Set, |curr| curr | ready)fromio.wake(ready) 📎 tokio/src/runtime/io/driver.rs:228-229。wake, and while iterating, for each fd event executesWakerandwake()。

internally takes out the。Waker::wake()for the corresponding direction and callspoll_readinessPhase four: task reschedulingReady,read()re-enqueues the task into the worker's local queue (discussed in the previous chapter). The worker polls the task again,

mermaid
sequenceDiagram
    participant Task as "任务 (worker 线程)"
    participant Reg as "Registration"
    participant SIO as "ScheduledIo"
    participant Drv as "Driver (I/O 线程)"
    participant OS as "epoll/kqueue"

    Task->>Reg: "poll_read_ready(cx)"
    Reg->>SIO: "poll_readiness(cx, Read)"
    SIO-->>Reg: "Pending (Waker 已存入读槽位)"
    Reg-->>Task: "Poll::Pending"
    Note over Task: 任务让出,worker 去跑别的任务
    Drv->>OS: "poll.poll(events, max_wait)"
    OS-->>Drv: "event(token=fd_ptr, READABLE)"
    Drv->>SIO: "set_readiness(Tick::Set, curr | READABLE)"
    Drv->>SIO: "wake(READABLE)"
    SIO->>Task: "Waker::wake() 重新入队"
    Note over Task: worker 再次 poll 该任务
    Task->>Reg: "poll_read_ready(cx)"
    Reg->>SIO: "poll_readiness(cx, Read)"
    SIO-->>Reg: "Ready(ReadyEvent{ready: READABLE})"
    Reg-->>Task: "Poll::Ready(Ok(ev))"
    Task->>Task: "read() 成功返回数据"

success.assume_readyCopy

TcpStream::new_accepted 📎 tokio/src/net/tcp/stream.rs:174-181Important branch:acceptoptimizationnew_acceptedis a noteworthy optimization.assume_ready(Ready::READABLE | Ready::WRITABLE) 📎 tokio/src/net/tcp/stream.rs:174-181。

assume_readyThe socket returned by📎 tokio/src/runtime/io/registration.rs:103-105is naturally writable and usually already holds the peer's first batch of bytes. If it waits for the driver's first event, under high load this event may be queued behind the events of all established connections, causing latency. SoWouldBlockdirectly callsWouldBlock,poll_io's commentsays: "A wrong guess costs one, which clears the readiness again." - the cost of guessing wrong is just one

loop will clear the readiness bit and wait again. This is an

optimistic guess + fast correction

design.DriverDesign consideration: why the I/O driver and scheduler are decoupledDriver[Design inference and architectural trade-offs]block_onFrom the source structure,Handleand worker threads are separate:

1. is placed in some dedicated location in the runtime (usually the:Handlethread or a dedicated I/O thread), while worker threads only holdmio::Registry. This decoupling brings several benefits:

2. Lock-free registrationholds a clone ofepoll_wait, and any worker can concurrently register a new fd without going back to the driver thread.

3. Centralized event waiting: only one thread blocks onScheduledIo, avoiding the thundering herd problem of multiple threads polling the same epoll fd simultaneously.Waker::wake(),wake()Short wakeup path

: after the driver receives an event, it directly operates onScheduledIoand callsset_readinessinternally pushes the task into the worker queue, without cross-thread message passing.poll_readinessThe cost is that

needs to handle concurrent access (is_shutdownandRUNTIME_SHUTTING_DOWN_ERROR

poll_readymay occur simultaneously), which is solved through atomic operations and internal locks.ev.is_shutdown 📎 tokio/src/runtime/io/registration.rs:155-171Production pitfalls:gone() 📎 tokio/src/runtime/io/registration.rs:265-267andRUNTIME_SHUTTING_DOWN_ERROR。

check

, and if true returnshutdown 📎 tokio/src/runtime/io/driver.rs:174-182will iterate over all registered and callio.shutdown(), setis_shutdownand wake all waiters. If this flag is not checked, a task may still try to read the socket after the runtime has already stopped scheduling, causing undefined behavior or hangs. In production, if you seeRUNTIME_SHUTTING_DOWN_ERROR, it usually means some task is still running after the runtime drop—check whether there arespawntasks that were not properly joined.

Another pitfall isderegister_source'sunpark 📎 tokio/src/runtime/io/driver.rs:328. If the driver is currently blocked inpoll, and at this point the lastRegistrationis dropped,unparkwill wake the driver. But if the driver is not in a blocked state (for example, it is handling other events),unparkonly makes the nextturnimmediately return📎 tokio/src/runtime/io/driver.rs:280-283. This semantics is explained in the documentation comments ofHandle::unpark.

Design thinking: The three key trade-offs of the Reactor

Trade-off one:Tokenuses pointers instead of indices。EXPOSE_IO.from_exposed_addr(token.0) 📎 tokio/src/runtime/io/driver.rs:220treatsmio::Tokendirectly as the address of*const ScheduledIo. This avoids maintaining aToken → ScheduledIomapping table, and lookup is O(1) and lock-free. The cost is that safety depends on strict lifetime management: the pointer must be released only after deregistration and after the driver no longer polls it📎 tokio/src/runtime/io/driver.rs:222-225。

Trade-off two: separate read/write Waker slots。RegistrationThe documentation📎 tokio/src/runtime/io/registration.rs:24-26says "A registration instance represents two separate readiness streams"—read and write each have an independentWakerslot. This allows the read task and write task of the same socket to register separately without interfering with each other. Butpoll_read_ready's comment📎 tokio/src/net/tcp/stream.rs:549-552reminds us: callingpoll_read_ready/poll_read/poll_peekmultiple times only keeps the lastWaker—there is only one slot for the read direction.

Trade-off three:events_busy's independent buffer. The test📎 tokio/src/runtime/io/driver.rs:364-386verifies this behavior:Driver::new(16, Some(2))creates a driver with busy capacity 2, and after registering 5 readable sources, non-blockingturnonly takes 2 events📎 tokio/src/runtime/io/driver.rs:375-376, leaving the remaining 3 in the kernel queue, and the next blockingturngets📎 tokio/src/runtime/io/driver.rs:379-380. This prevents a non-blocking poll from swallowing all events at once and causing subsequent polls to starve.

Chapter summary

This chapter traced the complete Reactor chain behindTcpStream::read:

  • Driver layer:Driverexclusivelymio::Poll,turnblocks waiting for events, usesEXPOSE_IOto convertTokenback into aScheduledIopointer, callsset_readiness + waketo triggerWaker。Handleprovides a registration entry point that can cross threads,unparkis used to interrupt blocking.
  • Registration layer:RegistrationholdsArc<ScheduledIo>,poll_readychecks the readiness bits or stores them inWaker,poll_iousesWouldBlockretry loop to handle false positives,try_io/async_ioserving synchronous and asynchronous scenarios respectively.
  • State layer:ScheduledIois the state slot for the fd, storing read/write readiness bits and dualWakerslots, and is the only bridge between events and tasks.

Chapter review questions

Q1: If inpoll_iotheWouldBlockbranch'sself.clear_readiness(ev)is deleted, in what scenario would it cause a task busy-loop? Why?

Reference analysis:poll_io's loop📎 tokio/src/runtime/io/registration.rs:173-192is called whenf()returnsWouldBlockclear_readiness(ev) 📎 tokio/src/runtime/io/registration.rs:187。evispoll_readyreturned byReadyEvent, containing the current readiness bits.clear_readinesswill clear these bits fromScheduledIo.

If not cleared, the next time the loop callspoll_ready → poll_readiness,ScheduledIostill retains the old "readable" bit,poll_readinesswill immediately returnReady(because the readiness bits are non-empty), and thenf()executesread()again; if the socket really has no data, it returnsWouldBlockagain, and the loop continues. Since the readiness bits are never cleared, this loop will never enterPending, and the task will keep occupying CPU polling.

Trigger scenarios: multiple tasks share the read direction of the same socket (althoughRegistrationdocumentation📎 tokio/src/runtime/io/registration.rs:28-33says at most two tasks, there is only one slot for the read direction), ortry_readandpoll_readare mixed. More commonly: after epoll reports readable, another thread reads the data first, and the current task'sread()returnsWouldBlock; at this point the readiness bits must be cleared, otherwise it will keep retrying.

Q2: add_sourceInregistry.register, why callregistrations.removewhen it fails? What happens if it is not called?

Reference analysis:add_source 📎 tokio/src/runtime/io/driver.rs:288-312firstregistrations.allocateallocatesScheduledIo 📎 tokio/src/runtime/io/driver.rs:293, thenregistry.registerregisters with the kernel📎 tokio/src/runtime/io/driver.rs:298. If registration fails,ScheduledIohas already been allocated but no fd is associated with it; if it is not removed, it will remain inRegistrationSetforever.

The comment📎 tokio/src/runtime/io/driver.rs:296-297explicitly says: "we should remove thescheduled_io from the registrations set if registering the source with the OS fails. Otherwise it will leak the scheduled_io."—this is a memory leak.

removeThe call to📎 tokio/src/runtime/io/driver.rs:300-303is wrapped in an unsafe block becauseScheduledIois part ofRegistrationSet, and the removal operation needs to ensure there are no other references. Consequences of the leak:RegistrationSetkeeps growing,Tokenspace is wasted, and eventually it may causeallocateto fail or memory exhaustion. In scenarios with high-frequency connection creation/destruction (such as short-connection servers), if the registration failure rate is high (for example, fd exhaustion), the leak will accelerate resource depletion.

Q3: deregister_sourceInunpark(), why isregistrations.deregisteronly called when

returns true? What problems would occur if it were called unconditionally?:deregister_source 📎 tokio/src/runtime/io/driver.rs:315-334Reference analysisregistry.deregister(source)The logic of📎 tokio/src/runtime/io/driver.rs:322is: firstregistrations.deregisterderegisters📎 tokio/src/runtime/io/driver.rs:315-334from the kernel, thenunpark() 📎 tokio/src/runtime/io/driver.rs:328。

registrations.deregistercleans up internal stateScheduledIo, and if it returns true, thenpollreturning true means this is the last reference,unparkis truly removed. At this point the driver may be blocked inmio::Wakerwaiting for events for this fd, but the fd has already been deregistered, and the kernel will no longer generate events.TOKEN_WAKEUPBy📎 tokio/src/runtime/io/driver.rs:280-283pushing apollevent into epoll

, it makesunparkreturn immediately, and the driver rechecks the registration set and may exit blocking.ScheduledIoIfTcpStreamis called unconditionally: every time a non-last reference is deregistered, the driver will be woken, causing unnecessary wakeups. In scenarios where many connections share the samesplitread-write halves), each drop of a half wakes the driver, increasing CPU overhead. More seriously

In this chapter, we dissected how Reactor translates epoll events into Waker wakeups: starting from TcpStream's poll_read_ready, going through Registration's registration and lookup, landing on ScheduledIo's readiness bits and Waker slots, and then the Driver locating and triggering wakeups by Token in the event loop. Key designs include: Token-as-pointer for O(1) lookup, read/write dual Waker slots supporting separated concurrent read/write, events_busy independent buffer preventing event starvation, and assume_ready optimistic guessing optimizing the accept scenario. At this point, the closed loop of I/O readiness notification is complete. But an async runtime also needs to handle another kind of "readiness" — time. In the next chapter, we will analyze the implementation of tokio::time::sleep and timeout: how timers are inserted into the timing wheel, how the timing wheel is leveled by expiration time, and how the driver calculates the timeout for the next park and triggers expired tasks. You will see the unified abstraction that "time is also an I/O event," and how start_paused and the test clock make time controllable in tests.

CHAPTER 06

Chapter 6: Time-Driven: How the Timing Wheel, Sleep, and Timeouts Are Woken

Project: tokio-rs/tokio · Book progress: Chapter 6 / 14 · Verification status: FACT line numbers truly anchored

In the previous chapter, we traced the complete path of TcpStream::read and saw how ScheduledIo translates epoll fd readiness events into Waker wakeups. But an async runtime also needs to handle another kind of "readiness": a sleep(100ms) Future must be woken after 100ms. This kind of event does not come from a kernel fd, but from "time itself." Tokio's design choice is to treat time as a kind of I/O event as well: the Driver struct has only one field, park: IoStack, which reuses the I/O driver's park/unpark mechanism. When the timing wheel calculates the "next expiration instant," the driver calls park_timeout to let the thread sleep until that instant; after being woken, it takes expired entries out of the timing wheel and triggers their Wakers. In this way, the scheduler only needs a unified park entry point to wait for both kinds of events: "fd readiness" and "timer expiration." This chapter answers three questions: How are timers inserted into the timing wheel? How is the timing wheel leveled by expiration time? How does the driver calculate the timeout for the next park and trigger expired tasks?

1. Timing Wheel: A Six-Level, 64-Slot Hashed Hierarchical Structure

Intuitive model

Imagine a mechanical clock: the second hand drives the minute hand through one revolution, and the minute hand drives the hour hand through one revolution. If there were only a second hand, representing "12 days later" would require counting a million ticks; but after layering, the second hand only handles precision within 64 seconds, the minute hand handles 64 minutes, and the hour hand handles 64 hours — each level only needs 64 slots to cover more than 2 years into the future.

Without layering, inserting a far-future timer would require either O(N) traversal or a huge array. The timing wheel uses "leveling by expiration time" to reduce both insertion and triggering to approximately O(1).

Memory layout and fields

WheelThe core fields of are only three📎 tokio/src/runtime/time/wheel/mod.rs:22-40:

rust
pub(crate) struct Wheel {
    elapsed: u64,                          // 自 wheel 创建以来经过的毫秒数
    levels: Box<[Level; NUM_LEVELS]>,      // 6 层,每层 64 槽
    pending: LinkedList<TimerShared>,      // 已到期、待触发的条目
}

NUM_LEVELS = 6,BITS_PER_LEVEL = 6(that is, 64 slots per level)📎 tokio/src/runtime/time/wheel/mod.rs:45-47。MAX_DURATION = 1 << (6 * 6) = 1 << 36milliseconds, about 2 years📎 tokio/src/runtime/time/wheel/mod.rs:50。

The granularity of the six levels, according to the documentation comments, is📎 tokio/src/runtime/time/wheel/mod.rs:22-40:

LevelSlot granularityCoverage range
01 ms64 ms
164 ms~4 s
2~4 s~4 min
3~4 min~4 hr
4~4 hr~12 day
5~12 day~2 yr

pendingis an intrusive linked list (LinkedList<TimerShared>), storing entries that have already been taken out of the wheel and are waiting to trigger their Wakers. Note that it isLinkedListrather thanVec: the entries themselves are embedded inTimerShared, so insertion/removal does not require allocation.

Scenario-driven: inserting a 100ms sleep

Whensleep(100ms)is first polled,Sleep::poll_elapsedwill constructTimer::newand callinit 📎 tokio/src/time/sleep.rs:436-440。initultimately callsHandle::reregister, and then callsWheel::insert。

insertThe first step is to check whether it has already expired📎 tokio/src/runtime/time/wheel/mod.rs:90-98:

rust
let when = unsafe { item.sync_when() };
if when <= self.elapsed {
    return Err((item, InsertError::Elapsed));
}

Ifwhenhas already fallen beforeelapsed(for example, the deadline has passed), return directlyElapsed, and the caller will immediately trigger the timer.

Otherwise, calculate which level the entry should be placed into📎 tokio/src/runtime/time/wheel/mod.rs:90-114:

rust
let level = self.level_for(when);
unsafe { self.levels[level].add_entry(item); }

level_foris the core of the leveling algorithm📎 tokio/src/runtime/time/wheel/mod.rs:276-289:

rust
fn level_for(elapsed: u64, when: u64) -> usize {
    const SLOT_MASK: u64 = (1 << BITS_PER_LEVEL) - 1;
    let masked = elapsed ^ when | SLOT_MASK;
    if masked >= MAX_DURATION {
        return NUM_LEVELS - 1;
    }
    masked.ilog2() as usize / BITS_PER_LEVEL
}

Hereelapsed ^ whenis used rather thanwhen - elapsed, which is an ingenious trick: the most significant bit of the XOR reflects "from which bit the two timestamps first differ," that is, "how coarse a granularity is needed to distinguish them."| SLOT_MASKforces the low 6 bits to 1, avoidingilog2calculating too small a level when they fall into the same slot.ilog2() / 6maps the bit width to the level number. If the XOR result exceedsMAX_DURATION(that is, more than 2 years), it is forcibly placed into the highest level — this is "fudge the timer into the top level."

For a 100ms sleep, assumingelapsedis close to 0,when ≈ 100,elapsed ^ when ≈ 100,ilog2(100) = 6,6 / 6 = 1, so it falls on level 1 (64ms granularity). This means it will wait in a slot on level 1 until time advances to that slot's boundary before being cascaded down to level 0.

Hierarchical cascading: process_expiration

Whenpoll(now)advances time,Wheel::pollwill repeatedly callnext_expirationandprocess_expiration 📎 tokio/src/runtime/time/wheel/mod.rs:142-166:

rust
pub(crate) fn poll(&mut self, now: u64) -> Option<TimerHandle> {
    loop {
        if let Some(handle) = self.pending.pop_back() {
            return Some(handle);
        }
        match self.next_expiration() {
            Some(ref expiration) if expiration.deadline <= now => {
                self.process_expiration(expiration);
                self.set_elapsed(expiration.deadline);
            }
            _ => {
                self.set_elapsed(now);
                break;
            }
        }
    }
    self.pending.pop_back()
}

process_expirationis responsible for "cascading" expired entries from one level down to the next, or (at level 0) marking them as pending📎 tokio/src/runtime/time/wheel/mod.rs:218-251:

rust
let mut entries = self.take_entries(expiration);
while let Some(item) = entries.pop_back() {
    match unsafe { item.mark_pending(expiration.deadline) } {
        Ok(()) => {
            self.pending.push_front(item);   // 真正到期
        }
        Err(expiration_tick) => {
            let level = level_for(expiration.deadline, expiration_tick);
            unsafe { self.levels[level].add_entry(item); }  // 下沉到更低层
        }
    }
}

mark_pendingis the key: it checks whether an entry's actual deadline has been reached. If reached, it returnsOk(()), and the entry enters thependinglinked list; if not yet reached (only the slot's boundary has been reached), it returnsErr(expiration_tick), and the entry is reinserted into a finer-grained level.

Note the point emphasized in the comments📎 tokio/src/runtime/time/wheel/mod.rs:219-228: all entries in the entire slot must be taken out first before processing, because some entries may be reinserted into the same slot (this happens when the insertion time exceedsMAX_DURATION, causing wraparound). If you take and insert simultaneously, you may fall into an infinite loop.

Calculation of the next expiration time

next_expirationscans from low level to high level, returning the first non-empty expiration point📎 tokio/src/runtime/time/wheel/mod.rs:169-191:

rust
fn next_expiration(&self) -> Option<Expiration> {
    if !self.pending.is_empty() {
        return Some(Expiration { level: 0, slot: 0, deadline: self.elapsed });
    }
    for (level_num, level) in self.levels.iter().enumerate() {
        if let Some(expiration) = level.next_expiration(self.elapsed) {
            debug_assert!(self.no_expirations_before(level_num + 1, expiration.deadline));
            return Some(expiration);
        }
    }
    None
}

Ifpendingis non-empty, it means there are expired entries waiting to be triggered, so it immediately returns the currentelapsedas the deadline (so the driver will park with a 0 timeout and come back immediately to process). Otherwise, it scans level by level, returning the deadline of the first slot with content.debug_assertverifies an invariant: a higher level cannot have an earlier expiration point than the current level.

mermaid
flowchart TD
    start["Wheel::poll(now)"] --> check_pending{"pending 非空?"}
    check_pending -->|是| pop["pop_back 返回 TimerHandle"]
    check_pending -->|否| next_exp{"next_expiration() 有到期点?"}
    next_exp -->|无| set_elapsed["set_elapsed(now) 后 break"]
    next_exp -->|有| cmp{"expiration.deadline <= now?"}
    cmp -->|否| set_elapsed
    cmp -->|是| proc["process_expiration(expiration)"]
    proc --> take["take_entries 取出整槽"]
    take --> mark{"item.mark_pending()"}
    mark -->|Ok 已到期| push_pending["pending.push_front(item)"]
    mark -->|Err 未到期| reinsert["level_for 后 add_entry 下沉"]
    push_pending --> set_elapsed2["set_elapsed(expiration.deadline)"]
    reinsert --> set_elapsed2
    set_elapsed2 --> check_pending
    set_elapsed --> pop2["pending.pop_back() 返回"]

---

II. The Driver's park loop: connecting the timer wheel to the I/O stack

Intuitive model

The timer wheel itself does not "run on its own." It needs an external loop to repeatedly ask it: "When is the next expiration?" Then it sleeps until that moment, and after waking up, advances time. This loop isDriver::park_internal. It translates "the timer wheel's next expiration" into apark_timeoutduration, handing it to the underlying I/O stack to sleep.

Without this loop, timers would never be triggered—the timer wheel is just a static data structure that needs someone to "turn" it.

Data structures: Driver and InnerState

Driverhas only one fieldpark: IoStack 📎 tokio/src/runtime/time/mod.rs:90-93. The real state is inHandle, distinguished via theInnerenum between the traditional implementation and the experimental implementation📎 tokio/src/runtime/time/mod.rs:95-127. The traditional implementation'sInnerStatecontains two fields📎 tokio/src/runtime/time/mod.rs:130-136:

rust
struct InnerState {
    next_wake: Option<NonZeroU64>,   // 承诺的最早唤醒时刻
    wheel: wheel::Wheel,
}

next_wakeusesNonZeroU64instead ofOption<u64>nesting, in order to leverage niche optimization—Option<NonZeroU64>andu64are the same size. It records "before which tick the driver promises to wake up," used duringreregisterto determine whetherunpark。

is_shutdownis neededAtomicBoolis an independent📎 tokio/src/runtime/time/mod.rs:90-93:Handle, and the comments explain why it was split out from the Mutexis_shutdownneeds to check

without locking the mutex. This is a typical "read-many, write-few" optimization—shutdown happens only once, but checks may be frequent.

park_internalScenario-driven: the complete flow of one park📎 tokio/src/runtime/time/mod.rs:213-256:

rust
fn park_internal(&mut self, rt_handle: &driver::Handle, limit: Option<Duration>) {
    let handle = rt_handle.time();
    let mut lock = handle.inner.lock();
    assert!(!handle.is_shutdown());

    let next_wake = lock.wheel.next_expiration_time();
    lock.next_wake = next_wake.map(|t| NonZeroU64::new(t).unwrap_or_else(|| NonZeroU64::new(1).unwrap()));
    drop(lock);

    match next_wake {
        Some(when) => {
            let now = handle.time_source.now(rt_handle.clock());
            let mut duration = handle.time_source.tick_to_duration(when.saturating_sub(now));
            if duration > Duration::from_millis(0) {
                if let Some(limit) = limit {
                    duration = std::cmp::min(limit, duration);
                }
                self.park_thread_timeout(rt_handle, duration);
            } else {
                self.park.park_timeout(rt_handle, Duration::from_secs(0));
            }
        }
        None => {
            if let Some(duration) = limit {
                self.park_thread_timeout(rt_handle, duration);
            } else {
                self.park.park(rt_handle);
            }
        }
    }

    handle.process(rt_handle.clock());
}

Copy

1. Step-by-step analysis::lock.wheel.next_expiration_time()Acquire lock, read next expirationOption<u64>returnslock.next_wake, i.e., the next expiration tick. At the same time, it writes it toreregister, for

2. to determine whether unpark is needed.:drop(lock)Release lock

3. must be before park, otherwise other threads cannot insert timers during park.:when.saturating_sub(now)Calculate park durationtick_to_durationobtains the remaining tick count,Durationconverts it to📎 tokio/src/runtime/time/mod.rs:228-230. The comments point out that this is actually rounded up to 1ms

4. , to avoid microsecond-level sleep being treated as zero-length by the OS.Handle limitlimit: if the caller passedpark_timeout(such asmin(limit, duration)'s explicit timeout), take

5. , ensuring it will not oversleep.Special caseduration == 0: ifpark_timeout(0)(already expired), use

6. to return immediately without actually sleeping.When there are no timersnext_wake: ifNoneislimit, if there ispark_thread_timeout(limit)thenpark。

7. , otherwise infinite:handle.process(clock)Process after waking up

advances the timer wheel and triggers expired entries.

processprocess_at_time: triggering expired entriesprocess_at_time 📎 tokio/src/runtime/time/mod.rs:296-337:

rust
pub(self) fn process_at_time(&self, mut now: u64) {
    let mut waker_list = WakeList::new();
    let mut lock = self.inner.lock();

    if now < lock.wheel.elapsed() {
        // 时间倒流保护
        now = lock.wheel.elapsed();
    }

    while let Some(entry) = lock.wheel.poll(now) {
        debug_assert!(unsafe { entry.is_pending() });
        if let Some(waker) = unsafe { entry.fire(Ok(())) } {
            waker_list.push(waker);
            if !waker_list.can_push() {
                drop(lock);
                waker_list.wake_all();
                lock = self.inner.lock();
            }
        }
    }

    lock.next_wake = lock.wheel.poll_at()
        .map(|t| NonZeroU64::new(t).unwrap_or_else(|| NonZeroU64::new(1).unwrap()));
    drop(lock);
    waker_list.wake_all();
}

Copy

  • Several key points: 📎 tokio/src/runtime/time/mod.rs:301-309Time reversal protectionnow < wheel.elapsed(): ifInstant, it means the clock has gone backwards. The comments point out this usually should not happen (Rust guaranteesnowmonotonicity), but it does happen in Linux VMs on Windows hosts, because std incorrectly trusts the hardware clock's monotonicity. The protection method is to clampelapsed。
  • to:WakeListBatch wakeup!can_push()collects Wakers, and when it is full (📎 tokio/src/runtime/time/mod.rs:319), it temporarily releases the lock, wakes up a batch, then reacquires the lock. The comments emphasize this is to avoid deadlock
  • . If a Waker is called while holding the lock, and the Waker tries to operate on the timer wheel (e.g., re-register a timer), it will deadlock.Update next_wakepoll_at(): after processing, recalculatenext_wake。

, update

reregister: re-registration and unparkSleep::resetWhenreregisteris called, the timer needs to be re-registered.📎 tokio/src/runtime/time/mod.rs:398-450:

rust
pub(self) unsafe fn reregister(&self, unpark: &IoHandle, new_tick: u64, entry: NonNull<TimerShared>) {
    let waker = unsafe {
        let mut lock = self.inner.lock();
        if unsafe { entry.as_ref().might_be_registered() } {
            lock.wheel.remove(entry);
        }
        let entry = entry.as_ref().handle();
        if self.is_shutdown() {
            unsafe { entry.fire(Err(crate::time::error::Error::shutdown())) }
        } else {
            entry.set_expiration(new_tick);
            match unsafe { lock.wheel.insert(entry) } {
                Ok(when) => {
                    if lock.next_wake.is_none_or(|next_wake| when < next_wake.get()) {
                        unpark.unpark();
                    }
                    None
                }
                Err((entry, crate::time::error::InsertError::Elapsed)) => unsafe {
                    entry.fire(Ok(()))
                },
            }
        }
    };
    if let Some(waker) = waker {
        waker.wake();
    }
}

Copynext_wakeKey logic: after successful insertion, if the new expiration time is earlier thanunpark.unpark(), call

to wake up the driver. This is because the driver may be sleeping until a later time and needs to be woken up early to recalculate the park duration.unparkNote thatis calledwhile holding the lock, whereaswaker.wake()is calledafter releasing the lock. The comments explain: the lock must be released before calling the Waker to avoid deadlock. But📎 tokio/src/runtime/time/mod.rs:441is different—it merely pushes an event into epoll and will not call back into user code, so calling it while holding the lock is safe.unparkCopy

mermaid
sequenceDiagram
    participant Sleep as Sleep::poll
    participant Handle as time::Handle
    participant Wheel as Wheel
    participant Driver as Driver::park_internal
    participant IoStack as IoStack

    Sleep->>Handle: reregister(unpark, new_tick, entry)
    Handle->>Handle: lock.inner.lock()
    Handle->>Wheel: wheel.remove(entry) [若已注册]
    Handle->>Wheel: wheel.insert(entry)
    Wheel-->>Handle: Ok(when)
    alt when < next_wake
        Handle->>IoStack: unpark.unpark()
    end
    Handle->>Handle: drop(lock)
    Handle-->>Sleep: 返回 waker (若有)

    Note over Driver: 另一线程
    Driver->>Handle: lock.inner.lock()
    Driver->>Wheel: next_expiration_time()
    Wheel-->>Driver: Some(when)
    Driver->>Driver: drop(lock)
    Driver->>IoStack: park_timeout(duration)
    IoStack-->>Driver: 被 unpark 或超时
    Driver->>Handle: process(clock)
    Handle->>Wheel: poll(now)
    Wheel-->>Handle: TimerHandle
    Handle->>Sleep: waker.wake()

---

Intuitive model

is the Future that users directly

Sleep,.await 的 Future,Timeoutis an adapter that wraps another Future. They themselves do not manage the timer wheel; they simply translate the "deadline" into a tick and delegate toTimerandHandle。

Sleep memory layout

Sleepuses thepin_project!macro to define📎 tokio/src/time/sleep.rs:221-227:

rust
pub struct Sleep {
    deadline: Instant,
    driver: scheduler::Handle,
    inner: Inner,
    #[pin]
    timer: Option<Timer>,
}

timerisOption<Timer>and carries#[pin]: before the first poll it isNone, and only on the first poll isTimercreated and registered. This "lazy initialization" avoids accessing the runtime whensleep()is called—sleep()can be called outside the runtime, as long as it is only actually registered at.await.

PinnedDropThe implementation ensures that the timer is canceled on drop📎 tokio/src/time/sleep.rs:230-235:

rust
impl PinnedDrop for Sleep {
    fn drop(this: Pin<&mut Self>) {
        let this = this.project();
        if let Some(timer) = this.timer.as_pin_mut() {
            timer.cancel(this.driver);
        }
    }
}

The complete flow of poll_elapsed

poll_elapsedisSleepthe core of📎 tokio/src/time/sleep.rs:396-454:

rust
fn poll_elapsed(self: Pin<&mut Self>, cx: &mut task::Context<'_>) -> Poll<Result<(), Error>> {
    ready!(crate::trace::trace_leaf());
    let mut this = self.project();

    // coop 预算
    let coop = ready!(crate::task::coop::poll_proceed(cx));

    let handle = this.driver;
    let timer = match this.timer.as_mut().as_pin_mut() {
        Some(timer) => timer,
        None => {
            let time_source = handle.driver().time().time_source();
            let deadline = time_source.deadline_to_tick(*this.deadline);
            let timer = Timer::new(handle, deadline);
            this.timer.set(Some(timer));
            let mut timer = this.timer.as_pin_mut().unwrap();
            timer.as_mut().init(handle, deadline);
            timer
        }
    };

    let result = timer.poll_elapsed(cx, handle).map(move |r| {
        coop.made_progress();
        r
    });
    result
}

Step by step:

1. coop budget check:poll_proceed(cx)consumes one unit of cooperative budget. If the budget is exhausted, returnPendingand yield execution. This is Tokio's mechanism for preventing a single task from starving other tasks.

2. Lazily create Timer: iftimerisNone, convertdeadlineinto a tick, createTimerand callinitto register it with the timer wheel.

3. Delegate to Timer::poll_elapsed: the actual expiration check is performed byTimer.

4. Mark progress on success:coop.made_progress()indicates that this poll made actual progress.

Timeout's poll: poll the value first, then poll the delay

TimeoutThe poll order of📎 tokio/src/time/timeout.rs:210-224:

rust
fn poll(self: Pin<&mut Self>, cx: &mut task::Context<'_>) -> Poll<Self::Output> {
    let me = self.project();
    let had_budget_before = coop::has_budget_remaining();

    // 先 poll 被包裹的 future
    if let Poll::Ready(v) = me.value.poll(cx) {
        return Poll::Ready(Ok(v));
    }

    match me.delay.as_pin_mut() {
        Some(delay) => poll_delay(had_budget_before, delay, cx).map(Err),
        None => Poll::Pending,
    }
}

Copy📎 tokio/src/time/timeout.rs:24-26The comment explicitly statesOk: the future is polled first, and only then is the timeout checked. So if the future completes without yielding, it may still return

poll_delayafter the timeout has passed. This is a design choice, not a bug.📎 tokio/src/time/timeout.rs:229-251:

rust
fn poll_delay(had_budget_before: bool, delay: Pin<&mut Sleep>, cx: &mut task::Context<'_>) -> Poll<Elapsed> {
    let delay_poll = || match delay.poll(cx) {
        Poll::Ready(()) => Poll::Ready(Elapsed::new()),
        Poll::Pending => Poll::Pending,
    };

    let has_budget_now = coop::has_budget_remaining();

    if let (true, false) = (had_budget_before, has_budget_now) {
        // 如果预算是被底层 future 耗尽的,用无约束预算 poll delay
        coop::with_unconstrained(delay_poll)
    } else {
        delay_poll()
    }
}

CopypollLogic: if there is still budget when enteringPending, but the budget is exhausted after polling value, that means value consumed the budget. At this point, if delay is polled with a constrained budget, delay may immediately returnwith_unconstrained, making it impossible to ever determine whether the timeout has been reached. So📎 tokio/src/time/timeout.rs:243-246。

is used to temporarily lift the budget restriction. The comment calls this "pathological cases"

timeouttimeout's deadline overflow handlingchecked_addThe function uses📎 tokio/src/time/timeout.rs:86-99:

rust
Timeout {
    value: future.into_future(),
    delay: match Instant::now().checked_add(duration) {
        Some(deadline) => Some(Sleep::new_timeout(deadline, trace::caller_location())),
        None => None,
    },
}

CopyInstant::now() + durationIfdelayoverflows (the duration is extremely large),NonebecomesPoll::Pending 📎 tokio/src/time/timeout.rs:222, and poll directly returns

---

. This is equivalent to "never time out," which is reasonable degradation behavior.

Design reflections and production pitfalls elapsed ^ whenWhy use XOR instead of subtraction to calculate the level?when - elapsedThe most significant bit ofelapseddirectly reflects "from which bit two timestamps first differ," which is exactly the measure of "how coarse a granularity is needed." Subtractionwhenwhenilog2is close to

has all high bits as 0, 📎 tokio/src/runtime/time/mod.rs:301-309will calculate too small a level. XOR naturally handles wraparound scenarios.InstantThe necessity of time-going-backward protectionInstant: Rust guaranteesnow = lock.wheel.elapsed()monotonicity, but the underlying OS may not. In a Linux VM on a Windows host, std trusts the hardware clock, causingset_elapsedto go backward. Tokio uses

to clamp it, avoiding 📎 tokio/src/runtime/time/mod.rs:319's assert failure.Sleep::resetBatch wakeups and deadlocksWakeList: calling Waker while holding the timer wheel lock is dangerous—the Waker may trigger the task to be polled again, which then calls

next_wake, trying to acquire the timer wheel lock again, causing a deadlock. 📎 tokio/src/runtime/time/mod.rs:130-136:Option<NonZeroU64>'s batching mechanism temporarily releases the lock when the lock is full, which is the standard "callback outside the lock" pattern.u64's niche optimizationNoneandNonZeroU64::new(t).unwrap_or_else(|| NonZeroU64::new(1).unwrap())are the same size, because 0 is used as the niche for📎 tokio/src/runtime/time/mod.rs:221. But tick 0 is a legal value, so the code uses

process_expirationto map 0 to 1 📎 tokio/src/runtime/time/wheel/mod.rs:219-228. This is a subtle boundary handling: tick 0 is treated as tick 1, causing at most 1ms of extra wakeup.MAX_DURATION's "take first, then process"

Timeout: the entire slot's entries must be taken out before processing, because entries exceeding 📎 tokio/src/time/timeout.rs:24-26will wrap around and be reinserted into the same slot. If you take and insert at the same time, it will loop infinitely.Ok's poll order traptimeout: the future is polled first, and the timeout is checked afterward. If the future is CPU-intensive and does not yield, it may still return

---

after the timeout. In production, do not rely on

to forcibly interrupt an uncooperative future.

1. Chapter summary(WheelThis chapter dismantled Tokio's three-layer time-driven structure:elapsed ^ whenTimer wheelpending): a six-level, 64-slot hierarchical hash structure, using the bit width ofprocess_expirationto determine the entry level, with insertion and triggering approximately O(1).

2. Driver(Driver::park_internalThe linked list stores expired entries,next_expiration_timeis responsible for cascading them down level by level.park_timeout): translates the timer wheel'sprocess_at_timeinto

3. duration, reusing the I/O stack's park/unpark.(Sleep / Timeout):Sleepadvances the timer wheel after wakeup, triggers Wakers in batches, and handles time-going-backward and deadlock protection.TimerUser APITimeoutlazily createswith_unconstrainedand registers it,

polls value first and then delay, usingnext_waketo handle the budget-exhaustion scenario.reregisterThe core design is that "time is also an I/O event": the driver has only one park entry point, waiting simultaneously for fd readiness and timer expiration.unparkrecords the promised wakeup time,

and when an earlier timer is inserted,Mutex、Semaphorewakes the driver to recalculate.

In the next chapter we will enter synchronization primitives:

and how channels implement asynchronous waiting. You will see how they reuse this chapter's Waker mechanism, and how "permit counting" and "wait queues" cooperate.Wheel::insertChapter reflection and self-testif when <= self.elapsedQ1: If inif when < self.elapsed(remove the equals sign), in what scenario would the timer never be triggered?

Reference analysis:when == self.elapsedmeans the timer's expiration time is exactly equal to the currently advanced time. The original code uses<=to judge it asElapsed, and the caller immediately triggers📎 tokio/src/runtime/time/wheel/mod.rs:96-98. If changed to<, this entry will be inserted into the level calculated bylevel_for(elapsed, when). Sinceelapsed ^ when == 0,masked = 0 | SLOT_MASK = 63,ilog2(63) = 5,5 / 6 = 0, it falls into level 0. But level 0'snext_expirationwill return a slot ofdeadline >= elapsed, andWheel::poll's condition isexpiration.deadline <= now. Ifnow == elapsed, the condition holds,process_expirationwill take out the entry,mark_pending(elapsed)checks whether the actual deadline has been reached—at this pointwhen == elapsed,mark_pendingreturnsOk, and the entry enters pending. So in fact it will still be triggered, but with an extra detour. The real risk is: ifelapsedhas already advanced pastwhen(when < elapsed), the original code returnsElapsedand triggers immediately, while after the change it is inserted into a slot that has already passed,next_expirationmay returndeadline < elapsed,set_elapsed's assertelapsed <= whenwill fail and panic📎 tokio/src/runtime/time/wheel/mod.rs:253-264. So this equals sign is the key boundary that prevents the assert from failing.

Q2: process_at_timeInWakeList, afterdrop(lock)is full, whywake_all()thenlockand then re-

? If this drop is removed, in what concurrency scenario would it deadlock?:WakeListReference analysis📎 tokio/src/runtime/time/mod.rs:318-325collects Wakers, and once full it must wake a batch to free up spaceself.inner.lock(). Ifwaker.wake()is called while holdingSleep::reset, the awakened task may immediately run on another thread (or the same thread's scheduler), callingSleep::poll_elapsedorHandle::reregister, and then callingreregister, while the first thingself.inner.lock() 📎 tokio/src/runtime/time/mod.rs:405does isstd::sync::Mutex. Sinceprocess_at_timeis not reentrant, the same thread will deadlock; even on a different thread, it will block untilprocess_at_timereleases the lock, whilewake_allis waiting for📎 tokio/src/runtime/time/mod.rs:319to return, forming a circular wait. The comment explicitly says "To avoid deadlock, we must do this with the lock temporarily dropped"while let Some(entry) = lock.wheel.poll(now). When re-locking after the drop, the timer wheel state may have been modified by other threads (such as a new timer being inserted), so

Q3: Timeout::pollwill continue to take entries from the new state, which is safe.had_budget_beforeInhas_budget_now, the combined judgment of(true, false)andwith_unconstrainedwhy is it only used when "there is budget on entry, and no budget after polling value"(false, true)? What if it were reversed

?:had_budget_beforeReference analysis📎 tokio/src/time/timeout.rs:208-208,has_budget_nowrecords📎 tokio/src/time/timeout.rs:239。(true, false)before polling value, and recordspoll_proceedafter polling value.Pendingmeans the budget was exhausted during polling value, indicating that value is a "budget consumer." At this point, if delay is polled with a restricted budget,with_unconstrainedwill immediately return📎 tokio/src/time/timeout.rs:247。(false, true), delay will never actually be checked, and timeout judgment becomes ineffective. So usingwith_unconstrainedto temporarily lift the restriction(false, false)cannot happen—the budget can only be consumed, not restored (unless explicitlyPending, but that is not the case here).poll_proceedmeans there was no budget on entry, at which point polling value may already have returned(true, true)(because

failed), and delay is also polled with a restricted budget, both pending, as expected.

CHAPTER 07

Back to top ↑

Next chapter: Chapter 7 → · Chapter 7: Synchronization Primitives: How Mutex, Semaphore, and Channels Implement Asynchronous Waiting · Project: tokio-rs/tokio

Book progress: Chapter 7 / 14

Verification status: FACT line numbers truly anchored

The previous chapter revealed how time is abstracted as a kind of I/O event, allowing timers and fd readiness to share the same park/unpark waiting entry point. However, when multiple tasks compete for the same lock or pass messages through channels, the object being waited on is no longer an fd or a clock, but another task's state change. This chapter enters the tokio::sync family to find out where a lock().await or recv().await actually stores the Waker when blocking, and how it is rescheduled when awakened.

std::sync::MutexWhy asynchronous Mutex cannot reuse std's implementationlock()Intuitive model: from "occupying the seat" to "yielding the seat"'swhen the lock is occupied willblock the current thread—the thread is suspended by the operating system until the lock is released. This is disastrous in an async runtime: a worker thread may drive hundreds or thousands of tasks at the same time, and if it blocks waiting for a lock, all the other tasks it carries come to a halt. The core requirement of an async Mutex is: when waiting for the lock,Pendingyield the thread

, register the fact that "I am waiting for this lock" into a queue, and then returnMutex, letting the executor run other tasks.Built entirely on top of a semaphore。

Data structures and memory layout

Mutex<T>The fields of are extremely minimal:

📎 tokio/src/sync/mutex.rs:133-138

rust
pub struct Mutex<T: ?Sized> {
    #[cfg(all(tokio_unstable, feature = "tracing"))]
    resource_span: tracing::Span,
    s: semaphore::Semaphore,
    c: UnsafeCell<T>,
}

The three fields each serve a distinct purpose:sis asemaphore with a permit count of 1,cisUnsafeCell<T>the protected data wrapped by . Note that heresemaphoreis an alias forbatch_semaphore📎 tokio/src/sync/mutex.rs:3-3, i.e., the underlying implementation, not thesync::Semaphorepublic wrapper layer.

MutexGuard<'a, T>only holds a reference toMutex:

📎 tokio/src/sync/mutex.rs:151-157

rust
pub struct MutexGuard<'a, T: ?Sized> {
    #[cfg(all(tokio_unstable, feature = "tracing"))]
    resource_span: tracing::Span,
    lock: &'a Mutex<T>,
}

There is a key design decision here:MutexGuard does not hold a semaphore permit object, only holds&Mutex. The action of releasing the lock happens inDrop, directly callingself.lock.s.release(1) 📎 tokio/src/sync/mutex.rs:959-961. This differs fromSemaphorePermitwhich holds apermits: usizecount and returns it on Drop—Mutex's permit count is always 1, so no counting is needed.

Send/SyncThe bounds of are worth examining separately:

📎 tokio/src/sync/mutex.rs:258-259

rust
unsafe impl<T> Send for Mutex<T> where T: ?Sized + Send {}
unsafe impl<T> Sync for Mutex<T> where T: ?Sized + Send {}

Synconly requiresT: Sendrather thanT: Sync—this is reasonable, because mutual exclusion guarantees that only one thread can touchTat a time. Transferring ownership ofTacross threads (Send) is sufficient; there is no need forTitself to be shareable (Sync). This is exactly whyMutex<T>can turn a non-SyncTintoSync.

Step-by-Step: A complete journey oflock().await

Scenario: Task A callsmutex.lock().await, and the lock is currently free.

Step one,lock()constructs an async block that firstself.acquire().await, and upon success constructsMutexGuard 📎 tokio/src/sync/mutex.rs:434-443。

Step two,acquire()directly delegates to the semaphore:

📎 tokio/src/sync/mutex.rs:655-663

rust
async fn acquire(&self) {
    crate::trace::async_trace_leaf().await;
    self.s.acquire(1).await.unwrap_or_else(|_| {
        unreachable!()
    });
}

unwrap_or_else(|_| unreachable!())This comment reveals the design constraint: Mutex never explicitly closes the semaphore and holds it exclusively, soacquirewill never returnErr. This eliminates the "semaphore closed" error path at the type level.

Step three, if the lock is occupied,s.acquire(1)returnsPending, and the current task's Waker is registered into the semaphore's wait queue.Where is the Waker stored?The answer lies inbatch_semaphore's wait queue (the source material for this chapter does not expand on that file, but its role is: each waiter holds a Waker, queued in FIFO order).

Step four, when task B, which holds the lock, releases it,MutexGuard::dropcallss.release(1) 📎 tokio/src/sync/mutex.rs:965-975, the semaphore hands the permit to the head waiter and wakes its Waker, task A is rescheduled,acquirereturnsOk, constructingMutexGuard。

The entire flow can be depicted with the following sequence diagram:

mermaid
sequenceDiagram
    participant TaskA as 任务 A
    participant Mutex as Mutex.s (batch_semaphore)
    participant TaskB as 任务 B (持锁者)
    participant Exec as Executor

    TaskA->>Mutex: acquire(1).await
    Mutex-->>TaskA: Pending (Waker 入队)
    TaskA->>Exec: 让出,调度其他任务
    Note over TaskB: 持有锁执行临界区
    TaskB->>Mutex: MutexGuard::drop -> release(1)
    Mutex->>TaskA: 唤醒队首 Waker
    Exec->>TaskA: 重新 poll
    TaskA->>Mutex: acquire(1) 重试
    Mutex-->>TaskA: Ok(()) 获得许可
    TaskA->>TaskA: 构造 MutexGuard

Design considerations: FIFO fairness and cancellation safety

The documentation explicitly states that Tokio's Mutex guarantees FIFO📎 tokio/src/sync/mutex.rs:20-22. This fairness comes from the underlying semaphore's queuing semantics. The cost of fairness is: alockbeing cancelled (e.g., losing inselect!) will cause you tolose your position in the queue 📎 tokio/src/sync/mutex.rs:415-419. This is not a bug, but an inevitability of FIFO queues—cancellation means removal from the queue, and re-lockrequires re-queuing.

Another counterintuitive design is thatdoes not poison(no poisoning)。std::sync::Mutexmarks itself as poisoned when the lock-holding thread panics, and subsequentlockreturnsErr. Tokio's Mutex does not do this: when the lock holder panics, the lock is released normally📎 tokio/src/sync/mutex.rs:122-125. The documentation warns that if the panic is caught, the protected data may be in an inconsistent state. This is a pragmatic trade-off in async scenarios—a panic in an async task usually means task termination, and a poisoning mechanism would only add complexity.

MutexGuard::mapThe series of methods is worth mentioning. It allows downgrading an entireMutexGuard<T>to aMappedMutexGuard<U>that only protects a certain subfield. In implementation, it first uses a closure to compute the subfield pointerdata, then throughskip_dropdecomposes the original guard into aMutexGuardInnerthat does not trigger Drop, and finally constructs a new guard📎 tokio/src/sync/mutex.rs:869-883。skip_dropusingManuallyDrop + ptr::readto transfer field ownership, avoidingDropbeing called twice📎 tokio/src/sync/mutex.rs:827-836. This is the classic Rust technique of "transferring ownership without triggering destruction."

Semaphore: How permit counting and wait queues implement backpressure

Intuitive model: Parking lot spaces

A semaphore is like a parking lot:acquireis driving in—if there's a space, you enter; if not, you queue at the entrance;releaseis driving out—when a space frees up, the car at the head of the queue is notified to enter. The permit count is the total number of spaces,acquire_many(n)is a large vehicle occupying n spaces.

Data structures and memory layout

The publicSemaphoreis just a thin wrapper around the underlyingbatch_semaphore::Semaphore:

📎 tokio/src/sync/semaphore.rs:427-432

rust
pub struct Semaphore {
    ll_sem: ll::Semaphore,
    #[cfg(all(tokio_unstable, feature = "tracing"))]
    resource_span: tracing::Span,
}

SemaphorePermit<'a>holds a semaphore reference and a permit count:

📎 tokio/src/sync/semaphore.rs:442-445

rust
pub struct SemaphorePermit<'a> {
    sem: &'a Semaphore,
    permits: usize,
}

permitsThe field is the key to understandingforget/merge/split.forgetsetspermitsto zero📎 tokio/src/sync/semaphore.rs:1193-1195, so that on Drop it returns 0 permits—equivalent to "permanently consuming" those permits.splitcuts n permits from the current permits for the new permit📎 tokio/src/sync/semaphore.rs:1260-1271。mergemerges another permit's count in, and asserts that both come from the same semaphore📎 tokio/src/sync/semaphore.rs:1230-1240。

[Design inference and architectural trade-offs]

MAX_PERMITSisusize::MAX >> 3 📎 tokio/src/sync/semaphore.rs:476-479. Why shift right by 3 bits? The underlyingbatch_semaphoreneeds to encode state flags (such as a closed flag) in the high bits, so the available permit count is limited to the low bits, leaving the high bits for flags. This is a common technique for packing "count + state" into a singleusize.

Step-by-Step: Permit flow of acquire and release

Scenario: The semaphore starts with 2 permits, task Aacquire(), task Bacquire_many(2)。

acquire()delegates toll_sem.acquire(1), and upon success constructsSemaphorePermit { permits: 1 } 📎 tokio/src/sync/semaphore.rs:614-631。acquire_many(2)similarly, but passes 2📎 tokio/src/sync/semaphore.rs:661-679。

If permits are insufficient,ll_sem.acquire(n)returnsPending, and the Waker is enqueued. There is a fairness detail here: the documentation points out that if the head of the queue is aacquire_many(5)and only 3 permits remain, even if aacquire(1)behind it could be satisfied immediately, it must wait—because the large vehicle at the head occupies the queue📎 tokio/src/sync/semaphore.rs:19-24. This is the cost of strict FIFO, avoiding starvation.

The release path is in Drop:

📎 tokio/src/sync/semaphore.rs:1402-1404

rust
impl Drop for SemaphorePermit<'_> {
    fn drop(&mut self) {
        self.sem.add_permits(self.permits);
    }
}

add_permitsdelegates toll_sem.release(n) 📎 tokio/src/sync/semaphore.rs:568-570, and the underlying layer returns the permit to the wait queue, waking waiters that can accumulate enough permits.

Regarding memory ordering, the documentation gives a strong guarantee: acquire, release, and close are allAcqReloperations, totally ordered with respect to each other, equivalent to those on a single atomic variableAcqRel 📎 tokio/src/sync/semaphore.rs:35-42. This means that a write that "writes data first and then releases the permit" is visible to a task that "acquires the permit later"—the semaphore can safely transfer data between tasks.

Design considerations: close and backpressure

close()causes all waiters to receiveAcquireError, and subsequenttry_acquirereturnsClosed 📎 tokio/src/sync/semaphore.rs:1161-1163. This is the foundation of graceful shutdown: when the receiver no longer needs data, closing the semaphore allows all blocked senders to fail and return immediately, instead of waiting forever.

The essence of backpressure is clearest in mpsc. As we will see in the next section, mpsc's capacity control is implemented with a semaphore whose permit count equals the buffer size.

Channel family: different trade-offs between waiter queues and Waker wakeups

Intuitive model: four kinds of channels, four waiting strategies

oneshotis a "one-shot envelope"—it can deliver only one message, and the sender does not wait (sendis synchronous), while the receiverawaitwaits for the message.mpscis a "bounded conveyor belt"—the sender waits when the belt is full, and the receiver waits when it is empty; capacity is controlled by a semaphore.broadcastandwatchare "broadcast loudspeakers"—one sender, multiple receivers, but the two handle "falling behind" in completely different ways.

The source material in this section focuses ononeshotandmpsc::bounded, and we will break them down one by one.

oneshot: a minimal handshake encoded with state bits

oneshot'sInnerstructure is the core of understanding its design:

📎 tokio/src/sync/oneshot.rs:386-409

rust
struct Inner<T> {
    state: AtomicUsize,
    value: UnsafeCell<Option<T>>,
    tx_task: Task,
    rx_task: Task,
}

stateis aAtomicUsize, using bit flags to encode the entire channel state. The four flag bits are defined at the end of the file:

📎 tokio/src/sync/oneshot.rs:1488-1505

rust
const RX_TASK_SET: usize = 0b00001;
const VALUE_SENT: usize = 0b00010;
const CLOSED: usize = 0b00100;
const TX_TASK_SET: usize = 0b01000;

valueisUnsafeCell<Option<T>>,tx_taskandrx_taskareTasktypes, internallyUnsafeCell<MaybeUninit<Waker>> 📎 tokio/src/sync/oneshot.rs:411-411. NoteMaybeUninit—the Waker may be uninitialized, and whether it is valid is determined by thestateinRX_TASK_SET/TX_TASK_SETbit📎 tokio/src/sync/oneshot.rs:396-399。

The essence of this design:VALUE_SENTThe bit not only indicates "the value has been sent," but also determines ownership of access toUnsafeCell. The comment is very explicit📎 tokio/src/sync/oneshot.rs:1491-1496: ifVALUE_SENTis set,UnsafeCellcan only be accessed by the receiver; if not set, it can only be accessed by the sender. This uses a single atomic bit to implement lock-free ownership transfer, avoiding an extra lock.

send's flow:

📎 tokio/src/sync/oneshot.rs:622-646

rust
pub fn send(mut self, t: T) -> Result<(), T> {
    let inner = self.inner.take().unwrap();
    inner.value.with_mut(|ptr| unsafe {
        *ptr = Some(t);
    });
    if !inner.complete() {
        unsafe {
            return Err(inner.consume_value().unwrap());
        }
    }
    Ok(())
}

First write the value intoUnsafeCell(at this pointVALUE_SENTis not set, so the receiver will not access it), then callcomplete()to try to setVALUE_SENT。complete()is a CAS loop:

📎 tokio/src/sync/oneshot.rs:1516-1549

rust
fn set_complete(cell: &AtomicUsize) -> State {
    let mut state = cell.load(Ordering::Relaxed);
    loop {
        if State(state).is_closed() {
            break;
        }
        match cell.compare_exchange_weak(
            state, state | VALUE_SENT, Ordering::AcqRel, Ordering::Acquire,
        ) {
            Ok(_) => break,
            Err(actual) => state = actual,
        }
    }
    State(state)
}

Why use CAS instead of a simplefetch_or? The comment explains it clearly📎 tokio/src/sync/oneshot.rs:1517-1529: if the channel is alreadyCLOSED, thenmust notsetVALUE_SENTagain. Because once it is set, the receiver will think it can accessUnsafeCell, while the sender is preparing to take the value back (consume_value), and simultaneous access from both sides would cause a data race. So the CAS loop breaks early when it seesCLOSED, without setting the bit.

complete()AfterRX_TASK_SETreturns, if the bit was successfully set and

📎 tokio/src/sync/oneshot.rs:1300-1315

rust
fn complete(&self) -> bool {
    let prev = State::set_complete(&self.state);
    if prev.is_closed() {
        return false;
    }
    if prev.is_rx_task_set() {
        unsafe {
            self.rx_task.with_task(Waker::wake_by_ref);
        }
    }
    true
}

Copypoll_recvThe receiver's

📎 tokio/src/sync/oneshot.rs:1317-1384

is the core of the state machine:is_complete()It first loads the state; ifconsume_valuethen directlyis_closed()return; ifErrreturnis_rx_task_set(); otherwise enter the "register Waker" branch. When registering, first checkwill_wake; if it is already set andis_complete()determines it is the same Waker, do not set it again; if different, first unset and then set. There is a subtle race handling here: after unset, if it is found thathas become true, the flag bit must be 📎 tokio/src/sync/oneshot.rs:1342-1344set back again

, otherwise the Waker will leak on Drop (because Drop relies on the flag bit to determine whether to drop the Waker).poll_closedThis "unset then set again" pattern also appears in📎 tokio/src/sync/oneshot.rs:839-848, and is the standard technique oneshot uses to handle concurrent wakeups.

mpsc::bounded: semaphore-driven backpressure

mpsc's capacity control is entirely delegated to the semaphore.channelThe function creates a semaphore whose permit count equals the buffer:

📎 tokio/src/sync/mpsc/bounded.rs:159-171

rust
pub fn channel<T>(buffer: usize) -> (Sender<T>, Receiver<T>) {
    assert!(buffer > 0, "mpsc bounded channel requires buffer > 0");
    let semaphore = Semaphore {
        semaphore: semaphore::Semaphore::new(buffer),
        bound: buffer,
    };
    let (tx, rx) = chan::channel(semaphore);
    let tx = Sender::new(tx);
    let rx = Receiver::new(rx);
    (tx, rx)
}

Semaphoreis an internal mpsc wrapper that holds both the underlying semaphore andbound(maximum capacity)📎 tokio/src/sync/mpsc/bounded.rs:176-179。boundis used formax_capacityqueries, whileavailable_permitsgives the current capacity📎 tokio/src/sync/mpsc/bounded.rs:591-593。

The send pathsendfirstreservethensend:

📎 tokio/src/sync/mpsc/bounded.rs:816-824

rust
pub async fn send(&self, value: T) -> Result<(), SendError<T>> {
    match self.reserve().await {
        Ok(permit) => {
            permit.send(value);
            Ok(())
        }
        Err(_) => Err(SendError(value)),
    }
}

reserveInternally callsreserve_inner(1), which first checksn > max_capacityand directly returns an error, thenacquire(n) 📎 tokio/src/sync/mpsc/bounded.rs:1272-1311. There is an ingeniousWakeReceiverOnDropguard here:

📎 tokio/src/sync/mpsc/bounded.rs:1286-1301

rust
struct WakeReceiverOnDrop<'a, T> {
    chan: &'a chan::Tx<T, Semaphore>,
}
impl<T> Drop for WakeReceiverOnDrop<'_, T> {
    fn drop(&mut self) {
        use chan::Semaphore;
        let semaphore = self.chan.semaphore();
        if semaphore.is_closed() && semaphore.is_idle() {
            self.chan.wake_rx();
        }
    }
}

The comment explains the motivation📎 tokio/src/sync/mpsc/bounded.rs:1279-1285: ifreserveis canceled after acquiring partial permits (for example,select!loses), the underlyingAcquirewill return these permits on Drop, butwill notnotify the receiver likePermitdoes. If the channel is already closed and idle at this point, the receiver may never receive the "channel closed" notification. This guard makes up for that wakeup on Drop. On success, usemem::forget(guard)to cancel the guard📎 tokio/src/sync/mpsc/bounded.rs:1306-1306, because the success path hasPermittake over the notification responsibility.

Permit's Drop does the same thing:

📎 tokio/src/sync/mpsc/bounded.rs:1732-1745

rust
impl<T> Drop for Permit<'_, T> {
    fn drop(&mut self) {
        use chan::Semaphore;
        let semaphore = self.chan.semaphore();
        semaphore.add_permit();
        if semaphore.is_closed() && semaphore.is_idle() {
            self.chan.wake_rx();
        }
    }
}

Permit::sendusesmem::forgetto skip Drop, avoiding returning permits📎 tokio/src/sync/mpsc/bounded.rs:1721-1728。

The receive pathrecvusespoll_fnto wrapchan.recv(cx) 📎 tokio/src/sync/mpsc/bounded.rs:243-246。poll_recvand directly delegates to📎 tokio/src/sync/mpsc/bounded.rs:650-652. The real waiter queue logic is in thechanmodule (not covered in this chapter), but it can be inferred: the receiver Waker is stored inchan::Rx, and is woken when the sendersend.

try_sendshows the non-blocking path:

📎 tokio/src/sync/mpsc/bounded.rs:924-934

rust
pub fn try_send(&self, message: T) -> Result<(), TrySendError<T>> {
    match self.chan.semaphore().semaphore.try_acquire(1) {
        Ok(()) => {}
        Err(TryAcquireError::Closed) => return Err(TrySendError::Closed(message)),
        Err(TryAcquireError::NoPermits) => return Err(TrySendError::Full(message)),
    }
    self.chan.send(message);
    Ok(())
}

try_acquireThe two errors ofClosedmap precisely toFulland

, distinguishing the two kinds of failure: "channel closed" and "buffer full."

Design considerations: cancellation safety and message loss📎 tokio/src/sync/mpsc/bounded.rs:776-784:sendThe mpsc documentation repeatedly emphasizes cancellation safetyselect!Whenloses in, the message will be discardedreserve. To avoid loss, you must usePermitto obtainsendand thenPermit—becausesendhas already reserved capacity,

recvis synchronous and will not be interrupted.📎 tokio/src/sync/mpsc/bounded.rs:199-204is cancellation-saferecv: ifselect!loses inrecv, it guarantees that no message is consumed. This is becausepoll_recv'sReady,Pendingreturns only when a message is actually obtained

oneshotand does not touch the queue whenReceiver.📎 tokio/src/sync/oneshot.rs:246-251'soneshotas a Future is also cancellation-safesend. But note:Err's

is synchronous, so there is no problem of "send being canceled"—either it is sent out, or

Pitfall 1: Using an async Mutex to protect pure data.The documentation explicitly recommends📎 tokio/src/sync/mutex.rs:26-36: if what is being protected is pure data (with no.awaitrequirement), usestd::sync::Mutexorparking_lotinstead, which are faster. The overhead of an async Mutex lies in the atomic operations of the semaphore and possible task scheduling. Only when you need to.awaitwhile holding the lock (for example, holding the lock to access a database connection) should you use an async Mutex.

Pitfall 2: Holding the lock across.awaitcauses deadlock.This is the most dangerous trap of the async Mutex. If task A holds the lock and then.awaitan event that requires task B to complete, while task B is waiting for this lock, a deadlock occurs.std::sync::MutexThe guard ofSendis not.await(in movable tasks), and the compiler will prevent holding acrossSend 📎 tokio/src/sync/mutex.rs:314-314; but the guard of an async Mutex is

, and the compiler will not stop you, so you need to ensure yourself that no circular wait is formed.reservePitfall 3:send。 Permitforgetting📎 tokio/src/sync/mpsc/bounded.rs:1732-1745's Drop will return the permit

, so capacity will not leak. But if the channel is already closed and idle, Drop will wake the receiver—this wakeup is necessary, otherwise the receiver might never receive the close notification.oneshotPitfall 4:poll'sPending。may falsely📎 tokio/src/sync/oneshot.rs:236-242The documentation statespoll: even if the message has been sent,Pendingmay still return

. This is not a bug, but a normal phenomenon under a concurrency race—the caller will be woken to retry, the message will not be lost, only delayed.forget_permitsPitfall 5: forget_permits(n)'s semantics.📎 tokio/src/sync/semaphore.rs:576-578Attempts to decrease n permits and returns the actual number decreased

. It does not block, nor does it wake waiters—it simply "swallows" permits. Used to dynamically shrink the semaphore capacity.

Chapter Summarytokio::syncThis chapter reveals's core pattern:。

  • MutexAll async wait primitives are built on "waiter queue + Waker wakeup", and the specific implementation of the queue varies by scenarioMutexGuardReuses a semaphore with a permit count of 1,release(1)only holds a reference, and on Drop
  • Semaphore, FIFO fair but does not poison.SemaphorePermitis permit count + wait queue,permitsusesforget/merge/split,MAX_PERMITScounting to support
  • oneshotright-shifting by 3 bits to reserve space for state flags.AtomicUsizeuses a singleVALUE_SENT's bit flags to encode state,UnsafeCellbits simultaneously determineCLOSED's access ownership, and the CAS loop prevents setting after
  • mpsc::bounded.WakeReceiverOnDropuses a semaphore whose permit count equals the buffer to implement backpressure,

the guard handles wakeup compensation on cancellation.

Chapter Review and Self-Testset_completeQ: Iffetch_or(VALUE_SENT)'s CAS loop were changed to a simple

, in what concurrency scenario would a data race be triggered?:set_completeReference Analysisfetch_orThe reason📎 tokio/src/sync/oneshot.rs:1517-1529uses a CAS loop instead ofVALUE_SENTis stated in the commentsCLOSED: it must checkfetch_orbefore settingclose(). If changed to an unconditionalCLOSED 📎 tokio/src/sync/oneshot.rs:1569-1574, consider this timing: the receiver first callssendto setfetch_or(VALUE_SENT), then the sender subsequentlyVALUE_SENTwrites the value andCLOSED. At this pointpoll_recvandis_complete()are set simultaneously, the receiver'sconsume_valuesees📎 tokio/src/sync/oneshot.rs:1325-1330as true and will callcomplete()to take the valueprev.is_closed(); while after the sender'sconsume_valuereturns, because📎 tokio/src/sync/oneshot.rs:1300-1315is true, it will callUnsafeCellto take the value backCLOSED. Both sides accessVALUE_SENTsimultaneously, a data race. The CAS loop breaks early upon discovering

Q: reserve_inner, without settingWakeReceiverOnDrop, thereby guaranteeing the invariant that "after closing, the sender has exclusive access."mem::forgetInforget, the

guard usesto skip on the success path; what happens if this📎 tokio/src/sync/mpsc/bounded.rs:1290-1298is removed?acquire(n)Reference AnalysisOk: the guard's Drop logic is "if the semaphore is already closed and idle, wake the receiver"Permit. On the success path,Permitreturnsreserve_inner, the caller obtains the permit and will constructis_idle, andPermitis responsible for the subsequent notification duty. If the guard is not removed, the guard will Drop when the function returns, and will additionally check once for "closed and idle"—but at this point the permit is already held by the caller ofmem::forget, so the semaphore is not idle (forgetis false), so in fact it will not wake twice. But more critically, the semantics are clear: the wakeup responsibility on the success path should be entirely borne byacquire, and the guard is only responsible for compensation on the "cancel/failure" path.Okexplicitly expresses the intent that "this path does not need the guard." IfPermitis removed and the semaphore happens to be in the boundary state of "closed and idle" (for example,

returnsMutexGuardbut the permit has not yet been taken over bySemaphorePermit), it may produce one extra wakeup—although it will not cause an error, it wastes one scheduling.

Q: Ifwere changed to hold a semaphore permit object (likeMutexGuard), what problems would be introduced?&MutexReference Analysisself.lock.s.release(1) 📎 tokio/src/sync/mutex.rs:959-961: currentlyMutexGuard::maponly holdsMappedMutexGuard, and on Drop calls📎 tokio/src/sync/mutex.rs:869-883. If changed to hold a permit object, several problems would be introduced. First,MappedMutexGuardthe series of methods need to decompose the guard into&Semaphore, protecting only the subfield📎 tokio/src/sync/mutex.rs:190-199. Under the current design,self.s.release(1) 📎 tokio/src/sync/mutex.rs:1252-1262only needs to holdMappedMutexGuardand the subfield pointerpermits: usize, and on DropMutexGuard. If the guard held a permit object, then map would have to transfer ownership of the permit object, andSend/Sync's field layout would become more complex. Second, the permit object usually carries aunsafe implcount, and for a Mutex this count is always 1, which is redundant. Third,📎 tokio/src/sync/mutex.rs:260-263'smap。

boundary is already precisely controlled throughtokio::sync, and holding a permit object would introduce additional trait constraints. The current design of "only holding a reference + manual release" is lighter and also easier to supportspawn_blockingAt this point, we have clearly seenblock_onhow, using the unified pattern of "waiter queue + Waker wakeup", it supports async waiting for Mutex, Semaphore, and various channels. But not all blocking can be made asynchronous—some operations (such as file system calls, CPU-intensive computation) will inherently block the thread. In the next chapter we will enter

The storage location of the Waker varies by primitive: Mutex/Semaphore store it in the underlying semaphore's wait queue, oneshot stores it in the Inner's tx_task/rx_task fields, and mpsc stores it in the chan module's send/receive queues. But the wakeup mechanism is unified: when state changes, the Waker is taken out and wake_by_ref is called, and the executor reschedules the task. At this point, the waiting and wakeup inside async primitives are clearly visible. However, not all code can be made async—the next chapter will explore how to bridge blocking operations with spawn_blocking, and how block_on drives Futures in non-async contexts.

CHAPTER 08

Chapter 8: Blocking and Bridging: spawn_blocking Thread Pool and the Boundaries of block_on

Project: tokio-rs/tokio · Book progress: Chapter 8 / 14 · Verification status: FACT line numbers genuinely anchored

In the previous chapter we saw that the key reason async Mutex and channels can wait without occupying a thread is that they store the Waker in the wait queue, and once the condition is satisfied the waker reschedules the task. But all of this presupposes that the task can voluntarily yield the thread when Pending. Once code calls std::fs::read, libsqlite3, or a pure CPU compression loop, it will monopolize the worker thread until it returns, during which all other tasks on that thread starve. Tokio's solution is to outsource such work to a separate blocking thread pool, and use block_on to drive Futures in non-async contexts. This chapter dissects these two boundaries.

8.1 Memory Layout of the Blocking Thread Pool: Inner and the Dual-Implementation Queue

Intuitive Model:spawn_blockingThe thread pool is like a restaurant's "outsourced helper pool." The front-of-house waiters (worker threads) only handle taking orders and delivering dishes; when they encounter a dish that needs slow stewing, they write a work order and toss it into the kitchen's pass-through window (queue), and the helpers (blocking threads) take orders from the window. Without this pool, the waiters would have to cook themselves, and the whole restaurant would grind to a halt.

Core Structures. The entire pool is held byBlockingPoolwhich stores only two things: a cloneableSpawner(submission entry) and ashutdown_rx(shutdown signal receiver)📎 tokio/src/runtime/blocking/pool.rs:20-23。Spawnerinternally isArc<Inner>, all submitters share the same state📎 tokio/src/runtime/blocking/pool.rs:26-28。

Inneris the entire state of the pool, and its fields are worth examining one by one📎 tokio/src/runtime/blocking/pool.rs:77-104:

  • inner_impl: InnerImpl: the implementation of queue + notification + lock topology, which is an enum withLockedandShardedtwo variants📎 tokio/src/runtime/blocking/pool.rs:107-110. This is the most critical abstraction in this chapter—it unifies the two topologies of "single-lock queue" and "sharded queue" under one interface.
  • thread_cap: usize: the upper limit on the number of threads, i.e.max_blocking_threads。
  • scheduler_threads: usize: the number of scheduler worker threads, used to subtract in metrics so thatnum_blocking_threadsonly counts blocking threads📎 tokio/src/runtime/blocking/pool.rs:455-460。
  • keep_alive: Duration: the idle thread survival duration, defaultKEEP_ALIVE = 10s 📎 tokio/src/runtime/blocking/pool.rs:231。
  • metrics: SpawnerMetrics: three atomic counters—num_threads、num_idle_threads、queue_depth 📎 tokio/src/runtime/blocking/pool.rs:31-35。
[Design Inference and Architectural Trade-offs]

Why use atomic counters instead of fields inside the lock? num_idle_threadsis read on the hot path ofspawn_task(to determine whether idle threads need to be woken); if it were hidden insideMutex, every submission would have to acquire the lock first and then read. By making itMetricAtomicUsize, the submission path can do a quick check first without holding the queue lock. The cost is that there is no atomicity guarantee between these counts and the queue state, so the code uses thenum_notifycounter to compensate—see below.

Thread Management State。ThreadManagementStateis extracted separately for reuse by both queue implementations📎 tokio/src/runtime/blocking/pool.rs:135-150:

  • shutdown: bool: shutdown flag.
  • shutdown_tx: Option<shutdown::Sender>: each worker thread holds a clone, and after all are droppedshutdown_rxreceives the notification.
  • last_exiting_thread: Option<JoinHandle<()>>: the handle of the last thread that exited due to timeout.
  • worker_threads: HashMap<usize, JoinHandle<()>>: the handles of all surviving workers.
  • worker_thread_index: usize: a monotonically increasing thread ID allocator.

last_exiting_threadThe design motivation of is clearly stated in the comments: a thread that exits due to timeout will join the previous thread that exited due to timeout, avoiding Valgrind false positives📎 tokio/src/runtime/blocking/pool.rs:135-150。worker_timed_outis exactly the implementation of this chained join—it removes its own handle and swaps out the oldlast_exiting_threadto return to the caller for joining📎 tokio/src/runtime/blocking/pool.rs:172-178。

Task Wrapping. What is stored in the queue isTask, which wraps aUnownedTask<BlockingSchedule>and aMandatoryflag📎 tokio/src/runtime/blocking/pool.rs:187-191。Mandatorydetermines whether the task is discarded or forcibly executed on shutdown:shutdown_or_run_if_mandatorycallsNonMandatorywhenshutdown(), and callsMandatorywhenrun() 📎 tokio/src/runtime/blocking/pool.rs:223-228. This is the difference betweenspawn_blocking(non-forced) andspawn_mandatory_blocking(forced, used by fs)📎 tokio/src/runtime/blocking/pool.rs:233-265。

Memory Layout of the Single-Lock Implementation。LockedImplis the most primitive topology: oneMutex<LockedInner>plus oneCondvar 📎 tokio/src/runtime/blocking/pool.rs:113-116。LockedInnercontainsVecDeque<Task>、num_notify: u32andthread_mgmt_state 📎 tokio/src/runtime/blocking/pool.rs:118-124. Note thatnum_notifyandthread_mgmt_stateare under the same lock, whilenum_idle_threadsis an atomic outside the lock—this hybrid layout of "part of the state inside the lock, part outside" is precisely the source of all the concurrency subtleties that follow.

8.2 Submission Path: From spawn_blocking to Thread Wakeup

Scenario: an async task callstokio::task::spawn_blocking(move || heavy_compute(data)), what happens at this moment?

Step 1: Boxing Decision and Task Construction。Spawner::spawn_blockingfirst measures the closure sizefn_size, then based onAutoBox::<F>::SHOULD_BOXdecides whether toBoxthe closure📎 tokio/src/runtime/blocking/pool.rs:359-389. This is Tokio's general "auto-box large Futures" strategy: box when the closure is too large, to avoid bloating the task struct.

Enteringspawn_blocking_inner, first allocate a task ID, then useblocking_taskto wrap the closure into a Future, and finally usetask::unownedto constructUnownedTaskandJoinHandle 📎 tokio/src/runtime/blocking/pool.rs:440-449. Note that what is returned here is the(JoinHandle<R>, Result<(), SpawnError>)tuple—the handle and the submission result are returned separately.

Step 2: Three Ways to Handle the Submission Result. Back inspawn_blocking, match onspawn_result📎 tokio/src/runtime/blocking/pool.rs:381-388:

  • Ok(()): normal, return the handle.
  • Err(ShuttingDown):does not panic, still returns a handle. The comment explains this is for compatibility—the handle will never resolve, but the caller won't crash because the runtime is shutting down.
  • Err(NoThreads(e)): the OS cannot create a thread and no one in the pool takes over, so it panics directly.

Step 3: Enqueue and wakeup decision。spawn_taskpasses theon_no_idleclosure toInnerImpl::spawn_task, and the concrete implementation decides when to call it📎 tokio/src/runtime/blocking/pool.rs:462-506. Look atLockedImpl::spawn_task's critical section📎 tokio/src/runtime/blocking/pool.rs:603-639:

rust
let mut locked = self.mutex.lock();

if locked.thread_mgmt_state.shutdown {
    task.task.shutdown();
    return Err(SpawnError::ShuttingDown);
}

locked.queue.push_back(task);
metrics.inc_queue_depth();

if metrics.num_idle_threads() == 0 {
    on_no_idle(&mut locked.thread_mgmt_state)?;
} else {
    metrics.dec_num_idle_threads();
    locked.num_notify += 1;
    self.condvar.notify_one();
}

There are two key points here. First, the shutdown check happens before enqueueing, and even if the task isMandatoryit is directlyshutdown()—the comment explains: it was only scheduled after shutdown began, so discarding it is legal📎 tokio/src/runtime/blocking/pool.rs:614-620. Second, the wakeup decision depends on the out-of-locknum_idle_threads: if it is 0, callon_no_idleto try to start a new thread; otherwise decrement the idle count and incrementnum_notify、notify_one。

num_notifyWhy must it exist?BecauseCondvarmay produce spurious wakeups. If onlynotify_oneis used without counting, a spuriously woken thread will mistakenly think there is a task to take, find the queue empty, and go back to sleep, while the thread that was actually woken may never receive the notification.num_notifyturns "legal wakeup" into a countable token: the submitter+1, and the woken side only considers the wakeup legal whennum_notify != 0and-1 📎 tokio/src/runtime/blocking/pool.rs:674-684。

Step 4: Start a new thread。on_no_idleThe closure executes📎 tokio/src/runtime/blocking/pool.rs:462-506while holding the queue lock. It first checksnum_threads == thread_cap, and if the upper limit is reached it returns directlyOk(())—the task stays in the queue waiting for an existing thread to handle it, which is backpressure. Otherwise cloneshutdown_tx, callspawn_threadto create a thread, and after success incrementnum_threads, incrementworker_thread_index, and insert the handle intoworker_threads。

spawn_threadUsethread::Builderto set the thread name and stack size, then spawn a closure: enter the runtime contextrt.enter(), callinner.run(id), and finally dropshutdown_tx 📎 tokio/src/runtime/blocking/pool.rs:508-528。

Fault tolerance for OS thread creation failure。spawn_threadmay fail. The code classifies the error📎 tokio/src/runtime/blocking/pool.rs:488-500: if it isWouldBlock(a temporary error, determined byis_temporary_os_thread_error) and there is already a blocked thread in the pool, then📎 tokio/src/runtime/blocking/pool.rs:750-752silently ignore—the task will eventually be taken by some currently busy thread. Otherwise return, which ultimately causes a panic.SpawnError::NoThreadsUse a control-flow diagram to summarize the decision branches of the submission path:

Copy

mermaid
flowchart TD
    call["Spawner::spawn_blocking(func)"] --> box{"AutoBox::SHOULD_BOX?"}
    box -->|是| boxed["Box::new(func)"]
    box -->|否| raw["func"]
    boxed --> inner["spawn_blocking_inner"]
    raw --> inner
    inner --> unowned["task::unowned -> Task + JoinHandle"]
    unowned --> spawn_task["InnerImpl::spawn_task"]
    spawn_task --> lock["LockedImpl: mutex.lock()"]
    lock --> shutting{"thread_mgmt_state.shutdown?"}
    shutting -->|是| discard["task.task.shutdown()"]
    discard --> err_sd["Err(ShuttingDown)"]
    shutting -->|否| push["queue.push_back(task)"]
    push --> idle{"num_idle_threads == 0?"}
    idle -->|是| on_no_idle["on_no_idle(thread_mgmt_state)"]
    on_no_idle --> cap{"num_threads == thread_cap?"}
    cap -->|是| backpressure["返回 Ok, 任务留队列"]
    cap -->|否| spawn_th["spawn_thread(shutdown_tx, rt, id)"]
    spawn_th --> th_ok{"spawn 成功?"}
    th_ok -->|是| reg["inc_num_threads, 注册 JoinHandle"]
    th_ok -->|否| tmp{"WouldBlock 且已有线程?"}
    tmp -->|是| ignore["忽略, 等忙碌线程取走"]
    tmp -->|否| err_nt["Err(NoThreads)"]
    idle -->|否| notify["dec_num_idle_threads, num_notify+=1, notify_one"]
    err_sd --> ret["返回 JoinHandle"]
    backpressure --> ret
    reg --> ret
    ignore --> ret
    err_nt --> panic_os["panic: OS can't spawn worker thread"]

Intuitive model

: each blocking thread is a "standby helper." When there are orders, it works continuously (BUSY); when there are none, it naps (IDLE), and if it naps longer thanit goes off duty (timeout exit). Without timeout reclamation, the pool would permanently retain all threads created at peak times, wasting memory and kernel scheduling overhead.keep_aliveMain loop structure

is a。LockedImpl::run_workerloop, internally alternating between the two phases BUSY and IDLE'main. Note: BUSY/IDLE here are📎 tokio/src/runtime/blocking/pool.rs:642-735phaseswithin the loop, not explicit enum states, so the following is described with a flowchart rather than a state diagram.BUSY phase

: the innercontinuously takes taskswhile let Some(task) = locked.queue.pop_front(). After taking one, decrement📎 tokio/src/runtime/blocking/pool.rs:655-661drop the lockqueue_depth,, execute, then reacquire the lock. The step of dropping the lock is crucial—a blocking task may run for a long time, and it must never be executed while holding the lock.task.run()IDLE phase

: the queue is empty, increment, setnum_idle_threads, then enter the wait loopis_counted_idle = true. The core is📎 tokio/src/runtime/blocking/pool.rs:663-696, and after it returns, check three things:condvar.wait_timeout(locked, keep_alive): legal wakeup. Decrement

1. num_notify != 0, setnum_notify(because the submitter has already decrementedis_counted_idle = false), break back to BUSYnum_idle_threads2. Not shut down and timed out: call📎 tokio/src/runtime/blocking/pool.rs:674-684。

to get the handle of the previous exited thread,worker_timed_outexit the loopbreak 'main3. Otherwise it is a spurious wakeup, continue waiting.📎 tokio/src/runtime/blocking/pool.rs:689-693。

Queue draining on shutdown

. Ifis true, enter the draining logicthread_mgmt_state.shutdown: pop tasks one by one, drop the lock, call📎 tokio/src/runtime/blocking/pool.rs:698-710—non-forced tasks are discarded, forced tasks execute as usual. Then break out of the main loop.task.shutdown_or_run_if_mandatory()Exit cleanup

. Before the thread exits, decrement. Ifnum_threads 📎 tokio/src/runtime/blocking/pool.rs:714is true, also decrementis_counted_idle, and usenum_idle_threadsto assert there is no underflowassert_ne!(prev_idle, 0). This assertion is a debug-time guardrail: once📎 tokio/src/runtime/blocking/pool.rs:716-726accounting goes wrong, it will panic immediately here instead of letting the error propagate silently.num_idle_threadsFinally, if shutting down and

(the last thread),num_threads == 0wakes up the shutdown initiator that may be waitingnotify_one. Return📎 tokio/src/runtime/blocking/pool.rs:728-730, andjoin_on_threadjoins before exitingInner::runShutdown handshake📎 tokio/src/runtime/blocking/pool.rs:755-771。

first calls。BlockingPool::shutdownto get all worker handlesbegin_shutdown, sets the shutdown flag, drops📎 tokio/src/runtime/blocking/pool.rs:310-312。LockedImpl::begin_shutdown, and wakes all waiting threadsshutdown_tx、notify_all. Then📎 tokio/src/runtime/blocking/pool.rs:740-745blocks waiting forshutdown_rx.wait(timeout)'s implementation is quite careful📎 tokio/src/runtime/blocking/pool.rs:324。

shutdown::Receiver::wait: first handle📎 tokio/src/runtime/blocking/shutdown.rs:37-70's fast path and return false directly; then calltimeout == 0to enter the blocking region, and if it fails and the current thread is panicking, return false; otherwise panic with the hint "cannot drop runtime in async context"try_enter_blocking_region(). Finally, depending on timeout, call📎 tokio/src/runtime/blocking/shutdown.rs:44-57orblock_on_timeoutto drive that oneshot.block_onThe mechanism of

shutdown_txis: each worker thread holds a clone ofArc<oneshot::Sender<()>>. After all threads exit, all clones are dropped,📎 tokio/src/runtime/blocking/shutdown.rs:12-14the count reaches zero,Arcis dropped,oneshot::SenderandReceiverreceives the notification. This is the classic pattern of "Receiver is woken after all Senders are dropped."

mermaid
sequenceDiagram
    participant App as "应用线程 (drop Runtime)"
    participant Pool as "BlockingPool::shutdown"
    participant Locked as "LockedImpl"
    participant Worker as "阻塞 worker 线程"
    participant Rx as "shutdown::Receiver"

    App->>Pool: shutdown(timeout)
    Pool->>Locked: begin_shutdown()
    Locked->>Locked: thread_mgmt_state.begin_shutdown() 设 shutdown=true, shutdown_tx=None
    Locked->>Worker: condvar.notify_all()
    Locked-->>Pool: Some((last_exited_thread, workers))
    Pool->>Rx: wait(timeout)
    Worker->>Worker: 从 wait_timeout 醒来, 见 shutdown=true
    Worker->>Worker: 排空队列 shutdown_or_run_if_mandatory()
    Worker->>Worker: dec_num_threads, 退出 run_worker
    Worker->>Worker: drop(shutdown_tx) 克隆
    Worker-->>Rx: 最后一个 Sender drop, oneshot 完成
    Rx-->>Pool: 返回 true
    Pool->>Worker: join 所有 worker 句柄

8.4 block_on: driving a Future in a non-async context

Intuitive model:block_onis the runtime's "front door." It turns the current thread into a temporary executor, repeatedly polling the passed-in Future until completion. Without it,mainfunctions cannot start any asynchronous code.

Entry and boxing。Runtime::block_onlikewise first measures the size, decides whether toSHOULD_BOXbased onBox::pin, then entersblock_on_inner 📎 tokio/src/runtime/runtime.rs:343-350。block_on_inner. Inside there are two conditionally compiled trace wrappers (taskdump and tracing), thenself.enter()enters the runtime context, and finally dispatches by scheduler type📎 tokio/src/runtime/runtime.rs:353-383:

rust
let _enter = self.enter();

match &self.scheduler {
    Scheduler::CurrentThread(exec) => exec.block_on(&self.handle.inner, future),
    Scheduler::MultiThread(exec) => exec.block_on(&self.handle.inner, future),
}

The two schedulers'block_onThe semantics are different, and the documentation states this very clearly.📎 tokio/src/runtime/runtime.rs:302-320:

  • Multi-threaded scheduler: the Future runs in the context of the I/O driver and timer,block_onand after returning, tasks that have already been spawned continue to run.
  • Current-thread scheduler:block_oncan be called concurrently by multiple threads; the first caller takes ownership of the I/O and timer drivers, and other threads "hook into" it. After the firstblock_oncompletes, other threads can "steal" the driver.block_onAfter returning, tasks that have already been spawned are suspended, and callingblock_onagain will resume them.

Key restriction: it cannot be called in an asynchronous context.. The documentation explicitly states thatblock_oncalling it in an asynchronous execution context will panic.📎 tokio/src/runtime/runtime.rs:321-324. The reason is straightforward:block_onit blocks the current thread until the Future completes. If the current thread itself is a worker thread, it will block the entire executor—this is exactlyspawn_blockingthe problem that is meant to solve, so the two are mutually exclusive.

Shutdown path。Runtime::dropdispatches by scheduler type📎 tokio/src/runtime/runtime.rs:506-521: the current-thread scheduler needs to firsttry_set_currententer the context and then shut down (ensuring tasks are dropped in the runtime context); the multi-threaded scheduler shuts down directly (the worker threads themselves are already in the context).shutdown_timeoutShut down the scheduler first, then shut down the blocking pool.📎 tokio/src/runtime/runtime.rs:457-461,shutdown_backgroundis equivalent toshutdown_timeout(Duration::from_nanos(0)) 📎 tokio/src/runtime/runtime.rs:494-496。

Design considerations, error recovery, and production pitfalls

Why doesspawn_blocking'sShuttingDownnot panic? 📎 tokio/src/runtime/blocking/pool.rs:383-384The comment says it is for compatibility considerations.spawn_blockingreturnsJoinHandlerather thanResult. If it panicked during shutdown, it would turn the predictable state of "the runtime is shutting down" into a crash. Returning a handle that never resolves means the callerawaitwill hang forever—but at this point the runtime has already shut down, and the entireblock_onwill also exit, so in practice it will not leak permanently.

max_blocking_threads's backpressure semantics. The default value is very large (512), becausespawn_blockingis often used for file I/O. But the documentation warns: when running CPU-intensive tasks, use a semaphore to limit concurrency, otherwise a large number of threads will be created📎 tokio/src/task/blocking.rs:94-100. Once the upper limit is reached, tasks queue in the queue, forming backpressure—but note that this backpressure only applies to the blocking pool and does not propagate back to the asynchronous scheduler.

spawn_blockingcannot be canceled. The documentation explicitly states:aborthas no effect on blocking tasks that have already started running; the tasks will continue to run to completion📎 tokio/src/task/blocking.rs:106-120. Only tasks that have not yet started may be prevented by abort. During shutdown, the runtime waits for all blocking tasks that have already started, andshutdown_timeoutafter the timeout, these threads will be leaked.

num_idle_threads's accounting trap。is_counted_idleThe existence of the flag shows that this count is very easy to get wrong. The submitting side decrementsnum_idle_threadswhen waking up, and the awakened side, after seeingnum_notify != 0, setsis_counted_idle = false, avoiding a duplicate decrement📎 tokio/src/runtime/blocking/pool.rs:679-682. If there is a bug in this path,assert_ne!(prev_idle, 0)will panic on exit📎 tokio/src/runtime/blocking/pool.rs:722-725. If you see "num_idle_threadsunderflowed on thread exit" in production, it means the pool's accounting logic has been broken.

[Design inference and architectural trade-offs]

last_exiting_threadThe cost of chained join. A thread exiting due to timeout will join the previous thread that exited due to timeout📎 tokio/src/runtime/blocking/pool.rs:172-178. This forms a join chain: each exiting thread must wait for the previous one to truly finish. In scenarios with high-frequency creation/destruction of blocking threads, this chain may become long, causing thread exit latency to accumulate. This is a trade-off made to avoid Valgrind false positives. Its impact in normal production environments is limited, but it is worth paying attention to under loads where threads frequently time out.

InnerImplThe meaning of enum abstraction. The comment explains that the behavior of theLockedvariant is exactly the same as before the refactor, while theShardedvariant reserves a symmetric slot for a future concurrent queue📎 tokio/src/runtime/blocking/pool.rs:537-539。spawn_task、run_worker、begin_shutdown. All three methods dispatch through the enum📎 tokio/src/runtime/blocking/pool.rs:548-582. This design of "enum dispatch + each variant owning its own critical section" means that adding a new queue topology does not require changing the caller.

Chapter summary

This chapter breaks down the two boundaries through which Tokio accommodates synchronous code.spawn_blockingdelivers closures to an independent blocking thread pool:Innerit holds the queue, thread limit, keep-alive duration, and atomic metrics;LockedImplit implements the queue with a single lock +Condvar,num_notifyand the counter compensates for spurious wakeups; workers cycle between BUSY/IDLE, and after idle timeout they exit via chained join;max_blocking_threadsonce the upper limit is reached, tasks queue up to form backpressure.block_ondrives Futures in a non-asynchronous context; the multi-threaded and current-thread schedulers have different semantics, and calling it in an asynchronous context is strictly forbidden. The shutdown path is triggered byshutdown_tx'sArccount reaching zero, which triggersoneshot, implementing the handshake of "waking the shutdown initiator after all workers exit."

Chapter reflection and self-test

Q1: If inLockedImpl::spawn_taskthe check forif metrics.num_idle_threads() == 0is changed to always true (that is, callingon_no_idleevery time), what will happen in a high-concurrency submission scenario? Why?

Reference analysis:on_no_idlechecksnum_threads == thread_cap, and if the upper limit has not been reached, it creates a new thread📎 tokio/src/runtime/blocking/pool.rs:471-487. If the check is always true, it will try to start a new thread even when there are idle threads, causing the thread count to rapidly hitthread_cap. More seriously, idle threads will not be woken bynotify_one(because theon_no_idlebranch is taken rather than theelsebranch'snum_notify += 1; notify_one 📎 tokio/src/runtime/blocking/pool.rs:627-636), and tasks in the queue may go unprocessed until some new thread starts and discovers that the queue is non-empty. This creates a false-deadlock state of "threads maxed out but tasks still queued." The purpose of the original check is precisely this: when there are idle threads, wake them first and avoid unnecessary thread creation.

Q2: LockedImpl::run_workerBefore executingtask.run()in the BUSY phase, it willdrop(locked) 📎 tokio/src/runtime/blocking/pool.rs:657-658. If thisdropis removed, in what scenario will deadlock be triggered?

Reference analysis:task.run()executes the user closure, and the closure itself may very well callspawn_blockingagain to submit a new task. The submission pathLockedImpl::spawn_task's first action isself.mutex.lock() 📎 tokio/src/runtime/blocking/pool.rs:612. If the worker holds the lock while executing the closure, the submission inside the closure will try to acquire the same lock, andstd::sync::MutexNon-reentrant, direct deadlock. Furthermore, holding the lock while executing long tasks will block all other submitters and workers from fetching tasks; even without deadlock, it will serialize the entire pool.drop(locked)It is necessary.

Q3: shutdown::Receiver::waitIntry_enter_blocking_region()Returns false when it fails and is currently panicking; otherwise panics📎 tokio/src/runtime/blocking/shutdown.rs:44-57. Why special-case panics? If this branch were removed, in what scenarios would problems arise?

Reference analysis:try_enter_blocking_regionFailure means the current context is asynchronous, and blocking is not allowed. Normally it should panic to tell the user, "cannot drop runtime in an asynchronous context." But if the current thread is already panicking (std::thread::panicking()is true), panicking again would cause a double panic, and Rust's default behavior is to abort the process directly. Scenario: the user drops a Runtime inside an asynchronous task, and that task itself is already panicking for some other reason; then the shutdown triggered by drop causes a second panic. Returning false lets shutdown give up waiting, avoiding process abort and preserving the chance for the user to see the original panic information. This is a typical "panic safety" handling.

The blocking thread pool and block_on define the capability boundaries of the asynchronous runtime: the former isolates work that cannot yield the thread onto dedicated threads, while the latter allows non-async entry points to drive Futures. But these two boundaries are often not handwritten in code—in the next chapter we will enter the world of macros and see how #[tokio::main], select!, and join! generate this runtime code at compile time.

CHAPTER 09

Chapter 9: The Magic of Macros: The Code Generation Behind #[tokio::main], select!, and join!

Project: tokio-rs/tokio · Book progress: Chapter 9 / 14 · Verification status: FACT line numbers are truly anchored

In the previous chapter we sawblock_onand how the blocking thread pool defines the capability boundaries of the asynchronous runtime, while users almost never handwrite these boundaries—they write#[tokio::main]、select!、join!, letting macros expand this boilerplate at compile time. Macros are Tokio's first layer of sugar for users, and also where runtime code is truly generated at compile time. This chapter focuses on thetokio-macroscrate andtokio/src/macros/select.rs, dissecting the three most commonly used macro expansion paths, with emphasis on answering one question: after macro expansion, what does the real call chain look like, and why must the cancellation-safety semantics ofselect!be watched separately.

9.1 #[tokio::main]: Rewriting async fn into Runtime::block_on

Intuitive model:#[tokio::main]It is like a "renovation authorization form." You hand over a bare room (async fn main), and it lays the plumbing and wiring for you (builds the Runtime), installs the doors and windows (enable_all), and finally moves your original furniture (the function body) in. Without it, everymainwould have to handwriteBuilder::new_multi_thread().enable_all().build().unwrap().block_on(...), and boilerplate would drown the business logic.

Data structures and memory layout

The macro itself does not produce runtime data structures, but the configuration it parses is placed into two structs.Configurationis the "mutable accumulator during parsing," and its fields are allOption, because attribute parameters may be omitted, repeated, or illegal📎 tokio-macros/src/entry.rs:74-84. Note thatworker_threads、start_paused、unhandled_panicall carrySpan—this is to locate errors at the line the user wrote, rather than inside the macro📎 tokio-macros/src/entry.rs:74-84。FinalConfigis the "validated immutable result,"flavoris no longerOption, becausebuild()has already useddefault_flavoras a fallback📎 tokio-macros/src/entry.rs:55-62。

RuntimeFlavorhas only three variants:CurrentThread、Threaded、Local 📎 tokio-macros/src/entry.rs:10-14。from_strdeliberately gives friendly errors for legacy names:single_threadindicates it should be calledcurrent_thread,basic_schedulerindicates it has been renamed,threaded_schedulerindicates it has been renamed📎 tokio-macros/src/entry.rs:17-27. This is a typical design of macros as the "user's first point of contact": error messages are documentation.

Step-by-Step expansion process

Scenario: the user writes#[tokio::main(flavor = "multi_thread", worker_threads = 4)] async fn main() { ... }。

Step one,mainThe entry first parses the item into the customItemFn 📎 tokio-macros/src/entry.rs:577-580. ThisItemFnis notsyn::ItemFn, but a parser implemented by Tokio itself, for the reason stated in the comments: it does not want to recursively parse the entire statement, and only performs lightweight parsing by "buffering by token tree and splitting on semicolons"📎 tokio-macros/src/entry.rs:720-764. This avoids the overhead of building a complete AST for the function body inside the macro.

Step two,build_configvalidates whether theasynckeyword exists; if missing, it reports "theasync keyword is missing" 📎 tokio-macros/src/entry.rs:346-349. Then it iterates over the attribute parameters and dispatchesworker_threads、flavor、start_paused、crate、unhandled_panic、nameto the corresponding setter📎 tokio-macros/src/entry.rs:369-399. Note thatcore_threadsis explicitly rejected with a message that it has been renamed📎 tokio-macros/src/entry.rs:379-382。

Step three,Configuration::buildperforms cross-field consistency validation. There are three key constraints here:worker_threadsonly allowsmulti_thread 📎 tokio-macros/src/entry.rs:197-217;start_pausedonly allowscurrent_thread/local 📎 tokio-macros/src/entry.rs:219-229;unhandled_paniclikewise only allowscurrent_thread/local 📎 tokio-macros/src/entry.rs:231-241. If the user selectsmulti_threadbut thert-multi-threadfeature is not enabled, the error message differs depending on whether flavor was explicitly specified📎 tokio-macros/src/entry.rs:209-216。

Step four,parse_knobsgenerates code. It first erasesasyncness 📎 tokio-macros/src/entry.rs:441, then chooses the builder starting point according to flavor:CurrentThread/LocalusesBuilder::new_current_thread(),ThreadedusesBuilder::new_multi_thread() 📎 tokio-macros/src/entry.rs:468-477。LocalThe special thing is that the build call isbuild_local(Default::default())rather thanbuild() 📎 tokio-macros/src/entry.rs:479-483. Then it appends chained.worker_threads(#v)、.start_paused(#v)、.unhandled_panic(...)、.name(#v) 📎 tokio-macros/src/entry.rs:485-497。

as neededlast_block:return #rt.enable_all().#build.expect("Failed building the Runtime").block_on(body) 📎 tokio-macros/src/entry.rs:509-522Step five, generate the final function body. The core isreturn. Note the explicit📎 tokio-macros/src/entry.rs:508。

, whose comment points to tokio-rs/tokio#4636, to fix a type inference issueasync #bodyStep six, the function body is wrapped as!and type-checked. On the non-test path, if the return type is notimpl Traitand does not containif false { let _: &dyn Future<Output = #output_type> = &body; }, it inserts📎 tokio-macros/src/entry.rs:551-571for a compile-time assertionpin!Pin the body on the stack and convert it toPin<&mut dyn Future>, with comments explaining this is to reduceblock_onthe compilation overhead of generic instantiation📎 tokio-macros/src/entry.rs:526-548。

mermaid
flowchart TD
    entry["main(args, item)"] --> parse_item{"syn::parse2(item) 成功?"}
    parse_item -->|否| err_ret["token_stream_with_error 返回原始 item + 编译错误"]
    parse_item -->|是| check_main{"ident == main 且有参数?"}
    check_main -->|是| err_args["报错: main 不能接受参数"]
    check_main -->|否| parse_args["AttributeArgs::parse_terminated"]
    parse_args --> build_cfg["build_config 校验 async 与各字段"]
    build_cfg --> cfg_ok{"config 构建成功?"}
    cfg_ok -->|否| fallback["parse_knobs(DEFAULT_ERROR_CONFIG) + 错误"]
    cfg_ok -->|是| knobs["parse_knobs 生成 Builder 链 + block_on"]
    knobs --> out["输出同步 fn main"]

Design considerations and production pitfalls

mainandtestshareparse_knobs, but the default flavor differs:testdefaults toCurrentThread,maindefaults toThreaded 📎 tokio-macros/src/entry.rs:91-94. This explains why#[tokio::test]defaults to single-threaded—tests usually don't need multiple cores, and single-threaded is easier to reproduce.

An easily overlooked pitfall: after macro expansion, each function call creates a new Runtime. The documentation explicitly warns that if the function is called frequently, you should switch to Builder to reuse the Runtime📎 tokio-macros/src/lib.rs:31-35. Using#[tokio::main]on an ordinary function is legal, but each call pays the cost of constructing a Runtime.

Another pitfall iscraterenaming. When the useruse tokio as tokio1, thetokio::runtime::Buildergenerated by default inside the macro cannot find the path, and you must explicitlycrate = "tokio1" 📎 tokio-macros/src/lib.rs:239-264。parse_knobsincrate_paththe default value ofIdent::new("tokio", ...) 📎 tokio-macros/src/entry.rs:456-462, which is exactly the root cause of errors in renaming scenarios.

9.2 select!: multi-branch polling, bitmask, and random fairness

Intuitive model:select!is like a "waiter watching multiple pickup windows at the same time." Whichever window serves food first, he takes that portion, and the queues at the other windows are discarded. Without it, users would have to hand-writepoll_fnto put multiple Futures into a tuple and poll them one by one, and also handle the logic of "once one branch is ready, the other branches should be discarded."

Data structures and memory layout

select!expands to generate a local module__tokio_select_util, containing an enumOutand a type aliasMask 📎 tokio/src/macros/select.rs:615-619。Outwhose variant names are_0、_1... one per branch, plus aDisabledindicating that all branches are disabled📎 tokio-macros/src/select.rs:33-39。Mask. The underlying type is dynamically selected according to the number of branches: ≤8 usesu8, ≤16 usesu16, ≤32 usesu32, ≤64 usesu64, and more than 64 directly panics📎 tokio-macros/src/select.rs:17-31. This bitmask isselect!the core state of : bit i being 1 means branch i has been disabled.

All Futures are stored in a tuplefutures, and each element is first converted viaIntoFuture::into_future📎 tokio/src/macros/select.rs:654-656. Note that herefutures_initis constructed first and theninto_futureone by one, with comments explaining this is to take advantage of temporary lifetime extension📎 tokio/src/macros/select.rs:641-646. Thenlet mut futures = &mut futures;downgrades the tuple to a mutable reference, avoidingpoll_fnthe closure taking ownership📎 tokio/src/macros/select.rs:658-662。

Step-by-Step polling flow

Put it into context:select! { v = stream1.next() => ..., v = stream2.next() => ..., else => break }。

Step one, macro entry rule matching. If there is abiased;prefix,start=0 📎 tokio/src/macros/select.rs:801-803; otherwisestartis a random expressionthread_rng_n(BRANCHES) 📎 tokio/src/macros/select.rs:805-809. This is the source of fairness described in the documentation as "randomly selecting a branch to check first by default"📎 tokio/src/macros/select.rs:61-65。

Step two, normalization. The tt-muncher normalizes each branch into the form(skip) pat = fut, if cond => handler,skip, where_is a sequence of📎 tokio/src/macros/select.rs:770-793。skipwhose length equals the number of branches before that branchfutures_init.$($skip)*, used both to generate tuple field accesscount!and to

compute the branch index.if $cStep three, precondition evaluation. For each branch'sdisabled |= 1 << index 📎 tokio/src/macros/select.rs:631-636, if false, then$fut. Note: even if a branch is disabled, its📎 tokio/src/macros/select.rs:39-41。

expression is still evaluated, it just won't be polledpoll_fnStep four, enter theready!(poll_budget_available(cx))closure. First check the cooperative budget:Pending 📎 tokio/src/macros/select.rs:664-667, and if the budget is exhausted, return directlyselect!. This ensures

will not monopolize the worker.for i in 0..BRANCHES,branch = (start + i) % BRANCHES 📎 tokio/src/macros/select.rs:680-685Step five, loopdisabled & mask == mask. For each branch: first checkcontinue 📎 tokio/src/macros/select.rs:694-699, and if disabled thenPin::new_unchecked; otherwise take that Future out of the tuple and wrap it with📎 tokio/src/macros/select.rs:701-707(safety depends on the Future being stored on the stack and not moved)Ready(out); poll it,disabled |= maskthen first📎 tokio/src/macros/select.rs:710-730。

and then match the patternoutStep six, pattern matching. If$bindmatchesPoll::Ready(Out::_i(out)) 📎 tokio/src/macros/select.rs:727-733, returncontinue; if it does not match,📎 tokio/src/macros/select.rs:44-47。

continue polling the other branches—this is exactly what step 5 of the documentation means by "if the pattern does not match, disable the current branch"is_pendingStep seven, end of loop. IfPendingis true, returnOut::Disabled 📎 tokio/src/macros/select.rs:740-745, otherwise all branches are invalid and returnmatch output. The outerOut::_imapsDisabledto the corresponding handler,elsemaps to📎 tokio/src/macros/select.rs:749-755。

mermaid
flowchart TD
    start["poll_fn 闭包被调用"] --> budget{"poll_budget_available(cx)?"}
    budget -->|否| pending_budget["返回 Pending"]
    budget -->|是| init["is_pending = false; start = $start"]
    init --> loop{"i < BRANCHES?"}
    loop -->|否| check_pending{"is_pending?"}
    check_pending -->|是| pending["返回 Pending"]
    check_pending -->|否| disabled_out["返回 Out::Disabled"]
    loop -->|是| branch["branch = (start+i) % BRANCHES"]
    branch --> is_disabled{"disabled & mask == mask?"}
    is_disabled -->|是| next_i["i += 1"]
    is_disabled -->|否| poll_fut["Pin::new_unchecked(fut).poll(cx)"]
    poll_fut --> poll_res{"Poll 结果?"}
    poll_res -->|Pending| set_pending["is_pending = true; i += 1"]
    poll_res -->|Ready| disable["disabled |= mask"]
    disable --> pat_match{"out 匹配 $bind?"}
    pat_match -->|否| next_i
    pat_match -->|是| ready_out["返回 Out::_i(out)"]
    next_i --> loop
    set_pending --> loop

Copy

Design considerations and production pitfallsVec<bool>?Why use a bitmask instead ofdisabled |= mask? A bitmask is a single integer on the stack, with no heap allocation, andselect!is a single instruction. For

on the hot path, this avoids heap access on every iteration.Why should a branch be disabled when the pattern does not match?select!This is the key difference betweenSome(v) = stream.next() => ...and a "simple race." Considerstream.next(), ifNonereturns📎 tokio/src/macros/select.rs:198-223。

(end of stream), the pattern does not match, and that branch is permanently disabled, avoiding infinite polling of an already-finished stream. The documentation example relies on exactly this semantics to collect two streams until both end:select!The true meaning of cancellation safetyread_exact、read_to_end、write_allOnce a branch is ready, the Futures of the other branches are dropped. If a dropped Future has already consumed data but has not yet returned, the data is lost. The documentation explicitly lists📎 tokio/src/macros/select.rs:119-124as not cancellation safeMutex::lock、Semaphore::acquire, while📎 tokio/src/macros/select.rs:126-133, due to queue fairness, cancellation loses the queue position.await. How to determine: find the.awaitpoint, and if restarting the function at📎 tokio/src/macros/select.rs:135-139。

ifis still correct, then it is cancellation safeThe race-condition trap of preconditionsif !sleep.is_elapsed(): the documentation gives a classic erroneous example—usingsleepto guard theis_elapsed()branch, butwhilemay become true between theselect!check and📎 tokio/src/macros/select.rs:336-376, causing the timeout to be missedif. The correct approach is to removesleep, let thebreak 📎 tokio/src/macros/select.rs:378-405。

biased;branch always participate in polling, and after timeoutthe cost of📎 tokio/src/macros/select.rs:67-74: random RNG has a CPU cost, and some scenarios require a deterministic polling orderbiased;. But📎 tokio/src/macros/select.rs:75-81。

leaves the responsibility for fairness to the user: if one branch is always ready, later branches will starve

9.3 join! and the engineering constraints of macro expansion:join!Intuitive modelselect!is like "waiting for all deliveries to arrive at the same time." UnlikeReady, which cancels the rest as soon as one arrives first, it aggregates thepoll_fnvalues of all Futures into a tuple. Without it, users would have to hand-write

to maintain the completion state of each Future.

join!The expansion of is also based on storing Futures in a tuple, but the state is not a bitmask—it's a tuple of "completed values." After each Future completes, its value is taken out and stored in the result tuple, and the corresponding slot is marked as completed. Unlikeselect!,join!does not drop incomplete Futures—it must wait for all Futures to complete before returning.

Step-by-Step Flow

join!The polling logic of shares the skeleton of "tuple storing Futures +select!driving" withpoll_fn, but the semantics are opposite:select!is "return as soon as any is ready,"join!is "return only when all are ready." Each round of poll iterates over all incomplete Futures; if any returnsPending, the wholePending; if allReady, then aggregate and return.

mermaid
flowchart LR
    subgraph input["输入"]
        f1["Future A"]
        f2["Future B"]
        f3["Future C"]
    end
    subgraph poll["poll_fn 驱动"]
        tuple["元组 (A, B, C)"]
        state["完成状态元组"]
    end
    subgraph output["输出"]
        result["(A::Output, B::Output, C::Output)"]
    end
    f1 --> tuple
    f2 --> tuple
    f3 --> tuple
    tuple --> state
    state -->|"全部 Ready"| result
    state -->|"任一 Pending"| pending["返回 Pending"]

Design Reflections and Production Pitfalls

join!The cancellation safety semantics of differ fromselect!:join!When is dropped, all incomplete Futures are dropped, which can likewise lose data. But sincejoin!does not actively cancel any branch, it will not, likeselect!, "cancel this branch because another branch is ready." The real risk lies injoin!being cancelled as a whole by an outerselect!or timeout.

join!The difference between andtry_join!is worth noting:try_join!returns immediately when any Future returnsErr, cancelling the remaining Futures, so it inheritsselect!'s cancellation safety risk.

Design Reflections

The Boundary of Macros as Compile-Time Code Generators。#[tokio::main]places configuration validation at compile time; illegal combinations (such asmulti_thread + start_paused) fail to compile directly rather than panicking at runtime. This is the core advantage of macros over Builders: errors are caught earlier.

Hybrid Architecture of Declarative Macros + Procedural Macros。select!The main body of ismacro_rules!, but two key pieces of logic are delegated to procedural macros:select_priv_declare_output_enumgenerates theOutenum and theMasktype📎 tokio-macros/src/lib.rs:658-660,select_priv_clean_patternclears theref/mut 📎 tokio-macros/src/lib.rs:666-668in the pattern. Why? The comments explain: declarative macros struggle to generate code that "dynamically selects an integer type based on the number of branches," and also struggle to do token-level cleaning in pattern positions📎 tokio/src/macros/select.rs:577-579。

clean_patternThe Necessity of。select!matchesoutin the form of&outagainst the pattern📎 tokio/src/macros/select.rs:727; if the user writesref v, it becomes&ref vcausing a type error.clean_patternrecursively deletesby_ref、mutability, as well as theReferenceof themutability 📎 tokio-macros/src/select.rs:68-73📎 tokio-macros/src/select.rs:100-103pattern. This is the compromise macros make between "user intuition" and "the borrow checker."

The Engineering Reality of the 64-Branch Limit。count!、count_field!、select_variant!The three macros each hand-write matching rules from 0 to 64📎 tokio/src/macros/select.rs:821-1017📎 tokio/src/macros/select.rs:1021-1217📎 tokio/src/macros/select.rs:1221-1414. The comments bluntly say "I'm not happy about it either"📎 tokio/src/macros/select.rs:816-817. This is the cost of declarative macros being unable to do arithmetic: you can only hardcode a mapping from token count to integer.

Chapter Summary

Chapter Reflections and Self-Test

Q1: select!'sdisabledbitmask is reinitialized toselect!every timeDefault::default() 📎 tokio/src/macros/select.rs:627is entered. If this line is moved inside thepoll_fnclosure, what happens in the scenario of "calling select! in a loop and some branch's pattern does not match"?

Reference Analysis:disabledIf initialized inside the closure, it would be reset on every poll, causing branches that were disabled in the previous round due to pattern mismatch to participate in polling again. ConsiderSome(v) = stream.next() => ...andstreamhas ended (returningNone); after the pattern mismatch, this branch should have been permanently disabled. Ifdisabledis reset, the next round of poll would poll this ended stream again; if the stream is not fused (i.e., polling again after completion may panic or return undefined behavior), problems arise. Even if the stream is fused, it wastes CPU repeatedly polling a stream that always returnsNone. The documentation explicitly says "Re-entering select! due to a loop clears the disabled state"📎 tokio/src/macros/select.rs:37-38, referring to re-entering theselect!macro (a new loop iteration), not multiple polls within the sameselect!.disabledmust be initialized outside the closure to maintain state across multiple polls within the sameselect!call.

Q2: select!After polling toReady(out), first executesdisabled |= maskand then matches the pattern📎 tokio/src/macros/select.rs:720-730. Ifdisabled |= maskis removed, what happens in the scenario where the pattern does not match and the Future immediately returnsReadyon every poll?

Reference Analysis: After removingdisabled |= mask, ifoutdoes not match$bind, the code goes tocontinueand continues polling other branches. But when the next round ofpoll_fnis called (e.g., polling again after another branch returnsPending), this branch is still not disabled and will be polled again. If the Future immediately returnsReadyon every poll and the value does not match the pattern, a livelock forms: "poll -> Ready -> mismatch -> continue -> other branches Pending -> return Pending -> poll again -> Ready again -> ...", spinning the CPU.disabled |= masksets the flag immediately afterReady, ensuring that even if the pattern does not match, the branch will not be polled again. Note that the flag is set before pattern matching, so both "Ready but pattern mismatch" and "Ready and pattern match" disable the branch—the former prevents livelock, the latter prevents double consumption.

Q3: parse_knobsinsertsif false { let _: &dyn Future<Output = #output_type> = &body; }on non-test paths for type checking📎 tokio-macros/src/entry.rs:557-561, but skips the check for types that return!or containimpl Trait. Why does📎 tokio-macros/src/entry.rs:551-556need to be skipped? What happens if the check is forced?impl TraitReference Analysis

At the return position is an "opaque type"; the compiler does not allow coercing it to:impl Trait, because&dyn Future<Output = impl Trait>requires a concrete type, whereasdyn 要求具体类型,而 impl TraitThe concrete type of is not visible outside the function. If a check is forcibly inserted, errors such as "the size for values of typeimpl Futurecannot be known at compilation time" or "cannot be made into an object" will be reported. The same applies to the type returned by!:!can be coerced to any type, but&dyn Future<Output = !>'sOutput = !itself may trigger the unstable feature issue of the never type. The cost of skipping the check is: if the user writesasync fn main() -> impl Traitbut the actual return type does not matchimpl Trait, the error will only be exposed atblock_on, and the error message may be less clear than with an explicit check. This is the trade-off between "completeness of compile-time checking" and "limitations of the type system."

The macro takes boilerplate code and compile-time validation off the user's hands, but what it generates is still ordinary Futures andpollcalls. In the next chapter, we will leave the macro's compile-time world and enter the runtime I/O abstraction layer, to see howAsyncRead/AsyncWritesplits a byte stream into frames, and how theFramedcodec framework works correctly underselect!'s cancellation-safety constraints.

#[tokio::main]The essence of is "configuration parsing + Builder chain generation +block_onwrapping." Configuration validation is completed at compile time, and the flavor determines the builder's starting point and build method.select!The core of is "store Futures in a tuple + record disabled branches with a bitmask + preserve fairness with a random starting point." A pattern mismatch disables the branch, and cancellation safety depends on whether the dropped Future can be restarted at.await.join!andselect!share the same skeleton but have opposite semantics: the former waits for all to complete, while the latter returns as soon as any one is ready. Together, the three demonstrate the core trade-off in Tokio's macro design: hand boilerplate code and compile-time validation to the macro, and leave the complexity of runtime semantics (especially cancellation safety) for users to understand explicitly. After understanding how macros generate runtime code, the next natural question is: when this code actually starts reading and writing byte streams, what abstractions does Tokio provide? Chapter 10 will analyzeAsyncRead/AsyncWriteand the codec framework, looking at howBufReader/BufWriterreduces system calls, howcopy_bidirectionaldrives bidirectional forwarding, and howFramedsplits a byte stream into frames, thereby answering "where is the abstraction boundary of asynchronous I/O?"

CHAPTER 10

Chapter 10: Streaming I/O Abstractions: AsyncRead/AsyncWrite and the Codec Framework

Project: tokio-rs/tokio · Book progress: Chapter 10 / 14 · Verification status: FACT line numbers are truly anchored

The previous chapter dissected the expansion process of tokio-macros, and we saw how #[tokio::main], select!, and join! take boilerplate code and compile-time validation off the user's hands. But what the macros generate is still ordinary Futures and poll calls—when these Futures actually start reading and writing bytes, the underlying abstractions Tokio provides are only two traits: AsyncRead and AsyncWrite. Their problem is that they are "too low-level": a single poll_read only guarantees "some bytes were read," not "a complete message was read." And the vast majority of protocols (HTTP, Redis, gRPC, custom RPC) are oriented toward "frames" rather than "byte streams." The core question this chapter aims to answer is: where should the abstraction boundary of asynchronous I/O be drawn? Tokio's answer has two layers: tokio::io provides byte-stream-level traits and utilities (BufReader/BufWriter/copy_bidirectional), and tokio-util's codec framework builds on top of it to provide frame-level Stream/Sink adaptation (Framed/LengthDelimitedCodec). Understanding the division of labor between these two layers means understanding "why almost all protocol implementations start with Framed."

1. AsyncRead/AsyncWrite: Why std::io::Read cannot be reused directly

Intuitive model

std::io::Read::readis "blocking pickup": you stand at the window and wait as long as the goods have not arrived, and the thread is suspended.AsyncRead::poll_readis "pickup by meal ticket": you ask "is it ready?", and if it is not ready (Poll::Pending), you go do something else first, while leaving a Waker so the system can notify you when the goods arrive. Without this trait, all asynchronous I/O would have to manually implementepollregistration and Waker mapping—this is exactly what the Reactor in Chapter 5 does, andAsyncReadis the unified facade it exposes to upper layers.

Data structures and memory layout

AsyncRead's definition is extremely concise, with only one method:

rust
pub trait AsyncRead {
    fn poll_read(
        self: Pin<&mut Self>,
        cx: &mut Context<'_>,
        buf: &mut ReadBuf<'_>,
    ) -> Poll<io::Result<()>>;
}

📎 tokio/src/io/async_read.rs:44-60

Each of the three parameters has its own considerations.self: Pin<&mut Self>rather than&mut self: becauseAsyncReadis often held by the Future generated byasync fn, and once a Future is polled it cannot be moved (self-reference),Pinis a contract enforced by the compiler.cx: &mut Context<'_>carries the Waker and is the transmission channel for the "meal pager."buf: &mut ReadBuf<'_>is Tokio's wrapper around&mut [u8]—it simultaneously records "filled length" and "uninitialized capacity," avoidingstd::io::ReadThat ambiguity of "returning the number of bytes read but the buffer may be uninitialized."

The documentation explicitly lists three return semantics📎 tokio/src/io/async_read.rs:15-32:Ready(Ok(()))indicates data has been writtenbuf, and the read amount is determined byReadBuf::filledthe length increment of; if the increment is 0, it is either EOF orbuf.remaining() == 0(zero-capacity buffer);Pendingindicates it is currently not readable but a wakeup has been registered;Ready(Err(e))is an underlying I/O error. Here is an easily overlooked trap:"read amount is 0" does not equal EOF—if the caller passes in a zero-capacity buffer,poll_readwill immediately returnReady(Ok(()))but nothing was read. If the upper layer treats "0 bytes" as EOF, it will misjudge the connection as closed.

Scenario-driven Walkthrough: from&[u8]read a segment of bytes

Consider the simplest implementation—for&[u8]'sAsyncRead:

rust
impl AsyncRead for &[u8] {
    fn poll_read(
        mut self: Pin<&mut Self>,
        _cx: &mut Context<'_>,
        buf: &mut ReadBuf<'_>,
    ) -> Poll<io::Result<()>> {
        let amt = std::cmp::min(self.len(), buf.remaining());
        let (a, b) = self.split_at(amt);
        buf.put_slice(a);
        *self = b;
        Poll::Ready(Ok(()))
    }
}

📎 tokio/src/io/async_read.rs:98-108

Step-by-step analysis:self.len()is the length of the remaining unread slice,buf.remaining()is the remaining capacity of the target buffer, take the smaller of the twoamt。split_at(amt)split the slice into "theato copy this time" and "the remainingb」。buf.put_slice(a)to be read"aintoReadBufand advance its filled pointer.*self = bAdvance the slice itself to the remaining part—this is&[u8]the key to being a "cursor": after each pollselfpoints to the unread part. Finally returnReady(Ok(())), because a memory slice is always "ready" and will notPending。

Note that_cxis ignored: a memory data source does not need a Waker. This contrasts with a network socket—the latter returnsPendingand registers read interest when there is no data.

io::Cursor<T>'s implementation adds an extra layer of bounds checking📎 tokio/src/io/async_read.rs:113-134: first takeposition(), ifpos > slice.len()(position out of bounds) directly returnReady(Ok(()))without panicking📎 tokio/src/io/async_read.rs:113-134. This is defensive design:Cursor's position can be set to any value by an externalset_position, and when out of bounds, treating it as "already read to the end" is more consistent with I/O semantics than panicking.

Design thinking: deref macros and the propagation of Pin

AsyncReadprovidesBox<T>、&mut T、Pin<P>with forwarding implementations. The first two generatederef_async_read!through the📎 tokio/src/io/async_read.rs:64-70macro, the core beingPin::new(&mut **self).poll_read(cx, buf)—dereferencePin<&mut Box<T>>toPin<&mut T>and then forward.Pin<P>'s implementation is more subtle📎 tokio/src/io/async_read.rs:87-93: it callscrate::util::pin_as_deref_mut(self), projectingPin<&mut Pin<P>>asPin<&mut P::Target>. This layer of projection is necessary; otherwise nestedPinwould cause a type mismatch.

[Design inference and architectural trade-offs]

The design motivation here is "zero-cost abstraction": forwarding implementations letBox<dyn AsyncRead>、&mut Tand other wrapper types avoid manually writingpoll_read, while keepingPinsemantics correct. The cost is that each forwarding layer introduces an indirect call, which the compiler can usually eliminate by inlining.

---

II. copy_bidirectional: the state machine for bidirectional forwarding

Intuitive model

copy_bidirectionalis a "bidirectional food runner": it watches both A→B and B→A directions at the same time, and whenever either side reads data, it writes it to the opposite side. Without it, implementing a TCP proxy would require manually writing twocopyFutures and combining them withselect!—andselect!'s cancellation safety constraint (Chapter 9) would cause data "read halfway then canceled" to be lost.copy_bidirectionaluses an explicit state machine to preserve the intermediate states of "read-write-close," thereby achieving cancellation safety.

Data structure and memory layout

The core is a three-state enum:

rust
enum TransferState {
    Running(CopyBuffer),
    ShuttingDown(u64),
    Done(u64),
}

📎 tokio/src/io/util/copy_bidirectional.rs:10-14

RunningholdsCopyBuffer(containing an 8KB buffer and read/write counts), indicating "currently moving data."ShuttingDown(u64)carries the number of bytes copied, indicating "the read side has reached EOF and is closing the write side."Done(u64)indicates "shutdown complete, recording the final byte count." This enum is the key to cancellation safety:if dropped at any moment, the state is preserved in the enum, and the next poll can continue from the breakpoint。

CopyBuffercomes fromcopy.rs, and the default size is determined byDEFAULT_BUF_SIZE(8KB)📎 tokio/src/io/util/copy_bidirectional.rs:76-88. Each direction holds an independentCopyBuffer, so the memory overhead is 16KB.

Scenario-driven Walkthrough: the complete lifecycle of one bidirectional forwarding

copy_bidirectional_implusepoll_fnto combine the state machines of the two directions:

rust
let mut a_to_b = TransferState::Running(a_to_b_buffer);
let mut b_to_a = TransferState::Running(b_to_a_buffer);
poll_fn(|cx| {
    let a_to_b = transfer_one_direction(cx, &mut a_to_b, a, b)?;
    let b_to_a = transfer_one_direction(cx, &mut b_to_a, b, a)?;
    let a_to_b = ready!(a_to_b);
    let b_to_a = ready!(b_to_a);
    Poll::Ready(Ok((a_to_b, b_to_a)))
})
.await

📎 tokio/src/io/util/copy_bidirectional.rs:127-151

Note the call order oftransfer_one_direction: first advance a→b, then advance b→a, both returningPoll。ready!The macro returns immediately when either direction is incompletePending—butthe state of the other direction has already been advanced. This is exactly what the comment emphasizes📎 tokio/src/io/util/copy_bidirectional.rs:143-144: even ifready!returns early, the other direction will still returnDone(count)on the next poll, without losing progress.

transfer_one_directionInternally is aloop, advancing by state:

rust
loop {
    match state {
        TransferState::Running(buf) => {
            let count = ready!(buf.poll_copy(cx, r.as_mut(), w.as_mut()))?;
            *state = TransferState::ShuttingDown(count);
        }
        TransferState::ShuttingDown(count) => {
            ready!(w.as_mut().poll_shutdown(cx))?;
            *state = TransferState::Done(*count);
        }
        TransferState::Done(count) => return Poll::Ready(Ok(*count)),
    }
}

📎 tokio/src/io/util/copy_bidirectional.rs:29-42

RunningIn the state, callpoll_copy, which internally loops "read a chunk, write a chunk" until the read side reaches EOF or the write side blocks. On EOF, return the total number copied, and the state transitions toShuttingDown。ShuttingDowncallpoll_shutdownto close the write side (send FIN), and after completion transition toDone。Donedirectly return the count.

The flowchart below shows the advancement logic and error branches of the unidirectional state machine:

mermaid
flowchart TD
    start["transfer_one_direction 进入 loop"] --> match_state{"当前 TransferState?"}
    match_state -->|Running| poll_copy["buf.poll_copy(cx, r, w)"]
    poll_copy --> copy_ready{"poll_copy 结果?"}
    copy_ready -->|Pending| ret_pending["返回 Poll::Pending<br/>状态保持 Running"]
    copy_ready -->|Err| ret_err["返回 Poll::Ready(Err)<br/>错误向上传播"]
    copy_ready -->|Ok(count)| to_shutdown["state = ShuttingDown(count)"]
    to_shutdown --> match_state
    match_state -->|ShuttingDown| poll_shutdown["w.poll_shutdown(cx)"]
    poll_shutdown --> shutdown_ready{"shutdown 结果?"}
    shutdown_ready -->|Pending| ret_pending2["返回 Poll::Pending<br/>状态保持 ShuttingDown"]
    shutdown_ready -->|Err| ret_err
    shutdown_ready -->|Ok| to_done["state = Done(count)"]
    to_done --> match_state
    match_state -->|Done| ret_done["返回 Poll::Ready(Ok(count))"]

Design thinking: why use an explicit state machine instead of async fn

[Design inference and architectural trade-offs]

Iftransfer_one_directionwere written asasync fn, the compiler would generate a Future whose internal state (CopyBuffer, copied count) is hidden inside the generated state machine. This is fine when used unidirectionally, butcopy_bidirectionalneeds towithin the same poll cycleadvance both directions simultaneously—if twoasync fnplusselect!were used, when either direction completes the other would be dropped, and its internal buffer and count would be lost, violating cancellation safety. An explicitTransferStateexposes the state on the stack,poll_fnand the state is still there each time it is re-entered, thus guaranteeing "recovery from the breakpoint after cancellation."

In error handling,poll_copyreturnsErrwhich will be immediately propagated upward through?. The documentation explicitly states📎 tokio/src/io/util/copy_bidirectional.rs:32: interrupted reads and writes will be retried, other errors are returned immediately, and📎 tokio/src/io/util/copy_bidirectional.rs:67-70partially read data may be lost(not written to the opposite side). This is a point to note in production environments:does not guarantee "either all succeed or all fail"; when an error occurs, half the data may already be in flight.copy_bidirectionaladditionally performs a zero-size assertion

copy_bidirectional_with_sizes, because a zero-capacity buffer would cause📎 tokio/src/io/util/copy_bidirectional.rs:99-125,因为零容量缓冲区会导致 poll_copyalways returnsReady(Ok(0))is misjudged as EOF, forming a busy loop.

---

III. Framed: Slicing a Byte Stream into Frames

Intuitive Model

Framedis a "sausage machine": the upstream is a continuous stream of water (AsyncRead/AsyncWrite), and the downstream is the cut sausage segments (Stream<Item = Frame> / Sink<Frame>)。Decoderis responsible for "cutting out a segment from the stream,"Encoderis responsible for "wrapping a segment into a stream." WithoutFramed, every protocol implementation would have to hand-write "buffer management + partial packet handling + sticky packet splitting"—this is exactly the repetitive labor that the codec framework aims to eliminate.

Data Structures and Memory Layout

Frameditself is just a thin wrapper:

rust
pub struct Framed<T, U> {
    #[pin]
    inner: FramedImpl<T, U, RWFrames>
}

📎 tokio-util/src/codec/framed.rs:38-41

The real state is inFramedImpl'sstate: RWFrames, containingread: ReadFrameandwrite: WriteFrametwo parts.ReadFrame's fields are visible inwith_capacity📎 tokio-util/src/codec/framed.rs:107-126:eof: bool(whether the read side is EOF),is_readable: bool(whether readable interest has been registered),buffer: BytesMut(read buffer),has_errored: bool(whether an error has occurred, to prevent repeated reads).WriteFrameFields📎 tokio-util/src/codec/framed.rs:119-122:buffer: BytesMut(write buffer),backpressure_boundary: usize(backpressure threshold).

backpressure_boundaryis the key to the backpressure mechanism: when the write buffer exceeds this threshold,poll_readywill returnPendinguntil the data is flushed, thereby applying backpressure to the upstreamSink. By default it equalscapacity 📎 tokio-util/src/codec/framed.rs:121, and can be adjusted viaset_backpressure_boundary📎 tokio-util/src/codec/framed.rs:271-273。

Scenario-Driven Walkthrough: Reading a Frame from a Socket

Framed'sStreamimplementation simply forwards toFramedImpl::poll_next 📎 tokio-util/src/codec/framed.rs:309-311. The real logic is inFramedImpl(this file is not provided in this chapter, but the call chain can be inferred fromFramed's interface):

1. poll_nextfirst checksread.bufferto see whether a complete frame already exists (by callingcodec.decode);

2. IfdecodereturnsSome(frame), produce it directly without touching the underlying I/O;

3. If it returnsNone(partial packet), checkread.eof: if already EOF and the buffer is non-empty, it means there is residual data that cannot be decoded, so return an error orNone;

4. Otherwise call the underlyingAsyncRead::poll_readto read more bytes intoread.buffer;

5. The bytes read are passed todecodeagain, looping until a frame is produced orPending。

This order of "decode first, then read" is important: it guarantees thata single read may produce multiple frames(sticky packets), anda frame may span multiple reads(partial packets).is_readableThe flag avoids repeatedly registering readable interest—if the previous poll already registered it and it was not ready, this time it directly returnsPendingwithout repeatedly calling the underlying layer.

Sink's implementation call chain📎 tokio-util/src/codec/framed.rs:315-338:start_sendcallscodec.encode(item, &mut write.buffer)to encode the frame into the write buffer;poll_flushflusheswrite.bufferto the underlyingAsyncWrite;poll_readycheckswrite.buffer.len() >= backpressure_boundary, and if it exceeds the threshold, flushes first and then returns ready.

The sequence diagram below showsFramed's cross-component collaboration in one "read-frame-write-frame" round trip:

mermaid
sequenceDiagram
    participant App as 应用层
    participant F as FramedImpl
    participant C as Decoder/Encoder
    participant IO as AsyncRead/AsyncWrite

    App->>F: poll_next(cx)
    F->>C: decode(&mut read.buffer)
    alt 缓冲中已有完整帧
        C-->>F: Some(frame)
        F-->>App: Poll::Ready(Some(frame))
    else 半包
        C-->>F: None
        F->>IO: poll_read(cx, &mut read.buffer)
        alt 数据就绪
            IO-->>F: Ready(Ok(()))
            F->>C: decode(&mut read.buffer)
            C-->>F: Some(frame) 或 None
        else 无数据
            IO-->>F: Pending
            F-->>App: Poll::Pending
        end
    end

    App->>F: start_send(frame)
    F->>C: encode(frame, &mut write.buffer)
    C-->>F: Ok(())
    App->>F: poll_flush(cx)
    F->>IO: poll_write(cx, &write.buffer)
    IO-->>F: Ready(Ok(n))
    F->>IO: poll_flush(cx)
    IO-->>F: Ready(Ok(()))

Cancellation Safety: Framed's Documentation Warning

Framed's documentation specifically lists cancellation safety semantics📎 tokio-util/src/codec/framed.rs:23-30:SinkExt::sendIf inselect!it is preempted by another branch and completes,the message is guaranteed not to have been sent, but the message itself is lost—becausesendinternally firstpoll_readythenstart_send, and if during thepoll_readystage it is dropped,itemhas already been consumed but not encoded. WhileStreamExt::nextis cancellation-safe: it only holds a reference to the underlying stream, and dropping it will not lose already-decoded frames.

[Design Inference and Architectural Trade-offs]

This asymmetry stems from the difference between the read and write paths: the state of the read path (read.buffer) is stored insideFramed, and droppingnextmerely abandons the action of "taking a frame," leaving the buffer unaffected; the state of the write path (the pendingitem) is onsend's Future stack, and dropping it means losing it. In production code, ifselect!is used insend, you must ensure the message can be resent or accept the loss.

Design Thinking:into_partsandmap_codec

Framedprovideinto_parts/from_partsfor "swapping the codec while preserving the buffer"📎 tokio-util/src/codec/framed.rs:290-298 📎 tokio-util/src/codec/framed.rs:155-166。map_codecis implemented based on this pair of methods📎 tokio-util/src/codec/framed.rs:221-234: firstinto_partssplit outio/codec/read_buf/write_buf, then use themapfunction to convert the codec, and finallyfrom_partsreassemble. This design allows buffered data to be preserved during protocol upgrades (such as switching from plaintext to TLS), avoiding re-reading.

FramedParts's_priv: ()field📎 tokio-util/src/codec/framed.rs:373-375is the "non-exhaustive struct" trick: private fields prevent external direct construction, forcing use ofnew/from_parts, thereby allowing fields to be added in the future without breaking compatibility.

---

IV. LengthDelimitedCodec: The State Machine of Length-Prefixed Encoding and Decoding

Intuitive Model

LengthDelimitedCodecis a specialized tool for "cutting sausages by length": it assumes that each frame is preceded by a fixed-byte-length length field, reading the length first and then the payload. Without it, implementing a length-prefixed protocol would require hand-writing the state machine "read 4 bytes → parse length → read N bytes → loop"—this is exactly what its internalDecodeStatedoes.

Data Structures and Memory Layout

rust
pub struct LengthDelimitedCodec {
    builder: Builder,
    state: DecodeState,
}

enum DecodeState {
    Head,
    Data(usize),
}

📎 tokio-util/src/codec/length_delimited.rs:451-457

DecodeStateis an explicit state machine:Headmeans "currently reading the length field,"Data(n)means "length n has been parsed, currently reading the payload." This state persists acrossdecodecalls, thereforeprogress is not lost in partial packet scenarios。

Builderholds all configuration📎 tokio-util/src/codec/length_delimited.rs:416-435:max_frame_len(default 8MB),length_field_len(default 4 bytes),length_field_offset(default 0),length_adjustment(default 0),num_skip(defaultNone, i.e.offset + len)、length_field_is_big_endian(default true).

Scenario-Driven Walkthrough: Decoding a Length-Prefixed Frame

decodeis the state machine entry point:

rust
fn decode(&mut self, src: &mut BytesMut) -> io::Result<Option<BytesMut>> {
    let n = match self.state {
        DecodeState::Head => match self.decode_head(src)? {
            Some(n) => {
                self.state = DecodeState::Data(n);
                n
            }
            None => return Ok(None),
        },
        DecodeState::Data(n) => n,
    };

    match self.decode_data(n, src) {
        Some(data) => {
            self.state = DecodeState::Head;
            src.reserve(self.builder.num_head_bytes().saturating_sub(src.len()));
            Ok(Some(data))
        }
        None => Ok(None),
    }
}

📎 tokio-util/src/codec/length_delimited.rs:579-603

HeadIn thedecode_headstate, callNone. If it returnsOk(None)(insufficient data), directly returnSome(n)to wait for more data; if it returnsData(n)。Data, the state transitions todecode_data(n, src)In thesplit_to(n)state, directly take n. Then callHead: if the buffer already has n bytes,Nonecuts out the frame, the state returns to

decode_head, and reserves space for the next frame header; otherwise return

rust
let head_len = self.builder.num_head_bytes();
let field_len = self.builder.length_field_len;

if src.len() < head_len {
    return Ok(None);
}

let n = {
    let mut src = Cursor::new(&mut *src);
    src.advance(self.builder.length_field_offset);
    let n = if self.builder.length_field_is_big_endian {
        src.get_uint(field_len)
    } else {
        src.get_uint_le(field_len)
    };

    if n > self.builder.max_frame_len as u64 {
        return Err(io::Error::new(
            io::ErrorKind::InvalidData,
            LengthDelimitedCodecError { _priv: () },
        ));
    }

    let n = n as usize;
    let n = if self.builder.length_adjustment < 0 {
        n.checked_sub(-self.builder.length_adjustment as usize)
    } else {
        n.checked_add(self.builder.length_adjustment as usize)
    };

    match n {
        Some(n) => n,
        None => {
            return Err(io::Error::new(
                io::ErrorKind::InvalidInput,
                "provided length would overflow after adjustment",
            ));
        }
    }
};

src.advance(self.builder.get_num_skip());
src.reserve(n.saturating_sub(src.len()));
Ok(Some(n))

📎 tokio-util/src/codec/length_delimited.rs:504-562

is the core parsing logic:src.len() >= head_lenCopyNone 📎 tokio-util/src/codec/length_delimited.rs:499-502Parse step by step: first checkCursor, and if insufficient, returnsrc. Useadvance/get_uintto wrapadvance(length_field_offset)so that📎 tokio-util/src/codec/length_delimited.rs:517operations can be performed without consuming the original buffer.field_lenSkip the header prefix📎 tokio-util/src/codec/length_delimited.rs:520-524。

. Read the length value ofbytes according to endiannessn > max_frame_lenKey defenseInvalidData: if📎 tokio-util/src/codec/length_delimited.rs:526-531, immediately return

errorchecked_sub/checked_add. This prevents a malicious peer from sending a frame with a "length field of 4GB" and causing memory exhaustion—this is the most classic DoS attack surface of length-prefixed protocols.📎 tokio-util/src/codec/length_delimited.rs:537-541Length adjustment usesInvalidInputerrors instead of panics.get_num_skip()Returnnum_skipor the defaultoffset + len 📎 tokio-util/src/codec/length_delimited.rs:1070-1073, skipping the remaining part of the header. Finally,reserve(n.saturating_sub(src.len()))reserve payload space📎 tokio-util/src/codec/length_delimited.rs:559— usingsaturating_subis becausesrcmay already contain part of the payload.

The flowchart below showsdecodethe complete decision path of:

mermaid
flowchart TD
    entry["decode(src)"] --> check_state{"self.state?"}
    check_state -->|Head| head["decode_head(src)"]
    head --> head_result{"结果?"}
    head_result -->|Ok(None)| ret_none1["返回 Ok(None)<br/>等待更多数据"]
    head_result -->|Err| ret_err1["返回 Err<br/>长度超限或溢出"]
    head_result -->|Ok(Some(n))| set_data["state = Data(n)"]
    set_data --> decode_data
    check_state -->|Data(n)| decode_data["decode_data(n, src)"]
    decode_data --> data_result{"src.len() >= n?"}
    data_result -->|否| ret_none2["返回 Ok(None)<br/>等待更多数据"]
    data_result -->|是| split["src.split_to(n)<br/>state = Head<br/>reserve 下一帧头部"]
    split --> ret_frame["返回 Ok(Some(frame))"]

Design consideration: max_frame_len clipping and overflow protection

Builder::adjust_max_frame_lenWhen constructing the codec, clipmax_frame_lento the maximum value representable by the length field📎 tokio-util/src/codec/length_delimited.rs:1075-1081。max_allowed_frame_lencomputemax_length_field_value + length_adjustment 📎 tokio-util/src/codec/length_delimited.rs:1083-1089, wheremax_length_field_valueuseschecked_shlto handlelength_field_len == 8shift overflow when📎 tokio-util/src/codec/length_delimited.rs:1091-1096. This clipping prevents contradictory configurations such as "length field is 2 bytes but max_frame_len is set to 1MB"—2 bytes can represent at most 65535, so after clipping max_frame_len becomes 65535.

Symmetric protection on the encoding path:encodecheckn > max_frame_lenreturnInvalidInput 📎 tokio-util/src/codec/length_delimited.rs:607-607, and length adjustment likewise useschecked_add/checked_sub 📎 tokio-util/src/codec/length_delimited.rs:620-631. Note that the adjustment direction during encoding is opposite to decoding: decoding is "read length ± adjustment = payload length", while encoding is "payload length ∓ adjustment = written length field"📎 tokio-util/src/codec/length_delimited.rs:620-624。

[Design inference and architectural trade-offs]

This symmetric design of "add when decoding, subtract when encoding" is intended to makelength_adjustmentsemantically unified: it represents "the difference between the length field value and the payload length". When the protocol's length field includes the header (such as Example 3),adjustment = -2, during decodingn - (-2) = n + 2yields the payload length, and during encodingpayload - (-2) = payload + 2writes back the length field.

---

Design consideration: the three levels of abstraction boundaries

Reviewing this chapter, Tokio's I/O abstraction presents a clear three-layer structure:

First layer: byte stream trait (AsyncRead/AsyncWrite). It only promises "read/write some bytes", not frame boundaries. This is the minimal interface, and any I/O source (socket, file, in-memory slice) can implement it. The cost is that upper layers must handle partial packets/sticky packets themselves.

Second layer: byte stream utilities (BufReader/BufWriter/copy_bidirectional). On top of the trait, they provide general capabilities such as "reducing system calls" and "bidirectional forwarding".copy_bidirectionalThe explicit state machine of

demonstrates how "cancel safety" is implemented at the utility layer—state is kept on the stack rather than inside the Future.Framed/Decoder/Encoder)Third layer: frame adaptation (Stream<Frame>/Sink<Frame>. It elevates byte streams toLengthDelimitedCodec, so protocol implementations only need to care about "frame encoding/decoding" rather than "buffer management".DecodeStateis the standard example of this layer, and itsmax_frame_lenstate machine and

protection are patterns that all length-prefixed protocols should reuse.

[Design inference and architectural trade-offs]tokio-utilThe division of these three layers is not accidental: it corresponds to the three gradients of "abstraction leakage". The lower the layer, the more general but harder to use; the higher the layer, the easier to use but more specialized. Tokio chooses to place "frames" as a first-class citizen intokiorather thantokiocore, because the definition of a frame varies by protocol—tokio-utilonly provides byte streams,Decoder/Encoder。

---

provides the frame framework, and specific protocols (HTTP/Redis/gRPC) implement

  • AsyncRead::poll_readin their respective crates.Pin<&mut Self> + Context + ReadBufChapter summarystd::io::Read::readusesReady(Ok(()))three parameters to replace
  • copy_bidirectional, turning "blocking wait" into "register Waker + return Pending".TransferStateAnd when the read amount is 0, it is necessary to distinguish between EOF and a zero-capacity buffer.Running/ShuttingDown/Doneusesselect!three-state enum (
  • Framed) to save intermediate state, so that bidirectional forwarding can still recover underAsyncRead/AsyncWritecancellation. When an error occurs, some data may be lost.Stream/Sink,ReadFrame/WriteFrameadaptsSinkExt::sendtoStreamExt::nextto separately manage read/write buffers and backpressure.
  • LengthDelimitedCodecis not cancel-safe (message loss),DecodeState(Head/Data(n)is cancel-safe.max_frame_lenuseschecked_add/checked_sub) state machine to handle partial packets,

protects against length-field DoS,

Q1: copy_bidirectionalprotects against adjustment overflow.transfer_one_directionChapter review questions and self-testTransferState::ShuttingDownInready!(w.as_mut().poll_shutdown(cx))?of*state = TransferState::Done(*count), if the

branch's:poll_shutdownis changed to directlyDone(skipping shutdown), in what scenarios would the peer connection fail to close normally?readReference analysisShuttingDownThe purpose of📎 tokio/src/io/util/copy_bidirectional.rs:35-39is to send a FIN packet to the peer, notifying it that "I have no more data on my side". If it is skipped andpoll_shutdownis directly transitioned to, the write side will not be closed, and the peer will keep waiting for data, forming a "half-open connection"—the peer may block forever onPendinguntil timeout. In a TCP proxy scenario, this causes connection leaks: the client has disconnected, but the proxy's connection to the backend remains open. The existence of theready!state in the source code

Q2: LengthDelimitedCodec::decode_headis precisely to ensure that the write side is explicitly closed after EOF. Note thatif n > self.builder.max_frame_len as u64itself may return📎 tokio-util/src/codec/length_delimited.rs:526-531(such as when the send buffer is full), so0xFFFFFFFFmust be used to wait rather than ignore it.length_adjustmentIn

, if the checkofnis removedusize, what consequences would be triggered if a malicious client sends a frame header with a length field ofdecode_data。decode_data(4GB)? Why must this check be beforesrc.len() < n?NoneReference analysisdecode_head: After removing the check,src.reserve(n.saturating_sub(src.len())) 📎 tokio-util/src/codec/length_delimited.rs:559will be converted tolength_adjustmentand passed tolength_adjustment. When checking-2, it returns0xFFFFFFFF - 2, but thechecked_subat the end of

At this point, we have clarified Tokio's two layers of abstraction between byte streams and message frames: tokio::io handles byte movement, while the codec framework in tokio-util handles frame splitting and encoding/decoding. The reason Framed becomes the starting point for protocol implementations is precisely that it encapsulates the high-frequency need to "read one complete message" into a reusable Stream/Sink adapter. But frames are only containers for data. When a protocol needs to handle dynamic task sets, structured cancellation, or more complex streaming composition, Framed alone is not enough. The next chapter will move into the extension mechanisms of tokio-stream and tokio-util, looking at how StreamExt combinators, StreamMap/JoinSet/TaskTracker, and CancellationToken reuse the underlying Waker and scheduling mechanisms to provide higher-level tools for asynchronous iteration and task management.

CHAPTER 11

Chapter 11: The Stream Ecosystem and Tooling Layer: Extension Mechanisms of tokio-stream and tokio-util

Project: tokio-rs/tokio · Book progress: Chapter 11 / 14 · Verification status: FACT line numbers are truly anchored

In the previous chapter, we broke down the byte-level mechanics of Framed: Decoder splits BytesMut into frames, Sink writes frames back, and the abstraction boundary of asynchronous I/O becomes clear. But frames are only containers for data, and real protocol implementations immediately encounter three problems that neither tokio::io nor Framed solves: asynchronous iteration—Framed implements Stream, but Stream only has poll_next, not next().await, filter, take, or merge, and hand-writing poll_fn is both verbose and prone to pitfalls in cancellation safety; dynamic task sets—a chat service needs to subscribe to N channels simultaneously, with channels joining and leaving at any time, while the number of branches in select! is fixed at compile time and cannot express a stream set that changes at runtime; structured cancellation—select! can cancel a single branch, but it cannot propagate the shutdown of the entire task tree, nor can it wait for all tasks to actually exit. tokio-stream and tokio-util were born precisely for these three things, and their key design principle is not to start from scratch: every combinator in StreamExt is just a wrapper around poll_next, StreamMap reuses the registration semantics of Waker, CancellationToken is built directly on top of tokio::sync::Notify, and TaskTracker encodes all state with an AtomicUsize. Understanding them is essentially understanding how to build zero-cost abstractions on top of the existing Waker and scheduling mechanisms. This chapter progresses through three layers: iteration, collections, and cancellation: first we look at how StreamExt turns poll_next into a composable iterator, then at how StreamMap and TaskTracker manage dynamic collections, and finally at how CancellationToken uses a tree to propagate cancellation signals to the entire task tree.

StreamExt: Turning poll_next into a composable iterator

Intuitive model

Streamis toFutureasIteratoris to values:Futureproduces "one value,"Streamproduces "a sequence of values." ButStreamdefines onlypoll_nextthis one primitive, just asIteratordefines onlynext. WithoutStreamExt, every filter, map, and truncate operation would require hand-writing apoll_fnclosure and manually managingPin—this was exactly the most painful part for early users of thefuturescrate.StreamExtThe role ofStreamis to giveIteratora combinator ecosystem like that of

. Without it, the disaster the system faces is not missing functionality, buta systemic collapse of cancellation safety: every hand-writtenpoll_fnmay lose an element that has already beenselect!when it is cancelled bypoll.

Data structures and memory layout

StreamExtis anextension trait, and it holds no data itself:

📎 tokio-stream/src/stream_ext.rs:106-106

rust
pub trait StreamExt: Stream {

All of its methods return aconcrete combinator struct, notBox<dyn Stream>. This is the key design:mapreturnsMap<Self, F>,filterreturnsFilter<Self, F>,takereturnsTake<Self>. These structs are all zero-heap-allocation generic wrappers, and the compiler can inline the entire chain into layers ofpoll_nextcalls.

Note the blanket impl of the trait:

📎 tokio-stream/src/stream_ext.rs:1213-1213

rust
impl<St: ?Sized> StreamExt for St where St: Stream {}

AnyStreamautomatically gains all combinators, with no manual implementation required.?Sizedallowsdyn Streamto also enjoy extension methods.

The module declarations of the combinators reveal the full capability surface of this trait:

📎 tokio-stream/src/stream_ext.rs:4-59

rust
mod all; use all::AllFuture;
mod any; use any::AnyFuture;
mod chain; pub use chain::Chain;
pub(crate) mod collect; use collect::{Collect, FromStream};
mod filter; pub use filter::Filter;
mod filter_map; pub use filter_map::FilterMap;
mod fold; use fold::FoldFuture;
mod fuse; pub use fuse::Fuse;
mod map; pub use map::Map;
mod map_while; pub use map_while::MapWhile;
mod merge; pub use merge::Merge;
mod next; use next::Next;
mod skip; pub use skip::Skip;
mod skip_while; pub use skip_while::SkipWhile;
mod take; pub use take::Take;
mod take_while; pub use take_while::TakeWhile;
mod then; pub use then::Then;
mod try_next; use try_next::TryNext;
mod peekable; pub use peekable::Peekable;

There is a noteworthy distinction here:next、try_next、all、any、fold、collectreturnsFuture(Next、TryNext、AllFuture...), because they consume the entire stream into a single value; whilemap、filter、takeand others returnStream, because they preserve the shape of the stream.nextThe return type ofNext<'_, Self>is

📎 tokio-stream/src/stream_ext.rs:144-149

rust
fn next(&mut self) -> Next<'_, Self>
where
    Self: Unpin,
{
    Next::new(self)
}

Self: UnpinCopynextThePinconstraint is deliberate:!Unpindoes not take ownership of the stream, only borrows it, and therefore cannotBox::pinthe stream. If the stream ispin_mut!, the user must first

📎 tokio-stream/src/stream_ext.rs:116-121

rust
/// Note that because `next` doesn't take ownership over the stream,
/// the [`Stream`] type must be [`Unpin`]. If you want to use `next` with
/// a [`!Unpin`](Unpin) stream, you'll first have to pin the stream. This can
/// be done by boxing the stream using [`Box::pin`] or
/// pinning it to the stack using the `pin_mut!` macro from the `pin_utils`
/// crate.

. The documentation explicitly points out this tradeoff:mergepolling of

mergeis the best example for understanding how combinators reuse Waker. It interleaves the outputs of two streams, andguarantees fairness— if both streams are ready simultaneously, it alternates outputs. The documentation specifically warns against chaining calls tomerge:

📎 tokio-stream/src/stream_ext.rs:319-321

rust
/// simultaneously, the merge stream alternates between them. This provides
/// some level of fairness. You should not chain calls to `merge`, as this
/// will break the fairness of the merging.

mergerequires that both streams have the sameItemtype:

📎 tokio-stream/src/stream_ext.rs:398-404

rust
fn merge<U>(self, other: U) -> Merge<Self, U>
where
    U: Stream<Item = Self::Item>,
    Self: Sized,
{
    Merge::new(self, other)
}

When the caller.next().awaitthe execution flow is as follows:

1. Next::pollcallsMerge::poll_next。

2. Mergeinternally maintains a boolean flag for "whose turn it was last time." It firstpollthe stream that did not produce last time; ifPendingthenpollthe other one.

3. If bothPending,MergereturnPendingbutboth streams' respective Wakers have been registered— either becoming ready will wake the current task.

4. If one stream returnsReady(None)(finished),Mergerecords that the stream has ended, and thereafter onlypollthe other stream until it also ends.

The key here is:Mergehas no Waker management logic of its own; it passescxas-is to the internal two streams'poll_next。Waker registration is entirely handled by the underlying streams,Mergeonly decides "whom to ask first this time." This is the literal meaning of "reusing the underlying Waker mechanism."

merge_size_hintsThe helper function demonstrates how combinators merge capacity hints:

📎 tokio-stream/src/stream_ext.rs:1216-1226

rust
fn merge_size_hints(
    (left_low, left_high): (usize, Option<usize>),
    (right_low, right_high): (usize, Option<usize>),
) -> (usize, Option<usize>) {
    let low = left_low.saturating_add(right_low);
    let high = match (left_high, right_high) {
        (Some(h1), Some(h2)) => h1.checked_add(h2),
        _ => None,
    };
    (low, high)
}

Note the choice ofsaturating_addandchecked_add: the lower bound uses saturating addition (better to underestimate than to overflow and panic), and the upper bound uses checked addition (if either is unknown, the whole is unknown). This is the typical way of handling thesize_hintcontract.

Design considerations: cancellation safety andchunks_timeout's panic protection

StreamExt's documentation annotates each method withCancel safety. Takingnextas an example:

📎 tokio-stream/src/stream_ext.rs:123-127

rust
/// # Cancel safety
///
/// This method is cancel safe. The returned future only
/// holds onto a reference to the underlying stream,
/// so dropping it will never lose a value.

nextis cancellation-safe because it only borrows the stream and does not consume elements —Nextwhen the future is dropped, the stream's own state is unchanged, and the nextnextwill re-poll。

But not all combinators are cancellation-safe.chunks_timeoutperforms parameter validation at construction time:

📎 tokio-stream/src/stream_ext.rs:1178-1185

rust
#[track_caller]
fn chunks_timeout(self, max_size: usize, duration: Duration) -> ChunksTimeout<Self>
where
    Self: Sized,
{
    assert!(max_size > 0, "`max_size` must be non-zero.");
    ChunksTimeout::new(self, max_size, duration)
}
[Design inference and architectural trade-offs]

#[track_caller]makes the panic location point to the caller rather than the library internals,assert!rejecting at construction timemax_size == 0. Why must it be checked at construction time? Ifmax_size == 0,ChunksTimeout's batching logic would fall into an infinite loop of "never accumulating a full batch" or produce empty batches, and such bugs are extremely difficult to locate at runtime. A construction-time panic moves the error to the earliest observable point.

timeoutThe difference betweentimeout_repeatingandtimeoutis also worth noting:returns an error after the timeout, but;timeout_repeatingcontinues polling the inner streamIntervalinstead, per

📎 tokio-stream/src/stream_ext.rs:985-1001

rust
/// Once a timeout error is received, no further events will be received
/// unless the wrapped stream yields a value (timeouts do not repeat).

📎 tokio-stream/src/stream_ext.rs:1071-1072

rust
/// Timeout errors will be continuously produced at the specified interval
/// until the wrapped stream yields a value.

---

Copy

StreamMap: Dynamic stream collections and fair polling

select!Intuitive modelStreamMap's number of branches is fixed at compile time. But the number of channels a chat service needs to subscribe to, or the number of connections a crawler needs to track, are only known at runtime.select!is a "runtime-addable/removablenext": it puts any number of streams into a collection, and each(key, value)returnsmpsc, telling you which stream the value came from. Without it, you could only stuff all streams into a single

channel, adding an extra layer of forwarding overhead.

StreamMapData structure and memory layoutVec:

📎 tokio-stream/src/stream_map.rs:204-208

rust
#[derive(Debug)]
pub struct StreamMap<K, V> {
    /// Streams stored in the map
    entries: Vec<(K, V)>,
}

Copy

📎 tokio-stream/src/stream_map.rs:38-44

rust
/// `StreamMap` is backed by a `Vec<(K, V)>`. There is no guarantee that this
/// internal implementation detail will persist in future versions, but it is
/// important to know the runtime implications. In general, `StreamMap` works
/// best with a "smallish" number of streams as all entries are scanned on
/// insert, remove, and polling. In cases where a large number of streams need
/// to be merged, it may be advisable to use tasks sending values on a shared
/// [`mpsc`] channel.
Copy

[Design inference and architectural trade-offs]HashMapWhy not useStreamMap? Because's core operation ispolling all streamsVec, not lookup by key.swap_remove's linear scan is CPU-cache-friendly, andHashMapis O(1). Ifpoll_nextwere used, eachinsertwould have to traverse hash buckets, with worse cache locality.removeand

insert's O(n) scan is acceptable under the "small-scale stream collection" assumption.

📎 tokio-stream/src/stream_map.rs:446-454

rust
pub fn insert(&mut self, k: K, stream: V) -> Option<V>
where
    K: Hash + Eq,
{
    let ret = self.remove(&k);
    self.entries.push((k, stream));

    ret
}

removeCopyswap_removeuses

📎 tokio-stream/src/stream_map.rs:471-483

rust
pub fn remove<Q>(&mut self, k: &Q) -> Option<V>
where
    K: Borrow<Q>,
    Q: Hash + Eq + ?Sized,
{
    for i in 0..self.entries.len() {
        if self.entries[i].0.borrow() == k {
            return Some(self.entries.swap_remove(i).1);
        }
    }

    None
}

Copy

StreamMapScenario-driven Walkthrough: poll_next_entry's random start point and cursor correctionpoll_next_entryThe core ofis. It starts polling from

📎 tokio-stream/src/stream_map.rs:515-550

rust
fn poll_next_entry(&mut self, cx: &mut Context<'_>) -> Poll<Option<(usize, V::Item)>> {
    let start = self::rand::thread_rng_n(self.entries.len() as u32) as usize;
    let mut idx = start;

    for _ in 0..self.entries.len() {
        let (_, stream) = &mut self.entries[idx];

        match Pin::new(stream).poll_next(cx) {
            Poll::Ready(Some(val)) => return Poll::Ready(Some((idx, val))),
            Poll::Ready(None) => {
                // Remove the entry
                self.entries.swap_remove(idx);

                // Check if this was the last entry, if so the cursor needs
                // to wrap
                if idx == self.entries.len() {
                    idx = 0;
                } else if idx < start && start <= self.entries.len() {
                    // The stream being swapped into the current index has
                    // already been polled, so skip it.
                    idx = idx.wrapping_add(1) % self.entries.len();
                }
            }
            Poll::Pending => {
                idx = idx.wrapping_add(1) % self.entries.len();
            }
        }
    }

    // If the map is empty, then the stream is complete.
    if self.entries.is_empty() {
        Poll::Ready(None)
    } else {
        Poll::Pending
    }
}

to guarantee fairness — if it always started from index 0, the first stream would starve the later ones:

Copy thread_rng_nThis code has three ingenious aspects, broken down one by one:FastRandFirst, the random start point.xorshift64+uses a thread-local

📎 tokio-stream/src/stream_map.rs:765-768

rust
/// Implement `xorshift64+`: 2 32-bit `xorshift` sequences added together.
/// Shift triplet `[17,7,16]` was calculated as indicated in Marsaglia's
/// `Xorshift` paper

fastrand_nalgorithm:% n:

📎 tokio-stream/src/stream_map.rs:787-792

rust
pub(crate) fn fastrand_n(&self, n: u32) -> u32 {
    // This is similar to fastrand() % n, but faster.
    // See https://lemire.me/blog/2016/06/27/a-fast-alternative-to-the-modulo-reduction/
    let mul = (self.fastrand() as u64).wrapping_mul(n as u64);
    (mul >> 32) as u32
}

uses Lemire's multiplicative modulo instead ofswap_removeCopySecond,idxthe cursor correction afterNone. When the stream at indexswap_removereturnsidxand is removed,moves the last element to. This moved element maystarthave already been polledidx < start && start <= self.entries.len()(if its original index was beforeidx = idx.wrapping_add(1) % len). The code usesidx == lento detect this case, and if so skips it (

). If the removed element was the last one (Poll::Pending), the cursor wraps around to 0.Third,Pending's semantics.

poll_nextIf a full traversal finds no stream ready and the collection is non-empty, it returnspoll_next_entry. At this point all streams' Wakers have been registered, and any becoming ready will wake it.

📎 tokio-stream/src/stream_map.rs:676-683

rust
fn poll_next(mut self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<Option<Self::Item>> {
    if let Some((idx, val)) = ready!(self.poll_next_entry(cx)) {
        let key = self.entries[idx].0.clone();
        Poll::Ready(Some((key, val)))
    } else {
        Poll::Ready(None)
    }
}

:ready!Copypoll_next_entryNote thePendingmacro: ifpoll_nextreturnsPending。K: Clone, the entirekey.clone()。

immediately returns

next_manyThe constraint comes from theStreamMaphere

📎 tokio-stream/src/stream_map.rs:581-583

rust
pub async fn next_many(&mut self, buffer: &mut Vec<(K, V::Item)>, limit: usize) -> usize {
    poll_fn(|cx| self.poll_next_many(cx, buffer, limit)).await
}

is the batch version of

📎 tokio-stream/src/stream_map.rs:573-578

rust
/// # Cancel safety
///
/// This method is cancel safe. If `next_many` is used as the event in a
/// [`tokio::select!`] statement and some other branch completes first,
/// it is guaranteed that no items were received on any of the underlying
/// streams.

Copynext_manyIts cancellation-safety guarantee is crucial:CopybufferWhy isbuffercancellation-safe? Because itbufferimmediately pushes elements into the caller-provided

poll_next_many, rather than buffering them internally. If the future is dropped, the already-pushed elements are still inpoll_next_entryand will not be lost. But this also means: when dropped,

📎 tokio-stream/src/stream_map.rs:597-666

rust
pub fn poll_next_many(
    &mut self,
    cx: &mut Context<'_>,
    buffer: &mut Vec<(K, V::Item)>,
    limit: usize,
) -> Poll<usize> {
    if limit == 0 || self.entries.is_empty() {
        return Poll::Ready(0);
    }

    let mut added = 0;

    let start = self::rand::thread_rng_n(self.entries.len() as u32) as usize;
    let mut idx = start;

    while added < limit {
        // Indicates whether at least one stream returned a value when polled or not
        let mut should_loop = false;

        for _ in 0..self.entries.len() {
            let (_, stream) = &mut self.entries[idx];

            match Pin::new(stream).poll_next(cx) {
                Poll::Ready(Some(val)) => {
                    added += 1;

                    let key = self.entries[idx].0.clone();
                    buffer.push((key, val));

                    should_loop = true;

                    idx = idx.wrapping_add(1) % self.entries.len();

                    if added == limit {
                        break;
                    }
                }
                Poll::Ready(None) => {
                    // Remove the entry
                    self.entries.swap_remove(idx);

                    // Check if this was the last entry, if so the cursor needs
                    // to wrap
                    if idx == self.entries.len() {
                        idx = 0;
                    } else if idx < start && start <= self.entries.len() {
                        // The stream being swapped into the current index has
                        // already been polled, so skip it.
                        idx = idx.wrapping_add(1) % self.entries.len();
                    }
                }
                Poll::Pending => {
                    idx = idx.wrapping_add(1) % self.entries.len();
                }
            }
        }

        if !should_loop {
            break;
        }
    }

    if added > 0 {
        Poll::Ready(added)
    } else if self.entries.is_empty() {
        Poll::Ready(0)
    } else {
        Poll::Pending
    }
}

's loop structure is more complex thanwhile added < limit's, because it must collect as many as possible within one round:forCopyshould_loop = trueThe outerlimitcombined with the inner

📎 tokio-stream/src/stream_map.rs:588-591

rust
/// * `Poll::Pending` if no items are available but the `StreamMap` is not empty.
/// * `Poll::Ready(count)` where `count` is the number of items successfully received and
///   stored in `buffer`. This can be less than, or equal to, `limit`.
/// * `Poll::Ready(0)` if `limit` is set to zero or when the `StreamMap` is empty.

size_hintThe implementation demonstrates how to aggregate capacity hints from multiple streams:

📎 tokio-stream/src/stream_map.rs:685-701

rust
fn size_hint(&self) -> (usize, Option<usize>) {
    let mut ret: (usize, Option<usize>) = (0, Some(0));

    for (_, stream) in &self.entries {
        let hint = stream.size_hint();

        ret.0 = ret.0.saturating_add(hint.0);

        match (ret.1, hint.1) {
            (Some(a), Some(b)) => ret.1 = a.checked_add(b),
            (Some(_), None) => ret.1 = None,
            _ => {}
        }
    }

    ret
}

Same asmerge_size_hintsthe same pattern: lower bound saturating add, upper bound checked add, if either is unknown then the whole is unknown.

Below is a flowchart depictingpoll_next_entrythe decision path of:

mermaid
flowchart TD
    start["poll_next_entry(cx)"] --> rand["start = thread_rng_n(len)"]
    rand --> loop{"遍历 len 次?"}
    loop -->|"未完成"| poll["Pin::new(stream).poll_next(cx)"]
    poll -->|"Ready(Some(val))"| ret_val["返回 Ready(Some((idx, val)))"]
    poll -->|"Ready(None)"| remove["entries.swap_remove(idx)"]
    remove --> wrap{"idx == entries.len()?"}
    wrap -->|"是"| set_zero["idx = 0"]
    wrap -->|"否"| check_swap{"idx < start && start <= len?"}
    check_swap -->|"是"| skip["idx = idx.wrapping_add(1) % len"]
    check_swap -->|"否"| loop
    set_zero --> loop
    skip --> loop
    poll -->|"Pending"| advance["idx = idx.wrapping_add(1) % len"]
    advance --> loop
    loop -->|"遍历完成"| empty{"entries.is_empty()?"}
    empty -->|"是"| ret_none["返回 Ready(None)"]
    empty -->|"否"| ret_pending["返回 Pending"]

---

TaskTracker: Encoding all state with a single AtomicUsize

Intuitive model

Graceful shutdown requires two things:Notifying tasks to stop(CancellationTokenis responsible for), andWaiting for tasks to actually exit(TaskTrackeris responsible for).TaskTrackeris like a "task counter + shutdown switch" hybrid: as long as there are still tasks running, orclose,wait()has not been called, it will not return. Without it, you could only useJoinSet, butJoinSetwould accumulate the return value of each task, and a long-running service would OOM.

Data structure and memory layout

TaskTrackeris aArcwrapper:

📎 tokio-util/src/task/task_tracker.rs:158-178

rust
pub struct TaskTracker {
    inner: Arc<TaskTrackerInner>,
}

/// Represents a task tracked by a [`TaskTracker`].
#[must_use]
#[derive(Debug)]
pub struct TaskTrackerToken {
    task_tracker: TaskTracker,
}

struct TaskTrackerInner {
    /// Keeps track of the state.
    ///
    /// The lowest bit is whether the task tracker is closed.
    ///
    /// The rest of the bits count the number of tracked tasks.
    state: AtomicUsize,
    /// Used to notify when the last task exits.
    on_last_exit: Notify,
}

This is the most ingenious memory layout in this chapter:A singleAtomicUsizesimultaneously encodes "whether closed" and "task count". The lowest bit is the closed flag, and the remaining bits are the task count (because the task count is+2each time, the lowest bit is always 0). This wayis_closed_and_emptyonly needs one atomic load:

📎 tokio-util/src/task/task_tracker.rs:216-222

rust
fn is_closed_and_empty(&self) -> bool {
    // If empty and closed bit set, then we are done.
    //
    // The acquire load will synchronize with the release store of any previous call to
    // `set_closed` and `drop_task`.
    self.state.load(Ordering::Acquire) == 1
}
[Design inference and architectural trade-offs]

state == 1means "closed bit is 1, count is 0". Why not use two atomic variables? Two variables require two loads, and cannot atomically determine "both conditions are satisfied at the same time". Single-variable encoding makesis_closed_and_emptya singleAcquireload, and on the fast path ofwaitno lock is needed.

Scenario-driven Walkthrough: the race between close and drop_task

Consider a typical scenario: the main thread callstracker.close(), while the last task is exiting (TaskTrackerToken::dropcallsdrop_task). The two may be concurrent, and it must be guaranteed that no matter which happens first,wait()can be woken up.

First look atset_closed:

📎 tokio-util/src/task/task_tracker.rs:225-249

rust
fn set_closed(&self) -> bool {
    // The AcqRel ordering makes the closed bit behave like a `Mutex<bool>` for synchronization
    // purposes. ...
    let state = self.state.fetch_or(1, Ordering::AcqRel);

    // If there are no tasks, and if it was not already closed:
    if state == 0 {
        self.notify_now();
    }

    (state & 1) == 0
}

fetch_or(1, AcqRel)atomically sets the closed bit and returns the old value. If the old value is 0 (previously not closed and no tasks), it means "after closing, empty + closed is immediately satisfied", so callnotify_now. The return value(state & 1) == 0indicates "this call actually changed the state".

Next look atdrop_task:

📎 tokio-util/src/task/task_tracker.rs:264-271

rust
fn drop_task(&self) {
    let state = self.state.fetch_sub(2, Ordering::Release);

    // If this was the last task and we are closed:
    if state == 3 {
        self.notify_now();
    }
}

fetch_sub(2, Release)decrements the count. If the old value is 3 (binary11: closed bit 1 + count 1), it means "this is the last task and it is already closed", so callnotify_now。

Race analysis of the two paths:

  • close executes first:set_closedsees the old value2(count 1, not closed), and does not notify. Thendrop_tasksees the old value3, and notifies. ✓
  • drop_task executes first:drop_tasksees the old value2(count 1, not closed), and does not notify. Thenset_closedsees the old value0(count 0, not closed), and notifies. ✓
  • Concurrent:fetch_orandfetch_subare atomic, so no matter the interleaving order, one of them will always see the "closed + empty" combination and notify. ✓

notify_nowThere is an easily overlookedAcquireload in:

📎 tokio-util/src/task/task_tracker.rs:274-285

rust
#[cold]
fn notify_now(&self) {
    // Insert an acquire fence. This matters for `drop_task` but doesn't matter for
    // `set_closed` since it already uses AcqRel.
    //
    // This synchronizes with the release store of any other call to `drop_task`, and with the
    // release store in the call to `set_closed`. That ensures that everything that happened
    // before those other calls to `drop_task` or `set_closed` will be visible after this load,
    // and those things will also be visible to anything woken by the call to `notify_waiters`.
    self.state.load(Ordering::Acquire);

    self.on_last_exit.notify_waiters();
}

Why doesdrop_taskuseReleaseinstead ofAcqRel? Becausedrop_task'sfetch_subonly needs to "make previous writes visible to subsequent readers" (Release semantics), and does not need to "see writes from other threads before this" (Acquire semantics). Butnotify_nowneeds Acquire to establish happens-before: ensuring that all cleanup work done before the task exits is visible to the code afterwait()returns. The result of thisloadis discarded purely for its memory-ordering side effect—this is a typical use of a "fence-style load" in Rust atomic operations.

Design thinking: wait's ABA resistance and TrackedFuture's drop semantics

waitreturns aTaskTrackerWaitFuture, which internally holdsNotified:

📎 tokio-util/src/task/task_tracker.rs:318-327

rust
pub fn wait(&self) -> TaskTrackerWaitFuture<'_> {
    TaskTrackerWaitFuture {
        future: self.inner.on_last_exit.notified(),
        inner: if self.inner.is_closed_and_empty() {
            None
        } else {
            Some(&self.inner)
        },
    }
}

Note theinnerfield: if it is already "closed and empty" at creation time, directly set it toNone,polland immediately returnReady. This is the fast path.

The documentation particularly emphasizes ABA resistance:

📎 tokio-util/src/task/task_tracker.rs:304-307

rust
/// The `wait` future is resistant against [ABA problems][aba]. That is, if the `TaskTracker`
/// becomes both closed and empty for a short amount of time, then it is guarantee that all
/// `wait` futures that were created before the short time interval will trigger, even if they
/// are not polled during that short time interval.

This guarantee comes from the semantics ofNotify::notified():Notifiedthe future registers itself as a "waiter" at creation time, so even ifnotify_waitersis called before it ispoll, it will still see the notification on its firstpoll.TaskTrackerWaitFuture::poll's implementation:

📎 tokio-util/src/task/task_tracker.rs:697-712

rust
fn poll(self: Pin<&mut Self>, cx: &mut Context<'_>) -> Poll<()> {
    let me = self.project();

    let inner = match me.inner.as_ref() {
        None => return Poll::Ready(()),
        Some(inner) => inner,
    };

    let ready = inner.is_closed_and_empty() || me.future.poll(cx).is_ready();
    if ready {
        *me.inner = None;
        Poll::Ready(())
    } else {
        Poll::Pending
    }
}

Eachpollfirst checksis_closed_and_empty(), thenpoll Notified. This order guarantees that even ifNotifiedis not woken for some reason, the state check can still serve as a fallback.

TrackedFuture's drop semantics are the core difference betweenTaskTrackerandJoinSet:

📎 tokio-util/src/task/task_tracker.rs:488-494

rust
/// The task is removed from the collection when it is dropped, not when [`poll`] returns
/// [`Poll::Ready`].

This means: even if the future has already returnedReady, as long asTrackedFutureitself has not been dropped,TaskTrackerstill considers the task to be alive. The documentation explains why this design is important:

📎 tokio-util/src/task/task_tracker.rs:33-35

rust
/// When a call to [`wait`] returns, it is guaranteed that all tracked tasks have exited and that
/// the destructor of the future has finished running. However, there might be a short amount of
/// time where [`JoinHandle::is_finished`] returns false.

TaskTrackerToken'sDropis the trigger point for count decrement:

📎 tokio-util/src/task/task_tracker.rs:670-672

rust
impl Drop for TaskTrackerToken {
    /// Dropping the token indicates to the [`TaskTracker`] that the task has exited.
    #[inline]
    fn drop(&mut self) {
        self.task_tracker.inner.drop_task();
    }
}

TrackedFutureBypin_project!packagingtokenandfuturetogether,token's drop automatically triggers the count decrement.spawn_blockingexplicitly manages the token:

📎 tokio-util/src/task/task_tracker.rs:452-464

At this point, StreamExt has turned poll_next into a composable iterator, StreamMap and TaskTracker give a dynamic task set a home, and CancellationToken uses a tree to propagate cancellation signals to the entire task tree. The common point of these three layers of extensions is that they do not introduce new scheduling primitives, but instead recombine existing mechanisms such as Waker, Notify, and atomic counting into higher-level abstractions. But a key question then emerges: when these combinators, task sets, and cancellation trees run concurrently on the same scheduler, how can we ensure that a task does not starve other tasks by not yielding for a long time? The next chapter will dive into Tokio's coop cooperative budget mechanism, looking at how each task consumes budget within a scheduling cycle, actively yields when exhausted, and how budget is passed through thread-local storage, thereby solving this classic problem.

CHAPTER 12

Chapter 12: Cooperative Scheduling and Budget: How the coop Mechanism Prevents Tasks from Starving the Scheduler

Project: tokio-rs/tokio · Book Progress: Chapter 12 / 14 · Verification Status: FACT line numbers genuinely anchored

In the previous chapter, we saw how tokio-stream and tokio-util reuse the underlying Waker and scheduling mechanisms to extend core capabilities. But no matter how many combinators are built, the core contradiction of an async runtime always exists: the scheduler must fairly distribute CPU time among multiple tasks, while tasks themselves are non-preemptive—once a Future's poll begins executing, the scheduler cannot interrupt it from the outside. If a task processes a hundred thousand messages in a single poll, or repeatedly awaits an always-ready Future in a loop, it will monopolize the worker thread, leaving other tasks on the same thread forever without a chance to be polled. This is the classic "task starves the scheduler" problem. Tokio's solution is not preemption, but cooperation: each task is allocated a limited budget within a scheduling cycle, resource operations consume the budget, and once the budget is exhausted, the task must voluntarily yield. This chapter dives deep into the implementation of this coop mechanism.

12.1 The Budget's Carrier: Thread-Local Storage and the Budget Struct

[Design Inference and Architectural Trade-offs]

If the scheduler is likened to the only waiter in a restaurant, and tasks are customers who keep ordering dishes, then the coop budget is the rule of "each customer can order at most N dishes"—the waiter doesn't need to forcibly interrupt the customer, but simply says "take a break, I'll serve the next one" after the customer has ordered N dishes. Without this rule, one talkative customer could paralyze the entire restaurant.

The budget must satisfy two constraints: first, it must be accessible from a call stack of arbitrary depth without passing parameters layer by layer; second, it must be able to distinguish "whether currently inside the Tokio runtime"—calling from outside the runtime should not be subject to budget constraints. Tokio chose to usepollthread-local storage (TLS)block_onto carry the budget, and manages it uniformly through themodule.The core type of the budget iscontext. Although the source code slice in this chapter does not directly provide the complete definition of

, its interface contract can be inferred from the usage sites ofcoop::Budget:coop.rsCopyworker.rsThree key APIs appear here:

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:695-795

rust
coop::budget(|| {
    // ... 轮询任务 ...
    task.run();
    // ...
    loop {
        // ...
        if !coop::has_budget_remaining() {
            // 预算耗尽,把 LIFO 任务推回队列
            core.run_queue.push_back_or_overflow(task, ...);
            return ControlFlow::Continue(core);
        }
        // ...
    }
})

queries the remaining budget, and the semantics ofcoop::budget(closure)andcoop::has_budget_remaining(), which will be seen later, are: upon entering the closure, reset the current thread's budget to a full value (default 128); during the closure's execution, all resource operations share this quota; upon exiting the closure, restore the outer budget.coop::stop()[Design Inference and Architectural Trade-offs]coop::set()。budgetThe budget value of 128 is an empirical value: it is large enough that a normal message-processing loop (say, processing a few dozen messages per poll) won't frequently trigger a yield; yet small enough that a runaway loop can perform at most 128 resource operations before being forced to yield, keeping latency within an acceptable range.

In TLS, it typically exists in the form of

. The outer semantics of

Budgetis "whether the current thread is in the Tokio runtime context":Cell<Option<Budget>>indicates not inside the runtime (e.g.,Optionoutside the runtime), in which case all budget checks pass through directly.None12.2 Budget Consumption Points: How Resource Operations Deductblock_onThe budget is not consumed out of thin air; only

resource operations

deduct it. So-called resource operations refer to those APIs that may be called in an infinite loop and interact with the outside world—channel's, I/O reads and writes,, etc. Takingsend/recvas an example, it is the common entry point for all send paths:yield_nowCopympsc::Sender::reserveBefore actually acquiring the semaphore permit,

📎 tokio/src/sync/mpsc/bounded.rs:1272-1311

rust
async fn reserve_inner(&self, n: usize) -> Result<(), SendError<()>> {
    crate::trace::async_trace_leaf().await;

    if n > self.max_capacity() {
        return Err(SendError(()));
    }
    // ... WakeReceiverOnDrop guard ...
    let guard = WakeReceiverOnDrop { chan: &self.chan };
    let result = self.chan.semaphore().semaphore.acquire(n).await;
    // ...
}

reserve_inner. This call, which appears to be merely for tracing, is actually one of the mounting points for budget deduction.crate::trace::async_trace_leaf()Internally calls a function likeasync_trace_leaf: if the budget is sufficient, deduct 1 and returncoop::poll_proceed; if the budget is exhausted, register a "yield" action—hand the current task's Waker to the scheduler, returnProceed, and let the task end early in this poll.PendingThis is the brilliance of coop:

budget exhaustion does not throw an error, but disguises "yield" as an ordinary. When the upper-level Future seesPending, it naturally returns; the scheduler re-enqueues the task, and when it is next scheduled, the budget has been reset, and the task continues from where it was interrupted. The entire process is completely transparent to business code.Pendingis the most straightforward manifestation of the budget mechanism; it does not consume budget, but

yield_nowactively triggers a yieldCopy:

📎 tokio/src/task/yield_now.rs:38-60

rust
pub async fn yield_now() {
    let mut yielded = false;
    poll_fn(|cx| {
        ready!(crate::trace::trace_leaf());

        if yielded {
            return Poll::Ready(());
        }

        yielded = true;

        // Don't wake the task immediately, as that would push it right back
        // onto the run queue and it could be polled again before other tasks
        // or the IO/timer driver get a chance to run. Instead, hand the waker
        // to the scheduler, which wakes deferred tasks only after it has run
        // out of ready tasks and polled the driver. When polled from outside
        // a Tokio runtime, the waker is woken immediately.
        context::defer(cx.waker());

        Poll::Pending
    })
    .await
}

. It does not directlycontext::defer(cx.waker()), but hands the Waker to the scheduler'swakedefer queue. Why? The source code comments explain it clearly: if woken immediately, the task would be pushed back onto the run queue right away and might be polled again before the I/O/timer driver runs, making the yield meaningless. The semantics of the defer queue is "wake these tasks only after the current worker has finished running ready tasks and has polled the driver."The defer queue is defined in the worker's

:ContextCopy

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:247-257

rust
pub(crate) struct Context {
    worker: Arc<Worker>,
    core: RefCell<Option<Box<Core>>>,
    /// Tasks to wake after resource drivers are polled. This is mostly to
    /// handle yielded tasks.
    pub(crate) defer: Defer,
}

deferCopy

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:613-621

rust
} else {
    // Wait for work
    core = if !self.defer.is_empty() {
        self.park_yield(core)
    } else {
        self.park(core)
    };
    core.stats.start_processing_scheduled_tasks();
}

If the defer queue is non-empty, the worker callspark_yield—parking with a 0 timeout, which drives I/O and timers, then wakes the tasks in defer. This guarantees that a "yielded" task is only rescheduled after the driver has run.

12.3 Establishment and Restoration of Budget Scope: run_task and block_in_place

The budget scope is established inrun_task. When each task is polled,coop::budgetwraps the entire polling process:

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:691-704

rust
// Make the core available to the runtime context
*self.core.borrow_mut() = Some(core);

// Run the task
coop::budget(|| {
    // ...
    task.run();
    // ...
})

coop::budgetOn entry, it sets the budget in TLS to full, and on exit, it restores it. This meanseach task gets a fresh budget every time it is polled. No matter how manyawaitresource operations the task performs internally, as long as a singlepollconsumes more than 128, it will be forced to yield.

But there is a subtle issue here: tasks in the LIFO slot are polled withinthe samebudgetclosure. Look atrun_task's loop:

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:709-750

rust
let mut lifo_polls = 0;

// As long as there is budget remaining and a task exists in the
// `lifo_slot`, then keep running.
loop {
    let mut core = match self.core.borrow_mut().take() {
        Some(core) => core,
        None => {
            return ControlFlow::Break(());
        }
    };

    let task = match core.lifo_slot.take() {
        Some(task) => task,
        None => {
            self.reset_lifo_enabled(&mut core);
            core.stats.end_poll();
            return ControlFlow::Continue(core);
        }
    };

    if !coop::has_budget_remaining() {
        core.stats.end_poll();
        // Not enough budget left to run the LIFO task, push it to
        // the back of the queue and return.
        core.run_queue.push_back_or_overflow(task, ...);
        debug_assert!(core.lifo_enabled);
        return ControlFlow::Continue(core);
    }
    // ...
}

The key point: tasks in the LIFO slotshare the outer task's budget. The comment at the beginning ofrun_tasksays: "Tasks from the LIFO slot inherit the 'parent''s limits". This is an intentional design—if every LIFO task reset the budget, then in a ping-pong scenario (task A wakes B, B wakes A), the two tasks would schedule each other indefinitely, the budget would always be reset, and the starvation problem would remain. Sharing the budget means A and B together can consume at most 128 resource operations, after which they must yield.

The LIFO slot itself also has an independent rate limiterMAX_LIFO_POLLS_PER_TICK:

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:756-766

rust
// Disable the LIFO slot if we reach our limit
//
// In ping-ping style workloads where task A notifies task B,
// which notifies task A again, continuously prioritizing the
// LIFO slot can cause starvation as these two tasks will
// repeatedly schedule the other. To mitigate this, we limit the
// number of times the LIFO slot is prioritized.
if lifo_polls >= MAX_LIFO_POLLS_PER_TICK {
    core.lifo_enabled = false;
    super::counters::inc_lifo_capped();
}

MAX_LIFO_POLLS_PER_TICKThe value of is 3:

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:263-263

rust
/// Value picked out of thin-air. Running the LIFO slot a handful of times
/// seems sufficient to benefit from locality. More than 3 times probably is
/// over-weighting. The value can be tuned in the future with data that shows
/// improvements.
const MAX_LIFO_POLLS_PER_TICK: usize = 3;

This isthe second line of defense: even if the budget has not been exhausted, the LIFO slot will be disabled after being prioritized 3 times in a row, and subsequent tasks will go through the normal queue. The budget governs "the total amount of resource operations", while the LIFO rate limiter governs "the number of times the same pair of tasks wake each other", and the two are complementary.

The budget scope has an important exception inblock_in_place.block_in_placehands over the worker core to another thread, and the current thread enters a blocked state. Blocking code is not subject to the budget, so it mustpausethe budget:

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:406-417

rust
if had_entered {
    // Unset the current task's budget. Blocking sections are not
    // constrained by task budgets.
    let _reset = Reset {
        take_core,
        budget: coop::stop(),
    };

    crate::runtime::context::exit_runtime(f)
} else {
    f()
}

coop::stop()returns the current budget and sets it toNone(that is, "not inside the runtime"),Reset'sDroprestores it after blocking ends:

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:374-397

rust
impl Drop for Reset {
    fn drop(&mut self) {
        with_current(|maybe_cx| {
            if let Some(cx) = maybe_cx {
                if self.take_core {
                    let core = cx.worker.core.take();
                    // ...
                    *cx_core = core;
                }

                // Reset the task budget as we are re-entering the
                // runtime.
                coop::set(self.budget);
            }
        });
    }
}

coop::set(self.budget)restores the budget previously saved bystop(). In this way,block_in_placesynchronous blocking code inside does not consume budget, nor does it mistakenly trigger a yield due to budget exhaustion; after blocking ends, the task continues executing with its original remaining budget.

The following diagram shows the complete control flow from a task being scheduled to yielding due to budget exhaustion:

mermaid
flowchart TD
    start["Context::run 主循环"] --> next["core.next_task()"]
    next --> has_task{"有本地任务?"}
    has_task -->|是| run_task["run_task(task, core)"]
    has_task -->|否| steal["core.steal_work()"]
    steal --> stolen{"窃取到任务?"}
    stolen -->|是| run_task
    stolen -->|否| defer_check{"defer 队列非空?"}
    defer_check -->|是| park_yield["park_yield: 驱动 IO/timer 后唤醒"]
    defer_check -->|否| park["park: 阻塞等待"]
    park_yield --> start
    park --> start

    run_task --> budget["coop::budget 建立满额预算"]
    budget --> poll["task.run() 轮询"]
    poll --> lifo_check{"lifo_slot 有任务?"}
    lifo_check -->|否| done["返回 ControlFlow::Continue"]
    lifo_check -->|是| budget_rem{"coop::has_budget_remaining()?"}
    budget_rem -->|否| push_back["push_back_or_overflow 推回队列"]
    push_back --> done
    budget_rem -->|是| lifo_limit{"lifo_polls >= 3?"}
    lifo_limit -->|是| disable["core.lifo_enabled = false"]
    lifo_limit -->|否| poll_lifo["task.run() 轮询 LIFO 任务"]
    disable --> poll_lifo
    poll_lifo --> lifo_check
    done --> start

In the diagram, two yield paths can be seen: when the budget is exhausted, the LIFO task is pushed back into the queue (push_back_or_overflow), and when the LIFO consecutive priority limit is exceeded, the LIFO slot is disabled. Both return to the main loop, giving the worker a chance to handle other tasks or the driver.

12.4 Design Considerations, Error Recovery, and Production Pitfalls

Why use TLS instead of explicit parameter passing?Budget checkpoints are scattered deep in various modules such as channel, I/O, and time. If passed explicitly, every API would need an extraBudgetparameter, polluting the entire public interface. TLS makes the budget completely transparent to business code, at the cost of one TLS access overhead per check. Tokio uses#[thread_local]or platform-specific fast TLS to reduce this overhead.

Interaction between budget exhaustion and cancellation safety.When budget exhaustion causesreserve_innerto returnPending, the task may be in some branch ofselect!. If another branch becomes ready at this point,select!will cancel the current branch—reserve_inner'sWakeReceiverOnDropguard will, on drop, check "the semaphore is closed and idle" and wake the receiver:

📎 tokio/src/sync/mpsc/bounded.rs:1286-1299

rust
struct WakeReceiverOnDrop<'a, T> {
    chan: &'a chan::Tx<T, Semaphore>,
}

impl<T> Drop for WakeReceiverOnDrop<'_, T> {
    fn drop(&mut self) {
        use chan::Semaphore;

        let semaphore = self.chan.semaphore();
        if semaphore.is_closed() && semaphore.is_idle() {
            self.chan.wake_rx();
        }
    }
}

The existence of this guard shows that budget-triggeredPendingand a true "no permit"Pendingmust behave consistently on the cancellation path; otherwise, the receiver may never receive the notification that "the channel has been closed".

Production pitfall: hidden latency caused by budget exhaustion.A common phenomenon is that a task suddenly slows down in processing messages, but CPU usage is not high. When troubleshooting, it is easy to suspect lock contention or I/O, but in reality the task may have processed more than 128 messages within a single poll, triggering a budget yield, and each yield goes through a complete cycle of "push back to queue → reschedule → drive poll". If message processing itself is fast, this scheduling overhead may account for a high proportion. The solution is to split large batch processing into multiplespawntasks, or explicitly insertyield_now。

into the loop. Boundary between budget andblock_in_place.As seen earlier,block_in_placewillcoop::stop()pause the budget. But note:coop::stop()is only called whenhad_enteredis true, that is, only when it is indeed on a runtime worker thread. Ifblock_in_placeis called outside the runtime,f()executes directly, and the budget state remains unchanged. This branch check is done inmaybe_move_runtime:

📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:424-464

rust
with_current(|maybe_cx| {
    match (
        crate::runtime::context::current_enter_context(),
        maybe_cx.is_some(),
    ) {
        (context::EnterRuntime::Entered { .. }, true) => {
            had_entered = true;
        }
        (
            context::EnterRuntime::Entered {
                allow_block_in_place,
            },
            false,
        ) => {
            if allow_block_in_place {
                had_entered = true;
                return Ok(());
            } else {
                return Err(
                    "can call blocking only when running on the multi-threaded runtime",
                );
            }
        }
        (context::EnterRuntime::NotEntered, true) => {
            return Ok(());
        }
        (context::EnterRuntime::NotEntered, false) => {
            return Ok(());
        }
    }
    // ...
})

The four combinations correspond respectively to: inside a worker thread,block_on's thread pool entry, nestedblock_in_place, outside the runtime. Only the first two need to pause the budget and hand over the core.

[Design Inference and Architectural Trade-offs]

The budget value is not configurable.From the source code, the full budget value is a hardcoded constant (128), not exposed asBuilderoption. This is intentional: the budget value affects the trade-off between scheduling fairness and throughput. If users were allowed to adjust it freely, it would be easy to tune a configuration where "too large a budget causes starvation" or "too small a budget causes scheduling overhead to explode." Tokio chooses to treat it as an internal invariant.

Chapter Summary

The coop mechanism uses a three-layer design to solve the fairness problem of a non-preemptive scheduler:

1. Budget carrier:coop::Budgetstored in TLS,Optionthe outer layer distinguishes inside/outside the runtime,coop::budgetestablishes a full-budget scope,coop::stop/coop::setsupports pause and resume (block_in_placescenario).

2. Consumption points: resource operations (channel send/receive, I/O,yield_now) throughcoop::poll_proceeddeduct budget, and when exhausted, disguise "yield" asPending, transparent to business logic.

3. Yield path:yield_nowthroughcontext::deferhands the Waker to the defer queue, ensuring rescheduling only occurs after the driver polls; LIFO slot tasks share the parent task's budget and haveMAX_LIFO_POLLS_PER_TICK = 3independent rate limiting.

The key insight of this mechanism is:fairness does not require preemption, only that "infinite loops" naturally break after a finite number of steps. The budget is the measure of this "finite number of steps."

Chapter Review and Self-Test

Q1: If inrun_taskthecoop::budgetLIFO loop inside the closure is changed to callcoop::budgetto reset the budget before each poll of a LIFO task, what happens in a ping-pong scenario (task A wakes B, B wakes A)? Why does the source code choose to let LIFO tasks share the parent task's budget?

Reference Analysis: The source code inrun_task's comments explicitly states "Tasks from the LIFO slot inherit the 'parent''s limits"📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:679-682. If each LIFO task reset the budget, then in an A→B→A→B ping-pong scenario, each poll would obtain a full budget, and the two tasks could schedule each other indefinitely, never yielding due to budget exhaustion. AlthoughMAX_LIFO_POLLS_PER_TICK = 3's rate limiting would disable the LIFO slot after 3 times📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:756-766, after LIFO is disabled the tasks go through the normal queue. If only A and B are in the queue, they would still be scheduled alternately, just without LIFO priority. Shared budget provides a fallback at the total resource operation level: A and B combined can consume at most 128 resource operations before they must yield, giving other tasks and the driver a chance. The two lines of defense are complementary and neither can be missing.

Q2: yield_nowusescontext::defer(cx.waker())instead ofcx.waker().wake_by_ref(). Supposedeferwere changed to directlywake. In a single-worker multi-task scenario, what would be the consequence of a task repeatedly callingyield_nowin a loop? Analyze in conjunction with the worker main loop'spark_yieldbranch.

Reference Analysis:yield_now's comments explain the reason: a direct wake would immediately push the task back onto the run queue, and it might be polled again before the I/O/timer driver runs📎 tokio/src/task/yield_now.rs:49-54. In a single-worker scenario, if a task repeatedlyyield_nowin a loop and directly wakes each time, the worker main loop'snext_taskwould immediately pick up this task and poll it again,park_yieldthe branch (responsible for driving I/O and timer)📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:613-621would never execute, because the defer queue is empty and the local queue always has tasks. The result is that I/O events and timers never get processed, and the entire runtime is "falsely alive"—tasks are running, but events from the outside world cannot make progress.deferThe queue ensures that a yielded task must wait until after the driver polls before being woken, thus leaving an execution window for the driver.

Q3: block_in_placeincoop::stop()sets the budget toNone,Reset::dropincoop::set(self.budget)restores. If insideblock_in_place's closurefthere is another call toblock_in_place(nested), what happens to the budget state?maybe_move_runtimeWhich branch of

handles this situation?Reference Analysisblock_in_place: Nestedmaybe_move_runtimeis handled by(context::EnterRuntime::NotEntered, true)in📎 tokio/src/runtime/scheduler/multi_thread/worker.rs:454-458thereturn Ok(())branchhad_entered. This branch directlyblock_in_place, without settingif had_entered, so the outercoop::stop()'sResetcheck is false, and it will not callf()again or create a newcoop::stop(). The comment states "This is a nested call to block_in_place (we already exited). All the necessary setup has already been done."—the outer layer has already paused the budget and handed over the core, and the inner layer only needs to directly executeNone. If the inner layerReset::dropagain, it would save the budget that is alreadyNoneone more time, and

on restore might restore the wrong value (

CHAPTER 13

Back to Top ↑

Next Chapter: Chapter 13 → · Chapter 13: Production Pitfalls and Boundary Conditions: Cancellation Safety, Panic Propagation, and Shutdown Ordering · Project: tokio-rs/tokio

In the previous chapter, we dissected the coop cooperative budget: each task has only a limited budget within one scheduling cycle, and once exhausted, it must yield, thereby preventing a single task from starving others. But the budget mechanism only solves the "fair scheduling" problem. In real production environments, there is another category of more insidious traps—cancellation safety, panic propagation, and shutdown ordering. When select! cancels a Future, when a task panic is caught, when the Runtime begins shutting down, the boundary behavior of the code often contradicts intuition. This chapter starts with cancellation safety, first examining what exactly is lost when a Future is dropped.

13.2 Panic Propagation: How JoinError Captures Crashes

Intuitive Model

A Tokio task panic does not crash the entire process (unless panic=abort); instead, it is caught, packaged intoJoinError, and returned throughJoinHandle::await. This is like an accident at a workstation on a factory assembly line: the safety net catches the worker, but the product is scrapped—what you get is an "accident report" rather than the product.

Data Structure and State

JoinHandle<T>'sFuture::Outputissuper::Result<T>, i.e.,Result<T, JoinError> 📎 tokio/src/runtime/task/join.rs:325。JoinErrorhas two forms: panic and cancelled. The documentation example demonstrates the panic scenario:

rust
let join_handle = tokio::spawn(async { panic!("boom"); });
let err = join_handle.await.unwrap_err();
assert!(err.is_panic());

📎 tokio/src/runtime/task/join.rs:121-127

The mechanism by which panic is caught is inRawTask's poll path: when a task is polled, it is wrapped withcatch_unwind. After a panic occurs, the payload is stored into the task's output slot, the state is marked as complete, and then the join waker is awakened.JoinHandle::pollWhat is read throughtry_read_outputisErr(JoinError::panic(payload))。

Scenario-Driven Walkthrough: Panic Propagation Chain

mermaid
sequenceDiagram
    participant App as 应用任务
    participant Worker as Worker 线程
    participant Raw as RawTask
    participant JH as JoinHandle

    App->>Worker: spawn(async { panic!("boom") })
    Worker->>Raw: poll 任务 Future
    Raw->>Raw: catch_unwind 捕获 panic
    Raw->>Raw: 存储 panic payload 到输出槽
    Raw->>Raw: state 标记 complete
    Raw->>JH: 唤醒 join waker
    JH->>App: await 返回 Err(JoinError::panic)

Key point: the panic payload is fully preserved,JoinErrorimplementsstd::error::Error, and you can retrieveinto_panic()throughBox<dyn Any + Send>, then usedowncast_ref::<&str>()to extract the panic message.

Design Considerations and Pitfalls

Pitfall 1:JoinHandle'sUnwindSafeis manually implemented.

rust
impl<T> UnwindSafe for JoinHandle<T> {}
impl<T> RefUnwindSafe for JoinHandle<T> {}

📎 tokio/src/runtime/task/join.rs:176-181

This is an unconditional implementation and does not requireT: UnwindSafe. Reason:JoinHandleitself does not holdT,TIn the heap allocation of the task, at panic time it has already been isolated bycatch_unwind. So even ifTis notUnwindSafe,JoinHandle, it is still safe.

Pitfall 2: Panic does not automatically propagate to the parent task.If task A spawns task B and B panics, A will not automatically be notified unless A awaits B'sJoinHandle. If A does not await, B's panic is silently swallowed. This is one of the most insidious sources of bugs in production environments.

Pitfall 3:spawn_blocking's panic is likewise caught.The worker of the blocking thread pool also wraps the task withcatch_unwind. After a panic, the thread does not die but returns to the pool to continue taking work. But if you holdMutexin a blocking task and do not release it on panic, it will cause lock poisoning—this isstd::sync::Mutex's inherent behavior, and Tokio does not intervene.

Pitfall 4: Panic during Runtime drop.If a task panics during Runtime drop,catch_unwindstill takes effect, but at this point the join waker may already be invalid, and the panic payload will be discarded. This is a subset of the shutdown ordering problem, which will be expanded in the next section.

13.3 Shutdown Ordering: Cleanup of Blocking Threads and I/O Resources

Intuitive Model

Runtime shutdown is like a restaurant closing: first let the front desk stop taking customers (stop accepting new tasks), then wait for the kitchen to finish the dishes at hand (async tasks run to the next yield point), and finally wait for outsourced helpers to finish up (blocking threads return). If the order is wrong, problems arise—for example, if you send the helpers away first, the kitchen's dishes will never be finished.

Data Structure and Shutdown Path

Runtime's three fields determine the shutdown order:

rust
pub struct Runtime {
    scheduler: Scheduler,
    handle: Handle,
    blocking_pool: BlockingPool,
}

📎 tokio/src/runtime/runtime.rs:97-106

DropImplementation:

rust
impl Drop for Runtime {
    fn drop(&mut self) {
        match &mut self.scheduler {
            Scheduler::CurrentThread(current_thread) => {
                let _guard = context::try_set_current(&self.handle.inner);
                current_thread.shutdown(&self.handle.inner);
            }
            Scheduler::MultiThread(multi_thread) => {
                multi_thread.shutdown(&self.handle.inner);
            }
        }
    }
}

📎 tokio/src/runtime/runtime.rs:506-521

Note:Droponly handlesscheduler,does not explicitly handleblocking_pool。blocking_pool's shutdown occurs in its ownDrop, triggered by field drop order afterRuntime::dropreturns. Field drop order is declaration order:scheduler → handle → blocking_pool. Therefore, the blocking pool is shut down last.

Butshutdown_timeoutexplicitly controls the order:

rust
pub fn shutdown_timeout(mut self, duration: Duration) {
    self.handle.inner.shutdown();
    self.blocking_pool.shutdown(Some(duration));
}

📎 tokio/src/runtime/runtime.rs:457-461

Firsthandle.inner.shutdown()notifies the scheduler and I/O driver to stop, thenblocking_pool.shutdown(Some(duration))waits for blocking tasks, at most waitingduration。

The underlying mechanism of blocking pool shutdown

blocking/shutdown.rsuses an ingenious oneshot channel:

rust
pub(super) struct Sender {
    _tx: Arc<oneshot::Sender<()>>,
}

pub(super) struct Receiver {
    rx: oneshot::Receiver<()>,
}

📎 tokio/src/runtime/blocking/shutdown.rs:13-19

Each blocking worker holds aSenderclone (internallyArc<oneshot::Sender>). When all workers exit and allSenderare dropped,Receiverreceives the notification.waitMethod:

rust
pub(crate) fn wait(&mut self, timeout: Option<Duration>) -> bool {
    use crate::runtime::context::try_enter_blocking_region;

    if timeout == Some(Duration::from_nanos(0)) {
        return false;
    }

    let mut e = match try_enter_blocking_region() {
        Some(enter) => enter,
        _ => {
            if std::thread::panicking() {
                return false;
            } else {
                panic!(
                    "Cannot drop a runtime in a context where blocking is not allowed. \
                    This happens when a runtime is dropped from within an asynchronous context."
                );
            }
        }
    };

    if let Some(timeout) = timeout {
        e.block_on_timeout(&mut self.rx, timeout).is_ok()
    } else {
        let _ = e.block_on(&mut self.rx);
        true
    }
}

📎 tokio/src/runtime/blocking/shutdown.rs:37-70

Step-by-step analysis:

1. timeout == Some(0)directly returns false—this isshutdown_background's path, without waiting.

2. try_enter_blocking_region()Attempts to enter the blocking region. If currently in an async context (such as dropping Runtime inside an async task), returnsNone。

3. When entry fails, if currently panicking, return false (do not panic while already panicking); otherwise panic with a clear error message.

4. If there is a timeout, useblock_on_timeout, returning false on timeout; if there is no timeout, wait indefinitely.

Complete Flow of Shutdown Ordering

mermaid
flowchart TD
    start["Runtime::drop 或 shutdown_timeout"] --> sched{"scheduler 类型?"}
    sched -->|CurrentThread| ct["try_set_current + current_thread.shutdown"]
    sched -->|MultiThread| mt["multi_thread.shutdown"]
    ct --> handle_drop["handle 字段 drop"]
    mt --> handle_drop
    handle_drop --> bp_drop["blocking_pool 字段 drop"]
    bp_drop --> bp_wait{"shutdown_timeout 已调用?"}
    bp_wait -->|是| explicit["blocking_pool.shutdown(Some(duration))"]
    bp_wait -->|否| implicit["BlockingPool::drop 默认等待"]
    explicit --> wait_check{"try_enter_blocking_region 成功?"}
    implicit --> wait_check
    wait_check -->|否且在 panic| skip["返回 false 不等待"]
    wait_check -->|否且不在 panic| panic_err["panic: Cannot drop a runtime in async context"]
    wait_check -->|是| block_on["block_on 等待所有 Sender drop"]

Design Considerations and Pitfalls

Pitfall 1: Dropping Runtime in an async context will panic.The error message is clear: "Cannot drop a runtime in a context where blocking is not allowed"📎 tokio/src/runtime/blocking/shutdown.rs:51-54. The solution is to useshutdown_background(), which is equivalent toshutdown_timeout(Duration::from_nanos(0)) 📎 tokio/src/runtime/runtime.rs:494-496, without waiting for blocking tasks.

Pitfall 2:shutdown_backgroundwill leak blocking tasks.The documentation explicitly warns "this may result in a resource leak (in that any blocking tasks are still running until they return)"📎 tokio/src/runtime/runtime.rs:470-472. Blocking tasks will continue running until they naturally return, but the Runtime has already been dropped, and the resources they hold may have become invalid.

Pitfall 3: I/O resources become invalid after the Runtime is dropped.The documentation states "Once the runtime has been dropped, any outstanding I/O resources bound to it will no longer function"📎 tokio/src/runtime/runtime.rs:52-54。is_rt_shutdown_errThe function is used to detect this kind of error📎 tokio/src/runtime/runtime.rs:585-593。

Pitfall 4:Dropwaits indefinitely by default.The documentation points out "TheDrop implementation waits forever for this」📎 tokio/src/runtime/runtime.rs:43-44. If a blocking task gets stuck (e.g., an infinite loop), dropping the Runtime will hang forever. In production, you should useshutdown_timeoutto set an upper limit.

13.4 Signal Handling and Multi-Runtime Conflicts

Intuitive Model

Unix signals are process-level, but Tokio'sSignalis bound to the Runtime. This is like a building sharing a single fire alarm bell, but each room installing its own independent receiver—the first person to install a receiver changed how the bell is wired, and everyone after can only share that change.

Data Structures and Global State

signal_enableis the entry point for registering signal handlers:

rust
fn signal_enable(signal: SignalKind, handle: &Handle) -> io::Result<()> {
    let signal = signal.0;
    if signal <= 0 || signal_hook_registry::FORBIDDEN.contains(&signal) {
        return Err(Error::other(format!(
            "Refusing to register signal {signal}"
        )));
    }

    handle.check_inner()?;

    let globals = globals();
    let siginfo = match globals.storage().get(signal as EventId) {
        Some(slot) => slot,
        None => return Err(io::Error::other("signal too large")),
    };

    siginfo
        .init
        .get_or_init(|| {
            unsafe { signal_hook_registry::register(signal, move || action(globals, signal)) }
                .map(|_| ())
                .map_err(|e| e.raw_os_error())
        })
        .map_err(|e| {
            e.map_or_else(
                || Error::other("registering signal handler failed"),
                || Error::from_raw_os_error,
            )
        })
}

📎 tokio/src/signal/unix.rs:266-296

Key points:

1. signal <= 0 || FORBIDDEN.contains(&signal)rejects illegal signals.

2. handle.check_inner()checks whether the signal driver is running—if the Runtime has been shut down, this will fail.

3. siginfo.init.get_or_init(...)usesOnceLockto ensure each signal registers an OS handler only once.get_or_initThe closure callssignal_hook_registry::register, which is a global, process-level registration.

4. The registered handler isaction(globals, signal), which does two things:globals.record_event(signal)records the event, then writes a byte to the pipe to wake up the driver📎 tokio/src/signal/unix.rs:252-259。

The Root Cause of Multi-Runtime Conflicts

globals()returns a process-level globalGlobals,OsExtraDataInsideUnixStreamthe pair is also global:

rust
pub(crate) struct OsExtraData {
    sender: UnixStream,
    pub(crate) receiver: UnixStream,
}

📎 tokio/src/signal/unix.rs:61-64

DefaultThe implementation creates a pair ofUnixStream 📎 tokio/src/signal/unix.rs:61-64. This pipe is globally unique, and all Runtimes' signal drivers share it.

Here's the problem:signal_enableInsidehandle.check_inner()checksthe current Runtime'ssignal driver. But the handler registered bysignal_hook_registry::registerisprocess-level, and it writes to theglobalpipe. If Runtime A registers SIGINT first, then Runtime B also registers SIGINT,get_or_initwill directly return the existingOk(()), without re-registering. But Runtime B's signal driver will read data from the global pipe—the two Runtimes will compete for bytes from the same pipe.

Scenario-Driven Walkthrough: Multi-Runtime Signal Contention

mermaid
sequenceDiagram
    participant OS as 操作系统
    participant Handler as 全局 signal handler
    participant Pipe as 全局 UnixStream pipe
    participant RtA as Runtime A 信号驱动
    participant RtB as Runtime B 信号驱动

    Note over RtA: signal(SIGINT) 注册
    RtA->>Handler: signal_hook_registry::register(SIGINT, action)
    Note over RtB: signal(SIGINT) 注册
    RtB->>Handler: get_or_init 返回已有 Ok,不重复注册
    OS->>Handler: 投递 SIGINT
    Handler->>Pipe: write(&[1])
    Pipe->>RtA: 可读事件
    Pipe->>RtB: 可读事件
    Note over RtA,RtB: 两个 Runtime 竞争读取,只有一个能读到字节

Design Reflections and Pitfalls

Pitfall 1: Signal handlers are never unloaded.The documentation explicitly warns "Once a signal handler is registered with the process the underlying libc signal handler is never unregistered"📎 tokio/src/signal/unix.rs:379-380. Even if theSignalinstance is dropped, subsequent signals will still be captured by Tokio, and the default behavior will not be restored📎 tokio/src/signal/unix.rs:338-340。

Pitfall 2: Signals get coalesced.The documentation states "beforepoll is called, all signal notifications are coalesced into one item returned from poll」📎 tokio/src/signal/unix.rs:312-315. If you receive 10 SIGINTs but only poll once, you'll only see one event. This is a characteristic of Unix signals themselves (standard signals are not queued); Tokio does not perform additional coalescing.

Pitfall 3: Signals may be lost under multiple Runtimes.Since the global pipe is read competitively by multiple Runtimes, one Runtime may read the byte while another waits forever. In production, you should handle signals in only one Runtime, or usesignal_hookto manage it yourself.

Pitfall 4:signalfunction panic conditions.The documentation states "This function panics if there is no current reactor set, or if thert feature flag is not enabled」📎 tokio/src/signal/unix.rs:398-405. Callingsignal()outside a Runtime will panic.

Pitfall 5:recv()cancel safety.The documentation guarantees "This method is cancel safe. If you use it as a branch intokio::select! and another branch completes first, then it is guaranteed that no signal is lost」📎 tokio/src/signal/unix.rs:423-427. This is because signal events are stored in the globalEventInfo,recv()only reads and does not consume the underlying state.

Design Reflections

The three topics in this chapter share one underlying pattern:Ownership of state determines the safety of cancellation/shutdown/signals。

  • JoinHandleis cancel-safe, because the output is on the heap, and the handle is just a reference.
  • Runtime shutdown order is sensitive, because the blocking pool and the scheduler shareHandle, wrong order will cause deadlock or panic.
  • Signals conflict across multiple Runtimes, because handlers and pipes are process-level global state, whileSignalis a Runtime-level view.

After understanding this pattern, the pitfall-avoidance checklist can be summarized into three principles:

1. Cancellation safety = state lives outside the Future.If the Future has an internal buffer, dropping it will lose data.JoinHandle、Signal::recv、tokio::sync::mpsc::Receiver::recvall satisfy this condition.

2. Shutdown order = reverse of dependency direction.Whoever depends on whom, shut down the depended-upon first. The scheduler depends on the I/O driver, so shut down the scheduler first; the blocking pool is independent, so shut it down last.

3. Global state = multi-instance conflict.Any process-level resource (signal handler, pipe, file descriptor table) will conflict under multiple Runtimes; either restrict to a single Runtime or use external synchronization.

Chapter Summary

Chapter Review Questions

Q1: If you removeJoinHandle::pollfromcoop::poll_proceed(cx), in what scenario would it cause other tasks to starve? Why doestry_read_outputitself not consume budget?

Reference Analysis:coop::poll_proceed(cx)consumes cooperative budget at📎 tokio/src/runtime/task/join.rs:325-325. If removed, a task that repeatedlyselect!multipleJoinHandlein a loop can poll all handles indefinitely within a single scheduling cycle, never returningPending, thereby starving other tasks on the same worker.try_read_outputitself does not consume budget, because it is just a memory read plus a possible waker store, involving no I/O or lock contention, with minimal overhead. The design intent of the budget mechanism is to constrain "operations that may run for a long time," not to charge for every poll. Note thatcoop.made_progress()only callsret.is_ready()when📎 tokio/src/runtime/task/join.rs:349-351, i.e., budget is only returned when output is actually obtained—this is to prevent operations that "polled but got no result" from accumulating budget consumption.

Q2:blocking/shutdown.rsIn thewaitmethod oftry_enter_blocking_region(), ifNonereturnsfalseand a panic is currently in progress, why choose to return

instead of continuing to wait? What would happen if changed to continue waiting?:try_enter_blocking_region()Reference AnalysisNoneReturning📎 tokio/src/runtime/blocking/shutdown.rs:44-57indicates that we are currently in an async context and blockingfalseis not allowed. If a panic is in progress at this time, the code chooses to return📎 tokio/src/runtime/blocking/shutdown.rs:47-49without waiting forblock_on. The reason is: panicking again during panic unwinding causes the process to abort (double panic). If changed to continue waiting, it would need to callblock_on, and in an async contextfalsewill panic—panicking during panic unwinding directly aborts the process, losing all diagnostic information. Returning

lets drop continue to completion, preserving the panic information. This is a "graceful degradation" design: an incomplete shutdown is better than a process crash.SignalQ3: Suppose you createSignalin Runtime A to listen for SIGTERM, then movesignal_enableto Runtime B for polling.handle.check_inner()Which Runtime does theSignalinside check? If Runtime A is dropped first, can the

in Runtime B still receive signals?:signal_enableReference Analysissignal()executes whenhandleis called, at which point📎 tokio/src/signal/unix.rs:398-405。check_inner()is Runtime A's📎 tokio/src/signal/unix.rs:275。Signalchecks Runtime A's signal driverRxFutureinternally iswatch::Receiver<()> 📎 tokio/src/signal/unix.rs:366-368, wrappingGlobals, and this receiver is registered on the globalEventInfo'srecord_event. If Runtime A is dropped, its signal driver stops reading data from the global pipe, but the global handler will stillEventInfoand write to the pipe. If Runtime B's signal driver is also running, it will read the pipe data and triggerSignal, thereby wakingSignal 's waker. So thein Runtime B maySignalstill receive signals, but it depends on whether Runtime B has a signal driver running. If Runtime B has no signal driver (e.g., signal feature not enabled or driver already shut down), no one reads the pipe data, and

will never be woken. This is the fragility of multi-Runtime signal handling.

Chapter Transitioncatch_unwindCancellation safety, panic propagation, shutdown order, signal conflicts—the common root of these four problems is the ambiguity of "state ownership" at async boundaries. Tokio provides engineering-usable answers by putting state on the heap, managing lifetimes with reference counting, usingGlobalsto isolate panics, and using a global

to share signal state. But these answers all have boundary conditions that must be explicitly handled in production.

At this point, we have covered the most error-prone boundary areas in Tokio production environments: cancellation safety relies on output being stored on the heap, and the atomicity of try_read_output; JoinHandle::drop does not cancel the task, while abort truly cancels it but has no effect on spawn_blocking; panics are caught by catch_unwind and packaged into JoinError, and are silently lost if not awaited; Runtime shutdown has a strict order, and dropping it in an async context will panic; signal handlers are process-level global state and are never unregistered once registered. Behind these rules are Tokio's repeated trade-offs between correctness and performance. In the next chapter, we will step away from specific mechanisms, review the origins of these trade-offs from an architectural perspective, and look ahead to where io_uring, driver refactoring, and custom executor interfaces will take Tokio.

CHAPTER 14

Chapter 14: Architectural Trade-offs and Future Evolution: From io_uring to Pluggable Drivers

Project: tokio-rs/tokio · Book progress: Chapter 14 / 14 · Verification status: FACT line numbers are truly anchored

In the previous chapter, we sorted out four types of production pitfalls: cancellation safety, panic propagation, shutdown order, and signal conflicts. They may seem scattered, but in fact they all point to the same architectural problem: how state ownership is clearly divided across asynchronous boundaries. And the way ownership is divided is precisely determined by the three lowest-level architectural decisions of the runtime—how tasks are scheduled, how I/O events are dispatched, and how concurrency correctness is verified. This chapter no longer digs into the implementation details of a specific function, but instead stands at the architectural level, reviews Tokio's trade-offs on these decisions, and follows the evolution clues already embedded in the official documentation and source code to see where io_uring, driver refactoring, and custom executor interfaces will take Tokio. After reading this chapter, you should be able to answer a practical question: when should you extend Tokio, and when should you bypass it.

1. Three Historical Trade-offs: Why It Is the Way It Is Now

Intuitive model

Imagine Tokio as a restaurant that has been open for ten years. The kitchen scheduling method (work-stealing), the separate staffing of food runners (separation of the I/O driver from the scheduler), and the kitchen hygiene inspection system (loom concurrency verification) were not all designed on the first day of opening, but gradually evolved as "more customers arrived and dishes became more complex." Only by understanding these evolutions can you judge which designs are forward-looking arrangements and which are historical baggage.

Trade-off 1: work-stealing instead of a global queue

[Design inference and architectural trade-off]

A global queue is the simplest to implement: all tasks go into oneMutex<VecDeque>, and worker threads contend for the lock to take tasks. But lock contention worsens as the number of cores increases, and cache locality is poor—which core a task is created on and which core it is executed on are completely random.

The trade-off of work-stealing is: each worker holds a local queue,spawnWhen pushing, it prioritizes the local queue (lock-free, cache-friendly), and only when the local queue is empty does it steal from the tail of another worker's queue. The cost is delayed load balancing, and stealing itself requires atomic operations and memory barriers. Tokio chose the latter because modern servers often have dozens of cores, and the cost of lock contention is far higher than the occasional stealing overhead.

[Design inference and architectural trade-off]

The boundary condition of this decision is:task granularity cannot be too fine. If each task only does a few microseconds of work, the overhead of stealing and scheduling will become disproportionately large. This is also why Tokio, in addition tospawn_blocking, also requires long tasks to activelyyield_now()—cooperative scheduling is essentially there to backstop work-stealing.

Trade-off 2: The I/O driver is independent of the scheduler

This is the most intriguing point in this chapter's source material. Look attokio/src/runtime/io/mod.rs's module structure:

📎 tokio/src/runtime/io/mod.rs:5-22

rust
mod driver;
use driver::{Direction, Tick};
pub(crate) use driver::{Driver, Handle, ReadyEvent};

mod registration;
pub(crate) use registration::Registration;

mod registration_set;
use registration_set::RegistrationSet;

mod scheduled_io;
use scheduled_io::ScheduledIo;

mod metrics;
use metrics::IoDriverMetrics;

use crate::util::ptr_expose::PtrExposeDomain;
static EXPOSE_IO: PtrExposeDomain<ScheduledIo> = PtrExposeDomain::new();

Note thatdriver、registration、scheduled_ioare three independent modules, and externally onlyDriver、Handle、ReadyEvent、Registrationthese types are exposed.ScheduledIoispub(crate)'s—it isPtrExposeDomainwrapped, used to expose raw pointers to concurrency checking under loom tests.

[Design inference and architectural trade-off]

Why is the I/O driver not directly embedded into the scheduler? Because their lifecycles and concurrency models are different. The scheduler cares about "which task should run," while the I/O driver cares about "which fd is ready." If coupled, then every adjustment to the scheduling strategy would require touching the I/O path, and vice versa. More importantly,block_onthe single-threaded runtime also needs an I/O driver, but does not need a work-stealing scheduler—separation allows the two runtimes to reuse the same I/O implementation.

Trade-off 3: Using loom for concurrency model checking

tokio/src/loom/mod.rsIt is only 14 lines, yet it reveals Tokio's verification strategy for concurrency correctness:

📎 tokio/src/loom/mod.rs:1-14

rust
//! This module abstracts over `loom` and `std::sync` depending on whether we
//! are running tests or not.

#![allow(unused)]

#[cfg(not(all(test, loom)))]
mod std;
#[cfg(not(all(test, loom)))]
pub(crate) use self::std::*;

#[cfg(all(test, loom))]
mod mocked;
#[cfg(all(test, loom))]
pub(crate) use self::mocked::*;

The key is the#[cfg(all(test, loom))]condition: only when bothtestandloomcfgs are enabled at the same time will themockedmodule replacestd. This means there is no loom code at all in production builds, with zero runtime overhead.

[Design inference and architectural trade-off]

The value of loom is that it can exhaustively enumerate "all possible orders of thread interleavings." LikeScheduledIoinAtomicUsize's read-modify-write,WaitersLinked list insertion and deletion—these might run a million times on real hardware without errors, but loom can construct an interleaving that triggers a race condition within seconds. The cost is slow test execution and high memory usage, so it can only be used for unit tests, not in production.

Design Reflections

These three trade-offs share a common characteristic:They all chose the "more complex but more scalable" approach, and confined the complexity internally. The complexity of work-stealing is hidden in the scheduler, the complexity of I/O driver is hidden inScheduledIo, and the complexity of loom is hidden in cfg conditions. The externally exposed API is alwaysspawn、TcpStream::readthese simple interfaces.

[Design Inference and Architectural Trade-offs]

This is also the first principle for judging "when to extend Tokio":If your needs can be expressed by existing APIs, don't touch the internal structures. Once you start depending onpub(crate)'s types ortokio_unstable's cfg, it means you've bound yourself to Tokio's internal implementation, and you'll pay the price when upgrading.

---

II. Driver Refactoring: From "One Waker One Direction" to "Arbitrary Interest Sets"

Intuitive Model

Early Tokio I/O types had a hard limitation:async fn read(&mut self)requires&mut self. This is like a restaurant with only one pickup window, where only one person can queue at a time—because the waker is stored inside the I/O resource, not in the Future corresponding to the operation.tokio/docs/reactor-refactor.mdfully documents the cause of this limitation and the refactoring plan.

Pain Points of the Old Architecture

The document states the problem right at the beginning:

📎 tokio/docs/reactor-refactor.md:16-20

rust
Currently, I/O types require `&mut self` for `async` functions. The reason for
this is the task's waker is stored in the I/O resource's internal state
(`ScheduledIo`) instead of in the future returned by the `async` function.
Because of this limitation, I/O types limit the number of wakers to one per
direction (a direction is either read-related events or write-related events).
[Design Inference and Architectural Trade-offs]

Storing the waker inside the resource means "one direction can only have one waiter." If you want to read and write the sameTcpStreamsimultaneously, you mustsplit()it into two halves, each holding an independent waker slot. This is whyTcpStream::split()exists—it's not an API design preference, but a direct constraint of the internal data structure.

New Architecture: Moving the Waker into the Future

The core idea of the refactoring is "moving the waker from the resource state into the operation Future," thereby supporting multiple wakers registered per operation:

📎 tokio/docs/reactor-refactor.md:22-25

rust
Moving the waker from the internal I/O resource's state to the operation's
future enables multiple wakers to be registered per operation. The "intrusive
wake list" strategy used by `Notify` applies to this case, though there are some
concerns unique to the I/O driver.

The newScheduledIostructure is as follows:

📎 tokio/docs/reactor-refactor.md:97-134

rust
#[derive(Debug)]
pub(crate) struct ScheduledIo {
    /// Resource's known state packed with other state that must be
    /// atomically updated.
    readiness: AtomicUsize,

    /// Tracks tasks waiting on the resource
    waiters: Mutex<Waiters>,
}

#[derive(Debug)]
struct Waiters {
    // List of intrusive waiters.
    list: LinkedList<Waiter>,

    /// Waiter used by `AsyncRead` implementations.
    reader: Option<Waker>,

    /// Waiter used by `AsyncWrite` implementations.
    writer: Option<Waker>,
}

// This struct is contained by the **future** returned by `readiness()`.
#[derive(Debug)]
struct Waiter {
    /// Intrusive linked-list pointers
    pointers: linked_list::Pointers<Waiter>,

    /// Waker for task waiting on I/O resource
    waiter: Option<Waker>,

    /// Readiness events being waited on. This is
    /// the value passed to `readiness()`
    interest: mio::Ready,

    /// Should not be `Unpin`.
    _p: PhantomPinned,
}

There are several elegant design points worth elaborating on:

First,readinessisAtomicUsize,waitersisMutex<Waiters>。Why not use a single lock to protect both? Becausereadiness's read operations are extremely frequent (checked on everyreadiness()call), while write operations only occur when mio events are received. Using atomic variables to make the read path lock-free is a typical read-write separation optimization.

Second,Waiteris an intrusive linked list node. pointers: linked_list::Pointers<Waiter>makesWaiteritself part of the linked list, without needing to allocate additional nodes._p: PhantomPinnedexplicitly marks it as notUnpin—because once an intrusive linked list node's address moves, the linked list breaks.

Third,readerandwriterthe twoOption<Waker>are forAsyncRead/AsyncWrite's use.The document explains the reason:

📎 tokio/docs/reactor-refactor.md:210-213

rust
The `AsyncRead` and `AsyncWrite` traits use a "poll" based API. This means that
it is not possible to use an intrusive linked list to track the waker.
Additionally, there is no future associated with the operation which means it is
not possible to cancel interest in the readiness events.
[Design Inference and Architectural Trade-offs]

This is a compromise coexistence of the old and new mechanisms:async fnThe path uses an intrusive linked list (supports multiple waiters, cancellable),pollThe path uses fixed slots (doesn't support cancellation, but is trait-compatible). This "coexistence of two mechanisms" is a typical cost of incremental refactoring.

Race Conditions and the Tick Mechanism

The trickiest problem in the refactoring is race conditions. The document gives a specific deadlock scenario:

📎 tokio/docs/reactor-refactor.md:175-175

rust
If care is not taken, if between `mio_socket.read(buf)` returning and
`clear_readiness(event)` is called, a readiness event arrives, the `read()`
function could deadlock. This happens because the readiness event is received,
`clear_readiness()` unsets the readiness event, and on the next iteration,
`readiness().await` will block forever as a new readiness event is not received.

The solution is to introduce a tick mechanism, splittingreadinessthisAtomicUsizeinto multiple bit segments:

📎 tokio/docs/reactor-refactor.md:199-199

code
| shutdown | generation |  driver tick | readiness |
|----------+------------+--------------+-----------|
|   1 bit  |   7 bits   +    8 bits    +  16 bits  |
[Design Inference and Architectural Trade-offs]

This bit segment layout is a classic case of "trading space for correctness."tickincrements on eachmio::poll(),ReadyEventcarries the tick at read time.clear_readiness()Only clears the ready state when the tick matches—if the tick doesn't match, it means new events arrived in the meantime, and it must not clear. This resolves the race between "clearing" and "new event arrival" within a single atomic read-modify-write.

The following flowchart depicts the decision path betweenreadiness()andclear_readiness():

mermaid
flowchart TD
    start["readiness(interest).await"] --> check_ready{"已知 readiness<br/>与 interest 有交集?"}
    check_ready -->|是| ret_event["返回 ReadyEvent<br/>携带当前 tick"]
    check_ready -->|否| wait["注册 Waiter 到<br/>ScheduledIo.waiters"]
    wait --> mio_poll["mio.poll() 收到事件<br/>tick 递增"]
    mio_poll --> notify["遍历 waiters<br/>interest 匹配者唤醒"]
    notify --> ret_event
    ret_event --> do_read["mio_socket.read(buf)"]
    do_read --> read_ok{"read 结果?"}
    read_ok -->|Ok| done["返回 Ok(v)"]
    read_ok -->|WouldBlock| clear["clear_readiness(event)"]
    read_ok -->|其他 Err| err["返回 Err(e)"]
    clear --> tick_match{"event.tick ==<br/>当前 readiness.tick?"}
    tick_match -->|是| clear_ok["清除 readiness 位"]
    tick_match -->|否| skip["跳过清除<br/>保留新事件"]
    clear_ok --> start
    skip --> start

The key branch in this diagram is attick_match: if the tick doesn't match,clear_readinessmust abandon the clear, otherwise it will lose the just-arrived event, causing the next round ofreadiness()to block permanently.

Cancelling Interest and Memory Leaks

The intrusive linked list brings a new problem: if the Future returned byreadiness()is dropped early, the linked list node must be removed. The document explicitly warns:

📎 tokio/docs/reactor-refactor.md:144-148

rust
The future returned by `readiness()` uses an intrusive linked list to store the
waker with `ScheduledIo`. Because `readiness()` can be called concurrently, many
wakers may be stored simultaneously in the list. If the `readiness()` future is
dropped early, it is essential that the waker is removed from the list. This
prevents leaking memory.
[Design Inference and Architectural Trade-offs]

This is exactly the manifestation of "cancellation safety" from the previous chapter at the I/O layer.readiness()'s Future must remove itself from the linked list in theDropimplementation, otherwise the node will remain inScheduledIopermanently, both leaking memory and being incorrectly woken when the next event arrives.

Design Reflections and Production Pitfalls

Why not useVec<Waker>but instead an intrusive linked list?The document gives the answer when discussing the&Resourceimplementation:

📎 tokio/docs/reactor-refactor.md:228-233

rust
It is only possible to implement `AsyncRead` and `AsyncWrite` for resource types
themselves and not for `&Resource`. Implementing the traits for `&Resource`
would permit concurrent operations to the resource. Because only a single waker
is stored per direction, any concurrent usage would result in deadlocks. An
alternate implementation would call for a `Vec<Waker>` but this would result in
memory leaks.
[Design Inference and Architectural Trade-offs]

Vec<Waker>The problem with

is: after a Future is dropped, the corresponding waker remains in the Vec and cannot be located for removal, and you only discover "this waker is already invalid" when the next event arrives. The intrusive linked list makes the node address the address of a field inside the Future, allowing precise removal on drop.:TcpStream::by_ref()Production Pitfall PointsTcpStreamRefTheread_waiterreturned bywrite_waiterholds two nodes,

📎 tokio/docs/reactor-refactor.md:238-244

rust
struct TcpStreamRef<'a> {
    stream: &'a TcpStream,

    // `Waiter` is the node in the intrusive waiter linked-list
    read_waiter: Waiter,
    write_waiter: Waiter,
}
:

CopyTcpStreamRef[Design Inference and Architectural Trade-offs]select!This means onceby_ref()is dropped, both waiter nodes become invalid simultaneously. If you useTcpStreamRef's reference across branches inTcpStream, be careful with lifetimes—select!cannot outlive

---

, nor can it be borrowed simultaneously across multiple

Intuition Model

Sometimes you don't want to use Tokio's scheduler, you just want to borrow its I/O and timers. This is like not wanting to dine in at a restaurant, but only using its takeout window.examples/custom-executor.rsThis demonstrates this "hybrid mode": usingfutures::executor::ThreadPoolfor scheduling, and Tokio for I/O.

Core mechanism: TokioContext

The key to the entire example isTokioContextthis wrapper type:

📎 examples/custom-executor.rs:51-54

rust
impl ThreadPool {
    fn spawn(&self, f: impl Future<Output = ()> + Send + 'static) {
        let handle = self.rt.handle().clone();
        self.inner.spawn_ok(TokioContext::new(f, handle));
    }
}
[Design Inference and Architectural Trade-offs]

TokioContext::new(f, handle)It binds the Future together with Tokio'sHandle. When the external executor polls this wrapper Future,TokioContextit first enters Tokio's runtime context (setting the thread-localHandle), then polls the innerf. This way,fwhen callingTcpListener::bind, it can find Tokio's I/O driver.

Look at the structure of the entire example:

📎 examples/custom-executor.rs:38-48

rust
static EXECUTOR: Lazy<ThreadPool> = Lazy::new(|| {
    // Spawn tokio runtime on a single background thread
    // enabling IO and timers.
    let rt = tokio::runtime::Builder::new_multi_thread()
        .enable_all()
        .build()
        .unwrap();
    let inner = futures::executor::ThreadPool::builder().create().unwrap();

    ThreadPool { inner, rt }
});
[Design Inference and Architectural Trade-offs]

Here the Tokio runtime is created butis notblock_ondriven—it merely "exists," providing the I/O driver and timers. The actual task scheduling is handled byfutures::executor::ThreadPool. In this mode, Tokio's worker threads are essentially spinning idle (waiting for I/O events), and task execution happens in futures' thread pool.

Data flow: a TcpListener::bind's cross-executor journey

mermaid
sequenceDiagram
    participant App as "应用 (main)"
    participant FE as "futures::ThreadPool"
    participant TC as "TokioContext"
    participant TR as "tokio::Runtime (后台线程)"
    participant IO as "I/O 驱动 (mio)"

    App->>FE: spawn_ok(TokioContext::new(f, handle))
    FE->>TC: poll(cx)
    TC->>TC: enter(handle) 设置线程局部上下文
    TC->>TC: f.poll(cx) 执行 TcpListener::bind
    TC->>TR: 通过 Handle 访问 I/O 驱动
    TR->>IO: Registration::new 注册 fd
    IO-->>TR: 注册完成
    TR-->>TC: 返回 Pending 或 Ready
    TC-->>FE: 返回 poll 结果
    Note over FE,TR: I/O 就绪时,Tokio 驱动唤醒 waker<br/>FE 重新调度该任务

The key to this sequence diagram is:The task's poll happens in the futures thread pool, but the waiting for I/O events happens in Tokio's background threads. The two are connected throughHandleand the waker.

Design Thinking: When to Bypass Tokio

[Design Inference and Architectural Trade-offs]

The very existence of this example is a signal: Tokio's architecture allows "using only the I/O driver, not the scheduler." The criteria can be summarized in three points:

1. If you need to integrate with an existing executor ecosystem(e.g., some frameworks mandatefutures::executor), usingTokioContextis the least invasive approach.

2. If you need complete control over scheduling policy(e.g., real-time systems requiring deterministic scheduling), Tokio's work-stealing doesn't meet the requirements, but its I/O driver is still usable.

3. If you just find Tokio's API complicated, then you shouldn't bypass it—TokioContextthe cross-executor boundary introduced by

Production Pitfalls:TokioContextInblock_onmode, Tokio runtime'sRuntime::shutdownis never called, meaningRuntime's cleanup logic won't trigger automatically. You must explicitly drop

before the program exits, otherwise the I/O driver's background threads may not shut down gracefully.

Relationship with io_uring

tokio/src/runtime/io/mod.rs[Design Inference and Architectural Trade-offs]

📎 tokio/src/runtime/io/mod.rs:1-4

rust
#![cfg_attr(
    not(all(feature = "rt", feature = "net", feature = "io-uring", tokio_unstable)),
    allow(dead_code)
)]

reveals how io_uring is integrated:feature = "io-uring"Copytokio_unstableNote thatandappear simultaneously. This means io_uring support is currentlyallow(dead_code)experimentalallow, and both unstable features must be enabled to compile.

indicates: when these features are not enabled, some code in the module won't be used, and the compiler will warn—suppressed with

.read/write[Design Inference and Architectural Trade-offs]ScheduledIoThe fundamental difference between io_uring and epoll is: epoll is "readiness notification," io_uring is "completion notification." The former requires the application to issuereadiness()system calls itself, while the latter has the kernel directly complete I/O and return results. This is a huge shock to Tokio's

---

model—

's semantics no longer apply under io_uring, requiring an entirely new "submit-complete" abstraction. This is also why io_uring support remains unstable: it's not as simple as adding a backend, but rather reconstructing the entire I/O driver abstraction layer.

Chapter Summary:

  • This chapter reviewed Tokio's three core trade-offs from an architectural perspective, and looked ahead at three evolution paths:
  • Historical Trade-offsblock_onwork-stealing trades scheduling complexity for multi-core scalability, with the boundary being that task granularity can't be too fine;
  • I/O driver is independent of the scheduler, allowing

and the multi-threaded runtime to reuse the same I/O implementation;(reactor-refactor.md):

  • loom completely disappears in production builds through cfg conditions, only exhaustively enumerating thread interleavings during testing.ScheduledIoDriver Refactoring
  • Moving the waker from insideAtomicUsizeto the operation Future, using an intrusive linked list to support multiple waiters;clear_readinessusing
  • AsyncRead/AsyncWrite's bitfield layout (shutdown/generation/tick/readiness) to eliminatereader/writer's race conditions;

because poll semantics can't use intrusive linked lists, retaining:

  • fixed slots as a compromise.tokio_unstableFuture Evolution
  • TokioContextio_uring requires a new "submit-complete" abstraction, currently protected by
  • ;

allows using only the I/O driver without the scheduler, but requires manual management of the Runtime lifecycle;

The criterion for judging "extend or bypass": if it can be expressed with existing APIs, don't touch internal structures.ScheduledIoChapter Reflection and Self-TestreadinessQ1: Intick'sclear_readinessbitfield layout, if the

field is reduced from 8 bits to 4 bits, in what scenarios would errors be triggered? Please analyze in conjunction with:tick's tick matching logic.mio::poll()Reference Analysis📎 tokio/docs/reactor-refactor.md:185-185。clear_readinessincrementsevent.tick == 当前 readiness.tickon each📎 tokio/docs/reactor-refactor.md:199-199, and only clears readiness bits whenReadyEvent. If tick only has 4 bits, then it wraps around every 16 polls. Suppose a certainclear_readinessPreviously, mio polled once more, and the tick wrapped around to 0. At this pointclear_readinessdiscovers the tick mismatch (15 != 0) and will incorrectly skip the clear—but in reality, no new events may have arrived during this period; the tick simply wrapped around. This causes the readiness bit to be permanently retained, and subsequentlyreadiness()returns immediately butreadstillWouldBlock, falling into a busy loop. The 8-bit tick is sufficient under normal load (a read-clear cycle completes within 256 polls), but under extreme high concurrency there is still a wraparound risk—this is an inherent boundary of the bitfield layout.

Q2: examples/custom-executor.rs, the Tokio runtime is created but neverblock_on. Ifrt.shutdown_timeout()is called at this point, what happens? Why does this example choose not to call it?

Reference Analysis:rt.shutdown_timeout()will wait for all tasks to complete and shut down the I/O driver. But in this example, the tasks actually run onfutures::executor::ThreadPoolon📎 examples/custom-executor.rs:51-54, and there are no tasks in the Tokio runtime—it only provides the I/O driver. Ifshutdown_timeoutis called, it will return immediately (since there are no tasks), but the I/O driver's background thread may still be running. The example chooses not to call it becauseEXECUTORis aLazystatic variable, handled by Rust's static destructor mechanism when the program exits. The real pitfall is: if the Future wrapped byTokioContextis still running andRuntimeis dropped, then I/O operations inside the Future will panic (runtime context not found). Production environments must ensure allTokioContextFutures complete before dropping the Runtime.

Q3: Suppose you want to add an io_uring-based I/O backend to Tokio. Based onreactor-refactor.mdinreadiness()'s semantics, which parts can be directly reused, and which must be rewritten?

Reference Analysis: What can be directly reused isRegistration's registration interface andScheduledIo'swaiterslinked list structure—they manage "who is waiting," which is independent of whether the underlying layer is epoll or io_uring. What must be rewritten isreadiness()'s semantics: under epoll it returns "fd ready," while under io_uring there is no concept of "ready"—only "submitted SQE completed."clear_readiness's tick mechanism also needs to be redesigned—io_uring's completion events carry their own user_data identifier, so no tick is needed to distinguish new events from old ones. The most fundamental change is:readiness()'s returned Future under io_uring should become "submit SQE and wait for CQE," which means theWaiterstructure needs to carry SQE parameters, not justinterest. This is also why io_uring support is protected bytokio_unstable📎 tokio/src/runtime/io/mod.rs:1-4—it is not replacing the backend, but changing the I/O driver's abstraction contract.

At this point, we have completed the climb from concrete pitfalls to architectural trade-offs. Looking back at the entire book, from Future's lazy evaluation to scheduler fairness, from cancellation safety to shutdown ordering, and then to this chapter's io_uring and pluggable drivers, all discussions revolve around one core: clearly delineating state ownership at asynchronous boundaries. Tokio's architecture is not set in stone—io_uring's zero-copy I/O, the decoupling of the driver layer, and the opening of custom executor interfaces are all pushing it toward a more flexible and efficient direction. When you close this book, I hope what remains is not a pile of API usage, but a set of judgment: knowing when to trust the runtime, when to intervene at the lower level, and how to avoid those combinations that bite in production environments. The async Rust ecosystem is still growing rapidly, and keeping track of source code and official documentation is more important than remembering any conclusion.

To understand any complex project, all you really need is a good book

This "Tokio Source Code Deep Dive: From Future to Production-Grade Async Runtime" was automatically compiled by AiReadCode by scanning the official open-source repository. Whether facing a large open-source masterpiece with hundreds of thousands of lines or a complex enterprise internal project, you can generate an equally well-organized exclusive monograph with one click.

Free Download AiReadCode Client Browse More Open-Source Books →
🇨🇳 Chinese · 🇺🇸 EN · 🇯🇵 Japanese · 🇰🇷 한국어 · 🌐 Traditional Chinese · 🇪🇸 ES · 🇩🇪 DE · 🇫🇷 FR · 🇧🇷 PT · 🇷🇺 RU