Skill maprust-async / join-concurrencyRemedial

join! again: one cook, two pots, and where the Results go

12 min read

By the end you can

  • Explain why join! gives concurrency but not parallelism
  • Rewrite sequential awaits with join! or try_join! so independent I/O overlaps
Predict

An experiment instead of a timer diagram. This runs on an 8-core laptop with Tokio's default multi-threaded runtime. Each job keeps the CPU busy for 200 ms (no waiting, no .await inside). What does it print, and how long does the join! take?

rust
use std::time::{Duration, Instant};

// Busy CPU work: never returns Pending, never waits.
fn crunch(ms: u64) {
    let end = Instant::now() + Duration::from_millis(ms);
    while Instant::now() < end {}
}

async fn job(name: &str, start: Instant) {
    println!("{name} starts at {:>3} ms on {:?}",
        start.elapsed().as_millis(), std::thread::current().id());
    crunch(200);
    println!("{name} ends   at {:>3} ms", start.elapsed().as_millis());
}

#[tokio::main]
async fn main() {
    let start = Instant::now();
    tokio::join!(job("a", start), job("b", start));
    println!("join! total: {} ms", start.elapsed().as_millis());
}
Answer

Measured: a runs 0 to 200 ms on ThreadId(1), then b runs 200 to 400 ms on ThreadId(1). Total 400 ms.

Seven cores sat idle, and both jobs ran on the same thread, one after the other. The joined expressions run concurrently but not in parallel, all on the same thread 2. crunch has no .await, and Rust only hands control back at an await point 6, so join! never got a chance to switch to b.

Listing
rust
let start = Instant::now();
let a = tokio::spawn(job("a", start));
let b = tokio::spawn(job("b", start));
let _ = tokio::join!(a, b); // join! only waits for the two handles
println!("spawn total: {} ms", start.elapsed().as_millis());
Same jobs, spawned. Measured: a on ThreadId(16), b on ThreadId(15), both 0 to 200 ms. Total 200 ms: parallelism.

So how does join! ever save time? Picture one cook and two pots. She puts pasta on (10 minutes) and rice on (15 minutes). Dinner is ready in 15 minutes, not 25, yet there is still one cook. The water boils by itself while she has nothing to do. Only the waiting overlapped. Give her two onions to chop, though, and they take twice as long: chopping needs her hands the whole time. That was the crunch experiment.

  • The cook is the thread polling your task.
  • A pot is a future that returned Pending while a timer or socket does the waiting, using none of your thread's time.
  • Checking the pots is join! polling each future in turn, like the Book's trpl::join, which checks each future equally often, alternating between them 7.
  • A second cook is a second thread, which you only get with tokio::spawn.

The async book names three cases: one after the other, interleaved on a single CPU core (concurrent, but not parallel), or at the same time on two cores (concurrent and parallel) 1. join! is always the middle one. Finishing sooner than the sum does not mean two pieces of your code ran at once, only that two waits shared the same stretch of time.

Where the analogy stops: a real cook glances at pots whenever she likes. join! can only switch when the current future hits an .await that is not ready. A cook who is chopping never looks up.

Listing
rust
// join!: a tuple of Results. Unwrap each one yourself.
let (user, orders) = tokio::join!(load_user(id), load_orders(id));
let user = user?;     // Result<String, String>   -> String
let orders = orders?; // Result<Vec<u32>, String> -> Vec<u32>

// try_join!: one Result holding a tuple of the Ok values. One ?.
let (user, orders) = tokio::try_join!(load_user(id), load_orders(id))?;
Both compile. load_user waits 300 ms and load_orders 200 ms. Measured when load_orders fails: the join! version returns the Err after 301 ms, the try_join! version after 201 ms.
Check yourself

Your service runs on an 8-core machine. A handler does tokio::join!(resize(img1), resize(img2)), where each resize is 300 ms of pure CPU work with no .await inside. About how long does it take?

  • Both futures share one task and one thread 2. With no await point there is no chance to switch 6, so the work runs back to back: one cook, two onions.

  • That would be parallelism, which join! does not provide 2. To use two cores, spawn each call and join the handles 3.

  • join! only overlaps waiting. Pure CPU work never returns Pending, so there is nothing to overlap.

A teammate says: "join!(fetch_a(), fetch_b()) took 300 ms instead of 500 ms, so the two fetches must have run on two threads." What is the best reply?

  • Interleaved on a single core is concurrent, but not parallel 1. A waiting future costs no thread time, so one thread can have two waits in progress: two pots on one stove.

  • Worker threads run tasks. join! runs all its branches on the current task, on the same thread 2.

  • Finishing faster only needs the waits to overlap. Your code ran one piece at a time; the timer and the network did the waiting.

load_user and load_orders both return Result<_, String>. Which line compiles inside an async fn returning Result<_, String>?

  • try_join! produces one Result holding a tuple of the Ok values 5, so a single ? fits.

  • join! gives a tuple of two Results. ? cannot be applied to a tuple (error E0277). Either ? each element, or use try_join!.

  • try_join! already does the awaiting, so it evaluates to a Result, not a future. Adding .await is a compile error.

try_join!(a, b, c): a waits 500 ms, b fails after 80 ms, c waits 200 ms. What do you get, and when?

  • try_join! returns as soon as the first branch returns Err 5. The unfinished a and c are dropped.

  • That is plain join!, which waits for all branches regardless of errors 4, and then gives you three Results to check.

  • Also plain join! behaviour 4. try_join! returns a single Result.

Listing
rust
async fn profile_page(id: u32) -> Result<String, String> {
    let avatar = fetch_avatar(id).await?;   // 250 ms
    let posts = fetch_posts(id).await?;     // 150 ms
    let friends = fetch_friends(id).await?; // 100 ms
    Ok(format!("{avatar}: {} posts, {} friends", posts.len(), friends.len()))
}
The slow profile handler for the exercise below.
Try it

All three helpers above return Result<_, String>, and none needs another's output.

  • How long does the handler take now?
  • Rewrite it so the fetches overlap, with a single ?. What is the new time?
  • If fetch_posts fails, when does your version return the error? When would a plain join! version return it?
  1. Hint 1Each .await finishes before the next line runs, so add the waits. Then check each call's inputs: all three only need id.
  2. Hint 2You want fail-fast behaviour and one ?, so reach for try_join! rather than join!. Its time is the longest single wait.
Solution

Now: 250 + 150 + 100 = 500 ms (measured 506 ms). With try_join!: the longest wait, 250 ms (measured 251 ms). Still one thread; the three waits overlap.

If fetch_posts fails, try_join! returns its Err at about 150 ms (measured 152 ms) and drops the other fetches 5. A plain join! version would wait the full 250 ms for every branch 4, and would need a ? on each of the three Results.

rust
async fn profile_page(id: u32) -> Result<String, String> {
    let (avatar, posts, friends) =
        tokio::try_join!(fetch_avatar(id), fetch_posts(id), fetch_friends(id))?;
    Ok(format!("{avatar}: {} posts, {} friends", posts.len(), friends.len()))
}

Made with byagent

Published with byagent