Skill maprust-async / join-concurrency

Running futures concurrently with join!

20 min read

By the end you can

  • Predict the total run time of two sequential .awaits versus the same futures in join!
  • Rewrite sequential awaits with join! or try_join! so independent I/O overlaps
  • Explain why join! gives concurrency but not parallelism
Warm-up

Review from async-await-basics

Your #[tokio::main] async fn main() calls a plain helper fn log_user() { let u = fetch_user().await; }. What happens?

  • .await is only allowed in async functions or blocks 10, checked per function body; rustc reports E0728: await is only allowed inside async functions and blocks. Make it async fn log_user() and call log_user().await.

  • Rust has no built-in runtime; it lets you choose one rather than providing one 12. .await only works inside async code 10, whose future a runtime like Tokio polls.

  • An async caller doesn't make the callee async; each body is checked on its own 10. Same as TypeScript: await in a non-async function is an error whoever calls it.

You delete #[tokio::main] from async fn main() and use only std. What happens, and why?

  • rustc reports E0752: main function is not allowed to be async. The reason is that async code needs a runtime, which main can start but is not itself 11. std has the Future trait but no executor.

  • std has no executor. Futures are inert until polled 6, and Rust leaves the runtime choice to you 12. That is why an async main is rejected 11.

  • Right that nothing would poll it 6, but the compiler stops earlier: Rust won't let main be async 10.

Roughly what does #[tokio::main] turn async fn main() { body } into?

  • The macro sets up the runtime for you 13; its documented equivalent is Builder::new_multi_thread().enable_all().build().unwrap().block_on(async { ... }) 14.

  • There is no compiler runtime 12. main can't be async, so the macro rewrites it into a sync main that starts Tokio 14.

  • The whole body becomes one async block driven by a single block_on 14. Inside it, each .await is a point where the task can pause and let the runtime do other work. Futures only overlap when you combine them, for example with join!, coming next; two .awaits in a row still run one after the other.

Predict

fetch_user takes 300 ms and fetch_orders takes 200 ms. Neither needs the other's result. Roughly how long does main take?

rust
use std::time::{Duration, Instant};
use tokio::time::sleep;

async fn fetch_user() -> String {
    sleep(Duration::from_millis(300)).await; // stands in for a network call
    "ana".to_string()
}

async fn fetch_orders() -> Vec<u32> {
    sleep(Duration::from_millis(200)).await;
    vec![1, 2, 3]
}

#[tokio::main]
async fn main() {
    let start = Instant::now();
    let user = fetch_user().await;
    let orders = fetch_orders().await;
    println!("{user} {orders:?} in {:?}", start.elapsed());
}
Answer

About 500 ms, the sum of the two. .await means "drive this future to completion before moving to the next line". fetch_orders() is not even created until fetch_user has finished, so the 200 ms wait cannot start early. This matches JavaScript: await a(); await b(); is sequential there too.

You already have the fix in your head from Node: Promise.all. In Tokio it is tokio::join!, which waits on multiple concurrent branches, returning when all branches complete 9. You list any number of futures, as long as the count is fixed when you write the code.

One difference from Promise.all: you don't put .await after it. The macro polls all the futures and does the awaiting for you, so the join! line evaluates straight to a tuple of results. That hidden .await is also why the join! macro must be used inside of async functions, closures, and blocks 9. Adding your own .await, as in tokio::join!(a(), b()).await, fails to compile with error E0277: the tuple of results is not a future.

If you read the Rust Book's chapter 17, it first uses trpl::join, a plain function. When you give it two futures, it produces a single new future whose output is a tuple containing the output of each future you passed in once they both complete 1. That one you do .await. Later it switches to the trpl::join! macro, which, like Tokio's, awaits an arbitrary number of futures where we know the number of futures at compile time 2, with no .await after it.

Listing
rust
#[tokio::main]
async fn main() {
    let start = Instant::now();
    let (user, orders) = tokio::join!(fetch_user(), fetch_orders());
    println!("{user} {orders:?} in {:?}", start.elapsed());
}
Same two calls, now overlapped, and no .await after join!. Measured: about 300 ms, the longer of the two.

The rule of thumb for timing:

  • Sequential .awaits: total time is the sum of the waits.
  • join!: total time is the longest single wait.

That rule holds when the futures spend their time waiting (network, disk, timers). The next sections show when it breaks.

Predict

A Node developer tries the JavaScript trick: create both first, await later. How long does this take?

rust
let start = Instant::now();
let u = fetch_user();   // create both futures first...
let o = fetch_orders();
let user = u.await;     // ...then await them
let orders = o.await;
println!("{:?}", start.elapsed());
Answer

Still about 500 ms. In JavaScript this trick works, because calling an async function starts it and hands you a Promise that is already running. In Rust, fetch_user() only builds a future. Futures alone are inert; they must be actively polled for the underlying computation to make progress 6. While u.await runs, nothing polls o, so its 200 ms timer has not even started. This is where the Promise.all analogy stops: in Rust, where you await decides what overlaps, not where you call.

Step through

What join! does with the two futures over time

  1. Step 1 of 4 · t = 0 ms

    join! polls each future once. fetch_user starts its 300 ms timer and returns Pending. join! does not stop at that Pending: it also polls fetch_orders, which starts its 200 ms timer and returns Pending too.

  2. Step 2 of 4 · Waiting

    Both futures are pending, so the .await hidden inside join! hands Pending back up, and the task's future (main's async body) returns Pending to Tokio. The task goes to sleep. Both timers are counting down during the same stretch of time. This is the overlap.

  3. Step 3 of 4 · t = 200 ms

    The orders timer fires and wakes the task. join! polls again. fetch_orders returns Ready(vec![1, 2, 3]), and join! stores that result. fetch_user is still Pending.

  4. Step 4 of 4 · t = 300 ms

    The user timer fires. join! polls fetch_user, which returns Ready("ana"). Every future is done, so the join! line evaluates to the tuple (user, orders) and main carries on to the next line.

Notice who did the work in that stepper: one task, and one thread polling it. join! never started a second thread. It gets its speed-up by switching to another future whenever the current one is waiting.

That is concurrency: several operations in progress at once, interleaved. Parallelism is different: parallelism is about computing on multiple processors (operations are parallel if they are literally happening at the same time) 3. The Tokio docs say it directly: by running all async expressions on the current task, the expressions are able to run concurrently but not in parallel 4.

For I/O this is all you need. A network request spends almost all of its time waiting for bytes, and a waiting future costs no CPU. Two waits can overlap on one thread just fine.

Your question
Predict

Same shape as before, but these futures use std::thread::sleep (a blocking call) instead of Tokio's sleep. How long does the join! take?

rust
let a = async { std::thread::sleep(Duration::from_millis(300)); "a" };
let b = async { std::thread::sleep(Duration::from_millis(200)); "b" };
let (x, y) = tokio::join!(a, b);
Answer

About 500 ms, not 300. std::thread::sleep freezes the whole thread. It does not return Pending, so join! gets no chance to switch to the other future. Both futures share that thread, so b waits for a to finish blocking: if one branch blocks the thread, all other expressions will be unable to continue 4. Blocking does not "only pause its own future". Each future is responsible for handing control back at .await points and for not blocking for long 7. Use tokio::time::sleep(...).await and the time drops back to 300 ms.

Real I/O calls return Result. With plain join! you get a tuple of Results, and every branch runs to the end even if one has already failed. try_join! is the Result-aware version: it waits on multiple concurrent branches, returning when all branches complete with Ok(_) or on the first Err(_) 5. On success you get a tuple of the Ok values, so one ? handles all the errors.

The Node comparison: try_join! behaves like Promise.all (fail fast), while join! is closer to Promise.allSettled (wait for everything). One difference: when try_join! returns early, the unfinished futures are dropped and stop running. A rejected Promise.all leaves the other promises running.

Listing
rust
let (user, orders, alerts) =
    tokio::try_join!(load_user(id), load_orders(id), load_alerts(id))?;
Each call returns Result<_, String>. One ? covers all three.
Listing
rust
let hits = std::sync::Mutex::new(0);
let one = async {
    sleep(Duration::from_millis(100)).await;
    *hits.lock().unwrap() += 1; // guard dropped at end of statement
};
let two = async {
    sleep(Duration::from_millis(50)).await;
    *hits.lock().unwrap() += 1;
};
tokio::join!(one, two);
println!("hits = {}", hits.lock().unwrap()); // hits = 2
Two joined futures bump one counter. Each lock is taken and released with no .await in between.
Worked example

Speeding up a dashboard handler

A handler loads a user (300 ms), their orders (200 ms) and their alert count (100 ms), then the user's team (150 ms), which needs user.team_id. Written with plain awaits it takes 750 ms. Make it as fast as possible without changing the helpers.

  1. Start from the sequential version. Add up the waits: 300 + 200 + 100 + 150 = 750 ms.

    rust
    async fn dashboard(id: u32) -> Result<String, String> {
        let user = load_user(id).await?;
        let orders = load_orders(id).await?;
        let alerts = load_alerts(id).await?;
        let team = load_team(user.team_id).await?;
        Ok(format!("{} ({team}): {} orders, {alerts} alerts", user.name, orders.len()))
    }
  2. Find the dependencies. Look at what each call takes as input. load_user, load_orders and load_alerts only need id, so they are independent. load_team needs user.team_id, so it must wait for load_user.

  3. Group the independent calls in one try_join!. They all return Result<_, String>, so a single ? handles them. This group now takes max(300, 200, 100) = 300 ms.

    rust
    let (user, orders, alerts) =
        tokio::try_join!(load_user(id), load_orders(id), load_alerts(id))?;
  4. Keep the dependent call as a normal .await after the group. Total: 300 + 150 = 450 ms, down from 750 ms.

    rust
    async fn dashboard(id: u32) -> Result<String, String> {
        let (user, orders, alerts) =
            tokio::try_join!(load_user(id), load_orders(id), load_alerts(id))?;
        let team = load_team(user.team_id).await?;
        Ok(format!("{} ({team}): {} orders, {alerts} alerts", user.name, orders.len()))
    }
  5. Check the failure case. If load_orders returns Err at 200 ms, try_join! returns that error at 200 ms. It does not wait the extra 100 ms for load_user, and the unfinished load_user future is dropped.

Takeaway

List each call's inputs. Calls that don't depend on each other go in one join! or try_join!. A call that needs an earlier result stays a plain .await after that result is available. Time = the longest wait in each group, summed across the groups.

Check yourself

a() waits 400 ms on the network and b() waits 100 ms. How long does let x = a().await; let y = b().await; take?

  • Each .await finishes before the next line runs, so the waits add up.

  • That is the join! time. Two awaits on adjacent lines do not overlap; b() is not even created until a is done.

  • Neither form finishes faster than its longest wait. Even join! must wait for a's 400 ms.

Why does tokio::join!(a(), b()) finish sooner than awaiting them one after the other, when both are network calls?

  • This is concurrency on one task: both futures are in progress, and their waiting overlaps 4.

  • All branches of join! run on the current task, on one thread 4. For separate threads you would tokio::spawn each one.

  • That is how JavaScript Promises work. Rust futures are inert until polled 6; join! starts polling them when the join! line itself runs, not when the futures are created.

Two joined futures each do 300 ms of heavy CPU hashing, with no .await inside. About how long does the join! take?

  • They share one thread and never yield, so the work runs back to back. join! gives concurrency, not parallelism.

  • That would need both to compute at the same time on two cores, which is parallelism 3. join! does not provide that.

  • Extra cores don't help: the futures in one join! are polled by one thread. Extra cores only help when the work is split into separate tasks or threads (for example tokio::spawn, or spawn_blocking for CPU-heavy work).

try_join!(a, b, c) where b returns Err after 50 ms, a would take 500 ms and c would take 300 ms. What happens?

  • try_join! returns on the first Err 5, and the futures it was still polling are dropped.

  • Plain join! would wait the full 500 ms and hand you all three Results; try_join! returns on the first Err 5.

  • try_join! only returns Ok when every branch succeeds.

Listing
rust
async fn checkout(session: u32, user: u32) -> Result<String, String> {
    let cart = get_cart(session).await?;         // 100 ms
    let address = load_address(user).await?;     // 250 ms
    let total = price_items(&cart).await?;       // 150 ms
    let in_stock = check_stock(&cart).await?;    // 200 ms
    Ok(format!("{} items, total {total}, in stock: {in_stock}, ship to {address}",
        cart.items.len()))
}
The slow checkout handler for the exercise below.
Try it

The checkout handler above is too slow. All four helpers return Result<_, String>, and the comments show how long each one waits.

  • First, say how long the current version takes.
  • Then rewrite it to be as fast as you can, and state the new total time.
  1. Hint 1The current version takes 100 + 250 + 150 + 200 = 700 ms. Now list each call's inputs: which calls need cart, and which don't?
  2. Hint 2get_cart and load_address are independent, so they can share a try_join!. price_items and check_stock both need cart but not each other, so they can share a second try_join!. That gets you 250 + 200 = 450 ms.
  3. Hint 3To go further: pricing and stock checks only need the cart, not the address. Put "get cart, then price and check stock" in its own async { ... } block, and join that block with load_address. Inside the block, end with Ok::<_, String>(...) so the compiler knows the error type for ?.
Solution

The original takes 700 ms. Two try_join! groups get it to 450 ms: try_join!(get_cart, load_address) takes 250 ms, then try_join!(price_items, check_stock) takes 200 ms.

The best version takes 300 ms. The cart branch takes 100 + max(150, 200) = 300 ms, and it runs alongside the 250 ms address load. The whole handler is still one task on one thread; it is just never waiting on one thing when it could be waiting on two.

rust
async fn checkout(session: u32, user: u32) -> Result<String, String> {
    // Branch 1: cart, then pricing + stock with overlapping waits (300 ms)
    let cart_work = async {
        let cart = get_cart(session).await?;
        let (total, in_stock) =
            tokio::try_join!(price_items(&cart), check_stock(&cart))?;
        Ok::<_, String>((cart, total, in_stock))
    };

    // Branch 2 runs alongside branch 1 (250 ms)
    let ((cart, total, in_stock), address) =
        tokio::try_join!(cart_work, load_address(user))?;

    Ok(format!("{} items, total {total}, in stock: {in_stock}, ship to {address}",
        cart.items.len()))
}

Made with byagent

Published with byagent