Why async exists: the cost of blocking a thread

 
 
  • Gérald Barré

This is the first post of a series about async/await in C#. Before looking at how it works, it is worth being precise about what problem it solves, because the usual explanations are either vague ("it makes your app faster") or wrong ("it runs your code in the background").

async/await does not make anything faster. A network call takes as long as it takes. What it changes is what your process does while waiting, and that turns out to matter a great deal.

#Waiting is not work

Consider a method that calls a web service and returns the response length:

C#
public static int GetLength(string url)
{
    using var client = new HttpClient();
    var response = client.Send(new HttpRequestMessage(HttpMethod.Get, url));
    return (int)response.Content.Headers.ContentLength.GetValueOrDefault();
}

While that request is in flight, the calling thread is inside a blocking call. It is not computing anything. It is not going to be woken up by anything except the network. It is parked, and the CPU is free to run something else.

But the thread is not free. It still exists, it still owns its stack, and, crucially, if it came from the thread pool it is no longer available to run other work.

That last point is the whole story.

Blocking

Async

Take a pool thread

Start the I/O

Return the thread to the pool

I/O completes

Continue on any pool thread

Take a pool thread

Start the I/O

Thread waits and does nothing

I/O completes

Continue on the same thread

Blocking

Async

Take a pool thread

Start the I/O

Return the thread to the pool

I/O completes

Continue on any pool thread

Take a pool thread

Start the I/O

Thread waits and does nothing

I/O completes

Continue on the same thread

In the blocking version, the thread is occupied for the whole duration of the I/O. In the async version, the thread is handed back as soon as the operation is started, and a thread is only needed again once the result is available. Nothing is running in the background in between.

#What a thread costs

The usual argument against threads is memory. A thread reserves a megabyte of stack address space by default, so ten thousand threads reserve ten gigabytes. That is true, and on 64-bit it matters less than it sounds because the memory is reserved rather than committed.

The cost that bites first is different, and it comes from the thread pool.

The .NET thread pool starts with a number of worker threads equal to the processor count. When every one of them is busy and more work is queued, it does not immediately create more. It adds threads gradually, using a hill-climbing algorithm that watches throughput to decide whether more threads actually help. That is the right behaviour for CPU-bound work, where extra threads past the core count only add context switching. It is exactly the wrong behaviour for threads that are blocked on I/O, because they are not consuming CPU and no amount of measuring throughput will reveal that adding threads would fix things.

The result is that blocked pool threads do not just waste memory. They throttle everything else the process wants to do.

#What that looks like

The sample queues 1000 operations that each wait 50 ms, once by blocking a thread and once by awaiting:

C#
// blocking: Thread.Sleep holds on to the thread for the whole wait
tasks[i] = Task.Run(() => Thread.Sleep(operationDuration));

// async: Task.Delay holds on to nothing
tasks[i] = Task.Delay(operationDuration);

On a 15-core machine:

Processors: 15
Thread pool minimum worker threads: 15

blocking  1000 operations of 50 ms
          elapsed:           3307 ms
          peak thread count: 16

async     1000 operations of 50 ms
          elapsed:           53 ms
          peak thread count: 16

Both versions describe the same 50 ms of waiting, 1000 times over, and both could in principle finish in a little over 50 ms. The async version does. The blocking version takes 62 times longer.

Look at the peak thread count: it is the same in both runs. This is the part that surprises people. The pool did not spin up 1000 threads to service the blocking work. It kept roughly one thread per core and ran the operations in waves of about 16, so 1000 operations of 50 ms became 63 sequential rounds. The concurrency you asked for silently became almost none.

#It only pays off under concurrency

That 62x number is not a general claim about async. The benchmark runs the same comparison at three concurrency levels:

MethodConcurrencyMeanRatioAllocatedAlloc Ratio
Blocking1051.08 ms1.00944 B1.00
Async1050.94 ms1.002024 B2.14
Blocking100244.78 ms1.048144 B1.00
Async10051.01 ms0.2217864 B2.19
Blocking1000720.64 ms1.0680144 B1.00
Async100051.78 ms0.08176264 B2.20
Benchmark project, measured with BenchmarkDotNet

At a concurrency of 10, on a machine with 15 cores, the two are identical. There are enough threads for everyone, so blocking them costs nothing measurable. Async is not faster here, and it allocates roughly twice as much.

At 100 the blocking version is already about 5 times slower, and at 1000 it is about 12 times slower. The async version takes the same 51 ms regardless, because it never needed a thread per operation in the first place.

So async is not a performance optimisation in the usual sense. It is a scalability one:

  • It does not reduce latency for a single operation. If anything it adds a little.
  • It does not reduce allocations. It adds some.
  • It removes the coupling between the number of operations in flight and the number of threads you need.

For a desktop application loading three files at startup, that trade is not obviously worth it, and sync I/O can be the better choice. For a server handling a thousand concurrent requests, it is the difference between working and not working.

#The other reason: responsiveness

There is a second, unrelated reason async exists, and it applies to applications with a UI.

WinForms, WPF and MAUI process their message loop on a single dedicated thread. Blocking that thread stops the message loop, which means the window stops repainting and stops responding to input. There is only ever one such thread, so "just use another one" is not an option: the work has to leave the UI thread, and the result has to come back to it.

async/await handles both halves of that, and it does it without callbacks. That is the part covered later in the series, when we get to SynchronizationContext.

#What to take away

  • Blocking a thread while waiting for I/O buys nothing. The thread does no work and cannot be reused.
  • The .NET thread pool deliberately grows slowly, so blocked pool threads throttle the whole process rather than being replaced.
  • Under low concurrency, async gains nothing measurable and costs a few allocations.
  • Under high concurrency, it is the difference between 51 ms and 3.3 seconds for the same work.
  • Async is about scalability and responsiveness, not speed.

The next post looks at what actually happens underneath, and why "there is no thread" is literally true: how the operating system reports that an I/O operation finished, and what the .NET thread pool does with that.

#Additional resources

Do you have a question or a suggestion about this post? Contact me!

Follow me:
Enjoy this blog?