How the C# compiler rewrites an async method

 
 
  • Gérald Barré

This post is part of the series 'Async/await in C#'. Be sure to check out the rest of the blog posts of the series!

async and await are not runtime features up to .NET 10. They are a compiler rewrite. When you mark a method async, the C# compiler removes the method body, moves it into a generated type, and replaces the original method with a few lines that create and start that type. Everything else, the scheduling, the exception handling, the value that eventually lands in the returned Task, is ordinary code you could have written yourself.

This post takes a small async method and looks at exactly what the compiler produced for it. Nothing here is guesswork: the sample is a real project, and the listings come from decompiling the assembly it builds.

#The method we are going to decompile

Here is the method. It is short, but it contains everything the compiler has to deal with: a parameter, locals that are still needed after an await, two await expressions of different types, a loop, and a try/finally.

C#
public static async Task<int> SumPageSizesAsync(string[] pages)
{
    var total = 0;
    var stopwatch = Stopwatch.StartNew();
    try
    {
        foreach (var page in pages)
        {
            var content = await LoadAsync(page);
            total += content.Length;
        }

        await Task.Yield();
        return total;
    }
    finally
    {
        Console.WriteLine($"Loaded {pages.Length} pages in {stopwatch.ElapsedMilliseconds}ms");
    }
}

To see what the compiler did with it, build the project and decompile the assembly. That is three commands, and they work on any project:

PowerShell
dotnet tool install --global ilspycmd --version 10.1.0.8386
dotnet build --configuration Release
ilspycmd ./bin/Release/net10.0/AsyncStateMachineSample.dll -lv CSharp4 > decompiled.cs

The last one is where the interesting part is. -lv CSharp4 sets the language version ILSpy targets, and async/await did not exist in C# 4. Asking for that version stops the decompiler from recognising the state machine and folding it back into the async method it came from, which is what a decompiler normally tries to do. Without the flag you get your own source back; with it you get the generated code as it really is.

Swap -lv CSharp4 for --ilcode to dump the IL instead:

PowerShell
ilspycmd ./bin/Release/net10.0/AsyncStateMachineSample.dll --ilcode > decompiled.il

The sample ships a decompile.ps1 that runs those commands for both Debug and Release, since the two differ in a way we get to at the end of the post.

#What is left of the original method

The body is gone. This is the whole method after compilation:

C#
[AsyncStateMachine(typeof(<SumPageSizesAsync>d__0))]
public static Task<int> SumPageSizesAsync(string[] pages)
{
    <SumPageSizesAsync>d__0 stateMachine = default(<SumPageSizesAsync>d__0);
    stateMachine.<>t__builder = AsyncTaskMethodBuilder<int>.Create();
    stateMachine.pages = pages;
    stateMachine.<>1__state = -1;
    stateMachine.<>t__builder.Start(ref stateMachine);
    return stateMachine.<>t__builder.Task;
}

Five things happen, in order:

  1. A state machine value is created.
  2. It gets a builder, AsyncTaskMethodBuilder<int>. The builder owns the Task that callers will await, and it is the only thing that ever completes it.
  3. The parameters are copied into the state machine. They are fields now, because the body that reads them is somewhere else.
  4. The state is set to -1, which means "not started, not suspended".
  5. Start runs the body, and the Task is returned.

The generated names, <SumPageSizesAsync>d__0 and <>t__builder, are not valid C# identifiers. That is deliberate: they cannot collide with anything you write, and you cannot reference them by accident.

The important consequence is the one people get wrong most often: calling an async method runs its body. Start invokes MoveNext synchronously on the calling thread, and the body executes until it hits an await on something that has not finished yet. Only then does control come back to the caller.

#The generated type

Here are the fields of the state machine, before we look at what runs:

C#
private struct <SumPageSizesAsync>d__0 : IAsyncStateMachine
{
    public int <>1__state;
    public AsyncTaskMethodBuilder<int> <>t__builder;
    public string[] pages;
    private int <total>5__2;
    private Stopwatch <stopwatch>5__3;
    private string[] <>7__wrap3;
    private int <>7__wrap4;
    private TaskAwaiter<string> <>u__1;
    private YieldAwaitable.YieldAwaiter <>u__2;
}

Every field is there for a reason, and the naming is systematic:

FieldWhat it is
<>1__stateWhere to resume. -1 running, -2 finished, 0 and up: one value per await that can suspend.
<>t__builderThe builder, which owns the returned Task.
pagesA parameter, kept under its original name.
<total>5__2, <stopwatch>5__3Locals that are still needed after an await, kept under their original name.
<>7__wrap3, <>7__wrap4The foreach state: the array and the index. The compiler needs them across the await inside the loop.
<>u__1, <>u__2One slot per awaiter type, holding the awaiter while the method is suspended.

Two things follow from this list.

A local becomes a field only if it is still needed after an await. A local used entirely between two await expressions stays a normal local. This is why an async method that captures a lot of state is more expensive than one that does not, and why moving work into a separate non-async method can shrink the state machine.

The awaiter slots are per awaiter type, not per await. LoadAsync returns Task<string>, so its awaiter is a TaskAwaiter<string>; Task.Yield() returns a YieldAwaitable, whose awaiter is a different type. Two types, two fields. Ten awaits on Task<string> would still share <>u__1.

#MoveNext, the method body

MoveNext is the original body, rewritten so it can be entered more than once. This is LoadAsync, the simpler of the two methods in the sample, decompiled in full:

C#
private void MoveNext()
{
    int num = <>1__state;
    string result;
    try
    {
        TaskAwaiter awaiter;
        if (num != 0)
        {
            awaiter = Task.Delay(1).GetAwaiter();
            if (!awaiter.IsCompleted)
            {
                num = (<>1__state = 0);
                <>u__1 = awaiter;
                <>t__builder.AwaitUnsafeOnCompleted(ref awaiter, ref this);
                return;
            }
        }
        else
        {
            awaiter = <>u__1;
            <>u__1 = default(TaskAwaiter);
            num = (<>1__state = -1);
        }
        awaiter.GetResult();
        result = page;
    }
    catch (Exception exception)
    {
        <>1__state = -2;
        <>t__builder.SetException(exception);
        return;
    }
    <>1__state = -2;
    <>t__builder.SetResult(result);
}

Read it as two paths through the same code.

First call. <>1__state is -1, so num != 0 and the method runs the code up to the await. It gets the awaiter and asks IsCompleted. If the operation has already finished, the if body is skipped entirely: no suspension, no callback, no allocation, execution simply continues to GetResult(). This is the fast path, and it is why an async method whose awaits all complete synchronously costs so little.

If it has not finished, three things happen and then the method returns: the state is set to 0 so the next call knows where to resume, the awaiter is stored in <>u__1 so it survives, and AwaitUnsafeOnCompleted asks the builder to call MoveNext again when the operation completes.

Second call. <>1__state is 0, so the else branch runs. The awaiter is read back out of the field, the field is cleared so it does not keep the awaiter alive, the state goes back to -1, and execution continues at GetResult(), which is exactly where the first call stopped.

That is the whole trick. await is a switch over a field, plus a callback registration. Every call to MoveNext takes one of these paths:

Start

IsCompleted true

IsCompleted false

awaiter completes

SetResult

SetException

Running

Suspended

Completed

Faulted

Start

IsCompleted true

IsCompleted false

awaiter completes

SetResult

SetException

Running

Suspended

Completed

Faulted

Two more details in that listing are worth naming.

The try/catch around everything is not from my code. The compiler adds it so that an exception thrown anywhere in the body becomes SetException instead of escaping. This is why an exception in an async method never surfaces at the call site: it is captured and placed on the returned Task, and it only reaches you when you await it.

The finally block from my source did not disappear either. It is still in SumPageSizesAsync, nested inside the compiler's own try, and the state machine also nulls out <stopwatch>5__3 on the way out so a completed state machine does not keep objects alive.

#Struct in Release, class in Debug

The sample decompiles both configurations, and they differ on the first line of the generated type:

C#
private sealed class <SumPageSizesAsync>d__0 : IAsyncStateMachine  // Debug
private struct <SumPageSizesAsync>d__0 : IAsyncStateMachine        // Release

In Release the state machine is a struct. It lives on the stack, and if the method never suspends it is never copied anywhere: an async method that completes synchronously allocates nothing for its state machine. The moment it does suspend, the builder boxes it onto the heap, because the stack frame is about to go away.

In Debug it is a class, allocated up front on every call. That is slower, but it means the object survives for the lifetime of the method, which is what the debugger needs in order to show you the locals while you step through.

That difference alone makes Debug measurements of async code misleading. Benchmark in Release.

#What it costs

The benchmark sums ten values three ways: a plain synchronous loop, an async method whose awaits all complete synchronously, and an async method that suspends on every iteration.

MethodMeanErrorStdDevRatioAllocated
Sync2.720 ns0.0746 ns0.1246 ns1.00-
AsyncCompletedSynchronously47.867 ns0.9465 ns1.2307 ns17.63216 B
AsyncSuspending8,820.771 ns249.2825 ns690.7606 ns3,249.651081 B
Benchmark project, measured with BenchmarkDotNet

The absolute numbers are specific to this machine, but the shape is the point.

Going async at all is not free even when nothing suspends: 216 bytes here, for the Task<int> objects the inner method returns. Suspending is a different order of magnitude, roughly 180 times slower than the synchronous-completion case, because each suspension means boxing the state machine, registering a continuation, and getting scheduled back onto a thread pool thread.

None of that is an argument against async. Ten additions are the wrong workload for it. It does say something useful though: async pays off when the thing you await is genuinely slow, and the overhead is real when it is not. A method that awaits something already in a cache goes through the fast path and costs little. A method that suspends on every element of a hot loop is worth restructuring.

#What to take away

  • The compiler moves the body of an async method into a generated type and leaves behind a stub that creates it, starts it, and returns the builder's Task.
  • Calling an async method runs its body synchronously until the first await that has not already completed.
  • Locals that outlive an await become fields. Locals that do not stay locals.
  • await compiles to: get the awaiter, check IsCompleted, continue inline if it is done, otherwise save the state and register a callback.
  • Exceptions are captured by the compiler's try/catch and put on the Task, which is why they surface at the await and not at the call.
  • In Release the state machine is a struct that only reaches the heap if the method suspends. In Debug it is always a class.

You now know enough to write the other end of that machinery yourself. The next post does exactly that: a thread pool and a Task built from nothing, in about a hundred and fifty lines.

#Additional resources

Do you have a question or a suggestion about this post? Contact me!

Follow me:
Enjoy this blog?