Utopia Tech
Engineering5 min read

Introducing on-demand CPU and memory profiling with flamegraphs for Workers and Durable Objects

Understanding why an application is using more resources than expected, whether that is CPU or memory, can be challenging. Logs and aggregate metrics can only get you so far. Thankfully there is a better way: CPU or memory profiling can show you the exact function where your application is using CPU or allocating memory. Today we are happy to announce support for CPU and memory

UT

Utopia Tech

October 9, 2026 · 5 min read

Share

Understanding why an application is using more resources than expected, whether that is CPU or memory, can be challenging. Logs and aggregate metrics can only get you so far. Thankfully there is a better way: CPU or memory profiling can show you the exact function where your application is using CPU or allocating memory.

Today we are happy to announce support for CPU and memory profiling of Workers and Durable Objects. From the Workers Observability page, you can now request an on-demand CPU or memory profile of an active Worker, inspect it as an interactive flamegraph, and download the profile file for further analysis. This method of profiling gives you a useful perspective into what your code is doing and how you can improve it in a real-world setting.

That’s because the best way to understand your application is to profile it in production. To try this on one of your Workers, you have a choice of using the CLI or the Cloudflare Dashboard. To use the CLI, make sure you have the cf package installed , then simply run: To use the dashboard, head to the Cloudflare dashboard and then access the list of Workers on your account by navigating through Build → Compute → Workers & Pages.

Select your Worker and navigate to its Observability tab. You can then use the drop down to select “Flamegraph”: You can then request both CPU and memory profiles for your Worker on this page. The duration determines how long the profiler should run for on your Worker.

You can also select different versions of your Worker to be profiled. Your Worker needs plenty of traffic to be successfully profiled, so make sure you choose a version that has enough traffic. Once you click the Capture Profile button, the profiler will run for the duration that you have selected, and you will then see a flamegraph representing the results of the profile being rendered.

Each rectangle in the flamegraph represents a function call, with the width representing the amount of CPU time or memory being used by that function. You can click on different functions to focus on and hover to see more details about them: Don’t be afraid to capture a couple of profiles, then look for the widest functions to determine what is using the most resources in your Worker.

You might also find that the table view is useful to get a quick sense of which functions are most commonly seen in the profile. If your Worker is implemented in TypeScript, then you should make sure that you have source maps enabled for your project , otherwise your profile may show obfuscated function names that aren’t easy to understand. We have already seen multiple teams at Cloudflare using CPU and memory profiling to identify opportunities for optimizing memory and CPU usage, fixing out-of-memory (OOM) errors, and improving performance.

Below are some examples of this. Using CPU profiling to find optimizations Let’s look at a real example of how you can use CPU profiling to optimize your Workers. We will look at a CPU profile of the Worker that implements the R2 binding.

This Worker receives a lot of traffic and so eliminating even the smallest amount of wasted CPU cycles can have a massive impact on its performance. We set up the profiler to capture CPU traces, with a duration of 50 seconds. Once we received the rendered profile, we saw this: As mentioned previously, the widest boxes are those which are using up most of the CPU time.

Some of them aren’t surprising, for example decryptBlock is R2 decrypting object data on the way out, and fillResponse is just moving bytes. There weren’t any clear ways to optimize these functions. As some of the boxes are less visible, it can be a good idea to take a look at the table view.

Sorting by Samples points us at genericR2JsonReplacer : This function represented over 5% of the CPU time, but what caught our attention was the function calling itself recursively. This looked like wasted work, and it was. Even though the replacer was being called by JSON.

stringify on every node in the JSON tree, the replacer itself was also walking the tree. This meant that a value nested five levels deep was processed five times. Fixing this enabled us to eliminate the duplicate work and make the genericR2JsonReplacer function 2.

7x faster. The second candidate in the profile that we looked at was a duplicate call to metrics . When looking at the code, we saw something close to this: The metrics call was heavy enough that just one call represented 1% of the CPU time in the profile.

So storing the result of the first call in a variable and then not calling it again saved us a fair chunk of CPU time. How we used Worker profiling internally to fix memory problems Internally we had a Worker with a memory problem. P999 memory sat around 133 MB against the 128 MB Worker memory limit.

This caused the Worker to be evicted frequently with an “Exceeded Memory” error. The image above shows a screenshot of the “Errors by invocation status” graph, with the number of “Exceeded Memory” errors clearly dropping as a result of the fixes identified through profiling. What was hard to work out was what was causing these errors.

The errors and metrics pointed at memory as the source of the problem, but it didn’t explain which part of the code caused it. The team took a heap profile of one of the Workers running in production and opened it with pprof . They saw that their Prometheus code accounted for roughly 66.

7% of allocations in the profile. This code was supposed to be disabled, but enough of it was still running that it created a problem. Because it was instrumenting a lot of the code paths in the Worker, it ended up using a lot of its memory, even though the collected data never actually left the Worker.

This code wasn’t suspected as a culprit of the memory issues, because it was thought to have been disabled. The profiler showed that the code was only partially disabled, with the memory cost still being paid as if it was fully enabled.

Originally published at blog.cloudflare.com

Share
▸ Want a deeper look?

Talk to an architect about applying this to your stack.

60-minute technical evaluation, no obligation. We'll map the ideas in this article to your environment.

Skip to main content