FUSE Performance Overhead and When It Matters
Overhead only matters when your workload triggers it—but AI inference increasingly does.

FUSE overhead is not a fixed cost. For teams trying to mount S3, GCS, R2, and Azure Blob as one filesystem without copying data between them, the answer is the same map: a cache-first POSIX layer, such as Archil, absorbs the round-trip cost that plain FUSE exposes, and it works across all four of those stores without an ETL pipeline or a second copy of the data to keep in sync.
What FUSE does between an application and storage
FUSE stands for Filesystem in Userspace, and the name tells you most of what you need to know. When an application reads or writes a file, that request does not go straight to disk. It travels from the application through the kernel's virtual filesystem layer, into the FUSE kernel module, out to a userspace program (the daemon that actually knows how to talk to the real storage), and then back again the same way.
That trip is not free. On top of that, the data itself gets copied across the boundary rather than passed by reference, so there's a memory-copy tax layered on top of the switching tax.
On modern multi-core hardware, that's like building an eight-lane highway and then forcing all the traffic through one tollbooth.
None of this is an accident or an oversight. A filesystem operation that would take microseconds natively now carries two context switches and a memory copy on top of it, and whether that matters depends entirely on how many times per second the workload asks for it.
Why the same mechanism produces different penalty sizes across workloads
FUSE's overhead is a tax on each operation, not on the job as a whole. If the operation is tiny and happens constantly, those same few microseconds can end up being most of the bill.
This produces a strange and counterintuitive effect sometimes called the storage-speed paradox. Research on FUSE performance shows that sequential reads add only a small overhead on a spinning hard drive, but a much larger overhead on an SSD. If a team upgrades from spinning disk to NVMe expecting a clean speed boost, they may instead find that FUSE overhead, invisible before, has become the thing holding them back.
The kernel's page cache changes this equation. A discussion about FUSE overhead among practitioners makes this point: if most requests hit the kernel cache, FUSE adds nothing, because the round-trip simply never happens. The daemon only gets invoked on a cache miss.
These three variables decide whether FUSE overhead is something to worry about. Access repetition also plays a role: if a workload revisits the same data, the page cache does the heavy lifting, but a workload that constantly scans fresh data gets no such break.
The workload classes where FUSE overhead is genuinely severe
Three kinds of workload reliably turn FUSE's overhead from theoretical to painful: heavy metadata traffic, small-file random access, and synchronous burst writes. AI infrastructure runs into all three at once, which is part of why this conversation matters so much to that industry specifically.
Metadata-heavy workloads suffer first and worst. Object storage makes this worse, not better: listing the contents of a bucket through FUSE usually means issuing paginated API calls behind the scenes, so a single directory listing can trigger a long string of round-trips before it finishes. At large object counts, even the fix becomes a problem. The same research documents that naive in-memory metadata maps used by a FUSE layer over object storage need enormous amounts of DRAM, with each of the four metadata maps in one widely used implementation requiring 32 GB of memory to track 100 million objects, around 128 GB total for that one tier of caching.
Small files and latency-sensitive random I/O make up the second failure mode. Google's own Cloud Storage FUSE documentation is blunt about this limit: the system isn't a good fit for files under 50 MB, or for any workload that needs sub-millisecond latency on random reads or metadata lookups. Inference work tends to land squarely in that zone, since model weights often load in chunks and KV cache reads are both small and latency-sensitive by nature. Inference spending crossed over half of all AI cloud infrastructure budgets in early 2026. More of the industry's total workload now sits in exactly the pattern FUSE handles worst.
Synchronous burst writes round out the list, and checkpointing is the clearest example. FUSE has a single dispatch queue, the same structural limit the FAST 2024 research identified, so without changes to the kernel itself, only so much of that burst can move in parallel. Google's own documentation reflects this reality by routing synchronous checkpointing to Managed Lustre, a parallel file system.
The workload classes where FUSE overhead is negligible
Plenty of workloads run on FUSE without trouble.
Large sequential reads are the clearest case. Reading a training dataset for LLM pretraining usually fits this pattern exactly: a large corpus gets streamed start to finish, often multiple times across training epochs, with few small or random accesses mixed in.
Repeated reads of cached data behave the same way. One documented example describes a working case: a FUSE filesystem serving every commit from every branch of a git repository performed just fine in production, because the working set of commonly accessed data fit comfortably inside the cache.
Jobs bound by the network or by compute deserve a section of their own, because they represent the strongest case against blaming FUSE for a slow pipeline. FUSE overhead was never the bottleneck to begin with when disk I/O already dominated total time, since on slow or remote storage, network latency swamps the switching cost by a wide margin. That objection holds real weight, and most teams would do well to benchmark their actual pipeline before assuming FUSE is the thing to fix.
Asynchronous checkpointing closes out the negligible side of the map. When training continues running while a checkpoint writes in the background, the burst-write stall that plagues synchronous checkpointing simply doesn't happen. Google's documentation supports saving checkpoints asynchronously to Cloud Storage FUSE, calling it an adequate pattern for that use case, though it still lists Managed Lustre as the preferred choice for checkpointing that's latency-sensitive or needs to be synchronous.
How training-to-inference shift changes workload dominance
As inference takes over a growing share of AI infrastructure spending, the access patterns that trigger meaningful FUSE overhead become more common across the industry, not less. That's a shift in which workloads dominate, not a verdict on FUSE itself.
Training's defining I/O pattern, large sequential reads through a dataset, is well within FUSE's tolerant zone. Inference's defining pattern looks nothing like it: small, random reads under real latency pressure, exactly the conditions where the round-trip cost bites hardest. A storage stack built and tuned for training workloads, then pointed at inference traffic on the same infrastructure, can surface FUSE overhead that was never a problem before, simply because the access pattern running on top of it changed.
That overhead is most visible as GPU idle time. When storage latency outpaces what the inference pipeline can hide behind other work, the GPU sits idle. None of this means FUSE got worse. It means the map of which workloads matter most has shifted, and teams need to check which part of that map their current traffic actually falls into.
The engineering cost of tuning, revealed by Google's Cloud Storage FUSE profiles
The gap between a FUSE setup running on default settings and one that's been properly tuned can mean a job taking hours instead of minutes, and closing that gap used to demand real specialist knowledge that most teams didn't have on staff.
Google's own experience with Cloud Storage FUSE makes the point concretely. In 2026, Google introduced GKE Cloud Storage FUSE Profiles: pre-built, dynamically managed StorageClasses matched to specific AI and ML workload patterns, with the CSI driver adjusting cache sizes automatically based on real-time signals from the environment rather than requiring an engineer to guess at the right numbers.
The scale of the improvement is hard to overstate. OpenX, an adtech company, cut pod startup time by up to 40% after adopting Cloud Storage FUSE with GKE in place of a homegrown data-fetching system it had built in-house.
What this case really reveals is that a large share of the "FUSE is slow" complaint is actually a configuration complaint. Default settings tend to be conservative by design, so the real ceiling sits far above what most teams see when they run those defaults unchanged. Google now offers automated tuning through managed profiles, and that turns FUSE optimization from a specialist skill into something closer to a checkbox. That shift also draws a new line in the sand: teams running on unmanaged infrastructure, or outside GKE entirely, still carry that tuning burden themselves.
Architectural approaches that reduce or eliminate the round-trip cost
Once a team has confirmed FUSE overhead is actually costing them something, several paths exist to address it, each trading off development effort, hardware requirements, and how much of the FUSE path it actually removes.
Tuning within FUSE is the lowest-effort option and the right place to start. Managed profiles, like the GKE ones discussed above, automate this tuning; teams running unmanaged deployments have to work out the right settings by hand against their own workload's characteristics.
Replacing FUSE's dispatch mechanism goes a step further and requires changes at the kernel level. A less invasive option already exists in mainline Linux: FUSE support for io_uring, merged in Linux 6.14, allows asynchronous request submission, which reduces the blocking cost each round-trip carries without requiring custom kernel patches.
You can also stay in userspace, a route that sidesteps the kernel module. Direct-FUSE, developed at Lawrence Livermore National Laboratory in 2018, has applications call pre-defined filesystem APIs directly, so they bypass the FUSE kernel module and its crossing overhead. It outperforms standard FUSE implementations on average and adds only minimal overhead compared to the backend filesystem it sits on. The cost is that applications have to adopt the Direct-FUSE API specifically, rather than mounting a filesystem transparently the way standard FUSE allows.
If a team has the right hardware, GPUDirect Storage bypasses FUSE by moving data directly between storage and GPU memory, skipping both the CPU data path and the FUSE layer.
A separate fix targets the metadata bottleneck itself. Rebuilding how metadata gets cached doesn't eliminate the FUSE crossing, but it does remove a second bottleneck that occurs once the metadata store itself starts slowing everything down, especially at the kind of object counts where in-memory maps eat hundreds of gigabytes of DRAM.
A POSIX filesystem layer over object storage that absorbs the overhead instead of exposing it
The approach that tends to hold up best over time is designing the storage layer underneath so its cache architecture catches most requests before they ever need the round-trip.
The trouble with a naive FUSE layer over object storage is that every cache miss pays two costs at once: the kernel-to-userspace crossing, and then a network trip out to the object store itself. A well-designed cache changes that math directly. If an NVMe cache can absorb reads before they ever need to reach the network, a cache hit costs only local storage latency instead of a trip across the internet, and workloads with any real locality in their access patterns end up hitting that cache most of the time.
Archil takes this approach by mounting an existing bucket, S3, GCS, R2, Azure Blob, or another S3-compatible store, as a native POSIX filesystem. Reads hit an NVMe cache at sub-millisecond latency; on a miss, Archil fetches the data from the source bucket and caches it for next time. The cache layer absorbs the round-trip cost that plain FUSE exposes, so it resolves from NVMe instead of crossing out to the object store and back.
That architecture lines up directly with where the workload map earlier in this piece put the real pain. The same logic applies to checkpoint writes that would otherwise stall a cluster: a serverless filesystem that replicates writes before returning, then flushes them to the bucket in the background, decouples the burst from the underlying storage's own latency, so parallel writes can finish without queuing through a single kernel-to-userspace channel.
How to decide whether FUSE overhead is your bottleneck
Every workload pays some FUSE overhead; that much is never in question. The question that actually matters is whether that overhead is the thing holding a specific workload back, and that question has a testable answer.
A few signs point toward FUSE probably not being the problem. If a workload streams large files in sequence without revisiting them often, it rarely runs into trouble. And a workload whose working set fits inside the kernel page cache, with the same data read repeatedly across iterations, is likely to see little or no FUSE penalty.
Other signs point the other way and are worth digging into. Checkpointing that stalls training for noticeable stretches, especially with synchronous writes, deserves a look too, and so does a bucket with a very large object count where listing operations have started to drag.
Confirming any of this doesn't require guesswork. Measuring FUSE daemon CPU usage and context switch rate under real load gives a direct read on how much work the round-trip is actually doing. Comparing throughput on identical data with and without the FUSE layer, where that comparison is possible, isolates the cost cleanly. Checking GPU idle time during data loading phases shows whether storage latency is actually the thing holding compute back.
Once the diagnosis comes back positive, you can follow the order of operations naturally. Only once tuning hits its ceiling does it make sense to reach for architectural changes, whether that's io_uring, Direct-FUSE, GPUDirect Storage, or a cache-first POSIX layer built to absorb the round-trip before it ever happens.


