EBS vs EFS vs S3 Performance Benchmarks for ML Workloads
Pick the right storage layer for each phase of training, not just the fastest option overall.

ML teams routinely pick storage by headline latency or IOPS numbers alone, then discover those numbers are irrelevant to the phase of the workload that is actually stalling their jobs. Then the training job stalls anyway, and nobody understands why. The numbers weren't wrong. They just answered a question nobody was actually asking.
EBS, EFS, and S3 are block, file, and object storage, three different architectures built to solve three different access problems. Ranking them on a single axis, like latency or throughput, produces a leaderboard that means nothing once real workloads hit it. SquareOps' 2025 AWS storage comparison guide put a number on the consequence: companies routinely overspend or underperform because they picked the wrong service for the job in front of them. Teams that are otherwise careful make this mistake regularly. It's a systematic pattern that occurs when storage gets chosen before the workload gets understood.
The fix is to stop asking "which is faster" and start asking "which bottleneck is actually killing this phase of the job." Teams whose ML infrastructure spans multiple clouds or requires on-premises data residency need more than AWS-native tier selection, because the storage layer itself needs to be cloud-agnostic. Each one stresses storage differently, and each one has a different right answer.
What each service is and where its performance profile breaks
Before any benchmark number means anything, it helps to know what each service was actually built to do. EBS, EFS, and S3 diverge first in how you're allowed to access them, and only second in how fast that access turns out to be. Get the access model straight and the performance numbers stop looking like magic and start looking like consequences.
EBS is block storage. It delivers sub-2ms latency with IOPS ceilings that reach up into the hundreds of thousands on io2 Block Express and gp3, a genuinely fast disk. One disk serves one machine. Share it across a training cluster and you're stuck building your own data-staging plumbing to get copies where they need to go.
EFS is a managed NFS file system. Many instances can mount it at once, it scales out to petabyte capacity without manual intervention, and small-file reads come back in under a millisecond with plenty of throughput and IOPS headroom to spare. Great for shared access. Expensive if you try to park a whole dataset there permanently.
S3 is object storage with no real ceiling on scale and eleven nines of durability, but latency runs 50 to 150 milliseconds for the Standard tier. S3 Express One Zone knocks that down to single digits. What actually caps S3's usefulness during training isn't total capacity, it's throughput per bucket and the API request-rate limits enforced per prefix, which put a hard ceiling on how many reads and writes can land in the same place per second.
The latency hierarchy runs instance store, then EBS, then S3 Express One Zone at 5 to 10 milliseconds, then S3 Standard at 50 to 150 milliseconds or slower, with EFS Standard's sub-millisecond reads sitting closer to EBS than to S3 Express One Zone. EFS doesn't slot in neatly between EBS and S3 Express One Zone the way people assume. Its sub-millisecond small-file reads put it closer to EBS territory than to S3 Express One Zone. None of these numbers is "the right one" in isolation. The right one depends entirely on which phase of the job is asking.
Dataset ingestion: where S3's throughput ceiling becomes the first bottleneck
Every training run starts the same way: pulling raw data off disk before a single forward pass can run. S3 is the correct long-term home for that data, but reading straight from it during active training is where the trouble starts.
The mechanism is straightforward once you see it. Pointing a few hundred GPUs at the same bucket makes them collectively hammer S3's per-prefix GET limits. Training jobs get throttled by exactly that mechanism, not by a lack of network bandwidth or a slow disk, but by a rate limit baked into how S3 organizes requests. Add egress cost on top: repeatedly pulling large datasets from S3 at standard AWS egress rates adds up fast, and the breakeven point for staging that data on local NVMe instead typically arrives within a handful of training iterations.
None of this means S3 is the wrong place to store a dataset. It means S3 is the wrong pipe to train directly against at scale. The established mitigation combines prefix sharding, which spreads objects across multiple prefixes to multiply the effective request rate ceiling, with multi-threaded prefetch pipelines, keeping S3 as the source of truth while relieving the hot path. SquareOps' AWS Storage Service Guide, published in January 2026, calls S3 "the default storage layer for training data". S3 is the settled choice as the data lake. What's left is how a training job reads from that lake without choking on it.
S3 Express One Zone, wired into SageMaker Model Training, helps narrow that gap. It delivers consistent single-digit millisecond request latency, which takes real weight off ingestion-heavy workloads. It leaves the per-prefix request ceiling in place. It just makes each individual request against that ceiling faster.
Checkpoint I/O: the phase where storage becomes the largest driver of idle compute time
Checkpointing is where storage choices turn directly into wasted GPU-hours, and the arithmetic is unforgiving. Saving a large model requires bandwidth that neither S3 Standard nor a stock EFS setup can deliver without stretching the write across minutes, and every one of those minutes is a minute of idle GPUs; at 8×H100 scale, S3 Standard's latency profile stretching multi-GB checkpoint writes this way leaves compute idle for a severe fraction of wall-clock training time without local staging.
Industry practice sets a target here: checkpoint overlap, the share of total training time eaten by checkpointing, should stay under ten percent. A training budget can't absorb checkpointing as a rounding error, so it has to stay under that budget.
EBS handles single-node checkpoint writes well: its IOPS ceiling and sub-2ms latency absorb the write burst without flinching. It just can't serve as the checkpoint store for a distributed run on its own, because nothing about it lets multiple nodes coordinate a shared write without an extra synchronization layer bolted on. EFS solves the multi-node access side of that problem, but its cost per GB-month rules it out as a permanent, petabyte-scale checkpoint archive. It works fine as a hot tier for checkpoints if the volume stays bounded.
FSx for Lustre, linked to S3, is built for exactly this phase. It now delivers over 1 TB/s of throughput with sub-millisecond latency, and connecting it to an S3 bucket produced an 83% performance improvement in subsequent training runs in an AWS 2025 benchmark. Shell's result at re:Invent 2025 is the clearest case: GPU utilization rose from below ninety percent to full utilization after moving checkpoint I/O to FSx for Lustre with EFA support, which bypasses the OS layer entirely. It's a significant tuning win. Paying for GPUs that compute instead of GPUs that wait comes down to this.
Distributed training reads: why per-node storage architecture determines cluster efficiency
Once a job moves past ingestion and checkpointing, it settles into its main phase: many nodes, reading data in parallel, continuously, for hours or days. That's a different stress test than either of the phases before it, and it exposes a limit that ingestion never touched directly.
EBS's single-instance attachment disqualifies it outright here. Sharing a volume across a cluster means duplicating data node by node, which is expensive to store and fragile to keep in sync. EFS solves the shared-access half of the problem cleanly, and its throughput scales with the size of the file system, but the cost, many times higher per GB than S3, pushes most teams to treat it as a hot tier during active runs rather than the permanent home for the dataset.
S3 stays the data lake of record, but its per-prefix request limits mean a large GPU cluster hammering one prefix runs out of API headroom before it runs out of network bandwidth. Prefix sharding remains the standard answer, though it requires organizing the dataset that way from the moment it's written, not bolted on after the fact. Meta FAIR's result with S3 Express One Zone, shown at re:Invent 2025, sustained very high aggregate transaction rates and proved that this ceiling isn't fixed. It moves depending on which S3 tier is in use and how the namespace is laid out.
The pattern that emerges across these sources is consistent: S3 holds the data permanently, FSx for Lustre or a similar NVMe-cached layer sits in front of it as the hot read tier, and training nodes pull from that local cache at close to memory bandwidth. In practical terms, that's the gap between GPUs running at full tilt and GPUs sitting there waiting for data that hasn't arrived yet.
Inference serving: the phase where latency consistency matters more than peak throughput
The three services aren't alternatives to each other: they're block, file, and object storage built for structurally different access patterns, so comparing them on a single axis produces a ranking that's meaningless in practice. Training rewards raw throughput. Serving rewards consistency: every request needs the model weights loaded fast and predictably, not in bulk and not eventually. That single shift in priority is what makes EBS the natural choice for hot-path model weights, with S3 taking over as the archive for models that get served rarely.
On a single inference server, EBS's sub-2ms latency and deep IOPS headroom load model weights without the tail-latency spikes that S3 Standard's 50 to 150 millisecond range would introduce the moment a cache misses. A single slow load might not matter for a training epoch running for hours. It matters a lot when a user is waiting on a response.
EFS earns its place when a fleet of inference instances all need access to the same model file. Shared-access semantics make that setup workable, as long as the per-GB cost pencils out against the size of the model being served. S3 Express One Zone, again through its SageMaker integration, narrows the latency gap enough to be a real option for model loading from object storage, provided the workload can tolerate occasional object-storage latency rather than the tighter consistency block storage guarantees.
The cost crossover here comes down to how often a model gets called. Frequently-served models justify EBS's higher per-GB price with the latency predictability it buys. Models served rarely, or archived at large scale, are better parked in S3 Standard-IA or Glacier tiers, which cut storage cost without touching hot-path performance at all.
How recent AWS releases shift the cost-performance frontier
Two releases from 2026 move the numbers behind this framework without changing the framework itself. They shift where the cost crossovers land in each phase. They don't erase the phases or the logic connecting them to storage type.
S3 Files, which reached general availability in April 2026, lets S3 buckets mount as NFS at S3 pricing instead of EFS throughput fees. A mid-2026 cost comparison put the difference in blunt terms: for 100 TB of storage, S3 Files runs dramatically cheaper per month than either EFS General Purpose or FSx Lustre ScratchFS 2, with EFS and FSx landing more than an order of magnitude higher on the monthly bill. For dataset-ingestion pipelines specifically, that collapses much of the cost penalty that used to push teams away from EFS, while preserving the NFS semantics those pipelines were built around. It doesn't yet replace EFS everywhere. Workloads that need multi-AZ write consistency and full POSIX semantics still need EFS itself.
FSx for Lustre Intelligent Tiering addresses a different pain point: operational burden. It delivers over 1 TB/s of throughput with sub-millisecond latency, and its automatic data management removes much of the manual tuning that used to make Lustre a tool only teams with dedicated HPC staff could run comfortably. Separately, S3 Express One Zone now scales to as many as 2 million GET transactions per second and hundreds of thousands of PUT transactions per second per directory bucket. That headroom changes the math on distributed training reads for any team willing to organize its dataset to take advantage of it.
What these releases don't touch is the structural logic underneath all of it. EBS still attaches to one instance at a time. EFS still carries a per-GB premium at large persistent scale. S3 Standard still carries the same latency profile for hot-path training reads that made it a poor direct pipe in the first place. The phase framework holds. The specific tier that wins inside each phase just got cheaper or faster, depending on which one you're looking at.
Where cloud-agnostic and self-hosted storage layers fit
Everything above assumes a team is operating entirely inside AWS. Plenty aren't. Teams running ML infrastructure across multiple clouds, or under requirements that data stay on-premises for sovereignty reasons, need something the AWS-native tiers alone can't fully provide. Picking the right AWS-native tier is still necessary work.
This is where S3-compatible object stores come in. They implement the Amazon S3 API directly, so application code, client libraries, and data pipelines built against S3 keep working without a rewrite, and the prefix-sharding and multi-threaded GET strategies developed for S3 apply directly.
The decision surface here is broader than a single vendor choice. Some platforms aggregate NVMe drives across a cluster into one global namespace, built for parallel I/O at tens of gigabytes per second. Others are positioned specifically for HPC and hybrid ML environments, where the requirement is low-latency access combined with compatibility across cloud and on-premises deployments at once. The evaluation criteria carry over directly from everything covered above: identify which phase is the actual bottleneck, then match the storage architecture to that phase, rather than to a single benchmark number on a spec sheet. That's the same discipline that applies inside AWS. The same discipline applies to a wider set of options: match storage architecture to the actual bottleneck, whether the choice is "which AWS tier" or "which storage architecture, anywhere."
Sources
- AWS S3 vs EBS vs EFS vs Glacier: Which Storage is Best in 2025? | by Ankush Madaan | SquareOps | Medium
- AWS Storage Service Guide 2025: S3 vs EBS vs EFS vs Glacier
- AI Training and Inference Storage Performance Requirements Benchmarked - SoftwareSeni
- Solved: I’ve benchmarked read latency of AWS S3, S3 Express, EBS and Instance store


