Est.

How a POSIX Filesystem Layer Is Built on Top of S3

Three S3 upgrades finally made it possible to layer filesystem semantics on top of object storage.

Senior Contributing Editor · · 11 min read
Cover illustration for “How a POSIX Filesystem Layer Is Built on Top of S3”
POSIX Filesystems · October 10, 2026 · 11 min read · 2,392 words

Building a real POSIX filesystem on top of S3 means solving a stack of problems that have nothing to do with each other on the surface: semantic mismatches, metadata coordination, caching, and write consistency. This piece walks through each layer, in order, so the engineering choices make sense.

S3 and POSIX describe different contracts for accessing data

S3 and a POSIX filesystem promise different things to the applications that use them, and that mismatch is the whole problem this piece exists to untangle.

S3 treats data as immutable blobs, each addressed by a flat key. You PUT a whole object. You GET a whole object. You DELETE it. There's no in-place modification, no native rename, no directory tree, and no locks. An object is a thing you replace entirely or not at all, the same way you can't edit one paragraph of a fax. You resend the whole page.

POSIX makes very different promises. Files live inside a hierarchy of directories. Processes can read and write at arbitrary byte offsets. Rename operations are atomic, meaning they either fully happen or don't happen at all, with nothing in between. Files can be locked so two processes don't collide. And every process that opens a file sees a consistent view of it, updated in real time.

That mismatch isn't academic. Apache Iceberg, Delta Lake, and Apache Hudi, the three table formats that run most modern data lakes, all exist because object storage has no atomic commit primitive. Each format builds its own bookkeeping layer (manifests, transaction logs, metadata pointers) just to fake the one guarantee POSIX gives away for free: that a write either fully succeeds or fully fails, with no in-between state visible to a reader.

The same crack runs through AI training infrastructure. Checkpoint frameworks often seek back into a file to update a header, append data incrementally as training progresses, or write a new file and rename it over the old one for safety. Every one of those patterns assumes a filesystem contract. None of them work against a raw S3 bucket, because S3 doesn't do in-place writes, incremental appends, or atomic renames. A checkpoint library built around those assumptions won't degrade gracefully on S3. It will corrupt data.

Why FUSE shims were the first response and why they failed structurally

The first fix engineers reached for was FUSE: software that intercepts filesystem calls and translates them into something else behind the scenes. For S3, that meant translating POSIX calls into S3 API calls. The idea sounds reasonable. The execution couldn't work, because FUSE can translate syntax, but it can't conjure semantics S3 never had.

S3FS-FUSE was one of the earliest serious attempts. It translated POSIX operations into S3 calls, and it worked fine if you were just one person browsing files casually. But under real production load, it broke in three structural ways baked into the approach itself.

The first failure was write amplification. S3 only supports whole-object PUTs, so a small 4KB random write anywhere inside a file forced a full re-upload of the entire object. A tiny edit to a large file could trigger an overwrite north of 100MB, turning a trivial write into a massive one.

The second was a directory-listing penalty. S3 has no native concept of a directory, only prefixes that look like paths. To fake an ls command, S3FS had to enumerate every object under a prefix, and at Uber, engineers observed a substantial fraction of data-loading time consumed by exactly these LIST operations.

The third was a consistency gap that made the whole approach dangerous. Before December 2020, S3 offered only eventual consistency for overwrite PUTs and DELETEs. A process could write new data and then read back stale data moments later. For checkpoint recovery in AI training, that's silent data corruption, and you won't see it until a model fails to resume training correctly weeks later.

Goofys tried a different angle: skip metadata emulation entirely, treat S3 purely as a blob store, and drop the local disk cache. It was lighter, but it introduced race conditions on concurrent writes and kept the same write-amplification problem, because the object model underneath hadn't changed. Trimming the shim doesn't fix the foundation.

Uber's engineering team reached the conclusion that matters most here: POSIX-on-S3 cannot be patched at the client layer. FUSE tools weren't failing because of bad engineering. They were failing because no amount of clever translation can manufacture metadata semantics or consistency guarantees that the storage service itself refuses to provide. That lesson pushed later production systems, Archil among them, to move those guarantees into a managed service layer instead of trying to bolt them on at the client, so existing applications can run unmodified without a separate SDK or a rewrite.

The S3-side changes that made a real POSIX layer possible

Production-grade POSIX-on-S3 became possible once S3 itself changed, not once clients got smarter. A handful of shifts in that period removed the structural failure modes that FUSE tools were never going to work around.

The first and most important: in December 2020, S3 introduced strong read-after-write consistency, at no extra cost. A write became immediately visible to every reader, everywhere. That single change is the prerequisite for any real filesystem semantics, and it directly fixed the checkpoint-recovery corruption risk that made early FUSE tools dangerous for training workloads.

The second shift addressed latency: S3 Express One Zone gave applications a low-latency storage class suited to the kind of fast, repeated access patterns that training and inference workloads generate, removing a performance ceiling that pure standard S3 couldn't clear.

The third shift, in 2023, addressed the root cause: S3's lack of real metadata support. S3's original flat namespace had no real concept of directories, only prefixes styled to look like paths. The new directory-bucket type added native directory support, which made commands like ls and mkdir fast without enumerating every object under a prefix. That's the exact tax that burned so much wall-clock time at Uber, gone at the source.

These three shifts are cumulative. Consistency removed the correctness blocker. Express One Zone removed the latency blocker. Directory buckets removed the metadata-cost blocker. Each one was necessary. None of them alone was sufficient. A storage layer with consistency but no fast metadata still chokes on ls. A layer with fast metadata but no consistency still risks corrupting a checkpoint. So all three had to move together before a real POSIX layer on top of S3 became worth building.

Diagram: Three S3 Changes That Made Real POSIX Possible. Visualizes: Show three cumulative S3 platform shifts that each removed a distinct blocker to production-grade POSIX semantics: (1) December 2020 — strong read-after-write consistency…

The two dominant architectural patterns for building the POSIX layer

Once S3 supplied the missing primitives, two architectural patterns emerged to build the POSIX layer on top of them. Neither is a toy. Both run in production at serious scale, and the choice between them comes down to how much operational control a team wants to keep for itself.

The first pattern separates the metadata engine from the object data store. A dedicated metadata engine (Redis, TiKV, FoundationDB, Postgres, or MySQL) holds the directory tree, file attributes, and the mapping from POSIX paths down to chunk IDs. The object backend holds the actual file content, split into fixed-size chunks and addressed by content hash. A FUSE client sits on top, translating POSIX syscalls into metadata reads from the engine and object reads from S3, backed by a local block cache on disk.

This pattern delivers what FUSE-on-S3 alone never could: atomic renames and metadata operations, guaranteed by the metadata engine's own transaction model, not faked at the client. The tradeoff is operational. The metadata engine becomes the most important dependency in the whole system, more important than S3 itself. If the metadata store goes down, the filesystem goes down, even if S3 is perfectly healthy. TiKV deployments at ByteDance and Xiaohongshu run billions of files this way, which proves the pattern scales. It also takes a team that's prepared to run and babysit a distributed metadata store, the kind of job that comes with its own on-call rotation.

The second pattern is a managed tiered architecture: a fast caching tier (commonly built on something like EFS) sits in front, with S3 behind it as the durable backing store. The caching tier handles POSIX operations immediately (rename, chmod, symlink, atomic writes) with full semantics, then asynchronously writes changes back to S3 in batches.

Writes land as durable on the caching tier right away, and every client mounting that same filesystem can see them. They propagate to S3 after a short write-inactivity window following the last write (commonly around 60 seconds). That window matters because it enables batching, so small writes get coalesced into larger objects, and that cuts S3 PUT costs. For AI training specifically, this fits naturally: checkpoint writes come in bursts followed by stretches of pure computation, so there's no rush to flush instantly.

The architectural pattern has traded the self-managed metadata engine for a managed service: no separate metadata store to operate, because the managed layer owns that job. The same three-layer shape (fast edge tier, metadata coordination, durable object backend) appears independently outside the US market too. Aliyun's CPFS+OSS combination uses a distributed, symmetric metadata server architecture that handles POSIX operations at sub-millisecond latency, with a tiered dataflow pushing cold data down to OSS at object-storage cost. Different company, different cloud, same conclusion.

That convergence is the real signal. Whether the metadata layer is run by the team or handed to a managed service, the winning architecture always keeps the POSIX-capable edge tier separate from the object-backed durable tier. No production system that actually works collapses those two into a single FUSE translation layer. The self-managed pattern is what teams reach for when they need multi-cloud portability or want to avoid being tied to one vendor's managed stack, since a self-run metadata engine can sit in front of any combination of object storage providers with the same logic underneath. The managed pattern trades that flexibility for far less operational overhead. Neither one is wrong. They're built for different priorities.

For teams that want a unified filesystem view spanning multiple clouds, without physically moving data off any of those object stores into a new silo, that separation of layers makes it possible: the metadata and caching layer can live independently of any single cloud's object store, mounting each one as a native filesystem without a copy step in between.

How Amazon S3 Files implements the managed tier pattern

Amazon S3 Files, announced April 7, 2026, is the clearest public example of the managed tiered pattern running in production, and its pricing model shows exactly where the underlying physics cost gets passed to the customer.

The mechanics start with networking. AWS provisions a mount target inside the customer's own VPC, and compute resources (EC2, ECS, EKS, or Lambda, with Amazon Bedrock AgentCore Runtime added in May 2026 through a separate bring-your-own file system integration) connect to it over NFS. The mount target gets its own ENI inside the customer's subnet, so a large file transfer through S3 Files won't starve other services that share the same instance. That's a small detail with a real payoff: one workload moving a lot of data won't choke out another workload's network path just because they happen to live on the same box.

Data materializes differently depending on file size. Files 128KB or smaller sync automatically into the fast file layer the moment their directory is accessed, which suits code repositories, config files, and build artifacts, the kind of small files where latency on every single access actually matters. Files larger than 128KB load lazily, only on first read.

Writes follow the same pattern as the general managed-tier architecture: they land first on the EFS-backed layer, immediately consistent to every client mounting that filesystem, then sync asynchronously back to S3. If two writes to the same object collide during that sync, the conflict gets handled through versioning or routed into a recovery namespace, so nothing silently overwrites something else. Files get evicted from the hot tier automatically too, based on a configurable retention window measured in days of inactivity.

The pricing model shows where the architecture becomes a line item. Hot-tier storage and small-file access are billed at EFS Performance-optimized Standard rates ($0.30 per GB of storage), because the hot tier is, physically, EFS. But large reads above a certain size skip the EFS layer entirely and go through a direct S3 GET at no S3 Files charge, so the cost structure forks depending on file size. That mirrors how Archil's own pricing separates cold storage, billed at the customer's own bucket rates, from a performance tier charged separately per GB, the same basic split between "fast tier" and "cheap tier" appearing in more than one managed product.

Two charges deserve attention because they compound quietly. Every data access operation carries a 32KB minimum, and metadata operations, like listing a directory or checking a file's attributes, are metered as a 4KB read. On a workload with millions of small metadata-heavy operations, those minimums add up fast, the digital equivalent of a vending machine that rounds every purchase up to a dollar. Renaming a directory compounds this further: S3 Files meters a rename as a separate copy-and-delete operation for every single object under that prefix, so moving a folder containing ten thousand files is ten thousand individual billed operations, not one.

Large files get a better deal under what AWS calls Read Bypass Mode. A Parquet file or similar large object read directly through S3 GET incurs no S3 Files data access charge at all, though a 4 KiB S3 Files metadata read still applies per operation. The surcharge for using the filesystem layer only applies to what's actually materialized inside the hot tier, not to everything that passes through the mount.

Finally, the namespace itself can be split into a large number of independent access points, each mapped to a specific directory path with its own IAM and POSIX permissions. It's built for multi-tenant and agent-based environments, where dozens or hundreds of isolated workloads need their own narrow view into a shared filesystem, so they don't step on each other's files. It's a pattern Archil applies as well, letting an S3 bucket or an existing system of record mount as a native filesystem for multiple agents at once, without an ETL pipeline or a second copy of the data sitting around to pay for and keep in sync.

Sources

  1. Orchestrating multi-agent AI architectures with Amazon S3 Files

More in POSIX Filesystems