Skip to content
High-Performance File Systems & Low-Level I/O: Inodes, File Descriptors, POSIX I/O, Page Cache, Journaling, `epoll` vs `kqueue`, and `io_uring` Asynchronous Ring Buffers

High-Performance File Systems & Low-Level I/O: Inodes, File Descriptors, POSIX I/O, Page Cache, Journaling, `epoll` vs `kqueue`, and `io_uring` Asynchronous Ring Buffers

What it is

A file system is a durable, named data model backed by kernels, block devices, and storage caches. POSIX provides the common file-descriptor interface, readiness APIs such as Linux epoll and BSD/macOS kqueue report when a descriptor can be attempted, and Linux io_uring submits operations through shared rings for completion-oriented I/O.

How it works

A path such as /srv/orders/day-1.json is resolved component by component. A dentry caches a parent-directory and name relationship, while an inode stores metadata and identifies the underlying object, such as a regular file. Different paths can reach the same inode through hard links. Opening a path creates a file descriptor in the process descriptor table; a descriptor references an open file description that stores state such as the current offset. fork copies descriptor-table entries that refer to the same open descriptions, whereas dup adds another descriptor to one existing description.

POSIX calls separate the identity of a descriptor from an operation. open obtains a descriptor, read and write copy bytes through the current offset, pread and pwrite supply an explicit offset, and lseek changes the current offset. POSIX describes the interface; it does not force every filesystem to store files the same way. Ext4, XFS, Btrfs, ZFS, network filesystems, and FUSE can all present similar calls with different consistency, durability, and performance behavior.

Linux’s page cache holds file-backed data pages that were read or written. A cache hit avoids a storage read, while a write usually updates the page cache before its data reaches durable media. fsync asks the kernel to complete a file’s required durability work, but the guarantee depends on the filesystem mode, mounted device, storage cache, and hardware failure model. Calling fsync does not make every file or directory update part of the same atomic transaction. Direct I/O can avoid or constrain the page cache for eligible aligned operations, at the cost of alignment, size, and application-buffering complexity.

A metadata journal records operations before or during a transaction so a crash during a multi-step update can be rolled forward or rejected. Write ordering and barriers determine how data-file updates, metadata, and device caches are coordinated. Journaling protects filesystem structure; backups and replication still determine recovery from media loss, corruption, or deletion.

When an application waits on many sockets, polling every descriptor wastes work on entries that are not ready. Linux epoll and BSD/macOS kqueue maintain registered interest sets and return ready descriptors. Linux commonly exposes one epoll instance for the process, while BSD and macOS can attach a kqueue instance to one or many descriptors; both are readiness mechanisms, and neither is a completion API. epoll supports level-triggered and edge-triggered operation; an edge-triggered consumer must drain or otherwise handle state until the next event. Readiness is an observation, not a reservation: another consumer can take the data, or an error, hangup, or priority condition can change the outcome before the application calls read or write. A nonblocking descriptor should therefore still be attempted, and EAGAIN or EWOULDBLOCK must return the descriptor to the wait set. End-of-file and errors are results, not proof that a read succeeded.

io_uring uses a submission queue for work descriptors called SQEs and a completion queue for completion entries called CQEs. User space and the kernel share mapped ring memory. A process can submit many filesystem, network, and timer operations in a batch, then consume completions. Some operations complete inline and others run asynchronously. Batching amortizes entry and exit costs, but queue depth, memory registration, filesystem support, and completion policy still determine throughput and tail latency.

A high-performance filesystem changes how layout, metadata, durability, and network distribution affect those calls. FFS-style allocation groups related metadata and data blocks to improve locality. Lustre and GPFS distribute files and strips across servers, relying on parallel metadata, striping, and client caching; an accelerator or parallel file system scales aggregate throughput but introduces server and fabric dependencies. NVMe reduces device queueing and latency, but the connected filesystem must still issue efficient, sufficiently large I/O. Direct I/O and io_uring are I/O paths, not filesystems: they can bypass or batch page-cache work while operating on a file hosted by a high-performance filesystem.

    flowchart TD
    App[Application] --> Buffered[read write pread or pwrite]
    App --> Direct[Aligned O_DIRECT I/O]
    App --> Submit[io_uring submission queue]
    Buffered --> VFS[Virtual filesystem layer]
    Direct --> VFS
    VFS --> Cache{Page-cache path}
    Cache -->|miss or direct| FS[Ext4 XFS FFS Lustre or GPFS]
    Submit --> KernelWork[Kernel I/O worker]
    KernelWork --> FS
    KernelWork --> Completion[io_uring completion queue]
    FS --> NVMe[Block layer and NVMe queue]
    NVMe --> Device[Storage device]
  

A blocking diagnostic can show the actual call path and timestamps:

findmnt -T /var/lib/orders
stat /var/lib/orders/day-1.json
strace -T -e trace=openat,read,write,pread64,pwrite64,fsync /usr/bin/find /var/lib/orders

Complexity

These bounds describe visible algorithmic work, not disk, network, or filesystem latency. Directory layout and the selected filesystem can make path lookup and write costs more expensive than the table alone suggests.

OperationWork boundImportant cost outside the bound
Descriptor-table lookupO(1)Cache and kernel locking
Path component lookupO(1) per component with an effective directory indexDirectory depth, collisions, and filesystem structure
Cached read or write of k bytesO(k)Memory bandwidth and copy behavior
Uncached read of k bytesO(k) plus storage I/OReadahead and device queueing
Scan r descriptors with pollO(r)Kernel transitions and scheduler behavior
Wait for an active interest setProportional to registrations and returned ready eventsepoll and kqueue internal data-structure details are implementation-dependent
Submit n SQEs and drain their completionsO(n) application workKernel execution, I/O, and completion batching
Copy between bounded k-byte buffersO(k) workNumber of copies dominates memory traffic

When to use

  • The application needs a durable namespace, shared file access, or random access to file data.
  • POSIX calls are sufficient and blocking behavior fits the service model.
  • A connection server needs a scalable way to wait for many descriptors.
  • You need to measure page-cache, syscall, journaling, and storage behavior before optimizing.
  • The Linux workload can benefit from batched completion-oriented operations.

Alternatives

  • Memory-mapped files — can reduce explicit copying and enable page-based access, but page faults, mapping lifetime, and cross-platform behavior complicate correctness.
  • Blocking threads or synchronous calls — are easy to reason about, but each active operation can consume a thread and a context switch.
  • select or poll — are portable and simple, but scanning the full descriptor set can waste CPU as concurrency grows.
  • Kernel-bypass networking — can reduce copies and queues for supported devices, but sacrifices portability, kernel scheduling, and some compatibility.
  • Object storage APIs — win for durable blobs accessed over a network, but they do not provide a local filesystem namespace or POSIX file semantics.

Related