Filesystem & VFS
The Virtual File System (VFS) is Linux’s abstraction layer that provides a unified interface to all filesystem types. Understanding VFS is crucial for debugging I/O issues and designing storage-aware systems.Prerequisites: System calls, process fundamentals
Interview Focus: File descriptors, VFS architecture, I/O paths, page cache
Time to Master: 4-5 hours
Interview Focus: File descriptors, VFS architecture, I/O paths, page cache
Time to Master: 4-5 hours
VFS Architecture
Core VFS Data Structures
The Four Pillars
- superblock
- inode
- dentry
- file
Represents a mounted filesystem:Key operations:
alloc_inode(): Allocate inodedestroy_inode(): Free inodewrite_inode(): Write inode to disksync_fs(): Sync filesystem
File Descriptors
File Descriptor Table
File Descriptor Limits
File Descriptor Inheritance
Path Lookup
The namei() Journey
Symlink Resolution
Page Cache
Page Cache Architecture
Page Cache Operations
Read-Ahead
Write Paths
Buffered vs Direct I/O
Write-Back and Dirty Pages
Filesystem Types
Virtual Filesystems
Disk Filesystems
Mount Namespaces and Bind Mounts
Interview Deep Dives
Q: Explain what happens when you run 'cat file.txt'
Q: Explain what happens when you run 'cat file.txt'
Complete flow:
-
Process creation: Shell forks, child execs
/bin/cat -
open() syscall:
- Path lookup via VFS (dcache, namei)
- Permission check (DAC + MAC)
- Allocate struct file
- Allocate file descriptor
- Return fd
-
read() syscall:
- fd → struct file → inode → address_space
- Check page cache for requested pages
- Cache miss: Submit I/O to block layer
- I/O completion: Copy to page cache
- Copy from page cache to user buffer
- Return bytes read
-
write() to stdout:
- fd 1 → terminal device
- TTY layer processes output
- Display on screen
-
close() + exit:
- Decrement file refcount
- Free fd in table
- Process exits
Q: What's the difference between fsync and fdatasync?
Q: What's the difference between fsync and fdatasync?
fsync():
- Flushes file data AND metadata to disk
- Metadata: size, mtime, block pointers
- Two writes: data blocks + inode
- Required for crash consistency
- Flushes file data to disk
- Only flushes metadata if needed for data access
- Skips non-essential metadata (atime, mtime)
- Faster for append-only patterns
Q: How does the kernel prevent file descriptor leaks?
Q: How does the kernel prevent file descriptor leaks?
Kernel mechanisms:
-
RLIMIT_NOFILE: Per-process limit on open fds
-
O_CLOEXEC: Close fd on exec
- Process exit: All fds automatically closed
- Use RAII (C++/Rust): Fd closed when object destroyed
- Close fds in error paths
- Use valgrind/lsof to detect leaks
Q: Why is 'ls' slow on directories with many files?
Q: Why is 'ls' slow on directories with many files?
Causes:Architectural solutions:
-
Directory reading: getdents() syscall reads directory entries
- Linear scan of directory file
- ext4 uses htree for lookup, but listing is still O(n)
-
stat() per file: ls -l stats every file
- Each stat is a separate syscall
- May require inode read from disk
-
Sorting: ls sorts output
- O(n log n) in memory
- Don’t put millions of files in one directory
- Use directory sharding: files/ab/cd/file.txt
- Use filesystem with better large directory support (XFS)
Performance Monitoring
Interview Deep-Dive
The page cache is consuming 90% of memory on a production server. Is this a problem? How does the kernel decide what to cache and what to evict?
The page cache is consuming 90% of memory on a production server. Is this a problem? How does the kernel decide what to cache and what to evict?
Strong Answer:
- A page cache using 90% of memory is normal and desirable. Linux uses all available free memory as page cache because unused memory is wasted memory. The page cache accelerates file reads from milliseconds (disk) to microseconds (memory). The key metric is not cache size but cache hit ratio. If the working set fits in cache and the hit ratio is 95%+, the system is performing optimally.
- It becomes a problem only when memory pressure causes the cache to evict pages that the application will need again soon, leading to major page faults (disk reads). The symptom is high
pgmajfaultcounts in/proc/vmstator highkswapdCPU usage. - The kernel uses a two-list LRU (Least Recently Used) algorithm to decide what to cache. Each page is on either the active or inactive list, and the lists are split by type (anonymous vs file-backed). When a file page is first read, it goes on the inactive list. If accessed again while on the inactive list, it is promoted to the active list. When memory reclaim is needed, kswapd scans the inactive list from the tail, evicting pages that have not been recently accessed. Active list pages are periodically demoted to the inactive list if not recently accessed.
- The
vm.vfs_cache_pressuresysctl tunes how aggressively the kernel reclaims dentry and inode caches relative to page cache. Thevm.swappinesstunes the preference for reclaiming anonymous pages (swap out) versus file pages (drop from cache). For a database server, I would setswappiness=10to preserve anonymous memory (heap, buffer pool) and prefer dropping file cache.
posix_fadvise() allow applications to influence page cache behavior, and when would you use it?Follow-up Answer:posix_fadvise()lets applications give hints to the kernel about their access patterns.POSIX_FADV_SEQUENTIALtells the kernel to increase read-ahead (prefetch more pages ahead of the current read position).POSIX_FADV_RANDOMdisables read-ahead, which is better for random-access patterns like database index lookups.POSIX_FADV_WILLNEEDasks the kernel to prefetch the specified range into cache (asynchronous read-ahead).POSIX_FADV_DONTNEEDtells the kernel that the specified range is no longer needed and can be evicted from cache. I would useDONTNEEDafter processing a large file in a streaming fashion (like log processing) to avoid polluting the cache with data that will never be re-read. I would useWILLNEEDbefore a database knows it will need specific pages for an upcoming query.
Compare ext4 and XFS for a write-heavy workload with many small files. What are the internal architectural differences that affect performance?
Compare ext4 and XFS for a write-heavy workload with many small files. What are the internal architectural differences that affect performance?
Strong Answer:
- For write-heavy workloads with many small files, XFS generally outperforms ext4 due to fundamental architectural differences in allocation and journaling.
- ext4 uses a bitmap-based block allocator: it searches free space bitmaps for available blocks. For many small files, this creates contention on the bitmap locks and can lead to fragmentation as the allocator struggles to find contiguous free space among many small allocations. ext4’s journaling (jbd2) writes metadata changes to a journal before committing them to their final locations, and the journal is a single sequential log shared by all operations.
- XFS uses B+ tree-based allocation groups. The filesystem is divided into independent allocation groups (AGs), each with its own B+ trees for free space, inodes, and extents. This parallelism means multiple threads can allocate blocks in different AGs simultaneously without contending on a single lock. XFS also uses delayed allocation aggressively: it defers block allocation until writeback time, which allows the allocator to make better decisions about contiguity.
- For journaling, XFS uses delayed logging: metadata changes are accumulated in memory and flushed in batches, reducing journal I/O. ext4’s jbd2 commits more frequently by default (every 5 seconds).
- The practical impact: on a server creating millions of small files per day, XFS’s allocation group parallelism and B+ tree-based free space tracking significantly reduce lock contention and provide more predictable latency. ext4 is better for simpler workloads and has wider tooling support (fsck is more mature).
- VFS defines a set of operation tables (
super_operations,inode_operations,file_operations,address_space_operations) that each filesystem implements. When user space callswrite(), VFS dispatches throughfile->f_op->write_iter()which is a function pointer set by the filesystem duringopen(). This is a single indirect function call overhead — roughly 1-2 nanoseconds on modern CPUs. Given that the actual I/O work takes microseconds to milliseconds, the VFS abstraction cost is negligible (less than 0.01% of total I/O time). The real benefit is enormous: applications, system calls, and kernel subsystems (page cache, memory mapping) work identically regardless of the underlying filesystem. This is why you can switch from ext4 to XFS by simply reformatting and remounting without changing any application code.
Next: I/O Subsystem →