one serves as the host PC for downloading firmware to and monitoring the output from the target board, and the other as the test machine.
The host identifies the SSD by performing read and write operations on its doorbell registers. It then sets up the Admin Submission Queue (ASQ) and Admin Completion Queue (ACQ) pair, notifying the SSD of the queue configuration via the corresponding doorbell registers. Subsequently, additional I/O Submission and Completion Queue (SQ/CQ) pairs are established by issuing queue creation commands through the admin queue pair. Once all required SQ/CQ pairs are properly initialized, the SSD is ready to process NVMe I/O requests.
The host places NVMe commands into the Submission Queue (SQ) and then updates the SQ tail pointer by writing to the corresponding doorbell register. The SSD subsequently fetches the commands and any associated data via PCIe DMA. Once the request has been processed, the SSD, if required, writes the resulting data to the host-provided buffer and, posts a completion entry to the Completion Queue (CQ).
The data cache employs a No Read Allocate policy for read misses and a Write Allocate policy combined with Write-Back for writes. The cache block size is 16 KB, aligned with the flash page size, whereas NVMe I/O requests operate on units of 4 KB (the logical sector size), with each request typically spanning one or more such sectors.
Consequently, when the data cache flushes a dirty cache block, it must consult the Logical-to-Physical (L2P) mapping table to identify which 4 KB sectors within the 16 KB page are valid. If only a subset of sectors is being updated, a Read-Modify-Write operation is performed to preserve the integrity of the unchanged sectors.
The L2P (Logical-to-Physical) mapping table is stored in flash memory and loaded into DRAM on demand. Each 16 KB translation page contains 4K physical page addresses (PPAs). The physical addresses of these translation pages themselves are stored in the Global Translation Directory (GTD).
Initially, the in-memory mapping table cache is empty. Translation pages are fetched into memory only when needed and are updated out-of-place: when modified, dirty pages are written back to new flash locations using a write-back policy, and the updated physical addresses are recorded in the GTD.
Physical page allocation and management are controlled by the block manager. Each flash plane maintains one open block, from which physical pages are allocated for logical pages mapped to that plane. The corresponding invalid page bitmap is updated atomically during page allocation. When the open block is exhausted, a new block is opened on the same plane.
The current garbage collection (GC) policy is intra-plane, online, and blocking. Specifically, when the block manager attempts to open a new block and finds insufficient free blocks available on the plane, it triggers an intra-plane GC operation. During this GC process, all other read and write requests directed to that plane are blocked until GC completes.
The Flash Interface Layer is responsible for handling data read, write, and erase operations on the flash chips. It first enqueues flash transactions into a per-chip pending queue and processes requests from each chip in a round-robin polling manner. Flash commands are issued through the ONFI controller.
The ONFI standard defines electrical interfaces, signaling protocols, and command sets to ensure interoperability among NAND flash chips from different manufacturers. Our ONFI flash controllers are implemented in FPGA, and currently support Asynchronous, NVDDR2, and NVDDR3 interfaces. The Asynchronous and NVDDR2 interfaces operate at 1.8 V, while the NVDDR3 interface operates at 1.2 V. When powered up at 1.8 V, the flash chip defaults to Asynchronous mode. The controller then issues a SET FEATURES command to switch the flash chip to NVDDR2 mode. When powered up at 1.2 V, the NVDDR3 interface is activated directly upon initialization.
Both ECC encoding and decoding are now fully integrated into the flash controller hardware. Once the controller is properly configured, it automatically performs ECC operations on all data during read and write transactions—without requiring software intervention. The ECC parity bits are stored alongside the user data in the out-of-band (OOB) area of each flash page.
However, when ECC decoding detects an uncorrectable error (or in some implementations, even a correctable multi-bit error), explicit error handling is required, and the correction procedure is handled by the code in ECC engine.
After data is written to flash memory, certain metadata structures in the SSD’s internal memory are generated or updated. These metadata must be persisted to non-volatile storage to ensure consistency and durability. When the host issues an fsync request, the SSD performs the following operations in sequence:
Upon system shutdown, the host sends a Shutdown Notification (as defined in the NVMe specification) to the SSD. In response, the SSD executes the same fsync-like persistence sequence to safely commit all in-memory metadata to flash before power is removed.
Important: The SSD board does not include power-loss protection capacitors. An unexpected power loss will result in data corruption or metadata inconsistency, as in-flight writes and unflushed metadata cannot be recovered.
NVMe requests residing in the Submission Queue (SQ) are fetched by the SSD into its internal memory via PCIe DMA. Each request is then split into 16 KB-aligned sub-requests, which are processed sequentially. Upon completion of all sub-requests, a completion entry is posted to the Completion Queue (CQ).
For each sub-request, the SSD first checks the data cache:
During both data write-back and data loading, the SSD consults the L2P mapping cache to translate logical addresses to physical flash locations:
Notably, data write-back requires allocation of a new flash page (due to the out-of-place update nature of NAND flash).
All flash read and write operations—for both data pages and translation pages—are performed by submitting requests to a dedicated request queue of an ARM core that runs the Flash Interface Layer. Completed flash operations are returned via a completion queue associated with that core.
Our SSD platform employs a multi-tasking concurrency model to process multiple NVMe requests simultaneously. Each task handles a single NVMe request, and tasks share critical resources using locks, semaphores, and other synchronization primitives.
Underlying this model is a lightweight coroutine framework. The main program initializes 16 task coroutines at startup. During the execution gaps between coroutine yields, the scheduler polls various queues—such as the NVMe Submission Queues (SQs) and flash request completion queues—and resumes the appropriate coroutines when their awaited events or data become available.