Our SSD development platform is built around the Xilinx Zynq® UltraScale+™ MPSoC as the core board. This device features a heterogeneous computing architecture, integrating programmable ARM Cortex® application processors with a high-performance FPGA fabric on a single chip. This combination enables flexible and efficient execution of both control-plane tasks (on the ARM cores) and data-plane acceleration (in the FPGA logic).

The core board is mounted onto a custom carrier (or base) board, which provides essential system-level support, including:
The SSD firmware runs on the ARM application processors, while flash translation and ECC offload are handled by dedicated co-processors.

The FPGA implements a NVMe-over-PCIe controller and ONFI-compliant NAND flash controllers. These controllers are connected with both on-board DDR memory and the ARM processor’s memory hierarchy, enabling efficient data movement across the heterogeneous compute fabric.
The NVMe-PCIe controller handles host communication via the PCIe interface, managing Submission and Completion Queues and performing DMA operations.
The ONFI flash controller manages low-level NAND operations, including read, program, and erase commands, timing control, and interface protocol compliance. Each flash controller corresponds to one flash memory channel, and the system currently supports up to eight channels.
The SSD board has three physical memroy regions:
Pins on the FPGA have been assigned for the flash chip interface, supporting up to 8 channels. Regarding interface protocols, NV-DDR2 is currently stably supported, while NV-DDR3 is supported experimentally. The board is shipped with the following flash chips:
The performance of pages at different locations within TLC NAND flash varies significantly. To provide a reference, internal bandwidth measurements for the SSD platform employing TLC chips (MT29F512G08EBHBF) are summarized in the table below: |Page latecncy|Lower-page|Upper-page|Extra-page|avg.| |—|—|—|—|—| |read|115us|130us|145us|130us| |program|600us|60us|1460us|706us|
| Internal bandwidth | Per-channel (single die) | Overall |
|---|---|---|
| read | 120MB/s | 480MB/s |
| program | 20MB/s | 80MB/s |
The FPGA operates at a frequency of 200 MHz and is configured to match the timing mode 7 of NV-DDR2/3. Each channel employs an 8-bit wide differential signaling interface, transferring data on both the rising and falling edges of the clock. Consequently, the theoretical bandwidth per channel is 400 MB/s.
All dies within a channel time-share the channel’s bandwidth. By increasing the number of LUNs (dies) per channel and leveraging interleaved die access (multi-LUN) along with multi-plane read/write operations, the effective bandwidth can approach this theoretical limit.
Under ideal conditions, with all eight channels operating simultaneously, the aggregate bandwidth ceiling reaches 3.2 GB/s.