On September 7, 2026, nand2mario described zSST, a SystemVerilog implementation of the 3dfx Voodoo Graphics, also known as SST-1. The project grew out of a month of work on z486_MiSTer and an interest in the 3D cards that followed the first half of the 1990s. Need for Speed II SE provided the original inspiration, with its textured graphics, fog, and fluid performance.

Glide triangle API
void grDrawTriangle(const GrVertex *a, const GrVertex *b, const GrVertex *c);

zSST is integrated with the z486 CPU and PC peripherals in z486 XL. The result is a DOS PC with Voodoo graphics implemented in the programmable logic of a Xilinx KV260 board. Tomb Raider runs with its original 3dfx renderer. zSST supports prepared triangles, texture filtering, mipmapping, depth and alpha tests, fog, blending, dithering, framebuffer access, and buffer swaps. It accepts both fixed-point and floating-point setup interfaces. Hardware testing currently focuses on Tomb Raider, while broader compatibility and later Voodoo generations remain future work.

The CPU and renderer run at 100 MHz on the KV260. Its logic, DSP blocks, on-chip memory, and DDR bandwidth accommodate the combined design. The DE10-Nano does not have enough space for the graphics addition. The KV260 uses its onboard DDR and needs no external SDRAM module.

The implementation follows the 1999 Glide source released by 3dfx and the SST-1 specification. 3dfx released Glide before NVIDIA acquired its core graphics assets in December 2000. The surviving Glide source, the SST-1 specification, 86Box, earlier MAME Voodoo work, and SpinalVoodoo traces supplied behavioral references and test material. The specification defines register behavior and rendering results, while leaving the original circuit structure open.

The accelerator has five main command registers. triangleCMD starts a prepared triangle, ftriangleCMD starts one through the floating-point setup interface, nopCMD flushes the pipeline and can reset statistics counters, fastfillCMD clears a clipped color or depth rectangle, and swapbufferCMD changes the displayed buffer immediately or at vertical retrace. Other registers carry coordinates, gradients, and render state. Memory-mapped regions handle texture uploads and direct framebuffer access.

Glide's grDrawTriangle call receives three vertices after the host CPU has transformed the geometry, calculated lighting, clipped the result, and projected it onto the screen. SST-1 has no hardware transform-and-lighting engine. The rasterizer covers the triangle, interpolates color, depth, and texture coordinates, and sends samples through the texture and framebuffer units. The texture unit fetches and filters texels. The framebuffer path performs visibility tests, fog, blending, and writes.

The original Voodoo divides this work between the Frame Buffer Interface, or FBI, and the Texture Mapping Unit, or TMU. At 50 MHz, its advertised peak is one textured, depth-tested output pixel per clock, or 50 million pixels per second. The pipeline overlaps several pixels at different stages. A pixel still takes multiple stages to finish. The filled pipeline can accept one pixel per clock when memory keeps up, which helped Voodoo deliver textured 3D games at 30 FPS or more.

Software can submit floating-point values, while most of the rendering path uses fixed-point arithmetic. The internal formats are 12.4 for screen X and Y, 12.12 for red, green, blue, and alpha, 20.12 for depth Z, 14.18 for texture S/W and T/W, and 2.30 for reciprocal W. The fvertex, fstart, and floating-point gradient registers accept IEEE single-precision values. SST-1 converts them to these fixed-point formats, after which interpolation mainly uses additions of precomputed increments.

For perspective-correct texturing, the TMU interpolates S/W, T/W, and 1/W, then divides the first two values by the third. It chooses a mip level and applies bilinear filtering to four neighboring texels. zSST uses a four-stage front end for perspective and level-of-detail calculations. Address generation, cache lookup, format decoding, and texture combining support palette-based and NCC-encoded textures.

The FBI pixel path has six registered stages. F0 selects sources, checks chroma key, and prepares Z/W values. F1 applies color and alpha combine functions. F2a performs alpha and depth tests and looks up the fog factor. F2b applies fog. F3 reconstructs the destination color and performs alpha blending. F4 converts to framebuffer precision, dithers, and applies write masks. The stage boundaries meet the FPGA clock target while preserving one-pixel-per-clock throughput.

Memory access became the main implementation challenge. The original Voodoo gives the FBI and TMU separate 64-bit memory paths. Four-way texture interleaving assigns banks according to the even or odd parity of the column and row, so every 2x2 texel window reaches all four banks. This supplies the four bilinear samples in parallel. The FBI uses a related interleaved layout for color and depth or alpha data. Its peak is one rendered pixel per clock and two pixels per clock for clears. At 50 MHz, each 64-bit path has theoretical bandwidth of 400 MB/s, for 800 MB/s in total, with the two paths dedicated to their respective jobs.

The KV260 has shared DDR behind AXI ports. Linux, the FPGA PC, and display scanout all compete for that memory. A 128-bit port at 100 MHz offers a theoretical 1.6 GB/s. Measurements reached 1,370 MiB/s for a 4 KiB read with one outstanding request and 189 MiB/s for a 64-byte read. The first data typically arrived after about 280 ns, or 28 clocks at 100 MHz, with occasional longer delays.

zSST uses an 8 KiB texture cache with 64-byte lines. It can keep eight cache-line fetches outstanding, parks misses in a replay queue, and uses a 64-entry reorder buffer to retire completed samples in their original order. Separate 4 KiB color and depth or alpha caches serve the framebuffer path. Write combiners pack neighboring 16-bit updates into 128-bit requests. Forwarding preserves dependencies when a later depth or blend operation needs a value that an earlier pixel has changed. The FBI and TMU share the renderer's HP2 AXI port. The PC uses HP0 and display scanout uses HP3.

Simulation tests model DDR timing, the TMU, FBI, shared arbitration, and the front end. Reads arrive after at least 26 clocks, writes are rate-limited, and the full-renderer tests allow 32 outstanding reads. At 100 MHz, zSST reaches 78.5 MPix/s for textured triangles and 72.8 MPix/s with depth testing and blending. Published 50 MHz Voodoo 1 estimates are 43 and 37 MPix/s for comparable feature groups. The workloads differ, so these figures show an approximate range rather than a measured speedup.

On the board, Tomb Raider Level 2 produced 237 displayed buffer swaps in about 20 seconds, or roughly 12 swaps per second. The measurement counts buffer swaps and does not use an engine FPS counter. Early results point to the CPU as the main limit, with shared DDR contention also possible. The 100 MHz z486 XL reaches 38.5 FPS in maximum-detail Doom and 8.1 FPS in Quake 1.06. Those results are about 20% above the 85 MHz DE10-Nano build, with gains of about 23% for Doom and 19% for Quake. A 512 KiB write-back L2 cache in UltraRAM helps the CPU use DDR.

In the integrated XCK26 build, zSST uses about 29,500 LUTs, 28,100 flip-flops, 14 RAMB36 blocks, 8 RAMB18 blocks, and 97 DSP slices. The combined PC and graphics design meets timing at 100 MHz. zSST and z486 XL are open source. The z486 XL release tagged z486_XL_20260906 includes the Linux support and application required to boot DOS disk images on a KV260. The project credits SpinalVoodoo for Glide traces and reference screenshots, 86Box for implementation references, and Fabien Sanglard's Voodoo 1 memory analysis as additional background.