SOFT IP / AI ACCELERATORS / CNN

CNN — Accelerator

Convolution accelerator pipeline, WIOWIZ IP and WIOWIZ VIP

256 INT8 MACs in a 16 by 16 convolution engine

A convolution accelerator built around a 16 by 16 systolic PE array of 256 INT8 multiply-accumulate units with INT32 accumulation. On-chip feature and weight buffers hold tensors, a line buffer forms sliding windows for kernels up to 7 by 7, and an AXI4 256-bit DMA with a tile engine and double buffer streams data. A layer sequencer walks the model while activation, pooling and batchnorm stages close the pipeline.

APB CSR, Config Registers and Layer Sequencer FEED & BUFFERS COMPUTE POST-PROCESS Feature Buffer4096 deep, double buffered Weight Buffer4096 deep, 4 banks Line Bufferwindows to 1024 wide 16 x 16 PE Array 256 INT8 MACsINT32 accumulateWS, OS and RS dataflowdepthwise and pointwise ActivationLUT 256, SiLU and ReLU Poolingmax, avg, global BatchNorm and Skipadd, concat, upsample AXI4 256-bit Master DMA, Tile Engine and Double Buffer (to and from DDR)
Clean architectural view. Real parameters from the WIOWIZ CNN Accelerator RTL.

Specifications

CategoryAI accelerator, convolution engine
Compute16 by 16 systolic PE array, 256 INT8 MAC units
PrecisionINT8 activations and weights, INT32 accumulator and bias, INT8 quantized output
ConvolutionKernels up to 7 by 7, stride up to 4, padding up to 3, up to 1024 input and output channels
Feature memoryOn-chip feature buffer, 4096 deep, double buffered, 128-bit word
Weight memoryOn-chip weight buffer, 4096 deep, 4 banks, with active and shadow prefetch
Sliding windowLine buffer for images up to 1024 pixels wide
DataflowWeight-stationary, output-stationary and row-stationary modes
Layer operationsConv2D, depthwise, pointwise, pooling, upsample, concat, residual add, batchnorm, detection head
Post-processingActivation LUT 256 deep (ReLU, ReLU6, SiLU, sigmoid, tanh, leaky), pooling (max, avg, global)
Host interfaceAPB CSR (4K register space, 32-bit registers), AXI-Stream feature, weight and result ports
Memory interfaceAXI4 256-bit master DMA (32-bit address), tile engine, double buffer, SRAM macro port
TelemetryLayer sequencer, 8 performance counters (48-bit), PE and bandwidth utilization
VerificationUVM VIP, golden-model scoreboard, functional coverage
OwnershipWIOWIZ IP and WIOWIZ VIP

Key features

256 INT8 MACs

A 16 by 16 systolic PE array with INT32 accumulation for dense convolution.

On-chip buffers

Double-buffered feature memory and a 4-bank weight buffer keep tensors on-die.

Tiled AXI4 DMA

A 256-bit master DMA with a tile engine and double buffer streams data to and from DDR.

Selectable dataflow

Weight, output and row stationary modes adapt to layer shape and reuse.

Fused post-processing

Activation LUT, pooling, batchnorm and skip paths run inline after compute.

UVM-verified

Golden-model scoreboard and coverage, WIOWIZ IP and VIP.

Applications

Edge vision inferenceObject detectionImage segmentationOn-device perceptionSoC AI subsystem
Request IP details