SOFT IP / AI ACCELERATORS / NPU

NPU, Neural Processing Unit

Systolic INT8 inference accelerator with a transformer sequencer, WIOWIZ IP and WIOWIZ VIP

A systolic INT8 array driven by a transformer sequencer

The WIOWIZ NPU is a systolic-array inference accelerator built around an INT8 multiply-accumulate array with INT32 accumulation and a weight-stationary dataflow. A hardware transformer sequencer walks attention, LayerNorm and feed-forward layers while a host command mailbox dispatches descriptors in order. Weights and activations stay on-chip in banked, double-buffered SRAM so compute and DMA overlap, and an AXI master moves tensors to and from external memory with the host path built on the WIOWIZ PCIe block. The IP ships as a family: arka is the frozen reference, v1 is the IP-sovereign design shipped today, and v2 extends the datapath to mixed INT4, INT8 and FP8 precision with state-space streaming, sparse attention and Mixture-of-Experts routing.

PCIe Host Interface, Command Mailbox and Transformer Sequencer FEED & MEMORY COMPUTE POST-PROCESS AXI Master DMAto external memory Weight SRAMbanked, parity checked Activation SRAMdouble buffered Systolic INT8MAC Array weight-stationaryINT32 accumulatetransformer sequencedstructured sparsity LayerNorminteger, inv-sqrt LUT Activation and GELULUT based Quantizerper-channel INT8 and INT4 L2 Datapath with SECDED ECC, Watchdog and Trace Bus
Clean architectural view. Real parameters from the WIOWIZ NPU RTL.

Specifications

CategoryAI accelerator, Neural Processing Unit
FamilyThree generations: arka frozen reference, v1 IP-sovereign design shipped today, v2 in specification
ComputeSystolic INT8 multiply-accumulate array, weight-stationary dataflow, INT32 accumulation
PrecisionINT8 activations and weights in v1; mixed INT4, INT8 and FP8 planned for v2
SequencerHardware transformer sequencer, WIOWIZ-authored in v1, walks attention, LayerNorm and feed-forward layers
On-chip memoryBanked weight and activation SRAM, double buffered for compute and DMA overlap
Command modelHost command mailbox with an in-order descriptor ring
Host interfacePCIe host path on the WIOWIZ PCIe block, APB CSR configuration
Memory interfaceAXI master DMA to external memory
Post-processingInteger LayerNorm with inverse-square-root LUT, activation and GELU LUT, per-channel quantization
ReliabilitySECDED ECC on the L2 datapath, per-bank parity, watchdog timer, built-in self-test, bank write-protect
SparsityStructured zero-skip, verified numerically identical to dense compute
v2 architectureState-space (SSM) streaming, sparse attention, Mixture-of-Experts routing, reconfigurable per-layer dataflow
VerificationNative FSimX, golden reference model, scoreboard, functional coverage, error-injection runs
IP sovereigntyEvery RTL line WIOWIZ-authored or referencing a sibling WIOWIZ IP or RISC-V core
OwnershipWIOWIZ IP and WIOWIZ VIP

Key features

Systolic INT8 array

A weight-stationary INT8 multiply-accumulate array with INT32 accumulation runs dense and batched GEMM.

Transformer sequencer

A hardware sequencer walks attention, LayerNorm and feed-forward layers with no host CPU in the compute loop.

Double-buffered SRAM

Banked on-chip weight and activation memory lets compute read one buffer while DMA fills the other.

IP-sovereign RTL

Every line is WIOWIZ-authored or references a sibling WIOWIZ IP or RISC-V core; the host path uses the WIOWIZ PCIe block.

Reliability core

SECDED ECC on the L2 datapath, per-bank parity, watchdog and built-in self-test guard the memory system.

Verified to a golden model

Native FSimX regression, a golden reference model and functional coverage check GEMM, attention, LayerNorm, FFN and sparsity.

Applications

Edge AI inferenceTransformer inferenceOn-device language modelsADAS perceptionSoC AI subsystem
Request IP details