Skip to content

Repository files navigation

SHA-3 Hardware Architectures for Post-Quantum Cryptography

This repository contains optimized Verilog HDL implementations of the Secure Hash Algorithm 3 (SHA-3) / Keccak permutation. These designs are structurally tuned to balance area, power, and throughput, making them ideal for accelerating Post-Quantum Cryptography (PQC) schemes like CRYSTALS-Kyber and Dilithium, which heavily rely on SHAKE functions.

These modules act as high-performance cryptographic accelerators and are well-suited for integration into complex hardware environments, such as secure 64-bit RISC-V System-on-Chip (SoC) designs utilizing Network-on-Chip (NoC) interconnects.

🏗️ Repository Structure

The project breaks down the Keccak-f[1600] permutation into its fundamental combinational steps and scales up into various architectural configurations:

1. Fundamental Step Functions

  • theta.v (θ): Achieves diffusion through parity XORing.
  • rho.v (ρ): Applies cyclic bit shifts within lanes.
  • pi.v (π): Rearranges lanes.
  • chi.v (χ): Introduces non-linearity.
  • iota.v (ι): Adds round-dependent constants.

2. Core Architectures

  • sha3_1round_pipeline.v: A single-round pipelined core. Pipeline registers are strategically inserted after the ρ (Rho) and ι (Iota) steps to isolate rotation-intensive paths and reduce critical path delay.
  • sha3_4round_pipeline.v: A partially unrolled core bundling 4 sequential rounds into a single block, utilizing intra-round pipelining to boost operating frequency.
  • sha3_6round_notpipelined.v: A wrapper module executing a sequence of rounds without intermediate pipelining, minimizing register area at the cost of lower throughput.
  • sha3_24round_pipeline.v: The top-level 24-round fully pipelined architecture, instantiating the 1-round pipeline to allow concurrent processing of multiple message blocks.

📊 Performance & Synthesis Results

These designs have been evaluated using ASIC implementation in the Cadence® Genus tool with a 180 nm technology slow library.

Key Findings:

  • Partially Unrolled & Pipelined: Offers the most balanced trade-off between resource utilization and performance. It achieves a higher operating frequency and better timing by approximately 68% compared to non-pipelined alternatives.
  • Iterative vs. Unrolled: Fully unrolled designs maximize throughput but consume excessive area. Iterative non-pipelined designs are compact but process slower.
  • Sub-pipelining Impact: Breaking the long combinational paths between the θ−ρ operations and π−χ−ι transformations effectively balances stage timing without breaking the hermetic sponge flow.

⚙️ Usage

These modules can be synthesized for both FPGA and ASIC targets. To use the top-level pipelined core, instantiate sha3_24round_pipeline.v and ensure all sub-modules (theta.v, rho.v, etc.) are included in your synthesis path.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages