Resource-efficient and latency-optimized hardware accelerator for Koblitz-curve point multiplication
Online published: 2026-07-09
Copyright
This paper investigates hardware acceleration architectures for point multiplication to address the computational efficiency issues of Koblitz curves in elliptic curve cryptography (ECC) systems. An optimized τNAF scalar conversion algorithm and its corresponding hardware structure are first proposed to reduce latency in the pre-computation phase. On this basis, two computation architectures with optimized pipeline efficiency are designed. The area-efficient architecture employs a compact four-stage pipeline with a single multiplier to improve hardware resource utilization. For different binary fields (GF(2163), GF(2283), and GF(2571)), the low-latency architecture adopts two-stage and three-stage pipeline designs respectively, and leverages dual parallel multipliers to reduce the clock cycles required for point addition. Experimental results on the Virtex-7 FPGA platform demonstrate that both architectures achieve substantial latency improvements across all three fields. Compared with state-of-the-art designs, the area-efficient architecture reduces latency by 58.111% over GF(2571), while the low-latency architecture reduces latency by 43.952% over GF(2163). The results demonstrate that optimizing pipeline scheduling and algorithm mapping can effectively balance the computational performance and resource consumption of ECC coprocessors, providing a practical reference for the design of high-performance cryptographic hardware.
Zhang Jinlei , Sun Weitian , Jiang Yujie , Wang An , Zhang Jingqi , Yang Xianming , Hao Yue . Resource-efficient and latency-optimized hardware accelerator for Koblitz-curve point multiplication[J]. Journal of Cybersecurity, 2026 , 4(3) : 64 -79 . DOI: 10.20172/j.issn.2097-3136.260606
表 1 基于 Virtex-7 的标量转换实现结果Table 1 Implementation result of scalar converter on Virtex-7 |
| m | 频率/MHz | CC | 时延/μs | #LUT | #FF | #Slice | ATP |
| 163 | 213 | 84 | 0.394 | 1096 | 503 | 299 | 118 |
| 283 | 162 | 144 | 0.889 | 1823 | 863 | 506 | 450 |
| 571 | 138 | 288 | 2.087 | 3550 | 1738 | 986 | 2058 |
表 2 基于 Virtex-7 的点乘实现结果Table 2 Implementation result of point multiplication on VIRTEX-7 |
| 域长 | 架构 | 流水线 | 频率/MHz | 周期数 | 时延/μs | LUT | Slice | ATP |
| 163 | 1M | 2S | 192.2 | 435 | 2.26 | 11007 | 3154 | 7137 |
| 3S | 293.3 | 435 | 1.483 | 11639 | 3379 | 5012 | ||
| 4S | 349.7 | 435 | 1.244 | 11791 | 3365 | 4186 | ||
| 5S | 400.0 | 543 | 1.358 | 11798 | 3297 | 4476 | ||
| 2M | 2S | 203.4 | 217 | 1.067 | 19448 | 5390 | 5750 | |
| 3S | 291.4 | 326 | 1.119 | 20474 | 5702 | 6380 | ||
| 4S | 350.9 | 435 | 1.240 | 20560 | 5762 | 7143 | ||
| 283 | 1M | 2S | 162.3 | 755 | 4.651 | 22637 | 6886 | 32025 |
| 3S | 245.7 | 755 | 3.073 | 23656 | 6622 | 20348 | ||
| 4S | 294.1 | 755 | 2.567 | 23445 | 6628 | 17014 | ||
| 5S | 337.8 | 943 | 2.791 | 25065 | 7258 | 20259 | ||
| 2M | 2S | 157.2 | 377 | 2.398 | 42737 | 11619 | 27859 | |
| 3S | 238.1 | 566 | 2.377 | 45685 | 12610 | 29976 | ||
| 4S | 305.8 | 755 | 2.469 | 44447 | 11961 | 29530 | ||
| 571 | 1M | 2S | 115.7 | 1523 | 13.159 | 61635 | 16716 | 219961 |
| 3S | 196.9 | 1523 | 7.737 | 65348 | 19401 | 150102 | ||
| 4S | 227.3 | 1523 | 6.701 | 63842 | 18850 | 126318 | ||
| 5S | 264.6 | 1903 | 7.193 | 65612 | 19998 | 143852 | ||
| 2M | 2S | 114.2 | 761 | 6.666 | 125526 | 33891 | 225930 | |
| 3S | 180.5 | 1142 | 6.327 | 124262 | 33982 | 214993 | ||
| 4S | 238.1 | 1523 | 6.397 | 123415 | 35921 | 229772 |
表 3 面积高效型架构点乘部分寄存器级调度方案Table 3 Register-level scheduling scheme of dot product for area-efficient architecture |
| 时钟 周期 | 乘法器 | 加法器 | 平方器 | ||||||
| MUL_0 | MUL_1 | ADD0_0 | ADD0_1 | ADD1_0 | ADD1_1 | SQR | |||
| 1 | x2(0) | Z1(0) | R1(−1) | G(−1) | — | — | — | ||
| 2 | RCH(−1) | R0(−1) | G(−1) | Y1(−1) | — | — | — | ||
| 3 | RCH(−1) | R0(−1) | MUL_r(−2) | J(−2) | — | — | R0(−1) | ||
| 4 | x2(−1) | Z3(−1) | MUL_r(0) | X1(0) | — | — | — | ||
| 5 | Z1(0) | R0(0) | MUL_r(−1) | R1(−1) | — | — | Z1(0) | ||
| 6 | y2(0) | R1(0) | MUL_r(−1) | R1(−1) | x2(−1) | y2(−1) | Z3(−1) | ||
| 7 | R2(−1) | R1(−1) | MUL_r(−1) | X3(−1) | — | — | R0(0) | ||
| 8 | RCH(−1) | R1(−1) | MUL_r(0) | Y1(0) | RCH(0) | R0(0) | RCH(0) | ||
| 9 | x2(1) | Z1(1) | R1(0) | G(0) | — | — | — | ||
| 10 | RCH(0) | R0(0) | G(0) | Y1(0) | — | — | — | ||
| 11 | RCH(0) | R0(0) | MUL_r(−1) | J(−1) | — | — | R0(0) | ||
| 12 | x2(0) | Z3(0) | MUL_r(1) | X1(1) | — | — | — | ||
| 13 | Z1(1) | R0(1) | MUL_r(0) | R1(0) | — | — | Z1(1) | ||
| 14 | y2(1) | R1(1) | MUL_r(0) | R1(0) | x2(0) | y2(0) | Z3(0) | ||
| 15 | R2(0) | R1(0) | MUL_r(0) | X3(0) | — | — | R0(1) | ||
| 16 | RCH(0) | R1(0) | MUL_r(1) | Y1(1) | RCH(1) | R0(1) | RCH(1) | ||
| 17 | x2(2) | Z1(2) | R1(1) | G(1) | — | — | — | ||
| 18 | RCH(1) | R0(1) | G(1) | Y1(1) | — | — | — | ||
| 19 | RCH(1) | R0(1) | MUL_r(0) | J(0) | — | — | R0(1) | ||
(−2)表示上上一轮;(−1)表示上一轮;(0)表示当前轮;(1)表示下一轮;(2)表示下下一轮。 |
表 4 实验结果对比Table 4 Comparison of experimental results |
| 曲线 | 文献 | 域(m) | 架构 | 设备 | 总计 | ||
| 时延/μs | #Slice | ATP | |||||
| Koblitz曲线 | [13] | 163 | — | V5 | 2.505 | 3670 | 9193 |
| — | V4 | 3.402 | 7732 | 26302 | |||
| 283 | — | V5 | 5.815 | 7738 | 44993 | ||
| 571 | — | V5 | 18.513 | 20291 | 375642 | ||
| [26] | 163 | — | V5 | 5.262 | 23975 | 126158 | |
| 本文 | 163 | 1M | V7 | 1.683 | 3631 | 6110 | |
| V5 | 1.975 | 4362 | 8615 | ||||
| V4 | 2.955 | 5571 | 16462 | ||||
| 2M | V7 | 1.347 | 6026 | 8115 | |||
| V5 | 1.404 | 6917 | 9712 | ||||
| V4 | 2.092 | 9019 | 18868 | ||||
| 283 | 1M | V7 | 3.455 | 7867 | 27184 | ||
| V5 | 3.548 | 9027 | 32029 | ||||
| V4 | 5.500 | 12379 | 68086 | ||||
| 2M | V7 | 3.279 | 14246 | 46713 | |||
| V5 | 3.346 | 14807 | 49544 | ||||
| V4 | 5.408 | 19597 | 105974 | ||||
| 571 | 1M | V7 | 7.511 | 20612 | 154821 | ||
| V5 | 7.755 | 24625 | 190962 | ||||
| V4 | 12.199 | 32630 | 398065 | ||||
| 2M | V7 | 7.071 | 38515 | 272325 | |||
| V5 | 7.457 | 42827 | 319380 | ||||
| V4 | 11.349 | 55726 | 632429 | ||||
| General曲线 | [19] | 163 | d=4 | V7 | 2.930 | 4435 | 12995 |
| d=6 | V7 | 2.450 | 5705 | 13977 | |||
| 283 | d=4 | V7 | 8.240 | 7096 | 58471 | ||
| d=6 | V7 | 6.750 | 8951 | 60419 | |||
| 571 | d=4 | V7 | 28.800 | 13789 | 397123 | ||
| d=6 | V7 | 23.020 | 16359 | 376584 | |||
| [27] | 163 | — | V7 | 2.068 | 8762 | 18122 | |
| 283 | — | V7 | 4.095 | 20451 | 83739 | ||
| 571 | — | V7 | 9.719 | 41974 | 407950 | ||
| [20] | 163 | — | V7 | 3.710 | 2437 | 9041 | |
| V5 | 5.220 | 2502 | 13062 | ||||
| 283 | — | V7 | 7.690 | 5493 | 42240 | ||
| V5 | 10.790 | 5640 | 60859 | ||||
| [28] | 163 | ∞ | V7 | 2.510 | 3422 | 8591 | |
| 1 | V7 | 4.832 | 3422 | 16536 | |||
| 283 | ∞ | V7 | 4.932 | 7983 | 39374 | ||
| 1 | V7 | 9.331 | 7983 | 74487 | |||
| 571 | ∞ | V7 | 10.851 | 20158 | 218732 | ||
| 1 | V7 | 20.377 | 20158 | 410763 | |||
| [29] | 163 | — | V5 | 3.900 | 3590 | 14002 | |
表 5 本文架构在 Virtex-7 上的功耗Table 5 Power consumption of the proposed architectures on Virtex-7 |
| 域(m) | 架构 | τNAF的功率/W | KP的功率/W | 总功率/W |
| 163 | 1M | 0.383 | 1.620 | 2.003 |
| 2M | 0.383 | 2.693 | 3.076 | |
| 283 | 1M | 0.399 | 2.979 | 3.378 |
| 2M | 0.399 | 8.679 | 9.078 | |
| 571 | 1M | 0.416 | 8.938 | 9.354 |
| 2M | 0.416 | 19.188 | 19.604 |
| 1 |
Garg S, Kaur K, Kaddoum G, et al. Toward secure and provable authentication for internet of things: realizing industry 4.0[J]. IEEE Internet of Things Journal, 2020, 7 (5): 4598- 4606.
|
| 2 |
王凯, 董建阔, 肖甫, 等. 面向物联网的认证密钥协商协议研究综述[J]. 网络空间安全科学学报, 2024, 2 (5): 2- 16.
Wang K, Dong J K, Xiao F, et al. Review of research on authentication key agreement protocols for internet of things[J]. Journal of Cybersecurity, 2024, 2 (5): 2- 16.
|
| 3 |
Koblitz N. Elliptic curve cryptosystems[J]. Mathematics of Computation, 1987, 48 (177): 203- 209.
|
| 4 |
Miller V S. Use of elliptic curves in cryptography[M]//Advances in Cryptology — CRYPTO ’85 Proceedings. Berlin, HeidelbergSpringer2007: 417-426.
|
| 5 |
Chen A C H, Lin B Y. Hybrid scheme of post-quantum cryptography and elliptic-curve cryptography for certificates ─ a case study of security credential management system in vehicle-to-everything communications[C]//Proceedings of the 2024 7th International Conference on Circuit Power and Computing Technologies (ICCPCT). Piscataway: IEEE Press, 2024: 426-430.
|
| 6 |
Koblitz N. CM-curves with good cryptographic properties[M]//Advances in Cryptology — CRYPTO ’91. Berlin, HeidelbergSpringer, 2007: 279-287.
|
| 7 |
Järvinen K U, Skyttä J O. Fast point multiplication on Koblitz curves: parallelization method and implementations[J]. Microprocessors and Microsystems, 2009, 33 (2): 106- 116.
|
| 8 |
Azarderakhsh R, Reyhani M A. High-performance implementation of point multiplication on Koblitz curves[J]. IEEE Transactions on Circuits and Systems II: Express Briefs, 2013, 60 (1): 41- 45.
|
| 9 |
Roy S S, Rebeiro C, Mukhopadhyay D. A parallel architecture for Koblitz curve scalar multiplications on FPGA platforms[C]//Proceedings of the 2012 15th Euromicro Conference on Digital System Design. Piscataway: IEEE Press, 2012: 553-559.
|
| 10 |
Wang T, Liu T T. ECC processor over the Koblitz curves with τ-NAF converter and square-square-add algorithm[C]//Proceedings of the 2020 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS). Piscataway: IEEE Press, 2020: 23-26.
|
| 11 |
Järvinen K U, Skyttä J O. High-speed elliptic curve cryptography accelerator for Koblitz curves[C]//Proceedings of the 2008 16th International Symposium on Field-Programmable Custom Computing Machines. Piscataway: IEEE Press, 2008: 109-118.
|
| 12 |
Vuillaume C, Okeya K, Takagi T. Short-memory scalar multiplication for Koblitz curves[J]. IEEE Transactions on Computers, 2008, 57 (4): 481- 489.
|
| 13 |
Li L J, Li S G. High-performance pipelined architecture of point multiplication on Koblitz curves[J]. IEEE Transactions on Circuits and Systems II: Express Briefs, 2018, 65 (11): 1723- 1727.
|
| 14 |
Zode P, Deshmukh R. Optimization of elliptic curve scalar multiplication using constraint based scheduling[J]. Journal of Parallel and Distributed Computing, 2022, 167, 232- 239.
|
| 15 |
Alharbi A R, Hazzazi M M, Jamal S S, et al. DCryp-unit: crypto hardware accelerator unit design for elliptic curve point multiplication[J]. IEEE Access, 2024, 12, 17823- 17835.
|
| 16 |
Rashid M, Sonbul O S, Zia M Y I, et al. Throughput/area-efficient accelerator of elliptic curve point multiplication over GF(2233) on FPGA[J]. Electronics, 2023, 12 (17): 3611.
|
| 17 |
Heidarpur M, Mirhassani M. An efficient and high-speed overlap-free karatsuba-based finite-field multiplier for FGPA implementation[J]. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2021, 29 (4): 667- 676.
|
| 18 |
Thirumoorthi M, Leigh A J, Heidarpur M, et al. Novel formulations of M-term overlap-free karatsuba binary polynomial multipliers and their hardware implementations[J]. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2023, 31 (10): 1509- 1522.
|
| 19 |
Zeghid M, Ahmed H Y, Chehri A, et al. Speed/area-efficient ECC processor implementation over GF(2m) on FPGA via novel algorithm-architecture co-design[J]. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2023, 31(8): 1192-1203.
|
| 20 |
Kumar N R, Shirisha C. ECC architecture over GF(2m) for resource-constrained applications[J]. AEU - International Journal of Electronics and Communications, 2020, 125: 153383.
|
| 21 |
Nadikuda P K G, Boppana L. An area-time efficient point-multiplication architecture for ECC over GF(2m) using polynomial basis[J]. Microprocessors and Microsystems, 2022, 91, 104525.
|
| 22 |
Solinas J A. Efficient arithmetic on Koblitz curves[M]// Towards a Quarter-Century of Public Key Cryptography. Boston, MA: Springer US, 2000: 125-179.
|
| 23 |
Meier W, Staffelbach O. Efficient multiplication on certain nonsupersingular elliptic curves[C]//Advances in Cryptology — CRYPTO’ 92. Berlin, Heidelberg: Springer, 1993: 333-344.
|
| 24 |
Brumley B B, Jarvinen K U. Conversion algorithms and implementations for Koblitz curve cryptography[J]. IEEE Transactions on Computers, 2010, 59 (1): 81- 92.
|
| 25 |
Adikari J, Dimitrov V S, Jarvinen K U. A fast hardware architecture for integer to τNAF conversion for Koblitz curves[J]. IEEE Transactions on Computers, 2012, 61 (5): 732- 737.
|
| 26 |
Tian X M, Ding R, Wu X J, et al. Hardware implementation of a cryptographically secure pseudo-random number generators based on Koblitz elliptic curves[C]//Proceedings of the 2020 IEEE 3rd International Conference on Electronics Technology (ICET). Piscataway: IEEE Press, 2020: 91-94.
|
| 27 |
Zhang J Q, Chen Z M, Ma M Z, et al. High-performance elliptic curve scalar multiplication architecture based on interleaved mechanism[J]. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2025, 33 (3): 757- 770.
|
| 28 |
Zhang J Q, Chen Z M, Ma M Z, et al. High-performance ECC scalar multiplication architecture based on comb method and low-latency window recoding algorithm[J]. IEEE Transactions on Very Large Scale Integration (VLSI) Systems, 2024, 32 (2): 382- 395.
|
| 29 |
Nadikuda P K G, Boppana L. Low area-time complexity point multiplication architecture for ECC over GF(2m) using polynomial basis[J]. Journal of Cryptographic Engineering, 2023, 13 (1): 107- 123.
|
/
| 〈 |
|
〉 |