设计一种基于数字迭代算法的多线程、高性能、可配置的超越函数的硬件单元,支持正余弦、反正切、求模长、指数和对数的计算,可配置4~24 bit定点小数精度。该设计使用SMIC 40 nm eFlash平台的标准单元库进行综合,最终实现了200 MHz的时钟频率,面积为301 074 μm2。
宋敏特
,
刘楠
,
茹占强
,
殷志珍
,
丁朋
,
王争光
,
程素珍
,
宋贺伦
. 面向工控MCU的超越函数单元设计[J]. 中国科学院大学学报, 2025
, 42(2)
: 260
-267
.
DOI: 10.7523/j.ucas.2023.009
The calculation of transcendental functions is one of the necessary steps in industrial control algorithms. As the complexity of industrial control systems increases, calculating the transcendental function by software approximation algorithm takes up a large number of CPU cycles, compressing the computational resources of real-time control algorithms and reducing the accuracy of closed-loop control. Equipped with hardware accelerating units, industrial microcontroller unit architecture becomes the preferred solution to solve this contradiction. In this paper, a multi-threaded, high-performance, configurable transcendental function unit based on digital iterative algorithms was designed, which supports trigonometric function, exponential, and logarithmic calculations. The design was synthesized by standard cell library of SMIC 40 nm eFlash platform, resulting in a clock frequency of 200 MHz and an area of 301 074 μm2.
[1] Kang Y G, Diaz Reigosa D. Dq-transformed error and current sensing error effects on self-sensing control[J]. IEEE Journal of Emerging and Selected Topics in Power Electronics, 2022, 10(2): 1935-1945. DOI: 10.1109/JESTPE.2021.3051942.
[2] Texas Instruments. Digital control library[DB/OL]. (2022-04-27)[2022-12-20].https://training.ti.com/node/1149587.
[3] Ramezannezhad A, Naderi P, Vandevelde L. A novel method for accuracy improvement of variable reluctance linear resolvers[J]. IEEE Sensors Journal, 2022, 22(19): 18409-18417. DOI: 10.1109/JSEN.2022.3199807.
[4] Rodriguez J, Lai J S, Peng F Z. Multilevel inverters: a survey of topologies, controls, and applications[J]. IEEE Transactions on Industrial Electronics, 2002, 49(4): 724-738. DOI: 10.1109/TIE.2002.801052.
[5] Kantabutra V. On hardware for computing exponential and trigonometric functions[J]. IEEE Transactions on Computers, 1996, 45(3): 328-339. DOI: 10.1109/12.485571.
[6] Wang D, Muller J M, Brisebarre N, et al. (M, p, k)-friendly points: a table-based method to evaluate trigonometric function[J]. IEEE Transactions on Circuits and Systems II: Express Briefs, 2014, 61(9): 711-715. DOI: 10.1109/TCSII.2014.2331094.
[7] Tang P T P. Table-lookup algorithms for elementary functions and their error analysis[C]//[1991] Proceedings 10th IEEE Symposium on Computer Arithmetic. June 26-28, 1991, Grenoble, France. IEEE, 2002: 232-236. DOI: 10.1109/ARITH.1991.145565.
[8] Zhu B Z, Lei Y W, Peng Y X, et al. Low latency and low error floating-point sine/cosine function based TCORDIC algorithm[J]. IEEE Transactions on Circuits and Systems I: Regular Papers, 2017, 64(4): 892-905. DOI: 10.1109/TCSI.2016.2631588.
[9] Koren I, Zinaty O. Evaluating elementary functions in a numerical coprocessor based on rational approximations[J]. IEEE Transactions on Computers, 1990, 39(8): 1030-1037. DOI: 10.1109/12.57042.
[10] Volder J E. The CORDIC trigonometric computing technique[J]. IRE Transactions on Electronic Computers, 1959, EC-8(3): 330-334. DOI: 10.1109/tec.1959.5222693.
[11] Walther J S. A unified algorithm for elementary functions[C]//Proceedings of the May 18-20, 1971, spring joint computer conference. May 18 - 20, 1971, Atlantic City, New Jersey. New York: ACM, 1971: 379-385. DOI: 10.1145/1478786.1478840.
[12] Meher P K, Valls J, Juang T B, et al. 50 years of CORDIC: algorithms, architectures, and applications[J]. IEEE Transactions on Circuits and Systems I: Regular Papers, 2009, 56(9): 1893-1907. DOI: 10.1109/tcsi.2009.2025803.
[13] Juang T B. Low latency angle recoding methods for the higher bit-width parallel CORDIC rotator implementations[J]. IEEE Transactions on Circuits and Systems II: Express Briefs, 2008, 55(11): 1139-1143. DOI: 10.1109/TCSII.2008.2002566.
[14] Rodrigues T K, Swartzlander E E. Adaptive CORDIC: using parallel angle recoding to accelerate rotations[J]. IEEE Transactions on Computers, 2010, 59(4): 522-531. DOI: 10.1109/TC.2009.190.
[15] Chandrakanth Y, Praveen Kumar M. Low latency & high precision CORDIC architecture using improved parallel angle recoding[C]//2011 International Conference on Signal Processing, Communication, Computing and Networking Technologies. July 21-22, 2011, Thuckalay, India. IEEE, 2011: 498-501. DOI: 10.1109/ICSCCN.2011.6024602.
[16] Mohamed S M, Sayed W S, Radwan A G, et al. FPGA implementation of reconfigurable CORDIC algorithm and a memristive chaotic system with transcendental nonlinearities[J]. IEEE Transactions on Circuits and Systems I: Regular Papers, 2022, 69(7): 2885-2892. DOI: 10.1109/TCSI.2022.3165469.
[17] Hoang T T, Nguyen X T, Le D H, et al. Low-power floating-point adaptive-CORDIC-based FFT twiddle factor on 65-nm silicon-on-thin-BOX (SOTB) with back-gate bias[J]. IEEE Transactions on Circuits and Systems II: Express Briefs, 2019, 66(10): 1723-1727. DOI: 10.1109/TCSII.2019.2928138.
[18] 吴庆达, 何书专, 潘红兵, 等. 32位定浮点数正余弦函数FPGA实现方法[J]. 微电子学与计算机, 2012, 29(1): 113-116. DOI: 10.19304/j.cnki.issn1000-7180.2012.01.027.
[19] 李天立, 尹韬, 魏星, 等. 高效单精度浮点三角函数计算电路结构与实现[J]. 微电子学与计算机, 2018, 35(12): 33-37. DOI: 10.19304/j.cnki.issn1000-7180.2018.12.007.
[20] Nguyen H T, Nguyen X T, Pham C K. A low-power hybrid adaptive CORDIC[J]. IEEE Transactions on Circuits and Systems II: Express Briefs, 2018, 65(4): 496-500. DOI: 10.1109/TCSII.2017.2732451.