The RZ/V2H is capable of responding to the evolution of AI and the sophisticated requirements of these applications. This article presents how the RZ/V2H solves heat generation problems, enables high-speed real-time processing, and achieves higher performance and lower power consumption for AI-equipped products.
Abstract:
As the working-age population shrinks due to declining birth rates and a growing elderly population, advanced artificial intelligence (AI) processing, such as environmental awareness, action decision-making, and motion control, will be necessary in various aspects of society, including factories, logistics, healthcare, urban service robots, and security cameras. Systems will need to handle this advanced AI processing in real time across various types of applications. Specifically, the system must be embedded within the device to enable rapid responses to its constantly changing environment. However, AI chips, while performing advanced AI processing in embedded devices, consume less power and are subject to strict limitations regarding heat generation.
To meet these market needs, Renesas has developed the DRP-AI3 (Dynamically Reconfigurable Processor for AI3) as an AI accelerator for high-speed AI inference processing, combining the low power consumption and flexibility required by cutting-edge devices. This reconfigurable AI accelerator processor technology, developed over many years, is embedded in the RZ/V series of MPUs designed for AI applications.
The RZ/V2H is a new high-end product in the RZ/V series, achieving approximately 10 times greater energy efficiency than its predecessors. The RZ/V2H is capable of responding to the evolving nature of AI and the sophisticated requirements of applications such as robotics. This article presents how the RZ/V2H addresses heat generation issues, enables high-speed real-time processing, and delivers higher performance and lower power consumption for AI-equipped products.
The DRP-AI3 accelerator efficiently processes debugged AI models.
As a typical technology for improving AI processing efficiency, debugging is available to omit calculations that do not significantly affect recognition accuracy. However, it is common for calculations that do not affect recognition accuracy to exist randomly in AI models. This creates a gap between the parallelism of hardware processing and the randomness of debugging, resulting in inefficient processing.
To solve this problem, Renesas optimized its unique DRP-based AI accelerator (DRP-AI) for debugging. By analyzing how debugging pattern characteristics and methods relate to recognition accuracy in typical image recognition AI models (CNN models), we identified an AI accelerator hardware structure capable of achieving both high recognition accuracy and efficient debugging, and applied it to the design of DRP-AI3. Furthermore, software was developed to reduce the size of the optimized AI models for DRP-AI3. This software converts the random debugging model configuration into highly efficient parallel computation, resulting in faster AI processing. In particular, Renesas' highly flexible debug support technology (N:M flexible debug technology), which can dynamically change the number of cycles in response to changes in the local debug rate in AI models, allows precise control of the debug rate according to the power consumption, operating speed, and recognition accuracy required by users.
Features of the heterogeneous architecture where DRP-AI3, DRP, and CPU operate cooperatively:
- Multithreaded and daisy-chained processing with the AI accelerator (DRP-AI3), DRP, and CPUs.
- High-speed, low-jitter robotic applications with DRP (dynamically reconfigurable hardwired logic hardware).
Service robots, for example, require advanced AI processing to recognize their surroundings. On the other hand, algorithm-based processing that doesn't use AI is also necessary to decide on and control the robot's behavior. However, current embedded processors (CPUs) lack sufficient resources to perform these different types of processing in real time. Renesas has solved this problem by developing a heterogeneous architecture technology that allows the Dynamically Reconfigurable Processor (DRP), the AI Accelerator (DRP-AI3), and the CPU to work together.
As shown in Figure 1, the Dynamically Reconfigurable Processor (DRP) can run applications while dynamically changing the circuit connection configuration of the chip's arithmetic units on each operating clock cycle, depending on the content to be processed. Because only the necessary arithmetic circuits are used, the DRP consumes less power than CPU processing and can achieve higher speeds. Furthermore, compared to CPUs, where frequent accesses to external memory due to cache leaks and other causes degrade performance, DRP can build the necessary data paths in hardware in advance, resulting in less performance degradation and less operating speed variation (jitter) due to memory accesses.
Distributed processing operations (DRPs) are particularly effective in processing data in a stream, such as image recognition, where parallelization and pipelining directly improve performance. On the other hand, programs like those for decision-making and robot behavior control require processing while conditions change and detailed processing in response to changes in the surrounding environment. Software CPU processing may be better suited for this than hardware processing, as in DRPs. It is important to distribute processing to the appropriate locations and operate in a coordinated manner. Renesas' heterogeneous architecture technology allows DRPs and CPUs to work together.

Figure 1: Characteristics of the dynamically reconfigurable flexible processor (DRP)
Figure 2 shows an overview of the MPU architecture and the AI accelerator (DRP-AI3). Robotic applications utilize a sophisticated combination of AI-based image recognition algorithms and non-AI-based decision and control algorithms. Therefore, a configuration with a DRP for AI processing (DRP-AI3) and a DRP for non-AI-based algorithms will significantly increase the performance of the robotic application.

Figure 2: Heterogeneous architecture configuration based on DRP-AI 3
Evaluation Results
(1) AI Model Processing Performance Evaluation
The RZ/V2H equipped with this technology has achieved a maximum of 8 TOPS (8 trillion sum-of-products operations per second) for AI accelerator processing performance. Furthermore, for AI models that have been debugged, the number of operation cycles can be reduced in proportion to the amount of debugging, resulting in AI model processing performance equivalent to a maximum of 80 TOPS compared to models before debugging. This is approximately 80 times higher than the processing performance of previous RZ/V products, a significant performance improvement that can adequately keep pace with the rapid evolution of AI (Figure 3).

Figure 3: Comparison of the maximum measured performance of DRP-AI3
On the one hand, as AI processing accelerates, image processing time based on non-AI algorithms, such as pre- and post-AI processing, is becoming a relative bottleneck. In AI-MPUs, part of the image processing program is offloaded to the DRP, which helps improve the overall system processing time. (Figure 4)

Figure 4: Heterogeneous architecture accelerates image recognition processing (as measured by the test chip)
In terms of energy efficiency, the AI accelerator's performance evaluation demonstrated the highest level of energy efficiency in the world (approximately 10 TOPS per watt) when running the main AI models. (Figure 5)

Figure 5: Energy efficiency of real AI models (measured by test chip)
We also demonstrated that the same real-time AI processing could be performed on an evaluation board equipped with the RZ/V2H, without forced ventilation, at temperatures comparable to those of competing products equipped with fans. (Figure 6)

Figure 6: Comparison of heat generation between an RZ/V2H board without forced ventilation and a GPU with a fan.
(2) Examples of applications with robots
For example, simultaneous localization and mapping (SLAM), one of the typical applications of robots, has a complex setup that requires multiple program processes for robot position recognition in parallel with environment recognition using AI processing. Renesas' DRP allows the robot to switch programs instantly, and parallel operation with an AI accelerator and the CPU has proven to be about 17 times faster than CPU-only operation, and reduces power consumption to 1/12 of the level of CPU-only operation.
In conclusion,
Renesas has developed RZ/V2H, a unique AI processor that combines the low power consumption and flexibility required by endpoints with processing capabilities for debugging AI models, and is 10 times more energy efficient (10 TOPS/W) than previous products.
Renesas will launch products in due course that address the evolution of AI, which is expected to become increasingly sophisticated, and will contribute to deploying systems that respond to endpoint products intelligently and in real time.
Author: Shingo Kojima, Principal Integrated Processing Engineer, Renesas Electronics
Related information
- RZ/V2H: https://www.renesas.com/rzv2h
- DRP-AI: Renesas' patented AI accelerator that combines high AI inference performance with low power consumption
