• Become a member
  • Log In
The Institution of Electronics
  • Home
  • About us
    • Our Objectives
    • Our History
    • Governance of the Institution
  • The Electron Magazine
    • 2024
      • 2024 – Winter
      • 2024 – Spring
      • 2024 – Summer
      • 2024 – Autumn
    • 2025
      • 2025 – Winter
      • 2025 – Spring
      • 2025 – Summer
      • 2025 – Autumn
    • 2026
      • 2026 – Winter
      • 2026 – Summer
  • Members
    • Membership Grades and Fees
    • Members’ Resources
      • The Electron Newsletter
      • The Archives
  • Education and Projects
    • National Electronics Competition
    • Student Members’ Projects
    • Arkwright Engineering Scholarships
  • News
  • Contact Us
  • Menu Menu
Uncategorised

The role of cache in AI processor design

Artificial intelligence (AI) is making its presence felt everywhere these days, from the data centers at the Internet’s core to sensors and handheld devices like smartphones at the Internet’s edge and every point in between, such as autonomous robots and vehicles. For the purposes of this article, we recognize the term AI to embrace machine learning and deep learning.

There are two main aspects to AI: training, which is predominantly performed in data centers, and inferencing, which may be performed anywhere from the cloud down to the humblest AI-equipped sensor.

AI is a greedy consumer of two things: computational processing power and data. In the case of processing power, OpenAI, the creator of ChatGPT, published the report AI and Compute, showing that since 2012, the amount of compute used in large AI training runs has doubled every 3.4 months with no indication of slowing down.

With respect to memory, a large generative AI (GenAI) model like ChatGPT-4 may have more than a trillion parameters, all of which need to be easily accessible in a way that allows to handle numerous requests simultaneously. In addition, one needs to consider the vast amounts of data that need to be streamed and processed.

Slow speed

Suppose we are designing a system-on-chip (SoC) device that contains one or more processor cores. We will include a relatively small amount of memory inside the device, while the bulk of the memory will reside in discrete devices outside the SoC.

The fastest type of memory is SRAM, but each SRAM cell requires six transistors, so SRAM is used sparingly inside the SoC because it consumes a tremendous amount of space and power. By comparison, DRAM requires only one transistor and capacitor per cell, which means it consumes much less space and power. Therefore, DRAM is used to create bulk storage devices outside the SoC. Although DRAM offers high capacity, it is significantly slower than SRAM.

As the process technologies used to develop integrated circuits have evolved to create smaller and smaller structures, most devices have become faster and faster. Sadly, this is not the case with the transistor-capacitor bit-cells that lie at the heart of DRAMs. In fact, due to their analog nature, the speed of bit-cells has remained largely unchanged for decades.

Having said this, the speed of DRAMs, as seen at their external interfaces, has doubled with each new generation. Since each internal access is relatively slow, the way this has been achieved is to perform a series of staggered accesses inside the device. If we assume we are reading a series of consecutive words of data, it will take a relatively long time to receive the first word, but we will see any succeeding words much faster.

This works well if we wish to stream large blocks of contiguous data because we take a one-time hit at the start of the transfer, after which subsequent accesses come at high speed. However, problems occur if we wish to perform multiple accesses to smaller chunks of data. In this case, instead of a one-time hit, we take that hit over and over again.

More speed

The solution is to use high-speed SRAM to create local cache memories inside the processing device. When the processor first requests data from the DRAM, a copy of that data is stored in the processor’s cache. If the processor subsequently wishes to re-access the same data, it uses its local copy, which can be accessed much faster.

It’s common to employ multiple levels of cache inside the SoC. These are called Level 1 (L1), Level 2 (L2), and Level 3 (L3). The first cache level has the smallest capacity but the highest access speed, with each subsequent level having a higher capacity and a lower access speed. As illustrated in Figure 1, assuming a 1-GHz system clock and DDR4 DRAMs, it takes only 1.8 ns for the processor to access its L1 cache, 6.4 ns to access the L2 cache, and 26 ns to access the L3 cache. Accessing the first in a series of data words from the external DRAMs takes a whopping 70 ns (Data source Joe Chang’s Server Analysis).

Figure 1 Cache and DRAM access speeds are outlined for 1 GHz clock and DDR4 DRAM. Source: Arteris

The role of cache in AI

There are a wide variety of AI implementation and deployment scenarios. In the case of our SoC, one possibility is to create one or more AI accelerator IPs, each containing its own internal caches. Suppose we wish to maintain cache coherence, which we can think of as keeping all copies of the data the same, with the SoCs processor clusters. Then, we will have to use a hardware cache-coherent solution in the form of a coherent interconnect, like CHI as defined in the AMBA specification and supported by Ncore network-on-chip (NoC) IP from Arteris IP (Figure 2a).

Figure 2 The above diagram shows examples of cache in the context of AI. Source: Arteris

There is an overhead associated with maintaining cache coherence. In many cases, the AI accelerators do not need to remain cache coherent to the same extent as the processor clusters. For example, it may be that only after a large block of data has been processed by the accelerator that things need to be re-synchronized, which can be achieved under software control. The AI accelerators could employ a smaller, faster interconnect solution, such as AXI from Arm or FlexNoC from Arteris (Figure 2b).

In many cases, the developers of the accelerator IPs do not include cache in their implementation. Sometimes, the need for cache wasn’t recognized until performance evaluations began. One solution is to include a special cache IP between an AI accelerator and the interconnect to provide an IP-level performance boost (Figure 2c). Another possibility is to employ the cache IP as a last-level cache to provide an SoC-level performance boost (Figure 2d). Cache design isn’t easy, but designers can use configurable off-the-shelf solutions.

Many SoC designers tend to think of cache only in the context of processors and processor clusters. However, the advantages of cache are equally applicable to many other complex IPs, including AI accelerators. As a result, the developers of AI-centric SoCs are increasingly evaluating and deploying a variety of cache-enabled AI scenarios.

Frank Schirrmeister, VP solutions and business development at Arteris, leads activities in the automotive, data center, 5G/6G communications, mobile, aerospace and data center industry verticals. Before Arteris, Frank held various senior leadership positions at Cadence Design Systems, Synopsys and Imperas.

Related Content

Verifying Cache Coherence
Cache Coherence Issues for Real-Time Multiprocessing
SoC design: When is a network-on-chip (NoC) not enough?
Efficient checks for cache-coherency verification in complex SoCs
Fast, Thorough Verification of Multiprocessor SoC Cache Coherency

<!–
googletag.cmd.push(function() { googletag.display(‘div-gpt-ad-native’); });
–>

The post The role of cache in AI processor design appeared first on EDN.

22 March 2024
http://institutionofelectronics.ac.uk/wp-content/uploads/2022/12/IOE_LOGO.png 0 0 http://institutionofelectronics.ac.uk/wp-content/uploads/2022/12/IOE_LOGO.png 2024-03-22 07:47:562024-03-22 07:47:56The role of cache in AI processor design

Latest news

  • Measurement bandwidth9 October 2026 - 13:56
  • Practical design for a multi-output flyback converter with improved cross-regulation9 October 2026 - 08:50
  • The multi-gig Ethernet migration: Motivations and implementations8 October 2026 - 13:29
  • GaN FET meets megawatt-scale demands7 October 2026 - 20:15
  • FPGA IP accelerates deterministic networking7 October 2026 - 20:15
  • Microphone array advances acoustic detection7 October 2026 - 20:15
  • Quad beamformer simplifies X-band radar design7 October 2026 - 20:15
  • Global-shutter sensor raises pixel density7 October 2026 - 20:15
  • An unbuttoned circuit for setting digital pots7 October 2026 - 13:10
  • Training and inference: Two faces of AI compute7 October 2026 - 11:08
IOE LOGO 2

Become a member

click here

Become a member

click here

Become a subscriber

click here

Become a sponsor

click here

© Copyright - The Institution of Electronics | Website by WHD Solutions
  • Link to LinkedIn
  • Link to Facebook
  • Link to X
Link to: Workarounds (and their tradeoffs) for integrated storage constraints Link to: Workarounds (and their tradeoffs) for integrated storage constraints Workarounds (and their tradeoffs) for integrated storage constraints Link to: 2-A Schottky rectifiers occupy tiny footprint Link to: 2-A Schottky rectifiers occupy tiny footprint 2-A Schottky rectifiers occupy tiny footprint
Scroll to top Scroll to top Scroll to top