DistantNews
Support us
AI Chip War Shifts to Inference: d-Matrix Tackles the 'Memory Wall'
๐Ÿ‡ฐ๐Ÿ‡ท South Korea /Technology

AI Chip War Shifts to Inference: d-Matrix Tackles the 'Memory Wall'

From Dong-A Ilbo · () Korean

Translated from Korean, summarized and contextualized by DistantNews.

At a glance

In-depth Named sources Context piece
  • AI chip company d-Matrix is shifting focus from AI model training to inference, aiming to reduce latency.
  • The company's 'Digital In-Memory Computing' architecture performs calculations directly within memory, bypassing traditional data transfer bottlenecks.
  • d-Matrix offers a platform including AI inference accelerators, high-speed data movers, and optimized software, positioning itself as a comprehensive inference platform provider.

The competitive landscape for artificial intelligence semiconductors is shifting, with the industry's focus moving from AI model training to inference. While the past few years saw intense competition centered on acquiring Graphics Processing Units (GPUs) for faster model training, the widespread daily use of AI following ChatGPT's emergence has pivoted the industry towards optimizing the performance of already-built models.

The focus of the industry has shifted from model training to inference. The speed of response is now crucial for AI applications.

โ€” Sid Sheth, CEO of d-MatrixExplaining the industry trend towards AI inference.

In the realm of inference, the key metric is no longer just computational power but the speed at which users receive answers. As applications like coding agents and real-time voice assistants become more prevalent, reducing latency, the time delay in receiving a response, has become a critical performance indicator. This delay is surprisingly not dictated by raw processing power but by memory bottlenecks. Large Language Models (LLMs) contain billions of parameters, requiring constant data retrieval from memory for each token generated. This data movement creates a "memory wall," limiting overall processing speed regardless of how fast the computational units are.

Silicon Valley-based d-Matrix directly addresses this challenge with its specialized AI inference chips. Instead of the conventional setup of placing GPUs alongside High Bandwidth Memory (HBM), d-Matrix employs a 'Digital In-Memory Computing' (DIMC) architecture. This design integrates computational capabilities directly within the memory, eliminating the need to move data to separate processing units. By performing calculations where the data resides, d-Matrix significantly reduces the time and energy consumed by data transfer.

Instead of moving data to the processing unit, we perform calculations where the data is located within the memory.

โ€” Sid Sheth, CEO of d-MatrixDescribing d-Matrix's Digital In-Memory Computing architecture.

The company, co-founded by CEO Sid Sheth and CTO Sudeep Bhoja, offers a comprehensive platform that includes its AI inference accelerator 'Corsair,' the high-speed data mover 'JetStream,' and the inference-optimized software 'Aviator.' This integrated approach positions d-Matrix not merely as a chip manufacturer but as a full-spectrum inference platform provider. Their strategy also involves a collaborative relationship with NVIDIA, advocating for a heterogeneous computing approach where GPUs handle large-scale parallel processing, and d-Matrix's accelerators manage memory-intensive, latency-sensitive tasks. This vision entails a future data center infrastructure where various processors, including CPUs, GPUs, and specialized accelerators like d-Matrix's, work together efficiently.

We are not trying to eliminate GPUs. Our strategy is to divide tasks based on strengths: GPUs for large-scale parallel computation and our accelerators for memory-intensive, latency-sensitive inference tasks.

โ€” Sid Sheth, CEO of d-MatrixDiscussing the company's approach to working alongside GPUs.
DistantNews Editorial

Originally published by Dong-A Ilbo in Korean. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.