top of page

Erik Chen

Artilux

* All members of the platform can watch the entire presentation.

 

Please register to become a member.

Erik Chen | Artilux: How does a wafer-integrated ASIC and modulated photo detector sandwich eliminate optical fibers in matrix multiplication?

47:13 - 48:51

Other snippets from this talk

Summary of the clip:

How does a wafer-integrated ASIC and modulated photo detector sandwich eliminate optical fibers in matrix multiplication?

The architecture is described as a sandwich. At the lower part, an ASIC first fetches data from memory, just like existing systems. Through wafer-and-wafer integration, a modulated photo detector is combined with that ASIC. Above it sits a huge array of light emitters, for example 1,000 by 1,000, a million of them. The ASIC sends an input signal through a DAC so the emitter can shine light. The key point: no fiber is needed; the parts can sync directly together.

What is wanted from the photons is the low-noise part. The photons go directly into the modulated photo detector, and how many lights are converted to an electric signal is determined by the modulation of the photo detector itself. That modulation is the weight. In this way, the multiplication is completed. The signal path is therefore emitter-to-detector, not through fiber.

For each pixel, the design includes in-pixel memory or capacitance. Each time the charge is accumulated, that is really just the communication. The description says it is straightforward. No fiber is present, everything is not even there, because the parts can sync directly together. The large emitter array sits above the integrated ASIC and modulated photo detector, completing the sandwich for matrix multiplication.

In this short video, you can learn:
* Artilux replaces fiber by wafer-and-wafer integrating a modulated photo detector with an ASIC beneath a huge light-emitter array.
* The photo detector's own modulation sets how many photons become electric signal, so it acts as the weight for multiplication.
* Each pixel uses in-pixel memory or capacitance to accumulate charge, which the clip describes as communication.

📋 **Clip Abstract** Artilux describes a sandwich architecture where an ASIC fetches memory data, drives a DAC, and shines a 1,000-by-1,000 emitter array onto a wafer-integrated modulated photo detector. Because the detector's modulation defines the weight, photons convert directly to electric signal, completing multiplication without fiber while in-pixel memory or capacitance accumulates charge.

About the speaker:
* Speaker: Erik Chen
* Company: Artilux
* Event: Eindhoven 2026
* Location: High Tech Campus, Eindhoven

#WaferAndWaferIntegration, #ModulatedPhotoDetector, #InPixelMemory, #DACDrivenEmitterArray, #OpticalComputing, #OptoelectronicComputing

This is a highlight of the presentation:

The Quest for Ultra-Low-Power Implementation of Transformer-Based LLMs: Opportunities for Optoelectronics

MicroLED Connect 2026

AR/VR Connect 2026

16-17 September 2026

High Tech Campus, Eindhoven

Organised By:

Khasha and Ron

Khasha and Ron

More Highlights from the same talk.

05:41 - 07:17

Where does the energy actually go when running large language models?

Where does the energy actually go when running large language models?

Erik Chen breaks LLM energy into three parts. The first is the compute itself, which splits into two: a linear part of linear algebra and matrix multiplication, including the big thousands-by-thousands matrix multiplied by thousands and thousands, and MAC — after a multiplication you need to do an accumulation. There are also vector-based operations. All of this is called linear compute, and for the compute part it is the most energy-dominant one.

The second part Chen calls nonlinear and peripheral compute: nonlinear functions for activation, softmax to determine the probability distribution, and normalization. These are smaller, but he states they are a non-negligible share of energy compute for this part. So the energy budget is not only matrix math: softmax probability distributions and normalization layers also draw power inside the model, alongside the linear algebra that dominates.

Between compute sits memory access. Chen describes a hierarchical memory from short distance at the bottom to longer distance at the top, from register to SRAM and all the way to DRAM storage. HBN sits near the same hierarchy as DRAM: DRAM stacked together with a really improved interface for high-bandwidth delivery. SRAM is closer to the chip and costs less energy to access, but its size is bigger — a memory hierarchy trade-off in designing a really useful IC.

In this short video, you can learn:
* Linear compute — matrix multiplication, MAC and vector-based operations — is the most energy-dominant part of LLM compute.
* Nonlinear and peripheral compute, covering activation, softmax probability distribution and normalization, is smaller but a non-negligible share of energy.
* Memory is hierarchical from register to SRAM to DRAM, with HBN near the DRAM hierarchy and SRAM cheaper to access but larger.

📋 **Clip Abstract** Erik Chen splits large language model energy into compute — a linear part of matrix multiplication with MAC and vector operations, plus smaller nonlinear and peripheral work such as activation, softmax and normalization — and memory access. The memory hierarchy runs from register to SRAM to DRAM storage, with HBN sitting near DRAM as stacked DRAM with an improved interface for high-bandwidth delivery, while SRAM is closer to the chip and costs less energy but is bigger.

About the speaker:
* Speaker: Erik Chen
* Company: Artilux
* Event: Eindhoven 2026
* Location: High Tech Campus, Eindhoven

#MemoryHierarchyTradeOff, #HighBandwidthDRAM, #SRAMvsDRAM, #SoftmaxNormalization, #LLMInference, #ComputeEnergy

19:12 - 20:32

Why is electrical copper interconnect failing at high data rates, and how does optics solve it?

Why is electrical copper interconnect failing at high data rates, and how does optics solve it?

At high data rates electrical transfer causes skin effect, so not the whole copper wire is used, only the skin wire. That causes extraordinary loss when communicating beyond large distances and at higher data rate. Equalization or coding concepts are needed, and these consume huge energy. Compared with optical, this is not the same order—one order, two orders of magnitude. This is the dominant energy consumption factor, making IO a significant system-level power consumption from a copper perspective.

Optical communication interconnect has been happening not just from long distance but approaching in-rack and even chip-to-chip. The major opportunity is using optical to move information rather than pushing large amounts of data electrically. And it is already happening. The speaker also asks whether the name from Micro LED Connect means connect as an application or connect for us to connect—question unresolved.

The transcript frames the interconnect problem as the easiest current issue. At high data rates, skin effect limits copper to the skin wire, causing extraordinary loss; equalization and coding add dominant energy consumption. Optics is presented as not the same order as electrical, with one or two orders of magnitude difference, and already moving from long-distance to in-rack and chip-to-chip. That makes optical IO the system-level opportunity for moving information.

In this short video, you can learn:
* Skin effect at high data rates means only the skin of a copper wire carries the signal, causing extraordinary loss for longer distances and higher data rates.
* Equalization and coding needed for electrical links consume huge energy, making IO a dominant system-level power consumption factor compared with optical.
* Optical communication interconnect is already moving from long distance toward in-rack and chip-to-chip, using optics to move information instead of pushing large amounts of data electrically.

📋 **Clip Abstract** Erik Chen explains that at high data rates copper suffers skin effect, so only the skin of the wire is used, causing extraordinary loss and requiring equalization or coding that consumes huge energy—one to two orders of magnitude compared with optical. Optical communication interconnect is therefore already advancing from long distance to in-rack and chip-to-chip, using optics to move information rather than pushing large amounts of data electrically.

About the speaker:
* Speaker: Erik Chen
* Company: Artilux
* Event: Eindhoven 2026
* Location: High Tech Campus, Eindhoven

#SkinEffect, #CopperInterconnect, #EqualizationCoding, #ChipToChipOptics, #OpticalCommunication, #Optoelectronics

13:59 - 15:49

Why is LLM prefill compute bound while decode is communication bound?

Why is LLM prefill compute bound while decode is communication bound?

Prefill is the phase where the model works out what the question is. Erik Chen describes it as the heavy matrix multiplication stage: thousands by thousands of entries, large matrix operations run over many input tokens, and the relationships between words being built out. It is energy intensive, and its cost scales with the square of the context length. Prefill is compute bound, so the ceiling is how many TOPS the hardware can deliver.

Decode is the answering phase, done one token at a time by autoregressive operation. The model repeatedly accesses the KV cache to deliver the next token to the user. Because it is one token at a time, decode scales linearly rather than quadratically. That makes decode communication bound: the question is already known, so the work becomes a memory read to find the weights and a write back. Chen calls this the contrasting bottleneck to prefill's compute bounding.

The two phases therefore demand different hardware even within inference. Prefill needs matrix throughput, measured in TOPS, over many input tokens; decode needs memory traffic for KV cache reads and weight read-back for a single token. Chen notes the hardware is "also already different" at the inference stage, and adds that hardware for inference and training is different again. One LLM, two workflows, two scaling laws.

In this short video, you can learn:
* Prefill performs thousands-by-thousands matrix multiplications over many input tokens and scales with the square of the context length, making it compute bound on delivered TOPS.
* Decode is autoregressive, produces one token at a time, scales linearly, and is communication bound because it repeatedly reads the KV cache and writes weights back.
* Inference alone already needs two different hardware profiles, and Chen stresses that inference hardware also differs from training hardware.

📋 **Clip Abstract** Erik Chen splits LLM inference into prefill, a compute-bound phase of thousands-by-thousands matrix multiplications over many input tokens that scales with the square of context, and decode, a communication-bound autoregressive phase producing one token at a time that scales linearly while repeatedly accessing the KV cache. Those two workflows already impose different hardware demands before training hardware is even considered.

About the speaker:
* Speaker: Erik Chen
* Company: Artilux
* Event: Eindhoven 2026
* Location: High Tech Campus, Eindhoven

#KVCache, #AutoregressiveDecode, #ComputeBoundPrefill, #MatrixMultiplication, #LLMInference, #AIAccelerators

More Snippets
CONTACT US

KGH Concepts GmbH

Mergenthalerallee 73-75, 65760, Eschborn

+49 17661704139

admin@techblick.com

TechBlick is owned and operated by KGH Concepts GmbH

Registration number HRB 121362

VAT number: DE 337022439

  • LinkedIn
  • YouTube

Sign up for our newsletter to receive updates on our latest speakers and events AND to receive analyst-written summaries of the key talks and happenings in our events.

Thanks for submitting!

© 2026 by KGH Concepts GmbH

bottom of page