Paolo Costa | Microsoft: Can spontaneous-emission microLEDs beat DFB lasers on picojoules per bit despite lower modulation speed?
26:50 - 28:16
Other snippets from this talk
Summary of the clip:
Can spontaneous-emission microLEDs beat DFB lasers on picojoules per bit despite lower modulation speed?
For a hyperscaler the decision is end-to-end: picojoule per bit, bandwidth density, reliability and latency matter more than the light source itself. Any technology that achieves those properties is acceptable. The DFB laser's problem is its lasing threshold, which applies to every laser and imposes a floor on how low power can go. Spontaneous emission is therefore one of the microLED's fundamental advantages.
Temperature sensitivity is the second challenge. Today a DFB-based solution often puts the laser outside the cooling domain, because the GPU itself dissipates three to four kilowatts. Once the laser sits outside, the light must be brought back in. And because DFB is edge emitting, you cannot take advantage of the full-area bandwidth that a surface emitter would offer.
The other side is maturity. DFB is a very mature, proven approach that can be modulated at high speed. So the comparison is not clear cut. That is precisely where microLEDs have a unique opportunity, if the ecosystem can actually arrive. What remains open is the trade between a proven high-speed laser and spontaneous emission's lower power floor.
In this short video, you can learn:
* A hyperscaler judges an optical link on picojoules per bit, bandwidth density, reliability and latency, not on the light source alone.
* Every laser carries a lasing threshold that sets a floor on power, while microLEDs emit spontaneously.
* DFB lasers are mature and fast, but edge emitting and temperature sensitive, so a three-to-four-kilowatt GPU package pushes them outside the cooling domain.
📋 **Clip Abstract** Asked how microLEDs compare with DFB lasers for data communication, the answer is end-to-end metrics: picojoule per bit, bandwidth density, reliability and latency. DFB wins on maturity and modulation speed, but its lasing threshold, temperature sensitivity and edge emission leave room for spontaneous-emission microLEDs.
About the speaker:
* Speaker: Paolo Costa
* Company: Microsoft
* Event: Eindhoven 2026
* Location: High Tech Campus, Eindhoven
#MicroLEDOpticalIO, #DFBLaser, #LasingThreshold, #SpontaneousEmission, #OpticalInterconnect, #AIInfrastructure
This is a highlight of the presentation:
From Displays to Data Centers: The Promise and Reality of MicroLED Optical I/O for Next-Generation AI Infrastructure
More Highlights from the same talk.
03:20 - 05:23
Why does each step up in data-center interconnect reach cost a 10x penalty in bandwidth density per picojoule?
Why does each step up in data-center interconnect reach cost a 10x penalty in bandwidth density per picojoule?
The interconnect community reads a simple graph: reach on the x-axis, from a few millimeters to 100 meters and beyond, and on the y-axis bandwidth density divided by picojoule per bit, where higher is better. Three technologies serve today's data centers. The speaker oversimplifies and warns not to fixate on exact values: the plot illustrates a trend, not precise benchmarks.
From the top left corner, parallel copper traces are wide and slow. They sit on a silicon interposer or package substrate and connect chip to chip or chip to memory. The power is extremely low, bandwidth density extremely high, but reach is only a few millimeters. Next, copper cables operate within the rack at very high speed: 100 gig, 200 gig, possibly 400 gig next generation, yet reach is about one meter.
Across racks, standard pluggable optics run at 200 gigabit per lane, with next generation probably 400 gigabit per lane, reaching hundreds of meters and beyond. Moving between these three technologies, the figure of merit drops by roughly one or two orders of magnitude. Public NVIDIA B200 data shows roughly a 10X gap every step. The speaker uses NVIDIA because it is public data, illustrating the trend.
In this short video, you can learn:
* Bandwidth density divided by picojoule per bit is the interconnect figure of merit, and higher is better.
* Parallel copper traces deliver extremely low power and high bandwidth density, but reach only a few millimeters on silicon interposers and package substrates.
* Copper cables reach about one meter at 100–400 gig, while standard pluggable optics reach hundreds of meters at 200 gigabit per lane, with roughly one to two orders of magnitude, or 10X, gap between each step.
📋 **Clip Abstract** Data-center interconnect spans a reach-versus-efficiency graph where bandwidth density divided by picojoule per bit falls by one to two orders of magnitude from parallel copper traces to copper cables to pluggable optics. Public NVIDIA B200 data shows roughly a 10X gap at every step, from millimeter-scale interposer links to 200-gigabit-per-lane optics reaching hundreds of meters.
About the speaker:
* Speaker: Paolo Costa
* Company: Microsoft
* Event: Eindhoven 2026
* Location: High Tech Campus, Eindhoven
#ParallelCopperTraces, #PluggableOptics, #SiliconInterposer, #NvidiaB200, #OpticalInterconnect, #DataCenterNetworking
13:43 - 15:23
Could going wide and slow remove serial conversion while simplifying 200-gigabit drivers down to 2-gigabit electronics?
Could going wide and slow remove serial conversion while simplifying 200-gigabit drivers down to 2-gigabit electronics?
Today’s pluggable optics stack includes SerDes power, analog front end, light sources such as lasers, VCSELs and MicroLED, plus microcontroller units; die-to-die is not there. The first industry move is co-packaged optics, bringing optics very close to the GPU. That uses a short die-to-die link at low power and can cut some SerDes, but SerDes remains a big fraction of the power. This is already happening, not just proposed.
The next step, advocated with community members, is to go wide and slow: remove serial conversion, avoid the 200-gigabit and 400-gigabit race, and go parallel. This simplifies electronics because a 2-gigabit driver, TIAs and related circuits are easier to design than 200-gigabit versions. Die-to-die now adds cost that pluggable optics did not have, but overall we still gain. Risk rises and this is not a done deal, though it might be possible.
Eventually, projecting a few years ahead with lots of risk, 3D stacking becomes affordable. That would remove the limitation on shoreline density and unlock the full aerial bandwidth. The path is staged: co-packaged optics first, then wide-and-slow parallel links, then 3D stacking. Each stage carries unknowns, but the direction aims to cut serial conversion and simplify the optical and electronic interface. The prize is a link architecture that scales without the serial-speed race.
In this short video, you can learn:
* Co-packaged optics brings optics close to the GPU, enabling short low-power die-to-die links but still leaving SerDes as a major power fraction.
* Wide-and-slow parallel links remove serial conversion and simplify 2-gigabit drivers and TIAs compared with 200-gigabit designs, though die-to-die cost appears.
* Projected 3D stacking could remove shoreline density limits and unlock full aerial bandwidth, but the path carries high risk and unknowns.
📋 **Clip Abstract** The clip compares today’s pluggable optics stack with co-packaged optics, which shortens die-to-die links and cuts some SerDes but leaves SerDes as a big power fraction. It argues for going wide and slow—removing serial conversion to simplify 2-gigabit electronics, accepting die-to-die cost, and projecting 3D stacking to remove shoreline density limits and unlock aerial bandwidth, while stressing high risk and unknowns.
About the speaker:
* Speaker: Paolo Costa
* Company: Microsoft
* Event: Eindhoven 2026
* Location: High Tech Campus, Eindhoven
#CoPackagedOptics, #WideAndSlow, #DieToDie, #ShorelineDensity, #OpticalInterconnect, #AIInfrastructure
07:04 - 08:27
Why can't HBM keep stacking forever, and what limits capacity beyond TSV bandwidth?
Why can't HBM keep stacking forever, and what limits capacity beyond TSV bandwidth?
Because the interconnect reaches only a few millimeters at high bandwidth and low power, memory must sit next to the GPU. Like Manhattan, the only way is up. Today's HBM is standard DRAM stacked in eight tiers. The next generation targets 12 tiers, possibly 16, and maybe around 20, but each step adds complexity and cost.
Adding tiers forces thinner layers to maintain planarity with the GPU, and cooling becomes harder: a cold plate sits at the top while HBM dies consume power at the bottom, so heat must pass through every layer. More layers also need more TSVs, so byte density falls and cost rises. This primarily increases capacity.
To increase bandwidth, however, you are limited by the number of TSVs. Recently announced work, and several startups, are looking at stacking HBM on top of a GPU. That shifts the geometry again, but the transcript's constraints remain: planarity, cooling through stacked die, TSV count, byte density, and cost all bound how far HBM can scale upward.
In this short video, you can learn:
* Today's HBM uses eight DRAM tiers, with next generations at 12, possibly 16, and maybe around 20, but complexity and cost rise with each layer.
* Adding layers demands thinner die for GPU planarity and makes cooling harder because heat from bottom HBM die must pass through a top cold plate.
* More tiers require more TSVs, lowering byte density and raising cost, while bandwidth remains limited by TSV count; startups are exploring HBM stacked on GPU.
📋 **Clip Abstract** The clip explains why memory must sit millimeters from the GPU and how HBM's upward stacking, from eight tiers toward 12, 16, or around 20, hits planarity, cooling, TSV, byte-density, and cost limits. It also notes that bandwidth is bounded by TSV count, prompting recently announced efforts to stack HBM on top of a GPU.
About the speaker:
* Speaker: Paolo Costa
* Company: Microsoft
* Event: Eindhoven 2026
* Location: High Tech Campus, Eindhoven
#HBMStacking, #ThroughSiliconVias, #GPUPlanarity, #ColdPlateCooling, #AdvancedPackaging, #MemoryBandwidth




