The autonomous driving field is divided over how vehicles should sense their environment. Tesla is advancing a vision-only stack for consumer Full Self-Driving and its commercial Robotaxi platform, while Alphabet’s Waymo supports a multimodal package that merges cameras, radar, and LiDAR.
At Y Combinator’s Startup School 2026, Waymo Co-CEO Dmitri Dolgov shared lessons from taking the company’s driverless robotaxis from a demo to a commercial product, explaining that Waymo fuses data from multiple physical sensing modalities to tackle physical AI in real traffic.
Fusing Cameras, Radar, and LiDAR for Redundancy
Waymo’s system is built around a multimodal sensor suite. Instead of relying solely on optical cameras, its vehicles combine high-resolution cameras, active LiDAR, and radar into a unified perception model.

Dolgov noted that cameras provide rich color and high resolution but can struggle in heavy glare, direct sunlight, complete darkness, or harsh weather. Active LiDAR supplies direct 3D structural measurements regardless of ambient lighting, and radar penetrates dense fog, rain, and snow while using Doppler to measure object velocity.
According to Waymo, a single-modality system leads to a safety curve that plateaus too early when targeting superhuman driving. Using multiple distinct sensor types creates overlapping coverage that protects against edge-case failures, such as debris partially blocking a camera lens.
Waymo vs. Tesla: Two Divergent Philosophies
Waymo’s multimodal approach contrasts with Tesla’s camera-only strategy. Waymo contends that a camera-only system may ramp more easily at first but ultimately hits a performance ceiling, whereas Tesla remains committed to one sensing modality and maintains it can solve edge cases by training on enough real-world data.

Tesla notoriously walked away from radar and LiDAR to go all-in on vision, asserting that because human drivers navigate using two eyes and a brain, self-driving vehicles should operate identically using cameras and neural networks. Instead of active sensors, Tesla models the physical world with vision alone, using cameras to reconstruct 3D space and processing billions of video frames to address long-tail edge cases.

This architectural split reflects different business models. Tesla’s vision-only system keeps hardware costs low, enabling driver-assist features to ship on millions of customer vehicles worldwide. By contrast, Waymo’s multimodal suite adds expensive hardware, but it aligns more readily with current regulatory frameworks thanks to its redundancies and demonstrated safety metrics.
Although humans can drive safely using visual perception alone, adding active radar and 3D LiDAR can introduce extra safety margins that may let autonomous systems react faster than humans and potentially drive faster than people can safely manage. As both companies expand unsupervised driverless operations, real-world intervention metrics will indicate which technical path reaches broad fleet deployment first. You can watch Dolgov’s full talk below:
![Waymo Explains Their Multimodal Approach to Autonomy [VIDEO]](http://teslahubs.com/cdn/shop/articles/radar.jpg?v=1785960136&width=1200)














Teilen:
Tesla Model Y L Accessories Expected to Arrive in North America
EVgo to Start Installing 500 kW V4 Tesla Superchargers