I am a research scientist in Adobe Nextcam team led by Marc Levoy. My work spans computational imaging and computer vision, with a lot of data collection and model & hardware joint tweaking.
I obtained my Ph.D. in EECS from UC Berkeley, where I was advised by Laura Waller and supported, in part, by Siebel Scholarship. I was also affiliated with the Berkeley Artificial Intelligence Research (BAIR). I was a research intern in Meta Reality Labs Research in 2022 and in Google Research in 2021. I received M.S. in CS from UCLA in 2019, where I was advised by Kyung Hyun Sung and Fabien Scalzo, and I also did college in CS and Math at UCLA.
Research
My research focuses on motion modeling and dynamic imaging. I jointly develop the optical system hardware and the computational methods to enhance our capabilities for observing and analyzing fast and complex dynamics. I work on various image modalities, including digital photography, optical microscopy, medical imaging, and most recently cryo-ET.
Optical sensing methods for dynamic imaging. I develop optical sensing hardware by enhancing signal encoding through dynamic sensors and novel contrast mechanisms, specifically designed for imaging moving samples [1, 2]. My work optimizes imaging hardware designs to improve signal quality and reduce acquisition time [3, 4]. Additionally, I have pioneered theoretical noise analysis for dynamic vision sensors, leveraging their unique properties for new applications [5].
Parsing dynamics and inverse problem optimization with machine learning. I work on inverse problem optimization algorithms that parse spatiotemporal correlations from unlabeled data, overcoming the limitations of frame-by-frame processing [6]. The neural space-time model significantly enhances temporal resolution, image quality [7], and multi-modal alignment [8] in image reconstruction . Furthermore, I am developing a large vision model for cryo-ET data processing tasks.
Recent work
Image generation
HDR-preserved instruction-based editing
Project Indigo’s AI Playground brings instruction-based generative editing directly into a camera workflow. Because leading generators return SDR images, our HDR-guided post-processing model uses the original capture to restore highlight brightness and dynamic range.
Original
Edited (Hover to see HDR)
M. Levoy, “An AI playground for the Project Indigo camera app,” Adobe Research, 2026.
[blog]
PhD work
Inverse problems · 3D+time representation
Parsing dynamics for science
Sequential computational imaging is vulnerable to motion artifacts, but it also records rich temporal structure. I developed neural space-time models that jointly recover a scene and its motion, improving image quality while resolving dynamics that frame-by-frame reconstructions miss.
Q. L. N. Nguyen, R. Cao, and L. Waller, “Multi-Modal Deformable Image Registration Using Untrained Neural Networks,” IEEE ISBI, pp. 1–5, 2025.
[paper]
R. Cao, N. S. Divekar, J. K. Nuñez, S. Upadhyayula, and L. Waller, “Neural space–time model for dynamic multi-shot imaging,” Nature Methods, 2024.
[paper][code][dataset][docs][news]
T. Chien, R. Cao, F. L. Liu, L. A. Kabuli, and L. Waller, “Space-time reconstruction for lensless imaging using implicit neural representations,” Optics Express, 2024.
[paper]
R. Cao, F. L. Liu, L.-H. Yeh, and L. Waller, “Dynamic Structured Illumination Microscopy with a Neural Space-time Model,” IEEE ICCP, 2022.
[paper][code][video]
Event cameras · novel signal
Understanding noise & denoising
Event cameras usually discard background noise and respond only to brightness changes. Noise2Image instead models photon-noise events as a signal, recovering static scene intensity that would otherwise be invisible to the sensor.
R. Cao*, D. Galor*, A. Kohli, J. L. Yates, and L. Waller, “Noise2Image: Noise-Enabled Static Scene Recovery for Event Cameras,” Optica, 2025.
[paper][code][dataset]* equal contribution
Computational optics · microscopy
Co-designing optics hardware and algorithms
Imaging hardware and reconstruction algorithms work best when they are designed together. These projects use optimized illumination, motion-based encoding, and self-calibration to improve resolution and simplify microscope acquisition.
S. Li, R. Cao, L. Waller, K. Monakhova, and S. Fridovich-Keil, “Sparse Light Field Sampling Improves Casual 3D and 4D Reconstruction,” ECCV Workshops, 2026.
[paper][project][dataset]
M. Zach, K.-C. Shen, R. Cao, M. Unser, L. Waller, and J. Dong, “Perturbative Fourier ptychographic microscopy for fast quantitative phase imaging,” Optics Express, vol. 33, no. 18, pp. 38984–38996, 2025.
[paper]
R. Cao, G. Meng, and L. Waller, “Sample motion for structured illumination fluorescence microscopy,” Optics Letters, 2025.
[paper][code]
R. Cao, M. Kellman, D. Ren, R. Eckert, and L. Waller, “Self-calibrated 3D differential phase contrast microscopy with optimized illumination,” Biomedical Optics Express, 2022.
[paper][code]
Earlier work
Magnetic resonance imaging
Medical AI & imaging
I developed deep learning methods for detecting, segmenting, and grading prostate cancer in multi-parametric MRI. The work paired whole-mount histopathology with large clinical imaging cohorts and compared model performance directly with radiologists.
R. Cao et al., “Performance of Deep Learning and Genitourinary Radiologists in Detection of Prostate Cancer Using 3-T Multiparametric Magnetic Resonance Imaging,” Journal of Magnetic Resonance Imaging, 2021.
[paper]
R. Cao et al., “Joint prostate cancer detection and Gleason score prediction in mp-MRI via FocalNet,” IEEE Transactions on Medical Imaging, 2019.
[paper][news]
R. Cao et al., “Prostate Cancer Detection and Segmentation in Multi-parametric MRI via CNN and Conditional Random Field,” IEEE ISBI, 2019.
[paper]
R. Cao et al., “Prostate cancer inference via weakly-supervised learning using a large collection of negative MRI,” ICCV Workshops, 2019.
[paper]
Mechanistic interpretability
Explaining deep convolutional networks
These projects study how convolutional networks encode object parts inside otherwise opaque feature maps. Explanatory graphs and active question-answering organize hidden activations into semantic structures that people can inspect and refine with very few annotations.
Q. Zhang, J. Ren, G. Huang, R. Cao, Y. N. Wu, and S.-C. Zhu, “Mining Interpretable AOG Representations From Convolutional Networks via Active Question Answering,” IEEE TPAMI, 2021.
[paper]
Q. Zhang, X. Wang, R. Cao, Y. N. Wu, F. Shi, and S.-C. Zhu, “Extraction of an Explanatory Graph to Interpret a CNN,” IEEE TPAMI, 2021.
[paper]
Q. Zhang, R. Cao, F. Shi, Y. N. Wu, and S.-C. Zhu, “Interpreting CNN Knowledge via an Explanatory Graph,” AAAI, 2018.
[paper][code]
Q. Zhang, R. Cao, Y. N. Wu, and S.-C. Zhu, “Mining Object Parts from CNNs via Active Question-Answering,” CVPR, 2017.
[paper]
Q. Zhang, R. Cao, Y. N. Wu, and S.-C. Zhu, “Growing Interpretable Part Graphs on ConvNets via Multi-Shot Learning,” AAAI, 2017.
[paper][code]