Research
I'm interested in developing learning algorithms that can effectively leverage diverse, multimodal data sources to enable robust robot learning. While robotics has been constrained by limited data compared to other AI domains, I believe the key lies not just in collecting more data, but in creating algorithms that can optimally extract knowledge from heterogeneous sources, including simulation, real-robot demonstrations, and internet videos.
Some questions I am interested in discussing and exploring:
- How can we design learning algorithms that effectively transfer knowledge across diverse domains and modalities?
- How can we enable robots to learn causal reasoning and abstract concepts from data to enable better generalization?
- What inductive biases and architectural choices enable efficient learning from sub-optimal demonstrations?
|
|
FloMo: 3D Scene Flow as a World Action Model Intermediate for Learning from Human Video
Jeremy A. Collins*, Namra Patel*, Ayush Agarwal*, Shitij Govil*, Animesh Garg
Under review at IEEE International Conference on Robotics and Automation (ICRA), 2027
paper /
website /
A world action model that renders 3D scene flow into the latent space of a pretrained video generation model and fine-tunes a single backbone to predict both motion and robot actions. Co-trained on egocentric human video and robot teleoperation data, FloMo shows that predicting motion instead of pixels yields better in-distribution success, far stronger out-of-distribution generalization, and higher sample efficiency.
|
|
COBALT: Crowdsourcing Robot Learning via Cloud-Based Teleoperation with Smartphones
Ayush Agarwal*, Ansh Gandhi*, Jeremy A. Collins, Omar Rayyan, Aryan Sarswat, Ranjani Koushik, Masoud Moghani, Ajay Mandlekar, Animesh Garg
IEEE International Conference on Robotics and Automation (ICRA), 2026
paper /
video /
code /
website /
A teleoperation platform that enables uninterrupted, concurrent data collection from multiple users worldwide using everyday devices like smartphones. This work not only significantly reduces teleoperation costs but also demonstrates the scalability of smartphone-based robot learning by successfully crowdsourcing 7500+ high-quality robot demonstrations from 50+ inexperienced teleoperators across nine countries.
|
|
Implementing Transformer Architectures for Audio Source Separation
Ayush Agarwal*, Brian Li*, Vinay Menon*, Neha Peddinti*, Yunbing Qian*, Devin Torres*
IEEE MIT Undergraduate Research Technology Conference (URTC), 2022
paper /
A novel transformer-based architecture that replaces BiLSTM blocks in the Open-Unmix model, achieving superior training efficiency and reduced inter-source interference on MUSDB18-HQ.
Figure adapted from Manilow, E., Seetharman, P., & Salamon, J. “Open Source Tools & Data for Music Source Separation” (2020)
|
|