Publications

A searchable list of some of my publications is below. You can also access my publications from the following sites.

My ORCID is ORCID iD icon

https://orcid.org/0000-0002-6236-2969

Publications:

Show all

Erik Wijmans, Irfan Essa, Dhruv Batra

How to Train PointGoal Navigation Agents on a (Sample and Compute) Budget Proceedings Article

In: International Conference on Autonomous Agents and Multi-Agent Systems, 2022.

Abstract | Links | BibTeX | Tags: computer vision, embodied agents, navigation

Vincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee, Irfan Essa, Dhruv Batra

Semantic MapNet: Building Allocentric SemanticMaps and Representations from Egocentric Views Proceedings Article

In: Proceedings of American Association of Artificial Intelligence Conference (AAAI), AAAI, 2021.

Abstract | Links | BibTeX | Tags: AAAI, AI, embodied agents, first-person vision

@inproceedings{2021-Cartillier-SMBASRFEV,

title = {Semantic MapNet: Building Allocentric SemanticMaps and Representations from Egocentric Views},

author = {Vincent Cartillier and Zhile Ren and Neha Jain and Stefan Lee and Irfan Essa and Dhruv Batra},

url = {https://arxiv.org/abs/2010.01191

https://vincentcartillier.github.io/smnet.html

https://ojs.aaai.org/index.php/AAAI/article/view/16180/15987},

doi = {10.48550/arXiv.2010.01191},

year  = {2021},

date = {2021-02-01},

urldate = {2021-02-01},

booktitle = {Proceedings of American Association of Artificial Intelligence Conference (AAAI)},

publisher = {AAAI},

abstract = {We study the task of semantic mapping -- specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map (`what is where?') from egocentric observations of an RGB-D camera with known pose (via localization sensors). Importantly, our goal is to build neural episodic memories and spatio-semantic representations of 3D spaces that enable the agent to easily learn subsequent tasks in the same space -- navigating to objects seen during the tour (`Find chair') or answering questions about the space (`How many chairs did you see in the house?'). 

Towards this goal, we present Semantic MapNet (SMNet), which consists of: (1) an Egocentric 

 

Visual Encoder that encodes each egocentric RGB-D frame, (2) a Feature Projector that projects egocentric features to appropriate locations on a floor-plan, (3) a Spatial Memory Tensor of size floor-plan length × width × feature-dims that learns to accumulate projected egocentric features, and (4) a Map Decoder that uses the memory tensor to produce semantic top-down maps. SMNet combines the strengths of (known) projective camera geometry and neural representation learning. On the task of semantic mapping in the Matterport3D dataset, SMNet significantly outperforms competitive baselines by 4.01-16.81% (absolute) on mean-IoU and 3.81-19.69% (absolute) on Boundary-F1 metrics. Moreover, we show how to use the spatio-semantic allocentric representations build by SMNet for the task of ObjectNav and Embodied Question Answering.},

keywords = {AAAI, AI, embodied agents, first-person vision},

pubstate = {published},

tppubtype = {inproceedings}

}

Erik Wijmans, Abhishek Kadian, Ari Morcos, Stefan Lee, Irfan Essa, Devi Parikh, Manolis Savva, Dhruv Batra

Decentralized Distributed PPO: Solving PointGoal Navigation Proceedings Article

In: Proceedings of International Conference on Learning Representations (ICLR), 2020.

Abstract | Links | BibTeX | Tags: embodied agents, ICLR, navigation, systems for ML

@inproceedings{2020-Wijmans-DDSPN,

title = {Decentralized Distributed PPO: Solving PointGoal Navigation},

author = {Erik Wijmans and Abhishek Kadian and Ari Morcos and Stefan Lee and Irfan Essa and Devi Parikh and Manolis Savva and Dhruv Batra},

url = {https://arxiv.org/abs/1911.00357

https://paperswithcode.com/paper/decentralized-distributed-ppo-solving},

year  = {2020},

date = {2020-04-01},

urldate = {2020-04-01},

booktitle = {Proceedings of International Conference on Learning Representations (ICLR)},

abstract = {We present Decentralized Distributed Proximal Policy Optimization (DD-PPO), a method for distributed reinforcement learning in resource-intensive simulated environments. DD-PPO is distributed (uses multiple machines), decentralized (lacks a centralized server), and synchronous (no computation is ever stale), making it conceptually simple and easy to implement. In our experiments on training virtual robots to navigate in Habitat-Sim, DD-PPO exhibits near-linear scaling -- achieving a speedup of 107x on 128 GPUs over a serial implementation. We leverage this scaling to train an agent for 2.5 Billion steps of experience (the equivalent of 80 years of human experience) -- over 6 months of GPU-time training in under 3 days of wall-clock time with 64 GPUs.

This massive-scale training not only sets the state of art on Habitat Autonomous Navigation Challenge 2019, but essentially solves the task --near-perfect autonomous navigation in an unseen environment without access to a map, directly from an RGB-D camera and a GPS+Compass sensor. Fortuitously, error vs computation exhibits a power-law-like distribution; thus, 90% of peak performance is obtained relatively early (at 100 million steps) and relatively cheaply (under 1 day with 8 GPUs). Finally, we show that the scene understanding and navigation policies learned can be transferred to other navigation tasks -- the analog of ImageNet pre-training + task-specific fine-tuning for embodied AI. Our model outperforms ImageNet pre-trained CNNs on these transfer tasks and can serve as a universal resource (all models and code are publicly available).},

keywords = {embodied agents, ICLR, navigation, systems for ML},

pubstate = {published},

tppubtype = {inproceedings}

}

Erik Wijmans, Julian Straub, Dhruv Batra, Irfan Essa, Judy Hoffman, Ari Morcos

Analyzing Visual Representations in Embodied Navigation Tasks Technical Report

no. arXiv:2003.05993, 2020.

Abstract | Links | BibTeX | Tags: arXiv, embodied agents, navigation

Huda Alamri, Vincent Cartillier, Abhishek Das, Jue Wang, Anoop Cherian, Irfan Essa, Dhruv Batra, Tim K. Marks, Chiori Hori, Peter Anderson, Stefan Lee, Devi Parikh

Audio Visual Scene-Aware Dialog Proceedings Article

In: IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2019.

Abstract | Links | BibTeX | Tags: computational video, computer vision, CVPR, embodied agents, vision & language

Huda Alamri, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das, Jue Wang, Irfan Essa, Dhruv Batra, Devi Parikh, Anoop Cherian, Tim K Marks, Chiori Hori

Audio Visual Scene-Aware Dialog (AVSD) Challenge at DSTC7 Technical Report

no. arXiv:1806.00525, 2018.

Abstract | Links | BibTeX | Tags: arXiv, embodied agents, multimedia, vision & language