Publications

A searchable list of some of my publications is below. You can also access my publications from the following sites.

My ORCID is ORCID iD icon

Publications:

Lijun Yu, José Lezama, Nitesh B. Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng, Vighnesh Birodkar, Agrim Gupta, Xiuye Gu, Alexander G. Hauptmann, Boqing Gong, Ming-Hsuan Yang, Irfan Essa, David A. Ross, Lu Jiang

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Proceedings Article

In: Proceedings of International Conference on Learning Representations (ICLR) , 2024.

Abstract | Links | BibTeX | Tags: AI, arXiv, computer vision, generative AI, google, ICLR

Karan Samel, Jun Ma, Zhengyang Wang, Tong Zhao, Irfan Essa

Knowledge Relevance BERT: Integrating Noisy Knowledge into Language Representation. Proceedings Article

In: AAAI workshop on Knowledge Augmented Methods for NLP (KnowledgeNLP-AAAI 2023), 2023.

Abstract | Links | BibTeX | Tags: AI, knowledge representation, NLP

Vincent Cartillier, Zhile Ren, Neha Jain, Stefan Lee, Irfan Essa, Dhruv Batra

Semantic MapNet: Building Allocentric SemanticMaps and Representations from Egocentric Views Proceedings Article

In: Proceedings of American Association of Artificial Intelligence Conference (AAAI), AAAI, 2021.

Abstract | Links | BibTeX | Tags: AAAI, AI, embodied agents, first-person vision

@inproceedings{2021-Cartillier-SMBASRFEV,

title = {Semantic MapNet: Building Allocentric SemanticMaps and Representations from Egocentric Views},

author = {Vincent Cartillier and Zhile Ren and Neha Jain and Stefan Lee and Irfan Essa and Dhruv Batra},

url = {https://arxiv.org/abs/2010.01191

https://vincentcartillier.github.io/smnet.html

https://ojs.aaai.org/index.php/AAAI/article/view/16180/15987},

doi = {10.48550/arXiv.2010.01191},

year  = {2021},

date = {2021-02-01},

urldate = {2021-02-01},

booktitle = {Proceedings of American Association of Artificial Intelligence Conference (AAAI)},

publisher = {AAAI},

abstract = {We study the task of semantic mapping -- specifically, an embodied agent (a robot or an egocentric AI assistant) is given a tour of a new environment and asked to build an allocentric top-down semantic map (`what is where?') from egocentric observations of an RGB-D camera with known pose (via localization sensors). Importantly, our goal is to build neural episodic memories and spatio-semantic representations of 3D spaces that enable the agent to easily learn subsequent tasks in the same space -- navigating to objects seen during the tour (`Find chair') or answering questions about the space (`How many chairs did you see in the house?'). 

Towards this goal, we present Semantic MapNet (SMNet), which consists of: (1) an Egocentric 

 

Visual Encoder that encodes each egocentric RGB-D frame, (2) a Feature Projector that projects egocentric features to appropriate locations on a floor-plan, (3) a Spatial Memory Tensor of size floor-plan length × width × feature-dims that learns to accumulate projected egocentric features, and (4) a Map Decoder that uses the memory tensor to produce semantic top-down maps. SMNet combines the strengths of (known) projective camera geometry and neural representation learning. On the task of semantic mapping in the Matterport3D dataset, SMNet significantly outperforms competitive baselines by 4.01-16.81% (absolute) on mean-IoU and 3.81-19.69% (absolute) on Boundary-F1 metrics. Moreover, we show how to use the spatio-semantic allocentric representations build by SMNet for the task of ObjectNav and Embodied Question Answering.},

keywords = {AAAI, AI, embodied agents, first-person vision},

pubstate = {published},

tppubtype = {inproceedings}

}