Acknowledgements
We are deeply grateful to Bach-Thuan Bui, Ang Cao, Robert Carrillo, Vikas
Chandra, Linzhuo Chen, Tianrun Chen, Yuchao Dai, Andrew Davison, Alexei
Efros, Haiwen Feng, David Forsyth, Jian Gao, Zhongrui Gui, Junlin Han,
Tengda Han, Kaiming He, Wenlong Huang, George Hulm,
Menglin Jia, Hanwen Jiang, Haian Jin, Linyi Jin, Angjoo Kanazawa, Nikhil
Keetha, Zihang Lai, Hongdong Li, Runjia Li, Weiyu Li, Zhengqi Li, Philipp
Lindenberger, Shaohui Liu, Shikun Liu, Iurii Makarov, Dmytro Mishkin,
Linfei Pan, Feike Postmes, Michaël Ramamonjisoa, Janahan Ramanan, Belal
Shaheen, Roman Shapovalov, You Shen, Yujun Shen, Kevin
Sheridan, Zifan Shi, Stanislaw Szymanowicz, Letian Wang, Qianqian Wang,
Yunnan Wang, Michael Wu, Rundi Wu, Shangzhe Wu, Junyu Xie, Yinghao Xu,
Nan Xue, Ceyuan Yang, Jihan Yang, Chuhan Zhang, Junyi Zhang, Tianyuan
Zhang, Kecheng Zheng, Yiran Zhong, and Andrew Zisserman for their
discussions and support.
We thank Mikhail Parkhomenko for kindly granting us permission to use the
film Transmission in our demo. We are especially grateful to
Tianhong Li for insightful discussions on why dense decoding heads should
be MLP-only (also see JiT). We are grateful to Bingyi Kang for discussions
on improving DPT. Discussions with Noah Snavely, Ben
Poole, Aleksander Holynski, and Jon Barron inspired parts of our
exposition in the Further Insights and Discussion sections. We
particularly appreciate
the detailed discussions with Haotong Lin and Yifan Wang on numerous
implementation details.
On a more personal note, this work is named Omega (Ω) because it is
Jianyuan's last PhD paper. In trying to make this work feel complete, he
kept adding new content and pursuing observations that he hoped to
understand more deeply, which inevitably led to delays for months. He is
deeply thankful to all coauthors for their patience, generosity, and
understanding throughout this process.