siyan824 commited on
Commit
08bc4c5
·
verified ·
1 Parent(s): 230c86c

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +73 -36
README.md CHANGED
@@ -33,15 +33,28 @@ Trained on approximately 8 million posed image pairs, <strong>Reloc3r</strong> a
33
  </p>
34
  <be>
35
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36
  ## TODO List
37
 
38
- - [x] Release pre-trained weights and inference code.
39
- - [x] Release evaluation code for ScanNet1500, MegaDepth1500 and Cambridge datasets.
40
- - [x] Release demo code for wild images and videos.
41
- - [ ] Release evaluation code for other datasets.
42
- - [ ] Release the accelerated version for visual localization.
43
- - [ ] Release Gradio Demo.
44
- - [ ] Release training code and data.
45
 
46
 
47
  ## Installation
@@ -57,7 +70,7 @@ cd reloc3r
57
  2. Create the environment using conda
58
  ```bash
59
  conda create -n reloc3r python=3.11 cmake=3.14.0
60
- conda activate reloc3r
61
  conda install pytorch torchvision pytorch-cuda=12.1 -c pytorch -c nvidia # use the correct version of cuda for your system
62
  pip install -r requirements.txt
63
  # optional: you can also install additional packages to:
@@ -65,7 +78,7 @@ pip install -r requirements.txt
65
  pip install -r requirements_optional.txt
66
  ```
67
 
68
- 3. Optional: Compile the cuda kernels for RoPE
69
  ```bash
70
  # Reloc3r relies on RoPE positional embeddings for which you can compile some cuda kernels for faster runtime.
71
  cd croco/models/curope/
@@ -73,33 +86,14 @@ python setup.py build_ext --inplace
73
  cd ../../../
74
  ```
75
 
76
- 4. Optional: Download the checkpoints [Reloc3r-224](https://huggingface.co/siyan824/reloc3r-224)/[Reloc3r-512](https://huggingface.co/siyan824/reloc3r-512). The pre-trained model weights will automatically download when running the evaluation and demo code below.
77
-
78
 
79
- ## Relative Pose Estimation on ScanNet1500 and MegaDepth1500
80
 
81
- Download the datasets [here](https://drive.google.com/drive/folders/16g--OfRHb26bT6DvOlj3xhwsb1kV58fT?usp=sharing) and unzip it to `./data/`.
82
- Then run the following script. You will obtain results similar to those presented in our paper.
83
- ```bash
84
- bash scripts/eval_relpose.sh
85
- ```
86
- <strong>Note:</strong> To achieve faster inference speed, set `--amp=1`. This enables evaluation with `fp16`, which increases speed from <strong>24 FPS</strong> to <strong>40 FPS</strong> on an RTX 4090 with Reloc3r-512, without any accuracy loss.
87
 
 
88
 
89
- ## Visual Localization on Cambridge
90
-
91
- Download the dataset [here](https://drive.google.com/file/d/1XcJIVRMma4_IClJdRq6rwBKX3ZPet5az/view?usp=sharing) and unzip it to `./data/cambridge/`.
92
- Then run the following script. You will obtain results similar to those presented in our paper.
93
- ```bash
94
- bash scripts/eval_visloc.sh
95
- ```
96
-
97
-
98
- ## Demo for Wild Images
99
-
100
- In the demos below, you can run Reloc3r on your own data.
101
-
102
- For relative pose estimation, try the demo code in `wild_relpose.py`. We provide some [image pairs](https://drive.google.com/drive/folders/1TmoSKrtxR50SlFoXOwC4a9aGr18h00yy?usp=sharing) used in our paper.
103
 
104
  ```bash
105
  # replace the args with your paths
@@ -112,9 +106,10 @@ Visualize the relative pose
112
  python visualization.py --mode relpose --pose_path ./data/wild_images/pose2to1.txt
113
  ```
114
 
115
- For visual localization, the demo code in `wild_visloc.py` estimates absolute camera poses from sampled frames in self-captured videos.
116
 
117
- <strong>Important</strong>: The demo uses the first and last frames as the database, which <strong>requires</strong> overlapping regions among all images. This demo does <strong>not</strong> support linear motion. We provide some [videos](https://drive.google.com/drive/folders/1sbXiXScts5OjESAfSZQwLrAQ5Dta1ibS?usp=sharing) as examples.
 
118
 
119
  ```bash
120
  # replace the args with your paths
@@ -128,9 +123,50 @@ python visualization.py --mode visloc --pose_folder ./data/wild_video/ids_poses/
128
  ```
129
 
130
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
131
  ## Citation
132
 
133
- If you find our work helpful in your research, please consider citing:
134
  ```
135
  @article{reloc3r,
136
  title={Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization},
@@ -148,4 +184,5 @@ Our implementation is based on several awesome repositories:
148
  - [Croco](https://github.com/naver/croco)
149
  - [DUSt3R](https://github.com/naver/dust3r)
150
 
151
- We thank the respective authors for open-sourcing their code.
 
 
33
  </p>
34
  <be>
35
 
36
+
37
+ ## Table of Contents
38
+
39
+ - [TODO List](#todo-list)
40
+ - [Installation](#installation)
41
+ - [Usage](#usage)
42
+ - [Evaluation on Relative Camera Pose Estimation](#evaluation-on-relative-camera-pose-estimation)
43
+ - [Evaluation on Visual Localization](#evaluation-on-visual-localization)
44
+ - [Training](#training)
45
+ - [Citation](#citation)
46
+ - [Acknowledgments](#acknowledgments)
47
+
48
+
49
  ## TODO List
50
 
51
+ - [x] Release pre-trained weights and inference code.
52
+ - [x] Release evaluation code for ScanNet1500, MegaDepth1500 and Cambridge datasets.
53
+ - [x] Release sample code for self-captured images and videos.
54
+ - [x] Release training code and data.
55
+ - [ ] Release evaluation code for other datasets.
56
+ - [ ] Release the accelerated version for visual localization.
57
+ - [ ] Release Gradio Demo.
58
 
59
 
60
  ## Installation
 
70
  2. Create the environment using conda
71
  ```bash
72
  conda create -n reloc3r python=3.11 cmake=3.14.0
73
+ conda activate reloc3r
74
  conda install pytorch torchvision pytorch-cuda=12.1 -c pytorch -c nvidia # use the correct version of cuda for your system
75
  pip install -r requirements.txt
76
  # optional: you can also install additional packages to:
 
78
  pip install -r requirements_optional.txt
79
  ```
80
 
81
+ 3. Optional: Compile the cuda kernels for RoPE
82
  ```bash
83
  # Reloc3r relies on RoPE positional embeddings for which you can compile some cuda kernels for faster runtime.
84
  cd croco/models/curope/
 
86
  cd ../../../
87
  ```
88
 
89
+ 4. Optional: Download the checkpoints [Reloc3r-224](https://huggingface.co/siyan824/reloc3r-224)/[Reloc3r-512](https://huggingface.co/siyan824/reloc3r-512). The pre-trained model weights will automatically download when running the evaluation and demo code below.
 
90
 
 
91
 
92
+ ## Usage
 
 
 
 
 
93
 
94
+ Using Reloc3r, you can estimate camera poses for images and videos you captured.
95
 
96
+ For relative pose estimation, try the demo code in `wild_relpose.py`. We provide some [image pairs](https://drive.google.com/drive/folders/1TmoSKrtxR50SlFoXOwC4a9aGr18h00yy?usp=sharing) used in our paper.
 
 
 
 
 
 
 
 
 
 
 
 
 
97
 
98
  ```bash
99
  # replace the args with your paths
 
106
  python visualization.py --mode relpose --pose_path ./data/wild_images/pose2to1.txt
107
  ```
108
 
109
+ For visual localization, the demo code in `wild_visloc.py` estimates absolute camera poses from sampled frames in self-captured videos.
110
 
111
+ > [!IMPORTANT]
112
+ > The demo simply uses the first and last frames as the database, which <strong>requires</strong> overlapping regions among all images. This demo does <strong>not</strong> support linear motion. We provide some [videos](https://drive.google.com/drive/folders/1sbXiXScts5OjESAfSZQwLrAQ5Dta1ibS?usp=sharing) as examples.
113
 
114
  ```bash
115
  # replace the args with your paths
 
123
  ```
124
 
125
 
126
+ ## Evaluation on Relative Camera Pose Estimation
127
+
128
+ To reproduce our evaluation on ScanNet1500 and MegaDepth1500, download the datasets [here](https://drive.google.com/drive/folders/16g--OfRHb26bT6DvOlj3xhwsb1kV58fT?usp=sharing) and unzip it to `./data/`.
129
+ Then run the following script. You will obtain results similar to those presented in our paper.
130
+ ```bash
131
+ bash scripts/eval_relpose.sh
132
+ ```
133
+
134
+ > [!NOTE]
135
+ > To achieve faster inference speed, set `--amp=1`. This enables evaluation with `fp16`, which increases speed from <strong>24 FPS</strong> to <strong>40 FPS</strong> on an RTX 4090 with Reloc3r-512, without any accuracy loss.
136
+
137
+
138
+ ## Evaluation on Visual Localization
139
+
140
+ To reproduce our evaluation on Cambridge, download the dataset [here](https://drive.google.com/file/d/1XcJIVRMma4_IClJdRq6rwBKX3ZPet5az/view?usp=sharing) and unzip it to `./data/cambridge/`.
141
+ Then run the following script. You will obtain results similar to those presented in our paper.
142
+ ```bash
143
+ bash scripts/eval_visloc.sh
144
+ ```
145
+
146
+
147
+ ## Training
148
+
149
+ We follow [DUSt3R](https://github.com/naver/dust3r) to process the training data. Download the datasets: [CO3Dv2](https://github.com/facebookresearch/co3d), [ScanNet++](https://kaldir.vc.in.tum.de/scannetpp/), [ARKitScenes](https://github.com/apple/ARKitScenes), [BlendedMVS](https://github.com/YoYo000/BlendedMVS), [MegaDepth](https://www.cs.cornell.edu/projects/megadepth/), [DL3DV](https://dl3dv-10k.github.io/DL3DV-10K/), [RealEstate10K](https://google.github.io/realestate10k/).
150
+
151
+ For each dataset, we provide a preprocessing script in the `datasets_preprocess` directory and an archive containing the list of [pairs](https://drive.google.com/drive/folders/193Lv5YB-2OVkqK3k6vnZJi36-FRPcAuu?usp=sharing) when needed. You have to download the datasets yourself from their official sources, agree to their license, and run the preprocessing script.
152
+
153
+ We provide a sample script to train Reloc3r with ScanNet++ on an RTX 3090 GPU
154
+ ```bash
155
+ bash scripts/train_small.sh
156
+ ```
157
+
158
+ To reproduce our training for Reloc3r-512 with 8 H800 GPUs, run the following script
159
+ ```bash
160
+ bash scripts/train.sh
161
+ ```
162
+
163
+ > [!NOTE]
164
+ > They are not strictly equivalent to what was used to train Reloc3r, but they should be close enough.
165
+
166
+
167
  ## Citation
168
 
169
+ If you find our work helpful in your research, please consider citing:
170
  ```
171
  @article{reloc3r,
172
  title={Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization},
 
184
  - [Croco](https://github.com/naver/croco)
185
  - [DUSt3R](https://github.com/naver/dust3r)
186
 
187
+ We thank the respective authors for open-sourcing their code.
188
+