Update README.md
Browse files
README.md
CHANGED
|
@@ -33,15 +33,28 @@ Trained on approximately 8 million posed image pairs, <strong>Reloc3r</strong> a
|
|
| 33 |
</p>
|
| 34 |
<be>
|
| 35 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 36 |
## TODO List
|
| 37 |
|
| 38 |
-
- [x] Release pre-trained weights and inference code.
|
| 39 |
-
- [x] Release evaluation code for ScanNet1500, MegaDepth1500 and Cambridge datasets.
|
| 40 |
-
- [x] Release
|
| 41 |
-
- [
|
| 42 |
-
- [ ] Release
|
| 43 |
-
- [ ] Release
|
| 44 |
-
- [ ] Release
|
| 45 |
|
| 46 |
|
| 47 |
## Installation
|
|
@@ -57,7 +70,7 @@ cd reloc3r
|
|
| 57 |
2. Create the environment using conda
|
| 58 |
```bash
|
| 59 |
conda create -n reloc3r python=3.11 cmake=3.14.0
|
| 60 |
-
conda activate reloc3r
|
| 61 |
conda install pytorch torchvision pytorch-cuda=12.1 -c pytorch -c nvidia # use the correct version of cuda for your system
|
| 62 |
pip install -r requirements.txt
|
| 63 |
# optional: you can also install additional packages to:
|
|
@@ -65,7 +78,7 @@ pip install -r requirements.txt
|
|
| 65 |
pip install -r requirements_optional.txt
|
| 66 |
```
|
| 67 |
|
| 68 |
-
3. Optional: Compile the cuda kernels for RoPE
|
| 69 |
```bash
|
| 70 |
# Reloc3r relies on RoPE positional embeddings for which you can compile some cuda kernels for faster runtime.
|
| 71 |
cd croco/models/curope/
|
|
@@ -73,33 +86,14 @@ python setup.py build_ext --inplace
|
|
| 73 |
cd ../../../
|
| 74 |
```
|
| 75 |
|
| 76 |
-
4. Optional: Download the checkpoints [Reloc3r-224](https://huggingface.co/siyan824/reloc3r-224)/[Reloc3r-512](https://huggingface.co/siyan824/reloc3r-512). The pre-trained model weights will automatically download when running the evaluation and demo code below.
|
| 77 |
-
|
| 78 |
|
| 79 |
-
## Relative Pose Estimation on ScanNet1500 and MegaDepth1500
|
| 80 |
|
| 81 |
-
|
| 82 |
-
Then run the following script. You will obtain results similar to those presented in our paper.
|
| 83 |
-
```bash
|
| 84 |
-
bash scripts/eval_relpose.sh
|
| 85 |
-
```
|
| 86 |
-
<strong>Note:</strong> To achieve faster inference speed, set `--amp=1`. This enables evaluation with `fp16`, which increases speed from <strong>24 FPS</strong> to <strong>40 FPS</strong> on an RTX 4090 with Reloc3r-512, without any accuracy loss.
|
| 87 |
|
|
|
|
| 88 |
|
| 89 |
-
|
| 90 |
-
|
| 91 |
-
Download the dataset [here](https://drive.google.com/file/d/1XcJIVRMma4_IClJdRq6rwBKX3ZPet5az/view?usp=sharing) and unzip it to `./data/cambridge/`.
|
| 92 |
-
Then run the following script. You will obtain results similar to those presented in our paper.
|
| 93 |
-
```bash
|
| 94 |
-
bash scripts/eval_visloc.sh
|
| 95 |
-
```
|
| 96 |
-
|
| 97 |
-
|
| 98 |
-
## Demo for Wild Images
|
| 99 |
-
|
| 100 |
-
In the demos below, you can run Reloc3r on your own data.
|
| 101 |
-
|
| 102 |
-
For relative pose estimation, try the demo code in `wild_relpose.py`. We provide some [image pairs](https://drive.google.com/drive/folders/1TmoSKrtxR50SlFoXOwC4a9aGr18h00yy?usp=sharing) used in our paper.
|
| 103 |
|
| 104 |
```bash
|
| 105 |
# replace the args with your paths
|
|
@@ -112,9 +106,10 @@ Visualize the relative pose
|
|
| 112 |
python visualization.py --mode relpose --pose_path ./data/wild_images/pose2to1.txt
|
| 113 |
```
|
| 114 |
|
| 115 |
-
For visual localization, the demo code in `wild_visloc.py` estimates absolute camera poses from sampled frames in self-captured videos.
|
| 116 |
|
| 117 |
-
|
|
|
|
| 118 |
|
| 119 |
```bash
|
| 120 |
# replace the args with your paths
|
|
@@ -128,9 +123,50 @@ python visualization.py --mode visloc --pose_folder ./data/wild_video/ids_poses/
|
|
| 128 |
```
|
| 129 |
|
| 130 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 131 |
## Citation
|
| 132 |
|
| 133 |
-
If you find our work helpful in your research, please consider citing:
|
| 134 |
```
|
| 135 |
@article{reloc3r,
|
| 136 |
title={Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization},
|
|
@@ -148,4 +184,5 @@ Our implementation is based on several awesome repositories:
|
|
| 148 |
- [Croco](https://github.com/naver/croco)
|
| 149 |
- [DUSt3R](https://github.com/naver/dust3r)
|
| 150 |
|
| 151 |
-
We thank the respective authors for open-sourcing their code.
|
|
|
|
|
|
| 33 |
</p>
|
| 34 |
<be>
|
| 35 |
|
| 36 |
+
|
| 37 |
+
## Table of Contents
|
| 38 |
+
|
| 39 |
+
- [TODO List](#todo-list)
|
| 40 |
+
- [Installation](#installation)
|
| 41 |
+
- [Usage](#usage)
|
| 42 |
+
- [Evaluation on Relative Camera Pose Estimation](#evaluation-on-relative-camera-pose-estimation)
|
| 43 |
+
- [Evaluation on Visual Localization](#evaluation-on-visual-localization)
|
| 44 |
+
- [Training](#training)
|
| 45 |
+
- [Citation](#citation)
|
| 46 |
+
- [Acknowledgments](#acknowledgments)
|
| 47 |
+
|
| 48 |
+
|
| 49 |
## TODO List
|
| 50 |
|
| 51 |
+
- [x] Release pre-trained weights and inference code.
|
| 52 |
+
- [x] Release evaluation code for ScanNet1500, MegaDepth1500 and Cambridge datasets.
|
| 53 |
+
- [x] Release sample code for self-captured images and videos.
|
| 54 |
+
- [x] Release training code and data.
|
| 55 |
+
- [ ] Release evaluation code for other datasets.
|
| 56 |
+
- [ ] Release the accelerated version for visual localization.
|
| 57 |
+
- [ ] Release Gradio Demo.
|
| 58 |
|
| 59 |
|
| 60 |
## Installation
|
|
|
|
| 70 |
2. Create the environment using conda
|
| 71 |
```bash
|
| 72 |
conda create -n reloc3r python=3.11 cmake=3.14.0
|
| 73 |
+
conda activate reloc3r
|
| 74 |
conda install pytorch torchvision pytorch-cuda=12.1 -c pytorch -c nvidia # use the correct version of cuda for your system
|
| 75 |
pip install -r requirements.txt
|
| 76 |
# optional: you can also install additional packages to:
|
|
|
|
| 78 |
pip install -r requirements_optional.txt
|
| 79 |
```
|
| 80 |
|
| 81 |
+
3. Optional: Compile the cuda kernels for RoPE
|
| 82 |
```bash
|
| 83 |
# Reloc3r relies on RoPE positional embeddings for which you can compile some cuda kernels for faster runtime.
|
| 84 |
cd croco/models/curope/
|
|
|
|
| 86 |
cd ../../../
|
| 87 |
```
|
| 88 |
|
| 89 |
+
4. Optional: Download the checkpoints [Reloc3r-224](https://huggingface.co/siyan824/reloc3r-224)/[Reloc3r-512](https://huggingface.co/siyan824/reloc3r-512). The pre-trained model weights will automatically download when running the evaluation and demo code below.
|
|
|
|
| 90 |
|
|
|
|
| 91 |
|
| 92 |
+
## Usage
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 93 |
|
| 94 |
+
Using Reloc3r, you can estimate camera poses for images and videos you captured.
|
| 95 |
|
| 96 |
+
For relative pose estimation, try the demo code in `wild_relpose.py`. We provide some [image pairs](https://drive.google.com/drive/folders/1TmoSKrtxR50SlFoXOwC4a9aGr18h00yy?usp=sharing) used in our paper.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 97 |
|
| 98 |
```bash
|
| 99 |
# replace the args with your paths
|
|
|
|
| 106 |
python visualization.py --mode relpose --pose_path ./data/wild_images/pose2to1.txt
|
| 107 |
```
|
| 108 |
|
| 109 |
+
For visual localization, the demo code in `wild_visloc.py` estimates absolute camera poses from sampled frames in self-captured videos.
|
| 110 |
|
| 111 |
+
> [!IMPORTANT]
|
| 112 |
+
> The demo simply uses the first and last frames as the database, which <strong>requires</strong> overlapping regions among all images. This demo does <strong>not</strong> support linear motion. We provide some [videos](https://drive.google.com/drive/folders/1sbXiXScts5OjESAfSZQwLrAQ5Dta1ibS?usp=sharing) as examples.
|
| 113 |
|
| 114 |
```bash
|
| 115 |
# replace the args with your paths
|
|
|
|
| 123 |
```
|
| 124 |
|
| 125 |
|
| 126 |
+
## Evaluation on Relative Camera Pose Estimation
|
| 127 |
+
|
| 128 |
+
To reproduce our evaluation on ScanNet1500 and MegaDepth1500, download the datasets [here](https://drive.google.com/drive/folders/16g--OfRHb26bT6DvOlj3xhwsb1kV58fT?usp=sharing) and unzip it to `./data/`.
|
| 129 |
+
Then run the following script. You will obtain results similar to those presented in our paper.
|
| 130 |
+
```bash
|
| 131 |
+
bash scripts/eval_relpose.sh
|
| 132 |
+
```
|
| 133 |
+
|
| 134 |
+
> [!NOTE]
|
| 135 |
+
> To achieve faster inference speed, set `--amp=1`. This enables evaluation with `fp16`, which increases speed from <strong>24 FPS</strong> to <strong>40 FPS</strong> on an RTX 4090 with Reloc3r-512, without any accuracy loss.
|
| 136 |
+
|
| 137 |
+
|
| 138 |
+
## Evaluation on Visual Localization
|
| 139 |
+
|
| 140 |
+
To reproduce our evaluation on Cambridge, download the dataset [here](https://drive.google.com/file/d/1XcJIVRMma4_IClJdRq6rwBKX3ZPet5az/view?usp=sharing) and unzip it to `./data/cambridge/`.
|
| 141 |
+
Then run the following script. You will obtain results similar to those presented in our paper.
|
| 142 |
+
```bash
|
| 143 |
+
bash scripts/eval_visloc.sh
|
| 144 |
+
```
|
| 145 |
+
|
| 146 |
+
|
| 147 |
+
## Training
|
| 148 |
+
|
| 149 |
+
We follow [DUSt3R](https://github.com/naver/dust3r) to process the training data. Download the datasets: [CO3Dv2](https://github.com/facebookresearch/co3d), [ScanNet++](https://kaldir.vc.in.tum.de/scannetpp/), [ARKitScenes](https://github.com/apple/ARKitScenes), [BlendedMVS](https://github.com/YoYo000/BlendedMVS), [MegaDepth](https://www.cs.cornell.edu/projects/megadepth/), [DL3DV](https://dl3dv-10k.github.io/DL3DV-10K/), [RealEstate10K](https://google.github.io/realestate10k/).
|
| 150 |
+
|
| 151 |
+
For each dataset, we provide a preprocessing script in the `datasets_preprocess` directory and an archive containing the list of [pairs](https://drive.google.com/drive/folders/193Lv5YB-2OVkqK3k6vnZJi36-FRPcAuu?usp=sharing) when needed. You have to download the datasets yourself from their official sources, agree to their license, and run the preprocessing script.
|
| 152 |
+
|
| 153 |
+
We provide a sample script to train Reloc3r with ScanNet++ on an RTX 3090 GPU
|
| 154 |
+
```bash
|
| 155 |
+
bash scripts/train_small.sh
|
| 156 |
+
```
|
| 157 |
+
|
| 158 |
+
To reproduce our training for Reloc3r-512 with 8 H800 GPUs, run the following script
|
| 159 |
+
```bash
|
| 160 |
+
bash scripts/train.sh
|
| 161 |
+
```
|
| 162 |
+
|
| 163 |
+
> [!NOTE]
|
| 164 |
+
> They are not strictly equivalent to what was used to train Reloc3r, but they should be close enough.
|
| 165 |
+
|
| 166 |
+
|
| 167 |
## Citation
|
| 168 |
|
| 169 |
+
If you find our work helpful in your research, please consider citing:
|
| 170 |
```
|
| 171 |
@article{reloc3r,
|
| 172 |
title={Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization},
|
|
|
|
| 184 |
- [Croco](https://github.com/naver/croco)
|
| 185 |
- [DUSt3R](https://github.com/naver/dust3r)
|
| 186 |
|
| 187 |
+
We thank the respective authors for open-sourcing their code.
|
| 188 |
+
|