
Who Lingbot Map is for#
Mapping a long indoor or outdoor walkthrough
Feed a handheld or vehicle video to windowed mode and render a flythrough of the point cloud. This suits surveyors, robotics teams, and researchers testing SLAM alternatives.
Skip if:
You need a textured mesh or a polished scan app for non-technical users.
Benchmarking streaming reconstruction
Use the released evaluation scripts for KITTI, Oxford Spires, and other datasets to compare a feed-forward method against optimization-based pipelines.
Skip if:
You do not have an NVIDIA GPU with enough memory for the checkpoint.
The problem it solves#
Reconstructing a 3D map from a long video normally relies on paid scanning services or offline optimization that runs after capture and struggles with drift on long trajectories. Teams that want to keep footage local, or need reconstruction to keep pace with incoming frames, have few open options that combine pose estimation and dense geometry in one streaming model.
How it solves it#
Geometric Context Transformer
One streaming architecture combines anchor context for coordinate grounding, a pose-reference window for dense geometric cues, and trajectory memory for long-range drift correction.
Paged KV cache streaming
With FlashInfer attention, inference runs at about 20 FPS on 518x378 frames over sequences longer than 10,000 frames. An SDPA fallback works without FlashInfer.
Interactive viser viewer
demo.py loads an image folder or video and opens a browser viewer at localhost:8080, with sky masking, confidence thresholds, and point size options.
Windowed mode for long sequences
For sequences over about 3,000 frames, windowed inference with a keyframe interval limits KV cache growth and helps when poses collapse.
Offline rendering pipeline
demo_render/batch_demo.py runs inference and produces a headless point-cloud flythrough MP4. The README shows a 25,000-frame, 13-minute indoor walkthrough.
Evaluation benchmark scripts
The repository includes evaluation scripts for datasets such as KITTI and Oxford Spires, so you can reproduce results on your own hardware.
Strengths and trade-offs#
Strengths
- Built for long sequencesTrajectory memory, keyframe intervals, and windowed mode target the drift and memory growth that hit streaming reconstruction on videos with thousands of frames.
- Permissive license, local controlApache-2.0 code and downloadable checkpoints on Hugging Face and ModelScope mean footage and models stay on your own machine.
- Active, widely watched repositoryThe project has 17,701 stars and 1,951 forks, and the maintainers have pushed fixes to the KV cache backends and added benchmarks and demos since April 2026.
Trade-offs
- -Needs a CUDA GPU and careful setupThe recommended stack is Python 3.10, PyTorch 2.8.0 with CUDA 12.8, and FlashInfer. The batch renderer also needs NVIDIA Kaolin and compiled CUDA extensions. Limited VRAM calls for flags like --offload_to_cpu.
- -Quality drops past the training rangeThe model trains with video RoPE on 320 views, so caching more views degrades results. It does not reset state by default, so very long distances can cause pose collapse until you switch to windowed mode.
- -Research code, not a productThere is no hosted service, mobile capture app, or mesh export workflow in the README. The project has 81 open issues, and the authors say a stronger long-sequence model is still in training.
Lingbot Map vs alternatives#
LingBot-Map vs Polycam
Polycam is a commercial scanning app. LingBot-Map is a research model you run yourself, so the two serve different people. With LingBot-Map you get Apache-2.0 code, downloadable weights, and full control over where footage is processed. You also take on a CUDA setup and command-line workflows.
Choose LingBot-Map when you need streaming pose and point-cloud output from long video and want it on your own hardware. Choose a commercial app when you want a finished capture experience and do not want to manage a GPU environment.
LingBot-Map vs Matterport
Matterport is a commercial platform for capturing and sharing spaces. LingBot-Map does not offer hosted tours, sharing, or managed capture. It reconstructs geometry from images or video and shows it in a local viser viewer or an offline MP4 render.
The README's indoor example covers a 13-minute, roughly 25,000-frame walkthrough, which shows the kind of footage the model is designed for. Teams that need a hosted deliverable for clients will still want a commercial platform.
Where the open approach wins
The trade is control against convenience. The code, the checkpoints, and the evaluation scripts are all available, so you can inspect results, change parameters such as keyframe interval, and reproduce benchmarks. The cost is setup effort and the research-grade state of the project.
Quick start#
Setup needs a CUDA GPU, a conda environment, and a downloaded checkpoint from Hugging Face.
```bash
git clone https://github.com/Robbyant/lingbot-map.git
conda create -n lingbot-map python=3.10 -y
pip install torch==2.8.0 torchvision==0.23.0 --index-url https://download.pytorch.org/whl/cu128
pip install -e .
pip install --index-url https://pypi.org/simple flashinfer-python
```What it's built on#
- Languages
- JavaScriptPython
FAQ#
What license does LingBot-Map use?
The GitHub metadata lists Apache-2.0, a permissive open source license that allows commercial use. Check the LICENSE file in the repository and the model pages for the terms that apply to the checkpoints you download.
What hardware does LingBot-Map need?
The README targets an NVIDIA GPU with CUDA 12.8 and PyTorch 2.8.0. FlashInfer is recommended for speed, and an SDPA fallback exists. For low memory, use --offload_to_cpu or --num_scale_frames 2. A community fork targets an 8 GB RTX 4060.
How do I run it on a long video?
Use windowed mode, for example python demo.py with --video_path, --mode windowed, --window_size 128, --overlap_keyframes 16, and --keyframe_interval 2. For very long clips, the offline pipeline in demo_render/batch_demo.py renders an MP4 instead of using the interactive viewer.
Where do I get the model weights?
The README links the lingbot-map checkpoint on Hugging Face (robbyant/lingbot-map) and ModelScope. A stage-1 checkpoint is also listed for loading into the VGGT model for bidirectional inference.
Similar open-source tools#
God's Eye View
Real-time global intelligence on a 3D globe
Vane
AI-powered answering engine for private, cited web searches
Dyad
Local, open-source AI app builder with your own API keys
Cloudflare Os
Open source AI workspace with sandboxed apps and Gatekeeper security
Effect
Typed errors, DI, and concurrency for TypeScript
Firebase Ios Sdk
Open source Apple SDK for Firebase auth, data, push and crashes

