EO-VLM: Benchmarking Vision-Language Models on Earth Observation

Overview

EO-VLM is a benchmarking framework for evaluating vision-language models on Earth observation tasks using satellite and UAV imagery. It provides a unified, reproducible evaluation pipeline across tasks including scene classification, visual question answering, counting, and temporal analysis. Evaluation uses the GEOBench-VLM dataset.

Command-line tools cover single-image evaluation, temporal (multi-frame) evaluation, and inspection of individual predictions.