krawler
A minimalist (geospatial) ETL.
Overview
Krawler aims at making the automated process of extracting and processing (geographic) data from heterogeneous sources easy. It can be viewed as a minimalist Extract, Transform, Load (ETL). ETL refers to a process where data is
- extracted from heterogeneous data sources (e.g. databases or web services);
- transformed in a target format or structure for the purposes of querying and analysis (e.g. JSON or CSV);
- loaded into a final target data store (e.g. a file system or a database).

ETL naturally leads to the concept of a pipeline: a set of processing functions (called hooks in krawler) connected in series, often executed in parallel, where the output of one function is the input of the next one. The execution of a given pipeline on an input dataset to produce the associated output is a job performed by krawler.
A set of introduction articles to krawler details:
Installation
Install the CLI globally with your preferred package manager:
pnpm add -g @kalisio/krawlernpm install -g @kalisio/krawleryarn global add @kalisio/krawlerOr pull the ready-to-use Docker image:
docker pull kalisio/krawlerSee the installation guide for usage as a module, in development mode and as a Docker container.
Documentation
The documentation is organized in two parts:
- Guides — a progressive walkthrough of the framework:
- Understanding Krawler — main concepts and architecture
- Installing Krawler — CLI, module and Docker setups
- Using Krawler — the job file, CLI options and healthcheck
- Extending Krawler — register your own stores, tasks, jobs and hooks
- Reference — the exhaustive API:
- Services — stores, tasks and jobs
- Hooks — the built-in processing functions
- Known issues — common pitfalls and their workarounds
If you intend to package a job image on top of Krawler, see also Building Krawler jobs.
What is inside?
Krawler is powered by Feathers and relies on a curated stack, notably:
- Feathers — the underlying services and hooks framework
- Lodash — JavaScript utility library, also used for templating
- gdal-async — Node.js bindings of GDAL / OGR used to process rasters and vectors
- got — used to manage HTTP requests
- js-yaml — used to process YAML files
- xml2js — used to process XML files
- Papa Parse — used to read and write CSV files
- abstract-blob-store — used to abstract storage backends
- node-postgres — used to manage PostgreSQL databases
- node-mongodb-native — used to manage MongoDB databases
License
Licensed under the MIT license.
Copyright (c) 2026 Kalisio.