Understanding Krawler
Krawler is powered by Feathers and relies on two of its main abstractions: services and hooks. We assume you are familiar with this technology.
Main concepts
Krawler manipulates three kinds of entities:
- a store defines where the extracted/processed data will reside,
- a task defines what data is to be extracted and how to query it,
- a job defines what tasks are to be run to fulfill a request (i.e. sequencing).
On top of this, hooks provide a set of functions that can typically be run before/after a task/job, such as a conversion after a download or task generation before a job run. More or less, this allows you to create a processing pipeline.
Regarding store management we rely on abstract-blob-store, which abstracts a lot of different storage backends (local file system, AWS S3, in-memory, etc.), and is already used by feathers-blob.
Global overview
The following figure depicts the global architecture and all concepts at play:

What is inside?
Krawler is made possible and mainly powered by the following stack:
- Feathers — the underlying services and hooks framework
- Lodash — a JavaScript utility library, also used for templating
- gdal-async — the Node.js bindings of GDAL / OGR used to process rasters and vectors
- js-yaml — used to process YAML files
- xml2js — used to process XML files
- Papa Parse — used to read and write CSV files
- abstract-blob-store — used to abstract storage
- got — used to manage HTTP requests
- node-postgres — used to manage PostgreSQL databases
- node-mongodb-native — used to manage MongoDB databases
Krawler ships as a native ES module ("type": "module") and targets Node.js >= 20.
Going further
Once the concepts are clear, head to the API reference to learn how each entity is configured:
- the services reference details stores, tasks and jobs,
- the hooks reference details all the built-in processing functions,
- the known issues cover advanced pipeline patterns (reusing hooks, parallelism, error handling).