Skip to content

Understanding Krawler ​

Krawler is powered by Feathers and relies on two of its main abstractions: services and hooks. We assume you are familiar with this technology.

Main concepts ​

Krawler manipulates three kinds of entities:

  • a store defines where the extracted/processed data will reside,
  • a task defines what data is to be extracted and how to query it,
  • a job defines what tasks are to be run to fulfill a request (i.e. sequencing).

On top of this, hooks provide a set of functions that can typically be run before/after a task/job, such as a conversion after a download or task generation before a job run. More or less, this allows you to create a processing pipeline.

Regarding store management we rely on abstract-blob-store, which abstracts a lot of different storage backends (local file system, AWS S3, in-memory, etc.), and is already used by feathers-blob.

Global overview ​

The following figure depicts the global architecture and all concepts at play:

Architecture

What is inside? ​

Krawler is made possible and mainly powered by the following stack:

Krawler ships as a native ES module ("type": "module") and targets Node.js >= 20.

Going further ​

Once the concepts are clear, head to the API reference to learn how each entity is configured: