Installing Krawler
As a Command-Line Interface (CLI)
Install it globally with your preferred package manager:
pnpm add -g @kalisio/krawlernpm install -g @kalisio/krawleryarn global add @kalisio/krawlerYou can now launch the CLI on a job file:
krawler jobfile.jsAs a module
As a dependency in another module/app:
pnpm add @kalisio/krawlernpm install @kalisio/krawler --saveyarn add @kalisio/krawlerKrawler is published as a native ES module, so import the symbols you need:
import { hooks, stores, tasks, jobs } from '@kalisio/krawler'In development mode
When contributing to Krawler or testing local changes, work from the krawler-ekosystem monorepo, which is managed with pnpm workspaces:
git clone https://github.com/kalisio/krawler-ekosystem
cd krawler-ekosystem
pnpm install
# Run the CLI from the workspace
pnpm --filter @kalisio/krawler exec krawler jobfile.jsEach job package under packages/krawler-<job> references the framework through the workspace, so a local change to @kalisio/krawler is immediately picked up by the jobs.
Please refer to the KDK documentation to set up your development environment.
As a Docker container
A ready-to-use image is published on Docker Hub:
docker pull kalisio/krawlerWhen using Krawler as a Docker container, the arguments to the CLI have to be provided through the ARGS environment variable, along with any other required variables and the data volume to make inputs accessible within the container and to get output files back:
docker run --name krawler --rm \
-v /mnt/data:/opt/krawler/data \
-e "ARGS=/opt/krawler/data/jobfile.js" \
-e S3_BUCKET=krawler \
kalisio/krawlerTIP
Job images are built on top of the published Krawler image. If you maintain a job in the monorepo, read Building Krawler jobs to understand how the Krawler base version is pinned and how images are released.