Skip to content

Using Krawler ​

The problem with hooks is that they are configured at application setup time and usually fixed during the whole application lifecycle. It means you would have to create an application instance for each pipeline you'd like to have — not so simple. This is the reason why Krawler is mainly used as a command-line utility (CLI), where each execution sets up a new application with a hooks pipeline according to the job to be done.

However, using the CLI, you can also launch it as a standard web application/API. You can then POST job or task requests to the exposed services, e.g. on localhost:3030/api/jobs.

Command-Line Interface ​

Internal API ​

The underlying implementation is managed by the global run(jobfile, options) function:

  • jobfile: a path to a local job file or a job file JSON object
  • options:
    • cron: a CRON pattern to schedule the job at regular intervals, e.g. */5 * * * * * will run it every 5 seconds; if not provided it will be run only once
    • run: force a first run on launch when scheduling a job with a cron pattern
    • proxy: proxy URL to be used for HTTP requests
    • proxyHttps: proxy URL to be used for HTTPS requests
    • user: user name to be used for authentication
    • password: user password to be used for authentication
    • debug: output debug messages
    • sync: activate the sync module with the given connection URI so that internal events can be listened to externally
    • port: port to be used by Krawler (defaults to 3030)
    • api: launch Krawler as a web service/API
    • apiPrefix: API prefix to be used when launching Krawler as a web service/API (defaults to /api)

This function is responsible for parsing the job definition, including all the required parameters, to call the underlying services with the relevant hooks configured (see below).

External API ​

The job file is the sole mandatory argument of the CLI, and options are read from the CLI arguments using shortcuts like this:

bash
krawler --user user_name -p password -P proxy_url --cron "*/5 * * * * *" path_to_jobfile.js

The available CLI flags are:

FlagDescription
-d, --debugVerbose output for debugging
-a, --apiSetup as a web app by exposing an API
-ap, --api-prefix [prefix]Change the API prefix (defaults to /api)
-po, --port [port]Change the port to be used (defaults to 3030)
-c, --cron [pattern]Schedule the job using a cron pattern
-r, --runForce a first run on launch when scheduling with a cron pattern
-P, --proxy [proxy]Proxy to be used for HTTP (and HTTPS)
-PS, --proxy-https [proxy]Proxy to be used for HTTPS
-u, --user [user]User name to be used for authentication
-p, --password [password]User password to be used for authentication
-s, --sync [uri]Activate the sync module with the given connection URI

A job file can be a JSON or JS file (it will be imported) and its structure is the following:

js
const job = {
  // Options for the job executor
  options: {
    workersLimit: 4,
    faultTolerant: true
  },
  // Store to be used for job output
  store: 'job-store',
  // Common options for all generated tasks
  taskTemplate: {
    // Store to be used for task output
    store: 'job-store',
    id: '<%= jobId %>-<%= taskId %>',
    type: 'xxx',
    options: {
      // ...
    }
  },
  // Hooks setup
  hooks: {
    // Tasks service hooks
    tasks: {
      // Hooks to be run after task creation
      after: {
        // Each entry is a hook name and an associated options object
        computeSomething: {
          hookOption: '...'
        }
      }
    },
    // Jobs service hooks
    jobs: {
      // Hooks to be run before job creation
      before: {
        generateTasks: {
          hookOption: '...'
        }
      },
      // Hooks to be run after job creation
      after: {
        generateOutput: {
          hookOption: '...'
        }
      }
    }
  },
  // The list of tasks to run if not generated by hooks
  tasks: [
    // ...
  ]
}

export default job

TIP

When running Krawler as a web API, note that only the hooks pipeline is mandatory in the job file. Indeed, job and task objects will then be sent by requesting the exposed web services.

Healthcheck ​

Healthcheck endpoint ​

When running Krawler as a cron job, note that it provides a healthcheck endpoint, e.g. on localhost:3030/api/healthcheck. The following JSON structure is returned:

  • isRunning: boolean indicating if the cron job is currently running
  • duration: last run duration in seconds
  • nbSkippedJobs: number of times the scheduled job has been skipped due to an on-going one
  • error: returned error object whenever the cron job has errored
  • nbFailedTasks: number of failed tasks for the last run of fault-tolerant jobs
  • nbSuccessfulTasks: number of successful tasks for the last run of fault-tolerant jobs
  • successRate: ratio of successful tasks / total tasks

The returned HTTP code is 500 whenever an error has occurred in the last run, 200 otherwise.

TIP

You can add your custom data to the healthcheck structure using the healthcheck hook.

Healthcheck command ​

For convenience, Krawler also includes a built-in healthcheck script that can be used e.g. by Docker. This script uses options similar to the CLI plus some specific ones:

  • debug: output debug messages
  • port: port used by Krawler (defaults to 3030)
  • api: indicates if Krawler has been launched as a web service/API
  • api-prefix: API prefix used when launching Krawler as a web service/API (defaults to /api)
  • success-rate: the success rate for fault-tolerant jobs to be considered as successful when greater or equal (defaults to 1)
  • max-duration: the maximum run duration in seconds for fault-tolerant jobs to be considered as failed if greater than (defaults to unset)
  • nb-skipped-jobs: the number of skipped runs for scheduled fault-tolerant jobs to be considered as failed (defaults to 3)
  • slack-webhook: Slack webhook URL to post messages on failure (defaults to process.env.SLACK_WEBHOOK_URL)
  • message-template: message template used on failure for console and Slack output (defaults to Job <%= jobId %>: <%= error.message %>)
  • link-template: link template used on failure for Slack output (defaults to an empty value)

TIP

Templates are generated with the healthcheck structure and environment variables as context. Learn more about templating.