Using Krawler
The problem with hooks is that they are configured at application setup time and usually fixed during the whole application lifecycle. It means you would have to create an application instance for each pipeline you'd like to have — not so simple. This is the reason why Krawler is mainly used as a command-line utility (CLI), where each execution sets up a new application with a hooks pipeline according to the job to be done.
However, using the CLI, you can also launch it as a standard web application/API. You can then POST job or task requests to the exposed services, e.g. on localhost:3030/api/jobs.
Command-Line Interface
Internal API
The underlying implementation is managed by the global run(jobfile, options) function:
- jobfile: a path to a local job file or a job file JSON object
- options:
- cron: a CRON pattern to schedule the job at regular intervals, e.g.
*/5 * * * * *will run it every 5 seconds; if not provided it will be run only once - run: force a first run on launch when scheduling a job with a cron pattern
- proxy: proxy URL to be used for HTTP requests
- proxyHttps: proxy URL to be used for HTTPS requests
- user: user name to be used for authentication
- password: user password to be used for authentication
- debug: output debug messages
- sync: activate the sync module with the given connection URI so that internal events can be listened to externally
- port: port to be used by Krawler (defaults to
3030) - api: launch Krawler as a web service/API
- apiPrefix: API prefix to be used when launching Krawler as a web service/API (defaults to
/api)
- cron: a CRON pattern to schedule the job at regular intervals, e.g.
This function is responsible for parsing the job definition, including all the required parameters, to call the underlying services with the relevant hooks configured (see below).
External API
The job file is the sole mandatory argument of the CLI, and options are read from the CLI arguments using shortcuts like this:
krawler --user user_name -p password -P proxy_url --cron "*/5 * * * * *" path_to_jobfile.jsThe available CLI flags are:
| Flag | Description |
|---|---|
-d, --debug | Verbose output for debugging |
-a, --api | Setup as a web app by exposing an API |
-ap, --api-prefix [prefix] | Change the API prefix (defaults to /api) |
-po, --port [port] | Change the port to be used (defaults to 3030) |
-c, --cron [pattern] | Schedule the job using a cron pattern |
-r, --run | Force a first run on launch when scheduling with a cron pattern |
-P, --proxy [proxy] | Proxy to be used for HTTP (and HTTPS) |
-PS, --proxy-https [proxy] | Proxy to be used for HTTPS |
-u, --user [user] | User name to be used for authentication |
-p, --password [password] | User password to be used for authentication |
-s, --sync [uri] | Activate the sync module with the given connection URI |
A job file can be a JSON or JS file (it will be imported) and its structure is the following:
const job = {
// Options for the job executor
options: {
workersLimit: 4,
faultTolerant: true
},
// Store to be used for job output
store: 'job-store',
// Common options for all generated tasks
taskTemplate: {
// Store to be used for task output
store: 'job-store',
id: '<%= jobId %>-<%= taskId %>',
type: 'xxx',
options: {
// ...
}
},
// Hooks setup
hooks: {
// Tasks service hooks
tasks: {
// Hooks to be run after task creation
after: {
// Each entry is a hook name and an associated options object
computeSomething: {
hookOption: '...'
}
}
},
// Jobs service hooks
jobs: {
// Hooks to be run before job creation
before: {
generateTasks: {
hookOption: '...'
}
},
// Hooks to be run after job creation
after: {
generateOutput: {
hookOption: '...'
}
}
}
},
// The list of tasks to run if not generated by hooks
tasks: [
// ...
]
}
export default jobTIP
When running Krawler as a web API, note that only the hooks pipeline is mandatory in the job file. Indeed, job and task objects will then be sent by requesting the exposed web services.
Healthcheck
Healthcheck endpoint
When running Krawler as a cron job, note that it provides a healthcheck endpoint, e.g. on localhost:3030/api/healthcheck. The following JSON structure is returned:
isRunning: boolean indicating if the cron job is currently runningduration: last run duration in secondsnbSkippedJobs: number of times the scheduled job has been skipped due to an on-going oneerror: returned error object whenever the cron job has errorednbFailedTasks: number of failed tasks for the last run of fault-tolerant jobsnbSuccessfulTasks: number of successful tasks for the last run of fault-tolerant jobssuccessRate: ratio of successful tasks / total tasks
The returned HTTP code is 500 whenever an error has occurred in the last run, 200 otherwise.
TIP
You can add your custom data to the healthcheck structure using the healthcheck hook.
Healthcheck command
For convenience, Krawler also includes a built-in healthcheck script that can be used e.g. by Docker. This script uses options similar to the CLI plus some specific ones:
- debug: output debug messages
- port: port used by Krawler (defaults to
3030) - api: indicates if Krawler has been launched as a web service/API
- api-prefix: API prefix used when launching Krawler as a web service/API (defaults to
/api) - success-rate: the success rate for fault-tolerant jobs to be considered as successful when greater or equal (defaults to
1) - max-duration: the maximum run duration in seconds for fault-tolerant jobs to be considered as failed if greater than (defaults to unset)
- nb-skipped-jobs: the number of skipped runs for scheduled fault-tolerant jobs to be considered as failed (defaults to
3) - slack-webhook: Slack webhook URL to post messages on failure (defaults to
process.env.SLACK_WEBHOOK_URL) - message-template: message template used on failure for console and Slack output (defaults to
Job <%= jobId %>: <%= error.message %>) - link-template: link template used on failure for Slack output (defaults to an empty value)
TIP
Templates are generated with the healthcheck structure and environment variables as context. Learn more about templating.