Best Practices for cron
This article is based on the author's own research and testing and does not guarantee the accuracy or completeness of the information provided.
If you use the information in this article, you do so at your own risk. The author assumes no responsibility for any issues or damages resulting from its use.
An issue occurred when temporarily stopping cron, which led to a discussion in a certain place about things such as change management. This is an overview of my best practices for cron.
* For now, I have just written down what I was thinking and what came to mind while listening to the discussion, so I may remember other things later or realize that "this was actually wrong after all."
* These are my personal cron tips, so feel free to customize them or optimize them for each project.
* This is a rewrite of an old article, so some parts may differ from the current environment.
Development
First, some practices from the development side (cron implementers/configurators).
Use /etc/cron.d
Assuming that it is supported by the environment you are using, use /etc/cron.d/ if possible.
* Do not edit crontab directly with crontab -e.
Specifically, place multiple crontab configuration files under /etc/cron.d/.
ls /etc/cron.d/
system-a
system-b
Example of system-a:
# system-a cron jobs
SHELL=/bin/bash
PATH=/sbin:/bin:/usr/sbin:/usr/bin
MAILTO=root
0 * * * * sysauser sysa-task01
30 * * * * sysauser sysa-task02
This provides the following benefits:
- crontab settings can be grouped in a meaningful way.
- By assigning appropriate permissions to individual configurations, it is also possible to prevent unnecessary modification of crontab settings for other systems.
- It works better when version-controlling them, just like cron shell scripts (described below).
Allowing individual users to have their own cron settings (users editing them with crontab -e) is also OK, but in many cases, I think there are relatively few use cases where individual users need to use cron.
Use shell scripts to execute tasks
For simple commands or processes that can be executed with a one-liner, it is possible to directly specify system or middleware commands in crontab.
However, when executing tasks managed by an organization, I recommend always creating a shell script to execute the task and implementing the processing in that file. Then, specify the shell script to be executed in crontab.
This provides the following benefits:
- The handling of environment variables differs when cron runs a command from when a user logs in, which can sometimes cause problems. By using a shell script, it becomes easy to set the environment variables required before executing commands, or to change them when necessary.
- By version-controlling the shell script itself, it becomes possible to track the history of changes to the processing, and to perform appropriate reviews by going through a PR/MR process.
- It also becomes easy to call another script for common processing (such as setting environment variables, notifications, or logging) from the script specified in cron before calling the actual processing.
Establish coding rules for scripts
I don't think it is necessary to make the rules too strict, but it is better to unify the level of description and processing as much as possible, for example by preparing an easy-to-understand template.
- Write comments at an appropriate level of detail.
- Turn values that may change into variables.
- Define common processing and use it. For example, make it a common script and call it.
Use version control
The reasons for using version control include, of course, being able to restore the previous version or check differences when a problem occurs. However, the following points should also be aimed for.
Deployment using CI/CD
At present, I think the number of projects that use CI/CD for application deployment is increasing, but Linux scripts, cron configurations, and so on should also be managed using a version control system and deployed using CI/CD.
This is also related to concepts proposed by The Twelve Factor App.Reviews using PR/MR
When an incident occurs related to a release or change, I feel like I have heard the following kinds of explanations many times:- One person did the work.
- There was no checking process, or it was not functioning.
- No review was performed.
And then, as a countermeasure: "We will make sure to do it properly from now on."
However, with this kind of countermeasure, even if you only put up a slogan, things often return to the way they were after a while. It is more effective to use the power of systems/tools by always using PR/MR (Pull Request/Merge Request) so that checks are always performed.
* GitLab should also allow you to restrict permissions for pushing. If you are going to require MRs, restrict direct pushes to designated branches, for example. Use tools to enforce the rules rather than relying on slogans.
Prevent multiple executions
When implementing an application, it is possible to use mechanisms provided by the processing language (such as Java) to prevent multiple executions. However, multiple executions can also be prevented by using flock, a standard Linux tool.
*/5 * * * * appuser flock -n /var/lock/cron-task.lock cron-task
Here is an actual example.
Script
Create a script containing the processing.
Create /root/scripts/sleep.sh as follows.
#!/bin/bash
INTERVAL=300
sleep $INTERVAL
echo "$INTERVAL sec slept"
It simply sleeps for the number of seconds specified by the interval and then outputs a message with echo. It is configured so that the job continues for 300 seconds (5 minutes).
cron configuration
For the cron configuration, create /etc/cron.d/test.
# put some comment here
* * * * * root flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh
flock prevents multiple executions. It runs at one-minute intervals.
Execution result
Output from /var/log/cron:
Feb 8 16:07:01 mosaos-dev CROND[7083]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:08:01 mosaos-dev CROND[7087]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:09:01 mosaos-dev CROND[7089]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:10:01 mosaos-dev CROND[7091]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:11:01 mosaos-dev CROND[7093]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:12:01 mosaos-dev CROND[7082]: (root) CMDOUT (300 sec slept)
Feb 8 16:12:01 mosaos-dev CROND[7096]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:13:01 mosaos-dev CROND[7100]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:14:01 mosaos-dev CROND[7102]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:15:01 mosaos-dev CROND[7104]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:16:01 mosaos-dev CROND[7108]: (root) CMD (flock -n /var/lock/cron-test-sleep.lock /root/scripts/sleep.sh)
Feb 8 16:17:01 mosaos-dev CROND[7095]: (root) CMDOUT (300 sec slept)
The cron job runs at one-minute intervals, but you can see that the actual processing (echo) occurs at five-minute intervals.
Use timeouts
Depending on the type of task, using a timeout can be useful.
Normally, jobs executed by cron run without a time limit, but this may not always be desirable.
For example, for a job that runs at regular intervals, if it does not finish within a certain amount of time, the job may be executed again while the previous execution is still running.
In such cases, try using timeout (/usr/bin/timeout) or something similar to limit the execution time.
*/5 * * * * appuser timeout 10s cron-task
As with flock, here is an actual example.
Script
#!/bin/bash
while true
do
date
sleep 1
done
Outputs date at one-second intervals.
cron configuration
Edit the /etc/cron.d/test file.
# put some comment here
* * * * * root timeout 10s /root/scripts/loop.sh
Timeout after 10 seconds.
Execution result
Feb 8 16:49:01 mosaos-dev CROND[7261]: (root) CMD (timeout 10s /root/scripts/loop.sh)
Feb 8 16:49:01 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:01 JST)
Feb 8 16:49:02 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:02 JST)
Feb 8 16:49:03 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:03 JST)
Feb 8 16:49:04 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:04 JST)
Feb 8 16:49:05 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:05 JST)
Feb 8 16:49:06 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:06 JST)
Feb 8 16:49:07 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:07 JST)
Feb 8 16:49:08 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:08 JST)
Feb 8 16:49:09 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:09 JST)
Feb 8 16:49:10 mosaos-dev CROND[7260]: (root) CMDOUT (2022年 2月 8日 火曜日 16:49:10 JST)
Feb 8 16:50:01 mosaos-dev CROND[7287]: (root) CMD (timeout 10s /root/scripts/loop.sh)
Feb 8 16:50:01 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:01 JST)
Feb 8 16:50:02 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:02 JST)
Feb 8 16:50:03 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:03 JST)
Feb 8 16:50:04 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:04 JST)
Feb 8 16:50:05 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:05 JST)
Feb 8 16:50:06 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:06 JST)
Feb 8 16:50:07 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:07 JST)
Feb 8 16:50:08 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:08 JST)
Feb 8 16:50:09 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:09 JST)
Feb 8 16:50:10 mosaos-dev CROND[7286]: (root) CMDOUT (2022年 2月 8日 火曜日 16:50:10 JST)
You can see that the processing is interrupted after 10 seconds. Depending on the task, terminating it midway may cause serious problems, so be careful about where you use this.
Operations
Operational practices.
Restrict permissions
When running a cron job, specify the execution user appropriately and restrict its permissions.
0 * * * * root cron-task
The above crontab configuration specifies that it should be run as the root user. However, when it is executed as root, it can be difficult to impose restrictions, even if, for example, the execution job is something that could destroy an important part of a system directory.
It is better to create a user with the minimum permissions required for each application and specify that user instead.
0 * * * * appuser cron-task
Do not discard output
*/5 * * * * appuser cron-task >/dev/null 2>&1
Do not redirect output to /dev/null.
When looking at examples on the Internet and copying them as they are, there may be cases where output is discarded without really understanding why.
However, this means throwing away clues that could be useful when a problem occurs, so unless there is a specific reason, do not discard the output.
Other
I have seen many environments where, in incident management, a ticket is closed after simply dealing with the incident without taking any measures to prevent it from happening again. In such environments, problems caused by simple carelessness are likely to continue occurring in the future.
If similar problems occur frequently, it may be better to consider introducing a platform (job management tool) that makes this kind of problem less likely to occur.If commercial management tools are expensive... open source is also an option. For example, DolphinScheduler.
I have never used it, but judging from the screenshots, it looks pretty good. If it really is what it claims to be, a "big data distributed workflow scheduling system," it may be usable for some time to come.If containerization will be promoted in the future, use CronJob in Kubernetes, for example.
Keep in mind that the optimal choice may differ depending on the platform.Even if there are best practices, they are meaningless if the individual developers and operations staff responsible for implementation do not follow them.
Rather than handling this at the project or individual level, the company as a whole should establish a certain level of policy and roll out the best practices across the organization.