Scaling a Cron Without Running It Twice: A Lease-Based Distributed Lock
Problem & Background
The architecture of the system is designed in a way that the main server is not responsible for handling cron jobs and it's execution. There is a separate pod, which is nothing but a replica of the API, that handles the entire cron and its execution. There are two identical copies of pods:
- API: which is responsible for handling all the API requests and a few other things.
- Cron: which is responsible for handling cron jobs and their execution.
The system is deployed on Kubernetes with HPA (Horizontal Pod Autoscaler) enabled. The main API server has HPA enabled so it can horizontally scale based on the threshold limits, but the cron has strictly one replica. That means it cannot horizontally scale. One of the main reasons we can't enable HPA directly on the cron is this: let's say if we have multiple replicas; the same operation will be executed multiple times because multiple schedulers will be running on each replica, and that leads to duplicate execution.
Ideally, the execution we want is that there should be a global scheduler, and all the replicas should be more like workers. So the cron should be scaled horizontally, and each replica should be responsible for executing the cron.
Approach
There are ideally two things that we could have done:
- Using distributed locking to ensure idempotent execution
- Delegating the scheduler outside of cron to a queue-based mechanism that our infra may support and then using each cron replica as a worker polling from the queue; that way, our workers are scalable based on KEDA, not based on HPA.
We chose option one, considering some trade-offs and the situation of the project. Although a better and long-term solution is still option 2, the reason behind that is: We are only worrying about the execution since we are delegating the scheduling part to a queue-based mechanism. (The queue-based mechanism in the project is a self-hosted open-source tool)
Implementation
We implemented a database table called cron_with_leases with:
jobId, primary key, storing the unique cron nameowner, stores pod ID (ID generated at the time of pod creation)expiredAt, stores the timestamp when the lease expiresacquiredAt, stores the timestamp when the lease was acquired
This table will have as many records as the total number of cron jobs. The way all the cron replicas will communicate with this table is to acquire a lease for the execution. Whoever comes first will register their pod ID as owner in this table. The rest of the replicas will see the combination of owner and expiredAt to identify if someone is running the cron execution or not.
When a pod wants to acquire the execution of a cron, it goes to the table with a row-level exclusive lock and immediately releases it once the update is done on the table. That way, there is no conflict when multiple replicas are trying to acquire the execution.
Heartbeat as a savior!
Along with the lease mechanism, we have a heartbeat. There are mainly 2 reasons for having a heartbeat:
- To build a self-healing system
- We don't know when the cron execution is going to complete because it highly depends on the operation being assigned to a cron job and the resources available to execute it
The responsibility of the heartbeat is to periodically check whether cron is actually executing the task and, if it is, then update the expired time with some margin. But along with this, it is also making sure that the cron server is not stuck at any point in time.
Alternatively, we could have implemented row-based locking instead of having a heartbeat, but there are some serious cases where this mechanism would have failed miserably. Here is how: If a cron replica acquired a lock for execution on the table and it went down, there is no way the next tick would pick up the execution of the cron (because the row is already being acquired and the next replicas are going to skip, assuming there is already execution going on).
Another advantage of having a heartbeat is to identify failure faster so that other replicas can take over the execution. Because the heartbeat is actually updating the expiredAt column, if a replica fails to update it, the next tick will take over the execution.
The Execution Flow
The execution flow is as follows:
- When the cron ticks, the replica is going to ask
cron_with_leasesto acquire an execution lock.- If the lock is acquired, the replica will execute the cron job and update the
expiredAtcolumn through the heartbeat. - If the lock is not acquired, the replica will skip the execution and wait for the next tick.
- If the lock is acquired, the replica will execute the cron job and update the
- If the execution completes before
expiredAt, it will just markexpiredAtas the current time. So that the next cron tick can take over the execution.
Impact
There are some operations that are being executed on the cron server that are resource-heavy and frequent. In certain conditions, we could vertically scale the pod, providing more resources, but at a certain point in time it would not be feasible.
The ideal alternative approach that we can take is horizontal scaling and ensuring idempotent execution and potentially avoiding duplicate execution. This mechanism would be really helpful while doing that. Before this task, we had a strict replica count of 1 on the cron server, but now we can auto-scale with HPA enabled.