← Dev Logs

Scaling a Cron Without Running It Twice: A Lease-Based Distributed Lock

scalingarchitecture

Problem & Background

The architecture of the system is designed in a way that the main server is not responsible for handling cron jobs and it's execution. There is a separate pod, which is nothing but a replica of the API, that handles the entire cron and its execution. There are two identical copies of pods:

  1. API: which is responsible for handling all the API requests and a few other things.
  2. Cron: which is responsible for handling cron jobs and their execution.

The system is deployed on Kubernetes with HPA (Horizontal Pod Autoscaler) enabled. The main API server has HPA enabled so it can horizontally scale based on the threshold limits, but the cron has strictly one replica. That means it cannot horizontally scale. One of the main reasons we can't enable HPA directly on the cron is this: let's say if we have multiple replicas; the same operation will be executed multiple times because multiple schedulers will be running on each replica, and that leads to duplicate execution.

Ideally, the execution we want is that there should be a global scheduler, and all the replicas should be more like workers. So the cron should be scaled horizontally, and each replica should be responsible for executing the cron.

Approach

There are ideally two things that we could have done:

  1. Using distributed locking to ensure idempotent execution
  2. Delegating the scheduler outside of cron to a queue-based mechanism that our infra may support and then using each cron replica as a worker polling from the queue; that way, our workers are scalable based on KEDA, not based on HPA.

We chose option one, considering some trade-offs and the situation of the project. Although a better and long-term solution is still option 2, the reason behind that is: We are only worrying about the execution since we are delegating the scheduling part to a queue-based mechanism. (The queue-based mechanism in the project is a self-hosted open-source tool)

Implementation

We implemented a database table called cron_with_leases with:

  • jobId, primary key, storing the unique cron name
  • owner, stores pod ID (ID generated at the time of pod creation)
  • expiredAt, stores the timestamp when the lease expires
  • acquiredAt, stores the timestamp when the lease was acquired

This table will have as many records as the total number of cron jobs. The way all the cron replicas will communicate with this table is to acquire a lease for the execution. Whoever comes first will register their pod ID as owner in this table. The rest of the replicas will see the combination of owner and expiredAt to identify if someone is running the cron execution or not.

When a pod wants to acquire the execution of a cron, it goes to the table with a row-level exclusive lock and immediately releases it once the update is done on the table. That way, there is no conflict when multiple replicas are trying to acquire the execution.

cron replicacron replicacron replicacron replicacron_with_leases
All four ask at once. One pod id lands in the row, and only that replica gets an answer.

Heartbeat as a savior!

Along with the lease mechanism, we have a heartbeat. There are mainly 2 reasons for having a heartbeat:

  1. To build a self-healing system
  2. We don't know when the cron execution is going to complete because it highly depends on the operation being assigned to a cron job and the resources available to execute it

The responsibility of the heartbeat is to periodically check whether cron is actually executing the task and, if it is, then update the expired time with some margin. But along with this, it is also making sure that the cron server is not stuck at any point in time.

Alternatively, we could have implemented row-based locking instead of having a heartbeat, but there are some serious cases where this mechanism would have failed miserably. Here is how: If a cron replica acquired a lock for execution on the table and it went down, there is no way the next tick would pick up the execution of the cron (because the row is already being acquired and the next replicas are going to skip, assuming there is already execution going on).

Another advantage of having a heartbeat is to identify failure faster so that other replicas can take over the execution. Because the heartbeat is actually updating the expiredAt column, if a replica fails to update it, the next tick will take over the execution.

The Execution Flow

The execution flow is as follows:

  • When the cron ticks, the replica is going to ask cron_with_leases to acquire an execution lock.
    • If the lock is acquired, the replica will execute the cron job and update the expiredAt column through the heartbeat.
    • If the lock is not acquired, the replica will skip the execution and wait for the next tick.
  • If the execution completes before expiredAt, it will just mark expiredAt as the current time. So that the next cron tick can take over the execution.
cron ticksask cron_with_leases for the locklock acquired?execute the jobheartbeat extends expiredAton finish, set expiredAt = nowskip this tickwait for the next ticknoyes
One tick. The lease decides whether this replica runs the job or stands down until the next tick.

Impact

There are some operations that are being executed on the cron server that are resource-heavy and frequent. In certain conditions, we could vertically scale the pod, providing more resources, but at a certain point in time it would not be feasible.

The ideal alternative approach that we can take is horizontal scaling and ensuring idempotent execution and potentially avoiding duplicate execution. This mechanism would be really helpful while doing that. Before this task, we had a strict replica count of 1 on the cron server, but now we can auto-scale with HPA enabled.