The Node's maintenance cron
A visit that ended has to start reading as ended. Until that happens the session is still open in the database: the day's session count is too high, average duration is wrong, and the 7-day rollup is computed over rows that are still going to change. None of that is a lost event — it is a number that has not settled yet.
The work that settles it is the Node's maintenance: a handful of short tasks that run on a seconds-scale interval, off the ingestion path. They have two possible clocks, and a Node is healthy with either one.
This page is for someone running their own Node, with shell access to the server, deciding whether a crontab line is worth adding.
What a maintenance job is
A maintenance job is a task that belongs to no request. It does not serve a visitor and does not answer an event: it tidies up the state that requests left behind. The Node has three, and each one handles a different problem.
| Job | What it does | Cadence |
|---|---|---|
session_sweep |
Closes sessions that have been inactive long enough to count as finished. | every 60 s |
session_reclaim |
Recovers sessions that started closing and stalled halfway — a process that died, a deploy at the wrong instant. The session goes back in the queue instead of staying stuck. | every 60 s |
session_reconsolidate |
Recounts a session that was already closed when something arrived late and flagged it for a recount. | every 300 s |
Each job is independent. It is triggered on its own, with its own deadline, its own batch and its own health record.
Two clocks, and the Node is healthy with both
The Node triggers maintenance in two ways:
- Traffic tick (
traffic_tick) — on by default, with no configuration at all. After ingesting an event, the Node checks whether any job is due and, if one is, runs it. This is how a freshly installed Node maintains itself from the first minute. - Local cron (
local_cron) — an authenticated call that your crontab makes, at whatever pace you set.
A Node whose crontab is slower than the interval — or that has no crontab at all — falls back to the traffic tick, and that is the design working, not a degradation.
Said plainly:
- a Node with no crontab at all still maintains itself, using visitor traffic as its clock;
- a slower crontab is not a misconfiguration — it just means more of the work rides on traffic;
- the cron exists so that a low-traffic Node does not have to wait for a visitor to show up.
The crontab is a recommendation because a timer is steadier than traffic on a quiet Node, not because the Node needs one to be correct.
What the cron improves — and what it does not
The cron does not make a session close any sooner.
A session only becomes eligible to close 1920 seconds (32 minutes) after its last activity: 1800 s of inactivity plus a 120 s grace. That floor belongs to the session model, not to the scheduler, and no crontab shortens it.
What changes is what happens after. A session becomes eligible at some arbitrary instant; what closes it is the next maintenance trigger. With the old 1800 s interval, a session that became eligible right after a pass had to wait for the following one — closing landed somewhere between roughly 32 and 62 minutes, and you had no way to know where. With a 60 s interval, that same session closes between 32 and 33 minutes.
That is a gain in predictability, not in speed. The floor is unchanged; what disappeared is the unpredictable wait sitting on top of it.
The recommended crontab
One line per minute, one line per job:
* * * * * curl -sS -X POST "https://YOUR-NODE/public/local-maintenance.php" -H "X-Master-Key: YOUR_MASTER_KEY" -H "Content-Type: application/json" -d '{"job":"session_sweep"}' > /dev/null
* * * * * curl -sS -X POST "https://YOUR-NODE/public/local-maintenance.php" -H "X-Master-Key: YOUR_MASTER_KEY" -H "Content-Type: application/json" -d '{"job":"session_reclaim"}' > /dev/null
*/5 * * * * curl -sS -X POST "https://YOUR-NODE/public/local-maintenance.php" -H "X-Master-Key: YOUR_MASTER_KEY" -H "Content-Type: application/json" -d '{"job":"session_reconsolidate"}' > /dev/null
Replace YOUR-NODE with your Node's domain and YOUR_MASTER_KEY with the MASTER_KEY value from its .env file — the same key you use to log into the Console.
Three rules that are not optional:
- The Master Key travels in the
X-Master-Keyheader — never in the URL, never in the body. A URL shows up in server logs, proxy logs and shell history; a header does not. - The body is exactly
{"job":"<name>"}. No other key is accepted: a body with extra fields is rejected. - One job per call. That is why there are three lines and not one.
* * * * * is the recommended cadence for the two 60 s jobs. session_reconsolidate runs every 300 s, so */5 * * * * is enough — calling it every minute does no harm, it simply answers not_due.
If your Node is installed so that the domain root already points at the public folder,
https://YOUR-NODE/local-maintenance.phpanswers as well. Both forms reach the same place; use whichever returns JSON instead of a404.
The endpoint contract
| Situation | Response |
|---|---|
| Correct call | 200 with the job envelope |
Any method other than POST |
405 · {"error":"method_not_allowed"} |
| Missing or wrong key | 403 · {"error":"forbidden_invalid_key"} |
| Malformed body, or unknown job | 422 · {"error":"job_not_registered"} |
| The job ran and failed | 500 with the envelope, "status":"failed" |
A successful call returns an envelope shaped like this:
{
"job": "session_sweep",
"status": "ran",
"trigger": "local_cron",
"claimed_at": 1757500000,
"next_eligible_at": 1757500060,
"duration_ms": 41,
"result": {
"claimed": 12,
"closed": 12,
"reconsolidated": 0,
"conflicts": 0,
"errors": 0,
"had_more": false
}
}
status tells you what happened:
status |
Meaning |
|---|---|
ran |
The job ran. result carries the numbers. |
not_due |
Not due yet. This is the normal answer on most minutes. |
deferred |
Another trigger got to the job first. Nothing to do. |
skipped_budget |
Less time was left than the smallest useful slice of work. The job returns on the next trigger. |
failed |
The job raised an error. Comes with 500. |
A not_due is not an error and should not alert on anything. A one-minute cron against a 60 s job spends most of its time answering exactly that.
Budgets: what the Node spends per trigger
Every maintenance pass is bounded by a time budget as well as a batch size. The budget depends on who is paying for it:
| Trigger | Budget | Who waits |
|---|---|---|
| the cron | 3000 ms | nobody — no visitor is attached |
| traffic, response already released (PHP-FPM) | 800 ms | nobody — the page has already been sent |
| traffic, response not yet released | 120 ms | the visitor, before the page goes out |
No single visitor ever pays more than the budget. When there is a backlog, the Node spends more triggers, not a longer one: the pass stops at its deadline, records that work was still pending, and re-arms the job to come back in seconds instead of waiting out the full interval.
And when nothing is due there is no meaningful cost: a request with no pending work costs one filesystem read and zero database queries.
Batches, validation, and what does not exist
Each job works in batches of up to 100 sessions per pass, and no pass ever exceeds the hard ceiling of 500, whatever value is asked for. Those two numbers, and the cadences in the table above, are fixed in the Node's code — there is nothing there for you to tune.
The three budgets accept an override in the Node's .env — ST_MAINT_BUDGET_LOCAL_CRON_MS, ST_MAINT_BUDGET_TRAFFIC_DETACHED_MS and ST_MAINT_BUDGET_TRAFFIC_INLINE_MS. Each has its own permitted range, and validation is explicit: a value that is not an integer, or that falls outside the range, goes back to the default instead of being applied. The Node's health report shows both side by side — what you wrote (configured), what is in force (effective) and why it was refused (fallback) — so an ignored setting never passes for an applied one.
There is no configuration UI and no remote configuration. None of this appears in the Console, and Supreme does not change these values on your Node. Cadence and batch size are fixed in the Node's code; the budgets live in your server's .env and nowhere else. If you went looking for the screen, it does not exist — this is a decision, not a missing screen.
Diagnostics: the maintenance health block
The authenticated health endpoint carries a maintenance block with per-job state:
curl -s "https://YOUR-NODE/public/api/health.php" -H "X-Master-Key: YOUR_MASTER_KEY"
The useful way to read it is to separate two questions that look like the same one:
1. What this server can do — probed at the moment of the call, under maintenance.runtime:
| Field | What it tells you |
|---|---|
sapi |
How PHP is running. fpm-fcgi is the only shape that can release the response before maintenance runs. |
terminator |
The function that releases the response early, when one exists. |
terminator_internal |
Whether that function is PHP's native one rather than a stand-in declared by other code. |
2. What actually happened last time — recorded, under maintenance.jobs.<job>.last:
| Field | What it tells you |
|---|---|
last_status |
ran, not_due, deferred, skipped_budget, failed — or never_run. |
last_trigger |
local_cron or traffic_tick: what triggered it. |
last_response_clock |
detached (the response had already gone out) or inline_fallback (it had not). |
last_response_reason |
Why that clock and not the other one. |
last_duration_ms |
How long the pass took. |
claimed / errors |
How many sessions the pass picked up, and how many errors occurred. |
They are kept apart on purpose. A server that can release the response early but whose last pass ran inline anyway is a real and normal state — it is what happens when the cron was the trigger, because there is no visitor waiting on a cron call. In that case last_response_reason reads local_cron_sync, and there is nothing to fix. Judging capability by the last record sends you hunting for a problem that is not there.
If a job never runs
last_status stuck at never_run, or a stale last_claimed_at, has a short list of causes:
- The job is switched off. Check
maintenance.jobs.<job>.enabled. Afalsecomes from an.envvariable —SESSION_SWEEP_JOB_ENABLED,SESSION_RECLAIM_JOB_ENABLEDorSESSION_RECONSOLIDATE_JOB_ENABLED— set to0. Absent means on. - There is no crontab and no traffic. Both clocks are stopped at once. A Node with no visitors and no crontab has nothing to trigger it. Call the endpoint once by hand and read the
statusthat comes back. - The marker file is not writable. Check
maintenance.jobs.<job>.pre_gate.status. Anunavailablemeans the Node cannot write tostorage/cache/— fix the directory permission and the job resumes on the next trigger.
The SESSION_SWEEP_TICK_ENABLED switch
This .env variable turns the traffic tick off entirely:
- absent — on. This is the default, and what you want in most cases.
0— off. No visitor request triggers any maintenance.
The cron keeps working when it is off. Turning the tick off and keeping the crontab is a legitimate setup: all maintenance moves onto the timer, and no visitor request pays for any of it. Turning the tick off without a crontab, on the other hand, leaves the Node with no clock at all — and then sessions really do stop closing.
maintenance.tick_enabled shows which of the two states the Node is in.