Controller lifecycle
The controller drives every Agent from creation to a terminal phase by reconciling its status
against the world. Reconciliation is a pure function with all I/O injected, which keeps it
unit-testable.
State machine
Section titled “State machine”stateDiagram-v2
[*] --> Pending : Agent created
Pending --> Running : Job created, status patched
Pending --> Failed : Station or AgentDefinition missing
Running --> Running : Job still running
Running --> Succeeded : Job exit 0
Running --> Failed : Job failed
Succeeded --> [*]
Failed --> [*]
note right of Running : on a terminal transition, prune old runs per history limits
Transitions
Section titled “Transitions”Pending → Running
- Resolve
Stationbyspec.stationRef; if missing, setphase = Failedwith a reason and stop. - Resolve
AgentDefinitionbystation.spec.agentDefRef; if missing, fail the same way. - Build the Job (see Agent runtime) and create it. Creation is idempotent. An already-existing Job is fine.
- Patch
status:phase = Running,jobName,startedAt.
Running → Succeeded | Failed
- Read the Job outcome. While it is still running, do nothing.
- On success, patch
phase = SucceededwithexitCodeandoutput(the truncated tail of the pod log). - On failure, patch
phase = FailedwithexitCode,failureReason, andoutput(the truncated tail). - After any terminal transition, prune history.
Every status patch carries the Agent’s resourceVersion as a precondition, so a write computed from
a stale read is rejected (409 Conflict) instead of clobbering a newer update. The watch and the poll
reconcile from independent snapshots, so this optimistic-concurrency guard is what keeps them from
lost-updating each other: a conflicted write is dropped and the next reconcile recomputes from fresh
state.
If the run pod was already garbage-collected when the controller reads back (so its captured stdout
is gone), reading it back is best-effort: the Agent still reaches its terminal phase — keeping any
Job-level failure reason, or recording run output unavailable: pod garbage-collected when there is
none — instead of leaving output silently empty or letting the read error leave the Agent stuck in
Running. If the Job itself was garbage-collected first (the controller was down or backlogged
past ttlSecondsAfterFinished), the outcome is unrecoverable, so the Agent is reported Failed with
run record unavailable: Job garbage-collected before its result was observed rather than being left
stuck in Running.
exitCode always agrees with the phase: a Failed run reports a non-zero code even when the pod’s
container status reads back as 0 (a GC race can leave it unpopulated), and a Succeeded run
reports 0.
Terminal → terminal is a no-op; reconciling a finished Agent does nothing.
Launch sequence
Section titled “Launch sequence”sequenceDiagram
autonumber
participant Caller
participant API as Kubernetes API
participant Ctrl as Controller
participant Job
participant Pod
participant Sink as Output sink
Caller->>API: create Agent (phase Pending)
API-->>Ctrl: watch event (ADDED)
Ctrl->>API: get Station + AgentDefinition
Ctrl->>API: create Job (owner ref to Agent)
Ctrl->>API: patch status = Running
Job->>Pod: schedule pod
Pod->>Pod: initContainer injects bundle
Pod->>Pod: supervisor runs agent process
Pod->>Sink: stream-json events
Ctrl->>API: read Job outcome
Ctrl->>API: patch status = Succeeded / Failed
Caller->>API: watchAgent polls status until terminal
Job ownership
Section titled “Job ownership”Each Job is created with an owner reference back to its Agent (controller: true,
blockOwnerDeletion: true). Deleting the Agent garbage-collects the Job. Jobs also set
ttlSecondsAfterFinished: 3600, so finished Jobs (and their pods) self-delete one hour after
completion while the Agent’s status retains the result. This is the window the controller has to
read the pod’s exit code and captured stdout back into status; one hour comfortably survives a
controller restart or backlog. The trade-off is that finished pods linger that long, but they are
terminal (only an etcd object plus node log disk, no CPU/memory), and history
pruning already cascade-deletes most of them sooner. Past the window the run is
reported terminal with a clear failureReason rather than silently losing its output.
History pruning
Section titled “History pruning”On every terminal transition the controller groups its cached Agents for the Station by phase, sorts
by completedAt (newest first), and deletes any beyond the Station’s
successfulRunsHistoryLimit / failedRunsHistoryLimit. This bounds how many finished Agents
accumulate without losing the most recent ones.
Watch + poll + cache
Section titled “Watch + poll + cache”The controller combines a low-latency watch with a ~15s poll, sharing one in-memory cache of
Agents. The watch seeds the cache and its starting resourceVersion with a full, paginated LIST
(?limit=&continue=), then resumes from that resourceVersion, applying each event to the cache
(ADDED/MODIFIED upsert, DELETED evict) and reconciling the changed Agent. When the watch
closes it resumes from the last resourceVersion seen — it does not replay the whole collection. A
410 Gone (the change history was compacted past our cursor) triggers a fresh paginated re-list. The
watch carries a server-side timeoutSeconds, so the API server periodically ends the long poll and
the loop reconnects — a silently dropped (half-open) connection can’t leave it blocked on a dead
socket. On a real error (API blip, TLS, 5xx) the loop backs off exponentially with jitter rather
than retrying every couple of seconds, so a degraded API server isn’t hammered in lockstep.
The poll runs independently every ~15s: a full paginated LIST that refreshes the cache and
reconciles every Agent. This is the safety net — it guarantees an Agent whose watch event was missed
or dropped is still reconciled, since the watch is the fast path but not a guarantee. Because
concurrency counts and history pruning read the cache rather than doing their own LIST per
reconcile, reconcile work is O(changed) even though the poll itself lists. The /metrics endpoint
exposes controller_resyncs_total (full LISTs) and controller_watch_reconnects_total.
Leader election
Section titled “Leader election”The controller runs two replicas for availability, but only one reconciles at a time. The
replicas contend for a single coordination.k8s.io/v1 Lease named
agent-controller; the holder is the leader and runs the watch + poll loop, while standbys stay
idle.
stateDiagram-v2
[*] --> Standby
Standby --> Leader : Lease absent, or held by us, or expired and we take it over
Leader --> Leader : renew the Lease every ~5s
Leader --> Standby : renewal lost (another replica wrote, or the API is unreachable)
note right of Leader : only the leader reconciles
Each tick (~5s) a replica reads the Lease, folds it into a running observation, and decides:
- no Lease → create one we hold;
- we hold it → renew its
renewTime; - someone else holds it, unrenewed past the 15s lease duration → take it over;
- someone else holds a still-valid Lease → stand by.
Expiry is judged against the replica’s own clock — the time since it first saw the current
renewTime — not the remote timestamp, so clock skew between replicas can’t trigger a premature
takeover. Writes carry the Lease’s resourceVersion as an optimistic-concurrency precondition, so
two standbys can’t both win a takeover and a leader that fails to renew (lost the race, or lost
the API server) immediately steps down and stops reconciling.
When the leader’s pod is deleted it stops renewing; a standby sees the renewTime stop changing and
takes over within the lease duration. The brief overlap a partition could cause is harmless anyway:
Job creation is keyed on the Agent name and idempotent (an existing Job is a 409, not a duplicate),
so two reconcilers never produce two Jobs for one Agent.