ARGOS LAB Start with an idea

16 / ROS 2 · HEARTBEATS · PROCESS INCARNATIONS

The process is back.
Does the observer believe it?

A ROS 2 process sends a heartbeat with its position every 100 ms. Interrupt it, restart it under the same logical identity, then watch two policies inspect the exact same callbacks. A reset sequence number can keep a running process looking silent.

Predict: can an observer tell a crashed process from a running process that stops publishing? What changes if a restart carries a new epoch number?

Fixed-timeout failure detection + message admission.

Understand the terms ↗
Failure detector
Fixed-timeout heartbeat detector

After the first accepted heartbeat, 400 ms without another produces suspect. An accepted heartbeat restores recent. Suspicion is an observer belief, not proof that a process died.

Admission policies
Sequence-only vs epoch / sequence

The baseline requires an increasing sequence number. The incarnation-aware rule also accepts a newer epoch, allowing the new process to begin again at sequence 0.

Architecture and timing
One logical agent → one observer

Separate ROS 2 processes on one host. One observer compares both policies on each callback. Each policy owns its own last-accepted receipt time and watchdog.

Execution and fidelity
Actual process runs · recorded replay

The runner records publication silence or a SIGKILL followed by a new process. The synthetic 3D reference continues; this is a software failure, with no aircraft crash or flight physics.

Information boundary: the observer receives heartbeat envelopes. It does not know the runner’s kill, launch or process-exit events. The interface shows those events separately as evaluator truth. Epochs are assigned by this trusted local runner; they are not a general identity or security protocol.

ONE CALLBACK STREAM · TWO ADMISSION POLICIES

Separate process truth from observer belief.

Loading trace…
RECORDED ROS 2 EXECUTIONValidating recorded events…

Playback moves a cursor through real recorded callbacks and watchdog decisions. It does not start or kill processes in this page.

Synthetic reference / process + observer memory
Runner truth: waitingObserver belief: waiting
● Solid drone / evaluator reference◇ Ghost / selected policy’s retained sample→ Spatial difference / evaluator only

Drag to orbit · scroll to zoom. The reference continues through software interruptions. The ghost changes only on an accepted heartbeat; no physical crash is simulated.

Observer decisions and runner events are recorded separately.

IDENTICAL CALLBACKS · DIFFERENT ACCEPTANCE RULES

Select the policy shown by the ghost.

Only an accepted heartbeat refreshes that policy’s watchdog.

OBSERVED CALLBACK / MESSAGE ADMISSION

Inspect the last heartbeat.

Recorded callback envelope

Recent callbacks, including ignored messages

Both policies see every listed callback. Select an envelope to seek to its actual receipt time. Future callbacks stay hidden.
Epoch / seqReceipt / sSelected policy

OBSERVER MEMORY / LOCAL WATCHDOG

Received is different from accepted.

Watchdog age uses local receipt time.

watchdog age = cursor − last accepted callback
retained sample age = cursor − sample generation
Ignored callbacks update raw receipt history, but neither the retained position nor that policy’s watchdog.

Recorded monitor states

State segments show recorded observer transitions up to the cursor. The runner’s process state is separate. A launch does not imply immediate DDS discovery or a received heartbeat.

RUNNER TRUTH / NOT AVAILABLE TO THE OBSERVER

The harness knows which process exists.

Observed events up to the cursor

    A living but silent publisher and an absent process can both make the observer suspect A1. PID changes and harness actions never reset the observer’s memory.

    Recording provenance, process identities and runtime

    PIDs belong to this saved run. Imported metadata is untrusted provenance: structural validation cannot authenticate that ROS produced a file.

    PREDICT · INSPECT · EXPLAIN

    One name can outlive one process.

    01 / NOMINAL

    Establish a recent-heartbeat baseline.

    Follow increasing sequence numbers under epoch 1. Both policies should accept the same envelopes. A heartbeat reports recent accepted communication, not completed work or overall health.

    02 / PUBLICATION SILENCE

    Suspected while still running.

    The same process temporarily stops publishing. Watch it become suspect, then recover under both rules without changing its PID or epoch. Missing heartbeats alone do not reveal the cause.

    03 / RESTART

    Sequence 0, a second time.

    The replacement keeps logical identity A1 but receives a new epoch. Compare the first returning callback: why does one rule ignore a message it has just received?

    FAILURE DETECTION / INCARNATION IDENTITY

    A timeout raises suspicion.
    An epoch names a new start.

    A heartbeat failure detector raises suspicion when accepted messages do not arrive within a deadline. With a fixed timeout, a long pause or delayed communication can look like a process failure. This experiment intentionally includes publication silence as that counterexample.

    The epoch identifies one incarnation of a logical agent. A sequence number orders messages within that incarnation. These rules are a small application policy; they do not implement SWIM, leader election, consensus or distributed recovery.

    sequence-only: accept if s > smax
    epoch-aware: accept if e > emax
    or if e = emax and s > smax
    suspect when t − tlast accepted receipt ≥ 400 ms
    Logical identity / A1
    A stable application name. Reusing the ROS node name or logical agent name does not preserve process memory, sequence counters, or distributed membership.
    Epoch / e
    A strictly newer incarnation marker, assigned by the runner here. On a higher epoch, the observer replaces its old sequence high-water mark. A lower epoch cannot resurrect old state.
    Sequence / s
    A counter starting at 0 in each process. The baseline ignores the epoch and initially rejects restarted messages until their sequence exceeds its pre-interruption high-water mark.
    Fixed timeout / 400 ms
    Only accepted messages reset each policy’s local receipt watchdog. Callback arrival and application acceptance are separate events. Watchdog checks run periodically, so recorded suspicion includes actual scheduling delay.

    Software state is not mission recovery.

    Accepting an epoch-2 heartbeat restores this observer’s recent-heartbeat state. It does not restore a task, controller, estimator, map or mission state. The reference curve keeps moving even while no process is publishing it.

    Timing has a declared scope.

    All processes share one host and a monotonic clock. Watchdog decisions need only local elapsed receipt time. Reported callback ages and process timing are measurements of these local runs, not guarantees for a wireless network.

    Record another actual ROS 2 run

    With Docker available, run npm run record:restart, then import local/ros2-restart.json. The runner controls disposable publisher processes and records the observer. The repository guide docs/lessons/16-process-restart.md specifies the experiment.

    Primary references: Chandra and Toueg, unreliable failure detectors; Aguilera, Chen and Toueg, crash-recovery; ROS 2 Jazzy nodes and Quality of Service settings. The detector shown here is a bounded fixed-timeout implementation, not a proof of a failure-detector class under arbitrary delays.