Establish a recent-heartbeat baseline.
Follow increasing sequence numbers under epoch 1. Both policies should accept the same envelopes. A heartbeat reports recent accepted communication, not completed work or overall health.
16 / ROS 2 · HEARTBEATS · PROCESS INCARNATIONS
A ROS 2 process sends a heartbeat with its position every 100 ms. Interrupt it, restart it under the same logical identity, then watch two policies inspect the exact same callbacks. A reset sequence number can keep a running process looking silent.
Predict: can an observer tell a crashed process from a running process that stops publishing? What changes if a restart carries a new epoch number?
After the first accepted heartbeat, 400 ms without another produces suspect. An accepted heartbeat restores recent. Suspicion is an observer belief, not proof that a process died.
The baseline requires an increasing sequence number. The incarnation-aware rule also accepts a newer epoch, allowing the new process to begin again at sequence 0.
Separate ROS 2 processes on one host. One observer compares both policies on each callback. Each policy owns its own last-accepted receipt time and watchdog.
The runner records publication silence or a SIGKILL followed by a new process. The synthetic 3D reference continues; this is a software failure, with no aircraft crash or flight physics.
Information boundary: the observer receives heartbeat envelopes. It does not know the runner’s kill, launch or process-exit events. The interface shows those events separately as evaluator truth. Epochs are assigned by this trusted local runner; they are not a general identity or security protocol.
ONE CALLBACK STREAM · TWO ADMISSION POLICIES
Playback moves a cursor through real recorded callbacks and watchdog decisions. It does not start or kill processes in this page.
Drag to orbit · scroll to zoom. The reference continues through software interruptions. The ghost changes only on an accepted heartbeat; no physical crash is simulated.
Observer decisions and runner events are recorded separately.
IDENTICAL CALLBACKS · DIFFERENT ACCEPTANCE RULES
Only an accepted heartbeat refreshes that policy’s watchdog.
OBSERVED CALLBACK / MESSAGE ADMISSION
| Epoch / seq | Receipt / s | Selected policy |
|---|
OBSERVER MEMORY / LOCAL WATCHDOG
watchdog age = cursor − last accepted callback
retained sample age = cursor − sample generation
Ignored callbacks update raw receipt history, but neither the retained position nor that policy’s watchdog.
State segments show recorded observer transitions up to the cursor. The runner’s process state is separate. A launch does not imply immediate DDS discovery or a received heartbeat.
RUNNER TRUTH / NOT AVAILABLE TO THE OBSERVER
A living but silent publisher and an absent process can both make the observer suspect A1. PID changes and harness actions never reset the observer’s memory.
PIDs belong to this saved run. Imported metadata is untrusted provenance: structural validation cannot authenticate that ROS produced a file.
PREDICT · INSPECT · EXPLAIN
Follow increasing sequence numbers under epoch 1. Both policies should accept the same envelopes. A heartbeat reports recent accepted communication, not completed work or overall health.
The same process temporarily stops publishing. Watch it become suspect, then recover under both rules without changing its PID or epoch. Missing heartbeats alone do not reveal the cause.
The replacement keeps logical identity A1 but receives a new epoch. Compare the first returning callback: why does one rule ignore a message it has just received?
FAILURE DETECTION / INCARNATION IDENTITY
A heartbeat failure detector raises suspicion when accepted messages do not arrive within a deadline. With a fixed timeout, a long pause or delayed communication can look like a process failure. This experiment intentionally includes publication silence as that counterexample.
The epoch identifies one incarnation of a logical agent. A sequence number orders messages within that incarnation. These rules are a small application policy; they do not implement SWIM, leader election, consensus or distributed recovery.
Accepting an epoch-2 heartbeat restores this observer’s recent-heartbeat state. It does not restore a task, controller, estimator, map or mission state. The reference curve keeps moving even while no process is publishing it.
All processes share one host and a monotonic clock. Watchdog decisions need only local elapsed receipt time. Reported callback ages and process timing are measurements of these local runs, not guarantees for a wireless network.
With Docker available, run npm run record:restart, then import local/ros2-restart.json. The runner controls disposable publisher processes and records the observer. The repository guide docs/lessons/16-process-restart.md specifies the experiment.
Primary references: Chandra and Toueg, unreliable failure detectors; Aguilera, Chen and Toueg, crash-recovery; ROS 2 Jazzy nodes and Quality of Service settings. The detector shown here is a bounded fixed-timeout implementation, not a proof of a failure-detector class under arbitrary delays.