Supervision and Fault Isolation
Overview
Section titled “Overview”PCP’s primary fault isolation strategy is process boundaries. The daemon, the compositor, and every Tier 1 application run in separate processes. A crash in one does not crash the others. The protocol does not rely on in-process supervisors, watchdog timers, or panic handlers for its main isolation guarantees. Those mechanisms exist, but they serve a different deployment mode.
This page covers two topics: how the primary model isolates failures by process boundary, and how the supervision library provides analogous guarantees for hosts that use the in-process embedding mode.
Process-Boundary Isolation
Section titled “Process-Boundary Isolation”The deployed PCP topology puts three classes of process on the system, each with a distinct failure domain:
| Process | Failure domain | What breaks when it crashes |
|---|---|---|
| Compositor | Surface authority | The desktop stops rendering. PCP is unaffected. |
| PCP Daemon | Capability authority | Clients lose capability access until the daemon restarts. The desktop keeps running. |
| Tier 1 apps | Application capability servers | Specific app capabilities become unavailable. The daemon detects the failure and degrades the app’s registration. |
These domains do not overlap. The compositor loads no PCP code, so a daemon bug cannot propagate into the rendering pipeline. The daemon loads no compositor code, so a rendering bug cannot corrupt the registry. Tier 1 apps load neither daemon nor compositor code, so failures stay contained within the app.
Daemon Failure
Section titled “Daemon Failure”If the PCP daemon crashes, clients lose their connection. System Intelligence stops receiving events and cannot invoke capabilities. CLI tools get connection errors. Tier 1 app servers lose their client but remain running.
Recovery is straightforward. The daemon restarts, rebinds its socket, and rebuilds the registry from static detection. The detection pipeline scans for running applications, classifies them, and repopulates the registry. The warm-start budget is defined in the performance SLOs: the daemon must be ready to accept connections within a bounded time after restart.
During restart, the compositor continues rendering. Applications continue running. The user’s desktop is fully interactive. The only visible effect is that System Intelligence and PCP-dependent tooling are temporarily unavailable.
Tier 1 Application Failure
Section titled “Tier 1 Application Failure”When a Tier 1 app crashes, its socket disappears. The daemon’s connection manager detects the closed socket and surfaces an AppNotReachable error for any pending or future invocations against that app.
The recovery runtime monitors repeated failures. An app that crashes repeatedly is degraded: its tier registration falls back to a lower tier (typically Tier 2 via AT-SPI2), and the daemon stops attempting Tier 1 connections until the app restarts successfully.
Other apps, the daemon, and the compositor are unaffected. The failure is scoped to the one application.
Compositor Failure
Section titled “Compositor Failure”The compositor is outside PCP’s blast radius entirely. It runs no PCP code, uses no PCP libraries, and connects to no PCP sockets. A compositor crash takes down the desktop but has no effect on the daemon or its internal state. When the compositor restarts, the daemon resumes serving clients.
Supervision Library
Section titled “Supervision Library”The protocol defines an in-process embedding mode for hosts that own their servers. In this mode, the daemon’s role (or part of it) runs inside a larger process rather than as a standalone service. Process boundaries no longer provide isolation, so the host needs another mechanism.
The core library ships a supervision module for this purpose. It provides the machinery that an in-process host needs to detect and recover from faults in its embedded PCP servers. The daemon itself does not use this module. It relies on process-boundary isolation instead.
Available Components
Section titled “Available Components”| Type | Purpose |
|---|---|
SupervisorState |
State machine tracking server health: Running, SuspectedHang, Restarting, Degraded |
SupervisorConfig |
Tunable parameters: heartbeat interval, max missed beats, max restart attempts, backoff base |
SupervisorEvent |
Notifications emitted on state transitions: hang detected, panic caught, restart initiated, degraded entered |
watchdog |
Heartbeat/missed-beat detection. The supervised server must call heartbeat() periodically. Missed beats trigger a hang declaration. |
MemoryBudget |
Byte-accurate allocation tracking. MemoryBudget::new(budget_mib) creates a budget. record_alloc and record_free track usage. is_over_budget and is_alert check thresholds. over_budget_count reports how many times the budget has been exceeded. |
TrackingAllocator |
Global allocation instrumentation. current_allocated() and peak_allocated() report live and peak usage. reset_peak() clears the high-water mark. |
health (HealthReport) |
Supervisor-level health snapshot combining watchdog state, memory usage, and component availability. |
PcpComponent::all() |
Enumerated list of budgeted components for consistent tracking across the supervised subsystem. |
How It Works
Section titled “How It Works”The supervision module wraps the in-process servers the same way the standalone daemon wraps the entire PCP role: with a state machine, a heartbeat contract, and a memory cap.
The supervised server calls heartbeat() on a regular interval. The watchdog monitors these calls and tracks missed beats. If the server misses too many beats in a row, the supervisor declares a hang and initiates a restart sequence: drop the server instance, create a new one, and resume normal operation.
Memory budget tracking prevents a runaway server from exhausting the host’s memory. Every allocation goes through record_alloc, and every deallocation through record_free. If the total exceeds the configured budget, new allocations fail with a budget-exceeded error. The server must handle this gracefully, typically by falling back to a lighter code path or returning an error to the caller.
The supervisor state machine follows a linear recovery path: Running to SuspectedHang to Restarting and back to Running. If restart fails repeatedly, the state enters Degraded and persists until the next retry opportunity. The host is notified through SupervisorEvent emissions at each transition.
Relationship to the Daemon
Section titled “Relationship to the Daemon”The daemon does not use the supervision module. It does not need it. The daemon runs as a standalone process, and process boundaries provide its isolation. If the daemon crashes, the OS restarts it. If a Tier 1 app crashes, the daemon detects the socket closure and degrades the app. No watchdog, no memory budget, no panic handler is needed inside the daemon for these scenarios.
The supervision library exists for a specific audience: hosts that embed PCP servers in-process and need the same quality of fault detection that process boundaries provide for free. A future native compositor host that embeds the entire PCP role would be the primary consumer.
Configuration
Section titled “Configuration”pub struct SupervisorConfig { /// Watchdog heartbeat interval. pub heartbeat_interval: Duration, // e.g., 2s /// Max missed heartbeats before hang declared. pub max_missed_heartbeats: u32, // e.g., 3 /// Max restart attempts before entering degraded mode. pub max_restart_attempts: u32, // e.g., 3 /// Delay between restart attempts (exponential backoff). pub restart_backoff_base: Duration, // e.g., 1s /// Total memory budget for supervised servers. pub memory_budget_bytes: usize, // e.g., 128 MiB /// Memory alert threshold (fraction of budget). pub memory_alert_threshold: f64, // e.g., 0.80}These defaults are generous for typical PCP workloads. The memory budget caps allocation well below what would threaten host stability. The heartbeat interval is long enough to tolerate brief pauses but short enough to detect genuine hangs within seconds. Exponential backoff prevents tight restart loops from thrashing the system.