blob: bb1f76f645fbed2cbc31828f9ceed91f16f5ece8 [file] [view]
# `fdomain-uart-driver` Target Transport Bridge Design
## 1. Overview
`fdomain-uart-driver` is the target-side counterpart to the host `ffx` UART
transport daemon (`//src/developer/ffx/tools/uart_driver`). It runs on the
Fuchsia target device and bridges a physical or virtual serial port
(`/dev/class/serial`) to the Remote Control Service
(`fuchsia.developer.remotecontrol.connector/Connector`).
The wire framing, CRC-8 header and CRC-32 payload checksum verification,
handshake state machine, and sliding-window Go-Back-N (`ProtocolId::ResendSP`)
reliability layer are shared with the host via `//src/developer/lib/uart_fpl`
and documented in [`//src/developer/ffx/tools/uart_driver/DESIGN.md`][host-design].
This document describes the target component's lifecycle, task pipeline, and
target-specific architectural choices.
---
## 2. Lifecycle & Task Architecture
### 2.1 Component Startup & Session Loop
1. **Boot Argument Gate**: On startup, `run_driver` queries
`fuchsia.boot/Arguments` for `dev.fdomain.uart` (default `false`). When
disabled, the component parks forever (`futures::future::pending()`) with
zero CPU or serial resource usage. When enabled, it reads
`dev.fdomain.uart.baud` (default `1000000`).
2. **Serial Discovery**: `open_serial_device` scans `/dev/class/serial` in
lexicographical order, selects a host-facing UART, and configures `8N1` with
`FlowControl::None` at the target baud rate.
3. **Cold-Start Reset Notification**: On its very first session negotiation
after component start (`is_initial_start`), `negotiate_session` transmits an
unsolicited `FrameType::Reset` frame (`session_id = 0`) before waiting for a
handshake. If the target rebooted while the host daemon had an active
session open, this immediately notifies the host to abandon stale sequence
state and initiate a fresh handshake.
4. **Handshake & Bridge Loop**: Once `run_target_handshake_with_initial_data`
negotiates `ProtocolId::ResendSP` and records the host-chosen `session_id`,
`run_bridge_session` runs four cooperating asynchronous tasks until the
session resets or the serial device errors.
### 2.2 Task Pipeline & Data Flow
```text
fuchsia.hardware.serial/Device (TX) fuchsia.hardware.serial/Device (RX)
▲ │
│ ▼ (SerialReader)
┌──────────────────────┐ ┌──────────────────────┐
│ writer_task │◄── AckTracker ────│ receiver_task │
│ (4 KiB batching) │ (latest-ACK slot) │ (ResendReceiver) │
└──────────────────────┘ └──────────────────────┘
▲ │ │
│ data_tx (bounded: 64) │ │
┌──────────────────────┐ │ │
│ sender_task │◄─── ack_tx (64) ──────┘ │
│ (ResendSender) │ (Cumulative ACKs) │
└──────────────────────┘ │
▲ │
│ sender_tx (bounded: 64) │
┌──────────────────────┐ │
│ coordinator_task │◄──────── serial_tx (bounded: 64) ────┘
│ (CoordinatorState) │ (In-order DATA / CLOSE)
└──────────────────────┘
│ ▲ ▲
│ │ └── internal_tx (bounded: 64)
│ │ (Registered / RegistrationFailed / WriterDone)
│ │
│ └── client_tx (bounded: 64, gated by total_queued < 64)
│
▼ writer_tx (bounded: 256)
channel_writer_task / channel_reader_task (per active channel_id)
│ ▲
▼ │
Zircon Stream Socket (RCS Connector.FdomainToolboxSocket)
```
The crate separates low-level serial I/O (`src/serial.rs`) from the Go-Back-N
`ResendSP` state machines (`src/receiver.rs` and `src/sender.rs`):
* **Why there is a `writer_task` on TX, but a `SerialReader` struct (not a
`reader_task`) on RX**:
* **RX (`SerialReader` $\rightarrow$ `receiver_task`)**: `receiver_task` is
the sole consumer of incoming serial bytes. Rather than spawning a separate
`reader_task` connected by an extra MPSC channel, `receiver_task` borrows
`&mut SerialReader` directly and calls `reader.next_frame().await`. When a
session terminates (for example, on `TargetDriverError::SessionIdChanged`),
`run_bridge_session` immediately calls `reader.take_unconsumed()` on that
`SerialReader` to recover any trailing bytes buffered in its `FrameParser`
for the next session's handshake.
* **TX (`sender_task` + `AckTracker` $\rightarrow$ `writer_task`)**: Multiple
producers emit frames to the serial device concurrently (`sender_task`
sends data and close frames, `receiver_task` sends duplicate
`NegotiateResp` frames via `data_tx`, and `receiver_task` signals priority
cumulative ACKs via `AckTracker`). A dedicated `writer_task` is required to
multiplex `AckTracker` ahead of `data_rx`, coalesce up to 4 KiB of frames
per `Device.Write` call, and retry transient write errors.
---
## 3. Non-Obvious Architectural Choices
### 3.1 Priority ACK Coalescing via `AckTracker`
In a half-automated or full-duplex Go-Back-N link over UART, enqueueing
outgoing `FrameType::Ack` frames into the same FIFO MPSC channel (`data_tx`) as
outgoing `FrameType::Data` frames causes two problems under load:
1. **Head-of-line ACK delay**: Up to 64 outgoing data frames
(`MAX_PAYLOAD_SIZE` each) could sit ahead of an ACK, stalling the host's
transmit window or triggering spurious host Go-Back-N timeouts.
2. **Redundant ACK traffic**: Because `ResendSP` ACKs are cumulative, sending
intermediate ACKs (`seq = 1`, `seq = 2`, ..., `seq = 8`) when `seq = 8` is
already known wastes UART bandwidth.
Instead, `receiver_task` and `writer_task` share an `AckTracker`
(`//src/developer/lib/uart_fpl`):
* `AckTracker` is a single-slot latest-value register (`Option<(u32, u8)>`)
paired with a `futures::task::AtomicWaker`. Calling
`ack_tracker.set_ack(session_id, seq)` overwrites any unsent older ACK in
$O(1)$ space and wakes `writer_task`.
* At the top of every iteration in `writer_task`, `ack_tracker.take_ack()` is
checked **before** polling `data_rx`, and `futures::select_biased!` polls
`ack_tracker.wait_ack()` ahead of `data_rx`.
* When `writer_task` does pull a frame from `data_rx`, `flush_frame_batch`
non-blockingly drains ready frames via `data_rx.try_recv()` up to
`WRITE_BATCH_CAPACITY` (4 KiB) into a single `Device.Write` FIDL call, then
immediately returns to the top of the loop to check `ack_tracker` before
flushing the next batch.
### 3.2 Two-Phase Asynchronous RCS Registration & Per-Channel Generations
When the host opens a new multiplexed channel, it does not send a separate
"channel open" handshake; it immediately sends the first `FrameType::Data`
frame carrying the new `channel_id`. On the target, connecting that channel to
RCS requires an asynchronous FIDL call (`Connector.FdomainToolboxSocket`).
* **Why registration is spawned off-loop (`ChannelState::Pending`)**: Awaiting
`fdomain_toolbox_socket` directly inside `coordinator_task` would stall the
entire multiplexer (blocking all other active channels and incoming serial
frames) while RCS spawns or binds the toolbox connection. Instead,
`open_pending_channel` spawns `spawn_channel_registration` in the background
and places the channel in `ChannelState::Pending`, staging any subsequent
`SerialData` frames for that `channel_id` in an in-memory buffer (capped at
`MAX_CHANNEL_STAGING_BYTES = 2 MiB`).
* **Graceful close-before-connect draining (`closing_writers`)**: Short-lived
channels may receive `SerialData` followed immediately by `SerialClose`
before `fdomain_toolbox_socket` finishes registering, or while staged bytes
are still waiting to be written to the Zircon socket. Rather than dropping
the buffered payload on `SerialClose`, `CoordinatorState` marks the channel
`closing: true`, flushes all staged buffers through `channel_writer_task`,
and holds the writer task in `closing_writers` until
`InternalEvent::WriterDone` confirms the socket has drained.
* **Generation-tagged events (`generation: u64`)**: Because registration and
socket teardown happen asynchronously, a host could close `channel_id = K`
and immediately reuse `channel_id = K` for a new connection while tasks from
the previous instance are still winding down. Every channel instance is
assigned a monotonically increasing `generation: u64`, and all
`InternalEvent` and `ClientEvent` messages carry `(channel_id, generation)`
so `CoordinatorState` ignores stale completions from superseded channel
instances.
### 3.3 Wait-Free Ingress Dispatch & Non-Blocking Channel Staging
To guarantee that a single stalled or slow RCS socket cannot deadlock the UART
link or block control frames (`Ack`, `Reset`, `NegotiateReq`):
1. **`receiver_task` never `.await`s on downstream MPSC channels**: Both
`serial_tx.try_send(event)` and `ack_tx.try_send(seq)` are synchronous. If
`serial_tx` is full because `coordinator_task` is busy, `receiver_task`
drops the incoming `DATA`/`CLOSE` frame **without** advancing
`ResendReceiver` and re-asserts `receiver.current_ack_seq()` on
`ack_tracker`. The host's Go-Back-N window naturally pauses and retransmits
the frame once `serial_tx` drains.
2. **`coordinator_task` never `.await`s on per-channel socket writers**:
`stage_or_send_connected` uses `sender.try_send(data)` on the per-channel
`writer_tx` queue (`CHANNEL_WRITER_QUEUE_CAPACITY = 256`). If a specific RCS
socket is slow and its queue fills, overflow chunks are staged in that
channel's `pending_incoming` queue up to `MAX_CHANNEL_STAGING_BYTES`
(2 MiB) and drained opportunistically via `drain_pending_incoming()` on
every coordinator turn. If a runaway channel exceeds
`MAX_CHANNEL_STAGING_BYTES`, only that channel is dropped and sent a
`SenderMessage::Close`, isolating healthy channels.
### 3.4 Round-Robin Egress Scheduling & Backpressure Gating
* **Fair Egress Multiplexing (`select_next_message`)**: Each active channel's
`channel_reader_task` chunks RCS socket reads into `MAX_PAYLOAD_SIZE` slices
and sends `ClientEvent::Data` to `coordinator_task`, which queues them per
`channel_id` in `CoordinatorState::outgoing`. When `sender_tx` has capacity
(`poll_send_next_queued`), `select_next_message` pops one frame at a time
across sorted `channel_id`s in round-robin order (`last_channel_id`). This
prevents a bulk transfer channel (such as `ffx target snapshot` or `ffx log`)
from monopolizing the Go-Back-N send window.
* **Gated `client_rx` Polling (`next_client_event`)**: When the UART link is
congested or unacknowledged frames fill `ResendSender`'s sliding window,
`sender_task` stops pulling from `sender_rx`. Once the total number of queued
frames across `CoordinatorState::outgoing` reaches
`DEFAULT_CLIENT_DATA_CAPACITY` (64), `next_client_event` returns
`futures::future::pending()`, suspending reads from `client_rx`. This
backpressures all `channel_reader_task` instances, which stop reading from
their Zircon sockets and propagate kernel socket buffer backpressure directly
to the RCS producer components.
### 3.5 Session Renegotiation, Duplicate Handshakes & Byte Carry-Over
* **Preserving Trailing Bytes (`take_unconsumed` & `requeue_frame`)**:
`fuchsia.hardware.serial/Device.Read` returns arbitrary byte chunks from the
driver ring buffer. A single read during handshake may contain the
`NegotiateReq` frame followed immediately by the first `Data` frames of the
new session. `run_target_handshake_with_initial_data` extracts any trailing
bytes from `FrameParser::take_unconsumed()` and seeds them into
`SerialReader::with_initial_data` so no post-handshake frames are lost.
Conversely, if `receiver_task` encounters a `NegotiateReq` with a new
`session_id` (`TargetDriverError::SessionIdChanged`), it puts that frame back
at the front of `SerialReader` via `requeue_frame` before returning to
`run_device_sessions`, allowing `negotiate_session` to process the new
handshake immediately without waiting for the host to time out and
retransmit.
* **Duplicate `NegotiateReq` Handling**: If the target's `NegotiateResp` frame
is corrupted on the wire during handshake, the target enters
`run_bridge_session` while the host times out and retransmits `NegotiateReq`
with the **same** `session_id`. When `receiver_task` sees a `NegotiateReq`
matching `current_session_id`, `handle_duplicate_negotiate_req` re-transmits
`NegotiateResp` (`NEGOTIATE_RESP_SEQ = 1`) through `data_tx` without tearing
down the active session.
### 3.6 Serial Device Discovery & Virtual UART Quirks
* **Device Class Filtering (`discover_serial_port`)**: A Fuchsia board may
expose multiple `/dev/class/serial` nodes, including internal board UARTs
(`Class::BluetoothHci`, `Class::KernelDebug`, `Class::Mcu`).
`discover_serial_port` probes `GetClass()` on each sorted path, immediately
selecting the first `Class::Generic` port, holding `Class::Console` as a
fallback (though no current Fuchsia serial driver reports `Class::Console`),
and ignoring internal peripheral UARTs.
* **Non-Fatal `SetConfig` Status (`open_serial_device`)**: Emulated UART
drivers such as `uart16550` on QEMU return an error status from `SetConfig`
if the port is already enabled or if `baud_rate > 115200`, even though the
underlying emulated byte pipe works at any speed. `open_serial_device` logs a
warning on non-zero `SetConfig` status rather than failing device
initialization.
* **Bounded Write Retries (`write_frame_to_device`)**: Transient driver buffer
saturation on `Device.Write` is retried up to `MAX_WRITE_ATTEMPTS`
(5 attempts with `50ms` delay) before tearing down the session and reopening
the device.
[host-design]: ../../ffx/tools/uart_driver/DESIGN.md