17 · Reliable data transfer (rdt) details
Slides 3-20 → 3-29 · HW3 P3
The setup
- Reliable data transfer matters at the application, transport, and link layers (a "top-10 networking topic").
- The channel underneath is unreliable. How bad it is (bit errors? losses?) decides how complex the rdt protocol must be.
The 4 interface functions
| Function | Side | Called by / does |
|---|---|---|
rdt_send() | sender | Called from above (the app) to hand data to rdt |
udt_send() | sender (and receiver, for ACKs) | Called by rdt to push a packet over the unreliable channel |
rdt_rcv() | receiver | Called when a packet arrives from the channel |
deliver_data() | receiver | Called by rdt to hand data up to the app |
Memory hook: rdt = reliable (the app's side), udt = unreliable (the channel's side).
Building it up, one problem at a time
| Channel problem | Fix |
|---|---|
| Bit errors (bits flipped) | Checksum to detect them (topic 14). ACK = "got it OK". NAK = "it had errors". Sender retransmits on NAK. |
| ACK/NAK itself gets corrupted | Sender doesn't know what happened, so it retransmits anyway. That can create a duplicate, so add a sequence number to every packet. The receiver discards duplicates (doesn't deliver them up). |
| Packets (data or ACKs) get lost | Sender waits a "reasonable" time for an ACK, then retransmits. Needs a countdown timer. The ACK must say which seq # it's ACKing. |
All of these are stop-and-wait: send one packet, wait for the response, repeat.
Why are seq #s 0 and 1 enough for stop-and-wait?
Only one packet is outstanding at a time. The receiver just needs to know: "is this the same packet again, or the next one?" One bit answers that. (Hence the "alternating-bit protocol".)
The receiver's state says whether it expects 0 or 1. Note: the receiver can't know whether its last ACK/NAK arrived OK at the sender.
If a packet is only delayed, not lost?
The timer fires and the sender retransmits. The receiver gets a duplicate, but the seq # already handles that: it discards it and re-ACKs.
The 4 stop-and-wait scenarios (slides 3-26, 3-27)
| Case | What happens |
|---|---|
| (a) No loss | pkt0 → ack0 → pkt1 → ack1 → pkt0 → ack0… Simple ping-pong. |
| (b) Packet loss | pkt1 is lost. No ACK comes back, so the sender times out and resends pkt1. Receiver gets it, sends ack1. Continue. |
| (c) ACK loss | Receiver gets pkt1 and sends ack1, but ack1 is lost. Sender times out, resends pkt1. Receiver sees seq 1 again → detects duplicate, discards it, re-sends ack1. |
| (d) Premature timeout / delayed ACK | ack1 is just slow. Sender times out and resends pkt1. Then the original ack1 arrives, and the sender moves on to send pkt0. The receiver detects the duplicate pkt1 and sends ack1 again. When that second ack1 reaches the sender, it ignores it (it's waiting for ack0 now). |
The final protocol (FSMs on 3-28, 3-29) in words
| Sender | Receiver |
|---|---|
| Wait for data from above → make pkt with seq 0 + checksum, send it, start timer. | Waiting for seq 0: if pkt is OK and seq 0 → deliver data, send ACK 0, now wait for seq 1. |
| Waiting for ACK0: if the reply is corrupt or ACK1 → ignore it. If timeout → resend, restart timer. | If pkt is corrupt or seq 1 (a duplicate) → send ACK 1 (the last good one), don't deliver. |
| If OK and ACK0 → stop timer, wait for the next data, which goes out with seq 1. | Mirror image for seq 1. |
Notice the final version uses only ACKs, no NAKs: re-ACKing the last good packet tells the sender "the one you just sent didn't make it", which does a NAK's job.
Performance
Stop-and-wait is correct but slow: one packet per RTT. That's topic 16 (U = 0.00027 on a 1 Gbps link), and why we need pipelining (topic 15).
Worked example: HW3 P3 (NAK-only protocol)
Q: A protocol uses only NAKs (no ACKs).
(a) The sender sends data only infrequently. Is NAK-only better than ACKs?
No, worse. With NAKs only, the receiver notices a lost packet only when the next packet arrives and it sees a gap in the seq #s. If data is infrequent, that next packet might come much later, so recovery is slow. With ACKs, the sender's timer catches the loss quickly.
(b) The sender has lots of data and few losses. Is NAK-only better?
Yes, better. Packets arrive constantly, so gaps are noticed almost immediately. And since losses are rare, NAKs are rare, so you avoid sending a huge number of ACKs (much less feedback traffic).
Quick check
1. What does each mechanism fix: checksum, ACK/NAK, seq #, timer?
Checksum: detects bit errors. ACK/NAK: tells the sender whether to retransmit. Seq #: lets the receiver spot duplicates. Timer: recovers from loss.2. Why add sequence numbers once ACKs/NAKs can be corrupted?
The sender retransmits when it can't read the ACK/NAK. If the original actually arrived fine, the receiver gets a duplicate. The seq # lets it detect and discard the duplicate.3. Why do 2 seq #s (0, 1) suffice for stop-and-wait?
Only one packet is outstanding at a time, so the receiver only needs to tell "same packet again" from "next packet".4. In scenario (c), ACK1 is lost. What does the receiver do when pkt1 arrives again?
Detects it's a duplicate (it expects seq 0 now), discards it without delivering, and re-sends ACK1.5. Which function does the app call to send data? Which does rdt call to put a packet on the channel?
App callsrdt_send(). rdt calls udt_send().