Hello document authors, Reading the document I have a few concerns/criticisms. Do either of you have any ideas regarding the following:
* A benefit of the approach in this document over RTBH, not called out, is that operators maintain visibility of whether the attack is still ongoing or not, and they can sample the attack traffic to analyse the attack and feed that into the DDoS mitigation system/provider (if they have one). * Because downgrading doesn't black-hole the destination, I could potentially downgrade a /24 or /48 rather than a host route, to avoid the existing chaos we have with RTBH validating host routes. However, if this is a shared prefix, then other customers in the same prefix could be affected by the DDoS. I'm back to downgrading host routes, and the major problem with RTBH is the validation part, not the traffic dropping part. This document doesn't change anything about the validation of host routes. How is this technique expected to be effective if what in my opinion is the major problem with RTBH routes (validation), also applies here? I realise this question seems somewhat out of scope, but I feel it is a genuine concern/problem when it comes to widespread adoption of the technique described in this document. * Related to the adoption problem; a new community has been allocated for DOWNGRADE. Another problem with RTBH is that many networks don't use the standardised community. Have you considered adding to this document, that networks MUST use the standardised community and not a custom community? Or, they can use a custom communities but MUST also in addition support the standardised one? * With RTBH, we auto detect the DDoS attacks and inject an RTBH route. With DOWNGRADE I can do the same thing. But, if I downgrade a prefix, I'm receiving all the attack traffic still, meaning I will have congested core links on the PE where the victim is connected, despite the automated mitigation. This should be fine because only the low priority attack traffic is being dropped due to the congestion, not genuine traffic. However, in order not be spammed with "packet loss on core link" alerts, I need my NMS to alert when any of QoS queues 0,2-7 have packet loss (but not for packet loss in queue 1). This level of detail is not available from all devices via gNMI and not supported by all NMSes. Have the authors considered these kinds of operational issues and do you have any recommendations for operators here? * I think the barrier to entry for this is technique quite a bit higher than RTBH. RTBH requires a bit of BGP policy and off you go. Technically, you could say that DOWNGRADE also only requires a bit of BGP policy too, to set the QoS class, but actually, I'll bet you a round of drinks many networks haven't got QoS deployed and/or it doesn't work very well on their devices (we've opened several separate bug cases with our vendor trying to get basic QoS to work). So if you don't already have QoS working on your network, I'd say the barrier to entry is quite a bit higher for DOWNGRADE than RTBH. Have you thought about recommendations for operators, i.e. "operators should have QoS deployed as per RFC xxx or BCP yyy, this technique builds on top of that QoS model"? * The main reason for using RTBH is when the attack bandwidth is simply too high (there is collateral damage to other customers/services). Because DOWNGRADE doesn't stop the traffic, the congestion is still present across multiple ASNs. I can de-prioritise the traffic in my network, but my direct peer who doesn't support DOWNGRADE, and is sending me the traffic, their link to me still congests, which affects all my customers relying on that link. With RTBH, I can send the RTBH route to my direct peer, they don't support RTBH, they forward it on to their peer, who does support RTBH, and the indirect peer drops the traffic, which saves the link capacity between me and my non-RTBH-supporting direct peer. With DOWNGRADE, my direct peer doesn't support it, they forward the DOWNGRADE route to their peer, they de-prioritise the traffic, but the capacity of their physical link to my direct peer is bigger than the capacity of the physical link between me and my direct peer, so the amount of traffic received at the link with my direct peer still causes it to congest and my other customers are still suffering. I would need to fall back to RTBH. Have you thought of recommending that DOWNGRADE be implemented in addition to RTBH to ensure operators have a fallback? Maybe DOWNGRADE becomes the option before RTBH. Any thoughts/feedback would be appreciated. James.
signature.asc
Description: OpenPGP digital signature
_______________________________________________ GROW mailing list -- [email protected] To unsubscribe send an email to [email protected]
