Hello document authors,

Reading the document I have a few concerns/criticisms. Do either of you have 
any ideas regarding the following:

* A benefit of the approach in this document over RTBH, not called out, is that 
operators maintain visibility of whether the attack is still ongoing or not, 
and they can sample the attack traffic to analyse the attack and feed that into 
the DDoS mitigation system/provider (if they have one).

* Because downgrading doesn't black-hole the destination, I could potentially 
downgrade a /24 or /48 rather than a host route, to avoid the existing chaos we 
have with RTBH validating host routes. However, if this is a shared prefix, 
then other customers in the same prefix could be affected by the DDoS. I'm back 
to downgrading host routes, and the major problem with RTBH is the validation 
part, not the traffic dropping part. This document doesn't change anything 
about the validation of host routes. How is this technique expected to be 
effective if what in my opinion is the major problem with RTBH routes 
(validation), also applies here? I realise this question seems somewhat out of 
scope, but I feel it is a genuine concern/problem when it comes to widespread 
adoption of the technique described in this document.

* Related to the adoption problem; a new community has been allocated for 
DOWNGRADE. Another problem with RTBH is that many networks don't use the 
standardised community. Have you considered adding to this document, that 
networks MUST use the standardised community and not a custom community? Or, 
they can use a custom communities but MUST also in addition support the 
standardised one?

* With RTBH, we auto detect the DDoS attacks and inject an RTBH route. With 
DOWNGRADE I can do the same thing. But, if I downgrade a prefix, I'm receiving 
all the attack traffic still, meaning I will have congested core links on the 
PE where the victim is connected, despite the automated mitigation. This should 
be fine because only the low priority attack traffic is being dropped due to 
the congestion, not genuine traffic. However, in order not be spammed with 
"packet loss on core link" alerts, I need my NMS to alert when any of QoS 
queues 0,2-7 have packet loss (but not for packet loss in queue 1). This level 
of detail is not available from all devices via gNMI and not supported by all 
NMSes. Have the authors considered these kinds of operational issues and do you 
have any recommendations for operators here?

*  I think the barrier to entry for this is technique quite a bit higher than 
RTBH. RTBH requires a bit of BGP policy and off you go. Technically, you could 
say that DOWNGRADE also only requires a bit of BGP policy too, to set the QoS 
class, but actually, I'll bet you a round of drinks many networks haven't got 
QoS deployed and/or it doesn't work very well on their devices (we've opened 
several separate bug cases with our vendor trying to get basic QoS to work). So 
if you don't already have QoS working on your network, I'd say the barrier to 
entry is quite a bit higher for DOWNGRADE than RTBH. Have you thought about 
recommendations for operators, i.e. "operators should have QoS deployed as per 
RFC xxx or BCP yyy, this technique builds on top of that QoS model"?

* The main reason for using RTBH is when the attack bandwidth is simply too 
high (there is collateral damage to other customers/services). Because 
DOWNGRADE doesn't stop the traffic, the congestion is still present across 
multiple ASNs. I can de-prioritise the traffic in my network, but my direct 
peer who doesn't support DOWNGRADE, and is sending me the traffic, their link 
to me still congests, which affects all my customers relying on that link. With 
RTBH, I can send the RTBH route to my direct peer, they don't support RTBH, 
they forward it on to their peer, who does support RTBH, and the indirect peer 
drops the traffic, which saves the link capacity between me and my 
non-RTBH-supporting direct peer. With DOWNGRADE, my direct peer doesn't support 
it, they forward the DOWNGRADE route to their peer, they de-prioritise the 
traffic, but the capacity of their physical link to my direct peer is bigger 
than the capacity of the physical link between me and my direct peer, so the 
amount of traffic received at the link with my direct peer still causes it to 
congest and my other customers are still suffering. I would need to fall back 
to RTBH. Have you thought of recommending that DOWNGRADE be implemented in 
addition to RTBH to ensure operators have a fallback? Maybe DOWNGRADE becomes 
the option before RTBH.

Any thoughts/feedback would be appreciated.
James.

Attachment: signature.asc
Description: OpenPGP digital signature

_______________________________________________
GROW mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to