Hi Job, Saku,

I like the draft, it solves a few problems with the existing RTBH based 
approach to tackling DDoS attacks:

* It doesn’t complete the DDoS attack so some genuine traffic may still get 
through.

* Visibility isn’t lost of the attack (is it still on going, how big is it, 
what kind of attack pattern is being used, etc.).

* It could de-risk / remove the validation issues we have today with trying to 
validate host routes for RTBH because I could just downgrade a /24 or /48, so 
we can stick to our existing route filtering practices and not have to do 
anything weird to try and validate host routes.


But I also see several problems that would need to be solved before I could 
implement this. Do either of you already have ideas on these (maybe you’ve 
thought of these already?):

* If I downgrade a /24 rather than a host route, then other customers could be 
affected by the DDoS, so that might not be that helpful actually. Then I’m back 
to downgrading host routes, and the major problem with RTBH is the validation 
part, not dropping the traffic. And this draft doesn’t change anything about 
the validation of host routes.

* With RTBH, we can auto detect the DDoS and inject an RTBH route. No need to 
wake up my on-call engineer. If I downgrade a prefix, I’m receiving all the 
attack traffic still, meaning I will have congested core links. But this would 
be fine because only the low priority attack traffic is being dropped due to 
the congestion, not genuine traffic. However, in order not to wake my on-call 
engineer due to “packet loss on core link” alerts, I need my NMS to alert when 
any of QoS queues 0,2-7 have packet loss (but not for packet loss in queue 1). 
This level of detail is not available from all devices via gNMI and not 
supported by all NMS. This could be tricky.

*  The barrier to entry for this is quite a bit higher than RTBH. RTBH requires 
a bit of BGP policy and off you go. Technically, you could say that DOWNGRADE 
also only requires a bit of BGP policy too, to set the QoS class, but actually, 
I’ll bet you a round of drinks many networks haven’t got QoS deployed and/or it 
doesn’t work very well on their devices (we’ve opened 3 separate bug cases with 
Arista trying to get basic QoS to work). So if you don’t already have QoS 
working on your network, I’d say the barrier to entry is quite a bit higher.

* The main reason for using RTBH is when the attack bandwidth is simply too 
high (there is collateral damage to other customers because I’m out of 
capacity, so I need to drop traffic to this prefix to protect my other 
customers). Because DOWNGRADE doesn’t stop the traffic, the congestion is still 
present across multiple ASNs. I can de-prioritise the traffic in my network, 
but my direct peer who doesn’t support DOWNGRADE, their link to me still 
congests, which affects all customers relying on that link. With RTBH, I can 
send the RTBH route to my direct peer, they don’t support RTBH, they forward it 
on to their peer, who does support RTBH, and they drop the traffic, which saves 
the link capacity between me and my direct peer. With DOWNGRADE, my direct peer 
doesn’t support it, they forward the DOWNGRADE route to their peer, they 
de-prioritise the traffic, but the capacity of their physical link to my direct 
peer is bigger than the capacity of the physical link between me and my direct 
peer, so the amount of traffic received at the link with my direct peer still 
causes it to congest and my other customers are still suffering. I would need 
to fall back to RTBH.

Maybe DOWNGRADE becomes an option before RTBH? One could try to use DOWNGRADE, 
if this doesn’t improve the situation enough we can RTBH.

Any thoughts/feedback are appreciated.

Cheers,
James.

Attachment: signature.asc
Description: OpenPGP digital signature

_______________________________________________
GROW mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to