Shraddha –

Please see inline.

From: Hegde, Shraddha <[email protected]>
Sent: Tuesday, July 14, 2026 2:25 AM
To: Les Ginsberg (ginsberg) <[email protected]>; Tony Przygienda 
<[email protected]>
Cc: [email protected]; [email protected]
Subject: Re: [Lsr] Re: New draft on Originator Sequence Number Checksum

Les,

Whatever "strategies" you can build on the receiver side as an experienced 
protocol developer, they aren't going to solve the problem. You cannot be too 
slow to react to an LSP change unless you are sure that there is another change 
related to the first change which is better to wait for.
[LES:] Well – here we disagree. But I am willing to listen to you. But you are 
going to have to provide a much more explicit example than you have so far. I 
assume you have some.

No LSP flooding mechanism change is being proposed in this draft. A new TLV is 
being proposed to carry additional information that gives a clue to the 
receiver, there is an important transaction across multiple fragments.

Can you explain what makes you think basic flooding mechanism change is 
proposed?
[LES:] You have changed the processing rules when an LSP is received. It is no 
longer sufficient to simply validate checksum/authentication. Additional 
processing is now required before what is a “valid LSP” can actually be used by 
the receiver.
The implications of this on implementations is significant.

   Les

Rgds
Shraddha


________________________________
From: Les Ginsberg (ginsberg) <[email protected]<mailto:[email protected]>>
Sent: 14 July 2026 11:03
To: Hegde, Shraddha <[email protected]<mailto:[email protected]>>; 
Tony Przygienda <[email protected]<mailto:[email protected]>>
Cc: [email protected]<mailto:[email protected]> 
<[email protected]<mailto:[email protected]>>; 
[email protected]<mailto:[email protected]> <[email protected]<mailto:[email protected]>>
Subject: RE: [Lsr] Re: New draft on Originator Sequence Number Checksum


Shraddha –



It is not that I don’t understand the problem.

It is that it is possible – with a smart LSP Generation implementation – to 
avoid moving a TLV from one LSP to another in most cases. And in the rare cases 
where a TLV move is unavoidable, there are strategies to make that hitless from 
the receiver’s POV.



Now, maybe I am overlooking some particular case – which is why I ask for some 
of the specific examples which convinced you that a major change to LSP 
flooding was required.



We are both experienced implementors of the protocol. So it is going to take 
more than the example you provide below to convince me that we have to change 
LSP flooding.

Please consider this in what you present.



    Les





From: Hegde, Shraddha <[email protected]<mailto:[email protected]>>
Sent: Monday, July 13, 2026 9:56 PM
To: Tony Przygienda <[email protected]<mailto:[email protected]>>; Les 
Ginsberg (ginsberg) <[email protected]<mailto:[email protected]>>
Cc: [email protected]<mailto:[email protected]>; 
[email protected]<mailto:[email protected]>
Subject: Re: [Lsr] Re: New draft on Originator Sequence Number Checksum



Hi Les,



The problem statement is described in the Introduction section with an example.

If its not clear, the text can be enhanced in next revision.



I understand the existing mechanisms of flooding the two fragments together and 
on the receiver applying spf-delay ensure multiple changes are accommodated.



Every mature ISIS implementation have these mechanisms implemented but its not 
sufficient to handle the fragment movement situation consistently all the time. 
Even if the sender floods the fragments together, there is no guarantee they 
are received together on the receiver. Even if the receiver applies a delay to 
process the packets there is no guarantee  both the fragments got received 
within that  delay timer.



Imagine Srv6 being enabled on a node A. Srv6 END.X SID will be advertised on an 
adjacency. TLV 22 moves from Fragment 1 to Fragment 3. Node A floods fragment 1 
and 3 together. On the receiver Fragment 1 arrives first. A will trigger the 
SPF based on Fragment 1 and a stale fragment 3. Node A triggers SPF/ TED update 
causing link to be removed in the computation. The impact  is routes might get 
deleted/ nexthop changed, Tunnels might get re-routed. In the next instance 
Fragment 3 arrives Routes would get re-added/ nexthop changed and Tunnels might 
get re-routed.





Rgds

Shraddha

















________________________________

From: Tony Przygienda <[email protected]<mailto:[email protected]>>
Sent: 14 July 2026 04:28
To: Les Ginsberg (ginsberg) <[email protected]<mailto:[email protected]>>
Cc: [email protected]<mailto:[email protected]> 
<[email protected]<mailto:[email protected]>>; Hegde, Shraddha 
<[email protected]<mailto:[email protected]>>; 
[email protected]<mailto:[email protected]> <[email protected]<mailto:[email protected]>>
Subject: Re: [Lsr] Re: New draft on Originator Sequence Number Checksum







On Tue, Jul 14, 2026 at 12:35 AM Tony Przygienda 
<[email protected]<mailto:[email protected]>> wrote:





On Mon, Jul 13, 2026 at 7:11 PM Les Ginsberg (ginsberg) 
<[email protected]<mailto:[email protected]>> wrote:

Tony –



No “mud-slinging” from me.



Les, ack





But as the owner of an implementation – and having chatted about your proposal 
with some colleagues – it isn’t obvious to me why such issues cannot be 
addressed at the originator (and in fact are already addressed).

Perhaps I am missing something – which is why I ask that you provide some 
detailed examples (not just the high level abstract that you have provided so 
far).



ack. the proposal was published last minute as initial draft. consider it as 
such. As I said, we do have a "much better" (in some metrics) solution but we 
judged it too complex in terms of practical deployment (for now). I argued for 
the OGID last minute as I said, you found it pretty quickly ;-) However, beyond 
that, anything you start to consider becomes really complex IME since it starts 
to build a "perfect transaction" algebra. But I'm more than interested in your 
suggestions of something relatively simple that works better than the proposal. 
I would suggest to look up the theory of "consistent cut snapshots" in 
distributed systems since this is basically the problem we're tackling (modulo 
colored channel we are thankfully spared, especially since flooding is anything 
but FIFO ;-) and is a well understood and solved problem in distributed 
databases since long time. But we do not have space for proper vector clock 
algebra in the encoding and hence the "we can only do a (good enough)" 
approximation here IMO rather than the full enchilada that would deliver a 
"perfect" solution.



Les, I realized having thought about the problem for quite a bit I ended up 
bits loose with my language here. So to be more precise: "vector clock" 
normally means "vector of single clocks of different distributed systems". 
However, this problem is really "congruent" once we consider each fragment of a 
node a "clock" with sequence number ticking and a consistent cut is where a 
remote node receives a "consistent cut"  of fragments (while there are no 
involved asynchronous channels within a node itself producing the vector). You 
already alured to it with the 6/4 problem, so to expand as an example



# is fragment nr. @ is seqnr, all within a single node



#1 @ 1 |

#2 @ 1 |  transaction 1



#3 @ 1

#4 @ 1 | transaction 2 -> which means it will encode transaction 1 clocks for 
#1 and #2 within it



#1 @ 2 | transaction 3 -> now the "vector clock is"  [ #1@2, #2@1, #3@1, #4@1 ]



so what we are really sending are the full vectors "mashed together as a hash" 
that represent consistent cuts. The space saving "mashing" prevents us from 
understanding what are the clock values or even the changed fragments we need 
or even omit the clocks not changed between transactions (SK techniques) but 
given the very small encoding space it seems to me it's about as well as we can 
do. A better approach would allow us to e.g. know that "we miss fragment #4" to 
have a consistent cut and prioritize the handling somehow etc but alas, vectors 
are not small (and if you think further you will start to derive the 
mathematically densest compression of differential vector clock here or think 
about "catalogs" but things get complex quickly or start to generate double 
volume of flooding in certain approaches).



 I hope that clarifies things a bit







Changing the fundamentals of flooding is far from a trivial undertaking. I 
would not want to invest in that until it is clear that it is needed.



no qualms here.



and here my English was slippery as well. When I said "no qualms" it was not 
"I'm more than happy to hack flooding anytime", it was more agreeing with you 
that any flooding behavior change needs very serious discussion, whether change 
is optional or not.



No'one was claiming this is the table from the mountain with the perfect 
solution cast in stone. So consider it a starting point and do argument away.



 "link flapping" we see despite all the implementation techniques (and more) is 
real and visible (although not very common) and  it can cause noticeable 
disturbances as you know from regular link flaps. This is grave enough. Further 
"examples" will IMO contribute nothing except red herring arguments about 
2ndary effects. The problem is always the same. multiple TLVs sprinkled around 
changing independently or something sliding into the next fragment due to 
sliding or re-packing. Everything after that is 2nd order computation on such 
inconsistent cuts and triggers lots of detrimental stuff further down the 
pipeline of instrumentation relying on IGPs view of topology state.



-- tony





   Les





From: Tony Przygienda <[email protected]<mailto:[email protected]>>
Sent: Monday, July 13, 2026 9:47 AM
To: [email protected]<mailto:[email protected]>
Cc: Les Ginsberg (ginsberg) <[email protected]<mailto:[email protected]>>; 
Shraddha Hegde <[email protected]<mailto:[email protected]>>; 
[email protected]<mailto:[email protected]>
Subject: Re: [Lsr] Re: New draft on Originator Sequence Number Checksum







On Mon, Jul 13, 2026 at 5:51 PM [email protected]<mailto:[email protected]> 
<[email protected]<mailto:[email protected]>> wrote:

comments inline



Sent from my iPhone



On Jul 13, 2026, at 11:08 AM, Les Ginsberg (ginsberg) 
<[email protected]<mailto:[email protected]>> wrote:



Tony/Shraddha –



Thanx for the replies.



As I see it, fundamentally you are trying to change the current flooding model 
– which is per LSP – to be “per LSP Batch”.

Not sure I want to go down that road…and there are certainly strategies 
associated with LSP Generation that can minimize the occurrence of this problem.



I think a better way to start this discussion would be for you folks to go into 
more detail on the problematic cases and why they cannot be addressed at the 
source.

I would hope/ask that you focus the discussion there to start with – rather 
than starting with a solution which even you folks admit needs significantly 
more work.



JD:  I agree.  This is the way the process is supposed to work.  Presenting an 
ill-considered solution and saying “voila” considered bad form.



@JD: the 'ill-considered' seems simply a bit of mud slinging given you probably 
don't know or think through all the likely problems encountered (although the 
ones Shraddha described are serious and for practical purposes detailed enough) 
and neither seem to have a better proposal to extend.



Further, the problems we see otherwise can be all simply described as "either 
something slid in a fragment or two independent fragments changed with 
inter-dependent information" and only one arrived on flooding and then some 
piece of logic hanging of IGP like SPF acted on it and triggered spurious 
triggers. So if you know your RFCs, deploy IGPs at scale and think a bit you'll 
quickly see where we have it today already and how we have ongoing 
proliferation of e.g. link information being sprinkled in lots of TLVs on an 
ongoing basis which for all practical purposes end up depend on each other. And 
that's AFAIS is really enough in the context of the problem encountered.



--- tony







Thanks.



    Les





From: Tony Przygienda <[email protected]<mailto:[email protected]>>
Sent: Monday, July 13, 2026 1:59 AM
To: Hegde, Shraddha <[email protected]<mailto:[email protected]>>
Cc: Les Ginsberg (ginsberg) <[email protected]<mailto:[email protected]>>; 
[email protected]<mailto:[email protected]>
Subject: Re: [Lsr] Re: New draft on Originator Sequence Number Checksum



for 3. Les, thanks for comments. Despite having implemented (obviously) pretty 
much e'thing you mention and even more clever things to deal with the fragment 
movement we still encounter the problem in the field and hence propose this 
TLV. Some affected operators may chime in. We do not only see link flaps but 
other problems manifesting due to fragment shifting BTW.



We have also a proposal dealing perfectly with the transaction order/no delay 
necessary etc but it's significantly more complex and we judged that this is 
the goldi locks between extreme complexity and something that should improve 
the situation noticeably from the current state. From the more complex proposal 
I suggested and we are discussing through possibly including the transaction ID 
 which will solve your example as in



fragment 4 -> transaction ID 1

fragment 6 -> transaction ID 2



so when 6 floods in first and then 4 the node knows that it has still to use 
the 6 to validate the checksum and 4 with ID 1 checksum does not matter.



tricky part are of course reboots/attacks and so on and very high frequency 
fragments being regenerated for which the timer is the simplified heuristic 
since the complete solution is practically speaking much more complex than this 
draft



-- tony



On Mon, Jul 13, 2026 at 6:49 AM Hegde, Shraddha 
<[email protected]<mailto:[email protected]>> 
wrote:

Hi Les,



Thanks for detailed review and valuable comments.



  1.  This is a valid scenario and as per the current proposal there will be a 
delay in computing SPF , eventually the timeout will cause the SPF to be 
triggered.

 This delay is not desirable and the proposal can be improved to do better in 
this case.

My thought is to indicate the fragment movement case in the TLV with a flag and 
let receiver apply delay only when needed.

Will update draft once window opens.

  1.  I agree when there are unrelated changes to multiple LSPs the existing 
mechanisms should be good enough

  1.  I don't agree the current mechanisms are good enough to handle the  
fragment movement case. There are a number of new TLVs/sub-TLVs being proposed 
in IETF and every time a new feature is to be enabled in the network, operator 
has the risk that it might cause fragment movement and cause disruptions in the 
network. The sender deterministically knows the fragment movement got triggered 
and can do better to let other nodes gracefully handle it.



Rgds

Shraddha



________________________________

From: Les Ginsberg (ginsberg) 
<[email protected]<mailto:[email protected]>>
Sent: Wednesday, July 8, 2026 5:33 AM
To: Hegde, Shraddha <[email protected]<mailto:[email protected]>>; 
[email protected]<mailto:[email protected]> <[email protected]<mailto:[email protected]>>
Subject: RE: New draft on Originator Sequence Number Checksum



Shraddha –



Some questions/comments regarding this draft.



1)I am struggling to see how what you propose works in the following scenario.



Node A currently has generated/flooded 10 LSPs at Level-1.

At Time T0, it updates LSP A.00-04 , recomputes the OSNC – which is “1234” – 
and includes that in the new version of A.00-04 and floods the LSP.



A short time later – say T0 + 20 ms  -  Node A now updates LSP.00-06 with 
information unrelated to the previous change to LSP #4.

It recomputes the OSNC – which is now “5678” – and includes that in the new 
version of A.00-06 and floods that LSP.



You cannot control the order in which the LSP updates are received as the LSPs 
are propagated through the area.



Some nodes may receive LSPs in the following order:

A.00-04(OSNC 1234)

A.00-06(OSNC 5678)



Nodes which try to validate their database against the received OSNC will 
succeed as they process each received LSP.



But some nodes may receive LSPS in the following order:

A.00-06(OSNC 5678)

A.00-04(OSNC 1234)



Nodes which receive the LSPs in this order will fail the OSNC validation.



The problem here is that the receivers do not know the order in which the LSPs 
were generated. So any set of LSPs which were not part of an “atomic update” on 
the originator will have different OSNC values and receivers don’t know which 
one is the latest.



How are you intending to address this problem??



2)The problem you are trying to address (updates to a single object – such as a 
neighbor) which impact multiple LSPs is not the only case in which updates to 
multiple LSPs may occur.

A node may (for example) add a new neighbor to one LSP and a new prefix to 
another LSP. There is far less motivation to treat these as an atomic update – 
yet you are imposing requirements on the receiver that these two unrelated 
updates cannot be processed by the receiver until/unless both are received.



3)The issue of movement of information about a single object is well known and 
was considered even in ISO 10589. Section 7.3.4.4 states:

“Once a particular adjacency has been assigned to a particular LSP Number, it 
is desirable that it not be moved to another LSP Number. This is because moving 
an adjacency from one LSP to another can cause temporary loss of connectivity 
to that system. This can occur if the new version of the LSP which originally 
contained information about the adjacency (which now does not contain that 
information) is propagated before the new version of the other LSP (which now 
contains the information about the adjacency).

…

If it is necessary to move an adjacency from one LSP to another, the SRMflags 
(see 7.3.15) for the two new LSPs shall be set as an atomic action.”



In addition, implementations have for many years included the capability to 
flood a limited number of LSPs “back-to-back” as a means of insuring that when 
other routers in the network trigger the Decision process, they would be more 
likely to have all the updated topology information.

Also, implementations provide the ability to delay the start of the Decision 
process by a modest amount of time (e.g., 50 milliseconds) to allow time for 
multiple LSPs to be received before a new SPT is calculated.

All of this has been demonstrated to reliably improve convergence.



So, while the concern you raise regarding moving an object such as a neighbor 
from one LSP to another is legitimate, it has already been considered in 
existing implementations.



Given the added complexity the draft’s solution imposes on receivers, the 
difficulties in dealing with cases when the order of receipt does not match the 
order of generation, and the significant potential for negatively impacting 
convergence – it is difficult to be enthused about this proposal.

It seems to me to be an example where  “striving for perfection is the enemy of 
the good”.



   Les





From: Hegde, Shraddha 
<[email protected]<mailto:[email protected]>>
Sent: Monday, July 6, 2026 10:05 PM
To: [email protected]<mailto:[email protected]>
Subject: [Lsr] New draft on Originator Sequence Number Checksum



WG,



A new draft is posted on Originator Sequence number







The proposal defines a new TOP level TLV that carries checksum of LSP-ID, 
sequence number and size of all the self originated LSPs. This is to be used by 
the receiver to get the information about the senders database. Pls review 
details from the draft below



https://datatracker.ietf.org/doc/draft-hegde-lsr-isis-osnc/<https://urldefense.com/v3/__https:/datatracker.ietf.org/doc/draft-hegde-lsr-isis-osnc/__;!!NpxR!jKnJ-5wj-zweoNDtMG8kZVwCP1dGlm2X1VJDpQULPfi9AxNixaFa45znEdKVp0w54EJC1KVHFS5y0HgMdufNTKoLohX_v_Vq$>



Shraddha Hegde

Distinguished Engineer

HPE Networking

E: [email protected]<mailto:[email protected]>

<image001.png>



_______________________________________________
Lsr mailing list -- [email protected]<mailto:[email protected]>
To unsubscribe send an email to [email protected]<mailto:[email protected]>

_______________________________________________
Lsr mailing list -- [email protected]<mailto:[email protected]>
To unsubscribe send an email to [email protected]<mailto:[email protected]>
_______________________________________________
Lsr mailing list -- [email protected]
To unsubscribe send an email to [email protected]

Reply via email to