[
https://issues.apache.org/jira/browse/OAK-6276?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16031222#comment-16031222
]
Vikas Saurabh commented on OAK-6276:
------------------------------------
bq. mixed message-producer delays that end up at the same message-consumer
would have the potential to delay each other. But I'd leave that as an
upper-level problem to solve, unrelated to this here.
Agreed that this is a problem for upper-levels. I was just wondering if we can
have some kind of assumptions and help in putting a lower bound on what we want
to do here. e.g. if we assume that all message must always get processed
serially (they are not meant to be processed independently) then we can
probably have delegated producers and consumers implemented by node-store
where-in the node-store is free to produce some milestone messages... the
consumer can consume messages and pass them along to delegate on "hitting a
milestone-message AND then verify that milestone message repo state is
available locally".
I mean some assumptions of what upper-layers can/want to do could simplify what
facilities we expose and maybe how are they exposed.
bq. I would assume that this mechanism shouldn't store anything and that it
shouldn't have to read anything - thus shouldn't be affected by GC in any way.
It should be a pure comparison based on whatever opaque string is returned
ideally. Or am I missing something?
What I meant NodeStore API has no semantics of ordering of revision - e.g.
unless you assume that X,Y,Z in \[rX-0-1, rY-0-2,rZ-0-3\] are timestamps, you
are always supposed to consult the store to find any details about that rev
vector (using DocNodeStore for simple sentences.. but the argument is valid for
NodeStore in general afaict).
We can, of course, bring in a marker interface NodeStoreWithOrderableRevision
and that should solve the problem. But, I think we don't need ordering - we
need visibility of revision \[0]. What I was trying to point out was that while
revisions-not-yet-seen (future revisions) are clearly something we need to
solve. But, along with that, we also need to assert some way that whatever
token thing is passed along can still work with revision that could be
re-written on a RevGC -- I haven't followed with current work on
document-node-store gc by Marcel... but seg-store, for example, would "lose" a
revision that can be peeked into after an compaction cycle.. I understand that
we don't plan to support seg-store... but, to me, it seems that we need to
assert some more guarantees from node-store (not sure what those guarantees
are... just that plain nodestore doesn't allow for this atm afaiu).
bq. But I think it could still be a vector of some kind?
Yes, vector of those contained-opaque-string should do. And, anyway, as you
already said current work on MultiplexingNodeStore shouldn't affect much. We
can tackle fancier multiplexing later :).
\[0]
Actually, thinking more about it - maybe we are ok with with ordering...
visibility etc are already a problem today and upper layers should be able to
handle gracefully if a posted-message in no longer valid on the consumer side.
So, a bunch of stuff I said after "don't need ordering" can probably be
ignored... leaving the thoughts for the record though.
> expose way to detect "eventual consistency delay"
> -------------------------------------------------
>
> Key: OAK-6276
> URL: https://issues.apache.org/jira/browse/OAK-6276
> Project: Jackrabbit Oak
> Issue Type: New Feature
> Components: api
> Reporter: Stefan Egli
>
> I have a requirement to support an external messaging channel (eg Kafka)
> between Oak-based instances of the same cluster. As part of handling those
> messages the target instance in some cases might have to access data from the
> repository.
> Now with DocumentNodeStore's eventual consistency that data might not
> 'travel' from the source to the target instance as fast as is the case for
> the external message.
> Therefore the need arises to be able to delay such messages (on the target
> instance) until the repository sees (at least) the data the source instance
> wrote when sending off the message.
> This ticket is to equip Oak with any feasible way for higher level code to
> generally speaking detect such an "eventual consistency delay".
> One simple idea that comes to mind is to expose the current _head revision
> vector_ (or that from a particular session, but that might not be required,
> ie be too complicated). The source instance could get the local head revision
> vector, piggyback that on the message, then that could be compared on the
> target instance with its head state. If that turns out to be older, then it
> could do a wait and retry. (Nicer would of course be if there would even be a
> call-back - but in theory that could also be implemented ontop of an
> Observer).
> One means to expose the head revision vector would be via a repository
> descriptor (which on access returns the current value, similar to [how
> discovery-lite does
> it|https://github.com/apache/jackrabbit-oak/blob/2634dbde9aedc2549f0512285e9abee5858b256f/oak-core/src/main/java/org/apache/jackrabbit/oak/plugins/document/DocumentDiscoveryLiteService.java#L246]).
> And the format could be normalized as eg {{longs}} (eg {{\[1496071927014,
> 1496071926243]}} (instead of {{\["r15c54d532ec-0-1", "r15c54d532ec-0-2"]}} to
> avoid leaking the revision format explicitly).
--
This message was sent by Atlassian JIRA
(v6.3.15#6346)