[
https://issues.apache.org/jira/browse/HDDS-16167?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Ivan Andika updated HDDS-16167:
-------------------------------
Summary: Ozone development best practice documentation (was: Create Ozone
development best practice documentation)
> Ozone development best practice documentation
> ---------------------------------------------
>
> Key: HDDS-16167
> URL: https://issues.apache.org/jira/browse/HDDS-16167
> Project: Apache Ozone
> Issue Type: Improvement
> Reporter: Ivan Andika
> Assignee: Ivan Andika
> Priority: Major
>
> In the course of Ozone project lifetime, we have encountered a lot of issues
> and learned some patterns to follow and anti patterns to avoid.
> We should start writing a best opiniated development practice of Ozone to
> prevent contributors falling to the same trap.
> We can split the best practices to per-component (e.g. RocksDB, Ratis, OM,
> SCM, DN, etc) or conceptual (Concurrency, etc). Note that the focus should be
> on Ozone related logic and should be concise, other general patterns like in
> "Effective Java" or "Java Concurrency in Practice" should be avoided unless
> they are directly related to Ozone issues. Things that can be enforced
> statically (e.g. PMD, etc) should also be avoided.
> Example best practices
> * RocksDB
> ** Do NOT hold a lock while doing a RocksDB range query (e.g. seek, iterate)
> ** Flatten a RocksDB schema so that each RocksDB entry size is bounded
> *** We have learned this with issues like very large MPU table entry (with
> multiple parts), Ozone key versions that refer to the previous versions,
> Deleted table that contains multiple versions of blocks
> ** Do NOT set exclusiveManualCompaction: This can cause write stalls
> *** Generally, don't invent a new
> ** RocksDB open is expensive
> *** For example, we need to learn this the hard way in the "Merge Container
> RocksDB in DN" improvement
> * OM
> ** OM can only be uniquely identified by using OM ID (which should be a
> UUID) instead of Node ID (which can diverge between client and server)
> * Ratis
> ** Do NOT block in the Ratis notification API since it will can cause Ratis
> to be stuck
> *** This includes Thread.join, etc
> *** We can rephrase this more formally to be something like Ratis
> notificaion API implementation need to be wait-free
> **** Wait-free can be designed formally: A method of an object
> implementation is _wait-free_ if every call finishes its execution in a
> finite number of steps. That is, if a thread with a pending invocation to a
> wait-free method keeps taking steps, it completes in a finite number of steps
> (From Art of Multiprocessor Programming)
> *** Do not submit another Ratis request in a Ratis notification API since
> this can cause Ratis deadlock
> ** Do NOT invent your own Ratis logic
> *** Like generated NotLeaderException
> ** Ratis implementation should replicate the state mutation not the command
> (see Leader Execution Framework)
> * General design
> ** Design decision need to have a precedent (either use a paper in the area,
> or copy approach in other production systems like HDFS, Ceph, etc)
> *** If the decision has no precedent, then it is a wholly new approach and
> the designer need to specify this deeply and the reviewers need to evaluate
> this decision deeper
> **** This is because a new approach is a new uncharted territory with
> unforeseen risks
> *** For example
> **** Ozone introduces a new Ozone Container which is an aggregation of blocks
> ***** AFAIK This is a whole new design with no precedent
> ***** The original designer does not seem to consider things like end-to-end
> deletion process, so we end up with orphan blocks issues with no long term
> solution
> ***** The ReplicationManager feature
> **** Ozone has a Raft based systems combined with eventually consistent
> heartbeats
> ***** AFAIK, this is only done in Ozone, we need to know why, we can compare
> with how HDFS reconcile this
> **** We have a Container Reconciliation feature that uses Merkle Tree,
> Merkle Tree is already used in DynamoDB paper, so there is a precedent
> ** Do NOT use Raft for a write pipeline
> *** We have learnt this the hard way in the past few years (QUASI_CLOSED,
> follower index lag logic, etc)
> ** Each class should have a concurrency model documentation (this is
> tentative in the future after we consolidate Ozone concurrency model)
> *** Use ThreadSafe, GuardedBy annotation
--
This message was sent by Atlassian Jira
(v8.20.10#820010)
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]