VladRodionov opened a new pull request, #8739:
URL: https://github.com/apache/hbase/pull/8739

   ## Summary
   
   Adds a persistence framework for stateful components in the pluggable HBase 
block cache
   architecture.
   
   Persistence is optional and component-local. A cache component participates 
by implementing
   `PersistentCacheComponent`; components that do not support persistence are 
unaffected.
   
   The change also introduces persistence orchestration and a Hadoop 
FileSystem-backed storage
   implementation so persistent topology, policy, and cache-engine state can be 
saved and restored
   independently.
   
   ## Changes
   
   * Add `PersistentCacheComponent` for cache components that support 
persistence.
     * `getPersistenceId()` provides a stable identifier for the component 
persistence format.
     * `save(OutputStream)` saves component-owned runtime state.
     * `restore(InputStream)` restores state into an already constructed 
component.
   * Persistence remains optional:
     * `CacheEngine`, `CacheTopology`, and `CachePlacementAdmissionPolicy` do 
not automatically
       implement persistence.
     * Concrete implementations opt in by implementing 
`PersistentCacheComponent`.
   * Add persistence infrastructure under the cache persistence package:
     * `CachePersistenceStorage`
     * `CachePersistenceOutput`
     * `CachePersistenceCoordinator`
     * `HadoopFsCachePersistenceStorage`
   * Add persistence orchestration through `CacheAccessService`.
   * `TopologyBackedCacheAccessService` delegates persistence to
     `CachePersistenceCoordinator`.
   * The coordinator discovers and persists topology, placement/admission 
policy, and individual
     cache engines independently.
   * Persistence storage keys include both component role and persistence 
identifier, for example:
     * `topology/<id>`
     * `policy/<id>`
     * `engine/l1/<id>`
     * `engine/l2/<id>`
   * Add a Hadoop `FileSystem` storage implementation supporting local FS, 
HDFS, and other Hadoop
     filesystem implementations.
   * Hadoop FS writes are staged through temporary files and published only 
after an explicit
     successful commit.
   * Failed or abandoned writes are aborted without replacing previously 
committed state.
   * Add unit tests covering persistence coordination, save/restore behavior, 
missing state,
     persistent and non-persistent components, staged publication, abort 
behavior, and preservation
     of previously committed state after save failures.
   
   ## Persistence Model
   
   Cache components are constructed normally from the current HBase 
configuration before persisted
   state is restored.
   
   Persistence therefore restores runtime state into existing objects; it does 
not construct or
   replace cache components.
   
   Current configuration remains authoritative. Restored state must be accepted 
subject to the
   current component configuration and runtime constraints.
   
   Persistence is local to each component. For example, a persistent topology 
saves only
   topology-owned state and does not recursively persist its cache engines. 
Engines and
   placement/admission policies are persisted independently by the coordinator.
   
   ## Stream and Storage Ownership
   
   `PersistentCacheComponent` operates only on caller-provided streams and must 
not close them.
   
   The persistence coordinator owns the persistence streams and is responsible 
for closing them.
   
   For writes, `CachePersistenceOutput` provides explicit commit/abort 
semantics:
   
       component.save(...)
             |
             v
       temporary state
             |
             +-- success -> commit -> publish
             |
             `-- failure -> abort  -> discard temporary state
   
   This prevents a partial component save from replacing previously valid 
persisted state.
   
   `HadoopFsCachePersistenceStorage` uses Hadoop `FileSystem`, keeping the 
component persistence API
   independent of local files, HDFS, object stores, or other Hadoop-compatible 
storage systems.


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to