jarredhj0214 opened a new issue, #13363:
URL: https://github.com/apache/gravitino/issues/13363

   ## Description
   
   Under concurrent authorization traffic, Gravitino's JCasbin authorization 
can observe an intermediate in-memory state during role policy reload. A 
request that should be allowed may be denied if it calls `enforce()` after the 
old role policies are removed and before the new policies are fully added back.
   
   We observed this as intermittent `loadTable` authorization failures. In the 
failing window, the same user and role relationship still had base privileges 
such as `USE_CATALOG` / `USE_SCHEMA`, but the table/schema level policy lookup 
for `SELECT_TABLE` temporarily had no hit policy. A retry shortly afterwards 
succeeded again.
   
   Example symptom:
   
   ```text
   User 'space-account-xxx' is not authorized to perform operation 'loadTable' 
on metadata 'dip_metalake.<catalog>.<schema>.<table>'
   ```
   
   ## Impact
   
   This can produce false authorization denials under high concurrency or 
during role policy reload/expiration. The user may retry successfully later, 
which makes the issue look like a transient permission gap and can trigger 
unnecessary permission repair operations.
   
   ## Root Cause
   
   Role policy reload is currently a multi-step update against the in-memory 
JCasbin enforcers:
   
   ```text
   remove old role policies
   add new role policies
   mark role as loaded
   ```
   
   `SyncedEnforcer` makes individual JCasbin method calls thread-safe, but it 
does not make a multi-call reload sequence atomically visible to concurrent 
authorization reads. A concurrent `enforce()` can therefore observe the 
temporary state where the role's policy rows have been cleared but not yet 
re-added.
   
   ## Proposed Fix
   
   Add an authorizer-level `ReentrantReadWriteLock` around the in-memory 
JCasbin enforcer state:
   
   - Use the read lock for authorization reads and policy inspections:
     - `enforce()`
     - active-role narrowed enforcement
     - deny policy scans
     - diagnostic policy inspection
   - Use the write lock for state mutations:
     - role policy replacement
     - role policy invalidation
     - loaded-role cache removal cleanup
     - user-role grouping bind/prune
   
   Keep DB and metadata resolution outside the write lock. The write lock 
should only protect in-memory enforcer state transitions. Normal authorization 
requests only take the shared read lock, so concurrent reads remain parallel.
   
   ## Validation
   
   A regression test can simulate this by holding the policy write lock, 
clearing a role's policy rows, starting multiple concurrent authorization 
reads, restoring the policy rows, and releasing the write lock. All concurrent 
reads should wait for the reload to complete and then succeed.
   
   Local validation command:
   
   ```bash
   ./gradlew :server-common:test --tests 
org.apache.gravitino.server.authorization.jcasbin.TestJcasbinAuthorizer
   ```
   
   The proposed regression passed locally with the targeted test class.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to