jrebmann opened a new pull request, #11708:
URL: https://github.com/apache/nifi/pull/11708

   <!-- Licensed to the Apache Software Foundation (ASF) under one or more -->
   <!-- contributor license agreements.  See the NOTICE file distributed with 
-->
   <!-- this work for additional information regarding copyright ownership. -->
   <!-- The ASF licenses this file to You under the Apache License, Version 2.0 
-->
   <!-- (the "License"); you may not use this file except in compliance with -->
   <!-- the License.  You may obtain a copy of the License at -->
   <!--     http://www.apache.org/licenses/LICENSE-2.0 -->
   <!-- Unless required by applicable law or agreed to in writing, software -->
   <!-- distributed under the License is distributed on an "AS IS" BASIS, -->
   <!-- WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or 
implied. -->
   <!-- See the License for the specific language governing permissions and -->
   <!-- limitations under the License. -->
   
   # Summary
   
   [NIFI-16359](https://issues.apache.org/jira/browse/NIFI-16359) - Git Flow 
Registry Clients: Cache commits per (path, branch) to remove per-process-group 
API load
   
   GitLabFlowRegistryClient against a self-hosted GitLab instance generates 
high load on the GitLab server as the number of versioned process groups grows. 
The load does not come from one expensive operation but from a 
per-process-group multiplier on the version-check path, compounded by a 
GitLab-specific full-history listing. This ticket proposes removing that 
multiplier at the shared base-class level so all Git-based registry clients 
benefit.
   
   ## Environment / Observation
   Self-hosted GitLab, many versioned process groups. Our infrastructure team 
flagged a high volume of GitLab REST API calls (repository/commits, 
repository/tree, repository/files) originating from NiFi. The volume scales 
with the number of versioned process groups, even when most of them are bound 
to the same flow at the same version.
   
   ## Root cause
   The dominant driver is a per-process-group multiplier, not the cost of a 
single listing:
   
   1. The periodic "up to date" check and version listing run per process 
group. In AbstractGitFlowRegistryClient, getLatestVersion(...), 
getFlowVersions(...) and the latest-version comparison in getFlowContents(...) 
all call repositoryClient.getCommits(filePath, branch). With N process groups 
bound to the same flow (identical filePath+branch), each cycle issues N 
identical repository/commits requests. There is no deduplication across process 
groups.
   
   2. repositoryClient is a singleton: getRepositoryClient(context) builds the 
client once (guarded by clientInitialized) and reuses the volatile instance. A 
cache placed on the base class is therefore shared across every process group 
using the component - i.e. the multiplier can be removed with a shared, 
in-memory cache.
   
   3. GitLab-specific amplification: GitLabRepositoryClient.getCommits(path, 
branch) uses the gitlab4j List overload that delegates to Pager.all(), which 
pages the entire commit history for the path (the client sets per_page=100). 
Each of the N calls above is thus itself several page requests.
   
   4. Content resolution (repository/tree, repository/files) is likewise 
performed per process group when materializing a version.
   
   ## Relationship to 
[NIFI-14837](https://issues.apache.org/jira/browse/NIFI-14837) / PR #10186
   [NIFI-14837](https://issues.apache.org/jira/browse/NIFI-14837) (fixed in 
2.6.0, PR #10186) improved the GitHub client by
   
   (a) limiting the commit listing to the first page and
   (b) adding an in-memory SHA->commit-detail cache. Two points:
   
   - Those changes were applied only to GitHubRepositoryClient. GitLab has 
neither, and additionally pages the full history (point 3 above).
   - PR #10186 reduced GitHub load from O(processGroups x commits) to 
O(processGroups): it removed the full-history paging and the per-commit N+1 
detail fetch. It did NOT remove the per-process-group multiplier - 
GitHubRepositoryClient.getCommits(path, branch) still issues a live getRef + 
first-page listing on every call, with no (path, branch) cache and no TTL. This 
simply was not visible at a small number of process groups. So GitHub scales 
linearly with process-group count as well.
   - The SHA->commit-detail cache is not the load-bearing part for GitLab. 
gitlab4j's commit listing returns fully-populated Commit objects (author, 
message, committed_date inline), so there is no per-commit detail call to 
eliminate. Porting that cache 1:1 to GitLab would add a dependency and 
complexity without reducing API calls.
   
   ## Proposed solution
   
   1. Primary, general fix - cache the commit listing at the base class. Add a 
short-TTL, size-bounded in-memory cache keyed by (filePath, branch) in 
AbstractGitFlowRegistryClient, wrapping the getCommits calls in 
getLatestVersion / getFlowVersions / getFlowContents. This collapses the 
per-process-group calls to one remote call per (flow, branch) per TTL window 
and benefits all Git clients (GitHub, GitLab, Bitbucket, Azure DevOps). TTL 
bounds the staleness for detecting a remote new version; invalidate eagerly on 
local writes (createContent, deleteContent) and on client (re)initialization so 
locally committed versions are visible immediately. TTL and size hard-coded 
initially; can be promoted to properties in a follow-up.
   
   2. GitLab-specific - bound the commit listing to the first page (a small 
COMMIT_PAGE_SIZE) instead of Pager.all(), to stop full-history paging and cap 
payload size, mirroring the listing limit from PR #10186.
   
   3. Optional follow-up - immutable content cache. For the repository/tree and 
repository/files load, add a cache keyed by (commitSha, path). Content at a 
fixed SHA is immutable, so this can be cached without a TTL.
   
   ##  Scope question for reviewers
   The primary fix lives in the shared base class and benefits all four Git 
clients, so this ticket's scope effectively broadens beyond GitLab. Should 
NIFI-16359 own the base-class caching (with GitLab as the motivating case), or 
should the base-class change be split into a dedicated framework ticket that 
per-client tickets reference?
   
   References
   [NIFI-14837](https://issues.apache.org/jira/browse/NIFI-14837), PR 
https://github.com/apache/nifi/pull/10186
   
   # Tracking
   
   Please complete the following tracking steps prior to pull request creation.
   
   ### Issue Tracking
   
   - [Apache NiFi Jira](https://issues.apache.org/jira/browse/NIFI) issue 
created
   
   ### Pull Request Tracking
   
   - Pull Request title starts with Apache NiFi Jira issue number, such as 
`NIFI-00000`
   - Pull Request commit message starts with Apache NiFi Jira issue number, as 
such `NIFI-00000`
   - Pull request contains [commits 
signed](https://docs.github.com/en/authentication/managing-commit-signature-verification/signing-commits)
 with a registered key indicating `Verified` status
   
   ### Pull Request Formatting
   
   - Pull Request based on current revision of the `main` branch
   - Pull Request refers to a feature branch with one commit containing changes
   
   # Verification
   
   Please indicate the verification steps performed prior to pull request 
creation.
   
   ### Build
   
   - [X] Build completed using `./mvnw clean install -P contrib-check`
     - [X] JDK 21
     - [X] JDK 25
   
   ### Licensing
   
   - [X] New dependencies are compatible with the [Apache License 
2.0](https://apache.org/licenses/LICENSE-2.0) according to the [License 
Policy](https://www.apache.org/legal/resolved.html)
   - [X] New dependencies are documented in applicable `LICENSE` and `NOTICE` 
files
   
   ### Documentation
   
   - [X] Documentation formatting appears as expected in rendered files


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to