[
https://issues.apache.org/jira/browse/NIFI-16379?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Pierre Villard updated NIFI-16379:
----------------------------------
Fix Version/s: 2.13.0
Resolution: Fixed
Status: Resolved (was: Patch Available)
> Restore GenerateFlowFile non-unique generation and batch write behavior
> -----------------------------------------------------------------------
>
> Key: NIFI-16379
> URL: https://issues.apache.org/jira/browse/NIFI-16379
> Project: Apache NiFi
> Issue Type: Test
> Affects Versions: 2.12.0
> Reporter: Joe Witt
> Assignee: Joe Witt
> Priority: Major
> Fix For: 2.13.0
>
> Time Spent: 20m
> Remaining Estimate: 0h
>
> NIFI-16322 changed GenerateFlowFile to stream generated content instead of
> allocating a full-size byte array. The change introduced substantial
> performance
> and behavioral regressions for the processor's primary load-testing use case.
> Before NIFI-16322, when Unique FlowFiles was false, GenerateFlowFile generated
> content once during scheduling and reused that content for every invocation.
> Each FlowFile in Batch Size was independently written to the Content
> Repository.
> After NIFI-16322, the processor stores only a random seed. Every invocation
> regenerates the complete payload through an 8 KB buffer, even though
> non-unique
> FlowFiles contain identical content. Text generation performs a
> Random.nextInt()
> operation for every generated byte. This substantially reduces throughput for
> the default non-unique configuration.
> NIFI-16322 also changed Batch Size behavior. The processor now writes the
> first
> FlowFile and creates the remaining FlowFiles using ProcessSession.clone(). As
> a
> result, a batch performs one Content Repository write instead of one write per
> FlowFile and produces clone behavior rather than independent generated
> content.
> This makes GenerateFlowFile unsuitable for measuring Content Repository write
> throughput and changes observable provenance behavior. DuplicateFlowFile
> already
> exists for workloads that require cloning.
> Restore the pre-NIFI-16322 behavior:
> - Generate and cache content once when Unique FlowFiles is false.
> - Independently write every FlowFile requested by Batch Size.
> - Remove the ProcessSession.clone() batch optimization.
> - Restore the documented higher-throughput behavior for non-unique content.
> - Retain the bounded File Size validator introduced by NIFI-16322 so values
> greater than Integer.MAX_VALUE remain invalid and cannot be truncated.
> Tests should verify:
> - Non-unique content remains identical across triggers.
> - Unique content differs between generated FlowFiles.
> - Batch Size produces independently written FlowFiles rather than
> shared-content
> clones.
> - File Size greater than Integer.MAX_VALUE remains invalid.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)