Kevin,

Py4J is the glue that bridges the NiFi JVM to the python process that runs 
python processors. So there’s a py4j server that is run in a separate python 
process. Then a py4j client that runs in the JVM. Py4J handles the RPC to make 
it possible to call the Python processors as if they were Java objects.

Thank
-Makr


> On Aug 13, 2026, at 3:37 PM, Kevin Doran <[email protected]> wrote:
> 
> Can someone clarify where Py4J is used today vs where subprocesses
> that invoke a configurable python executable are used? My
> understanding was that the Py4J was mostly restricted to the scripting
> processors, but if that is not the case I would be interested in
> hearing more.
> 
> On Thu, Aug 13, 2026 at 1:42 PM David Handermann
> <[email protected]> wrote:
>> 
>> Dan,
>> 
>> Thanks for the substantive reply and consideration of the various
>> issues surrounding support for native Python Processors.
>> 
>> As you highlighted, the scalability issue is foundational. Although
>> Processor authors can work with self-imposed constraints, the NiFi
>> framework does not provide the same level of safeguards. Any future
>> solution would need to account for these types of production-level
>> concerns.
>> 
>> Regarding supply chain security, you are correct that there are
>> potential paths forward on the packaging strategy. The framework could
>> remove support for runtime loading as a safety measure. Right now,
>> lack of pip or uv effectively disables runtime loading. Adjusting the
>> default project container configuration would strengthen the standard
>> position, and thus require Python Processors to be packaged using a
>> NAR. The project could provide tooling, similar to the NiFi NAR Maven
>> Plugin, but this broadens the maintenance scope in a meaningful way.
>> Thus before going in this direction, the future support strategy needs
>> to be considered.
>> 
>> The Py4J architecture remains a critical concern and high risk for
>> future maintainability. Considering an alternative architecture would
>> require a substantial redesign and rewrite of Python Processor
>> support, abandoning a majority of the current implementation. Current
>> tooling makes this more plausible in theory, but is a significant
>> undertaking that is difficult to justify without serious commitment on
>> the part of multiple project maintainers. To be clear, deprecating the
>> current implementation for removal does not preclude future redesign.
>> These two questions should be considered independently.
>> 
>> Deprecation and removal without replacement will have a different
>> impact depending on the current scope of usage. Simpler Python
>> Processors should be easier to convert to Java, but more complex
>> implementations would likely require something outside of NiFi, as you
>> mentioned. Although these are difficult questions, it ultimately comes
>> down to project focus and available maintenance cycles. Apache NiFi
>> has historically done many things. NiFi 2 served as a focusing
>> function. Going forward, maintaining project health and stability is
>> paramount, which may at times involve removing certain capabilities.
>> 
>> Regards,
>> David Handermann
>> 
>> On Wed, Aug 12, 2026 at 3:27 PM Daniel Weidman via dev
>> <[email protected]> wrote:
>>> 
>>> David,
>>> 
>>> Thanks for raising this directly rather than letting it drift, and for 
>>> laying
>>> out the reasoning rather than just the conclusion.
>>> 
>>> On scalability, your point lands. In my experience the memory behavior of a
>>> Python Processor is largely a function of how it is written, and disciplined
>>> implementations can be well behaved. But I recognize that is an argument 
>>> about
>>> good citizenship, not about enforcement, and it does not give an operator 
>>> the
>>> equivalent of a heap ceiling. That is a real gap. It does seem addressable 
>>> at
>>> the process supervision layer, through cgroup or rlimit bounds on spawned
>>> interpreters, rather than being inherent to Python as an extension language.
>>> 
>>> On security, I would offer a friendly amendment. The meaningful difference 
>>> does
>>> not seem to be Python versus Java, but when and where third party 
>>> dependencies
>>> land. A NAR resolves its tree at build time and ships self contained, so it 
>>> can
>>> be scanned, signed, and deployed into an air gapped environment, and 
>>> unpacking
>>> it executes nothing. The current Python path resolves dependencies at 
>>> runtime
>>> via pip on the NiFi host, and pip installation does execute packaging code.
>>> That is a real governance gap, and I will concede PyPI carries risks Maven
>>> Central does not.
>>> 
>>> But as you noted, the framework already supports loading Python Processors
>>> packaged as NARs. What is missing is project-provided tooling to build them.
>>> Sean's note about relying on a blog post to produce one seems telling. 
>>> Closing
>>> that gap looks like a smaller undertaking than removing the extension point,
>>> and it would address the supply chain concern directly rather than by
>>> subtraction.
>>> 
>>> I will admit I had not realized the current implementation still rests on 
>>> Py4J.
>>> That reframes a good deal of this for me, and it tracks with Matt's read 
>>> that
>>> Py4J is the long pole rather than the capability itself.
>>> 
>>> For what it is worth as an operator, the practical value of Python 
>>> Processors
>>> is not just iteration speed. It is that the people writing our transforms 
>>> are
>>> Python native. Removal without a replacement does not convert that work to
>>> Java, it pushes it out of NiFi entirely, to sidecars or external services,
>>> where, as Marc noted, we keep every one of the concerns you listed but lose
>>> NiFi's visibility into them.
>>> 
>>> I would rather see the runtime architecture reconsidered than the capability
>>> retired. If the deciding constraint is maintainer cycles, it would be 
>>> useful to
>>> know what a replacement would need in terms of effort and expertise. That 
>>> is a
>>> more actionable ask of the community than a deprecation vote.
>>> 
>>> Regards,
>>> Dan
>>> 
>>> 
>>> From: David Handermann <[email protected]>
>>> Date: Wednesday, August 12, 2026 at 9:10 AM
>>> To: [email protected] <[email protected]>
>>> Subject: Re: [DISCUSS] Future Support for Python Processors?
>>> 
>>> Thanks to everyone for the replies thus far!
>>> 
>>> The question about alternatives gets at the core issue, and the ideas
>>> around a more loosely coupled architecture provide some helpful
>>> possibilities.
>>> 
>>> At a basic level, simple deprecation and removal would leave the
>>> current set of options, including standard Java, Groovy scripting, or
>>> externalized Python scripts executed as a stream. None of these
>>> provide the same level of features for Python itself, which prompted
>>> the original native Python Processor implementation. Without a
>>> replacement, these alternatives would remain. With the rise of
>>> AI-enabled development, it is now much easier to go from Python to
>>> Java, but work is still required.
>>> 
>>> As to future strategies for Python integration, the current discussion
>>> around deprecation is more narrowly focused on the existing
>>> framework-level extension. Removal of the current implementation is
>>> not intended to halt future innovation. Although removal without an
>>> initial replacement might not be the perfect solution, it may be
>>> necessary given available maintainer cycles, and the substantive
>>> issues I described in the opening.
>>> 
>>> If there is sufficient interest and effort to propose and build a
>>> different solution, that would be a great outcome of this discussion.
>>> I've thought at different points about reworking the Script Engine
>>> integration to provide a cleaner path to external extension, so there
>>> are definitely opportunities in that direction.
>>> 
>>> With that, my primary concerns with the current implementation remain,
>>> but glad to consider future alternatives.
>>> 
>>> Regards,
>>> David Handermann
>>> 
>>> 
>>> On Wed, Aug 12, 2026 at 7:52 AM Matt Burgess <[email protected]> wrote:
>>>> 
>>>> I agree with Marc, the concerns of the implementation of the capability
>>>> should not warrant deprecation and future removal of the capability, but
>>>> certainly we want to ensure that anything we implement for NiFi doesn't
>>>> negatively the reputation of Apache NiFi as a secure, robust, reliable
>>>> system.
>>>> 
>>>> I think we should instead consider a new architecture, one
>>>> fully extensible so if the community decides to deprecate the capability,
>>>> it could still be restored using some extension point. This kind of generic
>>>> approach could allow other extensions not sanctioned by the Apache NiFi
>>>> community to coexist in the "NiFi+" deployment. Not to get off-topic but it
>>>> is somewhat analogous to exposing the Spring context such that alternate
>>>> implementations of the beans could be provided.
>>>> 
>>>> There are options to consider, I'm sure each with its own set of pros and
>>>> cons. Off the top of my head:
>>>> 
>>>> - Sidecar server responsible for executing Python scripts with FlowFile
>>>> input(s) and returning FlowFile output(s) directly from/to processors, like
>>>> Marc's suggestion but managed by NiFi
>>>> - Jep [1] and it's OS-independent counterpart jep-distro [2]
>>>> - Require Java bindings calling out to an embedded interpreter, such as
>>>> WrapPy4J (which uses Jep under the hood) [3]
>>>> 
>>>> I am far from a Python expert and I know there are several places in the
>>>> framework where the Python stuff has landed, so I don't expect any other
>>>> implementation to be any easier or take less time than the original, which
>>>> I'm sure took quite a lot of time and effort which is very much
>>>> appreciated. It seems that Py4J is the long pole in the tent for the
>>>> concerns, not the capability itself.
>>>> 
>>>> At the end of the day I think a discussion like this clearly shows the
>>>> commitment of the whole community to ensure Apache NiFi is all that we want
>>>> it to be and that we can introspect, improve, and collaborate to keep NiFi
>>>> awesome :)
>>>> 
>>>> Regards,
>>>> Matt
>>>> 
>>>> [1] 
>>>> https://urldefense.com/v3/__https://github.com/ninia/JEP?tab=readme-ov-file__;!!FtR4BK4x7WL3xYs!6mIniDSbJ62B5zII3b9g972tiJiM1K8M1E-1R8rvwlwl__KEfyNqtqWjM3xwIPbVawsv_VEsYzoNRToUxhft8Q3Rnnw$
>>>> [2] 
>>>> https://urldefense.com/v3/__https://github.com/icemachined/jep-distro?tab=readme-ov-file__;!!FtR4BK4x7WL3xYs!6mIniDSbJ62B5zII3b9g972tiJiM1K8M1E-1R8rvwlwl__KEfyNqtqWjM3xwIPbVawsv_VEsYzoNRToUxhft703BJKY$
>>>> [3] 
>>>> https://urldefense.com/v3/__https://github.com/313hemant313/wrappyj__;!!FtR4BK4x7WL3xYs!6mIniDSbJ62B5zII3b9g972tiJiM1K8M1E-1R8rvwlwl__KEfyNqtqWjM3xwIPbVawsv_VEsYzoNRToUxhftYE_fqvo$
>>>> 
>>>> On Wed, Aug 12, 2026 at 8:19 AM Marc P. <[email protected]> wrote:
>>>> 
>>>>> This is a feature I find personally valuable. I agree with several
>>>>> underlying issues, particularly resource management, Py4J, dependency
>>>>> handling, and the difficulty of maintaining a framework that spans both 
>>>>> the
>>>>> JVM and Python ecosystems.
>>>>> 
>>>>> Where I disagree is with the conclusion that these issues necessarily 
>>>>> point
>>>>> toward deprecating Python Processor support itself.
>>>>> 
>>>>> I think there is an important distinction between the value of Python as 
>>>>> an
>>>>> extension model and the architecture currently used to implement it.
>>>>> 
>>>>> To be solution oriented one way to look at this is to consider what users
>>>>> would do in the absence of Python Processor support. In many cases, the
>>>>> alternative is not to rewrite the functionality in Java. It is likely to
>>>>> become something like:
>>>>> 
>>>>> Apache NiFi -> HTTP/gRPC -> Python microservice -> NiFi
>>>>> 
>>>>> That architecture still has independent Python processes, Python 
>>>>> dependency
>>>>> management, memory and CPU consumption outside the JVM, and a Python
>>>>> software supply chain. We have not eliminated any of those concerns; we
>>>>> have moved them outside NiFi and made every operator build and manage that
>>>>> integration themselves.
>>>>> 
>>>>> There are also costs to that approach: additional network boundaries,
>>>>> deployment artifacts, authentication, service discovery, observability,
>>>>> retry semantics, and operational complexity. From a governance 
>>>>> perspective,
>>>>> NiFi may actually have less visibility and control over that external
>>>>> Python workload than it would if Python execution remained part of a
>>>>> defined NiFi extension model.
>>>>> 
>>>>> That makes me wonder whether the better direction is to reconsider the
>>>>> Python runtime architecture rather than Python support itself.
>>>>> 
>>>>> For example, instead of treating the current Py4J-based implementation as
>>>>> something that must be preserved indefinitely, we could define a stronger
>>>>> process boundary around Python execution. A future implementation could 
>>>>> use
>>>>> a well-defined IPC or RPC contract, potentially based on gRPC/protobuf,
>>>>> Unix domain sockets, shared memory, Arrow, or another mechanism designed
>>>>> around coarse-grained operations.
>>>>> 
>>>>> NiFi does not need to become a Python integrated application, but I think
>>>>> there is substantial value in allowing those workloads to participate
>>>>> directly in a governed NiFi dataflow without requiring users to construct 
>>>>> a
>>>>> separate microservice architecture for every Python integration.
>>>>> 
>>>>> For that reason, I would favor separating two questions:
>>>>> 
>>>>>   1. Is the current Python Processor implementation the architecture we
>>>>>   want to support long term?
>>>>>   2. Should NiFi continue to provide a first-class Python extension
>>>>>   capability?
>>>>> 
>>>>> I could see a strong argument that the answer to the first question is no,
>>>>> particularly if Py4J is becoming a maintenance risk. I am much less
>>>>> convinced that the answer to the second question should also be no.
>>>>> 
>>>>> With that said I distinctly understand that time is limited and we would
>>>>> all develop the coolest of things with infinite resources, so I yield to
>>>>> those with more influence, but would happy to contribute to a better
>>>>> architecture if others find it valuable.
>>>>> 
>>>>> Thanks,
>>>>> 
>>>>> Marc
>>>>> 
>>>>> On Wed, Aug 12, 2026 at 12:10 AM Joe Witt <[email protected]> wrote:
>>>>> 
>>>>>> You are totally fine to engage on this thread Sean. We appreciate it
>>>>>> 
>>>>>> 
>>>>>> On Tue, Aug 11, 2026 at 8:17 PM Sean Hunter via dev <[email protected]
>>>>>> 
>>>>>> wrote:
>>>>>> 
>>>>>>> Hi David,
>>>>>>> 
>>>>>>> As a user, rather than contributor, I think my primary concern is one
>>>>> of
>>>>>>> operations and implementation. What alternative would be available? We
>>>>>> used
>>>>>>> the 1.x version and were quite excited to move to 2.0 with its more
>>>>>> native
>>>>>>> integration that was easier to troubleshoot. We also do build a NAR as
>>>>>> part
>>>>>>> of our pipeline process, and I suspect since we used a blog as guidance
>>>>>> to
>>>>>>> develop that process, that others likely do as well. Our use case is a
>>>>>>> particularly nasty data transform (Windows logs of some sort,
>>>>> naturally)
>>>>>>> that’s reasonable to write and execute in Python. It’s readable, while
>>>>>>> still being performant enough for our needs, and seemingly not causing
>>>>>>> system stability issues for us.
>>>>>>> 
>>>>>>> I do understand and take your points. What I’d like to better
>>>>> understand
>>>>>>> are the options users have. Do you have any thoughts?
>>>>>>> 
>>>>>>> Thank you and have a great day!
>>>>>>> Sean
>>>>>>> 
>>>>>>> PS - Please let me know if this is too far off topic for the dev
>>>>> mailing
>>>>>>> list and I’ll take it elsewhere.
>>>>>>> 
>>>>>>> From: David Handermann <[email protected]>
>>>>>>> Date: Tuesday, August 11, 2026 at 8:38 PM
>>>>>>> To: [email protected] <[email protected]>
>>>>>>> Subject: [DISCUSS] Future Support for Python Processors?
>>>>>>> 
>>>>>>> Be Careful: Email is from an External Source
>>>>>>> Use caution when opening emails and do not click on links or
>>>>> attachments
>>>>>>> from unknown senders. For more resources, visit
>>>>>> http://security.cat.com/phishing .
>>>>>>> 
>>>>>>> 
>>>>>>> Team,
>>>>>>> 
>>>>>>> The introduction of support for Python Processors was one of the major
>>>>>>> new features of Apache NiFi 2. In light of experience, however, I
>>>>>>> believe it is time to reconsider future support for this extension
>>>>>>> strategy. The idea of removing this feature may raise some concerns,
>>>>>>> but it is important to consider the options based on scalability,
>>>>>>> security, and supportability.
>>>>>>> 
>>>>>>> One of the greatest benefits of Python-based Processors is rapid
>>>>>>> iteration. Although this is a strength during the development phase,
>>>>>>> it can become a serious weakness when moving to production.
>>>>>>> 
>>>>>>> NiFi is fundamentally a Java-based application, thus tuning resource
>>>>>>> allocation is a key part of maintaining a production system. Running a
>>>>>>> Python Processor requires one or more independent Python processes,
>>>>>>> each with their own memory consumption and CPU usage patterns. What
>>>>>>> might work at a small scale on a development machine breaks down when
>>>>>>> attempting to run at production scale, without careful tuning. This is
>>>>>>> challenging enough with Java-based components, but global settings
>>>>>>> like heap and non-heap memory provide clear boundaries. Python
>>>>>>> Processors, as currently implemented, do not support this level of
>>>>>>> control. Beyond process control, operations that require communicating
>>>>>>> between the NiFi JVM and Python processes involve socket
>>>>>>> communication, which can be expensive for frequent calls. Thus an
>>>>>>> apparent initial advantage of Python Processors becomes a major
>>>>>>> problem when moving beyond limited types of operations.
>>>>>>> 
>>>>>>> When it comes to security, Python Processors can involve more serious
>>>>>>> risks than bundled Java extensions.
>>>>>>> 
>>>>>>> With NiFi support for downloading dependencies at runtime, it is
>>>>>>> essential for an administrator to provide a secure supply chain
>>>>>>> strategy for retrieving these dependencies. Although the NiFi
>>>>>>> framework also supports packaging Python Processors as NAR bundles,
>>>>>>> this packaging capability is not currently part of the Apache NiFi
>>>>>>> project. The project could build on third party work, but as Python
>>>>>>> packaging is less opinionated, this is an additional maintenance
>>>>>>> question that quickly becomes more involved. This combination of
>>>>>>> supply chain security and lack of packaging support raises security
>>>>>>> challenges.
>>>>>>> 
>>>>>>> Supportability is always a question for any feature in an open source
>>>>>>> project, but Python Processors are a particular point of concern.
>>>>>>> Providing NiFi framework support requires a strong knowledge of the
>>>>>>> NiFi project itself, and a strong knowledge of Python. Project
>>>>>>> maintainers and contributors usually have expertise in particular
>>>>>>> areas, but it is rare to find both. More to the point, however, the
>>>>>>> Py4J project, which NiFi depends on for integration, has seen limited
>>>>>>> maintenance, with the last tagged version dating to January 2025.
>>>>>>> Without robust architecture support, the NiFi project cannot provide a
>>>>>>> strong framework for future implementation.
>>>>>>> 
>>>>>>> Taking these issues together, I believe the project should consider
>>>>>>> deprecating support for Python Processors. Similar to other
>>>>>>> deprecations, actual removal would not occur until a major release
>>>>>>> version, and that would be a separate conversation.
>>>>>>> 
>>>>>>> Rather than continuing to extend the life of Python Processor support,
>>>>>>> it seems better to invest limited project resources into maintaining
>>>>>>> and improving core functionality.
>>>>>>> 
>>>>>>> Regards,
>>>>>>> David Handermann
>>>>>>> Apache NiFi PMC Chair
>>>>>>> 
>>>>>>> 
>>>>>> 
>>>>> 
>>> 

Reply via email to