Hi  Chen Li / Team

Thanks for your valuable feedback.

I have installed Texera on a remote Linux server and created the 1 GB test
file on that server. Since I am accessing the Texera UI from my Windows
machine, I manually used Texera's multipart upload API from the Linux
server to upload the file as a Texera dataset.

During the upload, I initially encountered a file size limit issue, which I
resolved by increasing the configured limit. After that, the file was
uploaded successfully.

During execution, I observed that the *File Scan operator was reading the
file in small chunks and writing the data in Iceberg format*. I increased
the chunk size and restarted the execution.

However, the *Python UDF operator*, which reads one line at a time, parses
it as a CSV row, converts it into six structured fields, and outputs the
tuple, took around *11 hours* to process the 27 million records.

I am using *3 worker nodes*, and the Python UDF was also running with *3
workers*.

Is this expected behavior for Texera, or am I missing any configuration or
optimization that could improve the Python UDF performance?

[image: image.png]


Thanks & Regards
Rahul

On Fri, Aug 28, 2026 at 11:17 AM Chen Li <[email protected]> wrote:

> Rahul,
>
> Sorry for the late reply.  Currently the CSVFileScan operator doesn't
> support multiple workers to run in parallel.  So reading a large file may
> take time.  If you are interested in developing a parallel reader, you are
> very welcome to do so and the team can provide support.
>
> The Python UDF operator allows a user to specify the number of workers.
>
> The backend engine is doing in-memory computation.  Data transfer between
> operators can be pipelined.
>
> The backend storage layer supports large files using solutions such as
> MinIO, LakeFS, and Iceberg.
>
> The system had done a lot of optimization in uploading large files
> efficiently (e.g., in the range of TBs).
>
> Our team will be happy to answer questions related to your particular use
> case.
>
> Thank you,
>
> Chen Li
>
>
>
>
> On Thu, Aug 27, 2026 at 9:17 PM Rahul P <[email protected]> wrote:
>
> > Hi Team
> >
> > Can I please get an update on this.
> >
> > Thanks & Regards
> > Rahul
> >
> > On Thu, Aug 27, 2026 at 12:09 PM Rahul P <[email protected]>
> wrote:
> >
> > > Hi Team,
> > >
> > > I have installed Texera and am able to execute workflows successfully.
> I
> > > would like to understand how large-scale data processing is handled in
> > > Texera.
> > >
> > > If the input data size is very large, will Texera process it reliably
> > > without crashing, or are there any known limitations related to file
> > size,
> > > memory, or cluster resources?
> > >
> > > I also could not find any option to configure parallelism for the file
> > > fetch/file scan operator. How can parallelism be configured for this
> > > operator?
> > >
> > > Thanks & Regards
> > > Rahul
> > >
> >
>

Reply via email to