Hi Frank,

Thank you for the feedback and guidance.


This direction sounds good to me. I will follow up with the TsFile community to 
confirm the initial scope, maintenance responsibility, and compatibility 
testing expectations from our side.


As a first step, I plan to work on a small prototype based on a TsFile 
`InputFormat` and `InputEntityReader`, and reuse existing Druid `InputSource` 
implementations where possible.


Once the prototype is ready, I will include sample data, usage examples, 
dependency/license information, and a basic test plan as suggested, and share 
it with the Druid community for further review.


Thanks again.


Best regards,
Colin











At 2026-08-22 13:31:58, "Frank Chen" <[email protected]> wrote:
>Hi Colin,
>
>I think this file format integration is a good candidate for the
>community extension.
>The extensions-contrib area is an appropriate initial home, provided the
>TsFile community takes responsibility for maintenance and compatibility
>testing.
>Generally community extensions are maintained by their contributors rather
>than by Druid committers.
>
>Based on current design and implementation, an inputFormat and
>InputEntityReader for tsFile are enough to demonstrate this integration.
>Once a small prototype is ready, please include a sample TsFile, ingestion
>spec, expected output rows, supported TsFile version, dependency/license
>details, and the proposed test matrix.
>
>Thanks and regards.
>
>On Fri, Aug 14, 2026 at 3:49 PM ColinLee <[email protected]> wrote:
>
>>
>>
>> Hello Druid community,
>>
>>
>> My name is Colin, and I am from the Apache TsFile community.
>>
>>
>> We would like to discuss adding support for the Apache TsFile data format
>> to Apache Druid.
>>
>>
>> Apache TsFile is a columnar file format designed for time-series data and
>> is used as the underlying storage file format by Apache IoTDB. It can also
>> be integrated independently into other data-processing systems. TsFile
>> provides implementations in Java, C, C++, and Python. It supports chunked
>> storage, compression, encoding, time-range filtering, and column pruning
>> for efficient storage and access to time-series data.
>>
>>
>> Apache TsFile project: <https://github.com/apache/tsfile>
>>
>>
>> We noticed that Druid already supports data formats and data sources such
>> as Avro, ORC, Parquet, and Iceberg through its extension mechanism. We
>> believe a similar mechanism could enable Druid to read and ingest TsFile
>> data.
>>
>>
>> Based on our current understanding of Druid's extension mechanism, one
>> possible initial approach would be to implement a TsFile input format
>> extension:
>>
>>
>> 1. Implement Druid's `InputFormat` interface.
>> 2. Implement an appropriate `InputEntityReader` for the binary TsFile
>> format.
>> 3. Use the Apache TsFile Java Reader to read the files.
>> 4. Map a configured TsFile time column to Druid's timestamp field.
>> 5. Convert the remaining TsFile fields into rows consumed by Druid's
>> ingestion pipeline.
>> 6. Initially support batch ingestion of TsFile files from local
>> filesystems or object storage.
>>
>>
>> The initial version could focus on the core TsFile batch-ingestion path
>> and consider the following capabilities:
>>
>>
>> - Configuring the TsFile column to use as the Druid timestamp;
>> - Basic data type mapping;
>> - Column selection and time-range filtering;
>> - Basic schema discovery;
>> - Single-file and multi-file batch ingestion.
>>
>>
>> Before starting the implementation, we would appreciate feedback from the
>> Druid community on the following questions:
>>
>>
>> 1. Would the Druid community be interested in support for TsFile as an
>> input format or data source?
>> 2. For local files and object storage already supported by existing
>> `InputSource` implementations, would it be appropriate to reuse those
>> implementations and add a TsFile `InputFormat`? If TsFile requires
>> format-specific file discovery or split planning, would a dedicated
>> `InputSource` also be necessary?
>> 3. If we proceed with the implementation, would this extension be suitable
>> for the `extensions-contrib` directory in the `apache/druid` repository?
>> 4. Do you have any other recommendations regarding the initial scope or
>> data-model mapping?
>>
>>
>> If the community considers this direction reasonable, we are willing to
>> take responsibility for implementing, testing, and maintaining the
>> extension. We would also like to discuss with the Druid community where the
>> code should be hosted and maintained before we begin the implementation.
>>
>>
>> As a related reference, the DuckDB TsFile extension we recently
>> implemented has been accepted into DuckDB Community Extensions and can
>> query TsFile data directly:
>>
>>
>> - Extension repository: <https://github.com/ColinLeeo/tsfile-duckdb>
>> - Merged DuckDB Community Extensions PR: <
>> https://github.com/duckdb/community-extensions/pull/2435>
>>
>>
>> We look forward to hearing your feedback and suggestions.
>>
>>
>> Best regards,
>>
>>
>> Colin Lee
>>
>>
>> Apache TsFile
>>
>>

Reply via email to