Hi all, I recently tested `TsFileDataFrame` for data loading on a relatively large TsFile directory. The directory is about `1.1 GB` and contains `45,148` logical time series. The workload is close to a training data loading pattern: randomly select one logical series from many time series, then read a window such as `[offset, offset + length)`.
During profiling, I found a noticeable performance bottleneck in the current `TsFileDataFrame` implementation for this kind of random window access pattern. From the flame graph, a typical call chain looks like this: ``` Timeseries.__getitem__ -> _read_field_by_position -> reader.read_series_by_row(...) -> query_table/query_tree result set -> read_arrow_batch -> TableResultSet::get_next_tsblock -> DeviceOrderedTsBlockReader::has_next -> SingleDeviceTsBlockReader::ini -> TsFileIOReader::load_device_index_entry -> search_from_internal_node -> MetaIndexNode::device_deserialize_from ``` In other words, the Python-level `TsFileDataFrame` already knows information such as the logical series, its length, and which file it belongs to. However, when reading each window, the C++ query layer still behaves more like executing a full small query: ``` read device + measurement + offset/limit -> locate device metadata index -> deserialize meta index node -> locate timeseries index / chunk metadata -> initialize scan iterator / chunk reader -> start decoding pages ``` In this workload, the major hotspots in the flame graph include: ``` TsFileIOReader::load_device_index_entry 41.6% TsFileIOReader::search_from_internal_node 41.5% MetaIndexNode::device_deserialize_from 34.9% DeviceMetaIndexEntry::deserialize_from 27.6% ``` These hotspots suggest that the main cost of window reads is not only page decoding. A significant amount of time is also spent on device metadata index lookup, meta index node deserialization, and query/reader context initialization. I currently think there are two layers of problems. The first one is the problem inside a single TsFile file. Within the same TsFile file, when repeatedly accessing different offsets of the same or different time series, the current path still repeatedly constructs the query / reader context. For the `TsFileDataFrame` access pattern, `device + measurement` is already known at the beginning, and the main changing parameters afterwards are only `offset` and `limit`. Therefore, repeatedly locating the device metadata index, loading measurement metadata, and initializing scan iterators introduces considerable overhead. For this part, I plan to first try optimizing inside the single-file reader, for example: ``` - Cache device meta index nodes / measurement metadata; - Store the resolved device + measurement as a reader-bound access handle; - Reuse the existing series reader context when only offset/limit changes; - Avoid constructing a full query for every window. ``` The second one is the problem at the multi-file / Dataset layer. When `TsFileDataFrame` works over multiple TsFile files, the Python layer already maintains a global catalog: which file each logical series belongs to, its corresponding device / measurement, its length and time range, its data type, and the shard relationship between files. However, this information currently mainly lives in the Python `MetadataCatalog` inside `TsFileDataFrame`. When the actual query is executed, the `tsfilecpp` reader cannot directly use the information that has already been parsed by the upper layer. In other words, the lower-level reader still behaves more like handling an independent TsFile query, instead of using the series/file metadata already known by `TsFileDataFrame` to restore a lightweight access context. Therefore, in a random-access workload over multiple files, cache inside a single reader can only solve part of the problem. In the longer term, we may need a dataset/runtime layer that moves the catalog information maintained by `TsFileDataFrame` into a more compact native representation, and reuses this metadata when switching readers or accessing different files. This could reduce cold-start cost and repeated query construction. My current plan is to split the optimization into two stages: ``` 1. First solve the problem inside a single TsFile file: avoid reconstructing the query / reader context for every offset access. 2. Then solve the problem across multiple TsFile files: introduce a dataset runtime to manage the catalog, reader pool, access handle cache, and file access strategy. ``` I would like to discuss whether this direction makes sense, especially: ``` - For window-read workloads like TsFileDataFrame, should the reader layer provide a lighter per-series access handle? - At which layer should device metadata / measurement metadata caches live? - Should the catalog information already maintained by TsFileDataFrame have a native representation that can be reused by the lower-level reader/runtime? - Is a multi-file dataset runtime more appropriate in the TsFileDataFrame/dataset layer rather than in the low-level TsFile reader? - How can we improve access performance while still controlling the size of reader caches and file caches? ``` Any feedback or suggestions on this direction and the possible interface design would be very welcome. Best, Colin
