Hi all,

I recently tested `TsFileDataFrame` for data loading on a relatively large 
TsFile directory. The directory is about `1.1 GB` and contains `45,148` logical 
time series. The workload is close to a training data loading pattern: randomly 
select one logical series from many time series, then read a window such as 
`[offset, offset + length)`.

During profiling, I found a noticeable performance bottleneck in the current 
`TsFileDataFrame` implementation for this kind of random window access pattern.

From the flame graph, a typical call chain looks like this:

```
Timeseries.__getitem__
  -> _read_field_by_position
  -> reader.read_series_by_row(...)
  -> query_table/query_tree result set
  -> read_arrow_batch
  -> TableResultSet::get_next_tsblock
  -> DeviceOrderedTsBlockReader::has_next
  -> SingleDeviceTsBlockReader::ini
  -> TsFileIOReader::load_device_index_entry
  -> search_from_internal_node
  -> MetaIndexNode::device_deserialize_from
```

In other words, the Python-level `TsFileDataFrame` already knows information 
such as the logical series, its length, and which file it belongs to. However, 
when reading each window, the C++ query layer still behaves more like executing 
a full small query:

```
read device + measurement + offset/limit
  -> locate device metadata index
  -> deserialize meta index node
  -> locate timeseries index / chunk metadata
  -> initialize scan iterator / chunk reader
  -> start decoding pages
```

In this workload, the major hotspots in the flame graph include:

```
TsFileIOReader::load_device_index_entry        41.6%
TsFileIOReader::search_from_internal_node      41.5%
MetaIndexNode::device_deserialize_from         34.9%
DeviceMetaIndexEntry::deserialize_from         27.6%
```

These hotspots suggest that the main cost of window reads is not only page 
decoding. A significant amount of time is also spent on device metadata index 
lookup, meta index node deserialization, and query/reader context 
initialization.

I currently think there are two layers of problems.

The first one is the problem inside a single TsFile file.

Within the same TsFile file, when repeatedly accessing different offsets of the 
same or different time series, the current path still repeatedly constructs the 
query / reader context. For the `TsFileDataFrame` access pattern, `device + 
measurement` is already known at the beginning, and the main changing 
parameters afterwards are only `offset` and `limit`. Therefore, repeatedly 
locating the device metadata index, loading measurement metadata, and 
initializing scan iterators introduces considerable overhead.

For this part, I plan to first try optimizing inside the single-file reader, 
for example:

```
- Cache device meta index nodes / measurement metadata;
- Store the resolved device + measurement as a reader-bound access handle;
- Reuse the existing series reader context when only offset/limit changes;
- Avoid constructing a full query for every window.
```

The second one is the problem at the multi-file / Dataset layer.

When `TsFileDataFrame` works over multiple TsFile files, the Python layer 
already maintains a global catalog: which file each logical series belongs to, 
its corresponding device / measurement, its length and time range, its data 
type, and the shard relationship between files.

However, this information currently mainly lives in the Python 
`MetadataCatalog` inside `TsFileDataFrame`. When the actual query is executed, 
the `tsfilecpp` reader cannot directly use the information that has already 
been parsed by the upper layer. In other words, the lower-level reader still 
behaves more like handling an independent TsFile query, instead of using the 
series/file metadata already known by `TsFileDataFrame` to restore a 
lightweight access context.

Therefore, in a random-access workload over multiple files, cache inside a 
single reader can only solve part of the problem. In the longer term, we may 
need a dataset/runtime layer that moves the catalog information maintained by 
`TsFileDataFrame` into a more compact native representation, and reuses this 
metadata when switching readers or accessing different files. This could reduce 
cold-start cost and repeated query construction.

My current plan is to split the optimization into two stages:

```
1. First solve the problem inside a single TsFile file:
   avoid reconstructing the query / reader context for every offset access.

2. Then solve the problem across multiple TsFile files:
   introduce a dataset runtime to manage the catalog, reader pool,
   access handle cache, and file access strategy.
```

I would like to discuss whether this direction makes sense, especially:

```
- For window-read workloads like TsFileDataFrame, should the reader layer 
provide a lighter per-series access handle?
- At which layer should device metadata / measurement metadata caches live?
- Should the catalog information already maintained by TsFileDataFrame have a 
native representation that can be reused by the lower-level reader/runtime?
- Is a multi-file dataset runtime more appropriate in the 
TsFileDataFrame/dataset layer rather than in the low-level TsFile reader?
- How can we improve access performance while still controlling the size of 
reader caches and file caches?
```

Any feedback or suggestions on this direction and the possible interface design 
would be very welcome.

Best,
Colin

Reply via email to