[
https://issues.apache.org/jira/browse/KAFKA-20979?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]
Mickael Maison updated KAFKA-20979:
-----------------------------------
Affects Version/s: 4.1.0
> Potential data loss caused by time index truncation
> ---------------------------------------------------
>
> Key: KAFKA-20979
> URL: https://issues.apache.org/jira/browse/KAFKA-20979
> Project: Kafka
> Issue Type: Bug
> Affects Versions: 4.1.0
> Reporter: Mickael Maison
> Assignee: Mickael Maison
> Priority: Major
>
> On broker startup, I see these lines for many partitions:
> {noformat}
> [LocalLog partition=TOPIC-0, dir=LOG_DIR/TOPIC-0] Rolled new log segment at
> offset 83461 in 0 ms.
> [UnifiedLog partition=TOPIC-0, dir=LOG_DIR] Incremented log start offset to
> 83461 due to segment deletion
> [UnifiedLog partition=TOPIC-0, dir=LOG_DIR] Deleting segment
> LogSegment(baseOffset=83454, size=4483, lastModifiedTime=1786538669095,
> largestRecordTimestamp=0) due to log retention time 604800000ms breach based
> on the largest record timestamp in the segment
> {noformat}
> It's deleting many segments due to the retention limits. This has happened
> multiple times in the past few weeks. Using DumpLogSegments (on followers and
> by using file.delete.delay.ms to delay the actual deletion) I've been able to
> confirm that the retention time was not breached and that all records have
> valid timestamps (I first suspected some records with bad timestamps).
> In all cases it happened on a clean shutdown.
> Being unable to explain why a broker would incorrectly compute
> largestRecordTimestamp=0 I looked at the code and I think this behavior could
> be caused by invalid time index truncation.
> During clean shutdown, LogManager calls flush(true) then close() on each log.
> The flush() fsyncs the time index file at its pre-allocated size (10MB by
> default, zero-filled). The close() truncates the file to its actual size in
> AbstractIndex.resize() which calls RandomAccessFile.setLength(), but never
> fsyncs the truncation.
> If the broker restarts before the OS flushes the file metadata, the time
> index could revert to its fsynced pre-allocated state. The zero-filled tail
> is then read as valid entries with timestamp=0. Then when time-based
> retention triggers, it deletes the segment, incorrectly deleting valid,
> recent records.
--
This message was sent by Atlassian Jira
(v8.20.10#820010)