[ 
https://issues.apache.org/jira/browse/KAFKA-20979?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
 ]

Mickael Maison updated KAFKA-20979:
-----------------------------------
    Affects Version/s: 4.1.0

> Potential data loss caused by time index truncation
> ---------------------------------------------------
>
>                 Key: KAFKA-20979
>                 URL: https://issues.apache.org/jira/browse/KAFKA-20979
>             Project: Kafka
>          Issue Type: Bug
>    Affects Versions: 4.1.0
>            Reporter: Mickael Maison
>            Assignee: Mickael Maison
>            Priority: Major
>
> On broker startup, I see these lines for many partitions:
> {noformat}
> [LocalLog partition=TOPIC-0, dir=LOG_DIR/TOPIC-0] Rolled new log segment at 
> offset 83461 in 0 ms.
> [UnifiedLog partition=TOPIC-0, dir=LOG_DIR] Incremented log start offset to 
> 83461 due to segment deletion
> [UnifiedLog partition=TOPIC-0, dir=LOG_DIR] Deleting segment 
> LogSegment(baseOffset=83454, size=4483, lastModifiedTime=1786538669095, 
> largestRecordTimestamp=0) due to log retention time 604800000ms breach based 
> on the largest record timestamp in the segment
> {noformat}
> It's deleting many segments due to the retention limits. This has happened 
> multiple times in the past few weeks. Using DumpLogSegments (on followers and 
> by using file.delete.delay.ms to delay the actual deletion) I've been able to 
> confirm that the retention time was not breached and that all records have 
> valid timestamps (I first suspected some records with bad timestamps).
> In all cases it happened on a clean shutdown. 
> Being unable to explain why a broker would incorrectly compute 
> largestRecordTimestamp=0 I looked at the code and I think this behavior could 
> be caused by invalid time index truncation.
> During clean shutdown, LogManager calls flush(true) then close() on each log. 
> The flush() fsyncs the time index file at its pre-allocated size (10MB by 
> default, zero-filled). The close() truncates the file to its actual size in 
> AbstractIndex.resize() which calls RandomAccessFile.setLength(), but never 
> fsyncs the truncation.
> If the broker restarts before the OS flushes the file metadata, the time 
> index could revert to its fsynced pre-allocated state. The zero-filled tail 
> is then read as valid entries with timestamp=0. Then when time-based 
> retention triggers, it deletes the segment, incorrectly deleting valid, 
> recent records.



--
This message was sent by Atlassian Jira
(v8.20.10#820010)

Reply via email to