wangjingyu11 opened a new issue, #67619: URL: https://github.com/apache/doris/issues/67619
### Search before asking - [x] I had searched in the [issues](https://github.com/apache/doris/issues?q=is%3Aissue) and found no similar issues. ### Version ``` doris-4.0.8-rc02 Git commit: bc8ea1bac6d ``` The `4.0.8` and `4.0.8-rc02` tags both dereference to commit `bc8ea1bac6d`, so this is the official 4.0.8 GA build. - JDK: OpenJDK 17.0.17 - OS: Linux x86_64, RHEL 9 family, kernel 5.14 - CPU: AMD EPYC (family 19h model 10h) - RAM: 251 GiB physical; BE `mem_limit` = 90% = 225.99 GB - Storage: 4 local SSD data dirs (`storage_root_path`, `medium:SSD`), ~14 TB free - Deployment: 1 FE + 1 BE, single node, local (non-cloud) mode - `be.conf`: stock defaults apart from ports / `priority_networks` / `storage_root_path` ### What's Wrong? During a bulk `INSERT INTO ... SELECT` backfill into **Unique Key + Merge-on-Write** tables, the BE crashed **three times in ~2 hours** with `SIGSEGV` at fault address `0x8`. Every crash happens **while destroying elements of a container during cleanup**, but in **three unrelated subsystems**: | # | Time | Faulting frame | Subsystem | Process RSS at crash | |---|---|---|---|---| | A | T+0h | `ColumnReaderCache::~ColumnReaderCache` → `_List_base<CacheNode>::_M_destroy_node` → `shared_ptr` dispose | Segment cache | 54.3 GB / 226 GB | | B | T+1h | `LRUCache::insert` (`lru_cache.cpp:530`, inlined `delete (LRUCacheValueBase*)value`) | Storage page cache | 43.0 GB / 226 GB | | C | T+2h | `DorisRAMFSDirectory::~DorisRAMFSDirectory` → `__CLMap::clear()` → `RAMFile::~RAMFile` | Inverted index (CLucene) | **27.4 GB** / 226 GB | All three faulted at `@0x8` and all three are on the **tablet write / close path** (`TabletsChannel::close`, `calc_delete_bitmap*`, `BetaRowsetWriter::build`). Crashes A and B were mitigated by disabling the two caches involved (see *Anything Else*). The load then progressed further and crashed at C instead — a **different subsystem** that does not use `LRUCache` at all. This "whack-a-mole" behaviour is why I suspect a shared underlying lifetime / heap issue on the MoW write path rather than three independent bugs. #### Crash A — SegmentCache ``` *** SIGSEGV address not mapped to object (@0x8) received by PID <pid> (TID <tid>) *** 0# doris::signal::(anonymous namespace)::FailureSignalHandler(int, siginfo_t*, void*) at be/src/common/signal_handler.h:420 1# PosixSignals::chained_handler(int, siginfo_t*, void*) [clone .part.0] in libjvm.so 2# JVM_handle_linux_signal in libjvm.so 3# 0x00007F... in /lib64/libc.so.6 4# std::_Sp_counted_ptr<doris::segment_v2::ColumnReader*, (__gnu_cxx::_Lock_policy)2>::_M_dispose() at bits/shared_ptr_base.h:427 5# std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release_last_use_cold() at bits/shared_ptr_base.h:198 6# std::__cxx11::_List_base<doris::segment_v2::CacheNode, ...>::_M_destroy_node(...) at bits/stl_list.h:845 7# doris::segment_v2::ColumnReaderCache::~ColumnReaderCache() at be/src/olap/rowset/segment_v2/column_reader_cache.cpp:47 9# doris::segment_v2::Segment::~Segment() at be/src/olap/rowset/segment_v2/segment.cpp:237 11# std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release_last_use_cold() at bits/shared_ptr_base.h:198 12# doris::SegmentCache::CacheValue::~CacheValue() at be/src/olap/segment_loader.h:78 13# doris::LRUCache::insert(...) at be/src/olap/lru_cache.cpp:530 14# doris::SegmentCache::insert(...) at be/src/olap/segment_loader.cpp:46 15# doris::SegmentLoader::load_segment(...) at be/src/olap/segment_loader.cpp:83 16# doris::SegmentLoader::load_segments(...) at be/src/olap/segment_loader.cpp:94 17# doris::BaseTablet::lookup_row_key(...) at be/src/olap/base_tablet.cpp:507 18# doris::BaseTablet::calc_segment_delete_bitmap(...) at be/src/olap/base_tablet.cpp:712 19# doris::BaseTablet::calc_delete_bitmap(...) at be/src/olap/base_tablet.cpp:569 20# doris::CalcDeleteBitmapToken::submit_func<doris::BaseBetaRowsetWriter::_generate_delete_bitmap(int)::$_0>(...) at bits/std_function.h:292 21# doris::ThreadPool::dispatch_thread() at be/src/util/threadpool.cpp:623 22# doris::Thread::supervise_thread(void*) at be/src/util/thread.cpp:461 ``` #### Crash B — StoragePageCache (primary key index page) ``` *** SIGSEGV address not mapped to object (@0x8) received by PID <pid> (TID <tid>) *** 4# doris::LRUCache::insert(...) at be/src/olap/lru_cache.cpp:530 5# doris::StoragePageCache::insert(..., doris::MemoryTrackedPageWithPageEntity*, ...) at be/src/olap/page_cache.cpp:122 6# doris::segment_v2::PageIO::read_and_decompress_page_(...) at be/src/olap/rowset/segment_v2/page_io.cpp:264 7# doris::segment_v2::PageIO::do_read_and_decompress_page(...) at be/src/olap/rowset/segment_v2/page_io.h:162 8# doris::segment_v2::PageIO::read_and_decompress_page(...) at be/src/olap/rowset/segment_v2/page_io.cpp:276 9# doris::segment_v2::IndexedColumnReader::read_page(...) const at be/src/olap/rowset/segment_v2/indexed_column_reader.cpp:144 10# doris::segment_v2::IndexedColumnIterator::_read_data_page(...) at be/src/olap/rowset/segment_v2/indexed_column_reader.cpp:165 11# doris::segment_v2::IndexedColumnIterator::next_batch(...) at be/src/olap/rowset/segment_v2/indexed_column_reader.cpp:304 12# doris::MergeIndexDeleteBitmapCalculatorContext::_next_batch(unsigned int) at be/src/olap/delete_bitmap_calculator.cpp:81 13# doris::MergeIndexDeleteBitmapCalculatorContext::seek_at_or_after(doris::Slice const&) at be/src/olap/delete_bitmap_calculator.cpp:71 14# doris::MergeIndexDeleteBitmapCalculator::calculate_one(doris::RowLocation&) at be/src/olap/delete_bitmap_calculator.cpp:208 15# doris::MergeIndexDeleteBitmapCalculator::calculate_all(std::shared_ptr<doris::DeleteBitmap>) at be/src/olap/delete_bitmap_calculator.cpp:218 16# doris::BaseTablet::calc_delete_bitmap_between_segments(...) at be/src/olap/base_tablet.cpp:384 17# doris::BaseRowsetBuilder::submit_calc_delete_bitmap_task() in doris_be 18# doris::BaseDeltaWriter::submit_calc_delete_bitmap_task() at be/src/olap/delta_writer.cpp:196 19# doris::TabletsChannel::close(...) at be/src/runtime/tablets_channel.cpp:350 20# doris::LoadChannel::_handle_eos(...) at be/src/runtime/load_channel.cpp:221 22# doris::LoadChannelMgr::add_batch(...) at be/src/runtime/load_channel_mgr.cpp:179 24# doris::WorkThreadPool<false>::work_thread(int) at be/src/util/work_thread_pool.hpp:159 ``` #### Crash C — inverted index RAM directory (with both caches already disabled) ``` *** SIGSEGV address not mapped to object (@0x8) received by PID <pid> (TID <tid>) *** 4# lucene::store::RAMFile::~RAMFile() at contrib/clucene/src/core/CLucene/store/RAMDirectory.cpp:50 5# lucene::util::__CLMap<char const*, lucene::store::RAMFile*, ...>::clear() at contrib/clucene/src/core/CLucene/util/VoidMap.h:127 6# doris::segment_v2::DorisRAMFSDirectory::~DorisRAMFSDirectory() at be/src/olap/rowset/segment_v2/inverted_index_fs_directory.cpp:740 8# std::_Rb_tree<std::pair<long, std::string>, ..., std::shared_ptr<lucene::store::Directory>>::_M_erase(...) at bits/stl_tree.h:2590 10# doris::segment_v2::IndexFileWriter::finish_close() at be/src/olap/rowset/segment_v2/index_file_writer.cpp:253 11# doris::InvertedIndexFileCollection::finish_close() at be/src/olap/rowset/beta_rowset_writer.cpp:282 12# doris::SegmentFlusher::close() at be/src/olap/rowset/segment_creator.cpp:120 13# doris::SegmentCreator::close() at be/src/olap/rowset/segment_creator.cpp:472 14# doris::BaseBetaRowsetWriter::_close_file_writers() at be/src/olap/rowset/beta_rowset_writer.cpp:923 16# doris::BetaRowsetWriter::build(std::shared_ptr<doris::Rowset>&) at be/src/olap/rowset/beta_rowset_writer.cpp:972 17# doris::BaseRowsetBuilder::build_rowset() at be/src/olap/rowset_builder.cpp:248 18# doris::BaseDeltaWriter::build_rowset() at be/src/olap/delta_writer.cpp:186 20# doris::TabletsChannel::close(...) at be/src/runtime/tablets_channel.cpp:344 21# doris::LoadChannel::_handle_eos(...) at be/src/runtime/load_channel.cpp:221 23# doris::LoadChannelMgr::add_batch(...) at be/src/runtime/load_channel_mgr.cpp:179 25# doris::WorkThreadPool<false>::work_thread(int) at be/src/util/work_thread_pool.hpp:159 ``` `index_file_writer.cpp:253` is `_indices_dirs.clear();` — destroying the `std::map<..., std::shared_ptr<lucene::store::Directory>>` of per-index RAM directories. #### Notes on crash B `LRUCache::insert()` unlinks evicted entries under `_mutex`, then frees them outside it: ```cpp // be/src/olap/lru_cache.cpp (4.0.8, lines 526-532) // we free the entries here outside of mutex for // performance reasons while (to_remove_head != nullptr) { LRUHandle* next = to_remove_head->next; to_remove_head->free(); // <-- faulting line to_remove_head = next; } ``` `LRUHandle::free()` is inline in `lru_cache.h` and explicitly null-checks `value` (`if (value != nullptr) delete (LRUCacheValueBase*)value;`), so a plain null `value` does not explain the fault. In crash A the fault lands two frames deeper, in a `shared_ptr` control-block release — a `shared_ptr`'s control-block pointer sits at offset `0x8`, which matches the fault address. This `insert()` free-outside-mutex block is byte-identical in `branch-3.0` (line 453), `branch-3.1` (line 518) and `branch-4.0` (line 530). #### Ruled out | Hypothesis | Evidence against | |---|---| | OOM / memory limit | RSS at the three crashes: 54.3 / 43.0 / **27.4** GB against a 226 GB `mem_limit`; 170 GB system memory free. Highest query `PeakMemoryBytes` was 12.7 GB. | | Linux OOM killer | No kernel records; the process caught SIGSEGV in its own handler and unwound over several seconds. | | ECC / memory hardware | `amd64_edac` (F19h_M10h) loaded and monitoring; **zero** MCE/EDAC events in `dmesg` and `journalctl -k` across the whole uptime. | | Disk / I/O error | Zero `corrupt` / `checksum` / `Input/output error` in BE logs; zero SCSI/HBA errors in the kernel log. | | Query concurrency | Statements ran strictly **serially** — verified in `fe.audit.log`, each starts only after the previous finishes. | | Cache capacity auto-adjust thread | Process memory stayed at ~12-24% of `mem_limit`, far below the default `cache_capacity_reduce_mem_limit_frac` (0.6). | | Version accumulation / compaction backlog | Zero `too many versions` / `-235`; delete-bitmap batches ~416 K rows with `conflict rows: 0`. | | Non-GA build | `4.0.8` and `4.0.8-rc02` both point to `bc8ea1bac6d`. | ### What You Expected? The BE should not `SIGSEGV`. Under heavy load it should degrade gracefully — cancel the query, return `MEM_LIMIT_EXCEEDED`, reject the write — not take the whole process down, especially while using only ~12% of its configured memory limit. ### How to Reproduce? Reproduced three times in about two hours on the same node. **Target schema** (52 tables of this shape; names/columns redacted): ```sql CREATE TABLE db1.tbl_mow ( id VARCHAR(...), -- key ... 7 to 42 value columns ..., seq_id BIGINT, INDEX idx_a (`some_col`) USING INVERTED, INDEX idx_b (`text_col_1`) USING INVERTED PROPERTIES("lower_case"="true","parser"="chinese","parser_mode"="fine_grained","support_phrase"="true"), INDEX idx_c (`text_col_2`) USING INVERTED PROPERTIES("lower_case"="true","parser"="chinese","parser_mode"="fine_grained","support_phrase"="true") ) ENGINE=OLAP UNIQUE KEY(`id`) DISTRIBUTED BY HASH(`id`) BUCKETS 48 PROPERTIES ( "replication_allocation" = "tag.location.default: 1", "enable_unique_key_merge_on_write" = "true", "function_column.sequence_col" = "seq_id", "compression" = "LZ4", "storage_format" = "V2", "inverted_index_storage_format" = "V1" ); ``` Source tables are append-style tables with the same column set. **Workload:** ```sql -- one statement at a time, strictly serial, no concurrency INSERT INTO db1.tbl_mow (col1, ..., colN, seq_id) SELECT col1, ..., colN, seq_id FROM db1.tbl_source; ``` - 52 statements, ~3.27 billion source rows (~414 GB) in total. - Targets start **empty**; each statement is one transaction. - Single BE, so all tablet writes, delete-bitmap computation, inverted-index building and compaction land on one process. - Largest single statements: 256 M and 260 M rows, each with two fine-grained Chinese inverted indexes over long text columns. The common condition is a sustained, high-throughput MoW load that keeps the segment cache, the page cache and the inverted-index writers all busy for hours. ### Anything Else? **Possibly the same root cause as #64826** (`[Bug] BE SIGSEGV crash in bthread::TaskGroup::sched_to and Page Cache on 4.0.5-rc01`, still open). That report also shows `SIGSEGV @0x8` while an `LRUCache` frees evicted entries, reached from `Daemon::cache_adjust_capacity_thread` → `CacheManager::for_each_cache_refresh_capacity` → `LRUCache::set_capacity`. `set_capacity()` frees evicted handles the same way `insert()` does. **Core dumps unusable.** `systemd-coredump` truncated them (BE RSS 27-54 GB vs default `ProcessSizeMax=2G`), so only the BE's own symbolized handler output is available — that is what is quoted above. Happy to reconfigure and capture a full core if that would help. **Mitigations found, each blocking one manifestation** (source-level; each `if` below is the gate that stops the crashing call from being reached): ``` disable_segment_cache = true # segment_loader.cpp:74 -> blocks crash A disable_pk_storage_page_cache = true # primary_key_index.cpp:103 -> page_io.cpp:260 -> blocks crash B disable_storage_page_cache = true # base_tablet.cpp:104, segment.cpp:992 inverted_index_ram_dir_enable = false # inverted_index_fs_directory.cpp:927 -> blocks crash C ``` ```cpp // be/src/olap/segment_loader.cpp:74 if (use_cache && !config::disable_segment_cache) { _segment_cache->insert(...); } // crash A else { cache_handle->push_segment(std::move(segment)); } // be/src/olap/rowset/segment_v2/page_io.cpp:260 if (opts.use_page_cache && cache) { cache->insert(...); } // crash B else { *handle = PageHandle(page.get()); } // be/src/olap/rowset/segment_v2/inverted_index_fs_directory.cpp:927 if (config::inverted_index_ram_dir_enable && can_use_ram_dir) { /* DorisRAMFSDirectory */ } // crash C ``` Note `disable_pk_storage_page_cache` is `DEFINE_Bool` (needs a BE restart) while the other three are `DEFINE_mBool`. After disabling the two caches, the load got substantially further (43 tables in 13 minutes, versus 2 tables in 105 minutes before), then hit crash C. **Each config removes one entry point; it does not appear to fix the underlying problem.** I still have the environment available and can gather more detail (full unredacted stacks, `SHOW DATA` output, audit-log timings, a full core dump) on request. ### Are you willing to submit PR? - [ ] Yes I am willing to submit a PR! ### Code of Conduct - [x] I agree to follow this project's [Code of Conduct](https://www.apache.org/foundation/policies/conduct) -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected] --------------------------------------------------------------------- To unsubscribe, e-mail: [email protected] For additional commands, e-mail: [email protected]
