mrhhsg commented on code in PR #68032:
URL: https://github.com/apache/doris/pull/68032#discussion_r4119379351
##########
cloud/src/recycler/recycler.cpp:
##########
@@ -7828,6 +7835,50 @@ int InstanceRecycler::recycle_expired_stage_objects() {
return ret;
}
+int InstanceRecycler::recycle_expired_spill_objects() {
+ LOG_INFO("begin to recycle expired spill objects").tag("instance_id",
instance_id_);
+
+ int64_t start_time =
duration_cast<seconds>(steady_clock::now().time_since_epoch()).count();
+ RecyclerMetricsContext metrics_context(instance_id_,
"recycle_expired_spill_objects");
+
+ DORIS_CLOUD_DEFER {
+ int64_t cost =
+
duration_cast<seconds>(steady_clock::now().time_since_epoch()).count() -
start_time;
+ metrics_context.finish_report();
+ LOG_INFO("recycle expired spill objects, cost={}s",
cost).tag("instance_id", instance_id_);
+ };
+
+ int64_t expiration_time =
Review Comment:
Obsolete: the recycler spill sweep was removed in 422b839. Remote spill is
deleted only when its SpillFile is destroyed; crash residue is left to a bucket
lifecycle rule (documented at `spill_storage_type` in config.cpp).
##########
be/src/exec/spill/spill_file_manager.cpp:
##########
@@ -308,6 +438,99 @@ void SpillFileManager::gc(int32_t max_work_time_ms) {
}
}
+void SpillFileManager::_remote_gc(SpillDataDir* store) {
+ if (!store->ready()) {
+ // Retry about once a minute at the default 2s GC interval.
ensure_ready() only reads
+ // what the vault refresh thread and the FE heartbeat already brought
in.
+ if (_remote_not_ready_rounds++ % 30 != 0) {
+ return;
+ }
+ auto st = store->ensure_ready();
+ if (!st.ok()) {
+ LOG(WARNING) << "remote spill store is not ready yet: " << st;
+ return;
+ }
+ }
+ _report_remote_spill_stats(store);
+ if (!remote_startup_cleanup_pending()) {
+ return;
+ }
+ auto st = _remote_startup_cleanup(store);
+ if (st.ok()) {
+ _remote_startup_cleanup_pending.store(false,
std::memory_order_release);
+ } else {
+ LOG_EVERY_T(WARNING, 60) << "failed to clean up spill objects of
previous boots, will "
+ "retry: "
+ << st;
+ }
+}
+
+void SpillFileManager::flush_remote_spill_stats() {
+ for (auto& [path, store] : _spill_store_map) {
+ if (store->is_remote()) {
+ _report_remote_spill_stats(store.get(), /*final_report=*/true);
+ }
+ }
+}
+
+void SpillFileManager::_report_remote_spill_stats(SpillDataDir* store, bool
final_report) {
+ std::lock_guard<std::mutex> lock(_remote_report_mutex);
+ // About once a minute at the default 2s GC interval; a final report skips
the cadence.
Review Comment:
Obsolete: cumulative upload counters are no longer reported to the
meta-service. SHOW DATA shows the bytes currently held in object storage,
polled from the live BEs, so there is no cumulative total to lose.
##########
cloud/src/meta-service/meta_service.cpp:
##########
@@ -3530,6 +3530,143 @@ void
MetaServiceImpl::get_tablet_stats(::google::protobuf::RpcController* contro
}
}
+void MetaServiceImpl::report_spill_stats(::google::protobuf::RpcController*
controller,
+ const ReportSpillStatsRequest*
request,
+ ReportSpillStatsResponse* response,
+ ::google::protobuf::Closure* done) {
+ RPC_PREPROCESS(report_spill_stats, put);
Review Comment:
Obsolete: the meta-service `report_spill_stats` / `get_spill_stats` RPCs
were removed. SHOW DATA now polls the spill size from the BEs
(`RemoteSpillStatsPoller`), and `cloud/` is identical to master since 422b839.
--
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.
To unsubscribe, e-mail: [email protected]
For queries about this service, please contact Infrastructure at:
[email protected]
---------------------------------------------------------------------
To unsubscribe, e-mail: [email protected]
For additional commands, e-mail: [email protected]