nsivabalan commented on issue #19262:
URL: https://github.com/apache/hudi/issues/19262#issuecomment-5750570015

   Picking up the "table inspection/administration operations" half of this 
issue, here is a concrete plan for a set of **table health tools plus a skill** 
that let whoever operates a Hudi table find out whether it is progressing as 
expected. The intended usage is a scheduled run — daily is typical — whose 
output an operator (or an agent) acts on.
   
   ## Motivation
   
   Hudi 1.2.0 expanded what a table can *hold*, and the agentic lakehouse work 
here expands how you *talk to* it. Neither covers whether a table is actually 
**well**.
   
   That matters for this issue specifically: everything the gateway can see 
today goes through Trino SQL over the catalog. SQL cannot see timeline state, 
file-slice layout, metadata-table consistency, or record-level-index sizing. So 
an agent asked "is this table healthy?" or "what should I tune?" currently has 
nothing factual to stand on. These tools are meant to be that substrate — and 
they are equally useful standalone, from a cron job, with no gateway involved.
   
   ## What gets checked
   
   | Check | Question it answers |
   |---|---|
   | Compaction cadence | Are delta commits accumulating well past the 
configured trigger point? |
   | Cleaner progress | Is the cleaner keeping up with its retention policy? |
   | Archival cadence | Is the active timeline staying bounded? |
   | Savepoint blocking archival | Is a forgotten savepoint pinning the 
timeline? |
   | MDT sync + compaction lag | Is the metadata table caught up, and is its 
compaction lagging? |
   | MDT vs filesystem | Does metadata agree with what is actually on storage? |
   | Small files / micro-partitioning | Is the layout pathological? |
   | RLI sizing | Is the record-level index sized for the table's record count? 
|
   
   ## Three tools, no overlap
   
   | Tool | Question |
   |---|---|
   | `HoodieTableLayoutAnalyzer` (new) | Is the data laid out well? |
   | `HoodieTableHealthChecker` (new) | Are table services keeping up? |
   | `HoodieMetadataTableValidator` (exists) | Is the metadata correct? |
   
   The health checker wraps the existing metadata validator rather than 
reimplementing it.
   
   ## Design decisions worth flagging early
   
   **Writer properties are required, and there is no silent fallback to 
defaults.** Almost every question here is relative to how the table is written: 
whether fifty delta commits since the last compaction is fine depends on the 
configured trigger; whether a thousand instants in the active timeline is fine 
depends on the configured retention. Hudi has a default for each, so a tool 
*could* always produce an answer — but for a table written with non-default 
settings that answer is wrong in the worst direction, reporting healthy tables 
as unhealthy until operators stop trusting the tool.
   
   So properties come in via `--props` / `--hoodie-conf`, and a check whose 
governing property was not supplied reports `SKIPPED` naming that property, 
rather than guessing. An explicit `--apply-all-defaults` opts into evaluating 
against Hudi defaults. Every check echoes the effective configuration it used, 
so a verdict can always be read alongside the settings that produced it.
   
   **Read-only, and cheap.** Nothing schedules or runs a table service. The 
timeline-only checks need no Spark session at all.
   
   **Exit codes carry the verdict** (`0` healthy, `1` unhealthy, `2` error) so 
a scheduled run is usable without parsing output. `--output JSON` for machine 
consumption, human-readable table by default.
   
   **No new `hoodie.*` configs.** Thresholds ship as constants with documented 
rationale. They are deliberately generous — the goal is catching a service that 
has stopped or fallen badly behind, not policing normal scheduling jitter.
   
   ## Proposed PR sequence
   
   | PR | Scope |
   |---|---|
   | 1 | `HoodieTableHealthChecker` framework + compaction, savepoint, archival 
checks (timeline-only, no Spark) |
   | 2 | Cleaner check |
   | 3 | MDT sync/compaction lag + wrapper over the existing metadata validator 
|
   | 4 | RLI sizing, with a rebootstrap recommendation when undersized |
   | 5 | `HoodieTableLayoutAnalyzer` — small-file, micro-partition and 
hot-partition detectors, skew metrics, JSON output |
   | 6 | `hudi-table-health` skill under `hudi-agent-gateway/skills/`, 
mirroring the `hudi-architect` layout |
   
   PR 1 carries the framework so each later check is a self-contained addition 
plus tests; every PR stays small enough to review properly. PR 5 is independent 
of 1–4 and can land in any order.
   
   PR 6 is the piece that closes the loop for this issue: a skill that reads 
the JSON report, explains what each failure means, and says what to run next — 
the same structure as the `hudi-architect` skill (#19380), usable both inside 
the gateway and as a standalone Claude Code skill.
   
   ## Relationship to #19263
   
   This also lays groundwork for #19263 (specialized sub-agents for 
optimization and analysis). The optimization sub-agent described there — 
layout, compaction/clustering tuning, eventually closed-loop — needs to know 
what is actually wrong with a table before it can recommend anything, and today 
there is no way to find that out other than guessing.
   
   These tools are that sensing layer. The layout analyzer's detectors and skew 
metrics are the inputs a layout-tuning agent reasons over; the service checks 
tell it whether compaction and clustering are even keeping up; the RLI check 
already produces a concrete recommendation (rebootstrap when the index is 
undersized for the table's record count), which is the simplest possible 
instance of the optimization loop that issue describes.
   
   Worth deciding between the two issues where the boundary sits. My take: 
**detection and recommendation belong here as plain tools** (deterministic, 
testable, useful without an LLM in the loop), and **#19263 owns the agent that 
decides what to do about it and eventually acts.** Keeping the sensing 
deterministic also means the sub-agent's suggestions can be traced back to 
specific measured findings rather than model judgment.
   
   ## Status
   
   PR 1 is implemented and building against master. I will open it shortly; the 
rest follow in sequence.
   
   Feedback welcome on any of this — particularly the fail-fast stance on 
writer properties, and whether the check list above misses anything operators 
regularly get burned by.
   


-- 
This is an automated message from the Apache Git Service.
To respond to the message, please log on to GitHub and use the
URL above to go to the specific comment.

To unsubscribe, e-mail: [email protected]

For queries about this service, please contact Infrastructure at:
[email protected]

Reply via email to