nsivabalan commented on issue #19262: URL: https://github.com/apache/hudi/issues/19262#issuecomment-5750570015
Picking up the "table inspection/administration operations" half of this issue, here is a concrete plan for a set of **table health tools plus a skill** that let whoever operates a Hudi table find out whether it is progressing as expected. The intended usage is a scheduled run — daily is typical — whose output an operator (or an agent) acts on. ## Motivation Hudi 1.2.0 expanded what a table can *hold*, and the agentic lakehouse work here expands how you *talk to* it. Neither covers whether a table is actually **well**. That matters for this issue specifically: everything the gateway can see today goes through Trino SQL over the catalog. SQL cannot see timeline state, file-slice layout, metadata-table consistency, or record-level-index sizing. So an agent asked "is this table healthy?" or "what should I tune?" currently has nothing factual to stand on. These tools are meant to be that substrate — and they are equally useful standalone, from a cron job, with no gateway involved. ## What gets checked | Check | Question it answers | |---|---| | Compaction cadence | Are delta commits accumulating well past the configured trigger point? | | Cleaner progress | Is the cleaner keeping up with its retention policy? | | Archival cadence | Is the active timeline staying bounded? | | Savepoint blocking archival | Is a forgotten savepoint pinning the timeline? | | MDT sync + compaction lag | Is the metadata table caught up, and is its compaction lagging? | | MDT vs filesystem | Does metadata agree with what is actually on storage? | | Small files / micro-partitioning | Is the layout pathological? | | RLI sizing | Is the record-level index sized for the table's record count? | ## Three tools, no overlap | Tool | Question | |---|---| | `HoodieTableLayoutAnalyzer` (new) | Is the data laid out well? | | `HoodieTableHealthChecker` (new) | Are table services keeping up? | | `HoodieMetadataTableValidator` (exists) | Is the metadata correct? | The health checker wraps the existing metadata validator rather than reimplementing it. ## Design decisions worth flagging early **Writer properties are required, and there is no silent fallback to defaults.** Almost every question here is relative to how the table is written: whether fifty delta commits since the last compaction is fine depends on the configured trigger; whether a thousand instants in the active timeline is fine depends on the configured retention. Hudi has a default for each, so a tool *could* always produce an answer — but for a table written with non-default settings that answer is wrong in the worst direction, reporting healthy tables as unhealthy until operators stop trusting the tool. So properties come in via `--props` / `--hoodie-conf`, and a check whose governing property was not supplied reports `SKIPPED` naming that property, rather than guessing. An explicit `--apply-all-defaults` opts into evaluating against Hudi defaults. Every check echoes the effective configuration it used, so a verdict can always be read alongside the settings that produced it. **Read-only, and cheap.** Nothing schedules or runs a table service. The timeline-only checks need no Spark session at all. **Exit codes carry the verdict** (`0` healthy, `1` unhealthy, `2` error) so a scheduled run is usable without parsing output. `--output JSON` for machine consumption, human-readable table by default. **No new `hoodie.*` configs.** Thresholds ship as constants with documented rationale. They are deliberately generous — the goal is catching a service that has stopped or fallen badly behind, not policing normal scheduling jitter. ## Proposed PR sequence | PR | Scope | |---|---| | 1 | `HoodieTableHealthChecker` framework + compaction, savepoint, archival checks (timeline-only, no Spark) | | 2 | Cleaner check | | 3 | MDT sync/compaction lag + wrapper over the existing metadata validator | | 4 | RLI sizing, with a rebootstrap recommendation when undersized | | 5 | `HoodieTableLayoutAnalyzer` — small-file, micro-partition and hot-partition detectors, skew metrics, JSON output | | 6 | `hudi-table-health` skill under `hudi-agent-gateway/skills/`, mirroring the `hudi-architect` layout | PR 1 carries the framework so each later check is a self-contained addition plus tests; every PR stays small enough to review properly. PR 5 is independent of 1–4 and can land in any order. PR 6 is the piece that closes the loop for this issue: a skill that reads the JSON report, explains what each failure means, and says what to run next — the same structure as the `hudi-architect` skill (#19380), usable both inside the gateway and as a standalone Claude Code skill. ## Relationship to #19263 This also lays groundwork for #19263 (specialized sub-agents for optimization and analysis). The optimization sub-agent described there — layout, compaction/clustering tuning, eventually closed-loop — needs to know what is actually wrong with a table before it can recommend anything, and today there is no way to find that out other than guessing. These tools are that sensing layer. The layout analyzer's detectors and skew metrics are the inputs a layout-tuning agent reasons over; the service checks tell it whether compaction and clustering are even keeping up; the RLI check already produces a concrete recommendation (rebootstrap when the index is undersized for the table's record count), which is the simplest possible instance of the optimization loop that issue describes. Worth deciding between the two issues where the boundary sits. My take: **detection and recommendation belong here as plain tools** (deterministic, testable, useful without an LLM in the loop), and **#19263 owns the agent that decides what to do about it and eventually acts.** Keeping the sensing deterministic also means the sub-agent's suggestions can be traced back to specific measured findings rather than model judgment. ## Status PR 1 is implemented and building against master. I will open it shortly; the rest follow in sequence. Feedback welcome on any of this — particularly the fail-fast stance on writer properties, and whether the check list above misses anything operators regularly get burned by. -- This is an automated message from the Apache Git Service. To respond to the message, please log on to GitHub and use the URL above to go to the specific comment. To unsubscribe, e-mail: [email protected] For queries about this service, please contact Infrastructure at: [email protected]
